LiamVi

← All articles  ·  16 September 2026

Two upright bars standing on one shared baseline. The left bar, marked None, is tall and blue. The right bar, marked Some, is solid orange and reaches a little under one tenth of the left bar's height.

SERP overlap in keyword clustering: how often two keywords actually return the same pages

Across the 37 search-results records we re-pulled on 9 September 2026, 607 of the 666 pairs those keywords can form — every possible pairing of the 37 — share no URL at all, and 12 of the 666 reach the three shared results that four of the six guides we opened treat as the point where two keywords become one page. The records behind that count are not published yet, so what follows is the method rather than the dataset: where the records come from, what the guides recommend, and a check on two of your own keywords that needs nothing of ours.

The clearest single case is this article's own keyword: `serp overlap keyword clustering` and `keyword clustering` returned nine results each, and not one page appeared on both. We captured them on the same day, from the same location, and both phrases contain the words keyword clustering. A tool set to merge at three shared URLs would put them in separate clusters, with nothing to argue about.

This is a count of how often a condition is met. It is not a verdict on whether the rule is a good rule, and it is not a measurement of what clustering does to rankings, traffic or anything else. We have not tested them, and neither has any of the six pages we opened for this query.

What are we counting when we count SERP overlap?

The overlap in this article is the number of URLs that two keywords' Google results pages have in common. You get it by searching one keyword, searching the other, and counting the pages that appear on both lists. That is the number a SERP-based clustering tool is built on.

Two horizontal runs of nine identical navy chips, one above the other, the same length and starting from the same left margin. Below them, one empty orange slot the size of a single chip, marked Shared above it and Not one beneath it.
This article's own keyword and the phrase keyword clustering returned nine results each, captured on the same day from the same location, and not one page appeared on both lists. A result of zero is the ordinary answer rather than a failed search: it is what 607 of the 666 pairs in our 37 records did.

It is not the other overlap we publish, and the two have nothing to do with each other. Elsewhere on this site we report how far Google's results and ChatGPT's citations diverge — a comparison between a search engine and an assistant, with its own sample on our research page. That measurement says nothing about keyword clustering, and this one says nothing about AI citations. We are naming the collision once because we caused it: we used one word for two objects, and a reader coming from those articles would be right to check which is which.

The rule being tested here is short enough to state in a sentence. SERP-based clustering assumes that if two keywords return substantially the same pages, Google is treating them as one question, so you should answer them on one page. Set a threshold — three shared URLs, four, thirty per cent — and any pair above it gets merged. How to do that clustering well is covered at length by the guides ranking for this query, and we are not going to repeat them. The question none of the six we opened answers is how often the threshold is actually reached.

What threshold do the clustering guides recommend?

The six pages we could open for this query state the overlap threshold six different ways, and the values they land on run from three shared URLs to thirty per cent. Not one of the six publishes a sample, a test or any figure for how often its condition is met.

Six identical plates standing in one row, each carrying one threshold: 3 URLs, 3 to 4, 1 to 9, 1 to 10, 1 to 10, and 30 percent. Below them, a seventh plate of the same size, filled solid orange and completely empty, marked No test.
The six pages we opened for this query state the overlap threshold six different ways — three shared URLs; three to four in the top ten; a one-to-nine accuracy scale; a one-to-ten sensitivity setting; a one-to-ten slider with no value recommended; and thirty per cent — and not one of them publishes a sample or a test behind its number. All six read on 10 September 2026.

Nadia Mohamed's piece on keyword clustering and SERP overlap, published 8 July 2026, is the page most exactly on this question, and it gives the most careful version of the advice: "A threshold of 3 shared URLs is a sensible starting point for most sites; raise it when your clusters feel too loose." It also carries the most honest line of the six — "It is a dial, not a default."

Nightwatch's guide, by Aljaž Fajmut, states it as established practice: "A common threshold is 3–4 overlapping URLs in the top 10." The Nightwatch page shows no publication or update date at all, which we read the page for rather than assumed.

SE Ranking's guide, updated 25 November 2024 by Yulia Deda, ships it as a setting: "Keyword grouping accuracy (ranges from 1 to 9) indicates the minimum number of matching URLs for the queries," with the worked example "if you set it to 3, phrases will be grouped together only if they have 3 identical URLs in the search results."

KeyClusters' tool comparison, dated 6 April 2026, states the rule as a mechanism: "If two keywords return three or more of the same pages in the top ten, the tool treats them as belonging to the same cluster." It also offers a sensitivity setting from one to ten.

ContentGecko's free clustering page defines the parameter without committing to a value: "Minimum SERP overlap is the threshold that determines how many common URLs two keywords must share in their top 10 search results to be grouped together." The setting is a slider from one to ten.

Keywordly's clustering feature page states the only proportional version we found on the set: a "30% URL overlap threshold".

Two of the six also put that number at different points in a workflow, which changes what any count of overlap is a count of. Nightwatch's page puts the results page second: "start with intent and semantic clustering to build initial groups, then validate each cluster against the actual SERPs." KeyClusters' page puts it first — you "upload a CSV of keywords", set a sensitivity level, and "the tool fetches live Google results and groups keywords that share ranking URLs at or above your sensitivity threshold." How often the threshold fires depends on which of those two lists went in, and ours went in ungrouped: all 666 pairs that 37 keywords can form, rather than candidate groups somebody had already built by intent.

The rule arrives as three shared URLs, or three to four in the top ten, or a dial from one to nine, or a setting from one to ten, or thirty per cent, depending on which of the six pages you are reading. That spread is not itself a criticism, because reasonable people set a parameter differently. What it means for the reader is narrower and harder to shrug off: you are being handed a number, and none of the six pages handing it to you gives you a way to tell whether it is the right one.

Here is what we read each page for, and did not find. On all six we looked for a dataset, a stated sample size, a description of method, a comparison group testing one threshold against another, and a link to a validation study. On all six, all five were absent. What is there instead is external citation about something adjacent. Nadia Mohamed's page cites, according to that page, an Ahrefs study of three million searches about how many keywords a top-ranking page ranks for, and a Semrush example of a single page ranking for 2,200 keywords; neither of those figures tests a clustering threshold, and we did not open either source. SE Ranking's page shows a worked illustration using 200 exercise-related keywords at two different accuracy settings, which demonstrates what the setting does rather than whether the setting is right. Not one of the six carries a figure for how often two keywords' results actually overlap.

We should say plainly that these are careful pages. Three of the six run from roughly 3,400 to 4,700 words in our own capture of that results page, and they are detailed, specific practitioner work; the one most exactly on this question is also the one that hedges its own number honestly. The criticism here is of an unreported validation, not of a company or a writer.

One more scope note about that results page. Nine results were captured for this keyword and we could read six of them. A Reddit thread at the top of the set returned nothing to our fetcher, a LinkedIn post sits behind a sign-in wall, and a video page carries no readable text. We make no claim about what any of those three contains.

How often do two keywords actually return the same pages?

Two keywords in the same commercial category share no search result at all in 607 of the 666 pairs we compared, and only 12 of the 666 reach the three shared URLs the guides above treat as the merging point. The 666 is every pair 37 keyword records can form, not a shortlist of candidates somebody had already grouped; the records were captured on five dates, and the frame matters enough that it is spelled out below rather than left to a footnote.

Three solid navy bars of different heights standing on one shared baseline, marked Our records. An orange horizontal rule crosses the open ground above all three, marked Top ten. No bar reaches it: the empty band above each bar is a different size, widest above the shortest bar and narrowest above the tallest.
Not one of our 37 records holds a full top ten: 2 hold seven results, 18 hold eight and 17 hold nine, read back on 9 September 2026. The figure draws the shortfall rather than the counts — a nine-result record leaves one top-ten position unseen, an eight-result record two, a seven-result record three — so our count can only be lower than a top-ten rule would return, never higher.

Where the records come from, in enough detail to repeat. Before we write an article we capture the Google results page for its keyword and store it: the rank and the URL of each result, stamped with the date it was captured. On 9 September 2026 we read 37 of those stored records back, fetching no new results pages, so every capture date in this article is the date on the record. One further keyword we could name from our own record — `how to get citations from chatgpt` — had no stored record when we looked, so it is dropped from the count and named here rather than quietly absorbed.

The mechanical layer first, because none of it involves a judgement:

| Shared URLs between a pair | Pairs | |---|---| | none | 607 of 666 | | exactly one | 33 of 666 | | exactly two | 14 of 666 | | three or more | 12 of 666 | | four or more | 9 of 666 | | five or more | 3 of 666 |

Most pairs share nothing, and the fall-off after that is steep: 33 of the 666 pairs share exactly one URL, 14 share exactly two, and 12 share three or more. Fifty-nine of the 666 pairs have even one page in common, and only 12 of those 59 reach the threshold. The highest overlap anywhere in the set is 7, between `content optimization tool` and `seo content optimization tool` — two phrases that differ by one word.

Three limits belong right here, next to that table, because each of them changes what the numbers mean.

Not one of our 37 records holds a full top ten: 2 of them hold seven results, 18 hold eight and 17 hold nine. Four of the six pages above set their threshold explicitly in the top 10. A shallower set can only find fewer shared URLs than a deeper one, never more, so our count and a top-ten threshold are not the same measurement and ours is the more conservative of the two. The gap has a size we can name: 14 of the 666 pairs sit at exactly two shared URLs, one short of the line, and any of them could reach three in a place we never captured — a nine-result record leaves one top-ten position unseen, an eight-result record two, a seven-result record three. Our 12 of 666 is the count in the sets we captured, not the count the top-ten rule itself would return.

The captures fall on five dates, between 26 August and 9 September 2026, so a pair captured on different days is not a same-moment comparison — the results page could have moved in between. Taking only pairs captured on the same day leaves 140 pairs, of which 120 share nothing and 6 reach three or more.

That same-date subset is denser than the whole set — 6 of 140 against 12 of 666, more than twice the rate — and it cuts against our own headline. Keywords captured on the same day came out of the same planning batch, which means they are more closely related to each other by construction than two keywords picked at random would be. If anything, the same-date figure is the one that flatters the threshold, and we are reporting it because it does. That is the reading most favourable to the rule.

The frame, stated once and in full: these are captures made for one workspace, in one commercial category — AI search visibility and SEO tooling — at the location `United States`, one capture per results page, on five dates in August and September 2026. The count covers 37 records we can name from our own written record, out of the 41 our change log says we have built. The denominator is therefore what we could enumerate from that record, not everything we have built and certainly not keywords in general. Somebody clustering keywords in a different market should expect different results and should check their own.

Read the count as ours rather than as a neutral standard. We build a tool that reads ranking pages for a living, these records exist because we were buying something else with them, and we chose which keywords to plan articles for. The URLs in them are Google's and anybody can check a pair in five minutes; the selection of pairs is ours.

What does that rule say about our own pages?

The three-shared-URL rule, applied to our own catalogue, flags three pairs of our pages out of the 136 pairs it can form. Seventeen of the 37 captured keywords became LiamVi articles; 115 of their 136 pairs share nothing at all, and three sit at or above the threshold the pages above recommend:

| Our two keywords | Shared URLs | Captured | |---|---|---| | `best ai visibility analytics for search optimization` and `best ai visibility platforms with seo capabilities` | 5 | both 2026-09-01, same day | | `best ai search optimization platform for beginners` and `top user friendly ai search optimization tools` | 4 | both 2026-08-26, same day | | `best ai visibility platforms with seo capabilities` and `best ai visibility tool` | 3 | 2026-09-01 and 2026-09-03 |

We wrote each of those as two pages. Two of the three pairs were captured on the same day, so the date caveat above does not soften them, and the strongest of the three shares 5 of the 8 results each keyword returned.

We are not saying that writing those three pairs separately cost us anything. We have measured no consequence — not in rankings, not in traffic, not in citations — and we are not going to imply one we cannot show. What happened is narrower and more useful: we made each of those calls by reading the two results pages and judging that they answered different readers, and the rule, applied afterwards, would have merged all three pairs into one page each. Both of those are decisions. Only one of them is a number.

The other half of the same count points the other way. Six captured keywords never got a page of their own, and every one of them shares three or more URLs with an article we did write: `seo content optimization tool` shares 7 with our content optimization piece, `top rated ai visibility optimization software` shares 5 and 4 with two of ours, `people also ask seo` shares 4, `ai search engine optimization tools` shares 4, and two more sit at 3. A piece that printed only the three flags and left out those six would be performing contrition rather than reporting. The three-shared-URL rule disagreed with our own catalogue on 3 pairs and pointed the same way as it on 6 keywords we never wrote, which is a better record for the rule than the flagged count on its own suggests.

Whether those three pairs should have been one page each is a judgement about our own catalogue — what each page is for, who arrives at it, what it would lose by merging — and a threshold cannot make it for us. It can tell us where to look, which is a smaller and more honest job than making the decision for us.

How do you run this count on your own keywords?

Counting the shared URLs between two of your own keywords takes five steps, two browser tabs, no tool and no data of ours. The five steps below are our procedure rather than a standard, and they are shaped by what we look at all day.

  • Search the first keyword and write down the URLs of the results, in order. Take the first ten: the thresholds the guides publish are set against the top ten, and a shorter list can only find fewer shared URLs than a full one.
  • Search the second keyword in the same way, from the same place, on the same day.
  • Count the URLs that appear on both lists. That number is the SERP overlap, and it is the number every clustering threshold above is set against.
  • Compare it to whatever threshold your tool or your guide is using, and notice whether the tool tells you what its threshold is at all.
  • Write down the date. A results page is a photograph of one moment, and a pair that overlaps today may not next month.

A result of zero is the ordinary answer, not a mistake or a sign you searched badly: it is what 607 of the 666 pairs in our records did. If you get three or more you have found the uncommon case the rule was written for, and that is when it is worth reading both results pages properly to see whether Google really is answering one question twice.

The three questions we published for checking any claim about AI search work here too, pointed at a threshold instead of at a citation claim: what was measured, on how many of what; what came back null and did they publish it; and is this about the surface you care about. The longer version is in our write-up of generative engine optimization strategies.

Tools that do this clustering in bulk exist, and choosing one is outside what this article covers.

What this does not tell you

Rankings, traffic and citations are the three consequences of clustering that we have not measured at all. Whether merging two keywords at three shared URLs, or four, or thirty per cent, produces more of any of them than leaving the pages separate is untested by us, and we found no test of it on any of the six pages we read. We also have not measured what our own three flagged pairs cost us, if anything.

This is not a claim that the threshold is wrong. We counted how often its condition is met. Whether three shared results is a sensible line for deciding one page or two is a different question, and answering it would need an experiment we have not found published: two comparable sites, one clustering and one not, measured over a period long enough to matter.

Our statement about missing validation is scoped, and the scope is small. It covers the six pages we opened for this query, plus one web search that returned practical clustering guides rather than research with a stated sample. A search that finds nothing is not proof that nothing exists, and if somebody has published a real test of a clustering threshold we would like to read it.

The record itself is small and early: 37 keyword captures on five dates in a fortnight, under the frame set out above. We keep building these records every week. The set grows, we will re-count it, and when the answer changes we will publish the change.

Two limits belong to us rather than to the data. Our own scoring instrument has a published failure: tested against blind quality reviews run by independent reviewer agents, the drafts our meter rated highest were the ones the reviews rated lowest, at rho −0.61 across nine articles. That write-up is in our piece on content optimization tools. It belongs here because it is the standard we are asking you to hold this count to: one instrument of ours has already pointed the wrong way in public, and this one has been checked by nobody outside this desk. And the receipt for this count is not published yet: our research page, last updated 12 August 2026, carries our other samples and limits but does not carry this one. It should, and until it does the numbers above are the record.

There is a sibling count from the same records, taken from a different column: what the question box beside those same results pages returns, and whether those questions belong to the keyword they appear under. That is a separate article with its own figures, and this one does not repeat them.

What people actually ask us

The count changes how often you should expect this signal to fire, not whether you should use it — and that distinction is behind all four of the questions below. Each answer stands on its own.

Should I stop using SERP overlap to cluster keywords? No. It is a real signal and it is free to check. What our count adds is how rarely it fires: in 607 of the 666 keyword pairs we compared, the two results sets had no URL in common at all. How often it fires for you depends on what you put into it — a list already grouped by intent will overlap more often than every pair of an ungrouped one — and ours was ungrouped, so a workflow assuming overlap is common spent most of its time returning nothing. Use it as a flag for the uncommon case rather than as the engine of your site structure.

What number should I set the threshold to? We cannot tell you, and neither, on the evidence we could read, can the pages recommending one. The six pages we opened for this query stated the rule six different ways — three shared URLs; three to four in the top ten; a one-to-nine accuracy scale; three or more on a one-to-ten sensitivity setting; a one-to-ten slider with no value recommended; and a thirty per cent overlap threshold — and not one of them published a sample or a test behind its choice. The honest move is to pick a value, look at what it groups on your own keywords, and treat it the way Nadia Mohamed's page does: "a dial, not a default."

Does a zero mean I should write two pages? It means two captured results sets, taken on one day from one location, had no URL in common. That is information rather than an instruction, and it is not a reading of search intent: Google's own account of how it ranks results names "Context and settings""your location, past Search history, and Search settings" — alongside the meaning of the query, so two results pages can diverge for reasons that have nothing to do with who is asking. The decision still depends on things a URL count cannot see — whether you have enough to say for two pages, whether the two readers want different things, and what you would lose by merging. We have three pairs of our own where the rule says one thing and our reading said another, and we have not measured which was right.

Where does your data come from, and can I check it? It comes from the results pages we capture and store when we plan an article, each one holding the ranked URLs for a keyword with the date it was captured; 37 of them, read back on 9 September 2026, in one commercial category at the location `United States`. What you can reproduce today is the method: the underlying object is Google's own results for two keywords, which anybody can open and compare in five minutes on their own pair. What you cannot check is our 666-pair count itself, because the records behind it are not published yet. That is why we published the method rather than a single headline number.

Sources