Surfer SEO Review (2026): All 20 Fixture Keywords — Its Content Score Did Not Track Google Rank
All 20 fixture keywords on a seven-day Pro trial, tested 2026-08-24, re-run on 5 of them 2026-08-25.
Surfer is the first tool whose trial quota covered our whole fixture, so nothing here is a subset. Two things still need saying before the numbers. Its competitor table carries two score-like columns, a 0–100 content score per page and a 0–10 domain score; they measure different things, and this page reports each under its own name, never merged. And its term panel has no numeric importance, so the top 30 we score is the top 30 of Surfer’s own display order, where the other tools’ lists are ranked by their own importance values. Same rule — the tool’s own first 30 — different ordering source.
Verdict
partialSurfer works one keyword at a time. A Content Editor document crawls the SERP, lists ten competitors with position, word count and two scores, extracts the terms those pages use, and grades your draft against them as you write. The table is thorough. The score on top of it told us nothing about rank, and the term list under it changed overnight.
We found the pages it had scored inside the real Google (us) top 10 for each keyword and correlated its content score with actual position. On the URLs that matched exactly, the mean Spearman correlation was . Our ground truth holds some URLs that Ahrefs returned truncated; count those as matches and it is over 107 URLs. Against the rank Surfer crawled itself, over every scored competitor row, . Three framings, all within a tenth of a point of zero.
What sets this run apart from the other three is the day after. We paid five more documents to re-run five keywords ~26 hours later. The scores barely moved. The recommended terms did: on three of the five keywords, the day-two list overlapped the day-one list below our stability threshold, once while the competitor set stayed byte-for-byte identical. A score that repeats but does not predict, sitting on a term list that does not repeat.
Upstream, the crawl holds up. Its top 10 matched Google’s real top 10 on of domains. The first 30 terms of its panel, scored against the terms the ranking pages share, reached precision@30 and recall@30 .
you want one document to hand you the competitor set, each page’s word count and a term checklist, and you will judge coverage yourself — treating the term list as that day’s snapshot.
you plan to write toward the content score, or to treat the term list as a specification. The score sat on zero against Google, and a quarter of the terms can be different tomorrow.
Tested 2026-08-24 on the Pro trial, across all 20 keywords in our frozen fixture; consistency re-run 2026-08-25. Every number on this page is read out of the runs’ machine-readable reports or the verification sidecar beside them. See how it compares on our comparison table.
Key findings
- Surfer's content score did not track Google rank. Against the real Google (us) position of the 96 URLs that matched our ground truth exactly, the mean Spearman was −0.102 over 18 keywords; rescue the truth URLs Ahrefs returned truncated and it is −0.044 over 107 URLs. Same sign, still near zero. Against the rank Surfer crawled itself, −0.084 over 187 rows (tested 2026-08-24).
- The same table carries a second number, a 0–10 domain score, and it sat further from zero than the content score: +0.185 against real Google position over 101 URLs, +0.165 against Surfer's own ranking. 20 keywords, no significance test. We publish it as data, not as a conclusion.
- Its term list is not the same list the next day. A brand-new document for the same 5 keywords, 25.61–25.81 hours later, returned term sets whose overlap fell below our 0.7 stability threshold on 3 of the 5 (mean Jaccard 0.746, worst 0.524).
- The scores held still while the terms moved. Across 44 URLs present in both runs, the content score's mean absolute change was 1.43 points, 52% of the URLs did not move at all, and the domain score did not change on any of its 46 URLs.
- This is the first tool whose trial covered our whole fixture: all 20 keywords in one round, 20 documents, plus 5 more for the re-run. Its crawled top 10 matched Google's on 84.6% of domains.
What works
- One document per keyword returned everything at once: 79–80 recommended terms, a 10-row competitor table with a 0–100 score, a 0–10 domain score and a word count per URL.
- Its crawled top 10 matched the real Google (us) top 10 on 84.6% of domains. The competitor set it shows you is close to the one you are up against.
- Its scores repeat. On the day-two re-run, 52% of re-scored URLs kept the exact same content score and every domain score held. When a number moves, you probably moved it.
- The trial is wide enough for a real test: all 20 fixture keywords plus a 5-keyword consistency re-run, for $0.
What doesn’t
- The content score did not track rank: −0.102 against Google on a strict URL match, −0.044 on a lenient one, −0.084 against its own ranking. Near zero three ways.
- The first 30 terms of its panel reached precision@30 0.161 and recall@30 0.333 against the terms the ranking pages share — about one term in six is one those pages use, on its own keyword set.
- The term list moves overnight. Three of the five re-run keywords came back with a term set below our stability threshold; on one, 20 of 80 recommended terms fell out while the competitor set stayed identical.
- It exposes no numeric term importance. The panel's display order is the only ranking, so you cannot tell its tenth term from its thirtieth except by position.
- 13 of its 200 competitor rows carried no content score because Surfer's own fetch of the page failed, and its API sits above the trial tier, so every capture is UI work.
What Surfer is
The app lives at app.surferseo.com. The unit of work is a Content Editor document: one keyword, one country, one analysis. Each document returns a competitor table — ten URLs with Surfer’s crawled position, a 0–100 content score, a 0–10 domain score and a word count — a panel of recommended terms with suggested usage ranges, and an editor that grades your draft as you type. The trial let us create documents for the fixture and 5 more for the re-run, 25 of the plan’s nominal 30 a month.
We ran on the Pro trial, which ended 2026-08-31, and spent $0. The API sits on a higher tier than the trial, so the whole test went through the UI. Pricing was not recorded this round; see what we couldn’t test.
Does the content score predict rank?
negativeThe content score is the number the product asks you to raise, so we tested it against the thing you actually want. For each keyword we took the real Google (us) top 10 out of our ground truth, found those URLs in Surfer’s competitor table, and correlated the score it gave each page with that page’s real position. Positive means higher-scoring pages ranked higher. Near zero means the score carries no rank information.
We report that correlation twice, for the same reason as on the other reviews. Our ground truth came from Ahrefs, and some of its URLs came back cut short. A strict path match misses those rows; a lenient rule, which counts a truth URL as matched when it is the unique prefix of exactly one URL Surfer captured, rescues of them here, at the risk of pulling in a same-domain page that is not the one that ranked. On this run the two rules agree on the sign; the review still shows both, because neither is the corrected value.
The third framing is the tool grading its own homework: the same score against the position Surfer crawled itself, over all 187 rows that carried one. Bigger sample, easier test, same answer.
Part of: does a content score predict Google rank? Same measurement, every tool we have run it on.
| Keyword | URLs matched | ρ vs. Google, strict | Caveat |
|---|---|---|---|
| cost of living comparison | 3/9 | 0.500 | small sample, n = 3 |
| gpu comparison | 8/10 | -0.643 | |
| ai voice generator | 8/9 | -0.635 | |
| ai content detector | 8/8 | 0.192 | |
| best ai chatbot | 4/6 | 0.800 | small sample, n = 4 |
| ai presentation maker | 3/9 | -0.500 | small sample, n = 3 |
| ai for small business | 6/8 | -0.657 | |
| ai for coding | 1/6 | — | only 1 matched URL |
| ai for teachers | 6/8 | -0.290 | |
| best crm for small business | 0/7 | — | no URL overlap |
| best ai for writing | 3/5 | 0.500 | small sample, n = 3 |
| best ai for math | 4/5 | -0.800 | small sample, n = 4 |
| ai for customer service | 6/7 | -0.203 | |
| ai for business automation | 4/7 | -0.400 | small sample, n = 4 |
| best ai seo tools | 5/6 | 0.700 | |
| ai for marketing | 6/8 | -0.319 | |
| best ai writing tools | 4/6 | -0.400 | small sample, n = 4 |
| best laptop for programming | 6/7 | -0.232 | |
| best help desk software | 4/7 | 0.400 | small sample, n = 4 |
| digital nomad visa countries | 7/8 | 0.144 | |
| Mean over 18 keywords with a ρ | 96/146 | -0.102 |
← swipe the table sideways for the rest of the columns
best crm for small business produced no correlation at all: none of its 7 ground-truth URLs matched a row in Surfer’s table at the same path. Under the lenient match it gains a row, which is why that framing counts 19 keywords to the strict 18.
10 of the 20 keywords rest on four matched pages or fewer, and the report marks them so nobody quotes them alone. Drop every flagged keyword and the finding survives: 10 keywords, 66 pages, mean ρ −0.194. The per-keyword values run from -0.800 on best ai for math to 0.800 on best ai chatbot: 7 keywords positive, 11 negative. A mean that close to zero, built from swings that wide, is noise around zero, not a small effect.
| Framing | Keywords | Rows in those keywords | Mean ρ |
|---|---|---|---|
| Score vs. real Google (us) position, strict URL match | 18 | 96 | |
| Score vs. real Google (us) position, truncated truth URLs rescued | 19 | 107 | |
| Score vs. the rank Surfer crawled itself | 20 | 187 |
Unlike Scalenut’s run, the sign holds here: strict and lenient both land below zero, a move of 0.058 between them. That does not rescue the score. On the findings page we ask how far each figure sits from zero, and none of these three gets past ~0.1.
The day-two re-run caught it directly. On ai content detector, four of the ten competitor URLs changed position overnight — quillbot.com moved 2 → 1, gptzero.me 1 → 2 — and all ten content scores came back identical to the digit. The score does not move when the ranking moves, so it is computed from the page, not from the SERP. That cuts both ways: it cannot be accused of copying rank, and it visibly does not contain it.
We ran this measurement on Frase (-0.115 over 38 pages), NeuronWriter (-0.052 over 110 pages) and Scalenut (-0.071 over 32 pages) before this one. Surfer came in at −0.102. Different keyword sets and sample sizes, so no two of those numbers compete. What they share is where they land.
The other score in the table
Surfer’s competitor table carries a second number none of the other tools showed us: a 0–10 domain score per URL. It is a domain-level metric, not a judgment of the page’s text, so it is not the content score under another name and we did not fold it into anything. We did run the same correlation on it, because it was in the table.
| Score | vs. real Google (us), strict | vs. Surfer’s own rank |
|---|---|---|
| Content score (0–100, per page) | 96 pages, 18 keywords | 187 rows, 20 keywords |
| Domain score (0–10, per domain) | 101 pages, 18 keywords | 196 rows, 20 keywords |
The domain score sat further from zero than the content score, under both answer keys, and on the same side both times. We are not calling that a finding. It is 20 keywords, we ran no significance test, per-keyword values swing from −0.667 to +0.949, and +0.185 is still well short of anything we would call predictive. What the pair of rows is fair evidence for is narrower: on this sample, a number about the domain came closer to the ordering of a top 10 than a number about the page’s text. That is consistent with what every run on this site keeps showing — the things a content score cannot see are the things doing the ordering.
What we matched, and how
URLs are matched on the exact path, with the protocol, www, trailing slash and query string stripped. Same domain at a different path does not count. Of the 146 ground-truth web URLs for these keywords, appeared in Surfer’s tables at the same path. Not every matched row could be correlated: of the 200 competitor rows carried no content score, because Surfer’s own fetch of that page failed — expatistan and reddit pages mostly, the sites that block crawlers — which leaves 96 rows for the content-score correlation and 101 for the domain score. High domain overlap, of its top-10 domains matching Google’s, hides a lower URL overlap. The two are not the same claim.
The tables needed no cleaning: 0 duplicate rows, 0 rows without a position. Surfer’s SERP snapshot and our frozen truth were also not taken on the same day, so a URL that fails to match may be a crawl difference or a SERP that genuinely turned over; this report does not separate the two.
Do its terms match what ranks?
negativeSurfer publishes a real term panel, so this is the same measurement we ran on NeuronWriter, Clearscope and Scalenut, with one difference stated first. Those tools attach a numeric importance to each term, and their top 30 is the top 30 by that number. Surfer attaches none. Its only ranking is the order the panel displays, so our top 30 is its first 30 as shown. Both are “the tool’s own first 30”; the ordering source differs, and a comparison across tools inherits that difference.
We matched those 30 terms, on unigrams and bigrams, against the top 30 terms our ground truth extracted from the body text of the pages that rank. About one term in six is a term the ranking pages share, and the list covers a third of those shared terms. Across 20 keywords it found 200 shared terms in total.
| Measure | Value |
|---|---|
| Precision@30, mean over 20 keywords | |
| Recall@30, mean over 20 keywords | |
| Surfer SERP vs. real Google (us) top 10, domain overlap |
The spread across keywords is wide. Best case, best help desk software, recalled 0.500 of the shared terms. Worst case, best ai chatbot, recalled 0.067. Same tool, same settings, same day.
And the day matters more here than on any other review, because the consistency section below measured it: re-run a keyword a day later and the panel this test scored is not quite the panel you get. Read the P@30 and R@30 above as the hit rate of the 2026-08-24 list, dated like everything else on this page.
| Keyword | Terms offered | Terms taken | P@30 | R@30 |
|---|---|---|---|---|
| cost of living comparison | 80 | 30 | 0.151 | 0.367 |
| gpu comparison | 80 | 30 | 0.254 | 0.500 |
| ai voice generator | 79 | 30 | 0.211 | 0.400 |
| ai content detector | 80 | 30 | 0.273 | 0.400 |
| best ai chatbot | 79 | 30 | 0.030 | 0.067 |
| ai presentation maker | 79 | 30 | 0.182 | 0.400 |
| ai for small business | 79 | 30 | 0.123 | 0.300 |
| ai for coding | 80 | 30 | 0.133 | 0.267 |
| ai for teachers | 80 | 30 | 0.130 | 0.300 |
| best crm for small business | 80 | 30 | 0.148 | 0.300 |
| best ai for writing | 80 | 30 | 0.123 | 0.267 |
| best ai for math | 80 | 30 | 0.167 | 0.367 |
| ai for customer service | 80 | 30 | 0.141 | 0.300 |
| ai for business automation | 80 | 30 | 0.154 | 0.333 |
| best ai seo tools | 79 | 30 | 0.244 | 0.367 |
| ai for marketing | 80 | 30 | 0.077 | 0.167 |
| best ai writing tools | 80 | 30 | 0.121 | 0.267 |
| best laptop for programming | 80 | 30 | 0.154 | 0.400 |
| best help desk software | 80 | 30 | 0.238 | 0.500 |
| digital nomad visa countries | 79 | 30 | 0.160 | 0.400 |
| Mean, 20 keywords | — | — | 0.161 | 0.333 |
← swipe the table sideways for the rest of the columns
You can rebuild the ground truth yourself. Our Top 10 Term Finder runs the same extraction on any keyword you type: fetch the live top 10, pull the body text, keep the terms those pages share. Compare that against whatever your tool hands you.
Is it the same tool tomorrow?
splitHalf of it is. For 5 keywords we created a brand-new Content Editor document –25.81 hours after the first — new document, not a reopened one, because reopening only tests whether our extraction repeats, and each new document costs a trial credit. The five keywords are the same five the Frase and NeuronWriter consistency runs used, so the three measurements sit on the same words.
The scores held. Across 44 URLs present in both runs with a content score, the mean absolute change was points, the largest 9, and 23 of them (52%) did not change at all. The domain score did not move once: of 46 URLs, zero drift. Read a one-or-two-point move in either score as noise.
The term list did not hold. Overall Jaccard overlap between the two days’ recommended-term sets averaged with a worst case of 0.524, and 3 of the 5 keywords fell below the 0.7 threshold our methodology treats as unstable. In plainer units: on cost of living comparison, 20 of the first day’s 80 recommended terms were gone the next day.
| Keyword | Gap (h) | Term overlap | Competitor-set overlap | Score |Δ| mean / max | Scores unchanged |
|---|---|---|---|---|---|
| cost of living comparison | 25.76 | 0.600 | 1.000 | 2.00 / 5 | 1 of 8 |
| gpu comparison | 25.81 | 0.524 | 0.667 | 1.38 / 5 | 4 of 8 |
| ai voice generator | 25.67 | 0.629 | 0.667 | 3.88 / 9 | 1 of 8 |
| ai content detector | 25.71 | 1.000 | 1.000 | 0.00 / 0 | 10 of 10 |
| ai presentation maker | 25.61 | 0.975 | 1.000 | 0.50 / 2 | 7 of 10 |
| Mean, 5 keywords | 25.61–25.81 | 0.746 | 0.867 | 1.43 / 9 | 23 of 44 |
← swipe the table sideways for the rest of the columns
The clean cell is the one that settles where the churn comes from. On cost of living comparison, the two days’ competitor sets were identical — Jaccard 1.0, the reference frame did not move — and the term list still swapped 20 of its 80 entries, with the surviving terms shifting an average of 11 places. Nothing outside Surfer changed; the list did. The counter-case sits two rows down: ai content detector, competitor set also identical, term list Jaccard 1.0, every score to the digit. The same measurement can read “stable”, so when it reads “unstable” that is the tool, not our ruler.
Two of the five keywords (gpu comparison, ai voice generator) did see their competitor sets turn over (Jaccard 0.667), and for those the drift is confounded: a different SERP means a partly different analysis input, and we do not apportion the blame. The conclusion rests on the clean cells.
What this changes in practice: cite Surfer’s term list with a date on it. It is “the terms Surfer recommended on 2026-08-24”, not “the terms Surfer recommends”. The scores you can treat as repeatable — repeatably unrelated to rank, per the section above.
Five keywords, one repeat, ~26 hours. That is enough to show the term list can churn while its inputs hold still, and that the scores do not churn. It is not enough to say how often, how far, or whether a week does what a day did. Surfer also ticks five competitors into its guidelines by default; that default selection came back identical on 2 of 5 keywords, which moves what your draft is graded against — one reading, same caveat.
Length versus rank, from its own table
Every competitor row carries a word count beside its position, so we asked the question we ask of every table that does: do longer pages rank higher? Inside each keyword, Spearman between Surfer’s own position and its own word count, over 187 rows across 20 keywords: mean ρ , median +0.052, 11 keywords positive and 9 negative.
Both columns are Surfer’s, and its word count is undocumented, so this row sits beside our corpus and the other tools’ tables on the length-versus-rank finding as a fifth source and is never pooled with them. Five sources, five numbers, no average.
What we couldn’t test
The report lists as not run this round. Nothing on this page says anything about them.
Price, limits and speed. We recorded no price, no time-per-document benchmark and no paid-tier limits. The trial cost $0; that is the only cost figure on this page. We also never saw the trial’s real remaining balance — the account’s usage page rendered empty on three attempts across both days — so document counts here come from our own ledger, checked against the plan’s nominal 30 a month.
Factual accuracy of generated content. The trial includes AI writing credits and we spent none of them; every document ran in “write it yourself” mode. Nothing here speaks to what Surfer writes, only to how it analyses.
The AI Search panel. Alongside the SEO score, Surfer now shows a parallel “AI Search” analysis aimed at chatbots and AI overviews. We recorded that it exists and did not measure it. Its numbers appear nowhere on this page.
How we tested
Every tool runs against the same frozen fixture and the same ground truth (serp-truth-v1 + terms-truth-v1): the live Google (us) top 10 for each keyword, the body text of those pages, and the terms those pages share. Country, language and device stay constant across tools. The full procedure is on our methodology page.
Surfer’s API sits above the trial tier, so we drove the app through its UI and read the term panel and the competitor table out of the rendered editor’s component state rather than off the page’s markup, whose class names change with every build. One document per keyword, created with the country set to United States and left otherwise at defaults. The keyword and country on each capture were read back programmatically and checked against the fixture before anything was computed — a necessary step, because Surfer’s keyword field silently autocompletes: typing one keyword and confirming carelessly can create a document for a different one. Two extractor versions ran, 1 capture on v1, 19 captures on v1.1; the device setting exists only in the creation wizard, not in the captured state, so this page treats “desktop” as an operator record rather than captured data.
Before publishing, an independent script recomputed every headline number from the raw captures without importing the run’s code, and wrote the strict and lenient score-to-rank figures to a sidecar file that this page reads. The scripts are scripts/tests/content-opt/surfer/{extract_report.js,run.py}. First tested 2026-08-24, consistency re-run 2026-08-25, on the Pro trial, against app.surferseo.com.
Frequently asked questions
Does Surfer's content score predict Google rankings?
Not on the 96 matched pages across 18 keywords we tested 2026-08-24. Against the real Google (us) position of the URLs that matched our ground truth exactly, the mean Spearman correlation was −0.102. Our ground truth holds some URLs that Ahrefs returned truncated; rescue those and the figure is −0.044 over 107 URLs — same sign, still near zero. Against the rank Surfer crawled itself, over all 187 scored competitor rows, the mean was −0.084.
What about Surfer's domain score?
Its competitor table carries a second number, a 0–10 domain score. Correlated the same way, it came to +0.185 against real Google (us) position over 101 URLs and 18 keywords, and +0.165 against Surfer's own ranking. That is further from zero than the content score, and still short of anything we would call predictive. Twenty keywords, no significance test; we publish the numbers, not a conclusion.
Are Surfer's term suggestions stable from day to day?
Not reliably. We created a brand-new document for the same 5 keywords 25.61–25.81 hours after the first run. On 3 of the 5, the recommended-term list came back with a Jaccard overlap below our 0.7 stability threshold (mean 0.746, worst 0.524). The scores moved far less: across 44 URLs present in both runs, the content score's mean absolute change was 1.43 points and 52% of them did not change at all; the domain score did not move on any of its 46 URLs. Treat the term list as a snapshot with a date on it.
How good are Surfer's term suggestions?
Against the terms the ranking pages actually share, thin. We took the first 30 terms of its recommended panel — Surfer publishes no numeric importance, so its own display order is the ranking — and matched them on unigrams and bigrams against our ground-truth top 30. Precision@30 came to 0.161 and recall@30 0.333 across 20 keywords. It ran from recall 0.500 on "best help desk software" to 0.067 on "best ai chatbot". And the day-two re-run above says the list itself moves overnight.
What does Surfer cost?
We did not record a price on this run. The whole test ran on a seven-day Pro trial for $0: 20 documents for the full fixture, then 5 more for the consistency re-run, 25 in total against the plan's nominal 30 documents a month. Pricing, speed and limits are the cost test, and that test was not run, so this page carries no price rather than a guess.
How did you test Surfer?
Against the same frozen fixture and the same ground truth as every other tool here: the live Google (us) top 10 for each keyword, the body text of those pages, and the terms those pages share. Surfer's API sits on a higher tier than the trial, so we drove the app through its UI and read the term panel and the competitor table out of the rendered editor's component state. One Content Editor document per keyword returned both at once. The keyword and country on each capture were read back programmatically and checked against the fixture before anything was computed. An independent script then recomputed every headline number from the raw captures. Full method on our methodology page.
We have no affiliate relationship with Surfer SEO and have not applied for one, so nothing on this page earns us a commission. This test ran on a seven-day Pro trial and cost $0. Outbound links to Surfer SEO are marked nofollow. Where a page on this site does carry an affiliate link, it says so at the top. A commission cannot change a test result: the fixture is frozen, the extraction is scripted, and the raw captures are archived. More on who we are and how we make money.
Alternatives we tested
Other tools in this category we have run against the same fixture and the same ground truth. Where two runs covered different keywords, the comparison page says so instead of averaging across them.
- Frase — A research-to-draft pipeline with no term list, and a content score that did not predict Google rank in our sample. Surfer SEO vs Frase Tested 2026-08-18
- NeuronWriter — A weighted term list and a SERP that tracks Google's closely, under a content score that did not predict rank in our sample. Surfer SEO vs NeuronWriter Tested 2026-08-21
- Scalenut — A weighted term list and a scored competitor table, under a content score that did not track rank against Google or against its own SERP. Eight-keyword run. Surfer SEO vs Scalenut Tested 2026-08-21
- Clearscope — The widest term coverage we have measured, under letter grades that did not track even its own SERP. Three-keyword pilot; full run pending. Surfer SEO vs Clearscope Tested 2026-08-20
All 5 on one page, a row per tool and a column per measurement: best content optimization tools.