Finding · score vs. rank

No content score we have tested predicts where a page ranks on Google.

Every tool in this category gives you a number and tells you to raise it. So we tested the number. We took pages already ranking in the Google (us) top 10, ran each one through the tool’s own audit, and checked whether the score it returned matched the position the page holds. Across 4 tools, 48 keyword runs and 276 scored pages, no score got further from zero than −0.115, and a Clearscope pilot on its own SERP came to +0.108. Change the answer key, from the strict URL match to a lenient one, or to the tool’s own crawled rank, or swap in the separate domain-level score one tool exposes, and the furthest any figure on this page gets is +0.185. Zero means the score tells you nothing about rank.

What each tool scored

Spearman correlation between each tool’s own score and the real Google (us) position, averaged over the keywords that produced one. Every value is read from the run’s report file.
ToolScore testedMean ρPagesKeywordsFixtureTested
FraseSEO score−0.115385 of 20keywords-v12026-08-18
NeuronWritercontent score−0.05211018 of 20keywords-v12026-08-21
Scalenutcontent score−0.071327 of 8keywords-v1 ids 1–8, 8 of the 20 in keywords-v12026-08-21
Surfer SEOcontent score−0.1029618 of 20keywords-v12026-08-24
Total—276482026-08-24

← swipe the table sideways for the rest of the columns

Not a head-to-head

These runs cover different keywords, so the rows sit side by side rather than against each other. Frase ran 5 of the 20 keywords in keywords-v1. NeuronWriter ran 18 of the 20 keywords in keywords-v1. Scalenut ran 7 of the 8 keywords in keywords-v1 ids 1–8. Surfer SEO ran 18 of the 20 keywords in keywords-v1. A mean over one keyword set does not compare with a mean over another. This page reports what each score did on its own sample; it does not pick a winner.

How this was measured

We froze the keywords before touching any tool. For each one we pulled the live Google (us) organic top 10, kept the pages we could read, and treated their real positions as the answer key. Then we sent those same URLs through each tool’s audit with the keyword attached and recorded whatever score came back.

Within each keyword we correlated score against position, using Spearman. Read it this way: +1 means higher-scoring pages ranked higher every time, 0 means the score carries no rank information, and a negative value means it points the wrong way. We publish every keyword rather than the average alone, because an average near zero hides two different failures: steady noise, and violent disagreement.

The fixture, the ground truth and the coverage rules are all on the methodology page.

The direction changes from keyword to keyword

An average near zero has two possible causes. Either the score is slightly wrong in a consistent direction, which you could still work with, or it points a different way on every keyword, which you cannot. Every tool here is the second kind. Sort the keywords by correlation and every chart below crosses the axis: the same tool that tracked rank on one keyword ran against it on the next.

Frase

2 keywords came out positive, 3 negative, spanning −0.747 on ai presentation maker to +0.470 on ai content detector. Full per-keyword table and coverage on the Frase review.

Spearman correlation between Frase's SEO score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 38 pages across 5 keywords, tested 2026-08-18.−0.80+0.8← score ran against rank · score tracked rank →ai presentation maker−0.747gpu comparison−0.262cost of living comparison−0.086ai voice generator+0.049ai content detector+0.470mean −0.115
score tracked rank score ran against rank mean −0.115
Spearman correlation between Frase's SEO score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 38 pages across 5 keywords, tested 2026-08-18.

NeuronWriter

9 keywords came out positive, 9 negative, spanning −0.771 on ai for business automation to +0.800 on cost of living comparison. Full per-keyword table and coverage on the NeuronWriter review.

Spearman correlation between NeuronWriter's content score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 110 pages across 18 keywords, tested 2026-08-21.−0.80+0.8← score ran against rank · score tracked rank →ai for business automation−0.771best ai for math−0.600ai presentation maker−0.500best help desk software−0.400ai for marketing−0.395digital nomad visa countries−0.342best ai writing tools−0.316ai for customer service−0.086gpu comparison−0.042ai voice generator+0.083best ai chatbot+0.086best ai seo tools+0.100ai for small business+0.143ai for teachers+0.214best laptop for programming+0.257ai content detector+0.325best ai for writing+0.500cost of living comparison+0.800mean −0.052
score tracked rank score ran against rank mean −0.052
Spearman correlation between NeuronWriter's content score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 110 pages across 18 keywords, tested 2026-08-21.

Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 4), ai presentation maker (small sample, n = 3), best ai for writing (small sample, n = 3), best ai writing tools (small sample, n = 4), best help desk software (small sample, n = 4).

Scalenut

4 keywords came out positive, 3 negative, spanning −1.000 on gpu comparison to +0.800 on best ai chatbot. Full per-keyword table and coverage on the Scalenut review.

Spearman correlation between Scalenut's content score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 32 pages across 7 keywords, tested 2026-08-21.−1.00+1.0← score ran against rank · score tracked rank →gpu comparison−1.000ai for small business−0.800cost of living comparison−0.500ai content detector+0.036ai voice generator+0.464ai presentation maker+0.500best ai chatbot+0.800mean −0.071
score tracked rank score ran against rank mean −0.071
Spearman correlation between Scalenut's content score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 32 pages across 7 keywords, tested 2026-08-21.

Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 3), gpu comparison (small sample, n = 3), best ai chatbot (small sample, n = 4), ai presentation maker (small sample, n = 3).

Surfer SEO

7 keywords came out positive, 11 negative, spanning −0.800 on best ai for math to +0.800 on best ai chatbot. Full per-keyword table and coverage on the Surfer SEO review.

Spearman correlation between Surfer SEO's content score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 96 pages across 18 keywords, tested 2026-08-24.−0.80+0.8← score ran against rank · score tracked rank →best ai for math−0.800ai for small business−0.657gpu comparison−0.643ai voice generator−0.635ai presentation maker−0.500ai for business automation−0.400best ai writing tools−0.400ai for marketing−0.319ai for teachers−0.290best laptop for programming−0.232ai for customer service−0.203digital nomad visa countries+0.144ai content detector+0.192best help desk software+0.400cost of living comparison+0.500best ai for writing+0.500best ai seo tools+0.700best ai chatbot+0.800mean −0.102
score tracked rank score ran against rank mean −0.102
Spearman correlation between Surfer SEO's content score and real Google position, one bar per keyword. Bars right of the axis mean higher-scoring pages ranked higher. Google (us), 96 pages across 18 keywords, tested 2026-08-24.

Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 3), best ai chatbot (small sample, n = 4), ai presentation maker (small sample, n = 3), best ai for writing (small sample, n = 3), best ai for math (small sample, n = 4), ai for business automation (small sample, n = 4), best ai writing tools (small sample, n = 4), best help desk software (small sample, n = 4).

4 products built their scoring independently and landed in the same place. That points at the method rather than the implementation. A score computed from the text of a page cannot see links, site authority, click behaviour or intent match, and those decide most of the ordering inside a top 10.

None of this makes the tools useless. What they pull straight off a live SERP holds up: the competitor set, and the terms those pages share. The score is the part that did not survive.

One more tool, on a trial quota: Clearscope

Clearscope’s free trial allows three reports, so treat this as a pilot, not a row in the table above. The answer key differs too. Clearscope grades the ranking pages itself, inside its own competitor table, so we correlated its letter grades (A+ down to F, mapped to numbers) against the positions in that table rather than against our frozen ground truth. Same statistic, different answer key.

3 keywords, 50 graded pages, mean ρ +0.108. The grade missed in both directions: on gpu comparison the only A+ page sits at position 7, and on cost of living comparison the pages at positions 3 and 4 both graded F. A full run on the shared fixture replaces this block once one exists. Three keywords in, this scoring system looks like the others.

Spearman correlation between Clearscope's content grade (A+=12 to F=0) and the position in its own competitor table, one bar per keyword. 50 pages across 3 keywords, tested 2026-08-20. Pilot on a trial quota.−0.20+0.2← score ran against rank · score tracked rank →gpu comparison+0.052cost of living comparison+0.131ai voice generator+0.140mean +0.108
score tracked rank score ran against rank mean +0.108
Spearman correlation between Clearscope's content grade (A+=12 to F=0) and the position in its own competitor table, one bar per keyword. 50 pages across 3 keywords, tested 2026-08-20. Pilot on a trial quota.

Change the answer key and the answer stays on zero

The table at the top matches a tool’s competitor URLs to our ground truth on the exact path. Our ground truth came from Ahrefs, and some of its URLs came back cut short: bankrate.com/personal-finance where the page that ranks is bankrate.com/personal-finance/cost-of-living-calculator. A strict match misses those rows. A lenient rule, which counts a truth URL as matched when it is the unique prefix of exactly one URL the tool captured, rescues them, at the risk of pulling in a same-domain page that is not the one that ranked. Neither rule is the corrected value. We leave the ground truth alone and publish the correlation under both, for every run that has been recomputed that way, beside the figure each tool gets against the rank it crawled itself.

One score per tool, three answer keys. Read across a row to see how far the matching rule moves the number; do not read down a column, because the runs cover different keywords and the samples differ.
Toolvs. Google, strict matchvs. Google, truncated URLs rescuedvs. its own crawled rank
NeuronWriter · 2026-08-21−0.052110 pages, 18 keywords−0.013122 pages, 19 keywords, 12 rescued+0.105564 rows, 20 keywords
Scalenut · 2026-08-21−0.07132 pages, 7 keywords+0.03940 pages, 8 keywords, 8 rescued−0.100121 rows, 8 keywords
Surfer SEO · 2026-08-24−0.10296 pages, 18 keywords−0.044107 pages, 19 keywords, 11 rescued−0.084187 rows, 20 keywords

← swipe the table sideways for the rest of the columns

Scalenut’s sign flips between the two rules: −0.071 strict, +0.039 lenient, a move of 0.110 across zero. That is not a flaw to explain away. A correlation that changes direction when you change how a URL is matched is a correlation with no direction, which is why this page asks how far each figure sits from zero and not which way it points.

NeuronWriter’s figure in the main table stays as published, −0.052 on the strict match. Under the lenient rule it reads −0.013 over 122 pages; same sign, still inside the band.

Surfer SEO’s figure in the main table stays as published, −0.102 on the strict match. Under the lenient rule it reads −0.044 over 107 pages; same sign, still inside the band.

The last column is each tool grading its own homework, and it sits in the same place. Across every framing on this page — strict, lenient, each tool’s own rank, a pilot on its own table, and one tool's separate domain score below — no figure gets further from zero than +0.185.

The one number that did a little better was not a content score

Surfer SEO’s competitor table carries a second column none of the other tools showed us: a 0–10 score for the domain, not for the page’s text. Correlated the same way as everything above, it read +0.185 against real Google (us) position over 101 pages and 18 keywords, and +0.165 against the rank the tool crawled itself — further from zero than the same run’s content score (−0.102 and −0.084) on both answer keys, and on the positive side both times.

We are not promoting that to a finding. It is one run of 20 keywords, tested 2026-08-24, with no significance test and per-keyword values that swing across zero. What the pair of numbers is fair evidence for is only this: on this sample, a property of the domain came closer to the ordering of a top 10 than any property of the page’s text — which is consistent with the whole table above, not a way out of it. Full numbers on the Surfer SEO review.

What would overturn this

We would rather be corrected than quoted. Any of the following would change what this page says.

  • A positive mean that holds up. Mean ρ above +0.3 across at least 10 keywords and 100 pages, with the sign consistent per keyword rather than averaged into place.
  • The same tools, more keywords. Trial quotas cut these runs short. If a full-fixture rerun moved any mean off zero, this page is wrong about that tool.
  • A narrower claim from the vendor, tested. If a score only ranks drafts against one SERP and was never meant to rank pages already in it, say so and publish that test. We will rerun it on those terms.
  • A controlled ranking experiment. Our test asks whether the score describes pages that already rank; it cannot tell you whether raising a score moves one. Publish matched pages where only the score changed and the ones that rose scored higher, and this page becomes the weaker evidence.
  • A mistake in our own work. The reports, the keyword fixture and the ground truth are in the repo. Find an error in the matching or the statistic and we will rerun and correct the page. We have already found one in the ground truth, the truncated URLs above, and the answer held under both rules.

Update log

One row per tool added. Generated from the run dates in each report.
DateChange
2026-08-18Frase added: 5 keywords, 38 scored pages, mean ρ −0.115.
2026-08-19NeuronWriter added: 13 keywords, 81 scored pages, mean ρ +0.043, on 15 of the 20 fixture keywords.
2026-08-20Clearscope pilot added: 3 keywords, 50 graded pages on its own SERP snapshot, mean ρ +0.108. Trial quota; full run pending.
2026-08-21NeuronWriter topped up to all 20 fixture keywords and recomputed: 18 keywords, 110 scored pages, mean ρ −0.052.
2026-08-21Scalenut added: 7 keywords, 32 scored pages, mean ρ −0.071.
2026-08-24Surfer SEO added: 18 keywords, 96 scored pages, mean ρ −0.102.

This page changes whenever a tool with a content score finishes testing. The claim rests on the whole table, so one new row can overturn it.

Disclosure

Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.

Cite this

ToolVerdict. “No content score we have tested predicts where a page ranks on Google.” https://toolverdict.ai/findings/content-score-vs-rank. Last updated 2026-08-24.

The numbers here move when a run is added or redone, so quote the date alongside them.