No content score we have tested predicts where a page ranks on Google.
Every tool in this category gives you a number and tells you to raise it. So we tested the number. We took pages already ranking in the Google (us) top 10, ran each one through the tool’s own audit, and checked whether the score it returned matched the position the page holds. Across 4 tools, 48 keyword runs and 276 scored pages, no score got further from zero than −0.115, and a Clearscope pilot on its own SERP came to +0.108. Change the answer key, from the strict URL match to a lenient one, or to the tool’s own crawled rank, or swap in the separate domain-level score one tool exposes, and the furthest any figure on this page gets is +0.185. Zero means the score tells you nothing about rank.
What each tool scored
| Tool | Score tested | Mean ρ | Pages | Keywords | Fixture | Tested |
|---|---|---|---|---|---|---|
| Frase | SEO score | −0.115 | 38 | 5 of 20 | keywords-v1 | 2026-08-18 |
| NeuronWriter | content score | −0.052 | 110 | 18 of 20 | keywords-v1 | 2026-08-21 |
| Scalenut | content score | −0.071 | 32 | 7 of 8 | keywords-v1 ids 1–8, 8 of the 20 in keywords-v1 | 2026-08-21 |
| Surfer SEO | content score | −0.102 | 96 | 18 of 20 | keywords-v1 | 2026-08-24 |
| Total | — | 276 | 48 | 2026-08-24 |
← swipe the table sideways for the rest of the columns
These runs cover different keywords, so the rows sit side by side rather than against each other. Frase ran 5 of the 20 keywords in keywords-v1. NeuronWriter ran 18 of the 20 keywords in keywords-v1. Scalenut ran 7 of the 8 keywords in keywords-v1 ids 1–8. Surfer SEO ran 18 of the 20 keywords in keywords-v1. A mean over one keyword set does not compare with a mean over another. This page reports what each score did on its own sample; it does not pick a winner.
How this was measured
We froze the keywords before touching any tool. For each one we pulled the live Google (us) organic top 10, kept the pages we could read, and treated their real positions as the answer key. Then we sent those same URLs through each tool’s audit with the keyword attached and recorded whatever score came back.
Within each keyword we correlated score against position, using Spearman. Read it this way: +1 means higher-scoring pages ranked higher every time, 0 means the score carries no rank information, and a negative value means it points the wrong way. We publish every keyword rather than the average alone, because an average near zero hides two different failures: steady noise, and violent disagreement.
The fixture, the ground truth and the coverage rules are all on the methodology page.
The direction changes from keyword to keyword
An average near zero has two possible causes. Either the score is slightly wrong in a consistent direction, which you could still work with, or it points a different way on every keyword, which you cannot. Every tool here is the second kind. Sort the keywords by correlation and every chart below crosses the axis: the same tool that tracked rank on one keyword ran against it on the next.
Frase
2 keywords came out positive, 3 negative, spanning −0.747 on ai presentation maker to +0.470 on ai content detector. Full per-keyword table and coverage on the Frase review.
NeuronWriter
9 keywords came out positive, 9 negative, spanning −0.771 on ai for business automation to +0.800 on cost of living comparison. Full per-keyword table and coverage on the NeuronWriter review.
Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 4), ai presentation maker (small sample, n = 3), best ai for writing (small sample, n = 3), best ai writing tools (small sample, n = 4), best help desk software (small sample, n = 4).
Scalenut
4 keywords came out positive, 3 negative, spanning −1.000 on gpu comparison to +0.800 on best ai chatbot. Full per-keyword table and coverage on the Scalenut review.
Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 3), gpu comparison (small sample, n = 3), best ai chatbot (small sample, n = 4), ai presentation maker (small sample, n = 3).
Surfer SEO
7 keywords came out positive, 11 negative, spanning −0.800 on best ai for math to +0.800 on best ai chatbot. Full per-keyword table and coverage on the Surfer SEO review.
Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 3), best ai chatbot (small sample, n = 4), ai presentation maker (small sample, n = 3), best ai for writing (small sample, n = 3), best ai for math (small sample, n = 4), ai for business automation (small sample, n = 4), best ai writing tools (small sample, n = 4), best help desk software (small sample, n = 4).
4 products built their scoring independently and landed in the same place. That points at the method rather than the implementation. A score computed from the text of a page cannot see links, site authority, click behaviour or intent match, and those decide most of the ordering inside a top 10.
None of this makes the tools useless. What they pull straight off a live SERP holds up: the competitor set, and the terms those pages share. The score is the part that did not survive.
One more tool, on a trial quota: Clearscope
Clearscope’s free trial allows three reports, so treat this as a pilot, not a row in the table above. The answer key differs too. Clearscope grades the ranking pages itself, inside its own competitor table, so we correlated its letter grades (A+ down to F, mapped to numbers) against the positions in that table rather than against our frozen ground truth. Same statistic, different answer key.
3 keywords, 50 graded pages, mean ρ +0.108. The grade missed in both directions: on gpu comparison the only A+ page sits at position 7, and on cost of living comparison the pages at positions 3 and 4 both graded F. A full run on the shared fixture replaces this block once one exists. Three keywords in, this scoring system looks like the others.
Change the answer key and the answer stays on zero
The table at the top matches a tool’s competitor URLs to our ground truth on the exact path. Our ground truth came from Ahrefs, and some of its URLs came back cut short: bankrate.com/personal-finance where the page that ranks is bankrate.com/personal-finance/cost-of-living-calculator. A strict match misses those rows. A lenient rule, which counts a truth URL as matched when it is the unique prefix of exactly one URL the tool captured, rescues them, at the risk of pulling in a same-domain page that is not the one that ranked. Neither rule is the corrected value. We leave the ground truth alone and publish the correlation under both, for every run that has been recomputed that way, beside the figure each tool gets against the rank it crawled itself.
| Tool | vs. Google, strict match | vs. Google, truncated URLs rescued | vs. its own crawled rank |
|---|---|---|---|
| NeuronWriter · 2026-08-21 | −0.052110 pages, 18 keywords | −0.013122 pages, 19 keywords, 12 rescued | +0.105564 rows, 20 keywords |
| Scalenut · 2026-08-21 | −0.07132 pages, 7 keywords | +0.03940 pages, 8 keywords, 8 rescued | −0.100121 rows, 8 keywords |
| Surfer SEO · 2026-08-24 | −0.10296 pages, 18 keywords | −0.044107 pages, 19 keywords, 11 rescued | −0.084187 rows, 20 keywords |
← swipe the table sideways for the rest of the columns
Scalenut’s sign flips between the two rules: −0.071 strict, +0.039 lenient, a move of 0.110 across zero. That is not a flaw to explain away. A correlation that changes direction when you change how a URL is matched is a correlation with no direction, which is why this page asks how far each figure sits from zero and not which way it points.
NeuronWriter’s figure in the main table stays as published, −0.052 on the strict match. Under the lenient rule it reads −0.013 over 122 pages; same sign, still inside the band.
Surfer SEO’s figure in the main table stays as published, −0.102 on the strict match. Under the lenient rule it reads −0.044 over 107 pages; same sign, still inside the band.
The last column is each tool grading its own homework, and it sits in the same place. Across every framing on this page — strict, lenient, each tool’s own rank, a pilot on its own table, and one tool's separate domain score below — no figure gets further from zero than +0.185.
The one number that did a little better was not a content score
Surfer SEO’s competitor table carries a second column none of the other tools showed us: a 0–10 score for the domain, not for the page’s text. Correlated the same way as everything above, it read +0.185 against real Google (us) position over 101 pages and 18 keywords, and +0.165 against the rank the tool crawled itself — further from zero than the same run’s content score (−0.102 and −0.084) on both answer keys, and on the positive side both times.
We are not promoting that to a finding. It is one run of 20 keywords, tested 2026-08-24, with no significance test and per-keyword values that swing across zero. What the pair of numbers is fair evidence for is only this: on this sample, a property of the domain came closer to the ordering of a top 10 than any property of the page’s text — which is consistent with the whole table above, not a way out of it. Full numbers on the Surfer SEO review.
What would overturn this
We would rather be corrected than quoted. Any of the following would change what this page says.
- A positive mean that holds up. Mean ρ above +0.3 across at least 10 keywords and 100 pages, with the sign consistent per keyword rather than averaged into place.
- The same tools, more keywords. Trial quotas cut these runs short. If a full-fixture rerun moved any mean off zero, this page is wrong about that tool.
- A narrower claim from the vendor, tested. If a score only ranks drafts against one SERP and was never meant to rank pages already in it, say so and publish that test. We will rerun it on those terms.
- A controlled ranking experiment. Our test asks whether the score describes pages that already rank; it cannot tell you whether raising a score moves one. Publish matched pages where only the score changed and the ones that rose scored higher, and this page becomes the weaker evidence.
- A mistake in our own work. The reports, the keyword fixture and the ground truth are in the repo. Find an error in the matching or the statistic and we will rerun and correct the page. We have already found one in the ground truth, the truncated URLs above, and the answer held under both rules.
Update log
| Date | Change |
|---|---|
| 2026-08-18 | Frase added: 5 keywords, 38 scored pages, mean ρ −0.115. |
| 2026-08-19 | NeuronWriter added: 13 keywords, 81 scored pages, mean ρ +0.043, on 15 of the 20 fixture keywords. |
| 2026-08-20 | Clearscope pilot added: 3 keywords, 50 graded pages on its own SERP snapshot, mean ρ +0.108. Trial quota; full run pending. |
| 2026-08-21 | NeuronWriter topped up to all 20 fixture keywords and recomputed: 18 keywords, 110 scored pages, mean ρ −0.052. |
| 2026-08-21 | Scalenut added: 7 keywords, 32 scored pages, mean ρ −0.071. |
| 2026-08-24 | Surfer SEO added: 18 keywords, 96 scored pages, mean ρ −0.102. |
This page changes whenever a tool with a content score finishes testing. The claim rests on the whole table, so one new row can overturn it.
Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.
ToolVerdict. “No content score we have tested predicts where a page ranks on Google.” https://toolverdict.ai/findings/content-score-vs-rank. Last updated 2026-08-24.
The numbers here move when a run is added or redone, so quote the date alongside them.