No content score we have tested predicts where a page ranks on Google.
Every tool in this category gives you a number and tells you to raise it. So we tested the number. We took pages already ranking in the Google (us) top 10, ran each one through the tool’s own audit, and checked whether the score it returned matched the position the page holds. Across 2 tools, 18 keyword runs and 119 scored pages, no score got further from zero than −0.115. Zero means the score tells you nothing about rank.
What each tool scored
| Tool | Score tested | Mean ρ | Pages | Keywords | Fixture | Tested |
|---|---|---|---|---|---|---|
| Frase | SEO score | −0.115 | 38 | 5 of 20 | keywords-v1 | 2026-08-18 |
| NeuronWriter | content score | +0.043 | 81 | 13 of 15 | keywords-v1-sub15, 15 of the 20 in keywords-v1 | 2026-08-19 |
| Total | — | 119 | 18 | 2026-08-19 |
← swipe the table sideways for the rest of the columns
These runs cover different keywords, so the rows sit side by side rather than against each other. Frase ran 5 of the 20 keywords in keywords-v1. NeuronWriter ran 13 of the 15 keywords in keywords-v1-sub15. A mean over one keyword set does not compare with a mean over another. This page reports what each score did on its own sample; it does not pick a winner.
How this was measured
We froze the keywords before touching any tool. For each one we pulled the live Google (us) organic top 10, kept the pages we could read, and treated their real positions as the answer key. Then we sent those same URLs through each tool’s audit with the keyword attached and recorded whatever score came back.
Within each keyword we correlated score against position, using Spearman. Read it this way: +1 means higher-scoring pages ranked higher every time, 0 means the score carries no rank information, and a negative value means it points the wrong way. We publish every keyword rather than the average alone, because an average near zero hides two different failures: steady noise, and violent disagreement.
The fixture, the ground truth and the coverage rules are all on the methodology page.
The direction changes from keyword to keyword
An average near zero has two possible causes. Either the score is slightly wrong in a consistent direction, which you could still work with, or it points a different way on every keyword, which you cannot. Both tools are the second kind. Sort the keywords by correlation and every chart below crosses the axis: the same tool that tracked rank on one keyword ran against it on the next.
Frase
2 keywords came out positive, 3 negative, spanning −0.747 on ai presentation maker to +0.470 on ai content detector. Full per-keyword table and coverage on the Frase review.
NeuronWriter
8 keywords came out positive, 5 negative, spanning −0.771 on ai for business automation to +0.800 on cost of living comparison. Full per-keyword table and coverage on the NeuronWriter review.
Thin rows the run flagged, kept in the chart and marked here rather than dropped: cost of living comparison (small sample, n = 4), ai presentation maker (small sample, n = 3), best ai for writing (small sample, n = 3).
Two products built their scoring independently and landed in the same place. That points at the method rather than the implementation. A score computed from the text of a page cannot see links, site authority, click behaviour or intent match, and those decide most of the ordering inside a top 10.
None of this makes the tools useless. What they pull straight off a live SERP holds up: the competitor set, and the terms those pages share. The score is the part that did not survive.
What would overturn this
We would rather be corrected than quoted. Any of the following would change what this page says.
- A positive mean that holds up. Mean ρ above +0.3 across at least 10 keywords and 100 pages, with the sign consistent per keyword rather than averaged into place.
- The same tools, more keywords. Trial quotas cut both runs short. If a full-fixture rerun moved either mean off zero, this page is wrong about that tool.
- A narrower claim from the vendor, tested. If a score only ranks drafts against one SERP and was never meant to rank pages already in it, say so and publish that test. We will rerun it on those terms.
- A controlled ranking experiment. Our test asks whether the score describes pages that already rank; it cannot tell you whether raising a score moves one. Publish matched pages where only the score changed and the ones that rose scored higher, and this page becomes the weaker evidence.
- A mistake in our own work. The reports, the keyword fixture and the ground truth are in the repo. Find an error in the matching or the statistic and we will rerun and correct the page.
Update log
| Date | Change |
|---|---|
| 2026-08-18 | Frase added: 5 keywords, 38 scored pages, mean ρ −0.115. |
| 2026-08-19 | NeuronWriter added: 13 keywords, 81 scored pages, mean ρ +0.043. |
This page changes whenever a tool with a content score finishes testing. The claim rests on the whole table, so one new row can overturn it.
Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.