How much of what these tools tell you to write about is a term the ranking pages share.
The term list is the part of these products people actually use. You get thirty-odd words with importance scores and a brief that says to cover them. We built the answer key first — the terms the pages already ranking for a keyword have in common — then asked each tool for its top thirty, put both sides through one tokenizer, and counted the overlap. Two numbers per tool: how much of what it recommended was a shared term, and how much of the shared vocabulary it found. Frase 0.12 and 0.20; NeuronWriter 0.23 and 0.30; Scalenut 0.16 and 0.29; Surfer SEO 0.16 and 0.33. Nobody is close to a list you could take as read, and nobody is at zero either.
What each tool’s list contained
Precision reads the tool’s list: of everything its top thirty recommendations come to once they are normalised, what share the ranking pages actually share. Both sides run through the same tokenizer, which breaks an entry into single words and two-word phrases, so thirty multi-word suggestions are scored against every term they contain — usually more than thirty, and that expanded set is the denominator. Recall reads the answer key: of the thirty terms those pages share, what share the tool named. A tool can score well on one and badly on the other, and they answer different questions, so both are here and neither is combined into a grade.
| Tool | Precision@30 | Recall@30 | Keywords | Fixture | Tested |
|---|---|---|---|---|---|
| Frase · no term list | 0.120 | 0.203 | 20 of 20 | keywords-v1 | 2026-08-18 |
| Clearscope · pilot, n = 3 | 0.201 | 0.422 | 3 of 3 | keywords-v1 ids 1-3, 3 of the 20 in keywords-v1 | 2026-08-20 |
| NeuronWriter | 0.230 | 0.298 | 20 of 20 | keywords-v1 | 2026-08-21 |
| Scalenut | 0.158 | 0.292 | 8 of 8 | keywords-v1 ids 1–8, 8 of the 20 in keywords-v1 | 2026-08-21 |
| Surfer SEO | 0.161 | 0.333 | 20 of 20 | keywords-v1 | 2026-08-24 |
← swipe the table sideways for the rest of the columns
No total row. The runs cover different keyword sets captured on different days, so a mean across them would state a figure none of them measured.
These runs cover different slices of the same frozen fixture, so the rows sit side by side rather than against each other. Frase ran 20 of the 20 keywords in keywords-v1. Clearscope ran 3 of the 20 keywords in keywords-v1. NeuronWriter ran 20 of the 20 keywords in keywords-v1. Scalenut ran 8 of the 20 keywords in keywords-v1. Surfer SEO ran 20 of the 20 keywords in keywords-v1. The shorter runs are the fixture’s first ids, so each smaller sample sits inside the larger ones rather than beside them, and all five runs share the same first three keywords. A mean over one keyword set does not compare with a mean over another. This page reports what each list did on its own sample; it does not pick a winner.
No two of these runs happened on the same day — they span 2026-08-18 to 2026-08-24 — and a term list does not sit still. Re-running the same keywords a day later, we measured the tools dropping between 16.1% and 23.6% of their own recommended terms overnight, on average, and up to half of them on a single keyword. So a figure in the table is the accuracy of the list one tool handed us on one day, not a property of the product. Any gap between two rows carries that much churn before it carries anything about the tools.
Clearscope’s free trial allows 3 reports and all 3 were spent, so its 0.201 and 0.422 rest on three keywords — the first three in the fixture, not a sample of it. At that size one keyword moves the mean by more than the distance between most of the rows above. Its recall reads highest in the table; three keywords cannot establish that, and we are not reporting it as a result. A full run replaces this row when one exists. Details on the Clearscope review.
Frase does not hand you terms to cover, so there was nothing to score directly. Its row matches the closest thing the product does offer — the related keywords in its research panel — against the same answer key, which is why it sits lowest on precision. Read it as what happens when you use that list the way people use the others, not as Frase failing at a job it does not claim.
How this was measured
The answer key came first, before any tool was opened. For each fixture keyword we pulled the real Google (us) top 10, extracted the body text of the pages we could read, and took the terms those pages share more than the language at large — TF-IDF over the ranking set, cut at thirty. That list is the ground truth, and it is frozen: every tool is scored against the same file.
Then each tool’s own top thirty, ordered by whatever importance the product exposes. Matching is on single words and two-word phrases through one tokenizer, the same one for every run, so “living calculator” counts as a hit against the truth’s “living calculator” and the tools are not rewarded or punished for how they split a phrase. Precision is measured against what that expansion produces — the tokens in the tool’s thirty entries — rather than against the entries themselves.
The ground truth is a consensus of the pages that rank, not a list of ranking factors. A term missing from it is not a bad suggestion; it is a term the ranking pages do not share. And a tool that scored 1.00 on both would only have reproduced what its competitors already wrote. The fixture, the extraction rules and the truth file are on the methodology page.
Verbatim from each report, so you can see what its run actually matched.
- Frase — tested 2026-08-18
- Frase publishes no term list. Substitute: top-30 research.keywords[] (related keywords) matched against ground-truth top-30 terms on unigrams+bigrams. Also reports domain overlap between Frase's own SERP and our Google (us) top 10.
- Candidates offered per keyword before the top-30 cut: 12–40.
- Clearscope — tested 2026-08-20
- Top-30 terms by Clearscope importance order vs ground-truth top-30 (TF-IDF over real top-10 body text), unigram+bigram, same tokenizer/P-R as Frase/NeuronWriter runs.
- Terms offered per keyword before the top-30 cut: 35–46.
- NeuronWriter — tested 2026-08-21
- Top-30 terms by NeuronWriter's own importance score, matched against ground-truth top-30 terms (TF-IDF over real top-10 body text) on unigrams+bigrams. Same tokenizer and P/R definition as the Frase run, so the numbers are comparable.
- Terms offered per keyword before the top-30 cut: 73–100.
- Scalenut — tested 2026-08-21
- Top-30 key terms by Scalenut's own importance score (x/10), matched against ground-truth top-30 terms (TF-IDF over real top-10 body text) on unigrams+bigrams. Same tokenizer and P/R definition as the Frase / NeuronWriter / Clearscope runs, so the numbers are comparable. Ties on importance keep Scalenut's own list order.
- Terms offered per keyword before the top-30 cut: 61–70.
- Surfer SEO — tested 2026-08-24
- Top-30 of Surfer's own recommended term list ('SEO Entities to cover', included terms in Surfer's own display order — Surfer exposes no numeric importance), matched against ground-truth top-30 terms (TF-IDF over real top-10 body text) on unigrams+bigrams. Same tokenizer and P/R definition as the Frase / NeuronWriter / Scalenut / Clearscope runs.
- Terms offered per keyword before the top-30 cut: 79–80.
Keyword by keyword
The averages hide how wide the spread is inside one run. Every tool has keywords where its list lands and keywords where it barely touches the answer key. Which keywords those are differs by tool.
Frase
Recall ran from 0.433 on gpu comparison down to 0.000 on ai for customer service, across 20 keywords. Coverage rules and the full term diff sit on the Frase review.
| Keyword | Precision@30 | Recall@30 | Hits |
|---|---|---|---|
| ai for customer service | 0.000 | 0.000 | 0 |
| ai for marketing | 0.025 | 0.033 | 1 |
| best ai chatbot | 0.047 | 0.067 | 2 |
| best ai seo tools | 0.062 | 0.100 | 3 |
| ai for business automation | 0.071 | 0.133 | 4 |
| ai content detector | 0.104 | 0.167 | 5 |
| ai for small business | 0.093 | 0.167 | 5 |
| ai for teachers | 0.098 | 0.167 | 5 |
| best ai for writing | 0.100 | 0.167 | 5 |
| best help desk software | 0.128 | 0.167 | 5 |
| ai for coding | 0.100 | 0.200 | 6 |
| ai voice generator | 0.171 | 0.233 | 7 |
| best ai writing tools | 0.163 | 0.233 | 7 |
| best laptop for programming | 0.140 | 0.233 | 7 |
| ai presentation maker | 0.267 | 0.267 | 8 |
| best ai for math | 0.182 | 0.267 | 8 |
| cost of living comparison | 0.145 | 0.300 | 9 |
| best crm for small business | 0.179 | 0.333 | 10 |
| digital nomad visa countries | 0.188 | 0.400 | 12 |
| gpu comparison | 0.146 | 0.433 | 13 |
Clearscope: three keywords on a trial quota
Recall ran from 0.467 on gpu comparison down to 0.333 on cost of living comparison, across 3 keywords. Coverage rules and the full term diff sit on the Clearscope review.
| Keyword | Precision@30 | Recall@30 | Hits |
|---|---|---|---|
| cost of living comparison | 0.164 | 0.333 | 10 |
| gpu comparison | 0.197 | 0.467 | 14 |
| ai voice generator | 0.241 | 0.467 | 14 |
NeuronWriter
Recall ran from 0.500 on ai presentation maker down to 0.100 on best ai chatbot, across 20 keywords. Coverage rules and the full term diff sit on the NeuronWriter review.
| Keyword | Precision@30 | Recall@30 | Hits |
|---|---|---|---|
| best ai chatbot | 0.071 | 0.100 | 3 |
| best ai for writing | 0.095 | 0.133 | 4 |
| ai for business automation | 0.139 | 0.167 | 5 |
| ai for marketing | 0.139 | 0.167 | 5 |
| ai for teachers | 0.133 | 0.200 | 6 |
| ai for customer service | 0.120 | 0.200 | 6 |
| best ai seo tools | 0.207 | 0.200 | 6 |
| best ai for math | 0.269 | 0.233 | 7 |
| best ai writing tools | 0.132 | 0.233 | 7 |
| ai for coding | 0.243 | 0.300 | 9 |
| ai for small business | 0.244 | 0.333 | 10 |
| best crm for small business | 0.222 | 0.333 | 10 |
| gpu comparison | 0.379 | 0.367 | 11 |
| ai voice generator | 0.306 | 0.367 | 11 |
| digital nomad visa countries | 0.180 | 0.367 | 11 |
| ai content detector | 0.400 | 0.400 | 12 |
| best laptop for programming | 0.351 | 0.433 | 13 |
| cost of living comparison | 0.311 | 0.467 | 14 |
| best help desk software | 0.237 | 0.467 | 14 |
| ai presentation maker | 0.429 | 0.500 | 15 |
Scalenut
Recall ran from 0.400 on gpu comparison down to 0.133 on best ai chatbot, across 8 keywords. Coverage rules and the full term diff sit on the Scalenut review.
| Keyword | Precision@30 | Recall@30 | Hits |
|---|---|---|---|
| best ai chatbot | 0.057 | 0.133 | 4 |
| ai for small business | 0.101 | 0.233 | 7 |
| cost of living comparison | 0.096 | 0.267 | 8 |
| ai presentation maker | 0.138 | 0.300 | 9 |
| ai for coding | 0.141 | 0.300 | 9 |
| ai voice generator | 0.164 | 0.333 | 10 |
| ai content detector | 0.183 | 0.367 | 11 |
| gpu comparison | 0.387 | 0.400 | 12 |
Surfer SEO
Recall ran from 0.500 on gpu comparison down to 0.067 on best ai chatbot, across 20 keywords. Coverage rules and the full term diff sit on the Surfer SEO review.
| Keyword | Precision@30 | Recall@30 | Hits |
|---|---|---|---|
| best ai chatbot | 0.030 | 0.067 | 2 |
| ai for marketing | 0.077 | 0.167 | 5 |
| ai for coding | 0.133 | 0.267 | 8 |
| best ai for writing | 0.123 | 0.267 | 8 |
| best ai writing tools | 0.121 | 0.267 | 8 |
| ai for small business | 0.123 | 0.300 | 9 |
| ai for teachers | 0.130 | 0.300 | 9 |
| best crm for small business | 0.148 | 0.300 | 9 |
| ai for customer service | 0.141 | 0.300 | 9 |
| ai for business automation | 0.154 | 0.333 | 10 |
| cost of living comparison | 0.151 | 0.367 | 11 |
| best ai for math | 0.167 | 0.367 | 11 |
| best ai seo tools | 0.244 | 0.367 | 11 |
| ai voice generator | 0.211 | 0.400 | 12 |
| ai content detector | 0.273 | 0.400 | 12 |
| ai presentation maker | 0.182 | 0.400 | 12 |
| best laptop for programming | 0.154 | 0.400 | 12 |
| digital nomad visa countries | 0.160 | 0.400 | 12 |
| gpu comparison | 0.254 | 0.500 | 15 |
| best help desk software | 0.238 | 0.500 | 15 |
What these numbers do not say
None of this measures whether covering the terms helps a page rank. It measures agreement with the pages that already rank, which is what the products claim to model and not the same thing. Our score-versus-rank finding takes the second question and gets a flat answer.
A low precision is not automatically waste. Some of what a tool adds is intent it read off the SERP, or entities the truth’s frequency cut dropped, and a brief made only of consensus terms would produce a page identical to the ten already there. What the numbers do settle is the size of the editing job: on every run here, most of the list is yours to judge.
The truth file is ours. It comes from body text we could extract, which excludes the pages a crawler cannot read, and the pages it does include carry navigation and boilerplate. A different extractor would move every row in the table, though it would move them together.
What would overturn this
We would rather be corrected than quoted. Any of the following would change what this page says.
- A cleaner answer key. Ours is TF-IDF over extracted body text, boilerplate included. Strip the navigation, recompute the truth, rescore every run. If the ordering of the rows survives that, it means more than it does now.
- Same-day capture. Every row here was taken on its own date, and term lists move overnight. Run all the tools on one morning against one truth file and the differences between them may vanish.
- A full run where a pilot stands. Clearscope rests on a spent trial quota. A paid run over the whole fixture could put it anywhere in this table.
- A different cut than thirty. Thirty is our line, not a product’s. Tools that publish longer lists are penalised on precision by it; score at fifty and at the full list and publish all three.
- A vendor definition we are testing wrong. If a list is meant as candidates to choose from rather than terms to cover, precision is the wrong statistic and we will say so and publish recall alone.
- A mistake in our own work. The reports, the fixture, the truth file and the tokenizer are in the repo. Find an error in the matching and we will rerun it and correct the page.
Update log
| Date | Change |
|---|---|
| 2026-08-18 | Frase added: 20 keywords, precision@30 0.120, recall@30 0.203. |
| 2026-08-20 | Clearscope added: 3 keywords, precision@30 0.201, recall@30 0.422. Pilot on a trial quota. |
| 2026-08-21 | NeuronWriter added: 20 keywords, precision@30 0.230, recall@30 0.298. |
| 2026-08-21 | Scalenut added: 8 keywords, precision@30 0.158, recall@30 0.292. |
| 2026-08-24 | Surfer SEO added: 20 keywords, precision@30 0.161, recall@30 0.333. |
A tool joins this table when its term list has been scored against the frozen truth. The claim rests on the whole table, so one new row can change it.
Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.
ToolVerdict. “How much of what these tools tell you to write about is a term the ranking pages share.” https://toolverdict.ai/findings/term-list-precision. Last updated 2026-08-26.
Quote the tool with its date and its keyword count, never the table as a whole. There is no combined figure here, and each row is the list one tool returned on the day it was tested.