Finding · term accuracy

How much of what these tools tell you to write about is a term the ranking pages share.

The term list is the part of these products people actually use. You get thirty-odd words with importance scores and a brief that says to cover them. We built the answer key first — the terms the pages already ranking for a keyword have in common — then asked each tool for its top thirty, put both sides through one tokenizer, and counted the overlap. Two numbers per tool: how much of what it recommended was a shared term, and how much of the shared vocabulary it found. Frase 0.12 and 0.20; NeuronWriter 0.23 and 0.30; Scalenut 0.16 and 0.29; Surfer SEO 0.16 and 0.33. Nobody is close to a list you could take as read, and nobody is at zero either.

What each tool’s list contained

Precision reads the tool’s list: of everything its top thirty recommendations come to once they are normalised, what share the ranking pages actually share. Both sides run through the same tokenizer, which breaks an entry into single words and two-word phrases, so thirty multi-word suggestions are scored against every term they contain — usually more than thirty, and that expanded set is the denominator. Recall reads the answer key: of the thirty terms those pages share, what share the tool named. A tool can score well on one and badly on the other, and they answer different questions, so both are here and neither is combined into a grade.

Each tool’s top-30 term list against the ground-truth top 30 for the same keyword, averaged over the keywords its run covered. Every value is read from that run’s report file. Read down a row: the runs cover different keywords on different days, so the columns are not a league table.
ToolPrecision@30Recall@30KeywordsFixtureTested
Frase · no term list0.1200.20320 of 20keywords-v12026-08-18
Clearscope · pilot, n = 30.2010.4223 of 3keywords-v1 ids 1-3, 3 of the 20 in keywords-v12026-08-20
NeuronWriter0.2300.29820 of 20keywords-v12026-08-21
Scalenut0.1580.2928 of 8keywords-v1 ids 1–8, 8 of the 20 in keywords-v12026-08-21
Surfer SEO0.1610.33320 of 20keywords-v12026-08-24

← swipe the table sideways for the rest of the columns

No total row. The runs cover different keyword sets captured on different days, so a mean across them would state a figure none of them measured.

Not a head-to-head

These runs cover different slices of the same frozen fixture, so the rows sit side by side rather than against each other. Frase ran 20 of the 20 keywords in keywords-v1. Clearscope ran 3 of the 20 keywords in keywords-v1. NeuronWriter ran 20 of the 20 keywords in keywords-v1. Scalenut ran 8 of the 20 keywords in keywords-v1. Surfer SEO ran 20 of the 20 keywords in keywords-v1. The shorter runs are the fixture’s first ids, so each smaller sample sits inside the larger ones rather than beside them, and all five runs share the same first three keywords. A mean over one keyword set does not compare with a mean over another. This page reports what each list did on its own sample; it does not pick a winner.

Every row is one day’s list

No two of these runs happened on the same day — they span 2026-08-18 to 2026-08-24 — and a term list does not sit still. Re-running the same keywords a day later, we measured the tools dropping between 16.1% and 23.6% of their own recommended terms overnight, on average, and up to half of them on a single keyword. So a figure in the table is the accuracy of the list one tool handed us on one day, not a property of the product. Any gap between two rows carries that much churn before it carries anything about the tools.

Clearscope is a pilot, not a row like the others

Clearscope’s free trial allows 3 reports and all 3 were spent, so its 0.201 and 0.422 rest on three keywords — the first three in the fixture, not a sample of it. At that size one keyword moves the mean by more than the distance between most of the rows above. Its recall reads highest in the table; three keywords cannot establish that, and we are not reporting it as a result. A full run replaces this row when one exists. Details on the Clearscope review.

Frase has no term list

Frase does not hand you terms to cover, so there was nothing to score directly. Its row matches the closest thing the product does offer — the related keywords in its research panel — against the same answer key, which is why it sits lowest on precision. Read it as what happens when you use that list the way people use the others, not as Frase failing at a job it does not claim.

How this was measured

The answer key came first, before any tool was opened. For each fixture keyword we pulled the real Google (us) top 10, extracted the body text of the pages we could read, and took the terms those pages share more than the language at large — TF-IDF over the ranking set, cut at thirty. That list is the ground truth, and it is frozen: every tool is scored against the same file.

Then each tool’s own top thirty, ordered by whatever importance the product exposes. Matching is on single words and two-word phrases through one tokenizer, the same one for every run, so “living calculator” counts as a hit against the truth’s “living calculator” and the tools are not rewarded or punished for how they split a phrase. Precision is measured against what that expansion produces — the tokens in the tool’s thirty entries — rather than against the entries themselves.

The ground truth is a consensus of the pages that rank, not a list of ranking factors. A term missing from it is not a bad suggestion; it is a term the ranking pages do not share. And a tool that scored 1.00 on both would only have reproduced what its competitors already wrote. The fixture, the extraction rules and the truth file are on the methodology page.

Verbatim from each report, so you can see what its run actually matched.

Frase — tested 2026-08-18
Frase publishes no term list. Substitute: top-30 research.keywords[] (related keywords) matched against ground-truth top-30 terms on unigrams+bigrams. Also reports domain overlap between Frase's own SERP and our Google (us) top 10.
Candidates offered per keyword before the top-30 cut: 1240.
Clearscope — tested 2026-08-20
Top-30 terms by Clearscope importance order vs ground-truth top-30 (TF-IDF over real top-10 body text), unigram+bigram, same tokenizer/P-R as Frase/NeuronWriter runs.
Terms offered per keyword before the top-30 cut: 3546.
NeuronWriter — tested 2026-08-21
Top-30 terms by NeuronWriter's own importance score, matched against ground-truth top-30 terms (TF-IDF over real top-10 body text) on unigrams+bigrams. Same tokenizer and P/R definition as the Frase run, so the numbers are comparable.
Terms offered per keyword before the top-30 cut: 73100.
Scalenut — tested 2026-08-21
Top-30 key terms by Scalenut's own importance score (x/10), matched against ground-truth top-30 terms (TF-IDF over real top-10 body text) on unigrams+bigrams. Same tokenizer and P/R definition as the Frase / NeuronWriter / Clearscope runs, so the numbers are comparable. Ties on importance keep Scalenut's own list order.
Terms offered per keyword before the top-30 cut: 6170.
Surfer SEO — tested 2026-08-24
Top-30 of Surfer's own recommended term list ('SEO Entities to cover', included terms in Surfer's own display order — Surfer exposes no numeric importance), matched against ground-truth top-30 terms (TF-IDF over real top-10 body text) on unigrams+bigrams. Same tokenizer and P/R definition as the Frase / NeuronWriter / Scalenut / Clearscope runs.
Terms offered per keyword before the top-30 cut: 7980.

Keyword by keyword

The averages hide how wide the spread is inside one run. Every tool has keywords where its list lands and keywords where it barely touches the answer key. Which keywords those are differs by tool.

Frase

Recall ran from 0.433 on gpu comparison down to 0.000 on ai for customer service, across 20 keywords. Coverage rules and the full term diff sit on the Frase review.

Frase’s top 30 against the ground-truth top 30, one row per keyword, lowest recall first. Hits are the normalised terms that turn up in both lists. Tested 2026-08-18.
KeywordPrecision@30Recall@30Hits
ai for customer service0.0000.0000
ai for marketing0.0250.0331
best ai chatbot0.0470.0672
best ai seo tools0.0620.1003
ai for business automation0.0710.1334
ai content detector0.1040.1675
ai for small business0.0930.1675
ai for teachers0.0980.1675
best ai for writing0.1000.1675
best help desk software0.1280.1675
ai for coding0.1000.2006
ai voice generator0.1710.2337
best ai writing tools0.1630.2337
best laptop for programming0.1400.2337
ai presentation maker0.2670.2678
best ai for math0.1820.2678
cost of living comparison0.1450.3009
best crm for small business0.1790.33310
digital nomad visa countries0.1880.40012
gpu comparison0.1460.43313

Clearscope: three keywords on a trial quota

Recall ran from 0.467 on gpu comparison down to 0.333 on cost of living comparison, across 3 keywords. Coverage rules and the full term diff sit on the Clearscope review.

Clearscope’s top 30 against the ground-truth top 30, one row per keyword, lowest recall first. Hits are the normalised terms that turn up in both lists. Tested 2026-08-20, on a trial quota.
KeywordPrecision@30Recall@30Hits
cost of living comparison0.1640.33310
gpu comparison0.1970.46714
ai voice generator0.2410.46714

NeuronWriter

Recall ran from 0.500 on ai presentation maker down to 0.100 on best ai chatbot, across 20 keywords. Coverage rules and the full term diff sit on the NeuronWriter review.

NeuronWriter’s top 30 against the ground-truth top 30, one row per keyword, lowest recall first. Hits are the normalised terms that turn up in both lists. Tested 2026-08-21.
KeywordPrecision@30Recall@30Hits
best ai chatbot0.0710.1003
best ai for writing0.0950.1334
ai for business automation0.1390.1675
ai for marketing0.1390.1675
ai for teachers0.1330.2006
ai for customer service0.1200.2006
best ai seo tools0.2070.2006
best ai for math0.2690.2337
best ai writing tools0.1320.2337
ai for coding0.2430.3009
ai for small business0.2440.33310
best crm for small business0.2220.33310
gpu comparison0.3790.36711
ai voice generator0.3060.36711
digital nomad visa countries0.1800.36711
ai content detector0.4000.40012
best laptop for programming0.3510.43313
cost of living comparison0.3110.46714
best help desk software0.2370.46714
ai presentation maker0.4290.50015

Scalenut

Recall ran from 0.400 on gpu comparison down to 0.133 on best ai chatbot, across 8 keywords. Coverage rules and the full term diff sit on the Scalenut review.

Scalenut’s top 30 against the ground-truth top 30, one row per keyword, lowest recall first. Hits are the normalised terms that turn up in both lists. Tested 2026-08-21.
KeywordPrecision@30Recall@30Hits
best ai chatbot0.0570.1334
ai for small business0.1010.2337
cost of living comparison0.0960.2678
ai presentation maker0.1380.3009
ai for coding0.1410.3009
ai voice generator0.1640.33310
ai content detector0.1830.36711
gpu comparison0.3870.40012

Surfer SEO

Recall ran from 0.500 on gpu comparison down to 0.067 on best ai chatbot, across 20 keywords. Coverage rules and the full term diff sit on the Surfer SEO review.

Surfer SEO’s top 30 against the ground-truth top 30, one row per keyword, lowest recall first. Hits are the normalised terms that turn up in both lists. Tested 2026-08-24.
KeywordPrecision@30Recall@30Hits
best ai chatbot0.0300.0672
ai for marketing0.0770.1675
ai for coding0.1330.2678
best ai for writing0.1230.2678
best ai writing tools0.1210.2678
ai for small business0.1230.3009
ai for teachers0.1300.3009
best crm for small business0.1480.3009
ai for customer service0.1410.3009
ai for business automation0.1540.33310
cost of living comparison0.1510.36711
best ai for math0.1670.36711
best ai seo tools0.2440.36711
ai voice generator0.2110.40012
ai content detector0.2730.40012
ai presentation maker0.1820.40012
best laptop for programming0.1540.40012
digital nomad visa countries0.1600.40012
gpu comparison0.2540.50015
best help desk software0.2380.50015

What these numbers do not say

None of this measures whether covering the terms helps a page rank. It measures agreement with the pages that already rank, which is what the products claim to model and not the same thing. Our score-versus-rank finding takes the second question and gets a flat answer.

A low precision is not automatically waste. Some of what a tool adds is intent it read off the SERP, or entities the truth’s frequency cut dropped, and a brief made only of consensus terms would produce a page identical to the ten already there. What the numbers do settle is the size of the editing job: on every run here, most of the list is yours to judge.

The truth file is ours. It comes from body text we could extract, which excludes the pages a crawler cannot read, and the pages it does include carry navigation and boilerplate. A different extractor would move every row in the table, though it would move them together.

What would overturn this

We would rather be corrected than quoted. Any of the following would change what this page says.

  • A cleaner answer key. Ours is TF-IDF over extracted body text, boilerplate included. Strip the navigation, recompute the truth, rescore every run. If the ordering of the rows survives that, it means more than it does now.
  • Same-day capture. Every row here was taken on its own date, and term lists move overnight. Run all the tools on one morning against one truth file and the differences between them may vanish.
  • A full run where a pilot stands. Clearscope rests on a spent trial quota. A paid run over the whole fixture could put it anywhere in this table.
  • A different cut than thirty. Thirty is our line, not a product’s. Tools that publish longer lists are penalised on precision by it; score at fifty and at the full list and publish all three.
  • A vendor definition we are testing wrong. If a list is meant as candidates to choose from rather than terms to cover, precision is the wrong statistic and we will say so and publish recall alone.
  • A mistake in our own work. The reports, the fixture, the truth file and the tokenizer are in the repo. Find an error in the matching and we will rerun it and correct the page.

Update log

One row per tool added. Generated from the run dates in each report.
DateChange
2026-08-18Frase added: 20 keywords, precision@30 0.120, recall@30 0.203.
2026-08-20Clearscope added: 3 keywords, precision@30 0.201, recall@30 0.422. Pilot on a trial quota.
2026-08-21NeuronWriter added: 20 keywords, precision@30 0.230, recall@30 0.298.
2026-08-21Scalenut added: 8 keywords, precision@30 0.158, recall@30 0.292.
2026-08-24Surfer SEO added: 20 keywords, precision@30 0.161, recall@30 0.333.

A tool joins this table when its term list has been scored against the frozen truth. The claim rests on the whole table, so one new row can change it.

Disclosure

Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.

Cite this

ToolVerdict. “How much of what these tools tell you to write about is a term the ranking pages share.https://toolverdict.ai/findings/term-list-precision. Last updated 2026-08-26.

Quote the tool with its date and its keyword count, never the table as a whole. There is no combined figure here, and each row is the list one tool returned on the day it was tested.