Finding · overnight consistency

Whether the terms and the score a tool gives you today are the ones it gives you tomorrow.

You paste a keyword, the tool hands back a score and a list of words to use, and you write to it. Run the same keyword again tomorrow and you get a different list. We opened a brand-new analysis for the same 5 keywords at least a day after the first, in 3 tools, and diffed what came back. NeuronWriter dropped 23.6% of its own recommended terms on average, and anywhere from 9.2% to 57.4% keyword by keyword. Surfer SEO dropped 16.1% of its own recommended terms on average, and anywhere from 0.0% to 31.2% keyword by keyword. The scores held much steadier. Read the ranges rather than the averages: on one keyword a list came back word for word, on another more than half of it was gone.

Three separate runs, not a race

Each tool was re-run on its own days, and the measured gap between the two runs is not the same length in each case: Frase 2026-08-18 to 2026-08-19, 25.0–25.1 hours, NeuronWriter 2026-08-19 to 2026-08-20, 25.9–36.0 hours and some of those a lower bound, Surfer SEO 2026-08-24 to 2026-08-25, 25.6–25.8 hours. A longer wait gives a SERP more time to move, so the rows below sit beside each other and are never subtracted from one another. Five keywords, chosen because all three runs happened to cover them, cannot rank three products.

How this was measured

The second run is a new analysis in each product, not yesterday’s report reopened. Reopening tests whether the page still renders; it cannot see whether the tool would answer the same question the same way.

Two layers, kept apart. For the term list we lowercase both runs’ words and compare the sets, and we also count the plainer figure: of the words run one recommended, how many were gone by run two. For the score we take the URLs both runs scored and measure the absolute move on each one. A tool’s term list can churn while its scores sit still, which is what happened, so a single stability figure would hide the result.

Both tools with a term list were recomputed from their raw captures under one implementation rather than copied from their own reports, and the recomputed values match what each run published. Frase has no raw capture to recompute, so its row carries its report’s own figures and says so. The threshold below which we call a list changed, 0.7, is set in our method, not by any vendor.

Nothing new was fetched. The captures were already in the repo from the runs themselves, so this cost no quota and no API call. The script is scripts/tests/content-opt/cross/mine_findings.py --finding3, and the report it writes is data/reports/tests/cross-findings-2026-08-25.json.

The term lists moved

Two of the three tools publish a list of terms to cover. Neither returned the same list a day later, and neither missed by a consistent amount.

Term list on run one against the same keyword re-analysed on run two, averaged over 5 keywords inside each tool. Read across a row. The two runs sit on different days over different gaps, so the columns are not a comparison.
ToolRunsGap, hoursMean JaccardWorst keywordTerms dropped, meanRangeBelow 0.7
NeuronWriter2026-08-19to 2026-08-2025.9–36.0part lower bound0.6370.27223.6%9.2%57.4%2 of 5
Surfer SEO2026-08-24to 2026-08-2525.6–25.80.7460.52416.1%0.0%31.2%3 of 5

← swipe the table sideways for the rest of the columns

No combined row. The runs cover different days over different gaps, so an average across them would state a figure neither tool measured.

Why the range and not the mean

NeuronWriter lost 23.6% of its terms on average, but that average covers cost of living comparison, where only 7 of 76 words went, and gpu comparison, where 54 of 94 went. Surfer SEO lost 16.1% of its terms on average, but that average covers ai content detector, where the second run returned the list unchanged, and gpu comparison, where 25 of 80 went. Quote the average and a reader will plan for a list that is mostly right tomorrow. On the keywords where it is not, it is not close.

NeuronWriter, keyword by keyword

2 of 5 keywords came back under the 0.7 threshold: gpu comparison, ai presentation maker. The top-30 slice moved with the full list rather than staying put underneath it, so the churn is not confined to the tail. Full run on the NeuronWriter review.

NeuronWriter’s term list on 2026-08-19 against the same keyword re-analysed on 2026-08-20, one row per keyword, least alike first. Jaccard is the share of the two lists that overlap once you pool them: 1.000 means the second run returned the same words.
KeywordTerms, run 1Terms, run 2JaccardTop-30 JaccardDroppedCompetitor set
gpu comparison94930.2720.30454 of 9457.4%0.750
ai presentation maker87850.6540.62219 of 8721.8%0.735
ai content detector84890.7130.76512 of 8414.3%0.903
ai voice generator100990.7460.71415 of 10015.0%0.475
cost of living comparison76790.8020.6677 of 769.2%0.933

← swipe the table sideways for the rest of the columns

Surfer SEO, keyword by keyword

3 of 5 keywords came back under the 0.7 threshold: cost of living comparison, gpu comparison, ai voice generator. The top-30 slice moved with the full list rather than staying put underneath it, so the churn is not confined to the tail. Full run on the Surfer SEO review.

Surfer SEO’s term list on 2026-08-24 against the same keyword re-analysed on 2026-08-25, one row per keyword, least alike first. Jaccard is the share of the two lists that overlap once you pool them: 1.000 means the second run returned the same words.
KeywordTerms, run 1Terms, run 2JaccardTop-30 JaccardDroppedCompetitor set
gpu comparison80800.5240.62225 of 8031.2%0.667
cost of living comparison80800.6000.50020 of 8025.0%1.000
ai voice generator79790.6290.57918 of 7922.8%0.667
ai presentation maker79790.9750.9351 of 791.3%1.000
ai content detector80801.0000.9350 of 800.0%1.000

← swipe the table sideways for the rest of the columns

The scores barely moved

Same two runs, different question: on the URLs both runs scored, how far did the number travel? Much less than the term lists did. The three scores are three different 0–100 definitions, so the comparable thing is each tool’s own gap between its two runs, never the scores themselves.

Absolute move in each tool’s own score between two runs of the same keyword, over the URLs both runs scored. Read across a row: a mean of 1.4 on a 0–100 score means a typical page came back within a point or two. Where a run flags URLs whose crawl itself flipped, the second line in the cell is the same figure without them.
ToolScoreURLs comparedMean moveLargest moveIdenticalGap, hours
Frase (thin sample)seo_score50.0005 of 5100%25.0–25.1
NeuronWritercontent_score124121 excl. fetch flips4.473.58 excl.7019 excl.14 of 12411%25.9–36.0
Surfer SEOseo_score441.43923 of 4452%25.6–25.8

← swipe the table sideways for the rest of the columns

NeuronWriter: the crawl moved too

On 3 of NeuronWriter’s 124 URLs, the second run read a body more than three times longer or shorter than the first, or read nothing at all. The score on those rows follows whatever text came back, so they record an unsteady fetch rather than unsteady scoring, and NeuronWriter’s own report asks for the two to be counted apart. Both readings are in the table: a mean move of 4.47 and a largest of 70 across all 124, and 3.58 and 19 across the 121 that are left. Dropping them does not change where NeuronWriter lands: its scores still travelled further between runs than any other tool's here.

Frase’s zero is not a win

Frase returned the same score on every URL, and that row rests on 5 of them — one per keyword, against NeuronWriter’s 124 and Surfer SEO’s 44. An order of magnitude fewer rows buys an order of magnitude less confidence. The figure is here because leaving it out would be a choice too, not because it settles anything.

A changing top ten does not explain these cells

The obvious objection: these tools read the current top ten, the top ten changes overnight, so of course the output changes. So we pulled out the keywords where the competitor set barely moved — a set overlap of 0.9 or better, several of them identical — and looked only at those.

Read that overlap for what it measures. It compares the two runs’ competitor URLs after normalisation and nothing else — not the order they sat in, not the text on the pages behind them. It rules out one explanation, a different set of reference pages, and leaves the rest standing.

Every keyword whose competitor set overlapped 0.9 or better between the two runs, least alike term list first. The same reference URLs, and the term lists still differ.
ToolKeywordCompetitor setTerm JaccardTerms droppedMean score moveLargest
Surfer SEOcost of living comparison1.0000.60020 of 802.005
NeuronWriterai content detector0.9030.71312 of 842.148
NeuronWritercost of living comparison0.9330.8027 of 767.9370
Surfer SEOai presentation maker1.0000.9751 of 790.502
Surfer SEOai content detector1.0001.0000 of 800.000

← swipe the table sideways for the rest of the columns

The cleanest row is Surfer SEO on cost of living comparison. The two runs saw exactly the same competitor URLs, and the term list still swapped 20 of its 80 words. The URLs it was pointed at did not change; the output did. The most likely explanation left is what the tool did with those pages between the two runs.

Most likely, not proven. A matching URL set does not guarantee matching input: NeuronWriter’s run records 3 URLs whose second crawl read a body more than three times the length of the first, so the same address can hand a tool a different page. Only NeuronWriter’s run carries that check, so on the other rows we cannot rule the same thing out. What these rows do settle is that Google reshuffling the top ten is not the explanation.

The control sits in the same table, on the same tool — the rows where no more than 5% of the first run’s list dropped out. On ai presentation maker, Surfer SEO swapped 1 word of 79 (0.975), and its score moved by at most 2. On ai content detector, Surfer SEO returned the list word for word (1.000), and its score did not move at all. Same method, same day, same competitor URLs: when a run holds still, this measurement says so. The churn on the rows above is not an artefact of the ruler.

Part of this belongs to the keyword

gpu comparison was the worst keyword in both tools: NeuronWriter 0.272, Surfer SEO 0.524. Two products built independently landing on the same keyword points at the SERP behind it rather than at either analyser. Where the ranking pages are near-identical, ordinary crawl noise reshuffles which ones a tool leans on, and the recommended words follow. The steadiest keyword is not shared: NeuronWriter held best on cost of living comparison at 0.802, Surfer SEO held best on ai content detector at 1.000. Whatever a keyword contributes, it does not put the two tools in the same place.

Which one is steadier: we do not know

Surfer SEO has the higher mean Jaccard (0.746 against 0.637), yet it has 3 keywords under the threshold to NeuronWriter’s 2. One is split, a couple of keywords near-identical and the rest adrift; the other is middling everywhere. Pick whichever summary you like and you can name a different winner.

The gaps are not equal either. Frase waited 25.0–25.1 hours, NeuronWriter waited 25.9–36.0 hours, Surfer SEO waited 25.6–25.8 hours. Waiting longer gives the SERP more room to move, so part of any difference is the clock. Five keywords, no significance test, unequal intervals: this sample supports a statement about each tool and nothing about the order they belong in.

Date the list, or do not quote it

A term list is not a fact about a keyword. It is what one tool recommended on one day, and the same tool will hand you a different one tomorrow. Whenever you quote one — in a brief, in a report to a client, in an article about which tool covers more — carry the timestamp with it, the way you would with a rank check. Our own term-accuracy numbers on how accurate these term suggestions are follow that rule for the same reason: they score the list a tool gave us that day, and the noise measured here is baked into any comparison between two tools tested on different days.

The scores travel better. Within one tool, a move of a point or two between runs is background, so treat a small gain after an edit as nothing until it clears the tool’s own noise. That does not make the score useful for ranking, which is a separate question with its own answer.

What would overturn this

We would rather be corrected than quoted. Any of the following would change what this page says.

  • The same keywords, a wider run. Five keywords per tool is enough to show that churn happens and nowhere near enough to describe how often. Twenty keywords per tool, same method, and any of these ranges could tighten or widen.
  • Equal gaps. Our intervals run from 25.0 to 36.0 hours, and some of those are lower bounds. Re-run every tool at the same interval and the differences between them may shrink to nothing.
  • A vendor-side explanation we can test. If a list is meant to be resampled per run, or regenerated when a crawl refreshes, say so and we will report the churn as intended behaviour rather than instability. It would not change the number a reader has to plan around.
  • A steadier account of the same raw data. Our URL keys drop query strings, which merges a handful of distinct pages in one tool’s table; the report names the collisions. Re-key them and recompute. If the term Jaccard climbs above the threshold, this page is wrong.
  • A mistake in our own work. The raw captures, the script and the report are in the repo, and the two tools with a term list were recomputed independently of their own runs. Find an error in the set comparison and we will rerun it and correct the page.

Update log

One row per tool, dated by its second run. Generated from the report behind this page.
DateChange
2026-08-19Frase added: 5 keywords re-analysed after 25.0–25.1 hours, no term list, mean score move 0.00 over 5 URLs.
2026-08-20NeuronWriter added: 5 keywords re-analysed after 25.9–36.0 hours, mean term Jaccard 0.637, 23.6% of terms dropped, mean score move 4.47 over 124 URLs.
2026-08-25Surfer SEO added: 5 keywords re-analysed after 25.6–25.8 hours, mean term Jaccard 0.746, 16.1% of terms dropped, mean score move 1.43 over 44 URLs.

A tool gets a row here when it has been re-run at least a day apart, not when it is reviewed. Re-runs cost quota, so this table grows slowly.

Disclosure

Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.

Cite this

ToolVerdict. “Whether the terms and the score a tool gives you today are the ones it gives you tomorrow.https://toolverdict.ai/findings/score-stability. Last updated 2026-08-26.

Quote the tool with the figure and the dates of both runs. There is no combined number here to quote, and the churn shares are ranges, not estimates: 5 keywords per tool, recomputed by scripts/tests/content-opt/cross/mine_findings.py --finding3.