How we test AI content optimization tools
Every tool runs against the same keywords, is scored against the same ground truth, and is published with its coverage and its date attached. This page is the whole procedure, so you can decide for yourself whether any of our numbers are worth anything.
The fixture: 20 keywords, frozen
We chose twenty keywords once and froze them on 2026-08-18, before touching any tool. Eight are informational, nine commercial, three tool-seeking; none is a brand term. They stay fixed across every tool and every round, which is the only thing that makes two tools' numbers comparable. A fixture that changes between tests is an anecdote, not a benchmark.
Freezing the list also removes an obvious way to cheat. If we picked keywords after seeing a tool's output, we could make any tool look good or bad. The list is set, and it is the same list that produced the numbers on the Frase review.
The ground truth: what actually ranks
Tools claim to optimize “for the top 10”. So we build our own top 10 and check them against it. For each fixture keyword we take the live Google (us) organic results, fetch the body text of each result, and compute the terms those pages have in common. Country, language and device stay constant across every tool, and we rebuild the whole truth set rather than reuse it whenever the tests re-run.
One thing we found while building it shapes every number downstream. Across our twenty keywords the Google top 10 yielded about 150 organic text results rather than 200, roughly 7.5 per keyword. The rest of the slots go to AI overviews, People Also Ask blocks and video packs. Five of those results were videos and carry no body text to extract, leaving 145 web URLs in the truth set. Where a keyword's coverage is unusually thin, the tool page says so.
You can run the extraction yourself. Our free Top 10 Term Finder runs the same procedure on any keyword you type: fetch the live top 10, pull the body text, keep the terms those pages share. Reproduce the reference set and you can see what the tools are being graded against.
The five measurements
1. Term recommendation accuracy
Every tool in this category promises to tell you which words to put on the page. We take its recommended terms for each fixture keyword, cut the list at 30, and compare it with the 30 highest-signal terms on the pages that actually rank. We report precision@30, how many of its recommendations are real, and recall@30, how many of the real terms it found. Both are means over the twenty keywords. High precision with low recall means a cautious tool; the reverse means a noisy one.
2. Does the score predict rank
This is the hardest test and the one that decides whether a content score is worth acting on. For each keyword we take the pages already ranking in the top 10, run every one of them through the tool's own audit with that keyword attached, and record the score it returns. Within each keyword we compute the Spearman correlation between the score and the page's position. Read it like this: +1 means higher-scoring pages rank higher without exception, 0 means the score carries no rank information at all, and a negative value means the score points the wrong way. We report the per-keyword values and their mean, never the mean alone, because a mean near zero can hide steady noise or violent disagreement, and those are different failures.
3. Consistency
The same input, run again at least 24 hours later. For a tool with a term list we report the overlap between the two lists; below 0.7 we call it unstable. For a tool that only returns a score we report the mean and maximum absolute change in that score. A recommendation that moves overnight on its own cannot be acted on.
4. Cost, speed and limits
Recorded, not graded: time to first usable result, whether you can export the output, whether an API exists and on which tier, trial terms and quotas, and list prices. We re-check pricing monthly and stamp it with the date, because it is the number that goes stale fastest.
5. Factual accuracy of generated content
For tools that write, we generate drafts from fixture keywords, pull twenty checkable factual claims out of them, and trace each claim to a source. Each claim lands in one of four buckets: supported by a source, broadly right but stretched in the wording, no checkable source either way, or contradicted by the source. We publish the counts and, where it teaches something, the specific claim that failed. Numbers that sound precise with nothing behind them are the usual failure, so we name those one by one.
Coverage, and how we report it
Trials and quotas often cut a test short. When that happens we run the largest honest subset, state exactly what fraction of the truth set it covers, and say what the cap was. A page will tell you it audited 42 of 145 URLs across 5 of 20 keywords because the trial capped audited pages. It will not report the mean as though we had run the full set.
We report the failures inside a run too: requests the tool could not complete, and requests that completed but returned nothing usable. We subtract those from the sample before computing any statistic, and the page shows the arithmetic.
When a tool does not have the feature we are measuring
Categories drift. A tool that defined this category five years ago may have rebuilt itself into something else, and then a test written for the old shape produces a meaningless number. The rule is: substitute the closest thing the tool actually returns, state the substitution in the same breath as the result, and never let the substituted figure sit in a comparison table without a footnote.
Frase is the worked example. It no longer returns a term list with target frequencies at all, so the term test ran against its keyword research output instead. That measures a different thing, and a low score there means “does not do this job” rather than “does it badly”. We spell out the substitution and what it costs on the Frase page, and its consistency test uses the score-drift variant described above rather than list overlap.
Publication rules
- No term data and no score-versus-rank data, no page. A tool with only pricing and screenshots does not get published. This is why the tool list has empty rows.
- Every number carries its date and its source test. Nobody types a number onto a tool page by hand; the page reads it out of the machine-readable report the run produced.
- Estimates are never dressed as measurements. If we did not measure it, the page says “not tested” and explains what stopped us.
- Re-test cadence. The term and score-versus-rank tests re-run quarterly, pricing monthly. The date at the top of a tool page changes when they do.
- We keep the raw responses ourselves. Every API response, audit and generated draft stays on our side, so a published number survives the trial account that produced it.
What this method does not tell you
It does not tell you that following a tool's advice will move your rankings. That would need a controlled ranking experiment over months, which we do not run. What it does tell you is narrower: whether the tool's recommendations match what currently ranks, whether its score tracks rank on pages that already rank, whether it says the same thing twice, and whether what it writes is true.
Sample sizes are small by design: twenty keywords, fewer when a trial cuts the run short. We report direction and spread rather than treating a single keyword's coefficient as a finding, and we say so on the page each time.
Method version 1, last updated 2026-08-18. We version any change to the method; a page tested under an older version keeps its original numbers until we re-run it. More about who runs this and how it is funded on the about page.