Frase Review (2026): We Scored 38 Ranking Pages — Its Score Didn’t Predict Rank
Verdict
partialFrase runs keyword research, turns it into a brief, drafts the article and scores it 0–100. It is no longer a term-optimization tool: the rebuilt app at next.frase.io hands you no list of terms with target frequencies. We read all 83 paths of its public API spec looking for one. There is no terms or entities field anywhere.
Its content score did not predict Google rank on our sample. Across 5 keywords and 38 scored top-10 pages, the mean Spearman correlation between the SEO score and actual position was , and the sign flipped between keywords, from -0.75 to +0.47. Its own SERP data matched the real Google (us) top 10 on of domains.
The writing held up better than the scoring. Of 20 factual claims pulled from 5 generated drafts, were supported by a source, 2 were stretched, 2 could not be checked, and 1 was a fabricated number.
you want one tool that goes from a keyword to a publishable first draft with sourced outbound links.
you came for SERP term coverage, or you plan to treat its score as a rank forecast. On our data it does neither job.
Tested 2026-08-18 on the 7-day trial. Every number on this page comes from the run’s machine-readable report. See how it compares on our comparison table.
Key findings
- Across 5 keywords and 38 top-10 pages, Frase's SEO score correlated with Google rank at ρ = -0.115 (mean Spearman, tested 2026-08-18).
- Weak is not the whole of it. The sign flipped between keywords: -0.747 on “ai presentation maker”, +0.470 on “ai content detector”.
- Frase publishes no term list. None of the 83 paths in its public API spec has a terms or entities field, so we scored its keyword research instead: precision@30 0.120 and recall@30 0.203 against the terms ranking pages actually use.
- Frase's own SERP data matched the real Google (us) top 10 on 60.8% of domains, because it pulls up to 20 results and reports ranks past 10.
- Of 20 factual claims pulled from 5 generated drafts, 15 were supported by a source, 2 were stretched, 2 could not be checked, and 1 was a fabricated number.
What works
- A 7-day trial with no card, including 5 articles, 30 research runs, 50 audited pages, 250 API requests. That covered this entire test for $0.
- A real API on the trial tier. We drove the whole run by script and spent 198 of 250 requests.
- Drafting works. 5 drafts came back at 690–905 words in 65–107 seconds, with 15 of 20 checked claims traceable to a source.
- Page audits come back synchronously at a 9.4 s median, fast enough to put in a loop.
- Research exports to JSON or Markdown, drafts to JSON, HTML or Markdown. Nothing is trapped in the UI.
What doesn’t
- The content score did not predict rank on our sample: mean ρ -0.115 over 38 pages.
- No term list at all. There is no “add these words” output to act on, which is why most people buy a tool in this category.
- Its built-in SERP matched Google's real top 10 on only 60.8% of domains, so its competitor set is not the one you are competing against.
- Research runs queue serially. Twenty submitted together took 16 minutes to clear.
- All 5 drafts share one skeleton and push the target keyword into sentences verbatim. 1 of 20 claims was an invented figure.
What Frase is now
The current app lives at next.frase.io. Work goes one way: a research run gathers keywords, sub-topics and SERP results, you derive a brief from that research run, you generate a draft from the brief, and an audit scores a URL or a pasted draft from 0 to 100. You cannot feed a keyword straight into a brief, and you have to add a site to the workspace before research will run at all.
The 7-day trial asks for no card and includes 5 articles, 30 research runs, 50 audited pages, 250 API requests. Paid plans are Starter at per month billed annually ($49 month to month) and Pro at $103 per month billed annually. Prices checked 2026-08-18.
Does the score predict rank?
negativeThe score is the thing Frase asks you to act on, so we tested it directly. For each keyword we took the real Google (us) top 10, sent every URL to Frase’s page audit with that keyword attached, and recorded the SEO score it returned. Within each keyword we computed the Spearman correlation between the score and −position. A positive value means higher-scoring pages rank higher. A value near zero means the score carries no rank information.
Part of: does a content score predict Google rank? Same measurement, every tool we have run it on.
gpu comparison · 8 pages scored, 2 unscored · Spearman ρ -0.262 · Google (us), 2026-08-18
One dot per page in the live Google (us) top 10 for “gpu comparison” on 2026-08-18. If the score predicted rank, the dots would run from top-right down to bottom-left. Hover or focus a dot for the page behind it.
Table view — gpu comparison, 10 pages
| Rank | Page | SEO score | Words |
|---|---|---|---|
| 1 | GPU UserBenchmarks - 453 Graphics Cards Compared https://gpu.userbenchmark.com/ | 40 | 2536 |
| 2 | GPU Comparison https://technical.city/en/gpu | 62 | 388 |
| 3 | PassMark - Videocard Comparison https://www.videocardbenchmark.net/singleCompare.php | 41 | 335 |
| 4 | reddit.com https://www.reddit.com/r/nvidia/comments/1nki8qo/revised_and_expanded_gpu_performance_chart_for/ | — | — |
| 5 | techpowerup.com https://www.techpowerup.com/gpu-specs/ | — | — |
| 6 | GPU Benchmarks Hierarchy 2026 - Graphics Card Rankings | Tom's Hardware https://www.tomshardware.com/reviews/gpu-hierarchy,4388.html | 57 | 9540 |
| 7 | GPU comparison / graphics card comparison https://www.gpu-monkey.com/ | 82 | 1589 |
| 8 | AMD Radeon™ vs Nvidia RTX Graphics Card Gaming Benchmarks https://www.amd.com/en/products/graphics/gaming/gaming-benchmarks.html | 80 | 2627 |
| 9 | Graphics Card Comparison Chart https://www.logicalincrements.com/articles/graphicscardcomparison | 51 | 1405 |
| 10 | GPUs | GamersNexus http://gamersnexus.net/cat/gpus | 50 | 305 |
| Keyword | Pages scored | ρ SEO score | ρ overall score |
|---|---|---|---|
| gpu comparison | 8 | -0.262 | -0.156 |
| cost of living comparison | 6 | -0.086 | 0.147 |
| ai voice generator | 8 | 0.049 | 0.036 |
| ai presentation maker | 8 | -0.747 | -0.683 |
| ai content detector | 8 | 0.470 | 0.386 |
| Mean | 38 | -0.115 | -0.054 |
← swipe the table sideways for the rest of the columns
Knowing a page’s Frase score told us almost nothing about where it ranked. A page scoring 80 was no more likely to outrank a page scoring 40 than the other way round.
42 of 145 ground-truth web URLs were audited, across 5 of 20 keywords. The trial caps audited pages at 50, so we audited the five keywords with the most web results rather than all twenty. Three requests came back HTTP 500 because Frase could not fetch the page, and four finished but returned a null score because Frase extracted no body text. That leaves 38 pages with a usable score.
For ai presentation maker the correlation ran strongly negative, at -0.747. The pages ranking first and second on Google, Canva’s and Adobe’s own product pages, both scored 16 for SEO. Frase extracted 24 and 35 words of body text from them. Pages ranking fifth and eighth scored 80 and 83. The score is reading text volume and on-page structure. A brand landing page has little of either. A long affiliate listicle has plenty.
The direction is the finding here, not the exact coefficient. With 6–8 pages per keyword, any single keyword’s ρ is noisy. What survives the noise is a mean near zero and an unstable sign, so we have no evidence the score predicts rank.
Does it give the same score tomorrow?
splitHalf of it does. We sent 5 pages back through the audit hours after the first run: same URLs, same keywords, same endpoint. The SEO score came back identical on of 5. The GEO score moved on 3 of them, and elevenlabs.io lost points.
Read that next to the rank test. The SEO score repeats itself exactly and still doesn’t track rank, at a mean ρ of -0.115. It is stably wrong, and stable is worth something: two runs on the same page agree, so when the number moves you know you moved it. The GEO score can’t claim that much. Optimize a page against it and the same page can read 9 points lower the next day.
One URL per keyword: for each keyword the rank test covered, we re-audited the top-ranked page that had returned a score. That is 5 pages, not the 20-keyword fixture the rest of this page runs on.
| Keyword | Page re-audited | SEO | GEO | Overall |
|---|---|---|---|---|
| cost of living comparison | bankrate.com/personal-finance | 61 → 61 | 22 → 22 | 33 → 33 |
| gpu comparison | gpu.userbenchmark.com | 40 → 40 | 38 → 36 | 30 → 29 |
| ai voice generator | elevenlabs.io | 91 → 91 | 56 → 47 | 56 → 53 |
| ai content detector | quillbot.com/ai-content-detector | 83 → 83 | 70 → 67 | 61 → 60 |
| ai presentation maker | canva.com/Magic | 16 → 16 | 0 → 0 | 9 → 9 |
| Mean |Δ| / largest |Δ| | 5 pages | 0.0 / 0 | 2.8 / 9 | 1.0 / 3 |
← swipe the table sideways for the rest of the columns
Before → after on each score. First audit 2026-08-18, re-audit 2026-08-19, 25.0–25.1 hours apart.
So use the SEO score if you use either, and read a small GEO move as noise rather than progress.
5 pages, one repeat. That is enough to show the GEO score moves on its own and the SEO score doesn’t. It is not enough to say how often it moves, or how far it can go. Every score that moved went the same way, down. One repeat also can’t separate drift inside Frase from a page that changed under us.
Do its keyword suggestions match what ranks?
partialOur standard test compares a tool’s recommended terms against the terms the current top-10 pages actually use. Frase publishes no such list, so we substituted the closest thing it does return: the top 30 entries of its research keyword output, matched against our ground-truth top 30 terms on unigrams and bigrams. That compares keyword research against term coverage, which is not the same job. A low number here mostly means Frase does not do this job, not that it does it badly.
gpu comparison · 13 of 30 ground-truth terms matched · recall@30 0.433 · precision@30 0.146 · built from 7 readable pages · Google (us), 2026-08-18
Frase suggested · top 30 of 40
29 of 30 contain a term the ranking pages share
- gpu comparison33,100
- graphic card comparison9,900
- graphics card comparison chart5,400
- gpu comparison chart5,400
- graphics card ranking2,400
- graphics card benchmarks1,900
- gpu ranking1,600
- nvidia graphics cards ranked1,600
- gpu performance chart1,300
- gpu comparison tool1,000
- gpu performance comparison1,000
- nvidia gpu comparison1,000
- nvidia graphics card comparison880
- gpu benchmark list720
- video card comparison720
- how much vram do i need480
- best gpu for 1440p gaming390
- best gpu comparison site260
- gpu comparison benchmarks90
- best gpu under 300 dollars210
- how to compare gpus40
- gpu hierarchy 202690
- nvidia vs amd gpu comparison20
- gpu tier list for gaming10
- laptop gpu comparison260
- rtx 5090 vs rtx 4090 benchmarks10
- gpu performance per dollar50
- rtx vs gtx comparison10
- integrated vs dedicated gpu comparison10
- all graphics cards ranked140
Top-10 pages actually use · top 30
13 of 30 are covered by a suggestion
- rtx
- gpu
- radeon
- nvidia rtx
- benchmarks
- graphics
- nvidia
- amd
- benchmark results
- geforce rtx
- geforce
- cpu
- graphics card
- graphics cards
- cards
- ryzen
- gaming
- radeon radeon
- rtx geforce
- gre
- card graphics
- rtx radeon
- rtx graphics
- gpu comparison
- rtx super
- amd radeon
- radeon nvidia
- comparison
- amd ryzen
- gpus
Left column in the order Frase returned it; right column in frequency order from the frozen ground truth. We re-sort neither list and hide no misses: the grey on the right is the finding. Precision and recall come from the test report rather than being recomputed here.
Free · builds the right-hand column live from the current top 10, with the same extraction this test used.
Seven of the twenty fixture keywords: the five the score test audited, plus the best and worst cases named below. All twenty are in the run’s report file.
| Measure | Value |
|---|---|
| Precision@30, mean over 20 keywords | |
| Recall@30, mean over 20 keywords | |
| Frase SERP vs. real Google (us) top 10, domain overlap |

The spread is wide. Best case, digital nomad visa countries: recall 0.400, precision 0.188, with country names and visa phrasing coming through. Worst case, ai for customer service: precision and recall both 0.000. Not one of its 40 suggested keywords matched a high-frequency term on the pages that rank.
The two columns above show how it fails. On gpu comparison, 29 of 30 suggestions contain a term the ranking pages share. But nearly all of them contain the head word, so between them they cover only 13 of the 30 terms those pages actually use. The suggestions are mostly restating the query back at you.
The SERP overlap of 60.8% has a mechanical cause. Frase pulls up to 20 results and reports ranks past 10, so forum and retail URLs such as reddit.com and newegg.com enter its top-10 view of a SERP where Google’s actual top 10 has none. Check that against a real SERP before you trust its competitor set. Two keywords sat at the extremes: digital nomad visa countries matched 8 of 8, best crm for small business matched 1 of 7.
| Keyword | Suggestions | P@30 | R@30 | SERP overlap |
|---|---|---|---|---|
| cost of living comparison | 40 | 0.145 | 0.300 | 6/9 |
| gpu comparison | 40 | 0.146 | 0.433 | 5/10 |
| ai voice generator | 40 | 0.171 | 0.233 | 5/9 |
| ai content detector | 40 | 0.104 | 0.167 | 5/8 |
| best ai chatbot | 27 | 0.047 | 0.067 | 5/6 |
| ai presentation maker | 16 | 0.267 | 0.267 | 5/9 |
| ai for small business | 29 | 0.093 | 0.167 | 4/8 |
| ai for coding | 35 | 0.100 | 0.200 | 5/6 |
| ai for teachers | 21 | 0.098 | 0.167 | 4/8 |
| best crm for small business | 38 | 0.179 | 0.333 | 1/7 |
| best ai for writing | 40 | 0.100 | 0.167 | 4/5 |
| best ai for math | 21 | 0.182 | 0.267 | 3/5 |
| ai for customer service | 40 | 0.000 | 0.000 | 4/6 |
| ai for business automation | 29 | 0.071 | 0.133 | 4/7 |
| best ai seo tools | 33 | 0.062 | 0.100 | 4/6 |
| ai for marketing | 40 | 0.025 | 0.033 | 3/8 |
| best ai writing tools | 40 | 0.163 | 0.233 | 3/6 |
| best laptop for programming | 12 | 0.140 | 0.233 | 3/7 |
| best help desk software | 23 | 0.128 | 0.167 | 5/6 |
| digital nomad visa countries | 40 | 0.188 | 0.400 | 8/8 |
| Mean, 20 keywords | — | 0.120 | 0.203 | 60.8% |
← swipe the table sideways for the rest of the columns
You can rebuild the ground truth yourself. Our Top 10 Term Finder runs the same extraction on any keyword you type: fetch the live top 10, pull the body text, keep the terms those pages share.
How accurate is the generated content?
mostly sourcedWe generated 5 drafts through the full pipeline at the API’s minimum target of 500 words. They came back at 690–905 words in 64.8–106.5 seconds. From those drafts we pulled 20 checkable factual claims and tried to trace each one back to a source.
| Verdict on claim | Count |
|---|---|
| Supported by a source | 15 |
| Broadly right, wording stretched | 2 |
| No checkable source either way | 2 |
| Contradicted by the source | 1 |

The error was a number. The draft stated that Adobe Express turns a text prompt into a finished presentation in under two minutes. Adobe’s own page says “in minutes” and carries no such figure, so the two came from nowhere. Two of the unverifiable claims have the same shape: a specific-sounding range with nothing behind it.
All 5 drafts share one skeleton: three H2 sections, the last always titled “The Bottom Line”, each section opening with a bolded sentence. The target keyword goes into running prose verbatim, which reads badly. Every outbound link points at a page the research step had already fetched.
Frase scored its own drafts 75–81 for SEO and 69–74 for GEO. The pages that actually rank first for ai presentation maker scored 16.
Pricing, speed and limits
| Item | Measured |
|---|---|
| Single page audit, median | s, returned synchronously |
| Research run, fastest / median | 57.6 s / 510.4 s |
| 20 research runs submitted together | queued serially; the last one finished at 939.7 s, about 16 minutes |
| Draft generation | 64.8–106.5 s per draft |
| API | available on the trial tier; this run spent 198 of 250 requests and hit HTTP 429 twice while batching |
| Audit quota consumed | 42 of 50 audited pages, one per audit that completed. The three that failed with HTTP 500 were not charged (billing page, 2026-08-18) |
| Export | research to JSON or Markdown; drafts to JSON, HTML or Markdown |
| Trial | 7-day trial, no card: 5 articles, 30 research runs, 50 audited pages, 250 API requests |
| Price | Starter /mo annual, $49/mo monthly; Pro $103/mo annual |
← swipe the table sideways for the rest of the columns
The serial queue is the real constraint. One research run takes about a minute, which is fine. Twenty submitted at once do not run in parallel, so a batch of keywords is a coffee break.
What we couldn’t test yet
Full coverage of the score-versus-rank test. 5 of 20 keywords, capped by the trial’s 50 audited pages. A paid plan would let us score the full set, and the mean could move.
A contradiction in the scoring. Frase’s documentation says the score is computed live against SERP competitors, while its rescore endpoint says it makes zero external calls and returns in milliseconds. Both cannot describe the same score. Opening one of our drafts in the editor rescored it to a different number than the API had returned for the same text, which is the same puzzle from the other end. We have not worked out which description is right.
Next up: NeuronWriter, run against the same 20 keywords and the same ground truth.
How we tested
Every tool we test runs against the same frozen fixture of 20 keywords (keywords-v1) and the same ground truth (serp-truth-v1 + terms-truth-v1): the live Google (us) top 10 for each keyword, the body text of those pages, and the terms those pages share. Country, language and device are held constant across tools so the numbers stay comparable. The full procedure is on our methodology page.

We drove this run with a script rather than by hand, so the same commands reproduce it. We keep the raw API response for every audit, research run and draft on our side, so these numbers do not depend on Frase keeping our trial data. We re-run the score-versus-rank and term tests quarterly and re-check pricing monthly; the date at the top of this page changes when we do.
Last tested 2026-08-18, on the 7-day trial, against next.frase.io.
Frequently asked questions
Does Frase's content score predict Google rankings?
Not on the 38 ranking pages we scored across 5 keywords, tested 2026-08-18. We audited pages that already rank in the Google (us) top 10 and correlated Frase's SEO score with actual position. The mean Spearman correlation was -0.115, effectively zero, and the sign flipped between keywords: -0.747 on one, +0.470 on another. A page scoring 80 was no more likely to outrank a page scoring 40 than the other way round.
Does Frase still give a list of terms to include?
No. The rebuilt app at next.frase.io returns no term list with target frequencies. We read all 83 paths of its public API spec looking for one, and there is no terms or entities field anywhere. The closest thing it returns is a related-keyword list from its research step. Scored against the terms ranking pages actually use, that list reached precision@30 0.120 and recall@30 0.203 over 20 keywords.
Is Frase worth it in 2026?
It depends on which job you are buying. As a drafting pipeline it worked: five drafts came back at 690–905 words in 65–107 seconds, and 15 of 20 factual claims we pulled from them traced back to a source. As a way to decide what goes on a page, it did not. It publishes no term list, and its score did not track rank (-0.115 mean correlation). Buy it for the first job, not the second.
Does Frase have a free trial and what does it cost?
Yes. A 7-day trial with no card required, including 5 articles, 30 research runs, 50 audited pages, 250 API requests. Paid plans are Starter at $39 per month billed annually ($49 month to month) and Pro at $103 per month billed annually. Prices checked 2026-08-18.
How did you test Frase?
By script, against a frozen fixture of 20 keywords and a ground truth built the same day: the live Google (us) top 10 for each keyword, the body text of those pages, and the terms those pages share. We ran 198 API calls on the 7-day trial, audited 42 of 145 ground-truth URLs, and kept every raw response on our side. Total spend: $0. Full method on our methodology page.
We have no affiliate relationship with Frase and have not applied for one, so nothing on this page earns us a commission. Outbound links to Frase are marked nofollow. Where a page on this site does carry an affiliate link, it says so at the top. A commission cannot change a test result: the fixture is frozen, the run is scripted, and the raw responses are archived. More on who we are and how we make money.
Alternatives we tested
Other tools in this category we have run against the same fixture and the same ground truth. Where two runs covered different keywords, the comparison page says so instead of averaging across them.
- NeuronWriter — A weighted term list and a SERP that tracks Google's closely, under a content score that did not predict rank in our sample. Frase vs NeuronWriter Tested 2026-08-21
- Scalenut — A weighted term list and a scored competitor table, under a content score that did not track rank against Google or against its own SERP. Eight-keyword run. Frase vs Scalenut Tested 2026-08-21
- Surfer SEO — A full 20-keyword run: its content score did not track Google rank under either URL match, and a day-two re-run swapped terms on three of five keywords while the scores held still. Frase vs Surfer SEO Tested 2026-08-24
- Clearscope — The widest term coverage we have measured, under letter grades that did not track even its own SERP. Three-keyword pilot; full run pending. Tested 2026-08-20
All 5 on one page, a row per tool and a column per measurement: best content optimization tools.