Scalenut Review (2026): 8 Keywords, Two Answer Keys — Its Content Score Sat on Zero Against Both
8 keywords on a trial quota, tested 2026-08-21, two answer keys.
Scalenut’s free trial allows Content Optimizer analyses and each keyword costs one, so we ran ids 1–8 of keywords-v1 and spent the quota. Its tables gave us 8–18 scored competitor rows per keyword, against 25–30 on the NeuronWriter run, so the correlations on this page rest on fewer pages and are not read beside that one. The score-to-rank figure is also given twice, under a strict and a lenient URL match, because the sign changes between them. Both are below.
Verdict
partialScalenut works one keyword at a time. A Content Optimizer analysis crawls the SERP, lists the competitors with rank, word count and a content score, extracts the key terms those pages use and weights them by importance, then scores your draft against them. The table and the list are usable. The score on top of them is not.
We found the pages it had scored inside the real Google (us) top 10 for each keyword and correlated its content score with actual position. On the URLs that matched exactly, the mean Spearman correlation was . Our ground truth holds some URLs that Ahrefs returned truncated; count those as matches and it is over 40 URLs. Against the rank Scalenut crawled itself, over every competitor it listed, .
Three framings, none further than a tenth of a point from zero, and the sign depends on how you match a URL. That is what no rank information looks like. Writing toward a higher score bought no better position on our data.
Upstream of the score it holds up. Its crawled top 10 matched Google’s real top 10 on of domains, and it hands you 61–70 weighted key terms per keyword. Scored against the terms the ranking pages share, the top 30 of that list reached precision@30 and recall@30 . The list is a checklist, not a specification.
you want one analysis to hand you the competitor set, each page’s word count and a weighted term list, and you will judge coverage yourself.
you plan to write toward the content score. It sat on zero against Google and against the ranking the tool itself recorded.
Tested 2026-08-21 on the free trial, across 8 of the 20 keywords in our frozen fixture. Every number on this page is read out of the run’s machine-readable report or the verification sidecar beside it. See how it compares on our comparison table.
Key findings
- Scalenut's content score did not predict Google rank, under either answer key. Against the real Google (us) position of the 32 URLs that matched our ground truth exactly, the mean Spearman was −0.071 over 7 keywords (tested 2026-08-21).
- Rescue the ground-truth URLs Ahrefs returned truncated and the figure becomes +0.039 over 40 URLs and 8 keywords. The sign flips; the distance from zero does not. A correlation that changes direction with the URL-matching rule is telling you it has no direction.
- Grading its own homework gave the same answer. Against the rank Scalenut crawled itself, over all 121 competitor rows, the mean was −0.100: 3 keywords positive, 5 negative.
- Its term list is thin. The top 30 key terms by its own importance score reached precision@30 0.158 and recall@30 0.292 against the terms the ranking pages share, over 8 keywords. Its crawled top 10 matched Google's on 87.2% of domains.
- This is an eight-keyword run on a trial quota, with 8–18 scored competitor rows per keyword against 25–30 on the NeuronWriter run. Same measurement, smaller sample; the two ρ figures are not a head-to-head.
What works
- A real term list: 61–70 key terms per keyword, each with an importance score out of 10 and a suggested usage range.
- Every competitor row carries a rank, a URL, a word count and a score. One analysis gave us T1, T2 and a length-versus-rank table at once.
- Its crawled top 10 matched the real Google (us) top 10 on 87.2% of domains. The competitor set it shows you is close to the one you are up against.
- The trial is real: 8 full analyses, and this whole test cost $0.
What doesn’t
- The content score did not track rank: −0.071 against Google on a strict URL match, +0.039 on a lenient one, −0.100 against its own ranking. Near zero three ways.
- Its top 30 terms: precision@30 0.158, about one term in six, and recall@30 0.292. Lower on both than the NeuronWriter and Clearscope lists scored the same way, on their own keyword sets.
- Of the 65 ground-truth URLs we looked for, 23 never appeared in its tables and 10 of the 42 that did carried no score. 32 rows could be scored against Google.
- No API without a bearer token, so every analysis is UI work, and the trial's 8 analyses left nothing for a consistency re-run.
- The table shows the same number under two names, grade and optimiser score. Identical on all 121 rows that had a value, so there is one score here, not two.
What Scalenut is
The app lives at app.scalenut.com. The unit of work is a Content Optimizer analysis: one keyword, one country, one report. Each report returns a competitor table with rank, word count and content score per URL, a key-term list with an importance score out of 10 and a suggested usage range, and an editor that scores your draft as you write. The trial allows of those analyses and we used 8.
We ran on the free trial, which ends 2026-08-28. The API needs a bearer token we do not handle, so the whole test went through the UI. Pricing was not recorded this round; see what we couldn’t test.
Does the content score predict rank?
negativeThe score is the number the product asks you to work toward, so we tested it against the thing you want. For each keyword we took the real Google (us) top 10 out of our ground truth, found those URLs in Scalenut’s competitor table, and correlated the content score it gave each page with that page’s real position. Positive means higher-scoring pages ranked higher. Near zero means the score carries no rank information.
We report that correlation twice. Our ground truth came from Ahrefs, and some of its URLs came back cut short: bankrate.com/personal-finance where the page that ranks is bankrate.com/personal-finance/cost-of-living-calculator. A strict path match misses those rows. A lenient rule, which counts a truth URL as matched when it is the unique prefix of exactly one URL Scalenut captured, rescues of them on this run. Strict misses real truncations; lenient can pull in a same-domain page that is not the one that ranked. So we publish both and leave the ground truth as it is.
There is a third, easier framing, and we report it too: the same score against the rank Scalenut crawled itself, over every competitor it listed. That sample is about four times larger, but the tool is marking its own work and its ranks are not Google’s.
Part of: does a content score predict Google rank? Same measurement, every tool we have run it on.
| Keyword | URLs matched | ρ vs. Google, strict | Caveat |
|---|---|---|---|
| cost of living comparison | 3/9 | -0.500 | small sample, n = 3 |
| gpu comparison | 3/10 | -1.000 | small sample, n = 3 |
| ai voice generator | 7/9 | 0.464 | |
| ai content detector | 7/8 | 0.036 | |
| best ai chatbot | 4/6 | 0.800 | small sample, n = 4 |
| ai presentation maker | 3/9 | 0.500 | small sample, n = 3 |
| ai for small business | 5/8 | -0.800 | |
| ai for coding | 0/6 | — | no URL overlap |
| Mean over 7 keywords with a ρ | 32/65 | -0.071 |
← swipe the table sideways for the rest of the columns
ai for coding produced no correlation at all. Our ground truth holds only 6 web results for it, the rest of that top 10 being AI Mode and other non-page slots, and not one of those URLs matched a row in Scalenut’s table at the same path. The keyword contributes nothing to the strict mean. Under the lenient match it gains a row, which is why that framing counts 8 keywords to the strict 7.
5 of the 8 keywords rest on four matched pages or fewer, and the report marks them so nobody quotes them alone. Drop every flagged keyword and the finding survives: 3 keywords, 19 pages, mean ρ −0.100. Whichever subset you take, the answer sits on zero.
| Framing | Keywords | Rows in those keywords | Mean ρ |
|---|---|---|---|
| Score vs. real Google (us) position, strict URL match | 7 | 32 | |
| Score vs. real Google (us) position, truncated truth URLs rescued | 8 | 40 | |
| Score vs. the rank Scalenut crawled itself | 8 | 121 |
The middle row is not the true value and the top row is not the true value. They are the two ends of how far the truncation in our ground truth can move the answer, and the answer moves by 0.110 and crosses zero. On the findings page we ask how far each score sits from zero, not which way it points, and this is why.
The bottom row is the generous reading and it comes to nothing either. −0.100 over 121 rows, with 8–18 competitors per keyword, means the score barely tracks even the ordering the tool produced itself.
Scalenut’s competitor table exposes two score fields, grade and optimiser_score. We stored both. They were identical on all of the 121 rows that carried a value, Spearman 1.0 between them, so every correlation on this page is one score under one name. We do not know which label the Competition tab prints, and it does not matter to the numbers.
We ran this measurement on Frase (-0.115 over 38 pages) and NeuronWriter (-0.052 over 110 pages) before this one. Scalenut came in at −0.071 strict and +0.039 lenient. Different keyword sets and sample sizes, so no two of those numbers compete. What they share is where they land.
What we matched, and how
URLs are matched on the exact path, with the protocol, www, trailing slash and query string stripped. Same domain at a different path does not count. Of the 65 ground-truth web URLs for these keywords, appeared in Scalenut’s tables at the same path. Not every row in those tables carries a score: of its rows had none, and 10 of the matched ones were among them, which leaves 32 rows to correlate. High domain overlap, of its top-10 domains matching Google’s, hides a lower URL overlap. The two are not the same claim.
The tables needed no cleaning: 0 duplicate rows, 0 rows without a rank. What the run did flag is in our ground truth, not in Scalenut: the truncated URLs described above. That affects every tool scored against the same truth, and the lenient column exists to show how much.
Do its terms match what ranks?
negativeScalenut publishes a real key-term list, so this is the same measurement we ran on NeuronWriter and Clearscope. We took its top 30 terms by its own importance score and matched them, on unigrams and bigrams, against the top 30 terms our ground truth extracted from the body text of the pages that rank. Ties on importance keep Scalenut’s own order.
About one term in six is a term the ranking pages share, and the list covers under a third of those shared terms. Across 8 keywords it found 70 shared terms in total.
| Measure | Value |
|---|---|
| Precision@30, mean over 8 keywords | |
| Recall@30, mean over 8 keywords | |
| Scalenut SERP vs. real Google (us) top 10, domain overlap |
The spread across keywords is wide. Best case, gpu comparison, recalled 0.400 of the shared terms. Worst case, best ai chatbot, recalled 0.133. Same tool, same settings, same day.
Do not read these against the NeuronWriter or Clearscope figures as a head-to-head. Same definition, but this run covers 8 of the 20 fixture keywords, and the other runs covered different subsets.
| Keyword | Terms offered | Terms taken | P@30 | R@30 |
|---|---|---|---|---|
| cost of living comparison | 70 | 30 | 0.096 | 0.267 |
| gpu comparison | 70 | 30 | 0.387 | 0.400 |
| ai voice generator | 61 | 30 | 0.164 | 0.333 |
| ai content detector | 70 | 30 | 0.183 | 0.367 |
| best ai chatbot | 70 | 30 | 0.057 | 0.133 |
| ai presentation maker | 70 | 30 | 0.138 | 0.300 |
| ai for small business | 70 | 30 | 0.101 | 0.233 |
| ai for coding | 70 | 30 | 0.141 | 0.300 |
| Mean, 8 keywords | — | — | 0.158 | 0.292 |
← swipe the table sideways for the rest of the columns
You can rebuild the ground truth yourself. Our Top 10 Term Finder runs the same extraction on any keyword you type: fetch the live top 10, pull the body text, keep the terms those pages share. Compare that against whatever your tool hands you.
Length versus rank, from its own table
Every competitor row carries a word count beside its rank, so we asked the same question we ask of our own corpus: do longer pages rank higher? Inside each keyword, Spearman between Scalenut’s own rank and its own word count, over 126 rows across 8 keywords: mean ρ , median −0.107, 3 keywords positive and 5 negative.
Both columns are Scalenut’s, and its word count is undocumented, so this row sits beside our corpus and the other tools’ tables on the length-versus-rank finding and is never pooled with them. Four sources, four numbers, no average.
What we couldn’t test
The report lists as not run this round. Nothing on this page says anything about them.
Consistency. A re-run a day later costs another analysis per keyword, and all 8 went to coverage. We have that measurement for Frase and NeuronWriter, not for this tool. Treat every score here as one reading.
Price, limits and speed. We recorded no price, no time per analysis and no paid-tier limits. The trial cost $0; that is the only cost figure on this page.
Factual accuracy of generated content. The trial includes Article Writer credits and we spent none on drafts. Nothing here speaks to what Scalenut writes, only to how it analyses.
12 of the 20 fixture keywords. This run used keywords-v1 ids 1–8, the eight keywords that overlap most with the NeuronWriter and Clearscope runs. We captured every one of the eight, so nothing planned is missing; the other 12 were never run.
How we tested
Every tool runs against the same frozen fixture and the same ground truth (serp-truth-v1 + terms-truth-v1): the live Google (us) top 10 for each keyword, the body text of those pages, and the terms those pages share. Country, language and device stay constant across tools. The full procedure is on our methodology page.
Scalenut’s API needs a bearer token, which we do not handle, so we drove the app through its UI and read the competitor table and the term list out of the rendered report’s component state rather than off the page’s markup, whose class names change with every build. The keyword label on each capture came from the report’s breadcrumb, because the component state does not expose it; we checked all 8 against the fixture character by character before computing anything. One extractor version, 8 captures on v1. The captures recorded the country setting as US.
Before publishing, an independent script recomputed every headline number from the raw captures without importing the run’s code, and wrote the strict and lenient score-to-rank figures to a sidecar file that this page reads. The scripts are scripts/tests/content-opt/scalenut/{extract_report.js,run.py}. Last tested 2026-08-21, on the free trial, against app.scalenut.com.
Frequently asked questions
Does Scalenut's content score predict Google rankings?
Not on the 32 matched pages across 7 keywords we tested 2026-08-21, and we checked it two ways. Against the real Google (us) position of the URLs that matched our ground truth exactly, the mean Spearman correlation was −0.071. Our ground truth holds some URLs that Ahrefs returned truncated; rescue those and the figure moves to +0.039 over 40 URLs and 8 keywords. The sign flips and the size stays near zero, which is the point: a correlation that changes direction with the matching rule carries no rank information. Against the rank Scalenut crawled itself, over all 121 competitor rows, the mean was −0.100.
How good are Scalenut's term suggestions?
Thin. We took its top 30 key terms by its own importance score and compared them with the terms the current Google (us) top 10 actually share. Precision@30 came to 0.158 and recall@30 0.292 across 8 keywords: about one term in six is one the ranking pages share, and the list covers under a third of those shared terms. It ran from recall 0.400 on "gpu comparison" to 0.133 on "best ai chatbot".
What does Scalenut cost?
We did not record a price on this run. The whole test ran on the free trial, which allows 8 Content Optimizer analyses, and we spent 8 of them on 8 keywords for $0. Pricing, time per analysis and API access are the cost test, and that test was not run this round, so this page carries no price rather than a guess.
Why only eight keywords?
Because the trial allows 8 analyses and each keyword costs one. We spent all 8 on ids 1–8 of keywords-v1, the eight that overlap most with the other tools we have run, and kept none back for a consistency re-run. Eight keywords with 16–20 rows in each competitor table is a smaller sample than the NeuronWriter run, so the two rows do not belong side by side. Read this one on its own.
How did you test Scalenut?
Against the same frozen fixture and the same ground truth as every other tool here: the live Google (us) top 10 for each keyword, the body text of those pages, and the terms those pages share. Scalenut's API needs a bearer token we do not touch, so we drove the app through its UI and read the term list and the competitor table out of the rendered report's component state. The keyword label on each capture comes from the report's breadcrumb, checked character by character against the fixture. The table carries two score columns, grade and optimiser score; they were identical on all 121 rows that had a value, so the page reports one score. Full method on our methodology page.
We have no affiliate relationship with Scalenut and have not applied for one, so nothing on this page earns us a commission. This test ran on a free trial and cost $0. Outbound links to Scalenut are marked nofollow. Where a page on this site does carry an affiliate link, it says so at the top. A commission cannot change a test result: the fixture is frozen, the extraction is scripted, and the raw captures are archived. More on who we are and how we make money.
Alternatives we tested
Other tools in this category we have run against the same fixture and the same ground truth. Where two runs covered different keywords, the comparison page says so instead of averaging across them.
- Frase — A research-to-draft pipeline with no term list, and a content score that did not predict Google rank in our sample. Scalenut vs Frase Tested 2026-08-18
- NeuronWriter — A weighted term list and a SERP that tracks Google's closely, under a content score that did not predict rank in our sample. Scalenut vs NeuronWriter Tested 2026-08-21
- Surfer SEO — A full 20-keyword run: its content score did not track Google rank under either URL match, and a day-two re-run swapped terms on three of five keywords while the scores held still. Scalenut vs Surfer SEO Tested 2026-08-24
- Clearscope — The widest term coverage we have measured, under letter grades that did not track even its own SERP. Three-keyword pilot; full run pending. Tested 2026-08-20
All 5 on one page, a row per tool and a column per measurement: best content optimization tools.