Google does not use LSI keywords, and we measured what the lists sold under that name actually contain.
Google answered this one in public. “There's no such thing as LSI keywords -- anyone who's telling you otherwise is mistaken, sorry.” That was John Mueller, Google’s Search Advocate, on 2019-07-30. Most pages about the myth stop there, and fair enough: the question is settled. What those pages cannot tell you is what the lists sold under that name actually contain. We run those tools against a frozen answer key for a living, so we can: across the four full runs we have published, 0.12 to 0.23 of what the lists recommend is a term the ranking pages actually share, and the lists replace up to a quarter of themselves overnight. Whatever that is, it is not a stable semantic index. Here are both halves, with sources.
Where LSI comes from, and why it cannot be Google
Latent semantic indexing is a real technique with a short paper trail. Bell Communications Research filed the patent — “Computer information retrieval using latent semantic structure,” US4839853A — on September 15, 1988, and Deerwester and colleagues published the method in 1990. It factorises a term-document matrix for one fixed collection of documents, and the original work ran on corpora of a few thousand abstracts. Two properties matter here. The decomposition describes the collection it was computed on, so it goes stale the moment the collection changes. And a US utility patent runs twenty years from filing, so this one expired around 2008. Anyone has been free to build LSI into a search engine for over fifteen years. Nobody serving the live web does, because the live web is neither small nor fixed.
The people who read Google’s patents for a living agreed with the people who write them. Bill Slawski, who spent two decades documenting Google’s patent filings, put it in ten words: “LSI keywords do not use LSI, and are not keywords.” And the vendors concede the mechanism even while the name keeps selling. Surfer — one of the five tools in our table — concludes on its own blog that search engines do not use latent semantic indexing. The myth survives as a label, not as a claim anyone defends.
So what are the lists? We measured them
Call them LSI keywords, semantic keywords or NLP terms — functionally they are lists of terms derived from pages that rank for your keyword. That is testable. We built an answer key first: for each of our frozen fixture keywords, the terms the real Google (us) top 10 share more than the language at large. Then we asked each tool for its top thirty and counted the overlap through one tokenizer.
Across the four full runs, precision@30 ran from 0.120 to 0.230: of what the tools recommended, that share was a term the ranking pages actually have in common. Recall ran 0.203 to 0.333 — of the thirty terms those pages do share, that share made the tools’ lists. Per-tool tables, coverage rules and everything the numbers do not say are on the term accuracy finding. The short version: most of every list is not the ranking pages’ shared vocabulary, whatever the branding implies.
The overnight behaviour is the more telling measurement. A factorisation of a fixed corpus returns the same answer tomorrow. These lists do not: reopening the same keywords as brand-new analyses at least a day later, the tools dropped between 16.1% and 23.6% of their own recommended terms on average, and up to half on a single keyword — measured here. That is the signature of a feed tracking a moving SERP, not of a semantic index. Which is fine, and even useful. It is just not what the name says.
Does covering the list move rankings?
Separate question, measured separately. The same products score your draft against their list, and the pitch is that a higher score means better odds of ranking. We correlated those scores against the real positions of four tools’ own audits of pages already ranking in the top 10. No mean correlation got further from zero than −0.115. The per-keyword numbers, both URL-matching rules and the pilot are on the score-versus-rank finding.
What this page does not say
It does not say related vocabulary is useless. A page about yoga that never mentions poses is a strange page, and covering the subtopics readers expect is writing, not algorithm-pleasing. A term list can be a reasonable research shortcut; our precision numbers just size the editing job it hands you.
We should also disclose: we ship one of these lists ourselves. Our term finder extracts the terms the ranking pages share, and it says exactly that on the page — counts over the pages that rank today, no semantics claimed. The complaint here is not the list. It is the mechanism the name smuggles in.
What would overturn this
Two claims, two ways to be wrong.
- Google, on the record. An engineering statement that Google Search runs latent semantic indexing would overturn the first half outright. It would have to come from Google, dated and specific — the 2019 statement above is exactly that kind of source, pointing the other way.
- Our measurements failing. The second half rests on the term-accuracy and stability findings, and inherits every way those pages say they could be wrong: a cleaner answer key, same-day capture, a cut other than thirty. If those numbers move, this page moves with them.
- A vendor publishing its actual derivation. If a tool documented that its list comes from something other than pages ranking for the keyword, we would test that claim instead of inferring from behaviour, and rewrite the second section around it.
Update log
| Date | Change |
|---|---|
| 2026-08-26 | Published, drawing on term lists from 5 tools (runs 2026-08-18 to 2026-08-24) and the overnight re-runs behind the stability finding. |
Pages carrying affiliate links say so at the top. A commission cannot move a test result: the keyword fixture is frozen, the runs are scripted, and we keep the raw responses. More on who we are and how we make money.
ToolVerdict. “Google does not use LSI keywords, and we measured what the lists sold under that name actually contain.” https://toolverdict.ai/findings/lsi-keywords-myth. Last updated 2026-08-26.
The first half of the claim is Google’s record, not ours — cite Google’s statement directly. The measurements are ours; quote each with its tool, date and keyword count.