August 24, 2026·Martin Endara

The Read Set: Why Two AI Engines Can Look at the Same Web and See a Different Shortlist

Frontier models now retrieve on nearly every query — that argument is over. The variable nobody reports is how deep each one reads: a median of 50 sources for one model, 8 for another. Your visibility depends on which room your buyer walked into.

The first one sees fifty people. She works through the morning, the afternoon, the stragglers who came in on a friend’s tip. By the end of the day she has watched almost everyone who could plausibly play the part.

The second sees eight. Not out of laziness but out of method. She trusts her shortlist, she moves fast, and eight is enough to cast a role.

Two casting directors are filling the same role, in the same city, from the same pool of actors.

Now here is the part that matters. There is an actor who is, objectively, the twentieth-best fit for that role. In the first room, he gets seen, and he might win it on the day. In the second room, he does not lose the audition. He is never in the building. Nobody rejects him. Nobody has an opinion about him at all.

And he will never know which room he was in.

That is the current state of AI visibility, and almost nobody is measuring it.

The argument everyone is still having is over

For two years the entire GEO conversation has run on one assumption: the fight is to get retrieved. Make your content crawlable, structure it, publish llms.txt, make sure the engine can find and parse you when the question comes up.

That was the right fight. It is also, on current models, largely won by default.

A 2026 study of frontier models performing bibliographic lookup tasks measured how often each one actually invoked search rather than answering from memory. Across the corpus, search tools fired on 97.4% of entries. Claude Sonnet 4.6 searched on 99.9% of queries. Gemini 3 Flash on 97.6%. GPT-5 on 94.7%.

Compare that to the number still circulating in AI visibility decks: that 34% of Gemini answers and 24% of GPT-4o answers were generated without fetching any online content at all. That finding is real, it was published in Data & Policy, and it came from roughly 14,000 real conversation logs. It is also a measurement of GPT-4o and a pre-3 Gemini models that are, in this field’s timescale, archaeology.

I want to be careful about the scope of both, because scope is the whole game: the 97.4% figure comes from a specific task type structured bibliographic lookup, where a model has an obvious reason to go check. It is not a universal “frontier models always search” law, and anyone who quotes it that way is repeating the mistake in the other direction.

But the direction of travel is not ambiguous. Retrieval went from a coin flip to a default. Which means the interesting question moved.

The variable that replaced it

The same study measured something else, and this is the number I have not seen a single marketing team use.

Search depth varies across models by an order of magnitude. GPT-5 consulted a median of 50 sources per entry — mean 68.2, with a range running from 2 to 372, reflecting an architecture that issues many web queries per prompt and keeps pulling. Claude Sonnet 4.6 consulted a median of 10. Gemini 3 Flash, a median of 8.

Fifty versus eight.

Sit with what that does to a visibility strategy.

Your brand is not “visible” or “invisible” in AI search. Your brand occupies a position in a relevance ordering that the engine computes, and then the engine draws a line somewhere in that ordering and reads everything above it. Where it draws the line is a product architecture decision — how many tool calls the system is willing to spend, how aggressively it fans out the query, how much latency the product will tolerate.

If you sit near the top of the ordering, none of this touches you. You get read everywhere. Congratulations, and this article is not about you.

If you sit at position fifteen or twenty-five or forty, which is where most competent B2B companies actually sit in their category, your visibility is not a property of your content. It is a property of the read depth of whichever engine your buyer happened to open. You are cited in the deep room and structurally absent from the shallow one, on the same day, with the same page.

Now go look at any AI visibility dashboard. It will give you share of voice, citation rate, position, sentiment pooled across engines, or broken out by engine but presented on the same axis, as if a mention on a model that reads 8 sources and a mention on one that reads 50 were the same accomplishment.

They are not remotely the same accomplishment. Getting into the eight-source room is a far harder thing to do, and worth far more, than getting into the fifty-source one. No mainstream tool prices that difference, because no mainstream tool reports the depth.

The second clock

There is a layer underneath retrieval that moves on a completely different schedule, and it explains the most common complaint I hear from agencies.

Google DeepMind’s FACTS benchmark, introduced in late 2025, does something unusually honest: it splits factuality into separate dimensions rather than reporting one score. Two of them are Search accuracy when the model has web retrieval — and Parametric accuracy drawing on knowledge stored in the weights from training.

Gemini 3 Pro, the strongest overall performer on that benchmark, scored 83.8 on Search and 76.4 on Parametric. And its overall FACTS score across all four dimensions was 68.8 the highest of any model tested, which also means the best model available is wrong somewhere north of thirty percent of the time when you combine the slices.

An independent test in May 2026 ran 100 factual questions through current models and found Gemini 3.1 Pro scoring 91 out of 100 with live search enabled and 78 with search disabled. Small sample, and the author published the caveats, which is why I trust the direction more than the digits. But the shape is consistent with the benchmark: retrieval improves the answer, and it improves it on top of something the model already believes.

That something is the second clock.

When a company overhauls its site clean JSON-LD, an llms.txt that actually describes the entity, crawler access unblocked, comparison pages that answer the real buyer question, it is operating on the retrieval layer. That layer can move within a crawl cycle or two. Weeks, not quarters.

What that work does not touch is what the model already believes about the category from training: which vendors are the obvious names, which company is “the enterprise one,” which brand is a synonym for the category. That lives in the weights. It changes when the model changes.

“We fixed everything and nothing changed” is almost always these two clocks being mistaken for one. The retrieval fix worked. It just got applied to a query the model answered largely from a prior it formed before your fix existed and the prior is doing the ranking that decides whether you make the read set in the first place.

Four things this changes on Monday

Never pool engines. A pooled AI visibility score averages across systems whose read depth differs by an order of magnitude. It is the same error as averaging a thermometer and a barometer. Report per engine or report nothing.

Treat read depth as a measured variable, not a fact you memorize. The numbers above describe a specific cohort in a specific study window. Gemini 3.7 Flash and GPT-5.6 are already out; the depths will have moved. What is durable is that depth differs by engine and changes by release which makes it something to check each cycle, not a constant to write on a slide.

Spend where the room is small. Marginal gains in relevance ranking are worth far more against a shallow-read engine than a deep-read one. Moving from position 20 to position 12 changes nothing on a model reading fifty sources and changes everything on a model reading eight. If your buyers concentrate on one engine, that is where the ranking work pays.

Test both states. Run your category’s buyer questions with web search on, then again with it off or on a model configuration that leans parametric. The gap between those two answers is the most useful diagnostic in this whole discipline: it tells you what the weights already believe about your category, separate from what the engine can look up. If you look good with search and vanish without it, your position is rented. If you look good without it, you own something that survives the next retrieval change.

The part that worries me about Spanish

I will label this as reasoning rather than measurement, because I have not run the study yet.

Read depth interacts badly with source scarcity. In a market with a deep, well-indexed pool of relevant sources, an eight-source read set is a competitive sample the engine is choosing eight strong candidates out of hundreds. In a market where the indexed pool for a given B2B category is thin which is the situation across much of Spanish-language LATAM, as the citation source data from Spain already hints an eight-source read set may be most of what exists.

That cuts both ways, and that is what makes it interesting. Thin pools mean a small number of sources anchor nearly every answer, so displacement is brutally hard. It also means the barrier to entering that anchor set is lower than in a saturated English-language category, because there are simply fewer contenders in the room.

Which of those two effects dominates is an empirical question, and it is one of the things our LATAM sector studies are built to answer.

Back to the audition

The actor who never got seen did not have a content problem. He had a room-size problem, and he could not perceive it from where he was standing.

The whole discipline of AI visibility has been coaching people on their performance. Nobody is telling them how many chairs are in the room, or that the number is different in every room, or that it changed last month when the casting director got a new process.

Measure the room. Then decide how hard to work on the performance.


Definition, for the record: the read set is the ordered subset of retrieved sources a generative engine actually consults when composing an answer. Its size is determined by the engine’s retrieval architecture rather than by the querying user, and it varies across models by an order of magnitude. One 2026 study of bibliographic lookup tasks recorded a median of 50 sources consulted per entry for GPT-5 against 8 for Gemini 3 Flash. A brand ranked outside an engine’s read set is not evaluated and rejected; it is never retrieved, which is why per-engine reporting is required and pooled AI visibility scores are not interpretable.


Run it on your real domain. We measure your bleed in USD against a stated traffic base, per engine, in the language your buyers actually search in, and hand you the artifacts ready to deploy. Chairs counted.

Sources: BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation (arXiv 2604.03159) for search-invocation rates and per-model source-consultation depth · The Attribution Crisis in LLM Search Results, Data & Policy, Cambridge (arXiv 2508.00838), for the earlier-generation no-search figures · Google DeepMind FACTS benchmark (Search and Parametric dimensions), as reported in 2026 model evaluations · LumiChats, 100-question search-on/search-off comparison (May 2026).

← Back to blog