AI Search Research

How AI Systems Interpret and Rank Digital Entities

June 24, 2026 · 7 min read

Traditional search engines rank pages. AI systems assemble entities. When a language model answers a question about your company, it is not handing back a ranked list ... it is reconstructing an internal picture of who you are from every signal it has collected, then deciding whether that picture is solid enough to cite. The strength of that reconstruction is what determines whether you get named, hedged around, or left out.

That claim gets repeated a lot. We wanted to see what it looked like in server logs.

What 90 days of logs actually showed

We run a warehouse that ingests raw access logs across the properties we operate and study. Between 11 June and 8 September 2026, across 131 sites, it recorded 715,188 requests from AI crawlers. Not search engine crawlers. Not scanners. Requests from user agents belonging to AI companies, matched against a fingerprint table we maintain by hand.

Twenty-five distinct AI crawlers showed up, from sixteen vendors. Here is the top of the list:

  • Amazonbot (Amazon) ... 198,216 requests across 129 sites
  • ClaudeBot (Anthropic) ... 108,064 across 122 sites
  • GPTBot (OpenAI) ... 68,164 across 130 sites
  • Bytespider (ByteDance) ... 66,181 across 114 sites
  • meta-externalagent (Meta) ... 52,128 across 128 sites
  • OAI-SearchBot (OpenAI) ... 37,161 across 131 sites
  • Google-Extended (Google) ... 17,874 across 102 sites
  • PerplexityBot (Perplexity) ... 4,920 across 116 sites

The first thing that stood out was not the total. It was the ranking. If you had asked us to guess in advance, we would have put OpenAI first and Amazon somewhere in the middle. Amazon nearly doubled Anthropic and tripled OpenAI's main crawler. We do not have a confident explanation for that, and we are not going to invent one. It is worth knowing that the crawler you are optimizing your mental model around may not be the one actually reading your site.

The arrival dates matter as much as the volumes

Several of these crawlers did not exist in our data at the start of the window. Google-Extended first appeared on 30 July. Claude-SearchBot on 3 August. GrokBot on 11 August. DeepSeekBot on 22 July. Claude-User on 26 July.

Five new AI user agents began fetching pages across a portfolio in a single quarter. Whatever you built for last year's crawler list is already incomplete, which is an argument for structure that any parser can read rather than tuning for named agents.

What a log line proves, and what it does not

This is the part most write-ups skip, so we will be direct about it.

A crawler request proves a machine fetched a URL. It does not prove the content was ingested, retained, weighted, or ever surfaced in an answer. It does not prove attribution. We can see the fetch. We cannot see what happened after it, because none of these vendors publish that.

Two more caveats on our own numbers. First, log ingestion started on 11 June, so June is a partial month and any raw month-over-month figure that includes it is misleading. Second, the number of sites reporting into the warehouse grew from 76 to 129 across the window, so portfolio totals rose partly because coverage rose.

To get an honest trend we held the cohort fixed: the 97 sites that were already reporting in July. Across that same group, requests went from 217,918 in July to 383,053 in August. A 76 percent increase with the site count flat. That one we trust.

Volume is the boring number. Breadth is the interesting one.

Two sites in our data illustrate this better than any argument could.

One property took 107,278 AI crawler requests over the 90 days, the highest in the portfolio by a wide margin. It was visited by 11 distinct AI crawlers.

Another took 19,741 requests, roughly a fifth as many. It was visited by 23 distinct AI crawlers, the widest coverage of anything we operate.

Those are two different conditions wearing the same label. The first site is being crawled hard by a handful of systems. The second is being noticed by nearly every AI system we can identify. If your objective is to be citable across the systems people actually use, the second condition is the better one, and a dashboard that only reports total requests would rank it lower.

We now treat distinct crawler count as a first-class metric alongside volume. It is a rough proxy, not a proof, but it maps to the thing we actually care about: how many independent systems have decided this entity is worth reading.

The three properties that seem to matter

Nothing above tells you how a model decides to cite you. For that we are reasoning from how these systems are built, from vendor documentation, and from what we observe when we change something and watch the crawl behavior respond. Treat this section as informed inference rather than measurement.

Consistency

If your name, description, and stated relationships differ between your homepage, your structured data, and your third-party profiles, a system has to resolve a conflict before it can use you. Resolving conflicts costs confidence. A model that is uncertain which of two descriptions is correct is a model that hedges or omits.

This is the cheapest thing on the list to fix and the most commonly broken. Most organizations have never compared their schema markup to their About page. When we do that comparison for a new client, we find a contradiction more often than not.

Corroboration

A claim that appears only on your own domain is a claim with one source. The same claim echoed on an independent property, a directory, a partner site, or a profile is a claim with several. Models weight corroborated facts more heavily for the same reason people do.

This is why we design portfolios rather than single sites, and why the cross-references between properties are declared in structured data rather than left as decorative footer links.

Structure

Prose requires interpretation. Structured data does not. When a page states in machine-readable form that this organization has this founder, offers these services, published this article on this date, and relates to these other entities, a model does not have to infer any of it.

We rebuilt this site's own structured data in September 2026 and the before-and-after is instructive. Every page previously emitted four nodes: an Organization, a WebSite, a Person, and a generic WebPage that said almost nothing. It now emits an average of 12.5 nodes per page across all 57 URLs, including a federation network, a machine-readable data catalog, and page-specific entities. Same content. Same words. Substantially more of it stated rather than implied.

Why this favors structure over volume

Publishing more content does not repair an entity that is poorly defined. If a system cannot tell who you are, forty more articles give it forty more chances to be uncertain about the same thing.

The inverse is also true and more encouraging. A small, clearly defined, well-structured property can be read confidently by every AI system that finds it. Our own site took 1,873 AI crawler requests from 10 distinct crawlers in the window, which is modest. It is also a site with 57 pages. The ratio is what matters.

What we would actually do first

In order, because the order matters:

  1. Compare your descriptions. Put your homepage copy, your schema markup, and your two most prominent third-party profiles side by side. Fix every place they disagree. This costs an afternoon.
  2. Emit the facts you are currently implying. Who founded the organization, where it operates, what it offers, which entities it relates to. If a human can read it off your About page, a machine should be able to read it out of your markup.
  3. Check your logs. Not your analytics ... your raw access logs. Analytics platforms filter bots out by design, which means the entire audience described in this article is invisible in the tool most organizations use to decide what is working.
  4. Measure breadth, not just volume. Count distinct AI crawlers, and watch whether that number grows when you improve your structure.

What we still cannot see

We can measure who fetched what and when. We cannot measure what any of it produced downstream, and anyone claiming a clean attribution model between AI crawl activity and AI citations is ahead of the available evidence.

What we can say is narrower and still useful: a rapidly growing number of independent machine readers are fetching pages at a rate that dwarfs human visits on the same properties, they are readable systems that respond to explicit structure, and almost nobody is looking at the logs where this is visible.

That gap is the opportunity. It will not stay open indefinitely.

See this applied in the framework.

Explore the Signal Architecture Framework to see how this concept fits into a complete system.

Explore Our Framework →