Entity Engineering Research

From Keywords to Knowledge Graphs: The New Foundation of Visibility

May 27, 2026 · 6 min read

Keyword density tells a machine which words are on a page. Entity relationships tell it what the page means and how it connects to everything else. The second one is what search and AI systems now organize around, and the gap between the two shows up in a number most people never look at: the distance between how often you are seen and how often you are chosen.

The number that started this

Across 137 properties over 90 days, our warehouse recorded 1,895,698 search impressions and 14,110 clicks. That is a click-through rate of about 0.74 percent.

The obvious reading is that the titles need work. Sometimes that is true. But a portfolio-wide rate that low, sustained across very different subject areas, usually means something less flattering and more useful: a large share of those impressions are appearances where the system was not confident enough to place us anywhere a human would actually click.

Being matched as a string of text is cheap. Being selected as the answer requires the system to know what you are. Those are different achievements, and only one of them shows up in a keyword report.

What a knowledge graph actually is

Strip away the branding and a knowledge graph is a structured map of things and the relationships between them. People, organizations, products, places, concepts. Each one is a node. Each relationship is an edge. The organization node for a company connects to its founder, its services, its location, its parent company, its published work.

The practical consequence is worth stating plainly. When a system has a solid node for your organization, it does not have to work out who you are every time your name appears. It already knows, and it can spend its effort on the actual question instead.

When it does not have that node, every mention of you is a fresh problem to solve. Sometimes it solves it correctly. Sometimes it merges you with a similarly named company. Often it just goes with something it is more certain about.

Why keyword density stopped carrying the weight

Keyword targeting worked because early ranking systems had no better way to determine what a page was about. Counting words was a reasonable proxy for meaning when meaning could not be read directly.

That constraint is gone. Modern systems parse structured data, resolve named entities, and consult existing graph entries. A page that repeats a phrase fifty times and a page that states in machine-readable form who provides that service, what it includes, and which other entities it relates to are no longer competing on the same axis. The first is asserting relevance. The second is supplying facts.

None of this makes keywords useless. People still type words, and those words still have to appear on the page in a form a human recognizes. The change is in what the words are doing. They used to be the evidence. Now they are the entry point, and the evidence lives in the structure underneath.

A worked example, on ourselves

We rebuilt this site's structured data in September 2026, and the before-and-after is a clean illustration of the difference.

Before: every page emitted four nodes. An Organization, a WebSite, a Person for the founder, and a generic WebPage carrying a name, a description, and a URL. Technically valid. Almost entirely uninformative. Nothing stated what any given page was actually about, how it related to anything else, or what made it different from the other 56 pages emitting the same four nodes.

After: an average of 12.5 nodes per page across all 57 URLs. The organization now declares its founder, its address, its subject expertise, its service catalog, and its verified profiles elsewhere. Research articles emit a WebPage and an Article as separate linked nodes, with measured word counts, sections, and the internal entities each one references. Glossary terms emit DefinedTerm nodes inside a declared term set. The methodology page emits its five stages as ordered steps. Case studies declare the actual site each one is about.

Not one word of visible copy changed in that rebuild. What changed was how much of the meaning was stated rather than left to be inferred.

We are being careful about what we claim here. We are not going to tell you that number produced a ranking increase, because the change is recent and we do not have the follow-up window yet. What we can say is that a system reading the site now has substantially less to guess about, and the guessing is where visibility gets lost.

Three relationships worth stating before anything else

If you do nothing else structurally, state these three.

Who you are, once

An organization node with a stable identifier, a name, a description, and the profiles that corroborate it elsewhere. This is the anchor everything else attaches to. If you have four different descriptions of your company in circulation, none of them is doing this job.

Who is behind it

A person node for the founder or principal, connected to the organization, with links to the places that same person is described elsewhere. People are among the best-established entities in existing graphs, which means a clearly defined person can pull an ambiguous organization into focus.

What you are about

Explicit subject declarations. Not a keyword list. The actual concepts your organization has standing in, stated as entities rather than as prose adjectives. This is the connective tissue that lets a system place you in a topic without inferring it from word frequency.

You do not have to wait to be defined

A common assumption is that your entry in a knowledge graph is something a third party grants you. Partly true, and less limiting than it sounds. The systems building those graphs read your structured data, your consistency across properties, and your corroborating mentions. You are supplying most of the raw material either way. The only question is whether you are supplying it deliberately.

The organizations that get defined accurately are usually not the largest ones. They are the ones that said the same thing, in the same form, in enough places, early enough that there was never much ambiguity to resolve.

What to measure instead of density

Keyword density has a comforting property: it is easy to count. Entity clarity is harder to quantify, but these are workable proxies:

  • Impression-to-click distance. A large gap suggests you are being matched but not selected.
  • Distinct AI crawlers. How many independent systems fetch your site, not how often one does. Ours ranges from single digits to 23 across the portfolio.
  • Description drift. Count how many different descriptions of your organization exist across your own properties and profiles. The target is one.
  • Stated versus implied facts. For each claim on your About page, check whether a machine could read it out of your markup. Most cannot.

None of these is a ranking metric. All of them move earlier than rankings do, and all of them describe the thing that actually determines whether a system can use you with confidence.

See this applied in the framework.

Explore the Signal Architecture Framework to see how this concept fits into a complete system.

Explore Our Framework →