Methodology · Observe · Understand · Govern

How we observe AI agents reading your site.

Published May 2026
·
aater.ai

Every AI system that reads the web makes the same three structural checks: can it reach your content, can it read it, and is it informational or commercial. Then — the part that actually matters — its agents either visit and cite you, or they don't. Aater observes both: the structural gates (the free baseline) and the live agent behaviour.

The uncomfortable reality

Most sites fail before the first gate. They are not blocked, not penalised — simply invisible. AI systems move past them without recording a signal. The content exists. The classification does not.

New to the classification vocabulary? Read the Participation Lexicon first →

Why a standard matters

robots.txt told crawlers where to go.
This tells AI what it can use.

robots.txt established the protocol for crawler access. It is a declaration, not a measurement. It tells AI systems what they may crawl — not whether the content they find is actually usable, citable, or trustworthy.

Aater fills that gap with three measurable gates — reachability, legibility, and content type. Checkable facts, applied identically every time. The three gates are separate facts, not a score. Structure tells you what AI can reach and read; whether AI agents then visit and cite you is observed separately via Pulse.

Patent Pending (IN 202621062543)

robots.txt
Declares crawl permission
Does not measure content usability
sitemap.xml
Declares URL structure
Does not measure semantic value
schema.org
Declares content type
Does not measure authority or density
The three gates
Measures AI usability end-to-end
Classification Pipeline

Three gates. One derivation.

Gates are evaluated in strict sequence. Failure at any gate terminates evaluation and resolves the structural gate result immediately. No gate can be compensated for by performance at another.

Unfamiliar with gate terminology? See precise definitions in the Lexicon →

Gate 1
Reachability
Can AI reach it?
What is measured
  • DNS resolution and HTTP response
  • robots.txt policy — per-agent intersection
  • Content delivery completeness
  • Empty shell detection (200 OK ≠ content)
Outcomes
PassGate 2 →
DegradedGate 2 (constrained) →
BlockedBlocked — stop
Pass →
Gate 2
Legibility
Can AI read what it finds?
What is measured
  • Heading structure — hierarchy and coherence
  • Named entity density — distinct, non-UI tokens
  • Specific claim surface — verifiable assertions
  • Semantic depth vs render gap estimate
Outcomes
StructuredGate 3 →
PresentGate 3 →
IllegibleIllegible — stop
Pass →
Gate 3
Content Type
Is it informational or commercial?
What is measured
  • Declared schema — JSON-LD @type, og:type (publisher's own answer)
  • URL pattern — /blog/ /guide/ /vs/ vs /pricing/ /features/
  • Page structure — guide/reference vs product/CTA
  • Heading + CTA language cues
Outcomes
Informationalthe content type AI agents engage with differently
Commercialmentioned, rarely cited (the Tool Trap)
Derivation complete
Structural gates resolved
Apply the standard

Submit domain for derivation.

The same three gates. Applied in sequence. No human judgement enters the derivation.

Independent validation

What Google confirms — and what no vendor can.

Google's official AI optimisation guide independently describes the same structural floorour first two gates measure: be reachable, be legible. That floor is physics — if an AI agent can't reach or parse your content, no AI system can use it, whatever else is true.

But that guide describes Google's AI surface. No vendor publishes one that also covers ChatGPT, Claude, and Perplexity — they retrieve differently, and the signals above the floor differ by system. So the gates are a diagnostic of structural accessibility, not a citation predictor and not a universal ranking model. The only way to know how each agent treats your site is to observe it — which is what Pulse does, per model.

Gate 1 · Reachability
"Ensure your content is crawlable… Follow JavaScript SEO best practices."
Gate 2 · Legibility
"Use semantic HTML."
Gate 3 · Content Type
"non-commodity content that's helpful, reliable, and people-first."
Google also dismisses the shortcuts we deliberately do not gate on: “…Google Search itself doesn't use them” (llms.txt / AI text files), and “Structured data isn't required for generative AI search.” The gates measure what structurally determines access, not which signals AI marketing recommends. Our full analysis →
What the gates resolve to

Three checkable facts — not a single score.

We deliberately do not collapse the gates into one synthesized label. They are three separate facts, each checkable, each independent — that is the design. Structure and observed behaviour are different instruments: the gates measure what AI can reach and read; Pulse observes what AI agents actually do.

Reachability

Can automated systems access your content at all — HTTP response, robots.txt for the retrieval fetchers that drive citation, content completeness.

Legibility

Can they extract meaningful structure — heading hierarchy, named entities, a verifiable claim surface.

Content Type

Informational/guide vs commercial/product — read from declared schema, then URL and structure. Classifies which of your pages AI agents engage with differently (observed in Pulse); a structural observation, not a guarantee.

Structure is the free baseline — necessary, not sufficient. What AI agents actually do — visit, cite — is observed via Pulse →

Agent identity is verified where possible: a crawler is marked Verifiedonly when its connecting IP falls inside the operator's own published address ranges. Bots shown as Unverified may be legitimate crawlers whose operators do not publish IP ranges — absence of verification is not evidence of spoofing.

What the tamper-evident record seals — and what it deliberately excludes. The sealed ledger preserves observed agent traffic: verified bots, unverifiable bots, and spoofing attempts — a request impersonating an AI crawler is itself a fact worth preserving. It excludes operational telemetry(uptime monitors, health checks, internal probes), which is instrumentation of our own systems, not an observation of yours, and never enters the ledger. Derived figures (counts, shares) are computed from the ledger, never part of it. The canonical-event filter went live in June 2026; sealed records from before that date may include monitoring artifacts, so an evidence chain's authoritative window begins at the filter date, not at install.

Classification is not ranking.

The structural gate result is not a score. There is no weighted formula, no fuzzy weighting. Each gate either passes or fails. The derivation is deterministic — the same domain, the same signals, produces the same result every time.

Invisible failure is the norm.

Most sites that fail do not receive an error. They receive silence. AI systems move on without recording the visit. The site owner has no signal. This is the gap the standard exists to close.

State changes when you change.

Classification reflects current structural reality. Fix a robots.txt entry and reachability passes. Add structured data and your content becomes more legible. The standard measures the present, not history.

Apply the standard

Classification runs the same pipeline on every domain.

The three gates are applied in sequence. No domain receives special treatment. No human judgement enters the derivation. Enter your domain to see which gate your site passes and where the constraint sits.

Which gate your site fails (if any)
The specific signal that produced the outcome
Your structural gate result in the global ledger
Live Classification

Where does your domain stand?

Run a classification against the standard. See which gate your site fails — and why.

Enter your domain to see where your site fails this pipeline.

Versioning

A standard requires stability.

Every classification is tagged with the engine version that produced it. When the standard evolves, prior classifications remain interpretable. Agencies and publishers can cite a specific version with confidence that the classification is reproducible.

the structural gatesCurrentStructural-gate engine · Gate-based derivation
System behaviour

Structural gate results continuously shift as domains change structure and accessibility.

The observatory measures these transitions across the public web. Classification is not a snapshot. It is a continuous derivation — re-evaluated whenever structural signals change.

Beyond classification

Knowing your gate results is the beginning.

The audit tells you where you stand. Pulse shows what AI systems are actually doing on your site — which agents are visiting, at what frequency, and to what depth.