AgentShield AI Defense

Observatory

Not who we are — what we watch, and what counts as having watched it.

This is an independent observatory for the behaviour of automated clients on the web. It runs continuously on one domain, records every request it receives, and publishes what the record shows. It sells nothing, and it has no customers whose traffic it reports on. The only site under observation is this one.

What is observed

Four populations arrive at any public site, and they behave differently enough that lumping them together destroys the signal.

Corpus crawlers

Bulk collectors building training or index corpora. They arrive unprompted, on their own schedule, and sweep broadly. What they take may shape a model long after the visit — which is why the interval between publishing something and seeing it surface is worth measuring at all.

Retrieval agents

Maintainers of the indexes that AI products query while answering. They resemble traditional search crawlers, but the index they build feeds generated prose rather than a list of links.

User-triggered fetchers

Clients that fetch a page because a person asked an assistant something, right now, that required reading it. They arrive one URL at a time and track human activity patterns.

Autonomous and scripted clients

Everything else that is not a person with a browser: scanners, scrapers, headless automation, and agents acting on a goal. They arrived here within minutes of the domain going live, before anything linked to it.

The cycle

The observation cycle Observe, document, archive, analyze, publish, and return to observe. Each pass adds to a history that gives later observations their context. Observe requests as they arrive Document record, never edit Archive history accumulates Analyze versioned reading Publish findings in public History creates context
Each pass through the loop adds to a record that makes the next pass sharper. A single observation says little; the same observation against two years of history is evidence.

Declaration is not identity

Every automated client announces itself in a User-Agent string, and nothing verifies that announcement. Anyone can send any string. So the declaration is stored as a claim, and the behaviour is recorded separately — and the distance between the two is the measurement.

How a declared identity becomes evidence A client declares an identity in its user agent string. Separately, its behaviour is observed: whether it fetched robots.txt, whether it then took a disallowed path, whether it executed JavaScript. The gap between what was declared and what was done is the measurable signal. Declared identity User-Agent string a claim, never verified Observed behaviour what the client actually did recorded, re-checkable Compared declaration vs record Fetched robots.txt, then took a disallowed path Claimed a browser, never executed JavaScript Declared one crawler, behaved like another
Promise-keeping is measurable. A crawler that reads the rules and then ignores them has produced evidence about itself that no self-description can override — which is what makes behaviour, rather than declaration, the basis for trust.

This is where the observatory's thesis and its instrument meet. Trust is earned through observable, repeated behavioural evidence — and the first such evidence available to any website is whether a client that read the rules then followed them. A crawler that fetches robots.txt and afterwards takes a disallowed path has said something about itself that no self-description can retract.

What is refused

Admissible and inadmissible evidence Observed requests, request headers, timing and the appearance of a published canary string are admissible evidence. A model's own account of what it knows, and an inference stored as though it were an observation, are refused. ADMISSIBLE Observed request method, path, status, timing Request headers declared identity, stored as a claim Canary appearance coined string observed in the wild all re-checkable against the stored record REFUSED Model self-report the system under test describing itself Inference stored as fact “this was a bot” written into the record Unversioned conclusion cannot be re-derived later no layer may validate itself with its own output
The exclusion on the right is the expensive one. Asking a model what it knows about a site is the common way to measure AI visibility, and it is the system under test giving evidence about itself.

The refusal on the right is not fastidiousness. It removes the standard method for measuring AI visibility — prompting a model about a brand and recording the reply — because that is the system under test giving evidence about itself. What remains is slower and narrower, and it can be checked.

Current state

Observation began 2026-07-25. The record is small, and nothing here is presented as a finding about crawler behaviour in general yet. The lab shows the live counts, including how little there is; the methodology states what would make the figures wrong.

Reality marker for this page: asd-dofotu-pitega-5b116f · published 2026-07-25 11:15:37 UTC

This string is coined and appears nowhere else. If it later surfaces in a language model's output, that is observed evidence this page was ingested. What this is.