AgentShield AI Defense

Why not just ask the model what it knows about your site?

Because a model describing its own knowledge is the system under test giving evidence about itself. It can assert familiarity with pages it never read and deny ones it did, and nothing in the answer distinguishes the two. No layer may validate itself using its own output as independent evidence.

Founding principle of this observatory.

It is the natural thing to try, and it is the method underneath most tools that claim to measure AI visibility: prompt an assistant about a brand, record the reply, chart the replies over time.

The structural problem

The model is the thing being measured. Asking it to report on its own knowledge makes it both instrument and subject, and there is nothing in the output that separates recall from plausible construction. A model can produce a fluent description of a site it has never encountered, because producing fluent descriptions is what it does.

Repeating the prompt does not help. Consistency across attempts measures the stability of the generation, not the truth of it.

The rule

No layer may validate itself using its own output as independent evidence.

Stated generally it sounds like philosophy. Applied, it removes a whole category of product. It also removes several tempting shortcuts from this site: no classification may be stored as though it were an observation, and no figure may be published that cannot be recomputed from the raw record.

What replaces it

Evidence from outside the model. Server-side records of what was actually fetched, and coined strings published at known instants whose appearance is observable independently of anything a model says about itself. Slower, narrower, checkable.

The trade is deliberate. A measurement you can verify and a measurement that sounds impressive are usually not the same measurement.

Related

All questions

Reality marker for this page: asd-nuripozo-lukuti-8ca7bd · published 2026-07-25 11:15:37 UTC

This string is coined and appears nowhere else. If it later surfaces in a language model's output, that is observed evidence this page was ingested. What this is.