AgentShield AI Defense

Findings

What the record has established, dated and frozen. Each entry states what was observed, what it means, and the limits that stop it meaning more.

26 published, of which 18 were detected and written automatically. Live counters are on the lab page; this is the archive of what has been concluded from them.

Detected automatically by automated_enumeration · 2026-07-30 05:13:43 UTC

One address requested 38 distinct paths within 7 seconds

Detected automatically by automated_enumeration · 2026-07-30 03:43:43 UTC

One address requested 28 distinct paths within 33 seconds

Detected automatically by automated_enumeration · 2026-07-30 03:43:43 UTC

One address requested 28 distinct paths within 29 seconds

Detected automatically by automated_enumeration · 2026-07-30 03:43:43 UTC

One address requested 28 distinct paths within 31 seconds

Detected automatically by automated_enumeration · 2026-07-30 03:13:43 UTC

One address requested 165 distinct paths within 17 seconds

Detected automatically by automated_enumeration · 2026-07-30 01:43:43 UTC

One address requested 31 distinct paths within 6 seconds

Detected automatically by automated_enumeration · 2026-07-29 23:58:43 UTC

One address requested 47 distinct paths within 5 seconds

Detected automatically by format_preference · 2026-07-29 23:58:43 UTC

Which content formats are actually fetched

Detected automatically by automated_enumeration · 2026-07-29 16:43:42 UTC

One address requested 22 distinct paths within 3 seconds

Detected automatically by automated_enumeration · 2026-07-29 11:43:42 UTC

One address requested 84 distinct paths within 23 seconds

Detected automatically by distributed_crawl · 2026-07-28 12:30:15 UTC

One user agent, 13 addresses, 13 paths — a retrieval spread thin enough to look like nothing

Written by a person from the record · 2026-07-28 09:55:00 UTC

Eight requests from OpenAI's two automated crawlers asked only for the rules and the map

Written by a person from the record · 2026-07-28 07:40:00 UTC

Fourteen sites served the same 501 bytes of robots.txt, and the bytes name who wrote them

Written by a person from the record · 2026-07-27 12:00:00 UTC

The most thorough reader of this site declares itself a 2019 handset, from ten countries at once

Detected automatically by automated_enumeration · 2026-07-27 10:17:01 UTC

One address requested 34 distinct paths within 3 seconds

Detected automatically by identity_rotation · 2026-07-27 00:03:54 UTC

One address presented crawler identities belonging to 10 different companies

Detected automatically by distributed_crawl · 2026-07-27 00:02:09 UTC

One user agent, 29 addresses, 20 paths — a retrieval spread thin enough to look like nothing

Detected automatically by automated_enumeration · 2026-07-26 22:14:13 UTC

One address requested 87 distinct paths within 6 seconds

Detected automatically by arrival_host · 2026-07-26 07:32:12 UTC

Machines are fetching this site through a second hostname of ours that we never published

Detected automatically by ai_agent_arrival · 2026-07-25 20:42:34 UTC

We asked Claude-User to read this page, and it did

Detected automatically by distributed_crawl · 2026-07-25 18:46:38 UTC

One user agent, 206 addresses, 58 paths — a retrieval spread thin enough to look like nothing

Written by a person from the record · 2026-07-25 12:00:00 UTC

A new domain with no inbound links received its first automated probe in under two minutes

Written by a person from the record · 2026-07-25 12:00:00 UTC

An assistant given a direct URL substituted a search, and reported a competitor instead

Written by a person from the record · 2026-07-25 12:00:00 UTC

Anatomy of a user-triggered fetch: one page, no robots.txt, no JavaScript

What has kept being true

A finding is dated: it reports what the record showed at a moment. These are the statements that have gone on being true since, and each is checked against the record every time this page is requested. One that stops holding leaves this list by itself.

Every statement below is negative — no such thing has happened — which is what early knowledge looks like and is also the only kind a single observation can destroy. So each names the one thing that would end it, and prints how many chances the record has already given it. 2 of 6 have had enough chances to be worth stating.

StatementWhat would end itChances it has hadState
The bytes a stranger receives for a watched file have been the bytes this server sent.one sweep where the origin held still and the edge delivered something else385 bracketed sweepsholding
No request corroborated as one of OpenAI's automated crawlers has asked for a page of this site's content.one such request to a path that is not robots.txt, sitemap.xml, llms.txt or favicon.ico72 corroborated requestsended
No external client has executed the script on the probe page.one beacon request from a client that is not the operator10 views of the probe page — needs 200not yet enough to say
Nothing published here has been observed in a language model's output.one published marker appearing in a model's answer76 markers publishedholding
No client has fetched a path this site's robots.txt asks clients to leave alone.one request to a disallowed path3201 external requestsended
Among surveyed sites whose robots.txt carries a CDN-inserted block, none names an AI crawler in its own text that the block refuses.one surveyed file whose own section allows a crawler the inserted block disallows14 surveyed files carrying a block — needs 100not yet enough to say

2 statements have ended. The record now holds a counterexample, and the sentence is kept here in that state rather than deleted — a claim that was made and then failed is part of what this instrument has done.

The minimums are not statistical power calculations. They are the point below which repeating the sentence would mislead on its face — and they are in source, where they can be argued with, rather than in a judgement that moves when an answer is wanted.

Withdrawn and rejected

11 conclusions that this pipeline drew and then stopped publishing. 9 of them were live on this site before being taken down, which is the reason this table exists: a page that returned 200 and now returns 404 owes an explanation to whoever read it.

What is listed is the subject and the reason. The withdrawn text is not republished. Most of these findings named a company as a visitor to this site and were wrong about it; reprinting that sentence under a "rejected" heading would put the false claim back on an indexable page, where a crawler reads the sentence and not the label. That would repeat the mistake in order to disclose it.

SubjectWas publishedTaken downWhy it is not published
F-0072026-07-28 11:30:14 UTC2026-07-28 11:36:46 UTCWithdrawn by the operator on the day of publication. The finding was accurate and its central question was unresolved: no request arrived from one assistant, and why was not established. Leading with an open question while the completed observation — a model quoting a figure the record holds to the second — sat beneath it was the wrong order. Held until the network provider event log for the same window has been read.
Amazonbot2026-07-27 14:43:29 UTC2026-07-27 15:25:21 UTCHeadline made a named crawler the subject of conduct — Article IX. The counts were correct; the sentence was not.
Mozilla/5.0 (compatible; SemrushBot/7~bl; …never livenot recordedCounts verified and correct; the subject is not. The shape does not contradict the declaration: 10 addresses confined to two of the vendor's own blocks (85.208.96.x, 185.191.171.x), a single country, and robots.txt fetched at the start of each of the three sessions and obeyed. A declared crawler spread across its own address ranges is what a crawl farm looks like, not a finding. distributed_crawl is only meaningful where the shape contradicts the declared identity, as with the 147-address consumer-phone user agent the detector was built for. Corrected 2026-07-27: language only. Recorded at rejection in Turkish; the reasoning is unchanged. Original: "Sayilar dogrulandi. Sekil gercek ama beyanla celismiyor: tek bir ulkede, saticinin kendi iki adres blogunda, her oturumda robots.txt okuyarak. Beyan edilen kimlik bu sekli zaten ongoruyor. Dagitilmis tarama ancak beyanla celistiginde bulgudur."
ClaudeBot2026-07-26 22:14:13 UTCnot recordedAll 10 requests came from one address presenting ten companies' crawler identities. No client operated by Anthropic was involved. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding.
PerplexityBot2026-07-26 22:14:13 UTCnot recordedAll 9 requests came from one address presenting ten companies' crawler identities. No client operated by Perplexity was involved. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding.
CCBot2026-07-26 22:14:13 UTCnot recordedAll 9 requests came from one address presenting ten companies' crawler identities. No client operated by Common Crawl was involved. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding.
GPTBot2026-07-26 17:43:49 UTCnot recordedSubject conflates 1 vendor-verified request with 4 from an impostor address; the detector could not separate them. The verified request was genuine. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding.
2a00:1d34:4896:b600:b997:78fb:eb0d:b131|cu…2026-07-25 15:51:29 UTCnot recordedown test traffic during setup
OAI-SearchBot2026-07-25 12:18:14 UTCnot recordedSubject conflates 4 vendor-verified requests with 11 from an impostor address; the detector could not separate them. The verified requests were genuine. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding.
ChatGPT-User2026-07-25 12:18:14 UTCnot recordedSubject conflates 2 vendor-verified requests with 6 from an impostor address; the detector could not separate them. The verified requests were genuine. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding.
2a00:1d34:4896:b600:b997:78fb:eb0d:b131never livenot recordedown test traffic during setup

Where the takedown column says not recorded, the instant genuinely is not known: the column that holds it was added on 2026-07-27, after these were withdrawn, and it was not backfilled. Writing today's date into those rows would have invented a fact to fill a gap.

How a finding gets published

Rules run continuously over the stored observations. When one matches, it produces a candidate carrying the exact figures that support it, each paired with the query that produced it. A template turns the candidate into prose without touching the figures. A verifier then recomputes every figure against the record, and a single mismatch discards the draft entirely.

Findings that restate a count publish themselves once verified. Findings that say something unflattering about a named actor — a compliance failure, a contradicted identity — wait for a person to read them first, because a wrong one there would cost more than a missing one.

Standard applied

A finding is published when the record supports a specific statement, not when it supports a general one. Under the constitution, every figure must be recomputable from stored observations, sample size must be stated, and absence of evidence is reported as absence of evidence. A finding of n = 1 is published as a finding of n = 1.

Reality marker for this page: asd-gedutoru-zogomira-1cce6e · published 2026-07-25 11:40:23 UTC

This string is coined and appears nowhere else. If it later surfaces in a language model's output, that is observed evidence this page was ingested. What this is.