Findings
What the record has established, dated and frozen. Each entry states what was observed, what it means, and the limits that stop it meaning more.
26 published, of which 18 were detected and written automatically. Live counters are on the lab page; this is the archive of what has been concluded from them.
Detected automatically by automated_enumeration · 2026-07-30 05:13:43 UTC
One address requested 38 distinct paths within 7 seconds
Detected automatically by automated_enumeration · 2026-07-30 03:43:43 UTC
One address requested 28 distinct paths within 33 seconds
Detected automatically by automated_enumeration · 2026-07-30 03:43:43 UTC
One address requested 28 distinct paths within 29 seconds
Detected automatically by automated_enumeration · 2026-07-30 03:43:43 UTC
One address requested 28 distinct paths within 31 seconds
Detected automatically by automated_enumeration · 2026-07-30 03:13:43 UTC
One address requested 165 distinct paths within 17 seconds
Detected automatically by automated_enumeration · 2026-07-30 01:43:43 UTC
One address requested 31 distinct paths within 6 seconds
Detected automatically by automated_enumeration · 2026-07-29 23:58:43 UTC
One address requested 47 distinct paths within 5 seconds
Detected automatically by format_preference · 2026-07-29 23:58:43 UTC
Which content formats are actually fetched
Detected automatically by automated_enumeration · 2026-07-29 16:43:42 UTC
One address requested 22 distinct paths within 3 seconds
Detected automatically by automated_enumeration · 2026-07-29 11:43:42 UTC
One address requested 84 distinct paths within 23 seconds
Detected automatically by distributed_crawl · 2026-07-28 12:30:15 UTC
One user agent, 13 addresses, 13 paths — a retrieval spread thin enough to look like nothing
Written by a person from the record · 2026-07-28 12:16:53 UTC
893 requests in three minutes became 77% of a day's traffic, and obtained four files anyone can read
Written by a person from the record · 2026-07-28 12:03:12 UTC
Every request declaring an OpenAI identity was answered 200, and none has declared the browsing agent since 25 July
Written by a person from the record · 2026-07-28 09:55:00 UTC
Eight requests from OpenAI's two automated crawlers asked only for the rules and the map
Written by a person from the record · 2026-07-28 07:40:00 UTC
Fourteen sites served the same 501 bytes of robots.txt, and the bytes name who wrote them
Written by a person from the record · 2026-07-27 12:00:00 UTC
The most thorough reader of this site declares itself a 2019 handset, from ten countries at once
Detected automatically by automated_enumeration · 2026-07-27 10:17:01 UTC
One address requested 34 distinct paths within 3 seconds
Detected automatically by identity_rotation · 2026-07-27 00:03:54 UTC
One address presented crawler identities belonging to 10 different companies
Detected automatically by distributed_crawl · 2026-07-27 00:02:09 UTC
One user agent, 29 addresses, 20 paths — a retrieval spread thin enough to look like nothing
Detected automatically by automated_enumeration · 2026-07-26 22:14:13 UTC
One address requested 87 distinct paths within 6 seconds
Detected automatically by arrival_host · 2026-07-26 07:32:12 UTC
Machines are fetching this site through a second hostname of ours that we never published
Detected automatically by ai_agent_arrival · 2026-07-25 20:42:34 UTC
We asked Claude-User to read this page, and it did
Detected automatically by distributed_crawl · 2026-07-25 18:46:38 UTC
One user agent, 206 addresses, 58 paths — a retrieval spread thin enough to look like nothing
Written by a person from the record · 2026-07-25 12:00:00 UTC
A new domain with no inbound links received its first automated probe in under two minutes
Written by a person from the record · 2026-07-25 12:00:00 UTC
An assistant given a direct URL substituted a search, and reported a competitor instead
Written by a person from the record · 2026-07-25 12:00:00 UTC
Anatomy of a user-triggered fetch: one page, no robots.txt, no JavaScript
What has kept being true
A finding is dated: it reports what the record showed at a moment. These are the statements that have gone on being true since, and each is checked against the record every time this page is requested. One that stops holding leaves this list by itself.
Every statement below is negative — no such thing has happened — which is what early knowledge looks like and is also the only kind a single observation can destroy. So each names the one thing that would end it, and prints how many chances the record has already given it. 2 of 6 have had enough chances to be worth stating.
| Statement | What would end it | Chances it has had | State |
|---|---|---|---|
| The bytes a stranger receives for a watched file have been the bytes this server sent. | one sweep where the origin held still and the edge delivered something else | 385 bracketed sweeps | holding |
| No request corroborated as one of OpenAI's automated crawlers has asked for a page of this site's content. | one such request to a path that is not robots.txt, sitemap.xml, llms.txt or favicon.ico | 72 corroborated requests | ended |
| No external client has executed the script on the probe page. | one beacon request from a client that is not the operator | 10 views of the probe page — needs 200 | not yet enough to say |
| Nothing published here has been observed in a language model's output. | one published marker appearing in a model's answer | 76 markers published | holding |
| No client has fetched a path this site's robots.txt asks clients to leave alone. | one request to a disallowed path | 3201 external requests | ended |
| Among surveyed sites whose robots.txt carries a CDN-inserted block, none names an AI crawler in its own text that the block refuses. | one surveyed file whose own section allows a crawler the inserted block disallows | 14 surveyed files carrying a block — needs 100 | not yet enough to say |
2 statements have ended. The record now holds a counterexample, and the sentence is kept here in that state rather than deleted — a claim that was made and then failed is part of what this instrument has done.
The minimums are not statistical power calculations. They are the point below which repeating the sentence would mislead on its face — and they are in source, where they can be argued with, rather than in a judgement that moves when an answer is wanted.
Withdrawn and rejected
11 conclusions that this pipeline drew and then stopped publishing. 9 of them were live on this site before being taken down, which is the reason this table exists: a page that returned 200 and now returns 404 owes an explanation to whoever read it.
What is listed is the subject and the reason. The withdrawn text is not republished. Most of these findings named a company as a visitor to this site and were wrong about it; reprinting that sentence under a "rejected" heading would put the false claim back on an indexable page, where a crawler reads the sentence and not the label. That would repeat the mistake in order to disclose it.
| Subject | Was published | Taken down | Why it is not published |
|---|---|---|---|
| F-007 | 2026-07-28 11:30:14 UTC | 2026-07-28 11:36:46 UTC | Withdrawn by the operator on the day of publication. The finding was accurate and its central question was unresolved: no request arrived from one assistant, and why was not established. Leading with an open question while the completed observation — a model quoting a figure the record holds to the second — sat beneath it was the wrong order. Held until the network provider event log for the same window has been read. |
| Amazonbot | 2026-07-27 14:43:29 UTC | 2026-07-27 15:25:21 UTC | Headline made a named crawler the subject of conduct — Article IX. The counts were correct; the sentence was not. |
| Mozilla/5.0 (compatible; SemrushBot/7~bl; … | never live | not recorded | Counts verified and correct; the subject is not. The shape does not contradict the declaration: 10 addresses confined to two of the vendor's own blocks (85.208.96.x, 185.191.171.x), a single country, and robots.txt fetched at the start of each of the three sessions and obeyed. A declared crawler spread across its own address ranges is what a crawl farm looks like, not a finding. distributed_crawl is only meaningful where the shape contradicts the declared identity, as with the 147-address consumer-phone user agent the detector was built for. Corrected 2026-07-27: language only. Recorded at rejection in Turkish; the reasoning is unchanged. Original: "Sayilar dogrulandi. Sekil gercek ama beyanla celismiyor: tek bir ulkede, saticinin kendi iki adres blogunda, her oturumda robots.txt okuyarak. Beyan edilen kimlik bu sekli zaten ongoruyor. Dagitilmis tarama ancak beyanla celistiginde bulgudur." |
| ClaudeBot | 2026-07-26 22:14:13 UTC | not recorded | All 10 requests came from one address presenting ten companies' crawler identities. No client operated by Anthropic was involved. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding. |
| PerplexityBot | 2026-07-26 22:14:13 UTC | not recorded | All 9 requests came from one address presenting ten companies' crawler identities. No client operated by Perplexity was involved. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding. |
| CCBot | 2026-07-26 22:14:13 UTC | not recorded | All 9 requests came from one address presenting ten companies' crawler identities. No client operated by Common Crawl was involved. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding. |
| GPTBot | 2026-07-26 17:43:49 UTC | not recorded | Subject conflates 1 vendor-verified request with 4 from an impostor address; the detector could not separate them. The verified request was genuine. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding. |
| 2a00:1d34:4896:b600:b997:78fb:eb0d:b131|cu… | 2026-07-25 15:51:29 UTC | not recorded | own test traffic during setup |
| OAI-SearchBot | 2026-07-25 12:18:14 UTC | not recorded | Subject conflates 4 vendor-verified requests with 11 from an impostor address; the detector could not separate them. The verified requests were genuine. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding. |
| ChatGPT-User | 2026-07-25 12:18:14 UTC | not recorded | Subject conflates 2 vendor-verified requests with 6 from an impostor address; the detector could not separate them. The verified requests were genuine. Corrected 2026-07-27. The reason first recorded here was 'own test traffic', which was the console's pre-filled default and was not true of this finding. |
| 2a00:1d34:4896:b600:b997:78fb:eb0d:b131 | never live | not recorded | own test traffic during setup |
Where the takedown column says not recorded, the instant genuinely is not known: the column that holds it was added on 2026-07-27, after these were withdrawn, and it was not backfilled. Writing today's date into those rows would have invented a fact to fill a gap.
How a finding gets published
Rules run continuously over the stored observations. When one matches, it produces a candidate carrying the exact figures that support it, each paired with the query that produced it. A template turns the candidate into prose without touching the figures. A verifier then recomputes every figure against the record, and a single mismatch discards the draft entirely.
Findings that restate a count publish themselves once verified. Findings that say something unflattering about a named actor — a compliance failure, a contradicted identity — wait for a person to read them first, because a wrong one there would cost more than a missing one.
Standard applied
A finding is published when the record supports a specific statement, not when it supports a general one. Under the constitution, every figure must be recomputable from stored observations, sample size must be stated, and absence of evidence is reported as absence of evidence. A finding of n = 1 is published as a finding of n = 1.