33 entries, newest first, most recent 2026-09-29. Generated from the source document, so this page and the one we work from are the same text. Questions: abuse@agentcensus.io.
Breaking changes to the published API contract, announced here as ADR-028 requires: "needs to be announced as one, not shipped quietly." Everything else — new endpoints, new fields, bug fixes — lives in commit history and the per-item docs; this file is only for changes that break something a caller was relying on.
dataset_version 20.0: a 200 of JSON that is not a document of the requested kind is not a hitPer-mechanism publishing and probed fall, and errored rises, at every
HTTP document mechanism (/v1/overview, /v1/census; domains_publishing,
domains_probed and domains_error in the dataset export). Until this version
probeHTTP tested a response's content type and whether it parsed as JSON,
and nothing else, so any JSON served with a 200 at a mechanism's path counted
the domain as a publisher: an API error envelope such as {"code":500} at
/openapi.json, a status object at /.well-known/api-catalog, a request echo
at /.well-known/ard.json, and a draft-pro-adp-agent-discovery descriptor at
/.well-known/agent.json, which was even read as an A2A agent because it
carries a name
(#2171).
Migration 0186 is the boundary.
What changed. Each of the nine HTTP mechanisms — a2a, a2a_alt, ans,
ard, openapi, did_web, mcp, api_catalog,
oauth_protected_resource — now has a kind test (probe.ofKind) that names
the member making a document that kind: openapi or swagger, an id that is
a DID, linkset/links/resources, RFC 9728's required resource, an A2A
card's name or a listing (and no top-level protocol beginning ADP/), and so
on. A document that fails is recorded as outcome = error with the new
error_class wrong_kind: still a row (ADR-004), outside the denominator like
every error, with no agent and no stored payload.
What did NOT change. A document of the kind that a publisher left empty —
{"entries": []}, {"linkset": []}, an empty listing — is still a hit with
zero agents: the test reads for the member, not for entries in it. A host that
answers every path with the same document is still a miss with
wildcard_response, inside the denominator (ADR-062); that verdict outranks
this one. Invalid JSON is still parse_error, a non-JSON content type still a
miss, and llms_txt and agents_txt have no kind test. No outcome value is
added, no field is removed and no response shape changes. Nothing is
backfilled: a series spliced across this date compares two definitions of a
publisher.
observed.history[].actor on GET /api/v1/agents/{agentKey} is now a role
phrase, never a person. It used to carry the correcting member's email
address. It now reads "the verified owner of example.com", "the
operator-assigned owner of example.com", or "the owner of example.com (basis
not recorded)", composed from basis and the agent's domain. The address is
still recorded for audit and is not served. The same entry's summary no
longer ends with a "Verified owner of …" / "Operator-assigned owner of …"
sentence, because actor now says it. basis is unchanged. A caller that
keyed on actor as an identity has nothing to key on, and that is deliberate
(ADR-020 and
ADR-015, amendments of
this date; #2405).
composite.recommendedProfile is judged on the measured dimensionsThe field keeps its name and its four values, and changes what it means.
Inside TrustComposite (on GET /api/v1/agents/{agentKey}/trust, in search
results and on My agents), recommendedProfile is now the scoring engine's
cascade at the engine's thresholds, applied to the same measured dimensions the
composite's score averages. It used to be the engine's word verbatim, which
counts a signal that never reported as 0. On the same inputs the new value is
never lower than the old one. It is absent when nothing was measured, where the
engine said READ_ONLY.
The engine's word is not gone. It is composite.atdRecommendedProfile, and
the trust route also adds composite.unobservedSignals, the number of weighted
signals the engine scored with no observation. The top-level
recommendedProfile of the trust route and the insights trustProfile are
unchanged and still the engine's verbatim. A caller that wants the old reading
of the composite field reads atdRecommendedProfile
(ADR-041's amendment of this
date, #2408).
mechanism enum, and no published figure carries it yetA caller that switches exhaustively on mechanism will meet a value it has
never seen: mcp_endpoint (#2346). It is announced for the same reason
oauth_protected_resource was on 2026-09-15. It breaks a client that treats the
enum as closed. No field changed type, and the value sits immediately below
mcp in model.Mechanisms.
What it is. One bare GET of the conventional MCP streamable-HTTP path:
/mcp at the apex always, and mcp.<domain>/mcp and mcp.<domain>/ only
where the same pass's mcp_dns saw mcp.<domain> resolve and its control name
did not. No POST, no session, no credentials, and the resource_metadata URL in
a challenge is recorded and never requested. It counts as a hit only when the
answer is a 401 with a Bearer challenge, or a JSON-RPC 2.0 body, and a
nonsense path on the same host answers differently. The three new
labelVariant values are apex_mcp, mcp_host_mcp and mcp_host_root. It
mints no agent, so agentCount is always 0.
What deliberately does not change. It is probed and recorded, and not
published. No census total, cumulative figure, /v1/census cut or dataset
row includes it (model.UnpublishedMechanisms, and the NOT IN ('mcp_endpoint') in analytics/20, 21, 40 and 45, pinned by
services/cmd/compact/unpublished_test.go). Every published figure therefore
still spans the same fifteen mechanisms it did yesterday. Any response that
lists observations rather than figures can carry the value from today. No dataset_version
row accompanies this entry, because nothing published moves. It gets one on the
day it is published.
Six of the seven routes restricted on 2026-09-26 answer anyone again.
GET /api/v1/census/headline, /api/v1/census/labels,
/api/v1/census/protocols, /api/v1/census, /api/v1/status and
/api/v1/overview are back on the public tier, IP-limited and with no key, and
their response shapes are what they were before. The SPA's /census, /status
and /overview pages, their nav and footer links, and the census figure on the
front door are public again with them. The MCP agent's census_headline,
mechanism_adoption and discovery_labels tools answer with a figure again,
not a refusal.
What stays restricted. GET /api/v1/census/capability-change stays on the
platform-admin plane, since no public page reads it. The /census/methodology
page stays platform-admin only until the product documentation page is designed,
so the census page's links to it render for a platform admin only
(ADR-028's amendment of this date).
Seven routes that answered anyone now answer a signed-out caller with 401 and
a signed-in non-admin with 403. GET /api/v1/census/headline,
/api/v1/census/labels, /api/v1/census/protocols,
/api/v1/census/capability-change, /api/v1/census, /api/v1/status and
/api/v1/overview moved from the public tier to the platform-admin plane
(services/cmd/api/admin_census.go): a caller needs the platform-admin group
and a live elevation, and every read is audited. The SPA's /census,
/census/methodology, /status and /overview pages, and the census figure on
the front door, are gated the same way. At the operator's direction the census,
methodology, status and overview surfaces are restricted to platform admins in
the SPA and the API until a product documentation page is designed; the
underlying computation is unchanged
(ADR-028's amendment of that date).
What a caller loses, and what it keeps. Anything reading a figure or
datasetVersion off these routes anonymously now reads a refusal; the response
shapes are unchanged for an admin. The MCP agent's census_headline,
mechanism_adoption and discovery_labels tools stay listed and answer with an
isError carrying the 401, because the published capability digest pins the
tool list. Search, agent and domain records, and the /dataset/ export stay
public, and the export's manifest is now the public carrier of datasetVersion
(ADR-024's amendment of the same date).
A cap tightens to the number already published (#2210). The per-minute
window on the metered routes (anon_per_minute, search_per_minute) was
counted in each API task's memory. Prod runs two tasks behind a load balancer
with no stickiness, so a signed-out viewer or a session bursting across both
was admitted up to twice the published number before a 429 rate_limited.
The window is now one counter shared by every task, so the number in the
documentation is the number enforced. A client that paced itself to the old
effective ceiling will see 429s it did not see before; Retry-After still
says when to come back. API keys are unchanged and still counted per task.
In the other direction: re-reading a record already paid for today is now free on every task, not only on the one that charged it, so a signed-out reader's five record views are five records however their requests are balanced.
GET /api/v1/census has been telling callers that opted-out domains are
counted in the denominator. They are not, and never were. The caveat read
"Domains that opted out are counted in the denominator as skipped_optout and
never read"; 20_adoption_daily.sql filters suppressed = false before it
ranks, and the probed denominator is hit or miss only, so a suppressed domain
leaves the numerator, the denominator and every slice.
ADR-051 decided this; the
caveat was an inherited error that survived the decision.
Why it is in this file, which is for breaking changes. No field, shape or status code moves — the caveats array is free text, and the correction replaces one entry in place. What breaks is arithmetic done against the stated methodology. A caller reconstructing our rate, or adjusting it to remove the opt-out population it was told was included, was correcting for something that was not there, and would have moved the figure the wrong way. A published denominator convention is part of what a census means by a number, so a correction to it is announced rather than left to be discovered.
What is true, and what is unchanged. The refusal is still recorded — a
suppressed domain emits skipped_optout per mechanism on every later pass, and
that is published as a separate skipped counter, never as non-adoption. No
figure moves as a result of this change: the rollup already behaved this way, so
only the description of it changes. The correction names the old claim and
denies it, rather than quietly replacing the sentence, because a caller who read
the old text needs to know it was wrong and not merely reworded.
A cap loosens, and a client-side count of it stops matching ours (#1946).
ADR-040 §1.1 promises a signed-out
viewer five detail reads a day and defines a detail read as a record — an
agent's or a domain's, a domain's including "the agents there". The meter
charged a token per metered request, and since the records grew sub-resources
(/v1/domains/{d}/agents, /v1/domains/{d}/declarations,
/v1/agents/{k}/declarations, /v1/agents/{k}/documents/{source}) one page has
been two charges, Show more a third and a reload two more. Charging is now
once per record per viewer per UTC day, per that ADR's amendment of the same
date.
Why it is in this file, which is for breaking changes. Nothing a caller
could do before is refused: a caller that hit the wall after N requests now
gets further, and one under the cap is unaffected. What breaks is a count kept
on the client — repeated reads of one record no longer move X-RateLimit-Remaining,
so a caller predicting our number from its own request count will drift, in the
generous direction. That is announced rather than left to be discovered.
The headers were already specified for this, and the ADR sentence was not.
X-RateLimit-* are absent on a repeat, because it takes no token and this
platform does not publish a rate-limit figure it did not measure. That needed no
contract change: openapi.yaml's RateLimited response has always said the
three are "omitted, rather than guessed at, when no counter measured the
request" — it now names this case alongside the fail-open one. §1.1's looser
phrasing, "the same headers every metered response carries", is what the
amendment corrects. The 401 register_required refusal still carries all three,
as does every response that took a token.
usage_event gains a fifth status, admitted_repeat. It counts as a
request everywhere the four existing ones do — the /admin/usage rollup
special-cases only walled — so no published figure moves and no migration is
needed.
history on /v1/agents/{agentKey} returns fewer changed entries from the
next probe pass onward (#1880, migration
0162_a_history_row_records_whether_the_record_took_it.sql). No shape change:
no field added, removed or renamed, and kind still takes the same values. What
changes is which snapshots produce an entry at all.
The entries that go were never events. ADR-010 precedence has held a weaker
source's display_name, description, protocols, capabilities and status
at the strongest live contributor's values since #1812, while the SCD-2 history
row was written for every snapshot regardless — and the timeline was derived by
diffing consecutive rows. So each source's disagreement with whichever source
happened to precede it rendered as a change to the record. This platform's own
record carried seven changed entries from one pass; one described a write that
reached the record. The newest of the other six said the status had just become
UNVERIFIED directly beneath a record serving ACTIVE. Entries are now derived
from the effective record, so a snapshot the gate discarded produces no sentence
and the newest entry agrees with the record in the same response by construction.
Announced because a caller counting entries will measure a drop that is not a
drop in the world. Anything treating kind='changed' as a per-agent change
rate, a "recently churning" list, or an input to a freshness heuristic sees the
rate fall at every record with more than one contributing mechanism — sharply,
because the discarded writes outnumbered the accepted ones. The census figures
are untouched: this is a read-time derivation over agent_history, which no
export reads, and dataset_version therefore does not move (the migration
header says why at length).
Rows already written keep narrating the old way. The two new columns are
false on every existing row, which is what those rows can honestly say — the
verdict was never stored and is not recomputable. A record's next probe pass
writes marked rows over its newest end, which is the end anybody reads, so the
fix arrives per record rather than all at once.
protocols[] starts carrying the protocol an entry declares in the singular, and that figure can rise at ans and ardOne field binding, and protocols[] fills in on records this census already
holds (#1879). No shape change, no new field, no renamed one. rawAgent read
the plural protocols list and — for ARD — the type media type, and nothing
else; the singular protocol, which is what a live ANS index actually writes,
was bound nowhere. An entry declaring "protocol": "mcp" and no list therefore
parsed to an agent with a name, a description, its capabilities, and an empty
protocol list. It is now read and merged with the plural rather than preferred
over it, so an entry carrying both is read as declaring both and
NormalizeProtocols dedupes.
No adoption figure moves, and that is the first thing to read. A domain
publishing an ANS index or an ARD catalogue was already a hit with an agent —
this is not the error-to-hit widening of 16.0, no outcome class changes, no
denominator changes, and no row is written into a window that has already been
published. What changes is one attribute of an agent already counted.
/v1/census/protocols can rise at ans and ard, and the search filter moves
with it. That endpoint counts protocol tokens across agent records, so an agent
that contributed no token contributes one from its next pass onward; search's
protocol filter, which matches on the token, starts returning records it
excluded. A caller treating an empty protocols[] as the publisher declaring no
protocol — a completeness score, a "declares nothing" bucket, a facet count —
will watch it fill in. How much of the corpus is affected is not measured here
and this entry does not estimate it: what is measured is that this platform's own
front-page record served protocols: [] on 2026-09-20 while the document it was
read from said mcp, which is the defect arriving on the one domain whose
publisher can be asked.
contentHash moves once per affected agent. protocols joins the snapshot
hash, so the first pass after this ships closes the current SCD-2 row and opens a
new one for every agent whose entry declared a protocol in the singular — an AIDB
response span ending and a new one beginning, as the description widening of
2026-09-17 and the sorted-RRset hash change of 2026-09-15 both were. Nothing is
restated and nothing is backfilled; rows written before that pass hold what the
parser of the day read.
This entry deliberately carries no dataset_version row, and the precedent
cuts both ways. For: protocols means what it has always meant — the protocols
the publisher declared — and was merely incompletely populated, which is the
reading the 16.0 entry below applies to its own description half, and no
outcome class, denominator or already-published window moves. Against: the
closest analogue in this file took one. 8.0 (migration 0117,
DATA_MODEL §8) bundled the DNS-AID leaf's unread alpn —
another spelling that left protocols[] empty on records the census already held
— into a MAJOR bump, alongside two corrections that did move adoption. A version
row is also a migration, and this change takes no migration number. So the
question is put rather than answered here: whether the movement at
/v1/census/protocols earns a row is the reviewer's call rather than an agent's,
and 16.0 is the precedent for what taking one looks like after the fact — it
followed its fix rather than shipping with it.
What deliberately does not change. type stays the narrow reading it was
written to be: application/json names no protocol, and an entry's URL is still
never mined for one, because a declaration the publisher did not make is not an
observation. Nothing infers a protocol from the document's own kind either — an
A2A card and an MCP manifest still declare no protocol token unless they list
one, per model.WithProtocolVersion's rule that no entry is created (ADR-035).
dataset_version 16.0: a capability map that nests objects is a hit, not a parse errorThe version row for the entry below this one (#1830, migration
0155_a_capability_map_that_nests_objects_is_a_hit_not_an_error.sql). That entry
announced a published figure moving at mcp, a2a, a2a_alt, ard and ans,
and said in writing that it carried no dataset_version row: a version row is a
migration, that change took no migration number, and whether the movement needed
one was the reviewer's call. The call was to land the fix whole and let the row
follow. This is the row that follows.
Nothing in the dataset changes on this date. No table, column, index or
constraint; no row updated; no census_publication restated. What changes is
that every figure published from ff38713e onward carries a version saying the
definition of a hit at those five mechanisms is not the one 15.0 carried.
MAJOR, and both terms of the rate move. A card whose capabilities nests an
object under a capability flag used to fail json.Unmarshal and be recorded as
outcome='error' — which is outside the adoption denominator, so such a domain
was neither a hit nor probed for the rollup's purpose. Reading the shape
correctly makes it a hit and puts it in the denominator: domains_publishing
rises, domains_probed rises with it, domains_error falls. That is unlike
13.0, where a domain already in the denominator moved from miss to hit and the
denominator held still. A caller comparing a coverage rate across this date is
comparing two populations, not two numerators. The same pull request's
description half does not set the increment — description still means the
publisher's description and was merely incompletely populated — but numbering is
once per movement rather than once per change
(ADR-054 §6), so one row covers
both halves.
agentcensus.dev served moved figures under 15.0 for the interval between
ff38713e applying and 0155 applying, and a reader who pulled dev inside that
window holds figures labelled one version low. Prod's last promote
(2026-09-18T00:39Z) predates ff38713e, so agentcensus.io has not published a
moved figure at all and this row reaches it with the code it labels.
description starts carrying what publishers actually wrote, and one sentence this code wrote about api-catalog documents goes awaySix parsers stopped discarding description-shaped text the fetched documents
already carried (#1797). Nothing here changes a shape, adds a field or renames
one. It changes the CONTENT of description — and at a2a, a2a_alt, ard,
ans, mcp and agents_txt of capabilities — on records this census already
holds, from the next probe pass onward. Announced because three things break for
a caller reading those columns, and because the last of them is a published
figure moving for a reason that is about this instrument rather than the world.
A description that was empty will not be. 97.4% of agent records carried
none: llms_txt kept its H1 and threw the prose under it away, agents_txt read
Name: and ignored Description:, openapi read a title while every operation
below it carried a hand-written summary, an A2A card's skills[].description was
dropped, and api_catalog stored a sentence about the document instead of the
document's own. A caller treating an empty description as a property of the agent
— a completeness score, a "needs review" queue, a layout that hides the field —
will watch it fill in. Search moves with it: ranking here is relevance-only
(ADR-020 §1,
ADR-041 §6), so text that was never indexed
is now matchable at search_tsv weight C. No schema change was needed for that
and none was made — search_tsv is generated over agent and both columns
already feed it.
api_catalog stops describing itself. API catalog (resource list): 2 entries was a sentence this code composed and stored in the publisher's field.
It is now the fallback, used only where a document describes nothing at all, so a
caller parsing an entry count out of that string will find the publisher's prose
there instead. The count was never a contract; it is called out because it read
like one.
contentHash moves once per affected agent. A composed description is new
bytes, so the first pass after this ships closes the current SCD-2 row and opens
a new one for every agent whose document carried text the parser was not reading
— downstream, an AIDB response span ending and a new one beginning, exactly as
the sorted-RRset hash change of 2026-09-15 was. Nothing is restated and nothing
is backfilled; rows written before that pass stay valid and hold what the parser
of the day read.
And one published figure can rise at mcp, a2a, a2a_alt, ard and
ans. capabilities in MCP's own spelling is an object of objects
({"tools":{"listChanged":true}}), and until now a card written that way failed
json.Unmarshal and was recorded as an error row rather than a hit — at the
one mechanism where the capability map is most of what the card says. Error rows
leave the denominator, so those domains were neither hits nor probed for the
purpose of the adoption rollup, and reading the shape correctly turns each of
them into a hit with an agent. The same widening applies wherever a card nests an
object under a capability flag. This entry deliberately carried no
dataset_version row: a version row is a migration, this change took no
migration number, and whether the movement at those five mechanisms needed one
before it merged was the reviewer's call rather than an agent's — #1746's entry
below is the precedent for what taking one looks like. It needed one: the call
was to land whole and let the row follow, and it followed as dataset_version
16.0 in migration 0155, the entry above this one.
What deliberately does not change. No displayName moves: an over-wide read
lands at weight A, which is the #704/#644 harm, so the llms.txt title window
stays at ten lines while only the summary window widens. Every new read is
bounded by a named constant in services/internal/probe —
maxSnapshotDescription (4096 bytes, ATD's own MaxDescriptionLen),
maxDescriptionPiece (512), maxCapabilityNames (64) — and a cut is marked
[truncated] rather than left to read as a publisher's own short sentence. And
no policy about what the text MEANS: this is untrusted publisher input, storing
more of it widens the surface
ADR-052 governs, and that
ADR is Approved and wholly unbuilt. Nothing here withholds, flags or suppresses
anything; that decision stays where it was made.
labelVariant whose denominator is not the census, and one DNS hash that changes without the record changingTwo additions at dns_aid_v1, neither of which moves a published adoption
figure — which is why this entry leads with what it is not (#1784/#1785,
dataset_version 14.0). Announced anyway, because the first breaks a caller
that reads every variant's probed the same way and the second breaks a caller
that treats dnsAnswerSha256 as stable for an unchanged record.
A fifth labelVariant: _ans-badge. The Agent Name Service transparency-log
pointer at _ans-badge.<name>, beside the _ans record this census already
read. It is the first variant in this dataset that is not asked of every domain
probed: it is queried only where _ans itself answered, because a name this
census cannot justify asking is a name it does not ask
(ADR-006).
Two consequences, and the first is the one to read carefully. Its probed
counts the domains publishing _ans, not the domains probed — so it is a share
of ANS publishers and never a share of the census, and /v1/census's label cut
carries a caveat saying so beside the numbers. Second, dns_aid_v1 adoption
cannot move because of it: every domain that can produce an _ans-badge row is
already a dns_aid_v1 publisher via _ans, so the rollup's best-outcome collapse
per (day, mechanism) has nothing new to collapse. The row also mints no agent — a
receipt is not an agent, so agentCount is always 0 and enumerationBasis always
empty. The pointer URL is never dereferenced
(ADR-036: verify proofs, never display
badges); presence and the record's hash are the whole observation. Every row
written before today carries the empty string and stays valid — an addition, not a
rename, the same reading dataset_version 8.0 and 13.0 took.
dnsAnswerSha256 at dns_aid_v1's TXT labels is now computed over a sorted
RRset. A DNS RRset is unordered
(RFC 1035 §3.2.1) and
resolvers rotate it, so the previous hash could change without the zone changing —
which matters most at _ans, the one v1 label that routinely carries several
records, one per protocol. No adoption figure reads this column, so nothing
published moves. What does happen, exactly once, is that a domain whose
records are unchanged gets a different hash on its first pass after today, which
reads downstream as an AIDB response span ending and a new one beginning. The two
SVCB hash sites are deliberately left in arrival order. Nothing is restated and
nothing is backfilled.
mechanism enum, and adoption rises at mcp, a2a and llms_txt without anybody publishing anything newA caller that switches exhaustively on mechanism will meet a value it has
never seen — oauth_protected_resource — and a caller charting adoption at
mcp, a2a or llms_txt will see each of those three step up on this date for
a reason that is about this instrument and not about the world (#1746,
dataset_version 13.0). Neither is a shape change: the enum has always been
open-ended in the sense that values are appended, and no field changed type. It
is announced here because the first breaks a client that treats the enum as
closed, and the second is indistinguishable, from the outside, from an
ecosystem moving.
The new mechanism. oauth_protected_resource is the RFC 9728 OAuth 2.0
Protected Resource Metadata document at
/.well-known/oauth-protected-resource. It is its own signal rather than a
label_variant of mcp: mcp means an MCP document was served, and this is
not one, even though the MCP authorization specification is why most publishers
serve it. It appears second to last in model.Mechanisms, above only the DNS
beacon, because it names an endpoint and describes no agent — no name, no
description, no capabilities — so anywhere higher it would win an attribute
conflict against a card that actually described the agent. It has no history
before today, so its observation window is shorter than the other fourteen
mechanisms' and /v1/status's mechanismsFirstProbed says so; the census
caveat names it automatically for as long as that stays true.
Why three counts rise. mcp now also reads SEP-1649's
/.well-known/mcp/server-card.json, a2a also reads the A2A Registry's
/.well-known/agents.json, and llms_txt also reads /llms-full.txt. A
domain that has served any of those all along was a miss at that mechanism
yesterday and is a hit today. The denominator does not move — those domains
were already probed for those mechanisms, and the adoption rollup counts a
domain once per (day, mechanism), so a second or third path costs no
denominator. An adoption series computed across this date is two different
questions spliced together.
a2a's agentCount can now exceed 1. The listing at agents.json is
several cards and each is an agent, so it mints one snapshot per card where the
two single-card paths mint at most one. Because the listing is recognised by
shape rather than by path, an array of cards served at a single-card path now
parses too, where it previously produced a parse error.
New labelVariant values, on /v1/census's label cut and on
probe_observation: well_known_mcp_server_card, well_known_agent_json,
well_known_agents_json, llms_txt, llms_full_txt. Every row written before
today carries the empty string and stays valid — an addition, not a rename, the
same reading dataset_version 8.0 took when ard gained its second spelling.
llms_txt is the first text mechanism to carry a variant at all, so rows
for /llms.txt move from empty to llms_txt from this date.
dataset_version 12.0: a document about a site is not a record about an agentThe headline agent count is now the count of public.agent rows admitted
through a mechanism whose document describes an agent or names an endpoint,
not every non-INACTIVE row regardless of mechanism. Four
presence/documentation surfaces — llms_txt, agents_txt, dns_aid_bare,
mcp_dns — continue to count a domain in that mechanism's own
domains_publishing, unchanged, and no longer admit an agent to
/v1/census/headline and /v1/overview's agents, newAgents and
multiMechanism. This is
ADR-039 Amendment 3,
2026-09-15 — §1's admission test already refused these records; the headline
just never asked it. Migration 0144 is the boundary.
What was measured, read 2026-09-14:
| agents discovered (headline, before this date) | 157,264 |
declared no protocol but https |
156,422 (99.5%) |
existed only because a domain served /llms.txt |
135,800 |
existed only because a domain served /agents.txt |
6,146 |
| minted from those two surfaces alone | 141,946 (90.3%) |
GET /v1/domains/asana.com/agents returned one "agent" whose only mechanism
was llms_txt, name "Asana", no capabilities, no endpoint — and
parseLLMSTxt's own comment had said so since it was written: "UNVERIFIED:
llms.txt advertises documentation, not an agent endpoint."
Why. An llms.txt is a retrievable thing whose subject is a site's
documentation; an agents.txt is a policy about who may act on a site; a
resolving discovery hostname (dns_aid_bare, mcp_dns) is about a hostname.
None of the three has exactly one agent as its subject, so a figure that
counted them was not counting agents — it was counting domains that
published something discovery-shaped, under a name that says agents.
What changed. The headline agent count is the count of public.agent
rows with at least one agent_source row in the describing set — a2a,
a2a_alt, ans, ard, api_catalog, did_web, mcp, openapi,
dns_aid_v1, dns_aid_v2 — plus claimed (ADR-010 precedence 0).
services/cmd/api/main.go's describingMechanisms derives that set from
model.Mechanisms minus a named exclusion list, so a fifteenth mechanism is
counted by default and only a listed exclusion is not.
What did NOT change. No public.agent row is retired and no minting
logic in internal/probe changes — parseLLMSTxt is untouched. The rows
minted from the four surfaces stay: they are today's search corpus, and
/v1/search still returns them. Per-mechanism domains_publishing is
unchanged at all fourteen mechanisms, including the four this entry
concerns. mcpServers is unchanged — it is already protocol-derived — and
so is /census/protocols's agentsTotal, which is corpus-wide by its own
caveat and does not read this set.
What this does not decide. Whether the 141,946 rows already minted from
the four surfaces are retired, and what /v1/search should return for them,
is a separate change with its own dataset_version — removing them changes
what a query for "llms.co" finds. Until decided, they stay, and the census
caveat says so.
dataset_version 11.0: a zone that answers every name is not a DNS-AID publisherNo field changed shape. outcome = 'miss' now includes a case it previously
excluded, at dns_aid_v1's three underscore labels and at mcp_dns's
_mcp._tcp SRV label, and adoption falls on this date with no change in what
anybody publishes. This is the third correction announced here that moves a
figure down, and the largest of the three. It is not a new rule: it is
ADR-062 §1 — "a response that answers
everything is not a publisher", and "the control is not optional" — reaching
the two lookups dataset_version 3.0 left out. Migration 0143 is the
boundary.
What was measured, in production on 2026-09-13. dns_aid_v1 was publishing
926 domains that day and 8,111 cumulatively since 2026-09-05:
dns_aid_v1 hit domain-days since 2026-09-05 |
8,279 |
of those, on zones the bare-subdomain control flagged wildcard_response the same day |
7,800 (94%) |
distinct domains hitting all three underscore labels with the identical dns_answer_sha256 |
7,234 |
_mcp._tcp SRV hits in the same window, on such zones |
38 of 40 |
mcp. address-form hits on such zones — the surface that already asked a control |
0 of 2,100 |
Read live over DoH: slupca.com answers _index._agents., _agent., _ans.
and an invented control name with "bio=89d457271974df2ca8e491a5a8fcefd523a867c3";
gmfamily.com answers every label with its SPF record and an Afternic
verification string; jjosh.de has a wildcard CNAME to its apex, so every name
answers the apex's thirteen unrelated TXT records. The largest apparent
publishers were the same shape — ricicle.com at 362 agents on two labels,
cmasa.es at 197. None of them has ever published a DNS-AID record. The
corrected figure is roughly six v1 publishers a day and a few hundred
cumulatively, which agrees with the 2026-08-11 passive-DNS sweep's 29 confirmed
publishers internet-wide instead of exceeding it by two orders of magnitude.
Why the control that already existed did not catch it. This census has asked
agentcensus-control-should-nxdomain.<zone> since 3.0, and it asks it for an
address. A wildcard belongs to an RRset rather than to a zone (RFC 4592
§2.2.1): a zone with * IN TXT and no * IN A answers that control with
nothing and answers every discovery label anyway. The code said so out loud and
believed the opposite — "A zone with a wildcard TXT would have to be
publishing" — and * IN TXT is a hosting-panel default and an SPF-at-any-name
habit. The SRV half rested on "a wildcard A record answers an SRV query with
NODATA", which is true of an A wildcard and of nothing else: a * record set
can carry SRV, and a wildcard CNAME hands every query type to the target's own
records.
What changed. The same invented name is now asked in the record type of the
question it qualifies — A/AAAA as before, TXT for the three v1 discovery
labels, SRV for _mcp._tcp — memoised separately, each asked at most once per
zone and only once a lookup of that type has already answered. A domain
publishing none of these names is asked no control at all, which is nearly every
domain, so the request cost barely moves. /crawler.html publishes all three
lines and the reason.
What a discounted hit looks like, unchanged from 3.0 and 9.0:
outcome = 'miss' with error_class = 'wildcard_response', agent_count and
declared_agent_count 0, enumeration_basis cleared, no agent snapshot. miss
deliberately and never error — nothing failed, we asked and the zone answered,
and the answer was not about the name. The row is never dropped and the attempt
never skipped, so the denominator does not move. One thing is new: a demoted
_index._agents. answer is no longer handed to the v2 leaf fan-out, so the
queries derived from it stop being sent. That is a reduction in what this
crawler asks of a stranger's zone as well as a correction to a count.
The second half: a record has to look like a DNS-AID record. The control is
not the whole fix, because the rule it qualifies was also too loose. At a
dedicated label the probe counted any TXT string not on a short list of
recognisably-somebody-else's formats, which reads as a conservative under-read
and is the opposite — it counts whatever a zone happens to answer with. From
11.0 a record at _index._agents., _agent. or _ans. must carry a DNS-AID
shape: a v= version tag, an agents= inventory key, or — at
_index._agents. alone, per ADR-055's 2026-08-31 amendment — a keyless
<name>:<proto> entry. An unfamiliar version family still counts, as
generic, because the formats are still moving. The two halves are independent
and neither is redundant: a zone answering v=aid1 at every name passes the
shape rule and is caught by the control, and a single Afternic string at
_agent. in a zone with no wildcard passes the control and is caught by the
shape rule.
What is untouched. The apex label, which asks no control because a name that
exists in every zone is not a wildcard question and which has always required a
recognised signature. _mcp.<domain> and mcp.<domain> TXT, where a hit
requires v=mcp1 and a wildcard zone would have to publish that at every name.
model.Mechanisms stays fourteen.
Two limits, stated rather than glossed. History is not restated, and as
with 3.0 that is a limit rather than a preference: the control's answer on any
past day was never recorded on any row — ADR-062 §2 is still unbuilt — so which
historical hits were wildcard zones is not derivable from the warehouse.
Affected domains correct themselves as recrawl-sweep reaches them. And the
94% above is a measurement taken by hand against the curated lake on one day; it
is written into the caveat for that reason, and it is not a figure this
dataset can answer for itself.
authDeclared and authEnforced start having values, and authEnforced leaves the authenticated passTwo fields that have been absent from every response since they were
published begin appearing today, on every check rather than on a few — and
activeVerification.authenticated.authEnforced, which the 2026-08-25 entry
below says both keys carry, is now documented as belonging to the
uncredentialed pass alone. No caller sees a value change, because there
were no values: both fields read active_verification.auth_declared and
active_verification.auth_enforced, a column pair (migration 0005) bound
into the active-verify writer's INSERT and assigned by nothing before it, so
every row has held NULL in both for the life of the columns and the fields
were omitempty-dropped from every payload. A caller that branched on their
absence — and absence was the only state available — is the one to read this
(#1523).
What they mean now. authDeclared is the agent's own published posture:
whether its card names any auth scheme at all (agent.auth_declared,
ADR-065). It is NOT
NULL, which is why the field now appears on every behavior check rather
than on none — false, "declares no scheme", is a measurement and reads as
one. authEnforced is tier 1's own record of whether a call carrying no
credential was refused (active_verification.unauthenticated_refused);
false is the finding, an agent that declares auth and answers a stranger
anyway. Absent still means unmeasured — the a2a methods do not ask — and has
never meant "not enforced".
authEnforced is not a field of the authenticated pass. Only the
uncredentialed pass can put the question, so the authenticated key omits it
rather than answering one it did not ask, and the SPA's credentialed chip
already declined to render it. Nothing disappears from a live response, since
the field was absent there too; what changes is the contract, which used to
promise an identical field set under both keys.
The overlay answers for the uncredentialed pass only, both on the detail
surface and — new here, and previously unenforceable — on the row surfaces.
ADR-041 §3 and ADR-021's "two
passes, named separately, never merged" amendment put it there: these fields
mean what any caller could reproduce. So authDeclared: true, authEnforced: true says a stranger's call was refused, and it does not distinguish an agent
that enforces auth from one that refuses everyone. That separation needs the
credentialed pass, and it belongs to the registered signal, not to this
surface. Nothing about the fields ranks, filters or facets anything (ADR-047
§3.1, ADR-020 §1), and the composite does not move.
semanticAvailable now answers for the corpus, not only the queryA caller who was reading semanticAvailable: true as "this search's meaning
match is trustworthy" will start seeing false on prod today, for every
query, until the embed backlog drains (#1441). No caller update is
required — a truthful false is exactly what this field should have read
all along — but the value moving for every search at once, on a field
nobody changed the shape of, is the kind of thing a caller who alarms on it
or logs it should not have to discover for themselves.
Before this change, semanticAvailable reported only whether the query
embedded, which read true in production while agent_embedding held zero
rows — the embed backlog had not yet drained. It now reads true only when
the query embedded and the corpus holds at least one agent embedding.
The response also gains an optional semanticCoverage: {embeddedAgents, totalAgents} — two counts, never a ratio (ADR-011) — present only when the
corpus count was actually taken; absent, not {0, 0}, on an unmeasured
lookup. See SEARCH.md §"When it degrades".
dns_aid_bare and mcp_dns start counting IPv6-only publishersA host that published an AAAA record and no A record was recorded as a miss
on both presence-only DNS surfaces until this change, so part of any step up in
dns_aid_bare or mcp_dns after 2026-09-08 is publishers that were there all
along. No dataset_version moves — an IPv6-only publisher was always in scope
and was always meant to count — but the annotation is owed for the same reason
the entry below it is (ADR-011 §4, ADR-004).
dnsx.LookupHost asked for A alone. Its own documented contract said
A/AAAA, and its record filter accepted AAAA answers, so the code read as
though both were asked — but an A query never returns an AAAA record, so
that branch had never once fired (#1468). The two mechanisms whose entire signal
is presence took the loss: catalog./index. for dns_aid_bare, mcp. for
mcp_dns.
The lookup now asks for AAAA when, and only when, the A answer says the name
exists and has no A record. NXDOMAIN forecloses every type at once, and
NXDOMAIN is what those names are on almost every domain, so a pass sends the same
number of queries as before on the common path. /crawler.html publishes the
pair beside each of those qnames and states the condition.
Nothing is backfilled, and the size of the shift is not knowable in advance.
The affected rows are stored misses, and a miss derived from an A-only question
cannot be re-read as anything else — the record that would have contradicted it
was never asked for. What the shift equals is the number of catalog., index.
and mcp. hosts in the corpus reachable over IPv6 and not over IPv4, which is
precisely the quantity this code was unable to see. crawl_run.code_version
records which side of the change a partition is on.
dataset_version 10.0: endpoint_host is parsed out of the declared endpoint, not cut out of itA caller reading endpoint_host or endpoint_same_origin across this date is
comparing two readings of the same field, and the older one was wrong for three
shapes of card. Nothing about what a publisher publishes changed.
endpoint_host was produced by trimming https:// and http://, cutting at the
first / or :, and lowercasing the remainder. So:
| declared endpoint | published before | published from 10.0 |
|---|---|---|
https://alice@example.com/agent |
alice@example.com |
example.com |
https://[2001:db8::1]:443/agent |
[2001 |
2001:db8::1 |
wss://example.com/agent |
wss |
example.com |
The third is the one worth stating plainly: any scheme other than http or
https was published as the host, because its colon was then the first one
found. A2A v1.0 supportedInterfaces[] entries carry a transport beside the URL,
so a card declaring a WebSocket or gRPC endpoint is conformant, and this census
recorded such publishers as serving an agent at wss.
It travelled twice. endpoint_same_origin is computed by comparing this host
against the registrable domain probed, so a host the parser mangled could only
compare unequal — a publisher declaring its own host with userinfo attached was
published as pointing at a third party. That is a supply-chain claim about
somebody's deployment, produced by a parsing bug, and it is the second time this
field has needed correcting (2.0, migration 0090).
Migration 0125 is the boundary. Nothing is restated and nothing is
backfilled, and here that is a decision rather than a limit:
services/cmd/replay could re-derive both fields from stored payloads, and is
deliberately not asked to, because writing corrected values into a window that
has already been published is its own MAJOR bump — a consumer who queried March
before and after would hold two answers to one question. Affected agents correct
themselves as recrawl-sweep reaches them. How many there are is unmeasured;
it cannot be asked from outside the VPC.
What has not changed, and is not going to: a card declaring an internal
hostname or a private address — 10.0.0.5, localhost,
crm.corp.internal — is still published exactly as declared. This change makes
those values more accurate, never rarer. The census reports what a public card
says; the control against ever reaching such an address sits on collection,
where a declared endpoint is not contacted at all
(ADR-006,
ADR-021) and anything that tried would be
refused at dial time (SECURITY.md finding 1b).
A caller comparing the four DNS mechanisms across the resolver outage will see
them return to a level that may be higher than before it, and none of that
difference is adoption. No dataset_version moves — nothing about what the
dataset means changed — but the annotation is owed for the same reason the entry
below it is (ADR-011 §4, ADR-004).
Three things were suppressing dns_aid_v1, dns_aid_v2, dns_aid_bare and
mcp_dns, and only the first of them zeroed anything. The entry below is the first:
our own resolver was unreachable, so those four published probed 0. The second is
quieter and older. Unbound applies ratelimit per delegation point, and
ratelimit-for-domain — the exemption list that keeps the limit from applying where
no politeness is owed — is exact match, so the zones holding other zones'
nameserver addresses were never exempt. Resolving one nameserver set is a burst of up
to 2N queries at a single delegation point (13 root names over A and AAAA is 26,
against a limit of 20), and ratelimit-factor passes one query in ten, so the
symptom was never a stall: it was those four mechanisms missing a share of a busy
suffix while HTTP mechanisms on the same domains succeeded.
Correction, 2026-09-08 (#1470): the arithmetic in the paragraph above is wrong; the conclusion it supports is not. There is no 2N in this configuration. The resolver runs
do-ip6: no, which unbound copies into the iterator'ssupports_ipv6and which gates every AAAA query for a nameserver target, so one nameserver set costs N queries — the root's 13 names cost 13, which was under the limit of 20 rather than 26 over it. The exemptions are still owed, for aggregate rate rather than one pass's burst: at ~34 concurrent passes/s, every referral chain that misses the cache funnels through these same few delegation points, so our rate there scales with the pass rate and not with any zone's NS count. Nothing in this entry changes for a caller — the suppression was real, the level shift stands, and nothing is backfilled. Only the explanation of why the exemptions are required is corrected.
Seven exemptions ship on 2026-09-08 (root-servers.net, gtld-servers.net,
afilias-nst.{info,org}, cloudflare.com, domaincontrol.com, nsone.net).
The consequence for a caller is a level shift, not a rate correction: the four
DNS mechanisms' probed denominators recover to what they should always have been,
so both the counts and the rates derived from them may step up relative to any
partition before 2026-09-06 as well as relative to the outage. crawl_run.code_version
records the discontinuity on both sides, per ADR-004, and nothing is backfilled —
the suppressed queries were answered with SERVFAIL and there is nothing to re-derive.
The third is the largest and it was not a missing exemption — it was the limit
itself being set below the size of a single DNS pass. resolver_zone_ratelimit_qps
was 20, derived from an ADR-068 §2.2 sentence that said a pass is "about eight
authoritative queries in one burst". /crawler.html publishes the list and it is
fifteen name/type pairs across twelve distinct qnames — this entry first said
seventeen, which matches no version of that page (corrected 2026-09-08, #1470) —
and qname-minimisation turns each qname into one query per label below the zone
cut, all at the same delegation point the limit counts. The ceiling was never
derived from a count of names, in any case; it was derived from the measurement
below, and that measurement is unaffected.
Measured on dev 2026-09-08: ≥38 queries in one second at one publisher, against
a ceiling of 20. With ratelimit-factor passing one in ten, roughly 40% of every
affected zone's DNS pass returned SERVFAIL, on 708 distinct zones per ten minutes
— and those are error rows, which leave the denominator (ADR-004). Unlike the
second suppressor this one applied to ordinary publisher domains, not just busy
shared suffixes, so it touched a far larger share of the corpus. The number is 100
in both environments as of 2026-09-08.
What this means for a caller is the same level shift, larger. The four DNS
mechanisms' probed denominators were short by up to two-fifths on a large share
of zones for as long as the resolver has been serving, so the step up on both
counts and rates may be substantial and is entirely a suppressor lifting. Nothing
is backfilled, for the same reason: the refused queries were answered SERVFAIL and
there is nothing to re-derive.
All three suppressors are fixed only where the fix has been applied. Dev applies on merge; prod moves on a promote dispatch and has applied none of them, so prod's series carries all three effects until then.
Correction, 2026-09-22 (#2013): prod applied on 2026-09-09, the day after the paragraph above was written, and it carried all three. No deployment log is needed to place the discontinuity, which is the useful part: prod's own published series dates it.
dns_aid_v1read 87,012 probed / 151,909 errored on 09-07 and 44,125 / 183,869 on 09-08, then 225,592 / 360 on 09-09, and its error rate has stayed under 0.3% for the thirteen days since; a ceiling still at 20 would show as a ~40% error share, so the same apply carried the exemptions and the 100 as well. So the affected prod partitions are 09-07 and 09-08, and every one from 09-09 on is clean of all three. One thing not to read from this:probedmoves with the crawl pace, so 248,240 on 09-06 against 225,592 on 09-09 is not a level shift in either direction — the comparison this entry warns about has to be made within a mechanism against the error column, not across two days' denominators. Nothing is backfilled, for the reason the entry already gives.
A caller comparing across 2026-09-06/07 is comparing a measurement with a gap.
dns_aid_v1, dns_aid_v2, dns_aid_bare and mcp_dns publish 0 with probed 0 for the affected days, and step back to their real level afterwards. No
dataset_version moves, because nothing about what the dataset means changed —
this is an outage in the collection of it, annotated rather than smoothed over per
ADR-011 §4 and ADR-004.
Correction, 2026-09-22 (#2013):
probed 0is dev's shape and the lede above does not say so. The table below is labelledagentcensus.devon purpose, and on the public dataset the same fault reads as a two-day trough with a large error count rather than a hole.dns_aid_bareonagentcensus.io: 176 publishing / 248,197 probed on 09-06, then 79 / 86,990 on 09-07 and 26 / 43,978 on 09-08 with ~152,000 and ~184,000 errored, recovering to 168 / 224,925 on 09-09. The affected prod days are 09-07 and 09-08 only (see the 09-08 entry's correction). The reading this entry recommends survives the difference, but the instrument changes: on prod the tell is the error column, not a zero. And one thing stays unexplained rather than smoothed over — prod retained a 19% non-error share on 09-08, a full day on which this entry's own argument says a DNS-only mechanism with an unreachable resolver can report nothing at all. Which path answered those is an open question on #2013; no figure here is restated on account of it.
The switch to our own recursive resolver (ADR-068 §2) shipped two independent
HTTP/1.1 clients against a listener that speaks h2c and nothing else — HTTP/2 with
prior knowledge, no HTTP/1.1 and no upgrade path. The NLB's HTTP health check could
therefore never pass, so the target group held zero healthy targets and the NLB
refused every connection; and dnsx.Recursive itself was on Go's default transport,
which speaks HTTP/1.1 for an http:// URL, so even a healthy fleet would have
answered it nothing. Either defect alone produces this gap. Those four mechanisms are
answered entirely from DNS, and ADR-004 records an error as a mechanism not
observed (#1411).
As published on agentcensus.dev — dev, not the public dataset — the shape is
unmistakable once the two halves are put side by side:
| mechanism | 2026-09-06 | 2026-09-07 |
|---|---|---|
dns_aid_v1 |
741 / 125,929 | 0 / 0 |
dns_aid_bare |
128 / 125,944 | 0 / 0 |
mcp_dns |
164 / 125,971 | 0 / 0 |
dns_aid_v2 |
6 / 124,783 | 0 / 0 |
llms_txt |
6,162 / 106,282 | 2,268 / 33,664 |
openapi |
182 / 106,802 | 57 / 34,083 |
The HTTP mechanisms are unaffected and their lower 09-07 figures are just a
partial day. probed: 0 rather than a depressed rate is the tell, and it is
the reason this is recoverable as a reading even though the data is not: a
mechanism that reports nothing probed is visibly absent, not quietly low. Had the
resolver been slow rather than unreachable, the same fault would have produced
plausible non-zero figures and no way to tell.
Nothing is backfilled and nothing can be. The queries were never answered, so
services/cmd/replay has nothing to re-derive from. crawl_run.code_version
already records the discontinuity on both sides, per ADR-004.
Windows differ by environment, because each was switched separately: dev from its
2026-09-06 apply, prod from 2026-09-07 08:28:53Z — the creation time of prod's
resolver auto-scaling group, which is the same apply that set the probe's
DNS_RESOLVER_MODE=recursive. Both fleets were confirmed at zero healthy targets
and 51 launches for a desired capacity of 2 at the time of writing, so both were
affected from their respective switch onward.
Read agentcensus.io's own 09-07 row with that in mind, because it does not yet
show a gap and the reason is arithmetic rather than reassurance. Measured at
2026-09-07 against the published series, prod reports dns_aid_v1 169 / 28,580 and
mcp_dns 43 / 28,569 for 09-07, against 1,344 / 248,240 and 251 / 248,250 for
09-06 — a partial day in the same proportion to the HTTP mechanisms as the day
before it (a DNS-to-HTTP denominator ratio of 1.55 against 1.51), so nothing in
that row is depressed. That is because the partition is dominated by the hours
before 08:28Z. A split day is exactly the "plausible non-zero figure with no way
to tell" case this entry warns about, and it means the public series' first
visibly affected partition is 09-08 rather than 09-07. Do not read prod's 09-07 row
as evidence prod was unaffected.
The dev table above is dev's own measurement and is reproduced from the change that took it (#1417); dev's API is not reachable from outside the VPC, so those figures were not independently re-measured for this entry. Prod's were, at the time and against the published endpoint.
The gap is wider than the day the fix merged, and a caller should not infer the
end date from that merge. The fix landed on main on 2026-09-07, and its dev
apply failed before repointing anything: removing the old target group and
repointing the listener away from it cannot happen in one apply, so AWS refused the
delete and the whole apply aborted (ADR-068 §2b). Dev therefore stayed in the state
described above — and prod, which had never applied the fix at all, likewise. The
gap closes at the second apply in each environment, and this entry will not carry
an end date until each is observed serving. Until then, treat probed: 0 on those
four mechanisms as ongoing rather than historical.
/v1/status gains a fourth data-quality status, nodataTwo things a caller was relying on change. dataQuality[].status can now be a
value that was not in its enum, and the rows counted by warn are a smaller,
different set. No dataset_version moves and no census figure changes:
data_quality_check is the pipeline's report on itself, not part of the
published dataset.
analytics/30_quality_checks.sql judged all twenty-one checks with one CASE
whose first arm read WHEN observed IS NULL THEN 'warn'. A check whose input
partition was simply absent therefore reported the same status as a check
sitting at 84% of its threshold — in the only place either is reported. Those
want opposite responses: watch the second, go and look at the pipeline for the
first.
| status | means |
|---|---|
fail |
the check ran and its figure is over the threshold it asserts |
nodata |
new. The check's input was absent — no partition for the day, or an aggregate over an empty set. Evidence about the pipeline, not about the census |
warn |
the check ran and is within 80% of its threshold |
pass |
the check ran and the assertion holds |
What a caller has to do. Add a fourth branch. A client that switches on the
three old values and treats the default as an error will see nodata; one that
counted warn as "near breach" was over-counting, and the correction shows up as
warn falling and nodata appearing on the same day. nodata is not a mild
fail and must not be rendered as one — it says nothing at all about the census.
The order of the array changed too, and it was wrong before. /v1/status
sorted with ORDER BY status DESC, which sorts the status text: descending
order over the three old values is warn, pass, fail, so this endpoint
published its own failures underneath its own passes. The array is now
ranked worst first — fail, nodata, warn, pass — and the rank is explicit
rather than incidental to the spelling of the values.
nodata above warn on the same reasoning the alarms' treat_missing_data = "breaching" applies one layer up: a check that stops reporting has not passed.
History was relabelled, deliberately and exactly. Every stored warn row
with a NULL observed was a nodata row written under the old vocabulary, and
that is a biconditional — the 80% band and the breach arm both require a non-NULL
observed to be reached at all, so no other route existed to that pair.
Migration 0122 is the boundary; /v1/status serves the latest row per check,
so the relabelling is visible immediately rather than from tomorrow's run.
dataset_version 9.0: a redirect off the domain is not that domain's publicationNo field changed shape. outcome = 'miss' now includes a case it previously
excluded, at the ten mechanisms fetched over HTTP, and adoption falls on this
date with no change in what anybody publishes. This is the second correction
announced here that moves a figure down, and the second time the reason is
the same one stated differently: a response that is not about the thing we asked
is not evidence for it. In 3.0 the answer was not about the mechanism. Here it
is not about the domain.
The rule is ADR-062 §8; migration 0120 is
the boundary.
What was measured, on our own estate. agentcensus.app publishes nothing at
all — it is a CloudFront Function that 301s every path to agentcensus.io
(Decision #21). The census held it as a publisher of ten mechanisms:
ag_99963ee19c53, registrable_domain |
agentcensus.app |
| mechanisms recorded | a2a, a2a_alt, ans, ard, did_web, mcp, openapi, api_catalog, agents_txt, llms_txt |
| documents actually served by that domain | 0 |
Ten mechanisms attributed to a domain that serves one redirect, in a census whose headline figure is distinct publishing domains. Parked domains, brand aliases and country-code shadows of a main site are not a rare shape on the internet, and they are not evenly distributed either: a company large enough to publish an agent card is a company large enough to own the aliases.
Why it happened. internal/safefetch follows redirects deliberately,
re-validating every hop against the address filter and stripping credentials,
which is right. It has always recorded where it ended up, in
Response.FinalURL — and nothing had ever read that field. probeHTTP and
probeText built the observation from the target they were handed, gave the
parsers the domain they asked for, and stamped that domain on every snapshot.
The redirect was invisible to both. One place in the repository already got this
right: doFetchKeySet refuses a redirected DNS-AID key set outright, "because
the redirect target chooses which keys verify the publisher." The same sentence
holds one step weaker for a discovery document — the redirect target chooses
what the domain appears to publish.
What a discounted hit looks like. outcome = 'miss' with
error_class = 'offsite_redirect', agent_count 0, enumeration_basis cleared,
no agent snapshot and no stored payload. miss deliberately and never error:
nothing failed — we asked, a host answered, and the answer was somebody else's.
The row is never dropped and the attempt is never skipped, so the denominator
does not move: the domain was probed and stays probed, and only the numerator
changes.
The document is not reattributed to whoever served it, which is the
obvious-looking fix and is worse on two counts. It would write observations for a
domain that was never probed, so that domain would accrue mechanism hits while
its own probed denominator never moved — the incoherence ADR-011 exists to
prevent. And it would be an opt-out bypass: a domain on the shared crawl opt-out
register could have its documents entered into the census via any third party
willing to redirect to it, which makes the register advisory rather than binding.
So an off-domain redirect is a miss for the domain we asked, and nothing at all
for the domain that answered.
A redirect that stays inside the registrable domain is untouched. www. to
apex, http to https, a path rewrite, a regional subdomain — all still hits.
The unit of attribution is the registrable domain, everywhere, and the
overwhelming majority of redirects on the internet are of that kind. The DNS-AID
cap locator is also untouched: it is defined to be able to name a third
party's URL and is verified by digest against the record rather than by origin.
Two limits, stated rather than glossed. First, how large the affected
population is cannot be queried from this dataset. Rows written before today
never recorded FinalURL, so which of them redirected off-domain is not
derivable from the warehouse at all — it is unrecoverable, not merely uncounted.
Second, history is not restated: every row keeps the values it was written
with, and affected domains correct themselves as recrawl-sweep reaches them on
its rolling ~14-day window. No census_publication restatement row is owed under
ADR-018 §6, because nothing is recomputed — only rows written from this date
differ.
dataset_version 3.0: a host that answers everything is not a publisherNo field changed shape. outcome = 'miss' now includes a case it previously
excluded, at nine mechanisms, and adoption falls on this date with no change in
what anybody publishes. This is the first correction announced here that moves
a figure down, so it needs saying plainly: nothing was lost, and the numbers
before this date were larger because they counted answers that were not about
agents.
The rule is ADR-062; migration 0104 is
the boundary. Two negative controls have been in the probe, and described on
/crawler.html, for months — a well-known path we invented and a hostname we
invented, both asked so that a server or a zone answering everything can be
told from one that publishes something. Both had a hole.
HTTP: the comparison was bytes (#992). *.pages.dev answers every well-known
path with a Cloudflare request-echo — valid JSON, served as
application/json, carrying a per-request round-trip time — so no two responses
hashed alike and the control matched nothing. clientTcpRtt was 10 on the control
path and 13 on /.well-known/agent.json, two seconds apart: the same document
every time and never the same bytes. Over the seven days that made it visible in
dev, 96.2% of ard, did_web and ans hits parsed nothing, and the top hosts
hit 8 of 8 mechanisms on the same day. The control now keeps a shape beside its
hash — the sorted names of the body's top-level keys — and a hit matching either
is discounted. Keys only, because the echo's own length moves with the values
inside it.
DNS: one surface asked no control at all (ADR-062 §1). dns_aid_bare's hit
condition is that catalog.<domain> or index.<domain> resolves. Both are
generic English words, which is exactly what a catch-all answers, and the
mechanism credited every parking page and every platform wildcard as a publisher
— while mcp_dns, the other presence-only DNS surface asking the same question of
the same zones, had asked a control for months and recorded those zones as misses.
The two mechanisms disagreed, and the disagreement has a number on it.
Read in production on 2026-09-02, before the control existed:
publishingDomains, all mechanisms |
22,879 |
of which dns_aid_bare |
20,891 (91.4%) |
dns_aid_bare hit observations / distinct domains |
41,179 / 20,891 = 1.97 |
sampled domains where a nonsense label returned the same addresses as catalog./index. |
8 of 8 |
1.97 of two possible labels is the signature: both generic names answer for nearly
every domain the mechanism credits, which is what a catch-all does and not what a
publisher does. The sampled addresses were Sedo and Bodis parking ranges,
workers.dev, pages.dev, and shared-hosting catch-alls. The headline adoption
figure was 91% a measurement of wildcard DNS, and it existed for about eight
hours before anybody looked at it. Nothing alarmed: the twenty-check
data-quality suite ran clean at 5 warn, 0 fail, because
coverage.hit_without_agents — the check that found the HTTP half of this —
cannot see these rows at all, since dns_aid_bare sets a nonzero agent_count
from the address count.
What a discounted hit looks like now. outcome = 'miss' with
error_class = 'wildcard_response', agent_count 0, enumeration_basis cleared,
and no agent snapshot. miss deliberately and never error: nothing failed — we
asked, the host answered, and the answer was not about the mechanism. The row is
never dropped and the attempt is never skipped, so the denominator does not
move: the domain was probed and stays probed, and only the numerator changes. A
control that could not be asked — a timeout, a SERVFAIL, a refusal — is not a
wildcard, because unknown never becomes an accusation.
And a hit that survives the control now reports no agents (ADR-062 §7, the
amendment taken on the same day to settle the one question §6 had deferred).
dns_aid_bare set agent_count = len(addrs): the A records at a generic hostname
were published as agents, and a publisher who kept a second address for redundancy
was read as publishing a second agent. It is 0 now, which is what mcp_dns — the
same presence-only shape, asking the same question of the same zone in the same
pass — has recorded from the start. agent_count is agents parsed from this
response and a set of addresses parses to nothing; presence is evidence that a
name exists and is not a count of what is behind it. Hit counts and the
denominator do not move for this half — only the agent count attributed to that
mechanism, which was never a count of agents. It rides 3.0 rather than waiting for
a 4.0 because it is the same correction in the same direction on the same rows, and
a caller who has to consult two versions to learn what one mechanism's count means
is worse served than one who consults none.
The request cost does not move. The DNS control is still one extra lookup, now
on the minority of zones where any of mcp., catalog. or index.<domain>
answered. It is one and not three because the control is a question about the
zone: both presence-only surfaces share a single answer per target, whichever
reaches a resolving name first pays for it, and a zone with none of the three
names is not asked at all. /crawler.html states that rule and a test asserts it.
Two limits, stated rather than glossed. First, how large the wildcard
population is cannot be queried from this dataset. The control emits no
observation and its answer is recorded on no row, so "how many domains in the
denominator answered a control they should not have" is unanswerable — which is
exactly the number ADR-011 wants before a numerator correction is disclosed.
ADR-062 §2 fixes that going forward and is unbuilt; until it lands the population
is described in the census caveat and not counted. Second, history is not
restated. For a row already demoted the payload was dropped by design, and for a
historical hit the control's answer that day was never written down, so there is
nothing to re-derive from. Affected hosts are reached again on recrawl-sweep's
rolling ~14-day window and record the corrected classification when they are.
A prod-only mislabelling, fixed in the same change and worth knowing about
separately. The API selected the current dataset version with
ORDER BY effective_from DESC LIMIT 1, and every migration inserts
effective_from = current_date. On prod, where all migrations ran inside one
apply, every row tied and Postgres returned an arbitrary one: prod reported
1.1 — the second version ever declared — and stamped it on everything it
published. Dev, which accumulated its rows over weeks, was right by accident. The
read is now tied on created_at. A prod figure carrying dataset_version 1.1
is mislabelled, not computed under 1.1 semantics.
3.0, and not 2.1 or 2.0. 2.0 was claimed by migration 0090 for the
spec-conformance corrections, and this is a different statement about a different
thing: not "the probe stops asking for what no specification defines" but "a
response that is not about the mechanism is not evidence for it". Numbering is
once per movement rather than once per change, which is why the series is not
monotonic — 0090 took 2.0 and 0091, 0097 and 0098 then took 1.5, 1.6 and
1.7 after it.
publishingDomains is a count of domainsThe field did not change shape and its name did not change meaning. Its value was wrong, and correcting it moves the most quotable number on the site down by roughly half. Announced here rather than shipped quietly, because a caller who recorded the old figure has recorded something else.
adoption_daily is keyed on a mechanism: every grain='total' row is one
mechanism's total. GET /v1/census/headline and GET /v1/overview answered "how
many domains publish anything" by summing domains_publishing across those rows,
which counts a domain publishing three mechanisms three times. The figure was a
count of publishing signals, returned under a name that says domains,
rendered as domains on the census page, and described as domains in /llms.txt's
entry for census_headline. On 2026-09-01 it read 11,517 against a probed
population of 22,861.
The denominator in the same response was never wrong — it was max(), not a sum
— which is why nothing about the pair looked broken. A numerator on a different
population than its denominator is ADR-011's defect from the other side.
No gate could catch it. ADR-018 added
CHECK (domains_probed >= domains_publishing), enforced per row, and per row
it held. The defect was in a sum across rows, and a row constraint cannot see a
sum.
What changed. 20_adoption_daily.sql and 45_backfill_stratum.sql now count
the union where domain identities still exist, into two new adoption_daily
columns (migration 0100): distinct_publishing_domains and
distinct_probed_domains. Both are carried identically on every grain='total'
row for the day, because a cross-mechanism figure has no mechanism to key on —
so they are read with max() and never with sum(). Migration 0100 adds a CHECK
that the union is at least as large as the row's own mechanism count, which is a
per-row statement of the cross-row property that went unguarded.
dataset_version does not change. No field's meaning changed and no row in
the published dataset means anything different; an aggregation in the API was
wrong and is now right. Days computed before migration 0100 fall back to
max(domains_publishing), a strict lower bound on the union where the old
sum was an upper one — so historical days understate rather than overstate.
The per-mechanism figures are unaffected. /v1/census/mechanisms and the
totals block of /v1/census were always per-mechanism counts against a stated
denominator, and every tile on the census page still reads exactly as it did.
Whether the signal count is worth publishing alongside the domain count, under a
name that says what it is, is a separate open question (#1038 step 2).
dataset_version 2.0: mcp and mcp_dns mean more than they didNo field changed shape. Two enumerated values changed meaning, which is the
change a schema version cannot show you. mechanism = 'mcp' and
mechanism = 'mcp_dns' now include cases they previously excluded, so a query
that spans 2026-08-30 is comparing two definitions of what it means for a domain
to publish those mechanisms. Both counts rise, and the rise is this instrument
rather than the world.
What the probe was doing wrong. Four places where it cited a specification and then did something the specification does not describe:
/.well-known/mcp.json is defined by no specification. Not by the MCP
specification — the 2026-07-28 release defines no well-known discovery path at
all, and its discovery addition is a server/discover call, which is a call
and therefore something this crawler does not make. Not by SEP-1649
(/.well-known/mcp/server-card.json), not by SEP-2127
(/.well-known/mcp/server-cards.json), and not by
draft-serra-mcp-discovery-uri (/.well-known/mcp-server). The last two were
cited in our own code and on /crawler.html as defining it.version, where the draft requires
mcp_version._mcp.<domain> TXT and _mcp-key.<domain> — while asking mcp.<domain> for
addresses and never for text, at a name where zones publish the record.url. A2A v1.0 removed url,
protocolVersion, preferredTransport and additionalInterfaces in favour of
a single supportedInterfaces[], so a card written to the current
specification had no endpoint this probe could see — and, because the parser
gates on the card's name, still recorded as a clean hit with one parsed
agent. A row asserting the card was read, with the endpoint missing from it.What changed. /.well-known/mcp-server is probed beside /.well-known/mcp.json,
the two distinguished by label_variant. _mcp.<domain> TXT and the TXT at
mcp.<domain> are read, as two further mcp_dns label variants. Manifests are
read under both version spellings. A2A cards fall back to
supportedInterfaces[0].url for their endpoint and to that entry's
protocolVersion for their declared version, and supportedInterfaces is a
recognised card shape, so a v1.0 card no longer infers 0.3.
Three per-agent fields move for A2A cards written to v1.0.
source_identifier becomes the real endpoint instead of the invented
a2a://<domain>/<name>, which changes which rows entity resolution merges;
endpoint_same_origin stops being null, which changes the denominator of the
supply-chain statistic rather than any single answer in it; and
inferred_version reads 1.0. field_completeness moves for every A2A card in
both directions, because the optional-field list gained the 1.0 spelling. Each
of those is in content_hash, so every affected agent mints one new SCD-2 row on
its next probe — a one-time discontinuity, recorded as one.
a2a and a2a_alt still mean exactly the paths they always meant, and are
not renamed. They read backwards: a2a is the pre-0.3
/.well-known/agent.json and a2a_alt is the /.well-known/agent-card.json
that A2A registered in v0.3.0 and kept in v1.0. Renaming them would make one
identifier mean two things either side of a date, which is the failure
dataset_version exists to prevent rather than a use of it. What did change is
the ADR-010 precedence order: the registered path now outranks the pre-0.3 one,
so a domain serving both during a transition has the current card's name,
description, endpoint and version win the conflict instead of the stale copy's.
No count moves; a resolved agent's attributes can.
Nothing is backfilled and nothing can be. The paths and names were never
fetched, so services/cmd/replay has nothing to re-derive from. Prior waves are
neither re-probed nor re-stated, per ADR-004, and the discontinuity is annotated
rather than smoothed over, per ADR-011 §4.
agent.activeVerification is nested by passOne field, one level of nesting, and it is a disclosure fix rather than a feature. The per-agent active-verification summary — the only place an owner's result reaches a signed-out reader — used to be a flat object:
"activeVerification": { "lastCheckedAt": "2026-08-24", "method": "initialize", "outcome": "ok" }
It is now keyed by which pass produced the result:
"activeVerification": {
"unauthenticated": { "lastCheckedAt": "2026-08-24", "method": "initialize", "outcome": "ok" },
"authenticated": { "lastCheckedAt": "2026-08-24", "method": "tools/list", "outcome": "ok" }
}
The fields inside are unchanged. lastCheckedAt, method, outcome, and
the optional negotiatedVersion, versionMismatch and authEnforced mean
what they always meant, and both keys carry the identical set. A client that
read agent.activeVerification.outcome now reads
agent.activeVerification.unauthenticated?.outcome — the same value, from the
same pass, one level down. Nothing was renamed or removed.
Either key may be absent, and an absent one is a fact. A pass with no
published scheduled row simply has no key; the object itself is null when
neither has one, exactly as the flat field was null before. A client must not
read authenticated as a fallback for a missing unauthenticated, and this is
the whole reason the shape changed rather than the reason it is inconvenient.
What went wrong. ADR-021's 2026-08-25
tier-1b amendment added a second pass: the same methods and the same trigger,
sent with the owner's stored credential. Both write to one table, distinguished
only by which pass made the call, and this summary took the most recent row of
either. So a credentialed ok could be published to a reader who cannot
authenticate, asserting that an endpoint answers when it may refuse strangers
outright. The row was never wrong; reading it as the credential-free pass's
answer was.
Why the authenticated pass is published at all rather than suppressed, which
was the narrower fix: an agent that refuses strangers and answers its owner is a
materially different thing from one that answers nobody, and under suppression
both read as silence. The consent gate is unchanged — a verified claim, tier 1
enabled, publishAgentResults on, and the tenancy the row was recorded under.
This widens what is said, not who may say it. It remains true of both passes
that nothing sorts, filters, facets or scores on them
(ADR-020 §1).
One behaviour change with no shape change beside it. The
ADR-047 overlay's behavior
block now reads unauthenticated rows only, on the agent record, on search
results and on the org-plane agent list. Its signals — did the agent answer,
does it enforce auth, does the negotiated version match — mean something
different when the caller held a credential. That surface is account-gated, so
this is correctness rather than disclosure, but it is the same conflation, and
a caller may see a different row than they saw last week where a credentialed
check was the most recent one.
/v1/orgs/{slug}/entitlements returns an object, not an arrayOne endpoint, one shape change, and nothing else moves. The org-plane entitlements read used to answer with a bare JSON array of entitlements. It now answers with an object:
{ "entitlements": [ ... ], "bundles": [ ... ] }
entitlements is exactly the array that used to be the whole body — same
fields, same order, same meaning. A client that read body[0].capability now
reads body.entitlements[0].capability; nothing was renamed or removed.
bundles is derived, never stored, and gates nothing. A bundle is a name
over capabilities — "Premium — agent insights" is one product name over three
of them — added so a platform admin can grant a purchase in one audited action
instead of flipping three switches in the right order
(ADR-028 §5 as amended 2026-08-17,
ADR-036 §2). Each entry carries the
capabilities it names and a state of complete, partial or none, computed
from the rows on every read. An expired row does not count towards a bundle —
the derived state has to agree with what the gates do, and
tenant_has_capability() applies the same expiry — while the row itself stays
in entitlements with its expiresAt, so a lapsed grant is still
distinguishable from one that was never made. No authorisation decision
anywhere reads a bundle name, and a client that ignores the field entirely
has lost no information: the capability rows are still the whole of what a
tenant holds.
Why the shape changed rather than a second endpoint being added: a separate
/bundles read would be a second answer to "what does this organisation have",
and the first thing a second answer does is disagree with the first.
The admin write, PUT /v1/admin/orgs/{slug}/entitlements, returns the same
object for the same reason. It is otherwise unchanged, and it is still the way
to grant one capability on its own.
This is a reversal, not a feature. The entry below, dated 2026-08-11,
announced that /v1/search, /v1/agents/{agentKey} and /v1/domains/{domain}
had moved behind an account and been scoped to the domains that account's
organisation had claimed. That is undone.
ADR-040 supersedes
ADR-028 §1–2 and §7 for those three
routes, and it is recorded as a superseding ADR rather than a quiet amendment
because the thing being reversed is the tier boundary itself. Announcing it as
an addition would be the thing not to do here.
The claimed-domain scope is gone from all three routes, for every caller.
Anonymous, session or key: everyone searches every agent the census has
observed and reads any agent's or any domain's record. A signed-in caller who
has claimed nothing no longer gets 403 no_claims on them. 403 no_claims is
still the answer on the owner surfaces, where a claim still means what it
always meant — ownership, corrections, alerts, active verification. What a
claim never meant again is permission to find something.
Anonymous callers are back, and they are metered. No key, no session: 5
searches and 5 detail reads per viewer address per UTC day, counted exactly
across tasks. At the cap the answer is 401 with code: "register_required", plus registerUrl and resetsAt in the body and the
X-RateLimit-* headers on the response. It is a wall, not a slowdown — waiting
does not help and registering does — which is why it is not the 429 that
answers the per-minute window. If you branch on status alone, 401 now means
two things and code is what separates them: account_required is a route
that needs a caller, register_required is an allowance spent.
Every record is split into a published half and an observed half, and this
is the breaking part for a client reading the shape. What the subject
publishes stays where it is. What agentcensus derived by looking moves into one
additive observed object, which is absent — not null, not {} — for a
caller with no account:
Agent: status, firstSeen, lastSeen, primarySource, provenance and
history move to agent.observed. They were required fields; they are now
inside an object an anonymous response does not carry. mechanisms stays
where it is and changes meaning slightly: it is now every surface the agent
was found on, not a one-element list holding the primary source.status, firstSeen, lastSeen, primarySource and
match.similarity move to result.observed; match.kind stays out, because
how a query found a result is a fact about the query. Results also gain
limit and offset, and facets is present only for an account — absent,
not empty.DomainRecord: firstSeen, lastProbed, spans, dnsAidCheckedAt and
dnsAidChecks move to record.observed.Nothing was deleted and no field was renamed. A client that reads
agent.observed?.firstSeen works on both tiers; one that reads
agent.firstSeen gets undefined where it used to get a date.
Accounts gain a composite trust score.
ADR-041 reverses this platform's
standing refusal to emit one, under conditions that are the point of it: the
number is the integer mean of the dimensions ATD actually measured, an
unmeasured dimension is excluded rather than counted as zero, and it is never
served without its coverage — score, measured and of are required
together, and score is null when measured is 0. It appears on
/v1/agents/{agentKey}/trust as composite, and on an account's search results
inside observed.trust. It is never a sort key, a filter or a facet, on any
surface, and no census figure is computed from it. ATD still declines to emit
one; this number is ours and is derived at read time from what ATD already
sends.
/v1/agents/{agentKey}/trust is account-wide. Still never anonymous — trust
is observed information — but no longer narrowed to claimed domains, so an
account reads any agent's vector.
GET /v1/search gains offset, and page sizes are per tier. Ten results
and no paging with no account; 24–100 with an offset up to 2,000 for an
account. limit and offset are ignored rather than clamped on the anonymous
tier. status narrows results for an account only, because status is observed
information — sending it without an account is ignored, not refused.
The census:read-all key stops being a widening. It granted the platform
tenant a census-wide read of /v1/domains/{domain} while that route was
claim-scoped. Every reader now reads every domain, so there is nothing left for
it to grant. It remains the platform tenant's own credential.
A new entitlement, search_volume. Granted per organisation by a platform
admin. It raises caps and changes nothing else: no field, no facet and no
ordering differs for an organisation that holds it.
What changed. Three things a caller could have been relying on.
The 429 on the public census routes now carries code: "rate_limited", not
"quota_exhausted". GET /v1/census, /v1/census/headline,
/v1/census/labels, /v1/status, /v1/overview and /v1/example-agent refuse
a burst from one address with a window that reopens within the minute. That has
never been a spent quota, and calling it one left a caller unable to tell "come
back in a moment" from "come back tomorrow". If you branch on the code, this is
the line to change; the status is unchanged.
GET /v1/search, GET /v1/agents/{agentKey}, GET /v1/agents/{agentKey}/trust
and GET /v1/domains/{domain} now enforce a per-minute window as well as the
daily allowance. An API key had only a daily ceiling before. Over the window
the answer is 429 with code: "rate_limited" and a Retry-After header
saying how many seconds to wait. The default is generous relative to any
interactive use; a script that bursts is the thing it is for.
The trust route is metered. GET /v1/agents/{agentKey}/trust spends the
organisation's daily allowance like the other two detail reads. It was outside
the counter before, which was an oversight rather than an offer.
Every metered response now carries X-RateLimit-Limit. Alongside the
X-RateLimit-Remaining and X-RateLimit-Reset that were already there, so the
ceiling travels with the count and a client can render "3 of 5 left" from one
response instead of asking a second endpoint what 5 was. All three describe the
DAILY allowance. They are omitted rather than guessed at when no counter
measured the request — if our counter store is unreachable we admit the request
and say nothing about limits, because a header nothing measured is a lie.
Refusals from these routes carry both message and error. The same
sentence under both keys, plus code, plus resetsAt saying when the allowance
reopens. Two envelopes have existed in this API's history and a refusal that
carried only one of them reached clients reading the other as a bare status
code with the reason thrown away. Nothing is removed; message remains the one
to read, and api/openapi.yaml's Error schema still requires only it.
A new code exists that nothing returns yet: 401 register_required. It is
the anonymous daily wall of ADR-040 (Public metered search) §1.1 — a viewer
with no account has used its free searches or record views for the day. Its
body adds registerUrl (where to create an account) and resetsAt. It is
announced now, before any route answers anonymous callers, so a client can be
written against it rather than discovering it. It is a 401 and deliberately
not a 429: waiting does not help, and registering does.
The usage page's numbers change, downward. GET /v1/orgs/{slug}/usage
counted a keyed request on the search and detail routes twice — once in a
detached increment when the key was verified, once in the counter that enforces
the quota. So the effective ceiling on those routes was about half the 10,000 a
free organisation is told it has, X-RateLimit-Remaining understated what was
left, and the console over-reported. Counting now happens in exactly one place.
Expect a visible step down in the 30-day history on the day this deploys: the
older days were inflated and nothing can un-inflate them. The number now means
one thing — requests that spent this organisation's daily allowance — which is
the number the ceiling is compared against.
What did NOT change. Every route's tier: /v1/search and the three detail
routes still require a bearer token, and an anonymous request still gets
401 account_required with the announcement URL. No response body changed. The
census routes stay fully public. Signed-in sessions are now counted, which they
were not before, but no session that was being served stops being served.
Why. ADR-040 (Public metered search) §1.1 and §3. The limits substrate that will let search answer anonymous callers has to exist, and be observable, before anything is opened to them.
What changed. GET /v1/agents/{agentKey}, GET /v1/search, and
GET /v1/domains/{domain} now require a bearer token: an API key created in
an organisation's console, scoped to search:read and returning only agents
on domains that organisation has verified control of. A request without one
gets 401 account_required.
Why. ADR-028 §7. The per-agent card — Trust Vector, probe history, conformance detail — has a real cost to serve and identifies a specific agent; the census does not. Splitting the two is the whole decision, and leaving the split undocumented while the code enforced it would have been the "gate that only exists in the UI" ADR-028 explicitly rejects.
What did NOT change. GET /v1/census, GET /v1/census/headline,
GET /v1/status, GET /v1/overview, and GET /v1/example-agent stay fully
public and citable, with no key and no account. The aggregate dataset was
never the thing being gated — see ADR-028 §1.
The operator escape hatch. A gated or suppressed record stays reachable
by direct link, signed out, with the reason stated
(ADR-028 §3,
ADR-025 §1): an anonymous request for a
suppressed agent or domain returns 200 with a minimal
{ registrableDomain, recordStatus, reason, claimUrl } body instead of the
full record. An anonymous request for an ordinary (not suppressed or gated)
record still gets the plain 401. Proving control of the domain is the same
act as getting access to it.
Migration path.
Console → Keys) and send it as
Authorization: Bearer ac_live_....Every 401 these three routes return carries this page's URL, so a script
that broke against this change lands on the explanation rather than a bare
refusal.
Rate limits, unchanged by this entry but worth restating here since it ships
alongside it. Anonymous requests (the routes that stay public) are limited
by client IP. Keyed requests are limited by the organisation's quota,
enforced by the same conditional counter the console's usage page reads —
not an estimate reconciled later. See GET /v1/orgs/{slug}/usage.