Payer-report methodology
Anyone can re-run every check on this page against public chain data. We store the signatures and transaction hashes so that our verdicts are reproducible, not trusted.
A payer report is a signed, on-chain-proven statement by someone who actually paid a listed service, saying whether it delivered. It is the one signal in this directory that speaks to outcome from a real buyer — probes prove an endpoint answers correctly; on-chain volume proves people pay it; only a payer can say whether what they bought was any good. We also buy endpoints ourselves and publish every result at /stats; that proves delivery to us, where a report proves it to someone who needed it. Reports are not reviews and not an honor system: every accepted report has passed all of the checks below, and the evidence for each is stored.
The canonical message (v1)
Every report must carry payer_signature: an EIP-191
personal_sign over exactly this message
(newline-separated, no trailing newline):
nohumans.directory report v1 listing_id: <id> tx_hash: <0x...lowercase> ok: <true|false>
This is a published contract; it will only change with a version bump.
You do not need to construct it from documentation: POST a report without
a valid signature and the 400 response echoes the exact
message_to_sign for your request. The API teaches the format;
this page explains why it is shaped this way.
What is checked, and why
A real payment exists. tx_hash must be a successful
Base transaction whose logs contain an ERC-20 Transfer, on an asset the
listing's live 402 envelope accepts, to the listing's live
payTo. We read the listing's payment terms from the endpoint
itself at verification time — not from our own cached copy.
The payer is the Transfer log's sender — never the transaction sender. In x402's EIP-3009 flow the transaction is submitted by a facilitator; the wallet that actually paid appears only in the Transfer log. Any gate that read the transaction sender would verify the wrong party.
The report is signed by that payer. Transfer logs are public: anyone can harvest the hash of a payment they never made. The signature requirement means a report is bound to the paying wallet's key — the payment proves a purchase happened; the signature proves the report comes from the purchaser and not a bystander.
The payment is a real purchase, not dust. The transferred amount must be at least 90% of the listing's observed live price. Without this, a fraction of a cent would mint a "distinct payer".
The payment is recent and relevant. The transaction must be no more than 30 days old and must postdate the listing itself.
One report per transaction, one voice per payer. Each transaction hash can be reported exactly once, and a wallet's latest verified report supersedes its earlier ones. N payments are not N votes.
How reports enter ranking
Reports affect a listing's ranking only once 3+ distinct payers have reported on it, and then at 25% weight alongside the 75% probe-based score. Below that floor, reports are displayed but cannot move a listing — a single voice, however loud, is not a signal.
An economic property worth stating: faking a positive report requires actually paying the listing, which enriches the target of the fraud. Negative report-bombing likewise costs the attacker real money — paid to the victim.
Coverage and known gaps
Externally-owned accounts sign reports normally, including EIP-7702
delegated EOAs. ERC-1271 contract wallets (smart accounts) are not yet
supported — personal_sign recovery cannot verify them.
This is an open gap we state rather than hide; support is planned.
Our own wallets never count
Two wallets belong to this directory's operator: the scout
(0x54E163e9B8eDDa194D83F46AdD921bfA5fc5f4E0, which makes every
paid-verification purchase and announces itself with an
X-Verified-By header) and a test wallet
(0x2b5586cA1E28b6310Ada60d03Cd432493bEA2C49). Payer reports
filed from either are recorded for instrument history, tagged at ingestion,
and never displayed, never counted toward the 3-distinct-payer floor, and
never weighted into any ranking. Both addresses are public precisely so this
exclusion is checkable against the chain rather than taken on trust.
What the reputation score is
The 0..1 score on every listing is a recency-weighted probe
pass rate: computed only from this directory's own probe outcomes for that
listing, with newer probes weighted above older ones. It is re-derivable
from the probe history published on each listing page. It is a measurement
with a stated formula and input set — not an opinion, not editable by anyone
including us, not for sale, and deliberately never blended with paid
verification, payer reports, or on-chain activity, which are different
evidence classes answering different questions.
First-party listings: none, by policy
As of 2026-08-25 this directory lists nothing its operator sells. All
17 remaining first-party listings (labeled first_party) were
retired that day — reason first_party_retired in the changes
feed — after operating here since launch. While they existed they were
excluded from paid verification and their probe history ran identically
to everyone else's; their records remain visible as delisted history. A
directory operator should not be a seller in its own index.
One exception, stated so it cannot be mistaken for a listing.
Since 2026-09-15 a single fixture record (fixture-0001, origin
fixture) exists so the owner-only paid route, verify-now, can be bought
end to end in tests — it is the only way to prove that a seller who pays $3
gets a purchase queued, made and published. Its endpoint is a third party’s
cheap public route; it is not ours and sells nothing of ours. It is excluded from
discovery, the listing feed, export, both sitemaps, categories, resolve, the changes
feed and every public count, is never ranked, and holds no badge anyone can see. The
first test purchase through it found the runner marking a plan done with nothing
bought; that is why it exists.
Independence: what we will and won't do
Ranking is never for sale. No listing's position, status or score can be bought, and no arrangement of any kind will ever move one. The submission fee (below) buys a listing's entry, never a place in the order. The moment ranking can be paid for, every verdict on this site is worthless — including the ones that would still be honest.
Submission fee
Submitting is free for the first 5 listings per registered domain (api.example.com and www.example.com both count as example.com; on shared hosts such as vercel.app or workers.dev, each app's or account's own subdomain counts as its domain). Each further new listing costs $0.25 USDC on Base (since 2026-09-28; $0.50 before), paid over x402 on the same POST: over the allowance, an unpaid POST answers 402 with the terms, and the paid retry returns 201. The submission is validated before any charge, so a rejected submission is never charged. Edits, claims, an endpoint that is already listed, and re-listing an endpoint that was listed before are free. Listings already in the directory count toward the five and never pay. The fee buys the entry, never a status, a score or a place in the order. Since 2026-09-25.
Why: the fee keeps the index to endpoints someone stands behind, and
paying it over x402 shows the submitter’s own x402 client works. A fee-paid listing is probed, scored and ranked by exactly
the same rules as every other listing, and it earns or loses
verified the same way. Fees appear as a sales line in
/v1/stats (submission_fee_sales).
We publish; we do not partner with anyone we measure. Sellers listed here get the verdicts and nothing else: no joint reports, no data arrangements, no private exchanges of figures. This is not standoffishness — it protects a seller's good score as much as our neutrality, because neither can be explained away as a favour. Everything is CC-BY: replicate it, compare it against your own measurements, publish against it. No arrangement is needed for that, and none is available.
We publish findings, corrections, and enough method to check them — not our build. Every figure here states its source, its window, its controls and its limits, and the underlying facts are public chain data and public endpoints, so any of it can be checked independently. How the instrument is engineered is not part of that and is not published. Saying so plainly seems better than trimming quietly.
What we do not verify
Everything this directory verifies is endpoint behavior — correct 402s, live samples, stable payment destinations, real payments, and signed payer outcomes. We do not verify content accuracy beyond structure, the identity of whoever operates a domain today, or whether a payment-address rotation was legitimate (we detect it and reset trust; we do not judge it). A behavior-based verifier can be passed by a compromised domain that keeps behaving; that is a structural limit, not an oversight, and it is why every signal here says exactly what it measures and no more.
Reference client
Producing the signature from a private key (sign_report.mjs in the repo is the complete version):
import { secp256k1 } from "@noble/curves/secp256k1";
import { keccak_256 } from "@noble/hashes/sha3";
const msg = `nohumans.directory report v1
listing_id: ${listingId}
tx_hash: ${txHash.toLowerCase()}
ok: ${ok}`;
const prefixed = "\x19Ethereum Signed Message:\n" + msg.length + msg;
const digest = keccak_256(new TextEncoder().encode(prefixed));
const sig = secp256k1.sign(digest, privateKey);
// payer_signature = 0x + r + s + v(27/28)
Reusing this data
The directory data — listings, statuses, scores, probe outcomes and paid-verification results — is published under CC BY 4.0. Copy it, republish it, build on it. The only condition is attribution: credit nohumans.directory with a link to https://nohumans.directory.
Bulk access is free and needs no key:
GET /v1/export returns the whole catalog as NDJSON and
GET /v1/changes is the status-transition feed. Both state the
licence in the response. We would rather you took a clean copy from the
endpoint than scraped the pages.
A verdict here is a claim about behaviour observed at a point in time. If you republish one, say when it was observed — a status copied today and shown as current next month is our name on a number we no longer stand behind.
Corrections
A correction is a figure or claim we published that was wrong: what it was, what the right one is, and when it changed. Twenty-five since launch on 2026-07-24. The count is high because we point our instruments at each other — the probe log against the resolver, the verification badge against the paid census, the demand feed against the edit log — and publish whenever they disagree. Each seam gets checked once; the number of new ones should fall, and this line will say if it does not. Corrections stay on this page rather than being folded into the numbers. Rules that were tightened without a published figure being wrong are listed separately under Method changes.
2026-09-24 — descriptions of the free demand feed stated a delay and a
floor it no longer had (correction #25). The free feed at
/v1/demand has ended a week behind live since the paid tier was
switched on on 2026-09-11 (a day when it is off), and since 2026-09-12 09:54 UTC it
publishes a term after one searcher. Our llms.txt said “24h
delay”; the MCP tool what_agents_are_asking_for said the window
“ends a day behind live” and that a term needs 3 distinct client IPs; the
OpenAPI description of /v1/demand gave the same floor of 3. Correction
#22 fixed the floor on this page and said the machine-readable side had never been
wrong; the MCP tool and the OpenAPI description are machine-readable, and both said
3. The window was also
described as filling “over thirty days”: with the delay it is complete on
2026-10-19. And /privacy said terms logged between
2026-09-11 15:00 and the 09-12 amendment “stay under the two-searcher
rule”; the feed does not publish them at all. The feed itself always reported
its real delay_hours and k_anonymity_floor. Every
description now reads the floor and the delays from the constants the feed runs on,
and our daily check fails if a description disagrees with the live feed.
2026-09-23 — the paid demand feeds said thirty days and read fewer
(correction #24). GET /v1/demand/clusters (on sale since
2026-09-13) and GET /v1/demand/report (on sale since 2026-09-23)
grouped at most the 2,000 newest searches and labelled the result a 30-day
window. The 30-day window has held more than 2,000 searches since 2026-08-31, so
every sale of either route read less than it said: purchases on 2026-09-19 and
2026-09-20 reached back to 2026-09-05, about fourteen and a half of the thirty
days, and on 2026-09-23 the feed reached back to 2026-09-10. Fixed the same day:
a build now reads up to 10,000 searches, and every result states the oldest
search it read (window_starts_at) and whether the whole window fit
(window_complete); the report's page says “since” a date
instead of “over 30 days” whenever it did not. Both purchases of the
clusters feed by buyers other than us were refunded in full: 0x0d33729b…,
0xe3ad3416….
2026-09-23 — how many clean probes earn verified was
stated two ways, and each was wrong for part of the catalogue (correction #23).
Our agent documentation (llms.txt and the MCP server's instructions) said the badge
takes roughly sixteen consecutive clean probes, “not three”. The
sellers' FAQ said a listing whose payment address changed re-earns it on three
consecutive passes, in minutes. Both follow from one rule applied from different
starting points: a listing's score starts at the result of its first counted
probe and moves 10% toward each later one. A new listing whose first probe passes
starts at 1.0 and needs three clean probes: 279 of the 291 new listings verified in
the seven days to 2026-09-23 took three or four. A listing whose score starts at
zero — after a payment-address reset, a recovery from failing, or a failed
first probe — needs about sixteen: 22 of the 24 payment-address resets in
the same week took 11 to 20, and the 747 recoveries averaged 15. All three
documents now say which case is which. Also today: the probe tier that re-checks
a listing on every run now covers an unverified listing on a clean streak, so a
reset or recovering listing climbs to the badge in about half an hour rather than
hours (Method changes).
2026-09-22 — the last statement of the k-anonymity floor on this
page said two clients; it has been one since 2026-09-12. The 09-11 entry
below records the floor moving from three distinct clients to two. It was
amended again the next day: the two-client floor withheld 46 of 47 terms,
because agents phrase the same need differently and almost nothing is asked
twice verbatim, and what it withheld were generic category phrases that
identify nobody. One client has been enough since. The feed's
k_anonymity_floor field, the feed's own note and the code have all
said one throughout — the machine-readable side was never wrong. The
amendment was recorded in the source and never written onto this page, so the
most recent statement a reader could find here said two. The floor is one
distinct non-seller client IP; the protection is the content filter plus the
withheld-terms list, both unchanged. No published count changes. Found by
reading our published claims against the code that implements them, which the
same morning also produced a self-inflicted one: between 11:50 and 16:00 UTC
the feed briefly enforced a two-distinct-day rule on top_needs,
publishing 5 terms instead of 43 and adding a
needs_not_yet_persistent field, on the mistaken belief that a rule
described in older prose had never been implemented. It had been removed
deliberately on 09-11. Reverted the same day; the rules are unchanged and no
term was ever withheld, only moved into a second list. Twenty-second
correction.
2026-09-14 — 429 counted as a failure; 620 failing verdicts in seven days
were our own request rate. From 08-27 to 09-13 the prober treated HTTP 429 as a failed
probe. Five consecutive failures mark a listing failing, and failing listings get
scheduler priority, so a host that throttled us was probed harder and throttled more. In the
seven days to 09-13 18:30 UTC: 35,302 of 2,002,591 probes drew a 429 (1.76%); 620 transitions
to failing across 7 hosts and 52 listings had nothing but 429s in the five probes
behind them; 505 of those were agentdatum.com, 21 listings cycling in and out of
failing about two dozen times each, at 200 requests an hour from us of which 92%
were 429. A comparable host at 53 requests an hour returned none. Shipped 09-13: a 429 is
warn:rate_limited — recorded, excluded from the published pass rate, neither a
failure nor a pass for status — and a per-host cap of two requests per tick. In the 24 h
after, 91 transitions to failing had no five-429 run behind any of them. The 09-13
fix covered the GET only: a 429 on the POST retry was filed as not_payment_gated
until 09-14 (512 probes, 6 listings, no run older than two days), and now carries the same
label. The worked example, api.osf-master-server.com, is at
/state/controls/osf-429neutral.csv: under the
corrected rule 18 of its 19 “mixed” days read 0.96–1.00 and one, 08-26, reads
0.363. A peer running one request a day drew 429s from a single host on 35 of 37 days; at that
rate the same code is the host’s policy toward anonymous callers, which is why the label
names neither party.
2026-09-09 — we attributed a thirteen-point fall in the paid census to
one cause; it has three. Report #2 §1 said the 245 pre-flight-verb failures in
wave 7 were “why” the pass rate on payment-required endpoints read 41.4%
instead of 54.8%. On the same figures: removing those 245 from the denominator gives
45.9%; crediting them at the wave’s 89.8% delivery rate gives 50.3%. The verb
explains at most nine of the 13.4 points. The rest is the guard — 452 endpoints
refused before payment because their 402 declares parameters the scout cannot supply,
the largest outcome class after delivered, a rule written after wave 5
spent real money on 233 blind 400s — and the fact that supply selected on another
party’s evidence delivers less than supply that submitted itself. The paragraph is
corrected in place with a dated note; no figure changes. Eighteenth correction; found
the morning after publication, by reading our own sentence against our own table.
2026-09-06 — 48 listings on 25 hosts were failing for
validating their input. Since 2026-08-26 a route answering 400/422 on both
verbs has been labelled warn:params_likely_required with the stated
intent of not calling it broken — but the label was still a failed probe,
and five in a row set failing. Few listings tripped it until the
Bazaar import of 2026-09-05 brought in routes that validate a path parameter
before challenging; 33 of them on 13 hosts (omniterminal.app, grov.fun,
gateway.stride20k.com, agentic-jp.com ×3, and eight others) reached
failing at 19:49 UTC that day and stayed there about twelve hours.
The same evening, 15 more on 12 hosts: route templates (/tx/:hash,
/company/:number) that answer 404 to the literal placeholder, which
the absence rule of 2026-09-04 read as “no such route” — it is
“no such hash,” the same validation-before-payment in a third
costume. Fix: the class is neutral from today (Method changes, 2026-09-06), a
404/410 to a placeholder on a templated URL joins it, and the 48 are set back
to unverified with reason correction_17 in the changes
feed. They cannot verify until we probe them with a valid input —
that is the next prober item, and several of these sellers print the example in
their own error body. Seventeenth correction; the label was right and the
grade behind it was wrong.
2026-09-06 — we attributed $12.9M of USDC burns to one listing for
ten hours. A listing imported from the Bazaar on 2026-09-05 (19:05 UTC)
names the zero address, 0x000…000, as its payTo. The
on-chain watcher accepted it — it checked that a payTo was a 42-character
hex address, not that anyone could receive at it — and from the next tick
attributed every USDC transfer to the zero address on Base, which is where
burns go, to that listing: about $4.3M on 09-05 and $8.4M in the first hours of
09-06, from four senders. That listing’s public record carried those figures
as onchain_30d until 05:45 UTC on 09-06; the aggregate economy line
in our own morning check inflated by the same amount but is not published.
Fix: the watcher and the rollup now refuse sentinel addresses (zero, dead, and
the all-zero prefix), the eight index rows were removed, the rollup rewrote the
record with no on-chain figures, and the verdict for any listing whose live
challenge names a burn address is avoid with the address printed
— a well-formed 402 nobody can pay is worse than a broken one. Sixteenth
correction; an address is not a payee.
2026-09-03 — eleven listings were delisted in August for a limitation
of our prober, recorded as a fact about them. Between 2026-08-03 and 08-07 we
admin-delisted eleven listings with the reason “POST-only endpoint; probe
architecture only checks unpaid GET, this can never verify” — among
them api.venice.ai, api.exa.ai (two routes),
x402.tavily.com, api.arkm.com,
agents.allium.so, agents.chain.link and two
paysponge.com routes. Every one of those records answered
/v1/resolve with “delisted” for a month. The endpoints
were fine; the prober only spoke GET. Method-aware probing shipped 2026-09-03
10:05 UTC; all eleven were relisted the same morning (changes feed reason
relisted_method_aware) and ten verified within fifteen
minutes. The eleventh (agents.chain.link) answers 404 to both
verbs and is recorded as not_payment_gated, which is what its own
delist note predicted — it is listed as unverified, not
delisted, because that is what we observe. A twelfth POST-only delist
(mercator-entity-evidence.fly.dev) stays delisted: the seller asked
for it and replaced the route.
2026-09-03 — for ninety minutes the public probe count was a fifth of
the truth. From the 10:05 UTC deploy until 11:35, probes_24h on
this site read 21,449 instead of ~115,000, avg_probe_gap_min read
117 instead of 22, and probe_pass_rate_24h was wrong the same way.
Cause: the previous evening’s change to exclude neutral probes from the
pass rate rewrote IS NOT 'x' as NOT IN ('x','y'), and in
SQL NOT IN is never true for a NULL — every clean pass has a
NULL failure_reason, so every clean pass was dropped from the count.
No listing status, score or ranking was affected; only the three published
figures. Fixed 11:35 UTC. The deploy command now reads the public numbers back
against the database after every deploy and refuses to stay quiet when they
disagree by more than 5%.
2026-09-02 — 75 listings carried verified without a
payment challenge ever having been observed. verified is earned
by three consecutive passing probes. Until today a probe also passed when the
unpaid request was answered with a 2xx — “alive” — with a
warning noted but not published. So an endpoint that never once asked for
payment could hold the badge indefinitely, and 75 did: 75 listings on 14
hosts, none of which has ever returned a 402 on any probe, most of them
verified since 2026-07-24, the first day. 62 of the 75 are bare
domains with no path — homepages, which always answer 200. Our own paid
census had meanwhile labelled most of the same endpoints
no_payment_required: two instruments on this site disagreed about
the same listings and nobody joined them. Found after a working-group member
counted that 27% of hosts in his census do not serve their 402 to a bare GET,
39 of them answering 200 instead.
Today: all 75 set to unverified (changes feed reason
admin, 2026-09-02) — twice, because the first time the next
probe cycle re-verified 74 of them within minutes on another 2xx, which is the
rule doing exactly what it said; the transition now also requires that a 402
has been observed at least once, and the second flip held. From 2026-09-03: a pass requires a 402 —
on GET, or on POST when GET does not produce one; a 2xx with no challenge on
either verb is recorded as not_payment_gated and does not count.
The same rule restores the POST-only listings we had been reading as dead.
Fifteenth correction; the badge was easier to earn than the page said it was.
2026-09-02 — we set working endpoints to failing
because our own DNS lookup failed. The prober resolves every URL's DNS
before fetching it (an SSRF guard). A lookup that errored — the
resolver rate-limiting us, answering SERVFAIL, or not answering — was
recorded identically to a lookup that returned no records: as the endpoint's
failure. That verdict was cached per host for an hour, so one bad lookup for
a host with many listings failed all of them for the hour; and because
listings on a failure streak are re-probed every cycle, five failures
accumulated in about 25 minutes and the listing was set to
failing. It recovered on the next successful probe once the
cache expired — which is why it went unnoticed: the status went wrong
and right again inside an hour, eight times.
Measured from the probe log: eight episodes on six days between 2026-08-23
and 2026-09-01, 512 failing transitions on 64 verified
listings of one host that was up the whole time, and roughly 7,200 probe
failures (1.4% of probes in that window) that were ours. On an episode day
the published 24-hour pass rate was understated by about 1.5 points.
Historical probe rows keep their label: the code could not tell the two
cases apart, so neither can we after the fact. Those listings'
/history still shows the flips, because they happened; this
note is the context.
Since 2026-09-02 a lookup error is recorded as dns_lookup_error,
counts against nothing, is never cached, and is excluded from
probe_pass_rate_24h (stated in
probe_pass_rate_24h_basis). A genuine no-records answer is
cached for five minutes, not an hour. A listing now needs three consecutive
passes to leave failing, not one — the same evidence it
needs to earn verified. Fourteenth correction; found while
asking why 110 listings changed status 488 times in a day, and the answer
was the instrument, not the listings.
2026-09-01 — the public demand feed was publishing catalogue
enumeration, one seller's self-checks, and an IP count dressed as a client
count. Four findings against /v1/demand, all in one sitting,
all fixed the same day.
(1) The top ten “search terms” were the single letters a–j,
about 49 searches each — clients walking the catalogue alphabetically.
The endpoint had no minimum query length; our internal clustering had refused
short strings for exactly this reason since it was written. Now ≥ 5
characters.
(2) The single most-searched term, 196 of its 205 asks, came from one client
that also issued 193 edit requests against its own listing — its last
search and last edit one second apart. A seller checking its own ranking is
not demand. Searches from clients that also edit listings are now excluded
from the signal and published alongside it as
self_monitoring_searches, so the exclusion is visible rather
than silent.
(3) What we published as “distinct clients” is a salted hash of
the caller's IP — and one caller presented eight different IPs
within a single second. The count was never a count of callers. Fields
renamed (distinct_client_ips, buyer_client_ips),
the k-anonymity floor is documented as applying to IPs rather than
operators, and the note says to read every such number as a ceiling on
independent demand. This is a breaking rename on a public field, made
deliberately: the old name invited an inference the data cannot support.
(4) Because no IP-based count survives a distributed caller, ranking now
keys on persistence: a term appears only if searched on two or more
distinct days by non-seller clients, and lists rank by distinct days first.
A burst of any size occupies one day and does not rank — demonstrated
on our own 90-second, multi-IP test burst from 2026-08-15, which had been
occupying five of the top ten rows. Domain-shaped queries are now separated
into top_lookups (agents checking a specific service before
paying it) from top_needs (descriptive wants), by a stated
heuristic.
After all four: the top need is “real-time stock quotes sub-second
latency”, searched on 15 distinct days by 24 client IPs — the
first entry in this dataset that survives every artifact class we know
about. Thirteenth correction; the feed was measuring traffic and calling it
demand, and the loudest rows were the alphabet and ourselves.
Two of those four rules changed on 2026-09-11, and this paragraph is
left as written because it is the record of what was true then. The two-day
persistence requirement is gone: the feed now publishes every term two or more
independent clients asked, on however many days, because the product is
what agents ask for and a burst is a real thing that was asked. And
top_lookups is gone with it — a lookup was defined as a
dotted single token, which is a hostname, and hostnames are never published
under the rules now stated at /privacy, so the field
could only ever have been empty. Navigational lookups still happen and are
still counted in the totals; they are no longer itemised. The k-anonymity
floor moved from three distinct clients to two, and a content filter now
withholds any term containing a URL, a hostname, a wallet address or an email
address. Terms logged before that date were collected under the older rule and
are not republished under the new one.
2026-09-01 — “39 hosts offer different networks depending
on which channel the client reads” was computed without
normalization. Posted to the x402 domain-discovery working group on
2026-08-31. The comparison behind it was a raw string match, order-sensitive,
so ["base"] against ["eip155:8453"] counted as a
difference when it is the same chain spelled two ways. Recomputed on the same
run with v1 friendly names and chain aliases normalized: of 224 hosts
serving a readable challenge in both channels, 186 are identical, 36 have a
body that is a strict subset of the header, and 2 name something in the body
the header omits — one a genuine chain-set difference, one a
non-chain string where a chain identifier belongs. The corrected figures were
posted in the same thread. Getting there took two attempts: the first
normalization pass mapped bare solana to a mixed-case CAIP-2
constant while lowercasing everything else, and reported 41 body-only hosts,
39 of them case artifacts. Twelfth correction; the instrument compared
spellings, not chains — which is the same failure mode we have now seen
in three independent implementations in four days, and the reason the missing
canonical form matters more than the channel-precedence question it was
raised under.
2026-09-01 — public detail reads were publishing claim-token
hashes, submitter emails, and claim timestamps.
GET /v1/listings/:id, and the MCP tools
get_service_details and resolve_endpoint, all build
their record from one query which selected l.*. That published
every column the listings table has ever grown. A redaction directly below the
query removed claim_token — correct when written, when that
was the credential's only name. Migration 0021 (2026-08-30) split the
credential into claim_token_hash,
claim_token_hash_prev and claim_token_prev_at, and
added challenge_hash, challenge_hash_prev and
challenge_issued_at; l.* carried all six out, along
with submitter_email. Two exposure windows, and the longer
one is the more sensitive: submitter emails have been readable since this
service launched on 2026-07-19 — six weeks, 94 listings affected —
because l.* was in the detail query from the first commit. The
credential and challenge hashes were readable from 2026-08-30, when migration
0021 created them: 1,730 listings. There were 8,850 detail views by 147
distinct clients in the seven days before the fix. This contradicted two sentences on our privacy
page: that a submitter email is used only for status notifications and
disputes, and that we publish nothing identifying an individual person. Both
have been corrected rather than quietly reworded. The timestamps disclosed when
each seller last claimed or edited a listing — minor, but seller metadata
we published without saying so.
No action is required of sellers, and we are not rotating tokens. A
claim token is 24 random bytes from a CSPRNG, stored as a SHA-256 hash. The
published hashes cannot be reversed to a working credential at that entropy, by
anyone, at any budget — which is precisely what hashing at rest was there
for. Rotating 1,730 tokens would break every saved credential to fix a
vulnerability that does not exist, and we can reach only 94 of those sellers by
email; the rest would discover it by silently failing to edit, which is the
harm the eighth correction was about. If you would prefer a fresh token anyway,
the claim flow reissues one.
The fix is an explicit column allowlist rather than another redaction: a
delete-list cannot protect a column that did not exist when it was written, and
this one didn't. Found while capturing response shapes to write MCP output
schemas — we were reading our own output closely for an unrelated reason.
Eleventh correction; the leak was ours, and so was the design that made it
harmless.
2026-09-01 — what a null response_mime means.
response_mime shipped 2026-08-31, exposing the
mimeType an endpoint's own 402 challenge declares for a paid
response. This page said a null value meant the listing had “not yet
been re-observed since 2026-08-31”. That was true for the first
minutes and misleading afterwards. Of 1,480 listings re-observed since
the field went live, 65 declare a mimeType — 4.4%, and 63 of those
declare application/json. Null therefore means, almost always,
that the seller declares no response content type at all, not that we have
not looked. Two listings declare an empty string, which we now store as
null so that “declared” means declared. The field is read per
accepts-entry, so a listing offering several payment options could in
principle declare different types per option; we record the one on the
entry we parse. Wording corrected here, on /integrate and in
llms.txt. Caveat on our own reasoning: the field was motivated
by one client repeatedly searching for a plain-text conversion service and
never buying — every listing that client's query ranks returns null
here, so the field did not help it and cannot until those sellers declare
one. It is justified because we held the information and were not showing
it, not by that story. Tenth correction; the instrument was right and the
sentence around it was wrong.
2026-08-29, evening — seller edit tokens were being invalidated by
crawlers. GET /v1/listings/:id/claim issues a fresh claim
token on every request (so a lost token can be recovered), keeping the
previous one valid for an hour. Every listing page links to that URL and
the sitemap lists every listing page, so ordinary link-following crawlers
were requesting it continuously: in the 24 hours to 2026-08-29T10:00Z,
1,322 tokens on 431 hosts were rotated, about two a minute. The
consequence for sellers: a token from a submission or a claim, if not used
within roughly an hour, was probably already rejected by the time it was
tried — with no explanation. If your x-claim-token is refused,
request a new one via the claim flow (GET /v1/listings/:id/claim
explains it) and use it promptly; nothing about your listing was changed. Mitigation shipped
2026-08-29 ~09:50Z (robots.txt disallow on the claim route,
rel=nofollow on the link); rotations fell from about 120 an
hour to 7 in the following four and a half hours, a rate consistent with
sellers using the flow as intended. A design fix that separates the
challenge a stranger can request from the credential an owner holds is
planned, and this note will be updated when it ships. Eighth correction; a
GET with a side effect, which is on us.
Update, 2026-08-30: shipped. GET /claim is read-only;
challenges are minted by POST /claim/challenge into their own
columns and verified by POST /claim; the edit key is written
only at submission and on a successful proof. No unauthenticated request
can touch a seller's key any more. If your key from before 2026-08-29 no
longer works, the challenge flow reissues one.
2026-08-29, later — onchain_tx_count_30d has a ceiling of
150 that was never disclosed. The provider read pages at most three
times, 50 items a page, so no listing could ever show more than 150
transfers in 30 days, and onchain_unique_payers_30d and
onchain_volume_usd_30d are truncated with the count. Found
while testing a replacement source, which returned more transfers for one
address in 28 hours than our figure allowed for a month. As of publication
741 of 1,540 listings with on-chain rows sit exactly on 150; they
belong to 77 of 346 payment addresses (many listings share one). For those
77 addresses the sum of displayed volume is $1,346 — a floor, not a
figure; the true amount is unknown. Nothing above 150 exists, which is how a
cap looks. Corrected: onchain_count_capped is now exposed on the
listing record and documented; the numbers are left as they are, labelled,
rather than guessed upward. The cap goes away when the source is replaced.
Seventh correction; the instrument was our own paging loop.
Update, 2026-08-29 evening: the source is replaced. Figures now come
from our own index of USDC Transfer events on Base (onchain_source
'index'), no ceiling. Re-read against it, the ceiling had affected 1,014
of 1,499 listings, not only the 741 sitting exactly on 150: a capped read
lands at or below 150, because the 150 items were paged across all tokens
before the USDC filter. The largest case: a seller shown as 149 received
7,287,701 payments from 414 payers, $169,144 in the window — and at
the daily level the index keeps, it peaked at 1.2 million payments a day on
Aug 17–18 and fell 99% on Base from Aug 24 with its payTo unchanged. Base
USDC only: whether that traffic moved chains or stopped is not something this
index can say, and we don't.
2026-08-29 — on-chain activity figures were shown as fresh while
the provider was down. The third-party API behind
onchain_tx_count_30d, onchain_unique_payers_30d and
onchain_volume_usd_30d began returning HTTP 500
(“source: upstream”) for most addresses at 2026-08-28T10:30Z and
was still doing so at publication, for keyed and unkeyed calls alike, from
our infrastructure and from outside it. A fix shipped 2026-08-28 to stop
failed checks from starving the queue wrote a row on every failure that kept
the last good counts but stamped the check time as now. Result: on 990 of
1,540 listings with on-chain rows, figures last read before the outage
carried a timestamp saying they were read today. Corrected 2026-08-29:
onchain_checked_at is now the last successful read and is
null on those 990 (the true time was overwritten and is not retained);
onchain_last_attempt_at, onchain_provider_error and
onchain_provider_status are exposed so the state is visible
rather than inferred. 949 of the 990 carry real earlier figures, held; 41 had
never been read successfully and previously showed 0/0/0 — they now
return null. No figure was fabricated; the error was a timestamp asserting
freshness the data did not have. Checks continue at a reduced rate to detect
recovery. Until the provider recovers or is replaced, treat a null
onchain_checked_at as “pre-2026-08-28, exact time
unknown”.
Update, 2026-08-29 evening: replaced. Every active listing with a Base
payTo (1,499) now has its figures written from our own index every five
minutes (onchain_source 'index'); the held rows and null
timestamps described above no longer exist for those listings. 42 rows not
covered by the index keep their last earlier read, with a null
onchain_source.
2026-08-27, evening — “paid” in paid_but_status_4xx meant authorization attached, not money moved. The scout's non-200 path never consulted the chain; the verdict name asserted settlement its code did not check. Retroactive chain audit, validated with positive controls (the query finds all six known delivered settlements at the same addresses and window before its zeros are believed): wave 5's 99 such calls — 96 proven unsettled, 3 chain-confirmed, $0.0045 total; the first census's 405 4xx attempt-rows — 241 proven unsettled where a control passed, 124 unknown (stated as unknown), ceiling of 34 rows / ≤$0.452 possibly settled (upper bound, absorbs unproven-delivered settlements), 2 chain-confirmed. Published claims of “money gone” for this class are corrected on the census page. The scout will chain-verify settlement on rejected calls before any wave carries this verdict again; until then the verdict name is read as “authorization attached, request rejected”. Fourth correction of the day, each found by the instrument beneath the previous one.
2026-08-27, later the same day — the correction below itself
undercounted: our collector truncated 27% of the challenges. The census
tool behind the entry below capped stored headers at 4KB; 382 of 1,395
payment-required headers were cut mid-challenge and the cut-off challenges
were classified as declaring nothing — the same error shape as the
figure being corrected, one layer down, caught while preparing to name
specific sellers on the strength of it. Re-collected with the cap removed
and re-derived the same day: 82% (1,366 of 1,675) carry machine-readable
metadata in-band; 463 declare a field-level response-body schema; 934
declare their parameters (448 with required marked); 97 of wave 5's 99
paid-then-400 failures — not 46 — had declared the parameters
our buyer omitted, and only 2 declared nothing. Both corpora are
retained. The lesson, stated where it was earned: an instrument that
silently truncates its input will confidently measure absence.
Scope note, 2026-08-31: these census figures, and every per-listing
challenge field we publish (network, price, accepts_count, x402 version),
are readings of one channel per listing — the
PAYMENT-REQUIRED header when it parses, the JSON body
otherwise. We then measured both channels across all 503 listed hosts: 226
serve a challenge in both; 95 serve x402 v1 in the body and v2 in the
header; and after normalizing v1 network names against CAIP-2, 39 hosts
offer different networks depending on which channel a client reads (27
of the 39 are one operator's stack). The spec does not currently say which
channel is authoritative; until it does, a figure about "the endpoint's
terms" is a figure about the channel we read. Raised with the working group
2026-08-30/31 with per-host results retained. Ninth correction, of scope
rather than value — found by our own scan after one listing's
channels were caught contradicting each other.
2026-08-27 — "zero of 550 declare a checkable response schema" measured
our form, not the network. The census stated that no purchased endpoint
declares a machine-readable response schema and that conformance was therefore
unmeasurable network-wide, by anyone. That was true of one place: the
response_schema field of our own submission form. One unpaid
request to every active listing on 2026-08-27 (1,675 attempted, 1,639
collected) found 64% carry the extensions.bazaar block in the 402
challenge itself; 227 declare a field-level response-body schema; 841 an
output example; 460 typed parameter schemas; 36 hosts additionally serve
schemas at /.well-known/x402. Our own scout had parsed these
challenges 1,333 times and kept only the price. Related: 46 of wave 5's 99
paid-then-400 failures had declared the omitted parameters in-band; the scout
now refuses to pay when a challenge declares required parameters the call
omits. Corrected on the census page, its JSON twin, the docs index, and
here.
2026-08-25 — duplicate listings. Until today, listing uniqueness
compared raw endpoint URLs, so example.com/x and
www.example.com/x counted as two separate services. Four such
pairs existed, all from one seller, inflating listings_total by
4 and hosts_total by 1. URLs are now compared in canonical form
(www., trailing slash and tracking parameters collapse to one key), enforced
by a unique index over active listings. The four duplicates are delisted with
reason duplicate_canonical and appear in the changes feed; the
seller was told the same day.
2026-08-25 — on-chain payer counts.
onchain_unique_payers_30d was described as wallets that paid a
listing. It counts every incoming USDC transfer to the payment address, so
operator self-funding inflates it. It is now published as a ceiling on
customer count rather than a measurement of it, on every surface that carries
the field.
2026-08-21 — paid verification re-grades. Nine endpoints originally graded as failures were re-graded delivered after on-chain confirmation. The correction is visible in the outcome table at /stats rather than folded into the totals.
Date not recorded — a withdrawn rate. A published percentage of endpoints that accepted payment without settling on-chain was withdrawn because its denominator could not be reconstructed from our own records; raw counts with a stated sample are published instead. We did not record the date we withdrew it, which is why every correction above carries one.
Questions, corrections, disputes: hello@nohumans.directory · API reference · pre-spend check
Method changes
Rules that became stricter without a published figure being wrong. Each entry says what flipped, so a status change on a listing can be traced to the rule that caused it.
2026-09-03 — method-aware probing. A GET that does not
produce a 402 is retried as POST with an empty JSON body (retry set: 2xx and
4xx except 429; redirects, 5xx and transport errors are not retried); a 402 on
either verb is alive. Neither verb producing a 402 is
not_payment_gated: recorded on the probe, not a pass, and not a
failure either — it moves neither score nor status. Adopted from the
working group's measurement (356 of 1,312 hosts serve their 402 only to POST,
or serve 200 to GET). Observed on the first full pass (10:05–12:40
UTC): 87 listings on 23 hosts verify on POST — 70 that had been
failing (the GET-only estimate of 34 had missed every host whose
GET answered 404 rather than 405, including one with 24 listings), 10 relisted
from August (see Corrections), and 7 submitted that morning. No listing that
answered 405 to GET failed its POST retry. failing fell from 180
to 105. 95 listings on 22 hosts are not_payment_gated. Each listing
detail now carries verified_via (GET or
POST: the verb that produced the 402 on the latest passing probe,
observed, never declared) and a method_note when it is POST. The
submission check runs both verbs too; the warning that told POST-only sellers
to bring a GET-able route is withdrawn, and a new informational code
post_only_endpoint tells a seller their route verified on POST.
Rule added the same day: 14 continuous days not_payment_gated
— no probe on either verb has asked for payment for two weeks —
auto-delists with that reason (earliest possible effect 2026-09-17). The 30-day
rule is for endpoints that broke; this one is for listings that were never
paid endpoints.
2026-09-04 — two tightenings of the rule above, the morning after.
(1) 404 or 410 on both verbs is absence, not “present but ungated.”
Nothing answers at that path, so it fails (bad_status) exactly as it
did before 2026-09-03. Under the first day’s rule it was neutral, and neutral
probes never move status — so nine dead routes on three hosts kept
verified overnight instead of walking to failing. A
neutral class needs an absence check or absence becomes neutral. 405, 401,
400 and 2xx still count as an answer (403 did too, until the entry below), and
a 404 to GET with a 402 to POST is still POST-only and verifies. (2) Neutral listings leave the hot tier.
unverified listings are probed near every cycle so they can earn
the badge quickly; a listing whose probes are neutral cannot earn anything, and
94 of them took 18.5% of all probes on the first day. After ten consecutive
neutral probes a listing is re-checked at the same low rate as the long-dead.
Neither change touches a published figure; the first-day count of 95 on 22
hosts included the nine, so today’s figure is 86.
2026-09-05 — a 403 to the prober is a refusal, not an answer.
A 403 to the unpaid GET with no 402 on the POST retry is now
probe_refused: neutral like not_payment_gated (no
score, streak or status), but it neither starts nor breaks the 14-day
not_payment_gated run. Cause: one host with 61 listings answered
403 to two whole probe sweeps in a day (60 and 32 probes) and 402 to every
other probe — some 4,000 in the same 24 hours — and 402 to a plain
curl from a laptop. Under the 09-04 rule those two sweeps read as “present
but ungated,” and the morning brief listed all 61 as candidates for the
14-day delist. The clock had not started (a later 402 resets it), but the brief
said it had: its host table counted any neutral probe in 24 hours rather than
the latest probe, and that is fixed too. Third tightening of the neutral class
in three days, same lesson each time: a neutral class needs a check for the
case where we are the variable. Open question, deliberately not a rule yet: a
host that refuses us permanently. The brief now shows, for every listing whose
latest probe was refused, the days since a 402 was last seen on its host; if
that number grows on a real host, a rule follows the data. The public pass-rate
basis names three neutral classes from today. No published figure changes.
2026-09-06 — validation before payment is neutral, and a JSON 403 is
validation. A route that rejects the request before asking for money
(400/422 on both verbs) has been warn:params_likely_required since
2026-08-26, with the note that we “say so rather than call it ungated.”
It was still a failed probe: no score credit, and five in a row walked the
listing to failing. From today it is the fourth neutral class:
no score, streak or status, and it does not start the 14-day not-gated clock.
A 403 whose body is a JSON error (“Company number must be 6-8 digits …
Not charged. Example: …”) is the same design wearing the wrong status
code and is classified the same way; a 403 with an empty or non-JSON body stays
probe_refused. The prober now reads 403 bodies to tell them apart.
On a templated URL (/:param or {param}), 404/410 to the
literal placeholder is the same class, not absence — 165 of 239 imported
templates answer 402 to the placeholder; the rest validate first.
This one corrects a published figure — see correction #17.
2026-09-07 — templated routes are probed with the seller’s own
example. A route template that validates before charging answers the literal
placeholder with an error — and, often, with an example of a valid call in
the same body (“Example: /v1/company/SC311560/verdict”,
example_url, hint). From today the prober takes that
example, on the seller’s own origin only and never a guessed value, and
probes it once in the same cycle, GET then POST as usual. A 402 there is a pass,
and the probe row records probe_url — what we actually called.
No example in the body: the neutral class, as before. This is how the 48
listings of correction #17 can verify at all; whether they do is a fact about
their error bodies, and the morning check now counts it.
2026-09-07 — testnet listings leave the public counts. Seventy
listings answer their 402 with a testnet network (eip155:84532,
Base Sepolia, in two spellings). They verify correctly — the challenge is
well-formed and stable — and most arrived with the Bazaar import, where
“two distinct payers” is free on a testnet. A payment there carries no
value, so the same treatment as ephemeral hosts: probed, served, discoverable,
testnet: true on the record, excluded from
listings_total, verified and hosts_total,
badge reads testnet, and the paid verdict says caution
with the network printed. Not delisted: a test service is a real service. The
network list lives in one place and is published in the morning check.
2026-09-07 — a delisted route that answers 402 again comes back.
Until today a listing delisted by a probe rule (thirty days failing, or fourteen
days with neither verb asking for payment) was never probed again; a seller who
fixed the route had to resubmit, and eleven of them in August did not know to.
From today a probe-delisted listing stays in the probe queue for thirty days
after the delist, at the back — never ahead of a live listing — and a
parsed 402 returns it to unverified with reason
auto_relist_402 in the changes feed, earning the badge the normal
way from there. Listings delisted by policy or by their owner (our own routes,
duplicates, platform mirrors, self-delists, takedowns) are never relisted by a
probe: on 2026-09-06 all three “delisted here, live on the Bazaar”
cases were exactly those, and the rule was written to leave them alone.
2026-09-05 (b) — ephemeral tunnel hosts leave the public counts.
Listings whose host is a developer tunnel (lhr.life,
trycloudflare.com, ngrok, loca.lt,
serveo.net) are still probed, still served on
/v1/listings/:id and in discovery, and now carry
ephemeral: true on the record. They are excluded from
listings_total, verified and hosts_total,
and their badge reads ephemeral rather than a status. Cause: one
automated seller-onboarding loop — three listings per tunnel, a fresh tunnel
every 30–40 minutes, all three deleted within seconds, no contact address
— was 123 of yesterday’s 126 new listings and 77 of 110 pass-streak
verifications; 27 of its listings were live and counted at the moment of the
morning check, and a third-party crawler was indexing them minutes after they
were deleted. A verified badge on a host that will not exist in an
hour is a claim this directory cannot stand behind. This one changes a
published figure by rule: listings_total drops by the tunnel
listings live at deploy time (27 on 2026-09-05). The loop itself is welcome
— it is a seller testing the flow we document — and is not blocked.
2026-09-03 — kind=agent verifies exactly like
kind=api. The submission form accepts both kinds; verification
is identical: the endpoint must answer 402 to an unpaid GET or POST. An agent
that does not charge is not verifiable here and receives
endpoint_not_402 honestly. A separate verification path for
non-paid agents is not on the roadmap unless agent submissions arrive.
2026-09-14 — a sentinel payment address blocks and revokes
verified. A 402 that names a burn address (the zero address and
the two conventional dead addresses) as its payTo is a listing that cannot deliver,
whatever its probes say. The record has warned about it (payto_sentinel)
since 09-05; status now agrees: a listing whose current challenge names a sentinel is
held at unverified, a verified one is demoted with reason
payto_sentinel in the changes feed, and probes continue — a real
address earns status back through the normal streak. One listing held the badge on the
zero address on the day the rule shipped. Also today: the /v1/changes
note lists the reason values in use, derived from the table hourly rather than typed
(the typed list carried 8 of 38).
2026-09-16 — the pass-rate basis names five neutral classes.
Since 2026-09-13 (correction #19) a 429 to our unpaid GET, and since 09-14 to
the POST retry, has been excluded from the published 24-hour pass rate as
warn:rate_limited. The queries changed then; the basis sentence
served with the figure at /v1/stats did not, and went on naming four
classes for three days. The sentence and both queries are now built from one
list, so a class cannot reach one without the other. No published figure
changes.
2026-09-23 — probe cadence is tiered by doubt, and a sweep-wide timeout is ours.
(1) Until today every active listing had one re-probe target, 30 minutes. The
mean held (27 minutes) and the tail did not: on the morning of the change the
most overdue verified listing was last checked 185 minutes earlier. Most of the
~285,000 daily probes re-confirmed listings whose status was not in doubt. From
today each listing is in one tier, decided on every probe from what that probe
and the record show, and the tier is on the record: probe_tier,
probe_interval_s, and next_check_by with its
next_check_basis. The scheduler selects on the same stored
interval, so next_check_by is the rule, not an estimate.
The tiers:
- new tier, re-probed on every probe run (every 2 minutes): unverified and either new (fewer than 10 counted probes) or on a clean streak toward the badge, so a new, reset or recovering listing earns it in minutes rather than hours
- watch tier, re-probed every 30 minutes: unverified, failing (up to 10 in a row), failed in the last 24 h, payment terms changed in the last 7 days, or new or changed status in the last 7 days
- stable tier, re-probed every 2 hours: verified with no failure in the last 24 h and no change of terms or status in 7 days, or failing 11 to 50 times in a row; any failure moves it to the watch tier on that probe
- background tier, re-probed every 8 hours: probe-delisted and inside the 30-day relist window, on a tunnel host, challenge names a testnet, more than 10 neutral probes in a row, or failing more than 50 times in a row
A failure on any tier puts the listing in the watch tier on that probe. Sized
on the day against the live queue of 5,576 listings: 1,499 watch, 3,320
stable, 757 background, and the new tier a few dozen at a time — about
115,000 probes a day, down from about 285,000. This replaces the queue
priorities described in the 2026-09-02 and 2026-09-04 entries above; the
thresholds they introduced (ten neutral probes, fifty failures) carry over as
tier rules. Amended the same day: the new tier also covers an unverified
listing on a clean streak, whatever its probe count. A listing climbing back after
a payment-address reset or a recovery from failing needs about sixteen clean
probes (correction #23), and at the watch interval that took hours; it is now
re-checked every run while it climbs, and a failure moves it back to the watch tier.
Amended 2026-09-24: the tiers set how often a listing should be
checked, but we also never probe more than two of one host's listings in a run.
On a host with many listings those two rules disagree, and the record promised a
check the host limit could not keep: every late listing in the two fastest tiers
that morning, 199 of them, was on one host with 231 listings, checked about every
three hours against a stated two or thirty minutes. The stored interval is now the
larger of the tier's and the host's turn, and such a record says
probe_tier: host_limited with the reason. A host with one or two
listings is unaffected. Also 2026-09-24: the rule correction #23 states
(the score starts at the first counted probe's result) now holds for a listing
whose first probes were neutral. Before, neutral probes created its record with a
score of zero, so its first counted pass scored 0.1 and it climbed from zero, about
sixteen probes instead of three. A listing whose payment address was reset still
starts at zero, as stated.
(2) When at least half of the probes in one probe run time out, on at least 8
different hosts, those timeouts are recorded as warn:sweep_timeout:
no effect on score, streak or status, and excluded from the published pass rate
as its sixth neutral class. That is a failure of our own vantage point, not of
the listings. Amended the same day. The rule first shipped at 8%, sized
on the 2026-09-22 12:00–17:00 UTC window, on the premise that unrelated
endpoints timing out together is a fact about our path. That day’s own data
contradicted it: the timeouts were concentrated on networks on the US East
Coast (DigitalOcean, 64% of probes to 79 of 87 hosts run by 13 operators;
Amazon’s us-east-1, 28%, 23 operators), while hosts in Germany, on Google
Cloud and on Render stayed at 0–1%. For those listings the timeouts were
real failures to reach them, and a rule that set them aside would have hidden
them. At 50% the rule covers only a run where most of what we probe fails at
once; the 09-22 window peaked at 26% and would not trip it. No probe was recorded as warn:sweep_timeout under the first version. Flips: none;
the statuses set during the 09-22 window stand.
2026-09-24 — p50_latency_ms is a median. It was a
running average in which each new probe counted for half, so it mostly reflected
the last few probes. llms.txt said so; the API schemas did not, and the field’s
name says median. It could sit above the listing’s own p95: one read 1,553 ms
at p50 against 1,417 ms at p95 that morning. It is now the median of the
listing’s probes in the last seven days, from the same computation as p95 and
p99 (same window, timeouts included). A listing keeps its older value until its
next percentile refresh, and a new listing shows none until it has five probes in
the window. Flips: none; latency sets no status. The paid verdict reads p95, not p50.
2026-09-25 — submission fee. Submitting was free
and unlimited. The first 5 listings per registered domain stay
free; each further new listing is $0.50 USDC, paid over x402 on the same
POST /v1/listings. Validation runs before the charge, so a rejected
submission costs nothing. Listings already in the directory are not charged;
they count toward the five. Listings we index, claims and edits are free.
Flips: none; no status, score or order depends on whether a fee was paid.
Rule.
2026-09-26 — lookups of a named listing are listed
apart from demand. About a quarter of the searches in the demand window were a
service looking itself up: one term, a listed service’s own name, was searched
from over two hundred client IPs that searched nothing else. That is a service
checking it can be found, not someone shopping, and it led the needs ranking.
The rule: a term is a lookup of a named listing when it contains a service's own name (a word of 4 or more characters that appears in the listing names of exactly one seller, one registered domain, and in that seller's own hostname) and at least 80% of the client IPs that searched it searched nothing but that seller. Such terms move out of top_needs
(the free feed, /state/demand and the MCP demand tool) and
out of the ranked needs in the paid clusters and report, into
named_listing_lookups, with the same counts. Nothing is deleted or
withheld, and the privacy rules are unchanged. What it does not catch: a service
name that is not in its seller’s hostname, and a term that some clients also
search beside other things, which stays a need. The threshold is high on purpose,
because a buyer searching for a named service is demand. Flips: none; no
listing’s status, score or order depends on it.
2026-09-26 — listings say which payment protocols their
402 offers. The prober now also reads MPP offers (the “Payment” HTTP
authentication scheme, WWW-Authenticate: Payment) on every 402 it
receives. Before, it read only x402 terms: an endpoint offering MPP alone passed its
probe with no price shown, and one offering both showed only its x402 side. Discover
results and records now carry payment_protocols (x402,
mpp) and mpp_methods (as the endpoint names them: tempo, evm,
stripe and others); the record adds mpp_offers and mpp_note.
What does not change: paid verification and on-chain figures cover x402 on Base only,
so a listing that offers MPP alone can be probe-verified but not paid-verified, and a
Tempo recipient has no on-chain figures here. Flips: none; passing a probe never
depended on which protocol’s terms were read.
2026-09-28 — the submission fee is lower. A new listing beyond the first 5 per registered domain now costs $0.25 USDC ($0.50 since 2026-09-25). Nothing else about the rule changes. Flips: none; the fee still buys no status, score or place in the order. Rule.
2026-09-28 — the demand data is built in the background.
Every demand surface used to compute on the request that asked for it: the free feed whenever its hourly copy
expired, and both paid routes on every purchase, before the payment settled. That cost grew with every search,
and the paid clusters read at most 10,000 searches, so once the 30-day window outgrew that the oldest days fell
out (the response said so in window_complete). Now one build every 15 minutes
stores everything and requests only read it. What changes: /v1/demand/clusters and
/v1/demand/report are at most 15 minutes behind live instead of live, and say
when in generated_at; if the stored copy is missing or older than 45
minutes they answer 503 before any payment is asked for, instead of selling it. The report’s
answered_now is our search for the need, re-run about once a day for the 1,000 highest-ranked
needs (checked_at says when), not at the moment of purchase. Counts now cover every search in the window, with no row cap. A phrasing is grouped once,
when first seen, and keeps its need: grouping no longer changes as the window moves, so a need keeps its identity
from day to day. The threshold (similarity to the need’s first phrasing, no drifting centre) and the
privacy floor were unchanged by this move; the threshold was raised the same day, see
the regroup entry. A need is labelled by the phrasing the most distinct clients asked in the window
(its first phrasing is kept for grouping only: on the first build it was a wording asked once, labelling a need
116 clients asked in other words). Phrasings first seen since the last build are counted in texts_pending_grouping
until the next. A paid body lists at most 15,000 needs, highest-ranked first, and
counts any beyond that in needs_not_listed_beyond_max. The free feed is unchanged in content, still a
week behind, and now rebuilt hourly in the background, with one bound: our catalogue search beside each term
(catalogue_top) is run for the 1,500 highest-ranked terms, each about once a day
(supply_check.checked_top_terms); a term ranked below them is not_checked_yet, as the
feed says. It can be read in pages (?limit=&offset=,
with terms_total and next_offset); without them it is returned whole as before, until it
outgrows 5 MB, after which the first page is returned with the same paging
fields. The MCP tool what_agents_are_asking_for now returns one page (100 terms unless
limit says otherwise) instead of the whole feed, so its reply stays a size an agent can read; every term
is still published. Flips: none; no listing status depends on demand data.
2026-09-28 — scripted rotation is listed apart from demand.
From 2026-09-24 one caller on rotating addresses searched a fixed set of five generic terms: about 110 to 120
client IPs per term, almost all of which searched once and never again. Ranked as demand, those terms led the
paid report. An agent searching once is normal; a hundred addresses that each searched only once, all asking
the same term, is one caller. The rule: a term is listed apart as scripted rotation when at least 50 distinct client IPs asked it and at least 90% of them searched only once in the whole window (one caller on rotating addresses; an agent searching once is normal, a hundred single-use addresses asking the same term is not). Measured on the 30 days before it
shipped, it caught four of the five terms and nothing else (the fifth, “x402”, is under the
five-character floor that already keeps it out of the needs). Such terms move out of top_needs and out of the ranked
needs in the paid clusters and report, into scripted_rotation, with their counts and the share of
single-search addresses. Nothing is deleted or withheld. What it can catch by mistake: a real term asked
mostly by agents that search once from changing addresses; the list is published, so that would show.
Flips: none; no listing’s status, score or order depends on it.
2026-09-28 — needs regrouped at a stricter similarity (0.72 → 0.78). A phrase joins a need when its meaning is close enough to the need’s first phrasing. We checked every phrasing grouped under the old threshold, 0.72, by hand: of 422, 96 were in the wrong need (“news data feed” counted as a cryptocurrency price feed, “json repair” as web-page extraction), 259 in the right one, and 67 could not be called either way. Wrong joins were more than half of the clear cases between 0.72 and 0.75, a third between 0.78 and 0.80, and none above 0.85. Replayed over the same phrasings at 0.78, 31 of the 96 remain and 58 of 259 correct joins are split, and slightly more needs clear the privacy floor (65 against 59 in the replay; the live regroup also covers phrasings seen since). Wrong merges publish a wrong fact; splits publish the same demand in more rows, so we took the stricter setting and regrouped every phrasing from scratch under it. From about a minute after it started until it finished, the paid demand routes answered 503 and asked no payment. Needs listed before this date are not comparable one for one with needs after it. A rule that also looks at shared words did better on this data, but it was designed on the same phrasings, so it will be tested on phrasings first seen after this date before it is considered. The privacy floor and the ranking are unchanged. The first run, at 22:15 UTC, started some needs twice (“crypto price” beside “oracle data crypto price”, 0.86): a build grouped phrasings before the previous build’s new needs could be found by a search, and a fixed 90-second wait was the only guard. Paid copies built from 22:21 UTC until the rerun carried those duplicate rows. Grouping now also waits until the most recently stored need’s vector is found by a search (for at most 30 minutes, after which it goes on and the build records it), and the regroup is run again with that fix. Flips: none; no listing’s status, score or order depends on demand data.
2026-09-29 — grouping near the threshold: what we changed, and a correction. After the regroup, 9 phrasings had started needs of their own beside a need they scored 0.781–0.798 against, and 6 had joined one just under 0.78. We first put this down to the vector index’s approximate scores and, from this date, re-score every candidate at 0.73 or above exactly (a candidate that cannot be fetched leaves the phrasing waiting). That was not the cause. Checked against three exports, the embedding model itself returns a slightly different vector for the same text depending on which other texts share its batch (up to 0.033 in similarity), and the index’s search is approximate too: once it did not return a need scoring 0.90, and a duplicate was started. A second regroup today changed the count little (7 of 334). So near 0.78, a phrasing’s need can depend on those effects; the fix (one text per embedding call, and an exact search over stored needs in place of the index) is planned, and will be dated here when it ships. The threshold, the privacy floor and the ranking are unchanged. Flips: none; no listing’s status, score or order depends on demand data.
2026-09-29 — a verdict on our own routes would be labelled and capped.
A verifier that is not independent of the resource operator is no independent check at all (the x402
working group’s draft verifier schema reads such a verdict as inconclusive). This directory does not
list its own paid routes today, so no verdict was affected; the rule is in place for the day one of them
is listed, by us or by anyone submitting it. A listing whose host is this directory, or whose live challenge
pays our receiving address, gets self_referential: true, a first reason saying so, and a verdict
never better than caution. The rule looks at what the listing is, not at a list of ids. Flips: none.
Paid-verification outcomes
Every paid attempt records one outcome. /v1/stats reports each endpoint’s
latest non-dry attempt. The table is what each label means, whether USDC left our wallet,
and whether it sits in the denominator of pass_rate_payment_required. That denominator is every attempt except no_payment_required — refusals we chose, settlements we could not prove and our own client errors all count against the rate, which is why it reads low. our_limit reports how many of those are ours (payload errors, the per-call cap, unfilled templates) so a reader can take them out; report #2 §1 shows the rate three ways. Added 2026-09-14 after a reader asked
what settlement_unproven meant and found it defined nowhere.
| Outcome | What happened | USDC moved | In the published pass rate? |
|---|---|---|---|
delivered | Paid; the response validated; the settlement to the quoted payTo was found on Base. | yes | numerator and denominator |
settlement_unproven_onchain | Served correctly, but no settlement to the quoted payTo could be found on Base after checking every hash the response claimed and sweeping the payTo. Recorded as not delivered: we do not credit a delivery we cannot prove we paid for. A suffix (:local-skip, :first-can) names what the check could not do. | unknown | denominator, not credited |
settlement_unproven | The same outcome before 2026-08-29, when the on-chain check did not yet exist. | unknown | denominator, not credited |
declared_params_unmet | The 402 challenge declared required parameters the call did not carry; the scout refused before paying. Rows before 2026-09-09 carry this label. | no | denominator |
params_fillable / params_unknowable | The same refusal since 2026-09-09, split by whether a generic public value could have filled the gap. Fillable ones are re-tried with the value; unknowable ones (an id, a username, a wallet) are not, because an invented identifier buys a 400 on someone else’s record. | no | denominator |
requires_params_skipped | A route template with a placeholder and no example to fill it. | no | denominator; counted in our_limit |
no_payment_required | The route answered without a 402: nothing to verify with money. | no | no — not_applicable |
unparseable_402_quote | A 402 arrived but neither body nor PAYMENT-REQUIRED header yielded a readable price and payTo. | no | denominator |
quote_above_per_call_cap | The live price exceeded the scout’s per-call ceiling. | no | denominator; counted in our_limit |
price_drift | The live quote exceeded the listed price by more than the tolerance. Recorded on the listing page. | no | denominator |
pre_flight_status_NNN | The unpaid request returned NNN instead of 402. 404/410 is absence and counts; 400/422 is validation before payment and 405 is the wrong verb — called wrong, not broken. | no | denominator |
rejected_400_blind_request | The paid request answered 400: the endpoint rejected the input we sent. Usually the seller validates before settling and nothing is charged; the row’s tx_hash says when it did settle (9 of 128 in the 2026-09-01 audit, /state/the-400-you-paid-for). Label since 2026-09-01; earlier rows carry paid_but_status_400. | rarely — the row says | denominator |
paid_but_status_NNN | Payment accepted; the response was NNN, not 2xx. (paid_but_status_400 is the pre-2026-09-01 name of the row above.) | yes | denominator |
invalid_json / empty_body / stale_data | Paid, 2xx, but the body failed validation: not JSON; empty; or older than the freshness the listing declares. | yes | denominator |
timeout | No response within the scout’s limit. | no | denominator |
corrected_…_chain_verified_… | A row first recorded as settlement_unproven whose settlement was later found on Base — from a hash the body claimed (2026-08-21), from our own USDC index (2026-09-15), or from a direct query of Base for the scout’s USDC transfer to the seller’s address in the minutes after the purchase (2026-09-25; label source rpc). The scout’s chain check can run ahead of the block; the correction carries its date and source in the label, and the hash-chained log keeps the original row. | yes | numerator and denominator |
error_… | The client could not build or send the payment (unregistered scheme, EIP-712 domain mismatch, network). Ours, not theirs. | no | denominator; payload errors counted in our_limit |
The paid verdict
GET /v1/listings/:id/verdict is one of this directory’s 5 paid routes
(all listed with their prices in the x402 manifest): $0.005 USDC on Base, x402, settled through
the Coinbase facilitator before the response is sent, receipt in the
PAYMENT-RESPONSE header. Discovery, records and probes stay free. It returns one of pay,
caution, avoid, with the reasons printed, built only from things
this site already observes: probe history, the verb that produced the 402,
payTo age and stability, our own paid-delivery outcome, on-chain inflow, and
declared-vs-live price. No seller input, no ranking purchase, no exception
list. The rule, in the order it is applied (the worst level wins):
Paid verification, and what it does not buy (2026-09-10;
first offered 2026-09-08). A seller may pay to have their listing bought:
single, one real purchase within 24 hours, $3 USDC on Base; or
pack, three real purchases inside 30 days, the first within 24 hours and
the other two at times we choose, $5. Both only for listings quoting up to
$0.50 per call; above that we price by hand on request. The paid request must
carry the listing’s claim token — we sell purchases only to the
party that owns the route, so nobody can buy attempts against a competitor
hoping for a red badge; the unpaid request answers an ordinary 402 and anyone
may read the price. Every purchase is graded by the same code as a wave, enters
the rule below the same way, and is published whatever it finds: settlement
transaction, delivery outcome, failure reason, and the paid-depth rung —
including failed_rebuy when an endpoint that delivered before fails
a later purchase. The ledger row carries requested_by. A failure
whose cause is our own client does not count against the plan and is re-run.
Money buys when a purchase happens; it never buys the outcome, the rung, the
status, the score, the verdict, or a place in the order. Fulfilment is
manual in this phase and runs happen on weekdays; the record shows
purchases_done of purchases_total. Revenue lands in a
wallet separate from the scout’s; purchases a plan promises are funded
from it by hand before they run.
avoid if any of: status is failing; the latest probe was
not_payment_gated (neither verb asked for payment); the payment
destination is unstable (≥3 distinct payTo addresses in the last 10
probes); or our scout settled a real payment to it and did not receive a
valid response within the last 14 days (the settlement tx is printed).
caution if any of: the endpoint is operated by this directory itself
(its host is ours, or its live challenge pays our address) — our verdict on our
own routes is not an independent check, so it says self_referential: true
and is never better than caution (since 2026-09-29);
status is unverified; the payTo
changed within the last 7 days on a listing observed for longer than that; no
successful paid delivery is on record from our scout; probe score below 0.90;
p95 latency above 5,000 ms; or the live challenge price differs from the
declared price by more than 2%.
pay otherwise, with the reason “verified, delivering when paid, stable payTo, price as declared.”
The working group’s verifier object (since 2026-09-29).
The verdict also carries verification, the evidence object in the x402 working
group’s draft verifier schema (wg-domain-discovery
PR #6; the response names the revision it follows, verification_schema). The two answer different
questions: verification.verdict (true, false, inconclusive) is whether the claim about the
payment terms holds; pay, caution, avoid is our recommendation, and cites it. An endpoint
whose advertised terms hold but which took our payment and failed to deliver reads true
(probe-ok, evidence purchase) there and avoid here: both are right. The first rule that applies:
we operate the endpoint → inconclusive, self-referential; the latest probe got no
response (a timeout or a failed connection) → inconclusive, probe-unreachable; a price the
seller declared differs from the live one by more than 2% → false, terms-mismatch with the
amount in detail (not for listings we indexed ourselves: that price is ours, not the seller’s
claim); our scout’s purchase settled and delivered, and neither payTo nor price has changed since
→ true, settled-under-terms; otherwise, when the latest probe read the terms, →
true, probe-ok. evidence is the strongest level we hold: purchase when our scout
settled a payment under the current payTo and price (delivered or not), else probe when the terms
were ever read live, else none; never lowered to a fresher weaker check. observed_at is when
that evidence was obtained and checked_at our latest probe. digest is the accept we checked
(the cheapest; accepts_index says which), with extra as the challenge sent it (left out
when that object is over 2,000 characters, rather than shown in part). Where the draft
set has no code for what we saw (not probed yet; no terms ever read; a latest probe that got an answer
but not the terms, such as a 404, a 429 or a 200 without payment; or a multi-option challenge whose
index is not recorded yet), verification is null and verification_note says why: we do not use a
code that means something else. Except on our own routes, attestation points to a signed record at
/v1/attestations/:id: a compact JWS (EdDSA, the key in /.well-known/jwks.json) over the
object, the resource’s domain and ours, and valid_until (24 hours). It is left out on a
self-referential verdict, where a comparison of our own subdomains would look independent.
The verdict is a point-in-time reading of our instruments, not a guarantee;
the fields it is computed from are returned beside it so an agent can apply a
stricter rule of its own. Thresholds are listed here so that when one changes,
the change is a dated entry under Method changes like any other. Sales are
recorded (listing, payer, tx, verdict) and the count is published on
/v1/stats. Unpaid requests receive 402 with the
PaymentRequired object in the PAYMENT-REQUIRED
header and in the body, byte-identical.
2026-09-02 — resolver errors are ours. A DNS lookup that
errors (rate limit, SERVFAIL, transport) is dns_lookup_error:
recorded, never cached, no effect on score, streak or status. A genuine
no-records answer is cached 5 minutes, not 1 hour. probe_pass_rate_24h
excludes our own lookup errors and says so in probe_pass_rate_24h_basis.
Flips: none directly; ends the 512-transition loop in correction #14.
2026-09-02 — leaving failing takes three consecutive
passes, the same evidence as earning verified; one pass used
to suffice, and a listing passing one probe in ten cycled twelve times a day.
A listing in failing is re-probed every cycle for its first two
hours only, then at normal cadence. Flips: currently failing listings need
three passes from 2026-09-02 07:35 UTC to recover.
2026-09-02 — edit rejections are logged and explained. Every
rejected edit (400 or 401) is recorded with its reason and the field names sent.
An unchanged endpoint_url in a re-sent record is ignored rather
than refused. A 401 for a token retired by a re-claim names the retirement
date. Success returns applied and ignored. Flips:
none; a seller had sent a retired token 310 times with no way to learn why.
2026-09-01 — payment-address stability guard. Three or more
distinct payTos across a listing's last ten probes marks it
payto_unstable; the hijack reset is skipped and the detail
response tells agents to verify the address in the challenge they receive
immediately before paying. Flips: one listing (1,036 resets since 2026-08-19)
stops resetting; six legitimate one-time rotations catalogue-wide are
unaffected.
2026-09-01 — demand feed ranks by persistence. Minimum query length 5; searches from clients that also edit listings excluded and published alongside; a term appears only when searched on two or more distinct days by non-seller clients; IP counts named as IPs. Flips: the previous top ten (the alphabet, one seller's self-checks, our own test burst) leaves the feed.
state of the network · methodology · stats · sellers · integrate · terms · privacy
The paid purchase history
GET /v1/listings/:id/purchases — $0.01 USDC on Base, x402, same
settlement flow as the verdict. It returns every real purchase this directory has made
of one listing, oldest first: outcome, the live quote against the listed price, the
settlement transaction, JSON and schema validity, declared freshness, and a normalised
hash of the body (clock and request-id fields stripped at every level, so two
purchases of the same data hash the same). Consecutive purchases are compared as pairs:
delivery held, regressed, recovered or never; price
held or moved, with the size of the move so a floating quote is not read as a
repricing; body same, changed, or no_hash where a side predates
2026-09-10. The verdict is a rule applied to this series; this is the series. What is
left out: the seller’s response body, any third party’s verdict, and the
identity of anyone who paid for a verification. Each row carries its
record_hash and attestation_signer, so a reader can check it
against the published log. Nothing here can be bought or edited by a seller.
Since 2026-09-15.