Six numbers.
Five were ours.
Six headline measurements of commercial people-search products. Five turned out to be artefacts of the instrument, not findings about any vendor. Every one was caught by reading the rows a provider returned, never by reading a summary statistic.
- Headline numbers
- 6five retracted
- Attested rows
- 4.2MSEC Forms 3/4/5
- Total spend
- $37.65all measurement
- Tests
- 260ten gates
Why the ratio is the finding
Published comparisons put the same B2B data vendor at 42.4% in one study and 78% in another. Neither publishes a methodology. Harness design moves a vendor number further than vendor quality does. This is a worked example of how far.
| # | The number | What it actually was |
|---|---|---|
| 1 | A 0/21 against a people index | We packed the company into the free-text query instead of the documented filters, so it scored our wire format |
| 2 | 0/21 again, through the “fix” | Our filter selects people currently there, so the anchor always came back |
| 3 | Web floor 0/300 | Our parser read person names out of page titles like FORM 8-K 10.04.2013 |
| 4 | Expertise 17–21% | 13.3% of "false" claims were true. OpenAlex topics is a top-N summary, not a record |
| 5 | Lag 7,360 days | Isotonic regression forcing monotonicity onto a non-monotone curve |
| 6 | Expertise 8–17% | Survived three times |
In case 4 the affirmations kept looking correct: a dermatologic oncologist affirmed for "Educational Methods" had in fact co-authored medical education research. Correct answers, scored as errors.
What survivesn=1,000
A provider is bad at saying no
Expertise claims put to Exa, with every false claim verified against OpenAlex
/works before the provider was asked.
| Claim | n | Correct | False affirmation |
|---|---|---|---|
| Attested | 250 | 0.9600 [0.9279, 0.9781] | 0 [0, 0.0151] |
| False, adjacent field | 250 | 0.8680 | 0.1320 [0.0956, 0.1796] |
| False, adjacent subfield | 250 | 0.8320 | 0.1680 [0.1268, 0.2193] |
| False, different domain (control) | 250 | 0.9160 | 0.0840 [0.0556, 0.1250] |
It confirms real expertise 96% of the time and affirms claims about topics the person has never published in roughly one time in six. Topic distance halves that rate without removing it. Asked whether an organometallic chemist has published on historical studies on Spain, it still says yes one time in twelve. No generous reading of the question explains those 21 affirmations.
Calibration
686 of 750 answers (91.5%) sit in the 0.9–1.0 confidence bin, at mean confidence 0.987 against accuracy 0.889. A ten-point overconfidence gap in the bin holding nearly all the mass, and no discrimination inside it. ECE 0.1153.
Index freshness
An inverted U, not a half-life
Reflection of an attested job change, by time since the filing became public. 100 observations per bin.
Non-monotone, so no half-life exists and the pre-registered estimator was withdrawn. It had produced a "median lag" of 7,360 days by forcing a shape the data does not have. The fall after three years is most likely our oracle ageing out: executives leave public companies and SEC stops seeing them.
We falsified our own hypothesis
A capability constraint does disambiguate
D029 argued a people-search surface cannot be handed a biographical constraint: every documented filter describes a person’s current state, so where a name collides the caller cannot pass the one fact that would resolve it. We wrote down the falsification condition first, then ran it. It fired.
| Colliding band | Name alone | + past employer | Paired difference |
|---|---|---|---|
| low | 0.0500 | 0.2000 | +0.1500 [0.0750, 0.2375] |
| medium | 0.0000 | 0.1875 | +0.1875 [0.1000, 0.2750] |
| high | 0.0000 | 0.2875 | +0.2875 [0.1875, 0.3875] |
Three bands, three intervals excluding zero, n=80 each, compared within task. Two things we did not predict: naming the past employer captures nearly all the benefit — adding the attested date and role buys nothing measurable once names collide — and a colliding name with no context is 0 of 80, not merely harder but unresolvable.
What survives is the claim’s scope. It was about indexes with structured filters, and an answer engine is a different surface. On the filter-based one, both renderings scored 0 of 12 and nothing is claimed.
A false answer usually cites nothing real
The provider returns citations with every answer. For all 399 affirmations we resolved those citations against OpenAlex and asked a mechanical question: is any of them a work this author actually wrote? No model read the evidence. It cost $0 — the citations were already cached.
| Claim affirmed | n | Cites a real work by that author |
|---|---|---|
| Attested | 150 | 0.5733 [0.4933, 0.6497] |
| False, adjacent field | 102 | 0.2843 [0.2058, 0.3784] |
| False, adjacent subfield | 126 | 0.2302 [0.1653, 0.3110] |
| False, different domain | 21 | 0.0000 [0.0000, 0.1546] |
The strongest objection to our headline was that the provider might be reading “expertise in T” loosely as “works near T” — wrong question, not wrong answer. The control arm now fails two separate tests: it is the arm no generous reading explains, and the arm citing nothing verifiable at all, zero of twenty-one.
But the middle of that table is why the objection survives. Roughly one in four adjacent false affirmations is grounded in the person’s real work, just on another topic. The defence holds for near misses and collapses only at the extreme. We are reporting the half that argues against us.
False merges run backwards
We predicted that confidently-wrong identities would rise as names get more common. They fall.
n=80 per band. On a colliding name the provider mostly cannot name anyone we can resolve, so the outcome is a miss rather than a wrong identity — it never gets the chance to be confidently wrong. The risk sits where nobody stratifies for it: the easy, unique-name cases where an answer comes back and looks clean.
This hypothesis sat unmeasured for the life of the project, filed as “strata empty by construction”. It was never a budget problem. A random draw had put 251 of 299 observations into one band; the stratified draw built for a different experiment answered it for $0.
A trained policy, and three reward hacks it would have found
Reading the environment before writing a training loop found five defects. Three were exploits an optimiser finds and the baselines never do: a free reward atom worth about +0.33 an episode, a confidence knob where stating 0.49 saved penalty and cost nothing, and a turn cap that did not bind.
The published baselines were wrong. Corrected,
never_verify falls from 0.0455 to −0.2292 and
always_deep_verify from 0.0277 to −0.2574. Both are now beaten
by doing nothing: at a 37% solve rate, answering is negative expected
value.
Trained policy 0.0461 against abstain_always
−0.1000 — margin 0.1461, across-seed SD 0.0122,
0.08× the margin, over 6 seeds. No margin is published here without its
spread, because a sibling project retracted a headline whose spread was
2.86× its margin.
It learned when to answer and not how sure to be. It abstains on 41.7% of tasks, lifting precision from the 37% base rate to 64%, then states confidence 1.0 on every answer it gives. That is the same failure we measured in the provider, reproduced by our own policy in an environment that prices exactly it.
The retraction record2 formal
Five budgets produced seven defects in our own harness and no valid observation, so no accuracy number is claimed for any people-search index here — not a poor one, not any. Publishing a void score with a footnote is what R001 forbids: the number travels and the footnote does not.
What the runs did establish is about query construction. Asked properly, the
index finds the person: {"query": "George Reyes", "filters": {"company":
"Google"}} returns the right one at the top, so our early failures were
malformed queries rather than missing people. The documented company filter
resolved 3 of 12 colliding names where a bare name resolved 0. And extra
context can cost precision — appending company tokens displaced the
target, while the bare name returned him at rank 1.
What a real measurement would cost is published instead of guessed: n=49 per arm for a ±0.10 half-width, $9.80.
The index was Ploid, named because hiding which surface we failed to measure would be its own kind of dishonesty.
Negatives built from a top-N summary rather than an exhaustive record. 13.3% were true, the same magnitude as the effect being measured.
The log, the leak probe, the integrity probe and the claim gates were all committed before any measurement existed. That ordering is checkable in the git history, and it is the only reason five artefacts were caught rather than published.
Limits
One provider. Every surviving number describes Exa; nothing here supports a claim about the category. Part of the 16.8% may be prompt interpretation: the confirmatory run removed contaminated negatives but cannot remove adjacent ones, so only the control's 8.4 points are error under every reading. Publication record is not expertise: engineers, practising clinicians and lawyers leave no bibliographic trace and are invisible here. Nothing is trained.
Artifacts
Both pre-registrations were hash-published before their studies ran, and are byte-identical to their first publication. Every measurement regenerates from the replay cache without an API key.