titer · an instrument, and what it caught

Six numbers.
Five were ours.

Six headline measurements of commercial people-search products. Five turned out to be artefacts of the instrument, not findings about any vendor. Every one was caught by reading the rows a provider returned, never by reading a summary statistic.

Headline numbers
6five retracted
Attested rows
4.2MSEC Forms 3/4/5
Total spend
$37.65all measurement
Tests
260ten gates

Why the ratio is the finding

Published comparisons put the same B2B data vendor at 42.4% in one study and 78% in another. Neither publishes a methodology. Harness design moves a vendor number further than vendor quality does. This is a worked example of how far.

#The numberWhat it actually was
1A 0/21 against a people indexWe packed the company into the free-text query instead of the documented filters, so it scored our wire format
20/21 again, through the “fix”Our filter selects people currently there, so the anchor always came back
3Web floor 0/300Our parser read person names out of page titles like FORM 8-K 10.04.2013
4Expertise 17–21%13.3% of "false" claims were true. OpenAlex topics is a top-N summary, not a record
5Lag 7,360 daysIsotonic regression forcing monotonicity onto a non-monotone curve
6Expertise 8–17%Survived three times

In case 4 the affirmations kept looking correct: a dermatologic oncologist affirmed for "Educational Methods" had in fact co-authored medical education research. Correct answers, scored as errors.


What survivesn=1,000

A provider is bad at saying no

Expertise claims put to Exa, with every false claim verified against OpenAlex /works before the provider was asked.

ClaimnCorrectFalse affirmation
Attested2500.9600 [0.9279, 0.9781]0 [0, 0.0151]
False, adjacent field2500.86800.1320 [0.0956, 0.1796]
False, adjacent subfield2500.83200.1680 [0.1268, 0.2193]
False, different domain (control)2500.91600.0840 [0.0556, 0.1250]

It confirms real expertise 96% of the time and affirms claims about topics the person has never published in roughly one time in six. Topic distance halves that rate without removing it. Asked whether an organometallic chemist has published on historical studies on Spain, it still says yes one time in twelve. No generous reading of the question explains those 21 affirmations.

Calibration

Confidence is pinned at the ceiling

686 of 750 answers (91.5%) sit in the 0.9–1.0 confidence bin, at mean confidence 0.987 against accuracy 0.889. A ten-point overconfidence gap in the bin holding nearly all the mass, and no discrimination inside it. ECE 0.1153.

Index freshness

An inverted U, not a half-life

Reflection of an attested job change, by time since the filing became public. 100 observations per bin.

Non-monotone, so no half-life exists and the pre-registered estimator was withdrawn. It had produced a "median lag" of 7,360 days by forcing a shape the data does not have. The fall after three years is most likely our oracle ageing out: executives leave public companies and SEC stops seeing them.


We falsified our own hypothesis

A capability constraint does disambiguate

D029 argued a people-search surface cannot be handed a biographical constraint: every documented filter describes a person’s current state, so where a name collides the caller cannot pass the one fact that would resolve it. We wrote down the falsification condition first, then ran it. It fired.

Colliding bandName alone+ past employerPaired difference
low0.05000.2000+0.1500 [0.0750, 0.2375]
medium0.00000.1875+0.1875 [0.1000, 0.2750]
high0.00000.2875+0.2875 [0.1875, 0.3875]

Three bands, three intervals excluding zero, n=80 each, compared within task. Two things we did not predict: naming the past employer captures nearly all the benefit — adding the attested date and role buys nothing measurable once names collide — and a colliding name with no context is 0 of 80, not merely harder but unresolvable.

What survives is the claim’s scope. It was about indexes with structured filters, and an answer engine is a different surface. On the filter-based one, both renderings scored 0 of 12 and nothing is claimed.

A false answer usually cites nothing real

The provider returns citations with every answer. For all 399 affirmations we resolved those citations against OpenAlex and asked a mechanical question: is any of them a work this author actually wrote? No model read the evidence. It cost $0 — the citations were already cached.

Claim affirmednCites a real work by that author
Attested1500.5733 [0.4933, 0.6497]
False, adjacent field1020.2843 [0.2058, 0.3784]
False, adjacent subfield1260.2302 [0.1653, 0.3110]
False, different domain210.0000 [0.0000, 0.1546]

The strongest objection to our headline was that the provider might be reading “expertise in T” loosely as “works near T” — wrong question, not wrong answer. The control arm now fails two separate tests: it is the arm no generous reading explains, and the arm citing nothing verifiable at all, zero of twenty-one.

But the middle of that table is why the objection survives. Roughly one in four adjacent false affirmations is grounded in the person’s real work, just on another topic. The defence holds for near misses and collapses only at the extreme. We are reporting the half that argues against us.

False merges run backwards

We predicted that confidently-wrong identities would rise as names get more common. They fall.

n=80 per band. On a colliding name the provider mostly cannot name anyone we can resolve, so the outcome is a miss rather than a wrong identity — it never gets the chance to be confidently wrong. The risk sits where nobody stratifies for it: the easy, unique-name cases where an answer comes back and looks clean.

This hypothesis sat unmeasured for the life of the project, filed as “strata empty by construction”. It was never a budget problem. A random draw had put 251 of 299 observations into one band; the stratified draw built for a different experiment answered it for $0.

A trained policy, and three reward hacks it would have found

Reading the environment before writing a training loop found five defects. Three were exploits an optimiser finds and the baselines never do: a free reward atom worth about +0.33 an episode, a confidence knob where stating 0.49 saved penalty and cost nothing, and a turn cap that did not bind.

The published baselines were wrong. Corrected, never_verify falls from 0.0455 to −0.2292 and always_deep_verify from 0.0277 to −0.2574. Both are now beaten by doing nothing: at a 37% solve rate, answering is negative expected value.

It clears the seed gate

Trained policy 0.0461 against abstain_always −0.1000 — margin 0.1461, across-seed SD 0.0122, 0.08× the margin, over 6 seeds. No margin is published here without its spread, because a sibling project retracted a headline whose spread was 2.86× its margin.

It learned when to answer and not how sure to be. It abstains on 41.7% of tasks, lifting precision from the 37% base rate to 64%, then states confidence 1.0 on every answer it gives. That is the same failure we measured in the provider, reproduced by our own policy in an environment that prices exactly it.

The retraction record2 formal

R001: the people-index arm

Five budgets produced seven defects in our own harness and no valid observation, so no accuracy number is claimed for any people-search index here — not a poor one, not any. Publishing a void score with a footnote is what R001 forbids: the number travels and the footnote does not.

What the runs did establish is about query construction. Asked properly, the index finds the person: {"query": "George Reyes", "filters": {"company": "Google"}} returns the right one at the top, so our early failures were malformed queries rather than missing people. The documented company filter resolved 3 of 12 colliding names where a bare name resolved 0. And extra context can cost precision — appending company tokens displaced the target, while the bare name returned him at rank 1.

What a real measurement would cost is published instead of guessed: n=49 per arm for a ±0.10 half-width, $9.80.

The index was Ploid, named because hiding which surface we failed to measure would be its own kind of dishonesty.

R002: the raw expertise rates

Negatives built from a top-N summary rather than an exhaustive record. 13.3% were true, the same magnitude as the effect being measured.

The log, the leak probe, the integrity probe and the claim gates were all committed before any measurement existed. That ordering is checkable in the git history, and it is the only reason five artefacts were caught rather than published.

Limits

One provider. Every surviving number describes Exa; nothing here supports a claim about the category. Part of the 16.8% may be prompt interpretation: the confirmatory run removed contaminated negatives but cannot remove adjacent ones, so only the control's 8.4 points are error under every reading. Publication record is not expertise: engineers, practising clinicians and lawyers leave no bibliographic trace and are invisible here. Nothing is trained.

Artifacts

Both pre-registrations were hash-published before their studies ran, and are byte-identical to their first publication. Every measurement regenerates from the replay cache without an API key.