The machinery works. The questions didn't deserve you.
Placement ran on items derived from public LLM benchmarks. Audit that bank and its provenance falls apart: scraped test-prep, pipelined into a benchmark, deployed because it was convenient. Nobody along that chain ever asked whether an item was fit to measure a person. Roughly one in twelve to one in twenty knowledge items is outright broken — and the breakage concentrates at the top of the scale, because an ambiguous item and a hard item look identical to a difficulty estimator. The hardest rungs were the least trustworthy ones. That is the opposite of what a ladder is for.
So the whole bank is pulled from human measurement — not the flagged subset, all of it. It keeps exactly one job: calibrating models against each other, below.
Each source gets audited and signed off before a human sees it. An item that cannot survive scrutiny cannot be allowed to score a person — a bad item at the top of a scale doesn't merely add noise, it manufactures a false ceiling.
Maximum-likelihood Rasch fits on the benchmark bank (MMLU-Pro / GPQA / MATH-hard, 20 open-weight models, 2024-era), computed from per-item outcomes rather than reported headline scores. Valid for comparing these models to each other on this bank, and for nothing else. Not a human scale.
W is the Rasch scale used by the Woodcock–Johnson battery: 10 W = a factor of 3 in the
odds of solving an item (Woodcock & Dahl 1971; W = 500 + 9.1024·logit). It is
sample-free — it measures against item difficulty, not against a population — which is why a
child, an adult and a model can occupy the same column without anyone being converted into a
percentile of anyone else.
brainelo (BE) = 1000 + 19.0849·(W − 500), gauged so 400 BE is
one decade of odds, the way Elo is. Deviation IQ is a rank laundered through an assumed normal
curve; this is a quantity. The difference matters most exactly where the tails are — which is
where every existing instrument stops resolving and reports its own edge as though it were yours.