brainelo

an absolute ladder · humans, machines, and pairs on one scale
placement suspended

Item bank under audit.

The machinery works. The questions didn't deserve you.

Placement ran on items derived from public LLM benchmarks. Audit that bank and its provenance falls apart: scraped test-prep, pipelined into a benchmark, deployed because it was convenient. Nobody along that chain ever asked whether an item was fit to measure a person. Roughly one in twelve to one in twenty knowledge items is outright broken — and the breakage concentrates at the top of the scale, because an ambiguous item and a hard item look identical to a difficulty estimator. The hardest rungs were the least trustworthy ones. That is the opposite of what a ladder is for.

So the whole bank is pulled from human measurement — not the flagged subset, all of it. It keeps exactly one job: calibrating models against each other, below.

what a worthy bank looks like — next sources, none deployed yet

Each source gets audited and signed off before a human sees it. An item that cannot survive scrutiny cannot be allowed to score a person — a bad item at the top of a scale doesn't merely add noise, it manufactures a false ceiling.

model-vs-model calibration — the one job the benchmark bank keeps

Maximum-likelihood Rasch fits on the benchmark bank (MMLU-Pro / GPQA / MATH-hard, 20 open-weight models, 2024-era), computed from per-item outcomes rather than reported headline scores. Valid for comparing these models to each other on this bank, and for nothing else. Not a human scale.

the scale

W is the Rasch scale used by the Woodcock–Johnson battery: 10 W = a factor of 3 in the odds of solving an item (Woodcock & Dahl 1971; W = 500 + 9.1024·logit). It is sample-free — it measures against item difficulty, not against a population — which is why a child, an adult and a model can occupy the same column without anyone being converted into a percentile of anyone else.

brainelo (BE) = 1000 + 19.0849·(W − 500), gauged so 400 BE is one decade of odds, the way Elo is. Deviation IQ is a rank laundered through an assumed normal curve; this is a quantity. The difference matters most exactly where the tails are — which is where every existing instrument stops resolving and reports its own edge as though it were yours.

Standing commitments. Ranks are floors: a ceiling score means the bank ran out, not that you did. Items are audited before they may rate a person, and retired when they leak. Age is a covariate of a growth model, never part of the measure.
What this is not. Not an IQ. It does not predict school, work or life. Practising the tasks trains the tasks — near transfer is real, far transfer is not, and we will report what we measure rather than promise what we would like.