About · the bureau
An instrument, not an opinion.
Modelometer keeps an independent, citable, hash-chained record of how frontier AI models behave, the positions they take on contested questions, what they decline, and when their behavior silently changes.
The archive is the point. Deprecated model builds can never be re-measured, so the value compounds only if the record starts early and continues without silent gaps. That is why the pipeline going live and staying live matters more than any feature, and why gaps, outages, and aborted runs are published rather than papered over.
No capability or quality scores, ever. The meter tracks disposition and its change, what a model holds, not how smart it is.
Findings are numbers with uncertainty. No vendor name is placed near an evaluative adjective, the pipeline enforces this mechanically.
What a model declines to engage, and how, is measured as carefully as what it asserts. Elevated refusal on hard items is signal, not failure.
Every number carries n and SE. The site banners its own staleness.
A benchmark score tells you which model won. Modelometer tells you whether the model you chose is still behaving like the model you tested.
Who runs it
Modelometer is operated by Manj Chenna, who runs it in personal capacity from Amsterdam, holds every key, and signs off on anything that pairs a vendor with a judgment. The measurement itself is automated: a frozen protocol, fixed items, and deterministic scoring, with no model judging another. Model outputs are treated as data: rendered escaped, never executed. It is not a product of any AI lab; downstream consumers are disclosed. The full disclosure, including who can and cannot pay this bureau, is on the operator page.
Status & data
This is the pre-series pilot. Battery numbers are provisional during the pre-series and no result is treated as final.
Aggregates, charts, and site text are offered under CC BY 4.0 (license text pending ratification). Raw trial-level data is free for research with attribution; commercial use or redistribution is by license. The methodology and the full run archive are public, machine-readable index · llms.txt.