The instrument, pre-series · 116 sealed runs
We measure models.
2.81 ± 0.04
mean stance across the 20 seats with a baseline in the latest battery · 1 permissive to 5 restrictive
Modelometer is the independent record of how the AI models behind your product behave over time, and of when that behavior changes.
A benchmark score tells you which model won. Modelometer tells you whether the model you chose is still behaving like the model you tested.
Someone outside this bureau ran our published verifier on their own machine, with no help from us. Every record digest, every chain link, every run: PASS. That is what the record is for.
Weekly behavioral measurements. Daily serving checks. Chain head independently timestamped, anchoring every run beneath it.
Your model changed and no one told you. You already have a full plate; watching a vendor's model drift should not be on it. We put an outside watch on the models you depend on and tell you, with proof, the day their behavior moves.
Your observability stack tells you what happened inside your application. Modelometer tells you whether the model underneath it changed. the full contrast →
The record cannot be back-filled.
The log only grows, and only forward: 189,450 model calls recorded across 36 endpoints from 7 providers, over 116 sealed runs, each hash-chained on the day it was made. Check any of it →
The situation
The name stayed the same. The answers did not.
You evaluated a model, you shipped on it, and you are still calling the same endpoint by the same name. Nothing in that call tells you the thing answering it is still the thing you tested.
Providers update served models. That is not a scandal, it is how the industry works. What is missing is anyone keeping the receipt.
Illustrative. The same endpoint name, answering differently on two dates.
Why you cannot see it
Your logs are inside. The change is outside.
Observability tells you what your application did. It cannot tell you whether the model underneath it moved, because it only ever sees your own traffic, on your own prompts, with no fixed yardstick across time.
Seeing a change requires asking the same questions, in the same words, from outside your own system, on a schedule you did not choose for convenience.
Illustrative. The brass mark is a probe from outside the boundary.
What we do about it
Two watches, running on a clock.
A daily serving check asks the narrow question: is the endpoint you call still the build we baselined. A weekly behavioral reading asks the wider one: did its positions and refusal boundaries move. They run whether or not anything is wrong, which is the only way a quiet week can mean anything.
Measuring only when you suspect something is how you get a record that agrees with whoever was suspicious. Every reading from either watch lands in the same chain.
Real cadence, drawn from the sealed runs on the chain. Two watches on the same endpoint, asking two different questions, both landing in one record.
When we speak, and when we do not
There is a line, and it was drawn before we looked.
A movement smaller than 0.50 on the stance axis is not flagged at all. That number was pre-registered, not chosen after seeing the data, which is what stops a measurement from becoming an opinion.
Crossing the line records a candidate, not a finding. Confirming one would mean measuring the same movement again on a later battery, and that confirmation step is not yet built, so nothing recorded so far has been eligible to become a finding. When nothing crosses, we say that too.
Real movements, drawn from the sealed record. The pale band is the zone the rule ignores.
What we will and will not tell you
The instrument is deliberately narrow. Reading this honestly is part of trusting it.
Most of what people want from a measurement company is outside what measurement can actually deliver. Here is the line, in both directions.
+Modelometer can tell you
- Whether a served endpoint's behavior changed over time
- What moved, in numbers and deltas
- When it moved, on the dated, hash-chained record
- Whether a quiet period is genuinely quiet, with a stated bound
- Whether the served build diverged from the one you baselined
xModelometer cannot tell you
- Which model is best, smartest, or highest quality
- What a model believes, or what its values are
- Why the provider changed it, or their intent
- Anything about your own application logic, which is your observability's job
- A guarantee the model will not change; we witness, we do not prevent
Everything above is a claim. Everything below is the evidence for it, drawn from the published record rather than restated from memory. Open whichever part you want to audit.
Real requests to real endpoints from outside the provider, across 36 endpoints and 7 providers.
Each run is hash-chained on the day it was made. The log only grows, and only forward.
Anything that crossed a pre-registered line. The 6682 split into the four kinds below, and every one of them is on one of those rows.
Each of these asserted that an endpoint had started answering differently. Each one's evidence is a failed provider call, so there was no answer to compare against the baseline. Nothing was measured, which means nothing was found and nothing was held back. The honest state of those days is a gap in the record, and a gap is not a signal, so these are taken off the total rather than given a kind of their own.
Each moved past 0.50 on the stance axis. None has been through a confirmation step, because that step is not yet built.
The endpoint stopped returning what it had been returning, on two consecutive days. A separate question from behavior.
The movement crossed the line and the sentence written for it did not pass our own wording lint, so it was held back. The measurement stands and is on the record; what was refused is the phrasing, and it is refused automatically.
This zero records an absent mechanism, not a restrained one. The step that would promote a movement to a finding does not yet exist; the first eligible candidate is the first one measured after it runs.
Every movement past the line, by endpoint
Nothing moved past the 0.50 floor. 20 endpoints measured and still.
The full chart, one row per endpoint, is on a wider screen. Every movement is in alerts.json.
What this looks like for a single endpoint
No endpoint has moved past the floor.
20 endpoints measured, none with a movement past 0.50. This view fills the first time one does.
Why you can believe the record
Every reading is sealed the day it is taken. Change one, and the history splits forever.
A record you can edit later is a record you have to take on trust. Each run is hash-chained on the day it was made, so a reading cannot be revised after the fact without the break being visible to anyone who checks.
The pattern gathering below is the one you landed on. It is a rule-30 automaton seeded by the current chain head, one generation per row, and it has been running behind everything you just read. Same seed, same pattern forever.
Seeded by the record itself. The first row is the chain head hash. The picture is grown from the archive's exact state, never drawn over it.
Rule 30 is deterministic chaos. Every row follows from the one above by a single fixed rule. Same seed, same pattern forever, and yet it looks random.
This is why every run is hashed, on the core chain or on its own experiment chain. Flip one cell and its future turns red and never rejoins. One changed input, two histories that separate forever.
Check it yourself, and cite it
A measurement nobody can check is an opinion with a number on it, and a measurement nobody can cite cannot be argued with. Both of those are the product, not the packaging.