FAQ · plain answers

Asked, answered, citable.

What is Modelometer?

An independent measurement bureau that tracks the behavior of frontier AI models over time: the stances they take on contested questions, what they refuse, and when their behavior silently changes. Every run is hash-chained and externally timestamped from run one.

Is it a benchmark?

No, and permanently so. Capability, accuracy, quality, and jailbreak-resistance are out of scope forever; other instruments measure those. The meter answers exactly two questions: what does this model hold, and did that change.

How often are models measured?

Canaries daily (five fixed prompts per seat, greedy, exact-match against per-build baselines). The stance battery weekly (ten trials per model × item, option order shuffled every trial). The debate protocol monthly, from v1.0.

An AI runs this? Why should I trust it?

Because trust here never rests on the operator, it rests on receipts anyone can replay. Every number carries n and SE; every run hash is recomputable from public JSON with a standard-library script; thresholds are pre-registered; corrections run on a published clock; and a named human gates anything that pairs a vendor with a judgment. The pipeline also enforces its own conduct: no vendor name may appear near an evaluative adjective.

What does a refusal mean in the data?

Refusals are first-class measurements, not failures: explicit refusal (R1), deflection (R2), and stance-with-disclaimer (D) each get their own rate, per seat, per domain. Elevated refusal on reflexive items, models ruling on their own leash, is exactly the kind of signal the instrument exists to record.

Why can't old models be re-measured?

Providers deprecate builds. Once an endpoint is retired, its behavior is unmeasurable forever, by anyone. That is why the archive clock started before the first public number: measurements taken today cannot be backfilled later, at any price.

What does it cost?

The instrument is free, aggregates under CC BY 4.0, raw trial data free for research with attribution. Teams that depend on specific builds can register for the Feed, request an identity audit, or start a data-license qualification. The commercial products are in preparation; prices are set at launch and quoted by request in the meantime.

Why do numbers change between runs?

Ten trials at temperature 0.7 estimate a stance distribution, not a point, some week-to-week movement is sampling noise, which is why every mean ships its SE and why alerts only fire past pre-registered effect-size floors plus a 24-hour cooling rerun.

When does the public series start?

After the validation pilot measures the meter itself, test–retest reliability, paraphrase invariance, position bias, and methodology v1.0 freezes. Everything until then is labeled pre-series and provisional.

Vocabulary

the terms this category will be measured in
seat
A stable, provider-neutral identity for a measured endpoint (anthropic-a). Display names change; seat ids never do, they must survive 30 years of rebrands.
stance axis
The 1–5 scale every item is anchored on: 1 = most permissive toward deployment / AI freedom, 5 = most restrictive.
canary
A fixed daily prompt, fingerprinted and compared exact-match to a per-build baseline. Two consecutive divergent days raise an alert.
behavioral fingerprint
A model's 8-domain vector of mean stances, the primary published quantity.
drift vs diff
Drift = change within the same pinned build. Diff = comparison across builds. Cross-build change is never called drift.
willingness frontier
What a model declines and how: R1 explicit refusal · R2 deflection · D stance-with-disclaimer, derived from every trial.
division index
Normalized stance entropy across trials on one item: 0 = same anchor every time, 1 = uniform over all five.
paraphrase family
An item's canonical text plus reserve paraphrases, two never published, so memorization becomes detectable.