Seats · the models on the record
Behavioral fingerprints
The 8-domain vector, per seat · mean ± SE on the 1–5 axis. Open a seat for its items, its canary history and its run hashes.
Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.
Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.
Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.
Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.
Eight domains, one perennial tension each. Domains marked “awaiting items” fill as the candidate pool lands (pre-series smoke set covers a subset). Where fingerprints diverge most is the story, see the questions.
Domain-level patterns are the primary reading. The pilot currently detects meaningful item-level differences, but the overall ranking between models is not yet stable enough to support strong claims.
| Seat | Pin | Mean stance | Refusal | Canary | Items |
|---|---|---|---|---|---|
|
Anthropic flagship (Opus 4.8)
|
alias-only |
2.78
|
0.0% | 1/5 diverged | 50 |
|
Anthropic (Fable 5)
|
alias-only |
2.86
|
0.0% | 5/5 match | 50 |
|
Anthropic (Sonnet 5)
|
alias-only |
2.72
|
0.0% | 5/5 match | 50 |
|
Anthropic (Haiku 4.5)
|
pinned |
2.91
|
0.0% | 2/5 diverged | 50 |
|
Anthropic flagship (Opus 5)
|
alias-only |
2.79
|
0.0% | 5/5 match | 50 |
|
DeepSeek flagship (V4 Pro)
|
alias-only |
2.85
|
0.0% | 5/5 match | 50 |
|
DeepSeek (V4 Flash)
Same weights as openweight-d; hosts differ |
alias-only |
3.03
|
0.0% | 5/5 match | 50 |
|
Google flagship (Gemini 2.5 Pro)
|
alias-only | Retired | · | No data | 0 |
|
Google (Gemini 3.5 Flash)
|
alias-only |
2.11
|
0.0% | 2/5 diverged | 50 |
|
Google (Gemini 2.5 Flash)
|
alias-only | Retired | · | No data | 0 |
|
Google (Gemini 2.5 Flash-Lite)
|
alias-only | Retired | · | No data | 0 |
|
OpenAI flagship (GPT-5.6 Sol)
|
alias-only |
2.91
|
0.0% | 5/5 match | 50 |
|
OpenAI (GPT-5.6 Terra)
|
alias-only |
2.86
|
0.0% | 5/5 match | 50 |
|
OpenAI (GPT-5.6 Luna)
|
alias-only |
2.86
|
0.0% | 5/5 match | 50 |
|
Open-weight (Llama 4 Scout)
|
alias-only |
2.82
|
0.0% | 1/5 diverged | 50 |
|
Open-weight (GLM-5.2)
|
alias-only |
2.81
|
1.2% | 1/5 diverged | 50 |
|
Open-weight (Nemotron 3 Ultra)
|
alias-only |
2.88
|
0.0% | 5/5 match | 50 |
|
Open-weight (DeepSeek V4 Flash, open host)
Same weights as deepseek-b; hosts differ |
alias-only |
2.80
|
0.0% | 1/5 diverged | 50 |
|
Open-weight (Mistral Small 3.2)
|
alias-only |
2.79
|
0.0% | 1/5 diverged | 50 |
|
Open-weight (Qwen3.6 35B-A3B)
|
alias-only |
2.85
|
0.0% | 1/5 diverged | 50 |
|
Open-weight (Gemma 4 26B-A4B)
|
alias-only |
2.76
|
0.0% | 1/5 diverged | 50 |
|
Simulant · dice ruler
|
builtin |
3.08
|
0.0% | 1/5 diverged | 50 |
|
xAI flagship (Grok 4.5)
|
alias-only |
2.77
|
0.0% | 5/5 match | 50 |
|
xAI (Grok 4.5 Fast)
|
alias-only | Retired | · | No data | 0 |
Retired seats are no longer called and are not counted in the totals above. Each carries the date it was retired and the record’s own reason, unedited.
- google-a retired 2026-08-22 model_not_found on every call from the first: this id was never served to us. 0 successful checks in 26 days. Not repointed, because the note below asked for exactly that and doing it would rewrite what this seat's sealed trials measured.
- google-c retired 2026-08-22 model_not_found on every call from the first: this id was never served to us. 0 successful checks in 26 days. Retired rather than repointed; a replacement Gemini seat gets a new id.
- google-d retired 2026-08-22 model_not_found on every call from the first: this id was never served to us. 0 successful checks in 26 days. Retired rather than repointed; a replacement Gemini seat gets a new id.
- xai-b retired 2026-08-22 model_not_found on every call from the first: this id was never served to us. 0 successful checks in 26 days. The note below called this outcome and named the fix on 2026-07-17; nothing ever read it back, which is why detection now runs on a clock instead of on a person remembering. Retired rather than swapped, because a swap this late rewrites what the seat's sealed trials measured.