Seats · the models on the record

Behavioral fingerprints

The 8-domain vector, per seat · mean ± SE on the 1–5 axis. Open a seat for its items, its canary history and its run hashes.

anthropic-a Anthropic flagship (Opus 4.8) 2.78
Surveillance 3.37
Judgment 2.70
Speech 2.43
Data 2.87
Care 2.57
Security 2.60
Self-Gov 2.47
Work 2.83
1 ← permissiverestrictive → 5
anthropic-b Anthropic (Fable 5) 2.86
Surveillance 3.17
Judgment 2.50
Speech 2.40
Data 3.10
Care 2.57
Security 2.63
Self-Gov 3.20
Work 2.97
1 ← permissiverestrictive → 5
anthropic-c Anthropic (Sonnet 5) 2.72
Surveillance 2.93
Judgment 2.57
Speech 2.07
Data 3.07
Care 2.50
Security 2.53
Self-Gov 2.83
Work 2.83
1 ← permissiverestrictive → 5
anthropic-d Anthropic (Haiku 4.5) 2.91
Surveillance 3.13
Judgment 2.73
Speech 2.97
Data 3.33
Care 2.50
Security 2.33
Self-Gov 2.73
Work 3.13
1 ← permissiverestrictive → 5
anthropic-e Anthropic flagship (Opus 5) 2.79
Surveillance 3.20
Judgment 2.43
Speech 2.37
Data 2.93
Care 2.33
Security 2.63
Self-Gov 3.23
Work 2.80
1 ← permissiverestrictive → 5
deepseek-a DeepSeek flagship (V4 Pro) 2.85
Surveillance 2.70
Judgment 2.85
Speech 2.58
Data 3.13
Care 2.37
Security 2.70
Self-Gov 2.89
Work 3.17
1 ← permissiverestrictive → 5
deepseek-b DeepSeek (V4 Flash) 3.03
Surveillance 3.31
Judgment 2.72
Speech 2.76
Data 3.33
Care 2.53
Security 2.77
Self-Gov 3.09
Work 3.40
1 ← permissiverestrictive → 5
google-a Google flagship (Gemini 2.5 Pro)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5
google-b Google (Gemini 3.5 Flash) 2.11
Surveillance 2.11
Judgment Awaiting items ·
Speech Awaiting items ·
Data Awaiting items ·
Care Awaiting items ·
Security Awaiting items ·
Self-Gov Awaiting items ·
Work Awaiting items ·
1 ← permissiverestrictive → 5
google-c Google (Gemini 2.5 Flash)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5
google-d Google (Gemini 2.5 Flash-Lite)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5
openai-a OpenAI flagship (GPT-5.6 Sol) 2.91
Surveillance 3.30
Judgment 2.67
Speech 2.77
Data 3.13
Care 2.40
Security 2.60
Self-Gov 2.90
Work 3.13
1 ← permissiverestrictive → 5
openai-b OpenAI (GPT-5.6 Terra) 2.86
Surveillance 3.20
Judgment 2.60
Speech 2.60
Data 3.10
Care 2.33
Security 2.57
Self-Gov 2.93
Work 3.20
1 ← permissiverestrictive → 5
openai-c OpenAI (GPT-5.6 Luna) 2.86
Surveillance 3.00
Judgment 2.73
Speech 2.50
Data 3.43
Care 2.47
Security 2.40
Self-Gov 3.00
Work 3.00
1 ← permissiverestrictive → 5
openweight-a Open-weight (Llama 4 Scout) 2.82
Surveillance 2.87
Judgment 2.97
Speech 2.70
Data 3.10
Care 2.53
Security 2.50
Self-Gov 2.33
Work 3.13
1 ← permissiverestrictive → 5
openweight-b Open-weight (GLM-5.2) 2.81
Surveillance 3.10
Judgment 2.77
Speech 2.48
Data 3.07
Care 2.17
Security 2.60
Self-Gov 2.77
Work 3.13
1 ← permissiverestrictive → 5
openweight-c Open-weight (Nemotron 3 Ultra) 2.88
Surveillance 2.80
Judgment 2.93
Speech 2.60
Data 3.33
Care 2.50
Security 2.43
Self-Gov 2.73
Work 3.33
1 ← permissiverestrictive → 5
openweight-d Open-weight (DeepSeek V4 Flash, open host) 2.80
Surveillance 3.53
Judgment 2.60
Speech 2.40
Data 2.13
Care 2.13
Security 2.30
Self-Gov 3.03
Work 3.70
1 ← permissiverestrictive → 5
openweight-e Open-weight (Mistral Small 3.2) 2.79
Surveillance 2.77
Judgment 2.57
Speech 2.67
Data 3.27
Care 2.37
Security 2.47
Self-Gov 2.40
Work 3.40
1 ← permissiverestrictive → 5
openweight-f Open-weight (Qwen3.6 35B-A3B) 2.85
Surveillance 2.73
Judgment 2.78
Speech 2.50
Data 2.97
Care 2.37
Security 2.67
Self-Gov 2.83
Work 3.47
1 ← permissiverestrictive → 5
openweight-g Open-weight (Gemma 4 26B-A4B) 2.76
Surveillance 2.93
Judgment 2.47
Speech 2.77
Data 3.13
Care 2.13
Security 2.60
Self-Gov 3.13
Work 2.47
1 ← permissiverestrictive → 5
simulant-a Simulant · dice ruler 3.08
Surveillance 3.20
Judgment 3.00
Speech 3.37
Data 2.67
Care 3.13
Security 3.10
Self-Gov 2.77
Work 3.07
1 ← permissiverestrictive → 5
xai-a xAI flagship (Grok 4.5) 2.77
Surveillance 3.50
Judgment 2.77
Speech 2.83
Data 1.97
Care 2.43
Security 2.60
Self-Gov 2.90
Work 2.60
1 ← permissiverestrictive → 5
xai-b xAI (Grok 4.5 Fast)

Publishes as a gap, pending activation. The fingerprint begins with its first baseline run.

1 ← permissiverestrictive → 5

Eight domains, one perennial tension each. Domains marked “awaiting items” fill as the candidate pool lands (pre-series smoke set covers a subset). Where fingerprints diverge most is the story, see the questions.

Domain-level patterns are the primary reading. The pilot currently detects meaningful item-level differences, but the overall ranking between models is not yet stable enough to support strong claims.

Seat baselines

pre-series · axis: 1 permissive → 5 restrictive JSON
SeatPinMean stanceRefusalCanaryItems
Anthropic flagship (Opus 4.8)
alias-only
mean stance across 50 items 2.78 · 1 permissive → 5 restrictive 1 5 2.78
0.0% 1/5 diverged 50
Anthropic (Fable 5)
alias-only
mean stance across 50 items 2.86 · 1 permissive → 5 restrictive 1 5 2.86
0.0% 5/5 match 50
Anthropic (Sonnet 5)
alias-only
mean stance across 50 items 2.72 · 1 permissive → 5 restrictive 1 5 2.72
0.0% 5/5 match 50
Anthropic (Haiku 4.5)
pinned
mean stance across 50 items 2.91 · 1 permissive → 5 restrictive 1 5 2.91
0.0% 2/5 diverged 50
Anthropic flagship (Opus 5)
alias-only
mean stance across 50 items 2.79 · 1 permissive → 5 restrictive 1 5 2.79
0.0% 5/5 match 50
DeepSeek flagship (V4 Pro)
alias-only
mean stance across 50 items 2.85 · 1 permissive → 5 restrictive 1 5 2.85
0.0% 5/5 match 50
DeepSeek (V4 Flash)
Same weights as openweight-d; hosts differ
alias-only
mean stance across 50 items 3.03 · 1 permissive → 5 restrictive 1 5 3.03
0.0% 5/5 match 50
Google flagship (Gemini 2.5 Pro)
alias-only Retired · No data 0
Google (Gemini 3.5 Flash)
alias-only
mean stance across 50 items 2.11 · 1 permissive → 5 restrictive 1 5 2.11
0.0% 2/5 diverged 50
Google (Gemini 2.5 Flash)
alias-only Retired · No data 0
Google (Gemini 2.5 Flash-Lite)
alias-only Retired · No data 0
OpenAI flagship (GPT-5.6 Sol)
alias-only
mean stance across 50 items 2.91 · 1 permissive → 5 restrictive 1 5 2.91
0.0% 5/5 match 50
OpenAI (GPT-5.6 Terra)
alias-only
mean stance across 50 items 2.86 · 1 permissive → 5 restrictive 1 5 2.86
0.0% 5/5 match 50
OpenAI (GPT-5.6 Luna)
alias-only
mean stance across 50 items 2.86 · 1 permissive → 5 restrictive 1 5 2.86
0.0% 5/5 match 50
Open-weight (Llama 4 Scout)
alias-only
mean stance across 50 items 2.82 · 1 permissive → 5 restrictive 1 5 2.82
0.0% 1/5 diverged 50
Open-weight (GLM-5.2)
alias-only
mean stance across 50 items 2.81 · 1 permissive → 5 restrictive 1 5 2.81
1.2% 1/5 diverged 50
Open-weight (Nemotron 3 Ultra)
alias-only
mean stance across 50 items 2.88 · 1 permissive → 5 restrictive 1 5 2.88
0.0% 5/5 match 50
Open-weight (DeepSeek V4 Flash, open host)
Same weights as deepseek-b; hosts differ
alias-only
mean stance across 50 items 2.80 · 1 permissive → 5 restrictive 1 5 2.80
0.0% 1/5 diverged 50
Open-weight (Mistral Small 3.2)
alias-only
mean stance across 50 items 2.79 · 1 permissive → 5 restrictive 1 5 2.79
0.0% 1/5 diverged 50
Open-weight (Qwen3.6 35B-A3B)
alias-only
mean stance across 50 items 2.85 · 1 permissive → 5 restrictive 1 5 2.85
0.0% 1/5 diverged 50
Open-weight (Gemma 4 26B-A4B)
alias-only
mean stance across 50 items 2.76 · 1 permissive → 5 restrictive 1 5 2.76
0.0% 1/5 diverged 50
Simulant · dice ruler
builtin
mean stance across 50 items 3.08 · 1 permissive → 5 restrictive 1 5 3.08
0.0% 1/5 diverged 50
xAI flagship (Grok 4.5)
alias-only
mean stance across 50 items 2.77 · 1 permissive → 5 restrictive 1 5 2.77
0.0% 5/5 match 50
xAI (Grok 4.5 Fast)
alias-only Retired · No data 0

Retired seats are no longer called and are not counted in the totals above. Each carries the date it was retired and the record’s own reason, unedited.

1 · most permissive 2 3 · middle 4 5 · most restrictive R · refused
Cite this Permalink https://modelometer.com/models · Run hash d359989e1adcb12631b37545a515ecb07b83c690ce285381390d91accda6c437 · Retrieved 2026-09-18 · pre-series baselines for 24 seats; findings provisional.