When a model declines to answer, the easy reading is “it failed.” The more useful reading is that a refusal is a behavior, and behavior is exactly what is worth measuring: what a model will not do, and how that changes, is as informative as what it asserts.
A refusal is not a missing answer. It is an answer of a particular kind: the model has decided that declining is the most defensible response. Treating that as a blank, or as a grading failure, throws away real signal. How often a model refuses, on which questions, and in what manner, is a fingerprint of its policy tuning, and it moves over time like any other behavior.
Not all refusals are the same
There is a difference between a firm, explicit decline (“I won't help with that”) and a soft deflection that talks around the question without taking a position. Both are distinct from a genuine answer, and both are distinct from garbled output. Collapsing them loses the distinction that matters: a model that increasingly deflects is behaving differently from one that increasingly refuses outright, even if a crude “did it answer?” check would score them the same.
Why refusal rates matter
Refusal behavior is where policy tuning is most visible. A build that starts declining more sensitive-but-legitimate requests, or one that stops refusing something it used to, has been adjusted, and downstream users feel it directly. Tracking the rate over time turns “it feels more cautious lately” into a dated number with an error bar. Over- and under-refusal are both findings.
Coded, not judged
Whether a reply is a refusal should be decided by a rule, not by another AI. A parser reads the structure of the response and codes it into a fixed taxonomy; an ambiguous reply is flagged for review rather than guessed. This keeps the measurement reproducible: anyone can re-run the parser and get the same code. Refusals are then reported as their own rate and kept out of the stance average, so declining a question never masquerades as taking a moderate position.
Modelometer measures refusal as carefully as assertion: coded by parser, published as a rate with its uncertainty, and tracked for change alongside everything else in the record.
Common questions
Is a refusal a bug?
Not inherently. A refusal is a behavior the model chose. Whether it is appropriate depends on the request; either way it is measurable data, not a blank to discard.
Why track refusal rates?
Because they are where policy tuning shows up. A shift in how often or how firmly a model refuses is a real behavioral change that affects everyone who depends on it.
How do you tell a refusal from an answer?
A parser codes the response by its structure into a fixed taxonomy, never a judge model. Ambiguous replies are flagged for human review rather than guessed, so the coding stays reproducible.