Your LLM's 0.95 isn't a probability
Ask a language model for JSON with a confidence field and you get a tidy number back. Before you put a threshold on it, look at what that number is made of.
The 60-second version Open
- Asking for JSON fails in three ways: chatty wrappers, cut-off objects, and valid JSON with invented fields. The last one is the dangerous one.
- Constrained decoding guarantees the shape of the answer, never its truth. Validate every response like untrusted input.
- A confidence the model writes is text. In the book's test it used six values, and 98% were 0.9 or higher.
- A token's probability isn't an answer's probability: the same answer can start with many different tokens.
- Of three kinds of confidence on the same alerts, the written one barely ranked, the token probability was overconfident, and the decision model did both jobs best. Measure yours.
Kestrel's newest engineer has done the obvious thing. Every alert goes to a language model with a prompt asking for JSON: a verdict, a confidence, and a few fields read from the alert. Back comes "verdict": "malicious", "confidence": 0.95.
It looks like exactly what a decision system needs. Two questions are hiding inside it. Will the JSON always arrive intact? And can you put a threshold on that 0.95?
Three ways JSON goes wrong
Across 4,000 alerts in the book's test, 3,811 responses parsed as valid JSON. The rest failed in the three classic ways.
The model wrapped its answer in friendly chat ("Sure! Here is the JSON:"), because it was trained to be friendly. It stopped partway, leaving an unclosed brace. Or it produced perfectly valid JSON with an extra field it had invented: an attacker_country that appears nowhere in the alert.
The first two crash a naive parser, which is annoying but honest. The third is worse, because it parses fine and quietly puts made-up data into your system. Switching on JSON mode removes the first two problems. It doesn't remove the third.
The shape, not the truth
Constrained decoding goes further. A language model picks each token from a probability over every possible next token, so the decoder can simply refuse any token that would break the format. With it on, the output is guaranteed to match the grammar you gave it.
Look carefully at what that guarantees, though. The format. A constrained model can still pick "benign" for a real attack, with perfect syntax.
Constrained decoding guarantees the shape of the answer, never its truth.
So whatever the provider promises, validate. Parse each response into a strict type that rejects unknown fields and impossible values. If it fails, retry once. If it still fails, send the case to a person, never to the door where the machine acts on its own.
class Verdict(BaseModel):
model_config = ConfigDict(extra="forbid") # an invented field is an error
verdict: Literal["malicious", "benign"]
confidence: float = Field(ge=0, le=1)
def parse(raw):
try:
return Verdict.model_validate_json(raw)
except ValidationError:
return None # retry once, then send to a person"Be conservative in what you do, be liberal in what you accept from others." Jon Postel, 1980
Postel's advice helped the internet survive its own messiness. For model output that feeds decisions, flip half of it. Be liberal in what you parse, so a stray code fence doesn't crash you. Be strict in what you accept, so an invented field never reaches a decision. And remember that an alert description is text an attacker may have written, which the model then reads.
What's inside the 0.95
Now the quieter problem. Say the JSON is perfect. Can you put a threshold on the confidence?
That 0.95 is text. The model generated the characters "0", ".", "9", "5" the same way it generates any other text: by picking plausible tokens. It's often called verbalized confidence, and it has a shape you can see.

The model wrote only six different values, and 98% of them were 0.9 or higher. Many people have noticed the same pattern with real models: stated confidences cluster on round, high numbers. Whether a particular real model's stated confidence is well calibrated is genuinely mixed in the research. The only honest rule is to measure yours.
Three confidences, side by side
There are three different numbers you might use, and it's worth seeing them on the same alerts, all reading exactly the same raw text.
- The confidence the model wrote. The 0.95 in the JSON.
- The token probability. Many APIs return the probability the model gave each token it produced, so you can read off the probability of "malicious" as the first token of the verdict.
- A probability from a model built to produce one. A classifier, or a decision model whose output is a probability for each option.
The written confidence ranks alerts poorly, with an AUC of 0.66, because a handful of round numbers can't separate much. The token probability ranks much better, 0.84, but it's overconfident, with a calibration error of 0.053. The decision model ranks slightly better still, 0.85, with a calibration error of 0.030. The written confidence looks honest on paper, 0.024, but a number that can't rank alerts can't draw a useful line.
These numbers come from the book's mocks, built to follow failure patterns reported for real models. The chart shows how to compare confidences, not how any real model performs. Run the same comparison on your own models and data.
One essay a week on building AI that decides, with the code. Free.
A token isn't an answer
The token probability deserves a closer look, because it's so tempting. It's a real probability, straight from the model. Why isn't it the probability that "malicious" is the right answer?
Because the model isn't choosing between two answers. It's choosing between tens of thousands of possible next tokens.
The same answer can start with several tokens: "malicious", "Malicious", and " malicious" with a leading space. Some probability goes to the model starting a sentence, "The alert…", instead of answering at all. Longer labels are split across several tokens, so their probability is a product of several steps.
The probability of a token isn't the probability of an answer.
You can add up the right tokens and renormalise, and people do. But at that point you're building a decision model by hand, on top of a model built to write.
The other way to ask
Which brings us to a different way of asking. Instead of asking a writer to produce text in a decision-shaped format, ask a model whose native output is the decision. Declare the question and its possible answers as types, and get back the answer as data, with a probability for every option. Nothing to parse, nothing to invent, and a number that was trained to be a probability.
Whichever route you take, the last step is the same. Put the number through the check from last week's letter: group, count, and see whether its 0.95 happens 95% of the time.
Next in the letter: Six ways to make one decision.
Discussion