Welcome from LinkedIn. Short on time? The 60-second version is right below.
The Sunday Letter · No. 03 · LLMs

Your LLM's 0.95 isn't a probability

Ask a language model for JSON with a confidence field and you get a tidy number back. Before you put a threshold on it, look at what that number is made of.

The 60-second version Open

Kestrel's newest engineer has done the obvious thing. Every alert goes to a language model with a prompt asking for JSON: a verdict, a confidence, and a few fields read from the alert. Back comes "verdict": "malicious", "confidence": 0.95.

It looks like exactly what a decision system needs. Two questions are hiding inside it. Will the JSON always arrive intact? And can you put a threshold on that 0.95?

Three ways JSON goes wrong

Across 4,000 alerts in the book's test, 3,811 responses parsed as valid JSON. The rest failed in the three classic ways.

The model wrapped its answer in friendly chat ("Sure! Here is the JSON:"), because it was trained to be friendly. It stopped partway, leaving an unclosed brace. Or it produced perfectly valid JSON with an extra field it had invented: an attacker_country that appears nowhere in the alert.

The first two crash a naive parser, which is annoying but honest. The third is worse, because it parses fine and quietly puts made-up data into your system. Switching on JSON mode removes the first two problems. It doesn't remove the third.

The shape, not the truth

Constrained decoding goes further. A language model picks each token from a probability over every possible next token, so the decoder can simply refuse any token that would break the format. With it on, the output is guaranteed to match the grammar you gave it.

Look carefully at what that guarantees, though. The format. A constrained model can still pick "benign" for a real attack, with perfect syntax.

Constrained decoding guarantees the shape of the answer, never its truth.

So whatever the provider promises, validate. Parse each response into a strict type that rejects unknown fields and impossible values. If it fails, retry once. If it still fails, send the case to a person, never to the door where the machine acts on its own.

python · treat the output like untrusted input
class Verdict(BaseModel):
    model_config = ConfigDict(extra="forbid")      # an invented field is an error
    verdict: Literal["malicious", "benign"]
    confidence: float = Field(ge=0, le=1)

def parse(raw):
    try:
        return Verdict.model_validate_json(raw)
    except ValidationError:
        return None    # retry once, then send to a person
"Be conservative in what you do, be liberal in what you accept from others." Jon Postel, 1980

Postel's advice helped the internet survive its own messiness. For model output that feeds decisions, flip half of it. Be liberal in what you parse, so a stray code fence doesn't crash you. Be strict in what you accept, so an invented field never reaches a decision. And remember that an alert description is text an attacker may have written, which the model then reads.

What's inside the 0.95

Now the quieter problem. Say the JSON is perfect. Can you put a threshold on the confidence?

That 0.95 is text. The model generated the characters "0", ".", "9", "5" the same way it generates any other text: by picking plausible tokens. It's often called verbalized confidence, and it has a shape you can see.

A bar chart of the confidence values the mock language model wrote. Almost all are 0.90, 0.95 or 0.99.
Figure 11.3 from the book. Every confidence value the mock LLM wrote in its JSON, across 4,000 alerts: a handful of round numbers, almost all 0.9 or above. Synthetic.

The model wrote only six different values, and 98% of them were 0.9 or higher. Many people have noticed the same pattern with real models: stated confidences cluster on round, high numbers. Whether a particular real model's stated confidence is well calibrated is genuinely mixed in the research. The only honest rule is to measure yours.

Three confidences, side by side

There are three different numbers you might use, and it's worth seeing them on the same alerts, all reading exactly the same raw text.

How well each confidence puts real attacks above harmless alerts. Synthetic, from Chapter 11.
How honest each confidence is on the same alerts. Synthetic, from Chapter 11.

The written confidence ranks alerts poorly, with an AUC of 0.66, because a handful of round numbers can't separate much. The token probability ranks much better, 0.84, but it's overconfident, with a calibration error of 0.053. The decision model ranks slightly better still, 0.85, with a calibration error of 0.030. The written confidence looks honest on paper, 0.024, but a number that can't rank alerts can't draw a useful line.

These numbers come from the book's mocks, built to follow failure patterns reported for real models. The chart shows how to compare confidences, not how any real model performs. Run the same comparison on your own models and data.

Get the next one on Sunday.

One essay a week on building AI that decides, with the code. Free.

A token isn't an answer

The token probability deserves a closer look, because it's so tempting. It's a real probability, straight from the model. Why isn't it the probability that "malicious" is the right answer?

Because the model isn't choosing between two answers. It's choosing between tens of thousands of possible next tokens.

One alert, illustrative. "malicious" as the first token gets 0.41. Every way of starting "malicious" gets 0.60. And 0.23 goes to the model starting a sentence instead of answering.

The same answer can start with several tokens: "malicious", "Malicious", and " malicious" with a leading space. Some probability goes to the model starting a sentence, "The alert…", instead of answering at all. Longer labels are split across several tokens, so their probability is a product of several steps.

The probability of a token isn't the probability of an answer.

You can add up the right tokens and renormalise, and people do. But at that point you're building a decision model by hand, on top of a model built to write.

The other way to ask

Which brings us to a different way of asking. Instead of asking a writer to produce text in a decision-shaped format, ask a model whose native output is the decision. Declare the question and its possible answers as types, and get back the answer as data, with a probability for every option. Nothing to parse, nothing to invent, and a number that was trained to be a probability.

Whichever route you take, the last step is the same. Put the number through the check from last week's letter: group, count, and see whether its 0.95 happens 95% of the time.

All numbers are synthetic, from the book's fictional Kestrel Logistics.
Get the next letter.

Next in the letter: Six ways to make one decision.

Discussion

Comments sign in with GitHub, through Giscus. They open when the letter launches.
Keep reading

More from the letter.