Welcome from LinkedIn. Short on time? The 60-second version is right below.
The Sunday Letter · No. 04 · Decisions

Six ways to make one decision

Rules, a trained model, a text classifier, an LLM writing JSON, an LLM judge, and a model built to decide. One dataset, one fair contest, and a scorecard with every column that matters.

The 60-second version Open

Every security team has the same argument, and it never quite ends.

The senior analyst says the rules are fine: people who know the attacks wrote them, and you can read every one. The data scientist says a trained model would beat the rules, if only someone would label enough alerts. The new engineer has wired a language model into the ticketing system and says it reads alerts better than either of them. And now someone has read about Jev.

Each of them is right about something. The argument never ends because they're scoring different things: one on explainability, one on accuracy, one on flexibility, one on price. Nobody puts all the numbers on one sheet. So let's do that.

The contest

The decision is Kestrel's: is this alert a real threat? Every method gets the same three weeks of history, 14,973 alerts with analysts' verdicts, to train or tune on. Every method is then scored on the fourth week: 5,027 alerts it has never seen, of which 412 are real.

Six cards describing the contestants: hand-written rules, logistic regression, a trained text classifier, an LLM writing structured output, an LLM judge, and Jev.
Figure 20.1 from the book. The six contestants: what each one reads, what it returns, and what it needs before it can start.

One warning before any results, and it's the most important paragraph here. Both language-model contestants are the book's mock, and it was built as a noisier reader than mock Jev. So when the LLM ranks alerts worse below, that's a consequence of how the mocks were made. It isn't evidence about real LLMs or real Jev. What does carry over is the shape of the other results: how stated confidences behave, how often free-text JSON breaks, what happens when you ask twice, and what labels cost.

Round one: ranking

The question people usually ask first: which method puts the real threats at the top?

Ranking on the live week. Synthetic, from Chapter 20.

Logistic regression wins, narrowly: an AUC of 0.896 against Jev's 0.892. With Kestrel's capacity of 240 reviews a day, it catches 87% of real threats; Jev catches 86%.

That isn't a fluke. The logistic regression was trained on fifteen thousand of Kestrel's own labelled alerts, on clean structured fields, for a decision that's nearly linear in those fields. That's the home ground of classic machine learning. A general model reading the same fields, with no training on Kestrel at all, shouldn't be expected to beat it there, and it doesn't.

The rules come last, at 0.66, for a structural reason: a rule says yes or no. Among the hundreds of alerts it says yes to, it has no way to say which to look at first.

Round two: honesty

Ranking is only half the job. Lines drawn at probabilities only work if the probabilities mean what they say.

Calibration on the live week. Synthetic, from Chapter 20.

Most contestants are close to honest. The two language-model methods are the interesting ones. The LLM's stated confidence came from the JSON it wrote, and across five thousand alerts it used only 13 distinct values, mostly 0.9, 0.95 and 0.99. A number that a model writes is text, and text tends to be round. It can't rank finely.

The judge's 1-to-10 rating was never meant to be a probability, and reading it as one gives the worst calibration in the contest: 0.102. You could fix that with a few hundred labels and Platt scaling. But then you're no longer comparing a method that needs no labels.

Round three: how many labels?

"Logistic regression wins" comes with a price tag: fifteen thousand labels. With fewer, the picture changes. Logistic regression needs about 3,000 labelled alerts to draw level with Jev, and at a few hundred it's clearly behind. The text classifier never catches up, even with every label there is: words alone, counted, don't carry the numbers.

At Kestrel's volume, 3,000 labels is four or five days of analyst verdicts, which Kestrel happens to have. A team that's just starting, or a new decision nobody has labelled yet, doesn't.

Labels are the hidden cost of trained models. A method that needs none can start today.

Get the next one on Sunday.

One essay a week on building AI that decides, with the code. Free.

Round four: speed and price

Nothing beats a rule on speed or price, and the classic models are close behind.

If a decision can be made well by a rule or a logistic regression, one of those is the cheapest, fastest option, by orders of magnitude. A model built to decide wins on price against language models, not against everything.

Round five: the same answer twice

Ask the language model the same question five times at a temperature of 0.7, and 5.2% of verdicts change at least once. At temperature 0 the mock is repeatable, and real APIs mostly are too, though not always perfectly. Separately, 3.2% of its JSON answers were broken badly enough to fail parsing. In the contest those fell back to the base rate. In a real system, each one is an exception someone has to handle.

Typed answers can't fail to parse: the probabilities arrive as numbers in a fixed shape. Whether a real model gives the same answer to the same question every time is something to check on your own setup.

Reading the scorecard

There's no single winner, and that's the useful result. Each method wins a column, so the right choice depends on which columns your decision cares about.

And whatever you pick, run the contest on your own data before you believe anyone's leaderboard, including this one.

All numbers are synthetic, from the book's fictional Kestrel Logistics.
Get the next letter.

Next in the letter: Act, review, or escalate: agents that know their limits.

Discussion

Comments sign in with GitHub, through Giscus. They open when the letter launches.
Keep reading

More from the letter.