Six ways to make one decision
Rules, a trained model, a text classifier, an LLM writing JSON, an LLM judge, and a model built to decide. One dataset, one fair contest, and a scorecard with every column that matters.
The 60-second version Open
- Teams argue about rules versus models versus LLMs because each side scores a different column. Put every column on one sheet.
- On Kestrel's alerts, logistic regression ranked best (AUC 0.896), narrowly ahead of mock Jev (0.892).
- Its price is labels: it needed about 3,000 labelled alerts to draw level with a method that needs none.
- An LLM's written confidence used only 13 values, and 3.2% of its JSON didn't parse.
- If a rule or a logistic regression can make a decision well, it's the cheapest, fastest option by orders of magnitude.
Every security team has the same argument, and it never quite ends.
The senior analyst says the rules are fine: people who know the attacks wrote them, and you can read every one. The data scientist says a trained model would beat the rules, if only someone would label enough alerts. The new engineer has wired a language model into the ticketing system and says it reads alerts better than either of them. And now someone has read about Jev.
Each of them is right about something. The argument never ends because they're scoring different things: one on explainability, one on accuracy, one on flexibility, one on price. Nobody puts all the numbers on one sheet. So let's do that.
The contest
The decision is Kestrel's: is this alert a real threat? Every method gets the same three weeks of history, 14,973 alerts with analysts' verdicts, to train or tune on. Every method is then scored on the fourth week: 5,027 alerts it has never seen, of which 412 are real.

One warning before any results, and it's the most important paragraph here. Both language-model contestants are the book's mock, and it was built as a noisier reader than mock Jev. So when the LLM ranks alerts worse below, that's a consequence of how the mocks were made. It isn't evidence about real LLMs or real Jev. What does carry over is the shape of the other results: how stated confidences behave, how often free-text JSON breaks, what happens when you ask twice, and what labels cost.
Round one: ranking
The question people usually ask first: which method puts the real threats at the top?
Logistic regression wins, narrowly: an AUC of 0.896 against Jev's 0.892. With Kestrel's capacity of 240 reviews a day, it catches 87% of real threats; Jev catches 86%.
That isn't a fluke. The logistic regression was trained on fifteen thousand of Kestrel's own labelled alerts, on clean structured fields, for a decision that's nearly linear in those fields. That's the home ground of classic machine learning. A general model reading the same fields, with no training on Kestrel at all, shouldn't be expected to beat it there, and it doesn't.
The rules come last, at 0.66, for a structural reason: a rule says yes or no. Among the hundreds of alerts it says yes to, it has no way to say which to look at first.
Round two: honesty
Ranking is only half the job. Lines drawn at probabilities only work if the probabilities mean what they say.
Most contestants are close to honest. The two language-model methods are the interesting ones. The LLM's stated confidence came from the JSON it wrote, and across five thousand alerts it used only 13 distinct values, mostly 0.9, 0.95 and 0.99. A number that a model writes is text, and text tends to be round. It can't rank finely.
The judge's 1-to-10 rating was never meant to be a probability, and reading it as one gives the worst calibration in the contest: 0.102. You could fix that with a few hundred labels and Platt scaling. But then you're no longer comparing a method that needs no labels.
Round three: how many labels?
"Logistic regression wins" comes with a price tag: fifteen thousand labels. With fewer, the picture changes. Logistic regression needs about 3,000 labelled alerts to draw level with Jev, and at a few hundred it's clearly behind. The text classifier never catches up, even with every label there is: words alone, counted, don't carry the numbers.
At Kestrel's volume, 3,000 labels is four or five days of analyst verdicts, which Kestrel happens to have. A team that's just starting, or a new decision nobody has labelled yet, doesn't.
Labels are the hidden cost of trained models. A method that needs none can start today.
One essay a week on building AI that decides, with the code. Free.
Round four: speed and price
Nothing beats a rule on speed or price, and the classic models are close behind.
- Rules, logistic regression, text classifier: microseconds per decision, under $1 per million, on hardware you already own.
- Jev: about $21 per million decisions, in tens to hundreds of milliseconds.
- An LLM writing a small JSON answer: about $740 per million, and over a second each.
If a decision can be made well by a rule or a logistic regression, one of those is the cheapest, fastest option, by orders of magnitude. A model built to decide wins on price against language models, not against everything.
Round five: the same answer twice
Ask the language model the same question five times at a temperature of 0.7, and 5.2% of verdicts change at least once. At temperature 0 the mock is repeatable, and real APIs mostly are too, though not always perfectly. Separately, 3.2% of its JSON answers were broken badly enough to fail parsing. In the contest those fell back to the base rate. In a real system, each one is an exception someone has to handle.
Typed answers can't fail to parse: the probabilities arrive as numbers in a fixed shape. Whether a real model gives the same answer to the same question every time is something to check on your own setup.
Reading the scorecard
There's no single winner, and that's the useful result. Each method wins a column, so the right choice depends on which columns your decision cares about.
- You have thousands of labels and clean fields: a logistic regression is hard to beat, and it's the cheapest thing on the sheet.
- You're starting fresh, or the decision is new: a method that needs no labels lets you start today, and you can collect labels as you go.
- Your input is raw text: let a language model extract the fields, then make the decision on the fields. In the contest, Jev reading clean fields beat Jev reading the raw text, 0.892 against 0.856.
- You need to explain every outcome: rules still earn their place, often as a backstop beside a model.
And whatever you pick, run the contest on your own data before you believe anyone's leaderboard, including this one.
Next in the letter: Act, review, or escalate: agents that know their limits.
Discussion