Why 0.5 is almost always the wrong threshold
Every tutorial turns a probability into an action with one line at 0.5. On a real security queue, that line quietly lets thirty-seven threats a day through. Here is where the lines should go.
The 60-second version Open
- A probability isn't an action. Someone still has to decide what to do with 0.37 at 2:14 a.m.
- One line at 0.5 assumes two things that aren't true: that a missed threat and a false alarm cost the same, and that there are only two things you can do.
- There are three doors: act, review, escalate. Machines take the easy ends; people take the middle.
- Let the costs draw the lines. With Kestrel's numbers: act below 0.0016, escalate at 0.43.
- Check calibration before trusting any line, then fit the lines to the team you actually have.
It's 2:14 a.m. at Kestrel Logistics. An alert comes in. The model looks at it and says 0.37.
You know it's a probability. You may even know how to check whether it's honest. And now someone in the security team is standing behind you, asking the only question that matters at that hour: so what do we do with it? Close it? Put it in the queue for the morning? Wake someone up?
That step, from a number to an action, is the least glamorous part of any AI system. It's also where most of the money is won or lost.
The obvious answer
The obvious answer is the one every tutorial gives. If the probability is above 0.5, treat it as a threat. Below 0.5, it's fine.
Let's try it on a real week. Take three weeks of Kestrel's alerts as history and treat the fourth week as live, the way you would if you were switching this on for real.
one_line = pol.ThreeZonePolicy(low=0.5, high=0.5)
r = pol.evaluate(one_line, live.p, live.malicious)
print(f"auto-closed per day: {r['act'] / 7:.0f}")
print(f" of which real threats: {r['missed_by_automation'] / 7:.1f}")About 689 alerts a day close themselves, which sounds wonderful until you read the next line. Every day, roughly 37 of them were real threats: genuine phishing, genuine malware, a stranger with someone's password. The machine closed them on its own and nobody ever looked.
0.5 quietly assumes two things that aren't true here.
It assumes a missed threat and a false alarm cost the same, which is absurd for a security team. And it assumes there are only two things you can do with an alert. There are three.
Three doors
Walk around a real security operations centre and you'll see that every alert leaves through one of three doors.
The first door is the one nobody sees. The alert is closed, or handled by an automated playbook, and no human ever reads it. Most alerts should leave this way, because most alerts are noise. The second door leads to the queue: an analyst picks the alert up, pulls the logs, maybe messages the user, and makes the call. It costs about twelve minutes of a skilled person's time. The third door is the loud one. Someone's phone rings, at 2:14 a.m. if necessary, because this one can't wait.
Act, review, escalate. Machines take the easy ends; people take the middle.
The names are general on purpose. In customer support, act might be an automatic reply, review a human agent, escalate a supervisor. In content moderation, it might be auto-approve, human review, and legal.
Let the costs draw the lines
So where do the lines go? Not where they look nice. Where the money says. For every action, write down what it costs on average, given the probability P that the alert is real:
Act costs P × $10,000, the expected loss when a real threat slips through. Review costs $15 of an analyst's time, plus the small chance they miss it: P × 5% × $10,000. Escalate costs (1 − P) × $400, waking someone for nothing.
Draw all three as lines against P, and at every value pick the cheapest. The lines cross in exactly two places, and those crossings are your thresholds.
costs = pol.Costs(auto_close_miss=10_000, review_minutes=12,
analyst_per_hour=75, review_miss_rate=0.05,
escalate_false_alarm=400)
low, high = pol.cost_optimal_thresholds(costs)
print(f"act below {low:.4f}; escalate at {high:.2f} and above")That first number deserves a second look. It says: only let the machine close an alert on its own if it's more than 99.8% sure the alert is harmless. When a miss costs 650 times more than a review, that's what "cheap mistakes only" means.
One essay a week on building AI that decides, with the code. Free.
Check the number under the line
There's a catch. The cost lines are drawn as if P means exactly what it says: as if, among all the alerts the model scores at 0.001, one in a thousand really is a threat. If the model is overconfident, your "cheapest action" is quietly not the cheapest at all, and the machine is closing threats you never agreed to close.
So before moving a single line, check. On Kestrel's live week, the alerts the raw model called "under 2% likely" averaged a claimed 0.6%. The real rate was 1.4%. Its "almost certainly harmless" alerts were about two and a half times riskier than it said. After a simple fix fitted on the history weeks, the claim and reality land much closer: 0.76% against 1.00%.
Never put a threshold on a probability you haven't checked against recent outcomes.
Then reality walks in: the queue
Switch the cost-optimal policy on, and the live week sends about 638 alerts a day to review. At twelve minutes each, that's the full working day of about sixteen analysts. Kestrel has six. They can clear about 240 alerts a day, and on-call can absorb about 40 pages before the pages themselves become the problem.
A policy that ignores capacity isn't a policy. It's a wish. So ask a better question: among all the policies that fit the team you actually have, which is cheapest?
Fitted on history and checked on the live week, the answer is to act below 0.031 and escalate at 0.34. The machine closes about 432 alerts a day by itself, the analysts get 243, on-call gets 44 pages, and the number of real threats closed without a human falls from 37 a day to about 6.
Six is still not zero. Going from six analysts to eight cuts it to 2.7 a day, and with Kestrel's estimates, two extra salaries buy back roughly $34,000 of expected loss a day. The model doesn't get to make that call, and neither do you on your own. But now it's a decision with a visible price instead of a hunch.
That's the whole idea. A model gives you a number. A decision is that number, an honest check that the number means what it says, and lines drawn by what mistakes cost and who is there to catch them.
Next in the letter: When 0.8 should really mean 80%.
Discussion
Design preview: a sample thread showing how replies and the author's answers will look.
We use a single 0.5 cut-off on our fraud queue. How would you pick the escalate line if you don't know the cost of a false page?
Start from capacity instead: how many pages a day can on-call absorb? Put the line where that budget runs out, then revisit once you've priced a page.