Act, review, or escalate: agents that know their limits
A good model isn't a working decision system. The team you have, what's at stake, and what happens on the day the model is down decide whether an agent can safely act on its own.
The 60-second version Open
- A policy that ignores capacity is a wish. Kestrel's cost-optimal lines sent 638 alerts a day to six analysts who can review 240.
- Fitted to the team it has, the policy acts below 0.031 and escalates at 0.34. Threats closed with no human fall from 37 a day to about 6.
- Stakes override confidence: nothing on a critical asset closes itself, however sure the model is.
- Four habits make it trustworthy: version the policy, log every decision, fail towards review, and audit the act zone at random.
- Watch three numbers every week. When reality pulls away from the model, a person chooses the fix.
Every agent eventually reaches the same fork. It has looked at something and it has a number. Now it can act on its own, hand the case to a person, or call for help.
Act, review, escalate. The first letter drew those lines from what each mistake costs. That's the easy half. This letter is about the other half: what happens when you switch it on, with a real team, real stakes, and a model that will one day be down.
The queue is part of the model
Switch on Kestrel's cost-optimal policy, and the live week sends about 638 alerts a day to review. At twelve minutes each, that's the full working day of about sixteen analysts.
Kestrel has six. Each can clear about 40 alerts in a shift, so the queue can take 240 a day. On-call can absorb about 40 pages before the pages themselves become the problem. Anything beyond 240 simply doesn't get looked at, in some order nobody chose.
A policy that ignores capacity isn't a policy. It's a wish.
So ask a better question: among all the policies that fit the team you actually have, which is cheapest? The search is quick, because for any escalation line, the best review line is the one that fills the queue exactly and no more.
policy = pol.best_policy_with_capacity(
history.pc, history.malicious,
max_review_rate=240 / per_day, # 6 analysts
max_escalate_rate=40 / per_day, # on-call limit
costs=costs)
print(f"act below {{policy.low:.3f}}, escalate at {{policy.high:.2f}}")Fitted on history and checked on the live week, the machine closes about 432 alerts a day by itself, the analysts get 243, and on-call gets 44 pages. The number of real threats closed with no human involved falls from 37 a day, under one line at 0.5, to about 6.
What a person costs, and what they buy
About six a day is still not zero. And here is where it pays to stop thinking like a modeller and start thinking like the person who signs the budget.
Refit the policy for every team size and the trade-off becomes visible. Going from six analysts to eight cuts the threats closed without review from 6.1 a day to 2.7. With Kestrel's estimates, two extra salaries buy back roughly $34,000 of expected loss a day. The model doesn't get to make that call, and neither does any one person alone. But now it's a decision with a visible price instead of a hunch.
There's a second lever in the same chart. The same model reading raw alert text, instead of clean structured fields, does worse at every team size. Better input buys you the same thing as more analysts.
Stakes override confidence
A model can be 99% sure an alert on the payroll database server is harmless. The 1% still lives on the payroll database server. So the policy looks at what's at risk, as well as how likely.
def route(p, severity, criticality):
if p >= policy.high or (p >= 0.1 and severity >= 2):
return "escalate"
if p < policy.low and criticality < 3:
return "act"
return "review" # everything else gets a humanTwo details are worth copying. The act door has an extra lock: nothing on a critical asset closes itself, however confident the model is. And the last line sends everything that doesn't clearly qualify for another door to a person.
When in doubt, review.
One essay a week on building AI that decides, with the code. Free.
Four habits for running it for real
A policy that works in a notebook is maybe a third of the job. The rest is the day the model is down, the week the attackers change tactics, and the month an auditor asks why an alert was closed.

Write the policy down as code, and version it. Thresholds that live in someone's head change without anyone noticing. "Policy v3: act below 0.031, escalate at 0.34, fitted on 1–21 September, six analysts" is something you can review, test and roll back.
Log every decision. A probability isn't a reason, and that makes a single answer hard to explain. But you can make the system explainable: record what the model saw, what it said, which policy version turned that into an action, and when. Then any decision can be replayed and questioned later.
Fail safe towards review, never towards act. If the model times out, returns an error or is simply unavailable, the alert goes to a person. A decision system should never quietly close alerts because something it depends on went down.
def decide(alert, client):
try:
p = calibrated(client.score(alert))
action = policy.decide(p)
except ModelError as err:
p, action = None, "review"
log(alert, p, policy="v3", action=action)
return actionAudit the act zone, at random. This is the habit teams most often skip, and the one that matters most. Every alert that goes to review or escalation comes back with an answer, because a person decided. Alerts the machine closed come back with nothing. If you only learn from the alerts people looked at, you will never see the machine's own mistakes. So send a small random slice of auto-closed alerts, say 3%, to an analyst anyway. It's the only honest window into the act zone.
The week it broke
In week five, two things went wrong at once. The model expected about 7% of alerts to be real threats; the real figure was 12%. And the review queue climbed to about 278 a day against a capacity of 240, so work that nobody chose to skip started getting skipped.
Neither problem was visible from any single alert. Both were obvious on a daily chart. That's why the weekly check watches three numbers: the share of alerts going through each door, the queue against capacity, and predicted threats against confirmed ones.
When they pull apart, there are three responses, and a person should choose: refit calibration on the most recent week, move the act line down for the affected alert type for a while, or add review capacity until the campaign passes. And remember that thresholds are fitted to the past, while attackers live in the future. Random audits, keeping scores inside the team, and keeping critical assets out of the act zone all make the lines harder to game.
"The purpose of computing is insight, not numbers." Richard Hamming, 1962
An agent that knows its limits isn't one that's always right. It's one whose lines are drawn by cost and capacity, that sends doubt to a person, and that leaves a record anyone can question later.
Next in the letter: The Jevons paradox of decisions.
Discussion