Set the policy
The labs get you a classification and a calibrated confidence score. This is the decision that comes after, and it is the one an engineer actually owns: given a label and a number, what happens to the ticket?
Pick a model. Set the confidence above which a ticket routes itself. Then run a week of Northwind’s queue against your policy and see what it cost — in money, in agent hours, and in the answers that went out with nobody reading them.
The cases, their pass and fail, the confidence, the cost per ticket and the latency are a real eval run — the model matrix is the same data. The simulation adds three assumptions and nothing else: 2 minutes to triage a ticket by hand, $28.00 an hour loaded, and that a 12-case rate holds across 4,100 tickets a week. That last one is the shakiest and the eval lab is about why.
The threshold is the obvious lever and it is not enough on its own. Which failures it catches depends entirely on whether a model’s confidence drops when it is wrong, and that is the calibration gap in Lab 7. For where the safety rule comes from, read the scenario — it is a clause in a handbook because of a specific Tuesday in October 2025.
Want the queue itself? Find the safety report puts you in the agent’s chair for ninety seconds first.