Skip to main content

Lab 7 answers

Q1a. Do the two cost discoveries change the tier decision?

No, and saying why is the whole exercise.

They do not change the decision, because the decision was never about cost. Moving triage from Opus to Sonnet saves about $63 a month at warm rates ($111 versus $48). Fixing Haiku's cache would have taken it from ~$77 to ~$20. Against a $4,000 budget, all of those round to free. An argument that was already dominated by the accuracy column stays dominated by the accuracy column.

What they change is your confidence in every other number that table produced. The matrix contained two cost errors in one column and nothing flagged either: a fivefold cache miss on Haiku, and cold-write overhead that made Sonnet look barely cheaper than Opus. Not a test, not a review, not the person who wrote the "Haiku costs about half" sentence. The correct update is not about any one model. It is that this repo could print a wrong cost figure and ship it, which means the next wrong cost figure will also ship, and the next one may land on a decision where $63 is not the stake.

Notice also the shape of the errors. Neither made a tier look absurd. The Haiku miss made the cheapest model look reasonable, and the Sonnet cold writes made the middle tier look pointless. Errors that produce implausible output get caught. These produced plausible stories, which is why they survived.

Q1b. Where would this have been caught first?

A startup check, and the gap is not effort — it is timing and knowledge.

What it would have to know that a test does not: its own configuration at the moment it boots. The service knows its model and can measure its prefix in one free countTokens call. That pairing is the entire bug, and it is knowable before the first request is served. A test only ever asserts what someone already thought to check, and the whole failure here is that nobody thought to check the cheap tier — smoke would have caught this on day one if anyone had run it with TRIAGE_MODEL set, and for months nobody did.

Ranked by how early each instrument fires:

instrumentwhen it fireswhy it missed
startup checkboot, every deploynow exists — src/lib/preflight.ts, added because of this bug
CI matrixon the PR that changes the tiersmoke runs against the default model only
code reviewon the diffnobody memorizes per-model cache minimums
dashboardafter you are already payingcache_hit_rate flat at zero is unmissable if someone looks
smoke testonly when someone runs it with the right env varnobody did

The honest ranking puts the test fifth. That is worth sitting with, because the test is the thing this lab just spent a page improving. Making it fail loudly was right — it now tells the truth to whoever runs it — but a correct assertion nobody triggers is not a control. It is documentation with a CI badge.

The transferable rule: a check that requires someone to have already suspected the problem is not a control. Controls fire on their own schedule, not on your suspicion. Startup and CI qualify; a manual smoke run does not.

What was actually built. src/lib/preflight.ts runs on every boot, measures the frozen prefix against the configured model's minimum with one free countTokens call, and prints a loud block when caching is off. Three details are worth copying into your own version:

  1. It measures system[0] only — the frozen block that holds the breakpoint. Counting the whole request would include the volatile block and the user message, which sit after the breakpoint and do not count toward the minimum. That version passes while the real prefix falls short: a check that measures the wrong span is worse than no check, because it also removes the suspicion that something needs checking.
  2. It never blocks or throws. It fires after the listener is up and is not awaited. A diagnostic that can take production down is a worse bug than the one it diagnoses. On any failure — offline, bad key, API error — it says it could not check and returns, and willCache: false is always paired with an error so "unknown" is never read as "confirmed broken".
  3. It names the wrong remedy explicitly. The obvious reading of "prefix too short" is "make the prefix longer", which here means padding a legal document with ~1,300 tokens to win a discount. The message says not to.

Note what it still does not do: it warns, it does not exit non-zero. That is deliberate for a dev server and arguably wrong for a deploy pipeline. Deciding where that line sits for your own service is the actual exercise.

Q1. What does the headroom do to the tier argument?

It removes it. At 4,100 tickets a week the flagship costs about $111 a month warm (the twelve-case matrix projects $132, cold writes included) against a $4,000 budget. Sonnet saves roughly $63 of that — under 2% of the budget, in exchange for losing one to three more cases in twelve. The argument does not improve when you give it its best case, which is the sign that it was never a cost argument.

The general form is worth keeping: cost optimization only matters where cost is a binding constraint. Northwind's binding constraint is the mis-routing rate and the safety SLA, not the model bill. A team that optimizes the non-binding constraint has done work that cannot show up in any outcome they care about, and has spent accuracy to do it.

Where the argument would bite: raise volume to 400,000 tickets a week and Opus becomes ~$10,800/month, over budget, while Sonnet is ~$4,600, and the tradeoff is live again. Latency will not rescue it: in the 2026-09-15 run Sonnet and Opus are within a second of each other at p50 (2.7s versus 3.3s).

So the honest recommendation for this system is: ship the flagship, and put the effort you would have spent on tiering into the eval set instead.

Q2. Is it a 5% problem?

Much larger, and the accuracy number actively hides it.

The gold set is deliberately adversarial: twelve cases chosen because they sit where rules touch. Real traffic is not distributed that way — most tickets are "where is my package." So the cheap tier's aggregate accuracy on live traffic would be far better than 8/12 or 9/12 suggests, and a naive rollout would look fine.

But the cases it loses are not randomly drawn from the distribution. They are concentrated in exactly the population you built the system for. eval-04 is a child swallowing part of a product. Losing that case is not 8% of an accuracy score; it is the failure mode from the October 2025 incident on the scenario page, reproduced by choice.

To answer properly you would need the joint distribution: how often does live traffic hit a two-rule case, and what does being wrong on one cost? Northwind can estimate the first from a sample of the archive and the second from the handbook's SLA penalties. That is the measurement, and it is a business measurement rather than a model one.

The trap to name out loud: aggregate accuracy on a representative sample and aggregate accuracy on an adversarial sample answer different questions, and neither answers "what does this cost me."

Q3. What does a threshold catch when the gap is zero?

Roughly its base rate, which is to say nothing useful.

If confidence on wrong answers is drawn from the same distribution as confidence on right ones, then thresholding at 0.7 selects ~the same fraction of each. You escalate some wrong answers, you escalate just as many right ones, and you pay for a second call on all of them. The precision of the mechanism is the model's error rate — no better than escalating at random.

In the two earlier runs where Haiku's gap came out negative, it is worse than random: the threshold preferentially escalated cases the model had gotten right, while passing the wrong ones straight through. You are paying a premium to double-check the answers that did not need it.

The cost is not theoretical. Escalation on a low-signal model converges toward "run every ticket twice," at which point you are paying the cheap tier plus the flagship for a result no better than the flagship alone — strictly worse than just calling the flagship.

The rule that transfers: measure the calibration gap before you build anything that routes on confidence. It is one column in an eval you are already running, and it is the difference between a control and a costly no-op.

Q4. What message defeats pickModel?

Any high-stakes ticket written without high-stakes vocabulary. The routing signal is a keyword list, and the list can only match words the customer happens to use.

The canonical example is already in this course: "probably nothing, but the bottle cap cracked and my kid swallowed a bit of plastic." That one routes correctly, because "swallowed" and "kid" are both on the list. Now write it the way a worried, apologetic parent actually writes at 11pm: "Hi — the lid on the 32oz came apart and some of it ended up in my daughter's mouth. She seems okay. Just thought you should know." No "swallow," no "injury," no "child." So it routes to Sonnet, the cheap tier. Sonnet passed eval-04 in the 2026-09-15 run, but the router has just sent the case where being wrong is most expensive to the model whose confidence you trust least.

It does not need an attacker. The failure mode is politeness.

Why the failure is quiet: nothing errors. The router logs a confident reason ("no high-stakes language"), the cheap model returns a well-formed schema-valid classification with 0.9 confidence, and the ticket goes in the normal queue. Every component reports success. The only artifact is a meta.routed.reason in a log nobody reads, and you find out when the safety SLA is missed — which is precisely the three-day delay from the scenario.

Two consequences worth stating:

  1. A keyword router is a defence-in-depth layer, never the only layer. requires_human is decided by the model reading the whole message, not by pickModel, and that ordering is deliberate.
  2. Bias the router toward escalation. The HIGH_STAKES list is over-broad on purpose. A false positive costs a fraction of a cent; a false negative costs the incident.

Lab 8 makes the adversarial version of this explicit, where the message is written to route itself down deliberately.

Q5. What is escalation worth?

On the cheap tier: close to nothing, for the reasons in Q3. Its confidence does not predict its errors, so the trigger fires on the wrong population.

On Sonnet: possibly something, and you do not know yet. Across four earlier runs its gap was 0.20–0.30 — not flagship, but informative enough that a 0.7 threshold selects disproportionately for wrong answers, which would make sonnet → escalate to opus a defensible architecture. The 2026-09-15 run put it at 0.05, with a wrong answer at 0.90 confidence. Twelve cases cannot tell you which is the real number. Before shipping escalation from Sonnet, run the matrix five times and look at the distribution of the gap, not a single value.

The general shape: escalation is only as good as the calibration of the model you escalate from. It is not a safety net you can bolt onto any tier; it is a mechanism that consumes a signal, and you have to check the signal exists.

Which reframes Steps 4 and 5 as complementary rather than alternative. pickModel routes on the input and works regardless of the model's self-knowledge, so it is what protects the safety case. Escalation routes on the output and needs a calibrated model, so it is what catches genuine ambiguity. Neither substitutes for the other, and a system that ships only the second one has a hole exactly where it can least afford one.

Q6. Two models score 11/12. Is that the same number?

No, and there are three separate reasons, in increasing order of how much they should worry you.

Noise. This set moves by up to two cases run-to-run with nothing changed. Two observations of 11/12 are consistent with true rates anywhere from roughly 75% to 100%. A twelve-case set cannot resolve a difference smaller than about eight points, so "11/12 versus 11/12" is not evidence of equality — it is an absence of evidence about anything.

Composition. Even if both are truly 11/12, they may be missing different cases. One misses the deliberately ambiguous ticket; one misses the safety report. Same score, and you would ship them into different jobs.

What the score omits. Accuracy is one column. Two models tied on it can differ by a factor of ten in calibration gap — and Step 3 established that the gap decides whether you can build a control on top of the model at all. The tie is in the least informative column.

To distinguish them properly: run each five times and compare distributions rather than points ($0.45 and the cheapest real improvement available), read the disagreement matrix, and compare calibration gaps. If they still look identical, you have learned something useful — pick on latency or price and stop deliberating.

Extension notes​

Predict before you measure. Classification and prose generation load different capabilities, and there is no rule that says a model that classifies worse also writes worse. The cheap tiers may well write acceptable customer replies while being materially worse at deciding what the reply should say — in which case the correct architecture is not one tier for everything but a cheap drafter behind an expensive classifier, which is a conclusion you can only reach by measuring the two tasks separately.

Keep the judge pinned when you do it. The temptation to let each tier grade its own prose is strong and it destroys the comparison.

Note also what you will be measuring: the judge itself swings 1/4 to 3/4 on identical input. With a four-case sample you cannot distinguish tiers at all. Raising the sample is most of the work, and deciding how far to raise it is Lab 6 Q7 all over again.

Q7. What is the latency column worth here?

For Northwind's triage queue: almost nothing, and saying so is the point. 4,100 tickets a week arrive into a queue nobody reads in real time. Whether a classification takes 3 seconds or 30 makes no difference to any human — the tickets are processed faster than they arrive at every tier, and the constraint that binds is accuracy on the cases where two handbook rules interact. Shipping Sonnet to save 0.6 seconds nobody experiences, at the price of one to three cases in twelve, is the same mistake as shipping it to save $63 against a $4,000 budget. Same error, different column.

That is the transferable habit and it is worth stating flatly: optimize the constraint that binds. A number being real, measured, and printed in your table does not make it a decision input. Most of the columns in most model comparisons are like this, and the discipline is knowing which one is yours before you look.

Now change the product. Put the classifier in front of a person and latency stops being decoration:

  • The storefront support form. A customer watches the pipeline run while they wait. Here the p95 column matters (5.5s on Opus, 4.4s on Sonnet), and the classification is not what they came for — they came to file a ticket. The gap between tiers is small enough that the answer may be neither: stream a provisional result, or classify after submission. Measure p95 on the real form before deciding.
  • An agent-assist panel that classifies while a human reads the ticket. The budget is however long the person spends reading, which is a few seconds. p95 is the number that matters, not p50, because the failure is "the panel was still empty when I finished reading" and that happens in the tail.
  • A phone IVR routing a live caller. Nothing above is shippable; you need a different architecture, not a different tier.

Two things to carry out of that. Which percentile you read is part of the decision — p50 tells you what the experience usually is, p95 tells you how often it is bad, and an interactive surface is judged on the tail. And a system with both workloads should not pick one row: Northwind's real answer is the flagship on the batch queue and the cheap tier on the live form, which is pickModel from Step 4 doing exactly the job it exists for.

The measurement caveats matter before you quote any of this. These are four-in-flight, whole-request times through the local route, not time-to-first-token and not single-request figures. /v1/triage does not stream; a streaming surface would be judged on time-to-first-token instead, and /v1/draft would need its own measurement that this matrix does not make.

Was this page helpful?