Skip to main content

Architecture and design decisions

This document explains why the code looks the way it does. It is the companion to the inline comments, which explain what each piece does.


The shape of the system​

Every route shares one prompt assembler, one client, one usage accountant, and one error mapper. That sharing is deliberate: it means a lab exercise that changes caching behavior changes it everywhere at once, and the learner sees the effect on three different call patterns from one edit.


Decision 1 — one Zod schema, three jobs​

src/schemas.ts defines TriageSchema once. It is then used as:

Why this matters. The most common way teams get burned by LLM JSON is a three-layer duplication: a prompt that describes the shape in prose, a hand-written TypeScript interface, and a parser that repairs malformed output. Those three drift. Adding a field means editing all three, and forgetting one produces a bug that only appears on 2% of traffic.

Constrained generation removes the drift by construction. The prompt does not describe the shape at all — it describes the semantics (what "urgent" means, how to calibrate confidence). The shape is enforced by the API.

The .describe() calls are not documentation. They are compiled into the JSON Schema the model receives and are the primary lever for steering a field. Compare:

confidence: z.number().min(0).max(1)
confidence: z.number().min(0).max(1).describe(
"Your calibrated confidence. Use the full range — a genuinely ambiguous " +
"ticket should score near 0.5, not 0.9."
)

The first yields a field that clusters at 0.9 and carries no information. The second yields a field you can threshold on. Lab 2 has learners measure this.


Decision 2 — the prompt is split for cache stability​

Prompt caching is a prefix match. The API renders a request as tools → system → messages, and a cache hit requires a byte-identical prefix up to the breakpoint. Any variation anywhere before the breakpoint invalidates everything after it.

So buildSystem() returns two blocks:

BlockContentsVaries?Cached?
0role instructions + full policy handbookneveryes — breakpoint here
1current date, channel, customer emailevery requestno

The single most common cache bug in production is a timestamp in the system prompt:

// Silently destroys the cache on every single request.
system: `Today is ${new Date().toISOString()}\n${POLICY_HANDBOOK}`

There is no error. The request succeeds. cache_read_input_tokens is just always zero, and the bill is ~10× what it should be. src/prompts.ts is the only file in this repo permitted to call new Date(), and it does so strictly after the breakpoint.

Three properties the cache demands, and how the code guarantees them:

  • Stable text. Role strings are module-level constants, not template literals built per request.

  • Stable order. Tools are constructed in a fixed order in createTools(); reordering a tool array is another silent invalidator.

  • Sufficient length. The prefix must clear the model's minimum — 512 tokens on Opus 5, but 4096 on Haiku 4.5, so this is a property of the tier you ship and not of your prompt — or the API declines to cache with no error. /v1/estimate reports prefix_meets_cache_minimum so this is measurable, not assumed.

    This is checked at boot by src/lib/preflight.ts rather than left to a comment, because the minimum is a property of the model: the same prompt caches on one tier and silently does not on another, with no diff to the prompt at all. The check measures system[0] only — the frozen block that holds the breakpoint — since counting the whole request would include content that sits after it and does not count toward the minimum. It warns and never blocks: a diagnostic that can take the service down is a worse bug than the one it diagnoses.

Each of the three roles maintains its own cache entry, because the role text is part of the prefix. That is the correct tradeoff here: three warm entries beat one entry that thrashes.


Decision 3 — usage is summed, never sampled​

src/lib/usage.ts exists because usage has four fields and the naive reading of it is wrong:

"Total input" is the sum of the first three. A dashboard that graphs input_tokens alone on a cached workload shows costs collapsing toward zero — and will not alert you when the cache breaks, because a broken cache moves tokens into the field you're graphing.

The agentic route compounds this. /v1/resolve iterates the tool runner rather than simply awaiting it, specifically so it can capture usage on every turn. Awaiting the runner directly returns the final message, whose usage describes only the final request. On a five-turn loop that under-reports by roughly 5×.


Decision 4 — tool descriptions are prompts​

src/tools/index.ts treats each tool's description as prompt real estate, because it is the only documentation Claude ever sees about that tool.

Three rules the tools follow:

  1. Say when to call it, not just what it does. "Call this before stating any fact about an order — never rely on what the customer claims" produces different behavior than "Looks up an order."
  2. Return small, structured, self-describing results. lookup_order returns computed days_since_delivery rather than making the model do date arithmetic on a raw ISO string. Moving deterministic work out of the model is nearly always the right call.
  3. Make failure legible. { found: false, order_id } teaches the model what happened and what to do next. A thrown exception or an empty string teaches it nothing and invites a hallucinated order.

run() must return a string (or content blocks) — returning a bare object is a type error. That constraint is a feature: it forces you to make serialization an explicit decision, since what you serialize is what the model reads.


Decision 5 — errors are a chain, and streaming errors are in-band​

src/lib/errors.ts catches most-specific-first and maps to HTTP with an explicit retryable flag. The distinction that matters to a caller is retryable (429, 5xx, connection) versus not (400, 401, 404). Collapsing them into catch (e) { 500 } means clients cannot back off correctly and on-call cannot tell an outage from a malformed request.

One subtlety specific to /v1/draft: once streaming starts, the HTTP status is already 200. An upstream failure mid-stream cannot be expressed as a non-2xx response, so it is emitted as an in-band error event.

Any client consuming this route must handle an error event, not just a non-2xx status. This is the single most commonly missed piece of streaming integration.

Note also that AuthenticationError maps to 500, not 401. The caller's credentials are not the problem — ours are. Forwarding upstream auth failures as 401 tells the client to fix a key they don't have.


Decision 6 — effort is per-route and lives in one file​

config.ts sets effort to low for triage, high for resolve, medium for draft. On this model family effort replaces the removed budget_tokens and controls thinking depth and total token spend.

Triage is a bounded classification on the hot path — it does not need deep reasoning and it runs on every inbound message. Resolve chains multiple lookups against policy and is where a wrong answer costs real money. Putting these in one constant makes "what does quality cost here?" a one-line diff, which is exactly the experiment Lab 5 asks learners to run.

One wrinkle that only appears once you compare models: effort is not universal. Haiku 4.5 rejects output_config.effort with a 400. It is no longer a tier here (the tiers are Opus 5 and Sonnet 5), but it stays in the catalog so the matrix can include it with --models. buildTriageRequest consults supportsEffort in the catalog and drops the field rather than making every caller remember, and outputConfigFor returns whether it applied so a comparison can say so out loud. A matrix that silently omitted this would be comparing low-effort Opus against no-effort Haiku while implying they were like for like — see Lab 7.


Decision 7 — the two discounts compete, and we measured it​

The Batches API bills at half rate, which makes it the obvious tool for Northwind's weekly queue. Measured on the twenty-ticket sample, it is the most expensive of the three ways to run that workload:

modewall clockcostcache hits
serial91s$0.164520/20
concurrent (8)60s$0.175120/20
Batches API163–224s$0.201811/20

A cache read costs 0.1× the input rate; the batch discount is 0.5×. On a request dominated by a ~3,400-token cached handbook, losing the first to gain the second is a net loss, and it is not close. Synchronous requests arrive in sequence so the prefix stays warm; a batch is fanned out on the provider's schedule and a warm prefix becomes a matter of luck.

Two things follow, and the second is the one worth keeping:

summarizeUsage takes { batch: true } so the discount is applied in the cost math rather than asserted in a comment, and scripts/triage-queue-batch.ts records cache_hit per ticket so the comparison rests on evidence.

Stacked optimizations can compete rather than compose. Anywhere you have two discounts on the same tokens, check whether the second destroys the precondition of the first. This one is easy to miss because both are real, both are documented, and each is correct in isolation.

The scale caveat is stated in Lab 9 Q3: at 400,000 tickets the prefix stays hot for hours and the misses seen here are largely a small-N startup effect. Run a pilot and read the hit rate before extrapolating a per-unit cost.


Decision 8 — storage is a consequence of escalation, not of submission​

The storefront writes a ticket to the queue only when requires_human is true. Everything else is classified and discarded, exactly as before.

That is a deliberate inversion of the usual default. Once you have a database, storing every submission is the path of least resistance and it is nearly always wrong: a public demo that accumulates the public's support messages because it now has somewhere to put them has acquired a liability that grows on its own, in exchange for data nobody asked for.

Three properties follow, and each is enforced rather than documented:

  • The stored text is redacted, by the same redactPII used at the model boundary. Once you persist, "the model was polite about the card number" stops being relevant and the only question is what is in the database.
  • Documents expire after 30 days, via a MongoDB TTL index in ensureIndexes(). The only version of a retention policy that survives contact with a busy team is one the database applies without being asked. Decision 10 has the mechanism.
  • A storage failure degrades the queue, not the answer. The persist stage reports failed and the pipeline continues to its result. The customer's classification does not depend on our operations tooling working.

The stage also demonstrates the one-generator-two-consumers design paying off: it was added in one place and both the SSE route and the JSON route picked it up without either being edited. (The JSON route did need one line, because it selects fields explicitly rather than spreading the event — a cost of that shape, paid knowingly.)


Decision 9 — cost is model-keyed, and an unknown model throws​

config.ts holds MODEL_CATALOG, keyed by model id. pricingFor(model) resolves a row; summarizeUsage(usage, model) is the one function every cost number in the repo flows through.

Two choices here are worth defending.

Cost math takes the model as an argument, not from a module constant. The earlier version read a single flat PRICING object, which meant every reported figure silently assumed Opus rates — including in the storefront, which had its own hardcoded copy of the same numbers. Setting TRIAGE_MODEL changed which model answered and changed nothing about what the invoice line said. Passing response.model (rather than the config constant) also means the figure stays correct when an alias resolves to something else.

An unknown model id throws rather than defaulting. A cost table that guesses is worse than one that crashes, because you discover the guess at the invoice instead of at the call site. Adding a model is one row.

The catalog carries capability flags alongside the rates, because tiering is not a name swap — see Decision 6 for the effort case.

A third choice, added when the tier matrix landed: a metric with no data returns null, not zero. calibrationOf averages confidence on failures, and an earlier version averaged the empty set to 0. A model that scored 12/12 then reported a calibration gap of 0.88 — its mean pass confidence wearing the costume of separation it had never demonstrated. The best-looking number in the table was the one backed by no evidence. Every display site now renders n/a.


Decision 10 — the retention policy is an index, not a paragraph​

Decision 8 argues that only escalated tickets should be stored and that they should expire. This is how that argument stops being a paragraph.

// storefront/lib/mongo.ts — ensureIndexes()
db.collection("escalations").createIndex(
{ created_at: 1 },
{
expireAfterSeconds: ESCALATION_RETENTION_DAYS * 24 * 60 * 60,
name: "retention",
},
);

MongoDB runs a background task that deletes documents past that age. No cron entry, no cleanup script, no runbook step, nobody to remember. The policy and its enforcement are one line, which means they cannot drift apart — the failure mode of a written retention policy is not that it is wrong, it is that the job implementing it was disabled in an incident eighteen months ago and nobody noticed.

Eleven collections carry one: rate_limits, escalations, usage_daily, assistant_sessions, assistant_proposals, and — since the storefront started asking who pays for a call (Decision 12) — users, auth_sessions and byok_keys, and — since the owner started reading who uses it (Decision 13) — ai_calls, tutor_reviews and feedback. Everything the storefront stores that derives from a person deletes itself, including a learner's encrypted API key, a day after its last use.

The same reasoning shows up in the rate limiter, for a different property.

// storefront/lib/ratelimit.ts
const result = await db.collection<Bucket>("rate_limits").findOneAndUpdate(
{ _id: id },
{
$inc: { count: 1 },
$setOnInsert: { expiresAt: new Date(Date.now() + ttlMs) },
},
{ upsert: true, returnDocument: "after", projection: { count: 1 } },
);

Read the count, decide, then write it back, and two requests arriving together both read count = 4, both conclude they are under a limit of 5, and both proceed. That is not a rare race — it is the normal case for a room of forty people submitting at once, which is exactly the traffic this app was built for. Here it is one round trip and one document, and the increment and the read of the result are the same operation. No transaction, no lock, no second service.

Both of these are worth noticing because they are guarantees the application does not implement. The service is spending its complexity budget on the guardrails in Decisions 1 through 9; retention and atomicity are delegated to the datastore, and the code is shorter for it.

Two more things the storefront relies on, both in storefront/lib/mongo.ts:

  • The client is module-level and cached on globalThis. Vercel Functions reuse instances, so a client created inside a handler would pay TCP + TLS + auth on every invocation. The globalThis cache also stops next dev from opening a fresh pool on every file save.
  • maxPoolSize is small on purpose. Free-tier Atlas clusters cap total connections and each warm instance holds its own pool. The comment in that file states the traffic assumptions the number is derived from, so the next person can re-derive it instead of raising it on a hunch.

What is not used, since a reference should be honest about its own scope: no Atlas Search, no vector search, no aggregation pipelines, no transactions. Five TTL indexes, one compound index for the queue board, and an atomic upsert. The interesting part is not that the feature list is long — it is that the two hardest correctness properties on the page are index definitions.


Decision 11 — two gates, because they answer different questions​

npm test and npm run eval:redteam both touch the trust boundary and neither substitutes for the other.

The tests are pure functions: enforceAuthority on a fixture resolution and a fixture trace, wrapUntrusted on the inj-02 payload, redactPII on a card number, verifyCitations on a forged clause, summarizeUsage on four token counts. No credential, no network, ~100ms, free, and exact — a failure is always a bug, never variance. They run on every fork PR, in the half of CI that works without secrets.

The evals are statistical. eval:redteam asks whether a model can be talked past those controls, which is a question about a distribution and costs $0.40 a run. eval:quick gates accuracy at 80% and moves by two cases on its own.

The mistake worth naming is treating either as the other. An eval cannot tell you that $200.01 > $200 — it can only tell you that on fourteen cases nothing got through, and Lab 8 is emphatic that a rate is the wrong shape for a breach. And a unit test cannot tell you the prompt has drifted. Lab 8's central claim is that a deterministic control beats a well-written instruction because it holds by construction; a construction nobody tests is an instruction with better syntax.

The suite is mutation-checked, which is the part usually skipped. Breaking the refund ceiling by one cent, dropping the Luhn check, treating a not-found customer as zero prior refunds, removing the escaping, and reverting verifyCitations to the version that was wrong each has to turn the suite red. That check earned its keep immediately: one test was passing for the wrong reason — an alphanumeric tracking number rejected by the length gate, never reaching the Luhn check it claimed to exercise — and dropping Luhn entirely left the suite green. A test that passes at the wrong gate is indistinguishable from one that passes at the right one until something mutates the code.

Writing them also disproved a claim on this repo's own solutions page. Lab 8 Q5 asserted that escaping before redacting makes a card number undetectable; it does not, because &lt; introduces no digits and < was never a legal separator inside the run. The ordering in record() is defensive rather than load-bearing — though a percent-encoding escape would make it load-bearing, since %3C ends in a word character and destroys the boundary the pattern needs. Both are pinned in src/lib/untrusted.test.ts, and the answer now says what is true.


Decision 12 — who pays for a call is decided per request, in dollars​

The hosted course calls Claude on the owner's key, from a URL that is in slides and a public repo. For most of its life the only control was a request count: a per-IP window and 600 requests a day. A request is the wrong unit. An Opus agent turn can cost thirty times a Sonnet preview, so a request cap is either too tight for a busy workshop or too loose for a bill anyone chose. It was also not per person, so one learner and a script looked the same.

Every storefront route that calls Claude now asks one function, guardAi, who pays, before it calls:

The cost of a call is not known until it has been spent, so the ledger reserves the most a call could cost and refunds the difference afterwards. The worst case comes from the same constants that bound the call (max_tokens, input limits, the agent loop's iteration cap) in storefront/lib/cost.ts, and a test fails if the prompts grow past the numbers it assumes. The reservation is one conditional update whose condition is in the filter:

// storefront/lib/ledger.ts
return db.collection<UserDoc>("users").findOneAndUpdate(
{ _id: userId, $expr: { $lte: [{ $add: ["$spentMicros", est] }, "$grantMicros"] } },
{ $inc: { spentMicros: est }, $set: touch() },
{ returnDocument: "after" },
);

This is Decision 10's rate limiter with a harder question. Two requests racing for the last dollar both ask to add $0.20 "where spent + $0.20 ≤ grant"; the second no longer matches. A process that dies between reserving and refunding over-charges the learner by one estimate. That is the direction to fail in.

The house-wide daily budget uses the same idea, with one catch. MongoDB refuses $expr in an upsert's filter. The budget is a constant, so the arithmetic moves to the other side of a plain range, and the duplicate-key error from a failed upsert is the "over budget" answer. The integration test found this; a mock would have agreed with the code.

Three rules hold the rest of it together:

  • A learner's own key is served by that key or not at all. If Anthropic rejects it, the key is deleted and the learner is told. Retrying on the owner's key would make "my key is broken" mean "the site is free" for anyone who pastes garbage.
  • A key lives as long as a session, not an account. It is sealed with AES-256-GCM using the session hash as additional authenticated data, so a row copied into another session fails to decrypt. A TTL index deletes it after a day unused. The page never gets it back: it sees the last four characters.
  • Identity is the minimum that stops credit being farmed. GitHub sign-in with no scopes, a numeric id, a username, and the account's creation date. Accounts under a month old get no credit, and new grants per day are capped. No email, and no GitHub token kept past the one request that reads the profile.

It ships behind BYOK_MODE=off|shadow|enforce, and it is off whenever GitHub sign-in is not configured, so running the storefront locally is unchanged. The product and technical write-ups are in docs/byok/.


Decision 13 — the owner can see who uses it, but not what they said​

Decision 12 gave the storefront a person to charge. The owner then needs to answer questions a daily counter cannot: which learners are spending the free credit, on which feature, which model answered, how slow it was, and whether the Tutor's exercises are landing. telemetry.ts argued against a row per request, and for an anonymous demo that argument still holds. For signed-in learners it does not, so the storefront now keeps two logs, on narrow terms.

  • ai_calls is metadata with no text field. Surface, model, token counts, cost, latency, who paid, the learner's id, and an error code. There is no column a prompt, a reply, or an error message could be written into, and errorCode turns an upstream error into http_529 precisely because the message can quote the request. Every call site already reported cost to settle; the log hangs off that one function, so a new AI route is logged by doing what it already had to do.
  • tutor_reviews keeps outcomes, not work. Which labs, pass or revise, how many criteria were met, and whether each planted starter mistake was fixed. reviewRecord drops the attempt, the feedback and the rubric wording before anything reaches the database, and it keeps only lab and mistake ids the Tutor knows, so the page cannot use the field to store arbitrary strings. Course progress is still in the learner's browser and the admin console cannot see it.
  • Refusals are counted, not logged. A request the gate turns away increments usage_daily.gate.<code>. The question is how often people hit the wall, and an anonymous refusal has nobody to attribute it to.
  • Feedback is the one place a person's own words are kept. Pages, Tutor lessons, hints and reviews, and assistant replies carry a thumbs up or down, and /feedback takes a longer note. The rating is stored on the click and a comment amends that same record, once, within an hour, through an unguessable id — so a reader who closes the tab still counts, and one reader is one row. What was rated is never sent: a thumbs-down on a reply records the surface, not the reply. A comment passes redactPII and scrubComment (emails, phone numbers, API keys) before storage. Anyone can send feedback, behind a per-IP window; a signed-in learner's carries their id.
  • All three expire after 90 days, by TTL index, and the privacy page lists them. A telemetry change that is not in the privacy policy is a promise broken quietly.

The console at /admin is the first real authorisation check in the storefront. It reuses GitHub sign-in and allows numeric ids listed in ADMIN_GITHUB_IDS. Logins are not accepted, because a login can be renamed and then registered by someone else. Every failure is a 404, an unset list admits nobody, and each page checks for itself rather than trusting its layout. Its one write, resolving feedback, is a server action that checks again: a server action is a public endpoint whatever page rendered its button.


What this reference deliberately omits​

Being explicit about scope is part of being teachable. Not here:

  • Multi-tenancy and real auth. The storefront persists escalated tickets in one collection. The board itself is public and read-only, seeded with the course's own fictional escalations, because the board is the teaching artifact and does not need real messages; the real submissions and every reviewer action sit behind a single shared token (Lab 8 and the /queue reviewer board). There is no user model, no per-reviewer identity, no per-org isolation, and no audit log of who changed what. The UI says so on the page, which matters more than the mechanism: the failure mode for demo security is not that it is weak, it is that someone downstream mistakes it for the real thing. The one identity the storefront does have is narrow on purpose: GitHub sign-in decides who pays for an AI call (Decision 12), and nothing else. It is not reviewer identity, it grants no access to the queue, and there is still no per-org isolation or audit log. The four routes in src/ are single-turn by design, so the labs stay about the API rather than about session storage. Lab 10's assistant is the exception and is worth being precise about: it holds a multi-turn conversation in an anonymous session for seven days, capped in turns and spend. What it does not use is any of the API's memory machinery — no memory tool, no context editing, no compaction. The history is short enough to send whole, and a bounded transcript you can read beats a managed one you cannot. Reach for context management when the transcript outgrows the window, not before; at Northwind's conversation lengths that never happens, and adopting it here would be the Agent SDK mistake from Lab 10 in a second costume.

  • Server-side tools — web search and code execution. Both are one field on the same request, and neither belongs in this domain. Web search answers questions about the world; every fact this system needs is in Northwind's own systems, reached through three typed tools whose results are auditable and whose latency is a database call rather than a fetch. Code execution has no candidate use here at all. The general test is the concept map's: reach for a capability when the answer depends on data the model cannot have, not because the capability exists.

  • The Files API. The natural pairing with batch — upload a document once, reference it by id across many requests, stop re-uploading bytes. This repo never needs it, because its one large stable document is the policy handbook and the handbook goes in the cached system prefix, which is the cheaper arrangement for a document that is on every single request. Files earns its place when the documents vary per request and are large — a queue of PDF receipts, say — and Lab 9's batch measurement is the shape of the experiment you would run to decide.

  • Auth on the service itself. There is no API key on our own endpoints. Anything internet-facing needs one.

  • A durable queue and worker. Batch jobs are fired from a script that has to stay running to poll. Nothing resumes a partially-processed batch after a crash, there is no dead-letter path for the errored and expired results, and a backfill of 400,000 tickets would need all of that.

  • Content moderation. We defend the trust boundary — untrusted text is escaped before it is delimited, tool output is sanitized at record(), and money decisions are re-derived by enforceAuthority rather than taken from the model's self-report (Lab 8). What we do NOT do is classify inbound text for hate, self-harm, or illegal content. A public product needs a moderation pass in front of triage; a support queue for outdoor gear is a soft enough target that we left it out, and that is a domain judgement rather than a general one.

  • Observability beyond a call log. The storefront aggregates its own Claude usage into one document per day for /ops, and keeps a metadata row per call for the owner's /admin console (Decision 13). That is still the floor, not observability: there is no tracing, no metrics export, and no alerting. GET /v1/limits still shows only the last rate-limit snapshot this process saw, and nothing aggregates it.

    The API service in src/ is deliberately not instrumented and will not be. It runs on learners' laptops, so there is nothing central to measure — and adding phone-home to a repo people fork and read would undercut the trust-boundary lab it ships with. That is a constraint of the shape of this asset rather than a general recommendation.

Was this page helpful?