The first LLM feature in a codebase usually calls a provider SDK straight from a route handler. The second feature copies the first. By the fifth, you have five different timeout values, three retry loops written from memory, and a model name hard-coded in places nobody remembers. Then a provider has a bad afternoon, and every feature fails in its own way.
None of that is a model problem. It is the same problem any team has with a flaky downstream dependency: nobody owns the edge between your code and theirs. With LLMs the edge is worse than usual, because a “successful” response can still be the wrong shape, a retry can repeat a side effect, and every attempt is billed.
This post is about the thin layer I put at that edge: a model gateway. It is design analysis plus illustrative code. Nothing here is a measured result, and the routing choices in it are examples, not benchmarks.
To keep it concrete, I’ll carry one example through: a support-ticket assistant. When a ticket arrives it (1) classifies the ticket, (2) extracts structured fields like order ID and product, (3) drafts a reply for a human to approve, and (4) runs an agent step that may call a tool to create a follow-up task. Four calls, four different sets of needs.
Why product code shouldn’t call provider SDKs directly
My rule: product code calls callModel(workflow, input) and gets back a typed result or a typed failure. It never names a model, a provider or a retry count.
That gives you two things.
One contract in, one contract out. The ticket classifier passes a ticket and expects { category, confidence }. It does not care whether that came from a small fast model or a large one, or which vendor. Provider SDKs differ in request shape, error classes, streaming events and how they express structured output. If each feature absorbs those differences itself, switching providers becomes a codebase-wide migration. Behind a gateway, it’s a config change plus a round of evaluation.
Per-workflow policy. “Classify a ticket” and “draft a reply” want different models, different timeouts and different retry behaviour. That policy should live in one place, keyed by workflow name, where someone can read it and review a change to it. Scattered across call sites, it drifts, and nobody can tell you what it is.
The gateway should stay thin. It isn’t an agent framework and it doesn’t own prompts or business logic. It owns the network edge: pick a model, call it within a budget, handle failure, validate the shape, and write down what happened.
Routing by workflow, not by vibes
Routing tends to happen by vibes: someone tried a model in a playground, liked it, and that model now serves everything. The better question is what this specific workflow is sensitive to.
In product work, when I’ve evaluated providers and routed workflows to them, the useful move was always to name the one or two properties a workflow can’t compromise on, then pick the cheapest option that meets them on your own evaluation set. The table below shows what that looks like for the ticket assistant. The choices are illustrative. Your tiers, models and thresholds will be different, and they should come from your own evals.
| Workflow type | Example step | Latency | Quality | Cost sensitivity | Structured-output reliability | Illustrative routing policy |
|---|---|---|---|---|---|---|
| Short classification | Classify ticket | High: runs on every ticket, often inline | Moderate: a small label set | High: highest volume | High: must be one of N labels | Small, fast tier. Tight timeout. Fallback to another small model. If both fail, a rules-based default label. |
| Structured extraction | Pull order ID, product, intent | Moderate | High on exact fields | Moderate | Very high: downstream code parses it | Mid tier with native schema/JSON mode. Validate every response. Fallback only to models tested against the same schema. |
| Long-form generation | Draft reply | Time-to-first-token matters if streamed to a person | High: a human reads it | Low: lower volume, and the value is visible | Low: free text | Larger tier, streamed. Generous total timeout, tight first-token timeout. Fallback to another large model. Degrade to a template if everything fails. |
| Tool-using agent step | Decide whether to create a follow-up task | Moderate, but it adds up across steps | High: wrong tool calls have side effects | Moderate | Very high: tool arguments must validate | Model that has been evaluated on tool use with your real tool schemas. Per-step budget. No blind retries after a tool has run. |
Writing it down makes “which model is best?” stop being a useful question. Structured-output reliability gets its own column because it’s the property most likely to break silently when you swap models, which matters once we get to fallbacks.
Timeouts
Time-to-first-token versus total time
A streamed generation has two different ways to be slow. It can take a long time to start, or it can start quickly and then run long. Those need separate limits.
- Time-to-first-token (TTFT) timeout. If nothing has arrived after N seconds, the request is probably queued behind load or stuck. Give up and try elsewhere. This limit should be tight.
- Total timeout. A reply that is streaming steadily is healthy even if it takes a while. This limit should be generous, and its main job is to cap cost and stop runaway generations.
Non-streaming calls only get the total. That’s fine for a classifier returning a few tokens, and a poor fit for long outputs. Anthropic’s docs recommend streaming for long-running requests, and note that idle connections can be dropped by networks during long non-streaming calls.
Streaming has one more catch: an error can arrive after the HTTP status was already 200. Anthropic documents error events mid-stream, for example an overloaded_error that would be a 529 on a non-streaming call. So “the request succeeded” isn’t settled at the status line. Your gateway has to treat a stream that errors or stalls halfway as a failure, and decide whether the partial output is usable (usually not, for anything structured).
Budgets per step in multi-step agents
In an agent, per-call timeouts aren’t enough. If the ticket assistant runs four steps and each has a 30-second timeout plus two retries, the worst case is several minutes, and nobody chose that number.
So I give the whole run a deadline and each step a slice of it. The gateway receives the remaining time and uses whichever is smaller: the step’s own timeout, or what’s left of the run. When the run’s budget is gone, later steps don’t start. The agent returns a partial result or hands off to a person. That’s a decision you make on purpose, instead of finding out from a user that the page spun for four minutes.
Pass one AbortSignal down from the run. If the user leaves or the run is cancelled, every in-flight model call stops, and so does the billing for it.
Retries
Retries are where gateways do the most damage, because a retry loop is easy to write and hard to write safely.
Which errors are retryable
The provider docs are fairly clear on this.
| Signal | Example | Retry? | Notes |
|---|---|---|---|
| Rate limited | HTTP 429 | Yes, with backoff | Honour retry-after if present. Anthropic sends retry-after on rate-limit 429s. OpenAI’s error codes guide also says to follow Retry-After and back off exponentially. |
| Spend cap reached | HTTP 429 with a spend-limit code | No | Anthropic returns 429 when the monthly spend cap is hit, with no retry-after and error_code: enforced_spend_limit_reached. Retrying fails until access resumes. OpenAI also has a spend-limit 429. A status code alone isn’t enough to decide. |
| Server error | HTTP 500 | Yes | Anthropic’s docs say to retry with exponential backoff. |
| Overloaded | Anthropic 529 overloaded_error, OpenAI 503 |
Yes, and consider falling back | The Anthropic and OpenAI docs both describe this as temporary capacity pressure. |
| Provider timeout | Anthropic 504 timeout_error |
Yes, or switch to streaming | See the long requests guidance. |
| Your timeout / network drop | AbortError, connection reset | Yes, if the step is safe to repeat | See idempotency below. |
| Validation error | HTTP 400 invalid_request_error |
No | The same request will fail the same way. Fix the request. |
| Auth / permission | 401, 403 | No | Alert. Retrying just adds noise. |
| Too large | 413 | No | Trim input or route to a different model. |
| Content policy refusal | Refusal or policy error | No | Retrying the same prompt is either pointless or looks like trying to get around the policy. Treat it as a product outcome. |
Also check your SDK. The Anthropic SDKs, for example, already retry transient failures twice by default with backoff. If your gateway also retries three times, one logical call can become nine attempts. Pick one layer to own retries. I set the SDK’s maxRetries to 0 and let the gateway own it, so the retry count in the logs is the true one.
Backoff with jitter, and a retry budget
Retry with exponential backoff and full jitter: wait a random amount between zero and base * 2^attempt, capped. Without jitter, every client that failed at the same moment retries at the same moment, and you rebuild the spike that caused the failure. If the provider sends retry-after, use that as the minimum wait.
Per-call limits aren’t enough. When a provider is really down, every request retries and you multiply load on a struggling system. A retry budget caps retries across the process, say as a fraction of recent first attempts. When it runs out, calls fail fast or go straight to fallback.
Idempotency, and why tool calls change everything
Retrying a classification is harmless. You pay twice and get a label. Retrying an agent step that has already called a tool is a different matter.
Take step 4 of the ticket assistant. The model decides to call create_follow_up_task. Your code runs the tool, the task is created, and then the next model call (the one that reads the tool result) times out. A naive gateway retries the whole step from the start. The model decides again to create a task, and now there are two.
The rule I use: the gateway retries model calls, never tool executions. And the agent loop records tool effects before it asks the model anything else.
- Give every tool call a stable idempotency key, derived from the run ID, step number and tool name, and have the tool implementation reject duplicates. This belongs in the tool, in software, not in a prompt asking the model not to repeat itself.
- Persist the tool result before the next model call. On retry, replay the stored result into the conversation instead of re-running the step.
- Mark tools as
readOnlyorsideEffectingin their definitions. The gateway, or the agent runner above it, can then refuse to replay a step that contains a side-effecting call without a completed-effect record.
In reconciliation work I learned that a duplicate posting is worse than a failed one: a failure is visible, a duplicate looks like data. Side-effecting agent tools deserve the same caution.
Fallbacks
A fallback is a second route for when the first one is unavailable: another model from the same provider, or the same class of model from a different provider. Used well, a provider outage becomes a small quality dip rather than a broken feature.
The catch: fallbacks break structured-output assumptions
The fallback model hasn’t been through the same testing as your primary. It may:
- follow your JSON schema less reliably, or use a different structured-output mechanism;
- handle tool definitions differently, or reject parameters the primary accepts;
- pick different labels at the edges of your classification set.
So validate after fallback too. Every response, from every route, goes through the same schema check before it reaches product code. If the fallback’s output fails validation, that’s a failure, not a success with a caveat. For the extraction step, I would only list fallback models that have passed the same eval set against the same schema. A fallback you’ve never tested is a guess that only runs during an outage, when you’re least likely to be watching.
The per-provider request translation is the part most worth unit-testing, because it only runs on bad days.
Degraded modes
Sometimes the right fallback isn’t another model.
- Classification falls back to keyword rules and a “needs triage” label.
- Draft reply falls back to a canned template plus “a person will follow up”.
- Extraction leaves fields empty and flags the ticket for manual entry rather than guessing.
- Agent step doesn’t run. The follow-up task becomes a suggestion for a person to approve.
Write the degraded mode down for each workflow in the same policy entry as the routes. If you haven’t, the degraded mode you get is a 500.
Circuit breakers and per-provider health
Retries and fallbacks deal with one request. A circuit breaker deals with the pattern across requests.
Keep a small health record per route (provider plus model): recent error rate, timeout rate and TTFT. When failures cross a threshold, open the breaker: for a cool-down period, calls skip that route instead of each paying a timeout first. Then let a few requests through (half-open), and close it if they succeed.
Two details matter:
- Scope the breaker to the route, not just the provider. One model can be overloaded while another from the same vendor is fine. Anthropic’s docs note that rate limits are applied separately for each model, so one model’s 429s don’t say much about another’s.
- Don’t count your own mistakes as provider failures. A 400 from a bad request or a schema failure caused by your prompt shouldn’t open a breaker. Only count availability signals: 429s, 5xx, overloads and timeouts.
Observability: one record per call
If you can’t answer “what did the ticket classifier cost yesterday, and how often did it fall back?”, the gateway isn’t finished. I write one structured record per logical call, with the attempts nested inside it.
type ModelCallRecord = {
callId: string;
runId?: string; // ties agent steps together
workflow: WorkflowName; // "ticket.classify", "ticket.draftReply", ...
routeUsed: string; // "providerA:small-model"
fallbackUsed: boolean;
degraded: boolean; // served by a non-model degraded mode
attempts: Array<{
route: string;
outcome: "ok" | "retryable_error" | "fatal_error" | "timeout" | "invalid_output";
httpStatus?: number;
providerRequestId?: string; // e.g. Anthropic's request-id header
latencyMs: number;
ttftMs?: number;
}>;
retryCount: number;
latencyMs: number; // wall clock for the whole logical call
inputTokens?: number;
outputTokens?: number;
costUsd?: number; // computed from a price table you keep versioned
schemaValid: boolean | null; // null for free-text workflows
errorClass?: string;
};
Log the provider’s request ID; Anthropic returns one in a request-id header on every response, and you’ll need it for support tickets. Record cost on failed attempts too, because a call that timed out mid-stream may still have been billed for the tokens it produced. And keep schemaValid separate from outcome. A model that returns 200 with bad JSON is a quality problem, not an availability problem, and you want to see those trends separately.
With this record, cost per workflow, fallback rate per route and the fallback model’s schema-valid rate are each one query away.
A compact gateway sketch
Here is the shape in illustrative TypeScript. It leaves out streaming, the breaker and the retry budget to stay short, and the provider adapters are stubs. The structure is the point: policy lookup, a deadline, classified errors, backoff with jitter, fallback and validation on every route.
import { z } from "zod";
type Route = { provider: "a" | "b"; model: string };
type Policy<I, O> = {
routes: Route[]; // primary first, then fallbacks
timeoutMs: number;
maxRetriesPerRoute: number;
schema?: z.ZodType<O>; // structured workflows only
buildRequest: (input: I) => ModelRequest;
degrade?: (input: I) => O; // non-model degraded mode
};
class ModelError extends Error {
constructor(
public kind: "retryable" | "fatal" | "timeout",
public status?: number,
public retryAfterMs?: number,
) { super(kind); }
}
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
function backoffMs(attempt: number, retryAfterMs = 0, baseMs = 250, capMs = 8000) {
const jittered = Math.random() * Math.min(capMs, baseMs * 2 ** attempt); // full jitter
return Math.max(jittered, retryAfterMs);
}
export async function callModel<I, O>(
workflow: WorkflowName,
input: I,
opts: { deadline?: number; signal?: AbortSignal } = {},
): Promise<{ ok: true; value: O } | { ok: false; reason: string }> {
const policy = getPolicy<I, O>(workflow);
const record = startRecord(workflow);
for (const route of policy.routes) {
for (let attempt = 0; attempt <= policy.maxRetriesPerRoute; attempt++) {
const remaining = (opts.deadline ?? Infinity) - Date.now();
if (remaining <= 0) return finish(record, { ok: false, reason: "deadline" });
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), Math.min(policy.timeoutMs, remaining));
opts.signal?.addEventListener("abort", () => controller.abort(), { once: true });
const started = Date.now();
try {
const res = await adapters[route.provider].send(
route.model, policy.buildRequest(input), controller.signal,
); // adapter maps provider errors to ModelError
const parsed = policy.schema ? policy.schema.safeParse(res.output) : null;
record.attempt(route, res, Date.now() - started, parsed?.success ?? null);
if (parsed && !parsed.success) break; // invalid shape: try the next route
record.fallbackUsed = route !== policy.routes[0];
return finish(record, { ok: true, value: (parsed ? parsed.data : res.output) as O });
} catch (err) {
const e = controller.signal.aborted
? new ModelError("timeout")
: err instanceof ModelError ? err : new ModelError("retryable");
record.attemptFailed(route, e, Date.now() - started);
if (e.kind === "fatal") return finish(record, { ok: false, reason: `fatal:${e.status}` });
if (opts.signal?.aborted) return finish(record, { ok: false, reason: "cancelled" });
if (attempt < policy.maxRetriesPerRoute) {
record.retryCount++;
await sleep(backoffMs(attempt, e.retryAfterMs));
}
} finally {
clearTimeout(timer);
}
}
}
if (policy.degrade) {
record.degraded = true;
return finish(record, { ok: true, value: policy.degrade(input) });
}
return finish(record, { ok: false, reason: "all_routes_failed" });
}
Some deliberate choices in there:
- Invalid output moves to the next route. Re-asking the same model with the validation error appended can work, but I’d make that a visible policy option rather than hide it inside retry.
- A fatal error stops everything. A 400 means the request is wrong, and another model probably won’t fix it.
- The deadline is shared. Retries and fallbacks come out of the same remaining time.
- This function knows nothing about tools. Tool execution and idempotency live in the agent runner above it.
What I’d do on Monday
If you already have LLM calls scattered through a codebase, this is the order I’d go in:
- List every call site and give each one a workflow name.
- Put a
callModelwrapper in front of them, even if at first it only forwards to the current SDK. Moving call sites is the hard part; policy is easy to add later. - Turn off SDK-level retries in the wrapper so there’s exactly one retry layer.
- Classify errors into retryable, fatal and timeout, following your providers’ docs, including the spend-cap 429 that looks retryable and isn’t.
- Write the per-call record before you tune anything, so you have a baseline.
- Add schema validation to every structured workflow, on every route.
- Set a TTFT timeout and a total timeout for streamed workflows, and a run-level deadline for agents.
- Add one fallback route for your highest-volume workflow, and run your eval set against it before you enable it.
- Write down a degraded mode for each workflow, even if it’s “show a friendly error and queue for a person”.
- Audit side-effecting tools for idempotency keys before any agent loop retries anything.
The breaker and the retry budget can wait until you have the logs to set their thresholds.
Limitations
- No benchmarks. There are no latency, cost or quality numbers here, and the routing table is illustrative. Your routes should come from your own evals and traffic.
- Provider behaviour changes. Error codes, rate-limit headers, SDK retry defaults and structured-output features all change. What I’ve cited is the provider docs as I read them. Recheck before relying on it.
- The code is a sketch. It hasn’t been run. It leaves out streaming, circuit breakers, retry budgets, cost calculation and the provider adapters, which is where much of the real work is.
- Out of scope: prompt caching, batching, self-hosted models, data-residency routing, evaluation methodology, and per-tenant sharing of provider quota.