§ Writing13 min read

Sonnet 5.5 ends forced tool use. Here's how to keep structured output reliable

Sonnet 5.5 returns a 400 for tool_choice any or tool. Moving to auto plus strict tools changes how extraction fails, so the pipeline needs a missing-call check, a bounded retry, a fallback and an eval.

For a long time the dependable way to get JSON out of Claude was to define one tool, call it something like extract, and force it with tool_choice. The model had to call the tool, the tool’s input was your JSON, and you never parsed free text. Plenty of extraction pipelines are built on exactly that.

Claude Sonnet 5.5 launched on 28 September, and it rejects that request. According to the migration guide, tool_choice of type any or tool now returns a 400, and so do several other settings older pipelines lean on. The recommended replacement is tool_choice: auto with strict: true tools. That keeps the shape of the tool input guaranteed. It doesn’t guarantee the model calls the tool, or that it puts the right values in it.

This post is design analysis plus illustrative code, based on Anthropic’s docs as of 29 September 2026. I haven’t measured anything here. The running example is an invoice-extraction step: a supplier invoice arrives as text, and the pipeline needs supplier, invoice number, date, currency and amounts as typed fields. For routing, timeouts and retry classification in general, see the gateway post. This one only covers what changes for structured output.

What changed

The “before” column is not the same for every row. Some of these were already rejected by Sonnet 4.6 or Sonnet 5. Which rows affect you depends on the model you’re migrating from.

Parameter or pattern Before On Sonnet 5.5 What to do
tool_choice: {"type": "tool", ...} or {"type": "any"} “Every earlier model on this page accepts” them (Sonnet 5, 4.6, 4.5, 4, 3.7, Haiku 4.5) 400 invalid_request_error: tool_choice: type "tool" and "any" are not supported for this model. The same check runs on the token counting endpoint Send auto, mark the tool strict: true, and “say in the prompt when to use it”. On Bedrock, auto without strict, and validate in your code
tool_choice auto or none Accepted Still accepted. auto is the default Nothing
Non-default temperature, top_p, top_k Accepted on Sonnet 4.6 and earlier and Haiku 4.5 400 Remove them
Assistant prefill (for example, starting the reply with {) Accepted on Sonnet 4.5, Haiku 4.5 and older. Already rejected by Sonnet 4.6 and Sonnet 5 400: “This model does not support assistant message prefill. The conversation must end with a user message.” For output format, the guide says to use structured outputs, “or tools with enum fields for classification”
thinking: {"type": "enabled", "budget_tokens": N} Deprecated on Sonnet 4.6, the only option on Sonnet 4.5 and Haiku 4.5 400 Set an effort level. The guide says there’s no fixed mapping, “so run your evaluations at two or three levels”
thinking: {"type": "disabled"} How Sonnet 5 turned thinking off 400. The message points to between_tools Send thinking: {"type": "between_tools"}, which is accepted at low, medium and high effort only
No thinking field Ran without thinking on Sonnet 4.6 and earlier Runs adaptive thinking Read content blocks by type, pass thinking blocks back unchanged, revisit max_tokens
output_format Deprecated 400 without the structured-outputs-2025-11-13 beta header Move to output_config.format
Strict tools Optional Recommended replacement, with limits: “at most 20 strict tools” per request and additionalProperties: false on every object Mark only the tools that need it

Sonnet 5.5 isn’t the only model affected. The define-tools page lists Opus 5.5, Fable 5.1 and Mythos 5.1 as rejecting any and tool too. So “fall back to the newest Opus” doesn’t bring forced tool use back.

Why it matters for extraction pipelines

Forced tool use was never really about tools. It was a way to make the model’s whole reply a JSON object that you could parse. The failure you worried about was the shape: a string where you wanted an integer, a missing field, text around the JSON.

With auto plus strict, the shape problem mostly goes away. The strict tool use page says tool input “strictly follows the input_schema” and the tool name is always valid. What you lose is the guarantee that a call happens at all. auto, in the docs’ own words, “allows Claude to decide whether to call any provided tools or not”.

So the failure modes move:

Failure Forced tool, no strict auto + strict
Malformed or wrongly typed input The main risk Guaranteed away for the schema subset strict supports
No tool call, answer in text Not possible New, and has to be detected
Valid shape, wrong values Possible Still possible, and now the main risk
Refusal Possible Possible: HTTP 200 with stop_reason: "refusal", and output “may not match your schema”
Truncation Possible Possible, and more likely if you didn’t raise max_tokens, since it now also covers thinking

Two rows are easy to miss. First, a response with no tool call is still a 200. A pipeline that only checks the HTTP status and then reads content[0] will hand product code undefined, or a thinking block. Second, strict mode’s schema subset has gaps. The structured outputs page lists no support for numerical constraints like minimum, string constraints like minLength and pattern, oneOf, or conditional schemas. It also says enum casing isn’t guaranteed: you might get "Gbp" for "GBP". “Schema-valid” is weaker than “valid for my business rules”.

Should the extract tool still be a tool?

If your extract tool was only ever a JSON trick, look at JSON outputs (output_config.format) first. The errors page suggests it “when you need the response itself in a fixed JSON shape”, and it removes the “didn’t call the tool” failure entirely. Refusal and truncation still apply.

I’d keep a tool when the extraction sits inside an agent loop that has other tools, or when you want the model’s “this isn’t an invoice” decision to be a typed call rather than free text. The rest of this post assumes you keep one.

The migration pattern

The principle is the same one the rest of this site argues for: bound the model in software. With forced tool use gone, the prompt says when to call the tool, and the code decides whether the result counts.

Give the model a legitimate way out

A common reason a model skips an extraction tool is that the input doesn’t fit it: a credit note, a statement, a blurry scan. If the only structured option is record_invoice, the honest answer comes back as text and your pipeline has to guess what it means.

So I add a second tool, report_no_invoice, with a small enum of reasons. Both are strict. Now “no tool call” really is a failure, because the model always had a structured way to say no.

Strict tools plus Zod

The JSON Schema goes to the model. Zod holds the rules strict mode can’t express, and checks every route, including routes where strict isn’t available.

import { z } from "zod";

const upper = (v: unknown) => (typeof v === "string" ? v.toUpperCase() : v);

// Rules strict mode can't enforce (patterns, bounds, arithmetic) live here.
export const Invoice = z
  .object({
    supplierName: z.string().min(1),
    invoiceNumber: z.string().min(1),
    issueDate: z.string().regex(/^\d{4}-\d{2}-\d{2}$/),
    currency: z.preprocess(upper, z.enum(["GBP", "EUR", "USD"])), // enum casing isn't guaranteed
    subtotalMinor: z.number().int().nonnegative(),
    taxMinor: z.number().int().nonnegative(),
    totalMinor: z.number().int().nonnegative(),
  })
  .strict()
  .refine((i) => i.subtotalMinor + i.taxMinor === i.totalMinor, {
    message: "subtotal + tax must equal total",
  });
export type Invoice = z.infer<typeof Invoice>;

export const NoInvoice = z
  .object({ reason: z.enum(["not_an_invoice", "unreadable", "multiple_invoices"]) })
  .strict();

const int = { type: "integer" } as const;

export const TOOLS = [
  {
    name: "record_invoice",
    description:
      "Record the fields of the single supplier invoice in the document. " +
      "Amounts are integers in minor units (pence, cents). issueDate is YYYY-MM-DD.",
    input_schema: {
      type: "object",
      properties: {
        supplierName: { type: "string" },
        invoiceNumber: { type: "string" },
        issueDate: { type: "string", format: "date" },
        currency: { type: "string", enum: ["GBP", "EUR", "USD"] },
        subtotalMinor: int, taxMinor: int, totalMinor: int,
      },
      required: ["supplierName", "invoiceNumber", "issueDate", "currency",
                 "subtotalMinor", "taxMinor", "totalMinor"],
      additionalProperties: false,
    },
  },
  {
    name: "report_no_invoice",
    description: "Call this instead of record_invoice when the document is not exactly one readable supplier invoice.",
    input_schema: {
      type: "object",
      properties: { reason: { type: "string", enum: ["not_an_invoice", "unreadable", "multiple_invoices"] } },
      required: ["reason"],
      additionalProperties: false,
    },
  },
] as const;

Every property is required. That’s deliberate: the structured outputs page caps strict schemas at 24 optional parameters and 16 union-typed parameters across the request, and suggests making parameters required where possible.

Build the request per model, in the gateway

A gateway that sends one request body to several Claude models is the thing this release breaks. The same body is valid on Sonnet 5 and a 400 on Sonnet 5.5. So the tool mode becomes part of the route, and product code never sees it.

type ToolMode = "auto_strict" | "auto_plain" | "forced_strict";

type Route = { id: string; platform: "anthropic" | "bedrock"; model: string; toolMode: ToolMode };

// Illustrative. Take model IDs and per-platform capabilities from the provider docs.
export const INVOICE_ROUTES: Route[] = [
  { id: "sonnet-5-5", platform: "anthropic", model: "claude-sonnet-5-5", toolMode: "auto_strict" },
  { id: "sonnet-5", platform: "anthropic", model: "claude-sonnet-5", toolMode: "forced_strict" },
];

export function buildRequest(route: Route, messages: Message[]) {
  const strict = route.toolMode !== "auto_plain";
  return {
    model: route.model,
    max_tokens: 4096, // covers thinking as well as the tool call
    system: EXTRACTION_SYSTEM_PROMPT, // "Always respond by calling exactly one tool."
    messages,
    tools: TOOLS.map((t) => (strict ? { ...t, strict: true } : t)),
    // "any" = must call one of the two tools. Only send it to models that accept it.
    tool_choice: route.toolMode === "forced_strict" ? { type: "any" } : { type: "auto" },
    // No temperature, top_p, top_k or prefill on any route.
  };
}

The fallback is Sonnet 5 because the migration guide says the earlier models it covers accept forced tool choice, and Sonnet 5 is on the structured outputs supported-models list. The define-tools page says that on models that support forced tool use, any plus strict guarantees both a call and a valid shape. That’s the old guarantee, on a model that still offers it. It’s also the model you’re trying to leave, so treat it as a stopgap and watch how often it’s used.

Detect, correct once, fall back, validate

type Read =
  | { kind: "ok"; value: Invoice }
  | { kind: "no_invoice"; reason: string }
  | { kind: "no_tool_call" }
  | { kind: "invalid"; toolUseIds: string[]; issues: string[] }
  | { kind: "refusal" | "truncated" };

function readExtraction(res: ModelResponse): Read {
  if (res.stop_reason === "refusal") return { kind: "refusal" };
  if (res.stop_reason === "max_tokens") return { kind: "truncated" };

  const calls = res.content.filter((b) => b.type === "tool_use");
  if (calls.length === 0) return { kind: "no_tool_call" };
  const ids = calls.map((c) => c.id);
  if (calls.length > 1) return { kind: "invalid", toolUseIds: ids, issues: ["call exactly one tool"] };

  const [call] = calls;
  const parsed = call.name === "report_no_invoice"
    ? NoInvoice.safeParse(call.input)
    : Invoice.safeParse(call.input);
  if (!parsed.success) return { kind: "invalid", toolUseIds: ids, issues: parsed.error.issues.map((i) => i.message) };
  return call.name === "report_no_invoice"
    ? { kind: "no_invoice", reason: (parsed.data as z.infer<typeof NoInvoice>).reason }
    : { kind: "ok", value: parsed.data as Invoice };
}

function correction(read: Extract<Read, { kind: "no_tool_call" | "invalid" }>) {
  if (read.kind === "no_tool_call") {
    return "You replied without calling a tool. Respond only by calling record_invoice or report_no_invoice.";
  }
  // The model did call a tool, so the next turn must answer each call.
  return read.toolUseIds.map((id) => ({
    type: "tool_result", tool_use_id: id, is_error: true,
    content: `Rejected: ${read.issues.join("; ")}. Call exactly one tool with corrected values.`,
  }));
}

const MAX_CORRECTIONS = 1;

export async function extractInvoice(doc: string, routes = INVOICE_ROUTES) {
  const attempts: { route: string; outcome: Read["kind"] }[] = [];

  for (const route of routes) {
    const messages: Message[] = [{ role: "user", content: invoicePrompt(doc) }];

    for (let turn = 0; turn <= MAX_CORRECTIONS; turn++) {
      const res = await gateway.send(route, buildRequest(route, messages));
      const read = readExtraction(res);
      attempts.push({ route: route.id, outcome: read.kind });

      if (read.kind === "ok") return { kind: "invoice", value: read.value, attempts } as const;
      if (read.kind === "no_invoice") return { kind: "no_invoice", reason: read.reason, attempts } as const;
      if (read.kind === "refusal" || read.kind === "truncated") break; // a nudge won't fix these

      // Append-only: the model's turn goes back exactly as returned, thinking blocks included.
      messages.push({ role: "assistant", content: res.content });
      messages.push({ role: "user", content: correction(read) });
    }
  }
  return { kind: "needs_review", attempts } as const; // a person enters the fields
}

A few choices in there:

  • One correction, then move on. A second nudge on the same document mostly buys cost. The bound is config, not a loop that runs until it works.
  • The correction appends; it never edits. Sonnet 5.5 thinking blocks are signed over the conversation before them, and for newer accounts a replay after an edit to earlier history returns a 400.
  • Refusal and truncation skip the correction. For max_tokens stops the docs say to retry with a higher max_tokens; if that happens often, raise it in the route. Anthropic’s server-side fallback can handle refusals on the Claude API, but decide which layer owns that retry rather than running both.
  • Zod runs on every route, and attempts goes into the per-call record from the gateway post, so correction and fallback rates show up in the logs.

Platform differences

Claude API (first party). Strict tools and JSON outputs are available for Sonnet 5.5. Server-side refusal fallback is in beta on the Claude API only, and not on the Message Batches API.

Amazon Bedrock. The migration guide is explicit: structured outputs, “which include strict tool use, aren’t available for Claude Sonnet 5.5” on Bedrock. The instruction there is to “send auto without strict, say in the prompt when to call the tool, and validate the tool input in your code”. The structured outputs page’s Bedrock note lists only Opus 4.6, Sonnet 4.6, Sonnet 4.5, Opus 4.5 and Haiku 4.5, so a Sonnet 5 fallback on Bedrock would be forced but not strict. On Sonnet 5.5 via Bedrock you have neither guarantee, and Zod becomes the only check. Server-side refusal fallback isn’t available there either; the docs point to the SDK’s client-side middleware.

In config, that’s a separate route with toolMode: "auto_plain". Claude Platform on AWS is a different offering, which the structured outputs page lists as generally available without a model exception, so don’t let “AWS” become one flag.

Google Cloud. The structured outputs page lists Google Cloud as generally available with no model note, and the migration guide only makes an exception for Bedrock. I didn’t find a page that states Sonnet 5.5 strict-tool support on Google Cloud directly, so I’d confirm it with a test request before relying on it.

Testing the migration

Run the old configuration (Sonnet 5, forced) and the new one (Sonnet 5.5, auto plus strict) over the same set of documents. Anthropic’s own launch page is a reminder why: it notes that a pre-release deployment “had a bug that could degrade responses to requests that use structured outputs”, since fixed. Measure it yourself rather than assuming.

  1. Build a golden set of real, redacted documents with hand-checked fields, including credit notes, statements, two invoices in one PDF and poor scans.
  2. Tool-call rate, first turn and after one correction, per route. Forced tool use made this metric pointless; now it’s the first one to watch.
  3. Schema validity: the share of calls that pass Zod, tracked separately from strict’s guarantee.
  4. Content correctness per field against the golden set: exact match for IDs and amounts, normalised match for names. An aggregate hides “supplier name is wrong 1 time in 10”.
  5. False no-invoice rate: real invoices reported as not_an_invoice. The escape hatch is only safe if you measure it.
  6. Refusal, truncation and fallback rates, from the attempts log.
  7. Latency per successful extraction, p50 and p95, including corrections and fallbacks. Note that the first request with a new schema has extra latency while the grammar compiles.
  8. Cost per successful extraction: total spend across every attempt, correction and fallback, divided by extractions that passed validation and matched the golden set. Tokens per request is the wrong number here.
  9. Effort level. Levels are recalibrated on Sonnet 5.5, so run the set at two or three levels, including between_tools if you ran without thinking before.

What I’d do on Monday

  1. Grep for tool_choice, temperature, top_p, top_k, budget_tokens, "disabled", output_format and prefill, including token-counting calls.
  2. Pin routes that send forced tool use to a model that accepts it, so no alias moves them to Sonnet 5.5 before you’ve tested.
  3. Move request construction into the gateway as a per-route toolMode, with a unit test per route.
  4. Add the explicit “no” tool, make strict properties required, and put Zod behind every route.
  5. Add the missing-call check and one bounded correction, and log each attempt.
  6. Run both configurations over the golden set before switching traffic. Treat Bedrock as its own route with no forced call and no enforced shape.

Limitations

  • Facts as of 29 September 2026. Everything cited is Anthropic’s docs and launch page as I read them on that date, one day after launch. Model docs change, and feature availability per platform changes most often. Recheck before relying on any row of the table.
  • No measured results. I haven’t run these configurations. I’m not claiming anything about tool-call rates, accuracy, latency or cost on Sonnet 5.5, and nothing here should be read as a benchmark.
  • The code is a sketch. It hasn’t been run. gateway.send, Message, ModelResponse and the prompts are stand-ins, and SDK types will differ in detail.
  • Google Cloud and Microsoft Foundry aren’t confirmed here beyond what the structured outputs page says at platform level.
  • Out of scope: other providers’ structured-output features, JSON outputs in depth, streaming partial tool input, and batch extraction.

Sources