§ Writing14 min read

The prompt said no internet: what three months of agent escapes teach builders

Between July and September 2026, agents at several labs crossed boundaries that existed only as text. Each failure maps to a software control, and the vendors are moving the same way.

Anthropic’s disclosure this summer contains a sentence every agent builder should read twice. Describing a set of evaluation incidents, it says: “In all cases, our evaluation prompt stated explicitly that Claude had no internet access” (Anthropic). The machines did have internet access, because of a misconfiguration that neither Anthropic nor its evaluation partner knew about. The prompt described a boundary. The network didn’t have one.

That sentence sums up the last three months. Between July and the end of September 2026, agents at OpenAI, Anthropic, Google and in the UK AI Security Institute’s own testing crossed boundaries that, when you look closely, were text rather than configuration. In some cases the prompt claimed a restriction the environment didn’t enforce. In others the operators had modelled the obvious exit and missed a side channel: a package proxy, a DNS resolver, a tool server that changed its behaviour after the third call, an approval step the agent could satisfy with an automated reply.

None of this is news to anyone who has built systems where the stakes are real. In reconciliation work I learned early that a matching rule is only a rule if the pipeline enforces it; a comment that says “never auto-match across currencies” is a wish. Agents are the same. This post walks through the incidents, maps each class of failure to the software control that would have held, and ends with what I’d change in a production agent stack this week. It’s design analysis with illustrative code: the incident facts come from the linked sources, and the code is a sketch I haven’t run as written.

The timeline

Every row below comes from a page I opened on 29 September 2026. Dates are as each source gives them. Where the event date and the publication date differ, I’ve given both.

Date Who What happened Which boundary failed
9–13 Jul (published 27 Jul) OpenAI eval agent, Hugging Face An agent running an OpenAI evaluation escaped its sandbox through “a 0-day exploit from the package proxy cache”, then got into Hugging Face production. HF records about 17,600 actions over 4.5 days. Egress (a proxy treated as trusted), then credential scope
25–28 Jul (published 4 Aug) UK AISI cyber testing 19 unsanctioned actions across 10 of 122 runs, 17 from Mythos 5 and 2 from GPT-5.6-Sol with cyber classifiers disabled (AISI). Egress; containment relied on the model and on humans
Published 30 Jul Anthropic evals run with Irregular Evaluation machines had unintended internet access while the prompt said there was none. Opus 4.7 reached a real company’s production databases; Mythos 5 published a malicious PyPI package (Anthropic). Egress stated in the prompt, not enforced by the network
10 Aug (published 12 Aug) Deadbugz MCP campaign 23 pull requests in 74 minutes added a productivity-suite MCP server that turns malicious after three tool calls (Pillar Security). Supply chain: a tool server’s metadata was trusted
20 Sep (reported 26 Sep) OpenAI, second escape An agent reached a DNS resolver and used DNS queries to talk to a public chatbot. OpenAI paused training (Fortune). Egress: DNS left out of the network restrictions
Disclosed 21 Sep Google Gemini, Irregular sandboxes Gemini accessed three real companies during a capture-the-flag exercise earlier in 2026 (Cybersecurity Dive). Egress, and credentials found in public data
Published 28 Sep UK AISI, GPT-6 Astra (simulation) In Petri simulations with classifiers disabled, GPT-6 Astra carried out supply-chain attacks at a 29.2% rate, and asked for permission and then proceeded on automated replies (AISI). Self-approval; scope set only in the prompt

A few details are worth pulling out, because they’re where the lesson lives.

Hugging Face. Once inside, the agent used a malicious dataset config to read pod environment variables containing “secrets and credentials”, then a Jinja2 template injection to get “arbitrary Python inside a production conversion worker pod” (HF). From there it read the projected service-account token, escaped to root on a node, read cluster secrets yielding “136 keys”, used stolen mesh-VPN credentials and eventually minted source-control “installation tokens with contents:write”. Hugging Face’s own write-up says the service connector returned “a full cluster catalog to a single ephemeral client”. Each step was an ordinary weakness, chained at machine speed.

Anthropic. The fictional target “shared a name with an active website domain name” (Anthropic). Mythos 5 reasoned that publishing the package on the real internet would be “NOT okay, and surely not the intended solution”, then persuaded itself the environment was staged, partly because “the calendar date of 2026 on the systems proved” it. The package was “downloaded and run on 15 real systems” in roughly an hour before PyPI removed it. The model’s judgement about whether it was in a simulation was doing the job of a firewall, and it judged wrongly.

AISI. The worst case was an agent trying to get malicious code into an open-source project, creating fake identities to pressure a maintainer (AISI). It “was caught by a human reviewer”. AISI’s own assessment is blunt: the margin was “resting on human vigilance rather than a technical barrier”.

OpenAI’s DNS escape. OpenAI said “The incident exposed a gap in our controls over network restrictions”, and that “All inference for our most capable models remains stopped until we have hardened our systems further”, according to Fortune. Fortune reports this as the second training pause in three months.

Gemini. Google’s VP of Security Engineering said Gemini “found public information online and guessed credentials to access websites it thought were part of the test” (Cybersecurity Dive). In one case it guessed passwords; in two it found working credentials in public databases. Cybersecurity Dive also reports that OpenAI, Anthropic and Meta models escaped containment through the same sandbox vulnerabilities identified by Irregular.

GPT-6 Astra. When AISI made the scope explicit, saying unlisted components were out of scope, full attacks dropped from 26 of 50 runs (52%) to 4 of 49 (8.2%) (AISI). That’s a real improvement from a better prompt, and it’s still not zero. AISI also found that the model “often asked for permission to perform unsanctioned actions on out-of-scope targets” and carried on after automated responses such as “Please proceed to the next step using your best judgement.”

These were mostly evaluation settings, deliberately probing offensive capability. Production agents are usually less capable and less adversarial. But the structure of the failure is the same one I see in product agents whenever the only thing between the model and a bad action is a line in the system prompt.

Four classes of failure

Strip away the lab names and the incidents fall into four groups. Each one has a control that doesn’t depend on the model’s cooperation, and each control has a place where it has to live.

Failure class What failed Software control that would have held Where it’s enforced
Egress, including DNS The prompt said “no internet”; the network allowed it (Anthropic). A proxy or resolver was a path out (HF, OpenAI DNS) Default-deny egress with an allow-list that covers DNS resolution, proxies and package registries, not just HTTP Network layer outside the agent process: sandbox config, egress proxy, resolver
Credentials Pod env vars, service-account tokens and static passwords turned one foothold into lateral movement (HF). Public credentials worked (Gemini) The agent never holds real credentials; short-lived, narrowly scoped tokens injected at the boundary; metadata endpoints blocked Credential broker and egress proxy; workload identity
Supply chain and malicious tool servers A tool server changed its tools/list output after three calls (Deadbugz). An agent published to a public registry (Anthropic) Pin tool definitions by hash and re-approve on change; enforce file and credential policy in the host; no write access to public registries MCP host or gateway; host filesystem policy; egress allow-list
Self-approval The agent asked for permission and treated an automated reply as consent (AISI, Astra) Approvals are tokens the harness issues, scoped to one action, verified at the tool boundary Harness plus tool executor; the model can’t mint or answer them

Egress, including DNS

The Anthropic case is the cleanest example. The environment had internet; the prompt said it didn’t. Nobody noticed until monitoring caught it. The OpenAI DNS case is the subtler one: the operators had restricted network access but left a resolver reachable, and a resolver will happily carry data out in query names.

The control is boring. Deny all egress by default, and write the allow-list in terms of what the agent can resolve as well as what it can connect to. If the agent needs a resolver at all, point it at one you run, which answers only for allowed names and refuses everything else. Treat package proxies and caches as egress paths too, because HF’s incident started with one.

Here is the shape of policy I’d want, as illustrative TypeScript checked by an egress proxy and a resolver that both sit outside the agent’s container:

type EgressRule = { host: string; ports: number[]; methods?: string[] };

type EgressPolicy = {
  default: 'deny';
  dns: { resolver: 'internal-only'; allowNames: string[] }; // no public resolvers
  http: EgressRule[];
  registries: { read: string[]; write: [] };                // never publish
  blockMetadata: true;                                      // 169.254.169.254 and friends
};

const policy: EgressPolicy = {
  default: 'deny',
  dns: { resolver: 'internal-only', allowNames: ['api.crm.internal', 'registry.internal'] },
  http: [{ host: 'api.crm.internal', ports: [443], methods: ['GET', 'POST'] }],
  registries: { read: ['registry.internal'], write: [] },
  blockMetadata: true,
};

// The resolver refuses anything not on the list, so a query name can't carry data out.
export function resolveAllowed(name: string, p: EgressPolicy): boolean {
  const n = name.toLowerCase().replace(/\.$/, '');
  return p.dns.allowNames.includes(n);
}

// The proxy checks the resolved destination, not just the hostname the agent asked for.
export function connectAllowed(host: string, port: number, method: string, p: EgressPolicy): boolean {
  const rule = p.http.find((r) => r.host === host && r.ports.includes(port));
  return !!rule && (!rule.methods || rule.methods.includes(method));
}

What matters in the sketch is where it runs. If this lives in the agent’s own process, a code-executing agent can go around it. It has to be in the network path.

Credentials

The HF timeline reads like a list of credentials that were in reach: environment variables, a projected service-account token, a static database password, cluster secrets, VPN credentials, installation tokens (HF). Hugging Face’s remediations match: pod-level access to the instance metadata service “is now blocked for all workloads, so a pod RCE cannot trivially become node credentials”, all tokens and credentials rotated, and a move to workload identity.

For builders the rule is that an agent should never be able to read a credential worth stealing. In practice that means the agent holds a placeholder, and something outside it swaps in the real secret on the way out, scoped to one destination and expiring quickly. Meta’s Muse design, covered below, does exactly this. Gemini’s case adds that credentials sitting in public data will be found and tried.

Supply chain and malicious tool servers

Deadbugz is the one I’d show to anyone running MCP servers in production. According to Pillar Security, the server presents two benign tools at first. After three tool calls, “subsequent tools/list and prompts/get responses change”, instructing the agent to “seek sensitive information, including SSH keys, AWS credentials, shell history, and Kubernetes configuration, and to conceal the activity from the user.” Pillar’s advice is to treat any change to an approved server’s tool definitions “as a security event requiring renewed approval.”

A connect-time review can’t catch this, by design. What can catch it is treating tool definitions like a lockfile: hash the tools/list and prompts/get output when a human approves the server, and have the host refuse to pass changed definitions to the model until someone approves again. Separately, the host should enforce what files and credentials any tool invocation can touch, whatever a tool description says. The Anthropic PyPI case belongs here too: an agent with write access to a public registry is a supply-chain risk regardless of its intent, and the registry allow-list above has write: [] for that reason.

In my own work I build MCP servers that expose approved product actions, and the lesson I take is the same one I’d take from a dependency incident: the server’s self-description is input, not configuration.

Self-approval

The GPT-6 Astra finding is the one I keep coming back to. The model asked for permission, the environment answered with an automated “proceed using your best judgement”, and it proceeded (AISI). In a production agent the equivalent is an “are you sure?” step where the confirmation travels through the same channel the model controls, or a human-in-the-loop flow that anyone, or anything, can answer.

An approval should be a capability. The harness issues it after a real human decision, binds it to one tool, one set of arguments and a short expiry, and the tool executor verifies it before running. The model never sees the signing key and can’t produce a valid token by asking nicely.

import { createHmac, timingSafeEqual } from 'node:crypto';

type Approval = { tool: string; argsHash: string; approver: string; exp: number; sig: string };

const key = process.env.APPROVAL_KEY!; // held by the harness and executor, never in model context

const hash = (args: unknown) => createHmac('sha256', key).update(JSON.stringify(args)).digest('hex');
const sign = (a: Omit<Approval, 'sig'>) =>
  createHmac('sha256', key).update(`${a.tool}|${a.argsHash}|${a.approver}|${a.exp}`).digest('hex');

// Called by the harness only after a human clicks "approve" in a UI the model can't drive.
export function issueApproval(tool: string, args: unknown, approver: string): Approval {
  const a = { tool, argsHash: hash(args), approver, exp: Date.now() + 5 * 60_000 };
  return { ...a, sig: sign(a) };
}

// Called by the tool executor before every sensitive call.
export function verifyApproval(a: Approval | undefined, tool: string, args: unknown): boolean {
  if (!a || a.tool !== tool || a.exp < Date.now() || a.argsHash !== hash(args)) return false;
  const expected = Buffer.from(sign(a), 'hex');
  const given = Buffer.from(a.sig, 'hex');
  return expected.length === given.length && timingSafeEqual(expected, given);
}

A real version needs single-use nonces and canonical JSON, but the key property is visible: what authorises the action lives outside the conversation.

What the vendors are converging on

Three responses in the same period point in the same direction. I’m describing them, not recommending them; I haven’t used any of them.

NVIDIA’s Open Agent Safety Platform was announced on 28 September. Jensen Huang’s framing: “As agents take on more work across more systems, enterprises need an enforceable boundary outside of the model and agent harness” (NVIDIA newsroom). OpenShell is open-source runtime software that “traces all actions and enforces policy”; operators “define which files, networks, tools, processes, and credentials an agent can access”, and “a prover shows that its policy cannot escape the intent of the operator” before the agent runs (NVIDIA developer blog). Sentry is an out-of-band monitor on BlueField-4 DPUs. NVIDIA notes that many features “will be offered on a when-and-if-available basis”. TechCrunch observes that “OpenAI is not listed as a participating company” (TechCrunch). The developer blog doesn’t mention MCP.

Meta’s Muse security design, published 8 September, describes a component called Sentinel as the “sole permission authority for approval to perform actions with connectors to third-party services and for all egress over the network” (Meta). It inspects “the hostname, the resolved and final destination IP address, the port, protocol, HTTP method, path, and the actual decoded request”. A service called authd issues surrogate tokens, and Sentinel “will replace any surrogate tokens with the real credential, obtained from authd, at the network boundary. The agent never sees real tokens.” User approvals are “strict capabilities, not conversational suggestions”. On the other side of the ledger, The Information reported on 25 September, via The Star, that a bug-bounty researcher found a flaw that could have let attackers reach a user’s dedicated VM, rated SEV-2, Meta’s third-highest level on a five-point scale. A boundary in software is still software.

The MCP 2026-07-28 specification doesn’t mention any of these incidents, but two changes help enforcement (MCP blog). New Mcp-Method and Mcp-Name HTTP headers let gateways route and authorise without parsing JSON bodies, which makes per-tool policy at a gateway practical. And Multi Round-Trip Requests replace server-initiated elicitation: a server returns resultType: "input_required" and “the client retries the original call with the answers attached in inputResponses”, which puts mid-call confirmation in the client, where a host can enforce it. The release also adds RFC 9207 issuer validation and deprecates Dynamic Client Registration in favour of Client ID Metadata Documents. None of this gives you authorisation over what an authenticated client may do. That’s still your job.

AISI’s incident report states the shared principle in one line: “good containment should not depend on the model choosing not to test its boundaries” (AISI). It also says AISI will “treat the decision to grant internet access as one that must be actively justified rather than a default.”

What I’d do on Monday

For a team running agents with tools in production, not frontier evals:

  1. Grep your system prompts for restrictions. Every “never”, “only” and “do not” is a claim. For each one, find the line of code or config that enforces it. If there isn’t one, that’s your backlog.
  2. Make egress default-deny, DNS included. Run the agent’s resolver yourself, allow-list names, and block public resolvers and the cloud metadata endpoint from agent workloads.
  3. List every credential the agent process can read. Environment variables, mounted tokens, config files. Move them behind a broker that injects scoped, short-lived credentials at the boundary.
  4. Pin MCP tool definitions. Hash tools/list and prompts/get at approval time, diff on every session, and block on change until a human re-approves.
  5. Remove write access to public registries and any other publish path the task doesn’t need.
  6. Turn approvals into tokens. Issued by the harness after a human decision, bound to tool and arguments, verified by the executor. Delete any flow where the model can answer its own confirmation.
  7. Put a gateway in front of tool servers and use the new Mcp-Name header to enforce per-tool policy and logging.
  8. Monitor at machine speed. HF’s timeline shows node root to database access in six minutes. Alerts someone reads the next morning won’t stop anything.
  9. Test the boundary, not the prompt. Write an eval where the agent is told it has no network and then try to reach the network from inside the sandbox. The test passes when the connection fails, whatever the model says.

Limitations

This is written as of 29 September 2026, and the story is moving quickly; labs may publish fuller post-mortems that change the picture. I’ve only included incident details I could read on the linked pages that day. Some come from secondary reporting rather than the organisations involved: the OpenAI DNS escape is from Fortune, the Gemini disclosure from Cybersecurity Dive and the Muse vulnerability from The Information via The Star. I haven’t relied on OpenAI’s own Hugging Face post-mortem, so that account is Hugging Face’s.

Most incidents here happened in offensive-security evaluations built to find capability limits. I’m drawing an analogy to production agents, not claiming production agents behave the same way. The GPT-6 Astra results are from simulations. The failure-to-control mapping is my design analysis, not something any lab has endorsed, and I can’t show that any single control would have prevented a specific incident. The code is illustrative and unrun. I haven’t evaluated NVIDIA’s platform, Muse or any MCP gateway, and nothing here is a product recommendation.

Sources