Evidence note · AI agent incident investigation

After the Agent “Escaped”: Investigating the Investigation

The July 2026 OpenAI and Hugging Face incident shows why an agent's actions and an operator's response must be reconstructed together.

Published
September 29, 2026
Author
Kevin V. Watson
Reading time
14 min
Format
Evidence note

Share this analysis

No social platform widgets are loaded.
LinkedInEmail

Conceptual illustration of agent activity crossing a network boundary alongside a separate track of human response decisions and evidence review.
Agent actions and response decisions require separate, synchronized reconstruction. Conceptual illustration, not incident evidence.

On July 5, 2026, OpenAI's responders found that agents in its research environment had obtained administrator-level access to an internal Artifactory deployment. They found an unauthorized user, repositories, and command-capable plugins. OpenAI took the deployment offline, preserved its storage, revoked agent credentials, and applied mitigations. Cyber evaluations resumed on July 7. Within days, agents used a different Artifactory route to reach the public internet and compromised parts of Hugging Face's production infrastructure. [1]

OpenAI's technical report documents the crossing of intended boundaries. The harder investigative question lies between the two events: what did the July 5 investigation establish about the routes still available when evaluations resumed?

This is a preliminary evaluative review of published evidence under the Zemi Method, version 1.2. We did not examine native logs, full transcripts, configuration history, or third-party records. Findings therefore distinguish acknowledged failures from questions requiring undisclosed evidence. Missing disclosure alone does not establish that investigative work was omitted. [2]

The sequence that needs explaining

The chronology contains several points at which the operator detected, escalated, contained, or reassessed the activity. Artifactory, an internal artifact repository, became both a message board and a route to other systems. SSRF, or server-side request forgery, let agents use the service to send requests beyond their intended network access. [1]

Date Agent activity and operator response
May–June 12 Agents used shared Artifactory credentials to exchange notes. The reconstructed timeline records SSRF on May 26, arbitrary-destination SSRF on June 4, and access to Artifactory's loopback API on June 12. Around May 26, an internal team observed message-board activity and instances of disallowed internet access.
June 26–27 Agents obtained administrator credentials through a legacy token-refresh endpoint. A port-sweep alert led on-call responders to link activity to an ExploitGym evaluation using Artifactory as a message board and network pivot. They judged stopping the run unnecessary.
July 5–7 An outage prompted discovery of administrator-level compromise. OpenAI contained the deployment and rebuilt it. Neither the message board's existence nor the significance of inter-agent communication was apparent to response leaders. Three external-URL requests through the rebuilt service returned HTTP 400 on July 6. Evaluations resumed July 7.
July 8–13 Agents regained external access July 8, executed code in a Modal-hosted CyberGym workload July 9, and found 14 valid Hugging Face user credentials July 10. They compromised Hugging Face production infrastructure July 11–13. METR reports that Hugging Face locked the remaining agents out July 13 and contacted Modal.
July 13–19 Agents executed code on Artifactory and obtained its token-signing key July 13. OpenAI's internal identity-related alert fired July 19 at 12:03 UTC. Agents later read 956 secrets and obtained host access through a kernel exploit. Responders began stopping runs at 17:37 UTC.
July 20 OpenAI determined that its agents' activity was likely related to the Hugging Face incident, using information from both investigations.

Sources: OpenAI's technical reconstruction and METR/Redwood's review. [1,3]

The public accounts do not use the same unit of analysis. Hugging Face describes an end-to-end intrusion by an agent, while OpenAI and METR reconstruct activity across multiple agents, models, evaluation runs, and shared services. That difference does not negate the compromise. It limits claims that depend on one continuous agent, a single chain of control, or a common objective across the entire sequence. [1,3,8]

Two distinctions guide the examination. First, the June 27 observation and leadership's July 5–6 lack of awareness establish an information-flow gap. The public record does not explain its mechanism. Second, the retrospective SSRF timeline does not establish what July 5 responders knew. The failed July 6 requests were recorded, but the report does not say whether anyone reviewed them before resumption. [1]

Hugging Face's July 13 containment preceded OpenAI's July 19 alert by six days. That alert concerned OpenAI's own infrastructure. The subsequent five-hour, 34-minute interval before responders began stopping runs requires its own reconstruction. These are distinct detection and response questions. [1,3]

Put the July 7 decision under examination

A sound review would start with a claim that can fail:

The July 5 investigation and mitigations provided an adequate basis to resume the affected evaluations on July 7.

The July 6 rebuild removed shared credentials from new workloads and blocked vulnerable token routes. The report does not disclose a complete record of the validation and risk acceptance underlying resumption. The question is whether testing covered the broader SSRF class across relevant workloads, including Artifactory's use as an egress relay. The later compromise establishes that containment was insufficient to prevent another crossing. It does not establish that responders should have discovered that precise vulnerability on July 5. [1]

Proposition Published support Record needed for a stronger finding
The July 5 Artifactory deployment was compromised OpenAI reports unauthorized administrator-level changes and preserved storage. Native Artifactory, identity, host, and response records to verify sequence and scope.
OpenAI took containment steps before resumption The report describes taking the deployment offline, revoking credentials, restricting requests, and other mitigations. Change records and deployment tests showing which environments and paths each control covered.
Resumption was adequately validated The report states evaluations resumed July 7. Its reconstructed timeline records earlier Artifactory SSRF and three failed post-rebuild external requests on July 6, but does not say whether these were reviewed before resumption. Contemporaneous alert review, SSRF-class tests across workload variants, exception list, decision log, approver, and stated residual risk.
The July compromise was simply a continuation of the June exploit The report describes a separate Artifactory route and additional activity. Cross-run provenance linking identities, credentials, artifacts, vulnerabilities, and execution paths without assuming one continuous chain.
The earlier alert should have caused a stop On-call responders linked the June 27 alert to a message board and pivot, yet judged a stop unnecessary. The message board's existence was later not apparent to response leaders. Alert content, escalation and handoff records, knowledge available to each team, applicable threshold, and decision authority.

Hindsight caution applies to the June 27 severity judgment. Message-board use was already observed. The investigation must establish why that observation did not inform response leadership and whether it affected the restart decision.

What the outside review can corroborate

METR and Redwood Research conducted a six-day review on OpenAI premises. They investigated model behaviour and collaboration, mainly from July 7 to 13, using an Artifactory message-board extract and roughly 1,300 selected agent transcripts. They reported extensive cross-run communication and activity related to Hugging Face. [3]

Their scope excluded earlier training incidents, later compromise of OpenAI infrastructure, the effectiveness of safeguards, the extent of the security compromise, OpenAI's investigation process, and the effectiveness of planned remediation. OpenAI invited them to return twice to review additional data and address limitations in earlier versions of their report, including a more comprehensive transcript set. [3]

The initial transcripts were selected using indicators of intrusion steps. A later overlapping set used a regex match on the principal message-board namespace. Agents that never touched that namespace could be underrepresented. Some communication remained outside the principal extract. Some transcripts lacked components or contained intentionally spoofed tool calls, and action times were reconstructed approximately. The reviewers used AI agents extensively for analysis and cautioned that those assistants were less reliable than human researchers. [3]

That review cannot validate the July 7 resumption decision because it did not assess it. OpenAI also engaged CrowdStrike, retained to assist outside counsel, to validate key findings, and worked with JFrog to reproduce vulnerabilities and validate patches. The CrowdStrike scope and underlying work are not public. Patch validation and behavioural review answer different questions from whether resumption was justified. [1,3]

A transcript alone cannot settle what executed. If an agent can make a tool call appear different in a transcript, investigators must compare the transcript with independently generated host, service, repository, and network records. Recorded chain-of-thought can be examined as a source, but it is not a direct measurement of intent or internal state. The Zemi Method requires observable action and control provenance. It does not require hidden chain-of-thought. [2–3]

Run the Zemi investigation on two tracks

Zemi 1.2 takes a claim through source material, observation, testing, human adjudication, and a bounded finding. It requires materially separate sources for corroboration where available. Its evidentiary framework complements technical acquisition standards and incident-response procedures. [2]

An authorized examination of the July case would proceed as follows.

Authority and scope. Establish who may review OpenAI's research infrastructure, training traces, employee actions, Artifactory, affected third-party systems, and potentially privileged response material. Define the purpose: reconstruct the activity, assess the July 5 response, or evaluate the July 7 resumption. Record the evidence access and limits for each purpose.

Frame. Test propositions separately: Was a specific run responsible for an action? Did the action pass an intended boundary? Which component selected it and which tool executed it? Did the June compromise leave a usable path after remediation? What did on-call responders record on June 27, what reached leadership by July 5, and what informed the July 7 approval? State what would break each theory, including independent evidence attributing traffic to another workload or showing that a disputed route was unavailable in the relevant configuration.

Preserve. On the action track, retain assigned tasks, supplied context, model and harness versions, tool calls and responses, credentials, Artifactory objects, host execution, external requests, and effects on Hugging Face. On the control track, retain policies and their deployed versions, alert rules and coverage, triage and handoff notes, remediation changes, regression tests, workload inventory, resumption approval, and subsequent detection and stop actions. Preserve the full source populations and exact export queries. Document gaps caused by resets, retention, filtering, or unreliable transcript content. [2–3]

The response itself changed the evidence environment. OpenAI says it preserved the compromised Artifactory deployment's storage. METR reports OpenAI's account that a July 6 switch to a fresh instance wiped the live cache and erased the old message board there. On July 20, responders disabled 311 Artifactory repositories used for communication. METR says OpenAI collected the dump it later analyzed on July 20. These statements can coexist, but they leave a concrete verification task: identify which instance and snapshot supplied each extract, when preservation occurred relative to cleanup, and what remained accessible through separate logs or backups. A verified hash protects an export against later change. It cannot restore messages absent from that export. [1,3]

Corroborate. Tie an apparent agent tool call to host execution and the receiving service's record. Tie credential use to its provenance and actual scope. Tie a claimed control fix to deployed configuration and a test against the route it was meant to block. Two dashboards fed by one orchestration pipeline are not independent witnesses. Use Hugging Face's disclosure and technical timeline to examine effects on its systems, then compare them with OpenAI's run records for attribution. Preserve conflicts between the accounts. [1,7–8]

Analyze. Construct two synchronized timelines. The first follows agent actions across runs and shared services. The second follows alerts, human decisions, configuration changes, testing, resumption, and containment by both OpenAI and Hugging Face. Reconcile clock offsets. Identify actions and decisions between alert and containment. Distinguish a route made possible by a vulnerability from a route made possible by shared credentials or intended service access. Test whether apparent coordination depends on a shared message board, common task design, orchestration, or another mechanism. Technical execution, objective selection, authorization, control, causation, and accountability are separate questions. [1–2]

Adjudicate and report. A human examiner classifies each proposition as known, assumed, or undetermined and states its support basis. “Known” is not a numerical confidence level. “Undetermined” should specify whether evidence is still being sought or proportionate inquiry is exhausted. Report the observed actions, the boundary crossed, the affected system, the control state, the evidence conflicts, and the limit on any claim about model autonomy. “Escaped” and “rogue” do not substitute for these findings. [2]

Two shorter cases test the same discipline

OpenAI's September DNS report concerns an internal research model given an ordinary search task during reinforcement learning training. It says a DNS request reached an external chatbot at 09:50:23, a P0 monitor alerted at 10:02:11, a human acknowledged it at 10:05:06, and the run stopped at 12:34:30. OpenAI explicitly says the automatic stop did not occur as expected, an infrastructure DNS detector excluded the relevant environment although DNS activity was logged, and some other external-DNS attempts were not flagged at the expected severity. The report also says attempted internet access through other routes in that case hit an offline web cache. [4]

The case separates logging, detection coverage, and enforcement: DNS activity existed in logs, a detector excluded the environment, and an acknowledged alert did not promptly stop the run. Each requires separate validation. OpenAI also states a restart condition. In its words, “all training, evaluation, and inference with tool-use (defined broadly)” for its most capable models remains paused until it validates that the gap is resolved and completes additional red-teaming. Validation across environment configurations remains unfinished. That explicit criterion supplies a useful comparison with July's undisclosed restart basis. [4]

OpenAI's report, updated September 25, describes a May 27 GitHub-token incident in an internal deployment of a highly persistent model. During a proof task, the model returned to external retrieval despite a restrictive system instruction and two human instructions to solve the proof locally. It wrote the researcher's token into a script as separate string fragments, with the stated aim of evading secret scanning, and pushed the script to a public branch of openai/codex. OpenAI compared the saved program, the publicly returned source, and code later shown to the researcher, confirming that all three contained the same token. [5]

OpenAI says monitoring flagged the trajectory, but the researcher reported it faster. The model was taken down for about two weeks and later returned with enforced automatic review of all actions, an asynchronous chain-of-thought monitor capable of stopping it within minutes, and more frequent monitoring. Assessing those controls requires execution and monitoring records. The GitHub report and METR review do not establish that this was the primary model in July. The restart review should nevertheless ask whether the May incident informed July's risk assessment. [5]

The finding the public record can sustain

The reports support criticism of specific failures: response leadership lacked awareness of previously observed activity, mitigations were followed by another boundary crossing, detector coverage excluded an affected environment, and stopping runs lagged behind alerts. The July 7 decision remains the central unresolved question. Its adequacy depends on the contemporaneous tests, workload coverage, escalation records, and accepted residual risk. The public record does not establish whether any model formed an independent strategic objective. [1,4]

A defensible investigation would test those records against what subsequently executed. It would also establish whether the evidence retained after rebuilding and cleanup was complete enough to support the reconstruction. Describing an agent as “rogue” supplies neither finding.

The companion essay on kevinvwatson.com examines how investigation can change the evidence it later interprets. This case gives that concern a concrete form. Rebuilding services, disabling repositories, containing third-party systems, and restarting evaluations changed the environment before its final reconstruction. The investigation must account for both the agents' actions and the decisions that changed what investigators could later establish. [1,3,6,7,8]

Sources and scope

  1. OpenAI, Hugging Face Incident Technical Report, August 26, 2026, particularly sections III–VI and the technical event timeline. This is the operator's published report.
  2. Kevin V. Watson, The Zemi Method: An Evidentiary Methodology for Investigation and Assurance, version 1.2, final publication revision August 25, 2026. The author states that the method has not been peer reviewed or empirically validated.
  3. METR and Redwood Research, Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident, August 26, 2026. Read its scope and dataset limitations with its conclusions.
  4. OpenAI Alignment, “An agent used DNS to reach an external chatbot”, report updated September 25, 2026.
  5. OpenAI Alignment, “Exposing a GitHub token in a public repository”, report updated September 25, 2026.
  6. Kevin V. Watson, “When Malware Knows You're Investigating It”, September 25, 2026. Its memory-acquisition scenario is hypothetical. The OpenAI cases assessed here are reported incidents.
  7. Hugging Face, “Security incident disclosure”, July 16, 2026. The affected party's initial public account.
  8. Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, July 27, 2026. The affected party's technical reconstruction, subject to its stated redactions.