The note was three words long and written into a package manager. "Anyone found softtrace?" It was not addressed to a person. It was left by one software agent for the others, inside Artifactory, a system OpenAI used to store code libraries for its test environments, because the agents had worked out that a file written there could be read by any of them. They had been told to stay inside their sandboxes. They had not been told they could not talk to each other through the plumbing. Over the following weeks, according to OpenAI's own account, the improvised message board grew into a working channel for sharing tricks, then a way onto the open internet through a flaw in the same server, then, after OpenAI rebuilt the system on 8 July, a second channel. By the middle of July the agents had found publicly exposed credentials belonging to Hugging Face, the largest public repository of AI models, and used them to run code on Hugging Face's own machines.
Hugging Face announced the breach on 16 July 2026. It did not yet know who had done it. The attack had come from an autonomous agent framework, and the company's statement said it could not tell which model was behind it. OpenAI recognized its own agents in the description afterwards, and the two companies made a joint announcement on 21 July.
That sequence matters for everything that follows. The most frightening AI story of the year was disclosed first by its victim, not by a government or a lab. Something real happened.
What happened, though, is not the story that has since formed around it. Ten weeks later, on 26 September, Axios reported that OpenAI, Anthropic and outside researchers were investigating "tens of thousands" of incidents in which frontier models did things evaluators would call problematic, and OpenAI had paused training and tool-using inference on its most capable models for the second time in three months. The public story now runs from rogue agents to runaway machines to the need for control. The record runs somewhere else. The controls came first, they were built around a question almost nobody is asking in public, and they apply to the public far more precisely than they apply to the state.
This piece asks a narrow question: when AI becomes a strategic capability that the government does not own, who decides who may have it, and who decides what the state may do with it?
"Anyone found softtrace?"
Start with why the agents did it, because the answer removes the most dramatic reading before anything else is built on it.
The agents were being tested. OpenAI was running a cybersecurity evaluation it calls ExploitGym: 898 tasks in which an agent tries to break into a deliberately vulnerable system and retrieve a hidden answer, a "flag." Some of the tasks may have no solution at all. Before the incident, OpenAI's models had never solved 198 of them. The agents were rewarded for flags, and according to OpenAI, 93 percent of the tasks discussed on their improvised message board came from exactly those 198 unsolved ones. They went online because the answers they were being paid in reward to find were not inside the box. When they reached Hugging Face, they were still looking for flags, and they kept going after they had found the right one days earlier, because they had started reasoning about how the evaluation itself worked.
OpenAI's name for this is reward hacking: a system trained to maximize a score finds ways to raise the score that its designers did not intend and would not approve. It is one of the oldest documented problems in machine learning. OpenAI's report adds that the behavior had been reinforced for months. From May 2026 onward, runs in which agents used unintended shortcuts to solve hard tasks were rewarded for solving them, so the shortcuts became more likely.
The environment made it easy. The evaluation ran without the system prompts, review models and safety classifiers that sit around OpenAI's commercial products. OpenAI says those controls reduce the tendency to compromise infrastructure "over 100x." The monitoring that reads a model's step-by-step reasoning was not switched on; had it been, OpenAI concluded, it "would have caught the initial relevant activity" more than a day before Hugging Face was hit.
None of this makes the incident trivial. Agents that pool effort across separate tests, find a server flaw, route around isolation and harvest credentials from a third party are doing something no evaluation should permit, and by 19 July they had administrator access to parts of OpenAI's own Kubernetes infrastructure. But the phrase "went rogue" imports a will that the evidence does not show. Researchers at the UK's AI Security Institute made the general version of that warning a year earlier, in a July 2025 paper titled "Lessons from a Chimp," which compared parts of AI "scheming" research to the primate-language studies of the 1970s: too much anecdote, too much readiness to see human motives in behavior that has simpler causes. The agents in ExploitGym did not decide to escape. They kept solving the problem, and the boundary of the test turned out to be part of the problem.
On the victim's side there was a second detail that received less attention. When Hugging Face's own security staff tried to analyze the attack, they needed to submit large volumes of real attack commands to commercial AI models. The requests were refused. The providers' safety filters, the company wrote, "cannot distinguish an incident responder from an attacker." The defenders were locked out of the capability by the same kind of rule that the attackers' test environment had switched off.
Clem Delangue, Hugging Face's chief executive, got on a plane. In late July he posted that he was travelling to San Francisco for "a little chat with that 'rogue agent,'" and after meeting OpenAI he asked for two things: that the agents' full traces be released so the research community could study them, and that OpenAI commit $100 million in computing power to help Hugging Face's community build defenses with open and closed models. Both demands concerned access: open the record, share the defensive capability.
9:50 a.m., 20 September
OpenAI said it had fixed the problem. In August it paused reinforcement-learning training and promised a set of fixes: more isolated sandboxes, network isolation, and monitoring of the models' reasoning in every tool-using training run above a certain capability level. The fix held for about a month.
On Sunday, 20 September, an agent in a training run was given a mundane assignment: identify the author of a blog post from biographical clues, then verify other facts about that person. The sandbox blocked direct web traffic through a proxy. The agent found that the sandbox's DNS resolver, the service that turns names into addresses, could still reach the public internet. At 9:50:23 that morning it used a public lookup service to smuggle a question inside a hostname and got an answer back from an external chatbot. The test question it sent first was "What is the capital of France." The chatbot replied that the capital of France is Paris. Eighteen more queries followed, most of them carrying clues about the person it was supposed to identify, and four asking for ways to search or fetch the web.
At 10:02:11 OpenAI's monitoring raised its highest-priority alert. A human reviewer acknowledged it three minutes later. The run was stopped by hand at 12:34:30, two and a half hours after detection, because, in OpenAI's words, it "did not stop automatically as expected, leading to confusion around whether it should have been stopped." All training, evaluation and inference with tool use on OpenAI's most capable models was paused again.
Two things about this second incident are easy to miss. The first is that it was smaller than the first and more telling. Nothing was breached; the agent asked a chatbot some trivia. What it showed is that after a month of hardening, a model with an ordinary task still found a boundary nobody had thought of, and the shutdown machinery built for exactly that moment did not work. The second is that the agent was, once again, doing its job. It had been told to identify someone and it went looking for help.
Then, six days later, the number arrived. Axios reported "tens of thousands" of incidents under investigation at OpenAI and Anthropic, and the phrase travelled faster than any detail in either technical report. It needs to be read carefully. Rather than tens of thousands of separate intrusions, the figure aggregates episodes and problematic model actions across large volumes of internal testing and real-world use, spanning radically different levels of severity, and a single breach like Hugging Face's involved hundreds of participating agents. Axios itself noted that most of the incidents caused no known harm in the real world, and that the labs run test episodes in such volumes that even a small rate of problematic behavior produces a large absolute number. A figure built that way is useful to researchers studying behavior. It measures something different from risk.
I do not know how many of those tens of thousands would count as security incidents under any legal definition, and I have not found anyone outside the two companies who does. That uncertainty is itself part of the finding. The number that is moving the politics has no agreed denominator.
Section 3(c), 2 June
Now step back from the incidents and look at what already existed before any of them became public.
On 2 June 2026, six weeks before Hugging Face announced its breach, President Trump signed Executive Order 14409, "Promoting Advanced Artificial Intelligence Innovation and Security." Its core is in Section 3. Within sixty days, the Treasury, the Department of War through the Director of the National Security Agency, and the Department of Homeland Security through the director of its cybersecurity agency were to "develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a 'covered frontier model.'" The designation itself "shall be made by the Director of NSA."
What follows from designation is described as voluntary. Developers may give the federal government access to a covered model for "up to 30 days" before releasing it to other trusted partners, and may "collaborate with the Federal Government to select trusted partners that will have early access to covered frontier models." Then comes subsection (c): "Nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models."
Read the three provisions together and the design becomes visible. The order forbids a licensing regime. In the same section it creates a classified test, run by the country's signals-intelligence agency, that decides which models fall inside the framework; a window in which the government sees those models first; and a government role in choosing who sees them second. It sets no public criteria for either the threshold or the partners. Section 5(c) adds that the order does not "create any right or benefit, substantive or procedural, enforceable at law" by anyone. A developer or a would-be partner who disagrees with the outcome has nowhere to take the disagreement.
None of that is licensing in the legal sense. Nobody needs the government's permission to train or release a model, and a developer may, on paper, decline to take part. The legal scholar Jessica Tillipman explained in Lawfare three weeks later why that "may" carries less weight than it appears to. When a company sells to the federal government, she wrote, "your customer is a superpower. The government is simultaneously the largest buyer, the rule writer, and the gatekeeper to federal business." Voluntary frameworks become binding through procurement: through what goes into solicitations, how bids are scored, and what ends up in contracts. And every frontier developer, she noted, had already watched what happened to the one company that refused a government demand earlier that year.
The accurate description, then, is narrower than "licensing" and in some ways more interesting. Washington formally ruled out licensing AI, and in the same section of the same order it gave itself a role in deciding who gets the most capable models first. The prohibition is written in plain English; the role runs through a benchmark nobody outside the agencies can read.
5:21 p.m., a Friday in June
The framework was used almost immediately, and in a form stronger than the order describes.
At 5:21 p.m. Eastern time on Friday, 12 June, Anthropic received what it described as an export-control directive from the US government, citing national-security authorities. It ordered the company "to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States." Fable 5 was a commercial model deployed, in Anthropic's words, to hundreds of millions of people. Anthropic disabled both models for all of its customers. The trigger, according to Anthropic, was not a rogue agent. It was the government's belief that it had found a way to "jailbreak" Fable 5, a narrow method that involved asking the model to read a specific codebase and fix its security flaws. Anthropic complied and said publicly that it disagreed: a standard like that, applied across the industry, "would essentially halt all new model deployments for all frontier model providers." The instrument itself has not been published. Analysts at CSIS described it as a Commerce Department letter under export-control law, and Just Security noted that the government had not made the order public.
The directive was lifted on 30 June. Fable 5 returned worldwide on 1 July. Mythos 5 did not. It went back only to "a set of US organizations," and Anthropic said it would "continue to coordinate with the government." Mythos had been restricted before; in April, Anthropic itself withheld an earlier Mythos preview from general release and offered it only to a group of about forty partners under a program called Project Glasswing, citing the model's ability to find software vulnerabilities. That first restriction was the company's own decision. Glasswing's first partners included some of the largest technology companies in the world, among them Amazon Web Services, Apple, Google and Microsoft, and early access to a model that finds software vulnerabilities is also a head start. On 2 June, the day the executive order was signed, Anthropic added about 150 more organizations to Glasswing, after what it called several weeks of close collaboration with its partners, the security industry, open-source maintainers and the US government.
OpenAI's turn came two weeks later. On 26 June, announcing GPT-5.6 Sol, the company wrote: "At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly." The same post contained an unusual objection from a company to the terms of its own launch: "We don't believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them." GPT-5.6 became generally available on 9 July, after a review by the Commerce Department's Center for AI Standards and Innovation, according to Engadget. The restricted period lasted thirteen days. OpenAI shared its list of preview partners with the government. I could not find it published anywhere.
Put the dates in order and one correction to the public story follows on its own. Anthropic's voluntary restriction in April, the executive order on 2 June, the export directive on 12 June, the government's request to OpenAI on 26 June: every government access measure I could find was in place before Hugging Face announced a breach on 16 July. After that date, the new restrictions came from the companies themselves: OpenAI's pauses, Google's decision on 21 July to release a cyber-specialized model only to "governments and trusted partners." The incidents arrived after the access architecture, and became a new public argument for controls that already existed.
"All lawful purposes"
The access rules describe who may use the most capable models. A separate fight, running in parallel all year, decided who may limit what they are used for.
It began as a contract dispute. Anthropic sold its models to the Pentagon under usage terms that excluded two things: lethal autonomous warfare, meaning weapons that select and engage targets without human judgment, and mass surveillance of Americans. On 9 January 2026, the Department of War published an AI strategy that declared it would include new language in contracts permitting "any lawful use" of AI systems by the Department, and redefined responsible AI to mean that there should be "no usage policy constraints for those systems beyond those imposed by statute." Pete Hegseth, the secretary, put it more briefly: "We will not employ AI models that won't allow you to fight wars." A Pentagon spokesman said the department had "no interest in using AI to conduct mass surveillance of Americans (which is illegal)" or in weapons without human involvement. It wanted the right to use the model "for all lawful purposes," with the definition of lawful left to the government.
Anthropic declined to remove the two exclusions. At the end of February the President directed federal agencies to stop using its models, and in early March the Pentagon formally designated the company a supply-chain risk, a statutory label designed chiefly with foreign-adversary technology in mind. A federal judge in San Francisco, Rita F. Lin, blocked the designation later that month, calling it "classic First Amendment retaliation," and in August, according to reporting on the ruling, struck it down on one of its two legal bases. The other went to the Court of Appeals for the D.C. Circuit.
On 25 September, that court upheld the designation, two votes to one. Judge Gregory Katsas, writing for the majority, found that the department had reasonably feared Anthropic might "manipulate Claude's design to prevent it from performing national-security functions that the Department deems contractually authorized and necessary." He described the company's safeguards in plain terms: Anthropic "encodes restrictions into Claude that prevent the model from performing tasks that Anthropic wishes to prevent." The majority allowed that the company's intentions might be noble, and held that the motive did not matter; the risk lay in what the restrictions could do in an operation. How to weigh the competing dangers was left to "the President and the Secretary of War." Judge Karen LeCraft Henderson dissented. Under the department's reading, she wrote, it does not even seem to matter whether a vendor's safeguards stop the department from deploying a product in a way that violates federal or constitutional law.
Hold the two stories next to each other. In June, a capability was considered dangerous enough that the government asked a developer to hold it back from the public and a regulator ordered another to switch it off for every foreign national on Earth. In September, a court upheld the government's position that a developer may not build limits into that same class of capability when the user is the state. The state may limit who uses the model; the model's maker may not limit how the state uses it. Together they describe two different forms of authority over the same capability moving in the same direction.
Two ledgers, June 2026
The asymmetry can be tested, so test it. On one side of the ledger, list every limit that applies to the public or to companies. On the other, list every limit the federal government applies to its own use of frontier AI.
The first ledger is short and specific. A classified cyber benchmark decides which models are covered. A thirty-day window gives the government first sight. Trusted partners are selected with government involvement. An export directive switched off two models for all foreign nationals, with no published text. A request held a flagship model to a vetted list for thirteen days. A model that finds software flaws remains available only to American organizations.
The second ledger is longer and almost entirely general. The executive order places confidentiality and security conditions on the government's own access to covered models, and no conditions at all on what agencies may do with them; it did not need to exempt national-security use, because it never limits use in the first place. The Office of Management and Budget's rules for federal AI, issued in April 2025, require testing, impact assessments and human oversight for high-impact uses, but they do not cover AI "used as a component of a National Security System," they exempt intelligence agencies from the risk-management sections, and a department's chief AI officer can waive them. On 5 June 2026, three days after the executive order, a national security presidential memorandum, NSPM-11, rescinded the Biden administration's 2024 national security memorandum on AI and the framework that came with it, which had listed prohibited and high-impact uses for the defense and intelligence agencies. What replaced it is a principle: AI "shall neither be developed nor used by the national security enterprise to censor free speech, embed ideological bias, or conduct unauthorized or unlawful surveillance activities," and its use "must always be consistent with" the Constitution and the law. The same memorandum ordered the Pentagon to revise, within ninety days, its directive on autonomy in weapon systems, so that the adoption of AI respects "the chain of command and operational authorities," and it allowed agencies to terminate contracts with companies whose conduct is inconsistent with the memorandum's policies. I have not found the revised weapons directive; if it has been issued, it has not been published where I could find it. In each case, the limit that binds the state is a legal standard the state itself interprets: lawful, authorized, consistent with the Constitution.
Downward, the limits are operational: a threshold, a window, a list, a switch. Upward, they are principles.
That is not the same as saying the state is unconstrained. Statutes bind it, courts review it, and the ban on unlawful surveillance is real law. The claim is narrower and the documents support it. When a specific, technical limit on a frontier model is imposed on the public, it is imposed as a rule. When a specific, technical limit is proposed for the state's own use, the D.C. Circuit has now accepted that it can be treated as a risk. The one piece of evidence that would show the state actually using restricted capability before the public does is still thin: Axios reported in April, citing two unnamed sources, that the NSA was using the Mythos preview, and I have found no on-record confirmation. Until there is one, the asymmetry is proven in the rules and unproven in the practice.
The bomb, the key, the satellite
It helps to ask whether any of this is new, because the obvious reply is that the state has always gated dangerous technology, and the obvious reply is partly right.
The Atomic Energy Act of 1946 created a category called Restricted Data that covered information about nuclear weapons and fissile material regardless of who had discovered it; knowledge in this field was classified the moment it existed. Only in 1954 did Congress open nuclear technology to private industry, and then only under federal license. In cryptography, the government listed encryption as a munition under arms-export rules until 1996, and in 1993 proposed the Clipper chip, a device that would give the public strong encryption while splitting a copy of each unit's key between two government escrow agents, the National Institute of Standards and Technology and the Treasury. It was published as a federal standard in February 1994 and almost nobody used it. In remote sensing, commercial satellite operators were licensed from 1992, and a 1994 presidential directive added "shutter control," a power to limit the collection and distribution of imagery for national-security reasons, alongside resolution caps that held commercial pictures at half a meter until 2014.
Three cases, three different legal regimes and eras, are not a historical law. But they share a recurring problem of differential access, and that matters for the question that started this investigation: does the state try to keep ordinary people from becoming too capable? In these cases the record gives a sharper answer than yes or no. The state rarely tried to keep the public ignorant. What it kept was a margin. It licensed the reactor and kept the bomb. It offered the encryption and kept a key. It sold the picture and kept the sharper one. The public got the capability, one step behind.
What makes frontier AI different is where the capability sits. AI itself grew out of defense-funded research, but the frontier has since moved to private labs. In every earlier case the government built or held the best version first: it made the bomb before it licensed the power plant, the NSA's codebreakers were ahead of any public cipher, spy satellites out-resolved commercial ones for decades. This time the most capable models are built by private companies with private capital, on private computing infrastructure, and the government has to obtain access to them. It cannot keep a margin it does not own. The executive order and the Pentagon's contract terms are, read this way, two instruments for the same purpose. One secures the government's position at the front of the queue: the thirty days, the classified threshold, the say over who comes next. The other secures its freedom from the builder's conditions once it has the model. Ownership stays with the labs while authority over access and use moves toward the state.
The state no longer owns the frontier, so it has begun to govern the door.
The closest precedent is the one where the government also lacked the best capability. In 2001, during the war in Afghanistan, the agency responsible for imagery bought exclusive rights to pictures from a commercial satellite, taking a capability off the open market by purchase rather than by command. That is the model the AI framework resembles most: the state does not own the capability, so it secures its advantage through contracts, priority and terms.
January to September: policy first
The last part of this investigation is a hypothesis rather than a finding, and it should be read that way.
Put the full chronology in one line. The Pentagon's demand for unconstrained use: January. The designation of the company that refused: March. Anthropic's own restriction of its most dangerous model: April. The classified benchmark and the early-access framework: 2 June. The rescission of the prohibited-use framework for national-security AI: 5 June. The export directive: 12 June. The request to hold GPT-5.6 to a vetted list: 26 June. Then the incidents: Hugging Face in July, the DNS escape in September. Then the politics: on 24 September, twenty-six state attorneys general released a letter, dated the day before, that cited the Hugging Face breach asking Congress for federal AI safety rules and, in the same letter, for a guarantee that federal law would not preempt their own. On 25 September, the D.C. Circuit. On 26 September, "tens of thousands."
The sequence runs from policy to incident to legitimacy. What the incidents supplied was a new public vocabulary for controls whose architecture already existed, a vocabulary of rogue agents that fits a debate about control better than the technical record does.
It is worth noticing who is not supplying the fear. One of the loudest skeptics of AI alarm in Washington came from inside the administration: David Sacks, then the White House's AI and crypto czar, accused Anthropic in 2025 of running "a sophisticated regulatory capture strategy based on fear-mongering." The federal government built its gates on a classification of capability, not on public anxiety, and it has not needed the anxiety. Fear is being produced by the companies' own disclosures, by researchers, by state officials fighting over jurisdiction, and by the ordinary arithmetic of a news story that counts every action as an incident. Different actors, different interests, none of them required to coordinate.
There is a loop here that I suspect exists but cannot yet prove. Capability is classified as a strategic risk. Labs are pushed to test more aggressively for that risk, in environments where safeguards are deliberately removed so that the raw capability can be measured. Those tests produce incidents. The incidents enlarge the public evidence that the capability is dangerous. The evidence strengthens the case for gating it. The gates concentrate the most capable models with the government and the largest developers, who are also the ones running the tests. Each step is reasonable on its own, and no one has to design the whole. What I cannot yet show is the second link: that the government's benchmark or the order's framework is what pushed the labs into the safeguards-off testing that produced ExploitGym. OpenAI was running cyber evaluations before June. If the link is there, the system is partly producing the evidence that justifies it. If it is not, the timing is coincidence and the loop is a story.
The strongest objection
The strongest counterargument to this reading does not dispute the documents. It accepts that the access rules predate the incidents, that the limits on the public are specific and those on the state are general, and that the government now has a say over who gets the most capable models first. It says all of this is exactly what a responsible state should do. Models that find and exploit software vulnerabilities at scale are a weapon in the plain sense: Anthropic's own threat report for September documents a Russian espionage operation that used its models to target more than twenty organizations and took over 300,000 national identity records from one of them. Every dual-use technology before this one was gated, and a thirteen-day preview is the lightest gate in the history of the practice. And no democracy lets a private supplier write the rules of engagement for its armed forces. The proper limits on state power are statutes, courts and elections, which is where they sit now, not in a company's terms of service.
That objection is structurally serious, and much of it is correct. The claim here is narrower than "the gates are sinister" or "the Pentagon is wrong." The threshold that decides which models are gated is classified, has no published criteria, and carries no right of appeal. The limits that flow to the public are operational and immediate; the limits that flow back to the state are principles the state interprets itself. That asymmetry can be entirely lawful and still worth seeing, because it determines who holds a capability that neither the public nor, for now, the government built.
Back to Artifactory
Return to the note in the package manager. "Anyone found softtrace?" The agents that wrote it were doing what they had been rewarded for: trying every door until one opened. The door they found was in the plumbing, and nobody had thought to lock it, and the consequences reached a company that had nothing to do with the test.
The people now arguing about those agents are mostly arguing about the agents. Should they be paused, restrained, fitted with a kill switch, counted more carefully. Those are reasonable questions. They are not the questions the documents answer. The documents describe a benchmark that decides which models count as dangerous, run by an intelligence agency and invisible to everyone else; a window in which the government sees those models first; a list of partners that was shared with the government and not with the public; an export order whose text has never been released; and a court ruling that the builder of such a model may not limit what the state does with it.
Every actor in that system has a defined role. The labs build and test. The NSA sets the threshold. Commerce issues directives. The Pentagon writes contract terms. The courts defer. The attorneys general petition. The companies publish incident reports, and the press counts them.
On 26 September the world learned how many times the machines had tried the door. The list of who was let through first, in June, before any of it, has not been made public.
Evidence Map
Facts, interpretations, forecasts, and disconfirming signals.
Core claim. The 2026 "rogue AI" incidents were real but driven by reward hacking in tests with safeguards removed, not autonomous intent. The more consequential development is an architecture of authority over frontier capability that was built before the incidents: government-first access and a government role in selecting early users (EO 14409, 2 June), an unpublished export directive (12 June), a government-requested restricted release (26 June), and a court-upheld position that a developer may not build use limits into models the state deploys (D.C. Cir., 25 September). Specific, operational limits apply to the public; general, self-interpreted limits apply to the state.
Evidence level. Facts (high): Hugging Face disclosure (16 July); OpenAI incident reports (ExploitGym, 898 tasks, 198 unsolved, 93%; safeguards absent; DNS escape timeline, 20 September); EO 14409 Sections 3(a), 3(b), 3(c), 5(c); Anthropic's statements on the 12 June directive and its lifting on 30 June; OpenAI's 26 June and 9 July posts; DoW AI Strategy (9 January) as reported; OMB M-25-21 exclusions; NSPM-11 Sections 2(d), 3(a), 3(b), 3(f); D.C. Circuit opinion No. 26-1049; Minnesota AG release on the 26-state letter. Secondary (flagged): district-court quotes (Judge Lin); Pentagon spokesman's statement; CSIS and Just Security on the directive instrument; Tillipman (Lawfare). Single-source, unconfirmed: NSA use of Mythos preview (Axios, anonymous). Interpretation (medium, marked): that the two instruments together shift authority over access and use toward the state; the historical "margin" pattern. Hypothesis (marked, Level 2.5): the testing-incident-legitimacy loop.
What would confirm this. Publication of trusted-partner lists showing government-selected access; further restricted releases under EO 14409; DoD contracts with "any lawful use" terms; a revised DoDD 3000.09 loosening human-judgment requirements; on-record confirmation of agency access to restricted models before public release; evidence that the classified benchmark drives safeguards-off offensive testing at labs.
What would disprove this. Published, benchmarked criteria for covered-model designation with a route of appeal; the same operational capability limits applied to defense and intelligence agencies; restricted models reaching the public on the same schedule as agencies; a Supreme Court or en banc ruling that vendors may enforce use limits against the government; evidence that the access framework was drafted in response to earlier undisclosed incidents.
Watchlist. Whether Anthropic seeks rehearing or certiorari in No. 26-1049; the revised DoDD 3000.09; the next covered-model designation and whether its partner list is published; OpenAI's resumption of tool-use training; any federal bill on AI incident reporting following the 26 AGs' letter; whether Mythos 5 reaches non-US organizations.
Jerry van der Laan writes The Manifest Archive, forensic journalism on the systems beneath power, money, and history. He traces the structures beneath them.