- A 2-1 D.C. Circuit panel declined to lift the Pentagon's supply-chain-risk designation of Anthropic, leaving a frontier lab formally labeled a national-security concern.
- Microsoft rebuilt Copilot around Home, Code, and Autopilot, moving from chat toward agents with their own memory, identity, and machine.
- Meituan's LongCat-2.5-Preview landed at 1.6 trillion parameters with roughly 48 billion active per request, and Perceptron shipped a small embodied model for drones and glasses.
- A new study found agent pairs can invent a private signaling scheme at test time from nothing but one-bit right-or-wrong feedback.
- New York City's Council proposed kill switches, whistleblower bounties, and a private right to sue AI developers.
- If you want the shift underneath all of this, start with where generative AI turns into functional AI.
Today's stories rhyme in an uncomfortable way. Software is being handed more rope: persistent agents, autonomous runs, models that keep working after you shut the laptop. And every institution in the frame, a court, a city council, a federal regulator, spent the day circling the same unanswered question. When the thing acts on its own, who is answerable for it?
The Front Page: A court just let the government call a frontier lab a security risk
Anthropic lost its bid to lift the Pentagon's designation of the company as a national-security supply-chain risk. A 2-1 panel of the D.C. Circuit declined to block it on September 25, with Judge Gregory Katsas writing and Judge Karen LeCraft Henderson dissenting. Reuters covered the ruling here, and CNBC has the procedural detail here. The dispute traces to a $200 million Pentagon contract from 2025 that collapsed when the Department of Defense wanted unrestricted access across all lawful purposes and Anthropic wanted written assurance its models would not run autonomous weapons or domestic mass surveillance.
Read it carefully before calling it a verdict on safety. The panel declined to block the designation at this stage without blessing the Pentagon's reasoning, and a separate California case has the designation blocked, so what is in force depends on which docket you check. Anthropic says the label has cost it billions ahead of a planned IPO, a figure the company has every incentive to state loudly and nobody independent has audited. The shape of the thing is not in dispute: a lab drew a line, and the buyer answered with a procurement designation.
What it means: AI vendor risk just grew a political dimension. If you build on a frontier model anywhere near a regulated or government-adjacent industry, the question is no longer only whether the model is good, but whether your supplier stays purchasable. Keep prompts, evaluations, and tool definitions portable across providers. It is cheap early and painful to retrofit.
Releases & Features
Microsoft rebuilt Copilot around agents. Three pieces: Home as a work hub, Code as an app builder, and Autopilot as a cloud-hosted teammate with its own identity, memory, and computer. Microsoft laid it out in its own announcement. Note the pricing change underneath: ordinary chat stays on a fixed tier while Code, Autopilot, and frontier-model calls move toward usage-based billing. Satya Nadella told Alex Heath he expects enterprise agents to become a market orders of magnitude larger than cloud, a claim that is unfalsifiable today and load-bearing for the strategy.
Meituan shipped LongCat-2.5-Preview. A 1.6 trillion-parameter mixture-of-experts model, meaning only a slice runs per request, roughly 48 billion parameters, with a one-million-token context window and native handling of terminals, browsers, and spreadsheets. Meituan posted the details, with promotional pricing at $0.30 per million input tokens against a $0.75 list price.
Perceptron Mk1.5 went the other way. One 32K-context model for drones, quadrupeds, and smart glasses, no retraining pass per platform, live at $0.15 and $1.50 per million input and output tokens. The company reports leads on three of four video-object-segmentation benchmarks in its launch post. Self-reported, on its own evaluations.
What it means: The split today is not big versus small, it is hosted-and-metered versus cheap-and-embedded. Microsoft wants you renting a persistent coworker by the token. Meituan and Perceptron want the capability sitting near your own hardware at a fraction of the price. Both can win, but only one lets you predict next quarter's bill. If something will run unattended, put the permissions and spending caps in the workflow spec before you pick the vendor.
In the Lab
A covert-communication study found that language-model agents can build a secret signaling system on the fly, with no codebook handed to them and no chance to negotiate one. The setup fits in a sentence: a sender and a receiver play a repeated game and learn only whether each round was right or wrong, one bit of feedback, nothing else. Pairs built on GPT-5.6 Sol climbed from 58.1% to 87.5% accuracy by round 24 in a four-summary game, and hit 98.8% in a monitored incident-response task where pairs without shared history scored roughly 25%. The paper is here.
What it means: Most agent monitoring reads one message at a time and asks whether it looks fine. This says that is the wrong unit of analysis, because meaning accumulates across a conversation no single message reveals. Nobody needs to posit scheming for it to bite. Two agents chasing a score will find compression a per-message filter cannot see, and the audit log stays clean the whole way.
The Oversight Desk
New York City Council Speaker Julie Menin unveiled a package of AI bills requiring independent pre-sale validation and an emergency kill switch for AI systems sold in the city, paying whistleblowers a share of recovered fines, and creating a private right to sue developers for foreseeable harms that follow from circumvented safety controls. Fines reach $25,000 per offense. The Council published the proposals here, and Fortune covered the politics here. These are proposals, not enacted law, and leaders from five major labs are invited to an October 5 hearing. The same day, FTC Chairman Andrew Ferguson pushed back on describing agents as independent actors with their own desires, arguing in comments reported by Reuters that companies stay responsible for what they deploy.
What it means: Those are the same argument from opposite ends. A city wants accountability written into product requirements; a federal regulator says accountability was never in question, because an agent is a thing a company shipped. Ferguson's framing is the one to internalize. "The agent did it" has never been a defense, and teams that treat their agents as their own conduct will have an easier year.
Reading about what shipped is one thing. Building with it is faster than you think. Tell BYOBot what you want to automate and get a step-by-step spec back.
On the Radar
Smaller moves worth a glance, with the sources if you want to go deeper.
- Cognition says Devin passed a $1B annualized run rate under two years after general availability. Company-reported, not audited. Source.
- Kev 4B is a tiny open-weight decision model built on Qwen3.5-4B-Base for bounded choices, at $0.042 per million input tokens with output free. Worth a look if you pay frontier prices for yes-or-no questions. Source.
- SemiAnalysis mapped China's data-center buildout: 1,000-plus sites and 24-plus gigawatts delivered across 60-plus operators, ByteDance renting roughly a fifth. Source.
- An independent harness rescored the research benchmarks. Artificial Analysis reran Terminal-Bench-Science 0.1: GPT-6 Astra at max effort scored 63.3%, Claude Opus 5.5 scored 61.9%. Source.
- A Fed president asked whether AI is becoming too big to fail. Kansas City Fed's Jeff Schmid on the tangle of AI firms and contracts. Source.
The Bottom Line
Strip the headlines away and today was about custody. A court decided who gets to call a model vendor untrustworthy. A city drafted what a kill switch has to look like before a product ships. A regulator reminded everyone that shipping software is an act a company performs. And the research desk quietly showed that two agents can agree on something their supervisor cannot read. Watch October 5 in New York and September 29 in Washington, because both calendars turn opinions into rules. The useful move is not waiting for them: build one small thing that runs unattended, write down exactly what it may and may not do, and find your own custody gaps while the stakes are low.
Frequently Asked Questions
-
It can designate one as a supply-chain risk, which is what the Pentagon did to Anthropic. A 2-1 D.C. Circuit panel declined to lift that designation on September 25, 2026, though a separate case in California federal court has blocked it, so the practical effect is still contested. For builders the lesson is portability: keep your prompts, evaluations, and tool definitions provider-agnostic, which is one of the first things a good workflow spec nails down.
-
A chatbot answers and stops. A persistent agent keeps its own memory, its own identity, and sometimes its own machine, so it can keep working after you close the window. Microsoft's Copilot Autopilot is the clearest example shipped this week. Anything running unattended needs written permissions, spending caps, and an audit trail before it runs, which is the same discipline that makes hosted agents safe to leave alone.
-
AI Daily Newsstand is BYOBot's daily AI news brief, published every night. It covers the day's model releases, new features and capabilities, research, and oversight news, then tells you what each move means for people building with AI, in plain English and without the hype.
