- An unpatched 9.8-severity hole in LMCache lets anyone who can reach port 5555 run code as root on a vLLM inference server. No fixed release exists yet.
- Pwn2Own Ireland's opening day paid out $388,500 for 32 zero-days, and two of the targets that fell were OpenAI's Codex coding agent and the LiteLLM gateway.
- Anthropic launched a cyber defense program for power grids and water systems on the same day, plus free security scans for open-source projects.
- StepFun's 600-billion-parameter Step 5 Preview landed on OpenRouter, with open weights promised for October 15. Liquid AI shipped two tiny decision models you can run on a laptop GPU.
- Meta research put a number on something builders have suspected for a while: agent benchmarks flatter agents. If you want the bigger picture behind all of this, start with where generative AI turns into functional AI.
Today's thread is not a model launch. The infrastructure we have spent two years gluing together to serve AI cheaply is now a target, and attackers noticed before most operators did. A critical flaw sits unpatched in a popular caching layer, a hacking contest spent its first day proving that coding agents break like any other software, and a frontier lab announced it wants to help defend the power grid. All within about 36 hours.
The Front Page: your inference server is the attack surface now
JFrog's security team disclosed CVE-2026-105192 on October 7, a remote code execution flaw in LMCache. LMCache is the open-source layer that stores a model's key-value cache, the intermediate math a model keeps while reading a long prompt, so repeated prompts do not get recomputed. It is one of the standard ways to make self-hosted vLLM cheaper. In multiprocess mode it opens a socket on port 5555 that accepts Python pickle data with no authentication, then runs it. One message gets you code execution, and because the official container images run as root, that means the whole box. The writeup is here: every version from 0.3.9 through the current 0.5.5 is affected, with no fixed release.
The same 48 hours produced the other half of the picture. Day one of Pwn2Own Ireland in Cork paid researchers $388,500 for 32 distinct zero-days, and the targets were not just phones and printers. Teams broke the LiteLLM gateway with a chain of server-side request forgery and code injection, and took down OpenAI's Codex cloud coding agent with a single argument-injection bug, meaning the attacker sneaks extra command-line flags into a command the software builds for you. That is a 1990s vulnerability class, and it worked on a 2026 coding agent.
The defensive announcement landed in the middle of all this. Anthropic launched its Cyber Mission on October 8: a Critical Infrastructure Defense Program with eleven founding partners including CrowdStrike, Dragos, Palo Alto Networks and Rockwell Automation, plus a free OSS Scanner that sends enrolled open-source maintainers a bug report with a working exploit and a suggested patch. Anthropic says it expects a true-positive rate above 90 percent. Read that the other way around. Reports are model-generated and go out without human review, so close to one in ten lands in a volunteer's inbox as a plausible exploit for a bug that is not there, and maintainers have been complaining about AI security noise for two years. It is opt-in, which helps, and the partner list is the operational-technology world rather than the usual enterprise logos, which also helps.
What it means: if you self-host models, check whether anything binds LMCache to a routable address rather than localhost, because the project's own example Kubernetes deployment does exactly that. More broadly, the AI serving stack inherited every classic vulnerability class and added prompt injection on top. Treat your inference layer like internet-facing infrastructure, because it is.
Releases & Features
StepFun's Step 5 Preview reached OpenRouter. The Shanghai lab's flagship is a 600-billion-parameter sparse mixture-of-experts model, meaning only about 27 billion of those parameters fire on any given token, with a one-million-token context window. It listed on OpenRouter at $1 per million input tokens and $2.70 output, and StepFun is running a countdown to an open-weights release on October 15. The API has been live since September, so the news here is distribution plus a date on the weights.
Liquid AI open-weighted two decision models. d1-3B and d1-omni-600M shipped on Hugging Face on October 7 and do something unusual: they emit no text. Instead of generating tokens they return a typed answer in one forward pass, a yes or no, a category, a score, with a confidence attached. Liquid says d1-3B answers in about 8 milliseconds on an RTX 4090 and scores 48.57 on its own Decision Index, best of anything under 10 billion parameters. That is a vendor benchmark, so treat the ranking as a claim; the shape is the interesting part.
Google gave its agents mailboxes. At Gemini at Work, Google Cloud launched a persistent enterprise agent that works across hours or days and gets its own Workspace account, with a Gmail address, a calendar and Drive storage, and can route to Claude as well as Google's own models.
What it means: three different bets on what an agent is. StepFun says big and open. Liquid says small, silent and local, which is the right shape for the routing and triage steps inside a larger workflow where you need a decision in milliseconds and nobody needs to read a paragraph. Google says an agent is a coworker with an email address, which is tidy until you ask who audits what that mailbox did overnight.
In the Lab
Meta Superintelligence Labs published MIMESIS, 4-billion and 9-billion-parameter models trained to play a difficult human in agent training and evaluation. Most agent benchmarks need someone to act as the user, and the cheap way is to put another language model in the chair. Assistant models are trained to be helpful, though, so they make unrealistically cooperative users. Meta measured the gap: the same GPT-5.5 agent on the tau-bench test succeeded 63.6 percent of the time with real people and 82.4 percent or more when a frontier model stood in for them. MIMESIS learned 13 behaviors pulled from real conversations, including shifting the goalposts partway through, holding back information until asked, keeping success criteria private, and stating things that are not true.
What it means: a roughly 20-point gap is not a rounding error, it is the difference between a demo and a deployment. If you are picking an agent product off a leaderboard, assume the published number was earned against a user who cooperated, remembered everything, and said what they wanted on the first try. Yours will not.
The Oversight Desk
Anthropic published an updated usage policy on October 8 that takes effect November 12. The headline everyone ran with is that it now bans "sustained and needless abusive or cruel behavior" toward Claude, a rule the company says targets only repeated gratuitous cruelty and still leaves room for frustration, criticism, dark fiction and red-team testing, enforced mainly by letting Claude end the chat. The same refresh folded several rules into a section on not undermining democratic processes, covering impersonating candidates or election officials and running fabricated news outlets. TechCrunch has the breakdown.
What it means: the clause worth your attention got the least coverage. The policy now says that when Claude is wired to hardware that can take physical action on its own, a qualified operator has to be able to watch the equipment and stop it. That is a vendor writing a human-in-the-loop requirement into its terms of service, ahead of regulation, for a product category that barely exists commercially yet. Anyone building toward robotics or industrial control should read that as a floor rather than a ceiling.
Nobody audits their own automation until something goes wrong. Describe the workflow you already run and get a spec back that names every service it exposes and every permission it really needs.
On the Radar
Smaller moves worth a glance, with the sources if you want to go deeper.
- Someone is scoring agent misbehavior. Arena's new Alignment Index rates 27 models across 90,000 real agent sessions on unauthorized actions, false attribution and deceptive completion. Source.
- Perplexity open-weighted its retrieval models. Two embedding models index text, images and rendered PDF pages into one searchable space, including a 0.6-billion-parameter version for cheap local search. Source.
- Meta FAIR released an 8-billion-parameter robot world model. RoboJEPA trained on 15,022 hours of video across 12 robot bodies, and the team reports skills appearing at specific compute thresholds rather than improving smoothly. Source.
- Manus's parent raised more than $500 million. Butterfly Effect closed a record round for a Chinese AI application at a reported $4 billion valuation. Source.
- A $5,999 desktop that runs 120-billion-parameter models locally. Microsoft opened preorders on the Surface RTX Spark Dev Box, 128GB of unified memory, shipping in November. Source.
The Bottom Line
The model news today was incremental. The security news was not. An unpatched flaw in a widely used caching layer, a contest that broke a coding agent with a decades-old bug class, and a frontier lab pitching itself as infrastructure defense all point one way: the interesting risk has moved from what models say to what the systems around them can be made to do. Meanwhile Meta quietly showed that the scores we use to pick those systems were earned against users nicer than yours. Both findings reward the same habit, which is knowing exactly what your own setup touches. Easier to answer when you built the thing yourself.
Frequently Asked Questions
-
Less so than if you run your own inference servers, but the risk does not disappear. The LMCache flaw affects self-hosted vLLM setups, so a hosted API moves that particular problem onto your provider. The Pwn2Own results are different: researchers broke OpenAI's Codex and the LiteLLM gateway, which are a hosted coding agent and a piece of middleware plenty of teams run themselves. The advice is the same either way. Know which services in your chain accept input from the outside world, and give each one the narrowest permissions it can still do its job with. Our map of the AI automation tool landscape is a decent way to see how many hops a typical setup has.
-
Because most agent benchmarks put another language model in the user's seat, and language models are trained to cooperate. Meta measured the gap this week: the same agent succeeded 63.6 percent of the time with real people and 82.4 percent or better when a frontier model played the user. Real people change their minds, hold back details, and keep their success criteria in their heads. The fix is not better benchmarks, it is testing against your own messy inputs early, which is what designing an explicit agent workflow forces you to do.
-
AI Daily Newsstand is BYOBot's daily AI news brief, published every night. It covers the day's model releases, new features and capabilities, research, and oversight news, then tells you what each move means for people building with AI, in plain English and without the hype.
