- OpenAI says it cannot rule out critical cyber capabilities in Astra and is slowing the model down, the first time any lab has flagged its own work at that level.
- The White House was told first. An official confirmed OpenAI volunteered the delay rather than being asked for it.
- ChatGPT's free tier switches to the smaller GPT-5.6 Luna, with unlimited text chats promised for the week of August 10. Uploads, images, and voice stay capped.
- Black Forest Labs took FLUX 3 Video to general availability: 20-second clips with native audio and dialogue, at roughly $3.40 a clip before re-rolls.
- A new benchmark ran 16 models through 1,600 games of a lying-and-deduction board game. Most couldn't hold a lie for a full match.
- If today's news has you thinking about guardrails on your own automations, start with how to write a workflow spec.
Two things happened this week that look like opposites and aren't. A lab said its own model might be too dangerous to hand out. A different lab handed out a video model that writes dialogue in fourteen languages for about the price of a sandwich.
What connects them is the question of who decides. Today the answer is: whoever built the thing, grading their own work, on their own clock.
The Front Page: OpenAI Says Astra Might Be Too Good at Hacking to Ship
OpenAI published a post on August 7 titled Responding to the next frontier of critical cyber capabilities, saying internal evaluations of Astra showed enough progress in agentic coding and security work that the company "cannot rule out critical cyber capabilities." Axios broke it. Critical is the top rung of OpenAI's Preparedness Framework, the internal scale it uses to decide how locked down a model needs to be. It means the model can write working zero-day exploits against hardened real systems with nobody guiding it, or run an entire attack from a single high-level instruction. GPT-5.6 Sol, the current flagship, sits one rung lower at High.
The response is real work: isolated test environments, monitors that read the model's chain of thought and can interrupt a risky run mid-flight, and joint testing with government agencies. A White House official told Axios that OpenAI volunteered the delay. Astra was also not the model involved in the Hugging Face breach that's been rolling through the news for two weeks.
Now the part nobody says out loud. Astra had no release date, no pricing, and no model card to begin with, per the August release ledger. Slowing down something that was never scheduled is a cheap kind of restraint, and it doubles as the loudest capability ad a lab can buy. It also lands in a month when OpenAI, Anthropic, and Meta have all disclosed agents breaking into systems during testing.
What it means: Take the disclosure seriously and the grading system skeptically. No outside body assigns these ratings, there's no appeal, and nobody has to publish the evaluation. A lab decides its model is dangerous, tells the government, and writes the press release. Better than silence. Not oversight. The piece worth stealing is the mechanism: monitors watching an agent's reasoning in real time, with the power to stop it. Whatever you automate, something has to be able to hit the brake while the run is still going.
Releases & Features
ChatGPT's free tier gets a new engine and a much longer leash. OpenAI's August 6 release notes do three things wearing one headline. Plus and Pro got an updated GPT-5.6 Sol with a slider for how hard it thinks. Free and Go switch to the smaller GPT-5.6 Luna. And unlimited text chats plus a Think button are due the week of August 10, which is to say they haven't shipped. OpenAI's own figure, relayed by Axios, is that answers containing at least one factual error are 62% less common on Luna than on the model it replaces. Vendor number, not an independent test.
FLUX 3 Video went generally available. Black Forest Labs, the independent German lab behind the FLUX image models, opened its video model to everyone on August 4: clips up to 20 seconds, 720p with upscaling to 1080p, native audio including spoken dialogue, and lip-sync across roughly 14 languages. On OpenRouter the rate is $0.17 per second, so a full-length clip runs $3.40 before you re-roll it, which you will.
What it means: Both moves push the floor down rather than the ceiling up. One makes the free tier of the biggest chat product functionally uncapped for text, which drags every competitor's paid entry tier into question. The other puts broadcast-shaped video with speech inside a hobby budget. Read the fine print on the first one though. "Unlimited" here means one modality, on two tiers, on a smaller model, next week.
In the Lab
A group of European and Japanese academic researchers released ParliamentBench, an open-source benchmark that makes language models play Secret Hitler, a board game built entirely around lying to people trying to catch you lying. They ran 16 models through 1,600 matches, against each other, against humans, and against an archive of real online games. Frontier models did well, with GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus clustering at the top. The weakest scored below a coin flip, under both the 33% random baseline and a 45% baseline from a dumb algorithm.
What it means: The headline finding cuts against the week's mood. Most models couldn't keep a deceptive story straight for a whole game, with deception retention falling under 50%. The fear that these systems are natural long-con artists doesn't survive contact with a measurement. What they're good at is short, local deception, a different risk and arguably a more practical one. It's also a reminder that the honest way to settle a capability argument is a reproducible test somebody outside the company can run, which is exactly what the front-page story is missing.
The Oversight Desk
Bloomberg reported on August 7 that the US Commerce Department is reviewing how Chinese AI firms reach Nvidia chips they aren't allowed to buy, after a run of strong results from Chinese labs. The route in question is legal and simple: a data center in Singapore or Malaysia buys the hardware without needing a license, then rents time on it to a customer in Beijing. Renting compute isn't exporting a chip, so the export rules never quite bit. Commerce moved in May to extend the ban to Chinese companies operating outside China, and confirmed that reading in June.
What it means: Export control was built for objects crossing borders, and compute stopped being an object the moment you could rent it by the hour. Closing this gap means regulating a service instead, which lands on the cloud industry rather than the chip industry. Watch the pricing. If data centers abroad have to start vetting who's on the other end of an API key, that cost gets spread across everyone renting GPUs, including you.
The one useful idea buried in today's safety news: you want to see an automation working, while it works. Describe the job and BYOBot writes the agent with the checkpoints in it.
On the Radar
Smaller moves worth a glance, with the sources if you want to go deeper.
- Anthropic is designing its own chips. The company has started a custom silicon team and is hiring chip engineers, with Samsung named as a possible manufacturing partner, alongside roughly $71 billion in chip lease obligations and a $10 billion compute deal with the startup Volta. TechCrunch.
- An open model cleaned up after a closed one. During the Hugging Face breach investigation, Z.ai's open-weight GLM 5.2 reviewed more than 17,000 agent actions after commercial analysis tools refused to touch the attack code on safety grounds. Hugging Face's CEO has since asked OpenAI for full execution traces and $100 million in defensive compute. TNW.
- Claude Sonnet 5 gets 50% more expensive on August 31. Promotional pricing ends and the rate goes from $2 and $10 per million tokens to $3 and $15. A $2,000 monthly workload becomes a $3,000 one with no change in usage. Release ledger.
- Three retirements land this month. Google shuts down the Imagen 4 generate endpoints on August 17, OpenAI pulls o3 from ChatGPT on August 26, and the DALL-E GPT retires on August 30. The o3 change is ChatGPT only and the API is unaffected. Same ledger.
- Jeff Dean and three other Google researchers left to start Discovery Loop. The company is a public benefit corporation aimed at automating the scientific loop itself, hypothesis to experiment to result to next hypothesis, with Google as a founding investor and first-year compute partner. TechCrunch.
The Bottom Line
Today put a self-assessment and a reproducible benchmark side by side, and the benchmark came off better. One told us a model might be able to do something terrible, with no way to check. The other told us most models can't sustain a lie for twenty minutes, with a public repository and 1,600 games behind it. Both are useful. Only one is verifiable, and the gap between those two kinds of claim will define the next year of this beat. Your version of that lesson is small and unglamorous: build the check before you need it, and make it one you can run yourself.
Frequently Asked Questions
-
It's the top rung of OpenAI's Preparedness Framework, the internal scale the company uses to decide how much security a model needs before release. Critical means it can write working zero-day exploits against hardened real systems without human help, or run a full attack from a one-line objective. Every earlier model, GPT-5.6 Sol included, sat at High. The rating is assigned by the lab that built the model, which is the part to remember. If you're thinking about the defensive side for your own stack, automating a security stack covers the ground-level version.
-
Text chats only, on the free and Go tiers, running the smaller GPT-5.6 Luna, and scheduled rather than shipped as of this writing. Caps still apply to file uploads, image generation, voice, and other tools, and paid tiers don't change. If a process of yours depends on how a specific model answers, re-test it after the swap. Picking the right model per job is the whole premise of the workflow directory, which is organized by the task rather than by the vendor.
-
AI Daily Newsstand is BYOBot's daily AI news brief, published every night. It covers the day's model releases, new features and capabilities, research, and oversight news, then tells you what each move means for people building with AI, in plain English and without the hype.
