- Apple trained a China-only large language model with Alibaba's help, a break from leaning on Chinese partners' models to power Apple Intelligence there.
- DeepSeek shipped V4-Pro and raised some API prices by as much as 1,100%, with the new rates live on August 16.
- Google's Gemini 3.7 Flash landed at 75 cents per million input tokens, half the old Flash price, but only until December 31.
- Z.ai's GLM-5.3 posted 84.5% on a vulnerability-finding benchmark and the company is holding the open weights back for roughly two weeks.
- Writer says rebuilding the software around its agents, not the model inside them, cut cost per task by 41%. That only works if you can describe the job, which is what a written workflow gives you.
There isn't one AI market anymore. Friday's news was four answers to the same two questions: what does this cost, and who's allowed to sell it here. A phone company built a model for one country. One lab halved its prices with an expiry date attached. Another raised its elevenfold and said so out loud. A fourth shipped a model good enough at finding software holes that it decided to sit on the download link.
The Front Page: Apple Builds a Model for One Country
Apple has trained its own large language model for the Chinese market, with support from Alibaba, according to reporting published on August 14. The Verge and MacRumors both have the details, and the Reuters version ran here. Until now Apple planned to power Apple Intelligence in China with a partner's model, because the American ones it uses elsewhere, ChatGPT included, aren't available on the mainland. It has already registered an on-device generative AI service with Chinese regulators, and the rollout is expected within months.
The line worth staring at is that Apple is described as the first foreign company approved to offer its own AI model in China. That approval is the asset here, not the model. Nobody has said how much of the training Alibaba did, what data went in, or who can see the weights, and "with Alibaba's support" is carrying an enormous amount in a single sentence. Apple didn't do this for benchmark glory either. Huawei has spent two years making on-device AI the reason to buy a Chinese phone, and Apple can't sell a feature it isn't allowed to ship.
What it means: Compliance is turning into architecture. The comfortable idea that you build one product on one model and ship it everywhere is quietly ending, and the biggest name in consumer hardware just spent a research team to prove it. You won't train a regional model. You will, sooner than you think, need to run the same workflow on a different model in a different market, which is an argument for keeping the description of the work separate from whatever executes it.
Releases & Features
Google shipped Gemini 3.7 Flash. It arrived on August 13, three weeks after 3.6 Flash and still ahead of the delayed 3.5 Pro, at an introductory 75 cents per million input tokens and $3.75 per million output, roughly half the old Flash rate. Reuters has the launch and SiliconANGLE the developer angle. Google says coding scores jumped from 34.4% to 43.6% on FrontierCode 1.1 Main, which is Google's number on Google's chosen test. Read the small print on the price: the introductory rate runs to December 31, and on January 1 it doubles to $1.50 and $7.50.
DeepSeek went the other way. The Chinese lab released V4-Pro and raised API prices by as much as 1,100%, with the new rates taking effect at 16:00 UTC on August 16. Caixin broke it and Quartz has the breakdown. Peak-hour output on V4-Pro goes to $3.96 per million tokens from $0.87, cache-miss input to $1.32 from $0.435, and a new peak and off-peak split is meant to push work into the quiet hours. DeepSeek says it's allocating resources more sensibly. It's also still cheaper than most frontier alternatives, which is rather the point.
What it means: Two price moves in opposite directions inside 48 hours, from the company that made cheap models famous and the company that can most afford to give them away. Anyone budgeting on the assumption that inference gets cheaper forever just got two counterexamples. Google's discount has a printed expiry, DeepSeek's increase has a start time, and neither is a reason to panic. Both are a reason to know which steps of your work need which model.
In the Lab: The Harness, Not the Model
The most useful research of the day came from an unglamorous place. Writer, an enterprise AI company, published work arguing that the "harness," meaning the code that decides how an agent fetches context, calls tools, retries failures and carries conversation history, drives cost about as much as which model you picked. Rebuilding that layer alone, the company says, cut blended cost per task by 41% and finished tasks 44% faster across every model it tested, including Anthropic's and OpenAI's, without quality dropping. Writer's post has the claims, TechCrunch and VentureBeat the context.
What it means: This is a vendor's research about a vendor's product, so discount it. Discount it by half and it still says something awkward: much of what you're paying for isn't intelligence, it's a program calling a model badly. That's the same shape as Nvidia's routing numbers earlier this week, reached from a different direction, and it points somewhere cheerful. Your AI bill is a software problem, and software problems are fixable by people who don't run a frontier lab.
The Oversight Desk: An Open Model That Can Find Holes
Z.ai launched GLM-5.3 on August 14 and reported 84.5% on CyberGym, a benchmark for how well a model finds vulnerabilities in real software. That's a shade above the 83.8% and 83.6% it cites for Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol, though on writing working exploits the gap runs the other way, 54.4% against 78%. Decrypt covered the launch and Unite.AI the security side. Here's the part that matters: Z.ai bills itself as the open alternative, and it isn't releasing the weights yet. It says they go public around August 28, after safety testing, with a disclosure ledger tracking what testers find.
What it means: Finding a flaw and exploiting a flaw are the same skill pointed at opposite ends of the same problem, which is why this beat exists. Two weeks of extra testing is an honest answer, and it's unenforceable the moment the download goes live. Give the staging credit as a norm worth copying, then plan for the world after it. From late August, assume a competent attacker has a free model that reads your code as carefully as you do. The response isn't fear, it's making patching a scheduled job instead of an emergency.
If your AI bill is really a software problem, the fix starts with writing down what the software is meant to do. Describe the job you keep repeating and BYOBot will hand back the steps, so you can see where the money goes.
On the Radar
Smaller moves worth a glance, with the sources if you want to go deeper.
- OpenAI is previewing an Ultrafast tier. It runs GPT-5.6 Sol at up to 14 times standard speed and up to 750 output tokens per second, on Cerebras hardware rather than GPUs, in a limited preview with no price announced. OpenAI and Cerebras.
- Big Tech's AI purchase commitments are nearing $1.5 trillion. A Financial Times analysis puts Alphabet, Microsoft, Amazon, Nvidia, Oracle and Meta close to that figure on chips, data centers and energy, on top of a comparable pile of leases. Summary.
- France's tax agency was breached. The Finance Ministry confirmed Thursday night that taxpayer data was extracted, putting it at 678,000 users while a breach tracker estimates closer to 700,000. Report.
- SMIC is raising chip prices. China's largest foundry ran at 93.7% utilization last quarter, topped $3 billion in revenue for the first time and more than tripled profit to $479.2 million. Details.
The Bottom Line
Friday was four organizations pricing the same product and landing on four different numbers. Apple paid in engineering for permission to sell in one country. Google is paying in margin to win developers before January. DeepSeek decided its users will pay more, starting Sunday. Z.ai is paying two weeks of delay to look like a grown-up. None of that is a story about intelligence, and January 1 is the date to keep in your calendar. The habit that survives all of it is dull and cheap: know your work well enough to write it down, and which model runs it becomes something you can change your mind about.
Frequently Asked Questions
-
Both at once, depending on whose model you use. Google cut Gemini 3.7 Flash to 75 cents per million input tokens, but only through December 31, after which it doubles. DeepSeek raised some rates by as much as 1,100% starting August 16 and added peak and off-peak billing. Treat any quote you get today as temporary, and keep an eye on where the running costs land: our guide to hosted AI agents walks through what you end up paying for once something runs on a schedule.
-
Worry isn't the useful response, planning is. A model that spots a flaw in code helps the person patching it exactly as much as the person exploiting it, and GLM-5.3 puts that ability within reach of anyone who can download a file. For a small team the practical move is to point the same class of tool at your own code first and to schedule patching rather than react to it, which is the boring core of a security stack you can automate.
-
AI Daily Newsstand is BYOBot's daily AI news brief, published every night. It covers the day's model releases, new features and capabilities, research, and oversight news, then tells you what each move means for people building with AI, in plain English and without the hype.
