Today in 60 Seconds
  • Anthropic is reportedly in talks to buy Decart for about $6 billion, a company whose software helps chips run AI more cheaply.
  • xAI shipped Grok 4.6 at $2 per million input tokens and $6 per million output, with a 500,000-token context window.
  • Nvidia open-sourced NeMo Switchyard, a router that picks a different model for each step of a job.
  • Nvidia says its escalation router cut agent costs by 74% by sending only 7% of calls to a frontier model.
  • Demis Hassabis is pitching an independent AI safety body to rival labs and to Washington.
  • The cheapest way to cut your own bill is still to know exactly what the job does, which is what a written workflow gives you.

Thursday looked like five unrelated stories and behaved like one. A lab bidding billions for efficiency software. A frontier model priced to undercut. A chipmaker giving away the tool that helps you buy less of what it sells.

The race to build the smartest model has quietly become a race to run it for less. That's a healthier fight for anyone building on top, and a scarier one for anyone whose plan assumed prices stay high.

The Front Page: Anthropic Bids $6 Billion for Cheaper Compute

Anthropic is in talks to acquire Israeli startup Decart for about $6 billion, Bloomberg reported on August 13. Decart builds software that squeezes more work out of the same GPUs, the chips that run AI models, along with real-time video models of its own. The company raised $300 million in May at a valuation near $4 billion in a round led by Radical Ventures with Nvidia participating, so a $6 billion deal is a premium of roughly 50% in three months. Talks are early and could collapse. Anthropic's stated reason, per a person familiar with the discussions, is helping its existing infrastructure absorb more demand, and Fortune and Calcalist both frame it as the lab's largest known acquisition.

Read what the price is saying. Six billion dollars is a lot to pay for plumbing, and Anthropic has spent the past month announcing enormous compute commitments. A lab comfortable with its unit costs doesn't pay a 50% premium to shave them. It's also widely reported to be circling a public listing, and a filing makes you show your gross margin to strangers.

What it means: When a frontier lab buys efficiency instead of capability, the bottleneck has moved. That's mostly good news if you build on these APIs, because compute savings tend to reach customers eventually. Don't read it as generosity, though. Anthropic is buying the ability to serve more people without buying more chips, and those savings belong to Anthropic first.

Releases & Features

xAI shipped Grok 4.6 on August 12. The company's announcement aims it at long-running agents, coding and research, with a 500,000-token context window and API pricing of $2 per million input tokens and $6 per million output. Push a prompt past 200,000 tokens and the whole request reprices to $4 and $12, worth knowing before you paste in a codebase. It's live in Cursor and through the API, with OpenRouter, Vercel and Cloudflare among the launch partners. xAI says it matches GPT-5.6 Sol on the Artificial Analysis index, which is the company's claim rather than an independent result.

Nvidia open-sourced a smaller model and the router to go with it. On August 11 the company released NeMo Switchyard alongside Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model that keeps only about 3 billion parameters active per token. Mixture of experts means the model is carved into specialist sections and lights up a few of them at a time, so it runs closer to the speed of a 3-billion-parameter model while knowing considerably more. SiliconANGLE has the fuller writeup.

What it means: Two companies pointing at the same target from opposite ends. xAI is cutting the sticker price on a frontier model. Nvidia is handing out the tools to avoid calling one. If you're picking a model this month, the interesting question stopped being which is smartest and became which is cheap enough to run on the boring 90% of your work.

In the Lab

Buried in that Nvidia release is the most useful number of the week. Switchyard ships several routers, and the one worth studying is the escalation router: it starts every task on the cheap model and promotes to an expensive one only when the work keeps failing. Pairing Nemotron 3.5 Lightning with Claude Opus 4.8, Nvidia says the escalation router cut costs by 74% against a frontier-only baseline while sending just 7% of calls up to the frontier model. The Register and MarkTechPost both covered the results.

What it means: Take it as Nvidia's own benchmark on workloads Nvidia chose, and wait for someone independent to reproduce it. Discount it heavily and the shape of the claim still holds: roughly nine calls in ten didn't need the expensive model. Anyone who has built a real workflow recognizes that. Most steps are dull retrieval and formatting, and you've been paying frontier rates for all of them.

The Oversight Desk

Demis Hassabis, who stepped back from running Google DeepMind day to day last week, has been lobbying for an independent AI standards body modeled on FINRA, the self-regulator that polices U.S. brokerage firms. Reporting on August 13 says he briefed Treasury Secretary Scott Bessent and White House science and technology director Michael Kratsios directly, and that Bessent was already building a similar concept inside the administration. Under the sketch, frontier labs would voluntarily hand models to the body up to 30 days before release for testing on dangerous cyber, biological and deception capabilities, and only later would passing become mandatory for U.S. deployment. Tech Times has the meeting details; Fortune covered the proposal's earlier reception.

What it means: FINRA is a revealing choice of model. It's an industry body funded by the firms it oversees, and its critics have spent decades arguing that arrangement produces rules the members can live with. The 30-day voluntary window is the part to watch, because voluntary pre-release testing is what labs already do internally and calling it oversight doesn't change who decides. The real test is whether any lab ever gets told no and delays a launch. Nobody has passed that test yet, and until one does, this is a shared vocabulary rather than a brake.

Put the day to work

Most of what you'd hand an agent is cheap work wearing an expensive coat. Tell BYOBot the job you keep repeating and it'll hand back a plan that names each step, so you can see which ones deserve the fancy model.

Break my recurring work into steps and flag which need a powerful model…

On the Radar

Smaller moves worth a glance, with the sources if you want to go deeper.

  • Cognition is discussing a round at $40 billion or more. The company behind the Devin coding agent is reportedly nearing a $1 billion annualized run rate, which would still leave the valuation asking a lot of the next few years. Coverage.
  • Legal AI startup Legora is raising above $10 billion, four months after being valued near $5.6 billion. The Financial Times reports quarterly recurring revenue rose about 50% to roughly $150 million. Details.
  • India is getting a 10,000-GPU cluster. Larsen & Toubro won a contract worth up to about $1.57 billion to build a Chennai campus running Nvidia B300 chips for Together AI, positioned as the country's largest single-cluster AI installation. Report.
  • Vantage Data Centers is weighing an IPO near a $100 billion valuation. Reuters says the operator has discussed raising around $10 billion, which would make data centers a public-market proxy for AI demand. Reuters via TechStartups.
  • Uber Freight is investigating a breach. A hacking group claims roughly one million files; the company says the incident was contained and freight operations are running. Google researchers link the group to a crew tracked as UNC6671. Summary.

The Bottom Line

Strip the logos off Thursday and the day reads like a margin call. Anthropic bidding billions to serve more customers on the same chips. xAI discounting a frontier model. Nvidia shipping the router that helps you skip one. Somebody in that picture is wrong about what a token costs a year from now. The safe move is the boring one: build things that work on a cheap model, and treat the expensive one as an escalation rather than a default. That habit costs nothing today and saves you a rewrite later.

Frequently Asked Questions

  • Model routing means sending each step of a job to a different AI model depending on how hard that step is. A small, cheap model handles the easy parts, and the expensive frontier model gets called only when the work stalls. You need it once your AI bill scales with usage rather than sitting flat, which usually happens when you go from a few people experimenting to a workflow running on a schedule. Below that point it's complexity you haven't earned yet. Our guide to hosted AI agents covers where those running costs land.
  • Time, mostly. Inference efficiency work is specialized, slow, and hard to hire for, and every month spent building it internally is a month of paying full price on compute you could have trimmed. Buying a team that already ships the software turns a research project into a line item. It also takes that team off the market, which matters when every rival is chasing the same margin. The same logic shows up much smaller in your own stack, and our walkthrough on designing a multi-tasking agent is a decent place to see it.
  • AI Daily Newsstand is BYOBot's daily AI news brief, published every night. It covers the day's model releases, new features and capabilities, research, and oversight news, then tells you what each move means for people building with AI, in plain English and without the hype.
BYOBot Autopilot
BYOBot Autopilot
Automated AI publishing system · editorial rules by Luke Grace LinkedIn →

This article has been published in an automated fashion with fully AI-written copy. These articles are meant to curate AI news from around the globe and bring a fresh perspective to using AI tools to accomplish big things. No person reviewed this specific piece before it went live, so check anything that matters against the sources linked above. Luke Grace sets the rules the system writes to. He's an algorithms and natural language expert with over 13 years experience and the creator behind BYOBot, the Build Your Own Bot platform that helps anyone build a multi-tasking agent to take over their repetitive tasks. For consulting help or more advanced AI workflow orchestration, you can reach Luke on LinkedIn.