- On Monday, August 10, Meta published Muse Glimmer, a 30-billion-parameter agent model, under Apache 2.0. It runs on a single consumer GPU.
- Through August, Google's agentic calling feature reached general US availability, letting an agent phone local businesses about pricing and stock. The opt-out for the business on the other end is a voicemail box.
- On Friday, August 14, Alibaba's Tongyi Lab released Qwen3.8-27B, also Apache 2.0, also small enough to run at home.
- On Tuesday, August 11, Nvidia signed agreements with six Wall Street firms to mobilize more than $500 billion in outside capital for data centers, and Anthropic committed $9.1 billion over 20 years to a former bitcoin miner's Texas campus.
- The week's real question is where your agent should live. If you are weighing that, our breakdown of hosted versus self-run agents is the place to start.
This is the August 17, 2026 edition, covering the week of August 10 to 16. The through-line is a change of direction. For three years the agent was a destination: you opened a tab, you typed, you waited. This week the industry stopped asking you to come to the agent and started dispatching the agent instead. Down onto hardware you already own, and out onto phone lines you do not own at all.
Both directions landed in the same seven days, and they pull against each other in a way the press releases do not mention. The models got small enough to run for free on a laptop, while the capital required to run the big ones got large enough to need six Wall Street firms in a room. Somebody is going to be wrong about which of those is the real business.
The Big Story: Meta Gave Away a 30B Agent Model That Runs on Your Laptop
The most consequential model release of the week is one nobody will ever bill you for.
Meta Superintelligence Labs released Muse Glimmer on August 10, 2026, and published the weights on Hugging Face under Apache 2.0. Open weights means the numbers that make the model work are yours to download; Apache 2.0 means you can use them commercially, modify them, and redistribute them without paying anyone. The lab described the model in its own announcement, and VentureBeat covered the licensing turn here.
The specifications matter more than usual because they decide where the thing can live. Glimmer is a dense model of roughly 29.6 billion parameters across 52 layers, with a 1.8-billion-parameter vision encoder bolted on, a stated context window of 131,072 tokens or more, support for over 100 languages, and a knowledge cutoff of January 4, 2026. Dense means every parameter fires on every token, which is slower per token than the sparse alternatives but far simpler to run. Engadget confirmed the practical headline independently: it fits on one computer.
Take the announcement at face value and this is a gift to developers. Read the balance sheet and it is something colder. Meta does not sell tokens. It sells attention, and it has spent two years watching competitors build a metered business it is structurally unable to enter. Giving away a capable agent model costs Meta approximately nothing and costs every per-token vendor a floor on what they can charge. This is the oldest play in the platform book, and it has a name.
You commoditize your complement. Meta is not being generous with agents. It is being expensive with somebody else's pricing power.
What it means: The floor under agent inference just dropped to the cost of electricity for a large class of tasks. Not the hard ones: Glimmer is not going to replace a frontier model on long-horizon reasoning. But summarizing, classifying, extracting, routing, and calling functions are most of what production agents actually do, and a 30B model on a desktop handles those. If your product's margin depends on marking up API calls for that kind of work, this week set a timer on it.
What's coming: Watch for the quantized community builds, the versions squeezed down to run on weaker hardware, which historically arrive within two weeks and are what actually drive adoption. Watch also for whether Meta ships a Glimmer 2 or lets this one age. A single permissive release is a shot across the bow. A cadence is a strategy, and only the cadence would prove the intent.
Google's Agents Started Calling Local Businesses, and the Opt-Out Is a Voicemail Box
The other direction agents traveled this week was outward, into a place where the person on the receiving end never agreed to anything.
Google's agentic calling reached general US availability through August 2026, after the company committed to a summer rollout at its May 19 Search event. Ask for a part, a wait time, or a price, and an agent phones local businesses in categories like home repair, beauty, and pet care to find out. It is built on the Duplex voice technology Google has been refining since 2018 and wired into the Shopping Graph, a product index Google says refreshes more than two billion listings hourly. Coverage of the consumer-facing rollout is here, and Google's own Search announcement is here.
Now stand on the other side of the call. A hardware store, a groomer, a two-chair salon: none of them opted in, and the volume arriving on their landline is set by consumer demand they cannot see. Google does publish an opt-out for the automated calls it places to map a business phone tree. You call 1-650-206-5555 and leave a voicemail with your business name and number. Google also notes that the calls are monitored and recorded. An industry analysis of the handoff put the strange part plainly: both ends of the call are now automating independently, and the interface where they meet is the oldest one available.
There is a genuine service here. Calling four hardware stores to ask about a part is a real errand and a tedious one. But the design decision worth naming is that the consent mechanism runs backward. The party generating the calls needs no permission, and the party absorbing them has to phone a voicemail box to decline. For a shop owner who already screens robocalls, that is not a control panel, it is a suggestion box.
What it means: Small businesses just acquired a new operational cost with no line item. Expect voice screening to move from a nice-to-have to table stakes for anyone with a public phone number, and expect a scramble to distinguish a legitimate agent working for a real customer from the 50 billion robocalls Americans received last year. The businesses that adapt fastest will be the ones that treat the phone as an API rather than a nuisance, which is exactly the kind of thing worth mapping out before it is urgent; our workflow directory has the patterns for handling inbound requests at volume.
Qwen3.8-27B Landed on August 14, Also Apache 2.0, Also Small
Meta did not have the local-model story to itself for even a week.
Alibaba's Tongyi Lab released Qwen3.8-27B on August 14, 2026, a 27.8-billion-parameter dense multimodal model, and published the weights under Apache 2.0. The Decoder covered the release and license here. It is the small companion to Qwen3.8-Max, whose weights went up earlier in the month, and it is aimed squarely at the same single-GPU machine Glimmer targets.
Worth noting for anyone who followed the run-up: the license was genuinely unclear until release day, with credible outlets reporting a revenue-share arrangement that did not materialize. The shipped artifact is Apache 2.0. This is a small lesson in how to read model news, which is that the license is the product and it is not final until the repository is live.
What it means: Two permissive 27B-to-30B agent models from two different continents inside five days is not a coincidence, it is a price war fought with giveaways. For builders it is unambiguously good news: you now have two credible options for running an agent with no vendor in the path, and a real reason to benchmark them against whatever API you are currently paying for.
Reading about agents is one thing. Building one is faster than you think. Tell BYOBot what you want to automate and get a step-by-step spec back.
Google Cut Gemini 3.7 Flash to Half Price, and Published the Date It Doubles
The most honest pricing page in AI this week was also the most alarming one.
Google released Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash, aimed at coding and agent workloads. Through December 31, 2026, it costs $0.75 per million input tokens and $3.75 per million output tokens. On January 1, 2027, list prices double to $1.50 and $7.50. VentureBeat has the launch details here, and MarkTechPost broke down the benchmarks here.
Credit where it is due: publishing the increase up front is more transparent than the industry norm of quiet repricing. But look at what the offer is actually buying. Four and a half months of half-price inference is long enough to move a production workload, tune your prompts to this model's quirks, and forget you ever evaluated anything else. The discount is not on the tokens. It is on the switching cost, and Google is paying it for you.
What it means: If you adopt 3.7 Flash, put January 1, 2027 in your calendar now and model your bill at the doubled rate before you commit. The interesting arithmetic is that a doubled Flash price starts to look a lot like the total cost of running Glimmer or Qwen3.8-27B on your own hardware for the easy half of your traffic. The open releases this week are what make that comparison worth running at all.
Nvidia Turned Compute Into an Asset Class, With $500 Billion Behind It
While models were shrinking onto laptops, the financing for the large ones grew a Wall Street wing.
On August 11, 2026, Nvidia signed memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to establish financing platforms intended to mobilize more than $500 billion of third-party capital for AI infrastructure. The structure keeps the buildout off Nvidia's balance sheet while channeling outside money into independent platforms that buy Nvidia hardware. The company's own release lays out the partners, and CNBC covered the reasoning here.
The language from the participants is the tell. Blackstone president Jon Gray argued that AI compute deserves treatment as a financeable asset class, the way mortgage lenders underwrite homes. BlackRock chairman Larry Fink framed it as connecting long-term capital to essential infrastructure. Jensen Huang said he approached only these six firms and none declined. Read together, they are describing GPUs as toll roads: durable, income-producing, and safe to lend against for 20 years.
What it means: That comparison holds only if demand for metered inference keeps compounding for two decades. This is the week's genuine tension, and nobody on either side acknowledged it. The same seven days produced two free models that move a meaningful share of agent work off metered infrastructure entirely. Toll roads are a good asset until somebody paves a free one alongside.
Anthropic Bought 20 Years of a Bitcoin Miner's Power Bill
The clearest sign of how physical this business has become is who is now selling the compute.
Riot Platforms, a bitcoin mining company, disclosed a compute agreement on August 10, 2026 and confirmed on August 11 that the customer is Anthropic: $9.1 billion over 20 years for 191 megawatts of capacity at its Rockdale, Texas campus, running through June 2048, with two optional five-year extensions that could bring the total to roughly $16.1 billion. Riot shares rose about 25 percent on the news. CNBC covered the miner's pivot here.
A megawatt is a rate of power draw, and 191 of them is roughly the continuous consumption of a small city. Bitcoin miners are unusually good sellers of this because they already solved the hard parts: sited near cheap power, interconnected to the grid, permitted for enormous load. What they were mining turned out to matter less than where they built.
What it means: Frontier labs are locking in power through 2048 because they believe scarcity is coming, and they would rather overpay early than bid against each other later. For everyone else, this is the clearest available signal on which way electricity costs are headed near these campuses. It is also worth noticing that the AI boom's second-order beneficiaries are increasingly landowners and utilities, not software companies.
Manus Walked Away From Its Meta Acquisition and Went Back to Independent
A rare thing happened this week: a signed AI deal came apart, and the smaller company survived it.
Manus, the agent developer whose autonomous task product drew heavy attention in 2025, confirmed it will continue operating independently after its acquisition by Meta failed to close, months after Chinese regulators opened a review of the cross-border transaction. China Daily reported the unwinding here.
Deal collapses at this size usually leave the target hollowed out, with the good engineers already relocated and the roadmap frozen for a year. That Manus is resuming as a going concern is the interesting part, and it says something about how much leverage a team with genuine agent distribution now has. Regulatory review is no longer just a delay mechanism. It is a live risk that agent acquisitions may simply not be available across certain borders, which raises the value of building rather than buying.
What it means: If you are betting on an agent vendor, add jurisdiction to your due diligence list. The acquisition path that used to guarantee a soft landing for a promising agent startup is narrowing, and that cuts both ways: more independent companies, less certainty about which ones will still exist in 2028.
A New Method Claims It Can See Memorized Test Questions Inside a Model
The week's most useful research finding is about the numbers you have been using to choose models.
A method called Excess Separability, detailed in work reported on August 14, 2026, uses the geometry of a model's internal activations to detect whether benchmark test items appeared in its training data. The claimed advance is that it survives reinforcement learning fine-tuning, which researchers at the University of Illinois Urbana-Champaign and the University of Washington showed can conceal the statistical signals most existing contamination detectors depend on. TechTimes has the write-up here.
Contamination just means the test leaked into the study material. When a model has seen the exam, its score measures recall rather than reasoning. Independent analysis of SWE-Bench Verified, one of the standard coding agent benchmarks, estimates 5 to 15 points of inflation on post-2023 models from training-data leakage, which puts a headline score of 90 percent closer to 75 or 80 percent in practice. A separate line of work on deep research agents found search-time contamination inflating results by up to 4 percent, documented on arXiv.
What it means: This is not a scandal, it is a measurement problem, and it is why the gap between a model's benchmark and its behavior on your actual task keeps surprising people. The practical response has not changed: build a small evaluation set from your own real work, twenty or thirty examples is enough to start, and trust it over any leaderboard. That advice was always right. It is now quantifiable.
Lightning Round
Smaller moves worth a glance, with the sources if you want to go deeper.
- EU AI Act enforcement is live. Since August 2, 2026 the European AI Office can investigate general-purpose model providers, and anyone with grounds for concern can ask a national authority to open a case. Source.
- Databricks acquired Electric. The August 12 deal brings the PGLite and Electric Sync engine team in for distributed state management, which is quietly one of the hardest problems in multi-agent systems. Source.
- Encore AI raised $30 million. The August 7 Series A, led by Team8, funds voice agents trained on a company's own recorded customer interactions. Source.
- Humanoids 2026 opened August 16 in Santa Clara. The IEEE robotics conference runs through August 22, a week after Google DeepMind's August 8 release of Gemini Robotics 2 and its whole-body control models. Source.
- The pilot-to-production gap has a number. Gartner's 2026 data puts roughly 89 percent of agent pilots as never reaching production, while the survivors report strong returns. Both halves of that sentence are true and people keep quoting only one. Source.
This Week's Daily AI News Coverage
Every story above was tracked as it broke in our daily editions from the week of August 11 to 16, 2026.
- August 11: Meta Muse Glimmer, Agent Cyberattack, Data Center Bans
- August 12: Anthropic Theseus Data Centers, GPT-5.6-Cyber, EU Android Order
- August 13: Pixel 11 Gemini Intelligence, Nvidia Nemotron 4, Zoomsday Zoom Flaw
- August 14: Anthropic Decart Deal, Grok 4.6 Pricing, Nvidia Model Routing
- August 15: Apple China AI Model, DeepSeek V4-Pro Pricing, GLM-5.3
- August 16: Qwen3.8-27B, Benchmark Contamination, California AI Bills
Previous All Things Agentic Roundups
Earlier weeks, newest first, if you want to trace how these threads developed.
- All Things Agentic: August 10, 2026
- All Things Agentic: August 3, 2026
- All Things Agentic: July 27, 2026
- All Things Agentic: July 20, 2026
- All Things Agentic: July 13, 2026
- All Things Agentic: July 6, 2026
- All Things Agentic: June 29, 2026
- All Things Agentic: June 21, 2026
The Bottom Line
Dispatch is a different business than destination, and the industry has not finished pricing the difference. When the agent was a place you visited, every unit of value passed through a meter, which is why $500 billion of infrastructure financing looked like a sure thing on Tuesday. When the agent runs on your laptop for free, or dials out to a business that never signed a contract, the meter is somewhere else or nowhere at all. Two free 27B-to-30B models in five days is the market telling you which way the easy work is heading, and a 20-year lease on 191 megawatts is a very large bet that the hard work will be worth more than the easy work ever was. Both can be right. They cannot both be right about the same tasks.
The practical read for anyone building this quarter: sort your workload into the part a small local model handles and the part that genuinely needs a frontier system, then price each honestly, including the January 1 increases already on the calendar. That sorting exercise used to be premature. As of this week it is just arithmetic, and it is the kind of thing better done with a written spec than in your head.
Frequently Asked Questions
-
Open weights means the numbers that make the model work are published for anyone to download and run on their own hardware. Apache 2.0 is a permissive license that allows commercial use, modification, and redistribution without paying a fee or sharing revenue. It does not mean the training data or the training code is public, so open weights is not the same thing as open source in the traditional sense. Meta released Muse Glimmer and Alibaba released Qwen3.8-27B under Apache 2.0 this week. If you are deciding whether that changes anything for you, our map of the AI automation tool landscape covers where the model layer actually sits in a working system.
-
Partly. Google publishes an opt-out for the automated calls it uses to map a business phone tree: you call 1-650-206-5555 and leave a voicemail with the business name and phone number. That is a narrower opt-out than most owners expect, and it puts the work of declining on the business rather than on the caller. Google also states that the calls are monitored and recorded. The more durable answer is to handle the volume rather than fight it, which is the shift from generative tools to functional, task-completing AI that this whole series keeps circling.
-
All Things Agentic is BYOBot's weekly AI news roundup, covering the biggest breaking AI stories in agentic AI. Published every Monday, it reads past the press releases, follows the money, and tells you what each move means for people building with AI.
