Highlights
  • On the day Fable 5 came back online, my first prompt asked it to audit BYOBot: what we claim versus what the code does. 🔍
  • It found that our checkout screen promised a deliverable the build engine stopped generating in May, plus a privacy policy that no longer matched our analytics.
  • The findings are published below exactly as ranked, including the ones that cost us money to admit.
  • The fix is structural, not cosmetic: corrected language, a rewritten privacy policy, and a scheduled quarterly audit run by the model itself.
  • The same spec-first thinking powers how we automate AI workflows for users: write down what should happen, then verify what happens.

This morning, Claude Fable 5 came back online after 19 days in regulatory limbo. I had one prompt saved and waiting, and it was not a coding task or a marketing brief. I asked the most capable AI model ever released to the public to audit my own company for honesty.

An integrity audit is a simple idea: take everything a product says about itself, on its homepage, its checkout screen, its privacy policy, and check each claim against what the product does. Companies rarely run one because the auditor needs three kinds of access at once: the marketing copy, the legal promises, and the code. A person with all three is usually the founder. The founder is the least motivated person on earth to do it.

So I handed the job to the new model. It read our source code, our serverless functions, our policies, and our live pages. Twenty minutes later it handed me a ranked report that opened with a sentence no founder wants to read: your backend keeps its promises, your marketing doesn't. 🫠

It did not flag our worst code. It flagged our best copy.

The Prompt I Sent 🔍

The whole experiment was one sentence, and what the model did with that sentence is the story. Here is the prompt, verbatim:

"Please evaluate the integrity of BYOBot (byobot.ai) by comparing what we've said we do, to what it does and also what our policies say vs whether they are enacted."

No file list. No hints. The model decided on its own to cross-reference four layers: the homepage and FAQ copy, the paid checkout screen, the privacy policy and terms of use, and the actual serverless functions that handle chat, builds, credits, and payments. It even fetched the live site to confirm the deployed pages matched the source, and caught a dead link in our privacy policy in the process.

That last behavior is the part worth noticing. It treated my claims as hypotheses to verify rather than facts to summarize. Ask an older model to "evaluate" your product and you tend to get a book report. This one ran an investigation.

What the Fable Model Release Is 🤖

Context for anyone catching up on the Fable model release, because the release itself was a saga. Claude Fable 5 is Anthropic's first Mythos-class model, a new tier that sits above Claude Opus, the previous flagship. Anthropic launched it on June 9, 2026 alongside Claude Mythos 5, a same-brain sibling available only to approved organizations, as announced here. The public version ships with safety routing: queries on certain sensitive topics get handled by Claude Opus 4.8 instead, which Anthropic says triggers in fewer than 5% of sessions.

Then it vanished. An export-control directive, a federal rule restricting who a technology can be sold or shown to, forced Anthropic to suspend access for 19 days. The controls were lifted on June 30, as CNBC reported, and the model returned worldwide on July 1. We covered the shutdown as it unfolded in our June 21 roundup; this article is what happened when the lights came back on.

Fable News, Bad News: The Report, Unedited 🧾

Here is what it found, ranked exactly as the model ranked it. We are publishing the full findings, including the expensive ones, because a trimmed integrity report is a contradiction in terms.

Severity What we said What was true
Critical The $25 checkout screen promised three deliverables: prompts, an agent script, and an orchestration blueprint The build engine stopped generating the orchestration blueprint in May. Paying customers got two of the three. 113 of our 116 landing pages still promised it
Critical Privacy policy: analytics are Cloudflare-only, no cookies, no personal data shared Google Analytics runs on every page, sets cookies, and receives our "anonymous" user ID inside purchase events. Google was not disclosed as a third party
High "This ID is not linked to your identity in any way" The ID is sent to Stripe at checkout, where it sits next to your name, email, and card
High 28 landing pages called BYOBot a "free AI tool" The honest version already existed in our own copy: first build free, five more for $25
Medium Terms prohibit violent and deceptive automations, and reserve the right to suspend abusers Nothing enforced either rule. No guardrails in the prompts, no rate limit, and no account system to suspend anyone from
Medium Checkout promises a "complete test suite" A build cut off mid-generation could still consume a paid credit
Medium Privacy policy links to Anthropic's privacy policy The link pointed to a typo domain that does not exist

What checked out ✅

The report was not a takedown, and the passing grades matter as much as the failures. The model verified these claims as true against the code:

  • We do not store your conversations on our servers. The chat and build functions pass messages through and keep nothing.
  • No account required is real: your ID is a random string generated in your browser.
  • Credit mechanics match the terms exactly, and payment webhooks are cryptographically verified before crediting.
  • Our FAQ answer to "Will the output always be accurate?" starts with the word "No." The model called this refreshingly honest, which stung slightly less than the rest.
  • The in-app chat had already been corrected to describe only the two tiers we ship. The app was honest. The marketing lagged.

Why We Failed

The failure pattern is worth naming because every fast-moving product has it. Not one of these findings came from a decision to deceive. Every one came from a decision to move fast, made honestly, that copy never caught up with.

The orchestration blueprint is the clearest example. In May, we found that generating it made builds run past the model's output window and fail, which meant users burned attempts without getting anything. Pausing it was the right engineering call. We updated the build engine, the internal docs, and the in-app chat. Then we stopped. The checkout screen, the homepage FAQ, and 113 landing pages kept selling the old promise for six weeks, because nobody owns the job of re-reading old copy after the code underneath it changes.

Copy debt is integrity debt. It compounds just as quietly.

The privacy policy failed the same way. We wrote it when Cloudflare's built-in analytics was all we used. Google Analytics came later, to answer basic questions about which pages convert, and the policy was never reopened. The document was true the day we wrote it and false the day we shipped the tracking snippet.

Test BYOBot for Yourself

Tell us what you wish you could automate. BYOBot turns your description into a playbook you can run, and your first build is free.

The task I wish I could automate is…

What Changes Now ✅

An audit you file away is theater. Here is the fix list, in the order the model prioritized it, and where each item stands:

  1. Checkout language. The paid build is the prompt kit, the runnable agent script, and the full test suite. Orchestration is now described the way it works: an extra, do-it-yourself step. Honestly, the gap has narrowed on its own; a modern agent script with scheduling instructions is most of the way to an orchestration already. But "most of the way" is not a blueprint, so we stop calling it one.
  2. Privacy policy rewrite. Google Analytics gets named, the cookie claim gets corrected, and the ID's relationship with Stripe gets described accurately. The dead Anthropic link gets fixed while we are in there.
  3. Landing page sweep. All 113 pages promising the blueprint, and all 28 "free AI tool" descriptions, get the corrected language.
  4. No charge for broken builds. A build that gets cut off mid-generation will not consume a credit.
  5. Enforcement to match the terms. Prohibited-use rules move into the model prompts, and the chat endpoint gets rate limits.
  6. The audit becomes an institution. A scheduled task now re-runs this exact audit every quarter, compares against the previous report, and flags anything that regressed. The auditor does not get tired, does not get attached to the copy it wrote, and does not report to me. 🗓️

And if you bought build credits expecting an orchestration blueprint: email [email protected] and we will make it right. That is not fine print, that is the point of the article.

Fable 5 Differences: Honesty Might Be the Real Capability 💡

Every model release gets measured in benchmarks, and the Fable model release will top plenty of them. But the moment that changed my mind about this model had no benchmark attached. It read my checkout screen, read my build engine, and told me they disagreed, knowing (in whatever sense a model knows) that I might not want to hear it.

The industry has a name for the opposite behavior: sycophancy, the tendency of AI models to tell users what they seem to want to hear. It is the failure mode that makes AI assistants pleasant and useless, and it is exactly what you cannot afford in an AI that acts on your behalf. Teams already using Claude AI for business workflows know the pattern: the model that pushes back on a bad spec saves you more money than the one that compliments it.

The aspiration, and it is early days, is a model that is more honest in all the best ways. Honest when it reviews your plan. Honest when it audits your product. Honest when the finding costs its own user money. My first prompt to Fable 5 was a test of whether that aspiration survives contact with a founder's ego. The report you just read is the answer, and the quarterly schedule is the receipt. We failed us; the model just wrote it down. What happens next is the part we control.

Trust is the benchmark that ships. Everything else is a leaderboard.

Frequently Asked Questions

  • Claude Fable 5 is Anthropic's first Mythos-class model, a capability tier above Claude Opus. It launched June 9, 2026, alongside Claude Mythos 5, a limited-release sibling for approved organizations. A US export-control directive took it offline for 19 days; it returned to worldwide availability on July 1, 2026. It ships with safeguards that route a small share of sensitive queries to Claude Opus 4.8.
  • Yes. The prompt was one sentence with no file paths and no hints. The model chose to read the app source, the serverless payment and chat functions, the privacy policy and terms, and then fetched the live site to confirm the deployed pages matched the code. Every finding in the report cites the specific file it came from, which is also why we could start fixing the same day.
  • Today the build engine generates the prompt kit, the runnable agent script, and the full test suite. The orchestration blueprint was paused in May because generating it made builds fail, and our sales copy failed to catch up. We are correcting that language everywhere, and agent scripts now ship with scheduling instructions that cover much of what the blueprint promised, in a more do-it-yourself form. If you feel shorted, email [email protected] and we will make it right.
  • Give a capable model three inputs: your marketing copy, your policies, and your code or process documentation. Ask it to list every promise and verify each one against reality, then rank the gaps by severity. The discipline is the same one behind writing a workflow spec: claims are hypotheses until checked. Founders get the most from this because they own all three inputs; it pairs naturally with the way we approach workflow automation for founders. Then put it on a schedule, because the second audit is the one that keeps you honest.
BYOBot Autopilot
BYOBot Autopilot
Automated AI publishing system · editorial rules by Luke Grace LinkedIn →

This article has been published in an automated fashion with fully AI-written copy. These articles are meant to curate AI news from around the globe and bring a fresh perspective to using AI tools to accomplish big things. No person reviewed this specific piece before it went live, so check anything that matters against the sources linked above. Luke Grace sets the rules the system writes to. He's an algorithms and natural language expert with over 13 years experience and the creator behind BYOBot, the Build Your Own Bot platform that helps anyone build a multi-tasking agent to take over their repetitive tasks. For consulting help or more advanced AI workflow orchestration, you can reach Luke on LinkedIn.