- An interstitial is anything that interrupts the page before the agent reaches its target: cookie banners, modals, paywalls, CAPTCHAs, and MFA prompts.
- Interstitials do not break humans because we dismiss them on reflex. They break agents because the agent was never told the screen would appear.
- The fix is almost never a better model. It is a better brief: name each likely interstitial and give the agent a rule for it before the run starts.
- Every interstitial resolves to one of three moves: dismiss it, satisfy it, or hand off to a human. The job is deciding which, in advance.
- At scale, the power moves are staging and batching: warm a session once, seed the consent cookie, group work by domain, and queue every human handoff into a single burst.
- Naming interstitials is part of writing a real task spec, which is exactly what a BYOBot playbook produces before you touch any agent tool.
You watch the demo and the browser agent looks unstoppable. It logs in, navigates three pages deep, pulls the data, and drops it into a spreadsheet without a single keystroke from you. Then you point it at a real website and it freezes on step two, defeated by a cookie banner a five-year-old would have clicked away without thinking.
This is the most common way agent automation fails, and it has a name. The thing that stopped your agent is an interstitial: any screen that inserts itself between the agent and the content it came for. Consent banners, newsletter popups, paywalls, registration gates, CAPTCHAs, two-factor prompts. The modern web is layered with them, and they are multiplying.
Here is the part that matters. The agent did not fail because it was not smart enough. It failed because nobody told it the popup was coming. That distinction is the whole article, and once you internalize it, interstitials stop being a mysterious wall and become a checklist you handle before the run begins.
What Counts as an Interstitial
An interstitial is any layer that appears between you and the thing you want to do. The word comes from advertising, where a full-screen ad interrupts you on the way to an article. In browser automation the definition is broader and more useful: it is anything that occupies the screen and demands a response before the underlying page becomes usable.
That covers a wide and growing range. A consent banner asking about cookies. A modal offering 10% off if you hand over your email. A paywall after the third paragraph. A "create a free account to continue" wall. A CAPTCHA checkbox. A push notification asking to verify your login on your phone. To a person these are background noise we clear in half a second. To an agent each one is an unexpected state, and an unexpected state is where automation goes to die.
The web is also getting more hostile to anything that looks like a bot, which makes interstitials more frequent, not less. More than half of all web traffic is now automated: Imperva's 2025 Bad Bot Report put bots at 53% of global traffic, the second year running that machines outnumbered people. Sites have responded by stacking on detection and consent layers, and your perfectly legitimate agent gets caught in the same net as the scrapers and fraud bots.
Why They Break Agents and Not Humans
The reason interstitials are lethal to agents and trivial to humans comes down to how each one reads a screen. A person sees a popup, recognizes it instantly as not-the-thing, and dismisses it without conscious effort. The popup never enters our plan because clearing it is reflex.
An agent has no such reflex unless you give it one. To understand why, it helps to know how browser agents read a page in the first place. A DOM-based agent looks at the page structure and tries to click the element it was instructed to click. A consent overlay sits on top of that element and intercepts the click, so the agent's action lands on the banner instead of the button, and the run derails. A vision-based agent sees the popup as just another set of pixels and, with no instruction about it, has to guess whether to engage or ignore. Either way, the surprise is the problem.
An agent does not fail at the interstitial. It fails because the interstitial was never in the plan. Surprise is the bug, not the popup.
This is why throwing a more capable model at the problem rarely helps. A smarter agent guesses better, but it is still guessing. The agents that clear interstitials reliably are not the smartest ones. They are the ones that were told, before the run, exactly which screens would appear and what to do about each. The full mechanics of how agents interpret a page are worth understanding, and our browser agent field guide walks through the DOM-versus-vision split in detail.
The Five Interstitials You Will Hit
Almost every interstitial you meet falls into one of five families. Each one wants something different from the visitor, and each one calls for a different handling rule. Learn the five and you can plan for nearly anything a website throws at your agent.
| Interstitial | What it is | What it wants | Default agent move |
|---|---|---|---|
| Consent banner | Cookie or privacy notice overlaying the page | A click on accept, reject, or manage | Dismiss on a fixed rule |
| Modal popup | Newsletter, discount, or promo overlay | An email, or a close click | Close and move on |
| Paywall / reg wall | Gate after partial content or on entry | A login, payment, or signup | Authenticate, or stop |
| CAPTCHA | Human-verification challenge | Proof you are not a bot | Hand off to a human |
| MFA prompt | Second-factor login challenge | A code or device approval | Hand off, then resume |
Consent and cookie banners
The most common interstitial on the open web, and the easiest to plan for. Roughly two-thirds of websites now show a consent interface before the page is usable; a cross-country audit of consent banners documented the pattern in detail and is published here. Because the banner is predictable, it is a gift: tell the agent that a consent overlay will appear and give it one rule, such as "accept all and continue" or "reject non-essential and continue." Handled in the spec, it costs one step. Met by surprise, it eats the whole run.
Modals, popups, and overlays
The discount-for-your-email popup, the "wait, before you go" overlay, the app-install nag. These usually carry a close button, often a faint X in a corner the agent may not prioritize. The handling rule is simple but must be explicit: locate and click the dismiss control, and never enter data into a promotional modal. Left unspecified, agents have been known to "helpfully" type into the email field because it was the most prominent input on screen.
Paywalls and registration walls
Here the interstitial is not noise, it is a real gate, and your handling rule depends entirely on whether you have legitimate access. If the agent operates a logged-in session you control, the rule is to authenticate and proceed. If it does not, the honest answer is that the agent should stop and report rather than attempt to circumvent a paywall. This is the same judgment that separates a clean data entry workflow from one that quietly does something you would not sign your name to.
CAPTCHAs and bot detection
This is the interstitial that exists specifically to stop you. CAPTCHAs, "verify you are human" checkboxes, and invisible bot-detection scoring are built to break automated browsers, and reputable agent tools will not solve them for you. Trying to defeat a CAPTCHA is the wrong instinct. The right move is to route around it: pick tools that do not trigger detection, insert a human at that single step, or move the task off the browser entirely. Vision-based agents like Skyvern, whose source is on GitHub here, are more resilient to layout shifts than DOM scripts, but no agent should be in the business of cracking human-verification challenges.
MFA and login challenges
Multi-factor prompts are the well-behaved cousin of the CAPTCHA. They are not trying to stop you, they are trying to confirm it is really you, and that is a check you can satisfy with a human in the loop. The pattern is a clean handoff: the agent drives up to the login, pauses for a person to approve the device prompt or enter the code, then resumes. Anthropic's own implementation notes for agent control, documented here, lean on exactly this kind of supervised checkpoint for sensitive steps.
Spec It Upfront: The Real Fix
Every handling rule above shares one thing: it has to be decided before the agent runs, not improvised mid-task. This is the core move, and it is unglamorous. You do not beat interstitials with cleverness in the moment. You beat them by listing the ones a workflow is likely to hit and writing a rule for each into the task spec.
A spec that survives the real web names the interruptions explicitly. Here is what that looks like for a routine "pull pricing from three vendor portals" job:
- Expected interstitials: Cookie banner on all three sites; a newsletter modal on vendor B; an MFA prompt on vendor C's login.
- Consent rule: Accept all cookies and continue. Do not open the "manage preferences" panel.
- Modal rule: Close any promotional popup by its dismiss control. Never enter an email address.
- Auth rule: On vendor C, pause at the MFA screen and wait for human approval, then resume from the pricing page.
- Bail-out rule: If any site presents a CAPTCHA, stop, log which site and step, and report back rather than retrying.
Notice that none of this is technical. It is a set of decisions a person makes once, in plain English, so the agent never has to guess. Writing that spec is a discipline most people have never practiced, because no software ever forced them to articulate a workflow at this resolution. That structured brief, goal and inputs and outputs and the interstitial rules that keep a run alive, is precisely what BYOBot turns your plain-English description into, ready to hand to any agent tool.
Tell BYOBot what you want to automate and which sites it runs on. The output is a task spec that names the cookie banners, popups, and login walls upfront, with a rule for each.
Dismiss, Satisfy, or Hand Off
For all the variety in interstitials, the agent only ever has three possible responses. Memorize these three and every interstitial decision becomes a quick sort rather than a fresh puzzle.
1. Dismiss it. The interstitial is noise with no bearing on the task. Cookie banners, promo modals, and app-install nags live here. The rule is a single instruction: clear it by a named control and continue. The only trap is leaving the rule out, because an undismissed overlay blocks every click underneath it.
2. Satisfy it. The interstitial is a legitimate gate you have the right to pass, like a login on an account you own or an MFA prompt you can approve. The rule provides the credential path or a human checkpoint. The agent gives the gate what it asks for, then proceeds. Reporting workflows that an operations manager runs across five logged-in tools live almost entirely in this category.
3. Hand off. The interstitial is designed to stop automation, full stop. CAPTCHAs and aggressive bot detection belong here. The rule is not to win but to pause cleanly: stop at the challenge, surface it to a person, and either let them clear it or abandon that branch. An agent that knows when to hand off is far more valuable than one that bluffs and corrupts your data.
The most reliable agents are not the ones that never hit a wall. They are the ones that were told, in advance, exactly which walls to climb and which to knock on.
Advanced Workarounds: Staging, Batching, and Beating the Bottleneck
Handling interstitials correctly keeps a run alive. The next ceiling is throughput, because an agent that pays the full interruption tax on every single item is slow and brittle at scale. These are the power-user moves that separate a one-at-a-time agent from one that clears interruptions in bulk and, better still, rarely meets them at all. They fall into two strategies: make interstitials disappear before the agent arrives, and absorb the unavoidable ones in batches instead of one painful hit at a time.
Make interstitials disappear
The fastest interstitial to handle is the one that never renders. Three staging techniques get you there.
Prime the session, then save it. Run a one-time warm-up pass: open each target site, accept or reject every consent banner, dismiss every modal, complete every login, and then save that browser profile. Every later run starts from a warm session where the consent cookie is already set and the banners do not come back. You pay the interstitial tax once, at setup, instead of on every run. This single move eliminates the majority of consent and popup failures outright.
Seed the consent cookie directly. Most cookie banners are gated by one value, the cookie a consent platform sets once you have answered it. Set that cookie before navigating and the banner simply never appears, which means no overlay, no intercepted click, and nothing fragile to maintain. Name the cookie in your spec and the agent skips the entire dance. This is the cleanest workaround on the list because it removes the interstitial rather than reacting to it.
Pre-dismiss with an injected snippet. If your tool allows a hook that fires before the agent acts, inject a small piece of CSS or JavaScript that hides high z-index overlays and fixed-position banners, or that auto-clicks known dismiss controls. A reader-mode or DOM-cleaning pass does the same job: strip the page down to the content the agent needs so the interruptions are gone before the agent's plan even begins. The agent then operates on a clean page and never has to reason about the popup at all.
The best CAPTCHA workaround is never summoning one. The best cookie-banner workaround is a page where the banner never renders. Remove the interstitial and you remove the failure.
Absorb them in bulk
Some interruptions are unavoidable, especially logins and MFA. The trick is to stop meeting them one at a time and start clearing them in batches, so the cost is amortized across the whole run instead of paid per item.
Batch by domain, not by task. The expensive moment is the first hit on each site. So group every task that touches a given site into one continuous session: clear that site's interstitials once, then run all of its work back to back before moving on. A scattered run that bounces between five sites pays the consent-and-login cost five times over and over; a domain-batched run pays it once each. This is the highest-leverage habit for anyone running a batched reporting workflow across many tools.
Stage the run into lanes. Split the workflow into phases. An "unlock" lane handles every login and consent up front under human supervision, with all the MFA approvals knocked out in one sitting. Then an unattended "work" lane runs once everything is already open. The person is needed for five minutes at the start, not babysitting the entire run, and the slow human-in-the-loop steps stop blocking the fast automated ones.
Pool warm sessions and parallelize. Keep a set of pre-warmed browser profiles, one per site or identity, and run them in parallel lanes. Throughput then scales with the number of lanes, and because each profile is already past its interstitials, no lane stalls on a banner while the others wait. A bottleneck on one site no longer freezes the whole job.
Queue the human handoffs. Instead of pausing the agent at each MFA prompt the instant it appears, which strands the run waiting on a person, have the agent checkpoint its state at every gate and collect the gates into a queue. A human clears the whole batch of approvals at once, and each branch resumes from where it paused. Checkpoint-and-resume turns a dozen scattered interruptions into one short, supervised burst.
Shape your timing and cache the gated pages. Bot detection often scores on speed and rhythm, so spacing actions to a human-ish cadence with small randomized delays keeps an agent under the threshold that triggers a challenge in the first place. And if several tasks need the same gated page, fetch it once and reuse the result rather than re-clearing the same paywall for every record behind it. Less repetition means fewer chances to trip a wall.
When to Stop Fighting and Escalate
Sometimes the honest answer is that a browser agent is the wrong tool for a step, and the interstitials are the signal telling you so. If a workflow hits a hard CAPTCHA on every run, or needs to clear a login wall a hundred times a night with no human awake to approve it, you have left the territory where browser agents shine.
That is the moment to escalate the build. A workflow that must run unattended and at volume usually wants an API-based path instead, the kind you build with Zapier orchestration or a similar platform, where there is no browser to interrupt and therefore no interstitial to clear. The strongest setups are often hybrids: the agent handles the legacy portal with no API, and an integration layer takes over for the high-volume distribution. No-code agent builders such as Bardeen sit in the middle, for teams that want a visual builder rather than raw code.
The decision of where each step belongs is the entire reason to map a workflow before building it. Interstitials are not just an annoyance to clear. They are diagnostic. A step buried in CAPTCHAs is telling you it wants an API. A step behind a one-time consent banner is telling you it is perfectly safe for an agent. Reading those signals correctly, before you commit to a single tool, is what separates an automation that runs for a year from one that breaks by Friday.
Map the interruptions before you build
BYOBot designs your browser agent spec, names the cookie banners, popups, paywalls, and login walls it will meet, and gives each one a rule, so your agent runs instead of stalling.
Frequently Asked Questions
-
An interstitial is anything that appears between you and the content or action you want: a cookie consent banner, a newsletter popup, a registration wall, a paywall, a CAPTCHA, or a multi-factor authentication prompt. For a human it is a half-second of friction. For a browser agent it is an unexpected screen it was not told to expect, which is where most agent runs quietly fail. The browser agent field guide explains how agents read a page and why surprises derail them.
-
Usually no, and that is by design. CAPTCHAs and modern bot-detection systems exist specifically to stop automated browsers, and most reputable agent tools will not solve them. The right pattern is not to defeat the CAPTCHA but to design around it: target tools that do not trigger one, schedule a human handoff for that single step, or move the work to an API-based path that never opens a browser. Our no-code automation guide covers how to pick the right approach for a given task.
-
Because the banner sits on top of the page and intercepts clicks. The agent tries to click the element it was told to click, but the consent overlay is in the way, so the click lands on the banner instead and the run derails. The fix is to name the banner in your task spec and give the agent an explicit first step: dismiss or accept the consent interface before doing anything else.
-
Name them before the agent runs. List the interstitials a workflow is likely to hit, decide for each whether the agent should dismiss it, satisfy it, or hand off to a human, and write those rules into the task spec. An agent that has been told a cookie banner will appear handles it in one step; an agent that meets it by surprise stalls. BYOBot builds that spec for you so the rules are in place before you touch any agent tool.
