Highlights
  • The loop isn't the broken automation. It's the Monday morning spot check: opening six run histories one at a time to confirm the things you built are still alive.
  • Every alert you own is built on runs, and a workflow that stopped has no runs. Absence is the one event your platform cannot report, so the check has to come from a clock, not a trigger.
  • The prompt build turns your list of workflows into a monitoring contract: expected cadence, grace window, minimum plausible volume, and who gets told.
  • The script build is a watchdog with no AI in it at all. Each workflow pings a log as its last step, and the watchdog compares those pings against the contract and raises what's missing.
  • The schedule build puts the watchdog on a timer behind a dead man's switch, so silence anywhere gets caught within one cycle. That's how you automate the work that has to keep running instead of checking on it every Monday.

Welcome to The Loop. Every Wednesday we take one real, repetitive workflow that someone does by hand and unwind it into something you can run. This week the loop is the one you invented yourself after getting burned once: the recurring manual check that your automations are still working.

Here's the idea worth carrying out of this one. The boring task you repeat is a specification in disguise, and the task being repeated here is suspicion. Every Monday you open six tabs and ask each one the same question: did you run, and did you do about as much as usual? That question is a specification, and it is a very short one. The reason you're asking it by hand isn't that it's hard. It's that you've been waiting for your tools to volunteer the answer, and they can't, because they only know how to talk about runs that happened.

So let's sit next to somebody doing the Monday check by hand, price what the silence actually costs, and then build the thing that notices an absence instead of reporting a failure.

The Loop, in Full

Before you fix anything, you have to see the loop clearly. Picture Dana, who runs revenue operations alone at a forty person SaaS company. Over two years she's built nine automations that hold the business together: inbound demo requests from the website into the CRM, a nightly usage export into the billing sheet, trial expiry reminders, a Slack digest of closed-won deals, and five more. None of them is clever. All of them are load bearing. She built them herself in a no-code tool, which is exactly why this works at all and also why nobody else can check them.

Then a customer emailed asking why nobody had replied to the demo request they submitted nine days earlier. The form had been fine. The automation had stopped on day one, after the CRM's API deprecated a field, and the trigger had been returning nothing ever since. Zapier had not auto-paused it, because repeated errors trip that and there were no errors. Nothing was wrong. Nothing was happening.

So now there's a Monday ritual, and this is the fourth month of it:

  1. Open the automation platform, go to the run history, and sort by last run. Nine workflows, nine last-run timestamps to read and mentally compare against how often each one is supposed to fire.
  2. For the four that matter most, open the actual run list and eyeball the volume. Fourteen demo requests last week, eleven this week. Probably fine. Probably.
  3. Cross-check two of them against the destination. Count the rows that landed in the billing sheet, compare to the number of accounts, accept that the numbers are close enough.
  4. Check the second platform, because three of the nine live somewhere else, and its history screen works nothing like the first one.
  5. Write "all good" in the ops channel. Go do the actual job. Carry a low background hum of doubt until next Monday.

Forty minutes, and not one minute of it needs her judgment. She already knows what each workflow's normal looks like, and it hasn't changed in two years. What she's doing is reading timestamps and subtracting, which is the purest possible description of an automatable loop: the thinking happened once, and everything since has been arithmetic performed by a human being with a browser. This is the same trap that makes people abandon tools they were right to adopt, when the fix is to stand up a no-code automation that holds up rather than to stop building.

The Manual Tax

The real cost of a loop is never just the minutes on the clock. It's the minutes, plus the mistakes, plus the time the work sits waiting. This loop is unusual because the forty minutes a week is the cheapest part by a wide margin, and the ritual is specifically designed to catch something it mostly misses.

There's good evidence that the checking doesn't work. In the annual state of data quality survey run by Monte Carlo, 74% of data professionals reported that business stakeholders find problems first, all or most of the time, up from 47% the previous year. Three out of four teams, with monitoring tooling and dedicated engineers, learn their pipelines are broken from the person downstream who noticed the numbers looked wrong. The same survey put average time to detection at four hours or more for 68% of respondents. If that's the state of play where somebody is paid to watch, the Monday spot check on nine no-code workflows is not a safety net. It's a ritual.

Then add the two costs hiding behind the minutes:

  • Errors. The missed demo request is the expensive kind of error, because it never appears anywhere as a record. A failed run leaves a red entry in a log. A trigger that returns nothing leaves no trace at all, so there's no list to replay from and no way to count what you lost. Dana doesn't know how many other requests went nowhere in those nine days. She can't know. The evidence is the absence.
  • Latency. Nine days is the number to sit with. Not nine days of downtime, nine days of confident operation on top of a dead input: pipeline reviews run against a CRM missing a week of inbound, a forecast built from it, and a customer who concluded the company doesn't answer its own form. The detection gap, not the outage, is what does the damage.

An automation that fails loudly costs you an afternoon. One that stops quietly costs you whatever happens in the gap before somebody outside your company notices on your behalf.

So the goal isn't workflows that never break, because every one of them depends on an API somebody else is free to change on a Tuesday. The goal is to stop paying the tax: make absence a thing your system can detect, and compress nine days down to one cycle. Done once, that's the difference between automations you supervise every Monday and the freedom to automate a recurring job you actually depend on and then stop thinking about it.

Unwinding the Loop

Automating starts with description, not code. Describe the loop precisely enough that a machine could follow it without you in the room, and most of the work is done before a line is written. For this loop the description does something extra: writing down what each workflow's normal looks like is the artifact Dana has been carrying in her head for two years, and the reason nobody else has ever been able to cover for her on a Monday.

Here's the watchdog captured as a spec. The interesting parts are the decisions, not the steps:

Part of the loop What it is for this workflow
Trigger A clock, and nothing else. This is the whole point: the check must run on days the workflows don't, because the days they don't run are the days you care about.
Input A heartbeat log, one line per completed run, written by each workflow as its final step. Plus the contract: for every workflow, how often it should fire and what counts as a normal amount of work.
Decision: should this have run by now? Expected cadence plus a grace window sized to the job. A nightly export gets four hours. A form handler that fires on demand gets measured in business days of total silence instead.
Decision: is zero suspicious or normal? The question the Monday check keeps getting wrong. Zero orders on Sunday is Sunday. Zero orders on Tuesday is an outage. Baseline per day of week, and require consecutive days before escalating.
Decision: did it run and do nothing? The heartbeat arriving with a count of zero is a different failure from the heartbeat not arriving. One means the workflow is alive and its input died. The other means the workflow is dead. They need different messages and often different people.
Output One alert per silent workflow, naming it, the last confirmed run, the hours of silence, and the destination to check. Never a daily digest nobody reads.
Success signal The watchdog's own heartbeat, reported somewhere else. A monitor that can die quietly is the original problem wearing a new hat.

That last row is the one people leave out, and it's the one that makes the rest trustworthy. Notice also that two of the four decisions are judgment calls about your business rather than technical choices, which is why this can't be bought off a shelf and dropped in. If you'd rather answer questions than fill in a table, that's exactly how BYOBot turns a task into a spec, one question at a time, and the zero-is-normal row is where the conversation gets interesting.

Try It Now

Tell BYOBot which automations you check on by hand and it'll design the watchdog: the heartbeat, the contract, and the alert that fires on silence.

Help me build a watchdog that catches an automation that stopped…

The Build: One Loop, Three Ways to Run It

The same described loop becomes three builds depending on how much you want running on its own. Start at the top and move down only when you're ready. This loop is a rare case where the third build is the one that actually solves the problem, because a monitor you have to remember to run is not a monitor. The first two get you the contract and the logic that it runs.

The Prompt

Paste the list of what you've built into any AI chat tool with this. You're not asking it to monitor anything. You're asking it to interview you into a monitoring contract, which is the document that has lived in your head since the day you built the first workflow.

Here are the automations I run and what each one does:
[list them: name, platform, what it does, roughly how often it fires]

Help me write a monitoring contract. For each one, give me a row with:

1. Expected cadence, stated as a rule a clock can check
   ("by 06:00 daily", "at least once every 3 business days").

2. Grace window: how late is late, before anyone should be told.
   Size it to what the job costs when it's late, not to how long it takes.

3. Minimum plausible volume per run, and whether zero is ever normal.
   If zero is normal on some days, say which days.

4. The difference between "didn't run" and "ran and found nothing"
   for this specific workflow, and whether those are the same severity.

5. Where the output lands, so a human can verify independently.

6. Who to tell, and how loud.

Then do two things:

a) Flag any workflow where I cannot state a cadence, because those are
   the ones no clock can check and they need a different approach.

b) Rank them by what the first 24 hours of silence actually costs,
   and tell me which three are worth monitoring if I only do three.

Any of the big chat tools will handle it, so use whichever one you already have open. Part (b) is the one that earns its keep. Nine workflows feels like a project; three feels like this afternoon, and the three it picks will cover most of the real exposure. Writing the contract down before building anything is the same discipline as writing the workflow spec before you build, and it pays off the same way: the hard thinking is already done when you sit down to make it run.

The Script

This is an API job with no judgment in the hot path, so the honest build has no AI in it at runtime. Comparing a timestamp to a threshold is arithmetic, and it should be boring, free, and incapable of being creative at 3am. Put a model in here and you've bought a bill, a latency budget, and a monitor that can hallucinate a reassurance.

The one design choice worth arguing about is where the data comes from. You can poll each platform's run history, and the APIs exist: Zapier exposes Zap history with per-run statuses, and Make keeps a scenario history you can read. But you'll write a different integration per platform, you'll inherit their retention limits (Zapier guarantees sixty days), and you'll still be asking "what ran" when your question is "what didn't." The heartbeat inverts it. Add one HTTP call as the final step of every workflow, and now every tool you use reports in the same format and the watchdog is one small script forever.

// pseudo-shape of the watchdog BYOBot generates for you
// Each workflow's LAST step is: POST /beat { job, count }
const CONTRACT = readFile('monitoring-contract.json');   // from the prompt build
const now = clock.now();

for (const job of CONTRACT.jobs) {
  const last = beats.mostRecent(job.name);               // may be undefined

  // 1. SILENCE: the check a run-based alert structurally cannot make
  const due = job.nextDueAfter(last?.at ?? job.watchingSince);
  if (now > due + job.grace) {
    raise('SILENT', job, `last beat ${ago(last?.at)}, expected by ${fmt(due)}`);
    continue;                        // don't also complain about volume
  }

  // 2. RAN BUT EMPTY: alive workflow, dead input. Different failure.
  if (last.count === 0 && !job.zeroIsNormal(last.at)) {
    job.emptyStreak += 1;
    if (job.emptyStreak >= job.emptyStreakLimit) raise('NO_INPUT', job, last);
  } else {
    job.emptyStreak = 0;             // reset, or every quiet Sunday pages you
  }

  // 3. THIN: ran, found work, found less than this job has ever found
  if (last.count > 0 && last.count < job.minPlausible) raise('THIN', job, last);
}

// 4. PROVE YOU RAN. The watchdog's own heartbeat, sent elsewhere.
ping(DEAD_MANS_SWITCH_URL);

Steps one and four are the whole value. Step one is the check your platform can't make for you, because it's the only one framed around an event that didn't happen. Step four is what stops you from building a second thing that can die quietly. Note that the continue on line nine matters more than it looks: a silent job shouldn't also generate a volume alert, or your first real outage arrives as four confusing messages instead of one clear one. You don't write this from a blank page. BYOBot generates the full version with your contract, your alert routing, and the heartbeat snippet to paste into each workflow, which is where it starts paying for itself to automate the handoffs you keep checking on as a set rather than one at a time.

The Schedule

Putting the watchdog on a timer is where this loop is finally solved, and it's a few minutes of work: a cron entry, a GitHub Actions workflow on the free tier, or a scheduled scenario in whatever tool you already pay for. Run it hourly if you can. The cadence of the watchdog sets your floor on detection time, so hourly turns Dana's nine days into something she hears about the same morning.

Then close the loop on the watchdog itself, which is the ten minutes most people skip. Point step four at a hosted dead man's switch: Healthchecks.io and services like it exist for exactly this and invert the logic properly. Your script pings them on every run, and they alert you when a ping fails to arrive. The monitoring now lives on somebody else's infrastructure, so the scenario where your watchdog and your workflows die together stops being possible. That's the difference between an agent you check on and one that's genuinely running on its own while you do something else.

Here's how the three builds stack up, so you can pick your stopping point:

Build What runs it Setup Runs unattended?
The prompt You, interviewed into a monitoring contract by any AI chat tool Minutes No
The script A watchdog with no AI in it, comparing heartbeats against the contract An afternoon, once On demand
The schedule The same watchdog hourly, behind a hosted dead man's switch Ten extra minutes Yes

Steal This Build

Five lines, and they're the whole build. Copy them, swap in your own workflows, hand them to BYOBot or write them yourself this afternoon.

  • Trigger: a clock, never an event. The check has to run on the days your workflows don't.
  • Beat: one HTTP call as the final step of every workflow, carrying its name and how much work it just did.
  • Compare: expected cadence plus a grace window sized to what lateness costs, and a volume baseline that knows which days zero is normal.
  • Separate: "didn't run" and "ran and found nothing" are different failures. Different message, often a different person, never the same alert.
  • Prove: the watchdog reports its own heartbeat to a hosted switch, so the monitor cannot die the way the workflows did.

Think of this as the mirror image of the automation that breaks when the site changes, where the run at least fails loudly enough for somebody to see red. Here nothing goes red, so you have to go looking. The loop you just watched is nine no-code automations at a small SaaS company, but the shape is universal. Anything you expect to happen on a rhythm wants the same inversion: a backup that writes nightly, a report that lands on Monday, a sync that holds two systems level, a scraper that reads a page you don't own. In every one of those cases your tooling is built to describe failures and has no vocabulary at all for absence, so the check has to be driven by a clock and owned by you rather than by the platform. Learn that once and "it stopped a week ago and nobody knew" turns into an alert on the hour, which is what it looks like to automate a recurring workflow properly the first time.

Build Your Version

Tell BYOBot about a loop in your week

Describe the task you keep doing by hand and BYOBot will design the full playbook: the prompt, the script, and the schedule that runs it for you.

Frequently Asked Questions

  • Because most alerting is built on runs, and a workflow that stopped has no runs to report on. Platforms will email you when a step errors, and Zapier will auto-pause a Zap after repeated errors, but a trigger that quietly returns nothing is not an error. From the platform's point of view nothing went wrong, so there is nothing to send. Absence is the one event a run-based alert cannot describe, which is why the check has to live outside the workflow and be driven by a clock rather than by a trigger. That shift is the single idea behind everything else here, and it's the first thing to set up when you automate a job the business actually depends on.
  • A heartbeat is a one line HTTP call that your workflow makes as its very last step, saying it finished. A watchdog then checks the clock instead of the logs: if the expected ping did not arrive inside its grace window, something is wrong. This beats polling run history for two reasons. It works the same way across every tool you use, so you need one watchdog rather than one per platform, and it survives log retention limits, since a log you cannot read is a check you cannot make. It also means the monitoring doesn't care which tool you build in, which is useful if you ever want to move a workflow to a different tool without rebuilding the safety net.
  • Separate the two things you are measuring. The heartbeat says the workflow ran, which should happen on schedule whether or not there was any work to do, so a missing heartbeat is always worth an alert. Volume says how much work it found, and zero is often legitimate: no orders on Sunday, no tickets on a holiday. Give volume a baseline per day of the week and require a run of consecutive days before it escalates, and have it raise a question rather than an outage. The heartbeat is binary and loud. Volume is statistical and quiet. Get that split wrong and you'll mute the channel inside a month, which is a worse position than having no monitor at all.
  • A hosted dead man's switch, which is the one piece worth not building yourself. Services like Healthchecks.io invert the logic: your watchdog pings them, and they alert you when the ping fails to arrive. That means the thing responsible for noticing silence is itself monitored by something running on different infrastructure, so the failure mode where your monitoring and your workflows die together on the same box stops being possible. It is the cheapest insurance in the whole build and it takes about ten minutes.
BYOBot Autopilot
BYOBot Autopilot
Automated AI publishing system · editorial rules by Luke Grace LinkedIn →

This article has been published in an automated fashion with fully AI-written copy. These articles are meant to curate AI news from around the globe and bring a fresh perspective to using AI tools to accomplish big things. No person reviewed this specific piece before it went live, so check anything that matters against the sources linked above. Luke Grace sets the rules the system writes to. He's an algorithms and natural language expert with over 13 years experience and the creator behind BYOBot, the Build Your Own Bot platform that helps anyone build a multi-tasking agent to take over their repetitive tasks. For consulting help or more advanced AI workflow orchestration, you can reach Luke on LinkedIn.