← All incidents

majormonitoring-gapunverifiedcollected

A scheduled Claude Code job did nothing for 2.5 days while GitHub Actions stayed green (unverified report)

A GitHub Actions job that ran Claude Code every four hours to revise blog articles produced nothing for eight runs in a row over 2.5 days. Every run showed green.

Observed
Severity score
5/10
Blast radius
uptime
Tags
#claude-code#github-actions#usage-limit#exit-code#silent-failure#watchdog

Cause

The weekly usage limit ran out, and Claude printed a one-line limit message. The step ended in || true, pipefail was off so the pipeline's status came from tee, and alerts only fired on red. The watchdog meant to report blocked runs also ran on Claude and the same quota, so the work and its monitoring stopped together.

Consequence

No article was revised for 2.5 days (September 18 to 21, 2026). The author found out by chance, opening a status file that held eight consecutive usage-limit entries. Two other Claude-based jobs hit the same limit that day, and dashboards looked calm because nothing was changing.

Fix

The author moved the pass/fail decision out of shell into tested Python that classifies real error messages: auth errors turn red, usage limits stay green but warn. Output is checked by pattern and minimum size, results are counted per run, 'zero results' is kept apart from 'could not check', and alerts go out through plain curl so they no longer share the quota.

What happened

A freelance engineer ran Claude Code on a schedule: every four hours, a GitHub Actions workflow on a self-hosted runner called claude -p to rewrite blog articles for better search rankings. From September 18 to 21, 2026, the job ran eight times and changed nothing. All eight runs were green.

The chaos on the ground

Nobody was told. The author stumbled on the problem while opening a JSON file that was supposed to record blocked runs, and found eight entries in a row marked as a usage limit. The same day, two other jobs that relied on Claude, a weekly meeting summary and a review job, had also hit the limit. From the outside everything looked stable, which in hindsight was the symptom: nothing was changing because nothing was running.

Root cause

The weekly usage limit had run out, and Claude answered each call with a single line saying so. Three habits turned that into a green check. The command ended in || true, which throws away the exit code. The shell did not set pipefail, so the pipeline reported the status of tee, which succeeds. And notifications only went out when a run turned red.

There was a watchdog meant to catch exactly this, reading the status file and reporting problems. But it also used Claude, from the same weekly pool, so it went quiet at the same moment as the job it was watching.

The fix

The author took the decision about success and failure out of shell and into Python, where it can be tested against real error messages. Auth errors now fail the run; usage limits keep it green but raise a warning. Output has to match an expected pattern and exceed a minimum size, each run counts what it actually produced, and “zero results” is reported separately from “could not check”. Alerts are sent with plain curl, so a quota problem can no longer silence the alarm about the quota problem.

The same lesson shows up in an earlier entry: a green check means the process exited cleanly, not that the work happened. And a monitor that shares a failure mode with the thing it watches is not a monitor.

Sources