← All incidents

minormonitoring-gapunverifiedcollected

15 changes deferred "for a human to review" piled up with no one on the receiving end (unverified report)

Scheduled Claude Code tasks were gradually adjusting the author's automation and logging each change in a ledger. When the author counted, all 15 changes were still awaiting a verdict, and the oldest was five weeks past its deadline. The run logs showed a check recorded every time.

Observed
Severity score
4/10
Blast radius
none
Tags
#claude-code#scheduled-tasks#human-in-the-loop#handoff#ledger#review

Cause

The unattended checker, following the rule that it must not grade its own changes, handed every verdict to the weekly review. But the weekly review's procedure did not mention the ledger once, and the sending side had no wiring for where the ledger should go either. The sender kept recording that it had handed things off; the receiver did not know it was the receiver.

Consequence

Nine of the 15 changed the decision logic itself and should have been kept or reverted the following week. Regression checks ran almost every time and found nothing worse, but confirming whether each change helped, and deciding to keep or revert it, had stalled. For one change, the scheduled task meant to judge it was merged into another task two days before the deadline and disappeared, so it was checked zero times. A pending item about the same kind of drift in another ledger had been stuck for 42 days.

Fix

Keeping the rule that the unattended side never grades itself, the author proposed two changes and tested them in a small reproduction: when handing off, check that the receiver has an intake for it and fail if not; and give each hold an expiry, warning once it is more than 14 days past its verdict date. The author also recommends regularly counting the holds and the age of the oldest one, and searching for ledgers that name a task whenever an automation is retired or merged.

What happened

The author had scheduled Claude Code tasks gradually tuning their own automation. Each change went into a ledger as one line, and on a set date it was checked for whether things had got better or worse. Changes to decision logic also carried a rule: if things had got worse, revert automatically. One day the author counted the ledger. All 15 changes were still awaiting a verdict, and the oldest was five weeks past its deadline.

The chaos on the ground

The run logs recorded every check properly. Most runs deferred with the same reasoning: an unattended run does not grade its own changes, so the final call goes to the weekly review. Nine of the 15 changed the decision logic itself and should have been kept or reverted the following week. Regression checks ran almost every time and found nothing worse; what had stalled was confirming the effect and deciding to keep or revert. For one change, the scheduled task that was supposed to judge it had been merged into another task two days before the deadline and no longer existed, so it was checked zero times before the deadline passed. A pending item about the same kind of drift in a different ledger had been stuck for 42 days.

Root cause

A full-text search of the weekly review’s procedure found no mention of the ledger at all, and the sending side’s procedure had no wiring for where the ledger should go. The sender recorded “handed off” every time (and the sending task itself was handed over to another task along the way), while the receiver did not know it was the receiver. Per the author, each individual deferral looked correct, and reading either side alone made it look as if the other side was handling it. When the receiver was merged away, nobody updated the ledger lines that named it.

The fix

The author kept the constraint that the unattended side never grades its own changes and proposed two changes. When handing off, check that the receiver’s procedure declares an intake for the ledger, and if it does not, fail instead of deferring silently. Give each hold an expiry, and once it is more than 14 days past its verdict date, raise it as a separate problem: the holds are piling up. The receiver gets its intake written the same day and lists held items oldest first. Running this in a small reproduction, the author confirmed that a hand-off to a receiver without an intake stops, and that held items now appear on the receiver’s side in order of age. To prevent a repeat, the author also recommends regularly counting the holds and the age of the oldest one, and searching for ledgers that name a task whenever an automation is retired or merged.

“Send it to a human” is a sound constraint, but if nobody on the other end has an intake, the decision never gets made. Write a delegation into both the sender and the receiver, and watch the size of the pile rather than the individual log lines.

Sources