← All incidents

minorinfra-failureunverifiedcollected

An unattended transcription job died with "user cancelled", so the mechanical work was taken away from the LLM (unverified report)

A scheduled pipeline that turns interview recordings into meeting notes stopped twice in a row one morning with "user cancelled" before transcription began. It ran unattended; there was no human who could have cancelled it.

Observed
Severity score
3/10
Blast radius
uptime
Tags
#codex#claude-code#mcp#failover#transcription#scheduled-job

Cause

That morning Claude's usage quota had run out and the failover handed the job to Codex. Through codex exec, listing the MCP server's tools worked, but actually calling the transcription tool failed with "user cancelled MCP tool call". Calling the same server with the same tool and arguments directly, without Codex, succeeded. The author suspected the approval path but does not claim to know the internal cause.

Consequence

Transcription never started in either run and the pipeline stopped. Setting approval_policy="never" or changing the sandbox setting made no difference, and the author found no setting in their Codex CLI as of July 2026 that opened this path.

Fix

Instead of fixing the Codex configuration, the author turned the roughly 50-line MCP client written for the diagnosis into a deterministic transcription script. The LLM now receives only the finished transcript and handles speaker inference and structuring. Retries now redo only the broken chunks, and completed transcripts stay on disk so a rerun can skip them.

What happened

The author ran an automated pipeline, as a scheduled hidden process, that turns interview recordings into meeting notes. ffmpeg splits the recording into 30-second chunks, an agent CLI (Claude Code or Codex) calls a local Whisper server (Voicebox) over MCP once per chunk, and then structures the result into notes. One morning the job stopped twice in a row with “user cancelled”, before transcription had even started.

The chaos on the ground

No human was there to cancel anything, yet the log said the user had stopped it. The run ledger showed that Claude had been skipped that morning because its quota was exhausted, and Codex had taken over; only the Codex runs failed. In a minimal reproduction, asking Codex to list the tools worked, but having it call voicebox.transcribe failed with “user cancelled MCP tool call”. Setting approval_policy="never" or changing the sandbox setting gave the same result.

Root cause

The author took Codex out and sent the same arguments to the same server directly over MCP’s streamable HTTP. It succeeded immediately. That ruled out a fault in Voicebox itself and showed the problem depended on the call path through Codex. The author’s hypothesis was that non-interactive codex exec, with no one to confirm anything, was colliding with its approval path, but since the internals could not be observed, the cause was not pinned down any further.

The fix

The author stopped digging into the Codex configuration. Even if it worked, the pipeline would still ask an LLM, every time, to do a job with no judgment in it: pass 30-second audio files to a tool in order and join the results. That setup has other ways to break: chunks can be skipped without anyone noticing, the model can stop partway and say it will omit the rest, and each provider handles MCP connections and approvals differently. So the roughly 50-line MCP client written for the diagnosis became a dedicated transcription script, and the LLM now receives only the finished transcript. Because the LLM no longer calls MCP at all, this failure cannot recur whichever CLI the failover picks.

Mechanical work that needs no judgment is easier to check and to rerun when it is plain code rather than an agent’s tool call.

Sources