minorinfra-failureunverifiedcollected
An unattended transcription job died with "user cancelled", so the mechanical work was taken away from the LLM (unverified report)
A scheduled pipeline that turns interview recordings into meeting notes stopped twice in a row one morning with "user cancelled" before transcription began. It ran unattended; there was no human who could have cancelled it.
Cause
That morning Claude's usage quota had run out and the failover handed the job to Codex. Through codex exec, listing the MCP server's tools worked, but actually calling the transcription tool failed with "user cancelled MCP tool call". Calling the same server with the same tool and arguments directly, without Codex, succeeded. The author suspected the approval path but does not claim to know the internal cause.
Consequence
Transcription never started in either run and the pipeline stopped. Setting approval_policy="never" or changing the sandbox setting made no difference, and the author found no setting in their Codex CLI as of July 2026 that opened this path.
Fix
Instead of fixing the Codex configuration, the author turned the roughly 50-line MCP client written for the diagnosis into a deterministic transcription script. The LLM now receives only the finished transcript and handles speaker inference and structuring. Retries now redo only the broken chunks, and completed transcripts stay on disk so a rerun can skip them.
What happened
The author ran an automated pipeline, as a scheduled hidden process, that turns interview recordings into meeting notes. ffmpeg splits the recording into 30-second chunks, an agent CLI (Claude Code or Codex) calls a local Whisper server (Voicebox) over MCP once per chunk, and then structures the result into notes. One morning the job stopped twice in a row with “user cancelled”, before transcription had even started.
The chaos on the ground
No human was there to cancel anything, yet the log said the user had stopped
it. The run ledger showed that Claude had been skipped that morning because
its quota was exhausted, and Codex had taken over; only the Codex runs
failed. In a minimal reproduction, asking Codex to list the tools worked, but
having it call voicebox.transcribe failed with “user cancelled MCP tool
call”. Setting approval_policy="never" or changing the sandbox setting gave
the same result.
Root cause
The author took Codex out and sent the same arguments to the same server
directly over MCP’s streamable HTTP. It succeeded immediately. That ruled out
a fault in Voicebox itself and showed the problem depended on the call path
through Codex. The author’s hypothesis was that non-interactive codex exec,
with no one to confirm anything, was colliding with its approval path, but
since the internals could not be observed, the cause was not pinned down any
further.
The fix
The author stopped digging into the Codex configuration. Even if it worked, the pipeline would still ask an LLM, every time, to do a job with no judgment in it: pass 30-second audio files to a tool in order and join the results. That setup has other ways to break: chunks can be skipped without anyone noticing, the model can stop partway and say it will omit the rest, and each provider handles MCP connections and approvals differently. So the roughly 50-line MCP client written for the diagnosis became a dedicated transcription script, and the LLM now receives only the finished transcript. Because the LLM no longer calls MCP at all, this failure cannot recur whichever CLI the failover picks.
Mechanical work that needs no judgment is easier to check and to rerun when it is plain code rather than an agent’s tool call.