← All incidents

majortool-misuseverifiedfirsthand

An AI agent meant to restart one local server and SIGTERMed nine processes instead

An AI coding agent (Claude Code) set out to restart a single local server and ended up sending SIGTERM to all nine processes that were talking on 127.0.0.1.

Observed
Severity score
5/10
Blast radius
uptime
Tags
#lsof#kill#macos#ai-coding#process-management

Cause

Passing two -i flags to lsof makes the conditions OR together, so the port filter stops doing anything. The result was piped straight into kill via kill $(...), leaving no moment where anyone looked at what had matched.

Consequence

A home-built app launcher, a home-built sales dashboard, two test servers belonging to other sessions, a login-item app (presumed), and three unidentified processes were stopped. The launcher and the dashboard were restarted and recovered.

Fix

Before killing anything, narrow it down to a single PID, print it, and check the command name and working directory by eye. Never feed search results directly into kill with kill $(...).

What happened

The request was simple: restart the server for a home-built transcription app. The AI coding agent (Claude Code) put together a one-liner to find whatever was listening on that port and stop it:

kill $(lsof -tiTCP:<port> -sTCP:LISTEN -a [email protected] || ...)

It looks careful. It names the port, restricts to LISTEN, and limits itself to loopback. But what actually received SIGTERM was not one process. It was nine — every process that happened to be talking on 127.0.0.1.

The chaos on the ground

The list of casualties doubled as the incident report:

  • a home-built app launcher (along with its forwarding feature)
  • a home-built sales dashboard
  • two servers that another session was using for testing
  • an app running as a login item (presumed)
  • three processes nobody could identify

The last two items are the uncomfortable part. Even afterward, it was not possible to say exactly what had been killed. Once “presumed” and “unknown” show up in the damage report, whoever pulled the trigger can no longer fully account for the blast radius. The launcher and the dashboard were restarted and came back; what happened to the rest is not on record.

And some of those processes belonged to a different session entirely. Someone else’s test environment vanished for reasons that had nothing to do with the work at hand.

Root cause

The immediate cause was a misreading of lsof. With two -i flags, each condition is OR’ed, so the port number no longer narrows anything down. What came back was “everything on 127.0.0.1.”

The second cause was shape: the result went straight into kill via kill $(...). Search and execution were fused into one line, so neither the human nor the agent ever saw how many PIDs matched or what they were. Anyone can get a search expression wrong. This pattern turns that mistake directly into damage.

Killing a process can’t be undone, and it can hit resident apps and other sessions. That kind of operation was running as a single line with no checkpoint in it.

The fix

  • Before killing anything, narrow it to one PID and print it. Check the command name and working directory before stopping it. Read the output of lsof -iTCP:<port> -sTCP:LISTEN -P with your own eyes.
  • Never hand search results straight to kill with kill $(...).
  • When something else (a forwarder, for example) also holds the same port, pick only the PID you actually mean to stop. A port number alone does not tell you who the target is.

This was recorded as a standing rule for the AI agent: don’t trust that the search expression is correct; always insert a step where each target is looked at before anything is executed.

Sources

  • Firsthand account from the site operator. There is no public write-up to link to.