Docs

Why isn’t the agent’s word enough?

Why the agent can’t close its own task

Because an agent reporting on its own work is the one report you cannot check by reading it.

The failure this prevents

You leave an agent running. You come back to “Done. All tests pass.” You run the tests. Two fail, and one suite never ran at all.

The agent was not lying. It was guessing — it had no way to know, and nothing in the loop checked. Every agent CLI can produce that sentence, and none of them can be trusted to produce it only when it is true, because the sentence and the work come from the same process.

What Muster does instead

When the agent hands the turn back, Muster runs a command you approved, in the project folder, outside the agent’s session, and reads the exit code. Zero closes the task. Anything else blocks it, with the output attached.

The agent’s summary plays no part. Not as a tiebreaker, not as a hint. If the agent ran the tests itself mid-task, that does not count either — only the run after it hands back does, because only that run happens on the final state of the files.

Why an exit code and not something cleverer

An exit code is the one signal in this pipeline that no participant can talk its way around. It is produced by a process the agent does not control, on the files as they are, after the work is over.

Anything softer — a model judging the output, a heuristic over the diff — puts a guess back in the position the guess just failed in.

What it costs you

A task with no acceptance command cannot be closed automatically. That is not a gap; it is the same rule seen from the other side. No command means no verdict, so the task stops and waits for you instead of pretending it finished.

If that seems strict, consider the alternative: a queue that closes tasks nobody checked is a queue whose green ticks mean nothing, and a green tick that means nothing is worse than no tick at all.