> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fermata.run/llms.txt
> Use this file to discover all available pages before exploring further.

# When an agent fails

> What a failed agent's card tells you, the four ways forward, recovery for a whole halted run, and the per-piece policy that decides how much of this happens without you.

A failed agent does not sink the run. Its card names what happened and offers a way forward; the piece's worktree and branch keep every commit that landed before the failure, and the failed session keeps its transcript. What happens next is your call, made per agent on the card, per run on the action bar, or once per piece with the failure policy.

## What the card says

The card shows the failure text recorded on the session, and under it a kind label naming the cause. The kind is what decides which recovery makes sense.

| Kind | What it means |
| - | - |
| "Process crashed" | The Claude Code process died. Nothing to steer; re-run it. |
| "Agent never started" | The session never launched: a missing CLI, a worktree that is gone. Fix the cause first. |
| "Turn error · temporary" | The API refused the turn for a passing reason, like an overload. Repeating it can work. |
| "Turn error · content filter" | The response was blocked by a content filter. Repeating the same prompt verbatim will not help. |
| "Turn error · authentication" | The CLI could not authenticate. Fix your login, then retry. |
| "Turn error" | A turn error with no more specific class. |

A failure captured without a structured kind shows the reason alone; the card never guesses.

## The ways forward

* **Continue** leads the card, but only for a turn error whose session still holds a replayable transcript. It resumes that same session with a nudge to finish its work: the agent's own partial answer is still in the conversation, so steering it is cheaper and likelier to work than re-running the prompt that produced it.
* **Retry** runs the agent again cold, on a fresh session with the same brief.
* **Skip** moves past the agent and its dependents so the run can reach Review.
* **Mark Completed** accepts the agent's work as done without re-running it. This is for work that is already on the branch from an earlier attempt, where a retry finds nothing left to do and a skip would cascade through dependents that are not actually blocked.

**Open Agent View** opens the failed session's transcript; it stays available after the run ends, whatever you decided.

## Previous attempts

A retry re-arms the agent onto a fresh session, and the attempt it replaced stays one click away: the card links **Previous attempt**, or offers a **Previous attempts (N)** menu with the newest first. The evidence for why the first run failed is exactly what you want open while the second one runs.

## When the run halts

With failures outstanding, the action bar leads with recovery: **Retry All Failed (N)** re-runs every failed agent, and **Continue to Review (N failed) →** advances anyway, skipping whatever never finished. Advancing with agents unfinished is always reachable, but it is the escape hatch, not the recommendation; the verdict at [Review](/piece/review-and-pull-request) judges a run with holes in it.

## Run Again on a paused run

A paused run offers one more repair: a completed agent whose work never actually landed, like a turn that ended on a background wait, carries a **Run Again** link. It confirms first, because it re-runs work the run already counted as done. The dialog is titled "Run Agent Again?" and its confirm button reads "Run Task N Again"; the message says the agent "will be re-armed and runs again when the run continues", that "Its previous attempt stays in its history", and that "Tasks that depend on it and already finished are not re-run". If the review step already ran, it runs again after it.

## Inside the failed session

**Open Agent View** shows the same failure as a notice above the transcript, and the session itself is often the fastest way out.

* "The agent stopped", with a live composer. "Send a message to pick up where it left off." Your next message resumes the session; it is the same warm continue the card's button sends, with your own steer instead of a generic nudge.
* "The agent could not be started", when a launch failure hit a session that still holds its conversation. "The conversation is intact. Fix the cause, then try again." The **Try Again** button clears the recorded launch failure and hands the composer back; press it once you have fixed what stopped the launch.
* A session with nothing to replay says so: "There is no conversation to replay, so this session can't be continued. Start a new one."

## Decide it once: the failure policy

You do not have to be there for any of this. The "On Agent Failure" section of the Flow Configuration sheet, opened with **Configure Piece** in the piece inspector, sets what the run does on its own:

* **Halt & surface**: "Stop the branch and let me Retry or Skip the failed agent." The default; everything above is you resolving what it surfaced.
* **Skip dependents**: "Auto-skip the failed agent's dependents so the run reaches Review."
* **Auto-retry**: "Re-run a failed agent a few times before giving up." A **Max retries** stepper caps it, from 1 to 5.

The section locks while the run is live, like the rest of the sheet; the policy is a decision about the run, made before it. Where the sheet's other sections live is covered on [defaults](/control/defaults).

Whichever way the failure resolves, nothing is discarded along the way: completed agents, the branch, and the worktree are left untouched, and every attempt's transcript stays inspectable. The run carries what it has into [Review](/piece/review-and-pull-request).
