The circuit breaker
If an agent’s calls keep failing at the platform end, AgentValet pauses it. For five minutes its calls are refused with circuit_breaker_open. After that, the next call goes through as a probe: if it succeeds the breaker closes, if it fails the breaker re-opens for another five minutes.
The breaker is separate from the agent’s status. Tripping it does not suspend the agent — the agent stays active in the dashboard, and suspension remains something only you, a policy, or anomaly detection does. The breaker exists to stop an agent hammering a platform that is rejecting it, and it recovers on its own.
When it trips
The breaker is kept per agent, across all of that agent’s platforms and connections. Two kinds of failure count:
- Auth failure — the platform rejected the stored credential, AgentValet tried to refresh it, and the refresh didn’t fix it. The agent gets
reauth_failed. Trips at 3. - Upstream failure — the call couldn’t be completed at all: a network error, a timeout, or the credential vault unreachable. The agent gets a
502. Trips at 5.
Both feed one failure counter. The counter only survives while failures keep coming: if more than 10 minutes pass since the last failure, the next one starts the count again at 1.
What does not count:
- Bad signatures, expired JWTs, or a wrong key — those are refused before a call is attempted
- Permission denials (scope not granted, blocked by policy or guardrail, agent suspended or revoked)
pending_approvalreturns — the call was queued, not failed- Upstream error responses such as a
404or422— they’re relayed to the agent as-is. A401/403marks the connection as needing re-auth, but only counts once a refresh has also failed.
When it resets
Any successful call (a 2xx from the platform, including one that succeeded after a token refresh) resets the counter to zero and closes the breaker from any state.
The three states
| State | What happens to calls |
|---|---|
| Closed | Normal. Failures are counted. |
| Open | Every call is refused with circuit_breaker_open until 5 minutes have passed since it opened. |
| Half-open | The cooldown has elapsed. Calls go through as probes; the first success closes the breaker, a counted failure re-opens it for another 5 minutes. |
What the agent sees
While the breaker is open, calls return 403 with "error": "agent_halted", "reason": "circuit_breaker_open" and a halt object:
state: "breaker_open",recoverable_by: "time"cause—breaker_auth_failureorbreaker_upstream_failureretry_after_seconds— time left on the cooldown (at least 5 seconds)instructions— plain text telling the model to pause platform calls, tell its user why, wait, then retry one call as a probe
The agent doesn’t need you to bring it back. Waiting out the cooldown and probing is the intended recovery.
In the dashboard
Agents → [the agent] shows a Circuit Breaker card: the current state, a countdown while it’s open, the recent failure count, and the agent’s last few denied calls with a link to the audit log.
Reset circuit breaker (then Confirm) closes the breaker and zeroes the counter immediately, and writes a circuit_breaker_reset audit row. Use it once you’ve fixed the cause and don’t want to wait out the cooldown. Suspending and re-activating the agent does not reset the breaker.
Trips and recoveries are recorded in the audit log as circuit_breaker.tripped (with the failure class and count) and circuit_breaker.recovered.
Finding the cause
Open the agent’s audit log and look at the failures before the trip:
reauth_failedon one platform — the connection’s credential is dead. Look for Re-auth needed on the platform and reconnect it.502s across calls — the platform or the network path to it was down. These usually clear on their own once the platform recovers.
Resetting without fixing a dead credential just trips the breaker again after three more calls.
The thresholds, window and cooldown aren’t configurable per agent.