The circuit breaker is what protects your application from a persistently failing dependency. Where retries handle the occasional blip, the breaker handles the outage: once a service fails too often, the breaker "opens" and further calls fail fast — immediately, without touching the struggling service — until it has had time to recover.
This is the mechanism described in Michael Nygard's Release It! and popularized
by Martin Fowler. ExternalService implements it on top of the Erlang
:fuse library, but you never call :fuse
directly — the breaker is managed for you on every call.
Why fail fast?
When a dependency is down and you keep calling it, every caller blocks on timeouts, work piles up, and the failure spreads — a cascading failure. The breaker short-circuits that: after enough failures it stops you from even attempting the call, so callers get an immediate error they can handle (serve cached data, degrade gracefully, return 503) instead of hanging.
Unlike retries, which are per-call, the breaker is global to the service. If it trips, it trips for every caller in the system at once. That is precisely what makes it effective at preventing cascades.
Configuration
Configure the breaker with the :circuit_breaker option to use ExternalService or ExternalService.start/2:
use ExternalService,
circuit_breaker: [
tolerate: 5, # failures allowed within the window...
within: :timer.seconds(1), # ...this window, in milliseconds
reset: :timer.seconds(5) # stay open this long before resetting
]| Option | Default | Meaning |
|---|---|---|
:tolerate | 10 | Number of failed attempts tolerated within the :within window before the breaker opens. :infinity never opens. |
:within | 10_000 | Length of the failure-counting window, in milliseconds. |
:reset | 60_000 | Milliseconds to wait before the breaker resets (closes) after opening. |
:fault_injection | — | If set to a rate between 0.0 and 1.0, randomly fails that fraction of calls (for testing). |
So tolerate: 5, within: 1_000 means "open the breaker once there are more than
5 failed attempts inside any 1-second window." After opening, the breaker stays open
for :reset milliseconds, then closes again and calls resume under the same
monitoring.
The :circuit_breaker option (and every key within it) is optional. Omit it to
get the defaults above.
What counts as a failure?
The breaker is "melted" — pushed one step toward opening — on every call attempt that fails, where a failure is:
- the function returns
:retryor{:retry, reason}, or - the function returns a value matched by the
:retry_onpredicate, or - the function raises an exception whose type is listed in the
:retry_exceptionsretry option.
Melt and retry go together for exceptions
The :retry_exceptions retry option governs both whether a raised exception
is retried and whether it melts the breaker. An exception whose type is in
:retry_exceptions is retried and melts the breaker; an exception that is not
in :retry_exceptions is neither retried nor melted — it propagates to the
caller and leaves the breaker untouched.
Explicit :retry / {:retry, reason} return values, and results matched by the
:retry_on predicate, always melt the breaker — they are ways of asking for
another attempt.
Values your function simply returns — including its own {:error, reason} — are
successes as far as the breaker is concerned and do not melt it.
:tolerate counts attempts, not calls
The word doing the work above is attempt. :tolerate reads like "how many
failed calls before the breaker opens," but retries melt too: one call/3 with
max_attempts: 5 that fails throughout contributes 5 melts by itself.
So :tolerate and :max_attempts cannot be tuned independently. Measured with
this configuration:
use ExternalService,
circuit_breaker: [tolerate: 10, within: :timer.seconds(10)],
retry: [max_attempts: 5]failing call #1: returned RetriesExhausted, blown? false
failing call #2: returned RetriesExhausted, blown? false
failing call #3: returned CircuitBreakerOpen, blown? trueA breaker that reads as "open after 10 failures" opens during the third
failing call. (:fuse tolerates :tolerate melts and opens on the next, so
tolerate: 10 opens on the 11th melt; three calls × 5 attempts = 15 melts,
crossing 11 partway through the third.) On a service handling any real
concurrency that is a much shorter fuse than the numbers suggest — and the
breaker is global, so those calls fail everything else too.
Budget in attempts: pick :tolerate as roughly failing calls you'll accept
× :max_attempts. The Retries guide covers the same interaction
from the other side — why the breaker is not a reliable bound on retries.
When the breaker is open
A call made while the breaker is open does not invoke your function at all. Instead:
call/3returns{:error, %ExternalService.CircuitBreakerOpen{}},call!/3raisesExternalService.CircuitBreakerOpen, and- an
[:external_service, :circuit_breaker, :blown]telemetry event is emitted.
See Error handling for how to deal with these.
Introspecting and resetting
You can ask about the breaker's state at any time. With the module front door:
MyApp.Stripe.available?() #=> true when the breaker is closed
MyApp.Stripe.blown?() #=> true when the breaker is open
MyApp.Stripe.reset() #=> force the breaker closedOr with the functional API:
ExternalService.available?(:payments)
ExternalService.blown?(:payments)
ExternalService.all_available?([:payments, :inventory])
ExternalService.reset(:payments)A few semantics worth knowing:
available?/1istrueonly when the breaker is closed. A service that was never started reportsfalse— it is not "ready to use."blown?/1is the direct "is it open?" question. A service that was never started is not reported as blown (there is no breaker to be open); useavailable?/1when you want "ready to use" semantics.all_available?/1istrueonly if every listed service isavailable?/1— handy for guarding work that depends on several services.- Availability can change between the check and a subsequent call, so treat
these as best-effort signals, not guarantees. They let you bail out early;
they do not replace handling a
CircuitBreakerOpenerror from the call itself.
reset/1 forces the breaker closed immediately, discarding its recorded
failures. It is mainly useful in tests and in operational tooling ("we fixed the
upstream, stop failing fast now").
It resets only the breaker. A service's rate limiter is separate state, and
clearing it releases a burst at the service — rarely what someone closing a
breaker intended. When you do want both, ExternalService.reset_all/1 clears the
breaker and the limiter together:
ExternalService.reset_all(:payments)
MyApp.Stripe.reset_all()That is usually what a test setup block wants; see the Testing
guide.
Reporting a failure the library never saw
The breaker counts failures that happen inside call/3. Sometimes a service
fails somewhere else — a long-lived streaming connection drops, a webhook you
were expecting never arrives, a request made through a different client times
out. Those are real failures, and you can tell the breaker about them:
ExternalService.CircuitBreaker.melt(:payments)A melt counts toward the service's configured :tolerate exactly as an in-call
failure does, so enough of them will open the breaker — and with the
cluster breaker, open it across the cluster. This is the
mirror image of reset/1: one forces the breaker closed, the other pushes it
toward open.
Use it sparingly and only for genuine failures of the service. Melting on something that is not the service's fault will fail-fast traffic that would have succeeded.
When the service hangs
The breaker protects you against a service that fails. It has no answer for one that hangs.
If your function blocks, call/3 blocks with it. No failure has been observed,
so nothing melts the breaker, no retry is attempted, and no :stop telemetry is
emitted. Measured with tolerate: 1 while a slow call is in flight:
available?: true
blown?: falseThat is correct behavior — the call hasn't failed, it just hasn't finished — but it is the opposite of what "circuit breaker" leads people to expect. The pattern as usually described pairs a breaker with a timeout, and the timeout is what converts a hang into the failure the breaker counts. Without one, the most common real degradation — service up, responses crawling — is invisible: latency climbs, processes pile up, and nothing trips.
ExternalService does not impose a timeout
There is no :timeout option, and no retry option supplies one.
:expiry is evaluated between attempts,
so it cannot interrupt an attempt already running. Bounding a single attempt
is your responsibility, at the client.
In practice this is where it belongs, because your HTTP client already has the timeout that matters — it can actually abandon the socket, which the library could not do from the outside:
# Req
Req.get(url, receive_timeout: :timer.seconds(5))
# Finch
Finch.request(req, MyFinch, receive_timeout: :timer.seconds(5), pool_timeout: :timer.seconds(1))
# Tesla / Hackney
Tesla.get(client, path, opts: [adapter: [recv_timeout: :timer.seconds(5)]])Set the pool checkout timeout too, not just the receive timeout. Under saturation you block waiting for a connection before you ever send a request, and only the checkout timeout bounds that.
Once the client times out, the wrapped function returns or raises — which is an observable failure, so it melts the breaker and can be retried like any other:
def fetch(id) do
call fn ->
case Req.get(url(id), receive_timeout: :timer.seconds(5)) do
{:ok, %{status: status} = resp} when status < 500 -> {:ok, resp}
# Timeouts and 5xx are worth another attempt; both melt the breaker.
_ -> :retry
end
end
endWhy the library doesn't do this for you: running each attempt in a Task would
put a process on the hot path, which is the cost this library exists to avoid,
and it would not help as much as it appears. Killing the task stops you waiting
for the socket but does not cancel the request — the connection can stay checked
out, so the timeout masks saturation rather than relieving it. It would also move
your function off the calling process, so anything it reads from there — Logger
metadata, OpenTelemetry context, Ecto sandbox ownership — would silently change.
Note that a timeout bounds each call but does not bound how many are in flight at once. For that, see Concurrency Limiting — and note the two go together, since a concurrency limit whose calls never finish just fills up.
What this library does not bound
One thing, and it is the one above:
| Not bounded | Where it belongs |
|---|---|
| How long one attempt may take | Your HTTP client's receive and pool-checkout timeouts. |
Everything else has a bound available: failures with :tolerate, attempts with
:max_attempts/:expiry, call rate with :rate_limit, and calls in flight
with :concurrency.
The last two are easy to confuse. The rate limiter bounds how often calls
start, which is not the same as how many are running: at limit: 100, per: 1_000
against a service that slows to 10 seconds per call, roughly a thousand processes
end up parked in the same call, each holding a connection. That is what a
concurrency limit caps, and the breaker only opens once things
actually start failing — by which point the pressure is upstream of it.
Fault injection (for testing)
The :fault_injection option makes the breaker fail a random fraction of calls,
which is useful for exercising your own fallback and error-handling paths:
use ExternalService,
circuit_breaker: [tolerate: 5, within: 1_000, fault_injection: 0.25]This is a testing aid — leave it unset in production.
A breaker that never opens
tolerate: :infinity installs no breaker at all. Calls are never rejected,
melts are ignored, and the service holds no breaker state:
use ExternalService,
circuit_breaker: [tolerate: :infinity]Two situations want this.
In production, for a service where opening the breaker is worse than the
failures it would prevent — an idempotent write to a queue you'd rather keep
retrying, say — while you still want the retry and telemetry machinery. It says
that deliberately, where tolerate: 1_000_000 only says "not for a while".
In tests, because a breaker that holds no state cannot leak between them.
:tolerate is normally the awkward one: it is global to the service and nothing
resets it between tests, so a suite long enough to accumulate :tolerate melts
starts failing tests that have nothing to do with the breaker. A finite number
merely postpones that. See the Testing guide.
:infinity cannot be combined with :fault_injection — one promises the breaker
never opens and the other exists to open it — so start/2 raises rather than
letting either silently win:
** (ArgumentError) ExternalService.start(:payments, ...) sets both
circuit_breaker: [tolerate: :infinity] and :fault_injection, which contradict
each other.Choosing thresholds
There is no universally correct setting; it depends on the service's normal error rate and how costly a false trip is. Some rules of thumb:
- Set
:tolerate/:withinso the breaker tolerates normal transient noise but trips promptly on a real outage. Counting failures over a window (rather than consecutively) makes it robust to interleaved success and failure. - Set
:resetto roughly how long you expect a recovering service to need. Too short and you hammer a service that isn't ready; too long and you stay degraded after it has recovered. - Remember the breaker is global to the service. Size it for aggregate traffic, not a single caller.
Running on more than one node
"Global to the service" means global on this node. Each node counts its own failures and opens its own breaker, so in a cluster every node has to learn about an outage separately.
That is usually the behavior you want — a node with a bad network path should
stop calling the service without taking the rest of the cluster with it — but it
does mean slower convergence. ExternalService.CircuitBreaker.Cluster trades
that isolation for speed, propagating a trip to the other nodes:
use ExternalService,
circuit_breaker: [
tolerate: 5,
within: :timer.seconds(1),
backend: ExternalService.CircuitBreaker.Cluster
]See the Distributed Elixir guide for how it works and what the trade costs you.