CrowdControl.Provider behaviour (crowd_control v0.1.1)

Copy Markdown View Source

Behaviour for the infrastructure a CrowdControl.Backend.Sandboxd talks to.

A provider owns the lifecycle of a sandbox: acquiring one, handing back a reachable endpoint, reconnecting to one that outlived its session, releasing it, and enumerating the ones that are still alive. It owns nothing about bytes. CrowdControl.Backend.Sandboxd is the byte transport for every provider, because every provider runs the same in-sandbox agent (sandboxd) and speaks the same HTTP protocol to it.

That split is the point. CrowdControl.Backend already parameterizes where it provisions, but each substrate had to bring its own transport too — the Docker backend's FIFO/tee pair, the Kubernetes backend's exec stream. A VM has no exec API at all, so a fourth substrate meant a fourth transport. With one agent, a new substrate is provisioning code and nothing else.

Session --(Backend, unchanged)--> Backend.Sandboxd
                                      |
                        HTTP/1.1 + bearer token
                                      |
                                 sandboxd (in sandbox)

Backend.Sandboxd --(Provider, this module)--> Docker | Compose | Gce

Selecting a provider

CrowdControl.run("hello",
  backend:
    {CrowdControl.Backend.Sandboxd,
     provider: {CrowdControl.Provider.Docker, image: "crowd_control/sandbox:dev"}}
)

The {module, config} form merges config into the backend opts before acquire/1 is called, exactly as CrowdControl.Backend.resolve/1 does. Unlike :backend, :provider has no default: a provider decides where untrusted model-driven code runs, and guessing that is not a service this library provides.

Three load-bearing contracts

1. acquire/1 returns only when the agent has answered health

Not when the API call succeeded, not when the container is "created", not when the operation is DONE. Provisioning that reports success before the agent answers is the single largest source of flaky remote backends, and it is the default behaviour of the substrates underneath: gcp_compute's insert_and_wait/3 waits for the operation, never for the guest. Every implementation must poll GET /v1/health until it answers 200 or :ready_timeout elapses.

A failed acquire/1 must release whatever it created before returning. A leaked container is untidy; a leaked spot VM bills forever.

2. release/1 is idempotent, and "already gone" is success

CrowdControl.Session calls destroy/1 from both handle_cast(:eof, _) and terminate/2, and both can run for one session. A 404 from the substrate means the sandbox is gone, which is precisely what the caller asked for. Returning an error there turns a tidy shutdown into a crash.

3. The endpoint is never persisted

handle/0 goes into CrowdControl.Store and must survive :erlang.term_to_binary/1. CrowdControl.Provider.Endpoint.t/0 must not: it holds a derived token, a live tunnel pid, and a base_url whose port is assigned per-connection. Persisting it would make reattach fail after a node restart in the most confusing way available — a stale port that belongs to some other process.

So the handle persists the resource ({instance_name, zone}, {container_id}, {project_name}) and reconnect/1 rebuilds the path. The token is never persisted either; token/1 re-derives it from the session_key that CrowdControl.Store already holds.

Token derivation

token = Base.url_encode64(:crypto.mac(:hmac, :sha256, secret, session_key), padding: false)

where secret is config :crowd_control, :sandboxd_secret. See token/1. Nothing secret is written to disk, and reattach recomputes the token from the persisted session key. Rotating :sandboxd_secret therefore invalidates every live sandbox's token, and reattach fails closed with {:error, {:sandboxd, :unauthorized}}. That is the intended trade: the alternative is a live credential at rest in DETS.

What a provider must not do

Infer a security posture. If a sidecar needs egress, the caller says so; if a network must be reachable, the caller names it. CrowdControl.Backend.Docker already refuses to guess :network_mode (returning {:error, {:docker, :network_mode_required}}) and providers inherit that discipline.

Adding a provider: the Kubernetes mapping

CrowdControl.Provider.Kubernetes is deliberately not shipped yet, but the behaviour is graded on admitting it as provisioning code only. The mapping:

CallbackKubernetes implementation
acquire/1POST /api/v1/namespaces/{ns}/pods with sandboxd as the container command and the token in an env var, reusing Backend.Kubernetes' existing manifest hardening → poll GET …/pods/{name} for status.phase == "Running" → build the endpoint → poll GET /v1/health
reconnect/1rebuild the endpoint for the persisted {namespace, pod_name}. Nothing is re-created; only the path is
release/1DELETE …/pods/{name}, 404 = success
list_live/1GET …/pods?labelSelector=crowd-control-owner-hash=…, hand-paginated through metadata.continue exactly as Backend.Kubernetes.API.list_all/3 already does, because a truncated page makes the reaper prune live sandboxes
age_ms/1now - metadata.creationTimestamp
scrub/1keep {namespace, pod_name, session_key, owner}; drop the kubeconfig and every Req option

What writing that table changed about this behaviour

The reachability row did not fit, and the behaviour was wrong until it did.

A pod's agent port is not routable, so the endpoint has to be either a websocket port-forward (/portforward subresource, v4.channel.k8s.io — a transport, not ~200 lines) or the API server's pod proxy (GET …/pods/{name}:{port}/proxy/v1/health — a plain HTTP path, which is the ~200-line option). But the pod proxy consumes the authorization header for its own authentication, so a single token field cannot carry both credentials.

Hence CrowdControl.Provider.Endpoint.t/0 carries headers and req_options alongside token: a provider may state exactly how its transport is authenticated and configured, and Backend.Sandboxd.API merges those over its defaults. The one-header protocol assumption survives for Docker, Compose and GCE, where the agent is reached directly through a loopback port and authorization is free.

Summary

Types

Provider-opaque sandbox handle.

Callbacks

Create a sandbox and return a handle plus a reachable endpoint.

Age of the sandbox in milliseconds, or nil if unknown.

Every sandbox this provider can still see, owner-scoped.

Rebuild the endpoint for an existing sandbox.

Destroy the sandbox. Must be idempotent; already-gone is success.

Strip everything from handle that must not be persisted.

Functions

Age of handle via the provider's age_ms/1, or nil if it defines none.

Resolve the :provider option into {module, opts}.

Scrub handle via the provider's scrub/1, if it defines one.

Derive the agent token for session_key.

Types

handle()

@type handle() :: term()

Provider-opaque sandbox handle.

Persisted by CrowdControl.Store, so it must survive :erlang.term_to_binary/1 and must contain no credential, no live pid, and no ephemeral path. See scrub/1.

Callbacks

acquire(opts)

@callback acquire(opts :: keyword()) ::
  {:ok, handle(), CrowdControl.Provider.Endpoint.t()} | {:error, term()}

Create a sandbox and return a handle plus a reachable endpoint.

Must not return until GET /v1/health has answered 200, and must release anything it created if it cannot get there.

age_ms(handle)

(optional)
@callback age_ms(handle()) :: non_neg_integer() | nil

Age of the sandbox in milliseconds, or nil if unknown.

CrowdControl.Reaper uses it to spare sandboxes still inside the reap grace period: a sandbox younger than the grace window may belong to a session that has provisioned but not yet written its store record, and destroying it would be a race the caller cannot win.

Optional in the behaviour, mandatory in practice

Omitting it does not mean "no grace period". It means orphans are never collected at all, and the chain is worth spelling out because no single link looks wrong:

CrowdControl.Backend.Sandboxd exports age_ms/1, so the reaper always consults it → age_ms/2 returns nil for a provider that defines no callback → the reaper reads an unknown age as "too young to reap", because fail-open is the right default there (a missed reap costs one sweep interval; a wrong reap costs a live session).

So a provider without this callback leaks every orphan forever. Implement it. Both shipped Docker-shaped providers read it from the crowd_control.created_at label they set at create time.

list_live(opts)

@callback list_live(opts :: keyword()) :: {:ok, [handle()]} | {:error, term()}

Every sandbox this provider can still see, owner-scoped.

Backs CrowdControl.Reaper. Must paginate exhaustively: a truncated list makes the reaper treat live sandboxes as dead records and prune them.

reconnect(handle)

@callback reconnect(handle()) ::
  {:ok, handle(), CrowdControl.Provider.Endpoint.t()} | {:error, term()}

Rebuild the endpoint for an existing sandbox.

Called on reattach, where the handle came out of CrowdControl.Store and the endpoint did not exist. Returns a possibly-updated handle so a provider can refresh substrate state it learned while reconnecting.

release(handle)

@callback release(handle()) :: :ok

Destroy the sandbox. Must be idempotent; already-gone is success.

scrub(handle)

(optional)
@callback scrub(handle()) :: handle()

Strip everything from handle that must not be persisted.

Optional; the default is the handle untouched. Implement it whenever the handle carries substrate configuration, which for at least one provider holds a live token-provider argument.

Functions

age_ms(module, handle)

@spec age_ms(module(), handle()) :: non_neg_integer() | nil

Age of handle via the provider's age_ms/1, or nil if it defines none.

resolve(opts)

@spec resolve(keyword()) :: {module(), keyword()}

Resolve the :provider option into {module, opts}.

Accepts a bare module or a {module, config} tuple, merging config into the remaining opts. Mirrors CrowdControl.Backend.resolve/1 with one deliberate difference: there is no default provider.

iex> CrowdControl.Provider.resolve(provider: {CrowdControl.Provider.Docker, image: "x"})
{CrowdControl.Provider.Docker, [image: "x"]}

scrub(module, handle)

@spec scrub(module(), handle()) :: handle()

Scrub handle via the provider's scrub/1, if it defines one.

Returns the handle untouched for providers that do not.

token(session_key)

@spec token(String.t()) :: String.t()

Derive the agent token for session_key.

:sandboxd_secret must be configured; it is deliberately not defaulted or auto-generated, because a per-boot secret would silently break reattach across a node restart — the one thing the derivation exists to support.

config :crowd_control, sandboxd_secret: System.fetch_env!("CC_SANDBOXD_SECRET")