Spawn and manage Google Compute Engine instances from Elixir, over the Compute REST API. Pluggable auth, telemetry on every call, and an ergonomic builder for the (verbose) instance body.
[!WARNING]
spot: trueis the default. A bareGcpCompute.launch(config, "worker-1")provisions a Spot VM: cheap and preemptible, and when GCP preempts it the instance deletes itself along with everything on its disk (provisioningModel: "SPOT",automaticRestart: false,instanceTerminationAction: "DELETE"). That is deliberate — this library is built for disposable batch workers — but passspot: falsefor anything whose disk you care about.
Why REST and not gRPC? The Compute Engine API is REST/JSON only — it has no gRPC endpoint (it's the one notable exception among GCP APIs). So unlike
pubsub_grpc, there is no connection pool to run: a request is a handful of calls dominated by the 20–40s the VM takes to boot, where pooling buys nothing.GcpComputetalks to it withReq; aGcpCompute.Configis the only handle you pass around. The right place for gRPC +grpc_connection_poolis an agent running inside the machine you spawn — see Roadmap.
Installation
def deps do
[
{:gcp_compute, "~> 0.3.0"},
# Recommended token provider (optional dependency):
{:goth, "~> 1.4"}
]
endQuickstart
# 1. Start Goth to mint OAuth tokens (in your application supervision tree)
children = [{Goth, name: MyApp.Goth}]
# 2. Build a config once and reuse it
{:ok, config} =
GcpCompute.Config.production(
project: "my-project",
zone: "us-central1-a",
goth: MyApp.Goth
)
# 3. Spawn a cheap, self-deleting Spot VM and wait until it exists
{:ok, instance} =
GcpCompute.insert_instance_and_wait(config,
name: "worker-1",
machine_type: "e2-micro",
spot: true,
max_run_duration: 3600, # hard server-side TTL (seconds)
startup_script: "#!/bin/bash\necho ready > /tmp/ready",
labels: %{"owner" => "platform"}
)
GcpCompute.Instance.external_ip(instance) #=> "34.x.x.x"
# 4. Tear it down (and wait for the delete operation)
{:ok, _op} = GcpCompute.delete_instance_and_wait(config, "worker-1")Prefer to drive the lifecycle yourself? Every mutating call returns an
GcpCompute.Operation you can poll:
{:ok, op} = GcpCompute.insert_instance(config, name: "worker-1", machine_type: "e2-micro")
{:ok, done} = GcpCompute.wait_for_operation(config, op, timeout: :timer.minutes(3))Finding instances
# Exact name. Returns a tagged 404 when it does not exist — match, don't rescue.
case GcpCompute.get_instance(config, "worker-1") do
{:ok, vm} -> vm
{:error, %GcpCompute.Error{status: 404}} -> :not_found
end
# Everything in the config's zone (single page of items, token discarded).
{:ok, instances} = GcpCompute.list_instances(config)
# Server-side search. Note: instances.list takes GLOBS, not regexes.
{:ok, workers} = GcpCompute.list_instances(config, filter: ~s(name = "worker-*"))
{:ok, running} = GcpCompute.list_instances(config, filter: ~s(status = "RUNNING"))
{:ok, batch} = GcpCompute.list_instances(config, filter: ~s(labels.role = "batch"))
{:ok, both} = GcpCompute.list_instances(config, filter: ~s(labels.env = "prod" AND status = "RUNNING"))
# Explicit pagination when you need the token.
{:ok, %{items: page1, next_page_token: token}} =
GcpCompute.list_instances_page(config, max_results: 50)
{:ok, %{items: page2}} =
GcpCompute.list_instances_page(config, max_results: 50, page_token: token)Three sharp edges, all verified against the live API:
- Globs, not regexes.
name = "worker-*"works;name ~ "worker"is rejected with a 400Invalid list filter expression. order_byis narrow."name"and"creationTimestamp desc"are accepted;"name desc"is a 400.- Listing is per-zone. The
aggregatedListendpoint is not wrapped, so if you do not know an instance's zone you must iterate zones yourself. There is also nostream/auto-pagination helper —list_instances/2returns one page and discards the token, so loop onlist_instances_page/2for large zones.
Logs, tags and SSH
# "The VM's logs" = serial console output. This is where a :startup_script's
# output lands, and it is the only log surface the Compute API itself offers.
{:ok, %{contents: log, next: next}} = GcpCompute.instance_logs(config, "worker-1")
# Poll incrementally — the buffer is ~145 KB on a booted Debian image, so pass
# the previous :next back as :start rather than refetching the whole thing.
{:ok, %{contents: delta}} = GcpCompute.instance_logs(config, "worker-1", start: next)
# Ports 1..4; 1 carries boot and startup-script output. Anything else is
# rejected locally, before a request goes out.
{:ok, _} = GcpCompute.instance_logs(config, "worker-1", port: 2)Network tags vs labels are different things and both round-trip:
{:ok, vm} = GcpCompute.launch(config, "web-1",
tags: ["http-server", "ssh"], # firewall/route targets -> vm.tags
labels: %{"env" => "prod"} # key/value metadata -> vm.labels
)
vm.tags #=> ["http-server", "ssh"]
vm.labels #=> %{"env" => "prod"}Mutating either after creation (instances.setTags / setLabels) is not
wrapped: both take a fingerprint for optimistic concurrency, and hiding a
read-modify-write behind a helper would hide a lost-update race. The
fingerprints are in vm.raw if you want to call those endpoints yourself.
SSH. This library authorises keys but does not open connections:
{:ok, vm} = GcpCompute.launch(config, "worker-1",
ssh_keys: %{"deploy" => "ssh-ed25519 AAAAC3Nz… deploy@laptop"}
)
GcpCompute.Instance.external_ip(vm) #=> "34.x.x.x" — then use your own clientDeliberately no in-library SSH client: it would mean owning host-key
verification policy, key parsing and known_hosts — a larger security surface
than the rest of this package — and GCE SSH also involves OS Login, a separate
API. For "run a command and see the output", prefer :startup_script plus
instance_logs/3: no inbound access, no keys, no extra API.
Boot disks: image, snapshot, or an existing disk
A boot disk has exactly one source. Passing two is rejected before the request —
GCP's own answer is a 400 naming a JSON field you never typed
(Cannot specify both 'source' and 'initializeParams').
# 1. From an image (the default: debian-12 if you say nothing).
GcpCompute.launch(config, "worker-1", disk_size_gb: 50, disk_type: "pd-ssd")
# 2. From a snapshot — a NEW disk is created, so the snapshot is never mutated
# and any number of instances can boot the same one concurrently.
GcpCompute.launch(config, "worker-2", source_snapshot: "nightly-backup")
# 3. Attaching an EXISTING disk — nothing is created; it must be in the same
# zone and not already attached.
GcpCompute.launch(config, "worker-3", source_disk: "stateful-1")Bare names are qualified for you (nightly-backup →
global/snapshots/nightly-backup); pass a full path to cross projects.
autoDelete defaults differ, deliberately. A disk created from an image or
snapshot belongs to the instance and is deleted with it. A disk you attached
already existed, so deleting the instance leaves it alone. Override either way
with boot_disk_auto_delete:.
Snapshot and disk management is not wrapped: those are different resources
(disks, snapshots), and this package is the instance client. Booting from
one is an instance concern, which is why it lives here.
Spot VMs and preemption
:spot defaults to true, so the quickstart VM is preemptible and, with the
default instanceTerminationAction: "DELETE", deletes itself when preempted.
That is the right default for disposable batch workers and the wrong one if the
disk matters — pass spot: false then.
GcpCompute.Instance.spot?(vm) #=> true
# A preempted VM is DELETED, so `get/3` answers 404 and cannot tell you why it
# went away. The operation outlives the instance, and can:
case GcpCompute.instance_preemption(config, "worker-1") do
{:ok, nil} -> :not_preempted # also the answer for a name that never existed
{:ok, op} -> {:preempted_at, op.end_time}
end
# Or the whole audit trail — "what happened to my VM?"
{:ok, ops} = GcpCompute.instance_operations(config, "worker-1")
Enum.map(ops, &{&1.operation_type, &1.status_message})
#=> [{"compute.instances.preempted", "Instance was preempted."}, {"insert", nil}]
# Test your own preemption handling on purpose. This is a REAL preemption:
GcpCompute.simulate_maintenance_event(config, "worker-1")Measured against a live Spot VM: RUNNING → STOPPING at ~60 s → gone (404) at
~73 s. Budget minutes, not seconds.
There is no retry-on-preemption loop here: deciding whether to relaunch, where, and how often is orchestration, and it belongs in the supervised layer described under Roadmap rather than in a REST client.
Examples & guides
- Runnable examples in
examples/:stubbed_demo.exs— the full insert → poll → get flow offline, no GCP or credentials (elixir examples/stubbed_demo.exs).spawn_spot_vm.exs— a real (billable) Spot VM.- Live smoke tests (real project, real cents) — see
Testing without GCP for what each one is for:
smoke_test.exs,smoke_coverage.exs,smoke_advanced.exs.
- Guides: Getting Started · Configuring Machines · Testing
Configuration
GcpCompute.Config is validated with NimbleOptions. Three builders cover the
common cases:
# Production — tokens minted by Goth
{:ok, config} = GcpCompute.Config.production(project: "p", goth: MyApp.Goth)
# From application env
# config :my_app, :gcp_compute,
# project: "p", zone: "europe-west4-a",
# token_provider: {GcpCompute.TokenProvider.Goth, MyApp.Goth}
{:ok, config} = GcpCompute.Config.from_env(:my_app, :gcp_compute)
# Local / emulator / tests — a static token, no Goth required
{:ok, config} = GcpCompute.Config.local(project: "p", base_url: "http://localhost:8080/compute/v1")| Option | Default | Notes |
|---|---|---|
:project | — (required) | GCP project id. |
:zone | "us-central1-a" | Default zone; override per call with zone:. |
:token_provider | {TokenProvider.Goth, GcpCompute.Goth} | {module, arg} implementing GcpCompute.TokenProvider. |
:base_url | "https://compute.googleapis.com/compute/v1" | Point at the emulator or a proxy. Trusted app config — never derive from user input. |
:req_options | [] | Merged into every Req request (:retry, :adapter, …). |
:allow_insecure | false (local/1 sets true) | Permit a non-https:// :base_url. Off by default so a bearer token is never sent in cleartext. new/1 warns when it is on and the host is not loopback/RFC1918/internal. |
Pluggable auth
Auth is a behaviour, GcpCompute.TokenProvider, so the library never hard-depends
on Goth. Built-ins: TokenProvider.Goth (production), TokenProvider.Static
(tests/emulator). Bring your own (workload identity, metadata server, Vault) by
implementing one callback:
defmodule MyApp.MetadataToken do
@behaviour GcpCompute.TokenProvider
@impl true
def fetch_token(_arg), do: {:ok, %{token: fetch_from_metadata_server()}}
endConfigurable machines
GcpCompute.Instance.spec/1 turns friendly options into the Compute insert body
(defaults shown):
| Option | Default |
|---|---|
:machine_type | "e2-micro" |
:source_image | debian-cloud/.../debian-12 |
:disk_size_gb | 10 |
:spot | true (SPOT, no auto-restart, DELETE on terminate) |
:max_run_duration | nil (set seconds for a hard server-side TTL) |
:external_ip | true |
:startup_script, :metadata, :labels, :tags, :network, :subnetwork, :service_account, :scopes | — |
Need a field it doesn't cover? Pass a raw Compute body map to
GcpCompute.Instances.insert/3 instead.
Telemetry
Every API call is a :telemetry.span/3:
| Event | Metadata |
|---|---|
[:gcp_compute, :request, :start] | method, path, project |
[:gcp_compute, :request, :stop] | + result (:ok/:error), http_status |
[:gcp_compute, :request, :exception] | + kind, reason, stacktrace |
GcpCompute.Telemetry.attach_default_logger() # dev convenience
# or wire into Telemetry.Metrics / your reporterTesting without GCP
Tests stub the network via Req's native :adapter option — no Plug, no creds.
Req 0.7 takes a module, not a function:
defmodule MyStub do
def run(request) do
{request, %Req.Response{status: 200, body: %{"name" => "worker-1"}}}
end
end
config = GcpCompute.Config.local!(project: "test", req_options: [adapter: MyStub])See test/support/req_stub.ex for the routing-table helper used by the suite —
it asserts the bearer token on every stubbed request and fails loudly on an
unrouted one, so a dropped header or a mangled query cannot pass silently.
…and three that need real GCP
Stubs cannot catch everything, and it is worth knowing exactly where the line
is. An in-process adapter computes no headers, so it could not reproduce this:
the Compute API rejects a bodiless POST with HTTP 411 Length Required, which
broke launch/3, terminate/3 and every *_and_wait helper end to end while
the whole stub suite stayed green. Two live scenarios are the counterweight, and
between them they have found every bug the stubs could not.
examples/smoke_test.exs — the fast pre-publish check. Launches one
e2-micro Spot VM, verifies the parse of a real payload, tears it down, and
reports per-request timings from the library's own telemetry, including whether
a single operations.wait long poll exceeded Req's 15 s default
receive_timeout — the only real proof the long-poll fix works. An observed run
had one block for 51 s and succeed.
examples/smoke_coverage.exs — the full surface, for the parts whose
correctness only the real server can confirm, because a stub validates our
request construction against our own assumptions. 19 checks: pagination
(max_results/page_token genuinely paginating, not just our translation table
agreeing with itself), stop/start, spot: false, Operations.get/2,
wait_for_operation/3 called directly, a raw map body, :request_id, a
server-side :filter, and from_env/2 + config/2.
examples/smoke_advanced.exs — 30 checks covering the three things that need
real infrastructure to be falsifiable at all:
- Disks. Boot from a snapshot and from an attached existing disk, then prove
autoDelete: falsereally preserves a disk the instance did not create. No stub can tell you whether GCP accepts a bare snapshot name. - Network. Firewall rules, ports and tags, verified by opening real TCP
connections to the VM — the only test here that leaves both the process and
the Google API, and the only way to know
:tagsare load-bearing rather than decorative JSON. Expectations are derived from the project's live rule set, never hardcoded, because an auto-created default network shipsdefault-allow-sshand a custom VPC ships nothing. The sharp assertion is the difference between a timeout (packets dropped by a firewall) and:econnrefused(packets arrived, nothing listening) — conflating those two is how you convince yourself a firewall works when it does not. - Spot. Preempt a VM on purpose with
simulate_maintenance_event/3, watch it be deleted, then confirmpreemption/3still finds the evidence once the instance itself is gone.
The port assertions refuse to run until the VM's serial console reports the
test listener actually bound — read with instance_logs/3, so the scenario
dogfoods the library. That gate exists because it was needed: an early version
used nc, which Debian 12 images do not ship, and the dead listener passed as a
firewall result for two runs. :timeout cannot distinguish "firewall dropped it"
from "nothing was ever listening", so the listener has to be proven independently.
The tag check needs compute.firewalls.create
(roles/compute.securityAdmin). Without it, that one check SKIPs with the exact
grant command rather than passing quietly.
export GOOGLE_APPLICATION_CREDENTIALS=~/.config/gcloud/smoke.json
export SMOKE_PROJECT=my-project SMOKE_ZONE=us-central1-a
mix run examples/smoke_test.exs # ~1 min, 1 VM
mix run examples/smoke_coverage.exs # ~5 min, 3 VMs
mix run examples/smoke_advanced.exs # ~8 min, 4 VMs + a snapshot and a disk
# needs outbound TCP to the VM
All three cost fractions of a cent and clean up after themselves: e2-micro, a
10 GB disk, max_run_duration as a server-side hard TTL so GCP reclaims the
instance even if teardown fails, unique names per run, and a closing orphan
sweep. Run all three before every publish.
Roadmap
This package is the Compute client layer. A higher-level sandbox orchestration layer is planned on top of it:
gen_statemper sandbox (:provisioning → :probing → :running → :terminating)DynamicSupervisor+Registry, named profiles, per-user quotas- a reaper for TTL sweeps + orphan reconciliation (cloud resources outlive the BEAM)
- the gRPC hop: an in-VM agent reached via
grpc_connection_pool, lifecycle events published throughpubsub_grpc
License
MIT.