Five operator procedures, two of them irreversible. This is the operational half of ADR-0005 and its Amendment A; read the getting-started guide first if you have not stood a vault up yet.
Three of the five read differently depending on what shape the scope's key
has. If your provider is keyring-backed - Encryptor.Provider.Kms - read
"The shred and the rotate, per key shape"
before you run P2, P3 or P4. If it is Encryptor.Provider.GcpKms, read "The
GCP operator runbook" as well.
Every step below is labelled with who performs it:
- [encryptor] - a function this package ships.
- [your store] - an action on a key store this package does not own and cannot see. It defines no table, no migration, no repo and no transaction, and no function here takes one.
- [encryptor_ecto] - the downstream migrator's pass over application rows.
- [operator] - a decision, a deploy, or a wait.
Before anything: the vocabulary
encryptor and encryptor_ecto name the same three levels differently, and
the mismatch is dangerous rather than cosmetic. encryptor_ecto calls level 1
"the key-encrypting key" and level 2 a data key - and what this
package calls a data key is level 3, per message, generated by the engine, with
no row in any table at all.
An operator reading a downstream runbook against this package's names can plausibly run a level 2 rotation believing it is a level 1 rotation. That is the difference between a five-minute pass over thousands of key rows and a multi-hour pass over millions of application rows.
| This runbook | encryptor_ecto | Level | Rows rewritten | Function owner | Walk owner |
|---|---|---|---|---|---|
| Root rotation (P1) | R1 | 1 | one per scope per live version, in the key store | encryptor | the key store's package |
| Scope key rotation (P2) | R2 | 2 | every ciphertext for that scope | joint, see P2 | encryptor_ecto |
| Format or context change | R3 | none | every ciphertext the change covers | encryptor_ecto | encryptor_ecto |
| Crypto-shred (P3, P4) | R4 | 2 | one delete per shredded version | encryptor | the key store's package |
Data key rotation (level 3) is not in this runbook, because it is not an
operation an operator performs. It happens on its own every max_age, every
max_messages, every max_bytes, and on every cache recycle. The only lever
on it is the vault's :cache configuration, and a runbook step for it would be
a runbook step for restarting the vault.
R3 is not this package's at all. A change of format, algorithm or
encryption context is a re-encrypt that these functions cannot express:
rekey/2 preserves the context byte for byte, and the suite is vault
configuration rather than a per-call option. It belongs entirely downstream.
What this package ships, and what it only documents
Shipped, as functions:
Encryptor.Envelope.provision/3- mints a version (P2 step 1).Encryptor.Envelope.rewrap/2- rewraps one wrapping (P1 step 3).Encryptor.Vault.rekey/2- whatrewrap/2is built on.Encryptor.Vault.provision/2- asks the provider to provision the selector, on a provider that implements the optionalprovision/2callback (Encryptor.Vault.provision/2inlib/encryptor/vault.ex; theMyApp.Vault.provision/1thatuse Encryptor.Vaultgenerates is the same call with the vault module filled in).This is not the envelope-level mint, and the two are not interchangeable.
Encryptor.Envelope.provision/3mints 32 random bytes and hands you a wrapping to store (P2 step 1).Encryptor.Vault.provision/2creates whatever the provider's own backing authority needs before a selector can be resolved at all - forEncryptor.Provider.GcpKms, the scope'sCryptoKey, which the record makes a one-time act at the scope's mint (ADR-0007 decision 3). A wrap-provider deployment runs the vault-level one at onboarding; a deployment that stores its own wrappings runs the envelope-level one. Neither replaces the other, and a provider that does not export the callback answers{:not_provisionable, module}rather than raising. What comes back on success is what a store needs to rebuild the descriptor, keyed byscope_ref, never the plaintext key; persisting it is the host's, because this package owns no storage (ADR-0003 decision 9).Encryptor.Vault.suspend/2and its inverseEncryptor.Vault.reinstate/2- the vault-local deny gate of P5, written through the vault's:suspension_store(Encryptor.Vault.suspend/2andEncryptor.Vault.reinstate/2inlib/encryptor/vault.ex). They are deliberately not generated onto your vault module: they are an operator's verbs, invoked from a console or a release task against a named vault (ADR-0005 Amendment A decision 1).Encryptor.Provider.Static'skeys:option - the staged candidate list.
Documented, as procedures: all five below, including every step that touches a store this package does not own.
Not shipped, and deliberately:
- No
mixtask. Every procedure needs a store this package has no access to, so a task here could only be a task that takes a callback, which is a runbook with a worse interface. - No scheduler, no expiry, no automatic pruning. An automatic expiry would be a scheduled, unattended, silent data-destruction job.
- No walk over anything. No function here takes a repo, a query, a batch size, or a table.
- No
shred/2, noretire/2, norotate/2. Deleting a wrapping is aDELETEagainst your store, and wrapping it in a function here would imply this package knows what your store's copies are. It does not.
The window, and what closes it
A rotated-away version keeps decrypting until an operator explicitly deletes its wrapping. This is not a convenience; it is what makes rotation runnable at all. If a version stopped decrypting the moment a newer one became current, re-encrypting a scope's rows would require stopping writes for the duration of the pass.
The mechanism is already built into the provider contract. encryption_key/2
returns exactly one descriptor and decryption_keys/2 returns a list. Minting
version n+1 changes what encryption_key/2 answers; it does not change what
decryption_keys/2 answers except by prepending to it. The two are independent
by construction, so there is no ordering, no race, and no deploy in which a
version becomes uncurrent and undecryptable at the same instant.
Three properties follow, and all three are load-bearing:
- There is no upper bound. This package ships no timer, no expiry column, no TTL, no background sweeper, and no "retire versions older than N" helper. A version left live is live in five years. The window is as long as your policy says, and that policy lives in your compliance documentation, not in this package's configuration.
- There is a lower bound, and it is not zero. A shred is safe only after the rewrite pass has finished, the verification pass is green, and every running vault has stopped serving the retired version out of its materials cache. That last term is the one operators forget; see "Cache drainage" below.
- Widening the window costs a list walk and nothing else. A long candidate list costs one small decrypt per candidate on a cold cache; a warm materials cache saves those decrypts but never the provider lookup before them. Nothing about a long window is unsafe. The argument for closing it is that data whose key still exists is data that has not been shredded.
Rotation adds a name; a shred removes one. There is no third mechanism, and
pruning is manual. A live set is whatever your store answers
decryption_keys/2 with. Rotation calls provision/3 with version: n + 1
and you insert the row; the new version is current because it is the highest,
and the old one is live because its row is still there. A shred deletes the
row. There is no state between live and gone, no soft-delete this package
recognizes, and no retired_at column it reads. Whatever your store's notion
of liveness is, membership changes are the only rotation mechanism, and every
membership change is an operator action.
Pruning policy is count-based or age-based at your discretion. This package expresses no preference beyond the observation that many live versions cost one small decrypt each on a cold cache, which argues weakly for a count bound. What it does insist on is that pruning is performed by P3 or P4, with their preconditions, and never as a side effect of anything else.
Cache drainage
A vault's materials cache holds resolved materials for up to max_age seconds.
Every call asks the provider first, and a provider error ends the call before
the cache is consulted, so what a warm cache can outlive depends on what the
provider still answers:
- P3, the scope shred, does not wait on it for a provider that reads its
store on every call, as
Encryptor.Provider.GcpKmsdoes and anEncryptor.Provider.Functionclosure over your store does. Once the wrappings are gone the provider answers{:unknown_key, selector}, and the next call for the scope fails at once, warm cache or cold, on every node. A provider that keeps its own bounded cache delays P3 by that bound instead. - P4, the version retire, waits on it. The provider still answers, with a shorter list, and a warm entry for a message written under the retired version keeps decrypting it until the entry expires or the cache is dropped.
- P2 step 1, the mint, waits on it too, on the write side. A warm
encryption entry from before the mint keeps issuing data keys wrapped under
version n after
encryption_key/2has started answering n+1, until the entry expires, reaches itsmax_messagesormax_bytesbound, or the cache is dropped.
Draining has exactly two levers, and every procedure that needs one names it as an explicit step:
- wait
max_ageon every vault that serves the affected scope, or - restart those vaults.
There is no third option and this package cannot offer one. The engine's cache cannot be enumerated, measured, or selectively invalidated from outside, so a targeted invalidation is not available to be shipped. If the upstream engine grows a bounded, inspectable cache, these steps get sharper.
Cache drainage is part of the procedure, not a follow-up. A retire that stops at the delete is a retire that has not happened yet, and a rotation whose rewrite starts before the drain can rewrite rows under the version it is about to retire.
P1. Root rotation (R1, level 1)
Rotates the wrapping subkey. Touches no application data.
Preconditions
- [operator] The reference root is pinned in its own secret. If you followed the getting-started guide, both secrets were provisioned at install holding the same bytes and step 0 below does not apply to you. If your deployment has only one root secret, step 0 applies and it is the most dangerous step in this document.
- [your store] You can enumerate every wrapping for every live version.
- [your store] Updating one wrapping is a single-row write that can be retried.
Steps
0. [operator] Only if this deployment has a single root secret: copy the current root material into a new reference-root secret. Deploy and restart.
Copy it. Do not generate a fresh value. Nothing changes cryptographically; this step exists only so that step 2 does not move the reference. Generating a new value here changes every scope reference in the deployment, orphaning every stored key name and every header already written, for every scope. It is recoverable only by restoring the original root material. This is the single most destructive mistake available in this package, and it looks like a no-op.
[operator] Generate new root material and place it in the wrapping-root secret, alongside the outgoing value.
[operator] Configure the root vault's
Staticprovider withkeys: [new, old], newest first, with distinctnames. Restart the root vault. Reads now succeed under either subkey; writes use the new one.defmodule MyApp.RootVault do use Encryptor.Vault, otp_app: :my_app @impl true def init(config) do new_root = Base.decode64!(System.fetch_env!("MY_APP_WRAPPING_ROOT_KEY")) old_root = Base.decode64!(System.fetch_env!("MY_APP_WRAPPING_ROOT_KEY_PREVIOUS")) {:ok, Keyword.put(config, :provider, {Encryptor.Provider.Static, keys: [ [key: Encryptor.Envelope.root_subkey(new_root, "root-wrap"), namespace: "encryptor-root", name: "r/v2"], [key: Encryptor.Envelope.root_subkey(old_root, "root-wrap"), namespace: "encryptor-root", name: "r/v1"] ]})} end endThe root vault runs
cache: false, so the restart costs nothing but the restart. Two entries sharing anameis refused at start:{:invalid_config, :provider, :duplicate_key_names}.[encryptor] + [your store] Walk every live wrapping and call
Encryptor.Envelope.rewrap/2on it, persisting the result:for row <- MyApp.MerchantKeys.all_live() do {:ok, rewrapped} = Encryptor.Envelope.rewrap(MyApp.RootVault, row) MyApp.MerchantKeys.update_wrapping(row, rewrapped) endOrder does not matter and the pass is resumable.
rewrap/2is idempotent in effect and never in bytes: a wrapping already under the new subkey rewraps to different bytes (a fresh data key and a fresh IV go into every message) and an identical descriptor on unwrap. Only the:wrappedfield moves; every identity field is carried across unchanged, and so is the binding.[operator] Verify (below). Then remove the outgoing entry from
keys:and restart the root vault.[operator] Destroy the outgoing root material from the secret store, once backups taken before step 4 are out of retention or are known to be re-encrypted.
Verification
- Every wrapping's header names the new root key. The wrapping key's
nametravels in the clear in the encrypted data key's provider info, so this is a census over stored bytes rather than a decrypt pass -Encryptor.Message.describe/1reads it without a key. - A sample of scopes resolves:
decryption_keys/2returns descriptors, and a test decrypt of one known ciphertext per sampled scope succeeds. - After step 4, the outgoing root material no longer decrypts anything: a deliberate attempt with the old-only configuration in a scratch environment should fail.
Failure and rollback
Reversible at every step until step 5. A pass interrupted at step 3 leaves a
mixed population that both configurations can read, which is precisely why
step 2 precedes it. To roll back, reverse the order of keys: and re-run
step 3; the outgoing subkey becomes current again. The one irreversible action
is step 5, and it is separated from the rest by an explicit verification.
P2. Scope key rotation (R2, level 2)
Rotates one scope's master key. Every ciphertext for that scope is rewritten.
The seam runs between the key lifecycle (here) and the row walk (downstream).
Steps 1 and 4 are key-store operations and belong to this package; steps 2 and
3 are row operations and belong to encryptor_ecto.
Preconditions
- [your store] The scope has at least one live version and its current version resolves.
- [operator] The downstream migrator has a plan covering every table and column holding that scope's ciphertext. A column missed here is a column whose rows become unreadable at step 4. This is the failure mode the whole procedure exists to prevent.
- [operator] Enough time budget to complete the rewrite before the retire. The two are not required to be in the same maintenance session, and they should not be.
Steps
[encryptor] + [your store] Mint version n+1; insert the row. Version n+1 is now current for new writes; n remains live. The window opens here.
{:ok, wrapped} = Encryptor.Envelope.provision(MyApp.RootVault, merchant.id, reference_subkey: reference_subkey, namespace: "acme-merchant", version: 2 ) {:ok, _row} = MyApp.MerchantKeys.insert(wrapped)Both versions now resolve. New writes take v2; old rows still open under v1:
{:ok, [%Encryptor.Key.Aes{name: "t/" <> _ = v2}, %Encryptor.Key.Aes{name: v1}]} = MyApp.MerchantKeyProvider.decryption_keys(state, merchant.id) # v2 is "t/<scope_ref>/v2", v1 is "t/<scope_ref>/v1", newest first.[operator] Then drain the caches before step 2: wait
max_ageon every vault that serves the scope, counted from when the row is visible to every node's provider, or restart those vaults. Until the drain completes, a write whose encryption context finds a warm cache entry from before the mint is still wrapped under n, and so is a migrator rewrite, becauserekey/2's write half goes through the same cache. Those rows are ordinary rows under n: step 2 rewrites them because it starts after the drain.[encryptor_ecto] Run the migrator over the scope's rows. Nothing in this package participates. Every row it rewrites is re-encrypted under whatever
encryption_key/2now answers, which is n+1, without the migrator naming a version.[encryptor_ecto] Run the verification pass over the same rows - all of them, not a sample.
[operator] Only when step 3 is green, run P4 for version n.
Days may pass between steps 1 and 4, and the window stays open the whole time. Nothing forces step 4 to ever run.
Verification
- The migrator's verifier exits zero over all of the scope's rows.
- No row remains whose header names version n, per the same census as P1.
- The scope's
decryption_keys/2still contains both versions at this point. If it does not, something removed a version outside this procedure and P4 must not be run until that is understood.
Failure and rollback
Fully reversible until step 4. An interrupted rewrite leaves rows under both versions, all readable, and the pass resumes. Rolling back the rotation itself means making n current again, which is a question your store answers, and the already-rewritten rows stay readable either way because both versions are live. There is no state in this procedure, before step 4, from which data can be lost.
P3. Scope shred (R4, offboarding)
Destroys every version of one scope's master key. Irreversible.
Preconditions
- [operator] A recorded, human decision that this scope's data is to be
destroyed, with a reference to it in the change record. A shred is never
automated and never a cascade from a
DELETEon an accounts table. - [operator] Every legitimate consumer of that scope's data has been identified. After this procedure, no amount of key material recovers it.
- [operator] The backup implications below have been read and accepted.
Steps
- [your store] Confirm the scope reference resolves and enumerate the wrappings that are about to be destroyed. Record the count and the version numbers in the change record.
- [your store] Delete every wrapping for the scope. On a keyring-backed provider this step is not the shred - see "The shred and the rotate, per key shape" for the step that is, and "The GCP operator runbook" for step 2a on the GCP wrap-provider path.
- [operator] Drain the caches. For a provider that reads its store on
every call, step 2 already made the next call for the scope fail with
{:unknown_key, selector}, warm cache or cold, and this step is residency rather than readability: the scope's data keys stay in the cache table until the table is dropped, by a restart of those vaults or by the recycler's next drop (waitingmax_agedoes not remove them, because an entry is expired only when a read finds it). For a provider that keeps its own bounded cache, wait that bound: until it elapses, the scope stays readable on a node that cached it. - [your store] Delete the scope's ciphertext rows. See "What a shred does not destroy" below: whether this step is optional depends on whether the scope's attribution is itself personal data in your jurisdiction.
Verification
encrypt/2anddecrypt/2for that scope return{:error, %Encryptor.Error{reason: {:unknown_key, selector}}}, not:decrypt_failed. This is the acceptance test for the shred.- The key store returns no rows for the scope reference.
- A ciphertext known to belong to the scope does not decrypt from any node.
Failure and rollback
None after step 2. A partially completed step 2 leaves some versions live and some destroyed - some rows readable, some permanently not, and the two indistinguishable to the application. That is the worst state in this document. If step 2 fails partway, complete it rather than reverting it, and record what was destroyed.
P4. Version retire (R4, one version)
Closes the window opened by P2. Irreversible.
It is a separate procedure from P2 on purpose, so that it is a separate decision with its own preconditions rather than the last line of a longer procedure.
Preconditions
- [operator] P2 steps 2 and 3 completed and green, over all of the scope's rows, not a sample. This is the fence. There is no legitimate reason to retire a version whose rows have not been verifiably rewritten.
- [operator] The verification is recent enough that no traffic since could
have written under the retired version. It cannot have once P2 step 1's
drain completed before step 2 began: from the end of that drain every write
is wrapped under what
encryption_key/2answers, which is n+1 - so this precondition is about your confidence in the plan's coverage, not about the mechanism. - [operator] Cache drainage is understood to be part of this procedure, not a follow-up.
Steps
- [your store] Delete the wrapping for version n.
- [operator] Drain the caches, as P3 step 3.
Verification
decryption_keys/2for the scope no longer contains version n.- The application's error rate for that scope is unchanged. A rise in
:decrypt_failedmeans rows were missed, and they are now unrecoverable.
Failure and rollback
None.
P5. Suspend and reinstate (R-none, the third verb)
Denies every operation for one selector while leaving its wrappings untouched.
Reversible throughout. ADR-0005 Amendment A decision 1 defines the verb by
its observable: a selector is suspended when, on every vault that serves it,
encrypt/2, decrypt/2, rekey/2 and derive/2 fail with
{:key_unavailable, selector} and the wrappings that selector resolves to are
intact in the key store. The data is unreadable and not destroyed.
It exists because P3 was otherwise the only tool for "stop serving this scope now", and P3 cannot be undone. A suspended account, a disputed data licence, a subject-access hold, a customer migrating out and not yet gone: all of those want this and none of them wants a shred.
How far a suspension reaches depends on the vault's store. This is the
property most often got wrong. suspend/2 and reinstate/2 write through the
vault's :suspension_store, a {module, opts} pair (ADR-0010 decision 2).
Whichever store it is, the gate every call passes reads a per-node table and
never the store (Encryptor.Vault.Suspension.suspended?/2, ADR-0010 decision
3), so the deny costs one lookup and never waits on the store. What differs is
who sees the write:
The default store, Encryptor.Vault.Suspension.Store.Ets | A shared store the host implements | |
|---|---|---|
the node that ran suspend/2 | denies at once | denies at once, once the store accepted the write |
| every other node serving the vault | never, unless suspend/2 runs there too | denies within one :suspension_poll_interval of the write (milliseconds, default 5_000) |
| after a restart or a deploy | serves the scope again: the set went with the vault | denies it again, from the node's first read of the store |
| before the node has first read the store | does not arise | denies every scope with {:key_unavailable, selector} |
| the store cannot be reached | does not arise | a write answers {:suspension_store_unavailable, store} and changes nothing on this node; a failed refresh keeps the last set the node read |
That is ADR-0010 decision 6's table read with decision 7's failure rules, and
the code on each side: the default store's set is the node's table
(Encryptor.Vault.Suspension.Store.Ets), while under any other store a
refresher child reads the store at start and then once per poll interval
(Encryptor.Vault.Suspension.Refresher). Under a shared store the node's table
is created out of the gate's sight and put in place in one step when the store
first answers, so until then the gate finds no table and denies every scope
(Encryptor.Vault.Suspension.create/1 and suspended?/2).
Under the default store the suspension is node-local and volatile, and that
is a decision rather than a side effect (Amendment A decision 8, ADR-0010
decision 4). The set lives in an ETS table owned by the vault's Lifecycle
child, so it dies with the vault: a restarted vault serves the selector again,
and a host running four nodes has four vaults and must suspend on each. It also
dies with less than the vault: the vault's supervisor is :one_for_one, so a
Lifecycle child that crashes on its own is restarted with a new and empty
table while the cache and the provider keep serving, and every selector the set
held is served again with nothing else having stopped. Under a shared store the
same restart recreates the table out of the gate's sight, so every scope is
denied until the refresher has read the store again, which is what makes that store's suspension survive a
restart in fact and not only in storage (ADR-0010 decision 7). Between the
crash and the restart there is a window in which the vault answers nothing at
all, on either store: Lifecycle's terminate/2 erases the frozen
configuration on its way down, so every public entry point of the vault
answers {:vault_not_started, vault} until the child is back and has frozen it
again (Encryptor.Vault.ensure_started/2, through
Encryptor.Vault.Config.fetch/1), and so do suspend/2 and reinstate/2.
The gate's own read is not an entry point and answers no error: a call that
read the configuration before the crash and reaches the gate inside the window
finds no table, so it passes the gate under the default store and is refused
as suspended under a shared store
(Encryptor.Vault.Suspension.suspended?/2). The window is a restart long rather
than an outage, but a caller inside it sees a not-started error rather than a
served selector. Under the default store step 1 therefore runs on every
node, and again after every deploy, unless the provider locus of step 2 is
used instead or as well.
Either store binds only calls through this vault. A second client, code that reaches the key material without this vault, or a node configured with a different store is not denied by it. The provider locus of step 2 binds every holder of credentials to the key material (ADR-0010 decision 9).
Preconditions
- [operator] A recorded decision naming the selector, the reason, and - the part that is usually left out - who may lift it. A suspension with no named owner becomes a shred by neglect.
- [operator] Which column of the table above your vault is on is understood. Under the default store, the node fan-out and the post-deploy re-application are planned rather than remembered. Under a shared store, the poll interval is the delay you accept before the other nodes deny.
- [operator] Callers tolerate
{:key_unavailable, selector}.encryptor_ecto's per-owner filter shape (ece-ADR-0002 decision 11) is the shape that already works for this.
Steps
- [encryptor]
Encryptor.Vault.suspend(MyApp.Vault, selector). Under the default store, on every node. Under a shared store, once, on any node that shares it. An{:error, %Encryptor.Error{reason: {:suspension_store_unavailable, _}}}means the write was not confirmed: this node's set is unchanged, and a store that timed out may or may not hold it. Retry it; both verbs are idempotent. - [operator] Optionally, revoke at the provider's own backing authority - an IAM binding on the scope's key. This is the locus that binds more than this vault, it survives every restart, and it is the half that needs the change record.
There is no cache-drainage step. The deny gate sits at resolution, ahead of the
materials cache, so the very next call fails on a warm cache as on a cold one
(Amendment A decision 5). P3 needs no drain for the same reason, for a provider
that reads its store on every call; P4, whose provider still answers, is the
procedure that waits on one. Suspending does drop this vault's materials cache
as hygiene, because no partition-scoped eviction exists (decision 6), so every
other selector on the vault takes one cold miss; a vault configured
cache: false has no cache to drop and is unaffected.
Verification
encrypt/2anddecrypt/2for that selector return{:error, %Encryptor.Error{reason: {:key_unavailable, selector}}}- not:decrypt_failed, and not{:unknown_key, selector}, which would mean a shred rather than a suspension.- The key store still returns the scope's rows. If it does not, this was not a suspension.
- Run the check against every node; under a shared store, allow one poll
interval first.
suspend/2asks no provider, so suspending a typo'd selector succeeds quietly; this step is what catches it. - Under a shared store, watch
[:encryptor, :suspension, :changed]: it fires on everysuspend/2andreinstate/2, and on every refresh that changed the node's set, failed, or recovered after a failure. A node whose refreshes are failing is serving the last set it read.
Failure and rollback
Encryptor.Vault.reinstate(MyApp.Vault, selector) - on every node under the
default store, once under a shared store - plus restoring the provider binding
if step 2 was taken. reinstate/2 is total and idempotent: it succeeds on a
selector that was never suspended, and it evicts nothing, because nothing was
served under the suspension. Under a shared store it reaches the other nodes as
a suspension does, within one poll interval.
Under the default store, a partially applied step 1 leaves some nodes denying
and some serving, which shows up as an intermittent {:key_unavailable, _}.
Complete it rather than reverting it.
What reinstate/2 does not do is undo anything else. Reinstating a selector
whose wrappings were shredded while it was suspended restores the gate, and the
provider then answers {:unknown_key, selector}. There is no state in which it
recovers key material, and a suspension is not a backup.
The two failure shapes after a destruction
# After P4, on a row the plan missed. Indistinguishable from corruption:
MyApp.MerchantVault.decrypt(missed_row, key: merchant.id, encryption_context: ctx)
#=> {:error, %Encryptor.Error{reason: :decrypt_failed, operation: :decrypt}}
# After P3, the whole scope. Specific, and not an oracle:
MyApp.MerchantVault.decrypt(any_row, key: shredded.id, encryption_context: ctx)
#=> {:error, %Encryptor.Error{reason: {:unknown_key, "..."}, operation: :decrypt}}A whole-scope shred is loudly distinguishable from corruption: the provider finds no live row and fails at resolution time, before any ciphertext is examined, on a term that depends only on the caller's selector.
A single retired version is not distinguishable. A message written under a
retired version fails as :decrypt_failed, identical to a corrupted message, a
wrong-key message, and a context mismatch, because the whole decrypt-side
failure space collapses to one reason and this runbook does not reopen that.
The operational answer is the fence, not a new error term. P4's precondition is
a green verification over all of the scope's rows, so a :decrypt_failed after a retire means a
row the plan did not cover - a bug in the plan, not an ambiguity in the error.
What a shred does not destroy
Read this before promising anyone that your offboarding crypto-shreds.
A shred destroys plaintext, not attribution. Every message header carries
the scope's permanent pseudonym - its scope reference, in the encrypted data
key's name and in the encryption context under "tenant_ref" - and deleting
the wrapping does not touch it. Outsiders cannot resolve the pseudonym. The holder of the reference
subkey can, by guess-and-confirm, forever, in every retained backup: derive
the reference for a candidate scope id, compare it against the header, and the
match is a confirmation. The reference subkey is never rotated, so this does not
expire.
Consequently, P3 step 4's deletion of the scope's ciphertext rows is compliance-mandatory wherever the scope's attribution is itself personal data, not hygiene. And the shred claim must never be stated as full erasure.
A shred is only as good as the copies. Deleting a wrapping destroys the key in the primary store. It does nothing about:
- database backups taken while the wrapping existed,
- read replicas that have not yet received the delete,
- a logical dump on someone's laptop,
- a wrapping exported into a disaster-recovery vault,
- a snapshot of a running node's memory.
A host whose compliance story is "we crypto-shred on offboarding" has committed to a backup retention and replica story this package cannot see and does not verify. Stating the limit is the most it can honestly do.
A shred is wider than the columns this package encrypted. A scope master key is also the derivation root for any purpose-labelled subkey - a blind index, a search key - and those subkeys are recomputed on demand and never stored. They die with the master key. That is intended, and it means a shred reaches further than an inventory of encrypted columns would suggest.
The shred and the rotate, per key shape
Everything above P4 is written in raw-material terms: a scope master key is 32
bytes, you hold its wrapping, and deleting the wrapping destroys the key. That
is true of every material-source provider - Encryptor.Provider.Static,
Encryptor.Provider.Function, Encryptor.Provider.GcpKms, an Ecto-backed
wrapped-key table - and it is not true of a keyring-backed one.
Encryptor.Provider.Kms is keyring-backed: the data key is generated inside
AWS KMS and the wrapping key is never in your store, so there is no wrapping of
it for you to delete. ADR-0008 decision 4 is the record that reconciles the two
shapes, and it asks that its table be reproduced rather than paraphrased. It is
reproduced here in full (docs/adr/0008-aws-kms-keyring-backed.md, the table
under decision 4, "Rotation, the shred, and suspend, per shape");
Encryptor.Provider.Kms's moduledoc carries the same
reconciliation from the provider's side.
%Key.Aes{} (material source) | %Key.Kms{} (keyring-backed) | |
|---|---|---|
| version identity | name, minted by the provider (ADR-0002 d4) | the KMS key ARN, assigned by AWS |
| header provider id | the descriptor's namespace | "aws-kms", written by the engine |
| header provider info | name | the key ARN from the GenerateDataKey / Encrypt response |
| who holds the wrapping key | the host's key store, as a wrapped blob (ADR-0003 d2) | AWS KMS; nothing is stored |
| the data key is generated | by the engine, locally | inside KMS, by GenerateDataKey |
| ADR-0003's two-level envelope | yes | no - decision 7 |
| rotation (ADR-0005 R2, level 2) | mint a new name, prepend its descriptor, re-encrypt | point at a new KMS key, prepend its descriptor, re-encrypt |
| rotation that is invisible here | none | AWS KMS automatic key rotation: new backing material under the same ARN. Not R2, not R1, not an operation in this package's vocabulary at all |
dropping the identity from decryption_keys/2 | is the shred - the material exists nowhere else (ADR-0005 d3) | is not the shred - it hides the data from this vault while KMS can still decrypt it |
| the shred (ADR-0005 P3 step 2) | DELETE the wrapping from the key store | ScheduleKeyDeletion on the tenant's KMS key |
| irreversible | immediately, subject to backups of the store | after the KMS pending-deletion window; CancelKeyDeletion works inside it |
| does the shred survive a backup | only if every copy of the store was found (ADR-0005's residual) | yes - the key material was never in the backup |
| suspend (ADR-0005 Amendment A) | the vault-local deny gate, A3's first locus | the vault-local deny gate, and an IAM revoke on the key as A3's second locus |
The one row to read twice is the ninth. Dropping a version's identity from
decryption_keys/2 is the shred on the Aes shape and is not the shred
on the KMS shape - there it only hides the data from this vault, while KMS can
still decrypt it for anyone holding kms:Decrypt on the key, including from a
backup of your ciphertext. An operator who has internalised "delete the row and
it is shredded" will, on this path, have shredded nothing. On the KMS shape
the shred is ScheduleKeyDeletion on the scope's KMS key, and P3 step 2
reads that way rather than as a DELETE.
The pending-deletion window is not a reprieve to plan around. It is what
makes the KMS-path shred reversible for exactly as long as it lasts -
CancelKeyDeletion works inside it - and irreversible the moment it elapses.
P3's first precondition, a recorded human decision, is unchanged by its
existence.
The shred also gets stronger on this path, in one specific sense: the wrapping key was never in a backup of your store, so "a shred is only as good as the copies" stops being the binding constraint. It does not become full erasure. Attribution survives exactly as described above, and P3 step 4's row deletion stays as compliance-mandatory as it was.
The GCP operator runbook
Encryptor.Provider.GcpKms is a wrap-provider: it is a material source by
ADR-0002 decision 5's taxonomy - it decrypts a stored wrapped key and hands back
bytes - so every procedure above applies to it unchanged, with the two additions
below. It is not the keyring-backed shape of the table above.
The ring and the IAM bindings are provisioned out of band
ADR-0007 decision 3: Encryptor.Vault.provision/2 creates the scope's
CryptoKey and this package never creates the KeyRing and never writes
IAM. Both are a one-time, per-environment act by the operator, in Terraform,
the console, or gcloud, before any vault starts.
- Never
CreateKeyRing. A key ring cannot be deleted, so a package that created one would permanently enlarge your GCP project from inside a library call, on a path reachable with a typo'd scope id. - Never any IAM write. The provider's service account needs
cloudkms.cryptoKeyVersions.useToEncryptanduseToDecrypton the ring, pluscloudkms.cryptoKeys.createif it mints. Granting itself those would be a privilege-escalation surface with no upside; a deployment whose IAM is wrong fails loudly at the first call, which is the correct failure.
[operator] Before the first deploy of an environment: create the ring, grant the two use bindings (plus create, if the deployment mints scopes), and record the ring's fully qualified name in the change record. Nothing in this package will do it for you and nothing in it will tell you it is missing until the first call fails.
The ring is a destroy-time hazard in Terraform, not a create-time one
This is the operational note ADR-0007 decision 3 assigns to this guide.
google_kms_key_ring accepts a destroy and removes only the state entry.
The ring itself survives - key rings cannot be deleted - so a later re-apply
hits ALREADY_EXISTS on a resource that no terraform destroy can clear, and
the environment is stuck until someone imports or renames.
The two standard mitigations, and the choice between them is yours:
lifecycle { prevent_destroy = true }on thegoogle_kms_key_ring, or- keep the ring out of the application's Terraform state entirely, managed by the platform team beside the project itself.
The same hazard does not apply to the per-scope CryptoKey: those are created
by provision/2 at the scope's mint, not by Terraform. What they share is
permanence - a destroyed CryptoKey remains in the project forever, empty, and
ADR-0007 decision 3 names that as the cost, paid visibly.
P3 gains step 2a: destroy the scope's CryptoKey versions
ADR-0007's offboarding walk restates ADR-0005 P3 with one step added, numbered 2a so the rest keep their numbers. Preconditions are unchanged and still come first, in particular the recorded human decision.
| P3 step | On this provider |
|---|---|
| 1, enumerate the wrappings | unchanged |
| 2, delete every wrapping from the key store | unchanged, and still your DELETE |
2a, DestroyCryptoKeyVersion on every version of the scope's CryptoKey | new. After the destroy-scheduled window elapses, the scope's data is unreadable from any backup of the key store, because the key that would unwrap those wrappings no longer exists anywhere |
| 3, drain the caches | residency only: GcpKms reads the key store on every call, so step 2 already answers {:unknown_key, selector}; a restart drops the scope's cached data keys |
| 4, delete the scope's ciphertext rows | unchanged in mechanism, and still compliance-mandatory wherever the scope's attribution is itself personal data - destroying the GCP key does not remove the scope reference from retained headers |
P4 gains nothing: the scope's CryptoKey is shared across master-key versions,
so retiring version n is the wrapping delete and nothing else.
Two things step 2a does not change. This package still ships no shred/2 - the
store delete is your DELETE and the destroy is a GCP API call your runbook
makes - and GCP's scheduled destruction window is a delay, not a reprieve, on
exactly the reading the KMS-path window gets above.
Blast radius: what an attacker holding each combination can read
This is the trust boundary the design is shaped around.
| Attacker holds | Can read |
|---|---|
| Application ciphertext only | nothing |
| Ciphertext + the wrapped-key store | nothing |
| Ciphertext + the root key | nothing |
| Ciphertext + the wrapped-key store + the root key | everything, all scopes |
| Ciphertext + one scope's unwrapped master key | that scope, all versions that key covers |
| A running application process | everything it can currently resolve |
Four things follow, and each is a limitation as much as a property:
- The two-factor property is real but conditional. Separating the root key (environment, secrets manager) from the wrapped keys (database) means a database compromise alone - the overwhelmingly common one, from a backup, a read replica, a dump, an SQL injection - yields nothing. This is the design's main return.
- The running process is not protected against, at all. A process that can encrypt for a scope necessarily holds that scope's master key in memory, and one that can provision holds the root. Any attacker with code execution or memory read in the application defeats the hierarchy entirely. Moving the root to a key manager narrows this - the root material stops being in BEAM memory - but the unwrapped scope keys still are.
- A scope's compromise is bounded by version, not by time. One master key covers every message written while it was current. A scope compromised at version 3 exposes exactly the data encrypted under version 3, which is why rotation cadence is a real security parameter and not hygiene theatre.
- Shredding is only as good as the copies. See above.
Blast radius: what each step destroys
"Reversible" means reversible by an operator holding everything they held before the step, without recourse to backups.
| Procedure | Step | Destroys | Reversible | If performed wrongly |
|---|---|---|---|---|
| P1 | 0, copy reference root | nothing | yes | If a fresh value is generated instead of a copy: every scope reference changes, every stored name and every written header orphans. All scopes, all data, recoverable only by restoring the original root material. |
| P1 | 2, restart with keys: | nothing | yes | Wrong order: writes go under the outgoing subkey. Harmless while both are live. |
| P1 | 3, rewrap pass | nothing | yes | Partial pass leaves a mixed population, all readable. Resume or reverse. |
| P1 | 4, drop outgoing entry | nothing yet | yes, by re-adding it | Dropped before the pass finishes: unrewrapped wrappings stop unwrapping, so their scopes stop resolving. Recoverable by re-adding the entry. |
| P1 | 5, destroy old material | the ability to read pre-rotation backups | no | Backups taken before step 3 become unreadable. Bounded by backup retention. |
| P2 | 1, mint | nothing | yes | Minting twice concurrently can produce two version n+1 rows; the transaction that closes that race is your store's. |
| P2 | 1, drain after the mint | nothing | yes | Skipped: writes on a warm cache entry keep going under version n, and a rewrite started before the drain can rewrite rows under n. Step 3's census finds them; a P4 run on a census taken before the drain can orphan them. |
| P2 | 2, rewrite | nothing | yes | Downstream's compare-and-swap; a clobber is the failure it is designed against. |
| P3 | 2, delete all wrappings | one scope's entire dataset, everywhere | no | Wrong scope: that scope's data is permanently unreadable. This is the largest destructive action in the package and the reason P3's first precondition is a recorded human decision. |
| P3 | 3, drain caches | nothing | n/a | Skipped: for a provider that reads its store on every call nothing stays readable, because step 2 already answers {:unknown_key, selector}, but the scope's data keys stay resident in the cache table until a restart or the recycler's next drop. For a provider with its own bounded cache, the scope stays readable on running nodes until that bound elapses. |
| P4 | 1, delete one wrapping | every row still written under that version | no | Run before verification: exactly the rows the pass missed become permanently unreadable, and they surface as :decrypt_failed indistinguishable from corruption. |
| P5 | 1, suspend | nothing | yes, by reinstate/2 | Wrong selector: that scope's reads and writes fail loudly, at once on the node that ran it and on every node the step or the shared store reached. No data is lost and no window opens. Reinstate. |
| P5 | 2, revoke at the provider | nothing | yes, by restoring the binding | Wrong key: as above, durably, and it outlives a restart, so it is the half that needs the change record. |
| any | shredding a scope master key | that scope's derived subkeys too | no | Subkeys are recomputed from the master key and never stored, so a blind index, a search key, or any future purpose-labelled key dies with it. Intended, and it means a shred is wider than "columns encrypted by this package". |
| any | changing a vault's :slow_hash parameters | nothing directly | yes, by restoring the old parameters | Every index value written afterwards is hashed under the new parameters and stops matching values stored under the old ones, and nothing in this package notices: the output carries nothing about the parameters that produced it. It is an invalidating change in the same family as a :derivation_salt rotation, not a tuning knob to turn freely; the migration is encryptor_ecto's two-column dance under a new index version (ADR-0003 amendment B decision 6). |
The division of labour with encryptor_ecto
Stated from this side, so both packages' documentation agrees:
- The Ecto layer resolves a scope and passes it as
key:. It does not put a scope pair in:encryption_context- that pair is refused from a caller - and its missing-scope error fires before this package is called at all. - The Ecto layer supplies
tableandcolumnfrom its frozen declared values. Renaming a physical table or column while keeping the declared value pinned does not invalidate stored rows; changing the declared value is an R3 re-encrypt, not a rotation. - The Ecto layer enforces nothing. What makes those keys required is your
vault's
required_context: ["table", "column"]. Enforcement lives on the vault because it is a property of the vault, not of one type module: two schemas sharing a vault must not be able to disagree about how strictly their rows are bound. - P1 step 3's walk over live wrappings and P3 step 2's delete are operations on
the wrapped-key table, which lives in
encryptor_ecto. They are not migrator operations: they take no plan, no checkpoint, no compare-and-swap and no batch, and they never touch a ciphertext column of an application table.
Cadence
This package gives no numbers and no default, consistent with declining to ship a scheduler. What it will say is the argument: a scope's compromise is bounded by version rather than by time, so cadence is a security parameter with a reason attached rather than hygiene, and the right interval is a function of how much data one version is allowed to cover.
Records
ADR-0005 (rotation and crypto-shred) owns every procedure here, and its Amendment A owns P5, with ADR-0010 deciding where the suspended set is agreed and what each store's suspension reaches; ADR-0009 names the scope and pins the v1 spellings that still say tenant; ADR-0003 decision 10 owns the attacker table; ADR-0004 decision 10 owns the division of labour; ADR-0007 decision 3 owns the GCP out-of-band split and the destroy-time hazard, and ADR-0007 decision 8 owns P3 step 2a; ADR-0008 decision 4 owns the per-shape table, which is reproduced here rather than restated at that record's own request. Where this guide and a record disagree, the record wins and the disagreement is a bug in this guide.