The end-of-turn token figure, normalised from a session/prompt response
(#827).
Three places a runtime puts it, tried in order:
usage— the (unstable, as of protocol v1)Usageobject:inputTokens,outputTokens,cachedReadTokens,cachedWriteTokens,totalTokens,thoughtTokens. claude-agent-acp ≥ 0.6x fills it with the turn's accumulated API usage (reset when the turn starts) and codex-acp with the thread's last token count;_meta.quota.token_count— gemini-cli's own place for it, snake-cased and outside the protocol. Seefrom_meta_quota/1;_meta.inputTokens/_meta.outputTokens— where an adapter put it before the field had a name.
The result is one flat, string-keyed map (so it can go straight into a
JSON column): %{"input" => n, "output" => n} plus "cache_read" /
"cache_write" when reported. nil when the response carries none of them — a turn without a
usage, not a zero one.
Deliberately not derived from the usage_update notifications that
stream during a turn: those are context-window occupancy (used / size)
and a cumulative session cost, and what they mean per update differs by
runtime. The one figure recorded per turn is the one the runtime reports
when the turn ends.
Summary
Functions
gemini-cli's usage, from _meta.quota.token_count, or nil.
The turn's usage from a session/prompt result, or nil.
Types
@type t() :: %{required(String.t()) => non_neg_integer()}
Functions
gemini-cli's usage, from _meta.quota.token_count, or nil.
A workaround for one runtime, registered as
:gemini_usage_in_meta_quota in Managoat.Runtimes.Quirks; it is a public
function so that deleting it is what the registry notices.
gemini-cli leaves the protocol's usage field empty and reports the turn's
tokens under a vendor _meta extension of its own, snake-cased, on every
session/prompt return except cancelled:
%{"stopReason" => "end_turn",
"_meta" => %{"quota" => %{
"token_count" => %{"input_tokens" => 1234, "output_tokens" => 567},
"model_usage" => [...]}}}google-gemini/gemini-cli#24280 asked for the standard fields and was closed with no plans to add them, so reading this shape is the only way a host bills a gemini turn at all.
model_usage — the same counts split per model, for a turn that switched —
is deliberately not read: the figure recorded per turn is the turn's total,
and the host already knows which model it asked for.
No cache split. gemini's promptTokenCount includes cached tokens and
the adapter passes no cachedContentTokenCount through, so cached input
arrives inside "input" and a caller pricing it pays the base input rate
for it. That over-states rather than invents, which is the direction to err.
The turn's usage from a session/prompt result, or nil.