Raxol. UI. Rendering. PaintAuthority. ContentGuard
(Raxol v2.6.1)
View Source
Neutralizes control bytes inside agent/LLM-originated content before it
reaches PaintAuthority.InlineAuthority.seal/2 -- seal/2/
append_sealed/2 previously wrote iodata verbatim, so an embedded
\e[2J/\e[3J/CUP inside model output could wipe native scrollback
or repaint an already-sealed row FROM INSIDE THE CONTENT, defeating
every invariant the append path exists to hold, no matter how correct
the fill-down/seal-once machinery around it is.
This module is intentionally SHARED, not private to the append path:
Raxol.UI.Rendering.PaintAuthority.InlineAuthority.seal/2 is its first
caller, and the footer viewport's pinned-viewport repaint path (also fed
agent-controlled text) is expected to reuse it rather than growing a
second, divergent sanitizer.
The allowlist grammar
ASCII control structure is recognized byte-wise; the ONLY decode is at
the >= 0x80 boundary, by code point, to catch C1 controls (below) --
every other multi-byte UTF-8 sequence is re-emitted intact:
- Printable ASCII (
0x20..0x7E) and legitimate multi-byte UTF-8 (code points>= 0xA0) -- passed through unchanged. \t,\r,\n-- passed through unchanged (the C0 controls a line of legitimate content actually needs).- SGR (
CSI ... m) -- e.g.\e[1;31m,\e[0m-- passed through VERBATIM. This is the renderer's own styling vocabulary; content that legitimately carries color/style resets must keep working. - Every other C0 control byte (
0x00..0x1Fexcept the three above) and DEL (0x7F) -- stripped silently. These carry no printable residue worth preserving. - C1 control code points (
U+0080..U+009F) -- stripped, whether a raw single byte (0x9B= 8-bit CSI,0x9D= OSC,0x90= DCS, ...) or the UTF-8 form (0xC2 0x80..0xC2 0x9F); a lone invalid high byte is dropped too. These are the 8-bit siblings of the ESC-led sequences below -- a byte-wise>= 0x80pass-through would let them through raw and reopen the injection class from the C1 side (#616). - Any other ESC (
0x1B)-led sequence -- a non-SGR CSI (cursor moves, erases, ...), OSC (\e]...), DCS (\eP...), or a bare/ truncated ESC with no recognized introducer at all -- has ONLY its leading ESC byte stripped. Scanning resumes immediately after that ESC, so whatever printable bytes followed it (the[2Jin\e[2J, the]0;titlein an OSC, etc.) are re-scanned as ordinary text and, being printable, survive into the output.
"Visible-honest" neutralization, not silent deletion
A \e[2J with its ESC stripped becomes the literal, visible text
[2J -- four printable characters a terminal just prints, doing
nothing else. That is the deliberate choice this module makes over
silently deleting the whole sequence: a visible [2J fragment in the
sealed history is an honest record that content tried to inject a
control sequence and got stopped, whereas an invisible drop would look,
to anyone watching the terminal, exactly like the model simply never
said anything unusual. Any C0 control byte embedded in what would have
been that sequence's parameters (e.g. a stray BEL terminating an OSC)
IS silently dropped, per the point above -- there is no printable
residue for a non-printable control byte to leave behind.
Scope note: OSC 8 hyperlinks
OSC 8 ; params ; URI ST text OSC 8 ; ; ST (terminal hyperlinks) is
currently neutralized like any other OSC -- the \e]8;... introducer's
ESC is stripped, leaving the URI parameters as visible residue text
rather than a clickable link. Allowlisting OSC 8 specifically (parsing
its structure enough to keep the link semantics while still rejecting
every other OSC subcommand) is deliberately deferred; tracked as
backlog, not implemented here.
Summary
Functions
Sanitizes iodata per the allowlist grammar above, returning a binary.
Safe to call on a whole multi-line sealed block (not just a single
line) -- \r/\n inside are always allowed through, so line
boundaries are preserved exactly.