# Python

Run Python that arrives at request time. You need this page when you are
deciding whether the thing you want to run can run here, and when you are
setting limits: CPython needs several of them raised, and an adapter never
raises a ceiling behind your back.

## Run one

> **Where these modules come from.** `wasm_python_command`, `wasm_python` and
> the worker kernel they run on (`wasm_script_worker`) are installed with the
> application; you supply the CPython artifact.

<!-- check: run -->
<!-- check: needs python -->
```erlang
{ok, W} = wasm_script_worker:start_link(
            wasm_python_command,
            #{path => "test/fixtures/lang/python.wasm",
              limits => #{timeout => 300_000,
                          max_memory_pages => 4096,
                          max_host_calls => 1_000_000,
                          max_heap_words => 16 * 1024 * 1024,
                          fuel => 4_000_000_000}}),
{ok, #{result := #{~"answer" := 42}}} =
    wasm_script_worker:run(W, ~"def main(c):\n    return {'answer': c['value'] + 1}\n",
                           #{~"value" => 41}).
```

The tenant writes one function:

```python
def main(context):
    return {"answer": context["value"] + 1}
```

## Raise the ceilings knowingly

`wasm_limits:untrusted/0` is built for something much smaller than an
interpreter. Three of its defaults will not do:

| limit | why |
| --- | --- |
| `timeout` | a request is tens of seconds, not one |
| `max_memory_pages` | 16 MiB does not hold CPython |
| `max_heap_words` | the default **kills the runner** before the interpreter starts |
| `fuel` | 10,000,000 does not reach CPython's first line; 4,000,000,000 is measured to be enough |

**A bigger `max_heap_words` is slower, not safer.** Use 16M words: 64M and
256M were both measured slower, 85 s against 48 s for a request, because a
larger bound lets the heap grow before each collection and the collection then
costs more. A ceiling is not a target.

`test/fixtures/lang/PYTHON.md` has the rest of the numbers and what artifact
they were taken on.

## What you get

The language, and the interpreter's bundled standard library. The artifact this
is measured against embeds it: `sys.path` names a directory that is in no
preopen and `import json` works anyway, so the worker declares **one** mount.
An upstream build that ships `python.wasm` beside a `Lib` directory needs that
directory preopened read-only as a second mount.

The interpreter is started `-I -B -u`: isolated configuration, no `.pyc`
writes, no output buffering. The third is a bound rather than a preference,
because buffered output arrives in one burst at the end and the streaming limit
never sees it. Because `-I` implies `-P`, the work directory is **not** on
`sys.path`, and the bootstrap loads your module through
`importlib.util.spec_from_file_location` against an explicit path rather than
putting a tenant-supplied directory on the import path.

## What you do not get

**Arbitrary PyPI wheels, or any native extension module.** Anything with C in
it has to be built for `wasm32-wasip1` and linked into the interpreter.

**`subprocess`, native threading, or ordinary socket support.** This is not
something the runtime took away: **CPython itself disables them on the WASI
platform**, which is Tier 3 in
[PEP 11](https://peps.python.org/pep-0011/) and documented as lacking those
facilities in [CPython's WebAssembly platform
notes](https://docs.python.org/3/using/wasm.html).

**Pyodide compatibility.** Pyodide targets Emscripten and needs JavaScript
glue, a browser ABI and a package loader WASI preview 1 does not provide. On a
host whose engine is V8 that glue costs nothing; on the BEAM it is pure
liability. Nothing here should imply a Pyodide package runs unchanged.

**A network.** Denied by choice rather than missing. Granting it is a
`wasi_net` rule naming addresses and ports, and note what the knobs bound:
`max_sockets` caps the descriptors an instance holds **at once**, `timeout`
bounds **one blocking call**. Neither is a subrequest budget.

## It is slow, and the number is written down

53 to 76 s for one request, interpreted, on a lightly loaded machine, and 29
minutes for the conformance suite. That is not a
worker, and saying so is the point: `test/fixtures/lang/PYTHON.md` records it,
`wasm_worker_lang_SUITE` keeps the CPython groups out of its default run
because of it, and initialized runtime snapshots are the answer rather than
tuning.

A snapshot needs a **reactor** exporting `init()` and `handle()`. The artifact
measured here is a command with one `_start`, so it can never support one. The
next section is how to stop paying for that.

## Skip the interpreter start, with the reactor build

**88 ms a request instead of a minute or more.** CPython starts once when the
worker starts, and each request restores an image of that point. Build it, then
point a worker at it:

```sh
scripts/build-python-reactor.sh
```

<!-- check: run -->
<!-- check: fresh -->
<!-- check: needs python_reactor -->
```erlang
{ok, W} = wasm_script_worker:start_link(
            wasm_python,
            #{path => "test/fixtures/lang/py_reactor.wasm",
              lib  => "test/fixtures/lang/py_reactor_lib",
              capture_timeout => 300_000,
              limits => wasm_python:limits()}),
{ok, #{result := #{~"answer" := 42}}} =
    wasm_script_worker:run(W, #{source => <<"def main(c):\n"
                                       "    return {'answer': c['value'] + 1}\n">>,
                           context => #{~"value" => 41}}).
```

The tenant contract is unchanged: the same `main(context)`, the same JSON in
and out, the same capabilities.

Notes:

- **`start_link/2` takes about 14 seconds**, or 5 with the capture floor
  below, because that is one interpreter start. It happens once per worker, not
  once per request, and a host should start its workers before it starts taking
  traffic. It is also longer than the 60 s `capture_timeout` default, which is
  why the example raises it.
- **Isolation is unchanged.** A restore builds a *fresh* instance, so one
  request's module-level state never reaches the next.
- **Two paths, not one.** The module needs its standard library beside it, and
  `lib` is where you say so. The build script produces both.
- **The hash seed is in the image**, drawn once during initialisation and
  shared by every request that restores it. Re-seeding afterwards is not
  available: string hashes are already cached against the old secret, so
  rotation means recapturing. `docs/snapshots.md` covers what else an image
  freezes.
- **It needs a WASI SDK, binaryen's `wasm-opt` and about twenty minutes to
  build**, which the fetched command artifact does not.
  `test/fixtures/lang/PYTHON.md` has the pins and says why there is no
  checksum.
- **The standard library is precompiled.** Every module in `lib` ships with
  its `.pyc`, so an import your source makes loads bytecode instead of
  compiling the module in your request. The files are unchecked hash-based:
  the import never looks at the `.py`, so an edit to one has no effect until
  you rerun the build script.
- **Objects in the image are frozen out of the cyclic collector.** See
  [below](#the-image-is-frozen-out-of-the-collector).
- The numbers, their null experiment and where the time goes are in
  `test/audit/PERF.md`.

## The image is frozen out of the collector

The reactor ends its start with `gc.collect()` and `gc.freeze()`, and a
capture with an `entry` does the same again after the entry has run. Every
object alive at that point moves to the collector's permanent generation, so a
request's collections traverse only what the request allocated. Without it,
a collection due just after the capture was due in every request.

What that means for your code:

- **`gc.get_objects()` does not list objects from the image**, and
  `gc.get_freeze_count()` counts them. Objects your request creates are listed
  as usual.
- **Nothing in the image is collected by the cyclic collector.** Reference
  counting still frees an image object whose last reference goes, and its
  weakref callbacks run as usual. One kept alive only by a reference cycle is
  never freed, and its callbacks never run. It does not leak: the request's
  copy of the image is thrown away when the request ends.
- **Do not call `gc.unfreeze()`.** It hands the whole image back to the
  collector and the next collection pays for all of it, in that request.

## Call a fixed entry instead of sending a source

Use this when the code is the same on every request and only the context
changes: a host that runs one application per worker, for example. The worker
runs the code once, while it captures, and each request calls a function that
is already in the image, so nothing is compiled or imported per request.

<!-- check: run -->
<!-- check: fresh -->
<!-- check: needs python_reactor -->
```erlang
{ok, W} = wasm_script_worker:start_link(
            wasm_python,
            #{path => "test/fixtures/lang/py_reactor.wasm",
              lib  => "test/fixtures/lang/py_reactor_lib",
              entry => <<"import worker\n"
                         "def answer(c):\n"
                         "    return {'answer': c['value'] + 1}\n"
                         "worker.set_entry(answer)\n">>,
              capture_timeout => 300_000,
              limits => wasm_python:limits()}),
{ok, #{result := #{~"answer" := 42}}} =
    wasm_script_worker:run(W, #{context => #{~"value" => 41}}).
```

What to know:

- **`worker.set_entry(callable)` takes one callable, once.** The capture runs
  `entry` as `/main.py`, so a `main` it defines is called once with `None`. A
  worker whose `entry` never calls `set_entry` does not start, and a request
  that calls it again gets an `exception` error saying the entry is already
  set.
- **A request with no `source` calls the entry** with its context; one with a
  `source` runs that source, as on any reactor worker.
- **The entry's globals start fresh on every request.** Each request restores
  the image, so what one request's call changed is gone for the next.
- **The context arrives through an import**, `worker.context`, rather than as
  a staged file, so its size is bounded by `max_request_bytes` alone.
- **Changing the code means restarting the worker.** The entry is part of the
  image and of the image's version, so two workers with different entries
  never share a filed image.

The same request through `handle()` and through an entry, the compiled tier on,
alternating in one emulator on a loaded machine: the guest's call took 40.4 ms
and 2.1 ms. [Tuning a worker host](tuning.md) has the table.

## Give both processes a heap floor

CPython gains more from this than either other guest here, and it gains on both
halves: the start and the request. `wasm_python` sets the request runner's
floor for you, and the number depends on the tier your limits select:

| tier | limits | `runner_min_heap_words` |
| --- | --- | ---: |
| compiled | `compile => true`, `fuel => infinity` | 1,500,000 |
| interpreted | anything else | 1,000,000 |

The compiled tier wants more because at 1,000,000 its heap still grew once
mid-request, to 2.88 M words; at 1,500,000 it never grew, a request was 3 to
6% faster at p50, and a pool of ten answered 29 to 34% more requests against
17 to 24%. The interpreter gains nothing past 1,000,000. Both fit
under `wasm_python:limits/0` and under the untrusted preset with the worker's
headroom.

It does **not** set the capture's: add `capture_min_heap_words` yourself, with
the larger `max_heap_words` below.

<!-- check: run -->
<!-- check: fresh -->
<!-- check: needs python_reactor -->
```erlang
Limits = (wasm_python:limits())#{max_heap_words => 32 * 1024 * 1024},
{ok, W} = wasm_script_worker:start_link(
            wasm_python,
            #{path => "test/fixtures/lang/py_reactor.wasm",
              lib  => "test/fixtures/lang/py_reactor_lib",
              capture_timeout => 300_000,
              limits => Limits,
              capture_min_heap_words => 2_000_000}).
```

| | without | with |
| --- | ---: | ---: |
| `start_link/2`, capturing | 91 to 95 s | **17.4 s** |
| a request | 367 ms | **118 ms** |

Both rows are what the floor sweep measured, and the request row was taken
before the restore path stopped writing over the image's own zeros. With the
floors on, an interpreted request is **88 ms** now, and 35 ms once the
compiled tier has adopted. Read a pair of numbers from one row, never one from
each: they come from different runs. The start row predates the precompiled
standard library, which brought one capture to 13.6 s without the floor and
4.5 to 5.0 s with it.

Both processes keep almost nothing on their own Erlang heap, because the
module is a cache handle and the interpreter's memory is off-heap. The
collector sizes a heap from the live set, so it gives them the emulator's 233
words and then collects thousands of times through work that allocates
billions.

Three things to know:

- **The two options sit beside `root`, not inside `limits`.** A floor is not a
  bound, and one written into the map `wasm_python:limits/0` returns is
  ignored silently.
- **They are separate settings because they are separate processes.** CPython
  wants twice as much to capture as to answer, and a worker reading its image
  from `snapshot_dir` never captures at all, so it pays for the first and uses
  only the second.
- **Raise `max_heap_words` when you add the capture floor**, which is why the
  example above overrides it, and why the adapter sets no capture default: a
  default cannot know you raised the ceiling. `max_heap_words` bounds the peak
  and a floor raises the baseline that peak is measured from, so a ceiling
  that was comfortable without one can stop being comfortable with it.
  CPython at the adapter's own 16 M words is close enough to the edge that a
  floored capture dies **some** of the time: measured again on 0.7.0, four
  fresh starts at 2 M words, three captures died and one passed. A start that
  fails that way says `the capture died`, and names `max_heap_words` and the
  floor in its context so it is not a mystery.
- **CPython's request knee is five times QuickJS's**, which is why each
  adapter carries its own default. Pass `runner_min_heap_words` to change it,
  or `0` to turn it off. [The tuning guide](tuning.md) is how to find one for a
  different build.

## Errors

A tenant exception arrives as a value, with the traceback on stderr and the
message in the envelope:

```erlang
{error, #{class := adapter, kind := adapter_failure,
          msg := ~"boom",
          ctx := #{code := ~"exception", stdout := _, stderr := _}}}
```

`code` is a **binary**. Nothing a guest names becomes an atom, because the atom
table is node-wide and never reclaimed.
