PythonREPLTool Sandbox: Worker-Process Execution Model¶
Feature: FEAT-380 — Sandbox Hardening
Applies to: PythonREPLTool (parrot.tools.pythonrepl) and
PythonPandasTool (parrot.tools.pythonpandas)
Status: current as of this feature's merge (TASK-1939–1946)
PythonREPLTool no longer runs LLM-generated code in the host server
process. Each tool instance owns a persistent, per-instance worker
process — spawned, resource-limited, and torn down on timeout/crash —
that holds the REPL namespace and actually calls exec()/eval(). This
document describes that execution model, its failure modes, every deployment
knob, the namespace API that replaced direct .locals/.globals access,
and — read this if you deploy on Windows — the degraded guarantees
there.
What this is not. The worker is a resource-bounding sandbox, not an adversarial security boundary. It shares the kernel, network, and filesystem with the host process. It buys bounded resource consumption and blast-radius reduction (a runaway loop or a memory bomb kills the worker, never the server) — it does not buy containment against a deliberately malicious actor. Full container isolation (one container per session) is a documented future evolution, not implemented here.
1. Execution model¶
HOST (server process) WORKER (one per PythonREPLTool instance)
┌───────────────────────────┐ ┌──────────────────────────────┐
│ PythonREPLTool._execute() │ │ preexec: setrlimit( │
│ 1. sanitize_input() │ │ AS, CPU, NOFILE, CORE=0) │
│ 2. host gate (allowlist │ dedicated │ │
│ + AST denylist) ───────┼──── pipe ────►│ gate re-validated (defence │
│ — cheap reject, no │ (never │ in depth) before exec() │
│ round-trip if denied │ stdin/stdout)│ ns = worker's own locals │
│ 3. WorkerPool.acquire() │ │ _execute_code() — moved │
│ → WorkerHandle │◄──────────────┤ verbatim, unmodified │
│ 4. handle.execute(code) │ deadline_ms │ save_current_plot() writes │
│ enforces deadline_ms │ → SIGKILL │ to the shared output dir │
│ on the host side │ on timeout │ │
└───────────────────────────┘ └──────────────────────────────┘
shared output directory (plots, reports) — visible to both sides
- Host gate first.
PythonREPLTool._execute()runs the same allowlist + AST-denylist gate it always has, before the worker is ever contacted — denied code never starts a worker (cheap, no round-trip). - Worker re-validates. The worker independently re-runs the same gate
before calling
exec()— defence in depth in case a future caller reaches the worker without going through the host gate. _execute_code()moved, not rewritten. The method that actually runsexec()/eval()is unchanged; it now runs inside aPythonREPLToolinstance that lives in the worker process instead of the host.- One worker per tool instance. There is no broader "session" concept
in
PythonREPLTool— the tool instance itself is the isolation unit (matching its pre-existing per-instance.locals, never shared across instances). Lazy start: the worker is spawned on the first_execute()(or namespace-API) call, never in__init__. - Spawn only, never fork. Connection pools and the parent's
threads do not tolerate
fork(). - No automatic in-process fallback. If the worker cannot start,
_execute()returns an explicit error (see §2) — the tool never silently downgrades. The only way to run in-process is the explicit, loggedexecution_mode="inprocess"escape hatch (see §3b), chosen at construction time.
Readiness handshake (FEAT-500)¶
A freshly spawned worker is not usable yet: it still has to import the
parrot framework plus pandas and run the REPL bootstrap (~2.4 s on an idle
host, 12–14 s under heavy CPU contention — see
artifacts/logs/feat-500-bootstrap-profile.md).
So readiness is explicit:
- The worker writes exactly one
ReadyResponseframe — its pid plus the measuredbootstrap_ms— as the first frame on the control pipe, after its namespace is fully constructed and before it reads any request. It also logsrepl_worker: ready in <N> ms (...), entering service loop. - Host-side,
WorkerHandle.start()still returns as soon as the process is spawned, but it arms an internal readiness future: a background read of that first frame, bounded byWorkerConfig.bootstrap_timeout_ms. WorkerHandle._send()awaits readiness before writing anything, so every caller — namespace API,PythonPandasTool's DataFrame seeding,execute()— gets it for free. No request frame can reach a still-bootstrapping worker.WorkerPoolawaitshandle.wait_ready()before appending a spare to its prewarmed list, so "prewarmed" means ready, not merely spawned. TheWorkerPool: prewarmed worker ready (pid=..., bootstrap_ms=..., pool size=...)line is emitted only after the frame arrives.
await handle.wait_ready() and handle.is_ready are available to
integrators; awaiting readiness explicitly is optional (any request does it
implicitly).
Before FEAT-500 there was no handshake: the pool counted a still-booting worker as a ready spare, and the first namespace request into it timed out at a hard-coded 5 s and killed it — a self-sustaining restart loop on any host where bootstrap exceeded 5 s.
The one exception: execute_sync()¶
PythonREPLTool.execute_sync() is a separate, pre-existing synchronous
escape hatch that still calls _execute_code() in-process, unchanged.
It predates this feature and was deliberately left alone — hardening it
(routing it through the worker too) is future work. Anything calling
execute_sync() gets none of this document's guarantees (no rlimits, no
deadline, no isolation from the host process).
2. Failure modes¶
Every failure surfaces through the same return contract _execute() has
always had:
- Success → a plain
str. - Error →
{"status": "error" | "done_with_errors", "result": <text>, "error": <text>}.
| Cause | What happens | status |
result/error text |
|---|---|---|---|
| Code denied by the host or worker gate | Rejected before/without running | done_with_errors |
The SecurityError:/BlockedOperationError: message (unchanged from before this feature) |
| Runaway loop / hang, interruptible (FEAT-521) | Host's deadline_ms timer fires → SIGINT first (interrupt_before_kill=True, the default) → the worker converts it into a bounded reply and keeps its namespace |
error |
interrupted: exceeded deadline_ms=<N>; namespace preserved (partial side effects possible) — not a namespace-loss dict; known_vars is untouched |
| Runaway loop / hang, SIGINT-resistant or disabled | The SIGINT above gets no reply within interrupt_grace_ms (native code holding the GIL, or interrupt_before_kill=False) → deterministic SIGKILL (POSIX) / TerminateProcess (Windows) |
error |
REPL worker terminated (timeout: ...) — see below; names the observer's verdict and last sample (§2c) |
Allocation over rlimit_as_bytes |
Worker dies (often at import time under a very tight limit; see the calibration note in §3) | error |
REPL worker terminated (memory: ...) |
RSS over memory_hard_limit_bytes (FEAT-521) |
The host-side observer kills the worker deterministically, independent of rlimit_as_bytes/stderr — see §3c |
error |
REPL worker terminated (memory: RSS <measured> exceeded memory_hard_limit_bytes=<N>...) |
| Worker crash (segfault, OOM-killed, etc.) | Detected the next time the host tries to talk to it | error |
REPL worker terminated (crash: ...) |
| Concurrency ceiling reached | WorkerPool.acquire() raises immediately — no queueing |
error (wraps WorkerPoolExhaustedError) |
States the current ceiling and suggests raising WorkerConfig.max_workers |
Worker idle past idle_ttl_seconds |
Killed and unmapped by the pool's background sweep; next use spawns fresh | (not an error — the session's next call just gets a fresh, empty namespace) | — |
| Bootstrap timeout (FEAT-500) | bootstrap_timeout_ms expired without a ReadyResponse → the worker is killed (it never became a live worker) and every waiter gets a WorkerBootstrapError |
error on the execute() path (folded into the usual namespace-loss dict); a raised WorkerBootstrapError on the namespace API |
REPL worker pid=<pid> did not become ready within <N> ms (<cause>); stderr tail: <...> |
| Namespace-API timeout (FEAT-500) | namespace_timeout_ms expired on a non-exec request → NamespaceTimeoutError, worker left ALIVE, namespace preserved; the late reply is parked and drained before the next request |
(raises on the namespace API — no dict) | repl_worker[pid=<pid>]: '<op>' request did not answer within <N>s; the worker is still alive and the late reply will be drained on the next call |
Undrained straggler on the execute() path (FEAT-500) |
A reply parked by an earlier non-lethal timeout had still not arrived when execute()'s drain step ran. The worker is alive and its namespace is intact, so this is deliberately not reported as a namespace loss |
error |
repl_worker[pid=<pid>]: a reply from a previously timed-out request has still not arrived after <N>s; the worker is still alive and the reply stays queued for the next call — note the absence of the "ALL variables were lost" wording, which would be false here |
Every request is preceded by a drain step: if an earlier non-lethal
timeout left a reply in flight, the next request waits for that straggler
(bounded by its own budget) before writing, so a late reply can never be
handed to the wrong caller. execute() performs this drain too — which is
why it has its own row above — but its drain expiring never kills the worker.
Which timeouts kill the worker?¶
Only two, by design (FEAT-500 G2): a worker is expensive to replace and its namespace is the session's state, so losing them must be deliberate.
| Budget | Applies to | On expiry |
|---|---|---|
deadline_ms (+interrupt_grace_ms when interrupt_before_kill=True, +250 ms grace) |
execute() — running LLM code |
Two-stage as of FEAT-521 (see §2c "Two-stage deadline" below). Interrupt first if enabled and the worker is actually busy; SIGKILL + namespace-loss dict (cause="timeout") only as the fallback. |
bootstrap_timeout_ms |
waiting for the worker's first ReadyResponse |
Lethal. A process that cannot boot is not a live worker; killed + WorkerBootstrapError. FEAT-521: an optional bootstrap_stall_ms can fail this earlier — see §2c. |
namespace_timeout_ms |
get_var, set_var, list_vars, snapshot, reset, inject_dataframe |
Non-lethal. NamespaceTimeoutError; process and namespace survive. |
ping(timeout_s=...) |
health check only | Non-lethal. Returns False. |
memory_hard_limit_bytes (FEAT-521) |
continuous, independent of any in-flight request | Lethal. The observer kills the worker the poll cycle it crosses the limit — see §3c. |
2c. Observation & verdicts (FEAT-521)¶
Every live WorkerHandle owns a ProcessObserver (repl_worker/observer.py),
started right after the stdio drain task in start() and torn down in
kill(). It samples the child (CPU time, RSS, /proc state/wchan on
Linux, thread count) every observer_poll_ms (default 500 ms) via
psutil — host-side only; it never reads or writes the control pipe,
so no protocol reordering is possible (there is no heartbeat frame).
Verdicts¶
handle.observer.verdict() returns one of:
| Verdict | Meaning |
|---|---|
booting |
The observer has not yet been told the worker is ready (mark_idle() hasn't run) — the initial state for every worker. |
settled |
No request is in flight and CPU is flat — the worker is idle. |
computing |
A request is in flight and CPU has advanced within the trailing stall_window_ms window. |
stalled |
A request is in flight and CPU has been flat for the entire stall_window_ms window — busy is indistinguishable from hung until this fires. |
unavailable |
Non-POSIX host, or psutil could not resolve/read the process. Never blocks or fails a request — every codepath that consults the observer treats None/unavailable as "no information," not an error. |
handle.observer.describe() renders the current verdict plus the last
sample (cpu=, rss=, state=, wchan=, threads=) as one line — this
is what gets folded into every namespace-loss error and bootstrap-failure
message, so no error is ever blank about what the worker was doing.
Two-stage deadline (interrupt-before-kill)¶
On deadline_ms expiry, execute() no longer jumps straight to SIGKILL
when interrupt_before_kill=True (the default):
- SIGINT first — sent via
Popen.send_signal()on a dedicated lifecycle executor (never the pipe-reading one). Skipped entirely if the process is already dead, the host is non-POSIX, or the observer's verdict shows the worker isn't actually busy (settled/booting/unavailable— nothing to interrupt). - The worker's service loop converts the resulting
KeyboardInterruptinto a boundedExecResult(status="error")— the namespace survives (partial side effects from the interrupted snippet are possible and the message says so), and the reply lands on the same in-flight round-tripexecute()is already waiting on (no second pipe read). - If that reply doesn't land within
interrupt_grace_ms(native code holding the GIL, or a pathological allowlisted snippet) — deterministicSIGKILLfollows, exactly as before FEAT-521.
The caller is always answered within deadline_ms + interrupt_grace_ms +
250 ms grace, whichever path is taken — the same bound guarantee as
before, just with an extra namespace-preserving branch tried first. Set
interrupt_before_kill=False to restore the pre-FEAT-521 immediate-kill
behavior for every deadline breach.
Bootstrap diagnostics read from the observer, not a one-shot probe¶
_await_ready()'s timeout message no longer takes a one-shot /proc
snapshot at the moment of the kill — it reads the observer's sample ring,
which has been sampling since start():
REPL worker pid=31958 did not become ready within 30000 ms
(no ready frame within the bootstrap budget;
booting, cpu advanced 0.40 s in 30.0 s (starved));
stderr tail: ...
vs. a genuinely stuck child:
(no ready frame within the bootstrap budget;
booting, cpu flat since last sample, state=sleeping wchan=futex_wait_queue (stalled));
An optional bootstrap_stall_ms (default 0 = disabled) fails the
bootstrap earlier than bootstrap_timeout_ms once the observer sees
a sustained CPU-flat stall — a separate background task
(_watch_bootstrap_stall()) races the ordinary timeout path and resolves
WorkerBootstrapError first if it wins; both paths are safe to race since
resolving readiness is idempotent.
Namespace-loss error shape¶
The three worker-death causes (timeout / memory / crash) all produce the same structured shape, with the cause differentiated in the text:
{
"status": "error",
"result": "REPL worker terminated (timeout: execution exceeded deadline_ms=60000). ALL variables in this session were lost: x, y, df. You must recreate any state you need before retrying.",
"error": "REPL worker terminated (timeout: execution exceeded deadline_ms=60000). ALL variables in this session were lost: x, y, df. You must recreate any state you need before retrying."
}
The message always includes: the differentiated cause
(timeout/memory/crash), the list of variable names that existed
right before the kill (a cheap, names-only shadow the host maintains from
each successful call's new-variable list — never the values), and an
explicit instruction that the LLM must recreate state before retrying.
The session itself is still usable afterward — the next call transparently
gets a fresh worker.
3. Deployment configuration (WorkerConfig)¶
| Field | Default | Meaning |
|---|---|---|
rlimit_as_bytes |
12 GiB (12 * 1024**3) |
Virtual address space ceiling (RLIMIT_AS) applied to the worker. Applied by the worker itself in worker.main() (via apply_rlimits()) before any heavy import runs — not via Popen(preexec_fn=...). Empirically calibrated — see artifacts/logs/feat-380-rlimit-as-calibration.md for the measurements (peak observed VmPeak 5522.8 MB across a real bootstrap+500MB-load+merge+plot session, ×2 margin — predates FEAT-423's reduction of the REPL bootstrap import surface; actual footprint is now smaller, so this default is conservative). Re-run scripts/sdd/calibrate_rlimit_as.py after a pandas/numpy/pyarrow version bump (or to tighten this default post-FEAT-423). |
rlimit_cpu_seconds |
300 |
RLIMIT_CPU — a safety net if the host's own SIGKILL-on-timeout somehow failed to fire. |
rlimit_nofile |
256 |
RLIMIT_NOFILE — bounds file descriptors. |
deadline_ms |
60_000 |
Host-enforced wall-clock deadline per exec call. On expiry: SIGKILL + namespace-loss error (see §2). |
max_workers |
0 (→ max(4, cpu_count()), capped at 16) |
Concurrency ceiling across the pool. Reaching it makes acquire() raise immediately — no queueing. |
idle_ttl_seconds |
1800 (30 min) |
A session's worker idle past this is killed and unmapped by the pool's background sweep. |
prewarm_pool_size |
2 |
Idle, pre-booted spare workers (pandas/numpy already imported — the bootstrap import surface shrank as of FEAT-423) kept ready so a session's first call doesn't pay the 1–3s import cost. Since FEAT-500 a spare is only added to this pool after it signals readiness. |
bootstrap_timeout_ms |
30_000 |
FEAT-500. How long the host waits for a freshly spawned worker's ReadyResponse. Expiry is lethal (see §2). Do not lower this casually: bootstrap was measured at 12–14 s under 3× CPU oversubscription, and workers boot concurrently (1 session + prewarm_pool_size spares). Validated > 0. |
namespace_timeout_ms |
30_000 |
FEAT-500. Budget for every non-exec request (get_var/set_var/list_vars/snapshot/reset/inject_dataframe). Expiry is non-lethal: NamespaceTimeoutError, worker and namespace intact. Replaces the old hard-coded 5 s/10 s/30 s literals. Validated > 0. |
observer_poll_ms |
500 |
FEAT-521. How often the host samples a live worker's CPU/RSS/state via psutil (§2c). Never blocks a request — a disabled/unavailable observer just means less diagnostic detail. |
stall_window_ms |
5_000 |
FEAT-521. How long CPU must stay flat while a request is in flight before the verdict becomes stalled (§2c). |
bootstrap_stall_ms |
0 (disabled) |
FEAT-521. Fails the bootstrap early once the observer sees this many ms of CPU-flat stall — see §2c "Bootstrap diagnostics". 0 waits out the full bootstrap_timeout_ms as before. |
interrupt_before_kill |
True |
FEAT-521. Try SIGINT before SIGKILL on a deadline_ms breach (§2c "Two-stage deadline"). False restores the pre-FEAT-521 immediate-kill behavior. |
interrupt_grace_ms |
2_000 |
FEAT-521. How long the host waits for the SIGINT'd worker's reply before falling back to SIGKILL. Validated < deadline_ms (checked even when interrupt_before_kill=False). |
memory_soft_limit_bytes |
4 GiB (4 * 1024**3) |
FEAT-521. RSS threshold above which the next execute() result gets a one-line hint appended and the observer logs one WARNING per episode (§3c). 0 disables the soft limit. |
memory_hard_limit_bytes |
8 GiB (8 * 1024**3) |
FEAT-521. RSS threshold above which the observer kills the worker deterministically, cause memory (§3c). 0 disables the hard limit. Validated >= memory_soft_limit_bytes (when both enabled) and <= rlimit_as_bytes (when enabled) — empirically calibrated, see artifacts/logs/feat-521-memory-calibration.md. |
host_memory_reserve_bytes |
2 GiB (2 * 1024**3) |
FEAT-521. WorkerPool refuses to spawn or prewarm a new worker while psutil.virtual_memory().available (capped by cgroup v2 memory.max - memory.current when readable) is below this — §3c "Host memory reserve". |
RLIMIT_CORE = 0 is hardcoded, non-configurable — a core dump with live
DataFrames in memory is a data-exfiltration vector, not a tuning knob.
Tuning notes:
rlimit_as_bytesbounds virtual address space, not resident memory (Linux doesn't enforceRLIMIT_RSSat all — this is why the spec usesRLIMIT_AS). Setting it too low doesn't just reject huge allocations — it can crash the worker during its own bootstrap (numpy/pandas import), before any user code runs at all. Don't set it below the calibrated default without re-running the calibration script against your own pandas/numpy versions.max_workersbounds concurrent sessions, but each worker independently reserves up torlimit_as_bytesof virtual address space — planmax_workers × rlimit_as_bytesfor worst-case aggregate exposure on memory-constrained hosts, not just the per-worker number.prewarm_pool_sizeworkers count againstmax_workerstoo (the pool capssessions + prewarmed sparestogether).
3b. Execution modes — the inprocess escape hatch¶
PythonREPLTool (and therefore PythonPandasTool, PandasAgent, and every
agent that builds one internally) has two execution modes, fixed at
construction time:
| Mode | Where generated code runs | Sandbox | Selected by |
|---|---|---|---|
worker (default) |
a persistent child process (WorkerHandle/WorkerPool, everything in this document) |
rlimits, SIGKILL on deadline_ms, crash isolation, namespace-loss reporting |
default |
inprocess |
inside the host process, on the tool's own locals/globals, on its dedicated _repl_executor thread — the pre-FEAT-380 behaviour |
none of the above; the allowlist + AST denylist gate still applies | PythonREPLTool(execution_mode="inprocess") or PYTHON_REPL_EXECUTION_MODE=inprocess in the environment |
Resolution order: explicit constructor argument → PYTHON_REPL_EXECUTION_MODE
(read through navconfig, case-insensitive) → worker. Any other value raises
ValueError at construction. Selecting inprocess logs a WARNING naming
the tool, every time an instance is built.
# Deployment-wide kill switch — no agent code changes needed
export PYTHON_REPL_EXECUTION_MODE=inprocess
What inprocess keeps, so callers do not have to branch: the
{status, result, error} return contract, the namespace API
(get_var/set_var/list_vars/snapshot/inject_dataframe),
PythonPandasTool's DataFrame seeding and audit preview, and session
clones (each clone builds its own InProcessHandle over its own copied
namespace). WorkerPool is never instantiated in this mode.
What it gives up: a snippet that exceeds deadline_ms still returns the
bounded error to the caller, but the thread keeps running until the snippet
finishes (nothing can SIGKILL it) — the error text says so, and the
namespace is not reported as lost because nothing died; the handle stays
busy until that thread finishes, so a follow-up call gets a "still running"
error instead of mutating the namespace concurrently. A hard crash in
generated code (segfault, os._exit) takes the host down. And because
_execute_code captures output with redirect_stdout (process-global),
anything else the host prints to stdout during a snippet can leak into that
snippet's result — the pre-FEAT-380 behaviour, accepted as a documented
limitation of the hatch.
Intended use: a temporary, deployment-level escape hatch while the worker
path is being battle-tested on a given host, or hosts where spawning a
child interpreter is impossible. It is implemented as
parrot.tools.repl_worker.inprocess.InProcessHandle, a WorkerHandle
look-alike, so removing the hatch later is a one-branch change in
PythonREPLTool._get_worker_handle().
Instantiating a tool with a custom config¶
from parrot.tools.pythonrepl import PythonREPLTool
from parrot.tools.repl_worker.protocol import WorkerConfig
tool = PythonREPLTool(
worker_config=WorkerConfig(deadline_ms=15_000, max_workers=8),
)
3c. Memory guardrails (FEAT-521)¶
rlimit_as_bytes/RLIMIT_AS bounds virtual address space; it says
nothing about resident memory, and Linux does not enforce RLIMIT_RSS
at all — a cross join or a growing list of large arrays can inflate RSS
into the gigabytes while staying comfortably under the VA ceiling, driving
the host into swap. memory_soft_limit_bytes/memory_hard_limit_bytes
close that gap by having the host-side ProcessObserver watch RSS
directly (psutil.Process(pid).memory_info().rss), independent of
rlimit_as_bytes.
Soft breach (memory_soft_limit_bytes, default 4 GiB): the next
execute() result — whether a plain string or the result field of an
error dict — gets exactly one trailing line appended:
[REPL memory] RSS 4.30 GiB exceeds the 4.00 GiB soft limit — delete
DataFrames you no longer need (del name) before continuing.
(The error field itself is left unsuffixed — only result carries the
hint, so log/alerting code that keys off error sees the plain message.)
The observer also logs one WARNING per breach episode — re-armed only
once RSS drops below 90 % of the soft limit, so a worker sitting right
at the threshold doesn't spam a warning on every poll.
Hard breach (memory_hard_limit_bytes, default 8 GiB): the observer
kills the worker within one observer_poll_ms cycle. _classify_death()
consults the observer's recorded verdict before falling back to the
stderr-marker heuristic, so the memory cause is deterministic — it does
not depend on the killed process having written anything to stderr. An
idle worker over the hard limit is killed too (a leaked namespace is
still host memory); the pool's next acquire() for that session
transparently restarts it (a crash restart, cause named in the restart-loop
warning — see §4b).
Both defaults are empirically calibrated — see
artifacts/logs/feat-521-memory-calibration.md
for the measurements and the reconciliation with the approved 4 GiB soft /
8 GiB hard values.
Host memory reserve¶
WorkerPool also protects the host as a whole, independent of any
single worker's own limits: _top_up_prewarmed() and the spawn branch of
acquire() both check the effective available memory —
psutil.virtual_memory().available, capped by cgroup v2's
memory.max - memory.current when /sys/fs/cgroup/memory.max is readable
and not "max" (unlimited) — against host_memory_reserve_bytes (default
2 GiB):
- Below the reserve, prewarm top-up silently skips (DEBUG log) and
acquire()raisesWorkerPoolExhaustedErrornaming the pressure — but only when it would spawn. An existing live session and consuming an already-booted prewarmed spare are never blocked (neither one spawns anything new). - The maintenance sweep additionally evicts prewarmed spares (never
bound sessions) while under pressure, and logs the pool's aggregate RSS
(
WorkerPool.memory_summary()) at INFO every sweep whenever any worker is over its own soft limit.
summary = pool.memory_summary()
# {"workers": 3, "rss_total": 2415919104, "host_available": 8589934592}
Not covered¶
execution_mode="inprocess"(§3b) has no memory or observation guardrails — nothing can be measured or killed per snippet without process isolation.InProcessHandle.verdict()always reports"unavailable", and a debug log names the ignoredWorkerConfigfields the first time such a handle is built, so misconfiguration is visible without implying enforcement that doesn't exist.- Windows: see §5 below.
4. Namespace API (for integrators)¶
tool.locals / tool.globals are no longer the source of truth for
what the REPL namespace actually contains — the namespace lives in the
worker process. The host instance's .locals/.globals dicts still exist
and stay populated (for backward compatibility with any code not yet
ported), but they are a stale, construction-time snapshot — never
updated by code the worker actually executes. Reading them to discover
what a session computed will not see anything the LLM created.
Use the async namespace API instead:
value = await tool.get_var("my_dataframe")
await tool.set_var("previous_result", some_value)
names = await tool.list_vars()
snapshot = await tool.snapshot() # full, JSON-safe(ish) dump of the namespace
Any of these calls can raise NamespaceTimeoutError (FEAT-500) if the
worker does not answer within WorkerConfig.namespace_timeout_ms:
from parrot.tools.repl_worker import NamespaceTimeoutError, WorkerBootstrapError
try:
value = await tool.get_var("my_dataframe")
except NamespaceTimeoutError as exc:
# The worker is STILL ALIVE and its namespace is intact — retrying is
# legitimate. `exc` always carries a readable message naming the pid,
# the operation and the budget.
logger.warning("namespace read timed out: %s", exc)
It subclasses TimeoutError, so existing except TimeoutError handlers keep
working — but unlike the bare asyncio.TimeoutError this used to raise, its
str() is never empty. It is non-lethal: the process keeps running, the
namespace is preserved, and the straggling reply is drained before the next
request is written. WorkerBootstrapError (also always messaged) is raised
instead if the worker never finished booting at all.
There is no synchronous variant and no compatibility dict-proxy —
that was explicitly rejected (round-trip-per-key semantics and "looks live
but isn't" behavior would break silently and worse than an honest
AttributeError). Every call site that needs the namespace must be, or
become, async.
DataFrames specifically¶
await tool.inject_dataframe(name, df) pushes a pandas.DataFrame into the
worker via Arrow IPC over shared memory (falling back to pickle, with a
logged warning, only for dtypes Arrow can't represent) — cheaper than
set_var() for DataFrame-sized payloads, which always pickles. This is
what PythonPandasTool's own DataFrame seeding uses internally, and what
ToolManager.share_dataframe()/auto_push_to_pandas deliver through
transparently.
Snapshot semantics changed for WorkingMemoryToolkit wiring¶
BasicAgent auto-wires REPL/pandas tool namespaces into any registered
WorkingMemoryToolkit from configure() (its async setup hook — this
wiring needs to await tool.snapshot(), so it moved out of __init__,
which can't await). The wired dict is a snapshot frozen at wiring time,
not a live reference: DataFrames the agent loads after configure() runs
are not automatically visible through that registry entry. This is a
deliberate, spec-decided behavior change — a live cross-process reference
was never possible to begin with once the namespace moved into a separate
process.
4b. Restart-loop warning (FEAT-500)¶
A session whose worker keeps dying used to burn one replacement every few seconds in complete silence. The pool now counts restarts per session in a 60-second sliding window and, from the third one, logs:
WorkerPool: session 'pythonrepl-<uuid>' restarted 3 times in the last 60s —
possible restart loop (last worker exit code=-9, stderr tail='...')
FEAT-521: when the observer recorded a hard RSS breach on the dead worker, the warning names it explicitly instead of leaving you to guess from the exit code alone:
WorkerPool: session 'pythonrepl-<uuid>' restarted 3 times in the last 60s —
possible restart loop (last worker exit code=-9, stderr tail='...',
memory cause: rss=8993459200 limit=8589934592)
The per-session count is also readable programmatically:
pool.restart_count(session_id) # restarts in the last 60 s; observability only
exit_code, stderr_tail = handle.death_summary() # what the warning reports
Nothing branches on this value — it exists so the cause is visible. When you see it, check, in order:
- Bootstrap time on this host — see the procedure below. If
spawn→ready approaches
bootstrap_timeout_ms, spares are being killed for failing to boot; raise the budget or reduceprewarm_pool_size(fewer workers booting concurrently). - Host load — bootstrap is CPU/IO bound, and 12–14 s under 3× CPU oversubscription is measured, not hypothetical.
- The reported exit code, stderr tail, and memory cause —
-9means the host killed it (deadline, bootstrap timeout, or a FEAT-521 hard RSS breach — check for amemory cause:suffix, §3c); a traceback in the tail means the worker died on its own (e.g. an import failure, orRLIMIT_AStoo tight for pandas/numpy tommaptheir extensions). - The code being run — a genuine runaway loop hitting
deadline_mson every call is a correct restart loop; the warning is then telling you the LLM keeps submitting non-terminating code.
A related warning: an undersized shared executor¶
WorkerPool: the shared executor has 4 thread(s) but this pool can hold up to 6
live worker(s) (max_workers=4 + prewarm_pool_size=2), each of which can occupy
one thread for an entire blocking pipe read. ...
Each live worker can hold one thread of the shared executor for a whole
blocking pipe read — a request round-trip, or its readiness read while it
bootstraps. When there are fewer threads than possible live workers, requests
queue behind one another and the pool looks "slow" for reasons that have
nothing to do with the workers. Raise PythonREPLTool(executor_max_workers=…)
or lower prewarm_pool_size. The warning is advisory — nothing is clamped.
Process teardown is deliberately immune to this. WorkerHandle kills
workers on its own small, always-self-owned executor, never the shared one:
dispatching the SIGKILL to the same pool whose threads are parked on blocking
reads would mean the deadline_ms kill could never obtain a thread, while
freeing a thread required that kill to run — a deadlock in which the deadline
guarantee silently stops working. Keep that split if you refactor this class.
4c. Measuring worker bootstrap on your host (U3b)¶
Two independent measurements. The recorded baseline for both lives in
artifacts/logs/feat-500-bootstrap-profile.md.
Import cost — what the worker pays before it can serve anything:
python -X importtime -c "from parrot.tools.pythonrepl import PythonREPLTool" 2> importtime.log
sort -t'|' -k2 -n importtime.log | tail -25
On the reference dev box this totals 1.41 s, of which only ~0.22 s is
pandas — ≈80 % is parrot framework init (navconfig/vault/documentdb/events/
navigator auth) that the REPL child never uses. Trimming that surface is
deliberately out of scope for FEAT-500 and tracked as a follow-up spec.
Real spawn→ready — from a running server's logs, for one session:
grep -E "spawned worker pid=|repl_worker: ready|prewarmed worker ready|worker is dead|possible restart loop" server.log
A healthy cold start looks like this (one spawn, one ready, no restarts):
WorkerHandle: spawned worker pid=81228
repl_worker: ready in 2412 ms (max_workers config=0), entering service loop
WorkerPool: prewarmed worker ready (pid=81228, bootstrap_ms=2412, pool size=1)
bootstrap_ms is measured by the worker itself (monotonic, from main()
entry to readiness) and reported in the ReadyResponse frame, so it is the
number to compare against bootstrap_timeout_ms — no host-side clock
arithmetic needed. The WorkerHandle: spawned worker pid= and
prewarmed worker ready lines are at DEBUG level.
5. ⚠️ Windows degradation (read this before deploying on Windows)¶
On Windows, the worker gets none of the resource-limit guarantees this document otherwise describes.
| Guarantee | POSIX (Linux/macOS) | Windows |
|---|---|---|
| Separate process | ✅ | ✅ |
| Hard timeout enforcement | ✅ (SIGKILL) |
✅ (TerminateProcess, via subprocess.Popen.kill()) |
Memory ceiling (RLIMIT_AS) |
✅ | ❌ not enforced at all |
CPU ceiling (RLIMIT_CPU) |
✅ | ❌ not enforced at all |
File-descriptor ceiling (RLIMIT_NOFILE) |
✅ | ❌ not enforced at all |
No core dumps (RLIMIT_CORE=0) |
✅ | N/A (no Windows core-dump equivalent applied) |
| Process observation / verdicts (FEAT-521, §2c) | ✅ | ❌ ProcessObserver.verdict() always reports "unavailable" |
| Interrupt-before-kill (FEAT-521, §2c) | ✅ (interrupt_before_kill=True default) |
❌ interrupt() always returns False without sending anything — SIGINT semantics on Windows need a shared console and are unreliable; every deadline_ms breach goes straight to TerminateProcess, same as interrupt_before_kill=False on POSIX |
| Memory guardrails — soft hint / hard kill (FEAT-521, §3c) | ✅ | ❌ not enforced — no observer means no RSS sampling to trigger either one |
| Host memory reserve (FEAT-521, §3c) | ✅ (psutil.virtual_memory(), optionally cgroup v2) |
✅ psutil.virtual_memory() works cross-platform — this ONE guardrail still functions on Windows even though the per-worker ones above do not (cgroup v2 is Linux-only and silently contributes nothing there) |
resource.setrlimit() is POSIX-only. On import failure, the worker logs a
visible logger.warning ("rlimits are POSIX-only... running WITHOUT
memory/CPU/fd limits") and continues without any memory or CPU bound —
a runaway allocation or infinite loop on Windows is only stopped by the
host's deadline_ms timeout (which still works, since Popen.kill() maps
to TerminateProcess regardless of platform) or by Windows itself running
out of memory.
Do not deploy PythonREPLTool on Windows for untrusted/LLM-generated
code without understanding this gap. Windows Job Objects (which can
enforce memory/CPU/process-count limits, similar in spirit to POSIX
rlimits) are the natural next step and are tracked as future work — not
implemented in this feature.
6. History¶
- Module 1 (palliative, landed first, independently): replaced the
shared default
ThreadPoolExecutor(loop.run_in_executor(None, ...)) with a dedicated, bounded one (executor_max_workers, default 4) — so a runaway loop, even before the rest of this feature existed, could only exhaustPythonREPLTool's own thread pool, never the framework's shared one. That attribute (tool._repl_executor) still exists but is no longer on the code-execution path — the worker process replaced it as of Module 5. - Modules 2–9: the worker protocol,
WorkerHandle/WorkerPoollifecycle,PythonREPLToolintegration, the namespace-API port, Arrow DataFrame transport, and theRLIMIT_AScalibration this document reflects. - FEAT-500 — readiness handshake & non-lethal namespace timeouts: added
the
ReadyResponseframe plusWorkerHandle.wait_ready()/is_ready, so a worker is only used (and only counted as a prewarmed spare) once it has finished bootstrapping; made every non-exectimeout non-lethal (NamespaceTimeoutError, configurable vianamespace_timeout_ms, replacing hard-coded 5 s/10 s/30 s budgets that SIGKILLed the worker); addedbootstrap_timeout_msandWorkerBootstrapError; guaranteed every worker failure carries a readable message (no moreValueError('')reaching the LLM); and made restart loops visible (possible restart loopwarning +WorkerPool.restart_count()). Fixes a cold-start death spiral in which a host slower than the old 5 s budget could never produce a usable worker. Seesdd/specs/bug-workerpool-repl.spec.md. - Post-FEAT-500 —
inprocessescape hatch + bootstrap diagnostics: addedexecution_mode/PYTHON_REPL_EXECUTION_MODE(§3b,InProcessHandle) as an explicit, logged way to run the pre-worker in-process path while the worker pool is battle-tested; enrichedWorkerBootstrapErrorwith the child's/procstate, a thread-starvation cause, the stdout tail, and worker-side stage markers on stderr (§6 "Reading aWorkerBootstrapError"). - FEAT-521 — idle/busy detection & memory guardrails: added
ProcessObserver(§2c) giving every live worker a continuous, host-side busy/hung verdict (booting/settled/computing/stalled/unavailable) instead of the previous binary "reply arrived or deadline expired" signal. Made thedeadline_msbreach interruptible: SIGINT first (namespace-preserving),SIGKILLonly as the fallback (interrupt_before_kill, defaultTrue). Replaced the RLIMIT_AS-only memory story with RSS-based soft/hard guardrails (memory_soft_limit_bytes/memory_hard_limit_bytes, §3c) — a soft breach hints the LLM to free memory, a hard breach kills the worker deterministically (no more stderr-substring guessing for thememorycause). Added a host-wide memory reserve + prewarm-spare eviction under pressure (host_memory_reserve_bytes, cgroup v2-aware). Bootstrap diagnostics now read from the observer's continuous ring instead of a one-shot/procsnapshot, and an optionalbootstrap_stall_mscan fail a stalled bootstrap early. Seesdd/specs/repl-worker-idle-detection-memory-guardrails.spec.mdandartifacts/logs/feat-521-memory-calibration.md.
Reading a WorkerBootstrapError¶
Since the diagnostics pass that followed FEAT-500, the message a worker that
never sent its ready frame produces carries three extra facts. FEAT-521
changed the middle one: the process-state snapshot used to be a one-shot
/proc read taken at the moment of the kill; it now comes from the
observer's continuous sample ring (see §2c "Bootstrap diagnostics"), so it
can distinguish "never got CPU at all" (starved) from "got CPU, then
stopped advancing" (stalled) — a one-shot snapshot cannot tell those apart:
REPL worker pid=31958 did not become ready within 30000 ms
(no ready frame within the bootstrap budget;
booting, cpu flat since last sample, state=sleeping wchan=futex_wait_queue (stalled));
stderr tail: repl_worker[pid=31958] rlimits applied (...) | repl_worker[pid=31958] building namespace (...)
causedistinguishes a stuck child (no ready frame within the bootstrap budget) from thread starvation on the host (the readiness read never got an executor thread — all N thread(s) ... were busy): in the second case the worker may be perfectly healthy and the fix isexecutor_max_workers/prewarm_pool_size, not the worker.- The observer progress note (FEAT-521, replaces the old one-shot
process:snapshot) is one of two shapes:booting, cpu advanced <N> s in <M> s (starved)— the child never got scheduled — orbooting, cpu flat since last sample, state=<psutil status> wchan=<wchan> (stalled)(Linuxwchanonly) — it ran, then stopped making progress. A hugevmpeakin the earlier VA-based diagnostics meant thrashing againstrlimit_as_bytes; that check still exists (see §3) but is now independent of this line. stderr tailnow always contains the worker's own stage markers (interpreter up→rlimits applied→building namespace→namespace built ... sending ready frame), written unbuffered to stderr before any logging is configured, so the last line is the last stage the child reached.<empty>now means the interpreter never executed the first line ofworker.main()at all.
See also¶
sdd/specs/sandbox-hardening.spec.md— the full feature spec (design rationale, brainstorm decisions, acceptance criteria).docs/executors/docker-executor.md— a different, complementary isolation layer: routes an entire tool call (_execute()) to a remote Docker/K8s runtime via theparrot.tools.executorsframework. That mechanism relocates where_execute()runs; this document describes whatPythonREPLToolitself does within that call. The two can be combined.artifacts/logs/feat-500-bootstrap-profile.md— the worker bootstrap profile (import breakdown + spawn→ready timings) and the procedure for measuring both on your own host.artifacts/logs/feat-380-rlimit-as-calibration.md— theRLIMIT_AScalibration evidence referenced in §3.sdd/specs/repl-worker-idle-detection-memory-guardrails.spec.md— the FEAT-521 spec (observer design, two-stage deadline, memory guardrails, acceptance criteria) referenced throughout §2c/§3/§3c.artifacts/logs/feat-521-memory-calibration.md— thememory_soft_limit_bytes/memory_hard_limit_bytescalibration evidence referenced in §3c, plus the observation-overhead measurements for AC2.