Files
rvbox/docs/platform-and-operations.md
T

194 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RVBox v1 platform and operations contract
Daemon configuration uses strict TOML as specified in
[`configuration.md`](configuration.md); the annotated examples contain every
v1 knob and default.
## Unix-like clients
The client starts `sh` or `bash` in a new session/process group. Unix signals
address that group, so normally created descendants receive the signal too. On
orderly shutdown, or recovery after an unclean daemon failure, managed command
groups are terminated and marked interrupted because pipe capture cannot be
safely resumed.
Both command text and uploaded scripts execute from generated private files
beneath the effective CWD using exactly the selected executable (`sh FILE` or
`bash FILE`). No user-supplied filename becomes a filesystem path. The wrapper
file is removed during terminal cleanup.
Root-process exit begins a configurable 5-second drain grace period. RVBox waits
for the supervised tree and capture pipes, then terminates residual group/cgroup
members, drains to EOF, and only afterward emits the terminal lifecycle event.
If capture still cannot reach EOF, it closes the handles and emits explicit
incomplete-output metadata first. Shell-level detachment is not a supported way
to leave descendants running; callers use RVBox background mode instead.
Launch uses an internal blocked launcher rather than starting requested command
code directly. The launcher establishes its session/process group, reports its
identity, and waits on a private release/watchdog channel. The client durably
records `launch_prepared`, then durably records `launch_authorized`, and only
then sends the release token. The launcher creates the requested shell inside
that group and remains as a non-user-code watchdog until the tree exits. The
daemon keeps the channel open for that lifetime: EOF before authorization exits
without execution, while EOF after release terminates the group. Once
`launch_authorized` is durable, recovery never retries that UUID; an uncertain
launch is marked interrupted.
On Linux, create a per-command cgroup v2 for supervision even when no resource
profile was requested, whenever the daemon has a delegated writable cgroup.
Put the blocked launcher into that cgroup before release; use
`clone3(CLONE_INTO_CGROUP | CLONE_PIDFD)` where available, otherwise migrate the
still-blocked launcher through `cgroup.procs`. Persist the cgroup path, PID,
process group, `/proc/<pid>/stat` start time, and launch generation. A live
daemon uses the pidfd where available. Recovery uses `cgroup.kill` as the primary
tree-cleanup operation and verifies the recorded birth identity before any
PID/process-group fallback. Without cgroup delegation it uses the generic
watchdog/process-group fallback unless a requested profile requires cgroup
controls, in which case acceptance fails as unsupported. It never signals a
process based only on a persisted numeric PID or PGID.
Other Unix-like systems use the same launch barrier plus a watchdog control
channel whose EOF triggers process-group termination. Recovery validates the
platform's process-birth identity before signaling. Descendants that deliberately
create a new session may escape this generic fallback, so complete tree cleanup
outside Linux cgroup supervision is best-effort; the at-most-once launch
guarantee still applies.
Linux diagnostics sample `/proc/<pid>` and relevant children for state, CPU,
resident memory, I/O counters, CWD, and wait-channel information when readable.
These values may be unavailable due to permissions, kernel configuration, or a
short-lived process; absence is represented explicitly rather than fabricated.
Cgroup v2 profile limits are applied only when requested; a no-profile
supervisory cgroup imposes no resource limit.
## Windows clients
The minimum supported v1 Windows versions are Windows 10 and Windows Server
2016. For every command, create a non-inheritable Job Object, set
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`, and do not enable breakaway. The daemon
starts an RVBox per-command launcher suspended with `CREATE_NEW_CONSOLE`,
`CREATE_UNICODE_ENVIRONMENT`, and `EXTENDED_STARTUPINFO_PRESENT`; it assigns the
launcher atomically through `PROC_THREAD_ATTRIBUTE_JOB_LIST`. Only an explicit
standard-I/O and launcher-control handle list is inherited, and the Job handle
is never inherited.
The launcher invokes exactly the selected shell against the generated wrapper:
`cmd.exe /D /S /C` for a `.cmd` wrapper, or `powershell.exe` with `-NoLogo`,
`-NoProfile`, `-NonInteractive`, and `-File` for a `.ps1` wrapper. Application
paths and argument quoting are constructed by the Windows launcher, never by
concatenating an untrusted command line. There is no fallback between shells.
Persist and flush `launch_prepared` with the launcher PID,
`GetProcessTimes` creation `FILETIME`, and launch generation. Persist and flush
`launch_authorized` before calling `ResumeThread`. The launcher then starts the
requested `cmd` or `powershell` suspended in its console with
`CREATE_NEW_PROCESS_GROUP`, reports the shell PID/group through the private
control channel, connects the allowlisted pipes, and resumes it. This two-step
shape is required because `CREATE_NEW_PROCESS_GROUP` is ignored when combined
with `CREATE_NEW_CONSOLE`, and console control events reach only groups sharing
the caller's console. Failure after authorization terminates the Job and is
reported as interrupted; it never redispatches the UUID.
The launcher remains the in-console signal proxy and calls
`GenerateConsoleCtrlEvent(CTRL_BREAK_EVENT, shell_group_id)` on request. The
daemon retains the sole Job handle, so an unclean daemon exit closes the last
handle and terminates the launcher, shell, and descendants. Recovery never
kills by persisted PID alone; the PID/creation-time tuple is diagnostic evidence
for PID reuse or cleanup anomalies. Child processes normally join the Job.
Job Object limits enforce requested profiles and `KILL_ON_JOB_CLOSE` protects
against lost supervision.
Root-process exit begins the same drain grace period. Completion waits for the
Job Object to reach zero active processes; after the grace period RVBox
terminates the Job, drains its capture handles, records any incomplete-output
marker, and emits the terminal lifecycle event last.
Only `TERM`/`SIGTERM` and `KILL`/`SIGKILL` are accepted. `TERM` attempts `CTRL_BREAK_EVENT`
and waits 10 seconds, then calls Job Object termination if the job persists;
`SIGKILL` calls Job Object termination immediately. A console signal is
best-effort, so callers receive an explicit escalation result. Windows status
uses process and Job Object accounting APIs; it does not claim Linux-only
diagnostics such as an I/O wait channel.
## Storage and recovery
SQLite runs in WAL mode. Every append-only segment has a SQLite-owned
`committed_end_offset`. The writer validates and appends records, syncs the file
(grouped by the configured durability interval), and only then commits event
metadata plus the new offset in SQLite. An acknowledgement waits for both
steps. Therefore a crash can leave an uncommitted file tail, but cannot validly
acknowledge metadata whose bytes were not durable.
Startup acquires the instance lock and binds liveness/diagnostic endpoints, then
runs integrity and segment recovery asynchronously. A file longer than its
committed offset is safely truncated to that offset. A file shorter than the
offset, a checksum failure inside the committed range, or corrupt essential
metadata creates a durable scoped storage incident; affected output is marked
truncated/incomplete and affected active commands are interrupted when their
essential state cannot be trusted. Healthy scopes remain usable. Readiness is
false and mutations requiring an unrecovered or dirty scope return `UNAVAILABLE`,
but process startup, liveness, incident inspection, and unaffected work do not
wait for a full-store scan.
Safe repairs are attempted automatically and can also be requested online with
`rvc storage repair`. Irrecoverable loss stays dirty until explicitly accepted
with `rvc storage acknowledge`; an offline server has equivalent
`rvbox-server repair --data-dir ...` repair/list/acknowledge operations. Client
spool recovery follows the same committed-offset rule and exposes equivalent
offline `rvbox repair --state-dir ...` operations and local health diagnostics.
Resolving an incident clears derived dirty health but never erases the incident
or audit history as part of resolution. Unresolved compact incident records are
non-evictable; resolved incident/audit history follows the separate 100 MiB
rotation. Segment compression is Zstandard;
limits measure stored compressed bytes, while raw byte counts are reported
separately.
The system must reserve headroom before writes and use transactional metadata
updates. Storage-full, permission, and corruption failures are surfaced as
structured server/client health states and audit events. They must isolate the
affected command/session, reject work when needed, and keep the daemon's
heartbeat/control loops alive.
V1 state is plaintext at rest, including command text, scripts, environment
override values, stdin, and output. Private directory/file modes and dedicated
daemon accounts are deployment hygiene, not an application-level encryption
guarantee. Backups copy the same plaintext sensitivity. Encryption and external
key management are future-version work.
## Metrics, logging, and safe defaults
Both daemons should emit structured logs and metrics for session transitions,
heartbeat timeout, reconnect backoff, command state transitions, queue depth,
spool bytes, segment rotation/eviction, output loss markers, storage errors,
and protocol violations. Never emit stdin or raw output in normal daemon logs.
Recommended configuration defaults are: 10-second heartbeat idle period,
30-second liveness timeout, 1–60-second full-jitter reconnect backoff,
60-second stable-session reset, 5-minute one-shot live-conflict takeover grant,
16 running/100 queued commands per client, 1,000 server-queued commands per
target client and 10,000 server-wide,
15-minute queue TTL, 10 MiB per-command output window, 32 MiB total per command,
256 MiB per client/client daemon, 4 GiB server-wide command storage, 30-day
terminal retention, 100 MiB audit storage with quota-only rotation by default,
one million compact command tombstones, 64 KiB uncompressed stream chunk, 1 MiB
decoded agent envelope,
1 MiB/256 KiB per-command raw-output high/low watermarks, 8 MiB/4 MiB per-client
watermarks, 64 MiB/32 MiB server-wide watermarks, 1 MiB per-command and 8 MiB
per-session unacknowledged send windows, and 10 MiB raw script maximum. Control
gRPC accepts at most 16 MiB decoded requests; the JSON-RPC adapter accepts at
most 24 MiB HTTP bodies to allow protobuf JSON's base64 expansion while
retaining the same decoded field limits. Each active command reserves 64 KiB
within its quota for terminal/loss closeout metadata; protocol detail/reason and
incident-note text fields are individually limited to 4 KiB. Default emergency
filesystem free-space floors are 256 MiB on the server and 64 MiB on a client;
crossing one rejects new unreserved allocations even if the logical quota has
headroom. Already-reserved terminal/loss closeout remains writable while bytes
physically remain.
These bounds protect RVBox's own loops; they cannot make arbitrary child
commands harmless when no resource profile is requested. Operators should
enable resource profiles for untrusted or expensive workloads and keep nginx,
Unix-socket permissions, filesystem capacity, and service supervision correctly
configured.