docs: add initial RVBox design and protocol contracts
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
# RVBox v1 platform and operations contract
|
||||
|
||||
## Unix-like clients
|
||||
|
||||
The client starts `sh` or `bash` in a new session/process group. Unix signals
|
||||
address that group, so normally created descendants receive the signal too. On
|
||||
orderly shutdown, or recovery after an unclean daemon failure, managed command
|
||||
groups are terminated and marked interrupted because pipe capture cannot be
|
||||
safely resumed.
|
||||
|
||||
Linux diagnostics sample `/proc/<pid>` and relevant children for state, CPU,
|
||||
resident memory, I/O counters, CWD, and wait-channel information when readable.
|
||||
These values may be unavailable due to permissions, kernel configuration, or a
|
||||
short-lived process; absence is represented explicitly rather than fabricated.
|
||||
Cgroup v2 is used for requested resource profiles only when available.
|
||||
|
||||
## Windows clients
|
||||
|
||||
The client launches `cmd` or `powershell` in an appropriate dedicated console
|
||||
process group and assigns the root process to a per-command Job Object. Child
|
||||
processes normally join the Job Object. Job Object limits enforce requested
|
||||
profiles and `KILL_ON_JOB_CLOSE` protects against lost supervision.
|
||||
|
||||
Only `SIGTERM` and `SIGKILL` are accepted. `SIGTERM` attempts `CTRL_BREAK_EVENT`
|
||||
and waits 10 seconds, then calls Job Object termination if the job persists;
|
||||
`SIGKILL` calls Job Object termination immediately. A console signal is
|
||||
best-effort, so callers receive an explicit escalation result. Windows status
|
||||
uses process and Job Object accounting APIs; it does not claim Linux-only
|
||||
diagnostics such as an I/O wait channel.
|
||||
|
||||
## Storage and recovery
|
||||
|
||||
SQLite runs in WAL mode with integrity checking on startup. Output segments are
|
||||
written atomically, fsynced according to the configured durability interval, and
|
||||
indexed only after successful durable append. Startup scans/repairs incomplete
|
||||
tail records before accepting control requests. Segment compression is Zstandard;
|
||||
limits always measure stored compressed bytes, while clients expose raw byte
|
||||
counts separately.
|
||||
|
||||
The system must reserve headroom before writes and use transactional metadata
|
||||
updates. Storage-full, permission, and corruption failures are surfaced as
|
||||
structured server/client health states and audit events. They must isolate the
|
||||
affected command/session, reject work when needed, and keep the daemon's
|
||||
heartbeat/control loops alive.
|
||||
|
||||
## Metrics, logging, and safe defaults
|
||||
|
||||
Both daemons should emit structured logs and metrics for session transitions,
|
||||
heartbeat timeout, reconnect backoff, command state transitions, queue depth,
|
||||
spool bytes, segment rotation/eviction, output loss markers, storage errors,
|
||||
and protocol violations. Never emit stdin or raw output in normal daemon logs.
|
||||
|
||||
Recommended configuration defaults are: 10-second heartbeat idle period,
|
||||
30-second liveness timeout, 1–60-second full-jitter reconnect backoff,
|
||||
60-second stable-session reset, 16 running/100 queued commands per client,
|
||||
10 MiB per-command compressed window, 50 MiB per-client active spool and server
|
||||
history, 1 GiB server history, 64 KiB uncompressed stream chunk, 1 MiB decoded
|
||||
envelope, and 10 MiB script maximum.
|
||||
|
||||
These bounds protect RVBox's own loops; they cannot make arbitrary child
|
||||
commands harmless when no resource profile is requested. Operators should
|
||||
enable resource profiles for untrusted or expensive workloads and keep nginx,
|
||||
Unix-socket permissions, filesystem capacity, and service supervision correctly
|
||||
configured.
|
||||
Reference in New Issue
Block a user