docs: finalize rvbox v1 design and implementation plan
This commit is contained in:
+145
-16
@@ -1,5 +1,9 @@
|
||||
# RVBox v1 platform and operations contract
|
||||
|
||||
Daemon configuration uses strict TOML as specified in
|
||||
[`configuration.md`](configuration.md); the annotated examples contain every
|
||||
v1 knob and default.
|
||||
|
||||
## Unix-like clients
|
||||
|
||||
The client starts `sh` or `bash` in a new session/process group. Unix signals
|
||||
@@ -8,20 +12,99 @@ orderly shutdown, or recovery after an unclean daemon failure, managed command
|
||||
groups are terminated and marked interrupted because pipe capture cannot be
|
||||
safely resumed.
|
||||
|
||||
Both command text and uploaded scripts execute from generated private files
|
||||
beneath the effective CWD using exactly the selected executable (`sh FILE` or
|
||||
`bash FILE`). No user-supplied filename becomes a filesystem path. The wrapper
|
||||
file is removed during terminal cleanup.
|
||||
|
||||
Root-process exit begins a configurable 5-second drain grace period. RVBox waits
|
||||
for the supervised tree and capture pipes, then terminates residual group/cgroup
|
||||
members, drains to EOF, and only afterward emits the terminal lifecycle event.
|
||||
If capture still cannot reach EOF, it closes the handles and emits explicit
|
||||
incomplete-output metadata first. Shell-level detachment is not a supported way
|
||||
to leave descendants running; callers use RVBox background mode instead.
|
||||
|
||||
Launch uses an internal blocked launcher rather than starting requested command
|
||||
code directly. The launcher establishes its session/process group, reports its
|
||||
identity, and waits on a private release/watchdog channel. The client durably
|
||||
records `launch_prepared`, then durably records `launch_authorized`, and only
|
||||
then sends the release token. The launcher creates the requested shell inside
|
||||
that group and remains as a non-user-code watchdog until the tree exits. The
|
||||
daemon keeps the channel open for that lifetime: EOF before authorization exits
|
||||
without execution, while EOF after release terminates the group. Once
|
||||
`launch_authorized` is durable, recovery never retries that UUID; an uncertain
|
||||
launch is marked interrupted.
|
||||
|
||||
On Linux, create a per-command cgroup v2 for supervision even when no resource
|
||||
profile was requested, whenever the daemon has a delegated writable cgroup.
|
||||
Put the blocked launcher into that cgroup before release; use
|
||||
`clone3(CLONE_INTO_CGROUP | CLONE_PIDFD)` where available, otherwise migrate the
|
||||
still-blocked launcher through `cgroup.procs`. Persist the cgroup path, PID,
|
||||
process group, `/proc/<pid>/stat` start time, and launch generation. A live
|
||||
daemon uses the pidfd where available. Recovery uses `cgroup.kill` as the primary
|
||||
tree-cleanup operation and verifies the recorded birth identity before any
|
||||
PID/process-group fallback. Without cgroup delegation it uses the generic
|
||||
watchdog/process-group fallback unless a requested profile requires cgroup
|
||||
controls, in which case acceptance fails as unsupported. It never signals a
|
||||
process based only on a persisted numeric PID or PGID.
|
||||
|
||||
Other Unix-like systems use the same launch barrier plus a watchdog control
|
||||
channel whose EOF triggers process-group termination. Recovery validates the
|
||||
platform's process-birth identity before signaling. Descendants that deliberately
|
||||
create a new session may escape this generic fallback, so complete tree cleanup
|
||||
outside Linux cgroup supervision is best-effort; the at-most-once launch
|
||||
guarantee still applies.
|
||||
|
||||
Linux diagnostics sample `/proc/<pid>` and relevant children for state, CPU,
|
||||
resident memory, I/O counters, CWD, and wait-channel information when readable.
|
||||
These values may be unavailable due to permissions, kernel configuration, or a
|
||||
short-lived process; absence is represented explicitly rather than fabricated.
|
||||
Cgroup v2 is used for requested resource profiles only when available.
|
||||
Cgroup v2 profile limits are applied only when requested; a no-profile
|
||||
supervisory cgroup imposes no resource limit.
|
||||
|
||||
## Windows clients
|
||||
|
||||
The client launches `cmd` or `powershell` in an appropriate dedicated console
|
||||
process group and assigns the root process to a per-command Job Object. Child
|
||||
processes normally join the Job Object. Job Object limits enforce requested
|
||||
profiles and `KILL_ON_JOB_CLOSE` protects against lost supervision.
|
||||
The minimum supported v1 Windows versions are Windows 10 and Windows Server
|
||||
2016. For every command, create a non-inheritable Job Object, set
|
||||
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`, and do not enable breakaway. The daemon
|
||||
starts an RVBox per-command launcher suspended with `CREATE_NEW_CONSOLE`,
|
||||
`CREATE_UNICODE_ENVIRONMENT`, and `EXTENDED_STARTUPINFO_PRESENT`; it assigns the
|
||||
launcher atomically through `PROC_THREAD_ATTRIBUTE_JOB_LIST`. Only an explicit
|
||||
standard-I/O and launcher-control handle list is inherited, and the Job handle
|
||||
is never inherited.
|
||||
|
||||
Only `SIGTERM` and `SIGKILL` are accepted. `SIGTERM` attempts `CTRL_BREAK_EVENT`
|
||||
The launcher invokes exactly the selected shell against the generated wrapper:
|
||||
`cmd.exe /D /S /C` for a `.cmd` wrapper, or `powershell.exe` with `-NoLogo`,
|
||||
`-NoProfile`, `-NonInteractive`, and `-File` for a `.ps1` wrapper. Application
|
||||
paths and argument quoting are constructed by the Windows launcher, never by
|
||||
concatenating an untrusted command line. There is no fallback between shells.
|
||||
|
||||
Persist and flush `launch_prepared` with the launcher PID,
|
||||
`GetProcessTimes` creation `FILETIME`, and launch generation. Persist and flush
|
||||
`launch_authorized` before calling `ResumeThread`. The launcher then starts the
|
||||
requested `cmd` or `powershell` suspended in its console with
|
||||
`CREATE_NEW_PROCESS_GROUP`, reports the shell PID/group through the private
|
||||
control channel, connects the allowlisted pipes, and resumes it. This two-step
|
||||
shape is required because `CREATE_NEW_PROCESS_GROUP` is ignored when combined
|
||||
with `CREATE_NEW_CONSOLE`, and console control events reach only groups sharing
|
||||
the caller's console. Failure after authorization terminates the Job and is
|
||||
reported as interrupted; it never redispatches the UUID.
|
||||
|
||||
The launcher remains the in-console signal proxy and calls
|
||||
`GenerateConsoleCtrlEvent(CTRL_BREAK_EVENT, shell_group_id)` on request. The
|
||||
daemon retains the sole Job handle, so an unclean daemon exit closes the last
|
||||
handle and terminates the launcher, shell, and descendants. Recovery never
|
||||
kills by persisted PID alone; the PID/creation-time tuple is diagnostic evidence
|
||||
for PID reuse or cleanup anomalies. Child processes normally join the Job.
|
||||
Job Object limits enforce requested profiles and `KILL_ON_JOB_CLOSE` protects
|
||||
against lost supervision.
|
||||
|
||||
Root-process exit begins the same drain grace period. Completion waits for the
|
||||
Job Object to reach zero active processes; after the grace period RVBox
|
||||
terminates the Job, drains its capture handles, records any incomplete-output
|
||||
marker, and emits the terminal lifecycle event last.
|
||||
|
||||
Only `TERM`/`SIGTERM` and `KILL`/`SIGKILL` are accepted. `TERM` attempts `CTRL_BREAK_EVENT`
|
||||
and waits 10 seconds, then calls Job Object termination if the job persists;
|
||||
`SIGKILL` calls Job Object termination immediately. A console signal is
|
||||
best-effort, so callers receive an explicit escalation result. Windows status
|
||||
@@ -30,12 +113,36 @@ diagnostics such as an I/O wait channel.
|
||||
|
||||
## Storage and recovery
|
||||
|
||||
SQLite runs in WAL mode with integrity checking on startup. Output segments are
|
||||
written atomically, fsynced according to the configured durability interval, and
|
||||
indexed only after successful durable append. Startup scans/repairs incomplete
|
||||
tail records before accepting control requests. Segment compression is Zstandard;
|
||||
limits always measure stored compressed bytes, while clients expose raw byte
|
||||
counts separately.
|
||||
SQLite runs in WAL mode. Every append-only segment has a SQLite-owned
|
||||
`committed_end_offset`. The writer validates and appends records, syncs the file
|
||||
(grouped by the configured durability interval), and only then commits event
|
||||
metadata plus the new offset in SQLite. An acknowledgement waits for both
|
||||
steps. Therefore a crash can leave an uncommitted file tail, but cannot validly
|
||||
acknowledge metadata whose bytes were not durable.
|
||||
|
||||
Startup acquires the instance lock and binds liveness/diagnostic endpoints, then
|
||||
runs integrity and segment recovery asynchronously. A file longer than its
|
||||
committed offset is safely truncated to that offset. A file shorter than the
|
||||
offset, a checksum failure inside the committed range, or corrupt essential
|
||||
metadata creates a durable scoped storage incident; affected output is marked
|
||||
truncated/incomplete and affected active commands are interrupted when their
|
||||
essential state cannot be trusted. Healthy scopes remain usable. Readiness is
|
||||
false and mutations requiring an unrecovered or dirty scope return `UNAVAILABLE`,
|
||||
but process startup, liveness, incident inspection, and unaffected work do not
|
||||
wait for a full-store scan.
|
||||
|
||||
Safe repairs are attempted automatically and can also be requested online with
|
||||
`rvc storage repair`. Irrecoverable loss stays dirty until explicitly accepted
|
||||
with `rvc storage acknowledge`; an offline server has equivalent
|
||||
`rvbox-server repair --data-dir ...` repair/list/acknowledge operations. Client
|
||||
spool recovery follows the same committed-offset rule and exposes equivalent
|
||||
offline `rvbox repair --state-dir ...` operations and local health diagnostics.
|
||||
Resolving an incident clears derived dirty health but never erases the incident
|
||||
or audit history as part of resolution. Unresolved compact incident records are
|
||||
non-evictable; resolved incident/audit history follows the separate 100 MiB
|
||||
rotation. Segment compression is Zstandard;
|
||||
limits measure stored compressed bytes, while raw byte counts are reported
|
||||
separately.
|
||||
|
||||
The system must reserve headroom before writes and use transactional metadata
|
||||
updates. Storage-full, permission, and corruption failures are surfaced as
|
||||
@@ -43,6 +150,12 @@ structured server/client health states and audit events. They must isolate the
|
||||
affected command/session, reject work when needed, and keep the daemon's
|
||||
heartbeat/control loops alive.
|
||||
|
||||
V1 state is plaintext at rest, including command text, scripts, environment
|
||||
override values, stdin, and output. Private directory/file modes and dedicated
|
||||
daemon accounts are deployment hygiene, not an application-level encryption
|
||||
guarantee. Backups copy the same plaintext sensitivity. Encryption and external
|
||||
key management are future-version work.
|
||||
|
||||
## Metrics, logging, and safe defaults
|
||||
|
||||
Both daemons should emit structured logs and metrics for session transitions,
|
||||
@@ -52,10 +165,26 @@ and protocol violations. Never emit stdin or raw output in normal daemon logs.
|
||||
|
||||
Recommended configuration defaults are: 10-second heartbeat idle period,
|
||||
30-second liveness timeout, 1–60-second full-jitter reconnect backoff,
|
||||
60-second stable-session reset, 16 running/100 queued commands per client,
|
||||
10 MiB per-command compressed window, 50 MiB per-client active spool and server
|
||||
history, 1 GiB server history, 64 KiB uncompressed stream chunk, 1 MiB decoded
|
||||
envelope, and 10 MiB script maximum.
|
||||
60-second stable-session reset, 5-minute one-shot live-conflict takeover grant,
|
||||
16 running/100 queued commands per client, 1,000 server-queued commands per
|
||||
target client and 10,000 server-wide,
|
||||
15-minute queue TTL, 10 MiB per-command output window, 32 MiB total per command,
|
||||
256 MiB per client/client daemon, 4 GiB server-wide command storage, 30-day
|
||||
terminal retention, 100 MiB audit storage with quota-only rotation by default,
|
||||
one million compact command tombstones, 64 KiB uncompressed stream chunk, 1 MiB
|
||||
decoded agent envelope,
|
||||
1 MiB/256 KiB per-command raw-output high/low watermarks, 8 MiB/4 MiB per-client
|
||||
watermarks, 64 MiB/32 MiB server-wide watermarks, 1 MiB per-command and 8 MiB
|
||||
per-session unacknowledged send windows, and 10 MiB raw script maximum. Control
|
||||
gRPC accepts at most 16 MiB decoded requests; the JSON-RPC adapter accepts at
|
||||
most 24 MiB HTTP bodies to allow protobuf JSON's base64 expansion while
|
||||
retaining the same decoded field limits. Each active command reserves 64 KiB
|
||||
within its quota for terminal/loss closeout metadata; protocol detail/reason and
|
||||
incident-note text fields are individually limited to 4 KiB. Default emergency
|
||||
filesystem free-space floors are 256 MiB on the server and 64 MiB on a client;
|
||||
crossing one rejects new unreserved allocations even if the logical quota has
|
||||
headroom. Already-reserved terminal/loss closeout remains writable while bytes
|
||||
physically remain.
|
||||
|
||||
These bounds protect RVBox's own loops; they cannot make arbitrary child
|
||||
commands harmless when no resource profile is requested. Operators should
|
||||
|
||||
Reference in New Issue
Block a user