docs: finalize rvbox v1 design and implementation plan
This commit is contained in:
+171
-55
@@ -12,6 +12,9 @@ RVBox deliberately executes arbitrary commands with the identity, permissions,
|
||||
and base environment of the client daemon. It is therefore an administrative
|
||||
tool, not a multi-tenant remote-execution service.
|
||||
|
||||
Both daemons use strict TOML 1.0 configuration. The normative schema and fully
|
||||
annotated examples are in the [configuration contract](configuration.md).
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist
|
||||
@@ -26,10 +29,14 @@ tool, not a multi-tenant remote-execution service.
|
||||
|
||||
The server accepts a self-reported hostname as `client_id`; it is an opaque
|
||||
1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer
|
||||
registration for an ID replaces its prior live session. Consequently, a peer
|
||||
able to reach nginx can impersonate or take over a client ID. This is an
|
||||
accepted v1 limitation and deployments must restrict the endpoint to a trusted
|
||||
network.
|
||||
registration from the same durable client instance replaces its prior live
|
||||
session. A different instance is accepted normally when no session for that
|
||||
client ID is live. While one is live, a different instance is rejected unless
|
||||
an operator grants a one-shot override for that exact pending instance. This
|
||||
prevents accidental hostname collisions but is not authentication: a peer able
|
||||
to reach nginx and copy or guess the identifiers can still impersonate a
|
||||
client. This is an accepted v1 limitation and deployments must restrict the
|
||||
endpoint to a trusted network.
|
||||
|
||||
The Unix control socket is local-only and mode `0600`, owned by the server
|
||||
account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated;
|
||||
@@ -57,53 +64,113 @@ its dispatch loops.
|
||||
|
||||
1. The client connects over WSS and sends `ClientHello` with its client ID,
|
||||
protocol capability, OS/architecture, daemon version, current daemon CWD,
|
||||
supported shells, and a fresh reconnect UUID.
|
||||
supported shells, and a durable random client-instance UUID.
|
||||
2. The server accepts the current compatible protocol version, fences the
|
||||
previous connection for that ID, and returns a fresh server-issued
|
||||
`session_id` and monotonic `session_generation`.
|
||||
previous connection for that ID and same client instance, and returns a
|
||||
fresh server-issued `session_id` and monotonic `session_generation`. A live
|
||||
claim from a different client instance is rejected while the current session
|
||||
is live unless an operator has explicitly authorized that pending instance.
|
||||
With no live session, the new instance is accepted normally.
|
||||
3. Every client-to-server envelope and server dispatch is bound to that token.
|
||||
The server discards traffic from superseded sessions, including late output.
|
||||
4. The server queues work while a client is offline and dispatches it only when
|
||||
the active session advertises capacity. A replacement session immediately
|
||||
resumes non-terminal reconciliation.
|
||||
performs bidirectional reconciliation between the server's non-terminal set
|
||||
and the client's complete retained-command set. The server returns explicit
|
||||
local terminate/discard decisions before new dispatch begins.
|
||||
|
||||
## Command model
|
||||
|
||||
Each user request has a server-generated UUID (`issue_uuid`) and a durable
|
||||
request record: target client, request/issue timestamps, shell type, command
|
||||
text or script descriptor, CWD, environment overrides, resource-profile flags,
|
||||
and lifecycle state. The UUID is the end-to-end idempotency key. Transport is
|
||||
at-least-once, but the client durably remembers accepted UUIDs and never starts
|
||||
the same request twice.
|
||||
Each user request has a UUID (`issue_uuid`) and a durable request record: target
|
||||
client, request/issue timestamps, shell type, command text or script descriptor,
|
||||
CWD, environment overrides, resource-profile flags, and lifecycle state. `rvc`
|
||||
normally supplies this UUID as its optional `request_id`; the server generates
|
||||
one when it is omitted. Command and mutation identifiers are UUIDv7 values.
|
||||
Transport is at-least-once, but the client durably
|
||||
remembers accepted UUIDs and provides an at-most-once execution guarantee: it
|
||||
never authorizes the same request to execute twice. A crash in the launch window
|
||||
may interrupt a command before its requested code runs, but must never cause an
|
||||
automatic retry with an uncertain prior outcome.
|
||||
|
||||
After full terminal history is removed, each client retains a compact FIFO of
|
||||
the most recent 1,000,000 command tombstones containing binary UUIDv7,
|
||||
immutable request hash, and completion/acknowledgement time. The server retains
|
||||
the most recent 1,000,000 command tombstones globally as a first-line duplicate
|
||||
check. A matching UUID/hash returns structured `ALREADY_EXECUTED`; a matching
|
||||
UUID with different immutable content is a conflict. These ledgers have separate
|
||||
count-based budgets and do not retain command payload or output. Replay
|
||||
protection older than the retained client tombstone horizon is best-effort.
|
||||
|
||||
The server states are:
|
||||
|
||||
```text
|
||||
queued -> dispatched -> accepted -> running -> succeeded | failed | terminated
|
||||
\-------------------------------> cancelled
|
||||
queued -> dispatched | cancelled | expired
|
||||
dispatched -> queued | accepted | rejected | cancelled
|
||||
accepted -> running | rejected | cancelled
|
||||
running -> succeeded | failed | terminated | interrupted
|
||||
```
|
||||
|
||||
`accepted` means the client has durably accepted the request; `running` means
|
||||
the process has been launched. `cancelled` is used when it is stopped before
|
||||
launch. A kill racing launch is resolved by command revision: the client either
|
||||
acknowledges cancellation before launch or launches then immediately applies
|
||||
the requested signal, recording the race.
|
||||
Recovery/corruption handling may also move an affected `dispatched` or
|
||||
`accepted` command to `interrupted`; those are exceptional reconciliation
|
||||
transitions, not normal execution outcomes.
|
||||
|
||||
`accepted` means the client has durably admitted the request; `running` means
|
||||
the requested process has been launched. `cancelled` is used when it is stopped
|
||||
before launch. A permanent pre-launch validation, script-transfer, or process-
|
||||
preparation failure is terminal `rejected`; `failed` is reserved for code that
|
||||
actually launched. A transient
|
||||
capacity rejection returns to `queued` with backoff and remains subject to its
|
||||
queue TTL. The server may cancel work that has never been dispatched immediately.
|
||||
For dispatched or accepted work, it persists a higher command revision and
|
||||
sends the revisioned signal without prematurely declaring a terminal state. The
|
||||
client either returns a revisioned `cancelled` lifecycle before launch
|
||||
authorization or applies the signal after authorization and returns a
|
||||
revisioned signal result. Cancellation intent remains internal rather than
|
||||
adding a public lifecycle state.
|
||||
|
||||
Queued work has a configurable acceptance deadline, default 15 minutes; zero
|
||||
explicitly means no expiry. A command that was never dispatched becomes
|
||||
terminal `expired` at its deadline. A dispatched command whose acceptance is
|
||||
uncertain remains non-terminal and is displayed as expired pending
|
||||
reconciliation. Later client evidence updates the actual lifecycle. Acceptance
|
||||
or execution observed after the deadline creates an incident and is displayed
|
||||
as late-after-expiry; the server requests termination but continues recording
|
||||
the actual client-reported outcome.
|
||||
|
||||
Process launch has internal durable phases `launch_prepared` and
|
||||
`launch_authorized` between public `accepted` and `running`. The client creates
|
||||
the process behind an OS-specific execution barrier, durably records its process
|
||||
identity, durably authorizes launch, and only then releases requested command
|
||||
code. Once authorization is durable, an uncertain outcome is reconciled as
|
||||
`interrupted`, never by redispatching that UUID.
|
||||
|
||||
The client permits 16 concurrent processes and 100 pending commands by default.
|
||||
Those values are configurable and advertised to the server. A full queue causes
|
||||
a structured capacity error rather than creating unbounded work.
|
||||
The server separately defaults to at most 1,000 queued commands for one target
|
||||
client and 10,000 queued commands globally, subject to the stricter byte quotas.
|
||||
|
||||
### Execution contract
|
||||
|
||||
- Shell selection is explicit: Unix-like clients support `sh` and `bash`;
|
||||
Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`.
|
||||
Unsupported shells are rejected; no fallback occurs.
|
||||
- The client materializes `command_text` as a private generated wrapper and
|
||||
executes that file with exactly the selected shell: `sh`/`bash`, `cmd`, or
|
||||
PowerShell. This avoids transport quoting and Windows command-line limits;
|
||||
it never performs shell detection or fallback. Wrapper cleanup follows the
|
||||
same terminal rule as uploaded scripts.
|
||||
- A command receives the daemon account's permissions and startup environment,
|
||||
overlaid with the persisted `env_overrides` map. The effective CWD is the
|
||||
requested existing directory or the registered daemon CWD when omitted.
|
||||
- Each command is isolated into a process tree: a Unix session/process group or
|
||||
a Windows Job Object. A daemon that cannot supervise its children terminates
|
||||
them and reports interruption rather than claiming recovery it cannot make.
|
||||
- Terminal lifecycle means the supervised tree is empty, not only that its root
|
||||
shell exited. After root exit, RVBox allows a configurable 5-second descendant
|
||||
and output-drain grace period, then terminates residual descendants. It drains
|
||||
capture pipes and durably sequences all retained output (or an explicit
|
||||
`OutputIncomplete` marker) before emitting the terminal lifecycle event.
|
||||
- Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`,
|
||||
`MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in;
|
||||
the concrete administrator-configured limits are applied with cgroup v2 when
|
||||
@@ -115,18 +182,20 @@ a structured capacity error rather than creating unbounded work.
|
||||
Scripts are content-addressed uploads, not shell-escaped command strings. The
|
||||
server sends a descriptor containing SHA-256 and then ordered chunks (default
|
||||
maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary
|
||||
file beneath the effective CWD, executes it with the selected shell, and removes
|
||||
it after the command reaches a terminal state. Script content is not copied into
|
||||
the audit log; its digest and metadata are.
|
||||
file beneath the effective CWD, executes it with exactly the selected shell, and
|
||||
removes it after the command reaches a terminal state. The descriptor filename
|
||||
is display metadata only and is never used as a path component. Script content
|
||||
is not copied into the audit log; its digest and metadata are.
|
||||
|
||||
## Process control and diagnostics
|
||||
|
||||
On Unix, `kill` addresses the command's process group and accepts normal signal
|
||||
names/numbers supported by that client. On Windows, only `SIGTERM` and `SIGKILL`
|
||||
are valid. `SIGTERM` makes a best-effort `CTRL_BREAK_EVENT` delivery to the
|
||||
dedicated console group, waits 10 seconds, then terminates the Job Object if
|
||||
needed. `SIGKILL` immediately terminates the Job Object. The response reports
|
||||
the actual escalation outcome.
|
||||
On Unix, `kill` addresses the command's process group and accepts the portable
|
||||
v1 set `HUP`, `INT`, `TERM`, `KILL`, `USR1`, and `USR2` (including their
|
||||
`SIG`-prefixed CLI spellings). Arbitrary native signal numbers are not part of
|
||||
v1. On Windows, only `TERM` and `KILL` are valid. `TERM` makes a best-effort
|
||||
`CTRL_BREAK_EVENT` delivery to the dedicated console group, waits 10 seconds,
|
||||
then terminates the Job Object if needed. `KILL` immediately terminates the Job
|
||||
Object. The response reports the actual escalation outcome.
|
||||
|
||||
Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is
|
||||
emitted after the configurable default of 10 minutes without observable
|
||||
@@ -142,31 +211,73 @@ restart and reconciles only these; server-confirmed historical terminal commands
|
||||
are not re-reconciled. The client stores only active command state and output
|
||||
not acknowledged by the server.
|
||||
|
||||
Every execution-originated event has a strictly increasing `event_seq` scoped to
|
||||
one command. This includes lifecycle transitions, stdout/stderr chunks, stdin
|
||||
acknowledgments, resource snapshots, signals, and terminal events. The server
|
||||
preserves the sequence and additionally records receipt time. This is the
|
||||
canonical reconstruction order across interleaved streams and retries.
|
||||
Truncation is separate range metadata so it can truthfully describe missing
|
||||
event sequences without consuming one itself.
|
||||
Every transmitted execution event has a strictly increasing `event_seq` scoped
|
||||
to one command. The client first stores lifecycle transitions, stdout/stderr,
|
||||
stdin acknowledgements, resource snapshots, signals, and terminal state in a
|
||||
durable local order. It durably assigns wire sequences only as entries enter the
|
||||
bounded send window, after any unsent-output compaction. Once assigned, an event
|
||||
is pinned until acknowledged and retry content is immutable. A client-side
|
||||
truncation marker therefore consumes a normal sequence without creating a wire
|
||||
gap. Server-side retention exposes removed event ranges as query metadata
|
||||
without allocating client sequences. The server also records receipt time;
|
||||
`event_seq` remains the canonical transmitted order across streams and retries.
|
||||
|
||||
Output chunks are Zstandard-compressed before persistent quota accounting.
|
||||
Per-command history is a rolling compressed window (10 MiB default), so the
|
||||
oldest output segments for that command are removed first and a sequence-range
|
||||
truncation marker remains. This applies to active and terminal commands. When
|
||||
connected, the client first removes server-acknowledged segments. While offline,
|
||||
it must still honor both hard caps: it retains the newest tail, removes oldest
|
||||
unacknowledged compressed chunks when necessary, and records their exact missing
|
||||
ranges for durable reporting on reconnect.
|
||||
The terminal lifecycle event is always the final client event for a command.
|
||||
The control-plane follow wrapper emits any server-created retention metadata
|
||||
before returning that terminal event. Followers may therefore stop on terminal
|
||||
without missing subsequently sequenced stdout/stderr or known truncation data.
|
||||
|
||||
Each client also has a 50 MiB aggregate compressed spool cap for active,
|
||||
unacknowledged work. The server's matching per-client compressed-history cap is
|
||||
50 MiB; it evicts that client's oldest terminal command records as needed. The
|
||||
server-wide cap is 1 GiB; it evicts whole oldest terminal command records
|
||||
(metadata and output), never arbitrary stdout/stderr rows. Active commands are
|
||||
protected. If active commands alone consume a per-client budget, their oldest
|
||||
acknowledged output rotates by the per-command rule; pipes continue draining so
|
||||
a child cannot deadlock on output.
|
||||
All command-owned stored data is quota-accounted: execution metadata, script
|
||||
body, pending stdin, events, and output. Stored message/blob payloads are
|
||||
Zstandard-compressed; necessary SQLite index/state columns are charged by their
|
||||
encoded lengths plus a conservative versioned per-row/index overhead rather
|
||||
than pretending they are free. This logical accounting is deterministic across
|
||||
SQLite compaction. A separate filesystem free-space floor protects WAL,
|
||||
temporary files, tombstones, and accounting variance.
|
||||
The default total is 32 MiB per command, 256 MiB per client on both client and
|
||||
server, and 4 GiB server-wide. A separate 10 MiB rolling output window remains
|
||||
per command, and raw script input remains limited to 10 MiB. Client accounting
|
||||
also charges the raw generated execution wrapper/script file while it exists,
|
||||
even though the compressed durable source was already charged.
|
||||
|
||||
Each accepted active command reserves 64 KiB of its quota for bounded closeout
|
||||
metadata. Essential state for an accepted stdin/signal mutation is additionally
|
||||
reserved before that mutation succeeds. Essential active state is never silently
|
||||
rolled. Server output is evictable: the oldest retained chunks are removed first
|
||||
and sequence-range query metadata remains. The client removes acknowledged data,
|
||||
pins its bounded assigned send window, and may replace older unsequenced output
|
||||
with normally sequenced byte-loss markers. While offline or under sustained
|
||||
overload, it retains the newest tail and records exact known lost bytes. At
|
||||
client/server aggregate limits, terminal commands are evicted as whole UUID
|
||||
records in oldest server issue-time/UUIDv7 order first. If active data
|
||||
alone reaches a limit, output rotates or enters loss mode and new essential
|
||||
allocations are rejected with `CAPACITY_EXHAUSTED`; pipes continue draining.
|
||||
|
||||
Transient raw output also has bounded high/low watermarks before compression:
|
||||
1 MiB/256 KiB per command, 8 MiB/4 MiB per client daemon, and 64 MiB/32 MiB
|
||||
server-wide by default. Crossing a client high watermark enters loss mode;
|
||||
still-unsequenced output bytes may be discarded before compression and replaced
|
||||
in durable local order by a later `OutputTruncation`. Loss mode ends only below
|
||||
the corresponding low watermark. The server never drops an already sequenced
|
||||
client event: at its ingress high watermark it withholds acknowledgement and
|
||||
closes an overproducing session if bounded admission cannot continue, letting
|
||||
the durable client retry after the server backlog falls below its low watermark.
|
||||
Lifecycle,
|
||||
stdin acknowledgements, signal results, and truncation/incomplete markers use
|
||||
reserved capacity and are never treated as droppable output.
|
||||
|
||||
Independently of byte pressure, the server reclaims each whole terminal command
|
||||
and all command-owned data 30 days after its terminal time by default. Zero
|
||||
explicitly disables age rotation. Compact replay tombstones, audit records, and
|
||||
storage incidents remain under their separate retention policies.
|
||||
|
||||
Storage health is derived from durable incident records. Safe repairs resolve
|
||||
an incident automatically; known data loss remains dirty until an operator
|
||||
explicitly acknowledges it. Resolution clears the dirty health flag but does
|
||||
not erase records as part of that action. Unresolved compact records are
|
||||
non-evictable; resolved summaries and audit entries follow the separate 100 MiB
|
||||
audit/incident history rotation. Recovery and incident management are described
|
||||
in the platform contract and exposed through the control plane.
|
||||
|
||||
## Audit and timestamps
|
||||
|
||||
@@ -175,8 +286,13 @@ events: source/transport identity where available, target client, UUID, action,
|
||||
request time, result, and error. It records command text, environment-override
|
||||
names (not values), and script metadata/digest, but not duplicated
|
||||
stdin/stdout/stderr payloads. The persisted execution request necessarily keeps
|
||||
override values for dispatch/retry and must be access-controlled as sensitive
|
||||
data. Audit retention is configured independently of output retention.
|
||||
override values for dispatch/retry. In v1, command text, scripts, environment
|
||||
values, stdin, output, and other command-owned payloads are stored in plaintext;
|
||||
application-level encryption and key management are deferred to a future
|
||||
version. Audit storage uses compressed segments under a separate 100 MiB default
|
||||
quota and rotates complete oldest segments. Audit retention is configured
|
||||
independently of command retention; its default zero age limit means quota-only
|
||||
rotation.
|
||||
|
||||
All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client
|
||||
observed timestamps and server receipt timestamps are distinct; the latter is
|
||||
|
||||
Reference in New Issue
Block a user