Files
rvbox/docs/architecture.md
T

334 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RVBox v1 architecture
## Purpose and scope
RVBox is a reverse-connection remote command system. A client daemon (`rvbox`)
maintains a WebSocket connection to `rvbox-server`; the server persists command
state and exposes a local control plane to `rvc` and, optionally, JSON-RPC
callers. This design is the v1 contract for Go implementations and a later Rust
client implementation.
The first supported Go client is the Windows service client. The Linux client
is implemented afterward against the already proven shared runtime/protocol
and remains required for the complete v1 scope. It is explicitly deferred from
the current implementation effort: implement the Linux server and native
Windows client first, and do not begin Unix-like client implementation yet. The
server remains Linux-first.
RVBox deliberately executes arbitrary commands under locally selected
execution identities. It is therefore an administrative tool, not a multi-
tenant remote-execution service. On Windows the main service runs as
`LocalSystem` and maps the request's elevation intent plus current login state
to an effective user/session context. A peer that can successfully impersonate
the client's routing identity can request elevation, whose fallback may reach
SYSTEM. This makes the accepted self-reported-identity/network trust boundary
especially consequential; the explicit elevation bit and tray visibility are
intent/audit signals, not authorization controls.
Both daemons use strict TOML 1.0 configuration. The normative schema and fully
annotated examples are in the [configuration contract](configuration.md).
## Non-goals
- Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist
are not part of v1. TLS is terminated by the deployment's nginx instance.
- RVBox does not promise a definitive cross-platform “hung” determination.
- V1 does not provide arbitrary file transfer; `--script` transfers only the
temporary script required to execute that request.
- Resource limits are supported only when explicitly requested by an execution
profile; they are not imposed by default.
## Trust boundary
The server accepts a self-reported hostname as `client_id`; it is an opaque
1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer
registration from the same durable client instance replaces its prior live
session. A different instance is accepted normally when no session for that
client ID is live. While one is live, a different instance is rejected unless
an operator grants a one-shot override for that exact pending instance. This
prevents accidental hostname collisions but is not authentication: a peer able
to reach nginx and copy or guess the identifiers can still impersonate a
client. This is an accepted v1 limitation and deployments must restrict the
endpoint to a trusted network.
The Unix control socket is local-only and mode `0600`, owned by the server
account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated;
it defaults to loopback but can be bound elsewhere by configuration. Exposing
it to a network exposes full remote-command authority and is unsafe without an
external access-control layer.
## Components
```text
rvc -- gRPC/Unix socket -- rvbox-server -- WSS/nginx -- rvbox service -- shell
| | |
SQLite/WAL | local named pipe
| | |
compressed data +-- optional JSON-RPC +-- per-session tray
```
`rvc` is a thin control-plane client. The server owns durable command history,
dispatch, queueing, pagination, audit records, and session fencing. A client
owns active process supervision, unacknowledged output spooling, and safe
reconnection. Neither daemon lets a slow peer, command, or output stream block
its dispatch loops.
On Windows, SCM starts one machine-wide `LocalSystem` service before login. It
alone owns networking, durable state, logs, command admission, process handles,
and Job Objects. A separate unelevated tray may run in each logged-in session
and communicates only through an ACL-protected local named pipe; it is an
optional frontend and its exit or absence does not stop the service. Task
Scheduler is not used.
## Identity, sessions, and lifecycle
1. The client connects over WSS and sends `ClientHello` with its client ID,
protocol capability, OS/architecture, daemon version, configured daemon CWD/
Windows work root,
supported shells, and a durable random client-instance UUID.
2. The server accepts the current compatible protocol version, fences the
previous connection for that ID and same client instance, and returns a
fresh server-issued `session_id` and monotonic `session_generation`. A live
claim from a different client instance is rejected while the current session
is live unless an operator has explicitly authorized that pending instance.
With no live session, the new instance is accepted normally.
3. Every client-to-server envelope and server dispatch is bound to that token.
The server discards traffic from superseded sessions, including late output.
4. The server queues work while a client is offline and dispatches it only when
the active session advertises capacity. A replacement session immediately
performs bidirectional reconciliation between the server's non-terminal set
and the client's complete retained-command set. The server returns explicit
local terminate/discard decisions before new dispatch begins.
## Command model
Each user request has a UUID (`issue_uuid`) and a durable request record: target
client, request/issue timestamps, shell type, command text or script descriptor,
CWD, environment overrides, Windows elevation intent and effective execution
identity, resource-profile flags, and lifecycle state. `rvc`
normally supplies this UUID as its optional `request_id`; the server generates
one when it is omitted. Command and mutation identifiers are UUIDv7 values.
Transport is at-least-once, but the client durably
remembers accepted UUIDs and provides an at-most-once execution guarantee: it
never authorizes the same request to execute twice. A crash in the launch window
may interrupt a command before its requested code runs, but must never cause an
automatic retry with an uncertain prior outcome.
After full terminal history is removed, each client retains a compact FIFO of
the most recent 1,000,000 command tombstones containing binary UUIDv7,
immutable request hash, and completion/acknowledgement time. The server retains
the most recent 1,000,000 command tombstones globally as a first-line duplicate
check. A matching UUID/hash returns structured `ALREADY_EXECUTED`; a matching
UUID with different immutable content is a conflict. These ledgers have separate
count-based budgets and do not retain command payload or output. Replay
protection older than the retained client tombstone horizon is best-effort.
The server states are:
```text
queued -> dispatched | cancelled | expired
dispatched -> queued | accepted | rejected | cancelled
accepted -> running | rejected | cancelled
running -> succeeded | failed | terminated | interrupted
```
Recovery/corruption handling may also move an affected `dispatched` or
`accepted` command to `interrupted`; those are exceptional reconciliation
transitions, not normal execution outcomes.
`accepted` means the client has durably admitted the request; `running` means
the requested process has been launched. `cancelled` is used when it is stopped
before launch. A permanent pre-launch validation, script-transfer, or process-
preparation failure is terminal `rejected`; `failed` is reserved for code that
actually launched. A transient
capacity rejection returns to `queued` with backoff and remains subject to its
queue TTL. The server may cancel work that has never been dispatched immediately.
For dispatched or accepted work, it persists a higher command revision and
sends the revisioned signal without prematurely declaring a terminal state. The
client either returns a revisioned `cancelled` lifecycle before launch
authorization or applies the signal after authorization and returns a
revisioned signal result. Cancellation intent remains internal rather than
adding a public lifecycle state.
Queued work has a configurable acceptance deadline, default 15 minutes; zero
explicitly means no expiry. A command that was never dispatched becomes
terminal `expired` at its deadline. A dispatched command whose acceptance is
uncertain remains non-terminal and is displayed as expired pending
reconciliation. Later client evidence updates the actual lifecycle. Acceptance
or execution observed after the deadline creates an incident and is displayed
as late-after-expiry; the server requests termination but continues recording
the actual client-reported outcome.
Process launch has internal durable phases `launch_prepared` and
`launch_authorized` between public `accepted` and `running`. The client creates
the process behind an OS-specific execution barrier, durably records its process
identity, durably authorizes launch, and only then releases requested command
code. Once authorization is durable, an uncertain outcome is reconciled as
`interrupted`, never by redispatching that UUID.
The client permits 16 concurrent processes and 100 pending commands by default.
Those values are configurable and advertised to the server. A full queue causes
a structured capacity error rather than creating unbounded work.
The server separately defaults to at most 1,000 queued commands for one target
client and 10,000 queued commands globally, subject to the stricter byte quotas.
### Execution contract
- Shell selection is explicit: Unix-like clients support `sh` and `bash`;
Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`.
Unsupported shells are rejected; no fallback occurs.
- The client materializes `command_text` as a private generated wrapper and
executes that file with exactly the selected shell: `sh`/`bash`, `cmd`, or
PowerShell. This avoids transport quoting and Windows command-line limits;
it never performs shell detection or fallback. Wrapper cleanup follows the
same terminal rule as uploaded scripts.
- Unix commands receive the daemon account's permissions and startup
environment. Windows interprets the request's `elevated` boolean through a
login-aware hierarchy. With a usable active session, normal work uses a
deliberately non-elevated `active_user` token; elevated work tries
`active_user_elevated`, then `active_system`, then `local_system`. With no
usable active session, normal work uses `local_service` and elevated work uses
`local_system`. Fallback is allowed only during token selection before
`launch_prepared`, never after a process may have started. Requested elevation,
attempted contexts, effective token SID, session ID/session-owner SID, and
selection detail are persisted in command history.
- A command receives the base environment for its effective identity, overlaid
with the persisted `env_overrides` map. The effective CWD is the requested
existing accessible directory when supplied. When omitted it is the registered
daemon CWD on Unix or an ACL-isolated identity child beneath that registered
root on Windows.
- Each command is isolated into a process tree: a Unix session/process group or
a Windows Job Object. A daemon that cannot supervise its children terminates
them and reports interruption rather than claiming recovery it cannot make.
- Terminal lifecycle means the supervised tree is empty, not only that its root
shell exited. After root exit, RVBox allows a configurable 5-second descendant
and output-drain grace period, then terminates residual descendants. It drains
capture pipes and durably sequences all retained output (or an explicit
`OutputIncomplete` marker) before emitting the terminal lifecycle event.
- Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`,
`MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in;
the concrete administrator-configured limits are applied with cgroup v2 when
available on Linux and Job Object limits on Windows. Unsupported requested
controls are reported, never silently ignored.
### Script execution
Scripts are content-addressed uploads, not shell-escaped command strings. The
server sends a descriptor containing SHA-256 and then ordered chunks (default
maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary
file beneath the effective CWD, executes it with exactly the selected shell, and
removes it after the command reaches a terminal state. The descriptor filename
is display metadata only and is never used as a path component. Script content
is not copied into the audit log; its digest and metadata are.
## Process control and diagnostics
On Unix, `kill` addresses the command's process group and accepts the portable
v1 set `HUP`, `INT`, `TERM`, `KILL`, `USR1`, and `USR2` (including their
`SIG`-prefixed CLI spellings). Arbitrary native signal numbers are not part of
v1. On Windows, only `TERM` and `KILL` are valid. `TERM` makes a best-effort
`CTRL_BREAK_EVENT` delivery to the dedicated console group, waits 10 seconds,
then terminates the Job Object if needed. `KILL` immediately terminates the Job
Object. The response reports the actual escalation outcome.
Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is
emitted after the configurable default of 10 minutes without observable
progress. Linux enriches this with `/proc` state, CPU, memory, and I/O data;
Windows uses process and Job Object APIs where available. It is explicitly a
heuristic, not proof of an I/O stall.
## Durability, ordering, and retention
The server uses SQLite in WAL mode for metadata/indexes and append-only Zstandard
compressed segment files for output. It recovers non-terminal commands after a
restart and reconciles only these; server-confirmed historical terminal commands
are not re-reconciled. The client stores only active command state and output
not acknowledged by the server.
Every transmitted execution event has a strictly increasing `event_seq` scoped
to one command. The client first stores lifecycle transitions, stdout/stderr,
stdin acknowledgements, resource snapshots, signals, and terminal state in a
durable local order. It durably assigns wire sequences only as entries enter the
bounded send window, after any unsent-output compaction. Once assigned, an event
is pinned until acknowledged and retry content is immutable. A client-side
truncation marker therefore consumes a normal sequence without creating a wire
gap. Server-side retention exposes removed event ranges as query metadata
without allocating client sequences. The server also records receipt time;
`event_seq` remains the canonical transmitted order across streams and retries.
The terminal lifecycle event is always the final client event for a command.
The control-plane follow wrapper emits any server-created retention metadata
before returning that terminal event. Followers may therefore stop on terminal
without missing subsequently sequenced stdout/stderr or known truncation data.
All command-owned stored data is quota-accounted: execution metadata, script
body, pending stdin, events, and output. Stored message/blob payloads are
Zstandard-compressed; necessary SQLite index/state columns are charged by their
encoded lengths plus a conservative versioned per-row/index overhead rather
than pretending they are free. This logical accounting is deterministic across
SQLite compaction. A separate filesystem free-space floor protects WAL,
temporary files, tombstones, and accounting variance.
The default total is 32 MiB per command, 256 MiB per client on both client and
server, and 4 GiB server-wide. A separate 10 MiB rolling output window remains
per command, and raw script input remains limited to 10 MiB. Client accounting
also charges the raw generated execution wrapper/script file while it exists,
even though the compressed durable source was already charged.
Each accepted active command reserves 64 KiB of its quota for bounded closeout
metadata. Essential state for an accepted stdin/signal mutation is additionally
reserved before that mutation succeeds. Essential active state is never silently
rolled. Server output is evictable: the oldest retained chunks are removed first
and sequence-range query metadata remains. The client removes acknowledged data,
pins its bounded assigned send window, and may replace older unsequenced output
with normally sequenced byte-loss markers. While offline or under sustained
overload, it retains the newest tail and records exact known lost bytes. At
client/server aggregate limits, terminal commands are evicted as whole UUID
records in oldest server issue-time/UUIDv7 order first. If active data
alone reaches a limit, output rotates or enters loss mode and new essential
allocations are rejected with `CAPACITY_EXHAUSTED`; pipes continue draining.
Transient raw output also has bounded high/low watermarks before compression:
1 MiB/256 KiB per command, 8 MiB/4 MiB per client daemon, and 64 MiB/32 MiB
server-wide by default. Crossing a client high watermark enters loss mode;
still-unsequenced output bytes may be discarded before compression and replaced
in durable local order by a later `OutputTruncation`. Loss mode ends only below
the corresponding low watermark. The server never drops an already sequenced
client event: at its ingress high watermark it withholds acknowledgement and
closes an overproducing session if bounded admission cannot continue, letting
the durable client retry after the server backlog falls below its low watermark.
Lifecycle,
stdin acknowledgements, signal results, and truncation/incomplete markers use
reserved capacity and are never treated as droppable output.
Independently of byte pressure, the server reclaims each whole terminal command
and all command-owned data 30 days after its terminal time by default. Zero
explicitly disables age rotation. Compact replay tombstones, audit records, and
storage incidents remain under their separate retention policies.
Storage health is derived from durable incident records. Safe repairs resolve
an incident automatically; known data loss remains dirty until an operator
explicitly acknowledges it. Resolution clears the dirty health flag but does
not erase records as part of that action. Unresolved compact records are
non-evictable; resolved summaries and audit entries follow the separate 100 MiB
audit/incident history rotation. Recovery and incident management are described
in the platform contract and exposed through the control plane.
## Audit and timestamps
The server writes a separate durable audit trail for control actions and session
events: source/transport identity where available, target client, UUID, action,
request time, result, and error. It records command text, environment-override
names (not values), and script metadata/digest, but not duplicated
stdin/stdout/stderr payloads. The persisted execution request necessarily keeps
override values for dispatch/retry. In v1, command text, scripts, environment
values, stdin, output, and other command-owned payloads are stored in plaintext;
application-level encryption and key management are deferred to a future
version. Audit storage uses compressed segments under a separate 100 MiB default
quota and rotates complete oldest segments. Audit retention is configured
independently of command retention; its default zero age limit means quota-only
rotation.
All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client
observed timestamps and server receipt timestamps are distinct; the latter is
authoritative for server records, while `event_seq` is authoritative for order.