# RVBox v1 architecture ## Purpose and scope RVBox is a reverse-connection remote command system. A client daemon (`rvbox`) maintains a WebSocket connection to `rvbox-server`; the server persists command state and exposes a local control plane to `rvc` and, optionally, JSON-RPC callers. This design is the v1 contract for Go implementations and a later Rust client implementation. RVBox deliberately executes arbitrary commands with the identity, permissions, and base environment of the client daemon. It is therefore an administrative tool, not a multi-tenant remote-execution service. ## Non-goals - Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist are not part of v1. TLS is terminated by the deployment's nginx instance. - RVBox does not promise a definitive cross-platform “hung” determination. - V1 does not provide arbitrary file transfer; `--script` transfers only the temporary script required to execute that request. - Resource limits are supported only when explicitly requested by an execution profile; they are not imposed by default. ## Trust boundary The server accepts a self-reported hostname as `client_id`; it is an opaque 1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer registration for an ID replaces its prior live session. Consequently, a peer able to reach nginx can impersonate or take over a client ID. This is an accepted v1 limitation and deployments must restrict the endpoint to a trusted network. The Unix control socket is local-only and mode `0600`, owned by the server account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated; it defaults to loopback but can be bound elsewhere by configuration. Exposing it to a network exposes full remote-command authority and is unsafe without an external access-control layer. ## Components ```text rvc -- gRPC/Unix socket -- rvbox-server -- WSS/nginx -- rvbox client -- shell | | SQLite/WAL +-- optional HTTP JSON-RPC (debug/batch) | compressed output segment files + audit log ``` `rvc` is a thin control-plane client. The server owns durable command history, dispatch, queueing, pagination, audit records, and session fencing. A client owns active process supervision, unacknowledged output spooling, and safe reconnection. Neither daemon lets a slow peer, command, or output stream block its dispatch loops. ## Identity, sessions, and lifecycle 1. The client connects over WSS and sends `ClientHello` with its client ID, protocol capability, OS/architecture, daemon version, current daemon CWD, supported shells, and a fresh reconnect UUID. 2. The server accepts the current compatible protocol version, fences the previous connection for that ID, and returns a fresh server-issued `session_id` and monotonic `session_generation`. 3. Every client-to-server envelope and server dispatch is bound to that token. The server discards traffic from superseded sessions, including late output. 4. The server queues work while a client is offline and dispatches it only when the active session advertises capacity. A replacement session immediately resumes non-terminal reconciliation. ## Command model Each user request has a server-generated UUID (`issue_uuid`) and a durable request record: target client, request/issue timestamps, shell type, command text or script descriptor, CWD, environment overrides, resource-profile flags, and lifecycle state. The UUID is the end-to-end idempotency key. Transport is at-least-once, but the client durably remembers accepted UUIDs and never starts the same request twice. The server states are: ```text queued -> dispatched -> accepted -> running -> succeeded | failed | terminated \-------------------------------> cancelled ``` `accepted` means the client has durably accepted the request; `running` means the process has been launched. `cancelled` is used when it is stopped before launch. A kill racing launch is resolved by command revision: the client either acknowledges cancellation before launch or launches then immediately applies the requested signal, recording the race. The client permits 16 concurrent processes and 100 pending commands by default. Those values are configurable and advertised to the server. A full queue causes a structured capacity error rather than creating unbounded work. ### Execution contract - Shell selection is explicit: Unix-like clients support `sh` and `bash`; Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`. Unsupported shells are rejected; no fallback occurs. - A command receives the daemon account's permissions and startup environment, overlaid with the persisted `env_overrides` map. The effective CWD is the requested existing directory or the registered daemon CWD when omitted. - Each command is isolated into a process tree: a Unix session/process group or a Windows Job Object. A daemon that cannot supervise its children terminates them and reports interruption rather than claiming recovery it cannot make. - Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`, `MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in; the concrete administrator-configured limits are applied with cgroup v2 when available on Linux and Job Object limits on Windows. Unsupported requested controls are reported, never silently ignored. ### Script execution Scripts are content-addressed uploads, not shell-escaped command strings. The server sends a descriptor containing SHA-256 and then ordered chunks (default maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary file beneath the effective CWD, executes it with the selected shell, and removes it after the command reaches a terminal state. Script content is not copied into the audit log; its digest and metadata are. ## Process control and diagnostics On Unix, `kill` addresses the command's process group and accepts normal signal names/numbers supported by that client. On Windows, only `SIGTERM` and `SIGKILL` are valid. `SIGTERM` makes a best-effort `CTRL_BREAK_EVENT` delivery to the dedicated console group, waits 10 seconds, then terminates the Job Object if needed. `SIGKILL` immediately terminates the Job Object. The response reports the actual escalation outcome. Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is emitted after the configurable default of 10 minutes without observable progress. Linux enriches this with `/proc` state, CPU, memory, and I/O data; Windows uses process and Job Object APIs where available. It is explicitly a heuristic, not proof of an I/O stall. ## Durability, ordering, and retention The server uses SQLite in WAL mode for metadata/indexes and append-only Zstandard compressed segment files for output. It recovers non-terminal commands after a restart and reconciles only these; server-confirmed historical terminal commands are not re-reconciled. The client stores only active command state and output not acknowledged by the server. Every execution-originated event has a strictly increasing `event_seq` scoped to one command. This includes lifecycle transitions, stdout/stderr chunks, stdin acknowledgments, resource snapshots, signals, and terminal events. The server preserves the sequence and additionally records receipt time. This is the canonical reconstruction order across interleaved streams and retries. Truncation is separate range metadata so it can truthfully describe missing event sequences without consuming one itself. Output chunks are Zstandard-compressed before persistent quota accounting. Per-command history is a rolling compressed window (10 MiB default), so the oldest output segments for that command are removed first and a sequence-range truncation marker remains. This applies to active and terminal commands. When connected, the client first removes server-acknowledged segments. While offline, it must still honor both hard caps: it retains the newest tail, removes oldest unacknowledged compressed chunks when necessary, and records their exact missing ranges for durable reporting on reconnect. Each client also has a 50 MiB aggregate compressed spool cap for active, unacknowledged work. The server's matching per-client compressed-history cap is 50 MiB; it evicts that client's oldest terminal command records as needed. The server-wide cap is 1 GiB; it evicts whole oldest terminal command records (metadata and output), never arbitrary stdout/stderr rows. Active commands are protected. If active commands alone consume a per-client budget, their oldest acknowledged output rotates by the per-command rule; pipes continue draining so a child cannot deadlock on output. ## Audit and timestamps The server writes a separate durable audit trail for control actions and session events: source/transport identity where available, target client, UUID, action, request time, result, and error. It records command text, environment-override names (not values), and script metadata/digest, but not duplicated stdin/stdout/stderr payloads. The persisted execution request necessarily keeps override values for dispatch/retry and must be access-controlled as sensitive data. Audit retention is configured independently of output retention. All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client observed timestamps and server receipt timestamps are distinct; the latter is authoritative for server records, while `event_seq` is authoritative for order.