184 lines
9.4 KiB
Markdown
184 lines
9.4 KiB
Markdown
# RVBox v1 architecture
|
||
|
||
## Purpose and scope
|
||
|
||
RVBox is a reverse-connection remote command system. A client daemon (`rvbox`)
|
||
maintains a WebSocket connection to `rvbox-server`; the server persists command
|
||
state and exposes a local control plane to `rvc` and, optionally, JSON-RPC
|
||
callers. This design is the v1 contract for Go implementations and a later Rust
|
||
client implementation.
|
||
|
||
RVBox deliberately executes arbitrary commands with the identity, permissions,
|
||
and base environment of the client daemon. It is therefore an administrative
|
||
tool, not a multi-tenant remote-execution service.
|
||
|
||
## Non-goals
|
||
|
||
- Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist
|
||
are not part of v1. TLS is terminated by the deployment's nginx instance.
|
||
- RVBox does not promise a definitive cross-platform “hung” determination.
|
||
- V1 does not provide arbitrary file transfer; `--script` transfers only the
|
||
temporary script required to execute that request.
|
||
- Resource limits are supported only when explicitly requested by an execution
|
||
profile; they are not imposed by default.
|
||
|
||
## Trust boundary
|
||
|
||
The server accepts a self-reported hostname as `client_id`; it is an opaque
|
||
1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer
|
||
registration for an ID replaces its prior live session. Consequently, a peer
|
||
able to reach nginx can impersonate or take over a client ID. This is an
|
||
accepted v1 limitation and deployments must restrict the endpoint to a trusted
|
||
network.
|
||
|
||
The Unix control socket is local-only and mode `0600`, owned by the server
|
||
account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated;
|
||
it defaults to loopback but can be bound elsewhere by configuration. Exposing
|
||
it to a network exposes full remote-command authority and is unsafe without an
|
||
external access-control layer.
|
||
|
||
## Components
|
||
|
||
```text
|
||
rvc -- gRPC/Unix socket -- rvbox-server -- WSS/nginx -- rvbox client -- shell
|
||
| |
|
||
SQLite/WAL +-- optional HTTP JSON-RPC (debug/batch)
|
||
|
|
||
compressed output segment files + audit log
|
||
```
|
||
|
||
`rvc` is a thin control-plane client. The server owns durable command history,
|
||
dispatch, queueing, pagination, audit records, and session fencing. A client
|
||
owns active process supervision, unacknowledged output spooling, and safe
|
||
reconnection. Neither daemon lets a slow peer, command, or output stream block
|
||
its dispatch loops.
|
||
|
||
## Identity, sessions, and lifecycle
|
||
|
||
1. The client connects over WSS and sends `ClientHello` with its client ID,
|
||
protocol capability, OS/architecture, daemon version, current daemon CWD,
|
||
supported shells, and a fresh reconnect UUID.
|
||
2. The server accepts the current compatible protocol version, fences the
|
||
previous connection for that ID, and returns a fresh server-issued
|
||
`session_id` and monotonic `session_generation`.
|
||
3. Every client-to-server envelope and server dispatch is bound to that token.
|
||
The server discards traffic from superseded sessions, including late output.
|
||
4. The server queues work while a client is offline and dispatches it only when
|
||
the active session advertises capacity. A replacement session immediately
|
||
resumes non-terminal reconciliation.
|
||
|
||
## Command model
|
||
|
||
Each user request has a server-generated UUID (`issue_uuid`) and a durable
|
||
request record: target client, request/issue timestamps, shell type, command
|
||
text or script descriptor, CWD, environment overrides, resource-profile flags,
|
||
and lifecycle state. The UUID is the end-to-end idempotency key. Transport is
|
||
at-least-once, but the client durably remembers accepted UUIDs and never starts
|
||
the same request twice.
|
||
|
||
The server states are:
|
||
|
||
```text
|
||
queued -> dispatched -> accepted -> running -> succeeded | failed | terminated
|
||
\-------------------------------> cancelled
|
||
```
|
||
|
||
`accepted` means the client has durably accepted the request; `running` means
|
||
the process has been launched. `cancelled` is used when it is stopped before
|
||
launch. A kill racing launch is resolved by command revision: the client either
|
||
acknowledges cancellation before launch or launches then immediately applies
|
||
the requested signal, recording the race.
|
||
|
||
The client permits 16 concurrent processes and 100 pending commands by default.
|
||
Those values are configurable and advertised to the server. A full queue causes
|
||
a structured capacity error rather than creating unbounded work.
|
||
|
||
### Execution contract
|
||
|
||
- Shell selection is explicit: Unix-like clients support `sh` and `bash`;
|
||
Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`.
|
||
Unsupported shells are rejected; no fallback occurs.
|
||
- A command receives the daemon account's permissions and startup environment,
|
||
overlaid with the persisted `env_overrides` map. The effective CWD is the
|
||
requested existing directory or the registered daemon CWD when omitted.
|
||
- Each command is isolated into a process tree: a Unix session/process group or
|
||
a Windows Job Object. A daemon that cannot supervise its children terminates
|
||
them and reports interruption rather than claiming recovery it cannot make.
|
||
- Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`,
|
||
`MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in;
|
||
the concrete administrator-configured limits are applied with cgroup v2 when
|
||
available on Linux and Job Object limits on Windows. Unsupported requested
|
||
controls are reported, never silently ignored.
|
||
|
||
### Script execution
|
||
|
||
Scripts are content-addressed uploads, not shell-escaped command strings. The
|
||
server sends a descriptor containing SHA-256 and then ordered chunks (default
|
||
maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary
|
||
file beneath the effective CWD, executes it with the selected shell, and removes
|
||
it after the command reaches a terminal state. Script content is not copied into
|
||
the audit log; its digest and metadata are.
|
||
|
||
## Process control and diagnostics
|
||
|
||
On Unix, `kill` addresses the command's process group and accepts normal signal
|
||
names/numbers supported by that client. On Windows, only `SIGTERM` and `SIGKILL`
|
||
are valid. `SIGTERM` makes a best-effort `CTRL_BREAK_EVENT` delivery to the
|
||
dedicated console group, waits 10 seconds, then terminates the Job Object if
|
||
needed. `SIGKILL` immediately terminates the Job Object. The response reports
|
||
the actual escalation outcome.
|
||
|
||
Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is
|
||
emitted after the configurable default of 10 minutes without observable
|
||
progress. Linux enriches this with `/proc` state, CPU, memory, and I/O data;
|
||
Windows uses process and Job Object APIs where available. It is explicitly a
|
||
heuristic, not proof of an I/O stall.
|
||
|
||
## Durability, ordering, and retention
|
||
|
||
The server uses SQLite in WAL mode for metadata/indexes and append-only Zstandard
|
||
compressed segment files for output. It recovers non-terminal commands after a
|
||
restart and reconciles only these; server-confirmed historical terminal commands
|
||
are not re-reconciled. The client stores only active command state and output
|
||
not acknowledged by the server.
|
||
|
||
Every execution-originated event has a strictly increasing `event_seq` scoped to
|
||
one command. This includes lifecycle transitions, stdout/stderr chunks, stdin
|
||
acknowledgments, resource snapshots, signals, and terminal events. The server
|
||
preserves the sequence and additionally records receipt time. This is the
|
||
canonical reconstruction order across interleaved streams and retries.
|
||
Truncation is separate range metadata so it can truthfully describe missing
|
||
event sequences without consuming one itself.
|
||
|
||
Output chunks are Zstandard-compressed before persistent quota accounting.
|
||
Per-command history is a rolling compressed window (10 MiB default), so the
|
||
oldest output segments for that command are removed first and a sequence-range
|
||
truncation marker remains. This applies to active and terminal commands. When
|
||
connected, the client first removes server-acknowledged segments. While offline,
|
||
it must still honor both hard caps: it retains the newest tail, removes oldest
|
||
unacknowledged compressed chunks when necessary, and records their exact missing
|
||
ranges for durable reporting on reconnect.
|
||
|
||
Each client also has a 50 MiB aggregate compressed spool cap for active,
|
||
unacknowledged work. The server's matching per-client compressed-history cap is
|
||
50 MiB; it evicts that client's oldest terminal command records as needed. The
|
||
server-wide cap is 1 GiB; it evicts whole oldest terminal command records
|
||
(metadata and output), never arbitrary stdout/stderr rows. Active commands are
|
||
protected. If active commands alone consume a per-client budget, their oldest
|
||
acknowledged output rotates by the per-command rule; pipes continue draining so
|
||
a child cannot deadlock on output.
|
||
|
||
## Audit and timestamps
|
||
|
||
The server writes a separate durable audit trail for control actions and session
|
||
events: source/transport identity where available, target client, UUID, action,
|
||
request time, result, and error. It records command text, environment-override
|
||
names (not values), and script metadata/digest, but not duplicated
|
||
stdin/stdout/stderr payloads. The persisted execution request necessarily keeps
|
||
override values for dispatch/retry and must be access-controlled as sensitive
|
||
data. Audit retention is configured independently of output retention.
|
||
|
||
All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client
|
||
observed timestamps and server receipt timestamps are distinct; the latter is
|
||
authoritative for server records, while `event_seq` is authoritative for order.
|