docs: finalize rvbox v1 design and implementation plan
This commit is contained in:
@@ -0,0 +1,8 @@
|
||||
/bin/
|
||||
/dist/
|
||||
/coverage/
|
||||
*.db
|
||||
*.db-shm
|
||||
*.db-wal
|
||||
*.sock
|
||||
*.tmp
|
||||
@@ -11,6 +11,10 @@ Read the documents in this order:
|
||||
query/foreground semantics.
|
||||
4. [Platform and operations](platform-and-operations.md) — Unix/Windows
|
||||
contracts, recovery, storage safety, telemetry, and defaults.
|
||||
5. [Configuration contract](configuration.md) — strict TOML loading, shell
|
||||
resolution, cross-field validation, and annotated server/client examples.
|
||||
6. [Go implementation plan](implementation-plan.v1.md) — phased build order,
|
||||
package boundaries, storage/session/client details, tests, and release gates.
|
||||
|
||||
The wire authority is in [`../protos/rvbox/v1`](../protos/rvbox/v1):
|
||||
`common.proto` contains shared data types, `agent.proto` contains the
|
||||
|
||||
+171
-55
@@ -12,6 +12,9 @@ RVBox deliberately executes arbitrary commands with the identity, permissions,
|
||||
and base environment of the client daemon. It is therefore an administrative
|
||||
tool, not a multi-tenant remote-execution service.
|
||||
|
||||
Both daemons use strict TOML 1.0 configuration. The normative schema and fully
|
||||
annotated examples are in the [configuration contract](configuration.md).
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist
|
||||
@@ -26,10 +29,14 @@ tool, not a multi-tenant remote-execution service.
|
||||
|
||||
The server accepts a self-reported hostname as `client_id`; it is an opaque
|
||||
1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer
|
||||
registration for an ID replaces its prior live session. Consequently, a peer
|
||||
able to reach nginx can impersonate or take over a client ID. This is an
|
||||
accepted v1 limitation and deployments must restrict the endpoint to a trusted
|
||||
network.
|
||||
registration from the same durable client instance replaces its prior live
|
||||
session. A different instance is accepted normally when no session for that
|
||||
client ID is live. While one is live, a different instance is rejected unless
|
||||
an operator grants a one-shot override for that exact pending instance. This
|
||||
prevents accidental hostname collisions but is not authentication: a peer able
|
||||
to reach nginx and copy or guess the identifiers can still impersonate a
|
||||
client. This is an accepted v1 limitation and deployments must restrict the
|
||||
endpoint to a trusted network.
|
||||
|
||||
The Unix control socket is local-only and mode `0600`, owned by the server
|
||||
account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated;
|
||||
@@ -57,53 +64,113 @@ its dispatch loops.
|
||||
|
||||
1. The client connects over WSS and sends `ClientHello` with its client ID,
|
||||
protocol capability, OS/architecture, daemon version, current daemon CWD,
|
||||
supported shells, and a fresh reconnect UUID.
|
||||
supported shells, and a durable random client-instance UUID.
|
||||
2. The server accepts the current compatible protocol version, fences the
|
||||
previous connection for that ID, and returns a fresh server-issued
|
||||
`session_id` and monotonic `session_generation`.
|
||||
previous connection for that ID and same client instance, and returns a
|
||||
fresh server-issued `session_id` and monotonic `session_generation`. A live
|
||||
claim from a different client instance is rejected while the current session
|
||||
is live unless an operator has explicitly authorized that pending instance.
|
||||
With no live session, the new instance is accepted normally.
|
||||
3. Every client-to-server envelope and server dispatch is bound to that token.
|
||||
The server discards traffic from superseded sessions, including late output.
|
||||
4. The server queues work while a client is offline and dispatches it only when
|
||||
the active session advertises capacity. A replacement session immediately
|
||||
resumes non-terminal reconciliation.
|
||||
performs bidirectional reconciliation between the server's non-terminal set
|
||||
and the client's complete retained-command set. The server returns explicit
|
||||
local terminate/discard decisions before new dispatch begins.
|
||||
|
||||
## Command model
|
||||
|
||||
Each user request has a server-generated UUID (`issue_uuid`) and a durable
|
||||
request record: target client, request/issue timestamps, shell type, command
|
||||
text or script descriptor, CWD, environment overrides, resource-profile flags,
|
||||
and lifecycle state. The UUID is the end-to-end idempotency key. Transport is
|
||||
at-least-once, but the client durably remembers accepted UUIDs and never starts
|
||||
the same request twice.
|
||||
Each user request has a UUID (`issue_uuid`) and a durable request record: target
|
||||
client, request/issue timestamps, shell type, command text or script descriptor,
|
||||
CWD, environment overrides, resource-profile flags, and lifecycle state. `rvc`
|
||||
normally supplies this UUID as its optional `request_id`; the server generates
|
||||
one when it is omitted. Command and mutation identifiers are UUIDv7 values.
|
||||
Transport is at-least-once, but the client durably
|
||||
remembers accepted UUIDs and provides an at-most-once execution guarantee: it
|
||||
never authorizes the same request to execute twice. A crash in the launch window
|
||||
may interrupt a command before its requested code runs, but must never cause an
|
||||
automatic retry with an uncertain prior outcome.
|
||||
|
||||
After full terminal history is removed, each client retains a compact FIFO of
|
||||
the most recent 1,000,000 command tombstones containing binary UUIDv7,
|
||||
immutable request hash, and completion/acknowledgement time. The server retains
|
||||
the most recent 1,000,000 command tombstones globally as a first-line duplicate
|
||||
check. A matching UUID/hash returns structured `ALREADY_EXECUTED`; a matching
|
||||
UUID with different immutable content is a conflict. These ledgers have separate
|
||||
count-based budgets and do not retain command payload or output. Replay
|
||||
protection older than the retained client tombstone horizon is best-effort.
|
||||
|
||||
The server states are:
|
||||
|
||||
```text
|
||||
queued -> dispatched -> accepted -> running -> succeeded | failed | terminated
|
||||
\-------------------------------> cancelled
|
||||
queued -> dispatched | cancelled | expired
|
||||
dispatched -> queued | accepted | rejected | cancelled
|
||||
accepted -> running | rejected | cancelled
|
||||
running -> succeeded | failed | terminated | interrupted
|
||||
```
|
||||
|
||||
`accepted` means the client has durably accepted the request; `running` means
|
||||
the process has been launched. `cancelled` is used when it is stopped before
|
||||
launch. A kill racing launch is resolved by command revision: the client either
|
||||
acknowledges cancellation before launch or launches then immediately applies
|
||||
the requested signal, recording the race.
|
||||
Recovery/corruption handling may also move an affected `dispatched` or
|
||||
`accepted` command to `interrupted`; those are exceptional reconciliation
|
||||
transitions, not normal execution outcomes.
|
||||
|
||||
`accepted` means the client has durably admitted the request; `running` means
|
||||
the requested process has been launched. `cancelled` is used when it is stopped
|
||||
before launch. A permanent pre-launch validation, script-transfer, or process-
|
||||
preparation failure is terminal `rejected`; `failed` is reserved for code that
|
||||
actually launched. A transient
|
||||
capacity rejection returns to `queued` with backoff and remains subject to its
|
||||
queue TTL. The server may cancel work that has never been dispatched immediately.
|
||||
For dispatched or accepted work, it persists a higher command revision and
|
||||
sends the revisioned signal without prematurely declaring a terminal state. The
|
||||
client either returns a revisioned `cancelled` lifecycle before launch
|
||||
authorization or applies the signal after authorization and returns a
|
||||
revisioned signal result. Cancellation intent remains internal rather than
|
||||
adding a public lifecycle state.
|
||||
|
||||
Queued work has a configurable acceptance deadline, default 15 minutes; zero
|
||||
explicitly means no expiry. A command that was never dispatched becomes
|
||||
terminal `expired` at its deadline. A dispatched command whose acceptance is
|
||||
uncertain remains non-terminal and is displayed as expired pending
|
||||
reconciliation. Later client evidence updates the actual lifecycle. Acceptance
|
||||
or execution observed after the deadline creates an incident and is displayed
|
||||
as late-after-expiry; the server requests termination but continues recording
|
||||
the actual client-reported outcome.
|
||||
|
||||
Process launch has internal durable phases `launch_prepared` and
|
||||
`launch_authorized` between public `accepted` and `running`. The client creates
|
||||
the process behind an OS-specific execution barrier, durably records its process
|
||||
identity, durably authorizes launch, and only then releases requested command
|
||||
code. Once authorization is durable, an uncertain outcome is reconciled as
|
||||
`interrupted`, never by redispatching that UUID.
|
||||
|
||||
The client permits 16 concurrent processes and 100 pending commands by default.
|
||||
Those values are configurable and advertised to the server. A full queue causes
|
||||
a structured capacity error rather than creating unbounded work.
|
||||
The server separately defaults to at most 1,000 queued commands for one target
|
||||
client and 10,000 queued commands globally, subject to the stricter byte quotas.
|
||||
|
||||
### Execution contract
|
||||
|
||||
- Shell selection is explicit: Unix-like clients support `sh` and `bash`;
|
||||
Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`.
|
||||
Unsupported shells are rejected; no fallback occurs.
|
||||
- The client materializes `command_text` as a private generated wrapper and
|
||||
executes that file with exactly the selected shell: `sh`/`bash`, `cmd`, or
|
||||
PowerShell. This avoids transport quoting and Windows command-line limits;
|
||||
it never performs shell detection or fallback. Wrapper cleanup follows the
|
||||
same terminal rule as uploaded scripts.
|
||||
- A command receives the daemon account's permissions and startup environment,
|
||||
overlaid with the persisted `env_overrides` map. The effective CWD is the
|
||||
requested existing directory or the registered daemon CWD when omitted.
|
||||
- Each command is isolated into a process tree: a Unix session/process group or
|
||||
a Windows Job Object. A daemon that cannot supervise its children terminates
|
||||
them and reports interruption rather than claiming recovery it cannot make.
|
||||
- Terminal lifecycle means the supervised tree is empty, not only that its root
|
||||
shell exited. After root exit, RVBox allows a configurable 5-second descendant
|
||||
and output-drain grace period, then terminates residual descendants. It drains
|
||||
capture pipes and durably sequences all retained output (or an explicit
|
||||
`OutputIncomplete` marker) before emitting the terminal lifecycle event.
|
||||
- Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`,
|
||||
`MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in;
|
||||
the concrete administrator-configured limits are applied with cgroup v2 when
|
||||
@@ -115,18 +182,20 @@ a structured capacity error rather than creating unbounded work.
|
||||
Scripts are content-addressed uploads, not shell-escaped command strings. The
|
||||
server sends a descriptor containing SHA-256 and then ordered chunks (default
|
||||
maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary
|
||||
file beneath the effective CWD, executes it with the selected shell, and removes
|
||||
it after the command reaches a terminal state. Script content is not copied into
|
||||
the audit log; its digest and metadata are.
|
||||
file beneath the effective CWD, executes it with exactly the selected shell, and
|
||||
removes it after the command reaches a terminal state. The descriptor filename
|
||||
is display metadata only and is never used as a path component. Script content
|
||||
is not copied into the audit log; its digest and metadata are.
|
||||
|
||||
## Process control and diagnostics
|
||||
|
||||
On Unix, `kill` addresses the command's process group and accepts normal signal
|
||||
names/numbers supported by that client. On Windows, only `SIGTERM` and `SIGKILL`
|
||||
are valid. `SIGTERM` makes a best-effort `CTRL_BREAK_EVENT` delivery to the
|
||||
dedicated console group, waits 10 seconds, then terminates the Job Object if
|
||||
needed. `SIGKILL` immediately terminates the Job Object. The response reports
|
||||
the actual escalation outcome.
|
||||
On Unix, `kill` addresses the command's process group and accepts the portable
|
||||
v1 set `HUP`, `INT`, `TERM`, `KILL`, `USR1`, and `USR2` (including their
|
||||
`SIG`-prefixed CLI spellings). Arbitrary native signal numbers are not part of
|
||||
v1. On Windows, only `TERM` and `KILL` are valid. `TERM` makes a best-effort
|
||||
`CTRL_BREAK_EVENT` delivery to the dedicated console group, waits 10 seconds,
|
||||
then terminates the Job Object if needed. `KILL` immediately terminates the Job
|
||||
Object. The response reports the actual escalation outcome.
|
||||
|
||||
Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is
|
||||
emitted after the configurable default of 10 minutes without observable
|
||||
@@ -142,31 +211,73 @@ restart and reconciles only these; server-confirmed historical terminal commands
|
||||
are not re-reconciled. The client stores only active command state and output
|
||||
not acknowledged by the server.
|
||||
|
||||
Every execution-originated event has a strictly increasing `event_seq` scoped to
|
||||
one command. This includes lifecycle transitions, stdout/stderr chunks, stdin
|
||||
acknowledgments, resource snapshots, signals, and terminal events. The server
|
||||
preserves the sequence and additionally records receipt time. This is the
|
||||
canonical reconstruction order across interleaved streams and retries.
|
||||
Truncation is separate range metadata so it can truthfully describe missing
|
||||
event sequences without consuming one itself.
|
||||
Every transmitted execution event has a strictly increasing `event_seq` scoped
|
||||
to one command. The client first stores lifecycle transitions, stdout/stderr,
|
||||
stdin acknowledgements, resource snapshots, signals, and terminal state in a
|
||||
durable local order. It durably assigns wire sequences only as entries enter the
|
||||
bounded send window, after any unsent-output compaction. Once assigned, an event
|
||||
is pinned until acknowledged and retry content is immutable. A client-side
|
||||
truncation marker therefore consumes a normal sequence without creating a wire
|
||||
gap. Server-side retention exposes removed event ranges as query metadata
|
||||
without allocating client sequences. The server also records receipt time;
|
||||
`event_seq` remains the canonical transmitted order across streams and retries.
|
||||
|
||||
Output chunks are Zstandard-compressed before persistent quota accounting.
|
||||
Per-command history is a rolling compressed window (10 MiB default), so the
|
||||
oldest output segments for that command are removed first and a sequence-range
|
||||
truncation marker remains. This applies to active and terminal commands. When
|
||||
connected, the client first removes server-acknowledged segments. While offline,
|
||||
it must still honor both hard caps: it retains the newest tail, removes oldest
|
||||
unacknowledged compressed chunks when necessary, and records their exact missing
|
||||
ranges for durable reporting on reconnect.
|
||||
The terminal lifecycle event is always the final client event for a command.
|
||||
The control-plane follow wrapper emits any server-created retention metadata
|
||||
before returning that terminal event. Followers may therefore stop on terminal
|
||||
without missing subsequently sequenced stdout/stderr or known truncation data.
|
||||
|
||||
Each client also has a 50 MiB aggregate compressed spool cap for active,
|
||||
unacknowledged work. The server's matching per-client compressed-history cap is
|
||||
50 MiB; it evicts that client's oldest terminal command records as needed. The
|
||||
server-wide cap is 1 GiB; it evicts whole oldest terminal command records
|
||||
(metadata and output), never arbitrary stdout/stderr rows. Active commands are
|
||||
protected. If active commands alone consume a per-client budget, their oldest
|
||||
acknowledged output rotates by the per-command rule; pipes continue draining so
|
||||
a child cannot deadlock on output.
|
||||
All command-owned stored data is quota-accounted: execution metadata, script
|
||||
body, pending stdin, events, and output. Stored message/blob payloads are
|
||||
Zstandard-compressed; necessary SQLite index/state columns are charged by their
|
||||
encoded lengths plus a conservative versioned per-row/index overhead rather
|
||||
than pretending they are free. This logical accounting is deterministic across
|
||||
SQLite compaction. A separate filesystem free-space floor protects WAL,
|
||||
temporary files, tombstones, and accounting variance.
|
||||
The default total is 32 MiB per command, 256 MiB per client on both client and
|
||||
server, and 4 GiB server-wide. A separate 10 MiB rolling output window remains
|
||||
per command, and raw script input remains limited to 10 MiB. Client accounting
|
||||
also charges the raw generated execution wrapper/script file while it exists,
|
||||
even though the compressed durable source was already charged.
|
||||
|
||||
Each accepted active command reserves 64 KiB of its quota for bounded closeout
|
||||
metadata. Essential state for an accepted stdin/signal mutation is additionally
|
||||
reserved before that mutation succeeds. Essential active state is never silently
|
||||
rolled. Server output is evictable: the oldest retained chunks are removed first
|
||||
and sequence-range query metadata remains. The client removes acknowledged data,
|
||||
pins its bounded assigned send window, and may replace older unsequenced output
|
||||
with normally sequenced byte-loss markers. While offline or under sustained
|
||||
overload, it retains the newest tail and records exact known lost bytes. At
|
||||
client/server aggregate limits, terminal commands are evicted as whole UUID
|
||||
records in oldest server issue-time/UUIDv7 order first. If active data
|
||||
alone reaches a limit, output rotates or enters loss mode and new essential
|
||||
allocations are rejected with `CAPACITY_EXHAUSTED`; pipes continue draining.
|
||||
|
||||
Transient raw output also has bounded high/low watermarks before compression:
|
||||
1 MiB/256 KiB per command, 8 MiB/4 MiB per client daemon, and 64 MiB/32 MiB
|
||||
server-wide by default. Crossing a client high watermark enters loss mode;
|
||||
still-unsequenced output bytes may be discarded before compression and replaced
|
||||
in durable local order by a later `OutputTruncation`. Loss mode ends only below
|
||||
the corresponding low watermark. The server never drops an already sequenced
|
||||
client event: at its ingress high watermark it withholds acknowledgement and
|
||||
closes an overproducing session if bounded admission cannot continue, letting
|
||||
the durable client retry after the server backlog falls below its low watermark.
|
||||
Lifecycle,
|
||||
stdin acknowledgements, signal results, and truncation/incomplete markers use
|
||||
reserved capacity and are never treated as droppable output.
|
||||
|
||||
Independently of byte pressure, the server reclaims each whole terminal command
|
||||
and all command-owned data 30 days after its terminal time by default. Zero
|
||||
explicitly disables age rotation. Compact replay tombstones, audit records, and
|
||||
storage incidents remain under their separate retention policies.
|
||||
|
||||
Storage health is derived from durable incident records. Safe repairs resolve
|
||||
an incident automatically; known data loss remains dirty until an operator
|
||||
explicitly acknowledges it. Resolution clears the dirty health flag but does
|
||||
not erase records as part of that action. Unresolved compact records are
|
||||
non-evictable; resolved summaries and audit entries follow the separate 100 MiB
|
||||
audit/incident history rotation. Recovery and incident management are described
|
||||
in the platform contract and exposed through the control plane.
|
||||
|
||||
## Audit and timestamps
|
||||
|
||||
@@ -175,8 +286,13 @@ events: source/transport identity where available, target client, UUID, action,
|
||||
request time, result, and error. It records command text, environment-override
|
||||
names (not values), and script metadata/digest, but not duplicated
|
||||
stdin/stdout/stderr payloads. The persisted execution request necessarily keeps
|
||||
override values for dispatch/retry and must be access-controlled as sensitive
|
||||
data. Audit retention is configured independently of output retention.
|
||||
override values for dispatch/retry. In v1, command text, scripts, environment
|
||||
values, stdin, output, and other command-owned payloads are stored in plaintext;
|
||||
application-level encryption and key management are deferred to a future
|
||||
version. Audit storage uses compressed segments under a separate 100 MiB default
|
||||
quota and rotates complete oldest segments. Audit retention is configured
|
||||
independently of command retention; its default zero age limit means quota-only
|
||||
rotation.
|
||||
|
||||
All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client
|
||||
observed timestamps and server receipt timestamps are distinct; the latter is
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
# RVBox v1 configuration contract
|
||||
|
||||
RVBox v1 uses TOML 1.0 for daemon configuration. The normative annotated
|
||||
examples are [`examples/server.toml`](examples/server.toml) and
|
||||
[`examples/client.toml`](examples/client.toml). They list every supported v1
|
||||
knob, with the routing, persistence, and safety limits first in each section.
|
||||
|
||||
## Loading and precedence
|
||||
|
||||
- `rvbox-server --config PATH` and `rvbox --config PATH` load one UTF-8 TOML
|
||||
file. There is no implicit merge of multiple files and no hot reload in v1.
|
||||
- Precedence is compiled default, then TOML, then an explicitly supplied CLI
|
||||
flag. Flags exist for operationally important scalar keys; they use the same
|
||||
validation as TOML. RVBox does not implicitly import configuration from
|
||||
environment variables.
|
||||
- Unknown keys, duplicate keys/tables, type mismatches, invalid UTF-8, and
|
||||
values outside documented ranges are startup errors. Parsing never silently
|
||||
substitutes a default for a present invalid value.
|
||||
- Durations are quoted Go-style duration strings such as `"250ms"`, `"15m"`,
|
||||
and `"720h"`. Byte sizes and counts are base-10 TOML integers whose values are
|
||||
bytes; comments show the equivalent binary unit. URLs and paths are strings.
|
||||
- Relative paths are rejected for state, socket, CA, shell-executable, and
|
||||
allowed-CWD-root fields. `client.daemon_cwd` is resolved once at startup and
|
||||
then stored and advertised as an absolute path. The annotated client example
|
||||
uses Unix paths; a Windows deployment replaces `state_dir`, `daemon_cwd`, and
|
||||
relevant shell paths with absolute Windows paths. Shell fields for the other
|
||||
platform are syntax-checked but not resolved or advertised.
|
||||
- The daemon prints its effective configuration after validation, with no
|
||||
command data or TLS material. Since v1 stores command/environment payloads in
|
||||
plaintext, configuration output is hygiene rather than a secrecy guarantee.
|
||||
|
||||
## Limits and cross-field validation
|
||||
|
||||
Configuration may lower protocol and storage limits, but may not raise a hard
|
||||
wire ceiling above the v1 values in the examples. For every high/low watermark,
|
||||
`0 < low < high`; send windows must fit below their corresponding durable quota.
|
||||
The 64 KiB command closeout reserve must fit within the command quota, and the
|
||||
command quota must fit within the client and server tiers. Queue and byte counts
|
||||
must be positive except where a comment explicitly gives zero a disabling or
|
||||
indefinite meaning. `queue.max_per_client` may not exceed `queue.max_server`.
|
||||
|
||||
The server must reject external JSON-RPC binds unless `json_rpc.enabled=true`.
|
||||
Any enabled non-loopback bind produces a conspicuous warning but is permitted by
|
||||
the accepted v1 debugging contract. The control Unix socket always uses mode
|
||||
`0600`; it is not a configurable relaxation.
|
||||
|
||||
## Shell executable resolution
|
||||
|
||||
Each `ShellType` maps to one startup-validated absolute executable path from
|
||||
`[shells]`. The client canonicalizes the path, verifies that it names an
|
||||
executable regular file appropriate to the platform, and advertises only shells
|
||||
that passed validation. The configured platform default must be one of those
|
||||
shells. There is no fallback.
|
||||
|
||||
The resolved executable is independent of a command's `PATH` override. For
|
||||
example, a request for `SHELL_BASH` still launches the validated `/bin/bash`
|
||||
even when the request contains `PATH=/tmp/untrusted`; it never searches that
|
||||
directory for another `bash`. An administrator may intentionally select another
|
||||
implementation, such as an absolute `pwsh.exe` path for `SHELL_POWERSHELL`, but
|
||||
the selection remains fixed until daemon restart.
|
||||
|
||||
Command text and uploaded scripts are written to generated wrapper paths and
|
||||
passed to exactly this executable. The user-supplied script filename is display
|
||||
metadata only.
|
||||
|
||||
## Resource-profile composition
|
||||
|
||||
`LIGHT` is exclusive. Otherwise, a request may combine at most one CPU tier,
|
||||
one memory tier, and one disk tier. Thus `CPU_HEAVY + MEM_MEDIUM` is valid, while
|
||||
`CPU_MEDIUM + CPU_HEAVY` and `LIGHT + MEM_HEAVY` are invalid. A configured
|
||||
profile declares which controls are required. If the platform cannot apply a
|
||||
required control atomically before launch, the command is terminal `REJECTED`.
|
||||
|
||||
Profile names describe administrator-defined allowance classes: a `HEAVY` tier
|
||||
normally permits more resources than `MEDIUM`; RVBox does not invent numeric
|
||||
values. Zero for an individual numeric limit means that control is not requested
|
||||
by that profile, but every name in `required_controls` must have a nonzero,
|
||||
platform-applicable value. Linux disk limits use configured cgroup device
|
||||
major/minor keys. Windows applies the equivalent whole-Job rate control and
|
||||
ignores Linux device maps only when disk control is not declared required.
|
||||
|
||||
## Validation ownership
|
||||
|
||||
`internal/config` owns TOML DTOs, strict decoding, default application, flag
|
||||
overrides, canonicalization, and cross-field validation. It converts the parsed
|
||||
form into immutable domain configuration before listeners or child processes
|
||||
start. Network, storage, and supervisor packages receive only their relevant
|
||||
validated sub-configuration and never parse TOML themselves.
|
||||
+86
-21
@@ -12,32 +12,83 @@ an optional JSON-RPC 2.0 HTTP adapter for local debugging and batch automation.
|
||||
It has no authentication by design. Binding it beyond loopback is an explicit
|
||||
deployment choice and requires external protection.
|
||||
|
||||
gRPC can stream `RunCommandAndFollow` and `FollowCommand`. JSON-RPC remains
|
||||
simple: callers issue work, query command state, poll event/output pages after
|
||||
an event sequence, append stdin, close stdin, or signal a command. It does not
|
||||
invent a separate event-stream protocol.
|
||||
gRPC streams command history and live events through `FollowCommand`. JSON-RPC
|
||||
remains simple: callers issue work, query command state, poll event/output pages
|
||||
after an event sequence, append stdin, close stdin, or signal a command. It does
|
||||
not invent a separate event-stream protocol.
|
||||
|
||||
The JSON-RPC method names are the lower-camel protobuf operation names:
|
||||
`listClients`, `getClient`, `listCommands`, `getCommand`, `runCommand`,
|
||||
`appendStdin`, `closeStdin`, `signalCommand`, and `getOutput`. Parameters and
|
||||
results use protobuf JSON mapping (including base64 strings for `bytes` and UTC
|
||||
RFC 3339 strings for timestamps); JSON-RPC errors carry the corresponding
|
||||
`ControlError` code/data. `getOutput` and `getCommand` are the polling path for
|
||||
what gRPC exposes as follow streams.
|
||||
`appendStdin`, `closeStdin`, `signalCommand`, `getOutput`, and the three storage
|
||||
incident methods documented below. Parameters and results use protobuf JSON
|
||||
mapping (including base64 strings for `bytes` and UTC RFC 3339 strings for
|
||||
timestamps). Control response messages contain successful results only. gRPC
|
||||
failures use canonical non-OK status codes with structured RVBox details where
|
||||
needed; the JSON-RPC adapter maps the same domain errors to standard JSON-RPC
|
||||
error objects. `getOutput` and `getCommand` are the polling path for what gRPC
|
||||
exposes as follow streams.
|
||||
|
||||
`rvc stat CLIENT` also shows the active durable client-instance ID and the most
|
||||
recent different instance rejected while that client is live. An operator may
|
||||
run `rvc client takeover CLIENT INSTANCE-ID`; this creates a one-shot 5-minute
|
||||
authorization for that exact pending claim. Its matching reconnect consumes the
|
||||
authorization and fences the old session. If the old session is no longer live,
|
||||
the replacement connects normally without this command. The JSON-RPC method is
|
||||
`authorizeClientTakeover`.
|
||||
|
||||
Both transports enforce the same decoded field limits. A control gRPC request
|
||||
may be at most 16 MiB; a JSON-RPC HTTP body may be at most 24 MiB to accommodate
|
||||
base64 expansion of the 10 MiB script maximum. `ExecutionSpec` itself may be at
|
||||
most 768 KiB, which also keeps its agent dispatch below the 1 MiB envelope cap.
|
||||
|
||||
## CLI semantics
|
||||
|
||||
`rvc stat` maps to `ListClients`, `GetClient`, `ListCommands`, and `GetCommand`.
|
||||
History pages default to 20 commands and may request at most 100. Output pages
|
||||
default to 100 lines; a line is a display operation over ordered chunks, not a
|
||||
protocol boundary. Output can be filtered by stream and timestamped with the
|
||||
server's recorded client-observed timestamp plus stream name.
|
||||
History pages default to 20 commands and may request at most 100. Historical
|
||||
output uses opaque, byte-bounded cursors and may resume within an output event;
|
||||
it is not numbered or paginated by lines. The server returns uncompressed output
|
||||
slices with event sequence, byte offset, stream, observed timestamp, and server
|
||||
receipt timestamp. `rvc` may render line-oriented human output, but line
|
||||
boundaries are not storage or pagination boundaries.
|
||||
|
||||
`rvc run` creates a command. Foreground mode runs `RunCommandAndFollow`, which
|
||||
streams output and stops on a terminal event. `--background` uses `RunCommand`
|
||||
and returns the UUID immediately. Interrupting the CLI, timing out its local
|
||||
wait, or losing the local control connection never cancels remote work. The
|
||||
explicit `rvc kill` operation is the only termination path.
|
||||
`rvc run` always creates a durable command through unary `RunCommand`, which
|
||||
returns the UUID. Foreground mode then calls `FollowCommand` with that UUID,
|
||||
`after_event_seq=0`, and `include_existing=true`, streaming output until a
|
||||
terminal event. `--background` returns immediately after `RunCommand`.
|
||||
Interrupting the CLI, timing out its local wait, or losing the local control
|
||||
connection never cancels remote work. The explicit `rvc kill` operation is the
|
||||
only termination path. A foreground caller can resume `FollowCommand` after its
|
||||
last received event sequence without missing durable history.
|
||||
|
||||
`FollowCommandResponse` wraps either a client-sequenced `CommandEvent` or a
|
||||
server-created retention marker and exposes server receipt/recording time
|
||||
separately from client observation time. A retention marker does not advance the
|
||||
resume cursor. When retained history is incomplete, the server emits the relevant
|
||||
marker before any terminal event on that stream, so a follower never stops on
|
||||
terminal while believing truncated history was complete.
|
||||
|
||||
Mutating requests accept an optional `request_id`. `rvc` generates one per
|
||||
mutation and reuses it for transport retries; `--request-id` lets automation
|
||||
reuse it across CLI invocations. The CLI accepts that option globally and drops
|
||||
it for reads. For `RunCommand`, the supplied request ID is the command's
|
||||
`issue_uuid`. For stdin, close, signal, repair, and acknowledgement operations
|
||||
it deduplicates that action while `issue_uuid` or `incident_id` continues to
|
||||
identify the target. The server generates a request ID when omitted, preserving
|
||||
simple JSON-RPC use.
|
||||
|
||||
The server stores the mutation kind, target, immutable request hash, and result
|
||||
under that ID. An identical retry returns the original result; reuse with any
|
||||
different method, target, or content returns `CONFLICT`. The record is owned by
|
||||
the affected command or incident for quota and retention. A retry after that
|
||||
owner has been reclaimed cannot repeat the action: it returns retained
|
||||
tombstone information where available or `NOT_FOUND`.
|
||||
|
||||
`rvc run --queue-ttl` controls how long work may wait for server-confirmed
|
||||
acceptance and defaults to 15 minutes; zero means indefinite. `rvc stat`
|
||||
distinguishes terminal `Expired`, `Expired (awaiting reconciliation)`, and
|
||||
actual lifecycle with a late-after-expiry warning. It renders terminal
|
||||
`Rejected` with the client's structured validation/platform reason; `Failed`
|
||||
means the requested code actually launched.
|
||||
|
||||
`rvc append` turns a string into `StdinWrite` with `append_newline=true` unless
|
||||
the caller selects raw mode; `--file` supplies raw bytes; `--attach` streams
|
||||
@@ -53,6 +104,20 @@ must not be shown as a terminal state. Pagination response cursors are stable
|
||||
within their declared ordering (newest issue time for command lists; increasing
|
||||
`event_seq` for events/output).
|
||||
|
||||
The service returns `ControlError` codes for not found, offline, capacity,
|
||||
invalid request, unsupported platform feature, conflict, truncation, and
|
||||
internal/transient errors. It never encodes errors only as CLI text.
|
||||
Control failures are never embedded in otherwise-successful response messages.
|
||||
They use canonical gRPC status codes with structured RVBox details; JSON-RPC
|
||||
returns the corresponding JSON-RPC error object, and `rvc` maps the same domain
|
||||
error to a stable exit code rather than parsing text.
|
||||
|
||||
## Storage incidents
|
||||
|
||||
`rvc storage incidents` lists unresolved storage incidents by default and can
|
||||
include resolved history. `rvc storage repair INCIDENT` attempts only a known
|
||||
safe repair. `rvc storage acknowledge INCIDENT --note ...` accepts documented
|
||||
irrecoverable loss and clears that incident from dirty health. Both mutations
|
||||
use the same optional `--request-id` behavior as other mutations. An
|
||||
acknowledgement note is required and bounded to 4 KiB. Repair or acknowledgement
|
||||
changes incident state; it does not delete history as part of that action;
|
||||
resolved history later follows the independent audit/incident quota. The
|
||||
JSON-RPC equivalents are `listStorageIncidents`,
|
||||
`repairStorageIncident`, and `acknowledgeStorageIncident`.
|
||||
|
||||
@@ -0,0 +1,205 @@
|
||||
# RVBox v1 client example. Integer sizes are bytes; durations are quoted strings.
|
||||
|
||||
[client]
|
||||
# Reverse WebSocket endpoint exposed by nginx.
|
||||
server_url = "wss://rvbox.example.test/v1/agent"
|
||||
# Private durable accepted-command, event-spool, and tombstone root.
|
||||
state_dir = "/var/lib/rvbox"
|
||||
# Empty selects the local hostname; otherwise use an opaque 1-128 ASCII ID.
|
||||
client_id = ""
|
||||
# Default absolute CWD when a request omits cwd.
|
||||
daemon_cwd = "/"
|
||||
# Maximum simultaneously running supervised process trees.
|
||||
max_running_commands = 16
|
||||
# Maximum durably accepted commands waiting to start.
|
||||
max_queued_commands = 100
|
||||
# Grace for reserved terminal cleanup during orderly daemon shutdown.
|
||||
shutdown_grace = "30s"
|
||||
|
||||
[tls]
|
||||
# Optional PEM CA bundle; empty uses the operating-system trust store.
|
||||
ca_file = ""
|
||||
# Optional certificate name override; empty derives it from server_url.
|
||||
server_name = ""
|
||||
|
||||
[shells]
|
||||
# Default shell enum on Unix; the matching path must validate at startup.
|
||||
default_unix = "sh"
|
||||
# Default shell enum on Windows; the matching path must validate at startup.
|
||||
default_windows = "powershell"
|
||||
# Absolute executable used for SHELL_SH; empty marks it unsupported.
|
||||
sh = "/bin/sh"
|
||||
# Absolute executable used for SHELL_BASH; empty marks it unsupported.
|
||||
bash = "/bin/bash"
|
||||
# Absolute executable used for SHELL_CMD on Windows; empty marks it unsupported.
|
||||
cmd = "C:\\Windows\\System32\\cmd.exe"
|
||||
# Absolute executable used for SHELL_POWERSHELL; may instead point to pwsh.exe.
|
||||
powershell = "C:\\Windows\\System32\\WindowsPowerShell\\v1.0\\powershell.exe"
|
||||
# Absolute CWD roots permitted by local policy; empty allows any accessible path.
|
||||
allowed_cwd_roots = []
|
||||
|
||||
[network]
|
||||
# Send WebSocket Ping after this period without inbound activity.
|
||||
heartbeat_idle = "10s"
|
||||
# Close and reconnect after this total period without inbound activity.
|
||||
liveness_timeout = "30s"
|
||||
# Initial full-jitter reconnect backoff.
|
||||
reconnect_initial = "1s"
|
||||
# Maximum full-jitter reconnect backoff.
|
||||
reconnect_max = "60s"
|
||||
# Continuous session duration that resets reconnect backoff.
|
||||
stable_session_reset = "60s"
|
||||
# Timeout for DNS/TCP/TLS/WebSocket establishment.
|
||||
connect_timeout = "15s"
|
||||
# Deadline for an individual WebSocket data-frame write.
|
||||
write_deadline = "10s"
|
||||
|
||||
[storage]
|
||||
# Maximum rolling compressed output retained for one command (10 MiB).
|
||||
command_output_limit_bytes = 10485760
|
||||
# Maximum charged data, including raw execution wrapper, per command (32 MiB).
|
||||
command_total_limit_bytes = 33554432
|
||||
# Maximum charged command data across this client daemon (256 MiB).
|
||||
client_total_limit_bytes = 268435456
|
||||
# Compact completed-command replay ledger entry cap.
|
||||
tombstone_max_entries = 1000000
|
||||
# Per-active-command quota held for terminal and loss metadata (64 KiB).
|
||||
command_closeout_reserve_bytes = 65536
|
||||
# Reject new unreserved writes below this filesystem free space (64 MiB).
|
||||
free_space_floor_bytes = 67108864
|
||||
# Target size for a sealed append-only spool segment (256 KiB).
|
||||
segment_target_bytes = 262144
|
||||
# Maximum time a group-commit waits before fsync; acknowledgements wait too.
|
||||
durability_interval = "100ms"
|
||||
|
||||
[flow]
|
||||
# Enter per-command raw-output loss mode at this backlog (1 MiB).
|
||||
raw_output_command_high_bytes = 1048576
|
||||
# Leave per-command raw-output loss mode below this backlog (256 KiB).
|
||||
raw_output_command_low_bytes = 262144
|
||||
# Enter client-wide raw-output loss mode at this backlog (8 MiB).
|
||||
raw_output_client_high_bytes = 8388608
|
||||
# Leave client-wide raw-output loss mode below this backlog (4 MiB).
|
||||
raw_output_client_low_bytes = 4194304
|
||||
# Maximum assigned but unacknowledged encoded bytes per command (1 MiB).
|
||||
unacknowledged_per_command_bytes = 1048576
|
||||
# Maximum assigned but unacknowledged encoded bytes for the session (8 MiB).
|
||||
unacknowledged_per_session_bytes = 8388608
|
||||
|
||||
[execution]
|
||||
# Wait after root exit for descendants and capture EOF before forced cleanup.
|
||||
descendant_drain_grace = "5s"
|
||||
# Wait after Windows CTRL_BREAK before terminating the complete Job Object.
|
||||
windows_term_grace = "10s"
|
||||
# Mark suspected_hung after no observable progress for this duration.
|
||||
hung_threshold = "10m"
|
||||
# Interval for best-effort process/resource diagnostic snapshots.
|
||||
diagnostic_interval = "30s"
|
||||
# Maximum raw uploaded script or generated script body (10 MiB).
|
||||
max_script_bytes = 10485760
|
||||
# Maximum serialized command ExecutionSpec accepted from the server (768 KiB).
|
||||
max_execution_spec_bytes = 786432
|
||||
# Maximum decoded AgentEnvelope accepted from the server (1 MiB).
|
||||
max_agent_envelope_bytes = 1048576
|
||||
# Maximum uncompressed stdout/stderr chunk emitted to the protocol (64 KiB).
|
||||
max_raw_chunk_bytes = 65536
|
||||
# Maximum protocol detail/reason text encoded as UTF-8 (4 KiB).
|
||||
protocol_detail_max_bytes = 4096
|
||||
|
||||
[observability]
|
||||
# Loopback HTTP listener for local liveness, storage, and supervisor health.
|
||||
listen = "127.0.0.1:6902"
|
||||
# Liveness route.
|
||||
liveness_path = "/livez"
|
||||
# Readiness route; false while reconciliation or essential recovery is pending.
|
||||
readiness_path = "/readyz"
|
||||
# Prometheus metrics route.
|
||||
metrics_path = "/metrics"
|
||||
# Structured logging threshold: debug, info, warn, or error.
|
||||
log_level = "info"
|
||||
# Structured log encoding: json or text.
|
||||
log_format = "json"
|
||||
|
||||
# Resource profiles are administrator policy. These illustrative values are not
|
||||
# protocol guarantees. LIGHT is exclusive; otherwise combine at most one tier
|
||||
# from each cpu_*, mem_*, and disk_* family.
|
||||
|
||||
[profiles.light]
|
||||
# Advertise this profile when all required controls validate.
|
||||
enabled = true
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["cpu", "memory", "pids"]
|
||||
# CPU allowance as percent of one logical CPU; zero omits CPU control.
|
||||
cpu_percent = 50
|
||||
# Hard resident/commit memory allowance (512 MiB); zero omits memory control.
|
||||
memory_max_bytes = 536870912
|
||||
# Maximum processes in the supervised tree; zero omits process-count control.
|
||||
pids_max = 64
|
||||
# Windows whole-Job read rate; zero omits it.
|
||||
windows_io_read_bps = 0
|
||||
# Windows whole-Job write rate; zero omits it.
|
||||
windows_io_write_bps = 0
|
||||
# Linux cgroup read rates keyed by device major:minor.
|
||||
linux_io_read_bps = {}
|
||||
# Linux cgroup write rates keyed by device major:minor.
|
||||
linux_io_write_bps = {}
|
||||
|
||||
[profiles.cpu_medium]
|
||||
# Advertise the CPU_MEDIUM allowance class.
|
||||
enabled = true
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["cpu"]
|
||||
# CPU allowance as percent of one logical CPU.
|
||||
cpu_percent = 200
|
||||
|
||||
[profiles.cpu_heavy]
|
||||
# Advertise the CPU_HEAVY allowance class.
|
||||
enabled = true
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["cpu"]
|
||||
# CPU allowance as percent of one logical CPU.
|
||||
cpu_percent = 800
|
||||
|
||||
[profiles.mem_medium]
|
||||
# Advertise the MEM_MEDIUM allowance class.
|
||||
enabled = true
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["memory"]
|
||||
# Hard memory allowance (2 GiB).
|
||||
memory_max_bytes = 2147483648
|
||||
|
||||
[profiles.mem_heavy]
|
||||
# Advertise the MEM_HEAVY allowance class.
|
||||
enabled = true
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["memory"]
|
||||
# Hard memory allowance (8 GiB).
|
||||
memory_max_bytes = 8589934592
|
||||
|
||||
[profiles.disk_medium]
|
||||
# Advertise DISK_MEDIUM only after device/rate controls validate.
|
||||
enabled = false
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["io"]
|
||||
# Windows whole-Job read bandwidth allowance (100 MiB/s).
|
||||
windows_io_read_bps = 104857600
|
||||
# Windows whole-Job write bandwidth allowance (50 MiB/s).
|
||||
windows_io_write_bps = 52428800
|
||||
# Linux cgroup read allowances; replace 8:0 with an actual delegated device.
|
||||
linux_io_read_bps = { "8:0" = 104857600 }
|
||||
# Linux cgroup write allowances; replace 8:0 with an actual delegated device.
|
||||
linux_io_write_bps = { "8:0" = 52428800 }
|
||||
|
||||
[profiles.disk_heavy]
|
||||
# Advertise DISK_HEAVY only after device/rate controls validate.
|
||||
enabled = false
|
||||
# Controls that must be applied atomically or the request is rejected.
|
||||
required_controls = ["io"]
|
||||
# Windows whole-Job read bandwidth allowance (500 MiB/s).
|
||||
windows_io_read_bps = 524288000
|
||||
# Windows whole-Job write bandwidth allowance (250 MiB/s).
|
||||
windows_io_write_bps = 262144000
|
||||
# Linux cgroup read allowances; replace 8:0 with an actual delegated device.
|
||||
linux_io_read_bps = { "8:0" = 524288000 }
|
||||
# Linux cgroup write allowances; replace 8:0 with an actual delegated device.
|
||||
linux_io_write_bps = { "8:0" = 262144000 }
|
||||
@@ -0,0 +1,109 @@
|
||||
# RVBox v1 server example. Integer sizes are bytes; durations are quoted strings.
|
||||
|
||||
[server]
|
||||
# Durable SQLite, command payload, output, audit, and incident root.
|
||||
data_dir = "/var/lib/rvbox-server"
|
||||
# HTTP listener receiving WebSocket upgrades from nginx.
|
||||
agent_listen = "127.0.0.1:6899"
|
||||
# Exact WebSocket request path accepted on agent_listen.
|
||||
agent_path = "/v1/agent"
|
||||
# Local gRPC control socket used by rvc; RVBox forces mode 0600.
|
||||
control_socket = "/run/rvbox/server.sock"
|
||||
# Grace allowed for daemon workers to finish reserved closeout on shutdown.
|
||||
shutdown_grace = "30s"
|
||||
|
||||
[json_rpc]
|
||||
# Enable the intentionally unauthenticated JSON-RPC debugging adapter.
|
||||
enabled = false
|
||||
# HTTP bind for JSON-RPC; non-loopback use emits a prominent warning.
|
||||
listen = "127.0.0.1:6900"
|
||||
|
||||
[queue]
|
||||
# Default wait for server-confirmed client acceptance; "0s" means indefinite.
|
||||
default_ttl = "15m"
|
||||
# Maximum server-side queued commands for one client before rejection.
|
||||
max_per_client = 1000
|
||||
# Maximum server-side queued commands across all clients before rejection.
|
||||
max_server = 10000
|
||||
# Initial delay after a transient client rejection before redispatch.
|
||||
retry_initial = "1s"
|
||||
# Maximum full-jitter delay after repeated transient client rejection.
|
||||
retry_max = "30s"
|
||||
|
||||
[storage]
|
||||
# Maximum rolling compressed output retained for one command (10 MiB).
|
||||
command_output_limit_bytes = 10485760
|
||||
# Maximum charged data, including metadata/script/stdin/output, per command (32 MiB).
|
||||
command_total_limit_bytes = 33554432
|
||||
# Maximum charged command data for one target client on the server (256 MiB).
|
||||
client_total_limit_bytes = 268435456
|
||||
# Maximum charged command data across the server, excluding audit/tombstones (4 GiB).
|
||||
server_total_limit_bytes = 4294967296
|
||||
# Reclaim whole terminal commands after this age; "0s" disables age rotation.
|
||||
terminal_retention = "720h"
|
||||
# Separate compressed audit plus resolved-incident history budget (100 MiB).
|
||||
audit_limit_bytes = 104857600
|
||||
# Optional audit age rotation; "0s" keeps entries until the byte budget rolls them.
|
||||
audit_retention = "0s"
|
||||
# Compact global completed-command replay ledger entry cap.
|
||||
tombstone_max_entries = 1000000
|
||||
# Per-active-command quota held for terminal and loss metadata (64 KiB).
|
||||
command_closeout_reserve_bytes = 65536
|
||||
# Reject new unreserved writes below this filesystem free space (256 MiB).
|
||||
free_space_floor_bytes = 268435456
|
||||
# Target size for a sealed append-only payload/output segment (256 KiB).
|
||||
segment_target_bytes = 262144
|
||||
# Maximum time a group-commit waits before fsync; acknowledgements wait too.
|
||||
durability_interval = "100ms"
|
||||
# SQLite busy timeout before an operation returns a transient error.
|
||||
sqlite_busy_timeout = "5s"
|
||||
# Maximum operator incident acknowledgement note encoded as UTF-8 (4 KiB).
|
||||
incident_note_max_bytes = 4096
|
||||
# Maximum protocol error/detail/reason text encoded as UTF-8 (4 KiB).
|
||||
protocol_detail_max_bytes = 4096
|
||||
|
||||
[flow]
|
||||
# Stop admitting raw server validation/persistence work at this backlog (64 MiB).
|
||||
raw_output_high_bytes = 67108864
|
||||
# Leave raw-output loss mode only below this backlog (32 MiB).
|
||||
raw_output_low_bytes = 33554432
|
||||
# Maximum assigned but unacknowledged encoded bytes per command (1 MiB).
|
||||
unacknowledged_per_command_bytes = 1048576
|
||||
# Maximum assigned but unacknowledged encoded bytes per client session (8 MiB).
|
||||
unacknowledged_per_session_bytes = 8388608
|
||||
# Deadline for an individual WebSocket data-frame write.
|
||||
write_deadline = "10s"
|
||||
|
||||
[protocol]
|
||||
# Send WebSocket Ping after this period without inbound activity.
|
||||
heartbeat_idle = "10s"
|
||||
# Close a session after this total period without inbound activity.
|
||||
liveness_timeout = "30s"
|
||||
# One-shot live client-instance collision override lifetime.
|
||||
takeover_ttl = "5m"
|
||||
# Hard decoded AgentEnvelope ceiling (1 MiB).
|
||||
max_agent_envelope_bytes = 1048576
|
||||
# Hard serialized ExecutionSpec ceiling (768 KiB).
|
||||
max_execution_spec_bytes = 786432
|
||||
# Hard uncompressed stdout/stderr chunk ceiling (64 KiB).
|
||||
max_raw_chunk_bytes = 65536
|
||||
# Hard raw uploaded-script ceiling (10 MiB).
|
||||
max_script_bytes = 10485760
|
||||
# Hard decoded local gRPC request ceiling (16 MiB).
|
||||
max_control_request_bytes = 16777216
|
||||
# Hard JSON-RPC HTTP body ceiling including base64 expansion (24 MiB).
|
||||
max_json_rpc_body_bytes = 25165824
|
||||
|
||||
[observability]
|
||||
# Loopback HTTP listener for liveness, readiness, and Prometheus metrics.
|
||||
listen = "127.0.0.1:6901"
|
||||
# Liveness route, available before asynchronous storage recovery completes.
|
||||
liveness_path = "/livez"
|
||||
# Readiness route; false while required storage scopes are unavailable.
|
||||
readiness_path = "/readyz"
|
||||
# Prometheus metrics route.
|
||||
metrics_path = "/metrics"
|
||||
# Structured logging threshold: debug, info, warn, or error.
|
||||
log_level = "info"
|
||||
# Structured log encoding: json or text.
|
||||
log_format = "json"
|
||||
File diff suppressed because it is too large
Load Diff
+145
-16
@@ -1,5 +1,9 @@
|
||||
# RVBox v1 platform and operations contract
|
||||
|
||||
Daemon configuration uses strict TOML as specified in
|
||||
[`configuration.md`](configuration.md); the annotated examples contain every
|
||||
v1 knob and default.
|
||||
|
||||
## Unix-like clients
|
||||
|
||||
The client starts `sh` or `bash` in a new session/process group. Unix signals
|
||||
@@ -8,20 +12,99 @@ orderly shutdown, or recovery after an unclean daemon failure, managed command
|
||||
groups are terminated and marked interrupted because pipe capture cannot be
|
||||
safely resumed.
|
||||
|
||||
Both command text and uploaded scripts execute from generated private files
|
||||
beneath the effective CWD using exactly the selected executable (`sh FILE` or
|
||||
`bash FILE`). No user-supplied filename becomes a filesystem path. The wrapper
|
||||
file is removed during terminal cleanup.
|
||||
|
||||
Root-process exit begins a configurable 5-second drain grace period. RVBox waits
|
||||
for the supervised tree and capture pipes, then terminates residual group/cgroup
|
||||
members, drains to EOF, and only afterward emits the terminal lifecycle event.
|
||||
If capture still cannot reach EOF, it closes the handles and emits explicit
|
||||
incomplete-output metadata first. Shell-level detachment is not a supported way
|
||||
to leave descendants running; callers use RVBox background mode instead.
|
||||
|
||||
Launch uses an internal blocked launcher rather than starting requested command
|
||||
code directly. The launcher establishes its session/process group, reports its
|
||||
identity, and waits on a private release/watchdog channel. The client durably
|
||||
records `launch_prepared`, then durably records `launch_authorized`, and only
|
||||
then sends the release token. The launcher creates the requested shell inside
|
||||
that group and remains as a non-user-code watchdog until the tree exits. The
|
||||
daemon keeps the channel open for that lifetime: EOF before authorization exits
|
||||
without execution, while EOF after release terminates the group. Once
|
||||
`launch_authorized` is durable, recovery never retries that UUID; an uncertain
|
||||
launch is marked interrupted.
|
||||
|
||||
On Linux, create a per-command cgroup v2 for supervision even when no resource
|
||||
profile was requested, whenever the daemon has a delegated writable cgroup.
|
||||
Put the blocked launcher into that cgroup before release; use
|
||||
`clone3(CLONE_INTO_CGROUP | CLONE_PIDFD)` where available, otherwise migrate the
|
||||
still-blocked launcher through `cgroup.procs`. Persist the cgroup path, PID,
|
||||
process group, `/proc/<pid>/stat` start time, and launch generation. A live
|
||||
daemon uses the pidfd where available. Recovery uses `cgroup.kill` as the primary
|
||||
tree-cleanup operation and verifies the recorded birth identity before any
|
||||
PID/process-group fallback. Without cgroup delegation it uses the generic
|
||||
watchdog/process-group fallback unless a requested profile requires cgroup
|
||||
controls, in which case acceptance fails as unsupported. It never signals a
|
||||
process based only on a persisted numeric PID or PGID.
|
||||
|
||||
Other Unix-like systems use the same launch barrier plus a watchdog control
|
||||
channel whose EOF triggers process-group termination. Recovery validates the
|
||||
platform's process-birth identity before signaling. Descendants that deliberately
|
||||
create a new session may escape this generic fallback, so complete tree cleanup
|
||||
outside Linux cgroup supervision is best-effort; the at-most-once launch
|
||||
guarantee still applies.
|
||||
|
||||
Linux diagnostics sample `/proc/<pid>` and relevant children for state, CPU,
|
||||
resident memory, I/O counters, CWD, and wait-channel information when readable.
|
||||
These values may be unavailable due to permissions, kernel configuration, or a
|
||||
short-lived process; absence is represented explicitly rather than fabricated.
|
||||
Cgroup v2 is used for requested resource profiles only when available.
|
||||
Cgroup v2 profile limits are applied only when requested; a no-profile
|
||||
supervisory cgroup imposes no resource limit.
|
||||
|
||||
## Windows clients
|
||||
|
||||
The client launches `cmd` or `powershell` in an appropriate dedicated console
|
||||
process group and assigns the root process to a per-command Job Object. Child
|
||||
processes normally join the Job Object. Job Object limits enforce requested
|
||||
profiles and `KILL_ON_JOB_CLOSE` protects against lost supervision.
|
||||
The minimum supported v1 Windows versions are Windows 10 and Windows Server
|
||||
2016. For every command, create a non-inheritable Job Object, set
|
||||
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`, and do not enable breakaway. The daemon
|
||||
starts an RVBox per-command launcher suspended with `CREATE_NEW_CONSOLE`,
|
||||
`CREATE_UNICODE_ENVIRONMENT`, and `EXTENDED_STARTUPINFO_PRESENT`; it assigns the
|
||||
launcher atomically through `PROC_THREAD_ATTRIBUTE_JOB_LIST`. Only an explicit
|
||||
standard-I/O and launcher-control handle list is inherited, and the Job handle
|
||||
is never inherited.
|
||||
|
||||
Only `SIGTERM` and `SIGKILL` are accepted. `SIGTERM` attempts `CTRL_BREAK_EVENT`
|
||||
The launcher invokes exactly the selected shell against the generated wrapper:
|
||||
`cmd.exe /D /S /C` for a `.cmd` wrapper, or `powershell.exe` with `-NoLogo`,
|
||||
`-NoProfile`, `-NonInteractive`, and `-File` for a `.ps1` wrapper. Application
|
||||
paths and argument quoting are constructed by the Windows launcher, never by
|
||||
concatenating an untrusted command line. There is no fallback between shells.
|
||||
|
||||
Persist and flush `launch_prepared` with the launcher PID,
|
||||
`GetProcessTimes` creation `FILETIME`, and launch generation. Persist and flush
|
||||
`launch_authorized` before calling `ResumeThread`. The launcher then starts the
|
||||
requested `cmd` or `powershell` suspended in its console with
|
||||
`CREATE_NEW_PROCESS_GROUP`, reports the shell PID/group through the private
|
||||
control channel, connects the allowlisted pipes, and resumes it. This two-step
|
||||
shape is required because `CREATE_NEW_PROCESS_GROUP` is ignored when combined
|
||||
with `CREATE_NEW_CONSOLE`, and console control events reach only groups sharing
|
||||
the caller's console. Failure after authorization terminates the Job and is
|
||||
reported as interrupted; it never redispatches the UUID.
|
||||
|
||||
The launcher remains the in-console signal proxy and calls
|
||||
`GenerateConsoleCtrlEvent(CTRL_BREAK_EVENT, shell_group_id)` on request. The
|
||||
daemon retains the sole Job handle, so an unclean daemon exit closes the last
|
||||
handle and terminates the launcher, shell, and descendants. Recovery never
|
||||
kills by persisted PID alone; the PID/creation-time tuple is diagnostic evidence
|
||||
for PID reuse or cleanup anomalies. Child processes normally join the Job.
|
||||
Job Object limits enforce requested profiles and `KILL_ON_JOB_CLOSE` protects
|
||||
against lost supervision.
|
||||
|
||||
Root-process exit begins the same drain grace period. Completion waits for the
|
||||
Job Object to reach zero active processes; after the grace period RVBox
|
||||
terminates the Job, drains its capture handles, records any incomplete-output
|
||||
marker, and emits the terminal lifecycle event last.
|
||||
|
||||
Only `TERM`/`SIGTERM` and `KILL`/`SIGKILL` are accepted. `TERM` attempts `CTRL_BREAK_EVENT`
|
||||
and waits 10 seconds, then calls Job Object termination if the job persists;
|
||||
`SIGKILL` calls Job Object termination immediately. A console signal is
|
||||
best-effort, so callers receive an explicit escalation result. Windows status
|
||||
@@ -30,12 +113,36 @@ diagnostics such as an I/O wait channel.
|
||||
|
||||
## Storage and recovery
|
||||
|
||||
SQLite runs in WAL mode with integrity checking on startup. Output segments are
|
||||
written atomically, fsynced according to the configured durability interval, and
|
||||
indexed only after successful durable append. Startup scans/repairs incomplete
|
||||
tail records before accepting control requests. Segment compression is Zstandard;
|
||||
limits always measure stored compressed bytes, while clients expose raw byte
|
||||
counts separately.
|
||||
SQLite runs in WAL mode. Every append-only segment has a SQLite-owned
|
||||
`committed_end_offset`. The writer validates and appends records, syncs the file
|
||||
(grouped by the configured durability interval), and only then commits event
|
||||
metadata plus the new offset in SQLite. An acknowledgement waits for both
|
||||
steps. Therefore a crash can leave an uncommitted file tail, but cannot validly
|
||||
acknowledge metadata whose bytes were not durable.
|
||||
|
||||
Startup acquires the instance lock and binds liveness/diagnostic endpoints, then
|
||||
runs integrity and segment recovery asynchronously. A file longer than its
|
||||
committed offset is safely truncated to that offset. A file shorter than the
|
||||
offset, a checksum failure inside the committed range, or corrupt essential
|
||||
metadata creates a durable scoped storage incident; affected output is marked
|
||||
truncated/incomplete and affected active commands are interrupted when their
|
||||
essential state cannot be trusted. Healthy scopes remain usable. Readiness is
|
||||
false and mutations requiring an unrecovered or dirty scope return `UNAVAILABLE`,
|
||||
but process startup, liveness, incident inspection, and unaffected work do not
|
||||
wait for a full-store scan.
|
||||
|
||||
Safe repairs are attempted automatically and can also be requested online with
|
||||
`rvc storage repair`. Irrecoverable loss stays dirty until explicitly accepted
|
||||
with `rvc storage acknowledge`; an offline server has equivalent
|
||||
`rvbox-server repair --data-dir ...` repair/list/acknowledge operations. Client
|
||||
spool recovery follows the same committed-offset rule and exposes equivalent
|
||||
offline `rvbox repair --state-dir ...` operations and local health diagnostics.
|
||||
Resolving an incident clears derived dirty health but never erases the incident
|
||||
or audit history as part of resolution. Unresolved compact incident records are
|
||||
non-evictable; resolved incident/audit history follows the separate 100 MiB
|
||||
rotation. Segment compression is Zstandard;
|
||||
limits measure stored compressed bytes, while raw byte counts are reported
|
||||
separately.
|
||||
|
||||
The system must reserve headroom before writes and use transactional metadata
|
||||
updates. Storage-full, permission, and corruption failures are surfaced as
|
||||
@@ -43,6 +150,12 @@ structured server/client health states and audit events. They must isolate the
|
||||
affected command/session, reject work when needed, and keep the daemon's
|
||||
heartbeat/control loops alive.
|
||||
|
||||
V1 state is plaintext at rest, including command text, scripts, environment
|
||||
override values, stdin, and output. Private directory/file modes and dedicated
|
||||
daemon accounts are deployment hygiene, not an application-level encryption
|
||||
guarantee. Backups copy the same plaintext sensitivity. Encryption and external
|
||||
key management are future-version work.
|
||||
|
||||
## Metrics, logging, and safe defaults
|
||||
|
||||
Both daemons should emit structured logs and metrics for session transitions,
|
||||
@@ -52,10 +165,26 @@ and protocol violations. Never emit stdin or raw output in normal daemon logs.
|
||||
|
||||
Recommended configuration defaults are: 10-second heartbeat idle period,
|
||||
30-second liveness timeout, 1–60-second full-jitter reconnect backoff,
|
||||
60-second stable-session reset, 16 running/100 queued commands per client,
|
||||
10 MiB per-command compressed window, 50 MiB per-client active spool and server
|
||||
history, 1 GiB server history, 64 KiB uncompressed stream chunk, 1 MiB decoded
|
||||
envelope, and 10 MiB script maximum.
|
||||
60-second stable-session reset, 5-minute one-shot live-conflict takeover grant,
|
||||
16 running/100 queued commands per client, 1,000 server-queued commands per
|
||||
target client and 10,000 server-wide,
|
||||
15-minute queue TTL, 10 MiB per-command output window, 32 MiB total per command,
|
||||
256 MiB per client/client daemon, 4 GiB server-wide command storage, 30-day
|
||||
terminal retention, 100 MiB audit storage with quota-only rotation by default,
|
||||
one million compact command tombstones, 64 KiB uncompressed stream chunk, 1 MiB
|
||||
decoded agent envelope,
|
||||
1 MiB/256 KiB per-command raw-output high/low watermarks, 8 MiB/4 MiB per-client
|
||||
watermarks, 64 MiB/32 MiB server-wide watermarks, 1 MiB per-command and 8 MiB
|
||||
per-session unacknowledged send windows, and 10 MiB raw script maximum. Control
|
||||
gRPC accepts at most 16 MiB decoded requests; the JSON-RPC adapter accepts at
|
||||
most 24 MiB HTTP bodies to allow protobuf JSON's base64 expansion while
|
||||
retaining the same decoded field limits. Each active command reserves 64 KiB
|
||||
within its quota for terminal/loss closeout metadata; protocol detail/reason and
|
||||
incident-note text fields are individually limited to 4 KiB. Default emergency
|
||||
filesystem free-space floors are 256 MiB on the server and 64 MiB on a client;
|
||||
crossing one rejects new unreserved allocations even if the logical quota has
|
||||
headroom. Already-reserved terminal/loss closeout remains writable while bytes
|
||||
physically remain.
|
||||
|
||||
These bounds protect RVBox's own loops; they cannot make arbitrary child
|
||||
commands harmless when no resource profile is requested. Operators should
|
||||
|
||||
+101
-31
@@ -14,8 +14,10 @@ sizes before allocation/decompression.
|
||||
|
||||
The protobuf package is `rvbox.v1`. Registration negotiates a major/minor
|
||||
protocol range: incompatible majors are rejected; the highest shared minor is
|
||||
chosen. New fields are append-only. A peer ignores unknown optional fields but
|
||||
must respond with `PROTOCOL_ERROR` to an unknown required envelope feature.
|
||||
chosen. New fields are append-only. A peer ignores unknown optional fields. A
|
||||
new required behavior requires a negotiated minor-version change; an envelope
|
||||
with no recognized payload is a `PROTOCOL_ERROR` rather than an implicit
|
||||
“required unknown field” mechanism that protobuf cannot represent.
|
||||
|
||||
## Heartbeat and reconnect
|
||||
|
||||
@@ -33,11 +35,19 @@ reconciles only commands the server still considers non-terminal.
|
||||
|
||||
## Session fencing
|
||||
|
||||
`ClientHello` starts registration. `ServerWelcome` gives the selected version,
|
||||
random `session_id`, and `session_generation`. Except `ClientHello`, all
|
||||
envelopes carry those values. A new accepted registration fences and disconnects
|
||||
the prior one for the same client ID. The server accepts messages only from the
|
||||
current generation; command dispatches also identify their intended generation.
|
||||
`ClientHello` starts registration and includes a random `client_instance_id`
|
||||
generated once in the client state directory and retained across daemon restarts
|
||||
and reconnects. `ServerWelcome` gives the selected version, random `session_id`,
|
||||
and `session_generation`. Except `ClientHello`, all envelopes carry those
|
||||
values. A new accepted registration from the same instance fences and
|
||||
disconnects its prior session. A different instance claiming the same live
|
||||
client ID is rejected and recorded as the current pending claim unless an
|
||||
operator explicitly authorized that exact instance. The one-shot authorization
|
||||
expires after 5 minutes by default and is consumed by the matching reconnect.
|
||||
When no session for that client ID is live, a new instance is accepted normally.
|
||||
This is collision protection, not peer authentication. The
|
||||
server accepts messages only from the current generation; command dispatches
|
||||
also identify their intended generation.
|
||||
|
||||
## Reliable work flows
|
||||
|
||||
@@ -45,53 +55,113 @@ current generation; command dispatches also identify their intended generation.
|
||||
|
||||
The server persistently creates a command before dispatching `CommandDispatch`.
|
||||
It redelivers until it receives `CommandAccepted`. Clients durably deduplicate
|
||||
on `issue_uuid`. A client whose queue is full sends a capacity rejection.
|
||||
on `issue_uuid`. A client whose queue is full sends a transient capacity
|
||||
rejection, which the server requeues with backoff. A permanent validation or
|
||||
unsupported-platform rejection makes the server command terminal `rejected`;
|
||||
it is not retried or mislabeled as a launched-process failure.
|
||||
An exact UUID/request-hash hit in the compact tombstone ledger returns
|
||||
`CODE_ALREADY_EXECUTED`; the server suppresses dispatch and reconciles its stale
|
||||
state instead of representing the prior execution as a new rejection.
|
||||
|
||||
Following registration the server sends `ReconcileRequest` only for its
|
||||
non-terminal commands for that client. The client replies with a
|
||||
`ReconcileSnapshot` per requested known command, or `unknown_to_client`. It
|
||||
continues normal event retransmission from the server's last acknowledged event
|
||||
sequence. Terminal history already confirmed by the server is deliberately
|
||||
excluded.
|
||||
Following registration the server sends `ReconcileRequest` containing its
|
||||
non-terminal commands for that client. The client replies once with a complete,
|
||||
bounded `ReconcileSnapshot` of every command it still retains, including
|
||||
terminal-but-unacknowledged records and matching requested tombstones, then
|
||||
continues normal retransmission from the server's last acknowledged event
|
||||
sequence. The server does not dispatch new work until this snapshot is complete.
|
||||
For healthy client storage, absence means the client never durably accepted the
|
||||
UUID: server `queued`/`dispatched` work returns to `queued`, while absence of an
|
||||
`accepted`/`running` record is an invariant failure and becomes interrupted.
|
||||
A matching UUID/request-hash tombstone always suppresses replay. A client-known
|
||||
server-missing active command, or a contradiction with server-confirmed
|
||||
terminal history, is terminated locally after the server returns a durable
|
||||
`ReconcileResult` and creates a recovery incident rather than inventing server
|
||||
state. That result also identifies client-retained terminal records the server
|
||||
has already stored or deliberately tombstoned, allowing the client to discard
|
||||
them even when its last `EventAck` was lost. The result is idempotent and is
|
||||
written before any new dispatch on that session. A client
|
||||
with an unresolved essential-store incident must not claim a complete snapshot
|
||||
or accept work until repaired/acknowledged.
|
||||
|
||||
### Output and events
|
||||
|
||||
Client execution events use an increasing `event_seq`; retries reuse the same
|
||||
sequence and content. The server durably writes an event before sending
|
||||
`EventAck`. `EventAck` is cumulative through a sequence number. Output is
|
||||
Zstandard compressed with an explicit original-size field. Server storage can
|
||||
reuse the validated compressed bytes.
|
||||
Client execution entries first have a durable local order. They receive an
|
||||
increasing wire `event_seq` durably when admitted to the bounded send window;
|
||||
once assigned, the sequence/content is pinned until acknowledged and retries
|
||||
reuse it exactly. The server durably writes an event before sending `EventAck`.
|
||||
`EventAck` is cumulative through a sequence number. Output is Zstandard
|
||||
compressed with an explicit original-size field. Server storage can reuse the
|
||||
validated compressed bytes.
|
||||
|
||||
When rolling output removes old retained segments, the server writes
|
||||
`OutputTruncation` metadata containing the removed event/byte ranges. Queries
|
||||
must show that marker rather than silently presenting an apparently complete
|
||||
stream. An offline client that reaches a cap sends `ClientOutputTruncated`
|
||||
before replaying its retained tail on reconnect. The server records it, permits
|
||||
the named event-sequence gap, and exposes it in output queries. Truncation
|
||||
metadata is not an execution event and does not consume an `event_seq`.
|
||||
stream. Assigned but unacknowledged events occupy the pinned 1 MiB per-command
|
||||
send window and are not evicted. An offline client that reaches a cap replaces
|
||||
one or more still-unsequenced output runs in its durable local order with
|
||||
`OutputTruncation`; when admitted to the send window the marker receives the
|
||||
next normal `event_seq`. Its event-range fields are absent because the discarded
|
||||
bytes never had wire sequences. The normal cumulative `EventAck` acknowledges
|
||||
the marker. Server-created retention markers include their removed event range
|
||||
and remain query metadata because the server cannot allocate client sequences.
|
||||
|
||||
If a capture pipe cannot be drained to a provably complete EOF, the client emits
|
||||
a sequenced `OutputIncomplete` event before the terminal lifecycle event. It
|
||||
does not invent a missing sequence or byte count for bytes it never observed.
|
||||
|
||||
### Stdin and signals
|
||||
|
||||
`StdinWrite` is binary-safe and ordered by `write_seq`; the client durably
|
||||
deduplicates it and returns `StdinAck`. `append_newline` is true by default in
|
||||
the CLI but explicit on the wire. `CloseStdin` is a separate idempotent action.
|
||||
Signal and cancellation requests carry a command revision to settle start/kill
|
||||
races.
|
||||
deduplicates it and returns a sequenced stdin-acknowledgement command event.
|
||||
`append_newline` is true by default in the CLI but explicit on the wire.
|
||||
`CloseStdin` is a separate idempotent action. Signal and cancellation requests
|
||||
carry a command revision to settle start/kill races; the resulting lifecycle or
|
||||
signal-result event echoes that revision. Control-plane `request_id` values stay
|
||||
at the server and map to the assigned revision rather than crossing the agent
|
||||
protocol.
|
||||
|
||||
### Scripts
|
||||
|
||||
For a script command, dispatch first contains a `ScriptDescriptor`; then the
|
||||
server sends `ScriptChunk` messages and a commit. The client checks offset,
|
||||
chunk order, full length, and SHA-256 before it reports upload complete or
|
||||
launches the process. A retransmitted chunk is idempotent by offset/content.
|
||||
launches the process. `ScriptUploadStatus.received_bytes` is a cumulative
|
||||
durable progress acknowledgement: the server sends at most the command's
|
||||
unacknowledged window, waits for progress, and resumes from that offset. A
|
||||
retransmitted chunk is idempotent by offset/content.
|
||||
|
||||
Acceptance is durable queue admission, not proof that every later preparation
|
||||
step will succeed. A permanent script checksum/write error or failure to prepare
|
||||
the requested executable emits terminal `rejected` before any launch
|
||||
authorization. User cancellation in that interval emits `cancelled`; an
|
||||
uncertain crash window emits `interrupted`. None is mislabeled as process
|
||||
`failed`.
|
||||
|
||||
## Flow control and failure containment
|
||||
|
||||
No receive loop runs an executor, database write, decompressor, or slow socket
|
||||
operation inline. Each side has bounded staging queues. Durable spools are the
|
||||
source of truth and are charged to the 10 MiB/50 MiB/1 GiB compressed retention
|
||||
budgets described in the architecture document. A full staging queue pauses the
|
||||
related read/dispatch path and drains from disk; it never grows without bound.
|
||||
source of truth and are charged to the tiered unified command-storage budgets
|
||||
described in the architecture document. Senders keep a bounded unacknowledged
|
||||
window of encoded wire bytes (defaults: 1 MiB per command and 8 MiB per client
|
||||
session); unsent data
|
||||
waits durably. WebSocket Ping/Pong/Close plus fencing and protocol errors use a
|
||||
reserved priority lane and a dedicated writer with bounded data-frame size and
|
||||
write deadlines. A full essential ingress queue does not stall the sole
|
||||
WebSocket reader indefinitely:
|
||||
the receiver closes without acknowledgement and the durable sender retries
|
||||
after jittered reconnect. Droppable output uses the sequenced overload-loss
|
||||
path instead.
|
||||
|
||||
Before compression, raw output uses the architecture's high/low watermarks.
|
||||
Once a client high watermark is crossed, still-unsequenced droppable chunks are
|
||||
summarized in durable local order instead of consuming unbounded compression
|
||||
work. A server at its ingress high watermark does not acknowledge or discard an
|
||||
already sequenced event; it closes without acknowledgement if its bounded queue
|
||||
cannot admit the frame, and the client retries from durable spool. Fair
|
||||
scheduling prevents one verbose command from monopolizing workers. Reserved
|
||||
metadata capacity remains
|
||||
available to close each gap and emit the final lifecycle event.
|
||||
|
||||
Malformed protobuf, over-size payload, invalid compressed data, impossible
|
||||
sequence, bad session token, or protocol-version violation yields a structured
|
||||
|
||||
+39
-31
@@ -2,35 +2,34 @@ syntax = "proto3";
|
||||
|
||||
package rvbox.v1;
|
||||
|
||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||
|
||||
import "google/protobuf/timestamp.proto";
|
||||
import "rvbox/v1/common.proto";
|
||||
|
||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||
|
||||
// One AgentEnvelope is carried in one binary WebSocket message.
|
||||
message AgentEnvelope {
|
||||
// Empty only for ClientHello. All other envelopes are fenced to a session.
|
||||
string session_id = 1;
|
||||
uint64 session_generation = 2;
|
||||
string message_id = 3;
|
||||
|
||||
oneof payload {
|
||||
ClientHello client_hello = 10;
|
||||
ServerWelcome server_welcome = 11;
|
||||
CommandDispatch command_dispatch = 12;
|
||||
CommandAccepted command_accepted = 13;
|
||||
CommandEvent command_event = 14;
|
||||
EventAck event_ack = 15;
|
||||
StdinWrite stdin_write = 16;
|
||||
CloseStdin close_stdin = 17;
|
||||
SignalCommand signal_command = 18;
|
||||
ScriptChunk script_chunk = 19;
|
||||
ScriptCommit script_commit = 20;
|
||||
ClientCapacity client_capacity = 21;
|
||||
ReconcileRequest reconcile_request = 22;
|
||||
ReconcileSnapshot reconcile_snapshot = 23;
|
||||
AgentError error = 24;
|
||||
ClientOutputTruncated client_output_truncated = 25;
|
||||
ClientHello client_hello = 3;
|
||||
ServerWelcome server_welcome = 4;
|
||||
CommandDispatch command_dispatch = 5;
|
||||
CommandAccepted command_accepted = 6;
|
||||
CommandEvent command_event = 7;
|
||||
EventAck event_ack = 8;
|
||||
StdinWrite stdin_write = 9;
|
||||
CloseStdin close_stdin = 10;
|
||||
SignalCommand signal_command = 11;
|
||||
ScriptChunk script_chunk = 12;
|
||||
ScriptCommit script_commit = 13;
|
||||
ClientCapacity client_capacity = 14;
|
||||
ReconcileRequest reconcile_request = 15;
|
||||
ReconcileSnapshot reconcile_snapshot = 16;
|
||||
AgentError error = 17;
|
||||
ReconcileResult reconcile_result = 18;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -43,7 +42,8 @@ message ClientHello {
|
||||
string architecture = 5;
|
||||
string daemon_cwd = 6;
|
||||
repeated ShellType supported_shells = 7;
|
||||
string reconnect_uuid = 8;
|
||||
// Generated once and persisted in the client state directory.
|
||||
string client_instance_id = 8;
|
||||
uint32 max_running_commands = 9;
|
||||
uint32 max_queued_commands = 10;
|
||||
google.protobuf.Timestamp sent_at = 11;
|
||||
@@ -61,6 +61,7 @@ message CommandDispatch {
|
||||
uint64 target_session_generation = 3;
|
||||
google.protobuf.Timestamp issue_time = 4;
|
||||
ExecutionSpec spec = 5;
|
||||
google.protobuf.Timestamp queue_expiry_time = 6;
|
||||
}
|
||||
|
||||
message CommandAccepted {
|
||||
@@ -122,23 +123,30 @@ message ReconcileTarget {
|
||||
string issue_uuid = 1;
|
||||
uint64 last_server_event_seq = 2;
|
||||
uint64 command_revision = 3;
|
||||
bytes immutable_request_sha256 = 4;
|
||||
}
|
||||
|
||||
// Sent only for requests in ReconcileRequest; server-confirmed terminal history
|
||||
// is intentionally not requested.
|
||||
message ReconcileSnapshot {
|
||||
string issue_uuid = 1;
|
||||
bool known_to_client = 2;
|
||||
CommandLifecycle lifecycle = 3;
|
||||
uint64 last_client_event_seq = 4;
|
||||
uint64 command_revision = 5;
|
||||
// The complete set still retained by the client, including non-terminal and
|
||||
// terminal-but-unacknowledged commands.
|
||||
repeated ReconcileCommandState retained_commands = 1;
|
||||
}
|
||||
|
||||
// Sent before retained replay data when an offline client had to rotate
|
||||
// unacknowledged output to remain within its hard spool limits.
|
||||
message ClientOutputTruncated {
|
||||
message ReconcileCommandState {
|
||||
string issue_uuid = 1;
|
||||
OutputTruncation truncation = 2;
|
||||
CommandLifecycle lifecycle = 2;
|
||||
uint64 last_client_event_seq = 3;
|
||||
uint64 command_revision = 4;
|
||||
// True only for a matching UUID/request hash in the compact tombstone ledger.
|
||||
bool tombstoned = 5;
|
||||
bytes immutable_request_sha256 = 6;
|
||||
}
|
||||
|
||||
// Sent after the server durably compares a complete snapshot and before it
|
||||
// dispatches new work on the session. Repetition is idempotent.
|
||||
message ReconcileResult {
|
||||
repeated string terminate_local_issue_uuids = 1;
|
||||
repeated string discard_local_terminal_issue_uuids = 2;
|
||||
}
|
||||
|
||||
message AgentError {
|
||||
|
||||
@@ -2,11 +2,11 @@ syntax = "proto3";
|
||||
|
||||
package rvbox.v1;
|
||||
|
||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||
|
||||
import "google/protobuf/duration.proto";
|
||||
import "google/protobuf/timestamp.proto";
|
||||
|
||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||
|
||||
// An inclusive protocol-version range advertised during registration.
|
||||
message ProtocolRange {
|
||||
uint32 major = 1;
|
||||
@@ -52,6 +52,8 @@ enum CommandLifecycle {
|
||||
COMMAND_TERMINATED = 7;
|
||||
COMMAND_CANCELLED = 8;
|
||||
COMMAND_INTERRUPTED = 9;
|
||||
COMMAND_EXPIRED = 10;
|
||||
COMMAND_REJECTED = 11;
|
||||
}
|
||||
|
||||
enum StreamKind {
|
||||
@@ -109,16 +111,23 @@ message CommandRecord {
|
||||
ExecutionSpec spec = 5;
|
||||
CommandLifecycle lifecycle = 6;
|
||||
uint64 last_event_seq = 7;
|
||||
int32 exit_code = 8;
|
||||
optional int32 exit_code = 8;
|
||||
google.protobuf.Timestamp terminal_time = 9;
|
||||
bool output_truncated = 10;
|
||||
uint64 retained_compressed_bytes = 11;
|
||||
google.protobuf.Timestamp queue_expiry_time = 12;
|
||||
bool output_incomplete = 13;
|
||||
// Present when lifecycle is COMMAND_REJECTED.
|
||||
ControlError rejection = 14;
|
||||
uint64 command_revision = 15;
|
||||
bool late_after_expiry = 16;
|
||||
}
|
||||
|
||||
message LifecycleChange {
|
||||
CommandLifecycle lifecycle = 1;
|
||||
int32 exit_code = 2;
|
||||
optional int32 exit_code = 2;
|
||||
string detail = 3;
|
||||
uint64 command_revision = 4;
|
||||
}
|
||||
|
||||
// data is compressed according to compression. uncompressed_size is mandatory
|
||||
@@ -132,15 +141,16 @@ message OutputChunk {
|
||||
}
|
||||
|
||||
message ResourceSnapshot {
|
||||
uint64 resident_memory_bytes = 1;
|
||||
uint64 virtual_memory_bytes = 2;
|
||||
optional uint64 resident_memory_bytes = 1;
|
||||
optional uint64 virtual_memory_bytes = 2;
|
||||
google.protobuf.Duration cpu_time = 3;
|
||||
uint64 read_bytes = 4;
|
||||
uint64 write_bytes = 5;
|
||||
string process_state = 6;
|
||||
string wait_reason = 7;
|
||||
optional uint64 read_bytes = 4;
|
||||
optional uint64 write_bytes = 5;
|
||||
optional string process_state = 6;
|
||||
optional string wait_reason = 7;
|
||||
bool suspected_hung = 8;
|
||||
string diagnostic_detail = 9;
|
||||
optional string current_cwd = 10;
|
||||
}
|
||||
|
||||
enum OutputTruncationSource {
|
||||
@@ -148,19 +158,31 @@ enum OutputTruncationSource {
|
||||
OUTPUT_TRUNCATION_SOURCE_CLIENT_SPOOL = 1;
|
||||
OUTPUT_TRUNCATION_SOURCE_SERVER_COMMAND_WINDOW = 2;
|
||||
OUTPUT_TRUNCATION_SOURCE_SERVER_CLIENT_CAP = 3;
|
||||
OUTPUT_TRUNCATION_SOURCE_CLIENT_OVERLOAD = 4;
|
||||
OUTPUT_TRUNCATION_SOURCE_SERVER_GLOBAL_CAP = 5;
|
||||
}
|
||||
|
||||
// Persistent query metadata for a missing contiguous event range. It is not a
|
||||
// CommandEvent and therefore does not consume an event_seq.
|
||||
// Metadata for missing output. Server-side retention supplies the optional
|
||||
// event range. Client output discarded before wire sequence assignment omits
|
||||
// the range and carries this as a normally sequenced CommandEvent.
|
||||
message OutputTruncation {
|
||||
uint64 first_removed_event_seq = 1;
|
||||
uint64 last_removed_event_seq = 2;
|
||||
uint64 removed_compressed_bytes = 3;
|
||||
optional uint64 first_removed_event_seq = 1;
|
||||
optional uint64 last_removed_event_seq = 2;
|
||||
// Absent when bytes were discarded before compression.
|
||||
optional uint64 removed_compressed_bytes = 3;
|
||||
uint64 removed_uncompressed_bytes = 4;
|
||||
string reason = 5;
|
||||
OutputTruncationSource source = 6;
|
||||
}
|
||||
|
||||
// Indicates that capture ended without a provably complete byte stream. It is
|
||||
// sequenced before the terminal lifecycle event but does not claim an invented
|
||||
// event or byte range.
|
||||
message OutputIncomplete {
|
||||
repeated StreamKind streams = 1;
|
||||
string reason = 2;
|
||||
}
|
||||
|
||||
message StdinAcknowledgement {
|
||||
uint64 write_seq = 1;
|
||||
bool stdin_closed = 2;
|
||||
@@ -173,6 +195,7 @@ message SignalResult {
|
||||
bool graceful_delivery_attempted = 3;
|
||||
bool forced_termination_used = 4;
|
||||
string detail = 5;
|
||||
uint64 command_revision = 6;
|
||||
}
|
||||
|
||||
message ScriptUploadStatus {
|
||||
@@ -188,12 +211,14 @@ message CommandEvent {
|
||||
uint64 event_seq = 2;
|
||||
google.protobuf.Timestamp observed_at = 3;
|
||||
oneof payload {
|
||||
LifecycleChange lifecycle = 10;
|
||||
OutputChunk output = 11;
|
||||
ResourceSnapshot resource = 12;
|
||||
StdinAcknowledgement stdin_ack = 13;
|
||||
SignalResult signal_result = 14;
|
||||
ScriptUploadStatus script_status = 15;
|
||||
LifecycleChange lifecycle = 4;
|
||||
OutputChunk output = 5;
|
||||
ResourceSnapshot resource = 6;
|
||||
StdinAcknowledgement stdin_ack = 7;
|
||||
SignalResult signal_result = 8;
|
||||
ScriptUploadStatus script_status = 9;
|
||||
OutputTruncation output_truncation = 10;
|
||||
OutputIncomplete output_incomplete = 11;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -209,6 +234,7 @@ message ControlError {
|
||||
PROTOCOL_ERROR = 7;
|
||||
TRANSIENT = 8;
|
||||
INTERNAL = 9;
|
||||
CODE_ALREADY_EXECUTED = 10;
|
||||
}
|
||||
Code code = 1;
|
||||
string message = 2;
|
||||
|
||||
+112
-15
@@ -2,22 +2,27 @@ syntax = "proto3";
|
||||
|
||||
package rvbox.v1;
|
||||
|
||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||
|
||||
import "google/protobuf/duration.proto";
|
||||
import "google/protobuf/timestamp.proto";
|
||||
import "rvbox/v1/common.proto";
|
||||
|
||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||
|
||||
service Control {
|
||||
rpc ListClients(ListClientsRequest) returns (ListClientsResponse);
|
||||
rpc GetClient(GetClientRequest) returns (GetClientResponse);
|
||||
rpc ListCommands(ListCommandsRequest) returns (ListCommandsResponse);
|
||||
rpc GetCommand(GetCommandRequest) returns (GetCommandResponse);
|
||||
rpc RunCommand(RunCommandRequest) returns (RunCommandResponse);
|
||||
rpc RunCommandAndFollow(RunCommandRequest) returns (stream CommandEvent);
|
||||
rpc FollowCommand(FollowCommandRequest) returns (stream CommandEvent);
|
||||
rpc FollowCommand(FollowCommandRequest) returns (stream FollowCommandResponse);
|
||||
rpc AppendStdin(AppendStdinRequest) returns (AppendStdinResponse);
|
||||
rpc CloseStdin(CloseStdinRequest) returns (CloseStdinResponse);
|
||||
rpc SignalCommand(ControlSignalCommandRequest) returns (ControlSignalCommandResponse);
|
||||
rpc GetOutput(GetOutputRequest) returns (GetOutputResponse);
|
||||
rpc ListStorageIncidents(ListStorageIncidentsRequest) returns (ListStorageIncidentsResponse);
|
||||
rpc RepairStorageIncident(RepairStorageIncidentRequest) returns (RepairStorageIncidentResponse);
|
||||
rpc AcknowledgeStorageIncident(AcknowledgeStorageIncidentRequest) returns (AcknowledgeStorageIncidentResponse);
|
||||
rpc AuthorizeClientTakeover(AuthorizeClientTakeoverRequest) returns (AuthorizeClientTakeoverResponse);
|
||||
}
|
||||
|
||||
message ClientSummary {
|
||||
@@ -32,6 +37,10 @@ message ClientSummary {
|
||||
string daemon_version = 9;
|
||||
string daemon_cwd = 10;
|
||||
repeated ShellType supported_shells = 11;
|
||||
string client_instance_id = 12;
|
||||
// Most recently rejected different instance while this client is live.
|
||||
string pending_instance_id = 13;
|
||||
google.protobuf.Timestamp pending_instance_seen_at = 14;
|
||||
}
|
||||
|
||||
message ListClientsRequest {
|
||||
@@ -50,7 +59,6 @@ message GetClientRequest {
|
||||
|
||||
message GetClientResponse {
|
||||
ClientSummary client = 1;
|
||||
ControlError error = 2;
|
||||
}
|
||||
|
||||
message ListCommandsRequest {
|
||||
@@ -63,7 +71,6 @@ message ListCommandsRequest {
|
||||
message ListCommandsResponse {
|
||||
repeated CommandRecord commands = 1;
|
||||
string next_page_token = 2;
|
||||
ControlError error = 3;
|
||||
}
|
||||
|
||||
message GetCommandRequest {
|
||||
@@ -74,7 +81,6 @@ message GetCommandRequest {
|
||||
message GetCommandResponse {
|
||||
CommandRecord command = 1;
|
||||
ResourceSnapshot latest_resource = 2;
|
||||
ControlError error = 3;
|
||||
}
|
||||
|
||||
// For script execution, script_content contains the bytes whose descriptor is
|
||||
@@ -83,12 +89,15 @@ message RunCommandRequest {
|
||||
string target_client_id = 1;
|
||||
ExecutionSpec spec = 2;
|
||||
bytes script_content = 3;
|
||||
// Becomes issue_uuid when supplied; generated by the server when omitted.
|
||||
string request_id = 4;
|
||||
// Omitted uses the server default; zero explicitly requests no expiry.
|
||||
google.protobuf.Duration queue_ttl = 5;
|
||||
}
|
||||
|
||||
message RunCommandResponse {
|
||||
string issue_uuid = 1;
|
||||
CommandLifecycle lifecycle = 2;
|
||||
ControlError error = 3;
|
||||
}
|
||||
|
||||
message FollowCommandRequest {
|
||||
@@ -98,37 +107,47 @@ message FollowCommandRequest {
|
||||
bool include_existing = 4;
|
||||
}
|
||||
|
||||
message FollowCommandResponse {
|
||||
oneof item {
|
||||
CommandEvent event = 1;
|
||||
// Server retention metadata; does not advance the client event cursor.
|
||||
RetentionTruncation retention_truncation = 2;
|
||||
}
|
||||
// Set for a client event; server-created retention uses its own recorded_at.
|
||||
google.protobuf.Timestamp server_receipt_time = 3;
|
||||
}
|
||||
|
||||
message AppendStdinRequest {
|
||||
string client_id = 1;
|
||||
string issue_uuid = 2;
|
||||
bytes data = 3;
|
||||
bool append_newline = 4;
|
||||
string request_id = 5;
|
||||
}
|
||||
|
||||
message AppendStdinResponse {
|
||||
uint64 write_seq = 1;
|
||||
ControlError error = 2;
|
||||
}
|
||||
|
||||
message CloseStdinRequest {
|
||||
string client_id = 1;
|
||||
string issue_uuid = 2;
|
||||
string request_id = 3;
|
||||
}
|
||||
|
||||
message CloseStdinResponse {
|
||||
uint64 write_seq = 1;
|
||||
ControlError error = 2;
|
||||
}
|
||||
|
||||
message ControlSignalCommandRequest {
|
||||
string client_id = 1;
|
||||
string issue_uuid = 2;
|
||||
SignalKind signal = 3;
|
||||
string request_id = 4;
|
||||
}
|
||||
|
||||
message ControlSignalCommandResponse {
|
||||
uint64 command_revision = 1;
|
||||
ControlError error = 2;
|
||||
}
|
||||
|
||||
message GetOutputRequest {
|
||||
@@ -136,13 +155,91 @@ message GetOutputRequest {
|
||||
string issue_uuid = 2;
|
||||
repeated StreamKind streams = 3;
|
||||
uint64 after_event_seq = 4;
|
||||
uint32 page_size = 5;
|
||||
uint64 max_bytes = 5;
|
||||
string page_token = 6;
|
||||
}
|
||||
|
||||
// An uncompressed slice of one persisted output event. Opaque pagination may
|
||||
// resume within an event; event_byte_offset identifies that position.
|
||||
message OutputSlice {
|
||||
uint64 event_seq = 1;
|
||||
google.protobuf.Timestamp observed_at = 2;
|
||||
StreamKind stream = 3;
|
||||
bytes data = 4;
|
||||
uint64 event_byte_offset = 5;
|
||||
bool end_of_event = 6;
|
||||
google.protobuf.Timestamp server_receipt_time = 7;
|
||||
}
|
||||
|
||||
message RetentionTruncation {
|
||||
OutputTruncation truncation = 1;
|
||||
google.protobuf.Timestamp server_recorded_at = 2;
|
||||
}
|
||||
|
||||
message GetOutputResponse {
|
||||
repeated CommandEvent events = 1;
|
||||
repeated OutputSlice output = 1;
|
||||
string next_page_token = 2;
|
||||
bool output_truncated = 3;
|
||||
ControlError error = 4;
|
||||
repeated OutputTruncation truncations = 5;
|
||||
repeated RetentionTruncation truncations = 4;
|
||||
OutputIncomplete incomplete = 5;
|
||||
}
|
||||
|
||||
enum StorageIncidentState {
|
||||
STORAGE_INCIDENT_STATE_UNSPECIFIED = 0;
|
||||
STORAGE_INCIDENT_STATE_OPEN = 1;
|
||||
STORAGE_INCIDENT_STATE_REPAIRED = 2;
|
||||
STORAGE_INCIDENT_STATE_ACKNOWLEDGED = 3;
|
||||
}
|
||||
|
||||
message StorageIncident {
|
||||
string incident_id = 1;
|
||||
google.protobuf.Timestamp detected_at = 2;
|
||||
google.protobuf.Timestamp resolved_at = 3;
|
||||
StorageIncidentState state = 4;
|
||||
string scope = 5;
|
||||
string client_id = 6;
|
||||
string issue_uuid = 7;
|
||||
string summary = 8;
|
||||
bool data_loss = 9;
|
||||
bool automatically_repairable = 10;
|
||||
}
|
||||
|
||||
message ListStorageIncidentsRequest {
|
||||
bool include_resolved = 1;
|
||||
uint32 page_size = 2;
|
||||
string page_token = 3;
|
||||
}
|
||||
|
||||
message ListStorageIncidentsResponse {
|
||||
repeated StorageIncident incidents = 1;
|
||||
string next_page_token = 2;
|
||||
}
|
||||
|
||||
message RepairStorageIncidentRequest {
|
||||
string incident_id = 1;
|
||||
string request_id = 2;
|
||||
}
|
||||
|
||||
message RepairStorageIncidentResponse {
|
||||
StorageIncident incident = 1;
|
||||
}
|
||||
|
||||
message AcknowledgeStorageIncidentRequest {
|
||||
string incident_id = 1;
|
||||
string note = 2;
|
||||
string request_id = 3;
|
||||
}
|
||||
|
||||
message AcknowledgeStorageIncidentResponse {
|
||||
StorageIncident incident = 1;
|
||||
}
|
||||
|
||||
message AuthorizeClientTakeoverRequest {
|
||||
string client_id = 1;
|
||||
string client_instance_id = 2;
|
||||
string request_id = 3;
|
||||
}
|
||||
|
||||
message AuthorizeClientTakeoverResponse {
|
||||
google.protobuf.Timestamp expires_at = 1;
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user