docs: finalize rvbox v1 design and implementation plan
This commit is contained in:
@@ -0,0 +1,8 @@
|
|||||||
|
/bin/
|
||||||
|
/dist/
|
||||||
|
/coverage/
|
||||||
|
*.db
|
||||||
|
*.db-shm
|
||||||
|
*.db-wal
|
||||||
|
*.sock
|
||||||
|
*.tmp
|
||||||
@@ -11,6 +11,10 @@ Read the documents in this order:
|
|||||||
query/foreground semantics.
|
query/foreground semantics.
|
||||||
4. [Platform and operations](platform-and-operations.md) — Unix/Windows
|
4. [Platform and operations](platform-and-operations.md) — Unix/Windows
|
||||||
contracts, recovery, storage safety, telemetry, and defaults.
|
contracts, recovery, storage safety, telemetry, and defaults.
|
||||||
|
5. [Configuration contract](configuration.md) — strict TOML loading, shell
|
||||||
|
resolution, cross-field validation, and annotated server/client examples.
|
||||||
|
6. [Go implementation plan](implementation-plan.v1.md) — phased build order,
|
||||||
|
package boundaries, storage/session/client details, tests, and release gates.
|
||||||
|
|
||||||
The wire authority is in [`../protos/rvbox/v1`](../protos/rvbox/v1):
|
The wire authority is in [`../protos/rvbox/v1`](../protos/rvbox/v1):
|
||||||
`common.proto` contains shared data types, `agent.proto` contains the
|
`common.proto` contains shared data types, `agent.proto` contains the
|
||||||
|
|||||||
+171
-55
@@ -12,6 +12,9 @@ RVBox deliberately executes arbitrary commands with the identity, permissions,
|
|||||||
and base environment of the client daemon. It is therefore an administrative
|
and base environment of the client daemon. It is therefore an administrative
|
||||||
tool, not a multi-tenant remote-execution service.
|
tool, not a multi-tenant remote-execution service.
|
||||||
|
|
||||||
|
Both daemons use strict TOML 1.0 configuration. The normative schema and fully
|
||||||
|
annotated examples are in the [configuration contract](configuration.md).
|
||||||
|
|
||||||
## Non-goals
|
## Non-goals
|
||||||
|
|
||||||
- Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist
|
- Mutual TLS, client certificates, enrollment tokens, and a client-ID allowlist
|
||||||
@@ -26,10 +29,14 @@ tool, not a multi-tenant remote-execution service.
|
|||||||
|
|
||||||
The server accepts a self-reported hostname as `client_id`; it is an opaque
|
The server accepts a self-reported hostname as `client_id`; it is an opaque
|
||||||
1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer
|
1–128 ASCII-character routing/display key. Unknown IDs are accepted. A newer
|
||||||
registration for an ID replaces its prior live session. Consequently, a peer
|
registration from the same durable client instance replaces its prior live
|
||||||
able to reach nginx can impersonate or take over a client ID. This is an
|
session. A different instance is accepted normally when no session for that
|
||||||
accepted v1 limitation and deployments must restrict the endpoint to a trusted
|
client ID is live. While one is live, a different instance is rejected unless
|
||||||
network.
|
an operator grants a one-shot override for that exact pending instance. This
|
||||||
|
prevents accidental hostname collisions but is not authentication: a peer able
|
||||||
|
to reach nginx and copy or guess the identifiers can still impersonate a
|
||||||
|
client. This is an accepted v1 limitation and deployments must restrict the
|
||||||
|
endpoint to a trusted network.
|
||||||
|
|
||||||
The Unix control socket is local-only and mode `0600`, owned by the server
|
The Unix control socket is local-only and mode `0600`, owned by the server
|
||||||
account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated;
|
account. The optional HTTP JSON-RPC endpoint is intentionally unauthenticated;
|
||||||
@@ -57,53 +64,113 @@ its dispatch loops.
|
|||||||
|
|
||||||
1. The client connects over WSS and sends `ClientHello` with its client ID,
|
1. The client connects over WSS and sends `ClientHello` with its client ID,
|
||||||
protocol capability, OS/architecture, daemon version, current daemon CWD,
|
protocol capability, OS/architecture, daemon version, current daemon CWD,
|
||||||
supported shells, and a fresh reconnect UUID.
|
supported shells, and a durable random client-instance UUID.
|
||||||
2. The server accepts the current compatible protocol version, fences the
|
2. The server accepts the current compatible protocol version, fences the
|
||||||
previous connection for that ID, and returns a fresh server-issued
|
previous connection for that ID and same client instance, and returns a
|
||||||
`session_id` and monotonic `session_generation`.
|
fresh server-issued `session_id` and monotonic `session_generation`. A live
|
||||||
|
claim from a different client instance is rejected while the current session
|
||||||
|
is live unless an operator has explicitly authorized that pending instance.
|
||||||
|
With no live session, the new instance is accepted normally.
|
||||||
3. Every client-to-server envelope and server dispatch is bound to that token.
|
3. Every client-to-server envelope and server dispatch is bound to that token.
|
||||||
The server discards traffic from superseded sessions, including late output.
|
The server discards traffic from superseded sessions, including late output.
|
||||||
4. The server queues work while a client is offline and dispatches it only when
|
4. The server queues work while a client is offline and dispatches it only when
|
||||||
the active session advertises capacity. A replacement session immediately
|
the active session advertises capacity. A replacement session immediately
|
||||||
resumes non-terminal reconciliation.
|
performs bidirectional reconciliation between the server's non-terminal set
|
||||||
|
and the client's complete retained-command set. The server returns explicit
|
||||||
|
local terminate/discard decisions before new dispatch begins.
|
||||||
|
|
||||||
## Command model
|
## Command model
|
||||||
|
|
||||||
Each user request has a server-generated UUID (`issue_uuid`) and a durable
|
Each user request has a UUID (`issue_uuid`) and a durable request record: target
|
||||||
request record: target client, request/issue timestamps, shell type, command
|
client, request/issue timestamps, shell type, command text or script descriptor,
|
||||||
text or script descriptor, CWD, environment overrides, resource-profile flags,
|
CWD, environment overrides, resource-profile flags, and lifecycle state. `rvc`
|
||||||
and lifecycle state. The UUID is the end-to-end idempotency key. Transport is
|
normally supplies this UUID as its optional `request_id`; the server generates
|
||||||
at-least-once, but the client durably remembers accepted UUIDs and never starts
|
one when it is omitted. Command and mutation identifiers are UUIDv7 values.
|
||||||
the same request twice.
|
Transport is at-least-once, but the client durably
|
||||||
|
remembers accepted UUIDs and provides an at-most-once execution guarantee: it
|
||||||
|
never authorizes the same request to execute twice. A crash in the launch window
|
||||||
|
may interrupt a command before its requested code runs, but must never cause an
|
||||||
|
automatic retry with an uncertain prior outcome.
|
||||||
|
|
||||||
|
After full terminal history is removed, each client retains a compact FIFO of
|
||||||
|
the most recent 1,000,000 command tombstones containing binary UUIDv7,
|
||||||
|
immutable request hash, and completion/acknowledgement time. The server retains
|
||||||
|
the most recent 1,000,000 command tombstones globally as a first-line duplicate
|
||||||
|
check. A matching UUID/hash returns structured `ALREADY_EXECUTED`; a matching
|
||||||
|
UUID with different immutable content is a conflict. These ledgers have separate
|
||||||
|
count-based budgets and do not retain command payload or output. Replay
|
||||||
|
protection older than the retained client tombstone horizon is best-effort.
|
||||||
|
|
||||||
The server states are:
|
The server states are:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
queued -> dispatched -> accepted -> running -> succeeded | failed | terminated
|
queued -> dispatched | cancelled | expired
|
||||||
\-------------------------------> cancelled
|
dispatched -> queued | accepted | rejected | cancelled
|
||||||
|
accepted -> running | rejected | cancelled
|
||||||
|
running -> succeeded | failed | terminated | interrupted
|
||||||
```
|
```
|
||||||
|
|
||||||
`accepted` means the client has durably accepted the request; `running` means
|
Recovery/corruption handling may also move an affected `dispatched` or
|
||||||
the process has been launched. `cancelled` is used when it is stopped before
|
`accepted` command to `interrupted`; those are exceptional reconciliation
|
||||||
launch. A kill racing launch is resolved by command revision: the client either
|
transitions, not normal execution outcomes.
|
||||||
acknowledges cancellation before launch or launches then immediately applies
|
|
||||||
the requested signal, recording the race.
|
`accepted` means the client has durably admitted the request; `running` means
|
||||||
|
the requested process has been launched. `cancelled` is used when it is stopped
|
||||||
|
before launch. A permanent pre-launch validation, script-transfer, or process-
|
||||||
|
preparation failure is terminal `rejected`; `failed` is reserved for code that
|
||||||
|
actually launched. A transient
|
||||||
|
capacity rejection returns to `queued` with backoff and remains subject to its
|
||||||
|
queue TTL. The server may cancel work that has never been dispatched immediately.
|
||||||
|
For dispatched or accepted work, it persists a higher command revision and
|
||||||
|
sends the revisioned signal without prematurely declaring a terminal state. The
|
||||||
|
client either returns a revisioned `cancelled` lifecycle before launch
|
||||||
|
authorization or applies the signal after authorization and returns a
|
||||||
|
revisioned signal result. Cancellation intent remains internal rather than
|
||||||
|
adding a public lifecycle state.
|
||||||
|
|
||||||
|
Queued work has a configurable acceptance deadline, default 15 minutes; zero
|
||||||
|
explicitly means no expiry. A command that was never dispatched becomes
|
||||||
|
terminal `expired` at its deadline. A dispatched command whose acceptance is
|
||||||
|
uncertain remains non-terminal and is displayed as expired pending
|
||||||
|
reconciliation. Later client evidence updates the actual lifecycle. Acceptance
|
||||||
|
or execution observed after the deadline creates an incident and is displayed
|
||||||
|
as late-after-expiry; the server requests termination but continues recording
|
||||||
|
the actual client-reported outcome.
|
||||||
|
|
||||||
|
Process launch has internal durable phases `launch_prepared` and
|
||||||
|
`launch_authorized` between public `accepted` and `running`. The client creates
|
||||||
|
the process behind an OS-specific execution barrier, durably records its process
|
||||||
|
identity, durably authorizes launch, and only then releases requested command
|
||||||
|
code. Once authorization is durable, an uncertain outcome is reconciled as
|
||||||
|
`interrupted`, never by redispatching that UUID.
|
||||||
|
|
||||||
The client permits 16 concurrent processes and 100 pending commands by default.
|
The client permits 16 concurrent processes and 100 pending commands by default.
|
||||||
Those values are configurable and advertised to the server. A full queue causes
|
Those values are configurable and advertised to the server. A full queue causes
|
||||||
a structured capacity error rather than creating unbounded work.
|
a structured capacity error rather than creating unbounded work.
|
||||||
|
The server separately defaults to at most 1,000 queued commands for one target
|
||||||
|
client and 10,000 queued commands globally, subject to the stricter byte quotas.
|
||||||
|
|
||||||
### Execution contract
|
### Execution contract
|
||||||
|
|
||||||
- Shell selection is explicit: Unix-like clients support `sh` and `bash`;
|
- Shell selection is explicit: Unix-like clients support `sh` and `bash`;
|
||||||
Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`.
|
Windows supports `cmd` and `powershell`. Defaults are `sh` and `powershell`.
|
||||||
Unsupported shells are rejected; no fallback occurs.
|
Unsupported shells are rejected; no fallback occurs.
|
||||||
|
- The client materializes `command_text` as a private generated wrapper and
|
||||||
|
executes that file with exactly the selected shell: `sh`/`bash`, `cmd`, or
|
||||||
|
PowerShell. This avoids transport quoting and Windows command-line limits;
|
||||||
|
it never performs shell detection or fallback. Wrapper cleanup follows the
|
||||||
|
same terminal rule as uploaded scripts.
|
||||||
- A command receives the daemon account's permissions and startup environment,
|
- A command receives the daemon account's permissions and startup environment,
|
||||||
overlaid with the persisted `env_overrides` map. The effective CWD is the
|
overlaid with the persisted `env_overrides` map. The effective CWD is the
|
||||||
requested existing directory or the registered daemon CWD when omitted.
|
requested existing directory or the registered daemon CWD when omitted.
|
||||||
- Each command is isolated into a process tree: a Unix session/process group or
|
- Each command is isolated into a process tree: a Unix session/process group or
|
||||||
a Windows Job Object. A daemon that cannot supervise its children terminates
|
a Windows Job Object. A daemon that cannot supervise its children terminates
|
||||||
them and reports interruption rather than claiming recovery it cannot make.
|
them and reports interruption rather than claiming recovery it cannot make.
|
||||||
|
- Terminal lifecycle means the supervised tree is empty, not only that its root
|
||||||
|
shell exited. After root exit, RVBox allows a configurable 5-second descendant
|
||||||
|
and output-drain grace period, then terminates residual descendants. It drains
|
||||||
|
capture pipes and durably sequences all retained output (or an explicit
|
||||||
|
`OutputIncomplete` marker) before emitting the terminal lifecycle event.
|
||||||
- Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`,
|
- Resource profiles are composable flags (`LIGHT`, `CPU_MEDIUM`, `CPU_HEAVY`,
|
||||||
`MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in;
|
`MEM_MEDIUM`, `MEM_HEAVY`, `DISK_MEDIUM`, `DISK_HEAVY`). Profiles are opt-in;
|
||||||
the concrete administrator-configured limits are applied with cgroup v2 when
|
the concrete administrator-configured limits are applied with cgroup v2 when
|
||||||
@@ -115,18 +182,20 @@ a structured capacity error rather than creating unbounded work.
|
|||||||
Scripts are content-addressed uploads, not shell-escaped command strings. The
|
Scripts are content-addressed uploads, not shell-escaped command strings. The
|
||||||
server sends a descriptor containing SHA-256 and then ordered chunks (default
|
server sends a descriptor containing SHA-256 and then ordered chunks (default
|
||||||
maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary
|
maximum: 10 MiB). The client verifies the digest, writes an owner-only temporary
|
||||||
file beneath the effective CWD, executes it with the selected shell, and removes
|
file beneath the effective CWD, executes it with exactly the selected shell, and
|
||||||
it after the command reaches a terminal state. Script content is not copied into
|
removes it after the command reaches a terminal state. The descriptor filename
|
||||||
the audit log; its digest and metadata are.
|
is display metadata only and is never used as a path component. Script content
|
||||||
|
is not copied into the audit log; its digest and metadata are.
|
||||||
|
|
||||||
## Process control and diagnostics
|
## Process control and diagnostics
|
||||||
|
|
||||||
On Unix, `kill` addresses the command's process group and accepts normal signal
|
On Unix, `kill` addresses the command's process group and accepts the portable
|
||||||
names/numbers supported by that client. On Windows, only `SIGTERM` and `SIGKILL`
|
v1 set `HUP`, `INT`, `TERM`, `KILL`, `USR1`, and `USR2` (including their
|
||||||
are valid. `SIGTERM` makes a best-effort `CTRL_BREAK_EVENT` delivery to the
|
`SIG`-prefixed CLI spellings). Arbitrary native signal numbers are not part of
|
||||||
dedicated console group, waits 10 seconds, then terminates the Job Object if
|
v1. On Windows, only `TERM` and `KILL` are valid. `TERM` makes a best-effort
|
||||||
needed. `SIGKILL` immediately terminates the Job Object. The response reports
|
`CTRL_BREAK_EVENT` delivery to the dedicated console group, waits 10 seconds,
|
||||||
the actual escalation outcome.
|
then terminates the Job Object if needed. `KILL` immediately terminates the Job
|
||||||
|
Object. The response reports the actual escalation outcome.
|
||||||
|
|
||||||
Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is
|
Lifecycle state never asserts `hung`. A separate `suspected_hung` diagnostic is
|
||||||
emitted after the configurable default of 10 minutes without observable
|
emitted after the configurable default of 10 minutes without observable
|
||||||
@@ -142,31 +211,73 @@ restart and reconciles only these; server-confirmed historical terminal commands
|
|||||||
are not re-reconciled. The client stores only active command state and output
|
are not re-reconciled. The client stores only active command state and output
|
||||||
not acknowledged by the server.
|
not acknowledged by the server.
|
||||||
|
|
||||||
Every execution-originated event has a strictly increasing `event_seq` scoped to
|
Every transmitted execution event has a strictly increasing `event_seq` scoped
|
||||||
one command. This includes lifecycle transitions, stdout/stderr chunks, stdin
|
to one command. The client first stores lifecycle transitions, stdout/stderr,
|
||||||
acknowledgments, resource snapshots, signals, and terminal events. The server
|
stdin acknowledgements, resource snapshots, signals, and terminal state in a
|
||||||
preserves the sequence and additionally records receipt time. This is the
|
durable local order. It durably assigns wire sequences only as entries enter the
|
||||||
canonical reconstruction order across interleaved streams and retries.
|
bounded send window, after any unsent-output compaction. Once assigned, an event
|
||||||
Truncation is separate range metadata so it can truthfully describe missing
|
is pinned until acknowledged and retry content is immutable. A client-side
|
||||||
event sequences without consuming one itself.
|
truncation marker therefore consumes a normal sequence without creating a wire
|
||||||
|
gap. Server-side retention exposes removed event ranges as query metadata
|
||||||
|
without allocating client sequences. The server also records receipt time;
|
||||||
|
`event_seq` remains the canonical transmitted order across streams and retries.
|
||||||
|
|
||||||
Output chunks are Zstandard-compressed before persistent quota accounting.
|
The terminal lifecycle event is always the final client event for a command.
|
||||||
Per-command history is a rolling compressed window (10 MiB default), so the
|
The control-plane follow wrapper emits any server-created retention metadata
|
||||||
oldest output segments for that command are removed first and a sequence-range
|
before returning that terminal event. Followers may therefore stop on terminal
|
||||||
truncation marker remains. This applies to active and terminal commands. When
|
without missing subsequently sequenced stdout/stderr or known truncation data.
|
||||||
connected, the client first removes server-acknowledged segments. While offline,
|
|
||||||
it must still honor both hard caps: it retains the newest tail, removes oldest
|
|
||||||
unacknowledged compressed chunks when necessary, and records their exact missing
|
|
||||||
ranges for durable reporting on reconnect.
|
|
||||||
|
|
||||||
Each client also has a 50 MiB aggregate compressed spool cap for active,
|
All command-owned stored data is quota-accounted: execution metadata, script
|
||||||
unacknowledged work. The server's matching per-client compressed-history cap is
|
body, pending stdin, events, and output. Stored message/blob payloads are
|
||||||
50 MiB; it evicts that client's oldest terminal command records as needed. The
|
Zstandard-compressed; necessary SQLite index/state columns are charged by their
|
||||||
server-wide cap is 1 GiB; it evicts whole oldest terminal command records
|
encoded lengths plus a conservative versioned per-row/index overhead rather
|
||||||
(metadata and output), never arbitrary stdout/stderr rows. Active commands are
|
than pretending they are free. This logical accounting is deterministic across
|
||||||
protected. If active commands alone consume a per-client budget, their oldest
|
SQLite compaction. A separate filesystem free-space floor protects WAL,
|
||||||
acknowledged output rotates by the per-command rule; pipes continue draining so
|
temporary files, tombstones, and accounting variance.
|
||||||
a child cannot deadlock on output.
|
The default total is 32 MiB per command, 256 MiB per client on both client and
|
||||||
|
server, and 4 GiB server-wide. A separate 10 MiB rolling output window remains
|
||||||
|
per command, and raw script input remains limited to 10 MiB. Client accounting
|
||||||
|
also charges the raw generated execution wrapper/script file while it exists,
|
||||||
|
even though the compressed durable source was already charged.
|
||||||
|
|
||||||
|
Each accepted active command reserves 64 KiB of its quota for bounded closeout
|
||||||
|
metadata. Essential state for an accepted stdin/signal mutation is additionally
|
||||||
|
reserved before that mutation succeeds. Essential active state is never silently
|
||||||
|
rolled. Server output is evictable: the oldest retained chunks are removed first
|
||||||
|
and sequence-range query metadata remains. The client removes acknowledged data,
|
||||||
|
pins its bounded assigned send window, and may replace older unsequenced output
|
||||||
|
with normally sequenced byte-loss markers. While offline or under sustained
|
||||||
|
overload, it retains the newest tail and records exact known lost bytes. At
|
||||||
|
client/server aggregate limits, terminal commands are evicted as whole UUID
|
||||||
|
records in oldest server issue-time/UUIDv7 order first. If active data
|
||||||
|
alone reaches a limit, output rotates or enters loss mode and new essential
|
||||||
|
allocations are rejected with `CAPACITY_EXHAUSTED`; pipes continue draining.
|
||||||
|
|
||||||
|
Transient raw output also has bounded high/low watermarks before compression:
|
||||||
|
1 MiB/256 KiB per command, 8 MiB/4 MiB per client daemon, and 64 MiB/32 MiB
|
||||||
|
server-wide by default. Crossing a client high watermark enters loss mode;
|
||||||
|
still-unsequenced output bytes may be discarded before compression and replaced
|
||||||
|
in durable local order by a later `OutputTruncation`. Loss mode ends only below
|
||||||
|
the corresponding low watermark. The server never drops an already sequenced
|
||||||
|
client event: at its ingress high watermark it withholds acknowledgement and
|
||||||
|
closes an overproducing session if bounded admission cannot continue, letting
|
||||||
|
the durable client retry after the server backlog falls below its low watermark.
|
||||||
|
Lifecycle,
|
||||||
|
stdin acknowledgements, signal results, and truncation/incomplete markers use
|
||||||
|
reserved capacity and are never treated as droppable output.
|
||||||
|
|
||||||
|
Independently of byte pressure, the server reclaims each whole terminal command
|
||||||
|
and all command-owned data 30 days after its terminal time by default. Zero
|
||||||
|
explicitly disables age rotation. Compact replay tombstones, audit records, and
|
||||||
|
storage incidents remain under their separate retention policies.
|
||||||
|
|
||||||
|
Storage health is derived from durable incident records. Safe repairs resolve
|
||||||
|
an incident automatically; known data loss remains dirty until an operator
|
||||||
|
explicitly acknowledges it. Resolution clears the dirty health flag but does
|
||||||
|
not erase records as part of that action. Unresolved compact records are
|
||||||
|
non-evictable; resolved summaries and audit entries follow the separate 100 MiB
|
||||||
|
audit/incident history rotation. Recovery and incident management are described
|
||||||
|
in the platform contract and exposed through the control plane.
|
||||||
|
|
||||||
## Audit and timestamps
|
## Audit and timestamps
|
||||||
|
|
||||||
@@ -175,8 +286,13 @@ events: source/transport identity where available, target client, UUID, action,
|
|||||||
request time, result, and error. It records command text, environment-override
|
request time, result, and error. It records command text, environment-override
|
||||||
names (not values), and script metadata/digest, but not duplicated
|
names (not values), and script metadata/digest, but not duplicated
|
||||||
stdin/stdout/stderr payloads. The persisted execution request necessarily keeps
|
stdin/stdout/stderr payloads. The persisted execution request necessarily keeps
|
||||||
override values for dispatch/retry and must be access-controlled as sensitive
|
override values for dispatch/retry. In v1, command text, scripts, environment
|
||||||
data. Audit retention is configured independently of output retention.
|
values, stdin, output, and other command-owned payloads are stored in plaintext;
|
||||||
|
application-level encryption and key management are deferred to a future
|
||||||
|
version. Audit storage uses compressed segments under a separate 100 MiB default
|
||||||
|
quota and rotates complete oldest segments. Audit retention is configured
|
||||||
|
independently of command retention; its default zero age limit means quota-only
|
||||||
|
rotation.
|
||||||
|
|
||||||
All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client
|
All protocol timestamps are UTC `google.protobuf.Timestamp` values. Client
|
||||||
observed timestamps and server receipt timestamps are distinct; the latter is
|
observed timestamps and server receipt timestamps are distinct; the latter is
|
||||||
|
|||||||
@@ -0,0 +1,88 @@
|
|||||||
|
# RVBox v1 configuration contract
|
||||||
|
|
||||||
|
RVBox v1 uses TOML 1.0 for daemon configuration. The normative annotated
|
||||||
|
examples are [`examples/server.toml`](examples/server.toml) and
|
||||||
|
[`examples/client.toml`](examples/client.toml). They list every supported v1
|
||||||
|
knob, with the routing, persistence, and safety limits first in each section.
|
||||||
|
|
||||||
|
## Loading and precedence
|
||||||
|
|
||||||
|
- `rvbox-server --config PATH` and `rvbox --config PATH` load one UTF-8 TOML
|
||||||
|
file. There is no implicit merge of multiple files and no hot reload in v1.
|
||||||
|
- Precedence is compiled default, then TOML, then an explicitly supplied CLI
|
||||||
|
flag. Flags exist for operationally important scalar keys; they use the same
|
||||||
|
validation as TOML. RVBox does not implicitly import configuration from
|
||||||
|
environment variables.
|
||||||
|
- Unknown keys, duplicate keys/tables, type mismatches, invalid UTF-8, and
|
||||||
|
values outside documented ranges are startup errors. Parsing never silently
|
||||||
|
substitutes a default for a present invalid value.
|
||||||
|
- Durations are quoted Go-style duration strings such as `"250ms"`, `"15m"`,
|
||||||
|
and `"720h"`. Byte sizes and counts are base-10 TOML integers whose values are
|
||||||
|
bytes; comments show the equivalent binary unit. URLs and paths are strings.
|
||||||
|
- Relative paths are rejected for state, socket, CA, shell-executable, and
|
||||||
|
allowed-CWD-root fields. `client.daemon_cwd` is resolved once at startup and
|
||||||
|
then stored and advertised as an absolute path. The annotated client example
|
||||||
|
uses Unix paths; a Windows deployment replaces `state_dir`, `daemon_cwd`, and
|
||||||
|
relevant shell paths with absolute Windows paths. Shell fields for the other
|
||||||
|
platform are syntax-checked but not resolved or advertised.
|
||||||
|
- The daemon prints its effective configuration after validation, with no
|
||||||
|
command data or TLS material. Since v1 stores command/environment payloads in
|
||||||
|
plaintext, configuration output is hygiene rather than a secrecy guarantee.
|
||||||
|
|
||||||
|
## Limits and cross-field validation
|
||||||
|
|
||||||
|
Configuration may lower protocol and storage limits, but may not raise a hard
|
||||||
|
wire ceiling above the v1 values in the examples. For every high/low watermark,
|
||||||
|
`0 < low < high`; send windows must fit below their corresponding durable quota.
|
||||||
|
The 64 KiB command closeout reserve must fit within the command quota, and the
|
||||||
|
command quota must fit within the client and server tiers. Queue and byte counts
|
||||||
|
must be positive except where a comment explicitly gives zero a disabling or
|
||||||
|
indefinite meaning. `queue.max_per_client` may not exceed `queue.max_server`.
|
||||||
|
|
||||||
|
The server must reject external JSON-RPC binds unless `json_rpc.enabled=true`.
|
||||||
|
Any enabled non-loopback bind produces a conspicuous warning but is permitted by
|
||||||
|
the accepted v1 debugging contract. The control Unix socket always uses mode
|
||||||
|
`0600`; it is not a configurable relaxation.
|
||||||
|
|
||||||
|
## Shell executable resolution
|
||||||
|
|
||||||
|
Each `ShellType` maps to one startup-validated absolute executable path from
|
||||||
|
`[shells]`. The client canonicalizes the path, verifies that it names an
|
||||||
|
executable regular file appropriate to the platform, and advertises only shells
|
||||||
|
that passed validation. The configured platform default must be one of those
|
||||||
|
shells. There is no fallback.
|
||||||
|
|
||||||
|
The resolved executable is independent of a command's `PATH` override. For
|
||||||
|
example, a request for `SHELL_BASH` still launches the validated `/bin/bash`
|
||||||
|
even when the request contains `PATH=/tmp/untrusted`; it never searches that
|
||||||
|
directory for another `bash`. An administrator may intentionally select another
|
||||||
|
implementation, such as an absolute `pwsh.exe` path for `SHELL_POWERSHELL`, but
|
||||||
|
the selection remains fixed until daemon restart.
|
||||||
|
|
||||||
|
Command text and uploaded scripts are written to generated wrapper paths and
|
||||||
|
passed to exactly this executable. The user-supplied script filename is display
|
||||||
|
metadata only.
|
||||||
|
|
||||||
|
## Resource-profile composition
|
||||||
|
|
||||||
|
`LIGHT` is exclusive. Otherwise, a request may combine at most one CPU tier,
|
||||||
|
one memory tier, and one disk tier. Thus `CPU_HEAVY + MEM_MEDIUM` is valid, while
|
||||||
|
`CPU_MEDIUM + CPU_HEAVY` and `LIGHT + MEM_HEAVY` are invalid. A configured
|
||||||
|
profile declares which controls are required. If the platform cannot apply a
|
||||||
|
required control atomically before launch, the command is terminal `REJECTED`.
|
||||||
|
|
||||||
|
Profile names describe administrator-defined allowance classes: a `HEAVY` tier
|
||||||
|
normally permits more resources than `MEDIUM`; RVBox does not invent numeric
|
||||||
|
values. Zero for an individual numeric limit means that control is not requested
|
||||||
|
by that profile, but every name in `required_controls` must have a nonzero,
|
||||||
|
platform-applicable value. Linux disk limits use configured cgroup device
|
||||||
|
major/minor keys. Windows applies the equivalent whole-Job rate control and
|
||||||
|
ignores Linux device maps only when disk control is not declared required.
|
||||||
|
|
||||||
|
## Validation ownership
|
||||||
|
|
||||||
|
`internal/config` owns TOML DTOs, strict decoding, default application, flag
|
||||||
|
overrides, canonicalization, and cross-field validation. It converts the parsed
|
||||||
|
form into immutable domain configuration before listeners or child processes
|
||||||
|
start. Network, storage, and supervisor packages receive only their relevant
|
||||||
|
validated sub-configuration and never parse TOML themselves.
|
||||||
+86
-21
@@ -12,32 +12,83 @@ an optional JSON-RPC 2.0 HTTP adapter for local debugging and batch automation.
|
|||||||
It has no authentication by design. Binding it beyond loopback is an explicit
|
It has no authentication by design. Binding it beyond loopback is an explicit
|
||||||
deployment choice and requires external protection.
|
deployment choice and requires external protection.
|
||||||
|
|
||||||
gRPC can stream `RunCommandAndFollow` and `FollowCommand`. JSON-RPC remains
|
gRPC streams command history and live events through `FollowCommand`. JSON-RPC
|
||||||
simple: callers issue work, query command state, poll event/output pages after
|
remains simple: callers issue work, query command state, poll event/output pages
|
||||||
an event sequence, append stdin, close stdin, or signal a command. It does not
|
after an event sequence, append stdin, close stdin, or signal a command. It does
|
||||||
invent a separate event-stream protocol.
|
not invent a separate event-stream protocol.
|
||||||
|
|
||||||
The JSON-RPC method names are the lower-camel protobuf operation names:
|
The JSON-RPC method names are the lower-camel protobuf operation names:
|
||||||
`listClients`, `getClient`, `listCommands`, `getCommand`, `runCommand`,
|
`listClients`, `getClient`, `listCommands`, `getCommand`, `runCommand`,
|
||||||
`appendStdin`, `closeStdin`, `signalCommand`, and `getOutput`. Parameters and
|
`appendStdin`, `closeStdin`, `signalCommand`, `getOutput`, and the three storage
|
||||||
results use protobuf JSON mapping (including base64 strings for `bytes` and UTC
|
incident methods documented below. Parameters and results use protobuf JSON
|
||||||
RFC 3339 strings for timestamps); JSON-RPC errors carry the corresponding
|
mapping (including base64 strings for `bytes` and UTC RFC 3339 strings for
|
||||||
`ControlError` code/data. `getOutput` and `getCommand` are the polling path for
|
timestamps). Control response messages contain successful results only. gRPC
|
||||||
what gRPC exposes as follow streams.
|
failures use canonical non-OK status codes with structured RVBox details where
|
||||||
|
needed; the JSON-RPC adapter maps the same domain errors to standard JSON-RPC
|
||||||
|
error objects. `getOutput` and `getCommand` are the polling path for what gRPC
|
||||||
|
exposes as follow streams.
|
||||||
|
|
||||||
|
`rvc stat CLIENT` also shows the active durable client-instance ID and the most
|
||||||
|
recent different instance rejected while that client is live. An operator may
|
||||||
|
run `rvc client takeover CLIENT INSTANCE-ID`; this creates a one-shot 5-minute
|
||||||
|
authorization for that exact pending claim. Its matching reconnect consumes the
|
||||||
|
authorization and fences the old session. If the old session is no longer live,
|
||||||
|
the replacement connects normally without this command. The JSON-RPC method is
|
||||||
|
`authorizeClientTakeover`.
|
||||||
|
|
||||||
|
Both transports enforce the same decoded field limits. A control gRPC request
|
||||||
|
may be at most 16 MiB; a JSON-RPC HTTP body may be at most 24 MiB to accommodate
|
||||||
|
base64 expansion of the 10 MiB script maximum. `ExecutionSpec` itself may be at
|
||||||
|
most 768 KiB, which also keeps its agent dispatch below the 1 MiB envelope cap.
|
||||||
|
|
||||||
## CLI semantics
|
## CLI semantics
|
||||||
|
|
||||||
`rvc stat` maps to `ListClients`, `GetClient`, `ListCommands`, and `GetCommand`.
|
`rvc stat` maps to `ListClients`, `GetClient`, `ListCommands`, and `GetCommand`.
|
||||||
History pages default to 20 commands and may request at most 100. Output pages
|
History pages default to 20 commands and may request at most 100. Historical
|
||||||
default to 100 lines; a line is a display operation over ordered chunks, not a
|
output uses opaque, byte-bounded cursors and may resume within an output event;
|
||||||
protocol boundary. Output can be filtered by stream and timestamped with the
|
it is not numbered or paginated by lines. The server returns uncompressed output
|
||||||
server's recorded client-observed timestamp plus stream name.
|
slices with event sequence, byte offset, stream, observed timestamp, and server
|
||||||
|
receipt timestamp. `rvc` may render line-oriented human output, but line
|
||||||
|
boundaries are not storage or pagination boundaries.
|
||||||
|
|
||||||
`rvc run` creates a command. Foreground mode runs `RunCommandAndFollow`, which
|
`rvc run` always creates a durable command through unary `RunCommand`, which
|
||||||
streams output and stops on a terminal event. `--background` uses `RunCommand`
|
returns the UUID. Foreground mode then calls `FollowCommand` with that UUID,
|
||||||
and returns the UUID immediately. Interrupting the CLI, timing out its local
|
`after_event_seq=0`, and `include_existing=true`, streaming output until a
|
||||||
wait, or losing the local control connection never cancels remote work. The
|
terminal event. `--background` returns immediately after `RunCommand`.
|
||||||
explicit `rvc kill` operation is the only termination path.
|
Interrupting the CLI, timing out its local wait, or losing the local control
|
||||||
|
connection never cancels remote work. The explicit `rvc kill` operation is the
|
||||||
|
only termination path. A foreground caller can resume `FollowCommand` after its
|
||||||
|
last received event sequence without missing durable history.
|
||||||
|
|
||||||
|
`FollowCommandResponse` wraps either a client-sequenced `CommandEvent` or a
|
||||||
|
server-created retention marker and exposes server receipt/recording time
|
||||||
|
separately from client observation time. A retention marker does not advance the
|
||||||
|
resume cursor. When retained history is incomplete, the server emits the relevant
|
||||||
|
marker before any terminal event on that stream, so a follower never stops on
|
||||||
|
terminal while believing truncated history was complete.
|
||||||
|
|
||||||
|
Mutating requests accept an optional `request_id`. `rvc` generates one per
|
||||||
|
mutation and reuses it for transport retries; `--request-id` lets automation
|
||||||
|
reuse it across CLI invocations. The CLI accepts that option globally and drops
|
||||||
|
it for reads. For `RunCommand`, the supplied request ID is the command's
|
||||||
|
`issue_uuid`. For stdin, close, signal, repair, and acknowledgement operations
|
||||||
|
it deduplicates that action while `issue_uuid` or `incident_id` continues to
|
||||||
|
identify the target. The server generates a request ID when omitted, preserving
|
||||||
|
simple JSON-RPC use.
|
||||||
|
|
||||||
|
The server stores the mutation kind, target, immutable request hash, and result
|
||||||
|
under that ID. An identical retry returns the original result; reuse with any
|
||||||
|
different method, target, or content returns `CONFLICT`. The record is owned by
|
||||||
|
the affected command or incident for quota and retention. A retry after that
|
||||||
|
owner has been reclaimed cannot repeat the action: it returns retained
|
||||||
|
tombstone information where available or `NOT_FOUND`.
|
||||||
|
|
||||||
|
`rvc run --queue-ttl` controls how long work may wait for server-confirmed
|
||||||
|
acceptance and defaults to 15 minutes; zero means indefinite. `rvc stat`
|
||||||
|
distinguishes terminal `Expired`, `Expired (awaiting reconciliation)`, and
|
||||||
|
actual lifecycle with a late-after-expiry warning. It renders terminal
|
||||||
|
`Rejected` with the client's structured validation/platform reason; `Failed`
|
||||||
|
means the requested code actually launched.
|
||||||
|
|
||||||
`rvc append` turns a string into `StdinWrite` with `append_newline=true` unless
|
`rvc append` turns a string into `StdinWrite` with `append_newline=true` unless
|
||||||
the caller selects raw mode; `--file` supplies raw bytes; `--attach` streams
|
the caller selects raw mode; `--file` supplies raw bytes; `--attach` streams
|
||||||
@@ -53,6 +104,20 @@ must not be shown as a terminal state. Pagination response cursors are stable
|
|||||||
within their declared ordering (newest issue time for command lists; increasing
|
within their declared ordering (newest issue time for command lists; increasing
|
||||||
`event_seq` for events/output).
|
`event_seq` for events/output).
|
||||||
|
|
||||||
The service returns `ControlError` codes for not found, offline, capacity,
|
Control failures are never embedded in otherwise-successful response messages.
|
||||||
invalid request, unsupported platform feature, conflict, truncation, and
|
They use canonical gRPC status codes with structured RVBox details; JSON-RPC
|
||||||
internal/transient errors. It never encodes errors only as CLI text.
|
returns the corresponding JSON-RPC error object, and `rvc` maps the same domain
|
||||||
|
error to a stable exit code rather than parsing text.
|
||||||
|
|
||||||
|
## Storage incidents
|
||||||
|
|
||||||
|
`rvc storage incidents` lists unresolved storage incidents by default and can
|
||||||
|
include resolved history. `rvc storage repair INCIDENT` attempts only a known
|
||||||
|
safe repair. `rvc storage acknowledge INCIDENT --note ...` accepts documented
|
||||||
|
irrecoverable loss and clears that incident from dirty health. Both mutations
|
||||||
|
use the same optional `--request-id` behavior as other mutations. An
|
||||||
|
acknowledgement note is required and bounded to 4 KiB. Repair or acknowledgement
|
||||||
|
changes incident state; it does not delete history as part of that action;
|
||||||
|
resolved history later follows the independent audit/incident quota. The
|
||||||
|
JSON-RPC equivalents are `listStorageIncidents`,
|
||||||
|
`repairStorageIncident`, and `acknowledgeStorageIncident`.
|
||||||
|
|||||||
@@ -0,0 +1,205 @@
|
|||||||
|
# RVBox v1 client example. Integer sizes are bytes; durations are quoted strings.
|
||||||
|
|
||||||
|
[client]
|
||||||
|
# Reverse WebSocket endpoint exposed by nginx.
|
||||||
|
server_url = "wss://rvbox.example.test/v1/agent"
|
||||||
|
# Private durable accepted-command, event-spool, and tombstone root.
|
||||||
|
state_dir = "/var/lib/rvbox"
|
||||||
|
# Empty selects the local hostname; otherwise use an opaque 1-128 ASCII ID.
|
||||||
|
client_id = ""
|
||||||
|
# Default absolute CWD when a request omits cwd.
|
||||||
|
daemon_cwd = "/"
|
||||||
|
# Maximum simultaneously running supervised process trees.
|
||||||
|
max_running_commands = 16
|
||||||
|
# Maximum durably accepted commands waiting to start.
|
||||||
|
max_queued_commands = 100
|
||||||
|
# Grace for reserved terminal cleanup during orderly daemon shutdown.
|
||||||
|
shutdown_grace = "30s"
|
||||||
|
|
||||||
|
[tls]
|
||||||
|
# Optional PEM CA bundle; empty uses the operating-system trust store.
|
||||||
|
ca_file = ""
|
||||||
|
# Optional certificate name override; empty derives it from server_url.
|
||||||
|
server_name = ""
|
||||||
|
|
||||||
|
[shells]
|
||||||
|
# Default shell enum on Unix; the matching path must validate at startup.
|
||||||
|
default_unix = "sh"
|
||||||
|
# Default shell enum on Windows; the matching path must validate at startup.
|
||||||
|
default_windows = "powershell"
|
||||||
|
# Absolute executable used for SHELL_SH; empty marks it unsupported.
|
||||||
|
sh = "/bin/sh"
|
||||||
|
# Absolute executable used for SHELL_BASH; empty marks it unsupported.
|
||||||
|
bash = "/bin/bash"
|
||||||
|
# Absolute executable used for SHELL_CMD on Windows; empty marks it unsupported.
|
||||||
|
cmd = "C:\\Windows\\System32\\cmd.exe"
|
||||||
|
# Absolute executable used for SHELL_POWERSHELL; may instead point to pwsh.exe.
|
||||||
|
powershell = "C:\\Windows\\System32\\WindowsPowerShell\\v1.0\\powershell.exe"
|
||||||
|
# Absolute CWD roots permitted by local policy; empty allows any accessible path.
|
||||||
|
allowed_cwd_roots = []
|
||||||
|
|
||||||
|
[network]
|
||||||
|
# Send WebSocket Ping after this period without inbound activity.
|
||||||
|
heartbeat_idle = "10s"
|
||||||
|
# Close and reconnect after this total period without inbound activity.
|
||||||
|
liveness_timeout = "30s"
|
||||||
|
# Initial full-jitter reconnect backoff.
|
||||||
|
reconnect_initial = "1s"
|
||||||
|
# Maximum full-jitter reconnect backoff.
|
||||||
|
reconnect_max = "60s"
|
||||||
|
# Continuous session duration that resets reconnect backoff.
|
||||||
|
stable_session_reset = "60s"
|
||||||
|
# Timeout for DNS/TCP/TLS/WebSocket establishment.
|
||||||
|
connect_timeout = "15s"
|
||||||
|
# Deadline for an individual WebSocket data-frame write.
|
||||||
|
write_deadline = "10s"
|
||||||
|
|
||||||
|
[storage]
|
||||||
|
# Maximum rolling compressed output retained for one command (10 MiB).
|
||||||
|
command_output_limit_bytes = 10485760
|
||||||
|
# Maximum charged data, including raw execution wrapper, per command (32 MiB).
|
||||||
|
command_total_limit_bytes = 33554432
|
||||||
|
# Maximum charged command data across this client daemon (256 MiB).
|
||||||
|
client_total_limit_bytes = 268435456
|
||||||
|
# Compact completed-command replay ledger entry cap.
|
||||||
|
tombstone_max_entries = 1000000
|
||||||
|
# Per-active-command quota held for terminal and loss metadata (64 KiB).
|
||||||
|
command_closeout_reserve_bytes = 65536
|
||||||
|
# Reject new unreserved writes below this filesystem free space (64 MiB).
|
||||||
|
free_space_floor_bytes = 67108864
|
||||||
|
# Target size for a sealed append-only spool segment (256 KiB).
|
||||||
|
segment_target_bytes = 262144
|
||||||
|
# Maximum time a group-commit waits before fsync; acknowledgements wait too.
|
||||||
|
durability_interval = "100ms"
|
||||||
|
|
||||||
|
[flow]
|
||||||
|
# Enter per-command raw-output loss mode at this backlog (1 MiB).
|
||||||
|
raw_output_command_high_bytes = 1048576
|
||||||
|
# Leave per-command raw-output loss mode below this backlog (256 KiB).
|
||||||
|
raw_output_command_low_bytes = 262144
|
||||||
|
# Enter client-wide raw-output loss mode at this backlog (8 MiB).
|
||||||
|
raw_output_client_high_bytes = 8388608
|
||||||
|
# Leave client-wide raw-output loss mode below this backlog (4 MiB).
|
||||||
|
raw_output_client_low_bytes = 4194304
|
||||||
|
# Maximum assigned but unacknowledged encoded bytes per command (1 MiB).
|
||||||
|
unacknowledged_per_command_bytes = 1048576
|
||||||
|
# Maximum assigned but unacknowledged encoded bytes for the session (8 MiB).
|
||||||
|
unacknowledged_per_session_bytes = 8388608
|
||||||
|
|
||||||
|
[execution]
|
||||||
|
# Wait after root exit for descendants and capture EOF before forced cleanup.
|
||||||
|
descendant_drain_grace = "5s"
|
||||||
|
# Wait after Windows CTRL_BREAK before terminating the complete Job Object.
|
||||||
|
windows_term_grace = "10s"
|
||||||
|
# Mark suspected_hung after no observable progress for this duration.
|
||||||
|
hung_threshold = "10m"
|
||||||
|
# Interval for best-effort process/resource diagnostic snapshots.
|
||||||
|
diagnostic_interval = "30s"
|
||||||
|
# Maximum raw uploaded script or generated script body (10 MiB).
|
||||||
|
max_script_bytes = 10485760
|
||||||
|
# Maximum serialized command ExecutionSpec accepted from the server (768 KiB).
|
||||||
|
max_execution_spec_bytes = 786432
|
||||||
|
# Maximum decoded AgentEnvelope accepted from the server (1 MiB).
|
||||||
|
max_agent_envelope_bytes = 1048576
|
||||||
|
# Maximum uncompressed stdout/stderr chunk emitted to the protocol (64 KiB).
|
||||||
|
max_raw_chunk_bytes = 65536
|
||||||
|
# Maximum protocol detail/reason text encoded as UTF-8 (4 KiB).
|
||||||
|
protocol_detail_max_bytes = 4096
|
||||||
|
|
||||||
|
[observability]
|
||||||
|
# Loopback HTTP listener for local liveness, storage, and supervisor health.
|
||||||
|
listen = "127.0.0.1:6902"
|
||||||
|
# Liveness route.
|
||||||
|
liveness_path = "/livez"
|
||||||
|
# Readiness route; false while reconciliation or essential recovery is pending.
|
||||||
|
readiness_path = "/readyz"
|
||||||
|
# Prometheus metrics route.
|
||||||
|
metrics_path = "/metrics"
|
||||||
|
# Structured logging threshold: debug, info, warn, or error.
|
||||||
|
log_level = "info"
|
||||||
|
# Structured log encoding: json or text.
|
||||||
|
log_format = "json"
|
||||||
|
|
||||||
|
# Resource profiles are administrator policy. These illustrative values are not
|
||||||
|
# protocol guarantees. LIGHT is exclusive; otherwise combine at most one tier
|
||||||
|
# from each cpu_*, mem_*, and disk_* family.
|
||||||
|
|
||||||
|
[profiles.light]
|
||||||
|
# Advertise this profile when all required controls validate.
|
||||||
|
enabled = true
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["cpu", "memory", "pids"]
|
||||||
|
# CPU allowance as percent of one logical CPU; zero omits CPU control.
|
||||||
|
cpu_percent = 50
|
||||||
|
# Hard resident/commit memory allowance (512 MiB); zero omits memory control.
|
||||||
|
memory_max_bytes = 536870912
|
||||||
|
# Maximum processes in the supervised tree; zero omits process-count control.
|
||||||
|
pids_max = 64
|
||||||
|
# Windows whole-Job read rate; zero omits it.
|
||||||
|
windows_io_read_bps = 0
|
||||||
|
# Windows whole-Job write rate; zero omits it.
|
||||||
|
windows_io_write_bps = 0
|
||||||
|
# Linux cgroup read rates keyed by device major:minor.
|
||||||
|
linux_io_read_bps = {}
|
||||||
|
# Linux cgroup write rates keyed by device major:minor.
|
||||||
|
linux_io_write_bps = {}
|
||||||
|
|
||||||
|
[profiles.cpu_medium]
|
||||||
|
# Advertise the CPU_MEDIUM allowance class.
|
||||||
|
enabled = true
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["cpu"]
|
||||||
|
# CPU allowance as percent of one logical CPU.
|
||||||
|
cpu_percent = 200
|
||||||
|
|
||||||
|
[profiles.cpu_heavy]
|
||||||
|
# Advertise the CPU_HEAVY allowance class.
|
||||||
|
enabled = true
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["cpu"]
|
||||||
|
# CPU allowance as percent of one logical CPU.
|
||||||
|
cpu_percent = 800
|
||||||
|
|
||||||
|
[profiles.mem_medium]
|
||||||
|
# Advertise the MEM_MEDIUM allowance class.
|
||||||
|
enabled = true
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["memory"]
|
||||||
|
# Hard memory allowance (2 GiB).
|
||||||
|
memory_max_bytes = 2147483648
|
||||||
|
|
||||||
|
[profiles.mem_heavy]
|
||||||
|
# Advertise the MEM_HEAVY allowance class.
|
||||||
|
enabled = true
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["memory"]
|
||||||
|
# Hard memory allowance (8 GiB).
|
||||||
|
memory_max_bytes = 8589934592
|
||||||
|
|
||||||
|
[profiles.disk_medium]
|
||||||
|
# Advertise DISK_MEDIUM only after device/rate controls validate.
|
||||||
|
enabled = false
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["io"]
|
||||||
|
# Windows whole-Job read bandwidth allowance (100 MiB/s).
|
||||||
|
windows_io_read_bps = 104857600
|
||||||
|
# Windows whole-Job write bandwidth allowance (50 MiB/s).
|
||||||
|
windows_io_write_bps = 52428800
|
||||||
|
# Linux cgroup read allowances; replace 8:0 with an actual delegated device.
|
||||||
|
linux_io_read_bps = { "8:0" = 104857600 }
|
||||||
|
# Linux cgroup write allowances; replace 8:0 with an actual delegated device.
|
||||||
|
linux_io_write_bps = { "8:0" = 52428800 }
|
||||||
|
|
||||||
|
[profiles.disk_heavy]
|
||||||
|
# Advertise DISK_HEAVY only after device/rate controls validate.
|
||||||
|
enabled = false
|
||||||
|
# Controls that must be applied atomically or the request is rejected.
|
||||||
|
required_controls = ["io"]
|
||||||
|
# Windows whole-Job read bandwidth allowance (500 MiB/s).
|
||||||
|
windows_io_read_bps = 524288000
|
||||||
|
# Windows whole-Job write bandwidth allowance (250 MiB/s).
|
||||||
|
windows_io_write_bps = 262144000
|
||||||
|
# Linux cgroup read allowances; replace 8:0 with an actual delegated device.
|
||||||
|
linux_io_read_bps = { "8:0" = 524288000 }
|
||||||
|
# Linux cgroup write allowances; replace 8:0 with an actual delegated device.
|
||||||
|
linux_io_write_bps = { "8:0" = 262144000 }
|
||||||
@@ -0,0 +1,109 @@
|
|||||||
|
# RVBox v1 server example. Integer sizes are bytes; durations are quoted strings.
|
||||||
|
|
||||||
|
[server]
|
||||||
|
# Durable SQLite, command payload, output, audit, and incident root.
|
||||||
|
data_dir = "/var/lib/rvbox-server"
|
||||||
|
# HTTP listener receiving WebSocket upgrades from nginx.
|
||||||
|
agent_listen = "127.0.0.1:6899"
|
||||||
|
# Exact WebSocket request path accepted on agent_listen.
|
||||||
|
agent_path = "/v1/agent"
|
||||||
|
# Local gRPC control socket used by rvc; RVBox forces mode 0600.
|
||||||
|
control_socket = "/run/rvbox/server.sock"
|
||||||
|
# Grace allowed for daemon workers to finish reserved closeout on shutdown.
|
||||||
|
shutdown_grace = "30s"
|
||||||
|
|
||||||
|
[json_rpc]
|
||||||
|
# Enable the intentionally unauthenticated JSON-RPC debugging adapter.
|
||||||
|
enabled = false
|
||||||
|
# HTTP bind for JSON-RPC; non-loopback use emits a prominent warning.
|
||||||
|
listen = "127.0.0.1:6900"
|
||||||
|
|
||||||
|
[queue]
|
||||||
|
# Default wait for server-confirmed client acceptance; "0s" means indefinite.
|
||||||
|
default_ttl = "15m"
|
||||||
|
# Maximum server-side queued commands for one client before rejection.
|
||||||
|
max_per_client = 1000
|
||||||
|
# Maximum server-side queued commands across all clients before rejection.
|
||||||
|
max_server = 10000
|
||||||
|
# Initial delay after a transient client rejection before redispatch.
|
||||||
|
retry_initial = "1s"
|
||||||
|
# Maximum full-jitter delay after repeated transient client rejection.
|
||||||
|
retry_max = "30s"
|
||||||
|
|
||||||
|
[storage]
|
||||||
|
# Maximum rolling compressed output retained for one command (10 MiB).
|
||||||
|
command_output_limit_bytes = 10485760
|
||||||
|
# Maximum charged data, including metadata/script/stdin/output, per command (32 MiB).
|
||||||
|
command_total_limit_bytes = 33554432
|
||||||
|
# Maximum charged command data for one target client on the server (256 MiB).
|
||||||
|
client_total_limit_bytes = 268435456
|
||||||
|
# Maximum charged command data across the server, excluding audit/tombstones (4 GiB).
|
||||||
|
server_total_limit_bytes = 4294967296
|
||||||
|
# Reclaim whole terminal commands after this age; "0s" disables age rotation.
|
||||||
|
terminal_retention = "720h"
|
||||||
|
# Separate compressed audit plus resolved-incident history budget (100 MiB).
|
||||||
|
audit_limit_bytes = 104857600
|
||||||
|
# Optional audit age rotation; "0s" keeps entries until the byte budget rolls them.
|
||||||
|
audit_retention = "0s"
|
||||||
|
# Compact global completed-command replay ledger entry cap.
|
||||||
|
tombstone_max_entries = 1000000
|
||||||
|
# Per-active-command quota held for terminal and loss metadata (64 KiB).
|
||||||
|
command_closeout_reserve_bytes = 65536
|
||||||
|
# Reject new unreserved writes below this filesystem free space (256 MiB).
|
||||||
|
free_space_floor_bytes = 268435456
|
||||||
|
# Target size for a sealed append-only payload/output segment (256 KiB).
|
||||||
|
segment_target_bytes = 262144
|
||||||
|
# Maximum time a group-commit waits before fsync; acknowledgements wait too.
|
||||||
|
durability_interval = "100ms"
|
||||||
|
# SQLite busy timeout before an operation returns a transient error.
|
||||||
|
sqlite_busy_timeout = "5s"
|
||||||
|
# Maximum operator incident acknowledgement note encoded as UTF-8 (4 KiB).
|
||||||
|
incident_note_max_bytes = 4096
|
||||||
|
# Maximum protocol error/detail/reason text encoded as UTF-8 (4 KiB).
|
||||||
|
protocol_detail_max_bytes = 4096
|
||||||
|
|
||||||
|
[flow]
|
||||||
|
# Stop admitting raw server validation/persistence work at this backlog (64 MiB).
|
||||||
|
raw_output_high_bytes = 67108864
|
||||||
|
# Leave raw-output loss mode only below this backlog (32 MiB).
|
||||||
|
raw_output_low_bytes = 33554432
|
||||||
|
# Maximum assigned but unacknowledged encoded bytes per command (1 MiB).
|
||||||
|
unacknowledged_per_command_bytes = 1048576
|
||||||
|
# Maximum assigned but unacknowledged encoded bytes per client session (8 MiB).
|
||||||
|
unacknowledged_per_session_bytes = 8388608
|
||||||
|
# Deadline for an individual WebSocket data-frame write.
|
||||||
|
write_deadline = "10s"
|
||||||
|
|
||||||
|
[protocol]
|
||||||
|
# Send WebSocket Ping after this period without inbound activity.
|
||||||
|
heartbeat_idle = "10s"
|
||||||
|
# Close a session after this total period without inbound activity.
|
||||||
|
liveness_timeout = "30s"
|
||||||
|
# One-shot live client-instance collision override lifetime.
|
||||||
|
takeover_ttl = "5m"
|
||||||
|
# Hard decoded AgentEnvelope ceiling (1 MiB).
|
||||||
|
max_agent_envelope_bytes = 1048576
|
||||||
|
# Hard serialized ExecutionSpec ceiling (768 KiB).
|
||||||
|
max_execution_spec_bytes = 786432
|
||||||
|
# Hard uncompressed stdout/stderr chunk ceiling (64 KiB).
|
||||||
|
max_raw_chunk_bytes = 65536
|
||||||
|
# Hard raw uploaded-script ceiling (10 MiB).
|
||||||
|
max_script_bytes = 10485760
|
||||||
|
# Hard decoded local gRPC request ceiling (16 MiB).
|
||||||
|
max_control_request_bytes = 16777216
|
||||||
|
# Hard JSON-RPC HTTP body ceiling including base64 expansion (24 MiB).
|
||||||
|
max_json_rpc_body_bytes = 25165824
|
||||||
|
|
||||||
|
[observability]
|
||||||
|
# Loopback HTTP listener for liveness, readiness, and Prometheus metrics.
|
||||||
|
listen = "127.0.0.1:6901"
|
||||||
|
# Liveness route, available before asynchronous storage recovery completes.
|
||||||
|
liveness_path = "/livez"
|
||||||
|
# Readiness route; false while required storage scopes are unavailable.
|
||||||
|
readiness_path = "/readyz"
|
||||||
|
# Prometheus metrics route.
|
||||||
|
metrics_path = "/metrics"
|
||||||
|
# Structured logging threshold: debug, info, warn, or error.
|
||||||
|
log_level = "info"
|
||||||
|
# Structured log encoding: json or text.
|
||||||
|
log_format = "json"
|
||||||
File diff suppressed because it is too large
Load Diff
+145
-16
@@ -1,5 +1,9 @@
|
|||||||
# RVBox v1 platform and operations contract
|
# RVBox v1 platform and operations contract
|
||||||
|
|
||||||
|
Daemon configuration uses strict TOML as specified in
|
||||||
|
[`configuration.md`](configuration.md); the annotated examples contain every
|
||||||
|
v1 knob and default.
|
||||||
|
|
||||||
## Unix-like clients
|
## Unix-like clients
|
||||||
|
|
||||||
The client starts `sh` or `bash` in a new session/process group. Unix signals
|
The client starts `sh` or `bash` in a new session/process group. Unix signals
|
||||||
@@ -8,20 +12,99 @@ orderly shutdown, or recovery after an unclean daemon failure, managed command
|
|||||||
groups are terminated and marked interrupted because pipe capture cannot be
|
groups are terminated and marked interrupted because pipe capture cannot be
|
||||||
safely resumed.
|
safely resumed.
|
||||||
|
|
||||||
|
Both command text and uploaded scripts execute from generated private files
|
||||||
|
beneath the effective CWD using exactly the selected executable (`sh FILE` or
|
||||||
|
`bash FILE`). No user-supplied filename becomes a filesystem path. The wrapper
|
||||||
|
file is removed during terminal cleanup.
|
||||||
|
|
||||||
|
Root-process exit begins a configurable 5-second drain grace period. RVBox waits
|
||||||
|
for the supervised tree and capture pipes, then terminates residual group/cgroup
|
||||||
|
members, drains to EOF, and only afterward emits the terminal lifecycle event.
|
||||||
|
If capture still cannot reach EOF, it closes the handles and emits explicit
|
||||||
|
incomplete-output metadata first. Shell-level detachment is not a supported way
|
||||||
|
to leave descendants running; callers use RVBox background mode instead.
|
||||||
|
|
||||||
|
Launch uses an internal blocked launcher rather than starting requested command
|
||||||
|
code directly. The launcher establishes its session/process group, reports its
|
||||||
|
identity, and waits on a private release/watchdog channel. The client durably
|
||||||
|
records `launch_prepared`, then durably records `launch_authorized`, and only
|
||||||
|
then sends the release token. The launcher creates the requested shell inside
|
||||||
|
that group and remains as a non-user-code watchdog until the tree exits. The
|
||||||
|
daemon keeps the channel open for that lifetime: EOF before authorization exits
|
||||||
|
without execution, while EOF after release terminates the group. Once
|
||||||
|
`launch_authorized` is durable, recovery never retries that UUID; an uncertain
|
||||||
|
launch is marked interrupted.
|
||||||
|
|
||||||
|
On Linux, create a per-command cgroup v2 for supervision even when no resource
|
||||||
|
profile was requested, whenever the daemon has a delegated writable cgroup.
|
||||||
|
Put the blocked launcher into that cgroup before release; use
|
||||||
|
`clone3(CLONE_INTO_CGROUP | CLONE_PIDFD)` where available, otherwise migrate the
|
||||||
|
still-blocked launcher through `cgroup.procs`. Persist the cgroup path, PID,
|
||||||
|
process group, `/proc/<pid>/stat` start time, and launch generation. A live
|
||||||
|
daemon uses the pidfd where available. Recovery uses `cgroup.kill` as the primary
|
||||||
|
tree-cleanup operation and verifies the recorded birth identity before any
|
||||||
|
PID/process-group fallback. Without cgroup delegation it uses the generic
|
||||||
|
watchdog/process-group fallback unless a requested profile requires cgroup
|
||||||
|
controls, in which case acceptance fails as unsupported. It never signals a
|
||||||
|
process based only on a persisted numeric PID or PGID.
|
||||||
|
|
||||||
|
Other Unix-like systems use the same launch barrier plus a watchdog control
|
||||||
|
channel whose EOF triggers process-group termination. Recovery validates the
|
||||||
|
platform's process-birth identity before signaling. Descendants that deliberately
|
||||||
|
create a new session may escape this generic fallback, so complete tree cleanup
|
||||||
|
outside Linux cgroup supervision is best-effort; the at-most-once launch
|
||||||
|
guarantee still applies.
|
||||||
|
|
||||||
Linux diagnostics sample `/proc/<pid>` and relevant children for state, CPU,
|
Linux diagnostics sample `/proc/<pid>` and relevant children for state, CPU,
|
||||||
resident memory, I/O counters, CWD, and wait-channel information when readable.
|
resident memory, I/O counters, CWD, and wait-channel information when readable.
|
||||||
These values may be unavailable due to permissions, kernel configuration, or a
|
These values may be unavailable due to permissions, kernel configuration, or a
|
||||||
short-lived process; absence is represented explicitly rather than fabricated.
|
short-lived process; absence is represented explicitly rather than fabricated.
|
||||||
Cgroup v2 is used for requested resource profiles only when available.
|
Cgroup v2 profile limits are applied only when requested; a no-profile
|
||||||
|
supervisory cgroup imposes no resource limit.
|
||||||
|
|
||||||
## Windows clients
|
## Windows clients
|
||||||
|
|
||||||
The client launches `cmd` or `powershell` in an appropriate dedicated console
|
The minimum supported v1 Windows versions are Windows 10 and Windows Server
|
||||||
process group and assigns the root process to a per-command Job Object. Child
|
2016. For every command, create a non-inheritable Job Object, set
|
||||||
processes normally join the Job Object. Job Object limits enforce requested
|
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`, and do not enable breakaway. The daemon
|
||||||
profiles and `KILL_ON_JOB_CLOSE` protects against lost supervision.
|
starts an RVBox per-command launcher suspended with `CREATE_NEW_CONSOLE`,
|
||||||
|
`CREATE_UNICODE_ENVIRONMENT`, and `EXTENDED_STARTUPINFO_PRESENT`; it assigns the
|
||||||
|
launcher atomically through `PROC_THREAD_ATTRIBUTE_JOB_LIST`. Only an explicit
|
||||||
|
standard-I/O and launcher-control handle list is inherited, and the Job handle
|
||||||
|
is never inherited.
|
||||||
|
|
||||||
Only `SIGTERM` and `SIGKILL` are accepted. `SIGTERM` attempts `CTRL_BREAK_EVENT`
|
The launcher invokes exactly the selected shell against the generated wrapper:
|
||||||
|
`cmd.exe /D /S /C` for a `.cmd` wrapper, or `powershell.exe` with `-NoLogo`,
|
||||||
|
`-NoProfile`, `-NonInteractive`, and `-File` for a `.ps1` wrapper. Application
|
||||||
|
paths and argument quoting are constructed by the Windows launcher, never by
|
||||||
|
concatenating an untrusted command line. There is no fallback between shells.
|
||||||
|
|
||||||
|
Persist and flush `launch_prepared` with the launcher PID,
|
||||||
|
`GetProcessTimes` creation `FILETIME`, and launch generation. Persist and flush
|
||||||
|
`launch_authorized` before calling `ResumeThread`. The launcher then starts the
|
||||||
|
requested `cmd` or `powershell` suspended in its console with
|
||||||
|
`CREATE_NEW_PROCESS_GROUP`, reports the shell PID/group through the private
|
||||||
|
control channel, connects the allowlisted pipes, and resumes it. This two-step
|
||||||
|
shape is required because `CREATE_NEW_PROCESS_GROUP` is ignored when combined
|
||||||
|
with `CREATE_NEW_CONSOLE`, and console control events reach only groups sharing
|
||||||
|
the caller's console. Failure after authorization terminates the Job and is
|
||||||
|
reported as interrupted; it never redispatches the UUID.
|
||||||
|
|
||||||
|
The launcher remains the in-console signal proxy and calls
|
||||||
|
`GenerateConsoleCtrlEvent(CTRL_BREAK_EVENT, shell_group_id)` on request. The
|
||||||
|
daemon retains the sole Job handle, so an unclean daemon exit closes the last
|
||||||
|
handle and terminates the launcher, shell, and descendants. Recovery never
|
||||||
|
kills by persisted PID alone; the PID/creation-time tuple is diagnostic evidence
|
||||||
|
for PID reuse or cleanup anomalies. Child processes normally join the Job.
|
||||||
|
Job Object limits enforce requested profiles and `KILL_ON_JOB_CLOSE` protects
|
||||||
|
against lost supervision.
|
||||||
|
|
||||||
|
Root-process exit begins the same drain grace period. Completion waits for the
|
||||||
|
Job Object to reach zero active processes; after the grace period RVBox
|
||||||
|
terminates the Job, drains its capture handles, records any incomplete-output
|
||||||
|
marker, and emits the terminal lifecycle event last.
|
||||||
|
|
||||||
|
Only `TERM`/`SIGTERM` and `KILL`/`SIGKILL` are accepted. `TERM` attempts `CTRL_BREAK_EVENT`
|
||||||
and waits 10 seconds, then calls Job Object termination if the job persists;
|
and waits 10 seconds, then calls Job Object termination if the job persists;
|
||||||
`SIGKILL` calls Job Object termination immediately. A console signal is
|
`SIGKILL` calls Job Object termination immediately. A console signal is
|
||||||
best-effort, so callers receive an explicit escalation result. Windows status
|
best-effort, so callers receive an explicit escalation result. Windows status
|
||||||
@@ -30,12 +113,36 @@ diagnostics such as an I/O wait channel.
|
|||||||
|
|
||||||
## Storage and recovery
|
## Storage and recovery
|
||||||
|
|
||||||
SQLite runs in WAL mode with integrity checking on startup. Output segments are
|
SQLite runs in WAL mode. Every append-only segment has a SQLite-owned
|
||||||
written atomically, fsynced according to the configured durability interval, and
|
`committed_end_offset`. The writer validates and appends records, syncs the file
|
||||||
indexed only after successful durable append. Startup scans/repairs incomplete
|
(grouped by the configured durability interval), and only then commits event
|
||||||
tail records before accepting control requests. Segment compression is Zstandard;
|
metadata plus the new offset in SQLite. An acknowledgement waits for both
|
||||||
limits always measure stored compressed bytes, while clients expose raw byte
|
steps. Therefore a crash can leave an uncommitted file tail, but cannot validly
|
||||||
counts separately.
|
acknowledge metadata whose bytes were not durable.
|
||||||
|
|
||||||
|
Startup acquires the instance lock and binds liveness/diagnostic endpoints, then
|
||||||
|
runs integrity and segment recovery asynchronously. A file longer than its
|
||||||
|
committed offset is safely truncated to that offset. A file shorter than the
|
||||||
|
offset, a checksum failure inside the committed range, or corrupt essential
|
||||||
|
metadata creates a durable scoped storage incident; affected output is marked
|
||||||
|
truncated/incomplete and affected active commands are interrupted when their
|
||||||
|
essential state cannot be trusted. Healthy scopes remain usable. Readiness is
|
||||||
|
false and mutations requiring an unrecovered or dirty scope return `UNAVAILABLE`,
|
||||||
|
but process startup, liveness, incident inspection, and unaffected work do not
|
||||||
|
wait for a full-store scan.
|
||||||
|
|
||||||
|
Safe repairs are attempted automatically and can also be requested online with
|
||||||
|
`rvc storage repair`. Irrecoverable loss stays dirty until explicitly accepted
|
||||||
|
with `rvc storage acknowledge`; an offline server has equivalent
|
||||||
|
`rvbox-server repair --data-dir ...` repair/list/acknowledge operations. Client
|
||||||
|
spool recovery follows the same committed-offset rule and exposes equivalent
|
||||||
|
offline `rvbox repair --state-dir ...` operations and local health diagnostics.
|
||||||
|
Resolving an incident clears derived dirty health but never erases the incident
|
||||||
|
or audit history as part of resolution. Unresolved compact incident records are
|
||||||
|
non-evictable; resolved incident/audit history follows the separate 100 MiB
|
||||||
|
rotation. Segment compression is Zstandard;
|
||||||
|
limits measure stored compressed bytes, while raw byte counts are reported
|
||||||
|
separately.
|
||||||
|
|
||||||
The system must reserve headroom before writes and use transactional metadata
|
The system must reserve headroom before writes and use transactional metadata
|
||||||
updates. Storage-full, permission, and corruption failures are surfaced as
|
updates. Storage-full, permission, and corruption failures are surfaced as
|
||||||
@@ -43,6 +150,12 @@ structured server/client health states and audit events. They must isolate the
|
|||||||
affected command/session, reject work when needed, and keep the daemon's
|
affected command/session, reject work when needed, and keep the daemon's
|
||||||
heartbeat/control loops alive.
|
heartbeat/control loops alive.
|
||||||
|
|
||||||
|
V1 state is plaintext at rest, including command text, scripts, environment
|
||||||
|
override values, stdin, and output. Private directory/file modes and dedicated
|
||||||
|
daemon accounts are deployment hygiene, not an application-level encryption
|
||||||
|
guarantee. Backups copy the same plaintext sensitivity. Encryption and external
|
||||||
|
key management are future-version work.
|
||||||
|
|
||||||
## Metrics, logging, and safe defaults
|
## Metrics, logging, and safe defaults
|
||||||
|
|
||||||
Both daemons should emit structured logs and metrics for session transitions,
|
Both daemons should emit structured logs and metrics for session transitions,
|
||||||
@@ -52,10 +165,26 @@ and protocol violations. Never emit stdin or raw output in normal daemon logs.
|
|||||||
|
|
||||||
Recommended configuration defaults are: 10-second heartbeat idle period,
|
Recommended configuration defaults are: 10-second heartbeat idle period,
|
||||||
30-second liveness timeout, 1–60-second full-jitter reconnect backoff,
|
30-second liveness timeout, 1–60-second full-jitter reconnect backoff,
|
||||||
60-second stable-session reset, 16 running/100 queued commands per client,
|
60-second stable-session reset, 5-minute one-shot live-conflict takeover grant,
|
||||||
10 MiB per-command compressed window, 50 MiB per-client active spool and server
|
16 running/100 queued commands per client, 1,000 server-queued commands per
|
||||||
history, 1 GiB server history, 64 KiB uncompressed stream chunk, 1 MiB decoded
|
target client and 10,000 server-wide,
|
||||||
envelope, and 10 MiB script maximum.
|
15-minute queue TTL, 10 MiB per-command output window, 32 MiB total per command,
|
||||||
|
256 MiB per client/client daemon, 4 GiB server-wide command storage, 30-day
|
||||||
|
terminal retention, 100 MiB audit storage with quota-only rotation by default,
|
||||||
|
one million compact command tombstones, 64 KiB uncompressed stream chunk, 1 MiB
|
||||||
|
decoded agent envelope,
|
||||||
|
1 MiB/256 KiB per-command raw-output high/low watermarks, 8 MiB/4 MiB per-client
|
||||||
|
watermarks, 64 MiB/32 MiB server-wide watermarks, 1 MiB per-command and 8 MiB
|
||||||
|
per-session unacknowledged send windows, and 10 MiB raw script maximum. Control
|
||||||
|
gRPC accepts at most 16 MiB decoded requests; the JSON-RPC adapter accepts at
|
||||||
|
most 24 MiB HTTP bodies to allow protobuf JSON's base64 expansion while
|
||||||
|
retaining the same decoded field limits. Each active command reserves 64 KiB
|
||||||
|
within its quota for terminal/loss closeout metadata; protocol detail/reason and
|
||||||
|
incident-note text fields are individually limited to 4 KiB. Default emergency
|
||||||
|
filesystem free-space floors are 256 MiB on the server and 64 MiB on a client;
|
||||||
|
crossing one rejects new unreserved allocations even if the logical quota has
|
||||||
|
headroom. Already-reserved terminal/loss closeout remains writable while bytes
|
||||||
|
physically remain.
|
||||||
|
|
||||||
These bounds protect RVBox's own loops; they cannot make arbitrary child
|
These bounds protect RVBox's own loops; they cannot make arbitrary child
|
||||||
commands harmless when no resource profile is requested. Operators should
|
commands harmless when no resource profile is requested. Operators should
|
||||||
|
|||||||
+101
-31
@@ -14,8 +14,10 @@ sizes before allocation/decompression.
|
|||||||
|
|
||||||
The protobuf package is `rvbox.v1`. Registration negotiates a major/minor
|
The protobuf package is `rvbox.v1`. Registration negotiates a major/minor
|
||||||
protocol range: incompatible majors are rejected; the highest shared minor is
|
protocol range: incompatible majors are rejected; the highest shared minor is
|
||||||
chosen. New fields are append-only. A peer ignores unknown optional fields but
|
chosen. New fields are append-only. A peer ignores unknown optional fields. A
|
||||||
must respond with `PROTOCOL_ERROR` to an unknown required envelope feature.
|
new required behavior requires a negotiated minor-version change; an envelope
|
||||||
|
with no recognized payload is a `PROTOCOL_ERROR` rather than an implicit
|
||||||
|
“required unknown field” mechanism that protobuf cannot represent.
|
||||||
|
|
||||||
## Heartbeat and reconnect
|
## Heartbeat and reconnect
|
||||||
|
|
||||||
@@ -33,11 +35,19 @@ reconciles only commands the server still considers non-terminal.
|
|||||||
|
|
||||||
## Session fencing
|
## Session fencing
|
||||||
|
|
||||||
`ClientHello` starts registration. `ServerWelcome` gives the selected version,
|
`ClientHello` starts registration and includes a random `client_instance_id`
|
||||||
random `session_id`, and `session_generation`. Except `ClientHello`, all
|
generated once in the client state directory and retained across daemon restarts
|
||||||
envelopes carry those values. A new accepted registration fences and disconnects
|
and reconnects. `ServerWelcome` gives the selected version, random `session_id`,
|
||||||
the prior one for the same client ID. The server accepts messages only from the
|
and `session_generation`. Except `ClientHello`, all envelopes carry those
|
||||||
current generation; command dispatches also identify their intended generation.
|
values. A new accepted registration from the same instance fences and
|
||||||
|
disconnects its prior session. A different instance claiming the same live
|
||||||
|
client ID is rejected and recorded as the current pending claim unless an
|
||||||
|
operator explicitly authorized that exact instance. The one-shot authorization
|
||||||
|
expires after 5 minutes by default and is consumed by the matching reconnect.
|
||||||
|
When no session for that client ID is live, a new instance is accepted normally.
|
||||||
|
This is collision protection, not peer authentication. The
|
||||||
|
server accepts messages only from the current generation; command dispatches
|
||||||
|
also identify their intended generation.
|
||||||
|
|
||||||
## Reliable work flows
|
## Reliable work flows
|
||||||
|
|
||||||
@@ -45,53 +55,113 @@ current generation; command dispatches also identify their intended generation.
|
|||||||
|
|
||||||
The server persistently creates a command before dispatching `CommandDispatch`.
|
The server persistently creates a command before dispatching `CommandDispatch`.
|
||||||
It redelivers until it receives `CommandAccepted`. Clients durably deduplicate
|
It redelivers until it receives `CommandAccepted`. Clients durably deduplicate
|
||||||
on `issue_uuid`. A client whose queue is full sends a capacity rejection.
|
on `issue_uuid`. A client whose queue is full sends a transient capacity
|
||||||
|
rejection, which the server requeues with backoff. A permanent validation or
|
||||||
|
unsupported-platform rejection makes the server command terminal `rejected`;
|
||||||
|
it is not retried or mislabeled as a launched-process failure.
|
||||||
|
An exact UUID/request-hash hit in the compact tombstone ledger returns
|
||||||
|
`CODE_ALREADY_EXECUTED`; the server suppresses dispatch and reconciles its stale
|
||||||
|
state instead of representing the prior execution as a new rejection.
|
||||||
|
|
||||||
Following registration the server sends `ReconcileRequest` only for its
|
Following registration the server sends `ReconcileRequest` containing its
|
||||||
non-terminal commands for that client. The client replies with a
|
non-terminal commands for that client. The client replies once with a complete,
|
||||||
`ReconcileSnapshot` per requested known command, or `unknown_to_client`. It
|
bounded `ReconcileSnapshot` of every command it still retains, including
|
||||||
continues normal event retransmission from the server's last acknowledged event
|
terminal-but-unacknowledged records and matching requested tombstones, then
|
||||||
sequence. Terminal history already confirmed by the server is deliberately
|
continues normal retransmission from the server's last acknowledged event
|
||||||
excluded.
|
sequence. The server does not dispatch new work until this snapshot is complete.
|
||||||
|
For healthy client storage, absence means the client never durably accepted the
|
||||||
|
UUID: server `queued`/`dispatched` work returns to `queued`, while absence of an
|
||||||
|
`accepted`/`running` record is an invariant failure and becomes interrupted.
|
||||||
|
A matching UUID/request-hash tombstone always suppresses replay. A client-known
|
||||||
|
server-missing active command, or a contradiction with server-confirmed
|
||||||
|
terminal history, is terminated locally after the server returns a durable
|
||||||
|
`ReconcileResult` and creates a recovery incident rather than inventing server
|
||||||
|
state. That result also identifies client-retained terminal records the server
|
||||||
|
has already stored or deliberately tombstoned, allowing the client to discard
|
||||||
|
them even when its last `EventAck` was lost. The result is idempotent and is
|
||||||
|
written before any new dispatch on that session. A client
|
||||||
|
with an unresolved essential-store incident must not claim a complete snapshot
|
||||||
|
or accept work until repaired/acknowledged.
|
||||||
|
|
||||||
### Output and events
|
### Output and events
|
||||||
|
|
||||||
Client execution events use an increasing `event_seq`; retries reuse the same
|
Client execution entries first have a durable local order. They receive an
|
||||||
sequence and content. The server durably writes an event before sending
|
increasing wire `event_seq` durably when admitted to the bounded send window;
|
||||||
`EventAck`. `EventAck` is cumulative through a sequence number. Output is
|
once assigned, the sequence/content is pinned until acknowledged and retries
|
||||||
Zstandard compressed with an explicit original-size field. Server storage can
|
reuse it exactly. The server durably writes an event before sending `EventAck`.
|
||||||
reuse the validated compressed bytes.
|
`EventAck` is cumulative through a sequence number. Output is Zstandard
|
||||||
|
compressed with an explicit original-size field. Server storage can reuse the
|
||||||
|
validated compressed bytes.
|
||||||
|
|
||||||
When rolling output removes old retained segments, the server writes
|
When rolling output removes old retained segments, the server writes
|
||||||
`OutputTruncation` metadata containing the removed event/byte ranges. Queries
|
`OutputTruncation` metadata containing the removed event/byte ranges. Queries
|
||||||
must show that marker rather than silently presenting an apparently complete
|
must show that marker rather than silently presenting an apparently complete
|
||||||
stream. An offline client that reaches a cap sends `ClientOutputTruncated`
|
stream. Assigned but unacknowledged events occupy the pinned 1 MiB per-command
|
||||||
before replaying its retained tail on reconnect. The server records it, permits
|
send window and are not evicted. An offline client that reaches a cap replaces
|
||||||
the named event-sequence gap, and exposes it in output queries. Truncation
|
one or more still-unsequenced output runs in its durable local order with
|
||||||
metadata is not an execution event and does not consume an `event_seq`.
|
`OutputTruncation`; when admitted to the send window the marker receives the
|
||||||
|
next normal `event_seq`. Its event-range fields are absent because the discarded
|
||||||
|
bytes never had wire sequences. The normal cumulative `EventAck` acknowledges
|
||||||
|
the marker. Server-created retention markers include their removed event range
|
||||||
|
and remain query metadata because the server cannot allocate client sequences.
|
||||||
|
|
||||||
|
If a capture pipe cannot be drained to a provably complete EOF, the client emits
|
||||||
|
a sequenced `OutputIncomplete` event before the terminal lifecycle event. It
|
||||||
|
does not invent a missing sequence or byte count for bytes it never observed.
|
||||||
|
|
||||||
### Stdin and signals
|
### Stdin and signals
|
||||||
|
|
||||||
`StdinWrite` is binary-safe and ordered by `write_seq`; the client durably
|
`StdinWrite` is binary-safe and ordered by `write_seq`; the client durably
|
||||||
deduplicates it and returns `StdinAck`. `append_newline` is true by default in
|
deduplicates it and returns a sequenced stdin-acknowledgement command event.
|
||||||
the CLI but explicit on the wire. `CloseStdin` is a separate idempotent action.
|
`append_newline` is true by default in the CLI but explicit on the wire.
|
||||||
Signal and cancellation requests carry a command revision to settle start/kill
|
`CloseStdin` is a separate idempotent action. Signal and cancellation requests
|
||||||
races.
|
carry a command revision to settle start/kill races; the resulting lifecycle or
|
||||||
|
signal-result event echoes that revision. Control-plane `request_id` values stay
|
||||||
|
at the server and map to the assigned revision rather than crossing the agent
|
||||||
|
protocol.
|
||||||
|
|
||||||
### Scripts
|
### Scripts
|
||||||
|
|
||||||
For a script command, dispatch first contains a `ScriptDescriptor`; then the
|
For a script command, dispatch first contains a `ScriptDescriptor`; then the
|
||||||
server sends `ScriptChunk` messages and a commit. The client checks offset,
|
server sends `ScriptChunk` messages and a commit. The client checks offset,
|
||||||
chunk order, full length, and SHA-256 before it reports upload complete or
|
chunk order, full length, and SHA-256 before it reports upload complete or
|
||||||
launches the process. A retransmitted chunk is idempotent by offset/content.
|
launches the process. `ScriptUploadStatus.received_bytes` is a cumulative
|
||||||
|
durable progress acknowledgement: the server sends at most the command's
|
||||||
|
unacknowledged window, waits for progress, and resumes from that offset. A
|
||||||
|
retransmitted chunk is idempotent by offset/content.
|
||||||
|
|
||||||
|
Acceptance is durable queue admission, not proof that every later preparation
|
||||||
|
step will succeed. A permanent script checksum/write error or failure to prepare
|
||||||
|
the requested executable emits terminal `rejected` before any launch
|
||||||
|
authorization. User cancellation in that interval emits `cancelled`; an
|
||||||
|
uncertain crash window emits `interrupted`. None is mislabeled as process
|
||||||
|
`failed`.
|
||||||
|
|
||||||
## Flow control and failure containment
|
## Flow control and failure containment
|
||||||
|
|
||||||
No receive loop runs an executor, database write, decompressor, or slow socket
|
No receive loop runs an executor, database write, decompressor, or slow socket
|
||||||
operation inline. Each side has bounded staging queues. Durable spools are the
|
operation inline. Each side has bounded staging queues. Durable spools are the
|
||||||
source of truth and are charged to the 10 MiB/50 MiB/1 GiB compressed retention
|
source of truth and are charged to the tiered unified command-storage budgets
|
||||||
budgets described in the architecture document. A full staging queue pauses the
|
described in the architecture document. Senders keep a bounded unacknowledged
|
||||||
related read/dispatch path and drains from disk; it never grows without bound.
|
window of encoded wire bytes (defaults: 1 MiB per command and 8 MiB per client
|
||||||
|
session); unsent data
|
||||||
|
waits durably. WebSocket Ping/Pong/Close plus fencing and protocol errors use a
|
||||||
|
reserved priority lane and a dedicated writer with bounded data-frame size and
|
||||||
|
write deadlines. A full essential ingress queue does not stall the sole
|
||||||
|
WebSocket reader indefinitely:
|
||||||
|
the receiver closes without acknowledgement and the durable sender retries
|
||||||
|
after jittered reconnect. Droppable output uses the sequenced overload-loss
|
||||||
|
path instead.
|
||||||
|
|
||||||
|
Before compression, raw output uses the architecture's high/low watermarks.
|
||||||
|
Once a client high watermark is crossed, still-unsequenced droppable chunks are
|
||||||
|
summarized in durable local order instead of consuming unbounded compression
|
||||||
|
work. A server at its ingress high watermark does not acknowledge or discard an
|
||||||
|
already sequenced event; it closes without acknowledgement if its bounded queue
|
||||||
|
cannot admit the frame, and the client retries from durable spool. Fair
|
||||||
|
scheduling prevents one verbose command from monopolizing workers. Reserved
|
||||||
|
metadata capacity remains
|
||||||
|
available to close each gap and emit the final lifecycle event.
|
||||||
|
|
||||||
Malformed protobuf, over-size payload, invalid compressed data, impossible
|
Malformed protobuf, over-size payload, invalid compressed data, impossible
|
||||||
sequence, bad session token, or protocol-version violation yields a structured
|
sequence, bad session token, or protocol-version violation yields a structured
|
||||||
|
|||||||
+39
-31
@@ -2,35 +2,34 @@ syntax = "proto3";
|
|||||||
|
|
||||||
package rvbox.v1;
|
package rvbox.v1;
|
||||||
|
|
||||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
|
||||||
|
|
||||||
import "google/protobuf/timestamp.proto";
|
import "google/protobuf/timestamp.proto";
|
||||||
import "rvbox/v1/common.proto";
|
import "rvbox/v1/common.proto";
|
||||||
|
|
||||||
|
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||||
|
|
||||||
// One AgentEnvelope is carried in one binary WebSocket message.
|
// One AgentEnvelope is carried in one binary WebSocket message.
|
||||||
message AgentEnvelope {
|
message AgentEnvelope {
|
||||||
// Empty only for ClientHello. All other envelopes are fenced to a session.
|
// Empty only for ClientHello. All other envelopes are fenced to a session.
|
||||||
string session_id = 1;
|
string session_id = 1;
|
||||||
uint64 session_generation = 2;
|
uint64 session_generation = 2;
|
||||||
string message_id = 3;
|
|
||||||
|
|
||||||
oneof payload {
|
oneof payload {
|
||||||
ClientHello client_hello = 10;
|
ClientHello client_hello = 3;
|
||||||
ServerWelcome server_welcome = 11;
|
ServerWelcome server_welcome = 4;
|
||||||
CommandDispatch command_dispatch = 12;
|
CommandDispatch command_dispatch = 5;
|
||||||
CommandAccepted command_accepted = 13;
|
CommandAccepted command_accepted = 6;
|
||||||
CommandEvent command_event = 14;
|
CommandEvent command_event = 7;
|
||||||
EventAck event_ack = 15;
|
EventAck event_ack = 8;
|
||||||
StdinWrite stdin_write = 16;
|
StdinWrite stdin_write = 9;
|
||||||
CloseStdin close_stdin = 17;
|
CloseStdin close_stdin = 10;
|
||||||
SignalCommand signal_command = 18;
|
SignalCommand signal_command = 11;
|
||||||
ScriptChunk script_chunk = 19;
|
ScriptChunk script_chunk = 12;
|
||||||
ScriptCommit script_commit = 20;
|
ScriptCommit script_commit = 13;
|
||||||
ClientCapacity client_capacity = 21;
|
ClientCapacity client_capacity = 14;
|
||||||
ReconcileRequest reconcile_request = 22;
|
ReconcileRequest reconcile_request = 15;
|
||||||
ReconcileSnapshot reconcile_snapshot = 23;
|
ReconcileSnapshot reconcile_snapshot = 16;
|
||||||
AgentError error = 24;
|
AgentError error = 17;
|
||||||
ClientOutputTruncated client_output_truncated = 25;
|
ReconcileResult reconcile_result = 18;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -43,7 +42,8 @@ message ClientHello {
|
|||||||
string architecture = 5;
|
string architecture = 5;
|
||||||
string daemon_cwd = 6;
|
string daemon_cwd = 6;
|
||||||
repeated ShellType supported_shells = 7;
|
repeated ShellType supported_shells = 7;
|
||||||
string reconnect_uuid = 8;
|
// Generated once and persisted in the client state directory.
|
||||||
|
string client_instance_id = 8;
|
||||||
uint32 max_running_commands = 9;
|
uint32 max_running_commands = 9;
|
||||||
uint32 max_queued_commands = 10;
|
uint32 max_queued_commands = 10;
|
||||||
google.protobuf.Timestamp sent_at = 11;
|
google.protobuf.Timestamp sent_at = 11;
|
||||||
@@ -61,6 +61,7 @@ message CommandDispatch {
|
|||||||
uint64 target_session_generation = 3;
|
uint64 target_session_generation = 3;
|
||||||
google.protobuf.Timestamp issue_time = 4;
|
google.protobuf.Timestamp issue_time = 4;
|
||||||
ExecutionSpec spec = 5;
|
ExecutionSpec spec = 5;
|
||||||
|
google.protobuf.Timestamp queue_expiry_time = 6;
|
||||||
}
|
}
|
||||||
|
|
||||||
message CommandAccepted {
|
message CommandAccepted {
|
||||||
@@ -122,23 +123,30 @@ message ReconcileTarget {
|
|||||||
string issue_uuid = 1;
|
string issue_uuid = 1;
|
||||||
uint64 last_server_event_seq = 2;
|
uint64 last_server_event_seq = 2;
|
||||||
uint64 command_revision = 3;
|
uint64 command_revision = 3;
|
||||||
|
bytes immutable_request_sha256 = 4;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Sent only for requests in ReconcileRequest; server-confirmed terminal history
|
|
||||||
// is intentionally not requested.
|
|
||||||
message ReconcileSnapshot {
|
message ReconcileSnapshot {
|
||||||
string issue_uuid = 1;
|
// The complete set still retained by the client, including non-terminal and
|
||||||
bool known_to_client = 2;
|
// terminal-but-unacknowledged commands.
|
||||||
CommandLifecycle lifecycle = 3;
|
repeated ReconcileCommandState retained_commands = 1;
|
||||||
uint64 last_client_event_seq = 4;
|
|
||||||
uint64 command_revision = 5;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// Sent before retained replay data when an offline client had to rotate
|
message ReconcileCommandState {
|
||||||
// unacknowledged output to remain within its hard spool limits.
|
|
||||||
message ClientOutputTruncated {
|
|
||||||
string issue_uuid = 1;
|
string issue_uuid = 1;
|
||||||
OutputTruncation truncation = 2;
|
CommandLifecycle lifecycle = 2;
|
||||||
|
uint64 last_client_event_seq = 3;
|
||||||
|
uint64 command_revision = 4;
|
||||||
|
// True only for a matching UUID/request hash in the compact tombstone ledger.
|
||||||
|
bool tombstoned = 5;
|
||||||
|
bytes immutable_request_sha256 = 6;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sent after the server durably compares a complete snapshot and before it
|
||||||
|
// dispatches new work on the session. Repetition is idempotent.
|
||||||
|
message ReconcileResult {
|
||||||
|
repeated string terminate_local_issue_uuids = 1;
|
||||||
|
repeated string discard_local_terminal_issue_uuids = 2;
|
||||||
}
|
}
|
||||||
|
|
||||||
message AgentError {
|
message AgentError {
|
||||||
|
|||||||
@@ -2,11 +2,11 @@ syntax = "proto3";
|
|||||||
|
|
||||||
package rvbox.v1;
|
package rvbox.v1;
|
||||||
|
|
||||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
|
||||||
|
|
||||||
import "google/protobuf/duration.proto";
|
import "google/protobuf/duration.proto";
|
||||||
import "google/protobuf/timestamp.proto";
|
import "google/protobuf/timestamp.proto";
|
||||||
|
|
||||||
|
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||||
|
|
||||||
// An inclusive protocol-version range advertised during registration.
|
// An inclusive protocol-version range advertised during registration.
|
||||||
message ProtocolRange {
|
message ProtocolRange {
|
||||||
uint32 major = 1;
|
uint32 major = 1;
|
||||||
@@ -52,6 +52,8 @@ enum CommandLifecycle {
|
|||||||
COMMAND_TERMINATED = 7;
|
COMMAND_TERMINATED = 7;
|
||||||
COMMAND_CANCELLED = 8;
|
COMMAND_CANCELLED = 8;
|
||||||
COMMAND_INTERRUPTED = 9;
|
COMMAND_INTERRUPTED = 9;
|
||||||
|
COMMAND_EXPIRED = 10;
|
||||||
|
COMMAND_REJECTED = 11;
|
||||||
}
|
}
|
||||||
|
|
||||||
enum StreamKind {
|
enum StreamKind {
|
||||||
@@ -109,16 +111,23 @@ message CommandRecord {
|
|||||||
ExecutionSpec spec = 5;
|
ExecutionSpec spec = 5;
|
||||||
CommandLifecycle lifecycle = 6;
|
CommandLifecycle lifecycle = 6;
|
||||||
uint64 last_event_seq = 7;
|
uint64 last_event_seq = 7;
|
||||||
int32 exit_code = 8;
|
optional int32 exit_code = 8;
|
||||||
google.protobuf.Timestamp terminal_time = 9;
|
google.protobuf.Timestamp terminal_time = 9;
|
||||||
bool output_truncated = 10;
|
bool output_truncated = 10;
|
||||||
uint64 retained_compressed_bytes = 11;
|
uint64 retained_compressed_bytes = 11;
|
||||||
|
google.protobuf.Timestamp queue_expiry_time = 12;
|
||||||
|
bool output_incomplete = 13;
|
||||||
|
// Present when lifecycle is COMMAND_REJECTED.
|
||||||
|
ControlError rejection = 14;
|
||||||
|
uint64 command_revision = 15;
|
||||||
|
bool late_after_expiry = 16;
|
||||||
}
|
}
|
||||||
|
|
||||||
message LifecycleChange {
|
message LifecycleChange {
|
||||||
CommandLifecycle lifecycle = 1;
|
CommandLifecycle lifecycle = 1;
|
||||||
int32 exit_code = 2;
|
optional int32 exit_code = 2;
|
||||||
string detail = 3;
|
string detail = 3;
|
||||||
|
uint64 command_revision = 4;
|
||||||
}
|
}
|
||||||
|
|
||||||
// data is compressed according to compression. uncompressed_size is mandatory
|
// data is compressed according to compression. uncompressed_size is mandatory
|
||||||
@@ -132,15 +141,16 @@ message OutputChunk {
|
|||||||
}
|
}
|
||||||
|
|
||||||
message ResourceSnapshot {
|
message ResourceSnapshot {
|
||||||
uint64 resident_memory_bytes = 1;
|
optional uint64 resident_memory_bytes = 1;
|
||||||
uint64 virtual_memory_bytes = 2;
|
optional uint64 virtual_memory_bytes = 2;
|
||||||
google.protobuf.Duration cpu_time = 3;
|
google.protobuf.Duration cpu_time = 3;
|
||||||
uint64 read_bytes = 4;
|
optional uint64 read_bytes = 4;
|
||||||
uint64 write_bytes = 5;
|
optional uint64 write_bytes = 5;
|
||||||
string process_state = 6;
|
optional string process_state = 6;
|
||||||
string wait_reason = 7;
|
optional string wait_reason = 7;
|
||||||
bool suspected_hung = 8;
|
bool suspected_hung = 8;
|
||||||
string diagnostic_detail = 9;
|
string diagnostic_detail = 9;
|
||||||
|
optional string current_cwd = 10;
|
||||||
}
|
}
|
||||||
|
|
||||||
enum OutputTruncationSource {
|
enum OutputTruncationSource {
|
||||||
@@ -148,19 +158,31 @@ enum OutputTruncationSource {
|
|||||||
OUTPUT_TRUNCATION_SOURCE_CLIENT_SPOOL = 1;
|
OUTPUT_TRUNCATION_SOURCE_CLIENT_SPOOL = 1;
|
||||||
OUTPUT_TRUNCATION_SOURCE_SERVER_COMMAND_WINDOW = 2;
|
OUTPUT_TRUNCATION_SOURCE_SERVER_COMMAND_WINDOW = 2;
|
||||||
OUTPUT_TRUNCATION_SOURCE_SERVER_CLIENT_CAP = 3;
|
OUTPUT_TRUNCATION_SOURCE_SERVER_CLIENT_CAP = 3;
|
||||||
|
OUTPUT_TRUNCATION_SOURCE_CLIENT_OVERLOAD = 4;
|
||||||
|
OUTPUT_TRUNCATION_SOURCE_SERVER_GLOBAL_CAP = 5;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Persistent query metadata for a missing contiguous event range. It is not a
|
// Metadata for missing output. Server-side retention supplies the optional
|
||||||
// CommandEvent and therefore does not consume an event_seq.
|
// event range. Client output discarded before wire sequence assignment omits
|
||||||
|
// the range and carries this as a normally sequenced CommandEvent.
|
||||||
message OutputTruncation {
|
message OutputTruncation {
|
||||||
uint64 first_removed_event_seq = 1;
|
optional uint64 first_removed_event_seq = 1;
|
||||||
uint64 last_removed_event_seq = 2;
|
optional uint64 last_removed_event_seq = 2;
|
||||||
uint64 removed_compressed_bytes = 3;
|
// Absent when bytes were discarded before compression.
|
||||||
|
optional uint64 removed_compressed_bytes = 3;
|
||||||
uint64 removed_uncompressed_bytes = 4;
|
uint64 removed_uncompressed_bytes = 4;
|
||||||
string reason = 5;
|
string reason = 5;
|
||||||
OutputTruncationSource source = 6;
|
OutputTruncationSource source = 6;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Indicates that capture ended without a provably complete byte stream. It is
|
||||||
|
// sequenced before the terminal lifecycle event but does not claim an invented
|
||||||
|
// event or byte range.
|
||||||
|
message OutputIncomplete {
|
||||||
|
repeated StreamKind streams = 1;
|
||||||
|
string reason = 2;
|
||||||
|
}
|
||||||
|
|
||||||
message StdinAcknowledgement {
|
message StdinAcknowledgement {
|
||||||
uint64 write_seq = 1;
|
uint64 write_seq = 1;
|
||||||
bool stdin_closed = 2;
|
bool stdin_closed = 2;
|
||||||
@@ -173,6 +195,7 @@ message SignalResult {
|
|||||||
bool graceful_delivery_attempted = 3;
|
bool graceful_delivery_attempted = 3;
|
||||||
bool forced_termination_used = 4;
|
bool forced_termination_used = 4;
|
||||||
string detail = 5;
|
string detail = 5;
|
||||||
|
uint64 command_revision = 6;
|
||||||
}
|
}
|
||||||
|
|
||||||
message ScriptUploadStatus {
|
message ScriptUploadStatus {
|
||||||
@@ -188,12 +211,14 @@ message CommandEvent {
|
|||||||
uint64 event_seq = 2;
|
uint64 event_seq = 2;
|
||||||
google.protobuf.Timestamp observed_at = 3;
|
google.protobuf.Timestamp observed_at = 3;
|
||||||
oneof payload {
|
oneof payload {
|
||||||
LifecycleChange lifecycle = 10;
|
LifecycleChange lifecycle = 4;
|
||||||
OutputChunk output = 11;
|
OutputChunk output = 5;
|
||||||
ResourceSnapshot resource = 12;
|
ResourceSnapshot resource = 6;
|
||||||
StdinAcknowledgement stdin_ack = 13;
|
StdinAcknowledgement stdin_ack = 7;
|
||||||
SignalResult signal_result = 14;
|
SignalResult signal_result = 8;
|
||||||
ScriptUploadStatus script_status = 15;
|
ScriptUploadStatus script_status = 9;
|
||||||
|
OutputTruncation output_truncation = 10;
|
||||||
|
OutputIncomplete output_incomplete = 11;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -209,6 +234,7 @@ message ControlError {
|
|||||||
PROTOCOL_ERROR = 7;
|
PROTOCOL_ERROR = 7;
|
||||||
TRANSIENT = 8;
|
TRANSIENT = 8;
|
||||||
INTERNAL = 9;
|
INTERNAL = 9;
|
||||||
|
CODE_ALREADY_EXECUTED = 10;
|
||||||
}
|
}
|
||||||
Code code = 1;
|
Code code = 1;
|
||||||
string message = 2;
|
string message = 2;
|
||||||
|
|||||||
+112
-15
@@ -2,22 +2,27 @@ syntax = "proto3";
|
|||||||
|
|
||||||
package rvbox.v1;
|
package rvbox.v1;
|
||||||
|
|
||||||
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
import "google/protobuf/duration.proto";
|
||||||
|
import "google/protobuf/timestamp.proto";
|
||||||
import "rvbox/v1/common.proto";
|
import "rvbox/v1/common.proto";
|
||||||
|
|
||||||
|
option go_package = "github.com/rvbox/rvbox/gen/go/rvbox/v1;rvboxv1";
|
||||||
|
|
||||||
service Control {
|
service Control {
|
||||||
rpc ListClients(ListClientsRequest) returns (ListClientsResponse);
|
rpc ListClients(ListClientsRequest) returns (ListClientsResponse);
|
||||||
rpc GetClient(GetClientRequest) returns (GetClientResponse);
|
rpc GetClient(GetClientRequest) returns (GetClientResponse);
|
||||||
rpc ListCommands(ListCommandsRequest) returns (ListCommandsResponse);
|
rpc ListCommands(ListCommandsRequest) returns (ListCommandsResponse);
|
||||||
rpc GetCommand(GetCommandRequest) returns (GetCommandResponse);
|
rpc GetCommand(GetCommandRequest) returns (GetCommandResponse);
|
||||||
rpc RunCommand(RunCommandRequest) returns (RunCommandResponse);
|
rpc RunCommand(RunCommandRequest) returns (RunCommandResponse);
|
||||||
rpc RunCommandAndFollow(RunCommandRequest) returns (stream CommandEvent);
|
rpc FollowCommand(FollowCommandRequest) returns (stream FollowCommandResponse);
|
||||||
rpc FollowCommand(FollowCommandRequest) returns (stream CommandEvent);
|
|
||||||
rpc AppendStdin(AppendStdinRequest) returns (AppendStdinResponse);
|
rpc AppendStdin(AppendStdinRequest) returns (AppendStdinResponse);
|
||||||
rpc CloseStdin(CloseStdinRequest) returns (CloseStdinResponse);
|
rpc CloseStdin(CloseStdinRequest) returns (CloseStdinResponse);
|
||||||
rpc SignalCommand(ControlSignalCommandRequest) returns (ControlSignalCommandResponse);
|
rpc SignalCommand(ControlSignalCommandRequest) returns (ControlSignalCommandResponse);
|
||||||
rpc GetOutput(GetOutputRequest) returns (GetOutputResponse);
|
rpc GetOutput(GetOutputRequest) returns (GetOutputResponse);
|
||||||
|
rpc ListStorageIncidents(ListStorageIncidentsRequest) returns (ListStorageIncidentsResponse);
|
||||||
|
rpc RepairStorageIncident(RepairStorageIncidentRequest) returns (RepairStorageIncidentResponse);
|
||||||
|
rpc AcknowledgeStorageIncident(AcknowledgeStorageIncidentRequest) returns (AcknowledgeStorageIncidentResponse);
|
||||||
|
rpc AuthorizeClientTakeover(AuthorizeClientTakeoverRequest) returns (AuthorizeClientTakeoverResponse);
|
||||||
}
|
}
|
||||||
|
|
||||||
message ClientSummary {
|
message ClientSummary {
|
||||||
@@ -32,6 +37,10 @@ message ClientSummary {
|
|||||||
string daemon_version = 9;
|
string daemon_version = 9;
|
||||||
string daemon_cwd = 10;
|
string daemon_cwd = 10;
|
||||||
repeated ShellType supported_shells = 11;
|
repeated ShellType supported_shells = 11;
|
||||||
|
string client_instance_id = 12;
|
||||||
|
// Most recently rejected different instance while this client is live.
|
||||||
|
string pending_instance_id = 13;
|
||||||
|
google.protobuf.Timestamp pending_instance_seen_at = 14;
|
||||||
}
|
}
|
||||||
|
|
||||||
message ListClientsRequest {
|
message ListClientsRequest {
|
||||||
@@ -50,7 +59,6 @@ message GetClientRequest {
|
|||||||
|
|
||||||
message GetClientResponse {
|
message GetClientResponse {
|
||||||
ClientSummary client = 1;
|
ClientSummary client = 1;
|
||||||
ControlError error = 2;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
message ListCommandsRequest {
|
message ListCommandsRequest {
|
||||||
@@ -63,7 +71,6 @@ message ListCommandsRequest {
|
|||||||
message ListCommandsResponse {
|
message ListCommandsResponse {
|
||||||
repeated CommandRecord commands = 1;
|
repeated CommandRecord commands = 1;
|
||||||
string next_page_token = 2;
|
string next_page_token = 2;
|
||||||
ControlError error = 3;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
message GetCommandRequest {
|
message GetCommandRequest {
|
||||||
@@ -74,7 +81,6 @@ message GetCommandRequest {
|
|||||||
message GetCommandResponse {
|
message GetCommandResponse {
|
||||||
CommandRecord command = 1;
|
CommandRecord command = 1;
|
||||||
ResourceSnapshot latest_resource = 2;
|
ResourceSnapshot latest_resource = 2;
|
||||||
ControlError error = 3;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// For script execution, script_content contains the bytes whose descriptor is
|
// For script execution, script_content contains the bytes whose descriptor is
|
||||||
@@ -83,12 +89,15 @@ message RunCommandRequest {
|
|||||||
string target_client_id = 1;
|
string target_client_id = 1;
|
||||||
ExecutionSpec spec = 2;
|
ExecutionSpec spec = 2;
|
||||||
bytes script_content = 3;
|
bytes script_content = 3;
|
||||||
|
// Becomes issue_uuid when supplied; generated by the server when omitted.
|
||||||
|
string request_id = 4;
|
||||||
|
// Omitted uses the server default; zero explicitly requests no expiry.
|
||||||
|
google.protobuf.Duration queue_ttl = 5;
|
||||||
}
|
}
|
||||||
|
|
||||||
message RunCommandResponse {
|
message RunCommandResponse {
|
||||||
string issue_uuid = 1;
|
string issue_uuid = 1;
|
||||||
CommandLifecycle lifecycle = 2;
|
CommandLifecycle lifecycle = 2;
|
||||||
ControlError error = 3;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
message FollowCommandRequest {
|
message FollowCommandRequest {
|
||||||
@@ -98,37 +107,47 @@ message FollowCommandRequest {
|
|||||||
bool include_existing = 4;
|
bool include_existing = 4;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
message FollowCommandResponse {
|
||||||
|
oneof item {
|
||||||
|
CommandEvent event = 1;
|
||||||
|
// Server retention metadata; does not advance the client event cursor.
|
||||||
|
RetentionTruncation retention_truncation = 2;
|
||||||
|
}
|
||||||
|
// Set for a client event; server-created retention uses its own recorded_at.
|
||||||
|
google.protobuf.Timestamp server_receipt_time = 3;
|
||||||
|
}
|
||||||
|
|
||||||
message AppendStdinRequest {
|
message AppendStdinRequest {
|
||||||
string client_id = 1;
|
string client_id = 1;
|
||||||
string issue_uuid = 2;
|
string issue_uuid = 2;
|
||||||
bytes data = 3;
|
bytes data = 3;
|
||||||
bool append_newline = 4;
|
bool append_newline = 4;
|
||||||
|
string request_id = 5;
|
||||||
}
|
}
|
||||||
|
|
||||||
message AppendStdinResponse {
|
message AppendStdinResponse {
|
||||||
uint64 write_seq = 1;
|
uint64 write_seq = 1;
|
||||||
ControlError error = 2;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
message CloseStdinRequest {
|
message CloseStdinRequest {
|
||||||
string client_id = 1;
|
string client_id = 1;
|
||||||
string issue_uuid = 2;
|
string issue_uuid = 2;
|
||||||
|
string request_id = 3;
|
||||||
}
|
}
|
||||||
|
|
||||||
message CloseStdinResponse {
|
message CloseStdinResponse {
|
||||||
uint64 write_seq = 1;
|
uint64 write_seq = 1;
|
||||||
ControlError error = 2;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
message ControlSignalCommandRequest {
|
message ControlSignalCommandRequest {
|
||||||
string client_id = 1;
|
string client_id = 1;
|
||||||
string issue_uuid = 2;
|
string issue_uuid = 2;
|
||||||
SignalKind signal = 3;
|
SignalKind signal = 3;
|
||||||
|
string request_id = 4;
|
||||||
}
|
}
|
||||||
|
|
||||||
message ControlSignalCommandResponse {
|
message ControlSignalCommandResponse {
|
||||||
uint64 command_revision = 1;
|
uint64 command_revision = 1;
|
||||||
ControlError error = 2;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
message GetOutputRequest {
|
message GetOutputRequest {
|
||||||
@@ -136,13 +155,91 @@ message GetOutputRequest {
|
|||||||
string issue_uuid = 2;
|
string issue_uuid = 2;
|
||||||
repeated StreamKind streams = 3;
|
repeated StreamKind streams = 3;
|
||||||
uint64 after_event_seq = 4;
|
uint64 after_event_seq = 4;
|
||||||
uint32 page_size = 5;
|
uint64 max_bytes = 5;
|
||||||
|
string page_token = 6;
|
||||||
|
}
|
||||||
|
|
||||||
|
// An uncompressed slice of one persisted output event. Opaque pagination may
|
||||||
|
// resume within an event; event_byte_offset identifies that position.
|
||||||
|
message OutputSlice {
|
||||||
|
uint64 event_seq = 1;
|
||||||
|
google.protobuf.Timestamp observed_at = 2;
|
||||||
|
StreamKind stream = 3;
|
||||||
|
bytes data = 4;
|
||||||
|
uint64 event_byte_offset = 5;
|
||||||
|
bool end_of_event = 6;
|
||||||
|
google.protobuf.Timestamp server_receipt_time = 7;
|
||||||
|
}
|
||||||
|
|
||||||
|
message RetentionTruncation {
|
||||||
|
OutputTruncation truncation = 1;
|
||||||
|
google.protobuf.Timestamp server_recorded_at = 2;
|
||||||
}
|
}
|
||||||
|
|
||||||
message GetOutputResponse {
|
message GetOutputResponse {
|
||||||
repeated CommandEvent events = 1;
|
repeated OutputSlice output = 1;
|
||||||
string next_page_token = 2;
|
string next_page_token = 2;
|
||||||
bool output_truncated = 3;
|
bool output_truncated = 3;
|
||||||
ControlError error = 4;
|
repeated RetentionTruncation truncations = 4;
|
||||||
repeated OutputTruncation truncations = 5;
|
OutputIncomplete incomplete = 5;
|
||||||
|
}
|
||||||
|
|
||||||
|
enum StorageIncidentState {
|
||||||
|
STORAGE_INCIDENT_STATE_UNSPECIFIED = 0;
|
||||||
|
STORAGE_INCIDENT_STATE_OPEN = 1;
|
||||||
|
STORAGE_INCIDENT_STATE_REPAIRED = 2;
|
||||||
|
STORAGE_INCIDENT_STATE_ACKNOWLEDGED = 3;
|
||||||
|
}
|
||||||
|
|
||||||
|
message StorageIncident {
|
||||||
|
string incident_id = 1;
|
||||||
|
google.protobuf.Timestamp detected_at = 2;
|
||||||
|
google.protobuf.Timestamp resolved_at = 3;
|
||||||
|
StorageIncidentState state = 4;
|
||||||
|
string scope = 5;
|
||||||
|
string client_id = 6;
|
||||||
|
string issue_uuid = 7;
|
||||||
|
string summary = 8;
|
||||||
|
bool data_loss = 9;
|
||||||
|
bool automatically_repairable = 10;
|
||||||
|
}
|
||||||
|
|
||||||
|
message ListStorageIncidentsRequest {
|
||||||
|
bool include_resolved = 1;
|
||||||
|
uint32 page_size = 2;
|
||||||
|
string page_token = 3;
|
||||||
|
}
|
||||||
|
|
||||||
|
message ListStorageIncidentsResponse {
|
||||||
|
repeated StorageIncident incidents = 1;
|
||||||
|
string next_page_token = 2;
|
||||||
|
}
|
||||||
|
|
||||||
|
message RepairStorageIncidentRequest {
|
||||||
|
string incident_id = 1;
|
||||||
|
string request_id = 2;
|
||||||
|
}
|
||||||
|
|
||||||
|
message RepairStorageIncidentResponse {
|
||||||
|
StorageIncident incident = 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
message AcknowledgeStorageIncidentRequest {
|
||||||
|
string incident_id = 1;
|
||||||
|
string note = 2;
|
||||||
|
string request_id = 3;
|
||||||
|
}
|
||||||
|
|
||||||
|
message AcknowledgeStorageIncidentResponse {
|
||||||
|
StorageIncident incident = 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
message AuthorizeClientTakeoverRequest {
|
||||||
|
string client_id = 1;
|
||||||
|
string client_instance_id = 2;
|
||||||
|
string request_id = 3;
|
||||||
|
}
|
||||||
|
|
||||||
|
message AuthorizeClientTakeoverResponse {
|
||||||
|
google.protobuf.Timestamp expires_at = 1;
|
||||||
}
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user