192 lines
11 KiB
Markdown
192 lines
11 KiB
Markdown
# RVBox v1 agent protocol
|
|
|
|
The authoritative schemas are [`../protos/rvbox/v1/common.proto`](../protos/rvbox/v1/common.proto)
|
|
and [`../protos/rvbox/v1/agent.proto`](../protos/rvbox/v1/agent.proto). This
|
|
document specifies their use over WSS.
|
|
|
|
## Transport and compatibility
|
|
|
|
The nginx-terminated `wss://` connection carries exactly one serialized
|
|
`rvbox.v1.AgentEnvelope` in each binary WebSocket message. There is no extra
|
|
length prefix. A decoded envelope may not exceed 1 MiB. A chunk's uncompressed
|
|
payload may not exceed 64 KiB. Both peers validate declared and actual expanded
|
|
sizes before allocation/decompression.
|
|
|
|
The protobuf package is `rvbox.v1`. Registration negotiates a major/minor
|
|
protocol range: incompatible majors are rejected; the highest shared minor is
|
|
chosen. New fields are append-only. A peer ignores unknown optional fields. A
|
|
new required behavior requires a negotiated minor-version change; an envelope
|
|
with no recognized payload is a `PROTOCOL_ERROR` rather than an implicit
|
|
“required unknown field” mechanism that protobuf cannot represent.
|
|
|
|
## Heartbeat and reconnect
|
|
|
|
Both endpoints use the same algorithm. Any received valid WebSocket frame is
|
|
inbound activity. After 10 seconds without inbound activity, send a WebSocket
|
|
Ping. After 30 seconds without inbound activity, close the session and treat it
|
|
as dead. Pong processing is normal WebSocket behavior; it is not an application
|
|
message and never queues behind command traffic.
|
|
|
|
The client reconnects with full-jitter exponential backoff (initial 1 second,
|
|
cap 60 seconds). A session stable for 60 seconds resets the backoff. The server
|
|
does not reconnect; it waits for clients. A reconnect always registers again,
|
|
receives a new fencing generation, resends unacknowledged delivery/output, and
|
|
reconciles only commands the server still considers non-terminal.
|
|
|
|
## Session fencing
|
|
|
|
`ClientHello` starts registration and includes a random `client_instance_id`
|
|
generated once in the client state directory and retained across daemon restarts
|
|
and reconnects. `ServerWelcome` gives the selected version, random `session_id`,
|
|
and `session_generation`. Except `ClientHello`, all envelopes carry those
|
|
values. A new accepted registration from the same instance fences and
|
|
disconnects its prior session. A different instance claiming the same live
|
|
client ID is rejected and recorded as the current pending claim unless an
|
|
operator explicitly authorized that exact instance. The one-shot authorization
|
|
expires after 5 minutes by default and is consumed by the matching reconnect.
|
|
When no session for that client ID is live, a new instance is accepted normally.
|
|
This is collision protection, not peer authentication. The
|
|
server accepts messages only from the current generation; command dispatches
|
|
also identify their intended generation.
|
|
|
|
## Reliable work flows
|
|
|
|
### Dispatch and reconciliation
|
|
|
|
The server persistently creates a command before dispatching `CommandDispatch`.
|
|
It redelivers until it receives `CommandAccepted`. Clients durably deduplicate
|
|
on `issue_uuid`. A client whose queue is full sends a transient capacity
|
|
rejection, which the server requeues with backoff. A permanent validation or
|
|
unsupported-platform rejection makes the server command terminal `rejected`;
|
|
it is not retried or mislabeled as a launched-process failure.
|
|
`ExecutionSpec.elevated=false` requests normal privilege and true requests the
|
|
platform's elevated policy; non-Windows clients reject true in v1. Windows
|
|
resolves the effective token/session immediately before launch. With a usable
|
|
active user, false selects a verified non-elevated `ACTIVE_USER`; true tries
|
|
`ACTIVE_USER_ELEVATED`, `ACTIVE_SYSTEM`, then `LOCAL_SYSTEM`. With no usable
|
|
active user, false selects `LOCAL_SERVICE` and true selects `LOCAL_SYSTEM`.
|
|
Fallback is confined to token selection before `launch_prepared`; it is never a
|
|
process retry. Token/session attempts and effective identity are durable history.
|
|
An exact UUID/request-hash hit in the compact tombstone ledger returns
|
|
`CODE_ALREADY_EXECUTED`; the server suppresses dispatch and reconciles its stale
|
|
state instead of representing the prior execution as a new rejection.
|
|
|
|
Following registration the server sends `ReconcileRequest` containing its
|
|
non-terminal commands for that client. The client replies once with a complete,
|
|
bounded `ReconcileSnapshot` of every command it still retains, including
|
|
terminal-but-unacknowledged records and matching requested tombstones, then
|
|
continues normal retransmission from the server's last acknowledged event
|
|
sequence. The server does not dispatch new work until this snapshot is complete.
|
|
For healthy client storage, absence means the client never durably accepted the
|
|
UUID: server `queued`/`dispatched` work returns to `queued`, while absence of an
|
|
`accepted`/`running` record is an invariant failure and becomes interrupted.
|
|
A matching UUID/request-hash tombstone always suppresses replay. A client-known
|
|
server-missing active command, or a contradiction with server-confirmed
|
|
terminal history, is terminated locally after the server returns a durable
|
|
`ReconcileResult` and creates a recovery incident rather than inventing server
|
|
state. That result also identifies client-retained terminal records the server
|
|
has already stored or deliberately tombstoned, allowing the client to discard
|
|
them even when its last `EventAck` was lost. The result is idempotent and is
|
|
written before any new dispatch on that session. A client
|
|
with an unresolved essential-store incident must not claim a complete snapshot
|
|
or accept work until repaired/acknowledged.
|
|
|
|
### Output and events
|
|
|
|
Client execution entries first have a durable local order. They receive an
|
|
increasing wire `event_seq` durably when admitted to the bounded send window;
|
|
once assigned, the sequence/content is pinned until acknowledged and retries
|
|
reuse it exactly. The server durably writes an event before sending `EventAck`.
|
|
`EventAck` is cumulative through a sequence number. Output is Zstandard
|
|
compressed with an explicit original-size field. Server storage can reuse the
|
|
validated compressed bytes.
|
|
|
|
When rolling output removes old retained segments, the server writes
|
|
`OutputTruncation` metadata containing the removed event/byte ranges. Queries
|
|
must show that marker rather than silently presenting an apparently complete
|
|
stream. Assigned but unacknowledged events occupy the pinned 1 MiB per-command
|
|
send window and are not evicted. An offline client that reaches a cap replaces
|
|
one or more still-unsequenced output runs in its durable local order with
|
|
`OutputTruncation`; when admitted to the send window the marker receives the
|
|
next normal `event_seq`. Its event-range fields are absent because the discarded
|
|
bytes never had wire sequences. The normal cumulative `EventAck` acknowledges
|
|
the marker. Server-created retention markers include their removed event range
|
|
and remain query metadata because the server cannot allocate client sequences.
|
|
|
|
If a capture pipe cannot be drained to a provably complete EOF, the client emits
|
|
a sequenced `OutputIncomplete` event before the terminal lifecycle event. It
|
|
does not invent a missing sequence or byte count for bytes it never observed.
|
|
|
|
### Stdin and signals
|
|
|
|
`StdinWrite` is binary-safe and ordered by `write_seq`; the client durably
|
|
deduplicates it and returns a sequenced stdin-acknowledgement command event.
|
|
`append_newline` is true by default in the CLI but explicit on the wire.
|
|
`CloseStdin` is a separate idempotent action. Signal and cancellation requests
|
|
carry a command revision to settle start/kill races; the resulting lifecycle or
|
|
signal-result event echoes that revision. Control-plane `request_id` values stay
|
|
at the server and map to the assigned revision rather than crossing the agent
|
|
protocol.
|
|
|
|
### Scripts
|
|
|
|
For a script command, dispatch first contains a `ScriptDescriptor`; then the
|
|
server sends `ScriptChunk` messages and a commit. The client checks offset,
|
|
chunk order, full length, and SHA-256 before it reports upload complete or
|
|
launches the process. `ScriptUploadStatus.received_bytes` is a cumulative
|
|
durable progress acknowledgement: the server sends at most the command's
|
|
unacknowledged window, waits for progress, and resumes from that offset. A
|
|
retransmitted chunk is idempotent by offset/content.
|
|
|
|
Acceptance is durable queue admission, not proof that every later preparation
|
|
step will succeed. A permanent script checksum/write error or failure to prepare
|
|
the requested executable emits terminal `rejected` before any launch
|
|
authorization. User cancellation in that interval emits `cancelled`; an
|
|
uncertain crash window emits `interrupted`. None is mislabeled as process
|
|
`failed`.
|
|
|
|
On Windows, a context-selection rejection or the `running` lifecycle event
|
|
carries `WindowsExecutionIdentity`; later lifecycle events repeat it unchanged.
|
|
The server persists it into `CommandRecord`. If selection exhausted every
|
|
allowed context, `effective_context` and the effective/session identity fields
|
|
are absent. Session 0 effective contexts omit session fields.
|
|
Active contexts include both the target session/user SID and actual process-
|
|
token SID so `ACTIVE_SYSTEM` cannot be mistaken for execution as the desktop
|
|
owner. The ordered `attempted_contexts` and bounded `selection_detail` explain
|
|
fallback caused by a standard user, absent linked token, Administrator
|
|
Protection, or failed active-SYSTEM token construction. `selection_detail` and
|
|
every individual reason embedded in it are subject to the common 4 KiB detail
|
|
limit. Token selection and this event are downstream of durable acceptance but
|
|
upstream of requested code execution.
|
|
|
|
## Flow control and failure containment
|
|
|
|
No receive loop runs an executor, database write, decompressor, or slow socket
|
|
operation inline. Each side has bounded staging queues. Durable spools are the
|
|
source of truth and are charged to the tiered unified command-storage budgets
|
|
described in the architecture document. Senders keep a bounded unacknowledged
|
|
window of encoded wire bytes (defaults: 1 MiB per command and 8 MiB per client
|
|
session); unsent data
|
|
waits durably. WebSocket Ping/Pong/Close plus fencing and protocol errors use a
|
|
reserved priority lane and a dedicated writer with bounded data-frame size and
|
|
write deadlines. A full essential ingress queue does not stall the sole
|
|
WebSocket reader indefinitely:
|
|
the receiver closes without acknowledgement and the durable sender retries
|
|
after jittered reconnect. Droppable output uses the sequenced overload-loss
|
|
path instead.
|
|
|
|
Before compression, raw output uses the architecture's high/low watermarks.
|
|
Once a client high watermark is crossed, still-unsequenced droppable chunks are
|
|
summarized in durable local order instead of consuming unbounded compression
|
|
work. A server at its ingress high watermark does not acknowledge or discard an
|
|
already sequenced event; it closes without acknowledgement if its bounded queue
|
|
cannot admit the frame, and the client retries from durable spool. Fair
|
|
scheduling prevents one verbose command from monopolizing workers. Reserved
|
|
metadata capacity remains
|
|
available to close each gap and emit the final lifecycle event.
|
|
|
|
Malformed protobuf, over-size payload, invalid compressed data, impossible
|
|
sequence, bad session token, or protocol-version violation yields a structured
|
|
error where safe and closes that WebSocket session. It does not crash either
|
|
daemon. Network loss is normal and is handled through idempotent resend.
|