Files
rvbox/docs/implementation-plan.v1.md
T

1376 lines
78 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RVBox v1 Go implementation plan
## 1. Objective and implementation boundary
Implement the v1 contracts in the existing design documents as three Go
binaries:
- `rvbox-server`: durable server, WebSocket agent endpoint, local gRPC control
endpoint, optional JSON-RPC adapter, storage, and audit owner.
- `rvbox`: reverse-connecting client daemon, command supervisor, durable active
spool, and reconnect/reconciliation owner.
- `rvc`: local CLI over the server's Unix-domain gRPC socket.
The primary release target is a Linux server and Linux client. The Go client is
structured behind OS interfaces from the outset; Windows implementations of
process management, diagnostics, and resource controls are completed before a
Windows client is declared supported. Windows v1 requires Windows 10 or Windows
Server 2016 or newer. Do not claim Windows feature parity while those
implementations are absent.
This plan implements the already agreed v1 contract. In particular, it does
not add authentication, client enrollment, mutual TLS, or a client allowlist.
The deployment trust boundary remains nginx TLS termination and restricted
network access. The unauthenticated JSON-RPC listener remains disabled by
default and loopback-bound when enabled.
## 2. Development rules and definition of done
### 2.1 Local-environment rules
`../ENV.md` is binding for this repository:
- Do not install Go, Buf, `protoc`, SQLite tooling, or other heavy development
dependencies on the host.
- Put the toolchain, code generation, unit/integration testing, linting, and
local services in Docker/Docker Compose.
- Run container commands as UID/GID `1001:1001` so generated Go code, module
caches mounted into the project, and test artefacts remain host-owned.
- Do not delete Docker volumes, generated artifacts, or state directories as a
convenience cleanup action. Tests use dedicated temporary volumes/directories
and remove only the exact resources they created.
### 2.2 Completion standard for every phase
A phase is complete only when all of the following are true:
1. The public behavior and failure modes are covered by focused tests.
2. The package has context cancellation, bounded channels/queues, and no
network or disk operation on the WebSocket receive loop.
3. Structured logs and metrics exist for the new state transitions/failures.
4. Configuration defaults, validation, and errors are documented.
5. `docker compose run --rm toolchain make fmt lint test` passes, with no host
installation required. The lint target explicitly snapshots the accepted
first-draft Buf enum-prefix/service-suffix findings and fails on any other or
newly added finding; application/code lint has no such exception.
6. Any recovery or retention code is tested with a real temporary SQLite
database and filesystem rather than mocks alone.
## 3. Repository and build bootstrap (Phase 0)
### 3.1 Establish the repository layout
Create the following layout. Generated code is kept separate from handwritten
logic and is never manually edited.
```text
cmd/
rvbox-server/main.go
rvbox/main.go
rvc/main.go
internal/
agentproto/ # envelope validation, transport-neutral helpers
config/ # file/flag parsing and cross-field validation
domain/ # command state machine and typed errors
server/
control/ # gRPC implementation and JSON-RPC adapter
session/ # client registry, fencing, dispatch
store/ # SQLite metadata, segments, audit, migrations
client/
runtime/ # reconnect loop, transport, dispatcher
spool/ # active command/event/output durable spool
supervisor/ # OS-neutral interface
supervisor/unix/ # process groups, /proc, cgroup v2
supervisor/windows/# Job Objects and Windows diagnostics
observability/ # logging, metrics, health/readiness
testkit/ # clocks, fake transport, fault helpers
gen/go/rvbox/v1/ # generated protobuf/grpc code
protos/rvbox/v1/ # existing wire authority
docs/examples/ # fully annotated server/client TOML
deploy/
Dockerfile.toolchain
Dockerfile.runtime
compose.yaml
nginx/
systemd/
```
Use a single Go module rooted at the repository. Select and pin exact Go module
versions in `go.mod`; record reasons for non-standard dependencies in
`docs/dependency-decisions.md`. Prefer a pure-Go SQLite driver to avoid a C
toolchain for normal Linux/Windows client builds. Use a maintained,
context-aware WebSocket implementation and a Zstandard implementation that
supports bounded decompression. Do not rely on an archived WebSocket package.
### 3.2 Containerized developer tooling
Add:
- `deploy/Dockerfile.toolchain`: pinned Go base image plus Buf, `protoc`,
`protoc-gen-go`, `protoc-gen-go-grpc`, format/lint tools, and `make`.
- `deploy/Dockerfile.runtime`: minimal non-root image for server/client smoke
tests; it contains only built binaries and runtime CA/config assets.
- `deploy/compose.yaml`: a `toolchain` service with the repository mounted at
`/workspace`, run as `1001:1001`; optional `server`, `client`, `nginx`, and
fault-injection test services use isolated named volumes.
- `Makefile`: `generate`, `fmt`, `lint`, `test`, `test-race`, `test-integration`,
`build`, and `verify`. Each target invokes the container workflow; it must not
silently fall back to host tools.
- `buf.gen.yaml`: Go protobuf and gRPC generation paths under `gen/go`.
Generation is deterministic: `make generate` followed by `git diff --exit-code`
must be clean in CI. Update `buf.yaml`/`buf.gen.yaml` only with a matching
generation run.
### 3.3 Initial quality gates
Add CI (or a repository script ready for CI) that runs, in order:
1. `buf format --diff` and `buf lint`. Snapshot only the already accepted
enum-prefix/service-suffix findings so CI still fails if their set changes or
another warning appears. This pre-implementation cleanup defines the first
v1 compatibility baseline; enable `buf breaking` against that baseline
immediately after it merges, not against the obsolete draft.
2. Generation freshness check.
3. `go fmt`, `go vet`, static analysis, and unit tests.
4. Race tests for server/client concurrency packages.
5. Linux integration tests in Compose.
6. Cross-compilation/build verification for Windows client packages; Windows
runtime tests run on a Windows runner once available.
**Exit criteria:** all three empty `main` packages build in the toolchain
container; generated code is checked in; `make verify` works from a fresh clone.
## 4. Shared contracts, configuration, and state machine (Phase 1)
### 4.1 Preserve proto authority
Generate `common.proto`, `agent.proto`, and `control.proto` before writing
transport code. Do not fork request/response structs by hand. These schemas are
an unshipped first draft, so remove obsolete fields and compact their tags now
rather than preserving compatibility holes. Once Phase 0 generated APIs merge,
freeze field meanings: later additions use new tags, enum zero remains
unspecified, and removed tags are then reserved normally.
At the application boundary, validate what protobuf cannot express:
- `client_id` is configured hostname text, 1–128 ASCII characters.
- command and mutation IDs are canonical RFC 9562 UUIDv7 strings; all public
UUID parsing rejects empty, non-canonical, or wrong-version input.
- decoded envelopes are <= 1 MiB; output/stdin/script data validates every raw,
compressed, and declared size before allocation or decompression.
- a serialized `ExecutionSpec` is <= 768 KiB, decoded control gRPC requests are
<= 16 MiB, JSON-RPC HTTP bodies are <= 24 MiB, and decoded field limits remain
identical across the two control transports.
- shell type is explicit or assigned only to the documented platform default.
- CWD exists, is a directory, and is allowed by local daemon policy.
- an execution request has exactly one source: command text or script descriptor.
- script descriptors and payloads agree on SHA-256 and <= 10 MiB size.
- user-supplied script filenames are display-only basenames and never path
components; command text and scripts use generated private wrapper paths.
- profile flags are deduplicated and checked for configured supported limits;
`LIGHT` is exclusive, otherwise accept at most one CPU, one memory, and one
disk tier.
### 4.2 Domain package
Implement a pure-Go `internal/domain` package with no database/network imports.
It contains:
- the allowed command transition table and terminal-state predicate;
- `CommandRevision` compare-and-transition helpers for launch/cancel/signal
races;
- typed domain errors mapped to gRPC status/structured details, JSON-RPC error
objects, or agent-protocol `ControlError` as appropriate;
- event-sequence validation, duplicate equivalence checks, and declared gap
validation through `OutputTruncation` metadata;
- default values and hard-limit validation;
- stable ordering/pagination token encoders.
Model `launch_prepared` and `launch_authorized` as durable internal execution
phases without exposing them as additional public lifecycle enum values. No
recovery path may redispatch a UUID after `launch_authorized`; an uncertain
outcome becomes `interrupted`.
The normal legal state transitions are:
```text
queued -> dispatched -> accepted -> running -> succeeded|failed|terminated|interrupted
queued -> expired
dispatched -> rejected (permanent pre-acceptance failure)
dispatched -> queued (transient rejection with backoff)
accepted -> rejected (permanent upload/preparation failure before launch)
queued|dispatched|accepted -> cancelled (subject to revisioned race resolution)
dispatched|accepted -> interrupted (reconciliation/recovery invariant failure)
```
An attempted repeat event is harmless only if it has identical immutable
content. A conflicting duplicate UUID/sequence is a protocol error and fences
the bad session rather than rewriting history.
### 4.3 Configuration model
Use strict TOML 1.0 and implement the normative contract in
[`configuration.md`](configuration.md). Keep the annotated
[`examples/server.toml`](examples/server.toml) and
[`examples/client.toml`](examples/client.toml) synchronized with the Go config
structs and compiled defaults.
Implement configuration in this order:
1. Define separate decode-only TOML structs and immutable validated domain
structs. Every public TOML key has an explicit tag and documentation comment;
packages outside `internal/config` receive only a validated subsection.
2. Decode exactly one UTF-8 file with unknown-field and duplicate-key rejection.
Do not merge include files, expand environment variables, coerce strings to
numbers, or accept bare numeric durations.
3. Apply precedence deterministically: compiled defaults, file values, then
flags that were explicitly present. Generate flag bindings from one key
registry so omitted flags cannot accidentally overwrite TOML values.
4. Parse duration strings with overflow and non-negative checks. Parse byte/count
integers without floating point. Enforce v1 hard ceilings even when TOML asks
for more; operators may lower them but not create an incompatible wire peer.
5. Canonicalize data/state/socket/CA/shell/CWD-root paths to absolute paths.
Reject aliases that collapse distinct server data subdirectories and a
control socket outside its intended runtime-directory policy. Check existing
path ownership/type without claiming that private modes encrypt stored data.
6. Resolve and validate current-platform shells at client startup. Each
nonempty applicable path must be absolute and identify an executable regular
file; other-platform strings are syntax-checked but never advertised.
Advertise only validated entries and require the current platform default to
be one of them. Cache the canonical path and never re-resolve through a
command's overridden `PATH`. Revalidate identity before launch and reject
rather than fall back if the executable changed incompatibly.
7. Validate resource profiles by dimension. `LIGHT` is exclusive; otherwise
allow at most one `CPU_*`, one `MEM_*`, and one `DISK_*`. Each enabled profile
must give nonzero values for its `required_controls`; validate Linux device
keys as canonical `major:minor` values and Windows Job rate support before
advertising the profile.
8. Enforce cross-field invariants: low watermark below high, send window below
durable tier, closeout reserve below command quota, command quota below both
aggregate tiers, reconnect initial <= cap, heartbeat idle < liveness timeout,
positive page/queue limits, per-client server queue <= global server queue,
and zero only where the schema documents disable or indefinite behavior.
9. Print the normalized effective non-command configuration once after
validation. Log a conspicuous warning for enabled non-loopback JSON-RPC, but
retain the accepted behavior. Do not log TLS file contents or profile device
discovery details that reveal unrelated host paths.
10. Bind no command listener and launch no child until static configuration is
valid. Storage recovery remains asynchronous after this static gate as
specified in Phase 2.
Server TOML covers routing/listeners, queue/retry bounds, unified storage and
audit limits, group-commit behavior, flow windows/watermarks, protocol ceilings,
takeover TTL, and observability. Client TOML covers WSS/TLS routing, durable
identity/state, exact shell paths, allowed CWD roots, queue/concurrency,
reconnect/liveness, spool/flow limits, execution diagnostics/grace periods,
resource profiles, and observability.
Do not add or claim application-level at-rest encryption in v1. Command-owned
payloads are stored in plaintext under the private state directories; document
the equivalent sensitivity of backups and defer encryption/key management.
Hard defaults are: 10-second heartbeat-idle interval, 30-second liveness
timeout, 1–60-second full-jitter reconnect, reset after 60 stable seconds, 16
running/100 queued client commands, 1,000 server-queued commands per client and
10,000 server-wide, 5-minute one-shot live-conflict takeover grant, 15-minute
queue TTL, 64 KiB raw stream
chunk, 1 MiB decoded agent envelope, 768 KiB serialized execution spec, 16 MiB
decoded control request, 24 MiB JSON-RPC HTTP body, 10 MiB output window/raw
script cap, 32 MiB total per command, 256 MiB per client/client daemon, 4 GiB
server-wide command storage, 30-day terminal retention, 100 MiB audit storage
with quota-only rotation by default (zero age retention),
one million compact command tombstones, raw-output high/low watermarks of
1 MiB/256 KiB per command, 8 MiB/4 MiB per client, and 64 MiB/32 MiB server-wide,
plus unacknowledged send windows of 1 MiB per command and 8 MiB per session.
Reserve 64 KiB within every accepted command's quota for terminal/loss closeout
metadata and bound protocol detail/reason plus incident-note text to 4 KiB.
Use emergency filesystem free-space floors of 256 MiB server-side and 64 MiB
client-side by default.
Test both annotated example files through the production decoder. Add table tests
for every unknown key, invalid duration/size, high/low inversion, tier inversion,
zero semantic, noncanonical path/device, unsupported default shell, conflicting
profile combination, and flag-precedence case. Add a test proving a request-level
`PATH` override cannot change the chosen shell executable. Golden-test the
redacted effective configuration and keep its key order stable enough for
operators to compare deployments.
**Exit criteria:** state-machine and configuration tests cover all legal/illegal
transitions and boundary values; both examples parse to the documented defaults;
generated protos compile and are the only DTOs crossing process boundaries.
## 5. Server persistence and retention (Phase 2)
### 5.1 Storage ownership and directory safety
Implement `internal/server/store` as the sole writer for server state. On first
boot, create a private data directory (mode `0700`), SQLite database, segment
directory, and audit directory. Validate that paths resolve below the configured
data directory; never construct a path directly from unvalidated client IDs or
filenames. Use encoded UUID/client-ID path components.
Open SQLite with WAL enabled, foreign keys enabled, a busy timeout, and an
explicit durable-write policy. Run migrations transactionally. On startup:
1. acquire a single-instance lock;
2. bind liveness and diagnostic/control endpoints with readiness false;
3. run integrity/schema checks and committed-range recovery asynchronously;
4. safely truncate file bytes beyond SQLite's `committed_end_offset`, while a
missing/corrupt committed range creates a scoped durable incident;
5. mark prior live sessions disconnected and retain trustworthy non-terminal
commands for reconciliation; interrupt only affected commands whose
essential state cannot be trusted;
6. enable each healthy scope independently. Mutations against an unrecovered or
dirty scope return `UNAVAILABLE`; liveness, incident inspection, and
unaffected scopes remain available throughout best-effort recovery.
Represent recovery with explicit atomic scope states: `recovering`, `ready`,
`dirty_readable`, and `unavailable`. Liveness depends only on the process/event
loop; global readiness is true only when every required scope is ready, while an
RPC checks the narrower scopes it will touch. The WebSocket endpoint may accept
and close with a retryable storage error during recovery, but it must not issue a
durable `ServerWelcome`, accept events, or dispatch commands until the relevant
client/session scope is ready.
Run `PRAGMA quick_check`, migration checksum validation, and segment reconciliation
on a bounded recovery pool. Persist a startup incident in the database when it
is writable; otherwise expose an in-memory bootstrap incident through health and
require the offline repair tool. Never loop forever on one corrupt client or
segment. Record progress counters and the last failing path/offset without
putting payload bytes in logs.
### 5.2 Metadata schema and indexes
Use migrations to create at least the following tables; names may vary but their
invariants must not.
| Data | Required fields/invariants |
| --- | --- |
| `clients` | client ID, most-recent platform/capabilities/CWD/version, durable instance ID, current generation, connection/last-seen timestamps, latest rejected live-conflict instance/time, unified charged command-storage total |
| `sessions` | opaque session ID, client ID, durable client-instance UUID, generation, opened/fenced/closed times, close reason |
| `commands` | UUID, client ID, indexed issue/queue-expiry/terminal times, lifecycle, revision, exit result, last event sequence, retention status, and a Zstandard-compressed immutable execution-spec payload including plaintext environment values |
| `command_payloads` | command UUID, payload kind (script or other command-owned blob), Zstandard compression, raw/stored sizes, digest, inline bytes or validated segment reference |
| `command_events` | command UUID + event sequence unique key, observed and server receipt times, event type, payload metadata, immutable duplicate checksum |
| `output_segments` | command UUID, segment ordinal/path, `committed_end_offset`, min/max event sequence, stream mix, compressed/raw byte totals, checksum, created time |
| `output_truncations` | command UUID, removed event range/byte totals, source, reason, recorded time; non-overlapping ranges |
| `stdin_writes` | command UUID + write sequence unique key, compressed payload/raw/stored sizes/checksum, newline/close intent, acknowledged state |
| `control_mutations` | request UUIDv7 unique key, owner command/incident/client, method/target, immutable request hash, assigned write sequence/revision and stable result; conflicting reuse is rejected |
| `takeover_authorizations` | client ID + exact pending instance ID, creation/expiry/consumption times and authorizing request ID; one live, one-shot grant per client |
| `command_tombstones` | binary UUIDv7 unique key, immutable request hash, client ID, completion/acknowledgement time; compact global FIFO capped at 1,000,000 rows |
| `audit_events` | timestamp, source transport/principal where available, client/command IDs, action, outcome, error code, command text/script digest/env key names |
| `storage_incidents` | incident UUID, detected/resolved times, state, affected scope/client/command, summary, repairability and data-loss flags; unresolved rows derive dirty health |
| `schema_migrations` | applied version/checksum/time |
Index client command lists by `(client_id, issue_time DESC)`, active command
dispatch by `(client_id, lifecycle, issue_time)`, and output reads by
`(command_uuid, event_seq)`. Use cursor tokens based on the sorted key plus a
signature/version, not SQL offsets, so page results remain stable under writes.
Add foreign keys with explicit delete behavior and database constraints for
terminal timestamps, nonnegative byte counters, unique `(command,event_seq)` and
`(command,write_seq)`, one active session per client, and one unconsumed takeover
grant per client. Store UUIDs and SHA-256 values as fixed-size blobs internally;
canonical strings exist only at protocol/log boundaries. Canonical request hashes
use deterministic protobuf encoding of immutable fields after defaults and map
ordering are normalized; mutable lifecycle/receipt fields are excluded.
Every mutation transaction updates its owner record, all three applicable quota
counters, the idempotency row, and audit intent/result consistently. Identical
`request_id` plus hash returns the stored result without re-running side effects;
same ID with a different method, owner, or hash returns `CONFLICT`. Add migration
invariant queries that recompute charged totals and fail readiness if persisted
counters disagree until repaired.
### 5.3 Segment format and atomic append
Store large command, output, and audit payloads outside SQLite. An active segment
is append-only; seal it at the configured 256 KiB target and never append again.
Use an encoded UUID/ordinal filename created with exclusive create. Each record
contains fixed magic/version/header length/record length, command and event
identity, observed/receipt timestamps, payload kind/stream/compression, raw and
stored lengths, SHA-256 payload digest, header CRC, payload, and trailing record
CRC. All lengths are unsigned and validated against configured maxima before
allocation or seeking.
The append path is:
1. validate compressed bytes and bounded decompression;
2. choose/append the active command segment;
3. append and `fsync` (possibly as a configured group commit);
4. in one SQLite transaction, insert event/segment metadata, advance that
segment's `committed_end_offset`, update command sequence/counts, and record
receipt time;
5. only after durable success, return the cumulative `EventAck`.
The writer serializes appends per active segment but group-commits independent
commands together. Creating/renaming a segment also syncs its containing
directory before metadata can commit. A group-commit timer is a maximum batching
delay, never permission to acknowledge before file sync. SQLite uses a
documented synchronous mode sufficient for the selected durability promise;
tests assert the actual PRAGMA values after opening every connection.
On recovery, file bytes beyond the committed offset are unacknowledged tail and
are truncated. A short file or checksum failure inside the committed range is
never “repaired” by silently moving the database offset backward: mark the
affected output unavailable/truncated, preserve evidence, and create an
incident. Essential-state corruption interrupts only the affected command and
gates unsafe mutations.
Apply this committed-offset protocol to command blob and audit segments too.
For a missing committed output range, preserve all still-valid records, create
query-visible truncation/incomplete metadata, and never fabricate byte totals.
For a missing immutable execution spec, request hash, revision, or active stdin
state, mark the command essential state untrustworthy and interrupt it. Recovery
must be idempotent after a crash at every repair step.
If database or segment persistence fails, do not acknowledge the event. Surface
the server health failure, stop dispatching work as appropriate, and keep the
session/control loops responsive.
### 5.4 Retention transaction
Implement retention in a single serialized maintenance worker, never on the
WebSocket receive loop.
Define a versioned deterministic charge formula: compressed/encoded payload and
segment lengths plus a conservative fixed charge for each SQLite row/index
entry. Do not attempt to attribute shared SQLite pages after the fact. Reconcile
the logical counters transactionally and also reject allocations below the
configured 256 MiB server filesystem free-space floor, regardless of logical
quota headroom.
Implement one server `Reserve(owner, kind, chargedBytes)` path used by command
creation, script append, stdin, event append, and audit. Reuse the same contract
in the client spool, where it also covers raw execution-file preparation. Under
the store writer lock/transaction it checks, in order: hard field maximum,
per-command total and closeout reserve, per-client total, server/client-daemon
total, and filesystem floor. It either reserves all applicable tiers or none.
Release uses the same stored charge version; a future estimator change requires
a migration, not a silent reinterpretation of old rows.
1. Account all stored command-owned data, including request/script payload,
pending stdin, metadata/events, and output. The client additionally charges
generated raw execution files while they exist.
Reserve non-evictable active state
before accepting it, including 64 KiB per-command closeout headroom and the
bounded result for each accepted stdin/signal mutation; reject an allocation
that cannot fit.
2. Enforce the command's 10 MiB output window and 32 MiB total cap by rotating
oldest output segments and recording `OutputTruncation`; never roll essential
active state.
3. Enforce the 256 MiB per-client and 4 GiB server-wide caps by evicting terminal
commands as whole UUID records in oldest server issue-time/UUIDv7 order. If active data alone creates
pressure, rotate/drop output with markers and reject new essential state.
4. Reclaim every whole terminal command 30 days after authoritative terminal
time by default, independently of byte pressure; zero explicitly disables
age rotation. Preserve compact tombstones and separate audit/incidents.
5. Maintain audit events in compressed segments under their independent
100 MiB default, rotating complete oldest segments. Apply an independently
configurable audit age as a second trigger; zero, the default, disables age
rotation but not byte-budget rotation.
6. Keep unresolved storage incidents non-evictable. Resolved incident summaries
may rotate with their audit history under the separate 100 MiB budget, but
resolving dirty health never immediately rewrites or disguises recorded loss.
Retention selects candidates deterministically in `(server_issue_time,
issue_uuid)` order and rechecks terminal state plus generation inside the delete
transaction. Before deleting a command, close follower snapshots over its
retained ranges, ensure the compact tombstone exists, mark `evicting`, and move
all owned files to a same-filesystem deletion directory using collision-proof
names. A crash-safe sweeper rolls forward rows/files in either order. Never use
a glob or client-provided path for deletion.
For active pressure, first evict server-retained output and record exact
`RetentionTruncation`; then reject new unreserved mutations. Do not evict the
assigned client send window or essential lifecycle/revision/dedupe state. Age
rotation and byte-pressure eviction call the same whole-command primitive so
their crash behavior cannot diverge.
Deletion is recoverable: mark rows `evicting` transactionally, move segment
files to a same-filesystem tombstone location, commit metadata deletion, then
remove tombstones asynchronously. Startup completes or rolls forward stale
evictions deterministically.
Replay tombstones must never have a delete/insert gap. The server inserts the
command tombstone transactionally with authoritative terminal state before it
can acknowledge or later evict command history. On the client, receipt of an
acknowledgement through the terminal event atomically inserts or retains the
UUID/request-hash tombstone and removes the full command record. Enforce each
one-million-entry FIFO cap in that same transaction.
### 5.5 Incident resolution
Run safe, deterministic repairs automatically and through the control API.
`RepairStorageIncident` must refuse any action that would knowingly discard
committed data. `AcknowledgeStorageIncident` requires an operator note and is
the explicit path for accepting irrecoverable loss. Both are idempotent by
`request_id`. A repaired or acknowledged incident no longer contributes to the
derived dirty flag; resolution itself does not delete history. Bound notes to
4 KiB and retain resolved summaries/audit entries under their separate 100 MiB
rotation.
Provide the same store library behind offline `rvbox-server repair --data-dir`
list/repair/acknowledge modes. Provide equivalent local client-spool handling
through `rvbox repair --state-dir` and its local health diagnostics. Never open
a live store from the offline tool; acquisition of the same single-instance lock
is required.
Define stable incident kinds for uncommitted tail, missing committed bytes,
checksum mismatch, counter mismatch, SQLite integrity failure, failed eviction,
permission error, and disk exhaustion. An incident has immutable evidence plus
mutable state `OPEN -> REPAIRED|ACKNOWLEDGED`; terminal states never reopen, so a
recurrence creates a new incident ID. Scope keys identify global store, client,
command, segment, or audit without embedding raw filesystem paths in the public
API.
Safe repair may truncate only bytes beyond `committed_end_offset`, rebuild a
derivable counter/index, or roll forward a marked eviction. Anything that would
discard acknowledged bytes or immutable essential state requires acknowledgement
instead. Online repair acquires the same per-scope maintenance lock as recovery
and retention; offline repair additionally requires the process-wide instance
lock. Persist the repair/acknowledgement mutation and audit record before clearing
derived dirty health.
### 5.6 Store tests
Test real SQLite/filesystem cases: duplicate event idempotence; conflicting
duplicate rejection; crash between segment write and metadata transaction;
crash after metadata commit; corrupted final record; command-window rotation;
per-client/global whole-command eviction; active-command protection; cursor
stability; uncommitted-tail truncation; short/corrupt committed ranges; dirty
derivation and repair/acknowledgement; and disk-full/permission failure
simulation. Verify recovery endpoints become live promptly while unsafe scoped
mutations remain gated.
**Exit criteria:** a store restart preserves acknowledged data and never
acknowledges uncommitted data; all rolling, tiered unified-quota, 30-day
terminal-age, audit, and tombstone rules match the documented semantics.
## 6. Server sessions, WSS endpoint, and dispatch (Phase 3)
### 6.1 Agent endpoint
Expose an HTTP endpoint suitable for nginx WebSocket proxying. TLS may terminate
at nginx, but the server must verify WebSocket upgrade origin/path/configuration
and impose read limits itself. It accepts one binary protobuf envelope per
message and refuses text messages, fragmentation misuse, oversized frames,
invalid compressed payloads, and malformed/unknown mandatory messages.
Use four independently bounded activities per session:
- a read loop that validates/fragments/decode-dispatches only;
- a prioritized write loop with reserved Ping/Pong/Close/fencing/error capacity,
bounded data frames, and write deadlines;
- a dispatch worker that reads persisted command work;
- a persistence/event worker that writes through `store`.
No goroutine may hold a session-registry lock while doing database, filesystem,
compression, or WebSocket I/O. Every goroutine receives a session context and
must terminate on fencing/close.
Keep at most 1 MiB unacknowledged per command and 8 MiB per client session;
additional data waits in durable storage. If an essential ingress queue fills,
close without acknowledging rather than blocking the sole reader; the durable
sender retries after jittered reconnect. Ping/Pong remains sendable even while
the data window is full.
Implement the write scheduler as two bounded lanes owned by one socket writer:
the reserved control lane contains Ping/Pong/Close, welcome/fencing, protocol
errors, acknowledgements, and reconciliation control; the data lane contains
dispatch, script, stdin/signal, and events. Always drain a bounded burst of
control before one data frame so data still progresses. Impose write deadlines
and split data below the frame ceiling. A full control lane is a session-fatal
internal invariant; a full data lane leaves work durable and wakes the writer
later rather than allocating another goroutine.
The reader checks WebSocket type/size before protobuf decode, validates envelope
session fields before payload allocation, and routes only lightweight immutable
work descriptors. Persistence workers preserve per-command ordering while a fair
round-robin/deficit scheduler prevents a verbose command from monopolizing
decompression or disk. Unit-test queue capacities and goroutine shutdown with a
fake connection that never completes writes.
Apply the server 64 MiB/32 MiB ingress backlog as admission hysteresis, not as a
license to discard client-sequenced events. Above high, stop acknowledging new
output and close the producing session if the bounded descriptor queue cannot
accept its next frame; the client's durable spool retries it. Resume normal
admission only below low. Keep lifecycle/control reserve independent so another
client's heartbeats and command closeout remain responsive.
### 6.2 Registration and fencing
On `ClientHello`, validate client ID/capabilities/version and the durable random
client-instance UUID; choose the highest shared minor under major v1; persist
the registration; allocate a cryptographic random opaque session ID and
increment the client's generation transactionally. Fence the prior session for
the same client ID and instance before making the new one active. Accept a
different instance normally if no session for that client ID is live. Otherwise
record it as the pending claim and reject it unless an unexpired one-shot grant
matches that exact instance; consume the grant transactionally when accepting
the reconnect and fencing the old session. The default grant TTL is 5 minutes.
Send `ServerWelcome` with the outer envelope's canonical session fields.
Persist a rejected live-conflict claim before closing the candidate socket and
expose only its exact `client_instance_id` and observation time through
`GetClient`. `AuthorizeClientTakeover` verifies that this ID still matches the
pending claim, stores an idempotent UUIDv7 mutation and five-minute expiry, and
does not fence immediately. On matching reconnect, one transaction increments
generation, consumes the grant, closes the old session record, and installs the
new session; only then may the writer send welcome. A mismatched/expired grant is
never broadened to “the next instance.” If no session is live, accept a different
instance without creating or consuming a grant, as agreed.
All post-hello envelopes must match both session ID and generation. Stale or
unknown session traffic is discarded/audited and its WebSocket is closed. A
dispatch includes the intended generation, and the client must reject a mismatch.
Unknown client IDs are accepted by design; hostname is display/routing identity,
not an authentication assertion.
### 6.3 Heartbeat and liveness
Implement the identical policy at both ends: any received valid WebSocket frame
updates inbound activity; after 10 seconds silent send Ping; after 30 seconds
without inbound frame close the session. Pings/Pongs use WebSocket controls, not
protobuf envelopes. Make timers injectable/fakeable for deterministic tests.
Use a monotonic clock for elapsed activity/backoff and wall-clock UTC only for
persisted observation. Pong/any valid inbound frame updates activity in the read
loop without waiting for a database worker. Reset reconnect backoff only after a
continuous 60-second stable session. Add boundary tests at exactly 10 and 30
seconds, delayed writer tests proving Ping bypasses a full data lane, and clock
jump tests proving wall-clock adjustment cannot spuriously kill a session.
### 6.4 Queueing, dispatch, and reconciliation
`RunCommand` persists a `queued` command before returning. The dispatcher sends
to the active session while either an immediate running slot or a client queue
slot is available: `running < max_running || queued < max_queued`. Dispatch and
capacity updates are serialized per session so concurrent sends cannot
oversubscribe the advertised slots. It updates to `dispatched` before write and
retries after reconnect until a matching `CommandAccepted` arrives.
Maintain per-session shadow reservations for in-flight dispatches: reserve a
running slot when immediately available, otherwise a queued slot, before
enqueuing the frame. Reconcile the shadow with each `ClientCapacity` and
`CommandAccepted`; release it on definite rejection/session close. The
dispatcher uses one serialized worker per client plus a fair global ready queue,
so the `OR` capacity condition cannot oversubscribe and an offline/noisy client
cannot starve others.
Apply the configurable queue TTL (15 minutes by default; zero means indefinite)
until server-confirmed acceptance. A never-dispatched command becomes terminal
`expired`. A dispatched command with uncertain acceptance remains protected and
is rendered expired pending reconciliation rather than being prematurely made
terminal. Late client evidence advances the actual lifecycle, creates a
late-after-expiry incident, and triggers best-effort termination while retaining
subsequent actual outcome events.
Persist an absolute server-clock `queue_expiry_time` when creating the command;
never recompute it after configuration changes or retries. Drive expiry from a
database-backed ordered index and recheck lifecycle/revision transactionally.
Queued commands become `EXPIRED`; dispatched commands set a derived
late-pending-reconciliation indicator without writing a false terminal event.
On late acceptance, persist the incident and higher termination revision before
sending best-effort TERM, then retain every actual subsequent event.
Classify a negative `CommandAccepted` before changing lifecycle. Explicitly
transient capacity/storage pressure returns `dispatched -> queued` with bounded
jittered backoff and remains under the original TTL. Invalid execution data or
unsupported required platform controls become terminal `rejected` with the
structured reason persisted on `CommandRecord`. Reserve `failed` for a process
that reached `running`.
On replacement/reconnect, query non-terminal commands for that client and send
a `ReconcileRequest` with last durable event sequence/revision plus immutable
request hash. Require one complete client snapshot of all locally retained
commands, including
terminal-but-unacknowledged records and matching requested tombstones, and
compare the sets idempotently before enabling dispatch. With healthy client
storage, an absent `queued`/`dispatched` command returns to `queued`; an absent
`accepted`/`running` command is interrupted as an invariant failure. A matching
UUID/request-hash tombstone suppresses replay. The client terminates a local
active command absent from the server target set only after receiving the
server's durable `ReconcileResult`; either peer records a contradiction with
server-confirmed terminal state as a storage/recovery incident, and the server
does not invent missing state. The result also lists terminal local records safe
to discard when the server has stored or deliberately tombstoned them, covering
a lost final `EventAck` followed by server retention. Write this idempotent
result before dispatch. A client with unresolved essential-store corruption
cannot complete reconciliation or accept work.
Implement reconciliation as this explicit matrix:
| Server state | Client evidence | Durable result |
| --- | --- | --- |
| queued/dispatched | absent from healthy complete snapshot | return/remain queued; redelivery permitted |
| dispatched/accepted/running | matching retained record | keep highest valid revision and resume event/script delivery |
| any non-terminal | matching tombstone | interrupt/suppress replay and record stale-server incident |
| accepted/running | absent | interrupt and record client-state-loss incident |
| missing or contradictory terminal | client active | `ReconcileResult.terminate_local_issue_uuids` plus incident; invent no server command |
| missing/tombstoned or fully stored terminal | client terminal | `discard_local_terminal_issue_uuids` |
Validate matching UUID, immutable request hash, revision monotonicity, terminal
immutability, and event-sequence bounds for every row before applying any row.
Persist all server state/incident decisions in one reconciliation transaction,
then send one deterministic sorted `ReconcileResult` through the control lane.
The client durably records terminate/discard intent before reporting capacity.
If the result frame is lost, the next complete snapshot yields the same result;
dispatch is disabled until the current result has been written on the session.
Handle pre-start `kill` locally only for work that has never been dispatched.
For dispatched or accepted work, persist a higher revision and send the
revisioned signal while retaining the existing visible lifecycle until the
client resolves the race. Accept a revisioned `cancelled` lifecycle if launch
was not authorized; otherwise retain the revisioned signal result and actual
terminal outcome. The server maps control `request_id` to this revision/result
and never sends that request ID across the agent protocol.
### 6.5 Server session tests
Cover version negotiation, same-instance replacement, different-instance
pending claim, exact one-shot takeover consumption/expiry, stale-generation
late output, Ping/Pong inactivity thresholds, full client queue, dispatch retry
after lost acceptance, complete active/terminal-unacknowledged reconciliation,
lost/repeated `ReconcileResult`, permanent/transient dispatch rejection,
cancellation before launch, launch/cancel race, malformed envelope close, and a
slow client that cannot block another client or local control RPC.
**Exit criteria:** multiple simulated clients can register, replace each other,
receive bounded dispatches, and recover connection faults without duplicate
execution or server deadlock.
## 7. Client runtime, spool, and Unix supervisor (Phase 4)
### 7.1 Client state and reconnect runtime
Keep client state private (mode `0700`): durable accepted-command records,
active process metadata, command event journal, stdout/stderr spool segments,
stdin write acknowledgements, and outstanding script upload state. It contains
no terminal command history after server acknowledgement and local cleanup.
It retains a separate compact FIFO of the most recent 1,000,000 command
tombstones (binary UUIDv7, immutable request hash, and acknowledgement time) so
stale server state cannot replay recently completed work after full history is
removed.
The runtime has a long-lived supervisor/dispatcher and replaceable network
session. Generate the random client-instance UUID once in the private state
directory and retain it across daemon restarts and reconnects. Network loss must
not stop running commands or pipe readers. The
connection loop uses full-jitter exponential backoff from 1 second to 60
seconds, resetting only after 60 stable seconds. On each connection it sends a
hello with the durable client-instance UUID, waits for welcome, replays retained
sequenced loss events and other unacknowledged events in order, answers
reconciliation, and resumes dispatch.
All client queues are bounded. The disk spool, not an in-memory channel, is the
source of truth. Give stdout, stderr, status, signal results, and stdin
acknowledgements one durable command-local order. Assign and persist wire
`event_seq` only when an entry enters the bounded send window; an assigned entry
is pinned until acknowledgement and immutable on retry.
Create `client_instance_id` with a cryptographic UUID generator, write and fsync
a temporary file, atomically rename it, and sync the state directory before
first registration. Never regenerate it merely because the network/server
rejects a session. If the identity file is corrupt, enter dirty local health and
require the repair path rather than silently presenting a new installation.
Give the client store the same committed-offset segment discipline as the
server. Persist, at minimum, command phase/revision/hash, platform launch
identity, local event ordinal, optional assigned wire sequence, payload digest,
stdin write state, script received ranges, and last server acknowledgement.
Maintain separate monotonic `local_ordinal` and `event_seq` columns: insertion
assigns only local order; admission to the send window transactionally assigns
the next contiguous wire sequence and pins the row. Cumulative `EventAck`
atomically unpins/deletes acknowledged payload and updates the command cursor.
Structure reconnect as explicit states: `backoff -> connecting -> hello ->
reconciling -> active -> closing`. Only pipe/supervisor workers outlive the
replaceable network session. After welcome, send the complete snapshot, apply
`ReconcileResult` durably, replay assigned rows first, then assign/send local
rows in order. Advertise capacity and accept new dispatch only in `active`.
Cancellation of one session context must join all its readers/writers before a
new session can use their queues.
### 7.2 Output capture and offline caps
Create non-blocking readers for stdout and stderr immediately after process
start. Read raw bytes irrespective of newline boundaries, divide them into <=
64 KiB raw chunks, add observed UTC time/local order, Zstandard-compress, append
to the command spool, and enqueue only a durable reference to the sender.
Use one reader per pipe and a bounded shared raw-chunk scheduler. A reader never
waits for network I/O. Under normal load it transfers ownership of a bounded
buffer to a fair compressor, which returns buffers to a pool only after durable
append. Validate Zstandard frame size/checksum on readback before replay. Do not
log raw chunks on compression, checksum, or storage failure.
Bound work before compression as well as on disk. Defaults are high/low raw
backlogs of 1 MiB/256 KiB per command and 8 MiB/4 MiB for the client daemon; the
server compression/persistence tier independently uses 64 MiB/32 MiB globally
with the non-dropping admission behavior in Phase 3. Crossing a client high
watermark enters loss mode until the matching low watermark. Aggregate
discarded, still-unsequenced chunks into local-order/raw-byte loss records and
replace them with a `CLIENT_OVERLOAD` truncation marker; event-range and
compressed-byte counts are absent when sequencing/compression never occurred.
Use fair worker scheduling and reserved metadata capacity so lifecycle,
stdin/signal
acknowledgements, gap markers, and terminal state remain lossless.
Loss mode is hysteretic and scoped: crossing either command or daemon high
watermark marks affected raw chunks as dropped until both relevant backlogs fall
below their lows. Coalesce only adjacent dropped local ordinals with identical
cause/stream set; close a loss run before inserting any essential event. Persist
known raw-byte totals and observation interval in local metadata, then emit one
bounded `OutputTruncation`; omit event-range/compressed-byte fields if they were
never assigned/compressed. If even reserved marker persistence fails, mark the
command/store dirty and interrupt rather than silently claiming complete output.
Respect the local hard limits at all times:
- 10 MiB compressed rolling output window and 32 MiB total per command;
- 256 MiB aggregate command-owned storage per client daemon.
When connected, discard server-acknowledged oldest segments first. Assigned but
unacknowledged entries are pinned and bounded by the 1 MiB per-command send
window. When offline or an acknowledgement lags, rotate only still-unsequenced
output as needed and replace each removed run in durable local order with
`OutputTruncation` using `CLIENT_SPOOL`. The marker later receives a normal
wire sequence and is retained until cumulative `EventAck`; no wire sequence is
fabricated or skipped. The server's accepted truncation event is the auditably
visible explanation for the lost bytes.
Continue draining child pipes after caps are reached. A verbose child may lose
old or newly produced output, but it must not block the client daemon or
deadlock itself.
Test a continuously verbose process, alternating stdout/stderr, essential events
interleaved with dropped output, reconnect during loss mode, ack loss at every
send-window boundary, and terminal exit while above the high watermark. Assert
contiguous assigned sequences, exact known raw-byte totals, bounded heap/queue
size, and a terminal event after every marker/incomplete event.
### 7.3 Command acceptance and idempotency
Persist an accepted dispatch record keyed by `issue_uuid` before responding
`CommandAccepted`. Repeat dispatch with identical immutable fields returns the
same acceptance/revision and does not start another child; a conflicting repeat
is a protocol error. Enforce 16 running/100 queued defaults before acceptance.
After full terminal cleanup, also check the compact tombstone ledger: the same
UUID/hash is rejected as already executed, while changed immutable content is a
conflict.
Make acceptance a single durable state transition:
1. Validate the session generation, UUIDv7 syntax, immutable request hash,
command revision, shell/profile/CWD policy, script descriptor, and expiry.
2. Look up both the active-command table and tombstone ledger before reserving
capacity. An exact active duplicate returns its existing result; an exact
tombstone returns `CODE_ALREADY_EXECUTED`; either hash mismatch is a
permanent conflict. The server treats the former as stale-state
reconciliation/interruption, not as a newly rejected execution.
3. Reserve the running/queued slot and worst-case command metadata allocation in
the same local-store transaction that inserts the accepted command. Do not
count retransmission of an existing record twice.
4. Commit and fsync the accepted record, immutable hash, revision, script
cursors, and capacity counters before sending positive acceptance. A failure
before commit returns a transient rejection and leaves no partial command;
loss of the response is resolved by exact replay.
5. Classify policy/schema/hash/unsupported-profile failures as permanent and
local I/O/temporary-capacity failures as transient. Include a bounded,
operator-readable reason without command or environment contents.
For script dispatch, persist descriptor and chunks at validated offsets, verify
each chunk checksum and final SHA-256/length before launch, and keep the upload
restart-safe until terminal cleanup. Create the final temporary file with
owner-only permissions under the effective CWD using an atomic write/rename;
delete it after the terminal event is durable locally. Emit cumulative durable
`ScriptUploadStatus.received_bytes` progress often enough to advance the 1 MiB
server send window; the server never streams the full 10 MiB without progress
acknowledgement.
Accept a chunk only when its offset equals the durable contiguous prefix, or
when the entire byte range is an exact replay of already persisted bytes.
Reject holes, changed overlaps, arithmetic overflow, data past the declared
length, and chunks received after commit. Append and fsync bytes before
advancing `received_bytes`; after the final prefix, verify length and SHA-256,
reserve the raw-file quota, atomically materialize the wrapper, sync its parent
directory, and only then make the command launch-eligible. A permanent upload
failure durably ends the command as `REJECTED` before launch authorization and
releases every associated reservation.
Materialize ordinary `command_text` through the same private generated-file
machinery, while retaining its separate command-text metadata. Execute the file
with exactly `sh`/`bash`, `cmd.exe /D /S /C`, or `powershell.exe` with
`-NoLogo -NoProfile -NonInteractive -File` according to `shell_type`; never infer or
fallback to another shell. Build Windows application name/arguments with the
platform quoting routine rather than shell string concatenation. Treat `$0`,
`%0`, or equivalent exposing the generated wrapper name as documented behavior.
Resolve each shell setting once during configuration validation to a canonical
absolute executable path. Launch that exact path with an explicit argument
vector/application name; never search the daemon's `PATH`, honor a command's
environment overrides during resolution, follow a wrapper shebang, or select by
file extension. Re-stat the configured executable immediately before launch and
reject if it is no longer the validated regular executable. Add tests with an
attacker-controlled `PATH` entry and same-named fake shell to prove it cannot be
selected.
### 7.4 Unix process supervisor
Define a narrow interface used by the runtime:
```go
type Supervisor interface {
Start(context.Context, StartSpec) (Process, error)
Signal(context.Context, Process, SignalKind) (SignalOutcome, error)
Snapshot(context.Context, Process) (ResourceSnapshot, error)
StopAll(context.Context) error
}
```
The Unix implementation starts an RVBox launcher as leader of a new
session/process group; after authorization that launcher creates the selected
`sh`/`bash` child inside the same group and remains as its watchdog. It rejects
unsupported shell values. It applies daemon environment plus persisted
overrides, validates the CWD, connects stdin/stdout/stderr pipes, and records
launcher/root/group identities. Signal the full process group.
On orderly shutdown and unclean-start recovery, terminate surviving managed
groups and emit `interrupted` rather than pretending pipe monitoring survived.
Treat root exit as the start of a configurable 5-second tree/output drain grace
period. Wait for the cgroup/process group and capture readers; terminate residual
descendants after the grace period, drain to EOF, and emit incomplete-output
metadata as a sequenced `OutputIncomplete` event if handles still cannot be
drained. The terminal lifecycle event must be sequenced and spooled only after
all retained output/truncation events and is
always the final client event.
Implement launch through an internal blocked-launcher mode with a private
release/watchdog channel. The launcher remains alive as a non-user-code
supervisor for the process-tree lifetime. After the launcher reports ready,
persist and flush
its PID/process-group/platform birth identity as `launch_prepared`; persist and
flush `launch_authorized` before releasing requested command code. EOF before
release aborts without executing it. On Linux create a per-command cgroup v2
for supervision regardless of resource-profile selection, place the blocked
launcher into it before release (`clone3` with `CLONE_INTO_CGROUP` and
`CLONE_PIDFD` where available), and persist the cgroup path plus
`/proc/<pid>/stat` start time. Recovery prefers `cgroup.kill`; a process-group
fallback is allowed only after positive birth-identity verification. Other Unix
platforms use the watchdog/process-group fallback and document that descendants
which deliberately create a new session may escape it.
Encode launch as `accepted -> launch_prepared -> launch_authorized -> running`
and permit no shortcut:
1. Create the private cgroup/process group, pipes, wrapper, and close-on-exec
release/watchdog channel without executing user code.
2. Start the blocked launcher and obtain a positive ready message containing
the platform process identity; place and verify it in its command cgroup.
3. Commit and fsync `launch_prepared` with that identity. Recheck the latest
revision and pending cancellation while the launcher remains blocked.
4. If still executable, commit and fsync `launch_authorized`, then send the
one-byte release and keep the watchdog channel open until tree cleanup.
Record `running` only after the launcher reports successful requested-shell
creation/`exec`.
5. A crash before durable preparation cleans an untrusted orphan and may retry.
A prepared-but-unauthorized launcher must exit on channel EOF and may retry
only after positive death verification. After release, channel EOF makes the
launcher terminate its process group; an authorized but uncertain launch is
killed as a tree and ends `interrupted`, never launched again.
Make launch-journal and launcher-control records checksum-framed and bounded.
Fault-inject process death before and after every fsync, ready/release, and exec
acknowledgement; assert that user-visible side effects occur at most once and
that recovery cannot confuse a reused PID.
Implement the portable Unix set `HUP`, `INT`, `TERM`, `KILL`, `USR1`, and
`USR2`, accepting optional `SIG` prefixes in the CLI and mapping only through
the `SignalKind` enum. Reject arbitrary native numbers and unsupported names.
Never allow signal zero or arbitrary PID targeting. The process's group ID
comes only from durable supervisor metadata, never a request field.
### 7.5 Linux diagnostics and resource profiles
Poll active processes at a configurable interval outside pipe/network loops.
Read readable `/proc` values for CPU time, RSS, I/O, state, CWD, and wait reason;
aggregate only clearly associated group/child data. Missing/unreadable values
remain absent. Set `suspected_hung` only after default 10 minutes without
observable progress and label it diagnostic, not lifecycle.
Make requested resource-profile handling explicit. On Linux the supervisor
uses a delegated writable cgroup v2 for tree supervision whenever available,
even without a profile. Profiles add administrator-defined CPU/memory/disk/
process controls to that cgroup. Without delegation, no-profile execution uses
the watchdog/process-group fallback; a profile whose essential control cannot
be applied is rejected as unsupported. No profile means no resource restriction,
even when a supervisory cgroup exists.
Compile TOML resource profiles into immutable validated launch policies during
startup. Enforce `LIGHT` as exclusive; otherwise allow at most one CPU, one
memory, and one disk profile. On Linux, translate configured policy into the
per-command cgroup's `cpu.max`/`cpu.weight`, `memory.max`/`memory.high`,
`pids.max`, and per-device `io.max` controls as applicable. Write and read back
all essential controls while the launcher is blocked. If any essential control
is unsupported or cannot be applied, destroy the empty command cgroup and
permanently reject the dispatch; never run partially constrained.
Define diagnostic progress as a change in process-tree CPU ticks, cumulative
I/O counters, retained output/input activity, or lifecycle state. Persist only
the latest sample and no-progress start time, clear `suspected_hung` on the next
observed progress, and tolerate counter reset/process exit. Diagnostics must not
keep a command alive, change terminal status, or block the supervisor.
### 7.6 Client tests
Use fake transport/clock plus real subprocess tests for shell defaults, CWD/env
overlays, at-most-once duplicate dispatch, concurrent output with no newlines,
stdin ordering/close, process-group termination, reconnect/replay, offline
rolling truncation before sequence assignment, assigned-window pinning, script
progress/checksum failure/cleanup, atomic tombstone replacement, queue limits,
shutdown interruption, and `/proc` absence. Run race tests with multiple
commands and forced network churn.
Add a table-driven crash suite for every acceptance, script-upload, launch,
event-spool, terminal-acknowledgement, and tombstone-rotation commit point. Run
the same duplicate dispatch after each restart and assert exactly one of:
durable rejection before authorization, one supervised process, or an
`interrupted` uncertain launch—never a second execution.
**Exit criteria:** a Linux client can stay alive through server loss/restart,
execute up to its capacity at most once, preserve/replay bounded history, and
cleanly manage full command process trees.
## 8. End-to-end agent protocol (Phase 5)
Wire the server session layer and client runtime together before adding the CLI.
Use real protobuf bytes through an in-memory WebSocket test server first, then
through Docker Compose with nginx proxying a WSS endpoint.
Implement flows in this order:
1. hello/welcome/version/fencing and heartbeat;
2. persisted background command dispatch and lifecycle events;
3. stdout/stderr event persistence and cumulative acknowledgements;
4. reconnect replay, duplicate delivery, and non-terminal reconciliation;
5. queued cancellation and revisioned signal delivery;
6. ordered stdin/close acknowledgement;
7. verified chunked scripts;
8. capacity advertisements, diagnostics, and profile results.
At each step, add a failure-injection test that drops one frame at every
acknowledgement boundary. Required scenarios include lost `CommandAccepted`,
lost `EventAck`, stale old-session output after takeover, connection loss during
script transfer, server restart before/after event transaction, client restart
with managed process cleanup, and offline spool overflow. Verify command UUIDs
never execute twice and output gaps always appear as truncation metadata.
Build a deterministic scenario harness with fake wall/monotonic clocks,
controllable frame loss/reordering, daemon kill points, and bounded virtual disk
capacity. For each flow, run disconnect before delivery, after delivery/before
durable commit, after commit/before acknowledgement, and after acknowledgement.
Assert database invariants and quota counters after every restart, not merely
the visible CLI result. Include these cross-cutting cases:
- a queued command expires at the default 15-minute TTL while disconnected,
then a late client reports a real terminal result and `stat` displays both
expiry history and the updated terminal state;
- server per-command, per-client, and global quotas cross high and low
watermarks while assigned client events remain pinned;
- terminal age reclamation and command-level quota eviction create tombstones,
retention metadata, and audit entries without removing active work;
- client identity collision, exact-instance takeover expiry/consumption, and
an old connection attempting output after the new generation is committed;
- storage corruption found during asynchronous recovery, safe repair, explicit
acknowledgement of unavoidable loss, and continued service in healthy scopes;
- heartbeat control traffic remains serviceable while both data lanes are full
over a slow, intermittently writable WebSocket.
**Exit criteria:** Compose tests demonstrate a complete background command,
foreground follow, stdin interaction, signal, reconnect, and restart recovery
through the nginx WebSocket path.
## 9. Server control plane and `rvc` (Phase 6)
### 9.1 gRPC over Unix socket
Start a Unix socket only after removing an existing socket only when it is
proven to be a stale socket owned by this server account; do not blindly unlink
arbitrary paths. Set mode `0600` immediately. Implement the generated `Control`
service against domain/store/session interfaces, never directly against socket
state.
Control response messages represent success only and contain no embedded error
alternative. Return canonical non-OK gRPC statuses with structured RVBox
details; the JSON-RPC adapter maps the same domain error rather than inspecting
response payloads. Agent-protocol rejection envelopes continue using
`ControlError`.
Implement and test:
- `ListClients`, `GetClient`, `ListCommands`, `GetCommand`;
- `AuthorizeClientTakeover` for an exact displayed pending instance, with a
one-shot 5-minute grant and normal mutation idempotency;
- `RunCommand` for durable creation and immediate UUID return, followed by
`FollowCommand` for foreground ordered history plus live events;
- `FollowCommand` beginning at a supplied event sequence and returning a wrapper
around either a client event or non-sequenced server retention metadata;
- `AppendStdin`, `CloseStdin`, and `SignalCommand` with persisted idempotency;
- `GetOutput` with stream selection, byte-bounded opaque cursor pagination that
can resume within an event, distinct client-observed/server-receipt times, and
truncation/incomplete metadata;
- `ListStorageIncidents`, `RepairStorageIncident`, and
`AcknowledgeStorageIncident`; dirty health is derived from unresolved rows,
safe repair resolves automatically, and acknowledging loss preserves history.
Foreground context cancellation/timeouts stop only the local RPC stream. They
must never send a remote kill. The server uses domain/store authorization of
state transitions and an explicit command UUID for every mutation.
Foreground `rvc run` first completes idempotent unary `RunCommand`, then starts
`FollowCommand` from sequence zero with existing history included. Reconnects
resume after the last received client sequence; retention markers do not advance
that cursor and are emitted before a relevant terminal event. Do not add a
combined create-and-follow RPC whose stream can fail before returning the newly
created UUID.
For every mutation, canonicalize the validated protobuf request excluding
transport metadata, compute its immutable hash, and transact an idempotency row
keyed globally by `request_id`, with method included in the hashed/stored
identity. An exact replay returns the stored domain result
or stored error; a different hash returns `ALREADY_EXISTS`/request-conflict and
performs no work. Commit the idempotency row atomically with the mutation, retain
it at least as long as the referenced command/takeover/incident record, and
count its compressed storage against the owning command or audit scope. Reads
accept the field at the transport edge but deliberately drop it before domain
execution.
Define opaque output/page tokens as versioned, authenticated encodings of query
filters plus stable position (command, event sequence, byte offset, snapshot
boundary). Reject a token reused with changed filters. Pagination reads a
consistent upper boundary so concurrent appends do not duplicate/skip retained
bytes; if retention removes the resume point, return the next retained position
with explicit server-retention metadata.
### 9.2 CLI
Implement `rvc` with no direct server database access. Required commands are:
- `rvc stat [client-id] [issue-uuid] [--all --page --per-page]`;
- `rvc run [--background] [--cwd DIR] [--shell TYPE] [--env K=V]...
[--profile FLAG]... [--request-id UUID] [--queue-ttl DURATION] CLIENT COMMAND`;
- `rvc run --script PATH [same execution options] CLIENT`;
- `rvc append CLIENT UUID TEXT`, `--file PATH`, raw/no-newline mode, and
`--attach` stdin streaming;
- `rvc close-stdin CLIENT UUID`;
- `rvc kill [HUP|INT|TERM|KILL|USR1|USR2] CLIENT UUID` (optional `SIG` prefix;
default `TERM`);
- `rvc storage incidents`, `rvc storage repair INCIDENT`, and
`rvc storage acknowledge INCIDENT --note TEXT`;
- `rvc client takeover CLIENT INSTANCE-ID`;
- output selection, `--timestamped`, pagination, and `--follow` semantics.
Accept `--request-id` globally. Generate a UUIDv7 for every control mutation
unless supplied, reuse it across transport retries, and discard it for
read-only operations.
Create the request ID before dialing the Unix socket and retain it for all
automatic retries of that invocation. Validate user-supplied IDs as canonical
UUIDv7 strings. For `run`, map the same value directly to `issue_uuid`; for
other mutations keep it only as `request_id`. Do not generate a second
operation identifier. Print the issue UUID immediately after durable creation,
including before entering foreground follow mode.
Make `stat` render lifecycle and history independently: an expired queued
command is visibly marked `expired` with its expiry time/reason; if a late,
previously accepted client result later arrives, the actual terminal lifecycle
becomes current while the expiry contradiction remains in metadata/audit. Also
show dirty storage scope/incident ID and pending takeover instance,
output retained range, and whether terminal output is incomplete.
Render output without inventing line boundaries in stored data. Historical
pagination uses opaque cursors and byte limits; the display layer may buffer
partial lines for presentation and clearly prints retained range/truncation
notices. Map structured control errors to stable non-zero exit
codes while retaining machine-readable JSON output as a later optional CLI mode.
### 9.3 JSON-RPC adapter
Keep this adapter small and disabled by default. Use standard JSON-RPC 2.0 over
HTTP and protobuf JSON mapping for unary control request/response bodies.
Implement the documented lower-camel method names only, including storage
incident list/repair/acknowledge and `authorizeClientTakeover`. Bind default
`127.0.0.1:6900`; configuration may bind elsewhere but startup logs a conspicuous
unauthenticated-exposure warning. Do not add streaming or a second event model:
callers poll `getOutput`/`getCommand` by cursor.
Audit every control request with transport/source metadata where available,
action, target, result, and error. Audit logs include environment key names but
not override values or raw stdin/stdout/stderr.
Use the JSON-RPC envelope `id` only for JSON-RPC response correlation. Accept
the same optional protobuf `requestId` member as gRPC for domain idempotency;
when it is absent, generate a server UUIDv7. Consequently raw JSON-RPC remains
easy to use but only callers that persist/reuse `requestId` receive retry
idempotency. Map parse/invalid-request/method-not-found to standard JSON-RPC
codes and domain failures to one stable server-error code carrying the same
structured RVBox detail as gRPC.
**Exit criteria:** `rvc` can drive each documented example against the Compose
stack; the Unix socket has mode `0600`; JSON-RPC behavior matches gRPC unary
semantics and is off unless explicitly enabled.
## 10. Windows client implementation (Phase 7)
Keep Windows code in platform-specific files/build tags so Linux builds never
import Windows APIs. Implement the same supervisor interface and event/spool
runtime; only OS execution/diagnostics/resource enforcement differ.
1. Create a non-inheritable per-command Job Object, enable
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`, and do not enable breakaway.
2. Start an RVBox launcher suspended with `CREATE_NEW_CONSOLE`,
`CREATE_UNICODE_ENVIRONMENT`, and `EXTENDED_STARTUPINFO_PRESENT`. Assign it
to the Job at creation through `PROC_THREAD_ATTRIBUTE_JOB_LIST`; inherit only
an explicit standard-I/O/control handle list and never the Job handle.
3. Persist and flush `launch_prepared` with launcher PID, `GetProcessTimes`
creation `FILETIME`, and launch generation; persist and flush
`launch_authorized` before `ResumeThread`. The launcher then creates the
requested shell suspended in its console with `CREATE_NEW_PROCESS_GROUP`,
reports its PID/group, connects the allowlisted pipes, and resumes it. A
failure after authorization terminates the Job and becomes interrupted.
4. Keep the launcher as an in-console signal proxy. This separation is
mandatory: Windows ignores `CREATE_NEW_PROCESS_GROUP` when combined with
`CREATE_NEW_CONSOLE`, and `GenerateConsoleCtrlEvent` reaches only groups
sharing the caller's console. On daemon crash, closing the sole Job handle
terminates the launcher and complete tree; restart never signals by PID alone.
5. Expose Job Object accounting snapshots. Children normally remain in the Job.
6. Accept only `SIGTERM` and `SIGKILL`. For TERM, have the launcher call
`GenerateConsoleCtrlEvent(CTRL_BREAK_EVENT, shell_group_id)`, wait 10
seconds, then terminate the Job. For KILL, terminate the Job immediately.
Report graceful-attempt/escalation outcome in the event.
7. Apply requested profile limits through Job Object limits. If the requested
control cannot be applied, reject the dispatch with structured `UNSUPPORTED`.
8. Use available process/Job Object telemetry for CPU/RSS/I/O; do not emit a
Linux-style wait reason.
Implement the Windows launcher as a small mode of the same RVBox binary, not a
searchable external helper. Pass it an inherited, random per-launch control
pipe and fixed-size launch-generation token. Apply an explicit
`PROC_THREAD_ATTRIBUTE_HANDLE_LIST` so only stdin/stdout/stderr and that control
pipe cross creation; make every database, log, listener, Job, and unrelated
pipe handle non-inheritable. Frame ready/release/exec/error messages with length,
type, generation, and checksum, and reject any mismatched generation.
Open the configured shell executable by canonical absolute path and pass it as
the explicit `lpApplicationName`. Produce the command line with one reviewed
Windows argument-quoting routine and construct an explicit Unicode environment
block. Do not invoke `%COMSPEC%`, search `PATH`/file associations, or let command
environment overrides influence executable selection.
Associate the Job with an I/O completion port and use
`JOB_OBJECT_MSG_ACTIVE_PROCESS_ZERO` plus pipe EOF for tree/drain completion.
Persist the launcher and shell creation `FILETIME` identities before trusting
any PID. On recovery, reopen a process only to compare creation time and
generation evidence; if ownership cannot be proven, record a dirty incident
rather than terminating an unrelated reused PID.
Fault-inject daemon/launcher death around Job creation, attribute-list process
creation, prepared/authorized fsync, resume, requested-shell creation, and
terminal drain. Test both supported shells, paths with spaces/non-ASCII,
malicious `PATH`/`COMSPEC`, nested descendants, CTRL_BREAK refusal/escalation,
Job limit rejection, and inherited-handle leaks on supported Windows versions.
Run build checks on every platform and dedicated Windows integration tests for
shell selection, process-tree kill, forced termination, profile rejection,
reconnect, and output/spool behavior. No Windows release is supported until
these tests run on Windows CI.
## 11. Reliability, observability, and operational delivery (Phase 8)
### 11.1 Health, metrics, and logs
Expose separate liveness/readiness for the server. Liveness and incident
inspection start before asynchronous recovery. Global readiness requires all
required scopes recovered and writable, while scoped operations can proceed on
healthy scopes; no health state requires a client connection. Client health
reports reconnect state, spool health, unresolved incidents, and supervisor
health without leaking command output.
Publish counters/histograms/gauges for registrations/takeovers, stale messages,
heartbeat timeouts, reconnect duration, dispatch latency, command transitions,
queue depth, spool bytes, output compression/rotation/loss, segment eviction,
SQLite write latency/failure, script verification failure, and protocol errors.
Use stable labels with bounded cardinality—never UUID/client ID as a metric
label. Use structured logs and audit records for those identifiers instead.
### 11.2 Deployment assets
Provide:
- nginx configuration showing WebSocket upgrade proxying and TLS termination;
- systemd units for server/client with private state directories, restart
policy, working directory, file descriptor limits, and least privilege;
- example server/client configuration files with every default and an explicit
JSON-RPC exposure warning;
- a backup/restore procedure for SQLite plus output/audit segment directories;
- an upgrade procedure that stops dispatch safely, snapshots data, migrates,
and verifies recovery;
- a troubleshooting runbook for no client, stale session, spool full, output
truncation, storage full, and daemon restart.
Install the annotated TOML examples from `docs/examples/` and keep the most
operationally important identity/listener/state/TLS/quota keys first in each
section. Add `--check-config` to both daemons: it must run the production strict
decoder, defaulting, cross-field/profile/path/shell validation, print a redacted
normalized summary, and exit without opening stores/listeners or changing
state. CI parses both examples with this path on Linux; Windows CI additionally
validates the documented Windows path/shell variant.
Use `Delegate=yes` in the Linux client systemd unit when cgroup supervision is
enabled and create a private writable cgroup subtree for the service. Refuse a
configured mandatory resource profile when delegation/controllers are absent,
while still allowing unrestricted execution through the documented fallback.
Package upgrades must retain the previous TOML, run `--check-config` before
restart, and never silently ignore a newly unknown/removed key.
### 11.3 Security and regression review
Before v1 release, review path traversal, script temp-file permissions, command
logging, environment override redaction, malformed compression, decompression
bombs, oversized frames, Unicode/ASCII ID validation, Unix socket ownership,
JSON-RPC external bind warnings, SQL injection (all parameterized), segment
record corruption, and process-group PID reuse. Add fuzz tests for envelope
decode, compressed output validation, segment-tail recovery, pagination tokens,
and JSON-RPC parsing.
**Exit criteria:** release Compose/systemd/nginx assets exist; load/fault tests
show no unbounded memory/goroutine growth; the operations runbook reproduces
recovery and retention behavior.
## 12. Release gates and implementation order
The recommended merge order is deliberately vertical:
1. Phase 0: Docker-first repository/toolchain/proto generation.
2. Phase 1: domain/config validation and state machine.
3. Phase 2: SQLite/segments/audit/retention with recovery tests.
4. Phase 3: server registration, fencing, heartbeat, and persisted dispatch.
5. Phase 4: Linux client spool/supervisor and at-most-once execution.
6. Phase 5: full WSS protocol, fault injection, and nginx end-to-end tests.
7. Phase 6: gRPC Unix socket, `rvc`, and optional JSON-RPC.
8. Phase 7: Windows supervisor and Windows CI.
9. Phase 8: operational assets, stress/fuzz/recovery testing, release review.
Do not merge a later vertical slice by stubbing a durability/safety invariant.
For example: foreground mode may wait on a durable background command, but must
not bypass persistence; client output may be truncated under the documented
caps, but must never block a child pipe; and a reconnection may replay work,
but may never re-execute an already accepted UUID.
The v1 release is ready only after all release-gate tests pass on clean Docker
environments, Linux server/client end-to-end behavior matches the design docs,
Windows support is either fully tested or explicitly not shipped, and every
accepted limitation (self-reported identity and unauthenticated optional
JSON-RPC) is conspicuous in deployment documentation.