157 KiB
RVBox v1 Go implementation plan
1. Objective and implementation boundary
Implement the v1 contracts in the existing design documents as three Go binaries:
rvbox-server: durable server, WebSocket agent endpoint, local gRPC control endpoint, optional JSON-RPC adapter, storage, and audit owner.rvbox: reverse-connecting client daemon, command supervisor, durable active spool, and reconnect/reconciliation owner.rvc: local CLI over the server's Unix-domain gRPC socket.
The first supported client target is Windows connecting to a Linux server; Linux client support remains part of complete v1 but is explicitly deferred from the current implementation effort. Implement the Linux server and native Windows client first; do not begin Unix-like client code now. Build shared client runtime code behind OS interfaces without prematurely implementing the Unix supervisor. Windows v1 requires Windows 10 or newer, or Windows Server 2016 or newer. Desktop Experience is required only for the tray and active-session command contexts; the service and Session 0 execution contexts support headless Server Core.
The released rvbox.exe is one signed binary with explicit SCM-service, tray,
installer/configurator, per-command launcher, and signal-helper modes.
Installation is UAC-elevated and registers one machine-wide Automatic service
running as LocalSystem. That single main service owns WSS, all accepted
requests, durable state/history, logs, token policy, Job Objects, process
handles, and recovery before any user logs in. There is no separate persistent
worker or elevation broker.
The installer also registers an unelevated per-user tray through the machine-
wide Run key. A tray is only a named-pipe frontend to the service: it may show
health, open config/log, and request locally authorized service actions, but it
never opens the store or owns remote work. Closing it affects neither service
nor commands. Task Scheduler is not used anywhere in v1.
Use conventional machine-wide paths: %ProgramData%\RVBox\client.toml,
%ProgramData%\RVBox\state, and %ProgramData%\RVBox\logs\rvbox.log, with
state ACLs limited to Administrators and SYSTEM and separate read-only tray
access for config/log. The menu opens the exact config or current log via a safe
explicit file-opening path. SCM Automatic/Manual state is
the only service-startup source of truth and is not mirrored into TOML. Create
%ProgramData%\RVBox\work as a SYSTEM-owned root. For omitted CWDs, create an
ACL-isolated child for the selected user SID, LocalService, or SYSTEM context;
do not make one shared work directory writable by every identity. These
requirements are part of Phase 4, not post-v1 polish.
This plan implements the already agreed v1 contract. In particular, it does not add authentication, client enrollment, mutual TLS, or a client allowlist. The deployment trust boundary remains nginx TLS termination and restricted network access. The unauthenticated JSON-RPC listener remains disabled by default and loopback-bound when enabled.
2. Development rules and definition of done
2.1 Local-environment rules
../ENV.md is binding for this repository:
- Do not install Go, Buf,
protoc, SQLite tooling, or other heavy development dependencies on the host. - Put the toolchain, code generation, unit/integration testing, linting, and local services in Docker/Docker Compose.
- Run container commands as UID/GID
1001:1001so generated Go code, module caches mounted into the project, and test artefacts remain host-owned. - Do not delete Docker volumes, generated artifacts, or state directories as a convenience cleanup action. Tests use dedicated temporary volumes/directories and remove only the exact resources they created.
2.2 Completion standard for every phase
A phase is complete only when all of the following are true:
- The public behavior and failure modes are covered by focused tests.
- The package has context cancellation, bounded channels/queues, and no network or disk operation on the WebSocket receive loop.
- Structured logs and metrics exist for the new state transitions/failures.
- Configuration defaults, validation, and errors are documented.
docker compose run --rm toolchain make fmt lint testpasses, with no host installation required. The lint target explicitly snapshots the accepted first-draft Buf enum-prefix/service-suffix findings and fails on any other or newly added finding; application/code lint has no such exception.- Any recovery or retention code is tested with a real temporary SQLite database and filesystem rather than mocks alone.
2.3 Test layers and required coverage
Every test belongs to exactly one named layer. Do not call a test “unit” merely
because it is written with Go's testing package, and do not hide system setup
inside an ordinary package test.
Unit tests are hermetic, parallel-safe, and require no RVBox daemon, project
runtime/sidecar container, network listener, wall-clock delay, administrator
privilege, or shared machine state. They execute inside the toolchain container.
Use injected clocks, deterministic randomness/readers, in-memory fakes, and
t.TempDir(). A normal unit invocation must be safe to repeat and must leave no
repository artifact. Add table, property, race, and fuzz coverage as applicable
for:
- every legal and illegal lifecycle/revision transition, terminal predicate, reconciliation row, error mapping, and UUIDv7/request-hash idempotency rule;
- strict TOML decoding/defaulting/cross-field validation, path and shell resolution policy, resource-profile composition, redacted display, and all hard size/count/duration boundaries;
- protobuf envelope/field validation, compression limits, event sequencing, truncation/incomplete-output semantics, cursor encoding, and malformed input;
- quota charging, reservation, whole-command eviction choice, age rotation, tombstone ordering, audit rotation, and filesystem-floor admission decisions;
- dispatch capacity reservations, fairness, heartbeat/backoff boundaries, priority-lane scheduling, cancellation/signal races, and goroutine shutdown;
- client spool ordering, retransmission, acknowledgement, raw-output high/low-watermark loss conversion, script chunk validation, and crash-state decision logic;
- the pure Windows login-state/context-selection matrix, token-property validation, command-line quoting, environment overlay, path/ACL construction, named-pipe message validation, tray authorization decisions, and SCM desired- state transitions. Win32 calls themselves belong in native integration tests.
Unit tests use stable per-test seeds printed on failure. Timing assertions use
fake monotonic clocks; a small explicit set of real scheduler/deadline tests may
use bounded polling but never arbitrary sleeps. go test -shuffle=on and focused
-race jobs must pass repeatedly. A retry wrapper may collect flake evidence but
must never turn a failed assertion into a passing gate.
Integration tests exercise one real subsystem boundary while keeping remote peers controllable. Linux integration tests run in the project Compose stack and use real SQLite/WAL, append-only segment files, Unix sockets, HTTP/WebSocket listeners, compression, and compiled binaries where the boundary requires them. Native Windows integration tests run the compiled client/service/helper modes against a deterministic fake agent server and use real SCM, WTS/token APIs, named pipes, Job Objects, consoles, files, and registry entries. Integration tests may inject storage/network/process faults through documented harness controls, but may not mock the component whose contract is under test.
Required integration suites are store, server-session, control,
client-spool, windows-supervisor, windows-service-tray, and upgrade- recovery. They cover at minimum:
- all commit/crash boundaries, reopen/recovery, corrupt or short committed ranges, disk-full/permission failures, repair/acknowledgement, quota/retention, and concurrent readers/writers against real temporary storage;
- real gRPC/JSON-RPC/WebSocket serialization and limits, session fencing, dispatch/reconciliation, lost acknowledgements, slow peers, priority control traffic, pagination, and idempotent mutation replay;
- native Windows context selection for each supported login/token state, pre-launch fallback only, launcher authentication, process-tree ownership, stdout/stderr/stdin, signal escalation, shell/CWD/environment behavior, service/tray authorization, boot/logon/logoff, and cleanup after forced death.
End-to-end tests start the production-shaped Linux server, nginx TLS proxy,
local control plane, and the signed-or-CI-built native Windows service client.
They drive behavior only through rvc, supported service/tray controls, and
intentional fault-harness controls. They use real protobuf bytes over WSS and
real persistent stores; do not replace the Windows client with Wine, a Linux
client, or a protocol stub. Required named scenarios are:
smoke: install/start, registration, background and foreground commands, stdout/stderr/history, terminal status, and orderly uninstall preserving data;interactive: stdin append/raw/close, TERM then KILL behavior, scripts, CWD, environment overrides, and output pagination/follow;idempotency: repeatedrequest_id/issue_uuid, lost control responses, duplicate dispatch/events, and conflicting request hashes;reconnect: nginx/server/network interruption, old-session fencing, client restart, reconciliation, and no duplicate execution;retention: command/client/global quotas, offline spool loss markers, terminal-age reclamation, tombstones, audit rotation, and disk floors;expiry-and-incidents: offline 15-minute expiry, late client truth, corrupt storage, scoped dirty health, repair, and acknowledgement;windows-contexts: logged-out normal/elevated execution, active standard and split-token administrator sessions, ordered elevated fallbacks, ambiguous active sessions, and effective-identity status/audit output;service-tray: boot before login, per-session tray behavior, UAC-protected configuration actions, Explorer restart, service continuity, and log rotation;soak: bounded concurrency, slow/unstable network, verbose output, repeated reconnects, and resource-leak checks for the release gate.
Each layer has an explicit timeout budget. Unit and integration suites must be shardable by stable suite/case names. E2E scenarios are serial on a particular Windows host because they mutate machine-wide SCM/Run-key state; parallelism is allowed only across independently leased resettable VMs.
2.4 Reproducible and resumable test harness
Implement the following checked-in entry points; CI calls the same scripts that developers call rather than duplicating orchestration in workflow YAML:
scripts/test-unit [--package PATTERN] [--run REGEXP] [--race]
scripts/test-integration [--suite NAME|all] [--run-id ID] [--resume]
scripts/test-e2e [--scenario NAME|all] [--run-id ID] [--resume]
scripts/test-env doctor|coverage|status|logs|collect|recover|reuse|stop|reset|purge|gc ...
scripts/windows/test-host.ps1 Prepare|Status|Run|Collect|Stop|Reset
Scripts are thin, reviewed orchestration wrappers. scripts/test-unit invokes
the pinned toolchain Compose service as UID/GID 1001:1001. Integration and E2E
commands call a shared Go harness under test/harness so manifest parsing,
timeouts, process control, reporting, and cleanup logic are not reimplemented in
shell and PowerShell. The harness itself is built in the toolchain container;
no Go/Buf/protoc/SQLite SDK or package manager is installed on the host.
The first integration/E2E invocation creates a filesystem-safe random run ID,
or validates an explicitly supplied one, and records
.test-runs/<run-id>/run.toml. The manifest includes test layer/suite/scenario,
seed, repository commit and dirty-diff hash, tool/runtime image digests,
configuration/certificate hashes, Compose project name, allocated ports,
Windows VM/runner identity, phase state, and exact owned resource names. Store a
fsynced append-only step journal plus bounded artifacts beneath that directory.
Never record credentials, private keys, tokens, command environment values, or
unredacted payloads in the manifest/log bundle.
Integration binaries expose stable --list and --case names; the harness runs
one case process at a time and journals its setup, fault, reopen, invariant, and
cleanup checkpoints. E2E scenarios expose the same model at a larger step
granularity. SIGINT/SIGTERM/PowerShell cancellation handlers journal the
interruption and perform the bounded stop path. Emit both a human summary and
machine-readable JSON/JUnit results, preserving the original test failure even
when collection or cleanup also fails.
Use COMPOSE_PROJECT_NAME=rvbox-test-<run-id> and label every container, network,
and run-specific volume with the run ID and repository identity. Avoid fixed
host ports; allocate loopback ports once, persist them in the manifest, and
revalidate ownership before reuse. Hold a repository-local integration lock and
an exclusive Windows-host lease so two invocations cannot corrupt the same SCM,
registry, ProgramData, or port state. A per-run private CA/certificate is created
inside the exact run directory with restrictive permissions and destroyed by
that run's purge operation.
Pin base/runtime/fault-service images by digest and record the fully rendered
Compose configuration hash. scripts/test-env doctor is a read-only cold-start
preflight for Docker/Compose compatibility, UID/GID ownership, available CPU/
memory/disk, required cached-or-fetchable images, loopback port allocation, and
optional Windows-runner reachability/capability. It prints the exact missing
prerequisite and never installs host software or pulls an image unless the user
then runs a test command whose declared pull policy permits it.
Every setup and scenario step is named, bounded, and idempotent. After completing
a step, persist its input hash, outputs needed by later steps, and success marker.
--resume verifies the commit/diff hash, image digests, configuration, VM
identity, owned-resource labels, and last successful health checkpoint before
continuing. It refuses unsafe resume on mismatch and tells the operator to start
a new run; --resume never silently resets state. Fault injection uses persisted
named checkpoints rather than timing guesses, so a deliberate server/client kill
can be resumed and its post-restart invariant checks completed.
recover RUN_ID is distinct from resume: after verifying that the recorded
controller process/lease is dead, it reacquires the lock, reconciles the fsynced
journal with labelled Compose and Windows resources, completes or rolls back an
interrupted idempotent setup/cleanup step, and leaves the run in a declared
stopped-resumable or reset state. It never executes the next test assertion.
reuse RUN_ID creates a new run ID from the prior immutable image/configuration/
fixture definitions and reusable dependency caches, but allocates fresh ports,
PKI, SQLite/segment/spool storage, command IDs, and Windows test root. It never
copies prior mutable test state.
The Linux controller exposes bounded fault controls for frame drop/delay,
listener interruption, process termination at named durability checkpoints,
filesystem quota/error simulation, and monotonic/wall-clock advancement in the
deterministic peer. Native Windows test-host.ps1 validates administrator state,
OS build, UAC/policy fixture, active sessions, service identity, and test-root
ownership before running. It uses a test-specific ProgramData root and an
exclusive host lease; production-name SCM/Run-key installation tests run only in
a resettable VM snapshot. Passwords and VM access credentials come from the CI
secret store or interactive prompt, never arguments persisted in the manifest.
The E2E controller owns the canonical run manifest on Linux and passes the same run ID plus a one-run configuration bundle to the preconfigured Windows runner. The runner returns a checksummed step report and bounded artifacts through the CI transport or explicitly configured management channel; its credentials remain outside RVBox. Before executing a scenario, prove that Windows can reach the manifest's nginx WSS endpoint and trusts only the run's configured CA file, and that the Linux controller can observe service readiness. A lost runner is a stopped-resumable run, not permission to provision another Windows machine silently.
On success, integration/E2E runs collect a compact report and automatically stop
and remove their exact run-specific containers, networks, volumes, temporary
certificates, Windows service/Run entries, and scratch state. Reusable dependency
caches and shared immutable images remain. On failure or interruption, the
harness first collects bounded diagnostics, stops expensive processes/containers,
and preserves only the manifest plus state required for status, logs, or
--resume; keeping services running requires an explicit --keep-running.
scripts/test-env has these exact safety semantics:
doctor [--e2e]is read-only and validates local/container prerequisites and, when requested, the native Windows runner before allocating a run;coverageis read-only and validates the coverage inventory against listed unit/integration/E2E cases and their required host/layer metadata;status RUN_IDis read-only and reports steps, health, disk usage, owned Compose/Windows resources, and whether resume is valid;logs RUN_ID [COMPONENT]reads bounded/tail output;collect RUN_IDwrites a compressed, redacted diagnostic bundle without altering the environment;recover RUN_IDsafely adopts an abandoned owned environment and reaches a stable resumable or reset state;reuse RUN_IDcreates a fresh isolated run from the same immutable environment definition;stop RUN_IDstops only resources recorded in the validated manifest and preserves resumable state;reset RUN_IDremoves that run's runtime resources and scratch data but keeps its compact report/manifest; the run is no longer resumable;purge RUN_IDfirst prints the exact resolved targets and requires confirmation (or CI-only--yes), then removes that one run's remaining artifacts;gc --older-than DURATIONis dry-run by default, considers only validated completed/non-resumable RVBox manifests, and requires--execute; dependency caches require a separate--include-cachesconfirmation.
For deliberate low-disk cleanup, scripts/test-env purge --all enumerates every
validated RVBox run manifest and matching labelled resource, prints per-run and
total bytes, and does nothing until --execute plus confirmation. Reusable
dependency caches and test-only images are excluded unless separately requested
with --include-caches and --include-images; each inclusion is printed and
confirmed. This is the supported way to clear all RVBox test environments and
artifacts without performing a machine-wide Docker cleanup.
No script may invoke a global Docker builder/image/container/volume/system prune,
delete by an unvalidated glob, remove a repository/workspace root, or touch an
unlabelled resource. Resolve every deletion target, require it to be beneath
.test-runs or match both the persisted Compose project and RVBox labels, and
print it before deletion. Test cleanup must itself have unit tests plus an
integration test proving that unrelated containers, volumes, files, Windows
services, and registry entries survive reset/purge.
Keep the environment suitable for the resource-constrained host: build each immutable image once per source digest; use named Go module/build caches without copying dependency trees into the worktree; allow only one local integration or E2E run by default; set Compose CPU/memory/PID limits and suite deadlines; rotate component logs; cap collected logs/core/dumps; and run a disk-space preflight. Report exact cache/run/artifact usage and actionable cleanup commands whenever a preflight fails. Do not automatically delete caches or another run to make room.
Put overridable harness defaults in test/harness/defaults.toml, not scattered
environment variables. Initial local defaults are Go package parallelism 2; one
integration/E2E run; two aggregate runtime CPUs; 2 GiB aggregate runtime memory;
512 aggregate PIDs; 10 MiB collected log tail per component; 100 MiB total
failure bundle and 20 MiB success report; core/minidumps disabled unless a named
debug run enables them; 10-minute unit, 30-minute integration-suite, 45-minute
ordinary E2E-scenario, and 2-hour soak timeouts; and a 2 GiB free-space preflight
floor in addition to product-specific disk-floor cases. CI may explicitly raise
these budgets, but every run records and prints its effective values. Hitting an
artifact cap writes an explicit truncation manifest rather than silently
discarding diagnostics.
All scripts support --help, --dry-run for mutating environment operations,
stable nonzero exit codes, and a final summary giving the run ID, seed, report
path, resume command when applicable, and exact cleanup command. Document the
same workflow in docs/testing.md during Phase 0.
2.5 Traceable practical-scenario coverage catalogue
Create test/coverage.toml as the executable coverage inventory. Every concrete
case has a stable ID, class (happy, hostile-recovery, or race), owning
requirement/plan section, layer, suite, supported platforms, privilege/session
fixture, speed class, implementation phase, and case name. The ranges below are
shorthand; the inventory expands them into one row per case. scripts/test-env coverage rejects duplicate IDs, missing case implementations, cases that cannot
be listed by their suite, and normative requirements with no test. Test code
references the ID in its name or metadata, and failures print it.
Use line coverage as a warning signal, not as a substitute for scenarios. Nevertheless, pure policy/domain/protocol/config packages must maintain at least 90% statement coverage and core store/session/runtime packages at least 80%. Platform syscall wrappers report coverage separately and are judged primarily by native integration cases. Any excluded unreachable/generated/error branch needs a reviewed explanation in the inventory. CI reports coverage changes by package; a new critical branch may not lower coverage merely because the repository-wide percentage remains above threshold.
For every numeric/time/size limit, generate minimum, minimum+1, normal,
maximum-1, maximum, and maximum+1 cases where meaningful, plus zero,
negative decoded values where the source type permits them, integer overflow,
and unit-conversion boundaries. For every state machine, exhaustively generate
all state/event pairs and then add the concurrent interleavings listed below.
Pairwise generation covers independent dimensions such as shell, source type,
foreground/background, privilege intent, online/offline state, and output kind;
use a full Cartesian product only where interactions affect an invariant.
Happy-path cases
HP-CFG-01..12— Parse every annotated server/client/Windows TOML; apply compiled defaults and explicit flag precedence; normalize Unicode/space- containing absolute paths; validate Unix and Windows shell paths; advertise only current-platform shells; render redacted effective configuration; run--check-configwithout opening listeners/stores; create the first Windows template; start service live-but-not-ready for placeholder routing; restart into ready state after valid configuration; preserve config across upgrade.HP-PROTO-01..10— Negotiate exact and overlapping protocol ranges; ignore unknown optional fields; round-trip every envelope and control message; transfer maximum legal binary output/script chunks; Zstandard round-trip empty/compressible/incompressible data; accept Unicode command/output bytes; preserve client-observed and server-receipt timestamps; encode/decode every structured error and optional field without losing presence; store/forward the singleelevatedrequest bit unchanged while accepting Windows attempted/ effective contexts only as client-originated result metadata.HP-CTL-01..18— List/get clients and commands; run foreground/background command text and scripts; print UUID before following; resume follow; paginate commands and byte slices; select stdout/stderr; append text/raw/file stdin; close stdin; send TERM/KILL successfully to Windows and parse every portable signal name for later Unix delivery; inspect Windows identity, expiry, truncation, incomplete output, incidents, and takeover; repair and acknowledge incidents; exercise matching gRPC and JSON-RPC unary behavior; accept optional global--request-idand ignore it for reads.HP-IDEM-01..10— Generate UUIDv7 at the correct owner; accept externally supplied UUIDv7; return the same result for identical run/stdin/close/signal/ incident/takeover retries; map run request ID directly toissue_uuid; retain one mutation result through transport reconnect; suppress a matching command tombstone replay; preserve UUIDv7 FIFO order for equal-millisecond issuance.HP-SES-01..14— Register a new client; reconnect the same durable instance; fence its old connection; accept a different instance when none is live; authorize and consume an exact live-instance takeover; expire an unused grant; advertise/reconcile capacity; dispatch into running and queued slots; requeue a transient rejection; resume event and script delivery; reconcile active and terminal-unacknowledged records; discard a server-confirmed terminal client record; keep multiple clients fair.HP-CLIENT-01..14— Create and retain one durable client-instance ID; start with an empty spool; durably admit and acknowledge commands; advertise queue/ running capacity; dequeue fairly; assign and replay event sequences; compact acknowledged events; retain terminal-unacknowledged records; discard them only after reconciliation/acknowledgement; enforce local command/aggregate quotas; restart with clean recovery; reconnect with saved backoff/session inputs; stop all supervised Jobs on orderly service shutdown.HP-CMD-01..18— Executecmdand PowerShell command text and uploaded scripts; use foreground/background modes; apply requested/omitted CWD; materialize identity-scoped default CWD; overlay an empty/small/large legal environment; handle empty/multiline text, shell metacharacters, PowerShell Unicode, and space/non-ASCII executable/CWD/wrapper paths; return exit codes 0 and nonzero; record every public lifecycle; run 1 and configured maximum concurrent commands; queue then start work; execute short and descendant- producing commands; collect resource snapshots; apply every supported individual/composed resource profile before release.HP-WINCTX-01..14— With one active standard user run normal asACTIVE_USER; with a split-token administrator run normal filtered and elevated through the linked full token; accept an already-full admin token; build/verify a restricted medium token when UAC is off; run elevated asACTIVE_SYSTEMafter linked-token unavailability; useLOCAL_SYSTEMas final elevated fallback; with no active session run normal asLOCAL_SERVICEand elevated asLOCAL_SYSTEM; persist attempts/effective SIDs/session; omit session fields for Session 0; load/unload active profile; build the effective environment; access ACL-permitted drive, localhost, and network resources.HP-LAUNCH-01..16— Authenticate the per-command launcher pipe; create the launcher and shell suspended; assign the Job before release; persist/flush prepared then authorized state; resume exactly once; capture both streams; write and close stdin; keep descendants in the Job; observe root and complete- tree exit; drain pipes to EOF; deliver CTRL_BREAK through the verified helper; escalate TERM after grace; terminate immediately on KILL; enforce Job limits; remove wrappers/scripts after terminal acknowledgement; preserve effective identity on later lifecycle events.HP-OUT-01..14— Preserve binary/invalid-UTF-8 bytes, empty writes, partial lines, no-newline output, long lines, interleaved stdout/stderr, exact chunk boundaries, cumulative event acknowledgements, reconnect replay, byte-offset pagination within an event, live follow after historical output, normal drain, explicit incomplete marker, and normal client/server compression reuse.HP-STORE-01..20— Initialize/migrate SQLite; append inline and segmented payloads; group commit; reopen cleanly; query stable cursors; charge command/ client/server bytes; rotate output window; evict whole oldest terminal UUID; protect active commands; reserve closeout; reclaim terminal age; create and cap tombstones; rotate audit by bytes/age; maintain incident history; run safe repair; acknowledge loss; derive/clear dirty state; checkpoint WAL; backup and restore a quiesced store.HP-FLOW-01..12— Send heartbeat while idle; reset activity on any valid frame/Pong; reset reconnect backoff after stability; keep Ping/Pong/Close and acknowledgements ahead of saturated data; apply fair writer/persistence scheduling; cross high then low watermarks; retain assigned events; summarize unsequenced client loss; use closeout reserve; enforce per-command/session send windows; reconnect with deterministic full jitter.HP-SVC-01..16— Idempotently install/start/stop/restart/uninstall the SCM service; boot before login; preserve data on uninstall; change Automatic/ Manual in SCM; start one tray per session from the Run key; reconnect tray after Explorer restart; exit tray without stopping service; read status/log; perform an elevated configuration action; rotate service logs; operate without tray; run Session 0 on Server Core; attach console for human CLI modes; display native error/dialog when no parent console exists.HP-OPS-01..12— Expose liveness during recovery and readiness afterward; keep healthy scopes usable; emit bounded-cardinality metrics; rotate/redact logs and audits; report storage and cache usage; perform orderly shutdown; resume after service/server restart; validate deployment examples; upgrade with config check/snapshot/migration; restore prior version/config on an intentionally failed pre-migration check; generate checksums/package metadata.HP-HARNESS-01..14— Doctor a cold and cached environment; create/list a run; execute one case/scenario; collect JSON/JUnit/report; stop and resume; recover an interrupted controller; reuse immutable definitions into fresh state; reset; purge one run; dry-run and execute age GC; purge all confirmed RVBox runs; retain caches by default; explicitly clear caches/images; preserve file ownership and unrelated labelled fixtures.
Bad, hostile, and abnormal-recovery cases
BH-CFG-01..24— Reject missing/unknown/duplicate keys, wrong TOML types, invalid UTF-8, bare/overflow/negative durations, illegal zeros, limit and watermark inversions, incompatible profiles, relative paths, missing/wrong- type shell files, changed shell identity before launch, disallowed CWD, insecure state/config/work ownership or ACL, symlink/reparse-point escape, reserved device/ADS/NT-object path misuse, disallowed or inaccessible UNC paths, bad URL/TLS name, listener collision, malformed Windows quoting, and non-loopback JSON-RPC without explicit enable; prove no listener, store mutation, or child starts after static failure.BH-PROTO-01..28— Reject text WebSocket messages, invalid fragmentation, empty/unknown payload, incompatible major/minor, absent/wrong fencing fields, oversize envelope/spec/control/HTTP body/chunk/script, protobuf recursion/ length abuse, integer overflow, invalid enum/oneof combinations, malformed UUID/client ID, invalid declared compressed sizes, truncated/corrupt Zstandard, compression bombs, impossible event sequence/revision, changed duplicate payload, out-of-order/overlap/hole script chunks, post-commit chunks, SHA/size mismatch, invalid stdin sequence, and unsupported signal; close/fence only the offending peer and keep daemons alive.BH-TLSNET-01..18— Reject wrong CA/name/expired/not-yet-valid certificate, plaintext endpoint, wrong nginx path/upgrade, and unreachable proxy; recover from DNS refusal, TCP reset, TLS failure, proxy restart, half-open connection, packet loss/duplication/reordering, extreme latency, slow reads/writes, permanently blocked writer, write deadline, missing Pong, clock jump, and reconnect storm without goroutine/file-descriptor growth or data loss.BH-CTL-01..24— Reject malformed/empty/batch JSON-RPC where unsupported, unknown method, invalid params/base64/timestamp, oversized body, tampered or filter-mismatched cursor, page-size overflow, unknown client/command/incident, illegal lifecycle action, stdin after close/terminal, signal before eligible state, takeover for wrong/stale instance, repair of irreparable evidence, acknowledgement without/beyond note bound, reused request ID with changed method/target/content, read-only request ID leakage into domain, and client disconnect during follow; never cancel work merely because the caller exits.BH-SES-01..24— Reject stale/unknown generation, stale output after fencing, live different-instance collision without exact grant, mismatched/expired/ consumed grant, capacity lies/overflow, dispatch to wrong generation, permanent client rejection, reconnect snapshot missing accepted/running state, matching tombstone against stale server state, hash/revision/terminal contradiction, incomplete reconciliation snapshot, client dirty-store claim, lost/repeated result, offline target, queue capacity exhaustion, and one noisy/ corrupt client; isolate the scope and keep unrelated clients/control responsive.BH-CLIENT-01..28— Handle missing/corrupt/permission-denied client-instance identity, duplicate daemon lock, spool schema/checksum/segment corruption, uncommitted spool tail, missing acknowledged bytes, stale accepted record, incomplete prepared/authorized process metadata, unknown live process, local queue full, per-command/aggregate spool full, client disk floor, script temp orphan, tombstone overflow, event-sequence counter drift, impossible server Ack, dirty local storage, server rejecting registration/version, repeated connect refusal, shutdown deadline, service kill, and config/shell change across restart. Do not claim a complete reconciliation snapshot or accept work until essential local state is trustworthy or explicitly resolved.BH-IDEM-01..16— Reject noncanonical/non-v7 IDs, all-zero/wrong-length binary forms, same UUID with changed immutable execution spec/script bytes, same request ID across mutation kinds/targets, duplicate sequence with changed bytes, and replay older than full history but still in tombstones; define and test best-effort behavior after tombstone horizon expiry; never create a second process or mutation result on any retained duplicate.BH-SCRIPT-01..18— Treat filename as display only; reject path separators, traversal, reserved names, NUL/control abuse, excessive metadata, changed replay bytes, sparse/overlapping writes, premature commit, digest mismatch, quota exhaustion, disk-full, permission/share violation, wrapper replacement, symlink/reparse race, shell executable replacement, and cancellation during upload; clean only owned temporary files and preserve durable rejection detail.BH-WINCTX-01..26— Handle no linked token, standard user elevation request, Administrator Protection/approval-only policy, disabled/removed account, locked/disconnected/logging-off session,WTSQueryUserTokenfailure, multiple ambiguous active sessions, console-session change, session-ID reuse with a new logon SID, token SID/session/integrity/type mismatch, restricted-token creation failure, profile load/unload failure, environment block failure, LocalService logon failure,TokenSessionIdfailure, active-SYSTEM failure, final LocalSystem failure, explicit CWD denial, identity-work-root tampering, absent mapped drive/ HKCU/network credentials, and Session 0 interactive API failure. Follow only the allowed pre-launch chain and emit a bounded decision record with no effective context if exhausted.BH-LAUNCH-01..34— Reject spoofed/remote/early launcher or signal-helper pipe clients, wrong PID/creation time/SID/session/generation/checksum/length, inherited unrelated handles, launcher outside Job, breakaway attempt, shell/ wrapper/CWD/profile validation failure, invalid or case-colliding Windows environment keys, oversized environment block, unencodable wrapper content, process creation/assignment/resume failure, helper attach/delivery refusal, PID reuse, root disappearance, child holding pipes forever, descendant escape attempt, output read/write error, stdin broken pipe, Job accounting failure, unsupported Job limit, service death before/after preparation/authorization, launcher death at every handshake, and incomplete EOF. Before authorization reject/cancel safely; after durable authorization terminate/interrupt and never redispatch.BH-OUTFLOW-01..24— Survive output flood from one/many commands, tiny writes, incompressible data, invalid bytes, compressor error, client raw backlog, offline spool full, pinned send window full, server ingress full, command/ client/global quota full, physical disk floor, loss-marker reservation pressure, missing EOF, corrupt spool record, repeated Ack, Ack beyond sent sequence, server retention removing resume point, and slow follower. Preserve lifecycle/control progress and expose every loss/incomplete range honestly.BH-STORE-01..36— Recover uncommitted tails and marked eviction; detect short committed files, checksum/frame corruption, missing segment, wrong file type, path replacement, SQLite corruption/lock/busy/readonly, WAL/shm loss, schema or migration checksum mismatch, partial migration, fsync/rename/directory-sync failure, ENOSPC at every append/closeout/audit phase, counter/reservation drift, orphan temp/tombstone files, interrupted audit/age rotation, tombstone cap, filesystem free-space breach, permission change, backup inconsistency, and restore mismatch. Gate only affected scopes, expose incidents promptly, and require acknowledgement before clearing known loss.BH-RET-01..20— Reject new command when its own cap cannot fit; reject one client/global admission when active/non-evictable data alone fills the tier; evict terminal commands with no output; charge scripts/stdin/metadata/history/ idempotency data; never partly evict a command; preserve active commands, tombstones, audit, unresolved incidents, assigned client events, and closeout reserve; handle no eligible victim; retain truncation metadata; apply 30-day and independent audit age/byte rotation exactly once.BH-SVC-01..28— Reject non-admin install/config/uninstall, untrusted binary/ config path, direct internal service/launcher/helper mode, wrong SCM launch, duplicate/corrupt service registration, insecure ProgramData ACL, malicious tray pipe client, spoofed admin claim, cross-session tray action, oversized/ stalled IPC, second tray mutex, config/log path substitution, Run-key quoting attack, startup config invalidity, network unavailable at boot, tray absent/ crashed, Explorer absent, service start timeout, forced stop with active Jobs, uninstall interruption, log disk-full/rotation sharing violation, and Server Core without GUI. Keep service live/not-ready where designed and never let tray failure own daemon state.BH-SEC-01..22— Exercise command/script path traversal, SQL metacharacters, log/terminal escape injection, Unicode confusables in IDs, environment-value redaction, secret-looking payloads in metrics/logs/audits, malicious PATH/ COMSPEC/PATHEXT/file association, unsafe executable replacement, named-pipe ACL bypass, remote pipe access, inherited-handle leakage, decompression allocation abuse, cursor forgery, JSON-RPC bind warning, unauthenticated JSON-RPC authority as the explicit v1 behavior, and least-access state/work/log ACLs. Tests assert the documented insecure boundary rather than pretending authentication exists.BH-OPS-01..20— Handle asynchronous recovery timeout, one dirty scope, metrics scrape during churn, log sink failure, full audit budget, invalid health path, shutdown deadline, kill during shutdown, restart loop, incompatible/failed upgrade, downgrade attempt, backup destination full, restore with wrong config, clock rollback/forward/DST, hostname/client-ID change, and missing OS diagnostic data. Liveness stays meaningful, readiness stays conservative, and absence is never fabricated as zero/healthy data.BH-HARNESS-01..26— Reject invalid/traversal run ID, corrupt/truncated/foreign manifest, commit/image/config/VM mismatch, live lock stealing, stale lock with live owner, unlabelled or wrong-repository Docker resource, symlinked run root, path outside.test-runs, Windows resource outside test root, secret in report, oversized artifact, no disk, unavailable image/runner, port collision, lost runner, interrupted setup/collect/cleanup, failed Docker/SCM removal, unsafe resume, reuse attempting mutable-state copy, GC of resumable run, purge without confirmation, and--allwith unrelated resources. Fail closed and print an exact recovery/cleanup instruction without invoking global prune.
Obscure race, crash-window, and interleaving cases
All race cases run under deterministic schedule hooks first, then selected cases run repeatedly with the Go race detector or native process concurrency. A case must assert durable state, process count/identity, quota counters, emitted event order, and leaked goroutine/handle/file counts—not only the CLI exit status.
RC-DOM-01..18— Concurrent identical/conflicting command creation; mutation retry while original commits; cancel versus queue dispatch, acceptance, preparation, authorization, running, and terminal commit; signal versus cancel/ terminal; stdin write versus close/terminal; late resource snapshot versus terminal; duplicate terminal events; command revision increments from multiple controllers; UUIDv7 generation in one millisecond and across clock rollback.RC-SES-01..20— Same-instance reconnects cross; old read/write loops race fencing; different-instance reconnect races grant creation/expiry/consumption; session closes while dispatch reserves capacity; capacity update crosses acceptance/rejection; reconnect crosses queue expiry; stale output arrives before/after new generation commit; reconciliation result is lost while new dispatch wakes; two clients contend for global fairness; heartbeat timeout crosses Pong/read activity/write deadline/server shutdown.RC-CLIENT-01..22— Duplicate dispatch crosses durable local admission; acceptance acknowledgement crosses service crash; queue dequeue crosses cancel/ shutdown; output/event sequence assignment crosses Ack/reconnect/compaction; terminal acknowledgement crosses local cleanup/tombstone insertion; server discard result crosses retransmission; local quota reservation crosses script/ stdin/output writes; reconnect crosses config restart; identity/spool recovery crosses connection startup; two commands become runnable as capacity changes; stop-all crosses a new prepared Job. Preserve at-most-once execution, monotonic event order, and accurate advertised capacity.RC-STORE-01..28— Event append races Ack/query/retention; terminal commit races output drain and whole-command eviction; eviction races pagination, follow, idempotency retry, and age rotation; tombstone insertion races full- record deletion/replay; audit rotation races incident resolution; repair races startup recovery/retention/backup; command/client/global reservations cross exact caps concurrently; closeout consumes its reserve while disk floor trips; group commit crosses process kill; WAL checkpoint crosses readers/shutdown; two failures create incidents for one scope without reopening resolved history.RC-OUT-01..20— stdout/stderr readers race root exit/descendant exit; child inherits pipe while grace expires; raw backlog crosses high/low repeatedly; compressor completion reorders across commands but not within one command; Ack arrives during reconnect/spool compaction; output assignment races client loss conversion; server retention marker races live follow; terminal event races final output/incomplete marker; slow follower disconnects while page read crosses segment eviction.RC-SCRIPT-01..12— Duplicate chunks arrive concurrently; commit crosses last fsync, cancellation, quota eviction pressure, client restart, digest failure, and launch eligibility; cleanup crosses terminal acknowledgement/service death; wrapper identity is swapped between validation and creation/revalidation; identical upload resumes from competing old/new sessions.RC-WINCTX-01..18— User logs on/off, locks/unlocks, disconnects/reconnects, or changes console session during enumeration/token construction/revalidation; session ID is reused with a different logon SID; multiple RDP sessions become active/ambiguous; linked-token policy changes; profile unload races Job exit; active-user CWD disappears or ACL changes before release. Re-enumerate at most once as specified, never retarget silently, and never fall back afterlaunch_prepared.RC-LAUNCH-01..30— Cancellation/service stop/client crash at every instruction boundary among durable acceptance, token selection, Job/pipe creation, suspended launcher creation, Job assignment, pipe authentication, suspended shell report, prepared fsync, authorized fsync, release send, resume, running acknowledgement, root exit, drain, terminal fsync, EventAck, cleanup, and tombstone insertion. Duplicate dispatch after each restart must yield zero or one process, never two.RC-SIGNAL-01..16— TERM/KILL/cancel collide with launcher connection, shell resume, root exit, descendant-only Job, helper pipe authentication,AttachConsole, CTRL_BREAK delivery, grace timeout, Job termination, service shutdown, PID reuse, and repeated signal revision. Record the actual attempted/ escalated result and never signal an unrelated PID/console.RC-SVC-01..20— SCM start/stop/restart/uninstall cross recovery, command admission, prepared launch, and active Jobs; two installers/configurators race; tray startup crosses logon/logoff/Explorer restart/service restart; multiple trays contend for the per-session mutex; pipe disconnect crosses a mutating helper result; log rotation crosses tray open/read; config edit crosses explicit restart; Automatic/Manual change crosses SCM query.RC-OPS-01..14— Health/metrics/read RPCs cross recovery state transitions, dirty resolution, retention, and shutdown; log/audit rotation crosses process crash; backup crosses group commit and dispatch pause; upgrade crosses queued/ active commands and rollback; wall-clock changes cross queue TTL/retention while monotonic heartbeat remains correct.RC-HARNESS-01..18— Controller dies before/after manifest fsync, resource creation/label recording, Windows lease, fault checkpoint, result write, artifact collection, stop/reset/purge, and lock release; two cleanup commands race; status/collect runs during cleanup; resume crosses a late old controller; GC crosses a newly resumed run. Recovery must converge idempotently and never delete an unrelated resource.
For each crash-window family, implement a loop that enumerates named hooks rather than hand-selecting a few attractive points. Persist the hook name before triggering death, restart from the real store/spool, replay the same external request, and check the invariant matrix. A newly added durable write or external side effect must add a hook and coverage row before merging.
2.6 Native Windows test-host requirements
Native Windows is mandatory for the Phase 4/5 gates. Development can begin with unit tests and cross-compilation before a host is connected, but the Windows supervisor/service/tray implementation cannot be called complete without it. Prefer disposable or snapshot-resettable VMs over a personal workstation.
Provision at least one primary interactive host with:
- current Windows 11 Pro/Enterprise, Desktop Experience, Explorer,
cmd.exe, Windows PowerShell 5.1, UAC enabled, and all current updates captured in a named clean snapshot; - at least 4 vCPU, 8 GiB RAM, and 60 GiB free disk recommended (2 vCPU, 4 GiB, and 30 GiB free is the minimum smoke lane), with reboot and snapshot-revert authority;
- one local standard account and one traditional split-token local administrator, plus a separate automation principal able to install/control the test service;
- a preconfigured CI runner, OpenSSH, or WinRM management channel reachable only from the test controller; credentials live outside the repository and run manifest;
- bidirectional reachability to the Linux nginx test endpoint, stable DNS or a supplied address, time synchronization, and permission to transfer the CI-built binary/config/CA into a dedicated test root;
- no valuable user data, credentials, mapped drives, or production services, because tests deliberately change SCM/Run-key state, ACLs, UAC fixtures, logon sessions, kill processes, fill bounded scratch storage, reboot, and revert.
Maintain snapshot/policy variants for logged-out, standard-user active,
split-token-admin active, UAC disabled/already-full admin, and Administrator
Protection when supported. The harness must restore the baseline after any
variant that changes machine policy. Native multi-session/ambiguous-session
validation requires a later Windows Server Desktop Experience or RDS fixture
with multiple simultaneous WTSActive sessions. It is explicitly deferred for
the current Windows-first v1 milestone; keep its exhaustive selector unit test
mandatory and mark only that native case blocked. Add a Server Core VM for the
headless Session 0 release smoke. Before publishing, also exercise the oldest
supported Windows 10 or Server 2016 baseline; it need not be the everyday runner.
A physical Windows machine is optional. It is useful for an additional real
display/audio/DDC command smoke, but RVBox only guarantees correct token/session/
process execution—not success of arbitrary vendor hardware APIs—so physical
hardware is not a release blocker. test-host.ps1 Prepare must inventory the
host against this checklist and refuse destructive suites unless the machine is
explicitly marked disposable/resettable and the clean snapshot identity is
recorded.
2.6.1 Provisioned headless VirtualBox smoke lane
A resettable Windows lane is already provisioned on the SSH host alias
helium-remote. It is the first native execution target and is suitable for
unit-adjacent native checks, supervisor/service/tray smoke, and protocol E2E
work. It is deliberately a minimum smoke lane, not a replacement for the
larger Windows release matrix above: it has one active user, no standard-user
fixture, no ambiguous multi-session fixture, and no Server Core variant.
| Item | Value |
|---|---|
| Hypervisor | VirtualBox 7.2.16r174877 on the Arch Linux helium-remote host |
| VM name / UUID | rvbox-win10-test / 6cdc114f-71e5-4167-a394-e922e14e6f5c |
| VM group / config | /RVBox/Tests; /home/cabbage/VirtualBox VMs/RVBox/Tests/rvbox-win10-test/rvbox-win10-test.vbox |
| Guest OS | Windows 10 Pro 22H2, build 19045.2006, en-US, BIOS boot |
| Resources | 2 vCPU, 4096 MiB RAM, 32 MiB VRAM, 40 GiB dynamically allocated VDI |
| Disk / source media | /home/cabbage/VMs/rvbox-win10-test.vdi; source ISO /media/Data2/Downloaded/Win10_22H2_English_x64.iso (Windows image index 6) |
| Devices/network | NAT NIC; audio, USB, 3D, clipboard, drag-and-drop, and VRDE disabled; last observed guest IPv4 was 10.0.2.15 (DHCP observation only, never a management identity) |
| Guest control | Guest Additions installed and verified (GuestAdditionsRunLevel=3) |
| Test account | local rvboxtest; one active console session (session 1); split-token local administrator; Guest Control was verified with whoami, whoami /groups, and query user |
| Baseline | snapshot baseline-disk-first (UUID 9430a9a4-754a-4b22-beaa-8dfd90043f5b), current known-good snapshot; parent baseline-clean is retained |
| Initial state | VM is normally left powered off; restore the baseline before each destructive run |
The guest password, SSH key, and any host account secret are test secrets. Keep
them in the operator/CI secret store or a mode-600 password file outside the
repository; never put them in this plan, a command-line argument, a run
manifest, or collected logs. VBoxManage guestcontrol supports
--passwordfile; prefer that option over an inline password. Every operator or
agent must set RVBOX_TEST_GUEST_PASSWORD_FILE to the absolute path of that
host-side file before a native run. The account name and VM metadata above are
not credentials.
For this provisioned lane, the approved host-only credential-file location is
/home/cabbage/.local/share/rvbox-secrets/rvbox-win10-test.password. It must
remain mode 0600, is never read into a repository process, and is supplied to
Guest Control only with --passwordfile. The current VirtualBox 7.2.16
Guest Control build supports --wait-stdout and --wait-stderr, but not
--wait-exit; adapters must not pass that unsupported flag.
For a guest cmd.exe command payload, this Guest Control build requires
--unquoted-args so the controller preserves Windows backslashes and the
single /c command payload. Treat each Guest Control process as a bounded,
owned resource: record its guest session/PID, close it after completion, and
use closeprocess before reset if an output wait does not complete. Do not use
this compatibility rule to assemble untrusted command text; it is solely the
adapter transport for prevalidated test commands and exact executable paths.
Use a local, non-secret environment description when operating the lane. The host alias must resolve through the operator's SSH config; another controller may substitute an equivalent target, but must record the resulting host/VM identity in the run manifest:
export RVBOX_TEST_VBOX_HOST=helium-remote
export RVBOX_TEST_VBOX_VM=rvbox-win10-test
export RVBOX_TEST_VBOX_SNAPSHOT=baseline-disk-first
export RVBOX_TEST_GUEST_USER=rvboxtest
export RVBOX_TEST_GUEST_PASSWORD_FILE=/secure/outside-repo/rvbox-win10-test.password
All host lifecycle commands run on Helium through SSH; all guest commands run
through VirtualBox Guest Control. Do not assume that VBoxManage is installed
on the Linux controller or that WinRM/OpenSSH is enabled inside this fixture.
The future Windows harness may use a different management backend, but this
lane's backend must validate fixed VM names/UUIDs and use argument vectors (not
shell-concatenated user input):
# Read-only identity/state check.
ssh "$RVBOX_TEST_VBOX_HOST" \
'VBoxManage showvminfo "rvbox-win10-test" --machinereadable'
# Restore only while powered off, then boot without a GUI.
ssh "$RVBOX_TEST_VBOX_HOST" \
'VBoxManage snapshot "rvbox-win10-test" restore "baseline-disk-first"'
ssh "$RVBOX_TEST_VBOX_HOST" \
'VBoxManage startvm "rvbox-win10-test" --type headless'
# Poll these properties before a test; GuestAdditionsRunLevel=3 is required
# for Guest Control, and LoggedInUsers/NoLoggedInUsers selects the session case.
ssh "$RVBOX_TEST_VBOX_HOST" \
'VBoxManage guestproperty enumerate "rvbox-win10-test"'
The harness must wait for the VM to report running, then poll Guest
Properties until Guest Additions is ready and the requested login fixture is
observed. NAT address 10.0.2.15 was observed during provisioning but is DHCP
state, not an identity or a stable endpoint; use Guest Control for management
and discover any test networking separately. A failed readiness poll is a
stopped-resumable run, not permission to start a second VM with the same name.
For a guest command, create/use the secret file on the host where
VBoxManage executes and pass explicit executable/argument vectors. For
example, the following verifies the effective account without opening a shell
through PATH:
ssh "$RVBOX_TEST_VBOX_HOST" \
'VBoxManage guestcontrol "rvbox-win10-test" run \
--username "rvboxtest" \
--passwordfile "/secure/outside-repo/rvbox-win10-test.password" \
--exe "C:\\Windows\\System32\\whoami.exe" \
--wait-stdout --wait-stderr --'
Copy CI-built binaries/configuration to a host staging directory with scp,
then use guestcontrol ... copyto into a run-specific guest test root. Collect
only bounded, redacted artifacts with guestcontrol ... copyfrom before reset.
Do not use shared folders or expose the service store to the tray; this VM has
no shared-folder contract.
Use this shutdown/reset sequence for every native run:
- Stop admission and ask the guest/service to shut down cleanly through Guest
Control; wait for
VMState=poweroff. - If Guest Control is unavailable, send
VBoxManage controlvm ... acpipowerbuttonand poll. Usecontrolvm ... poweroffonly for a hung, disposable test; it intentionally loses guest state. - Collect diagnostics while the VM is still available, then restore
baseline-disk-firstand verify the snapshot UUID/current marker. - Leave the VM powered off after cleanup. Never delete either baseline
snapshot, unregister the VM, or modify
win10_dev(that name refers to a stale unregistered configuration with a missing disk on this host).
For headless diagnosis, capture a bounded screenshot on Helium and copy it to the controller:
ssh "$RVBOX_TEST_VBOX_HOST" \
'VBoxManage controlvm "rvbox-win10-test" screenshotpng /tmp/rvbox-win10-test.png'
scp "$RVBOX_TEST_VBOX_HOST:/tmp/rvbox-win10-test.png" ./artifacts/
Keyboard scancodes are a last-resort recovery aid for BIOS prompts, UAC, or
logon when Guest Control cannot reach the guest; they are not a test-control
API. Prefer deterministic in-guest commands and restore the snapshot after any
manual interaction. When reprovisioning is unavoidable, use the same ISO,
Windows image index 6, 2-vCPU/4-GiB/40-GiB profile, NAT/no-audio/no-USB device
profile, Guest Additions, and a newly generated test password. VirtualBox
unattended installation may attach both the original ISO and an auxiliary VISO;
ensure the original media is bootable (or inject the BIOS key), detach install
media after setup, set boot1=disk, and take a fresh named baseline only after
Guest Additions and Guest Control have been verified.
This fixture is now available for native testing, so remove any generic “Windows host unavailable” skip only when the harness can acquire this VM's exclusive lease and perform the reset/health checks above. Keep the deferred native multi-session/ambiguous-session, Server Core, and older-build cases explicitly represented as separate blocked matrix entries until their own resettable fixtures exist.
3. Repository and build bootstrap (Phase 0)
3.1 Establish the repository layout
Create the following layout. Generated code is kept separate from handwritten logic and is never manually edited.
cmd/
rvbox-server/main.go
rvbox/main.go
rvc/main.go
internal/
agentproto/ # envelope validation, transport-neutral helpers
config/ # file/flag parsing and cross-field validation
domain/ # command state machine and typed errors
server/
control/ # gRPC implementation and JSON-RPC adapter
session/ # client registry, fencing, dispatch
store/ # SQLite metadata, segments, audit, migrations
client/
runtime/ # reconnect loop, transport, dispatcher
spool/ # active command/event/output durable spool
supervisor/ # OS-neutral interface
supervisor/windows/# first target: tokens/sessions, launcher, Jobs, diagnostics
supervisor/unix/ # later target: process groups, /proc, cgroup v2
windowsservice/ # SCM install/control and service host
windowstray/ # per-session named-pipe frontend and notification icon
observability/ # logging, metrics, health/readiness
testkit/ # clocks, fake transport, fault helpers
gen/go/rvbox/v1/ # generated protobuf/grpc code
protos/rvbox/v1/ # existing wire authority
docs/examples/ # fully annotated server/client TOML
scripts/
test-unit # hermetic toolchain-container entry point
test-integration # resumable component-suite entry point
test-e2e # resumable production-shaped scenario entry point
test-env # doctor/recover/reuse/cleanup by exact run ID
windows/test-host.ps1 # native Windows host lifecycle adapter
test/
coverage.toml # requirement-to-case inventory with stable IDs
harness/ # shared manifest, journal, orchestration, and reporting
defaults.toml # local resource/time/artifact budgets
integration/ # cross-package/real-resource suite definitions
e2e/scenarios/ # named black-box scenario definitions
fixtures/ # small deterministic non-secret fixtures
deploy/
Dockerfile.toolchain
Dockerfile.runtime
compose.yaml
compose.test.yaml # isolated integration/E2E/fault-injection profiles
nginx/
systemd/
Use a single Go module rooted at the repository. Select and pin exact Go module
versions in go.mod; record reasons for non-standard dependencies in
docs/dependency-decisions.md. Prefer a pure-Go SQLite driver to avoid a C
toolchain for normal Linux/Windows client builds. Use a maintained,
context-aware WebSocket implementation and a Zstandard implementation that
supports bounded decompression. Do not rely on an archived WebSocket package.
For Windows, prefer golang.org/x/sys/windows plus narrow reviewed wrappers for
SCM, tokens/WTS, Job, console, named-pipe, shell, ACL, and notification-area APIs. Do not
adopt an unmaintained tray abstraction merely to reduce Win32 code. Pin the
resource compiler used for icon/version manifests and keep signing credentials
out of build images and repository state.
3.2 Containerized developer tooling
Add:
deploy/Dockerfile.toolchain: pinned Go base image plus Buf,protoc,protoc-gen-go,protoc-gen-go-grpc, format/lint tools, andmake.deploy/Dockerfile.runtime: minimal non-root image for server/client smoke tests; it contains only built binaries and runtime CA/config assets.deploy/compose.yaml: atoolchainservice with the repository mounted at/workspace, run as1001:1001; optionalserver,client,nginx, and fault-injection test services use isolated named volumes.deploy/compose.test.yaml: test-only profile overrides, health checks, resource limits, run-ID labels, ephemeral PKI, fault proxy, and artifact/state mounts. It accepts only values produced by the validated harness manifest and never uses an implicit default project name.Makefile:generate,fmt,lint,test,test-race,test-coverage,test-integration,test-e2e,test-status,build, andverify.testaliases the unit layer; integration/E2E targets delegate to the checked-in scripts and print their run ID. Each target invokes the container workflow and must not silently fall back to host tools.buf.gen.yaml: Go protobuf and gRPC generation paths undergen/go.
Ignore .test-runs/ and any local harness lock file in Git while retaining a
tracked .test-runs/README only if an in-tree explanation is useful. The
harness creates the directory with host UID/GID 1001:1001; test binaries never
write generated dependencies into the repository tree.
Generation is deterministic: make generate followed by git diff --exit-code
must be clean in CI. Update buf.yaml/buf.gen.yaml only with a matching
generation run.
3.3 Initial quality gates
Add CI (or a repository script ready for CI) that runs, in order:
buf format --diffandbuf lint. Snapshot only the already accepted enum-prefix/service-suffix findings so CI still fails if their set changes or another warning appears. This pre-implementation cleanup defines the first v1 compatibility baseline; enablebuf breakingagainst that baseline immediately after it merges, not against the obsolete draft.- Generation freshness check.
- Coverage-inventory validation: every required stable ID maps to a listed case and every normative critical invariant has at least one case.
go fmt,go vet, static analysis, and unit tests with per-package coverage.- Race tests for server/client concurrency packages.
- Linux server/storage/session integration tests in Compose; no Linux client supervisor is required at this stage.
- Cross-compilation checks for Windows packages in the container plus native Windows build/runtime smoke tests on a Windows runner. A Windows runner is a Phase 0 prerequisite, not a later optional enhancement.
Provide two native Windows lanes before Phase 4: a clean noninteractive runner for unit/Job/process/storage tests, and a resettable interactive VM with Explorer, a split-token administrator, UAC enabled, and Desktop Experience for tray/WTS/ token-session tests. Add repeatable test users/policy fixtures for a standard user, no logged-in user, multiple active sessions where supported, UAC disabled, and Administrator Protection enabled when the current Windows image exposes it. Tests that require the interactive desktop may be a release-candidate gate rather than per-commit, but cannot be replaced by mocks or a service-account runner. Add a Server Core service/Session-0 smoke lane. For the Windows-first v1 milestone, use the provisioned Windows 10 Pro 22H2 fixture as the compatibility baseline. Older Windows and Server validation is explicitly deferred to a later compatibility pass; its absence does not delay current implementation or native smoke coverage.
Use one declared test matrix and the same harness entry points everywhere:
- every change runs format/lint/generation, all unit tests, selected race tests,
Linux
store/server-session/controlintegration suites, Windows cross- compilation, and the native noninteractive Windows smoke integration suite; - scheduled/nightly runs execute the full race/fuzz budgets, all Linux and
native Windows integration suites, plus
smoke,idempotency,reconnect, andretentionE2E scenarios; - release candidates run every E2E scenario, including the resettable interactive Windows VM matrix, Server Core smoke, destructive fault cases, upgrade/recovery, and soak.
CI allocates a run ID per job, always invokes collect after failure, and invokes
reset in an unconditional finalizer. Upload only the bounded redacted report,
manifest, and relevant logs; never upload full command payload stores by
default. A failed cleanup is a visible job failure with its exact manual cleanup
command, not a warning hidden behind the original test result.
Exit criteria: all three empty main packages build in the toolchain
container; generated code is checked in; make verify works from a fresh clone;
the test scripts can create, interrupt, inspect, resume, collect, reset, and purge
one sample run without affecting a labelled unrelated fixture.
4. Shared contracts, configuration, and state machine (Phase 1)
4.1 Preserve proto authority
Generate common.proto, agent.proto, and control.proto before writing
transport code. Do not fork request/response structs by hand. These schemas are
an unshipped first draft, so remove obsolete fields and compact their tags now
rather than preserving compatibility holes. Once Phase 0 generated APIs merge,
freeze field meanings: later additions use new tags, enum zero remains
unspecified, and removed tags are then reserved normally.
At the application boundary, validate what protobuf cannot express:
client_idis configured hostname text, 1–128 ASCII characters.- command and mutation IDs are canonical RFC 9562 UUIDv7 strings; all public UUID parsing rejects empty, non-canonical, or wrong-version input.
- decoded envelopes are <= 1 MiB; output/stdin/script data validates every raw, compressed, and declared size before allocation or decompression.
- a serialized
ExecutionSpecis <= 768 KiB, decoded control gRPC requests are <= 16 MiB, JSON-RPC HTTP bodies are <= 24 MiB, and decoded field limits remain identical across the two control transports. - shell type is explicit or assigned only to the documented platform default.
elevated=falseis normal privilege; non-Windows clients reject true in v1. Windows applies the documented login-aware context hierarchy and records every attempted/effective context. Fallback ends beforelaunch_prepared.- active Windows contexts bind one revalidated session ID, logon SID, and user SID; process creation is never retried under a different context.
- CWD exists, is a directory, and is allowed by local daemon policy.
- an execution request has exactly one source: command text or script descriptor.
- script descriptors and payloads agree on SHA-256 and <= 10 MiB size.
- user-supplied script filenames are display-only basenames and never path components; command text and scripts use generated private wrapper paths.
- profile flags are deduplicated and checked for configured supported limits;
LIGHTis exclusive, otherwise accept at most one CPU, one memory, and one disk tier.
4.2 Domain package
Implement a pure-Go internal/domain package with no database/network imports.
It contains:
- the allowed command transition table and terminal-state predicate;
CommandRevisioncompare-and-transition helpers for launch/cancel/signal races;- typed domain errors mapped to gRPC status/structured details, JSON-RPC error
objects, or agent-protocol
ControlErroras appropriate; - typed Windows pre-launch errors for a failed normal execution context or an exhausted elevated-context chain; ambiguity and per-attempt token/policy failures remain bounded selection detail rather than separate terminal codes;
- event-sequence validation, duplicate equivalence checks, and declared gap
validation through
OutputTruncationmetadata; - default values and hard-limit validation;
- stable ordering/pagination token encoders.
Model launch_prepared and launch_authorized as durable internal execution
phases without exposing them as additional public lifecycle enum values. No
recovery path may redispatch a UUID after launch_authorized; an uncertain
outcome becomes interrupted.
The normal legal state transitions are:
queued -> dispatched -> accepted -> running -> succeeded|failed|terminated|interrupted
queued -> expired
dispatched -> rejected (permanent pre-acceptance failure)
dispatched -> queued (transient rejection with backoff)
accepted -> rejected (permanent upload/preparation failure before launch)
queued|dispatched|accepted -> cancelled (subject to revisioned race resolution)
dispatched|accepted -> interrupted (reconciliation/recovery invariant failure)
An attempted repeat event is harmless only if it has identical immutable content. A conflicting duplicate UUID/sequence is a protocol error and fences the bad session rather than rewriting history.
4.3 Configuration model
Use strict TOML 1.0 and implement the normative contract in
configuration.md. Keep the annotated
examples/server.toml and
examples/client.toml synchronized with the Go config
structs and compiled defaults. Keep
examples/client.windows.toml as the runnable
Windows-first subset using those same defaults.
Implement configuration in this order:
- Define separate decode-only TOML structs and immutable validated domain
structs. Every public TOML key has an explicit tag and documentation comment;
packages outside
internal/configreceive only a validated subsection. - Decode exactly one UTF-8 file with unknown-field and duplicate-key rejection. Do not merge include files, expand environment variables, coerce strings to numbers, or accept bare numeric durations.
- Apply precedence deterministically: compiled defaults, file values, then flags that were explicitly present. Generate flag bindings from one key registry so omitted flags cannot accidentally overwrite TOML values.
- Parse duration strings with overflow and non-negative checks. Parse byte/count integers without floating point. Enforce v1 hard ceilings even when TOML asks for more; operators may lower them but not create an incompatible wire peer.
- Canonicalize data/state/socket/CA/shell/CWD-root paths to absolute paths. Reject aliases that collapse distinct server data subdirectories and a control socket outside its intended runtime-directory policy. Check existing path ownership/type without claiming that private modes encrypt stored data.
- Resolve and validate current-platform shells at client startup. Each
nonempty applicable path must be absolute and identify an executable regular
file; other-platform strings are syntax-checked but never advertised.
Advertise only validated entries and require the current platform default to
be one of them. Cache the canonical path and never re-resolve through a
command's overridden
PATH. Revalidate identity before launch and reject rather than fall back if the executable changed incompatibly. - Validate resource profiles by dimension.
LIGHTis exclusive; otherwise allow at most oneCPU_*, oneMEM_*, and oneDISK_*. Each enabled profile must give nonzero values for itsrequired_controls; validate Linux device keys as canonicalmajor:minorvalues and Windows Job rate support before advertising the profile. - Enforce cross-field invariants: low watermark below high, send window below durable tier, closeout reserve below command quota, command quota below both aggregate tiers, reconnect initial <= cap, heartbeat idle < liveness timeout, positive page/queue limits, per-client server queue <= global server queue, and zero only where the schema documents disable or indefinite behavior.
- Print the normalized effective non-command configuration once after validation. Log a conspicuous warning for enabled non-loopback JSON-RPC, but retain the accepted behavior. Do not log TLS file contents or profile device discovery details that reveal unrelated host paths.
- Bind no command listener and launch no child until static configuration is valid. Storage recovery remains asynchronous after this static gate as specified in Phase 2.
Server TOML covers routing/listeners, queue/retry bounds, unified storage and audit limits, group-commit behavior, flow windows/watermarks, protocol ceilings, takeover TTL, and observability. Client TOML covers WSS/TLS routing, durable identity/state, exact shell paths, allowed CWD roots, queue/concurrency, reconnect/liveness, spool/flow limits, execution diagnostics/grace periods, resource profiles, and observability. Windows client configuration additionally covers rotating service-log output. SCM startup type, service/launcher/tray identities, execution-context selection, conventional paths, and secure ACLs are mandatory policy rather than TOML knobs.
Do not add or claim application-level at-rest encryption in v1. Command-owned payloads are stored in plaintext under the private state directories; document the equivalent sensitivity of backups and defer encryption/key management.
Hard defaults are: 10-second heartbeat-idle interval, 30-second liveness timeout, 1–60-second full-jitter reconnect, reset after 60 stable seconds, 16 running/100 queued client commands, 1,000 server-queued commands per client and 10,000 server-wide, 5-minute one-shot live-conflict takeover grant, 15-minute queue TTL, 64 KiB raw stream chunk, 1 MiB decoded agent envelope, 768 KiB serialized execution spec, 16 MiB decoded control request, 24 MiB JSON-RPC HTTP body, 10 MiB output window/raw script cap, 32 MiB total per command, 256 MiB per client/client daemon, 4 GiB server-wide command storage, 30-day terminal retention, 100 MiB audit storage with quota-only rotation by default (zero age retention), one million compact command tombstones, raw-output high/low watermarks of 1 MiB/256 KiB per command, 8 MiB/4 MiB per client, and 64 MiB/32 MiB server-wide, plus unacknowledged send windows of 1 MiB per command and 8 MiB per session. Reserve 64 KiB within every accepted command's quota for terminal/loss closeout metadata and bound protocol detail/reason plus incident-note text to 4 KiB. Use emergency filesystem free-space floors of 256 MiB server-side and 64 MiB client-side by default. The Windows service installs as Automatic; service logs rotate at 10 MiB and retain five sealed files.
Test all annotated example files through the production decoder on their target
platforms. Add table tests for every unknown key, invalid duration/size,
high/low inversion, tier inversion, zero semantic, noncanonical path/device,
unsupported default shell, conflicting profile combination, and flag-precedence
case. Add a test proving a request-level PATH override cannot change the
chosen shell executable. Golden-test the redacted effective configuration and
keep its key order stable enough for operators to compare deployments.
Exit criteria: state-machine and configuration tests cover all legal/illegal transitions and boundary values; all examples parse to the documented defaults; generated protos compile and are the only DTOs crossing process boundaries.
5. Server persistence and retention (Phase 2)
5.1 Storage ownership and directory safety
Implement internal/server/store as the sole writer for server state. On first
boot, create a private data directory (mode 0700), SQLite database, segment
directory, and audit directory. Validate that paths resolve below the configured
data directory; never construct a path directly from unvalidated client IDs or
filenames. Use encoded UUID/client-ID path components.
Open SQLite with WAL enabled, foreign keys enabled, a busy timeout, and an explicit durable-write policy. Run migrations transactionally. On startup:
- acquire a single-instance lock;
- bind liveness and diagnostic/control endpoints with readiness false;
- run integrity/schema checks and committed-range recovery asynchronously;
- safely truncate file bytes beyond SQLite's
committed_end_offset, while a missing/corrupt committed range creates a scoped durable incident; - mark prior live sessions disconnected and retain trustworthy non-terminal commands for reconciliation; interrupt only affected commands whose essential state cannot be trusted;
- enable each healthy scope independently. Mutations against an unrecovered or
dirty scope return
UNAVAILABLE; liveness, incident inspection, and unaffected scopes remain available throughout best-effort recovery.
Represent recovery with explicit atomic scope states: recovering, ready,
dirty_readable, and unavailable. Liveness depends only on the process/event
loop; global readiness is true only when every required scope is ready, while an
RPC checks the narrower scopes it will touch. The WebSocket endpoint may accept
and close with a retryable storage error during recovery, but it must not issue a
durable ServerWelcome, accept events, or dispatch commands until the relevant
client/session scope is ready.
Run PRAGMA quick_check, migration checksum validation, and segment reconciliation
on a bounded recovery pool. Persist a startup incident in the database when it
is writable; otherwise expose an in-memory bootstrap incident through health and
require the offline repair tool. Never loop forever on one corrupt client or
segment. Record progress counters and the last failing path/offset without
putting payload bytes in logs.
5.2 Metadata schema and indexes
Use migrations to create at least the following tables; names may vary but their invariants must not.
| Data | Required fields/invariants |
|---|---|
clients |
client ID, most-recent platform/capabilities/CWD/version, durable instance ID, current generation, connection/last-seen timestamps, latest rejected live-conflict instance/time, unified charged command-storage total |
sessions |
opaque session ID, client ID, durable client-instance UUID, generation, opened/fenced/closed times, close reason |
commands |
UUID, client ID, indexed issue/queue-expiry/terminal times, lifecycle, revision, exit result, last event sequence, retention status, a Zstandard-compressed immutable execution-spec payload including plaintext environment values, and optional Windows selection attempts/effective identity |
command_payloads |
command UUID, payload kind (script or other command-owned blob), Zstandard compression, raw/stored sizes, digest, inline bytes or validated segment reference |
command_events |
command UUID + event sequence unique key, observed and server receipt times, event type, payload metadata, immutable duplicate checksum |
output_segments |
command UUID, segment ordinal/path, committed_end_offset, min/max event sequence, stream mix, compressed/raw byte totals, checksum, created time |
output_truncations |
command UUID, removed event range/byte totals, source, reason, recorded time; non-overlapping ranges |
stdin_writes |
command UUID + write sequence unique key, compressed payload/raw/stored sizes/checksum, newline/close intent, acknowledged state |
control_mutations |
request UUIDv7 unique key, owner command/incident/client, method/target, immutable request hash, assigned write sequence/revision and stable result; conflicting reuse is rejected |
takeover_authorizations |
client ID + exact pending instance ID, creation/expiry/consumption times and authorizing request ID; one live, one-shot grant per client |
command_tombstones |
binary UUIDv7 unique key, immutable request hash, client ID, completion/acknowledgement time; compact global FIFO capped at 1,000,000 rows |
audit_events |
timestamp, source transport/principal where available, client/command IDs, action, outcome, error code, command text/script digest/env key names |
storage_incidents |
incident UUID, detected/resolved times, state, affected scope/client/command, summary, repairability and data-loss flags; unresolved rows derive dirty health |
schema_migrations |
applied version/checksum/time |
Index client command lists by (client_id, issue_time DESC), active command
dispatch by (client_id, lifecycle, issue_time), and output reads by
(command_uuid, event_seq). Use cursor tokens based on the sorted key plus a
signature/version, not SQL offsets, so page results remain stable under writes.
Add foreign keys with explicit delete behavior and database constraints for
terminal timestamps, nonnegative byte counters, unique (command,event_seq) and
(command,write_seq), one active session per client, and one unconsumed takeover
grant per client. Store UUIDs and SHA-256 values as fixed-size blobs internally;
canonical strings exist only at protocol/log boundaries. Canonical request hashes
use deterministic protobuf encoding of immutable fields after defaults and map
ordering are normalized; mutable lifecycle/receipt fields are excluded.
Every mutation transaction updates its owner record, all three applicable quota
counters, the idempotency row, and audit intent/result consistently. Identical
request_id plus hash returns the stored result without re-running side effects;
same ID with a different method, owner, or hash returns CONFLICT. Add migration
invariant queries that recompute charged totals and fail readiness if persisted
counters disagree until repaired.
5.3 Segment format and atomic append
Store large command, output, and audit payloads outside SQLite. An active segment is append-only; seal it at the configured 256 KiB target and never append again. Use an encoded UUID/ordinal filename created with exclusive create. Each record contains fixed magic/version/header length/record length, command and event identity, observed/receipt timestamps, payload kind/stream/compression, raw and stored lengths, SHA-256 payload digest, header CRC, payload, and trailing record CRC. All lengths are unsigned and validated against configured maxima before allocation or seeking.
The append path is:
- validate compressed bytes and bounded decompression;
- choose/append the active command segment;
- append and
fsync(possibly as a configured group commit); - in one SQLite transaction, insert event/segment metadata, advance that
segment's
committed_end_offset, update command sequence/counts, and record receipt time; - only after durable success, return the cumulative
EventAck.
The writer serializes appends per active segment but group-commits independent commands together. Creating/renaming a segment also syncs its containing directory before metadata can commit. A group-commit timer is a maximum batching delay, never permission to acknowledge before file sync. SQLite uses a documented synchronous mode sufficient for the selected durability promise; tests assert the actual PRAGMA values after opening every connection.
On recovery, file bytes beyond the committed offset are unacknowledged tail and are truncated. A short file or checksum failure inside the committed range is never “repaired” by silently moving the database offset backward: mark the affected output unavailable/truncated, preserve evidence, and create an incident. Essential-state corruption interrupts only the affected command and gates unsafe mutations.
Apply this committed-offset protocol to command blob and audit segments too. For a missing committed output range, preserve all still-valid records, create query-visible truncation/incomplete metadata, and never fabricate byte totals. For a missing immutable execution spec, request hash, revision, or active stdin state, mark the command essential state untrustworthy and interrupt it. Recovery must be idempotent after a crash at every repair step.
If database or segment persistence fails, do not acknowledge the event. Surface the server health failure, stop dispatching work as appropriate, and keep the session/control loops responsive.
5.4 Retention transaction
Implement retention in a single serialized maintenance worker, never on the WebSocket receive loop.
Define a versioned deterministic charge formula: compressed/encoded payload and segment lengths plus a conservative fixed charge for each SQLite row/index entry. Do not attempt to attribute shared SQLite pages after the fact. Reconcile the logical counters transactionally and also reject allocations below the configured 256 MiB server filesystem free-space floor, regardless of logical quota headroom.
Implement one server Reserve(owner, kind, chargedBytes) path used by command
creation, script append, stdin, event append, and audit. Reuse the same contract
in the client spool, where it also covers raw execution-file preparation. Under
the store writer lock/transaction it checks, in order: hard field maximum,
per-command total and closeout reserve, per-client total, server/client-daemon
total, and filesystem floor. It either reserves all applicable tiers or none.
Release uses the same stored charge version; a future estimator change requires
a migration, not a silent reinterpretation of old rows.
- Account all stored command-owned data, including request/script payload, pending stdin, metadata/events, and output. The client additionally charges generated raw execution files while they exist. Reserve non-evictable active state before accepting it, including 64 KiB per-command closeout headroom and the bounded result for each accepted stdin/signal mutation; reject an allocation that cannot fit.
- Enforce the command's 10 MiB output window and 32 MiB total cap by rotating
oldest output segments and recording
OutputTruncation; never roll essential active state. - Enforce the 256 MiB per-client and 4 GiB server-wide caps by evicting terminal commands as whole UUID records in oldest server issue-time/UUIDv7 order. If active data alone creates pressure, rotate/drop output with markers and reject new essential state.
- Reclaim every whole terminal command 30 days after authoritative terminal time by default, independently of byte pressure; zero explicitly disables age rotation. Preserve compact tombstones and separate audit/incidents.
- Maintain audit events in compressed segments under their independent 100 MiB default, rotating complete oldest segments. Apply an independently configurable audit age as a second trigger; zero, the default, disables age rotation but not byte-budget rotation.
- Keep unresolved storage incidents non-evictable. Resolved incident summaries may rotate with their audit history under the separate 100 MiB budget, but resolving dirty health never immediately rewrites or disguises recorded loss.
Retention selects candidates deterministically in (server_issue_time, issue_uuid) order and rechecks terminal state plus generation inside the delete
transaction. Before deleting a command, close follower snapshots over its
retained ranges, ensure the compact tombstone exists, mark evicting, and move
all owned files to a same-filesystem deletion directory using collision-proof
names. A crash-safe sweeper rolls forward rows/files in either order. Never use
a glob or client-provided path for deletion.
For active pressure, first evict server-retained output and record exact
RetentionTruncation; then reject new unreserved mutations. Do not evict the
assigned client send window or essential lifecycle/revision/dedupe state. Age
rotation and byte-pressure eviction call the same whole-command primitive so
their crash behavior cannot diverge.
Deletion is recoverable: mark rows evicting transactionally, move segment
files to a same-filesystem tombstone location, commit metadata deletion, then
remove tombstones asynchronously. Startup completes or rolls forward stale
evictions deterministically.
Replay tombstones must never have a delete/insert gap. The server inserts the command tombstone transactionally with authoritative terminal state before it can acknowledge or later evict command history. On the client, receipt of an acknowledgement through the terminal event atomically inserts or retains the UUID/request-hash tombstone and removes the full command record. Enforce each one-million-entry FIFO cap in that same transaction.
5.5 Incident resolution
Run safe, deterministic repairs automatically and through the control API.
RepairStorageIncident must refuse any action that would knowingly discard
committed data. AcknowledgeStorageIncident requires an operator note and is
the explicit path for accepting irrecoverable loss. Both are idempotent by
request_id. A repaired or acknowledged incident no longer contributes to the
derived dirty flag; resolution itself does not delete history. Bound notes to
4 KiB and retain resolved summaries/audit entries under their separate 100 MiB
rotation.
Provide the same store library behind offline rvbox-server repair --data-dir
list/repair/acknowledge modes. Provide equivalent local client-spool handling
through rvbox repair --state-dir and its local health diagnostics. Never open
a live store from the offline tool; acquisition of the same single-instance lock
is required.
Define stable incident kinds for uncommitted tail, missing committed bytes,
checksum mismatch, counter mismatch, SQLite integrity failure, failed eviction,
permission error, and disk exhaustion. An incident has immutable evidence plus
mutable state OPEN -> REPAIRED|ACKNOWLEDGED; terminal states never reopen, so a
recurrence creates a new incident ID. Scope keys identify global store, client,
command, segment, or audit without embedding raw filesystem paths in the public
API.
Safe repair may truncate only bytes beyond committed_end_offset, rebuild a
derivable counter/index, or roll forward a marked eviction. Anything that would
discard acknowledged bytes or immutable essential state requires acknowledgement
instead. Online repair acquires the same per-scope maintenance lock as recovery
and retention; offline repair additionally requires the process-wide instance
lock. Persist the repair/acknowledgement mutation and audit record before clearing
derived dirty health.
5.6 Store unit and integration tests
Unit-test record framing/checksums, size arithmetic, quota reservations, eviction ordering, age/tombstone/audit rotation decisions, cursor boundaries, incident-state transitions, and fault-point state-machine outcomes without opening SQLite. Run these with shuffled order, deterministic seeds, and race coverage for the in-process ownership/maintenance-lock coordinators.
The store integration suite uses a real temporary SQLite database, WAL, and
segment filesystem for: duplicate event idempotence; conflicting duplicate
rejection; crash between segment write and metadata transaction; crash after
metadata commit; corrupted final record; command-window rotation; per-client/
global whole-command eviction; active-command protection; cursor stability;
uncommitted-tail truncation; short/corrupt committed ranges; dirty derivation
and repair/acknowledgement; and bounded disk-full/permission failure simulation.
Every case records a named fault checkpoint in the shared run journal and can be
resumed at its reopen/invariant-verification step. Verify recovery endpoints
become live promptly while unsafe scoped mutations remain gated.
Exit criteria: a store restart preserves acknowledged data and never acknowledges uncommitted data; all rolling, tiered unified-quota, 30-day terminal-age, audit, and tombstone rules match the documented semantics.
6. Server sessions, WSS endpoint, and dispatch (Phase 3)
6.1 Agent endpoint
Expose an HTTP endpoint suitable for nginx WebSocket proxying. TLS may terminate at nginx, but the server must verify WebSocket upgrade origin/path/configuration and impose read limits itself. It accepts one binary protobuf envelope per message and refuses text messages, fragmentation misuse, oversized frames, invalid compressed payloads, and malformed/unknown mandatory messages.
Use four independently bounded activities per session:
- a read loop that validates/fragments/decode-dispatches only;
- a prioritized write loop with reserved Ping/Pong/Close/fencing/error capacity, bounded data frames, and write deadlines;
- a dispatch worker that reads persisted command work;
- a persistence/event worker that writes through
store.
No goroutine may hold a session-registry lock while doing database, filesystem, compression, or WebSocket I/O. Every goroutine receives a session context and must terminate on fencing/close.
Keep at most 1 MiB unacknowledged per command and 8 MiB per client session; additional data waits in durable storage. If an essential ingress queue fills, close without acknowledging rather than blocking the sole reader; the durable sender retries after jittered reconnect. Ping/Pong remains sendable even while the data window is full.
Implement the write scheduler as two bounded lanes owned by one socket writer: the reserved control lane contains Ping/Pong/Close, welcome/fencing, protocol errors, acknowledgements, and reconciliation control; the data lane contains dispatch, script, stdin/signal, and events. Always drain a bounded burst of control before one data frame so data still progresses. Impose write deadlines and split data below the frame ceiling. A full control lane is a session-fatal internal invariant; a full data lane leaves work durable and wakes the writer later rather than allocating another goroutine.
The reader checks WebSocket type/size before protobuf decode, validates envelope session fields before payload allocation, and routes only lightweight immutable work descriptors. Persistence workers preserve per-command ordering while a fair round-robin/deficit scheduler prevents a verbose command from monopolizing decompression or disk. Unit-test queue capacities and goroutine shutdown with a fake connection that never completes writes.
Apply the server 64 MiB/32 MiB ingress backlog as admission hysteresis, not as a license to discard client-sequenced events. Above high, stop acknowledging new output and close the producing session if the bounded descriptor queue cannot accept its next frame; the client's durable spool retries it. Resume normal admission only below low. Keep lifecycle/control reserve independent so another client's heartbeats and command closeout remain responsive.
6.2 Registration and fencing
On ClientHello, validate client ID/capabilities/version and the durable random
client-instance UUID; choose the highest shared minor under major v1; persist
the registration; allocate a cryptographic random opaque session ID and
increment the client's generation transactionally. Fence the prior session for
the same client ID and instance before making the new one active. Accept a
different instance normally if no session for that client ID is live. Otherwise
record it as the pending claim and reject it unless an unexpired one-shot grant
matches that exact instance; consume the grant transactionally when accepting
the reconnect and fencing the old session. The default grant TTL is 5 minutes.
Send ServerWelcome with the outer envelope's canonical session fields.
Persist a rejected live-conflict claim before closing the candidate socket and
expose only its exact client_instance_id and observation time through
GetClient. AuthorizeClientTakeover verifies that this ID still matches the
pending claim, stores an idempotent UUIDv7 mutation and five-minute expiry, and
does not fence immediately. On matching reconnect, one transaction increments
generation, consumes the grant, closes the old session record, and installs the
new session; only then may the writer send welcome. A mismatched/expired grant is
never broadened to “the next instance.” If no session is live, accept a different
instance without creating or consuming a grant, as agreed.
All post-hello envelopes must match both session ID and generation. Stale or unknown session traffic is discarded/audited and its WebSocket is closed. A dispatch includes the intended generation, and the client must reject a mismatch. Unknown client IDs are accepted by design; hostname is display/routing identity, not an authentication assertion.
6.3 Heartbeat and liveness
Implement the identical policy at both ends: any received valid WebSocket frame updates inbound activity; after 10 seconds silent send Ping; after 30 seconds without inbound frame close the session. Pings/Pongs use WebSocket controls, not protobuf envelopes. Make timers injectable/fakeable for deterministic tests.
Use a monotonic clock for elapsed activity/backoff and wall-clock UTC only for persisted observation. Pong/any valid inbound frame updates activity in the read loop without waiting for a database worker. Reset reconnect backoff only after a continuous 60-second stable session. Add boundary tests at exactly 10 and 30 seconds, delayed writer tests proving Ping bypasses a full data lane, and clock jump tests proving wall-clock adjustment cannot spuriously kill a session.
6.4 Queueing, dispatch, and reconciliation
RunCommand persists a queued command before returning. The dispatcher sends
to the active session while either an immediate running slot or a client queue
slot is available: running < max_running || queued < max_queued. Dispatch and
capacity updates are serialized per session so concurrent sends cannot
oversubscribe the advertised slots. It updates to dispatched before write and
retries after reconnect until a matching CommandAccepted arrives.
Maintain per-session shadow reservations for in-flight dispatches: reserve a
running slot when immediately available, otherwise a queued slot, before
enqueuing the frame. Reconcile the shadow with each ClientCapacity and
CommandAccepted; release it on definite rejection/session close. The
dispatcher uses one serialized worker per client plus a fair global ready queue,
so the OR capacity condition cannot oversubscribe and an offline/noisy client
cannot starve others.
Apply the configurable queue TTL (15 minutes by default; zero means indefinite)
until server-confirmed acceptance. A never-dispatched command becomes terminal
expired. A dispatched command with uncertain acceptance remains protected and
is rendered expired pending reconciliation rather than being prematurely made
terminal. Late client evidence advances the actual lifecycle, creates a
late-after-expiry incident, and triggers best-effort termination while retaining
subsequent actual outcome events.
Persist an absolute server-clock queue_expiry_time when creating the command;
never recompute it after configuration changes or retries. Drive expiry from a
database-backed ordered index and recheck lifecycle/revision transactionally.
Queued commands become EXPIRED; dispatched commands set a derived
late-pending-reconciliation indicator without writing a false terminal event.
On late acceptance, persist the incident and higher termination revision before
sending best-effort TERM, then retain every actual subsequent event.
Classify a negative CommandAccepted before changing lifecycle. Explicitly
transient capacity/storage pressure returns dispatched -> queued with bounded
jittered backoff and remains under the original TTL. Invalid execution data or
unsupported required platform controls become terminal rejected with the
structured reason persisted on CommandRecord. Reserve failed for a process
that reached running.
On replacement/reconnect, query non-terminal commands for that client and send
a ReconcileRequest with last durable event sequence/revision plus immutable
request hash. Require one complete client snapshot of all locally retained
commands, including
terminal-but-unacknowledged records and matching requested tombstones, and
compare the sets idempotently before enabling dispatch. With healthy client
storage, an absent queued/dispatched command returns to queued; an absent
accepted/running command is interrupted as an invariant failure. A matching
UUID/request-hash tombstone suppresses replay. The client terminates a local
active command absent from the server target set only after receiving the
server's durable ReconcileResult; either peer records a contradiction with
server-confirmed terminal state as a storage/recovery incident, and the server
does not invent missing state. The result also lists terminal local records safe
to discard when the server has stored or deliberately tombstoned them, covering
a lost final EventAck followed by server retention. Write this idempotent
result before dispatch. A client with unresolved essential-store corruption
cannot complete reconciliation or accept work.
Implement reconciliation as this explicit matrix:
| Server state | Client evidence | Durable result |
|---|---|---|
| queued/dispatched | absent from healthy complete snapshot | return/remain queued; redelivery permitted |
| dispatched/accepted/running | matching retained record | keep highest valid revision and resume event/script delivery |
| any non-terminal | matching tombstone | interrupt/suppress replay and record stale-server incident |
| accepted/running | absent | interrupt and record client-state-loss incident |
| missing or contradictory terminal | client active | ReconcileResult.terminate_local_issue_uuids plus incident; invent no server command |
| missing/tombstoned or fully stored terminal | client terminal | discard_local_terminal_issue_uuids |
Validate matching UUID, immutable request hash, revision monotonicity, terminal
immutability, and event-sequence bounds for every row before applying any row.
Persist all server state/incident decisions in one reconciliation transaction,
then send one deterministic sorted ReconcileResult through the control lane.
The client durably records terminate/discard intent before reporting capacity.
If the result frame is lost, the next complete snapshot yields the same result;
dispatch is disabled until the current result has been written on the session.
Handle pre-start kill locally only for work that has never been dispatched.
For dispatched or accepted work, persist a higher revision and send the
revisioned signal while retaining the existing visible lifecycle until the
client resolves the race. Accept a revisioned cancelled lifecycle if launch
was not authorized; otherwise retain the revisioned signal result and actual
terminal outcome. The server maps control request_id to this revision/result
and never sends that request ID across the agent protocol.
6.5 Server session unit and integration tests
Unit-test version selection, registry/generation transitions, takeover grant matching, capacity-shadow accounting, dispatch fairness, TTL/revision races, reconciliation matrix decisions, heartbeat/backoff boundaries, queue saturation, and writer-lane scheduling with fake clocks, stores, and sockets.
The server-session integration suite runs the compiled server with real
SQLite/filesystem state and real WebSocket protobuf frames. Cover same-instance
replacement, different-instance pending claim, exact one-shot takeover
consumption/expiry, stale-generation late output, Ping/Pong inactivity
thresholds, full client queue, dispatch retry after lost acceptance, complete
active/terminal-unacknowledged reconciliation, lost/repeated ReconcileResult,
permanent/transient dispatch rejection, cancellation before launch, launch/
cancel race, malformed-envelope close, server restart, and a slow client that
cannot block another client or local control RPC. Use the harness fault proxy
and named acknowledgement checkpoints rather than sleeps or probabilistic packet
loss.
Exit criteria: multiple simulated clients can register, replace each other, receive bounded dispatches, and recover connection faults without duplicate execution or server deadlock.
7. Windows-first client runtime, spool, and supervisor (Phase 4)
7.1 Client state and reconnect runtime
Keep client state private: mode 0700 on Unix; on Windows disable inherited
ACLs and grant only SYSTEM and Administrators the required access. Store durable
accepted-command records, active process metadata, command event journal,
stdout/stderr spool segments, stdin write acknowledgements, and outstanding
script upload state. It contains
no terminal command history after server acknowledgement and local cleanup.
It retains a separate compact FIFO of the most recent 1,000,000 command
tombstones (binary UUIDv7, immutable request hash, and acknowledgement time) so
stale server state cannot replay recently completed work after full history is
removed.
The runtime has a long-lived supervisor/dispatcher and replaceable network session. Generate the random client-instance UUID once in the private state directory and retain it across daemon restarts and reconnects. Network loss must not stop running commands or pipe readers. The connection loop uses full-jitter exponential backoff from 1 second to 60 seconds, resetting only after 60 stable seconds. On each connection it sends a hello with the durable client-instance UUID, waits for welcome, replays retained sequenced loss events and other unacknowledged events in order, answers reconciliation, and resumes dispatch.
All client queues are bounded. The disk spool, not an in-memory channel, is the
source of truth. Give stdout, stderr, status, signal results, and stdin
acknowledgements one durable command-local order. Assign and persist wire
event_seq only when an entry enters the bounded send window; an assigned entry
is pinned until acknowledgement and immutable on retry.
Create client_instance_id with a cryptographic UUID generator, write and fsync
a temporary file, atomically rename it, and sync the state directory before
first registration. Never regenerate it merely because the network/server
rejects a session. If the identity file is corrupt, enter dirty local health and
require the repair path rather than silently presenting a new installation.
Define durable-file primitives per platform and use them everywhere. Unix uses
file fsync, atomic same-filesystem rename, then parent-directory fsync.
Windows writes a same-directory temporary file with restrictive ACLs, calls
FlushFileBuffers, atomically installs it with ReplaceFileW or
MoveFileExW(MOVEFILE_REPLACE_EXISTING | MOVEFILE_WRITE_THROUGH) as applicable,
and flushes a directory handle where supported. Treat an unsupported/failed
required flush as a storage error; do not acknowledge durable state based only
on buffered writes. Test power-loss boundaries on NTFS, not only process exit.
Give the client store the same committed-offset segment discipline as the
server. Persist, at minimum, command phase/revision/hash, platform launch
identity, local event ordinal, optional assigned wire sequence, payload digest,
stdin write state, script received ranges, and last server acknowledgement.
Maintain separate monotonic local_ordinal and event_seq columns: insertion
assigns only local order; admission to the send window transactionally assigns
the next contiguous wire sequence and pins the row. Cumulative EventAck
atomically unpins/deletes acknowledged payload and updates the command cursor.
Structure reconnect as explicit states: backoff -> connecting -> hello -> reconciling -> active -> closing. Only pipe/supervisor workers outlive the
replaceable network session. After welcome, send the complete snapshot, apply
ReconcileResult durably, replay assigned rows first, then assign/send local
rows in order. Advertise capacity and accept new dispatch only in active.
Cancellation of one session context must join all its readers/writers before a
new session can use their queues.
The current implementation checkpoint is intentionally split at this seam:
internal/client/agent.RunOnce owns the reconnect/session reader and remains
the sole live-session writer; internal/client/spool owns the SQLite source of
truth; and internal/client/agent.Executor owns process lifetime independently
of the WebSocket context. A bounded event-notification channel wakes the active
writer to assign and send newly appended events, while reconnect replay uses
the same spool rows and a per-session sent cursor. Every accepted dispatch is
stored with its deterministic execution specification (and, for scripts, its
descriptor reservation) before acknowledgement.
Add a schema launch barrier to every implementation of the executor. Persist
prepared before entering the supervisor, persist authorized before the OS
release boundary, and clear it only when a terminal lifecycle transition is
committed. Startup recovery must convert any non-terminal authorized row to
one interrupted event before registering a new network session. This is the
at-most-once fence for a crash between process release and the first running
event; it is not a substitute for verifying a native process creation identity.
7.2 Output capture and offline caps
Create non-blocking readers for stdout and stderr immediately after process start. Read raw bytes irrespective of newline boundaries, divide them into <= 64 KiB raw chunks, add observed UTC time/local order, Zstandard-compress, append to the command spool, and enqueue only a durable reference to the sender.
Use one reader per pipe and a bounded shared raw-chunk scheduler. A reader never waits for network I/O. Under normal load it transfers ownership of a bounded buffer to a fair compressor, which returns buffers to a pool only after durable append. Validate Zstandard frame size/checksum on readback before replay. Do not log raw chunks on compression, checksum, or storage failure.
Bound work before compression as well as on disk. Defaults are high/low raw
backlogs of 1 MiB/256 KiB per command and 8 MiB/4 MiB for the client daemon; the
server compression/persistence tier independently uses 64 MiB/32 MiB globally
with the non-dropping admission behavior in Phase 3. Crossing a client high
watermark enters loss mode until the matching low watermark. Aggregate
discarded, still-unsequenced chunks into local-order/raw-byte loss records and
replace them with a CLIENT_OVERLOAD truncation marker; event-range and
compressed-byte counts are absent when sequencing/compression never occurred.
Use fair worker scheduling and reserved metadata capacity so lifecycle,
stdin/signal
acknowledgements, gap markers, and terminal state remain lossless.
Loss mode is hysteretic and scoped: crossing either command or daemon high
watermark marks affected raw chunks as dropped until both relevant backlogs fall
below their lows. Coalesce only adjacent dropped local ordinals with identical
cause/stream set; close a loss run before inserting any essential event. Persist
known raw-byte totals and observation interval in local metadata, then emit one
bounded OutputTruncation; omit event-range/compressed-byte fields if they were
never assigned/compressed. If even reserved marker persistence fails, mark the
command/store dirty and interrupt rather than silently claiming complete output.
Respect the local hard limits at all times:
- 10 MiB compressed rolling output window and 32 MiB total per command;
- 256 MiB aggregate command-owned storage per client daemon.
When connected, discard server-acknowledged oldest segments first. Assigned but
unacknowledged entries are pinned and bounded by the 1 MiB per-command send
window. When offline or an acknowledgement lags, rotate only still-unsequenced
output as needed and replace each removed run in durable local order with
OutputTruncation using CLIENT_SPOOL. The marker later receives a normal
wire sequence and is retained until cumulative EventAck; no wire sequence is
fabricated or skipped. The server's accepted truncation event is the auditably
visible explanation for the lost bytes.
Continue draining child pipes after caps are reached. A verbose child may lose old or newly produced output, but it must not block the client daemon or deadlock itself.
Test a continuously verbose process, alternating stdout/stderr, essential events interleaved with dropped output, reconnect during loss mode, ack loss at every send-window boundary, and terminal exit while above the high watermark. Assert contiguous assigned sequences, exact known raw-byte totals, bounded heap/queue size, and a terminal event after every marker/incomplete event.
7.3 Command acceptance and idempotency
Persist an accepted dispatch record keyed by issue_uuid before responding
CommandAccepted. Repeat dispatch with identical immutable fields returns the
same acceptance/revision and does not start another child; a conflicting repeat
is a protocol error. Enforce 16 running/100 queued defaults before acceptance.
After full terminal cleanup, also check the compact tombstone ledger: the same
UUID/hash is rejected as already executed, while changed immutable content is a
conflict.
Make acceptance a single durable state transition:
- Validate the session generation, UUIDv7 syntax, immutable request hash, command revision, shell/profile/CWD policy, script descriptor, and expiry.
- Look up both the active-command table and tombstone ledger before reserving
capacity. An exact active duplicate returns its existing result; an exact
tombstone returns
CODE_ALREADY_EXECUTED; either hash mismatch is a permanent conflict. The server treats the former as stale-state reconciliation/interruption, not as a newly rejected execution. - Reserve the running/queued slot and worst-case command metadata allocation in the same local-store transaction that inserts the accepted command. Do not count retransmission of an existing record twice.
- Commit and fsync the accepted record, immutable hash, revision, script cursors, and capacity counters before sending positive acceptance. A failure before commit returns a transient rejection and leaves no partial command; loss of the response is resolved by exact replay.
- Classify policy/schema/hash/unsupported-profile failures as permanent and local I/O/temporary-capacity failures as transient. Include a bounded, operator-readable reason without command or environment contents.
For script dispatch, persist descriptor and chunks at validated offsets, verify
each chunk checksum and final SHA-256/length before launch, and keep the upload
restart-safe until terminal cleanup. Create the final temporary file with
owner-only permissions under the effective CWD using an atomic write/rename;
delete it after the terminal event is durable locally. Emit cumulative durable
ScriptUploadStatus.received_bytes progress often enough to advance the 1 MiB
server send window; the server never streams the full 10 MiB without progress
acknowledgement.
Accept a chunk only when its offset equals the durable contiguous prefix, or
when the entire byte range is an exact replay of already persisted bytes.
Reject holes, changed overlaps, arithmetic overflow, data past the declared
length, and chunks received after commit. Append and fsync bytes before
advancing received_bytes; after the final prefix, verify length and SHA-256,
reserve the raw-file quota, atomically materialize the wrapper, sync its parent
directory, and only then make the command launch-eligible. A permanent upload
failure durably ends the command as REJECTED before launch authorization and
releases every associated reservation.
On Windows create wrappers with an explicit non-inherited ACL, exclusive random basename, and delete-sharing suitable for terminal cleanup; use the common Windows durable replacement primitive rather than POSIX rename assumptions.
Materialize ordinary command_text through the same private generated-file
machinery, while retaining its separate command-text metadata. Execute the file
with exactly sh/bash, cmd.exe /D /S /C, or powershell.exe with
-NoLogo -NoProfile -NonInteractive -File according to shell_type; never infer or
fallback to another shell. Build Windows application name/arguments with the
platform quoting routine rather than shell string concatenation. Treat $0,
%0, or equivalent exposing the generated wrapper name as documented behavior.
Resolve each shell setting once during configuration validation to a canonical
absolute executable path. Launch that exact path with an explicit argument
vector/application name; never search the daemon's PATH, honor a command's
environment overrides during resolution, follow a wrapper shebang, or select by
file extension. Re-stat the configured executable immediately before launch and
reject if it is no longer the validated regular executable. Add tests with an
attacker-controlled PATH entry and same-named fake shell to prove it cannot be
selected.
7.4 Supervisor contract and Windows process implementation
Define a narrow interface used by the runtime:
type Supervisor interface {
Start(context.Context, StartSpec) (Process, error)
Signal(context.Context, Process, SignalKind) (SignalOutcome, error)
Snapshot(context.Context, Process) (ResourceSnapshot, error)
StopAll(context.Context) error
}
StartSpec includes the validated elevated intent; a successful Process
exposes an immutable effective WindowsExecutionIdentity after preparation,
and a context-selection error carries the same record without an effective
context. The runtime persists the intent, attempted contexts, selection detail,
and optional effective identity. It emits the record with a pre-launch
rejection or with running before reporting requested code as started.
Implement Windows code in platform-specific files so non-Windows builds never
import Windows APIs. Keep launch phases identical across platforms:
accepted -> launch_prepared -> launch_authorized -> running, with no shortcut.
The first native adapter is now required to expose this contract through
internal/client/supervisor.Supervisor: the non-Windows adapter is test-only,
while the Windows implementation must perform token selection and Job setup
inside the same Start call. It may return only after the child has been
assigned to its kill-on-close Job and released; all token/session attempts must
be represented in the returned immutable identity. A failed start clears the
pre-launch barrier and produces one rejected lifecycle event; an uncertain
authorized row is never retried as a fresh process.
Implement one exhaustive token selector; do not scatter token fallback across launch code:
- Reject
elevated=trueon non-Windows in v1. On Windows, enumerate WTS sessions immediately before preparation. Prefer the physical console only when it isWTSActivewith a valid user token; otherwise accept exactly oneWTSActiveinteractive session. Zero candidates means logged out. Multiple candidates without a preferred console are recorded as ambiguous and treated as no usable active user, never selected nondeterministically. - With a usable active user and
elevated=false, attempt onlyACTIVE_USER. Use an existing filtered/standard token. If Windows supplies only a full administrator token, callCreateRestrictedTokenwith LUA/max-privilege restriction, make administrator/privilege-bearing groups deny-only, set and verify medium integrity, and retain the user's SID/session. If a verified non-elevated token cannot be built, reject; do not use LocalService or SYSTEM while that selected active session remains valid. - With a usable active user and
elevated=true, attempt contexts in this exact order:ACTIVE_USER_ELEVATED,ACTIVE_SYSTEM,LOCAL_SYSTEM.ACTIVE_USER_ELEVATEDqueriesTokenElevationType, uses a linked full token for a traditional limited administrator, and accepts an already-full admin token. A standard user/missing linked token recordsCODE_ELEVATION_UNAVAILABLE; Administrator Protection or another just-in-time approval policy records a bounded policy-specific attempt reason. V1 shows no UAC/Hello prompt and continues toACTIVE_SYSTEM. ACTIVE_SYSTEMduplicates the service SYSTEM token, setsTokenSessionIdto the selected session whileSeTcbPrivilegeis enabled only around that call, verifies SYSTEM SID/session, and hides all windows. A pre-preparation failure records its error and continues to Session 0LOCAL_SYSTEM. This fallback intentionally bypasses user-scoped approval because the installed service already holds SYSTEM and must be conspicuous in status/audit history. The finalLOCAL_SYSTEMfallback supplies elevation but no interactive-desktop access; session-scoped commands may launch successfully and then fail in the ordinary way in Session 0.- With no usable active user,
elevated=falseattempts onlyLOCAL_SERVICE: passwordlessLogonUserWforNT AUTHORITY\LocalServicewithLOGON32_LOGON_SERVICE, required SIDS-1-5-19, and session 0. Do not use a restricted SYSTEM token as a substitute.elevated=trueattempts onlyLOCAL_SYSTEMby duplicating the service token and requiring SIDS-1-5-18plus session 0. - Fallback decisions cover token/session capability only and all complete
before
launch_prepared. Never fall back after launcher creation, shell creation, CWD/executable/profile validation failure, authorization, or an uncertain outcome. If the selected user logs out during selection, close all provisional handles, enumerate once more, and apply the resulting row; never switch silently to another user. - Verify the final user SID, session ID, token type, elevation state, and integrity level against the selected context. Close every token/profile handle on all paths. Persist the request's elevation bit, ordered attempted contexts with a total bounded selection detail, optional effective context/ user SID, and target session/user/logon SIDs; never persist token handles or credentials. If every context fails, emit that decision record with no effective context on the terminal pre-launch rejection.
Build the base Unicode environment with CreateEnvironmentBlock for the
effective token, then apply validated request overrides deterministically.
Load/unload an active user's profile only when required and keep it loaded until
the complete Job exits. Session 0 contexts use their service-account profile and
cannot see interactive mapped drives. Validate an explicit CWD and wrapper ACL
access while impersonating the effective token. For an omitted CWD, treat
%ProgramData%\RVBox\work (or configured daemon_cwd) as a SYSTEM-owned root
and create/open an ACL-isolated child keyed by the selected user SID,
LocalService, or SYSTEM context. Reject insecure owners, inherited write grants,
and reparse points on reuse. Generated wrappers grant only SYSTEM and the
effective token SID the required access.
- Create a non-inheritable per-command Job Object, enable
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, and do not enable breakaway. - Resolve the hierarchy's effective token/session as specified above. Create
unpredictable local named pipes for control/stdin/stdout/stderr with
PIPE_REJECT_REMOTE_CLIENTSand an ACL restricted to SYSTEM and the effective token SID. Do not rely on inherited service handles: Windows prohibits normal handle inheritance across Terminal Services sessions. - Create a stateless per-command launcher mode of the canonical
rvbox.exesuspended throughCreateProcessAsUser, with no inherited handles and an opaque pipe-channel identifier. Assign this still-suspended launcher to the empty Job, then resume it. Verify the connecting pipe client's PID, creationFILETIME, token SID, session ID, and generation before sending any launch material. - The launcher opens only the verified pipes and creates the requested shell
suspended with
CREATE_NEW_CONSOLE,CREATE_SUSPENDED,CREATE_UNICODE_ENVIRONMENT, andEXTENDED_STARTUPINFO_PRESENT. UseSTARTF_USESHOWWINDOW/SW_HIDE; pass only standard-I/O handles withPROC_THREAD_ATTRIBUTE_HANDLE_LIST. Job membership propagates because breakaway is disabled. Report shell PID/creation time and wait for release. - Persist and flush
launch_preparedwith requested elevation, attempted/ effective contexts and reasons, token SID, session/user/logon SIDs when active, launcher/shell birth identities, Job generation, and console identity. Recheck revision/cancellation, WTS session identity, executable identity, CWD access, and requested Job limits while requested code remains suspended. - Persist and flush
launch_authorized, then send the authenticated release message. The launcher resumes the shell and acknowledges the attempt; recordrunningwith effective identity only after that acknowledgement. Failure or crash after authorization terminates the Job and becomes interrupted, never redispatched. - The service retains the sole Job handle. Control EOF before authorization makes the launcher exit without executing; EOF afterward terminates the Job. Service exit therefore kills launcher, shell, and descendants in every crash window. Recovery never signals or kills by PID alone.
Do not combine CREATE_NEW_PROCESS_GROUP with CREATE_NEW_CONSOLE: Windows
ignores the former, and it is unnecessary because every command owns a distinct
hidden console. For TERM, start a short-lived private signal-helper mode of the
same canonical rvbox.exe using the command's effective token/session. Pass the
target identity through a separate unpredictable, PID/token/session-verified
local named pipe rather than inherited cross-session handles or the command
line. Prefer the root shell while it is live;
otherwise select a live PID from the Job's process list and verify its creation
FILETIME/generation before use. The helper calls AttachConsole(target_pid),
installs a control handler that consumes its own CTRL_BREAK, then calls
GenerateConsoleCtrlEvent(CTRL_BREAK_EVENT, 0) so only processes attached to
that command's console receive it. It reports a checksum-framed result and
detaches. The helper receives neither the Job handle nor command stdio.
Apply PROC_THREAD_ATTRIBUTE_HANDLE_LIST to every same-session child creation.
Every database, log, listener, Job, unrelated pipe, and SCM handle is non-
inheritable. Length/type/generation/checksum-frame launcher/signal-helper
control messages and reject mismatched generations.
Open the configured shell by canonical absolute path and pass it as explicit
lpApplicationName. Use one reviewed Windows argument-quoting routine and an
explicit Unicode environment block. Never invoke %COMSPEC%, search
PATH/file associations, or let command environment overrides choose the
executable.
Associate the Job with an I/O completion port. Use
JOB_OBJECT_MSG_ACTIVE_PROCESS_ZERO plus capture-pipe EOF to establish complete
tree/output drain. Persist shell creation FILETIME before trusting a PID.
Recovery may reopen a process only to compare creation time/generation;
inability to prove ownership creates a dirty incident rather than risking
termination of a reused PID.
Accept only TERM/SIGTERM and KILL/SIGKILL. TERM runs the verified
console-attach helper, waits the configured 10 seconds, then terminates the Job
if necessary. KILL terminates it immediately. Emit the actual attach, delivery,
wait, and escalation outcome. Root exit starts the common 5-second drain grace;
the terminal lifecycle event remains last.
Compile requested TOML profiles before launch and apply CPU, memory, active
process, and supported I/O rate limits to the empty Job. Read back every required
limit before authorization. Permanently reject UNSUPPORTED rather than run
partially constrained. Use process/Job accounting for CPU/RSS/I/O and omit
Linux-only wait reasons.
7.5 Windows desktop host and lifecycle
Build the executable with the Windows GUI subsystem so service/helper/tray modes
never flash an unwanted console. Human-invoked --help, --check-config,
install, uninstall, and configuration modes call
AttachConsole(ATTACH_PARENT_PROCESS), reopen the inherited standard handles,
and use UTF-8 terminal diagnostics when attached; define native-dialog plus exit-
code behavior when no console exists and test both paths. install-service and
uninstall-service use ShellExecuteEx(..., "runas", ...) only when their caller
is not already elevated, validate the canonical executable/config paths, and
make idempotent SCM changes. The installed service command line selects only the internal
service mode and exact config path. At service entry, verify SCM startup rather
than accepting that mode from an ordinary process invocation. Hold one machine-
wide service/state lock and report SCM start/stop checkpoints within their
deadlines; lengthy spool recovery remains asynchronous and does not block
service start.
Register the tray with the machine-wide Run key, not Task Scheduler and not
service-created cross-session injection. Use one ACL-protected mutex per WTS
session so each logged-in user gets at most one tray while multiple RDP/console
sessions remain independent. Run its Win32 message pump on one locked OS thread,
recreate the icon after TaskbarCreated, and use bounded nonblocking IPC/UI
channels. The tooltip shows connection/readiness/dirty state without payloads.
Exit closes only that tray. The service continues through tray exit, user
logoff, Explorer restart, and periods with no interactive user.
Expose a versioned, length-bounded local named-pipe tray protocol with
PIPE_REJECT_REMOTE_CLIENTS, explicit SYSTEM/Administrators/interactive-user
ACLs, peer PID/session/token inspection, deadlines, and no command payloads.
The tray never reads the spool/database. Read-only status may be served to a
local interactive user. For a mutating service/config action, the tray launches
the narrow canonical configure-service mode with UAC; that helper sends one
authenticated request and exits. The service rechecks the helper token rather
than trusting a tray-supplied administrator claim.
Resolve defaults under %ProgramData%\RVBox; create the config, state, and log
directories with explicit Administrators/SYSTEM ACLs and reject insecure
existing objects. Create %ProgramData%\RVBox\work as a SYSTEM-owned root and
create identity-scoped child directories on demand with non-inherited ACLs for
SYSTEM plus the effective non-SYSTEM SID; no command identity receives state-
directory access. Give tray users only the documented read access to config/
current logs. On first installation, copy the annotated Windows
TOML template and leave the service live/not-ready until routing is valid. Open config and Open log use exact resolved regular paths—never an unquoted shell
command—and show actionable errors. Config edits are administrator-only and
take effect after explicit service restart; v1 has no hot reload or TOML rewrite.
Install the service with SERVICE_AUTO_START. configure-service maps the tray
Automatic/Manual control directly to SCM start type; there is no
windows.start_on_boot TOML key or second desired-state store. Stop/restart is a
separate administrator action and reports that stopping the service terminates
all supervised Jobs. Uninstall first stops admission, performs bounded orderly
shutdown, removes the Run entry and service registration, and preserves config/
state/log data unless a separately confirmed purge operation is designed later.
Use bounded rotating service file logs because SCM has no useful interactive
stderr. Flush warnings/errors promptly, redact as elsewhere, and make Open log
target the current file after rotation. Add native icon/version/service metadata;
release signing and checksum publication are Phase 8 gates.
7.6 Windows-first client unit and native integration tests
Unit-test shared runtime/spool behavior with fake transport/clock/storage fault points. Unit-test the exhaustive login/elevation selection function with token- property DTOs rather than Win32 handles, and separately test Windows quoting, environment, ACL descriptor, named-pipe frame, SCM transition, and tray- authorization helpers. Cover at-most-once duplicates, reconnect/replay, offline truncation, script validation, tombstone rotation, queue limits, and shutdown decisions. Run race tests with multiple commands and forced logical network churn.
The native windows-supervisor and windows-service-tray integration suites use
real Windows subprocesses and APIs for both shells, CWD/environment overlays,
concurrent output without newlines, stdin ordering/close, Job tree termination,
script materialization/cleanup, service shutdown interruption, and every token/
session context. They run through scripts/windows/test-host.ps1 under the
exclusive host lease and journal every machine-wide mutation for reset.
Add a table-driven crash suite for every acceptance, script-upload, launch,
event-spool, terminal-acknowledgement, and tombstone-rotation commit point. Run
the same duplicate dispatch after each restart and assert exactly one of:
durable rejection before authorization, one supervised process, or an
interrupted uncertain launch—never a second execution.
Fault-inject daemon death around Job creation, suspended shell creation,
launcher pipe authentication, prepared/authorized fsync, release/resume, and
terminal drain; fault the signal helper before/after attach, control-event
delivery, and result. Test paths with spaces/non-ASCII, malicious
PATH/COMSPEC, nested descendants, hidden-console behavior, CTRL_BREAK
refusal/escalation, Job-limit rejection, cross-session output/stdin, and
inherited-handle leaks.
Add a native table for every hierarchy row: active standard user, traditional
split-token administrator, already-full administrator/UAC-off token, enabled
Administrator Protection where available, logged-out machine, ambiguous RDP
sessions, and logout between selection/revalidation. Assert normal active-user
execution is medium/non-admin; elevated attempts are ordered
ACTIVE_USER_ELEVATED -> ACTIVE_SYSTEM -> LOCAL_SYSTEM; no-user normal/elevated
use LOCAL_SERVICE/LOCAL_SYSTEM; and attempted/effective identities are
persisted and displayed. Inject failure at each token attempt and prove fallback
occurs only before launch_prepared, never as a process retry.
Test idempotent UAC installation/uninstallation, SCM Automatic/Manual/start/ stop/restart, preservation of state on uninstall, service boot before login, service continuity through logon/logoff/tray exit, one tray per session, all tray actions and authorization checks, malicious named-pipe clients, Explorer restart, Run-key registration/removal, first-install bootstrap, ProgramData/work ACLs, Server Core headless behavior, and log rotation on each supported Windows version. Assert no Task Scheduler object is created or required.
Exit criteria: the native Windows client can stay alive through server loss/restart, execute up to capacity at most once, preserve/replay bounded history, manage complete Job trees, and satisfy tray/elevation/autostart behavior on Windows CI, including all execution-hierarchy rows. Reaching this gate does not automatically start the deferred Unix-client work; that requires an explicit later implementation decision.
8. End-to-end agent protocol (Phase 5)
Wire the server session layer and Windows client runtime together before adding the CLI. Use real protobuf bytes through an in-memory test transport first. Then run Linux server/nginx/storage containers and connect a native Windows CI runner to their short-lived WSS endpoint; do not substitute Wine or a cross-compiled binary that was never executed on Windows.
Implement flows in this order:
- hello/welcome/version/fencing and heartbeat;
- persisted background command dispatch and lifecycle events;
- stdout/stderr event persistence and cumulative acknowledgements;
- reconnect replay, duplicate delivery, and non-terminal reconciliation;
- queued cancellation and revisioned signal delivery;
- ordered stdin/close acknowledgement;
- verified chunked scripts;
- capacity advertisements, diagnostics, and profile results.
At each step, add a failure-injection test that drops one frame at every
acknowledgement boundary. Required scenarios include lost CommandAccepted,
lost EventAck, stale old-session output after takeover, connection loss during
script transfer, server restart before/after event transaction, client restart
with managed process cleanup, and offline spool overflow. Verify command UUIDs
never execute twice and output gaps always appear as truncation metadata.
Use the shared resumable harness from Section 2.4 with fake wall/monotonic
clocks in its deterministic peer, controllable frame loss/reordering, daemon
kill points, and bounded virtual disk capacity. scripts/test-e2e --scenario NAME runs one case; its printed --run-id can be passed back with --resume
after an intentional or accidental interruption. For each flow, run disconnect
before delivery, after delivery/before
durable commit, after commit/before acknowledgement, and after acknowledgement.
Assert database invariants and quota counters after every restart, not merely
the visible CLI result. Include these cross-cutting cases:
- a queued command expires at the default 15-minute TTL while disconnected,
then a late client reports a real terminal result and
statdisplays both expiry history and the updated terminal state; - server per-command, per-client, and global quotas cross high and low watermarks while assigned client events remain pinned;
- terminal age reclamation and command-level quota eviction create tombstones, retention metadata, and audit entries without removing active work;
- client identity collision, exact-instance takeover expiry/consumption, and an old connection attempting output after the new generation is committed;
- storage corruption found during asynchronous recovery, safe repair, explicit acknowledgement of unavoidable loss, and continued service in healthy scopes;
- heartbeat control traffic remains serviceable while both data lanes are full over a slow, intermittently writable WebSocket.
Exit criteria: a native Windows client against the Compose-hosted Linux server demonstrates a complete background command, history/follow behavior, stdin interaction, signal, reconnect, and restart recovery through nginx WSS. This is the first supported-client milestone. It satisfies a prerequisite for the deferred Linux supervisor work but does not automatically authorize it.
9. Server control plane and rvc (Phase 6)
9.1 gRPC over Unix socket
Start a Unix socket only after removing an existing socket only when it is
proven to be a stale socket owned by this server account; do not blindly unlink
arbitrary paths. Set mode 0600 immediately. Implement the generated Control
service against domain/store/session interfaces, never directly against socket
state.
Control response messages represent success only and contain no embedded error
alternative. Return canonical non-OK gRPC statuses with structured RVBox
details; the JSON-RPC adapter maps the same domain error rather than inspecting
response payloads. Agent-protocol rejection envelopes continue using
ControlError.
Implement and test:
ListClients,GetClient,ListCommands,GetCommand;AuthorizeClientTakeoverfor an exact displayed pending instance, with a one-shot 5-minute grant and normal mutation idempotency;RunCommandfor durable creation and immediate UUID return, followed byFollowCommandfor foreground ordered history plus live events;FollowCommandbeginning at a supplied event sequence and returning a wrapper around either a client event or non-sequenced server retention metadata;AppendStdin,CloseStdin, andSignalCommandwith persisted idempotency;GetOutputwith stream selection, byte-bounded opaque cursor pagination that can resume within an event, distinct client-observed/server-receipt times, and truncation/incomplete metadata;ListStorageIncidents,RepairStorageIncident, andAcknowledgeStorageIncident; dirty health is derived from unresolved rows, safe repair resolves automatically, and acknowledging loss preserves history.
Foreground context cancellation/timeouts stop only the local RPC stream. They must never send a remote kill. The server uses domain/store authorization of state transitions and an explicit command UUID for every mutation.
Foreground rvc run first completes idempotent unary RunCommand, then starts
FollowCommand from sequence zero with existing history included. Reconnects
resume after the last received client sequence; retention markers do not advance
that cursor and are emitted before a relevant terminal event. Do not add a
combined create-and-follow RPC whose stream can fail before returning the newly
created UUID.
For every mutation, canonicalize the validated protobuf request excluding
transport metadata, compute its immutable hash, and transact an idempotency row
keyed globally by request_id, with method included in the hashed/stored
identity. An exact replay returns the stored domain result
or stored error; a different hash returns ALREADY_EXISTS/request-conflict and
performs no work. Commit the idempotency row atomically with the mutation, retain
it at least as long as the referenced command/takeover/incident record, and
count its compressed storage against the owning command or audit scope. Reads
accept the field at the transport edge but deliberately drop it before domain
execution.
Define opaque output/page tokens as versioned, authenticated encodings of query filters plus stable position (command, event sequence, byte offset, snapshot boundary). Reject a token reused with changed filters. Pagination reads a consistent upper boundary so concurrent appends do not duplicate/skip retained bytes; if retention removes the resume point, return the next retained position with explicit server-retention metadata.
9.2 CLI
Implement rvc with no direct server database access. Required commands are:
rvc stat [client-id] [issue-uuid] [--all --page --per-page];rvc run [--background] [--cwd DIR] [--shell TYPE] [--env K=V]... [--profile FLAG]... [--request-id UUID] [--queue-ttl DURATION] CLIENT COMMAND;rvc run --script PATH [same execution options] CLIENT;rvc append CLIENT UUID TEXT,--file PATH, raw/no-newline mode, and--attachstdin streaming;rvc close-stdin CLIENT UUID;rvc kill [HUP|INT|TERM|KILL|USR1|USR2] CLIENT UUID(optionalSIGprefix; defaultTERM);rvc storage incidents,rvc storage repair INCIDENT, andrvc storage acknowledge INCIDENT --note TEXT;rvc client takeover CLIENT INSTANCE-ID;- output selection,
--timestamped, pagination, and--followsemantics.
Accept --request-id globally. Generate a UUIDv7 for every control mutation
unless supplied, reuse it across transport retries, and discard it for
read-only operations.
Create the request ID before dialing the Unix socket and retain it for all
automatic retries of that invocation. Validate user-supplied IDs as canonical
UUIDv7 strings. For run, map the same value directly to issue_uuid; for
other mutations keep it only as request_id. Do not generate a second
operation identifier. Print the issue UUID immediately after durable creation,
including before entering foreground follow mode.
Make stat render lifecycle and history independently: an expired queued
command is visibly marked expired with its expiry time/reason; if a late,
previously accepted client result later arrives, the actual terminal lifecycle
becomes current while the expiry contradiction remains in metadata/audit. Also
show dirty storage scope/incident ID and pending takeover instance,
output retained range, and whether terminal output is incomplete.
Render output without inventing line boundaries in stored data. Historical pagination uses opaque cursors and byte limits; the display layer may buffer partial lines for presentation and clearly prints retained range/truncation notices. Map structured control errors to stable non-zero exit codes while retaining machine-readable JSON output as a later optional CLI mode.
9.3 JSON-RPC adapter
Keep this adapter small and disabled by default. Use standard JSON-RPC 2.0 over
HTTP and protobuf JSON mapping for unary control request/response bodies.
Implement the documented lower-camel method names only, including storage
incident list/repair/acknowledge and authorizeClientTakeover. Bind default
127.0.0.1:6900; configuration may bind elsewhere but startup logs a conspicuous
unauthenticated-exposure warning. Do not add streaming or a second event model:
callers poll getOutput/getCommand by cursor.
Audit every control request with transport/source metadata where available, action, target, result, and error. Audit logs include environment key names but not override values or raw stdin/stdout/stderr.
Use the JSON-RPC envelope id only for JSON-RPC response correlation. Accept
the same optional protobuf requestId member as gRPC for domain idempotency;
when it is absent, generate a server UUIDv7. Consequently raw JSON-RPC remains
easy to use but only callers that persist/reuse requestId receive retry
idempotency. Map parse/invalid-request/method-not-found to standard JSON-RPC
codes and domain failures to one stable server-error code carrying the same
structured RVBox detail as gRPC.
9.4 Control-plane tests
Unit-test CLI parsing/rendering/exit-code mapping, global --request-id
generation/drop rules, protobuf-to-domain validation, gRPC/JSON-RPC error
mapping, cursor signing/filter binding, byte slicing, follow-resume behavior, and
redaction. Golden output must cover expiry contradictions, Windows attempted/
effective identities, truncation/incomplete output, dirty incidents, and pending
takeover without depending on terminal width or locale.
The control integration suite runs the compiled server and rvc against a
real mode-0600 Unix socket plus the real optional HTTP adapter and SQLite
store. Exercise every unary method, follow cancellation/reconnect, pagination
during concurrent appends/retention, identical/conflicting mutation retries,
JSON/protobuf size limits, disabled/default/non-loopback JSON-RPC behavior, and
server restart between mutation commit and response. Phase 5 E2E scenarios then
repeat the user-visible command paths against the real Windows client rather
than treating component integration as sufficient.
Exit criteria: rvc can drive each documented example against the Compose
stack; the Unix socket has mode 0600; JSON-RPC behavior matches gRPC unary
semantics and is off unless explicitly enabled; the control integration suite
can be interrupted, resumed, and reset through the shared run ID.
10. Deferred Linux client implementation (Phase 7; still required for full v1)
This phase is design-only in the current effort. Do not implement it now. Retain these tasks so the later Unix-client work completes the already defined v1 contract without weakening the shipped Windows behavior.
Begin this phase only after both an explicit later implementation decision and the Windows Phase 4/5 release gates. Keep Unix code in platform-specific files/ build tags and reuse the proven runtime/store/protocol contracts without changing their wire semantics to suit Linux.
The Unix supervisor starts an RVBox launcher as leader of a new session/process
group; after authorization the launcher creates the selected sh/bash child
inside the same group and remains its watchdog. It applies daemon environment
plus persisted overrides, validates the CWD, connects stdin/stdout/stderr, and
records launcher/root/group identities. On shutdown or unclean-start recovery,
terminate managed groups and emit interrupted rather than claiming pipe
monitoring survived.
Implement launch with a private release/watchdog channel. Persist and fsync
launcher PID/process-group/platform birth identity as launch_prepared; recheck
revision/cancellation, persist and fsync launch_authorized, then release user
code. Keep the channel open through tree cleanup. EOF before authorization exits
without execution; EOF after release terminates the group. An authorized but
uncertain launch is killed and interrupted, never retried.
On Linux create a per-command cgroup v2 for supervision whenever a delegated
writable cgroup exists, even without a resource profile. Put the blocked
launcher into it before release using clone3(CLONE_INTO_CGROUP | CLONE_PIDFD) where available; otherwise migrate only the still-blocked launcher
through cgroup.procs. Persist cgroup path, pidfd-derived identity where
available, PID/process group, /proc/<pid>/stat start time, and launch
generation. Recovery prefers cgroup.kill; a PID/process-group fallback is
allowed only after positive birth-identity verification.
Other Unix platforms use the same barrier plus watchdog/process-group fallback. Document that descendants deliberately creating a new session may escape that fallback. The at-most-once authorization rule still applies.
Treat root exit as the configured 5-second tree/output drain grace. Wait for
cgroup/process-group emptiness and capture EOF; terminate residual descendants,
drain, and sequence OutputIncomplete if EOF remains unprovable. Emit terminal
lifecycle last.
Implement portable HUP, INT, TERM, KILL, USR1, and USR2, accepting
optional SIG prefixes and mapping only through SignalKind. Reject native
numbers, signal zero, unsupported names, and arbitrary PID targets. Derive group
identity only from durable supervisor metadata.
Poll readable /proc data outside pipe/network loops for CPU time, RSS, I/O,
state, CWD, and wait reason, aggregating only clearly associated descendants.
Missing data remains absent. Define diagnostic progress as changes in tree CPU
ticks, cumulative I/O, retained input/output activity, or lifecycle. Clear
suspected_hung on progress and never let diagnostics alter lifecycle.
Compile profile policy before launch. A no-profile supervisory cgroup imposes
no limit. Profiles apply required cpu.max/cpu.weight,
memory.max/memory.high, pids.max, and per-device io.max values while the
launcher is blocked, then read them back. Without required delegation/control,
permanently reject as UNSUPPORTED; never run partly constrained.
Run native Linux integration/race/crash tests mirroring the shared Windows
suite: exact shell resolution, CWD/env, duplicate dispatch, output/stdin,
process-group/cgroup tree cleanup, launcher death at every durable boundary,
reconnect/replay, truncation, scripts, tombstones, queue limits, /proc
absence, profile failure, and shutdown interruption. Add explicit tests for
pidfd/clone3 availability fallbacks and deliberate setsid escape behavior.
Exit criteria: the Linux client passes the same protocol/durability suite as Windows, plus cgroup/process-group tests, without weakening the already shipped Windows behavior or changing the v1 wire contract.
11. Reliability, observability, and operational delivery (Phase 8)
11.1 Health, metrics, and logs
Expose separate liveness/readiness for the server. Liveness and incident inspection start before asynchronous recovery. Global readiness requires all required scopes recovered and writable, while scoped operations can proceed on healthy scopes; no health state requires a client connection. Client health reports reconnect state, spool health, unresolved incidents, and supervisor health without leaking command output.
Publish counters/histograms/gauges for registrations/takeovers, stale messages, heartbeat timeouts, reconnect duration, dispatch latency, command transitions, queue depth, spool bytes, output compression/rotation/loss, segment eviction, SQLite write latency/failure, script verification failure, and protocol errors. Use stable labels with bounded cardinality—never UUID/client ID as a metric label. Use structured logs and audit records for those identifiers instead.
11.2 Deployment assets
Provide:
- nginx configuration showing WebSocket upgrade proxying and TLS termination;
- systemd units for server/client with private state directories, restart policy, working directory, file descriptor limits, and least privilege;
- example server/client configuration files with every default and an explicit JSON-RPC exposure warning;
- Windows
rvbox.exeGUI/icon/version resources, SCM service installation, per-session tray registration, first-run TOML template, ACL setup, rotating- file-log support, code-signing hook, and SHA-256 manifest; - a backup/restore procedure for SQLite plus output/audit segment directories;
- an upgrade procedure that stops dispatch safely, snapshots data, migrates, and verifies recovery;
- a troubleshooting runbook for no client, stale session, spool full, output truncation, storage full, and daemon restart.
Install the annotated TOML examples from docs/examples/ and keep the most
operationally important identity/listener/state/TLS/quota keys first in each
section. Add --check-config to both daemons: it must run the production strict
decoder, defaulting, cross-field/profile/path/shell validation, print a redacted
normalized summary, and exit without opening stores/listeners or changing
state. Linux CI parses the server and all-knob client reference without starting
a client. Native Windows CI parses and normalizes client.windows.toml with the
production Windows path/shell/ACL checks.
Use Delegate=yes in the Linux client systemd unit when cgroup supervision is
enabled and create a private writable cgroup subtree for the service. Refuse a
configured mandatory resource profile when delegation/controllers are absent,
while still allowing unrestricted execution through the documented fallback.
Package upgrades must retain the previous TOML, run --check-config before
restart, and never silently ignore a newly unknown/removed key.
11.3 Security and regression review
Before v1 release, review path traversal, script temp-file permissions, command logging, environment override redaction, malformed compression, decompression bombs, oversized frames, Unicode/ASCII ID validation, Unix socket ownership, JSON-RPC external bind warnings, SQL injection (all parameterized), segment record corruption, process-group PID reuse, UAC relaunch argument handling, SCM service/Run-key executable and config quoting/ownership, ProgramData ACLs, tray file-opening paths, global single-instance behavior, and Windows inherited handles. Add fuzz tests for envelope decode, compressed output validation, segment-tail recovery, pagination tokens, and JSON-RPC parsing.
Exit criteria: release Compose/systemd/nginx assets exist; load/fault tests show no unbounded memory/goroutine growth; the operations runbook reproduces recovery and retention behavior.
12. Release gates and implementation order
The currently authorized merge order is deliberately vertical:
- Phase 0: Docker-first repository/toolchain/proto generation.
- Phase 1: domain/config validation and state machine.
- Phase 2: SQLite/segments/audit/retention with recovery tests.
- Phase 3: server registration, fencing, heartbeat, and persisted dispatch.
- Phase 4: shared client spool plus Windows supervisor/tray and at-most-once execution on native Windows CI.
- Phase 5: full WSS protocol and fault injection from native Windows to the Compose-hosted Linux server/nginx stack.
- Phase 6: gRPC Unix socket,
rvc, and optional JSON-RPC. - The Windows-applicable Phase 8 operational, packaging, stress, and release gates needed for the Linux-server/Windows-client milestone.
Stop there for the current effort. When Unix-client implementation is explicitly started later, continue with Phase 7 (Linux supervisor/cgroup and client parity), then the remaining Phase 8 Unix operational/release gates to complete v1.
Do not merge a later vertical slice by stubbing a durability/safety invariant. For example: foreground mode may wait on a durable background command, but must not bypass persistence; client output may be truncated under the documented caps, but must never block a child pipe; and a reconnection may replay work, but may never re-execute an already accepted UUID.
The current Windows-client milestone is releasable only after Phases 0–6 plus its applicable Phase 8 packaging/security gates pass on native Windows and a clean Linux server environment. Windows support cannot be marked optional or replaced by cross-compilation-only checks. Stop the current implementation effort at that milestone; Phase 7 remains deferred until explicitly started later. Full v1 is ready only after that later Linux client also passes the common protocol/ durability suite and Linux-specific cgroup/process tests. Every release must conspicuously document self-reported identity, the elevated Windows execution authority, and unauthenticated optional JSON-RPC.