docs: define reproducible v1 test strategy

This commit is contained in:
2026-08-28 17:29:37 +00:00
parent 4b18bfc147
commit c03a3edef2
+384 -31
View File
@@ -84,6 +84,266 @@ A phase is complete only when all of the following are true:
6. Any recovery or retention code is tested with a real temporary SQLite
database and filesystem rather than mocks alone.
### 2.3 Test layers and required coverage
Every test belongs to exactly one named layer. Do not call a test “unit” merely
because it is written with Go's `testing` package, and do not hide system setup
inside an ordinary package test.
**Unit tests** are hermetic, parallel-safe, and require no RVBox daemon, project
runtime/sidecar container, network listener, wall-clock delay, administrator
privilege, or shared machine state. They execute inside the toolchain container.
Use injected clocks, deterministic randomness/readers, in-memory fakes, and
`t.TempDir()`. A normal unit invocation must be safe to repeat and must leave no
repository artifact. Add table, property, race, and fuzz coverage as applicable
for:
- every legal and illegal lifecycle/revision transition, terminal predicate,
reconciliation row, error mapping, and UUIDv7/request-hash idempotency rule;
- strict TOML decoding/defaulting/cross-field validation, path and shell
resolution policy, resource-profile composition, redacted display, and all
hard size/count/duration boundaries;
- protobuf envelope/field validation, compression limits, event sequencing,
truncation/incomplete-output semantics, cursor encoding, and malformed input;
- quota charging, reservation, whole-command eviction choice, age rotation,
tombstone ordering, audit rotation, and filesystem-floor admission decisions;
- dispatch capacity reservations, fairness, heartbeat/backoff boundaries,
priority-lane scheduling, cancellation/signal races, and goroutine shutdown;
- client spool ordering, retransmission, acknowledgement, raw-output
high/low-watermark loss conversion, script chunk validation, and crash-state
decision logic;
- the pure Windows login-state/context-selection matrix, token-property
validation, command-line quoting, environment overlay, path/ACL construction,
named-pipe message validation, tray authorization decisions, and SCM desired-
state transitions. Win32 calls themselves belong in native integration tests.
Unit tests use stable per-test seeds printed on failure. Timing assertions use
fake monotonic clocks; a small explicit set of real scheduler/deadline tests may
use bounded polling but never arbitrary sleeps. `go test -shuffle=on` and focused
`-race` jobs must pass repeatedly. A retry wrapper may collect flake evidence but
must never turn a failed assertion into a passing gate.
**Integration tests** exercise one real subsystem boundary while keeping remote
peers controllable. Linux integration tests run in the project Compose stack and
use real SQLite/WAL, append-only segment files, Unix sockets, HTTP/WebSocket
listeners, compression, and compiled binaries where the boundary requires them.
Native Windows integration tests run the compiled client/service/helper modes
against a deterministic fake agent server and use real SCM, WTS/token APIs,
named pipes, Job Objects, consoles, files, and registry entries. Integration
tests may inject storage/network/process faults through documented harness
controls, but may not mock the component whose contract is under test.
Required integration suites are `store`, `server-session`, `control`,
`client-spool`, `windows-supervisor`, `windows-service-tray`, and `upgrade-
recovery`. They cover at minimum:
- all commit/crash boundaries, reopen/recovery, corrupt or short committed
ranges, disk-full/permission failures, repair/acknowledgement, quota/retention,
and concurrent readers/writers against real temporary storage;
- real gRPC/JSON-RPC/WebSocket serialization and limits, session fencing,
dispatch/reconciliation, lost acknowledgements, slow peers, priority control
traffic, pagination, and idempotent mutation replay;
- native Windows context selection for each supported login/token state,
pre-launch fallback only, launcher authentication, process-tree ownership,
stdout/stderr/stdin, signal escalation, shell/CWD/environment behavior,
service/tray authorization, boot/logon/logoff, and cleanup after forced death.
**End-to-end tests** start the production-shaped Linux server, nginx TLS proxy,
local control plane, and the signed-or-CI-built native Windows service client.
They drive behavior only through `rvc`, supported service/tray controls, and
intentional fault-harness controls. They use real protobuf bytes over WSS and
real persistent stores; do not replace the Windows client with Wine, a Linux
client, or a protocol stub. Required named scenarios are:
1. `smoke`: install/start, registration, background and foreground commands,
stdout/stderr/history, terminal status, and orderly uninstall preserving data;
2. `interactive`: stdin append/raw/close, TERM then KILL behavior, scripts, CWD,
environment overrides, and output pagination/follow;
3. `idempotency`: repeated `request_id`/`issue_uuid`, lost control responses,
duplicate dispatch/events, and conflicting request hashes;
4. `reconnect`: nginx/server/network interruption, old-session fencing, client
restart, reconciliation, and no duplicate execution;
5. `retention`: command/client/global quotas, offline spool loss markers,
terminal-age reclamation, tombstones, audit rotation, and disk floors;
6. `expiry-and-incidents`: offline 15-minute expiry, late client truth, corrupt
storage, scoped dirty health, repair, and acknowledgement;
7. `windows-contexts`: logged-out normal/elevated execution, active standard and
split-token administrator sessions, ordered elevated fallbacks, ambiguous
active sessions, and effective-identity status/audit output;
8. `service-tray`: boot before login, per-session tray behavior, UAC-protected
configuration actions, Explorer restart, service continuity, and log rotation;
9. `soak`: bounded concurrency, slow/unstable network, verbose output, repeated
reconnects, and resource-leak checks for the release gate.
Each layer has an explicit timeout budget. Unit and integration suites must be
shardable by stable suite/case names. E2E scenarios are serial on a particular
Windows host because they mutate machine-wide SCM/Run-key state; parallelism is
allowed only across independently leased resettable VMs.
### 2.4 Reproducible and resumable test harness
Implement the following checked-in entry points; CI calls the same scripts that
developers call rather than duplicating orchestration in workflow YAML:
```text
scripts/test-unit [--package PATTERN] [--run REGEXP] [--race]
scripts/test-integration [--suite NAME|all] [--run-id ID] [--resume]
scripts/test-e2e [--scenario NAME|all] [--run-id ID] [--resume]
scripts/test-env doctor|status|logs|collect|recover|reuse|stop|reset|purge|gc ...
scripts/windows/test-host.ps1 Prepare|Status|Run|Collect|Stop|Reset
```
Scripts are thin, reviewed orchestration wrappers. `scripts/test-unit` invokes
the pinned toolchain Compose service as UID/GID `1001:1001`. Integration and E2E
commands call a shared Go harness under `test/harness` so manifest parsing,
timeouts, process control, reporting, and cleanup logic are not reimplemented in
shell and PowerShell. The harness itself is built in the toolchain container;
no Go/Buf/protoc/SQLite SDK or package manager is installed on the host.
The first integration/E2E invocation creates a filesystem-safe random run ID,
or validates an explicitly supplied one, and records
`.test-runs/<run-id>/run.toml`. The manifest includes test layer/suite/scenario,
seed, repository commit and dirty-diff hash, tool/runtime image digests,
configuration/certificate hashes, Compose project name, allocated ports,
Windows VM/runner identity, phase state, and exact owned resource names. Store a
fsynced append-only step journal plus bounded artifacts beneath that directory.
Never record credentials, private keys, tokens, command environment values, or
unredacted payloads in the manifest/log bundle.
Integration binaries expose stable `--list` and `--case` names; the harness runs
one case process at a time and journals its setup, fault, reopen, invariant, and
cleanup checkpoints. E2E scenarios expose the same model at a larger step
granularity. SIGINT/SIGTERM/PowerShell cancellation handlers journal the
interruption and perform the bounded stop path. Emit both a human summary and
machine-readable JSON/JUnit results, preserving the original test failure even
when collection or cleanup also fails.
Use `COMPOSE_PROJECT_NAME=rvbox-test-<run-id>` and label every container, network,
and run-specific volume with the run ID and repository identity. Avoid fixed
host ports; allocate loopback ports once, persist them in the manifest, and
revalidate ownership before reuse. Hold a repository-local integration lock and
an exclusive Windows-host lease so two invocations cannot corrupt the same SCM,
registry, ProgramData, or port state. A per-run private CA/certificate is created
inside the exact run directory with restrictive permissions and destroyed by
that run's purge operation.
Pin base/runtime/fault-service images by digest and record the fully rendered
Compose configuration hash. `scripts/test-env doctor` is a read-only cold-start
preflight for Docker/Compose compatibility, UID/GID ownership, available CPU/
memory/disk, required cached-or-fetchable images, loopback port allocation, and
optional Windows-runner reachability/capability. It prints the exact missing
prerequisite and never installs host software or pulls an image unless the user
then runs a test command whose declared pull policy permits it.
Every setup and scenario step is named, bounded, and idempotent. After completing
a step, persist its input hash, outputs needed by later steps, and success marker.
`--resume` verifies the commit/diff hash, image digests, configuration, VM
identity, owned-resource labels, and last successful health checkpoint before
continuing. It refuses unsafe resume on mismatch and tells the operator to start
a new run; `--resume` never silently resets state. Fault injection uses persisted
named checkpoints rather than timing guesses, so a deliberate server/client kill
can be resumed and its post-restart invariant checks completed.
`recover RUN_ID` is distinct from resume: after verifying that the recorded
controller process/lease is dead, it reacquires the lock, reconciles the fsynced
journal with labelled Compose and Windows resources, completes or rolls back an
interrupted idempotent setup/cleanup step, and leaves the run in a declared
`stopped-resumable` or `reset` state. It never executes the next test assertion.
`reuse RUN_ID` creates a new run ID from the prior immutable image/configuration/
fixture definitions and reusable dependency caches, but allocates fresh ports,
PKI, SQLite/segment/spool storage, command IDs, and Windows test root. It never
copies prior mutable test state.
The Linux controller exposes bounded fault controls for frame drop/delay,
listener interruption, process termination at named durability checkpoints,
filesystem quota/error simulation, and monotonic/wall-clock advancement in the
deterministic peer. Native Windows `test-host.ps1` validates administrator state,
OS build, UAC/policy fixture, active sessions, service identity, and test-root
ownership before running. It uses a test-specific ProgramData root and an
exclusive host lease; production-name SCM/Run-key installation tests run only in
a resettable VM snapshot. Passwords and VM access credentials come from the CI
secret store or interactive prompt, never arguments persisted in the manifest.
The E2E controller owns the canonical run manifest on Linux and passes the same
run ID plus a one-run configuration bundle to the preconfigured Windows runner.
The runner returns a checksummed step report and bounded artifacts through the
CI transport or explicitly configured management channel; its credentials
remain outside RVBox. Before executing a scenario, prove that Windows can reach
the manifest's nginx WSS endpoint and trusts only the run's configured CA file,
and that the Linux controller can observe service readiness. A lost runner is a
stopped-resumable run, not permission to provision another Windows machine
silently.
On success, integration/E2E runs collect a compact report and automatically stop
and remove their exact run-specific containers, networks, volumes, temporary
certificates, Windows service/Run entries, and scratch state. Reusable dependency
caches and shared immutable images remain. On failure or interruption, the
harness first collects bounded diagnostics, stops expensive processes/containers,
and preserves only the manifest plus state required for `status`, `logs`, or
`--resume`; keeping services running requires an explicit `--keep-running`.
`scripts/test-env` has these exact safety semantics:
- `doctor [--e2e]` is read-only and validates local/container prerequisites and,
when requested, the native Windows runner before allocating a run;
- `status RUN_ID` is read-only and reports steps, health, disk usage, owned
Compose/Windows resources, and whether resume is valid;
- `logs RUN_ID [COMPONENT]` reads bounded/tail output; `collect RUN_ID` writes a
compressed, redacted diagnostic bundle without altering the environment;
- `recover RUN_ID` safely adopts an abandoned owned environment and reaches a
stable resumable or reset state; `reuse RUN_ID` creates a fresh isolated run
from the same immutable environment definition;
- `stop RUN_ID` stops only resources recorded in the validated manifest and
preserves resumable state;
- `reset RUN_ID` removes that run's runtime resources and scratch data but keeps
its compact report/manifest; the run is no longer resumable;
- `purge RUN_ID` first prints the exact resolved targets and requires confirmation
(or CI-only `--yes`), then removes that one run's remaining artifacts;
- `gc --older-than DURATION` is dry-run by default, considers only validated
completed/non-resumable RVBox manifests, and requires `--execute`; dependency
caches require a separate `--include-caches` confirmation.
For deliberate low-disk cleanup, `scripts/test-env purge --all` enumerates every
validated RVBox run manifest and matching labelled resource, prints per-run and
total bytes, and does nothing until `--execute` plus confirmation. Reusable
dependency caches and test-only images are excluded unless separately requested
with `--include-caches` and `--include-images`; each inclusion is printed and
confirmed. This is the supported way to clear all RVBox test environments and
artifacts without performing a machine-wide Docker cleanup.
No script may invoke a global Docker builder/image/container/volume/system prune,
delete by an unvalidated glob, remove a repository/workspace root, or touch an
unlabelled resource. Resolve every deletion target, require it to be beneath
`.test-runs` or match both the persisted Compose project and RVBox labels, and
print it before deletion. Test cleanup must itself have unit tests plus an
integration test proving that unrelated containers, volumes, files, Windows
services, and registry entries survive reset/purge.
Keep the environment suitable for the resource-constrained host: build each
immutable image once per source digest; use named Go module/build caches without
copying dependency trees into the worktree; allow only one local integration or
E2E run by default; set Compose CPU/memory/PID limits and suite deadlines; rotate
component logs; cap collected logs/core/dumps; and run a disk-space preflight.
Report exact cache/run/artifact usage and actionable cleanup commands whenever a
preflight fails. Do not automatically delete caches or another run to make room.
Put overridable harness defaults in `test/harness/defaults.toml`, not scattered
environment variables. Initial local defaults are Go package parallelism 2; one
integration/E2E run; two aggregate runtime CPUs; 2 GiB aggregate runtime memory;
512 aggregate PIDs; 10 MiB collected log tail per component; 100 MiB total
failure bundle and 20 MiB success report; core/minidumps disabled unless a named
debug run enables them; 10-minute unit, 30-minute integration-suite, 45-minute
ordinary E2E-scenario, and 2-hour soak timeouts; and a 2 GiB free-space preflight
floor in addition to product-specific disk-floor cases. CI may explicitly raise
these budgets, but every run records and prints its effective values. Hitting an
artifact cap writes an explicit truncation manifest rather than silently
discarding diagnostics.
All scripts support `--help`, `--dry-run` for mutating environment operations,
stable nonzero exit codes, and a final summary giving the run ID, seed, report
path, resume command when applicable, and exact cleanup command. Document the
same workflow in `docs/testing.md` during Phase 0.
## 3. Repository and build bootstrap (Phase 0)
### 3.1 Establish the repository layout
@@ -117,10 +377,23 @@ internal/
gen/go/rvbox/v1/ # generated protobuf/grpc code
protos/rvbox/v1/ # existing wire authority
docs/examples/ # fully annotated server/client TOML
scripts/
test-unit # hermetic toolchain-container entry point
test-integration # resumable component-suite entry point
test-e2e # resumable production-shaped scenario entry point
test-env # doctor/recover/reuse/cleanup by exact run ID
windows/test-host.ps1 # native Windows host lifecycle adapter
test/
harness/ # shared manifest, journal, orchestration, and reporting
defaults.toml # local resource/time/artifact budgets
integration/ # cross-package/real-resource suite definitions
e2e/scenarios/ # named black-box scenario definitions
fixtures/ # small deterministic non-secret fixtures
deploy/
Dockerfile.toolchain
Dockerfile.runtime
compose.yaml
compose.test.yaml # isolated integration/E2E/fault-injection profiles
nginx/
systemd/
```
@@ -148,11 +421,22 @@ Add:
- `deploy/compose.yaml`: a `toolchain` service with the repository mounted at
`/workspace`, run as `1001:1001`; optional `server`, `client`, `nginx`, and
fault-injection test services use isolated named volumes.
- `deploy/compose.test.yaml`: test-only profile overrides, health checks,
resource limits, run-ID labels, ephemeral PKI, fault proxy, and artifact/state
mounts. It accepts only values produced by the validated harness manifest and
never uses an implicit default project name.
- `Makefile`: `generate`, `fmt`, `lint`, `test`, `test-race`, `test-integration`,
`build`, and `verify`. Each target invokes the container workflow; it must not
silently fall back to host tools.
`test-e2e`, `test-status`, `build`, and `verify`. `test` aliases the unit
layer; integration/E2E targets delegate to the checked-in scripts and print
their run ID. Each target invokes the container workflow and must not silently
fall back to host tools.
- `buf.gen.yaml`: Go protobuf and gRPC generation paths under `gen/go`.
Ignore `.test-runs/` and any local harness lock file in Git while retaining a
tracked `.test-runs/README` only if an in-tree explanation is useful. The
harness creates the directory with host UID/GID `1001:1001`; test binaries never
write generated dependencies into the repository tree.
Generation is deterministic: `make generate` followed by `git diff --exit-code`
must be clean in CI. Update `buf.yaml`/`buf.gen.yaml` only with a matching
generation run.
@@ -187,8 +471,28 @@ or a service-account runner. Add a Server Core service/Session-0 smoke lane.
Exercise at least the oldest supported Windows baseline and one current desktop
release before publishing.
Use one declared test matrix and the same harness entry points everywhere:
- every change runs format/lint/generation, all unit tests, selected race tests,
Linux `store`/`server-session`/`control` integration suites, Windows cross-
compilation, and the native noninteractive Windows smoke integration suite;
- scheduled/nightly runs execute the full race/fuzz budgets, all Linux and
native Windows integration suites, plus `smoke`, `idempotency`, `reconnect`,
and `retention` E2E scenarios;
- release candidates run every E2E scenario, including the resettable
interactive Windows VM matrix, Server Core smoke, destructive fault cases,
upgrade/recovery, and soak.
CI allocates a run ID per job, always invokes `collect` after failure, and invokes
`reset` in an unconditional finalizer. Upload only the bounded redacted report,
manifest, and relevant logs; never upload full command payload stores by
default. A failed cleanup is a visible job failure with its exact manual cleanup
command, not a warning hidden behind the original test result.
**Exit criteria:** all three empty `main` packages build in the toolchain
container; generated code is checked in; `make verify` works from a fresh clone.
container; generated code is checked in; `make verify` works from a fresh clone;
the test scripts can create, interrupt, inspect, resume, collect, reset, and purge
one sample run without affecting a labelled unrelated fixture.
## 4. Shared contracts, configuration, and state machine (Phase 1)
@@ -595,16 +899,24 @@ and retention; offline repair additionally requires the process-wide instance
lock. Persist the repair/acknowledgement mutation and audit record before clearing
derived dirty health.
### 5.6 Store tests
### 5.6 Store unit and integration tests
Test real SQLite/filesystem cases: duplicate event idempotence; conflicting
duplicate rejection; crash between segment write and metadata transaction;
crash after metadata commit; corrupted final record; command-window rotation;
per-client/global whole-command eviction; active-command protection; cursor
stability; uncommitted-tail truncation; short/corrupt committed ranges; dirty
derivation and repair/acknowledgement; and disk-full/permission failure
simulation. Verify recovery endpoints become live promptly while unsafe scoped
mutations remain gated.
Unit-test record framing/checksums, size arithmetic, quota reservations,
eviction ordering, age/tombstone/audit rotation decisions, cursor boundaries,
incident-state transitions, and fault-point state-machine outcomes without
opening SQLite. Run these with shuffled order, deterministic seeds, and race
coverage for the in-process ownership/maintenance-lock coordinators.
The `store` integration suite uses a real temporary SQLite database, WAL, and
segment filesystem for: duplicate event idempotence; conflicting duplicate
rejection; crash between segment write and metadata transaction; crash after
metadata commit; corrupted final record; command-window rotation; per-client/
global whole-command eviction; active-command protection; cursor stability;
uncommitted-tail truncation; short/corrupt committed ranges; dirty derivation
and repair/acknowledgement; and bounded disk-full/permission failure simulation.
Every case records a named fault checkpoint in the shared run journal and can be
resumed at its reopen/invariant-verification step. Verify recovery endpoints
become live promptly while unsafe scoped mutations remain gated.
**Exit criteria:** a store restart preserves acknowledged data and never
acknowledges uncommitted data; all rolling, tiered unified-quota, 30-day
@@ -789,15 +1101,24 @@ was not authorized; otherwise retain the revisioned signal result and actual
terminal outcome. The server maps control `request_id` to this revision/result
and never sends that request ID across the agent protocol.
### 6.5 Server session tests
### 6.5 Server session unit and integration tests
Cover version negotiation, same-instance replacement, different-instance
pending claim, exact one-shot takeover consumption/expiry, stale-generation
late output, Ping/Pong inactivity thresholds, full client queue, dispatch retry
after lost acceptance, complete active/terminal-unacknowledged reconciliation,
lost/repeated `ReconcileResult`, permanent/transient dispatch rejection,
cancellation before launch, launch/cancel race, malformed envelope close, and a
slow client that cannot block another client or local control RPC.
Unit-test version selection, registry/generation transitions, takeover grant
matching, capacity-shadow accounting, dispatch fairness, TTL/revision races,
reconciliation matrix decisions, heartbeat/backoff boundaries, queue saturation,
and writer-lane scheduling with fake clocks, stores, and sockets.
The `server-session` integration suite runs the compiled server with real
SQLite/filesystem state and real WebSocket protobuf frames. Cover same-instance
replacement, different-instance pending claim, exact one-shot takeover
consumption/expiry, stale-generation late output, Ping/Pong inactivity
thresholds, full client queue, dispatch retry after lost acceptance, complete
active/terminal-unacknowledged reconciliation, lost/repeated `ReconcileResult`,
permanent/transient dispatch rejection, cancellation before launch, launch/
cancel race, malformed-envelope close, server restart, and a slow client that
cannot block another client or local control RPC. Use the harness fault proxy
and named acknowledgement checkpoints rather than sleeps or probabilistic packet
loss.
**Exit criteria:** multiple simulated clients can register, replace each other,
receive bounded dispatches, and recover connection faults without duplicate
@@ -1223,14 +1544,23 @@ stderr. Flush warnings/errors promptly, redact as elsewhere, and make `Open log`
target the current file after rotation. Add native icon/version/service metadata;
release signing and checksum publication are Phase 8 gates.
### 7.6 Windows-first client tests
### 7.6 Windows-first client unit and native integration tests
Use fake transport/clock tests for shared runtime logic and real Windows
subprocess tests for both shells, CWD/env overlays, at-most-once duplicates,
Unit-test shared runtime/spool behavior with fake transport/clock/storage fault
points. Unit-test the exhaustive login/elevation selection function with token-
property DTOs rather than Win32 handles, and separately test Windows quoting,
environment, ACL descriptor, named-pipe frame, SCM transition, and tray-
authorization helpers. Cover at-most-once duplicates, reconnect/replay, offline
truncation, script validation, tombstone rotation, queue limits, and shutdown
decisions. Run race tests with multiple commands and forced logical network
churn.
The native `windows-supervisor` and `windows-service-tray` integration suites use
real Windows subprocesses and APIs for both shells, CWD/environment overlays,
concurrent output without newlines, stdin ordering/close, Job tree termination,
reconnect/replay, offline truncation, script verification/cleanup, tombstone
rotation, queue limits, and shutdown interruption. Run race tests with multiple
commands and forced network churn.
script materialization/cleanup, service shutdown interruption, and every token/
session context. They run through `scripts/windows/test-host.ps1` under the
exclusive host lease and journal every machine-wide mutation for reset.
Add a table-driven crash suite for every acceptance, script-upload, launch,
event-spool, terminal-acknowledgement, and tombstone-rotation commit point. Run
@@ -1297,9 +1627,12 @@ script transfer, server restart before/after event transaction, client restart
with managed process cleanup, and offline spool overflow. Verify command UUIDs
never execute twice and output gaps always appear as truncation metadata.
Build a deterministic scenario harness with fake wall/monotonic clocks,
controllable frame loss/reordering, daemon kill points, and bounded virtual disk
capacity. For each flow, run disconnect before delivery, after delivery/before
Use the shared resumable harness from Section 2.4 with fake wall/monotonic
clocks in its deterministic peer, controllable frame loss/reordering, daemon
kill points, and bounded virtual disk capacity. `scripts/test-e2e --scenario
NAME` runs one case; its printed `--run-id` can be passed back with `--resume`
after an intentional or accidental interruption. For each flow, run disconnect
before delivery, after delivery/before
durable commit, after commit/before acknowledgement, and after acknowledgement.
Assert database invariants and quota counters after every restart, not merely
the visible CLI result. Include these cross-cutting cases:
@@ -1321,7 +1654,8 @@ the visible CLI result. Include these cross-cutting cases:
**Exit criteria:** a native Windows client against the Compose-hosted Linux
server demonstrates a complete background command, history/follow behavior,
stdin interaction, signal, reconnect, and restart recovery through nginx WSS.
This is the first supported-client milestone and gates Linux supervisor work.
This is the first supported-client milestone. It satisfies a prerequisite for
the deferred Linux supervisor work but does not automatically authorize it.
## 9. Server control plane and `rvc` (Phase 6)
@@ -1449,9 +1783,28 @@ idempotency. Map parse/invalid-request/method-not-found to standard JSON-RPC
codes and domain failures to one stable server-error code carrying the same
structured RVBox detail as gRPC.
### 9.4 Control-plane tests
Unit-test CLI parsing/rendering/exit-code mapping, global `--request-id`
generation/drop rules, protobuf-to-domain validation, gRPC/JSON-RPC error
mapping, cursor signing/filter binding, byte slicing, follow-resume behavior, and
redaction. Golden output must cover expiry contradictions, Windows attempted/
effective identities, truncation/incomplete output, dirty incidents, and pending
takeover without depending on terminal width or locale.
The `control` integration suite runs the compiled server and `rvc` against a
real mode-`0600` Unix socket plus the real optional HTTP adapter and SQLite
store. Exercise every unary method, follow cancellation/reconnect, pagination
during concurrent appends/retention, identical/conflicting mutation retries,
JSON/protobuf size limits, disabled/default/non-loopback JSON-RPC behavior, and
server restart between mutation commit and response. Phase 5 E2E scenarios then
repeat the user-visible command paths against the real Windows client rather
than treating component integration as sufficient.
**Exit criteria:** `rvc` can drive each documented example against the Compose
stack; the Unix socket has mode `0600`; JSON-RPC behavior matches gRPC unary
semantics and is off unless explicitly enabled.
semantics and is off unless explicitly enabled; the control integration suite
can be interrupted, resumed, and reset through the shared run ID.
## 10. Deferred Linux client implementation (Phase 7; still required for full v1)