From c03a3edef2c3b08a0fa1135e648873bfbe0ca061 Mon Sep 17 00:00:00 2001 From: cabbage Date: Fri, 28 Aug 2026 17:29:37 +0000 Subject: [PATCH] docs: define reproducible v1 test strategy --- docs/implementation-plan.v1.md | 415 ++++++++++++++++++++++++++++++--- 1 file changed, 384 insertions(+), 31 deletions(-) diff --git a/docs/implementation-plan.v1.md b/docs/implementation-plan.v1.md index 4804244..ca23ce9 100644 --- a/docs/implementation-plan.v1.md +++ b/docs/implementation-plan.v1.md @@ -84,6 +84,266 @@ A phase is complete only when all of the following are true: 6. Any recovery or retention code is tested with a real temporary SQLite database and filesystem rather than mocks alone. +### 2.3 Test layers and required coverage + +Every test belongs to exactly one named layer. Do not call a test “unit” merely +because it is written with Go's `testing` package, and do not hide system setup +inside an ordinary package test. + +**Unit tests** are hermetic, parallel-safe, and require no RVBox daemon, project +runtime/sidecar container, network listener, wall-clock delay, administrator +privilege, or shared machine state. They execute inside the toolchain container. +Use injected clocks, deterministic randomness/readers, in-memory fakes, and +`t.TempDir()`. A normal unit invocation must be safe to repeat and must leave no +repository artifact. Add table, property, race, and fuzz coverage as applicable +for: + +- every legal and illegal lifecycle/revision transition, terminal predicate, + reconciliation row, error mapping, and UUIDv7/request-hash idempotency rule; +- strict TOML decoding/defaulting/cross-field validation, path and shell + resolution policy, resource-profile composition, redacted display, and all + hard size/count/duration boundaries; +- protobuf envelope/field validation, compression limits, event sequencing, + truncation/incomplete-output semantics, cursor encoding, and malformed input; +- quota charging, reservation, whole-command eviction choice, age rotation, + tombstone ordering, audit rotation, and filesystem-floor admission decisions; +- dispatch capacity reservations, fairness, heartbeat/backoff boundaries, + priority-lane scheduling, cancellation/signal races, and goroutine shutdown; +- client spool ordering, retransmission, acknowledgement, raw-output + high/low-watermark loss conversion, script chunk validation, and crash-state + decision logic; +- the pure Windows login-state/context-selection matrix, token-property + validation, command-line quoting, environment overlay, path/ACL construction, + named-pipe message validation, tray authorization decisions, and SCM desired- + state transitions. Win32 calls themselves belong in native integration tests. + +Unit tests use stable per-test seeds printed on failure. Timing assertions use +fake monotonic clocks; a small explicit set of real scheduler/deadline tests may +use bounded polling but never arbitrary sleeps. `go test -shuffle=on` and focused +`-race` jobs must pass repeatedly. A retry wrapper may collect flake evidence but +must never turn a failed assertion into a passing gate. + +**Integration tests** exercise one real subsystem boundary while keeping remote +peers controllable. Linux integration tests run in the project Compose stack and +use real SQLite/WAL, append-only segment files, Unix sockets, HTTP/WebSocket +listeners, compression, and compiled binaries where the boundary requires them. +Native Windows integration tests run the compiled client/service/helper modes +against a deterministic fake agent server and use real SCM, WTS/token APIs, +named pipes, Job Objects, consoles, files, and registry entries. Integration +tests may inject storage/network/process faults through documented harness +controls, but may not mock the component whose contract is under test. + +Required integration suites are `store`, `server-session`, `control`, +`client-spool`, `windows-supervisor`, `windows-service-tray`, and `upgrade- +recovery`. They cover at minimum: + +- all commit/crash boundaries, reopen/recovery, corrupt or short committed + ranges, disk-full/permission failures, repair/acknowledgement, quota/retention, + and concurrent readers/writers against real temporary storage; +- real gRPC/JSON-RPC/WebSocket serialization and limits, session fencing, + dispatch/reconciliation, lost acknowledgements, slow peers, priority control + traffic, pagination, and idempotent mutation replay; +- native Windows context selection for each supported login/token state, + pre-launch fallback only, launcher authentication, process-tree ownership, + stdout/stderr/stdin, signal escalation, shell/CWD/environment behavior, + service/tray authorization, boot/logon/logoff, and cleanup after forced death. + +**End-to-end tests** start the production-shaped Linux server, nginx TLS proxy, +local control plane, and the signed-or-CI-built native Windows service client. +They drive behavior only through `rvc`, supported service/tray controls, and +intentional fault-harness controls. They use real protobuf bytes over WSS and +real persistent stores; do not replace the Windows client with Wine, a Linux +client, or a protocol stub. Required named scenarios are: + +1. `smoke`: install/start, registration, background and foreground commands, + stdout/stderr/history, terminal status, and orderly uninstall preserving data; +2. `interactive`: stdin append/raw/close, TERM then KILL behavior, scripts, CWD, + environment overrides, and output pagination/follow; +3. `idempotency`: repeated `request_id`/`issue_uuid`, lost control responses, + duplicate dispatch/events, and conflicting request hashes; +4. `reconnect`: nginx/server/network interruption, old-session fencing, client + restart, reconciliation, and no duplicate execution; +5. `retention`: command/client/global quotas, offline spool loss markers, + terminal-age reclamation, tombstones, audit rotation, and disk floors; +6. `expiry-and-incidents`: offline 15-minute expiry, late client truth, corrupt + storage, scoped dirty health, repair, and acknowledgement; +7. `windows-contexts`: logged-out normal/elevated execution, active standard and + split-token administrator sessions, ordered elevated fallbacks, ambiguous + active sessions, and effective-identity status/audit output; +8. `service-tray`: boot before login, per-session tray behavior, UAC-protected + configuration actions, Explorer restart, service continuity, and log rotation; +9. `soak`: bounded concurrency, slow/unstable network, verbose output, repeated + reconnects, and resource-leak checks for the release gate. + +Each layer has an explicit timeout budget. Unit and integration suites must be +shardable by stable suite/case names. E2E scenarios are serial on a particular +Windows host because they mutate machine-wide SCM/Run-key state; parallelism is +allowed only across independently leased resettable VMs. + +### 2.4 Reproducible and resumable test harness + +Implement the following checked-in entry points; CI calls the same scripts that +developers call rather than duplicating orchestration in workflow YAML: + +```text +scripts/test-unit [--package PATTERN] [--run REGEXP] [--race] +scripts/test-integration [--suite NAME|all] [--run-id ID] [--resume] +scripts/test-e2e [--scenario NAME|all] [--run-id ID] [--resume] +scripts/test-env doctor|status|logs|collect|recover|reuse|stop|reset|purge|gc ... +scripts/windows/test-host.ps1 Prepare|Status|Run|Collect|Stop|Reset +``` + +Scripts are thin, reviewed orchestration wrappers. `scripts/test-unit` invokes +the pinned toolchain Compose service as UID/GID `1001:1001`. Integration and E2E +commands call a shared Go harness under `test/harness` so manifest parsing, +timeouts, process control, reporting, and cleanup logic are not reimplemented in +shell and PowerShell. The harness itself is built in the toolchain container; +no Go/Buf/protoc/SQLite SDK or package manager is installed on the host. + +The first integration/E2E invocation creates a filesystem-safe random run ID, +or validates an explicitly supplied one, and records +`.test-runs//run.toml`. The manifest includes test layer/suite/scenario, +seed, repository commit and dirty-diff hash, tool/runtime image digests, +configuration/certificate hashes, Compose project name, allocated ports, +Windows VM/runner identity, phase state, and exact owned resource names. Store a +fsynced append-only step journal plus bounded artifacts beneath that directory. +Never record credentials, private keys, tokens, command environment values, or +unredacted payloads in the manifest/log bundle. + +Integration binaries expose stable `--list` and `--case` names; the harness runs +one case process at a time and journals its setup, fault, reopen, invariant, and +cleanup checkpoints. E2E scenarios expose the same model at a larger step +granularity. SIGINT/SIGTERM/PowerShell cancellation handlers journal the +interruption and perform the bounded stop path. Emit both a human summary and +machine-readable JSON/JUnit results, preserving the original test failure even +when collection or cleanup also fails. + +Use `COMPOSE_PROJECT_NAME=rvbox-test-` and label every container, network, +and run-specific volume with the run ID and repository identity. Avoid fixed +host ports; allocate loopback ports once, persist them in the manifest, and +revalidate ownership before reuse. Hold a repository-local integration lock and +an exclusive Windows-host lease so two invocations cannot corrupt the same SCM, +registry, ProgramData, or port state. A per-run private CA/certificate is created +inside the exact run directory with restrictive permissions and destroyed by +that run's purge operation. + +Pin base/runtime/fault-service images by digest and record the fully rendered +Compose configuration hash. `scripts/test-env doctor` is a read-only cold-start +preflight for Docker/Compose compatibility, UID/GID ownership, available CPU/ +memory/disk, required cached-or-fetchable images, loopback port allocation, and +optional Windows-runner reachability/capability. It prints the exact missing +prerequisite and never installs host software or pulls an image unless the user +then runs a test command whose declared pull policy permits it. + +Every setup and scenario step is named, bounded, and idempotent. After completing +a step, persist its input hash, outputs needed by later steps, and success marker. +`--resume` verifies the commit/diff hash, image digests, configuration, VM +identity, owned-resource labels, and last successful health checkpoint before +continuing. It refuses unsafe resume on mismatch and tells the operator to start +a new run; `--resume` never silently resets state. Fault injection uses persisted +named checkpoints rather than timing guesses, so a deliberate server/client kill +can be resumed and its post-restart invariant checks completed. + +`recover RUN_ID` is distinct from resume: after verifying that the recorded +controller process/lease is dead, it reacquires the lock, reconciles the fsynced +journal with labelled Compose and Windows resources, completes or rolls back an +interrupted idempotent setup/cleanup step, and leaves the run in a declared +`stopped-resumable` or `reset` state. It never executes the next test assertion. +`reuse RUN_ID` creates a new run ID from the prior immutable image/configuration/ +fixture definitions and reusable dependency caches, but allocates fresh ports, +PKI, SQLite/segment/spool storage, command IDs, and Windows test root. It never +copies prior mutable test state. + +The Linux controller exposes bounded fault controls for frame drop/delay, +listener interruption, process termination at named durability checkpoints, +filesystem quota/error simulation, and monotonic/wall-clock advancement in the +deterministic peer. Native Windows `test-host.ps1` validates administrator state, +OS build, UAC/policy fixture, active sessions, service identity, and test-root +ownership before running. It uses a test-specific ProgramData root and an +exclusive host lease; production-name SCM/Run-key installation tests run only in +a resettable VM snapshot. Passwords and VM access credentials come from the CI +secret store or interactive prompt, never arguments persisted in the manifest. + +The E2E controller owns the canonical run manifest on Linux and passes the same +run ID plus a one-run configuration bundle to the preconfigured Windows runner. +The runner returns a checksummed step report and bounded artifacts through the +CI transport or explicitly configured management channel; its credentials +remain outside RVBox. Before executing a scenario, prove that Windows can reach +the manifest's nginx WSS endpoint and trusts only the run's configured CA file, +and that the Linux controller can observe service readiness. A lost runner is a +stopped-resumable run, not permission to provision another Windows machine +silently. + +On success, integration/E2E runs collect a compact report and automatically stop +and remove their exact run-specific containers, networks, volumes, temporary +certificates, Windows service/Run entries, and scratch state. Reusable dependency +caches and shared immutable images remain. On failure or interruption, the +harness first collects bounded diagnostics, stops expensive processes/containers, +and preserves only the manifest plus state required for `status`, `logs`, or +`--resume`; keeping services running requires an explicit `--keep-running`. + +`scripts/test-env` has these exact safety semantics: + +- `doctor [--e2e]` is read-only and validates local/container prerequisites and, + when requested, the native Windows runner before allocating a run; +- `status RUN_ID` is read-only and reports steps, health, disk usage, owned + Compose/Windows resources, and whether resume is valid; +- `logs RUN_ID [COMPONENT]` reads bounded/tail output; `collect RUN_ID` writes a + compressed, redacted diagnostic bundle without altering the environment; +- `recover RUN_ID` safely adopts an abandoned owned environment and reaches a + stable resumable or reset state; `reuse RUN_ID` creates a fresh isolated run + from the same immutable environment definition; +- `stop RUN_ID` stops only resources recorded in the validated manifest and + preserves resumable state; +- `reset RUN_ID` removes that run's runtime resources and scratch data but keeps + its compact report/manifest; the run is no longer resumable; +- `purge RUN_ID` first prints the exact resolved targets and requires confirmation + (or CI-only `--yes`), then removes that one run's remaining artifacts; +- `gc --older-than DURATION` is dry-run by default, considers only validated + completed/non-resumable RVBox manifests, and requires `--execute`; dependency + caches require a separate `--include-caches` confirmation. + +For deliberate low-disk cleanup, `scripts/test-env purge --all` enumerates every +validated RVBox run manifest and matching labelled resource, prints per-run and +total bytes, and does nothing until `--execute` plus confirmation. Reusable +dependency caches and test-only images are excluded unless separately requested +with `--include-caches` and `--include-images`; each inclusion is printed and +confirmed. This is the supported way to clear all RVBox test environments and +artifacts without performing a machine-wide Docker cleanup. + +No script may invoke a global Docker builder/image/container/volume/system prune, +delete by an unvalidated glob, remove a repository/workspace root, or touch an +unlabelled resource. Resolve every deletion target, require it to be beneath +`.test-runs` or match both the persisted Compose project and RVBox labels, and +print it before deletion. Test cleanup must itself have unit tests plus an +integration test proving that unrelated containers, volumes, files, Windows +services, and registry entries survive reset/purge. + +Keep the environment suitable for the resource-constrained host: build each +immutable image once per source digest; use named Go module/build caches without +copying dependency trees into the worktree; allow only one local integration or +E2E run by default; set Compose CPU/memory/PID limits and suite deadlines; rotate +component logs; cap collected logs/core/dumps; and run a disk-space preflight. +Report exact cache/run/artifact usage and actionable cleanup commands whenever a +preflight fails. Do not automatically delete caches or another run to make room. + +Put overridable harness defaults in `test/harness/defaults.toml`, not scattered +environment variables. Initial local defaults are Go package parallelism 2; one +integration/E2E run; two aggregate runtime CPUs; 2 GiB aggregate runtime memory; +512 aggregate PIDs; 10 MiB collected log tail per component; 100 MiB total +failure bundle and 20 MiB success report; core/minidumps disabled unless a named +debug run enables them; 10-minute unit, 30-minute integration-suite, 45-minute +ordinary E2E-scenario, and 2-hour soak timeouts; and a 2 GiB free-space preflight +floor in addition to product-specific disk-floor cases. CI may explicitly raise +these budgets, but every run records and prints its effective values. Hitting an +artifact cap writes an explicit truncation manifest rather than silently +discarding diagnostics. + +All scripts support `--help`, `--dry-run` for mutating environment operations, +stable nonzero exit codes, and a final summary giving the run ID, seed, report +path, resume command when applicable, and exact cleanup command. Document the +same workflow in `docs/testing.md` during Phase 0. + ## 3. Repository and build bootstrap (Phase 0) ### 3.1 Establish the repository layout @@ -117,10 +377,23 @@ internal/ gen/go/rvbox/v1/ # generated protobuf/grpc code protos/rvbox/v1/ # existing wire authority docs/examples/ # fully annotated server/client TOML +scripts/ + test-unit # hermetic toolchain-container entry point + test-integration # resumable component-suite entry point + test-e2e # resumable production-shaped scenario entry point + test-env # doctor/recover/reuse/cleanup by exact run ID + windows/test-host.ps1 # native Windows host lifecycle adapter +test/ + harness/ # shared manifest, journal, orchestration, and reporting + defaults.toml # local resource/time/artifact budgets + integration/ # cross-package/real-resource suite definitions + e2e/scenarios/ # named black-box scenario definitions + fixtures/ # small deterministic non-secret fixtures deploy/ Dockerfile.toolchain Dockerfile.runtime compose.yaml + compose.test.yaml # isolated integration/E2E/fault-injection profiles nginx/ systemd/ ``` @@ -148,11 +421,22 @@ Add: - `deploy/compose.yaml`: a `toolchain` service with the repository mounted at `/workspace`, run as `1001:1001`; optional `server`, `client`, `nginx`, and fault-injection test services use isolated named volumes. +- `deploy/compose.test.yaml`: test-only profile overrides, health checks, + resource limits, run-ID labels, ephemeral PKI, fault proxy, and artifact/state + mounts. It accepts only values produced by the validated harness manifest and + never uses an implicit default project name. - `Makefile`: `generate`, `fmt`, `lint`, `test`, `test-race`, `test-integration`, - `build`, and `verify`. Each target invokes the container workflow; it must not - silently fall back to host tools. + `test-e2e`, `test-status`, `build`, and `verify`. `test` aliases the unit + layer; integration/E2E targets delegate to the checked-in scripts and print + their run ID. Each target invokes the container workflow and must not silently + fall back to host tools. - `buf.gen.yaml`: Go protobuf and gRPC generation paths under `gen/go`. +Ignore `.test-runs/` and any local harness lock file in Git while retaining a +tracked `.test-runs/README` only if an in-tree explanation is useful. The +harness creates the directory with host UID/GID `1001:1001`; test binaries never +write generated dependencies into the repository tree. + Generation is deterministic: `make generate` followed by `git diff --exit-code` must be clean in CI. Update `buf.yaml`/`buf.gen.yaml` only with a matching generation run. @@ -187,8 +471,28 @@ or a service-account runner. Add a Server Core service/Session-0 smoke lane. Exercise at least the oldest supported Windows baseline and one current desktop release before publishing. +Use one declared test matrix and the same harness entry points everywhere: + +- every change runs format/lint/generation, all unit tests, selected race tests, + Linux `store`/`server-session`/`control` integration suites, Windows cross- + compilation, and the native noninteractive Windows smoke integration suite; +- scheduled/nightly runs execute the full race/fuzz budgets, all Linux and + native Windows integration suites, plus `smoke`, `idempotency`, `reconnect`, + and `retention` E2E scenarios; +- release candidates run every E2E scenario, including the resettable + interactive Windows VM matrix, Server Core smoke, destructive fault cases, + upgrade/recovery, and soak. + +CI allocates a run ID per job, always invokes `collect` after failure, and invokes +`reset` in an unconditional finalizer. Upload only the bounded redacted report, +manifest, and relevant logs; never upload full command payload stores by +default. A failed cleanup is a visible job failure with its exact manual cleanup +command, not a warning hidden behind the original test result. + **Exit criteria:** all three empty `main` packages build in the toolchain -container; generated code is checked in; `make verify` works from a fresh clone. +container; generated code is checked in; `make verify` works from a fresh clone; +the test scripts can create, interrupt, inspect, resume, collect, reset, and purge +one sample run without affecting a labelled unrelated fixture. ## 4. Shared contracts, configuration, and state machine (Phase 1) @@ -595,16 +899,24 @@ and retention; offline repair additionally requires the process-wide instance lock. Persist the repair/acknowledgement mutation and audit record before clearing derived dirty health. -### 5.6 Store tests +### 5.6 Store unit and integration tests -Test real SQLite/filesystem cases: duplicate event idempotence; conflicting -duplicate rejection; crash between segment write and metadata transaction; -crash after metadata commit; corrupted final record; command-window rotation; -per-client/global whole-command eviction; active-command protection; cursor -stability; uncommitted-tail truncation; short/corrupt committed ranges; dirty -derivation and repair/acknowledgement; and disk-full/permission failure -simulation. Verify recovery endpoints become live promptly while unsafe scoped -mutations remain gated. +Unit-test record framing/checksums, size arithmetic, quota reservations, +eviction ordering, age/tombstone/audit rotation decisions, cursor boundaries, +incident-state transitions, and fault-point state-machine outcomes without +opening SQLite. Run these with shuffled order, deterministic seeds, and race +coverage for the in-process ownership/maintenance-lock coordinators. + +The `store` integration suite uses a real temporary SQLite database, WAL, and +segment filesystem for: duplicate event idempotence; conflicting duplicate +rejection; crash between segment write and metadata transaction; crash after +metadata commit; corrupted final record; command-window rotation; per-client/ +global whole-command eviction; active-command protection; cursor stability; +uncommitted-tail truncation; short/corrupt committed ranges; dirty derivation +and repair/acknowledgement; and bounded disk-full/permission failure simulation. +Every case records a named fault checkpoint in the shared run journal and can be +resumed at its reopen/invariant-verification step. Verify recovery endpoints +become live promptly while unsafe scoped mutations remain gated. **Exit criteria:** a store restart preserves acknowledged data and never acknowledges uncommitted data; all rolling, tiered unified-quota, 30-day @@ -789,15 +1101,24 @@ was not authorized; otherwise retain the revisioned signal result and actual terminal outcome. The server maps control `request_id` to this revision/result and never sends that request ID across the agent protocol. -### 6.5 Server session tests +### 6.5 Server session unit and integration tests -Cover version negotiation, same-instance replacement, different-instance -pending claim, exact one-shot takeover consumption/expiry, stale-generation -late output, Ping/Pong inactivity thresholds, full client queue, dispatch retry -after lost acceptance, complete active/terminal-unacknowledged reconciliation, -lost/repeated `ReconcileResult`, permanent/transient dispatch rejection, -cancellation before launch, launch/cancel race, malformed envelope close, and a -slow client that cannot block another client or local control RPC. +Unit-test version selection, registry/generation transitions, takeover grant +matching, capacity-shadow accounting, dispatch fairness, TTL/revision races, +reconciliation matrix decisions, heartbeat/backoff boundaries, queue saturation, +and writer-lane scheduling with fake clocks, stores, and sockets. + +The `server-session` integration suite runs the compiled server with real +SQLite/filesystem state and real WebSocket protobuf frames. Cover same-instance +replacement, different-instance pending claim, exact one-shot takeover +consumption/expiry, stale-generation late output, Ping/Pong inactivity +thresholds, full client queue, dispatch retry after lost acceptance, complete +active/terminal-unacknowledged reconciliation, lost/repeated `ReconcileResult`, +permanent/transient dispatch rejection, cancellation before launch, launch/ +cancel race, malformed-envelope close, server restart, and a slow client that +cannot block another client or local control RPC. Use the harness fault proxy +and named acknowledgement checkpoints rather than sleeps or probabilistic packet +loss. **Exit criteria:** multiple simulated clients can register, replace each other, receive bounded dispatches, and recover connection faults without duplicate @@ -1223,14 +1544,23 @@ stderr. Flush warnings/errors promptly, redact as elsewhere, and make `Open log` target the current file after rotation. Add native icon/version/service metadata; release signing and checksum publication are Phase 8 gates. -### 7.6 Windows-first client tests +### 7.6 Windows-first client unit and native integration tests -Use fake transport/clock tests for shared runtime logic and real Windows -subprocess tests for both shells, CWD/env overlays, at-most-once duplicates, +Unit-test shared runtime/spool behavior with fake transport/clock/storage fault +points. Unit-test the exhaustive login/elevation selection function with token- +property DTOs rather than Win32 handles, and separately test Windows quoting, +environment, ACL descriptor, named-pipe frame, SCM transition, and tray- +authorization helpers. Cover at-most-once duplicates, reconnect/replay, offline +truncation, script validation, tombstone rotation, queue limits, and shutdown +decisions. Run race tests with multiple commands and forced logical network +churn. + +The native `windows-supervisor` and `windows-service-tray` integration suites use +real Windows subprocesses and APIs for both shells, CWD/environment overlays, concurrent output without newlines, stdin ordering/close, Job tree termination, -reconnect/replay, offline truncation, script verification/cleanup, tombstone -rotation, queue limits, and shutdown interruption. Run race tests with multiple -commands and forced network churn. +script materialization/cleanup, service shutdown interruption, and every token/ +session context. They run through `scripts/windows/test-host.ps1` under the +exclusive host lease and journal every machine-wide mutation for reset. Add a table-driven crash suite for every acceptance, script-upload, launch, event-spool, terminal-acknowledgement, and tombstone-rotation commit point. Run @@ -1297,9 +1627,12 @@ script transfer, server restart before/after event transaction, client restart with managed process cleanup, and offline spool overflow. Verify command UUIDs never execute twice and output gaps always appear as truncation metadata. -Build a deterministic scenario harness with fake wall/monotonic clocks, -controllable frame loss/reordering, daemon kill points, and bounded virtual disk -capacity. For each flow, run disconnect before delivery, after delivery/before +Use the shared resumable harness from Section 2.4 with fake wall/monotonic +clocks in its deterministic peer, controllable frame loss/reordering, daemon +kill points, and bounded virtual disk capacity. `scripts/test-e2e --scenario +NAME` runs one case; its printed `--run-id` can be passed back with `--resume` +after an intentional or accidental interruption. For each flow, run disconnect +before delivery, after delivery/before durable commit, after commit/before acknowledgement, and after acknowledgement. Assert database invariants and quota counters after every restart, not merely the visible CLI result. Include these cross-cutting cases: @@ -1321,7 +1654,8 @@ the visible CLI result. Include these cross-cutting cases: **Exit criteria:** a native Windows client against the Compose-hosted Linux server demonstrates a complete background command, history/follow behavior, stdin interaction, signal, reconnect, and restart recovery through nginx WSS. -This is the first supported-client milestone and gates Linux supervisor work. +This is the first supported-client milestone. It satisfies a prerequisite for +the deferred Linux supervisor work but does not automatically authorize it. ## 9. Server control plane and `rvc` (Phase 6) @@ -1449,9 +1783,28 @@ idempotency. Map parse/invalid-request/method-not-found to standard JSON-RPC codes and domain failures to one stable server-error code carrying the same structured RVBox detail as gRPC. +### 9.4 Control-plane tests + +Unit-test CLI parsing/rendering/exit-code mapping, global `--request-id` +generation/drop rules, protobuf-to-domain validation, gRPC/JSON-RPC error +mapping, cursor signing/filter binding, byte slicing, follow-resume behavior, and +redaction. Golden output must cover expiry contradictions, Windows attempted/ +effective identities, truncation/incomplete output, dirty incidents, and pending +takeover without depending on terminal width or locale. + +The `control` integration suite runs the compiled server and `rvc` against a +real mode-`0600` Unix socket plus the real optional HTTP adapter and SQLite +store. Exercise every unary method, follow cancellation/reconnect, pagination +during concurrent appends/retention, identical/conflicting mutation retries, +JSON/protobuf size limits, disabled/default/non-loopback JSON-RPC behavior, and +server restart between mutation commit and response. Phase 5 E2E scenarios then +repeat the user-visible command paths against the real Windows client rather +than treating component integration as sufficient. + **Exit criteria:** `rvc` can drive each documented example against the Compose stack; the Unix socket has mode `0600`; JSON-RPC behavior matches gRPC unary -semantics and is off unless explicitly enabled. +semantics and is off unless explicitly enabled; the control integration suite +can be interrupted, resumed, and reset through the shared run ID. ## 10. Deferred Linux client implementation (Phase 7; still required for full v1)