3.2 KiB
RVBox Linux-server operations runbook
This runbook covers the v1 Linux server and native Windows clients. It does not describe a Unix-like client, which is outside v1.
Deploy and verify
Use the checked-in Compose deployment as the sole Linux-server runtime. Start
from the annotated server.toml and retain
json_rpc.enabled = false unless performing loopback-only debugging.
For Compose, follow the setup in
deploy/production/README.md, set a pinned
RVBOX_SERVER_IMAGE, RVBOX_TLS_CERT, and RVBOX_TLS_KEY, then run:
docker compose -f deploy/production/compose.yaml config
docker compose -f deploy/production/compose.yaml up -d
docker compose -f deploy/production/compose.yaml ps
The first command must succeed before any containers start. /livez shows that
the process is up; /readyz becomes successful only after durable recovery.
The public endpoint accepts only wss://HOST/v1/agent. Do not publish port
6900 or add a proxy route for JSON-RPC.
Backup and restore
Stop dispatch before copying data: stop the server gracefully, confirm it is
down, then copy the entire configured server.data_dir tree. That tree contains
SQLite, WAL/SHM state, retained command segments, audit segments, and incidents;
backing up SQLite alone is incomplete.
To restore, keep the server stopped, move the failed directory aside without
deleting it, restore the complete backup with ownership restricted to the
server account, and run rvbox-server --check-config followed by a normal
start. Keep readiness under observation. A failed recovery leaves readiness
false and records an incident; do not delete segments to force readiness.
Upgrade and rollback
- Record the running image/binary digest and
rvc statoutput. - Stop the server gracefully so no new dispatch is accepted.
- Take a full data-directory backup as above.
- Install the new image/binary without changing configuration, and run
--check-configbefore starting it. - Start, wait for
/readyz, and verifyrvc statplus one known client reconnect. - If readiness or storage recovery fails, stop, restore the prior binary/image and full data directory, then start the known-good version. Preserve logs and the failed copy for diagnosis.
Common incidents
- No client / stale session: verify nginx has WebSocket
101entries for/v1/agent, server readiness is true, and the Windows service is running. Uservc stat CLIENT_ID; do not restart the client merely to clear history. - Spool or storage full:
CAPACITY_EXHAUSTEDis intentional admission protection. Inspect command retention and free space, allow terminal-age or quota rotation to reclaim eligible data, or enlarge the owned filesystem. Never manually remove live SQLite/WAL/segment files. - Output truncated: query command status and output history for its explicit loss/truncation markers. The command may still have completed correctly.
- Server restart / dirty health: wait for readiness and inspect the incident record. Use the documented repair/acknowledgement controls only after preserving evidence; a late client report is authoritative and is not rewritten to match an earlier provisional state.