Files
rvbox/docs/operations-runbook.md
T

3.5 KiB

RVBox Linux-server operations runbook

This runbook covers the v1 Linux server and native Windows clients. It does not describe a Unix-like client, which is outside v1.

Deploy and verify

Use either the checked-in Compose deployment or the systemd unit, never both for the same server data directory. Start from the annotated server.toml and retain json_rpc.enabled = false unless performing loopback-only debugging.

For Compose, follow the setup in deploy/production/README.md, set a pinned RVBOX_SERVER_IMAGE, RVBOX_TLS_CERT, and RVBOX_TLS_KEY, then run:

docker compose -f deploy/production/compose.yaml config
docker compose -f deploy/production/compose.yaml up -d
docker compose -f deploy/production/compose.yaml ps

The first command must succeed before any containers start. /livez shows that the process is up; /readyz becomes successful only after durable recovery. The public endpoint accepts only wss://HOST/v1/agent. Do not publish port 6900 or add a proxy route for JSON-RPC.

For systemd, create the rvbox service account, install the binary and deploy/systemd/rvbox-server.service, place a root:rvbox owned 0640 /etc/rvbox/server.toml, then run systemctl daemon-reload and systemctl enable --now rvbox-server. The unit performs --check-config before every start.

Backup and restore

Stop dispatch before copying data: stop the server gracefully, confirm it is down, then copy the entire configured server.data_dir tree. That tree contains SQLite, WAL/SHM state, retained command segments, audit segments, and incidents; backing up SQLite alone is incomplete.

To restore, keep the server stopped, move the failed directory aside without deleting it, restore the complete backup with ownership restricted to the server account, and run rvbox-server --check-config followed by a normal start. Keep readiness under observation. A failed recovery leaves readiness false and records an incident; do not delete segments to force readiness.

Upgrade and rollback

  1. Record the running image/binary digest and rvc stat output.
  2. Stop the server gracefully so no new dispatch is accepted.
  3. Take a full data-directory backup as above.
  4. Install the new image/binary without changing configuration, and run --check-config before starting it.
  5. Start, wait for /readyz, and verify rvc stat plus one known client reconnect.
  6. If readiness or storage recovery fails, stop, restore the prior binary/image and full data directory, then start the known-good version. Preserve logs and the failed copy for diagnosis.

Common incidents

  • No client / stale session: verify nginx has WebSocket 101 entries for /v1/agent, server readiness is true, and the Windows service is running. Use rvc stat CLIENT_ID; do not restart the client merely to clear history.
  • Spool or storage full: CAPACITY_EXHAUSTED is intentional admission protection. Inspect command retention and free space, allow terminal-age or quota rotation to reclaim eligible data, or enlarge the owned filesystem. Never manually remove live SQLite/WAL/segment files.
  • Output truncated: query command status and output history for its explicit loss/truncation markers. The command may still have completed correctly.
  • Server restart / dirty health: wait for readiness and inspect the incident record. Use the documented repair/acknowledgement controls only after preserving evidence; a late client report is authoritative and is not rewritten to match an earlier provisional state.