Files
rvbox/docs/operations-runbook.md
T

69 lines
3.2 KiB
Markdown

# RVBox Linux-server operations runbook
This runbook covers the v1 Linux server and native Windows clients. It does not
describe a Unix-like client, which is outside v1.
## Deploy and verify
Use the checked-in Compose deployment as the sole Linux-server runtime. Start
from its Compose-specific
[`server.toml.example`](../deploy/production/server.toml.example) and retain
`json_rpc.enabled = false` unless performing loopback-only debugging.
For Compose, follow the setup in
[`deploy/production/README.md`](../deploy/production/README.md), set a pinned
`RVBOX_SERVER_IMAGE`, `RVBOX_TLS_CERT`, and `RVBOX_TLS_KEY`, then run:
```sh
docker compose -f deploy/production/compose.yaml config
docker compose -f deploy/production/compose.yaml up -d
docker compose -f deploy/production/compose.yaml ps
```
The first command must succeed before any containers start. `/livez` shows that
the process is up; `/readyz` becomes successful only after durable recovery.
The public endpoint accepts only `wss://HOST/v1/agent`. Do not publish port
6900 or add a proxy route for JSON-RPC.
## Backup and restore
Stop dispatch before copying data: stop the server gracefully, confirm it is
down, then copy the entire configured `server.data_dir` tree. That tree contains
SQLite, WAL/SHM state, retained command segments, audit segments, and incidents;
backing up SQLite alone is incomplete.
To restore, keep the server stopped, move the failed directory aside without
deleting it, restore the complete backup with ownership restricted to the
server account, and run `rvbox-server --check-config` followed by a normal
start. Keep readiness under observation. A failed recovery leaves readiness
false and records an incident; do not delete segments to force readiness.
## Upgrade and rollback
1. Record the running image/binary digest and `rvc stat` output.
2. Stop the server gracefully so no new dispatch is accepted.
3. Take a full data-directory backup as above.
4. Install the new image/binary without changing configuration, and run
`--check-config` before starting it.
5. Start, wait for `/readyz`, and verify `rvc stat` plus one known client
reconnect.
6. If readiness or storage recovery fails, stop, restore the prior binary/image
and full data directory, then start the known-good version. Preserve logs
and the failed copy for diagnosis.
## Common incidents
- **No client / stale session:** verify nginx has WebSocket `101` entries for
`/v1/agent`, server readiness is true, and the Windows service is running.
Use `rvc stat CLIENT_ID`; do not restart the client merely to clear history.
- **Spool or storage full:** `CAPACITY_EXHAUSTED` is intentional admission
protection. Inspect command retention and free space, allow terminal-age or
quota rotation to reclaim eligible data, or enlarge the owned filesystem.
Never manually remove live SQLite/WAL/segment files.
- **Output truncated:** query command status and output history for its explicit
loss/truncation markers. The command may still have completed correctly.
- **Server restart / dirty health:** wait for readiness and inspect the
incident record. Use the documented repair/acknowledgement controls only
after preserving evidence; a late client report is authoritative and is not
rewritten to match an earlier provisional state.