68 lines
3.2 KiB
Markdown
68 lines
3.2 KiB
Markdown
# RVBox Linux-server operations runbook
|
|
|
|
This runbook covers the v1 Linux server and native Windows clients. It does not
|
|
describe a Unix-like client, which is outside v1.
|
|
|
|
## Deploy and verify
|
|
|
|
Use the checked-in Compose deployment as the sole Linux-server runtime. Start
|
|
from the annotated [`server.toml`](examples/server.toml) and retain
|
|
`json_rpc.enabled = false` unless performing loopback-only debugging.
|
|
|
|
For Compose, follow the setup in
|
|
[`deploy/production/README.md`](../deploy/production/README.md), set a pinned
|
|
`RVBOX_SERVER_IMAGE`, `RVBOX_TLS_CERT`, and `RVBOX_TLS_KEY`, then run:
|
|
|
|
```sh
|
|
docker compose -f deploy/production/compose.yaml config
|
|
docker compose -f deploy/production/compose.yaml up -d
|
|
docker compose -f deploy/production/compose.yaml ps
|
|
```
|
|
|
|
The first command must succeed before any containers start. `/livez` shows that
|
|
the process is up; `/readyz` becomes successful only after durable recovery.
|
|
The public endpoint accepts only `wss://HOST/v1/agent`. Do not publish port
|
|
6900 or add a proxy route for JSON-RPC.
|
|
|
|
## Backup and restore
|
|
|
|
Stop dispatch before copying data: stop the server gracefully, confirm it is
|
|
down, then copy the entire configured `server.data_dir` tree. That tree contains
|
|
SQLite, WAL/SHM state, retained command segments, audit segments, and incidents;
|
|
backing up SQLite alone is incomplete.
|
|
|
|
To restore, keep the server stopped, move the failed directory aside without
|
|
deleting it, restore the complete backup with ownership restricted to the
|
|
server account, and run `rvbox-server --check-config` followed by a normal
|
|
start. Keep readiness under observation. A failed recovery leaves readiness
|
|
false and records an incident; do not delete segments to force readiness.
|
|
|
|
## Upgrade and rollback
|
|
|
|
1. Record the running image/binary digest and `rvc stat` output.
|
|
2. Stop the server gracefully so no new dispatch is accepted.
|
|
3. Take a full data-directory backup as above.
|
|
4. Install the new image/binary without changing configuration, and run
|
|
`--check-config` before starting it.
|
|
5. Start, wait for `/readyz`, and verify `rvc stat` plus one known client
|
|
reconnect.
|
|
6. If readiness or storage recovery fails, stop, restore the prior binary/image
|
|
and full data directory, then start the known-good version. Preserve logs
|
|
and the failed copy for diagnosis.
|
|
|
|
## Common incidents
|
|
|
|
- **No client / stale session:** verify nginx has WebSocket `101` entries for
|
|
`/v1/agent`, server readiness is true, and the Windows service is running.
|
|
Use `rvc stat CLIENT_ID`; do not restart the client merely to clear history.
|
|
- **Spool or storage full:** `CAPACITY_EXHAUSTED` is intentional admission
|
|
protection. Inspect command retention and free space, allow terminal-age or
|
|
quota rotation to reclaim eligible data, or enlarge the owned filesystem.
|
|
Never manually remove live SQLite/WAL/segment files.
|
|
- **Output truncated:** query command status and output history for its explicit
|
|
loss/truncation markers. The command may still have completed correctly.
|
|
- **Server restart / dirty health:** wait for readiness and inspect the
|
|
incident record. Use the documented repair/acknowledgement controls only
|
|
after preserving evidence; a late client report is authoritative and is not
|
|
rewritten to match an earlier provisional state.
|