Tested, not promised.
High availability that is not tested is assumed broken. We cut the power of our own machines, restore our own backups, scan our own ports, and publish what we find, failures included.
Power cuts.
One machine at a time is switched off for about five minutes, as a power or hardware failure would. Meanwhile every service is polled every two seconds from outside, the way a person would use it. The table shows the latest run on each machine of our own environment (core) and of the demo environment.
| Machine | Tested | Powered off | Sign-in down | Longest other outage | Whole again after power on |
|---|---|---|---|---|---|
| core · node-aheld the database | 2 October 2026 | 5 min | 26 s | Monitoring 64 s | 340 s |
| core · node-b | 29 September 2026 | 5 min | 2 s | SIEM 97 s | 188 s |
| core · witness | 29 September 2026 | 5 min | none | none | 69 s |
| demo · node-aheld the database | 2 October 2026 | 30 min | 20 s | SIEM 31 s | 222 s |
| demo · node-bheld the database | 2 October 2026 | 4 min | 333 s | SIEM 56 s | 230 s |
| demo · witness | 29 September 2026 | 5 min | none | none | 90 s |
When a run finds something, we fix it and run it again. On 2 October 2026, powering off the database's machine kept sign-in down for the whole five minutes: the network's control plane waited on connections to the machine that was gone. With shorter network timeouts, sign-in came back within about 20 seconds in the next run. A machine whose latest run came before a fix shows that run until it is tested again.
Restores.
A backup that was never restored is a hope. The latest backup is restored into a throwaway database, its log replayed and checked, after every change that touches backups. Backup age and size are watched hourly, and a person is alerted when they look wrong.
Ports.
Each of our nodes scans every TCP port of the other machines, and of every customer's public address, every day. Any open port pages a person. The only thing that may answer from the internet is WireGuard, and only to a packet carrying a key it knows.
Builds.
xsenv-net is built reproducibly: the same release always gives the same bytes. Every file is signed with two keys, and the installers check the signature before they install anything. You can check it yourself:
curl -fsSLO https://xsenv.com/releases/signing-key.pub
v=$(curl -fsSL https://xsenv.com/releases/latest.txt)
curl -fsSLO https://xsenv.com/releases/$v/xsnet-linux-amd64
curl -fsSLO https://xsenv.com/releases/$v/xsnet-linux-amd64.sig
ssh-keygen -Y verify -f <(echo "xsnet-release $(cat signing-key.pub)") -I xsnet-release \
-n xsnet-release -s xsnet-linux-amd64.sig < xsnet-linux-amd64
Signing keys: SSH, P-256 for OpenSSL.
Leaving.
Each customer holds a recovery key we never see. With it alone they can open a handover package of their environment: its code, its credentials, the way to its backups. We rehearse that handover with the key and nothing else, and list whatever is still ours.