Reliability
Speed is the smaller half of a payment benchmark. Each run here generates load, most of them break something on purpose, and all of them end with a verifier that checks the outcome against what every client actually asked for.
A run passes only if all four of these hold.
Each request a client made is replayed under its original key and must return what the client saw. Requests that never got an answer are replayed until they have one.
Each wallet must hold exactly what those outcomes sum to, with nothing left reserved.
The ledger's own invariant check must find nothing.
The day is settled, every shop must be paid exactly its net, and the ledger, the settlement records and the bank must match.
The five runs in which no request failed, including three with a dependency taken away. Median and 99th-percentile payment latency, in milliseconds.
"Requests failed" counts attempts that got no response or a server error; clients retry them under the same key. "Unresolved at the end" counts intents whose client ran out of retries, which the verifier then resolves to a single outcome each.
| Run | What is done to it | Result | Rate | Requests failed | p50 / p99 ms | Unresolved at the end |
|---|---|---|---|---|---|---|
| Steady | Mixed traffic at a constant rate | PASS | 149.9/s | 0 | 44 / 63 | 0 |
| Peak | Twice the steady rate, past saturation | PASS | 178.7/s | 92.7% | 355 / 2,923 | 4,092 |
| Hot merchant | Half of all payments go to one shop | PASS | 149.9/s | 0 | 56 / 173 | 0 |
| Retry storm | Clients give up after 40 ms and retry | PASS | 99.9/s | 36.7% | 4 / 39 | 0 |
| Ledger killed | The ledger is killed twice under load | PASS | 99.9/s | 38.4% | 252 / 1,727 | 4,005 |
| Payment service killed | Killed twice under load | PASS | 82.2/s | 49.5% | 54 / 2,816 | 1,080 |
| Merchant service killed | Down for 15 s; payments continue from cache | PASS | 99.9/s | 0 | 47 / 366 | 0 |
| Kafka frozen | For 25 s; events must arrive afterwards | PASS | 99.8/s | 0 | 48 / 80 | 0 |
| DynamoDB frozen | For 12 s; money-moving requests are refused | PASS | 100.0/s | 60.9% | 48 / 2,378 | 773 |
| Redis frozen | For 12 s; QR payments stop, the rest continues | PASS | 92.3/s | 0 | 44 / 1,411 | 0 |
| Bank frozen | For 12 s; top-ups wait, then complete once | PASS | 100.0/s | 0 | 50 / 191 | 0 |
| Database frozen | TiDB for 8 s; requests stall and resolve | PASS | 99.9/s | 7.0% | 46 / 2,327 | 636 |
| Everything | Six faults in 90 s: two kills of the ledger, one of the payment service, Kafka, the bank and Redis frozen | PASS | 92.4/s | 40.6% | 248 / 2,634 | 5,026 |
A check that has never failed proves nothing. So a bug was planted: the payment service was rebuilt to ignore idempotency keys, and the retry storm was run against it.
Retried top-ups and transfers were applied more than once, and the verifier caught every one.
Each was paid something other than its true net.
Every duplicate was a perfectly balanced, perfectly reconciled movement of money that nobody had asked for. Only comparing outcomes with intents finds that kind of fault.
These are single-machine numbers. Every service and every store ran in containers on one modest computer, one instance of each. They are useful relative to each other, not as a forecast of production throughput. With a single instance, a killed service is simply gone until it restarts, so the failure drills measure correctness under failure and not availability. Past roughly 150 to 300 requests a second on that machine most requests time out, and it stays correct.
Each run's full report is written by the tool that ran it, and all thirteen can be reproduced in about twenty minutes.