ChecklistFrontier R&D
A validity checklist for benchmarking BFT consensus on your own terms
Published throughput and finality figures for consensus protocols come from someone else's hardware, network and transaction mix. If your consortium will run on different ones, those figures are a starting hypothesis at best. This checklist sets out what a benchmark needs before its results can inform a choice: a workload model from your own traffic, realistic topology and latency, injected faults, tail-aware metrics, fair tuning of every candidate and complete reporting.
On this page
- Why published consensus benchmarks rarely transfer
- Workload model built from your own traffic
- Topology, network emulation and fault scenarios
- Metrics and run discipline for consensus tests
- Configuring each protocol family fairly
- Hypothetical example: a trade-document consortium
- What to publish alongside benchmark results
- Questions and answers
- Sources
Why published consensus benchmarks rarely transfer
Byzantine fault-tolerant protocols are sensitive to conditions that papers and vendors choose for reasons of their own: committee size, message sizes, latency between replicas, batching and whether anything fails during the run. PBFT tolerates f faulty replicas out of 3f+1 and exchanges messages among all replicas in its prepare and commit phases2, so its costs change sharply as committees grow. HotStuff was designed so that a correct leader drives agreement with communication linear in the number of replicas3, which moves the bottleneck to leader placement and signature handling. A benchmark run elsewhere cannot tell you which effect dominates in your deployment.
A consortium choosing a design it will operate for years needs measurements under its own load profile, one of the example questions behind our Frontier R&D practice1. The checklists below are the conditions any such benchmark, ours or anyone else's, should meet.
Workload model built from your own traffic
Topology, network emulation and fault scenarios
Linux traffic control with netem can add delay, jitter, loss, duplication, reordering and rate limits to each replica link4, which makes geographic conditions repeatable in a lab.
Metrics and run discipline for consensus tests
Configuring each protocol family fairly
A comparison is fair only when each candidate is configured as its maintainers would recommend for your conditions. These settings skew results most often.
| Protocol family | Settings that dominate results | Common fairness trap |
|---|---|---|
| PBFT family | Batch size, checkpoint interval and view-change timeouts | Testing small committees, where all-to-all messaging looks cheap |
| HotStuff family | Leader rotation, pacemaker timeouts and the signature aggregation scheme | Placing every leader in the best-connected region |
| CometBFT (Tendermint) | Propose and commit timeouts, block size limits and mempool settings | Leaving default timeouts unchanged, which can add deliberate delay between blocks |
| Hashgraph-style aBFT | Gossip frequency, event creation rate and node connectivity | Testing a network so small that gossip-based ordering looks trivially fast |
Take recommended settings from each candidate's maintainers or documentation, record every change from the defaults, and give every candidate the same tuning effort.
Hypothetical example: a trade-document consortium
What to publish alongside benchmark results
Questions and answers
Why not rely on the throughput figures vendors publish?
Vendor figures are measured under conditions chosen to show a protocol well: often one region, small transactions, a tuned committee and no faults. They can be honest for those conditions, but your consortium will have different geography, payloads and failures. Use published figures to choose which candidates to test, then measure the shortlist under your own workload.
How many consensus protocols should we benchmark?
Usually two or three, chosen after a desk review of governance, licensing, maturity and operating requirements. Each candidate needs careful, fair configuration, so a long list spreads tuning effort thin and weakens the comparison. A short list benchmarked thoroughly, with faults injected, is worth more than a long list tested under ideal conditions.
Is network emulation good enough, or do we need a real multi-region deployment?
Emulation is the right tool for comparison, because every candidate sees identical latency, jitter and loss and runs can be repeated. For the final candidate, a short confirmation run on real multi-region infrastructure is worth adding to catch what emulation misses, such as noisy neighbours and routing changes.
Should finality be measured at the protocol or at the client?
Both, but decide on client-observed finality, because that is what applications and users experience. Protocol-level timings explain where the time goes, for example in proposal, voting or gossip. Define finality precisely for each candidate, since probabilistic and deterministic finality are different guarantees and should not be compared as if they were the same.
Sources
- Frontier R&D: questions that justify research — ColdAI
- Practical Byzantine Fault Tolerance (Castro and Liskov, OSDI 1999) — MIT Laboratory for Computer Science · checked 10 October 2026
- HotStuff: BFT Consensus in the Lens of Blockchain (Yin, Malkhi, Reiter, Golan Gueta, Abraham) — arXiv · checked 10 October 2026
- tc-netem(8): Network Emulator — Debian manpages (iproute2) · checked 10 October 2026