BranchBench
ReleasedMeasuring whether branchable databases can sustain the branching patterns created by exploratory AI agents.
Background
Exploratory agents create, mutate, compare, and discard many speculative states. BranchBench asks whether today’s branchable databases can handle that workload shape efficiently and reliably.
What it evaluates
- Branch creation, switching, connection, and deletion
- Deep and wide speculative branch structures
- Reads, joins, mutations, and schema changes on branch-local state
- Pruning behavior and concurrent exploration
Dataset and task taxonomy
The benchmark includes five ready-to-run workflows with configurable branch depth, fanout, mutation type, schemas, and query templates. New workflows and database backends can be added through protobuf configuration.
Evaluation methodology
Workloads report operation-level and end-to-end performance while checking whether each backend completes the requested workflow correctly. Parameters can be scaled to locate failure and timeout boundaries.
Representative task
Create a configurable tree of speculative database branches, execute branch-local analytic and mutation workloads, prune losing branches, and report latency and successful completion.
Results
The paper and accompanying blog discuss results across systems including Neon and Dolt.