ScarfBench
Benchmark for Agentic Enterprise Java Migration
ScarfBench (Self-Contained Application Refactoring) is a benchmark suite for evaluating AI agents on enterprise Java application migration across frameworks. It covers 102 real applications spanning Jakarta EE, Quarkus, and Spring — from focused layer-specific demos to full production-grade apps — with 1,331 tests across 6 architectural layers. All implementations are manually converted and developer-verified, providing a rigorous testbed for agentic code transformation.
Contributors
Rahul Krishna, Bridget McGinn, Raju Pavuluri