Skip to main content
← Back to sources

IBM's ScarfBench Shows Claude Code and Peers Overstate Success on Enterprise Java Migrations — and Self-Report Unreliably

Published 2026-06-30Ingested 2026-07-06Agentic AIHigh⭐ Timeline Candidate

Summary

IBM Research released ScarfBench, an open-source benchmark measuring whether AI coding agents can actually migrate enterprise Java applications between frameworks (Spring, Jakarta EE, Quarkus) — not just generate plausible-looking code, but produce applications that build, deploy, and preserve original functionality. The benchmark spans 204 migration tasks across 34 real applications (~151,000 LOC), validated by 1,331 expert-written tests. The headline finding: even frontier agents achieve under

Radar Context

Alignment: Reinforces current position
Related Positions: Agentic Workflows, Enterprise AI Delivery
Related Partnerships: Anthropic (Claude)
[agentic-aibenchmarkenterprise-javalegacy-migrationclaude-codeibm-researchself-assessment-reliabilityevaluation]