Real-SWE benchmark exposes weak AI coding agents
The top result in the Real-SWE benchmark resolved 38.8% of private-code tasks. Six of the ten analyzed tasks recorded success rates below 15%, including one task with zero successful runs across 64 attempts.