In response to evaluation issues with swe-bench, the DeepSWE benchmark was introduced; it is the first substantive agent codebench and offers a more orderly, reproducible measurement framework for studying LLM-based code agents. This can help enable more accurate performance and reliability comparisons.
DeepSWE: a new, more realistic agent code benchmark in place of swe-bench's problems
In response to evaluation issues with swe-bench, the DeepSWE benchmark was introduced; it is the first substantive agent codebench and offers a more orderly, reproducible measurement framework for studying LLM-based code agents.


