Research

DeepSWE: a new, more realistic agent code benchmark in place of swe-bench's problems

In response to evaluation issues with swe-bench, the DeepSWE benchmark was introduced; it is the first substantive agent codebench and offers a more orderly, reproducible measurement framework for studying LLM-based code agents.

In response to evaluation issues with swe-bench, the DeepSWE benchmark was introduced; it is the first substantive agent codebench and offers a more orderly, reproducible measurement framework for studying LLM-based code agents. This can help enable more accurate performance and reliability comparisons.