How many bugs can LMs find & fix in large codebases?
Given a real repository, an agent must discover & repair as many bugs as they can.
Agents are not given any hint about the type of bug or its location.
- Website: https://swesweep.com
- Paper: https://swesweep.com/paper
Note
The code in this repo is only a thin wrapper around Harbor, however see the warning below regarding harbor version
Warning
- Harbor version: Harbor 0.23 does not yet support the separate verifier environments and collect hooks used by these tasks. The project therefore pins a compatible revision from Harbor's source repository until Harbor 0.24 is released.
- Final score: We build a micro-average of the bugs resolved (equivalent to a weighted average of the individual task scores). See evaluation notes.
We recommend uv for managing Python environments.
git clone https://github.com/facebookresearch/swe-sweep.git
cd swe-sweep
uv sync # or pip install .Verify your setup:
sweep infra doctorDevelopment setup
Clone the repository and install the editable package with its dev dependencies:
git clone https://github.com/facebookresearch/swe-sweep.git
cd swe-sweep
uv sync --extra testRun the test suite (matches CI, which tests on Python 3.12 and 3.13):
uv run pytestRun the linting/formatting hooks:
uvx pre-commit run --all-filesPass a task name and a unified diff against that task's base commit:
sweep eval <task> <path/to/submission.diff> # see raw harbor command belowRaw harbor command
uv run harbor run \
--path tasks/pandora-bench__dateutil \
--agent swesweep.patch_agent:PatchAgent \
--agent-kwarg "patch_path=$(realpath /path/to/model.patch)" \
--jobs-dir jobs \
--n-concurrent 1 \
--yes
The command runs the corresponding Harbor task with a small patch-applying agent. Harbor
then evaluates the resulting checkout in the task's separate verifier environment. Results
are written under jobs/ by default.
How evaluation works
Evaluation follows the following pseudo-code:
reset_to_base_commit()
apply(agent_patch)
reset(test_files)
build_if_needed()
base_commit_results = run_visible_suite() # test -> pass/fail
subtask_results = {} # subtask -> {test -> pass/fail}
for subtask in subtasks:
apply(subtask.test_patch)
subtask_results[subtask.id] = run(subtask.hidden_tests)
revert(subtask.test_patch)
if any_new_failures(base_commit_results):
task_score = 0
else:
task_score = sum(passed_subtasks) / len(subtasks)The benchmark score is calculated as
total number of bugs solved across tasks / total number of bugs =
= mean(number of bugs in task * task score)
Summarize one or more directories containing graded Harbor trials:
sweep info path/to/graded-solutions --per-task
sweep info path/to/graded-solutions --json@misc{lieret2026swesweep,
title = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
Yuxiang Wei and Jeremy Ma and Sten Sootla and
Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
Rui Hou and Ofir Press and John Yang},
year = {2026},
note = {Preprint},
url = {https://github.com/facebookresearch/swe-sweep}
}SWE-sweep is licensed under the terms of the license found in LICENSE.
