Real DeepSeek V4 Flash · 12.8 s replay · 120 fps video
- Small, but complete. About 1,700 lines of Python: nine tools, context compaction, request retries, streaming responses, and session memory.
- Tools defined with Pydantic. Typed inputs, generated JSON Schema, and validation before execution make tools easier to compose and orchestrate.
- A practical baseline. Evaluated on SWE-bench Verified and Terminal-Bench 2.1 with DeepSeek V4 Flash. See the results below.
- Built for learning. Follow the agent loop, tools, compaction, and retry policy in ordinary Python. The optional TUI is a single file you can read and modify.
Historical, self-reported results with DeepSeek V4 Flash. The latest source changes have not been re-evaluated on these benchmarks.
| Benchmark | Solved / attempts | Score |
|---|---|---|
| SWE-bench Verified | 401 / 500 | 80.2% |
| Terminal-Bench 2.1 | 309 / 445 | 69.44% |
The figure inserts our score into the public Terminal-Bench 2.1 leaderboard as of September 4, 2026; 13th of 19 is an illustrative position, not an official rank. The Terminal-Bench result averages five attempts on each of 89 tasks, not pass@5. See evaluation details, recorded failures, and reproduction limits.
A tool combines a Pydantic input model, a Python function, and a ToolDefinition.
After installing and exporting DEEPSEEK_API_KEY (see Get started),
pass your definitions to the agent to choose which tools it can use:
from pydantic import BaseModel, ConfigDict, Field
from mini_harness.agent import DeepSeekAgent
from mini_harness.tool.box import TOOLS, ToolDefinition
class CountWordsInput(BaseModel):
model_config = ConfigDict(extra="forbid")
text: str = Field(min_length=1, description="Text to count words in.")
def count_words(args: CountWordsInput) -> str:
return str(len(args.text.split()))
count_words_tool = ToolDefinition(
name="count_words",
description="Count whitespace-separated words in text.",
parameters=CountWordsInput,
function=count_words,
risky=False,
)
agent = DeepSeekAgent([*TOOLS, count_words_tool])The agent uses model_json_schema() to describe each tool to the model.
Before dispatch, the tool executor calls
model_validate_json() to validate its arguments; invalid calls return an error
for the agent to correct. One definition keeps the schema, validation, and
execution connected. See the Pydantic model documentation.
Install uv if you do not have it, then clone the repository. Python 3.12+ is required; uv can install it.
# macOS / Linux — install uv once, then restart your terminal
curl -LsSf https://astral.sh/uv/install.sh | shgit clone https://github.com/mini-harness/mini-harness.git
cd mini-harness
uv python install 3.12
uv sync --lockedSet your DeepSeek API key in the shell,
then launch the TUI. No .env file is needed.
export DEEPSEEK_API_KEY="your-api-key"
uv run tui.pyEnter sends a message · Ctrl+N starts a session · F2 opens sessions ·
Ctrl+E expands tools · Ctrl+C cancels · Ctrl+Q quits.
The agent reads the project workspace; local file tools write to sandbox/.
The TUI asks before shell, sandbox, or subagent calls. Sessions stay in .local/ and are
ignored by Git. Cancelling a task keeps file edits already completed.
To run generated code with run_sandbox, install and start
Docker, then pull its Python image once:
docker pull python:3.12-slimrun_sandbox is a Pydantic-defined tool accepting command and
timeout_seconds (1–300, default 30). It runs in a disposable Python 3.12
container with networking disabled, a read-only system, and limits of one CPU,
256 MB RAM, and 64 processes. Only sandbox/ is mounted at /workspace; changes
there persist. For example, {"command": "python hello.py"} runs
sandbox/hello.py. No host environment variables are forwarded into the container.
The host run_bash tool remains available and is not isolated.
For the plain terminal interface:
uv run --locked mini-harnessRun these commands from the repository root after installation. The adapters in
bench/ connect mini-harness to Harbor and SWE-bench cloud scoring.
The examples use OpenAI; for DeepSeek, export DEEPSEEK_API_KEY and change the
model to deepseek/deepseek-v4-flash.
With Docker running, evaluate one SWE-bench Verified task using the Harbor adapter:
export OPENAI_API_KEY="your-openai-api-key"
PYTHONPATH=. uv run --with "harbor==0.20.0" harbor run \
-d swebench-verified \
--agent bench.adapter:MiniHarnessAgent \
--model openai/gpt-4.1 --env docker --n-tasks 1 -n 1Harbor runs the agent and verifier, saving results and ATIF trajectories under
jobs/. Remove --n-tasks 1 to evaluate the full dataset; -n controls concurrency.
See Harbor's custom-agent guide.
The cloud entry point uses Modal to run the agent, exports its
patches, then submits them to sb-cli for scoring. Set OPENAI_API_KEY as above,
obtain a verified sb-cli API key,
and sign in to Modal:
export SWEBENCH_API_KEY="your-verified-sb-cli-api-key"
uv run --with-requirements bench/requirements-sb-cli.txt modal setup
uv run --with-requirements bench/requirements-sb-cli.txt \
python -m bench.sb_cli run \
--job-dir jobs/sb-cli-smoke --run-id mini-harness-smoke \
--model openai/gpt-4.1 --n-tasks 1Add --dry-run to the final command to preview without launching a job. Replace
--n-tasks 1 with --all for all 500 tasks, using a new job directory and run ID
for each run. Predictions are saved to jobs/sb-cli-smoke/predictions.json and
reports to jobs/sb-cli-smoke/sb-cli-reports/. The generate, export, and
submit subcommands also let you run each stage separately.