Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -45% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 30% | 0% |
DeepSWE (deepswe.datacurve.ai) is a 113-task Harbor-compatible coding-agent benchmark. It runs via Pier (Harbor fork) driving mini-swe-agent (model-agnostic). Any model reachable through OpenRouter can be scored.
bashwhich uv git docker || echo "MISSING: install uv, git, docker" docker info >/dev/null 2>&1 || echo "MISSING: Docker daemon not running (Pier's default sandbox)" echo "OPENROUTER_API_KEY set? ${OPENROUTER_API_KEY:+YES}"
Docker must be running — Pier sandboxes each task in Docker by default (--env modal for cloud instead).
OPENROUTER_API_KEY must already be present in the environment. If it is unset, ask the user to configure their preferred secret-management path; do not read shell startup files, print secrets, or invent a key.
bashgit clone https://github.com/datacurve-ai/deep-swe && cd deep-swe uv tool install datacurve-pier # PyPI (preferred) # or: uv tool install git+https://github.com/datacurve-ai/pier # pier bundles mini-swe-agent as the --agent driver
Run all pier commands from inside deep-swe/, using relative -p tasks/....
mini-swe-agent has a native OpenRouter model class. Both routes below use OPENROUTER_API_KEY and the OpenRouter slug (vendor/model, e.g. minimax/minimax-m3):
Route A — native OpenRouter class (preferred, hits openrouter.ai/api/v1 directly):
bashpier run -p deep-swe/tasks --agent mini-swe-agent \ --model minimax/minimax-m3 --model-class openrouter
Route B — LiteLLM provider prefix (fallback; same key):
bashpier run -p deep-swe/tasks --agent mini-swe-agent \ --model openrouter/minimax/minimax-m3
Notes:
export MSWEA_COST_TRACKING=ignore_errors.pier run --help and mini --help.Always validate end-to-end wiring on a single task before spending tokens on the corpus:
bashpier run -p deep-swe/tasks/<task-id> --agent mini-swe-agent \ --model minimax/minimax-m3 --model-class openrouter # list available task ids: ls deep-swe/tasks
Pass criteria: run completes, model returns actions (not auth/format errors), a score/trajectory is emitted. If it 401s → key wrong. If "provider not provided"/"model not mapped" → fix slug or switch route.
bashpier run -p deep-swe/tasks --agent mini-swe-agent \ --model minimax/minimax-m3 --model-class openrouter \ --n-tasks 10 --sample-seed 0
bashpier run -p deep-swe/tasks --agent mini-swe-agent \ --model minimax/minimax-m3 --model-class openrouter # add `--env modal` to run in parallel Modal sandboxes (needs Modal configured)
jobs/<run>/<trial_id>/. Inspect with pier view jobs/<run>, pier analyze jobs/<run>, or pier critique run jobs/<run>.| Symptom | Cause | Fix | |---|---|---| | HTTP 401 | bad/missing key | re-export OPENROUTER_API_KEY | | "LLM Provider NOT provided" | missing slug prefix | use Route B openrouter/... or Route A with --model-class openrouter | | "model isn't mapped"/cost error | unknown cost for model | export MSWEA_COST_TRACKING=ignore_errors | | unknown flag | version drift | check pier run --help |
davidondrej/skills; verify local paths, tools, credentials, and agent features before acting.Other measured skills in the registry, with their headline benchmark lift.