Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deterministic benchmark harnesses for public template exemplars. Use when scoring generated project outputs against benchmark manifests, refreshing the default template smoke manifest, checking publication-readiness rubrics, or adding bounded no-network readiness checks for public template outputs.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -62% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -62% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -50% | 0% |
Use infrastructure.benchmark for small, deterministic readiness benchmarks over public template exemplar outputs. The module reads real files, applies explicit manifest checks and optional weighted rubrics, and emits JSON or Markdown score reports.
bashuv run python -m infrastructure.benchmark --repo-root . uv run python -m infrastructure.benchmark \ --repo-root . \ --output-json /tmp/template_benchmark.json \ --output-markdown /tmp/template_benchmark.md uv run python -m infrastructure.benchmark \ --repo-root . \ --write-default-manifest
failed checks, weights, and partial scores remain inspectable.
never shrink it to make a readiness run green.
BenchmarkManifestBenchmarkCheckResultBenchmarkScoreRubricScoreRubricSetload_benchmark_manifestrun_benchmark_manifestscore_project_against_manifestscore_rubricscores_to_dictscores_to_markdownwrite_default_manifestPair this skill with README.md and AGENTS.md for module rules and validation commands.
Other measured skills in the registry, with their headline benchmark lift.