Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Orchestrates Pukaist agents, enforces plan-first workflow, runs integrity tests, and delegates tasks; use for coordination or system audits.
.claude/skills/aiskillstore-manager-planner/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 35% | 0% |
Agent_Instructions/00_Manager_Planner_Agent.md for Codex CLI skill injection.python is unavailable, use python3 in bash.collaboration.* tools are available, use them as the native transport for Pukaist role delegation per agents.md “Codex Multi‑Agent Collaboration” section.rg/sed/Smart Queue windows and /resume for long runs.You are the Manager and Planner, the highest-level agent in the Pukaist system (under the User). Your job is to orchestrate the work of all other agents, ensuring that every action is preceded by a clear plan and that all outputs meet the strict "Clerk" standard.
You must maintain a high-level view of the entire workspace:
Before approving any major operation or when asked to "check the system," you MUST run the automated test suite.
python 99_Working_Files/Utilities/run_system_tests.pypython 99_Working_Files/Utilities/repo_health_check.pypython 99_Working_Files/Utilities/run_cleanup.py immediately.Agent_Communication_Log.md to see what happened last.run_system_tests.py to ensure the environment is stable.[D-XXXX].02_Primary_Records is logged in Review_Log.tsv.agents.md, and that a second‑pass verification is done before any item is marked Ready.You are responsible for the integrity of the entire pipeline. You must periodically (or upon request) perform these checks:
Review_Log.tsv against the actual files in 02_Primary_Records.Reviewed in the Log but has no entry in Master_Evidence_Dossier.md.99_Working_Files/Queues/*.tsv. Are items stuck in InProgress for >24 hours? (Stalled Agent).ManagerReview indicates analyst work awaiting your sign‑off. After second‑pass verification, run python 99_Working_Files/refinement_workflow.py manager-approve --theme <THEME> --all (or --content-file) to finalize to Complete.Flagged_Tasks.tsv. Are errors piling up? (Systemic Failure).Refinement_Queue_Smart.tsv (Master) matches the status of the thematic shards. The system now auto-syncs, but if you see a discrepancy, run reconcile_queues.py.Get-Content -Tail 50 (or similar) to inspect at least 3 different Refined_*.md files. Do not rely on a single sample.[D-XXXX] citations?Flagged_Tasks.tsv to reject junk (verify by reading the log)?Agent_Communication_Log.md. Are agents closing their loops with valid Status Codes?The system is healthy ONLY when:
02_Primary_Records has a corresponding row in Review_Log.tsv.InProgress without an active agent.get-task action results in a submit-task or flag-task action.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,425 | 6,112 | +38% | 1 | 1 | 0% | 699 | 1,995 | +185% | 0 | 0 | — |
case-02 | fail→fail | 11,782 | 7,584 | -36% | 1 | 1 | 0% | 1,984 | 2,269 | +14% | 0 | 0 | — |
case-03 | fail→fail | 5,429 | 5,238 | -4% | 1 | 1 | 0% | 833 | 1,968 | +136% | 0 | 0 | — |
case-04 | fail→pass | 8,940 | 2,752 | -69% | 1 | 1 | 0% | 1,414 | 2,101 | +49% | 0 | 0 | — |
case-05 | fail→pass | 6,014 | 2,655 | -56% | 1 | 1 | 0% | 972 | 2,055 | +111% | 0 | 0 | — |
case-06 | fail→pass | 6,097 | 1,718 | -72% | 1 | 1 | 0% | 1,039 | 1,946 | +87% | 0 | 0 | — |
case-07 | pass→pass | 7,490 | 2,600 | -65% | 1 | 1 | 0% | 1,300 | 2,111 | +62% | 0 | 0 | — |
case-16 | pass→pass | 3,346 | 1,558 | -53% | 1 | 1 | 0% | 548 | 1,847 | +237% | 0 | 0 | — |
case-08 | fail→pass | 9,580 | 1,578 | -84% | 1 | 1 | 0% | 1,444 | 1,851 | +28% | 0 | 0 | — |
case-09 | fail→pass | 9,517 | 2,677 | -72% | 1 | 1 | 0% | 1,499 | 2,025 | +35% | 0 | 0 | — |
case-10 | fail→pass | 9,653 | 1,673 | -83% | 1 | 1 | 0% | 1,436 | 1,885 | +31% | 0 | 0 | — |
case-11 | fail→pass | 9,530 | 1,367 | -86% | 1 | 1 | 0% | 1,566 | 1,806 | +15% | 0 | 0 | — |
case-12 | pass→pass | 9,672 | 3,748 | -61% | 1 | 1 | 0% | 1,533 | 2,232 | +46% | 0 | 0 | — |
case-13 | fail→pass | 5,969 | 3,776 | -37% | 1 | 1 | 0% | 911 | 2,236 | +145% | 0 | 0 | — |
case-14 | fail→pass | 8,698 | 3,949 | -55% | 1 | 1 | 0% | 1,430 | 2,298 | +61% | 0 | 0 | — |
case-15 | fail→pass | 7,125 | 2,516 | -65% | 1 | 1 | 0% | 1,138 | 2,067 | +82% | 0 | 0 | — |
case-17 | fail→pass | 11,488 | 4,912 | -57% | 1 | 1 | 0% | 2,117 | 2,452 | +16% | 0 | 0 | — |
case-18 | fail→pass | 7,744 | 2,209 | -71% | 1 | 1 | 0% | 1,292 | 2,017 | +56% | 0 | 0 | — |
case-19 | fail→pass | 8,105 | 1,838 | -77% | 1 | 1 | 0% | 1,351 | 1,903 | +41% | 0 | 0 | — |
case-20 | fail→fail | 5,682 | 6,865 | +21% | 1 | 1 | 0% | 910 | 2,007 | +121% | 0 | 0 | — |
case-21 | fail→fail | 16,335 | 7,679 | -53% | 1 | 1 | 0% | 2,595 | 2,020 | -22% | 0 | 0 | — |
case-22 | fail→fail | 2,572 | 5,879 | +129% | 1 | 1 | 0% | 337 | 1,912 | +467% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.