Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Notion-as-source-of-truth dispatch board for running your work like an AI agency. One Tasks database is the source of truth; tasks flow Suggestion through Discussion, To-Do, In Progress, and Done with subtasks, recurring cadences, dependencies, and template subtrees. Batch execution fans approved To-Do rows out to parallel agents with per-task model selection. Use when capturing chat to Notion, running the To-Do queue, suggesting, approving, or discussing tasks, or coordinating multi-task batche
.claude/skills/jeremylongshore-agency-os/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 225% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -10% | 0% |
Notion-as-source-of-truth dispatch board. One Tasks database, one Hub page, one page per Corpus, one page each for General Guidance and Resources. The skill mutates Notion via the Notion MCP (mcp__*__notion-* tools); only references/notion-pointers.json is committed to git.
Skill name decision: the skill is named agency-os (matching the repo). All commands are /agency-os <cmd>. This is the single plugin entry point; there is no agency-os/notion sub-namespace. If you embed this plugin alongside others, prefix commands with agency-os to avoid collisions.
agency-os turns a single Notion database into a multi-status dispatch board for AI work. The model is intentionally narrow:
selection, and ownership. No parallel kanban tools.
that every task consults.
In Progress → Done. The dedup gate at each transition prevents accidental re-execution.
run fans approved To-Do rows out to parallel agents. Each task carriesits own model selection (Haiku for cheap fan-out, Sonnet for default, Opus for hard reasoning) and respects declared dependencies.
The skill is stateless on disk — the only committed artifact is references/notion-pointers.json (database/page IDs). All runtime state lives in Notion.
For the full architecture (status flow, sync protocol, workspace structure, pointer/cache format), see references/architecture.md.
npx -y @notionhq/notion-mcp-server (declare in .mcp.json)
NOTION_TOKEN) with read+write access to yourworkspace. Add .env to .gitignore — never commit the token.
references/architecture.md § "Workspace structure" for the schema). The /agency-os init command can scaffold this for you.
scripts/query-tasks.py helper.First-time setup:
bashccpi install agency-os # then in Claude Code: /agency-os init --harness=basic --haiku=cost-tier --sonnet=default --opus=hard-reasoning
The skill is a CLI surface over Notion. Three usage patterns:
/agency-os <cmd> [args]. Seereferences/commands.md for the full reference of 19 commands (init, scaffold, suggest, discuss, log, add-subtask, approve, start, refresh, run, done, kill, next, status, list, show, update, move, plus launch alias).
the corresponding command. Examples in references/natural-language.md.
/agency-os run [--go] fans the entire To-Do queueout to parallel agents with per-task model selection. See ## Examples below for the canonical flow.
Status flow is enforced — you cannot skip a stage. Every command performs a sync preflight to ensure your local view of Notion is current (see references/architecture.md § "Sync — preflight on every command").
When drafting any user-facing copy (READMEs, blog posts, launch surfaces), apply the positioning brief at references/positioning.md before writing.
Every command returns to chat with:
✅ <action> or ⚠️ <reason> (one line, scannable)Batch run additionally emits:
The skill fails closed on five well-defined cases (full details in references/architecture.md § "Status flow — the dedup gate"):
| Condition | Behavior | |---|---| | Notion API auth fails | Halt, print "NOTION_TOKEN missing or invalid", exit 1 | | Database/page ID drift (pointers stale) | Halt, print "Run /agency-os refresh", exit 1 | | Status-flow violation (e.g. approve on a Suggestion) | Halt with the required prerequisite step quoted | | Dependency cycle detected during run | Halt, list the cycle, exit 1 | | Task missing required model selection | Halt, print "Run /agency-os update <id> --model <tier>" |
The skill never silently corrects state in Notion — every fix is an explicit command the operator must run.
Capture a chat insight as a Suggestion:
textUser: add a suggestion: refactor the auth flow to use the new token cache Skill: → /agency-os suggest "refactor the auth flow to use the new token cache" ✅ Created Suggestion #t-2026-05-23-001 in corpus "platform" Next: /agency-os discuss t-2026-05-23-001
Approve and run a batch:
textUser: approve t-2026-05-23-{001..003} then run the queue Skill: ✅ Approved 3 tasks → To-Do /agency-os run --go → fanning to 3 parallel agents... ✅ Done: 2 | ⚠️ Blocked on deps: 1 | Total spend: ~$0.04
More examples and the full command catalog are in references/commands.md.
references/architecture.md — status flow, sync protocol, workspace schemareferences/commands.md — full CLI reference (19 commands)references/natural-language.md — chat-to-command translation tablereferences/positioning.md — canonical brief for user-facing copyreferences/general-guidance.md — shared operating principles applied to every taskreferences/notion-pointers.json — pointer file scaffold (database/page IDs)references/task-page-template.md — Notion page template for new tasksreferences/corpus-template.md — Notion page template for a Corpusreferences/config-template.json — default per-task model routingscripts/query-tasks.py — optional Python helper for offline introspection| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 4,724 | 5,205 | +10% | 1 | 1 | 0% | 826 | 2,682 | +225% | 0 | 0 | — |
case-02 | fail→fail | 9,522 | 5,099 | -46% | 1 | 1 | 0% | 1,608 | 2,016 | +25% | 0 | 0 | — |
case-03 | fail→fail | 9,145 | 5,648 | -38% | 1 | 1 | 0% | 1,772 | 2,192 | +24% | 0 | 0 | — |
case-04 | pass→pass | 6,655 | 2,833 | -57% | 1 | 1 | 0% | 1,363 | 2,223 | +63% | 0 | 0 | — |
case-05 | pass→pass | 3,131 | 1,993 | -36% | 1 | 1 | 0% | 596 | 2,107 | +254% | 0 | 0 | — |
case-06 | pass→fail | 6,620 | 4,223 | -36% | 1 | 1 | 0% | 1,187 | 2,419 | +104% | 0 | 0 | — |
case-07 | fail→pass | 9,850 | 2,738 | -72% | 1 | 1 | 0% | 1,818 | 2,177 | +20% | 0 | 0 | — |
case-08 | fail→pass | 9,920 | 1,448 | -85% | 1 | 1 | 0% | 1,570 | 1,915 | +22% | 0 | 0 | — |
case-09 | fail→pass | 10,319 | 2,503 | -76% | 1 | 1 | 0% | 1,633 | 2,165 | +33% | 0 | 0 | — |
case-10 | pass→pass | 8,365 | 2,600 | -69% | 1 | 1 | 0% | 1,404 | 2,205 | +57% | 0 | 0 | — |
case-11 | fail→pass | 13,818 | 2,204 | -84% | 1 | 1 | 0% | 2,316 | 2,081 | -10% | 0 | 0 | — |
case-12 | fail→pass | 6,632 | 1,247 | -81% | 1 | 1 | 0% | 1,107 | 1,934 | +75% | 0 | 0 | — |
case-13 | fail→pass | 8,219 | 1,554 | -81% | 1 | 1 | 0% | 1,535 | 1,942 | +27% | 0 | 0 | — |
case-14 | pass→pass | 11,041 | 1,657 | -85% | 1 | 1 | 0% | 2,113 | 1,995 | -6% | 0 | 0 | — |
case-15 | fail→pass | 5,174 | 1,608 | -69% | 1 | 1 | 0% | 894 | 2,009 | +125% | 0 | 0 | — |
case-16 | fail→pass | 11,981 | 1,310 | -89% | 1 | 1 | 0% | 2,657 | 1,945 | -27% | 0 | 0 | — |
case-17 | fail→pass | 20,988 | 2,516 | -88% | 1 | 1 | 0% | 1,829 | 2,237 | +22% | 0 | 0 | — |
case-18 | pass→pass | 6,416 | 4,297 | -33% | 1 | 1 | 0% | 1,059 | 2,568 | +142% | 0 | 0 | — |
case-19 | fail→fail | 5,893 | 1,480 | -75% | 1 | 1 | 0% | 980 | 1,961 | +100% | 0 | 0 | — |
case-20 | fail→pass | 5,550 | 1,207 | -78% | 1 | 1 | 0% | 994 | 1,920 | +93% | 0 | 0 | — |
case-21 | fail→pass | 3,862 | 1,323 | -66% | 1 | 1 | 0% | 622 | 1,940 | +212% | 0 | 0 | — |
case-22 | pass→pass | 8,688 | 1,779 | -80% | 1 | 1 | 0% | 1,440 | 1,995 | +39% | 0 | 0 | — |
case-23 | fail→pass | 24,439 | 1,177 | -95% | 1 | 1 | 0% | 4,690 | 1,911 | -59% | 0 | 0 | — |
case-24 | fail→pass | 12,498 | 2,001 | -84% | 1 | 1 | 0% | 2,254 | 2,076 | -8% | 0 | 0 | — |
case-25 | fail→pass | 6,619 | 2,751 | -58% | 1 | 1 | 0% | 1,252 | 2,191 | +75% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 23 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +56 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.