Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design an MCP server for a product — the tool surface, auth model, and safety boundaries that make it genuinely usable by AI agents. Use when asked to spec an MCP server, expose a product to agents, design tools for Claude or other MCP clients, or review why an existing MCP server performs badly. Produces a complete server spec: a small task-shaped toolset with agent-tested descriptions, auth and scoping decisions, error design, and an explicit not-exposed list.
.claude/skills/mohitagw15856-mcp-server-spec/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 26 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 11% | 0% |
Every SaaS is shipping an MCP server; most dump their REST API as forty tools and wonder why agents flail. This skill designs the server as what it actually is: a user interface for a non-human user — few tools, task-shaped, with descriptions written for a model deciding under uncertainty.
Ask for (if not already provided):
has_more; keep any response under ~2k tokens by default with an opt-in for detail."date must be YYYY-MM-DD" beats 400 Bad Request. Every error names the parameter at fault and the fix.Agent jobs served: the 5-8 tasks] · Tool count: n] · Auth: model + scoping]
Tools | Tool | Description (as shipped) | Key params | Returns | Risk class | |---|---|---|---|---|
Gated actions: which tools require confirmation params, and the expected client behaviour]
Never exposed: capability → reason] (one line each; this list is reviewed like an API contract)
Error design: the error shape + 3 example messages]
Test plan: 10-15 realistic agent prompts spanning the jobs; run against a real client; a tool whose description gets misselected twice gets rewritten, not documented around]
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 29,787 | 22,580 | -24% | 1 | 1 | 0% | 6,292 | 5,595 | -11% | 0 | 0 | — |
case-02 | fail→pass | 29,520 | 27,043 | -8% | 1 | 1 | 0% | 6,297 | 6,586 | +5% | 0 | 0 | — |
case-03 | fail→fail | 29,596 | 29,990 | +1% | 1 | 1 | 0% | 5,682 | 6,956 | +22% | 0 | 0 | — |
case-04 | pass→pass | 11,165 | 17,572 | +57% | 1 | 1 | 0% | 2,473 | 4,898 | +98% | 0 | 0 | — |
case-05 | pass→pass | 13,248 | 14,487 | +9% | 1 | 1 | 0% | 2,621 | 4,309 | +64% | 0 | 0 | — |
case-06 | pass→pass | 22,819 | 22,150 | -3% | 1 | 1 | 0% | 5,671 | 6,383 | +13% | 0 | 0 | — |
case-07 | fail→pass | 14,094 | 13,449 | -5% | 1 | 1 | 0% | 2,927 | 3,710 | +27% | 0 | 0 | — |
case-08 | fail→pass | 9,572 | 6,882 | -28% | 1 | 1 | 0% | 1,629 | 2,437 | +50% | 0 | 0 | — |
case-09 | fail→pass | 21,069 | 14,457 | -31% | 1 | 1 | 0% | 2,229 | 2,898 | +30% | 0 | 0 | — |
case-10 | pass→pass | 9,953 | 22,050 | +122% | 1 | 1 | 0% | 1,773 | 2,526 | +42% | 0 | 0 | — |
case-11 | pass→pass | 10,126 | 10,304 | +2% | 1 | 1 | 0% | 1,727 | 2,880 | +67% | 0 | 0 | — |
case-12 | pass→pass | 25,591 | 21,324 | -17% | 1 | 1 | 0% | 1,901 | 3,029 | +59% | 0 | 0 | — |
case-13 | pass→pass | 29,958 | 20,836 | -30% | 1 | 1 | 0% | 2,954 | 3,253 | +10% | 0 | 0 | — |
case-14 | pass→pass | 11,804 | 12,243 | +4% | 1 | 1 | 0% | 2,481 | 3,303 | +33% | 0 | 0 | — |
case-15 | fail→fail | 14,496 | 17,017 | +17% | 1 | 1 | 0% | 2,598 | 4,196 | +62% | 0 | 0 | — |
case-16 | fail→fail | 16,466 | 14,254 | -13% | 1 | 1 | 0% | 2,920 | 3,624 | +24% | 0 | 0 | — |
case-17 | pass→pass | 16,188 | 11,815 | -27% | 1 | 1 | 0% | 3,087 | 3,649 | +18% | 0 | 0 | — |
case-18 | pass→pass | 11,423 | 10,813 | -5% | 1 | 1 | 0% | 2,117 | 3,112 | +47% | 0 | 0 | — |
case-19 | pass→pass | 7,289 | 5,322 | -27% | 1 | 1 | 0% | 1,369 | 2,156 | +57% | 0 | 0 | — |
case-20 | fail→pass | 10,257 | 5,236 | -49% | 1 | 1 | 0% | 1,883 | 2,097 | +11% | 0 | 0 | — |
case-21 | fail→pass | 12,217 | 10,747 | -12% | 1 | 1 | 0% | 2,104 | 3,007 | +43% | 0 | 0 | — |
case-22 | pass→pass | 11,229 | 6,436 | -43% | 1 | 1 | 0% | 1,939 | 2,047 | +6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.