Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. The quality of an MCP server is measured by how well it enables LLMs to accomplish real-world tasks.
.claude/skills/sickn33-mcp-builder/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 282% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 9% | 0% |
Modified in AAS on 2026-09-05: version-scoped examples and bounded evaluation.
Create MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. The quality of an MCP server is measured by how well it enables LLMs to accomplish real-world tasks.
Creating a high-quality MCP server involves four main phases:
API Coverage vs. Workflow Tools: Balance comprehensive API endpoint coverage with specialized workflow tools. Workflow tools can be more convenient for specific tasks, while comprehensive coverage gives agents flexibility to compose operations. Performance varies by client—some clients benefit from code execution that combines basic tools, while others work better with higher-level workflows. Start with the minimum operations needed for the user's authorized workflow. Expand coverage from observed gaps.
Tool Naming and Discoverability: Clear, descriptive tool names help agents find the right tools quickly. Use consistent prefixes (e.g., github_create_issue, github_list_repos) and action-oriented naming.
Context Management: Agents benefit from concise tool descriptions and the ability to filter/paginate results. Design tools that return focused, relevant data. Some clients support code execution which can help agents filter and process data efficiently.
Actionable Error Messages: Error messages should guide agents toward solutions with specific suggestions and next steps.
Navigate the MCP specification:
Start with the sitemap to find relevant pages: https://modelcontextprotocol.io/sitemap.xml
Then fetch specific pages with .md suffix for markdown format (e.g., https://modelcontextprotocol.io/specification/2025-11-25/index).
Key pages to review:
Choose the project-compatible stack:
Load framework documentation:
For TypeScript (recommended):
https://raw.githubusercontent.com/modelcontextprotocol/typescript-sdk/main/README.mdFor Python:
https://raw.githubusercontent.com/modelcontextprotocol/python-sdk/main/README.mdUnderstand the API: Review the service's API documentation to identify key endpoints, authentication requirements, and data models. Use available read-only research tools as needed.
Tool Selection: List the exact operations and permissions needed by the task. Keep write tools separate and require authorization at the server boundary.
See language-specific guides for project setup:
Create shared utilities:
For each tool:
Input Schema:
Output Schema:
outputSchema where possible for structured datastructuredContent in tool responses (TypeScript SDK feature)Tool Description:
Implementation:
Annotations:
readOnlyHint: true/falsedestructiveHint: true/falseidempotentHint: true/falseopenWorldHint: true/falseReview for:
TypeScript:
npm run build to verify compilationnpx @modelcontextprotocol/inspectorPython:
python -m py_compile your_server.pySee language-specific guides for detailed testing approaches and quality checklists.
After implementing your MCP server, create comprehensive evaluations to test its effectiveness.
Load ✅ Evaluation Guide for complete evaluation guidelines.
Use evaluations to test whether LLMs can effectively use your MCP server to answer realistic, complex questions.
To create effective evaluations, follow the process outlined in the evaluation guide:
Ensure each question is:
Create an XML file with this structure:
xml<evaluation> <qa_pair> <question>Find discussions about AI model launches with animal codenames. One model needed a specific safety designation that uses the format ASL-X. What number X was being determined for the model named after a spotted wild cat?</question> <answer>3</answer> </qa_pair> <!-- More qa_pairs... --> </evaluation>
Load these resources as needed during development:
https://modelcontextprotocol.io/sitemap.xml, then fetch specific pages with .md suffixhttps://raw.githubusercontent.com/modelcontextprotocol/python-sdk/main/README.mdhttps://raw.githubusercontent.com/modelcontextprotocol/typescript-sdk/main/README.md@mcp.toolserver.registerToolUse for a new MCP tool contract, a transport/client compatibility defect, or a review of a server's bounded input/output and permission behavior. For an existing server, inspect its implementation, lockfile and actual protocol negotiation before changes.
Record the SDK version, protocol revision, runtime, intended client, exact tool names, service scopes and permitted test data. The bundled Python evaluator targets the SDK v1 API; the reference guides identify version-sensitive sketches. Do not combine v1 package imports with a newer SDK guide or assume a draft specification is deployed.
For a read-only document search server, use three fixture documents and one denied tenant. Confirm initialization, tools/list (including pagination), search, empty results, unknown tools, malformed inputs, bounded oversized responses and denied access. Run the exact packaged command in the intended client. Record the observed calls and outcomes. A build or Inspector probe alone is not a real-client result.
Expected handoff: the smallest working contract, reproducible fixtures and commands, client/version evidence, known limits and separate authorization for any service writes. The bundled evaluator is optional and makes billable Anthropic API calls; it sends the questions, selected tool schemas and tool results to that provider. Use only approved data and endpoints, an explicit model, and explicitly reviewed read-only tool names. It does not infer safety from tool annotations. See the evaluation guide for limits.
prompt-injection resistance or general reliability. Test those boundaries directly.
in returned documents or expand permissions because a tool suggests it.
customer workflow. Pin their source snapshots; closed historical records can change.
stdio connection launches the specified executable; inspect it and use a safe fixture.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 23,283 | 17,441 | -25% | 1 | 1 | 0% | 4,050 | 5,265 | +30% | 0 | 0 | — |
case-02 | pass→pass | 24,498 | 25,578 | +4% | 1 | 1 | 0% | 5,393 | 7,889 | +46% | 0 | 0 | — |
case-03 | fail→pass | 41,971 | 12,739 | -70% | 1 | 1 | 0% | 1,120 | 4,280 | +282% | 0 | 0 | — |
case-04 | pass→pass | 15,336 | 15,696 | +2% | 1 | 1 | 0% | 3,299 | 5,349 | +62% | 0 | 0 | — |
case-05 | pass→pass | 17,825 | 13,762 | -23% | 1 | 1 | 0% | 3,004 | 4,295 | +43% | 0 | 0 | — |
case-06 | pass→pass | 16,547 | 16,159 | -2% | 1 | 1 | 0% | 3,810 | 5,746 | +51% | 0 | 0 | — |
case-07 | pass→pass | 16,809 | 9,579 | -43% | 1 | 1 | 0% | 2,696 | 3,483 | +29% | 0 | 0 | — |
case-08 | pass→pass | 9,212 | 4,656 | -49% | 1 | 1 | 0% | 1,575 | 2,766 | +76% | 0 | 0 | — |
case-09 | fail→pass | 7,446 | 2,190 | -71% | 1 | 1 | 0% | 1,412 | 2,428 | +72% | 0 | 0 | — |
case-10 | fail→pass | 8,854 | 1,953 | -78% | 1 | 1 | 0% | 1,509 | 2,336 | +55% | 0 | 0 | — |
case-11 | pass→pass | 14,123 | 8,457 | -40% | 1 | 1 | 0% | 2,388 | 3,519 | +47% | 0 | 0 | — |
case-12 | fail→fail | 7,302 | 3,916 | -46% | 1 | 1 | 0% | 1,381 | 2,760 | +100% | 0 | 0 | — |
case-13 | pass→pass | 23,522 | 4,049 | -83% | 1 | 1 | 0% | 4,652 | 2,797 | -40% | 0 | 0 | — |
case-14 | pass→pass | 12,069 | 7,384 | -39% | 1 | 1 | 0% | 2,158 | 3,311 | +53% | 0 | 0 | — |
case-15 | fail→pass | 11,750 | 2,741 | -77% | 1 | 1 | 0% | 2,190 | 2,387 | +9% | 0 | 0 | — |
case-16 | fail→pass | 13,124 | 6,256 | -52% | 1 | 1 | 0% | 2,165 | 2,951 | +36% | 0 | 0 | — |
case-17 | fail→pass | 13,266 | 7,186 | -46% | 1 | 1 | 0% | 2,455 | 3,309 | +35% | 0 | 0 | — |
case-18 | pass→pass | 3,317 | 2,624 | -21% | 1 | 1 | 0% | 523 | 2,447 | +368% | 0 | 0 | — |
case-19 | pass→pass | 3,624 | 1,596 | -56% | 1 | 1 | 0% | 595 | 2,288 | +285% | 0 | 0 | — |
case-20 | fail→pass | 4,740 | 3,327 | -30% | 1 | 1 | 0% | 766 | 2,536 | +231% | 0 | 0 | — |
case-21 | pass→pass | 9,791 | 1,607 | -84% | 1 | 1 | 0% | 1,612 | 2,281 | +42% | 0 | 0 | — |
case-22 | fail→pass | 15,158 | 7,775 | -49% | 1 | 1 | 0% | 2,577 | 3,282 | +27% | 0 | 0 | — |
case-23 | pass→pass | 4,943 | 3,783 | -23% | 1 | 1 | 0% | 882 | 2,581 | +193% | 0 | 0 | — |
case-24 | fail→pass | 11,803 | 3,682 | -69% | 1 | 1 | 0% | 2,090 | 2,623 | +26% | 0 | 0 | — |
case-25 | fail→pass | 7,461 | 1,618 | -78% | 1 | 1 | 0% | 1,229 | 2,271 | +85% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 24 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +44 percentage points is the difference between those two pass rates over the 24 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.