Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build Claude tool use (function calling) workflows with the Messages API. Use when implementing tool use, function calling, agent loops, or building AI assistants that interact with external systems. Trigger with phrases like "claude tool use", "anthropic function calling", "claude tools", "agent loop anthropic", "tool_use blocks".
.claude/skills/jeremylongshore-anth-core-workflow-a/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 250% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 112% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 287% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 69% | 0% |
Implement Claude's tool use capability where the model can call functions you define. Claude returns tool_use content blocks with structured JSON inputs; your code executes the function and returns tool_result blocks. This is the foundation for building AI agents.
anth-install-auth setuppythonimport anthropic client = anthropic.Anthropic() tools = [ { "name": "get_weather", "description": "Get current weather for a city. Use when the user asks about weather conditions.", "input_schema": { "type": "object", "properties": { "city": { "type": "string", "description": "City name, e.g. 'San Francisco, CA'" }, "units": { "type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature units" } }, "required": ["city"] } }, { "name": "search_database", "description": "Search product database by query string. Returns matching products.", "input_schema": { "type": "object", "properties": { "query": {"type": "string"}, "max_results": {"type": "integer", "default": 10} }, "required": ["query"] } } ]
pythonmessage = client.messages.create( model="claude-sonnet-4-20250514", max_tokens=1024, tools=tools, messages=[{"role": "user", "content": "What's the weather in Tokyo?"}] ) # Claude responds with stop_reason="tool_use" # message.content contains both text and tool_use blocks: # [ # {"type": "text", "text": "I'll check the weather for you."}, # {"type": "tool_use", "id": "toolu_01A...", "name": "get_weather", # "input": {"city": "Tokyo", "units": "celsius"}} # ]
pythondef execute_tool(name: str, input_data: dict) -> str: """Route tool calls to actual implementations.""" if name == "get_weather": # Call your weather API return '{"temp": 22, "condition": "partly cloudy", "humidity": 65}' elif name == "search_database": return '{"results": [{"name": "Widget A", "price": 29.99}]}' raise ValueError(f"Unknown tool: {name}") # Extract tool_use blocks and execute tool_results = [] for block in message.content: if block.type == "tool_use": result = execute_tool(block.name, block.input) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, # Must match the tool_use block id "content": result }) # Continue conversation with tool results follow_up = client.messages.create( model="claude-sonnet-4-20250514", max_tokens=1024, tools=tools, messages=[ {"role": "user", "content": "What's the weather in Tokyo?"}, {"role": "assistant", "content": message.content}, {"role": "user", "content": tool_results} ] ) print(follow_up.content[0].text) # "The current weather in Tokyo is 22°C and partly cloudy with 65% humidity."
pythondef run_agent(user_message: str, tools: list, max_turns: int = 10) -> str: """Run an agentic loop that handles multiple sequential tool calls.""" messages = [{"role": "user", "content": user_message}] for _ in range(max_turns): response = client.messages.create( model="claude-sonnet-4-20250514", max_tokens=4096, tools=tools, messages=messages ) # If Claude is done (no more tool calls), return final text if response.stop_reason == "end_turn": return next( (b.text for b in response.content if b.type == "text"), "" ) # Process tool calls messages.append({"role": "assistant", "content": response.content}) tool_results = [] for block in response.content: if block.type == "tool_use": result = execute_tool(block.name, block.input) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": result }) messages.append({"role": "user", "content": tool_results}) return "Max turns reached"
tool_use / tool_result message threading| Error | Cause | Solution | |-------|-------|----------| | invalid_request_error: tool schema invalid | Malformed input_schema | Validate against JSON Schema spec | | tool_use_id mismatch | Result ID doesn't match tool_use ID | Copy block.id exactly | | Claude ignores tools | Description too vague | Add clear "Use when..." descriptions | | Infinite loop | Claude keeps calling tools | Add max_turns guard + tool_choice: {"type": "auto"} |
python# Let Claude decide (default) tool_choice={"type": "auto"} # Force Claude to use a specific tool tool_choice={"type": "tool", "name": "get_weather"} # Force Claude to use any tool (must call at least one) tool_choice={"type": "any"}
For a support assistant, define lookup_order with an order_id string and return a structured order status from the application database. Send the tool result back using the original block.id; a successful run either returns an end_turn response with the status in plain language or requests the next tool needed to answer the user. Keep max_turns bounded so an unavailable dependency fails predictably instead of looping.
For streaming with tools, see anth-core-workflow-b.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 6,320 | 6,005 | -5% | 1 | 1 | 0% | 1,432 | 3,035 | +112% | 0 | 0 | — |
case-02 | pass→pass | 2,598 | 2,310 | -11% | 1 | 1 | 0% | 535 | 2,070 | +287% | 0 | 0 | — |
case-03 | pass→pass | 8,950 | 6,773 | -24% | 1 | 1 | 0% | 1,802 | 3,042 | +69% | 0 | 0 | — |
case-04 | pass→pass | 4,045 | 3,196 | -21% | 1 | 1 | 0% | 756 | 2,319 | +207% | 0 | 0 | — |
case-05 | pass→pass | 7,044 | 6,881 | -2% | 1 | 1 | 0% | 1,393 | 2,936 | +111% | 0 | 0 | — |
case-06 | pass→pass | 8,952 | 5,433 | -39% | 1 | 1 | 0% | 1,848 | 2,861 | +55% | 0 | 0 | — |
case-07 | fail→pass | 3,442 | 3,342 | -3% | 1 | 1 | 0% | 625 | 2,189 | +250% | 0 | 0 | — |
case-08 | pass→pass | 3,315 | 2,591 | -22% | 1 | 1 | 0% | 728 | 2,133 | +193% | 0 | 0 | — |
case-09 | pass→pass | 3,466 | 2,897 | -16% | 1 | 1 | 0% | 642 | 2,179 | +239% | 0 | 0 | — |
case-10 | pass→pass | 4,549 | 2,041 | -55% | 1 | 1 | 0% | 914 | 1,984 | +117% | 0 | 0 | — |
case-11 | pass→pass | 3,436 | 2,069 | -40% | 1 | 1 | 0% | 587 | 2,078 | +254% | 0 | 0 | — |
case-12 | pass→pass | 11,588 | 6,264 | -46% | 1 | 1 | 0% | 2,305 | 2,914 | +26% | 0 | 0 | — |
case-13 | fail→pass | 10,195 | 5,144 | -50% | 1 | 1 | 0% | 1,664 | 2,549 | +53% | 0 | 0 | — |
case-14 | pass→pass | 9,172 | 4,411 | -52% | 1 | 1 | 0% | 1,582 | 2,473 | +56% | 0 | 0 | — |
case-15 | pass→pass | 8,537 | 4,344 | -49% | 1 | 1 | 0% | 1,728 | 2,533 | +47% | 0 | 0 | — |
case-16 | pass→pass | 6,724 | 6,293 | -6% | 1 | 1 | 0% | 1,289 | 2,858 | +122% | 0 | 0 | — |
case-17 | pass→pass | 8,693 | 6,726 | -23% | 1 | 1 | 0% | 1,724 | 3,097 | +80% | 0 | 0 | — |
case-18 | pass→pass | 4,851 | 3,597 | -26% | 1 | 1 | 0% | 1,012 | 2,326 | +130% | 0 | 0 | — |
case-19 | pass→pass | 3,919 | 2,683 | -32% | 1 | 1 | 0% | 688 | 2,123 | +209% | 0 | 0 | — |
case-20 | fail→fail | 12,894 | 12,572 | -2% | 1 | 1 | 0% | 2,871 | 4,358 | +52% | 0 | 0 | — |
case-21 | fail→fail | 5,759 | 3,052 | -47% | 1 | 1 | 0% | 1,170 | 2,241 | +92% | 0 | 0 | — |
case-22 | fail→fail | 18,161 | 15,776 | -13% | 1 | 1 | 0% | 3,101 | 4,732 | +53% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.