Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Stream Claude responses, use system prompts, handle multi-turn conversations, Use when working with model-inference patterns. and process structured output with the Messages API. Trigger with "anthropic streaming", "claude messages api", "claude inference", "stream claude response".
.claude/skills/jeremylongshore-clade-model-inference/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 215% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 289% | 0% |
The Messages API is the only inference endpoint. Every Claude interaction goes through client.messages.create(). This skill covers streaming, system prompts, vision, and structured output.
clade-install-authclade-hello-worldtypescriptimport Anthropic from '@claude-ai/sdk'; const client = new Anthropic(); const stream = client.messages.stream({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, messages: [{ role: 'user', content: 'Write a haiku about TypeScript.' }], }); for await (const event of stream) { if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') { process.stdout.write(event.delta.text); } } const finalMessage = await stream.finalMessage(); console.log('\n\nTokens:', finalMessage.usage);
typescriptconst message = await client.messages.create({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, messages: [{ role: 'user', content: [ { type: 'image', source: { type: 'base64', media_type: 'image/png', data: fs.readFileSync('screenshot.png').toString('base64'), }, }, { type: 'text', text: 'Describe what you see in this image.' }, ], }], });
typescriptconst message = await client.messages.create({ model: 'claude-sonnet-4-20250514', max_tokens: 1024, system: `Respond with valid JSON only. Schema: { "summary": string, "sentiment": "positive"|"negative"|"neutral", "confidence": number }`, messages: [{ role: 'user', content: 'Analyze: "This product exceeded my expectations!"' }], }); const result = JSON.parse(message.content[0].text); // { summary: "Very positive review", sentiment: "positive", confidence: 0.95 }
pythonimport anthropic client = anthropic.Anthropic() with client.messages.stream( model="claude-sonnet-4-20250514", max_tokens=1024, messages=[{"role": "user", "content": "Write a haiku about Python."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) print(f"\nTokens: {stream.get_final_message().usage}")
Message object with content, usage, stop_reasonmessage_start — message metadatacontent_block_start — new content block beginningcontent_block_delta — incremental text (text_delta) or tool input (input_json_delta)message_delta — final stop_reason and usagemessage_stop — stream complete| Error | Cause | Solution | |-------|-------|----------| | overloaded_error (529) | Anthropic API temporarily overloaded | Retry with exponential backoff; use client.messages.create with built-in retries | | rate_limit_error (429) | Exceeded RPM or TPM | Check retry-after header. See clade-rate-limits | | invalid_request_error | Image too large or bad format | Max 20 images per request. Supported: PNG, JPEG, GIF, WebP. Max 5MB each |
| Parameter | Type | Description | |-----------|------|-------------| | model | string | Required. Model ID (e.g. claude-sonnet-4-20250514) | | max_tokens | int | Required. Maximum output tokens (1–8192 typical) | | messages | array | Required. Alternating user/assistant messages | | system | string | Optional. System prompt for behavior/persona | | temperature | float | Optional. 0.0–1.0, default 1.0 | | top_p | float | Optional. Nucleus sampling threshold | | stop_sequences | string] | Optional. Custom stop strings | | stream | boolean | Optional. Enable SSE streaming |
See Step 1 (streaming), Step 2 (vision with base64 images), and Step 3 (structured JSON output) above. Python streaming example included.
See clade-embeddings-search for tool use and function calling patterns.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 7,456 | 5,565 | -25% | 1 | 1 | 0% | 1,483 | 2,333 | +57% | 0 | 0 | — |
case-02 | fail→pass | 13,414 | 11,365 | -15% | 1 | 1 | 0% | 2,640 | 3,308 | +25% | 0 | 0 | — |
case-03 | pass→pass | 7,480 | 6,348 | -15% | 1 | 1 | 0% | 1,485 | 2,503 | +69% | 0 | 0 | — |
case-04 | pass→pass | 9,753 | 6,570 | -33% | 1 | 1 | 0% | 1,800 | 2,476 | +38% | 0 | 0 | — |
case-05 | pass→pass | 10,145 | 5,118 | -50% | 1 | 1 | 0% | 1,932 | 2,300 | +19% | 0 | 0 | — |
case-10 | pass→pass | 4,607 | 2,293 | -50% | 1 | 1 | 0% | 639 | 1,700 | +166% | 0 | 0 | — |
case-06 | pass→pass | 13,001 | 5,840 | -55% | 1 | 1 | 0% | 2,138 | 2,304 | +8% | 0 | 0 | — |
case-07 | fail→fail | 15,850 | 9,403 | -41% | 1 | 1 | 0% | 2,577 | 2,843 | +10% | 0 | 0 | — |
case-08 | pass→pass | 5,462 | 3,776 | -31% | 1 | 1 | 0% | 998 | 1,840 | +84% | 0 | 0 | — |
case-09 | pass→pass | 4,892 | 7,771 | +59% | 1 | 1 | 0% | 897 | 2,144 | +139% | 0 | 0 | — |
case-11 | pass→pass | 8,369 | 3,220 | -62% | 1 | 1 | 0% | 738 | 1,757 | +138% | 0 | 0 | — |
case-12 | pass→pass | 2,893 | 4,497 | +55% | 1 | 1 | 0% | 513 | 1,969 | +284% | 0 | 0 | — |
case-13 | pass→pass | 3,100 | 2,795 | -10% | 1 | 1 | 0% | 504 | 1,656 | +229% | 0 | 0 | — |
case-14 | pass→pass | 7,482 | 4,570 | -39% | 1 | 1 | 0% | 1,204 | 2,000 | +66% | 0 | 0 | — |
case-15 | pass→pass | 7,445 | 6,521 | -12% | 1 | 1 | 0% | 1,461 | 2,254 | +54% | 0 | 0 | — |
case-16 | fail→pass | 6,424 | 3,625 | -44% | 1 | 1 | 0% | 1,112 | 1,925 | +73% | 0 | 0 | — |
case-17 | fail→pass | 3,152 | 2,149 | -32% | 1 | 1 | 0% | 516 | 1,624 | +215% | 0 | 0 | — |
case-18 | fail→pass | 2,631 | 2,339 | -11% | 1 | 1 | 0% | 429 | 1,669 | +289% | 0 | 0 | — |
case-19 | pass→pass | 15,807 | 13,049 | -17% | 1 | 1 | 0% | 3,116 | 3,886 | +25% | 0 | 0 | — |
case-20 | pass→pass | 10,293 | 6,497 | -37% | 1 | 1 | 0% | 1,864 | 2,511 | +35% | 0 | 0 | — |
case-21 | pass→pass | 8,444 | 6,124 | -27% | 1 | 1 | 0% | 1,680 | 2,469 | +47% | 0 | 0 | — |
case-22 | pass→pass | 5,212 | 1,839 | -65% | 1 | 1 | 0% | 787 | 1,616 | +105% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.