Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible Responses API with typed output blocks (reasoning, message, function_call, web_search_call). Covers request shape, streaming, differences from /chat/completions, supported venice_parameters subset, and E2EE behavior.
.claude/skills/sediman-agent-venice-responses/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 120% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 81% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 46% | 0% |
POST /api/v1/responses is Venice's OpenAI-compatible Responses endpoint. It returns a structured, typed output array instead of a single message.content string — ideal for agents that need to separate reasoning, messages, tool calls, and built-in tool events.
> Alpha. Access is gated behind the responsesApiEnabled flag on Bearer API keys (staff-only during beta). x402 wallet auth bypasses this flag — you can pay per request without the flag. Schemas may change.
output[] with typed type: "reasoning" | "message" | "function_call" | "web_search_call" blocks) for a client library that expects it.Otherwise use venice-chat — it has more features, more models, and full Venice parameters.
/chat/completions| Limitation | Detail | |---|---| | Stateless | No conversation persistence across requests. Send the full history each call. | | E2EE models default to rejection | E2EE-capable models return 400 unless you pass venice_parameters.enable_e2ee: false (TEE-only mode). For end-to-end encrypted inference with E2EE headers, use /chat/completions. | | Subset of venice_parameters | character_slug, enable_e2ee, enable_web_search, enable_web_scraping, enable_web_citations, include_venice_system_prompt, include_search_results_in_stream are supported. strip_thinking_response, disable_thinking, enable_x_search are not wired through in Alpha. | | Access gated by feature flag | Bearer keys without responsesApiEnabled get 401. x402 requests are allowed (pay-per-call). |
Same as the rest of the API — either Authorization: Bearer <key> or X-Sign-In-With-X: <SIWE>. See venice-auth.
bashcurl https://api.venice.ai/api/v1/responses \ -H "Authorization: Bearer $VENICE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "zai-org-glm-5-1", "input": "Explain why the sky is blue in one paragraph." }'
input accepts:
chat/completions message parts) for multi-turn or multimodal history.json{ "id": "resp_abc123", "object": "response", "created_at": 1735689600, "model": "zai-org-glm-5-1", "status": "completed", "output": [ { "type": "reasoning", "id": "rs_1", "summary": ["I considered Rayleigh scattering..."], "encrypted_content": "..." }, { "type": "message", "id": "msg_1", "status": "completed", "role": "assistant", "content": [{ "type": "output_text", "text": "The sky is blue because...", "annotations": [{ "type": "url_citation", "url": "https://example.com/rayleigh", "title": "Rayleigh scattering", "start_index": 42, "end_index": 99 }] }] }, { "type": "function_call", "id": "fc_1", "call_id": "call_abc", "name": "get_weather", "arguments": "{\"city\":\"Paris\"}", "status": "completed" }, { "type": "web_search_call", "id": "ws_1", "status": "completed" } ], "usage": { "input_tokens": 20, "input_tokens_details": {"cached_tokens": 0}, "output_tokens": 80, "output_tokens_details": {"reasoning_tokens": 40}, "total_tokens": 100 } }
Top-level status ∈ completed | failed | in_progress | cancelled. On failed, error.code and error.message are populated.
| type | Purpose | |---|---| | reasoning | Thought process from reasoning models. summary[] holds human-readable text; encrypted_content holds opaque signatures — round-trip verbatim for multi-turn tool calls. | | message | Main text output. content[].type === "output_text", plus annotations[] for url_citation entries from web search. | | function_call | Tool call: name, stringified-JSON arguments, call_id. | | web_search_call | Sentinel showing the built-in web_search tool fired; use alongside url_citation annotations on messages. |
Match tool outputs back by call_id when continuing the turn.
| Field | Notes | |---|---| | model | Required. Model ID, trait, or compatibility mapping. Feature suffixes allowed (see venice-chat). | | input | Required. String or input-items array. To set system/developer context, include a leading message with role: "system"/"developer" in the input array. | | tools | Array of {type:"function",function:{...}} or built-in {type:"web_search"} — availability depends on the model. | | tool_choice | "auto" / "required" / "none" / {type:"function",function:{"name":"..."}}. | | reasoning.effort | Reasoning effort hint for thinking models ("low" \| "medium" \| "high"). | | temperature, top_p, max_output_tokens, n, stop, seed, prompt_cache_key | Standard generation controls — translated to /chat/completions equivalents server-side. | | stream | Boolean. SSE response with typed events (response.created, response.output_item.added, response.output_text.delta, response.completed, …). | | venice_parameters | Subset listed above. Example: {"character_slug":"alan-watts","enable_web_search":"on"}. |
Fields commonly found in OpenAI's Responses API that are not in Venice's Alpha schema (and silently ignored or rejected by Zod): instructions, metadata, parallel_tool_calls, response_format, store, previous_response_id, background. For response_format / JSON-schema structured output, use /chat/completions.
With stream: true, the response is an SSE stream of typed events. Typical flow:
event: response.created
event: response.output_item.added # type=reasoning
event: response.reasoning.delta
event: response.output_item.added # type=message
event: response.content_part.added
event: response.output_text.delta
event: response.output_text.delta
event: response.output_item.done
event: response.completedConsume events in order and reconstruct output[] client-side; the shape on response.completed matches the non-streamed response exactly.
400 — bad request; also returned when an E2EE-capable model is used without venice_parameters.enable_e2ee: false.401 — auth failed, or Bearer key lacks responsesApiEnabled, or the model is Pro-only and you're on an INFERENCE key / x402 wallet.402 — insufficient balance. Bearer → { error: "INSUFFICIENT_BALANCE" }. x402 → PAYMENT_REQUIRED with topUpInstructions and siwxChallenge (see venice-x402).429 — rate-limited.500 — inference failed.X-Balance-Remaining is on 200 responses when using x402 auth; PAYMENT-REQUIRED header on 402.
messages → pass as input (string, or typed array with leading {role:"system"|"developer", content:"..."}).venice_parameters.character_slug → supported; pass inside venice_parameters or as a model feature suffix (:character_slug=alan-watts).venice_parameters.enable_web_search → pass inside venice_parameters, or append :enable_web_search=on to the model ID, or add {"type":"web_search"} to tools.venice_parameters.strip_thinking_response / disable_thinking → not supported on /responses in Alpha; stay on /chat/completions for these./chat/completions. For TEE-only inference on an E2EE-capable model, pass venice_parameters.enable_e2ee: false here.response_format / JSON-schema structured output → stay on /chat/completions.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | pass→pass | 8,208 | 3,628 | -56% | 1 | 1 | 0% | 1,315 | 3,023 | +130% | 0 | 0 | — |
case-01 | fail→pass | 11,535 | 9,594 | -17% | 1 | 1 | 0% | 1,958 | 4,300 | +120% | 0 | 0 | — |
case-02 | fail→fail | 10,815 | 3,834 | -65% | 1 | 1 | 0% | 1,793 | 3,060 | +71% | 0 | 0 | — |
case-03 | pass→pass | 11,106 | 5,032 | -55% | 1 | 1 | 0% | 1,823 | 3,262 | +79% | 0 | 0 | — |
case-04 | fail→pass | 12,734 | 3,874 | -70% | 1 | 1 | 0% | 1,960 | 3,132 | +60% | 0 | 0 | — |
case-05 | fail→pass | 11,602 | 3,938 | -66% | 1 | 1 | 0% | 2,074 | 3,149 | +52% | 0 | 0 | — |
case-06 | pass→pass | 13,500 | 4,834 | -64% | 1 | 1 | 0% | 2,203 | 3,266 | +48% | 0 | 0 | — |
case-08 | pass→pass | 10,870 | 3,009 | -72% | 1 | 1 | 0% | 1,871 | 2,896 | +55% | 0 | 0 | — |
case-09 | fail→pass | 10,322 | 4,069 | -61% | 1 | 1 | 0% | 1,690 | 3,053 | +81% | 0 | 0 | — |
case-10 | fail→pass | 13,449 | 4,476 | -67% | 1 | 1 | 0% | 2,153 | 3,143 | +46% | 0 | 0 | — |
case-11 | fail→pass | 10,383 | 5,207 | -50% | 1 | 1 | 0% | 1,844 | 3,351 | +82% | 0 | 0 | — |
case-12 | fail→pass | 12,502 | 4,029 | -68% | 1 | 1 | 0% | 1,966 | 3,064 | +56% | 0 | 0 | — |
case-13 | fail→pass | 15,123 | 3,213 | -79% | 1 | 1 | 0% | 2,246 | 2,905 | +29% | 0 | 0 | — |
case-14 | pass→pass | 6,489 | 3,085 | -52% | 1 | 1 | 0% | 991 | 2,863 | +189% | 0 | 0 | — |
case-15 | fail→pass | 12,763 | 4,316 | -66% | 1 | 1 | 0% | 2,049 | 3,128 | +53% | 0 | 0 | — |
case-16 | fail→pass | 9,130 | 3,006 | -67% | 1 | 1 | 0% | 1,374 | 2,901 | +111% | 0 | 0 | — |
case-17 | pass→pass | 11,652 | 5,540 | -52% | 1 | 1 | 0% | 1,949 | 3,419 | +75% | 0 | 0 | — |
case-18 | pass→pass | 4,394 | 2,804 | -36% | 1 | 1 | 0% | 759 | 2,837 | +274% | 0 | 0 | — |
case-19 | pass→pass | 9,646 | 1,800 | -81% | 1 | 1 | 0% | 1,544 | 2,616 | +69% | 0 | 0 | — |
case-20 | pass→pass | 10,863 | 3,680 | -66% | 1 | 1 | 0% | 1,668 | 2,933 | +76% | 0 | 0 | — |
case-21 | fail→pass | 16,460 | 4,043 | -75% | 1 | 1 | 0% | 2,684 | 3,147 | +17% | 0 | 0 | — |
case-22 | fail→pass | 13,069 | 3,982 | -70% | 1 | 1 | 0% | 1,951 | 3,112 | +60% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.