Install any skill in seconds. Free to start, no credit card required.
Get Started Free →You are an expert in Cloudflare Workers AI, the serverless AI inference platform running on Cloudflare's global network. You help developers run LLMs, embedding models, image generation, speech-to-text, and translation models at the edge with zero cold starts, pay-per-use pricing, and integration with Workers, Pages, and Vectorize — enabling AI features without managing GPU infrastructure.
.claude/skills/terminalskills-cloudflare-ai/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 163% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 84% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 53% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 75% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 39% | 0% |
You are an expert in Cloudflare Workers AI, the serverless AI inference platform running on Cloudflare's global network. You help developers run LLMs, embedding models, image generation, speech-to-text, and translation models at the edge with zero cold starts, pay-per-use pricing, and integration with Workers, Pages, and Vectorize — enabling AI features without managing GPU infrastructure.
typescript// src/worker.ts — AI-powered API at the edge export default { async fetch(request: Request, env: Env): Promise<Response> { const url = new URL(request.url); // Text generation (LLM) if (url.pathname === "/api/chat") { const { messages } = await request.json(); const response = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { messages, max_tokens: 1024, temperature: 0.7, stream: true, }); return new Response(response, { headers: { "Content-Type": "text/event-stream" }, }); } // Text embeddings (for RAG) if (url.pathname === "/api/embed") { const { text } = await request.json(); const embeddings = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: Array.isArray(text) ? text : [text], }); return Response.json({ embeddings: embeddings.data }); } // Image generation if (url.pathname === "/api/generate-image") { const { prompt } = await request.json(); const image = await env.AI.run("@cf/stabilityai/stable-diffusion-xl-base-1.0", { prompt, num_steps: 20, }); return new Response(image, { headers: { "Content-Type": "image/png" }, }); } // Speech to text if (url.pathname === "/api/transcribe") { const audioData = await request.arrayBuffer(); const result = await env.AI.run("@cf/openai/whisper", { audio: [...new Uint8Array(audioData)], }); return Response.json({ text: result.text }); } // Translation if (url.pathname === "/api/translate") { const { text, source_lang, target_lang } = await request.json(); const result = await env.AI.run("@cf/meta/m2m100-1.2b", { text, source_lang, target_lang, }); return Response.json({ translated: result.translated_text }); } return new Response("Not Found", { status: 404 }); }, };
typescript// RAG pipeline: Embed → Store in Vectorize → Query → Generate export default { async fetch(request: Request, env: Env): Promise<Response> { const { question } = await request.json(); // Step 1: Embed the question const queryEmbedding = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: [question], }); // Step 2: Search Vectorize const matches = await env.VECTORIZE.query(queryEmbedding.data[0], { topK: 5, returnMetadata: "all", }); // Step 3: Generate answer with context const context = matches.matches.map(m => m.metadata?.text).join("\n\n"); const answer = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { messages: [ { role: "system", content: `Answer based on this context:\n${context}` }, { role: "user", content: question }, ], }); return Response.json({ answer: answer.response, sources: matches.matches.map(m => ({ text: m.metadata?.text, score: m.score })), }); }, };
bash# Create Workers project npm create cloudflare@latest my-ai-app # wrangler.toml [ai] binding = "AI" [[vectorize]] binding = "VECTORIZE" index_name = "my-index" # Deploy npx wrangler deploy
stream: true for LLM responses; first token in ~200ms at the edge@cf/ models; Llama 3.1, Mistral, Stable Diffusion, Whisper, BGE all available| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 7,465 | 6,796 | -9% | 1 | 1 | 0% | 1,413 | 2,603 | +84% | 0 | 0 | — |
case-02 | pass→pass | 11,566 | 9,494 | -18% | 1 | 1 | 0% | 2,018 | 3,093 | +53% | 0 | 0 | — |
case-03 | pass→pass | 7,052 | 5,002 | -29% | 1 | 1 | 0% | 1,267 | 2,212 | +75% | 0 | 0 | — |
case-04 | pass→pass | 10,965 | 9,077 | -17% | 1 | 1 | 0% | 2,359 | 3,276 | +39% | 0 | 0 | — |
case-05 | pass→pass | 4,161 | 3,827 | -8% | 1 | 1 | 0% | 754 | 2,060 | +173% | 0 | 0 | — |
case-06 | pass→pass | 9,627 | 9,044 | -6% | 1 | 1 | 0% | 2,080 | 3,131 | +51% | 0 | 0 | — |
case-07 | fail→fail | 13,188 | 9,464 | -28% | 1 | 1 | 0% | 2,863 | 3,333 | +16% | 0 | 0 | — |
case-08 | pass→pass | 4,394 | 4,693 | +7% | 1 | 1 | 0% | 745 | 2,166 | +191% | 0 | 0 | — |
case-09 | pass→pass | 11,765 | 7,303 | -38% | 1 | 1 | 0% | 2,531 | 2,760 | +9% | 0 | 0 | — |
case-10 | pass→pass | 12,478 | 11,068 | -11% | 1 | 1 | 0% | 2,771 | 3,572 | +29% | 0 | 0 | — |
case-11 | fail→pass | 4,091 | 4,345 | +6% | 1 | 1 | 0% | 827 | 2,171 | +163% | 0 | 0 | — |
case-12 | pass→pass | 2,979 | 4,040 | +36% | 1 | 1 | 0% | 503 | 1,985 | +295% | 0 | 0 | — |
case-13 | pass→pass | 6,449 | 3,566 | -45% | 1 | 1 | 0% | 1,162 | 1,975 | +70% | 0 | 0 | — |
case-14 | pass→pass | 12,509 | 9,235 | -26% | 1 | 1 | 0% | 2,316 | 3,057 | +32% | 0 | 0 | — |
case-15 | pass→pass | 10,633 | 10,190 | -4% | 1 | 1 | 0% | 1,939 | 3,261 | +68% | 0 | 0 | — |
case-16 | pass→pass | 9,988 | 8,191 | -18% | 1 | 1 | 0% | 1,702 | 2,897 | +70% | 0 | 0 | — |
case-17 | pass→pass | 5,798 | 4,076 | -30% | 1 | 1 | 0% | 1,001 | 2,001 | +100% | 0 | 0 | — |
case-18 | pass→pass | 6,315 | 5,961 | -6% | 1 | 1 | 0% | 1,195 | 2,380 | +99% | 0 | 0 | — |
case-19 | pass→pass | 6,625 | 4,807 | -27% | 1 | 1 | 0% | 1,213 | 2,304 | +90% | 0 | 0 | — |
case-20 | pass→pass | 4,272 | 6,445 | +51% | 1 | 1 | 0% | 830 | 2,045 | +146% | 0 | 0 | — |
case-21 | pass→pass | 8,889 | 8,361 | -6% | 1 | 1 | 0% | 1,593 | 2,764 | +74% | 0 | 0 | — |
case-22 | pass→pass | 6,407 | 5,685 | -11% | 1 | 1 | 0% | 1,193 | 2,414 | +102% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.