Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Serverless vector database at the edge with Cloudflare Vectorize. Use when: building semantic search on Cloudflare Workers, RAG pipelines at the edge, low-latency vector similarity search, or storing and querying embeddings without managing a separate vector database.
.claude/skills/terminalskills-cloudflare-vectorize/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 7 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 159% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 60% | 0% |
Cloudflare Vectorize is a globally distributed vector database built into the Cloudflare Workers platform. It stores high-dimensional vectors (embeddings) and supports fast approximate nearest-neighbor search — all at the edge, with no separate infrastructure to manage.
Key features:
Use Wrangler CLI to create an index. Specify the embedding dimensions and distance metric:
bash# For BAAI/bge-base-en-v1.5 (768 dims, cosine similarity) npx wrangler vectorize create my-index \ --dimensions=768 \ --metric=cosine # For OpenAI text-embedding-3-small (1536 dims) npx wrangler vectorize create my-index \ --dimensions=1536 \ --metric=cosine # Euclidean and dot-product are also supported npx wrangler vectorize create my-index \ --dimensions=384 \ --metric=euclidean
wrangler.tomltomlname = "my-worker" main = "src/index.ts" compatibility_date = "2024-09-23" [[vectorize]] binding = "VECTORIZE_INDEX" index_name = "my-index"
typescriptexport interface Env { VECTORIZE_INDEX: VectorizeIndex }
Each vector needs a unique string id and a values array matching the index dimensions:
typescriptexport default { async fetch(request: Request, env: Env): Promise<Response> { const vectors: VectorizeVector[] = [ { id: "doc-001", values: [0.1, 0.2, 0.3, /* ... 768 total */], metadata: { title: "Introduction to Cloudflare", url: "/docs/intro" }, }, { id: "doc-002", values: [0.4, 0.5, 0.6, /* ... */], metadata: { title: "Workers AI Overview", url: "/docs/workers-ai" }, }, ] const result = await env.VECTORIZE_INDEX.insert(vectors) // result.count = number of vectors inserted return Response.json({ inserted: result.count }) }, }
typescriptexport default { async fetch(request: Request, env: Env): Promise<Response> { const { queryVector, topK = 5 } = await request.json() as { queryVector: number[] topK?: number } const results = await env.VECTORIZE_INDEX.query(queryVector, { topK, returnMetadata: true, // include metadata in results returnValues: false, // skip returning raw vector values }) // results.matches is sorted by score (highest = most similar) return Response.json({ matches: results.matches.map(m => ({ id: m.id, score: m.score, metadata: m.metadata, })) }) }, }
Filter results to a subset before computing similarity — useful for multi-tenant or categorized data:
typescriptconst results = await env.VECTORIZE_INDEX.query(queryVector, { topK: 10, returnMetadata: true, filter: { category: { $eq: "documentation" }, }, }) // Compound filter const filtered = await env.VECTORIZE_INDEX.query(queryVector, { topK: 5, returnMetadata: true, filter: { language: { $eq: "en" }, published: { $eq: true }, }, })
Supported filter operators: $eq, $ne, $lt, $lte, $gt, $gte, $in
Use namespaces to isolate data for different tenants or categories within a single index:
typescript// Insert with namespace await env.VECTORIZE_INDEX.insert([{ id: "tenant-a-doc-1", values: embedding, metadata: { text: "Document content..." }, namespace: "tenant-a", }]) // Query within a namespace const results = await env.VECTORIZE_INDEX.query(queryVector, { topK: 5, returnMetadata: true, namespace: "tenant-a", })
typescript// Get vectors by ID const vectors = await env.VECTORIZE_INDEX.getByIds(["doc-001", "doc-002"]) // Upsert (insert or update) await env.VECTORIZE_INDEX.upsert([{ id: "doc-001", values: newEmbedding, metadata: { updated: true }, }]) // Delete by ID await env.VECTORIZE_INDEX.deleteByIds(["doc-001", "doc-002"])
Complete RAG pipeline — embed query, search Vectorize, generate answer with LLM:
typescriptexport interface Env { AI: Ai VECTORIZE_INDEX: VectorizeIndex } export default { async fetch(request: Request, env: Env): Promise<Response> { const { question } = await request.json() as { question: string } // 1. Embed the user's question const embeddingResult = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: [question], }) const queryVector = embeddingResult.data[0] // 2. Find relevant documents const searchResults = await env.VECTORIZE_INDEX.query(queryVector, { topK: 3, returnMetadata: true, }) const context = searchResults.matches .map(m => m.metadata?.text as string) .filter(Boolean) .join("\n\n") // 3. Generate answer with context const answer = await env.AI.run("@cf/meta/llama-3-8b-instruct", { messages: [ { role: "system", content: `Answer the question using only the provided context.\n\nContext:\n${context}`, }, { role: "user", content: question }, ], max_tokens: 512, }) return Response.json({ answer: answer.response, sources: searchResults.matches.map(m => ({ id: m.id, score: m.score, url: m.metadata?.url, })), }) }, }
For indexing large document collections, batch inserts for efficiency:
typescriptasync function indexDocuments( documents: Array<{ id: string; text: string; metadata: Record<string, unknown> }>, env: Env, batchSize = 100 ) { for (let i = 0; i < documents.length; i += batchSize) { const batch = documents.slice(i, i + batchSize) // Embed batch const embeddingResult = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: batch.map(d => d.text), }) // Prepare vectors const vectors: VectorizeVector[] = batch.map((doc, idx) => ({ id: doc.id, values: embeddingResult.data[idx], metadata: { ...doc.metadata, text: doc.text }, })) // Insert batch await env.VECTORIZE_INDEX.insert(vectors) console.log(`Indexed ${i + batch.length}/${documents.length} documents`) } }
bash# List all indexes npx wrangler vectorize list # Describe an index (dimensions, metric, vector count) npx wrangler vectorize info my-index # Delete an index npx wrangler vectorize delete my-index # Get vectors by ID (for debugging) npx wrangler vectorize get-vectors my-index --ids=doc-001,doc-002
cosine distance for normalized text embeddings (BAAI, OpenAI); use euclidean or dot-product only when your model specifically recommends it.metadata so you can return it with search results without a separate database lookup.insert() call — batch larger datasets.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,514 | 8,991 | -22% | 1 | 1 | 0% | 2,383 | 4,122 | +73% | 0 | 0 | — |
case-02 | pass→pass | 5,398 | 2,991 | -45% | 1 | 1 | 0% | 1,045 | 2,760 | +164% | 0 | 0 | — |
case-03 | pass→pass | 2,748 | 2,350 | -14% | 1 | 1 | 0% | 527 | 2,685 | +409% | 0 | 0 | — |
case-08 | fail→pass | 8,534 | 5,360 | -37% | 1 | 1 | 0% | 1,826 | 3,515 | +92% | 0 | 0 | — |
case-04 | fail→pass | 10,402 | 5,754 | -45% | 1 | 1 | 0% | 1,950 | 3,382 | +73% | 0 | 0 | — |
case-05 | fail→pass | 5,719 | 3,481 | -39% | 1 | 1 | 0% | 1,141 | 2,957 | +159% | 0 | 0 | — |
case-06 | fail→pass | 13,763 | 9,036 | -34% | 1 | 1 | 0% | 2,371 | 3,795 | +60% | 0 | 0 | — |
case-07 | pass→pass | 7,070 | 5,253 | -26% | 1 | 1 | 0% | 1,495 | 3,335 | +123% | 0 | 0 | — |
case-09 | pass→pass | 3,665 | 4,052 | +11% | 1 | 1 | 0% | 759 | 3,029 | +299% | 0 | 0 | — |
case-10 | fail→pass | 5,200 | 1,718 | -67% | 1 | 1 | 0% | 810 | 2,576 | +218% | 0 | 0 | — |
case-11 | pass→pass | 2,138 | 2,175 | +2% | 1 | 1 | 0% | 306 | 2,627 | +758% | 0 | 0 | — |
case-12 | pass→pass | 10,704 | 7,839 | -27% | 1 | 1 | 0% | 1,804 | 3,627 | +101% | 0 | 0 | — |
case-13 | pass→pass | 3,074 | 2,355 | -23% | 1 | 1 | 0% | 613 | 2,684 | +338% | 0 | 0 | — |
case-14 | fail→pass | 4,043 | 2,205 | -45% | 1 | 1 | 0% | 760 | 2,645 | +248% | 0 | 0 | — |
case-15 | pass→pass | 7,715 | 2,543 | -67% | 1 | 1 | 0% | 1,237 | 2,726 | +120% | 0 | 0 | — |
case-16 | fail→pass | 9,156 | 6,740 | -26% | 1 | 1 | 0% | 1,931 | 3,430 | +78% | 0 | 0 | — |
case-17 | pass→pass | 10,159 | 2,650 | -74% | 1 | 1 | 0% | 1,437 | 2,756 | +92% | 0 | 0 | — |
case-18 | fail→pass | 6,711 | 2,853 | -57% | 1 | 1 | 0% | 1,211 | 2,690 | +122% | 0 | 0 | — |
case-19 | pass→pass | 3,476 | 3,649 | +5% | 1 | 1 | 0% | 631 | 2,793 | +343% | 0 | 0 | — |
case-20 | pass→pass | 3,088 | 1,765 | -43% | 1 | 1 | 0% | 496 | 2,532 | +410% | 0 | 0 | — |
case-21 | fail→pass | 10,832 | 9,838 | -9% | 1 | 1 | 0% | 2,182 | 3,978 | +82% | 0 | 0 | — |
case-22 | pass→pass | 8,198 | 3,929 | -52% | 1 | 1 | 0% | 1,764 | 3,003 | +70% | 0 | 0 | — |
case-23 | pass→pass | 8,608 | 6,322 | -27% | 1 | 1 | 0% | 1,732 | 3,566 | +106% | 0 | 0 | — |
case-24 | pass→pass | 6,780 | 4,229 | -38% | 1 | 1 | 0% | 1,404 | 3,023 | +115% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +42 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.