Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill guides you through building a production RAG pipeline or persistent
.claude/skills/pinecone-rag/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 46 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | — | — |
| case-11 | ✗→✓ | ▲ Improved | — | — |
| case-12 | ✗→✓ | ▲ Improved | — | — |
| case-17 | ✗→✓ | ▲ Improved | — | — |
| case-08 | ✗→✓ | ▲ Improved | — | — |
This skill guides you through building a production RAG pipeline or persistent agent memory system using Pinecone. Follow the workflow from start to finish — don't skip steps or jump to code before understanding what the user actually needs.
Before writing any code, identify which of these two use cases applies:
A — RAG over documents: User wants to index a corpus (PDFs, docs, code, web pages) and retrieve relevant chunks to ground LLM responses.
B — Agent memory: User wants an agent to remember facts, decisions, or context across sessions or across multiple agents sharing a knowledge base.
The setup is similar but the namespace strategy and retrieval patterns differ. If the user hasn't said, ask: "Is this for document retrieval, agent memory, or both?" Then follow the relevant workflow below.
Pick the index type before writing any code. Getting this wrong means re-creating the index later.
Serverless (recommended for most cases)
pythonfrom pinecone import Pinecone, ServerlessSpec pc = Pinecone(api_key="PINECONE_API_KEY") if "my-index" not in pc.list_indexes().names(): pc.create_index( name="my-index", dimension=1536, # must match your embedding model exactly metric="cosine", spec=ServerlessSpec(cloud="aws", region="us-east-1") ) index = pc.Index("my-index")
Pod-based (for consistent high-throughput production)
pythonfrom pinecone import PodSpec pc.create_index( name="my-index-prod", dimension=1536, metric="cosine", spec=PodSpec(environment="us-east1-gcp", pod_type="p1.x1") )
Dimension quick reference — match this exactly to your embedding model: | Model | Dimension | |---|---| | text-embedding-3-small | 1536 | | text-embedding-3-large | 3072 | | voyage-3 / voyage-multimodal-3 | 1024 | | BAAI/bge-large-en-v1.5 | 1024 | | intfloat/multilingual-e5-large (Arabic, Malay, Chinese) | 1024 |
> Checkpoint: Index exists, dimension matches embedding model, index.describe_index_stats() returns without error.
Always batch upserts — never upsert one vector at a time.
pythonfrom openai import OpenAI client = OpenAI() def embed(texts: list[str]) -> list[list[float]]: res = client.embeddings.create(model="text-embedding-3-small", input=texts) return [r.embedding for r in res.data] def upsert_docs(index, docs: list[dict], namespace: str = "default"): """docs = [{"id": "...", "text": "...", "metadata": {...}}]""" BATCH = 100 for i in range(0, len(docs), BATCH): batch = docs[i:i + BATCH] vecs = [ { "id": d["id"], "values": emb, "metadata": {**d.get("metadata", {}), "text": d["text"]} } for d, emb in zip(batch, embed([d["text"] for d in batch])) ] index.upsert(vectors=vecs, namespace=namespace)
Always store the original text in metadata — this avoids a second lookup at retrieval time.
> Checkpoint: index.describe_index_stats() shows vector count > 0 in the > target namespace.
pythondef search(index, query: str, top_k: int = 5, namespace: str = "default", filter: dict = None) -> list[dict]: [q_emb] = embed([query]) results = index.query( vector=q_emb, top_k=top_k, namespace=namespace, include_metadata=True, filter=filter ) return [{"text": m.metadata["text"], "score": m.score, "id": m.id} for m in results.matches]
Use hybrid when the domain has precise terms that semantic search misses: legal citations, medical codes, product SKUs, API method names.
pythonfrom pinecone_text.sparse import BM25Encoder bm25 = BM25Encoder().default() bm25.fit([d["text"] for d in docs]) # fit once on your corpus def hybrid_search(index, query: str, top_k: int = 5, alpha: float = 0.7): """alpha=1.0 is pure dense; alpha=0.0 is pure sparse.""" dense = [v * alpha for v in embed([query])[0]] sparse_raw = bm25.encode_queries(query) sparse = { "indices": sparse_raw["indices"], "values": [v * (1 - alpha) for v in sparse_raw["values"]] } return index.query(vector=dense, sparse_vector=sparse, top_k=top_k, include_metadata=True).matches
python# Exact match results = index.query(vector=emb, filter={"source": {"$eq": "confluence"}}) # Combined filter results = index.query(vector=emb, filter={ "$and": [ {"category": {"$eq": "engineering"}}, {"language": {"$in": ["en", "ar"]}} ] })
> Checkpoint: A test query returns relevant results with scores > 0.7 for > clearly matching content.
pythondef rag_answer(index, question: str, namespace: str = "default", model: str = "gpt-4o-mini") -> str: hits = search(index, question, top_k=5, namespace=namespace) context = "\n\n".join(h["text"] for h in hits) return client.chat.completions.create( model=model, messages=[ { "role": "system", "content": ( "Answer using only the provided context. " "If the answer isn't in the context, say so.\n\n" f"Context:\n{context}" ) }, {"role": "user", "content": question} ] ).choices[0].message.content
Use namespaces to isolate each agent's or user's memories completely. Namespace per agent prevents memory bleed across users or sessions.
pythonimport time, hashlib def remember(index, agent_id: str, content: str, memory_type: str = "fact"): """Store a memory for an agent.""" mem_id = hashlib.md5( f"{agent_id}{content}{time.time()}".encode() ).hexdigest() [emb] = embed([content]) index.upsert( vectors=[{ "id": mem_id, "values": emb, "metadata": { "text": content, "type": memory_type, "timestamp": time.time(), "agent_id": agent_id } }], namespace=f"agent_{agent_id}" ) def recall(index, agent_id: str, query: str, top_k: int = 5) -> list[str]: """Recall relevant memories for an agent.""" return [h["text"] for h in search(index, query, top_k=top_k, namespace=f"agent_{agent_id}")] def forget(index, agent_id: str): """Wipe all memories for an agent (e.g., on user request).""" index.delete(delete_all=True, namespace=f"agent_{agent_id}")
Run a quick smoke test before integrating into the larger system:
python# Smoke test upsert_docs(index, [ {"id": "t1", "text": "Pinecone is a vector database for semantic search."}, {"id": "t2", "text": "RAG combines retrieval with language model generation."}, ]) hits = search(index, "What is Pinecone?") assert hits[0]["score"] > 0.7, f"Expected high similarity, got {hits[0]['score']}" print("Smoke test passed:", hits[0]["text"])
> Checkpoint: Smoke test passes. End-to-end: index → upsert → query → > LLM response works without errors.
len(embed(["test"])[0]) matchesthe index dimension before your first upsert.
"text" in metadata,you'll need a second lookup to get the actual content at query time.
prevents cross-tenant data leaks that are hard to fix later.
build good term frequencies. Fit on at least a few hundred documents.
Use a different approach when:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.