Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Designs and implements production-grade RAG systems by chunking documents, generating embeddings, configuring vector stores, building hybrid search pipelines, applying reranking, and evaluating retrieval quality. Use when building RAG systems, vector databases, or knowledge-grounded AI applications requiring semantic search, document retrieval, context augmentation, similarity search, or embedding-based indexing.
.claude/skills/jeffallan-rag-architect/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 121% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 17% | 0% |
For each step, validate before moving on (see checkpoints below).
Load detailed guidance based on context:
| Topic | Reference | Load When | |-------|-----------|-----------| | Vector Databases | references/vector-databases.md | Comparing Pinecone, Weaviate, Chroma, pgvector, Qdrant | | Embedding Models | references/embedding-models.md | Selecting embeddings, fine-tuning, dimension trade-offs | | Chunking Strategies | references/chunking-strategies.md | Document splitting, overlap, semantic chunking | | Retrieval Optimization | references/retrieval-optimization.md | Hybrid search, reranking, query expansion, filtering | | RAG Evaluation | references/rag-evaluation.md | Metrics, evaluation frameworks, debugging retrieval |
pythonfrom langchain.text_splitter import RecursiveCharacterTextSplitter # Evaluate chunk_size on your domain data — never use 512 blindly splitter = RecursiveCharacterTextSplitter( chunk_size=800, chunk_overlap=100, separators=["\n\n", "\n", ". ", " "], ) chunks = splitter.create_documents( texts=[doc.page_content for doc in raw_docs], metadatas=[{"source": doc.metadata["source"], "timestamp": doc.metadata.get("timestamp")} for doc in raw_docs], )
Checkpoint: assert all(c.metadata.get("source") for c in chunks), "Missing source metadata"
pythonfrom openai import OpenAI import qdrant_client from qdrant_client.models import VectorParams, Distance, PointStruct client = OpenAI() qdrant = qdrant_client.QdrantClient("localhost", port=6333) # Create collection qdrant.recreate_collection( collection_name="knowledge_base", vectors_config=VectorParams(size=1536, distance=Distance.COSINE), ) def embed_chunks(chunks: list[str], model: str = "text-embedding-3-small") -> list[list[float]]: response = client.embeddings.create(input=chunks, model=model) return [r.embedding for r in response.data] # Idempotent upsert with deduplication via deterministic IDs import hashlib, uuid points = [] for i, chunk in enumerate(chunks): doc_id = str(uuid.UUID(hashlib.md5(chunk.page_content.encode()).hexdigest())) embedding = embed_chunks([chunk.page_content])[0] points.append(PointStruct(id=doc_id, vector=embedding, payload=chunk.metadata)) qdrant.upsert(collection_name="knowledge_base", points=points)
Checkpoint: assert qdrant.count("knowledge_base").count == len(set(p.id for p in points)), "Deduplication failed"
pythonfrom qdrant_client.models import Filter, FieldCondition, MatchValue, SparseVector from rank_bm25 import BM25Okapi def hybrid_search(query: str, tenant_id: str, top_k: int = 20) -> list: # Dense retrieval query_embedding = embed_chunks([query])[0] tenant_filter = Filter(must=[FieldCondition(key="tenant_id", match=MatchValue(value=tenant_id))]) dense_results = qdrant.search( collection_name="knowledge_base", query_vector=query_embedding, query_filter=tenant_filter, limit=top_k, ) # Sparse retrieval (BM25) corpus = [r.payload.get("text", "") for r in dense_results] bm25 = BM25Okapi([doc.split() for doc in corpus]) bm25_scores = bm25.get_scores(query.split()) # Reciprocal Rank Fusion ranked = sorted( zip(dense_results, bm25_scores), key=lambda x: 0.6 * x[0].score + 0.4 * x[1], reverse=True, ) return [r for r, _ in ranked[:top_k]]
Checkpoint: assert len(hybrid_search("test query", tenant_id="demo")) > 0, "Hybrid search returned no results"
Load provider API keys from environment variables or a secrets manager; never commit them to source code.
pythonimport os import cohere co = cohere.Client(os.environ["COHERE_API_KEY"]) def rerank(query: str, results: list, top_n: int = 5) -> list: docs = [r.payload.get("text", "") for r in results] reranked = co.rerank(query=query, documents=docs, top_n=top_n, model="rerank-english-v3.0") return [results[r.index] for r in reranked.results]
python# Run precision@k and recall@k against a labeled evaluation set # python evaluate.py --metrics precision@10 recall@10 mrr --collection knowledge_base from ragas import evaluate from ragas.metrics import context_precision, context_recall, faithfulness, answer_relevancy from datasets import Dataset eval_dataset = Dataset.from_dict({ "question": questions, "contexts": retrieved_contexts, "answer": generated_answers, "ground_truth": ground_truth_answers, }) results = evaluate(eval_dataset, metrics=[context_precision, context_recall, faithfulness, answer_relevancy]) print(results)
Checkpoint: Target context_precision >= 0.7 and context_recall >= 0.6 before moving to LLM integration.
When designing RAG architecture, deliver:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 34,006 | 29,008 | -15% | 1 | 1 | 0% | 6,225 | 7,637 | +23% | 0 | 0 | — |
case-02 | fail→pass | 57,886 | 28,848 | -50% | 1 | 1 | 0% | 6,218 | 8,086 | +30% | 0 | 0 | — |
case-03 | fail→fail | 35,717 | 28,584 | -20% | 1 | 1 | 0% | 6,214 | 8,082 | +30% | 0 | 0 | — |
case-04 | pass→pass | 9,731 | 10,668 | +10% | 1 | 1 | 0% | 1,587 | 3,532 | +123% | 0 | 0 | — |
case-05 | pass→pass | 22,981 | 17,227 | -25% | 1 | 1 | 0% | 2,958 | 4,756 | +61% | 0 | 0 | — |
case-06 | pass→pass | 15,844 | 17,812 | +12% | 1 | 1 | 0% | 2,869 | 5,417 | +89% | 0 | 0 | — |
case-07 | fail→pass | 9,481 | 10,922 | +15% | 1 | 1 | 0% | 1,751 | 3,875 | +121% | 0 | 0 | — |
case-21 | pass→pass | 8,384 | 4,959 | -41% | 1 | 1 | 0% | 1,630 | 2,704 | +66% | 0 | 0 | — |
case-08 | pass→pass | 14,970 | 14,376 | -4% | 1 | 1 | 0% | 2,843 | 4,697 | +65% | 0 | 0 | — |
case-09 | pass→pass | 15,706 | 22,381 | +42% | 1 | 1 | 0% | 3,201 | 6,261 | +96% | 0 | 0 | — |
case-10 | pass→pass | 12,324 | 13,232 | +7% | 1 | 1 | 0% | 2,347 | 4,394 | +87% | 0 | 0 | — |
case-11 | fail→fail | 14,623 | 13,879 | -5% | 1 | 1 | 0% | 2,688 | 4,766 | +77% | 0 | 0 | — |
case-12 | fail→pass | 16,648 | 14,586 | -12% | 1 | 1 | 0% | 2,904 | 4,669 | +61% | 0 | 0 | — |
case-13 | fail→pass | 19,870 | 13,734 | -31% | 1 | 1 | 0% | 3,793 | 4,437 | +17% | 0 | 0 | — |
case-14 | fail→pass | 17,026 | 17,664 | +4% | 1 | 1 | 0% | 2,694 | 5,534 | +105% | 0 | 0 | — |
case-15 | fail→pass | 17,095 | 16,175 | -5% | 1 | 1 | 0% | 3,659 | 5,533 | +51% | 0 | 0 | — |
case-16 | fail→fail | 14,529 | 10,540 | -27% | 1 | 1 | 0% | 2,641 | 3,703 | +40% | 0 | 0 | — |
case-17 | fail→pass | 35,070 | 28,395 | -19% | 1 | 1 | 0% | 6,097 | 7,322 | +20% | 0 | 0 | — |
case-18 | pass→pass | 22,584 | 21,243 | -6% | 1 | 1 | 0% | 3,718 | 5,540 | +49% | 0 | 0 | — |
case-19 | pass→pass | 17,243 | 18,066 | +5% | 1 | 1 | 0% | 2,951 | 5,128 | +74% | 0 | 0 | — |
case-20 | fail→pass | 19,118 | 13,968 | -27% | 1 | 1 | 0% | 4,004 | 4,985 | +25% | 0 | 0 | — |
case-22 | pass→pass | 12,012 | 9,976 | -17% | 1 | 1 | 0% | 2,105 | 3,558 | +69% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.