Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Enhancement-overlay SOP for adding sparse (BM25 / keyword) retrieval alongside dense (embedding) retrieval. Activate when a calling agent is building, reviewing, or debugging a retrieval pipeline whose corpus contains exact-match tokens — identifiers, error codes, SKUs, API/function names, proper nouns, citations, rare jargon — that pure dense embedding silently misses. Encodes the single decision rule (**hybrid is traffic-driven, not theoretical: add sparse only when the query share that depend
.claude/skills/agentsope-agentsop-hybrid-retrieval/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 132% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 194% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 218% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 401% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 466% | 0% |
> Third-person operating model for a coder agent that owns retrieval recall on a > corpus where both meaning and exact tokens matter. The audience is the LLM > agent writing or reviewing retrieval code — not an end user.
> One sentence: Dense captures meaning, sparse captures exact tokens; hybrid > wins when both matter — but only fuse them when traffic actually carries > exact-match queries, and tune the blend per query type or hybrid loses to dense.
Activate this skill when any of the following holds:
error codes (ERR_SSL_PROTOCOL), SKUs / part numbers (A1-2293-X), API or function names (as_query_engine), proper nouns, legal/medical citations (42 U.S.C. § 1983), version strings, rare jargon, ticket IDs.
"the right document exists but dense retrieval ranks it below fuzzy near-misses".
dense-only (index.as_retriever(...) / similarity_search(...) with no sparse leg).
embedding model, chunk size) per the llamaindex]] optimization ladder — hybrid is the next rung.
QueryFusionRetriever,EnsembleRetriever, or vector-store-native hybrid (Qdrant/Weaviate/Pinecone).
Do not activate when:
lexical-identity share <5% — adding BM25 doubles index footprint for no gain.
(llamaindex]] Stage 3 step 4): baseline and measure before fusing.
Three principles. Violating any of them is why teams "try hybrid and conclude it didn't help".
Dense embedding models destroy lexical identity by pooling token representations: querying a specific error string yields a vector that captures "document about SSL errors" rather than "document containing this exact string". BM25 does the inverse — it scores against an inverted index of exact tokens and is blind to synonyms and paraphrase. (TianPan, Hybrid search in production, 2026; cited in llamaindex]] Dilemma 2.)
> Operational corollary: the symptom "I pasted the exact code and got nothing" > is not a bug in the embedding model — it is the embedding model working as > designed. The fix is a second retriever that indexes tokens, not a better > embedding.
Whether to add sparse is decided by the query-type distribution of real traffic, not by a belief that "more retrievers = better". The decision threshold is the lexical share: the fraction of queries whose correct answer hinges on an exact token. Below ~5% → dense-only. 5–50% → hybrid. Above ~50% (code, logs, legal) → invert to sparse-first with dense as a rerank signal. (Cited in llamaindex]] OP-04 / Dilemma 2.)
alpha is the dense↔sparse blend (alpha=1 → pure dense, alpha=0 → pure sparse). A semantic query wants high alpha; a lexical-identity query wants low alpha. A single global alpha picked to help lexical queries hurts the semantic slice — which is exactly why a flat alpha "loses to pure dense" and teams wrongly conclude hybrid failed. Tune alpha per query type, or route per type and pick alpha per route. (LlamaIndex alpha-tuning blog; llamaindex]] Dilemma 2 结果.)
> The fusion method (RRF vs weighted) and the fusion parameter (alpha) are two > separate decisions. RRF is score-scale-robust and parameter-light; weighted fusion > is tunable but requires score normalization. Choosing neither deliberately is the > third silent failure.
Four stages. Each gates the next. This overlay assumes a working dense baseline + eval loop already exists (see llamaindex]] Stages 1–2); do not start here.
ABC-123", "the function named foo_bar") / mixed.
< 5% → dense-only. Stop; hybrid is over-engineering. Record the decision.5–50% → add hybrid (Stages 1–3).> 50% (code search, log search, legal citation lookup) → invert: BM25-first,dense as a fallback / rerank signal.
that never appears verbatim in any chunk can't be recovered by BM25 either).
Artifact: a 3-line decision note co-located with the retriever recording the lexical share and the chosen shape.
pythonfrom llama_index.core.retrievers import QueryFusionRetriever from llama_index.retrievers.bm25 import BM25Retriever dense = index.as_retriever(similarity_top_k=10) # embedding leg sparse = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=10)
Hard rules:
set silently drops recall (it can only return tokens it has).
candidates to combine. Narrow after fusion (and after any reranker).
alongside the index, or use a vector store with native hybrid (Stage 1-alt).
Stage 1-alt — vector-store-native hybrid. If the vector store does sparse internally (Qdrant sparse_vectors, Weaviate hybrid(alpha=...), Pinecone sparse-dense, pgvector + ts_rank), prefer it: one round-trip, one consistency domain, no separate BM25 index to keep in sync. Use the framework fusion path only when the store has no native hybrid.
Pick the fusion method deliberately (OP-02):
python# Reciprocal Rank Fusion — score-scale-robust, parameter-light. Good default. fused = QueryFusionRetriever( [dense, sparse], mode="reciprocal_rerank", # RRF: combine by rank, ignore raw score scales similarity_top_k=8, num_queries=1, # set >1 to also fan out query rewrites use_async=True, ) # Relative-score / weighted fusion — tunable blend; requires score normalization. fused = QueryFusionRetriever( [dense, sparse], mode="relative_score", retriever_weights=[0.6, 0.4], # ~ alpha=0.6 toward dense similarity_top_k=8, )
Default to RRF unless you have a labeled set to tune weights on — RRF sidesteps the dense-cosine vs BM25-score scale mismatch that breaks naive weighted sums.
Stage-0 sample.
retriever_weights) at {0.0, 0.25, 0.5, 0.75, 1.0} on eachtype separately, scoring recall / MRR / hit-rate per type.
alpha (dense-leaning), mixed in between.
route by query type and apply per-route alpha (hand off classification to agentsop-query-routing]] if present), or split into two retrievers selected per query.
slice vs the dense baseline. If it regresses semantic, the alpha is wrong, not hybrid.
> The most common false negative: tuning one global alpha, watching semantic recall > drop, and reverting to dense. The correct read is "alpha was global; tune per type".
Format: Trigger / Action / Output / Evidence.
exact-match tokens, (b) traffic references them, (c) lexical share ≥5%, (d) those tokens appear verbatim in chunks. All four must hold.
outcome.
not theoretical").
mode="reciprocal_rerank") — rank-based, robust to thedense-cosine vs BM25-score scale gap, no weight to tune. Use weighted / relative-score only when you have a labeled set to fit weights and have normalized scores. Never naive-sum un-normalized scores.
QueryFusionRetriever modes;LangChain EnsembleRetriever (RRF default).
BM25Retriever.from_defaults(nodes=...) over the same nodeset as the dense index; persist/rebuild it as a deployment artifact alongside the vector index. Set per-leg similarity_top_k wide (10–20).
candidates to fuse.
BM25Retriever docs; llamaindex]] OP-04 AddHybridBM25.Prefetch +fusion, Weaviate hybrid(query, alpha=...), Pinecone sparse-dense vectors, pgvector full-text + vector) instead of a separate BM25 index. One round-trip, one consistency domain.
sparse-dense docs; llamaindex]] OP-04 ("or vendor hybrid (Qdrant/Milvus alpha)").
{0, 0.25, 0.5, 0.75, 1.0} on labeledsubsets per query type, not globally. Lexical → low alpha, semantic → high.
semantic.
Dilemma 2 ("tune per type, not globally — otherwise hybrid underperforms dense").
with the per-type alpha (or pure dense / pure sparse). Hand classification to a router (agentsop-query-routing]] if available).
mechanism to apply it.
lookup).
signal to catch paraphrase, not as the lead leg. Equivalent to alpha pinned low.
dense recovers the semantic minority.
signal"); TianPan production write-up.
rises vs dense baseline AND semantic-slice metrics do not regress. Gate merge on both. A lexical lift bought with a semantic regression is not a win.
困境: Dense is the modern default; BM25 looks like "the old keyword thing". Adding hybrid doubles the index footprint, adds a BM25 index to keep in sync, and introduces alpha tuning. Is the complexity justified, or is a better embedding model enough?
约束:
names, rare jargon (TianPan 2026, quoted in Principle 1) — and a better embedding model does not fix it; lexical loss is a property of pooled representations.
决策步骤:
(Stage 0).
<5% → dense-only; the complexity is not justified.5–50% → add hybrid; tune alpha per type.>50% (legal, code, logs) → invert: BM25-first, dense as rerank signal.{0, 0.25, 0.5, 0.75, 1.0} on labeled per-type subsets.结果: Hybrid lifts the lexical slice with no degradation on the semantic slice — if alpha is tuned per type. A single global alpha often loses to pure dense on semantic queries, which is why teams sometimes wrongly conclude "hybrid didn't help". The decision is traffic-driven, not theoretical. (Verbatim from llamaindex]] Dilemma 2.)
可提取的操作: OP-01, OP-05, OP-07, OP-08.
困境: After wiring hybrid, a single alpha must serve both a user pasting ERR_TLS_CERT_INVALID (wants exact-token match, low alpha) and a user asking "why is my connection failing?" (wants meaning, high alpha). Picking alpha=0.5 helps neither fully; picking alpha to win lexical regresses semantic, and vice versa.
约束:
alpha=1 = pure dense, alpha=0 = pure sparse; the optimum differs by querytype, not by corpus.
that underperforms pure dense on the semantic majority — the classic "hybrid hurt us" report.
rankings.
决策步骤:
pin it (cheapest).
(OP-06) — classify first, blend second.
>50% lexical corpora, skip the balancing act: invert to BM25-first (OP-07).结果: Per-type alpha (or per-type routing) makes both slices peak simultaneously. The error to avoid is treating alpha as a single global hyperparameter — that is the documented cause of "hybrid underperforms dense". (LlamaIndex alpha-tuning blog; llamaindex]] Dilemma 2.)
可提取的操作: OP-05, OP-06, OP-08.
| # | Anti-pattern | Why it's wrong | Correct move | |---|---|---|---| | A1 | Adding hybrid by default on a corpus that is purely semantic | Doubles index footprint and ops surface for no recall gain; <5% lexical traffic | OP-01: gate on lexical share; dense-only is the right answer for semantic corpora | | A2 | Picking a single global alpha for mixed traffic | Helps lexical, regresses semantic (or vice versa) → "hybrid hurt us" | OP-05/OP-06: tune alpha per query type, or route per type | | A3 | Ignoring the fusion method — naive-summing dense cosine + BM25 score | Score scales are incomparable; the larger-scale leg dominates arbitrarily | OP-02: use RRF (rank-based) or normalize before weighting | | A4 | BM25 leg built over a different/stale node set than the dense index | Sparse can only return tokens it indexed; silent recall loss | OP-03: build both legs over the same nodes; version them together | | A5 | "We added hybrid, recall on semantic dropped, hybrid is bad" | The conclusion is wrong; the alpha was global | OP-05: re-evaluate per type before reverting | | A6 | Final cut taken before fusion (narrow per-leg top_k=3) | Fusion has too few candidates to combine; near-misses never surface | Retrieve wide per leg (10–20), narrow after fusion / rerank | | A7 | Separate BM25 index when the vector store has native hybrid | Extra index to keep in sync; two consistency domains | OP-04: prefer native hybrid (Qdrant/Weaviate/Pinecone/pgvector) | | A8 | Hybrid as a substitute for a reranker (or vice versa) | Different jobs: hybrid widens candidate recall, rerank reorders | Hybrid then rerank (llamaindex]] OP-03); they compose, don't replace | | A9 | Shipping hybrid with no per-slice eval | A lexical lift can hide a semantic regression | OP-08: gate on both slices vs the dense baseline |
optimization, not a starting point (llamaindex]] Stage 3).
metadata filters — fix the cheaper knob first (llamaindex]] optimization ladder).
recover what isn't indexed; fix ingestion, not retrieval.
not a recall problem (llamaindex]] OP-03 AddReranker).
QueryFusionRetriever([...]) with no mode= set and no comment on fusion choice.alpha / retriever_weights with no per-type eval behind it.BM25Retriever.from_defaults(nodes=other_nodes) where other_nodes differs fromthe dense index's node set.
similarity_top_k equal to the final desired count (no headroom for fusion).dense_scores + bm25_scores summed without normalization.How "dense + sparse, fused" looks across the common stacks. Cross-links llamaindex]].
QueryFusionRetrieverpythonfrom llama_index.core.retrievers import QueryFusionRetriever from llama_index.retrievers.bm25 import BM25Retriever dense = index.as_retriever(similarity_top_k=10) sparse = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=10) fused = QueryFusionRetriever( [dense, sparse], mode="reciprocal_rerank", # or "relative_score" / "dist_based_score" retriever_weights=[0.6, 0.4], # used by weighted modes; ~alpha toward dense similarity_top_k=8, num_queries=1, # >1 also fans out query rewrites )
Fusion modes: reciprocal_rerank (RRF, default-recommended), relative_score, dist_based_score (the latter two are weighted). (llamaindex]] OP-04.)
EnsembleRetrieverpythonfrom langchain.retrievers import EnsembleRetriever from langchain_community.retrievers import BM25Retriever bm25 = BM25Retriever.from_texts(texts); bm25.k = 10 dense = vectorstore.as_retriever(search_kwargs={"k": 10}) ensemble = EnsembleRetriever( retrievers=[bm25, dense], weights=[0.4, 0.6], # weighted RRF blend )
EnsembleRetriever fuses with Reciprocal Rank Fusion under the hood; weights biases the RRF contribution per retriever (the LangChain analogue of alpha).
pythondef rrf(result_lists, k=60, top_k=8): scores = {} for results in result_lists: # each = ranked list of doc ids for rank, doc_id in enumerate(results): scores[doc_id] = scores.get(doc_id, 0) + 1.0 / (k + rank + 1) return sorted(scores, key=scores.get, reverse=True)[:top_k] fused = rrf([dense_ids, bm25_ids]) # rank-based, no score normalization
RRF combines by rank, so dense-cosine and BM25 raw scores never need to share a scale — the reason it is the safe default (OP-02). k≈60 is the canonical constant (Cormack et al., 2009).
| Store | Mechanism | Alpha knob | |---|---|---| | Qdrant | sparse + dense vectors, Prefetch + FusionQuery(fusion=RRF) | server-side fusion (RRF/DBSF) | | Weaviate | collection.query.hybrid(query=..., alpha=0.5) | alpha (0=BM25, 1=dense) — the canonical alpha | | Pinecone | sparse-dense vectors in one index | alpha weighting on the sparse/dense split | | pgvector | tsvector full-text + vector distance, combined in SQL | hand-weighted in the ORDER BY expression | | Milvus | hybrid search with WeightedRanker / RRFRanker | ranker choice + weights |
Prefer native hybrid when available (OP-04): one round-trip, one consistency domain, no separate BM25 index to keep in sync. Weaviate's alpha is the literal parameter the llamaindex]] alpha-tuning blog generalizes.
https://www.llamaindex.ai/blog/llamaindex-enhancing-retrieval-performance-with-alpha-tuning-in-hybrid-search-in-rag-135d0c9b8a00
QueryFusionRetriever / BM25Retriever module guides:https://developers.llamaindex.ai/python/framework/
https://developers.llamaindex.ai/python/framework/optimizing/basic_strategies/basic_strategies/
https://tianpan.co/blog/2026-04-12-hybrid-search-production-bm25-dense-embeddings
EnsembleRetriever (RRF) docs:https://python.langchain.com/docs/how_to/ensemble_retriever/
https://qdrant.tech/documentation/concepts/hybrid-queries/
alpha):https://docs.weaviate.io/weaviate/search/hybrid
https://docs.pinecone.io/guides/data/understanding-hybrid-search
references/R1-source-evidence.md — extraction provenance: which llamaindex]]SKILL / R3 passages each section, OP, and dilemma derive from.
intermediate/operation_candidates.json — machine-readable operation list.[[llamaindex]] — parent RAG SOP; this overlay extracts and standalone-izes itsStage-3 step-4 hybrid recipe + Dilemma 2 (alpha tuning per query type).
[[langchain]] — EnsembleRetriever equivalent.[[agentsop-query-routing]] — mechanism for per-query-type alpha (OP-06).| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 27,429 | 28,758 | +5% | 1 | 1 | 0% | 4,698 | 10,907 | +132% | 0 | 0 | — |
case-02 | fail→fail | 35,218 | 30,663 | -13% | 1 | 1 | 0% | 6,230 | 12,349 | +98% | 0 | 0 | — |
case-03 | pass→pass | 11,478 | 11,035 | -4% | 1 | 1 | 0% | 1,697 | 8,499 | +401% | 0 | 0 | — |
case-04 | fail→pass | 20,133 | 15,929 | -21% | 1 | 1 | 0% | 3,127 | 9,208 | +194% | 0 | 0 | — |
case-05 | pass→pass | 7,991 | 6,583 | -18% | 1 | 1 | 0% | 1,396 | 7,896 | +466% | 0 | 0 | — |
case-14 | pass→pass | 10,761 | 11,155 | +4% | 1 | 1 | 0% | 1,743 | 8,570 | +392% | 0 | 0 | — |
case-06 | pass→pass | 12,465 | 11,269 | -10% | 1 | 1 | 0% | 1,870 | 8,511 | +355% | 0 | 0 | — |
case-07 | pass→pass | 16,736 | 12,952 | -23% | 1 | 1 | 0% | 2,787 | 8,812 | +216% | 0 | 0 | — |
case-08 | pass→pass | 9,299 | 9,230 | -1% | 1 | 1 | 0% | 1,491 | 8,169 | +448% | 0 | 0 | — |
case-09 | pass→pass | 20,621 | 15,442 | -25% | 1 | 1 | 0% | 3,200 | 9,269 | +190% | 0 | 0 | — |
case-10 | pass→pass | 20,501 | 14,348 | -30% | 1 | 1 | 0% | 3,086 | 9,076 | +194% | 0 | 0 | — |
case-11 | fail→pass | 17,593 | 12,887 | -27% | 1 | 1 | 0% | 2,776 | 8,840 | +218% | 0 | 0 | — |
case-12 | pass→pass | 23,354 | 14,081 | -40% | 1 | 1 | 0% | 2,528 | 9,320 | +269% | 0 | 0 | — |
case-13 | pass→pass | 13,410 | 13,525 | +1% | 1 | 1 | 0% | 2,143 | 8,804 | +311% | 0 | 0 | — |
case-15 | pass→pass | 13,880 | 12,070 | -13% | 1 | 1 | 0% | 2,153 | 8,742 | +306% | 0 | 0 | — |
case-16 | pass→pass | 9,070 | 6,085 | -33% | 1 | 1 | 0% | 1,495 | 7,682 | +414% | 0 | 0 | — |
case-17 | pass→pass | 15,391 | 9,331 | -39% | 1 | 1 | 0% | 2,525 | 8,281 | +228% | 0 | 0 | — |
case-18 | pass→pass | 14,421 | 13,154 | -9% | 1 | 1 | 0% | 2,606 | 8,920 | +242% | 0 | 0 | — |
case-19 | pass→pass | 8,212 | 3,596 | -56% | 1 | 1 | 0% | 1,452 | 7,332 | +405% | 0 | 0 | — |
case-20 | pass→pass | 13,170 | 7,588 | -42% | 1 | 1 | 0% | 2,157 | 8,037 | +273% | 0 | 0 | — |
case-21 | pass→pass | 16,952 | 20,667 | +22% | 1 | 1 | 0% | 2,735 | 9,899 | +262% | 0 | 0 | — |
case-22 | pass→pass | 19,002 | 23,962 | +26% | 1 | 1 | 0% | 3,074 | 10,640 | +246% | 0 | 0 | — |
case-23 | pass→pass | 12,153 | 11,648 | -4% | 1 | 1 | 0% | 1,873 | 8,559 | +357% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +13 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.