Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Vector search via embeddings_* (large-scale HNSW) and ruvllm_hnsw_* (WASM router for ≤11 hot patterns), with RaBitQ 1-bit quantization for 32× memory reduction
.claude/skills/ruvnet-vector-search/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 49% | 0% |
Two distinct vector-search paths live in this plugin. Pick the right one — they're not interchangeable.
| Path | Tool family | Backing | Capacity | Latency | |------|-------------|---------|----------|---------| | Large-scale corpus | embeddings_* | @claude-flow/memory HNSW (Rust/Native) | up to millions of vectors | ~1.9× at N=20k, ~3.2×–4.7× at N=5k vs brute-force (measured; recall@10 ≈ 0.99). ANN wins above the crossover | | Hot-path router | ruvllm_hnsw_* | WASM-backed router (v2.0.1) | ~11 patterns max (ruvllm-tools.ts:58) | sub-ms; designed for high-priority routing, not corpus search |
The "12,500×" headline applies to the large-scale embeddings_search path. The WASM router is not that path.
| Need | Path | |---|---| | Search a corpus of N ≥ 500 documents | embeddings_search | | Memory-constrained corpus (≥5,000 vectors) | RaBitQ quantized — see "Quantized search" below | | Compare two strings | embeddings_compare | | Hierarchical / taxonomic data | embeddings_hyperbolic (Poincare ball) | | Route a query to one of ≤11 hot patterns | ruvllm_hnsw_route | | Cross-namespace search | memory_search_unified |
mcp__plugin_ruflo-core_ruflo__embeddings_status to verify the embedding engine.mcp__plugin_ruflo-core_ruflo__embeddings_init if not active.mcp__plugin_ruflo-core_ruflo__embeddings_generate for text input.mcp__plugin_ruflo-core_ruflo__embeddings_search with the query.mcp__plugin_ruflo-core_ruflo__embeddings_compare to measure similarity.mcp__plugin_ruflo-core_ruflo__memory_search_unified for cross-namespace.For corpora ≥5,000 vectors and/or memory-constrained environments, use the RaBitQ 1-bit quantization workflow. Below 5,000 vectors the rebuild cost outweighs the savings — use the standard path instead.
| Step | Tool | Purpose | |---|---|---| | 1 | embeddings_init | Engine warm | | 2 | embeddings_rabitq_build | One-time build of the 1-bit index after corpus is loaded | | 3 | embeddings_rabitq_search | Hamming-prefilter returns top-N candidate IDs (cheap) | | 4 | embeddings_search | Optional exact rerank on the candidate set (full-precision) | | 5 | embeddings_rabitq_status | Index health, memory footprint, build time |
> Note: embeddings_rabitq_search returns candidate IDs only — the rerank in step 4 is the user's responsibility (mirrors the docstring at embeddings-tools.ts:911). Without rerank, results are approximate; with rerank, you get full-precision quality at 32× lower memory.
HNSW exposes three knobs that trade recall against latency. The "12,500×" headline assumes defaults; tune deliberately for your workload:
| Profile | efSearch | M | When to use | |---------|-----------|-----|-------------| | recall-first | 200 | 32 | Pattern recall during planning; quality matters more than ms | | balanced (default) | 64 | 16 | General-purpose semantic recall | | latency-first | 16 | 8 | Hot-path routing where p99 latency matters |
efSearch is passed via ruvllm_hnsw_create (ruvllm-tools.ts:64). M is registry-level today; raise as a follow-up if it should be MCP-tunable. efConstruction defaults to 200 in the lite index (hnsw-index.ts:537).
For routing a small number of high-priority patterns:
mcp__plugin_ruflo-core_ruflo__ruvllm_hnsw_create — create the WASM index (cap ~11)mcp__plugin_ruflo-core_ruflo__ruvllm_hnsw_add — add a patternmcp__plugin_ruflo-core_ruflo__ruvllm_hnsw_route — route an incoming queryThis is not a corpus index. Treat it as a fast classifier over a curated set of patterns.
For hierarchical data (code trees, org charts), use mcp__plugin_ruflo-core_ruflo__embeddings_hyperbolic which maps to Poincare ball space. Distance is geodesic, not cosine.
bashnpx @claude-flow/cli@latest embeddings search --query "authentication patterns" npx @claude-flow/cli@latest embeddings init npx @claude-flow/cli@latest memory search --query "your query"
Measured numbers (source: scripts/benchmark-intelligence.mjs, ruvector NAPI backend; recall@10 ≈ 0.99). The older "150×–12,500×" figures were brute-force-fallback artifacts and have been retired — see project CLAUDE.md "V3 Performance Targets".
| Method | Measured speedup vs brute-force | |--------|---------------------------------| | Brute-force scan | Baseline | | HNSW (N=5,000) | ~3.2×–4.7× faster | | HNSW (N=20,000) | ~1.9× faster | | HNSW (below crossover, small N) | ties/loses vs brute-force | | RaBitQ quantization | 32× memory reduction; 0.60 ms/query at N≈14.7k | | ruvllm_hnsw_route (n≤11) | sub-ms per route, fixed cost |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | fail→pass | 10,698 | 2,335 | -78% | 1 | 1 | 0% | 1,956 | 1,983 | +1% | 0 | 0 | — |
case-06 | fail→pass | 13,173 | 4,769 | -64% | 1 | 1 | 0% | 2,248 | 2,406 | +7% | 0 | 0 | — |
case-07 | fail→pass | 22,547 | 3,554 | -84% | 1 | 1 | 0% | 1,982 | 2,181 | +10% | 0 | 0 | — |
case-08 | fail→pass | 8,089 | 8,725 | +8% | 1 | 1 | 0% | 1,681 | 2,117 | +26% | 0 | 0 | — |
case-01 | fail→fail | 12,423 | 6,803 | -45% | 1 | 1 | 0% | 2,359 | 1,930 | -18% | 0 | 0 | — |
case-02 | fail→fail | 15,339 | 5,666 | -63% | 1 | 1 | 0% | 3,524 | 2,034 | -42% | 0 | 0 | — |
case-03 | fail→pass | 14,321 | 13,612 | -5% | 1 | 1 | 0% | 2,842 | 4,236 | +49% | 0 | 0 | — |
case-04 | fail→pass | 21,198 | 9,312 | -56% | 1 | 1 | 0% | 4,087 | 3,478 | -15% | 0 | 0 | — |
case-05 | fail→pass | 15,950 | 9,461 | -41% | 1 | 1 | 0% | 2,812 | 3,250 | +16% | 0 | 0 | — |
case-09 | fail→pass | 12,097 | 2,503 | -79% | 1 | 1 | 0% | 2,006 | 2,027 | +1% | 0 | 0 | — |
case-10 | fail→fail | 5,101 | 2,346 | -54% | 1 | 1 | 0% | 852 | 2,026 | +138% | 0 | 0 | — |
case-11 | fail→pass | 6,006 | 3,299 | -45% | 1 | 1 | 0% | 1,011 | 2,170 | +115% | 0 | 0 | — |
case-12 | pass→pass | 8,685 | 1,903 | -78% | 1 | 1 | 0% | 1,644 | 1,798 | +9% | 0 | 0 | — |
case-13 | fail→pass | 6,302 | 5,517 | -12% | 1 | 1 | 0% | 1,128 | 1,888 | +67% | 0 | 0 | — |
case-15 | fail→pass | 30,502 | 7,626 | -75% | 1 | 1 | 0% | 1,309 | 3,022 | +131% | 0 | 0 | — |
case-16 | fail→pass | 14,264 | 2,017 | -86% | 1 | 1 | 0% | 2,722 | 1,910 | -30% | 0 | 0 | — |
case-17 | fail→pass | 10,002 | 2,332 | -77% | 1 | 1 | 0% | 1,962 | 1,992 | +2% | 0 | 0 | — |
case-18 | pass→pass | 7,217 | 2,920 | -60% | 1 | 1 | 0% | 1,646 | 2,092 | +27% | 0 | 0 | — |
case-19 | pass→pass | 8,751 | 5,340 | -39% | 1 | 1 | 0% | 1,564 | 2,659 | +70% | 0 | 0 | — |
case-20 | pass→pass | 6,452 | 5,189 | -20% | 1 | 1 | 0% | 1,508 | 2,545 | +69% | 0 | 0 | — |
case-21 | pass→pass | 3,174 | 3,396 | +7% | 1 | 1 | 0% | 605 | 2,192 | +262% | 0 | 0 | — |
case-22 | pass→pass | 2,600 | 2,481 | -5% | 1 | 1 | 0% | 460 | 1,971 | +328% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.