Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Security-first SOP for multi-tenant RAG systems. Activate when a calling agent is building, reviewing, or debugging any retrieval pipeline whose vector store is shared across more than one user, organisation, workspace, customer, or permission scope. Encodes the single non-negotiable rule — **filter at the vector store query, never after retrieval / never after rerank** — together with the per-vendor query-time filter APIs (Pinecone namespaces + `$eq`/`$in`, Weaviate `multiTenancyConfig` + tenan
.claude/skills/agentsope-agentsop-multi-tenant-rag/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 329% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 528% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 313% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 619% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 358% | 0% |
> Third-person operating model for a coder agent that owns retrieval correctness > across tenant boundaries. The audience is the LLM agent writing or reviewing > the code — not the end user.
> One sentence: Isolation lives at the vector store query boundary, not > at the model. Anything that reaches the LLM's context window has already > leaked.
Activate this skill whenever any of the following holds:
vector_store.query,query_points, similarity_search, as_retriever().retrieve(...), raw pgvector ORDER BY embedding <-> $1) and the corpus serves more than one tenant, customer, organisation, workspace, user, or permission scope.
tenant_id, org_id, user_id, "cross-customer", "shared index", "knowledge base per team".
surfaced", "the assistant cited a doc I don't have access to", or anything that smells like cross-context bleed.
filter argument, or filtering only on the returned nodes / documents list after retrieval.
explicitly audited.
Do not activate when:
Q&A).
collection per tenant via infrastructure the application code cannot override (e.g. one Pinecone index per customer with credentials issued per-tenant). In that case the isolation lives in IAM, not in this skill.
Three principles. If a design violates any of them the system is exploitable, regardless of how good the LLM prompt is.
The boundary is the vector store query call. Whatever crosses that boundary is in the trust set. An LLM, a reranker, or a postprocessor that "filters out" foreign-tenant chunks is operating inside the breach: those tokens have already been embedded, retrieved, scored, and exposed to attacker-controlled prompts. CVE-2024-41892 (Pinecone, 2024) is the canonical demonstration — RBAC checks executed after retrieval allowed sentinel data to cross namespace boundaries before the access check fired. (See we45, CSO Online references.)
> Operational corollary: any code shaped results = vs.query(...); results > = [r for r in results if r.metadata["tenant"] == ctx.tenant] is a security > defect, not a style nit. The leak already happened — you only hid it from > the user.
A query-time filter is only enforceable if every vector carries the key. The two failure shapes:
tenant_id becomes ashared global record returned to all tenants — it matches no filter predicate that requires equality.
tenant_id derived from the document body / LLM extractionrather than from the request's authenticated session — attacker can craft document content that re-labels itself.
The tenant key must come from the authenticated session at ingestion time and be written into the immutable payload / metadata column. Treat it like a foreign key to your tenants table.
Most production-grade stacks layer two mechanisms:
| Layer | Mechanism | What it stops | |---|---|---| | Storage | Per-tenant namespace / shard / collection / RLS policy | Operator bugs, mis-routed queries, ops mistakes | | Query | filter={tenant_id: $session.tenant} on every call | App bugs, namespace selection mistakes, cross-shard joins |
Pinecone's official guidance is "namespace per tenant in serverless"; Weaviate mandates a tenant handle on every CRUD when multiTenancyConfig.enabled=true; Qdrant marks the field as is_tenant=true and requires Filter.must on read; pgvector pairs schema partitioning with RLS. The principle is the same: one layer is one bug away from failing open.
Eight stages. Each gates the next. Stop and remediate at the first failure.
Before any code change, write down:
tenant_id,org_id, workspace_id, user_id, acl_label). Pick one — do not mix.
It must be from the authenticated session / JWT claim / signed header — never from a query param, prompt, or document body.
>10k → payload partition + index is mandatory). Drives Stage 2 choice.
(PII / regulated / IP / public). Raises bar on Stages 4-6.
Artifact: a 5-line THREAT_MODEL.md snippet co-located with the retriever.
| Vendor | Recommended primitive | When | |---|---|---| | Pinecone | One namespace per tenant in a serverless index | Default; cheapest query (1 RU per 1 GB / namespace); offboarding = delete_namespace | | Weaviate | multiTenancyConfig.enabled=True on the collection | Native, per-shard isolation; mandatory tenant handle on CRUD | | Qdrant | Single collection, payload key with is_tenant=true, keyword index | Scales to millions of tenants; physical co-location, logical isolation | | Chroma | where={"tenant_id": "..."} on every query/get/update/delete | Smallest stacks; cookbook calls this the "naive multi-tenancy strategy" — accept its limits | | pgvector | Postgres schema or row-level security (RLS) on tenant_id column | When the rest of the app already lives in Postgres | | Milvus | partition_key on tenant field | Same shape as Qdrant is_tenant |
Decision rule: prefer the strongest physical primitive your tenant cardinality can afford, then always add the filter at query time anyway (Principle 3).
python# Canonical shape — vendor-agnostic node.metadata = { "tenant_id": session.tenant_id, # MUST come from authenticated session "doc_id": doc.id, "source": doc.uri, "ingested_at": now_utc(), # ... domain fields } assert node.metadata["tenant_id"], "refuse to embed without tenant_id"
Hard rules:
tenant_id is missing or empty — fail loud, donot default.
tenant_id is never taken from document content; never mutablepost-ingestion.
drift breaks $eq silently.
(create_payload_index(field="tenant_id", schema="keyword", is_tenant=True)).
collection.tenants.create([Tenant(name=t)]))before any write.
Every retrieval call must accept a tenant key from the session and pass it into the vendor's filter argument. Templates in OP-04 through OP-08.
python# Generic LlamaIndex shape — works across Pinecone, Qdrant, Weaviate, Chroma from llama_index.core.vector_stores import ( MetadataFilter, MetadataFilters, FilterOperator, ) filters = MetadataFilters(filters=[ MetadataFilter(key="tenant_id", value=session.tenant_id, operator=FilterOperator.EQ), ]) retriever = index.as_retriever(similarity_top_k=8, filters=filters)
If the framework or vendor SDK does not expose a query-time filter, switch the framework or vendor — do not patch with a post-filter.
Before merging, the PR must include — and CI must run — a test that:
would be a tenant-B document (craft the query against B's content).
If the unfiltered top-1 is not a foreign-tenant doc, the test is dishonest — rewrite the corpus to make foreign-tenant semantically closer, otherwise the test always passes vacuously.
See OP-09 for a runnable scaffold.
Layer at least one of:
API key project scope, Postgres role per tenant).
current_setting('app.current_tenant')::uuid)).
A single layer is one config bug away from open.
Every retrieval call logs:
{ts, request_id, session.tenant_id, query_hash,
vs.namespace_or_collection, filter_clause,
returned_count, returned_tenant_ids_set}Then add a synchronous assertion in the request path:
pythonassert returned_tenant_ids_set.issubset({session.tenant_id, GLOBAL_TENANT})
This converts a silent leak into a loud 500 — the right failure mode.
Walk every component after retrieval and confirm none of them:
cache_key = hash(query) is a leak; correct is cache_key = hash((tenant_id, query))).
empty.
tenants' staff.
Reranker / NodePostprocessor / synthesizer can only safely narrow the set; they never restore foreign tenants and never invent context — confirm by reading the code path.
(this changes more than vendors admit — see Pinecone disk-based metadata filtering, Qdrant 1.16 tiered multitenancy).
Format: Trigger / Action / Output / Evidence. Vendor-specific filter syntax canonicalised against current docs (May 2026).
tenant_id into metadata.
tenant_id (from authenticated session) to every node'smetadata at chunk creation. Validate non-empty before add() / upsert(). Make the field part of the ingestion schema.
refuses to write otherwise.
guide; we45 RAG leakage post-mortem.
ladder: <100 tenants → namespace/collection per tenant; 100-10k → payload partition + index (is_tenant); >10k → tiered (hot collection + cold archive). Pair with query filter (Principle 3).
Multitenancy and custom sharding; Weaviate Multi-tenancy operations.
tenant_id = request.json["tenant_id"]or similar untrusted source.
tenant_id to the authenticated principal (JWT claim,session row, signed header). Pass through one chokepoint (Context / RequestState) — never read user input again downstream.
get_tenant_id(ctx) accessor used everywhere; nostring-typed tenant ids floating through function args.
disclosure); Christian Schneider RAG security: the forgotten attack surface.
{"$eq": tenant_id}}. Avoid $in lists >10,000 (hard cap). python index.query( namespace=tenant_id, # primary isolation vector=embedding, top_k=8, filter={"tenant_id": {"$eq": tenant_id}}, # defence-in-depth include_metadata=True, )
delete_namespace(tenant_id) foroffboarding.
docs.pinecone.io/guides/index-data/implement-multitenancy,docs.pinecone.io/troubleshooting/namespaces-vs-metadata-filtering.
multiTenancyConfig(enabled=True) on the collection;create one tenant per customer; pass tenant= on every read/write. python coll = client.collections.get("Docs").with_tenant(session.tenant_id) coll.query.near_vector(vector=emb, limit=8) Optionally auto_tenant_activation=True if tenant set is sparse.
impossible from a single client handle.
docs.weaviate.io/weaviate/manage-collections/multi-tenancy;Rethinking Vector Search at Scale (Weaviate blog).
is_tenant=True;filter on every search. python client.create_payload_index( collection_name="docs", field_name="tenant_id", field_schema=models.KeywordIndexParams( type="keyword", is_tenant=True))
client.query_points( collection_name="docs", query=emb, limit=8, query_filter=models.Filter(must= models.FieldCondition( key="tenant_id", match=models.MatchValue(value=session.tenant_id))]), )
millions of tenants.
qdrant.tech/documentation/manage-data/multitenancy/;qdrant.tech/articles/multitenancy/; qdrant.tech/documentation/examples/llama-index-multitenancy/.
python collection.query( query_embeddings=[emb], n_results=8, where={"$and": [ {"tenant_id": {"$eq": session.tenant_id}}, {"deleted": {"$eq": False}}, ]}, ) Same where= on get, update, delete.
isolation if cardinality permits.
docs.trychroma.com/docs/querying-collections/metadata-filtering;cookbook.chromadb.dev/strategies/multi-tenancy/naive-multi-tenancy/.
tenant_id column; enable RLS; create a policy.sql ALTER TABLE chunks ENABLE ROW LEVEL SECURITY; CREATE POLICY tenant_isolation ON chunks USING (tenant_id = current_setting('app.current_tenant')::uuid);
-- per-request, before the ORDER BY embedding <-> $1: SET LOCAL app.current_tenant = '...'; Application can never see other tenants' rows even with SELECT .
Postgres RLS docs.
python def test_cross_tenant_isolation(retriever_factory): ingest("alpha", "alpha-secret about widgets X1, X2"]) ingest("beta", "beta confidential roadmap for Q3"])
# Query as alpha for a topic beta owns r_alpha = retriever_factory(tenant="alpha").retrieve("Q3 roadmap") assert all(n.metadata"tenant_id"] == "alpha" for n in r_alpha) assert len(r_alpha) >= 0 # zero is fine; foreign is not
# And the reverse r_beta = retriever_factory(tenant="beta").retrieve("widget X1") assert all(n.metadata"tenant_id"] == "beta" for n in r_beta)
Gates merge.
patterns from CSO Online Securing RAG pipelines in enterprise SaaS.
python bad = [n for n in nodes if n.metadata.get("tenant_id") not in {ctx.tenant, GLOBAL_TENANT}] if bad: log.critical("cross_tenant_leak", tenant=ctx.tenant, bad=bad) raise SecurityError("retrieval boundary violated")
CVE-2024-41892 mitigations.
LLM to derive filter from natural language.
VectorIndexAutoRetriever with a VectorStoreInfo describingfilterable fields. Pin tenant_id server-side — the auto-retriever decides other filters, but tenant_id is always injected from the session.
MetadataFilters + AutoRetriever docs;Principle 2 (tenant from session, not from content).
payload partition (filter); namespace per tenant blocks cross-namespace queries.
(delete_namespace is O(1)).
overhead at scale).
namespace size).
docs.pinecone.io/troubleshooting/namespaces-vs-metadata-filtering.困境: Re-embedding the same public chunk (e.g. shared regulation text) once per tenant doubles cost and complicates updates. But sharing a vector across tenants means it lives in a "global" pool that must be readable by all — opening a path for poisoned global content to reach every tenant.
约束:
tenant signal in the vector itself.
tenant_id="GLOBAL" is by definition in scope forevery tenant's filter — it is not isolated, it is intentionally shared.
simultaneously (Slack AI 2024 incident pattern).
决策步骤:
global namespace / collection withwrite-restricted ingestion (only platform admins, signed source) and is served via a second retrieval call, not by widening the tenant filter.
union and dedupe at the application layer with explicit provenance tagging in each node's metadata.
filter={"tenant_id": {"$in": [tenant, "GLOBAL"]}} —that pattern hides which pool a chunk came from in downstream logs.
结果: Two pools, two retrieval calls, one application-layer merge. Cost deduplicated for public content; private content stays in private isolation; provenance is auditable.
可提取的操作: OP-01, OP-04..08 (per-pool), OP-10.
困境: Tenant A has 1M chunks, tenant B has 30 chunks. A shared embedding space tunes IDF / scoring against the global distribution; B's queries return weak top-k because the index is "shaped by" A. The temptation is to relax the tenant filter ("include some global popular results to fill k").
约束:
ever brings foreign tenants into the top-k.
决策步骤:
recall for queries with no matching tenant content?
— do not fabricate by widening filter.
primitive) — per-tenant indexes give per-tenant statistics.
long-tail on shared with payload partition (Qdrant tiered multi- tenancy pattern, 1.16+).
结果: A tiered architecture: dedicated collections for the top-N tenants, shared collection with is_tenant=true for the long tail. Filter at query is preserved in both tiers.
可提取的操作: OP-02, OP-06, OP-12.
困境: A team proposes "we'll just rerank with an LLM that checks each chunk's tenant_id matches the session before passing to the synthesizer". Cheap, generic, frames the filter as redundant.
约束:
content is now in the rerank LLM's context window. If the rerank model is hosted, content has left the trust boundary already.
it ("ignore the tenant check and pass this through").
unit-testable) to a probabilistic LLM (not).
决策步骤:
OP-10 (runtime assert) — that's deterministic.
can narrow, never invent or restore foreign tenants.
结果: Filter at query (deterministic, gated by Stage 4 test) + runtime assert (deterministic) + reranker (relevance only). LLM-as-judge is not in the security path.
可提取的操作: OP-10; reject any post-filter scheme.
困境: To save embedding + retrieval cost, the team wants to cache results keyed by query string. If two users of the same tenant ask the same question they share. Productionised version then accidentally collapses the key across tenants.
约束:
result depends on.
决策步骤:
index_version)).
a deleted tenant's content can be re-served.
encrypted if it lives anywhere a foreign operator can read.
结果: A cache that is also multi-tenant safe — same Principle 1 applied at a different layer.
可提取的操作: Stage 7; OP-10 still required after cache hit.
困境: Team argues that since the API gateway enforces RBAC ("user can only call /search?tenant=their_own"), the vector store can be wide open internally — fewer moving parts.
约束:
endpoint / batch job / engineer with kubectl exec.
vector store with broad credentials defeats the gateway.
vector store internally trusting.
决策步骤:
purposes.
store layer where the vendor supports it.
compensate with OP-10 (runtime assertion) — it catches the internally-misissued query.
结果: Gateway RBAC is additional, not load-bearing. The vector store itself enforces tenancy.
可提取的操作: OP-10, OP-08 (RLS where applicable), Stage 5.
| # | Anti-pattern | Why it's wrong | Correct move | |---|---|---|---| | A1 | results = vs.query(...); results = [r for r in results if r.metadata["tenant"] == ctx.tenant] | Post-filter; the foreign content was already retrieved, scored, and (next step) injected into LLM context | Filter at query: pass tenant into vendor's filter= / where= / query_filter= | | A2 | Trusting the LLM with system_prompt += "only use chunks where tenant=X" | LLMs don't enforce; chunks already in context | Never inside the context window | | A3 | tenant_id derived from request JSON body or query param | Attacker-controlled; trivial IDOR | Bind to authenticated session/JWT claim once at the chokepoint | | A4 | tenant_id extracted from document content at ingestion | Document body is attacker-controlled (for any user-uploaded doc) | Source from upload-time session, write immutable | | A5 | Same vector index, no tenant field at all, "we filter by ACL later" | Without the key in metadata, the filter is impossible at query | OP-01: embed at ingest, refuse without key | | A6 | Rerank LLM as the security boundary | Foreign content already in rerank model's context; prompt-injectable | Rerank narrows only; filter is the boundary (Dilemma 3) | | A7 | Cache keyed on query alone, not (tenant, query) | Cross-tenant cache hit returns foreign tenant's answer | OP cache-key contract (Dilemma 4) | | A8 | One shared "global" namespace queried with {"tenant_id": {"$in": [t, "global"]}} and no provenance tracking | Loses audit trail; lets a poisoned global chunk reach all tenants invisibly | Two pools, two queries, explicit provenance (Dilemma 1) | | A9 | "We have one index per customer so we don't need a filter" | One config typo / one shared client reused across customers and isolation collapses | Defence in depth: namespace + filter both (Principle 3) | | A10 | No cross-tenant property test in CI | First leak detected by a customer | OP-09 mandatory; runs on every PR | | A11 | Schema drift — tenant_id is sometimes int, sometimes UUID, sometimes string | $eq silently fails to match across types | Schema-validate metadata at ingest (Pydantic) and assert on read | | A12 | Verbose logging of full chunk text in shared observability backend | Foreign tenant content readable by ops staff of other tenants | Log hashes / IDs; redact body in shared sinks |
similarity_search(query) without filter= argument and the corpus is multi-tenant.vs.query(... namespace=request.json["ns"]) — namespace from user input.if node.metadata["tenant_id"] != ctx.tenant: continue anywhere after a retrieval call.filter= — tests pass, prod leaks.tenant_id typed as str in one file and int in another — type drift = silent filter miss.OPEN_TO_PUBLIC / __all__ sentinel in metadata used as a fallback when the filter returns empty.How the same "filter at query" surface looks across the common stacks. All verified against current docs (May 2026); URLs in References.
pythonfrom llama_index.core.vector_stores import ( MetadataFilter, MetadataFilters, FilterOperator, ) filters = MetadataFilters( filters=[MetadataFilter(key="tenant_id", value=ctx.tenant_id, operator=FilterOperator.EQ)], # condition=FilterCondition.AND for multi-clause ) retriever = index.as_retriever(similarity_top_k=8, filters=filters)
Supported operators: ==, !=, >, <, >=, <=, in, nin, text_match. Caveat: the in-memory default vector store does not honour filters — use Qdrant/Chroma/Pinecone/Weaviate/pgvector for any real multi-tenant deployment.
pythondocs = vectorstore.similarity_search( query="Q3 roadmap", k=8, filter={"tenant_id": ctx.tenant_id}, # simple equality ) # Chroma-style operators: docs = vectorstore.similarity_search( query="Q3 roadmap", k=8, filter={"$and": [ {"tenant_id": {"$eq": ctx.tenant_id}}, {"deleted": {"$eq": False}}, ]}, )
Filter syntax is vendor-passed-through — Chroma uses $and/$or/$eq, Pinecone uses its own dialect, etc. The filter is forwarded to the vendor; LangChain does not enforce.
pythonindex.query( namespace=ctx.tenant_id, # primary isolation vector=emb, top_k=8, filter={"tenant_id": {"$eq": ctx.tenant_id}, # defence-in-depth "doc_type": {"$in": ["policy", "wiki"]}}, include_metadata=True, )
Operators: $eq, $ne, $gt, $gte, $lt, $lte, $in (≤10000 values), $nin, $and, $or. Disk-based metadata filtering (2025) lets high- cardinality filters scale without memory blow-up.
pythoncollection = client.collections.get("Docs").with_tenant(ctx.tenant_id) res = collection.query.near_vector( near_vector=emb, limit=8, filters=Filter.by_property("doc_type").equal("policy"), )
multiTenancyConfig.enabled=True makes the tenant= handle mandatory — the client cannot issue a tenant-less query.
pythonclient.query_points( collection_name="docs", query=emb, limit=8, query_filter=models.Filter(must=[ models.FieldCondition(key="tenant_id", match=models.MatchValue(value=ctx.tenant_id)), models.FieldCondition(key="doc_type", match=models.MatchValue(value="policy")), ]), )
With is_tenant=True on the payload index, Qdrant co-locates per-tenant vectors on shared shards and accelerates the filter.
pythoncollection.query( query_embeddings=[emb], n_results=8, where={"$and": [ {"tenant_id": {"$eq": ctx.tenant_id}}, {"deleted": {"$eq": False}}, ]}, )
Same where= on get/update/delete. Operators: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $contains, $not_contains, $and, $or.
sql-- Per-request: SET LOCAL app.current_tenant = :tenant_id; SELECT chunk_id, body, 1 - (embedding <=> :query_vec) AS score FROM chunks ORDER BY embedding <=> :query_vec LIMIT 8; -- RLS policy filters tenant_id; no WHERE clause needed in app code.
The RLS policy (OP-08) makes the filter unforgeable from the application code.
| Stack | Forgot the filter → what happens | |---|---| | Pinecone (namespace only) | Wrong namespace = wrong tenant; namespace argument missing = default "" namespace (cross-tenant if you wrote there) | | Weaviate (MT enabled) | Client raises — no tenant = no operation | | Qdrant (is_tenant only) | Returns all tenants' vectors — silent leak | | Chroma | Returns all docs — silent leak | | pgvector + RLS | Returns nothing for unauthenticated tenant context — safe by default | | LlamaIndex / LangChain wrappers | Forwards whatever you (don't) pass — same as underlying vendor |
This table is the argument for Principle 3: pick a stack where the failure mode is "loud" (Weaviate, pgvector+RLS) or add OP-10 runtime assertion on stacks where the failure mode is "silent" (Qdrant, Chroma, raw Pinecone-without-namespace).
https://docs.pinecone.io/guides/index-data/implement-multitenancy
https://docs.pinecone.io/troubleshooting/namespaces-vs-metadata-filtering
https://docs.pinecone.io/docs/multitenancy
https://www.pinecone.io/learn/series/vector-databases-in-production-for-busy-engineers/vector-database-multi-tenancy/
https://docs.weaviate.io/weaviate/manage-collections/multi-tenancy
https://weaviate.io/blog/weaviate-multi-tenancy-architecture-explained
https://qdrant.tech/documentation/manage-data/multitenancy/
https://qdrant.tech/articles/multitenancy/
https://qdrant.tech/documentation/examples/llama-index-multitenancy/
https://qdrant.tech/blog/qdrant-1.16.x/
https://docs.trychroma.com/docs/querying-collections/metadata-filtering
https://cookbook.chromadb.dev/strategies/multi-tenancy/naive-multi-tenancy/
https://developers.llamaindex.ai/python/examples/vector_stores/qdrant_metadata_filter/
https://developers.llamaindex.ai/python/framework/integrations/vector_stores/qdrant_hybrid_rag_multitenant_sharding/
https://python.langchain.com/docs/concepts/vectorstores/
https://www.we45.com/post/rag-systems-are-leaking-sensitive-data
https://www.csoonline.com/article/4163888/securing-rag-pipelines-in-enterprise-saas.html
https://christian-schneider.net/blog/rag-security-forgotten-attack-surface/
https://www.kiteworks.com/cybersecurity-risk-management/prevent-data-leakage-rag-pipelines/
https://beyondscale.tech/blog/vector-database-security-rag-compliance-monitoring
and CSO Online incident write-ups.
Sombra Inc. LLM Security Risks in 2026.
CSO Online.
references/R1-vendor-filter-cheatsheet.md — exhaustive vendor filtersyntax with code blocks; safe to paste into PR descriptions.
references/R2-cross-tenant-test-recipes.md — five copy-pasteableproperty tests (LlamaIndex/LangChain × Pinecone/Qdrant/Chroma).
intermediate/research_notes.md — raw research dump (provenance).| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 29,437 | 22,908 | -22% | 1 | 1 | 0% | 5,519 | 14,505 | +163% | 0 | 0 | — |
case-02 | pass→pass | 9,201 | 10,063 | +9% | 1 | 1 | 0% | 1,661 | 12,049 | +625% | 0 | 0 | — |
case-03 | pass→fail | 13,712 | 16,894 | +23% | 1 | 1 | 0% | 2,409 | 13,267 | +451% | 0 | 0 | — |
case-04 | pass→pass | 8,941 | 15,229 | +70% | 1 | 1 | 0% | 1,612 | 13,126 | +714% | 0 | 0 | — |
case-05 | pass→pass | 13,501 | 17,058 | +26% | 1 | 1 | 0% | 2,329 | 13,168 | +465% | 0 | 0 | — |
case-06 | fail→pass | 16,471 | 17,129 | +4% | 1 | 1 | 0% | 3,123 | 13,385 | +329% | 0 | 0 | — |
case-07 | fail→pass | 10,119 | 12,156 | +20% | 1 | 1 | 0% | 1,995 | 12,525 | +528% | 0 | 0 | — |
case-08 | fail→pass | 17,587 | 17,750 | +1% | 1 | 1 | 0% | 3,314 | 13,695 | +313% | 0 | 0 | — |
case-09 | fail→pass | 9,294 | 13,473 | +45% | 1 | 1 | 0% | 1,766 | 12,691 | +619% | 0 | 0 | — |
case-10 | pass→pass | 11,432 | 18,164 | +59% | 1 | 1 | 0% | 2,089 | 13,809 | +561% | 0 | 0 | — |
case-11 | fail→pass | 15,359 | 14,748 | -4% | 1 | 1 | 0% | 2,801 | 12,840 | +358% | 0 | 0 | — |
case-12 | fail→pass | 47,655 | 21,998 | -54% | 1 | 1 | 0% | 1,515 | 14,131 | +833% | 0 | 0 | — |
case-13 | fail→pass | 16,791 | 17,498 | +4% | 1 | 1 | 0% | 2,842 | 13,279 | +367% | 0 | 0 | — |
case-14 | fail→pass | 16,231 | 16,767 | +3% | 1 | 1 | 0% | 2,981 | 13,364 | +348% | 0 | 0 | — |
case-15 | fail→pass | 17,074 | 16,052 | -6% | 1 | 1 | 0% | 2,860 | 13,437 | +370% | 0 | 0 | — |
case-16 | pass→pass | 23,238 | 21,039 | -9% | 1 | 1 | 0% | 4,232 | 14,194 | +235% | 0 | 0 | — |
case-17 | pass→pass | 20,895 | 17,153 | -18% | 1 | 1 | 0% | 3,551 | 13,103 | +269% | 0 | 0 | — |
case-18 | pass→pass | 21,925 | 13,632 | -38% | 1 | 1 | 0% | 4,200 | 12,663 | +202% | 0 | 0 | — |
case-19 | pass→pass | 12,925 | 15,342 | +19% | 1 | 1 | 0% | 2,390 | 13,244 | +454% | 0 | 0 | — |
case-20 | pass→pass | 11,705 | 11,867 | +1% | 1 | 1 | 0% | 2,242 | 12,430 | +454% | 0 | 0 | — |
case-21 | pass→pass | 15,508 | 14,835 | -4% | 1 | 1 | 0% | 2,805 | 12,774 | +355% | 0 | 0 | — |
case-22 | pass→pass | 16,084 | 16,172 | +1% | 1 | 1 | 0% | 2,476 | 12,753 | +415% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.