Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Emit privacy-safe receipts for context selection, deferral, hydration, compaction, pruning, delegation, usage attribution, and boundary handoffs.
.claude/skills/aiskillstore-context-receipts/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 154% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 87% | 0% |
Use this skill when an agent workflow claims to save context by selecting, deferring, hydrating, summarizing, compacting, pruning, delegating, attributing usage, or isolating context.
The job is not to log the private content. The job is to emit a small receipt that lets a reviewer answer:
> what crossed the context boundary, what stayed out, and what audit gap remains?
Never include raw prompts, raw tool schemas, raw tool arguments, raw tool results, raw skill bodies, memory bodies, secrets, customer names, or full transcripts in the receipt.
Prefer:
audit_gap field when the receipt proves routing but not semantic correctness.For MCP Tool Search, lazy tool loading, or progressive disclosure, emit enough evidence to answer these seven checks:
Minimal JSONL event names:
jsonl{"event":"mcp.tool_index.loaded","loaded_server_count":12,"loaded_tool_index_count":84,"full_schema_count":0,"suppressed_tool_count":84,"raw_schema_copied":false,"startup_token_bucket":"lt_1k"} {"event":"mcp.tool_search.performed","query_hash":"sha256:...","query_category":"repo_search","candidate_tool_count":5,"selected_tool_id":"github.search_code","raw_query_copied":false} {"event":"mcp.tool_definition.loaded","tool_id":"github.search_code","hydrate_reason":"selected_after_tool_search","suppressed_tool_count":83,"definition_token_bucket":"1k_2k","raw_schema_copied":false} {"event":"mcp.tool_call.completed","tool_id":"github.search_code","args_hash":"sha256:...","result_token_bucket":"2k_4k","raw_args_copied":false,"raw_result_copied":false,"status":"ok"}
For MCP dynamic discovery, gateways, admin/Purview-style audit trails, or runtime tool catalogs, separate discovery from activation:
accepted, blocked_by_rai, blocked_by_xpia, schema_invalid, or entitlement_filtered;Minimal JSON shape:
json{ "receipt_type": "pluribus.mcp_tool_surface_diff_receipt.v1", "runtime_discovery": { "trigger": "turn_start|admin_refresh|tool_search|manual_refresh", "before_catalog_hash": "sha256:...", "after_catalog_hash": "sha256:..." }, "summary": { "discovered_count": 3, "activated_count": 1, "withheld_count": 1, "blocked_count": 1 }, "privacy": { "raw_schemas_copied": false, "raw_prompts_copied": false, "raw_results_copied": false }, "audit_gap": "proves tool-surface boundary, not semantic usefulness" }
For GraphRAG, memory, code search, transcript review, or baseline-first workflows, separate retrieval from attention:
Minimal JSON shape:
json{ "receipt_type": "pluribus.context_attention_receipt.v1", "required_context_ids": ["ctx:auth-boundary", "ctx:migration-plan"], "delivered_context_ids": ["ctx:auth-boundary", "ctx:migration-plan"], "acknowledged_before_plan_ids": ["ctx:auth-boundary", "ctx:migration-plan"], "cited_before_edit_ids": ["ctx:auth-boundary"], "missing_context_stop": "stop_before_edit", "privacy": { "raw_context_copied": false, "raw_transcript_copied": false }, "audit_gap": "proves required context was acknowledged/cited, not that the edit is correct" }
For skills, rules, AGENTS.md overlays, or instruction files, answer:
Minimal event names:
context.skill.registry.index.loadedcontext.skill.registry.skill.readcontext.skill.registry.skill.injectedcontext.input.loadedcontext.input.candidate_suppressedFor role-specific subagents or per-agent MCP configs, prove the policy boundary before debugging model quality:
Minimal JSONL event names:
jsonl{"event":"subagent.mcp_policy.applied","subagent_role":"testing","available_server_count":2,"available_servers_hash":"sha256:...","excluded_server_count":5,"excluded_servers_hash":"sha256:...","policy_source":"role_config","raw_server_names_copied":false} {"event":"subagent.context_boot.evaluated","subagent_role":"testing","loaded_tool_definition_count":0,"deferred_tool_definition_count":48,"startup_token_bucket":"50k_75k","raw_schema_copied":false,"audit_gap":"proves injection boundary, not tool relevance"}
For subagents that should inherit MCP through ToolSearch, distinguish policy, declaration, and runtime filtering:
tools: declaration wildcard, explicit include, or exclusion style?ToolSearch declared and was it actually exposed in the subagent tool surface?Minimal JSONL event names:
jsonl{"event":"subagent.toolsearch.propagation.evaluated","spawn_path":"Task","tools_declaration_shape":"enumerated_include","toolsearch_declared":false,"toolsearch_exposed":false,"mcp_servers_available_bucket":"0","deferred_tool_definitions_bucket":"0","filtered_by":"frontmatter_tools_policy_or_runtime_filter","raw_tool_schemas_copied":false} {"event":"subagent.toolsearch.matrix.completed","tested_axis":"tools_frontmatter_shape","audit_gap":"proves ToolSearch exposure, not semantic tool relevance or runtime call success"}
For semantic code search, repo RAG, or MCP tools such as Claude Context, separate "search returned" from "agent context loaded":
Minimal JSONL event names:
jsonl{"event":"code.index.snapshot.used","snapshot_id_hash":"sha256:...","codebase_path_hash":"sha256:...","indexed_chunk_count_bucket":"over_1k","raw_codebase_path_copied":false} {"event":"code.search.performed","query_hash":"sha256:...","query_category":"auth_debug","candidate_count_bucket":"over_1k","raw_query_copied":false} {"event":"code.search.result.returned","rank":1,"chunk_id_hash":"sha256:...","chunk_text_hash":"sha256:...","path_hash":"sha256:...","score_bucket":"high","stale":false,"raw_code_copied":false} {"event":"context.input.loaded","kind":"retrieved_code_chunks","loaded_chunk_count":3,"suppressed_chunk_count":2,"suppression_reasons":["duplicate","stale_snapshot_chunk"],"raw_code_copied":false}
For /usage, /context, /doctor, or other context-budget breakdowns, map each displayed category to evidence that can be reviewed without exposing private content:
Minimal JSONL event names:
jsonl{"event":"context.usage.window.measured","window":"current_session","total_token_bucket":"100k_150k","raw_prompts_copied":false} {"event":"context.usage.category.attributed","category":"mcp_server","component_hash":"sha256:...","loaded_token_bucket":"10k_25k","deferred_definition_count":42,"hydrated_definition_count":3,"raw_schema_copied":false} {"event":"context.usage.breakdown.completed","categories":["skills","subagents","plugins","mcp_server"],"audit_gap":"proves attribution buckets, not whether each component was necessary"}
For context-cleaning, pruning, compaction, or doctor/guard tools, answer:
Minimal JSONL event names:
jsonl{"event":"context.prune.started","prescription":"balanced","trigger":"manual_dry_run","before_token_bucket":"150k_200k","raw_transcript_copied":false} {"event":"context.prune.strategy.evaluated","strategy":"tool-output-trim","candidate_bucket":"10_25","changed_bucket":"5_10","protected_bucket":"1_5","raw_tool_output_copied":false} {"event":"context.prune.completed","after_token_bucket":"75k_100k","backup_verified":true,"protected_summary_count":2,"raw_text_copied":false,"audit_gap":"proves pruning/protection counts, not semantic disposability"}
For failed compaction, also prove transaction safety:
Minimal JSONL event names:
jsonl{"event":"context.compaction.summary.attempted","summary_call_status":"failed_rate_limited","candidate_summary_available":false,"raw_error_copied":false} {"event":"context.compaction.rollback.completed","swap_committed":false,"original_context_preserved":true,"deferred_tool_registry_restored":true,"system_reminder_queue_restored":true,"replayed_system_reminder_count":0} {"event":"context.compaction.transaction.completed","status":"rolled_back","authoritative_state":"pre_compaction_context","post_tokens_recorded_as_success":false,"raw_context_copied":false}
For subagents, manager agents, or child workers, answer:
Minimal event names:
subagent.delegation.requestedsubagent.tool_output.capturedsubagent.summary.returnedparent.context_budget.evaluatedA receipt is useful if a maintainer can debug one of these failures without seeing private content:
A receipt is not enough if it only says “Tool Search enabled” or “used subagent”. It must prove the boundary behavior.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | pass→pass | 17,958 | 14,824 | -17% | 1 | 1 | 0% | 2,323 | 6,057 | +161% | 0 | 0 | — |
case-21 | pass→pass | 17,273 | 15,345 | -11% | 1 | 1 | 0% | 1,859 | 5,334 | +187% | 0 | 0 | — |
case-22 | pass→pass | 13,443 | 16,628 | +24% | 1 | 1 | 0% | 1,410 | 5,561 | +294% | 0 | 0 | — |
case-14 | fail→pass | 11,684 | 10,712 | -8% | 1 | 1 | 0% | 2,316 | 4,556 | +97% | 0 | 0 | — |
case-01 | fail→pass | 14,390 | 14,453 | +0% | 1 | 1 | 0% | 3,002 | 5,531 | +84% | 0 | 0 | — |
case-02 | fail→pass | 14,670 | 9,507 | -35% | 1 | 1 | 0% | 1,777 | 4,515 | +154% | 0 | 0 | — |
case-03 | fail→pass | 11,454 | 9,686 | -15% | 1 | 1 | 0% | 2,144 | 4,376 | +104% | 0 | 0 | — |
case-19 | fail→pass | 12,454 | 10,465 | -16% | 1 | 1 | 0% | 2,395 | 4,474 | +87% | 0 | 0 | — |
case-04 | fail→pass | 13,727 | 12,083 | -12% | 1 | 1 | 0% | 1,636 | 5,097 | +212% | 0 | 0 | — |
case-05 | fail→pass | 11,917 | 15,116 | +27% | 1 | 1 | 0% | 2,470 | 5,544 | +124% | 0 | 0 | — |
case-06 | fail→pass | 15,976 | 13,336 | -17% | 1 | 1 | 0% | 3,377 | 5,846 | +73% | 0 | 0 | — |
case-07 | fail→pass | 18,813 | 12,173 | -35% | 1 | 1 | 0% | 2,155 | 4,976 | +131% | 0 | 0 | — |
case-08 | pass→pass | 10,852 | 7,038 | -35% | 1 | 1 | 0% | 1,079 | 5,020 | +365% | 0 | 0 | — |
case-09 | fail→pass | 16,665 | 5,467 | -67% | 1 | 1 | 0% | 2,297 | 4,680 | +104% | 0 | 0 | — |
case-10 | fail→pass | 22,742 | 17,346 | -24% | 1 | 1 | 0% | 2,723 | 5,827 | +114% | 0 | 0 | — |
case-11 | fail→pass | 27,304 | 8,938 | -67% | 1 | 1 | 0% | 4,186 | 5,304 | +27% | 0 | 0 | — |
case-12 | fail→pass | 11,083 | 3,256 | -71% | 1 | 1 | 0% | 908 | 4,252 | +368% | 0 | 0 | — |
case-13 | fail→pass | 45,711 | 12,784 | -72% | 1 | 1 | 0% | 4,443 | 5,131 | +15% | 0 | 0 | — |
case-15 | pass→pass | 22,795 | 15,755 | -31% | 1 | 1 | 0% | 2,968 | 5,468 | +84% | 0 | 0 | — |
case-16 | fail→pass | 11,156 | 13,038 | +17% | 1 | 1 | 0% | 1,156 | 5,111 | +342% | 0 | 0 | — |
case-17 | fail→pass | 16,246 | 12,583 | -23% | 1 | 1 | 0% | 1,866 | 5,133 | +175% | 0 | 0 | — |
case-18 | pass→pass | 19,525 | 7,477 | -62% | 1 | 1 | 0% | 2,391 | 4,920 | +106% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +73 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.