Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Author, review, or evolve a ContextOS Context Pack and its compiler fixtures, policies, tools, decisions, evidence gates, memory rules, and evaluator targets. Use when the requested artifact is a Context Pack or compiled-context scenario, not for generic prompt writing.
.claude/skills/contextosai-contextos-context-pack/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 38% | 0% |
Produce a versioned workflow contract that compiles into bounded, attributable context. A Context Pack is not a prompt template and must not become a data dump.
When a ContextOS repository is available:
AGENTS.md and repository guidance;ContextPack and related types, published JSON Schema, canonical scenario fixture, compiler, and Context Pack documentation;For the layer checklist, binding graph, and compiler invariants, read references/pack-checklist.md.
Start with the governed decision, not prose:
decision_key;Then bind policy rules, approval gates, tool permissions, evidence requirements, identity namespaces, memory rules, and budgets to that decision. Every reference must resolve in the same pack or a named registry.
Create all ten required layers using the current schema. Use stable, neutral identifiers. Keep raw customer data in evidence stores, secrets in a vault, mutable state in the run/session, and adapter code outside the pack.
Exercise at least these cases when the implementation supports them:
Inspect the complete CompiledContext, not only the prompt. Confirm manifests, runtime controls, context provenance, evidence gates, omissions, and the context-ledger hash all agree.
ActionRisk dimensions are checked conjunctively. The legacy approval mode is a lossy compatibility projection, not the risk model.decision_binding, approval gate, adapter/capability, evidence name, and evaluator intent resolves exactly.pack_id@pack_version is immutable. Overlays are ordered, bounded, validated, and part of replay lineage.Return the pack or patch, a binding summary, scenario results, compatibility/versioning decision, and any unresolved gaps. Run the repository's targeted compiler, context-admission, schema, and spec-drift tests when applicable; state exactly which checks were skipped.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 39,492 | 63,715 | +61% | 1 | 1 | 0% | 8,272 | 9,068 | +10% | 0 | 0 | — |
case-02 | fail→pass | 40,417 | 33,453 | -17% | 1 | 1 | 0% | 8,254 | 7,849 | -5% | 0 | 0 | — |
case-03 | fail→pass | 32,327 | 37,777 | +17% | 1 | 1 | 0% | 6,219 | 8,161 | +31% | 0 | 0 | — |
case-04 | pass→fail | 16,867 | 39,778 | +136% | 1 | 1 | 0% | 3,126 | 8,936 | +186% | 0 | 0 | — |
case-05 | pass→pass | 16,016 | 20,784 | +30% | 1 | 1 | 0% | 2,961 | 4,770 | +61% | 0 | 0 | — |
case-06 | fail→fail | 11,270 | 32,858 | +192% | 1 | 1 | 0% | 1,825 | 6,595 | +261% | 0 | 0 | — |
case-07 | pass→pass | 15,572 | 17,353 | +11% | 1 | 1 | 0% | 2,998 | 4,258 | +42% | 0 | 0 | — |
case-08 | pass→pass | 21,247 | 16,960 | -20% | 1 | 1 | 0% | 3,072 | 3,244 | +6% | 0 | 0 | — |
case-09 | pass→pass | 12,587 | 12,574 | -0% | 1 | 1 | 0% | 2,015 | 2,445 | +21% | 0 | 0 | — |
case-10 | pass→pass | 15,554 | 15,698 | +1% | 1 | 1 | 0% | 2,410 | 3,440 | +43% | 0 | 0 | — |
case-11 | pass→pass | 16,536 | 17,838 | +8% | 1 | 1 | 0% | 2,889 | 3,511 | +22% | 0 | 0 | — |
case-12 | pass→pass | 21,836 | 17,699 | -19% | 1 | 1 | 0% | 3,257 | 3,667 | +13% | 0 | 0 | — |
case-13 | fail→pass | 16,698 | 24,043 | +44% | 1 | 1 | 0% | 2,617 | 4,662 | +78% | 0 | 0 | — |
case-14 | pass→pass | 19,065 | 14,131 | -26% | 1 | 1 | 0% | 3,009 | 3,311 | +10% | 0 | 0 | — |
case-15 | pass→pass | 12,134 | 11,703 | -4% | 1 | 1 | 0% | 2,151 | 2,571 | +20% | 0 | 0 | — |
case-16 | pass→pass | 15,203 | 19,900 | +31% | 1 | 1 | 0% | 2,504 | 3,836 | +53% | 0 | 0 | — |
case-17 | pass→pass | 22,205 | 19,811 | -11% | 1 | 1 | 0% | 3,426 | 3,901 | +14% | 0 | 0 | — |
case-18 | fail→pass | 15,138 | 13,706 | -9% | 1 | 1 | 0% | 2,096 | 2,981 | +42% | 0 | 0 | — |
case-19 | fail→pass | 14,333 | 14,208 | -1% | 1 | 1 | 0% | 2,081 | 2,865 | +38% | 0 | 0 | — |
case-20 | fail→pass | 17,969 | 11,744 | -35% | 1 | 1 | 0% | 2,921 | 2,364 | -19% | 0 | 0 | — |
case-21 | pass→pass | 13,701 | 11,476 | -16% | 1 | 1 | 0% | 2,181 | 2,226 | +2% | 0 | 0 | — |
case-22 | fail→pass | 15,352 | 19,888 | +30% | 1 | 1 | 0% | 2,533 | 3,878 | +53% | 0 | 0 | — |
case-23 | fail→pass | 12,984 | 9,278 | -29% | 1 | 1 | 0% | 1,852 | 2,196 | +19% | 0 | 0 | — |
case-24 | pass→pass | 16,611 | 13,818 | -17% | 1 | 1 | 0% | 2,411 | 2,786 | +16% | 0 | 0 | — |
case-25 | fail→pass | 16,114 | 13,669 | -15% | 1 | 1 | 0% | 2,605 | 2,781 | +7% | 0 | 0 | — |
case-26 | fail→pass | 18,990 | 10,795 | -43% | 1 | 1 | 0% | 2,390 | 2,274 | -5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 26 cases were attempted. The headline lift of +35 percentage points is the difference between those two pass rates over the 26 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.