Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Change the ContextOS canonical specification and typed reference implementation without semantic drift, conflicting conventions, or unnecessary scope. Use for runtime-model, terminology, compiler, or cross-plane spec changes in the ContextOS repository; not for ordinary site styling.
.claude/skills/contextosai-contextos-spec-steward/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -47% | 0% |
Make the smallest coherent change to the public spec and its typed reference. New patterns belong in the spec only after validation in a working system; label proposals instead of presenting unproven concepts as canonical.
AGENTS.md plus repository architecture guidance.Read references/change-map.md for semantic touch sets and validation routing.
Determine whether the request is:
Write down the invariant being changed and the compatibility expectation before editing. For schema-level work, apply the same compatibility-first sequence across types, producers, schemas, examples, docs, and tests.
src/lib/contextos/ is the small deterministic typed reference, not a production runtime.ActionRisk dimensions remain independent. ApprovalMode is only a compatibility projection.Use the coupling map as a maximum coherent set, not a checklist of files to touch. Update only surfaces whose observable contract changes. Keep canonical IDs and examples sourced from the scenario fixture rather than copying stale variants.
When adding a concept, define:
Call out non-canonical extensions explicitly. Do not add a type solely because prose benefits from a convenient noun.
Run the narrowest relevant tests first, then typecheck/lint/full tests/build in proportion to the surface. Inspect the diff for accidental terminology changes and verify published routes/assets when touched.
Report changed invariants, coupled surfaces actually updated, compatibility decision, checks run, and anything unverified.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,051 | 6,414 | +27% | 1 | 1 | 0% | 256 | 950 | +271% | 0 | 0 | — |
case-02 | fail→fail | 5,712 | 5,588 | -2% | 1 | 1 | 0% | 271 | 903 | +233% | 0 | 0 | — |
case-03 | fail→fail | 5,878 | 6,460 | +10% | 1 | 1 | 0% | 230 | 947 | +312% | 0 | 0 | — |
case-04 | fail→pass | 13,449 | 10,370 | -23% | 1 | 1 | 0% | 2,052 | 2,128 | +4% | 0 | 0 | — |
case-05 | fail→fail | 20,063 | 6,471 | -68% | 1 | 1 | 0% | 2,505 | 1,061 | -58% | 0 | 0 | — |
case-06 | fail→fail | 13,958 | 5,971 | -57% | 1 | 1 | 0% | 2,028 | 1,029 | -49% | 0 | 0 | — |
case-07 | fail→fail | 37,099 | 16,404 | -56% | 1 | 1 | 0% | 692 | 869 | +26% | 0 | 0 | — |
case-08 | fail→pass | 13,546 | 7,024 | -48% | 1 | 1 | 0% | 1,787 | 1,636 | -8% | 0 | 0 | — |
case-09 | pass→fail | 15,860 | 6,443 | -59% | 1 | 1 | 0% | 2,339 | 927 | -60% | 0 | 0 | — |
case-10 | pass→pass | 11,885 | 8,276 | -30% | 1 | 1 | 0% | 1,768 | 1,939 | +10% | 0 | 0 | — |
case-11 | fail→pass | 13,185 | 14,724 | +12% | 1 | 1 | 0% | 2,189 | 1,665 | -24% | 0 | 0 | — |
case-12 | fail→fail | 9,602 | 5,234 | -45% | 1 | 1 | 0% | 1,487 | 841 | -43% | 0 | 0 | — |
case-13 | fail→pass | 12,292 | 2,353 | -81% | 1 | 1 | 0% | 1,970 | 950 | -52% | 0 | 0 | — |
case-14 | pass→pass | 11,314 | 7,491 | -34% | 1 | 1 | 0% | 1,838 | 1,330 | -28% | 0 | 0 | — |
case-15 | fail→pass | 11,555 | 3,372 | -71% | 1 | 1 | 0% | 1,994 | 1,051 | -47% | 0 | 0 | — |
case-16 | pass→fail | 16,602 | 15,347 | -8% | 1 | 1 | 0% | 2,722 | 2,774 | +2% | 0 | 0 | — |
case-17 | fail→pass | 8,516 | 3,471 | -59% | 1 | 1 | 0% | 1,153 | 1,254 | +9% | 0 | 0 | — |
case-18 | fail→pass | 13,269 | 9,805 | -26% | 1 | 1 | 0% | 2,043 | 1,528 | -25% | 0 | 0 | — |
case-19 | pass→pass | 13,220 | 3,423 | -74% | 1 | 1 | 0% | 1,702 | 1,132 | -33% | 0 | 0 | — |
case-20 | pass→fail | 8,939 | 4,879 | -45% | 1 | 1 | 0% | 1,445 | 861 | -40% | 0 | 0 | — |
case-21 | pass→pass | 9,221 | 9,909 | +7% | 1 | 1 | 0% | 1,521 | 2,475 | +63% | 0 | 0 | — |
case-22 | pass→pass | 10,128 | 5,895 | -42% | 1 | 1 | 0% | 1,872 | 1,606 | -14% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 13 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 13 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.