Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Every TicketData price signal for any event, performer, or venue, plus a local price-history store for trend math, drift alerts, and cross-event comparisons no ticket tool ships. Trigger phrases: `get-in price for`, `ticketdata price for`, `when should I buy tickets to`, `price history for`, `is this ticket cheap right now`, `which date is cheapest for`, `use ticketdata`, `run ticketdata`.
.claude/skills/mvanhorn-pp-ticketdata/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 383% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 302% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 3137% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 346% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 340% | 0% |
This skill drives the ticketdata-pp-cli binary. You must verify the CLI is installed before invoking any command from this skill. If it is missing, install it first:
$HOME/.local/bin on macOS/Linux and %LOCALAPPDATA%\Programs\PrintingPress\bin on Windows:bash npx -y @mvanhorn/printing-press-library install ticketdata --cli-only
ticketdata-pp-cli --version$PATH for the agent/runtime that will invoke this skill.If the npx install fails (no Node, offline, etc.), fall back to a direct Go install (requires Go 1.26.6 or newer). This installs into $GOPATH/bin (default $HOME/go/bin), so add that directory to $PATH instead:
bashgo install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/ticketdata/cmd/ticketdata-pp-cli@latest
If --version reports "command not found" after install, the runtime cannot see the binary directory on $PATH. Do not proceed with skill commands until verification succeeds.
TicketData tracks each event's get-in price (the lowest all-in resale price across marketplaces), its full price history, a forecast, and 3/7/14/30-day change. This CLI exposes all of it as scriptable, agent-native JSON and keeps the raw price series in local SQLite, so stats can compute historical lows and volatility, drift can alert when a floor drops to your target, board can rank your whole watchlist, and compare can pick the cheapest date, none of which the website surfaces.
Use this CLI when you need TicketData's get-in price, price history, forecast, or N-day change for specific events, performers, or venues, and especially when you want to track a set of events over time, compute historical price stats, alert on price drops, or compare events, all offline and scriptable. It is ideal for agents answering 'is this ticket cheap right now', 'when should I buy', or 'which of these dates is cheapest'.
Do not use this CLI for:
events get and complete the purchase there.These capabilities aren't available in any other tool for this API.
watch — Track a set of events locally so one sync re-fetches all their current prices._Reach for this first: everything downstream (board, drift, stats, compare) reads the watchlist the CLI syncs._
bash ticketdata-pp-cli watch add 22323960
board — One sortable table of every watched event: get-in price, N-day change, forecast direction, and where today sits in its own history._Use it for a whole-watchlist snapshot; for what changed since last sync use drift, for one event's distribution use stats._
bash ticketdata-pp-cli board --sort change --agent
drift — Diff the two most recent snapshots per watched event, flag floors that moved past a threshold, and fire price-target hits._Use it for what moved since the last sync and for buy-when-it-hits-my-price alerts._
bash ticketdata-pp-cli drift --threshold 10 --target 22323960=150 --agent
stats — Historical low/high, median, current percentile, volatility, and the weekday the floor is typically lowest, from the local price series._Use it to judge whether today's get-in price is actually low for one event and which day tends to be cheapest._
bash ticketdata-pp-cli stats 22323960 --agent
compare — Rank multiple watched events, or all of one performer's watched events, by get-in price and percent change._Use it to pick the cheapest city or date; for a single event's own history use stats._
bash ticketdata-pp-cli compare --performer ariana-grande --agent
zones — Rank an event's zones by current get-in price and by each zone's drop versus its own history to surface the underpriced section._Use it to pick a section by value; for the plain section-name catalog use events sections._
bash ticketdata-pp-cli zones 22323960 --agent
search — Local full-text search across synced events, performers, and venues, returning multiple matches._Use it to browse offline matches; for the single canonical resolve of a name use performers search or venues search._
bash ticketdata-pp-cli search "ariana" --type performers --agent
This CLI uses Chrome-compatible HTTP transport for browser-facing endpoints. It does not require a resident browser process for normal API calls.
events — Look up ticket events, their get-in price, price history, and section catalog
ticketdata-pp-cli events get — Get an event's current get-in price, forecast, N-day change, marketplace links, and venue/performer detailticketdata-pp-cli events history — Get the full get-in-price time series for an event (hundreds of points), plus per-zone series and on/presale datesticketdata-pp-cli events sections — List the section names catalogued for an eventperformers — Look up performers (artists, teams) and resolve names to canonical performer pages
ticketdata-pp-cli performers get — Get a performer's stats (upcoming events, avg/min/max get-in price) and resale/social linksticketdata-pp-cli performers search — Resolve a search query to the best-matching performervenues — Look up venues and resolve names to canonical venue pages
ticketdata-pp-cli venues get — Get a venue's stats (upcoming events, location)ticketdata-pp-cli venues search — Resolve a search query to the best-matching venueWhen you know what you want to do but not which command does it, ask the CLI directly:
bashticketdata-pp-cli which "<capability in your own words>"
which resolves a natural-language capability query to the best matching command from this CLI's curated feature index. Exit code 0 means at least one match; exit code 2 means no confident match — fall back to --help or use a narrower query.
bashticketdata-pp-cli watch add 22323960 && ticketdata-pp-cli sync && ticketdata-pp-cli stats 22323960 --agent
Track the event, pull its history, and see where today's get-in price sits in its own distribution.
bashticketdata-pp-cli drift --target 22323960=150 --threshold 8 --agent
Flags target hits and any floor that moved more than 8% since the last sync.
bashticketdata-pp-cli compare --performer ariana-grande --agent
Ranks all of the performer's watched events by get-in price and percent change.
bashticketdata-pp-cli events get 22323960 --agent --select tickets.get_in_price,tickets.number_of_listings,forecast_value,price_trend.direction
Pulls only the decision-relevant fields from the verbose event response using dotted select paths.
bashticketdata-pp-cli zones 22323960 --agent
Ranks the event's zones by current get-in price and by each zone's drop versus its own history.
No account or API key. TicketData's public data API is open; this CLI reads it directly.
Run ticketdata-pp-cli doctor to verify setup.
Add --agent to any command. Expands to: --json --compact --no-input --no-color --yes.
--select keeps a subset of fields. Dotted paths descend into nested structures; arrays traverse element-wise. Critical for keeping context small on verbose APIs:bash ticketdata-pp-cli events get mock-value --agent --select id,name,status
--dry-run shows the request without sendingCommands that read from the local store or the API wrap output in a provenance envelope:
json{ "meta": {"source": "live" | "local", "synced_at": "...", "reason": "..."}, "results": <data> }
Parse .results for data and .meta.source to know whether it's live or local. A human-readable N results (live) summary is printed to stderr only when stdout is a terminal AND no machine-format flag (--json, --csv, --compact, --quiet, --plain, --select) is set — piped/agent consumers and explicit-format runs get pure JSON on stdout.
Agents should treat the CLI's path resolver as part of the runtime contract:
--home <dir> for one invocation, or set TICKETDATA_HOME=<dir> to relocate all four path kinds under one root.TICKETDATA_CONFIG_DIR, TICKETDATA_DATA_DIR, TICKETDATA_STATE_DIR, TICKETDATA_CACHE_DIR.--home, TICKETDATA_HOME, XDG (XDG_CONFIG_HOME, XDG_DATA_HOME, XDG_STATE_HOME, XDG_CACHE_HOME), then platform defaults.config contains settings like config.toml and profiles. data contains credentials.toml, data.db, cookies, and auth sidecars. state contains persisted queries, jobs, and teach.log. cache contains regenerable HTTP/cache files.credentials.toml under the data dir. Existing legacy config.toml secrets are read for compatibility and leave config.toml on the first auth write.ticketdata-pp-cli doctor --fail-on warn to surface path and credential-location warnings. agent-context exposes a schema v4 paths block for agents that need the resolved dirs.json { "mcpServers": { "ticketdata": { "command": "ticketdata-pp-mcp", "env": { "TICKETDATA_HOME": "/srv/ticketdata" } } } }
Fleet precedence: an inherited per-kind env var overrides an explicit --home for that kind. Use TICKETDATA_HOME or per-kind vars as durable fleet levers, and use --home only for a single invocation. Relocation is not reversible by unsetting env vars; move files manually before clearing TICKETDATA_HOME, or doctor will not find credentials left under the former root.
This CLI ships a self-capturing learning loop. The CLI does its own bookkeeping: every invocation is journaled locally, a failed flag followed by a corrected retry auto-derives a flag_alias candidate, and a teach on a query family without a playbook auto-synthesizes a playbook_candidate from the session's journal. Your job is judgment only: recall first, act on surfaced candidates, teach the final answer, playbook amend when you observe a correction. You never record failures by hand.
recall before any discoveryBefore list/search/drill commands on a new user question, run:
bashticketdata-pp-cli recall "<user's question>" --agent
The response envelope:
json{ "query": "...", "normalized": "<normalized form>", "query_entities": ["..."], "found": true | false, "match_score": 0.0, "results": [ { "resource_id": "...", "resource_type": "...", "venue": "...", "confidence": 2, "entity_match": "exact|partial|unknown", "source": "taught|preseed|pattern", "warnings": ["..."] } ], "mismatches": [ /* only when --debug-mismatches */ ], "warnings": [ /* top-level */ ], "candidates": [ { "id": 12, "class": "flag_alias | playbook_candidate", "summary": "...", "sightings": 3, "last_seen": "...", "rationale": "...", "next_action": ["<trial command>", "ticketdata-pp-cli learnings confirm 12"] } ], "playbook": { "query_family": "...", "playbook": { "steps": [ { "cmd": "<command with {slot} substitution>", "purpose": "..." } ], "entity_slots": ["$ENTITY"], "expected_tool_calls": 3 }, "slots_resolved": { "$ENTITY": { "token": "<live token>", "canonical": "<canonical>" } }, "notes": "<workarounds + gotchas for this query family>" }, "notes": "<duplicate surface for non-playbook callers>" }
Empty-store short-circuit: if the store has no learnings, playbooks, or candidates yet (recall finds nothing and learnings list and learnings candidates are both empty), skip recall for the rest of this session instead of taxing every query; resume recall-first once something has been taught.
Read candidates, playbook, notes, results[0], and warnings in that order:
if Candidates present (warnings include "candidates_present"):
-> candidates are try-then-confirm, never facts. Follow each candidate's
two-step next_action verbatim: run the trial command first, then run
`learnings confirm <id>` only after the trial verified the behavior.
Reject a wrong candidate with `learnings reject <id>`.
-> NEVER re-teach something recall surfaced as a candidate; confirm or
reject that candidate instead of teaching a duplicate.
-> candidates ride alongside playbooks and resource hits, not instead of
them; continue with the branches below after acting on them.
if Playbook present:
-> READ Playbook.notes verbatim FIRST (workarounds + gotchas the CLI surface doesn't expose)
-> replay Playbook.steps in order, substituting Playbook.slots_resolved entries
for the entity slot tokens. If a step's slot is unresolved, fall back to
discovery for that step only.
-> the Playbook's expected_tool_calls is a budget; if you find yourself running
materially more, record the divergence via `ticketdata-pp-cli playbook amend`
at end-of-session.
elif Notes present (no Playbook):
-> read Notes verbatim before any discovery step; they carry known gotchas
for this query family even when no structured choreography exists yet.
elif Found AND Results[0].EntityMatch == "exact" AND Results[0].Confidence >= 2:
-> skip discovery; fetch live data for Results[*].ResourceID in parallel
elif Found AND Results[0].EntityMatch == "partial":
-> candidate hint, NOT a hit; read the resource title to validate before trusting
elif (any row in Mismatches[] when --debug-mismatches was passed):
-> treat as cold start; the stored learning is for a different entity
(different canonical resolved from query_entities)
else: // Found == false, no playbook, no notes
-> cold start; run discovery normally; teach the answer afterward (Step 4).
If the family has no playbook yet, that teach auto-synthesizes a
playbook candidate from this session's journal - you do not need to
record one by hand.Playbook and Notes are orthogonal to the per-resource path. A recall response can carry both a Playbook AND a Results[] hit - use both: the Playbook tells you which choreography to run; the resource hits short-circuit specific steps. Default to skipping mismatches; pass --debug-mismatches only when investigating cold-start surprises.
Candidate judgment details: learnings confirm <id> prints the candidate's full payload before materializing it - check that the printed payload matches the behavior you verified. learnings reject <id> tombstones the derivation signature so the same candidate does not resurface. The envelope carries only the few candidates worth acting on now; ticketdata-pp-cli learnings candidates lists the full open set.
Graceful degradation: if learnings confirm is an unknown command, you are driving an older binary - ignore the candidates guidance and follow the rest of the protocol.
warningslow_confidence: row exists at confidence<2. Treat as a hint, not a skip-discovery hit.resource_not_in_store: the local store doesn't have the resource the learning points at. The match validator couldn't classify entities — direct-fetch and re-evaluate.cross_alias_match (per-result): the row was taught under a different alias and matched the live query's canonical via entity_lookups (e.g., a "USA" teach satisfying a "United States" recall). Trust the resource_id.similar_shape_different_entity:<canonical> (top-level): a structurally matching row exists but its canonical entity differs from the live query's. Treated as cold start; the warning carries the conflicting canonical as a hint, but the row is NOT promoted into Results.ambiguous_alias (top-level): a single query entity resolved to multiple canonicals (e.g., "Cards" → Arizona Cardinals + St. Louis Cardinals). Surface the ambiguity from context before committing to a resource.candidates_present (top-level): the envelope carries a candidates section. Handle it via the candidates branch in Step 2 before anything else.lookup_refresh_available (top-level): an entity in the query has no lookup row yet, but synced data could provide one. Run ticketdata-pp-cli sync to refresh entity lookups.no_learnings_for_query_family: the table had no rows above the Jaccard floor. Pure cold start.teach & after finalizing your response - alwaysTeaching is unconditional. After resolving a query the store could not answer, background-teach the final resource mapping - no call-count threshold, no judging whether it was "worth" learning. The teach is the anchor of the loop: it triggers playbook synthesis for a family without a playbook, and same-referent phrasings fold into one family so near-duplicate teaches do not fragment the store. Fire it after assembling your user-facing response but BEFORE emitting it, with a shell & so the call returns immediately:
bashticketdata-pp-cli teach --query "<user's question>" --resource-type <type> --resource <id1> --resource <id2> # (append shell `&` to background it)
Silent on success. Errors only land in teach.log under the resolved state dir. Teach the most specific resource - if the user asked a broad question and you walked through parent records to find the specific answer, teach the leaf id, not the parent. The CLI uses seeded entity_lookups for cross-alias resolution at recall time, so a teach under one alias (e.g., "Niners") satisfies future queries under another alias (e.g., "49ers", "San Francisco") automatically.
PII rule: teach the structural question with identifiers stripped - never include names, emails, phone numbers, account ids, or other personal identifiers in taught queries or notes. The CLI scans teach queries for obvious email/phone shapes and warns, but does not block; strip before teaching rather than relying on the warning.
You do not need to decide whether a session "deserves" a playbook: a teach on a family without one auto-synthesizes a playbook_candidate from the session's journal, and the next session judges it via confirm/reject. Attach explicit playbook flags only when you already hold choreography worth recording verbatim - workarounds the CLI didn't surface (silently-dropped flags, undocumented params, pagination tricks, payload gotchas). Prefer the integrated one-call form - record the resource learning and the playbook in the same teach invocation:
bash# Common case: record both the resource learning AND the playbook in one call. ticketdata-pp-cli teach \ --query "<user's question>" \ --resource <id> \ --playbook-file ~/playbooks/<shape>.json \ --playbook-notes-file ~/playbooks/<shape>-notes.md # (append shell `&` to background it) # Alternate: playbook-only (no resource to record alongside). ticketdata-pp-cli teach-playbook \ --query "<user's question>" \ --playbook-file ~/playbooks/<shape>.json \ --notes-file ~/playbooks/<shape>-notes.md
Playbook files are JSON with steps, entity_slots, expected_tool_calls. Notes files are markdown carrying the gotchas verbatim. File-free callers (MCP-only agents) pass the same content inline: --playbook-json and --playbook-notes on the integrated teach form, --playbook-json and --notes on teach-playbook. On the integrated teach form, the playbook flags are optional - omit them entirely for a resource-only teach. On the standalone teach-playbook form, at least one of the playbook and notes flags must be set; both empty is rejected. Playbooks are keyed on the structural query family (entities stripped) so a recipe taught from one entity-shaped query applies to every other query of the same shape, with slots_resolved binding the live query's canonical at recall time.
When you DO find a playbook on a future recall, treat it as ground truth: replay the steps with slots_resolved substitutions, skip the discovery that the choreography already documents, and read notes before any step.
playbook amend & when your debug response identifies a correctionIf your debug-protocol response identifies a concrete correction the notes or playbook should know — a workaround, an undocumented endpoint shape, a stale field name, observed schema drift, an empty-payload fallback — fire playbook amend BEFORE emitting your user-facing response. Same fire-and-forget posture as teach.
bashticketdata-pp-cli playbook amend \ --query "<exact recall query string>" \ --add-note "<your concrete correction>" # (append shell `&` to background it)
What counts as worth amending: a behavior you OBSERVED this session that future-you would benefit from knowing. Examples worth amending:
{meta, results}, payload nested two levels deeper than the docs claim).What does NOT belong in notes:
The amend command appends to the family's existing notes with a timestamped marker ([amend YYYY-MM-DDTHH:MMZ]: <text>). Multiple amends accumulate; the audit trail is visible. If no playbook exists yet for the family, amend creates a notes-only one (so cold-start corrections still land).
playbook amend notes are designed to potentially flow upstream as shared knowledge in future versions of the Printing Press. Keep them clean of user-identifying content so the upstream-contribution path stays open without retroactive scrubbing:
If a correction is only meaningful with user-specific context, it belongs in a personal note, not in the playbook amend.
ticketdata-pp-cli learnings stats reports recall hit rate, teach-to-reuse, playbook resolution rate, and candidate confirm/reject counts from the local learn_events table. Rates are null until they have a denominator; everything stays on this machine. Use it to check whether the loop is earning its keep for this CLI.
--no-learn on a single command short-circuits both recall and the teach write path. Use for deterministic agent flows or tests that must not be affected by accumulated learnings.TICKETDATA_NO_LEARN=true in the environment globally disables the pipeline.When you (or the agent) notice something off about this CLI, record it:
ticketdata-pp-cli feedback "the --since flag is inclusive but docs say exclusive"
ticketdata-pp-cli feedback --stdin < notes.txt
ticketdata-pp-cli feedback list --json --limit 10Entries are stored locally as feedback.jsonl under the resolved data dir. They are never POSTed unless TICKETDATA_FEEDBACK_ENDPOINT is set AND either --send is passed or TICKETDATA_FEEDBACK_AUTO_SEND=true. Default behavior is local-only.
Write what surprised you, not a bug report. Short, specific, one line: that is the part that compounds.
Every command accepts --deliver <sink>. The output goes to the named sink in addition to (or instead of) stdout, so agents can route command results without hand-piping. Three sinks are supported:
| Sink | Effect | |------|--------| | stdout | Default; write to stdout only | | file:<path> | Atomically write output to <path> (tmp + rename) | | webhook:<url> | POST the output body to the URL (application/json or application/x-ndjson when --compact) |
Unknown schemes are refused with a structured error naming the supported set. Webhook failures return non-zero and log the URL + HTTP status on stderr.
A profile is a saved set of flag values, reused across invocations. Use it when a scheduled agent calls the same command every run with the same configuration - HeyGen's "Beacon" pattern.
ticketdata-pp-cli profile save briefing --json
ticketdata-pp-cli --profile briefing events get mock-value
ticketdata-pp-cli profile list --json
ticketdata-pp-cli profile show briefing
ticketdata-pp-cli profile delete briefing --yesExplicit flags always win over profile values; profile values win over defaults. agent-context lists all available profiles under available_profiles so introspecting agents discover them at runtime.
| Code | Meaning | |------|---------| | 0 | Success | | 2 | Usage error (wrong arguments) | | 3 | Resource not found | | 5 | API error (upstream issue) | | 7 | Rate limited (wait and retry) | | 10 | Config error |
Parse $ARGUMENTS:
help, or --help → show ticketdata-pp-cli --help outputinstall → ends with mcp → MCP installation; otherwise → see Prerequisites above--agent)bash go install github.com/mvanhorn/printing-press-library/library/media-and-entertainment/ticketdata/cmd/ticketdata-pp-mcp@latest
bash claude mcp add ticketdata-pp-mcp -- ticketdata-pp-mcp
claude mcp listwhich ticketdata-pp-cliIf not found, offer to install (see Prerequisites at the top of this skill).
--agent flag:bash ticketdata-pp-cli <command> [subcommand] [args] --agent
ticketdata-pp-cli <command> --help.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 45,537 | 67,425 | +48% | 1 | 1 | 0% | 1,811 | 7,583 | +319% | 0 | 0 | — |
case-02 | fail→fail | 8,069 | 6,516 | -19% | 1 | 1 | 0% | 966 | 7,479 | +674% | 0 | 0 | — |
case-03 | fail→fail | 7,344 | 6,549 | -11% | 1 | 1 | 0% | 1,053 | 7,364 | +599% | 0 | 0 | — |
case-04 | fail→fail | 3,792 | 7,302 | +93% | 1 | 1 | 0% | 493 | 7,543 | +1430% | 0 | 0 | — |
case-05 | fail→fail | 12,776 | 5,521 | -57% | 1 | 1 | 0% | 1,200 | 7,349 | +512% | 0 | 0 | — |
case-06 | pass→pass | 4,702 | 6,854 | +46% | 1 | 1 | 0% | 856 | 8,075 | +843% | 0 | 0 | — |
case-07 | fail→fail | 15,101 | 36,682 | +143% | 1 | 1 | 0% | 3,145 | 7,510 | +139% | 0 | 0 | — |
case-08 | fail→fail | 12,634 | 6,880 | -46% | 1 | 1 | 0% | 1,977 | 7,454 | +277% | 0 | 0 | — |
case-09 | fail→fail | 8,315 | 4,914 | -41% | 1 | 1 | 0% | 1,195 | 7,331 | +513% | 0 | 0 | — |
case-10 | fail→fail | 5,122 | 7,124 | +39% | 1 | 1 | 0% | 194 | 7,450 | +3740% | 0 | 0 | — |
case-11 | pass→pass | 18,531 | 3,441 | -81% | 1 | 1 | 0% | 3,059 | 7,757 | +154% | 0 | 0 | — |
case-12 | fail→fail | 6,228 | 6,220 | -0% | 1 | 1 | 0% | 1,138 | 7,385 | +549% | 0 | 0 | — |
case-13 | fail→pass | 7,860 | 5,081 | -35% | 1 | 1 | 0% | 1,531 | 7,394 | +383% | 0 | 0 | — |
case-14 | fail→fail | 21,161 | 5,393 | -75% | 1 | 1 | 0% | 3,400 | 7,316 | +115% | 0 | 0 | — |
case-15 | fail→pass | 11,739 | 2,769 | -76% | 1 | 1 | 0% | 1,886 | 7,583 | +302% | 0 | 0 | — |
case-16 | fail→fail | 1,497 | 5,659 | +278% | 1 | 1 | 0% | 222 | 7,395 | +3231% | 0 | 0 | — |
case-17 | fail→fail | 2,624 | 6,734 | +157% | 1 | 1 | 0% | 431 | 7,478 | +1635% | 0 | 0 | — |
case-18 | fail→pass | 1,498 | 4,516 | +201% | 1 | 1 | 0% | 228 | 7,380 | +3137% | 0 | 0 | — |
case-19 | fail→pass | 10,957 | 5,091 | -54% | 1 | 1 | 0% | 1,648 | 7,348 | +346% | 0 | 0 | — |
case-20 | fail→fail | 6,456 | 4,538 | -30% | 1 | 1 | 0% | 1,023 | 7,312 | +615% | 0 | 0 | — |
case-21 | fail→fail | 9,636 | 6,346 | -34% | 1 | 1 | 0% | 978 | 7,451 | +662% | 0 | 0 | — |
case-22 | pass→fail | 10,697 | 5,060 | -53% | 1 | 1 | 0% | 1,679 | 7,387 | +340% | 0 | 0 | — |
case-23 | fail→fail | 17,707 | 6,016 | -66% | 1 | 1 | 0% | 2,196 | 7,414 | +238% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 6 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +13 percentage points is the difference between those two pass rates over the 6 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.