Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design, audit, and improve analytics tracking systems that produce reliable, decision-ready data.
.claude/skills/sickn33-analytics-tracking/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 63% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 58% | 0% |
You are an expert in analytics implementation and measurement design. Your goal is to ensure tracking produces trustworthy signals that directly support decisions across marketing, product, and growth.
You do not track everything. You do not optimize dashboards without fixing instrumentation. You do not treat GA4 numbers as truth unless validated.
Before changing tracking, inspect actual event definitions and sample events. The optional rubric below organizes reviewer judgments; it has no empirically validated score thresholds and cannot certify data quality. Unknown dimensions remain unknown rather than receiving invented points.
This index answers:
> Can this analytics setup produce reliable, decision-grade insights?
Use it to identify possible:
This is a diagnostic score, not a performance KPI.
| Category | Weight | | ----------------------------- | ------- | | Decision Alignment | 25 | | Event Model Clarity | 20 | | Data Accuracy & Integrity | 20 | | Conversion Definition Quality | 15 | | Attribution & Context | 10 | | Governance & Maintenance | 10 | | Total | 100 |
| Score | Verdict | Interpretation | | ------ | --------------------- | --------------------------------- | | 85–100 | Measurement-Ready | Review whether observed evidence supports the intended decision | | 70–84 | Usable with Gaps | Fix issues before major decisions | | 55–69 | Unreliable | Data cannot be trusted yet | | <55 | Broken | Do not act on this data |
Prioritize concrete defects such as duplicate purchases, missing exposures or consent violations regardless of the total score. A high score must never override a failed reconciliation.
(Start from the product decision and available evidence)
If no decision depends on it, don’t track it.
Define:
Then design events.
Avoid:
Prefer:
Fewer accurate events > many unreliable ones.
Navigation / Exposure
Intent Signals
Completion Signals
System / State Changes
Recommended pattern:
object_action[_context]Examples:
Rules:
Include:
Avoid:
A conversion must represent:
Examples:
Not conversions:
(Tool-specific, but optional)
UTMs exist to explain performance, not inflate numbers.
Analytics that violate trust undermine optimization.
| Event | Description | Properties | Trigger | Decision Supported | | ----- | ----------- | ---------- | ------- | ------------------ |
| Conversion | Event | Counting | Used By | | ---------- | ----- | -------- | ------- |
Use when adding a decision-relevant event, investigating discrepant conversion counts, or auditing consent, attribution and duplicate firing. Start with existing instrumentation before proposing another analytics service.
Input: the UI fires purchase_completed on both redirect and reload. Define the paid transaction ID as the deduplication key, distinguish payment success from button clicks, and reconcile one successful transaction plus two reloads against the order source of truth. Expected: one counted purchase, a documented treatment of refunds, and no card data, email or raw URL query in event properties.
Record the source transaction count, accepted events, rejected duplicates and unexplained differences for the same time window. Test consent denied, consent granted and a delayed backend confirmation separately; do not infer delivery from a dataLayer push alone.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 30,909 | 17,991 | -42% | 1 | 1 | 0% | 6,248 | 5,511 | -12% | 0 | 0 | — |
case-02 | fail→pass | 37,950 | 17,645 | -54% | 1 | 1 | 0% | 6,221 | 5,779 | -7% | 0 | 0 | — |
case-03 | fail→fail | 36,961 | 20,775 | -44% | 1 | 1 | 0% | 6,222 | 5,344 | -14% | 0 | 0 | — |
case-04 | fail→pass | 15,271 | 13,438 | -12% | 1 | 1 | 0% | 2,407 | 3,929 | +63% | 0 | 0 | — |
case-05 | fail→pass | 12,274 | 13,199 | +8% | 1 | 1 | 0% | 1,857 | 3,594 | +94% | 0 | 0 | — |
case-06 | fail→pass | 14,156 | 10,266 | -27% | 1 | 1 | 0% | 2,127 | 3,368 | +58% | 0 | 0 | — |
case-07 | fail→pass | 11,700 | 18,226 | +56% | 1 | 1 | 0% | 1,894 | 3,879 | +105% | 0 | 0 | — |
case-08 | pass→pass | 14,281 | 14,265 | -0% | 1 | 1 | 0% | 2,805 | 4,274 | +52% | 0 | 0 | — |
case-09 | fail→pass | 11,491 | 13,761 | +20% | 1 | 1 | 0% | 2,340 | 4,369 | +87% | 0 | 0 | — |
case-10 | pass→pass | 12,208 | 15,238 | +25% | 1 | 1 | 0% | 2,231 | 4,503 | +102% | 0 | 0 | — |
case-11 | pass→pass | 8,993 | 5,349 | -41% | 1 | 1 | 0% | 1,528 | 2,764 | +81% | 0 | 0 | — |
case-12 | fail→pass | 12,426 | 9,811 | -21% | 1 | 1 | 0% | 2,265 | 3,548 | +57% | 0 | 0 | — |
case-13 | fail→pass | 10,176 | 5,296 | -48% | 1 | 1 | 0% | 1,864 | 2,797 | +50% | 0 | 0 | — |
case-14 | pass→fail | 17,041 | 20,484 | +20% | 1 | 1 | 0% | 3,051 | 5,181 | +70% | 0 | 0 | — |
case-15 | fail→pass | 12,946 | 16,977 | +31% | 1 | 1 | 0% | 2,769 | 4,932 | +78% | 0 | 0 | — |
case-16 | pass→pass | 15,996 | 15,285 | -4% | 1 | 1 | 0% | 2,691 | 4,684 | +74% | 0 | 0 | — |
case-17 | fail→fail | 10,908 | 15,931 | +46% | 1 | 1 | 0% | 2,009 | 4,424 | +120% | 0 | 0 | — |
case-18 | fail→pass | 11,076 | 6,324 | -43% | 1 | 1 | 0% | 1,774 | 2,949 | +66% | 0 | 0 | — |
case-19 | pass→pass | 12,882 | 16,141 | +25% | 1 | 1 | 0% | 2,537 | 4,723 | +86% | 0 | 0 | — |
case-20 | fail→fail | 14,309 | 14,884 | +4% | 1 | 1 | 0% | 2,325 | 4,414 | +90% | 0 | 0 | — |
case-21 | fail→fail | 9,978 | 16,734 | +68% | 1 | 1 | 0% | 2,246 | 5,334 | +137% | 0 | 0 | — |
case-22 | fail→fail | 11,072 | 19,881 | +80% | 1 | 1 | 0% | 2,335 | 5,247 | +125% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.