Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Machine learning and LLM engineering judgment, distilled from a stronger model - invoke when DECIDING whether/how to use ML or an LLM for a task (prompt vs RAG vs fine-tune vs classical); working with training/eval data or labels; building or reviewing evals for models and LLM features; designing RAG, structured output, or agent pipelines; or diagnosing why a model/LLM feature underperforms. Method-selection ladder, data and leakage discipline, eval-as-spec rules, LLM-era craft, and a trap catal
.claude/skills/telagod-ml/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 3% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -4% | 0% |
| case-05 | ✓→✓ | = Same ✓ | -14% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 7% | 0% |
Rule content lives in the five files below; this SKILL.md only routes (doctrine/04-maintenance.md governs edits to this bundle too).
| You are about to… | Read (in this folder) | |---|---| | Decide whether ML/an LLM is warranted, and which method rung to use | approach.md | | Touch a dataset, labels, or splits; suspect a score is too good | data.md | | Define success, build/judge an eval, or assess someone's metric claim | evals.md | | Build with LLMs: prompts, RAG, structured output, agents, model choice | llm.md | | Diagnose an underperforming model or LLM feature | data.md §1 first (read real failures), then llm.md §3 if RAG, traps.md to name the pattern | | Review an ML project's health; name why a claim or pipeline smells wrong | traps.md |
A new ML feature usually runs approach.md (interrogate + pick the rung) → evals.md §1 (eval BEFORE build) → data.md → then llm.md if the rung is LLM-shaped → skim traps.md §C before finalizing any launch or monitoring plan.
Modeling and evaluation judgment. The serving infrastructure around a model is ordinary backend (backend bundle: APIs, queues, operate.md); experiment execution discipline is methods (investigate/verify); whether to delegate → doctrine.
The eval is the spec; anything unmeasured is folklore. Look at the data with your own eyes (data.md §1), climb the method ladder from the cheapest rung (approach.md §3), and treat every surprising score as leakage until disproven (data.md §2). The failure mode of this field is not bad models — it is unearned confidence in numbers.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | pass→pass | 15,458 | 12,214 | -21% | 1 | 1 | 0% | 2,201 | 2,277 | +3% | 0 | 0 | — |
case-04 | pass→pass | 17,632 | 13,871 | -21% | 1 | 1 | 0% | 2,536 | 2,430 | -4% | 0 | 0 | — |
case-05 | pass→pass | 17,605 | 12,910 | -27% | 1 | 1 | 0% | 2,681 | 2,319 | -14% | 0 | 0 | — |
case-01 | fail→pass | 22,134 | 18,168 | -18% | 1 | 1 | 0% | 3,326 | 2,874 | -14% | 0 | 0 | — |
case-02 | pass→pass | 23,066 | 22,516 | -2% | 1 | 1 | 0% | 3,651 | 3,919 | +7% | 0 | 0 | — |
case-03 | pass→pass | 18,741 | 24,041 | +28% | 1 | 1 | 0% | 3,244 | 4,608 | +42% | 0 | 0 | — |
case-07 | pass→pass | 14,073 | 12,962 | -8% | 1 | 1 | 0% | 2,110 | 2,301 | +9% | 0 | 0 | — |
case-08 | pass→pass | 13,979 | 11,100 | -21% | 1 | 1 | 0% | 1,848 | 1,903 | +3% | 0 | 0 | — |
case-09 | pass→pass | 12,868 | 9,952 | -23% | 1 | 1 | 0% | 1,959 | 1,859 | -5% | 0 | 0 | — |
case-10 | pass→pass | 16,204 | 12,943 | -20% | 1 | 1 | 0% | 2,742 | 2,386 | -13% | 0 | 0 | — |
case-11 | pass→pass | 13,577 | 12,118 | -11% | 1 | 1 | 0% | 2,207 | 2,308 | +5% | 0 | 0 | — |
case-12 | pass→pass | 10,637 | 10,425 | -2% | 1 | 1 | 0% | 1,783 | 2,115 | +19% | 0 | 0 | — |
case-13 | pass→pass | 11,126 | 10,303 | -7% | 1 | 1 | 0% | 1,817 | 2,182 | +20% | 0 | 0 | — |
case-14 | pass→pass | 11,213 | 11,153 | -1% | 1 | 1 | 0% | 1,603 | 2,053 | +28% | 0 | 0 | — |
case-15 | pass→pass | 18,129 | 14,715 | -19% | 1 | 1 | 0% | 2,562 | 2,570 | +0% | 0 | 0 | — |
case-16 | pass→pass | 14,210 | 12,822 | -10% | 1 | 1 | 0% | 2,144 | 2,286 | +7% | 0 | 0 | — |
case-17 | pass→pass | 17,574 | 11,860 | -33% | 1 | 1 | 0% | 2,553 | 2,287 | -10% | 0 | 0 | — |
case-18 | pass→pass | 9,933 | 11,085 | +12% | 1 | 1 | 0% | 1,569 | 2,049 | +31% | 0 | 0 | — |
case-19 | pass→pass | 10,615 | 7,365 | -31% | 1 | 1 | 0% | 1,400 | 1,480 | +6% | 0 | 0 | — |
case-20 | pass→pass | 11,293 | 8,620 | -24% | 1 | 1 | 0% | 1,699 | 1,681 | -1% | 0 | 0 | — |
case-21 | pass→pass | 9,398 | 8,939 | -5% | 1 | 1 | 0% | 1,361 | 1,725 | +27% | 0 | 0 | — |
case-22 | pass→pass | 15,431 | 11,774 | -24% | 1 | 1 | 0% | 2,253 | 2,209 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.