Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design or audit the uncertainty-quantification, out-of-distribution (OOD) detection, and selective-prediction layer of a medical-imaging model framed for deployment — so a clinical-use claim carries calibrated per-case uncertainty (MC-dropout / deep ensemble / conformal / Bayesian), an OOD guard validated on a held-out OOD set, an abstention rule at a pre-specified operating point, and uncertainty checked under distribution shift. Emits an uncertainty manifest and a deterministic gate that flags
.claude/skills/aperivue-uncertainty-imaging/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -35% | 0% |
A medical-imaging model framed for deployment must say more than "class 1, 0.87". It needs a calibrated uncertainty on each case, an out-of-distribution (OOD) guard validated on data known to be out-of-distribution, and — if it abstains — a pre-specified operating point. The failures are predictable and reviewer-visible: a clinical-use claim built on point predictions, conformal intervals quoted without ever measuring their coverage, an "OOD detector" evaluated only on in-distribution data, a deep ensemble whose members share a seed, and uncertainty validated only in-distribution when deployment sees scanner/site/case-mix shift. This skill designs that layer and audits an existing one (Gal 2016; Lakshminarayanan 2017; Angelopoulos & Bates; Ovadia 2019; DECIDE-AI).
It is the deployment-safety companion in the model-engineering lane: /model-evaluation computes the held-out metrics and calibration, and uncertainty-imaging covers the uncertainty / OOD / abstention machinery a deployment claim rests on. It integrates MAPIE (conformal), captum, and pretrained OOD scorers; it does not reimplement them and never runs a model on real patient data.
unsure, or off-distribution?"
shift checks right before submission.
/model-evaluation then/analyze-stats.
/model-scaffold (+ /model-validation)./explainability./radiomics-ml + /analyze-stats.all — add MC-dropout, a deep ensemble, conformal prediction, or a Bayesian estimate.
can fail on clinical data — measure achieved coverage on a held-out calibration/test set.
until you evaluate on data known to be out-of-distribution (different scanner / site / pathology).
underestimates epistemic uncertainty.
pass is identical and the estimate collapses to a point prediction.
pre-specify the coverage / risk operating point.
degrades under shift, so report it on shifted / external data.
the strongest default when a calibration set is available. Validate empirical coverage.
best-quality epistemic uncertainty, at K× cost.
See references/uncertainty_guide.md.
on a held-out OOD set (different scanner/site/pathology) and report detection AUROC + the operating point.
target coverage or risk; report the risk–coverage curve.
Report calibration / coverage on shifted or external data, not in-distribution only (Ovadia 2019).
json{ "task": "classification", "deployment_claim": true, "uncertainty_method": "conformal", "coverage_target": 0.90, "coverage_validated": true, "ood_method": "mahalanobis", "ood_heldout_set": "external-ood-cohort", "selective_prediction": true, "selective_target": 0.95, "calibration_under_shift": true }
bashpython3 scripts/check_uncertainty_reporting.py --manifest uncertainty_manifest.json --strict
Verdicts: POINT_PREDICTION_NO_UNCERTAINTY, CONFORMAL_NO_COVERAGE_VALIDATION, OOD_NO_HELDOUT_SET (Major); ENSEMBLE_NOT_INDEPENDENT, MCDROPOUT_DISABLED_AT_INFERENCE, SELECTIVE_NO_TARGET, NO_CALIBRATION_UNDER_SHIFT (Minor). Audits the declared spec at design/report time; it complements /model-evaluation's executed calibration/subgroup metrics.
/model-evaluation — the point predictor's held-out metrics + calibration this layer sits on top of./analyze-stats — calibration curve / risk–coverage plotting for the report./check-reporting — TRIPOD+AI / DECIDE-AI deployment-monitoring items./model-validation — the DECIDE-AI monitoring seam (the deployment-time counterpart of the splitaudit).
reported number comes from the researcher's executed code — never invented. This skill designs and audits the uncertainty spec; it does not run a model on real patient data.
clinical data (CONFORMAL_NO_COVERAGE_VALIDATION).
check_uncertainty_reporting.py. Theverdict is reproduced deterministically, never asserted from prose.
or OOD library or claim results for one.
scripts/check_uncertainty_reporting_challenge/ ships a synthetic weak/strong uncertainty-manifest pair with a network-free verify.sh wired into the skill's validation commands.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,854 | 4,631 | -83% | 1 | 1 | 0% | 4,671 | 2,115 | -55% | 0 | 0 | — |
case-02 | fail→fail | 29,983 | 8,553 | -71% | 1 | 1 | 0% | 5,877 | 2,019 | -66% | 0 | 0 | — |
case-03 | fail→fail | 35,599 | 6,226 | -83% | 1 | 1 | 0% | 6,779 | 2,068 | -69% | 0 | 0 | — |
case-04 | pass→pass | 18,659 | 9,438 | -49% | 1 | 1 | 0% | 3,612 | 3,455 | -4% | 0 | 0 | — |
case-05 | pass→fail | 24,725 | 18,051 | -27% | 1 | 1 | 0% | 4,124 | 4,596 | +11% | 0 | 0 | — |
case-06 | pass→pass | 25,017 | 22,109 | -12% | 1 | 1 | 0% | 4,889 | 6,004 | +23% | 0 | 0 | — |
case-07 | fail→fail | 20,727 | 7,943 | -62% | 1 | 1 | 0% | 3,489 | 2,168 | -38% | 0 | 0 | — |
case-08 | fail→fail | 16,795 | 8,125 | -52% | 1 | 1 | 0% | 2,735 | 2,179 | -20% | 0 | 0 | — |
case-09 | fail→pass | 14,233 | 13,235 | -7% | 1 | 1 | 0% | 2,509 | 4,165 | +66% | 0 | 0 | — |
case-10 | fail→fail | 11,793 | 8,270 | -30% | 1 | 1 | 0% | 2,056 | 3,226 | +57% | 0 | 0 | — |
case-11 | fail→pass | 12,421 | 6,579 | -47% | 1 | 1 | 0% | 1,985 | 2,872 | +45% | 0 | 0 | — |
case-12 | fail→pass | 14,423 | 13,198 | -8% | 1 | 1 | 0% | 2,310 | 3,990 | +73% | 0 | 0 | — |
case-13 | fail→pass | 37,972 | 10,749 | -72% | 1 | 1 | 0% | 2,952 | 3,805 | +29% | 0 | 0 | — |
case-14 | fail→pass | 26,969 | 11,106 | -59% | 1 | 1 | 0% | 5,376 | 3,481 | -35% | 0 | 0 | — |
case-15 | fail→pass | 11,659 | 3,321 | -72% | 1 | 1 | 0% | 1,637 | 2,298 | +40% | 0 | 0 | — |
case-16 | fail→pass | 18,500 | 10,944 | -41% | 1 | 1 | 0% | 3,229 | 3,618 | +12% | 0 | 0 | — |
case-17 | pass→pass | 27,253 | 8,577 | -69% | 1 | 1 | 0% | 2,051 | 3,235 | +58% | 0 | 0 | — |
case-18 | pass→pass | 12,254 | 11,114 | -9% | 1 | 1 | 0% | 2,018 | 3,547 | +76% | 0 | 0 | — |
case-19 | pass→pass | 16,963 | 8,796 | -48% | 1 | 1 | 0% | 2,895 | 3,126 | +8% | 0 | 0 | — |
case-20 | pass→pass | 21,775 | 16,518 | -24% | 1 | 1 | 0% | 3,507 | 4,534 | +29% | 0 | 0 | — |
case-21 | fail→pass | 18,047 | 13,508 | -25% | 1 | 1 | 0% | 2,957 | 3,923 | +33% | 0 | 0 | — |
case-22 | pass→fail | 18,523 | 12,879 | -30% | 1 | 1 | 0% | 3,007 | 3,874 | +29% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 17 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.