Install any skill in seconds. Free to start, no credit card required.
Get Started Free →YOLO detection, segmentation, classification, and pose estimation setup, training workflow, evaluation metrics, and XAI verification
.claude/skills/aeren23-yolo-pipeline/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 212% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 39% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 111% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 47% | 0% |
What output do you need?
├── "What objects are here and WHERE?" (bounding boxes)
│ └── ✅ Detection — model('img.jpg') with yolov8n.pt
│
├── "Exact pixel-level shape of each object"
│ └── ✅ Segmentation — model('img.jpg') with yolov8n-seg.pt
│
├── "What IS this image overall?" (single label)
│ └── ✅ Classification — model('img.jpg') with yolov8n-cls.pt
│
├── "What pose is this person in?" (17 keypoints)
│ └── ✅ Pose Estimation — model('img.jpg') with yolov8n-pose.pt
│
└── "Rotated/angled objects" (ships, aircraft)
└── ✅ Oriented Bounding Boxes (OBB) — yolov8n-obb.pt| Type | Question | Output | Can Count Individuals? | |------|----------|--------|------------------------| | Semantic | "What’s in the scene?" | One color per class | ❌ No (5 sheep = one green blob) | | Instance (YOLO) | "Which objects where?" | Unique ID per object | ✅ Yes (5 sheep = 5 different colors) |
> YOLO does Instance Segmentation — each object gets its own mask and identity.
| Dataset | Source | Classes | Typical Use | |---------|--------|---------|-------------| | COCO | Microsoft | 80 | General object detection | | ImageNet | Stanford | 1000 | Image classification | | DOTAv1 | Wuhan Univ. | 15 | Aerial/satellite OBB |
| Condition | Classic CV | YOLO/DL | |-----------|-----------|----------| | High contrast, uniform light | ✅ Works | Overkill | | Shadows, uneven lighting | ❌ Fails | ✅ Robust | | Touching/overlapping objects | ❌ Fails | ✅ Handles | | Diverse viewpoints | ❌ Unreliable | ✅ Generalizes | | Need to detect 80+ classes | ❌ Impractical | ✅ Built-in |
[center_x, center_y, width, height] (normalized 0-1).txt per image): [class_index] [center_x] [center_y] [width] [height] Example: 1 0.45 0.60 0.30 0.40 All coordinates normalized to 0-1 range.
dataset/ ├── train/ │ ├── COVID/ ← folder name = class label │ └── Normal/ └── val/ ├── COVID/ └── Normal/
| Parameter | What It Controls | Guidance | |-----------|-----------------|----------| | Learning Rate | Step size for weight updates | Too high = overshoots, too low = slow convergence | | Epochs | Full passes through dataset | More is not always better (overfitting risk) | | Batch Size | Images per training step | Limited by GPU memory. 16-32 typical | | imgsz | Input image resolution | 640 default. Higher = better accuracy, slower |
Don't train from scratch. Use pre-trained weights:
pythonfrom ultralytics import YOLO model = YOLO('yolov8n.pt') # Pre-trained on COCO model.train(data='my_data.yaml', epochs=100, imgsz=640)
> Transfer Learning = standing on giants' shoulders. The model already knows edges, textures, shapes from millions of images. You only retrain the final layers for your specific task.
| File | Contains | Use For | |------|----------|--------| | best.pt | Weights at best validation accuracy | ✅ Production / inference | | last.pt | Weights at final epoch | Resume interrupted training only |
> Always use best.pt for inference. last.pt may have started overfitting.
| | Predicted Positive | Predicted Negative | |---|---|---| | Actually Positive | TP (correct detection) | FN (missed detection — WORST in medical) | | Actually Negative | FP (false alarm) | TN (not used in object detection) |
> TN is not used in object detection systems — "correctly identifying nothing" is not meaningful.
IoU = Area of Overlap / Area of Union| Metric | Threshold | Strictness | |--------|-----------|------------| | mAP@0.50 | 50% overlap required | Standard | | mAP@0.50:0.95 | Average across 50%-95% | Very strict, comprehensive |
Precision = TP / (TP + FP) — "Of all detections, how many are correct?"
Recall = TP / (TP + FN) — "Of all real objects, how many did we find?"
F1 = 2 × (P × R) / (P + R) — Harmonic meanIf training loss keeps dropping but validation loss plateaus or rises → overfitting. Solutions:
DL models are black boxes — they give answers but don't explain why.
| Heatmap Color | Meaning | |---------------|--------| | 🔴 Hot (Red/Yellow) | Model's decision focus — high attention | | 🔵 Cold (Blue/Purple) | Irrelevant to model's decision |
Why it matters in medicine:
pythonfrom ultralytics import YOLO model = YOLO('yolov8n.pt') results = model( source='image.jpg', # or '0' for webcam, or video path show=True, # display results save=True, # save to disk conf=0.25, # minimum confidence threshold )
| Light Type | Setup | Best For | |-----------|-------|----------| | Array/Screen | Front-facing flat | Surface defects | | Ring Light | Around camera lens | Small parts, uniform light | | Back Light | Behind object | Silhouette, hole detection | | Bar Light | Angled beam | Directional surface inspection | | Dome Light | Surrounds object | Eliminating reflections on metal/shiny surfaces |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,330 | 4,605 | -72% | 1 | 1 | 0% | 921 | 2,876 | +212% | 0 | 0 | — |
case-02 | pass→pass | 11,565 | 5,232 | -55% | 1 | 1 | 0% | 2,191 | 3,040 | +39% | 0 | 0 | — |
case-03 | pass→pass | 8,038 | 3,677 | -54% | 1 | 1 | 0% | 1,271 | 2,685 | +111% | 0 | 0 | — |
case-04 | pass→pass | 8,721 | 4,504 | -48% | 1 | 1 | 0% | 2,069 | 3,041 | +47% | 0 | 0 | — |
case-05 | pass→pass | 7,457 | 3,434 | -54% | 1 | 1 | 0% | 1,502 | 2,794 | +86% | 0 | 0 | — |
case-06 | pass→pass | 7,104 | 2,365 | -67% | 1 | 1 | 0% | 1,490 | 2,437 | +64% | 0 | 0 | — |
case-11 | pass→pass | 3,722 | 2,983 | -20% | 1 | 1 | 0% | 782 | 2,596 | +232% | 0 | 0 | — |
case-07 | pass→pass | 9,307 | 3,578 | -62% | 1 | 1 | 0% | 1,880 | 2,699 | +44% | 0 | 0 | — |
case-08 | pass→pass | 12,648 | 6,649 | -47% | 1 | 1 | 0% | 2,508 | 3,293 | +31% | 0 | 0 | — |
case-09 | pass→pass | 3,059 | 2,679 | -12% | 1 | 1 | 0% | 618 | 2,514 | +307% | 0 | 0 | — |
case-10 | pass→pass | 9,146 | 6,172 | -33% | 1 | 1 | 0% | 1,703 | 3,113 | +83% | 0 | 0 | — |
case-12 | pass→pass | 9,208 | 5,913 | -36% | 1 | 1 | 0% | 1,855 | 3,096 | +67% | 0 | 0 | — |
case-13 | pass→pass | 13,675 | 5,719 | -58% | 1 | 1 | 0% | 2,298 | 3,014 | +31% | 0 | 0 | — |
case-14 | pass→pass | 5,728 | 2,360 | -59% | 1 | 1 | 0% | 1,052 | 2,401 | +128% | 0 | 0 | — |
case-15 | pass→pass | 3,346 | 3,311 | -1% | 1 | 1 | 0% | 679 | 2,626 | +287% | 0 | 0 | — |
case-21 | pass→pass | 11,550 | 9,851 | -15% | 1 | 1 | 0% | 2,606 | 4,075 | +56% | 0 | 0 | — |
case-16 | pass→pass | 5,936 | 2,508 | -58% | 1 | 1 | 0% | 1,168 | 2,461 | +111% | 0 | 0 | — |
case-17 | pass→pass | 4,643 | 2,878 | -38% | 1 | 1 | 0% | 883 | 2,590 | +193% | 0 | 0 | — |
case-18 | pass→pass | 10,104 | 7,457 | -26% | 1 | 1 | 0% | 1,883 | 3,344 | +78% | 0 | 0 | — |
case-19 | pass→pass | 12,373 | 5,608 | -55% | 1 | 1 | 0% | 2,195 | 3,005 | +37% | 0 | 0 | — |
case-20 | pass→pass | 12,316 | 10,595 | -14% | 1 | 1 | 0% | 2,796 | 4,199 | +50% | 0 | 0 | — |
case-22 | fail→pass | 10,805 | 13,175 | +22% | 1 | 1 | 0% | 2,425 | 4,319 | +78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.