Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manage unified memory and thermals during long-running ML jobs on NVIDIA DGX Spark. Use when planning memory headroom for a training run on GB10, when a job OOMs on unified memory, or when monitoring temperature and power during multi-hour training.
.claude/skills/wshobson-spark-memory-thermal-ops/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 32% | 0% |
DGX Spark's GB10 chip has one 128GB unified memory (UMA) pool shared by CPU and GPU, and a sustained power ceiling well below its rated figure. Both break discrete-GPU assumptions: headroom isn't what nvidia-smi reports, and a run that starts fast will slow down mid-job with nothing misconfigured. This skill covers planning memory headroom, working an actual OOM, and watching thermals across a long job. For launch-time failure modes (ABI mismatches, flash-attn, playbook breakage), see spark-training-gotchas — this skill assumes the job starts.
| Situation | Do this | |---|---| | Planning headroom before launch | Budget against free -g, not nvidia-smi — see UMA Memory Model | | Job OOMs on unified memory | Work the OOM Ladder in order: flush, then batch/pack, then method downgrade | | Throughput drops mid-run | Check the power/temp log before assuming a config bug — see Thermal Monitoring | | Trainer + inference server both wanted | Run one at a time — see Concurrent Workloads |
before launch — will this model, method, and batch/pack combination fit.
remediation order matters — what to try first, second, third.
multi-hour job, deciding whether a slowdown is thermal throttling or something else.
inference server (vLLM, Ollama) on the same box.
Spark has no separate GPU VRAM — the GPU and CPU share one 128GB pool. Two consequences:
nvidia-smi and cudaMemGetInfounderreport pressure — or report nothing at all. Both report CUDA-allocator-visible memory, not the pool's actual state — a box can show headroom in nvidia-smi and still OOM, because page-cache and mmap'd pages the allocator doesn't see consume the same pool. On some driver/setups, the memory query returns [N/A], [N/A] outright instead of a number — a script grepping for a numeric value there gets nothing, not a misleading undercount (see spark-training-gotchas gotcha G3).
steady state. Loading safetensors weights mmaps the file, then copies into CUDA tensors — for a window during load, both the mmap'd pages and the CUDA copy count against the pool at once. A model that fits while training can still OOM during load if headroom was sized for the post-load footprint instead of this doubled transient.
Plan and diagnose with free -g, not nvidia-smi:
bashfree -g | awk 'NR==2 {print "free:", $4, "GB"}'
Rule of thumb: take that free figure, subtract a few GB for OS/driver overhead, and budget against the result — not the 128GB spec number. The worksheet in references/uma-accounting.md accepts parameter count, dtype, and method as input, and returns a memory estimate to compare against known anchors.
Before launch, work through these in order:
free -g; subtract OS/driver overheadfor the budget.
activations from references/uma-accounting.md.
QLoRA, 27B LoRA, 9B full FT), not the estimate alone.
with shorter packing or a smaller batch — cheaper than hitting the OOM Ladder mid-run.
A sanity check of the worksheet formula against the ≈40GB anchor:
pythonparams = 70e9 weights_gb = params * 0.5 / 1e9 # NF4, step 1 adapter_gb = 0.5 # step 5, negligible total_gb = weights_gb + adapter_gb # + activations print(f"{total_gb:.0f}GB before activations")
Weights alone land near the ≈40GB anchor — a plan estimating far above that for the same model class is a signal to recheck dtype and method.
When a job OOMs on unified memory, work this ladder in order. Each step is more disruptive than the last — don't skip ahead: reducing batch size is never step 1.
previous run or a large dataset read often accounts for GB of the "missing" headroom. This costs nothing but a rerun and doesn't touch the job's configuration:
bash sync; echo 3 > /proc/sys/vm/drop_caches
Needs root; a between-run reset, not a mid-training step. See spark-training-gotchas (gotcha G3) for the full diagnostic behind this step.
after a flush fails to free enough headroom, cut batch size or packing length — the first step that changes what the run does. Prefer packing length first; it drives activation footprint more directly at long context.
QLoRA. If flushing and shrinking batch/pack still OOM, drop the method a tier — bf16 LoRA is next, not the reverse. QLoRA's bitsandbytes dequantization buffers are transient CUDA-side allocations that can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller. A QLoRA OOM is not proof the model doesn't fit.
Fall back further (smaller model, multi-Spark) only after all three steps and the job still won't fit.
Multi-hour runs push into Spark's sustained power ceiling, well under the rated figure — expected platform behavior, not a symptom to explain away:
training logs, not after a slowdown is noticed — every 30-60 seconds correlates a throughput drop with a thermal event. Keep the CSV output format assets/thermal-sample.sh writes, so timestamps line up against the log:
bash bash assets/thermal-sample.sh 30 thermal.log
cap, not a configuration bug. Don't re-tune batch size or precision to "fix" a plateau that's the box behaving normally under load. If temperature climbs while power stays flat under the rated 240W figure, that's the signature to recognize.
letting a run silently slow down unrecorded. A run whose per-step time doubles two hours in should show that in the log, correlated against the thermal sample at that timestamp. Full throttling diagnostics: spark-training-gotchas (gotcha G4).
Because the 128GB pool is global, eviction happens without either process's logs showing an OOM:
near-capacity workloads — an uncapped trainer and inference server (vLLM, Ollama) compete for the same pool. A small, capped workload doesn't: a <4GB LoRA fine-tune coexists fine alongside vLLM capped at gpu-memory-utilization<=0.5 — check the other process's cap, not just its presence, before stopping it.
under uncapped/near-capacity contention, and vice versa — neither logs an error, so a slow run or lost KV cache is a contention symptom to check for. Stop unrelated uncapped servers before a long or full-pool run.
Check for GPU-resident processes first:
bashps aux | grep -E 'vllm|ollama|trl|axolotl' | grep -v grep
This procedure complements spark-training-gotchas (gotchas G3, G4, G6) — that skill covers launch-time failures; this one, the running job.
Memory math worksheets: references/uma-accounting.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 27,941 | 15,125 | -46% | 1 | 1 | 0% | 6,206 | 5,133 | -17% | 0 | 0 | — |
case-02 | fail→pass | 17,788 | 12,601 | -29% | 1 | 1 | 0% | 3,326 | 4,295 | +29% | 0 | 0 | — |
case-03 | fail→pass | 26,428 | 14,290 | -46% | 1 | 1 | 0% | 4,383 | 4,513 | +3% | 0 | 0 | — |
case-04 | pass→pass | 11,506 | 5,604 | -51% | 1 | 1 | 0% | 2,101 | 3,190 | +52% | 0 | 0 | — |
case-05 | fail→pass | 16,467 | 6,536 | -60% | 1 | 1 | 0% | 2,996 | 3,313 | +11% | 0 | 0 | — |
case-06 | fail→fail | 12,719 | 7,581 | -40% | 1 | 1 | 0% | 2,138 | 3,536 | +65% | 0 | 0 | — |
case-07 | fail→pass | 12,607 | 5,397 | -57% | 1 | 1 | 0% | 2,310 | 3,049 | +32% | 0 | 0 | — |
case-08 | pass→pass | 14,508 | 9,224 | -36% | 1 | 1 | 0% | 2,854 | 3,975 | +39% | 0 | 0 | — |
case-09 | pass→pass | 15,560 | 5,645 | -64% | 1 | 1 | 0% | 2,708 | 3,097 | +14% | 0 | 0 | — |
case-10 | fail→pass | 11,785 | 5,179 | -56% | 1 | 1 | 0% | 2,148 | 3,176 | +48% | 0 | 0 | — |
case-11 | fail→pass | 16,568 | 7,674 | -54% | 1 | 1 | 0% | 3,380 | 3,480 | +3% | 0 | 0 | — |
case-12 | fail→pass | 22,145 | 5,613 | -75% | 1 | 1 | 0% | 1,616 | 3,104 | +92% | 0 | 0 | — |
case-13 | fail→pass | 13,732 | 8,530 | -38% | 1 | 1 | 0% | 2,316 | 3,505 | +51% | 0 | 0 | — |
case-14 | fail→pass | 9,757 | 3,304 | -66% | 1 | 1 | 0% | 1,915 | 2,676 | +40% | 0 | 0 | — |
case-15 | fail→pass | 16,823 | 7,261 | -57% | 1 | 1 | 0% | 2,861 | 3,591 | +26% | 0 | 0 | — |
case-16 | pass→pass | 19,358 | 9,349 | -52% | 1 | 1 | 0% | 3,665 | 3,668 | +0% | 0 | 0 | — |
case-17 | pass→pass | 11,448 | 4,371 | -62% | 1 | 1 | 0% | 2,072 | 2,899 | +40% | 0 | 0 | — |
case-18 | pass→pass | 13,556 | 5,147 | -62% | 1 | 1 | 0% | 2,727 | 3,049 | +12% | 0 | 0 | — |
case-19 | fail→pass | 16,788 | 7,178 | -57% | 1 | 1 | 0% | 2,998 | 3,322 | +11% | 0 | 0 | — |
case-20 | pass→pass | 11,437 | 7,627 | -33% | 1 | 1 | 0% | 2,507 | 3,639 | +45% | 0 | 0 | — |
case-21 | pass→pass | 13,381 | 11,104 | -17% | 1 | 1 | 0% | 2,608 | 4,135 | +59% | 0 | 0 | — |
case-22 | pass→pass | 12,358 | 8,155 | -34% | 1 | 1 | 0% | 2,200 | 3,714 | +69% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.