Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results. Use when running TAO training/eval/inference jobs on an on-prem or DGX SLURM cluster. Trigger phrases include "run on SLURM", "submit sbatch", "DGX SLURM cluster", "Pyxis/Enroot container", "Lustre dataset".
.claude/skills/nvidia-tao-run-on-slurm/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 96% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 83% | 0% |
Remote GPU compute platform for clusters managed by SLURM. Jobs are submitted from the TAO service or SDK host to a login node over SSH, staged on a shared filesystem, submitted with sbatch, and executed with srun container support.
Use SLURM when the user has access to a managed GPU cluster, shared Lustre storage, and scheduler-owned GPU allocation. Do not use SLURM for local files that exist only on the agent machine; data and outputs must be reachable from the cluster.
Confirm SLURM_USER and SLURM_HOSTNAME are exported and passwordless SSH to a login host works (ssh -o BatchMode=yes). Optionally install the TAO SDK wrapper for Job handles + S3 wrapping (nvidia-tao-sdk[slurm], on public PyPI). For private nvcr.io images, install ~/.config/enroot/.credentials on the cluster once per (cluster, user): Pyxis/Enroot does not read NGC_KEY from the job env, and without persistent credentials, auth-gated pulls fail with "Could not process JSON input" at job startup. Install it via the printf | ssh heredoc so the NGC_KEY value never lands in shell history, intermediate files, or chat output; never cat/echo the value.
If a preflight check fails, the agent prompts the user to authorize the install/fix via Bash. Pip-installable Python requirements are the exception: install them automatically, then rerun preflight.
See references/slurm-ssh-credentials.md for the full preflight script, the enroot-credentials heredoc, prerequisite key setup (keypair, ssh-copy-id, known_hosts, container key mounts, 2FA handling), and the SSH failure remediation prompt.
Use shared-filesystem URIs, not local or file:// paths; tao-core rejects local/file paths for remote backends.
lustre:///absolute/path for user-provided datasets on Lustre.slurm:// paths may appear in microservices metadata and are converted toLustre paths before the container starts.
Accept either dataset roots (model skills map them to required files) or direct spec-key paths. After SSH succeeds and before generating scripts, test -e each required dataset path from the login host; if it fails, stop and ask for corrected paths or staged data rather than producing scripts that fail in the first training job. See references/slurm-ssh-credentials.md for root vs. direct-spec modes, backend details, and the results-dir default.
tao-core runs TAO containers through Pyxis/Enroot:
<job_dir>/specs, <job_dir>/env, and <job_dir>/meta.
srun -n1 -p <conversion_partition> enroot import.
<job_dir>/sbatch/job_<job_id>.sbatch.sbatch --export=ALL <script>.srun --container-image=<image> --container-mounts=/lustre.Accepted image formats: /path/to/image.sqsh, registry#image:tag, docker://registry#image:tag, and ordinary registry/image:tag (converted to Pyxis form when needed). SQSH conversion is cached by image name; for :latest images the cached SQSH is reused unless force_reconvert_latest is enabled.
squeue/sacct;TAO terminal status comes from status.json in the shared results folder.
any non-terminal job (PENDING, RUNNING, or otherwise). Do not stop after a fixed elapsed time such as 30 minutes; long queue waits are normal on shared GPU partitions.
monitoring is enabled. A final response is a detach action; use it only if the user asked to detach/stop or the job reached terminal state.
<job_dir>/slurm-logs/<slurm_job_name>-<slurm_job_id>/main.out and .err.
backend_details.slurm_metadata.slurm_job_id and runningscancel <slurm_job_id> over SSH. Treat missing or already terminated jobs as successful cancellation.
Status mapping:
PENDING -> PendingRUNNING or COMPLETING -> RunningCOMPLETED -> check status.jsonFAILED, BOOT_FAIL, DEADLINE, OUT_OF_MEMORY, NODE_FAIL -> retry iflogs match retriable infrastructure patterns, otherwise Error
CANCELLED, PREEMPTED, REVOKED -> CanceledTIMEOUT -> ErrorSUSPENDED, STOPPED -> PausedAsk for these in the SLURM intake; see references/slurm-ssh-credentials.md for the full credential list, microservices schema keys, and defaults.
default polar,polar3,polar4,grizzly, treated as 4-hour queues.
non-interactive public-key auth. Ask for this first in remediation; prefer it over the SSH_AUTH_SOCK agent-socket fallback.
/lustre/fsw/portfolios/edgeai/users/<your-dir> (your per-user Lustre dir).
#SBATCH --account.Do not ask for SLURM_ACCOUNT or SLURM_BASE_RESULTS_DIR in the initial intake unless the user says their site requires an account, wants a custom results root, or the workflow cannot proceed without overriding defaults.
Defaults from tao-core:
num_nodes: 1num_gpus: 4max_num_gpus_per_node: 8cpus_per_task: 16time_hours: 4timeout_hours: 3.8max_time_hours: 4container_mounts: /lustreuse_requeue: trueuse_sqsh: trueWhen generating launchers or wrapper scripts for SLURM, set the wall-time defaults explicitly from the packaged platform resource defaults:
bashexport SLURM_TIME_HOURS="${SLURM_TIME_HOURS:-4}" export SLURM_TIMEOUT_HOURS="${SLURM_TIMEOUT_HOURS:-3.8}"
Do not default to 12 hours on SLURM. If the user supplies a longer SLURM_TIME_HOURS, verify that the selected partition supports it before submitting. For the packaged default partition list polar,polar3,polar4,grizzly, reject requests above 4 hours and ask for a different partition only if the user actually wants a longer wall time.
When num_gpus is greater than or equal to max_num_gpus_per_node, the handler treats the request as exclusive per node and computes additional nodes from total GPU count when necessary.
For multi-node jobs (num_nodes > 1), the SDK builds the sbatch directives and exports the PyTorch-distributed rendezvous env vars automatically: WORLD_SIZE, NUM_GPU_PER_NODE, NODE_RANK, MASTER_ADDR, and MASTER_PORT (29500). TAO entrypoints read WORLD_SIZE + NUM_GPU_PER_NODE and build torchrun internally. Cosmos-RL has special multi-node role handling for controller, policy, and rollout workers.
Use Lustre, not S3, for SLURM job inputs. The GPU allocation starts the moment the job is dispatched, so a long s3:// download at the top of the script burns the allocation, can get the job killed for GPU-idle, and is billed either way. Stage training data on the shared filesystem first and reference it as lustre:///.... S3/HF/NGC pre-fetch is fine for small auxiliary inputs (checkpoints, configs), not training datasets. K8s/Brev do not share this scheduler-idle constraint.
Auto-retry of infrastructure failures (NODE_FAIL, BOOT_FAIL, NCCL transport timeouts, CUDA driver init failures, GPU/IB link-down, OOM-killer node reaping, Xid errors) is automatic in the SDK, with a stable user-facing Job.id across retries. Plain training failures surface immediately so a broken spec does not consume the retry budget. #SBATCH --requeue is enabled by default via SLURM_USE_REQUEUE=true.
See references/slurm-container-execution.md for the full multi-node env-var/sbatch directive detail and table, cluster requirements, the optional TAO SDK path (SlurmSDK, build_entrypoint, ActionWorkflow) with code, the Lustre-not-S3 rule in full, and the failure-mode checklist; references/slurm-execution-sdk.md covers the MAX_JOB_RETRIES retry budget. When the SDK is in scope, read tao-skill-bank:tao-run-platform for the SlurmSDK kwarg reference.
references/slurm-ssh-credentials.md — preflight script, SSH/key setup,enroot credentials, full credential list, backend details, storage rules, SSH remediation prompt.
references/slurm-container-execution.md — container execution steps,monitoring, status mapping, cancellation, multi-node detail, SDK use, Lustre-not-S3, auto-retry, failure modes.
references/slurm-preflight-storage.md — extended preflight/storage notes.references/slurm-execution-sdk.md — extended execution/SDK notes.references/detailed-guide.md — navigation map for the split references.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | pass→pass | 7,562 | 7,242 | -4% | 1 | 1 | 0% | 1,513 | 3,977 | +163% | 0 | 0 | — |
case-01 | fail→pass | 15,253 | 8,825 | -42% | 1 | 1 | 0% | 3,454 | 4,612 | +34% | 0 | 0 | — |
case-02 | fail→pass | 14,591 | 9,439 | -35% | 1 | 1 | 0% | 2,825 | 4,570 | +62% | 0 | 0 | — |
case-04 | pass→pass | 10,855 | 7,669 | -29% | 1 | 1 | 0% | 2,102 | 4,215 | +101% | 0 | 0 | — |
case-05 | fail→fail | 10,799 | 11,601 | +7% | 1 | 1 | 0% | 2,293 | 5,253 | +129% | 0 | 0 | — |
case-06 | fail→fail | 10,835 | 6,987 | -36% | 1 | 1 | 0% | 2,184 | 4,055 | +86% | 0 | 0 | — |
case-07 | fail→pass | 12,085 | 4,419 | -63% | 1 | 1 | 0% | 1,727 | 3,380 | +96% | 0 | 0 | — |
case-08 | fail→pass | 12,596 | 5,695 | -55% | 1 | 1 | 0% | 2,518 | 3,697 | +47% | 0 | 0 | — |
case-09 | fail→pass | 10,417 | 3,812 | -63% | 1 | 1 | 0% | 1,866 | 3,408 | +83% | 0 | 0 | — |
case-10 | fail→pass | 10,520 | 3,941 | -63% | 1 | 1 | 0% | 1,928 | 3,379 | +75% | 0 | 0 | — |
case-11 | fail→pass | 9,325 | 2,984 | -68% | 1 | 1 | 0% | 1,699 | 3,172 | +87% | 0 | 0 | — |
case-12 | fail→pass | 13,200 | 2,677 | -80% | 1 | 1 | 0% | 1,847 | 3,120 | +69% | 0 | 0 | — |
case-13 | fail→pass | 7,783 | 2,867 | -63% | 1 | 1 | 0% | 1,383 | 3,130 | +126% | 0 | 0 | — |
case-19 | fail→pass | 10,138 | 2,070 | -80% | 1 | 1 | 0% | 1,838 | 2,928 | +59% | 0 | 0 | — |
case-14 | fail→pass | 5,039 | 2,956 | -41% | 1 | 1 | 0% | 995 | 3,194 | +221% | 0 | 0 | — |
case-15 | fail→pass | 9,616 | 4,218 | -56% | 1 | 1 | 0% | 1,791 | 3,375 | +88% | 0 | 0 | — |
case-16 | pass→pass | 10,724 | 3,528 | -67% | 1 | 1 | 0% | 1,932 | 3,221 | +67% | 0 | 0 | — |
case-17 | fail→pass | 10,359 | 4,849 | -53% | 1 | 1 | 0% | 1,940 | 3,505 | +81% | 0 | 0 | — |
case-18 | fail→pass | 5,642 | 4,763 | -16% | 1 | 1 | 0% | 1,027 | 2,919 | +184% | 0 | 0 | — |
case-20 | fail→pass | 9,500 | 2,237 | -76% | 1 | 1 | 0% | 1,656 | 3,007 | +82% | 0 | 0 | — |
case-21 | fail→pass | 5,876 | 4,134 | -30% | 1 | 1 | 0% | 1,160 | 3,024 | +161% | 0 | 0 | — |
case-22 | pass→fail | 9,696 | 2,724 | -72% | 1 | 1 | 0% | 1,725 | 3,022 | +75% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +68 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.