Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when adding, modifying, optimizing, or debugging CuTile autotuning code. Trigger signals: `exhaustive_search` / `replace_hints` / `hints_fn` / `cuda.tile.tune` in code, `autotune` in filenames, or correctness/performance issues in autotuned CuTile kernels. Covers: tune-once/cache/launch pattern, per-architecture configs (sm80–sm120), parameter space design (tile sizes, occupancy, num_ctas), and 7 common pitfalls with solutions.
.claude/skills/nvidia-tilegym-cutile-autotuning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 134% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 202% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 91% | 0% |
Add autotuning to CuTile kernels using the exhaustive_search API with tune-once/cache/direct-launch pattern.
Follow the decision tree to classify the kernel, design a search space, implement the tune-once/cache/launch pattern, and validate performance.
references/kernel-type-templates.md; prune to ≤ 30 configs in the final code via arch filters (directed exploration probes may temporarily exceed this — see Design Philosophy)exhaustive_search + cache + ct.launch following the Step-by-Step Workflow; handle in-place writes with split-buffer if neededDISABLE_AUTOTUNE=1references/search-strategies.md| What are you trying to do? | Go to | |---|---| | Add autotune to a new kernel (most common) | Quick Reference below → Workflow: Adding Autotune → references/kernel-type-templates.md (pick by kernel type: T1=elementwise, T2=in-place, T3=matmul, T4=persistent, T5=FMHA, T6=FP8, T7=grouped GEMM, T8=varlen attention, T9=dual-GEMM fusion) | | Debug: data corruption / wrong results after first run | Pitfall #1 (In-Place Kernel) | | Debug: autotune taking 5+ minutes | Pitfall #2 (Compilation Timeout) | | Debug: search space generator returning zero configs | Pitfall #5 first; also check arch filters, size guards, and num_ctas constraints | | Optimize an existing autotune config | Workflow: Optimizing an Existing Config |
Most CuTile kernels (elementwise, reduction, LayerNorm) need only occupancy tuning. Copy this pattern:
pythonfrom types import SimpleNamespace from cuda.tile.tune import exhaustive_search import cuda.tile as ct import torch def _my_autotune_configs(): for occ in [1, 2, 4, 8]: yield SimpleNamespace(occupancy=occ) # Module-level cache: tune once, launch fast forever after _autotune_cache = {} def my_op(x, output): stream = torch.cuda.current_stream() NUM_SM = torch.cuda.get_device_properties(x.device).multi_processor_count # Cache key: anything that affects optimal config (use str() for device) cache_key = (x.shape, x.dtype, str(x.device)) if cache_key not in _autotune_cache: configs = list(_my_autotune_configs()) result = exhaustive_search( configs, stream, grid_fn=lambda cfg: (min(NUM_SM * cfg.occupancy, M), 1, 1), kernel=my_kernel, args_fn=lambda cfg: (x, output, ...), hints_fn=lambda cfg: {"occupancy": cfg.occupancy}, ) best_cfg = result.best.config tuned_kernel = my_kernel.replace_hints(occupancy=best_cfg.occupancy) _autotune_cache[cache_key] = (best_cfg, tuned_kernel) # cache BOTH cfg, tuned_kernel = _autotune_cache[cache_key] grid = (min(NUM_SM * cfg.occupancy, M), 1, 1) ct.launch(stream, grid, tuned_kernel, (x, output, ...))
Key rules:
exhaustive_search runs only on first call per shape; subsequent calls use cached config + ct.launch with zero overheadexhaustive_search requires a Sequence (list/tuple) — convert generators with list()When to use this pattern: Kernel has fixed block size (not tile-size tunable). Includes: elementwise (SwiGLU, GeGLU), reduction (RMSNorm, LayerNorm), RoPE, and persistent kernels with heuristic block sizes (grouped GEMM).
For complex kernels (matmul with tile sizes, FMHA, FP8 with num_ctas), read the full guide below + kernel-type-templates.md.
> ⚠️ Three pitfalls catch almost everyone — check before submitting: > - replace_hints on hot path? → Cache BOTH config AND kernel object from exhaustive_search. Calling replace_hints() every invocation recompiles (100–500× slower) → Pitfall #7 > - In-place kernel (writes back to input tensor)? → MUST use split-buffer pattern during search → Pitfall #1 > - Search space empty? → Check arch filters and num_ctas constraints → Pitfall #5
> Minimum coverage: On sm100+, FMHA/matmul/varlen search spaces must include both num_ctas=1 and num_ctas=2. For core dimensions (tile sizes, occupancy), keep at least 2 distinct values even if unsure which is better — let exhaustive_search decide.
> When to stop tuning: A mean speedup in 0.98, 1.02] means your current search space isn't helping — but doesn't mean no config will help. Before stopping, check whether you've covered the key dimensions for this kernel type (consult references/kernel-type-templates.md). If the search space already covers the template's recommended dimensions and the best result is still noise-floor, then stop — further micro-adjustments won't help. If key dimensions are missing (e.g., never tried num_ctas=2 for a dual-GEMM kernel), expand the search space rather than giving up. > > Once correctness tests pass and the autotuned kernel shows speedup over the fixed-config baseline, stop — do not re-run to "confirm". GPU kernel timing fluctuates ±5–10 % between invocations due to clock scaling and OS scheduling; a subsequent timing dip does not mean your code is wrong. > > To improve speedup, only modify the autotune search space (configs, tile sizes, occupancy, num_ctas). Do not modify other code (Python wrapper, stream management, etc.) to chase speedup — kernel performance is determined by the config selection, not by host-side code.
references/ docs. For in-place kernels, also read Pitfall #1.references/ docs.5-step summary: Classify kernel → Design search space (parameter-space-design.md) → Implement using template (kernel-type-templates.md) → Validate with A/B test → Check Pitfall Checklist.
Reading references: Read only the reference relevant to your kernel type — e.g., for FMHA, read the Template 5 section in references/kernel-type-templates.md; for hardware constraints, read only the target architecture's section. Avoid reading all references end-to-end when a targeted lookup suffices.
Build a small, precise search space bottom-up — not a large space trimmed down. CuTile compilation is much heavier than Triton (~0.5-1s per config), so the final code should contain ≤ 30 configs. The approach is: classify the kernel type first, then construct only the relevant configs for that type and architecture.
Directed exploration during development: If the initial template configs yield speedup < 1.0, you may run a temporary larger probe (30–100 configs) via bash + python3 -c to identify which dimensions matter — but this probe must be directional, not a blind cartesian product. Use the kernel type classification to decide which dimensions to vary (e.g. for dual-GEMM, probe num_ctas × occupancy while fixing tile sizes; for FMHA, probe TILE_M × num_ctas while fixing TILE_N). Once the probe identifies the winning region, lock the final code's search space to ≤ 8 top candidates. Do NOT write the large probe into the source file — it is a one-shot diagnostic tool.
All kernels should have autotuning added. The question is not whether to autotune, but what dimensions to search:
What type of kernel is this?
├── Compute-bound (matmul, GEMM, FMHA) → Does it have multiple tunable dimensions (tile sizes)?
│ ├── YES → Is it a fused multi-GEMM kernel (dual-GEMM, e.g. Linear+GLUAct)?
│ │ ├── YES → Template 9: low occupancy (1–2), conservative tiles (2× SHMEM/register pressure)
│ │ └── NO → Full search: TILE_M × TILE_N × (TILE_K) × occupancy × num_ctas
│ │ (see matmul/FMHA templates in kernel-type-templates.md)
│ └── NO → Occupancy-only search: [1, 2, 4, 8]
│ (see Quick Reference above)
├── Balanced (LayerNorm, reduction + compute) →
│ Occupancy-only search: [1, 2, 4, 8]
│ Expected benefit: 2-15%
└── Memory-bound (CE Loss, pure elementwise) →
Occupancy-only search: [1, 2, 4, 8]
Expected benefit: 0-15% (varies by kernel; zero-cost after tuning)Why memory-bound kernels only search occupancy (not num_ctas or tile sizes):
num_ctas has zero benefit: num_ctas > 1 enables TMA multicast, where multiple CTAs share tile data in shared memory (e.g., matmul A/B tiles reused across CTAs). Memory-bound kernels use per-element ct.gather/ct.scatter with no tile reuse — multi-CTA cooperation adds overhead with no data sharing benefit.> Evidence — CE Loss experiment: A 12-config search (occupancy × num_ctas) on Cross-Entropy Loss yielded only 2.5% gain (0.79x → 0.81x vs Triton). The num_ctas dimension contributed nothing; the result was reverted because compilation cost outweighed the marginal benefit. Occupancy-only (4 configs) achieves the same result at 3x less compilation time.
Note on memory-bound kernels: Adding occupancy-only autotune is always worthwhile because:
Occupancy controls how many CTAs run concurrently per SM. Use this as a starting point when designing the occupancy search space:
| Occupancy Range | Best For | Example Kernels | |-----------------|----------|-----------------| | 1–4 | Compute-bound (heavy math) | Complex transforms, matmul | | 4–8 | Balanced (GEMM, TMA) | Matrix multiply, FMHA | | 8–16 | Memory-bound (reductions) | Softmax, LayerNorm | | 16–32 | Very light (copies, casts) | Type conversions, elementwise |
Use these ranges to seed your initial search space. For occupancy-only kernels, [1, 2, 4, 8] covers most cases — see Quick Reference above.
See references/api-reference.md for the full exhaustive_search API surface — current signature, TuningResult, the tune-once/cache/launch pattern, replace_hints, kernel hints, search_space design, and grid_fn patterns.
See references/workflow.md for the end-to-end workflow — adding autotune to a new kernel, handling existing multi-architecture configs, integration with torch.autograd.Function, cross-backend config transfer (Triton → CuTile), and optimizing an existing config.
See references/pitfalls.md for the full list of common pitfalls — in-place data corruption, compilation timeout, cold-cache performance skew, NCU profiling interference, search_space generator exhaustion, FP8 precision loss, and replace_hints recompilation on hot paths.
This skill covers only autotune configuration: search space design, exhaustive_search invocation, caching, and ct.launch with tuned hints. It does not modify kernel code.
In scope (autotune config):
exhaustive_search() calls and result handlingkernel.replace_hints() for applying tuned hintsct.launch() with tuned kernelDISABLE_AUTOTUNE fallback pathOut of scope (kernel code modifications — do NOT make these changes):
After adding autotuning, the following kernel-level optimizations may yield additional gains. These are outside the scope of this skill — mention them to the user as potential next steps, but do not implement them as part of autotuning:
flush_to_zero=True + rounding_mode=APPROX can provide 34-72% improvement for FMHA-class kernels (set via environment variables TILEIR_ENABLE_FTZ=1 TILEIR_ENABLE_APPROX=1 or in kernel code). Causal chain: larger tiles initially decrease performance by 18-43% due to subnormal handling overhead; enabling FTZ+APPROX rescues this and flips the result to +34-72%. Math flags are therefore a prerequisite for large-tile configs to be effective on FMHA-class kernels.slice_hint, buffer_depth, copy_config — requires modifying kernel IR codect.load) instead of ct.gather; removing unnecessary bounds checks (check_bounds=False when safe)padding_value parameter instead of manual ct.where masking; removing safe_offsKey differences: Triton uses @triton.autotune decorator with Config(...) objects; CuTile uses exhaustive_search() with SimpleNamespace configs + separate cache + ct.launch. CuTile has no num_warps/num_stages (compiler decides) — only tile sizes + occupancy + num_ctas. CuTile compilation is heavier (keep ≤30 configs in final code). CuTile cache is user-managed in-memory (no automatic persistence). CuTile separates args_fn (kernel args) from hints_fn (compiler hints).
| Category | Document | Content | |----------|----------|---------| | API Reference | api-reference.md | exhaustive_search signature, TuningResult, tune-once/cache/launch pattern, replace_hints, kernel hints, search_space design, grid_fn patterns | | Workflow | workflow.md | End-to-end workflow: adding autotune to a new kernel, multi-architecture configs, torch.autograd.Function integration, Triton→CuTile transfer, optimizing existing configs | | Pitfalls | pitfalls.md | Common pitfalls: in-place corruption, compilation timeout, cold-cache skew, NCU interference, search_space exhaustion, FP8 precision, replace_hints recompilation | | Parameter Design | parameter-space-design.md | Per-kernel-type parameter spaces, cross-arch patterns, grid_fn patterns, pruning rules | | Search Strategies | search-strategies.md | Exhaustive search, A/B test methodology, DISABLE_AUTOTUNE pattern | | Templates | kernel-type-templates.md | Copy-paste autotune templates for 8 kernel types | | Hardware | hardware-constraints.md | Per-architecture constraints, tile size ranges, num_ctas rules, TMA requirements |
Key files: ops/cutile/matmul.py (matmul autotune), ops/cutile/attention.py (FMHA autotune), suites/unsloth/cutile/ct_ops.py (shared autotune_configs() occupancy=1,2,4,8]), suites/unsloth/cutile/swiglu.py (elementwise example), suites/unsloth/cutile/rope_embedding.py (split-buffer pattern), suites/unsloth/cutile/grouped_gemm.py (persistent GEMM, occupancy-only).
Each example shows the before → after pattern: fixed_launch.py (hardcoded ct.launch) and autotuned_launch.py (refactored to tune-once/cache/launch).
| Directory | Kernel | Autotune Pattern | Complexity | Key Teaching Point | |-----------|--------|-----------------|------------|-------------------| | assets/examples/01_rmsnorm_occupancy_only/ | RMSNorm (reduction) | Occupancy-only [1,2,4,8] | Low | Most common pattern — no tile tuning, just find best occupancy. Grid = NUM_SM * cfg.occupancy. Not in-place. | | assets/examples/02_matmul_full_search/ | GEMM C=A@B | Full: TILE_M/N/K + occupancy + num_ctas (sm90+) | High | Compute-bound kernel with multiple tunable dimensions. args_fn passes tile sizes as ct.Constant[int]. grid_fn depends on cfg. ≤30 configs. | | assets/examples/03_rope_inplace_splitbuffer/ | RoPE embedding (in-place) | Occupancy-only, with split-buffer | Medium | In-place kernel MUST use split-buffer during search to avoid corruption. Search writes to scratch; final ct.launch uses real in-place args. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→pass | 34,146 | 12,992 | -62% | 1 | 1 | 0% | 3,274 | 7,684 | +135% | 0 | 0 | — |
case-17 | pass→pass | 14,871 | 8,883 | -40% | 1 | 1 | 0% | 2,527 | 6,546 | +159% | 0 | 0 | — |
case-18 | pass→pass | 17,839 | 7,663 | -57% | 1 | 1 | 0% | 2,851 | 6,184 | +117% | 0 | 0 | — |
case-01 | fail→pass | 14,755 | 12,055 | -18% | 1 | 1 | 0% | 3,246 | 7,604 | +134% | 0 | 0 | — |
case-02 | fail→pass | 17,263 | 22,780 | +32% | 1 | 1 | 0% | 3,491 | 10,555 | +202% | 0 | 0 | — |
case-03 | fail→pass | 25,292 | 14,903 | -41% | 1 | 1 | 0% | 5,675 | 8,314 | +47% | 0 | 0 | — |
case-04 | fail→pass | 20,417 | 16,085 | -21% | 1 | 1 | 0% | 4,496 | 8,578 | +91% | 0 | 0 | — |
case-06 | pass→pass | 10,614 | 10,779 | +2% | 1 | 1 | 0% | 2,239 | 7,441 | +232% | 0 | 0 | — |
case-07 | pass→pass | 9,584 | 5,875 | -39% | 1 | 1 | 0% | 1,770 | 6,042 | +241% | 0 | 0 | — |
case-08 | fail→pass | 37,111 | 10,973 | -70% | 1 | 1 | 0% | 1,890 | 7,159 | +279% | 0 | 0 | — |
case-09 | fail→pass | 16,253 | 11,655 | -28% | 1 | 1 | 0% | 3,242 | 7,210 | +122% | 0 | 0 | — |
case-10 | pass→pass | 9,276 | 5,968 | -36% | 1 | 1 | 0% | 1,638 | 5,965 | +264% | 0 | 0 | — |
case-11 | pass→pass | 9,823 | 7,327 | -25% | 1 | 1 | 0% | 1,738 | 6,418 | +269% | 0 | 0 | — |
case-12 | pass→pass | 15,849 | 10,630 | -33% | 1 | 1 | 0% | 3,088 | 7,386 | +139% | 0 | 0 | — |
case-13 | fail→pass | 41,839 | 13,641 | -67% | 1 | 1 | 0% | 4,231 | 7,528 | +78% | 0 | 0 | — |
case-14 | fail→pass | 11,914 | 13,385 | +12% | 1 | 1 | 0% | 2,486 | 7,893 | +217% | 0 | 0 | — |
case-15 | pass→pass | 12,456 | 6,393 | -49% | 1 | 1 | 0% | 2,391 | 6,168 | +158% | 0 | 0 | — |
case-16 | pass→pass | 14,819 | 10,289 | -31% | 1 | 1 | 0% | 3,177 | 7,163 | +125% | 0 | 0 | — |
case-19 | pass→pass | 12,675 | 6,931 | -45% | 1 | 1 | 0% | 2,309 | 6,014 | +160% | 0 | 0 | — |
case-20 | fail→pass | 7,622 | 7,890 | +4% | 1 | 1 | 0% | 1,337 | 6,624 | +395% | 0 | 0 | — |
case-21 | pass→pass | 8,163 | 11,222 | +37% | 1 | 1 | 0% | 1,378 | 7,089 | +414% | 0 | 0 | — |
case-22 | fail→pass | 13,374 | 9,855 | -26% | 1 | 1 | 0% | 2,476 | 7,062 | +185% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.