Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Add a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, __init__.py exports, test creation, and benchmark in tests/benchmark. Use when adding, creating, or implementing a new cuTile operator/kernel in TileGym, or when asking how to register a new cuTile op.
.claude/skills/nvidia-tilegym-adding-cutile-kernel/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 75% | 12 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 90% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 134% | 0% |
End-to-end workflow for adding a new operator (e.g., my_op) with cuTile backend.
MUST follow these rules strictly:
completed after finishing, in_progress when startingcompleted with a note, do NOT silently skipMUST copy this checklist to TodoWrite at the start:
- [ ] Step 1: Register dispatch interface in ops.py
- [ ] Step 2: Implement cuTile backend
- [ ] Step 3: Register in __init__.py (cutile)
- [ ] Step 4: Add tests
- [ ] Step 5: Add benchmark to tests/benchmark
- [ ] Step 6: Verify (run pytest + lint)File: src/tilegym/ops/ops.py
Add a @dispatch function — this is the single entry point for all backends.
python@dispatch( "my_op", ) def my_op( input: torch.Tensor, out: Optional[torch.Tensor] = None, **kwargs: Any, ): """ Description of my_op. Args: input: Input tensor out: Optional preallocated output tensor **kwargs: Additional arguments for backend-specific configurations Returns: torch.Tensor """ raise NotImplementedError(f"my_op is not implemented for {get_current_backend()}")
Key rules:
NotImplementedError**kwargs for backend-specific parametersReference: See existing ops in src/tilegym/ops/ops.py (e.g., silu_and_mul, softmax)
File: src/tilegym/ops/cutile/my_op.py
The file structure follows this template:
pythonimport torch import cuda.tile as ct from tilegym.backend import register_impl @ct.kernel def my_op_kernel_ct(x, output, n_elements: ct.Constant[int], BLOCK_SIZE: ct.Constant[int]): bid = ct.bid(0) indices = bid * BLOCK_SIZE + ct.arange(0, BLOCK_SIZE) x_val = ct.gather(x, indices) # ... compute ... ct.scatter(output, indices, result) @register_impl("my_op", backend="cutile") def my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs) -> torch.Tensor: n = input.numel() if out is None: out = torch.empty_like(input) grid = ((n + 1023) // 1024,) ct.launch(stream, grid, kernel, (some args, ...)) return out
Reference: src/tilegym/ops/cutile/silu_and_mul.py
__init__.py (CRITICAL)Missing this step means the cuTile backend implementation never gets loaded.
File: src/tilegym/ops/cutile/__init__.py
Add inside if is_backend_available("cutile"): block (alphabetically):
pythonfrom . import my_op
And in the function import section:
pythonfrom .my_op import my_op
And add "my_op" to __all__.
File: tests/ops/test_my_op.py
CRITICAL: Always import from tilegym.ops, NEVER from tilegym.ops.cutile.my_op.
pythonimport pytest import torch from tilegym.backend import is_backend_available, set_backend from .. import common _backends = ["cutile"] class Test_MY_OP(common.PyTestCase): @staticmethod def reference(input): """Reference implementation using PyTorch.""" return torch.some_reference(input) @pytest.mark.parametrize("shape, dtype", [ ((1024,), torch.float16), ((1024, 512), torch.float32), ((64, 64, 64), torch.bfloat16), ]) @pytest.mark.parametrize("backend", _backends) def test_op(self, shape, dtype, backend, arch): if backend == "cutile" and not is_backend_available("cutile"): pytest.skip("Cutile backend not available") try: set_backend(backend) except Exception as e: pytest.skip(f"Backend is not supported: {e}") self.setUp() from tilegym.ops import my_op A = torch.randn(*shape, dtype=dtype, device="cuda") self.assertCorrectness( my_op, self.reference, {"input": A}, atol=1e-3, rtol=1e-3, )
Key patterns:
_backends = ["cutile"]test_op: use set_backend(backend) with try-except, call self.setUp()Reference: tests/ops/test_silu_and_mul.py
Below is the common errors.
1. Missing _backends list (inside class)
2. test_op / test_op_xxx — missing @pytest.mark.parametrize("backend", _backends), backend parameter, and tilegym.is_backend_available / tilegym.set_backend patternFile: tests/benchmark/bench_my_op.py
Key rules from benchmark_rules.md:
tilegym.ops.my_op(a, b, ..., backend=backend) — do not use set_backend.ALL_BACKENDS (include at least cutile and torch), filter with get_supported_backends().reference_my_op(...) and register it: register_impl("my_op", "torch")(reference_my_op).create_benchmark_config() to build triton.testing.Benchmark configs (e.g. by shape/dtype).@triton.testing.perf_report([...]) on bench_my_op(...); inside the bench function: correctness check with torch.testing.assert_close(fn(), ref(), ...), then ms = triton.testing.do_bench(fn) (or do_bench_cudagraph), compute GB/s or TFLOPS, and return the metric.if __name__ == "__main__": bench_my_op.run(print_data=True).Template structure:
pythonimport torch import triton import triton.testing import tilegym from tilegym.backend import is_backend_available, register_impl ALL_BACKENDS = [ ("cutile", "cuTile", ("orange", "-")) if is_backend_available("cutile") else None, ("torch", "PyTorch", ("green", "-")), ] def get_supported_backends(): return [p for p in ALL_BACKENDS if p is not None] def reference_my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs): """Reference implementation using PyTorch.""" ... register_impl("my_op", "torch")(reference_my_op) def create_benchmark_config(datatype, ...): available_backends = get_supported_backends() if not available_backends: return None backends, names, styles = zip(*available_backends) return triton.testing.Benchmark( x_names=["M"], # or other dimension names x_vals=[...], line_arg="backend", line_vals=list(backends), line_names=list(names), styles=list(styles), ylabel="GB/s", # or TFLOPS plot_name="my-op-...", args={"datatype": datatype, ...}, ) @triton.testing.perf_report([ create_benchmark_config(datatype, ...) for datatype in [torch.float16, torch.float32] for ... in [...] ]) def bench_my_op(M, backend, datatype, ..., device="cuda"): x = torch.randn(..., dtype=datatype, device=device) fn = lambda: tilegym.ops.my_op(x, backend=backend) ref = lambda: reference_my_op(x) torch.testing.assert_close(fn(), ref(), rtol=1e-2, atol=1e-2) ms = triton.testing.do_bench(fn) # or do_bench_cudagraph(fn) # Compute metric (e.g. GB/s or TFLOPS) from ms and problem size return metric if __name__ == "__main__": bench_my_op.run(print_data=True)
Benchmark Plot Names: Must include -TFLOPS or -GBps suffix
plot_name=f"persistent-layer-norm-M{num_rows}-{dtype_name}-GBps"bash# Run tests pytest tests/ops/test_my_op.py -v # Run benchmark (optional) python tests/benchmark/bench_my_op.py # Lint pre-commit run -a
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,984 | 8,466 | +113% | 1 | 1 | 0% | 309 | 2,898 | +838% | 0 | 0 | — |
case-02 | fail→fail | 26,007 | 5,105 | -80% | 1 | 1 | 0% | 4,799 | 2,589 | -46% | 0 | 0 | — |
case-08 | fail→pass | 8,588 | 4,492 | -48% | 1 | 1 | 0% | 1,506 | 3,250 | +116% | 0 | 0 | — |
case-03 | fail→fail | 4,557 | 9,286 | +104% | 1 | 1 | 0% | 211 | 3,027 | +1335% | 0 | 0 | — |
case-04 | fail→pass | 13,037 | 3,533 | -73% | 1 | 1 | 0% | 1,624 | 3,083 | +90% | 0 | 0 | — |
case-05 | fail→pass | 11,291 | 3,377 | -70% | 1 | 1 | 0% | 2,182 | 3,067 | +41% | 0 | 0 | — |
case-06 | fail→pass | 11,070 | 5,008 | -55% | 1 | 1 | 0% | 1,985 | 3,290 | +66% | 0 | 0 | — |
case-07 | fail→pass | 8,437 | 5,082 | -40% | 1 | 1 | 0% | 1,467 | 3,434 | +134% | 0 | 0 | — |
case-09 | fail→pass | 20,463 | 3,124 | -85% | 1 | 1 | 0% | 1,861 | 3,008 | +62% | 0 | 0 | — |
case-10 | fail→pass | 9,794 | 3,647 | -63% | 1 | 1 | 0% | 1,586 | 3,017 | +90% | 0 | 0 | — |
case-11 | fail→pass | 12,847 | 6,075 | -53% | 1 | 1 | 0% | 2,508 | 3,538 | +41% | 0 | 0 | — |
case-12 | fail→pass | 14,985 | 2,210 | -85% | 1 | 1 | 0% | 2,587 | 2,704 | +5% | 0 | 0 | — |
case-13 | fail→pass | 11,283 | 3,933 | -65% | 1 | 1 | 0% | 1,968 | 3,232 | +64% | 0 | 0 | — |
case-14 | fail→pass | 21,051 | 2,574 | -88% | 1 | 1 | 0% | 1,105 | 2,782 | +152% | 0 | 0 | — |
case-15 | fail→fail | 11,172 | 2,631 | -76% | 1 | 1 | 0% | 1,589 | 2,815 | +77% | 0 | 0 | — |
case-16 | pass→pass | 8,296 | 4,015 | -52% | 1 | 1 | 0% | 1,572 | 3,249 | +107% | 0 | 0 | — |
case-17 | fail→pass | 10,203 | 4,325 | -58% | 1 | 1 | 0% | 1,828 | 3,279 | +79% | 0 | 0 | — |
case-18 | pass→pass | 12,273 | 5,638 | -54% | 1 | 1 | 0% | 2,025 | 3,456 | +71% | 0 | 0 | — |
case-19 | fail→pass | 11,298 | 3,795 | -66% | 1 | 1 | 0% | 1,935 | 3,141 | +62% | 0 | 0 | — |
case-20 | fail→pass | 8,688 | 4,444 | -49% | 1 | 1 | 0% | 1,437 | 3,239 | +125% | 0 | 0 | — |
case-21 | pass→pass | 18,195 | 13,652 | -25% | 1 | 1 | 0% | 3,712 | 5,167 | +39% | 0 | 0 | — |
case-22 | pass→pass | 13,275 | 14,187 | +7% | 1 | 1 | 0% | 2,681 | 5,225 | +95% | 0 | 0 | — |
case-23 | pass→pass | 14,873 | 12,091 | -19% | 1 | 1 | 0% | 2,807 | 4,571 | +63% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 19 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +61 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.