Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Converts cuTile Python GPU kernels (@ct.kernel) to cuTile.jl Julia equivalents. Handles kernel syntax translation, 0-indexed to 1-indexed conversion, broadcasting differences, memory layout (row-major to column-major), type system mapping, and launch API differences. Use when converting, porting, or translating cuTile Python kernels to Julia cuTile.jl, or debugging/optimizing existing Julia cuTile translations.
.claude/skills/nvidia-tilegym-converting-cutile-to-julia/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 58% | 0% |
Convert @ct.kernel Python kernels to Julia function ... end cuTile.jl kernels.
translations/workflow.mdMethodError, IRError, numerical mismatch) → references/debugging.mdreferences/api-mapping.md + references/critical-rules.mdreferences/testing.mdJulia kernels are standalone — no Python bridge, no pytest integration. The Julia sub-project lives in julia/ at the repo root with its own Project.toml for dependency management.
julia/ # Self-contained Julia sub-project
├── Project.toml # Dependencies: CUDA.jl, cuTile.jl, NNlib.jl, Test
├── kernels/ # cuTile.jl kernel implementations
│ ├── add.jl # ← Ground-truth: 1D element-wise with alpha scaling (tensor+tensor, tensor+scalar)
│ ├── matmul.jl # ← Ground-truth: 2D tiled MMA, standard Julia layout (M,K)×(K,N)→(M,N)
│ └── softmax.jl # ← Ground-truth: 3 strategies (TMA, online, chunked) using ct.load/ct.store
└── test/ # Julia-native tests (using Test stdlib)
├── runtests.jl # Test runner entry point
├── test_add.jl
├── test_matmul.jl
└── test_softmax.jlGround-truth reference: Always consult julia/kernels/*.jl and julia/test/*.jl for patterns that compile and pass tests. These are the canonical examples of working cuTile.jl code.
julia/kernels/<op>.jl with cuTile.jl kernel + bridge function(s)translations/workflow.md Phase 2)references/api-mapping.md + references/critical-rules.md)julia/test/test_<op>.jl using Test stdlib + NNlib.jl for referenceinclude(...) in julia/test/runtests.jlpython <skill-dir>/scripts/validate_cutile_jl.py <file.jl>julia --project=julia/ julia/test/runtests.jlFull conversion checklist with post-conversion verification → translations/workflow.md
The most dangerous translation errors. Full rules (17 total) in references/critical-rules.md.
| # | Pitfall | One-line fix | |---|---------|-------------| | 1 | ct.full() doesn't exist in Julia | Use fill(val, shape), zeros(T, dims...), or ones(T, dims...) | | 2 | max(a, b) on tiles → IRError | Use max.(a, b) (broadcast dot) | | 3 | IRError / MethodError mentioning IRStructurizer | Compiler bug — file upstream with minimal reproducer | | 4 | ct.launch arg order silently wrong | Args are positional — match kernel signature exactly | | 5 | ct.load with order — index positions wrong | order remaps BOTH shape AND index (Critical Rule 16) |
Side-by-side Python → Julia conversions matching the released Julia kernels in julia/kernels/. Each directory contains cutile_python.py (before) and cutile_julia.jl (after).
| # | Example | Key Patterns | When to Reference | |---|---------|-------------|-------------------| | 01 | add | 1D ct.load/ct.store, alpha scaling, scalar broadcast, fill/zeros, keyword load/store | Starting point; basic TMA + element-wise patterns | | 02 | matmul | muladd, TF32 conversion, K-loop with for, 2D swizzle, standard Julia layout, ct.@compiler_options | MMA / tensor core operations | | 03 | softmax | Persistent scheduling, for loops, gather/scatter, padding_mode, multi-pass | Large-tensor reduction patterns |
These match the released kernels in julia/kernels/ (add.jl, matmul.jl, softmax.jl). The examples are simplified teaching versions — always consult julia/kernels/*.jl for the canonical, tested implementations.
| Category | Document | Content | |----------|----------|---------| | Workflows | translations/workflow.md | Full conversion workflow with todo list, validation loop, checklist | | Rules | references/critical-rules.md | 17 Critical Rules for cuTile Python → Julia conversion | | API | references/api-mapping.md | Python↔Julia bidirectional API mapping + kernel patterns | | Testing | references/testing.md | Julia-native test patterns, tolerances, failure diagnosis | | Debugging | references/debugging.md | Julia-specific error diagnosis + IR debug commands | | Scripts | scripts/validate_cutile_jl.py | Static validation for Julia anti-patterns (run it) | | Ground Truth | julia/kernels/*.jl + julia/test/*.jl | Actual working implementations in the codebase |
Prerequisite — Julia: this skill requires the Julia version declared in julia/Project.toml under [compat] julia. If julia --version is missing or older than that, install from the official Julia site at <https://julialang.org/install/> following the verified installer instructions for your OS. Resume below once julia --version is compatible.
Then, from the repo root:
bash# Install Julia dependencies declared in julia/Project.toml julia --project=julia/ -e 'using Pkg; Pkg.instantiate()' # Run tests julia --project=julia/ julia/test/runtests.jl
Requirements:
julia/Project.toml under [compat] julia)julia/Project.toml: CUDA.jl, cuTile.jl, NNlib.jl, Test| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-19 | pass→pass | 13,158 | 7,636 | -42% | 1 | 1 | 0% | 2,403 | 3,168 | +32% | 0 | 0 | — |
case-05 | pass→pass | 6,197 | 3,349 | -46% | 1 | 1 | 0% | 1,136 | 2,376 | +109% | 0 | 0 | — |
case-01 | fail→fail | 15,545 | 4,336 | -72% | 1 | 1 | 0% | 3,732 | 1,905 | -49% | 0 | 0 | — |
case-02 | fail→fail | 3,935 | 33,329 | +747% | 1 | 1 | 0% | 211 | 8,246 | +3808% | 0 | 0 | — |
case-03 | fail→fail | 5,095 | 5,829 | +14% | 1 | 1 | 0% | 239 | 1,870 | +682% | 0 | 0 | — |
case-04 | pass→pass | 11,628 | 9,340 | -20% | 1 | 1 | 0% | 2,442 | 3,567 | +46% | 0 | 0 | — |
case-06 | pass→pass | 8,757 | 5,462 | -38% | 1 | 1 | 0% | 1,813 | 2,835 | +56% | 0 | 0 | — |
case-07 | pass→pass | 11,483 | 5,507 | -52% | 1 | 1 | 0% | 2,176 | 2,738 | +26% | 0 | 0 | — |
case-08 | pass→pass | 10,961 | 5,403 | -51% | 1 | 1 | 0% | 1,883 | 2,642 | +40% | 0 | 0 | — |
case-09 | fail→pass | 16,046 | 5,592 | -65% | 1 | 1 | 0% | 2,608 | 2,637 | +1% | 0 | 0 | — |
case-10 | fail→pass | 9,112 | 1,866 | -80% | 1 | 1 | 0% | 1,737 | 2,023 | +16% | 0 | 0 | — |
case-11 | pass→pass | 10,705 | 4,488 | -58% | 1 | 1 | 0% | 1,790 | 2,410 | +35% | 0 | 0 | — |
case-12 | fail→pass | 13,332 | 7,793 | -42% | 1 | 1 | 0% | 2,478 | 3,147 | +27% | 0 | 0 | — |
case-13 | fail→pass | 10,140 | 3,262 | -68% | 1 | 1 | 0% | 1,971 | 2,398 | +22% | 0 | 0 | — |
case-14 | fail→pass | 7,967 | 3,079 | -61% | 1 | 1 | 0% | 1,438 | 2,273 | +58% | 0 | 0 | — |
case-15 | fail→pass | 3,448 | 1,224 | -65% | 1 | 1 | 0% | 627 | 1,868 | +198% | 0 | 0 | — |
case-16 | fail→pass | 10,344 | 3,460 | -67% | 1 | 1 | 0% | 1,778 | 2,304 | +30% | 0 | 0 | — |
case-17 | fail→pass | 9,992 | 3,093 | -69% | 1 | 1 | 0% | 1,666 | 2,126 | +28% | 0 | 0 | — |
case-18 | pass→pass | 15,467 | 10,619 | -31% | 1 | 1 | 0% | 2,770 | 3,873 | +40% | 0 | 0 | — |
case-20 | fail→pass | 5,490 | 2,102 | -62% | 1 | 1 | 0% | 1,039 | 1,993 | +92% | 0 | 0 | — |
case-21 | pass→pass | 16,330 | 8,290 | -49% | 1 | 1 | 0% | 2,843 | 3,381 | +19% | 0 | 0 | — |
case-22 | fail→pass | 10,156 | 3,195 | -69% | 1 | 1 | 0% | 1,557 | 2,191 | +41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.