Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run unit, integration, performance, and benchmark tests for gz-sim, filter by name, and debug a single failing case. Trigger when the user asks to test, rerun a test, debug a flaky test, or check that a change didn't regress something.
.claude/skills/harunkurtdev-gz-test/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 385% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-09 | ✓→✓ | = Same ✓ | -7% | 0% |
gz-sim ships three kinds of tests, all wired through CTest:
| Kind | Source root | Target prefix | |------|-------------|---------------| | Unit (next to source) | src/**/<Name>_TEST.cc | UNIT_<Name> | | Integration | test/integration/<name>.cc | INTEGRATION_<name> | | Performance | test/performance/<name>.cc | PERFORMANCE_<name> | | Benchmark | test/benchmark/<name>.cc | BENCHMARK_<name> | | Plugins (helpers) | test/plugins/ | built as test fixtures, not run directly |
shctest --test-dir build --output-on-failure -j"$(nproc)"
CI requires this to be clean before merge.
sh# only integration tests ctest --test-dir build -R '^INTEGRATION_' --output-on-failure # one specific suite ctest --test-dir build -R '^UNIT_EntityComponentManager$' -V # exclude slow/flaky ones during local iteration ctest --test-dir build -E 'PERFORMANCE_|BENCHMARK_'
This is the fastest way to debug a single failing case, because you get full gtest filter / repeat / shuffle flags:
sh./build/bin/INTEGRATION_diff_drive_system \ --gtest_filter='DiffDriveTest.*' \ --gtest_repeat=10 \ --gtest_shuffle
Most GUI tests need a display. Wrap with xvfb:
shxvfb-run -s '-screen 0 1280x1024x24' \ ctest --test-dir build -R '^INTEGRATION_' --output-on-failure
cmake --build build -j$(nproc) --target INTEGRATION_foo../build/bin/INTEGRATION_foo --gtest_filter=Bar.*.-DCMAKE_CXX_FLAGS="-fsanitize=thread").
sh gdb --args ./build/bin/INTEGRATION_foo --gtest_filter=Bar.NameOfCase
test/worlds/ — many integrationtests load a world from there.
sh./build/bin/BENCHMARK_each --benchmark_min_time=2 \ --benchmark_out=/tmp/each.json --benchmark_out_format=json
Avoid running benchmarks under load — they're noisy on shared machines.
shctest --test-dir build -R '^PYTHON_'
These exercise the pybind11 bindings.
Tell the user what you ran and which counts came back (e.g. "148 pass / 0 fail / 2 skipped"). Don't claim success without numbers.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,975 | 3,827 | -52% | 1 | 1 | 0% | 1,393 | 1,441 | +3% | 0 | 0 | — |
case-02 | fail→fail | 8,466 | 3,815 | -55% | 1 | 1 | 0% | 1,531 | 1,369 | -11% | 0 | 0 | — |
case-03 | fail→fail | 8,373 | 2,610 | -69% | 1 | 1 | 0% | 1,264 | 1,052 | -17% | 0 | 0 | — |
case-04 | fail→pass | 3,214 | 5,797 | +80% | 1 | 1 | 0% | 360 | 1,746 | +385% | 0 | 0 | — |
case-05 | fail→fail | 11,061 | 2,820 | -75% | 1 | 1 | 0% | 1,866 | 1,126 | -40% | 0 | 0 | — |
case-06 | fail→fail | 12,206 | 2,761 | -77% | 1 | 1 | 0% | 2,223 | 1,281 | -42% | 0 | 0 | — |
case-07 | fail→pass | 7,766 | 3,014 | -61% | 1 | 1 | 0% | 1,326 | 1,285 | -3% | 0 | 0 | — |
case-08 | fail→fail | 10,945 | 5,271 | -52% | 1 | 1 | 0% | 1,776 | 1,577 | -11% | 0 | 0 | — |
case-09 | pass→pass | 11,300 | 6,538 | -42% | 1 | 1 | 0% | 2,059 | 1,924 | -7% | 0 | 0 | — |
case-10 | fail→pass | 4,600 | 2,210 | -52% | 1 | 1 | 0% | 843 | 1,142 | +35% | 0 | 0 | — |
case-11 | pass→pass | 9,041 | 2,392 | -74% | 1 | 1 | 0% | 1,472 | 1,055 | -28% | 0 | 0 | — |
case-12 | fail→fail | 12,190 | 4,902 | -60% | 1 | 1 | 0% | 1,920 | 1,598 | -17% | 0 | 0 | — |
case-13 | pass→pass | 10,831 | 5,579 | -48% | 1 | 1 | 0% | 1,842 | 1,646 | -11% | 0 | 0 | — |
case-14 | fail→fail | 5,267 | 2,230 | -58% | 1 | 1 | 0% | 841 | 1,081 | +29% | 0 | 0 | — |
case-15 | pass→pass | 10,464 | 2,749 | -74% | 1 | 1 | 0% | 1,902 | 1,256 | -34% | 0 | 0 | — |
case-16 | pass→pass | 8,820 | 1,924 | -78% | 1 | 1 | 0% | 1,492 | 1,093 | -27% | 0 | 0 | — |
case-17 | pass→pass | 7,315 | 2,918 | -60% | 1 | 1 | 0% | 1,166 | 1,184 | +2% | 0 | 0 | — |
case-18 | fail→pass | 7,503 | 2,163 | -71% | 1 | 1 | 0% | 1,132 | 1,069 | -6% | 0 | 0 | — |
case-19 | pass→pass | 11,749 | 5,833 | -50% | 1 | 1 | 0% | 1,923 | 1,747 | -9% | 0 | 0 | — |
case-20 | pass→pass | 10,213 | 6,460 | -37% | 1 | 1 | 0% | 1,833 | 1,876 | +2% | 0 | 0 | — |
case-21 | pass→pass | 8,859 | 4,428 | -50% | 1 | 1 | 0% | 1,525 | 1,462 | -4% | 0 | 0 | — |
case-22 | pass→pass | 6,473 | 4,172 | -36% | 1 | 1 | 0% | 1,259 | 1,561 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.