Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when writing tests for EmbodiChain modules, including observation functors, reward functors, solvers, sensors, environments, or any Python module
.claude/skills/dexforce-add-test/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 180% | 0% |
Write tests following EmbodiChain's conventions and patterns.
Tests mirror the source tree under tests/:
embodichain/lab/sim/motion/solvers/pytorch_solver.py → tests/sim/motion/solvers/test_pytorch_solver.py
embodichain/lab/gym/envs/managers/rewards.py → tests/gym/envs/managers/test_reward_functors.py
embodichain/toolkits/graspkit/pg_grasp/foo.py → tests/toolkits/test_pg_grasp.py
embodichain_tasks/embodichain_tasks/manipulation/push_cube.py → tests/gym/envs/tasks/test_push_cube.pyRules:
test_<module>.pyembodichain/ structure under tests/__init__.py files in new tests/ subdirectories if neededUse when: testing functors, utility functions, pure math, config validation — anything that doesn't need a SimulationManager.
python# ---------------------------------------------------------------------------- # Copyright (c) 2021-2026 DexForce Technology Co., Ltd. # # Licensed under the Apache License, Version 2.0 (the "License"); # ... # ---------------------------------------------------------------------------- from __future__ import annotations import pytest import torch from embodichain.my_module import my_function def test_expected_output(): result = my_function(input_value) assert result == expected_value def test_edge_case(): result = my_function(edge_input) assert result is not None
Use when: tests need SimulationManager, GPU setup, or must run in a specific order. Share state via setup_method/teardown_method.
python# ---------------------------------------------------------------------------- # Copyright (c) 2021-2026 DexForce Technology Co., Ltd. # # Licensed under the Apache License, Version 2.0 (the "License"); # ... # ---------------------------------------------------------------------------- from __future__ import annotations import pytest import torch from embodichain.lab.sim import SimulationManager, SimulationManagerCfg class TestMySimComponent: def setup_method(self): config = SimulationManagerCfg(headless=True, device="cpu") self.sim = SimulationManager(config) # ... setup ... def teardown_method(self): self.sim.destroy() SimulationManager.flush_cleanup_queue() def test_basic_behavior(self): result = self.sim.do_something() assert result == expected_result def test_raises_on_bad_input(self): with pytest.raises(ValueError): self.sim.do_something(bad_input)
Use the narrowest test type that proves the behavior:
tests/conftest.py automatically classifies conventional CUDA, renderer, andreal-simulation tests. Add @pytest.mark.gpu, @pytest.mark.slow, or @pytest.mark.requires_sim explicitly only when the test's node id/source cannot reveal that requirement (for example, a hidden CUDA helper or an end-to-end toolkit test).
GPU tests are skipped by default to keep normal test runs within the shared VRAM budget. Run them explicitly and serially:
bash# Default suite: GPU-marked tests are skipped. pytest tests/ # Dedicated GPU suite. Do not add -n unless pytest-xdist is installed. pytest tests/ --run-gpu -m gpu
For backend/device matrices, run the complete contract on one representative configuration and use small (one environment, low-resolution) smoke tests for the remaining configurations. Always destroy a real SimulationManager and flush its cleanup queue in teardown.
Choose the test surface by ownership layer:
| Changed boundary | Primary test location | |---|---| | Language schema, strict decoder, AST/compiler | tests/lab/task_program/ | | Semantic scene/profile/call/effect contracts | tests/lab/task_program/semantics/ | | Configured integration, catalog, simulation assembly | tests/gym/envs/task_program/ | | Gym bridge and episode completion | tests/gym/envs/task_program/, tests/gym/envs/test_embodied_env_task_program.py | | Packaged task components/deployments | tests/test_task_program_package_data.py, task-layout/config tests | | Lightweight RL environment/trainer routing | tests/learning/ |
For JSON/YAML configuration, test through the same strict loader and component composition used by production. A yaml.safe_load() assertion alone does not prove closed fields, relative component paths, physical scene targets, contracts, catalog coverage, or program preflight.
Add negative coverage for the earliest ownership boundary being changed: unknown fields, missing component files, duplicate inline/component ownership, absent simulation_uid, contract mismatch, unknown Semantic Calls, or missing registered lowerers. Keep live-simulation qualification separate from provider-free decode/compiler tests.
For a complete configured Task Program deployment, use $add-task-program's read-only inspector as a focused smoke check before adding a heavier environment test.
Most functor tests don't need a live simulation. Use mock objects following the pattern in tests/gym/envs/managers/test_reward_functors.py:
pythonfrom unittest.mock import MagicMock, Mock class MockSim: """Mock simulation for functor tests.""" def __init__(self, num_envs: int = 4): self.num_envs = num_envs self.device = torch.device("cpu") self._rigid_objects: dict = {} def get_rigid_object(self, uid: str): return self._rigid_objects.get(uid) def add_rigid_object(self, obj): self._rigid_objects[obj.uid] = obj class MockEnv: """Mock environment for functor tests.""" def __init__(self, num_envs: int = 4): self.num_envs = num_envs self.device = torch.device("cpu") self.sim = MockSim(num_envs)
Key points for mock objects:
num_envs and device attributes (functors use these)MagicMock(uid="...") for SceneEntityCfg parametersAsk the user:
Map the source path to test path:
embodichain/<subpath>/<module>.py → tests/<subpath>/test_<module>.pyCheck if the test file already exists — append new test classes/functions if so.
dotdigraph test_style { rankdir=LR; "Needs SimulationManager?" -> "Class style" [label="yes"]; "Needs SimulationManager?" -> "pytest style" [label="no"]; "Tests share state/order?" -> "Class style" [label="yes"]; "Tests share state/order?" -> "pytest style" [label="no"]; }
Use the appropriate template (pytest or class style above).
Rules:
from __future__ import annotations — after header, before importstest_<scenario> (descriptive, not just test_foo)if __name__ == "__main__" BlockInclude this for tests that support optional visual/interactive debugging:
pythonif __name__ == "__main__": # For visual debugging: set is_visual=True when calling env methods test_obj = TestMyComponent() test_obj.setup_method() # ... manually run test logic ...
bash# Single file pytest tests/<subpath>/test_<module>.py -v # Single test function pytest tests/<subpath>/test_<module>.py::test_expected_output -v # GPU-specific test pytest tests/<subpath>/test_<module>.py --run-gpu -m gpu -v # Single test class method pytest tests/<subpath>/test_<module>.py::TestMyClass::test_basic_behavior -v
blackbashblack tests/<subpath>/test_<module>.py
| Convention | Rule | |-----------|------| | File header | Apache 2.0 copyright block (same 15 lines as source) | | File naming | test_<module>.py | | Function naming | test_<scenario> | | from __future__ | Required after header | | Magic numbers | Define as named constants with explanatory comments | | Simulation tests | Initialize/teardown in setup_method/teardown_method | | CUDA coverage | Use @pytest.mark.gpu; run with --run-gpu -m gpu | | Long integration | Use @pytest.mark.slow; keep it out of normal PR runs | | Pure-logic tests | Use mock objects, no real sim | | SceneEntityCfg | Use MagicMock(uid="...") in tests | | Assertions | assert, pytest.approx, torch.allclose, pytest.raises | | Entry block | if __name__ == "__main__" for visual debugging support |
| Mistake | Fix | |---------|-----| | Missing Apache header on test file | Copy the 15-line copyright block | | Using real SimulationManager for functor tests | Use MockEnv/MockSim — much faster, no GPU needed | | Hardcoded numbers without explanation | Define as EXPECTED_DISTANCE = 0.5 # cube at origin, target at (0.5, 0, 0) | | Testing multiple concepts in one function | Split into separate test_<scenario> functions | | Forgetting cleanup | Call self.sim.destroy() and SimulationManager.flush_cleanup_queue() in teardown | | Using full matrices | Use one full representative case and low-resource smoke coverage elsewhere | | Not running black on test file | CI checks all files including tests |
| Action | Command | |--------|---------| | Run default tests | pytest tests/ | | Run GPU tests | pytest tests/ --run-gpu -m gpu | | Run single file | pytest tests/<path>/test_<name>.py -v | | Run single test | pytest tests/<path>::test_<name> -v | | Run with print output | pytest -s tests/<path>/test_<name>.py | | Format | black tests/<path>/test_<name>.py |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 58,421 | 18,315 | -69% | 1 | 1 | 0% | 3,270 | 6,487 | +98% | 0 | 0 | — |
case-02 | fail→pass | 20,561 | 15,292 | -26% | 1 | 1 | 0% | 4,044 | 5,613 | +39% | 0 | 0 | — |
case-03 | fail→pass | 25,648 | 32,508 | +27% | 1 | 1 | 0% | 5,215 | 7,527 | +44% | 0 | 0 | — |
case-04 | fail→pass | 18,631 | 14,488 | -22% | 1 | 1 | 0% | 3,113 | 5,751 | +85% | 0 | 0 | — |
case-05 | fail→fail | 14,437 | 16,424 | +14% | 1 | 1 | 0% | 2,734 | 6,054 | +121% | 0 | 0 | — |
case-06 | fail→pass | 8,785 | 10,092 | +15% | 1 | 1 | 0% | 1,585 | 4,437 | +180% | 0 | 0 | — |
case-07 | fail→pass | 18,107 | 13,724 | -24% | 1 | 1 | 0% | 3,171 | 5,499 | +73% | 0 | 0 | — |
case-08 | fail→fail | 23,405 | 18,779 | -20% | 1 | 1 | 0% | 3,838 | 6,596 | +72% | 0 | 0 | — |
case-09 | fail→pass | 14,516 | 7,675 | -47% | 1 | 1 | 0% | 2,129 | 3,878 | +82% | 0 | 0 | — |
case-10 | fail→pass | 9,830 | 11,957 | +22% | 1 | 1 | 0% | 1,757 | 4,961 | +182% | 0 | 0 | — |
case-11 | fail→pass | 20,009 | 16,264 | -19% | 1 | 1 | 0% | 3,473 | 5,898 | +70% | 0 | 0 | — |
case-12 | fail→pass | 11,793 | 9,171 | -22% | 1 | 1 | 0% | 1,699 | 4,561 | +168% | 0 | 0 | — |
case-13 | fail→pass | 14,892 | 10,831 | -27% | 1 | 1 | 0% | 2,408 | 4,659 | +93% | 0 | 0 | — |
case-14 | fail→pass | 13,519 | 15,692 | +16% | 1 | 1 | 0% | 2,233 | 5,330 | +139% | 0 | 0 | — |
case-15 | fail→fail | 19,523 | 17,169 | -12% | 1 | 1 | 0% | 3,874 | 6,035 | +56% | 0 | 0 | — |
case-16 | fail→pass | 14,427 | 12,136 | -16% | 1 | 1 | 0% | 2,289 | 4,638 | +103% | 0 | 0 | — |
case-17 | fail→pass | 18,076 | 14,653 | -19% | 1 | 1 | 0% | 2,914 | 5,276 | +81% | 0 | 0 | — |
case-18 | fail→pass | 8,047 | 10,888 | +35% | 1 | 1 | 0% | 1,389 | 4,405 | +217% | 0 | 0 | — |
case-19 | pass→fail | 19,444 | 16,659 | -14% | 1 | 1 | 0% | 3,446 | 5,897 | +71% | 0 | 0 | — |
case-20 | pass→pass | 18,593 | 21,824 | +17% | 1 | 1 | 0% | 3,930 | 7,431 | +89% | 0 | 0 | — |
case-21 | pass→fail | 19,618 | 24,048 | +23% | 1 | 1 | 0% | 3,984 | 8,140 | +104% | 0 | 0 | — |
case-22 | fail→pass | 14,708 | 3,362 | -77% | 1 | 1 | 0% | 989 | 3,304 | +234% | 0 | 0 | — |
case-23 | pass→pass | 17,862 | 14,418 | -19% | 1 | 1 | 0% | 3,028 | 5,430 | +79% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +61 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/3/2026 | +48% |
| gemini-3.6-flash | verified | 8/27/2026 | +50% |
| gemini-3.6-flash | verified | 8/22/2026 | +9% |
Other measured skills in the registry, with their headline benchmark lift.