Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when creating a new task environment for EmbodiChain, including expert demonstration tasks, RL tasks or any EmbodiedEnv subclass
.claude/skills/dexforce-add-task-env/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 108% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 24% | 0% |
Own the task identity, task-first layout, physical environment, registration boundary, and solution routing. Do not assume every new task needs a Python environment subclass.
Classify the requested solution before creating files. Read only the matching reference files.
| Prompt signal | Route | Required reference | |---|---|---| | handwritten, scripted, demo segments, create_demo_segments, custom Python planning | Handwritten expert trajectory | references/handwritten-expert.md | | Task Program, program.yaml, integration.yaml, declarative task, Semantic Call, MLLM-generated program | Task Program expert trajectory | references/task-program-expert.md | | Generic expert trajectory or expert demo with no implementation signal | Resolve handwritten versus Task Program from required behavior; ask only if still ambiguous | One selected expert reference | | RL, PPO, GRPO, APG, policy, reward learning, training config | Reinforcement learning | references/rl.md | | Task/environment/scene only, with no requested solution | Environment-only baseline | This file only |
Routing rules:
keep all artifacts under one task-first directory.
expert trajectory and provides no signal forhandwritten versus Task Program, determine whether the requested behavior is declarative and supported by existing Semantic Calls. Ask only when that choice materially changes the requested deliverables.
implementation.
$add-task-program. New reusable robot/skill-profile declarations are owned by $add-embodiment-component.
Read agent_context/MAP.yaml, then load env-framework. Also load:
task-programs for the Task Program route;rl-learning for the RL route; andrequested environment.
Verify paths and configuration fields against the current source of truth. Do not infer schemas from older task examples alone.
Resolve from the prompt or nearby conventions:
<category_path>: a task family plus optional subdomain, such asmanipulation/tableware;
<task_name>: snake case;Organize by task identity, never by solution method. Use:
textembodichain_tasks/embodichain_tasks/<category_path>/<task_name>.py embodichain_tasks/configs/tasks/<category_path>/<task_name>/
Keep a single task Python entry point flat at <task_name>.py. Do not create a same-named Python package, scenario package, mdp package, or a task-local task_program Python package when configuration and manager functors express the behavior.
Choose the narrowest representation that supports the requested routes.
For new task-first compositions, prefer:
text<task config>/env.yaml
This is a component, not a runnable deployment. It owns:
environment_id;physics: default|newton backend and its optional,backend-matching physics_config;
simulation; andenv.It must not contain a runnable id, Task Program fields, semantic scene bindings, a robot, or sensors. Task Program semantic roots and affordances belong to integration.yaml.scene_binding.
Prefer one thin deployment per embodiment or execution variant:
text<task config>/task.<variant>.yaml
It owns the runnable id and selects reusable components. A typical physical deployment selects:
yamlid: MyTask-v1 environment: component: env.yaml embodiment: component: ../../../components/embodiments/<embodiment>.yaml
The original inline Gym format remains supported. When extending an existing inline env.json or env.yaml, preserve that representation unless the user asked for component extraction. Every inline runnable config declares exactly one physics: default|newton backend. Never select a component and repeat its owned inline fields, including physics or physics_config, in the same deployment. Use separate environment files for backend-specific settings.
Create embodichain_tasks/embodichain_tasks/<category_path>/<task_name>.py when the route requires import-owned behavior or registration:
A supported configuration-defined Task Program is the exception: it normally omits the task module and dynamically registers the common EmbodiedEnv while loading its runnable deployment.
Keep @register_env or @register_learning_env in the task-named module. Task discovery recursively imports these modules, so category __init__.py files do not need per-task re-exports.
After the shared baseline exists, follow the selected specialized reference:
references/handwritten-expert.mdreferences/task-program-expert.mdreferences/rl.mdReuse existing manager functors and registered components. Invoke $add-functor only when a missing observation, reward, event, action, dataset, or randomization term is required. Use $add-test for test structure.
At minimum:
embodichain list-task when task-discovery metadata changed; andDo not claim an expert trajectory is qualified from schema validation alone. Physical expert behavior requires an environment run with its validators and persisted completion result. Do not claim an RL task works from config parsing alone; at least construct/reset the selected environment and run a minimal trainer-routing smoke test when dependencies permit.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,954 | 8,674 | +9% | 1 | 1 | 0% | 171 | 1,903 | +1013% | 0 | 0 | — |
case-02 | fail→fail | 32,982 | 8,053 | -76% | 1 | 1 | 0% | 6,684 | 2,044 | -69% | 0 | 0 | — |
case-03 | fail→fail | 6,429 | 10,986 | +71% | 1 | 1 | 0% | 263 | 1,989 | +656% | 0 | 0 | — |
case-04 | pass→fail | 20,302 | 8,414 | -59% | 1 | 1 | 0% | 4,047 | 2,014 | -50% | 0 | 0 | — |
case-05 | fail→fail | 42,340 | 12,476 | -71% | 1 | 1 | 0% | 8,229 | 2,013 | -76% | 0 | 0 | — |
case-06 | fail→fail | 38,581 | 7,958 | -79% | 1 | 1 | 0% | 5,871 | 2,088 | -64% | 0 | 0 | — |
case-07 | pass→pass | 19,027 | 12,200 | -36% | 1 | 1 | 0% | 2,987 | 3,768 | +26% | 0 | 0 | — |
case-08 | fail→pass | 12,545 | 7,389 | -41% | 1 | 1 | 0% | 1,780 | 2,665 | +50% | 0 | 0 | — |
case-09 | pass→pass | 17,137 | 6,415 | -63% | 1 | 1 | 0% | 2,579 | 2,609 | +1% | 0 | 0 | — |
case-10 | pass→pass | 13,585 | 5,449 | -60% | 1 | 1 | 0% | 2,001 | 2,478 | +24% | 0 | 0 | — |
case-11 | fail→pass | 12,966 | 9,183 | -29% | 1 | 1 | 0% | 1,884 | 3,163 | +68% | 0 | 0 | — |
case-12 | fail→fail | 20,446 | 3,918 | -81% | 1 | 1 | 0% | 3,035 | 2,219 | -27% | 0 | 0 | — |
case-13 | fail→pass | 10,694 | 7,043 | -34% | 1 | 1 | 0% | 1,724 | 2,652 | +54% | 0 | 0 | — |
case-14 | fail→pass | 9,582 | 6,650 | -31% | 1 | 1 | 0% | 1,263 | 2,629 | +108% | 0 | 0 | — |
case-15 | fail→fail | 10,501 | 5,609 | -47% | 1 | 1 | 0% | 1,507 | 2,512 | +67% | 0 | 0 | — |
case-16 | fail→fail | 22,521 | 2,604 | -88% | 1 | 1 | 0% | 2,270 | 1,912 | -16% | 0 | 0 | — |
case-17 | fail→pass | 19,396 | 14,990 | -23% | 1 | 1 | 0% | 3,158 | 3,905 | +24% | 0 | 0 | — |
case-18 | fail→pass | 9,512 | 8,888 | -7% | 1 | 1 | 0% | 1,403 | 3,087 | +120% | 0 | 0 | — |
case-19 | fail→pass | 22,089 | 2,400 | -89% | 1 | 1 | 0% | 1,363 | 1,874 | +37% | 0 | 0 | — |
case-20 | fail→pass | 15,927 | 4,254 | -73% | 1 | 1 | 0% | 2,301 | 2,156 | -6% | 0 | 0 | — |
case-21 | fail→pass | 15,746 | 5,939 | -62% | 1 | 1 | 0% | 1,964 | 2,204 | +12% | 0 | 0 | — |
case-22 | fail→pass | 12,532 | 6,210 | -50% | 1 | 1 | 0% | 1,690 | 2,529 | +50% | 0 | 0 | — |
case-23 | fail→fail | 8,280 | 5,853 | -29% | 1 | 1 | 0% | 1,180 | 2,438 | +107% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/3/2026 | — |
| gemini-3.6-flash | verified | 8/27/2026 | +57% |
| gemini-3.6-flash | verified | 8/22/2026 | +23% |
Other measured skills in the registry, with their headline benchmark lift.