Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn an object or character reference image into a quality-gated, animation-ready procedural Three.js model built in code. Use for image-to-3D reconstruction, detail-accurate object rebuilds, stylized/likeness-maximized human characters, sculpt specs, and staged code generation.
.claude/skills/img2threejs-img2threejs/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 147% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 182% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 300% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 372% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 287% | 0% |
Rebuild the object visible in a reference image as a code-only procedural Three.js model, gated by a staged sculpting pipeline and an AI-vision self-correction loop. This is reconstruction-by-code, not photogrammetry, mesh extraction, or downloaded art packs. That promise governs how the model is built — it says nothing about which file formats it can subsequently be exported to; an explicitly-selected emission target (--target <kind>) is a terminal, whole-artifact transform of the already-built model, verified to its own stated limit, never a second way to build one.
Agent-agnostic: works under Claude Code, Codex, or OpenCode. Wherever this doc says "agent vision" or "agent browser tool", use whatever the host provides — native image reading, a browser MCP (playwright/chrome-devtools), the project preview, or a user-supplied screenshot.
This file is the always-loaded router: it holds the order of operations and every hard rule as one line. The full contract behind each rule lives in the grimoire/ or docs/ file that rule names — read the named file at the moment you reach that stage, not before.
Keep one checkout of this repository and let every host enter it through a symlink, so Claude and Codex execute the same code instead of drifting apart:
text~/.claude/skills/img2threejs -> <your checkout> ~/.codex/skills/img2threejs -> <your checkout>
The user attaches/points to an object image and wants a procedural Three.js model, a reconstruction/animation/destruction plan, a sculpt spec, or code. Also for material studies, action-ready props, game objects, botanical/mechanical parts, and stylized reconstructions.
Sculpt from a photo, in order — never one-shot a mesh:
python3 forge/next.py --state .img2threejs/state.json [<spec>] first, at every start,resume, and before every correction iteration. It reports the ordered checklist, exact next command, evidence status, and bounded correction-loop status; it never replaces the spec/pass gates. Obey a hard stop; never continue from memory.
grimoire/intake/validation_rubric.md).qualityContract before any code.identity-defining feature is wrong even when the global score looks fine.
State explicitly when output is approximate/stylized/low-poly. A single image cannot reveal hidden sides or guarantee exact geometry — say so instead of faking confidence.
Conversation context is disposable; .img2threejs/state.json is the local checklist authority. Initialize once per reconstruction, then gate every step through it:
bashpython3 forge/state.py init --state .img2threejs/state.json --reference <img> --profile <generic|character|installed-domain> --spec object-sculpt-spec.json python3 forge/next.py --state .img2threejs/state.json [object-sculpt-spec.json] python3 forge/state.py mark <step-id> --state .img2threejs/state.json --evidence <path>
next.py prints the current step, pass, incomplete mandatory steps, exact next command, andloop/max. Exit code 3 or status=stopped is a hard stop: report the reason and request input. Never bypass it by reconstructing progress from chat history.
skipped only with --reason —silent omission is forbidden. Loop counts derive from reviewHistory actions (refine-spec/refine-code), not agent memory. Defaults: 3 corrections per pass, 6 total.
modules (character) and installed plugins (cs2, animated-character from plugin-character) register identically, and forge/state.py init names what is available. A profile adds mandatory gates without changing the core order -- a domain plugin typically requires an authoritative classification, an intake manifest, and a machine-readable domain review before AI review; character requires the character contracts and landmark evidence; animated-character (requires the installed plugin-character) adds all of character plus the nine Stage R steps (grimoire/readiness/animation_contract.md). Pick it whenever the rig must MOVE — on character the Stage R gates are absent and the build completes without ever running them, which is how animation used to ship broken. Its order is load-bearing: repair the mesh, freeze it, bind additively, then verify parity. Every profile records suitability, projection applicability, and material-evidence applicability. The state file is a resumability index, not visual evidence: renders, specs, review history, and deterministic gates remain the authoritative artifacts.
(default: real-time browser prop with interactive performance)
an explicit request for the user/vision provider to supply one; heuristic detection alone is not enough to select a geometry adapter
Run scripts from the skill root (forge/...). Pure Python 3.10+ stdlib, no pip installs. Full flags: grimoire/scripts.md. Never let a script score visuals — that is the agent's job.
protocol in grimoire/intake/image_analysis.md — identify/classify, decompose macro→meso→micro, map part relationships, name materials in PBR terms, list identity-defining features, and flag what the single view hides. Observation before inference; controlled 3D vocabulary; 3D object-space not 2D image-space. Then probe local images: forge/stage1_intake/probe_image.py <image> (metadata only, not a visual check). 1a. Local Spec Search — after image analysis, before writing or refining a spec, pull local domain evidence (anatomy/PBR/wear/geometry/runtime/physics) rather than inventing it: python3 forge/stage2_spec/new_pre_spec_assessment.py "Name" --image <img> --out assessment.json (auto-runs BM25 over the core_3d collection, or the collection a declared domain contributes -- the collection is NEVER guessed from the target name; writes a localSpecSearch bundle that new_sculpt_spec.py --assessment carries into the spec). Full query-expansion recipe (bilingual terms, focused search_specs.py retrieval, cache rules): grimoire/intake/local_spec_search.md. MUST read it before retrying an incomplete or domain-specific query. 1b. Domain intake — when a domain plugin serves the item, complete its intake steps before pre-spec authoring (admission, heuristic signal, classification, family/route resolution). MUST read the contract its step names, completely, before creating the manifest or running pre-spec assessment. 1c. Optional fidelity evidence adapters — only when they improve an observed weak point; the stdlib core remains authoritative. Thin/complex masks → local SAM2; character face/pose → MediaPipe; weak front/back cues → Depth Anything V2 (forge/stage1_intake/run_vision_adapter.py <segment|landmarks|depth> ...; every adapter emits provenance; monocular depth is relative only). MCP-only scene mutations never count as implementation — write the proven change back to the spec or TypeScript, rebuild, recapture. Full adapter + MCP routing and authority boundaries: docs/integrations/reference_fidelity_tooling.md.
forge/stage2_spec/new_pre_spec_assessment.py "Name" --image <img> --complexity <simple|moderate|complex|ultra-complex> --out assessment.json. Rules: grimoire/intake/quality_contract.md. Set objectClass.primaryDomain (object | character | hybrid) and fill the seeded detailInventory (its targetMinDetails scales with complexity). A domain plugin may raise these floors through its augmentation -- the merge clamps, so a plugin can never lower one (a skin's finish/wear/hardware IS the item, so such a domain is held to the top fidelity bar. Author procedural GEOMETRY but route the FINISH through the projection path in step 2c — a procedural finish for a patterned skin (Doppler/Gamma/Marble/Fade) reads visibly wrong against the reference. A domain plugin ships its own finish rulebook and texture-acquisition guide; read what its checklist steps name. 2b. Detail inventory (do not skip for detailed subjects) — scan zones and enumerate every identity-defining small detail (gloss, bevel, fasteners, linework, contours, stains): forge/stage1_intake/build_detail_inventory.py <image> --mode grid-3x3 --out-dir <dir> --out di.json. Each detail MUST map to a component.localFeatures or material.localOverrides entry — never prose only. Taxonomy + 3D-term recipes: grimoire/intake/detail_inventory.md. 2c. Projection-first fidelity (characters AND reference-matched surfaces — painted skins, decals, painted patterns) — when the goal is matching a specific reference's surface, put the photo's own pixels on the mesh instead of approximating them procedurally. This is the single biggest fidelity lever; a procedural material for a patterned surface is the #1 reconstruction failure. Recipe (grimoire/character/likeness_maximization.md — its two levers generalize past characters): solve the camera (stage1_intake/solve_camera_pose.py → referenceCamera), de-light the reference (stage1_intake/delight_albedo.py, hard requirement — de-lighting is what makes projection safe), then project the de-lit crop and bake it into UVs (stage3_build/bake_projected_texture.py --mesh-id <id>). For a painted skin the projected de-lit crop IS the finish — no procedural Doppler material. For characters, first capture landmarks (stage1_intake/extract_landmarks.py --out anatomy.json), fill preSpecAssessment.anatomy, route grimoire/character/reconstruction.md. A single view cannot show hidden sides — report per-region confidence and request more views when it matters. Character sub-routes, in order — decide what parts exist before shaping any, and shape the head before the hair that sits on it:
grimoire/character/structure_decomposition.mdgrimoire/character/head_construction.md (what the likeness gate reads against)grimoire/character/stylized_hair_threejs.md + parameter contract ingrimoire/character/threejs_hair_parameter_contract.json. Lock topology only after the silhouette review passes: material tuning cannot repair wrong lock topology. 2d. Reference-free humanoid — a generic figure with no reference image has nothing to measure, so fill anatomy from public canon: forge/stage2_spec/humanoid_proportions.py <spec> --style-heads 8 --in-place. It writes anatomy.source: "canon-table" so canon is never mistaken for measurement, refuses to run when the spec names a reference image, and names anything the corpus does not supply rather than interpolating it.
forge/stage2_spec/new_sculpt_spec.py "Name" --image <img> --assessment assessment.json --augmentation spec-augmentation.json --domain <profile> --out object-sculpt-spec.json (the checklist step carries the resolved flags). Replace generic starter featureReviewTargets with the object's real identity-defining systems (≤5 critical, ≤3 important per pass); for characters add anatomy-proportion, face-landmark-placement, pose-silhouette, outfit-and-palette. Use 3D-graphics terms only (grimoire/glossary/3d_vocabulary.md), never "nice/smooth/shiny". Classify every component's topologyClass/topologyRationale per grimoire/intake/surface_topology.md before picking a primitive — this is what prevents a continuous organic form from being picked as a box.
extract reference PBR evidence, both per crop (verify the crop is on the part you think it is):
forge/stage1_intake/analyze_texture.py <crop> --spec spec.json --material-id <id> --in-placeclassifies the finish, extracts the gradient palette, and writes doc-grounded MeshPhysicalMaterial scalars onto the material. Recipes + Three.js texture/PBR rules: grimoire/build/threejs_texture_reference.md. Rule of thumb: solid albedo for flat paint, real reference crop for patterned finishes.
forge/stage1_intake/extract_pbr_evidence.py <crop> --out-dir <dir> --material-id <id> --target-threshold 0.7.Confidence < 0.7 is a stop/refine-input signal, not a pass. It is inference, not inverse rendering.
forge/stage1_intake/material_region_analysis.py --manifest regions.json --out-dir material-evidence --out material-analysis.json,resolve each assignment from docs/materials/material-reference.json, wire it in with forge/stage2_spec/apply_material_analysis.py.
forge/stage4_review/material_views.py),compare visible-footprint crops (material_comparator.py), apply only bounded material-scoped corrections (material_feedback.py), and record the blocking result (material_gate.py).
forge/stage2_spec/validate_sculpt_spec.py object-sculpt-spec.json then --strict-quality. Strict blocks shallow specs (a complex object with one root, no repetition systems, no local overrides, no micro groups is NOT implementation-ready even if JSON validates).
forge/stage3_build/orchestrate_passes.py status object-sculpt-spec.json forge/stage3_build/generate_threejs_factory.py object-sculpt-spec.json --out src/createObjectModel.ts The generator is fail-closed: strict-quality must pass before it writes any factory, and a future --pass-id fails until prior passes are reviewed continue. If blocked, preserve the BLOCKED artifact and refine the subject-specific spec; do not substitute a generic template. The local state adds --force only for a new pass or refine-spec; refine-code edits the current artifact without regenerating it. Before overwriting, carry valid hand refinement back into the spec; generated code must not be the only copy of reconstruction decisions. 6a. Hitting a triangle budget. performanceBudget.targetTriangles selects a tessellation tier for every primitive with segment counts (low ≤6k, standard ≤60k, else hero) and caps implicit-surface sampling grids. Where a tier is not precise enough, add geometryDescriptor.decimate: {"targetRatio": 0.4} to that component — a quadric collapse in the generated factory, run before skin binding so weights are computed on surviving vertices. It keeps position only (normals recomputed), so it is refused on an authored/unwrapped uvStrategy. Offline LOD tiers: forge/stage3_build/decimate.py <mesh.json> --ratio <r> --json.
7a. Off-axis and placement gates — a single review viewpoint is not evidence about the model. Capture a turntable, not one frame, and run all three; each catches a defect class the older gates pass by construction (a hole through a skull, a hat at hip height and a floating charm all survived eight front-only review rounds): forge/stage4_review/turntable_gate.py --capture 0=front.png --capture 90=right.png --capture 180=rear.png --capture 270=left.png --json node runtime/scripts/export_mesh_geometry.mjs --url <preview> --out meshes.json then forge/stage4_review/self_intersection.py meshes.json --json forge/stage4_review/attachment_anchor.py object-sculpt-spec.json --measured measured.json --json All three exit 0 clean / 1 gate failure / 2 error. A failure blocks continue even when the global fidelity score passes. Read sampledVertexCount / unmeasuredAttachments / missingAzimuths before believing a clean verdict: each names what the gate did not look at.
grimoire/review/gates_reference.md and grimoire/review/self_correction.md completely. Run forge/stage4_review/diagnose_render.py and record the passing Tier 1 result with --spec object-sculpt-spec.json --pass-id <pass> --in-place; for non-planar forms also run forge/stage4_review/diagnose_render_multi_angle.py with the fixed view and at least two meaningful orbit views. Then run forge/stage3_build/orchestrate_passes.py check object-sculpt-spec.json --pass-id <pass>.
forge/stage4_review/make_comparison_sheet.py --reference <img> --render <shot> --out cmp.png --json.
forge/stage4_review/append_review.py object-sculpt-spec.json --pass-id <pass> --fidelity <0-1> --action <continue|refine-spec|refine-code|request-input|stop> --summary "..." --render-screenshot <shot> --comparison-image cmp.png --ai-vision-score <0-1> --layer-scores-json '{...}' --feature-reviews-json <f.json> --in-place. When a domain plugin contributes a review gate, produce its versioned report first with the command that plugin's review step names, then attach it with --domain-review-json <report>.json --review-scene-json <the plugin's scene fixture>. The checklist step carries the resolved paths. A failed family, painted-region, projection-coverage, critical-detail, or orbit gate blocks continue even when the global score passes. See the plugin's own review-gate documentation.
state gate before another correction or pass: forge/stage3_build/orchestrate_passes.py sync object-sculpt-spec.json --in-place python3 forge/next.py --state .img2threejs/state.json object-sculpt-spec.json.
forge/stage4_review/check_part_coverage.py --spec object-sculpt-spec.json --manifest parts.json and verify the action-ready hierarchy. Mark part-coverage and action-ready only with evidence.
When the user supplies a GLB as an intermediate reference, the browser-rendered GLB is the structural and visual baseline for an independently authored procedural factory. The raw GLB is never pixel evidence and its topology/materials are never copied into the factory. Before any factory edit — full contract in grimoire/build/python_threejs_render_bridge.md, machine-readable schema in docs/specs/render-profile.v2.schema.json (+ example; fail-closed validation):
forge/stage1_intake/probe_glb.py first. A merged one-node/one-mesh asset is insufficientfor semantic labels; request a multipart GLB or a browser semantic-ID pass before claiming exact regions.
render-profile.v2 (forge/stage4_review/validate_render_profile.py) used byboth the GLB and procedural routes. Region IDs are subject-specific, never inherited from the example profile; declare the required set in extensions.requiredSemanticRegions so omission is a hard validation error.
beauty, alpha-silhouette, semantic-id, depth,normal, roughness-material-id); score with forge/stage4_review/compare_region_passes.py. Missing semantic-ID data blocks per-region confidence rather than falling back to whole-image scores.
with floating primitives when the region's silhouette requires a continuous surface.
→ materials → lighting; recapture the full pass set after each group and record the changed group, hashes and score. Never combine groups when diagnosing improvement.
Before any visual review or continue decision, MUST read the full gate-by-gate contract in grimoire/review/gates_reference.md (Divine Eye, VLM rescue, multi-angle, interior difference, chirality, hair, domain review, bounded correction, Divine Eye fitting, screenshot feedback, assembly, attachment, material, detail inventory, rig payload, character track). In short:
grimoire/intake/validation_rubric.md, check_reference_admission.py).divine_eye.py is deterministic-first; the VLM (vlm_gate.py) is a gated last layer, neverconsulted on a hard-gate failure.
diagnose_render_multi_angle.py).interior_difference.py). Silhouette IoU reads~11% of figure cells: a model with its face deleted scored the same 0.8803 as the finished face.
-l/-r pair is a MIRROR, not a rotation — hard at spec time (validate_chirality). A pairwrong the same way on both sides still passes, and needs medial_lateral_bias vs a reference.
scalp_exposure.py is HARD and runs on geometry before any render; hair_gate.pyis soft and subordinate to it. A coverage shortfall never authorises widening the masses.
identity feature, so their boundaries are gated on geometry: vertex_region_gate.py. Never a texture — this pipeline emits code; the shape predicates live in _shared/vertex_paint.py.
swept_arc_gate.py: silhouette IoUpasses a straight cone occupying roughly the right cells.
stage5_rig/validate_rig_payload.py) before binding aTHREE.Skeleton; it proves payload integrity only, never pose stress or likeness.
grimoire/readiness/animation_contract.md;the checklist steps and gate come from the installed plugin-character -- stage5_rig/ remains in this repo as the emitter's library, not the checklist authority). A clip that exists is not a clip that plays: only G1 (maxSampledBindingDelta <= 2^-23) separates the two, and a gate whose input is missing reports unevaluated, never a pass. Bind at IDENTITY in attached mode and take the display offset from the mesh bounds alone; loop is decided by poseReturn, never by travel.
hard stop. correction_loop.py may stop earlier on repeated defects, oscillation, or plateau.
continue requires a render + comparison sheet + AI-vision score ≥ threshold, every criticalfeature ≥ its own threshold (grimoire/feedback/render_capture.md).
(check_part_coverage.py, grimoire/build/geometry_patterns.md).
grimoire/readiness/action_rigging.md, grimoire/readiness/joint_attachment.md, grimoire/feedback/shading_realism.md, grimoire/intake/quality_contract.md, grimoire/intake/validation_rubric.md.
After every pass, decide exactly one: continue | refine-spec | refine-code | request-input | stop. refine-spec fixes a wrong/missing/shallow spec (re-validate, don't patch code around it); refine-code fixes geometry/material/lighting that doesn't match a sound spec. Before making the decision, MUST read the root-cause guide + fidelity scale in grimoire/review/self_correction.md, record the decision, and re-run the local state gate.
Small features need a different instrument. Divine Eye's SSIM/tonal/edge signals run on a 64×64 luma grid, so a detail a few pixels wide is absent before any comparison happens. When fidelity depends on individual tears, spars, fangs or eyes, use the four-tier microscope: grimoire/review/divine_eye_microscope.md. Two empirically established rules from it: measure fidelity on a component's visible footprint (full frame minus a component-hidden frame), never on an isolation render; and never colour-gate a concave feature, where a dark ratio captures cavity shading rather than material.
Report what changed each pass with evidence (exact values/coordinates), name what still doesn't match, and never claim "done" when only "improved". A passing gate is not proof of 3D realism. Full rule + examples: grimoire/review/self_correction.md.
A left/right pair is a reflection, never a rotation: negate the lateral axis and nothing else, (x, y, z) → (-x, y, z). With forward: +Z, Y up and a right-handed frame, the character's own left is +X. The convention lives as code in forge/_shared/chirality.py (CHARACTER_LEFT_SIGN), with two different gates for the two defects that shipped from getting it wrong: validate_chirality catches a rotation-mistaken-for-reflection at spec time, and medial_lateral_bias vs a reference catches a pair that is wrong the same way on both sides. Reflecting also inverts triangle winding — flip it back on the mirrored side or flatShading lights the limb as though lit from behind. Full write-up with the measured defects: grimoire/scripts.md ("Left and right").
Hair has its own subsystem because it has failure modes no other gate can see. Full contract, measurements and non-goals: docs/HAIR_PIPELINE.md. The hard rules:
(u, v), never absolute positions (hard validation error).standProud is enforced by the generator, not advisory.scalp_exposure.py is a HARD gate on geometry before any render; a coverage shortfall never onits own authorises widening the masses.
shell, not locks; strand impression comes from faceting andmaterial (hair.human.code-only), since this skill emits no textures.
plane-card is rejected for hair (needs an alpha texture this skill cannot emit).A domain plugin makes a run exact where this pipeline would otherwise infer. It contributes its own checklist steps and gates, and publishes a spec-augmentation.json that this pipeline pulls at spec-authoring. With no plugin serving the item, nothing here changes: author the skeleton and infer the shape from the reference, as for any other object.
what each step names -- a plugin ships its own contract, and it governs its own domain.
blocked run: it is the generic path, and the reconstruction proceeds by inference.
Plugins are installed and managed by the img2 harness (img2threejs/img2), a separate dependency-free CLI (Node launcher, Python core). This skill never installs anything itself: when state.py init names a profile as unavailable, name the img2 add command that installs it and stop — never vendor domain logic instead.
bashimg2 add img2threejs/plugin-<id> --ref <tag> # install a plugin at a tag (e.g. plugin-cs2) img2 doctor # what each host resolves: base-skill path + version, plugin gates img2 capabilities --from-kind image --to-kind glb # which installed plugin serves an edge
Setup and the full CLI reference live in the README quick-start and the harness repo's docs.
Subdivision runtime tests compile generated TypeScript against the showcase checkout. Set IMG2THREEJS_SHOWCASE_ROOT to that checkout; without it, local runtime-only tests skip with an actionable message while static contracts still run. CI should set IMG2THREEJS_REQUIRE_SHOWCASE=1 to turn a missing showcase checkout into a test failure.
bashIMG2THREEJS_SHOWCASE_ROOT=/path/to/img2threejs-showcase python3 forge/tests/test_subdivision.py IMG2THREEJS_SHOWCASE_ROOT=/path/to/img2threejs-showcase python3 -m unittest discover -s forge/tests IMG2THREEJS_SHOWCASE_ROOT=/path/to/img2threejs-showcase python3 forge/tests/test_showcase_tsc_smoke.py
TypeScript + plain Three.js unless the project uses a wrapper. Group factory createObjectNameModel(spec, options), reconstruction data kept separate from renderer objects, deterministic seeds for all procedural noise. Prefer primitives / Shape extrude / curve+tube / instancing / displacement / generated canvas textures before any external art. Full geometry & material recipes + hard-won failure patterns: grimoire/build/geometry_patterns.md.
When Python is requested for character rendering, use it as a deterministic job/evidence layer around the browser Three.js runtime: camera-batch manifests, source/output hashes, readiness and settle checks, screenshot persistence, masks, diagnostics, and comparison packaging. The target Three.js browser route remains the rendering authority. Do not silently replace the procedural TypeScript factory with Blender/VRM/GLB output. Full routing, manifest fields, and failure rules: grimoire/build/python_threejs_render_bridge.md.
Use grimoire/readiness/standard_character_pipeline.md for character work. Beta owns the strict sculpt/build/review gates; alpha owns deterministic camera manifests, browser screenshot evidence and UniRig-shaped rig validation. CharacterGen, Tripo, VRM and other neural/asset systems are opt-in adapters with source, checkpoint, license, coordinate conversion and output hashes. They never silently replace the procedural TypeScript factory. Image-to-mesh systems emit a static mesh with no skeleton, so their output is never animation-ready however good it looks. Executable entry points: forge/stage4_review/render_bridge.py and scripts/capture_threejs_playwright.py (init → browser capture → validate → diagnose; capture must operate on the real showcase/browser route and leave readable PNGs in the workspace).
geometry strategy, material/lighting recipe, animation/destruction feasibility, plan + risks.
a narrower target. "This cannot reach the requested fidelity from this image" is a valid result.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 45,366 | 17,182 | -62% | 1 | 1 | 0% | 7,843 | 8,685 | +11% | 0 | 0 | — |
case-02 | fail→fail | 27,220 | 16,876 | -38% | 1 | 1 | 0% | 3,382 | 8,565 | +153% | 0 | 0 | — |
case-03 | fail→fail | 47,449 | 15,677 | -67% | 1 | 1 | 0% | 8,251 | 8,600 | +4% | 0 | 0 | — |
case-04 | pass→pass | 23,487 | 18,424 | -22% | 1 | 1 | 0% | 3,102 | 10,638 | +243% | 0 | 0 | — |
case-05 | fail→pass | 28,668 | 14,508 | -49% | 1 | 1 | 0% | 3,929 | 9,716 | +147% | 0 | 0 | — |
case-06 | pass→pass | 20,208 | 13,454 | -33% | 1 | 1 | 0% | 2,156 | 9,638 | +347% | 0 | 0 | — |
case-07 | fail→pass | 26,827 | 12,063 | -55% | 1 | 1 | 0% | 3,329 | 9,380 | +182% | 0 | 0 | — |
case-08 | fail→pass | 22,648 | 18,081 | -20% | 1 | 1 | 0% | 2,631 | 10,520 | +300% | 0 | 0 | — |
case-09 | fail→pass | 18,166 | 11,963 | -34% | 1 | 1 | 0% | 1,983 | 9,368 | +372% | 0 | 0 | — |
case-10 | fail→pass | 19,637 | 10,518 | -46% | 1 | 1 | 0% | 2,387 | 9,245 | +287% | 0 | 0 | — |
case-11 | pass→pass | 21,598 | 10,549 | -51% | 1 | 1 | 0% | 2,368 | 9,201 | +289% | 0 | 0 | — |
case-12 | fail→pass | 17,634 | 10,314 | -42% | 1 | 1 | 0% | 1,881 | 9,160 | +387% | 0 | 0 | — |
case-13 | fail→pass | 19,191 | 11,041 | -42% | 1 | 1 | 0% | 2,056 | 9,302 | +352% | 0 | 0 | — |
case-14 | fail→pass | 22,826 | 9,914 | -57% | 1 | 1 | 0% | 3,025 | 9,026 | +198% | 0 | 0 | — |
case-15 | fail→pass | 24,609 | 12,660 | -49% | 1 | 1 | 0% | 3,078 | 9,345 | +204% | 0 | 0 | — |
case-16 | fail→pass | 13,914 | 10,769 | -23% | 1 | 1 | 0% | 1,334 | 9,305 | +598% | 0 | 0 | — |
case-17 | fail→pass | 20,482 | 11,467 | -44% | 1 | 1 | 0% | 2,674 | 9,411 | +252% | 0 | 0 | — |
case-18 | fail→pass | 20,417 | 12,641 | -38% | 1 | 1 | 0% | 2,266 | 9,576 | +323% | 0 | 0 | — |
case-19 | fail→pass | 19,369 | 14,498 | -25% | 1 | 1 | 0% | 2,183 | 9,939 | +355% | 0 | 0 | — |
case-20 | pass→pass | 34,685 | 9,515 | -73% | 1 | 1 | 0% | 2,205 | 9,115 | +313% | 0 | 0 | — |
case-21 | fail→fail | 21,723 | 18,385 | -15% | 1 | 1 | 0% | 3,133 | 8,754 | +179% | 0 | 0 | — |
case-22 | fail→fail | 26,276 | 37,966 | +44% | 1 | 1 | 0% | 4,580 | 14,987 | +227% | 0 | 0 | — |
case-23 | fail→fail | 29,406 | 29,405 | -0% | 1 | 1 | 0% | 4,346 | 12,624 | +190% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 19 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +57 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/24/2026 | +45% |
| gemini-3.6-flash | verified | 8/11/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.