Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user asks to create, scaffold, or edit Jupyter notebooks (`.ipynb`) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script `new_notebook.py` to generate a clean starting notebook.
.claude/skills/davila7-jupyter-notebook/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -40% | 0% |
Create clean, reproducible Jupyter notebooks for two primary modes:
Prefer the bundled templates and the helper script for consistent structure and fewer JSON mistakes.
.ipynb notebook from scratch.experiment.tutorial.bashexport CODEX_HOME="${CODEX_HOME:-$HOME/.codex}" export JUPYTER_NOTEBOOK_CLI="$CODEX_HOME/skills/jupyter-notebook/scripts/new_notebook.py"
User-scoped skills install under $CODEX_HOME/skills (default: ~/.codex/skills).
Identify the notebook kind: experiment or tutorial. Capture the objective, audience, and what "done" looks like.
Use the helper script to avoid hand-authoring raw notebook JSON.
bashuv run --python 3.12 python "$JUPYTER_NOTEBOOK_CLI" \ --kind experiment \ --title "Compare prompt variants" \ --out output/jupyter-notebook/compare-prompt-variants.ipynb
bashuv run --python 3.12 python "$JUPYTER_NOTEBOOK_CLI" \ --kind tutorial \ --title "Intro to embeddings" \ --out output/jupyter-notebook/intro-to-embeddings.ipynb
Keep each code cell focused on one step. Add short markdown cells that explain the purpose and expected result. Avoid large, noisy outputs when a short summary works.
For experiments, follow references/experiment-patterns.md. For tutorials, follow references/tutorial-patterns.md.
Preserve the notebook structure; avoid reordering cells unless it improves the top-to-bottom story. Prefer targeted edits over full rewrites. If you must edit raw JSON, review references/notebook-structure.md first.
Run the notebook top-to-bottom when the environment allows. If execution is not possible, say so explicitly and call out how to validate locally. Use the final pass checklist in references/quality-checklist.md.
assets/experiment-template.ipynb and assets/tutorial-template.ipynb.Script path:
$JUPYTER_NOTEBOOK_CLI (installed default: $CODEX_HOME/skills/jupyter-notebook/scripts/new_notebook.py)tmp/jupyter-notebook/ for intermediate files; delete when done.output/jupyter-notebook/ when working in this repo.ablation-temperature.ipynb).Prefer uv for dependency management.
Optional Python packages for local notebook execution:
bashuv pip install jupyterlab ipykernel
The bundled scaffold script uses only the Python standard library and does not require extra dependencies.
No required environment variables.
references/experiment-patterns.md: experiment structure and heuristics.references/tutorial-patterns.md: tutorial structure and teaching flow.references/notebook-structure.md: notebook JSON shape and safe editing rules.references/quality-checklist.md: final validation checklist.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,374 | 7,733 | -68% | 1 | 1 | 0% | 3,161 | 1,279 | -60% | 0 | 0 | — |
case-02 | fail→fail | 22,087 | 7,659 | -65% | 1 | 1 | 0% | 3,974 | 1,239 | -69% | 0 | 0 | — |
case-03 | fail→fail | 16,746 | 8,955 | -47% | 1 | 1 | 0% | 3,163 | 1,475 | -53% | 0 | 0 | — |
case-04 | pass→pass | 10,488 | 8,346 | -20% | 1 | 1 | 0% | 1,422 | 2,319 | +63% | 0 | 0 | — |
case-05 | pass→pass | 7,557 | 4,713 | -38% | 1 | 1 | 0% | 1,247 | 1,677 | +34% | 0 | 0 | — |
case-06 | pass→pass | 11,735 | 8,272 | -30% | 1 | 1 | 0% | 2,012 | 2,378 | +18% | 0 | 0 | — |
case-07 | pass→pass | 17,153 | 12,902 | -25% | 1 | 1 | 0% | 2,771 | 2,998 | +8% | 0 | 0 | — |
case-08 | pass→pass | 8,369 | 3,827 | -54% | 1 | 1 | 0% | 1,371 | 1,556 | +13% | 0 | 0 | — |
case-09 | fail→pass | 10,649 | 2,515 | -76% | 1 | 1 | 0% | 1,978 | 1,264 | -36% | 0 | 0 | — |
case-10 | fail→fail | 3,755 | 2,344 | -38% | 1 | 1 | 0% | 622 | 1,198 | +93% | 0 | 0 | — |
case-11 | fail→pass | 8,991 | 2,275 | -75% | 1 | 1 | 0% | 1,427 | 1,305 | -9% | 0 | 0 | — |
case-12 | fail→pass | 7,971 | 2,571 | -68% | 1 | 1 | 0% | 1,264 | 1,355 | +7% | 0 | 0 | — |
case-13 | fail→pass | 7,624 | 3,829 | -50% | 1 | 1 | 0% | 1,264 | 1,601 | +27% | 0 | 0 | — |
case-14 | fail→pass | 12,456 | 1,794 | -86% | 1 | 1 | 0% | 1,959 | 1,177 | -40% | 0 | 0 | — |
case-15 | fail→pass | 12,448 | 2,762 | -78% | 1 | 1 | 0% | 1,950 | 1,266 | -35% | 0 | 0 | — |
case-16 | fail→pass | 9,302 | 1,678 | -82% | 1 | 1 | 0% | 1,308 | 1,144 | -13% | 0 | 0 | — |
case-17 | fail→pass | 11,208 | 2,573 | -77% | 1 | 1 | 0% | 1,801 | 1,260 | -30% | 0 | 0 | — |
case-22 | pass→pass | 7,825 | 2,084 | -73% | 1 | 1 | 0% | 1,191 | 1,252 | +5% | 0 | 0 | — |
case-18 | pass→pass | 8,920 | 4,094 | -54% | 1 | 1 | 0% | 1,443 | 1,696 | +18% | 0 | 0 | — |
case-19 | pass→pass | 10,504 | 2,814 | -73% | 1 | 1 | 0% | 1,671 | 1,398 | -16% | 0 | 0 | — |
case-20 | pass→pass | 12,638 | 9,295 | -26% | 1 | 1 | 0% | 1,930 | 2,260 | +17% | 0 | 0 | — |
case-21 | pass→pass | 11,862 | 7,514 | -37% | 1 | 1 | 0% | 1,962 | 2,007 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.