Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page or the docs site to click through it, screenshot it, or reproduce a report. Covers the four dev entry points and their ports, which data root each writes to, and the four ways a healthy setup looks broken.
.claude/skills/prism-shadow-penguin-harness-manual-test/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 6 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -3% | 0% |
Node >= 24. dev:* runs dev-prebuild.mjs first (keeps pnpm install current, prebuilds workspace deps); pnpm desktop runs a full pnpm -r build, so it is slow to start.
| Command | Open | Data root | | --- | --- | --- | | pnpm dev | http://localhost:7365 | ~/.penguin/dev-data | | pnpm desktop | its own window | ~/.penguin/dev-data | | pnpm dev:landing | http://localhost:7366 | none (static) | | pnpm dev:docs | http://localhost:7367 | none (static) |
Other fixed ports (packages/core/src/internal/ports.ts): 7364 installed server, 7368 dev backend, 7369 pnpm penguin web (data root ~/.penguin/dev-data-cli — its own, so it can serve while an Agent it hosts runs pnpm dev; the lock is per root). On a shared box, ss -tln before assuming one is free; PORT= inline moves it.
The user's installed app, server and CLI all use ~/.penguin/data — their real Agents, Sessions and keys. Never point a dev run there. Both surfaces print the root they took (Data root: …, [shell] dev instance … on data root …); read it rather than assume.
7368 shows a stale app. The dev backend also serves packages/web/dist — the last pnpm -r build, not what Vite is serving. Screenshot 7365, never 7368.
curl returns 502 but the server is fine. A shell http_proxy routes loopback through the proxy. Use curl --noproxy '*'. Browsers and the server itself are unaffected.
/api on 127.0.0.1 returns 401. That address is reserved as the Workspace-preview host. Use localhost.
The server exits 3 saying the data root is in use. Another instance — usually the user's own desktop app — holds <root>/server.lock. The lock is per root, not per port, so PORT= will not get past it. Use a different root; do not kill their process.
Never export it. scripts/run-with-env.mjs applies each VAR=value only when unset, so an exported value silently redirects every dev script — an exported ~/.penguin/data puts pnpm dev on the user's real data, where the lock is already held. It now names any default the environment displaced before running, so read that line if it appears.
When you need your own root, prefix the one command:
shPENGUIN_HOME=~/.penguin/dev-data-<topic> pnpm dev
Unset it rather than blanking it: run-with-env reads empty as unset, but resolveRoot() uses ?? and takes an empty string literally.
A fresh root seeds admin with a random penguin-<4 digits> password, printed once at startup. PENGUIN_SEED_ADMIN_PASSWORD pins it, which is how the e2e harness logs in.
Changes to core or skills do not reach a running dev server — web/server consume snapshot copies that re-sync only when that package's build runs. Restart pnpm dev.
Stop what you started. Leave anything the user was already running alone.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 19,520 | 16,337 | -16% | 1 | 1 | 0% | 2,145 | 2,695 | +26% | 0 | 0 | — |
case-02 | fail→pass | 20,317 | 7,656 | -62% | 1 | 1 | 0% | 2,459 | 2,204 | -10% | 0 | 0 | — |
case-03 | fail→pass | 21,444 | 15,284 | -29% | 1 | 1 | 0% | 2,405 | 2,385 | -1% | 0 | 0 | — |
case-04 | pass→pass | 17,713 | 17,277 | -2% | 1 | 1 | 0% | 2,280 | 2,869 | +26% | 0 | 0 | — |
case-05 | pass→pass | 21,156 | 14,234 | -33% | 1 | 1 | 0% | 2,655 | 2,276 | -14% | 0 | 0 | — |
case-06 | pass→pass | 19,580 | 24,401 | +25% | 1 | 1 | 0% | 2,435 | 4,125 | +69% | 0 | 0 | — |
case-07 | fail→pass | 12,558 | 10,963 | -13% | 1 | 1 | 0% | 1,893 | 1,698 | -10% | 0 | 0 | — |
case-08 | fail→pass | 13,177 | 11,975 | -9% | 1 | 1 | 0% | 2,038 | 1,985 | -3% | 0 | 0 | — |
case-09 | fail→fail | 16,463 | 16,667 | +1% | 1 | 1 | 0% | 1,727 | 2,820 | +63% | 0 | 0 | — |
case-10 | fail→pass | 16,136 | 6,948 | -57% | 1 | 1 | 0% | 1,894 | 1,923 | +2% | 0 | 0 | — |
case-11 | fail→pass | 18,587 | 9,503 | -49% | 1 | 1 | 0% | 1,917 | 1,593 | -17% | 0 | 0 | — |
case-12 | fail→pass | 18,365 | 8,891 | -52% | 1 | 1 | 0% | 2,288 | 1,516 | -34% | 0 | 0 | — |
case-13 | pass→pass | 15,987 | 9,883 | -38% | 1 | 1 | 0% | 1,867 | 1,675 | -10% | 0 | 0 | — |
case-14 | pass→pass | 17,745 | 7,833 | -56% | 1 | 1 | 0% | 1,906 | 2,184 | +15% | 0 | 0 | — |
case-15 | fail→pass | 9,725 | 11,049 | +14% | 1 | 1 | 0% | 1,703 | 1,917 | +13% | 0 | 0 | — |
case-16 | fail→pass | 18,843 | 10,156 | -46% | 1 | 1 | 0% | 1,957 | 1,737 | -11% | 0 | 0 | — |
case-17 | fail→pass | 21,410 | 8,310 | -61% | 1 | 1 | 0% | 3,428 | 1,301 | -62% | 0 | 0 | — |
case-18 | fail→fail | 132,305 | 7,876 | -94% | 1 | 1 | 0% | 6,445 | 1,324 | -79% | 0 | 0 | — |
case-19 | pass→pass | 10,849 | 9,184 | -15% | 1 | 1 | 0% | 1,570 | 1,342 | -15% | 0 | 0 | — |
case-20 | fail→pass | 13,705 | 4,468 | -67% | 1 | 1 | 0% | 1,180 | 1,482 | +26% | 0 | 0 | — |
case-21 | fail→pass | 21,723 | 5,084 | -77% | 1 | 1 | 0% | 2,574 | 1,558 | -39% | 0 | 0 | — |
case-22 | fail→pass | 23,775 | 9,436 | -60% | 1 | 1 | 0% | 4,064 | 1,497 | -63% | 0 | 0 | — |
case-23 | fail→pass | 16,693 | 7,776 | -53% | 1 | 1 | 0% | 1,866 | 1,212 | -35% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +65 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.