Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run the Macro app on Cursor Cloud and pick up edits. Use when asked to run the app, open it in the browser, test a frontend or backend change, hot reload, or before touching stack.sh, rebuild.sh, frontend.sh, or just stack.
.claude/skills/macro-inc-run-app/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -44% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -28% | 0% |
The app is http://localhost:3000/app. Open it in the VM browser. The user's laptop cannot reach these URLs.
| Goal | Command | | --- | --- | | Start | bash .cursor/stack.sh | | After a Rust edit | bash .cursor/rebuild.sh | | After an apps/web edit | save the file |
The .cursor scripts re-enter the pinned nix shell. For any other docker, just, bun, or doppler command, prefix with nix develop /workspace --command.
Do not run just run_local. That is the laptop TUI.
bashbash .cursor/stack.sh
Let the first run finish. It compiles whatever changed since the environment bake. Do not kill the Nix build, copy binaries out of /nix/store, or replace this with hand-rolled docker compose.
Log in with any email. The login API returns the code. Codes also land in Mailpit at http://localhost:8025.
bashbash .cursor/rebuild.sh
Volumes stay. If the log says binaries are unchanged, the mounts were left alone.
Do not run just stack up or stack.sh --fresh to pick up an edit. up deletes the stack and recreates it. Use --fresh only when you mean to wipe.
cargo build and just build do not remount containers. rebuild.sh is the remount path.
apps/web editSave the file. Vite watches apps/web. If the page does not update, read ~/.cursor-cloud/frontend-dev.log for hmr update.
If http://localhost:3000/app stops responding, restart the dev server:
bashbash .cursor/frontend.sh stop && bash .cursor/frontend.sh
Use the same restart after a frontend dependency or Vite config change.
Do not run just stack update --frontend. That builds a static bundle. This Cloud stack does not serve one. On a --no-frontend stack that flag then tells you to re-run just stack up, which wipes volumes.
Bring the stack up, log in, and demonstrate the change in the UI. Unit tests and SQL probes are not a walkthrough.
If the page says "unable to lookup identity providers", FusionAuth is in maintenance mode. docker ps shows macro-fusionauth-1 as unhealthy. curl localhost:9011/api/status returns 302.
bashdocker restart macro-fusionauth-1
Wait until the container is healthy.
If the page loads and API calls fail, the backend is down. Run just stack status --json, then bash .cursor/stack.sh.
Do not chase agent_harness_service restart loops when DOPPLER_TOKEN is absent. That loop is expected.
bashjust stack status --json docker compose -p macro logs -f <service>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,509 | 7,350 | +33% | 1 | 1 | 0% | 149 | 933 | +526% | 0 | 0 | — |
case-02 | fail→fail | 4,976 | 7,085 | +42% | 1 | 1 | 0% | 441 | 975 | +121% | 0 | 0 | — |
case-03 | fail→fail | 7,851 | 7,260 | -8% | 1 | 1 | 0% | 148 | 1,032 | +597% | 0 | 0 | — |
case-04 | fail→pass | 9,730 | 3,397 | -65% | 1 | 1 | 0% | 1,501 | 1,107 | -26% | 0 | 0 | — |
case-05 | fail→pass | 10,451 | 4,627 | -56% | 1 | 1 | 0% | 1,484 | 1,339 | -10% | 0 | 0 | — |
case-06 | fail→pass | 12,910 | 3,411 | -74% | 1 | 1 | 0% | 2,121 | 1,189 | -44% | 0 | 0 | — |
case-07 | fail→pass | 12,417 | 3,065 | -75% | 1 | 1 | 0% | 1,536 | 1,101 | -28% | 0 | 0 | — |
case-08 | fail→pass | 12,957 | 4,751 | -63% | 1 | 1 | 0% | 1,985 | 1,427 | -28% | 0 | 0 | — |
case-09 | fail→fail | 12,729 | 3,751 | -71% | 1 | 1 | 0% | 1,981 | 1,140 | -42% | 0 | 0 | — |
case-10 | fail→fail | 8,179 | 3,508 | -57% | 1 | 1 | 0% | 1,320 | 1,219 | -8% | 0 | 0 | — |
case-11 | fail→fail | 9,421 | 51,423 | +446% | 1 | 1 | 0% | 1,555 | 997 | -36% | 0 | 0 | — |
case-12 | fail→pass | 11,979 | 2,711 | -77% | 1 | 1 | 0% | 1,750 | 1,014 | -42% | 0 | 0 | — |
case-13 | fail→pass | 9,274 | 3,602 | -61% | 1 | 1 | 0% | 1,373 | 1,119 | -18% | 0 | 0 | — |
case-14 | fail→fail | 17,975 | 6,561 | -63% | 1 | 1 | 0% | 2,820 | 958 | -66% | 0 | 0 | — |
case-15 | fail→pass | 11,997 | 2,821 | -76% | 1 | 1 | 0% | 1,765 | 987 | -44% | 0 | 0 | — |
case-16 | fail→fail | 3,351 | 3,862 | +15% | 1 | 1 | 0% | 492 | 1,186 | +141% | 0 | 0 | — |
case-17 | pass→pass | 16,230 | 3,414 | -79% | 1 | 1 | 0% | 2,214 | 1,115 | -50% | 0 | 0 | — |
case-18 | fail→pass | 14,608 | 5,341 | -63% | 1 | 1 | 0% | 2,258 | 1,422 | -37% | 0 | 0 | — |
case-19 | fail→pass | 13,067 | 6,171 | -53% | 1 | 1 | 0% | 1,938 | 1,056 | -46% | 0 | 0 | — |
case-20 | fail→pass | 14,053 | 5,014 | -64% | 1 | 1 | 0% | 2,294 | 1,398 | -39% | 0 | 0 | — |
case-21 | fail→fail | 10,973 | 3,929 | -64% | 1 | 1 | 0% | 1,610 | 1,173 | -27% | 0 | 0 | — |
case-22 | pass→pass | 11,069 | 9,738 | -12% | 1 | 1 | 0% | 1,758 | 2,141 | +22% | 0 | 0 | — |
case-23 | pass→fail | 10,172 | 7,577 | -26% | 1 | 1 | 0% | 1,439 | 877 | -39% | 0 | 0 | — |
case-24 | pass→pass | 11,759 | 9,595 | -18% | 1 | 1 | 0% | 1,959 | 2,203 | +12% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 18 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +42 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.