Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Before declaring any task complete, actually verify the outcome. Run the code. Test the fix. Check the output. AI-generated code optimizes for plausible-looking output, not verified-correct output. Use when completing code changes, bug fixes, or any task where correctness matters.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 397% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 347% | 0% |
| case-02 | ✓→✗ | ▼ Worse | -33% | 0% |
| case-03 | ✓→✗ | ▼ Worse | -47% | 0% |
| case-09 | ✓→✗ | ▼ Worse | -14% | 0% |
AI-generated code comes from pattern-matching on training data. Something can look syntactically perfect, follow best practices, and still be wrong. The model optimizes for "looks right" not "works right." Verification is a separate cognitive step that must be explicitly triggered. This skill closes the loop between implementation and proof.
Limitations this skill addresses:
1. Generation vs Execution Code is generated but not run during generation. Confidence comes from "this looks like working code" not from "this was executed and the result observed."
2. Pattern-Matching Blindness Code that matches common patterns feels correct. But subtle bugs hide in the gaps between patterns. Off-by-one errors. Wrong variable names. Missing edge cases. These "look right" but aren't.
3. Confidence-Correctness Gap High confidence in output doesn't correlate with actual correctness. The agent is often most confident when most wrong, because the wrong answer pattern-matched strongly.
4. No Feedback Loop Code is generated sequentially. There's no natural "go back and check" step. Without explicit verification, errors compound silently.
ALWAYS verify before declaring complete:
Code Changes:
Fixes:
Claims:
Before saying any of these:
STOP. You haven't verified yet.
| Change Type | Verification | |-------------|--------------| | New code | Run it with test input | | Bug fix | Reproduce original bug, confirm fixed | | Function change | Call the function, check output | | Config change | Restart service, test affected feature | | Build script | Run the build | | API endpoint | Make a request | | UI change | Describe what user should see, or screenshot |
bash# Don't just write the test - run it python -m pytest tests/test_new_feature.py # Don't just fix the code - prove the fix python -c "from module import func; print(func(edge_case))" # Don't just update config - verify it loads node -e "console.log(require('./config.js'))"
Verified:
What I changed:
- Added input validation to user_signup()
How I verified:
- Ran: python -c "from auth import user_signup; user_signup('')"
- Expected: ValidationError
- Got: ValidationError("Email required")
Proof that it works. Done.Minimal test that proves basic functionality:
bash# After writing a new function python -c "from new_module import new_func; print(new_func('test'))"
If this crashes, you're not done.
After fixing a bug, trigger the original failure:
bash# Bug was: crash on empty input python -c "from module import func; func('')" # Should not crash anymore
Before claiming code is complete:
bash# Does it at least compile/parse? python -m py_compile new_file.py npm run build cargo check
After changes that affect multiple components:
bash# Start the service npm run dev & # Hit the affected endpoint curl http://localhost:3000/affected-route # Check for expected response
Before declaring done:
[ ] Code runs without syntax errors
[ ] Function produces expected output on happy path
[ ] Edge case handled (empty, null, boundary)
[ ] Original bug/issue no longer reproduces
[ ] No new errors introduced (check logs)
[ ] Build passes (if applicable)User: "Fix the divide by zero error"
*edits code to add check*
Done! Added a check for zero before dividing. The function now returns
None if the divisor is zero.This is unverified. The code was never run.
User: "Fix the divide by zero error"
*edits code to add check*
Let me verify the fix:
$ python -c "from math_utils import safe_divide; print(safe_divide(10, 0))"
None
$ python -c "from math_utils import safe_divide; print(safe_divide(10, 2))"
5.0
Verified:
- Division by zero returns None (was: crashed)
- Normal division still works
- Fix confirmed.User: "Update the production config"
I've updated config/production.json with the new API endpoint.
Verification status: UNABLE TO VERIFY
- This is production config
- I can't safely test against production
- Recommend: Test in staging first, or have ops verify after deploy
Change is made but unverified. Proceed with caution.When "Done!" is said without verifying, it's not reporting a fact. It's making a prediction based on pattern-matching. Sometimes that prediction is wrong.
Verification converts prediction into observation. It's the difference between "this should work" and "this works."
One is a guess. One is proof.
Prove it.
Other measured skills in the registry, with their headline benchmark lift.