Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill provides a comprehensive verification and quality assurance system that ensures code quality and correctness through:
.claude/skills/verification-quality-assurance/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 255% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 277% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 1579% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 827% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 462% | 0% |
This skill provides a comprehensive verification and quality assurance system that ensures code quality and correctness through:
> Shipped vs. aspirational. The concrete, in-CI verification stack — the 6 regression-guard jobs + the witness manifest + the tool-discoverability audit — is real and runs on every push. The truth-scoring / auto-rollback / WebSocket-dashboard surface described later in this doc is partly shipped (ruflo verify runs the witness checks) and partly design — treat the "CI Guards" section below as the authoritative current state.
Ruflo's regression protection is three layers, all gated before publish. Authoritative reference: verification/README.md.
| Layer | What | CI job(s) in .github/workflows/v3-ci.yml | ADR | |---|---|---|---| | 1 — install/behavioral smoke | Exercise user-visible failure modes against a real build | smoke-install-no-bsqlite (npm install on platforms w/o prebuilds), plugin-hooks-smoke (#1859/#1862 — hook flag parsing), mcp-protocol-smoke (#1874 — HTTP MCP wire format), memory-import-smoke (#1883/#1884 — WSL path + key sanitization), mcp-roundtrip-smoke (#1889 paired-tool round-trip + #1863 cli-no-crash + ADR-095 G2 consensus-transport) | ADR-102 | | 1 — discoverability gate | Every MCP tool description must answer "use this over native when?" | tool-descriptions-audit — scripts/audit-tool-descriptions.mjs, baseline at verification/mcp-tool-baseline.json (monotone-decreasing: noGuidance / tooShort / duplicates) | ADR-112 | | 2 — cryptographic witness | Every documented fix's load-bearing marker must still be present in dist; Ed25519-signed, per-OS bundles | witness-verify (ubuntu/macos/windows) — plugins/ruflo-core/scripts/witness/verify.mjs against verification/<os>/manifest.md.json | ADR-103 | | 3 — temporal history | When was a regression introduced | verification/<os>/history.jsonl + history.mjs (summary / regressions / timeline) | ADR-103 |
bash# Tool-description discoverability audit (ADR-112) node scripts/audit-tool-descriptions.mjs # fails if any baseline count rises node scripts/audit-tool-descriptions.mjs --update-baseline # lock the new floor after a fix lands # Behavioral smokes (each builds what it needs; safe to run individually) node plugins/ruflo-core/scripts/test-hooks.mjs "node $PWD/v3/@claude-flow/cli/bin/cli.js" node plugins/ruflo-core/scripts/test-mcp-protocol.mjs node plugins/ruflo-core/scripts/test-memory-import.mjs node plugins/ruflo-core/scripts/test-mcp-roundtrips.mjs # #1889 paired-tool round-trip node plugins/ruflo-core/scripts/test-cli-no-crash.mjs # #1863 unhandled-exception class node plugins/ruflo-core/scripts/test-consensus-transport.mjs # ADR-095 G2 consensus transport # Witness manifest — regenerate + verify node scripts/regen-witness.mjs node plugins/ruflo-core/scripts/witness/verify.mjs --manifest verification/macos/manifest.md.json # Temporal history node plugins/ruflo-core/scripts/witness/history.mjs --history verification/macos/history.jsonl summary node plugins/ruflo-core/scripts/witness/history.mjs --history verification/macos/history.jsonl regressions
plugins/ruflo-core/scripts/test-<name>.mjs. Pattern: static dist-scan first (fast, always completes), behavioral probe second with an internal timeout + a process-level watchdog so CI never hangs. Add a step to the relevant job in v3-ci.yml.scripts/audit-<name>.mjs that scans, counts violations, and fails if the count exceeds a monotone-decreasing baseline in verification/<name>-baseline.json. Support --update-baseline. Add a CI job; wire it into witness-verify needs[] if it should gate publish.{ id, desc, file, marker } to verification/witness-fixes.json, run node scripts/regen-witness.mjs. The marker must be a substring the fix specifically creates (not a generic pattern like 'function').npx ruflo@alpha)@noble/ed25519 (for the witness verifier — a single runtime dep, npm i @noble/ed25519)bash# View current truth scores npx ruflo@alpha truth # Run verification check npx ruflo@alpha verify check # Verify specific file with custom threshold npx ruflo@alpha verify check --file src/app.js --threshold 0.98 # Rollback last failed verification npx ruflo@alpha verify rollback --last-good
Display comprehensive quality and reliability metrics for your codebase and agent tasks.
Basic Usage:
bash# View current truth scores (default: table format) npx ruflo@alpha truth # View scores for specific time period npx ruflo@alpha truth --period 7d # View scores for specific agent npx ruflo@alpha truth --agent coder --period 24h # Find files/tasks below threshold npx ruflo@alpha truth --threshold 0.8
Output Formats:
bash# Table format (default) npx ruflo@alpha truth --format table # JSON for programmatic access npx ruflo@alpha truth --format json # CSV for spreadsheet analysis npx ruflo@alpha truth --format csv # HTML report with visualizations npx ruflo@alpha truth --format html --export report.html
Real-time Monitoring:
bash# Watch mode with live updates npx ruflo@alpha truth --watch # Export metrics automatically npx ruflo@alpha truth --export .claude-flow/metrics/truth-$(date +%Y%m%d).json
Example dashboard output:
📊 Truth Metrics Dashboard
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Overall Truth Score: 0.947 ✅
Trend: ↗️ +2.3% (7d)
Top Performers:
verification-agent 0.982 ⭐
code-analyzer 0.971 ⭐
test-generator 0.958 ✅
Needs Attention:
refactor-agent 0.821 ⚠️
docs-generator 0.794 ⚠️
Recent Tasks:
task-456 0.991 ✅ "Implement auth"
task-455 0.967 ✅ "Add tests"
task-454 0.743 ❌ "Refactor API"Truth Scores (0.0-1.0):
1.0-0.95: Excellent ⭐ (production-ready)0.94-0.85: Good ✅ (acceptable quality)0.84-0.75: Warning ⚠️ (needs attention)<0.75: Critical ❌ (requires immediate action)Trend Indicators:
Statistics:
Execute comprehensive verification checks on code, tasks, or agent outputs.
File Verification:
bash# Verify single file npx ruflo@alpha verify check --file src/app.js # Verify directory recursively npx ruflo@alpha verify check --directory src/ # Verify with auto-fix enabled npx ruflo@alpha verify check --file src/utils.js --auto-fix # Verify current working directory npx ruflo@alpha verify check
Task Verification:
bash# Verify specific task output npx ruflo@alpha verify check --task task-123 # Verify with custom threshold npx ruflo@alpha verify check --task task-456 --threshold 0.99 # Verbose output for debugging npx ruflo@alpha verify check --task task-789 --verbose
Batch Verification:
bash# Verify multiple files in parallel npx ruflo@alpha verify batch --files "*.js" --parallel # Verify with pattern matching npx ruflo@alpha verify batch --pattern "src/**/*.ts" # Integration test suite npx ruflo@alpha verify integration --test-suite full
The verification system evaluates:
bash# Get structured JSON output npx ruflo@alpha verify check --json > verification.json # Example JSON structure: { "overallScore": 0.947, "passed": true, "threshold": 0.95, "checks": [ { "name": "code-correctness", "score": 0.98, "passed": true }, { "name": "security", "score": 0.91, "passed": false, "issues": [...] } ] }
Automatically revert changes that fail verification checks.
Basic Rollback:
bash# Rollback to last known good state npx ruflo@alpha verify rollback --last-good # Rollback to specific commit npx ruflo@alpha verify rollback --to-commit abc123 # Interactive rollback with preview npx ruflo@alpha verify rollback --interactive
Smart Rollback:
bash# Rollback only failed files (preserve good changes) npx ruflo@alpha verify rollback --selective # Rollback with automatic backup npx ruflo@alpha verify rollback --backup-first # Dry-run mode (preview without executing) npx ruflo@alpha verify rollback --dry-run
Rollback Performance:
Create detailed verification reports with metrics and visualizations.
Report Formats:
bash# JSON report npx ruflo@alpha verify report --format json # HTML report with charts npx ruflo@alpha verify report --export metrics.html --format html # CSV for data analysis npx ruflo@alpha verify report --format csv --export metrics.csv # Markdown summary npx ruflo@alpha verify report --format markdown
Time-based Reports:
bash# Last 24 hours npx ruflo@alpha verify report --period 24h # Last 7 days npx ruflo@alpha verify report --period 7d # Last 30 days with trends npx ruflo@alpha verify report --period 30d --include-trends # Custom date range npx ruflo@alpha verify report --from 2025-01-01 --to 2025-01-31
Report Content:
Run interactive web-based verification dashboard with real-time updates.
bash# Launch dashboard on default port (3000) npx ruflo@alpha verify dashboard # Custom port npx ruflo@alpha verify dashboard --port 8080 # Export dashboard data npx ruflo@alpha verify dashboard --export # Dashboard with auto-refresh npx ruflo@alpha verify dashboard --refresh 5s
Dashboard Features:
Set verification preferences in .claude-flow/config.json:
json{ "verification": { "threshold": 0.95, "autoRollback": true, "gitIntegration": true, "hooks": { "preCommit": true, "preTask": true, "postEdit": true }, "checks": { "codeCorrectness": true, "security": true, "performance": true, "documentation": true, "bestPractices": true } }, "truth": { "defaultFormat": "table", "defaultPeriod": "24h", "warningThreshold": 0.85, "criticalThreshold": 0.75, "autoExport": { "enabled": true, "path": ".claude-flow/metrics/truth-daily.json" } } }
Adjust verification strictness:
bash# Strict mode (99% accuracy required) npx ruflo@alpha verify check --threshold 0.99 # Lenient mode (90% acceptable) npx ruflo@alpha verify check --threshold 0.90 # Set default threshold npx ruflo@alpha config set verification.threshold 0.98
Per-environment thresholds:
json{ "verification": { "thresholds": { "production": 0.99, "staging": 0.95, "development": 0.90 } } }
GitHub Actions:
yamlname: Quality Verification on: [push, pull_request] jobs: verify: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Install Dependencies run: npm install - name: Run Verification run: | npx ruflo@alpha verify check --json > verification.json - name: Check Truth Score run: | score=$(jq '.overallScore' verification.json) if (( $(echo "$score < 0.95" | bc -l) )); then echo "Truth score too low: $score" exit 1 fi - name: Upload Report uses: actions/upload-artifact@v3 with: name: verification-report path: verification.json
GitLab CI:
yamlverify: stage: test script: - npx ruflo@alpha verify check --threshold 0.95 --json > verification.json - | score=$(jq '.overallScore' verification.json) if [ $(echo "$score < 0.95" | bc) -eq 1 ]; then echo "Verification failed with score: $score" exit 1 fi artifacts: paths: - verification.json reports: junit: verification.json
Run verification automatically during swarm operations:
bash# Swarm with verification enabled npx ruflo@alpha swarm --verify --threshold 0.98 # Hive Mind with auto-rollback npx ruflo@alpha hive-mind --verify --rollback-on-fail # Training pipeline with verification npx ruflo@alpha train --verify --threshold 0.99
Enable real-time verification during collaborative development:
bash# Pair with verification npx ruflo@alpha pair --verify --real-time # Pair with custom threshold npx ruflo@alpha pair --verify --threshold 0.97 --auto-fix
Monitor codebase continuously during development:
bash# Watch directory for changes npx ruflo@alpha verify watch --directory src/ # Watch with auto-fix npx ruflo@alpha verify watch --directory src/ --auto-fix # Watch with notifications npx ruflo@alpha verify watch --notify --threshold 0.95
Send metrics to external monitoring systems:
bash# Export to Prometheus npx ruflo@alpha truth --format json | \ curl -X POST https://pushgateway.example.com/metrics/job/claude-flow \ -d @- # Send to DataDog npx ruflo@alpha verify report --format json | \ curl -X POST "https://api.datadoghq.com/api/v1/series?api_key=${DD_API_KEY}" \ -H "Content-Type: application/json" \ -d @- # Custom webhook npx ruflo@alpha truth --format json | \ curl -X POST https://metrics.example.com/api/truth \ -H "Content-Type: application/json" \ -d @-
Automatically verify before commits:
bash# Install pre-commit hook npx ruflo@alpha verify install-hook --pre-commit # .git/hooks/pre-commit example: #!/bin/bash npx ruflo@alpha verify check --threshold 0.95 --json > /tmp/verify.json score=$(jq '.overallScore' /tmp/verify.json) if (( $(echo "$score < 0.95" | bc -l) )); then echo "❌ Verification failed with score: $score" echo "Run 'npx ruflo@alpha verify check --verbose' for details" exit 1 fi echo "✅ Verification passed with score: $score"
Verification Speed:
Rollback Speed:
Dashboard Performance:
Low Truth Scores:
bash# Get detailed breakdown npx ruflo@alpha truth --verbose --threshold 0.0 # Check specific criteria npx ruflo@alpha verify check --verbose # View agent-specific issues npx ruflo@alpha truth --agent <agent-name> --format json
Rollback Failures:
bash# Check git status git status # View rollback history npx ruflo@alpha verify rollback --history # Manual rollback git reset --hard HEAD~1
Verification Timeouts:
bash# Increase timeout npx ruflo@alpha verify check --timeout 60s # Verify in batches npx ruflo@alpha verify batch --batch-size 10
Verification commands return standard exit codes:
0: Verification passed (score ≥ threshold)1: Verification failed (score < threshold)2: Error during verification (invalid input, system error)npx ruflo@alpha pair - Collaborative development with verificationnpx ruflo@alpha train - Training with verification feedbacknpx ruflo@alpha swarm - Multi-agent coordination with quality checksnpx ruflo@alpha report - Generate comprehensive project reports/docs/truth-scoring.md/docs/verification-criteria.md/examples/verification//docs/api/verification.md| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 9,043 | 6,726 | -26% | 1 | 1 | 0% | 1,914 | 6,793 | +255% | 0 | 0 | — |
case-02 | fail→pass | 14,651 | 13,756 | -6% | 1 | 1 | 0% | 1,856 | 7,002 | +277% | 0 | 0 | — |
case-03 | fail→pass | 4,144 | 4,708 | +14% | 1 | 1 | 0% | 366 | 6,145 | +1579% | 0 | 0 | — |
case-04 | fail→pass | 35,518 | 2,199 | -94% | 1 | 1 | 0% | 625 | 5,795 | +827% | 0 | 0 | — |
case-05 | fail→pass | 6,014 | 2,171 | -64% | 1 | 1 | 0% | 1,010 | 5,679 | +462% | 0 | 0 | — |
case-06 | fail→pass | 10,684 | 2,146 | -80% | 1 | 1 | 0% | 1,716 | 5,656 | +230% | 0 | 0 | — |
case-07 | fail→pass | 7,891 | 1,758 | -78% | 1 | 1 | 0% | 1,332 | 5,638 | +323% | 0 | 0 | — |
case-08 | fail→pass | 8,734 | 3,068 | -65% | 1 | 1 | 0% | 1,643 | 5,873 | +257% | 0 | 0 | — |
case-09 | fail→pass | 15,399 | 2,075 | -87% | 1 | 1 | 0% | 2,476 | 5,662 | +129% | 0 | 0 | — |
case-10 | fail→pass | 8,704 | 1,868 | -79% | 1 | 1 | 0% | 1,670 | 5,661 | +239% | 0 | 0 | — |
case-11 | fail→pass | 8,948 | 3,161 | -65% | 1 | 1 | 0% | 1,634 | 5,838 | +257% | 0 | 0 | — |
case-12 | fail→pass | 5,831 | 1,455 | -75% | 1 | 1 | 0% | 1,003 | 5,532 | +452% | 0 | 0 | — |
case-13 | fail→fail | 9,700 | 1,653 | -83% | 1 | 1 | 0% | 1,639 | 5,555 | +239% | 0 | 0 | — |
case-14 | fail→fail | 6,528 | 1,559 | -76% | 1 | 1 | 0% | 1,375 | 5,566 | +305% | 0 | 0 | — |
case-15 | fail→pass | 12,417 | 2,992 | -76% | 1 | 1 | 0% | 2,195 | 5,845 | +166% | 0 | 0 | — |
case-16 | fail→pass | 10,982 | 1,956 | -82% | 1 | 1 | 0% | 1,793 | 5,602 | +212% | 0 | 0 | — |
case-17 | fail→pass | 9,341 | 3,395 | -64% | 1 | 1 | 0% | 1,586 | 5,786 | +265% | 0 | 0 | — |
case-18 | fail→pass | 14,731 | 5,210 | -65% | 1 | 1 | 0% | 2,343 | 6,288 | +168% | 0 | 0 | — |
case-19 | fail→pass | 7,801 | 2,502 | -68% | 1 | 1 | 0% | 1,481 | 5,834 | +294% | 0 | 0 | — |
case-20 | fail→fail | 7,518 | 4,957 | -34% | 1 | 1 | 0% | 1,490 | 6,208 | +317% | 0 | 0 | — |
case-21 | fail→fail | 10,442 | 7,028 | -33% | 1 | 1 | 0% | 1,827 | 6,449 | +253% | 0 | 0 | — |
case-22 | fail→fail | 11,365 | 10,683 | -6% | 1 | 1 | 0% | 2,200 | 7,316 | +233% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +77 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.