Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build, rebuild, audit, and compare QQ mailbox invoice ground truth datasets for this repository using the existing truth-building scripts and evidence artifacts. Use when the task is to generate a QQ mailbox ground truth set, validate QQ batch output against truth, review QQ mailbox invoice evidence, or investigate QQ-specific invoice extraction pitfalls in this project.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -49% | 0% |
Use this project-level skill for QQ mailbox ground truth work in this repository.
Read references/workflow.md and references/pitfalls.md first. Read references/case-study-qq-20260201-20260311.md and references/project-artifacts.md when you need a validated example, current artifact paths, or prior evidence.
build_truth_dataset.py and audit_email_truth.py. Do not create a parallel truth-building flow unless the user explicitly asks for one.references/workflow.md for the build and validation sequence.references/pitfalls.md as the default debug checklist when counts or fields look wrong.references/case-study-qq-20260201-20260311.md only as a worked example and evidence sample.references/project-artifacts.md to find the current canonical manifests, reports, and diagnostics.truth_manifest.json for machine comparison.ground_truth_report.md for human review.pending_review_count = 0 before calling a dataset final.Other measured skills in the registry, with their headline benchmark lift.