Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. Distinct from public benchmarks (MMLU/HumanEval/GSM8K via lm-evaluation-harness) which measure GENERAL capability. Only a held-out domain set predicts whether THIS system works on YOUR data. Collect real examples, label, hold out (never train/prompt on it), size 50-200, version it, refresh on drift.
.claude/skills/agentsope-agentsop-domain-eval-set/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 144% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 268% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 288% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 753% | 0% |
> "Compiled program beats baseline on a held-out test set (not the val set used in optimization)." > — DSPy SOP exit criterion dspy.ai/learn/optimization/overview/]
> "Build the eval loop before optimizing anything. Every subsequent change must be gated on these numbers." > — LlamaIndex SOP Stage 2
This is an ENHANCE overlay skill. It produces one artifact — a versioned, sealed, human-labeled set of 50–200 examples drawn from your domain — that other skills consume: [[agentsop-regression-gate]] enforces it on every PR, [[agentsop-metric-design]] defines the scoring function applied to each example, and [[lm-evaluation-harness]] runs the complementary public-capability axis. The core claim: public benchmarks tell you the model is smart in general; only a held-out domain set tells you it works on your task. The latter is the one that predicts production.
Activate when any of these is true:
LLM/RAG/agent system and the only evidence is vibes, a demo, or a public benchmark number. You need a quantitative answer on the real distribution.
"92% on MMLU" or "passes HumanEval" to justify go-live. That measures general capability, not your task fit (AP-1). Force a domain set into the decision.
no domain test set exists yet to gate against. You must build the set before [[agentsop-regression-gate]] can do its job.
may be small while the domain gap is large, or vice versa. Only your held-out set tells you which.
(refresh, OP-DE06) or it never reflected the domain (rebuild from real traffic).
Do NOT activate for:
MMLU/GSM8K?" → that is [[lm-evaluation-harness]], not this skill.
ships. Don't build a benchmark for a script you'll delete tomorrow.
schema validity gives ≥95% of signal) — the "eval set" is just running the oracle; you don't need curated held-out examples. Don't gold-plate.
Two orthogonal axes, constantly confused:
| Axis | What it measures | Tool | Predicts production? | |---|---|---|---| | General capability | Reasoning, knowledge, coding in general, on shared public tasks | [[lm-evaluation-harness]] (MMLU, HumanEval, GSM8K, TruthfulQA) | No — a proxy at best | | Domain task fit | Whether the system answers your users on your data | this skill (held-out domain set) | Yes — this is the signal |
A model can score 90% on MMLU and 40% on your insurance-claims triage. A model can score below SOTA on HumanEval and be perfect at your internal codebase's patterns. The public number and the domain number are nearly uncorrelated once you're past a basic capability floor. The public bench is a sanity check; the domain set is the decision.
Three corollaries (each maps to an SOP stage):
(tickets, queries, logs, transactions), stratified, with edge cases pulled deliberately. Auto-generated QA pairs (LlamaIndex DatasetGenerator) are a fine bootstrap, but a model can ace generated questions and still fail real user phrasing. Generated sets do not replace a real held-out set (§7).
never pasted into a prompt as a few-shot demo, never used to pick chunk size or reranker, never in the fine-tune data. The moment it leaks, the number is inflated and meaningless (AP-2, OP-DE07). Per DSPy: the test set must be distinct from the val set used in optimization dspy.ai/learn/optimization/overview/].
training" dspy.ai/learn/optimization/overview/] and differences are noise. The set is small enough to label by hand and large enough to detect ~5–10pp regressions and to slice by segment.
0. Confirm activation (§1) — is the question "does this work on OUR data"?
1. COLLECT — sample real domain examples; stratify; pull edge cases (OP-DE01)
2. LABEL — gold answer / reference / pass-fail; 2 annotators on subset (OP-DE02)
3. HOLD OUT — split train/dev/test; SEAL the test split (OP-DE03)
4. SIZE — land at 50-200; per-segment counts (OP-DE04)
5. VERSION — hash + date + rubric; freeze as an artifact (OP-DE05)
6. LEAK-AUDIT — diff held-out vs demos / train / fine-tune data (OP-DE07)
7. PAIR — report alongside public bench; gate on the domain set (OP-DE08)
(later) REFRESH on domain shift (OP-DE06)Pull from where the real distribution lives: support tickets, search/query logs, user transcripts, transaction records, bug reports. Stratify so the set covers the production mix — by query type (lookup / summary / compare), by segment (tenant, language, product area), by difficulty. Then deliberately over-sample edge cases and known failures — the head of the distribution is easy; the tail is where systems break.
Target a raw pool ≥ 2× the final size (you'll drop ambiguous items in labeling). Record provenance and timestamp per example (needed later for drift refresh).
Exit: a candidate pool ≥ 2× target, with provenance, spanning the real mix.
Attach ground truth per example: a gold answer, an acceptable reference response (not "the unique correct" one for open-ended tasks — see [[agentsop-metric-design]]), or a pass/fail label. For RAG, also label the gold passage so RetrieverEvaluator(["mrr","hit_rate"]) can run LlamaIndex OP-10].
Have two annotators label a subset, measure agreement, resolve disagreements, and drop genuinely ambiguous items — an example two experts can't agree on will only add noise. Record the rubric. (This is the data-side analogue of DSPy's "human-validate the metric on ≥20 spot-checks" discipline DSPy Case C].)
Exit: labeled set with inter-annotator agreement noted, rubric recorded, ambiguous items logged as rejected.
Split into train / dev / test. The test (held-out) split is sealed:
Store it in a separate file/location with an access note. Per DSPy, the exit-gate test set must be "distinct from the val set used in optimization" dspy.ai/learn/optimization/overview/]. The dev split is what you tune against; the test split is the one number you trust at decision time.
Exit: sealed held-out test split + train/dev splits; access policy written.
segment (each slice needs its own ≥~30 to be meaningful).
dspy.ai/learn/optimization/overview/].
Size up (toward 200, or split into per-segment sets each ~50) when you need per-segment confidence. LlamaIndex's DatasetGenerator default of num=50 sits at the low end of this band — fine to bootstrap, then curate.
Freeze the set as a versioned artifact — eval_v1.jsonl plus a manifest with a content hash, creation date, and the labeling rubric. Score every model / prompt / retriever change against the same version; keep a results table keyed by (eval_version, system_version); bump only on a deliberate refresh, never silently. DSPy ships program.json as a versioned artifact dspy.ai/tutorials/saving/]; LlamaIndex versions indices as deployment artifacts (SOP Stage 5) — the eval set deserves the same rigor.
Before any release, and whenever few-shot demos or fine-tune data are assembled, diff the held-out set against (a) prompt few-shot demos, (b) fine-tune / training data, (c) the optimizer trainset. Any overlap = contamination → the held-out number is inflated and worthless (AP-2). Remove the overlap or rebuild the split — the same provenance discipline as [[agentsop-metric-design]]'s calibration receipt (OP-M10).
Run [[lm-evaluation-harness]] for the capability floor (sanity check: is the model fundamentally competent?). Run the domain held-out set for the decision. Report both side by side. If they disagree, the domain set wins the go/no-go. Hand the sealed set to [[agentsop-regression-gate]] to enforce on every subsequent PR.
Domains drift: new product line, new user segment, seasonal change. When held-out scores stop tracking production complaints, refresh (OP-DE06): add fresh real examples from recent traffic, retire stale ones, re-label edge cases production surfaced, bump the version, keep the old version for back-comparison. Cadence: quarterly or on any major domain change, whichever comes first. (This mirrors LlamaIndex's live-corpus reconciliation, A10.)
Each operation: Trigger → Action → Output Evidence]. Full Trigger/Action/ Output/Evidence form in intermediate/operation_candidates.json.
real inputs (logs/tickets/queries/transactions), stratify by type/segment/ difficulty, over-sample edge cases → raw pool ≥2× target with provenance. DSPy dev-set discipline; LlamaIndex OP-10 eval-from-corpus]
per item; two annotators on a subset, resolve disagreement, drop ambiguous, record rubric; for RAG label the gold passage → curated labeled set with agreement noted. DSPy Case C ≥20 spot-checks; LlamaIndex RetrieverEvaluator]
the test split (never to optimizer, never as few-shot demo, never to pick chunking/reranker/model, never in fine-tune data) → sealed test + train/dev. DSPy "held-out distinct from val"; Case A step 4]
100–200 = detect ~5–10pp regressions + per-segment slices; <30 = noise) → sized set with per-segment counts. DSPy "30 min, 200+ for MIPROv2"; LlamaIndex num=50]
eval_v1.jsonl + manifest(hash, date, rubric); score every change vs the same version; results keyed by (eval_version, system_version); bump only on deliberate refresh → versioned artifact. DSPy program.json versioning; LlamaIndex versioned indices]
→ add fresh recent-traffic examples, retire stale, re-label edge cases, bump version, keep old for comparison (quarterly or on major change) → new version + drift log. LlamaIndex live-corpus reconciliation A10]
diff held-out vs few-shot demos, fine-tune data, optimizer trainset; any overlap = contamination → remove or rebuild → leak-audit report (0 overlap). DSPy held-out-distinct rule; metric-design provenance OP-M10]
treat public bench as capability floor/sanity check, require the domain held-out set as the decision gate; report both, on disagreement the domain set wins → two-axis report gated on domain. [[lm-evaluation-harness]] covers public, not your domain]
困境: A team wants to ship a contract-review assistant. They have thousands of raw contracts but only ~25 examples a lawyer has labeled with gold answers. 25 < the 50 floor and well below the 30 "memorizing, not training" line dspy.ai/learn/optimization/overview/]. They're tempted to (a) skip the held-out set and ship on MMLU/legal-bench numbers, or (b) auto-generate 200 QA pairs with LlamaIndex DatasetGenerator and call that the held-out set.
约束: Lawyer labeling time is the bottleneck (~$$/hour, scarce). Public legal benchmarks exist but don't reflect this firm's contract templates. Auto-generated questions risk testing "what the corpus says" rather than "what real reviewers ask".
决策步骤:
capability floor, not proof the assistant handles these contracts.
DatasetGenerator (LlamaIndex Stage 2) gives a cheap dev set for iteration — but it is synthetic, so it cannot be the trusted held-out number (§7 caveat).
the lawyer label the 50 hardest real examples (stratified, edge-case-heavy, OP-DE01/02) rather than 200 easy generated ones. 50 real-labeled > 200 synthetic for the decision gate.
ambiguous ones rather than padding the count.
synthetic dev set; report the go/no-go on the 50 real held-out.
source of new labeled examples.
结果: A 50-example human-labeled, sealed held-out set built from the hardest real contracts predicts production far better than 200 synthetic questions or any public legal benchmark. The synthetic set still earns its keep — as the dev set you tune against, never as the number you trust.
可提取的操作: OP-DE01, OP-DE02, OP-DE03, OP-DE04. Lesson: spend scarce labels on a small REAL held-out set; let synthetic generation cover the dev set; never let a public bench be the gate.
困境: A support-triage classifier shows 0.91 on eval_v1 (built 9 months ago) and every PR passes [[agentsop-regression-gate]]. Yet production accuracy collapsed and users are escalating. The eval set says everything is fine.
约束: eval_v1 is versioned and trusted; nobody wants to "move the goalposts". The domain shifted — a new product line generates a third of current tickets, and none of those ticket types existed when eval_v1 was built. Rebuilding costs annotator time.
决策步骤:
compare against eval_v1's segment counts. The new product line is ~33% of live traffic and 0% of the eval set → the eval set no longer represents the domain. The green score is measuring an obsolete distribution.
stale. (Compare metric-design AP-8: changing the yardstick mid-stream without re-grounding.)
line and recent escalations — label them, retire ticket types that no longer occur, and build eval_v2.
eval_v1 for back-comparison. Re-score thecurrent system on eval_v2: it drops to 0.63 — now matching reality.
[[agentsop-regression-gate]] on eval_v2. Add a drift check to therefresh cadence: quarterly, compare live segment mix vs eval segment mix; if any segment drifts >X%, trigger a refresh.
结果: The "green-but-on-fire" gap was a stale held-out set, not a model regression. A versioned refresh (eval_v2) restored the eval as a true production predictor; the back-comparison against eval_v1 documented exactly how much the domain moved.
可提取的操作: OP-DE06 RefreshOnDomainShift, OP-DE05 VersionTheSet. Lesson: a held-out set is a snapshot of a moving distribution. Schedule drift checks; an old green score can be the most dangerous number you have.
| # | Anti-pattern | Why it's wrong | Fix | |---|---|---|---| | AP-1 | Public bench as proxy for domain performance ("92% MMLU → ship it") | Public benches measure general capability; near-uncorrelated with task fit past a floor | Build a domain held-out set; gate on it (OP-DE08) | | AP-2 | Eval set leaks into prompt / training / trainset | Held-out number is inflated and meaningless; you're testing on the train set | Seal it; leak-audit before release (OP-DE03, OP-DE07) | | AP-3 | Set too small to be significant (<30 examples) | "Memorizing, not training" dspy.ai/learn/optimization/overview/]; variance swamps signal | Target 50–200 (OP-DE04) | | AP-4 | Synthetic-only held-out (auto-generated QA is the test set) | Tests "what the corpus says", not real user phrasing; flatters the system | Synthetic = dev set bootstrap only; real-labeled = held-out (§7, Dilemma 1) | | AP-5 | Never refreshing as the domain drifts | Green scores on an obsolete distribution; "green but on fire" (Dilemma 2) | Schedule drift checks; refresh + version (OP-DE06) | | AP-6 | Unversioned set silently edited | Can't compare across system versions; results table is meaningless | Hash + date + rubric; bump on deliberate refresh (OP-DE05) | | AP-7 | No stratification / edge cases (only easy head-of-distribution) | Passes eval, fails the tail where systems actually break | Stratify by segment/type; over-sample edge cases (OP-DE01) | | AP-8 | Tuning chunk size / reranker / model against the held-out set | That makes it a val set, not held-out; the trust is gone | Tune on dev; touch held-out only at decision time (OP-DE03) |
model is best at reasoning?" → [[lm-evaluation-harness]] (MMLU/GSM8K/etc.), not this skill. This skill is for your task, not the leaderboard.
validity gives ≥95% of signal). The "eval set" is just running the oracle on inputs — you don't need curated human-labeled held-out examples. Don't gold-plate.
cost of building and labeling a real set has no payoff.
cold start). Bootstrap with synthetic + public benches transparently, label as soon as pilot traffic appears, and treat early numbers as provisional.
reference answer) — building the set is necessary but not sufficient. Pair with [[agentsop-metric-design]] to define a defensible, calibrated scoring function.
When does each kind of eval set apply? They are complementary axes, not substitutes — a mature pipeline uses all three.
| Concept | Held-out domain set (this skill) | [[lm-evaluation-harness]] (public) | LlamaIndex DatasetGenerator (synthetic) | |---|---|---|---| | What it measures | Task fit on your data | General capability | Coverage of your corpus's content | | Data source | Real traffic, human-labeled | Public academic datasets (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag) | LLM-generated QA from your docs | | Size | 50–200 | thousands (fixed by benchmark) | arbitrary (default num=50) | | Contamination risk | You control it (leak-audit) | High — public benches leak into pretraining | Low (your private corpus) but synthetic | | Predicts production? | Yes (the decision gate) | No (capability floor / sanity check) | Partially (dev-set iteration, not the gate) | | When to use | Go/no-go on shipping to your users; per-PR regression gate | Model selection on raw capability; academic reporting; training-progress tracking | Bootstrap a dev set fast before you've labeled real data | | Invocation | eval_vN.jsonl + scoring fn from [[agentsop-metric-design]] | lm_eval --tasks mmlu,gsm8k,... | DatasetGenerator.from_documents(docs).generate_dataset_from_nodes(num=50) |
Decision rubric:
Q1. Are you deciding whether to SHIP / SWITCH on YOUR users' data?
YES → held-out domain set is the gate (this skill). Public bench = sanity check only.
Q2. Are you comparing raw model capability or reporting academic numbers?
YES → lm-evaluation-harness (MMLU/HumanEval/GSM8K). Not this skill.
Q3. Do you have NO real labeled data yet but a corpus exists?
YES → DatasetGenerator to bootstrap a DEV set; label real held-out as soon as traffic appears.
Q4. Is there an objective oracle (tests/schema/exact-match)?
YES → run the oracle; no curated set needed.
DEFAULT → build + version a 50-200 real held-out set; gate via [[agentsop-regression-gate]];
score via [[agentsop-metric-design]]; pair with [[lm-evaluation-harness]] for the floor.Combination patterns:
[[agentsop-regression-gate]]: this skill produces the sealed,versioned set; regression-gate enforces it on every PR (chunking / embedding / prompt / model change). Division of labor: produce vs enforce.
[[agentsop-metric-design]]: this skill defines what's in the set;metric-design defines how each example is scored (decomposed sub-judges, bool-during-compile/float-during-eval, human-calibrated, length-penalized). A set with no defensible scoring function is half a benchmark.
[[lm-evaluation-harness]]: report both axes side by side(OP-DE08). Public bench answers "is the model competent?"; the domain set answers "does it work for us?". On disagreement, the domain set wins go/no-go.
"compiled program beats baseline on a held-out test set (not the val set)" dspy.ai/learn/optimization/overview/]. The DSPy trainset/valset come from the non-held-out splits.
Opinionated default: build the held-out set in plain jsonl (transparent, diffable, hashable), label it with humans on the hardest real examples, seal it, version it, and treat the public-benchmark number as a sanity check you report but never gate on.
references/R1-source-evidence.md — verbatim source quotes (DSPy held-outdiscipline, LlamaIndex eval-loop, lm-evaluation-harness public scope)
intermediate/operation_candidates.json — 8 operations in Trigger / Action /Output / Evidence form
Cross-links: [[lm-evaluation-harness]] (public-benchmark axis), [[agentsop-regression-gate]] (per-PR enforcement), [[agentsop-metric-design]] (scoring function).
Citations: dspy.ai/learn/optimization/overview/], dspy.ai/learn/optimization/optimizers/], dspy.ai/learn/evaluation/metrics/], dspy.ai/tutorials/saving/], developers.llamaindex.ai/python/framework-api-reference/evaluation/], llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5], ~/.claude/skills/lm-evaluation-harness/SKILL.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 36,249 | 23,134 | -36% | 1 | 1 | 0% | 6,005 | 10,225 | +70% | 0 | 0 | — |
case-02 | fail→fail | 25,946 | 18,398 | -29% | 1 | 1 | 0% | 4,370 | 9,971 | +128% | 0 | 0 | — |
case-03 | fail→pass | 25,866 | 20,439 | -21% | 1 | 1 | 0% | 4,146 | 10,119 | +144% | 0 | 0 | — |
case-04 | fail→pass | 23,444 | 11,171 | -52% | 1 | 1 | 0% | 3,740 | 8,544 | +128% | 0 | 0 | — |
case-05 | fail→pass | 15,170 | 10,811 | -29% | 1 | 1 | 0% | 2,303 | 8,467 | +268% | 0 | 0 | — |
case-06 | pass→pass | 9,548 | 5,528 | -42% | 1 | 1 | 0% | 1,471 | 7,648 | +420% | 0 | 0 | — |
case-07 | fail→pass | 14,971 | 13,235 | -12% | 1 | 1 | 0% | 2,271 | 8,810 | +288% | 0 | 0 | — |
case-08 | pass→pass | 13,101 | 12,244 | -7% | 1 | 1 | 0% | 2,068 | 8,781 | +325% | 0 | 0 | — |
case-09 | pass→pass | 10,381 | 9,483 | -9% | 1 | 1 | 0% | 1,697 | 8,306 | +389% | 0 | 0 | — |
case-10 | pass→pass | 13,654 | 13,103 | -4% | 1 | 1 | 0% | 2,137 | 8,916 | +317% | 0 | 0 | — |
case-11 | pass→pass | 16,926 | 12,377 | -27% | 1 | 1 | 0% | 2,485 | 8,806 | +254% | 0 | 0 | — |
case-12 | pass→pass | 16,953 | 13,943 | -18% | 1 | 1 | 0% | 2,631 | 9,103 | +246% | 0 | 0 | — |
case-13 | pass→pass | 16,416 | 13,855 | -16% | 1 | 1 | 0% | 2,561 | 9,041 | +253% | 0 | 0 | — |
case-14 | pass→pass | 13,025 | 10,366 | -20% | 1 | 1 | 0% | 1,892 | 8,385 | +343% | 0 | 0 | — |
case-15 | fail→pass | 20,983 | 9,736 | -54% | 1 | 1 | 0% | 986 | 8,413 | +753% | 0 | 0 | — |
case-16 | pass→pass | 15,808 | 16,067 | +2% | 1 | 1 | 0% | 2,307 | 9,035 | +292% | 0 | 0 | — |
case-17 | pass→pass | 11,850 | 8,145 | -31% | 1 | 1 | 0% | 1,900 | 8,156 | +329% | 0 | 0 | — |
case-18 | pass→pass | 14,755 | 12,920 | -12% | 1 | 1 | 0% | 2,164 | 8,794 | +306% | 0 | 0 | — |
case-19 | pass→pass | 12,977 | 8,039 | -38% | 1 | 1 | 0% | 2,098 | 8,093 | +286% | 0 | 0 | — |
case-20 | pass→pass | 12,588 | 9,465 | -25% | 1 | 1 | 0% | 1,915 | 8,349 | +336% | 0 | 0 | — |
case-21 | fail→pass | 35,666 | 13,678 | -62% | 1 | 1 | 0% | 1,115 | 9,208 | +726% | 0 | 0 | — |
case-22 | pass→pass | 11,242 | 11,480 | +2% | 1 | 1 | 0% | 1,832 | 8,712 | +376% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.