Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when targeting Conference and Labs of the Evaluation Forum (CLEF) or deciding whether a computer-science manuscript fits this venue. Encodes conference fit, framing, evidence bar, submission-cycle checks, rebuttal posture, and desk-reject risks for evaluation forum.
.claude/skills/brycewang-stanford-conference-and-labs-of-the-evaluation-forum/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 63% | 0% |
Conference and Labs of the Evaluation Forum (CLEF) is a top computer-science conference venue for IR/NLP evaluation campaigns, shared tasks, multilingual retrieval, and benchmark methodology. It rewards a paper whose value comes from benchmark design, shared-task analysis, or rigorous evaluation. Treat this skill as a fit / venue-selection / re-framing tool for conference submission strategy, not as a substitute for the current year's CFP, author kit, ethics policy, or submission portal.
Because CS conferences change deadlines, templates, page limits, review workflow, artifact rules, AI-use policy, and rebuttal formats every cycle, always verify the live official instructions before making a submission-ready recommendation. Start from the official source anchor recorded for this venue in ../../resources/conference-roster.md and ../../resources/official-source-map.md.
forum and the author can say why CLEF reviewers are the primary audience, not merely a convenient deadline.
information-interaction-and-retrieval (CHIIR), acm-ieee-joint-conference-on-digital- libraries (JCDL), acm-conference-on-designing-interactive-systems (DIS), acm- interactive-surfaces-and-spaces (ISS). Break ties by contribution type, evidence shape, reviewer community, and the current official CFP from www.clef-initiative.eu.
If the paper misses CLEF's bar, compare against annual-meeting-of-the-association-for-computational-linguistics / conference-on-empirical-methods-in-natural-language-processing / north-american-chapter-of-the-association-for-computational-linguistics / european-chapter-of-the-association-for-computational-linguistics. Re-route based on contribution type, not prestige: theory to a theory venue, systems to a systems venue, application-heavy work to a domain venue, and early ideas to workshops or shorter tracks when the official CFP supports them.
text[Fit] High / Medium / Low (one-line reason) [Target] Conference and Labs of the Evaluation Forum (CLEF) [Contribution type] algorithm / theory / system / dataset / benchmark / empirical / design / security / other [Main evidence gap] <single most important missing proof, experiment, study, artifact, or policy check> [Official items to re-check] CFP / author kit / deadline / format / anonymity / ethics / AI-use / artifact / rebuttal / camera-ready [Top rejection risk] <venue-specific risk> [Re-route suggestion] <better-matched conference or journal if not a fit>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | pass→pass | 18,103 | 9,850 | -46% | 1 | 1 | 0% | 2,220 | 2,729 | +23% | 0 | 0 | — |
case-11 | fail→pass | 17,979 | 10,549 | -41% | 1 | 1 | 0% | 2,076 | 2,709 | +30% | 0 | 0 | — |
case-01 | fail→pass | 15,960 | 12,965 | -19% | 1 | 1 | 0% | 1,807 | 3,100 | +72% | 0 | 0 | — |
case-02 | fail→pass | 19,899 | 10,906 | -45% | 1 | 1 | 0% | 2,376 | 2,805 | +18% | 0 | 0 | — |
case-03 | fail→pass | 22,544 | 14,494 | -36% | 1 | 1 | 0% | 2,239 | 3,371 | +51% | 0 | 0 | — |
case-04 | pass→fail | 13,315 | 16,272 | +22% | 1 | 1 | 0% | 1,339 | 3,589 | +168% | 0 | 0 | — |
case-05 | pass→fail | 18,123 | 18,574 | +2% | 1 | 1 | 0% | 2,313 | 4,201 | +82% | 0 | 0 | — |
case-06 | pass→pass | 15,812 | 13,571 | -14% | 1 | 1 | 0% | 1,757 | 2,959 | +68% | 0 | 0 | — |
case-07 | fail→fail | 21,282 | 10,941 | -49% | 1 | 1 | 0% | 2,355 | 2,770 | +18% | 0 | 0 | — |
case-08 | pass→pass | 16,444 | 12,830 | -22% | 1 | 1 | 0% | 1,621 | 2,632 | +62% | 0 | 0 | — |
case-09 | fail→fail | 19,749 | 11,416 | -42% | 1 | 1 | 0% | 2,922 | 2,869 | -2% | 0 | 0 | — |
case-12 | fail→fail | 23,281 | 15,174 | -35% | 1 | 1 | 0% | 3,013 | 2,840 | -6% | 0 | 0 | — |
case-13 | pass→pass | 16,571 | 12,432 | -25% | 1 | 1 | 0% | 2,606 | 3,210 | +23% | 0 | 0 | — |
case-14 | fail→fail | 15,300 | 13,270 | -13% | 1 | 1 | 0% | 2,304 | 3,106 | +35% | 0 | 0 | — |
case-15 | pass→pass | 18,729 | 12,451 | -34% | 1 | 1 | 0% | 1,987 | 2,910 | +46% | 0 | 0 | — |
case-16 | fail→pass | 18,085 | 6,405 | -65% | 1 | 1 | 0% | 1,677 | 2,735 | +63% | 0 | 0 | — |
case-17 | pass→fail | 19,481 | 12,870 | -34% | 1 | 1 | 0% | 2,394 | 3,186 | +33% | 0 | 0 | — |
case-18 | pass→pass | 16,580 | 6,773 | -59% | 1 | 1 | 0% | 1,883 | 2,895 | +54% | 0 | 0 | — |
case-19 | pass→pass | 19,956 | 13,846 | -31% | 1 | 1 | 0% | 2,684 | 3,205 | +19% | 0 | 0 | — |
case-20 | pass→pass | 13,928 | 12,790 | -8% | 1 | 1 | 0% | 1,425 | 3,076 | +116% | 0 | 0 | — |
case-21 | pass→pass | 17,130 | 7,024 | -59% | 1 | 1 | 0% | 2,425 | 2,938 | +21% | 0 | 0 | — |
case-22 | fail→fail | 11,511 | 11,481 | -0% | 1 | 1 | 0% | 1,760 | 2,857 | +62% | 0 | 0 | — |
case-23 | pass→pass | 19,257 | 13,048 | -32% | 1 | 1 | 0% | 2,147 | 2,785 | +30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.