Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate the success criteria for a task or question, then review work against them. Given a task, goal, or open-ended question, decompose it into scenarios, evaluation perspectives, and fine-grained weighted YES/NO criteria using the Recursive Expansion Tree (RET) method; if work is supplied, score it criterion-by-criterion and surface what is missing or could be better. Use when asked to self-review or check your own work, judge whether a task is done well or completely, build a definition-of-
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 118% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 130% | 0% |
| case-09 | ✓→✓ | = Same ✓ | 122% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 53% | 0% |
Review the actual work the user means against the goal it was meant to satisfy. Default to a concise qualitative assessment. Scoring and the full Qworld Recursive Expansion Tree (RET) are opt-in.
"eval current work", that sentence is an instruction to review. The task being judged is the preceding user goal; the work is the current result, implementation, draft, plan, or progress.
eval with scoring. Eval, evaluate, review, assess, andcheck mean qualitative review unless the user explicitly asks for a score, grade, rating, points, weighted rubric, Qworld, RET, or a numeric scale.
evidence is unavailable, say what could not be verified.
and rubric construction internal unless the user asked for those artifacts.
checks from the request, stated constraints, acceptance criteria, and risks.
Separate three objects before reviewing:
examine.
Resolve the work product in this order:
about code changes.
Interpret deictic phrases such as "current work", "this work", "what we have", "刚才的 工作", and "当前工作" using that order. In a multi-turn conversation, the last user message is usually the evaluation instruction, not the original goal.
Proceed without asking when the goal and work can be recovered confidently. Ask one short clarifying question only when there is no reviewable work or when multiple plausible targets would produce materially different reviews.
| Mode | Trigger | Default output | |---|---|---| | Qualitative review (default) | "eval/review/check current work", "is this done?", "what is missing?" | Evidence-backed findings, strengths, gaps, fixes, and a completion verdict; no numbers | | Checklist | Explicit request for definition of done, success criteria, or completeness checklist without work to review | Task-specific checklist; no points | | Rubric | Explicit request for evaluation criteria or a rubric, but no scoring request | Binary or observable criteria grouped as must/should/could; no points | | Scored evaluation | Explicit request for score, grade, rating, points, weighted criteria, LLM-as-judge scoring, Qworld, RET, or a numeric scale | Evidence-backed scored review using the requested scale or RET |
The phrase "evaluate this" alone selects qualitative review, even when a work product is present. The presence of work never turns scoring on by itself.
Requests to create or run evals, implement a grader, write evaluation tests, or build a benchmark are engineering tasks, not self-review requests. Do not route those requests to this workflow merely because they contain the word eval.
two sentences. Prefer explicit acceptance criteria over inferred preferences.
supplied artifact. For code, inspect relevant implementation and verification evidence; do not judge from a summary alone when the files are available.
to judge correctness, completeness, user intent, risks, and verification. Do not print a large rubric unless asked.
each finding, cite the evidence and explain the consequence.
with generic praise.
important gaps.
complete, mostly complete, partiallycomplete, not complete, or unable to verify, with a short reason. Do not convert the verdict into a number.
Adapt the headings to the task and omit empty sections:
For code review, prioritize actionable defects and regressions over summaries. Cite file paths and tight line ranges when possible. If no problems are found, say so directly and name any residual verification gaps.
must, should, and could for importance when prioritization helps.met, partially met, not met, ornot verifiable, with brief evidence. These labels are not scores.
derivation.
Use this mode only when the request contains an unambiguous scoring signal listed above.
custom scheme, read and follow references/ret-scored-evaluation.md.
criteria outside the original goal.
number alone is not a useful review.
| User request | Correct interpretation | |---|---| | "Eval current work." | Review the current result against the preceding goal; qualitative, no score | | "Evaluate whether we finished the original request." | Inspect current artifacts and give a completion verdict; no score | | "What is missing from this implementation?" | Findings-first implementation review; no score | | "Make a definition-of-done checklist for this feature." | Checklist mode; no score | | "Create an evaluation rubric for these answers." | Unscored rubric unless weights or grading are requested | | "Score this answer from 1 to 10." | Scored mode on the requested scale | | "Apply Qworld/RET to grade these responses." | Full scored RET mode | | "Create an eval suite for this agent." | Out of scope for this skill; treat as an eval-engineering task |
files alone.
progress and label unfinished parts rather than pretending the final result exists.
framework dump.
Other measured skills in the registry, with their headline benchmark lift.