Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a draft, paper, or paper-like report is substantial enough for an independent skeptical audit before finalization, rebuttal, or revision routing.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | — | — |
| case-21 | ✗→✓ | ▲ Improved | — | — |
| case-15 | ✗→✓ | ▲ Improved | — | — |
| case-16 | ✗→✓ | ▲ Improved | — | — |
| case-08 | ✗→✓ | ▲ Improved | — | — |
Use this skill when the quest already has a substantial draft, paper, or paper-like report and now needs an independent, skeptical, evidence-grounded audit.
This is not the same as ordinary write. It is also not the same as rebuttal.
write turns accepted evidence into a narrative.review audits that narrative like a harsh but constructive expert reviewer.rebuttal responds to concrete external reviewer pressure that already exists.artifact.interact(kind='milestone', reply_mode='threaded', ...) update that says what the main risks are, what should be fixed next, and whether the next route is writing, experiment, or claim downgrade.bash_exec.review is an auxiliary audit skill for paper-like deliverables.
It should convert “the draft feels almost done” into a durable, skeptical, technically grounded review workflow:
write, analysis-campaign, baseline, scout, or decisionDefault review stance: independent audit before celebration. Do not treat “looks polished” as “is defensible”.
paper/draft.md, report draft, or paper-like manuscript already existsrebuttalstartup_contract.review_followup_policy is present, honor it:audit_onlyauto_execute_followupsuser_gated_followupsstartup_contract.manuscript_edit_mode = latex_required, treat the provided LaTeX tree or paper/latex/ as the writing surface when manuscript revision is needed.latex_required is requested, do not pretend the manuscript was edited; produce LaTeX-ready replacement text and an explicit blocker note instead.Use, in roughly this order:
evaluation_summary blocks from recent main experiments and analysis slicesIf the draft/result state is still unclear, open intake-audit first before continuing the review workflow. Before proposing extra experiments, read those structured evaluation_summary blocks first so you do not request work that the recorded evidence already resolved. If the user provided draft files or manuscript bundles directly, first normalize them into durable quest-visible paths before planning experiments or section-level revisions.
The review pass should usually leave behind:
paper/review/review.mdpaper/review/revision_log.mdpaper/review/experiment_todo.mdpaper/paper_experiment_matrix.md when more evidence is still neededpaper/paper_experiment_matrix.json when more evidence is still neededUse the templates in references/ when needed:
review-report-template.mdrevision-log-template.mdexperiment-todo-template.mdAudit at least these dimensions:
Before writing the review itself, make the audit explicit.
Identify:
C1, C2, C3If novelty, related-work coverage, or field positioning is unclear:
scoutDo not request new experiments just to answer a literature-positioning question.
Write paper/review/review.md using references/review-report-template.md.
The review should be:
At minimum, the review report should cover:
If helpful, include an internal conservative overall judgment or score, but do not pretend numerical precision when evidence is still unstable.
Write paper/review/revision_log.md using references/revision-log-template.md.
For each serious issue, record:
finalizestartup_contract.manuscript_edit_mode = latex_requiredOnly if more evidence is truly needed, write paper/review/experiment_todo.md using references/experiment-todo-template.md.
When the paper still lacks experimental support, also create or revise:
paper/paper_experiment_matrix.mdpaper/paper_experiment_matrix.jsonTreat the matrix as the paper-facing master plan and paper/review/experiment_todo.md as only the current execution frontier or review-facing subset.
Each TODO item should include:
exp_id in the paper experiment matrixDo not write a vague “run more ablations” list. Each TODO item should be concrete enough to turn into analysis-campaign slices or a baseline recovery task. The matrix should be broader than the TODO list and should classify the full paper-facing experiment space, not just analysis work. When building or revising that matrix, explicitly consider:
Do not assume the paper only needs “analysis experiments”. Do not assume case studies belong in the required set. If efficiency or cost could become a reviewer-facing strength or concern, put that into the matrix explicitly.
For the matrix, each row should usually record:
exp_idtierexperiment_typestatusfeasibility_nowclaim_idshighlight_idsresearch_questionhypothesiscomparatorsmetricsminimal_success_criterionpaper_placementpromotion_rulenext_actionThe matrix should also keep a short highlight hypotheses block. Do not rely on prose intuition for the method's best selling point; if a likely highlight matters, it should have a corresponding validation row in the matrix.
Before treating the experiments section as stable, require that every currently feasible matrix row that is not merely optional or dropped is either:
When extra evidence is truly needed, use the shared supplementary-experiment protocol:
artifact.create_analysis_campaign(...)artifact.record_analysis_slice(...)Do not invent a separate review-only experiment workflow.
After the review artifacts are durable:
writescoutbaselineanalysis-campaigndecisionDo not stop immediately after writing the review if the next route is already clear.
When startup_contract.review_followup_policy = auto_execute_followups:
analysis-campaignbaselinewritepaper/review/revision_log.mdpaper/review/experiment_todo.mdWhen startup_contract.review_followup_policy = user_gated_followups:
When startup_contract.review_followup_policy = audit_only:
If manuscript revision is required, make the delta explicit:
If startup_contract.manuscript_edit_mode = copy_ready_text:
paper/review/revision_log.md or a nearby revision notewriteIf startup_contract.manuscript_edit_mode = latex_required:
Open additional skills only when the review workflow requires them:
intake-auditscoutbaselineanalysis-campaignwritefigure-polishdecisionUse these tools deliberately:
artifact.record(payload={'kind': 'decision', ...})artifact.create_analysis_campaign(...)artifact.record_analysis_slice(...)artifact.submit_paper_outline(mode='revise', ...)artifact.submit_paper_bundle(...)artifact.interact(...)Stage-start requirement:
memory.list_recent(scope='quest', limit=5)memory.search(...) for:Stage-end requirement:
memory.write(...)Useful tags include:
stage:reviewtype:paper-reviewtype:revision-plantype:experiment-gaptype:claim-downgradereview is successful when:
write, analysis-campaign, baseline, scout, or finalize without ambiguityThe goal is not to sound severe. The goal is to make the next revision step technically clear and evidence-bound.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 12 counted toward the lift figure. The other 10 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 12 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.