Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop llm" or "llm review".
.claude/skills/wanshuiyin-auto-review-loop-llm/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 161% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 67% | 0% |
Autonomously iterate: review → implement fixes → re-review, until the external reviewer gives a positive assessment or MAX_ROUNDS is reached.
review-stage/AUTO_REVIEW.md (cumulative log) (fall back to `./AUTO_REVIEW.md` for legacy projects)This skill uses any OpenAI-compatible API for external review via the llm-chat MCP server.
Add to ~/.codex/settings.json:
json{ "mcpServers": { "llm-chat": { "command": "/usr/bin/python3", "args": ["/Users/yourname/.codex/mcp-servers/llm-chat/server.py"], "env": { "LLM_API_KEY": "your-api-key", "LLM_BASE_URL": "https://api.deepseek.com/v1", "LLM_MODEL": "deepseek-chat" } } } }
| Provider | LLM_BASE_URL | LLM_MODEL | |----------|--------------|-----------| | OpenAI | https://api.openai.com/v1 | gpt-4o, o3 | | DeepSeek | https://api.deepseek.com/v1 | deepseek-chat, deepseek-reasoner | | MiniMax | https://api.minimax.io/v1 | MiniMax-M3 | | Kimi (Moonshot) | https://api.moonshot.cn/v1 | moonshot-v1-8k, moonshot-v1-32k | | ZhiPu (GLM) | https://open.bigmodel.cn/api/paas/v4 | glm-4, glm-4-plus | | SiliconFlow | https://api.siliconflow.cn/v1 | Qwen/Qwen2.5-72B-Instruct | | 阿里云百炼 | https://dashscope.aliyuncs.com/compatible-mode/v1 | qwen-max | | 零一万物 | https://api.lingyiwanwu.com/v1 | yi-large |
Primary: MCP Tool
mcp__llm-chat__chat:
message: |
[Review prompt content]
model: "deepseek-chat"
system: "You are a senior ML reviewer..."Fallback: curl
bashcurl -s "${LLM_BASE_URL}/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${LLM_API_KEY}" \ -d '{ "model": "${LLM_MODEL}", "messages": [ {"role": "system", "content": "You are a senior ML reviewer..."}, {"role": "user", "content": "[review prompt]"} ], "max_tokens": 4096 }'
Persist state to review-stage/REVIEW_STATE.json after each round:
json{ "round": 2, "status": "in_progress", "last_score": 5.0, "last_verdict": "not ready", "pending_experiments": [], "timestamp": "2026-03-15T10:00:00" }
Write this file at the end of every Phase E (after documenting the round).
On completion, set "status": "completed".
review-stage/REVIEW_STATE.json for recovery (fall back to `./REVIEW_STATE.json` if not found — legacy path)If MCP available:
mcp__llm-chat__chat:
system: "You are a senior ML reviewer (NeurIPS/ICML level)."
message: |
[Round N/MAX_ROUNDS of autonomous review loop]
[Full research context: claims, methods, results, known weaknesses]
[Changes since last round, if any]
1. Score this work 1-10 for a top venue
2. List remaining critical weaknesses (ranked by severity)
3. For each weakness, specify the MINIMUM fix
4. State clearly: is this READY for submission? Yes/No/Almost
Be brutally honest. If the work is ready, say so clearly.If MCP NOT available:
bashcurl -s "${LLM_BASE_URL}/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${LLM_API_KEY}" \ -d '{ "model": "${LLM_MODEL}", "messages": [ {"role": "system", "content": "You are a senior ML reviewer (NeurIPS/ICML level)."}, {"role": "user", "content": "[Full review prompt]"} ], "max_tokens": 4096 }'
CRITICAL: Save the FULL raw response verbatim. Then extract:
STOP: If score >= 6 AND verdict ∈ {"ready", "almost"} (exact — "not ready" does NOT qualify)
Priority: metric additions > reframing > new experiments
Monitor remote experiments
Append to review-stage/AUTO_REVIEW.md:
markdown## Round N (timestamp) ### Assessment (Summary) - Score: X/10 - Verdict: [ready/almost/not ready] - Key criticisms: [bullet list] ### Reviewer Raw Response <details> <summary>Click to expand full reviewer response</summary> [Paste the COMPLETE raw response here — verbatim, unedited.] </details> ### Actions Taken - [what was implemented/changed] ### Results - [experiment outcomes, if any] ### Status - [continuing to round N+1 / stopping]
Write review-stage/REVIEW_STATE.json with current state.
review-stage/REVIEW_STATE.json status to "completed"cat << 'EOF' > file) to write in chunks. Do NOT ask the user for permission — just do it silently.mcp__llm-chat__chat:
system: "You are a senior ML reviewer (NeurIPS/ICML level)."
message: |
[Round N/MAX_ROUNDS of autonomous review loop]
## Previous Review Summary (Round N-1)
- Previous Score: X/10
- Previous Verdict: [ready/almost/not ready]
- Previous Key Weaknesses: [list]
## Changes Since Last Review
1. [Action 1]: [result]
2. [Action 2]: [result]
## Updated Results
[paste updated metrics/tables]
Please re-score and re-assess:
1. Score this work 1-10 for a top venue
2. List remaining critical weaknesses (ranked by severity)
3. For each weakness, specify the MINIMUM fix
4. State clearly: is this READY for submission? Yes/No/Almost
Be brutally honest. If the work is ready, say so clearly.> Follow these shared protocols for all output files: > - Output Versioning Protocol — write timestamped file first, then copy to fixed name > - Output Manifest Protocol — log every output to MANIFEST.md > - Output Language Protocol — respect the project's language setting
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | fail→pass | 10,543 | 4,120 | -61% | 1 | 1 | 0% | 2,065 | 2,889 | +40% | 0 | 0 | — |
case-01 | fail→fail | 6,030 | 4,534 | -25% | 1 | 1 | 0% | 228 | 2,364 | +937% | 0 | 0 | — |
case-16 | pass→pass | 12,698 | 6,414 | -49% | 1 | 1 | 0% | 2,111 | 3,332 | +58% | 0 | 0 | — |
case-21 | fail→pass | 7,600 | 5,555 | -27% | 1 | 1 | 0% | 1,155 | 3,019 | +161% | 0 | 0 | — |
case-02 | fail→fail | 13,731 | 4,938 | -64% | 1 | 1 | 0% | 2,727 | 2,333 | -14% | 0 | 0 | — |
case-03 | fail→fail | 23,918 | 4,259 | -82% | 1 | 1 | 0% | 5,668 | 2,529 | -55% | 0 | 0 | — |
case-04 | fail→fail | 3,484 | 16,367 | +370% | 1 | 1 | 0% | 520 | 2,837 | +446% | 0 | 0 | — |
case-05 | fail→fail | 2,546 | 5,193 | +104% | 1 | 1 | 0% | 255 | 2,358 | +825% | 0 | 0 | — |
case-06 | fail→fail | 9,106 | 93,053 | +922% | 1 | 1 | 0% | 1,656 | 2,413 | +46% | 0 | 0 | — |
case-07 | pass→pass | 8,930 | 2,848 | -68% | 1 | 1 | 0% | 1,537 | 2,634 | +71% | 0 | 0 | — |
case-08 | fail→pass | 13,623 | 7,628 | -44% | 1 | 1 | 0% | 2,081 | 3,574 | +72% | 0 | 0 | — |
case-09 | fail→pass | 9,358 | 7,565 | -19% | 1 | 1 | 0% | 1,758 | 3,397 | +93% | 0 | 0 | — |
case-10 | pass→pass | 7,550 | 2,690 | -64% | 1 | 1 | 0% | 1,266 | 2,644 | +109% | 0 | 0 | — |
case-12 | fail→pass | 8,687 | 2,636 | -70% | 1 | 1 | 0% | 1,529 | 2,554 | +67% | 0 | 0 | — |
case-13 | fail→pass | 13,184 | 2,139 | -84% | 1 | 1 | 0% | 2,323 | 2,464 | +6% | 0 | 0 | — |
case-14 | fail→pass | 7,988 | 2,146 | -73% | 1 | 1 | 0% | 1,268 | 2,452 | +93% | 0 | 0 | — |
case-15 | pass→pass | 12,266 | 3,132 | -74% | 1 | 1 | 0% | 1,837 | 2,695 | +47% | 0 | 0 | — |
case-17 | pass→pass | 13,162 | 3,177 | -76% | 1 | 1 | 0% | 1,439 | 2,595 | +80% | 0 | 0 | — |
case-18 | fail→fail | 15,979 | 1,802 | -89% | 1 | 1 | 0% | 1,385 | 2,367 | +71% | 0 | 0 | — |
case-19 | fail→pass | 9,733 | 2,306 | -76% | 1 | 1 | 0% | 1,608 | 2,466 | +53% | 0 | 0 | — |
case-20 | pass→pass | 9,662 | 4,274 | -56% | 1 | 1 | 0% | 1,528 | 2,764 | +81% | 0 | 0 | — |
case-22 | fail→pass | 6,783 | 1,877 | -72% | 1 | 1 | 0% | 1,199 | 2,494 | +108% | 0 | 0 | — |
case-23 | pass→pass | 8,017 | 4,665 | -42% | 1 | 1 | 0% | 1,211 | 2,949 | +144% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.