Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Verify research idea novelty against recent literature. Use when user says "查新", "novelty check", "有没有人做过", "check novelty", or wants to verify a research idea is novel before implementing.
.claude/skills/wanshuiyin-novelty-check/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✓→✗ | ▼ Worse | -13% | 0% |
| case-03 | ✓→✗ | ▼ Worse | -64% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -20% | 0% |
| case-15 | ✓→✗ | ▼ Worse | 5% | 0% |
| case-16 | ✓→✗ | ▼ Worse | -3% | 0% |
Check whether a proposed method/idea has already been done in the literature: $ARGUMENTS
gpt-6-astra — Model used via a secondary Codex agent. Must be an OpenAI model (e.g., gpt-6-astra, o3, gpt-4o)codex — Default: Codex xhigh reviewer. Use --reviewer: oracle-pro only when explicitly requested; if Oracle is unavailable, warn and fall back to Codex xhigh.Given a method description, systematically verify its novelty:
For EACH core claim, search using ALL available sources:
WebSearch):Call REVIEWER_MODEL via spawn_agent (spawn_agent) with xhigh reasoning:
reasoning_effort: xhighPrompt should include:
Copy this block verbatim into the reviewer's briefing; the report in Phase D is judged under it too.
=== NOVELTY VERDICT LIMITS (these bound how you judge, never how widely you search) ===
Search exhaustively; judge calibrated. Two failures waste months equally:
passing an idea a published paper already contains, and killing a viable idea
because the territory has neighbors.
1. Proximity is information, not a verdict. Someone working nearby goes in the
report; it is not by itself a reason to reject.
2. ABANDON has exactly one qualification: a specific published paper already
contains this result — name that paper. No named paper, no ABANDON.
3. Crowded-but-deltaed is PROCEED: state the delta in one sentence a reviewer
could verify. Thin or contested delta is PROCEED WITH CAUTION — say what
would make it carry, not why it should die. CAUTION is not a safe middle:
if you cannot name the specific thing that makes the delta thin, the
verdict is PROCEED.
4. Concurrent or competing work is not a veto. That is a race — report it and
let the user decide whether to run it.
5. A direct attack on a central problem is legitimate novelty when nobody has
executed it well. "This area is hot" does not mean "this area is taken."
6. This check is an early gate, never the last one — more triage, pilots, or
external review still stand between any idea and a paper, whatever order
this run uses. A wrongly passed idea dies cheaply at one of them; a wrongly
killed idea is never seen again. When torn between two verdicts, choose the
more permissive one.
Say plainly when an idea clears the check. Do not manufacture overlap.Output a structured report:
markdown## Novelty Check Report ### Proposed Method [1-2 sentence description] ### Core Claims 1. [Claim 1] — Closest: [paper] — What stays unknown or different: [delta] 2. [Claim 2] — Closest: [paper] — What stays unknown or different: [delta] ... ### Closest Prior Work | Paper | Year | Venue | Overlap | Key Difference | |-------|------|-------|---------|----------------| ### Overall Novelty Assessment - Score: X/10 (anchor: 5/10 = has clear neighbors but a defensible delta worth a pilot; reserve 1-3 for results a named published paper already contains) - Recommendation: PROCEED / PROCEED WITH CAUTION / ABANDON (per the verdict limits: crowded-but-deltaed ground is PROCEED; ABANDON must name the paper) - Key differentiator: [what makes this unique, if anything] - Risk: [what a reviewer would cite as prior work] ### Suggested Positioning [State the delta honestly in one sentence a reviewer could verify]
abandoned because the territory has neighbors. Be brutally honest in both directions — and when an idea clears the check, say so plainly.
individual claim rates LOW — judge the idea, not each claim in isolation. Known parts arranged to reveal something unknown are novel.
non-obvious interaction, failure mode, or insight. Judge the revelation, not the template.
After each spawn_agent or optional oracle-pro reviewer call, save the trace following ../shared-references/review-tracing.md. Write files directly to .aris/traces/novelty-check/<date>_run<NN>/ and record searched claims, closest papers, reviewer route, raw response, and final novelty decision. Respect the --- trace: parameter when present (default: full).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→fail | 17,848 | 21,816 | +22% | 1 | 1 | 0% | 1,909 | 2,319 | +21% | 0 | 0 | — |
case-14 | pass→fail | 23,362 | 20,281 | -13% | 1 | 1 | 0% | 2,432 | 2,126 | -13% | 0 | 0 | — |
case-01 | fail→fail | 37,342 | 19,311 | -48% | 1 | 1 | 0% | 5,005 | 2,053 | -59% | 0 | 0 | — |
case-02 | fail→fail | 36,257 | 19,397 | -47% | 1 | 1 | 0% | 4,764 | 2,068 | -57% | 0 | 0 | — |
case-03 | pass→fail | 45,302 | 18,487 | -59% | 1 | 1 | 0% | 5,976 | 2,131 | -64% | 0 | 0 | — |
case-04 | pass→fail | 25,309 | 20,833 | -18% | 1 | 1 | 0% | 2,863 | 2,297 | -20% | 0 | 0 | — |
case-05 | fail→fail | 25,749 | 24,666 | -4% | 1 | 1 | 0% | 3,171 | 2,197 | -31% | 0 | 0 | — |
case-06 | pass→pass | 18,455 | 24,584 | +33% | 1 | 1 | 0% | 2,045 | 4,478 | +119% | 0 | 0 | — |
case-07 | fail→fail | 22,360 | 20,303 | -9% | 1 | 1 | 0% | 2,595 | 2,180 | -16% | 0 | 0 | — |
case-08 | fail→fail | 28,450 | 20,217 | -29% | 1 | 1 | 0% | 3,761 | 2,152 | -43% | 0 | 0 | — |
case-09 | fail→fail | 20,071 | 18,847 | -6% | 1 | 1 | 0% | 2,127 | 2,118 | -0% | 0 | 0 | — |
case-10 | fail→fail | 30,260 | 21,604 | -29% | 1 | 1 | 0% | 4,170 | 2,061 | -51% | 0 | 0 | — |
case-11 | fail→fail | 45,400 | 14,474 | -68% | 1 | 1 | 0% | 6,083 | 1,981 | -67% | 0 | 0 | — |
case-12 | fail→fail | 34,948 | 18,376 | -47% | 1 | 1 | 0% | 4,494 | 2,080 | -54% | 0 | 0 | — |
case-15 | pass→fail | 20,228 | 20,049 | -1% | 1 | 1 | 0% | 2,186 | 2,287 | +5% | 0 | 0 | — |
case-16 | pass→fail | 18,579 | 20,622 | +11% | 1 | 1 | 0% | 2,092 | 2,032 | -3% | 0 | 0 | — |
case-17 | fail→fail | 25,221 | 20,436 | -19% | 1 | 1 | 0% | 3,118 | 2,233 | -28% | 0 | 0 | — |
case-18 | fail→fail | 24,599 | 18,126 | -26% | 1 | 1 | 0% | 3,150 | 1,926 | -39% | 0 | 0 | — |
case-19 | pass→pass | 20,633 | 12,201 | -41% | 1 | 1 | 0% | 2,187 | 2,590 | +18% | 0 | 0 | — |
case-20 | pass→pass | 18,252 | 26,911 | +47% | 1 | 1 | 0% | 2,077 | 5,138 | +147% | 0 | 0 | — |
case-21 | pass→fail | 22,372 | 26,614 | +19% | 1 | 1 | 0% | 2,534 | 3,198 | +26% | 0 | 0 | — |
case-22 | pass→fail | 23,380 | 22,072 | -6% | 1 | 1 | 0% | 3,036 | 2,265 | -25% | 0 | 0 | — |
case-23 | fail→fail | 6,983 | 11,305 | +62% | 1 | 1 | 0% | 299 | 2,532 | +747% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 4 counted toward the lift figure. The other 19 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. A headline lift is not published for this run.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/1/2026 | +17% |
| gemini-3.6-flash | verified | 8/11/2026 | +45% |
Other measured skills in the registry, with their headline benchmark lift.