Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Read-only audit of `.tex`, `.qmd`, or `.md` text for AI-voice tells — boilerplate transitions ("Moreover", "Furthermore", "It is important to note that"), AI-cliché lexicon ("delve", "navigate the complexities", "tapestry", "robust framework"), em-dash overuse, symmetric paragraph shapes, tricolon abuse, hedging stacking, "not only X but also Y" frames, and formulaic openers. Produces a report; does NOT rewrite. Use when user says "humanize", "does this sound like AI?", "check for AI tells", "de
.claude/skills/pedrohcgs-humanize/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 203% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 113% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 171% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 218% | 0% |
/humanize — AI-voice audit (detect-and-flag)Read the target file (or all paper-like files), audit for the canonical AI-voice tells in academic prose, and write a structured report. The skill does not rewrite. The author edits.
Referees and editors increasingly recognise AI-generated prose. The tells are not stylistic preferences — they're statistically conspicuous patterns the LLM training distribution produces at higher rates than human academic writers. Five reasons to audit before submission:
--rewrite mode. Auto-rewriting AI tells degrades prose quality (cross-vendor research finding); the author preserves voice by editing manually./review-paper for argument structure, identification, citations./proofread for grammar, typos, overflow, citation format./verify-claims for Chain-of-Verification fact-checking of citations and numeric claims./humanize is the voice lens. Run it alongside the others — none of them substitute.
.bib, .R, or other non-prose files — the detectors are tuned for academic prose.The humanize-auditor agent checks these category groups:
High-confidence AI tells when they appear sentence-initial or mid-paragraph as connective tissue:
Moreover, / Furthermore, / Additionally, / In addition,It is important to note that / It is worth noting that / Notably,In conclusion, / In summary, / To summarise,On the other hand, (when not contrasting two named things)Building on this, / Building upon this,As we can see, / As is evident, / Indeed, (stacked)Severity: HIGH if more than 1 per 1000 words. MED if 1 per 2000 words. LOW if rare but present.
Words and phrases statistically over-represented in LLM output relative to academic prose:
Severity: HIGH on a paper's first three pages (abstract, intro). MED elsewhere.
Severity: MED. Em-dashes are a legitimate authorial choice; flag overuse, not all use.
Paragraphs with the same micro-architecture: topic sentence → three examples → summarising clause. Repeated across consecutive paragraphs is the AI tell — not the shape itself.
Detection: flag any three-paragraph window where each paragraph fits the topic→examples→summary cadence.
Severity: MED if 3-paragraph window; HIGH if 5+ paragraph stretch.
"X, Y, and Z" three-element lists are a legitimate rhetorical device. Tells are:
Severity: LOW if rare; MED if patterned.
Stacked epistemic hedges in single sentences:
Severity: HIGH — these are almost never authorial choices; they're LLM uncertainty-management.
Used sparingly, this is a legitimate construction. AI tells:
Severity: MED.
Severity: LOW unless every section starts this way.
Long chains of compound modifiers as a paragraph signature:
Severity: LOW.
Severity: HIGH — these read as AI-generated promotional copy; referees will react badly.
$ARGUMENTS starts with a filename: audit that file only.$ARGUMENTS is all: audit all .qmd, .tex, .md files in Slides/, Quarto/, root, and master_supporting_docs/..bib, .R, .py, code files, and any file under scripts/.--severity flag (default: report all).--severity low → report all findings.--severity med → suppress LOW findings.--severity high → report only HIGH findings.humanize-auditor agent with the 10 detection categories. line N | category | severity | current text | suggested rewrite or "remove"
quality_reports/audits/humanize_<filename>_report.md. Include:| Situation | Do | |---|---| | When you've drafted prose with AI assistance | Run /humanize before submission. Pair with /proofread (grammar) and /verify-claims (citations). | | When you wrote in your own voice | Run /humanize anyway — your own prose drifts toward LLM patterns after long sessions of AI-assisted work. | | Submission-ready review | /review-paper --peer [journal] --variance 3 for substance, /humanize for voice, /verify-claims for facts. |
--rewrite modeWe deliberately do not ship /humanize --rewrite. Cross-vendor research (Cursor / Aider community findings; cited in the v1.9.0 plan) finds that auto-rewriting prose to strip AI tells degrades quality more often than it improves it — the rewriter introduces its own AI tells. The detect-and-flag pattern preserves authorial voice; the cost is your editing time, which is exactly the cost we want to pay.
If you find yourself reaching for an auto-rewriter, that's the signal to rewrite the paragraph from scratch — not to patch the tells one by one.
quality_reports/audits/humanize_<filename>_report.md (that subdirectory is gitignored).If voice-profile.md exists at the repo root, read it first. A habit the author has declared deliberate — frequent em-dashes, first person, a particular connective — is not a finding. Flagging a documented preference as an AI tell is a false positive, and false positives erode the report's authority faster than misses do.
Build one with /voice-profile. This skill says what to remove; that one says what to write toward.
/humanize finds surface tells — boilerplate transitions, the AI-cliché lexicon, hedging stacks, symmetric paragraph shapes. Fixing them improves readability, which is worth doing whoever wrote the text.
It does not make prose stop reading as machine-generated to a detector. An article polished through several rounds of surface de-AI-ing was submitted to Pangram, a neural AI-text detector, and came back 100% AI-written. Those detectors classify on the token-level statistics of LLM generation, which survive any transformation the model applies — because every transformation is still LLM-generated text.
So: a clean report here means the prose reads well. It does not mean it reads human. If provenance matters, the author writes the load-bearing sentences and measures with a real detector. See writing-with-ai.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 64,574 | 65,044 | +1% | 1 | 1 | 0% | 379 | 3,016 | +696% | 0 | 0 | — |
case-02 | fail→fail | 4,793 | 37,551 | +683% | 1 | 1 | 0% | 317 | 3,006 | +848% | 0 | 0 | — |
case-03 | fail→fail | 3,503 | 6,096 | +74% | 1 | 1 | 0% | 329 | 3,100 | +842% | 0 | 0 | — |
case-04 | pass→pass | 2,551 | 10,332 | +305% | 1 | 1 | 0% | 352 | 4,472 | +1170% | 0 | 0 | — |
case-05 | fail→fail | 3,106 | 6,254 | +101% | 1 | 1 | 0% | 485 | 3,026 | +524% | 0 | 0 | — |
case-06 | fail→fail | 11,078 | 5,570 | -50% | 1 | 1 | 0% | 616 | 3,073 | +399% | 0 | 0 | — |
case-07 | pass→pass | 11,694 | 9,019 | -23% | 1 | 1 | 0% | 1,697 | 4,205 | +148% | 0 | 0 | — |
case-08 | pass→pass | 12,791 | 8,306 | -35% | 1 | 1 | 0% | 1,807 | 3,993 | +121% | 0 | 0 | — |
case-09 | pass→pass | 14,217 | 18,830 | +32% | 1 | 1 | 0% | 1,877 | 5,590 | +198% | 0 | 0 | — |
case-10 | pass→pass | 11,504 | 5,217 | -55% | 1 | 1 | 0% | 1,773 | 3,576 | +102% | 0 | 0 | — |
case-11 | fail→pass | 6,646 | 3,491 | -47% | 1 | 1 | 0% | 1,099 | 3,328 | +203% | 0 | 0 | — |
case-12 | fail→pass | 11,499 | 3,920 | -66% | 1 | 1 | 0% | 1,902 | 3,332 | +75% | 0 | 0 | — |
case-13 | pass→pass | 11,306 | 5,563 | -51% | 1 | 1 | 0% | 1,849 | 3,688 | +99% | 0 | 0 | — |
case-14 | fail→pass | 10,789 | 8,135 | -25% | 1 | 1 | 0% | 1,608 | 3,421 | +113% | 0 | 0 | — |
case-15 | fail→pass | 7,815 | 3,482 | -55% | 1 | 1 | 0% | 1,245 | 3,370 | +171% | 0 | 0 | — |
case-16 | fail→pass | 14,860 | 5,052 | -66% | 1 | 1 | 0% | 1,116 | 3,553 | +218% | 0 | 0 | — |
case-17 | fail→pass | 11,154 | 4,169 | -63% | 1 | 1 | 0% | 1,965 | 3,400 | +73% | 0 | 0 | — |
case-18 | fail→pass | 7,966 | 3,061 | -62% | 1 | 1 | 0% | 1,274 | 3,280 | +157% | 0 | 0 | — |
case-19 | fail→pass | 9,694 | 3,852 | -60% | 1 | 1 | 0% | 1,451 | 3,383 | +133% | 0 | 0 | — |
case-20 | fail→pass | 9,071 | 2,372 | -74% | 1 | 1 | 0% | 1,594 | 3,172 | +99% | 0 | 0 | — |
case-21 | fail→pass | 10,494 | 5,643 | -46% | 1 | 1 | 0% | 1,669 | 3,617 | +117% | 0 | 0 | — |
case-22 | pass→pass | 10,480 | 5,059 | -52% | 1 | 1 | 0% | 1,731 | 3,585 | +107% | 0 | 0 | — |
case-23 | fail→pass | 7,908 | 3,479 | -56% | 1 | 1 | 0% | 1,242 | 3,335 | +169% | 0 | 0 | — |
case-24 | pass→pass | 11,743 | 7,661 | -35% | 1 | 1 | 0% | 1,851 | 3,838 | +107% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 19 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +46 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +55% |
Other measured skills in the registry, with their headline benchmark lift.