Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Remove the discourse-level (structural) signs of AI writing that survive surface editing: stated lessons and moral-of-the-story closers, tidy single-track arcs, embodied-emotion performance ("chest tightened"), vague allusions instead of named references, unbroken linear structure, and shape convergence across pieces. Grounded in the StoryScope study (Russell et al. 2026): narrative structure alone detects AI text at 93.2% F1, and professional stylistic rewriting moved detection only 1.6 points.
.claude/skills/nulightjens-structural-humanizer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 66% | 0% |
Read this first. The humanizer skill fixes words: "delve", em dashes, rule of three, negative parallelism. This skill fixes what survives that pass: the structure. The two are different jobs, run in sequence. Surface pass first, structural pass second.
Why this layer matters more. StoryScope (Russell et al. 2026, arXiv:2604.03136) classified 61,608 stories from humans and 5 LLMs using only discourse-level features, with all style features withheld: 93.2% detection accuracy. Then they ran AI text through LAMP, a professional span-level rewriting framework that removes cliche, purple prose, and redundant exposition (functionally, a surface humanizer). Detection dropped 1.6 points. Meanwhile the surface layer is decaying on its own: GPT 5.4 already slashed em-dash usage, and fine-tuning drops stylistic detection from 97% to 3%. The durable fingerprint is structural, and fixing it requires structural rewrites, not word swaps. Full findings with numbers: references/storyscope-findings.md.
Do not replace one default with another. If every piece now opens mid-scene, names three feelings, and ends unresolved, that is a new detectable cluster. The study's deepest finding is convergence: all five AI models occupy one tight region of structural space while humans are dispersed and rare. Rarity IS the human signal.
So: pick 1-2 structural interventions per piece, vary them across pieces, and be able to say why this piece got this shape. Never apply the whole menu at once.
Run these one at a time (aspect-based checking found 95% of issues in the study's own pipeline vs 68% for one mega-pass). Numbers are human vs AI rates from the study.
AI states its lesson. Narrator explains the theme 77% of the time vs 52% for humans; themes are moralized ~20% harder; everything ties back to one central point. In content: the takeaway sentence, "What this means for you", the thesis restated at every section end, every example dutifully interpreted. Fix: state the point once, where it lands hardest. Cut every restatement. Let at least one example sit uninterpreted. Trust the reader.
AI writes single-track: unbroken causal chain, no subplots (79% vs 57%), everything resolved, protagonist-choice endings. Humans digress, loop, and leave threads open (thematically parallel tangents: 42% vs 21%; ambivalent endings far more common). Fix options: one tangent that only obliquely relates; one question raised and explicitly not answered; stop before the resolution.
The single largest gap in the study: AI performs emotion through the body and atmosphere 81% of the time vs 38% for humans ("chest tightened", "breath caught", "the lamplight dimmed"). Humans just name it: explicit emotion labels 29% vs 8%. Fix: say the feeling plainly ("honestly, it scared me", "I was pissed"). Reserve embodied detail for the one moment that earns it. Yes, this contradicts classic writing advice. Classic writing advice is now a machine signature.
Humans name real things: specific texts, people, brands, places, prices (explicit named references 47% vs 24%). AI stays at vague allusion (72% vs 50%) and avoids naming real brands or works. Fix: "a popular productivity book" becomes "Deep Work". "An expert" gets a name. "Recently" gets a date. Add the price, the version number, the city.
Humans acknowledge the reader (direct address 28% vs 7%; fourth-wall permeability 67% vs 39%). "AI writes as though no one is watching." Content marketing already uses "you" constantly, so the transferable move is acknowledging the writing itself: "I know how this sounds", "skip this section if you already run ads", "you're probably skimming, so here's the number". Use sparingly; it is a spice.
Does this piece have the same skeleton as your last three? Same opener type, same arc, same closer? That is the cluster forming. Compare against recent pieces and break the pattern before publishing.
lesson is stated (and how many times), time structure (linear or not), what gets resolved, tangent count, emotion moments and their mode, named vs vague references. Audit the outline, not the prose. (This is the study's own method: structural tells hide from prose-level reading.)
(see references/genre-calibration.md), different from the last piece.
just polish sentences; that is the other skill's job.
python3 scripts/structural_scan.py <file> catches the pattern-matchableslice (embodied-emotion cliches, takeaway markers, vague allusions, uniformity).
vary it.
in the close.
two-thirds through.
"Remember the $80/month from the top? That was the cheap part."
Do not tie it back explicitly.
Most drafts here come from Claude, whose fingerprint is the most distinctive of all five models. If the draft is Claude: flat event escalation (uniform intensity throughout; fix by varying stakes and energy across the piece), the epilogue habit (a wrap-up coda after the natural ending; cut it and end earlier), reverent, quiet endings (occasionally end on the spike or unresolved). GPT drafts over-index on distant retrospective framing ("years later, I realize") and social/gossip mechanics. Gemini produces the tidiest endings; kill the bow on top.
It does not fix vocabulary or punctuation (run humanizer). It does not impose a voice (that is jens-blog-writer / solo-scale-writer / the Nick Saraev templates). It does not make text undetectable; nothing does. And one honest caveat: StoryScope studied ~5,000-word fiction. The transfer to short nonfiction is an inference, but the core pattern (over-explanation, tidiness, linearity, convergence to one default shape) is exactly what independent analyses of nonfiction AI slop keep finding, and the short-text subset that transfers cleanly is audits 1, 3, 4, and 6.
distilled: all 30 core features with rates, fingerprints, robustness results, caveats.
and interventions apply per genre (LinkedIn / course lesson / blog / email).
the grep-able tells. Designed to later run as a hook.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 14,888 | 12,621 | -15% | 1 | 1 | 0% | 2,202 | 3,847 | +75% | 0 | 0 | — |
case-02 | pass→pass | 7,472 | 8,842 | +18% | 1 | 1 | 0% | 1,124 | 3,007 | +168% | 0 | 0 | — |
case-03 | pass→pass | 6,001 | 5,486 | -9% | 1 | 1 | 0% | 830 | 2,762 | +233% | 0 | 0 | — |
case-04 | fail→pass | 11,669 | 10,014 | -14% | 1 | 1 | 0% | 1,626 | 3,281 | +102% | 0 | 0 | — |
case-05 | fail→pass | 14,671 | 10,887 | -26% | 1 | 1 | 0% | 2,285 | 3,580 | +57% | 0 | 0 | — |
case-06 | pass→pass | 13,066 | 8,831 | -32% | 1 | 1 | 0% | 1,656 | 3,347 | +102% | 0 | 0 | — |
case-07 | pass→pass | 14,375 | 10,924 | -24% | 1 | 1 | 0% | 2,142 | 3,452 | +61% | 0 | 0 | — |
case-08 | fail→pass | 18,135 | 13,717 | -24% | 1 | 1 | 0% | 2,395 | 3,790 | +58% | 0 | 0 | — |
case-09 | pass→pass | 10,431 | 7,961 | -24% | 1 | 1 | 0% | 1,744 | 3,145 | +80% | 0 | 0 | — |
case-10 | fail→fail | 12,415 | 9,720 | -22% | 1 | 1 | 0% | 1,688 | 3,279 | +94% | 0 | 0 | — |
case-11 | fail→pass | 12,837 | 5,813 | -55% | 1 | 1 | 0% | 1,462 | 2,739 | +87% | 0 | 0 | — |
case-12 | fail→pass | 17,068 | 12,664 | -26% | 1 | 1 | 0% | 2,265 | 3,753 | +66% | 0 | 0 | — |
case-13 | fail→pass | 10,258 | 2,340 | -77% | 1 | 1 | 0% | 1,653 | 2,272 | +37% | 0 | 0 | — |
case-14 | fail→fail | 17,560 | 16,290 | -7% | 1 | 1 | 0% | 2,405 | 4,169 | +73% | 0 | 0 | — |
case-15 | pass→pass | 6,853 | 9,041 | +32% | 1 | 1 | 0% | 1,063 | 3,149 | +196% | 0 | 0 | — |
case-16 | fail→fail | 11,267 | 10,732 | -5% | 1 | 1 | 0% | 1,504 | 3,374 | +124% | 0 | 0 | — |
case-17 | fail→pass | 30,651 | 12,805 | -58% | 1 | 1 | 0% | 2,175 | 3,781 | +74% | 0 | 0 | — |
case-18 | fail→fail | 14,521 | 14,051 | -3% | 1 | 1 | 0% | 2,092 | 3,942 | +88% | 0 | 0 | — |
case-19 | pass→pass | 12,525 | 13,246 | +6% | 1 | 1 | 0% | 1,844 | 3,761 | +104% | 0 | 0 | — |
case-20 | fail→fail | 10,233 | 11,221 | +10% | 1 | 1 | 0% | 1,542 | 3,315 | +115% | 0 | 0 | — |
case-21 | fail→pass | 17,423 | 22,267 | +28% | 1 | 1 | 0% | 2,773 | 3,740 | +35% | 0 | 0 | — |
case-22 | pass→pass | 13,786 | 7,307 | -47% | 1 | 1 | 0% | 1,980 | 2,967 | +50% | 0 | 0 | — |
case-23 | pass→pass | 12,773 | 3,414 | -73% | 1 | 1 | 0% | 1,852 | 2,436 | +32% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +35 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.