Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Split and read long documents chapter-by-chapter for structured analysis
.claude/skills/brycewang-stanford-large-document-reader/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 114% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 50% | 0% |
Split long documents (books, reports, theses, legal filings, technical manuals) into structured chapters or sections for systematic, chapter-by-chapter reading and analysis within LLM context windows.
Large Language Models have finite context windows, and even models with 100K+ token limits can lose accuracy on information buried in the middle of very long inputs. Academic researchers frequently work with documents that exceed practical context limits: doctoral theses (200+ pages), government reports, book-length monographs, legal case compilations, and multi-volume technical standards.
This skill provides a systematic approach to splitting large documents into semantically meaningful chapters or sections, maintaining cross-references between parts, and reading each section with full comprehension. Rather than naive fixed-size chunking that breaks mid-sentence or mid-argument, this approach respects document structure -- headings, chapter breaks, section markers, and logical boundaries.
The result is a structured reading experience where each chapter is analyzed in full context, summaries are maintained across sessions, and the reader can navigate directly to any section of interest. This is especially valuable for literature reviews, systematic reviews, and comprehensive document analysis tasks.
Documents should be split at the highest-level structural boundary that keeps each chunk within the target size:
| Priority | Boundary Type | Markers | |----------|--------------|---------| | 1 | Part/Volume | PART I, Volume 2, page breaks with Roman numerals | | 2 | Chapter | Chapter 1, CHAPTER, numbered headings level 1 | | 3 | Section | 1.1, Section, headings level 2 | | 4 | Subsection | 1.1.1, headings level 3 | | 5 | Paragraph break | Double newline, indentation change | | 6 | Sentence boundary | Period + space + capital letter |
pythondef split_document(text, max_tokens=8000, overlap_tokens=200): """Split document respecting structural boundaries.""" # Step 1: Detect document structure chapters = detect_chapters(text) if not chapters: # Fallback: split by sections chapters = detect_sections(text) if not chapters: # Fallback: split by paragraphs with size limit chapters = split_by_paragraphs(text, max_tokens) # Step 2: Merge small adjacent sections merged = merge_small_sections(chapters, min_tokens=500) # Step 3: Split oversized sections final = [] for chapter in merged: if count_tokens(chapter.text) > max_tokens: sub_parts = split_by_paragraphs(chapter.text, max_tokens) for i, part in enumerate(sub_parts): final.append(Section( title=f"{chapter.title} (Part {i+1})", text=part, index=len(final) )) else: chapter.index = len(final) final.append(chapter) # Step 4: Add overlap for continuity for i in range(1, len(final)): final[i].context_prefix = get_last_n_tokens( final[i-1].text, overlap_tokens ) return final
pythonimport re CHAPTER_PATTERNS = [ r'^#{1,2}\s+.+', # Markdown H1/H2 r'^Chapter\s+\d+', # "Chapter 1" r'^\d+\.\s+[A-Z]', # "1. Introduction" r'^PART\s+[IVX]+', # "PART III" r'^\\(chapter|section)\{', # LaTeX commands r'^\f', # Form feed (page break) ] def detect_chapters(text): sections = [] current_title = "Preamble" current_start = 0 for match in re.finditer('|'.join(CHAPTER_PATTERNS), text, re.MULTILINE): if match.start() > current_start: sections.append(Section( title=current_title, text=text[current_start:match.start()].strip() )) current_title = match.group().strip() current_start = match.start() sections.append(Section(title=current_title, text=text[current_start:].strip())) return sections
Read the table of contents, introduction, and conclusion first to build a mental model of the document's argument structure:
1. Extract and display Table of Contents
2. Read Introduction (typically Chapter 1)
3. Read Conclusion (typically last chapter)
4. Generate a document map: chapter titles + estimated page counts
5. Identify key themes and argumentsProcess each chapter with a standardized analysis template:
For each chapter:
- Chapter title and position in document
- Key arguments or findings (3-5 bullet points)
- Methodology described (if applicable)
- Data or evidence presented
- Connections to previous chapters
- Open questions or points for follow-up
- Notable quotes or passages (with page/section references)After all chapters are read, generate cross-cutting analyses:
- Thematic summary across all chapters
- Argument progression map
- Methodology comparison (if multiple studies)
- Contradiction or tension identification
- Gap analysis relative to research questionsFor documents that take multiple sessions to read, maintain a reading state file:
json{ "document": "thesis_smith_2024.pdf", "total_sections": 24, "completed": [0, 1, 2, 3, 4, 5], "current": 6, "summaries": { "0": "Preamble: Defines scope of study on...", "1": "Chapter 1: Introduction to the problem of...", "2": "Chapter 2: Literature review covering..." }, "themes": ["data governance", "algorithmic fairness", "institutional trust"], "open_questions": [ "How does the author reconcile findings in Ch3 with Ch5?" ] }
| Format | Tool | Notes | |--------|------|-------| | PDF | pdfplumber, PyMuPDF | Extract text with layout awareness | | EPUB | ebooklib | Chapters are HTML files in the spine | | DOCX | python-docx | Headings define structure | | LaTeX | Regex on \chapter, \section | Native structure markers | | HTML | BeautifulSoup | Split on <h1>, <h2> tags | | Plain text | Heuristic detection | Use blank lines, indentation, page breaks |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 22,941 | 50,888 | +122% | 1 | 1 | 0% | 3,856 | 4,757 | +23% | 0 | 0 | — |
case-02 | fail→pass | 18,306 | 45,257 | +147% | 1 | 1 | 0% | 2,853 | 4,062 | +42% | 0 | 0 | — |
case-03 | pass→pass | 13,047 | 9,849 | -25% | 1 | 1 | 0% | 1,933 | 3,731 | +93% | 0 | 0 | — |
case-04 | pass→pass | 15,518 | 9,351 | -40% | 1 | 1 | 0% | 2,211 | 3,429 | +55% | 0 | 0 | — |
case-05 | fail→fail | 13,762 | 14,028 | +2% | 1 | 1 | 0% | 2,522 | 3,923 | +56% | 0 | 0 | — |
case-06 | pass→pass | 6,495 | 4,874 | -25% | 1 | 1 | 0% | 940 | 2,426 | +158% | 0 | 0 | — |
case-07 | pass→pass | 8,455 | 4,446 | -47% | 1 | 1 | 0% | 1,421 | 2,539 | +79% | 0 | 0 | — |
case-08 | fail→pass | 15,137 | 15,710 | +4% | 1 | 1 | 0% | 2,629 | 4,019 | +53% | 0 | 0 | — |
case-09 | fail→fail | 16,122 | 39,873 | +147% | 1 | 1 | 0% | 2,103 | 3,403 | +62% | 0 | 0 | — |
case-10 | fail→pass | 12,225 | 2,395 | -80% | 1 | 1 | 0% | 2,001 | 2,263 | +13% | 0 | 0 | — |
case-11 | fail→pass | 14,428 | 14,227 | -1% | 1 | 1 | 0% | 2,036 | 4,350 | +114% | 0 | 0 | — |
case-12 | pass→pass | 14,476 | 11,894 | -18% | 1 | 1 | 0% | 2,202 | 3,433 | +56% | 0 | 0 | — |
case-13 | fail→fail | 7,767 | 16,760 | +116% | 1 | 1 | 0% | 1,100 | 2,302 | +109% | 0 | 0 | — |
case-14 | pass→fail | 8,877 | 2,283 | -74% | 1 | 1 | 0% | 1,419 | 2,174 | +53% | 0 | 0 | — |
case-15 | pass→pass | 11,925 | 12,727 | +7% | 1 | 1 | 0% | 1,902 | 3,617 | +90% | 0 | 0 | — |
case-16 | fail→pass | 18,521 | 16,360 | -12% | 1 | 1 | 0% | 2,958 | 4,441 | +50% | 0 | 0 | — |
case-17 | pass→pass | 15,457 | 13,700 | -11% | 1 | 1 | 0% | 2,459 | 4,002 | +63% | 0 | 0 | — |
case-18 | fail→pass | 14,155 | 10,622 | -25% | 1 | 1 | 0% | 1,938 | 3,680 | +90% | 0 | 0 | — |
case-19 | pass→pass | 11,630 | 11,201 | -4% | 1 | 1 | 0% | 1,982 | 3,445 | +74% | 0 | 0 | — |
case-20 | fail→fail | 1,843 | 2,301 | +25% | 1 | 1 | 0% | 249 | 2,185 | +778% | 0 | 0 | — |
case-21 | pass→pass | 3,057 | 3,673 | +20% | 1 | 1 | 0% | 654 | 2,416 | +269% | 0 | 0 | — |
case-22 | pass→pass | 3,661 | 8,134 | +122% | 1 | 1 | 0% | 667 | 3,327 | +399% | 0 | 0 | — |
case-23 | fail→pass | 14,811 | 2,284 | -85% | 1 | 1 | 0% | 2,020 | 2,312 | +14% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +26 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.