Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.
.claude/skills/lingxling-pdf/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✓→✓ | = Same ✓ | 419% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 58% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 88% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 162% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 146% | 0% |
This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read forms.md and follow its instructions.
pythonfrom pypdf import PdfReader, PdfWriter # Read a PDF reader = PdfReader("document.pdf") print(f"Pages: {len(reader.pages)}") # Extract text text = "" for page in reader.pages: text += page.extract_text()
pythonfrom pypdf import PdfWriter, PdfReader writer = PdfWriter() for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]: reader = PdfReader(pdf_file) for page in reader.pages: writer.add_page(page) with open("merged.pdf", "wb") as output: writer.write(output)
pythonreader = PdfReader("input.pdf") for i, page in enumerate(reader.pages): writer = PdfWriter() writer.add_page(page) with open(f"page_{i+1}.pdf", "wb") as output: writer.write(output)
pythonreader = PdfReader("document.pdf") meta = reader.metadata print(f"Title: {meta.title}") print(f"Author: {meta.author}") print(f"Subject: {meta.subject}") print(f"Creator: {meta.creator}")
pythonreader = PdfReader("input.pdf") writer = PdfWriter() page = reader.pages[0] page.rotate(90) # Rotate 90 degrees clockwise writer.add_page(page) with open("rotated.pdf", "wb") as output: writer.write(output)
pythonimport pdfplumber with pdfplumber.open("document.pdf") as pdf: for page in pdf.pages: text = page.extract_text() print(text)
pythonwith pdfplumber.open("document.pdf") as pdf: for i, page in enumerate(pdf.pages): tables = page.extract_tables() for j, table in enumerate(tables): print(f"Table {j+1} on page {i+1}:") for row in table: print(row)
pythonimport pandas as pd with pdfplumber.open("document.pdf") as pdf: all_tables = [] for page in pdf.pages: tables = page.extract_tables() for table in tables: if table: # Check if table is not empty df = pd.DataFrame(table[1:], columns=table[0]) all_tables.append(df) # Combine all tables if all_tables: combined_df = pd.concat(all_tables, ignore_index=True) combined_df.to_excel("extracted_tables.xlsx", index=False)
pythonfrom reportlab.lib.pagesizes import letter from reportlab.pdfgen import canvas c = canvas.Canvas("hello.pdf", pagesize=letter) width, height = letter # Add text c.drawString(100, height - 100, "Hello World!") c.drawString(100, height - 120, "This is a PDF created with reportlab") # Add a line c.line(100, height - 140, 400, height - 140) # Save c.save()
pythonfrom reportlab.lib.pagesizes import letter from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreak from reportlab.lib.styles import getSampleStyleSheet doc = SimpleDocTemplate("report.pdf", pagesize=letter) styles = getSampleStyleSheet() story = [] # Add content title = Paragraph("Report Title", styles['Title']) story.append(title) story.append(Spacer(1, 12)) body = Paragraph("This is the body of the report. " * 20, styles['Normal']) story.append(body) story.append(PageBreak()) # Page 2 story.append(Paragraph("Page 2", styles['Heading1'])) story.append(Paragraph("Content for page 2", styles['Normal'])) # Build PDF doc.build(story)
IMPORTANT: Never use Unicode subscript/superscript characters (₀₁₂₃₄₅₆₇₈₉, ⁰¹²³⁴⁵⁶⁷⁸⁹) in ReportLab PDFs. The built-in fonts do not include these glyphs, causing them to render as solid black boxes.
Instead, use ReportLab's XML markup tags in Paragraph objects:
pythonfrom reportlab.platypus import Paragraph from reportlab.lib.styles import getSampleStyleSheet styles = getSampleStyleSheet() # Subscripts: use <sub> tag chemical = Paragraph("H<sub>2</sub>O", styles['Normal']) # Superscripts: use <super> tag squared = Paragraph("x<super>2</super> + y<super>2</super>", styles['Normal'])
For canvas-drawn text (not Paragraph objects), manually adjust font the size and position rather than using Unicode subscripts/superscripts.
bash# Extract text pdftotext input.pdf output.txt # Extract text preserving layout pdftotext -layout input.pdf output.txt # Extract specific pages pdftotext -f 1 -l 5 input.pdf output.txt # Pages 1-5
bash# Merge PDFs qpdf --empty --pages file1.pdf file2.pdf -- merged.pdf # Split pages qpdf input.pdf --pages . 1-5 -- pages1-5.pdf qpdf input.pdf --pages . 6-10 -- pages6-10.pdf # Rotate pages qpdf input.pdf output.pdf --rotate=+90:1 # Rotate page 1 by 90 degrees # Remove password qpdf --password=mypassword --decrypt encrypted.pdf decrypted.pdf
bash# Merge pdftk file1.pdf file2.pdf cat output merged.pdf # Split pdftk input.pdf burst # Rotate pdftk input.pdf rotate 1east output rotated.pdf
python# Requires: pip install pytesseract pdf2image import pytesseract from pdf2image import convert_from_path # Convert PDF to images images = convert_from_path('scanned.pdf') # OCR each page text = "" for i, image in enumerate(images): text += f"Page {i+1}:\n" text += pytesseract.image_to_string(image) text += "\n\n" print(text)
pythonfrom pypdf import PdfReader, PdfWriter # Create watermark (or load existing) watermark = PdfReader("watermark.pdf").pages[0] # Apply to all pages reader = PdfReader("document.pdf") writer = PdfWriter() for page in reader.pages: page.merge_page(watermark) writer.add_page(page) with open("watermarked.pdf", "wb") as output: writer.write(output)
bash# Using pdfimages (poppler-utils) pdfimages -j input.pdf output_prefix # This extracts all images as output_prefix-000.jpg, output_prefix-001.jpg, etc.
pythonfrom pypdf import PdfReader, PdfWriter reader = PdfReader("input.pdf") writer = PdfWriter() for page in reader.pages: writer.add_page(page) # Add password writer.encrypt("userpassword", "ownerpassword") with open("encrypted.pdf", "wb") as output: writer.write(output)
| Task | Best Tool | Command/Code | |------|-----------|--------------| | Merge PDFs | pypdf | writer.add_page(page) | | Split PDFs | pypdf | One page per file | | Extract text | pdfplumber | page.extract_text() | | Extract tables | pdfplumber | page.extract_tables() | | Create PDFs | reportlab | Canvas or Platypus | | Command line merge | qpdf | qpdf --empty --pages ... | | OCR scanned PDFs | pytesseract | Convert to image first | | Fill PDF forms | pdf-lib or pypdf (see forms.md) | See forms.md |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 3,635 | 3,836 | +6% | 1 | 1 | 0% | 554 | 2,877 | +419% | 0 | 0 | — |
case-01 | fail→fail | 48,962 | 39,986 | -18% | 1 | 1 | 0% | 8,246 | 8,685 | +5% | 0 | 0 | — |
case-02 | pass→pass | 18,085 | 11,013 | -39% | 1 | 1 | 0% | 2,545 | 4,025 | +58% | 0 | 0 | — |
case-03 | pass→pass | 13,426 | 11,500 | -14% | 1 | 1 | 0% | 2,448 | 4,608 | +88% | 0 | 0 | — |
case-05 | pass→pass | 7,051 | 2,808 | -60% | 1 | 1 | 0% | 1,047 | 2,748 | +162% | 0 | 0 | — |
case-06 | pass→pass | 9,929 | 7,389 | -26% | 1 | 1 | 0% | 1,498 | 3,689 | +146% | 0 | 0 | — |
case-07 | fail→fail | 7,378 | 6,287 | -15% | 1 | 1 | 0% | 1,328 | 3,395 | +156% | 0 | 0 | — |
case-08 | pass→pass | 33,968 | 3,689 | -89% | 1 | 1 | 0% | 1,235 | 3,034 | +146% | 0 | 0 | — |
case-17 | pass→pass | 2,508 | 2,699 | +8% | 1 | 1 | 0% | 430 | 2,542 | +491% | 0 | 0 | — |
case-09 | pass→pass | 4,810 | 3,836 | -20% | 1 | 1 | 0% | 944 | 3,000 | +218% | 0 | 0 | — |
case-10 | pass→pass | 5,498 | 2,083 | -62% | 1 | 1 | 0% | 892 | 2,594 | +191% | 0 | 0 | — |
case-11 | pass→pass | 20,532 | 14,546 | -29% | 1 | 1 | 0% | 3,036 | 4,451 | +47% | 0 | 0 | — |
case-12 | pass→pass | 3,965 | 2,615 | -34% | 1 | 1 | 0% | 714 | 2,679 | +275% | 0 | 0 | — |
case-13 | pass→pass | 4,134 | 3,397 | -18% | 1 | 1 | 0% | 729 | 2,668 | +266% | 0 | 0 | — |
case-14 | pass→pass | 8,645 | 5,208 | -40% | 1 | 1 | 0% | 1,820 | 3,301 | +81% | 0 | 0 | — |
case-15 | pass→pass | 5,517 | 2,695 | -51% | 1 | 1 | 0% | 1,005 | 2,748 | +173% | 0 | 0 | — |
case-16 | pass→pass | 4,802 | 2,945 | -39% | 1 | 1 | 0% | 1,003 | 2,812 | +180% | 0 | 0 | — |
case-18 | fail→fail | 7,618 | 4,842 | -36% | 1 | 1 | 0% | 1,056 | 3,122 | +196% | 0 | 0 | — |
case-19 | fail→fail | 4,361 | 6,256 | +43% | 1 | 1 | 0% | 866 | 3,195 | +269% | 0 | 0 | — |
case-20 | pass→pass | 15,043 | 13,541 | -10% | 1 | 1 | 0% | 2,273 | 4,880 | +115% | 0 | 0 | — |
case-21 | pass→pass | 11,810 | 10,699 | -9% | 1 | 1 | 0% | 2,244 | 4,291 | +91% | 0 | 0 | — |
case-22 | pass→pass | 12,968 | 13,269 | +2% | 1 | 1 | 0% | 2,494 | 4,140 | +66% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of 0 percentage points is the difference between those two pass rates over the 22 comparable cases.
Other measured skills in the registry, with their headline benchmark lift.