Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Apply handwriting OCR to digitize historical and archival documents
.claude/skills/brycewang-stanford-handwriting-recognition-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 40% | 0% |
A skill for applying handwriting text recognition (HTR) to digitize historical documents, archival manuscripts, and handwritten research notes. Covers HTR platforms, image preprocessing, model training, post-correction, and integration into digital humanities research workflows.
Printed Text OCR:
- Characters are standardized and uniform
- Well-solved problem (>99% accuracy on clean scans)
- Tools: Tesseract, ABBYY FineReader, Adobe Acrobat
Handwriting Text Recognition (HTR):
- Characters vary by writer, mood, pen, era
- Much harder -- typically 85-95% character accuracy
- Requires training on specific handwriting styles
- Tools: Transkribus, Kraken, HTR-Flor, Google Cloud Vision
Challenges specific to historical documents:
- Faded ink, bleed-through, stains, tears
- Archaic letterforms and abbreviations
- Multiple hands in one document
- Non-standard orthography
- Mixed languages and scriptsPricing note: Transkribus uses a credit-based pricing model. A limited free tier is available, but processing large volumes of pages requires purchasing credits.
Transkribus is the leading platform for historical HTR.
Workflow:
1. Upload document images
2. Automatic layout analysis (detect text regions and baselines)
3. Manual correction of layout (if needed)
4. Apply a pre-trained HTR model (or train your own)
5. Review and correct transcription
6. Export as TEXT, PAGE XML, TEI, DOCX, or PDF
Pre-trained models:
- Noscemus GM (general model for Latin scripts)
- English Writing M1 (18th-19th century English)
- German Kurrent models
- Dutch, French, Italian, Spanish models available
Training a custom model:
- Requires ~15,000-25,000 words of ground truth (manually transcribed)
- Can start with a pre-trained base model and fine-tune
- Training takes 1-8 hours depending on dataset size| Tool | Type | Strengths | |------|------|----------| | Transkribus | Cloud platform | Best for historical documents, active community | | Kraken | Open source (Python) | Flexible, scriptable, custom training | | eScriptorium | Open source (web) | Based on Kraken, collaborative interface | | Google Cloud Vision | API | Good for modern handwriting, many languages | | Azure AI Vision | API | Competitive with Google for modern text | | HTR-Flor | Open source | Research-focused, PyTorch-based |
pythonfrom PIL import Image, ImageFilter, ImageEnhance def preprocess_document_image(image_path: str, output_path: str) -> dict: """ Preprocess a document scan for optimal HTR performance. Args: image_path: Path to the input scan output_path: Path to save the preprocessed image """ img = Image.open(image_path) # Convert to grayscale img = img.convert("L") # Enhance contrast enhancer = ImageEnhance.Contrast(img) img = enhancer.enhance(1.5) # Remove noise img = img.filter(ImageFilter.MedianFilter(size=3)) # Binarize (convert to black and white) threshold = 128 img = img.point(lambda x: 255 if x > threshold else 0, "1") img.save(output_path) return { "original": image_path, "processed": output_path, "steps_applied": [ "Grayscale conversion", "Contrast enhancement (1.5x)", "Median filter (noise removal)", "Binarization (threshold=128)" ], "additional_steps_if_needed": [ "Deskewing (correct rotation)", "Dewarping (correct page curvature)", "Bleed-through removal", "Background normalization" ] }
Resolution: 300-400 DPI for most documents
600 DPI for fine handwriting or damaged originals
Color: Grayscale usually sufficient; color for illuminated MSS
Format: TIFF (lossless) for archival; PNG for working copies
Lighting: Even, diffused light; avoid shadows and glare
Flatness: Use a book cradle or V-shaped scanner for bound volumes
Calibration: Include a color/grayscale chart for batch consistencypythondef post_correction_workflow(raw_transcription: str, dictionary: set, confidence_threshold: float = 0.8) -> dict: """ Post-correction strategy for HTR output. Args: raw_transcription: Raw OCR/HTR text output dictionary: Set of valid words for the document's language/period confidence_threshold: Below this, flag for manual review """ words = raw_transcription.split() flagged = [] corrected = [] for word in words: clean = word.strip(".,;:!?()[]") if clean.lower() in dictionary: corrected.append(word) else: flagged.append({ "word": word, "position": len(corrected), "suggestion": "Manual review needed" }) corrected.append(word) return { "total_words": len(words), "flagged_words": len(flagged), "estimated_accuracy": 1 - len(flagged) / max(len(words), 1), "flagged": flagged[:20], "correction_strategies": [ "Dictionary-based spell checking (period-appropriate dictionary)", "N-gram language model for context-aware correction", "Crowdsourcing (Zooniverse, FromThePage)", "Double-keying (two independent transcribers, compare)", "AI-assisted correction with human verification" ] }
1. Transcribe documents using HTR
2. Correct and validate transcriptions
3. Encode in TEI-XML for digital editions
4. Apply NLP for named entity recognition, topic modeling
5. Link entities to knowledge bases (Wikidata, VIAF)
6. Publish as a searchable digital archive
Tools for TEI encoding:
- oXygen XML Editor (standard for digital humanities)
- TEI Publisher (web-based publishing platform)
- FromThePage (collaborative transcription with TEI export)Report Character Error Rate (CER) and Word Error Rate (WER) on a held-out test set. CER below 5% is generally considered production-quality for historical documents. Always compare against a manually created ground truth. Report accuracy separately for different document types, hands, or time periods if your corpus is heterogeneous.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 22,941 | 40,991 | +79% | 1 | 1 | 0% | 3,447 | 3,442 | -0% | 0 | 0 | — |
case-02 | fail→fail | 53,186 | 12,660 | -76% | 1 | 1 | 0% | 3,833 | 3,756 | -2% | 0 | 0 | — |
case-03 | fail→fail | 27,338 | 55,169 | +102% | 1 | 1 | 0% | 4,468 | 6,482 | +45% | 0 | 0 | — |
case-04 | pass→fail | 14,865 | 16,253 | +9% | 1 | 1 | 0% | 2,169 | 3,741 | +72% | 0 | 0 | — |
case-05 | pass→pass | 15,366 | 16,100 | +5% | 1 | 1 | 0% | 2,467 | 4,209 | +71% | 0 | 0 | — |
case-06 | fail→pass | 15,754 | 12,554 | -20% | 1 | 1 | 0% | 2,258 | 3,667 | +62% | 0 | 0 | — |
case-07 | pass→pass | 19,734 | 30,980 | +57% | 1 | 1 | 0% | 2,617 | 4,253 | +63% | 0 | 0 | — |
case-08 | pass→pass | 17,529 | 12,594 | -28% | 1 | 1 | 0% | 1,967 | 3,419 | +74% | 0 | 0 | — |
case-09 | fail→pass | 17,870 | 10,433 | -42% | 1 | 1 | 0% | 2,541 | 3,516 | +38% | 0 | 0 | — |
case-10 | fail→pass | 16,198 | 15,053 | -7% | 1 | 1 | 0% | 2,869 | 3,923 | +37% | 0 | 0 | — |
case-11 | pass→pass | 19,006 | 15,669 | -18% | 1 | 1 | 0% | 2,738 | 3,917 | +43% | 0 | 0 | — |
case-12 | pass→pass | 12,969 | 9,707 | -25% | 1 | 1 | 0% | 2,346 | 3,030 | +29% | 0 | 0 | — |
case-13 | pass→pass | 14,905 | 15,137 | +2% | 1 | 1 | 0% | 2,801 | 4,527 | +62% | 0 | 0 | — |
case-14 | fail→pass | 18,023 | 14,562 | -19% | 1 | 1 | 0% | 2,867 | 4,025 | +40% | 0 | 0 | — |
case-15 | pass→pass | 18,318 | 22,627 | +24% | 1 | 1 | 0% | 2,562 | 5,091 | +99% | 0 | 0 | — |
case-16 | fail→pass | 10,809 | 10,795 | -0% | 1 | 1 | 0% | 1,817 | 3,485 | +92% | 0 | 0 | — |
case-17 | pass→pass | 12,494 | 7,880 | -37% | 1 | 1 | 0% | 1,802 | 3,008 | +67% | 0 | 0 | — |
case-18 | pass→pass | 17,697 | 14,814 | -16% | 1 | 1 | 0% | 2,961 | 4,249 | +43% | 0 | 0 | — |
case-19 | pass→pass | 11,861 | 9,837 | -17% | 1 | 1 | 0% | 2,006 | 3,300 | +65% | 0 | 0 | — |
case-20 | pass→pass | 16,446 | 16,853 | +2% | 1 | 1 | 0% | 2,792 | 4,534 | +62% | 0 | 0 | — |
case-21 | pass→pass | 27,663 | 21,508 | -22% | 1 | 1 | 0% | 3,951 | 5,204 | +32% | 0 | 0 | — |
case-22 | pass→pass | 19,977 | 20,881 | +5% | 1 | 1 | 0% | 3,427 | 5,621 | +64% | 0 | 0 | — |
case-23 | pass→pass | 18,380 | 15,998 | -13% | 1 | 1 | 0% | 2,579 | 4,452 | +73% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +22 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.