Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use PathML for local, research-only computational pathology workflows: load and tile slides, build preprocessing and QC pipelines, manage h5path data, quantify multiplex images, construct spatial graphs, and plan bounded model inference.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 225% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 123% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 82% | 0% |
Use PathML for local computational pathology research. It is beta research software, not a validated medical device, diagnostic system, clinical decision support tool, or substitute for a pathologist. Do not use outputs to diagnose, grade, stage, or treat a patient.
Pathology files may contain faces, labels, accession numbers, patient identifiers, DICOM tags, filenames, or linked clinical data. Before processing:
analysis workspace.
patient_id, slide_id, and specimen_id values. Do not putdirect identifiers in filenames, logs, .h5path labels, model cards, or reports.
pathml==3.0.5, published 2026-03-24.PyPI does not declare Requires-Python and still has a stale 3.8 classifier, so use the release statement and test the exact environment.
no artifacts for them as of this review. v3.0.7 updates Torch/TorchVision/ torch-geometric and ONNX export code. Do not mix those source dependencies with the 3.0.5 wheel.
/latest identifies itself as 3.0.5. Examples here were checkedagainst the v3.0.5 tag and PyPI wheel metadata, not unversioned snippets.
licensing options; review upstream terms before redistribution.
Use Python 3.11 unless the project has tested another supported interpreter:
bashuv venv --python 3.11 source .venv/bin/activate uv pip install "pathml==3.0.5" python -c "import importlib.metadata as m; print(m.version('pathml'))"
PathML 3.0.5 declares no package extras: do not use pathml[all]. Its base distribution pins a large scientific/ML stack, including Torch 2.8.0, ONNX 1.17.0, ONNX Runtime 1.17.x, OpenSlide Python 1.3.1, python-bioformats 4.1.0, and python-javabridge 4.0.4.
Install native prerequisites before the uv command:
bash# Debian/Ubuntu sudo apt-get install openslide-tools gcc g++ libblas-dev liblapack-dev openjdk-17-jdk # macOS brew install openslide openjdk@17 # Windows OpenSlide option documented upstream vcpkg install openslide
Java/Bio-Formats is needed for the broad multidimensional format backend. OpenSlide handles common brightfield WSI formats more efficiently. CUDA is optional and must match the pinned PyTorch build; follow PyTorch's platform selector rather than guessing a CUDA wheel. See references/image_loading.md.
PathML 3.0.5 uses slide convenience classes and SlideData.run(). It does not provide SlideData.from_slide(), and Pipeline does not have run():
pythonfrom pathml.core import HESlide from pathml.preprocessing import BoxBlur, Pipeline, TissueDetectionHE slide = HESlide("data/pseudonymous_slide.svs", backend="openslide") pipeline = Pipeline( [ BoxBlur(kernel_size=5), TissueDetectionHE(mask_name="tissue", min_region_size=5000), ] ) slide.run( pipeline, distributed=False, tile_size=512, tile_stride=512, level=0, tile_pad=False, ) slide.write("derived/pseudonymous_slide.h5path")
Start with a bounded manual sample before a full run:
pythonfrom itertools import islice for tile in islice(slide.generate_tiles(shape=512, stride=512, level=0), 8): pipeline.apply(tile) assert tile.masks["tissue"].shape[:2] == tile.image.shape[:2]
Tiles use (i, j) = (row, column) coordinates at the selected pyramid level. For OpenSlide, PathML maps them to level-0 coordinates internally. Record the level and downsample; convert to (x, y) or micrometres explicitly downstream.
allowlisted technical metadata, and remove identifiers.
generating overlapping tiles, graphs, normalization references, or features.
stain behavior, edge padding, and empty-mask cases on representative training slides. Do not tune from test slides.
(i, j), downsample, MPP,mask names, QC decisions, and failed/skipped tiles.
instance labels, node-feature alignment, graph edges, and cell-to-tissue assignments.
loading unknown pickle checkpoints. Keep predictions linked to slide/tile coordinates and stitch overlaps with a documented rule.
stain, parameters, seeds, split manifest, model card, exclusions, and QC.
Do not instantiate download-capable classes or set dataset download=True unless the user explicitly opts in after receiving the endpoint and disclosure:
SegmentMIFRemote downloads an ONNX file fromhttps://huggingface.co/pathml/test/resolve/main/mesmer.onnx at construction, then runs inference locally. Stable source does not upload image pixels. The request still discloses network metadata such as IP address and headers and creates temp.onnx; there is no built-in checksum or offline flag.
SegmentMIF imports local DeepCell Mesmer, but DeepCell modelinitialization may need separately provisioned weights. It is not a PathML extra and is not the preferred stable API.
RemoteTestHoverNet downloads a model from Hugging Face.PanNukeDataModule(download=True) contacts Warwick; DeepFocusDataModulecontacts Zenodo. Both default to download=False.
Before any future hosted prediction call, state the exact destination, pixel channels/regions, metadata, identifiers, retention, legal basis, and safeguards; obtain explicit consent; and never send PHI by default. Prefer reviewed, checksummed local model artifacts and local inference.
model.eval() means evaluation mode for modules; it is not Python'sdangerous built-in evaluator. Never use Python dynamic evaluation or execution.
pathml.py, torch.py, onnx.py, or after standardlibraries; shadow modules can silently change imports.
EntityDataset loads .pt objects with weights_only=False. Neveropen an untrusted graph/checkpoint. Treat pickle-based pipelines and .pt files as executable code.
expected input/output schema, file size, and runtime limits; use isolation for third-party models.
All helpers reject URLs and symlinks, cap inputs/work, use strict JSON, avoid network access, and require no PathML import for --help:
bashpython scripts/slide_manifest.py validate --manifest manifest.csv --root . python scripts/slide_manifest.py inspect --slide data/example.svs --root . python scripts/plan_pipeline.py --width 100000 --height 80000 --tile-size 512 --stride 512 python scripts/image_qc.py synthetic --width 256 --height 256 python scripts/validate_spatial_schema.py graph --input graph.json --root . python scripts/validate_spatial_schema.py multiplex --input cells.csv --root . python scripts/plan_inference.py --tile-count 4000 --batch-size 16 --height 256 --width 256
The inference planner reads numbers or a bounded JSON model card only; it never imports a model framework or opens a checkpoint.
references/image_loading.md — slide classes, backends, formats, levels,coordinates, technical metadata, and privacy.
references/preprocessing.md — stable transforms, masks/QC, stain processing,pipeline execution, and leakage prevention.
references/data_management.md — .h5path, manifests, datasets, provenance,splits, and safe downloads.
references/multiparametric.md — multidimensional layout, CODEX/Vectra,quantification, AnnData, DeepCell/Mesmer, and network disclosure.
references/graphs.md — instance maps, feature alignment, KNN/RAG/HACT graphs,spatial units, schemas, and validation.
references/machine_learning.md — HoVer-Net/HACTNet, local ONNX inference,batching, checkpoint trust, evaluation, and model provenance.
All checked 2026-07-23:
https://doi.org/10.1158/1541-7786.MCR-21-0665
https://doi.org/10.1016/j.labinv.2025.104220
Other measured skills in the registry, with their headline benchmark lift.