Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Process, clean, compare, and search tandem mass spectra with matchms. Use for MS/MS file I/O, metadata harmonization, peak filtering, spectral similarity, library matching, score matrices, and molecular-similarity networks. Use pyopenms instead for LC-MS feature detection or proteomics pipelines.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 76% | 0% |
Matchms is a Python package for importing, cleaning, processing, and comparing tandem mass spectra. This skill targets matchms 0.33.1, released 2026-06-08, and corrects several breaking API changes that older tutorials do not reflect.
Use matchms for:
Do not use matchms as a replacement for:
protein quantification — use pyopenms
proof of identity
Create or activate an environment, then install the release used by this skill:
bashuv pip install "matchms==0.33.1"
Verify the runtime:
bashuv run python -c "import matchms; print(matchms.__version__)"
Matchms 0.33.1 supports Python 3.10-3.14 and installs RDKit as a regular dependency. The old matchms[chemistry] extra is not part of the current package metadata.
coverage, ion mode, peak counts, and identifier fields.
a deliberate requirement.
Keep metadata enrichment separate when reference annotations are richer.
require_* filters return None.Modified and neutral-loss scores require valid precursor_mz.
len(references) * len(queries) before scoring. A sparse resultcontainer does not automatically avoid computing every requested pair.
score name, number of matched peaks when available, and candidate metadata.
agreement, ion/adduct compatibility, and orthogonal evidence.
These points prevent the most common failures from pre-0.33 examples:
ModifiedCosineGreedy or ModifiedCosineHungarian; ModifiedCosine wasremoved in 0.32.0.
add_losses(). It was removed in 0.27.0; usespectrum.losses, spectrum.compute_losses(...), or NeutralLossesCosine directly.
SpectrumProcessor is not callable. Use process_spectrum() orprocess_spectra().
process_spectra() returns (processed_spectra, processing_report).Scores.scores is a StackedSparseArray, often with separate structuredfields such as CosineGreedy_score and CosineGreedy_matches.
scores_by_query() returns (reference_spectrum, score_record) pairs, notreference indices.
spectra in parameter names. The legacy spelling spectrums isdeprecated.
See references/migration.md for a complete old-to-current mapping.
pythonfrom matchms import SpectrumProcessor, calculate_scores from matchms.filtering import ( default_filters, normalize_intensities, require_minimum_number_of_peaks, select_by_relative_intensity, ) from matchms.importing import load_spectra from matchms.similarity import ModifiedCosineGreedy def load_and_process(path): spectra = [default_filters(spectrum) for spectrum in load_spectra(path)] processor = SpectrumProcessor( [ normalize_intensities, (select_by_relative_intensity, {"intensity_from": 0.01}), (require_minimum_number_of_peaks, {"n_required": 5}), ] ) processed, _ = processor.process_spectra( spectra, progress_bar=False, create_report=False, ) return processed references = load_and_process("library.msp") queries = load_and_process("queries.mgf") metric = ModifiedCosineGreedy(tolerance=0.02) scores = calculate_scores( references=references, queries=queries, similarity_function=metric, ) score_name = "ModifiedCosineGreedy_score" matches_name = "ModifiedCosineGreedy_matches" for query in queries: ranked = scores.scores_by_query(query, name=score_name, sort=True) for reference, values in ranked[:5]: print( query.get("spectrum_id", query.get("id")), reference.get("compound_name", reference.get("spectrum_id")), float(values[score_name]), int(values[matches_name]), )
SpectrumProcessor automatically orders built-in filters according to matchms's filter order. The aggregate default_filters callable is not in that registry, so run it first as above or expand its nine component filters. Inspect processor.processing_steps and preserve it with results.
Similarity classes expose pair() for one reference/query pair. Cosine-family results are structured NumPy scalars:
pythonfrom matchms.similarity import CosineGreedy result = CosineGreedy(tolerance=0.02).pair(reference, query) similarity = float(result["score"]) matched_peaks = int(result["matches"])
Use calculate_scores() for matrix-oriented methods such as FlashSimilarity; its single-pair path is supported but intentionally not the optimized path.
CosineGreedy — standard peak cosine with greedy peak assignment.CosineHungarian — exact assignment; slower, useful for benchmarks.CosineLinear — current linear-scaling cosine implementation.ModifiedCosineGreedy — permits precursor-delta-shifted matches; common foranalog search.
ModifiedCosineHungarian — exact modified-cosine assignment.NeutralLossesCosine — compares losses computed from precursor and fragments.BlinkCosine — fast BLINK-style cosine approximation for larger matrices.FlashSimilarity — optimized matrix scoring using spectral entropy or cosinewith fragment, neutral-loss, or hybrid matching.
BinnedEmbeddingSimilarity — binned spectral vectors and optional approximatenearest-neighbor indexing.
PrecursorMzMatch, ParentMassMatch, MetadataMatch — candidate masks ormetadata constraints, not rich spectral scores.
FingerprintSimilarity — molecular-structure similarity; it is not spectralsimilarity and requires fingerprints prepared from valid structures.
Read references/similarity.md before choosing a fast method, combining scores, or interpreting structured outputs.
For all-vs-all scoring of one collection, set is_symmetric=True:
pythonscores = calculate_scores( references=spectra, queries=spectra, similarity_function=CosineGreedy(tolerance=0.02), array_type="sparse", is_symmetric=True, )
For a precursor-gated search, compute and filter PrecursorMzMatch first, then calculate the spectral metric only on retained coordinates through Pipeline or Scores.calculate(...). See references/workflows.md.
Do not choose a universal "identification threshold." Score distributions depend on preprocessing, mass accuracy, collision conditions, library quality, and metric. At minimum, retain both score and matched-peak count for cosine-family methods.
scripts/library_search.py provides a reproducible query-versus-library search with current score extraction, pair-count limits, preprocessing, and CSV output:
bashuv run python scripts/library_search.py \ queries.mgf library.msp hits.csv \ --metric modified \ --tolerance 0.02 \ --top-k 10 \ --min-score 0.6 \ --min-matches 5
Run --help for fast metrics, preprocessing options, identifier fields, overwrite control, and the explicit large-matrix override.
pythonimport numpy as np from matchms import Spectrum spectrum = Spectrum( mz=np.array([100.0, 150.0, 200.0]), intensities=np.array([0.2, 1.0, 0.4]), metadata={"spectrum_id": "query-1", "precursor_mz": 250.5}, ) print(spectrum.peaks.mz) print(spectrum.get("precursor_mz")) losses = spectrum.compute_losses(loss_mz_from=5.0, loss_mz_to=200.0) spectrum.plot() spectrum.plot_against(reference_spectrum)
Read only the reference needed for the task:
references/importing_exporting.md — formats, return types, generic I/O,mzSpecLib, score serialization, and pickle safety
references/filtering.md — current filter catalog, clone/None semantics,default filters, ordering, and SpectrumProcessor
references/similarity.md — all current similarity classes, outputs,candidate masking, performance, and interpretation
references/workflows.md — library search, sparse gating, Pipeline, networks,plotting, and provenance
references/migration.md — breaking changes and deprecated APIsreferences/sources.md — authoritative docs, release notes, user guides, andscientific publications used for this refresh
Scores value is a plain float; inspect score_names.Other measured skills in the registry, with their headline benchmark lift.