Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Preprocesses cell-free DNA sequencing data including adapter trimming, alignment optimized for short fragments, and UMI-aware duplicate removal using fgbio. Applies cfDNA-specific quality thresholds and fragment length filtering. Use when processing plasma cfDNA sequencing data before downstream analysis.
.claude/skills/bio-cfdna-preprocessing/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | — | — |
| case-17 | ✗→✓ | ▲ Improved | — | — |
| case-01 | ✗→✓ | ▲ Improved | — | — |
| case-02 | ✗→✓ | ▲ Improved | — | — |
| case-07 | ✗→✓ | ▲ Improved | — | — |
Reference examples tested with: BWA 0.7.17+, fgbio 2.1+, matplotlib 3.8+, numpy 1.26+, pysam 0.22+, samtools 1.19+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Preprocess my cfDNA sequencing data" → Process cell-free DNA reads with UMI extraction, consensus calling, and error suppression for sensitive variant detection.
fgbio FastqToBam → fgbio GroupReadsByUmi → fgbio CallMolecularConsensusReadsPreprocess cell-free DNA sequencing data with UMI-aware deduplication.
| Factor | Requirement | Rationale | |--------|-------------|-----------| | Collection tube | Streck (7 days) or EDTA (6 hrs) | Prevents cell lysis | | Processing time | ASAP or per tube specs | Minimizes genomic DNA contamination | | Hemolysis | Avoid | Releases cellular DNA | | Storage | -80C after extraction | Prevents degradation |
bash# fgbio 3.0+ (actively maintained) # Step 1: Extract UMIs from reads and annotate fgbio ExtractUmisFromBam \ --input raw.bam \ --output with_umis.bam \ --read-structure 3M2S+T 3M2S+T \ --molecular-index-tags ZA ZB \ --single-tag RX # Step 2: Align with BWA-MEM # Use -Y for soft-clipping (preserves UMIs) bwa mem -t 8 -Y reference.fa with_umis.bam | \ samtools view -bS - > aligned.bam # Step 3: Group reads by UMI fgbio GroupReadsByUmi \ --input aligned.bam \ --output grouped.bam \ --strategy adjacency \ --edits 1 \ --min-map-q 20 # Step 4: Call molecular consensus reads fgbio CallMolecularConsensusReads \ --input grouped.bam \ --output consensus.bam \ --min-reads 2 \ --min-input-base-quality 20 # Step 5: Filter consensus reads fgbio FilterConsensusReads \ --input consensus.bam \ --output filtered_consensus.bam \ --ref reference.fa \ --min-reads 2 \ --max-read-error-rate 0.05 \ --min-base-quality 30
Goal: Run the complete cfDNA UMI-consensus pipeline from raw BAM to error-suppressed consensus reads in a single Python function call.
Approach: Chain fgbio operations (UMI extraction, grouping, consensus calling, filtering) with BWA alignment, handling intermediate files and cleanup within the function.
pythonimport subprocess import pysam from pathlib import Path def preprocess_cfdna(input_bam, output_bam, reference, read_structure='3M2S+T 3M2S+T', min_reads=2, threads=8): ''' Full cfDNA preprocessing pipeline with fgbio. Args: input_bam: Input BAM with UMIs in reads output_bam: Output consensus BAM reference: Reference FASTA path read_structure: UMI read structure min_reads: Minimum reads per UMI group threads: CPU threads ''' work_dir = Path(output_bam).parent prefix = Path(output_bam).stem # Extract UMIs with_umis = work_dir / f'{prefix}_umis.bam' subprocess.run([ 'fgbio', 'ExtractUmisFromBam', '--input', input_bam, '--output', str(with_umis), '--read-structure', read_structure, '--single-tag', 'RX' ], check=True) # Align aligned = work_dir / f'{prefix}_aligned.bam' cmd = f'bwa mem -t {threads} -Y {reference} {with_umis} | samtools view -bS - > {aligned}' subprocess.run(cmd, shell=True, check=True) # Sort sorted_bam = work_dir / f'{prefix}_sorted.bam' pysam.sort('-@', str(threads), '-o', str(sorted_bam), str(aligned)) # Group by UMI grouped = work_dir / f'{prefix}_grouped.bam' subprocess.run([ 'fgbio', 'GroupReadsByUmi', '--input', str(sorted_bam), '--output', str(grouped), '--strategy', 'adjacency', '--edits', '1' ], check=True) # Consensus calling consensus = work_dir / f'{prefix}_consensus.bam' subprocess.run([ 'fgbio', 'CallMolecularConsensusReads', '--input', str(grouped), '--output', str(consensus), '--min-reads', str(min_reads) ], check=True) # Filter consensus subprocess.run([ 'fgbio', 'FilterConsensusReads', '--input', str(consensus), '--output', output_bam, '--ref', reference, '--min-reads', str(min_reads) ], check=True) return output_bam
pythonimport pysam import numpy as np import matplotlib.pyplot as plt def analyze_fragment_sizes(bam_path, max_size=500): '''Analyze cfDNA fragment size distribution.''' bam = pysam.AlignmentFile(bam_path, 'rb') sizes = [] for read in bam.fetch(): if read.is_proper_pair and not read.is_secondary and read.template_length > 0: if read.template_length <= max_size: sizes.append(read.template_length) bam.close() # cfDNA signature: peak at ~167bp (mononucleosome) # Shorter fragments (90-150bp) enriched in ctDNA sizes = np.array(sizes) print(f'Fragments analyzed: {len(sizes)}') print(f'Median size: {np.median(sizes):.0f} bp') print(f'Mode: {np.bincount(sizes).argmax()} bp') return sizes
| Metric | Threshold | Notes | |--------|-----------|-------| | Modal fragment size | 150-180 bp | Peak ~167 bp indicates good cfDNA | | UMI families >= 2 reads | > 50% | Sufficient for consensus | | Mean base quality | >= 30 | After consensus | | Mapping quality | >= 20 | Exclude multi-mappers |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-17 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.