Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Population genetics analysis — allele frequencies (gnomAD, 1000 Genomes), Hardy-Weinberg equilibrium testing, Fst between populations, GWAS associations, evolutionary constraint scores. Use for cross-population variant comparison, ancestry-aware allele frequency lookups, and population-level evolutionary analysis.
.claude/skills/mims-harvard-tooluniverse-population-genetics/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 300% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 176% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 283% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 543% | 0% |
| case-02 | ✓→✗ | ▼ Worse | 252% | 0% |
MC Strategy: Population genetics MC questions often test whether you know a specific theorem or result. COMPUTE the answer first (use popgen_calculator.py or write Python), then match to options. Don't try to reason about which option "sounds right."
Analyze population-level genetic variation, allele frequencies, GWAS associations, clinical significance, and evolutionary constraints using ToolUniverse tools.
Activate this skill when the user asks about:
Query gnomAD/1000Genomes/GWAS Catalog FIRST for allele frequencies and associations. Preferred: use the PopGen_hwe_test, PopGen_fst, PopGen_inbreeding, and PopGen_haplotype_count tools for HWE, Fst, inbreeding, and haplotype calculations. Fallback: run popgen_calculator.py directly. For theoretical problems (delta-q, drift, LD decay), apply the formulas in the Theoretical Reasoning section below.
| Tool | Key Parameters | Notes | |------|---------------|-------| | gnomad_search_variants | query (REQUIRED) | Resolve rsID to variant_id format "CHR-POS-REF-ALT" | | gnomad_get_variant | variant_id (REQUIRED), dataset | Population frequencies. Default dataset: gnomad_r3; use gnomad_r4 for latest | | gnomad_get_gene_constraints | gene_symbol (REQUIRED) | pLI, o/e ratios. May timeout -- retry once | | MyVariant_query_variants | query (REQUIRED) | Aggregated: ClinVar + dbSNP + gnomAD + CADD. Uses hg19 coordinates | | EnsemblVEP_annotate_rsid | variant_id (REQUIRED) | Functional consequence, SIFT, PolyPhen. Param is "variant_id" NOT "rsid" | | EnsemblVEP_variant_recoder | variant_id (REQUIRED) | Convert between rsID/HGVS/VCF/SPDI | | gwas_get_snps_for_gene | gene_symbol (REQUIRED) | All GWAS SNPs for a gene | | gwas_search_associations | query (REQUIRED) | GWAS for a disease/trait (NOT gene name -- use gwas_get_snps_for_gene for genes) | | gwas_get_variants_for_trait | trait (REQUIRED) | Variants associated with a trait | | ClinVar_search_variants | gene, condition, significance | At least one filter required | | RegulomeDB_query_variant | rsid (REQUIRED) | Regulatory scoring (1a=strongest to 7=minimal) |
"CHR-POS-REF-ALT" (no "chr" prefix). Always resolve rsIDs via gnomad_search_variants first.gwas_get_snps_for_gene for gene-based lookups.gwas_get_snps_for_gene instead.{data, metadata}, or {error}). Handle all three.Variant frequency: gnomad_search_variants -> gnomad_get_variant(dataset="gnomad_r4") -> MyVariant_query_variants (1000G pop breakdowns) -> EnsemblVEP_annotate_rsid
GWAS for disease: gwas_search_associations -> gwas_get_variants_for_trait -> gnomad_get_variant for top hits -> EuropePMC_search_articles
Gene characterization: gnomad_get_gene_constraints -> gwas_get_snps_for_gene -> ClinVar_search_variants -> PubMed_search_articles
Pathogenicity assessment: EnsemblVEP_annotate_rsid -> MyVariant_query_variants (CADD, ClinVar) -> gnomad_get_variant (frequency) -> RegulomeDB_query_variant (if non-coding)
These formulas are needed for quantitative population genetics problems. Work through step by step, showing intermediate values.
For a recessive deleterious allele (fitness: AA=1, Aa=1, aa=1-s):
delta_q = -s * q^2 * p / (1 - s * q^2)where p = freq(A), q = freq(a), s = selection coefficient.
For dominant deleterious (AA=1, Aa=1-s, aa=1-s):
delta_q = -s * q * p / (1 - s * q * (2 - q))For heterozygote advantage (AA=1-s1, Aa=1, aa=1-s2):
equilibrium: q_hat = s1 / (s1 + s2)Example: plug in s1 and s2 from the question; q_hat = s1/(s1+s2).
Selection against recessives is slow at low q because most a alleles hide in heterozygotes. Time to reduce q from q0 to qt: t ~ (1/qt - 1/q0) / s generations.
Variance in allele frequency per generation: Var(delta_p) = pq / (2Ne)
Probability of fixation of a new neutral mutation: 1/(2Ne)
Time to fixation (given it fixes): ~4Ne generations for neutral alleles
Heterozygosity decay: H_t = H_0 (1 - 1/(2Ne))^t
After t generations, fraction of heterozygosity lost ~ 1 - e^(-t/(2Ne))
Effective population size (Ne) adjustments:
Drift vs selection: Drift dominates when |s| < 1/(2Ne). A variant with s=0.01 behaves neutrally in a population of Ne < 50.
D = freq(AB) - freq(A)freq(B), where A and B are alleles at two loci.
Decay with recombination: D_t = D_0 (1 - r)^t, where r = recombination fraction, t = generations.
Half-life of LD: t_half = -ln(2) / ln(1-r) ~ 0.693/r generations (for small r).
r-squared (normalized LD): r^2 = D^2 / (pA pa pB pb). Range 0-1.
Expected r^2 in finite population at equilibrium: Er^2] = 1 / (1 + 4Ner) (for drift-recombination balance).
Practical implications:
For alleles A (freq p) and a (freq q=1-p): expected genotypes AA=p^2, Aa=2pq, aa=q^2.
Chi-square test: df=1 (2 alleles). Preferred: use PopGen_hwe_test tool. Fallback: popgen_calculator.py --type hwe --AA N1 --Aa N2 --aa N3.
Causes of HWE departure: non-random mating, selection, migration, drift, genotyping error. Excess homozygotes -> inbreeding or population structure (Wahlund effect). Excess heterozygotes -> overdominant selection or negative assortative mating.
For n SNPs between two inbred (homozygous) strains:
python3 -c "...") for these counts. Never enumerate by hand.Equilibrium frequency of a deleterious allele:
PopGen_fst tool. Fallback: popgen_calculator.py --type fst --p1 X --p2 Y --n1 N1 --n2 N2For any genetics cross problem, follow these steps IN ORDER. Do not skip steps.
For bacterial conjugation and Hfr mapping problems:
These are specific patterns that have caused reasoning failures in hard genetics questions. Review before answering genetics MCQs.
For "necessarily true" questions about PGS and heritability: a statement is necessarily true only if it holds when V_D=0 AND when V_D=V_G. Test the extremes.
Do NOT guess path signs from general knowledge. Signs may differ from well-known systems. Follow this protocol:
Compute chi-square from the expected ratio given in the question. Compare to chi-square-critical at df = (number of phenotype classes - 1). Pick the answer choice with the highest chi-square, but also check which pattern is biologically diagnostic of the alternative hypothesis.
LD block boundaries at recombination hotspots are a source of GWAS false localization — strong signal in the block does not guarantee the causal variant is in the block.
Duplex sequencing (unique molecular identifiers + double-strand consensus) detects alleles at 0.01% frequency — far below standard NGS even at 80X depth. Simply increasing read depth does NOT help for ultra-rare variants because the Illumina error rate (~0.1%) masks variants rarer than ~1% regardless of depth. Error correction methods (UMIs, duplex consensus) are needed to distinguish true rare variants from sequencing errors.
Script: skills/tooluniverse-population-genetics/scripts/popgen_calculator.py
Preferred: Use ToolUniverse tools (via MCP/SDK) instead of the script when possible:
PopGen_hwe_test tool -- HWE chi-square test. Fallback: popgen_calculator.py --type hwePopGen_fst tool -- Weir-Cockerham Fst. Fallback: popgen_calculator.py --type fstPopGen_inbreeding tool -- Inbreeding coefficient from pedigree. Fallback: popgen_calculator.py --type inbreedingPopGen_haplotype_count tool -- Expected haplotype diversity. Fallback: popgen_calculator.py --type haplotypesFallback script modes (all require --type):
hwe: --AA N --Aa N --aa N -- chi-square HWE test with p-valuefst: --p1 F --p2 F --n1 N --n2 N -- Weir-Cockerham Fstinbreeding: --pedigree TYPE --generations G -- F from pedigree (self, full-sib, half-sib, first-cousin, etc.)haplotypes: --snps N --generations G --recomb_rate R -- expected haplotype diversity| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,415 | 5,878 | -71% | 1 | 1 | 0% | 4,021 | 6,003 | +49% | 0 | 0 | — |
case-21 | fail→fail | 11,411 | 5,862 | -49% | 1 | 1 | 0% | 2,391 | 6,939 | +190% | 0 | 0 | — |
case-02 | pass→fail | 8,455 | 6,528 | -23% | 1 | 1 | 0% | 1,733 | 6,095 | +252% | 0 | 0 | — |
case-03 | fail→pass | 68,183 | 3,607 | -95% | 1 | 1 | 0% | 1,571 | 6,279 | +300% | 0 | 0 | — |
case-04 | fail→pass | 12,353 | 3,775 | -69% | 1 | 1 | 0% | 2,300 | 6,341 | +176% | 0 | 0 | — |
case-05 | fail→pass | 7,531 | 2,199 | -71% | 1 | 1 | 0% | 1,561 | 5,981 | +283% | 0 | 0 | — |
case-06 | pass→pass | 4,389 | 3,913 | -11% | 1 | 1 | 0% | 1,074 | 6,464 | +502% | 0 | 0 | — |
case-07 | pass→pass | 8,280 | 4,191 | -49% | 1 | 1 | 0% | 1,897 | 6,493 | +242% | 0 | 0 | — |
case-08 | pass→pass | 15,168 | 11,462 | -24% | 1 | 1 | 0% | 2,787 | 7,772 | +179% | 0 | 0 | — |
case-09 | pass→pass | 13,701 | 12,183 | -11% | 1 | 1 | 0% | 2,700 | 7,977 | +195% | 0 | 0 | — |
case-14 | pass→pass | 4,393 | 4,149 | -6% | 1 | 1 | 0% | 947 | 6,416 | +578% | 0 | 0 | — |
case-10 | pass→pass | 7,807 | 9,088 | +16% | 1 | 1 | 0% | 1,834 | 7,685 | +319% | 0 | 0 | — |
case-11 | pass→pass | 12,508 | 9,451 | -24% | 1 | 1 | 0% | 2,309 | 7,378 | +220% | 0 | 0 | — |
case-12 | fail→pass | 6,364 | 27,460 | +331% | 1 | 1 | 0% | 1,136 | 7,307 | +543% | 0 | 0 | — |
case-13 | pass→pass | 5,778 | 6,589 | +14% | 1 | 1 | 0% | 1,010 | 6,731 | +566% | 0 | 0 | — |
case-20 | pass→pass | 15,547 | 16,183 | +4% | 1 | 1 | 0% | 2,824 | 8,575 | +204% | 0 | 0 | — |
case-15 | pass→pass | 5,630 | 6,934 | +23% | 1 | 1 | 0% | 1,348 | 7,254 | +438% | 0 | 0 | — |
case-16 | pass→pass | 5,585 | 4,104 | -27% | 1 | 1 | 0% | 1,353 | 6,463 | +378% | 0 | 0 | — |
case-17 | pass→pass | 13,463 | 20,532 | +53% | 1 | 1 | 0% | 3,155 | 10,167 | +222% | 0 | 0 | — |
case-18 | pass→pass | 13,829 | 15,626 | +13% | 1 | 1 | 0% | 2,506 | 8,591 | +243% | 0 | 0 | — |
case-19 | pass→pass | 11,104 | 13,623 | +23% | 1 | 1 | 0% | 2,220 | 8,072 | +264% | 0 | 0 | — |
case-22 | pass→pass | 6,486 | 7,029 | +8% | 1 | 1 | 0% | 1,220 | 7,015 | +475% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.