Huan Fan http://fanhuan.github.io 2026-08-20T08:55:02+00:00 huan.fan@wisc.edu 2017 fontiers in Plant Science Tranbarger http://fanhuan.github.io/en/2026/08/17/Tranbarger-2017/ 2026-08-17T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/08/17/Tranbarger-2017 Link To The Article I

Title: Transcriptome Analysis of Cell Wall and NAC Domain Transcription Factor Genes during Elaeis guineensis Fruit Ripening: Evidence for Widespread Conservation within Monocot and Eudicot Lineages

The oil palm (Elaeis guineensis), a monocotyledonous 

cotyledon: 子叶

species in the family Arecaceae, has an extraordinarily oil rich fleshy mesocarp, and presents an original model to examine the ripening processes and regulation in this particular monocot fruit. Histochemical analysis 

Histochemical analysis(组织化学分析) is the use of chemical reactions performed in situ on tissue sections to localize specific molecules — carbohydrates, lipids, proteins, phenolics, nucleic acids, or enzyme activities — while preserving the spatial organization of the tissue. The core idea is that you’re not homogenizing the sample and assaying a bulk extract; you’re asking where in the tissue something is.

Tissue is usually fixed and sectioned via cryosection(冷冻切片), paraffin(石蜡切片), or resin(树脂切片), and treated with a reagent that yields a colored or fluorescent product where the target is present. Classic examples are:

  • PAS (periodic acid–Schiff) — magenta(品红) for polysaccharides/starch
  • Sudan/Nile red(苏丹红) — neutral lipids and oil bodies
  • Toluidine(甲苯胺) blue O — metachromatic(changes color when it binds to specific chemicals); distinguishes lignified(木质化) vs. pectinaceous(富含果胶的) walls
  • Phloroglucinol-HCl / Mäule — lignin and lignin subtype
  • Enzyme histochemistry — a substrate is converted by endogenous enzyme into an insoluble precipitate (e.g., peroxidase, phosphatase activity)
    and cell parameter measurements revealed cell wall and middle lamella expansion and degradation during ripening and in response to ethylene. 
    

    The middle lamella(胞间层/中胞层,主要是果胶 pectin) is a layer that cements together the primary cell walls of two adjoining plant cells.

    Cell wall related transcript profiles suggest a transition from synthesis to degradation is under transcriptional control during ripening, in particular a switch from cellulose, hemicellulose, and pectin synthesis to hydrolysis and degradation.
    

    Now we need to understand what are “cell wall related transcript profiles”. In the “Mesocarp Transcriptome Data Mining” section in “Materials and Methods”,

    Transcriptome data of the developing mesocarp previously clustered (clusters A, B, C, and D) was searched for transcripts with expression profiles that either increase or decrease during the burst of ethylene production observed during oil palm fruit ripening between 100 and 160 DAP (Supplementary Table 1; Tranbarger et al., 2011).
    

    So first of all, the transcript needs to responding to ethylene. Then they talked about how the transcripts were annotated:

    The BLAST2GO and InterProScan web services with BLASTX using an E-value cutoff of 1e-5 were used to annotate the gene sets (Altschul et al., 1990; Zdobnov and Apweiler, 2001; Götz et al., 2008).
    

    Now some filtering:

  1. “Cell wall sequences were identified by searching the GO annotated sequences for InterPro accessions and key words related to cell wall processes”
  2. “by searching (TBLASTX) the 454 sequence database with known candidates related to cell wall biosynthesis and degradation.”

This ends up with “A total of 75 transcripts for cell wall related activities were found to be differentially expressed in the mesocarp during ripening, 63% of which have expression peaks at 140 and 160 DAP (including EgPG4) concomitant with the ethylene burst as measured previously (Tranbarger et al., 2011).”

The data provide evidence for the transcriptional activation of expansin, polygalacturonase, mannosidase, beta-galactosidase, and xyloglucan endotransglucosylase/hydrolase proteins in the ripening oil palm mesocarp, suggesting widespread conservation of these activities during ripening for monocotyledonous and eudicotyledonous fruit types.

These are hand-picked genes that are important in the ripening of tomato and banana.

Profiling of the most abundant oil palm polygalacturonase (EgPG4) and 1-aminocyclopropane-1-carboxylic acid oxidase (ACO) transcripts during development and in response to ethylene demonstrated both are sensitive markers of ethylene production and inducible gene expression during mesocarp ripening, and provide evidence for a conserved regulatory module between ethylene and cell wall pectin degradation."

First of all, what is EgPG4? “Recent studies by our group identified a polygalacturonase (EgPG4) highly induced by ethylene in oil palm fruit abscission zone cells and associated with the cell separation and fruit abscission (Roongsattham et al., 2012, 2016).” While in this 2016 Paper, “Previous studies revealed that oil palm fruit AZ cell walls are rich in unmethylated pectin and that PG activity and the EgPG4 transcript is highly expressed in the AZ in response to ethylene (Henderson and Osborne, 1994; Henderson et al., 2001; Roongsattham et al., 2012).”. Indeed, this (2012 Paper)[https://link.springer.com/article/10.1186/1471-2229-12-150] was all about the PG gene family in oil palm. Also in this paper, EgPG4 was reported to be “the most highly induced in the fruit base, with a 700–5000 fold increase during the ethylene treatment.” among all the PGs.

Secondly, “profiling” here means real-time RT-PCR. This paper did not generate any new RNAseq data.

“A comprehensive analysis of NAC transcription factors confirmed at least 10 transcripts from diverse NAC domain clades are expressed in the mesocarp during ripening, four of which are induced by ethylene treatment, with the two most inducible (EgNAC6 and EgNAC7) phylogenetically similar to the tomato NAC-NOR master-ripening regulator.”

Overall, the results provide evidence that despite the phylogenetic distance of the oil palm within the family Arecaceae from the most extensively studied monocot banana fruit, it appears ripening of divergent monocot and eudicot fruit lineages are regulated by evolutionarily conserved molecular physiological processes.

]]>
When a GRM Says Cousins Are Clones — Wahlund Meets Rare-Allele Weighting http://fanhuan.github.io/en/2026/07/21/Wahlund-GRM/ 2026-07-21T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/07/21/Wahlund-GRM In the last post I talked myself into fastGWA and its sparse GRM: compute all the pairwise relationships, then zero out everything below a cutoff (default 0.05) so only close kin survive as a random effect, while PCs handle population structure. Clean division of labour. So I went and built the thing. And the moment I looked at it, it was obviously broken — in a way that turned out to be the same lesson as my founders post, wearing a different costume.

This post is about why my GRM reported pairs of individuals as being more than identical twins, and why the culprit was not a bug but a property of pooling divergent populations into one relationship matrix — the Wahlund effect, amplified by how GCTA weights variants.

1. The symptom: a “sparse” GRM that wasn’t sparse

I built the GRM from my LD-pruned, --maf 0.01 set (~425k SNPs, ~2000 oil palm samples) and sparsified at the default 0.05:

gcta --bfile pruned_set --make-grm --out grm_full --threads 20
gcta --grm grm_full --make-bK-sparse 0.05 --out grm_sparse --threads 20

The log cheerfully reported 922,287 pairs retained. For ~2000 individuals there are only ~2 million possible pairs — so ~45% of all pairs survived a “keep only close relatives” cutoff. A sparse GRM that keeps half the matrix is not sparse, and no cohort has half its pairs related at 0.05+.

So I dumped the full GRM (--make-grm-gz) and actually looked at the numbers. The GCTA .grm.gz is four columns — i, j, #SNPs, value — with i==j being the diagonal. Two things fell out:

  • Diagonals should be ~1 (a diagonal is 1 + F, so 1 plus the inbreeding coefficient). Mine had a mean of 1.35 and a maximum of 5.96. A diagonal of 6 implies F ≈ 5. Inbreeding coefficients live in [0, 1]. This is not a number biology can produce.
  • Off-diagonals (the actual relatedness values) topped out at 4.5, with 8,820 pairs reporting relatedness above 1.0. Relatedness caps at ~1 for clones/MZ twins and ~0.5 for parent–offspring. My matrix was calling thousands of pairs five times more related than identical twins.

Those impossible values are the tell. This isn’t “my population is unusually related.” It’s a distorted matrix.

2. The fingerprint: sort the diagonal by population

The clue was in who had the crazy diagonals. My dataset is a germplasm collection: a big breeding population coded by parent numbers (081.081, 161.161, TS1, TS3 …) plus a scatter of named exotic origins (Angola, Deli, Ghana, Nigeria, Tanzania, AVROS, Ekona…). When I averaged the GRM diagonal within each group, it split into two clean tiers:

Group mean diagonal n
Tanzania 4.52 14
AVROS 3.95 3
Nigeria 3.53 19
AGO (Angola) 3.28 115
Ghana 2.59 22
Deli 2.16 82
TS3 1.08 170
081.081 0.83 191
161.161 0.80 196
TS1 0.53 61

The exotic origins — the small, genetically distinct groups — are the ones blowing up. The large breeding population sits right around a sane ~1 (or even below). The inflation isn’t random noise or bad samples; it’s structured by population membership. That points at one thing: allele frequencies.

3. Why rare-in-the-pool alleles detonate the GRM

Here is GCTA’s GRM entry between individuals j and k, summed over variants i with pooled allele frequency p_i:

A_jk = (1/M) Σ_i  (x_ij − 2p_i)(x_ik − 2p_i) / [ 2 p_i (1 − p_i) ]

Stare at the denominator: 2 p_i (1 − p_i). When a variant is rare in the pooled sample, p_i is tiny, so the denominator is tiny, so every term for that variant is divided by a very small number — its contribution explodes. This is deliberate: rare-variant sharing is stronger evidence of recent common ancestry, so GCTA up-weights it. The weighting is only trustworthy, though, when p_i actually describes the individuals you’re applying it to.

Now bring in the structure. GCTA computes each p_i across the whole pooled sample, which is ~75% breeding population. Take a variant that is nearly absent in the breeding majority but common, even fixed, inside Tanzania. Pooled, its p_i comes out small. But every Tanzanian is homozygous for that “rare” allele — so (x − 2p) is large and it gets divided by a tiny 2p(1−p). Two Tanzanians both homozygous for it contribute a gigantic positive product to their pairwise A_jk, and each contributes a gigantic term to their own diagonal. Multiply across all such variants and you get diagonals of 4–6 and within-origin “relatedness” above 1.

The individuals aren’t clones. The frequencies are computed on the wrong reference population, and the inverse-frequency weighting turns that mismatch into enormous numbers.

4. The name for it: the Wahlund effect

There’s a classic population-genetics name for the root cause. The Wahlund effect: when you pool subpopulations that have different allele frequencies and treat them as one HWE population, you see a deficit of heterozygotes / excess of homozygotes relative to what pooled allele frequencies predict — even if every subpopulation is in perfect HWE internally. Structure masquerades as inbreeding.

A GRM diagonal is essentially a measurement of an individual’s homozygosity relative to pooled-HWE expectation. So the Wahlund excess-homozygosity lands straight on the diagonal as spurious inbreeding, and the same frequency mismatch lands on the off-diagonals as spurious within-group relatedness. My “F ≈ 5” individuals aren’t inbred — they belong to a subpopulation whose allele frequencies look nothing like the pool I forced them into.

It’s worth noticing this is the second way I’ve been burned by fake inbreeding. Back in How Low is Low? the villain was low coverage: miss the second allele of a heterozygote and you call it homozygous, so observed heterozygosity (H-obs) sags and F inflates as a sequencing artifact. Here the villain is population structure: pool subpopulations and the heterozygote deficit inflates F as a base-population artifact. Same symptom — excess homozygosity, an inbreeding coefficient that’s too high — two completely unrelated causes. When F looks wrong, “is it the reads or is it the base population?” is now a question I know to ask.

This is the exact same lever as my founders / base-population headache, just downstream. There the question was whose allele frequencies define --maf; here it’s whose allele frequencies define the GRM weights. Same denominator problem, different tool. The base population you (implicitly) choose is doing all the work.

5. Two things that did not fix it

Adding samples back. I first ran this with 1,996 samples, then noticed I’d expected ~2,023 and regenerated with 2,022. The distortion was identical (diag max 5.96, off-diag max 4.5, ~45% surviving the cutoff either way). Of course it was — the problem was never the sample count, it was which populations were pooled.

Tuning the sparse cutoff. The instinct is to raise --make-bK-sparse from 0.05 to 0.1 or 0.2 to trim the flood of pairs. But raising the cutoff on a distorted matrix just discards a different arbitrary slice of a matrix whose numbers are wrong. It treats the symptom (too many surviving pairs) and ignores the disease (the values themselves are inflated by frequency mismatch). It would also throw the PCA off the same cliff — and sure enough, this is exactly why my PC1 eigenvalue (585) dwarfed PC2 (222): PC1 was separating the exotic origins from the breeding population, i.e. it was reporting the very structure the GRM was choking on.

6. The twist: the “broken” matrix still did one job

Here’s where I nearly made things worse. Having seen diagonals of 6 and relatedness above 1, my gut said this matrix is garbage, throw it out — I even talked myself into dropping the GRM from a downstream family-based mixed model entirely.

That was wrong, and my own earlier notes said so. When I fed this same distorted GRM into the mixed model as a variance component, it controlled genomic inflation just fine — the direct-effect λ came out around 0.95. Drop the GRM instead, and λ blew up past 6.

Why does a matrix full of impossible numbers still work? Because a mixed model doesn’t lean on the GRM’s magnitudes — it leans on its block structure: who clusters with whom. And the block structure is correct. The rare-allele weighting inflated the sizes of the within-origin similarities, but it didn’t scramble which individuals are similar — Tanzanians still look like Tanzanians. So as a device for saying “don’t be surprised these individuals’ phenotypes resemble each other,” the distorted GRM is still telling the truth.

What the distortion does wreck is anything that reads the magnitudes literally:

  • Heritability / variance components — with diagonals inflated, the model mis-partitions variance (in one trait my family-relatedness term collapsed to exactly zero while the GRM ate 43%). Any h² from this is fiction.
  • “Is this pair related?” — the off-diagonals are uninterpretable as kinship.
  • Per-SNP stability — a fraction of SNPs blow up (tiny/huge effective N) and need filtering.

So “broken” was too strong. The right question isn’t is this GRM valid? — it’s valid for what?

7. So what you do depends on what you need it for

  • If you need the magnitudes — heritability, kinship, “how related are these two” — the matrix is unusable as-is, and fixing it is a design decision, not a flag. Stop pooling: restrict to the intended study population (my breeding-cross groups, sane ~1 diagonals) and rebuild, or go within-population / stratified and meta-analyze, so 2p(1−p) means what GCTA thinks it means. No --maf, cutoff, or sample juggling fixes a matrix that pooled populations it shouldn’t have.
  • If you only need structure control in a mixed model, the distorted GRM may still calibrate — but prove it with the QQ plot / λ, don’t assume. The tempting cleanup — rebuild the GRM on common variants (MAF ≥ 0.05) to tame the diagonals — seems principled, but for family data it can backfire badly. Test it before you trust it; here’s what happened when I did.

8. The cleanup that backfired

The obvious next move — the one I’d half-talked myself into — is to rebuild the GRM from common variants only (MAF ≥ 0.05). The logic feels airtight: the explosion came from rare alleles hitting the tiny 2p(1−p) denominator, so drop the rare alleles and the diagonals should settle back to ~1. I built it and ran it. It failed twice over.

First, it didn’t even clean the matrix. The extreme spikes came down (max diagonal 6 → 2.5), but the baseline barely moved — mean diagonal 1.35 → 1.27, still thousands of impossible off-diagonals, still ~43% of pairs above the cutoff. By now the inflation isn’t coming from rare alleles; it’s coming from differentiated common alleles — an allele at 8% pooled but 40% inside one origin sails through a MAF ≥ 0.05 filter and carries the full Wahlund signal. You can’t threshold away structure that lives in the common variants.

Second — and worse — it broke the analysis. Feeding the common-only GRM into the mixed model, the variance-component fit blew up with “Factor is exactly singular” for exactly my well-measured traits — the ones that had fit fine with the rare-variant GRM. The reason is the lesson I won’t forget: in family data, rare variants are what distinguish siblings. My families are huge full-sib crosses (100–200 sibs each). Strip the rare variants and the remaining common-only genotypes make those sibs look nearly identical — near-duplicate rows in the GRM — so the matrix goes rank-deficient and the model can’t invert it. The “noise” I was filtering out was carrying the information that told relatives apart.

So the intuitive fix was worse than the disease: it left the structural distortion in place and destroyed the within-family signal. I reverted to the rare-variant-inclusive GRM — impossible diagonals and all — and everything converged again. The ugly matrix was the working one.

9. Key takeaway

Before you trust a GRM — sparse or dense — look at the numbers, not just the log line:

gcta --grm grm_full --make-grm-gz --out check
# diagonals should sit near 1; off-diagonals near 0 with a thin
# right tail at ~0.125 / 0.25 / 0.5. Anything above 1 is impossible.
zcat check.grm.gz | awk '$1==$2{d[$4>1.25]++} $1!=$2 && $4>1{imposs++}
  END{print "diag>1.25:", d[1]+0, " off-diag>1:", imposs+0}'

If diagonals run well above 1 and off-diagonals climb past 1, you’re seeing the Wahlund effect refracted through GCTA’s inverse-frequency weighting — a fancy way of saying your allele frequencies were computed on the wrong population. But don’t over-correct the way I almost did: a matrix like this is worthless as kinship or heritability, yet can still earn its keep as structure control in a mixed model, because block structure survives what magnitude doesn’t. So don’t ask “is my GRM broken?” — ask “broken for which job?”, and confirm the answer with λ.

And resist the reflexive cleanup. Filtering to common variants felt like the principled fix and made things strictly worse — it left the real (common-allele) structure untouched and threw away the rare variants that tell siblings apart, so the GRM went singular. Twice now the tidy-looking move (drop the GRM; drop the rare variants) was the wrong one, and the ugly, rare-variant-inclusive, impossible-on-paper matrix was the one that actually worked. The GRM, like --maf, like PCA, is only as meaningful as the base population you feed it — so choose that population on purpose, know which of its numbers you’re leaning on, and check what a “cleanup” is quietly throwing away before you trust it.

]]>
mlma-loco Or fastGWA http://fanhuan.github.io/en/2026/06/23/fastGWA/ 2026-06-23T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/23/fastGWA I had been running my mixed-model GWAS with GCTA’s --mlma-loco for a while, quite happily, until I tried to be clever about the GRM and ran straight into a wall. Sorting out why led me to fastGWA (Jiang et al., Nature Genetics 2019), and to finally understanding a word I’d been nodding along to for years without really getting: polygenicity. So this post is half a tooling decision and half a concept I should have learned ages ago.

1. The wall I hit with --mlma-loco

A mixed-model GWAS fits, for every individual, roughly this:

phenotype = (effect of the SNP being tested) + g + e

That middle term g is a per-individual genetic value modeled through a GRM (genetic relationship matrix) — a big matrix of “how genetically similar is every pair of individuals.” It’s the thing that lets the model say these two people resemble each other genetically, so don’t get excited when their phenotypes resemble each other either. It’s how you stop relatedness and population structure from manufacturing fake associations.

--mlma-loco builds that GRM in a clever leave-one-chromosome-out way: when testing SNPs on chromosome 1, it builds the GRM from chromosomes 2–N (and so on), so a SNP is never used to correct itself. Good design.

Here was my problem. I have tens of millions of variants, and I did not want to build the GRM from all of them — partly cost, partly because a GRM doesn’t need every redundant SNP in an LD block. I wanted to feed --mlma-loco a nice LD-pruned, common-SNP set for the GRM, while still testing every variant. And --mlma-loco simply won’t let you: it takes one --bfile and uses it for both the GRM and the association tests. The GRM SNPs and the test SNPs are welded together.

(There’s a manual workaround — build per-chromosome GRMs yourself and stitch LOCO back together with --mlma-subtract-grm — but it’s fiddly, and it turned out I was solving the wrong problem.)

2. Enter fastGWA, with a different philosophy

fastGWA was built for biobank-scale data (hundreds of thousands of samples), where a dense GRM is simply too big to compute or store. Its trick is a sparse GRM: compute the relationships, then zero out every relationship below a cutoff (default 0.05). What survives is only close relatives — sibs, half-sibs, parent–offspring, cousins. Everyone else is treated as unrelated.

That sounds reckless until you see the division of labour:

  • Population structure → handled by principal components as fixed-effect covariates (you supply them).
  • Close family relatedness → handled by the sparse GRM random effect.

So fastGWA splits the job that --mlma-loco does with one dense GRM into two specialised pieces. And crucially for me: the GRM is a separate input from the genotypes you test. You build the sparse GRM once, from whatever (pruned) set you like, and point the association step at the full set:

# sparse GRM, built once, from a pruned common-SNP set
gcta64 --bfile grm_snps --make-grm --out grm_full
gcta64 --grm grm_full --make-bK-sparse 0.05 --out grm_sparse

# test the FULL set; GRM supplied separately
gcta64 --fastGWA-mlm --bfile full_set --grm-sparse grm_sparse \
       --qcovar pcs.txt --pheno pheno.txt --out result

The exact decoupling I’d been fighting --mlma-loco for is just… how the tool works. The GRM is built once instead of rebuilt per chromosome, and LOCO stops being a worry — because a sparse GRM barely contains the tested SNP’s signal in the first place, there’s nothing meaningful to leak.

3. The thing I finally understood: polygenicity

Here’s where I had to stop and learn something. The difference between the dense GRM and the sparse GRM is really a difference in how much polygenicity they model. So what is it?

A trait is polygenic when it’s shaped by many variants — hundreds to thousands — each nudging the trait by a tiny amount, instead of a few big-effect genes. Basically every complex quantitative trait is like this. The causal alleles are sprinkled all over the genome and none of them is a smoking gun.

That g term above is the polygenic background: the summed effect of all those tiny variants for a given individual. We never estimate them one by one. Instead we say: two people who are genetically similar overall (high GRM value) probably share a lot of those small alleles, so their g should be similar too. Modeling g does two jobs — it removes confounding (structure/relatedness that would otherwise fake associations) and it buys power (it soaks up variance explained by the rest of the genome, so the SNP you’re testing stands out against a quieter background).

Now the dense-vs-sparse difference becomes clear:

  • Dense GRM (--mlma-loco) uses every pairwise similarity, including the faint, diffuse resemblance among people who aren’t relatives at all. So g captures the full genome-wide polygenic background.
  • Sparse GRM (fastGWA) keeps only close-kin similarities. Its g is really just family resemblance. The diffuse polygenic sharing across the whole sample is not in the random effect — fastGWA leans on the PCs to handle the structure part of it.

So when people say fastGWA “doesn’t model the full polygenic background,” that’s what they mean: it deliberately drops the diffuse genome-wide term and trusts your PCs to mop up structure instead.

4. So which should I use?

My case: a quantitative trait, a few thousand individuals (not biobank-scale), lots of real family structure from pedigree, and several distinct populations. I already add PCs to my model either way. Pulling it apart:

  • I’m not in the regime where fastGWA is required — at a few thousand samples a dense GRM is perfectly affordable. So speed isn’t the deciding factor.
  • My biggest confounders are family relatedness (handled by both) and gross population structure (handled by my PCs either way). Those do almost all of the work.
  • What I’d give up by going sparse is only the diffuse polygenic-background term — usually a modest power difference, not a correctness problem.
  • What I’d gain is the clean GRM/test-set separation, a GRM built once, and a tool that’s literally designed around family data via its sparse GRM.

For me that trade comes out in fastGWA’s favour, mostly because it solves my original problem by design instead of by workaround. The one discipline it demands: since fastGWA relies on the PCs to control structure rather than a dense GRM, I need enough PCs and I need to actually check the QQ plot / λ_GC to confirm structure is controlled. With a dense GRM that check is a bit more forgiving; with fastGWA it’s on me.

A quick way to hold the whole decision in your head:

  --mlma-loco fastGWA
GRM dense, rebuilt per chromosome sparse, built once
Polygenic background full (diffuse + family) family only
Structure control dense GRM (+ PCs) PCs (+ sparse GRM for kin)
GRM vs test set same --bfile separate inputs
Sweet spot moderate N, want full polygenic modeling large N, or you want GRM/test decoupled

5. Takeaway

I went looking for a way to feed --mlma-loco a pruned GRM and a full test set, and the real answer wasn’t a clever flag — it was a different tool with a different theory of where confounding comes from. --mlma-loco says “one dense GRM models everything, including polygenic background.” fastGWA says “let PCs handle structure, let a sparse GRM handle family, and don’t bother modeling diffuse polygenicity at all.” For biobank scale that’s a necessity; for my structured, pedigree-heavy, moderate-N quantitative dataset it’s just a cleaner fit — as long as I keep an eye on my PCs.

]]>
The MultiSuSiE Paper http://fanhuan.github.io/en/2026/06/22/MultiSuSiE-Paper/ 2026-06-22T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/22/MultiSuSiE-Paper Today’s paper is Rossen 2026 Nature Genetics, MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data. I am reading is because I use SuSiE for fine mapping.

  1. The data is all public, from All of Us. Impressive dataset. For a genome size of 3 billion bp, “all of Us identified more than 1 billion genetic variants, including more than 275 million previously unreported genetic variants, more than 3.9 million of which had coding consequences”. They also have longitudinal electronic health record which allowed them to evaluate 3,724 genetic variants associated with 117 diseases and found high replication rates across both participants of European ancestry and participants of African ancestry. You can start a genotyping company based on those 3724 variants if you work with people from those ancestries.
  2. In simulation, the authors used a balanced design (36K * 3) that matches the size of the European cohort (109K). This is capped by the 36K Latino-ancestry dataset.
  3. The concept of calibration. “To assess calibration, we compared the empirical FDR to (1 − PIP threshold), a conservative FDR upper bound (as in ref. 12), as well as (1 − mean PIP), the expected FDR (which has been reported to be slightly miscalibrated in previous fine-mapping simulations.”

Let’s break down this sentence. Calibration is to see whether the predicted FDR match the empirical FDR (FP/(FP+TP) measured in simulation where truth is known. Two ways of calibration are mentioned.

The first one is empirical FDR vs. (1-PIP threshold). Ref 12 is Weissbord 2020 Nature Genetics. In it, FDR is defined as “the proportion of false positives among SNPs with posterior causal probability (posterior inclusion probability (PIP)) above a given threshold (for example, PIP > 0.95), aggregating the results across all simulations”. This is to say, only SNPs with PIP > 0.95 are considered in the calculation of FDR, so in theory the FDR calculated this way should be much lower than the FDR calculated based on all the SNPs tested.

The second one is empirical FDR vs (1 - mean PIP). Here 1 - mean(PIP) is known as the expected FDR. To understand why, we need to firstly understand PIP, or posterior inclusion probability.

What is posterior inclusion probability? In the SuSiE paper, Wang 2020 Journal of the Royal Statistical Society Series B: Statistical Methodology, it is defined as (equation 2.3)

\[\text{PIP}_j := \Pr(b_j \neq 0 \mid X, y)\]

that is, the given the data (genotypes $X$ and phenotypes $y$.), what is probability that variable $j$ has a non-zero effect ($b_j \neq 0$).

Now, let’s try to understand why 1-mean(PIP) is the expected FDR.

As threshold of PIP(0.95) is almost always higher than mean(PIP), the first one is less tolerant of higher FDR (therefore a more conservative/lower upper bound).

Interestingly, Ref 12 is received on 28 October 2019 and Accepted on 02 October 2020. It cited SuSiE, which is also published in 2020, but it was submitted on 01 December 2018, and only accepted on 01 May 2020, from a different lab.

]]>
The Hard Question http://fanhuan.github.io/en/2026/06/19/The-Hard-Question/ 2026-06-19T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/19/The-Hard-Question When I was doing my first year of postdoc, I was supposed to extend my OPT (optional practical training), when I learnt that my university is no longer listed on e-verify and cannot hire me under a OPT any more. My OPT expires in about 40 days. I need to find another employer that is on e-verify before that in order to stay in the US. A friend very nicely referred me to a start-up company and I’ve got to meet the team. I remember two things from those interviews. The first one is this very professional lady, who worked for a prominent company, telling me that she joined this start-up due to family considerations. The other thing I remember was one of the questions from the CEO:”How do you distinguish rare variants from sequencing error?”

I don’t quite remember how I answered. In fact, to this day, I don’t know the answer.

Today let’s make some attempts at least try to understand the problem that we are facing.

Individual level vs. Population level

First of all, one need to understand that rare variants is a population-level concept, while sequencing error is at read level, and can be minimized at individual level.

Quality score

One obvious tool is quality score. Quality score of that base at fastq level, qulity score of the variant. If this variant got good coverage,

]]>
Order of Things http://fanhuan.github.io/en/2026/06/19/Order-Of-Things/ 2026-06-19T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/19/Order-Of-Things I was doing some QC of variants called using bcftools filter. Here are the three versions I had so far.

First Version

bcftools filter -g3 -G10 | bcftools view -i 'QUAL>20 & MAF>0.0000001' -m2 -M2

Firstly let’s explain each option separately.

  • -g3: Filter SNPs within 3 base pairs of an indel (the default) or any combination of indel,mnp,bnd,other,overlap. This is because a SNP this close to an indel might be incorportated into this indel as another allele.
  • -G10: Filter clusters of indels separated by 10 or fewer base pairs allowing only one to pass. Because one of the indels might be a result of the other. They should be considered together.
  • QUAL>20: this is the QUAL column of the vcf file. It is phred-scaled, and the higher the better. QUAL = 20 means only 1% of the chance that this variant is false.
  • MAF>0.00000001: Minor allele frequency for filtering rare variants. But here the number I set is so low, as long as you have 1 allele count, you will stay. So basically this is filtering for monomorphic/invariant calls.
  • -m2: minimum 2 alleles
  • -M2: maximum 2 alleles. So only bi-allelaic stays.

The problem with this version is that along the way, I lost a SNP because it was close to an indel that had low QUAL. Because -g3 -G10 goes first, I filtered that SNP with high QUAL for an indel with low QUAL. Therefore I reordered things in the next version.

Second Version

bcftools view -i 'QUAL>20 & MAF>0.0000001' -m2 -M2 | bcftools filter -g3 -G10

One obvious thing to improve is that MAF>0.0000001 the misleading. Might as well use 0 instead.

Now the problem becomes, should we filter the multiallellic (-m2 -M2) and monomorphic (MAF>0) variants upfront? For multiallellic indels, they might be true and with high quality scores, and they could be used in -g3 and -G10, and be discarded later. For the monomorphic indels, it could be everyone 1/1, monomorphic for the ALT — a fixed difference from the reference. This could mean that there is a mistake in the assembly and the indel is still true. I actually do not have a strong preference of whether this should go earlier or later. How do you think?

Third Version

bcftools view -i 'QUAL>20' -m2 -M2 | bcftools filter -g3 -G10 | bcftools view -i `MAF>0` -m2 -M2

What would you have done differently?

]]>
GTF Format and UTR Prediction http://fanhuan.github.io/en/2026/06/16/UTR-Prediction/ 2026-06-16T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/16/UTR-Prediction I was trying to prioritize some variants manually, after all the GWAS tests and fine mapping, to see whether the mutation it predicts in the protein is in the relevant domain and can cause actual structual changes. However when I wrote a script to generate the aa sequence with this mutation, the aa sequence was the same. What is going on?

This is the full annotation of this variant through SnpEff with some editing so it is not a real variant anymore. But the problem is true.

ANN=G|missense_variant|MODERATE|START_CODON_1_1000_1002|g12345|transcript|g12345.t1|protein_coding|11/15|c.721A>G|p.His241Asp|1826/2931|721/2562|241/853||WARNING_TRANSCRIPT_MULTIPLE_STOP_CODONS,G|synonymous_variant|LOW|START_CODON_1_1000_1002|g12345|transcript|g12345.t2|protein_coding|11/14|c.1803A>G|p.Val601Val|1803/2682|1803/2682|601/893||,G|intragenic_variant|MODIFIER|GENE_1_1000_28598|GENE_1_1000_28598|gene_variant|GENE_1_1000_28598|||n.23957A>G||||||,G|non_coding_transcript_variant|MODIFIER|TRANSCRIPT_1_1000_28252|null.22587|transcript|TRANSCRIPT_1_1000_28252|pseudogene||||||||,G|non_coding_transcript_variant|MODIFIER|TRANSCRIPT_1_1000_28598|null.22586|transcript|TRANSCRIPT_1_1000_28598|pseudogene||||||||

We can see that it is very long. There are multiple transcripts that it is involved in, separated by comma (,). Let’s put them into a table:

# Field Meaning 1 2 3 4 5
1 Allele the ALT allele being annotated G G G G G
2 Annotation effect, as a Sequence Ontology term missense_variant synonymous_variant intragenic_variant non_coding_transcript_variant non_coding_transcript_variant
3 Impact HIGH / MODERATE / LOW / MODIFIER MODERATE LOW MODIFIER MODIFIER MODIFIER
4 Gene_Name gene symbol START_CODON_1_1000_1002 START_CODON_1_1000_1002 GENE_1_1000_28598 TRANSCRIPT_1_1000_28252 TRANSCRIPT_1_1000_28598
5 Gene_ID gene identifier g12345 g12345 GENE_1_1000_28598 null.22587 null.22586
6 Feature_Type transcript, gene_variant, etc. transcript transcript gene_variant transcript transcript
7 Feature_ID transcript/feature identifier g12345.t1 g12345.t2 GENE_1_1000_28598 TRANSCRIPT_1_1000_28252 TRANSCRIPT_1_1000_28598
8 BioType protein_coding, pseudogene, etc. protein_coding protein_coding   pseudogene pseudogene
9 Rank/Total exon (or intron) rank / total 11/15 11/14      
10 HGVS.c nucleotide change (coding coords) c.721A>G c.1803A>G n.23957A>G    
11 HGVS.p amino-acid change p.His241Asp p.Val601Val      
12 cDNA_pos/len position in cDNA / cDNA length 1826/2931 1803/2682      
13 CDS_pos/len position in CDS / CDS length 721/2562 1803/2682      
14 AA_pos/len residue position / protein length 241/853 601/893      
15 Distance distance to feature (for intergenic)          
16 Errors/Warnings annotation QC messages WARNING_TRANSCRIPT_MULTIPLE_STOP_CODONS        

So one variant produced five annotations. Two questions jump out:

  1. why are there five entries for what is really one gene with only two transcripts
  2. a missense is nice but why is g12345.t1 flagged with WARNING_TRANSCRIPT_MULTIPLE_STOP_CODONS?

Where entries 3, 4, 5 come from

The strings pseudogene, null., TRANSCRIPT_1_… and GENE_1_… do not appear anywhere in the GTF. SnpEff fabricated entries 3–5 at database-build time because of how the BRAKER GTF writes its parent lines. Compare the gene/transcript lines with their child features:

gene        … 1000 28598 … +  .  g12345                                       ← bare ID
transcript  … 1000 28598 … +  .  g12345.t1                                    ← bare ID
CDS         … 1000 1449 … +  0  transcript_id "g12345.t1"; gene_id "g12345"; ← proper attributes
exon        … 1000 1449 … +  .  transcript_id "g12345.t1"; gene_id "g12345";

SnpEff’s GTF parser expects key "value"; attribute pairs. The gene and transcript lines in the raw BRAKER gtf output instead put a bare ID in column 9, so the parser extracts no ID and falls back to synthetic, coordinate-based markers. As a result SnpEff parses the locus twice:

  • the properly-attributed CDS/exon/intron lines correctly rebuild g12345 with transcripts .t1 and .t2 (entries 1 and 2) because there are proper “transcript_id” and “gene_id” as the key.
  • each unparseable gene line becomes a childless phantom gene GENE_chr_start_endintragenic_variant (entry 3), and each unparseable transcript line becomes a phantom transcript TRANSCRIPT_chr_start_end wrapped in an invented null.N gene. With no CDS attached, these default to pseudogene / non_coding_transcript_variant (entries 4 and 5).

This is genome-wide: every gene and transcript line in the file uses the bare-ID format, so the database ends up with a lot of GENE_* and null.* phantom records. The fix is to normalise the attribute column (give the parent lines real gene_id "…"; transcript_id "…"; fields) and rebuild. Entries 3–5 then disappear, leaving only the two real transcripts. Here normalise means converting records to a canonical/standard form in data/file engineering.

The real puzzle: the g12345.t1 warning

With the phantoms gone we are left with entries 1 and 2 — the same variant on two transcripts of the same gene: missense (p.His241Asp) on .t1 and synonymous (p.Val601Val) on .t2. I mentioned in the beginning that when I tried to apply this mutation to the protein sequence, the sequence did not change. Why did the missense go missing? Is it related to the WARNING_TRANSCRIPT_MULTIPLE_STOP_CODONS that .t1 carries?

Let’s take a look at the gtf of this gene. The difference between the two transcripts is that .t1 carries UTR features added by stringtie2utr and .t2 does not. Sorting .t1’s features by position shows the problem:

# transcript g12345.t1  
AUGUSTUS       start_codon     1000  1002
AUGUSTUS       CDS             1000  1449
   …
AUGUSTUS       CDS             20367  20499   (phase 2)
stringtie2utr  five_prime_UTR  20555  20577   ← a 5' UTR ~19 kb INTO the coding region
AUGUSTUS       CDS             20578  20731   (phase 1)
   …
AUGUSTUS       stop_codon      28250  28252
stringtie2utr  three_prime_UTR 28253  28598
# transcript g12345.t2 has the same CDS but no UTR lines

A 5′ UTR is by definition upstream of the start codon. This one sits in the middle of the CDS, downstream of the start codon and inside an intron. When SnpEff reads an explicit five_prime_UTR, it uses it to set the coding start — the new start codon would be right after the 5′ UTR ends. So SnpEff treats everything from the true start codon (1000) up to 20577 as untranslated and begins translation at 20578, in the wrong frame. In this way, the mutation would result in missense and multiple stop codons thus the warning.

The coordinates also confirm it. Counting the variant (genomic position 23957) from the bogus coding start at 20578:

20578–20731 (154) + 20967–21047 (81) + 21373–21495 (123)
+ 22568–22738 (171) + 22819–22953 (135) + 57 bases into 23901–23996
= 721  →  codon 241

That is precisely the c.721A>G | p.His241Asp of entry 1 — computed against the broken frame. Counting the same variant from the correct start codon (1000), the way .t2 does, gives c.1803A>G, codon 601, at the wobble (3rd) position — a synonymous A→G that leaves Val unchanged. That is entry 2.

Mystery solved.

This misplaced-UTR problem is not a one-off either; I found hundreds of them in my gtf disrupting gene models the same way.

Where the bad UTR comes from — and the fix

As suggested by the source column in the gtf, the UTRs were added by stringtie2utr.py (a BRAKER helper that decorates a gene model with UTRs inferred from a StringTie assembly). This itself is a long story. In theory we should be able to do UTR prediction in AUGUSTUS with the option --UTR=on. However it will return with error and it is a known issue in BRAKER3. Katherine the author suggested us to use stringtie2utr.py as a workaround. The flaw is in how it builds the UTRs. First, merge_features adds a StringTie exon to the gene whenever that exon overlaps a BRAKER CDS and is not shorter than it — so a StringTie exon a few bases longer than the coding exon it covers (one that pokes into the flanking intron) gets merged in. Then compute_utr_features walks each exon and, for the single CDS segment it overlaps, carves whatever sticks out into a UTR:

cds_start = int(overlapping_cds[0][3])   # the LOCAL overlapping CDS, not the gene's coding start
if strand == "+":
    if start < cds_start:                  # exon pokes out 5' of THAT CDS segment
        utr5 =  "five_prime_UTR" 
        utr5 = utr5.replace(str(end), str(cds_start - 1))
    if end > cds_end:                      # … or 3' of it
        utr3 =  "three_prime_UTR" 

The comparison is purely local — an exon against the one CDS it happens to overlap — with no check that the exon is the transcript’s first (or last) one, nor if the resulting UTR falls outside the gene’s overall coding span as it should.

Apparently there is now BRAKER4, but unfortunately it seems like the UTR prediction is still done through stringtie2utr.py. While improved on many fronts, merge_features still merges over-long StringTie exons, and the UTR carving still uses only local exon-to-CDS comparison — so the root cause is unchanged. This is actually intentional as the author mentioned in one of the related issues that since the gene structure is ab initio, why should we trust it over RNAseq evidence? Btw I just came across this deep learning gene prediction tool called Helixer that may be worth trying.

The fix

So now we need to fix two things:

  1. UTRs that fall within coding regions.
  2. Bare gene/transcript ID.

We can write a post-processing script to fix both. See postprocess_braker_gtf.py as an example. Or we can fix the UTR part in string2utr.py, see an updated version in my git repo. Along the way I found a third problem from the raw braker output. There is always an redundant mRNA line for each transcript that was generated from GeneMark, and it does not include the UTRs that were predicted. postprocess_braker_gtf.py will simply remove those lines.

Do we have to redo the whole annotation? No.

A reasonable worry at this point is that fixing the GTF means re-running everything from scratch. It doesn’t. All three cleanups touch only the gene, transcript, mRNA, and UTR lines — they never modify a CDS feature therefore the cds and proteins are unchanged, thus the annotation. However we do need to update the SnpEff database and re-run the annotation of variants.

]]>
Screen http://fanhuan.github.io/en/2026/06/15/Screen/ 2026-06-15T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/15/Screen I usually use screen to manage parallel tasks so I can keep track of the cmd I used for each task. Of course you should also keep them either local or in a notebook somewhere in case your machine is restarted and you will lose all your screens all at once. Sometimes I have more than 10 screens and I lose track of them. Usually I will do screen -r to list out all the screens so I know the exact name of the screen that I’d like to attach to. Recently, I’ve run into the situation that screen -r would just hang. For the screens whose names I can remember, there was no problem. I can attach them by doing screen -r abc. So what is going on?

After diagnosing with AI, it turns out that one of my screens was in the T status, or was stopped. I do remember stopping jobs within that screen, but I don’t remember and don’t know how to stop a screen. But anyways, since screen -ls was also hanging, there are two helpful cmds that you can use to see what is going on.

The first one is ls -la /run/screen/S-$USER. This one will allow you to see the full names of all the screens that you have started. This way, if the screen you want to go back to is OK (as in not in T status etc.), you will be able to see their full names and attach back by doing screen -r abc.

Of course we also want to identify the root cause. This is a snippet of code that AI asked me to run:

for s in /run/screen/S-$USER/*; do
  p=${s##*/}; p=${p%%.*}
  st=$(cut -d' ' -f3 /proc/$p/stat 2>/dev/null)
  wc=$(cat /proc/$p/wchan 2>/dev/null)
  printf '%-8s %-22s %-4s %s\n' "$p" "${s##*/}" "${st:-GONE}" "$wc"
done

This would retrieve the PID of the screens and return their status. In my case, all my screens were in the mode of S/do_select (sleeping on the socket waiting for a client) except one being on T/do_signal_stop. This means the screen daemon was hit with a job-control stop signal (SIGSTOP/SIGTSTP) and is suspended. What to do? You just need to resume this job by kill -CONT $PID. Note that the PID is the number attached to the front of the screen name when you created it.

After resuming this job, you will be able to do screen -r and screen -ls without hang.

So screen has created a lot of problem for me so far. Something it gets stuck, and you can use ctrl + A + Q to exit.

I haven’t decided whether to move to tmux.

]]>
Fine Mapping http://fanhuan.github.io/en/2026/06/03/Fine-Mapping/ 2026-06-03T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/06/03/Fine-Mapping What is fine mapping

Due to LD, a lot of SNPs in the same region will be showing the same genotype-phenotype correlation. Fine mapping is the process of narrowing down to the causual variant by distinguishing the hitchhiking ones.

Why do we need fine mapping?

I can see two senarios.

  1. When we have WGS data, it is usually not necessary to test very variant (precisely due to LD). Testing SNP A will give you almost the same results from testing SNP B if they are tightly linked, or in LD. You can do some LD pruning and test the representatives. However, the representatives were chosen at random and you might have actually removed the causual variant from the testing dataset.

  2. Even if we have tested every single variant, again, you will need a way to distinguish the causual ones versus the hitchikers.

So one thing worth pointing out is that the fine mapping will be carried out on the full dataset.

Tools to use

Currently I am using “Sum of Single Effects” (SuSiE). It’s R realization is called susieR. The original model is described in Wang et al. 2020. This year, a newer version called MultiSuSiE where multi-ancestry is accomodated was publised.

]]>
The Hidden Importance of Founders in PLINK Analysis http://fanhuan.github.io/en/2026/05/15/Plink-Founders/ 2026-05-15T00:00:00+00:00 Huan Fan http://fanhuan.github.io/en/2026/05/15/Plink-Founders Recently I starting doing family-based GWAS using SNIPAR. This means I need to know the relationship between the samples in my analysis. Previously I only have info on two families which makes the majority of the data that I am working on, and I just treated the rest as un-related. But I know that is not true. In order to increase the sample size, I used KING, a kinship inference tool to predict the possible relationships based on SNP data. Then I check with the breeders to see whether they agree with those relationships. So now in my dataset, a lot of individuals have derived hypothetical PID or MID (parental or maternal ID), just to suggest their full or half sibling relationships.

Then I just went ahead to do my usually data preparation using PLINK until I realized some problem, and it centers around this concept called founder.

It is basically anyone with 0 0 in the PID (column 3) and MID (columns 4) of the .fam file. Meaning, we do not have information on who their parents are. Thus they are founders themselves. Since they might not be the actual founders from their population, therefore it is Not a biological concept but purely a pedigree bookkeeping artifact. If there is no pedigree info in the whole dataset, then everyone becomes a founder. In our family data where we have grandparents, F1 and F2, only the grandparents are founders.

2. Why founder matters

By default, PLINK calculates allele frequencies based on founders only. Meaning, if we do a --maf 0.05 filtering, say for variant chr1_10000_A_T, in the founders it is all A, but maybe there are a lot of copies of T in non_founders, this variant will still be considered not meeting the maf cutoff and filtered out. This makes sense when we do have the parents or grandparents in the dataset, since mendelianly they should have all the alleles of their offsprings. But in my dataset, due to the include of hypothetical PID and MID, it would be a huge lose if only “founders” are considered. Using only founders approximates sampling independent chromosomes from the base population.

Beyond presence and absense of alleles, this is also related to how allele frequencies in this population should be calculated. Allele frequency estimation assumes you’ve drawn N independent chromosomes from the population. Since related individuals share alleles IBD, they are not independent observations. If you genotype a parent and then genotype their three children, you’re partly re-counting the parent’s alleles three more times — the children’s genotypes are predictable from the parent’s. The “effective sample size” is the count of independent draws, which is far smaller than the raw count. Using raw count makes you think your estimate is more precise than it is, and it lets a few large families dominate. This concept is realted to base population that we talked about before. So founders makes the base population.

At that point, I thought it only affects certain plink functions such as --maf or --hwe. Not until today did I realized that by default, any feature of PLINK is based on the base population or the founders. OK so the first conclusion of today is, in PLINK, founders are the base population.

3. Analyses silently affected by founder status

Basically any analysis. You need to be very careful about whether you want to just use the founders (if your pedigree in the .fam file is correct), or all the individuals (turn on --nonfounders). Sometimes you also do not want to do the latter if your dataset is heavily biased by some families like I do. Here is a limited summary table for features I usually use. But again, only founders are used for allele freq calculation by default for any featuer, any!

Flag What uses founders Consequence if few founders
--freq Frequency computed from founders only Inaccurate MAF
--maf Filters based on founder frequencies Wrong variants removed/retained
--hwe HWE test on founders only Underpowered or wrong results
--pca (PLINK 1.9) GRM built from founders only Fails if N_founders < 20 or has duplicates
--pca approx (PLINK 2) Allele freqs from founders Hard error if N_founders < 50
--indep-pairwise / LD pruning r² computed from founders only Over-pruning when few founders (spurious LD from small N)
--genome / IBD Uses founder allele frequencies Biased IBD estimates

4. How did I discover this silent scary behavior?

  1. Like I said in the beginning, after adding all those PID and MID, there are very few founders left in my dataset, and I noticed that a lot more SNPs were filtered out under the same --maf.

  2. Then I realized that it also affects the LD prunning because under the same parameters (--indep-pairwise 500 50 0.8 ), higher percentage of SNPs were found in LD/heavier prunning.

  3. Eventually, PCA failed:

  • PLINK 1.9 --pca: silent failure with cryptic GRM error (“Failed to extract eigenvector(s) from GRM”), probably a singularity problem.
  • PLINK 2 --pca approx: explicit error (“less than 50 founders available to impute allele frequencies”)

Both errors have the same root cause: the GRM and allele frequency estimation are operating on fewer than 50 individuals for a dataset with thousands of samples.

5. Solutions and tradeoffs

--nonfounders: Usually this is an easy problem to fix by turning on this option and use all individuals in the dataset. This indeed retained slightly more SNPs (less than 10%) using all thousands of individuals, however still significantly less than the previous batch with only hundreds of individuals.

  • --freq + --read-freq: pre-compute frequencies from a representative subset, then feed them in — most principled for mixed datasets. Four-step workflow:
    1. Pre-filter without --maf (apply --geno and --mind only)
    2. Define a representative subset: include all true founders (PID=0, MID=0) plus one individual per unique (PID, MID) pair among non-founders. This ensures every independent lineage contributes exactly once — full siblings collapse to one representative, but half-siblings (who share only one parent and thus have different (PID, MID) combinations) each get their own representative.
    3. Compute frequencies from that subset: plink --bfile ... --keep <subset> --freq --nonfounders --out ...--nonfounders is required here because the subset includes non-founders (e.g., the half-sib representatives); without it PLINK falls back to the 27 true founders.
    4. Apply MAF filter using the pre-computed frequencies: plink --bfile ... --read-freq <freq_file> --maf 0.005 --make-bed --out ...
  • Remove relatives first for LD pruning: use --rel-cutoff (can try third degree: 0.125 or second degree: 0.25) + --make-founders (required when parents are absent from the kept subset) + --indep-pairwise; apply the resulting prune list to the full dataset. For highly structured multi-population datasets, population structure will still inflate LD — per-population pruning followed by taking the union of kept variants is the most principled approach.

    A note on consistency between MAF and LD representative selection: It is natural — and correct — to use different criteria at the two stages. For MAF estimation, the pedigree-based approach (one per unique PID/MID pair) is optimal because it uses known family structure to ensure independent lineage representation; half-siblings are included because their distinct (PID, MID) pairs represent genuinely different crosses. For LD estimation, a kinship cutoff (e.g., 0.125) uses empirical relatedness to prevent shared haplotype blocks from inflating apparent LD; half-siblings (IBD ≈ 0.25) are excluded by this threshold. The LD stage being stricter about relatedness than the MAF stage is the safe direction and is not a methodological inconsistency.

  • --bad-freqs: override (not recommended — hides the problem)

Attempt 1 — default (a couple dozens of founders): retained only ~4.5% of variants vs ~12.4% for a previous version where we assigned hundreds of founders. Noisy r² from small N causes spurious high-LD calls and over-pruning.

Attempt 2 — --nonfounders (all individuals in thousands): retained even fewer variants (~4.2%). This is counterintuitive — more individuals, yet worse results. The explanation requires understanding two distinct sources of r² inflation:

  • Attempt 1 suffers from small-N noise: with only ~27 individuals, r² estimates are imprecise and systematically upward-biased (r² is bounded at 0, so random errors can only push it higher, never lower). Some truly unlinked variants get flagged as in LD by chance.

  • Attempt 2 suffers from kinship-induced pseudo-LD: related individuals share long IBD haplotype blocks. Two variants sitting on the same shared haplotype will co-occur systematically across all members of a family — not because of actual LD in the population, but because of shared ancestry. Within a pruning window, PLINK cannot distinguish this from real LD and prunes accordingly. This is especially bad when you have a lot of related samples in your dataset.

In my case, the kinship inflation turns out to be larger than the small-N noise inflation, so going from a couple of dozens of founders to thousands of related individuals makes things worse. Therefore we need to remove relatives first — you need a dataset where r² reflects actual population LD, not shared ancestry.

At first I tried to get a unrelated subset using --rel-cutoff 0.125 , but again only a couple of dozens of individuals are left. leaving 14 — worse than the original 27 founders. This is because 2nd-degree relatedness is pretty common in my dataset. Then I tried a lower cut off --rel-cutoff 0.25 (remove only 1st-degree + duplicates), now we have a few hundreds remaining. You then need to make all of them founders (--make-founders)

Attempt 5 — add --make-founders: promotes all individuals with absent parents to founder status. This is necessary whenever you use --keep to subset a pedigree dataset. Still retained fewer variants than expected (~3.1%), because population structure (many divergent populations) inflates within-window r² regardless of relatedness.

Validation: despite all this, PCA eigenvectors computed before and after LD pruning showed >0.99 correlation — confirming that for PCA, the exact pruning strategy matters little in practice.

5.5 A second worry, a wrong turn, and what the lever actually is

After all that founder agonizing, I hit a related worry. My dataset is dominated by two big populations (TS1 and TS3), with a bunch of smaller ones trailing behind. My fear: even with --nonfounders turned on, a global --maf 0.005 is computed by pooling everyone together. So a variant that is common inside a small population but rare across the whole pool gets dropped — exactly the variants I thought I’d most want to keep.

My first instinct was that --maf was the wrong tool and I should switch to a count threshold, --mac. The reasoning felt clean: what actually destabilizes an association test is the minor allele count — how many copies enter the regression — not the frequency, so filter on the thing you actually care about. I was fairly convinced. So I ran it.

It made almost no difference. --mac 20 --nonfounders returned essentially the same variant set as --maf 0.005 --nonfounders (a hair fewer, in fact). And once I saw that, the reason was obvious and a little embarrassing: on a single pooled sample, a frequency is a count. With ~2000 samples, --maf 0.005 means a minor allele count of about 0.005 × 2 × 2000 ≈ 20. So --maf 0.005 and --mac 20 are the same threshold written two different ways. They can only diverge at the boundary, and on how each treats missingness (--mac is slightly stricter on high-missingness sites, which is why it kept a touch fewer). Switching frequency-for-count could never have fixed population imbalance — I’d been comparing a tool to itself.

So what is the lever? It’s the denominator — who counts as the base population — not the form of the threshold. That’s the whole lesson of this post, and it’s the one knob that actually moves variants in and out:

  • founders only (a couple dozen people): noisy estimate, over-removes — the broken case.
  • all individuals (--nonfounders, or equivalently --mac on everyone): the pooled frequency. Repairs the over-removal.
  • one representative per independent lineage (the --read-freq subset trick from section 5): weights each lineage once, so the pooled denominator no longer drowns out small populations — retains the most variants.

That last one looks like the answer to my imbalance worry, and as an estimator of allele frequency it is the principled choice. But here’s the catch I only saw after running everything: the extra variants the lineage-weighted set keeps are, by construction, the ones with very few actual copies in the full sample. They survive only because dividing by a small denominator inflates their frequency. For a pooled GWAS, those are exactly the underpowered variants — there genuinely aren’t enough copies in the data I’m analyzing to test them stably.

Which dissolves the original worry rather than solving it. In a pooled analysis, a variant that is rare in the pool is untestable in the pool — no matter how common it is inside some small population. That isn’t a filtering bug to engineer around; it’s a property of pooling. If those small-population variants are biologically interesting, the answer is a stratified or population-specific analysis (where you’d filter within that population), not a cleverer global filter.

So my actual conclusion, after the wrong turn: for everything analyzed together in GCTA and SNIPAR, use all individuals as the base population and a stringency around --maf 0.005 / --mac 20 (they’re the same thing — pick whichever you find clearer; --mac is marginally more honest about missingness). For SNIPAR’s family-based tests, where the effective number of independent units is smaller than the raw N, leaning a bit more conservative (--mac 30) is reasonable. Reserve the lineage-weighted subset for when you want an unbiased frequency estimate, not for deciding which variants enter the test.

And the meta-lesson: I almost shipped a fix to a problem the fix couldn’t touch, because the reasoning sounded right. Running it was what corrected me.

6. Key takeaway

Always check your founder count before running any frequency-dependent analysis:

grep "founders" your.log

If you have a pedigree-filled .fam file and few founders, every downstream result is quietly wrong unless you intervene. The --hwe case is worth special attention: HWE violations are expected in related samples, so filtering on HWE in a pedigree dataset silently removes valid markers.

]]>