A preprint published on Aug. 25, 2026 in bioRxiv (Unexpected D-tour Ahead: Why the D-Statistic, applied to Humans, Measures Mutation Rate Variation and Not Neanderthal Introgression by William Amos and Eran Elhaik, bioRxiv 2024.12.31.630954; doi: https://doi.org/10.1101/2024.12.31.630954) suggests that the tool used to evaluate admixtures, introgression and sharing of alleles between sister species may be flawed.
The abstract states that D statistic "relies on two untested assumptions: mutation rate constancy between groups and a negligible contribution from recurrent mutations" in the paper, the authors argue that both assumptions are invalid!
They also report that "heterozygosity modulates mutation rate" and that it is this, the difference in heterozygosity that is the critical factor.
So, first of all, we will try to explain the D-statistic.
D-statistic, what is it?
The D-statistic or ABBA-BABA test was created by mathematician Nicholas Patterson (b. 1947) to evaluate if there had been gene flow between Neanderthals and humans (Patterson, 2012). It has been used since then to study admixtures and introgressions in different species.
This test looks into four populations, P1 and P2 that are closely related, P3 also on the same tree, but on a different branch, and a distant outgroup, O. The computer program, because that is what it is, compares the alleles of each population, which can come in two variants: A, the "ancestral" ones, found in the outgroup, and B, the "derived" ones (mutated, or more recent variants). There can be two patterns. One is the "ABBA" pattern, where P1 and the outgroup share the A (ancestral) variant while P2 and P3 have the mutated or derived B variant. The other is the "BABA" pattern, where P1 and P3 have the derived "B" variant while P2 and the Outgroup share the ancestral "A" allele. See the image below as a reference.
The outgroup, helps discriminate between ancestral (A) and derived (B) alleles.
ABBA and BABA patterns should appear with equal probability if there is no interbreeding, which results (see the formula below) in a value of zero for the D-statistic. (the numerator of the equation would be zero).
The reason for this is something called Incomplete lineage sorting or "ILS", which assumes that the ancestral alleles are distributed in a random manner among the descent of a species as it evolves and branches into new ones, so it is possible that some descendants inherit some alleles but not others. Interbreeding could add alleles to a branch that hadn't inherited them.
Therefore, if there has been gene-flow (introgressions), there will be a higher number of ABBA patterns than BABA ones, signfiying a positive value for D, and therefore, hybridization, with gene flow between P2 and P3.
In the equation, C stands for the counts of ABBA or BABA being observed (value = 1) or not (value = 0) at a given site in the genome (i), the symbol Σ means "sum", so you add up all the counts and perform the subtraction in the numerator, and the addition in the denominator.
However, there are some limitations. For instance, there may be an unknown extinct, and therefore unsampled species that introgressed into P3, which has not been included in the tree. Or "population stucture" could lead to discordant readings of the ABBA-BABA sequences. By "structure" it should be understood that a population is not homogeneous, it is composed of different subpopulations with significant differences between them regarding their alleles.
This example may help clarify the subject: P1 is Sub-Saharan African San, P2 is an Englishman, P3 is a Neanderthal, and O, the outgroup is a chimpanzee. So, we will look at different sites and identify which base is found in the DNA:
But, apparently this method is flawed.
The incorrect assumptions
Amos and Elhaik argue that the two basic assumptions of the D statistic are wrong. Let's follow their logic:
- Mutation Rate is not constant. In genetic studies, it is taken for granted that mutation rates are invariable. They don't change. However, research has shown that this is not the case. Genetic sequences from family lineages have shown that mutation rates vary. Furthermore, the authors argue that mutation rates vary between primate species, human populations, and along chromosomes.
- Recurrent Mutations are common. Genetics assume that mutations in the same site are extremely rare. But this paper propses that they are far more common than we believe. A given single nucleotide polymorphism, or SNP, may flip from one to another base, with "at least two mutations since the most recent common ancestral base."
Mutation rates are different between Africans and non-Africans
The authors acknowledge that there are different mutation rates between modern Africans and people living elsewhere. They attribute it to enzyms that regulate the replication of DNA or, that heterozygosity affects it (see my post "Mutation rate is Faster in Africa"). So, the Out of Africa event, with its bottleneck reduced the heterozygosity of non-Africans and caused a drop in their mutation rate:
"Moreover, we find that differences in heterozygosity strongly predict relative mutation rate, with higher heterozygosity in Africa predicting a higher mutation rate in Africans"
The authors also point out that introgression is almost always found in genes that are subjected to "strong selection", and they reason that "Strong selection appears to impact local mutation rates... and we suggest a simple mechanism whereby selection impacts heterozygosity which in turn drives changes in mutation rate."
I must admit that their case is strong. Mutation rates are variable, and tools like the D statistic may be strongly influenced by facts like recurrent mutations and variable rates. We should not forget that these "statistics" are just models, tools, they are not facts, they are instruments to help us understand how the real world works. All models are fallible and can be improved.
Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall ©







