Translate

Guide to Patagonia's Monsters & Mysterious beings

I have written a book on this intriguing subject which has just been published.
In this blog I will post excerpts and other interesting texts on this fascinating subject.

Austin Whittall


Showing posts with label molecular clock. Show all posts
Showing posts with label molecular clock. Show all posts

Friday, April 24, 2026

Higher mutation accumulation in Europeans vs. Africans


Another paper mentioning different mutation rates in Africans and non-Africans!


The paper by Mallick, S., Li, H., Lipson, M. et al. The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature 538, 201–206 (2016). https://doi.org/10.1038/nature18964, reported that "Our analysis reveals key features of the landscape of human genome variation, including that the rate of accumulation of mutations has accelerated by about 5% in non-Africans compared to Africans since divergence."


The branch-length issue


This reminded me of a recent post I wrote about Short branch lengths in Africans for Y-chromosomes. Branch lengths are linked to accumulated mutations (the lenght of a branch is the number of mutations in it) so if we start from the fork where Africans and non-Africans split, the branch of Africans is shoreter because it has accumulated fewer mutations, while the non-African one is longer, as it has more accumulated mutations. In that post I asked "...but we are all the same age and equally distant from our common ancestor. So why do the Africans have fewer mutations? Do Eurasians accumulate more mutations?"


I also went back to reread a recent post Mutation rate is faster in Africa where I mentioned different studies suggesting that a higher diversity in Africans (measured by their heterozygosity) promoted a higher mutation rate or μ. But Mallick, Li and Lipson et al. in their 2016 paper suggest otherwise. I quote them below and highlight their findings, which seem to baffle them:


"More mutation accumulation in non-Africans than in Africans
The SGDP data provide an opportunity to compare the rates at which mutations have accumulated across populations. We restricted our analyses to samples for which our genotypes are likely to be most reliable ... We pooled samples by region to increase power, and for all pairs of regions, computed the expected number of positions where, if we picked a random chromosome from both, region A would mismatch chimpanzee and region B would be identical to chimpanzee (or vice versa). If the rate of accumulation of mutation has been the same since the two populations diverged, these numbers are expected to be equal. However, when we compute the ratio of mutations on one lineage or the other since separation, we find a subtle (average of 0.5%) but significant excess of mutations in nonAfricans relative to sub-Saharan Africans. Because any difference must reflect events since non-African / African population divergence which is a less than a tenth of average genetic divergence, this implies a greater difference in mutation accumulation rates since population divergence (~5%). We were concerned that these results might be biased by the fact that the human genome reference sequence is more closely related to non-Africans than to Africans, or by higher levels of heterozygosity in Africans, as both these issues could make detection of divergent sites in Africans more difficult. However, we replicated the findings after remapping to chimpanzee, which is equally distant to all present populations, and after restricting analyses to the X chromosome in males (males only have a single X chromosome, and so this procedure avoids bias due to different error rates in detecting heterozygous genotypes in populations with different rates of heterozygosity). These observations are most likely to be explained by acceleration in the rate of mutation accumulation in non-Africans, since the same signal appears in comparisons to sub-Saharan Africans related in different ways to non-Africans. It is known that the rate of CCT>CTT mutations differs across human populations. However, this particular mutation class was found to be enriched relative to Africans in Europeans but not in East Asians, and thus cannot explain our signal. One of several possible explanations for these findings is a decrease in the generation interval in non-Africans compared to Africans since separation...
"


What is going on?


Mallick, S., Li, H., Lipson, M. et al. affirm that mutations accumulate at a higher rate in non-Africans than within Africa, this does not mean that mutation rates (μ) are different, it means that the mutations are fixed differently. It is also higher among Europeans than East Asians. To explain it they propse that non-Africans have a shorter "generation time": they are mating at a younger age than Africans so they accumulate more generations in a given span of time, and therefore more mutations than Africans in the same period.

Generation Time

I am surprised because other research has shown that longer generation times lead to more mutations because "each additional year of paternal age results in an average of 3.9 × 10−10 more mutations per base per generation (Source) and hunter-gatherer people nowadays have longer generation ages (32.3 years for fathers) than sedentary groups and because older fathers accumulate more mutations, if Africans are mating at an older age, there will be more mutations. There is something that isn't adding up here!


A similar viewpoint was reported by Wang and Obbard, 2023: "Our analysis also shows that mutation rates increase significantly with increasing generation time... The relationship we observe between generation time and per-generation mutation rate could therefore be a consequence of either a greater number of cell divisions or of accumulating damage over time." Which makes sense.


However, not all agree, research conducted by Lewin and Eyre-Walker, 2025 confirms that mutation rate (μ) and generation time are inversely correlated (longer generation time = lower mutation rate; and shorter generation time = higher mutation rate).


Due to these conflicting findings, until consensus is reached, for the time being I will leave generation times out of the matter and look for other plausible causes for the higher number of mutations in non-Africans vs. Africans.


Other causes explaining higher accumulation of mutations


Effective population size or Ne. Wang and Obbard, 2023, also notice that "populations with larger Ne tend to have a lower mutation rate even after accounting for their shorter generation times." Africans are said to have a larger initial Ne due to the bottleneck effect that affected those leaving Africa, a small subset of the large original population. The authors suggest that if "... species with small Ne tend to have a longer generation time, and a longer generation time causes higher mutation rates then a higher μ in species with low Ne could be driven by a mechanistic generation-time effect." Since Eurasians seem to have both factors (small Ne and low generation times) is seems logical that their mutation rate is higher.


Besides the effective population size and generation time, there are more possible explanations for the shorter branch in Africans and the longer one in Eurasians are: natural selection, that removes noxious mutations, so looking back from the present, the mutations never seem to have taked place because they were not fixed. Reduced DNA repair mechanisms, one group has a less efficient repair mechanism for mutations and these tend to accumulate in comparison to another group with a more efficient repair system.


The increased TCC→TTC mutation rate in Europeans


This factor is mentioned in Mallick, S., Li, H., Lipson, M. et al. as a possible explanation. But, what does it really mean? I will quote from Harris and Pritchard, 2017, who studied the matter.


Our DNA is made up of two backbones, the intertwined helixes linked by "steps" like a ladder, made from bases called Guanine (G), Cytosine (C), Thymine (T) and Adenine (A). Guanine always links to Adenine G—A, and Thymine with Cytosine (C—T) bonds. Looking at the steps of one side of the helix you will see a sequence like "ATCGATTGAGCTCTAG", and opposing it, on the other strand the complementary bases: "GCTAGCCAGATCTCGA".


Research has shown that "European people experience more mutations within certain DNA motifs (specifically, the DNA sequences ‘TCC’, ‘TCT’, ‘CCC’ and ‘ACC’) than Africans or East Asians do." Why?


Harris and Pritchard propose that "the rate of TCC→TTC mutations increased dramatically ∼15,000 years ago and decreased again ∼2000 years ago... [and] hypothesize that this mutation pulse may have been caused by a mutator allele that drifted up in frequency starting 15,000 years ago, but that is now rare or absent from present day populations." They go on to explain the cause: " At this time, we cannot exclude a role for nongenetic factors such as changes in life history or mutagen exposure in driving these signals. However, given the sheer diversity of the effects reported here, it seems parsimonious to us to propose that most of this variation is driven by the appearance and drift of genetic modifiers of mutation rate."

So it seems that it is due to a chance appearance of genes that regulate mutation rates.


A curious yet interesting fact is that the same effect of TCC→TTC mutation increase is observed in East Asian cattle! It appeared in two separate mammal groups, indicine cattle, derived from the Bos taurus indicus and humans but outside of Africa (Talenti, et al., 20216)


A challenge to the "stable" molecular clock


Harris, 2015 also looks into the TCC→TTC subject and says that explaining the cause is beyond the scope of the paper. However, Harris concludes that "Even if the overall European mutation rate increase was small, it adds to a growing body of evidence that molecular clock assumptions break down on a faster timescale than generally assumed during population genetic analysis. It was once assumed that the human lineage’s mutation rate had changed little since we shared a common ancestor with chimpanzees, but this assumption is losing credibility due to the conflict between direct mutation rate estimates and molecular-clock-based estimates. Although this conflict might have arisen from a gradual decrease in the rate of germline mitoses per year as our ancestors evolved longer generation times, the results of this paper indicate that another force may have come into play: change in the mutation rate per mitosis. If the mutagenic spectrum was able to change during the last 60,000 years of human history, it might have changed numerous times during great ape evolution and beforehand."


I agree, mutation rates are variable, and conclusions based on a constant rate will be wrong.


Notice how different papers find opposite effects (faster mutation rates in Africans, or in Europeans), and don't quite understand the reason!



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Monday, March 30, 2026

Western and Eastern Neanderthals Were Very Divergent


New research published last week revealed the genetic makeup of a Neanderthal man (Denisova 17) who lived in the Denisova cave in Altai, 110,000 years ago. The paper reached some surprising conclusions.


This is the paper: D. Massilani,S. Peyrégne, et al. A high-coverage Neandertal genome from the Altai Mountains reveals population structure among Neandertals, Proc. Natl. Acad. Sci. U.S.A. 123 (13) e2534576123, https://doi.org/10.1073/pnas.2534576123 (2026).


Key findings:


  • Population Replacement. The Denisova 17 Neanderthal man was closer to the older (~120 kya) Altai Neanderthal from Denisova (a woman known as Denisova 5) than to more recent European Neanderthals (the Vindija Cave woman, from Croatia and the Goyet female) and to the Neanderthal woman from Chagyrskaya Cave in the Altai region (80 kya, called Chagyrskaya 8 (Chag 8 for short). Suggesting that more recent Neanderthals replaced the older population in Altai.
  • Denisovan admixture. The 120 and 110 kya Neanderthals from Denisova Cave (Denisova 17 and Denisova 5) had Denisovan admixture. This is not seen in the later Neanderthals of Western Europe or Chagyrskaya Cave.
  • The estimated age of the common ancestor of the Y chromosome of Denisova 17 and modern human beings is ~395 ± 44 kya.
  • The Eastern Neanderthals lived in groups of less than 50 individuals, these were smaller, and also more isolated than those of Western Eurasian Neanderthals.
  • Replacement. The Western Neanderthals moved eastwards across Eurasia and replaced the older Eastern Neanderthals. But there was no admixture so the authors conjecture that both populations never met "perhaps because the older Neandertal population disappeared before the younger population from the west appeared in the Altai Mountains... Western-derived Neandertals are thought to have replaced the Eastern Neandertals in the Altai Mountains, and presumably elsewhere, sometime between ~110,000 and ~70,000 y ago."
  • Interaction with modern humans: comparing non-African modern human genes with Neanderthal genes and counting the matches, the authors find that Western Neanderthals, especially the Vindija type were closest to the Neanderthals who admixed with humans. Those matching Eastern Neanderthals were four-times smaller. I wonder if modern humans leaving Africa mated with Western Neanderthals, who may have carried Eastern alleles? or perhaps they also intermingled with Eastern Neanderthals.

Mutation Rate and Divergence


The most interesting part is how divergent both Western and Eastern Neanderthal were, when compared to modern human beings. The paper states that:


"... it is striking that the allele frequency differentiation between Eastern Neandertals (D5 and D17) and Western Neandertals (Vi33.19 and others) (FST = 0.30, 95% CI: 0.29 to 0.31) exceeds that of even the most differentiated pairs of present-day populations, such as the Mbuti of Central Africa and the Papuan Highlanders of New Guinea (FST = 0.27, 95% CI: 0.26 to 0.27). The divergence between Mbuti and Papuan is estimated to have occurred 130 to 220 kya, resulting in separate genetic drift along the two lineages over 260 to 440 ky. The divergence between Eastern and Western Neandertals occurred about 35 ky before D5 and D17 and about 80 ky before Vi33.19 lived resulting in separate genetic drift along the two Neandertal lineages over about 115 ky. This suggests that Neandertal populations reached greater levels of differentiation over shorter timescales than modern humans did. This is also illustrated by the modest differentiation between a ~45,000-y-old modern human genome from Siberia (Ust’Ishim) and present-day populations (FST = 0.052, 95% CI: 0.049 to 0.056)."


The authors suggest that the small bands of Neanderthals underwent higher genetic drift, that led to high frequencies of different alleles in the two populations. So even though they were not far apart in Eurasia, they were, nevertheless isolated from each other. The point is that "Neandertal populations accumulated allele frequency differences more rapidly than the ancestors of present-day human groups." So much for the molecular clock and its regular ticking rate.


In my previous posts on the MUC19 allele that introgressed into humans from Denisovans via Neanderthals wo were not related to the Altai Neanderthal (Denisova 5), I wondered if the Neanderthals who were ancestral to the Chagyrskaya and Vindija individuals but not to Denisova 5, the Altai Neanderthal were the ones who had admixed with modern humans. This paper seems to suggest it was them.


If there was an Early peopling of America by Neanderthals, it was surely done by the Older Eastern Neanderthal Group.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Sunday, March 22, 2026

Timeline, and age of the Y Chromosome


This post will look into the "age" or timeline of our Y chromosome. There are different methods used to estimate how this chromosome changes over time by accumulating mutations (*). This leads to a mutation rate expressed in mutations per site per year. It can be calculated by comparing the Y chromosomes of two different men, or a modern human and an archaic one, or another ape and humans, and count the mutations by which they differ, if we know when their lineages split.


(*) Actually, to be correct, a mutation is a chance change in a base pair, an error in the transcription of our genetic code. A substitution is when a mutation becomes fixed within a population, one that has not been erased by genetic drift, or natural selection. For instance, a chance mutation may lead to a "C" to appear instead of a "G", but and then another chance mutation could flip it to a "T" in that case looking at the original and the final versions we'd see one substitution, and imagine one mutation, but there were 2 mutations. Another example would be a flip back, to "C" (back-mutation or reversion), where we'd see no mutation or substitution, but in fact, there was one mutation.


The formula to calculate the Time to the Most Recent Common Ancestor (TMRCA) is shown below. Where the TMRCA in years; k is the number of base pair differences between both men, 2 is a factor added because the difference corresponds to the divergence in both men, half for each one and we don't want to count them twice, L is the total sequence length sampled in which k was detected; and μ is the mutation rate.


TMRCA (years) = k (mutations) / μ (mutations/site · year) · L (sites)· 2


From which the mutation rate can be calculated by shuffling terms as shown:


μ (mutations/site · year) = k (mutations) / TMRCA (years) · L (sites)· 2


A real life example comparing Chimpanzee and Human base pair divergences found in Shen et al., (2000) is the following: assumed TMRCA: 4,900,000 years, acutal values measured (the length sampled in base pairs) L: 38,568, (the number of divergent base pairs) k: 470, calculation of μ shown below:


μ (mutations/site · year) = 470k (mutations) / 4,900,000 (years) · 38,568 (sites)· 2


μ (mutations/site · year) = 1.25 × 10−9.


Assumptions and Data


As you can see, the formula is logic, and straightforward, we take a sample of a DNA strand from a Y Chromosome with a length of L base pairs (my intro to Y Chromosome post explains the basic terms) in both subjects, and then count the differences we observe (k), knowing how long ago both individuals shared their last common ancestor (TMRCA) we calculate the mutation rate μ.


The same method can be applied to any chunk of nuclear DNA including the X chromosome and any of the non-sexual chromosomes or autosomal DNA, as well as the mtDNA. The interesting part is that they all mutate at different rates.


Y chromosome mutation rates (μ)


Over the course of the years, different studies using different methods have attempted to calculate the mutation rate of the Y chromosome in humans. The table below shows some of these studies. The values of μ are given in mutations per site per year. For non-scientists, note that expressing a number multiplied by 10-9 is another way of writing: "divided by 109" and 10 to the ninth power is 1000,000,0000. This means that 1.24 x 10-9 = 1.24 / 1000,000,000 = 0.00000000124 (very small indeed!).


The difference between the first value reported by Thompson et al, and the figure given by Francalcci et al is 2.34 times!


There are big discrepancies in these μ values


Using one or the other to estimate the age of a given specimen would result in one figure being over twice the age of the other one.


The different methods employed to calculate μ involve different assumptions and they all have their shortcomings. In the following commentary I will follow the excellent work of Wang CC, Gilbert MT, Jin L, Li H. (2014) ( Evaluating the Y chromosomal timescale in human demographic and lineage dating. Investig Genet. 2014 Sep 10;5:12. doi: 10.1186/2041-2223-5-12. PMID: 25215184; PMCID: PMC4160915).


The data used for comparing divegencies in the Y chromosome's base pairs can come from different sources: Ancient DNA, samples taken from prehistoric human remains which have been dated by using radiocarbon or other methods. Genealogical, using samples from a certain family for which the genealogy has been confirmed and dated, these usually involve short time spans and few generations. Archaeological Events, taking an estimated date for an event, such as the peopling of America or the settlement in a given region in Europe, and applied to samples from that date.


Confounding Factors


As mentioned in my post on phylogenetic tree branch lengths (a factor that shows that mutation rates are not constant), the value of μ should be considered as an estimation, and not something written in stone. Trombetta et al., (2015) point out that "variants" (mutations) appear at different rates across Y chromosome haplogroups, geographic locations, and time: "... we observed a remarkable heterogeneity in the distribution of variants, not only across different regions, but also across lineages and different times. Hg A00 stands out for showing strong associations with almost all genomic features considered. In the rest of the tree Hg's A0, A1a, A2'3 and B differ from Hg's DE, FC and R, and ancient branches differ from recent ones. The two levels are not entirely independent, as far as recent branches are enriched in lineages belonging to Hg's DE and R. It is possible that different social habits, lifestyles and environmental conditions experienced by populations harbouring different haplogroups resulted in systematic variations of the generation time and average paternal age at conception." They attribute this to older or younger age of the fathers when their children are conceived —older dads have more mutated sperm, and the effects of environment on DNA mutations.


Chimpanzees and Men


When comparing human and chimp Y chromosomes, we are not only separated by a gulf of 5 to 7 million years of separate evoluton, the evolution itself has been different in both species. The chimpanzee Y chromosome is much smaller than that of humans. it lost roughly one-third of its genes in the MSY, or male-specific Y region of their Y chromosomes, compared to men.


Kuroki, Y., Toyoda, A., Noguchi, H. et al. (2006) noted that there is a greater divergence in the sequence of human Y chromosome vs. chimpanzee Y chromosome than between the whole genomes of both species (1.78% and 1.23% respectively).


The reliable dating of the Chimpanzee-Human split is still being debated, and figures range from 4.2 to 12.5 million years ago. A factor of three!


There are also structural differences in the shape of our and the chimp's Y chromosome which complicates the alignment of segments for comparison. Finally, chimpanzees and modern humans have different pair-bonding sexual behaviors (Hughes et al, 2013). Schaller et al., (2010) note that receptive females copulate with multiple male partners creating selective pressure towards the male fertility genes in the chimp's Y chromosome. Monogamous pair bonding in humans lacks this intense selective force. This alters the rate of the mutational clock. The pair bonding in hominins can be seen in the reduced sexual dimorphism in australopithecines and the loss of sperm competition adaptations, suggesting less male-to-male strife, and the growth of cooperation as a means for reproductive success (Gavrilets, 2012).


Genealogical Methods


Pedigree-based studies look at men belonging to the same family, sharing the same paternal lineage, and whose birth dates are known, or at least, how many generations separate them. This provides a well defined dating.


Xue et al., (2009) studied 13 generations of men in haplogroup O3a. However, there are some potential sources of error in this method: Since mutation rates are random, variable, and unpredictable (the statistical term for this is highly stochastic), are we sure that 13 generations (approx. 390 years) is a long enough interval.


Another is the haplogroup itself, we have mentioned that haplogroups accrue mutations at different rates. Is haplogroup O3a a good reference for all other haplogroups?


Finally, if selection and genetic drift act mutations eliminating some of them over longer timescales, the number of mutations found in a genealogical study will be smaller than the actual one detected in longer scales.


Mutation Rates adjusted for autosomal mutation rates


This method was developed by Mendez F., et al., (2013) when they dated the extremely ancient A00 haplogroup Y chromosome, found in the Mbo people of Cameroon, Africa. The μ used by Mendez team was based on a study conducted in Iceland that calculated the autosomal mutation rate by analyzing the divergence between parents and their children. This implies several unverified assumptions: autosomal and Y chromosome mutation rates are similar (they are not), substitution rates and mutation rates are equivalent (they are not). They also used a generation time that spanned from 20 to 40 years when life expectancy for men in Cameroon in 1950 was below 40! (Source). Other authors using genealogical data, like Boattini et al., (2019) found generation lengths of 33.57 years. Hunter-gatherer groups in Equatorial Africa 150,000 years ago probably began mating at the age of 15. How can we know for sure? (See Elhaik E, Tatarinova TV, Klyosov AA, Graur D., (2013) and their critique to Mendez et al.


Older generation times lead to more mutations, and an overestimated TRMCA. A00's age is far shorter than the one put forward by Mendez et al.


Archaeological Data


In my posts, I have mentioned many ancient DNA samples taken from the remains of prehistoric, ancestral human beings and hominins (Yana River, Mal'ta and Ust'-Ishim). They provide a certain age with reliable radiocarbon-dating, and if the DNA is not degraded or contaminated, a reliable count of mutations.


Fu, et al., (2016) compared the remains of the Ust'-Ishim man with those of men alive today, and looked for "missing" mutations, those that appeared in modern men after Ust'-Ishim died. The team calculated a mutation rate of 0.76 × 10-9


Human Migration approximations


Assuming dates for certain migratory events, like the peopling of America, placed at around 15,000 years ago, the dates for certain haplogroups found exclusively in America can be set close to that event. This has been used to calculate μ for splits between Asian and Amerindian lineages, or between Amerindian groups within America, obtaining a μ of 0.820 × 10-9 (Poznick et al., (2013)).


A similar method was applied by Francalacci et al., (2013) to men native to the island of Sardinia in Europe, peopled around 7,700 years ago, and using the mutations detected in a sample of Sardinian men to calculate μ for their haplogroup (I2a1a). The value obtained was very low compared to those shown in the Table further up: 0.530 × 10-9.


Can we be certain that the haplogroup diversified in its current location in Sardinia or America? Could it have taken place earlier (European mainland, or Siberia, respectively). Do the current Sardininans belong to the group that reached the island 7.7 kya? Or did they arrive later?


Implications


The sum of these factors show that environment, generation times, sexuality (pair bonding), natural selection, and genetic drift can promote or reduce mutation rates. Studies have a 2.4-fold variability, ranging from 1.24 × 10-9 Thompson et al., (2000) to 0.53 × 10-9 Francalacci et al., (2013), and these are the mean values, the confidence intervals are even wider. Supposing a threefold difference, what some estimate as a divergence taking place between two Y chromosome haplogroups 50,000 years ago during the final Out of Africa migration, may have taken place within Africa 150,000 years ago! And the A00 split instead of taking place 250,000 years ago may reflect one that ocurred 750,000 years ago.


The possibility that our most ancient ancestors had more chimp-like behavior would have implied a faster mutation rate during that period, followed by a slower rate later on. Even pair-bonding during the patrilineal, sedentary, agricultural period of the past 8,000 years may have slowed down mutation rates in comparison to the matrilineal hunter-gatherer period that preceeded it.


genetic mutation rates
Mutation Rates and their impact on TMRCA inference. Copyright © 2026 by Austin Whittall

The image above shows how mutation rates can influence the age of the TMRCA. For the same measured value of "m" mutations marked by the gray line, the use of different mutation rates μ influences the depth or age to the TMRCA. A quick mutation rate like the blue one (μ1) accumulates mutations quicker and takes less time to reach the detected "m" mutations, so its TMRCA is younger (T1), a slower mutation rate (red line) like μ2 takes longer to accumulate the "m" mutations and its T2 is longer. Finally, a variable mutation rate like μ3 (green line) that was faster in the distant past, and slower later on will have an even longer and older timeline (T3).


For those interested in maths, the slope of each curve marks the mutation rate dm/dt = μ(t), steeper curves mean quicker mutation rates. It an analogy of speed as the differential of space over differential of time.


I believe that the μ3, variable mutation rate is closer to reality, reflecting variable social-sexual-cultural patterns of the small and egalitarian promiscuous hunter-gatherer groups of early hominins and human evolution. Later settled societies with paternal monogamy and private property led to slower mutation rates. This of course would imply an older root for the human Y-chromosome haplogroups, in Africa, prior to our migration into Eurasia.




Fall has begun here in Buenos Aires, in the Southern Hemisphere! Lucky you who are now entering Spring!



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Friday, March 20, 2026

Short Branch Lengths (Y chromosome)


In my last post I mentioned the issue of shorter branches for contemporary Africans in the Y-chromosome phylogenetic tree. This means that starting from the fork that leads on one side to Africans, and the other to non-Africans, the latter contains more mutations than the former, but we are all the same age and equally distant from our common ancestor. So why do the Africans have fewer mutations? Do Eurasians accumulate more mutations? Are the branches built incorrectly? This post will try to shed some light on this matter.


Y-chromosomes and haplogroups


The accepted haplogroup structure for chromosome Y, just like that of mtDNA, is rooted in Africa, where the most basal lineages are found.


Using the phylogenetic tree analogy, all other variants, found outside of Africa are branches that stem from this African origin. Outlier branches, even closer to the root, include our ancestor-relatives, the Denisovans and Neanderthals.


Back in 2014 I posted about Neanderthal Y chromosomes, and used the following image, which I have updated to add Denisovans.


hominin Y chromosome haplo tree

The Denisovan and Neanderthal Y-chromosomes were studied by Martin Petr et al. (2020) in their paper The evolutionary history of Neanderthal and Denisovan Y chromosomes (Science 369, 1653-1656 (2020). doi:10.1126/science.abb6460 🔒- free access on Biorxiv🔓), which I will comment in depth in a future post. The authors of this paper mention that their Y-chromosome phylogenetic trees display shorter branch lengths for Africans.


This is interesting! They state "Importantly, we discovered that the branch-lengths in Africans are as much as 13% shorter compared to non-Africans (Figure S7.3), which is consistent with significant branch length variability discovered in previous studies and suggested to be a result of various demographic and selection processes."


Below is Figure S7.3 mentioned above. You can see that all these African samples have ratios, except for the S_Mbuti_1 sample, that are lessr than 1, meaning the branches are shorter than the European ones. Furthermore, the most diverged samples (A00) are even shorter :


branch length african vs non-african y chromosome phylo trees
Original caption:Branch length differences between African Y chromosomes and a panel of 13 non-African Y chromosomes. Ratios were calculated by creating an alignment of chimpanzee, African and non-African Y chromosomes and taking the ratio of the number of derived alleles observed in an African (x-axis) and the number of derived alleles in each of the individual non-Africans (dots, Table S7.1). “A00” represents a merge of sequences of two lower coverage Y chromosomes, A00-1 and A00-2 (Table S4.3). Fig S7.3 in Petr et al. (2020)

The branch lengths refer to the number of accumulated mutations in the branches of phylogenetic-trees. Africans have fewer mutations than non-Africans, so their branches are shorter, yet they are supposedly older! This is an anomaly, because it impliles a slower mutation rate in Africa, or a quicker one outside of Africa. The explanation offered by the authors is a classic one. This explanation is that leaving Africa caused population bottlenecks and forced adaptation to new environments which speed up mutations, or so the theory goes! Below is Fig. S1.7 from this paper.


y chromosome phylo tree
Branches. Fig. S1.7

The values of the branches a, d, e, and f are given in the paper's Table S7.1 and are the following (I adapted the image and included a new column, a+d the branch leading to non-Africans, which, as you can see, has more mutations than the African ones -compare the values of a+d with f.


branch lengths of Y chromosome phylo tree
Branch lengths. Table S7.1

The difference seems small but it is significant. Furthermore since Ust'Ishim, who died 45,000 years ago, non-Africans added an average of d-e mutations, ~200 of them. Africans added ~180-190 mutations. Hence, the "shorter branch" issue.


Shorter or Longer?


However, an earlier paper that studied Neanderthal and H. sapiens Y chromosomes by Mendez F, Poznik G, Castellano S, Bustamante C, (2016) (The Divergence of Neandertal and Modern Human Y Chromosomes. The American Journal of Human Genetics, 98, 728-734) showed different branch lengths, but with an opposite skew! This work included two figures (Fig. 1B, and Fig. 2) which I have combined and adapted in the image below. (the filters are different regions used to compare the DNA strands, some are more restrictive than others).


Neanderthal and human Y chromosome phylo tree

The branch lengths leading to the most divergent Africans with haplogroup A00, Mbo people from Cameroon, has a length e, which is longer than the one leading to the Reference (European men), branch d. But both share the same root. Why have the Mbo men accumulated more mutations than Europeans during the same time span?


This paper calculates the split age for both Modern Human branches (Mbo and Europeans) at 280 thousand years ago (kya), and dates the Neanderthals split at ∼588 kya. The Neanderthal man that was analyzed, died ∼49,000 years ago, in El Sidrón, Spain, and is located on branch f. His lineage contains 49,000 years of fewer mutations because we mutated while he remained static, yet, the total line f contains far more mutations than either modern human line: the A00 (a+e) or European lineage (a+d), who, by the way have had an added 50 ky of mutations on them!


This shows that the Neanderthal Y chromosome mutated faster than Homo sapiens Y chromosome, or that the timeline calculated in the paper is inaccurate.


Back and Recurring mutations


The paper noted that "The 17 sites that are incompatible with the tree are principally due to recurrent and back mutations". So these are not as infrequent as imagined.


Reference Bias


Janet Kelso, co-author of Petr et al.'s paper investigated branch lengths and published her research in 2024: Resolving the source of branch length variation in the Y chromosome phylogeny, Yaniv Swiel, Janet Kelso, Stéphane Peyrégne. bioRxiv 2024.07.05.602100; doi: https://doi.org/10.1101/ 2024.07.05.602100.


This paper admits that population size, and reproductive age, accumulated deleterious mutations due to bottlenecks in the out of Africa group, may play a role, but the main cause of branch length differences is the reference human Y chromosome used for comparison, that lacks mutations that appear in more diverged haplogroups: "branch length variation amongst human Y chromosomes cannot solely be explained by differences in demographic or biological processes. Instead, reference bias results in mutations being missed on Y chromosomes that are highly diverged from the reference used for alignment."


Reference bias is an error caused by using a certain benchmark (in this case the reference haplogroup, which is European, known as the Homo sapiens (human) genome assembly GRCh37 (hg19) from the Genome Reference Consortium), that favors genetic "reads" that match it, over those in alternative alleles. The reference Y haplogroup is R1b.



Comment on A00, the most ancient Y chromosome


For those interested in the deepest root of Y-chromosomes, the one named A00, you can find the original paper describing it by Mendez F., et al., (2013) (An African American Paternal Lineage Adds an Extremely Ancient Root to the Human Y Chromosome Phylogenetic Tree. AJHG, Vol 92:3 3, 7 March 2013, pp 454-459, https://doi.org/10.1016/j.ajhg.2013.02.002). An interesting critique to the findings, especially the extreme old age of this "basal" root, can be found in this paper: Elhaik E, Tatarinova TV, Klyosov AA, Graur D., (2013). The 'extremely ancient' chromosome that isn't: a forensic bioinformatic investigation of Albert Perry's X-degenerate portion of the Y chromosome. (Eur J Hum Genet. 2014 Sep;22(9):1111-6. doi: 10.1038/ejhg.2013.303. Epub 2014 Jan 22. PMID: 24448544; PMCID: PMC4135414).


San, the oldest humans?


Sometimes the media, and websites mention "the oldest" or "the earliest" people pointing at the Mbo or the Khoisan (San) groups, but in fact nobody alive nowadays is "older" than other populations. We have all been evolving since the first Homo sapiens appeared. We are all equally distant from him or her, nobody is closer or more similar to those original modern humans.


This is why I dislike phylogenetic trees like the one shown below (source) that implies a direct link from the ancient root to nowadays for the San people, and a series of steps to a short fork for Asians and Europeans. (Hss: H. sapiens, Hsnn: Neanderthals, Hsnd: Denisovan)


human phylo tree

When I read that the Khoisan separated from all other humans 150,000 years ago, I get the impression that it is a false statement. The Khoisan were not isolated since then, they also have admixture of other humans, but having lived in isolation in the deep past, and admixing with other diverse, divergent, isolated groups, they acquired a higher diversity themselves, as a population, while humans living outside of Africa lost diversity due to bottlenecks and founder effects. But the genes we retained in America, Asia, Oceania and Europe are mostly as old as the ones found in Africans.



Back to differing branch lengths


y chromosome different haplogroup branch lengths

Hallast P, Batini C, Zadik D, et al. (2015). (The Y-chromosome tree bursts into leaf: 13,000 high-confidence SNPs covering the majority of known clades. Molecular Biology and Evolution. 2015 Mar;32(3):661-673. DOI: 10.1093/molbev/msu327. PMID: 25468874; PMCID: PMC4327154. 🔓) mentioned that "Different clades within the tree show subtle but significant differences in branch lengths to the root." Fig. 3 in this paper (above is part of the figure) gives a clear image on how the branch lengths differ.


The tips of all haplogroups should all align, justified on the right side, as all the tips are contemporary, however, they have different lengths. I took R2 as the reference and drew a black vertical line. This makes the shorter branches stand out: haplogroups A, B, H, I1, Q, and R, and also the longer ones like C, G, J, or T. As you can see in the image above (I recommend visiting Fig 3 following the link, because it has far more detail than the simplified version I included above.)


Replication timing


A very thorough analysis on the causes of branch length differences can be found in Qiliang Ding , Ya Hu , Amnon Koren , Andrew G Clark, (2021). Mutation Rate Variability across Human Y-Chromosome Haplogroups. Molecular Biology and Evolution, Vol 38:3, March 2021, pp 1000–1005, https://doi.org/10.1093/molbev/msaa268.🔓.


The paper used data from over 1,700 men and "uncovered substantial variation (up to 83.3%) [in the] mutation rate among haplogroups. This rate positively correlates with phylogenetic branch length, indicating that interhaplogroup mutation rate variation is a likely cause of branch length heterogeneity."


The authors remarked that "Previous studies suggested that branch length heterogeneity might be caused by nongenetic factors, for example, paternal age variation across populations, acting over many generations. Another possibility is variation in mutation rate among Y-chromosome haplogroups.... [but] It was suggested that variation in Y-chromosome mutation rate across haplogroups was unlikely (Jobling and Tyler-Smith 2017)."


They disagree with the nongenetic factors and with Jobling and Tyler-Smith's dismissal of varying mutation rates, and prove that both are mistaken. This paper confirms that something known as replication timing varies across haplogroups, and this difference is linked to higher mutation rates (later replication causing more mutations than early replication timing).


Replication timing is the sequence in which the DNA of a chromosome is duplicated during cellular division. It involves unwinding and unzipping the DNA strand in a specific orer, in different places, some of them simultaneously.


Due to these differing mutation rates, branch lengths are different, and this impacts on the timing and dating of haplogroups. The paper's supplementary file states that the divergence time of haplogroups E1b, R1a, and R1b may be underestimated, while that of haplogroup B is overestimated, as the former have shorter branches, and the latter, longer ones. See Fig. 3 C and D in the paper.


The explanation sounds good, but why do different haplogroups have different replication timing? Alas, no answer is provided!


Population factors


Nevertheless, Barbieri, C., Hübner, A., Macholdt, E. et al. (2016) (Refining the Y chromosome phylogeny with southern African sequences. Hum Genet 135, 541–553 (2016). https://doi.org/10.1007/s00439-016-1651-0 🔓) attribute branch length in Southern African haplogroups to paternal age: "there is pronounced variation in branch length between major haplogroups; in particular, haplogroups associated with Bantu speakers have significantly longer branches. Technical artifacts cannot explain this branch length variation, which instead likely reflects aspects of the demographic history of Bantu speakers, such as recent population expansion and an older average paternal age. The influence of demographic factors on branch length variation has broader implications both for the human Y phylogeny and for similar analyses of other species." (Sure! it affects the calculation of dates along the branches of phylogenetic trees!).


This paper finds "The shortest branches in the Y chromosome phylogeny are for haplogroups A and B... E1b1a lineages have significantly longer branches than E1b1b or E2 lineages." Taking a look at the mutations marked along the phylogenetic tree shown in the paper's Fig 1, it confirms the comment branch lengths variability (below is the number of mutations from the tip to the root at the A2—T node).


  • A2a: 17
  • A2b: 7
  • A2c:22
  • A3b1b: 21
  • B2B1: 113
  • E1b1a: 208
  • E1b1b: 138
  • E2: 105

These people, living today have an extremely wide variation in mutation numbers between their common ancestor at the A2—T root and themselves: 7 to 208 mutations!! They are all Africans, and should be equally distant to the R1b reference genome, meaning that Kelso's reference bias does not apply in this case. This could be due to paternity age (older men have more mutations in their sperm as they sire children and pass on mutations in their Y chromosomes to their sons), or to the different replication times of different haplogroups.


T Naidoo et al., (2020) in their analysis of Khoe-San men in South Africa also found the branch issue: " Branch Length Heterogeneity Several earlier studies (Scozzari et al. 2014; Hallast et al. 2015; Barbieri et al. 2016) found evidence of branch length heterogeneity among Y-chromosome haplogroups, and provided possible reasons for its occurrence. We also noted significant differences in branch length heterogeneity among the major African haplogroups (supplementary tables S2 and S3, Supplementary Material online). A reduced mean branch length for haplogroup A, noted previously by Scozzari et al. (2014), was again apparent from our data. Although most major haplogroups differed significantly (with the exception of the E1b1a subclades), we found that haplogroup B did not appear to have as reduced a mean branch length, relative to haplogroup E, as found previously (Hallast et al. 2015; Barbieri et al. 2016). Within haplogroup E, E1b1b1 was found to have the highest mean branch length; though this may have been due to a lower sample size compared with haplogroup E1b1a." It seems to me, as a layman, that the branch length issue perplexes even the smartest scholars.


Closing comments


This post shows that scholars don't agree on why the African branches, the most diverged, and "archaic", leading to the root, and origin of our H. sapiens species, contain fewer mutations than those found in Eurasian people. Since the basis of calculating the splits between modern humans and archaic relatives like Neanderthals and Denisovans is the assumption that there is a "mutation clock" that ticks at a regular pace, so if we know the ticking rate, and the number of mutations, we can calculate when species split from others, and people diverged from others. Short branches on supposedly ancient lineages are incongruent.


We are all equally ancient, Africans, Eurasians, and Americans, yet we have accumulated mutations in our Y chromosome at different rates. This is something that should be clearly analyzed. Software issues, methodology, sampling, reference bias, replication times, older reproductive ages, larger population sizes, bottlenecks, etc. have been put forward to explain this anomaly. None of these answers seems satisfactory. Chromosome Y is peculiar, it is small, and critical; any mutations here can have disruptive effects. We are overlooking something. When we find it, we will know why some branches are longer than others.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Saturday, March 7, 2026

Early Peopling of America ~40 kya


This will be a very short post. I came across a paper published in Nature in 2020, by V. M. Cabrera (Counterbalancing the time-dependent effect on the human mitochondrial DNA molecular clock. BMC Evol. Biol. 20, 1–9 , 2020) which, as you can see by its title is about the uneven rate at which mutations accumulate in mitochondrial genomes, it can run faster or slower and the research cited by Cabrera attributes it to "changes in the effective population size of the human populations." Another factor mentioned by Cabrera are " "transient polymorphisms... play a slowdown role in the evolutionary rate deduced from haplogroup intraspecific trees".


Transient polymorphisms


These are short-lived variants that through selection take over and displace another. Later, they are themselves replaced, hence their name (transient). An example is the classical moths of England, they came in two varieties, light colored and dark. Before the industrial revolution, the light ones prevailed, as the dark ones were very visible (easy preys) on the pale tree trunks. Then, when the industrial soot of coal furnaces darkened the bark, the situation changed, and the dark ones were less visible. Soon they displaced the ligher colored moths (dark allele replaced pale one). However, in the 1970s, with pollution abating measures in place, the bark of trees became pale once again, and the light colored moths now prevail over the dark ones.


Regarding variability as time passes, Cabrera mentions an interesting fact: "For example, an ancient mtDNA study has corroborated empirically the persistence of an ancestral M lineage unaltered along a period of more than 8000 years" and adds that "Consequently, a lineage could remain immutable for several generations while identical lineages in the same population suffer one or several mutational changes in the same time interval."


This situation is shown in the paper's Figure 2, shown below, which very clearly shows how mutations can be overlooked when building a tree, and its effect on the mutation rate that is inferred from it. Cabrera describes the figure as follows: "We represent such a scenario in Fig. 2A as a maternal genealogy. Notice that the ancestral lineage can give rise to offspring in different generations throughout its existence in the population. In this way, when the population is sampled after n generations, we can find, in addition to the ancestral lineage (e), lineages derived from it that have accumulated significant mutational differences in their branches (f, h). However, this fact is not reflected in the tree built from the same sample (Fig. 2B) because, irrespective of the generation in which they appeared, all derived lineages sprout at the same time from the ancestral node (e). This difference between genealogies and trees has notable consequences. On the one hand, it could explain the lack of mutation rate homogeneity between lineages found in intraspecific haplogroup.""


phylo and genealogical trees
Comparison between a genealogy (A) and a tree (B) constructed from the same sample (a to h). White circles are individuals with the ancestral lineage, black circles are individuals with additional mutations. Inter-circle segments represent generations and crosses on the segments represent mutations.. Figure 2 in Cabrera.

Cabrera then calculates the ages of the different mtDNA haplogroups, and finds a relatively old age for the entry of humans into America:


"Finally, although we proposed a unique migration for the colonization of the Americas around 40,000 years ago (20) which is directly or indirectly supported by archaeological dates, it seems possible that this first migration, signaled by the ages of haplogroups A2 and B2 was followed afterwards by a second wave, also before the Last Glacial Maximum, marked by haplogroups C1, D1, D4h3a and X2a around 27,500 years ago (Table 3).


Note 1. The figures from Table 3 are these: C1: 30 ky, D1: 32 ky, D4h3a: 26 ky, and X2a: 22 ky.


Note 2. Cabrera in (20) is citing himself! (Cabrera, V. M., 2020, Counterbalancing the time-dependent effect on the human mitochondrial DNA molecular clock. BMC Evol. Biol. 20, 1–9.)


This second paper by Cabreera, cited above as (20), states that "Finally, the time of human expansion to the American Continent deduced from the haplogroup B2 phylogeny was approximately 37,000 ya (Table 1, and Table S11). This age supports a pre-Clovis occupation of the New World, well before the last glacial maximum."


I am happy to see someone who is pushing the boundaries towards an older date, 40-37 ky ago. This is positive. It will open the door for others to use this information to validate their findings of an even earlier date for the arrival of modern humans to America.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Thursday, March 5, 2026

Molecular Clock (revisited Again)


Following my recent post on reversions, also known as back mutations, which shows that mutation rates are higher than expected due to back-mutations, I decided to update the information about the errors in "molecular clocks."


Behar DM, et al. A "Copernican" reassessment of the human mitochondrial DNA tree from its root, published in 2012 (Am J Hum Genet. 2012 Apr 6;90(4):675-84. doi: 10.1016/j.ajhg.2012.03.002. Erratum in: Am J Hum Genet. 2012 May 4;90(5):936. PMID: 22482806; PMCID: PMC3322232), included a section titled "Indications for Violation of the Molecular Clock", which compared the number of mutations from the base (called RSRS or Reconstructed Sapiens Reference Sequence) along haplogroup lineages. They noticed a wide variation, from 42 to 71 substitutions with a mean of just over 57. This is a wide spread of data, clearly the clock by which mutations ocurr is clicking at different rates for different haplogroups.


In line with my recent post (mutation rates are faster in Africa), Behar et al. noticed that there were 77 mutations in the L2b1a mtDNA variant (supposedly old, and ancient, and only found within Africa). But... M2b1b and M7b3a, found outside of Africa, and "daughters" of the African L clades, has 71 mutations. Below I quote this paper:


"The accepted notion of a molecular clock means that contemporary mtDNA haplotypes should show statistically insignificant differences in the number of accumulated mutations from the RSRS. Triggered by the suggested change in the reference sequence that facilitates substitution counts from the ancestral root, we further evaluated this hypothesis. The range of substitution counts separating contemporary mitogenomes belonging to major haplogroups from the RSRS is shown in Figure S2.
The mean distance is 57.1 substitutions, the median is 56 and the empirical standard deviation is 5.9. Widely different distances ranging from 41 substitutions in some L0d1a1 mitogenomes to 77 in some L2b1a mitogenomes are observed. Interestingly, the ranges of substitution counts within haplogroups M and N, which are hallmarks of the relatively recent out-of-Africa exodus of humans, are also very large. For example, within M there are two mitogenomes with 43 substitutions (in M30a and M44) and two mitogenomes with as many as 71 substitutions (in M2b1b and M7b3a). This is especially striking because the path from the RSRS to the root of M already contains 39 substitutions. Hence, the difference between the M root and its M44 descendant is only four substitutions (two in the coding region and two in the control region) as compared to 32 substitutions in the M2b1b and M7b3a mitogenomes.
These observations raise the possibility that the tree in general, and haplogroup M in particular, might not adhere uniformly to the assumed molecular clock, under which substitutions occur at a fixed rate on all branches of the tree over time. We evaluated this scenario by performing generalized likelihood ratio tests of the molecular clock by using PAML33 on subsets of samples from the entire tree, on haplogroup L2 (following past evidence of clock violations in this haplogroup40) and on the sister haplogroups M and N. Our results demonstrate violations of the molecular clock in M (0.00015 ≤ p value ≤ 0.0003 for χ2 GLR test in three different analyses) and give mixed results for the entire tree (p = 0.005 and p = 0.018 for two analyses, which might be sensitive to the parts of the tree randomly sampled) and L2 (GLR χ2 p value = 5 × 10−5 and p value = 0.033 for two analyses) and borderline results in N (GLR χ2 p value = 0.049 and p value = 0.054 in two analyses). We are currently unable to offer well-founded explanations for these findings, which remain the scope of future studies.
"


The researchers couldn't explain why!


Generation Times or Generation interval


One explanation for variability in mutation rates is a gradual decrease in the rate of germ-line mitoses per year in the human lineage caused by longer generation times. This is known as the "hominoid slowdown hypthesis", first proposed by Morris Goodman in the 1950s.


The "generation length" or interval, (time between birth and reproduction) vary from one species to another. The concept of the slowdown is that creatures with a short generation time, go through many more generations per unit time than animals with a long generation time (like humans). Humans have longer generations than chimps and other old and new world monkeys. The same can be noted for humans and mice.


The explanation is simple: take an animal, like a dog, which can produce a litter per year after the age of 1. If allowed to mate on a yearly basis, it will have produced 12 generations in 12 years (its lifespan). So if each for each generation there are n random mutations taking at take place in the germline (ova and sperm) dogs will have an accumulated mutation number of 10n mutations after ten years. Considering human generation times of 29 years, a human being will only have produced one generation, with only n mutations in 29 years (assuming both species undergo the same random number of mutations per generation), while dogs will have undergone 29n mutations.


Can we be certain that Neanderthals, Denisovans, Homo erectus, or even H. sapiens who lived 200,000 years ago, had generation times shorter, equal, or longer than ours? Were they 25, 29, or 30 years long?


One 2006 paper states that "Using 15 years as the generation time for chimpanzees and ancient humans and 20 years for that of modern humans, the estimated time of the evolution of long generation time in the modern humans is approximately one million years."


Another paper using introgressed Neanderthal segments in Eurasians (the paper is very interesting!) calculated that East and West Eurasians had different generation times: "differences in the generation interval across Eurasia, by up 10–20%, over the past 40,000 years... we estimate that this difference corresponds to a 2.68 or 3.39 years shorter generation interval in West Eurasia if East Asian mean generation time was 28 or 32 years respectively."


Richard Wang et al., (2023) explored the subject in depth "Our analyses of whole-genome data reveal an average generation time of 26.9 years across the past 250,000 years, with fathers consistently older (30.7 years) than mothers (23.2 years). Shifts in sex-averaged generation times have been driven primarily by changes to the age of paternity, although we report a substantial increase in female generation times in the recent past. We also find a large difference in generation times among populations, reaching back to a time when all humans occupied Africa." The image below shows how generation time changes over time, and region. The image caption reads "Fig. 3. Change in generation interval across different human populations. Generation intervals were estimated in ancestors of four major continental human populations included in the 1000 Genomes Project; sex-averaged generation intervals are shown here as smoothed by loess (see fig. S6 for full results). Confidence intervals for each population were obtained by bootstrapping, as in Fig. 2. The inset shows results from including polymorphisms that date back to 78,000 generations ago; note that age estimates of mutations in the very distant past have decreased accuracy (15). AFR, Africa; EAS, East Asia; EUR, Europe; SAS, South Asia.


chart, generation time function of antiquity
Genertion time evolution over time, by region. Fig. 3 in Wang et al. (2023)

Note: In case you wonder why the graph shows South Asians and Eurasians 10000 generations ago (roughly 300,000 years ago) a time when modern humans were just originating inside of Africa, the paper points out that "While the continental labels for each population are used across the span of the analysis, note that beyond roughly 2000 generations ago, all non-African populations were likely located in Africa and show little differentiation among themselves; coalescence among all ancestral populations living in Africa does not occur until more than 10,000 generations ago."


The paper quantifies the time and its impact on mutation rates: "The dominating pattern across the past 10,000 generations is a significantly shorter sex-averaged generation interval for East Asian, European, and South Asian populations— 20.1 ± 3.9, 20.6 ± 3.8, and 21.0 ± 3.7 years—compared to the African population, 26.9 ± 3.5 years. The estimated generation times do not converge between populations until we expand our analysis to include periods older than 10,000 generations ago (Fig. 3, inset)... The large difference in generation times between populations suggests that different time scales are needed to estimate events outside of Africa (20 to 21 years per generation) versus those in Africa (27 years per generation). These results are consistent with the prediction of a shorter generation time in non-Africans, based on the observation of a slightly elevated per-year mutation rate in these populations."


Environmental Factors and Mutations


In a recent post I discussed a hypothesis that suggested a climatic influence on mtDNA mutations. Today I will mention the effects of the Ultraviolet Radiation (UV) in sunlight on mutation rates.


Research by Kelley Harris (2015) looked at European mutation rates, and found "Europeans experience higher rates of a specific mutation type that has known associations with UV light exposure." It isn't recent either, it is very old: "rate acceleration seems to have occurred between 25,000 and 60,000 y ago, not long after Europeans diverged from Asians."


This is not trivial, Harris goes on to explain its significance: "Even if the overall European mutation rate increase was small, it adds to a growing body of evidence that molecular clock assumptions break down on a faster timescale than generally assumed during population genetic analysis. It was once assumed that the human lineage’s mutation rate had changed little since we shared a common ancestor with chimpanzees, but this assumption is losing credibility owing to the conflict between direct mutation rate estimates and molecular-clock-based estimates."


Harris then argues that "the results of this paper indicate that another force may have come into play: change in the mutation rate per mitosis." In the case of germline mutation rates, those passed on to the next generation because they take place in the germ cells, there are several drivers that can cause them:


Paternal Age: the older the father, more chances that mutations have accumulated in the spermatogonial stem cells due to repeated mitosis. Replication errors: mistakes in copying the DNA sequence (chunks are deleted, repeated, substituted). Methylation: DNA methylation (addition of a methyl —CH3) influences mutations, etc.


Harris attributes it to the light skin of Europeans and UV radiation. Yet wonders about the mechanism: "
The question remains how UV could affect germ-line cells that are generally shielded from solar radiation
" (there is an answer provided, folate deficienty due to UV depletion that causes mutations).


Closing Comments


We have seen that mutation rates are affected by climate, sunlight, generation times, and also natural selection. This means that we should be cautious when interpreting data using molecular clocks.


Soojin Yi, Darrell L. Ellsworth, and Wen-Hsiung Li (2002) express this very clearly "Therefore, application of a molecular clock to estimate divergence dates should be exercised with great caution even in relatively closely related taxa. That is, a molecular clock calibrated for some lineages may not be applicable to other lineages because the assumption of rate constancy among lineages may not hold, as shown in the case of higher primates. Furthermore, the rate estimated from one genomic region may not be applicable to another region because the mutation rate varies among genomic regions."



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Saturday, February 21, 2026

Back Mutations are common & more frequent than previously stated


Reversions, also known as back mutations are considered rare in biology. Basically, what it means is an initial mutation in the geneome, reverses back to its original (wild-type) form through a second mutation that restores the base in the DNA sequence.


For example, a "chunk" of DNA could contain the following bases: ACGCTG and, a random, chance mutation replaces the cytosine (C) for an adenine (A) ACGATG and a second mutation restores (reverses) the original situation ACGCTG.


This has important consequences, first of all, if we look at the ancient sample and the most recent one, there is no way we will ever know that it mutated and reverse by back-mutating, (both have the same sequence: ACGATG so how can we know if there was a back and forth flip in between both samples?)... unless we find an sample of someone in the same line, or a parallel lineage with the first (derived) mutation, but not the reversion,which seems a very unlikely situation.


Assuming that all mutations are forward oriented and never reverse, may overlook mutations that were reversed. Since coalescence time and dates of lineage splits are based on mutations, if we overlook the reversions, we will miss out on the actual mutations (to and fro), counting zero when in fact there were two mutations.


So we will assume that mutation rates are lower than they really are by missing out these back-and-forth mutations.


If we overlook reversions we will assume there were only n mutations per a given amount of years, while there were actually m mutations: "n" that we see (for instance, there is a G instead of an A at a certain locus), and "p" mutations that flipped forth and another "p" that flipped back. n is therefore smaller than m; m = n+2p. So the mutation rate is higher than assumed.


However, the Neutral Theory of genetic evolution does not consider this alternative, it requires No back-mutations. Changes can only happen in one direction A → G. Which will never again flip back G → A, and No Recurrence there can't be multiple mutations at identical loci in different lineages.

Are they Common?


William Amos (2020) suggests that "back-mutations are far commoner than has been previously assumed" he adds that "Back-mutations are ‘silent' because they create the original ancestral allele, but can reasonably be assumed to occur about twice as often as triallelic SNPs are generated (two transitions are approximately twice as likely as one transition and one transversion). Triallelic SNPs are coded ‘MULTI-ALLELIC' rather than ‘SNP' in the 1000 genome data and are often ignored, but I counted 257,827 occurrences across all autosomes, implying over half a million sites carrying back-mutations. Moreover, this is probably an underestimate because the 1000 genomes data are low coverage and rely on extensive imputation which will often cause rare third alleles to go undetected. Equally, conservative curation will tend to remove third alleles that lack strong support. Note, triallelic sites are unlikely to be generated mainly by sequencing errors because only 1% of these sites carry a singleton as the rarest allele. This analysis is not intended to provide an accurate estimate of the back-mutation rate, but instead simply to demonstrate that large numbers of back-mutations do exist to the extent that models of evolution that rely on back-mutations occurring in appreciable numbers should not be dismissed a priori."


Research by Anke Fähnrich et al., (2023) on the North and Eastern African mtDNA shows that some mutations that serve as markers appear time and time again, the authors consider some as "Shared back mutations", others are simply repeat mutations. The paper shows them in a tree for "L0a1 and (b) L2a1" and clarifies that "We highlight with magenta, gray and turquoise boxes those variants that indicate that a different phylogenetic tree may better explain the samples from North and East Africa." These trees can be seen in the image below. The paper adds that "Shared back mutations (magenta) denote that parental haplotypes may be missing in PhyloTree. Mutations repeatedly observed in a subtree (gray) suggest that child haplogroups are missing. Variants that occur in multiple samples and differ from variants defining a parental haplogroup (turquoise) suggest that different variant combinations and haplogroup specifications may better explain North and East African mtDNA sequences". It also marks them with an "@" as "assumed back mutation or missing mutation". This goes to show that markers are shared across different haplogroups.


phylo tree mtDNA
Figure 10, haplo tree mtDNA L. Source

A similar situation was reported by Neil Howell, Joanna L Elson, D M Turnbull, and Corinna Herrnstadt (2004) who were investigating the oldest mtDNA haplogroups L0 and L2. Besides finding oddities in the trees, ages, etc., they noted multiple reversions: "The L0a outgroup sequence carries C alleles at nucleotides 16189 and 16192, whereas the L2a ancestral sequence is predicted to carry C and T, respectively, at these sites. The 16189 site subsequently undergoes mutation on four occasions (three forward and one reverse relative to the outgroup sequence), whereas the 16192 site undergoes reversion on five occasions. Thus, both sites appear to have relatively high rates of mutation, a result that has been observed in previous studies (Excoffier and Yang 1999; Meyer, Weiss, and von Haeseler 1999; Howelland Bogolin Smejkal 2000) and in the L2a networks of Salas et al. (2002)... The ancestral L2a sequence carries a C:T transitional nucleotide 16519, which undergoes reversion on three occasions. These results are not surprising and this site has long been recognized to have a high mutation rate."


This seems to define a "mutation hotspot", that I mentioned in a previous post (Laguna de los Pampas 10,000 BP remains in Argentina).



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Friday, February 20, 2026

mtDNA variants and Natural Selection


The chance mutations that are fixed in the DNA of our mitochondria and accumulate there have been used to trace the spread of human beings across the globe. Passed on in a matrilineal form, we all receive the mtDNA from our mother's ovum. Our father's sperm does not carry any mitochondria. Randmo mutations gradually accumulate so they serve as markers in the mtDNA and specific markers define haplogroups.


A paper suggests that these random mutations are then shaped by the forces of Natural Selection. (D. Mishmar, E. Ruiz-Pesini, P. Golik, V. Macaulay, A.G. Clark, S. Hosseini,M. Brandon, K. Easley, E. Chen, M.D. Brown, R.I. Sukernik, A. Olckers, & D.C. Wallace, /2003) Natural selection shaped regional mtDNA variation in humans, Proc. Natl. Acad. Sci. U.S.A. 100 (1) 171-176, https://doi.org/10.1073/pnas.0136972100).


They note that although mutations arise in a random way in the mtDNA, as they have an effect on the mitochondria which produce the body's cells energy and regulate cellular metabolism by producing the energy-rich molecule adenosine triphosphate (ATP), they may be a target of natural selection. They state that "Natural selection shaped regional mtDNA variation in humans."


mitochondria
Mitochondria the body's powerhouse. Copyright © 2026 by Austin Whittall

Molecular clock affected


The fact that mutations are not neutral, and are acted upon by natural selection, implies that the assumptions on which the mtDNA molecular clock are based, are flawed. The paper warns: "If selection has played an important role in the radiation of human mtDNA lineages, then the rate of mtDNA molecular clock may not have been constant throughout human history. If this is the case, then conjectures about the timing of human migrations may need to be reassessed."


The molecular clock based on mtDNA is based on an axiom: genes accumulate new mutations in a clock-like manner, so knowing the rate at which mutations take place (i.e. 3 mutations per 10,000 years), and measuring the average amount of mutations that have appeared since a particular node on a phylogenetic tree (9 mutations), allows us to date the node: 30,000 years. And from there date other nodes based on the number of mutations and the mutation rate.


This is reasonable as long as the mutation rate is constant. But if it varies, then it will provide incorrect dates.


Positive selection could affect the mutation pattern similar and cause an acceleration in the mutation speed. (Further reading on the mtDNA clock: Eva-Liis Loogväi, Toomas Kivisild, Tõnu Margus, Richard Villems (2009))


mtDNA and Selection


After a long stasis in Africa where the L haplogroup is found, humans moved into Eurasia and two branches, or clades, M and N formed outside of Africa and comprise all the mtDNA diversity in the rest of the world. M and N are derived from the African haplogroup L3. And the split is supposed to have taken place around 55-70 kya, during the Out of Africa Event.


Interestingly, M is basically absent in the Middle East, yet it is found in Ethiopia, Southern Arabia and in India and East Asia, suggesting to some a Southern route of migration out of the Horn of Africa across Bab el Mandeb and Hormuz straits. However, a paper published in 2018 by Vicente M Cabrera, Patricia Marrero, Khaled K Abu-Amero, and Jose M Larruga, suggests that both M and N originated in Southeast Asia and migrated westwards. In the case of N haplogroup, it was believed to have formed in the area that links the Levant and Africa and that it appeared in humans taking a northern route out of Africa into Eurasia. But this paper suggests that N originated in Southeast Asia, and moved west across Asia towards Africa. The authors argue that "If one accepts that basal L3 lineages (M, N) evolved independently in southeastern Asia and not in Africa or near the borders of the African continent where the remaining L3 lineages expanded, one is confronted with the question of where the basal trunk of L3 evolved. A gravitating midpoint between eastern Africa and southeastern Asia would situate the origin of L3 in inner Asia."


The paper then states:


"L3 exited from Africa as a pre-L3 lineage that evolved as basal L3 in inner Asia. From there, it expanded, returning to Africa as well as expanding to southeastern Asia, giving rise to the African L3 branches in eastern Africa and the M and N L3 Eurasian branches in southeastern Asia, respectively. This model, which implies an earlier exit of modern humans out of Africa, has been tested against independent results from other disciplines...."


The paper includes the following maps as its Figure 1, and the caption reads: "Geographic origin and dispersion of mtDNA L haplogroups: a Sequential expansion of L haplogroups inside Africa and exit of the L3 precursor to Eurasia. b Return to Africa and expansion to Asia of basal L3 lineages with subsequent differentiation in both continents. The geographic ranges of Neanderthals, Denisovans and Erectus are estimates only."



The paper adds that the "early return and subsequent expansion inside Africa of carriers of L3... haplogroup might help explain, the Neanderthal introgression detected in the western African Yoruba and in northern African Tunisian Berbers." (see my recent post on Neanderthals in Africa).


The authors assume anatomically modern humans left Africa in an early migration 125 kya , met with Neanderthals in south-central Asia, admixed and as the climate worsened ~75kya, the humans moved west and returned to Africa (with the L3 variant with them and it diversified there), and they also moved east reaching SE Asia and China.


Selection and Diversification


Getting back to Mishmar et al., they argue that in Eurasia the M and N lineages spread across the continent in different lineages: A, C, D, and G. Which have a "striking regional variation, traditionally attributed to genetic drift. However, it is not easy to account for the fact that [these lineages] show a 5-fold enrichment from central Asia to Siberia". They argue that this enrichment is the result of natural selection acting as people left their traditional environment (warm, tropical, or temperate climates) and advanced into harsher and colder continental climates in Central and Northern Asia.


The researchers analyzed 104 complete mtDNA sequences from across the world and found that the African haplogroups more or less followed the neutral model, but American, European, Siberian and Asians didn't, they deviated from it. They found that the ATP6 gene, which is a "conserved" mtDNA protein had the highest variation in its amino acid sequences. "Conserved" means that it has remained mostly unchanged over the ages and among individuals and species because it has a low tolerance for mutations, because it is critical for cellular function. So, why would it present so many mutations?


To find out why, they compared the ratios of mutations for the ATP6 gene in different climate zones (arctic, tropical, and temperate) and found that it was highly variable in mtDNAs from the Arctic. Another mtDNA protein called cytochrome b which helps move electrons and create a proton gradient, essential for cellular energy production, was particularly variable in the temperate zones. Another protein, cytochrome oxidase I (or COX1), which also plays a vital role in electron transport, was more variable in the tropical areas. The authors concluded that "selection may have played a role in shaping human regional mtDNA variation and that one of the selective influences was climate."


They then downplay the effects of founder effects arguing as follows:


"..there are striking differences in the nature of the mtDNAs found in different geographic regions. Previously, these marked differences in mtDNA haplogroup distribution were attributed to founder effects, specifically the colonizing of new geographic regions by only a few immigrants that contributed a limited number of mtDNAs.
However, this model is difficult to reconcile with the fact that northeastern Africa harbors all of the African-specific mtDNA lineages as well as the progenitors of the Eurasia radiation, yet only two mtDNA lineages (macrohaplogroups M and N) left northeastern Africa to colonize all of Eurasia and also that there is a striking discontinuity in the frequency of haplogroups A, C, D, and G between central Asia and Siberia, regions that are contiguous over thousands of kilometers.
Rather than Eurasia and Siberia being colonized by a limited number of founders, it seems more likely that environmental factors enriched for certain mtDNA lineages as humans moved to the more northern latitudes.
Natural selection has been hypothesized to explain anomalies in the branch lengths of certain European and African mtDNA lineages.
"


However, a paper by Taku Amu and Martin Brand (2007), disagrees with this concept, and states that there were no differences between the mitochondrial energy management in Arctic or Tropical populations, and that the mutations which were expected to lower coupling efficiency leading to more heat generation in colder climates wasn't detected, and in fact, "Contrary to the predictions of this hypothesis, mitochondria from Arctic haplogroups had similar or even greater coupling efficiency than mitochondria from tropical haplogroups."


More recent research by Jukka Kiiskilä et al (2021) also notes that mtDNA variants are under natural selection and that different mtDNA haplogroups exert a different effect on the physical performance in athletes! the paper looked at Finnish military conscripts and reported that "Following a standard-dose training period, excellence in endurance performance was less frequent among subjects with haplogroups J or K than among subjects with non-JK haplogroups."


Takayuki Nishimura and Shigeki Watanuki (2014) studied mtDNA haplogroup D vs. non-D groups regarding body warmth, and found that "[Non shivering thermogenesis] NST was greater in winter, and that the D group exhibited greater NST than the non-D group during winter...no significant differences in rectal and skin temperatures were found between groups in either season. Therefore, it was supposed that mitochondrial DNA haplogroups had a greater effect on variation in energy expenditure involving NST than they had on insulative responses... individuals from the D group exhibited greater winter values of ΔVO2 than individuals from the non-D group." So, mtDNA haplogroup D subjects had higher oxygen uptake (ΔVO2), meaning their body was "burning" more oxygen but not shivering or increasing the temperature. This suggests an efficient use of energy to heat the core only, and it has a clear mtDNA haplogroup component to it.


Interestingly, Haplogroup D seems to enhance energy burn (without shivering), and without increasing external temperature. From an engineering point of view this is great, since the ΔT or temperature differential between a body and its surroundings impacts directly on the energy loss (Q) the body experiences: Q = U · A · ΔT (where "A" is the area that transfers heat loss, and "U" is a heat transfer coefficient). So this is why the study didn't notice differences in skin or rectal temperatures.


Closing Comments


If random mtDNA mutations somehow provide an adaptative advantage (efficient energy use to keep warm in cold climates), and natural selection acts upon it, then the "neutral" theory is mistaken, and the molecular clock used to calculate dates is also wrong.


Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 
Hits since Sept. 2009:
Copyright © 2009-2025 by Austin Victor Whittall.
Todos los derechos reservados por Austin Whittall para esta edición en idioma español y / o inglés. No se permite la reproducción parcial o total, el almacenamiento, el alquiler, la transmisión o la transformación de este libro, en cualquier forma o por cualquier medio, sea electrónico o mecánico, mediante fotocopias, digitalización u otros métodos, sin el permiso previo y escrito del autor, excepto por un periodista, quien puede tomar cortos pasajes para ser usados en un comentario sobre esta obra para ser publicado en una revista o periódico. Su infracción está penada por las leyes 11.723 y 25.446.

All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted in any form or by any means - electronic, mechanical, photocopy, recording, or any other - except for brief quotations in printed reviews, without prior written permission from the author, except for the inclusion of brief quotations in a review.

Please read our Terms and Conditions and Privacy Policy before accessing this blog.

Terms & Conditions | Privacy Policy

Patagonian Monsters - https://patagoniamonsters.blogspot.com/