Translate

Guide to Patagonia's Monsters & Mysterious beings

I have written a book on this intriguing subject which has just been published.
In this blog I will post excerpts and other interesting texts on this fascinating subject.

Austin Whittall


Showing posts with label mutations. Show all posts
Showing posts with label mutations. Show all posts

Thursday, July 30, 2026

Genealogical research, some interesting thoughts


For several years I have been exploring my family tree, building on the work of my maternal grandfather and my father who pieced toegether information they had into two branches of ancestors.


It is far easier now than when my grandfather did his research in the 1980s, and my dad did his in the mid 1990s. In those days you had to go to the sites where the records were preserved (Register Office, cemeteries, churches — for baptism, burial, and marriage records). Now, online websites, offer free or paid services that let you browse document collections online. Government offices have searchable data (census, weddings, births, deaths, migration, etc.)


So, building on the work of my predecessors, I have added more information and found interesting data about my family.


It also helped me put in perspective things like generation times, the size of families, the terrible mortality rates of past ages, and the effects of these factors on the genetic mixup that I carry.


Each generation dupicates my direct ancestors (two parents, four grandparents, eight great grandparents, sixteen great great grandparents, and so on), by the time you have gone back seven generations to, say, James Whitehall, born in 1741, you have 128 great great great great great grandparents. Each of them contributed genes to their offspring, but even though you'd reckon that you received 1/128th of their genetic material (or 0.78%), it is very likely that you did not get any!


This is because genetic material is sliced into chunks, and recombined. Although on average a child gets 50% of its genes from each parent it may not be exactly 50 percent because these chunks aren't equal. So, alleles are passed on in a random way. So as you go further back in time, the chances are that a given ancestor's genes may have got lost on their way to you. Below I quote Bird, 2025 to clarify this fact:


"For example, while each of a person's four grandparents is expected to contribute to 1/4th of their autosomal genome, ten generations (~250–300 years) ago, a person can have at most 210 = 1024 ancestors. Thus each of these more distant ancestors is expected to contribute to < 1/1000th of their autosomal genome, with a ∼58% chance of no detectable relationship at all. At fourteen generations ago (∼420 years), this probability goes up to <95% (Donnelly 1983; Ralph and Coop 2013). For this reason, our genome is necessarily only a snapshot of a fraction of a person's complete set of ancestors and carries no information about the ancestors that did not contribute any DNA to this particular descendant."


I am certain that my mother's mtDNA is in my body's cells, and that it came from her mother, and so forth, along that matrilineal line. But, many mothers' mtDNA never got a chance to pass on because they only had sons, their daughters dying before they managed to reproduce.


Regarding the man, I am also certain that I carry the Y chromosome along a patrilineal line which may or may not be that of those who passed on the Whittall surname given the proclivity to stray that we humans have when it comes to making babies. A study by Guerrini, 2022 published in Cell, found that "3% of the total sample learned that the person who they thought was their biological parent is not." So, who provided your Y chromosomes may not be too certain.


My X chromosome came from my mom, but... which of her two X chromosomes did I inherit? The one she got from her Dad, and therefore, originating in his mother, or the one that came from my maternal grandmother, which in turn could have come from either of her parents... Complexity increases as we move deeper into the past.


Regarding the autosomal DNA, I carry a cocktail of alleles in my chromosomes, passed down from past generations, but, as mentioned further up, some ancestors are not represented, and their signature has been lost. This is like a bottleneck. Genes got lost despite the fact that there is a continuous line of descent between those ancestors and me.


Generation times vary. For instance James Whitehall, is 7 generations from me, and this spans 218 years between our birth dates: 31.14 years per generation. Taking another random ancestor on my maternal line is James Paterson, born in 1781, seven generations away, 178 years, 25.42 years per generation. I didn't average all 128 ancestors of those seven generations ago, but as you can see, the spread is large; roughly 28 ± 3 years.


A study published in Science by Wang et al., 2023 found "an average generation time of 26.9 years across the past 250,000 years" which seems more or less similar to the one of my family tree. However, the paper states that "The average human generation interval was at a recent minimum of 24.9 ± 3.5 years at ~250 generations ago (6.4 ka ago), roughly concurrent with the historic rise of early civilizations. Before this, it had declined from a peak of 29.8 ± 4.1 years at ~1400 generations ago (38 ka ago), just before the beginning of the Last Glacial Maximum." So, bearing this in mind, my family seems to have a longer generation time, roughly 10% longer than this paper's value for the recent past.


Finally, I was surprised at how people died at a young age. Women bore children from age 20 to their mid 40s, as many as 14 children, and most died in their youth. Women and men lost their spouses and remarried in their 30s and 40s, some did so even twice! The world was deadly in those days before vaccines, antiseptics and antibiotics. These deaths also erased many genes that my distant cousins and great great uncles and auntes carried.


So, just looking back at my family, I wonder how can we be so certain about the genetic diversity of some populations, generation times, admixtures, bottlenecks, when chance and probabilities rule how genes are passed down from our distant ancestors, even in a healthy population in Modern Scotland, England, and Ireland.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Sunday, March 22, 2026

Timeline, and age of the Y Chromosome


This post will look into the "age" or timeline of our Y chromosome. There are different methods used to estimate how this chromosome changes over time by accumulating mutations (*). This leads to a mutation rate expressed in mutations per site per year. It can be calculated by comparing the Y chromosomes of two different men, or a modern human and an archaic one, or another ape and humans, and count the mutations by which they differ, if we know when their lineages split.


(*) Actually, to be correct, a mutation is a chance change in a base pair, an error in the transcription of our genetic code. A substitution is when a mutation becomes fixed within a population, one that has not been erased by genetic drift, or natural selection. For instance, a chance mutation may lead to a "C" to appear instead of a "G", but and then another chance mutation could flip it to a "T" in that case looking at the original and the final versions we'd see one substitution, and imagine one mutation, but there were 2 mutations. Another example would be a flip back, to "C" (back-mutation or reversion), where we'd see no mutation or substitution, but in fact, there was one mutation.


The formula to calculate the Time to the Most Recent Common Ancestor (TMRCA) is shown below. Where the TMRCA in years; k is the number of base pair differences between both men, 2 is a factor added because the difference corresponds to the divergence in both men, half for each one and we don't want to count them twice, L is the total sequence length sampled in which k was detected; and μ is the mutation rate.


TMRCA (years) = k (mutations) / μ (mutations/site · year) · L (sites)· 2


From which the mutation rate can be calculated by shuffling terms as shown:


μ (mutations/site · year) = k (mutations) / TMRCA (years) · L (sites)· 2


A real life example comparing Chimpanzee and Human base pair divergences found in Shen et al., (2000) is the following: assumed TMRCA: 4,900,000 years, acutal values measured (the length sampled in base pairs) L: 38,568, (the number of divergent base pairs) k: 470, calculation of μ shown below:


μ (mutations/site · year) = 470k (mutations) / 4,900,000 (years) · 38,568 (sites)· 2


μ (mutations/site · year) = 1.25 × 10−9.


Assumptions and Data


As you can see, the formula is logic, and straightforward, we take a sample of a DNA strand from a Y Chromosome with a length of L base pairs (my intro to Y Chromosome post explains the basic terms) in both subjects, and then count the differences we observe (k), knowing how long ago both individuals shared their last common ancestor (TMRCA) we calculate the mutation rate μ.


The same method can be applied to any chunk of nuclear DNA including the X chromosome and any of the non-sexual chromosomes or autosomal DNA, as well as the mtDNA. The interesting part is that they all mutate at different rates.


Y chromosome mutation rates (μ)


Over the course of the years, different studies using different methods have attempted to calculate the mutation rate of the Y chromosome in humans. The table below shows some of these studies. The values of μ are given in mutations per site per year. For non-scientists, note that expressing a number multiplied by 10-9 is another way of writing: "divided by 109" and 10 to the ninth power is 1000,000,0000. This means that 1.24 x 10-9 = 1.24 / 1000,000,000 = 0.00000000124 (very small indeed!).


The difference between the first value reported by Thompson et al, and the figure given by Francalcci et al is 2.34 times!


There are big discrepancies in these μ values


Using one or the other to estimate the age of a given specimen would result in one figure being over twice the age of the other one.


The different methods employed to calculate μ involve different assumptions and they all have their shortcomings. In the following commentary I will follow the excellent work of Wang CC, Gilbert MT, Jin L, Li H. (2014) ( Evaluating the Y chromosomal timescale in human demographic and lineage dating. Investig Genet. 2014 Sep 10;5:12. doi: 10.1186/2041-2223-5-12. PMID: 25215184; PMCID: PMC4160915).


The data used for comparing divegencies in the Y chromosome's base pairs can come from different sources: Ancient DNA, samples taken from prehistoric human remains which have been dated by using radiocarbon or other methods. Genealogical, using samples from a certain family for which the genealogy has been confirmed and dated, these usually involve short time spans and few generations. Archaeological Events, taking an estimated date for an event, such as the peopling of America or the settlement in a given region in Europe, and applied to samples from that date.


Confounding Factors


As mentioned in my post on phylogenetic tree branch lengths (a factor that shows that mutation rates are not constant), the value of μ should be considered as an estimation, and not something written in stone. Trombetta et al., (2015) point out that "variants" (mutations) appear at different rates across Y chromosome haplogroups, geographic locations, and time: "... we observed a remarkable heterogeneity in the distribution of variants, not only across different regions, but also across lineages and different times. Hg A00 stands out for showing strong associations with almost all genomic features considered. In the rest of the tree Hg's A0, A1a, A2'3 and B differ from Hg's DE, FC and R, and ancient branches differ from recent ones. The two levels are not entirely independent, as far as recent branches are enriched in lineages belonging to Hg's DE and R. It is possible that different social habits, lifestyles and environmental conditions experienced by populations harbouring different haplogroups resulted in systematic variations of the generation time and average paternal age at conception." They attribute this to older or younger age of the fathers when their children are conceived —older dads have more mutated sperm, and the effects of environment on DNA mutations.


Chimpanzees and Men


When comparing human and chimp Y chromosomes, we are not only separated by a gulf of 5 to 7 million years of separate evoluton, the evolution itself has been different in both species. The chimpanzee Y chromosome is much smaller than that of humans. it lost roughly one-third of its genes in the MSY, or male-specific Y region of their Y chromosomes, compared to men.


Kuroki, Y., Toyoda, A., Noguchi, H. et al. (2006) noted that there is a greater divergence in the sequence of human Y chromosome vs. chimpanzee Y chromosome than between the whole genomes of both species (1.78% and 1.23% respectively).


The reliable dating of the Chimpanzee-Human split is still being debated, and figures range from 4.2 to 12.5 million years ago. A factor of three!


There are also structural differences in the shape of our and the chimp's Y chromosome which complicates the alignment of segments for comparison. Finally, chimpanzees and modern humans have different pair-bonding sexual behaviors (Hughes et al, 2013). Schaller et al., (2010) note that receptive females copulate with multiple male partners creating selective pressure towards the male fertility genes in the chimp's Y chromosome. Monogamous pair bonding in humans lacks this intense selective force. This alters the rate of the mutational clock. The pair bonding in hominins can be seen in the reduced sexual dimorphism in australopithecines and the loss of sperm competition adaptations, suggesting less male-to-male strife, and the growth of cooperation as a means for reproductive success (Gavrilets, 2012).


Genealogical Methods


Pedigree-based studies look at men belonging to the same family, sharing the same paternal lineage, and whose birth dates are known, or at least, how many generations separate them. This provides a well defined dating.


Xue et al., (2009) studied 13 generations of men in haplogroup O3a. However, there are some potential sources of error in this method: Since mutation rates are random, variable, and unpredictable (the statistical term for this is highly stochastic), are we sure that 13 generations (approx. 390 years) is a long enough interval.


Another is the haplogroup itself, we have mentioned that haplogroups accrue mutations at different rates. Is haplogroup O3a a good reference for all other haplogroups?


Finally, if selection and genetic drift act mutations eliminating some of them over longer timescales, the number of mutations found in a genealogical study will be smaller than the actual one detected in longer scales.


Mutation Rates adjusted for autosomal mutation rates


This method was developed by Mendez F., et al., (2013) when they dated the extremely ancient A00 haplogroup Y chromosome, found in the Mbo people of Cameroon, Africa. The μ used by Mendez team was based on a study conducted in Iceland that calculated the autosomal mutation rate by analyzing the divergence between parents and their children. This implies several unverified assumptions: autosomal and Y chromosome mutation rates are similar (they are not), substitution rates and mutation rates are equivalent (they are not). They also used a generation time that spanned from 20 to 40 years when life expectancy for men in Cameroon in 1950 was below 40! (Source). Other authors using genealogical data, like Boattini et al., (2019) found generation lengths of 33.57 years. Hunter-gatherer groups in Equatorial Africa 150,000 years ago probably began mating at the age of 15. How can we know for sure? (See Elhaik E, Tatarinova TV, Klyosov AA, Graur D., (2013) and their critique to Mendez et al.


Older generation times lead to more mutations, and an overestimated TRMCA. A00's age is far shorter than the one put forward by Mendez et al.


Archaeological Data


In my posts, I have mentioned many ancient DNA samples taken from the remains of prehistoric, ancestral human beings and hominins (Yana River, Mal'ta and Ust'-Ishim). They provide a certain age with reliable radiocarbon-dating, and if the DNA is not degraded or contaminated, a reliable count of mutations.


Fu, et al., (2016) compared the remains of the Ust'-Ishim man with those of men alive today, and looked for "missing" mutations, those that appeared in modern men after Ust'-Ishim died. The team calculated a mutation rate of 0.76 × 10-9


Human Migration approximations


Assuming dates for certain migratory events, like the peopling of America, placed at around 15,000 years ago, the dates for certain haplogroups found exclusively in America can be set close to that event. This has been used to calculate μ for splits between Asian and Amerindian lineages, or between Amerindian groups within America, obtaining a μ of 0.820 × 10-9 (Poznick et al., (2013)).


A similar method was applied by Francalacci et al., (2013) to men native to the island of Sardinia in Europe, peopled around 7,700 years ago, and using the mutations detected in a sample of Sardinian men to calculate μ for their haplogroup (I2a1a). The value obtained was very low compared to those shown in the Table further up: 0.530 × 10-9.


Can we be certain that the haplogroup diversified in its current location in Sardinia or America? Could it have taken place earlier (European mainland, or Siberia, respectively). Do the current Sardininans belong to the group that reached the island 7.7 kya? Or did they arrive later?


Implications


The sum of these factors show that environment, generation times, sexuality (pair bonding), natural selection, and genetic drift can promote or reduce mutation rates. Studies have a 2.4-fold variability, ranging from 1.24 × 10-9 Thompson et al., (2000) to 0.53 × 10-9 Francalacci et al., (2013), and these are the mean values, the confidence intervals are even wider. Supposing a threefold difference, what some estimate as a divergence taking place between two Y chromosome haplogroups 50,000 years ago during the final Out of Africa migration, may have taken place within Africa 150,000 years ago! And the A00 split instead of taking place 250,000 years ago may reflect one that ocurred 750,000 years ago.


The possibility that our most ancient ancestors had more chimp-like behavior would have implied a faster mutation rate during that period, followed by a slower rate later on. Even pair-bonding during the patrilineal, sedentary, agricultural period of the past 8,000 years may have slowed down mutation rates in comparison to the matrilineal hunter-gatherer period that preceeded it.


genetic mutation rates
Mutation Rates and their impact on TMRCA inference. Copyright © 2026 by Austin Whittall

The image above shows how mutation rates can influence the age of the TMRCA. For the same measured value of "m" mutations marked by the gray line, the use of different mutation rates μ influences the depth or age to the TMRCA. A quick mutation rate like the blue one (μ1) accumulates mutations quicker and takes less time to reach the detected "m" mutations, so its TMRCA is younger (T1), a slower mutation rate (red line) like μ2 takes longer to accumulate the "m" mutations and its T2 is longer. Finally, a variable mutation rate like μ3 (green line) that was faster in the distant past, and slower later on will have an even longer and older timeline (T3).


For those interested in maths, the slope of each curve marks the mutation rate dm/dt = μ(t), steeper curves mean quicker mutation rates. It an analogy of speed as the differential of space over differential of time.


I believe that the μ3, variable mutation rate is closer to reality, reflecting variable social-sexual-cultural patterns of the small and egalitarian promiscuous hunter-gatherer groups of early hominins and human evolution. Later settled societies with paternal monogamy and private property led to slower mutation rates. This of course would imply an older root for the human Y-chromosome haplogroups, in Africa, prior to our migration into Eurasia.




Fall has begun here in Buenos Aires, in the Southern Hemisphere! Lucky you who are now entering Spring!



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Saturday, March 7, 2026

Early Peopling of America ~40 kya


This will be a very short post. I came across a paper published in Nature in 2020, by V. M. Cabrera (Counterbalancing the time-dependent effect on the human mitochondrial DNA molecular clock. BMC Evol. Biol. 20, 1–9 , 2020) which, as you can see by its title is about the uneven rate at which mutations accumulate in mitochondrial genomes, it can run faster or slower and the research cited by Cabrera attributes it to "changes in the effective population size of the human populations." Another factor mentioned by Cabrera are " "transient polymorphisms... play a slowdown role in the evolutionary rate deduced from haplogroup intraspecific trees".


Transient polymorphisms


These are short-lived variants that through selection take over and displace another. Later, they are themselves replaced, hence their name (transient). An example is the classical moths of England, they came in two varieties, light colored and dark. Before the industrial revolution, the light ones prevailed, as the dark ones were very visible (easy preys) on the pale tree trunks. Then, when the industrial soot of coal furnaces darkened the bark, the situation changed, and the dark ones were less visible. Soon they displaced the ligher colored moths (dark allele replaced pale one). However, in the 1970s, with pollution abating measures in place, the bark of trees became pale once again, and the light colored moths now prevail over the dark ones.


Regarding variability as time passes, Cabrera mentions an interesting fact: "For example, an ancient mtDNA study has corroborated empirically the persistence of an ancestral M lineage unaltered along a period of more than 8000 years" and adds that "Consequently, a lineage could remain immutable for several generations while identical lineages in the same population suffer one or several mutational changes in the same time interval."


This situation is shown in the paper's Figure 2, shown below, which very clearly shows how mutations can be overlooked when building a tree, and its effect on the mutation rate that is inferred from it. Cabrera describes the figure as follows: "We represent such a scenario in Fig. 2A as a maternal genealogy. Notice that the ancestral lineage can give rise to offspring in different generations throughout its existence in the population. In this way, when the population is sampled after n generations, we can find, in addition to the ancestral lineage (e), lineages derived from it that have accumulated significant mutational differences in their branches (f, h). However, this fact is not reflected in the tree built from the same sample (Fig. 2B) because, irrespective of the generation in which they appeared, all derived lineages sprout at the same time from the ancestral node (e). This difference between genealogies and trees has notable consequences. On the one hand, it could explain the lack of mutation rate homogeneity between lineages found in intraspecific haplogroup.""


phylo and genealogical trees
Comparison between a genealogy (A) and a tree (B) constructed from the same sample (a to h). White circles are individuals with the ancestral lineage, black circles are individuals with additional mutations. Inter-circle segments represent generations and crosses on the segments represent mutations.. Figure 2 in Cabrera.

Cabrera then calculates the ages of the different mtDNA haplogroups, and finds a relatively old age for the entry of humans into America:


"Finally, although we proposed a unique migration for the colonization of the Americas around 40,000 years ago (20) which is directly or indirectly supported by archaeological dates, it seems possible that this first migration, signaled by the ages of haplogroups A2 and B2 was followed afterwards by a second wave, also before the Last Glacial Maximum, marked by haplogroups C1, D1, D4h3a and X2a around 27,500 years ago (Table 3).


Note 1. The figures from Table 3 are these: C1: 30 ky, D1: 32 ky, D4h3a: 26 ky, and X2a: 22 ky.


Note 2. Cabrera in (20) is citing himself! (Cabrera, V. M., 2020, Counterbalancing the time-dependent effect on the human mitochondrial DNA molecular clock. BMC Evol. Biol. 20, 1–9.)


This second paper by Cabreera, cited above as (20), states that "Finally, the time of human expansion to the American Continent deduced from the haplogroup B2 phylogeny was approximately 37,000 ya (Table 1, and Table S11). This age supports a pre-Clovis occupation of the New World, well before the last glacial maximum."


I am happy to see someone who is pushing the boundaries towards an older date, 40-37 ky ago. This is positive. It will open the door for others to use this information to validate their findings of an even earlier date for the arrival of modern humans to America.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Thursday, November 13, 2014

On Mutations, mutation rates and Ust'-Ishim


A week ago I read the Supplementary Information on the 45 ky old Ust'-Ishim gene sequencing and also several posts that dealt with this very interesting paper.


I was intrigued by two points: one was their estimation of human mutation rates and the other was the autosomal "diversity" of modern and ancient humans as shown in SI 12. So I decided to write a post on each of those subjects. Today's post looks into the question of "mutation rates", the diversity issue will be the subject of a future post.


Molecular clocks and mutations


It is no secret that I am very skeptical about molecular clocks which tick with a constant rate and therefore allow us to measure the timing of past events such as splits between species or the dates when a given haplogroup appeared. So please read on with this in mind: I don't trust molecular clocks.


The key issues in today's post are:


  1. Mutations in humans take place at different rates. Yes, mtDNA, Y chromosome and autosomal DNA mutate at different rates when compared to each other, and that is ok, and explainable. What I mean by different rates is that when we compare the same kind of genetic material: mtDNA against mtDNA or autosomal DNA against autosomal DNA, in different human samples, the rates are different.
  2. Genetic mutations grow at different paces in men and women.
  3. The DNA of our closest relative, the chimpanzee mutates at a different rate, when compared to ours.
  4. More genetic diversity may not mean "older" populations but simply quicker mutation rates that accumulate in a population of the same age as a less diverse one.
  5. The concept of a molecular clock is unsustainable.

The "molecular clock" that is used to estimate past events in human evolution is based on the simple assumption that: mutations take place in our DNA at a certain fixed pace and that by comparing the quantity of mutations by which specimen A and specimen B differ, we can calculate the time that has elapsed since they shared a common ancestor.


Neat. But what exactly does it mean? and even more important: is it really a clock?


First of all, lets see what a mutation is:


We are all familiar DNA and its role in inheritance. DNA is a polymer that is found inside each and every one of our cells, inside the cellular nucleus. It is shaped like two intertwining coils (a "double helix") and these spirals are stabilized by molecules known as "nucleobases" (or "bases" for short) which link the coils together.


One base is attached to one of the coils and another base is attached to the other coil; and both bond together to form a "base pair".


Fortunately these bases come in four kinds: adenine (abbreviated A), cytosine (C), guanine (G) and thymine (T) and they link up in a very simple manner:


A only bonds with T and C only links up to G. So the "steps" of the ladder that joins the two spiral strands is made up of base pairs such as:


....
AT
TA
CG
CG
AT
GC
TA
....


During replication, the DNA strands within the nucleus of the cell, "unwind", unzipping each strand. The exposed bases attract a new pair and form the other complementary spiral. The "unpaired" Adenine will attract thymine and link to it while the guanine will bond to cytosine and so forth. This bonding takes place on each of the unzipped strands, so from one (1) initial DNA molecule, two (2) "identical" copies are obtained.


Copies and errors, causes


Well, not exactly "identical", there are some sequence errors, and this takes place even though cells have proofreading abilities and mismatch repair mechanisms that compare original and copied DNAs.


These errors in the transcription mean that the "daughter" DNAs are not an exact replica of the "mother" DNA. So maybe an AT is lost or replaced by a GC or a duplicate is inserted so TA becomes TA TA...


In other words these are natural spontaneously occuring mutations.


Other external factors known as "mutagens" can interact with the DNA and alter the nucleotide sequence producing mutations:


  • High Energy Electromagnetic radiation. For instance X-rays or Ultraviolet light (UV).
  • Oxidizing agents (or free-radicals).
  • Chemicals, such as alkylating agents, heavy metals, solvents, monomers, agrochemicals.

These factors can disrupt the sequence of a DNA strand resulting in a "new" (mutated) chunk of genetic information, either due to deletion or addition of base pairs.


There is also a process known as "Recombination" by which two DNA molecules merge in certain sections and produce a totally new variety of DNA.


Where (and when) do mutations take place?


Those terrible summer sun-burns that I suffered as a child back in the 70s, and the accumulated doses of UV radiation that my skin has received over the years, may result in a mutation that causes skin cancer (I keep my fingers crossed and visit my dermatologist once a year just in case).


The same could be said about the X-rays that have zapped me at my dentist or during my medical check-ups or the cosmic rays that incessantly criss-cross my body. Some pesticides that I ingested via fruits or cereals, those glasses of red malbec (and its metabolized by-products) or the cigarrettes that I used to smoke are also packed with mutagens... which disrupt my DNA here and there.


But all of these mutations which are taking place in some of the 40 trillion cells that make up my body would only affect me and therefore would not be passed on to future generations unless they took place in certain cells and, during a specific time frame which differs for men and women.


The new DNA information created by mutations would pass on to my progeny only if it mutated inside my "germ cells" (sperm in my case since I am a man or, for women: their ova).


Mutations that are inherited: Meiosis


Normally our cells reproduce by a process known as "Mitosis", and its outcome are two cells, each carrying the same genetic information that the mother cell had, as well as the same number of chromosomes.


We humans are dipolid organisms and as such, our cells carry two homologous copies of each chromosome, one inherited from our mother, the other from our father. Our cells therefore have 23 pairs of chromosomes of which 22 pairs are autosomes, one pair are the sex chromosomes (the X and Y chromosomes, paired XX in women and XY in men ). The grand total is 46 chromosomes per normal human cell


But our "germ cells" are different, they can only carry half of the genetic information of each parent so that the combination of the father's sperm and the mother's ovum with their chromosomes add up to exactly the full number of chromosomes.


This process of germ cell formation is known as "Meiosis", and it takes place differently in males and females:


  • Females. Meiosis in females is known as "oogonia", and consists of a series of divisions of the "original" oogonium with the complete set of chromosomes, the outcome is an an ovum with half the quantity of chromosomes.
    Meiosis in females takes place during the formation of the embryo (after the fourth week of pregnancy), as soon as the primordial germ cells migrate to the ovary, and they will lie dormant inside a protective follicle until the woman reaches puberty, when her menstrual cycle begins.
  • Males. The process in males is known as "spermatogenesis", and takes place in a continuous manner, after puberty, until death. Meiosis produces spermatozoa in the seminiferous tubes inside the testicles.

Implications


Since male germ cells are produced in a constant manner, the different mutagens that interact with an individual during is whole adult life, (meiosis is a continuous process in men) may cause mutations. Furthermore, males have a very poor DNA repair mechanism, so these mutations are more likely accumulate, without being "fixed", and therefore more likely to get transmitted to their offspring.


The female ova, on the other hand are produced during fetal growth, and are placed in hybernation for many years until the onset of puberty. But despite this long period during which external mutagens could interact with the ova's DNA producing mutations, females have a very efficient repair mechanism for postmeiotic stages which can repair DNA until after fertilization. [3]


This means that sperm accumulate more mutations than ova, and men transmit more mutations to their offspring than women do. This has been corroborated by separate studies both in humans and in chimps:


Chimpanzees, our closest relatives have different mutation rates


Chimps are our closest primate relatives, and a recent paper (Venn et al., 2014) [1] found that "mutation rates and patterns differ between [our] closely related species", Venn reported that male chimpanzees pass on between seven and eight times more mutations to their offspring than do female chimps. This means that roughly 88% of the mutations found in their offspring have a paternal origin and 12% are maternal.


Also, the older the father, the more the mutations in the paternal genes ageing adds "three mutations per year of father's age"[1].


In humans on the other hand "every additional year of father’s age contribut[es] two mutations across the genome and males contribut[e] three to four times as many mutations as females." [1], so males provide between 75 and 80% of the mutations and females 20 - 25%. /p>

This increased mutation contribution by males was reported in a genetic study by Campbell et al., 2012 [7] which found among Hutterite families that 85% of the new mutations were of paternal origin. Which is higher than those mentioned by Venn for humans and very close to those of chimps.


Campbell reported the following figures:


SNV mutation rate: 1.20 × 10-8, (95% confidence interval 0.89 – 1.43 × 10-8) mutations per basepair per generation. And 0.96×10-8 for the most recent generation.


The lower mutation rate for the latest generation was justified by "the relatively young age of the father of the trios analyzed here (21–30 years old at the time of the child’s birth)" [7], meaning that a younger father had accumulated less mutations than an older one.


Calculating Mutation Rate


The mutation rate can be calculated using the following formula:


Mutation Rate = # of mutations observed ⁄ (# of generations x # of base pairs sequenced) [a]


By comparing the discrepancies in the gene sequences of two related individuals separated by a given time span the mutation rate can be calculated (i.e. father - son pairs or comparisons of the DNA between living individuals and that sequenced from his ⁄ her ancestors).


Roach et al., (2010) [8] analyzed the full genome sequence of a family (two children and their parents) and calculated a mutation rate of 1.1 x 10-8 per position per haploid genome.


But what does this mean? Look at it this way: humans have about 6 x 109 base pairs (six billion), so it is very straightforward to work out the number of mutations that will appear in a child, inherited from its parents. They are (see [a] above) directly proportional to the number of bases, the number of generations - in this case = 1 - and the mutation rate. So, using [a] we can calculate:


# of mutations observed = Mutation Rate x # of generations x # of base pairs


So, replacing the terms with actual numbers:


# of mutations = 1.1 x 10-8 x 1 x 6 x 109


# of mutations = 66 (the new mutations in a child, compared to its parents).


Comments


Out of these 66 mutations, roughtly 80% (or 53 are parental, the other 13 maternal). Since paternal mutations grow at a rate of 2 per year [1], we can see that the child of an "old" dad aged 45 would receive 2 x (45-20) = 50 "extra" mutations in its genome in comparison to the child of a "young" 20 year old father.


So the baby of "old" dad would have 66 + 50 = 106 mutations while the baby of the "young" dad would have only 66 mutations.


Looking at a society where older parenting prevails (young males are not successful in mating with the women, or they die off before bearing children, or their children die off before reaching maturity) we would find that mutations would have accumulated at 106 mutations⁄generation. While a society where young males exclude older ones from bearing children, the mutations would accumulate at a rate of 60 mutations⁄generation.


After "n" generations the situation would be:


Old men society: n x 106 mutations. Time span: n x 40.
Mutations per year: n x 106 ⁄ n x 40 = 106⁄40 = 2.65


Young men society: n x 60 mutations. Time span: n x 20.
Mutations per year: n x 60 ⁄ n x 20 = 60⁄20 = 3.00


On a "per generation" basis there are 76.7% more mutations in the "old men" society, but on a "per year" basis, the mutation rate is 13.2% higher in the "young men" society.


This should be a word of caution when using "generations" to gauge ancient events. The conclusions will be very different in one case or the other.


The variability of Mutation Rates


The problem is that the mutaton rates are quite "variable". A paper by Wang, J. et al., (2012) [9], sequenced individual sperm cells in a 40 year-old individual. They obtained mutation rates of 2.0 to 3.8 x 10-8, which, are different from other values measured in other studies:


Hutterites (mentioned above): their "SNV mutation rate [was] 1.20 × 10-8 (95% confidence interval 0.89-1.43 × 10-8) mutations per base pair per generation." [7]


Pedigree. Xue et al., (2009) [6] compared the mutations detected in the Y chromosme of two members of the same family separated by a span of 13 generations. They reported that "The mutation rate is ... 1.0 × 10-9 mutations ⁄ nucleotide ⁄ year (95% CI: 3.0 × 10-10 – 2.5 × 10-9), or 3.0 × 10-8 mutations ⁄ nucleotide ⁄ generation (95% CI: 8.9 × 10-9 – 7.0 × 10-8)" [6] .


Just look at Xue et al.'s enormous Confidence Interval: it is almost one order of magnitude, that is the upper limit is nearly 10 times the value of the lower limit! This is like saying that we estimate the weight of the stone to be 10 pounds, with a CI of 3 lb - 25 lb.


The range between the minimum and maximum values of these studies goes from 0.89 to 7 x 10 -8 mutations ⁄ base pair ⁄ generation. A big window indeed.


Variability between families


Additional proof of the variability of mutation rates comes from a paper (Conrad et al., 2011) [5] which confirms that there is "considerable variation in mutation rates within and between families". The authors compared female and male germline and non-germline de novo mutation rates and found that "in one family [...] 92% of germline DNMs were from the paternal germline, whereas, in contrast, in the other family, 64% of DNMs were from the maternal germline." [5]


Against the constancy of mutation rates


We could imagine that additional research will refine those values and come up with a more reliable rate, but there is another problem: the mutation rate is not constant. It fluctuates accelerating and slowing down over time.


A paper by Amos W., (2013) points out a that "tendency for Africans to have diverged more from chimpanzees than non-Africans is unexpeced under classical theory." [4] Since we all derived from chimps, and have had the same time to accumulate mutations, why do Africans appear more distinct?


Applying the formula [a] it is quite simple to see that if the number of mutations is higher for Africans, then either the number of generations and ⁄ or the mutation rate must be higher for them than among non-Africans. Amos finds no reason to imagine a shorter generation time in Africa (shorter duration means more generations in a given time span). Clearly something is influencing the mutation rates in Africa.


So for Amos, this implies that the anomaly may be due to two reasons "local effects that vary across the genome due, for example, to natural selection, and genome-wide effects arising from a mutator allele impacting mutation rate or demographic influences that alter generation time." [4], the paper suggests that the mechanism that is acting to distort mutation rates is known as the "heterozygote instability" (HI) hypothesis.


Under the HI hypothesis:


mutation rate increases at and near heterozygous sites where the two homologous chromosomes differ in sequence [...] [4]


This means that "gene conversion events focused on heterozygous sites during meiosis locally increase the mutation rate [and] As humans left Africa they lost variability, which, if HI operates, should have reduced the mutation rate in non-Africans. [4]


In other words, the bottle necks that decimated humans (and also their Neanderthal and Denisovan) predecessors as they moved out of Africa and across Eurasia, led to reduced diversity and this in turn decelerated their mutation rate in comparison to those that remained in the African homeland whose diversity remained higher.


"Under the HI hypothesis, this demographically-induced reduction in heterozygosity should create a parallel reduction in mutation rate such that Africans have diverged more than non-Africans from their common ancestor" [4]


I will go over this in detail in my next post on "African diversity vs. non-African lack of diversity", but focusing on this post's subject, how does this impact on mutation rates?


Simple: mutation rate grow with increasing heterozygosity. Also "When population size is constant, smaller populations will experience lower mutation rates than related larger populations" [4].


Summary on mutation rates


DNA mutates at different rates:


  • In male or female germ cells
  • In different families
  • In less diverse populations vs highly heterozygous populations (HI hypothesis)
  • In chimpanzees (vs. humans)
  • In older men's sperm vs. younger men's sperm
  • In large populations vs. small populations

And as we will see below, ancestral DNA (obtained from the remains of ancient humans) show different rates when compared to the pedigree rates calculated by using sequences of recent modern families.


Ust'-Ishim and its estimates on mutation rates


And now, we get to the paper on the Ust'-Ishim remains (Fu, et al., 2014) [2]. Besides a wealth of data on admixture and mtDNA & Y chromosome haplogroups also deals with mutation rates, and reaches some very interesting conclusions (Below I will refer to the Supplementary Information freely available online):


Autosomal Mutation Rates Estimates


The paper measured how many mutations are "missing" in this 45 ky old individual when compared to contemporary humans. The logic behind this calculation is that we kept on evolving during that period of time and accumulated new mutations (see SI 15). Since the bone was carbon dated (41,410 ± 960 BP or 45,000 cal BP) and the substitutions can be measured, the calculation was relatively simple.


As expected, the DNA of Ust'-Ishim is around 0.6% shorter than the A-Panel (a low-coverage of 24 - 32%) modern humans and 0.35% for B-Panel samples (with a higher coverage of 35 -42%). So the autosomal mutation rates are: "0.80-0.91 × 10-9 ⁄ bp ⁄year for panel-A and 0.44-0.63 × 10-9 ⁄ bp ⁄ year for the B-panel." [2]


The authors therefore "estimate a nuclear mutation rate of 0.44 -0.63 × 10-9 ⁄ site ⁄ year, which is lower than the value that has been widely used in the past (1 × 10-9)." they do point out that the "lower quality A-panel individuals give significantly different results, indicating that this measure is sensitive to quality differences between the compared genomes" [2]


Indeed "different", the values differ by a factor of 2! Notice that they are also giving their figures in mutations per site per Year, further up, the figures we mentioned were "per Generation". Conversion from one to other depends on the time span assigned to a Generation which can range from 19 to 40 years!


Comparing Ust'-Ishim (ancestral) and Xue's pedigree values [4] there is a two-fold spread between the minimum and maximum values.


  • 0.44 - 0.63 x 10-9 ⁄ bp ⁄ year (Ust'-Ishim)
  • 0.30 - 2.50 x 10-9 ⁄ bp ⁄ year (Xue's data)

Mitochondrial DMA mutation rates


Using the mtDNA (SI 8), Fu et al., (which by the way, the mtDNA "appears to be most closely related to the direct R sub-clades R* (P, B, F, T, J)" [2]), they also estimate a mutation rate of "2.53 × 10-8 substitutions per site per year (95% HPD: 1.76 -3.23 × 10-8) for the complete mtDNA" [2]. Notice that it differs from the autosomal mutation rate calculated above by a factor of about 50 corroborating that mtDNA mutates rapidly.


Y-chromosome mutation rate


The team also sequenced Ust'-Ishim's Y chromosome and found that it "clusters with the K(xLT) haplogroup." [2], they estimated its mutation rate as "0.76 × 10-9 substitutions per site per year (95% HPD: 0.67-0.86 × 10-9)". Which was higher than the rate reported in the controversial paper by Mendez et al. (2013) which discovered a new Y chromosome lineage (A00) with a Most Recent Common Ancestor (TMRCA) of 338 ky, far older than the oldest anatomically modern human fossils.


Table S9.1, shows that the (TMRCA) for all Y-Chromosomes as 153 ky old (range: 132-175 ky).[2], since the mutation rate used is higher, the TMRCA is much more recent than the figure calculated by Mendez (153 ky vs. 338 ky).


A novel calculation of the mutation rate


Fu et al., devised a new method of calculating the mutation rate "assuming the population size history of the ancient sample is identical to that of present humans prior to the death of the archaic individual" [2] , and estimated a muation rate of 0.43×10-9 per site per year, with a 95% CI (0.38×10-9 - 0.49×10-9).


This value coincides with their estimate based on another method and is much lower than other previous values. The fact that the mutation rate is lower and therefore slower means that more time is necessary to accumulate the mutations that we carry, in other words it suggests an older date for the split between modern and ancient humans.


Another recent paper on Mutation Rates


A few days after Fu's paper, another one (Rieux et al., 2014) [10] was published, it reported an improved calibration of the mtDNA clock, and calculated a TMRC for modern humans as 143 ky (95% CI 112 - 180 ky), which agrees pretty well with Fu's estimate.


They ratified the "acceleration of substitution rates in recent times" (for mtDNA mutation rates) which is explained by the " 'time-dependency of molecular rates' hypothesis, which postulates an acceleration over recent times in coding sequences due to the time needed for selection to purge slightly deleterious mutations" [10].


They used two types of estimations, one based on the dates of the fossil remains of ancient humans (tip-based-estimates) and another calibrated on nodes, which mark peopling events such as the peopling of America, New Zealand, Madagascar, etc. They noticed that:


tip-based rate estimates are slower (by a factor of 0.63-fold) than the ones obtained using internal node calibration by Endicott and Ho (2008) but are faster (by a factor of ~1.5-fold) than previous fossil-calibrated rates
[...]
the variance over individually calibrated substitution rates is 11 times smaller for tips than internal nodes. Moreover, all of the 21 substitution rates estimated from aAMH sequences had overlapping 95% HPD (fig. 3). The situation is strikingly different for node-based calibrations, where substitution rate estimates strongly depended on the demographic episode used for dating, with only four out of ten individually calibrated rates having overlapping HPDs.
These results strongly suggest that tip calibration estimates are far more consistent than internal node-based ones. However, tip-based calibration also point to slower mean substitution rates than those based on internal nodes. Thus, one important question we need to answer is whether tipbased calibrations are affected by some systematic bias that might lead to slow (and homogeneous) substitution rates. [10]


In other words, the "nodes" or estimated dates of demographic events have a larger variance than those of the "tips" (reliably dated fossil remains). The 95% HPD interval obtained for each independent ancient sequence overlapped the others while the "nodes" only overlapped in 40% of the cases. (See their Fig. 3). The ancient remains are therefore more reliable than the estimated dates for peopling events! (something I have written about several times suggesting an ancient peopling of America). Additionally the fossils indicate a slower mutation rate meaning a more ancient date for all events (African - non-African split, Sapiens - Neanderthal split, etc.).


Rieux et al. recognize that carbon dated remains are much more reliable than the estimations of dates regarding peopling events (allow me to quote them extensively) :


the uncertainty around dated nodes is far more complex and multifactorial and is likely to lead to different degrees of reliability associated to each node. First, there is generally considerable uncertainty associated with the age of the colonization⁄migration event including error in the dating of the archaeological, anthropological, and historical evidence. The age of the oldest evidence for human presence is unlikely to coincide exactly with the demographic expansion. Very generally, we would predict to see a delay in the appearance of traces of human presence after the expansion of AMHs into any new area (Signor and Lipps 1982).
[...]
Second, even if a demographic event had been accurately dated, the age of the node in the phylogenetic tree might not coincide with it for a number of reasons (Edwards and Beerli 2000; Ho and Phillips 2009; Balloux 2010; Firth et al. 2010; Crandall et al. 2012). For instance, the phylogenetic node of interest may correspond to the most recent common ancestor (MRCA) of the sampled sequences rather than the split of the population of interest.
[...]
the population might have experienced a reduction in size later on, so that the TMRCA could coincide with this subsequent population bottleneck. We could think of additional scenarios and the situation would become even more complex if we considered a possible effect of natural selection. To summarize, node calibration can be affected bymany sources of error, and it is thus nearly impossible to model the age uncertainty around nodes satisfyingly. [10]


In other words, the nodes will give "later" dates for cases like America, subjected to a drastic bottleneck. Even so, I am quite happy to see the dates that they estimated using ancient genomes gave a much older date for the peopling of America than the usual "orthodox" 13 - 17 ky:


time peopling America
Table 2 in [10]. Coalescence Times for Major Haplogroups Involved in the Colonization/Migration Events Considered.

Yet the authors are aware of their "early" dating and try to explain the "discrepancy" to conform to orthodoxy: "However, in the case of the Canary Islands, Remote Oceania, New Zealand, and the Americas, the estimated coalescence times were systematically older than the archaeological evidence. Potential explanations for such discrepancies include ancestral polymorphism in the founding population or complex demographic histories involving multiples wavesof colonists" [10].


Finally their dating of the Divergence between Humans and Chimpanzees was 4.14 Ma (95% HPD 2.991 - 5.448). Which they admit "... may appear too young when compared with the dates that are generally derived from the fossil record." [10] . The authors attempt to explain this "recent" date as due to (i.e. "...more complex speciation scenarios where an initial split was followed by an extended period of gene flow before the final separation..."). But I do not think that it is due to a flaw in their method, but to the dissimilar mutation rates of humans and chimps, as pointed out by Venn et al., in June 2014, [1]:


Under a model in which the mutation rate increases linearly with parental age, the rate of neutral substitution is the ratio of the average number of mutations inherited per generation to the average parental age. We predict the neutral substitution rate to be ~0.46 × 10-9 per base pair (bp) per year in chimpanzees, compared to estimates in humans of ~0.51 × 10-9 bp-1 year-1 (9). These results are consistent with near-identical levels of lineage-specific sequence divergence (12) but surprising given the differences in paternal age effect. In the intersection of the autosomal genome accessible in this study and regions where human and chimpanzee genomes can be aligned with high confidence, the rate is slightly lower (0.45 × 10-9 bp-1 year-1) and the level of divergence is 1.2% (13), implying an average time to the most common ancestor of 13 million years, assuming uniformity of the mutation rate over this time (95% ETPI 11 to 17 million years; table S11). [1]


This is in agreement with the estimate of Langergraber et al., (2012) who dated the Pan-Homo divergence to between 6.78 and 13.45 Ma.


An earlier human-chimp split renders useless the calculations that calibrate molecular clocks based on that event. If it took place 13 Ma instead of 6, the clock's ticking rate has to be adjusted and all events derived from such a clock would actually be much older than currently accepted. Including the peopling of America.


More to follow, on the differing diversity of humans Africans and non-Africans


Sources
[1] Oliver Venn, et al., (2014). Strong male bias drives germline mutation in chimpanzees. Science 13 June 2014: Vol. 344 no. 6189 pp. 1272-1275 DOI: 10.1126/science.344.6189.1272
[2] Fu, Q. and many others. (2014). Genome sequence of a 45,000-year-old modern human from western Siberia. Nature, 514, 445-450. doi:10.1038/nature13810
[3] Andrew J. Wyrobek et al., (2007) Assessing Human Germ-Cell Mutagenesis in the Postgenome Era: A Celebration of the Legacy of William Lawson (Bill) Russell. Environ Mol Mutagen. Mar 2007; 48(2): 71–95. doi: 10.1002/em.20284
[4] William Amos, (2013) Variation in Heterozygosity Predicts Variation in Human Substitution Rates between Populations, Individuals and Genomic Regions. April 30, 2013DOI: 10.1371/journal.pone.0063048
[5] Donald F Conrad, et al., (2011). Variation in genome-wide mutation rates within and between human families. Nature Genetics 43, 712–714 (2011) doi:10.1038/ng.862. Published online 12 June 2011
[6] Yali Xue et al., (2009). Human Y Chromosome Base-Substitution Mutation Rate Measured by Direct Sequencing in a Deep-Rooting Pedigree. Curr Biol. Sep 15, 2009; 19(17): 1453–1457. doi: 10.1016/j.cub.2009.07.032
[7] Catarina D Campbell, et al., (2012). Estimating the human mutation rate using autozygosity in a founder population. Nature Genetics 44, 1277–1281 (2012) doi:10.1038/ng.2418
[8] Roach JC, (2010). Analysis of genetic inheritance in a family quartet by whole-genome sequencing. Science. 2010 Apr 30;328(5978):636-9. doi: 10.1126/science.1186802. Epub 2010 Mar 10
[9] Wang J, Fan HC, Behr B, Quake SR, (2012). Genome-wide single-cell analysis of recombination activity and de novo mutation rates in human sperm. Cell. 2012 Jul 20;150(2):402-12. doi: 10.1016/j.cell.2012.06.030
[10] Adrien Rieux, et al., (2014). Improved Calibration of the Human Mitochondrial Clock Using Ancient Genomes. Molecular Biology and Evolution, 2014, 2780-2792, DOI: 10.1093/molbev/msu222


Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2014 by Austin Whittall © 

Friday, May 23, 2014

Generation time is not 25 years


This post is just to provide some additional data regarding the "length" or "duration" of a generation. As seen in yesterday's post (which criticises Y chromosome mutation rates), calculations are based on a 25 year generational interval, which when input into the formulas used to calculate the ages of Y chrmosome lineages will give an incorrect date if, (and that is what we will clarify in this post) generations are either longer or shorter than 25 years.


Generations last more than 25 years


A long term study by Nancy Howell among the !Kung of Namibia and Botswana revealed that this contemporary hunter-gatherer group of people are relatively old at the time of bearing children: For women it averages 25.5 years, but for men (and this is important when it gets down to Y chromosomes): 31 to 38 years averaging 34.5 years. These people are very similar to the pre-agricultural society of our distant ancestors. [1]


A paper (Matsumura and Forster, 2008) [2] found that, among Eskimos, the father-son interval is 32.1 years. And point out (bold face is mine) : "The majority of the previous studies assumed that the generation time for mitochondrial DNA and Y-chromosomal DNA is 20 and 25 years, respectively (e.g. Harpending & Rogers 2000). We suggest that a higher value, 25–30 for mtDNA and 30–35 years for Y-chromosomal DNA, should be used in genetic inference." [2].


Fenner (2005) [3] indicates a male generation length of 31 to 32 years. [3]


So, it seems that 25 is too short a time for male generations, the value is at least 30 years, and perhaps higher. In polygynoous societies, older men would have monopolized women and the generation length have been even longer.


Sources


[1] N. Howell, (2009). Demography of the Dobe Kung. Transaction Publishers.
[2] Matsumura S, Forster P., (2008). Generation time and effective population size in Polar Eskimos Proc. R. Soc. B 7 July 2008 vol. 275 no. 1642 1501-1508
[3] Fenner JN, (2005) Cross-cultural estimation of the human generation interval for use in genetics-based population divergence studies. American journal of physical anthropology 128. doi: 10.1002/ajpa.20188


Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2014 by Austin Whittall © 

Thursday, May 22, 2014

Y Chromosome mutation rates


In my previous post I pointed out the differences found between the ages of different branches of the Y chromosome's Q haplogroup, and how papers tend to date Native American lineages so that they coincide with the date that mainstream science considers the correct one for peopling America: ca. 15 kya.


This made me wonder what certainty do we have of their accuracy, or how precise are these dating methods. The answer is surprising!: not very precise.


The complexity encountered in dating Y chromosome lineages is summarized very well by Chuan-Chao Wang and Li Hui (2014): "... Different time estimation methods use different algorithms and assumptions, thus alternative methods probably fit more or less well with sequence data in time estimations. In addition, the best-fit mutation model might vary for different STRs... some specific lineages might have their own unique best-fit STR mutation rates for time estimation." [4]


In other words, it is extremely fuzzy. Today's post will look into the issue of "dates", the calculations of haplogroup ages, of TMRC, and the fallacy of a "clock" behind Y chromosome mutations.


The complexities behind the Y chromosome


The Y chromosome is particular because it is passed along from father to son, basically unchanged -excepting random mutations- in a long line that links all modern men to an "ancestral Adam" who lived in Africa in the distant past of mankind and from which all Y chromosome haplogroups derive.


Y is one of the sex chromosomes found in mammals, and obviously, humans; the other is the X chromosome. An X chromosome and a Y chromosome ("XY") pair detrmine a male, a double X ("XX"), a woman.


Just like all other chromosomes, the Y chromosome also mutates: chance mutations and natural selection act upon it and create small differences that, if not negative (that is, killing its carrier), are passed on to the next generations.


Y chromosome's high mutation rates


The Y chromosome mutation rate is much higher than that of autosomes because it is restricted to the male germ line, and there, most cell divisions occur by meiosis [1]: Sperm is formed in a process of cellular division known as gametogenesis, inside the testis, and it is during this process, where mutations may take place in a way that can affect the future generations (if a Y chromosome in any other cell of the body mutates -i.e. a cell in the liver-, it will have no impact on the offspring of the bearer of the mutation).


Men produce sperm from pubrerty to death, while women are born with a given number of ovum one of which matures monthly from puberty till menopause. This means that sperm are subjected to many more rounds of cell divisions and may accumulate more chance mutations. If the men are older, the chances are even higher.


Unlike the X chromosome, excepting small regions at the telomers (tips), the Y chromosome cannot undergo recombination (where mutated parts are replaced with other "healthier" ones). This means that most of the Y chromosome (95% of it) forms a non-combining region where Single Nucleotide Polymorphisms (SNP) mutations accumulate without being "repaired".


This non-recombining situation arose because X and Y chrmosomes do not recombine among each other (as do the X chromosome pairs in women), to preserve them from gaining harmful genes from the opposite sex. Allowing Y chromosome to preserve male-specific genes.


These mutations allow geneticists to trace lineages and paternity by comparison. They would also allow calculating the age of lineages by comparing differences that accumulated in each line and the rate at which they accumulate. But this is easier said than done.


Calculating ages of Y-chromosome lineages


The key element in dating lineages or haplogroups is to know the mutation rate, and there are basically two methods for calculating it:


I. Direct Measurement or Pedigree estimates: Take two individuals, related by descent and identify the mutations in their Y-chromosome. As the time span that separates them is known (either in years or in generations), the mutation rate can be calculated directly. (it is a value given in: mutations per nucleotide per generation).


The direct measurement (Yali Xue et al., 2009) [1] of the substitutions in the Y chromosome of two related men, separated by 13 generations gave a "mutation-rate measurement of 3.0 × 10-8 mutations/nucleotide/generation... 1.0 × 10-9 mutations/nucleotide/year " [1].


The published human-chimpanzee comparisons are "2.3 × 10-8 – 6.3 × 10-8 mutations/nucleotide/generation... depending on the generation and split times assumed" [1].


The uncertainty is highlighted by the very ample confidence interval values (95% CI) 8.9 × 10-9 – 7.0 × 10-8 mutations/nucleotide/generation obtained.


II. Evolutionary estimates: they use STR polymorphisms or Microsatellites (defined by SNPs). These can be easily genotyped. So taking the microsatellite variation within Y chromosome lineages and knowing the historical dates of certain key events in these lineages history, a mutation rate can be calculated.


As an example, I will folllow the very cited paper by Zhivotovsky et al., (2004) [2], which has a lot of assumptions, plenty of formula and maths. I am an engineer and love maths, but I will spare you the details. Those interested can check the paper (see Statistical Analysis in [2]).


The calculated average "effective mutation rate" (w), was between 0.000312 and 0.000454 per 25 years for Polynesians and Gypsies respectively, however (and these are the things that surprise me!), these values are "adjusted" because they were considered underestimates. The correcting factor ASD0 or average squared difference was applied and voilá, a mutation rate w of 0.000705±0.000332 and 0.000725±0.000187 is obtained for Maori - Cook islanders and Bulgarian Gypsies respectively.


As can be seen the "adjustment" roughly doubled w (it increased 2.25 times in Polynesians and 1.59 times in Gypsies). Furthermore, the error bars are enormous (47% and 26% for each population).


These two values and another one estimated for "global" loci were then averaged resulting in the "magic number" most quoted, cited and used in current genetics papers: "an effective mutation rate at an average Y chromosome short-tandem repeat locus as 6.9×10-4 per 25 years" [2].


Different mutation rates


As we can see the values calculated with each method (Pedigree and Evolutionary) are very different, and applying them to calculate ages of lineages will give very differing results.


Being an engineer with a scientific point of view, I believe that the real values are those that are measured, and that the theory should provide a good model that explains reality and sets of equations or formulae that can be applied with some simple parameters to obtain results that are very similar to reality.


The "laws" of mechanics are used because they are a reasonable model that fit the every day world and gives accurate predictions and practical results (you can design a car or a plane to withstand stress and accelerations, calculate the trajectory of a missile with precision, etc). But when it comes to "laws" in genetics, it seems things are much more blurred and lack precision.


Let's look at possible factors that may explain these differences in mutation rates:

  • Frequent mutations might occur within the few generations used in pedigree studies, while slowly mutating loci only become significant over a longer time interval. [2]
  • Evolutionary calculations use statistics of current variation which include reverse mutation of old alleles as well as forward mutation to new alleles; and these reverse mutation would reduce the number of alleles. On the other hand, Pedigree estimates count mutations on a per-meiosis basis so reversals are counted as new alleles. [2]

I would add that the mechanisms working here are not clearly understood so the model fails to replicate reality.


Factors that distort the estimations


Software and assumptions


Another factor to take into account when calculating the age of different haplotypes are the assumptions behind the calculations.


Modern geneticists employ software that runs simulations (i.e. rho statistics with Network, Bayesian analysis with Batwing), which are fed with these assumptions: weight assigned to different STR variants, exclusion of certain loci (those considered ambiguous or with multi nucleotide repeats), generation time, population sizes, mutation rates (which as seen above are also shrouded in uncertainties), and "others". [3]


Among these "others" are the assumptions that, after populations split, no further migration occurs between them, [3] or, for instance, that there is an exponential growth from an initally population with a constant size "N" [4]. These may not be true, as we will see below, together with other causes


Evolutionary Rate and Repeat unit size


Evolution rate is lower for STRs that have an increased repeat unit size (that is, "n" has more nucleotides).[6][7] In other words, penta or hexa nucleotides mutate slower (3.45 x 10-4 per 25 year generation) than tri or tetra markers (6.9×10-4 per 25 years -the figure given by Zhivotovsky et al., (2004) [2]). [7]


This is because (Dupuya et al., 2004) [8] there are "relatively more gains in short alleles and more losses in long alleles.". [8]


These mutation rates yields different coalescence dates for haplogroups; for instance the age estimate for haplogroup CF clade based on tri/tetra marker results is 42.2 ky which is much lower than 64.7 ky estimated with penta/hexa markers. [7]


The fact that mutation rate depends on allele size means that the different haplogroups (which are characterized by different and specific STRs) will mutate at different rates when compared to each other. [8] Yielding incorrect coalescence dates when compared.


More Factors that influence Y chromosome estimates


When comparing mtDNA timelines (these are based on women) and the male Y chromosome datings, differing patterns appear. These are due to:


1. Genetic drift. It acts strongly upon Y chromosomes: many males don't have sons (they may have only daughters, or die before reproducing) so their Y chromosome is not passed on, and is lost from the gene pool, reducing diversity. [9]


2. Polygyny (having more than one wife at a time) This custom would lead to a small number of males to spread their genes (including their Y chromosome) among a disproportionately large number of children. While others are excluded from the reproductive cycle and their Y chromosomes are lost. [9]


In our recent evolutionary past, humans lived in polygynous, extended families. Where male longevity (>50) would allow them to reproduce up to high ages via younger women, situation which is not found in monogamous societies where menopause effectively cuts off older men's reproductive cycle. Older male sperm may also accumulate more mutations than younger sperm, adding more diversity to the gene pool.


3. Lower effective Male Population Size. The higher Male mortality Rate and the reproductive sucess of males (i.e. due to polygyny) are factors that reduces Y chromosome diversity in populations compared to mtDNA and autosomes. [9] This is seen in the higher level of X chromosome (females) variability compared to that of Y chromosome (males). [10]


In their estimate, Zhivotovsky et al., (2004) [2] consider male and female population as equal, but they are not. And this influences the data on ratio of variance at Y chromosome STRs to that of autosomal STR loci. This ratio varies from 1.14 in "sub-Saharan African hunters" to 0.51 among "American farmers" (the global average is close to 1); and this is due to less males per female in the latter population. This lower ratio leads to a lower mutation rate.


4. Migration. Is an important cause of gene flow within a population. It will lead to overestimation of the accumulated STR variance used in evolutionary calculations.


If migrants admixing with a population are of the same haplogroup they cannot be told apart from the original population, so mutation rates would be overestimated for the admixed population.


The gender mix is also important: if more men migrate than women, this will influence the Y to autosomal STR variance as discussed above. [2] Patrilocality (the residence of a newly married couple with the husband's family or tribe) and Matrilocality (the opposite situation) also alters mtDNA to Y chromosome variance.


5. Generation times. "In present-day hunter-gatherer societies generation time is estimated to be approximately 32 and 26 years for males and females, respectively" [11] which is different to the 25 years postulated by Zhivotovsky et al., (2004) [2]. It may seem trivial but if a generation is 32 years instead of 25, the estimates will vary considerably 10 ky can actually mean 12.8 ky. Historical generation times as calculated by pedigree estimates may be very different from those of our evolutionary past.


My next post "Generation time is not 25 years", gives some sources and data to prove it is at least 30 years for males.


6. Variation in founding populations. The Y-STR variation of the founding population at time of arrival in a geographic region is taken into account in evolutionary estimates [2], if variation is lower, the mutation rate will increase and, for higher variation mutation rate will be lower. So if the founding male population has a substantial diversity it will lead to an incorrect (lower) divergence time calculation. [2]


7. Positive Selection. Natural selection also acts upon men, and will increase frequency of a given lineage if it is more benefical for those carrying it. [9] Or, may I add, it will also benefit Y chromosomes piggybacking on individuals with some other allele favored by selection.


8. Expansion and bottlenecks. Genetic diversity between two populations that shared the same original genetic structure may be due to expansion of one of them: because random mutations will arise more frequently in a larger population simply because there are more sperm cells in which they can arise. This will increase the diversity of the larger population. [9]


A bottleneck will have exactly the opposite effect: a paucity in genetic diversity of the decreasing population as lineages become extinct. [11]


Amerindians


When considering Native Americans we must look back towards their Paleo-Indian ancestors and see how some of the assumptions mentioned above apply to them:


They were not small isolated groups with a closed-shared ancestry. Instead they were dinamic groups that had fluid contacts and exchange between each other and their ancestral populations back in Asia. [12]


They were not a "neutral" system where mutations accrete regularly, they were instead subject to positive selection, war, disease, famine which modified the clock's rate of ticking. [12]


Last but not least is the sampling bias when studying populations. Most are not drawn in a random manner from large populations. Instead they come from tiny samples from small villages where the groups are mostly composed by relatives with shared ancestry. This of course modifies the basic premises of coalescent methods and leads to shorter coalescence times than the actual ones.


Another factor is that the current genes found in a population may not actually represent the historic or even the prehistoric mix of that population [12]. Amerindians suffered a severe bottleneck after the discovery and conquest of America (after 1492 CE) which wiped out many lineages (who knows how many Y chromosome or mtDNA haplogroups disappeared during this period?).


Anzic-1 remains


The remains of a Clovis youth from Montana, US (Anzic-1), which are 12.6 ky old, were typed (Rasmussen et al., 2014) [5] and found to belong to Q-L54*(xM3).


The paper indicates that they then calculated the date of divergence between haplogroups Q-L54*(xM3) (Anzic-1) and Q-M3 of contemporary Native Americans. It is a simple rule of three calculation:


They notice that Anzic-1 had 12 traversions (mutations) while modern ones have on average 48.7, then these 36.7 additional traversions must have arisen during the 12,600 years that elapsed between Anzic-1's death and today: so 12.6 x 48.7 ⁄ 36.7 = divergence date, which happened 16.72 kya.


Of course, to make it statistically neater for the paper, they then "implemented a Poisson process model for mutations on the tree and used the constrOptim() function in R to compute a maximum likelihood TMRCA estimate of 16.9 ky. We then repeated this for 100,000 bootstrap simulations to yield a 95% confidence interval of 13.0–19.7 ky." [5]. The outcome ratified their previous simple calculation.


Below is part B of their Extended Data Figure 2: [5]


Fig 2. Adapted from [5]

The figure's original caption reads: "Each branch is labelled by an index and the number of transversion SNPs assigned to the branch (in brackets). Terminal taxa (individuals) are also labelled by population, ID and haplogroup. Branches 21 and 25 represent the most recent shared ancestry between Anzick-1 and other members of the sample. Branch 19 is considerably shorter than neighbouring branches, which have had an additional ~12,600 years to accumulate mutations."


Cross checked and doubts


I checked this value using the transversions indicated in their figure.


So I took the values in brackets and added them up for each individual, the sum is shown on the far right in green (Q-M3 individuals) and red (Q-L54 ones). At the top is an example of the calculation. The sum is referred to the split that takes place at branch 26 (marked with the vertical green line).


As an example individual at branch 0, MXL NA19682, has 8 + 2 + 8 + 3 + 21 = 42 transversions.


For Q-M3 individuals I calculate an average of 40.33 extra transversions in moderns vs. Anzik-1 and an age of 17.96 ky. Using only the Q-L54 individual's values the average is 44.3 transversions and the age is 17.29 ky, using all modern values the figurs are 41.54 transversions and 17.72 ky. They differ slightly from the 16.9 calculated in [5].


Weird Maths or incorrect assumptions


The odd thing is that when the same methodology is applied to the Saqqaq remains (Branch 27), the age estimation goes awry!:


The paper mentions the Palaeo-Eskimo Saqqaq "sequence had a relatively high missing rate of 0.24 and is divergent with respect to the other hgQ lineages in the sample, its singleton branch should more properly be considered to be of length 71 (54 / 0.76) transversions" [5].


So when we take the age of Saqqaq (4 ky), its transversions from the root at the split of branch 28 (which are 71), and calculate the amount of transversions for modern samples (by adding the 31 that correspond to branch 26, to the previously calculated figures), we obtain an average for all moderns of 72.93 transversions, so the difference that accumulated over 4,000 years is only 1.93 transversions, which leads to: 4.0 x 72.93 ⁄ 1.93 = divergence happened 152 kya! Yes, one hundred and fifty two thousand years ago.


Furthermore the distance in transversions from the baseline (the green line in the figure above) ranges from 26 (on branch 3) to 52 (on branch 12), that is, twice the amount. But all belong to modern human populations, why would one group accumulate twice the quantity of transversions than another? the difference of 26 is 26/36.7 = 70.8% of those accumulated by Anzic-1, and if we apply the same criteria 0.708 x 12,600 y = 8,926 years should separate these populations. But no, they are contemporary. In other words, the amount of transversions does not reflect age as a direct proportion.


This clearly indicates that better calculation methods for Y chromosome lineage dating are necessary.


Subjects for future posts: no Y chromosome from Neanderthals is found in modern humans. Did Q haplogroup originate in America?. Where did the Q hg found in UK and Scandinavia come from?


Sources


[1] Yali Xue et al., (2009). Human Y Chromosome Base-Substitution Mutation Rate Measured by Direct Sequencing in a Deep-Rooting Pedigree. Curr Biol. Sep 15, 2009; 19(17): 1453–1457, doi: 10.1016/j.cub.2009.07.032
[2] Lev A. Zhivotovsky, et al., (2004). The Effective Mutation Rate at Y Chromosome Short Tandem Repeats, with Application to Human Population-Divergence Time. Am J Hum Genet. Jan 2004; 74(1): 50–61. doi: 10.1086/380911
[3] Matthew C. Dulik, et al., (2012). Mitochondrial DNA and Y Chromosome Variation Provides Evidence for a Recent Common Ancestry between Native Americans and Indigenous Altaians. Am J Hum Genet. Mar 9, 2012; 90(3): 573. doi: 10.1016/j.ajhg.2012.02.003
[4] Chuan-Chao Wang and Li Hui, (2014). Comparison of Y-chromosomal lineage dating using either evolutionary or genealogical Y-STR mutation rates. bioRxiv posted online May 3, 2014. doi: http://dx.doi.org/10.1101/004705
[5] Morten Rasmussen, et al., (2014). The genome of a Late Pleistocene human from a Clovis burial site in western Montana. Nature 506, 225–229 (13 February 2014) doi:10.1038/nature13025
[6] Mari Järve, Lev A. Zhivotovsky, et al., (2009). Decreased Rate of Evolution in Y Chromosome STR Loci of Increased Size of the Repeat Unit. PLoS One. 2009; 4(9): e7276. doi: 10.1371/journal.pone.0007276
[7] Järve M, Zhivotovsky LA, Rootsi S, Help H, Rogaev EI, et al. (2009). Decreased Rate of Evolution in Y Chromosome STR Loci of Increased Size of the Repeat Unit. PLoS ONE 4(9): e7276. doi:10.1371/journal.pone.0007276
[8] B. Myhre Dupuya, M. Stenersena, , A.G. Flønesa, T. Egelandb and B. Olaisena, (2004). Y-chromosomal microsatellite mutation rates: differences in mutation rate between and within loci. International Congress Series 1261 (2004) 76 – 78 doi:10.1016/S0531-5131(03)01791-6
[9] Cuan-Chao Wang, Li Jin, Hui Li1, Natural selection on human Y chromosomes. arxiv.org
[10] Michael F. Hammer, Fernando L. Mendez, Murray P. Cox, August E. Woerner, Jeffrey D. Wall, (2008). Sex-Biased Evolutionary Forces Shape Genomic Patterns of Human Diversity. PLoS Genetics doi:10.1371/journal.pgen.1000202
[11] Labuda D, Yotova V, Lefebvre J-F, Moreau C, Utermann G, et al., (2013). X-Linked MTMR8 Diversity and Evolutionary History of Sub-Saharan Populations. PLoS ONE 8(11): e80710. doi:10.1371/journal.pone.0080710
[12] Peter N. Jones, American Indian mtDNA, Y Chromosome genetic data and the peoping of North America, Bauu Institute, 2004.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2014 by Austin Whittall © 
Hits since Sept. 2009:
Copyright © 2009-2025 by Austin Victor Whittall.
Todos los derechos reservados por Austin Whittall para esta edición en idioma español y / o inglés. No se permite la reproducción parcial o total, el almacenamiento, el alquiler, la transmisión o la transformación de este libro, en cualquier forma o por cualquier medio, sea electrónico o mecánico, mediante fotocopias, digitalización u otros métodos, sin el permiso previo y escrito del autor, excepto por un periodista, quien puede tomar cortos pasajes para ser usados en un comentario sobre esta obra para ser publicado en una revista o periódico. Su infracción está penada por las leyes 11.723 y 25.446.

All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted in any form or by any means - electronic, mechanical, photocopy, recording, or any other - except for brief quotations in printed reviews, without prior written permission from the author, except for the inclusion of brief quotations in a review.

Please read our Terms and Conditions and Privacy Policy before accessing this blog.

Terms & Conditions | Privacy Policy

Patagonian Monsters - https://patagoniamonsters.blogspot.com/