Translate

Guide to Patagonia's Monsters & Mysterious beings

I have written a book on this intriguing subject which has just been published.
In this blog I will post excerpts and other interesting texts on this fascinating subject.

Austin Whittall


Showing posts with label genetic drift. Show all posts
Showing posts with label genetic drift. Show all posts

Thursday, February 26, 2026

Human Genetic Diversity: Some Maps


The global human diversity map that I included in my previous post came without any reference to source or scientific backing, so it isn't reliable. I decided to search for something better and came across the three maps shown below.


They come from Figure 2, in the 2009 paper by Romero, I., Manica, A., Goudet, J. et al. How accurate is the current picture of human genetic variation?. Heredity 102, 120–126 (2009). https://doi.org/10.1038/hdy.2008.89


genetic diversity map
Interpolation of estimates of genetic diversity (HS) for short tandem repeats (STRs) (a), indels (b) and single nucleotide polymorphisms (SNPs) (c). The intensity of the red colour represents the genetic diversity obtained with an inverse distance-weighted (IDW) interpolation on landmasses. Blue dots represent the 54 populations from the H971 subset of the HGDP-CEPH data set. (d) The difference in genetic diversity between African and European populations for the three classes of markers. Error bars report standard deviation.. Fig. 2 in Romero, Manica, Goudet et al.

The map does not include data for Australia, South America appears with the lowest diversity for all three indicators, SNPs, indels, and SNPs, however, Africa has a low scoer in indels as you can see in (b).


I also found the source of the original map, it was published by Luca Pagani, as his thesis (online here, Through the layers of the Ethiopian genome: a survey of human genetic variation based on genome-wide genotyping and re-sequencing data. July 2013, DOI:10.17863/CAM.13969, Thesis for: PhDAdvisor: Toomas Kivisild). The map (shown again, below) is captioned "Pattern of genetic diversity in worldwide human populations. The distribution of STR diversity in worldwide human populations, adapted from the literature (Colonna et al. 2011), shows a higher diversity in African populations and a decline with the distance from Africa (each black dot represents a sampled population). The observed pattern fits with the proposed single African origin with subsequent migrations out of Africa proposed by Stringer and Andrews in 1988."



Colonna et al, (2011), cited by Pagani has several maps in his paper as Fig. 1 (Colonna, V., Pagani, L., Xue, Y. et al. A world in a grain of sand: human history from genetic data. Genome Biol 12, 234 (2011). https://doi.org/10.1186/gb-2011-12-11-234), where Pagani was a co-author.


I wonder what the plesiosaur in the Pacific stands for (?) it appears in the four maps of the 2011 paper.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Friday, February 20, 2026

mtDNA variants and Natural Selection


The chance mutations that are fixed in the DNA of our mitochondria and accumulate there have been used to trace the spread of human beings across the globe. Passed on in a matrilineal form, we all receive the mtDNA from our mother's ovum. Our father's sperm does not carry any mitochondria. Randmo mutations gradually accumulate so they serve as markers in the mtDNA and specific markers define haplogroups.


A paper suggests that these random mutations are then shaped by the forces of Natural Selection. (D. Mishmar, E. Ruiz-Pesini, P. Golik, V. Macaulay, A.G. Clark, S. Hosseini,M. Brandon, K. Easley, E. Chen, M.D. Brown, R.I. Sukernik, A. Olckers, & D.C. Wallace, /2003) Natural selection shaped regional mtDNA variation in humans, Proc. Natl. Acad. Sci. U.S.A. 100 (1) 171-176, https://doi.org/10.1073/pnas.0136972100).


They note that although mutations arise in a random way in the mtDNA, as they have an effect on the mitochondria which produce the body's cells energy and regulate cellular metabolism by producing the energy-rich molecule adenosine triphosphate (ATP), they may be a target of natural selection. They state that "Natural selection shaped regional mtDNA variation in humans."


mitochondria
Mitochondria the body's powerhouse. Copyright © 2026 by Austin Whittall

Molecular clock affected


The fact that mutations are not neutral, and are acted upon by natural selection, implies that the assumptions on which the mtDNA molecular clock are based, are flawed. The paper warns: "If selection has played an important role in the radiation of human mtDNA lineages, then the rate of mtDNA molecular clock may not have been constant throughout human history. If this is the case, then conjectures about the timing of human migrations may need to be reassessed."


The molecular clock based on mtDNA is based on an axiom: genes accumulate new mutations in a clock-like manner, so knowing the rate at which mutations take place (i.e. 3 mutations per 10,000 years), and measuring the average amount of mutations that have appeared since a particular node on a phylogenetic tree (9 mutations), allows us to date the node: 30,000 years. And from there date other nodes based on the number of mutations and the mutation rate.


This is reasonable as long as the mutation rate is constant. But if it varies, then it will provide incorrect dates.


Positive selection could affect the mutation pattern similar and cause an acceleration in the mutation speed. (Further reading on the mtDNA clock: Eva-Liis Loogväi, Toomas Kivisild, Tõnu Margus, Richard Villems (2009))


mtDNA and Selection


After a long stasis in Africa where the L haplogroup is found, humans moved into Eurasia and two branches, or clades, M and N formed outside of Africa and comprise all the mtDNA diversity in the rest of the world. M and N are derived from the African haplogroup L3. And the split is supposed to have taken place around 55-70 kya, during the Out of Africa Event.


Interestingly, M is basically absent in the Middle East, yet it is found in Ethiopia, Southern Arabia and in India and East Asia, suggesting to some a Southern route of migration out of the Horn of Africa across Bab el Mandeb and Hormuz straits. However, a paper published in 2018 by Vicente M Cabrera, Patricia Marrero, Khaled K Abu-Amero, and Jose M Larruga, suggests that both M and N originated in Southeast Asia and migrated westwards. In the case of N haplogroup, it was believed to have formed in the area that links the Levant and Africa and that it appeared in humans taking a northern route out of Africa into Eurasia. But this paper suggests that N originated in Southeast Asia, and moved west across Asia towards Africa. The authors argue that "If one accepts that basal L3 lineages (M, N) evolved independently in southeastern Asia and not in Africa or near the borders of the African continent where the remaining L3 lineages expanded, one is confronted with the question of where the basal trunk of L3 evolved. A gravitating midpoint between eastern Africa and southeastern Asia would situate the origin of L3 in inner Asia."


The paper then states:


"L3 exited from Africa as a pre-L3 lineage that evolved as basal L3 in inner Asia. From there, it expanded, returning to Africa as well as expanding to southeastern Asia, giving rise to the African L3 branches in eastern Africa and the M and N L3 Eurasian branches in southeastern Asia, respectively. This model, which implies an earlier exit of modern humans out of Africa, has been tested against independent results from other disciplines...."


The paper includes the following maps as its Figure 1, and the caption reads: "Geographic origin and dispersion of mtDNA L haplogroups: a Sequential expansion of L haplogroups inside Africa and exit of the L3 precursor to Eurasia. b Return to Africa and expansion to Asia of basal L3 lineages with subsequent differentiation in both continents. The geographic ranges of Neanderthals, Denisovans and Erectus are estimates only."



The paper adds that the "early return and subsequent expansion inside Africa of carriers of L3... haplogroup might help explain, the Neanderthal introgression detected in the western African Yoruba and in northern African Tunisian Berbers." (see my recent post on Neanderthals in Africa).


The authors assume anatomically modern humans left Africa in an early migration 125 kya , met with Neanderthals in south-central Asia, admixed and as the climate worsened ~75kya, the humans moved west and returned to Africa (with the L3 variant with them and it diversified there), and they also moved east reaching SE Asia and China.


Selection and Diversification


Getting back to Mishmar et al., they argue that in Eurasia the M and N lineages spread across the continent in different lineages: A, C, D, and G. Which have a "striking regional variation, traditionally attributed to genetic drift. However, it is not easy to account for the fact that [these lineages] show a 5-fold enrichment from central Asia to Siberia". They argue that this enrichment is the result of natural selection acting as people left their traditional environment (warm, tropical, or temperate climates) and advanced into harsher and colder continental climates in Central and Northern Asia.


The researchers analyzed 104 complete mtDNA sequences from across the world and found that the African haplogroups more or less followed the neutral model, but American, European, Siberian and Asians didn't, they deviated from it. They found that the ATP6 gene, which is a "conserved" mtDNA protein had the highest variation in its amino acid sequences. "Conserved" means that it has remained mostly unchanged over the ages and among individuals and species because it has a low tolerance for mutations, because it is critical for cellular function. So, why would it present so many mutations?


To find out why, they compared the ratios of mutations for the ATP6 gene in different climate zones (arctic, tropical, and temperate) and found that it was highly variable in mtDNAs from the Arctic. Another mtDNA protein called cytochrome b which helps move electrons and create a proton gradient, essential for cellular energy production, was particularly variable in the temperate zones. Another protein, cytochrome oxidase I (or COX1), which also plays a vital role in electron transport, was more variable in the tropical areas. The authors concluded that "selection may have played a role in shaping human regional mtDNA variation and that one of the selective influences was climate."


They then downplay the effects of founder effects arguing as follows:


"..there are striking differences in the nature of the mtDNAs found in different geographic regions. Previously, these marked differences in mtDNA haplogroup distribution were attributed to founder effects, specifically the colonizing of new geographic regions by only a few immigrants that contributed a limited number of mtDNAs.
However, this model is difficult to reconcile with the fact that northeastern Africa harbors all of the African-specific mtDNA lineages as well as the progenitors of the Eurasia radiation, yet only two mtDNA lineages (macrohaplogroups M and N) left northeastern Africa to colonize all of Eurasia and also that there is a striking discontinuity in the frequency of haplogroups A, C, D, and G between central Asia and Siberia, regions that are contiguous over thousands of kilometers.
Rather than Eurasia and Siberia being colonized by a limited number of founders, it seems more likely that environmental factors enriched for certain mtDNA lineages as humans moved to the more northern latitudes.
Natural selection has been hypothesized to explain anomalies in the branch lengths of certain European and African mtDNA lineages.
"


However, a paper by Taku Amu and Martin Brand (2007), disagrees with this concept, and states that there were no differences between the mitochondrial energy management in Arctic or Tropical populations, and that the mutations which were expected to lower coupling efficiency leading to more heat generation in colder climates wasn't detected, and in fact, "Contrary to the predictions of this hypothesis, mitochondria from Arctic haplogroups had similar or even greater coupling efficiency than mitochondria from tropical haplogroups."


More recent research by Jukka Kiiskilä et al (2021) also notes that mtDNA variants are under natural selection and that different mtDNA haplogroups exert a different effect on the physical performance in athletes! the paper looked at Finnish military conscripts and reported that "Following a standard-dose training period, excellence in endurance performance was less frequent among subjects with haplogroups J or K than among subjects with non-JK haplogroups."


Takayuki Nishimura and Shigeki Watanuki (2014) studied mtDNA haplogroup D vs. non-D groups regarding body warmth, and found that "[Non shivering thermogenesis] NST was greater in winter, and that the D group exhibited greater NST than the non-D group during winter...no significant differences in rectal and skin temperatures were found between groups in either season. Therefore, it was supposed that mitochondrial DNA haplogroups had a greater effect on variation in energy expenditure involving NST than they had on insulative responses... individuals from the D group exhibited greater winter values of ΔVO2 than individuals from the non-D group." So, mtDNA haplogroup D subjects had higher oxygen uptake (ΔVO2), meaning their body was "burning" more oxygen but not shivering or increasing the temperature. This suggests an efficient use of energy to heat the core only, and it has a clear mtDNA haplogroup component to it.


Interestingly, Haplogroup D seems to enhance energy burn (without shivering), and without increasing external temperature. From an engineering point of view this is great, since the ΔT or temperature differential between a body and its surroundings impacts directly on the energy loss (Q) the body experiences: Q = U · A · ΔT (where "A" is the area that transfers heat loss, and "U" is a heat transfer coefficient). So this is why the study didn't notice differences in skin or rectal temperatures.


Closing Comments


If random mtDNA mutations somehow provide an adaptative advantage (efficient energy use to keep warm in cold climates), and natural selection acts upon it, then the "neutral" theory is mistaken, and the molecular clock used to calculate dates is also wrong.


Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Sunday, February 15, 2026

Neutral Theory of Genetic Evolution and Out Of Africa


The main backing for the Out of Africa theory is the Genetic Neutrality Theory.


The arguments of an African origin of modern humans and our dispersal across the globe is supported by the high genetic diversity found in modern African populations, with lower diversity elsewhere, and a gradient or cline in diversity that reflects less diversity as distance to the African homeland increases. Both of these factors are expected according to the Neutrality Theory.


Starting with a highly diverse population, if a small group from that population migrates (into Eurasia), it can only feasibly carry with it a sub-sample of the original diversity. This is known as a Founder Effect, the founders of a new population carry fewer genes than the population from which they split from.


This happened time and time again, as sub-sub-groups split from the main population and moved into Europe, Eastern, Northern, and Southern Asia, Melanesia, Australia, Polynesia, and across North America, and then, into South America.


The Neutral Theory states that each split reduces genetic diversity.


Genetic Heterozygosity


Heterozygosity is a measure of diversity. Each person receives genes from their parents, that code for different proteins and produce traits. Those who have two different varriants (alleles) of a specific gene, one inherited from each parent are heterozygous. If the alleles are identical, they are homozygous.


The image below shows two parents (both are heterozygous) each carrying two different variants A and a. The probability for passing them on to the next generation is simple there are four possible combinations, each has a 25% probability of occurring:


heterozygosity and homozygosity
Hetero and Homozygosity. Copyright © 2026 by Austin Whittall

The chances are that two of the offspring will carry Aa alleles, and will therefore be heterozygous, while the other two will receive the same allele from each parent and be either AA or aa, carrying two identical copies. This makes them homozygous.


As we can see, a population that is 100% heterozygous as become 50% homozygous and 50% heterozygous. All the possible combinations of those homozygous and heterozygous genes are shown below:

allele combinations
Combinations of alleles. Copyright © 2026 by Austin Whittall

As you can see 25% of each variant (AA, aa, Aa, and aA). So why would heterozygosity decrease? Suppose only aa homozygous couples mate, the chance of this happening is 1 in 16, or AA mate, again, 1 in 16. So 2:16 or, 1:8 chance of only homozygous mating and offspring. But... if these offspring meet and mate aa with AA, they would have a 100% heterozygous descent. This is true for large populations, but for smaller groups the founder effects and bottlenecks can reduce the allele diversity.


Genetic Bottlenecks


The argument of loss of heterozygosity, or its equivalent, increase in homozygosity is based on genetic bottlenecks, where a small sub-population splits and carries with it the homozygous variant, say only aa or only AA. Losing the possibility of reintroducing the lost allelle. This is a 1 in 16 chance.


Other causes of heterozygosity loss are natural catastrophes, war, and disease. But, why would such events affect the heterozygous individuals more than the homozygous. Wouldn't they be random, and therefore have an equal chance of impacting on hetero- and homozygous individuals?


Regarding the root population. There is the chance that the root from which a population split off from suffered some event that eliminated a large swath of it, while the migrating sub-population in another geographic location was not affected by it. Wouldn't that lower the heterozygosity of the basal group and make the sub-population appear as "enriched"?


Genetic Drift


Both Founder effect and Bottlenecks are part of process called Genetic Drift. As we saw, genetic drift takes place when random events, by chance modify which alleles passed on by parents to their offspring. They also include not only non-reproduction of certain individuals due to war, disease, natural catastrophes, but also loss of genetic variation due to people who don't reproduce because they die before mating, choose not to do so, etc. Genetic Drift isn't driven by evolution. The random changes may or may nor provide adaptations to a changing environment, so they may or not be acted upon by the forces of natural selection.


A sub-population may lose certain alleles, or others may become Fixed reaching a 100% frequency in the population due to chance events.


Mutations and Natural Selection


Random mutations take place, and modify the alleles, natural selection may also work, favoring the survival of individuals with alleles that provide adaptative benefits.


But, what about mutations, that happen by chance, that have a deleterious effect? Some mutations may have harmful consequences. The Neutral theory says that some deleterious mutations may rise to high frequencies in small populations due to fixation promoted by genetic drift. But, why wouldn't people carrying unfavorable genes be affected by natural selection, causing them and their descent to die out?


The Neutral Theory of Molecular Evolution


It was the creation of Motoo Kimura, who in 1968 proposed that at a molecular level, mutations are caused by random genetic drift. These mutations are neutral from a selective point of view. They aren't affected by natural selection.


Kimura has been criticized, for instance Kern and Hahn (2018), argue that modern, genome-scale data demonstrates far more evidence of adaptive evolution than the neutral theory allows, suggesting that natural selection (both positive and negative) shapes much of the genome.


As mutations take place by chance, the probability of them being neutral, deleterious, or beneficial would seem equivalent. So, why assume they are neutral? A beneficial mutation even if it is rare would confer an evolutionary advantage for those carrying it, and modify the population beyond what neutral models suggest.


Linked Selection. The loci (or addresses) that mark the location (locus) of a gene in our DNA isn't independent and isolated. Some genes or DNA sequences located close together on the same chromosome are inherited together, as a unit, during meiosis (linked chromosomes).


Selective Sweep is when an allele that improves the fitness of its carrier increases in frequency due to natural selection, is accompanied (hitchhiking) by other genes linked to it by physical proximity on the DNA strand are also increased in frequency even though they may be neutral. Finally, Background Selection is similar and has the opposite effect: deleterious alleles are removed by natural selection and neighboring neutral alleles are lost too, due to physical proximity to the harmful variants.


These examples show that "neutrality" is not necessarily true.


Molecular Clock


Kimura's theory states that neutral mutations took place at a constant speed, accumulating over time at the same pace. However, this is not true.


However mutations don't appear in a uniform manner in all loci along the genome, they arise unequally, and the probability of fixation depends on where they arise in the genome. This modifies how the clock ticks (Source). Furthermore, substitutions depend on population size, and generation overlap (Source).


Generation time is also an important factor: is it 20 or 30 years? 25? or 18? Over 10,000 generations this means a time scale that can vary from 180,000 to 300,000 years!


Back-Mutations and Recurrence Not Allowed


Kimura's theory, at least when applied in practice, has three axioms that are not true:

  1. Infinite sites, it assumes that each mutation takes place at a site that has never mutated before.
  2. No back-mutations, changes happen in one direction A → G. Which will never again flip back G → A
  3. No Recurrence, in practice there are multiple mutations that take place at the same site. The neutral theory does not accept it, there can't be multiple mutations at identical loci in different lineages.

A paper gives a great example of why and how a back-mutation can have positive effects (here showing how a base C = Cytosine mutates to T = Thymine and back):


"...simple back-mutation is expected to generate slightly advantageous mutations. For example, let us imagine that a site is fixed for C, and that a new T mutation occurs that is slightly deleterious with a disadvantage of −s. Let us imagine that this T mutation spreads through the population and becomes fixed. If a new C mutation then occurs at this site, it will be slightly advantageous with an advantage of +s, unless the relative fitnesses of the C and T alleles have changed. Such a change in fitness could occur because of a change in the environment or the fixation of mutations at other sites which have epistatic interactions with the alleles at a site of interest."


Americas: Great Dying


Regarding Amerindian diversity, we know that up to 90%, or more, of the Native Americans died during the century that followed European "discovery". Disease, war, famine, social disruption, force labor, etc. killed tens of millions of Amerindians. Lineages died out, massively. This is the unique and most massive genocide (albeit unplanned) in the history of humanity. How can we know the number of unique, diverse, divergent alleles that were wiped out during this event? In 1491, America probably presented a far more diverse genetic structure than it does now.

And this brings us to the other point: African "diversity".


African Diversity... is it real?


Finally, and this will be the subject of a future post, do modern Africans reflect the genetic makeup of ancient Africa 100,000 or 75,000 years ago? Is a modern Nigerian, Gambian, Angolan African representative of the ancient population from which the Out of Africa migrants split? Have other events taken place within Africa, isolated from the sub-population that migrated into Eurasia? Admixture with archaic hominins after the OOA event, admixture between many separate and formerly isolated hunter gatherer sub-populations could have led to a modern highly diverse African population, while the original OOA root was far less diverse.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Saturday, May 30, 2015

Unique Amerindian Genetic Trait


My previous post dealt with the anomalous prevalence of Alzheimer's Disease among American Natives, today's deals with another "unique" Amerindian genetic trait, that extends to what in USA are known as Latinos (people with mixed ancestry that includes Native Americans): one that protects against breast cancer.


Breast Cancer rates by race USA
Breast Cancer incidence by Race USA. From [1]

The table above clearly shows how American Natives and Latinos have the lowest incidence of Breast Cancer among American women.


The cause according to a paper [2] by Laura Fejerman et al.,(2014) is a mutation in chromosome 6: "Here we carry out a genome-wide association study of breast cancer in Latinas and identify a genome-wide significant risk variant, located 5′ of the ​Estrogen Receptor 1 gene (​ESR1; 6q25 region). The minor allele for this variant is strongly protective (rs140068132: odds ratio (OR) 0.60, 95% confidence interval (CI) 0.53–0.67, P=9 × 10−18), originates from Indigenous Americans and is uncorrelated with previously reported risk variants at 6q25."


This mutation is the reason that "Latina women, those with a high proportion of Indigenous American ancestry are at a lower risk of developing breast cancer..." [2].


This mutation must have appeared in America otherwise the purported ancestors of Amerindians (as per the Out of Africa theory) would also carry this variant. By the way, the prevalence of Cancer among Amerindians is almost 1/3 of that found among White American women and half of that found among African American women. The Asian Americans' ratio is also almost twice that of American Natives. (these are supposedly the closest genetic relatives to Amerindians).


Is this also due to a bottleneck? or is did it appear during the "Beringian standstill"?


What does the genome of Neanderthal or Denisova tell us about this mutation? I have tried to find information but have not found anything. It may be a mutation inherited from them. Found only in America.


But what about Papuans, who have a high proportion of Denisovan genes? I found two papers (here) and (here) which inform extremely low levels of cance: roughly 8 to 20 times lower than the ratio among Ameridians"!: from 1958 to 1988, the incidence of breast cancer was betwenn 6.9 and 2.4 per 100,000 women.


Do Papuan women have a genetic mutation that protects them too? or is it just lifestyle? Or are these numbers not adjusted by age?


I found another interesting source (global Cancer atlas) which lists cancer prevalence among all human populations. I selected Breast Cancer Incidence and got this map:



Clearly this differs from the other information: dark blue= EU, Australia, America and NZ, Argentina... countries with a high prevalence of White Europeans that eat beef. And low prevalence in "poor" countries where fatty foods are not so common... Asia, Africa, Bolivia. The quality of the data is also variable, ranging from "A" in the US to "C" in China or "G" in Bolivia (19.2 per 100,000 cases) so it makes me wonder how reliable this information is.


Anway, the intersting point is the mutation in Chromosome 6 found among Native American women.

Sources


[1] Zhang and Olopade in Hereditary Breast Cancer. Edited by Caludine Isaacs, T.Rebbeck. pp.234
[2] Laura Fejerman, et al.,(2014). Genome-wide association study of breast cancer in Latinas identifies novel protective variants on 6q25, Nature Communications 5, Article number: 5260 doi:10.1038/ncomms6260



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2014 by Austin Whittall © 

Tuesday, June 24, 2014

Biases in Genetic Models that are generally overlooked


Before continuing with the Y chromosome C haplogroup and its peculiar global range, I want to dedicate today's post to "Models" in general (and those used in genetics in particular), and how they shape the way we believe things work.


As an engineer I have a clear notion that a model is a simplification of the real world, we make assumptions and find that models allow us to make predictions that are reasonably similar to the real world. This allows us to build a bridge that will not collapse and yet use a minimum amount of steel, or shoot a cannon ball from A and hit a target in B. The models can be more complex (relativity is taken into account to make your GPS work correctly and quantum mechanics in all things electronic), but all of them are just that: a "model", an approximation to reality, not reality itself, merely a simplification of the real world.


Models and Genetics a critical overview


Let's review some causes of errors and the hidden biases in genetic models and methods:


Populations


Some studies use strange populations such as "MXL, Mexican Ancestry from Los Angeles USA" or "ACB, African Carribbeans in Barbados", [10] which are far less significant than a native aboriginal individual in his or her homeland: such as an "Alakaluf from Magallanes, Chile" for instance.


What conclusions can be reached on the peopling of America by looking at the genome of a person living in LA, of Mexican ancestry? Mexico has several native groups, later overalaid by Spanish conquistadors (Spanish are themselves a mixture of many ethnic groups... aboriginal Iberians, Basques, Celts, Romans Carthaginians, Arab invaders, Goths, etc.), the African slaves they brought over to work in their plantations and a touch of other European and Southern and Mesoamerican groups. In other words the Mexicans are very admixed population. MXL are actually irrelevant or former slaves of African origin living in Barbados!


Then we have studies that ignore the New World completely. Samplings cover populations in Africa and Eurasia (sometimes including Australia and PNG), seldom the Americas. This of course limits the usefulness of these studies and maybe conceals interesting findings that will remain ignored until American samples are contrasted with Old World ones.


The Sampling within the populations


Once the populations have been identified (with all the caveats mentioned above), its ancestry is studied by means of a sample of individuals that in theory (but not in practice) is drawn from it in a random manner.


In other words a sample of size "n" is taken from a population with a size "N". In most populations "N" is several orders of magnitude lager than "n" (imagine a sample of 200 Italians from Tuscany out of the 61 million people living in Italy). This implies that we may have a sampling bias, and leave out some unique or even critical genetical sequences that appear at low frequencies in a given population.


In other cases the sample is not a random one (Ascertainment bias); for instance, in the case of small tribes: the sample "n" is small because "N" is also small and perhaps encompasses a clan or family group. Thus diversity is low and, as we will see below, calculations based on these samples will be affected by this sampling bias.


Ascertainment bias also arises when researchers take samples from databases in which some populations are missing, or some regions are under represented. Or from samples obtained from volunteers (i.e. a University class) which are definitively non-random samples.


The Typing of Haplogroups in those samples


Once we have sample of our population, then we sequence the Y chromosomes (or in the case of mtDNA, the mitochondrial DNA) for known markers (remember this: known, we will get back to it later) and type the individuals based on these markers. We don't read the whole sequence of 60 million base pairs in a Y chromosome and compare them all. Instead we choose certain ones which we believe are markers and base our analysis on these markers. This introduces another ascertainment bias: the choice of markers is not random.


Look at it this way, we have a million books printed in different languages, we take a sample of fifty from the pile of books and compare them based on a few of the words printed in them: If we find the word "Chapter" we will place them in the English group, "Capítulo" in the Spanish one, "Haupstück" in the German one. All the other words are ignored. We may even find a book printed in Armenian that by chance carries the word "Chapter" and we will place it in the "English" group and ignore that an "Armenian" group even exists. We may have not sampled a Chinese book and also ignore its existence.


Yes, the analogy is faulty (well it is a model after all), the markers we look for in genetics are placed in specific locations within a chromosome. We compare the markers in specific positions with each other to define similarity. In the books example the analogy falls through because the word Chapter can appear on any page and not on a specific page. Yet, what I wanted to point out is that even though there are millions of base pairs in a chromosome, we only use a handful to define haplogroups.

Let's look at this in depth:


Y chromosome SNPs, the Haplogroups


First of all, the Y chromosomes are sequenced and the Single Nucleotide Polymorphism (SNP) that identifies a given haplogroup (hg) is identified. An example of an SNP is shown below:


  • AAGCCTA - ancestral
  • AAGCTTA - derived

The fifth nucleotide, the "C" in the ancestral variant, mutated to a "T" in the derived variant. This SNP ocurrs in a specific location. These mutations are assumed to be apparently rare, and once they happen, they remain in the DNA and are inherited by all descendants of the mutated individual. This allows us to build trees based on which SNPs are present in the populations.


The SNP that identifies the whole Hg C is marker RPS4Y711. The different haplotypes also have their own markers: M8 (for C1), M38 (for C2), M217 (for C3), M93 (for C3a), and so on.


But, as I mentioned in a previous post, it is possible that these SNP mutations revert (in that post I mention the Motala remains sequenced for a Q haplotype which lacked a key marker but had all the others), so this may also introduce additional errors.


Haplotypes and Paragroups


So, in our sample which already may have a sampling bias, we will find that most individuals will belong to a given haplogroup (i.e. Y chromosome hg C) and to a given haplotype (i.e. if they tested positive for marker M38 we will be confident that they belong to the C2 haplotype).


But... Some cases will not test positive for the known markers (i.e. M8 -excluding C1, M217 -excluding C3, M347 -excluding C4 or M356 -excluding C5), so we will assume that they form a paragroup that includes all C-other lineages, and we will identify it as C* (with an asterisk).


So these pargroups clump together a large variety of haplotypes for which we have not yet identified specific "unique" markers which would set them apart as new haplotypes.


Hidden diversity


This is not a trivial matter. For instance paragroup C* is found all across the Eastern edge of the Old World, from Australia, across Indonesia, China, Japan and India, to Bering at frequencies that range from 0.3% to 10% of the population. (some are even higher: C3* among Mongols reaches a frequency of 18% [6]).


To assume that they all belong to the same group is erroneous, the paragroup masks subclades (haplotypes) which have not yet been discovered: Maybe C* in Australia is a not yet identified C8 hg, while C* in Central Asia is a yet undiscovered C9 hg for example. In other words, there is hidden diversity out there waiting to be discovered.


Bias in the choice of the markers


The SNPs are not chosen in a random manner, they are defined by geneticists to type haplogroups. This means that the real distribution of polymorphisms in a population may differ from those shown in a study. The reason is simple: The genotyping arrays (or chips)used to identify the markers contain biased sets of pre-ascertained SNPs. These SNPs tend to be older than the majority of the SNPs in a given population, and are found in many populations besides the one being studied. These hand-picked SNPs act as sieves, classifying the samples and causing alterations such as "[shifts in] allele frequency distributions ... towards intermediate frequency alleles", furthermore, "estimates of linkage disequilibrium are modified" [11].


In other words, this increases the frequency of the most commonly polymorphic loci and eliminates other markers (loci that are less polymorphic in the screening panel). This ascertainment bias in the SNP arrays strongly skews the estimates of genetic diversity by ignoring those that are not included in them.


Variety within the Haplotypes


Once the haplogroup and haplotype have been identified, we can take a look at the Microsatellites to check out even further diversity. Microsatellites are repeats (2 to 6 nucleotides long) that repeat "n" times ( n= 5 to 100). An example: the sequence "AT" repeated 25 times in a row (this is expressed as follows: (AT)25).


These microsatellites are found across species and mutate quicker than single point mutations (SNPs) and for this reason they are used as markers to define the subclades within haplotypes. An example would be a mutation in (AT)25 to (AT)24 or to (AT)26.


In the case of our Y chromosome, we will use special microsatellites known as Short Tandem Repeats (STR). These are named as "DYS-Number" (i.e. DYS393) which indicates a position for the STR. The STR, is a series of repeats of a dinucleotide (two nucleotides).


Below is a real example, for haplogroup C, for four individuals, two from Colombia, one Korean and a Kalkh:


genetic sequence
A real set of STRs for four individuals. Copyright © 2014 by Austin Whittall

We can see in the image above that some DYS markers differ in the quantity of repeats (they are shaded pale blue and yellow).


The Evolutionary sequence


Now come the questions: Does the Korean derive from the Colombians or is it the other way round? What about the Mongolian? Note the differences at DYS392 and DYS393 between Koreans and Colombians (in yellow) and the difference between Colombians and Mongolians at the other DYSs (pale blue).


Without an A priori theory you can't answer that question just by looking at the repeats. Some additional assumptions are necessary, and computer programs are used to build the phylogenetic trees that link these individuals.


We could assume that humans came from Asia and peopled America, so the Americans are more recent than Asians. And as Koreans are Asians, they predate the Colombians so they accumulated mutations, passing from 11 to 12 in DYS392 and 13 to 15 in DYS393. So far so good, but then we have a Khalkh, from Mongolia, who are also Asians, but have less accumulated mutations than the Colombians or the Koreans. This way of comparing STRs is faulty.


Building a Phylogenetic Tree


The individuals are placed on phylogenetic trees using other assumptions that consider the differences between individuals, and are based on the different STRs, but it uses a different reasoning process to the one used above.


Distorting elements


Nevertheless we must remember that each DYS may have its own mutation rate, so if there are several DYS that differ, then they all have to be considered to calculate the "distance" between individuals.


An additional complication is that the "real" evolutionary history of any given set of individuals may differ from the "inferred" evolutionary history. As can be seen in the following image where only three (3) mutations out of twelve (12) are detected during analysis. The other nine (9) are ignored and remain undetected. This affects the estimations on divergence since the mutations are underestimated and the split time between the ancestor and its descendants is underestimated:


missing mutations
How mutations are Underestimated. Adapted from Fig. 2 in [4]

So "corrections" are introduced such as the Jukes-Cantor Model [4] which fiddle with the equations used in the models to make them fit better to reality.


So, how are the differences calculated?


Computer algorithms add even more assumptions


Algorithms are used. They are run on computers and compare the individuals in a pair-wise manner. There are several algorithms, each with their pros and cons. The two basic classes are:


  • Distance-Based Methods - Neighbor Joining (NJ), which initially form an unresolved star-like tree and compare the branch length sum in a pairwise manner. It then groups as "close" those that have the minimum sum. This pair is linked in a branch and then the process begins again, iterating until all individuals have been grouped.
  • Maximum Parsimony method, uses certain features (substitutions in the sequences) to work out a most likely evolutionary relationship among individuals. It builds the tree using the least substitutions chain from the common ancestor to the individuals being located on the tree. So it scores each possible option and minimiizes the mutation number to buid the tree.

The trees are then rooted by comparing them with some outgroup species (ie. chimpanzees are used for human and hominin comparisons). Of course this requires the assumption that molecular clocks are valid and that the divergence date with the outgroup species is well known (more on this below - see clock ticking out of time).


An example: Batwing


Batwing [8] (acronym which stands for Bayesian Analysis of Trees with Internal Node Generation), is a widely used computer program for analysis of genetic data. It has some implicit assumptions that I list below which are not mentioned in the papers that use it, but which impact on the outcome of the program's analysis. By the way, the authors of the program clearly point out that "Natural populations are unlikely to satisfy BATWING's modelling assumptions" [8].


  • The data is a random sampling from the population (we have seen above it is not usually the case)
  • The population is panmitic (not frequent in human groups)
  • Splitting between populations is instantaneous (actually it takes plenty of time)
  • There is no subsequent migration events between populations (there is always posterior admixture due to migrations between populations that have split)

Batwing uses different mutation models, but the default setting is the Stepwise Mutation Model or SSM, which we analyse in detail below.


Comment, the TMRCA (time to most recent common ancestor), Ť, is calculated under the Simple SSM model using the expression: Ť = Δ ⁄ μ.


Where Δ is the average squared difference in the number of repeats between all sampled Y chromosome and the founder haplotype, averaged over STR loci, and μ is the Simple SSM mean mutation rate per generation averaged over loci. But, if, as we will see below SSM is not very reliable, then how can clade age estimates be reliable?.


The Stepwise Mutation Model


This Stepwise Mutation Model (SSM) [2] was proposed in 1973 by Ohta and Kimura and has been widely adopted as the model for microsatellite evolution:


Microsatellites are believed to evolve neutrally: natural selection does not influence the number of repeats so, the SMM premise is: "In one generation the repeat number can only increase or decrease by one, and the probability is equal".


But this assumption is not exactly so for several reasons:

  • Actually, the probability of mutation is larger for longer microsatellites [1][3].
  • A long set of repeats (n larger than 20) may cause physical instability in the microsatellites and hamper its further growth, actually leading to their contraction. [3][2]
  • Some microsatellites are interrupted and have lower mutation rates.
  • The repeat unit also influences mutation rates: dinucleotides mutate slower than tetranucleotides.
  • The motif of the dinucleotide (i.e. TG vs. TA) also plays a role: certain motifs are much longer than others.
  • Variable mutation rates (those that change the repeat by more than 1) are not uncommon and happen about 15 to 22% of the time [3], in other words, the model ignores a big chunk of mutations.

Add to this that insertions or deletions next to the microsatellites also influence their lenghts. [3]


Even the "neutrality" of satellites is questionable since some repeats take place in promoter regions and may influence protein building [3] and thus be subject to natural selection. Some microsatellite repeats have been linked to certain diseases (myotonic dystrophy, Huntingtons' disease, etc.) making their neutrality doubtful too.


There are also "point mutations" that interrupt a repeat; an example: (AT)18 may suffer a chance point mutation where "A" mutates to "G" in position 10, causing the new sequence to be: (AT)9 GT (AT)8. Transormation which may go undetected in sequence analysis, altering the mutation rate estimates.


Panmitic populations


Last but not least, the SMS model assumes that the individuals come from a random sample form a single panmitic population of constant size "N", and this is not the case. [5] Application to expanding populations or those with mixing due to migration may provide different results.


A panmictic population allows random mating without any restrictions of any kind (due to age, genes, behavior, social, environment, etc.), which is seldom the case in human populations, past or present.


Migrations


When an ancestral population splits into two groups, they are subjected to two opposing processes (see image below):


drift and mutation in splitting populations
How genetic drift and migration affect allele frequencies. Copyright © 2014 by Austin Whittall

  • Genetic Drift. It arises because Ne (number of effective breeders) which contribute their genes to the next generation is smaller than the total population, so they pass on their genes only and since this is a random sampling process, the frequency of these genes will differ from that of the previous generation. The smaller Ne, the larger the drift. This effect accumulates with each successive generation and separates the diverging subpopulations as time passes.

  • Migration. Exchanges between the populations as they diverge will limit the drift, keeping them similar. The proportion of migrants "m" if larger will have a higer impact on stability.

Comment on drift and lack of migration: The "Beringian Standstill" was invented to justify the strangely unique American haplogroups, completely absent in the purported Asian homeland of the Native Americans.
The Standstill theory first suggested by Bonatto & Salzano (1997) and perfected by Tamm et al., 2007, is based on a one-in-a-million "Founding effect" that isolates a group of "founding fathers" in Beringia, cut off from their Asian relatives and from the vast empty Americas by ice sheets for about 15,000 years. During this long period of time they mutated their Asian mtDNA and NRY haplogroups into new ones and then in a quick wave covered America swiftly so as not to allow new diversity to arise.
Furthermore, their Asian relatives all died off, leaving no trace on the Asian side of Beringia.
Yes, I know it sounds improbable, yet even though the odds are against this kind of event, several papers apart from Tamm et al, support the theory.


But let's get back to the Stepwise Mutation Model: it is quite weak, to put it mildly.


Summary: SMM is unreliable


All these phenomena make SMM a very rough approximation to reality, yet it is used as if it was 100% reliable!


Just as an example of this lack of reliability is the quote below (Nebel a., et al., 2001) [7]:


"the behaviour of DYS388 appears to be inconsistent with the SMM, as was shown in two populations of Middle Eastern origin. Additionally, another widely used microsatellite, DYS392, has recently been demonstrated to deviate from the SMM" [7]


The Clock that ticks out of time


I have posted on the useless genetic clocks in the past, so I will not bore you, just highlight my previous objections to clocks:


Divergence from Chimps. Scientists devise clocks to calculate mutation rates. To do so we estimate the divergence dates of the human line from the chimpanzee line. But the date of this event is uncertain, and has been increasing since the 1970s from an estimated 5 Mya to 6 - 7 Mya, and in June 2014, to 13 Mya [9], this recent change should surely impact on the dating of human origins!


Assumptions are also made regarding the duration of a generation (what can we know about how long a generation was 50 kya? did females mature earlier or later? what about males? was it 27 years or 35?). Population sizes and their trends (expansion, migration, admixture with other groups as well as bottlenecks and founder effects) also should also be factored in.


When those clocks are calibrated against real mutation rates measured in (again a discrete sample) familes over the last few hundred years strong discrepancies arise: these family (pedigree) calculated mutation rates usuall differ from the former ones (evolutionary). But these differences remain unexplained in the papers. They merely show both figures but avoid explaining the causes (i.e. the mutational clock does not tick at a regular pace).


The mutation rates are also calibrated against the estimated dates for the peopling of certain regions based on the information provided by archaeology. i.e. 40 kya for Australia or 17-20 kya for America. However just by looking at the published error margins we can see the uncertainty involved in these calculations.


Diversity is taken as an indicator of antiquity so if Region X has a large variety of haplotypes while Region Z has fewer, population in Z is assumed to be younger. But, actually what happens is that if the people in "Z" are a subset of population from Region "X", the fact that they are a subset means that they will have less diversity than the original group. This does not mean that they are more recent, it means that they left certain genes behind. Add to this the pressure of natural Selection (and chance i.e. genetic drift) and the genes of certain individuals within the subset at "Z" will get lost too. So if we measure "X" against "Z" by their diversity we would incorrectly judge "Z" to be more recent, when they are really just as ancient as "X".


Frequency, Migrations and antiquity


Often the current distributions of haplogroups occur at differing frequencies in certain territories. What does this mean? That the less frequent "A" hg is a recent arrival of a small group carrying it, entering the territory of the prevailing "B" hg.? Did "A" exist in the same population as "B", but in very low frequencies, and those have been maintained or even decreased?


Or is "A" an ancient colonizer that suffered attrition over thousands of years and has been gradually losing ground to better equipped newcomers with hg. "B"? Maybe "A" and "B" were found in equal proportions in the original colonizers but "B" grew due to genetic drift or natural selection...


Questions like those are seldom asked or answered in mainstream papers. It is clear that the choice of the correct answer requires an in depth analysis which is not found in the academic literature (I have read tens of papers and these matters are not even addressed).


The low diversity among Amerindians is always invariably attributed to a founder effect or a bottleneck during the peopling of America event. The massive death of millions of Natives (virtually a genocide) during the process of discovery and conquest of the New World between 1492 and 1560 is ignored. Disease and war acted selectively wiping out tribes without leaving a trace of them, but this issue is simply ignored and the "lack of diversity" is assumed to be due to the original peopling event some 15 kya.


Dogma

The unidirectional migratory route from Africa to the World has some inconsistencies which can only be explained by back-migrations. These into Africa migrations are reluctantly accepted by orthodoxy but, fortunately, are gradually altering the OoA picture with a more parsimonious explanation. Sometimes I get the feeling that OoA is supported because it is politically correct and assuages the guilt complex of the Western world for the tragic crimes of Slavery and colonialism perpetrated against Africa.


Two issues requiring a serious review are: The East Siberian void of putative ancestors to the Amerindians, which remains unexplained and The "Beringian standstill" justification for Amerindian uniqueness, which also requires a critical analysis due to its improbability.


Closing comments


What I have tried to express in today's post is that there are many assumptions underlying the "facts" expressed by mainstream geneticists regarding human diversity and evolution.


Models are simple representations of reality and not reality itself. They should be taken as such and not as truth written in stone.


Algorithms and simulations run on models are only as reliable as the models they are based on. And we have seen the flaws in some of these models and programs. Flaws that introduce errors in their output yet are not explicitly mentioned in the papers that basethemselves on them.


The complexities in the statistical assumptions mentioned in papers (those pages or paragraphs, full of equations that you skip when reading a paper) mask some very evident biases that skew the results and produce patterns that do not correctly reflect reality, which is richer and much more varied than what these papers show us.


Sources


[1] Esra Ruzgar and Kayhan Erciyes, Phylogenetic Tree Construction for Y-DNA, Haplogroups.
[2]Amke Caliebe et al., (2010). A Markov chain description of the stepwise mutation model : Local and global behaviour of the allele process. Journal of Theoretical Biology 266(2010)336–342
[3] Peter Calabrese and Raazesh Sainudiin, (2004) Models of Microsatellite Evolution
[4] Yan Li, Phycs498BIO Assignment 2, How to Build a Phylogenetic Tree
[5] Valdes, Ana M. Slatkin M. and Freimer N., (1993). Allele Frequencies at Microsatellite Loci: The Stepwise Mutarion Model Revisited. Genetics 133: 737-749 March 1993
[6] Boris Malyarchuk, et al., (2010). Phylogeography of the Y-chromosome haplogroup C in northern Eurasia. Annals of Human Genetics (2010) 00,1–8 doi: 10.1111/j.1469-1809.2010.00601.x
[7] Nebel A., et al.,(2001). Haplogroup-specific deviation from the stepwise mutation model at the microsatellite loci DYS388 and DYS392. Eur J Hum Genet. 2001 Jan;9(1):22-6
[8] Ian Wilson, David Balding and Mike Weale, (2003), Batwing User Guide. See pt. 1.2.
[9] Oliver Venn et al., (2014). Strong male bias drives germline mutation in chimpanzees. Science 13 June 2014: Vol. 344 no. 6189 pp. 1272-1275 DOI: 10.1126/science.344.6189.1272
[10] www.1000genomes.org.
[11]Lachance J, Tishkoff SA. et al., (2013). SNP ascertainment bias in population genetic analyses: why it is important, and how to correct it. Bioessays. 2013 Sep;35(9):780-6. doi: 10.1002/bies.201300014. Epub 2013 Jul 9.
<(p>


Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2014 by Austin Whittall © 
Hits since Sept. 2009:
Copyright © 2009-2025 by Austin Victor Whittall.
Todos los derechos reservados por Austin Whittall para esta edición en idioma español y / o inglés. No se permite la reproducción parcial o total, el almacenamiento, el alquiler, la transmisión o la transformación de este libro, en cualquier forma o por cualquier medio, sea electrónico o mecánico, mediante fotocopias, digitalización u otros métodos, sin el permiso previo y escrito del autor, excepto por un periodista, quien puede tomar cortos pasajes para ser usados en un comentario sobre esta obra para ser publicado en una revista o periódico. Su infracción está penada por las leyes 11.723 y 25.446.

All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted in any form or by any means - electronic, mechanical, photocopy, recording, or any other - except for brief quotations in printed reviews, without prior written permission from the author, except for the inclusion of brief quotations in a review.

Please read our Terms and Conditions and Privacy Policy before accessing this blog.

Terms & Conditions | Privacy Policy

Patagonian Monsters - https://patagoniamonsters.blogspot.com/