Translate

Guide to Patagonia's Monsters & Mysterious beings

I have written a book on this intriguing subject which has just been published.
In this blog I will post excerpts and other interesting texts on this fascinating subject.

Austin Whittall


Showing posts with label admixture graphs. Show all posts
Showing posts with label admixture graphs. Show all posts

Saturday, April 18, 2026

Another paper on Introgressions (April 2026)


Continuing with the wide variety of introgression / admixture papers published over the past few years, today I add a new preprint (in Biorxiv, and therefore not peer-reviewed) published a few days ago, on April 12, 2026: Inferring hominin history with recurrent gene flow from single unphased genomes and a two-locus statistic. Nicholas W Collier, Simon Gravel, Aaron P Ragsdale. bioRxiv 2026.04.11.717825; doi: https://doi.org/10.64898/2026.04.11.717825


Through the use of a very particular statistical model (described at the beginning of the paper, and well over my statistical abilities to understand), and genetic analysis of the autosomal DNA, the authors suggest a population structure and admixture, and population sizes for modern humans, Neanderthals, Denisovans, and super-archaics that mix to and fro over the past million years. The paper assumes "a fixed mutation rate of 1.3 × 10−8 per bp per generation and a generation time of 29 years" (I have previously posted about mutation rate, its variability, and generation times and the combined effect of them on calculating timelines.


The paper reports the following events and dates:


  • Neandertal-Denisovan Common ancestor or ND lived from 779 to 726 kya and lasted for ~50.000 years.
  • Ancestral Neanderthals or AN that around 123 kya split into Altai people in Siberia - who later became extinct, and the Western Neandrthals or WN
  • Anatomically Modern Humans or AMH introgressed into AN 250 kya ago, and 110 kya into the western Neanderthal (WN) group, which later evolved into the Croatian Vindija and Chagyrskaya (Siberia) lineages.
  • Denisovans received gene flow from a ghost lineage, a "Superarchaic" S that may be Homo erectus, it had split from our ancestors 2 million years ago.
  • They reckon that the Ust’Ishim people from East Central Siberia dated to around 45 ky were the first humans in Eurasia to split from the other branches after the Out of Africa Event.

The arrows in the chart show the introgression: "broken one-headed arrows denote instantaneous gene flow events; solid double-headed arrows denote continuous gene flow." The percentages, and population sizes (Ne) are also represented:


Figure 6: Early hominin history in Eurasia with recurrent gene flow. From Nicholas W Collier, Simon Gravel, Aaron P Ragsdale, 2026.

The timeline is the following:


TND→AMH (ky) AMH–ND split time 798 CI: 748 – 827
TAN→Den (ky) AN–Denisova split time 688 CI: 639 – 734
TWN→Alt (ky) WN–Altai split time 123 CI: 117 – 137
TCha→Vin (ky) Chagyrskaya–Vindija split time 60.5 CI: 57.3 – 67.7
TYor→OOA (ky) Yoruba–OOA split time 56.9 CI: 53.6 – 60
TOOA→BE (ky) OOA–BE split time 54.7 CI: 49.7 – 57.6
TLos→Stu (ky) Loschbour→Stuttgart admixture time 29.4 CI: 13.2 – 35.9


The authors conclude that "Using these advances, we inferred a demographic model that broadly explained observed H2 patterns and integrated major supported features in hominin evolution, including recurrent interbreeding between Neanderthals and AMH, introgression from a distantly-related, unsampled lineage to Denisovans, and population structure in western Eurasian AMH."


There is no Denisovan to AMH admixture in this model, it seems to only focus on Northern and Western Eurasians, and does not consider Eastern, Southern or Southeastern Asians and Oceanians.


Effective Populations


I found the Effective population sizes to be of interest (the Ne). As you can see, the Ancient basal root at the top of the image (A) has a large population from which the Superarchaics (S), the modern humans (AMH) split from and conserve a large population size, the ancestor of Neanderthals and Denisovans (ND) has a tiny population and remain that size, so do the N, Denisovans, and the original Out of Africa migration group (bottleneck). The Yoruba people retain a large population.


The paper says that the original Ancestral population had an Ne of 16500 individuals (CI: 15900 – 17300) and then it says "We fixed the effective size of the Superarchaic lineage (S) to 20,000" So the superarchaics splitting from the Ancestral line into Eurasia didn't suffer a bottleneck? Why?


Yet the other groups splitting from the Ancestral group did! The authors explain this large Superarchaic population size as follows: "We justified fixing the population size of S with the observation that changing the effective size of a ghost lineage which makes a small ancestry contribution to a sampled lineage has a negligible effect on E[H2]." So, their model and formulation allows these unrealistic assumptions.


The upper part of the image further down, shows how the Superarchaic introgression into Denisovans affects the population sizes and dates, their model calculates an outcome with a minor impact on effective populations or the timelines.


Interestingly, they note that small effective population sizes may be an artifact, because they can be "plausibly explained by geographic population structure. With spatial structure, recent ancestors are expected to live in closer proximity, and to therefore have a higher probability of sharing parents, than ancient ancestors. Strong structure therefore causes recent coalescence rates to be larger than ancient ones,a pattern which is interpreted as a small recent effective size in a panmictic model."


The paper also notes that "using a lower mutation rate inflated effective size and time parameters, while a higher rate diminished them." They show tables with the effects of different mutation rates as can be seen in the lower part of the image below. The image shows the effects of Superarchaic introgression into Denisovans and the effect of different mutation rates on Ages and Ne of the hominin clades. The original can be seen in tables S8 and S9 in the Supplementary Information of this paper:


hominin population structures

The impact of a slower mutation rate can be seen in the older split between ND and our lineage and an older ND split into Denisovans and Neanderthals, but does not affect on more recent events. The impact of mutation rates on the effective population size (Ne) is also variable, some populations have a bigger effective population (A, AMH, Altai, Vindija, Chagyrskaya, Denisovan) while others smaller (ND, Yoruba).


I have already posted about mutation rates and how it impacts on Ne and heterozygosity, and mentioned the same effect reported by the authors of this paper: For a given heterozygosity, lower mutation rates increase effective population size. I also posted about the effect of mutation rates on dating splits, lower rates lead to deeper (older) split dates.


However, this paper does not explain why the same mutation rate affects Ne in opposite ways (there is a complex explanation about their model in the Appendix that mentions some effects on the Effective Population size). It does admit that the model, like all models is a simplification of reality: "... we made many approximations to simplify our models. We treated populations as discrete entities, with random mating, piecewise-constant sizes, and instantaneous divergence. Some of these assumptions allow us to model the evolution of HR statistics, while others are useful for formally testing tractable demographic models. Of course, the true evolutionary history includes unmodeled populations, continuously fluctuating population sizes, population structure induced by the spatial distribution of individuals, and variable migration rates... We also made a number of simplifying biological assumptions. We assumed that the genome-wide average germline mutation rate was constant across all lineages throughout the modeled period... We also assumed a constant generation time for all lineages throughout the studied period."


Interesting work.

Introgression Index


Visit my index post, with all the introgression posts in one single place.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Monday, February 23, 2026

Templeton's critique (Admixture, Statistical Tools)


American geneticist and statistician Alan R. Tempeleton (b. 1947), at Washington University in St. Louis, U.S., has been an outspoken and vocal critic of the use of Bayesian statitstics to support the Out Of African theory, and the replacement of ancient humans in Asia by our more recent African ancestors.


He holds a master's degree in statistics, so he is an expert in that field, unlike most geneticists (by the way, he also holds a doctorate in human genetics). This places him in the unique position to understand the genetics underlying the black-box algorithms used by researchers. He knows the statistical tools.


Templeton has criticized the incorrect use of Bayesian calculations and the mathematical and formal errors reproduced in different research publications.


His arguments sound solid, and professional. He disputes the missuse of specific statistical tools, as only an expert can. Below are some excerpts from his work.


One paper published in 2023 (Templeton A. The importance of gene flow in human evolution. Hum Popul Genet Genom 2023; 3(3):0005. https://doi.org/10.47248/hpgg2303030005) states the following:


"By the 1980’s CE, the paleontological record had convinced most scientists that the ancestors of humans had first evolved in Africa and then spread out into Eurasia in the early Pleistocene as Homo erectus. However, there was no consensus on what happened next. Three major models emerged by the latter half of the 20th century: the out-of-Africa replacement (OAR) model [1], the candelabra model of racial isolates [2], and the multiregional model [3,4]. Both the OAR and candelabra models posit that the expansion of Homo erectus into Eurasia results in independently evolving populations with no or extremely little genetic interchange. The OAR model in addition assumes a more modern form of humans, Homo sapiens, first evolved in Africa followed by an expansion into Eurasia, where the more modern humans completely replaced the archaic inhabitants of Eurasia. In both of these models, human evolution is dominated by splits into isolated lineages, followed by mostly independent evolution within the isolates. The OAR model in addition posits that the African isolate evolved into a form that expanded into Eurasia where it drove to extinction all the archaic Eurasians without genetic interchange with them. There is no or little role for gene flow in human evolution under these two models: rather, human evolution is dominated by splits, isolation, and extinction of lineages. Weidenreich’s multiregional model takes an opposite position on the importance of these evolutionary forces. There are no splits or isolates in his model because all human populations are interconnected by gene flow.
...
Despite the extensive evidence for gene flow and the lack of evidence of highly isolated evolutionary lineages, much of the human evolutionary literature is still full of “splits’, “divergence times of populations”, and pictures of human evolutionary trees showing separate branches leading to modern day Europeans, Asians, and Africans. These “splits”, “divergence times”, and “trees” are typically estimated with computer programs that will automatically yield a population tree regardless of whether or not the underlying data has a tree-like structure.
"


This is a clear description of the OOA hypothesis, and the expansion of hominins from Africa into Eurasia. Then he mentions the artifacts and models created with computer programs (trees, splits and dats for the forks of the branches of those trees). This is a novel point of view. It shows how we frame our ideas in models that then restrict how our thoughts can evolve in the future. A generation of geneticists thinks in terms of trees, splits, divergences, and coalescence dates.


In 1995, Templeton developed a statistical method called Nested Clade Phylogeographic Analysis or (NCPA) which uses genetic and geographic data to study how a population evolved (more about it in this paper). He mentions this method in his 2023 article, which continues below:


"... aDNA studies have confirmed the most controversial conclusion from NCPA that there was limited genetic interchange between the expanding out-of-Africa population with Eurasian populations. Moreover, genomic studies have revealed much genetic interchange and movement of human populations over the last 100,000 years, as reviewed in. One common method for achieving such inferences is to assume an evolutionary tree for the populations being sampled, and then calculate from the sequence data various statistical tests such as ABBA/BABA or several other alternatives... These statistics are tests of the null hypothesis that the underlying data do indeed come from a tree of populations. Rejection of this null hypothesis indicates that genetic interchange occurred that violated the assumed tree-like structure. When these tests reject a tree-like structure, often an admixture event of genetic interchange is assumed to have occurred to explain the rejection of the null hypothesis of a tree. For example, Figure 4... presents a typical visualization of this type of analysis."


Figure 4 is reproduced below.


Regarding the term "null hypothesis" used by Templeton, it is a statistical term used when one analyzes if the difference between two features of a population are due to chance, sampling errors, or some unknown real variable. To do so, one defines a null hypothesis (in this case, the differences between populations are due to a tree-like structure) and then performs rigorous statistical tests to compare the samples to see if chance or a real effect is responsible for any differences (usually a Student's t-test, created by English statistician William Sealy Gosset (1876–1937), aka Student), one also specifies a probability (the significance level or p), usually 5%, which sets the rejection threshold. If the calculated probability value from the t-test is lower than the significance level, one can reject the null hypothesis and accept that the results are not caused by random chance alone. In Templeton's example the null hypothesis was that tree populations explain human populations' genetics. But statistical analysis rejects this hypothesis meaning that some other genetic factors have intervened that are not tree-like. To save the day, geneticists have concocted admixture events into their trees to explain away the null hypothesis rejection.


Fig. 4 in Templeton
Figure 4. A simplified version of Figure 8 from [51] that shows the estimated gene flow between populations of modern humans and various archaic populations. Gene flow from modern humans into archaic hominins was not estimated.
Source

Tempelton rejects these admixture patches and the concept itself:


"There are two serious problems with this analytical approach. First, these test statistics have an identifiability problem as they cannot distinguish between a single, virtually instantaneous admixture event, versus multiple, recurring admixture events, versus continuous gene flow, or versus gene flow with isolation by distance.
Hence, figures such as Figure 4 are visually misleading as they imply a degree of knowledge that is not truly available from the test results. Drawing a trellis between populations with gene flow would have been equally justified for Figure 4. Other workers have used fossils from different time periods and/or analytical techniques that make use of the length of introgressed genomic segments to infer recurrent and frequent genetic interchange between Neanderthals and modern humans in Eurasia from 100,000 ybp to 37,000 ybp, and perhaps as far back as 270,000 ybp.
Hence, Figure 4 is not only visually misleading, it displays a false narrative for Neanderthals and modern humans. The only conclusion that is justifiable from these ABBA/BABA and similar analyses is the falsification of the null hypothesis that these populations are interrelated through an evolutionary tree. The null hypothesis of human population evolutionary trees has been falsified again and again since the mid-Pleistocene (as reviewed in Chapter 7 in [19]), and this is not surprising as NCPA already indicated that since the mid-Pleistocene human population structure has been dominated by population movements and/or individual dispersal coupled with interbreeding, with no significant role for splits and isolation
"


Then he attacks the statistical logic behind the analysis that leads to these trees, and points out how researchers use computing tools to suit the outcome they are searching for, and don't quite grasp the meaning of the output of these programs:


"The second and more serious reason why figures such as Figure 4 are misleading is that the analytical method of starting with a tree and then adding connections to reflect gene flow is that this approach is statistically inconsistent when the actual relationship of the populations is a more complicated network than a simple tree.
Statistical inconsistency means that the estimators do not converge to the true answer with increasing amounts of data; indeed, the more data you have, the more likely you will have the wrong answer. Patently, inconsistency, just like incoherence, is a highly undesirable statistical property. This inconsistency is illustrated by the work of Pugach et al. They analyzed genomic data from Siberian populations for which some prior demographic historical information was available. Using the TreeMix program that starts with an evolutionary tree of the populations followed by adding on admixture events as needed, they found that the “TreeMix results were not easy to interpret and seem to contradict well-accepted aspects of human population history.” They then analyzed the same data with SpaceMix, a program that does not assume an underlying evolutionary tree. In contrast to TreeMix, the SpaceMix results fit the genomic data better and without contradictions to well-accepted aspects of the population history. SpaceMix indicated a history that included isolation by distance, long-distance dispersals, and multiple admixture events – all of which violate the assumption of a population evolutionary tree. Because of inconsistency, the credibility of figures such as Figure 4 is highly questionable in human evolutionary studies.
"


Templeton then attacks the use of trees and the lack statistical knowledge or interpretation by scholars:


"Despite the almost universal rejection of tree-like structures in human evolution since the mid-Pleistocene, some workers in this area still construct population trees using programs that will always generate a tree no matter how bad the fit is to a tree, treating each human population as an isolate on its own branch of the tree without any indication of any genetic interchange between branches... These population trees are typically presented without any statistical assessment of how well the tree fits the underlying genetic data. I tested the tree given in [63], and rejected a tree-like structure with a p-value < 10-200. To say the least, this is an abysmal fit, and the utility of such poor-fitting trees to gain insight into human evolution and population structure is highly questionable."


I agree with Templeton, and in the past posted on these tools (Some thoughts about the tools used in genetic admixture analysis), remarking that a craftsman is only as good as his tools.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall © 

Thursday, November 20, 2025

Some thoughts about the tools used in genetic admixture analysis


Often, while reading a paper on populations and how different human or archaic groups contributed to their genetics, I wonder about the validity of the computer program tools used by the authors of these papers.


Basically these are computer programs that sift through certain data and apply the border conditions defined by the authors, and then just do the numbers


As an engineeer I am aware of the power and usefulness of advanced statistics. And it is clear that f2, f3, and f4 statistics are great tools in theory, but when they come out of a black box as an output based on some unknown program, can anthropologists, and archaeologists be certain of their validity? Just as I am not trained in their base field, they are not trained in advanced statistics. Mind you, in University I had very tough statistics courses and had to understand both the theory and the practical aspects for using them. I did the numbers, not a program, and had to interpret the data myself.


Today I read an interesting paper: Robert Maier, Pavel Flegontov, Olga Flegontova, Ulaş Işıldak, Piya Changmai, David Reich (2023) On the limits of fitting complex models of population history to f-statistics eLife 12:e85492. https://doi.org/10.7554/eLife.85492. It looks into these questions and reaches some interesting conclusions.


The authors are promoting a new tool, the "findGraphs tool within a software package, ADMIXTOOLS 2, which is a reimplementation of the ADMIXTOOLS software with new features and large performance gains." In the paper, they point out the caveats in current tools. Below are excerpts from this paper:


  • "Conclusions
    Sampling AG
    [admixture graph]space is a useful method for modeling population histories, but finding robust and accurate models can be challenging. As we demonstrated by revisiting a handful of published AGs and re-analyzing the datasets used to fit them, f-statistics are usually insufficient for identifying uniquely fitting AG models, making it necessary to incorporate other sources of evidence. This provides a challenge to previous approaches for automated model building"
  • "A challenge for fitting AG [admixture graph] models is that they are often not uniquely constrained by the data, with many providing equally good fits to the f2-, f3-, and f4-statistics used to constrain them within the limits of statistical resolution."
    This means that the preconceived notions of the authors define which model is the one they like most, and may exclude other models!
  • "As we show in our discussion of case studies, the simple models explored with an exhaustive approach can lead to misleading conclusions about population history because not including additional populations can blind users to additional mixture events that occurred (and whose existence is revealed by examining data from additional populations). Furthermore, models with additional admixture events that are qualitatively different to the best-fitting parsimonious graph and that capture the true history, will sometimes be completely missed when constraining the number of gene flows."
    Again, the choices of the authors of a paper influence the outcome, and their conclusions!
  • For manually built AGs, the sitution is similar, they are based on the intuition, knowledge, and why not, the biases of the builders: "Most AGs in the literature have been constructed manually... often acknowledging the existence of alternative models by presenting plausible models side-by-side, and this approach has been the basis for many claims about population history... A strength of this approach is that it takes advantage of human judgment and outside knowledge about what graphs best fit the history of the human or animal populations being analyzed. This external information is powerful as it can incorporate non-genetic evidence such as geographic plausibility and temporal ordering of populations or linguistic similarity, or other genetic data such as estimates of population split times, or shared Y chromosomes, or rejection of proposed scenarios based on joint analysis of much larger numbers of populations than can reasonably be analyzed within a single AG. Thus, while manual approaches explore many orders of magnitude fewer topologies than automatic approaches often do, they still may provide inferences about population history that are more useful than those provided by automatic approaches. These methods’ strength is also their weakness: by relying on intuition, following a manual approach has the potential to validate the biases users have as to what types of histories are most plausible (these may be the only types of histories that will be carefully explored). This can blind users to surprises: to profoundly different topologies that may correspond more closely to the true history, and we discuss examples of this in the Results section."

Maier et al's paper then looked into several published articles and sifted the data in a different way, with surprising outcomes: "we also identified many additional graphs that fit the data not significantly worse than the published ones. In every example, some of these graphs have topologies that are qualitatively different in important ways from those of the published graphs. Features such as which populations are admixed or unadmixed, direction of gene flow, or the order of split events, if not constrained a priori, are generally not the same between alternative fitting models for the same populations... for all of the publications except one (Shinde et al., 2019), there are alternative equally-well-or-better-fitting graphs..."


In other words they found different admixture graphs with the same statistical stregnth as the ones published in several papers, but with different proportions of admixture, different populations, gene flow directions, etc.


They give many examples and detail the differences they note. I chose one in particular as it mentions a population often used in relation to the peopling of America, Mal'ta:


"Sikora et al. came to the following striking conclusion relying on the "Western" AG (Table 2): the Mal’ta (MA1_ANE) lineage received a gene flow from the Caucasus hunter–gatherer (CHG) lineage. However, in our findGraphs exploration this direction of gene flow (CHG → Mal’ta) was supported by two of the 29 topologies, and the opposite gene flow direction (from the Mal’ta and East European hunter–gatherer lineages to CHG) was supported by the remaining 27 plausible topologies (Figure 4—source data 3). The highest-ranking plausible topology (Figure 4c) has a fit that is not significantly different from that of the simplified published model with six admixture events (p-value = 0.392). We note that the gene flow direction contradicting the graph by Sikora et al. was supported by published qpAdm analyses (Lazaridis et al., 2016; Narasimhan et al., 2019), and qpAdm is not affected by the same model degeneracy issues that are the focus of this study. Considering the topological diversity among models that are temporally plausible, conform to robust findings about relationships between modern and archaic humans, and fit nominally better than the published model, we conclude that the direction of the Mal’ta-CHG gene flow cannot be resolved by AG analysis."


The following graph is an example of the original Sikora AG and the one suggested by the authors as "nominally better fitting."


population admixture graphs
Population admixture graphs, originally published (Left), better fit (Right). Online

These remarks are concerning, the wrong flow of genetic input is serious!


Closing Comments


So my fears about the validity of these graphs is confirmed. The way they are created depends on the ability, biases, preconceptions, and willingness to adhere to previous findings. These then come together and shape the outcome which then conforms to these preconceptions. There is a high risk of Confirmation Bias!


And black boxes that provide answers after you provide an input require human criteria to assess the reasonability of the answer. I remember a professor at University when I was studying Industrial Engineering in the late 1970s. We all had our Casio scientific calculators and could calculate numbers with many decimals, while in the past you needed logarithms or slide rules to run calculations. He said, before you do the numbers you have to have an idea of what the answer is going to be, what use is finding an answer with five decimals if you input the wrong numbers and you are off target by an order of magnitude? (i.e. you calculated 3.567844, yet the answer should be somewhere between 35 and 36). Common sense based on sound theoretical and practical knowledge are paramount.


Tools are only as good as the craftsman that uses them. See a rather basic course for users of these programs on this webpage it tells them which data to input and how to use an F-statistics tool. And this paper on how to interpret the outputs.



Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2025 by Austin Whittall © 
Hits since Sept. 2009:
Copyright © 2009-2025 by Austin Victor Whittall.
Todos los derechos reservados por Austin Whittall para esta edición en idioma español y / o inglés. No se permite la reproducción parcial o total, el almacenamiento, el alquiler, la transmisión o la transformación de este libro, en cualquier forma o por cualquier medio, sea electrónico o mecánico, mediante fotocopias, digitalización u otros métodos, sin el permiso previo y escrito del autor, excepto por un periodista, quien puede tomar cortos pasajes para ser usados en un comentario sobre esta obra para ser publicado en una revista o periódico. Su infracción está penada por las leyes 11.723 y 25.446.

All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted in any form or by any means - electronic, mechanical, photocopy, recording, or any other - except for brief quotations in printed reviews, without prior written permission from the author, except for the inclusion of brief quotations in a review.

Please read our Terms and Conditions and Privacy Policy before accessing this blog.

Terms & Conditions | Privacy Policy

Patagonian Monsters - https://patagoniamonsters.blogspot.com/