In a recent post I discussed a study that questioned the validity of D-statistics; today I will explore two research papers that look into PCA or Principal Components Analysis, a statistical tool widely used in genetics and anthropology, which, according to two papers, may be biased, subjective, and flawed.
What is PCA?
Principal Components Analysis or PCA is a tool used to analyze data. It can be simplified as a method that concentrates, or reduces this data into two dimensions, where the data can be plotted and patterns, or structure identified. Points closer to one another should be similar and associated, those that are more distant in the plot are assumed to be more weakly (or not at all) linked.
You can obtain different information from samples taken from a population, and if there are, say 15 different variables measured for a given sample, it is very complex to pair up the variables and compare the samples to reach a conclusion.
Suppose you are studying bacteria, both pathogens and benign ones, and have defined fifteen variables, collecting data for each of them (length, width, thickness, rugosity, motility, lifespan, etc.). Your goal is to evaluate the influence of these factors on bacterial virulence. It would be impossible unless you could somehow simplify the data without losing information. Imagine if you could concentrate those 15 variables (or 15 dimensions) into just two, it would be much easier to visualize differences or similarities for your analysis.
PCA does just that. Using statistical formulas and computer power it condenses these variables into two main ones, "Principal Components" PC1 and PC2 without losing information, and plots them in a 2-dimensional chart.
Below is a simplified graph showing how three variables, that form a three dimensional space can be compressed into two Principal Components. The points (one sample is equal to one point), are graphed into a three dimensional space, with the three variables, x, y, and z defining their position.
As you can see, the first line (blue) is drawn along the region with most variability in the samples. A second line is drawn perpendicular to it (red). The First Principal Component is PC-1, the other is the Second Principal Component or PC-2. The data is theb "transformed" into a new, 2-D chart, where each point is one of the original samples from the original dataset.
I won't go into statistic, but the transformation is based on evaluating how the variance of the variables of the dataset (how "far" they are from the mean value of the total set), and how they change together (covariance). An example of covariance could be: alchol intake in oz. and driving ability (the higher the intake, the lower the ability to drive).
This analysis results in one main Component PC-1, which reproduces the highest variance, followed by all other perpendicular components ranked by importance. Of these, the second (PC-2) component is chosen.
We should take into account that since PC-1 is more important than PC-2 the differences along the PC-1 axis are greater than those along the PC-2 axis even if the "distance" between points is the same.
Below is a real example from a paper on human population genetics.
The Africans seem to be most distant from all other groups and compact (more similar with each other), while West Asians are widely spread out along PC-1 (showing intra-population differences); Oceanians seem to form another compact group.
The method is faulty
Thhhe first article criticizing this PCA method is Principal Components Analysis fails to recover phylogenetic structure in hominins, by Levi Y. Raskin, et al., bioRxiv doi: https://doi.org/10.1101/2025.10.31.685754 🔓 (it has been published now in the American Journal of Biological Anthropology doi: 10.1002/ajpa.70306; Vol 190, no. 3: e70306. https://doi.org/10.1002/ajpa.70306.🔒). The authors analyzed studies using "morphometric data" (lengths, widths, angles, distances, thicknesses of bones, skulls, specimens) and found that "...PCA trees inferred from traditional morphometric data were identical to the sampled tree in 0.11% of datasets when we only considered PC axes 1 and 2, and in 2.9% of datasets when we considered all axes. No PCA tree inferred from any of the 2,400,000 shape datasets was identical to the sampled tree, regardless of the number of axes."
Only 11 out of 1000 trees produced by PCA methods with 2 dimensions were valid! And using more dimensions, the success ratio was less than three percent!
This questions the validity of conclusions reached using PCA methods: "Phylogenetic interpretations of the hominin fossil record based on proximity in PC space are inherently flawed and likely to be erroneous. Arguments in the hominin systematics literature based on PCA should therefore be reevaluated using phylogenetically-informed alternatives."
The authors are terminant. Those were best case scenarios. In real studies, bias and missing data could make matters even worse. The paper concludes that "PCA is fundamentally ill-suited and unreliable for inferring hominin evolutionary relationships from morphological data. Studies of hominin systematics that incorporate, and especially those that rely on, PCA as a means of inferring morphological and phylogenetic affinity should therefore be revisited."
This is relevant because many PCA charts using morphometric data assume that those hominins who are closer to each other in the chart, are related in the family tree and belong to a similar lineage. But it seems that this may not be the case. Below is an example of a morphometric PCA (Fig. 6 in Source); in the caption the text says "... High-scoring test subjects cluster together, suggesting cranial morphological similarity between them...":
Another Critical Voice: PCA and its use in genetic studies
An earlier paper published Aug. 29, 2022 by Elhaik E. Principal Component Analyses (PCA)-based findings in population genetic studies are highly biased and must be reevaluated (Sci Rep. 2022; 12(1):14683. doi: 10.1038/s41598-022-14395-4. PMID: 36038559; PMCID: PMC9424212), questioned the validity of conclusions obtained by using PCA. Elhaik is highly critical about its applicability to genetic studies and suggests that they are unreliable, and lead to false interpretations.
"...We demonstrate that PCA results can be artifacts of the data and can be easily manipulated to generate desired outcomes. PCA adjustment also yielded unfavorable outcomes in association studies. PCA results may not be reliable, robust, or replicable as the field assumes. Our findings raise concerns about the validity of results reported in the population genetics literature and related fields that place a disproportionate reliance upon PCA outcomes and the insights derived from them. We conclude that PCA may have a biasing role in genetic investigations and that 32,000-216,000 genetic studies should be reevaluated."
This work looks into the population PCAs, like the one shown furhter up (Africans, Europeans, Asians, etc). It notes that the PCA charts place the populations along the corners of an imaginary triangle: Africans, Europeans, and East Asians on each vertex of the triangle. And this is assumed to be the result of the migration out of Africa, followed by separate divergence and mutations. A conclusion widely accepted by geneticists. However, Elhaik says it isn't so: "Here we show that the appearance of continental populations at the corners of a triangle is an artifact of the sampling scheme since variable sample sizes can easily create alternative results as well as alternative “clines”."
This study states that "PCA can be used to generate conflicting and absurd scenarios, all mathematically correct but, obviously, biologically incorrect and cherry-pick the most favorable solution. This is an example of how vital a priori knowledge is to PCA. It is thereby misleading to present one or a handful of PC plots without acknowledging the existence of many other solutions, let alone while not disclosing the proportion of explained variance."
The manipulation implied by the author is that the outcome of the PCA chart is controlled by the scholar who already "knows" what the chart should reflect "By manipulating the choice of populations, sample sizes, and markers, experimenters can create multiple conflicting scenarios with real or imaginary historical interpretations, cherry-pick the one they like, and adopt circular reasoning to argue that PCA results support their explanation".
Closing comments
PCA is a valid statistical tool, with ample use in marketing, segmentation of customers, in facial recognition, in finance, climate modelling, medicine, and engineering (predictive maintenance, quality control). But, it seems that it use in the fields of anthropology and genetics is not robust or reliable.
Patagonian Monsters - Cryptozoology, Myths & legends in Patagonia Copyright 2009-2026 by Austin Whittall ©








No comments:
Post a Comment