Showing posts with label Central Asians. Show all posts
Showing posts with label Central Asians. Show all posts

Wednesday, October 26, 2011

'eurasia7' calculator

This calculator was made with 196 different populations and 2,659 individuals, including 518 project participants. The following Dodecad populations do not have 5 individuals yet, so they are included in the OTHERS_D generic category:
Algerian_D, North_African_Jews_D, Slovenian_D, Mixed_Scandinavian_D, Danish_D, Moroccan_D, Tunisian_D, Serb_D, Austrian_D, Saudi_D, Pakistani_D, Tatar_Various_D, Palestinian_D, Greek_Italian_D, Romanian_D, Swiss_German_D, Szekler_D, Mandaean_D, Azeri_D, Czech_D, Georgian_D, Belgian_D, Latvian_D, Estonian_D, Bangladesh_D, Yemenese_D, Sri_Lanka_D, Hungarian_D, Basque_D, Udmurt_D, Egyptian_D
As always, I encourage people with 4 grandparents from the same country or ethnic group of Eurasia, North or East Africa to contact me (do not send data!) for possible inclusion in the Project. If I have overlooked any such individuals, drop me a line (my e-mail address is at the bottom of the blog). I usually start a new _D population whenever individuals with 4 grandparents from the same group are submitted, but I may have missed some.

Note that all individuals from the reference populations have also been included, including outliers; you should be aware of this when reading the population averages, and consult the Outliers tab in the v3 spreadsheet for some instances of outliers.
Due to image size restrictions in Picasa, the labels are not visible well. A large version of the above plot can be found in the download bundle.

The seven ancestral populations inferred at this level of resolution are:
  • Sub_Saharan
  • West_Asian
  • Atlantic_Baltic
  • East_Asian
  • Southern
  • South_Asian
  • Siberian
As usual, you should take these names as useful labels, and interpret them in conjunction with the components' distribution in different populations, and their Fst distances, both of which can be found in the spreadsheet.

The table of Fst distances:


Below you can see a neighbor-joining tree based on inter-population Fst distances:
The first six dimensions of a multi-dimensional scaling of the same:





Calculator Files:

  • The spreadsheet contains population averages, the table of Fst distances, and individual results for included Project participants.
  • The download RAR file (Google Docs or Sendspace) contains all the files needed to run the calculator. You must download and install DIYDodecad 2.1 first. In order to run the calculator, you follow the instructions of the README file, but type 'eurasia7' instead of 'dv3'.

Terms of use: 'eurasia7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Technical Details:

The calculator is built using allele frequencies of K=7 ancestral components inferred by ADMIXTURE 1.21 analysis of 2,659 individuals. Markers included in the source datasets, as well as the Family Finder and 23andMe (as of Oct 21) platforms were included. The marker set was thinned of markers with less than 99.5% genotype rate and less than 0.5% minor allele frequency. Linkage-disequilibrium based pruning was carried out with a window size of 250 SNPs, advanced by 25 SNPs and R-squared greater than 0.4. A total of 164,990 SNPs remained after these filtering steps.

All relevant populations available to me, and genotyped at a sufficient number of markers were included. Inclusion of the Kalash population resulted in a population-specific component at K=7, and hence their admixture components were inferred a posteriori. Their proportions are consistent with previous results, showing them to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components.

Friday, July 1, 2011

Results up to DOD764 are posted (+portraits, Indo-Iranians etc.)

The results can be found in the spreadsheet

Submission to the Project is currently closed, and of course I encourage participants who have not already done so to leave a message in the ancestry thread.

This completes the results for all Project participants who joined during the latest submission opportunity.

The population averages are finalized -for the time being- but I will occasionally update the _D populations as more participants join the Project and/or I discover cases of fraud in terms of ancestry self-reporting.

Population Portraits

Finally, the population portraits have been uploaded (here and here). For example, here is the Nganassan one, showing three distinctive outliers:

A colorful view of the Nepalese, showing the co-existence of South-Asian-like and East-Asian-like individuals:
Note, that there are also some portraits of populations not included in the averages. For example here are the Onge:
The Onge from the Indian Ocean are outside the area covered by the populations used to create the Dodecad v3, and show mixed "South Asian", "South-East Asian" affiliations. They are probably a good example of case #4.

Indo-Iranian Origins

Here is the population portrait of the Kurds:
I have long noticed that all Indo-Iranian populations possess some of the "South Asian" component. The origin of that component is difficult to ascertain, as it is a composite of "North Indian" and "South Indian" ancestral components, related to West Asians and Onge respectively.

What also seems interesting is that the "South Asian" component is closer to the "West Asian" one with respect to all other West Eurasian components, while many South Asian individuals have substantial levels of the "West Asian" component itself.

The occurrence of "South Asian" in non-negligible levels seems to track the Indo-Iranian world quite well: it is found at about 1/10 in Iranians and Kurds, and also occurs widely in Central Asia, where its true ancient levels were probably much higher due to the substantial presence of east Eurasian elements in the area today. It even occurs at non-trace levels in people who have been part of historical Persian empires such as those from the eastern Caucasus (compare Lezgins and Azerbaijan Jews with Georgians and Adygei, and Iranians/Kurds with Turks, Cypriots, Syrians, and Armenians).

These patterns can be well-explained, I believe, if we accept that Indo-Iranians are partially descended not only from the early Proto-Indo-Europeans of the Near East, but also from a second element that had conceivable "South Asian" affiliations. The most likely candidate for the "second element" is the population of the Bactria Margiana Archaeological Complex (BMAC). The rise and demise of the BMAC fits well with the relative shallowness of the Indo-Iranian language family and its 2nd millennium BC breakup, and has been assigned an Indo-Iranian identity on other grounds by its excavator. As climate change led to the decline and abandonment of BMAC sites, its population must have spread outward: to the Iranian plateau, the steppe, and into South Asia, reinforcing the linguistic differentiation that must have already began over the extensive territory of the complex.

The proposed Indo-Iranian homeland, transitional between the West and the South would explain both:
  • the presence of the "West Asian" component in South Asians (contrast e.g., Kashmiri Pandits with other Indians and south Indian Brahmins with non-Brahmin south Indians), and also
  • the "South Asian" component in Iranians and Iranian-admixed Central Asian Turkic speakers
In their westward march, the Iranians would acquire an excess of West and Southwest Asian components (which would reduce their "South Asian" one), while in their southward march, the Indo-Aryans would acquire an excess of the South Asian component (which would reduce their "West Asian" one).

Thursday, November 18, 2010

How Turkish are Anatolians? revisiting the question

In 2005, I estimated the Y-chromosome heritage of Turkic speakers on modern Anatolians at 11%.

In the same year, I estimated the Mongoloid admixture in Anatolian Turks at 6.2% on the basis of Y-chromosome and mtDNA. This is not inconsistent with the previous percentage, as the Turks, when they arrived in Anatolia were almost certainly of mixed Caucasoid-Mongoloid heritage.

Surprisingly, the maternal contribution from East Eurasia seems higher than the paternal one, on the basis of uniparental markers. But, that is not so surprising if one considers that the Turks who arrived to Anatolia were to a degree descended from Turkicized groups of Iranian steppe nomads bearing Caucasoid patrilineages. Already we have ancient DNA evidence of groups in Central Asia with Caucasoid patrlineages (R1a1) and mixed Caucasoid-Mongoloid mtDNA.

In 2007, some Turkish researchers estimated, using Alu polymorphisms, the Central Asian admixture in Turks at 13%, quite close to my own estimate, and, given the observation that there is more Mongoloid mtDNA than Mongoloid Y-chromosomes in modern Anatolians, the slight difference of 2% is probably taken care of.

In October, I estimated the Mongoloid admixture in Turks at 5.5%, quite close to the 6.2% arrived in 2005 using Y-chromosomes and mtDNA. Subsequently, my K=10 Dodecad analysis (spreadsheet) arrived at 6.7% sum of "East Asian" and "Northeast Asian" components. The slight increase is not surprising, as the K=10 analysis included a greater sampling of Mongoloid diversity.

In ISBA4, another group of Turkish researchers arrived at a 13% estimate for the nomadic Turkic element in modern Anatolian Turks.

Finally, my K=15 analysis has revealed 7.9% "eastern" components in Turks. Given that the "Central Siberian" component is equidistant from Caucasoids and Mongoloids, this translates into about 7.2% East Eurasian admixture. Again, the slightly larger result can be accounted by the sampling of even greater Mongoloid diversity, from the previously unsampled Siberia.

Summary

Y-chromosome, mtDNA, and autosomal DNA analysis by myself and by Turkish researchers all point to 6-7% of Turkish genetic heritage being specifically east Eurasian in origin, and about 1/7 of their genetic heritage coming from Central Asia.

ADMIXTURE analysis

I received an e-mail from a Turkish participant in the project, who wondered whether the K=15 analysis was supportive of much higher demographic influence of Central Asian Turks in the current Turkish population.

In particular, a back-of-the-envelope calculation of "eastern" components in Turks and Uzbeks led him to the conclusion that this was at least 20% and probably more.

Thus, I decided to perform direct ADMIXTURE analysis of Turks and Uzbeks to see what the estimate of Central Asian admixture in Turks actually is.

In the above figure, there are (left-to-right): 1 Dodecad Project 50% Turk-50% Laz showing no Central Asian admixture, 3 Dodecad Project Turks, 19 Turks from Behar et al. (2010), followed by Uzbeks (blue, with some seemingly admixed individuals), followed by 15 Dodecad Project Greeks and Armenians (red).

Behar et al. (2010) Turks have 15.4% Central Asian admixture; if we add the 3 Dodecad Project Turks to the sample, this becomes 14.4%. I'll be happy to tell the three Turks in the Project their individual proportions if they e-mail me.

In conclusion, this analysis too provides an estimate of the Central Asian component in Turks similar to all the ones listed in the beginning.

Conclusion

Estimating the precise genetic identity of nomadic Turks at their time of arrival in Anatolia is difficult to achieve. First of all, modern Anatolian Turks are a subset of recent Anatolians; second, there is the problem of how many Iranian-speakers were absorbed by the westward migrating Turks from Central Asia, and when; also, what was the impact of the Mongol expansion in Central Asia after Turks had already reached the west, and later what were the impacts of Chinese and Russian expansion in the Eurasian heartland.

It's all a big puzzle, but, for the time being 5-7% East Eurasian admixture in modern Anatolian Turks and about 1/7th of their heritage coming from Central Asia seems like a reasonable estimate.