Showing posts with label Iranian. Show all posts
Showing posts with label Iranian. Show all posts

Friday, August 10, 2012

fastIBD analysis of East/Central Eurasians and select West Eurasians


Individuals from the following populations have been included in this analysis:
Philippines_D Turkish_D Iranian_D Russian_D Finnish_D Turkish_Cypriot_D Ukrainian_D Belorussian_D Chinese_D Korean_D Japanese_D Tatar_Various_D Kazakh_D Szekler_D Hungarian_D Estonian_D Azeri_D Udmurt_D Mixed_Turkic_D 
These were analyzed in a context of a complete set of Central/East Eurasian populations; West Eurasian populations included were mostly Uralic and Turkic speaking groups, and a few others (such as East Slavs or Iranians).

A few quick points:
  • fastIBD was run with default parameters over a dataset of 627 individuals/255020 SNPs
  • fastIBD identifies segments of relatively recent origin that are shared by individuals. These results should not be construed as measures of overall genetic similarity or origins. Rather, they suggest which populations have exchanged genes in the relative recent past.
With that said, you can get:
  • Spreadsheet of numeric results, showing sharing (in centi-Morgans, cM)
  • Population-level graphical results, showing an ordering of other populations based on mean IBD sharing.
IBD sharing was assessed only for populations with 5+ individuals.

The following heat map allows for a quick appraisal of populations sharing an excess of IBD sharing (read row-by-row)

And, a few visualizations of mean IBD sharing:

Notice high levels of within-population IBD sharing for Finns, consistent with a population that experienced expansion from a small number of founders (small ancestral population size).
Compare with Turks, who are a much more diverse population.
These two plots (you can check the spreadsheet for exact numbers) indicate different sources for the East Eurasian element in Turks and Finns. 

The top eastern populations for Turks are: Turkmen, Chuvash, Uzbek, Uygur, all of which are Turkic speakers, followed by Hazara, Yukagir, and Selkup.  For Finns, there is high degree of sharing with various Siberian groups of different languages, including Uralic Selkups (16.4cM) and Nganassan (9.6cM). Turks share less with these Uralic speakers (6.4 and 2.8cM respectively). So, these are strong hints of common shared ancestry within the Turkic and Uralic language families.

The Chuvash population is also quite interesting, as it shares more with Selkup and Nganassan, contrasting with other Turkic speakers. This makes excellent sense, and is in agreement with other recent findings:
Results from this study maintain that the Chuvash are not related to Altaic or Mongolian populations along their maternal line, thus supporting the “Elite” hypothesis that their language was imposed by a conquering group —leaving Chuvash mtDNA largely of Eurasian origin. Their maternal markers appear to most closely resemble Finno-Ugric speakers rather than Turkic speakers.
Sources of data are listed at the bottom left of this blog.

Wednesday, February 15, 2012

Correspondence between ChromoPainter clusters and ADMIXTURE components in Balkans/West Asia

I took the 25 different inferred clusters from my recent ChromoPainter analysis, and calculated their normalized median components in terms of the K12b calculator. This is a quite useful exercise, since it can show in what sense clusters are different from each other.




Here are two ways in which you may use this correspondence.

1. Different clusters of a single population

For example, the Turks with partial Balkan ancestry tend to belong to pop10, whereas those of Anatolian ancestry to pop13, and those from northeastern Anatolia to pop22. If we compare the admixture proportions of these three groups, we notice e.g.,

  • An excess of Atlantic_Baltic and North_European in pop10
  • An excess of Caucasus in pop22
Or, there is a group of 5 Iranians that belong to pop12, whereas the overwhelming majority of Iranians and Kurds belong to pop21. Strikingly, pop12 differs from all other populations in having substantial levels of East_African and Sub_Saharan. So, it seems that fineSTRUCTURE was able to infer that some Iranian individuals had this feature in common. These individuals were already evident in the Iranian population portrait (right), but fineSTRUCTURE was able to group them even though there were no African populations in the ChromoPainter analysis; presumably, the software was able to detect that these individuals shared a set of chunks that were quite different than is the norm for the Balkan/West Asian area.

2. Related clusters


fineSTRUCTURE grouped the different populations in a tree structure. For example, it grouped pop18, the "North Balkan" cluster with pop23, the "Bulgarian-Romanian" one.

Looking at the admixture proportions, we can tell that the two clusters do indeed seem quite similar, but there are some differences, e.g., an excess of North_European in pop18, and an excess of Caucasus in pop23. This makes sense given the geographical origin of individuals belonging to the two clusters.

Saturday, January 14, 2012

fastIBD analysis of Balkans/West Asia

Now that I've discovered a way to boost Clusters Galore analysis even further by using fastIBD, I will start experimenting with different regional populations. This analysis took about 5 hours to complete, so it appears to be quite practical.

For my first experiment, I carry out an analysis of various populations from the Balkans and West Asia.

Clusters Galore

27 different clusters were inferred with 17 MDS dimensions. Some interesting findings:
  • For the first time there emerge a couple of clusters that appear to be quite specific to Armenians (#2 and #3). 
  • Similarly, Assyrians are broken to a few clusters that appear fairly specific to them  (#9-11)
  • Georgians are split into three clusters, one of which (#14) is linked with the neighboring Abkhasians, who in turn have their own exclusive cluster (#25)
  • The cluster modal in Greeks (#6) includes 14 of 19 Greek participants, and a few Greeks are also in the Balkan cluster (#8) and an Iranian-Turkish cluster (#4)
  • The Behar Cypriot sample also splits into two, and the few Turkish Cypriot participants link to one of them (#13)
  • The Ossetian project participant links to one of the three North_Ossetian clusters
  • The major Balkan cluster (#8) still defies resolution. I am certain, however, that structure in this cluster will be uncovered with more participation. MCLUST adapts the cluster size and shape, and a "big", inclusive cluster spanning the Balkans appears more parsimonious than smaller clusters centered on the different groups. With larger participation, I anticipate that regional structure will be uncovered in the Balkans as well.
I cannot stress the importance of participation strongly enough. When groups have more participants, it is possible to both:

  1. Discover group-specific clusters, by identifying what is common between members of groups
  2. Discover within-group clusters, by identifying what is different between members of groups
For example, the great participation of Armenians in the Project has now allowed me to discover structure within the Armenian population. It appears, that cluster #2 corresponds to a more "western" Armenian group, and #3 to a more "eastern" one, with some overlap between the two.

Inter-population IBD


You can also see a visual representation of inter-population IBD:

I have only included populations with 5+ participants in this representation. Reddish shades express high IBD sharing; bluish ones low one. The heatmap has been scaled by row.

As you might expect, values across the diagonal are "reddish", since individuals within populations tend to have high IBD sharing with each other.

A few features "pop out" of the screen. Going from top to bottom:
  • Intra-Iranic sharing
  • Intra-Armenian sharing
  • Intra-Balkan sharing
  • Georgian-Abkhaz sharing
You can probably get more out of the figure, but these appear to be the most salient features.

Results for Project Participants


The results can be found in the spreadsheet, and include:
  • Probabilities of assignment in each of the 27 clusters of the Clusters Galore analysis
  • Z-scores of IBD between each individual and each of the 20 populations with 5+ participants. Higher values mean more IBD sharing. Note that Z-scores have been calculated for each row, hence each participant must scan his own row to find populations with an excess (+) or deficiency (-) of IBD sharing, and people should not compare across different rows.
Last but not least, I want to remind new project participants to leave a message in the Information about Project samples thread. Your comment will not appear immediately, since comment moderation is on, and also note that there are multiple pages of comments. 


If you haven't joined the Project yet, I encourage you to do so if you are eligible.

Sunday, September 25, 2011

Yunusbayev et al. (2011) data assessed with Dodecad v3

I have acquired the data from the recent Yunusbayev et al. (2011) paper on the Caucasus. This includes the following populations:
  • Kurds_Y 6
  • Bulgarians_Y 13
  • Ukranians_Y 20
  • Mordovians_Y 15
  • Armenians_Y 16
  • Abhkasians_Y 20
  • Balkars_Y 19
  • North_Ossetians_Y 15
  • Chechens_Y 20
  • Nogais_Y 16
  • Kumyks_Y 14
  • Turkmens_Y 15
  • Tajiks_Y 15
It is a valuable new addition to the Project, and it is commendable that it has been made publicly and easily available so swiftly after the appearance of the Yunusbayev et al. (2011) paper.

To get the ball rolling on the new Yunusbayev et al. data, I will map the new populations onto the Dodecad v3 components; they will be added to the Dodecad v3 spreadsheet as they are calculated.

I have been laboriously designing a new global (including Amerindians and Australasians) Dodecad X1 experimental calculator with 3,010 individuals for a few weeks now, but I guess I will now have to reboot it with 3,214.

Together with some other new data I recently discovered, I now have 9,799 individuals (some duplicates from different sources) in my global database. My Dodecad dataset of 511 individuals from a single country or ethnic group isn't too shabby either. Let's hope for a new data release that will push the data collection above the magic 10,000.

UDDATE:

I have added the first 7 populations to the spreadsheet; the others are being calculated as we speak. Most of them seem in line with expectations, but the Abkhasian sample has one outlier individual (abh27), and has thus been placed in the "Outliers" tab of the spreadsheet; a new set of admixture proportions, minus that outlier individual, will be calculated anew:

UPDATE II: The population portraits have been uploaded to Google Docs as a rar file (Sendspace mirror). Average admixture results have all been entered to the spreadsheet.

Friday, July 1, 2011

Results up to DOD764 are posted (+portraits, Indo-Iranians etc.)

The results can be found in the spreadsheet

Submission to the Project is currently closed, and of course I encourage participants who have not already done so to leave a message in the ancestry thread.

This completes the results for all Project participants who joined during the latest submission opportunity.

The population averages are finalized -for the time being- but I will occasionally update the _D populations as more participants join the Project and/or I discover cases of fraud in terms of ancestry self-reporting.

Population Portraits

Finally, the population portraits have been uploaded (here and here). For example, here is the Nganassan one, showing three distinctive outliers:

A colorful view of the Nepalese, showing the co-existence of South-Asian-like and East-Asian-like individuals:
Note, that there are also some portraits of populations not included in the averages. For example here are the Onge:
The Onge from the Indian Ocean are outside the area covered by the populations used to create the Dodecad v3, and show mixed "South Asian", "South-East Asian" affiliations. They are probably a good example of case #4.

Indo-Iranian Origins

Here is the population portrait of the Kurds:
I have long noticed that all Indo-Iranian populations possess some of the "South Asian" component. The origin of that component is difficult to ascertain, as it is a composite of "North Indian" and "South Indian" ancestral components, related to West Asians and Onge respectively.

What also seems interesting is that the "South Asian" component is closer to the "West Asian" one with respect to all other West Eurasian components, while many South Asian individuals have substantial levels of the "West Asian" component itself.

The occurrence of "South Asian" in non-negligible levels seems to track the Indo-Iranian world quite well: it is found at about 1/10 in Iranians and Kurds, and also occurs widely in Central Asia, where its true ancient levels were probably much higher due to the substantial presence of east Eurasian elements in the area today. It even occurs at non-trace levels in people who have been part of historical Persian empires such as those from the eastern Caucasus (compare Lezgins and Azerbaijan Jews with Georgians and Adygei, and Iranians/Kurds with Turks, Cypriots, Syrians, and Armenians).

These patterns can be well-explained, I believe, if we accept that Indo-Iranians are partially descended not only from the early Proto-Indo-Europeans of the Near East, but also from a second element that had conceivable "South Asian" affiliations. The most likely candidate for the "second element" is the population of the Bactria Margiana Archaeological Complex (BMAC). The rise and demise of the BMAC fits well with the relative shallowness of the Indo-Iranian language family and its 2nd millennium BC breakup, and has been assigned an Indo-Iranian identity on other grounds by its excavator. As climate change led to the decline and abandonment of BMAC sites, its population must have spread outward: to the Iranian plateau, the steppe, and into South Asia, reinforcing the linguistic differentiation that must have already began over the extensive territory of the complex.

The proposed Indo-Iranian homeland, transitional between the West and the South would explain both:
  • the presence of the "West Asian" component in South Asians (contrast e.g., Kashmiri Pandits with other Indians and south Indian Brahmins with non-Brahmin south Indians), and also
  • the "South Asian" component in Iranians and Iranian-admixed Central Asian Turkic speakers
In their westward march, the Iranians would acquire an excess of West and Southwest Asian components (which would reduce their "South Asian" one), while in their southward march, the Indo-Aryans would acquire an excess of the South Asian component (which would reduce their "West Asian" one).

Monday, May 30, 2011

More Zombies: Ancestral North Indians and Ancestral South Indians reborn

In my previous post I showed how synthetic individuals corresponding to ADMIXTURE ancestral components can be created and used. This was made possible by the fact that ADMIXTURE outputs allele frequencies for its components, which can be utilized to create a population of random genotypes with the same allele frequencies.

A more difficult task is to create such "zombie" individuals when there are no allele frequencies at hand. A prime example of this is the paper by Reich et al. (2009) on the two ancestral components in Indians: Ancestral South Indians (ASI) and Ancestral North Indians (ANI). The paper provides admixture estimates for these two components in present-day "Indian Cline" groups, but no allele frequencies for these components: we only knew that ANI was closely related to West Eurasians, and ASI formed a clade with the Onge from the Indian Ocean.

Both ANI and ASI are extinct (in pure form) populations, and they are blended (in varying proportions) in modern day Indians, with highest ANI occurring in the Northwest and among upper caste groups, and highest ASI among South Indian tribal and low caste populations.

As I was thinking of ways to extend the "zombie" approach, it occurred to me that there is a fairly involved way to extract the ANI/ASI allele frequencies from the available evidence:

If f(ANI) and f(ASI) are the allele frequencies at a locus for ANI and ASI, and an admixed population P has x fraction of ancestry from ANI and 1-x from ASI, then its allele frequency is expected to be:

x*f(ANI)+(1-x)*f(ASI) = f(P)

I have marked (in bold) the known variables. Obviously, this equation does not hold in practice, because of sampling error, uncertainty in the estimation of x, as well as genetic drift that may affect the allele frequencies of the admixed population.

Nonetheless, we do not only have one equation of this sort, but 18, since Reich et al. (2009) provides ANI/ASI estimates for 18 different Indian Cline populations. We can thus fit a linear regression to recover f(ANI) and f(ASI).

This is exactly what I did; there are two important caveats:
  • because most of the Reich et al. (2009) populations are very small, f(P) is expected to be very noisy. I thus grouped the Indian Cline populations into five groups (based on increasing ANI, and making sure that each one had >15 individuals), and calculated admixture proportions (x's) and allele frequencies (f(P)'s) on these groups.
  • linear regression coefficients (the f(ANI) and f(ASI) estimates) may be less than 0 or more than 1, which makes no biological sense, so these were fixed to 0 and 1 in a few cases whenever that was the case (~5% of markers)
All of this required a bit of thinking and work, so I was very skeptical that it would work; given sampling/admixture estimation errors/limitations of regression/random creation of individuals, the whole process from input data to output "zombies" passed through so many layers, that it could very well lead to nonsense.

Nonetheless, there is power in numbers, and I was hopeful that this might work. If it did, I could have synthesized ANI and ASI populations to play with and use pretty much like regular populations in a variety of experiments.

Validation of synthetic ANI/ASI populations

I generated 25 ANI and 25 ASI individuals using the above-described method. There are 119,588 SNPs in these populations.

To validate them, I ran supervised ADMIXTURE using these ANI/ASI individuals as ancestral populations, and all the Indian Cline populations as test data. The results can be seen below:
Although the estimates for some populations (e.g., Chenchu: 31 vs. 40.7%) are substantially off, the median error is 1%, and the average error is 2.4%. Overall, it does appear that the synthetic ANI/ASI individuals are fairly good standins for their (extinct) populations.

Ancestral North Indians

I included ANI together with the 4 West Eurasian components of the Dodecad Project in an MDS plot:
Also, a neighbor-joining tree:
Putting ANI/ASI to work: Romanian Gypsies

I have previously detected 2 individuals in the Behar et al. (2010) Romanian sample that are likely to be of Roma (Gypsy) heritage. Here is a supervised admixture of the Romanian sample using the ANI/ASI components:
The previously detected individuals do possess both ANI and ASI components, indeed these are:

18.1, 15.3
16.9, 16.4

in the two individuals, which might be useful in constraining geographically the origin of European Gypsies along the Indian Cline.

Putting ANI/ASI to work: Iranians

Iranians generally show affinity to South Asians. Is this affinity related to the common Indo-Iranian background of Iranians and Indo-Aryans, or, is it, perhaps, due to the absorption of South Asian population elements during Iran's long imperial past?

The ANI/ASI components in the Iranians and Iranian_D samples are:

11.7, 7.5
12.0, 6.9

Compared to the previously described Romanian Gypsies, the South Asian component in Iranians tends to be clearly tilted towards ANI.

Friday, April 8, 2011

Structure in West Asian Indo-European groups (part 2)

I will occasionally revisit old posts such as Structure in West Asian Indo-European groups to take advantage of new population samples from project submitters. This time around, I included our first Kurdish and Azeri participants, and limited the analysis to populations for which I had large numbers of markers (the final total is a ~132k pruned set of markers).

I also included Greeks, Caucasian populations (Georgians and Lezgins), and Levantine Arabs (Syrians-Lebanese) who frame this region from the West, North, and South respectively, as well as Assyrians who are interspersed in West Asia as an ethno-religious minority.

Here are dimensions 1 and 2 of the multidimensional scaling plot:

The two Caucasian groups (South and Northeast Caucasian Georgians and Lezgins respectively) form their own clusters. So do the Iranians and the Syro-Lebanese.

(As always population labels are placed on population averages, and _D denoted Dodecad Project populations)

Curiously, many linguists assert a close relationship of Greek and Armenian within the Indo-European language family. Turks speak an Altaic language due to migration of a numerically small population element, but their pre-Turkish genetic ancestors were Anatolian speakers, Greeks, Armenians, and Iranians, i.e., primarily Indo-Europeans.

Dimensions 1 and 3:

Dimension 3 contrasts Greeks from West Asian groups. Notice also the presence of 3 Armenians at the bottom of the plot, these are outliers of the Behar et al. Armenian sample.

Notice that the Azeri_D sample clustered with Iranians in dimension 2 and with Turks in dimension 3. This is not very surprising, as Azeris speak a Turkic language, but also have clear Iranian antescendants. The Kurd_D sample, on the other hand, clusters consistently with Iranians.

The variability of the Greek_D population sample along dimension 3 is also quite interesting. This could reflect variable levels of influence of extra-Greek European/Anatolian population elements on the basic Greek stock. Greeks who possess 23andMe or FTDNA Illumina population data are strongly encouraged to join the Project to help us better determine regional variation within the ethnic Greek population.

The Clusters Galore analysis results are as follows (9 clusters/3 MDS dimensions retained) :

In brief, the modal populations in each cluster are:
  1. Greeks
  2. Iranians
  3. Turks
  4. Syrians and Lebanese
  5. Armenians and Georgians
  6. Armenians and Assyrians
  7. Syrians and Iranians
  8. Lezgins
  9. Georgians
I will be happy to provide to all Dodecad Project members from the _D populations with their individual results. If you send me e-mail at dodecad@gmail.com, I will send you a line with your probabilities of assignment in each of the 9 clusters, as well as the 3 co-ordinates in the first 3 MDS dimensions plotted above.

I strongly encourage individuals from West Asia, the Balkans, and Italy to contact me for inclusion in the Project (send e-mail first, not data). While submission to the Project is currently closed, I usually accept data from these regions.

Thursday, December 30, 2010

Structure in West Asian Indo-European groups

Indo-European languages used to be represented by several branches in West Asia. The most ancient Anatolian languages (Hittite, Luwian, and Palaic) became extinct by ancient times, and its descendants (e.g. Lycian and Lydian) soon thereafter. Greek and Phrygian were added from the Balkans, and so did Armenian, which, according to tradition was derived from the Phrygian.

Unrelated to all these languages, to the east, were the Iranian speakers, the major branches of which extant today in West Asia are Kurdish and Persian (Farsi).

The boundary between the western Indo-Europeans and the east ones changed several times in history. The apex of Iranian power came early when the Achaemenids subjugated the entirety of Asia Minor, but the Iranian expansion was halted in Europe during the Persian Wars of the 5th c. BC. In the next century a reversal of fortunes resulted in the conquest of the Persian Empire by the Greeks of Alexander. The Hellenistic kingdoms lasted for centuries thereafter, but eventually succumbed and new Greco-Persian kingdoms arose that led to the Parthian revival which clashed with the new western power, Rome. The limit between East and West survived with victories and defeats on either side well into medieval times as the eastern Romans fought the Sassanids, the successor dominant power in the Iranian world. Heraclius ended the centuries-old struggle when he defeated the Sassanids in Mesopotamia, but the whole affair became irrelevant as soon thereafter the new power of the Arabs destroyed the Sassanids and threatened the Roman Empire itself which managed to survive for a few more centuries, before eventually succumbing to the Crusaders and Turks, but both of these probably had a minor effect on the local population. Greeks and Armenians continued to exist in Anatolia until the 20th century, with only remnants of them remaining in Asia Minor today.

The data

To study the relationship between the various West Asian Indo-European groups, I gathered an Iranian sample (from Behar et al.), an Iraqi Kurdish one (from Xing et al.), an Armenian one (from Behar et al.), as well as an Armenian one from the Dodecad Project. I have also included the Behar et al. Turkish sample, and a new Turkish sample from the Dodecad Project.

Below are the first two dimensions of the MDS plot.

It appears that Kurds are not particularly closely related to their linguistic cousins, the Iranians. Neither are they very close to Turks and Armenians; the latter appear very close in most of my analyses, with the main difference being a small Mongoloid component in the former, which is not visible here due to the lack of Mongoloid reference populations.

The distinctiveness of the Kurds is also evident in the ADMIXTURE analysis:

The high blue component distinguishes Kurds from both Iranians and Armenians/Turks. Iranians have slightly more of it, suggesting a somewhat closer relationship. The difference between the small Dodecad Turkish sample and the Behar et al. one is suggestive of heterogeneity within Turks, so it is important to be aware of this. Hopefully, if more West Asian individuals join the project, we will be able to discover patterns of regional variation between and within different ethnic groups.

Sunday, December 19, 2010

Fine-scale admixture in Europe (Dagestan/Basque/Sardinian components)

Wanting to see whether the Dagestan mystery would extend into Europe, I carried out an ADMIXTURE analysis including all my European populations. Once again, as this is done on only ~30k markers, a little noise on the low-level components is expected.

Admixture proportions can be found in the spreadsheet

Notably there is now both a Sardinian and a Basque centered cluster; the latter was formerly (in the standard K=10 analysis) split between "Southern European" and "Northern European". The Urkarah, Lezgin, and Stalskoe samples show the highest presence of the "blue" component, which I label, once again, Dagestan. Note, however, that you should not compare admixture proportions across ADMIXTURE runs for components that happen to be labeled the same (homonymous). Certainly "this" Dagestan is related to the "previous" Dagestan component, but do not assume they are identical.

Here is the Fst distance matrix between the 7 components:


Discussion

The most notable thing about this figure is the relative absence of the West Asian component in the periphery of Europe. The lowest values are seen in Basque, Sardinian, Orcadian, White Utahns, Lithuanians, Finns, and Scandinavians (in no particular order).

It is worthwhile to order the European populations in terms of their Dagestan component. Excluding the populations of the Caucasus, these are, in ascending order: Basque (0.7%), Sardinian, Cypriot, Belorussian, South Italian/Sicilian, Lithuanian, Tuscans, Portuguese, Greek (3.8%), Vologda Russian, Romanian, Finnish, Spaniards, North Italian, Dodecad Spaniards, Dodecad Russian, Chuvash, Hungarian, French (7.9%), German, Scandinavian, White Utahn, Orcadian (12.6%).

Interpreting this pattern is not easy, but it does seem that this component seems to have a V-like distribution, achieving its maximum in Caucasus and its environs, then undergoing a diminution, and achieving a secondary (lower) frequency mode in NW Europe.

The surprising appearance of the homonymous Dagestan component in India suggests a widespread presence of a common ancestry element. The West Asian element, by comparison seems to have a more normal /\-like distribution around its center in Anatolia-Caucasus-Iran region. It does reach the Atlantic coast, but is lacking in Scandinavia and Finland, and also in India itself.

This is just a piece of a broader puzzle, and the picture is not yet clear. However, we can tentatively say that whatever brought the "Dagestan" component to India was not a unidirectional process, but also brought a similar population element to western Europe.