Showing posts with label Finnic. Show all posts
Showing posts with label Finnic. Show all posts

Friday, August 10, 2012

fastIBD analysis of East/Central Eurasians and select West Eurasians


Individuals from the following populations have been included in this analysis:
Philippines_D Turkish_D Iranian_D Russian_D Finnish_D Turkish_Cypriot_D Ukrainian_D Belorussian_D Chinese_D Korean_D Japanese_D Tatar_Various_D Kazakh_D Szekler_D Hungarian_D Estonian_D Azeri_D Udmurt_D Mixed_Turkic_D 
These were analyzed in a context of a complete set of Central/East Eurasian populations; West Eurasian populations included were mostly Uralic and Turkic speaking groups, and a few others (such as East Slavs or Iranians).

A few quick points:
  • fastIBD was run with default parameters over a dataset of 627 individuals/255020 SNPs
  • fastIBD identifies segments of relatively recent origin that are shared by individuals. These results should not be construed as measures of overall genetic similarity or origins. Rather, they suggest which populations have exchanged genes in the relative recent past.
With that said, you can get:
  • Spreadsheet of numeric results, showing sharing (in centi-Morgans, cM)
  • Population-level graphical results, showing an ordering of other populations based on mean IBD sharing.
IBD sharing was assessed only for populations with 5+ individuals.

The following heat map allows for a quick appraisal of populations sharing an excess of IBD sharing (read row-by-row)

And, a few visualizations of mean IBD sharing:

Notice high levels of within-population IBD sharing for Finns, consistent with a population that experienced expansion from a small number of founders (small ancestral population size).
Compare with Turks, who are a much more diverse population.
These two plots (you can check the spreadsheet for exact numbers) indicate different sources for the East Eurasian element in Turks and Finns. 

The top eastern populations for Turks are: Turkmen, Chuvash, Uzbek, Uygur, all of which are Turkic speakers, followed by Hazara, Yukagir, and Selkup.  For Finns, there is high degree of sharing with various Siberian groups of different languages, including Uralic Selkups (16.4cM) and Nganassan (9.6cM). Turks share less with these Uralic speakers (6.4 and 2.8cM respectively). So, these are strong hints of common shared ancestry within the Turkic and Uralic language families.

The Chuvash population is also quite interesting, as it shares more with Selkup and Nganassan, contrasting with other Turkic speakers. This makes excellent sense, and is in agreement with other recent findings:
Results from this study maintain that the Chuvash are not related to Altaic or Mongolian populations along their maternal line, thus supporting the “Elite” hypothesis that their language was imposed by a conquering group —leaving Chuvash mtDNA largely of Eurasian origin. Their maternal markers appear to most closely resemble Finno-Ugric speakers rather than Turkic speakers.
Sources of data are listed at the bottom left of this blog.

Saturday, January 8, 2011

ADMIXTURE analysis with Dodecad Populations (update #2)

Thanks to all the participants of the Project, the number of populations has increased, and so have sample sizes within pre-existing populations in the Project. There are now 17 populations with at least 5 individuals in the Project:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian, Russian, Spanish, British, Irish, Turkish, N_Italian, Balkans
Below are the K=10 ADMIXTURE results with these populations:

Admixture proportions can be found in the spreadsheet.

The fact that the addition of 17 populations and 143 individuals to the core set of 36 populations and 692 individuals results in the same 10 ancestral components testifies to the stability of this solution. Hopefully, within 2011 I will develop an even better comparison set to work with.

Another test of the validity of the analysis is comparison of independent samples of the same populations:
Ashkenazi, Armenian, Spanish, Turkish, N_Italian
I have a sample of Dodecad Project members for each of the above, as well as a published population. A way to measure the concordance between the two is to calculate the correlation coefficient (rounded to the 3rd decimal point):
  • Ashkenazi Jews: 0.999
  • Armenians: 0.988
  • Spanish: 0.998
  • Turkish: 0.995
  • N_Italian: 0.996
The concordance is remarkable.

I have also made a RAR of "population portraits". It is important to do this to determine whether minor ancestral components represent population-wide phenomena or are limited to a few individuals.

For example, here are the Turks of the Dodecad project:
The sample is a bit more varied than the sample included in Behar et al:
This probably underscores the importance of broad coverage of large countries and ethnic groups, as I have discovered recently in my analysis of 9 different populations of Pakistan.

Another new population are the Irish, presenting a picture of remarkable homogeneity:
Here is the population portrait for the Balkans, which consists of non-Greek, non-Roma inhabitants of the Balkans:
This appears quite varied; hopefully more Balkan project participants will allow me to split this into additional sample populations.

Finally, here is a portrait of the Ashkenazi population, which appears quite similar to the Behar et al. one:
A very interesting thing about this population is the existence of small slices of "East Asian" and "Northeast Asian" components totalling about 1.5% in almost all individuals. In my opinion this testifies to some type of old minor absorption, as it is fairly evenly spread in the population.

If you haven't joined the project yet, feel free to submit your sample during this opportunity.

Friday, November 19, 2010

ADMIXTURE analysis with Dodecad Populations (update #1)

Repeating the previous analysis with additonal populations of Dodecad Project members and/or modified sample sizes for pre-existing ones:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian
Admixture proportions can be found in the spreadsheet. Dodecad Project populations in italics.

Populations portraits can be found in the RAR. For example, here are the ones for Dodecad Project Ashkenazi and Behar et al. (2010) Ashkenazi Jews:
and here is a Portrait of the Portuguese:

Saturday, November 6, 2010

ADMIXTURE analysis with Dodecad Populations

Thanks to our project members, I am now able to repeat the ADMIXTURE analysis which launched this product, with 5 new populations, composed entirely of Project members. I have made a spreadsheet where you can find admixture proportions for the original 36 populations, as well as the 5 new ones:


The new populations are: Assyrians, Scandinavians, Greeks, Finns, and South Italians/Sicilians. As more people sign up to the project, more populations will pass the 5-person threshold, and will be included in the admixture analysis, and the admixture proportions of existing ones will be fine-tuned. Moreover, some of the populations, such as Scandinavians, may be split into e.g., Swedes, Danes, and Norwegians.

Stay tuned for a notice about future opportunities to submit your data.

Wednesday, November 3, 2010

Analysis of Germanic, Slavic, Finnic, and Baltic Europeans


I have taken 25 HapMap-3 White Americans, 5 Dodecad Project Finns, 25 HGDP-CEPH Russians from Vologda, 12 Dodecad Project continental Germanics (Scandinavians and Germans), 10 Behar et al. (2010) Lithuanians, and 9 Behar et al. (2010) Belorussians and 3 Dodecad Project Northern Slavs, in order to study genetic structure in Northern Europe.

The MDS plot has the following color coding: CEU=blue; Finns=red; HGDP Russians=grey; Dodecad Germanics="green"; Belorussians and Dodecad North Slavs="black"; Lithuanians="magenta".

The grouping of samples by populations is apparent: Germanic speakers (blue+green), Balto-Slavs (magenta+black), Finns (red), with North Russians occupying an intermediate position between Finns and the other Slavs.


ADMIXTURE analysis with K=3 confirms this observation: dark blue corresponds to the Germanic cluster; green to the Finnish one, and light blue to the Balto-Slavic one.

Below are population averages:

CEU are of the Germanic component, with a little Balto-Slavic.

Finns are of the Finnic component, with a little Germanic. Note, however, that this does not mean that they do not share substantial ancestry with Balto-Slavs! The Finnic component itself is composite, comprising of both East and West Eurasian elements, the latter of which could very well represent a common element between Finns and Balto-Slavs.

Vologda Russians seem to be primarily Finnic in origin, but with some distinct Balto-Slavic component.

The Germanic group is primarily in the Germanic component, with some Finnic; without revealing any information, I'll just say that this is contributed primarily by 3 Dodecad Project members who deviate towards Finns and whose ADMIXTURE analysis shows a higher than expected Northeast Asian component. Their outlier status is also visible in the MDS plot.

The Balto-Slavic component reaches its maximum in Lithuanians, while my Slavic sample is primarily Balto-Slavic, but shows both Finnic and Germanic admixture.

The study of structure in northern Europe is important not only for the region itself, but also for the world at large. The northern European component occurs at a lower frequency outside North/Central Europe and knowledge about its structure in its homeland may elucidate its origins in other regions.