Showing posts with label Germanic. Show all posts
Showing posts with label Germanic. Show all posts

Saturday, January 21, 2012

fastIBD analysis of Central/Eastern Europe

Please refer to the previous analysis on the Balkans/West Asia for more information about the interpretation of this type of analysis.

Clusters Galore


The Clusters Galore can be found in the spreadsheet. After inspection of the 23 clusters inferred with 21 dimensions, they could be described as:

  1. Mordvin
  2. East Slavic
  3. Polish-Ukrainian
  4. East Balkan
  5. Vologda Russians
  6. Lithuanian
  7. Central European (combining many groups with small sample sizes)
  8. A couple of related (?) individuals
  9. Anatolian
  10. Greek
  11. Chuvash
  12. Ossetian
  13. A couple of related individuals
  14. A couple of related individuals
  15. Balkar
  16. A couple of related individuals
  17. Chechen
  18. Kumyk
  19. A couple of related individuals
  20. Adygei
  21. Lezgin #1 (main)
  22. Lezgin #2
  23. Lezgin #3
If you belong to a population with few other participants, you might end up latching onto a cluster dominated by a bigger group. This does not mean that your population is not distinctive, only that there are not enough samples to reveal its distinctiveness if it exists.

Inter-Population IBD


Results for Dodecad Participants

Results can be found in the spreadsheet.

If you have joined the Project, please consider leaving a comment in the Information about Project samples thread. That will help others make better sense of their results, e.g., if you find that you belong in the same cluster with some other individual, you might want to know something about their origins.

UPDATE: I have added the IBD sharing matrix.See here on how to use it.

Saturday, January 8, 2011

ADMIXTURE analysis with Dodecad Populations (update #2)

Thanks to all the participants of the Project, the number of populations has increased, and so have sample sizes within pre-existing populations in the Project. There are now 17 populations with at least 5 individuals in the Project:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian, Russian, Spanish, British, Irish, Turkish, N_Italian, Balkans
Below are the K=10 ADMIXTURE results with these populations:

Admixture proportions can be found in the spreadsheet.

The fact that the addition of 17 populations and 143 individuals to the core set of 36 populations and 692 individuals results in the same 10 ancestral components testifies to the stability of this solution. Hopefully, within 2011 I will develop an even better comparison set to work with.

Another test of the validity of the analysis is comparison of independent samples of the same populations:
Ashkenazi, Armenian, Spanish, Turkish, N_Italian
I have a sample of Dodecad Project members for each of the above, as well as a published population. A way to measure the concordance between the two is to calculate the correlation coefficient (rounded to the 3rd decimal point):
  • Ashkenazi Jews: 0.999
  • Armenians: 0.988
  • Spanish: 0.998
  • Turkish: 0.995
  • N_Italian: 0.996
The concordance is remarkable.

I have also made a RAR of "population portraits". It is important to do this to determine whether minor ancestral components represent population-wide phenomena or are limited to a few individuals.

For example, here are the Turks of the Dodecad project:
The sample is a bit more varied than the sample included in Behar et al:
This probably underscores the importance of broad coverage of large countries and ethnic groups, as I have discovered recently in my analysis of 9 different populations of Pakistan.

Another new population are the Irish, presenting a picture of remarkable homogeneity:
Here is the population portrait for the Balkans, which consists of non-Greek, non-Roma inhabitants of the Balkans:
This appears quite varied; hopefully more Balkan project participants will allow me to split this into additional sample populations.

Finally, here is a portrait of the Ashkenazi population, which appears quite similar to the Behar et al. one:
A very interesting thing about this population is the existence of small slices of "East Asian" and "Northeast Asian" components totalling about 1.5% in almost all individuals. In my opinion this testifies to some type of old minor absorption, as it is fairly evenly spread in the population.

If you haven't joined the project yet, feel free to submit your sample during this opportunity.

Friday, November 19, 2010

ADMIXTURE analysis with Dodecad Populations (update #1)

Repeating the previous analysis with additonal populations of Dodecad Project members and/or modified sample sizes for pre-existing ones:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian
Admixture proportions can be found in the spreadsheet. Dodecad Project populations in italics.

Populations portraits can be found in the RAR. For example, here are the ones for Dodecad Project Ashkenazi and Behar et al. (2010) Ashkenazi Jews:
and here is a Portrait of the Portuguese:

Saturday, November 6, 2010

ADMIXTURE analysis with Dodecad Populations

Thanks to our project members, I am now able to repeat the ADMIXTURE analysis which launched this product, with 5 new populations, composed entirely of Project members. I have made a spreadsheet where you can find admixture proportions for the original 36 populations, as well as the 5 new ones:


The new populations are: Assyrians, Scandinavians, Greeks, Finns, and South Italians/Sicilians. As more people sign up to the project, more populations will pass the 5-person threshold, and will be included in the admixture analysis, and the admixture proportions of existing ones will be fine-tuned. Moreover, some of the populations, such as Scandinavians, may be split into e.g., Swedes, Danes, and Norwegians.

Stay tuned for a notice about future opportunities to submit your data.

Wednesday, November 3, 2010

Analysis of Germanic, Slavic, Finnic, and Baltic Europeans


I have taken 25 HapMap-3 White Americans, 5 Dodecad Project Finns, 25 HGDP-CEPH Russians from Vologda, 12 Dodecad Project continental Germanics (Scandinavians and Germans), 10 Behar et al. (2010) Lithuanians, and 9 Behar et al. (2010) Belorussians and 3 Dodecad Project Northern Slavs, in order to study genetic structure in Northern Europe.

The MDS plot has the following color coding: CEU=blue; Finns=red; HGDP Russians=grey; Dodecad Germanics="green"; Belorussians and Dodecad North Slavs="black"; Lithuanians="magenta".

The grouping of samples by populations is apparent: Germanic speakers (blue+green), Balto-Slavs (magenta+black), Finns (red), with North Russians occupying an intermediate position between Finns and the other Slavs.


ADMIXTURE analysis with K=3 confirms this observation: dark blue corresponds to the Germanic cluster; green to the Finnish one, and light blue to the Balto-Slavic one.

Below are population averages:

CEU are of the Germanic component, with a little Balto-Slavic.

Finns are of the Finnic component, with a little Germanic. Note, however, that this does not mean that they do not share substantial ancestry with Balto-Slavs! The Finnic component itself is composite, comprising of both East and West Eurasian elements, the latter of which could very well represent a common element between Finns and Balto-Slavs.

Vologda Russians seem to be primarily Finnic in origin, but with some distinct Balto-Slavic component.

The Germanic group is primarily in the Germanic component, with some Finnic; without revealing any information, I'll just say that this is contributed primarily by 3 Dodecad Project members who deviate towards Finns and whose ADMIXTURE analysis shows a higher than expected Northeast Asian component. Their outlier status is also visible in the MDS plot.

The Balto-Slavic component reaches its maximum in Lithuanians, while my Slavic sample is primarily Balto-Slavic, but shows both Finnic and Germanic admixture.

The study of structure in northern Europe is important not only for the region itself, but also for the world at large. The northern European component occurs at a lower frequency outside North/Central Europe and knowledge about its structure in its homeland may elucidate its origins in other regions.