This was done on the same dataset as the previous fastIBD analysis.
The population assignments:
The heatmap, showing relationship between inferred populations:
The principal components analysis:
The correspondence between inferred populations and K12b components:
Results for Project participants can be found in this spreadsheet; remember than in the chunkcounts tabs, columns represent donor and rows recipient populations.
Showing posts with label Italians. Show all posts
Showing posts with label Italians. Show all posts
Sunday, March 11, 2012
ChromoPainter/fineSTRUCTURE analysis of Italy/Balkans/Anatolia
Monday, March 5, 2012
fastIBD analysis of Italy/Balkans/Anatolia
I have included the new Turkish data from Hodoğlugil & Mahley (2012) in this analysis. Additionally, there are now 5 participants in the Serb_D and Turkish_Cypriot_D sub-populations, as well as a Bosnian Muslim. There are now project participants from many Balkan countries, although Albania, the fYROM, and Croatia remain as "black holes" in the map.
Remember that the tree groups similar populations together, and for each row in the matrix, the red end of the spectrum indicates lots of IBD sharing, and the blue end low IBD sharing. Additionally, I have now calculated the median IBD sharing, which is more resistant in the presence of potential relatives in the data.
Still, I am hopeful that there will be more project participants from currently under-represented populations. I have already started processing the same dataset with ChromoPainter (which takes much longer), and hopefully that analysis will be posted at the end of this week or the beginning of the next one.
First, the heatmap of inter-population IBD:
The results appear fairly reasonable, with the Balkan, Anatolian, and Italian populations of the title forming separate branches, and the mainland Greek sample joining with Central/South Italians and Sicilians.
The Clusters Galore can be seen below; 28 clusters were inferred with 21 dimensions:
Results for Project participants can be found in the spreadsheet, and include the probabilities that each ID is assigned to each of the 28 clusters, as well as the Z-scores comparing each individual against all populations with 5+ individuals. The Z-score should be read as follows: for each row, high values indicate a high degree of IBD sharing, while low values indicate a low degree of IBD sharing.
Of course, I encourage Project participants to leave a message in the Information about Project samples thread.
Tuesday, January 17, 2012
fastIBD analysis of Iberia, France, Italy, Balkans, Anatolia and European Jews
On the heels of the previous analysis of Balkans/West Asia, a new experiment on a different set of populations. Please refer to the earlier post for some thoughts/explanations about this type of analysis, I'll stick to "just the data" for this post.
Clusters Galore
24 clusters inferred with 17 MDS dimensions.
The Galore analysis provides increased resolution within Iberia (#6-9, 11), Italy, and the Ashkenazi Jewish group (#14-16).
The Iberian results are particularly interesting, showing the power of this approach compared to the one with unlinked data. There appear to be:
Inter-Population IBD
Results for Project Participants
The results can be found in the spreadsheet.
Clusters Galore
24 clusters inferred with 17 MDS dimensions.
The Galore analysis provides increased resolution within Iberia (#6-9, 11), Italy, and the Ashkenazi Jewish group (#14-16).
The Iberian results are particularly interesting, showing the power of this approach compared to the one with unlinked data. There appear to be:
- a Spanish Basque (#6),
- French Basque (#11) cluster, as well as
- a Portuguese/Galician/Castilla Y Leon (#9) cluster, and
- a complementary Castilla La Manch/Cantabria/Andalucia/Murcia (#7) cluster, and
- a smaller Aragon/Cataluna cluster (#8).
Inter-Population IBD
Results for Project Participants
The results can be found in the spreadsheet.
Tuesday, May 24, 2011
Italy, the Balkans, and Anatolia
Here is a PCA plot of Italian, Balkan, and Anatolian samples, together with some reference populations. I've removed 2 Roma-admixed Romanians and 3 Northern European-admixed Armenians from the Behar et al. (2010) set.
Below are MCLUST results:
Below you can see the shape of the 7 clusters:
With increasing sample sizes, I expect the validity and distinctiveness of the various clusters to improve; as you can see, there are several samples bordering different clusters or far away from most of them.
If you write to me with your ID, I will send you your results: cluster assignment followed by the two co-ordinates, so that you can locate yourself on the plot.
Saturday, January 8, 2011
ADMIXTURE analysis with Dodecad Populations (update #2)
Thanks to all the participants of the Project, the number of populations has increased, and so have sample sizes within pre-existing populations in the Project. There are now 17 populations with at least 5 individuals in the Project:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian, Russian, Spanish, British, Irish, Turkish, N_Italian, BalkansBelow are the K=10 ADMIXTURE results with these populations:
Admixture proportions can be found in the spreadsheet.
The fact that the addition of 17 populations and 143 individuals to the core set of 36 populations and 692 individuals results in the same 10 ancestral components testifies to the stability of this solution. Hopefully, within 2011 I will develop an even better comparison set to work with.
Another test of the validity of the analysis is comparison of independent samples of the same populations:
Ashkenazi, Armenian, Spanish, Turkish, N_ItalianI have a sample of Dodecad Project members for each of the above, as well as a published population. A way to measure the concordance between the two is to calculate the correlation coefficient (rounded to the 3rd decimal point):
- Ashkenazi Jews: 0.999
- Armenians: 0.988
- Spanish: 0.998
- Turkish: 0.995
- N_Italian: 0.996
I have also made a RAR of "population portraits". It is important to do this to determine whether minor ancestral components represent population-wide phenomena or are limited to a few individuals.
For example, here are the Turks of the Dodecad project:
The sample is a bit more varied than the sample included in Behar et al:
This probably underscores the importance of broad coverage of large countries and ethnic groups, as I have discovered recently in my analysis of 9 different populations of Pakistan.
Another new population are the Irish, presenting a picture of remarkable homogeneity:
Here is the population portrait for the Balkans, which consists of non-Greek, non-Roma inhabitants of the Balkans:
This appears quite varied; hopefully more Balkan project participants will allow me to split this into additional sample populations.
Finally, here is a portrait of the Ashkenazi population, which appears quite similar to the Behar et al. one:
A very interesting thing about this population is the existence of small slices of "East Asian" and "Northeast Asian" components totalling about 1.5% in almost all individuals. In my opinion this testifies to some type of old minor absorption, as it is fairly evenly spread in the population.
If you haven't joined the project yet, feel free to submit your sample during this opportunity.
Friday, November 19, 2010
ADMIXTURE analysis with Dodecad Populations (update #1)
Repeating the previous analysis with additonal populations of Dodecad Project members and/or modified sample sizes for pre-existing ones:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian
Admixture proportions can be found in the spreadsheet. Dodecad Project populations in italics.Populations portraits can be found in the RAR. For example, here are the ones for Dodecad Project Ashkenazi and Behar et al. (2010) Ashkenazi Jews:
and here is a Portrait of the Portuguese:
Labels:
Assyrians,
Dodecad,
Experiments,
Finnic,
Germanic,
Greeks,
Iberians,
Italians,
Jews,
South Asians
Tuesday, November 9, 2010
Multidimensional scaling in Italy, the Balkans, Anatolia, and the Caucasus + Lezgin ADMIXTURE surprise
On the left you can see an MDS plot of several population groups from Italy, the Balkans, Anatolia, and the Caucasus. This combines data of Dodecad Project members with published samples. I have placed the labels manually over the main point blobs. There is also a single Bulgarian sample who falls between the 'm' and the 'a' in 'Romanians'.I had previously studied the distinctiveness of Caucasus populations, and now I have added Turks, Cypriots and populations from further West. I am still not satisfied with my Balkan samples (I have 2 Slovenians, 2 Serbs and 1 Bulgarian), so I encourage Balkan participants to contact me for possible inclusion in the Project.
When I turned to ADMIXTURE, a little mystery emerged, for which I have currently no explanation:
Two main components emerged, a light blue "Italo-Balkan" one that seems deficient in West Asia, and red "Cypriot" one that is deficient in West Balkan Slavs and the Caucasus. The three Caucasus populations, each form their own distinctive cluster (green, yellow, blue), and a magenta low-frequency component emerges at K=6, which is why I stopped the analysis at this K. Results for K=5 were similar, minus this low-frequency component.
Here is the big puzzle: my Bulgarian, 2 Serbs, 2 Slovenians, all show unambiguous membership in the green "Lezgin" cluster. Out of all the Caucasus components, this is the only one that seems to have a Balkan connection. While one could argue that this might reflect Neolithic farmers, as it has been argued that they spoke a North Caucasian language, the same "Lezgin" component is insignificant in Greeks and Mixed-Greeks, Southern Italians/Sicilians and Italian (other).
Is this some signal of a population that once inhabited the northern arc of the Black sea, from the Balkans to the Caucasus? This might find some support in the possession by both Lezgins (and Balkan Slavs) of a "North European" component, but the Adygei, who similarly possess such a component show no special affinity with Balkan Slavs. Below is the Lezgin K=10 portrait:


If anyone has any (pre)-historical scenario that might account for this unexpected affinity, feel free to write to me or leave a comment.
Saturday, November 6, 2010
ADMIXTURE analysis with Dodecad Populations
Thanks to our project members, I am now able to repeat the ADMIXTURE analysis which launched this product, with 5 new populations, composed entirely of Project members. I have made a spreadsheet where you can find admixture proportions for the original 36 populations, as well as the 5 new ones:
The new populations are: Assyrians, Scandinavians, Greeks, Finns, and South Italians/Sicilians. As more people sign up to the project, more populations will pass the 5-person threshold, and will be included in the admixture analysis, and the admixture proportions of existing ones will be fine-tuned. Moreover, some of the populations, such as Scandinavians, may be split into e.g., Swedes, Danes, and Norwegians.
Stay tuned for a notice about future opportunities to submit your data.
Tuesday, November 2, 2010
Analysis of Greeks, Italians, Cypriots, and European Jews
Here are the first two dimensions of a multidimensional scaling in which the following are included:- Behar et al. (2010) Cypriots, and Ashkenazi/Sephardi Jews
- HGDP-CEPH North Italians and Tuscans
- Dodecad Project Greeks and Part-Greeks (N=7)
- Dodecad Project Italians (N=14)
- Dodecad Project Ashkenazi Jews (N=2)
The Jewish/non-Jewish groups can be distinguished by the first axis, with value greater than 0 including almost all Jews and less than 0 including almost all non-Jews. A group of Sephardic Jews, while still distinguishable, nonetheless approaches the non-Jewish individuals. According to Behar et al. (2010):
Samples from the Bulgarian and Turkish communities were combined in the analysis and are referred to as Sephardi Jews.
The separability of Ashkenazi Jews from non-Jewish Europeans had been previously reported, e.g., by Tian et al. (2008), from which the figure on the left is presented.The current post adds Sephardic Jews into the picture; these are also distinguishable from non-Jewish Europeans, but appear to be somewhat closer to them -- at least in the first 2 dimensions of an MDS plot.
While the MDS plot gives a nice overview of this distinctiveness, to assess the possible admixture between these various groups, we must turn to ADMIXTURE.
Indeed, there are two groups: dark blue (European Jews), and light blue (Cypriots, Italians, Greeks). I don't want to put too much emphasis on this, as there are obviously many factors at play, e.g., (i) admixture with other elements, not represented in this study (e.g., Spaniards or Anatolians for Sephardi Jews / North-Central Europeans for Ashkenazi Jews / Syrians, or Phoenicians, or Arabs for Southern Europeans), or (ii) inability to detect admixture accurately among closely related groups (the two clusters are at an Fst=0.029); previously I had seen some residual admixture among Tuscans and White Americans
Nonetheless, the ADMIXTURE analysis confirms that (i) non-Jewish Greeks (including Cypriots) and Italians are distinguishable from European Jews, and (ii) Sephardi Jews are closer to these Europeans than Ashkenazi are.
PS: Idea for this post by recent Italianthro post.
Labels:
Cypriots,
Experiments,
Greeks,
Italians,
Jews
Subscribe to:
Posts (Atom)























