This was done on the same dataset as the previous fastIBD analysis.
The population assignments:
The heatmap, showing relationship between inferred populations:
The principal components analysis:
The correspondence between inferred populations and K12b components:
Results for Project participants can be found in this spreadsheet; remember than in the chunkcounts tabs, columns represent donor and rows recipient populations.
Showing posts with label Balkans. Show all posts
Showing posts with label Balkans. Show all posts
Sunday, March 11, 2012
ChromoPainter/fineSTRUCTURE analysis of Italy/Balkans/Anatolia
Monday, March 5, 2012
fastIBD analysis of Italy/Balkans/Anatolia
I have included the new Turkish data from Hodoğlugil & Mahley (2012) in this analysis. Additionally, there are now 5 participants in the Serb_D and Turkish_Cypriot_D sub-populations, as well as a Bosnian Muslim. There are now project participants from many Balkan countries, although Albania, the fYROM, and Croatia remain as "black holes" in the map.
Remember that the tree groups similar populations together, and for each row in the matrix, the red end of the spectrum indicates lots of IBD sharing, and the blue end low IBD sharing. Additionally, I have now calculated the median IBD sharing, which is more resistant in the presence of potential relatives in the data.
Still, I am hopeful that there will be more project participants from currently under-represented populations. I have already started processing the same dataset with ChromoPainter (which takes much longer), and hopefully that analysis will be posted at the end of this week or the beginning of the next one.
First, the heatmap of inter-population IBD:
The results appear fairly reasonable, with the Balkan, Anatolian, and Italian populations of the title forming separate branches, and the mainland Greek sample joining with Central/South Italians and Sicilians.
The Clusters Galore can be seen below; 28 clusters were inferred with 21 dimensions:
Results for Project participants can be found in the spreadsheet, and include the probabilities that each ID is assigned to each of the 28 clusters, as well as the Z-scores comparing each individual against all populations with 5+ individuals. The Z-score should be read as follows: for each row, high values indicate a high degree of IBD sharing, while low values indicate a low degree of IBD sharing.
Of course, I encourage Project participants to leave a message in the Information about Project samples thread.
Wednesday, February 15, 2012
Correspondence between ChromoPainter clusters and ADMIXTURE components in Balkans/West Asia
I took the 25 different inferred clusters from my recent ChromoPainter analysis, and calculated their normalized median components in terms of the K12b calculator. This is a quite useful exercise, since it can show in what sense clusters are different from each other.
Here are two ways in which you may use this correspondence.
1. Different clusters of a single population
For example, the Turks with partial Balkan ancestry tend to belong to pop10, whereas those of Anatolian ancestry to pop13, and those from northeastern Anatolia to pop22. If we compare the admixture proportions of these three groups, we notice e.g.,
2. Related clusters
fineSTRUCTURE grouped the different populations in a tree structure. For example, it grouped pop18, the "North Balkan" cluster with pop23, the "Bulgarian-Romanian" one.
Looking at the admixture proportions, we can tell that the two clusters do indeed seem quite similar, but there are some differences, e.g., an excess of North_European in pop18, and an excess of Caucasus in pop23. This makes sense given the geographical origin of individuals belonging to the two clusters.
1. Different clusters of a single population
For example, the Turks with partial Balkan ancestry tend to belong to pop10, whereas those of Anatolian ancestry to pop13, and those from northeastern Anatolia to pop22. If we compare the admixture proportions of these three groups, we notice e.g.,
- An excess of Atlantic_Baltic and North_European in pop10
- An excess of Caucasus in pop22
2. Related clusters
fineSTRUCTURE grouped the different populations in a tree structure. For example, it grouped pop18, the "North Balkan" cluster with pop23, the "Bulgarian-Romanian" one.
Looking at the admixture proportions, we can tell that the two clusters do indeed seem quite similar, but there are some differences, e.g., an excess of North_European in pop18, and an excess of Caucasus in pop23. This makes sense given the geographical origin of individuals belonging to the two clusters.
Labels:
Anatolia,
Balkans,
ChromoPainter,
Experiments,
Greeks,
Iranian,
Slavic,
Turkic
Tuesday, February 14, 2012
ChromoPainter/fineSTRUCTURE analysis of Balkans/West Asia
I have carried out a ChromoPainter/fineSTRUCTURE analysis of Balkans/West Asia. This is a slightly different dataset than the one used in the previous fastIBD analysis of the same region. It also took much longer (about a week, with two CPUs dedicated to the task) to complete, so it is not something that can be done routinely.
Technical details (skip if you want)
413 individuals from 33 populations were studied, on 258,100 SNPs, after --geno 0.03 --maf 0.01 filters were applied. Data were phased in Beagle with the default 10 iterations. Genetic maps from the HapMap were used. fineSTRUCTURE was used on ChromoPainter output, with 500,000 burnin/runtime iterations each.
25 Inferred Populations
fineSTRUCTURE imposes a tree structure on a number of inferred populations. The following heatmap shows this tree structure; columns represent donor populations, rows, recipient ones.
There was a total of 25 populations, labeled pop0, pop1, ..., pop24.
The following table summarizes how many individuals from each original population were assigned to each inferred population:
I will limit myself to populations which include Dodecad Project members:
I have also used the PCA feature of fineSTRUCTURE to carry out principal components analysis. I am plotting the first two dimensions of this PCA, using my own visualization code that places labels in the average position on the plane:
Results
Results for Project participants are included in the spreadsheet.
The raw chunkcounts for all 413x413 individuals can be found here.
Technical details (skip if you want)
413 individuals from 33 populations were studied, on 258,100 SNPs, after --geno 0.03 --maf 0.01 filters were applied. Data were phased in Beagle with the default 10 iterations. Genetic maps from the HapMap were used. fineSTRUCTURE was used on ChromoPainter output, with 500,000 burnin/runtime iterations each.
25 Inferred Populations
fineSTRUCTURE imposes a tree structure on a number of inferred populations. The following heatmap shows this tree structure; columns represent donor populations, rows, recipient ones.
There was a total of 25 populations, labeled pop0, pop1, ..., pop24.
The following table summarizes how many individuals from each original population were assigned to each inferred population:
I will limit myself to populations which include Dodecad Project members:
- pop6 includes a Project North Ossetian, as well as all Yunusbayev et al. North Ossetians
- pop7 is mainly Armenian
- pop16 is also mainly Armenian; it would be interesting to see whether this bipartite division of Armenians is in agreement with the one inferred in the previous fastIBD analysis
- pop8 is mainly Greek, and appears to be "continental Greek"; it also includes some other Balkan individuals
- pop14 is also Greek, and includes a variety of people with ancestry from Crete, the Aegean, Cyprus, Asia Minor, Cappadocia, and the Pontus as well as continental Greek. It could be labeled "eastern Greek"
- pop11 is Cypriot, including the single 100% Greek Cypriot of the Project, all 3 100% Turkish Cypriots, as well as a Turkish individual of partial Turkish_Cypriot ancestry
- pop10 is Turkish, and includes people with some ancestry from the Balkans, as well as Anatolia. It could be labelled "Balkan Turkish"
- pop13 is also Turkish, and seems to include people with ancestry exclusively from Anatolia, including almost all the Behar et al. Turks
- pop15 is Assyrian; some Assyrians also fall on the aforementioned pop16 which includes mainly Armenians
- pop18 could be labelled "North Balkan"; there is probably structure to be uncovered within this cluster, once more participants from the Balkans join the Project
- pop20 is "Georgian-Abkhazian"
- pop21 is "Kurdish-Iranian"
- pop22 could be labeled "Northeastern Anatolia" or (more classically) "Pontus-Colchis". It appears to unite various individuals from Northeastern Turkey and neighboring Georgia, having Karadeniz Turkish, Armenian, Pontic Greek, and Kartvelian ancestry. I strongly encourage participants from this region to join the Project, especially Pontic Greeks, as there are no 100% Pontic Greeks currently in the Project.
- pop23 is "Bulgarian-Romanian" mainly, and also includes one Serb. Once again, I emphasize that the power of this approach using haplotypes depends on participation, so I encourage all people from the Balkans to consider joining the Project.
I have also used the PCA feature of fineSTRUCTURE to carry out principal components analysis. I am plotting the first two dimensions of this PCA, using my own visualization code that places labels in the average position on the plane:
Results
Results for Project participants are included in the spreadsheet.
- Population matrix, shows how many individuals from each population were assigned to each cluster
- Z score population matrix, shows the normalized number of "chunks" from each donor population (columns) to each recipient (row). Do not compare across rows! The way to read this table is the following: for each row, higher values indicate more sharing. For example, the "Cypriots" population has pop11 as its main donor.
- Individual assignments: the pop number that all Project and reference IDs were assigned to
- Individual Chunkcounts: the number of chunks copied from its donor population (column) to each individual
- Individual PCA: your PCA co-ordinates that can help you find your dot on the Principal Components Analysis graphic (see above)
The raw chunkcounts for all 413x413 individuals can be found here.
Labels:
Anatolia,
Balkans,
Caucasus,
ChromoPainter,
Results
Saturday, January 14, 2012
fastIBD analysis of Balkans/West Asia
Now that I've discovered a way to boost Clusters Galore analysis even further by using fastIBD, I will start experimenting with different regional populations. This analysis took about 5 hours to complete, so it appears to be quite practical.
For my first experiment, I carry out an analysis of various populations from the Balkans and West Asia.
Clusters Galore
27 different clusters were inferred with 17 MDS dimensions. Some interesting findings:
Inter-population IBD
You can also see a visual representation of inter-population IBD:
I have only included populations with 5+ participants in this representation. Reddish shades express high IBD sharing; bluish ones low one. The heatmap has been scaled by row.
As you might expect, values across the diagonal are "reddish", since individuals within populations tend to have high IBD sharing with each other.
A few features "pop out" of the screen. Going from top to bottom:
Results for Project Participants
The results can be found in the spreadsheet, and include:
If you haven't joined the Project yet, I encourage you to do so if you are eligible.
For my first experiment, I carry out an analysis of various populations from the Balkans and West Asia.
Clusters Galore
27 different clusters were inferred with 17 MDS dimensions. Some interesting findings:
- For the first time there emerge a couple of clusters that appear to be quite specific to Armenians (#2 and #3).
- Similarly, Assyrians are broken to a few clusters that appear fairly specific to them (#9-11)
- Georgians are split into three clusters, one of which (#14) is linked with the neighboring Abkhasians, who in turn have their own exclusive cluster (#25)
- The cluster modal in Greeks (#6) includes 14 of 19 Greek participants, and a few Greeks are also in the Balkan cluster (#8) and an Iranian-Turkish cluster (#4)
- The Behar Cypriot sample also splits into two, and the few Turkish Cypriot participants link to one of them (#13)
- The Ossetian project participant links to one of the three North_Ossetian clusters
- The major Balkan cluster (#8) still defies resolution. I am certain, however, that structure in this cluster will be uncovered with more participation. MCLUST adapts the cluster size and shape, and a "big", inclusive cluster spanning the Balkans appears more parsimonious than smaller clusters centered on the different groups. With larger participation, I anticipate that regional structure will be uncovered in the Balkans as well.
- Discover group-specific clusters, by identifying what is common between members of groups
- Discover within-group clusters, by identifying what is different between members of groups
Inter-population IBD
You can also see a visual representation of inter-population IBD:
I have only included populations with 5+ participants in this representation. Reddish shades express high IBD sharing; bluish ones low one. The heatmap has been scaled by row.
As you might expect, values across the diagonal are "reddish", since individuals within populations tend to have high IBD sharing with each other.
A few features "pop out" of the screen. Going from top to bottom:
- Intra-Iranic sharing
- Intra-Armenian sharing
- Intra-Balkan sharing
- Georgian-Abkhaz sharing
Results for Project Participants
The results can be found in the spreadsheet, and include:
- Probabilities of assignment in each of the 27 clusters of the Clusters Galore analysis
- Z-scores of IBD between each individual and each of the 20 populations with 5+ participants. Higher values mean more IBD sharing. Note that Z-scores have been calculated for each row, hence each participant must scan his own row to find populations with an excess (+) or deficiency (-) of IBD sharing, and people should not compare across different rows.
If you haven't joined the Project yet, I encourage you to do so if you are eligible.
Sunday, September 25, 2011
Yunusbayev et al. (2011) data assessed with Dodecad v3
I have acquired the data from the recent Yunusbayev et al. (2011) paper on the Caucasus. This includes the following populations:
- Kurds_Y 6
- Bulgarians_Y 13
- Ukranians_Y 20
- Mordovians_Y 15
- Armenians_Y 16
- Abhkasians_Y 20
- Balkars_Y 19
- North_Ossetians_Y 15
- Chechens_Y 20
- Nogais_Y 16
- Kumyks_Y 14
- Turkmens_Y 15
- Tajiks_Y 15
It is a valuable new addition to the Project, and it is commendable that it has been made publicly and easily available so swiftly after the appearance of the Yunusbayev et al. (2011) paper.
To get the ball rolling on the new Yunusbayev et al. data, I will map the new populations onto the Dodecad v3 components; they will be added to the Dodecad v3 spreadsheet as they are calculated.
I have been laboriously designing a new global (including Amerindians and Australasians) Dodecad X1 experimental calculator with 3,010 individuals for a few weeks now, but I guess I will now have to reboot it with 3,214.
Together with some other new data I recently discovered, I now have 9,799 individuals (some duplicates from different sources) in my global database. My Dodecad dataset of 511 individuals from a single country or ethnic group isn't too shabby either. Let's hope for a new data release that will push the data collection above the magic 10,000.
UDDATE:
I have added the first 7 populations to the spreadsheet; the others are being calculated as we speak. Most of them seem in line with expectations, but the Abkhasian sample has one outlier individual (abh27), and has thus been placed in the "Outliers" tab of the spreadsheet; a new set of admixture proportions, minus that outlier individual, will be calculated anew:
UPDATE II: The population portraits have been uploaded to Google Docs as a rar file (Sendspace mirror). Average admixture results have all been entered to the spreadsheet.
Monday, September 5, 2011
'bat' calculator (Balkans-Anatolia-Turkic)
I have decided to make a new calculator for DIYDodecad that may be useful for individuals from the Balkans and Anatolia. You can download it from here at Google Docs (or here from sendspace). The terms of use are the same as for DIYDodecad v 2.0. To run it, you simply extract the contents of the RAR file in your working directory, and type bat.par whenever you typed dv3.par in the instructions.
The reference populations can be seen below. I have included all available Balkan populations, as well as Turks and Armenians. Moreover, I have included all available Turkic populations.
The marker set is the same as used in Dodecad v3. Three components emerge: one centered in the northern Balkans, one in eastern Anatolia, and one present in various proportions among all Turkic populations (see Turkic cline).

The components have been named accordingly, but please note that they do not necessarily reflect recent ancestors. For example, it is a good hypothesis that the Anatolia component was present in the Balkans even in ancient times, so one need not seek a recent Anatolian ancestor to explain its presence in a Balkan individual. Similarly for the Balkans component in Anatolia, which may reflect the diverse Balkan peoples that have settled in Anatolia since the dawn of history, so a present-day inhabitant of Anatolia need not seek a recent Balkan ancestor.
Likewise, the Turkic component is only part of the genetic makeup of the Turkic speakers who arrived in Anatolia, since those probably also carried West Eurasian population elements picked up en route from Siberia to Anatolia.
The way to interpret your results is to see whether you have an excess or deficiency of any component relative to your ethnic group. For example, an Anatolian Greek may have a higher Anatolia/Balkans ratio than a Balkan Greek and likewise for a Balkan vs. Anatolian Turk; the latter may also have a variable Turkic component which will reflect differential Central Asian input.
Tuesday, August 30, 2011
Balkan averages (August 2011)
Since my last call for more participation from the Balkans, I was able to create a new Bulgarian_D sample of 5 participants. Together with the Greek_D sample, the Balkans_D sample of non-Greek, non-Bulgarian project members, the Behar et al. (2010) Romanians, and the Xing et al. (2010) Slovenians (the latter on a smaller number of markers), we are beginning to get a better feel of genetic variation in the Balkans. There have been several other averages that have been adjusted with more participation; all of them can be seen in the Dodecad v3 spreadsheet.
The table below shows the major components (>1%) in the available Balkan populations.
The Bulgarian average as it stands seems reasonably close to the Romanian one, and is characterized by balanced West/East European components; in this balance it resembles Greeks, who, however, have lower levels of both components and higher levels of the Mediterranean/West Asian/Southwest Asian components.
Slovenians contrast with Hungarians in having reverse West/East European levels, and with their neighboring Italians in having quite a bit more of the East European component, and quite a bit less of the Mediterranean one. Bulgarians/Romanians contrast with Slavic groups from eastern Europe in having less East European and more Mediterranean/West_Asian.
Hopefully before long, more participation from the western/central Balkans (Serbs, Croats, Bosnians, Montenegrins, Albanians, Slav Macedonians) will allow us to fill more holes in our understanding of the genetic landscape of Southeastern Europe.
Tuesday, May 24, 2011
Italy, the Balkans, and Anatolia
Here is a PCA plot of Italian, Balkan, and Anatolian samples, together with some reference populations. I've removed 2 Roma-admixed Romanians and 3 Northern European-admixed Armenians from the Behar et al. (2010) set.
Below are MCLUST results:
Below you can see the shape of the 7 clusters:
With increasing sample sizes, I expect the validity and distinctiveness of the various clusters to improve; as you can see, there are several samples bordering different clusters or far away from most of them.
If you write to me with your ID, I will send you your results: cluster assignment followed by the two co-ordinates, so that you can locate yourself on the plot.
Saturday, January 8, 2011
ADMIXTURE analysis with Dodecad Populations (update #2)
Thanks to all the participants of the Project, the number of populations has increased, and so have sample sizes within pre-existing populations in the Project. There are now 17 populations with at least 5 individuals in the Project:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian, Russian, Spanish, British, Irish, Turkish, N_Italian, BalkansBelow are the K=10 ADMIXTURE results with these populations:
Admixture proportions can be found in the spreadsheet.
The fact that the addition of 17 populations and 143 individuals to the core set of 36 populations and 692 individuals results in the same 10 ancestral components testifies to the stability of this solution. Hopefully, within 2011 I will develop an even better comparison set to work with.
Another test of the validity of the analysis is comparison of independent samples of the same populations:
Ashkenazi, Armenian, Spanish, Turkish, N_ItalianI have a sample of Dodecad Project members for each of the above, as well as a published population. A way to measure the concordance between the two is to calculate the correlation coefficient (rounded to the 3rd decimal point):
- Ashkenazi Jews: 0.999
- Armenians: 0.988
- Spanish: 0.998
- Turkish: 0.995
- N_Italian: 0.996
I have also made a RAR of "population portraits". It is important to do this to determine whether minor ancestral components represent population-wide phenomena or are limited to a few individuals.
For example, here are the Turks of the Dodecad project:
The sample is a bit more varied than the sample included in Behar et al:
This probably underscores the importance of broad coverage of large countries and ethnic groups, as I have discovered recently in my analysis of 9 different populations of Pakistan.
Another new population are the Irish, presenting a picture of remarkable homogeneity:
Here is the population portrait for the Balkans, which consists of non-Greek, non-Roma inhabitants of the Balkans:
This appears quite varied; hopefully more Balkan project participants will allow me to split this into additional sample populations.
Finally, here is a portrait of the Ashkenazi population, which appears quite similar to the Behar et al. one:
A very interesting thing about this population is the existence of small slices of "East Asian" and "Northeast Asian" components totalling about 1.5% in almost all individuals. In my opinion this testifies to some type of old minor absorption, as it is fairly evenly spread in the population.
If you haven't joined the project yet, feel free to submit your sample during this opportunity.
Tuesday, November 9, 2010
Multidimensional scaling in Italy, the Balkans, Anatolia, and the Caucasus + Lezgin ADMIXTURE surprise
On the left you can see an MDS plot of several population groups from Italy, the Balkans, Anatolia, and the Caucasus. This combines data of Dodecad Project members with published samples. I have placed the labels manually over the main point blobs. There is also a single Bulgarian sample who falls between the 'm' and the 'a' in 'Romanians'.I had previously studied the distinctiveness of Caucasus populations, and now I have added Turks, Cypriots and populations from further West. I am still not satisfied with my Balkan samples (I have 2 Slovenians, 2 Serbs and 1 Bulgarian), so I encourage Balkan participants to contact me for possible inclusion in the Project.
When I turned to ADMIXTURE, a little mystery emerged, for which I have currently no explanation:
Two main components emerged, a light blue "Italo-Balkan" one that seems deficient in West Asia, and red "Cypriot" one that is deficient in West Balkan Slavs and the Caucasus. The three Caucasus populations, each form their own distinctive cluster (green, yellow, blue), and a magenta low-frequency component emerges at K=6, which is why I stopped the analysis at this K. Results for K=5 were similar, minus this low-frequency component.
Here is the big puzzle: my Bulgarian, 2 Serbs, 2 Slovenians, all show unambiguous membership in the green "Lezgin" cluster. Out of all the Caucasus components, this is the only one that seems to have a Balkan connection. While one could argue that this might reflect Neolithic farmers, as it has been argued that they spoke a North Caucasian language, the same "Lezgin" component is insignificant in Greeks and Mixed-Greeks, Southern Italians/Sicilians and Italian (other).
Is this some signal of a population that once inhabited the northern arc of the Black sea, from the Balkans to the Caucasus? This might find some support in the possession by both Lezgins (and Balkan Slavs) of a "North European" component, but the Adygei, who similarly possess such a component show no special affinity with Balkan Slavs. Below is the Lezgin K=10 portrait:


If anyone has any (pre)-historical scenario that might account for this unexpected affinity, feel free to write to me or leave a comment.
Subscribe to:
Posts (Atom)

























