Showing posts with label Caucasus. Show all posts
Showing posts with label Caucasus. Show all posts

Sunday, March 11, 2012

ChromoPainter/fineSTRUCTURE analysis of Italy/Balkans/Anatolia

This was done on the same dataset as the previous fastIBD analysis.

The population assignments:



The heatmap, showing relationship between inferred populations:


The principal components analysis:



The correspondence between inferred populations and K12b components:



Results for Project participants can be found in this spreadsheet; remember than in the chunkcounts tabs, columns represent donor and rows recipient populations.

Tuesday, February 14, 2012

ChromoPainter/fineSTRUCTURE analysis of Balkans/West Asia

I have carried out a ChromoPainter/fineSTRUCTURE analysis of Balkans/West Asia. This is a slightly different dataset than the one used in the previous fastIBD analysis of the same region. It also took much longer (about a week, with two CPUs dedicated to the task) to complete, so it is not something that can be done routinely.

Technical details (skip if you want)


413 individuals from 33 populations were studied, on 258,100 SNPs, after --geno 0.03 --maf 0.01 filters were applied. Data were phased in Beagle with the default 10 iterations. Genetic maps from the HapMap were used. fineSTRUCTURE was used on ChromoPainter output, with 500,000 burnin/runtime iterations each.

25 Inferred Populations


fineSTRUCTURE imposes a tree structure on a number of inferred populations. The following heatmap shows this tree structure; columns represent donor populations, rows, recipient ones.


There was a total of 25 populations, labeled pop0, pop1, ..., pop24.

The following table summarizes how many individuals from each original population were assigned to each inferred population:


I will limit myself to populations which include Dodecad Project members:

  • pop6 includes a Project North Ossetian, as well as all Yunusbayev et al. North Ossetians
  • pop7 is mainly Armenian
  • pop16 is also mainly Armenian; it would be interesting to see whether this bipartite division of Armenians is in agreement with the one inferred in the previous fastIBD analysis
  • pop8 is mainly Greek, and appears to be "continental Greek"; it also includes some other Balkan individuals
  • pop14 is also Greek, and includes a variety of people with ancestry from Crete, the Aegean, Cyprus, Asia Minor, Cappadocia, and the Pontus as well as continental Greek. It could be labeled "eastern Greek"
  • pop11 is Cypriot, including the single 100% Greek Cypriot of the Project, all 3 100% Turkish Cypriots, as well as a Turkish individual of partial Turkish_Cypriot ancestry
  • pop10 is Turkish, and includes people with some ancestry from the Balkans, as well as Anatolia. It could be labelled "Balkan Turkish"
  • pop13 is also Turkish, and seems to include people with ancestry exclusively from Anatolia, including almost all the Behar et al. Turks
  • pop15 is Assyrian; some Assyrians also fall on the aforementioned pop16 which includes mainly Armenians
  • pop18 could be labelled "North Balkan"; there is probably structure to be uncovered within this cluster, once more participants from the Balkans join the Project
  • pop20 is "Georgian-Abkhazian"
  • pop21 is "Kurdish-Iranian"
  • pop22 could be labeled "Northeastern Anatolia" or (more classically) "Pontus-Colchis". It appears to unite various individuals from Northeastern Turkey and neighboring Georgia, having Karadeniz Turkish, Armenian, Pontic Greek, and Kartvelian ancestry. I strongly encourage participants from this region to join the Project, especially Pontic Greeks, as there are no 100% Pontic Greeks currently in the Project.
  • pop23 is "Bulgarian-Romanian" mainly, and also includes one Serb. Once again, I emphasize that the power of this approach using haplotypes depends on participation, so I encourage all people from the Balkans to consider joining the Project.
Principal Components Analysis


I have also used the PCA feature of fineSTRUCTURE to carry out principal components analysis. I am plotting the first two dimensions of this PCA, using my own visualization code that places labels in the average position on the plane:


Results


Results for Project participants are included in the spreadsheet.

  • Population matrix, shows how many individuals from each population were assigned to each cluster
  • Z score population matrix, shows the normalized number of "chunks" from each donor population (columns) to each recipient (row). Do not compare across rows! The way to read this table is the following: for each row, higher values indicate more sharing. For example, the "Cypriots" population has pop11 as its main donor.
  • Individual assignments: the pop number that all Project and reference IDs were assigned to
  • Individual Chunkcounts: the number of chunks copied from its donor population (column) to each individual
  • Individual PCA: your PCA co-ordinates that can help you find your dot on the Principal Components Analysis graphic (see above)
Averaged results were included only for populations with >=5 members.
The raw chunkcounts for all 413x413 individuals can be found here.

Saturday, January 14, 2012

fastIBD analysis of Balkans/West Asia

Now that I've discovered a way to boost Clusters Galore analysis even further by using fastIBD, I will start experimenting with different regional populations. This analysis took about 5 hours to complete, so it appears to be quite practical.

For my first experiment, I carry out an analysis of various populations from the Balkans and West Asia.

Clusters Galore

27 different clusters were inferred with 17 MDS dimensions. Some interesting findings:
  • For the first time there emerge a couple of clusters that appear to be quite specific to Armenians (#2 and #3). 
  • Similarly, Assyrians are broken to a few clusters that appear fairly specific to them  (#9-11)
  • Georgians are split into three clusters, one of which (#14) is linked with the neighboring Abkhasians, who in turn have their own exclusive cluster (#25)
  • The cluster modal in Greeks (#6) includes 14 of 19 Greek participants, and a few Greeks are also in the Balkan cluster (#8) and an Iranian-Turkish cluster (#4)
  • The Behar Cypriot sample also splits into two, and the few Turkish Cypriot participants link to one of them (#13)
  • The Ossetian project participant links to one of the three North_Ossetian clusters
  • The major Balkan cluster (#8) still defies resolution. I am certain, however, that structure in this cluster will be uncovered with more participation. MCLUST adapts the cluster size and shape, and a "big", inclusive cluster spanning the Balkans appears more parsimonious than smaller clusters centered on the different groups. With larger participation, I anticipate that regional structure will be uncovered in the Balkans as well.
I cannot stress the importance of participation strongly enough. When groups have more participants, it is possible to both:

  1. Discover group-specific clusters, by identifying what is common between members of groups
  2. Discover within-group clusters, by identifying what is different between members of groups
For example, the great participation of Armenians in the Project has now allowed me to discover structure within the Armenian population. It appears, that cluster #2 corresponds to a more "western" Armenian group, and #3 to a more "eastern" one, with some overlap between the two.

Inter-population IBD


You can also see a visual representation of inter-population IBD:

I have only included populations with 5+ participants in this representation. Reddish shades express high IBD sharing; bluish ones low one. The heatmap has been scaled by row.

As you might expect, values across the diagonal are "reddish", since individuals within populations tend to have high IBD sharing with each other.

A few features "pop out" of the screen. Going from top to bottom:
  • Intra-Iranic sharing
  • Intra-Armenian sharing
  • Intra-Balkan sharing
  • Georgian-Abkhaz sharing
You can probably get more out of the figure, but these appear to be the most salient features.

Results for Project Participants


The results can be found in the spreadsheet, and include:
  • Probabilities of assignment in each of the 27 clusters of the Clusters Galore analysis
  • Z-scores of IBD between each individual and each of the 20 populations with 5+ participants. Higher values mean more IBD sharing. Note that Z-scores have been calculated for each row, hence each participant must scan his own row to find populations with an excess (+) or deficiency (-) of IBD sharing, and people should not compare across different rows.
Last but not least, I want to remind new project participants to leave a message in the Information about Project samples thread. Your comment will not appear immediately, since comment moderation is on, and also note that there are multiple pages of comments. 


If you haven't joined the Project yet, I encourage you to do so if you are eligible.

Wednesday, October 26, 2011

'eurasia7' calculator

This calculator was made with 196 different populations and 2,659 individuals, including 518 project participants. The following Dodecad populations do not have 5 individuals yet, so they are included in the OTHERS_D generic category:
Algerian_D, North_African_Jews_D, Slovenian_D, Mixed_Scandinavian_D, Danish_D, Moroccan_D, Tunisian_D, Serb_D, Austrian_D, Saudi_D, Pakistani_D, Tatar_Various_D, Palestinian_D, Greek_Italian_D, Romanian_D, Swiss_German_D, Szekler_D, Mandaean_D, Azeri_D, Czech_D, Georgian_D, Belgian_D, Latvian_D, Estonian_D, Bangladesh_D, Yemenese_D, Sri_Lanka_D, Hungarian_D, Basque_D, Udmurt_D, Egyptian_D
As always, I encourage people with 4 grandparents from the same country or ethnic group of Eurasia, North or East Africa to contact me (do not send data!) for possible inclusion in the Project. If I have overlooked any such individuals, drop me a line (my e-mail address is at the bottom of the blog). I usually start a new _D population whenever individuals with 4 grandparents from the same group are submitted, but I may have missed some.

Note that all individuals from the reference populations have also been included, including outliers; you should be aware of this when reading the population averages, and consult the Outliers tab in the v3 spreadsheet for some instances of outliers.
Due to image size restrictions in Picasa, the labels are not visible well. A large version of the above plot can be found in the download bundle.

The seven ancestral populations inferred at this level of resolution are:
  • Sub_Saharan
  • West_Asian
  • Atlantic_Baltic
  • East_Asian
  • Southern
  • South_Asian
  • Siberian
As usual, you should take these names as useful labels, and interpret them in conjunction with the components' distribution in different populations, and their Fst distances, both of which can be found in the spreadsheet.

The table of Fst distances:


Below you can see a neighbor-joining tree based on inter-population Fst distances:
The first six dimensions of a multi-dimensional scaling of the same:





Calculator Files:

  • The spreadsheet contains population averages, the table of Fst distances, and individual results for included Project participants.
  • The download RAR file (Google Docs or Sendspace) contains all the files needed to run the calculator. You must download and install DIYDodecad 2.1 first. In order to run the calculator, you follow the instructions of the README file, but type 'eurasia7' instead of 'dv3'.

Terms of use: 'eurasia7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Technical Details:

The calculator is built using allele frequencies of K=7 ancestral components inferred by ADMIXTURE 1.21 analysis of 2,659 individuals. Markers included in the source datasets, as well as the Family Finder and 23andMe (as of Oct 21) platforms were included. The marker set was thinned of markers with less than 99.5% genotype rate and less than 0.5% minor allele frequency. Linkage-disequilibrium based pruning was carried out with a window size of 250 SNPs, advanced by 25 SNPs and R-squared greater than 0.4. A total of 164,990 SNPs remained after these filtering steps.

All relevant populations available to me, and genotyped at a sufficient number of markers were included. Inclusion of the Kalash population resulted in a population-specific component at K=7, and hence their admixture components were inferred a posteriori. Their proportions are consistent with previous results, showing them to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components.

Friday, September 30, 2011

'euro7' calculator

I am releasing a new calculator for Europeans, including their immediate neighboring populations around the Black Sea (Caucasus and Anatolia). The calculator can be used with DIYDodecad

There are additional African and Far-Asian population controls, so, in principle, the calculator could be used by non-Europeans/Anatolians/Caucasians, although I would be less confident of their results. For example, people of South Asian ancestry may obtain a Far-Asian result if they use this calculator, due to the deep affinity of Ancestral South Indians with East Asians. Other West Eurasians and West Eurasian-admixed peoples, not from the studied regions (e.g., Arabians or East Africans) will have their West Eurasian components mapped onto the ones used in this calculator.

'euro7' uses 7 ancestral components:
  • Caucasus
  • Northwestern
  • Northeastern
  • Southeastern
  • African
  • Far_Asian
  • Southwestern
These names represent 7 ancestral populations inferred by ADMIXTURE, and have been chosen based on the geographical regions where each of them achieves its maximum representation. You should always refer to A note of caution on admixture estimates, Interpretation of ADMIXTURE results: component sharing, as well as the average population values in the spreadsheet when interpreting your individual results.

The distribution of these 7 components can be seen in the barplot on the top left, and precise admixture proportions can be found in the spreadsheet. Note that additional samples have been used to infer these components, but as these come from Dodecad populations with less than 5 participants, I am not reporting average values for them, as per the usual project policy.

Here is the neighbor-joining tree based on the Fst divergences between the 7 ancestral components:
Instructions:

You can download the calculator RAR from here (Google docs; File->Download original), or here (sendspace).

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'euro7' instead of 'dv3' in these instructions.

Terms of use: 'euro7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Calculators released by the Dodecad Project:

Sunday, September 25, 2011

Yunusbayev et al. (2011) data assessed with Dodecad v3

I have acquired the data from the recent Yunusbayev et al. (2011) paper on the Caucasus. This includes the following populations:
  • Kurds_Y 6
  • Bulgarians_Y 13
  • Ukranians_Y 20
  • Mordovians_Y 15
  • Armenians_Y 16
  • Abhkasians_Y 20
  • Balkars_Y 19
  • North_Ossetians_Y 15
  • Chechens_Y 20
  • Nogais_Y 16
  • Kumyks_Y 14
  • Turkmens_Y 15
  • Tajiks_Y 15
It is a valuable new addition to the Project, and it is commendable that it has been made publicly and easily available so swiftly after the appearance of the Yunusbayev et al. (2011) paper.

To get the ball rolling on the new Yunusbayev et al. data, I will map the new populations onto the Dodecad v3 components; they will be added to the Dodecad v3 spreadsheet as they are calculated.

I have been laboriously designing a new global (including Amerindians and Australasians) Dodecad X1 experimental calculator with 3,010 individuals for a few weeks now, but I guess I will now have to reboot it with 3,214.

Together with some other new data I recently discovered, I now have 9,799 individuals (some duplicates from different sources) in my global database. My Dodecad dataset of 511 individuals from a single country or ethnic group isn't too shabby either. Let's hope for a new data release that will push the data collection above the magic 10,000.

UDDATE:

I have added the first 7 populations to the spreadsheet; the others are being calculated as we speak. Most of them seem in line with expectations, but the Abkhasian sample has one outlier individual (abh27), and has thus been placed in the "Outliers" tab of the spreadsheet; a new set of admixture proportions, minus that outlier individual, will be calculated anew:

UPDATE II: The population portraits have been uploaded to Google Docs as a rar file (Sendspace mirror). Average admixture results have all been entered to the spreadsheet.

Friday, April 8, 2011

Structure in West Asian Indo-European groups (part 2)

I will occasionally revisit old posts such as Structure in West Asian Indo-European groups to take advantage of new population samples from project submitters. This time around, I included our first Kurdish and Azeri participants, and limited the analysis to populations for which I had large numbers of markers (the final total is a ~132k pruned set of markers).

I also included Greeks, Caucasian populations (Georgians and Lezgins), and Levantine Arabs (Syrians-Lebanese) who frame this region from the West, North, and South respectively, as well as Assyrians who are interspersed in West Asia as an ethno-religious minority.

Here are dimensions 1 and 2 of the multidimensional scaling plot:

The two Caucasian groups (South and Northeast Caucasian Georgians and Lezgins respectively) form their own clusters. So do the Iranians and the Syro-Lebanese.

(As always population labels are placed on population averages, and _D denoted Dodecad Project populations)

Curiously, many linguists assert a close relationship of Greek and Armenian within the Indo-European language family. Turks speak an Altaic language due to migration of a numerically small population element, but their pre-Turkish genetic ancestors were Anatolian speakers, Greeks, Armenians, and Iranians, i.e., primarily Indo-Europeans.

Dimensions 1 and 3:

Dimension 3 contrasts Greeks from West Asian groups. Notice also the presence of 3 Armenians at the bottom of the plot, these are outliers of the Behar et al. Armenian sample.

Notice that the Azeri_D sample clustered with Iranians in dimension 2 and with Turks in dimension 3. This is not very surprising, as Azeris speak a Turkic language, but also have clear Iranian antescendants. The Kurd_D sample, on the other hand, clusters consistently with Iranians.

The variability of the Greek_D population sample along dimension 3 is also quite interesting. This could reflect variable levels of influence of extra-Greek European/Anatolian population elements on the basic Greek stock. Greeks who possess 23andMe or FTDNA Illumina population data are strongly encouraged to join the Project to help us better determine regional variation within the ethnic Greek population.

The Clusters Galore analysis results are as follows (9 clusters/3 MDS dimensions retained) :

In brief, the modal populations in each cluster are:
  1. Greeks
  2. Iranians
  3. Turks
  4. Syrians and Lebanese
  5. Armenians and Georgians
  6. Armenians and Assyrians
  7. Syrians and Iranians
  8. Lezgins
  9. Georgians
I will be happy to provide to all Dodecad Project members from the _D populations with their individual results. If you send me e-mail at dodecad@gmail.com, I will send you a line with your probabilities of assignment in each of the 9 clusters, as well as the 3 co-ordinates in the first 3 MDS dimensions plotted above.

I strongly encourage individuals from West Asia, the Balkans, and Italy to contact me for inclusion in the Project (send e-mail first, not data). While submission to the Project is currently closed, I usually accept data from these regions.

Sunday, December 19, 2010

Fine-scale admixture in Europe (Dagestan/Basque/Sardinian components)

Wanting to see whether the Dagestan mystery would extend into Europe, I carried out an ADMIXTURE analysis including all my European populations. Once again, as this is done on only ~30k markers, a little noise on the low-level components is expected.

Admixture proportions can be found in the spreadsheet

Notably there is now both a Sardinian and a Basque centered cluster; the latter was formerly (in the standard K=10 analysis) split between "Southern European" and "Northern European". The Urkarah, Lezgin, and Stalskoe samples show the highest presence of the "blue" component, which I label, once again, Dagestan. Note, however, that you should not compare admixture proportions across ADMIXTURE runs for components that happen to be labeled the same (homonymous). Certainly "this" Dagestan is related to the "previous" Dagestan component, but do not assume they are identical.

Here is the Fst distance matrix between the 7 components:


Discussion

The most notable thing about this figure is the relative absence of the West Asian component in the periphery of Europe. The lowest values are seen in Basque, Sardinian, Orcadian, White Utahns, Lithuanians, Finns, and Scandinavians (in no particular order).

It is worthwhile to order the European populations in terms of their Dagestan component. Excluding the populations of the Caucasus, these are, in ascending order: Basque (0.7%), Sardinian, Cypriot, Belorussian, South Italian/Sicilian, Lithuanian, Tuscans, Portuguese, Greek (3.8%), Vologda Russian, Romanian, Finnish, Spaniards, North Italian, Dodecad Spaniards, Dodecad Russian, Chuvash, Hungarian, French (7.9%), German, Scandinavian, White Utahn, Orcadian (12.6%).

Interpreting this pattern is not easy, but it does seem that this component seems to have a V-like distribution, achieving its maximum in Caucasus and its environs, then undergoing a diminution, and achieving a secondary (lower) frequency mode in NW Europe.

The surprising appearance of the homonymous Dagestan component in India suggests a widespread presence of a common ancestry element. The West Asian element, by comparison seems to have a more normal /\-like distribution around its center in Anatolia-Caucasus-Iran region. It does reach the Atlantic coast, but is lacking in Scandinavia and Finland, and also in India itself.

This is just a piece of a broader puzzle, and the picture is not yet clear. However, we can tentatively say that whatever brought the "Dagestan" component to India was not a unidirectional process, but also brought a similar population element to western Europe.

Tuesday, November 9, 2010

Multidimensional scaling in Italy, the Balkans, Anatolia, and the Caucasus + Lezgin ADMIXTURE surprise

On the left you can see an MDS plot of several population groups from Italy, the Balkans, Anatolia, and the Caucasus. This combines data of Dodecad Project members with published samples. I have placed the labels manually over the main point blobs. There is also a single Bulgarian sample who falls between the 'm' and the 'a' in 'Romanians'.

I had previously studied the distinctiveness of Caucasus populations, and now I have added Turks, Cypriots and populations from further West. I am still not satisfied with my Balkan samples (I have 2 Slovenians, 2 Serbs and 1 Bulgarian), so I encourage Balkan participants to contact me for possible inclusion in the Project.

When I turned to ADMIXTURE, a little mystery emerged, for which I have currently no explanation:

Two main components emerged, a light blue "Italo-Balkan" one that seems deficient in West Asia, and red "Cypriot" one that is deficient in West Balkan Slavs and the Caucasus. The three Caucasus populations, each form their own distinctive cluster (green, yellow, blue), and a magenta low-frequency component emerges at K=6, which is why I stopped the analysis at this K. Results for K=5 were similar, minus this low-frequency component.

Here is the big puzzle: my Bulgarian, 2 Serbs, 2 Slovenians, all show unambiguous membership in the green "Lezgin" cluster. Out of all the Caucasus components, this is the only one that seems to have a Balkan connection. While one could argue that this might reflect Neolithic farmers, as it has been argued that they spoke a North Caucasian language, the same "Lezgin" component is insignificant in Greeks and Mixed-Greeks, Southern Italians/Sicilians and Italian (other).

Is this some signal of a population that once inhabited the northern arc of the Black sea, from the Balkans to the Caucasus? This might find some support in the possession by both Lezgins (and Balkan Slavs) of a "North European" component, but the Adygei, who similarly possess such a component show no special affinity with Balkan Slavs. Below is the Lezgin K=10 portrait:
If anyone has any (pre)-historical scenario that might account for this unexpected affinity, feel free to write to me or leave a comment.

Friday, November 5, 2010

Analysis of Armenians, Lezgins, Georgians, and Adygei

I have taken Behar et al. (2010) Armenians, Georgians, and Lezgins, and HGDP-CEPH Adygei, together with the three Dodecad Project members scoring highest in the "West Asian" component in order to study the structure of Caucasus populations.


The first two dimensions of the MDS plot can be seen on the left. The color coding is as follows: DOD155: green; DOD011: red; DOD049: orange; Adygei: grey; Lezgins: magenta; Georgians: black; Armenians: blue.

It is fairly obvious that Adygei, speakers of a NW Caucasian language, can be distinguished from Lezgins, speakers of a NE Caucasian language, and from Georgians, speakers of a S Caucasian language. Armenians can be distinguished from Georgians, although they appear to be closer to them. The three project members all appear to be in the broad Georgian-Armenian cluster.


Turning to ADMIXTURE, analysis with K=3 confirmed the existence of the afore-mentioned clusters, with Adygei (dark blue), Lezgins (green), and Georgians/Armenians (light blue) forming their own distinctive clusters.

Going one step further at K=4, convergence became more difficult, but, nonetheless, the clusters produced by ADMIXTURE were also informative. Below, you can see the individual-level ADMIXTURE results:
While much "noisier", perhaps due to both the presence of actual admixture, and the difficulty of inferring structure in relatively closely related populations, four clusters emerge: light blue (Adygei), dark green (Lezgins), dark blue (Georgians), and light green (Armenians). The clusters become more apparent if we look at averages:
It is fascinating that four different populations, representative of four different language families (Indo-European and S/NE/NW Caucasian) can be distinguished in the Caucasus. With the exception of Armenians, none of the other populations have linguistic cousins in the world-at-large, although it has been speculated that the extent of Caucasian languages was once much broader than it is today, and some have even suggested that the Neolithic proto-farmers who dispersed into Europe were bearers of such languages.

It will be extremely interesting to study these clusters in other individuals from the Caucasus and Anatolia, as well as to see what is their relationship to the "West Asian" component which appears to be widespread in Eurasia.