Showing posts with label Results. Show all posts
Showing posts with label Results. Show all posts

Sunday, August 12, 2012

fastIBD analysis of Africans and African Americans

Individuals from the following populations have been included in this analysis:
African_American_D Somali_D Moroccan_D Algerian_D North_African_Jews_D Tunisian_D East_African_Various_D Yoruba_D Sudan_D Egyptian_D Chad_D
These were analyzed in the context of a large set of African populations. CEU European Americans were also added to account for the European admixture present in some African American individuals.
This is the first time I have included African American Dodecad participants in this type of analysis.

A few quick points:
  • fastIBD was run with default parameters over a dataset of 679 individuals/255020 SNPs
  • fastIBD identifies segments of relatively recent origin that are shared by individuals. These results should not be construed as measures of overall genetic similarity or origins. Rather, they suggest which populations have exchanged genes in the relative recent past.
With that said, you can get:
  • Spreadsheet of numeric results, showing median sharing (in centi-Morgans, cM)
  • Population-level graphical results, showing an ordering of other populations based on median IBD sharing.

IBD sharing was assessed only for populations with 5+ individuals.

The following heat map allows for a quick appraisal of populations sharing an excess of IBD sharing (read row-by-row). The grouping of populations by language group and/or region is clearly manifested. There are some interesting details that jump off the screen (but do consult the spreadsheet for details). For example, notice that: 
  • within the Bantu group (Bantu_NE, LWK/Luhya, and Bantu_S), only the South Bantu have an excess of IBD sharing with San.
  • Of the North Africans, Egyptans show an excess of IBD sharing with Tigray
  • Notice that of the Ethiopians/East Africans it is the Omotic speaking Wolayta that seem to especially share IBD with the Ari people who are also Ethiopian Omotic speakers.




Some visualizations (see graphical results above for full set):

Mozabites showing a high degree of within-population IBD sharing, and secondarily with other NW African groups.

The Dodecad Project Somali sample shows high degree of sharing within itself and also with the Pagani et al. Somali and Ethiopian Somali samples, and then with various other East African groups.
Sources of data are listed at the bottom left of this blog.

Friday, August 10, 2012

fastIBD analysis of East/Central Eurasians and select West Eurasians


Individuals from the following populations have been included in this analysis:
Philippines_D Turkish_D Iranian_D Russian_D Finnish_D Turkish_Cypriot_D Ukrainian_D Belorussian_D Chinese_D Korean_D Japanese_D Tatar_Various_D Kazakh_D Szekler_D Hungarian_D Estonian_D Azeri_D Udmurt_D Mixed_Turkic_D 
These were analyzed in a context of a complete set of Central/East Eurasian populations; West Eurasian populations included were mostly Uralic and Turkic speaking groups, and a few others (such as East Slavs or Iranians).

A few quick points:
  • fastIBD was run with default parameters over a dataset of 627 individuals/255020 SNPs
  • fastIBD identifies segments of relatively recent origin that are shared by individuals. These results should not be construed as measures of overall genetic similarity or origins. Rather, they suggest which populations have exchanged genes in the relative recent past.
With that said, you can get:
  • Spreadsheet of numeric results, showing sharing (in centi-Morgans, cM)
  • Population-level graphical results, showing an ordering of other populations based on mean IBD sharing.
IBD sharing was assessed only for populations with 5+ individuals.

The following heat map allows for a quick appraisal of populations sharing an excess of IBD sharing (read row-by-row)

And, a few visualizations of mean IBD sharing:

Notice high levels of within-population IBD sharing for Finns, consistent with a population that experienced expansion from a small number of founders (small ancestral population size).
Compare with Turks, who are a much more diverse population.
These two plots (you can check the spreadsheet for exact numbers) indicate different sources for the East Eurasian element in Turks and Finns. 

The top eastern populations for Turks are: Turkmen, Chuvash, Uzbek, Uygur, all of which are Turkic speakers, followed by Hazara, Yukagir, and Selkup.  For Finns, there is high degree of sharing with various Siberian groups of different languages, including Uralic Selkups (16.4cM) and Nganassan (9.6cM). Turks share less with these Uralic speakers (6.4 and 2.8cM respectively). So, these are strong hints of common shared ancestry within the Turkic and Uralic language families.

The Chuvash population is also quite interesting, as it shares more with Selkup and Nganassan, contrasting with other Turkic speakers. This makes excellent sense, and is in agreement with other recent findings:
Results from this study maintain that the Chuvash are not related to Altaic or Mongolian populations along their maternal line, thus supporting the “Elite” hypothesis that their language was imposed by a conquering group —leaving Chuvash mtDNA largely of Eurasian origin. Their maternal markers appear to most closely resemble Finno-Ugric speakers rather than Turkic speakers.
Sources of data are listed at the bottom left of this blog.

Tuesday, June 12, 2012

'K10a' calculator

The 'K10a' calculator represents an intermediate stage between the K7 and K12 analyses released so far from the Project. The following components have been inferred:
  • Palaeoafrican 
  • South_Asian 
  • West_Asian 
  • Southeast_Asian 
  • Sub_Saharan 
  • Atlantic_Baltic
  • Red_Sea 
  • East_Asian 
  • Mediterranean 
  • Siberian 
There are a couple of points of interest; first, the Red_Sea component related Arabians with East Africans. At a higher level of resolution the "Southwest_Asian" and "East_African" (K12)  components emerge. The "Red_Sea" component is not very closely related to any other components, but is somewhat related to the "Mediterranean" and "Atlantic_Baltic" components.

So, using the different calculators of the Dodecad Project, we first have (K7) a contrast between Africa and West Eurasia, then a signal of the shared ancestry between Arabia and East Africa (K10), and finally, strong signals of local ancestry in the two regions.

Second, the Mediterranean component here is modal in Sardinians as usual, but also projects into North Africa. Again, this is intermediate between K7 which shows a predominance of West Eurasian ancestry in North Africa + an African component, and K12 in which there are "Atlantic_Med" and "Northwest_Afican" regional components.

These are strong hints that the West Eurasian element in Africa differs between NW and E Africa. In the former region, it is most related to Sardinians, and in the latter it is most related to Arabians. Of course, ultimately the two elements are related to each other.

Table of Fst distances between components:


MDS plots of the first few dimensions:


Downloads: 
Project participants can find their results in the spreadsheet. Non-participants can use DIYDodecad to calculate their results, but they should place all the calculator files in the same directory as the DIYDodecad software, and replace 'dv3' with 'K10a' in all the instructions of the README file.

Component labels are indicative, and you should compare your results against the normalized median results for different populations included in the spreadsheet.

Terms of Use

You are free to use 'K10a', including all downloaded files for any non-commercial purpose, as long as you attribute them to the Dodecad Project and to Dienekes Pontikos as follows:

The 'K10a' admixture calculator is courtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project; more information here.

Saturday, June 9, 2012

'weac2' calculator

I have made a new version of the 'weac' calculator (West Eurasian cline). This is based on a large Old World dataset at K=7 and includes the following ancestral components:
  • Palaeoafrican 
  • Atlantic_Baltic 
  • Northeast_Asian 
  • Near_East 
  • Sub_Saharan 
  • South_Asian 
  • Southeast_Asian 

The West Eurasian cline is formed between the Near_East and Atlantic_Baltic components.

Here is the table of Fst distances between components:

MDS plots of the first few dimensions:


Downloads:
Project participants can find their results in the spreadsheet. Non-participants can use DIYDodecad to calculate their results, but they should place all the calculator files in the same directory as the DIYDodecad software, and replace 'dv3' with 'weac2' in all the instructions of the README file.

(NOTE: Some  IDs may have wrong results in the spreadsheet because of a misalignment of IDs with results; I'll fix this and update this notice. UPDATE: Results should be correct in spreadsheet now - 9 Jun 2012)


Component labels are indicative, and you should compare your results against the normalized median results for different populations included in the spreadsheet.

Terms of Use

You are free to use 'weac2', including all downloaded files for any non-commercial purpose, as long as you attribute them to the Dodecad Project and to Dienekes Pontikos as follows:

The 'weac2' admixture calculator is courtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project; more information here.

Friday, April 27, 2012

Estimating your Gök4-related ancestry


I have taken Table S15 from Skoglund et al. (2012), and the Dodecad Project K7b admixture proportions in order to investigate possible relationships.

In Table S15 the authors estimate the Neolithic farmer ancestry in several populations on the basis of a single Neolithic individual from the Funnel Beaker (TRB) culture which was found in a megalithic burial in Gökhem parish.

Most of these populations are already part of the Dodecad Ancestry Project, except the three Swedish samples; given the intermediacy of the Central_Sweden sample, I have decided to use my Swedish_D sample of Project participants as a stand-in for it.

Below, you can see a scatterplot relating Gök4-related ancestry with K7b "Southern" component:



The correlation between the two variables is very strong (R-squared = 0.93).

Dodecad Project participants who already have K7b results, as well as customers of DTC testing companies who can use DIYDodecad together with the K7b calculator can approximately estimate their Gök4-related ancestry by plugging in their "Southern" value (in %) into the following equation:

Gök4-related ancestry = 1.721*Southern+19.736

I anticipate that when I am able to study the Neolithic Swedish genomes directly, the Neolithic farmer from Sweden will turn up "Southern" in a K=7 resolution experiment.

Sunday, March 11, 2012

ChromoPainter/fineSTRUCTURE analysis of Italy/Balkans/Anatolia

This was done on the same dataset as the previous fastIBD analysis.

The population assignments:



The heatmap, showing relationship between inferred populations:


The principal components analysis:



The correspondence between inferred populations and K12b components:



Results for Project participants can be found in this spreadsheet; remember than in the chunkcounts tabs, columns represent donor and rows recipient populations.

Monday, March 5, 2012

fastIBD analysis of Italy/Balkans/Anatolia

I have included the new Turkish data from Hodoğlugil & Mahley (2012) in this analysis. Additionally, there are now 5 participants in the Serb_D and Turkish_Cypriot_D sub-populations, as well as a Bosnian Muslim. There are now project participants from many Balkan countries, although Albania, the fYROM, and Croatia remain as "black holes" in the map.

Still, I am hopeful that there will be more project participants from currently under-represented populations. I have already started processing the same dataset with ChromoPainter (which takes much longer), and hopefully that analysis will be posted at the end of this week or the beginning of the next one.

First, the heatmap of inter-population IBD:

Remember that the tree groups similar populations together, and for each row in the matrix, the red end of the spectrum indicates lots of IBD sharing, and the blue end low IBD sharing. Additionally, I have now calculated the median IBD sharing, which is more resistant in the presence of potential relatives in the data.

The results appear fairly reasonable, with the Balkan, Anatolian, and Italian populations of the title forming separate branches, and the mainland Greek sample joining with Central/South Italians and Sicilians.

The Clusters Galore can be seen below; 28 clusters were inferred with 21 dimensions:



Results for Project participants can be found in the spreadsheet, and include the probabilities that each ID is assigned to each of the 28 clusters, as well as the Z-scores comparing each individual against all populations with 5+ individuals. The Z-score should be read as follows: for each row, high values indicate a high degree of IBD sharing, while low values indicate a low degree of IBD sharing.

Of course, I encourage Project participants to leave a message in the Information about Project samples thread.

Tuesday, February 14, 2012

ChromoPainter/fineSTRUCTURE analysis of Balkans/West Asia

I have carried out a ChromoPainter/fineSTRUCTURE analysis of Balkans/West Asia. This is a slightly different dataset than the one used in the previous fastIBD analysis of the same region. It also took much longer (about a week, with two CPUs dedicated to the task) to complete, so it is not something that can be done routinely.

Technical details (skip if you want)


413 individuals from 33 populations were studied, on 258,100 SNPs, after --geno 0.03 --maf 0.01 filters were applied. Data were phased in Beagle with the default 10 iterations. Genetic maps from the HapMap were used. fineSTRUCTURE was used on ChromoPainter output, with 500,000 burnin/runtime iterations each.

25 Inferred Populations


fineSTRUCTURE imposes a tree structure on a number of inferred populations. The following heatmap shows this tree structure; columns represent donor populations, rows, recipient ones.


There was a total of 25 populations, labeled pop0, pop1, ..., pop24.

The following table summarizes how many individuals from each original population were assigned to each inferred population:


I will limit myself to populations which include Dodecad Project members:

  • pop6 includes a Project North Ossetian, as well as all Yunusbayev et al. North Ossetians
  • pop7 is mainly Armenian
  • pop16 is also mainly Armenian; it would be interesting to see whether this bipartite division of Armenians is in agreement with the one inferred in the previous fastIBD analysis
  • pop8 is mainly Greek, and appears to be "continental Greek"; it also includes some other Balkan individuals
  • pop14 is also Greek, and includes a variety of people with ancestry from Crete, the Aegean, Cyprus, Asia Minor, Cappadocia, and the Pontus as well as continental Greek. It could be labeled "eastern Greek"
  • pop11 is Cypriot, including the single 100% Greek Cypriot of the Project, all 3 100% Turkish Cypriots, as well as a Turkish individual of partial Turkish_Cypriot ancestry
  • pop10 is Turkish, and includes people with some ancestry from the Balkans, as well as Anatolia. It could be labelled "Balkan Turkish"
  • pop13 is also Turkish, and seems to include people with ancestry exclusively from Anatolia, including almost all the Behar et al. Turks
  • pop15 is Assyrian; some Assyrians also fall on the aforementioned pop16 which includes mainly Armenians
  • pop18 could be labelled "North Balkan"; there is probably structure to be uncovered within this cluster, once more participants from the Balkans join the Project
  • pop20 is "Georgian-Abkhazian"
  • pop21 is "Kurdish-Iranian"
  • pop22 could be labeled "Northeastern Anatolia" or (more classically) "Pontus-Colchis". It appears to unite various individuals from Northeastern Turkey and neighboring Georgia, having Karadeniz Turkish, Armenian, Pontic Greek, and Kartvelian ancestry. I strongly encourage participants from this region to join the Project, especially Pontic Greeks, as there are no 100% Pontic Greeks currently in the Project.
  • pop23 is "Bulgarian-Romanian" mainly, and also includes one Serb. Once again, I emphasize that the power of this approach using haplotypes depends on participation, so I encourage all people from the Balkans to consider joining the Project.
Principal Components Analysis


I have also used the PCA feature of fineSTRUCTURE to carry out principal components analysis. I am plotting the first two dimensions of this PCA, using my own visualization code that places labels in the average position on the plane:


Results


Results for Project participants are included in the spreadsheet.

  • Population matrix, shows how many individuals from each population were assigned to each cluster
  • Z score population matrix, shows the normalized number of "chunks" from each donor population (columns) to each recipient (row). Do not compare across rows! The way to read this table is the following: for each row, higher values indicate more sharing. For example, the "Cypriots" population has pop11 as its main donor.
  • Individual assignments: the pop number that all Project and reference IDs were assigned to
  • Individual Chunkcounts: the number of chunks copied from its donor population (column) to each individual
  • Individual PCA: your PCA co-ordinates that can help you find your dot on the Principal Components Analysis graphic (see above)
Averaged results were included only for populations with >=5 members.
The raw chunkcounts for all 413x413 individuals can be found here.

Tuesday, January 31, 2012

'K12b' and 'K7b' calculators

I am releasing two new calculators with K=12 and K=7 components, named 'K12b' and 'K7b'. You can scroll down to the bottom if you are just interested in the downloads, or read on.


New Features

The new 'K12b' calculator is an update of the previous K12a one, that was inferred using all the new samples submitted during the last submission opportunity. The 12 components are still roughly the same, although their allele frequencies may have changed by a bit, so existing participants can expect to have slightly altered results, and new participants in the Project more so, since their data are now contributing to the creation of the new tool. Non-participants can, of course, use the new calculator with DIYDodecad.

I have also taken the opportunity to do some minor tweaks. I am releasing population portraits for K12b (which were lacking in K12a); I've changed my visualization code so that the sample IDs of non-Dodecad populations can now be seen in the barplots. This may be useful for anyone else using these reference populations, by quickly identifying potential outliers in them.

I have also decided to use normalized median admixture proportions for the populations. For example, if 5 individuals in a population have 0, 0, 0.2, 0.5, 10.0% of a particular component, then the average is 2.14%, but the median is 0.2%. By using the median, the proportions become less susceptible to the presence of outliers (such as the 10%). However, if the median is calculated over every component separately, it is no longer guaranteed that the components will add up to 100%; this can be addressed by re-normalizing them (scaling them by a constant factor) so that they do. I believe that use of the normalized median will not only give better proportions that are less susceptible to outliers, but will also improve results of the new Dodecad Oracle for K12b.

At the same time I am also releasing 'K7b' which is an update of the existing 'eurasia7' calculator and which has been built on exactly the same dataset as 'K12b' but at a lower (K=7) level of detail.

Information on K7b


Information spreadsheet.

Normalized median admixture proportions barplot for all included populations (a high resolution version of this is included in the download bundle):


Table of Fst divergences:

Neighbor-joining tree (based on above):

Information on K12b


Information spreadsheet.


Normalized median admixture proportions barplot for all included populations (a high resolution version of this is included in the download bundle):

Table of Fst divergences:

Neighbor-joining tree (based on above):
Multidimensional Scaling Plots of K12b and K7b


I have created MDS plots using synthetic individuals representing the 12 ancestral components of K12b and the 7 ancestral components of K7b. By including both in the same plot, one gets an idea of the relationship of the components at different resolution. The first 10 dimensions can be seen below:

Here is a blowup of the main West Eurasian groups from the plot of the first two dimensions:

Some observations:

  • The Atlantic_Med component which is bi-modal in Basques and Sardinians occupies the apex of the figure; this makes sense, since Southwest Europe is quite distant (along land routes) to both Asia and Africa.
  • The Caucasus component is surrounded by most of the others; this is consistent with my theory elaborated in The womb of nations: how West Eurasians came to be.
  • The Atlantic_Baltic component (from K=7) is intermediate between the Atlantic_Med and North_European components.
  • Similarly, the West_Asian component (from K=7) is intermediate between the Caucasus and Gedrosia components; the Gedrosia component diverges in the direction of the Asian groups (not shown in this figure), and in particular of South Asians. This divergence can also be seen in the plot of dimension #3.
  • The Northwest_African component diverges in the direction of Sub-Saharan Africans.

Technical Details


A dataset of 268 populations/3,115 individuals was assembled. A total of 265,519 SNPs are in common in the various source datasets as well as the 23andMe v2/v3 and Family Finder platforms. Iterative removal of distant relatives was performed by removing one individual from each pair within a population if that pair had a RATIO of 2.5 or greater or more than the mean and two standard deviations in IBD analysis performed in PLINK 1.07. A total of 2,675 individuals remained. 4 individuals were removed for low genotyping rate (less than 97%). 264,328 SNPs remained after removal of SNPs with less than 97% genotyping rate or 1% minor allele frequency. 166,770 SNPs remained after linkage-based disequilibrium pruning (--indep-pairwise 200 25 0.4). The final set thus consisted of 2,671 individuals/268 populations/166,770 SNPs. Ancestral populations (components) were inferred using ADMIXTURE 1.21, with K=7 and K=12 and default parameters.

No individuals were removed from the source datasets, except in the case of the Armenians_Y sample, where one individual (ID: armenia3) was dropped because he/she was the same as a Dodecad Project participant.

Downloads


K7b population portraits, spreadsheet, and DIYDodecad files.
K12b population portraits, spreadsheet, and DIYDodecad files.

Dodecad Oracle (K12b edition) can be downloaded from here. Please read the instructions of the previous Oracle on how to use this tool. Note that the number of populations is now 223.

To use either calculator with DIYDodecad, with your 23andMe or Family Finder data, follow the instructions in the README file, but substitute 'K12b' or 'K7b' for 'dv3'.

Project participant results for both K7b and K12b are found in the spreadsheets in the Individual Results tab.

Terms of Use


You are free to use K12b and K7b, including all downloaded files for any non-commercial purpose, as long as you attribute them to the Dodecad Project and to Dienekes Pontikos as follows:

The [K7b/K12b] admixture calculator is courtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project; more information here.

Saturday, January 21, 2012

fastIBD analysis of Afroasiatic groups (Jews, Arabs, Assyrians, Berbers, Somalis, Amharas, etc.)

Please refer to the previous analysis on the Balkans/West Asia for more information about the interpretation of this type of analysis.

I am very pleased with the way this analysis of Afroasiatic groups has turned out, revealing an exceptional degree of resolution. I invite individuals from the Near East and Africa who are eligible, to submit their data, so that they can be included in future runs of this kind.

Clusters Galore


45 clusters were inferred with 29 dimensions.


I can't comment on all 45 clusters, so I'll just limit myself to the ones that are significantly represented among Project participants: 1. Ashkenazi, 4. Assyrian/Mandaean, 6. Somali, 7. Moroccan, 8. Algerian/Tunisian, 9. Sephardic, 10. Morocco Jews, 11. Iran/Iraq Jews, 12. Non-Jewish Ethiopians, 13. Saudi, 14. Arab #1, 15. Arab #2, 16. Egyptian

Inter-Population IBD


Results for Project Participants


The results can be found in the spreadsheet.

I have also added the full IBD sharing matrix which lists how many Morgans of sequence are estimated to be IBD with probability greater than 10^-6 between all pairs of individuals.

You can google any non-Project sample IDs to get some more information about their origin. For example, GSM536710 is an Iraqi Jew who shares about half his genome with GSM536714, also an Iraqi Jew. These two samples are almost certainly first-degree relatives. Or, GSM537032, a Samaritan shares 740-1,480cM with the other 2 Samaritans, an exceptional amount in this small and probably highly inbred population.

You can manipulate this matrix in R. After you download it and unzip it, you can load it into R as follows:

X<-read.table('afroasiatic_ibd_sharing.txt',row.names=1,header=T)

Then, you can, for example, sort the IBD sharing for a particular individual, as follows:

sort(X['DOD026',])

fastIBD analysis of Central/Eastern Europe

Please refer to the previous analysis on the Balkans/West Asia for more information about the interpretation of this type of analysis.

Clusters Galore


The Clusters Galore can be found in the spreadsheet. After inspection of the 23 clusters inferred with 21 dimensions, they could be described as:

  1. Mordvin
  2. East Slavic
  3. Polish-Ukrainian
  4. East Balkan
  5. Vologda Russians
  6. Lithuanian
  7. Central European (combining many groups with small sample sizes)
  8. A couple of related (?) individuals
  9. Anatolian
  10. Greek
  11. Chuvash
  12. Ossetian
  13. A couple of related individuals
  14. A couple of related individuals
  15. Balkar
  16. A couple of related individuals
  17. Chechen
  18. Kumyk
  19. A couple of related individuals
  20. Adygei
  21. Lezgin #1 (main)
  22. Lezgin #2
  23. Lezgin #3
If you belong to a population with few other participants, you might end up latching onto a cluster dominated by a bigger group. This does not mean that your population is not distinctive, only that there are not enough samples to reveal its distinctiveness if it exists.

Inter-Population IBD


Results for Dodecad Participants

Results can be found in the spreadsheet.

If you have joined the Project, please consider leaving a comment in the Information about Project samples thread. That will help others make better sense of their results, e.g., if you find that you belong in the same cluster with some other individual, you might want to know something about their origins.

UPDATE: I have added the IBD sharing matrix.See here on how to use it.

Thursday, January 19, 2012

fastIBD analysis of South Asia

Please refer to the previous analysis on the Balkans/West Asia for more information about the interpretation of this type of analysis.

Clusters Galore


The Clusters Galore analysis can be found in the spreadsheet. 59 clusters were inferred with 47 MDS dimensions. The very fine-scale structure (I only considered the first 50 dimensions, but many more seemed significant than in any previous experiment) is probably the result of the size of the South Asian population, as well as the practice of endogamy associated with the caste system. High intra-population IBD sharing is also evident in the following (notice how well-defined the diagonal is):

Inter-Population IBD




Results for Dodecad participants

They can be found in the spreadsheet. Many Project participants belong to a population with 1 or 2 individuals, so cluster #1 seems to be a generalized catch-all for many such individuals. Individuals from he two sub-populations that I've identified recently Iyer_D, and Jatt_D all belong to the same cluster. The Iyer_D cluster (#4) also seems to include the Iyengar project participants as might be expected.

It is also interesting how all Dodecad participants fall in just 7 of the 59 clusters. This goes to show how truly diverse people from the Indian subcontinent are. I fully expect that with more participation further structure will be revealed, since it seems that due to endogamy it only takes a few participants from each ethnic group for a specific cluster pertaining to that group to be identified. So, I invite people from South Asia to join the Project during this submission opportunity.

Tuesday, January 17, 2012

fastIBD analysis of Iberia, France, Italy, Balkans, Anatolia and European Jews

On the heels of the previous analysis of Balkans/West Asia, a new experiment on a different set of populations. Please refer to the earlier post for some thoughts/explanations about this type of analysis, I'll stick to "just the data" for this post.

Clusters Galore




24 clusters inferred with 17 MDS dimensions.

The Galore analysis provides increased resolution within Iberia (#6-9, 11), Italy, and the Ashkenazi Jewish group (#14-16).

The Iberian results are particularly interesting, showing the power of this approach compared to the one with unlinked data. There appear to be:

  • a Spanish Basque (#6), 
  • French Basque (#11) cluster, as well as 
  • a Portuguese/Galician/Castilla Y Leon (#9) cluster, and 
  • a complementary Castilla La Manch/Cantabria/Andalucia/Murcia (#7) cluster, and 
  • a smaller Aragon/Cataluna cluster (#8). 
There is overlap between these clusters, but the geographical contrasts are quite evident. I did not go through the results of Spanish Project participants (all the Portuguese fall in the Galician cluster, and our Basque member in the Basque cluster as expeccted), so it would be interesting to hear whether they fall in the cluster(s) which exist in their regions of origin.

Inter-Population IBD




Results for Project Participants


The results can be found in the spreadsheet.

Saturday, January 14, 2012

fastIBD analysis of Balkans/West Asia

Now that I've discovered a way to boost Clusters Galore analysis even further by using fastIBD, I will start experimenting with different regional populations. This analysis took about 5 hours to complete, so it appears to be quite practical.

For my first experiment, I carry out an analysis of various populations from the Balkans and West Asia.

Clusters Galore

27 different clusters were inferred with 17 MDS dimensions. Some interesting findings:
  • For the first time there emerge a couple of clusters that appear to be quite specific to Armenians (#2 and #3). 
  • Similarly, Assyrians are broken to a few clusters that appear fairly specific to them  (#9-11)
  • Georgians are split into three clusters, one of which (#14) is linked with the neighboring Abkhasians, who in turn have their own exclusive cluster (#25)
  • The cluster modal in Greeks (#6) includes 14 of 19 Greek participants, and a few Greeks are also in the Balkan cluster (#8) and an Iranian-Turkish cluster (#4)
  • The Behar Cypriot sample also splits into two, and the few Turkish Cypriot participants link to one of them (#13)
  • The Ossetian project participant links to one of the three North_Ossetian clusters
  • The major Balkan cluster (#8) still defies resolution. I am certain, however, that structure in this cluster will be uncovered with more participation. MCLUST adapts the cluster size and shape, and a "big", inclusive cluster spanning the Balkans appears more parsimonious than smaller clusters centered on the different groups. With larger participation, I anticipate that regional structure will be uncovered in the Balkans as well.
I cannot stress the importance of participation strongly enough. When groups have more participants, it is possible to both:

  1. Discover group-specific clusters, by identifying what is common between members of groups
  2. Discover within-group clusters, by identifying what is different between members of groups
For example, the great participation of Armenians in the Project has now allowed me to discover structure within the Armenian population. It appears, that cluster #2 corresponds to a more "western" Armenian group, and #3 to a more "eastern" one, with some overlap between the two.

Inter-population IBD


You can also see a visual representation of inter-population IBD:

I have only included populations with 5+ participants in this representation. Reddish shades express high IBD sharing; bluish ones low one. The heatmap has been scaled by row.

As you might expect, values across the diagonal are "reddish", since individuals within populations tend to have high IBD sharing with each other.

A few features "pop out" of the screen. Going from top to bottom:
  • Intra-Iranic sharing
  • Intra-Armenian sharing
  • Intra-Balkan sharing
  • Georgian-Abkhaz sharing
You can probably get more out of the figure, but these appear to be the most salient features.

Results for Project Participants


The results can be found in the spreadsheet, and include:
  • Probabilities of assignment in each of the 27 clusters of the Clusters Galore analysis
  • Z-scores of IBD between each individual and each of the 20 populations with 5+ participants. Higher values mean more IBD sharing. Note that Z-scores have been calculated for each row, hence each participant must scan his own row to find populations with an excess (+) or deficiency (-) of IBD sharing, and people should not compare across different rows.
Last but not least, I want to remind new project participants to leave a message in the Information about Project samples thread. Your comment will not appear immediately, since comment moderation is on, and also note that there are multiple pages of comments. 


If you haven't joined the Project yet, I encourage you to do so if you are eligible.

Wednesday, January 11, 2012

Clusters Galore (fastIBD edition) for some northern European participants

You can find some new Clusters Galore results here (scroll down for spreadsheet link). The new methodology described in that post has made it possible to infer even finer-level population structure than "classical" Clusters Galore.

Sunday, December 11, 2011

Participant results for 'K12a' calculator

The participant results can be found in the "Individual Results" tab of the K12a spreadsheet.
You can read more about the K12a calculator at my other blog; if you are not a Project participant, you can also find a DIY version of it there, which can be used in conjunction with DIYDodecad 2.1.

Wednesday, October 26, 2011

'eurasia7' calculator

This calculator was made with 196 different populations and 2,659 individuals, including 518 project participants. The following Dodecad populations do not have 5 individuals yet, so they are included in the OTHERS_D generic category:
Algerian_D, North_African_Jews_D, Slovenian_D, Mixed_Scandinavian_D, Danish_D, Moroccan_D, Tunisian_D, Serb_D, Austrian_D, Saudi_D, Pakistani_D, Tatar_Various_D, Palestinian_D, Greek_Italian_D, Romanian_D, Swiss_German_D, Szekler_D, Mandaean_D, Azeri_D, Czech_D, Georgian_D, Belgian_D, Latvian_D, Estonian_D, Bangladesh_D, Yemenese_D, Sri_Lanka_D, Hungarian_D, Basque_D, Udmurt_D, Egyptian_D
As always, I encourage people with 4 grandparents from the same country or ethnic group of Eurasia, North or East Africa to contact me (do not send data!) for possible inclusion in the Project. If I have overlooked any such individuals, drop me a line (my e-mail address is at the bottom of the blog). I usually start a new _D population whenever individuals with 4 grandparents from the same group are submitted, but I may have missed some.

Note that all individuals from the reference populations have also been included, including outliers; you should be aware of this when reading the population averages, and consult the Outliers tab in the v3 spreadsheet for some instances of outliers.
Due to image size restrictions in Picasa, the labels are not visible well. A large version of the above plot can be found in the download bundle.

The seven ancestral populations inferred at this level of resolution are:
  • Sub_Saharan
  • West_Asian
  • Atlantic_Baltic
  • East_Asian
  • Southern
  • South_Asian
  • Siberian
As usual, you should take these names as useful labels, and interpret them in conjunction with the components' distribution in different populations, and their Fst distances, both of which can be found in the spreadsheet.

The table of Fst distances:


Below you can see a neighbor-joining tree based on inter-population Fst distances:
The first six dimensions of a multi-dimensional scaling of the same:





Calculator Files:

  • The spreadsheet contains population averages, the table of Fst distances, and individual results for included Project participants.
  • The download RAR file (Google Docs or Sendspace) contains all the files needed to run the calculator. You must download and install DIYDodecad 2.1 first. In order to run the calculator, you follow the instructions of the README file, but type 'eurasia7' instead of 'dv3'.

Terms of use: 'eurasia7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Technical Details:

The calculator is built using allele frequencies of K=7 ancestral components inferred by ADMIXTURE 1.21 analysis of 2,659 individuals. Markers included in the source datasets, as well as the Family Finder and 23andMe (as of Oct 21) platforms were included. The marker set was thinned of markers with less than 99.5% genotype rate and less than 0.5% minor allele frequency. Linkage-disequilibrium based pruning was carried out with a window size of 250 SNPs, advanced by 25 SNPs and R-squared greater than 0.4. A total of 164,990 SNPs remained after these filtering steps.

All relevant populations available to me, and genotyped at a sufficient number of markers were included. Inclusion of the Kalash population resulted in a population-specific component at K=7, and hence their admixture components were inferred a posteriori. Their proportions are consistent with previous results, showing them to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components.