Showing posts with label Dodecad. Show all posts
Showing posts with label Dodecad. Show all posts

Tuesday, October 23, 2012

'globe10' calculator

As part of the on-going analysis of the world dataset, I am releasing the 'globe10' calculator, which is based on the K=10 analysis. This calculator includes the following ancestral components:
  • Amerindian
  • West_Asian
  • Australasian
  • Palaeo_African
  • Neo_African
  • Siberian
  • Southern
  • East_Asian
  • Atlantic_Baltic
  • South_Asian
The names may be the same as the ones from previous calculators released by the Project, but you should always consult the spreadsheet to see how they might differ. In this case, inclusion of Amerindian, Australasian populations, African hunter-gatherers, dealing with the Paniya issue, and inclusion of data of Schlebusch et al. (2012), and  Pagani et al. (2012), have all combined to change components in subtle ways, although their modalities remain largely unchanged, and hence so do the names.

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'globe10' instead of 'dv3' in these instructions. You can consult the spreadsheet for proportions of the 10 components in different world populations.

Terms of use: 'globe10', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Saturday, October 13, 2012

Geno 2.0 data request

If anyone has received results from the Geno 2.0 test of the Genographic Project and want to share it with me, feel free to send it at dodecad@gmail.com. I will not distribute it or share it with anyone. I want to see what SNPs are tested, what format the data is in, and what is its intersection with other available datasets. This way, I can update my DIYDodecad software so that Geno 2.0 testees can use the various calculators released by the project to get an alternative ancestry assessment.

In time, and if there is interest, I may release additional calculators that make use of the particular SNP set tested by Geno 2.0.

Saturday, August 11, 2012

On the so-called "Calculator Effect"

The genome blogger Polako recently announced a calculator effect (May 2012) affecting admixture estimates:
However, many people are getting skewed results, despite doing everything right. For instance, users from the UK often come out much more continental European than they should. Some of them actually believe that this is because they're genetically more Norman or Saxon than the average Brit. Nope, the real reason is what I call the "calculator effect". This is when the algorithm produces different results for people who are part of the original ADMIXTURE runs that set up the allele frequencies used by the calculators, than those who aren't, even though both sets of users are of exactly the same origin, and should expect basically identical results.
This, however, was described by myself many months prior, in Novemeber 2011, following up on observations made during my first analysis of Yunusbayev et al. Armenians in September 2011. It has been listed in the Technical Stuff at the bottom of this blog ever since.

I had observed at the time that the newly available Yunusbayev et al. Armenian sample appeared more "European" using the Dodecad v3 calculator tool, which had been built using the Project Armenians (Armenian_D) as well as the Armenian sample of Behar et al.

I then explained why this was happening, and released new versions of the Dodecad tools, such as K12a, and K12b, and more recently K10a as new scientific and project participant samples became available.

Polako also proposes a "solution" to the problem:
I actually designed my Eurogenes ancestry tests for Gedmatch with this problem in mind, by only using academic references to source the allele frequencies. This means that test results for Eurogenes project members and non-members are directly comparable. Perhaps other genome bloggers can eventually do the same?
The only effect of this "solution" is to ensure that there is a "calculator effect" for everyone using his tools. For example, if he uses only published Finns and Lithuanians to build his calculator, then every Finn and Lithuanian who takes his test will wonder why he is "different" from the published Finns and Lithuanians, because they will all suffer a "calculator effect" with respect to the reference populations. So, perhaps they will all be on equal footing with respect to each other, but their results will all be biased because of the issue I had identified.

Moreover, their results will never improve as more people join his Project, because these new people will not be included in newer versions of calculators: all users of DIY Eurogenes tools will continue to receive sub-par results. Well, small consolation, at least they'll all receive comparable sub-par results.

The solution to this problem was also described in my original post, and it's not an unimaginative quick fix of biasing everyone's results with respect to the reference populations:
What can we do to solve this problem? Sample, sample, sample. There is no shortcut. The gross details of the genetic landscape (such as the relationship between major continental groups) are easy to infer, but the details will always have room for improvement.
It is only by adequate sampling, that is by including more and more people, rather than excluding even the ones we have, that ever more accurate admixture estimators can be devised. As sample sizes grow (= more scientists publish their data, and more people join projects such as this one), allele frequencies of the different components will become ever more secure, and deviations of individuals who did not contribute to the inference of the genetic components will converge to zero.

I am already quite confident that inclusion biases amount to only a few percent for Dodecad Project tools and only for the closely related components (e.g., West Asian vs. North European); as mentioned in my original post, these biases are trivial for more distantly related components (e.g., European vs. East Asian).

And, the way to further reduce biases that do persist is to foster participation, rather than consign everyone to a sort of fossilized mediocrity, excluding whole populations of active direct-to-consumer customers (e.g., Norwegians, or Assyrians, or Iraqis, or Germans, or Koreans, or, ...) on the basis that no "academic reference" has made dense genotype data on them freely and publicly accessible.

Friday, April 27, 2012

Estimating your Gök4-related ancestry


I have taken Table S15 from Skoglund et al. (2012), and the Dodecad Project K7b admixture proportions in order to investigate possible relationships.

In Table S15 the authors estimate the Neolithic farmer ancestry in several populations on the basis of a single Neolithic individual from the Funnel Beaker (TRB) culture which was found in a megalithic burial in Gökhem parish.

Most of these populations are already part of the Dodecad Ancestry Project, except the three Swedish samples; given the intermediacy of the Central_Sweden sample, I have decided to use my Swedish_D sample of Project participants as a stand-in for it.

Below, you can see a scatterplot relating Gök4-related ancestry with K7b "Southern" component:



The correlation between the two variables is very strong (R-squared = 0.93).

Dodecad Project participants who already have K7b results, as well as customers of DTC testing companies who can use DIYDodecad together with the K7b calculator can approximately estimate their Gök4-related ancestry by plugging in their "Southern" value (in %) into the following equation:

Gök4-related ancestry = 1.721*Southern+19.736

I anticipate that when I am able to study the Neolithic Swedish genomes directly, the Neolithic farmer from Sweden will turn up "Southern" in a K=7 resolution experiment.

Monday, February 6, 2012

Other testing companies

The Dodecad Project is not affiliated with any genetic testing companies. Until now, I have included Project participants from 23andMe and FamilyTreeDNA "Family Finder" tests, but it has come to my attention that there are new players in the field, such as Ancestry.com (see post on Your Genetic Genealogist) and Lumigenix (see post on GenomesUnzipped).

If you have data from any company entering this field, please contact me at dodecad@gmail.com (do not send data right away!). That way, I can find out how many markers are in common between the new tests and my existing datasets, and figure out how easy it will be to convert them for use in the Project and in DIYDodecad.

Tuesday, January 3, 2012

Submission opportunity (January 2012)

Who is eligible


Anyone who:

  • has 23andMe or Family Finder autosomal data,
  • is not related to any other Project participants,
  • has 4 grandparents from the same African, European, or Asian ethnic group or country (e.g., 4 Albanian grandparents, 4 grandparents born in Ethiopia, 4 Kazakh grandparents, etc.)
Any ineligible submissions will be blacklisted. Do not send data if you do not meet the eligibility criteria.

What to send


Send your compressed autosomal data (ending .zip or .gz) that you can download from your testing company.
Send to dodecad@gmail.com as an attachment, and include in your e-mail as much information about your ancestry as you can (e.g., birthplace of grandparents, spoken languages, practiced religions, ethnic affiliation, etc.). Samples without adequate ancestral information will be ignored.

Data Privacy Statement


Your raw data or genealogical information will not be shared or distributed in any manner, and it will not be analyzed for any other purpose than assessment of ancestry (i.e., not for any physical or health-related traits). It will be identified by a unique ID, known to you and me, and results will be posted in the blog using that ID. I will continue to analyze your data for ancestry, and new results will be posted using that same ID. Also, I will report aggregate results for populations with at least 5 participants.

What you will receive


I will add you to the K12a spreadsheet of the K12a calculator. You will also be eligible to participate in future data analyses, and newer results will be posted in this blog with your ID.

Clarification (added on 6 Jan, 2012): The results which you will receive will be based on the K12a calculator whose components were inferred in December 2011, and hence included only those who had submitted their data up to that time. As new members of the Project, your data will be used for the development of the next version of the admixture analysis, and this will -in all likelihood- lead to a subtle redrawing of the ancestral components and different ancestral proportions (see technical note).

By participating in the Project, you help better draw both the basic ancestral components underlying genomic variation in Africa/Europe/Asia, and create more robust samples of different populations. This is helpful both to the Project, and to yourself, because it helps you get increasingly better results with newer versions of the analysis. All newer analysis tools are announced on this blog.

End of Submission Opportunity


The end of this submission opportunity will be announced on this blog.

Sunday, December 11, 2011

Dodecad Oracle (K12a edition)

I have created a new version of the Dodecad Oracle for use with the K12a calculator.

You can refer to the original Dodecad Oracle for detailed usage instructions.

(The only difference in the use of the program is that the number of populations is 204, so make sure to use this if you plan to remove any reference populations, as mentioned in the instructions)

In short:
  • you first load the file DodecadOracleK12a.RData in R. You can do this by double-clicking on this file in Windows, or using the File->Load Workspace menu. In Linux, you can use the "load" command, e.g., load('/home/ubuntu/Desktop/DodecadOracleK12a.RData')
  • You then enter commands at the command prompt
Some examples:

Comparing a population against other populations

DodecadOracle("Somali_D")
[,1] [,2]
[1,] "Somali_D" "0"
[2,] "Ethiopian_Jews" "12.3049"
[3,] "Ethiopians" "12.3309"
[4,] "Sandawe_He" "38.2093"
[5,] "MKK25" "40.7983"
[6,] "Egyptans" "63.2307"
[7,] "Yemenese" "69.1628"
[8,] "Moroccans" "72.6233"
[9,] "Jordanians" "73.1838"
[10,] "Palestinian" "74.2867"

Comparing a population against 2-way population mixes:

DodecadOracle("Pathan",mixedmode=T)
[,1] [,2]
[1,] "Pathan" "0"
[2,] "79.5% Sindhi + 20.5% Lezgins" "3.948"
[3,] "82% Sindhi + 18% Chechens_Y" "4.0251"
[4,] "16.7% Adygei + 83.3% Sindhi" "4.5471"
[5,] "83.4% Sindhi + 16.6% Balkars_Y" "4.6487"
[6,] "80.8% Sindhi + 19.2% Kumyks_Y" "4.7067"
[7,] "83.7% Sindhi + 16.3% North_Ossetians_Y" "4.8352"
[8,] "80.9% Sindhi + 19.1% Nogais_Y" "4.8821"
[9,] "66.6% Sindhi + 33.4% Tajiks_Y" "5.6708"
[10,] "86.4% Sindhi + 13.6% Georgians" "6.2927"

Comparing an individual against populations

DodecadOracle(c(8.4, 0, 2.8, 6, 2.2, 0.1, 40.3, 25.9, 0.3, 11.9, 1.5, 0.5))
[,1] [,2]
[1,] "Iranian_D" "2.2405"
[2,] "Kurd_D" "3.8092"
[3,] "Kurds_Y" "5.4945"
[4,] "Iranians" "6.634"
[5,] "Uzbekistan_Jews" "12.8957"
[6,] "Turks" "17.3173"
[7,] "Turkmens_Y" "17.7316"
[8,] "Iranian_Jews" "18.14"
[9,] "Assyrian_D" "18.8968"
[10,] "Azerbaijan_Jews" "18.9444"

Comparing an individual against 2-way population mixes

DodecadOracle(c(28, 0.8, 1.6, 49.9, 1.9, 0, 10.6, 4.1, 0, 2.4, 0, 0.6),mixedmode=T)
[,1] [,2]
[1,] "47.7% French_D + 52.3% Mordovians_Y" "2.5849"
[2,] "48.3% French + 51.7% Mordovians_Y" "2.6012"
[3,] "36.3% Spaniards + 63.7% Mordovians_Y" "2.9985"
[4,] "36% Spanish_D + 64% Mordovians_Y" "3.0577"
[5,] "65.9% Russian_D + 34.1% Spaniards" "3.0923"
[6,] "35.9% IBS + 64.1% Mordovians_Y" "3.0943"
[7,] "40% French + 60% Ukranians_Y" "3.1662"
[8,] "66.4% Russian_D + 33.6% IBS" "3.2359"
[9,] "24.5% Swedish_D + 75.5% Hungarians" "3.3021"
[10,] "39.3% French_D + 60.7% Ukranians_Y" "3.4046"


The numbers to the right of each result represent the "goodness" of the match; the lower, the better. If you wanted to list the top-30 results, in any of the above commands, you would enter, e.g.,

DodecadOracle(c(28, 0.8, 1.6, 49.9, 1.9, 0, 10.6, 4.1, 0, 2.4, 0, 0.6),mixedmode=T, k=30)

If you recently joined the Project, please consider leaving a brief comment in the Information about Project Samples thread.

Participant results for 'K12a' calculator

The participant results can be found in the "Individual Results" tab of the K12a spreadsheet.
You can read more about the K12a calculator at my other blog; if you are not a Project participant, you can also find a DIY version of it there, which can be used in conjunction with DIYDodecad 2.1.

Wednesday, October 26, 2011

'eurasia7' calculator

This calculator was made with 196 different populations and 2,659 individuals, including 518 project participants. The following Dodecad populations do not have 5 individuals yet, so they are included in the OTHERS_D generic category:
Algerian_D, North_African_Jews_D, Slovenian_D, Mixed_Scandinavian_D, Danish_D, Moroccan_D, Tunisian_D, Serb_D, Austrian_D, Saudi_D, Pakistani_D, Tatar_Various_D, Palestinian_D, Greek_Italian_D, Romanian_D, Swiss_German_D, Szekler_D, Mandaean_D, Azeri_D, Czech_D, Georgian_D, Belgian_D, Latvian_D, Estonian_D, Bangladesh_D, Yemenese_D, Sri_Lanka_D, Hungarian_D, Basque_D, Udmurt_D, Egyptian_D
As always, I encourage people with 4 grandparents from the same country or ethnic group of Eurasia, North or East Africa to contact me (do not send data!) for possible inclusion in the Project. If I have overlooked any such individuals, drop me a line (my e-mail address is at the bottom of the blog). I usually start a new _D population whenever individuals with 4 grandparents from the same group are submitted, but I may have missed some.

Note that all individuals from the reference populations have also been included, including outliers; you should be aware of this when reading the population averages, and consult the Outliers tab in the v3 spreadsheet for some instances of outliers.
Due to image size restrictions in Picasa, the labels are not visible well. A large version of the above plot can be found in the download bundle.

The seven ancestral populations inferred at this level of resolution are:
  • Sub_Saharan
  • West_Asian
  • Atlantic_Baltic
  • East_Asian
  • Southern
  • South_Asian
  • Siberian
As usual, you should take these names as useful labels, and interpret them in conjunction with the components' distribution in different populations, and their Fst distances, both of which can be found in the spreadsheet.

The table of Fst distances:


Below you can see a neighbor-joining tree based on inter-population Fst distances:
The first six dimensions of a multi-dimensional scaling of the same:





Calculator Files:

  • The spreadsheet contains population averages, the table of Fst distances, and individual results for included Project participants.
  • The download RAR file (Google Docs or Sendspace) contains all the files needed to run the calculator. You must download and install DIYDodecad 2.1 first. In order to run the calculator, you follow the instructions of the README file, but type 'eurasia7' instead of 'dv3'.

Terms of use: 'eurasia7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Technical Details:

The calculator is built using allele frequencies of K=7 ancestral components inferred by ADMIXTURE 1.21 analysis of 2,659 individuals. Markers included in the source datasets, as well as the Family Finder and 23andMe (as of Oct 21) platforms were included. The marker set was thinned of markers with less than 99.5% genotype rate and less than 0.5% minor allele frequency. Linkage-disequilibrium based pruning was carried out with a window size of 250 SNPs, advanced by 25 SNPs and R-squared greater than 0.4. A total of 164,990 SNPs remained after these filtering steps.

All relevant populations available to me, and genotyped at a sufficient number of markers were included. Inclusion of the Kalash population resulted in a population-specific component at K=7, and hence their admixture components were inferred a posteriori. Their proportions are consistent with previous results, showing them to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components.

Friday, October 21, 2011

Eurogenes is upset

Eurogenes seems to be upset this week, first throwing a tantrum at Dr. McDonald and then at myself. You can probably find the cached text in Google for some time, although Eurogenes has deleted his anti-McDonald tantrum, and changed the verbiage on the one directed against me on advice of some more cool-headed people. Here is the epilogue of his original anti-Dienekes rant:
Dienekes, you've got a spreadsheet online showing all sorts of weird things. You need stop being a prat, and do something about it ASAP.
Eurogenes' animus towards me is not surprising for those who have followed our interactions since the old days. Of course he is benefiting from my work (I have pointed him towards data he didn't know existed, he is using DIYDodecad, as well as the 1000Genomes data extracted with my code by the MDLP), so one would think that if he had any criticism against me, he would at least express it in a more dignified way.

Of course, being rude, ungrateful and mean-spirited does not mean one is wrong! So, what has Eurogenes actually discovered?

He noted the high Ukrainian West/East European ratio produced by Dodecad v3, and objected to my idea that Ukrainians were transitional to the Balkans and the Caucasus. Actually, according to the PCA plot of the Yunusbayev et al. (2011) paper, they are transitional, being situated toward both the Balkans and the Caucasus, relative to Belorussians/Lithuanians, i.e., the populations that generally show peaks of East European-related components. This is also supported by the ADMIXTURE analysis that reveals Ukrainians to possess a Caucasus-centered component largely lacking in other Eastern Slavs, but shared with Balkan/Caucasus populations.

Should I have not tested the new Yunusbayev data with Dodecad v3 and reported their results? Of course not. When one has a measuring instrument, one uses it on new data to test its performance and reports what he sees. This is exactly what I have done. At the same time, one uses the new data to create new measuring instruments that have been trained using all available data, which is also what I have done with euro7 and the upcoming Dodecad v4.

To make matters worse, Eurogenes suggests that my euro7 analysis agrees with his K=10 which was presented two weeks later. So, apparently, I am posting correct information about Ukrainians 2 weeks before he does, and this means that I am turning around to his way of thinking rather than vice versa. Go figure.

Eurogenes continues with his posting of supposed MDS/PCA plots supporting his thesis. Actually, what he has posted are plots based on metric distances in the space of admixture proportions; these are not genetic distances because e.g., a +/- 1% difference in a Sub-Saharan component results in the same Euclidean distance difference as a +/-1% in a European one, although the former affects genetic distance much more strongly than the latter. Metric distances are fine to quickly determine closeness of samples in the space of admixture proportions, but they are certainly no substitute for real genetic distances. I have already linked above with evidence that Ukrainians are transitional to the Balkans and the Caucasus relative to the Yunusbaeyev et al. populations.

I am also, apparently, accused of neglecting to point out the deficiencies of Dodecad v3, and I am invited by Eurogenes to retract it completely! This proposal is equivalent to the idea that we should burn old topographic maps that were based on measurements with sticks, ropes, and trigonometers, because we can now measure distances with laser beams. And, it is funny indeed that I am supposedly neglecting the deficiencies of Dodecad v3 when, 3 weeks before the Eurogenes rant, I post exactly what its limitations are, and how it can be made better.

It is unfortunate that Eurogenes has chosen to go down that path. Envy is not a good guide to behavior, and perhaps, instead of relishing at the prospect of putting others down, he could spend a little more time inventing something of his own.

As for myself, I will continue to work on my tools, and to encourage cross-pollination between different projects for the benefit of all.

UPDATE: In a newer post, Eurogenes attempts to justify his mishandling of MDS, by suggesting that he presented results based on raw SNP data. This is of course nonsense, since Eurogenes does not have the raw SNP data of the Dodecad populations. He is comparing apples and oranges by comparing plots made on raw data with those made in the space of admixture proportions. Furthermore, his supposed findings have no bearing on the Yunusbayev et al. ADMIXTURE and PCA results, posted above.

Sunday, September 25, 2011

Yunusbayev et al. (2011) data assessed with Dodecad v3

I have acquired the data from the recent Yunusbayev et al. (2011) paper on the Caucasus. This includes the following populations:
  • Kurds_Y 6
  • Bulgarians_Y 13
  • Ukranians_Y 20
  • Mordovians_Y 15
  • Armenians_Y 16
  • Abhkasians_Y 20
  • Balkars_Y 19
  • North_Ossetians_Y 15
  • Chechens_Y 20
  • Nogais_Y 16
  • Kumyks_Y 14
  • Turkmens_Y 15
  • Tajiks_Y 15
It is a valuable new addition to the Project, and it is commendable that it has been made publicly and easily available so swiftly after the appearance of the Yunusbayev et al. (2011) paper.

To get the ball rolling on the new Yunusbayev et al. data, I will map the new populations onto the Dodecad v3 components; they will be added to the Dodecad v3 spreadsheet as they are calculated.

I have been laboriously designing a new global (including Amerindians and Australasians) Dodecad X1 experimental calculator with 3,010 individuals for a few weeks now, but I guess I will now have to reboot it with 3,214.

Together with some other new data I recently discovered, I now have 9,799 individuals (some duplicates from different sources) in my global database. My Dodecad dataset of 511 individuals from a single country or ethnic group isn't too shabby either. Let's hope for a new data release that will push the data collection above the magic 10,000.

UDDATE:

I have added the first 7 populations to the spreadsheet; the others are being calculated as we speak. Most of them seem in line with expectations, but the Abkhasian sample has one outlier individual (abh27), and has thus been placed in the "Outliers" tab of the spreadsheet; a new set of admixture proportions, minus that outlier individual, will be calculated anew:

UPDATE II: The population portraits have been uploaded to Google Docs as a rar file (Sendspace mirror). Average admixture results have all been entered to the spreadsheet.

Thursday, September 15, 2011

Third-party tools based on the Dodecad Project

Gedmatch.com has made available some tools based on Dodecad v3 as bundled in DIYDodecad 2.0. In addition to the regular admixture analysis, there is a chromosome painting, and an option to compare 2 kits. This should be quite useful to Mac users, who can't use DIYDodecad at present. Gedmatch.com requires upload of your data to the server side, which provides the benefit of the other tools of the site, but may not be ideal for people with privacy concerns for whom the DIY tool was partly built.

Note that because admixture estimation is expensive computationally, the Gedmatch.com tools are slightly less accurate than DIYDodecad because of more lax termination criteria. This should not be a problem for the major components of one's ancestry, but may be for the minor ones. Moreover, convergence is achieved with a different number of iterations for different genotype files, so accuracy may vary.

Two other genome bloggers have released their own calculators that can be run with DIYDodecad. Magnus Ducatus Lituaniae has released MDLP based on its K=7 analysis. Eurogenes has released a K=14 imaginatively named test calculator for Eurasian data.

I keep a list of calculators for DIYDodecad in the DIYDodecad 2.0 page.

I neither endorse nor am I affiliated with any third-party tools, but I encourage readers to try them out; the more the merrier.

Monday, September 5, 2011

'bat' calculator (Balkans-Anatolia-Turkic)

I have decided to make a new calculator for DIYDodecad that may be useful for individuals from the Balkans and Anatolia. You can download it from here at Google Docs (or here from sendspace). The terms of use are the same as for DIYDodecad v 2.0. To run it, you simply extract the contents of the RAR file in your working directory, and type bat.par whenever you typed dv3.par in the instructions.

The reference populations can be seen below. I have included all available Balkan populations, as well as Turks and Armenians. Moreover, I have included all available Turkic populations.

The marker set is the same as used in Dodecad v3. Three components emerge: one centered in the northern Balkans, one in eastern Anatolia, and one present in various proportions among all Turkic populations (see Turkic cline).


The components have been named accordingly, but please note that they do not necessarily reflect recent ancestors. For example, it is a good hypothesis that the Anatolia component was present in the Balkans even in ancient times, so one need not seek a recent Anatolian ancestor to explain its presence in a Balkan individual. Similarly for the Balkans component in Anatolia, which may reflect the diverse Balkan peoples that have settled in Anatolia since the dawn of history, so a present-day inhabitant of Anatolia need not seek a recent Balkan ancestor.

Likewise, the Turkic component is only part of the genetic makeup of the Turkic speakers who arrived in Anatolia, since those probably also carried West Eurasian population elements picked up en route from Siberia to Anatolia.

The way to interpret your results is to see whether you have an excess or deficiency of any component relative to your ethnic group. For example, an Anatolian Greek may have a higher Anatolia/Balkans ratio than a Balkan Greek and likewise for a Balkan vs. Anatolian Turk; the latter may also have a variable Turkic component which will reflect differential Central Asian input.

Tuesday, August 23, 2011

Populations in need of 5 participants

Submission to the Project is currently closed, but I am often willing to include new members if they contact me (dodecad@gmail.com) before sending their data with some information about their ancestry.

I am most likely to accept new participants from:
  • Greece, the Balkans, Italy, and West/Central Asia
  • Under-represented populations
  • Populations that are a few members short of reaching the 5-person mark, after which I can calculate an average for them.
I typically don't accept new participants of multiple ancestries; I've made DIYDodecad for just that case.

Here is a list of populations that are short 1-2 participants:

Algerian_D 4
North_African_Jews_D 4
Slovenian_D 4
Bulgarian_D 4
Danish_D 3
Moroccan_D 3
Tunisian_D 3
Mixed_Scandinavian_D 3
Serb_D 3
Austrian_D 3
Saudi_D 3
Pakistani_D 3
Tatar_Various_D 3
Palestinian_D 3

Any individuals from the Balkans are strongly encouraged to contact me, as the Balkans_D sample size of 17 can soon be broken down into specific populations if a few more individuals from different Balkan populations join the Project.

I also encourage new members to post their information in the ancestry thread.

Monday, July 25, 2011

Do-It-Yourself Dodecad v 1.0

(UDPATE: There is a new 2.0 version of the software)

I have decided to release a Do-It-Yourself calculator (and mirror), for several reasons:
  1. So that people who don't want to send me their data can still get their results
  2. So that people can estimate admixture proportions in all their relatives, as relatives can't be accepted in the Project
  3. So that people of mixed ancestry can get their results, as there have been limited opportunities for them to submit their data to the Project so far
  4. Most importantly, so that I won't have to do-it-myself ;-)
You need a Windows or Linux 32bit/64bit machine to run DIYDodecad. The instructions should be easy to follow, but if you encounter any bugs or have any problems, feel free to leave a comment or write to me (dodecad@gmail.com).

Of course, I will continue to ask for people to send me their data in the future: the calculator is made possible in part because of their contributions. Project participants have added benefits, such as the more specialized Clusters Galore or regional analyses.

If you are a project participant, you can still try DIYDodecad; you will get slightly different results than the ones you already have, because DIYDodecad does not use the same "random seed" as ADMIXTURE, and has a different default convergence criterion (maximum log-likelihood change of 1e-6 between successive EM iterations). You will also need the program in the future, as more "calculators" will be disseminated for it.

So, if you don't want to/can't join the Project, you can still get your Dodecad v3 results; you can also try the Dodecad Oracle with them. Also, feel free to leave a comment in this post with your results.

Related: some background on the creation of DIYDodecad

Tuesday, July 19, 2011

The Dodecad Oracle v1

Here is a little fun tool that tests the Dodecad v3 admixture proportions of an individual against all the reference populations, but also against the best pairwise combinations of these populations.

You need to install R to use it, and then download the program and double click on the file DodecadOracleV1.RData that can be found within the rar file. You will then be faced with a command prompt where you can enter the following commands:

Examining which populations are available

Just enter

X[,1]

You will see a list of 227 populations. You can use these population IDs in the next section.

Which populations are closest to a particular population?

Enter:

DodecadOracle("British_D")
[,1] [,2]
[1,] "British_D" "0"
[2,] "British_Isles_D" "0.9798"
[3,] "Cornwall_1KG" "1.1533"
[4,] "Kent_1KG" "2.265"
[5,] "Irish_D" "3.7643"
[6,] "Dutch_D" "4.5354"
[7,] "Mixed_Germanic_D" "6.8971"
[8,] "Norwegian_D" "11.3111"
[9,] "Orkney_1KG" "12.4652"
[10,] "Orcadian" "12.8195"

If you want to find e.g., the top-30 populations, rather than just the top-10, enter:

DodecadOracle("British_D", k=30)

Which populations are closer to a particular individual?

Enter the admixture proportions of the individual (from the "Individual results" tab of the spreadsheet) as follows:

DodecadOracle(c(4.6, 16.7, 33.6, 0, 23.2, 0.4, 0.6, 1.6, 0.7, 14.1, 4.5, 0.2))
[,1] [,2]
[1,] "Ashkenazi_D" "3.7908"
[2,] "Ashkenazy_Jews" "4.1473"
[3,] "Morocco_Jews" "6.338"
[4,] "S_Italian_Sicilian_D" "12.5443"
[5,] "Sephardic_Jews" "13.5067"
[6,] "C_Italian_D" "14.4554"
[7,] "Sicilian_D" "14.7469"
[8,] "S_Italian_D" "15.748"
[9,] "Tuscan_X" "15.9981"
[10,] "O_Italian_D" "16.1474"

Once again, you can specify k=30, if you desire the 30 top matching populations instead of the default 10.

Mixed Mode

You use mixed mode by adding mixedmode=T in any of the commands. The program then considers all pairs of populations, and for each one of them calculates the minimum distance to the sample in consideration, and the admixture proportions that produce it; population pairs where the distance to one of the two populations is smaller than to any admixture of the two are ignored.

Example:

DodecadOracle("Pathan",mixedmode=T)
[,1] [,2]
[1,] "Pathan" "0"
[2,] "84.8% Pakistani + 15.2% Urkarah" "1.075"
[3,] "84% Pakistani + 16% Stalskoe" "1.1555"
[4,] "63.9% TN_Brahmin + 36.1% Urkarah" "1.6669"
[5,] "32.4% Urkarah + 67.6% Meghawal" "2.3516"
[6,] "56.3% INS + 43.7% Urkarah" "2.4901"
[7,] "11.5% Adygei + 88.5% Pakistani" "2.6245"
[8,] "82.4% Sindhi + 17.6% Stalskoe" "2.6318"
[9,] "62.9% AP_Brahmin + 37.1% Urkarah" "2.7322"
[10,] "11.2% Lezgins + 88.8% Pakistani" "2.7749"

The mixed mode should be used with caution, and it shows, more than anything else, how similar apparent "mixes" can be achieved by different combinations of ancestry. Nonetheless, it may prove somewhat useful. For example, there is a suggestion in the above results, that Pathans can be viewed as a mix of other South Asian populations and populations from the eastern Caucasus, a suggestion that was arrived at independently by the Project using different methods.

Here is another example:

DodecadOracle("Assyrian_D",mixedmode=T)
[,1] [,2]
[1,] "Assyrian_D" "0"
[2,] "83.9% Armenians_16 + 16.1% Yemen_Jews" "1.7829"
[3,] "89.1% Armenian_D + 10.9% Saudis" "2.1624"
[4,] "84.3% Armenians_16 + 15.7% Saudis" "2.2884"
[5,] "88.9% Armenian_D + 11.1% Yemen_Jews" "2.2983"
[6,] "83.8% Armenian_D + 16.2% Bedouin" "4.1579"
[7,] "72.2% Armenian_D + 27.8% Syrians" "4.1841"
[8,] "23.4% Georgians + 76.6% Iraq_Jews" "4.2418"
[9,] "76.2% Armenians_16 + 23.8% Bedouin" "4.332"
[10,] "61.5% Armenians_16 + 38.5% Syrians" "4.4019"

This reaffirms the close relationship of Assyrians to Armenians that has been noticed in the project and by others, and it also shows that Assyrians differ from Armenians in a Southwestern Asian direction, consistent with their Semitic language.

Or, African Americans:

DodecadOracle("ASW",mixedmode=T)
[,1] [,2]
[1,] "ASW" "0"
[2,] "81.3% Hausa + 18.7% N._European" "2.3891"
[3,] "18.4% Orkney_1KG + 81.6% Hausa" "2.4031"
[4,] "18.5% Argyll_1KG + 81.5% Hausa" "2.4268"
[5,] "18.4% Orcadian + 81.6% Hausa" "2.4657"
[6,] "80.5% Igbo + 19.5% N._European" "2.5031"
[7,] "80.6% Brong + 19.4% N._European" "2.523"
[8,] "18.6% CEU + 81.4% Hausa" "2.5938"
[9,] "19.1% Argyll_1KG + 80.9% Brong" "2.6197"
[10,] "19% Orkney_1KG + 81% Brong" "2.6274"

I don't know that much about the slave trade, but I believe that Ghana was an important part of it?

Another thing to watch, is that some populations tend to have more than one sample available, so they appear to be mixtures of themselves, which is not really very informative, e.g., Spanish_D

DodecadOracle("Spanish_D",mixedmode=T)
[,1] [,2]
[1,] "Spanish_D" "0"
[2,] "7.9% French_Basque + 92.1% IBS" "0.8713"
[3,] "68.9% IBS + 31.1% Spaniards" "1.0377"
[4,] "98.8% IBS + 1.2% Irish_D" "1.2959"
[5,] "1.2% British_Isles_D + 98.8% IBS" "1.3018"
[6,] "1.2% British_D + 98.8% IBS" "1.3019"
[7,] "99% IBS + 1% Norwegian_D" "1.3046"
[8,] "1.2% Cornwall_1KG + 98.8% IBS" "1.3048"
[9,] "98.8% IBS + 1.2% Kent_1KG" "1.3142"
[10,] "2.2% French_D + 97.8% IBS" "1.3179"

To deal with these problems, you must "edit" the X matrix if you want to exclude some populations. For example, if you want to exclude "Spaniards" and "IBS", you must enter:

X <- X[setdiff(1:227,which(X[,1]=="IBS" | X[,1]=="Spaniards")),]

but notice, that you must relaunch the program, if you want to get the original matrix, or alternatively save it like this:

Z<-X

and then retrieve it like this:

X<-Z

Monday, June 27, 2011

Submission opportunity... is OVER

The latest submission opportunity is now over. All those who meet the eligibility criteria and who sent me their data before this announcement will get their IDs/have their data processed. Do NOT send unsolicited data after this time, as they will most likely be ignored.

Wednesday, June 22, 2011

Dodecad v3: population averages

With Dodecad v3 it is possible to use any population as test data in supervised ADMIXTURE analysis, and extract its admixture proportions in terms of the 12 ancestral components.

So, I have set up an automated job that will do just that: use pretty much every population available to me under only a few conditions:
  1. Each population must have at least 5 individuals
  2. It must have the same 166,462 SNPs on which the test is based
  3. It must not be from a group not covered by the test (e.g., Australo-Melanesians or Native Americans)
By my count, I have 141 different populations that meet these requirements. Each population is run on its own in a supervised ADMIXTURE analysis together with the 600-strong synthetic set (50 per ancestral component). These are ideal conditions to produce a high-quality comparative reference set.

The average admixture proportions of different populations will be put in this spreadsheet as they are calculated, which will probably take several days.

Tuesday, June 21, 2011

The design of Dodecad v3

Dodecad v2 was short-lived, as I discovered a way to improve it shortly after I announced it.

The first step was to carry out an extensive K=3 ADMIXTURE analysis of about 130 different populations and about 2,000 individuals from Europe, Asia, and Africa. Using the allele frequency results of this analysis I was able to create the most comprehensive synthetic individuals to represent West Eurasians, Asians, and Sub-Saharan Africans.

Subsequently, I carried out an analysis of East Eurasian populations using the West Eurasian/Sub-Saharan synthetic individuals as controls, as well as an analysis of Sub-Saharan populations using the West Eurasian/Asian individuals as controls.

In East Eurasia, I was able to infer the existence of two components, one centered in the extreme northeast, another in the southeast, with many other populations arrayed between these two extremes:


In Sub-Saharan Africa, the primary division was between San, Mbuti, and Biaka Pygmies (whom I have called "Palaeo-Africans") and the rest (Yoruba, Mandenka, and Bantu, "Neo-Africans"):


Now, I had four synthetic "framing populations": Neo-Africans, Palaeo-Africans, Northeast Asians and Southeast Asians, created from hundreds of individuals from several different populations:
  1. I did not have to choose a particular population (e.g., Chinese) to represent East Asia
  2. I did not have to aggregate individuals from populations with variable levels of non-East Asian admixture
I now used my South Asian populations, together with Neo-African, West Eurasian, Northeast and Southeast Asian controls to extract a South Asian specific component:



Armed with these 5 synthetic "framing" populations, I carried out a K=12 analysis with my West Eurasian, South Asian, and North/East African populations (1,247 individuals; 69 populations):

And, finally, I generated 50 synthetic individuals from each of the 12 inferred components to create a dataset of 600 individuals that will be the basis of Dodecad v3.

Below is the table of Fst divergences:



The following MDS plots show the first 10 dimensions of variation of these individuals:

Finally, here is a neighbor-joining tree of the 12 components:
(to be continued)

Wednesday, June 8, 2011

Dodecad v2

This is an announcement of the new generation of Dodecad ancestry analysis. In comparison to the standard K=10 used since the beginning of the Project:
  1. Participants' data are now used to enrich the set of reference populations and to help define new ancestral components
  2. Rather than choosing arbitrary reference populations, I employ a very large set of individuals to capture allele frequencies and then create synthetic individuals ("panmictic zombies") that embody these frequencies; more on this below.
  3. Results for unrelated project participants will be reported in a separate post, using my new technique of converting unsupervised ADMIXTURE runs into supervised ones. Hence, Project participants can expect to receive new K=12 results; moreover, the fact that this will be done in supervised mode means that it is no longer necessary to process samples in small batches of 10 or so. All current unrelated participants will receive their results in one go, and only future submissions will be processed in batches.
This analysis utilizes results from Project participants (populations with _D endings), as well as synthetic individuals summarizing allele frequencies of East Eurasians, Sub-Saharan Africans, and South Indians (populations with _Z endings)

The framing populations (_Z)

The following _Z populations were included:
  • Sub_Saharan_Z: Bantu, Yoruba, Mandenka, San, and Pygmies from HGDP-CEPH
  • South_Indian_Z: North Kannadi, Sakilli from Behar et al. (2010), AP_Madiga, AP_Mala, TN_Dalit from Xing et al. (2010), Bhil, Chenchu, Kurumba, Satnami, Madiga, Mala, Kamsali, Onge, Great_Andamanese from Reich et al. (2009)
  • Sino_Tibetan_Z: Yizu, Naxi, Han, Tujia from HGDP-CEPH
  • Altaic_Z: Tu, Xibo, Mongola, Daur, Hezhen, Oroqen, Yakut from HGDP-CEPH, and Evenk, Buryat from Rasmussen et al. (2010)
  • Siberian_Other_Z: Selkup, Ket, Yukagir, Nganasan, Koryak, Chuckchi from Rasmussen et al. (2010)
  • Southeast_Asian_Z: Dai, Lahu, Miaozu, Cambodians from HGDP-CEPH, Khmer-Cambodian, Thai from Xing et al. (2010), and Singapore Malay from the Singapore Genome Variation Project
The 12 inferred ancestral components

Results of the ADMIXTURE analysis defining the new K=12 components of the Project can be seen below:
Raw proportions can be found in a spreadsheet. There are also population portraits in a zip file, showing individual-level variation.

The 12 components are:
  • West_Asian
  • East_European
  • West_European
  • East_Asian
  • Mediterranean
  • Northwest_African
  • North_Eurasian
  • Arabian
  • Inner_Asian
  • Sub_Saharan
  • East_African
  • South_Indian
Once again, I have tried to make these as neutral and appropriate as possible, but don't forget that they are simply descriptive labels to aid memory. For example, the Arabian component is centered on Saudis, Yemenese, and Yemen Jews, the Inner Asian component on the Altaic synthetic population, and so on.

The Fst divergences between the 12 components can be seen in the spreadsheet and also below:

A different way of showing them is via a neighbor-joining tree. Note, however, that this is not a replacement for the Fst table above which alone fully preserves the inter-population relationships:
We can also plot the first few MDS dimensions using synthetic individuals from the 12 components; again, these capture variation only partially:



What comes next?

Hopefully quite soon, I will:
  1. Report new v2 results for all project participants
  2. Report new v2 proportions for many other populations not included here
Project members who still haven't received their results (during the ongoing submission opportunity) can expect to receive K=10 standard results, and they will receive their new v2 results later.