Showing posts with label DIYDodecad. Show all posts
Showing posts with label DIYDodecad. Show all posts

Friday, November 30, 2012

Geno 2.0 patch for DIYDodecad

(See important update at the end of this post)

People who have tested using the Genographic Project's Geno 2.0 test can now use the DIYDodecad tool with their data. The raw data download from this test has a slightly different format than the ones from 23andMe and Family Finder, so it is necessary to convert your data in a format that DIYDodecad can interpret.

So, after you have downloaded and extracted the DIYDodecad software as per its instructions, you should also download a couple of extra files into your working directory; these files are included in this patch:

  • standardize.r which replaces the standardize.r in the DIYDodecad software bundle, and allows you to convert your Geno 2.0 formatted data
  • hgdp.base.txt which includes additional information about SNP markers that is not found in your Geno 2.0 raw data download, and which is necessary to complete the conversion process.
Once these two files have been extracted into your working directory, the process of using DIYDodecad is exactly the same as for any other user of the software.

The only difference is that at the step where you convert your data using the standardize command (see DIYDodecad README file), you will use the command:


standardize('johndoe.csv', company='geno2')

where johndoe.csv is your unzipped raw data download. This will write a genotype.txt file in the working directory, and you can proceed the rest of the way as per the instructions.

You can use all ancestry calculators released by the Project (or indeed other projects); the most recent one is globe13

You should be aware, that because the Geno 2.0 test includes a smaller number of SNPs, and because globe13 and other calculators were developed using the common SNP set of 23andMe and Family Finder, the analysis using globe13 will only include ~34 thousand SNPs and will be "noisier" than usual. In the future, I might develop new calculators that make use of the SNP set of the Geno 2.0 test itself.

PS: Feel free to post a comment below if you experienced any difficulty converting your data; also thanks to CeCe Moore for graciously sharing a raw data file with me, which allowed me to build this converter.

UPDATE:

Apparently, the data format has been changed for some Geno 2.0 data downloads.
If your data includes a [Header] ... [Data] preamble followed by a list of 5 comma-separated values, ignore this.
If it includes a header "SNP,Chr,Allele1,Allele2" followed by a list of 4 comma-separated values, you should follow the instructions as above, but use company='geno2new' instead.

Wednesday, October 31, 2012

'globe13' participant results

Project participant results for the globe13 calculator can be found in the spreadsheet. Population median results and Fst divergences are also included.

Below, you can see the first two dimensions of an MDS plot of the 13 components:

A neighbor-joining tree of the 13 components based on the Fst divergences:
I have also created a TreeMix plot using Palaeo_African as an outgroup, and allowing as many as 5 migration edges:
The actual tree is:


((West_African:0.00448794,(East_African:0.00506576,(((((East_Asian:0.0173284,Siberian:0.00732773):0.0027852,(Amerindian:0.026174,Arctic:0.0118342):0.00742092):0.0114738,Australasian:0.0488974):0.00266559,South_Asian:0.00734044):0.008089,(Southwest_Asian:0.00541405,((West_Asian:0.00620657,North_European:0.00657599):0.00311587,Mediterranean:0.00798949):0.00650328):0.0118925):0.0299627):0.00597674):0.00671186,Palaeo_African:0.0215931);
0.0640319 NA NA NA Palaeo_African:0.0215931 Australasian:0.0488974
0.270468 NA NA NA Australasian:0.0488974 East_Asian:0.0173284
0.185213 NA NA NA South_Asian:0.00734044 ((West_Asian:0.00620657,North_European:0.00657599):0.00311587,Mediterranean:0.00798949):0.00650328
0.129883 NA NA NA North_European:0.00657599 Amerindian:0.026174
0.138757 NA NA NA Arctic:0.0118342 (West_Asian:0.00620657,North_European:0.00657599):0.00311587

Monday, October 29, 2012

'globe13' calculator

The globe13 calculator is based on the K=13 analysis. It includes the following components:


  • Siberian
  • Amerindian
  • West_African
  • Palaeo_African
  • Southwest_Asian
  • East_Asian
  • Mediterranean
  • Australasian
  • Arctic
  • West_Asian
  • North_European
  • South_Asian
  • East_African

Fst divergences between ancestral components can be found here.

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'globe13' instead of 'dv3' in these instructions. You can consult the spreadsheet for proportions of the 13 components in different world populations.

Terms of use: 'globe13', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Tuesday, October 23, 2012

'globe10' calculator

As part of the on-going analysis of the world dataset, I am releasing the 'globe10' calculator, which is based on the K=10 analysis. This calculator includes the following ancestral components:
  • Amerindian
  • West_Asian
  • Australasian
  • Palaeo_African
  • Neo_African
  • Siberian
  • Southern
  • East_Asian
  • Atlantic_Baltic
  • South_Asian
The names may be the same as the ones from previous calculators released by the Project, but you should always consult the spreadsheet to see how they might differ. In this case, inclusion of Amerindian, Australasian populations, African hunter-gatherers, dealing with the Paniya issue, and inclusion of data of Schlebusch et al. (2012), and  Pagani et al. (2012), have all combined to change components in subtle ways, although their modalities remain largely unchanged, and hence so do the names.

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'globe10' instead of 'dv3' in these instructions. You can consult the spreadsheet for proportions of the 10 components in different world populations.

Terms of use: 'globe10', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Friday, October 19, 2012

'globe4' calculator

Patterson et al. (2012) recently published evidence for admixture in northern Europeans between a population resembling modern Sardinians (and the Neolithic Tyrolean Iceman, whose genome was published earlier this year), and, surprisingly Native Americans. The authors attribute the Amerindian-like ancestry element to a North Eurasian population that spawned Native Americans, and which also contributed ancestry to northern Europeans. They propose two possibilities for the origin of this admixture: (i) the Mesolithic Europeans resembled Amerindians, or (ii) there was an influx of Amerindian-like populations from the east during late prehistory. A palimpsest of these two processes may explain parts of the observed signal of admixture.

In a recent K=4 admixture experiment, I demonstrated that ADMIXTURE software produces an Amerindian ancestral component that closely tracks the signal of admixture using the D-statistic test. I have decided to make this test available for download and use with DIYDodecad.

The test has four ancestral populations:
  • European
  • Asian
  • African
  • Amerindian
It is important to remember that some of these components track different aspects of ancestry that is better resolved at higher resolution. There are also populations that "don't fit well" in this 4-partite scheme (e.g., certain African or Australasian populations).

For example, the Amerindian component of this test may indicate (i) real recent Native American ancestry, (ii) East Eurasian ancestry found in Siberia and East Asia, (iii) the common signal of admixture differentiating most European groups from Sardinians and Near Eastern Caucasoid groups. Similarly, the Asian component may indicate Australasian, South Asian, or East Eurasian ancestry. And, the European component tracks the ancestry of individuals from West Eurasia in general, although it reaches is maximum in Sardinians.

This test may, however, be useful to Old World individuals who want to get an idea about the signal of admixture discovered by Patterson et al., so I decided to make it available. For individuals who don't suspect recent Amerindian or Siberian/East Asian ancestry, and who don't belong to populations with recent such ancestry, the Amerindian component will most likely represent the aforementioned signal.

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'globe4' instead of 'dv3' in these instructions. You can consult the spreadsheet for proportions of the 4 components in different world populations.

Terms of use: 'globe4', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Saturday, October 13, 2012

Geno 2.0 data request

If anyone has received results from the Geno 2.0 test of the Genographic Project and want to share it with me, feel free to send it at dodecad@gmail.com. I will not distribute it or share it with anyone. I want to see what SNPs are tested, what format the data is in, and what is its intersection with other available datasets. This way, I can update my DIYDodecad software so that Geno 2.0 testees can use the various calculators released by the project to get an alternative ancestry assessment.

In time, and if there is interest, I may release additional calculators that make use of the particular SNP set tested by Geno 2.0.

Tuesday, June 12, 2012

'K10a' calculator

The 'K10a' calculator represents an intermediate stage between the K7 and K12 analyses released so far from the Project. The following components have been inferred:
  • Palaeoafrican 
  • South_Asian 
  • West_Asian 
  • Southeast_Asian 
  • Sub_Saharan 
  • Atlantic_Baltic
  • Red_Sea 
  • East_Asian 
  • Mediterranean 
  • Siberian 
There are a couple of points of interest; first, the Red_Sea component related Arabians with East Africans. At a higher level of resolution the "Southwest_Asian" and "East_African" (K12)  components emerge. The "Red_Sea" component is not very closely related to any other components, but is somewhat related to the "Mediterranean" and "Atlantic_Baltic" components.

So, using the different calculators of the Dodecad Project, we first have (K7) a contrast between Africa and West Eurasia, then a signal of the shared ancestry between Arabia and East Africa (K10), and finally, strong signals of local ancestry in the two regions.

Second, the Mediterranean component here is modal in Sardinians as usual, but also projects into North Africa. Again, this is intermediate between K7 which shows a predominance of West Eurasian ancestry in North Africa + an African component, and K12 in which there are "Atlantic_Med" and "Northwest_Afican" regional components.

These are strong hints that the West Eurasian element in Africa differs between NW and E Africa. In the former region, it is most related to Sardinians, and in the latter it is most related to Arabians. Of course, ultimately the two elements are related to each other.

Table of Fst distances between components:


MDS plots of the first few dimensions:


Downloads: 
Project participants can find their results in the spreadsheet. Non-participants can use DIYDodecad to calculate their results, but they should place all the calculator files in the same directory as the DIYDodecad software, and replace 'dv3' with 'K10a' in all the instructions of the README file.

Component labels are indicative, and you should compare your results against the normalized median results for different populations included in the spreadsheet.

Terms of Use

You are free to use 'K10a', including all downloaded files for any non-commercial purpose, as long as you attribute them to the Dodecad Project and to Dienekes Pontikos as follows:

The 'K10a' admixture calculator is courtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project; more information here.

Saturday, June 9, 2012

'weac2' calculator

I have made a new version of the 'weac' calculator (West Eurasian cline). This is based on a large Old World dataset at K=7 and includes the following ancestral components:
  • Palaeoafrican 
  • Atlantic_Baltic 
  • Northeast_Asian 
  • Near_East 
  • Sub_Saharan 
  • South_Asian 
  • Southeast_Asian 

The West Eurasian cline is formed between the Near_East and Atlantic_Baltic components.

Here is the table of Fst distances between components:

MDS plots of the first few dimensions:


Downloads:
Project participants can find their results in the spreadsheet. Non-participants can use DIYDodecad to calculate their results, but they should place all the calculator files in the same directory as the DIYDodecad software, and replace 'dv3' with 'weac2' in all the instructions of the README file.

(NOTE: Some  IDs may have wrong results in the spreadsheet because of a misalignment of IDs with results; I'll fix this and update this notice. UPDATE: Results should be correct in spreadsheet now - 9 Jun 2012)


Component labels are indicative, and you should compare your results against the normalized median results for different populations included in the spreadsheet.

Terms of Use

You are free to use 'weac2', including all downloaded files for any non-commercial purpose, as long as you attribute them to the Dodecad Project and to Dienekes Pontikos as follows:

The 'weac2' admixture calculator is courtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project; more information here.

Friday, April 27, 2012

Estimating your Gök4-related ancestry


I have taken Table S15 from Skoglund et al. (2012), and the Dodecad Project K7b admixture proportions in order to investigate possible relationships.

In Table S15 the authors estimate the Neolithic farmer ancestry in several populations on the basis of a single Neolithic individual from the Funnel Beaker (TRB) culture which was found in a megalithic burial in Gökhem parish.

Most of these populations are already part of the Dodecad Ancestry Project, except the three Swedish samples; given the intermediacy of the Central_Sweden sample, I have decided to use my Swedish_D sample of Project participants as a stand-in for it.

Below, you can see a scatterplot relating Gök4-related ancestry with K7b "Southern" component:



The correlation between the two variables is very strong (R-squared = 0.93).

Dodecad Project participants who already have K7b results, as well as customers of DTC testing companies who can use DIYDodecad together with the K7b calculator can approximately estimate their Gök4-related ancestry by plugging in their "Southern" value (in %) into the following equation:

Gök4-related ancestry = 1.721*Southern+19.736

I anticipate that when I am able to study the Neolithic Swedish genomes directly, the Neolithic farmer from Sweden will turn up "Southern" in a K=7 resolution experiment.

Tuesday, January 31, 2012

'K12b' and 'K7b' calculators

I am releasing two new calculators with K=12 and K=7 components, named 'K12b' and 'K7b'. You can scroll down to the bottom if you are just interested in the downloads, or read on.


New Features

The new 'K12b' calculator is an update of the previous K12a one, that was inferred using all the new samples submitted during the last submission opportunity. The 12 components are still roughly the same, although their allele frequencies may have changed by a bit, so existing participants can expect to have slightly altered results, and new participants in the Project more so, since their data are now contributing to the creation of the new tool. Non-participants can, of course, use the new calculator with DIYDodecad.

I have also taken the opportunity to do some minor tweaks. I am releasing population portraits for K12b (which were lacking in K12a); I've changed my visualization code so that the sample IDs of non-Dodecad populations can now be seen in the barplots. This may be useful for anyone else using these reference populations, by quickly identifying potential outliers in them.

I have also decided to use normalized median admixture proportions for the populations. For example, if 5 individuals in a population have 0, 0, 0.2, 0.5, 10.0% of a particular component, then the average is 2.14%, but the median is 0.2%. By using the median, the proportions become less susceptible to the presence of outliers (such as the 10%). However, if the median is calculated over every component separately, it is no longer guaranteed that the components will add up to 100%; this can be addressed by re-normalizing them (scaling them by a constant factor) so that they do. I believe that use of the normalized median will not only give better proportions that are less susceptible to outliers, but will also improve results of the new Dodecad Oracle for K12b.

At the same time I am also releasing 'K7b' which is an update of the existing 'eurasia7' calculator and which has been built on exactly the same dataset as 'K12b' but at a lower (K=7) level of detail.

Information on K7b


Information spreadsheet.

Normalized median admixture proportions barplot for all included populations (a high resolution version of this is included in the download bundle):


Table of Fst divergences:

Neighbor-joining tree (based on above):

Information on K12b


Information spreadsheet.


Normalized median admixture proportions barplot for all included populations (a high resolution version of this is included in the download bundle):

Table of Fst divergences:

Neighbor-joining tree (based on above):
Multidimensional Scaling Plots of K12b and K7b


I have created MDS plots using synthetic individuals representing the 12 ancestral components of K12b and the 7 ancestral components of K7b. By including both in the same plot, one gets an idea of the relationship of the components at different resolution. The first 10 dimensions can be seen below:

Here is a blowup of the main West Eurasian groups from the plot of the first two dimensions:

Some observations:

  • The Atlantic_Med component which is bi-modal in Basques and Sardinians occupies the apex of the figure; this makes sense, since Southwest Europe is quite distant (along land routes) to both Asia and Africa.
  • The Caucasus component is surrounded by most of the others; this is consistent with my theory elaborated in The womb of nations: how West Eurasians came to be.
  • The Atlantic_Baltic component (from K=7) is intermediate between the Atlantic_Med and North_European components.
  • Similarly, the West_Asian component (from K=7) is intermediate between the Caucasus and Gedrosia components; the Gedrosia component diverges in the direction of the Asian groups (not shown in this figure), and in particular of South Asians. This divergence can also be seen in the plot of dimension #3.
  • The Northwest_African component diverges in the direction of Sub-Saharan Africans.

Technical Details


A dataset of 268 populations/3,115 individuals was assembled. A total of 265,519 SNPs are in common in the various source datasets as well as the 23andMe v2/v3 and Family Finder platforms. Iterative removal of distant relatives was performed by removing one individual from each pair within a population if that pair had a RATIO of 2.5 or greater or more than the mean and two standard deviations in IBD analysis performed in PLINK 1.07. A total of 2,675 individuals remained. 4 individuals were removed for low genotyping rate (less than 97%). 264,328 SNPs remained after removal of SNPs with less than 97% genotyping rate or 1% minor allele frequency. 166,770 SNPs remained after linkage-based disequilibrium pruning (--indep-pairwise 200 25 0.4). The final set thus consisted of 2,671 individuals/268 populations/166,770 SNPs. Ancestral populations (components) were inferred using ADMIXTURE 1.21, with K=7 and K=12 and default parameters.

No individuals were removed from the source datasets, except in the case of the Armenians_Y sample, where one individual (ID: armenia3) was dropped because he/she was the same as a Dodecad Project participant.

Downloads


K7b population portraits, spreadsheet, and DIYDodecad files.
K12b population portraits, spreadsheet, and DIYDodecad files.

Dodecad Oracle (K12b edition) can be downloaded from here. Please read the instructions of the previous Oracle on how to use this tool. Note that the number of populations is now 223.

To use either calculator with DIYDodecad, with your 23andMe or Family Finder data, follow the instructions in the README file, but substitute 'K12b' or 'K7b' for 'dv3'.

Project participant results for both K7b and K12b are found in the spreadsheets in the Individual Results tab.

Terms of Use


You are free to use K12b and K7b, including all downloaded files for any non-commercial purpose, as long as you attribute them to the Dodecad Project and to Dienekes Pontikos as follows:

The [K7b/K12b] admixture calculator is courtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project; more information here.

Wednesday, October 26, 2011

'eurasia7' calculator

This calculator was made with 196 different populations and 2,659 individuals, including 518 project participants. The following Dodecad populations do not have 5 individuals yet, so they are included in the OTHERS_D generic category:
Algerian_D, North_African_Jews_D, Slovenian_D, Mixed_Scandinavian_D, Danish_D, Moroccan_D, Tunisian_D, Serb_D, Austrian_D, Saudi_D, Pakistani_D, Tatar_Various_D, Palestinian_D, Greek_Italian_D, Romanian_D, Swiss_German_D, Szekler_D, Mandaean_D, Azeri_D, Czech_D, Georgian_D, Belgian_D, Latvian_D, Estonian_D, Bangladesh_D, Yemenese_D, Sri_Lanka_D, Hungarian_D, Basque_D, Udmurt_D, Egyptian_D
As always, I encourage people with 4 grandparents from the same country or ethnic group of Eurasia, North or East Africa to contact me (do not send data!) for possible inclusion in the Project. If I have overlooked any such individuals, drop me a line (my e-mail address is at the bottom of the blog). I usually start a new _D population whenever individuals with 4 grandparents from the same group are submitted, but I may have missed some.

Note that all individuals from the reference populations have also been included, including outliers; you should be aware of this when reading the population averages, and consult the Outliers tab in the v3 spreadsheet for some instances of outliers.
Due to image size restrictions in Picasa, the labels are not visible well. A large version of the above plot can be found in the download bundle.

The seven ancestral populations inferred at this level of resolution are:
  • Sub_Saharan
  • West_Asian
  • Atlantic_Baltic
  • East_Asian
  • Southern
  • South_Asian
  • Siberian
As usual, you should take these names as useful labels, and interpret them in conjunction with the components' distribution in different populations, and their Fst distances, both of which can be found in the spreadsheet.

The table of Fst distances:


Below you can see a neighbor-joining tree based on inter-population Fst distances:
The first six dimensions of a multi-dimensional scaling of the same:





Calculator Files:

  • The spreadsheet contains population averages, the table of Fst distances, and individual results for included Project participants.
  • The download RAR file (Google Docs or Sendspace) contains all the files needed to run the calculator. You must download and install DIYDodecad 2.1 first. In order to run the calculator, you follow the instructions of the README file, but type 'eurasia7' instead of 'dv3'.

Terms of use: 'eurasia7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Technical Details:

The calculator is built using allele frequencies of K=7 ancestral components inferred by ADMIXTURE 1.21 analysis of 2,659 individuals. Markers included in the source datasets, as well as the Family Finder and 23andMe (as of Oct 21) platforms were included. The marker set was thinned of markers with less than 99.5% genotype rate and less than 0.5% minor allele frequency. Linkage-disequilibrium based pruning was carried out with a window size of 250 SNPs, advanced by 25 SNPs and R-squared greater than 0.4. A total of 164,990 SNPs remained after these filtering steps.

All relevant populations available to me, and genotyped at a sufficient number of markers were included. Inclusion of the Kalash population resulted in a population-specific component at K=7, and hence their admixture components were inferred a posteriori. Their proportions are consistent with previous results, showing them to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components.

Friday, September 30, 2011

'euro7' calculator

I am releasing a new calculator for Europeans, including their immediate neighboring populations around the Black Sea (Caucasus and Anatolia). The calculator can be used with DIYDodecad

There are additional African and Far-Asian population controls, so, in principle, the calculator could be used by non-Europeans/Anatolians/Caucasians, although I would be less confident of their results. For example, people of South Asian ancestry may obtain a Far-Asian result if they use this calculator, due to the deep affinity of Ancestral South Indians with East Asians. Other West Eurasians and West Eurasian-admixed peoples, not from the studied regions (e.g., Arabians or East Africans) will have their West Eurasian components mapped onto the ones used in this calculator.

'euro7' uses 7 ancestral components:
  • Caucasus
  • Northwestern
  • Northeastern
  • Southeastern
  • African
  • Far_Asian
  • Southwestern
These names represent 7 ancestral populations inferred by ADMIXTURE, and have been chosen based on the geographical regions where each of them achieves its maximum representation. You should always refer to A note of caution on admixture estimates, Interpretation of ADMIXTURE results: component sharing, as well as the average population values in the spreadsheet when interpreting your individual results.

The distribution of these 7 components can be seen in the barplot on the top left, and precise admixture proportions can be found in the spreadsheet. Note that additional samples have been used to infer these components, but as these come from Dodecad populations with less than 5 participants, I am not reporting average values for them, as per the usual project policy.

Here is the neighbor-joining tree based on the Fst divergences between the 7 ancestral components:
Instructions:

You can download the calculator RAR from here (Google docs; File->Download original), or here (sendspace).

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'euro7' instead of 'dv3' in these instructions.

Terms of use: 'euro7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Calculators released by the Dodecad Project:

Wednesday, September 21, 2011

'weac' calculator


This new calculator places individuals on the West Eurasian cline. This cline is the first-order description of variation in West Eurasians, with populations from northern and western Europe falling on one end, and those from the Near East on the other.

On the left, you can see the populations on which the calculator is based, sorted on their average "Atlantic-Baltic" component. The raw data can be found in the spreadsheet.

Note that the main purpose of the calculator is to place European and Near Eastern samples on the West Eurasian cline, and to do so, some African and East Eurasian populations are used as controls. Other types of ancestry (e.g., South Asian or Amerindian) may register as Far-Asian in the context of this test.

You can download the calculator RAR from here (Google docs), or here (sendspace).

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'weac' instead of 'dv3' in these instructions.

Terms of use: 'weac', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Sunday, September 18, 2011

Do-It-Yourself Dodecad v 2.1

DIYDodecad v 2.1 allows incomplete genotype files to be used, i.e., genotype
files that do not include all expected SNP markers used in a calculator. This
is useful to individuals having older genotype files from their testing
companies, and allows the tool to be used with any type of genotype data, and
not only the Illumina platforms currently used by 23andMe and FamilyTreeDNA.

There is a minimum requirement of at least 100 usable SNPs, i.e., SNPs that are
in the genotype file and do not have no-calls.

If you had previously followed the instructions carefully, and got an "end of file reached" error, this was most likely due to your genotype file lacking some of the expected markers used in the calculator. Version 2.1 should work for you.

You can download it from here (Google Docs, File->Download Original), or here (Sendspace). Uncompress DIYDodecad2.1.rar to a local directory on your computer, and follow the instructions in the README.txt file.

Past versions: 2.0, 1.0

Thursday, September 15, 2011

Third-party tools based on the Dodecad Project

Gedmatch.com has made available some tools based on Dodecad v3 as bundled in DIYDodecad 2.0. In addition to the regular admixture analysis, there is a chromosome painting, and an option to compare 2 kits. This should be quite useful to Mac users, who can't use DIYDodecad at present. Gedmatch.com requires upload of your data to the server side, which provides the benefit of the other tools of the site, but may not be ideal for people with privacy concerns for whom the DIY tool was partly built.

Note that because admixture estimation is expensive computationally, the Gedmatch.com tools are slightly less accurate than DIYDodecad because of more lax termination criteria. This should not be a problem for the major components of one's ancestry, but may be for the minor ones. Moreover, convergence is achieved with a different number of iterations for different genotype files, so accuracy may vary.

Two other genome bloggers have released their own calculators that can be run with DIYDodecad. Magnus Ducatus Lituaniae has released MDLP based on its K=7 analysis. Eurogenes has released a K=14 imaginatively named test calculator for Eurasian data.

I keep a list of calculators for DIYDodecad in the DIYDodecad 2.0 page.

I neither endorse nor am I affiliated with any third-party tools, but I encourage readers to try them out; the more the merrier.

Wednesday, September 14, 2011

'africa9' calculator

I have devised a new calculator targeted specifically for Africans. Admixture proportions in the reference panel, Fst distances between the K=9 components, as well as individual results for Project participants from the North_Africa_D, North_African_Jews_D and East_African_D populations can be seen in the spreadsheet.

The calculator combines data from Henn et al. (2011), HGDP, and Behar et al. (2010). As a result, the number of SNPs is small: there is probably noise in the minor components, but the major components of one's ancestry should be well-defined.

It should be used only by Africans and African-West Eurasian admixed individuals. It is not meant for people with additional admixture (e.g., South/East Asian or Native American).

You can download the calculator RAR from here (Google docs), or here (sendspace).

You need to extract the contents of the RAR file to the working directory of DIYDodecad. You use it by following exactly the instructions of the DIYDodecad README, but always type 'africa9' instead of 'dv3' in these instructions.

Terms of use: 'africa9', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

NB: Note that the components of 'africa9' do not necessarily have the same meaning as the same-named components you might have seen elsewhere. Refer to the spreadsheet for the admixture proportions and Fst distances between components. For example, the NW African is substantially removed from other West Eurasian components in Dodecad v3 but equidistant from Europe and SW_Asia in 'africa9'. I also advise that you read Interpretation of ADMIXTURE results: component sharing

Monday, September 5, 2011

'bat' calculator (Balkans-Anatolia-Turkic)

I have decided to make a new calculator for DIYDodecad that may be useful for individuals from the Balkans and Anatolia. You can download it from here at Google Docs (or here from sendspace). The terms of use are the same as for DIYDodecad v 2.0. To run it, you simply extract the contents of the RAR file in your working directory, and type bat.par whenever you typed dv3.par in the instructions.

The reference populations can be seen below. I have included all available Balkan populations, as well as Turks and Armenians. Moreover, I have included all available Turkic populations.

The marker set is the same as used in Dodecad v3. Three components emerge: one centered in the northern Balkans, one in eastern Anatolia, and one present in various proportions among all Turkic populations (see Turkic cline).


The components have been named accordingly, but please note that they do not necessarily reflect recent ancestors. For example, it is a good hypothesis that the Anatolia component was present in the Balkans even in ancient times, so one need not seek a recent Anatolian ancestor to explain its presence in a Balkan individual. Similarly for the Balkans component in Anatolia, which may reflect the diverse Balkan peoples that have settled in Anatolia since the dawn of history, so a present-day inhabitant of Anatolia need not seek a recent Balkan ancestor.

Likewise, the Turkic component is only part of the genetic makeup of the Turkic speakers who arrived in Anatolia, since those probably also carried West Eurasian population elements picked up en route from Siberia to Anatolia.

The way to interpret your results is to see whether you have an excess or deficiency of any component relative to your ethnic group. For example, an Anatolian Greek may have a higher Anatolia/Balkans ratio than a Balkan Greek and likewise for a Balkan vs. Anatolian Turk; the latter may also have a variable Turkic component which will reflect differential Central Asian input.

Saturday, September 3, 2011

No affiliation to Gedmatch.com

A thread at the 23andMe forum suggests that the Dodecad project is somehow related to Gedmatch.com. Apart from the fact that the admin of Gedmatch.com, for whatever reason, chose to co-opt the exact names of the 12 ancestral components of the Dodecad project for his own purposes, creating unnecessary confusion between the two projects.

I would like to state that I have absolutely no relation whatsoever to that outfit, and that use of any DIYDodecad files without my permission is forbidden without attribution, as stated in the DIYDodecad README file.

UPDATE: The admin of Gedmatch.com has kindly asked for permission, and I've replied that he is entirely free to use the DIYDodecad materials, provided that:
  1. It is clearly visible to someone considering using the utility that it is based on Dodecad v3
  2. It is also clear that the results may differ from the same-named components of Dodecad v3 produced by the Dodecad Project
I consider the misunderstanding to be over, and hopefully the tool will be online again soon.

Tuesday, August 23, 2011

How to make your own calculator for DIYDodecad

As I have explained in the README file of DIYDodecad, it is possible to use the software to create and distribute new calculators, based on different marker sets/ancestral populations.

(The following discussion will only be useful to other genome bloggers, or people who have experience with ADMIXTURE software).

Currently, DIYDodecad is distributed together with the 'dv3' calculator ("Dodecad v3"). This consists of a set of files:

dv3.par (The parameter file that tells DIYDodecad what to expect and what to do)
dv3.alleles (Allele names and variants)
dv3.12.F (Allele frequencies for 12 ancestral populations)
dv3.txt (Names for 12 ancestral populations)

I will now explain how you can use PLINK and ADMIXTURE to create your own calculator.

(1) Running ADMIXTURE

In the following discussion, I will assume that you have your dataset in binary PLINK format (bed/bim/fam files), that it has 123,456 markers, and you run ADMIXTURE regularly for 7 populations, e.g.:

./admixture test.bed 7

CAVEAT! The 123,456 markers must be included in the commercial platform you are targeting your calculator for. So, before you run ADMIXTURE, you must make sure that test.bed includes only markers for your chosen platform (e.g., 23andMe v3). I will assume that you have the list of markers from your commercial platform in a file (one per line), e.g., 23andMeV3.txt. You must then first do:

./plink --bfile test --extract 23andMeV3.txt --make-bed --out test

You can repeat this with other commercial marker sets, so that in the end your "test" dataset on which you run ADMIXTURE only has commercially available markers that your targeted audience will possess in their genotype files.

Actually, my main personal working sequence is to:
  • Merge (--merge-list) all reference datasets in PLINK with a --geno flag
  • Extract (--extract) commercial markers that form the intersection of 23andMe v3/v3 and Family Finder (Illumina)
  • Do linkage-disequilibrium based pruning (--indep-pairwise)
  • Finally run ADMIXTURE
It's better to do LD-based pruning after commercial marker pruning, since doing it in reverse may disrupt the physical spacing of the markers identified by --indep-pairwise.

After ADMIXTURE finishes its run, it will output a file called test.7.P; this is the allele frequencies file that you will use for your calculator, but you have to modify the order of the alleles! We will do this later.

(2) Preparing the test.alleles file

First, run the following command:

./plink --bfile test --freq --out test

This will produce a test.frq file which will be the basis of the dv3.alleles file. In R, do the following:

X<-read.table('test.frq', header=T)[, 2:4]

This will basically load the SNP names and minor/major alleles into the X table. We now identify the alphabetical order of the SNPs:

ORDER <- order(X[,1])

And, now we re-order X, so that SNPs are ordered alphabetically:

X <- X[ORDER,]

and, we save this as the test.alleles file

write.table(X, file='test.alleles', quote=F, row.names=F, col.names=F)

(3) Preparing the test.7.F file

The test.7.P file can be prepared as follows:

X <- read.table('test.7.P')
X <- X[ORDER, ]
write.table(X, file='test.7.F', quote=F, row.names=F, col.names=F)

Note that in this example test.7.P contains the output of ADMIXTURE, and test.7.F will contain the same output, but with rows re-ordered in the same way as the test.alleles file.

(4) Preparing the test.txt file

You do that with an editor; just pick whatever names you want for your 7 ancestral populations, which, of course, should be in the same order as the corresponding frequency columns output by ADMIXTURE.

(5) Preparing test.par file

Again with your editor, for this example:

1d-7
7
genotype.txt
123456
test.txt
test.7.F
test.alleles
verbose
genomewide

(6) Instructions to users

Do NOT distribute the DIYDodecad software itself, rather direct your users to the Dodecad Project download page (e.g., here, for the current 2.0 version of the software). This will ensure both compliance with the terms of use of the software, and also that users have access to the most up-to-date version.

You only have to distribute test.par, test.alleles, test.7.F, and test.txt.

Your users will follow exactly the same sequence of actions as described in the Dodecad README.txt file, with the only difference that they should type 'test', rather than 'dv3' whenever it is needed.

Hopefully more genome bloggers will decide to release calculators based on their ADMIXTURE runs to the wider public. There are several reasons to do this:
  • Reduced workload
  • Wider distribution of your work in the community, since, due to privacy concerns, not everyone is willing to share their data
  • Ability to study the utility/validity of inferred components on test data and by persons other than the discoverer
  • Ability to use the advanced bychr, byseg, and target modes with your calculators

Friday, August 19, 2011

A few comments on the use of DIYDodecad 2.0

Here are some observations that might be useful to people, especially for the new byseg and target modes:

1. Finding the origin of shared segments

Until now, when you had a segment match with another customer in your testing company, you had no idea what was the origin of the shared segment. Suppose, for example, that a Russian and a German share some sequence in a region X. This could be:
  • Russian-like ancestry in the German individual
  • German-like ancestry in the Russian individual
  • Third party ancestry in both individuals
Using the new modes, if the German saw an excess of Eastern European (relative to his usual average), then he'd pick the first scenario; if he saw nothing unremarkable, the second; if an excess of some component rare in both Russians and Germans (e.g., West_Asian), the third.

This is extremely important, as there is a noticeably confirmation bias in some individuals of interpreting the unusual as evidence of exotic ancestry. For example, an individual in search of Jewish ancestry may interpret segment matches with Jews as evidence for that ancestry: if he sees high Southwest_Asian ancestry in such segments, then that's a reasonable interpretation, but the shared segments could very well be interpreted as non-Jewish ancestry in the Jewish individual, if, they happen to be, e.g., East_European.

2. With parents' DNA

It is important to remember that each region includes both paternal and maternal DNA and you got a random draw of the segments inherited by their parents (your grandparents).

So, if you try to figure out where your region X came from, remember that it came from two places. So, if you see an unusual combination (e.g., Northeast_Asian + Northwest_African) that doesn't correspond well to any known population, this may mean that you got half of it from one parent, and the other half from the other.

Note also, that while on genomewide analysis a child's results will often be intermediate (but not necessarily so) in his ancestral components between his parents, this is not the case when looking at small segments. Suppose parent A is 50% West_Asian and 50% Mediterranean in a particular region, and parent B is 50% West_Asian and 50% West_European in the other region.

Then the child may end up with West_Asian near 100% in that region (if he happens to inherit the West_Asian segments from both parenets) or near 0% (if he happens to inherit the Mediterranean/West_European ones).

3. With Dodecad Oracle

In general, I discourage the use of Dodecad Oracle with chromosome or segment results. For two reasons:
  • Small segments may appear more mixed than they are, because there may not be any informative SNPs in a particular region to distinguish between some of the ancestral components. So, the scale of the noise may be higher. As an experiment, you can average your segments, weighted by either the number of SNPs or their physical length, and you will come up with something close to your "genomewide" average, that will, however, be off, because of this factor.
  • From a different perspective, segments may appear less mixed, because it is less likely that you got genetic material from all ancestral populations in a small section of your DNA. Your genomewide admixture may have several non-zero components, but you are unlikely to have many non-zero components in a small region (barring the aforementioned noise), and you could very well see >80% percentages in some of them that are very typical of a particular ancestral component.