Showing posts with label Experiments. Show all posts
Showing posts with label Experiments. Show all posts

Friday, April 27, 2012

Estimating your Gök4-related ancestry


I have taken Table S15 from Skoglund et al. (2012), and the Dodecad Project K7b admixture proportions in order to investigate possible relationships.

In Table S15 the authors estimate the Neolithic farmer ancestry in several populations on the basis of a single Neolithic individual from the Funnel Beaker (TRB) culture which was found in a megalithic burial in Gökhem parish.

Most of these populations are already part of the Dodecad Ancestry Project, except the three Swedish samples; given the intermediacy of the Central_Sweden sample, I have decided to use my Swedish_D sample of Project participants as a stand-in for it.

Below, you can see a scatterplot relating Gök4-related ancestry with K7b "Southern" component:



The correlation between the two variables is very strong (R-squared = 0.93).

Dodecad Project participants who already have K7b results, as well as customers of DTC testing companies who can use DIYDodecad together with the K7b calculator can approximately estimate their Gök4-related ancestry by plugging in their "Southern" value (in %) into the following equation:

Gök4-related ancestry = 1.721*Southern+19.736

I anticipate that when I am able to study the Neolithic Swedish genomes directly, the Neolithic farmer from Sweden will turn up "Southern" in a K=7 resolution experiment.

Wednesday, February 15, 2012

Correspondence between ChromoPainter clusters and ADMIXTURE components in Balkans/West Asia

I took the 25 different inferred clusters from my recent ChromoPainter analysis, and calculated their normalized median components in terms of the K12b calculator. This is a quite useful exercise, since it can show in what sense clusters are different from each other.




Here are two ways in which you may use this correspondence.

1. Different clusters of a single population

For example, the Turks with partial Balkan ancestry tend to belong to pop10, whereas those of Anatolian ancestry to pop13, and those from northeastern Anatolia to pop22. If we compare the admixture proportions of these three groups, we notice e.g.,

  • An excess of Atlantic_Baltic and North_European in pop10
  • An excess of Caucasus in pop22
Or, there is a group of 5 Iranians that belong to pop12, whereas the overwhelming majority of Iranians and Kurds belong to pop21. Strikingly, pop12 differs from all other populations in having substantial levels of East_African and Sub_Saharan. So, it seems that fineSTRUCTURE was able to infer that some Iranian individuals had this feature in common. These individuals were already evident in the Iranian population portrait (right), but fineSTRUCTURE was able to group them even though there were no African populations in the ChromoPainter analysis; presumably, the software was able to detect that these individuals shared a set of chunks that were quite different than is the norm for the Balkan/West Asian area.

2. Related clusters


fineSTRUCTURE grouped the different populations in a tree structure. For example, it grouped pop18, the "North Balkan" cluster with pop23, the "Bulgarian-Romanian" one.

Looking at the admixture proportions, we can tell that the two clusters do indeed seem quite similar, but there are some differences, e.g., an excess of North_European in pop18, and an excess of Caucasus in pop23. This makes sense given the geographical origin of individuals belonging to the two clusters.

Monday, October 31, 2011

Origin of Kalash inferred with Eurogenes K=10 "test" calculator

Vasishta, is asking Eurogenes for help in demonstrating that the Kalash have Northern-European-specific segments:
Yes. He keeps citing the Kalash as proof that the Indo-Iranians were an almost exclusively a West-Asian like population, even though I personally think the mainly West Asian-South Asian assortment of the Kalash in his analyses might be an artifact of their inbreeding and isolation, thus confusing ADMIXTURE. Zack's K=11 at Harappa has shown that the Kalash display around 22% of the component modal in Lithuanians. Yet, he ignores the North/Eastern European admixture in Northwest Indians and North Indian Brahmins (in his own analyses at that!). Interestingly enough the aforementioned groups tend to score a sliver of Northeast European admixture in Dr.Doug McDonald's analyses, with the top matches for that sliver usually being Lithuanians, Russians and Finns; in that order. It (NEU) is even found in frequencies of around 4-6% in Dravidian-speaking southern Brahmins. As much as I hate to say it, he is indeed rather stubborn and has somewhat of an underlying agenda.

David, I think you should look into proving that the Kalash do indeed have some NEU-specific segments. I would be super-surprised if they didn't, given that more mixed populations south of their geographical area display it themselves.

It appears that Vasishta disagrees with me because he "personally thinks" that the admixture proportions of the Kalash are due to inbreeding and the limitations of ADMIXTURE.

Does he cite any studies or make any argument why ADMIXTURE would remove precisely the component that he is so eager to be present? No. While genetic drift in an isolated population could indeed lead to the loss of genetic diversity, there is no reason to think that this would lead preferentially to the loss of Northern-European segments. It is strange that Vasishta accuses me of bias and yet, at the same time, invokes the magic of some unspecified flaw of ADMIXTURE for the loss of his favorite component.

Vasishta invokes the Harappa Ancestry Project K=11 admixture analysis in support of his idea that the Kalash have 22% of the component modal in Lithuanians. However, he neglects to mention that at K=11 there is no West-Asian or Caucasus centered component in the HAP analysis, but rather only "European" (modal in Lithuanians) and "SW Asian" (modal in Yemen Jews). It is indeed strange that he accuses me of bias for providing evidence about the relationship of the Kalash with West Asia, while at the same time, showing preference for a level of analysis where such a component is lacking.

The West Eurasian cline between Arabia and Northeastern Europe is evident in the 'weac' admixture analysis, where the European-centered component (Atlantic-Baltic) is present in populations such as Assyrians and Armenians whereas it is lacking at the appropriate level of resolution. Therefore, the fact that the Kalash show "European" admixture at the level of Europe vs. Near East does not mean that they ought to show such admixture at the level of Europe vs. West Asia/Caucasus vs. Arabia.

One of the benefits of DIYDodecad has been the availability of data from projects that have hitherto been black boxes. In the interest of transparency, I have taken the Eurogenes K=10 "test" calculator and repeated my analysis of the Kalash, that had been previously shown by me to be a fairly simple West/South Asian mix. I could have waited for him to get around to it, but since he's quick on the talk and slow on the trigger, I decided to do it for him.

The admixture proportions of the Kalash, according to the Eurogenes K=10 are: 40.3% S_Asian, 58.7% W_Asian, 0.9% N_E_Euro, 0.1% N_Asian, and hence the analysis based on the Eurogenes K=10 components confirms the analysis based on my eurasia 7, "showing the Kalash to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components."

Eurogenes alleges, not without his usual charm, that:
Dienekes has a keen eye for things he wants to see. But he hasn't yet noticed that in all accurate analyses, there's significant Eastern European admix in North India. His monocle got fogged up in that instance.
Let us consider some pertinent facts: the Indian peninsula has been invaded multiple times from Central Asia, a process that continued long after the establishment of the Indo-Aryans during the 2nd millennium BC. Eurogenes may want to think that the "Eastern European" admixture in South Asia dates to his mythological Polish Indo-Europeans, galloping across the steppes on their horses, but there is, at present, no particular reason to think that this is the case

Furthermore, the "West Asian" component as a fraction of the "West Asian" + "Atlantic-Baltic" component reaches a minimum of 77% in the Pathans in populations from the northern parts of the Indian subcontinent. His own monocle is surely in greater need of de-fogging if I miss the 23% and he misses the >77%.

Indeed, the Europe vs. Caucasus ratio in Indian subcontinental populations is similar to that found in people from the Middle East and Caucasus region. It is not surprising that Eurogenes has abandoned his search for North European components in South Asia, going as far as reconstructing Ancestral North Indians as "Northern Europeans". Needless to say, he was wrong. The West Eurasian ancestry of the population of the Indian subcontinent is similar to that found in modern West Asian populations, not Slavs.

Eurogenes promises:
This shouldn't be too difficult. I'll use Dienekes' calculator for the job, and then check the results with LAMP.How poetic.
Been there, done that. It will be fun to see what "Northern European" components he will be able to squeeze out of the 0.9% N_E_Euro component that my software, in conjunction with his "test" calculator produces.

Why are the Kalash important?

There are three reasons why the Kalash are important in the study of Eurasian prehistory:
  1. Their mountainous habitat contributed to isolation and relative immunity from historical population movements
  2. Their non-Islamic religion has definitely preserved them from recent gene inflow
  3. Their language is unique within the Indo-Aryan family, and it often considered today as part of a separate Dardic family of Indo-Iranian in addition to the more populous Iranian and Indo-Aryan families.
The Kalash are crucial for those interested in the origins of Indo-Iranians, and the fact that they are, indeed, a simple West/South Asian mix is not without significance for that question.

UPDATE:

Here is the result of a PCA analysis of the Kalash together with 50 synthetic individuals from each of the S_Asian, W_Asian, and N_E_Euro components of Eurogenes K=10 "test". This was calculated with smartpca with numoutlieriter set to 0.

It is evident that the Kalash appear to fall on the S_Asian to W_Asian line, and toward the W_Asian pole, consistent with being a population of those two origins, with the W_Asian component predominating.

UPDATE II:

As mentioned in the eurasia7 post, the Kalash tend to form population-specific components in ADMIXTURE analyses, so they are generally not included in my runs. So, I run the K=7 analysis again, but this time I included the Kalash. Here are the top populations of the component that was modal in the Kalash:

[186,] "Kurd_D" "50.2"
[187,] "Kurds_Y" "50.7"
[188,] "Armenian_D" "50.9"
[189,] "Armenians_Y" "51.2"
[190,] "Adygei" "51.5"
[191,] "Chechens_Y" "53"
[192,] "North_Ossetians_Y" "53.2"
[193,] "Lezgins" "54.4"
[194,] "Georgians" "59.8"
[195,] "Georgian_D" "60.1"
[196,] "Abhkasians_Y" "60.5"
[197,] "Kalash" "63.2"

Here are their exact admixture proportions in this unsupervised ADMIXTURE run:

Kalash N=23
East_Asian: 0.5
Atlantic_Baltic: 1.5
South_Asian: 32.9
Sub_Saharan: 0.0
Southern: 0.0
Siberian: 1.8
West Asian: 63.2

UPDATE III (November 22): Eurogenes estimates that there is 4% "Northeast European" admixture for Kalash individual HGDP00302. He managed to avoid the creation of a Kalash-specific component by including only a single Kalash individual in an ADMIXTURE experiment.

The Kalash do tend to create their own Kalash-specific component, and a good way to avoid such a component is to include each of them individually, and repeat the analysis 23 times. An alternative, and less time consuming way, is to create a single synthetic individual using the allele frequencies of the Kalash population as a whole. Even simpler, one could randomly pick a single individual (such as HGDP00302), but at the risk of picking an individual that has either much more or much less than average a particular type of ancestry.

Below are the admixture proportions of all the 23 Kalash individuals from the unsupervised ADMIXTURE run of UPDATE II. Individual HDGP00302 is 4th of 23 in terms of their "Atlantic_Baltic" component that peaks in Lithuanians (3%). The Kalash have 1.5% "Atlantic_Baltic" on average (median=1%, standard deviation=2.1%).

ID East_Asian Atlantic_Baltic South_Asian Sub_Saharan Southern Siberian West_Asian
HGDP00279 0.007 0.081 0.361 0 0 0.031 0.521
HGDP00307 0.004 0.059 0.336 0 0 0.018 0.583
HGDP00315 0.019 0.036 0.338 0 0 0 0.606
HGDP00302 0.006 0.03 0.337 0 0 0.02 0.608
HGDP00311 0.014 0.029 0.325 0 0 0.021 0.611
HGDP00285 0 0.027 0.319 0 0 0.019 0.635
HGDP00333 0 0.02 0.324 0 0 0.018 0.638
HGDP00277 0 0.016 0.334 0 0 0.021 0.63
HGDP00298 0.012 0.016 0.325 0 0 0.016 0.631
HGDP00281 0.011 0.015 0.332 0 0 0.01 0.633
HGDP00304 0.007 0.012 0.329 0 0 0.013 0.638
HGDP00290 0.007 0.01 0.325 0 0 0.021 0.637
HGDP00274 0.007 0.004 0.341 0 0 0.013 0.635
HGDP00309 0.007 0 0.317 0 0 0.019 0.656
HGDP00330 0 0 0.335 0 0 0.026 0.639
HGDP00319 0.011 0 0.328 0 0 0.01 0.651
HGDP00288 0.004 0 0.339 0 0 0.013 0.644
HGDP00286 0 0 0.329 0 0 0.018 0.653
HGDP00313 0 0 0.351 0 0 0.015 0.634
HGDP00328 0 0 0.31 0 0 0.023 0.667
HGDP00267 0 0 0.332 0 0 0.022 0.647
HGDP00326 0 0 0.307 0 0 0.03 0.663
HGDP00323 0.002 0 0.304 0 0 0.013 0.68

Wednesday, October 26, 2011

'eurasia7' calculator

This calculator was made with 196 different populations and 2,659 individuals, including 518 project participants. The following Dodecad populations do not have 5 individuals yet, so they are included in the OTHERS_D generic category:
Algerian_D, North_African_Jews_D, Slovenian_D, Mixed_Scandinavian_D, Danish_D, Moroccan_D, Tunisian_D, Serb_D, Austrian_D, Saudi_D, Pakistani_D, Tatar_Various_D, Palestinian_D, Greek_Italian_D, Romanian_D, Swiss_German_D, Szekler_D, Mandaean_D, Azeri_D, Czech_D, Georgian_D, Belgian_D, Latvian_D, Estonian_D, Bangladesh_D, Yemenese_D, Sri_Lanka_D, Hungarian_D, Basque_D, Udmurt_D, Egyptian_D
As always, I encourage people with 4 grandparents from the same country or ethnic group of Eurasia, North or East Africa to contact me (do not send data!) for possible inclusion in the Project. If I have overlooked any such individuals, drop me a line (my e-mail address is at the bottom of the blog). I usually start a new _D population whenever individuals with 4 grandparents from the same group are submitted, but I may have missed some.

Note that all individuals from the reference populations have also been included, including outliers; you should be aware of this when reading the population averages, and consult the Outliers tab in the v3 spreadsheet for some instances of outliers.
Due to image size restrictions in Picasa, the labels are not visible well. A large version of the above plot can be found in the download bundle.

The seven ancestral populations inferred at this level of resolution are:
  • Sub_Saharan
  • West_Asian
  • Atlantic_Baltic
  • East_Asian
  • Southern
  • South_Asian
  • Siberian
As usual, you should take these names as useful labels, and interpret them in conjunction with the components' distribution in different populations, and their Fst distances, both of which can be found in the spreadsheet.

The table of Fst distances:


Below you can see a neighbor-joining tree based on inter-population Fst distances:
The first six dimensions of a multi-dimensional scaling of the same:





Calculator Files:

  • The spreadsheet contains population averages, the table of Fst distances, and individual results for included Project participants.
  • The download RAR file (Google Docs or Sendspace) contains all the files needed to run the calculator. You must download and install DIYDodecad 2.1 first. In order to run the calculator, you follow the instructions of the README file, but type 'eurasia7' instead of 'dv3'.

Terms of use: 'eurasia7', including all files in the downloaded RAR file is free for non-commercial personal use. Commercial uses are forbidden. Contact me for non-personal uses of the calculator.

Technical Details:

The calculator is built using allele frequencies of K=7 ancestral components inferred by ADMIXTURE 1.21 analysis of 2,659 individuals. Markers included in the source datasets, as well as the Family Finder and 23andMe (as of Oct 21) platforms were included. The marker set was thinned of markers with less than 99.5% genotype rate and less than 0.5% minor allele frequency. Linkage-disequilibrium based pruning was carried out with a window size of 250 SNPs, advanced by 25 SNPs and R-squared greater than 0.4. A total of 164,990 SNPs remained after these filtering steps.

All relevant populations available to me, and genotyped at a sufficient number of markers were included. Inclusion of the Kalash population resulted in a population-specific component at K=7, and hence their admixture components were inferred a posteriori. Their proportions are consistent with previous results, showing them to be a "West Asian" population (62.4%) with substantial "South Asian" admixture (37.1%), and near-complete absence of any other genetic components.

Thursday, October 20, 2011

Comparing different ADMIXTURE runs using Zombies

My idea of using zombies with ADMIXTURE is the gift that keeps on giving. Remember that "zombies" are synthetic individuals created from ADMIXTURE output, representing the K inferred ancestral components. They can be viewed as hypothetical ancestral individuals representing each of these K components without any admixture from any of the others.

An interesting problem that often comes up is to compare across different ADMIXTURE runs. I can think of at least three different applications of this:
  1. To compare components across different K; for example, how does a "West Asian"-centered component at K=5 differ from a similarly-centered component at K=12?
  2. To compare components across different datasets; for example, how does a "West Asian"-centered component inferred from an existing dataset (e.g., the current Dodecad v3) differ from a "West Asian"-centered one from a new dataset (e.g., the upcoming Dodecad v4, which will also be trained on the valuable new populations of Yunusbayev et al. 2011)
  3. To compare components across different projects; there has been a proliferation of different ancestry projects since the launching of Dodecad nearly a year ago, and since all of them slightly different individuals/SNPs/terminology, it is quite useful to be able to gauge how one component from one project maps onto other components in other projects.
As proof of concept, I took the MDLP calculator from the Magnus Ducatus Lituaniae Project and generated 50 zombies for each of its 7 ancestral components:
  1. Scandinavian
  2. Volga_Region
  3. Altaic
  4. Celto_Germanic
  5. Caucassian_Anatolian_Balkanic
  6. Balto_Slavic
  7. North_Atlantic
I then inferred the ancestry of the MDLP zombies using Dodecad v3, and vice versa. Since Dodecad v3 also includes populations (e.g., Africans) not considered by MDLP, I did not try to map those onto MDLP.


I will comment on the MDLP-to-dv3 mapping:
  1. The MDLP "Scandinavian" component appears to be West/East European with a little Mediterranean and a little Northeast Asian
  2. The MDLP "Volga_Region" component appears to be East European with some Northeast Asian
  3. The MDLP "Altaic" component is West Asian+Northeast Asian+Southeast Asian. Note that in Dodecad v3, the Northeast Asian component peaks at Chukchi, Nganasan, and Koryak, and most other east Eurasian populations have much less of it
  4. The MDLP "Celto-Germanic" component is (surprisingly) Mediterranean-dominated. One possible interpretation is that in the context of MDLP this captures one aspect of the difference between Southwestern and Northeastern Europe -higher Mediterranean in the former-, whereas the...
  5. ... MDLP "North-Atlantic" component seems to be entirely West European, and is capturing a different aspect of east-west variation in Europe.
  6. The MDLP "Balto-Slavic" appears the reverse of the "Celto-Germanic" with lower Mediterranean and reversed East/West European
  7. Finally, the MDLP "Caucassian_Anatolian_Balkanic" component is predictably mainly West Asian, but with a little Mediterranean and Southwest Asian as well
A different way of comparing the different components is to include them all in a joint MDS plot, or calculate various types of distances between them (e.g., Fst).

For example, the first couple of dimensions are dominated by the African/Asian components of Dodecad v3 that are not present in MDLP. Notice, however, the position of "Altaic", right where one might expect to find it between West and East Eurasians.

Limiting ourselves to only European populations, we obtain:

It appears that the "North_Atlantic" component may be centered on a small number of related individuals.

I encourage other genome bloggers to try their own hand at comparing their components with those of other projects, or even their own. This process will be made possible if people using ADMIXTURE follow the simple instructions to convert their output for use with DIYDodecad.

Once Dodecad v4 is off the ground, and if I find time to fully automate the process, I will perhaps try to map all my past calculators (i.e., the initial K=10, Dodecad v3, 'bat', 'euro7', 'weac', 'africa9') onto the new golden standard of the Project.

PS: This analysis was done on ~63k SNPs in common between MDLP and Dodecad v3

Tuesday, July 19, 2011

The Dodecad Oracle v1

Here is a little fun tool that tests the Dodecad v3 admixture proportions of an individual against all the reference populations, but also against the best pairwise combinations of these populations.

You need to install R to use it, and then download the program and double click on the file DodecadOracleV1.RData that can be found within the rar file. You will then be faced with a command prompt where you can enter the following commands:

Examining which populations are available

Just enter

X[,1]

You will see a list of 227 populations. You can use these population IDs in the next section.

Which populations are closest to a particular population?

Enter:

DodecadOracle("British_D")
[,1] [,2]
[1,] "British_D" "0"
[2,] "British_Isles_D" "0.9798"
[3,] "Cornwall_1KG" "1.1533"
[4,] "Kent_1KG" "2.265"
[5,] "Irish_D" "3.7643"
[6,] "Dutch_D" "4.5354"
[7,] "Mixed_Germanic_D" "6.8971"
[8,] "Norwegian_D" "11.3111"
[9,] "Orkney_1KG" "12.4652"
[10,] "Orcadian" "12.8195"

If you want to find e.g., the top-30 populations, rather than just the top-10, enter:

DodecadOracle("British_D", k=30)

Which populations are closer to a particular individual?

Enter the admixture proportions of the individual (from the "Individual results" tab of the spreadsheet) as follows:

DodecadOracle(c(4.6, 16.7, 33.6, 0, 23.2, 0.4, 0.6, 1.6, 0.7, 14.1, 4.5, 0.2))
[,1] [,2]
[1,] "Ashkenazi_D" "3.7908"
[2,] "Ashkenazy_Jews" "4.1473"
[3,] "Morocco_Jews" "6.338"
[4,] "S_Italian_Sicilian_D" "12.5443"
[5,] "Sephardic_Jews" "13.5067"
[6,] "C_Italian_D" "14.4554"
[7,] "Sicilian_D" "14.7469"
[8,] "S_Italian_D" "15.748"
[9,] "Tuscan_X" "15.9981"
[10,] "O_Italian_D" "16.1474"

Once again, you can specify k=30, if you desire the 30 top matching populations instead of the default 10.

Mixed Mode

You use mixed mode by adding mixedmode=T in any of the commands. The program then considers all pairs of populations, and for each one of them calculates the minimum distance to the sample in consideration, and the admixture proportions that produce it; population pairs where the distance to one of the two populations is smaller than to any admixture of the two are ignored.

Example:

DodecadOracle("Pathan",mixedmode=T)
[,1] [,2]
[1,] "Pathan" "0"
[2,] "84.8% Pakistani + 15.2% Urkarah" "1.075"
[3,] "84% Pakistani + 16% Stalskoe" "1.1555"
[4,] "63.9% TN_Brahmin + 36.1% Urkarah" "1.6669"
[5,] "32.4% Urkarah + 67.6% Meghawal" "2.3516"
[6,] "56.3% INS + 43.7% Urkarah" "2.4901"
[7,] "11.5% Adygei + 88.5% Pakistani" "2.6245"
[8,] "82.4% Sindhi + 17.6% Stalskoe" "2.6318"
[9,] "62.9% AP_Brahmin + 37.1% Urkarah" "2.7322"
[10,] "11.2% Lezgins + 88.8% Pakistani" "2.7749"

The mixed mode should be used with caution, and it shows, more than anything else, how similar apparent "mixes" can be achieved by different combinations of ancestry. Nonetheless, it may prove somewhat useful. For example, there is a suggestion in the above results, that Pathans can be viewed as a mix of other South Asian populations and populations from the eastern Caucasus, a suggestion that was arrived at independently by the Project using different methods.

Here is another example:

DodecadOracle("Assyrian_D",mixedmode=T)
[,1] [,2]
[1,] "Assyrian_D" "0"
[2,] "83.9% Armenians_16 + 16.1% Yemen_Jews" "1.7829"
[3,] "89.1% Armenian_D + 10.9% Saudis" "2.1624"
[4,] "84.3% Armenians_16 + 15.7% Saudis" "2.2884"
[5,] "88.9% Armenian_D + 11.1% Yemen_Jews" "2.2983"
[6,] "83.8% Armenian_D + 16.2% Bedouin" "4.1579"
[7,] "72.2% Armenian_D + 27.8% Syrians" "4.1841"
[8,] "23.4% Georgians + 76.6% Iraq_Jews" "4.2418"
[9,] "76.2% Armenians_16 + 23.8% Bedouin" "4.332"
[10,] "61.5% Armenians_16 + 38.5% Syrians" "4.4019"

This reaffirms the close relationship of Assyrians to Armenians that has been noticed in the project and by others, and it also shows that Assyrians differ from Armenians in a Southwestern Asian direction, consistent with their Semitic language.

Or, African Americans:

DodecadOracle("ASW",mixedmode=T)
[,1] [,2]
[1,] "ASW" "0"
[2,] "81.3% Hausa + 18.7% N._European" "2.3891"
[3,] "18.4% Orkney_1KG + 81.6% Hausa" "2.4031"
[4,] "18.5% Argyll_1KG + 81.5% Hausa" "2.4268"
[5,] "18.4% Orcadian + 81.6% Hausa" "2.4657"
[6,] "80.5% Igbo + 19.5% N._European" "2.5031"
[7,] "80.6% Brong + 19.4% N._European" "2.523"
[8,] "18.6% CEU + 81.4% Hausa" "2.5938"
[9,] "19.1% Argyll_1KG + 80.9% Brong" "2.6197"
[10,] "19% Orkney_1KG + 81% Brong" "2.6274"

I don't know that much about the slave trade, but I believe that Ghana was an important part of it?

Another thing to watch, is that some populations tend to have more than one sample available, so they appear to be mixtures of themselves, which is not really very informative, e.g., Spanish_D

DodecadOracle("Spanish_D",mixedmode=T)
[,1] [,2]
[1,] "Spanish_D" "0"
[2,] "7.9% French_Basque + 92.1% IBS" "0.8713"
[3,] "68.9% IBS + 31.1% Spaniards" "1.0377"
[4,] "98.8% IBS + 1.2% Irish_D" "1.2959"
[5,] "1.2% British_Isles_D + 98.8% IBS" "1.3018"
[6,] "1.2% British_D + 98.8% IBS" "1.3019"
[7,] "99% IBS + 1% Norwegian_D" "1.3046"
[8,] "1.2% Cornwall_1KG + 98.8% IBS" "1.3048"
[9,] "98.8% IBS + 1.2% Kent_1KG" "1.3142"
[10,] "2.2% French_D + 97.8% IBS" "1.3179"

To deal with these problems, you must "edit" the X matrix if you want to exclude some populations. For example, if you want to exclude "Spaniards" and "IBS", you must enter:

X <- X[setdiff(1:227,which(X[,1]=="IBS" | X[,1]=="Spaniards")),]

but notice, that you must relaunch the program, if you want to get the original matrix, or alternatively save it like this:

Z<-X

and then retrieve it like this:

X<-Z

Wednesday, June 22, 2011

Dodecad v3: population averages

With Dodecad v3 it is possible to use any population as test data in supervised ADMIXTURE analysis, and extract its admixture proportions in terms of the 12 ancestral components.

So, I have set up an automated job that will do just that: use pretty much every population available to me under only a few conditions:
  1. Each population must have at least 5 individuals
  2. It must have the same 166,462 SNPs on which the test is based
  3. It must not be from a group not covered by the test (e.g., Australo-Melanesians or Native Americans)
By my count, I have 141 different populations that meet these requirements. Each population is run on its own in a supervised ADMIXTURE analysis together with the 600-strong synthetic set (50 per ancestral component). These are ideal conditions to produce a high-quality comparative reference set.

The average admixture proportions of different populations will be put in this spreadsheet as they are calculated, which will probably take several days.

Tuesday, June 21, 2011

The design of Dodecad v3

Dodecad v2 was short-lived, as I discovered a way to improve it shortly after I announced it.

The first step was to carry out an extensive K=3 ADMIXTURE analysis of about 130 different populations and about 2,000 individuals from Europe, Asia, and Africa. Using the allele frequency results of this analysis I was able to create the most comprehensive synthetic individuals to represent West Eurasians, Asians, and Sub-Saharan Africans.

Subsequently, I carried out an analysis of East Eurasian populations using the West Eurasian/Sub-Saharan synthetic individuals as controls, as well as an analysis of Sub-Saharan populations using the West Eurasian/Asian individuals as controls.

In East Eurasia, I was able to infer the existence of two components, one centered in the extreme northeast, another in the southeast, with many other populations arrayed between these two extremes:


In Sub-Saharan Africa, the primary division was between San, Mbuti, and Biaka Pygmies (whom I have called "Palaeo-Africans") and the rest (Yoruba, Mandenka, and Bantu, "Neo-Africans"):


Now, I had four synthetic "framing populations": Neo-Africans, Palaeo-Africans, Northeast Asians and Southeast Asians, created from hundreds of individuals from several different populations:
  1. I did not have to choose a particular population (e.g., Chinese) to represent East Asia
  2. I did not have to aggregate individuals from populations with variable levels of non-East Asian admixture
I now used my South Asian populations, together with Neo-African, West Eurasian, Northeast and Southeast Asian controls to extract a South Asian specific component:



Armed with these 5 synthetic "framing" populations, I carried out a K=12 analysis with my West Eurasian, South Asian, and North/East African populations (1,247 individuals; 69 populations):

And, finally, I generated 50 synthetic individuals from each of the 12 inferred components to create a dataset of 600 individuals that will be the basis of Dodecad v3.

Below is the table of Fst divergences:



The following MDS plots show the first 10 dimensions of variation of these individuals:

Finally, here is a neighbor-joining tree of the 12 components:
(to be continued)

Thursday, June 9, 2011

Cornwall, Kent, and Orkney

I just finished extracting the regional samples (Kent, Cornwall, and Orkney) from the 1000 Genomes GBR sample, and I made a quick experiment to put them in context of other West European populations (Irish, German, Dutch, French, and Scandinavian).
I will probably try to integrate some of these to the new version of Dodecad, and try some other things, so perhaps v2 may not be the next stage of the Project. Hopefully the extra wait will be worth it.

UPDATE:

Here is also a supervised ADMIXTURE analysis with the standard K=10 components. Please note that as this is not done with the same SNPs as the standard K=10 results, they are not comparable directly to other K=10 results of the Project.

Wednesday, June 8, 2011

Dodecad v2

This is an announcement of the new generation of Dodecad ancestry analysis. In comparison to the standard K=10 used since the beginning of the Project:
  1. Participants' data are now used to enrich the set of reference populations and to help define new ancestral components
  2. Rather than choosing arbitrary reference populations, I employ a very large set of individuals to capture allele frequencies and then create synthetic individuals ("panmictic zombies") that embody these frequencies; more on this below.
  3. Results for unrelated project participants will be reported in a separate post, using my new technique of converting unsupervised ADMIXTURE runs into supervised ones. Hence, Project participants can expect to receive new K=12 results; moreover, the fact that this will be done in supervised mode means that it is no longer necessary to process samples in small batches of 10 or so. All current unrelated participants will receive their results in one go, and only future submissions will be processed in batches.
This analysis utilizes results from Project participants (populations with _D endings), as well as synthetic individuals summarizing allele frequencies of East Eurasians, Sub-Saharan Africans, and South Indians (populations with _Z endings)

The framing populations (_Z)

The following _Z populations were included:
  • Sub_Saharan_Z: Bantu, Yoruba, Mandenka, San, and Pygmies from HGDP-CEPH
  • South_Indian_Z: North Kannadi, Sakilli from Behar et al. (2010), AP_Madiga, AP_Mala, TN_Dalit from Xing et al. (2010), Bhil, Chenchu, Kurumba, Satnami, Madiga, Mala, Kamsali, Onge, Great_Andamanese from Reich et al. (2009)
  • Sino_Tibetan_Z: Yizu, Naxi, Han, Tujia from HGDP-CEPH
  • Altaic_Z: Tu, Xibo, Mongola, Daur, Hezhen, Oroqen, Yakut from HGDP-CEPH, and Evenk, Buryat from Rasmussen et al. (2010)
  • Siberian_Other_Z: Selkup, Ket, Yukagir, Nganasan, Koryak, Chuckchi from Rasmussen et al. (2010)
  • Southeast_Asian_Z: Dai, Lahu, Miaozu, Cambodians from HGDP-CEPH, Khmer-Cambodian, Thai from Xing et al. (2010), and Singapore Malay from the Singapore Genome Variation Project
The 12 inferred ancestral components

Results of the ADMIXTURE analysis defining the new K=12 components of the Project can be seen below:
Raw proportions can be found in a spreadsheet. There are also population portraits in a zip file, showing individual-level variation.

The 12 components are:
  • West_Asian
  • East_European
  • West_European
  • East_Asian
  • Mediterranean
  • Northwest_African
  • North_Eurasian
  • Arabian
  • Inner_Asian
  • Sub_Saharan
  • East_African
  • South_Indian
Once again, I have tried to make these as neutral and appropriate as possible, but don't forget that they are simply descriptive labels to aid memory. For example, the Arabian component is centered on Saudis, Yemenese, and Yemen Jews, the Inner Asian component on the Altaic synthetic population, and so on.

The Fst divergences between the 12 components can be seen in the spreadsheet and also below:

A different way of showing them is via a neighbor-joining tree. Note, however, that this is not a replacement for the Fst table above which alone fully preserves the inter-population relationships:
We can also plot the first few MDS dimensions using synthetic individuals from the 12 components; again, these capture variation only partially:



What comes next?

Hopefully quite soon, I will:
  1. Report new v2 results for all project participants
  2. Report new v2 proportions for many other populations not included here
Project members who still haven't received their results (during the ongoing submission opportunity) can expect to receive K=10 standard results, and they will receive their new v2 results later.

Monday, June 6, 2011

Panmictic zombies

I had previously developed a new way of choosing "framing" populations for ADMIXTURE analyses. Such populations are necessary in order to tease out genetic contributions from outside one's region of interest.

I proposed to create a meta-population which included a single individual from a large number of populations (e.g., East Eurasians). Use of such a meta-population has two interesting properties:
  1. It solves the problem of "which" population to choose (e.g., Han, Miaozu, She, Mongol ?) as a framing reference: the meta-population captures features of all candidate populations
  2. It avoids the generation of population-specific clusters in the "framing" individuals, as no two individuals from a single population are included!
There is, however, a problem with the technique as I first described it: it only uses a single individual from each population to compose the meta-population! Hence, it is potentially sensitive to the presence of outliers, and, in any case, it throws away most of the data.

More recently, I proposed the use of "zombies" from allele frequency data output by ADMIXTURE. These zombies are, in a sense, the opposite, of what I am trying to do here, since they represent ancestral components that exist in mixed form in present-day individuals.

Instead, we can generate "panmictic zombies" by composing a dataset of all individuals from a region of interest; we then calculate allele frequencies over the combined set, and then generate synthetic individuals based on these allele frequencies.

This technique has several advantages:
  1. It is extremely resilient to outliers, as the presence of a few outliers only shifts allele frequencies by a little, and no actual outliers are included in the "panmictic zombie" population
  2. It amortizes the full set of individuals and hence does not depend on the random sample one chooses from each population
  3. It avoids the creation of population-specific clusters
  4. It speeds up the technique I introduced for converting unsupervised ADMIXTURE runs to supervised ones substantially: populations framing the region of interest (e.g., East Eurasians, Sub-Saharan Africans, South Asians, in the case of West Eurasia) can be "folded" into a number of panmictic zombie populations a priori.
Point #4 is extremely important for practitioners:
  1. It is great not to include every single East Eurasian sample in ADMIXTURE analyses when you are trying to infer patterns of variation in Europe; this is a much better solution than the ad hoc approach adopted by some of ignoring East Eurasia altogether when studying patterns of variation in Europe!
  2. It is great not to worry after several hours of ADMIXTURE analysis whether upping K by +1 will finally produce added resolution in your region of interest, or split, e.g., Mbuti from Biaka Pygmies, which is hardly of relevance if one is trying to study East Asian or European variation
Panmictic zombies can be further fine-tuned: the allele frequencies can be calculated in many different ways:
  1. Over all individuals
  2. Averaged over all population averages (to account for different sample sizes)
  3. Weighted average over all populations (to account for different demographic sizes of source populations)
A first experiment

The following MDS plot shows a population ("Synthetic", red) generated from a sample of different HGDP East Eurasian populations.
It's important to note that while "Synthetic" appears to be closer to the Tu population, that does not mean that it is interchangeable with the Tu!

The "Synthetic" population is much more diverse, as it encompasses parts (alleles) from all the different populations of the set, that, because of the averaging process happen to coincide with the Tu in the first two dimensions of the MDS plot.

Saturday, June 4, 2011

Projecting Pakistan populations on West Eurasian PCA

In a first post I showed that ADMIXTURE output allele frequencies could be used to create synthetic individuals corresponding to the ancestral components ("zombies"), and that these artificial populations could be used for both performance, and to avoid the creation of population-specific clusters in ADMIXTURE run. I was hence, able to infer the composition of several idiosyncratic populations in terms of the K=10 components of the Dodecad Project.

In a second post, I showed that "zombies" could be created even in the absence of allele frequencies, if one had admixture proportions only for the ancestral components. I was thus able to reconstruct synthetic individuals corresponding to the ANI/ASI of Reich et al. (2009). I was further able to confirm the West Asian origin of Ancestral North Indians. In a subsequent post, I used these synthetic ANI/ASI populations on groups of Pakistan, showing the main West Asian/ANI origin of the Caucasoid component in South Asia. Moreover, I confirmed that the Ancestral South Indians are related (but distantly) to the Onge from the Indian Ocean.

In this post, I run principal components analysis on the Pakistan populations; the Hazara were excluded because of their high East Eurasian admixture. Here is the unsupervised PCA:


First, you notice that the first dimension is dominated by the Kalash, a very distinctive population because of its long-term isolation. The second dimension is dominated by a Sindhi outlier, which, if you consult a Sindhi population portrait from a previous experiment, is revealed to be of substantial Sub-Saharan admixture.

Obviously, this is no good, as our first two dimensions are not anthropologically interesting. If we are interested in learning about the origins of populations, knowing that there are a few Sindhi individuals with Sub-Saharan admixture, or that the Kalash are highly isolated is not helpful.

We can run PCA again, but this time we project populations of interest onto the PCA plot of the West Eurasian control populations:
It is fairly obvious that the populations of Pakistan fall on the South Asia-West Asia line. There are small deviations from the cline:
  • Balochis and Brahuis deviate towards the SW Asian component, which is consistent with their ADMIXTURE results.
  • The position of the non-Indo-European Burusho and Indo-Aryan Sindhi populations on either side of the cline is consistent with a little SW Asian component in the Sindhi and a little North European component in the Burusho, which pull them away from the cline in the expected directions.
Moreover, the relative position of the Pakistan populations along this cline is preserved.

Using the West Eurasian "zombies" is thus, not only useful for ADMIXTURE, but also for principal components analysis; in the latter it is helpful because:
  1. It avoids domination by very isolated/inbred populations and/or outliers
  2. It is possible to create synethic "zombie" population with absolutely equal sample sizes, hence removing a source of bias (some residual bias may persist, e.g., if one used a component centered on 5 "real" individuals to create a "zombie" population of 100, then the effective sample is not really 100)


Thursday, June 2, 2011

Ancestral South Indian (ASI) in context

I have taken the synthetic ASI population together with 25 HapMap-3 Chinese (CHB), 16 HGDP Papuans, and 9 Reich et al. (2009) Onge from the Andaman Islands to determine its relationships with other Eurasian populations.

Below is an MDS plot which shows that ASI does not appear to be particularly close to any of the other populations.

I have also ran supervised K=3 ADMIXTURE analysis that treated the ASI population as test data and CHB, Onge, Papuan as parental populations; the ASI turned out 100% "Onge", consistent with the idea that ASI is distantly related to Onge, although closer than with the other two populations.

It should be noted, however, that the similarity of ASI to Onge is not unexpected, since:
  • Onge was used by Reich et al. (2009) to infer admixture proportions of Indian Cline populations, which were (in turn):
  • used by myself to infer allele frequencies of ASI, and then:
  • used by myself to create a synthetic population of ASI individuals.
So, the Onge-ness of ASI is contingent upon the accuracy of Reich et al. (2009), but, anyway, the population of my ASI "zombies" seem to pass a second test of being reasonable standins for ASI in the sense of that paper.