Showing posts with label Jews. Show all posts
Showing posts with label Jews. Show all posts

Saturday, January 21, 2012

fastIBD analysis of Afroasiatic groups (Jews, Arabs, Assyrians, Berbers, Somalis, Amharas, etc.)

Please refer to the previous analysis on the Balkans/West Asia for more information about the interpretation of this type of analysis.

I am very pleased with the way this analysis of Afroasiatic groups has turned out, revealing an exceptional degree of resolution. I invite individuals from the Near East and Africa who are eligible, to submit their data, so that they can be included in future runs of this kind.

Clusters Galore


45 clusters were inferred with 29 dimensions.


I can't comment on all 45 clusters, so I'll just limit myself to the ones that are significantly represented among Project participants: 1. Ashkenazi, 4. Assyrian/Mandaean, 6. Somali, 7. Moroccan, 8. Algerian/Tunisian, 9. Sephardic, 10. Morocco Jews, 11. Iran/Iraq Jews, 12. Non-Jewish Ethiopians, 13. Saudi, 14. Arab #1, 15. Arab #2, 16. Egyptian

Inter-Population IBD


Results for Project Participants


The results can be found in the spreadsheet.

I have also added the full IBD sharing matrix which lists how many Morgans of sequence are estimated to be IBD with probability greater than 10^-6 between all pairs of individuals.

You can google any non-Project sample IDs to get some more information about their origin. For example, GSM536710 is an Iraqi Jew who shares about half his genome with GSM536714, also an Iraqi Jew. These two samples are almost certainly first-degree relatives. Or, GSM537032, a Samaritan shares 740-1,480cM with the other 2 Samaritans, an exceptional amount in this small and probably highly inbred population.

You can manipulate this matrix in R. After you download it and unzip it, you can load it into R as follows:

X<-read.table('afroasiatic_ibd_sharing.txt',row.names=1,header=T)

Then, you can, for example, sort the IBD sharing for a particular individual, as follows:

sort(X['DOD026',])

Tuesday, January 17, 2012

fastIBD analysis of Iberia, France, Italy, Balkans, Anatolia and European Jews

On the heels of the previous analysis of Balkans/West Asia, a new experiment on a different set of populations. Please refer to the earlier post for some thoughts/explanations about this type of analysis, I'll stick to "just the data" for this post.

Clusters Galore




24 clusters inferred with 17 MDS dimensions.

The Galore analysis provides increased resolution within Iberia (#6-9, 11), Italy, and the Ashkenazi Jewish group (#14-16).

The Iberian results are particularly interesting, showing the power of this approach compared to the one with unlinked data. There appear to be:

  • a Spanish Basque (#6), 
  • French Basque (#11) cluster, as well as 
  • a Portuguese/Galician/Castilla Y Leon (#9) cluster, and 
  • a complementary Castilla La Manch/Cantabria/Andalucia/Murcia (#7) cluster, and 
  • a smaller Aragon/Cataluna cluster (#8). 
There is overlap between these clusters, but the geographical contrasts are quite evident. I did not go through the results of Spanish Project participants (all the Portuguese fall in the Galician cluster, and our Basque member in the Basque cluster as expeccted), so it would be interesting to hear whether they fall in the cluster(s) which exist in their regions of origin.

Inter-Population IBD




Results for Project Participants


The results can be found in the spreadsheet.

Saturday, January 8, 2011

ADMIXTURE analysis with Dodecad Populations (update #2)

Thanks to all the participants of the Project, the number of populations has increased, and so have sample sizes within pre-existing populations in the Project. There are now 17 populations with at least 5 individuals in the Project:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian, Russian, Spanish, British, Irish, Turkish, N_Italian, Balkans
Below are the K=10 ADMIXTURE results with these populations:

Admixture proportions can be found in the spreadsheet.

The fact that the addition of 17 populations and 143 individuals to the core set of 36 populations and 692 individuals results in the same 10 ancestral components testifies to the stability of this solution. Hopefully, within 2011 I will develop an even better comparison set to work with.

Another test of the validity of the analysis is comparison of independent samples of the same populations:
Ashkenazi, Armenian, Spanish, Turkish, N_Italian
I have a sample of Dodecad Project members for each of the above, as well as a published population. A way to measure the concordance between the two is to calculate the correlation coefficient (rounded to the 3rd decimal point):
  • Ashkenazi Jews: 0.999
  • Armenians: 0.988
  • Spanish: 0.998
  • Turkish: 0.995
  • N_Italian: 0.996
The concordance is remarkable.

I have also made a RAR of "population portraits". It is important to do this to determine whether minor ancestral components represent population-wide phenomena or are limited to a few individuals.

For example, here are the Turks of the Dodecad project:
The sample is a bit more varied than the sample included in Behar et al:
This probably underscores the importance of broad coverage of large countries and ethnic groups, as I have discovered recently in my analysis of 9 different populations of Pakistan.

Another new population are the Irish, presenting a picture of remarkable homogeneity:
Here is the population portrait for the Balkans, which consists of non-Greek, non-Roma inhabitants of the Balkans:
This appears quite varied; hopefully more Balkan project participants will allow me to split this into additional sample populations.

Finally, here is a portrait of the Ashkenazi population, which appears quite similar to the Behar et al. one:
A very interesting thing about this population is the existence of small slices of "East Asian" and "Northeast Asian" components totalling about 1.5% in almost all individuals. In my opinion this testifies to some type of old minor absorption, as it is fairly evenly spread in the population.

If you haven't joined the project yet, feel free to submit your sample during this opportunity.

Friday, November 19, 2010

ADMIXTURE analysis with Dodecad Populations (update #1)

Repeating the previous analysis with additonal populations of Dodecad Project members and/or modified sample sizes for pre-existing ones:
Assyrian, Scandinavian, Greek, Finnish, S_Italian_Sicilian, Ashkenazi, German, Indian, Portuguese, Armenian
Admixture proportions can be found in the spreadsheet. Dodecad Project populations in italics.

Populations portraits can be found in the RAR. For example, here are the ones for Dodecad Project Ashkenazi and Behar et al. (2010) Ashkenazi Jews:
and here is a Portrait of the Portuguese:

Monday, November 8, 2010

Portraits of the populations

Averages give an overview of different populations' makeup, but they may mask important variation. Thus, I have decided to present the individual-variation of the various populations included in the Project. You can get a RAR file with all populations here.

Below, I will highlight a few ones. I apologize about the quality of the graphics, but these were generated by a script I wrote which doesn't seem to have the presentation controls of normal "Save As".

First, the Saudis:

You can see, for example, that the West African admixture in this population comes pretty much from 3 individuals, with one of them having a substantial chunk of it. Some individuals have South European or West Asian admixture, while many belong 100% to the Southwest Asian component.

Next, the Ashkenazi Jews:

You can see that some individuals have an excess of the North European component, while others have almost none.

Next, the Burusho:

These appear very homogeneous, with what variation we might expect from random hereditary processes or the limitations of admixture inference.

Next, the Romanians

I've commented on the presence of 2 probable Roma individuals in this population before, and this seems quite evident in this plot.

Finally, the Gujarati:

It is evident that the North European and East Asian element is limited to a few individuals, with most of them appearing to be a "South Asian" and "West Asian" blend. One might speculate that these represent individuals admixed with other subcontinental populations where these two components are more important.

Tuesday, November 2, 2010

Analysis of Greeks, Italians, Cypriots, and European Jews


Here are the first two dimensions of a multidimensional scaling in which the following are included:
  • Behar et al. (2010) Cypriots, and Ashkenazi/Sephardi Jews
  • HGDP-CEPH North Italians and Tuscans
  • Dodecad Project Greeks and Part-Greeks (N=7)
  • Dodecad Project Italians (N=14)
  • Dodecad Project Ashkenazi Jews (N=2)
The following color coding is used: Greeks: blue; Italians: green; Cypriots: magenta; Sephardic Jews: grey; Ashkenazi Jews: red.

The Jewish/non-Jewish groups can be distinguished by the first axis, with value greater than 0 including almost all Jews and less than 0 including almost all non-Jews. A group of Sephardic Jews, while still distinguishable, nonetheless approaches the non-Jewish individuals. According to Behar et al. (2010):
Samples from the Bulgarian and Turkish communities were combined in the analysis and are referred to as Sephardi Jews.

The separability of Ashkenazi Jews from non-Jewish Europeans had been previously reported, e.g., by Tian et al. (2008), from which the figure on the left is presented.

The current post adds Sephardic Jews into the picture; these are also distinguishable from non-Jewish Europeans, but appear to be somewhat closer to them -- at least in the first 2 dimensions of an MDS plot.

While the MDS plot gives a nice overview of this distinctiveness, to assess the possible admixture between these various groups, we must turn to ADMIXTURE.

Indeed, there are two groups: dark blue (European Jews), and light blue (Cypriots, Italians, Greeks). I don't want to put too much emphasis on this, as there are obviously many factors at play, e.g., (i) admixture with other elements, not represented in this study (e.g., Spaniards or Anatolians for Sephardi Jews / North-Central Europeans for Ashkenazi Jews / Syrians, or Phoenicians, or Arabs for Southern Europeans), or (ii) inability to detect admixture accurately among closely related groups (the two clusters are at an Fst=0.029); previously I had seen some residual admixture among Tuscans and White Americans

Nonetheless, the ADMIXTURE analysis confirms that (i) non-Jewish Greeks (including Cypriots) and Italians are distinguishable from European Jews, and (ii) Sephardi Jews are closer to these Europeans than Ashkenazi are.

PS: Idea for this post by recent Italianthro post.