Showing posts with label Romanians. Show all posts
Showing posts with label Romanians. Show all posts

Saturday, January 21, 2012

fastIBD analysis of Central/Eastern Europe

Please refer to the previous analysis on the Balkans/West Asia for more information about the interpretation of this type of analysis.

Clusters Galore


The Clusters Galore can be found in the spreadsheet. After inspection of the 23 clusters inferred with 21 dimensions, they could be described as:

  1. Mordvin
  2. East Slavic
  3. Polish-Ukrainian
  4. East Balkan
  5. Vologda Russians
  6. Lithuanian
  7. Central European (combining many groups with small sample sizes)
  8. A couple of related (?) individuals
  9. Anatolian
  10. Greek
  11. Chuvash
  12. Ossetian
  13. A couple of related individuals
  14. A couple of related individuals
  15. Balkar
  16. A couple of related individuals
  17. Chechen
  18. Kumyk
  19. A couple of related individuals
  20. Adygei
  21. Lezgin #1 (main)
  22. Lezgin #2
  23. Lezgin #3
If you belong to a population with few other participants, you might end up latching onto a cluster dominated by a bigger group. This does not mean that your population is not distinctive, only that there are not enough samples to reveal its distinctiveness if it exists.

Inter-Population IBD


Results for Dodecad Participants

Results can be found in the spreadsheet.

If you have joined the Project, please consider leaving a comment in the Information about Project samples thread. That will help others make better sense of their results, e.g., if you find that you belong in the same cluster with some other individual, you might want to know something about their origins.

UPDATE: I have added the IBD sharing matrix.See here on how to use it.

Saturday, January 14, 2012

fastIBD analysis of Balkans/West Asia

Now that I've discovered a way to boost Clusters Galore analysis even further by using fastIBD, I will start experimenting with different regional populations. This analysis took about 5 hours to complete, so it appears to be quite practical.

For my first experiment, I carry out an analysis of various populations from the Balkans and West Asia.

Clusters Galore

27 different clusters were inferred with 17 MDS dimensions. Some interesting findings:
  • For the first time there emerge a couple of clusters that appear to be quite specific to Armenians (#2 and #3). 
  • Similarly, Assyrians are broken to a few clusters that appear fairly specific to them  (#9-11)
  • Georgians are split into three clusters, one of which (#14) is linked with the neighboring Abkhasians, who in turn have their own exclusive cluster (#25)
  • The cluster modal in Greeks (#6) includes 14 of 19 Greek participants, and a few Greeks are also in the Balkan cluster (#8) and an Iranian-Turkish cluster (#4)
  • The Behar Cypriot sample also splits into two, and the few Turkish Cypriot participants link to one of them (#13)
  • The Ossetian project participant links to one of the three North_Ossetian clusters
  • The major Balkan cluster (#8) still defies resolution. I am certain, however, that structure in this cluster will be uncovered with more participation. MCLUST adapts the cluster size and shape, and a "big", inclusive cluster spanning the Balkans appears more parsimonious than smaller clusters centered on the different groups. With larger participation, I anticipate that regional structure will be uncovered in the Balkans as well.
I cannot stress the importance of participation strongly enough. When groups have more participants, it is possible to both:

  1. Discover group-specific clusters, by identifying what is common between members of groups
  2. Discover within-group clusters, by identifying what is different between members of groups
For example, the great participation of Armenians in the Project has now allowed me to discover structure within the Armenian population. It appears, that cluster #2 corresponds to a more "western" Armenian group, and #3 to a more "eastern" one, with some overlap between the two.

Inter-population IBD


You can also see a visual representation of inter-population IBD:

I have only included populations with 5+ participants in this representation. Reddish shades express high IBD sharing; bluish ones low one. The heatmap has been scaled by row.

As you might expect, values across the diagonal are "reddish", since individuals within populations tend to have high IBD sharing with each other.

A few features "pop out" of the screen. Going from top to bottom:
  • Intra-Iranic sharing
  • Intra-Armenian sharing
  • Intra-Balkan sharing
  • Georgian-Abkhaz sharing
You can probably get more out of the figure, but these appear to be the most salient features.

Results for Project Participants


The results can be found in the spreadsheet, and include:
  • Probabilities of assignment in each of the 27 clusters of the Clusters Galore analysis
  • Z-scores of IBD between each individual and each of the 20 populations with 5+ participants. Higher values mean more IBD sharing. Note that Z-scores have been calculated for each row, hence each participant must scan his own row to find populations with an excess (+) or deficiency (-) of IBD sharing, and people should not compare across different rows.
Last but not least, I want to remind new project participants to leave a message in the Information about Project samples thread. Your comment will not appear immediately, since comment moderation is on, and also note that there are multiple pages of comments. 


If you haven't joined the Project yet, I encourage you to do so if you are eligible.

Monday, September 5, 2011

'bat' calculator (Balkans-Anatolia-Turkic)

I have decided to make a new calculator for DIYDodecad that may be useful for individuals from the Balkans and Anatolia. You can download it from here at Google Docs (or here from sendspace). The terms of use are the same as for DIYDodecad v 2.0. To run it, you simply extract the contents of the RAR file in your working directory, and type bat.par whenever you typed dv3.par in the instructions.

The reference populations can be seen below. I have included all available Balkan populations, as well as Turks and Armenians. Moreover, I have included all available Turkic populations.

The marker set is the same as used in Dodecad v3. Three components emerge: one centered in the northern Balkans, one in eastern Anatolia, and one present in various proportions among all Turkic populations (see Turkic cline).


The components have been named accordingly, but please note that they do not necessarily reflect recent ancestors. For example, it is a good hypothesis that the Anatolia component was present in the Balkans even in ancient times, so one need not seek a recent Anatolian ancestor to explain its presence in a Balkan individual. Similarly for the Balkans component in Anatolia, which may reflect the diverse Balkan peoples that have settled in Anatolia since the dawn of history, so a present-day inhabitant of Anatolia need not seek a recent Balkan ancestor.

Likewise, the Turkic component is only part of the genetic makeup of the Turkic speakers who arrived in Anatolia, since those probably also carried West Eurasian population elements picked up en route from Siberia to Anatolia.

The way to interpret your results is to see whether you have an excess or deficiency of any component relative to your ethnic group. For example, an Anatolian Greek may have a higher Anatolia/Balkans ratio than a Balkan Greek and likewise for a Balkan vs. Anatolian Turk; the latter may also have a variable Turkic component which will reflect differential Central Asian input.

Tuesday, August 30, 2011

Balkan averages (August 2011)

Since my last call for more participation from the Balkans, I was able to create a new Bulgarian_D sample of 5 participants. Together with the Greek_D sample, the Balkans_D sample of non-Greek, non-Bulgarian project members, the Behar et al. (2010) Romanians, and the Xing et al. (2010) Slovenians (the latter on a smaller number of markers), we are beginning to get a better feel of genetic variation in the Balkans. There have been several other averages that have been adjusted with more participation; all of them can be seen in the Dodecad v3 spreadsheet.

The table below shows the major components (>1%) in the available Balkan populations.

The Bulgarian average as it stands seems reasonably close to the Romanian one, and is characterized by balanced West/East European components; in this balance it resembles Greeks, who, however, have lower levels of both components and higher levels of the Mediterranean/West Asian/Southwest Asian components.

Slovenians contrast with Hungarians in having reverse West/East European levels, and with their neighboring Italians in having quite a bit more of the East European component, and quite a bit less of the Mediterranean one. Bulgarians/Romanians contrast with Slavic groups from eastern Europe in having less East European and more Mediterranean/West_Asian.

Hopefully before long, more participation from the western/central Balkans (Serbs, Croats, Bosnians, Montenegrins, Albanians, Slav Macedonians) will allow us to fill more holes in our understanding of the genetic landscape of Southeastern Europe.

Monday, May 30, 2011

More Zombies: Ancestral North Indians and Ancestral South Indians reborn

In my previous post I showed how synthetic individuals corresponding to ADMIXTURE ancestral components can be created and used. This was made possible by the fact that ADMIXTURE outputs allele frequencies for its components, which can be utilized to create a population of random genotypes with the same allele frequencies.

A more difficult task is to create such "zombie" individuals when there are no allele frequencies at hand. A prime example of this is the paper by Reich et al. (2009) on the two ancestral components in Indians: Ancestral South Indians (ASI) and Ancestral North Indians (ANI). The paper provides admixture estimates for these two components in present-day "Indian Cline" groups, but no allele frequencies for these components: we only knew that ANI was closely related to West Eurasians, and ASI formed a clade with the Onge from the Indian Ocean.

Both ANI and ASI are extinct (in pure form) populations, and they are blended (in varying proportions) in modern day Indians, with highest ANI occurring in the Northwest and among upper caste groups, and highest ASI among South Indian tribal and low caste populations.

As I was thinking of ways to extend the "zombie" approach, it occurred to me that there is a fairly involved way to extract the ANI/ASI allele frequencies from the available evidence:

If f(ANI) and f(ASI) are the allele frequencies at a locus for ANI and ASI, and an admixed population P has x fraction of ancestry from ANI and 1-x from ASI, then its allele frequency is expected to be:

x*f(ANI)+(1-x)*f(ASI) = f(P)

I have marked (in bold) the known variables. Obviously, this equation does not hold in practice, because of sampling error, uncertainty in the estimation of x, as well as genetic drift that may affect the allele frequencies of the admixed population.

Nonetheless, we do not only have one equation of this sort, but 18, since Reich et al. (2009) provides ANI/ASI estimates for 18 different Indian Cline populations. We can thus fit a linear regression to recover f(ANI) and f(ASI).

This is exactly what I did; there are two important caveats:
  • because most of the Reich et al. (2009) populations are very small, f(P) is expected to be very noisy. I thus grouped the Indian Cline populations into five groups (based on increasing ANI, and making sure that each one had >15 individuals), and calculated admixture proportions (x's) and allele frequencies (f(P)'s) on these groups.
  • linear regression coefficients (the f(ANI) and f(ASI) estimates) may be less than 0 or more than 1, which makes no biological sense, so these were fixed to 0 and 1 in a few cases whenever that was the case (~5% of markers)
All of this required a bit of thinking and work, so I was very skeptical that it would work; given sampling/admixture estimation errors/limitations of regression/random creation of individuals, the whole process from input data to output "zombies" passed through so many layers, that it could very well lead to nonsense.

Nonetheless, there is power in numbers, and I was hopeful that this might work. If it did, I could have synthesized ANI and ASI populations to play with and use pretty much like regular populations in a variety of experiments.

Validation of synthetic ANI/ASI populations

I generated 25 ANI and 25 ASI individuals using the above-described method. There are 119,588 SNPs in these populations.

To validate them, I ran supervised ADMIXTURE using these ANI/ASI individuals as ancestral populations, and all the Indian Cline populations as test data. The results can be seen below:
Although the estimates for some populations (e.g., Chenchu: 31 vs. 40.7%) are substantially off, the median error is 1%, and the average error is 2.4%. Overall, it does appear that the synthetic ANI/ASI individuals are fairly good standins for their (extinct) populations.

Ancestral North Indians

I included ANI together with the 4 West Eurasian components of the Dodecad Project in an MDS plot:
Also, a neighbor-joining tree:
Putting ANI/ASI to work: Romanian Gypsies

I have previously detected 2 individuals in the Behar et al. (2010) Romanian sample that are likely to be of Roma (Gypsy) heritage. Here is a supervised admixture of the Romanian sample using the ANI/ASI components:
The previously detected individuals do possess both ANI and ASI components, indeed these are:

18.1, 15.3
16.9, 16.4

in the two individuals, which might be useful in constraining geographically the origin of European Gypsies along the Indian Cline.

Putting ANI/ASI to work: Iranians

Iranians generally show affinity to South Asians. Is this affinity related to the common Indo-Iranian background of Iranians and Indo-Aryans, or, is it, perhaps, due to the absorption of South Asian population elements during Iran's long imperial past?

The ANI/ASI components in the Iranians and Iranian_D samples are:

11.7, 7.5
12.0, 6.9

Compared to the previously described Romanian Gypsies, the South Asian component in Iranians tends to be clearly tilted towards ANI.

Monday, November 8, 2010

Portraits of the populations

Averages give an overview of different populations' makeup, but they may mask important variation. Thus, I have decided to present the individual-variation of the various populations included in the Project. You can get a RAR file with all populations here.

Below, I will highlight a few ones. I apologize about the quality of the graphics, but these were generated by a script I wrote which doesn't seem to have the presentation controls of normal "Save As".

First, the Saudis:

You can see, for example, that the West African admixture in this population comes pretty much from 3 individuals, with one of them having a substantial chunk of it. Some individuals have South European or West Asian admixture, while many belong 100% to the Southwest Asian component.

Next, the Ashkenazi Jews:

You can see that some individuals have an excess of the North European component, while others have almost none.

Next, the Burusho:

These appear very homogeneous, with what variation we might expect from random hereditary processes or the limitations of admixture inference.

Next, the Romanians

I've commented on the presence of 2 probable Roma individuals in this population before, and this seems quite evident in this plot.

Finally, the Gujarati:

It is evident that the North European and East Asian element is limited to a few individuals, with most of them appearing to be a "South Asian" and "West Asian" blend. One might speculate that these represent individuals admixed with other subcontinental populations where these two components are more important.