Tuesday, November 30, 2010

Outliers in the Dodecad Project (23andMe data)

As promised, I have started to investigate outliers among Dodecad Project members. I used NNclean as implemented in the prabclus package to find data points that had a great distance to their nearest neighbor among either Dodecad Project members or the standard 692-individual panel I use in the Galore analysis.

To make a long story short, here are the IDs identified as outliers:
"DOD157" "DOD168" "DOD169" "DOD036" "DOD048" "DOD088" "DOD034" "DOD030" "DOD060" "DOD132" "DOD128" "DOD175"
An outlier is someone who is not very close to any other individual and hence does not really "cluster" with anyone. Thus, it is recommended to remove outliers prior to clustering, as otherwise they will form makeshift clusters that don't really have a good meaning.

Looking at the individual spreadsheet reveals that many of these outliers have very unusual ancestry. This falls under two categories:
  1. Recent admixture between geographically separated populations
  2. Being the only member from an unsampled population
In the first case, admixed individuals fall in the "empty space" between their parental clusters, and thus do not cluster with anyone else, unless a person with a similar type of admixture happens to also be in the dataset.

In the second case, there are no members of the individual's group. Sometimes, if a group is close enough to another, this is not a problem, but there are many distinctive population groups for which that is not the case.

While outliers will be removed from some analyses, their outlier status will continue to be evaluated as new reference populations, or Dodecad Project members are added.

What's next for Clusters Galore analysis

The first few runs of the Clusters Galore analysis have proven quite successful; I've posted another one on the HGDP panel in my other blog.

Now, it is time to assess the results and see what improvements can be made. I see a few avenues for improvement:

Outliers

Clusters, by definition, are composed of at least 2 individuals. Individuals who are the only representatives of their populations (e.g. if a Pygmy or an Icelandic+Armenian mix) will, by necessity, attach themselves to the closest cluster (e.g., to Yoruba, or to some Central European population), even though they are not necessarily close to that population.

Outlier detection is a difficult problem, but I will try some ideas on how to tackle it.

Phantom clusters

mclust is resilient to phantom clusters, i.e., clusters of "misfits" who don't belong in any other populations but are banded together erroneously by the algorithm. That is inevitable in an automated procedure, especially one that is pushing the limits of ancestry inference. Phantom clusters are, by their nature, transient, so there are some ideas on how to avoid them and how to focus on very robust and repeatable clusters.

Typicality

Being part of a cluster tells you nothing about how "typical" a member of the cluster you are, i.e., how close to the average. This problem is exacerbated by the fact that the clusters inferred by mclust may have varying shape, size, and orientation.

Nonetheless there are ideas on how to quantify members' typicality, and I will explore them. Please note that typicality is not necessarily the same as "purity". For example, an elongated cluster of African Americans will have typical members with 20% European admixture, but the "purest" African Americans will have 0% European admixture and be very atypical of their group as a whole. Similarly, typical Turks have 5-6% East Eurasian admixture, but people with 10% East Eurasian admixture are less typical, but more likely to be descended from central Asian Turkic people.

Any new technique will have its birth pains, and hopefully myself and others will help identify them and resolve them.

Monday, November 29, 2010

Galore analysis improved, plus K=56 results of Clusters Galore analysis for Dodecad Project members (up to DOD236)

For background, please read the post on the K=48 analysis and links therein.

As I mentioned in the previous posts, my technique depends on the use of MDS to reduce the dimensionality of genomic data from 177,000 SNPs or so to a few dozen dimensions capturing most of the variance.

Subsequently, mclust a state of the art clustering algorithm is applied on the MDS representation: this iterates between choices of K, the number of clusters, trying clusters of different shape, volume, and orientation, and chooses the optimal clustering, maximizing the Bayes Information Criterion. In simpler terms, it finds as much detail as possible in the data but penalizes too ornate models and avoids finding "ghost" clusters that are not really supported.

These are clusters derived from data of unlabeled individuals. The only human input into the process is the number of MDS dimensions to retain.

In my previous K=48 analysis, I retained 30 dimensions, but I also noted that this is not really optimal. Choosing more, or less, dimensions might lead to even better resolution (higher K).

More dimensions = more possible ways to distinguish between individuals, but also, possibly, more noise, as individuals might not be "clustered" in them.

Fewer dimensions = less possible ways to distinguish between individuals, but also, possibly, less noise from the uninformative higher dimensions.

Thus, the question arises: how many dimensions to retain?

Here is a plot of the optimal number of clusters inferred, depending on how many dimensions I chose to retain:
As you can see, when a few dimensions are retained, relatively few clusters are inferred, while as the number of dimensions goes beyond a certain point, the number of clusters starts to decrease again, as more noise is added (*)

The number of clusters peaks (see figure) at 16 and 22 dimensions retained; both of these produce 56 different clusters in the optimal solution.

Here are the results for Dodecad Project members (up to DOD236) with K=56 and 16 dimensions retained. In comparison to the previous K=48 analysis, we are now able to:
  1. Split CEU White Utahns (#1) from French (#15)
  2. Split CEU White Utahns (#1) from continental Germanics (#14)
  3. Split French (#15) from Spaniards (#2)
  4. Split Armenians (#7) from Turks (#19)
  5. Split Slavs (#23) from Balts (#26)
  6. Split Cypriots (#30) from Sephardic Jews (#21)
The most astonishing finding is, however, at least for me, the emergence of a cluster (#16) comprised in the great majority by people from Greece and Southern Italy, with very few individuals from elsewhere. Notice that #16 has absolutely no representation in the reference populations, which lack South Italians and Greeks.

Once again, I urge participants to help themselves and others by leaving a comment in the ancestry thread.

(*) The change is not, however, smooth. The more general problem is to choose which dimensions to retain, rather than choosing how many of the first ones to retain. The first few dimensions of MDS capture a decreasing portion of variance, but the data are not guaranteed to be "split" in them. However, this is a much harder problem, as we have to figure out (i) how many dimensions to retain, and (ii) which ones. Even if we fix (i), by choosing to retain, e.g., 10 dimensions, we still have to choose which 10: this is close to half a billion different combinations of which 10 to choose from the total of 38 possible candidate ones.

Results for FFD048 to FFD055 posted

This concludes the number of Family Finder individuals who sent me their data by the deadline.

Admixture proportions can be found in the spreadsheet

All populations:

Individual bars:

Results of Clusters Galore analysis for Dodecad Project members (up to DOD236), K=48

This is the result of the new type of ancestry analysis I have recently devised. For background, please read:
In total 894 individuals were included in this analysis, 202 Dodecad Project members and 692 from the published references. 30 MDS dimensions were retained, and mclust was run with a maximum number of clusters = 60. In the optimal solution, 48 clusters were inferred.

At the beginning of the results spreadsheet are the 202 Dodecad Project members and their probabilities of belonging to the 48 clusters. This is followed by the 36 reference populations, listing how many individuals from each one were assigned to each of the 48 clusters.

To help you interpret these results, you might want to consult the individual and population ADMIXTURE spreadsheets, as well as the information Project members have chosen to reveal about themselves. Feel free to add your own information in that thread.

Here are the results for some people who have chosen to reveal information about themselves:
  • Spencer Wells is DOD162 and he has 28% probability of being in the CEU (White Utahn)/French cluster #1, and 72% probability of being in cluster #3 which is CEU (White Utahn) specific (in the reference populations) and in which 29 Project members are also assigned (almost all of them Northwestern Europeans)
  • Razib of Gene Expression is DOD075 and is assigned in cluster #15, which includes Gujarati and North Kannada individuals
  • pconroy, submitted three samples, DOD097 is Sicilian fall in cluster #9 to which many Cypriots and Sephardic Jews belong, and many Project members of South Italian/Sicilian background; DOD098 and DOD099 are his Irish parents: DOD098 is almost evenly split between #1 and #3 (like Wells), and DOD099 is 97% in #3
  • Adriano Squecco DOD139 is North Italian and is in cluster #2, in which 25/25 reference Tuscans and 10/12 reference North Italians belong
  • Lacko DOD083 is 100% southern Polish and he falls in cluster #20, in which all reference Belorussians and Lithuanians fall
  • Mike Maddi DOD021 is Sicilian and is in cluster #9 like DOD097, showing some probability (13%) of also being in #2 like Adriano Squecco
  • An Anonymous Pole falls in #20 like Lacko and the reference Slavs
  • Ilmari DOD003 and Ari DOD131 are Finns and fall in cluster #13. This is an interesting one, as it does not occur at all in the reference populations; I'll let you guess what population it's centered on.
  • Eastara's mother (DOD025) is Bulgarian and falls in cluster #10, which is centered in Romanians in the reference populations, but note that there are also 13 Dodecad Project members who fall in it, many of them from different parts of the Balkans. I hope more people from the Balkans will contact me for inclusion in the project, as I am sure that finer-scale can be achieved there with increased participation.
  • Basar (DOD049) is half-Anatolian Turk and half-Laz. He falls in the cluster encompassing Armenians and Turks in the reference populations.
  • Bubba (DOD066) is North German with a pinch of Danish and falls in cluster #3
  • afpjr (DOD014) is half Greek and half Italian/Sicilian; he falls in cluster #9 (95%)
  • Francesc (DOD217) is Catalan and falls in cluster #17, centered on Spaniards in the reference populations.
The Project Greeks and Mixed Greeks fall in clusters #2, 9, 10 tying them to Italy and the Balkans, as might be expected. I hope more Greeks will decide to participate in the Project, so we can discover more interesting patterns in our population.

I cannot stress enough how revealing non-identifying ancestry information in the relevant thread will help both yourselves and others make better sense of these and future results.

There are clusters composed entirely of Dodecad Project members (e.g., the aforementioned one), and others which are centered on one or two reference populations, but encompass a wider variety of non-represented populations. So, please take the time to leave a comment in the ancestry thread.

This is not the end of the story. There are more clusters to be discovered in the data; the inclusion of 200+ new samples in this analysis has caused new clusters to appear and distinctions that were previously detected to "fold back" (e.g., between Armenians and Turks). I am currently investigating how the choice of number of MDS dimensions to retain affects the number of inferred clusters in the optimal solution.

Sunday, November 28, 2010

Submission of Family Finder data is now CLOSED

Thank you all for submitting your data; the remaining results will be posted in the blog over the next week or so.

If you have submitted your data in time, but did not receive an ID yet, you will.

If you want to be alerted for future opportunities, and to keep up with the progress of the Project, please subscribe to the feed.

Clusters galore: less is more, or, pushing the limits of ancestry inference

I had already hinted in my previous post on my new technique that retaining all MDS dimensions might add noise to the analysis, and I was hopeful that even finer resolution could be achieved with fewer dimensions.

In the spreadsheet you can see the optimal solution if one retains only 10 dimensions in which case 45 clusters are inferred. Previously, I retained 47 dimensions, and got only 35 clusters in the optimal solution: less is more.

In comparison to the previous analysis, I can detect some interesting changes:
  1. Spaniards and Portuguse are split from Tuscans and are joined by some French and North Italians; the rest of the French stay with White Utahns, and the rest of the North Italians stay with Tuscans.
  2. Romanians too get their own cluster
  3. Turks are split from Assyrians/Armenians.
  4. Germans are split from Scandinavians, with 1 sample from either population going to the other population.
  5. The relationship between Cypriots and South Italians is retained, but most of the Greeks (many of whom were borderline between the Tuscan and South Italian cluster) go the Tuscan way.
I am now studying how to choose the optimal number of MDS dimensions to retain, so I will not report any individual data about this to project participants. I just wanted to let everyone share in the excitement.

Results for FFD035 to FFD047 posted

Individual proportions can be found in the spreadsheet

All populations:

Individual bars:

Clusters galore with Dodecad populations

Number of individuals assigned to each cluster can be found in the spreadsheet. Populations in italics are composed entirely from Dodecad Project members.

Please read the post in Dienekes' Anthropology Blog to see what this type of analysis means.

47 MDS dimensions were retained, and the optimal number of clusters was 35. Retaining less or more dimensions may alter this number, as after a certain point extra dimensions only contribute noise to the analysis; this is a matter of investigation.

It is hardly practical to comment on all 35 clusters, so I will limit myself to a few observations:
  • Turks, Armenians, and Assyrians fall in cluster #1
  • Scandinavians, White Utahns, Germans, and some French fall in cluster #2
  • Portuguese, French, North Italians, Tuscans, Spaniards, and Romanians fall in cluster #3
  • Greeks, South Italians/Sicilians, Cypriots, and Sephardic Jews from Bulgaria and Turkey fall in cluster #4 (but see note)
  • Finns fall in cluster #5
  • Almost all Ashkenazi Jews fall in #6
  • All Dodecad Project Russians, plus reference Lithuanians and Belorussians fall in #9
8 Greeks fall in cluster #4 and 2 in cluster #3. However, many of the ones who fall in #4 also have some non-trivial probability of falling in #3. Probabilities for all other clusters are less than 0.1%. All Project Greeks can write to me to learn their exact probabilities.

Of course, it should be noted that:
  1. If two populations can be perfectly distinguished from each other, then there are genetic differences between them (they split from each other some time ago, they underwent different types of admixture, etc.) allowing the clustering algorithm to detect their differentiation
  2. If two populations cannot be distinguished from each other, this does not mean that they are not indistinguishable in principle; it does mean, however that through either common ancestry or very similar patterns of admixture, they have become quite similar to each other in the Eurasian context.
If you are a Dodecad Project member (23andMe data) from one of the populations in italics and are wondering which cluster you fall in, first check whether all individuals from your population fall in the same cluster, in which case you already know the answer.

Otherwise, you may write to me, with your DOD number, and I'll tell you.

Results for FFD020 to FFD033 posted

Admixture proportions can be found in the spreadsheet

All populations:

Individual bars:


Saturday, November 27, 2010

Results for FFD003 to FFD019 posted

UPDATE: Color-coding problem fixed.

Admixture proportions can be found in the spreadsheet

All populations:


Individual bars:


Friday, November 26, 2010

Results for DOD223 to DOD236 posted

Admixture proportions can be found in the spreadsheet

All populations:
Individual bars: