Genome-Wide Association Studies for the Rest of Us: Panel Discussion 2
Transcript source: creator-uploaded captions on YouTube, unedited. Paragraph breaks and timestamps added by Vidleaf.
[0:00] [Dr. Teri Manolio] So again, while people are sort of getting their coffee metabolized and that, David, it would be interesting to hear a little bit more, if you can expand a bit, on this issue of choice of controls. And I know there was some discussion at the break that we really ought to be thinking seriously about convenience controls and how one goes about, and so there may be others who want to speak on that. But in terms of the convenience controls that have been done so far, can you talk a little bit more about what the strengths and weaknesses of those might be?
[0:31] [Dr. David Hunter] So I think, the strengths are that, you know, these are large-scale samples, and often we don't have, for a particular phenotype, studies that are as large as 3,000 controls. In the Wellcome Trust example, you know, we might have a best case study for a rarer disease of just a few hundred. So in that context, probably the best advice would be to genotype the controls you have that meet the best practice epidemiologic methods, but be aware that there may be other populations you can import to essentially increase your number of controls.
[1:11] And that will give you some limited capacity to compare allele prevalence in the best practice design with the in silico controls. And the other issue that Elizabeth brought up is that as the platforms change, the meaning of these in silico controls is going to change. So as you move from 100k to 317k to 550k and as you import these control groups, you're going to have a subset of SNPs that aren't represented on those controls. So again, I think if it's possible, you'd always like to have some of your own controls, just so you can compare a little prevalence, recognizing that you're not going to have maximum statistical power by any means.
[1:49] [Dr. Teri Manolio] Debbie. [Dr. Deborah Nickerson] Yeah, I think it's really important, too, to think about comparing your allele frequencies to the HapMap, and to the other controls that may exist, because they do point out some quality metrics, sometimes, that maybe you're not aware of. The other thing that I would suggest, that I found is interesting, is that most of the studies have not even genotyped sex. And one of the discrepancies that we found in working with individuals is sample mix-ups over time.
[2:25] And you don't really know that they exist, and actually just getting a genetic test for sex among your samples will show you where there may have been a small batch that went in and people -- years ago, they reversed tubes, and this has just been carried on in your study over time. So actually getting some aims done on your study and some indication of the sex of your samples will help you to go forward with some of these large-scale studies, because you have an idea of quality and integrity of your samples.
[3:05] [Dr. Teri Manolio] Those are excellent points, Debbie, and I think David can give examples of breast cancer studies where unexpected men were included, and prostate cancer studies where unexpected women were included, and outbreaks of twins in duplicate samples, and -- [Dr. David Hunter] It actually does make a very good point that the platforms are now so good, you know, not for every SNP, but across the board, the completion rates are 99 percent, concordance rates at 99.5 percent, that the biggest source of error may be this repository in other areas. And so having a fingerprint very early in the life of the sample -- people used to use the microsatellites, now people seem very happy with a sequenomyoplex [spelled phonetically] about 30 SNPs, something like that.
[3:45] Having a fingerprint very early in the study allows you to track these errors if somebody turns a plate upside down. There are other ways of catching that, but you catch it unambiguously if a couple of samples are flipped if you have an expected fingerprint from very early on in the life of the sample. [Dr. Teri Manolio] Front microphone, please. [Female Speaker] Thanks very much for a really interesting day so far. And I have a question related to a point that Elizabeth made. I think it's really important that the genotypes are actually phenotypes themselves, and they're actually not -- they're not categorical; they're actually just two quantitative allele intensities, and I wonder -- it's a technical and a statistical question -- if people are starting to look at using that as the exposure measure, and if anybody has experience in studies with that.
[4:36] [Dr. Teri Manolio] That's a really interesting question. This is one of those questions, you know, that could be pursued in datasets. I don't know that there is a lot of that, and maybe some of the genomicists in the group could comment. There have been some interesting, you know, explorations of samples that fail whole-genome amplification, for example. There are certain characteristics of people that will fail that and that may cluster in families, so there might be genes related to that. Debbie, did you want to comment on whether these phenotypes -- the variant genotypes that one sees from these are somehow possible to use as outcomes?
[5:12] [Dr. Deborah Nickerson] You can use them. And actually, the copy number variation is a great way. That was found by weird clustering. So, looking at these outlier genotypes that you get where you see multiple clusters, and actually people are finding these differences. For example, in some of these platforms, they're typing over unknown SNPs, and that causes, actually, four clusters or five clusters to occur.
[5:43] And you're getting genotype information on that unknown SNP and you can go back and mine that information. They may be thrown out. So, I think that actually looking at some of these things are really important. But, they're not the first thing that you look at, right? [Dr. Teri Manolio] Good point. Mike? [Mike] Yeah, thank you for two great talks. So this is a question about meta-analysis projects, and the effect of using different platforms in studies in a meta-analysis.
[6:15] I'd like to hear more about that. And also, the different choice of control groups in the studies in a meta-analysis, and how that actually might make you have pause in some of these meta-analysis projects. [Dr. Teri Manolio] So, maybe Elizabeth -- could you comment on the different platforms? [Dr. Elizabeth Pugh] I can comment a little, and I think Laura Scott will probably talk about it a little bit more later. It really depends on how you're going into your analysis. There are different SNPs on each of these platforms. So, if you're combining data where you have different SNPs in the two studies, then you're going to have to either infer the genotypes from the other platform, or calculate haplotypes using the SNPs from both, and use the haplotypes, not the individual SNPs, or look at linkage to a region and see whether -- or association to a particular region or haplotype.
[7:06] So, it makes your analysis a little bit more complex, but it's certainly possible, and there are a number of groups out there who have done this very successfully. So, it's not nearly as critical as people thought it might have been a few years ago to have exactly the same product and exactly the same SNPs, because a number of groups have shown that it's very possible to do the analysis; it's just another technical challenge along the way. [Dr. Teri Manolio] And Laura will probably comment as well, but certainly in the diabetes studies, those three diabetes studies that I showed, when they wanted to combine their data, they were combining it across the Affymetrix and Illumina platforms, and they used imputation algorithms to do this.
[7:43] I think one of the take-home lessons from that is that imputation algorithms are really very good, but they're not perfect, like nothing is. And once one looks at an imputation and sees that it looks pretty good, you do have to go back and actually genotype the SNP that you're inferring. So -- [Dr. David Hunter] Just being aware, also, that there are gaps in the platforms, and so, I guess, cognitive function is one of the Wellcome Trust disease groups, but oops, there's no good surrogate for APOE4 on that generation of the Affy chip. And you know, you can get these printouts where there's an infomatic attachment of a gene name to the RS numbers, but if you're really after a very specific SNP, you really have to confirm that there's something in good LD with that.
[8:24] [Dr. Teri Manolio] And I think your second question, Mike, was one for David. You might want to repeat it if -- [Mike] This is to do variation in the control groups in studies in a meta-analysis. [Dr. David Hunter] So again, you know, I don't think there's much evidence so far that sort of reasonably well designed case control studies, compared with reasonably well designed case control studies, compared with sort of hospital-based versus population controls, are leading to radically different answers or even perceptibly different answers for these stronger SNPs. In a meta-analysis, nonetheless, I think you'd always want to, you know, run within strata defined by design or some measure of quality, just to confirm that.
[9:03] And as we get to the lower relative risks, you'd expect that if we're introducing any methodologic noise, that's going to make it harder to discern those. But, you know, I think we do have to be aware that the sort of traditional reasons for concern about certain designs related to information bias and selection bias, you know, may not be so pressing in this area, if there's no relationship between participation rate and genotype, et cetera. So we have to be careful, and probably in a meta-analysis you want to stratify.
[9:37] But so far, I think, you know, the concerns we have with self-reported data and retrospective versus prospective data, you know, may not apply in such strength to these studies. [Mike] Thank you. [Dr. Teri Manolio] And certainly when I was a baby epidemiologist, they told us first, be terrified of case control studies, which I don't think was great advice, but -- and also, if you're going to do them, consider using multiple control groups. And one of the big advantages, now, of data sharing and of databases like CGEMS and others, that have their control allele frequencies widely available, is that you can look at multiple groups and see, "Is my control group somewhat comparable to these others, and if not, why not?"
[10:16] And also, do your findings sort of hold up? Several of the platform developers are actually developing large control databases that they're making widely available. So, Illumina just announced theirs, and that one will be going into dbGaP as soon as some of us do the work that's needed to get it in there, myself included. But at any rate, those will be available. There has been, for many years, discussions about trying to get the [unintelligible] genotype frequencies, because that's a nationally representative sample, into some kind of an accessible database, and the work is going forward on that as well.
[10:50] [Male Speaker] This is a basic question. How do you define a population? And then second, when the -- in some of the papers that have come out in the last two years, we know very well different populations, different regions of the genome undergo different selection pressures. So therefore, the population history differs from human -- [unintelligible]. So when we put these things together -- that is, mixing populations and then seeing the effect of these mixed populations, and then the question is, what are the chances of finding spurious results in the pooled population?
[11:39] Third, -- well, I'll stop there. [Dr. Teri Manolio] Yeah, why don't you stop there? [laughs] We'll come back if we can. So who would like to tackle that one? [Dr. David Hunter] So very quickly -- I mean, clearly this is a major concern. You know, I think the genomic control and other methods give us some confidence that we can approach this statistically, and anyone would advise that you collect self-reported information about ancestry, however that's defined, you know, country of origin or whatever, to allow yourself to restrict or stratify on the basis of that information.
[12:16] And then the other key thing, obviously, is that a lot of these studies are obviously happening in populations of European ancestry. There's no guarantee that that SNP will also be associated with the same phenotype in other ancestries. And that has to be looked at carefully. And then finally, as the costs come down, as the sample sets increase, there are going to be associations that are private to particular populations that will only ever be discovered if we look in those populations. [Male Speaker] Just a comment along those lines.
[12:48] I prefer that some of the people that are doing GWAS studies give importance to a concept called Walland Effect [spelled phonetically] in population genetics. But there's an important concept. [Dr. Teri Manolio] And maybe you'd take a second to explain what you mean by that? [Male Speaker] Walland Effect -- long time ago, a population geneticist named Walland suggested that by putting the -- so-called mixing these populations, it affects the genotypic frequency's spuriousness, and when you use those elevated genotypic frequencies, it could affect the association studies.
[13:26] [Dr. Teri Manolio] Well, and I think -- you know, that is widely recognized and people are considering it. Perhaps they're considering it a little bit too much, and perhaps Laura will comment. But what we've tended to do is to focus on very homogeneous populations, and somewhat to the detriment of being able to describe risk in the general U.S. population. Laura? [Dr. Laura Scott] So there are a couple interesting things. One thing that I've already started to see happen is that if different groups are collecting case samples, and they're comparing those to existing control samples, once you go to try to do a meta-analysis, you actually don't have independent samples.
[14:06] And so, you're going to have to go back and redo analysis, either in some way allocating controls, or taking into account that you've actually used the same samples twice. The other thing that's interesting to think about -- sorry, this is having a problem -- is that when you have epidemiology studies that have matched particular variables, such as if you had a study of type 2 diabetes that had matched on body mass index, what we're finding is that sometimes genes that work to cause diabetes through body mass index actually can't be found in those epidemiological studies that were very tightly matched.
[14:49] And so, in those cases you might actually have advantages to having population-based controls. [Dr. Teri Manolio] Very good points. And just before Elizabeth, maybe I might ask the AV person if he would kindly come out and put up Laura's slides, because we're just going to finish up here. And then Elizabeth's comment. [Dr. Elizabeth Pugh] Just -- David asked me one question and I thought I'd just mention it. The question was what about pooled samples. So, one possibility that people did, especially before genotyping costs started coming down, was to take a pool of cases and a pool of controls and compare estimated allele frequencies between them.
[15:26] That's still being done. It's a little bit challenging technically, both in making sure that you take very good care when you're building the DNA pool so that you have equal amounts of DNA from each individual, and then it's a little bit tricky as far as the interpretation. You can't just run a genotyping clustering algorithm and get genotypes. You have to try to estimate the proportion of each of the genotype frequencies. So it's still possible. It's certainly a consideration, and it's obviously a much lower budgetary cost. But you need to work with somebody who really knows what they're doing both on the lab end and on the analysis end.
[16:00] [Dr. Teri Manolio] That's an excellent point, and this was one study design that I think was used initially when genotype costs were so very high. And of course, another disadvantage of it is that you don't have individual genotype data then to analyze -- you know, to look at your very interesting folks.
Open in the Vidleaf workbench
Search the transcript, select lines, copy quotes with timestamps, translate.
Attribution
"Genome-Wide Association Studies for the Rest of Us: Panel Discussion 2" by National Human Genome Research Institute (https://www.youtube.com/@genometv), licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/). Source video: https://www.youtube.com/watch?v=-JBpDZu-qxw. This page is a text transcript of the video with paragraph breaks and timestamps added; the creator is not affiliated with and does not endorse Vidleaf.
Are you the creator or a rights holder? Request a correction or removal: copyright@vidleaf.app (see About these pages).
Last updated