TRANSCRIPT · CC BY 3.0

Genome-Wide Association Studies for the Rest of Us: Panel Discussion 4

National Human Genome Research Institute · Published · 15 min · English · License: CC BY 3.0 · Source: watch on YouTube

Transcript source: creator-uploaded captions on YouTube, unedited. Paragraph breaks and timestamps added by Vidleaf.

[0:00] [Dr. Daniel Levy] Let me begin, if I can, by asking Bob a question. You mentioned in your presentation that good quality science is made up of several components, and one that you emphasized, actually three times, was replication, replication, replication. As we go from five to 10 to 20 to 50 to 100 genome-wide association studies, and as the numbers of individuals and cases and controls and populations studied gets larger and larger and larger, it really -- it would be wonderful to have the tools and the opportunity to conduct in silico replication.

[0:40] Are we near the point where we can have such tools, and implement them effectively to allow for an expensive replication? [Dr. Robert Hoover] Well, I -- I think that David addressed that a little bit in his presentation -- and actually Jim's is relevant here, too, that the -- in fact, two of the recent publications on prostate cancer used the NCI results as a -- as another replication, since they were on the Web.

[1:18] So I think there -- as more and more of the actual results of these get posted, you -- you have an ability to -- to replicate instantaneously and in your findings. [Dr. Daniel Levy] What about for more global tools, to allow for replication across trades? I imagine dgGaP might give some of that capability? [Dr. Jim Ostell] Well, we'd love to, but we're not sure how to do it. We -- we do assume that people will be doing in silico replications, particularly when there's access to things like large cohort studies, like Framingham, because there, there is a large number of phenotype measures that are available that might be appropriate to lots of these studies, because they do physical exams on people.

[2:09] And so even though the exact measure may not be exactly the same, since they're typed on large number of SNPs, since you can impute across platforms, you could probably certainly find out if you were consistent or not. [Dr. Daniel Levy] We have questions from Teri Manolio and Debbie Nickerson. [Dr. Teri Manolio] One thing that I think people would find very useful in these large databases is a -- is a very simple, very, you know, obvious way to cite them. When we're, you know, trying to -- trying to look at these, and then acknowledge the use of them in papers and other places, it would be just grand to say, you know, "To cite this, and you should if you're using the data at all, here's what you put in your reference list."

[2:51] Is that available on the CGEMS Web site? [Dr. Robert Hoover] I mean, people have -- they did, in fact, reference where they -- where they got the replication, and they -- they referenced the Web site. [Dr. Jim Ostell] In -- also to answer that on the dgGaP one, in the authorized access download, you agree to data use conditions. And typically, it's things like, "I'll only study eye disease," so it's consistent with the consent, or "I promise not to try to re-identify individuals."

[3:22] But also usually there's a few sentences in there which say, "And I will cite this study, not by citing the database, but by citing the study -- the original PIs." A different type of citation is when you publish, having analyzed a lot of this sort of stuff. Those accession numbers are really for that. They don't really credit the database, but it's for tracking the data. [Dr. Daniel Levy] Debbie? [Dr. Debbie Nickerson] So I'm curious, Jim, when you take the phenotype data, will you take only normalized phenotypes, or will you get raw as well?

[3:58] Because a lot of these large studies are going to be a consortium of 10 epidemiological studies that are out there, looking across maybe 20 to 30,000 people at a time. The one thing I could think of is CARe. But those phenotypes will be normalized, and it's kind of like the difference between a raw genotype and -- you know, having raw versus processed -- [Dr. Jim Ostell] Well, just like the genotypes, we'd like to have both.

[4:33] And so, we encourage our submitters to submit both. When they try to submit only one or the other, we get back to them and encourage them. But of course, this is voluntary, so we can't really -- if they really only want to submit the derived variable and won't reveal the -- the raw data, we can't really force it. I would say, for the most part, most studies are willing to share it all.

[5:03] [Dr. Debbie Nickerson] I guess I was just wondering if it's kind of like, "I'll give you what the sequence is, but I won't give you the trace." [Dr. Jim Ostell] Yes. [Dr. Debbie Nickerson] So you won't know if it's really junky, and whether the base calls are really true. [Dr. Jim Ostell] That's exactly right. So -- that's the problem. So we have had extensive discussion with some submitters who didn't want to submit the background information, and like I said, in most cases it resolved that they found that they could. In a few cases, they found they couldn't.

[5:34] [Dr. Laura Scott] So, I have a follow-up on Debbie's question, which is there's a lot of accumulated knowledge that happens when people do a study, right? And when you go into a new study, it can sometimes take, I don't know, a year to really get -- if you're that quick, to get up to speed on all the nuances. So is there a place in dbGaP where -- where someone can say, "These are the things you need to watch out for," or how do you help people avoid making mistakes that someone more familiar with the data might not have made?

[6:10] [Dr. Jim Ostell] Well, that's a very good question. The original submitter can supply whatever documentation or comments they're able or willing to make. And as you know, though, a lot of these things are sort of, "Well, something's funny happening here," you talked to the PI, they say, "Oh yeah, well, five years ago we changed exactly how we were doing that, and now we're doing something else." And this just has those same problems. I mean, if it's really still just sitting in the head of the PI, that's where it is, and -- however, you know, these are not anonymous, so, you know, the name of the group that supplied it and the people working on the project are there, and if someone -- our assumption is that if someone really starts to get into one of these and finds interesting things, they're very likely to want to contact the PI and collaborate, and that seems like a good thing.

[7:08] In fact, I would guess that the PIs may find themselves with more collaborators than they want. As opposed to losing their data, it's going to be the opposite. [Dr. Laura Scott] Yeah, I think that may be true. Just another comment into the comments about the candidate -- the candidate gene SNPs, and the seeming lack of replication. And I guess I went into doing this genome-wide association, and having done candidate gene studies before, and thinking, you know, the genome-wide is just going to find so much more. We're going to learn that this candidate gene approach really -- really didn't work very well.

[7:41] And what's amazing is that, especially for traits like lipids, the things that are coming up at the top are the known candidate genes, and that's true also for diabetes. So I guess I would see almost the glass more full for candidate genes, maybe not in taking the whole spectrum, but in terms of actually have been able to pick out of the whole genome things that were associated was pretty -- pretty amazing. [Dr. Robert Hoover] I agree with you, in terms of the -- the diabetes area, and lipids that did a much better job than the cancer people did.

[8:16] I'm talking my own view of the cancer perspective. We've done an absolutely terrible job, including myself, with candidate genes, and so that's why I'm so pleased with GWAS results, because we can actually now start to look at gene environment interactions for genes that look to be important. [Dr. Daniel Levy] Bob, if I could return to the theme of replication, early today, and actually several times today, we heard about the Wellcome Trust Case Control Consortium, and it was a real tour de force: 14,000 cases divided among seven diseases, 3,000 controls.

[8:53] This was a paper that was published in a top-tier journal. But I don't remember seeing replication in there. Does it surprise you that a paper was published in a top-tier journal using genome-wide association without replication? [Dr. Robert Hoover] David, do you have any more -- since -- member of the group. [Dr. David Hunter] [Inaudible] paper [inaudible] impress -- they published studies already, so this was supposed to be their [inaudible], and you can look forward to a whole series of [inaudible] replication -- [Dr. Daniel Levy] Actually, my question wasn't directed at the Wellcome Trust Consortium, it was a question about journals and criteria for publication in the absence of replication.

[9:46] Debbie, you had a comment on it? [Dr. Debbie Nickerson] Well, just that most of the things that were in that paper were replicated prior to -- by prior publications. So it is true that there were some things in there that were not replicated, but for the most part, most of the things that were reported had been reported before. [Dr. Daniel Levy] Thank you. Raju Govindaraju has been waiting patiently. [Dr. Raju Govindaraju] Well, this is for Jim Ostell. I believe you mentioned that the dbGaP, what I see from this stage it is geared for individual -- primarily individual investigators, so they could apply.

[10:23] What if, on the other hand, you know, no individual investigators -- probably except a few -- is clever enough to look at the multiple aspects of the data, and then what if multiple investigators get together as a -- as a group, and then supposing they want to take this data else -- to different places, is there a mechanism built in for -- to allow this group to apply, and then get the data, and then take the data elsewhere, and then again come back, and then provide the results?

[10:57] Is there any? [Dr. Jim Ostell] Well, this is a -- this has been a hard thing for the NIH policy people to deal with, in terms of, you know, this broader sharing of -- of individual data like this. And so the compromise position right now is even though there's a lower bar for getting at this data than traditionally was before, the -- each application has to come from one institution, because what they're looking for is the validation of an accepted -- it's actually the institution that's taking responsibility for the behavior of the investigator, the same as with a grant.

[11:38] And so a group of collaborate -- you can list collaborators on the application, but only collaborators from the same institution, because the signing official is who is signing off on this. And if you're going to be collaborating with people from another institution, they need to also apply, so that their institution also signs off, as a separate application. One of the things you agree to is to not redistribute the data.

[12:09] And that essentially -- you're agreeing that you will hold it, you'll protect it, you're not going to put it somewhere where people can take it, or Google can index it. So it is a little clumsy if you're going to want to collaborate very flexibly, but this is a big step forward, and that's sort of the compromise position for now. [Dr. Raju Govindaraju] Thank you. [Dr. Teri Manolio] Just a -- [Dr. Jim Ostell] Teri? [Dr. Teri Manoli] A quick question, because it is 2:00 and we'll need to move on, but in terms of the -- the replication requirements that initially had been put in the literature -- and journal editors were really quite, you know, strong in requiring this -- it became kind of obvious that sometimes people were sort of splitting their samples, and doing -- doing other things that may not have been as scientifically productive and useful.

[12:53] And so replication -- while it's very important, there may be other reasons not to, and that's somewhat covered in the Nature paper that came out, just the page before the -- the Wellcome Trust Case Control Consortium. And there, there really was a group of diseases that had many things in common that the editors felt was -- was much more important to get published than waiting for the replications. Thank you. [low audio] [Female Speaker] One more. [Male Speaker] Quick. [Female Speaker] Yeah, I just want to say this is all very exciting, but there is still a bit of trepidation I think several people feel.

[13:25] So it's exciting, but it's also scary, and it -- it has to do with just maintaining the confidentiality of our study participants, and, you know, it is conceivable that there are some situations where a genotype could be used to infer identity, even if name, and you know, the traditional identification factors aren't in there. So I just want to kind of get it out there, that there still is issue, and -- and I think part of the issue is as well that there is confusion about what exactly is going to be required, as both a data provider and a data user.

[13:56] So there's still -- I think NIH is coming up with a policy of some sort, so I wonder if there's any hint that we could get on kind of what's the guideline for when data needs to be contributed, and what kind of data? Is it everything on the questionnaire? I mean, you could really go on and on, what kind of clinical data, so is there any information we could get on that? [Dr. Jim Ostell] Well, I hear it's covered in the next two talks, and I can also say that I -- that policy is being finalized now, I understand, and they're talking about releasing it in August, I believe now.

[14:35] So I don't want to jump the gun, because a lot -- a lot of those decisions are being made by people who aren't me -- [laughter] -- who are lawyers.

Open in the Vidleaf workbench

Search the transcript, select lines, copy quotes with timestamps, translate.

Open in the workbench →

Attribution

"Genome-Wide Association Studies for the Rest of Us: Panel Discussion 4" by National Human Genome Research Institute (https://www.youtube.com/@genometv), licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/). Source video: https://www.youtube.com/watch?v=phfz0qDZgVo. This page is a text transcript of the video with paragraph breaks and timestamps added; the creator is not affiliated with and does not endorse Vidleaf.

Are you the creator or a rights holder? Request a correction or removal: copyright@vidleaf.app (see About these pages).

Last updated