Transcription of Gene set enrichment analysis: A knowledge-based …
1 Gene set enrichment analysis: A knowledge-basedapproach for interpreting genome-wideexpression profilesAravind Subramaniana,b, Pablo Tamayoa,b, Vamsi K. Moothaa,c, Sayan Mukherjeed, Benjamin L. Eberta,e,Michael A. Gillettea,f, Amanda Paulovichg, Scott L. Pomeroyh, Todd R. Goluba,e, Eric S. Landera,c,i,j,k, and Jill P. Mesirova,kaBroad Institute of Massachusetts Institute of Technology and Harvard, 320 Charles Street, Cambridge, MA 02141;cDepartment of Systems Biology, Alpert536, Harvard Medical School, 200 Longwood Avenue, Boston, MA 02446;dInstitute for Genome Sciences and Policy, Center for Interdisciplinary Engineering,Medicine, and Applied Sciences, Duke University, 101 Science Drive, Durham, NC 27708;eDepartment of Medical Oncology, Dana Farber Cancer Institute,44 Binney Street, Boston, MA 02115;fDivision of Pulmonary and Critical Care Medicine, Massachusetts General Hospital, 55 Fruit Street, Boston, MA 02114;gFred Hutchinson Cancer Research Center, 1100 Fairview Avenue North, C2-023, Box 19024, Seattle, WA 98109-1024.
2 HDepartment of Neurology, Enders260, Children s Hospital, Harvard Medical School, 300 Longwood Avenue, Boston, MA 02115;iDepartment of Biology, Massachusetts Institute ofTechnology, Cambridge, MA 02142; andjWhitehead Institute for Biomedical Research, Massachusetts Institute of Technology, Cambridge, MA 02142 Contributed by Eric S. Lander, August 2, 2005 Although genomewide RNA expression analysis has become aroutine tool in biomedical research, extracting biological insightfrom such information remains a major challenge. Here, we de-scribe a powerful analytical method called Gene Set EnrichmentAnalysis (GSEA) for interpreting gene expression data. The methodderives its power by focusing on gene sets, that is, groups of genesthat share common biological function, chromosomal location, orregulation.
3 We demonstrate how GSEA yields insights into severalcancer-related data sets, including leukemia and lung , where single-gene analysis finds little similarity betweentwo independent studies of patient survival in lung cancer, GSEA reveals many biological pathways in common. The GSEA method isembodied in a freely available software package, together with aninitial database of 1,325 biologically defined gene expression analysis with DNA microarrays hasbecome a mainstay of genomics research (1, 2). The challengeno longer lies in obtaining gene expression profiles, but rather ininterpreting the results to gain insights into biological a typical experiment, mRNA expression profiles are generatedfor thousands of genes from a collection of samples belonging toone of two classes, for example, tumors that are sensitive to a drug.
4 The genes can be ordered in a ranked listL,according to their differential expression between the classes. Thechallenge is to extract meaning from this common approach involves focusing on a handful of genes atthe top and bottom ofL( , those showing the largest difference)to discern telltale biological clues. This approach has a few majorlimitations.(i) After correcting for multiple hypotheses testing, no individualgene may meet the threshold for statistical significance, because therelevant biological differences are modest relative to the noiseinherent to the microarray technology.(ii) Alternatively, one may be left with a long list of statisticallysignificant genes without any unifying biological theme. Interpre-tation can be daunting and ad hoc, being dependent on a biologist sarea of expertise.
5 (iii) Single-gene analysis may miss important effects on processes often affect sets of genes acting in concert. Anincrease of 20% in all genes encoding members of a metabolicpathway may dramatically alter the flux through the pathway andmay be more important than a 20-fold increase in a single gene.(iv) When different groups study the same biological system, thelist of statistically significant genes from the two studies may showdistressingly little overlap (3).To overcome these analytical challenges, we recently developeda method called Gene Set enrichment Analysis (GSEA) thatevaluates microarray data at the level of gene sets. The gene sets aredefined based on prior biological knowledge, , published infor-mation about biochemical pathways or coexpression in previousexperiments.
6 The goal of GSEA is to determine whether membersof a gene setStend to occur toward the top (or bottom) of the listL, in which case the gene set is correlated with the phenotypic used a preliminary version of GSEA to analyze data frommuscle biopsies from diabetics vs. healthy controls (4). The methodrevealed that genes involved in oxidative phosphorylation showreduced expression in diabetics, although the average decrease pergene is only 20%. The results from this study have been indepen-dently validated by other microarray studies (5) and byin vivofunctional studies (6).Given this success, we have developed GSEA into a robusttechnique for analyzing molecular profiling data. We studied itscharacteristics and performance and substantially revised andgeneralized the original method for broader this paper, we provide a full mathematical description of theGSEA methodology and illustrate its utility by applying it to severaldiverse biological problems.
7 We have also created a softwarepackage, calledGSEA-Pand an initial inventory of gene sets(Molecular Signature Database, MSigDB), both of which are of considers experiments with genomewideexpression profiles from samples belonging to two classes, labeled1 or 2. Genes are ranked based on the correlation between theirexpression and the class distinction by using any suitable metric(Fig. 1A).Given anaprioridefined set of genesS( , genes encodingproducts in a metabolic pathway, located in the same cytogeneticband, or sharing the same GO category), the goal of GSEA is todetermine whether the members ofSare randomly distributedthroughoutLor primarily found at the top or bottom. We expectFreely available online through the PNAS open access : ALL, acute lymphoid leukemia; AML, acute myeloid leukemia;ES, enrich-ment score; FDR, false discovery rate; GSEA, Gene Set enrichment Analysis; MAPK, mitogen-activated protein kinase; MSigDB, Molecular Signature Database; NES, normalized enrich-ment Commentary on page and contributed equally to this whom correspondence may be addressed.
8 E-mail: 2005 by The National Academy of Sciences of the cgi doi October 25, 2005 vol. 102 no. 43 15545 15550 GENETICSSEE COMMENTARY that sets related to the phenotypic distinction will tend to show thelatter are three key elements of the GSEA method:Step 1: Calculation of an enrichment calculate an enrich-ment score (ES) that reflects the degree to which a setSisoverrepresented at the extremes (top or bottom) of the entireranked listL. The score is calculated by walking down the listL,increasing a running-sum statistic when we encounter a gene inSand decreasing it when we encounter genes not magnitudeof the increment depends on the correlation of the gene with thephenotype. The enrichment score is the maximum deviation fromzero encountered in the random walk; it corresponds to a weightedKolmogorov Smirnov-like statistic (ref.)
9 7 and Fig. 1B).Step 2: Estimation of Significance Level estimate thestatistical significance (nominalPvalue) of theESby using anempirical phenotype-based permutation test procedure that pre-serves the complex correlation structure of the gene expressiondata. Specifically, we permute the phenotype labels and recomputetheESof the gene set for the permuted data, which generates a nulldistribution for empirical, nominalPvalue of theobservedESis then calculated relative to this null , the permutation of class labels preserves gene-genecorrelations and, thus, provides a more biologically reasonableassessment of significance than would be obtained by 3: Adjustment for Multiple Hypothesis an entiredatabase of gene sets is evaluated, we adjust the estimated signif-icance level to account for multiple hypothesis testing.
10 We firstnormalize theESfor each gene set to account for the size of the set,yielding a normalized enrichment score (NES). We then control theproportion of false positives by calculating the false discovery rate(FDR) (8, 9) corresponding to eachNES. The FDR is the estimatedprobability that a set with a givenNESrepresents a false positivefinding; it is computed by comparing the tails of the observed andnull distributions for details of the implementation are described in theAppendix(see alsoSupporting Text, which is published as supporting infor-mation on the PNAS web site).We note that the GSEA method differs in several important waysfrom the preliminary version (seeSupporting Text). In the originalimplementation, the running-sum statistic used equal weights atevery step, which yielded high scores for sets clustered near themiddle of the ranked list (Fig.)