Example: dental hygienist

edgeR: differential analysis of sequence read count data ...

edger : differential analysisof sequence read count dataUser s GuideYunshun Chen1,2, Davis McCarthy3,4, Matthew Ritchie1,2,Mark Robinson5, and Gordon Smyth1,61 Walter and Eliza Hall Institute of Medical Research, Parkville, Victoria, Australia2 Department of Medical Biology, University of Melbourne, Victoria, Australia3St Vincent s Institute of Medical Research, Fitzroy, Victoria, Australia4 Melbourne Integrative Genomics, University of Melbourne, Victoria, Australia5 Institute of Molecular Life Sciences and SIB Swiss Institute of Bioinformatics, University ofZurich, Zurich, Switzerland6 School of Mathematics and Statistics, University of Melbourne, Victoria, AustraliaFirst edition 17 September 2008 Last revised 20 April 2022 Contents1 Introduction.

2.15Gene set testing.....25 2.16Clustering, heatmaps etc ... MD, and Smyth, GK (2008). Small sample estimation of negative binomial dis-persion,withapplicationstoSAGEdata. Biostatistics 9,321–332. ... All other questions or problems concerning edgeR should be …

Tags:

  Analysis, Question, Differential, Sequence, Read, Edger, Differential analysis of sequence read

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of edgeR: differential analysis of sequence read count data ...

1 edger : differential analysisof sequence read count dataUser s GuideYunshun Chen1,2, Davis McCarthy3,4, Matthew Ritchie1,2,Mark Robinson5, and Gordon Smyth1,61 Walter and Eliza Hall Institute of Medical Research, Parkville, Victoria, Australia2 Department of Medical Biology, University of Melbourne, Victoria, Australia3St Vincent s Institute of Medical Research, Fitzroy, Victoria, Australia4 Melbourne Integrative Genomics, University of Melbourne, Victoria, Australia5 Institute of Molecular Life Sciences and SIB Swiss Institute of Bioinformatics, University ofZurich, Zurich, Switzerland6 School of Mathematics and Statistics, University of Melbourne, Victoria, AustraliaFirst edition 17 September 2008 Last revised 20 April 2022 Contents1 Introduction.

2 Scope.. Citation.. How to get help.. Quick start..102 Overview of capabilities.. Terminology.. Aligning reads to a genome.. Producing a table of read counts.. Reading the counts from a file.. Pseudoalignment and quasi-mapping.. The DGEList data class.. Filtering.. Normalization.. is only necessary for sample-specific effects.. depth.. library sizes.. content.. length.. normalization, not transformation..162edgeR User s Negative binomial models.. coefficient of variation (BCV).. BCVs.. negative binomial.. The classic edger pipeline: pairwise comparisons betweentwo or more groups.. Estimating dispersions.. Testing for DE genes.

3 More complex experiments (glm functionality).. Generalized linear models.. Estimating dispersions.. Testing for DE genes.. What to do if you have no replicates.. differential expression above a fold-change threshold.. Gene ontology (GO) and pathway analysis .. Gene set testing.. Clustering, heatmaps etc.. Alternative splicing.. CRISPR-Cas9 and shRNA-seq screen analysis .. Bisulfite sequencing and differential methylation analysis ..273 Specific experimental designs.. Introduction.. Two or more groups.. approach.. approach.. and contrasts.. more traditional glm approach.. ANOVA-like test for any differences..343edgeR User s Experiments with all combinations of multiple factors.

4 Each treatment combination as a group.. interaction formulas.. effects over all times.. at any time.. Additive models and blocking.. samples.. effects.. Comparisons both between and within subjects..414 Case studies.. RNA-Seq of oral carcinomas vs matched normal tissue.. in the data.. and normalization.. exploration.. design matrix.. the dispersion.. expression.. ontology analysis .. Setup.. RNA-Seq of pathogen inoculated arabidopsis with batch .. samples.. the data.. and normalization.. exploration.. design matrix.. the dispersion.. expression.. Profiles of Yoruba HapMap individuals..594edgeR User s the data.. and normalization.. the dispersion.

5 Expression.. set testing.. RNA-Seq profiles of mouse mammary gland.. alignment and processing.. loading and annotation.. and normalization.. exploration.. design matrix.. the dispersion.. expression.. testing.. Gene ontology analysis .. Gene set testing.. Setup.. differential splicing after Pasilla knockdown.. samples.. alignment and processing.. loading and annotation.. and normalization.. exploration.. design matrix.. the dispersion.. expression.. Alternative splicing.. Setup.. Acknowledgements.. CRISPR-Cas9 knockout screen analysis .. processing.. and data exploration.. design matrix and dispersion estimation..905edgeR User s representation analysis .

6 Set tests to summarize over multiple sgRNAs targeting thesame gene.. Bisulfite sequencing of mouse oocytes.. in the data.. and normalization.. exploration.. design matrix.. the dispersion.. methylation analysis at CpG loci.. counts in promoter regions.. methylation in gene promoters.. Setup.. Time course RNA-seq experiments of Drosophila .. object.. annotation.. and normalization.. exploration.. design matrix.. the dispersion.. course trend analysis .. 1146 Chapter ScopeThis guide provides an overview of the Bioconductor package edger for differential expres-sion analyses of read counts arising from RNA-Seq, SAGE or similar technologies [32].

7 Thepackage can be applied to any technology that produces read counts for genomic particular interest are summaries of short reads from massively parallel sequencing tech-nologies such as Illumina , 454 or ABI SOLiD applied to RNA-Seq, SAGE-Seq or ChIP-Seqexperiments, pooled shRNA-seq or CRISPR-Cas9 genetic screens and bisulfite sequencingfor DNA methylation studies. edger provides statistical routines for assessing differentialexpression in RNA-Seq experiments or differential marking in ChIP-Seq package implements exact statistical methods for multigroup experiments developed byRobinson and Smyth [34,35]. It also implements statistical methods based on generalizedlinear models (glms), suitable for multifactor experiments of any complexity, developed byMcCarthy et al.

8 [25], Lund et al. [23], Chen et al. [3] and Lun et al. [22]. Sometimes werefer to the former exact methods asclassicedgeR, and the latter asglmedgeR. Howeverthe two sets of methods are complementary and can often be combined in the course of adata analysis . Most of the glm functions can be identified by the letters glm as part of thefunction name. The glm functions can test for differential expression using either likelihoodratio tests[25, 3] or quasi-likelihood F-tests [23,22].A particular feature of edger functionality, both classic and glm, are empirical Bayes methodsthat permit the estimation of gene-specific biological variation, even for experiments withminimal levels of biological can be applied to differential expression at the gene, exon, transcript or tag level.

9 Infact, read counts can be summarized by any genomic feature. edger analyses at the exon levelare easily extended to detect differential splicing or isoform-specific differential guide begins with brief overview of some of the key capabilities of package, and thengives a number of fully worked case studies, from counts to lists of CitationThe edger package implements statistical methods from the following User s GuideRobinson, MD, and Smyth, GK (2008). Small sample estimation of negative binomial dis-persion, with applications to SAGE , 321 the idea of sharing information between genes by estimating the negativebinomial variance parameter globally across all genes. This made the use of negativebinomial models practical for RNA-Seq and SAGE experiments with small to moderatenumbers of replicates.

10 Introduced the terminologydispersionfor the variance parame-ter. Proposed conditional maximum likelihood for estimating the dispersion, assumingcommon dispersion across all genes. Developed an exact test for differential expressionappropriate for the negative binomially distributed counts. Despite the official publica-tion date, this was the first of the papers to be submitted and accepted for , MD, and Smyth, GK (2007). Moderated statistical tests for assessing differencesin tag , 2881 empirical Bayes moderated dispersion parameter estimation. This is a crucialimprovement on the previous idea of estimating the dispersions from a global model, be-cause it permits gene-specific dispersion estimation to be reliable even for small dispersion estimation is necessary so that genes that behave consistentlyacross replicates should rank more highly than genes that do , MD, McCarthy, DJ, Smyth, GK (2010).