Transcription of Incorporation by MMLV-Derived Reverse …
1 Base Preferences in Non- templated NucleotideIncorporation by MMLV-Derived Reverse TranscriptasesPawel Zajac a, Saiful Islam, Hannah Hochgerner b, Peter L nnerberg, Sten Linnarsson*Laboratory for Molecular Neurobiology, Department of Medical Biochemistry and Biophysics, Karolinska Institutet, Stockholm, SwedenAbstractReverse transcriptases derived from Moloney Murine Leukemia Virus (MMLV) have an intrinsic terminal transferaseactivity, which causes the addition of a few non- templated nucleotides at the 3 end of cDNA, with a preference forcytosine. This mechanism can be exploited to make the Reverse transcriptase switch template from the RNAmolecule to a secondary oligonucleotide during first-strand cDNA synthesis, and thereby to introduce arbitrarybarcode or adaptor sequences in the cDNA. Because the mechanism is relatively efficient and occurs in a singlereaction, it has recently found use in several protocols for single-cell RNA sequencing.
2 However, the base preferenceof the terminal transferase activity is not known in detail, which may lead to inefficiencies in template switching whenstarting from tiny amounts of mRNA. Here, we used fully degenerate oligos to determine the exact base preferenceat the template switching site up to a distance of ten nucleotides. We found a strong preference for guanosine at thefirst non- templated nucleotide , with a greatly reduced bias at progressively more distant positions. Based on thisresult, and a number of careful optimizations, we report conditions for efficient template switching for cDNAamplification from single : Zajac P, Islam S, Hochgerner H, L nnerberg P, Linnarsson S (2013) Base Preferences in Non- templated nucleotide Incorporation by MMLV-Derived Reverse transcriptases . PLoS ONE 8(12): e85270. : Luis Men ndez-Arias, Centro de Biolog a Molecular Severo Ochoa (CSIC-UAM), SpainReceived July 29, 2013; Accepted November 26, 2013; Published December 31, 2013 Copyright: 2013 Zajac et al.
3 This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permitsunrestricted use, distribution, and reproduction in any medium, provided the original author and source are : This work was supported by the Swedish Foundation for Strategic Research (MDB09-0052) and the European Research Council (261063/BRAINCELL). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the interests: The authors have declared that no competing interests exist.* Email: a Current address: Illumina Inc., San Diego, California, USA b Current address: Department of Clinical Neuroscience, Karolinska Institutet, Stockholm, SwedenIntroductionThe introduction of second-generation massively parallelsequencing has revolutionized many areas in biological andmedical research. One area greatly benefiting from thesequencing revolution is transcriptomics, where the massiveoutput provided by modern sequencers has shed new light onthe complexity of the RNA landscape.
4 RNA studies usingmassively parallel sequencing are performed using a group ofmethods collectively termed RNA sequencing, or simply RNA-seq[1,2]. Briefly, in RNA-seq the sample of interest is isolated,the RNA is extracted and converted via a number of enzymaticsteps into a sequencing library, to a format suitable forsequencing. Generally, this includes the introduction ofappropriate adaptors to the ends of the molecules and sizeselection to arrive at homogenous fragment sizes fromtranscripts of different lengths. The library is thereaftersequenced, the obtained reads aligned to the genome andtranscriptome with further analysis outlining the transcriptionallandscape. The benefits of massive sequencing, as comparedto the traditionally used microarrays, include a broad dynamicrange, high accuracy and that no a priori information about theRNA is different polymerases show nontemplated nucleotideaddition at the 3 end of an extended template[3].
5 For example,HIV Reverse transcriptase[4] and Taq DNA polymerase bothpreferentially incorporate adenosines, whereas reversetranscriptases from the Moloney murine leukemia virus(MMLV), preferentially introduce cytosines[5].Template switching is a mechanism by which reversetranscriptases (RTs) of the Moloney murine leukemia virus(MMLV) family can switch template from the RNA molecule to asecondary oligonucleotide, called the template-switchingoligonucleotide (TSO), during cDNA synthesis. The templateswitch is enabled by the terminal transferase activity of theMMLV RTs that adds a couple of nucleotides in a template-independent fashion upon reaching the terminus of the RNAmolecule. The TSO can transiently anneal to these protrudingbases by virtue of a complementary ribonucleotide stretch. TheRT then switches template from the RNA to the TSO andcontinues with the cDNA synthesis.
6 In this manner, arbitrarysequences for instance amplification handles or adapters PLOS ONE | 2013 | Volume 8 | Issue 12 | e85270can be incorporated at both ends of the final cDNA molecule: atthe 3 end by tailing the oligo(dT) primer and at the 5 end bydesigning a suitable TSO. For this reason, template switchingis indeed employed in a number of commonly used protocols[6][7].In a number of instances, including development and cancer,studies at the single-cell level can provide information that islost when analyzing whole populations of cells. Population-levelstudies generate average data over the entire cell pool that canmask contributions of important individual cells or cell types. Itis therefore not surprising that the last couple of years, largelydue to technological advances, have seen a great increase inthe number of single-cell studies. For example, genomes ofsingle cancer cells from a breast cancer have been sequencedto investigate the evolution and progress of this disease [8].
7 Furthermore, gene expression analysis of single circulatingmelanoma cells was used to identify potential biomarkers [9].Single-cell studies are challenging because of the scarcity ofthe starting material. Preferably, the protocols should be simplewith minimal numbers of enzymatic steps and of the straightforwardness and ease ofimplementation, template switching has found use in single-cellRNA seq library preparation. For example, a slightly modifiedversion of Clontech s SMARTer Ultra Low RNA Kit wasrecently combined with the standard Illumina library preparationprocedure, as well as with the transposon-based Nexteralibrary construction technique, in SMART-Seq an RNA-seqprotocol capable of transcriptome profiling of single cells [9].Similarly, we have reported Single-cell tagged reversetranscription (STRT), an RNA-seq approach using templateswitching to introduce barcode and amplification sequences[10,11].
8 With STRT up to 96 single cells can be profiled inparallel, translating into reduced cost and increasedthroughput. Briefly, after cell isolation and lysis, each cell smRNA is converted to cDNA, which is simultaneously labeledwith a cellular barcode sequence and also receives a universalamplification handle. The Incorporation of these two features isa result of template-switching events. Consequently, templateswitching is the core of the approach. Following the barcoding,all reactions can be pooled and from this step processed in asingle reaction tube. The Reverse transcription / templateswitching and pooling are followed by PCR amplification wherethe cDNA molecules are amplified using a single universalprimer. The obtained full-length cDNA library is thenreformatted to an Illumina sequencing library using standardmethods. In addition to the streamlined procedure, STRT retains strand information and, moreover, the template-switching event occurs predominantly at the 5 -ends oftranscripts meaning that the position of the transcription startsite (TSS) can be studies have explored the parameters of and theconditions for template switching.
9 Some investigations havefocused on the nature and number of the incorporatednucleotides. For instance, it was demonstrated thatsupplementing the reaction with manganese ions increased thenumber of incorporated nucleotides from a single residue to 3-4[12 ]. Different MMLV Reverse transcriptases , predominantlySuperScript II (SSII; Life / Invitrogen) and SuperScript III (SSIII;Life / Invitrogen), have been examined for their template-switching proficiency [7] [ 13,14]. The TSO has also beenstudied. It has, for example, been shown that 5 -modificationsof this secondary oligonucleotide can minimize the formation ofproducts carrying tandem copies of the TSO [15]. 3 -blocking ofthe TSO has been demonstrated to eliminate spurious primingof the TSO during cDNA synthesis and PCR amplification [16].Finally, the artifacts generated by the template-switchingmechanism have also been the subject of research.
10 A recentstudy put forward the process of strand invasion wherebypremature template switching, at the positions where thesequences of the RNA molecule and the 3 -part of the TSO arethe same, generates truncated, but correctly tagged, cDNAmolecules [17].In this article, we use fully degenerate TSOs to directlydetermine the base preference of the terminal transferaseactivity of SuperScript II. In addition, we report carefuloptimizations of all the relevant parameters in the reaction,which leads to an integrated set of guidelines for efficienttemplate switching in the context of cDNA amplification fromsingle cells. The STRT method served as the framework for theanalyses. However, as has been described above, templateswitching is incorporated into a number of RNA seq protocols,and therefore the results of this work are applicable and MethodsThe scope of the investigations presented in this articlenecessitated a variety of experimental setups.