3.2 Transcription, Translation & Genetic Code
Key Takeaways
- Eukaryotic RNA polymerases display distinct transcriptional targets and α-amanitin sensitivities: Pol I synthesizes 28S, 18S, and 5.8S rRNAs (resistant); Pol II synthesizes pre-mRNA, snRNAs, and miRNAs (highly sensitive); and Pol III synthesizes tRNA and 5S rRNA (moderately sensitive).
- Pre-mRNA maturation requires co-transcriptional 5' m7G capping via an unusual 5'-to-5' triphosphate bridge, 3' polyadenylation driven by the AAUAAA signal, and two-step spliceosomal transesterification at conserved GU-AG intron motifs.
- The genetic code is a degenerate triplet code where 61 sense codons specify 20 amino acids; non-Watson-Crick pairing at the third codon base is governed by Crick's wobble hypothesis.
- Pre-analytical RNA integrity assessment via capillary electrophoresis (28S:18S rRNA ratio of ~2.0 and high RIN scores) is a critical quality control prerequisite for clinical RT-qPCR and RNA sequencing.
3.2 Transcription, Translation & Genetic Code
Quick Summary: Gene expression converts double-stranded DNA templates into functional proteins through transcription, RNA processing, and translation. Eukaryotic transcription is segregated among three specialized RNA polymerases, with RNA Polymerase II synthesizing pre-mRNAs under the control of the core promoter and multi-subunit general transcription factors (TFIIA–TFIIH). Pre-mRNA transcripts undergo obligate co-transcriptional processing: 5' 7-methylguanosine capping, spliceosomal intron excision at invariant GU–AG dinucleotides, and 3' cleavage and polyadenylation driven by the AAUAAA signal. The genetic code is a degenerate triplet code read on the ribosome, where peptide bond formation is catalyzed by rRNA ribozymes.
1. Transcription Machinery: Prokaryotic vs. Eukaryotic Systems
Transcription is the enzymatic synthesis of complementary single-stranded RNA from a deoxyribonucleic acid template strand. Transcription proceeds strictly in the 5' $\rightarrow$ 3' direction, reading the template DNA strand in the 3' $\rightarrow$ 5' direction.
Prokaryotic Transcription Mechanics
Prokaryotes utilize a single multi-subunit RNA polymerase enzyme:
- Core Enzyme ($\alpha_2\beta\beta'\omega$): Catalyzes phosphodiester bond formation but cannot independently recognize promoter sequences.
- Holoenzyme ($\alpha_2\beta\beta'\omega\sigma$): Association of the core enzyme with a transient regulatory sigma ($\sigma$) factor confers sequence-specific promoter binding. In E. coli, the primary housekeeping factor $\sigma^{70}$ recognizes:
- -35 Element: Consensus sequence
5'-TTGACA-3' - -10 Element (Pribnow Box): Consensus sequence
5'-TATAAT-3'
- -35 Element: Consensus sequence
- Termination Modalities:
- Rho-Independent (Intrinsic) Termination: Transcribed RNA forms a stable, self-complementary GC-rich hairpin stem-loop followed immediately by a run of 6–8 uracil (U) residues. The weak rU-dA hybrid base pairs cause the stalled polymerase to dissociate.
- Rho-Dependent Termination: Hexameric ATP-dependent Rho factor helicase binds cytosine-rich rut (Rho utilization) sites on the nascent transcript, translocates along the RNA, and unwinds the RNA-DNA hybrid at the stalled polymerase.
Eukaryotic RNA Polymerase Specialization
Eukaryotes express three distinct nuclear RNA polymerases differentiated by their subcellular localization, transcript specificity, and vulnerability to the fungal toxin $\alpha$-amanitin (derived from the death cap mushroom, Amanita phalloides).
| Enzyme | Subcellular Localization | Primary Transcripts Synthesized | Relative Transcript Output | $\alpha$-Amanitin Sensitivity |
|---|---|---|---|---|
| RNA Polymerase I | Nucleolus | 45S pre-rRNA (cleaved into 28S, 18S, and 5.8S rRNAs) | ~80% of cellular RNA | Insensitive / Resistant (uninhibited) |
| RNA Polymerase II | Nucleoplasm | Pre-mRNA, most snRNAs (U1, U2, U4, U5), microRNAs (miRNAs), lncRNAs | ~5% of cellular RNA | Highly Sensitive (inhibited at <1 $\mu$g/mL) |
| RNA Polymerase III | Nucleoplasm | tRNAs, 5S rRNA, U6 snRNA, 7SL RNA (SRP component) | ~15% of cellular RNA | Moderately Sensitive (inhibited at 10–100 $\mu$g/mL) |
| POLRMT | Mitochondria | Polycistronic mitochondrial heavy/light strand RNAs | Variable | Insensitive to $\alpha$-amanitin; sensitive to ethidium bromide |
┌──────────────────────────────┐
│ Eukaryotic RNA Polymerases │
└──────────────┬───────────────┘
┌───────────────────────────────────┼───────────────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ RNA Polymerase I │ │ RNA Polymerase II│ │RNA Polymerase III│
├──────────────────┤ ├──────────────────┤ ├──────────────────┤
│ • Nucleolus │ │ • Nucleoplasm │ │ • Nucleoplasm │
│ • 28S, 18S, 5.8S │ │ • Pre-mRNA, miRNA│ │ • tRNA, 5S rRNA │
│ • α-Amanitin RES │ │ • α-Amanitin SEN │ │ • α-Amanitin MOD │
└──────────────────┘ └──────────────────┘ └──────────────────┘
RNA Polymerase II Pre-Initiation Complex (PIC)
Initiation of protein-coding genes by RNA Polymerase II requires the stepwise assembly of General Transcription Factors (GTFs) at the core promoter:
- TFIID Binding: The TATA-Binding Protein (TBP) subunit of TFIID binds the minor groove of the TATA box (consensus
5'-TATAAA-3', located at -25 to -30 relative to the +1 transcription start site), introducing an 80° bend in the DNA. TBP-Associated Factors (TAFs) recognize the Initiator (Inr) and Downstream Promoter Element (DPE). - Sequential Assembly: TFIIA stabilizes the TFIID-DNA interaction; TFIIB binds and bridges TFIID with RNA Pol II, establishing the exact transcription start site.
- Recruitment of Pol II & TFIIF: TFIIF escorts RNA Pol II to the promoter.
- Final Assembly: TFIIE binds, recruiting TFIIH.
- Promoter Clearance (TFIIH Dual Function): TFIIH possesses ATP-dependent DNA helicase activity (XPB/XPD subunits) to melt promoter DNA into an open transcription bubble, and cyclin-dependent kinase activity (CDK7 subunit) that phosphorylates serine residues (specifically Ser5) on the heptapeptide repeat (
Tyr1-Ser2-Pro3-Thr4-Ser5-Pro6-Ser7) of the C-Terminal Domain (CTD) of RNA Pol II. CTD phosphorylation releases Pol II from the PIC to initiate elongation.
2. Eukaryotic Post-Transcriptional Pre-mRNA Processing
In eukaryotic organisms, nascent transcripts undergo three coupled maturation events before nucleocytoplasmic export:
Nascent Pre-mRNA: 5' ─── Exon 1 ─── Intron 1 ─── Exon 2 ─── Intron 2 ─── Exon 3 ─── 3'
│
Co-transcriptional Maturation
▼
Mature mRNA: m7G-ppp-5' ─── Exon 1 ─ Exon 2 ─ Exon 3 ─── AAAAAAAAAAAAAAAAAAAA-3' (Poly-A Tail)
1. 5' 7-Methylguanosine (m7G) Capping
Occurs co-transcriptionally when the nascent transcript reaches 20–30 nucleotides in length:
- An RNA triphosphatase removes the $\gamma$-phosphate from the 5' terminal nucleotide.
- Guanylyltransferase condenses GMP from GTP with the 5'-diphosphate, creating an unusual 5'-to-5' triphosphate linkage ($5^{\prime}\text{-ppp-}5^{\prime}$).
- Guanine-N7-methyltransferase adds a methyl group from S-adenosylmethionine (SAM) to nitrogen-7 of the terminal guanine.
- Biological Role: Protects mRNA from 5' $\rightarrow$ 3' exonucleases (XRN1/XRN2), facilitates nuclear export via the Cap-Binding Complex (CBC), and serves as the primary recognition docking site for the eukaryotic translation initiation factor eIF4E.
2. 3' Cleavage & Polyadenylation
- As Pol II transcribes the 3' untranslated region (UTR), it encounters the consensus polyadenylation signal sequence:
- Cleavage and Polyadenylation Specificity Factor (CPSF) binds AAUAAA, while Cleavage Stimulation Factor (CstF) binds a downstream GU-rich or U-rich tract.
- Endonucleolytic cleavage occurs 10–30 nucleotides downstream of the AAUAAA motif.
- Poly(A) Polymerase (PAP) synthesizes a polyadenylate tail of 200–250 adenine residues in a template-independent reaction utilizing ATP.
- Biological Role: Poly(A)-Binding Protein (PABP) coats the tail, promoting nuclear export, translational circularization via eIF4G interaction, and protection against 3' deadenylases.
3. Pre-mRNA Splicing & Spliceosomal Catalysis
Splicing is the precise excision of non-coding introns and ligation of coding exons, orchestrated by the spliceosome—a dynamic ribonucleoprotein complex comprising five small nuclear RNAs (U1, U2, U4, U5, and U6 snRNAs) complexed with Sm core proteins to form snRNPs.
Conserved Intron Consensus Sequences
- 5' Splice Donor Site: Invariant
5'-GU-3'dinucleotide at the 5' intron boundary (consensus:5'-MAG|GURAGU-3'). - 3' Splice Acceptor Site: Invariant
5'-AG-3'dinucleotide at the 3' intron boundary (consensus:5'-YNCAG|G-3'). - Branch Point Sequence: Located 18–40 nucleotides upstream of the 3' splice site; contains an invariant branch point Adenosine (A) residue.
- Polypyrimidine Tract: A pyrimidine-rich tract ($Y_n$, where $Y = \text{U or C}$) situated between the branch point and the 3' AG acceptor.
Two-Step Transesterification Chemistry
Splicing proceeds through two isoenergetic transesterification reactions without external ATP hydrolysis for the phosphodiester transfers:
- First Transesterification: The $2^{\prime}\text{-OH}$ group of the invariant branch point adenosine executes a nucleophilic attack on the phosphodiester bond at the 5' splice donor ($GU$) site. This cleaves the 5' exon-intron boundary and forms an unusual $2^{\prime}\text{-}5^{\prime}$ phosphodiester bond, generating a free 5' exon and a branched intron-exon lariat intermediate.
- Second Transesterification: The newly liberated $3^{\prime}\text{-OH}$ group of the 5' exon executes a nucleophilic attack on the phosphodiester bond at the 3' splice acceptor ($AG$) site. This ligates the two exons via a standard $3^{\prime}\text{-}5^{\prime}$ phosphodiester bond and releases the excised intron lariat, which is subsequently linearized by RNA debranching enzyme (Dbr1) and degraded.
Step 1: 5' Exon ─[5' GU]───────── (A) ───[3' AG]─ Exon 3'
▲ │
└─────────────┘ (2'-OH attacks 5' splice site)
Step 2: 5' Exon-3'OH + Lariat Intron ───[3' AG]─ Exon 3'
│ ▲
└──────────────────────────────┘ (3'-OH attacks 3' splice site)
▼
Exon 1 ─── Exon 2 + Excised Lariat Intron
3. The Genetic Code & Wobble Hypothesis
The genetic code translates a four-letter nucleic acid alphabet into a twenty-letter amino acid alphabet via triplet nucleotide units termed codons.
Hallmarks of the Genetic Code
- Triplet Nature: 64 total codons ($4^3$). 61 codons specify amino acids (sense codons), while 3 codons specify translation termination (nonsense / stop codons).
- Degenerate (Redundant): Multiple codons encode the same amino acid (e.g., Leucine, Serine, and Arginine are each encoded by 6 distinct codons). Only Methionine (
AUG) and Tryptophan (UGG) are specified by single unique codons. - Non-Overlapping & Commaless: Codons are read sequentially in contiguous groups of three without punctuation or skipped nucleotides.
- Universal (with exceptions): Shared across almost all extant life, with minor variations in human mitochondrial DNA (mtDNA:
UGA= Tryptophan instead of Stop;AUA= Methionine instead of Isoleucine;AGA/AGG= Stop instead of Arginine).
| Codon Classification | Codon Triplet Sequences | Amino Acid / Function |
|---|---|---|
| Initiation (Start) Codon | 5'-AUG-3' | Encodes Methionine (Met) in eukaryotes; $N$-formylmethionine (fMet) in bacteria and mitochondria |
| Termination (Stop) Codons | 5'-UAA-3' (Ochre)<br>5'-UAG-3' (Amber)<br>5'-UGA-3' (Opal) | Recognized by protein Release Factors (eRF1 in eukaryotes; RF1/RF2 in bacteria); zero cognate tRNAs exist |
Crick's Wobble Hypothesis
While cells translate 61 sense codons, most organisms express fewer than 45 distinct tRNA species. Francis Crick proposed the Wobble Hypothesis: base-pairing rules are strictly enforced at codon positions 1 and 2 (pairing with tRNA anticodon positions 3 and 2), but steric conformational flexibility at the 3' position of the mRNA codon (pairing with the 5' position of the tRNA anticodon) allows non-standard Watson-Crick hydrogen bonding.
| tRNA Anticodon 5' Base | mRNA Codon 3' Base Recognized | Pairing Mechanism / Notes |
|---|---|---|
| C | G | Strict Watson-Crick pairing |
| A | U | Strict Watson-Crick pairing |
| U | A or G | Wobble allows single tRNA to decode two purine-ending codons |
| G | C or U | Wobble allows single tRNA to decode two pyrimidine-ending codons |
| I (Inosine) | U, C, or A | Adenosine deaminated by adenosine deaminase acting on tRNA (ADAT); tri-wobble decoding |
4. Ribosomal Translation Mechanics
Translation converts mRNA sequence into a polypeptide chain across four concerted stages: tRNA charging, initiation, elongation, and termination.
Ribosome Architecture
- Prokaryotic 70S Ribosome: Composed of a 50S large subunit (23S rRNA [peptidyl transferase ribozyme], 5S rRNA, 31 proteins) and a 30S small subunit (16S rRNA [decoding center], 21 proteins).
- Eukaryotic 80S Ribosome: Composed of a 60S large subunit (28S rRNA [peptidyl transferase ribozyme], 5.8S rRNA, 5S rRNA, 47 proteins) and a 40S small subunit (18S rRNA [decoding center], 33 proteins).
Translation Cycle Stages
- tRNA Aminoacylation (Charging): 20 specific aminoacyl-tRNA synthetases activate amino acids using ATP to form an aminoacyl-AMP intermediate, followed by esterification to the 2'- or 3'-OH of the invariant terminal 3'-CCA acceptor stem of cognate tRNA. Synthetases possess proofreading hydrolytic sites that deacylate mispaired amino acids, maintaining an error rate below $10^{-4}$.
- Translation Initiation:
- Prokaryotes: The 30S small subunit aligns its 16S rRNA with the Shine-Dalgarno consensus sequence (
5'-AGGAGG-3') located ~8 nucleotides upstream of the start codon. - Eukaryotes: The 40S small subunit, pre-loaded with Met-tRNA$_i^{\text{Met}}$ and eIF2-GTP (ternary complex), docks at the 5' m7G cap via eIF4F (eIF4E, eIF4G, eIF4A helicase) and scans unidirectionally 5' $\rightarrow$ 3' through the 5' UTR until encountering the Kozak consensus motif: Hydrolysis of eIF2-GTP triggers 60S large subunit joining, positioning Met-tRNA$_i^{\text{Met}}$ directly into the ribosomal P (Peptidyl) site.
- Prokaryotes: The 30S small subunit aligns its 16S rRNA with the Shine-Dalgarno consensus sequence (
- Elongation:
- A (Aminoacyl) Site: Elongation factor eEF1A-GTP delivers incoming aminoacyl-tRNA complementary to the codon. Correct codon-anticodon pairing stimulates GTP hydrolysis and factor release.
- Peptide Bond Formation: The catalytic 28S rRNA ribozyme in the 60S subunit catalyzes peptide bond synthesis, transferring the growing peptide chain from the P-site tRNA to the $\alpha$-amino group of the A-site amino acid.
- Translocation: Elongation factor eEF2-GTP hydrolyzes GTP to translocate the ribosome exactly 3 nucleotides downstream along the mRNA. Uncharged tRNA moves from P $\rightarrow$ E (Exit) site for discharge, while peptidyl-tRNA moves from A $\rightarrow$ P site.
- Termination: Upon positioning a stop codon (
UAA,UAG,UGA) in the A site, eukaryotic Release Factor 1 (eRF1), which structurally mimics a tRNA molecule, binds the stop codon. eRF3-GTP stimulates peptidyl transferase to transfer the polypeptide to a water molecule rather than an amino acid, hydrolyzing the ester bond and releasing the finished protein. Ribosome recycling factor (ABCE1) dissociates the 80S complex.
5. Translational Regulation & Non-Coding RNAs
Gene expression is dynamically modulated post-transcriptionally by non-coding regulatory RNAs:
- microRNAs (miRNAs): Endogenous ~21–23 nucleotide non-coding RNAs that mediate post-transcriptional gene silencing.
- Transcribed by RNA Pol II as primary transcripts (pri-miRNAs) containing local stem-loop hairpins.
- The nuclear microprocessor complex (Drosha-DGCR8) crops pri-miRNA into a ~70 nt precursor hairpin (pre-miRNA).
- Exported to the cytoplasm via Exportin-5.
- Cleaved in the cytoplasm by the RNase III endonuclease Dicer into an asymmetric ~22 bp double-stranded miRNA duplex.
- The mature guide strand is loaded into the RNA-Induced Silencing Complex (RISC) containing an Argonaute (AGO2) core protein. The miRNA seed region (nucleotides 2–8 at the 5' end) binds complementary sequences in the 3' UTR of target mRNAs, directing translational repression or exonucleolytic deadenylation and mRNA decay.
6. Pre-Analytical RNA Quality Metrics & Reverse Transcription in the Clinical Laboratory
Because RNA is susceptible to ubiquitous environmental ribonucleases (RNases) and spontaneous alkaline 2'-OH auto-hydrolysis, pre-analytical validation is mandatory in molecular pathology.
Assessing Total RNA Integrity: The Bioanalyzer & RIN Score
In clinical next-generation sequencing and microarray workflows, microfluidic capillary electrophoresis (e.g., Agilent 2100 Bioanalyzer, TapeStation) analyzes ribosomal RNA peaks to quantify degradation:
- 28S:18S Ribosomal Ratio: Intact eukaryotic total RNA displays two prominent electrophoretic peaks corresponding to the 28S (~5.0 kb) and 18S (~1.9 kb) rRNAs. Intact, un-degraded RNA yields a 28S:18S peak area ratio of ~2.0:1. Progressive degradation causes collapse of the 28S peak prior to 18S degradation, shifting the ratio below 1.0.
- RNA Integrity Number (RIN): A software algorithm calculating an objective numerical score from 1 (completely degraded) to 10 (completely intact) based on the entire electropherogram baseline, 28S/18S peak heights, and low-molecular-weight degradation products.
- RIN $\ge$ 7.0–8.0: Required for whole-transcriptome RNA-Seq and full-length cDNA cloning.
- DV200 Metric: For highly degraded formalin-fixed paraffin-embedded (FFPE) clinical specimens where RIN scores are artificially low, the DV200 (percentage of RNA fragments >200 nucleotides) is utilized; a DV200 $\ge$ 30–50% is required for library preparation.
Intact RNA (RIN 10): Degraded RNA (RIN 2):
Fluorescence Fluorescence
│ 18S 28S │
│ │ │ │ Degradation
│ │ │ │ "Smear"
│ ┌┴┐ ┌┴┐ │ ┌───┐
│ Marker│ │ │ │ │Marker│
└──┴─────┴─┴────┴─┴───▶ └──┴───┴───────────────▶
Size (nt) Size (nt)
Reverse Transcription (RT) Priming Strategies for cDNA Synthesis
Converting RNA into complementary DNA (cDNA) via viral reverse transcriptases (M-MLV or AMV RT) requires targeted selection of priming chemistries:
| Priming Strategy | Primer Mechanism | Advantages | Disadvantages / Clinical Selection |
|---|---|---|---|
| Oligo(dT) Primers | 12–18 nt synthetic deoxythymidine tract annealing to mature eukaryotic mRNA poly(A) tails | Selectively enriches for protein-coding mRNAs; yields full-length cDNA from high-quality total RNA | Fails on prokaryotic RNA, non-polyadenylated transcripts (e.g., microRNAs, histones), and fragmented FFPE RNA with severed tails |
| Random Hexamers / Octamers | Degenerate 6-mer or 8-mer oligonucleotides annealing indiscriminately along transcripts | Primes all RNA classes (mRNA, rRNA, tRNA, viral genomes); superior for fragmented or FFPE RNA | Primes abundant ribosomal RNA, reducing sequencing library coverage efficiency for low-abundance targets |
| Gene-Specific Primers (GSP) | Custom oligonucleotide designed to anneal exclusively to the target transcript of interest | Delivers maximum analytical sensitivity and specificity; minimal background | Restricted to single-target assays; required for clinical viral load assays (e.g., HIV-1, HCV quantitative RT-qPCR) |
Which eukaryotic RNA polymerase is responsible for synthesizing ribosomal 5S rRNA and transfer RNAs (tRNAs), and shows intermediate sensitivity to inhibition by α-amanitin?
During pre-mRNA splicing catalyzed by the spliceosome, what chemical event initiates the first transesterification step?
A clinical molecular laboratory assesses total RNA extracted from a frozen tumor resection using capillary electrophoresis prior to RNA sequencing. Which result represents high-quality, intact total RNA suitable for whole-transcriptome sequencing?