Machine Traces of Discovery Paths #3 — Lab Notebook
Evidence for the Deep Dive essay Machine Traces of Discovery Paths #3 — The Path to CRISPR-Cas9 (Jennifer Doudna), A Comprehensive Analysis and Reconstruction.
This Lab Notebook is the verifiable source record behind the #3 Deep Dive. It reconstructs, from the primary literature, the path that carried Jennifer Doudna from engineering catalytic RNA (1989) to the programmable dual-RNA-guided Cas9 (2012). It mirrors the method of #2 (the iPS path): a nine-paper skeleton read at the level of full text (Methods and Results), with each experiment carrying a verbatim quotation and each protocol value tagged by provenance — [P] for a value printed in the paper, [S]/[M] for a value the paper omits and that must be supplied from standard practice.
Scope rule (this is Doudna's trajectory). Every node is a paper Doudna authored. Foundational works she did not author — Mojica's CRISPR loci, Jansen (2002, cas genes/naming), Barrangou & Horvath (2007, adaptive immunity), Brouns et al. (2008, "small CRISPR RNAs guide antiviral defence" — confirmed: no Doudna in the author list), Marraffini & Sontheimer (2008, DNA is the target), Deltcheva & Charpentier (2011, tracrRNA), Sapranauskas/Gasiūnas/Šikšnys (2011–2012, parallel Cas9) — are credited as influences (see the influence section), not as nodes. A trace that read as a solo effort would be wrong.
Honest coverage note (read before trusting a quote). Unlike #2, whose nine papers were read from a single clean full-text file, the papers here arrived as a mix of PDF qualities. Extraction quality is recorded per paper and summarised in Evidence 4: the PMC author-manuscript papers (H7 Csy4, H8 Cascade, H9 Cas9) and the 2001 IRES and 2006 Dicer reprints extracted as clean machine-readable text; the 1989 origin paper (H1) was a scan recovered by rendering pages to images and transcribing by vision; the 1991 (H2) and 1996 (H3) Science papers exist only as scanned reprints (there was no born-digital typesetting in that era), and were recovered at high fidelity by rendering the pages to images and transcribing them by vision — the same route used for H1 — which eliminated the earlier (OCR-uncertain) flags; and the 1998 HDV ribozyme (H4), added after the initial build, is a clean born-digital copy read in full. Where a value or quote could not be verified, it is flagged rather than presented as certain.
The skeleton of the path
The nine papers, all Doudna-authored, fall into four phases:
| # | Year | Phase | Paper (short) | Doudna position |
|---|---|---|---|---|
| H1 | 1989 | Engineering catalytic RNA | RNA-catalysed synthesis of complementary-strand RNA | first |
| H2 | 1991 | Engineering catalytic RNA | A multisubunit ribozyme (catalyst + template) | first |
| H3 | 1996 | RNA structure | Crystal structure of a group I ribozyme domain (P4-P6) | last |
| H4 | 1998 | RNA structure | Crystal structure of a hepatitis delta virus ribozyme | last |
| H5 | 2001 | RNA directs a machine | HCV IRES RNA reshapes the 40S ribosomal subunit | middle |
| H6 | 2006 | Guide-RNA machinery | Structural basis for dsRNA processing by Dicer | last |
| H7 | 2010 | CRISPR | Sequence/structure-specific RNA processing by Csy4 | last |
| H8 | 2011 | CRISPR | Structures of the RNA-guided surveillance complex (Cascade) | senior co-author |
| H9 | 2012 | CRISPR | A programmable dual-RNA-guided DNA endonuclease (Cas9) | co-senior/corresp. |

The chain of Jennifer Doudna's path to CRISPR-Cas9, 1989–2012: nine papers, four phases, the carry-forwards, and the split/fuse mirror.
The nine-paper skeleton as a chain. Unlike the iPS path's reagent lineage, the material through-line here is a method and a set of hands — the RNA-crystallography pipeline from 1996 and shared personnel (Zhou across H3·H4·H6·H7·H8; Jinek H7→H9; Wiedenheft H7→H8). Evidence 2 traces each arrow.
The reconstruction reads each paper in the order: what question it set, what experiments it ran, what they produced, what it handed to the next paper — with the values printed in the paper and a verbatim quotation as evidence.
Machine Traces #3 — Evidence 1 of 4 · Mechanism-level chain
The mechanism-level reconstruction of the nine-paper path, read from the full text (Methods/Results) of each paper (H1 recovered by vision-OCR of a scan; H4 added and read after the initial build). Each experiment carries a verbatim quotation confirmed against the source; each protocol value carries a provenance tag.
H1 — Engineering the Tetrahymena ribozyme into a template-directed RNA ligase that builds a complementary strand
Doudna JA, Szostak JW. RNA-catalysed synthesis of complementary-strand RNA. Nature 1989;339(6225):519-522. PMID 2660003. (Doudna first author, Szostak lab.)
The question this paper set out to answer — Could a ribozyme be made to act as the beginnings of an RNA replicase? The self-replicating RNA hypothesized to bootstrap life from prebiotic chemistry demands an RNA that copies an arbitrary template. Prior polymerase experiments with the Tetrahymena group I intron (Zaug/Cech; Been/Cech) added nucleotides onto a primer, but only ~15 nt and beyond the template's end, with three unmet requirements: the template had to be a separate molecule from the enzyme, polymerization had to be template-directed, and the intron's built-in nucleotide/base-pairing specificity had to be overcome so an arbitrary sequence could be copied. Doudna and Szostak asked whether a modified Tetrahymena ribozyme could overcome all three problems and catalyse formation of an RNA strand complementary to an external template.
Experiments and results
H1.1 — Cleavage of an isolated P1 substrate by the modified ribozyme (independent template)
What was done: They synthesized the isolated P1 stem-loop as a separate molecule and a modified ribozyme spanning stems P2–P9 (missing P1, P9.1, P9.2, and the 3′ intron-exon junction; T7 runoff transcript of pJD1100 cut with NheI). They tested site-specific guanosine attack on the free P1 substrate (Fig. 1, Fig. 2).
What came out: The enzyme catalysed site-specific attack of guanosine on the isolated P1 stem, but the K_m for free P1 was very high (>0.1 mM), reflecting weak enzyme–substrate interaction because the core intron has few sequence/size requirements for recognizing P1.
Evidence: "We found that this enzyme RNA catalyses the site-specific attack of guanosine on the isolated P1 stem (Fig. 2), but that the K_m for free P1 was very high (>0.1 mM)."
H1.2 — Reverse reaction: template-directed regeneration of intact P1 (two-oligonucleotide ligation)
What was done: They synthesized two RNA oligonucleotides corresponding to the products of guanosine attack on P1, differing from wild-type P1 only in a U·G base pair at the cleavage site. These anneal into a primer + partial-hairpin complex with a guanosine extruded at the gap; incubation with the modified ribozyme was tested for the reverse (ligation) reaction (Fig. 3).
What came out: Extremely efficient regeneration of intact P1 with release of free guanosine — the reverse of Fig. 1a. Reaction essentially complete after one hour, with about 250 turnovers of substrate per enzyme.
Evidence: "Incubation of this complex with the modified ribozyme resulted in the extremely efficient regeneration of intact P1, with the release of free guanosine".
H1.3 — Base-pair requirement at the ligation junction
What was done: Four primers ending in A, C, G or U and four partial hairpins with A, C, G or U opposite the last primer base were made; all 16 primer–hairpin combinations were tested for ligation (Fig. 4a,b).
What came out: Unlike the cleavage reaction, only U·G and C·G combinations gave efficient ligation, with A·G to a lesser extent — the reaction retained an inherent sequence/geometry specificity.
Evidence: "only the U·G and C·G base combinations allow efficient ligation. Ligation occurs to a lesser extent with the A·G combination."
H1.4 — Spermidine overcomes the sequence specificity
What was done: The same primer–template combinations were tested under varied conditions, including addition of 5 mM spermidine; ten related polyamines were compared (Fig. 4c).
What came out: 5 mM spermidine gave efficient ligation of all substrate complexes with either Watson–Crick or wobble (U·G, C·A, A·G) junctions. Spermidine was most effective (putrescine/spermine lesser), so a faithful all-Watson–Crick copy of a template should be possible.
Evidence: "This led to the efficient ligation of all the substrate complexes with either Watson–Crick or wobble (U·G, C·A or A·G) base pairs at the ligation junction."
H1.5 — Template-directed ligation of a primer and ligator on a separate template
What was done: Dispensing with the P1 loop, they aligned two short oligonucleotides (5′ = primer, 3′ = ligator beginning with an unpaired G) on a longer third template oligonucleotide. Seven primer/ligator/template combinations were designed: three with U·G junctions, four with Watson–Crick junctions; products characterized by T1 digestion (Fig. 5).
What came out: The ribozyme ligated primer and ligator in a template-dependent manner in every case. Turnovers ranged from 100 (15-min U·G reactions) to 30–100 (60 min for C·G, G·C, A·U) to a low of 5 in 60 min for U·A. Expected T1 fragments were seen — ligation independent of sequence.
Evidence: "in each case, the ribozyme catalysed the ligation of the primer and the ligator in a template-dependent manner (Fig. 5)."
H1.6 — Ligation of multiple aligned oligonucleotides → complementary-strand synthesis
What was done: Three sets of templates/oligonucleotides (Fig. 6a): (1) four identical oligos on one long template, U·G junctions; (2) four different oligos on one template (to exclude a dissociation/reannealing artifact), U·G junctions; (3) five small RNAs on a template with a different Watson–Crick base pair at each of the four junctions. Equimolar template and ribozyme (both 5 µM).
What came out: Products from ligation of two, three or all four oligos appeared within 5 min; full-length/short-product ratio rose over time with release of free GMP. The four-different-oligo case behaved like the first, excluding the reannealing model; full-length product sequence confirmed by dideoxy sequencing with reverse transcriptase. The five-RNA all-Watson–Crick template gave ~5% full-length (fully complementary) product plus shorter products; reaction was completely dependent on template and enzyme.
Evidence: "We observed ~5% ligation to full-length product, with larger amounts of shorter products also accumulating (Fig. 6d). The reaction was completely dependent on template and enzyme."
What it made possible next
This paper delivers, for the first time, an engineered, guide-programmable RNA catalyst: a truncated Tetrahymena group I intron (P2–P9) recast as a template-directed RNA ligase whose product is set entirely by base-pairing of substrate oligonucleotides to a separate external template rather than to the intron's own internal guide sequence. The template/guide logic — an RNA enzyme whose specificity comes from Watson–Crick pairing to an external strand, decoupled from the catalytic core — is the conceptual seed that runs forward through the lab's ribozyme-engineering program to the 1991 multisubunit ribozyme (H2) and beyond. The reagents established here (the sunY/Tetrahymena intron constructs, e.g. plasmid pJD1100; the P1-stem, primer, ligator and template oligonucleotide design; spermidine to erase intrinsic specificity) become the toolkit for dissecting and re-assembling group I introns into engineered, multi-component catalysts. It frames the concrete next problems — increase full-length efficiency, defeat template secondary structure, and separate product from template — that the subsequent work takes up on the road to a self-replicating RNA.
Protocol extraction — [P] values printed in the paper
- [P] Modified ribozyme: Tetrahymena intron spanning stems P2–P9, missing P1 stem-loop, P9.1, P9.2, and 3′ intron-exon junction; made by T7 RNA polymerase runoff transcription of plasmid pJD1100 digested with NheI.
- [P] Ribozyme purification: 6% polyacrylamide / 7 M urea gel or Sephadex G-75 spin column; phenol extraction, ethanol precipitation; stored in distilled water at ~−80 °C.
- [P] K_m for free P1 substrate: very high, >0.1 mM (100 µM).
- [P] Fig. 2 (P1-like substrate cleavage) reactions: 10 mM NH4Cl, 20 mM MgCl2, 30 mM Tris-HCl pH 7.4, 1 mM aurin tricarboxylic acid ("aurin trichloroacetic acid" as printed — OCR/typo in paper), 0.2 µM enzyme, 200 µM [α-32P]GTP, P1 RNA titrated 2–277 µM, 5 µl reaction; incubated 58 °C for 20 min.
- [P] Fig. 2 stop/analysis: equal volume 90% formamide, 10 mM Tris-HCl pH 8.0, 1 mM EDTA, 0.2% bromophenol blue and xylene cyanol; 15% acrylamide/7 M urea gel. Major product band = 27 nt.
- [P] Fig. 3 (P1 regeneration) reactions: 10 mM NH4Cl, 20 mM MgCl2, 30 mM Tris-HCl pH 7.4, 1 mM aurin tricarboxylic acid, 0.2 µM enzyme, 50 µM each RNA substrate oligonucleotide, 5 µl total; 58 °C; time course 0, 15, 30, 45, 60 min. 5′ primer internally labelled with 0.5 µCi/µl [α-32P]GTP in transcription.
- [P] P1 regeneration kinetics: essentially complete after 1 hour; ~250 turnovers of substrate per enzyme.
- [P] Base-pair test (Fig. 4): 20 µM oligonucleotides, 58 °C for 60 min (otherwise as Fig. 2/3 conditions).
- [P] Spermidine: 5 mM added; ligates all Watson–Crick and wobble (U·G, C·A, A·G) junctions. Ten polyamines tested; spermidine most effective, putrescine and spermine lesser.
- [P] Two-oligo ligation turnovers (Fig. 5): 100 in 15 min for U·G substrates; 30–100 in 60 min for C·G, G·C, A·U; low of 5 in 60 min for U·A.
- [P] Fig. 5 conditions: 0.2 µM enzyme, 20 µM oligonucleotides, 5 mM spermidine; primer RNAs 5′-labelled with [γ-32P]ATP + polynucleotide kinase; stopped with 3 vol 9.3 M urea, 33 mM EDTA; 20% polyacrylamide, 90% formamide gel.
- [P] Multiple-oligonucleotide ligation (Fig. 6): both enzyme and substrate at 5 µM (equimolar template and ribozyme); products of 2/3/4-oligo ligation visible within 5 min; ~5% full-length product for the five-RNA all-Watson–Crick template; 20% polyacrylamide / 90% formamide gels. Time courses: Fig. 6b lanes 0/5/30/60/90 min; Fig. 6c lanes 0/15/30/45/60 min; Fig. 6d lanes 0/30/60/120 min.
- [P] Product-strand verification: full-length product sequence confirmed by dideoxy sequencing with reverse transcriptase.
- [P] Substrate oligonucleotide sequences printed: e.g. P1 primer 5′-GGGAGAGGCU...(OH) / template strand 3′-CUCCGG (Figs 3a, 4a); Fig. 6a repeating unit GGAGUAGCAU aligned on template 3′-CCUCAUCGUG.
- [P] Received 14 April; accepted 8 May 1989.
Gaps requiring [S]/[M]
- [S] Exact sequence of the modified ribozyme transcript and the full pJD1100 construct map / cloning details (given only by reference and Fig. 1b schematic).
- [S] Precise ribozyme concentration in the Fig. 6 "equimolar" multiple-ligation reactions is stated as 5 µM, but molar ratios of the several distinct substrate oligonucleotides to each other are not individually tabulated.
- [S] Absolute rate constants / k_cat for cleavage and ligation (only turnover counts and qualitative K_m > 0.1 mM given).
- [S] Quantitative yields for individual Fig. 5 and Fig. 6b/6c reactions beyond the single ~5% full-length figure for Fig. 6d.
- [S] Full T1-digestion fragment sizes used to characterize each ligated product (stated only as "expected fragments were observed").
- [S] Composition of "RNA oligonucleotide size markers" (M lanes) and exact marker nucleotide lengths beyond the labelled ladder positions (15/20/28/33/34/38/42/43 nt).
- [M] Reason for the printed reagent name "aurin trichloroacetic acid" — almost certainly aurintricarboxylic acid (an RNase inhibitor); needs confirmation, flagged as likely typo in the original.
Anomalies / caveats the authors flag
- Weak enzyme–substrate interaction: very high K_m (>0.1 mM) for free P1, because the core intron imposes few sequence/size requirements on P1 recognition. The authors argue this weakened interaction (template not base-paired to the enzyme) is the simplest explanation for the rapid turnover seen in ligation.
- Intrinsic base-pair specificity of the ligation reaction (only U·G, C·G efficient; A·G weak) — an obstacle to copying an arbitrary template, overcome only by adding 5 mM spermidine.
- Inefficiency of full-length complementary-strand synthesis: only ~5% full-length product from the five-oligonucleotide all-Watson–Crick template, with shorter products dominating.
- Requirement for high Mg2+ (20 mM MgCl2) and a specific substrate secondary structure (the extruded/unpaired 5′ guanosine on the ligator, paired-stem geometry at the junction).
- Authors explicitly list unresolved refinements for autocatalytic replication: increase efficiency of full-length synthesis, overcome deleterious template secondary structure, and find a way to separate product and template strands.
H2 — A multisubunit sunY ribozyme that both catalyzes and templates complementary-strand RNA synthesis
Doudna JA, Couture S, Szostak JW. A multisubunit ribozyme that is a catalyst of and template for complementary strand RNA synthesis. Science 1991;251(5001):1605-1608. PMID 1707185. (Doudna first author, Szostak lab.)
The question this paper set out to answer — RNA molecules capable of self-replication are postulated to have been important in the early evolution of life, but an RNA replicase probably requires a folded structure to carry out efficient catalysis, while an optimal template is unstructured. These conflicting requirements might be accommodated by a multisubunit RNA replicase, whose dissociated components serve as templates and whose assembled complex serves as the polymerase. Group I self-splicing introns catalyze phosphodiester exchange reactions like those of DNA and RNA polymerases, and the Tetrahymena intron catalyzes complementary-strand RNA (cRNA) synthesis on external templates — but very inefficiently (less than 1% of a 40-nt template copied to full length), and self-replication of a ribozyme as large as the Tetrahymena intron (413 nt) is impossible at that efficiency. The paper asks whether smaller, more efficient forms of the self-splicing sunY intron from bacteriophage T4 can be dissected into short trans-acting subunits that reassemble into an active ribozyme and can template-direct the assembly of oligonucleotides into a complementary strand — a model for how prebiotically synthesized oligonucleotides might assemble into a complex capable of self-replication.
Experiments and results
H2.1 — Stabilizing the sunY catalytic domain with td-intron base changes (JD929, 180 nt)
What was done: Because small sunY deletion derivatives (down to 164 nt) retained catalytic activity but were unstable (active only at high monovalent cation and Mg), certain bases in a P1-lacking sunY derivative were changed to those found at the corresponding positions in the larger td intron of phage T4; the sunY-td chimera was assayed in a template-directed oligonucleotide ligation reaction with a separate P1-like substrate.
What came out: The td nucleotide changes gave a three- to fourfold enhancement of the rate of ligation, yielding the more stable derivative JD929 (180 nt) used in all subsequent experiments.
Evidence: "The td nucleotide changes resulted in a three-to fourfold enhancement of the rate of ligation."
H2.2 — Interrupting JD929 at loops L6 and L8 into three trans-acting subunits (fragments A, B, C)
What was done: The RNA chain was interrupted at the non-conserved loops L6 and L8 (chosen because their sizes/sequences are not phylogenetically conserved and interruption at L6 in the td intron does not inactivate splicing, and because these sites yield fragments with minimal secondary structure and substantial complementarity to promote template function and assembly), producing three fragments made by in vitro transcription of synthetic DNA templates.
What came out: Fragment A (5' end through L6), fragment B (3' side of P6a to L8), and fragment C (3' side of P8 through P9) are 59, 75, and 43 nucleotides long, respectively, and assemble into an active three-subunit ribozyme; a 60-min reaction lacking any single fragment (−A, −B, −C) showed the fragment is required.
Evidence: "Fragments A, B, and C are 59, 75, and 43 nucleotides in length, respectively."
H2.3 — Comparing single-chain versus three-subunit ribozyme ligation activity
What was done: The single-chain sunY-td ribozyme and the corresponding three-subunit ribozyme were assayed for template-directed oligonucleotide ligation activity of the P1-like substrate over a time course (0, 15, 30, 45, 60 min) on a 20% acrylamide–7 M urea gel.
What came out: The multisubunit enzyme was about twofold less active at 42°C than the single-chain enzyme and had a lower temperature optimum (37° versus 45°C), a difference attributed to the multisubunit enzyme being only partially assembled into active complexes.
Evidence: "The multisubunit enzyme was approximately twofold less active at 42°C than the corresponding single-chain enzyme and had a lower temperature optimum (37° versus 45°C)."
H2.4 — Template-directed assembly of multiple oligonucleotides into a full-length complementary strand
What was done: A 36-nt RNA was used as a template for cRNA synthesis; four short RNA oligomers (one 9-nt and three 10-nt) that anneal to it were made by solid-phase phosphoramidite chemistry, and their ligation into full-length cRNA (three separate phosphodiester exchange reactions, with release of the 5' guanosine of each 10-nt fragment) was measured over time for three ribozymes: the Tetrahymena intron, the sunY-td single-chain ribozyme, and the sunY-td multisubunit ribozyme.
What came out: The yield of full-length cRNA was markedly greater for the sunY-catalyzed reactions than for the Tetrahymena reactions; the efficiency of Tetrahymena-catalyzed ligations fell off precipitously as substrate chain length increased. The 27-nt product varied between experiments but was always lower in abundance than the other products.
Evidence: "The yield of full-length cRNA was markedly greater for the sunY catalyzed reactions than for reactions catalyzed by the Tetrahymena ribozyme."
H2.5 — The ribozyme templates synthesis of a strand complementary to one of its own subunits (fragment C)
What was done: Five oligomer substrates complementary to the 43-nt fragment C were synthesized, and a time course (up to 16 hours) was generated for their ligation catalyzed by the multisubunit enzyme, including a lane with no added template beyond fragment C already in the enzyme complex (lane 7), and a lane with fragment C at 10 µM.
What came out: The multisubunit ribozyme catalyzed synthesis of a strand complementary to its own fragment C subunit; some ligation occurred even in the reaction with no additional template (lane 7), implying that a fraction of complexes were unfolded enough to allow annealing of the anti-sense oligomer substrates, yet the complex was stable enough that its function was not completely inhibited by their presence. Full-length and intermediate product identities were confirmed by base-specific RNA sequencing.
Evidence: "One of the multisubunit sunY derivatives catalyzed the synthesis of a strand of RNA complementary to one of its own subunits."
H2.6 — Inefficiency / anomaly of the multisubunit system
What was done: Across the ligation assays the multisubunit ribozyme's activity, assembly, and product yields were compared with the single-chain ribozyme and Tetrahymena enzyme.
What came out: The multisubunit enzyme is less efficient than the single-chain enzyme and is only partially assembled into active complexes; nonetheless, the sunY-derived multisubunit ribozyme is the smallest and one of the most efficient RNA catalysts then identified for cRNA synthesis, and because its individual subunits are short, each could itself serve as a template for complementary-strand synthesis.
Evidence: "one explanation for these differences is that the multisubunit enzyme is only partially assembled into active complexes."
What it made possible next
This engineered, dissectible group I intron system crystallized the central tension the field would next have to resolve: the same RNA must simultaneously fold into a precise three-dimensional catalyst and be copied as an unstructured template. Doudna's demonstration that a ~180-nt sunY ribozyme could be split into three short subunits that reassemble, and that one subunit could be copied by the assembled complex, made the folded-versus-copied trade-off concrete and motivated the question of exactly how such a small ribozyme achieves its tertiary architecture. That question — what holds a group I intron's domains together in three dimensions — is precisely what the 1996 Tetrahymena P4-P6 domain crystal structure (H3) set out to answer, moving the program from functional dissection toward atomic-resolution ribozyme architecture.
Protocol extraction — [P] values printed in the paper
- [P] Standard ligation reaction (ref. 20): tris pH 7.4, 30 mM; NH4Cl, 10 mM; MgCl2, 50 mM; KCl, 0.4 M; intron (enzyme) RNA, 1 µM; substrate RNA, 20 µM of each.
- [P] Reaction procedure: enzyme RNA incubated in the buffer mix at the reaction temperature for 20 min; reactions initiated by addition of substrate RNA; incubation at 42°C for indicated times.
- [P] Stop/loading mix: equal volume of formamide (90%), tris (10 mM), EDTA (1 mM), xylene cyanol (0.1%), and bromophenol blue (0.1%); products analyzed on denaturing polyacrylamide gels and quantitated on a Betagen beta scanner.
- [P] Fig. 3 (multiple-oligonucleotide cRNA assembly) conditions: as ref. 20 except ethanol (10%) included, MgCl2 concentration 20 mM, all RNAs at 10 µM; RNA enzymes preincubated 20 min in buffer at 37°C before adding substrates; incubated at 37°C; analyzed on 20% acrylamide–70% formamide gels; far-left lane contains a 31-nt marker; time points lanes 1–5 = 0, 2, 4, 8 h (lane 5 = 8 h, no enzyme RNA).
- [P] Fig. 4 (copying fragment C) conditions: as refs. 20, 21 except 20 µM fragment C used; all other RNAs at 10 µM; incubation times lane 1 = 0, lane 2 = 2 h, lane 3 = 4 h, lane 4 = 8 h, lane 5 = 16 h, lane 6 = 16 h (no preincubation), lane 7 = 16 h with fragment C at 10 µM.
- [P] Fig. 2 ligation time course: gel is 20% acrylamide–7 M urea; lanes 0, 15, 30, 45, 60 min; far-right lanes are 60-min reactions with fragment A, B, or C omitted.
- [P] RNA/fragment lengths: JD929 = 180 nt; fragment A = 59 nt; fragment B = 75 nt; fragment C = 43 nt; sunY derivatives as small as 164 nt retain activity; sunY intron ~200 nt; Tetrahymena intron = 413 nt; cRNA-assembly template = 36 nt; oligomers = one 9-nt and three 10-nt.
- [P] RNA preparation: intron/substrate RNA prepared by T7 RNA polymerase transcription of plasmids (or synthetic DNA oligonucleotide templates) digested with BamHI and purified by denaturing polyacrylamide gel electrophoresis; smaller substrate ("primer") 5'-end labeled with T4 polynucleotide kinase and γ-[32P]ATP.
- [P] Chemically synthesized oligomers (Fig. 3): solid-phase RNA phosphoramidite chemistry (phosphoramidites from Milligen-Biosearch); purified by anion-exchange HPLC on a Dionex NA-100 column with acetonitrile (10%) and ammonium acetate, pH 5.6 (0.01 to 2 M) gradient.
Gaps requiring [S]/[M]
- [S] Exact sequences of the three trans-acting fragments A, B, and C are shown only as secondary-structure diagrams (Fig. 3A / Fig. 1A), not as clean linear strings in text — precise nucleotide sequences require reading the figure or supplementary/methods sources.
- [S] Precise sequence of the 36-nt template and the four annealing oligomers (Fig. 3A) and the five fragment-C-complementary oligomers (Fig. 4A) are given only in figure diagrams; a machine-usable list is not in the running text.
- [S] Quantitative ligation yields / rate constants (fold differences beyond "twofold" and "three- to fourfold") and the beta-scanner quantitation values are not tabulated numerically.
- [S] Detailed T7 transcription reaction composition (NTP concentrations, template amounts, T7 units) and T4 PNK labeling reaction composition are not specified in this reprint (cited to refs. 11, 20).
- [M] The mechanistic basis for "partial assembly" of the multisubunit complex and the exact fraction of active complexes is inferred, not measured.
H3 — Crystal structure of the P4-P6 domain of the Tetrahymena group I intron: principles of RNA packing
Cate JH, Gooding AR, Podell E, Zhou K, Golden BL, Kundrot CE, Cech TR, Doudna JA. Crystal structure of a group I ribozyme domain: principles of RNA packing. Science 1996;273(5282):1678-1685. PMID 8781224. (Doudna last author / her lab, with Cech.)
The question this paper set out to answer — RNA can both encode genetic information and catalyze reactions, but how a single-stranded polynucleotide folds into a compact, solvent-inaccessible three-dimensional structure capable of catalysis was essentially unknown, since prior atomic-resolution RNA structures were limited to small molecules (tRNAs ~76 nt, hammerhead ribozymes ~50 nt). This work sought to determine, at atomic resolution, the crystal structure of a large (160-nt) independently folding RNA — the P4-P6 domain of the Tetrahymena thermophila group I self-splicing intron — to reveal the long-range tertiary interactions, metal-binding sites, and packing principles that allow large RNAs (group I/II introns, RNase P, ribosomal and spliceosomal RNAs) to assemble into complex globular folds.
Experiments and results
H3.1 — Synthesis, crystallization, and space group of the 160-nt P4-P6 RNA. What was done: The 160-nt P4-P6 domain was transcribed in vitro with T7 RNA polymerase, gel purified, and crystallized by vapor diffusion. What came out: Crystals belong to space group P2₁2₁2₁ with two P4-P6 molecules per asymmetric unit and diffract anisotropically (2.8 to 2.5 Å). Evidence: "We present here the x-ray crystal structure of the 160-nt P4-P6 domain from the Tetrahymena group I intron at 2.8 Å resolution."
H3.2 — Crystallization conditions and role of divalent metals. What was done: Crystals were grown from defined buffer with Mg²⁺, spermine, and cobalt hexammine, using MPD as precipitant. What came out: Divalent (Mg²⁺) ions are required for folding, consistent with a direct metal role in structure formation. Evidence: "Crystals of the RNA were grown in 60 mM potassium cacodylate (pH 6), 30 mM magnesium chloride, 0.3 mM spermine, and 0.2 to 1.0 mM cobalt hexammine chloride by vapor diffusion with methylpentanediol as a precipitant."
H3.3 — Structure solution by MAD/SIR phasing with heavy-atom derivatives. What was done: Phases were obtained by multiwavelength anomalous diffraction (MAD) and single isomorphous replacement (SIR) using an osmium derivative, combined with a cobalt hexammine derivative, followed by density modification. What came out: A high-quality experimental map allowed correct nucleotide positioning; the register was confirmed by 5-iodouracil difference Fouriers. Evidence: "The crystal structure of the P4-P6 domain RNA was solved by multiwavelength anomalous diffraction (MAD) and single isomorphous replacement (SIR) with the use of an osmium derivative."
H3.4 — Refined model and overall architecture. What was done: The model was built and refined against the Cohex 1 data set with X-PLOR. What came out: The current model contains 154 nt in each molecule plus 28 metals and six waters; the domain comprises two helical regions packing side by side with dimensions ~110 × 50 × 25 Å. Evidence: "The current model consists of 154 nt in molecule A, 154 nt in molecule B, and a total of 28 metals and six waters."
H3.5 — Two coaxially stacked helical stacks and the ~150° bend. What was done: The fold of the paired regions was analyzed relative to the predicted secondary structure. What came out: Helices P6b, P6a, P6, P4, and P5 form a straight column on one side; P5b and P5a stack on the other; a sharp bend juxtaposes the two halves so the P5abc extension packs against the conserved core. Evidence: "A bend of ~150° between helices P5 and P5a allows the P5abc extension to interact with one helical face of the conserved core region."
H3.6 — The A-rich bulge, its corkscrew backbone, and two-Mg²⁺ metal core. What was done: The A-rich bulge in P5a and its interactions were resolved in the electron density. What came out: The bulge backbone makes a corkscrew turn with flipped-out bases contacting P4 and the three-helix junction; closely spaced phosphate oxygens directly coordinate two Mg²⁺ ions forming an approximate helical-symmetry axis from A183 to A187. Evidence: "These phosphate oxygens directly coordinate two magnesium ions, clearly visible in the experimental electron density map."
H3.7 — Adenosine-rich corkscrew as a two-Mg²⁺ clamp into the P4 minor groove. What was done: The tertiary contact between the bulge and P4 was mapped, with A186 examined at the three-helix junction. What came out: The first two adenosines bind the P4 minor groove and the last two bind pockets at the junction; A186 nestles in a pocket formed by the C137·G181 pair and G164. Evidence: "a two-Mg²⁺–coordinated adenosine-rich corkscrew plugs into the minor groove of a helix."
H3.8 — GAAA tetraloop–11-nt tetraloop-receptor interaction. What was done: The docking of the GAAA (L5b) tetraloop into the conserved 11-nt receptor (J6a/6b) was resolved. What came out: The loop docks in the minor groove of P5b/P6a at ~30°, with each adenine making base-specific hydrogen bonds, explaining the sequence specificity of the interaction. Evidence: "a GAAA hairpin loop binds to a conserved 11-nucleotide internal loop."
H3.9 — The adenosine platform motif. What was done: The conformation of adjacent adenosines in the receptor internal loop was analyzed. What came out: Two adjacent adenosines stack side by side across the helix, forming an "adenosine platform" that kinks the backbone and opens the receptor minor groove for loop docking. Evidence: "adjacent adenosines in the receptor internal loop that stack across the helix, forming an adenosine platform motif."
H3.10 — Ribose zippers and metal-mediated close packing of backbones. What was done: Interhelical backbone contacts in the A-rich bulge and tetraloop regions were examined. What came out: Pairs of riboses hydrogen bond via shared 2′-hydroxyl/O2 contacts ("ribose zippers"), and hydrated Mg²⁺ bridges phosphates 7–8 Å apart, enabling snug side-by-side helical packing. Evidence: "Pairs of riboses interact by hydrogen bonding, forming 'ribose zippers' in the A-rich bulge and the GAAA tetraloop long-range contacts."
H3.11 — Overall folding, backbone accessibility, and agreement with prior biochemistry. What was done: Crystallographic solvent accessibility of C4′ atoms was compared to Fe(II)-EDTA cleavage data. What came out: Protected C4′ atoms lie on the buried interior between the two halves; the crystallographic and biochemical accessibilities correlate well, validating the fold. Evidence: "There is good correlation between the backbone accessibility determined biochemically and crystallographically."
H3.12 — Comparison to prior RNA structures and implications. What was done: The structure was placed in the context of earlier tRNA and hammerhead ribozyme structures and modeling. What came out: The domain assembles a complex globular fold from only four similar building blocks, providing the first detailed view of the compactness of large-RNA folding relevant to ribozymes, the spliceosome, and the ribosome. Evidence: "The structure indicates the extent of RNA packing required for the function of large ribozymes, the spliceosome, and the ribosome."
What it made possible next
This structure established the RNA-crystallography pipeline that defined the Doudna lab's structure-first program: in-vitro T7 transcription of a discrete folded domain, heavy-atom phasing using hexammine derivatives (osmium and cobalt hexammine as isomorphous/anomalous scatterers), MAD/SIR + density modification, and refinement of a large folded RNA. It also produced a reusable vocabulary of RNA tertiary motifs — the GAAA tetraloop/11-nt receptor, the adenosine platform, the A-rich bulge/two-metal corkscrew, ribose zippers, and metal-mediated backbone packing — that recur throughout later RNA structures. It set up the lab's next targets, notably the 1998 hepatitis delta virus (HDV) ribozyme crystal structure (H4), by proving that discrete catalytic RNA domains could be crystallized and solved at atomic resolution and that metal-phosphate coordination and 2′-hydroxyl networks are the organizing principles of RNA folding.
Protocol extraction — [P] values printed in the paper
- [P] Molecule/domain length: 160-nt P4-P6 domain of the Tetrahymena thermophila group I intron (two guanosines added at 5′ end for T7 transcription).
- [P] Reported resolution: 2.8 Å (crystals diffract anisotropically 2.8 to 2.5 Å); refinement resolution 8.0–2.5 Å.
- [P] Space group: P2₁2₁2₁, two P4-P6 molecules in the asymmetric unit.
- [P] Crystallization: 60 mM potassium cacodylate (pH 6), 30 mM MgCl₂, 0.3 mM spermine, 0.2 to 1.0 mM cobalt hexammine chloride, MPD (methylpentanediol) as precipitant, vapor diffusion.
- [P] Cryo/transfer solution (Table 1): 25% MPD, 100 mM potassium cacodylate (pH 6.0), 50 mM MgCl₂, 0.5 mM spermine, 10% isopropanol, 0.035 to 0.07 mM cobalt hexammine chloride; flash-frozen in liquid propane cooled with liquid nitrogen.
- [P] Heavy atoms / phasing: osmium hexammine triflate (osmium derivative, substituted for cobalt hexammine), cobalt hexammine derivative; MAD + SIR; osmium data collected at peak (λ1) and first inflection (λ2) of the osmium L-III absorption edge; 5-iodouracil derivatives (nucleotides 241, 253, 258, 259) for register confirmation.
- [P] Data (Table 1): Oshex λ1 20.0–2.8 Å, 63,718 unique, R_sym 4.8; Oshex λ2 20.0–2.9 Å, R_sym 4.3; Cohex1 18.0–2.5 Å, 42,836 unique, R_sym 4.5; Cohex2 20.0–2.8 Å, 31,828 unique, R_sym 6.9.
- [P] Model: 154 nt in molecule A, 154 nt in molecule B, 28 metals, six waters; total 6,824 atoms (Table 1: number of atoms N = 6,824).
- [P] Refinement statistics (Table 1, X-PLOR 3.8, Cohex1): working set N = 34,551; test set N = 1,850; R-factor = 0.242; R-free = 0.285; rms bond = 0.010 Å; rms angle = 1.27°; mean FOM overall 0.71 (after final DM).
- [P] Metal/geometry values: ~28 metals modeled; phosphate oxygen–metal ion distances ~2.2 Å on average; two Mg²⁺ coordinate the A-rich bulge corkscrew (axis A183→A187); interhelical phosphate oxygens 7–8 Å apart bridged by hydrated Mg²⁺; closest bulge phosphate oxygens ~3 Å apart; overall molecular dimensions ~110 × 50 × 25 Å; interhelical bend ~150°; tetraloop docks at ~30°.
- [P] Software: DENZO/SCALEPACK (processing), MLPHARE (MIR/anomalous phasing), DM (density modification), O (model building), X-PLOR 3.8 (refinement), RIBBONS/MidasPlus (figures).
Gaps requiring [S]/[M]
- [S] No PDB accession code is printed in the paper (structure deposition; a full account of structure determination and refinement was noted as forthcoming — refs 34 "Cate & Doudna, Structure, in press" and 52 "in preparation").
- [S] No overall data-collection temperature stated in main text beyond Table 1 note (−160°C for Patterson data); per-derivative completeness/redundancy for all sets only partially tabulated.
- [S] Precise number and identity/coordination of all 28 metals (which are Mg²⁺ vs monovalent/cobalt hexammine) not individually enumerated in this paper ("details of other metal binding interactions... are not yet completely known").
- [S] The monovalent-ion contribution to the A-rich bulge metal core is not quantified here (only Mg²⁺ are explicitly resolved and described).
- [S] Exact overall figure-of-merit and phasing power per shell for every derivative not fully given; some Table 1 cells are blank by design.
- [M] Quantitative free-energy/thermodynamic contribution of each tertiary motif (tetraloop-receptor, A-rich bulge) to domain stability is not measured in this structural paper.
H4 — Crystal structure of a hepatitis delta virus ribozyme: a nested double pseudoknot builds a protein-enzyme-like active-site cleft
Ferré-D'Amaré AR, Zhou K, Doudna JA. Crystal structure of a hepatitis delta virus ribozyme. Nature 1998;395(6702):567-574. PMID 9783582. (Doudna last author / her lab.)
The question this paper set out to answer — The HDV ribozyme is the fastest known natural self-cleaving RNA and the only catalytic RNA required for a human pathogen's viability, yet it stood out from all other ribozymes: it needs no specific metal ion, stays active in 5 M urea or 18 M formamide, and requires only a single nucleotide 5' of the scissile bond. Secondary-structure models (Perrotta and Been) proposed a pseudoknot, and biochemistry had mapped conserved residues, but there was no three-dimensional structure to explain how this compact RNA achieves protein-like catalytic rates or how it activates its nucleophile without an obligate metal. This paper set out to crystallize and solve the atomic structure of a self-cleaved (product-form) genomic HDV ribozyme, to reveal its fold and read out the architecture of its active site directly.
Experiments and results
H4.1 Engineering a U1A-protein binding site to crystallize the ribozyme (crystallization chaperone)
- What was done: Because naked HDV ribozyme RNA gave poorly ordered crystals, the solvent-exposed P4 stem — dispensable for catalysis — was engineered to carry a high-affinity binding site for the small, basic RNA-binding domain of the U1A spliceosomal protein (U1A-RBD), and the ribozyme was co-crystallized as an RNA–protein complex. Self-cleavage assays (±U1A-RBD, ±BSA) confirmed activity was unaffected.
- What came out: The alteration of P4 and the binding of U1A-RBD did not impair self-cleavage; the co-crystal of the product RNA with a selenomethionyl U1A-RBD diffracted X-rays to 2.3 Å. The authors present this as a general RNA crystallization technique.
- Evidence: "by engineering the RNA to bind a small, basic protein without affecting ribozyme activity."
- Evidence: "The alteration of P4 did not adversely affect the self-cleavage activity of the HDV ribozyme, nor did the binding of U1A-RBD"
H4.2 Crystallizing the self-cleaved (product-form) ribozyme and solving it by MAD
- What was done: A 72-nucleotide self-cleaved (product) genomic HDV ribozyme was transcribed (plasmid pDU9, T7 polymerase, run-off then self-cleavage), complexed 1:1 with selenomethionyl U1A-RBD, and the structure solved by four-wavelength multiwavelength anomalous diffraction (MAD) from selenium sites, then refined at 2.3 Å. Product (rather than precursor) RNA was used because much in-vitro-transcribed ribozyme is misfolded, and product/precursor structures are known to be similar.
- What came out: An interpretable experimental map (MAD phases from four Se sites, phase extension to 2.7 Å) into which nearly all RNA and protein residues were built; the final model refined to R-free 28.4% (Rwork 28.0%) at 2.3 Å in space group R32.
- Evidence: "We obtained crystals of a 72-nucleotide, self-cleaved form of the genomic HDV ribozyme that diffract X-rays to 2.3 A˚ resolution"
H4.3 The overall fold: five helices in a nested double pseudoknot forming two coaxial stacks
- What was done: Interpreted the global architecture, comparing the crystallographic helix arrangement to the Perrotta–Been secondary-structure model and to hydroxyl-radical footprinting.
- What came out: Beyond the four known paired regions P1–P4, the structure revealed a new two-base-pair helix P1.1 formed by nucleotides of J1/4 and loop L3 (thought to be single-stranded), which introduces a second pseudoknot. The five segments form two parallel coaxial stacks joined side-by-side by five strand-crossovers; this makes HDV the fourth RNA fold ever solved (after tRNA, hammerhead, and the P4-P6 domain).
- What came out: Disrupting either pseudoknot cripples activity — deleting or disrupting the four conserved P1.1 nucleotides reduces activity three to five orders of magnitude.
- Evidence: "the compact catalytic core comprises five helical segments connected as an intricate nested double pseudoknot."
- Evidence: "P1, P1.1 and P4 stack nearly co-axially, whereas P2 and P3 form a second co-axial stack"
H4.4 A deep, solvent-inaccessible active-site cleft around the 5'-hydroxyl leaving group
- What was done: Located the active site via the 5'-hydroxyl leaving group of the self-scission reaction, mapped the cleft-lining elements (P1 substrate helix, J4/2, the P3–L3 niche, the P3–P1 strand-crossover arch, P1.1 pedestal), and cross-checked against hydroxyl-radical protection and crosslinking data.
- What came out: The convoluted fold buries the 5'-hydroxyl deep in a cleft whose walls are formed by biochemically important backbone and base groups — an arrangement the authors liken to protein enzymes. The substrate-bearing P1 helix is positioned by helical crossovers at its top and by stacking of the G1·U37 wobble on P1.1 at its bottom; the wobble's backbone distortion pushes the 5'-hydroxyl deeper into the core.
- Evidence: "The compact, convoluted fold buries the 59-hydroxyl leaving group of the self-cleavage reaction deep within an active-site cleft"
- Evidence: "surrounded by biochemically important backbone and base functional groups in a manner reminiscent of protein enzymes."
H4.5 C75 positioned at the active site and proposed as the general base
- What was done: Traced the J4/2 "trefoil" turn that arranges G74, C75, G76; located C75 relative to the 5'-hydroxyl; integrated mutagenesis (C75 cannot be substituted), crosslinking (deoxythiouridine at −2 crosslinks to C75), and the RNA's lack of pH dependence / metal specificity.
- What came out: C75 projects deep into the core, with its Watson–Crick face within hydrogen-bonding distance of the 5'-hydroxyl; a hydrogen-bond network and negative electrostatic environment could perturb its pKa. The authors propose C75 acts as the general base activating the −1 2'-hydroxyl nucleophile — catalysis read directly from structure.
- Evidence: "C75 projects deep into the core of the ribozyme"
- Evidence: "we propose that C75 acts as the general base that activates the 29-hydroxyl group of nucleotide −1 for nucleophilic attack."
- Evidence: "C75 cannot be mutated to any other nucleotide without completely abolishing ribozyme activity"
H4.6 No specific metal ion in catalysis — the ion as a charge sink
- What was done: Searched the maps for tightly bound metals; considered phosphorothioate-substitution data (pro-Rp at cleavage site and at C22 sensitive, not rescued by Mn2+) and the ribozyme's known lack of specific-metal and pH requirements.
- What came out: No tightly bound metal ions were found; the structure is stabilized by base-pairing, stacking and non-canonical base/backbone contacts. The model contains only two sulphate and three magnesium ions. Metal is inferred not to activate the nucleophile directly but to act as a charge sink.
- Evidence: "No tightly bound metal ions were found in the HDV ribozyme"
- Evidence: "the ion does not participate directly in activating the nucleophile, but instead acts as a charge sink."
What it made possible next
This structure carried the Doudna lab's structure-first RNA program forward from a large, loosely packed domain (the 1996 P4-P6 group I intron, H3) to a compact catalytic active site, showing that catalytic mechanism could be read directly off a folded RNA — here nominating C75 as a general acid/base years before pH-rate and rescue experiments confirmed nucleobase catalysis. Methodologically it extended and generalized the lab's toolkit: heavy-atom/hexammine and MAD phasing (cobalt(III)- and osmium(III)-hexammine derivatives, selenomethionine MAD), the U1A-RBD protein as a reusable crystallization chaperone that adds ordered surface to "naked" RNA, and Kaihong Zhou at the bench across both structures. That protein-as-chaperone trick and the discipline of solving discrete, functionally intact RNA/RNP assemblies fed directly into the lab's move to larger ribonucleoprotein machines — most immediately the 2001 HCV IRES–40S ribosome work (H5) — where reading how an RNA docks and organizes a protein partner was the central problem.
Protocol extraction — [P] values printed in the paper
- [P] RNA construct = 72-nucleotide self-cleaved (product-form) genomic HDV ribozyme; homogeneous 5', heterogeneous 3' termini; from plasmid pDU9 (pUC19 derivative), T7 run-off transcription then self-cleavage
- [P] Resolution = refined at 2.3 Å (final); MAD phasing / initial map at 2.7–2.9 Å with phase extension to 2.7 Å
- [P] Space group = R32; cell a ≈ 108.7–109.35 Å, c ≈ 190.4–190.9 Å (crystals I–III)
- [P] Crystallization chaperone = RNA-binding domain of U1A spliceosomal protein (U1A-RBD); site engineered into the P4 stem; a 98-residue double-mutant U1A-RBD (Y31H, Q36R), selenomethionyl form; A1-98 construct; complex 1:1 with RNA at 0.5 mM in 1.25 mM MgCl2
- [P] Phasing = four-wavelength MAD from selenomethionine (four Se sites; three well-ordered, Met97 site disordered); programs SOLVE/SHARP/SOLOMON, model building in O, refinement in CNS
- [P] Heavy-atom / hexammine additives = 0.2 mM cobalt(III) hexammine chloride (CoHex; crystals I, III); 0.4 mM osmium(III) hexammine chloride triflate (OsHex; crystal II, not specifically bound); MgCl2
- [P] Metal ions in final model = two sulphate and three magnesium ions (no tightly bound metals; nineteen waters)
- [P] Catalytic residue = C75 (Watson–Crick face within hydrogen-bonding distance of the 5'-hydroxyl; proposed general base); cleavage-site pair = G1·U37 wobble
- [P] Refinement statistics = R-free 28.4% (Rwork 28.0%) at 2.3 Å; 17,042 work / 1,861 free reflections; rms bond 0.009 Å, rms angle 1.43°; 2,366 non-hydrogen atoms
- [P] Crystallization drop = 12.5–14% PEG-MME 2000, 100 mM Tris-HCl pH 7.0, 200–250 mM Li2SO4, 4 mM spermine-HCl; trigonal trapezohedra at 25 °C over several weeks, to 0.8 × 0.6 × 0.6 mm
- [P] Data collection = NSLS beamline X4A (MAD, crystals I/II, R-axis IV); CHESS beamline F-1 (crystal III, Quantum-4 CCD); DENZO/SCALEPACK, CCP4
- [P] PDB accession = 1drz
Gaps requiring [S]/[M]
- [S] Full account of experimental details, exhaustive refinement, and crystal-contact analysis explicitly deferred to a separate manuscript ("A.R.F. and J.A.D., manuscript in preparation")
- [S] Direct proof of the catalytic mechanism — C75-as-general-base is proposed from position + biochemistry, not demonstrated here (no measured C75 pKa, no pH-rate/chemical-rescue experiment in this paper)
- [S] Precursor / transition-state structure not determined — only the product (self-cleaved) form solved; precursor geometry inferred from similarity arguments
- [S] Role of metal ions in catalysis unresolved ("do not reveal how metal ions might participate"); identity/occupancy of physiological ions beyond the three modeled Mg2+ not established
- [S] Positions/conformations of the mobile, disordered nucleotides (U23, C26, U27) and the exact trajectory of the −1/upstream substrate strand in the precursor
H5 — An RNA that reshapes the ribosome: cryo-EM of the HCV IRES bound to the 40S subunit
Spahn CMT, Kieft JS, Grassucci RA, Penczek PA, Zhou K, Doudna JA, Frank J. "Hepatitis C virus IRES RNA-induced changes in the conformation of the 40S ribosomal subunit." Science 2001;291(5510):1959-62. PMID 11239155. (Doudna = sixth/middle author; the HCV IRES RNA is from the Doudna lab, in collaboration with Joachim Frank's cryo-EM group at the Wadsworth Center.)
The question this paper set out to answer — Cap-independent translation initiation by viral RNAs relies on an internal ribosome entry site (IRES), a structured RNA element in the 5′ UTR that recruits the 40S ribosomal subunit and eIF3 without a 5′ cap or most canonical initiation factors. Prior models cast the HCV IRES as a passive scaffold onto which the 40S subunit docks. The authors asked, at the structural level, how the HCV IRES RNA engages the 40S subunit: where it binds, which of its domains contact which parts of the subunit, and — critically — whether the IRES merely provides a landing platform or actively alters the conformation of the ribosomal machine to position the coding RNA for initiation.
Experiments and results
H5.1 — Control cryo-EM reconstruction of the vacant rabbit 40S subunit
What was done: Single-particle cryo-EM reconstruction of the vacant 40S ribosomal subunit from rabbit reticulocytes, from zero-tilt images, as a reference/control map.
What came out: A ~20 Å map showing the classical body/head/platform division; the first eukaryotic ribosomal-subunit reconstruction from zero-tilt images, with resolution improved over an earlier random-conical-tilt reconstruction, and comparing well with the 40S portion of a yeast 80S reconstruction.
Evidence: "It represents the first reconstruction of a eukaryotic ribosomal subunit from zero-tilt images and has a greatly improved resolution compared with a previous random-conical tilt reconstruction"
H5.2 — Cryo-EM of the full-length HCV IRES RNA bound to the 40S subunit
What was done: Reconstructed the complex of full-length HCV IRES RNA (nt 40–372, genotype 1b, including 30 nt of coding RNA) bound to the rabbit 40S subunit; computed a difference map versus the vacant 40S.
What came out: IRES density appears as an extended structure on the solvent side of the subunit consistent with A-form helices; binding induces a pronounced conformational change (reorientation of the head/beak relative to the body, reshaping of the platform), and closes the mRNA binding cleft.
Evidence: "IRES binding induces a pronounced conformational change in the 40S subunit and closes the mRNA binding cleft, suggesting a mechanism for IRES-mediated positioning of mRNA in the ribosomal decoding center."
H5.3 — Domain II deletion mutant to localize the source of the conformational change
What was done: Reconstructed the complex of a truncated IRES (nt 119–372, domain II deleted; binds 40S with near-wild-type affinity) with the 40S subunit and compared with the full-length complex.
What came out: Density missing in the truncated complex unambiguously identifies domain II (which contacts the head/edge of the platform near the E site); the conformational change seen with full-length IRES is absent with the ΔdII mutant, so the change is a consequence of domain II's interaction near the E site.
Evidence: "The conformational change induced by the binding of full-length HCV IRES RNA is not observed when the domain II deletion mutant is bound to the 40S subunit"
H5.4 — Assignment of IRES structural domains by correlation with chemical/enzymatic probing
What was done: Correlated previously mapped tertiary-structure probing/protection data with the cryo-EM density to assign IRES domains (IIIa/b/c, IIId/e/f, junctions) to features of the map.
What came out: The globular density on the back of the platform contains domain IIId/e/f; extended density corresponds to IIIb/IIIa/IIIc, with IIIb projecting away from the surface exactly where eIF3 is known to bind — so even without eIF3, domain IIIb is pre-positioned to recruit it.
Evidence: "even in the absence of eIF3, the IRES RNA domain IIIb is properly positioned to bind this factor." (OCR-uncertain)
What it made possible next
This is the lab's first structural demonstration of an RNA element actively reprogramming the conformation of a large protein–RNA machine — "the first reported example of a structured viral RNA that binds to, and actively manipulates, the structure of a cellular machine." The IRES is not a passive scaffold; a defined RNA domain drives a mechanical change in the ribosome that positions coding RNA in the decoding center. This reframes an RNA as a conformational effector acting on a large ribonucleoprotein complex — a conceptual bridge from the Doudna lab's RNA-folding/tertiary-structure work toward machines in which an RNA guides and reconfigures a protein partner: forward to Dicer/RNAi substrate-processing structural work (2006) and, conceptually, to the CRISPR RNA-guided effector complexes in which a guide RNA directs and conformationally licenses a large protein for its target.
Protocol extraction — [P] values printed in the paper
- [P] Cryo-EM overall resolution: ~20 Å (stated); per-map FSC 0.5 cutoff: 22.7 Å (vacant 40S), 19.8 Å (IRES–40S), 21.9 Å (IRESΔdII–40S); under the 3σ criterion: 15.9 Å, 14.5 Å, 15.4 Å respectively.
- [P] Particles: 18,801 (vacant 40S), 20,939 (IRES–40S), 13,613 (IRESΔdII–40S), from 69 / 74 / 61 selected micrographs respectively.
- [P] Microscope: Philips F20 (FEI/Philips); magnification 50,000 ±2%; low-dose; defocus range 1.7–5.7 μm. (Accelerating voltage/kV not printed.)
- [P] Scanning: Highscan drum scanner at 1069 dpi → pixel size 4.78 Å on the object scale; data processed in SPIDER; CTF-corrected 3D reconstruction, 16–23 defocus groups per data set.
- [P] IRES constructs: full-length = nt 40–372 of HCV IRES genotype 1b (includes 30 nt coding RNA); deletion mutant = nt 119–372 (domain II deleted).
- [P] Sample prep: 40S from rabbit reticulocytes; folding/binding buffer 20 mM tris-HCl, 100 mM potassium acetate, 200 mM KCl, 2.5 mM MgCl2, 1 mM DTT; RNA annealed 75°C 1 min; complex incubated 37°C 15 min; final complex ~500 nM; cryo grids applied at ~50 nM, shock-frozen in liquid ethane, stored at −80°C.
Gaps requiring [S]/[M]
- [S] Accelerating voltage (kV) of the Philips F20 microscope — not printed in the extracted text.
- [S] Electron dose (e⁻/Ų) under "low-dose conditions" — not given numerically.
- [S] Number of micrographs recorded before selection (only the selected counts 69/74/61 are stated).
- [S] Temperature of imaging/vitrification stage and grid type — not specified in extracted text.
- [S] Exact number of defocus groups per individual data set (given only as a range, 16–23).
- [M] Quantitative magnitude of the head-rotation/conformational change (angle or displacement) — described qualitatively only.
H6 — Dicer is a molecular ruler: the crystal structure that showed how dsRNA is measured and cut into siRNAs
MacRae IJ, Zhou K, Li F, Repic A, Brooks AN, Cande WZ, Adams PD, Doudna JA. "Structural basis for double-stranded RNA processing by Dicer." Science 2006;311(5758):195-8. PMID 16410517. (Doudna = last/corresponding author; her lab, UC Berkeley / HHMI / LBNL.)
The question this paper set out to answer — Dicer initiates RNA interference by chopping long double-stranded RNA into small fragments (~21–25 nt siRNAs/miRNAs) of a strikingly uniform length, ideal to guide sequence-specific silencing. Several models had been proposed for how Dicer produces fragments of this specific size, but structural information was lacking. The authors set out to determine the crystal structure of an intact, fully active Dicer enzyme to reveal, at the molecular level, how the machine recognizes a dsRNA end and cleaves a defined distance from it — i.e., how product length is set.
Experiments and results
H6.1 — A minimal, highly active Dicer in Giardia
What was done — Identified a Giardia intestinalis open reading frame encoding the PAZ plus tandem RNase III domains characteristic of Dicer but lacking the N-terminal helicase, C-terminal dsRBD, and extended interdomain regions; expressed the recombinant protein and ran an in vitro dsRNA cleavage time course (Fig. 1).
What came out — The recombinant enzyme has robust, Mg2+-dependent dicing activity, producing discrete ~25–27 nt fragments in intervals of ~25 nt, indicating processive cleavage from the helical end similar to human Dicer.
Evidence: "A recombinant form of this protein possesses robust dicing activity in vitro" and "The RNA fragments produced by Giardia Dicer are 25 to 27 nucleotides long."
H6.2 — Crystal structure of full-length Dicer ("hatchet" architecture)
What was done — Determined the crystal structure of full-length Giardia Dicer at 3.3 Å resolution (Fig. 2A; table S1).
What came out — An elongated, hatchet-shaped molecule: the two RNase III domains form the blade, the PAZ domain forms the base of the handle, and a long "connector" a helix runs through the handle linking PAZ to RNase IIIa; a flat surface extends along one face.
Evidence: "We determined the crystal structure of the full-length Giardia Dicer at 3.3 Å resolution" and it "takes on a shape resembling a hatchet."
H6.3 — Two-metal-ion RNase III active sites forming an intramolecular dimer
What was done — The two RNase III domains were examined; crystals were derivatized with ErCl3 (and grown with MnCl2), and anomalous difference electron density maps were inspected at the catalytic sites (Fig. 2B).
What came out — The two RNase III domains form an internal heterodimer; each active site contains a pair of metals (M1 between four conserved acidic residues, M2 adjacent), with Er3+–Er3+ distances of ~4.2 Å (IIIa) and ~5.5 Å (IIIb), supporting a conserved two–metal-ion mechanism of catalysis.
Evidence: "the anomalous difference electron density map ... revealed a pair of Er3+ cations in the active site of each RNase III domain" (OCR-uncertain).
H6.4 — PAZ domain: a conserved 3′-overhang-binding pocket plus a Dicer-specific loop
What was done — Superposed the Giardia Dicer PAZ domain on human Argonaute1 PAZ and compared electrostatic surfaces (Fig. 3).
What came out — The domains share the same OB-like fold and 3′ two-nucleotide RNA-binding pocket, but Dicer carries a large conserved extended loop (absent in Argonaute) rich in basic residues that reshapes the surface around the overhang pocket, likely affecting RNA recognition/hand-off.
Evidence: "the two domains share the same overall fold and 3′ two-nucleotide RNA binding pocket."
H6.5 — The molecular-ruler model of siRNA length
What was done — Measured the distance from the RNase IIIa active site to the PAZ 3′-overhang pocket and built a model docking ideal A-form dsRNA using the metal-ion pairs to anchor the scissile phosphates (Fig. 4).
What came out — The PAZ-to-RNase III distance is ~65 Å, matching 25 base pairs of dsRNA; exactly 25 nucleotides lie between the PAZ-bound 3′ end and the RNase IIIa scissile phosphate — so the connector-helix length sets product size.
Evidence: "Thus, Dicer is a molecular ruler that measures and cleaves ~25 nucleotides from the end of a dsRNA."
H6.6 — Giardia Dicer functions as an intact Dicer in vivo
What was done — Expressed Giardia Dicer in a Schizosaccharomyces pombe dcrD (Dicer-deletion) strain and assayed thiabendazole (TBZ) sensitivity and centromeric transcript silencing (Fig. 5).
What came out — Giardia Dicer partially rescued TBZ sensitivity and restored silencing of aberrantly transcribed centromeric regions, showing the minimal enzyme is sufficient to act as an intact Dicer and that Dicer architecture/mechanism is evolutionarily conserved.
Evidence: "these results demonstrate that Giardia Dicer is sufficient to function as an intact Dicer in vivo."
What it made possible next
This paper delivered the founding structural logic of small-RNA biogenesis: a defined dsRNA end is clamped by the PAZ domain, and a fixed physical distance (~65 Å, ~25 bp) to the RNase III active sites acts as a ruler that dictates guide-RNA length. That "measure-from-the-end, cut-at-a-set-distance" principle, and the demonstration that appended RNA-binding modules confer specificity on an otherwise nonspecific RNase, is exactly the conceptual and methodological toolkit Doudna's lab carried into bacterial CRISPR RNA processing — Csy4 cleaving pre-crRNA at a defined position (2010) and the Cascade surveillance complex measuring and presenting a guide RNA (2011). It cemented the lab's approach of solving the crystal structure of an intact RNA-processing machine to explain how guide RNAs of exact length are generated, the precondition for programmable, guide-directed nucleases.
Protocol extraction — [P] values printed in the paper
- [P] Organism / source: Dicer from Giardia intestinalis (a minimal Dicer: PAZ + tandem RNase IIIa/IIIb only).
- [P] Resolution: full-length structure solved at 3.3 Å (table S1).
- [P] Molecular ruler distance: ~65 Å between the RNase IIIa active site and the PAZ 3′-overhang-binding pocket, matching ~25 dsRNA base pairs / exactly 25 nucleotides.
- [P] Connector: a long "connector" a helix links PAZ to RNase IIIa; contains a conserved proline (~11 residues from the predicted N terminus of the helix) that kinks it toward RNase IIIa; helix length sets product size.
- [P] Metal ions: two metals per RNase III active site (two–metal-ion mechanism); Er3+ (M1 between four conserved acidic residues, M2 adjacent); Er3+–Er3+ distances ~4.2 Å (RNase IIIa) and ~5.5 Å (RNase IIIb); Mn2+ observed at M1 (and some M2) sites; catalysis is Mg2+-dependent (Mn2+, Ni2+, Co2+ also support activity).
- [P] Product length: Giardia Dicer produces ~25–27 nt fragments; discrete intermediates spaced ~25 nt.
- [P] Deposition: coordinates and structure factors in PDB, accession code 2FFL.
- [P] Data collection: Advanced Light Source beam lines 8.2.1 and 8.2.2 (Lawrence Berkeley National Lab).
Gaps requiring [S]/[M]
- [S] Space group and unit-cell parameters — not printed in the main text; reside in table S1 (Supporting Online Material).
- [S] Detailed crystallization conditions, phasing method, and refinement statistics (R/Rfree) — in SOM Materials and Methods, not in the extracted text.
- [S] Precise Er3+/Mn2+ derivatization protocol and metal-soak concentrations — main text notes only "high concentrations of MnCl2"; exact values not given.
- [S] Full protein construct boundaries, purification, and in vitro assay conditions (buffer, temperature, substrate) — in SOM.
- [S] S. pombe strain genotype, plasmid/expression details, and RT-PCR primers for the in vivo rescue — in SOM.
H7 — Csy4, the CRISPR endoribonuclease that reads sequence and structure to cut pre-crRNA
Haurwitz RE, Jinek M, Wiedenheft B, Zhou K, Doudna JA. Sequence- and structure-specific RNA processing by a CRISPR endonuclease. Science 2010 Sep 10;329(5997):1355-1358. doi:10.1126/science.1192272. PMID 20829488. (Doudna = last/corresponding author — her lab's first CRISPR paper.)
The question this paper set out to answer — CRISPR loci confer prokaryotic adaptive immunity by producing short CRISPR-derived RNAs (crRNAs) that home to invading viruses and plasmids; central to this is the post-transcriptional processing of long precursor transcripts (pre-crRNAs) into mature ~60-nucleotide crRNAs. In Pseudomonas aeruginosa UCBPP-PA14 (Pa14), a Yersinia-subtype CRISPR/Cas system carries six Cas genes, but which protein performs this maturation cleavage, where and how it cuts, and how it discriminates CRISPR transcripts from the bulk of cellular RNA were all unknown. This paper set out to identify the endoribonuclease responsible for crRNA biogenesis in this subtype and to determine — at atomic resolution — the recognition and catalytic mechanism that makes its processing both sequence- and structure-specific.
Experiments and results
H7.1 Identifying Csy4 as the crRNA-processing endoribonuclease among the six Cas proteins
- What was done: Each of the six Pa14 Cas proteins was recombinantly expressed and individually tested for endoribonuclease activity against an in vitro transcribed Pa14 pre-crRNA.
- What came out: Only Csy4 produced sequence-specific pre-crRNA processing, identifying it as the enzyme responsible for crRNA biogenesis in the Yersinia subtype.
- Evidence: "we concluded that Csy4 is the endoribonuclease responsible for crRNA biogenesis (Fig. 1B)."
H7.2 Metal-independent cleavage at the base of the repeat stem-loop yielding 60-nt crRNAs
- What was done: Csy4 was incubated with in vitro transcribed Pa14 pre-crRNA over a time course (30 s to 5 min) in buffer with no added metal, with 2.5 mM MgCl2, or with 2.5 mM EDTA; products were resolved on denaturing PAGE. Cleavage of S. thermophilus pre-crRNA (a distinct repeat stem-loop) was tested as a specificity control.
- What came out: Cleavage was rapid and metal-ion-independent, occurring within the repeat at the base of a predicted stem-loop and generating 60-nt crRNAs (32-nt spacer flanked by 8-nt 5' and 20-nt 3' repeat sequence). Csy4 did not cleave the S. thermophilus transcript.
- Evidence: "CRISPR transcript cleavage is a rapid, metal ion-independent reaction, as observed for crRNA processing within two other CRISPR/Cas subtypes."
H7.3 Specificity: no binding to bulk cellular RNA; co-purification of a ~19-nt protected fragment
- What was done: Csy4 (highly basic, pI 10.2) was expressed in E. coli alone or co-expressed with a synthetic Pa14 CRISPR transcript (eight repeats, seven spacers); the affinity-purified protein was assayed for associated nucleic acids by denaturing PAGE.
- What came out: Despite its basic charge, Csy4 did not associate with endogenous cellular nucleic acids; when co-expressed with the CRISPR transcript it co-purified with a protected ~19-nt crRNA fragment, underscoring the specificity of recognition.
- Evidence: "the protein co-purified with a protected ~19-nucleotide crRNA fragment (Fig. 1C)."
H7.4 A 16-nt minimal substrate and an absolute 2'-hydroxyl requirement for cleavage
- What was done: Csy4 binding and cleavage were assayed in vitro against RNA oligonucleotides spanning different regions of the 28-nt Pa14 repeat, and against a minimal substrate carrying a 2'-deoxyribonucleotide substitution immediately upstream of the cleavage site.
- What came out: A 16-nt fragment (the repeat-derived stem-loop plus one downstream nucleotide) was sufficient for cleavage. Cleavage required a 2'-hydroxyl on the ribose immediately upstream of the scissile bond; a 2'-deoxy substitution abolished cleavage but not binding.
- Evidence: "Csy4 activity requires the presence of a 2'-hydroxyl group on the ribose immediately upstream of the cleavage site."
H7.5 The 1.8 Å Csy4–RNA co-crystal structure: an arginine-rich helix clamping the major groove
- What was done: Csy4 (wild-type and catalytically active S22C mutant) was co-crystallized with the non-cleavable 16-nt minimal RNA (2'-deoxy at G20); three crystal forms were solved to 2.3, 2.6 and 1.8 Å.
- What came out: Csy4 is a two-domain protein (N-terminal ferredoxin-like domain, C-terminal RNA-binding domain) that clamps the RNA hairpin between its body and an arginine-rich helix (α3, residues 108–120) inserted into the major groove. Phe155 stacks on the C6–G20 closing base pair, anchoring the ssRNA–dsRNA junction; Arg102 and Gln104 make sequence-specific major-groove contacts to G20 and A19; helix arginines (Arg114/115/118/119, His120) contact the phosphate backbone.
- Evidence: "The RNA hairpin is clamped into a highly basic groove between the main body of the protein and an arginine-rich helix."
H7.6 Catalytic residues His29 and Ser148: mutants abolish cleavage without disrupting binding
- What was done: Point mutants of conserved active-site residues (His29, Ser148, Tyr176) and recognition residues (Arg102, Gln104, Phe155) were tested for in vitro cleavage of pre-crRNA (5 min at 25°C) and for RNA binding; His29 was also mutated to lysine.
- What came out: His29 and Ser148 are invariant and flank the scissile phosphate; mutating either abolished cleavage without disrupting RNA binding. Tyr176→Phe retained activity (Tyr176 orients His29). Arg102Ala and Phe155Ala impaired processing (substrate orientation) while Gln104Ala did not. His29Lys partially restored activity, implying His29 acts as a proton donor to the 5' leaving group; Ser148 likely positions/activates the 2'-OH via a 2'-3' cyclic mechanism.
- Evidence: "Mutations of His 29 or Ser 148 (to alanine and cysteine, respectively) completely abolished cleavage activity without disrupting RNA binding."
What it made possible next
This paper established the first mechanistic and structural picture of how a CRISPR protein recognizes and matures a guide RNA — the founding result of the Doudna lab's CRISPR program. It showed that crRNA maturation is achieved by a single protein reading both the sequence and the fold of the repeat stem-loop, discriminating cognate CRISPR transcripts from the cellular RNA pool through major-groove base contacts plus electrostatic backbone clamping. The central image it introduced — a protein clamping a structured guide RNA in a highly basic groove, with an arginine-rich helix threaded into the major groove and an aromatic residue stacking on the closing base pair — became a template for understanding CRISPR RNA processing and effector recognition more broadly. It fed directly into the lab's 2011 structural work on the multi-subunit Cascade surveillance complex (crRNA-guided recognition of foreign nucleic acid) and, more generally, into the logic that would culminate in RNA-guided Cas9: a protein that binds a guide RNA and uses it to license sequence-specific nucleic-acid cleavage. It also handed forward Csy4 itself as a programmable, sequence-specific RNA-cleaving tool.
Protocol extraction — [P] values printed in the paper
- [P] Organism = Pseudomonas aeruginosa UCBPP-PA14 (Pa14), Yersinia-subtype CRISPR/Cas
- [P] Cas gene content = six Cas genes (Cas1, Cas3, Csy1–4) flanked by two CRISPR elements
- [P] Repeat length = 28-nt near-identical direct repeats; spacer length = ~32-nt unique
- [P] Mature crRNA = 60-nt (32-nt phage-derived spacer + 8-nt 5' repeat + 20-nt 3' repeat flanks)
- [P] Cleavage time course = 30 s, 1 min, 5 min
- [P] Metal conditions tested = no exogenous metal; 2.5 mM MgCl2; 2.5 mM EDTA
- [P] Mutant cleavage assay conditions = 5 min at 25°C
- [P] Product analysis = acid phenol-chloroform extraction, 15% denaturing PAGE, SYBR Gold staining
- [P] Csy4 isoelectric point = pI 10.2
- [P] Synthetic CRISPR construct for specificity test = 8 repeat sequences + 7 identical spacer sequences
- [P] Co-purified protected fragment = ~19-nt crRNA
- [P] Minimal cleavable substrate = 16-nt (repeat stem-loop + 1 downstream nucleotide)
- [P] Catalysis-abolishing modification = 2'-deoxyribonucleotide at G20 (immediately upstream of cleavage site)
- [P] Crystal forms/resolutions = 2.3 Å, 2.6 Å, 1.8 Å (WT + two S22C-mutant forms)
- [P] Catalytically active crystallization mutant = S22C
- [P] Csy4 domain boundaries = N-terminal ferredoxin-like domain (residues 1–94); C-terminal domain (residues 95–187)
- [P] Arginine-rich helix = α3, residues 108–120
- [P] DALI rmsd vs CasE = 3.8 Å (over N-terminal 111 Cα); vs Cas6 = 3.9 Å (over 104 Cα)
- [P] Sequence identity to CasE/Cas6 = <10%
- [P] RNA stem base pairing = nucleotides 6–10 pair with 16–20 (A-form stem)
- [P] Pentaloop = GUAUA, sheared G11–A15 pair, extruded U14
- [P] Base-stacking residue = Phe155 (stacks on C6–G20 base pair)
- [P] Major-groove sequence-specific contacts = Arg102 → G20; Gln104 → A19; Arg115 → G6 base
- [P] Phosphate-backbone contacts (helix) = Arg114, Arg115, Arg118, Arg119, His120 → phosphates of nucleotides 7–12
- [P] Scissile phosphate = between G20 and C21
- [P] Active-site residues = His29 and Ser148 (both invariant); Ser148 4.6 Å from 2' ribose carbon of G20; His29 and Gln149 backbone amide contact scissile phosphate
- [P] Additional conserved residue near scissile phosphate = Tyr176
- [P] Mutations abolishing cleavage (binding intact) = His29Ala, Ser148Cys, Arg102Ala, Phe155Ala (severely impaired)
- [P] Mutations not disrupting activity = Tyr176Phe, Gln104Ala
- [P] Partial rescue mutant = His29Lys
- [P] Proposed catalytic mechanism = via 2'-3' cyclic intermediate/product (His29 = proton donor to 5' leaving group)
- [P] X-ray beamlines = ALS 8.2.2 and 8.3.1 (Lawrence Berkeley National Laboratory)
- [P] Published as = Science 2010 Sep 10;329(5997):1355–1358
Gaps requiring [S]/[M] — items not stated in the main text
- [S] Csy4 expression/purification protocol (vector, tag, host strain, induction, chromatography) — "data not shown"
- [S] In vitro transcription conditions for pre-crRNA substrate (polymerase, template, yields)
- [S] Cleavage/binding reaction buffer composition, salt, pH, enzyme:RNA ratios, and RNA amounts
- [S] Binding-assay method and quantitative Kd values for wild-type vs mutants (referenced as fig. S1, not in main text)
- [S] Crystallization conditions, cryoprotection, space groups, and refinement statistics (table S1, not in main text)
- [S] Sequences of the RNA oligonucleotide panel used to map the minimal substrate and base-pair requirements (data not shown; fig. S7)
- [S] Sequence alignment / homolog set defining "invariant" His29 and Ser148 (fig. S6)
- [S] Number of biological/technical replicates and any statistical treatment of cleavage quantitation
- [S] Exact affinity resin/tag and elution conditions for the E. coli co-expression specificity experiment
H8 — Cryo-EM structures of Cascade reveal an RNA-guided surveillance complex
Wiedenheft B, Lander GC, Zhou K, Jore MM, Brouns SJJ, van der Oost J, Doudna JA, Nogales E. "Structures of the RNA-guided surveillance complex from a bacterial immune system." Nature 2011;477(7365):486-489. PMID 21938068. (Doudna = co-corresponding senior author with Nogales, cryo-EM; Wiedenheft/Lander co-first.)
The question this paper set out to answer — In E. coli, crRNAs are known to be loaded into a 405-kDa 11-subunit ribonucleoprotein called Cascade that is required for phage protection, but the arrangement of its subunits and, crucially, the mechanism by which it recognizes a complementary target were unknown. How is the guide crRNA physically displayed within a protein scaffold, how does an invading nucleic acid engage it, and how does target recognition translate into a signal for destruction? The authors used single-particle cryo-EM to solve subnanometre structures of Cascade before and after target binding to answer these questions.
Experiments and results
H8.1 — Cryo-EM structure of unbound Cascade (~8 Å)
What was done: Single-particle cryo-EM reconstruction of the intact 11-subunit Cascade complex, with rigid-body docking of two crystallized subunits and known stoichiometry.
What came out: A "sea-horse"-shaped architecture with a helical backbone of six CasC subunits; the crRNA lies in a groove along the concave CasC surface, protected yet displayed for base pairing; resolution resolved secondary-structure elements in all components.
Evidence: "the crRNA is displayed along a helical arrangement of protein subunits that protect the crRNA from degradation while maintaining its availability for base pairing".
H8.2 — 3' end: CasE forms the head bound to the crRNA stem-loop
What was done: Fitted a Thermus thermophilus CasE:stem-loop co-crystal structure into the head density.
What came out: CasE (the endoribonuclease) caps the head, remaining bound to the 3' stem-loop of the mature 61-nt crRNA, which protrudes as a "beak."
Evidence: "CasE remains bound to the 39 [3'] end of the mature crRNA and the RNA stem–loop protrudes like a 'beak' from the head".
H8.3 — 5' handle hooked at the CasA/CasC6 tail; programmed helical capping
What was done: Traced the crRNA path and analyzed CasC helical symmetry, including a helical reconstruction of C1–5.
What came out: The 5' end forms a hook in a pocket between CasC6 and CasA; C1–5 form a right-handed helix (135 Å pitch) but C6's distal domain is rotated ~160°, breaking the helix and exposing the 5' spacer region.
Evidence: "The 5′ end of the crRNA terminates within the tail of the complex, forming a hook-like structure in a pocket between CasC6 (C6...) and CasA".
H8.4 — Target-bound structure (~9 Å) with a 32-nt complementary RNA
What was done: Solved cryo-EM structure of Cascade bound to a 32-nucleotide ssRNA complementary to the crRNA spacer.
What came out: The crRNA:target duplex is not one contiguous helix but five short duplex segments of 4–5 bp each, connected by non-helical pinch points at individual CasC subunits.
Evidence: "we observe density consistent with five short duplex segments, each accommodating four or five base pairs of double-stranded RNA".
H8.5 — Target binding triggers a concerted conformational change
What was done: Compared unbound and target-bound maps; corroborated with limited proteolysis on RNA and DNA targets.
What came out: Coordinated movements — CasE rotates ~15°, both CasB subunits slide ~17 Å toward the tail, CasA rotates ~30° (hinged at CasD), and the 5' hook / C6 distal domain is disrupted.
Evidence: "we observed a concerted conformational change in the locations and orientations of CasE, CasB and CasA".
H8.6 — Seed sequence has high target-binding affinity (tiling gel-shift assays)
What was done: Native electrophoretic mobility shift assays with a series of 16-nt target DNAs tiling across the crRNA in 8-nt steps.
What came out: Highest affinity for targets covering the 5' "seed" region (nucleotides 1–5, 7–8), with affinity decreasing toward the 3' end — the whole spacer is accessible but the seed is preferred.
Evidence: "we observed high-affinity interactions for targets that include the seed region, and that binding affinities decrease with increasing steps in the 3′ direction".
H8.7 — A model for surveillance and Cas3 recruitment
What was done: Synthesized the structural and binding data into a mechanistic model.
What came out: Seed-first binding nucleates duplex formation propagating 3' in 4–5-bp increments, shortening the crRNA and triggering the conformational change that may signal Cas3 recruitment for target destruction.
Evidence: "This conformational rearrangement may serve as a signal that recruits a trans-acting nuclease (Cas3) for destruction of invading nucleic-acid sequences".
What it made possible next
This work establishes, at structural resolution, the central principle that a CRISPR guide RNA is held by a protein scaffold that both protects it and presents it for Watson–Crick base pairing with an invading target — and that target recognition begins at a high-affinity "seed" and then propagates, driving a conformational change that licenses nuclease action. These three ideas — a protein-displayed crRNA that base-pairs a complementary target, a seed region that governs recognition specificity, and RNA-guided surveillance as programmable targeting — are precisely the conceptual scaffolding that Doudna's lab carried forward. Cascade is a multi-subunit (Type I) machine; the drive to understand and simplify RNA-guided targeting points directly toward the single-protein Cas9 of 2012, where a programmable guide RNA directs a nuclease to a complementary DNA target.
Protocol extraction — [P] values printed in the paper
- [P] Complex: 405-kDa ribonucleoprotein; 11 subunits, stoichiometry CasA1B2C6D1E1crRNA1 (1 CasA, 2 CasB, 6 CasC, 1 CasD, 1 CasE); 61-nucleotide crRNA; encoded by 8 cas genes; CRISPR = 29-nt repeats with 32-nt spacers.
- [P] Target-bound complex assembled with a 32-nt ssRNA complementary to the crRNA spacer.
- [P] CasC helix: first five subunits (C1–5) right-handed, pitch 135 Å; C6 distal domain rotated ~160°.
- [P] Target-binding motions: CasE rotates ~15°, CasB dimer moves ~17 Å, CasA rotates ~30°.
- [P] Resolutions (Fourier shell correlation): unbound 8.8 Å (FSC 0.5) / 7.7 Å (FSC 0.143); target-bound 9.2 Å (0.5) / 8.0 Å (0.143).
- [P] Final particle counts: 275,573 (unbound) and 176,090 (target-bound); initial automatic selections 498,137 / 389,166; images collected 2,370 (unbound) / 1,406 (target-bound).
- [P] Seed sequence = crRNA nucleotides 1–5 and 7–8; gel-shift tiling with 16-nt target DNAs in 8-nt steps.
- [P] Cryo-EM: Tecnai F20 Twin at 120 keV, nominal mag 100,000× (1.15 Å/pixel specimen level), low dose ~20 e−/Ų, defocus −0.8 to −2.5 µm, Gatan 4,000×4,000 CCD (15-µm pixel); box 288×288; final pixel 2.3 Å (binned ×2).
- [P] Expression: E. coli BL21(DE3), 0.5 mM IPTG at OD600 = 0.5, overnight at 16 °C; N-terminal Strep-II tag on CasB.
- [P] Lysis buffer: 100 mM Tris pH 8.0, 300 mM KCl, 1 mM EDTA, 1 mM TCEP, 5% glycerol. Gel filtration buffer: 25 mM HEPES pH 7.5, 100 mM KCl, 1 mM TCEP.
- [P] Target-bound prep: fivefold molar excess oligoribonucleotide, 37 °C for 15 min.
- [P] EMSA buffer: 25 mM HEPES pH 7.5, 100 mM KCl, 1 mM TCEP, 1% glycerol, 1 mM MgCl2, 1 mg/ml tRNA; 5' 32P-labeled ssDNA; 15 min at 37 °C; 6% polyacrylamide gels.
- [P] Limited proteolysis: 30 µM trypsin, 3.7 µM Cascade, 25 °C, 100 µl; 12% SDS-PAGE.
- [P] Grids: C-flats, glow-discharged 60 s; FEI Vitrobot, 4 °C, 100% humidity, 3-s blot; ~1.2 mg/ml sample. EMDB accessions 5314 (unbound), 5315 (target-bound).
- [P] Software: LEGINON, APPION, ACE2/CTFFIND, IMAGIC (MSA–MRA), FINDEM, EMAN/EMAN2/SPARX, SPIDER, UCSF Chimera.
Gaps requiring [S]/[M]
- [S] The physiological target in vivo is DNA, but the target-bound structure used a 32-nt ssRNA "to achieve maximal target site occupancy and sample homogeneity"; the paper infers similarity from limited proteolysis rather than solving a DNA-bound structure.
- [S] Cas3 recruitment is proposed as a model ("may serve as a signal"); no structural or biochemical demonstration of Cas3 binding to the conformationally changed complex is provided here.
- [S] At ~8–9 Å resolution, base pairs and the crRNA:target register are inferred from density width and rigid-body docking, not directly resolved; atomic-level interactions require higher-resolution [M] methods.
- [S] The protospacer-adjacent motif (PAM) role in target licensing is noted only via the proteolysis substrate design and not structurally resolved.
H9 — Cas9 as a programmable dual-RNA-guided DNA endonuclease (the CRISPR-Cas9 discovery)
Jinek M, Chylinski K, Fonfara I, Hauer M, Doudna JA, Charpentier E. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Science 2012;337(6096):816-21. PMID 22745249. (Doudna = co-senior/corresponding author with Charpentier; both marked * correspondence.)
The question this paper set out to answer — Type II CRISPR/Cas systems were known to provide bacteria adaptive immunity, and the Cas9 protein was hypothesized to be involved in both crRNA maturation and crRNA-guided silencing of foreign DNA, but Cas9's direct role in target DNA destruction had not been investigated. The paper set out to test whether and how the S. pyogenes Cas9 endonuclease cleaves target DNA, what RNA components it requires, how it recognizes and cuts its target, and — critically — whether the natural two-RNA guide could be simplified into a single engineered RNA to create a programmable DNA-cutting tool for genome editing.
Experiments and results
H9.1 — Cas9 requires both crRNA and tracrRNA (plus Mg2+) to cleave DNA
What was done: Purified S. pyogenes Cas9 was tested for cleavage of plasmid and short linear dsDNA bearing a protospacer complementary to a mature crRNA plus a bona fide PAM, with and without tracrRNA.
What came out: crRNA alone did not support cleavage; adding tracrRNA (which base-pairs the crRNA repeat) triggered cleavage. The reaction required Mg2+ and a crRNA complementary to the target; a non-cognate crRNA did not cut.
Evidence: "mature crRNA alone was incapable of directing Cas9-catalyzed plasmid DNA cleavage... However, addition of tracrRNA... triggered Cas9 to cleave plasmid DNA."
H9.2 — Cleavage is site-specific, blunt, 3 bp upstream of the PAM
What was done: Sequencing/primer-extension and end-labeled marker mapping of plasmid and short dsDNA cleavage products.
What came out: Plasmid cleavage produced blunt ends three base pairs upstream of the PAM. On duplexes, the complementary strand is cut 3 bp upstream of the PAM; the non-complementary strand is cut within 3-8 bp upstream (then trimmed 3'-5'). Cleavage rates 0.3-1 min-1, multiple-turnover enzyme.
Evidence: "Plasmid DNA cleavage produced blunt ends at a position three base pairs upstream of the PAM sequence."
H9.3 — Two nuclease domains: HNH cuts complementary strand, RuvC cuts non-complementary strand; mutants nick
What was done: Cas9 variants with inactivating point mutations in the catalytic residues of either the HNH or RuvC-like domain were purified and assayed on plasmid and strand-radiolabeled dsDNA.
What came out: Mutant Cas9 produced nicked open-circular plasmid (each domain cuts one strand); wild-type produced linear DNA. HNH cleaves the complementary strand, RuvC-like cleaves the non-complementary strand.
Evidence: "the Cas9 HNH domain cleaves the complementary DNA strand, while the Cas9 RuvC-like domain cleaves the non-complementary DNA strand."
H9.4 — tracrRNA is required for target DNA binding; minimal RNA regions defined
What was done: EMSA of catalytically inactive Cas9 binding target DNA ± crRNA/tracrRNA; cleavage tests with truncated tracrRNA and crRNA.
What came out: tracrRNA substantially enhanced target DNA binding (little binding with Cas9 or Cas9-crRNA alone). A truncated tracrRNA (nt 23-48) still supported robust cleavage; crRNA could lose its 3'-terminal 10 nt but not its 5' end.
Evidence: "Addition of tracrRNA substantially enhanced target DNA binding by Cas9, whereas little specific DNA binding was observed with Cas9 alone or Cas9-crRNA."
H9.5 — The crRNA's 5'-terminal 20-nt segment programs the target; a 3' seed sequence governs specificity
What was done: Analysis of predicted tracrRNA:crRNA structure; protospacer point-mutation and mismatch series scored by in vivo plasmid maintenance and in vitro cleavage.
What came out: The 5'-terminal 20 nucleotides of the crRNA (variable in sequence) are available for target DNA binding. Mutations near the PAM/cleavage site were not tolerated, defining a 3' "seed"; ≥13 contiguous PAM-proximal base pairs required, up to six 5' mismatches tolerated.
Evidence: "the 5'-terminal 20 nucleotides of the crRNA, which vary in sequence in different crRNAs, are available for target DNA binding."
H9.6 — Target requires a PAM (NGG); both G's needed, licenses R-loop formation
What was done: Transformation assays and in vitro cleavage/EMSA with dsDNA duplexes carrying mutations in the PAM on either or both strands (protospacers 2 and 4).
What came out: The S. pyogenes PAM is an NGG consensus; cleavage was especially sensitive to non-complementary-strand PAM mutations; both G's required; PAM mutation reduced Cas9-RNA affinity for target. PAM required only for dsDNA (not ssDNA), licensing strand invasion/R-loop.
Evidence: "the PAM conforms to an NGG consensus sequence, containing two G:C base pairs that occur one base pair downstream of the crRNA binding sequence."
H9.7 — THE key engineering result: a single chimeric guide RNA directs sequence-specific cleavage
What was done: crRNA and tracrRNA were fused into a single chimeric RNA (target-recognition sequence at the 5' end followed by a hairpin retaining tracrRNA:crRNA base-pairing, joined with a GAAA tetraloop); tested on plasmid and short dsDNA.
What came out: The longer chimeric RNA guided Cas9 cleavage like the dual RNA, cutting at the identical position. (Shorter chimera worked poorly, showing nt 5-12 beyond the base-pairing region matter.)
Evidence: "The dual-tracrRNA:crRNA, when engineered as a single RNA chimera, also directs sequence-specific Cas9 dsDNA cleavage."
H9.8 — Reprogramming the guide to cut chosen sites (GFP)
What was done: Five different chimeric guide RNAs were engineered to target the GFP gene and tested against a GFP-carrying plasmid in vitro.
What came out: In all five cases Cas9 cleaved the plasmid at the correct target site, showing rational chimeric-RNA design is robust and can target essentially any sequence bearing an adjacent GG.
Evidence: "In all five cases, Cas9 programmed with these chimeric RNAs efficiently cleaved the plasmid at the correct target site."
What it made possible next
The discovery: a single programmable, RNA-guided DNA-cutting enzyme. By fusing the natural dual tracrRNA:crRNA into one single-guide RNA (sgRNA/chimera), Cas9 becomes an efficient, versatile, and programmable endonuclease that can be directed to cut any chosen dsDNA sequence simply by changing the ~20-nt target sequence in the guide (with only a GG/NGG-PAM constraint). This established the molecular basis and toolkit for RNA-programmed genome editing, positioning Cas9 as an alternative to ZFNs and TALENs and setting up subsequent demonstration of genome editing in eukaryotic cells.
Protocol extraction — [P] values printed in the paper
- [P] Enzyme: Cas9 from Streptococcus pyogenes, purified from an overexpression system (fig. S2).
- [P] Cofactor: cleavage requires magnesium (Mg2+).
- [P] Guide lengths: mature crRNA-sp2 = 42 nt; tracrRNA = 75 nt (also active constructs nt 4-89 and nt 23-89); minimal functional tracrRNA retains nt 23-48; crRNA tolerates loss of 3'-terminal 10 nt but not 5' end.
- [P] Target-programming segment: 5'-terminal 20 nucleotides of the crRNA base-pair with the target.
- [P] PAM: NGG consensus (two G:C bp), one bp downstream of the crRNA-binding sequence on the target DNA; both G nucleotides required.
- [P] Nuclease-domain mutants: RuvC-like = D10A; HNH = H840A (double mutant D10A/H840A is catalytically dead, used for binding/EMSA). Mutants nick DNA (open-circular product); wild-type linearizes.
- [P] Cleavage position: blunt ends 3 bp upstream of PAM (complementary strand at 3 bp; non-complementary strand within 3-8 bp upstream, then 3'-5' trimmed).
- [P] Kinetics: single-turnover rates 0.3-1 min-1; multiple-turnover enzyme.
- [P] Single-guide/chimera design: 3' end of crRNA fused to 5' end of tracrRNA via a GAAA tetraloop linker; five GFP-targeting chimeras validated.
- [P] Seed/specificity: ≥13 contiguous PAM-proximal crRNA:DNA base pairs needed for efficient cleavage; up to 6 contiguous 5' mismatches tolerated.
Gaps requiring [S]/[M]
- [S] Exact reaction buffer composition, Mg2+ concentration, pH, temperature, and incubation times are in the Supplementary Materials and Methods, not the main text.
- [S] Precise catalytic-residue rationale/full domain-boundary definitions and ortholog panel details are in supplementary figures (fig. S7, S11).
- [S] Full chimeric-RNA and GFP-target sequences, oligonucleotide/protospacer sequences, and protein purification protocol reside in Supplementary Tables S1-S3 and Methods.
- [M] The paper is entirely in vitro (and bacterial transformation assays); demonstration of activity/editing in eukaryotic/mammalian cells is not shown here and is left for later work.
Sentences explicitly framing the result as programmable / for genome editing:
- "Our study further demonstrates that the Cas9 endonuclease family can be programmed with single RNA molecules to cleave specific DNA sites, thereby raising the exciting possibility of developing a simple and versatile RNA-directed system to generate dsDNA breaks (DSBs) for genome targeting and editing."
- "highlights the potential to exploit the system for RNA-programmable genome editing." (Abstract)
- "We propose an alternative methodology based on RNA-programmed Cas9 that could offer considerable potential for gene targeting and genome editing applications." (Conclusions)
- "the possibility of a single RNA-guided Cas9 is appealing due to its potential utility for programmed DNA cleavage and genome editing."
Machine Traces #3 — Evidence 2 of 4 · The chain — what each paper handed to the next
For each transition, the specific reagent, method, or concept one paper handed forward, why the next step could not have proceeded without it, and the provenance on both sides (the "enabled-next" statement in the earlier reconstruction and the point of use in the later one).
The chain — what each paper handed to the next
Read as ideas, the nine papers form a single argument: specificity can be programmed into a macromolecule by a base-pairing guide, and the machine that reads that guide can be understood — and then simplified — by solving its structure. Read at the level of Methods, the material lineage is thinner than the iPS path's reagent chain (there is no single physical molecule passed hand to hand from 1989 to 2012). Instead the concrete through-line is a method and a set of hands: the RNA-crystallography pipeline established in 1996 (heavy-atom/hexammine phasing, crystallisation of discrete folded RNA/RNP domains) and specific personnel — Kaihong Zhou crystallises across H3, H4, H6, H7, H8; Martin Jinek runs from Csy4 (H7) to Cas9 (H9); Blake Wiedenheft from Csy4 (H7) to Cascade (H8). That the Doudna path is carried by transferable concepts + a shared structural pipeline rather than by a physical reagent lineage is itself a finding, and a contrast with the iPS path.
H1 → H2 (1989 → 1991)
- Handed forward: the engineered, guide-programmable ribozyme system — a truncated group I intron recast as a template-directed ligase whose product is set by base-pairing to a separate external template — plus the intron-construct toolkit (sunY/Tetrahymena constructs; primer/ligator/template oligonucleotide design; spermidine to erase intrinsic specificity).
- Why H2 needed it: the multisubunit ribozyme of 1991 is the same engineered-intron system taken one step further — dissected into separable subunits that reassemble.
- Provenance: from — "an engineered, guide-programmable RNA catalyst … whose product is set entirely by base-pairing of substrate oligonucleotides to a separate external template." to — H2 divides the sunY-td derivative (JD929) into three fragments that assemble into active catalytic complexes.
H2 → H3 (1991 → 1996)
- Handed forward: the guiding tension that a ribozyme must be folded to catalyse yet unfolded to be copied, and the demonstration that a modular ribozyme assembles from separate subunits but only inefficiently — pointing squarely at the need to see how the catalytic core folds in three dimensions.
- Why H3 needed it: the inefficiency of multisubunit assembly is the motivation for solving the folded structure of a group I intron domain.
- Provenance: from — "the need to understand how the catalytic core folds and packs in three dimensions." to — H3 crystallises the P4-P6 domain of the Tetrahymena intron.
H3 → H4 (1996 → 1998)
- Handed forward: the structure-first program and a reusable RNA-crystallography pipeline — heavy-atom/hexammine phasing (osmium/cobalt hexammine in major-groove sites), the strategy of solving discrete independently folding domains, and a vocabulary of tertiary motifs (GAAA tetraloop/receptor, A-rich bulge, ribose zippers, metal-bridged packing). Kaihong Zhou at the bench.
- Why H4 needed it: the HDV ribozyme structure applies the same phasing/crystallisation strategy to a compact catalytic RNA.
- Provenance: from — "the MAD/heavy-atom-hexammine phasing strategy for RNA … carried forward into the 1998 crystal structure of the hepatitis delta virus (HDV) ribozyme." to — H4 determines the HDV ribozyme fold (a nested double pseudoknot) and reads catalysis directly from structure (C75 as the proposed general base; no specific catalytic metal).
H4 → H5 (1998 → 2001)
- Handed forward: two transferable methods and one stance. (i) The U1A-protein crystallization chaperone — engineering a dispensable helix (P4) to bind a small basic protein so an otherwise poorly ordered RNA crystallizes — a general trick for solving difficult RNA/RNP targets. (ii) The consolidated MAD/heavy-atom phasing pipeline (Kaihong Zhou at the bench) now proven on a compact catalytic RNA. (iii) The stance that catalytic mechanism can be read directly off a folded structure (C75 positioned as general base) — confidence that emboldens tackling an RNA that acts on a machine.
- Why H5 needed it: with the lab's RNA-structure methods mature and a protein-assisted crystallization/EM mindset in hand, the next target is an RNA that changes the shape of a large ribonucleoprotein — the ribosome.
- Provenance: from — the U1A-chaperone technique and the "structure gives mechanism" confidence. to — H5 uses cryo-EM (with the Frank lab) to image the HCV IRES bound to the 40S subunit, an RNA acting on a large machine.
H5 → H6 (2001 → 2006)
- Handed forward: the concept of an RNA element as a conformational effector on a large ribonucleoprotein — "the first reported example of a structured viral RNA that binds to, and actively manipulates, the structure of a cellular machine" — reframing RNA from passive scaffold to active director of a protein machine.
- Why H6 needed it: it primes the move to a machine (Dicer) whose job is to process RNA into guides, and toward machines an RNA directs.
- Provenance: from — "an RNA guides and reconfigures a protein partner." to — H6 solves the structure of Dicer, the guide-RNA-generating machine.
H6 → H7 (2006 → 2010)
- Handed forward: the founding structural logic of small-RNA biogenesis — a defined nucleic-acid end is clamped by a binding module and a fixed distance acts as a ruler setting product length ("Dicer is a molecular ruler"), and appended RNA-binding modules confer specificity on an otherwise nonspecific nuclease — plus the method of crystallising an intact RNA-processing machine.
- Why H7 needed it: Csy4 is exactly this problem in bacteria — a protein that recognises and cuts a repeat to release a mature guide RNA at a defined position.
- Provenance: from — "the conceptual and methodological toolkit Doudna's lab carried into bacterial CRISPR RNA processing — Csy4 (2010)." to — H7 shows Csy4 cleaving pre-crRNA at the base of the repeat stem-loop.
H7 → H8 (2010 → 2011)
- Handed forward: the image of a protein clamping a structured guide RNA in a basic groove to license sequence-specific recognition/cleavage (arginine-rich helix in the major groove; aromatic stacking on the closing base pair); Csy4 itself as a programmable RNA-cleaving tool; and shared personnel (Jinek, Wiedenheft).
- Why H8 needed it: Cascade is the multi-subunit generalisation — a protein scaffold that holds a crRNA and presents it for target base-pairing.
- Provenance: from — "a protein that binds a guide RNA and uses it to license sequence-specific nucleic-acid cleavage … fed directly into the 2011 Cascade structure." to — H8's cryo-EM of Cascade displaying the crRNA.
H8 → H9 (2011 → 2012)
- Handed forward: the central CRISPR principle at structural resolution — a protein-displayed crRNA base-pairs a complementary target from a high-affinity seed, and this triggers a conformational change that licenses nuclease action — plus the explicit drive to simplify RNA-guided targeting from a multi-subunit machine toward a single protein.
- Why H9 needed it: Cas9 is the single-protein realisation of the same logic; the seed/target-pairing/conformational-licensing principle is exactly what Cas9 executes.
- Provenance: from — "the drive to understand and simplify RNA-guided targeting points directly toward the single-protein Cas9 of 2012." to — H9 shows Cas9 cutting DNA under crRNA:tracrRNA guidance, then under a single fused guide.
The confluence at 2012
The Cas9 result did not come from the Doudna node-chain alone. Three external lines converged on it: Charpentier's tracrRNA (Deltcheva et al. 2011) supplied the second RNA that Cas9 requires; the field's demonstrations that CRISPR targets DNA (Marraffini & Sontheimer 2008; Garneau et al. 2010) fixed the substrate; and parallel Cas9 biochemistry (Sapranauskas 2011; Gasiūnas/Šikšnys 2012) was underway independently. Doudna's own chain supplied what none of these did: two decades of how a protein reads and is directed by a structured guide RNA, and the reflex to simplify — which produced the paper's decisive engineering move, the single-guide RNA.
Machine Traces #3 — Evidence 3 of 4 · The engineering reduction — two RNAs to one
Where the iPS path's #3-analogous crux was the 24 → 10 → 4 factor narrowing, the crux of the Cas9 paper is a reduction of components: from a two-enzyme-domain protein guided by two separate RNAs, down to one protein guided by one engineered RNA. This section reconstructs that reduction from H9's Results.
The reduction of the 2012 paper — from two RNAs to a single guide
The 2012 paper establishes Cas9 as an RNA-guided DNA endonuclease and then strips it to its minimal programmable form, in a sequence of dissections:
Stage 1 — Cas9 needs both RNAs. Purified S. pyogenes Cas9 cleaves target dsDNA only in the presence of both the mature crRNA (which carries the ~20-nt guide) and the tracrRNA. Neither RNA alone suffices; tracrRNA is not merely a maturation factor but a required cofactor of the cleavage reaction.
Stage 2 — two domains, two strands. Cas9 carries two nuclease domains, HNH and RuvC-like. Active-site point mutants separate their jobs: an HNH mutant (H840A) and a RuvC mutant (D10A) each convert Cas9 from a double-strand cutter into a nickase. HNH cleaves the DNA strand complementary to the guide; RuvC cleaves the displaced non-complementary strand. The cut falls a fixed short distance (3 bp) upstream of the PAM.
Stage 3 — the target rules: guide + PAM. Cleavage requires Watson–Crick complementarity between the ~20-nt guide segment of the crRNA and one DNA strand, and a short protospacer-adjacent motif (PAM, NGG) in the target. Changing the guide sequence retargets the enzyme; destroying complementarity or the PAM abolishes cleavage.
Stage 4 — the reduction: two RNAs → one. The natural dual tracrRNA:crRNA is re-engineered into a single chimeric guide RNA (sgRNA) by fusing the 3′ end of the crRNA to the 5′ end of the tracrRNA with a loop. This single RNA is sufficient to program Cas9 for sequence-specific dsDNA cleavage.
- Evidence (verbatim, H9): "The dual-tracrRNA:crRNA, when engineered as a single RNA chimera, also directs sequence-specific Cas9 dsDNA cleavage."
Stage 5 — programmability demonstrated. Guides were designed against chosen sequences (including sites in a GFP gene), and Cas9 cut where the guide directed — establishing that a single ~20-nt sequence change reprograms the enzyme to any matching, PAM-flanked target.
What this reduction shows is the same move seen across the whole path, executed one last time: a functional machine is dissected into separable parts and then recombined into the minimal programmable unit. In 1991 (H2) Doudna split one ribozyme into three trans-acting subunits; in 2012 she and Charpentier fused two natural RNAs into one. Split and fuse are the same design instinct — control specificity through a modular, base-pairing guide — and the sgRNA is where it lands.
Machine Traces #3 — Evidence 4 of 4 · Resolution: what full text added, and coverage
This section records what reading full text (rather than abstracts) added, and — because the source PDFs varied in quality — the honest per-paper coverage.
Resolution and coverage
Per-paper extraction coverage
| Paper | Source quality | How read | Experiments | [P] | [S]/[M] | Verbatim fidelity |
|---|---|---|---|---|---|---|
| H1 1989 | scan, no born-digital exists | pdftoppm → vision | 6 | 18 | 6 + 1[M] | high (verbatim) |
| H2 1991 | scan, no born-digital exists | pdftoppm → vision | 6 | 10 | 4 + 1[M] | high (verbatim) |
| H3 1996 | scan, no born-digital exists | pdftoppm → vision | 12 | 11 | 5 + 1[M] | high (verbatim) |
| H4 1998 | clean born-digital | Drive text | 6 | 12 | 5 | high (verbatim) |
| H5 2001 | clean reprint | Drive text | 4 | 6 | 6 | high |
| H6 2006 | clean reprint | Drive text | 6 | 8 | 5 | high |
| H7 2010 | PMC author manuscript | Drive text | 6 | 34 | 9 | high (verbatim) |
| H8 2011 | PMC author manuscript | Drive text | 7 | 15 | 4 | high (verbatim) |
| H9 2012 | PMC author manuscript | Drive text | 8 | 10 | 3 + 1[M] | high (verbatim) |
| Total | 9 of 9 read | ~61 | ~125 | ~50 |
Coverage: all 9 nodes read at full or near-full text, and verbatim quotation is reliable across the corpus — only two individual sentences (one each in H5 and H6, from clean-text reads) retain a precautionary (OCR-uncertain) tag. The three pre-digital papers (H1, H2, H3) have no born-digital version and were recovered by rendering the scanned pages to images and transcribing by vision; this route (used first for H1) removed the earlier (OCR-uncertain) flags that the connector's raw text layer had forced on H2 and H3. This is reported, not imputed.
What full text added over an abstract-level trace
- The path is conceptual, not material. Full text confirms there is no physical reagent handed 1989→2012; the through-line is the RNA-crystallography pipeline and shared personnel (Zhou, Jinek, Wiedenheft). An abstract-level trace would have missed that the "chain" is a method and a set of hands, not a molecule.
- The recurring "split/fuse a guide-programmable RNA" habit is visible only in Methods. The 1991 three-subunit dissection and the 2012 sgRNA fusion are mirror-image operations; the mirror is legible only when both papers' construct designs are read.
- The CRISPR turn is a field-entry, not an internal anomaly. Unlike the iPS path (where a liver-tumour surprise bent the path), Doudna's corpus shows no internal experimental anomaly forcing the jump into CRISPR around 2006–2008; the redirect came from outside the publication record (a collaboration approach). The corpus shows what she did and when; motive belongs to her essays and lectures and is outside PubMed — flagged accordingly.
- "Clamp/measure-and-cut" is one idea across four papers. Dicer's ruler (H6), Csy4's stem-loop clamp (H7), Cascade's crRNA display (H8) and Cas9's guide+PAM licensing (H9) are the same mechanistic template at increasing generality — a continuity invisible at abstract level.
Induced thought patterns
Reported only where a habit recurs across ≥3 independent transitions (anything appearing once is an anecdote).
- Programmable specificity through a base-pairing guide, decoupled from the catalytic core. The single most consistent habit: H1 (external template sets the product) → H2 (separable guide/catalyst subunits) → H6 (guide RNAs of defined length) → H7/H8 (crRNA reads sequence+structure) → H9 (a ~20-nt guide reprograms the enzyme). Specificity is always delegated to Watson–Crick pairing, never hard-wired into the protein/core.
- Split-and-fuse modularity. Take a working machine apart into separable parts and recombine it into the minimal unit: isolate the P1 substrate (H1), split the intron into three subunits (H2), fuse two RNAs into one sgRNA (H9). The same instinct, run forward and backward.
- Structure-first mechanism. Six of nine nodes are structures (H3, H4, H5, H6, H7, H8). The lab's reflex is: to understand or engineer an RNA machine, solve its folded structure. The 1996 crystallography pipeline is the material spine of the whole path.
- Clamp/measure-and-cut. A binding module anchors a fixed feature of a nucleic acid to license a precise catalytic event: PAZ measures the dsRNA end (H6), Csy4 clamps the repeat stem-loop (H7), Cascade displays the crRNA (H8), Cas9's PAM+seed license the cut (H9).
- Simplify toward a programmable tool. Repeatedly the endpoint is not just understanding but a minimal, reprogrammable device — the engineered ligase (H1), Csy4 as a tool (H7), and the sgRNA-programmed Cas9 (H9).
Influences credited (not nodes on Doudna's path)
- Cech & Zaug; Kruger et al. (1982); Guerrier-Takada & Altman (1983) — the discovery that RNA can be a catalyst (group I intron; RNase P). The premise of H1–H4.
- Mojica et al. (2005) — CRISPR spacers match foreign genetic elements; PAM. Jansen et al. (2002) — cas genes and the CRISPR name.
- Barrangou & Horvath et al. (2007) — CRISPR confers adaptive immunity (Danisco).
- Brouns et al. (2008, Science) — small crRNAs guide antiviral defence; Cascade. (Confirmed: Doudna is not an author — an influence, not a node.)
- Marraffini & Sontheimer (2008); Garneau et al. (2010) — CRISPR/Cas targets DNA.
- Deltcheva, Chylinski, Charpentier et al. (2011) — tracrRNA and crRNA maturation; the second RNA that Cas9 requires. Charpentier is co-corresponding author of H9.
- Sapranauskas et al. (2011); Gasiūnas, Barrangou, Horvath & Šikšnys (2012) — parallel demonstration of Cas9 as an RNA-guided DNA nuclease.
Bibliographic details for the nine nodes were verified against PubMed; influence citations above are recorded from domain knowledge and should be PubMed-reconciled before publication.
References (nine nodes; verified against PubMed E-utilities, 2026-08)
- (Doudna & Szostak, 1989) Doudna JA, Szostak JW. RNA-catalysed synthesis of complementary-strand RNA. Nature. 1989;339(6225):519–522. PMID 2660003.
- (Doudna, Couture & Szostak, 1991) Doudna JA, Couture S, Szostak JW. A multisubunit ribozyme that is a catalyst of and template for complementary strand RNA synthesis. Science. 1991;251(5001):1605–1608. PMID 1707185.
- (Cate et al., 1996) Cate JH, Gooding AR, Podell E, Zhou K, Golden BL, Kundrot CE, Cech TR, Doudna JA. Crystal structure of a group I ribozyme domain: principles of RNA packing. Science. 1996;273(5282):1678–1685. PMID 8781224.
- (Ferré-D'Amaré, Zhou & Doudna, 1998) Ferré-D'Amaré AR, Zhou K, Doudna JA. Crystal structure of a hepatitis delta virus ribozyme. Nature. 1998;395(6702):567–574. PMID 9783582.
- (Spahn et al., 2001) Spahn CMT, Kieft JS, Grassucci RA, Penczek PA, Zhou K, Doudna JA, Frank J. Hepatitis C virus IRES RNA-induced changes in the conformation of the 40S ribosomal subunit. Science. 2001;291(5510):1959–1962. PMID 11239155.
- (Macrae et al., 2006) Macrae IJ, Zhou K, Li F, Repic A, Brooks AN, Cande WZ, Adams PD, Doudna JA. Structural basis for double-stranded RNA processing by Dicer. Science. 2006;311(5758):195–198. PMID 16410517.
- (Haurwitz et al., 2010) Haurwitz RE, Jinek M, Wiedenheft B, Zhou K, Doudna JA. Sequence- and structure-specific RNA processing by a CRISPR endonuclease. Science. 2010;329(5997):1355–1358. PMID 20829488.
- (Wiedenheft et al., 2011) Wiedenheft B, Lander GC, Zhou K, Jore MM, Brouns SJJ, van der Oost J, Doudna JA, Nogales E. Structures of the RNA-guided surveillance complex from a bacterial immune system. Nature. 2011;477(7365):486–489. PMID 21938068.
- (Jinek et al., 2012) Jinek M, Chylinski K, Fonfara I, Hauer M, Doudna JA, Charpentier E. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science. 2012;337(6096):816–821. PMID 22745249.
Provenance
- Subject: Jennifer A. Doudna → the 2012 dual-RNA-guided Cas9.
- Corpus construction: PubMed E-utilities via the sandboxed WebFetch tool (direct API access blocked by proxy; the discovery-path-tracer kernel could not run live). Author query
Doudna JA[au]returned 434 records spanning 1987–2026; the nine-paper skeleton was selected by hand and every node's title/authors/journal/PMID verified individually against PubMed. Doudna is a rare surname; homonym contamination is minimal and no disambiguation scoring was required. - Node authorship: all nine nodes confirmed Doudna-authored. Brouns et al. (2008) checked and excluded (no Doudna in author list) → influence, not node.
- Full-text source: the nine PDFs supplied by the user in Google Drive folder CRISPR-Path. Read via the Google Drive connector's text extraction, except the pre-digital scanned papers H1, H2 and H3 (no born-digital version exists for that era; each recovered by rendering pages to images → vision transcription). H4 was supplied after the initial build and read as clean born-digital text. Extraction quality per paper is tabulated in Evidence 4.
- Verbatim status: quotations reproduced from clean text (born-digital) or clean page-image vision transcription (H1, H2, H3) for all nine papers; the earlier (OCR-uncertain) flags on H2/H3 were removed after re-reading those scans by page-image vision. Two individual quotes (one each in H5 and H6) retain a precautionary (OCR-uncertain) tag from the original clean-text reads and should be treated as provisional. No independent second-pass word-match audit was run (the #2 "word-match 1.00" check was not reproduced here).
- Attribution is inference, not ground truth, for the causal links: chronological adjacency plus each paper's own "handed-forward" statement supports the chain, but the CRISPR-entry motive (2006–2008) is not in the publication record and is flagged as outside-PubMed.
- Influences: listed from domain knowledge; not all PubMed-reconciled.
This is the Lab Notebook for Machine Traces of Discovery Paths #3 — The Path to CRISPR-Cas9 (Jennifer Doudna). Nine-paper skeleton, all nine read at full/near-full text (H1, H2, H3 recovered by page-image vision transcription of pre-digital scans).
How to cite
Kitano, H. (2026). Lab Notebook for
Machine Traces of Discovery Paths #3 - The Path to CRISPR-Cas9 (Jennifer Doudna), A Comprehensive Analysis and Reconstruction, The Discovery Engine .
ORCID: 0000-0002-3589-1953
https://orcid.org/0000-0002-3589-1953
First published: August 16, 2026
Hiroaki Kitano