The Discovery Trace — Architectural Implications from CRISPR-Cas9
The Discovery Engine — Machine Traces of Discovery Paths — Architectural Implications
Disclaimer — Discovery Trace Series. The Discovery Trace series is an experimental series of articles attempting the semi- or fully-automatic reconstruction of the paths to major discoveries. Multiple AI tools have been deployed to reconstruct each "Discovery Path." To evaluate the progress of the technology, an initial set of articles is published prior to rigorous verification, and may therefore contain errors and omissions. Content will be updated as corrections are made, with notes on publication history and a correction log.
The first Architectural Implications entry proposed that a discovery path has two levels: a lower operation level, where operations act on objects in two pairs — Computational → Data (dry) and Protocol → Reagent (wet) — and an upper cognitive cycle that decides which operations to run and reads what they return: Hypothesis → Experimental Design → (operations) → Interpretation → Gap → next Hypothesis. Four translations sit on the cycle: Experimental Design (hypothesis → operations, downward), Interpretation Design (choosing, before reading, what will count as the signal), Interpretation (results → meaning, upward), and the step from a Gap to the next Hypothesis. The Gap comes in kinds — quantitative-deviation (a difference on a known axis), unanticipated-observation (a result off the predicted space), and confirmation (the prediction holds) — and turns can be of different work: verification vs optimization.
That structure was drawn from the iPS path. This companion reads the Doudna → CRISPR-Cas9 path (Machine Traces #3 https://thediscoveryengine.ai/machine-traces-of-discovery-paths-3-lab-notebook, nodes H1–H9) through the same lens, to test the structure on a second, very different discovery and to see what it must be extended to hold.
The headline: the structure fits, but three features of this path stretch it — (1) the operation level is not dry-then-wet but a tightly coupled wet+dry unit (structural biology); (2) the path's decisive redirect does not come from a Gap at all but from outside the cycle (an external opportunity), which the single-cycle diagram has no arrow for; and (3) because the path is structure-first, its Gaps skew toward confirmation/consolidation rather than anomaly — the surprises are fewer but pivotal.

The two additions the CRISPR case makes to the Discovery-Trace structure: a second, exogenous inbound arrow to the Hypothesis, and coupled wet+dry operation.
The full two-level structure is unchanged from the first Architectural Implications entry (see "The Structure of a Discovery Trace"). The CRISPR case adds exactly two things. (1) The Hypothesis gains a second, exogenous inbound arrow: the decisive turn came from outside the cycle, not from a Gap. (2) For structure-driven work the dry (Computational→Data) and wet (Protocol→Reagent) pairs fire as one coupled operation — the wet product exists only to feed the dry one. Everything else is as in AI-1.
The operation level of a structure-driven path: a coupled wet+dry unit
In the iPS path the two operation pairs often fired separately — a computational screen (digital differential display) narrowed a space, then a wet protocol (gene targeting) acted in it. The Doudna path is dominated by structural biology, where the two pairs fire as one coupled operation: a Protocol → Reagent step grows a crystal or vitrifies a complex, and a Computational → Data step (phasing, refinement, single-particle reconstruction) turns diffraction or images into a structure. Six of nine nodes (H3, H4, H5, H6, H7, H8) are this coupled unit. The Discovery Trace's two pairs are still the right primitives, but for this path they should be read as a bound pair, not alternatives — the wet product (a crystal) exists only to be consumed by the dry operation (a map), and neither half is interpretable alone. This is the operation-level signature of a structure-first program.
The path, turn by turn

The nine turns of the CRISPR-Cas9 path, each coloured by the kind of Gap it produced, with the exogenous-opportunity arrow between H6 and H7.
The nine nodes as nine turns of one cycle. Left-border colour marks the Gap kind (quantitative-deviation, unanticipated-observation, confirmation, optimization); the badge marks the dominant operation pair (WET, or WET+DRY coupled). The dashed panel between H6 and H7 marks the exogenous redirect into bacterial CRISPR.
Each row is one turn of the cycle: the Hypothesis it opened with, the Experimental Design that turned it into operations, which operation pair dominated, what Interpretation returned, and the kind of Gap the turn produced (which sets up the next Hypothesis).
| Turn | Hypothesis | Experimental Design | Dominant operation pair | Interpretation → Gap kind |
|---|---|---|---|---|
| H1 1989 | A modified group I intron can catalyse template-directed synthesis of a complementary strand (an RNA-replicase seed). | Truncate the intron (P2–P9); supply an external template + oligonucleotides; assay copying/ligation. | Wet (Protocol→Reagent), dry sequence design | It works but is markedly inefficient (high K_m, low yield). → quantitative-deviation |
| H2 1991 | The intron can be split into separable trans-acting subunits that reassemble into an active complex. | Interrupt at loops L6/L8 → three fragments; assay assembly + complementary-strand synthesis. | Wet | Assembles and copies a subunit, but again inefficient; exposes the tension folded-to-catalyse vs unfolded-to-be-copied. → quantitative-deviation + a conceptual obstacle |
| H3 1996 | The folded structure of a group I intron domain will reveal how a large RNA packs. | Crystallise P4-P6; phase with osmium/cobalt hexammine; solve. | Coupled wet+dry | A vocabulary of tertiary motifs (tetraloop–receptor, A-platform, ribose zipper). The prediction (structure will show packing) holds. → confirmation (consolidating; builds the method) |
| H4 1998 | The HDV ribozyme's fold explains its fast, metal-independent catalysis. | U1A-chaperone crystallisation; MAD; locate the active site via the 5′-OH leaving group. | Coupled wet+dry | Nested double pseudoknot; C75 as general base; no tightly bound catalytic metal (against metalloribozyme expectation); a new helix P1.1 where single strand was assumed. → unanticipated-observation (no metal; new pseudoknot) |
| H5 2001 | The HCV IRES has a defined structure that positions it on the ribosome. | Cryo-EM of the IRES–40S complex. | Coupled wet+dry | The IRES actively reshapes the ribosome (closes the mRNA cleft) — not a passive scaffold. → unanticipated-observation (RNA manipulates a machine) — the conceptual pivot |
| H6 2006 | Dicer's structure explains how it makes fixed-length small RNAs. | Crystallise Dicer; test the "ruler" model. | Coupled wet+dry | PAZ clamps one end; a fixed distance to the RNase III sites sets guide length ("molecular ruler"). Prediction holds. → confirmation (hands forward clamp/measure) |
| H7 2010 | A Cas protein processes pre-crRNA in a sequence/structure-specific way. | Screen six Cas proteins for cleavage; co-crystallise Csy4–RNA; mutate catalytic residues. | Coupled wet+dry | Csy4 clamps the repeat stem-loop; His29/Ser148 catalysis; the authors call the major-groove recognition "unexpected." → confirmation (which protein) + unanticipated-observation (recognition mechanism) |
| H8 2011 | How is the crRNA displayed, and how is a target recognised, in Cascade? | Cryo-EM of Cascade ± target RNA. | Coupled wet+dry | crRNA displayed along the backbone; seed-first recognition; target binding triggers a concerted conformational change. → confirmation + mechanistic refinement (seed model) |
| H9 2012 | Cas9 is an RNA-guided DNA endonuclease whose specificity is programmable. | Reconstitute cleavage in vitro; dissect the RNA requirement; HNH/RuvC domain mutants; map PAM; then fuse the two RNAs into one. | Wet (biochemistry) + dry design | Both RNAs required; two domains cut two strands; guide+PAM target; a single fused sgRNA suffices. → confirmation of the hypothesis + an optimization/compression turn (the sgRNA) |
Reading the turns
The quantitative-deviation Gaps open the path (H1, H2). Like much of any path, the earliest turns produce a number that falls short — the engineered ribozyme copies, but inefficiently. Those deviations are what push from "make it work" toward "understand why it barely works," which becomes the structural program. Note the between-turn move here is itself a redirection: the Gap at H2 is answered not by another biochemistry cycle but by changing the operation level to structure (H3). The Discovery Trace handles this as a new Experimental Design choice — a legitimate, non-automatic decision to switch how the question is asked.
The unanticipated-observation Gaps bend the path (H4, H5, H7). These are the turns that could not have come from a subtraction. H4's "no tightly bound metal" contradicted the reigning metalloribozyme picture; H5's IRES reshaping the ribosome appeared on an axis the hypothesis had not named (the experiment asked "where does it sit," the answer was "it moves the machine"); H7's recognition mechanism is flagged "unexpected" in the authors' own words. As the Architectural Implications piece argued, these off-axis results are the hardest for any schema to hold and the most consequential — and here they are the rungs by which the lab climbs from RNA-as-structure to RNA-as-director-of-machines.
The confirmation turns consolidate and build method (H3, H6, H8). A structure-first program produces many turns whose prediction simply holds — the structure does reveal the packing, the ruler, the display. These are not failures of drama; they are how a technique lineage is laid down. This is the operation-level counterpart of the essay's finding that the CRISPR path's through-line is a method, not a molecule: confirmation turns are where a method is proven and handed forward.
The final turn is a compression (H9). Exactly as the iPS path ended in an optimization (24 → 4 factors), the Cas9 turn ends not in a new Gap but in a reduction: two RNAs fused into one programmable guide. In the cycle's terms this is an optimization turn — searching a small design space for the minimal sufficient configuration — and, as in iPS, it is the moment the path yields a tool.
The hypothesis, evolving across the path

Doudna's path as one evolving question, deepening across the career and broken once by an exogenous jump, over a continuous base-pairing-guide thread.
The path as one evolving question. It deepens — external template → modular guide → programmable guide — and is broken exactly once, by the exogenous jump into bacterial CRISPR (not by a Gap). The base-pairing-guide thread runs continuous underneath. Unlike the iPS path, which branched at an internal anomaly and rejoined, this one is linear-and-deepening with a single discontinuity.
Read as one continuous question, the path runs:
Can an RNA be engineered to copy a template? (H1–H2) → How does a catalytic RNA fold and work? (H3–H4) → Can an RNA direct and reshape a large machine? (H5) → How is a guide RNA made and measured? (H6) → How does a protein read and display a guide RNA to find a target? (H7–H8) → Can one protein, guided by one programmable RNA, cut any chosen DNA? (H9)
The single thread that survives every turn is specificity delegated to a base-pairing guide — the abstraction that deepens from "external template" (H1) to "programmable sgRNA" (H9). Unlike the iPS hypothesis, which branched at the 1995 gap into a mechanistic and a target question that later rejoined, the Doudna hypothesis is more nearly linear-and-deepening — the same idea re-asked at higher abstraction — with one discontinuity the internal thread cannot explain (below).
Where the framework must stretch: the redirect that is not a Gap
The Architectural Implications cycle assumes the next Hypothesis is generated by the Gap — background knowledge points a returned result toward one new question. The Doudna path contains a turn that this arrow does not cover: the move from RNA/RNAi structure (through H6) into bacterial CRISPR (H7), around 2006–08. There is no Gap in the H5–H6 results that forces bacterial immunity; the corpus shows no such internal anomaly. The redirect came from outside the cycle — an external opportunity (a collaboration/problem carried in from another field).
This is a genuine extension of the framework. Alongside the intra-cycle path Gap → next Hypothesis, a discovery path can receive an exogenous input to the next-Hypothesis step: an opportunity that arrives independent of any returned result. The single-cycle diagram needs a second inbound arrow to the Hypothesis node — one from the Gap (endogenous), one from outside (exogenous). This maps directly onto the Meta-Trace's two turn engines: internal-anomaly turns are the Gap→Hypothesis arrow firing; external-opportunity turns are the exogenous arrow firing. (And the coming traces suggest two more inbound kinds — a constraint-driven push, Karikó; an agenda-setting self-redefinition, Sharpless.)
So the Discovery Trace's most open question — what directs a Gap toward one next Hypothesis? — gains, from the CRISPR case, a prior question: did the next Hypothesis even come from the Gap? Sometimes it comes from the world.
How an exogenous input becomes a hypothesis
Saying a turn "came from outside" is not yet a mechanism. An exogenous input is not a result, so it cannot generate a hypothesis the way a Gap does — a Gap is a returned value that already points somewhere. An opportunity, a constraint, or an agenda specifies only a new problem, target, or goal; on its own it is inert. It becomes a hypothesis only by docking onto the scientist's accumulated repertoire — the transferable capabilities the Transfer Trace records (technique, concepts, reagents, people). The hypothesis is the product of two factors:
Hypothesis = (external opening) × (prepared repertoire).

How an exogenous input becomes a hypothesis: an opening docks onto the scientist's repertoire to yield a testable prediction.
An opening is not a result, so it cannot fire alone. It becomes a hypothesis only at the dock with the prepared repertoire — opening × repertoire → hypothesis. The three exogenous sub-types (opportunity, constraint, agenda) dock by different rules; the instant of recognition keeps an irreducible tacit core.
The opening supplies the new target; the repertoire supplies the means and the framing that make that target testable. For CRISPR, the external opening was an approach around 2006 about a puzzling bacterial repeat system — nothing in Doudna's own results pointed there. What converted it into hypotheses was the repertoire she carried: twenty years of RNA structural biology, the base-pairing-guide concept, the crystallography pipeline, and the guide-RNA-processing framing from Dicer. Their intersection produced the lab's first CRISPR hypotheses in her own terms — e.g. "a Cas protein processes its guide RNA sequence- and structure-specifically, and its mechanism can be read from a structure" (H7). The opportunity chose the target; the repertoire chose the question. This is also why the opening was seized by this scientist and not another: only a repertoire already shaped to convert it could fire the arrow — Pasteur's prepared mind, stated mechanically.
The three exogenous sub-types dock by different rules. An opportunity docks as "apply my method M to the new target X" (CRISPR). A constraint docks as "what change lets X satisfy the imposed requirement C?" — the shape of the modified-nucleoside mRNA turn, where the requirement of non-immunogenicity generated the hypothesis that naturally modified nucleosides evade innate sensing. An agenda docks by deduction from a self-set goal — "if synthesis should be a few near-perfect modular reactions, then this reaction is a candidate" (click chemistry). In each, the exogenous input is the same kind of thing (a non-result opening) and the generative step is the same (docking onto repertoire); only the docking rule differs.
And the honest limit: this specifies the form of the arrow (opening × repertoire → hypothesis) but not the full selection — why this repertoire recognised this opening as tractable, at that moment, lives in conversations and intuition outside the publication record. The corpus shows the dock happened (the CRISPR papers run on the prior structural methods); it cannot show the instant of recognition. Like the Gap → next-Hypothesis step it mirrors, the exogenous arrow is specifiable in structure yet keeps an irreducible tacit core.
Interpretation Design and Experimental Design on this path
- Interpretation Design (choosing the criterion before reading): the Cas9 work fixes, in advance, that in-vitro cleavage of a defined dsDNA will stand for "function," and that the ~20-nt complementary segment is the variable that will stand for "programmability." Those choices determine what the assay can see — and, unlike the iPS Fbx15 proxy (which came apart from the state it stood for), here the proxy (DNA cut) is the target function, which is part of why the readout translated so cleanly into a tool.
- Experimental Design (the downward translation) is unusually visible at H9: the same hypothesis ("Cas9 is programmable") could have been tested many ways; the decision to fuse crRNA+tracrRNA into a single chimera is a design act, not a deduction — the executable plan that the first stage of any automated discovery would have to generate. It is the mirror of the 1991 design act (splitting one RNA into three), twenty-one years apart.
What the CRISPR case adds to the Architectural Implications framework
- Coupled operation pairs. For structure-driven discovery the dry and wet pairs are not alternatives but a bound wet+dry unit (crystal→map; particles→reconstruction). The two primitives stand; their coupling is a mode the iPS path did not foreground.
- An exogenous inbound arrow to the Hypothesis. The decisive redirect came from outside the cycle (opportunity), not from a Gap. The framework needs a second arrow into the next-Hypothesis step — which is exactly the Meta-Trace's turn typology seen from inside the cycle.
- Confirmation-heavy Gap distribution is not a weakness. A structure-first path spends many turns on confirming/consolidating Gaps; that is how a technique lineage (the path's real handoff) is built. Gap kind correlates with handoff medium: anomaly-rich paths tend to hand forward reagents; confirmation-rich paths tend to hand forward methods.
- The compression turn recurs. Both traced paths end in an optimization/reduction (24→4; two RNAs→one). "Compression marks invention" is visible at the cycle level as a distinct optimization turn, not a verification one.
Two paths are not a theory. But where iPS gave the cycle its shape, CRISPR gives it its first stress test — and tells us precisely which arrows to add before the next path is read.
Companion to Machine Traces of Discovery Paths #3 (https://www.thediscoveryengine.ai/machine-traces-of-discovery-paths-3-the-path-to-crispr-cas9-jennifer-doudna). Reads the Doudna → CRISPR-Cas9 nodes (H1–H9; see the #3 Lab Notebook https://thediscoveryengine.ai/machine-traces-of-discovery-paths-3-lab-notebook) through the two-level Discovery Trace of the first Architectural Implications entry. Terminology and the Gap taxonomy follow that entry.
Hiroaki Kitano
How to cite
Kitano, H. (2026). The Discovery Trace — Architectural Implications from CRISPR-Cas9. The Discovery Engine.
https://thediscoveryengine.ai/the-discovery-trace-architectural-implications-from-crispr-cas9/
Hiroaki Kitano ORCID: 0000-0002-3589-1953
Publish: 16 August 2026