Connecting Distant Dots

Share
Connecting Distant Dots
Connecting Distant Dots: What appears distant may be adjacent in structure.

The third essay in the Discovery Engine series.

Discovery often comes from connecting distant dots — observations, concepts, and fields that no one had thought to join. This essay looks closely at three such connections — Darwin, CRISPR, and iPS — and asks how each was actually made, and whether a differently built system could make them again.


1. Discovery despite the limits

The previous essay set out the limits of human-led science — the information horizon, the dimensionality limit, the coarse representations, the career-shaped concentration, the siloing, the two long tails. And yet, working under exactly those limits, human scientists have made the discoveries that define modern biology and medicine. That is the puzzle this essay takes up: not why human-led discovery falls short, but how the landmark discoveries happened in spite of the limits — and whether a differently built system could make them happen again.

The way to get traction is to look closely at real cases. This essay takes three — Darwin's theory of natural selection, the genome-editing tool built from CRISPR, and the reprogramming of adult cells into stem cells (iPS) — and asks of each the same two questions: how was the connection actually made, and could an autonomous system have made it, or where would it get stuck? The answers differ from case to case, and the differences are the point.

2. Darwin: a principle carried across a field

By the late 1830s Darwin was already convinced that species change — the Beagle's biogeography, the work of animal breeders, and the distribution of forms across islands had pressed the conclusion on him. What he lacked was a mechanism: some force that would do in nature what a breeder does by hand, preserving certain variations and discarding others. In October 1838 he read, for amusement, Thomas Malthus's essay on human population — an argument from political economy that populations grow faster than the resources that feed them, so that struggle and death are the inevitable check. By Darwin's own account, the mechanism arrived in that moment (Darwin, 1958): if more individuals are born than can survive, favourable variations would tend to be preserved and unfavourable ones destroyed, and the result would be new species. He had carried a principle from one field into another that shared none of its vocabulary, and in doing so turned a pile of observations into a theory (Figure 1).

Could a machine have made that move? The materials were public — Malthus in one literature, Darwin's observations in another — and a system that read widely enough would have had both in view. But having both in view is not the same as seeing that one answers the other. Nothing in an essay on the economics of human hunger announces that it holds the missing mechanism for the diversity of life; the connection is stated in neither domain, and there is no shared term to follow from one to the other.

That recognition is the bottleneck. It is not retrieval — the link existed nowhere to be retrieved — and it is not the filling-in of a gap within a settled field. It is the perception of a correspondence between domains that no one had thought to compare. A system can be handed both texts and still not see it; seeing it is the discovery.

Path-to-Theory-of-Evolution.jpg

Figure 1. The path to natural selection. Cross-domain exploration — the Beagle's observations, breeders' artificial selection, and Malthus's political economy — converging on a substrate-independent mechanism (variation + inheritance + limited resources → selection), followed by roughly two decades of systematic evidence-gathering and theory-building.

3. CRISPR: a defense mechanism seen as a tool

The making of CRISPR genome editing is a relay that ran for a quarter of a century across fields that had no reason to talk to each other (Figure 2). In 1987, while sequencing an ordinary E. coli gene — one governing how a metabolic enzyme is processed — Yoshizumi Ishino and colleagues noticed a strange run of repeated sequences just past it: regular, evenly interrupted, unlike anything then known. They had no way to pursue it; the paper states plainly that the sequences matched nothing elsewhere in prokaryotes and that their biological significance was unknown (Ishino et al., 1987). The observation sat almost uncited for nearly two decades. In 2005, three groups working separately found that the spacers between the repeats had been captured from viruses and plasmids, making the array a genetic record of past infections (Mojica et al., 2005; Bolotin et al., 2005; Pourcel et al., 2005); two of them — Mojica's and Bolotin's — went further, proposing what the record was for: a bacterial immune system. The hypothesis was confirmed in, of all places, the dairy industry, by a food-cultures company whose business depended on keeping yogurt and cheese starters alive against bacteriophage (Barrangou et al., 2007). The last ingredient arrived from a study of Streptococcus pyogenes, a human pathogen, where Emmanuelle Charpentier found an unpredicted small RNA — tracrRNA — that proved essential to the machinery (Deltcheva et al., 2011). Then Charpentier and Doudna saw that the whole apparatus — a protein steered to a chosen DNA sequence by a pair of RNA guides — was a programmable cutting tool: change the guide, change the target, in any genome at all (Jinek et al., 2012).

What followed the recognition was engineering, and it was fast. Jinek's group had already fused the two natural guide RNAs into a single molecule; within months, codon-optimized versions of Cas9 carrying nuclear-localization signals were made to cut precisely in human and mouse cells, several targets at a time (Cong et al., 2013; Mali et al., 2013), and the optimization has not stopped since — better guides, fewer off-target cuts, a wider range of targetable sites. That half of the story is systematic search and refinement, the kind of work that scales.

The other half does not. The striking thing is that by 2011 every piece was in the public literature — the repeats, the immune function, the molecular components — and the leap still had to be made. A system reading the microbiology exhaustively would have found everything the field knew and still not found the genome-editing tool, because the editing reading of the mechanism was nowhere in the microbiology to be read. The recognition that the pieces composed a tool was an act performed on top of them, not a fact retrievable among them.

Two bottlenecks stand out. The first is the cross-field recognition itself — seeing a bacterial defense as an engineering primitive, a correspondence between domains that shared no surface. The second is quieter: the relay turned on unplanned encounters — an accident in a sequencing project, an immune system found in a cheese culture, a small RNA found in a pathogen — none of which a program aimed at "build a genome editor" would have undertaken. The discovery depended on decades of work that was not looking for it.

Path-to-CRISPR-Genome_Editing.jpg

Figure 2. The path to CRISPR genome editing. A quarter-century relay across unrelated fields — an E. coli sequencing accident, the 2005 foreign-origin and immune hypotheses, a dairy-industry proof, and tracrRNA found in a human pathogen — resolving into the recognition of programmable targeting, after which engineering and optimization turned the mechanism into a tool.

The path to induced pluripotent stem cells did not begin as a question about cell identity. It began in cardiovascular medicine. Shinya Yamanaka's early work concerned APOBEC-1, an enzyme that edits messenger RNA, studied as a way into cholesterol metabolism. Engineered to overexpress it in the liver, his transgenic animals developed something no one had been studying for: liver cancer (Yamanaka et al., 1995). Asking why — which messenger RNAs the enzyme had wrongly edited in those tumors — led to a previously unknown gene, NAT1, a fundamental repressor of protein translation (Yamanaka et al., 1997). And NAT1 proved essential for something remote from both cholesterol and cancer: deleted from embryonic stem cells, it left them unable to differentiate — locked in the embryonic state (Yamanaka et al., 2000). A question that began in lipid metabolism had arrived, by way of an unsought tumor, at the molecular control of cellular identity itself (Figure 3).

What followed had two parts of very different character. The first was exploration — open-ended, curiosity-driven, and shaped by chance. Working forward from the unsought clue about NAT1, Yamanaka's program kept asking what holds a cell in its identity, until it reached a usable idea: that a small set of factors keeps a mature cell as it is, and might, if overridden, return it to the embryonic state. The second was search and optimization — systematic and goal-directed: twenty-four candidate genes, drawn from the FANTOM database of transcripts characteristically active in embryonic stem cells, were introduced together; then, by withdrawing them one at a time — a leave-one-out screen — the set was narrowed to the four now called the Yamanaka factors (Takahashi & Yamanaka, 2006).

The two answer the reproducibility question differently. The search is exactly what an autonomous system would do well — a defined candidate set, a clear assay, tireless elimination — and the history says as much: the twenty-four-gene experiment was only a pilot to validate the selection system, and the team's actual plan, had it failed, was unbiased large-scale screening of cDNA libraries (Yamanaka, 2026), exhaustive search of the kind a machine runs without tiring. But the search existed only because the exploration came first — and exploration, in the sense that mattered here, is not a search at all. Nothing in a cholesterol program contains the goal reprogram a differentiated cell; the destination was reached because an unwanted result — a tumor where a lipid effect was expected — was not discarded as failure, and because the mind that met it was free to follow the anomaly out of its own project and recognize, in a fact about a translational repressor, a fact about cell identity.

The bottleneck here has a particular character: serendipity in the precise sense — the trained recognition of meaning in what one was not looking for. A goal-directed search cannot reach it by construction, because the goal that mattered was not the one being pursued. This does not make it unreachable forever; it reframes what a discovering system would need — not only an optimizer driving toward a fixed goal, but an explorer wide enough, and unconstrained enough by prior notions of significance, to turn up unsought results and register when one of them matters. That capacity, again, was the iPS team's own reserve plan.

The Path to iPS Cells
Path-to-iPS_Cell.jpg

Figure 3. The path to induced pluripotent stem cells. An unsought result — a tumor where a lipid effect was expected — and the open-ended exploration of cell-identity control, followed by the systematic search and optimization (FANTOM candidate genes, a leave-one-out screen) that isolated the four Yamanaka factors.

5. The underlying dynamics of discoveries

Across all three cases the same pattern appears, and it is sharper than "connect two dots." Discovery begins when a deeper structure is recognized beneath domains that look unrelated; only afterward does the work become systematic search and refinement. What Darwin took from Malthus was not an analogy but a substrate-independent logic — wherever there is variation, inheritance, and limited resources, selection follows, in an economy or an ecology alike. What Charpentier and Doudna saw was not bacterial immunity but programmable, information-guided intervention. What Yamanaka's path uncovered was not a fact about stem cells but a principle of state transitions in a regulatory network. In each, the move that mattered was the recognition that two distant things are instances of one more abstract pattern.

This reframes what the "distance" between the domains really is. They look unrelated, but the separation is often an artifact of how human knowledge is filed — into departments, journals, vocabularies, and careers — rather than a fact about the world; our partitions are simply coarser than the structure of reality. Population pressure does govern finite biological populations; a bacterial nuclease is a programmable cutting tool; a translational repressor does sit in the machinery of cell identity. Sometimes the correspondence was there all along, hidden only by fields that do not read each other; sometimes it had never been conceptualized at all. Either way, the boundary that had to be crossed was one we had drawn, not one nature had — the structural misalignment of the previous essay seen from the other side: the boundaries on the map are not the seams of the world. The real seams are the deeper patterns that run beneath the fields, the ones of which the separate domains are only instances.

That reframing changes the reproducibility question. The second phase — search and optimization — is exactly what a tireless, unbounded system does well, and in the iPS case the discoverers had planned to run the exhaustive version themselves. The first phase is the open one, and it cuts two ways. A system that does not inherit the human partition of knowledge — that reads across every field at once, unbound by department or career — is in principle better placed than any specialist to see that two distant literatures describe the same structure. But the number of possible pairings is astronomical and almost all are empty, so the difficulty is not having both domains in view; it is detecting the rare correspondence that matters — and in two of our three cases that detection turned on an encounter, an accident in a sequencing run or an unwanted tumor, that no goal-directed program would have arranged. Whether a machine can find such structure reliably, and not by luck, is genuinely unsettled.

What this implies for machine discovery is both humbling and encouraging. Connecting distant dots takes many forms — a new cellular function, an intervention, a reframing, an untried combination — and some may follow patterns a system could learn. But the domains nature entails are many, and the pairings between them grow combinatorially: far too many to try, almost all empty. The way through is unlikely to be enumerating pairings; it is to change the partition itself — to stop carving knowledge into the fields humans happen to have built and re-cluster it around the deeper patterns the cases point to, so that things now treated as distant fall under one structure. That is what tames the combinatorial blow-up: once two distant things sit under a single pattern, the space of pairings collapses, and the search and optimization that follow are exactly what a machine does well.

Put the two together and the subject of the next essay comes into view: a system that does not merely re-cluster what we already know but can generate new dimensions in which to describe reality — and so carry us to regions of knowledge we have no path to today. Building it would do more than produce discoveries. It would force into the open what discovery actually is — a question whose answer, in the end, does not turn on whether the discoverer is a human or a machine.


AI Co-Authorship Disclosure

This essay was written by Hiroaki Kitano in collaboration with Claude (Anthropic, primary), Gemini (Google, secondary), and ChatGPT (OpenAI). Final editorial judgment is mine. The process layer is documented in the companion Lab Notebook. Figures were generated with ChatGPT, from initial prompts drafted by Claude and under detailed direction from the author, with cross-checking by Claude and Gemini.


References

  • Barrangou, R., Fremaux, C., Deveau, H., Richards, M., Boyaval, P., Moineau, S., Romero, D. A., & Horvath, P. (2007). CRISPR provides acquired resistance against viruses in prokaryotes. Science, 315(5819), 1709–1712.
  • Bolotin, A., Quinquis, B., Sorokin, A., & Ehrlich, S. D. (2005). Clustered regularly interspaced short palindrome repeats (CRISPRs) have spacers of extrachromosomal origin. Microbiology, 151(8), 2551–2561.
  • Cong, L., Ran, F. A., Cox, D., Lin, S., Barretto, R., Habib, N., Hsu, P. D., Wu, X., Jiang, W., Marraffini, L. A., & Zhang, F. (2013). Multiplex genome engineering using CRISPR/Cas systems. Science, 339(6121), 819–823.
  • Darwin, C. (1958). The Autobiography of Charles Darwin, 1809–1882 (N. Barlow, Ed.). Collins. (Original work written 1876.)
  • Deltcheva, E., Chylinski, K., Sharma, C. M., Gonzales, K., Chao, Y., Pirzada, Z. A., Eckert, M. R., Vogel, J., & Charpentier, E. (2011). CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. Nature, 471(7340), 602–607.
  • Ishino, Y., Shinagawa, H., Makino, K., Amemura, M., & Nakata, A. (1987). Nucleotide sequence of the iap gene, responsible for alkaline phosphatase isozyme conversion in Escherichia coli, and identification of the gene product. Journal of Bacteriology, 169(12), 5429–5433.
  • Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, J. A., & Charpentier, E. (2012). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, 337(6096), 816–821.
  • Mali, P., Yang, L., Esvelt, K. M., Aach, J., Guell, M., DiCarlo, J. E., Norville, J. E., & Church, G. M. (2013). RNA-guided human genome engineering via Cas9. Science, 339(6121), 823–826.
  • Mojica, F. J. M., Díez-Villaseñor, C., García-Martínez, J., & Soria, E. (2005). Intervening sequences of regularly spaced prokaryotic repeats derive from foreign genetic elements. Journal of Molecular Evolution, 60(2), 174–182.
  • Pourcel, C., Salvignol, G., & Vergnaud, G. (2005). CRISPR elements in Yersinia pestis acquire new repeats by preferential uptake of bacteriophage DNA, and provide additional tools for evolutionary studies. Microbiology, 151(3), 653–663.
  • Takahashi, K., & Yamanaka, S. (2006). Induction of pluripotent stem cells from mouse embryonic and adult fibroblast cultures by defined factors. Cell, 126(4), 663–676.
  • Yamanaka, S., Balestra, M. E., Ferrell, L. D., Fan, J., Arnold, K. S., Taylor, S., Taylor, J. M., & Innerarity, T. L. (1995). Apolipoprotein B mRNA-editing protein induces hepatocellular carcinoma and dysplasia in transgenic animals. Proceedings of the National Academy of Sciences, 92(18), 8483–8487.
  • Yamanaka, S., Poksay, K. S., Arnold, K. S., & Innerarity, T. L. (1997). A novel translational repressor mRNA is edited extensively in livers containing tumors caused by the transgene expression of the apoB mRNA-editing enzyme. Genes & Development, 11(3), 321–333.
  • Yamanaka, S., Zhang, X.-Y., Maeda, M., Miura, K., Wang, S., Farese, R. V., Jr., Iwao, H., & Innerarity, T. L. (2000). Essential role of NAT1/p97/DAP5 in embryonic differentiation and the retinoic acid pathway. The EMBO Journal, 19(20), 5533–5541.
  • Yamanaka, S. (2026). Two decades of induced pluripotent stem cell research: From discovery to diverse applications. Cell Stem Cell, 33(3). https://doi.org/10.1016/j.stem.2026.02.003

Read more

科学的発見における人間の認知的・行動的限界

科学的発見における人間の認知的・行動的限界

本シリーズの第一論考は、機械は発見できるかと問うた。その問いに答えるには、それに先立つ別の問いを正面から引き受けなければならない。すなわち、人間が主導する発見は、本当のところどれほどのものなのか。本稿は、現在の状況がどこで限界に達しているかを、人間の認知構造と社会構造という「構造」の問題として論じる。 1. 人間の発見が限界に達するところ 本シリーズの第一論考は、ある厳しい観察で締めくくられた。今日、生み出される知識の量は、どんな個人の知性も、どんな制度も追いつけないほど膨大になっている。また、人間の認知構造の限界、研究を行う社会的環境からの研究テーマの選択に関する影響などいろいろな問題が、人間による科学研究の限界の要因になっている。その帰結が構造的ミスアライメントである。これは個々の科学者の失敗ではなく、人間の発見というものが全体としてどう組織化されているかの不整合なのである。構造的ミスアライメントとは、人間の認知と行動の構造と、それが捉えようとしている自然の構造との、相互の食い違いを指す。一人の科学者が複雑な現実を丸ごと受け止めることはできない。その人が研究するものは、

By Hiroaki Kitano