Machine Traces of Discovery Paths #1 - The Path to iPS Discovery, A Seed- and Abstract-Based Reconstruction
Machine Traces of Discovery Paths — Autonomous reconstructions of how landmark discoveries were reached, traced by AI from the primary literature.
Disclaimer — Discovery Trace Series. The Discovery Trace series is an experimental series of articles attempting the semi- or fully-automatic reconstruction of the paths to major discoveries. Multiple AI tools have been deployed to reconstruct each "Discovery Path." To evaluate the progress of the technology, an initial set of articles is published prior to rigorous verification, and may therefore contain errors and omissions. Content will be updated as corrections are made, with notes on publication history and a correction log.
This is the first entry in a series in which an AI system, rather than a human author, reconstructs the path to a landmark discovery directly from the primary literature. It is not a new telling of how induced pluripotent stem (iPS) cells were discovered; the interpretive analysis of that path — what its transitions mean, and which of them a machine could perform — is the subject of the companion essays Connecting Distant Dots (Post #3) and Can a Machine Connect Distant Dots? (Post #4). What this entry documents is different and narrower: an AI system was pointed at the literature and asked to recover the path on its own, and this is the record of what it produced, how, and where it stopped.
What was done
An AI research system (Claude Science) was given one objective — trace Shinya Yamanaka's path to the 2006 discovery of iPS cells — and instructed to build the answer only from the primary literature. It assembled 275 disambiguated PubMed records (1990–2026), parsed 1,982 reference entries from his pre-2008 papers, and retrieved Europe PMC citation counts. From that corpus it reconstructed a causal chain, identified the prior work the chain depended on, extracted the recurring moves within it, and — for one hinge experiment — reconstructed the bench protocol from the paper's Methods. Every result below is traceable to the code and queries that produced it.
The value of this entry is not the chain itself — that story is told in Posts #3 and #4 — but four things those essays do not contain: the fact that the reconstruction was performed autonomously by a machine; the provenance and corpus behind it; a set of named, reusable moves the system extracted from the record; and a measurement, from one hinge paper, of the gap between a published protocol and an executable one.
Method
The task decomposed into five operations, each executed against public databases and logged with its output.
- Name disambiguation. PubMed returns 1,038 records for "Yamanaka S." Several researchers share the name, including a Kyoto Prefectural cardiologist, a Gunma pathologist, a Tohoku thoracic surgeon, and an Osaka City University materials physicist. The system filtered by per-author affiliation strings, iterative co-author networks, and topic. 275 records survived. One intermediate step failed — a PubMed record returned a null field, raising a TypeError — and was recovered by re-parsing affiliation per author rather than per paper.
- Corpus assembly. The 275 records were ordered chronologically and grouped by institution and topic.
- Reference parsing. Reference lists were extracted for 53 of 69 pre-2008 papers, limited by Europe PMC coverage; older J Biol Chem and Japanese-language items were the main gaps.
- Citation retrieval. Europe PMC citation counts were attached to each paper.
- Pattern extraction. Recurring methodological moves were identified from the sequence of papers and their stated results.
The chain the machine reconstructed
The system recovered a continuous chain from clinical pharmacology to iPS cells, with four documented hinge points. The narrative below is the machine's output, not a fresh retelling; Posts #3 and #4 analyse what these transitions mean.

From a failed cholesterol experiment to induced pluripotency: Yamanaka's twelve-year chain
Phase 1 — Osaka City University, 1990–1993. Nine papers on canine renal and vascular pharmacology (endothelin, nitric oxide synthase inhibition, thromboxane), median citations ~20. No content anticipates pluripotency; the relevant feature is a phenotype-first, whole-animal orientation.
Phase 2 — Gladstone / UCSF, 1994–1997. Working on apolipoprotein B mRNA editing, the 1995 PNAS transgenic experiment was built to lower LDL cholesterol. It did, but the animals unexpectedly developed liver dysplasia and hepatocellular carcinoma (Yamanaka, S. et al., 1995). The 1996 J Biol Chem paper resolved the mechanism as off-target hyperediting (Yamanaka, S. et al., 1996); the 1997 Genes Dev paper identified NAT1, a translational repressor, via differential display on the tumours (Yamanaka, S. et al., 1997) (hinge points 1–2).
Phase 3 — Osaka, 1997–1999. Mostly middle-author positions on another lab's cardiovascular work, with two first-author exceptions. The authorship pattern is consistent with a documented difficult interval.
Phase 4 — NAIST, 2000–2003. The 2000 EMBO J NAT1 knockout (Yamanaka, S. et al., 2000) found that null ES cells were normal except that they could not differentiate (hinge point 3). Three 2003 papers established the platform: the Fbx15 paper in Mol Cell Biol (Tokuzawa, Y. et al., 2003), the ERas paper in Nature (Takahashi, K. et al., 2003), and the Nanog paper in Cell (Mitsui, K. et al., 2003). The common method was to identify genes expressed specifically in ES cells and preimplantation embryos and test them individually — the ECAT catalogue.
Phase 5 — 2006. Takahashi & Yamanaka introduced 24 ECAT candidates into fibroblasts carrying the Fbx15 selection cassette and reduced them to four factors: Oct3/4, Sox2, c-Myc, Klf4 (Takahashi, K. & Yamanaka, S., 2006) (hinge point 4). Each component traces to a prior phase: Fbx15 selection to 2003, the candidate list to the ECAT catalogue, ES-cell culture competence to the NAT1 knockout.
Phase 6 — 2007 onward. Subsequent papers removed successive limitations: Nanog selection for germline competence (Okita, K. et al., 2007), human iPS cells (Takahashi, K. et al., 2007), and derivations without c-Myc (Nakagawa, M. et al., 2008), without viral vectors (Okita, K. et al., 2008), and integration-free (Okita, K. et al., 2011).
In the citation record, the 2006 and 2007 reprogramming papers stand orders of magnitude above everything else in the corpus.

The 2006–07 reprogramming papers stand orders of magnitude above everything else in the record
The prior work the chain depended on
Parsing the reference lists identifies the external work cited most often across his pre-2008 papers (count = number of his papers citing it).
| n | Work | Role in the chain |
|---|---|---|
| 12 | Evans & Kaufman 1981; Martin 1981 | Existence of mouse ES cells |
| 12 | Thomson et al. 1998 | Human ES cells |
| 11 | Nichols et al. 1998; Niwa et al. 2000 | Oct3/4 as dose-dependent pluripotency determinant |
| 9 | Chambers et al. 2003 | Nanog by expression cloning |
| 8 | Avilion et al. 2003 | Sox2 in early lineages |
| 7 | Morita et al. 2000 | Plat-E retroviral packaging (delivery of 24 factors) |
| 6 | Yuan et al. 1995 | Oct/Sox synergy (cis-logic behind Fbx15) |
| 5 | Cowan et al. 2005 | ES fusion reprograms somatic nuclei |
The moves the machine extracted
Beyond the chain, the system identified four moves that recur across it. Each is documented at a specific point in the record rather than inferred generally. These named moves are the analytic output of this trace, and the part that is portable to other cases.
- Anomaly promotion — an off-hypothesis result is made the next project, rather than explained away as an artifact. The 1995 liver tumours and the 2000 differentiation block were both off-hypothesis results that became the next project.
- Productive failure — an experiment fails at its stated goal but yields a result worth chasing. The 1995 APOBEC-1 experiment succeeded at lowering cholesterol yet was pursued for the tumours it also produced.
- Negative result as reagent — a finding that turns out to be "dispensable" is retained as a usable tool rather than discarded. Fbx15, dispensable and functionless, became a selection marker; Nanog, reported dispensable in the 2006 screen, later became the stringent marker that yielded germline-competent iPS cells.
- Question inversion — when a direct question saturates, its inverse is asked. The shift from "what maintains pluripotency" to "what can install it" turned a catalogue of dispensable genes into the 24-factor candidate list.
These four are the set observed in this particular path. They are not proposed as necessary conditions for discovery in general. A different discovery may exhibit a different set of moves; this single case does not determine which moves generalise. The working expectation is that discovery paths have structure of this kind, but that the structure is varied rather than uniform — a claim that only additional traces can test, which is the purpose of this series.
The protocol layer
The trace so far recovers the shape of the discovery — who did what, when, and on whose prior work. It does not recover how each experiment was actually performed. As a first, proof-of-concept attempt at reconstructing this layer, the system was given the full text of a single hinge paper — the 2000 EMBO J NAT1 knockout that turned the work toward stem-cell biology — and asked to extract its protocol from the Methods alone, marking every element the paper omits rather than filling it in. This is the first experimental attempt at the protocol layer; the same protocol extraction applied to the other hinge papers is planned for later entries in the series.
Two things came out of this, and the second matters more than the first.
The first is that the protocol is largely extractable. The system recovered the targeting construct — the 1.2 kb and 5 kb homology arms, the pgk-neo and pgk-tk selection cassettes, the SalI linearization — the genotyping primers and their diagnostic band sizes, the reagents and their stated concentrations, and the procedure step by step: the electroporation, the high-G418 selection that yielded 29 of 64 null colonies, the retinoic-acid differentiation assays, the teratoma histology. Where a step was cited to an earlier paper rather than restated, the system flagged the reference rather than inventing the method.
The second is the finding that bears on automation: the published paper does not contain an executable protocol. What it contains is a protocol addressed to an expert reader, who is expected to supply the rest from tacit knowledge, and the omissions are not incidental. The composition of the ES-cell medium — base medium, serum, supplements — is never given. The feeder-withdrawal assay that produced the founding observation, that null cells behaved abnormally off feeders, has no stated duration, no seeding density, and no scoring rubric. The retinoic-acid concentration is stated for two experiments and left unstated for two others, at two different values. Reading critically, the system also flagged a printed LIF concentration as almost certainly a units error, and noted that the central claim stands without a genetic rescue, without karyotype analysis, and without any statistical test.
To make the experiment runnable, those omissions have to be filled, and the reconstruction fills them under an explicit tag: [P] for values stated in the paper, [S] for values supplied from period-standard practice so the protocol runs, and [M] for modern additions the original lacked. The tags are the point. The [S] and [M] layer is a direct measure of the distance between what the literature records and what the experiment actually requires — the tacit layer that expert practitioners carry and papers omit.
This is a second frontier, alongside the one named in Can a Machine Connect Distant Dots? (Post #4). That essay located the hard step of discovery in the forward generation of a redirecting framing. The protocol layer locates a different gap, on the execution side: automating the second stage — an experiment system carrying out AI-generated protocols — requires generating the tacit layer that the published record does not contain, not merely retrieving protocols from it. The full extraction and the tagged bench reconstruction for this one paper are in the Evidence documents (Machine Traces of Discovery Paths #1 — Lab Notebook https://thediscoveryengine.ai/machine-traces-of-discovery-paths-1-lab-notebook/ ), which make the [P]/[S]/[M] boundary concrete for a single hinge experiment.
Limits of the reconstruction
The reconstruction is explicit about its boundaries, and the boundaries are as informative as the results.
- It recovers what, not why. The record shows the sequence of experiments and results. It does not contain the reasoning behind each pivot. The moves listed above are patterns observed in the record, not evidence of intent.
- Coverage is partial. Reference lists were available for 53 of 69 pre-2008 papers. Book chapters, Japanese-language reviews not indexed in PubMed, and the Nobel lecture are outside the corpus. The protocol layer was reconstructed for one hinge paper, not the whole chain.
- Disambiguation is probabilistic. 275 of 1,038 records were retained on affiliation and co-author evidence; misclassification at the margins is possible.
What this trace does and does not show
This entry demonstrates two things concretely. First, the retrospective reconstruction of a discovery's path — assembling the corpus, recovering the causal chain, and extracting recurring moves — can be performed autonomously by a machine, with provenance, for a known discovery. Second, the protocol layer of one hinge experiment can be extracted from a paper and the gap between that published protocol and an executable one can be measured, item by item, as an explicit [P]/[S]/[M] boundary.
It does not cross either of two frontiers. On the discovery side, it does not perform the forward act that Post #4 isolates as the open question: generating the redirecting framing for a discovery not yet made. On the execution side, it does not generate the tacit protocol layer that turning a published method into a runnable one requires; it only measures that layer's size for a single case. Both frontiers are what the series is meant to probe: running the same procedure over further discoveries — CRISPR and others, in later entries — will test how general, and how varied, the structure of these paths and their protocol gaps turn out to be.
Data and evidence
The reconstruction is based on three datasets, to be released in the project repository: the full disambiguated corpus of 275 records (1990–2026); the most-cited prior work; and the key papers annotated by phase and hinge point. The figures are generated from these datasets. The full trace report, the analysis findings, and — for the NAT1 (2000) hinge experiment — the structured protocol extraction and the tagged bench reconstruction are published as the Evidence documents for this entry in the Lab Notebook; each result there is traceable to the query and code that produced it, and each protocol value carries its provenance.
References
Cited in author–year form; listed alphabetically by first author, then year. Bibliographic fields (authors, volume, issue, pages, DOI) retrieved from PubMed E-utilities and reconciled against the stored corpus.
Primary papers named in the text (Yamanaka's own work)
(Mitsui, K. et al., 2003) Mitsui K, Tokuzawa Y, Itoh H, Segawa K, Murakami M, Takahashi K, et al. The homeoprotein Nanog is required for maintenance of pluripotency in mouse epiblast and ES cells. Cell. 2003;113(5):631-42. doi:10.1016/s0092-8674(03)00393-3. PMID: 12787504.
(Nakagawa, M. et al., 2008) Nakagawa M, Koyanagi M, Tanabe K, Takahashi K, Ichisaka T, Aoi T, et al. Generation of induced pluripotent stem cells without Myc from mouse and human fibroblasts. Nat Biotechnol. 2008;26(1):101-6. doi:10.1038/nbt1374. PMID: 18059259.
(Okita, K. et al., 2007) Okita K, Ichisaka T, Yamanaka S. Generation of germline-competent induced pluripotent stem cells. Nature. 2007;448(7151):313-7. doi:10.1038/nature05934. PMID: 17554338.
(Okita, K. et al., 2008) Okita K, Nakagawa M, Hyenjong H, Ichisaka T, Yamanaka S. Generation of mouse induced pluripotent stem cells without viral vectors. Science. 2008;322(5903):949-53. doi:10.1126/science.1164270. PMID: 18845712.
(Okita, K. et al., 2011) Okita K, Matsumura Y, Sato Y, Okada A, Morizane A, Okamoto S, et al. A more efficient method to generate integration-free human iPS cells. Nat Methods. 2011;8(5):409-12. doi:10.1038/nmeth.1591. PMID: 21460823.
(Takahashi, K. et al., 2003) Takahashi K, Mitsui K, Yamanaka S. Role of ERas in promoting tumour-like properties in mouse embryonic stem cells. Nature. 2003;423(6939):541-5. doi:10.1038/nature01646. PMID: 12774123.
(Takahashi, K. & Yamanaka, S., 2006) Takahashi K, Yamanaka S. Induction of pluripotent stem cells from mouse embryonic and adult fibroblast cultures by defined factors. Cell. 2006;126(4):663-76. doi:10.1016/j.cell.2006.07.024. PMID: 16904174.
(Takahashi, K. et al., 2007) Takahashi K, Tanabe K, Ohnuki M, Narita M, Ichisaka T, Tomoda K, et al. Induction of pluripotent stem cells from adult human fibroblasts by defined factors. Cell. 2007;131(5):861-72. doi:10.1016/j.cell.2007.11.019. PMID: 18035408.
(Tokuzawa, Y. et al., 2003) Tokuzawa Y, Kaiho E, Maruyama M, Takahashi K, Mitsui K, Maeda M, et al. Fbx15 is a novel target of Oct3/4 but is dispensable for embryonic stem cell self-renewal and mouse development. Mol Cell Biol. 2003;23(8):2699-708. doi:10.1128/MCB.23.8.2699-2708.2003. PMID: 12665572.
(Yamanaka, S. et al., 1995) Yamanaka S, Balestra ME, Ferrell LD, Fan J, Arnold KS, Taylor S, et al. Apolipoprotein B mRNA-editing protein induces hepatocellular carcinoma and dysplasia in transgenic animals. Proc Natl Acad Sci U S A. 1995;92(18):8483-7. doi:10.1073/pnas.92.18.8483. PMID: 7667315.
(Yamanaka, S. et al., 1996) Yamanaka S, Poksay KS, Driscoll DM, Innerarity TL. Hyperediting of multiple cytidines of apolipoprotein B mRNA by APOBEC-1 requires auxiliary protein(s) but not a mooring sequence motif. J Biol Chem. 1996;271(19):11506-10. doi:10.1074/jbc.271.19.11506. PMID: 8626710.
(Yamanaka, S. et al., 1997) Yamanaka S, Poksay KS, Arnold KS, Innerarity TL. A novel translational repressor mRNA is edited extensively in livers containing tumors caused by the transgene expression of the apoB mRNA-editing enzyme. Genes Dev. 1997;11(3):321-33. doi:10.1101/gad.11.3.321. PMID: 9030685.
(Yamanaka, S. et al., 2000) Yamanaka S, Zhang XY, Maeda M, Miura K, Wang S, Farese RV Jr, et al. Essential role of NAT1/p97/DAP5 in embryonic differentiation and the retinoic acid pathway. EMBO J. 2000;19(20):5533-41. doi:10.1093/emboj/19.20.5533. PMID: 11032820.
Prior work (most-cited, from reference-list analysis)
(Avilion, A. A. et al., 2003) Avilion AA, Nicolis SK, Pevny LH, Perez L, Vivian N, Lovell-Badge R. Multipotent cell lineages in early mouse development depend on SOX2 function. Genes Dev. 2003;17(1):126-40. doi:10.1101/gad.224503. PMID: 12514105.
(Chambers, I. et al., 2003) Chambers I, Colby D, Robertson M, Nichols J, Lee S, Tweedie S, et al. Functional expression cloning of Nanog, a pluripotency sustaining factor in embryonic stem cells. Cell. 2003;113(5):643-55. doi:10.1016/s0092-8674(03)00392-1. PMID: 12787505.
(Cowan, C. A. et al., 2005) Cowan CA, Atienza J, Melton DA, Eggan K. Nuclear reprogramming of somatic cells after fusion with human embryonic stem cells. Science. 2005;309(5739):1369-73. doi:10.1126/science.1116447. PMID: 16123299.
(Evans, M. J. & Kaufman, M. H., 1981) Evans MJ, Kaufman MH. Establishment in culture of pluripotential cells from mouse embryos. Nature. 1981;292(5819):154-6. doi:10.1038/292154a0. PMID: 7242681.
(Martin, G. R., 1981) Martin GR. Isolation of a pluripotent cell line from early mouse embryos cultured in medium conditioned by teratocarcinoma stem cells. Proc Natl Acad Sci U S A. 1981;78(12):7634-8. doi:10.1073/pnas.78.12.7634. PMID: 6950406.
(Morita, S. et al., 2000) Morita S, Kojima T, Kitamura T. Plat-E: an efficient and stable system for transient packaging of retroviruses. Gene Ther. 2000;7(12):1063-6. doi:10.1038/sj.gt.3301206. PMID: 10871756.
(Nichols, J. et al., 1998) Nichols J, Zevnik B, Anastassiadis K, Niwa H, Klewe-Nebenius D, Chambers I, et al. Formation of pluripotent stem cells in the mammalian embryo depends on the POU transcription factor Oct4. Cell. 1998;95(3):379-91. doi:10.1016/s0092-8674(00)81769-9. PMID: 9814708.
(Niwa, H. et al., 2000) Niwa H, Miyazaki J, Smith AG. Quantitative expression of Oct-3/4 defines differentiation, dedifferentiation or self-renewal of ES cells. Nat Genet. 2000;24(4):372-6. doi:10.1038/74199. PMID: 10742100.
(Thomson, J. A. et al., 1998) Thomson JA, Itskovitz-Eldor J, Shapiro SS, Waknitz MA, Swiergiel JJ, Marshall VS, et al. Embryonic stem cell lines derived from human blastocysts. Science. 1998;282(5391):1145-7. doi:10.1126/science.282.5391.1145. PMID: 9804556.
(Yuan, H. et al., 1995) Yuan H, Corbi N, Basilico C, Dailey L. Developmental-specific activity of the FGF-4 enhancer requires the synergistic action of Sox2 and Oct-3. Genes Dev. 1995;9(21):2635-45. doi:10.1101/gad.9.21.2635. PMID: 7590241.
This entry was produced with AI. The literature reconstruction and the protocol extraction were performed by Claude Science working autonomously from public scientific databases and the source papers, with every step logged and reproducible; this write-up was drafted from that work with Claude. The record shows what was done and can be checked against the underlying data.
Hiroaki Kitano
How to cite
Kitano, H. (2026). Machine Traces of Discovery Paths #1 - The Path to iPS Discovery, A Seed- and Abstract-Based Reconstruction, The Discovery Engine .
ORCID: 0000-0002-3589-1953
https://orcid.org/0000-0002-3589-1953
First published: August 14, 2026