Human Cognitive and Behavioral Limitations in Scientific Discovery

Share
Human Cognitive and Behavioral Limitations in Scientific Discovery

The second essay in the Discovery Engine series.

The first essay in this series asked whether machines can discover. Before that question can be answered, a prior one must be faced honestly: how good, really, is human-led discovery? This essay gives an account — structural, not personal — of where the present arrangement reaches its limits.


1. Where human discovery reaches its limits

The first essay in this series closed on a difficult observation: there is now more to know than any mind or institution can keep up with. The result is structural misalignment — not the failure of any individual scientist, but a misalignment in how human discovery is organized as a whole. Structural misalignment is the incompatibility between the structures of human cognition and behavior and the structures of nature, which those minds are trying to understand. A human scientist cannot take in a complex reality as a whole; what they study is what they have chosen to study, or what their methods will let them see.

The first essay asked a question — can machines discover? This one asks a quieter one: what kind of science are we doing now, and what are its limits? Those limits come in two kinds. The first lives inside the individual scientist — in the capacity of a single mind to reach the relevant knowledge, to represent a complex reality faithfully, and to reason across that reality without distortion. The second lives outside the individual, in the incentive structures and institutions that decide what gets studied, funded, published, and believed. The two are not in competition; they operate together, and the science we actually have is the science that runs with both active at once.

None of this is a charge against scientists. Their limits are not matters of character. They are what arises when finite minds, working through finite institutions, face a reality larger and more complex than either. To name those limits is the first step toward designing past them.


2. PART I — Cognitive Limitations

The first kind of limitation is the one a scientist carries inside. Each of the limits that follow narrows what a single mind can see.

2.1 The information horizon

Begin with the simplest limit: there is too much. The scientific literature has grown past what any individual, or any realistic team, can read even in its own subfield. Papers now appear at a rate of millions per year. Past a certain distance from where one stands, relevant knowledge is unreachable in a working lifetime.

The biomedical literature alone now adds more than 1.5 million articles per year (González-Márquez et al., 2024). The natural reply is that no one needs to read them all — one reads only what is important. But who decides which are the important ones? Importance is rarely visible on the surface of a paper; it often becomes legible only after the connection has been made, and the very act of filtering for relevance is bounded by the finite attention that filtering was meant to spare. The limit is therefore not volume itself but a volume that cannot be triaged in advance: knowledge accumulates continuously, and one cannot know which paper, two streets over from one's own subfield, contains the missing piece.

Call this an information horizon, by analogy with the event horizon of physics — the surface beyond which no signal can reach an outside observer. The same condition holds here: once the volume of information generated exceeds the reading capacity of any one mind, the information beyond that threshold continues to exist and accumulate, but no signal from it can arrive in time to be used. The limit is not that a scientist reads too slowly; it is that the readable has outgrown what reading can do. Everything that follows happens inside the horizon — among the knowledge a scientist can reach.

2.2 The information gap

Within reach, a second limit appears: even what is published does not, by itself, deliver what the paper claims to say. The gap meant here is not the one between distant findings — that one belongs to the next essay — but something more basic. A paper does not carry everything its sentences appear to. What gets through is the explicit text; what does not is the tacit knowledge a competent reader silently supplies.

The clearest demonstration is at the level of a single sentence. Consider a statement from a yeast-biology paper: in response to mating pheromones, the Far1–Cdc24 complex is exported from the nucleus by Msn5 (Shimada, Gulli & Peter, 2000). To represent this one statement faithfully — to draw it as a molecular interaction another researcher, or a machine, could reason over — one must supply what the sentence does not say. Is the Far1–Cdc24 complex present in the nucleus to begin with? Is it exported regardless of how it has been post-translationally modified, or only in certain forms? Does Msn5 itself relocate as part of the event? None of this is written; all of it is required, and a competent yeast biologist fills it in without noticing. The Systems Biology Graphical Notation is the standard graphical notation developed for this purpose, designed to surface what natural-language descriptions leave implicit (Le Novère et al., 2009) — and even SBGN diagrams cannot draw the molecule's prior nuclear localization unless someone knows to draw it. The unwritten layer lives outside the sentence, inside the reader.

fig2_1_information_gap.jpg
[Figure 2.1 — The information gap inside a single statement: a natural-language sentence at top; an SBGN diagram beneath rendering Far1, Cdc24, Msn5, and the nucleus boundary; and beside the diagram, the questions the sentence does not answer (nuclear presence? modification state? Msn5 relocation?).]

The same incompleteness appears at the level of method. A protocol section reports the steps that distinguish one experiment from others; it does not report the thousand small judgements that any practitioner in the field is assumed to know — the timing of a transition, the feel of a confluent monolayer, the choice between two equally defensible buffer pHs, the lab's standing convention for handling reagents that have sat too long on the bench. Two laboratories receiving the same printed protocol will not, in general, perform the same experiment. Some of what is called a reproducibility problem (which we return to in §4) is in fact a transmission problem: the protocol document is not the protocol, in the same way the map is not the territory. What the paper carries is the explicit residue of an act of investigation whose tacit substance lives in laboratories, not on pages.

Both examples point at the same failure. Scientific writing is a compression that assumes a competent decompressor, and the compression rate is high. For a human in the right subfield, decompression is automatic and the gap is invisible. For anything reading science at scale, the unwritten layer is exactly what does not arrive — and what does not arrive is a substantial fraction of what science actually knows.

2.3 The representation problem

Suppose the knowledge is reached and connected. There remains the question of how it is represented — and here human cognition imposes a cost.

Human conceptual categories are coarse. They are low-dimensional, shaped by what researchers can hold in memory and communicate to one another. When the underlying reality is high-dimensional and nonlinear, the category — the map — fails to match the territory. Alfred Korzybski's formulation is the most economical: the map is not the territory (Korzybski, 1933). The trouble is that we are obliged to reason with maps drawn at the resolution of the human mind.

fig2_2_non_linear_object.jpg
[Figure 2.2 — How to describe a non-linear object: an irregular, closed (non-linear) region, and the open question of which axes to project it onto, how many dimensions to retain, and how fine to make the grid — a coarse partition versus a fine one over the same shape.]

The cost shows up in two distinct acts of representation. The first is how biology models its own mechanisms. A living signaling system is a dense web of feedback, cross-talk, and nonlinear response; the representation a researcher can hold and publish is typically a linear chain — A activates B, which activates C. The pathway diagram standard in molecular biology is exactly such a compression. The low-dimensional idiom organized the field for decades not because anyone believed cells were linear, but because the linear story was the one a mind could carry.

The second is how reality is partitioned into categories in the first place, and its cost is concrete. When a single label covers a family of distinct underlying conditions — Marfan syndrome is the often-cited case, with roughly a quarter of patients waiting years for correct identification (EURORDIS, 2007, cited in Kitano, 2016) — the coarse category mismatches the territory, and the mismatch propagates downstream into trials run on mixed populations, where a real effect in one true subtype is washed out across the rest. The same coarse-category failure that flattens a feedback network into an arrow also flattens a heterogeneous reality into a single label.

The same failure recurs at the scale of the genome, and it returns with force in §3.4. For decades much of the non-coding genome was filed under the inherited category "junk DNA." The error was not in noticing that some of the genome is genuinely without function — much of it is. The error was that a coarse, protein-centric label licensed inattention to the functional regulatory elements embedded within those regions, where a large share of the genetic signal for human traits turns out to lie. The representation problem and the long tail of §3 are, at bottom, the same failure seen from two sides: a category inherited for convenience, mistaken for a finding.

2.4 The dimensionality limit

Why must our categories be coarse? Because the reasoning process itself cannot traverse a high-dimensional, nonlinear, feedback-laden system unaided. This is the deepest of the cognitive limits, and the most rooted in biological constraint.

Working memory holds only a handful of distinct items at once — the classic estimates range from about seven (Miller, 1956) down to roughly four (Cowan, 2001). Whatever the exact figure, it is small, and it is a hard bound on how many interacting variables a person can hold in active relation. Living systems do not respect that bound: behavior emerges from the simultaneous interaction of dozens or hundreds of components, and the interesting behavior lives precisely in the interactions, not the parts.

This is not the whole story, and the honest version is stronger for admitting it. Humans partly circumvent the bound — through external notation, mathematics, simulation, and the division of cognitive labor across collaborators. A field can collectively hold what no member holds alone. But these aids extend the limit; they do not abolish it. The mathematics must still be devised, the simulation still interpreted by a mind subject to the same four-slot register, and the larger the system, the more the burden of integration falls back on individuals who cannot, in the end, hold it all at once. The coarse categories of §2.3 are therefore not a lazy choice but a forced one, and forced limits of this kind are what a differently built system is positioned to relieve.

2.5 Cognitive biases in scientific reasoning

The reasoning process is not only bounded; it is systematically skewed. Confirmation bias, anchoring, availability, premature closure, representativeness — these are not isolated incidents but consistent biases, well catalogued in clinical reasoning (Tversky & Kahneman, 1974), and they bend scientific judgment in predictable directions. They act not only on individual judgments but on entire fields.

Consider the lock-and-key picture of protein–ligand binding (Fischer, 1894). The model survived as the dominant framework for nearly a century. Confirmation bias kept it intact: results consistent with rigid complementarity were celebrated, while evidence of protein flexibility was absorbed as "interesting exceptions" rather than as evidence that the category itself was wrong. Anchoring kept the field tethered to the original metaphor even after induced-fit (Koshland, 1958) had loosened it; the rigid-shape picture remained the default mental image long after the data had moved on. Availability filtered what got studied: proteins amenable to crystallization — and therefore to static-shape visualization — were the proteins the field knew, while disordered and dynamic proteins were systematically under-represented, not because they were unimportant but because they were invisible to the available method. The result was a category framework that selected the data, and the data confirmed the framework. The induced-fit refinement (Koshland, 1958), and later the conformational-selection picture (Boehr et al., 2009), eventually re-categorized binding as a dynamic event; but the rigid-shape framework had filtered decades of inquiry before the re-categorization arrived. They are, in part, a consequence of how human categories work: as George Lakoff argued, our categories are not neutral mirrors of the world but structured by human cognition itself (Lakoff, 1987), so the very act of categorizing imports a prior.

Picture an irregular reality and several observers attempting to describe it. Each, constrained by a different prior and vantage, imposes a rectilinear partition over the same continuous shape. The partitions do not coincide; none captures the true contour; and each observer is convinced their partition is the shape of the thing. The same reality yields different representations to different minds — and the same observed expression can be produced by different underlying realities, so that agreement on the data does not guarantee agreement on the world. Bias makes the map not merely coarse but tilted, in ways the mapmaker cannot easily detect from the inside.

fig2_3_reality_vs_cognition.jpg
[Figure 2.3 — Reality vs. Human Cognition: an organic, cloud-shaped region (reality) overlaid by two offset, partly overlapping dashed rectangles ("Human Cognitive Representation 1" and "2"), neither matching the cloud's contour nor each other.]

2.6 The minority report problem

The final cognitive limit operates even when everything else has gone right. Suppose the signal is reached, connected, faithfully represented, and correctly perceived. It can still be discarded — because it sits in the minority.

Make it concrete. Suppose ninety-nine percent of the reports on a given pair of molecules conclude "A activates B," and one percent conclude "A inhibits B." What should be done with that one percent? The reflexive answer is to treat it as noise and ignore it — and often that is right, because most outliers are noise. But notice that the reflexive answer never asks the question that actually matters: what is the quality of the minority report? The right discriminator is not the size of the minority but its independence. If all ninety-nine confirming reports trace back to a single laboratory, a single technique, or a single founding assumption, their numerical weight is an illusion — they are one observation repeated, and the lone dissenter from a different lab may carry more true information than all of them. If, instead, the ninety-nine come from genuinely diverse labs, methods, and model systems, the consensus is robust and the outlier probably is noise.

The failure here is therefore not perception but valuation, and "weight by count" is the wrong instrument: it counts repetitions instead of weighing independent evidence. The minority signal is seen and then dismissed for its rarity, when the right question concerns the structure of the majority that excludes it. The moment valuation depends on the structure of the community rather than the content of a single mind, we have left individual cognition and entered the territory of incentives and institutions. That is the hinge into Part II.


3. PART II — Social and Behavioral Limitations

The second kind of limitation lives outside the individual mind, in the structures that govern what the enterprise, as a collective, chooses to do. Part I asked what a single scientist can perceive. Part II asks what the system as a whole permits and rewards.

3.1 Career incentives, risk aversion, and the concentration bottleneck

Start without cynicism, because cynicism misidentifies the mechanism. Scientists are moved by two things at once: a viable career, and a sincere wish to do work that matters. Both push toward the same choice — to study targets already known or strongly presumed important. Working on a well-characterized gene is both safer for a career and a reasonable bet about where the important biology lies.

But individually rational choices, aggregated, produce a collective outcome with a power-law shape. A few targets receive the overwhelming share of attention; a vast number receive almost none. Su and Hogenesch (2007) named the mechanism — the information-scientist's principle of least effort: one studies what is easy to study, and what is easy to study is, increasingly, what has already been studied. This is the preferential attachment that generates scale-free networks elsewhere (Barabási & Albert, 1999).

The magnitude is stark. Roughly three-quarters of biomedical research remains on the tenth of human genes already known before the genome was mapped (Edwards et al., 2011); research clusters on about two thousand of the roughly twenty thousand protein-coding genes (Stoeger et al., 2018). And it correlates weakly with present importance: the most-studied tenth are, by one estimate, only about four times more physiologically important than the least-studied tenth, yet studied on the order of two thousand times as much. The skew tracks the experimental tractability of the 1980s and 1990s — which genes happened to be approachable then — not present significance.

Nor is the bottleneck attention alone. It is also a structural condition. Funding follows a power-law distribution similar to that of publications (Su & Hogenesch, 2007), so Robert Merton's "Matthew effect" — unto those who have, more shall be given (Merton, 1968) — is visible in the money, not merely inferred from citations. And it reaches careers: graduate students and postdocs who work on poorly characterized genes have markedly lower odds — by one estimate, roughly halved — of becoming independent investigators (Stoeger et al., 2018). The tail is not merely unstudied. It is rational to avoid.

Here I will show how much opportunity is being missed by overlooking the tail, and why this matters. Consider TREM2, a receptor on the brain's immune cells. It was first cloned around 2001, then sat nearly flat in the literature for over a decade; its publication count exploded only after 2013, when two papers reported that rare TREM2 variants roughly double the risk of late-onset Alzheimer's disease (Guerreiro et al., 2013; Jonsson et al., 2013), at which point preferential attachment took over and the field rushed in (Qian et al., 2024). Here is what this means. First, the importance was present the entire time — the risk those variants conferred was a biological fact in 2001 no less than in 2013. Second, attention through that decade was allocated almost independently of that latent importance; the gene sat in the tail. Third, a single triggering finding caused an abrupt reallocation. Therefore the allocation mechanism was not tracking latent biological significance — if it had been, the reallocation would not have waited twelve years for an accident to trigger it. The lag is evidence about the mechanism: attention is assigned by tractability and precedent, and importance is discovered late, by accident, in proportion to neither.

3.2 Collective evaluation bias

The TREM2 story is one of omission: no one looked. There is a second, distinct failure, in which people do look — and judge wrong.

The development of chimeric antigen receptor (CAR) T-cell therapy is the cleanest example. Through the 1990s and 2000s, Carl June's pursuit of engineered T-cells to treat cancer was widely regarded as fringe; funding was hard to secure and the work hard to publish, because the dominant framework treated cell therapy as marginal. The verdict was not a failure to perceive the work — it was perceived, assessed, and found wanting. Then, around 2017, the first CAR-T therapies won approval and moved toward standard of care for certain blood cancers, vindicating roughly two decades of work the community had collectively undervalued.

The contrast with §3.1 is structural. TREM2 is "no one looked." CAR-T is "people looked and judged wrong." Both are failures, but at different stages: the first of coverage, the second of evaluation. The community, acting as a collective evaluator, regularly fails to recognize transformative directions in progress — not because its members are foolish, but because evaluation is anchored to the prevailing paradigm, and the most consequential work is frequently the work that does not fit it.

3.3 Institutional contingency

Behind both incentive and evaluation lies a deeper fact: what science pursues is path-dependent. The questions a field takes up are shaped by which institutions happened to exist, which programs happened to be funded, which methods happened to mature first, and who happened to hold influence when a direction was set. The same reality could have been investigated in a different order — or not at all — under a different history.

This generalizes §3.1 and §3.2 from incentive to contingency, and it carries a consequence that points ahead. Science is organized — by discipline, journal, department, society — and that organization, whatever its virtues, actively keeps distant fields distant. The boundaries that make a discipline legible are the same boundaries that make cross-disciplinary contact rare. This is also the institutional aspect of what hinders the cognitive breakthrough of joining distant bodies of knowledge. The structure that makes such joining rare is itself a behavioral limitation, and the next essay takes it up directly.

3.4 The empirical signature — the two long tails

Everything argued so far leaves a measurable trace in the literature itself.

The first long tail is one of scientific attention. Repeatedly select human genes at random, and plot for each how many publications have been issued about it. The distribution is not a gentle gradient: a handful of genes — TP53, EGFR — dominate the field, while dozens sit near the floor, named but barely studied. In one such chart, across roughly a hundred randomly selected human genes, 96.8% of the publications fall on just the top ten (Kitano, 2025). The sample is random, not curated to dramatize the point, and the concentration appears anyway. The underlying law is that references per gene decay as a power law (Su & Hogenesch, 2007); at the extreme, in a 2007 census the single most common publication count for a human gene entry was zero. This distribution likely reflects the experimental tractability of earlier eras and preferential-attachment dynamics, among other factors (Stoeger et al., 2018).

fig3_1_long_tail.jpg
Figure 3-1: Long tail distribution of publications on randomly picked 100 human genes — 96.8% of publications are on the top 10 genes.

The phenomenon has an established name. Edwards et al. (2011) call it the Harlow-Knapp (H-K) effect — "the propensity of the biomedical and pharmaceutical research communities to focus their activities, as quantified by the number of publications and patents, on a small fraction of the proteome." Their analysis of the human protein kinome reveals how deeply the effect persists. Of the 518 human kinases, the same small set drew 84% of citations before the kinome was charted in 2002, 77% from 2003 to 2008, and 74% in 2009 — even as the full human genome sequence arrived, and even as unbiased screens kept turning up important new family members. The genome was precisely the new capability that ought to have redistributed attention. It did not. The lag is not a lag in tools. It is a lag in allocation.

The second long tail is one of reality. Biology is itself long-tailed — at multiple scales — and its tail does not coincide with the head of attention. Within a single cell, the topology of metabolic networks is governed by power laws in metabolite degrees: a handful of metabolic hubs and a long tail of locally important reactions (Tanaka, 2005). Within a single medicine, mass-spectrometry of maoto, a traditional Japanese (Kampo) herbal remedy for flu-like symptoms, finds the pharmacological activity distributed across a long tail of compounds — numerous components of broad-spectrum activity acting on multiple molecular targets with weak to moderate intensity (Nishi et al., 2017); the term long-tail drugs names this class explicitly (Kitano, 2007). Within a single disease, causal drivers follow the same distribution: a uniform analysis of more than a thousand prostate cancers found ninety-seven significantly mutated genes, seventy never previously implicated in the disease, most altered in fewer than three to five percent of cases — discoverable only because the cohort was large enough to see them (Armenia et al., 2018). Across diseases, the shape holds again: thousands of rare Mendelian disorders, individually uncommon and collectively vast, a domain its own students describe in the explicit language of the long tail (Shen et al., 2015). At every scale at which biology has been seriously measured, the head is small, the tail is long, and the tail is where most of the action lives.

A clarification, because "long-tailed reality" can be heard too loosely. The claim is not that every biologically important quantity follows one and the same distribution — effect sizes, disease frequencies, mutational spectra, and tractability are different distributions with different shapes. The claim is narrower and more defensible: causal significance is diffusely distributed, spread well beyond the narrow region where attention concentrates. Importance does not live only in the head.

Here is the synthesis. There are two long tails — one of attention, one of reality. Structural misalignment names the gap between them — between where attention concentrates and where causal significance actually lies. Attention's head rests on a small set of long-known, tractable targets; reality's important causes are spread thin across the tail. A human-led enterprise, doing everything its incentives ask of it, is in effect optimizing the wrong tail. That is the one-sentence form of the entire diagnosis.

This is an interpretation, not a measurement, and its strongest objection has to be granted plainly. The objection is that the concentration is not misalignment but rationality under constraint. A tractable target is one where progress is actually possible; concentrating force is how cumulative science gets done at all; and much of the tail was not irrationally ignored but genuinely inaccessible until enabling technologies — sequencing, large cohorts, computation — matured. There is no investigating all tails uniformly, and the exploration–exploitation tradeoff is real. Nor is concentration merely a tax; it is also how scientific communities build the depth, the standards, the shared tools, and the cumulative confidence on which any field depends. A hundred labs converging on TP53 is how the kinase-inhibitor industry, the structural-biology pipeline, and the clinical-trial infrastructure for oncology came into existence — none of which would have been built by uniform investigation of an undifferentiated tail. The charge here is not that science failed to be omniscient.

Granting this does not dissolve the misalignment; it sharpens what the misalignment is. The TREM2 evidence (§3.1) is the same diagnosis at the single-gene scale: the importance was present and the tools to find it existed for years before attention arrived. What the kinome H-K data show across hundreds of proteins, TREM2 shows for one.

Let me state this more precisely. The concentration tracks a real feature of biology, but it then overreads it. Biological networks genuinely are scale-free, and hub genes like TP53 are highly pleiotropic and do exert outsized downstream control. So studying hubs intensively is not itself a mistake.

What, then, is the misalignment? It lies somewhere more subtle. There is a tacit assumption running through current science — that if we just look at the hubs, we can understand and steer the system. But there are places where the hubs' control runs out, and those places lie in the tail. Even granting that studying hubs is correct, the important objects of study that hubs alone cannot capture lie in great number in the tail.

In other words: the problem is not the focused, deep investigation of hubs (exploitation) itself. The current system has well-developed methods for investigating hubs deeply, but it does not yet have, alongside them, methods for exploring the long tail at scale. Human-led science, under its constraints, already exploits efficiently. What it cannot do, under those same constraints, is explore the long tail at the scale the tail demands. To address this, we do not need to stop the concentrated investment that is working — we need to establish a methodology that efficiently explores the tail and produces discoveries from it.


4. Twilight zone reasoning

It would be convenient if Part I and Part II were separable problems, solved one at a time. They are not. They co-occur, and the actual regime is one in which a bounded, biased mind works inside a contingent, incentive-warped institution — at once. The symptoms compound, and the visible one is reproducibility: by one well-known assessment, a large fraction of preclinical findings in a particular sample could not be reproduced (Prinz, Schlange & Asadullah, 2011). The figure is contested in scope and should not be inflated into a slogan, but its behavioral causes are not in doubt — negative results are buried (the file-drawer problem; Rosenthal, 1979) and novelty is rewarded over replication, so the record is skewed at the source.

And yet science works. Knowledge accumulates not from a sequence of clean, individually correct papers but from a noisy collective — error-prone reports, biased datasets, arbitrary interpretations — that nonetheless converges, imperfectly and over time, on something usable and true. Hiroaki Kitano named this twilight zone reasoning (Kitano, 2016).

How does a noisy aggregate converge at all? The mechanism is redundancy across independent error. A robust biological truth tends to leave traces in many places at once — different assays, organisms, laboratories, and theoretical vantages — and while each approach carries its own bias, the biases are not identical, so a real effect survives the intersection of many flawed views where an artifact does not. Convergence is, in effect, error-cancellation by uncorrelated mistakes. The mechanism is real and powerful. It is also inefficient: it spends enormous redundant labor and tolerates long latencies — the TREM2 decade is a case in miniature, where the truth survived but only after an accident. That a system this noisy works at all is the achievement. That it works this slowly is the opening. The question that follows is whether a differently optimized system could navigate the same regime more efficiently.


5. Implication

Parts I and II describe a set of limitations most of which are the kind a differently built system is positioned to relieve. Section 4 names the combined regime in which they operate and the price it pays in time. What this essay has not yet drawn is the boundary between the limitations a differently built system can address and the one capability that may resist scale — the analogical leap that, when historians look back, is what they tend to call the central act of discovery. That boundary is the work of the next essay.

There is a precedent for the kind of relief this essay implies. Long-tailed problems have been "opened up" before. Internet retail made the vast tail of niche products reachable by lowering the cost of access; next-generation sequencing did the same for the long tail of rare disease, turning thousands of individually neglected disorders into a tractable domain (Shen et al., 2015). In each case the tail did not shrink — the means of reaching it changed. The implication of this essay is that the long tail of discovery itself may be a problem of the same shape: not a tail to be deplored, but a tail to be opened.

That is where the diagnosis ends. The next essay asks which parts of it a differently built system can actually reach.


AI Co-Authorship Disclosure

This essay was developed collaboratively between the human author and AI systems, consistent with the disclosure practice established in the first essay of this series. The argument, structure, source selection, and final judgments are the author's; AI systems contributed to drafting, source synthesis, fact-checking, and iterative refinement.


References

  • Anderson, C. (2006). The Long Tail: Why the Future of Business Is Selling Less of More. Hyperion.
  • Armenia, J., Wankowicz, S. A. M., Liu, D., Gao, J., Kundra, R., Reznik, E., et al. (2018). The long tail of oncogenic drivers in prostate cancer. Nature Genetics, 50(5), 645–651. https://doi.org/10.1038/s41588-018-0078-z
  • Barabási, A.-L., & Albert, R. (1999). Emergence of scaling in random networks. Science, 286(5439), 509–512.
  • Boehr, D. D., Nussinov, R., & Wright, P. E. (2009). The role of dynamic conformational ensembles in biomolecular recognition. Nature Chemical Biology, 5(11), 789–796. https://doi.org/10.1038/nchembio.232
  • Cowan, N. (2001). The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24(1), 87–114.
  • Edwards, A. M., Isserlin, R., Bader, G. D., Frye, S. V., Willson, T. M., & Yu, F. H. (2011). Too many roads not taken. Nature, 470(7333), 163–165. https://doi.org/10.1038/470163a
  • EURORDIS (2007). The Voice of Rare Disease Patients: Survey on the delay in diagnosis of eight rare diseases in Europe ("EurordisCare2"). [cited in Kitano, 2016]
  • Fischer, E. (1894). Einfluss der Configuration auf die Wirkung der Enzyme. Berichte der deutschen chemischen Gesellschaft, 27(3), 2985–2993. https://doi.org/10.1002/cber.18940270364
  • González-Márquez, R., Schmidt, L., Schmidt, B. M., Berens, P., & Kobak, D. (2024). The landscape of biomedical research. Patterns, 5(6), 100968. https://doi.org/10.1016/j.patter.2024.100968
  • Guerreiro, R., Wojtas, A., Bras, J., et al. (2013). TREM2 variants in Alzheimer's disease. New England Journal of Medicine, 368(2), 117–127.
  • Jonsson, T., Stefansson, H., Steinberg, S., et al. (2013). Variant of TREM2 associated with the risk of Alzheimer's disease. New England Journal of Medicine, 368(2), 107–116.
  • Kitano, H. (2007). A robustness-based approach to systems-oriented drug design. Nature Reviews Drug Discovery, 6(3), 202–210. https://doi.org/10.1038/nrd2195
  • Kitano, H. (2016). Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery. AI Magazine, 37(1), 39–49. https://doi.org/10.1609/aimag.v37i1.2642
  • Kitano, H. (2025). PubMed publication distribution across human genes [figure]. © Hiroaki Kitano.
  • Korzybski, A. (1933). Science and Sanity: An Introduction to Non-Aristotelian Systems and General Semantics.
  • Koshland, D. E. (1958). Application of a theory of enzyme specificity to protein synthesis. Proceedings of the National Academy of Sciences USA, 44(2), 98–104.
  • Lakoff, G. (1987). Women, Fire, and Dangerous Things: What Categories Reveal About the Mind. University of Chicago Press.
  • Le Novère, N., Hucka, M., Mi, H., Moodie, S., Schreiber, F., Sorokin, A., et al. (2009). The Systems Biology Graphical Notation. Nature Biotechnology, 27(8), 735–741. https://doi.org/10.1038/nbt.1558
  • Merton, R. K. (1968). The Matthew effect in science. Science, 159(3810), 56–63.
  • Miller, G. A. (1956). The magical number seven, plus or minus two. Psychological Review, 63(2), 81–97.
  • Nishi, A., Ohbuchi, K., Kushida, H., Matsumoto, T., Lee, K., Kuroki, H., Nabeshima, S., Shimobori, C., Komokata, N., Kanno, H., Tsuchiya, N., Zushi, M., Hattori, T., Yamamoto, M., Kase, Y., Matsuoka, Y., & Kitano, H. (2017). Deconstructing the traditional Japanese medicine "Kampo": compounds, metabolites and pharmacological profile of maoto, a remedy for flu-like symptoms. npj Systems Biology and Applications, 3, 32. https://doi.org/10.1038/s41540-017-0032-1
  • Prinz, F., Schlange, T., & Asadullah, K. (2011). Believe it or not: how much can we rely on published data on potential drug targets? Nature Reviews Drug Discovery, 10(9), 712. https://doi.org/10.1038/nrd3439-c1
  • Qian, M., Zhong, J., Lu, Z., Zhang, W., Weng, M., Zhang, K., & Jin, Y. (2024). Bibliometric analysis of TREM2 (2001–2022): Trends, hotspots and prospects in human disease. International Journal of Medical Sciences, 21(10), 1852–1865. https://doi.org/10.7150/ijms.96851
  • Rosenthal, R. (1979). The "file drawer problem" and tolerance for null results. Psychological Bulletin, 86(3), 638–641.
  • Shen, T., Lee, A., Shen, C., & Lin, C. J. (2015). The long tail and rare disease research: the impact of next-generation sequencing for rare Mendelian disorders. Genetics Research, 97, e15. https://doi.org/10.1017/S0016672315000166
  • Shimada, Y., Gulli, M.-P., & Peter, M. (2000). Nuclear sequestration of the exchange factor Cdc24 by Far1 regulates cell polarity during yeast mating. Nature Cell Biology, 2(2), 117–124. https://doi.org/10.1038/35000073
  • Stoeger, T., Gerlach, M., Morimoto, R. I., & Nunes Amaral, L. A. (2018). Large-scale investigation of the reasons why potentially important genes are ignored. PLOS Biology, 16(9), e2006643.
  • Su, A. I., & Hogenesch, J. B. (2007). Power-law-like distributions in biomedical publications and research funding. Genome Biology, 8(4), 404. https://doi.org/10.1186/gb-2007-8-4-404
  • Tanaka, R. (2005). Scale-rich metabolic networks. Physical Review Letters, 94(16), 168101. https://doi.org/10.1103/PhysRevLett.94.168101
  • Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124

Read more

科学的発見における人間の認知的・行動的限界

科学的発見における人間の認知的・行動的限界

本シリーズの第一論考は、機械は発見できるかと問うた。その問いに答えるには、それに先立つ別の問いを正面から引き受けなければならない。すなわち、人間が主導する発見は、本当のところどれほどのものなのか。本稿は、現在の状況がどこで限界に達しているかを、人間の認知構造と社会構造という「構造」の問題として論じる。 1. 人間の発見が限界に達するところ 本シリーズの第一論考は、ある厳しい観察で締めくくられた。今日、生み出される知識の量は、どんな個人の知性も、どんな制度も追いつけないほど膨大になっている。また、人間の認知構造の限界、研究を行う社会的環境からの研究テーマの選択に関する影響などいろいろな問題が、人間による科学研究の限界の要因になっている。その帰結が構造的ミスアライメントである。これは個々の科学者の失敗ではなく、人間の発見というものが全体としてどう組織化されているかの不整合なのである。構造的ミスアライメントとは、人間の認知と行動の構造と、それが捉えようとしている自然の構造との、相互の食い違いを指す。一人の科学者が複雑な現実を丸ごと受け止めることはできない。その人が研究するものは、

By Hiroaki Kitano