A molecule can bind beautifully, score well in QSAR models, show a reassuring early toxicity profile and even produce encouraging Phase II data. It can still fail to become a medicine.
Much of the AI drug-discovery conversation is organised around the milestones that are easiest to measure: finding a target, predicting binding, generating a molecule and reducing obvious toxicity. These are meaningful advances. They can make the earliest stages of discovery faster and more systematic.
But scientific promise is not the same as clinical or commercial value. A medicine must survive the far harder journey through late-stage trials, regulatory scrutiny, patient heterogeneity, real-world care and reimbursement. Phase III is a different beast. The decisive question is not simply whether AI can design an interesting molecule, but whether that molecule can become a safe, effective, durable and economically viable medicine.
That leads me to a deeper question: are we optimising what is easiest to calculate, or what ultimately determines whether a therapy works?
Disease is not a target. It is a state.
The familiar drug-discovery story is reductionist: identify the gene, protein or pathway driving a disease, perturb it and restore health. That model has produced important medicines, but complex diseases rarely behave like isolated switches. Cancer, neurodegeneration and immune disorders emerge from interactions among genomic regulation, cellular programmes, proteins, metabolites, tissue structure, environment and clinical history.
A more useful representation may be a hidden biological state—a position in a high-dimensional landscape that captures the combined condition of the system. What we measure in practice, from sequencing and imaging to pathology and clinical records, provides incomplete and noisy views of that state.
The AI question then changes. Instead of asking, “Which biomarker correlates with this disease?”, we ask, “What underlying biological state best explains the evidence available for this patient?” The target has not disappeared, but it has become part of a system rather than the system itself.
Health is not one point on the map
If disease is a state, we also need a richer definition of health. There is no single universal healthy profile. Normal biology varies by tissue, age, sex, ancestry, environment and many other covariates. Health may be better represented as a reference landscape containing many biologically valid states.
A patient’s disease could then be described as a departure from the relevant part of that landscape. Conceptually, severity becomes the distance between the patient’s inferred state and the nearest plausible healthy state—not a straight line through an abstract space, but a biologically meaningful route.
This is an attractive idea, but it is still a hypothesis that must be tested. A mathematically elegant distance is not automatically a clinically useful measure. It would need to track pathological burden, treatment response and outcomes in independent patient cohorts.
Missing information is part of the problem
Real patients rarely arrive with every modality measured at the same time. One person may have genomics and blood tests; another may have imaging, pathology and a fragmented clinical history. A useful model must reason across what is present without pretending that missing information was observed.
This is where diffusion-based generative modelling becomes interesting. Rather than forcing one confident answer, a model could infer a distribution of plausible completions for missing modalities and use that distribution to reason about the underlying disease state. The output should not be “this is what the missing test would have shown.” It should be “these are the states consistent with the evidence, and this is how uncertain we remain.”
That distinction matters. In biology, uncertainty is not an inconvenience to hide; it can determine whether we measure again, intervene cautiously or refrain from acting.
From molecule generation to therapeutic navigation
If disease is represented as a changing state, treatment becomes a navigation problem. The goal is not merely to hit a target. It is to move the biological system from its present state toward a healthy region while respecting what is therapeutically reachable and safe.
The theoretically shortest route may not be possible with the therapies we have. The clinically useful route may require a sequence of interventions, intermediate biological waypoints and repeated measurements. A treatment could work initially, then stop moving the patient in the desired direction. At that point, the system should help us ask whether to continue, switch agents, combine mechanisms or reconsider the plan entirely.
This reframing connects discovery with development. Molecule design becomes one component of a larger loop: measure, infer, intervene, reassess and adapt. It also makes combination and sequential therapy central rather than exceptional.
Why Phase III changes the AI question
Binding affinity, QSAR predictions and early toxicity estimates are useful local signals. Phase III tests a much broader proposition: whether an intervention produces meaningful benefit across a diverse population, over time, against an existing standard of care and at an acceptable level of risk.
A system-state model will not make late-stage trials easy. But it may help us ask better questions earlier. Are responders and non-responders in different biological regions? Does treatment actually move the inferred state toward health? Which intermediate changes predict durable benefit? Where does the available therapeutic repertoire fail to provide a reachable path?
Answering those questions requires longitudinal, paired patient data—measurements before, during and after intervention—not only larger libraries of molecules. That is a harder data problem, but it is closer to the problem that patients, developers and health systems ultimately need solved.
Where this idea could fail
The data bottleneck is enormous. Few cohorts combine molecular, imaging and clinical modalities from the same patient at the same time, and longitudinal interventional data are rarer still.
Healthy reference landscapes can also encode bias if they do not represent diverse populations and contexts. Generated completions can create dangerous confidence if possibilities are presented as measurements. A visually compelling latent space can be biologically meaningless. And no architecture—however sophisticated—substitutes for prospective experimental and clinical validation.
These are not footnotes. They are conditions that determine whether the idea becomes useful science or merely an elegant diagram.
Questions worth asking
The questions I keep returning to are:
- What should AI ultimately predict: a molecule’s score, a patient’s biological state or a therapeutic trajectory?
- How should models communicate inferred possibilities without turning them into apparent facts?
- What longitudinal data partnerships would be required to learn and test state transitions responsibly?
- Could movement toward a healthier biological state become useful mechanistic evidence without being mistaken for patient benefit?
AI has already changed how we search chemical and biological possibility. The next frontier may be to connect those possibilities to the dynamic reality of disease—to understand not only what molecule we can make, but where a patient is, where health lies and which safe path might connect the two.
That is a much more ambitious goal. It is also a conversation worth having now.
Working on multimodal biology, therapeutic development or the evidence needed to connect the two?