AI in Drug Discovery: Milestones, Limitations and the Push for Clinically Relevant Data
AI is transforming drug discovery across target ID, molecule design and trial optimization. Milestones include regulatory qualification of digital-twin controls and positive Phase 2a data, though 90-95% of candidates still fail.
Artificial intelligence is being applied across the drug discovery pipeline, from biological target identification and generative molecule design to laboratory automation and clinical trial optimization, with the field now organized into four distinct layers. Recent milestones include a European regulatory qualification for a digital-twin methodology that shrinks placebo arms in clinical trials and positive Phase 2a data for an AI-designed drug candidate.
The clinical trial optimization layer, the newest and least proven part of the stack, includes Unlearn.AI, whose PROCOVA methodology generates digital twin control patients to reduce the size of a trial's placebo arm. The method has been formally qualified by European regulators for Phase 2 and 3 trials, with confirmed alignment from the FDA, and has reportedly reduced required control-arm sizes by roughly 30 percent.
In generative molecule design, Insilico Medicine's rentosertib program has become one of the field's clearest end-to-end successes, moving from an AI-identified target to an AI-generated molecule and into positive Phase 2a data. Isomorphic Labs, spun out of Google DeepMind and carrying the AlphaFold lineage, has partnerships valued at more than $1.7 billion with Eli Lilly and roughly $1.2 billion with Novartis, on top of a $600 million external raise. Xaira Therapeutics launched in 2024 with more than $1 billion in backing, the largest AI-biotech launch on record, while Chai Discovery, Genesis Therapeutics and Iambic Therapeutics compete on structure prediction, translational chemistry and generative design respectively.
At the target identification layer, Recursion Pharmaceuticals applies deep learning to millions of microscope images of cellular assays to map disease biology at scale and has merged with Exscientia, while Insitro combines genomics, imaging and clinical data to stratify heterogeneous diseases into subtypes. The laboratory automation layer includes Emerald Cloud Lab, Strateos and Arctoris, which offer remote, roboticized laboratories as an on-demand service; companies citing automated design-make-test-analyze cycles report compressing the path from initial hypothesis to a preclinical candidate to roughly 13 to 18 months, against a traditional industry average nearer three to four years.
Despite these advances, between 90% and 95% of drug candidates still fail during development, typically after years of investment. In oncology, the problem is not the volume of data or the computational power applied, but the degree to which the data reflect tumor biology in real patients. Patient-derived data are highly relevant but constrained, while model systems such as cell lines, organoids and patient-derived xenografts offer scalability at the cost of patient context. Genomic profiling has not translated into broad clinical success, with only a small proportion of patients eligible for existing precision medicines based on genomic biomarkers alone, and transcriptional signatures are frequently weak predictors of protein expression or therapeutic response. There is a growing effort to integrate functional data — measurements of how tumors respond to drugs — with molecular profiling to identify biomarker signatures that are predictive of phenotype.
Researchers also emphasize that biomedical AI carries higher stakes than AI in many commercial settings, because errors can affect scientific conclusions, drug development decisions or patient care. Biomedical data are often noisy and incomplete, and AI predictions still need experimental and clinical validation. Data integration is a related challenge: drug discovery is increasingly multimodal, spanning small molecules, protein degraders, antibody-drug conjugates, RNA therapeutics, and cell and gene therapies, each generating unique datasets across chemistry, biology and translational research. Many organizations rely on disconnected informatics systems, and knowledge graphs have emerged as a way to connect fragmented data and evidence into a unified, traceable framework. One biomedical knowledge graph solution integrates datasets including MetaBase, Cortellis Drug Discovery Intelligence and OFF-X, encompassing more than 1 million chemical compounds, more than 1 million drugs, nearly 87,000 biological entities, and thousands of disease, pathway, biomarker and safety concepts.
In a five-week collaboration between EU-OPENSCREEN partner sites, a postdoctoral researcher from CIB-CSIC in Madrid joined a research group at Karolinska Institutet in Sweden to explore AI- and machine-learning-based approaches for de novo compound design. The work focused on synthon-based design strategies and Thompson sampling approaches to generate and prioritise molecular candidates, and resulted in a new series of potential Mac1 inhibitors — a key target in viral replication and immune response evasion — providing a starting point for future synthesis and in vitro testing.