Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice.
The 30-second version
- What. The widely quoted claim that “AI-discovered drugs succeed at roughly twice the rate” traces back, in practice, to a single aggregation by one BCG-adjacent author team (Jayatunga et al., Drug Discovery Today, 2024). Its Phase 1 “80–90%” is 21 of 24 molecules — a small sample (n=24) with no confidence interval reported. A firm-computed Wilson 95% CI on 21/24 spans roughly [69%, 96%], a 27-percentage-point width, so the point estimate is statistically fragile.
- So what. The “doubling” of overall success (from ~5–10% to ~9–18%) is not an observation — it is a forward projection that multiplies the observed Phase 1/2 rates by an assumed traditional Phase 3 rate (this “observed doubling” claim is refuted). Phase 2 efficacy is indistinguishable from historical drugs (~40% vs ~37%). Of 24 AI-discovered targets, only 3–9 are genuinely novel (first-in-class); the rest are chemistry optimized on already-validated targets, so attribution of success to “AI” versus “a validated target” is not separable in the data.
- Now what. Two zeros seal the verdict: zero AI-discovered drugs are FDA-approved (marketed), and zero studies statistically separate the AI contribution from traditional drug success. The honest position is suspend judgment — design/chemistry looks partially demonstrated (Phase 1), but clinical outcomes and attribution are unverified. Watch for survivorship bias and small-sample fragility. Verification here is PARTIAL: 9 confirmed / 1 refuted / 3 unverified.
[demonstration gap] A high Phase 1 pass rate is not the same as clinical efficacy or approval. Where the numbers are single-source, company-self-reported, or a projection rather than an observation (the Phase 1 “80–90%”, the “9–18% doubling”, higher vendor-reported rates), they are attributed and are not headlined as established fact.
The five-minute read
One table, one sample
The claim that “AI drugs (nearly) double the success rate” is one of the load-bearing narratives of the AI-drug-discovery boom. On inspection it rests on a narrow base: a single BCG-adjacent author team, a Phase 1 sample of n=24 (21 successes), and — for the parts that impress most — figures reported by the companies themselves. The thesis of this part: the success-rate advantage is a mixture of partial evidence and a large narrative premium. The evidence is confined to Phase 1 (design/chemistry), and even that is exposed to a small sample and self-reporting. The purpose here is not to argue “AI drugs fail,” but to sharpen the honest evidence frame before any success-rate claim is treated as an asset.
Where the number comes from, and how it was counted
The circulating “Phase 1 80–90%” figure converges on essentially one paper — Jayatunga, Ayers, Bruens, Jayanth and Meier, “How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons,” Drug Discovery Today (June 2024). One author’s contact address is meier.chris@bcg.com, i.e. the team is adjacent to BCG (Boston Consulting Group), not an independent academic consortium (peer-reviewed publication is a fact, but that is distinct from interest neutrality). What the paper actually counts: of AI-derived molecules that had completed Phase 1 by late 2023, 21 of 24 succeeded (≈87.5%). Phase 2 is reported at ~40% on an even smaller sample, with no confidence interval. The denominator is unstable too: the population is ~300 AI-native biotechs and 67 clinical candidates, of which 24 are AI-discovered targets and only 3–9 are genuinely novel — most “AI-derived” work is chemistry optimization on top of already-validated targets.
Phase 1 advantage is real but narrow; Phase 2 erases it
Read carefully, the phase pattern is precise. Phase 1 measures druglikeness, safety and PK — that is, design and chemistry, exactly what generative/physics models are good at, so a Phase 1 edge is partially true (under the small-sample, self-reported caveat). Phase 2 measures efficacy — whether the target and the biology actually move the disease — and there the AI-derived rate (~40%) is indistinguishable from the historical rate (~37%). So “AI designs molecules well” is supported; “AI produces clinical success” is not. The Nature editorial (October 2023) is explicit that many of these numbers are self-reported: “such claims have come from the companies themselves. Until they can be independently verified, some caution is in order.”
[diagram: AI-derived vs traditional success rate by phase]
Phase 1 (design/chemistry: druglikeness, safety, PK)
AI-derived ██████████████████████ ~80-90% (21/24, n=24, no CI)
Traditional ████████████ ~40-65%
-> firm-computed Wilson 95% CI on 21/24 ~ [69%, 96%] (width 27pp)
Phase 2 (efficacy: does the biology move the disease?)
AI-derived ██████████ ~40% (smaller sample)
Traditional █████████ ~37%
-> advantage disappears: not distinguishable
Overall clinical PoS (PROJECTED, not observed)
AI-derived ████ 9-18% <- assumes traditional Ph3 rate
Traditional ██ 5-10%
--------------------------------------------------------------
Two zeros: AI-derived FDA approvals = 0; studies separating
AI contribution from traditional success = 0
Deep dive
1. Background — one paper carries the weight
Part 0 of this series posed the question directly: is “AI drugs raise the success rate” evidence or narrative, and is there any case that statistically separates the AI contribution from traditional drug success? This part answers it by dissecting the source, sample, definition and survivorship structure of the success-rate data. The core fact of the background is that a claim this widely cited rests on a remarkably thin base — a single aggregation whose Phase 1 figure is 21 of 24 molecules, reported without a confidence interval, by a team adjacent to a consulting firm. That does not make the number false; it makes it fragile, and it means the burden of independent verification has not been met.
2. What this deep dive newly establishes
Core answer: the success-rate advantage is a composite — a narrow, partially demonstrated Phase 1 edge (design/chemistry) plus company self-reporting plus a projection plus the halo of already-validated targets — and only the narrow Phase 1 piece is independently and statistically established.
- Single-source aggregation — CONFIRMED: the circulating Phase 1 “80–90%” converges on Jayatunga et al., Drug Discovery Today 2024, a BCG-adjacent team (contact
meier.chris@bcg.com), continuous with the earlier framing paper “AI in small-molecule drug discovery: a coming wave?” (NRDD 2022). Peer-reviewed, but not an independent academic consortium. - Small sample — CONFIRMED (fragility): Phase 1 = 21/24 successes; Phase 2 ~40% on a smaller sample; no CI reported. A firm-computed Wilson 95% CI on 21/24 is roughly [69%, 96%] — a 27-point width, illustrating that “80–90%” is a point estimate that stays open into the high-60s. This computation is the firm’s, to illustrate fragility, not a number from the original paper.
- Unstable definition / denominator — CONFIRMED: population ~300 AI-native biotechs, 67 clinical candidates, 24 AI-discovered targets, of which only 3–9 are genuinely novel (first-in-class). “AI-derived” differs by company (target discovery only? molecule design only? end-to-end?), and most of it is chemistry on already-validated targets — so the denominator itself is shifting.
- No independent verification — CONFIRMED (editorial/high): the Nature editorial (October 2023) warns that the claims come from the companies themselves and that caution is warranted until independently verified. A 2025 critical review (Pharmaceuticals, MDPI) adds that publication bias in training data (over-sampling of successful compounds and well-studied targets) and data leakage inflate reported performance.
3. Strengths and limits of the methodology
Strengths: the phase-by-phase decomposition is clean and interpretable — Phase 1 isolates design/chemistry, Phase 2 isolates efficacy, and the projection is explicitly separated from observation. The boundary is preserved honestly: the Phase 1 edge is not dismissed, it is qualified. Limits: several key quantities are unverified — the precise Phase 2 sample size and the stability of the ~40% figure (small sample, no CI); the higher, company-self-reported success rates that circulate below the peer-reviewed number; and whether the 80–90% would hold as the sample grows (a forward question). The refuted item is specific: the claim of an observed doubling of overall success — the 9–18% figure is a projection that assumes the traditional Phase 3 rate, consistent with the paper’s own description; only the “observed doubling” reading is rejected.
4. Neighbouring domains — the attribution problem
The frontier AI-derived successes stand on already-validated targets, which makes the attribution problem concrete. Zasocitinib (Nimbus/Schrödinger FEP+, licensed to Takeda; TYK2; Phase 3) shows strong late-stage efficacy — but TYK2 was already validated by deucravacitinib, so the computational contribution is selectivity (chemistry), not target discovery. RLY-4008 lirafugratinib (Relay; FGFR2) is a precision-target agent, and FGFR2 is likewise an established oncology driver. Success here is a product of two factors — (A) was the target right (biology) times (B) did the molecule hit it well (chemistry) — and AI/physics engines mostly contribute (B), while the success-rate data does not separate (A) from (B). Success on a validated target (A already true) is therefore contaminated by the favorability of target selection, not a clean read on the AI contribution.
5. Commercialization and market context (survivorship bias, companies)
Two mechanisms compound. First, survivorship bias: discontinued candidates (e.g. EXS-21546, DSP-1181, REC-994, REC-2282, SGR-2921, BEN-2293, per Part 0) can drop out of the numerator depending on the counting date and the definition; count only those that reached Phase 2 and the rate is structurally inflated. Second, small sample times survivorship: in n=24, including or excluding a few failures can pull “80–90%” down toward the 60s. On current public data, the number of studies that statistically separate “the AI contribution changed a clinical hard outcome” from traditional drug success is zero — doing so would require a head-to-head of AI-designed versus traditionally-designed candidates on the same target and indication, or a stratified success rate for the 3–9 novel targets, and the sample (3–9) is below what statistics require.
| Item | Firm verdict | Evidence tier |
|---|---|---|
| “Phase 1 success 80–90%” | Partial Yes — 21/24 (single BCG-adjacent aggregation, n=24, no CI) | peer-review, small sample |
| Phase 1 edge = design/chemistry demonstrated | Yes (confined) — the measured object is druglikeness/PK | interpretation, med |
| Phase 2 efficacy advantage | No — ~40% ≈ ~37%, not distinguishable | secondary/med |
| “Doubling of success (9–18%)” | Refuted (not observed) — forward projection assuming traditional Ph3 rate | stated in paper, re-read |
| Independent verification of rates | No — Nature editorial warning, largely self-reported | editorial/high |
| AI contribution separated from traditional | No (0 studies) — validated-target halo, survivorship bias, small sample | synthesis |
| AI-derived FDA-approved drugs | 0 | fact |
Company context (factual, neutral, no buy/sell implication): the “doubling to 9–18%” narrative circulates as a valuation input for listed AI-drug names (Recursion RXRX, Schrödinger SDGR, Relay RLAY) and for private financing rounds (Insilico, Nimbus and others), while the number itself rests on a projection, a small sample and self-reporting. There is a direction mismatch worth noting: the most heavily capitalized player (Isomorphic Labs) has zero clinical assets, yet the success-rate narrative keeps justifying capital formation. The firm treats company-self-reported success rates as claims and independent, peer-reviewed aggregations as evidence. Stock-specific merit is separated into an investment layer; no single-name conclusion is asserted.
6. The skeptic’s bottom line
- verified-clean: the facts that (i) the Phase 1 21/24 = 80–90% comes from a single BCG-adjacent paper, (ii) Phase 2 ~40% ≈ traditional ~37%, (iii) the Nature editorial calls for independent verification, and (iv) AI-derived approvals = 0 — all cross-checked against primary and peer-reviewed sources, confirmed.
- proceed-with-caveats (conditional): the “Phase 1 edge” statement passes only when the small-sample (n=24), no-CI, BCG-attribution and self-reported caveats are placed alongside it and it is not extended to efficacy or overall success; any listed-company valuation mention must be separated into an investment layer.
- hold: (1) the completed narrative “AI has (clinically) doubled drug success rates” — the doubling is a projection, Phase 2 is indistinguishable, approvals are zero; (2) headlining Phase 1 80–90% as an unconditional absolute, ignoring the [69%, 96%] CI; (3) describing company-self-reported rates as if independently verified; (4) attributing validated-target success to the “net effect of AI.”
- Honest position: suspend judgment. Design/chemistry is partially demonstrated (Phase 1); clinical outcomes and attribution are unverified. This is neither “AI drugs succeed more” nor “AI drugs fail” — it is a claim whose only independently established piece is a narrow one.
7. What to watch (falsifiable predictions)
- As the sample grows (into the hundreds of molecules), the Phase 1 advantage should narrow, and in an independent aggregation that strips out the self-reported share, the 80–90% should fall to the low-70s or below. Falsifiable via a large-scale re-aggregation by unaffiliated researchers.
- Restricted to novel targets (the 3–9), AI-derived candidates should show Phase 2/3 success rates significantly lower than the validated-target group (i.e. success had been riding on target validation). Falsifiable via stratified analysis once the sample accumulates.
- Within the next 24 months, the number of peer-reviewed studies that separate the AI contribution via a same-target, same-indication head-to-head of AI-designed versus traditionally-designed candidates will remain zero. Falsifiable by the appearance of such a study.
References
- Jayatunga, Madura K. S., Margaret Ayers, Lotte Bruens, Dhruv Jayanth, and Christoph Meier. 2024. “How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons.” Drug Discovery Today (June). https://www.sciencedirect.com/science/article/pii/S135964462400134X
- Jayatunga, Madura K. S., et al. 2022. “AI in small-molecule drug discovery: a coming wave?” Nature Reviews Drug Discovery. https://www.nature.com/articles/d41573-022-00025-1
- Nature editorial. 2023. “AI’s potential to accelerate drug discovery needs a reality check.” Nature (October). https://www.nature.com/articles/d41586-023-03172-6
- “AI in Small-Molecule Drug Discovery: A Critical Review of Methods, Applications, and Real-World Outcomes.” 2025. Pharmaceuticals 18(9):1271. https://www.mdpi.com/1424-8247/18/9/1271
- “AI in pharmaceutical development: hype or panacea?” (cites the Nature editorial). European Pharmaceutical Review. https://www.europeanpharmaceuticalreview.com/article/251835
- QuantumBiospace. “How successful are AI-discovered drugs in clinical trials?” (summary of Jayatunga 2024: 300 AI-native biotechs, 67 clinical, 24 AI-discovered targets, 3–9 novel). https://www.quantumbiospace.ai/news/how-successful-are-ai-discovered-drugs-in-clinical-trials/
Source knowledge asset: knowledge-base/deep-dives/ai-drug-clinical-readout/part4-attribution-success-rate.md (generated 2026-07-12, verification PARTIAL — 9 confirmed / 1 refuted / 3 unverified). This draft inherits the figures and source attributions of the original part and creates no new figures or sources. The Wilson 95% confidence interval on 21/24 is firm-computed to illustrate small-sample fragility and is not reported in the original paper.
Disclosure
This post is for information only and is not investment advice. The author holds no position in, and no financial interest in, the listed companies mentioned (Recursion, Schrödinger, Relay Therapeutics and others).
COI note: The primary source for the aggregated success rates (Jayatunga et al. 2024) has several authors affiliated with BCG (Boston Consulting Group); this is stated as fact — it is a consulting/industry-adjacent source and is treated separately from fully independent academic verification. References to listed AI-drug companies (Recursion RXRX, Schrödinger SDGR, Relay RLAY) and private ones (Insilico, Nimbus, Isomorphic Labs) are factual and neutral, with no relative-merit or buy/sell implication. All quantitative success-rate claims are attributed to their original source as “vendor / author / preprint claims”; the “9–18% doubling” is explicitly a forward projection, not an observation, and the Phase 1 “80–90%” carries small-sample (n=24, no CI) and self-reported caveats.
Leave a comment