Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice.
The 30-second version
- What. This part dissects individual AI-designed clinical readouts along a single axis: primary endpoint vs. headline. The decisive fact of framing asymmetry: rentosertib (headlined “positive”) and BEN-2293 (headlined “failed”) both had safety as their Phase 2a primary endpoint, and both met it. What split the headlines was the direction of the secondary efficacy narrative, not the primary result.
- So what. Read against the endpoint structure, the pattern is consistent: rentosertib’s efficacy signal (FVC +98.4 mL at 60 mg QD) is a secondary / exploratory outcome in a small, 12-week study, and its discontinuations cluster in the highest-efficacy, highest-dose arms (16/71 = 22.5% overall). EXS-21546 was not an efficacy-readout failure — it was discontinued on target-pharmacology grounds (A2A receptor coverage / therapeutic index). Across four clinical-stage AI-derived candidates, there is essentially no case whose primary endpoint was efficacy, and zero cases where AI’s contribution has been separately shown to move a hard efficacy outcome.
- Now what. Hold the verdict narrowly: what AI has demonstrated here is design (a safe, clinic-ready molecule); what remains unproven is efficacy (that the molecule moves the disease). The metric to watch is not a “positive/failed” press headline but whether a follow-up trial promotes efficacy to a primary endpoint under placebo control and reproduces it.
[demo-gap note] Every efficacy number cited here (rentosertib FVC, REC-4881 polyp burden, SGR-1505 ORR, BEN-2293 subgroup EASI) is a secondary / exploratory or small-sample early signal, not a confirmed efficacy result. rentosertib n=71, BEN-2293 n=91, REC-4881 evaluable n≈11–12 (single-arm, no placebo). BEN-2293’s ≥20% BSA efficacy signal is a post-hoc subgroup, not a pre-specified result. The “AI designed the molecule” fact and “the molecule works” claim cannot be attributed to the same evidence.
The five-minute read
The headline is decided by the secondary endpoint, not the primary
AI-derived Phase 2a trials typically set safety as the primary endpoint — that is, they test whether the designed molecule is a tolerable, clinic-viable drug. Efficacy sits in the secondary / exploratory tier. So the real success/failure, and the press framing, are decided one tier down. The sharpest illustration is a matched pair: rentosertib (Insilico) and BEN-2293 (BenevolentAI) both met their safety primary endpoint, yet one was reported “positive Phase 2a” and the other “flunks / misses Phase 2a.” The firm’s discipline — what did the primary endpoint actually test? — applies here most sharply.
rentosertib: the frontier of end-to-end AI design, but the primary endpoint is safety
GENESIS-IPF (Phase 2a, double-blind, placebo-controlled, n=71, 22 China sites, 12 weeks; four arms: placebo 17 / 30 mg QD 18 / 30 mg BID 18 / 60 mg QD 18; Nature Medicine 2025, 31:2602–2610). Primary endpoint = safety (TEAE rate) — met; the authors describe the drug as “safe and well-tolerated.” The efficacy signal is secondary: forced vital capacity (FVC) change at 12 weeks ran placebo −20.3 mL, 30 mg QD −27.0 mL, 30 mg BID +19.7 mL, and 60 mg QD +98.4 mL (95% CI 10.9–185.9). A dose-response tendency is visible, but this is a small-sample (≈18/arm), 12-week secondary outcome, and the authors themselves conclude only “safety confirmed, dose-dependent lung-function signal, large/long-term confirmation needed.” The bottleneck is tolerability: 16/71 (22.5%) discontinued early, with dropouts concentrated in the highest-efficacy arms (30 mg BID 6, 60 mg QD 6), largely from GI events plus mild liver-enzyme elevation — an efficacy-tolerability trade-off. AI contribution here is genuinely end-to-end (target discovery via PandaOmics on the novel target TNIK, plus molecular design via Chemistry42).
Vendor / company narrative vs. endpoint structure, and a matched failure
EXS-21546 (Exscientia, A2A antagonist; IGNITE-AI Phase 1/2 in RCC/NSCLC with anti-PD-1, co-developed with Evotec) was discontinued in 2023 — but the company’s stated reason was that “sufficient and sustained A2A receptor occupancy within the therapeutic range” was hard to achieve, i.e. a target-pharmacology / therapeutic-index judgment, not an efficacy-readout failure. BEN-2293 (BenevolentAI, pan-Trk, atopic dermatitis; Phase 2a, n=91, topical BID 28 days) met its safety primary endpoint but missed statistical significance on its secondary efficacy endpoints (EASI / NRS) in the ITT population; a post-hoc ≥20% BSA subgroup showed an EASI signal (ITT p=0.0296, PP p=0.0427) that is hypothesis-generating only. The decisive symmetry: BEN-2293’s primary endpoint (safety) was met exactly as rentosertib’s was — only the secondary efficacy direction differed, yet the headlines diverged sharply.
[diagram: same primary endpoint, opposite headlines]
"Was the AI-designed Phase 2a positive or failed?"
CANDIDATE PRIMARY ENDPOINT HEADLINE WHAT ACTUALLY SPLIT IT
─────────────────────────────────────────────────────────────────────
rentosertib safety -> MET "positive" 2ndary FVC +98.4mL
(Insilico) small n, 12wk, 22.5% D/C
BEN-2293 safety -> MET "failed" 2ndary EASI/NRS miss (ITT)
(BenevolentAI) post-hoc >=20% BSA only
─────────────────────────────────────────────────────────────────────
EXS-21546 (Ph1/2, halted) "discontinued" NOT efficacy failure —
(Exscientia) A2A coverage / therap. index
─────────────────────────────────────────────────────────────────────
REC-4881 (Ph1b/2, polyp) "positive" single-arm, evaluable n~11-12
SGR-1505 (Ph1, activity) "signal" ORR 22%, 45 pts, early
Verdict: primary endpoint = safety/design in every case.
"positive/failed" is decided by the SECONDARY efficacy
tier + company communication. AI-attributable efficacy = 0.
Deep dive
1. Background — the clinical version of the demo-gap
Part 0 of this series established the pattern (AI-derived candidates advantaged at Phase 1, ordinary at Phase 2). This part supplies the micro-evidence. The framing must be precise: the claim is not “AI drug design does not work.” It is that AI has demonstrated the design / safety tier (a molecule that can enter and survive the clinic), while the efficacy tier (that the molecule moves the disease) remains unverified. This is the clinical analogue of the demo-gap seen in the computing and bio-FM series: the primary endpoint passes, but the real bottleneck sits one tier down.
2. What this deep dive newly establishes
Core answer: across four clinical-stage AI-derived candidates, the primary endpoint is safety/pharmacology rather than efficacy, and every “positive/failed” verdict is decided in the secondary/exploratory efficacy tier — where no case cleanly attributes the outcome to AI.
- rentosertib (Insilico, TNIK, IPF) — CONFIRMED: primary endpoint = safety (met); FVC by arm (placebo −20.3 / 30 QD −27.0 / 30 BID +19.7 / 60 QD +98.4 mL) is secondary/exploratory; 16/71 (22.5%) discontinued (GI + mild liver-enzyme elevation), clustered in the higher-dose arms. AI timeline ≈18 months to lead nomination, ≈30 months to Phase 1 completion. TNIK is a novel target, so this is the frontier case where AI reached into target discovery — but design did not remove the hepatotoxicity tail, and efficacy is a signal, not a confirmed result. Arm-level ALT/AST precision figures and confirmed DILI cases are UNVERIFIED-precision (not disclosed in the summary sources); the Nature Medicine full text sits behind an access wall, read via summary cross-reference.
- EXS-21546 (Exscientia, A2A antagonist) — CONFIRMED: IGNITE-AI Phase 1/2 (RCC/NSCLC + anti-PD-1), discontinued in 2023 on the company’s own statement that adequate, sustained A2A target coverage within the therapeutic range was not achievable. This is a strategic discontinuation on pharmacology grounds, not an efficacy-readout failure; it accompanied a pipeline downsizing. It shows that even good AI design fails when the target biology is narrow — and A2A is an industry-common immuno-oncology target, not an AI-unique discovery. The exact month of discontinuation is UNVERIFIED.
- BEN-2293 (BenevolentAI, pan-Trk, atopic dermatitis) — CONFIRMED: n=91 (49 drug / 42 placebo), topical BID 28 days, top-line 2023-04-04. Primary endpoint = safety/tolerability (met, “safe and well-tolerated”); secondary EASI / NRS efficacy failed to reach significance in ITT; a post-hoc ≥20% BSA subgroup showed EASI signal (ITT p=0.0296, PP p=0.0427, day-28 interaction p=0.0237). One secondary source described the primary endpoint as EASI/NRS efficacy — refuted by the businesswire top-line (primary = safety; EASI/NRS are secondary). The lead failure was followed by large-scale restructuring.
- Balancing the positive side — REC-4881 (Recursion, allosteric MEK1/2, FAP) — CONFIRMED: TUPELO Phase 1b/2, 4 mg QD; at 12 weeks, 75% of evaluable had reduced total polyp burden, median −43% (n=12); at 25 weeks (12 weeks off-therapy), 82% (9/11) sustained reduction, median −53% vs baseline. FAP has no approved drug (high unmet need). But this is a single-arm, small-sample (evaluable n≈11–12) signal; the target (MEK1/2) is not novel; and the 1H26 FDA-engagement registration path is forward-looking (UNVERIFIED). Recursion frames this as “first clinical validation of Recursion OS” — a company narrative to hold at arm’s length.
- SGR-1505 (Schrödinger, MALT1) — CONFIRMED (inherited from Part 0): Phase 1 in B-cell malignancies, ORR 22% across 45 patients (EHA 2025) — an early activity signal on a known target.
3. Strengths and limits of the methodology
Strengths: each clinical event, primary-endpoint character, arm-level efficacy number and discontinuation reason is cross-checked against primary press, peer review and IR filings; vendor / company efficacy numbers are attributed and separated from confirmed facts. Limits (scope honesty): the efficacy signals for rentosertib, REC-4881 and SGR-1505 are all small-sample and short-term, and REC-4881 is single-arm (no placebo). The deeper lesson is that the endpoint tier dominates the headline: because the primary endpoint is safety, the same “met primary” fact coexists with opposite “positive/failed” narratives depending on the secondary efficacy direction and company communication. Attributing AI’s design contribution separately from target biology on a hard efficacy outcome is currently not possible for any of these cases.
4. Neighbouring domains
These cases feed Part 4 (attribution / success-rate integrity) three quantitative questions. (a) The denominator of success-rate tallies — whether discontinuations like EXS-21546 (pharmacology halt) and BEN-2293 (efficacy miss) are counted governs the survivorship bias in the widely cited ~40% Phase 2 success rate. (b) The definition of “AI-derived” — rentosertib (end-to-end, including target discovery) vs. REC-4881 / SGR-1505 (known targets, chemistry / patient-matching) vs. EXS-21546 (common target) represent entirely different grains of “AI contribution,” making the denominator unstable. (c) Quantifying the framing asymmetry — counting cases where the same safety primary endpoint yields divergent “positive/failed” headlines gives a measure of headline reliability. This connects to the series-wide lens (also in the computing and bio-FM parts): the real bottleneck sits next to the headline, not in it.
5. Commercialization and market context (TRL, actors — not stock guidance)
| Candidate (company) | Primary endpoint | Headline | Reality (primary vs. secondary) | Where AI contributed |
|---|---|---|---|---|
| rentosertib (Insilico, private) | safety — MET | “positive Phase 2a” | primary safety met; efficacy (FVC +98.4 mL) is secondary/exploratory; small n, 12 wk; 22.5% discontinuation | End-to-end (TNIK discovery + design). Frontier, but efficacy unconfirmed |
| EXS-21546 (Exscientia, absorbed by Recursion) | (Ph1/2, halted) | “discontinued” | NOT an efficacy failure — strategic halt on A2A target coverage / therapeutic index | Good design, but narrow target biology was the bottleneck |
| BEN-2293 (BenevolentAI, listed) | safety — MET | “flunks / misses Phase 2a” | primary safety met; efficacy (EASI/NRS) failed in ITT; only post-hoc ≥20% BSA signal | Knowledge-graph target discovery. Efficacy unverified |
| REC-4881 (Recursion, listed) | (Ph1b/2, polyp burden) | “positive” | polyp median −43% (12 wk) / −53% (25 wk); single-arm, evaluable n≈11–12 | Phenomics matching; target (MEK) known |
| SGR-1505 (Schrödinger, listed) | (Ph1, safety/activity) | “activity signal” | ORR 22% / 45 pts, early activity | FEP+ chemistry optimization; target (MALT1) known |
Deployment / TRL readiness: all five are clinical-stage assets whose efficacy readout is either secondary or not yet reported. Of four clinical-stage AI-derived candidates, essentially none had efficacy as a primary endpoint (a Phase 2a / Phase 1 characteristic), and there is no case where AI’s contribution has been separately shown to move a hard efficacy outcome — that quantitative attribution question is deferred to Part 4. No single-winner or stock conclusion is asserted; company quantitative claims (rentosertib FVC, REC-4881 polyp burden, Recursion “first OS validation”) are attributed as company / preprint claims and kept separate from the confirmed endpoint facts. Negative facts (discontinuations, hepatotoxicity signal) are stated factually and must not be read as stock implications.
6. The skeptic’s bottom line
- verified-clean: each clinical event, primary-endpoint character, arm-level efficacy figure and success/failure — cross-confirmed against primary press, peer review and IR (rentosertib primary=safety / FVC by arm / 16-71 discontinuation; EXS-21546 A2A therapeutic-index halt; BEN-2293 primary=safety met / EASI-NRS secondary ITT miss; REC-4881 polyp median −43% / −53%).
- proceed-with-caveats (conditional): (1) rentosertib, REC-4881 and SGR-1505 efficacy signals must carry their small-sample, short-term (and, for REC-4881, single-arm) conditions and must not be extended to confirmed efficacy; (2) BEN-2293’s ≥20% BSA signal must be labelled post-hoc; (3) listed-company descriptions (BenevolentAI, Recursion, Schrödinger) carry no stock implication and are separated into the investment layer.
- HOLD: (1) headlining rentosertib’s “positive Phase 2a” as an efficacy result (the primary endpoint was safety); (2) any “AI has validated clinical efficacy” completed narrative — zero approvals, all efficacy secondary; (3) framing EXS-21546 as an “efficacy failure” — it was a pharmacology (target-coverage) halt.
- Verdict (inherited): the firm view — “for AI-derived clinical readouts, the primary endpoint is safety/design; efficacy remains a secondary signal, and ‘positive/failed’ is decided by framing” — is robust (proceed-with-caveats). Escalated to skeptic (framing asymmetry, small samples, attribution) and to Principal (neutral frame for listed pipelines).
7. What to watch (falsifiable predictions)
- If rentosertib’s follow-up (Phase 2b / US, larger, placebo-controlled) promotes FVC to a primary efficacy endpoint and reproduces it statistically, the “efficacy is still only a secondary signal” verdict is refuted. Falsifiable by Insilico disclosure.
- If REC-4881 reproduces polyp-burden reduction significantly under placebo control at expanded sample and enters an FDA registration path, the “single-arm small-sample signal” reservation is relaxed. Falsifiable by Recursion IR.
- If, over the next 24 months, two or more AI-derived candidates pass a Phase 2/3 with efficacy set as the primary endpoint, the “AI trials keep efficacy in the secondary tier” pattern is refuted. Falsifiable by trial registries / disclosures.
References
- Zhavoronkov, Alex, et al. 2025. “rentosertib (INS018_055), an AI-discovered TNIK inhibitor for IPF: GENESIS-IPF Phase 2a.” Nature Medicine 31:2602–2610, 2025-06 (PMID 40461817). https://www.nature.com/articles/s41591-025-03743-2 (DOI 10.1038/s41591-025-03743-2; access-walled full text, read via summary cross-reference; PMC meta: PMC12353801)
- Insilico Medicine / PR Newswire / EurekAlert. 2025. “Insilico’s rentosertib Phase 2a published in Nature Medicine.” 2025-06. eurekalert.org/news-releases/1086096 (arm-level FVC / discontinuation / AI timeline summary: intuitionlabs.ai)
- Fierce Biotech. 2023. “Exscientia chops down pipeline; EXS-21546 discontinued.” fiercebiotech.com (IGNITE-AI Phase 1/2, RCC/NSCLC + anti-PD-1: Business Wire 20221128005185; clinicaltrialsarena.com)
- BenevolentAI. 2023. “BEN-2293 Phase 2a top-line results.” Business Wire, 2023-04-04. businesswire.com/news/home/20230404006090 (EASI/NRS miss, post-hoc ≥20% BSA: dermatologytimes.com; clinicaltrialsarena.com)
- Recursion Pharmaceuticals. 2025. “REC-4881 TUPELO Phase 1b/2 polyp-burden update (FAP).” GlobeNewswire / IR, 2025-05-04. ir.recursion.com
- Schrödinger. 2025. “SGR-1505 (MALT1) Phase 1 results.” EHA 2025 (inherited from Part 0). ir.schrodinger.com
Source knowledge asset: knowledge-base/deep-dives/ai-drug-clinical-readout/part3-readout-cases.md (generated 2026-07-12, VERIFIED — 12 confirmed / 1 refuted / 3 unverified). This draft inherits the figures and source attributions of the original part and creates no new figures or numbers.
Disclosure
This post is for information only and is not investment advice. The author holds no position in, and no financial interest in, the entities mentioned (Insilico Medicine, Exscientia, BenevolentAI, Recursion, Schrödinger, Evotec, Takeda and others).
COI note: This document describes listed AI-drug developers (Recursion, Schrödinger, BenevolentAI) and an absorbed pipeline (Exscientia) as well as private companies (Insilico Medicine) in a factual, neutral technology context, with no buy/sell implication. Company / preprint quantitative claims (rentosertib FVC +98.4 mL, REC-4881 polyp burden, SGR-1505 ORR, Recursion “first OS validation”) are attributed as company / preprint claims and kept separate from the confirmed primary-endpoint facts; they are not presented as demonstrated efficacy. Negative facts (discontinuations, a hepatotoxicity signal, an efficacy miss) are stated factually and must not be read as negative characterizations of, or stock implications for, any single actor.
Leave a comment