The AI protein-design landscape — structure prediction is near-solved at the headline, but generative de novo design’s real bottleneck is the in-silico-metric-to-wet-lab-function gap

Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice, not medical advice. All success-rate, hit-rate and efficiency figures are attributed to the reporting paper or company; some are company blogs or preprints rather than peer-reviewed data (noted inline). In-silico design metrics (pLDDT, ipTM, self-consistency) are kept strictly separate from wet-lab-validated function (expression, binding, catalysis). Quotes are 150 characters or fewer.

The 30-second version

  • What. AI protein work splits into two waves. Structure prediction — AlphaFold2 (2021), AlphaFold3 (2024, co-folding) and ESMFold (2023) — has, at the headline, largely solved sequence-to-3D-structure for single static structures; the 2024 Nobel Prize in Chemistry recognized this (Hassabis/Jumper for prediction, Baker for computational design). Generative de novo design — RFdiffusion (backbone generation), ProteinMPNN (sequence/inverse folding) and Chroma (programmable generation) — now builds proteins that do not exist in nature and has validated binders and enzymes in the wet lab (crystallography/cryo-EM), so it is past a pure demo.
  • So what. The headline (single-structure prediction) is close to solved; the real bottleneck is elsewhere — the gap between high in-silico design metrics and low wet-lab experimental success. Generative pipelines score impressively on pLDDT/ipTM/AlphaFold2 self-consistency, yet a large share of filter-passing candidates still fail on expression, non-specific binding or undetectable affinity. The central falsifiable question is whether generative design is a discovery engine or a low-hit-rate hypothesis generator — and the evidence leans toward a metric-reality gap (self-consistency predicts “folds,” not “function”). As of late 2025 there are zero approved de novo AI-designed protein drugs; every figure above is a lab or in-silico metric, not a clinical outcome.
  • Now what. The decisive test is independent, prospective, leakage-controlled experimental success rates from non-originator groups, plus whether in-silico-metric-to-wet-lab correlation generalizes to novel targets. An independent evaluation across six targets (bioRxiv 2025) reported many RFdiffusion binders failing (low expression, non-specific, undetectable affinity), and the RFdiffusion-antibody paper (Nature 2025) itself states success rates are “currently low.” Until those readouts arrive, “AI has solved protein design / is a discovery engine” is an overstatement to be avoided.

The five-minute read

Two waves — prediction is near-solved, generative design is validated but rate-limited

AI protein work reproduces the firm’s recurring lens — “the headline is the starting point; the real bottleneck is in the outcome layer” — more cleanly than almost any other AI-times-Bio interface. The first wave, structure prediction, has largely solved sequence-to-structure at a practical level: AlphaFold2 broke through at CASP14, AlphaFold3 added co-folding (protein plus ligand, nucleic acid and ions), and ESMFold predicts fast without alignments. What remains unsolved sits outside the single static structure — disorder, conformational ensembles, protein-protein interfaces and ligands (covered in Part 1). The second wave, generative de novo design, builds proteins from scratch: RFdiffusion generates backbones, ProteinMPNN designs sequences (inverse folding), and Chroma generates under programmable constraints. These have produced real binders and enzymes confirmed by crystallography and cryo-EM — this is past a demo (it clears analysis-standards Section 2).

The headline is near-solved; the bottleneck is metric-to-function translation

The pipeline is: generate thousands to tens of thousands of candidates, filter them in-silico (AlphaFold2 self-consistency — does the designed sequence re-fold to the intended backbone?), then make only a small subset (dozens to hundreds). This filtering raised success rates substantially — the de novo binder work (Cao 2022) reports that AF2/RF filtering improved success roughly tenfold — but filter-passing candidates still drop out on expression failure, non-specific binding and undetectable affinity. So the reported “success rate” is relative to candidates tested, not candidates generated, and swings widely by target and epitope. The RFdiffusion-antibody paper (Nature 2025) itself states success rates are “currently low”; antibodies stack CDR-loop, humanization and expression problems on top of the binder bottleneck. The honest reading is that in-silico filtering raises hit rate but does not remove the wet-lab bottleneck.

Layer Status (2026-07) Verdict
Structure prediction (single structure) AF2 (Nature 2021), AF3 co-folding (Nature 2024), ESMFold (Science 2023); 2024 Nobel Chemistry Near-solved (demonstrated)
Prediction, unsolved edge (disorder, ensemble, PPI, ligand) Baseline advantage disappears on antibody/nucleic-acid; data-leakage debate (bio-FM Part 1) Unsolved
Design engine reproducibility (code/weights) RFdiffusion, ProteinMPNN, Chroma all open on GitHub Open, a strength
De novo binder / enzyme wet-lab validation Cao 2022 (hundreds characterized), RFdiffusion cryo-EM match; serine hydrolase kcat/Km up to 2.2×10^5 M^-1 s^-1 (below natural) Validated (not a demo)
Generative-design experimental success rate Independent 6-target eval — many failures (bioRxiv 2025); RFdiffusion-antibody “currently low” (Nature 2025) Low — bottleneck
In-silico metric → wet-lab translation AF2 filter ~10x improvement, but filter-passing candidates still drop out; target-dependent Partial, target-dependent
Approved de novo AI-designed protein drug Zero (end of 2025); Absci ABS-101 at Phase 1 None yet
“The headline is near-solved” does not mean “generative design is a proven, high-success discovery engine.” All success-rate and efficiency figures are attributed to the reporting paper or company, and are labeled peer-reviewed (Nature/Science), preprint (bioRxiv) or company blog. In-silico metrics (pLDDT, ipTM, AF2 self-consistency) are surrogates, not wet-lab-validated function. These tools were reported on different benchmarks, targets and metrics — the rows are not head-to-head comparisons.

Deep dive

1. Background — three layers: prediction, design engine, function

The landscape resolves into three layers, and the key is how they chain together.

  • Layer 1 — Structure prediction: sequence to 3D structure, including co-folding (protein plus ligand plus nucleic acid). Tools: AlphaFold2/3, ESMFold, RoseTTAFold. Near-solved for single structures; unsolved for disorder, ensembles, PPI and ligands (Part 1).
  • Layer 2 — Design engine: backbone generation plus sequence design (inverse folding). Tools: RFdiffusion and RFdiffusion-AA, Chroma, ProteinMPNN and LigandMPNN, Genie. De novo generation, with in-silico filters selecting candidates (Part 2).
  • Layer 3 — Function: real binders, enzymes, antibodies and therapeutics. Examples: de novo binder (Cao 2022), de novo serine hydrolase (2025), RFdiffusion-antibody (2025). Wet-lab success rate is the bottleneck (Parts 3-4).

The core structure: Layer 2’s output is filtered in-silico by Layer 1 (self-consistency — does the designed sequence re-fold to the original backbone?), then made in Layer 3. This pipeline raised success rates substantially (Cao 2022: AF2/RF filtering ~10x), but many filter-passing candidates still fail in the wet lab on expression, non-specific binding or undetectable affinity (Section on the bottleneck below). In-silico filtering lifts hit rate; it does not remove the wet-lab bottleneck.

2. What this landscape establishes — tools, contributions and publication venue (attributed)

Principle: publication venue is labeled peer-reviewed (Nature/Science) vs preprint (bioRxiv) vs company; success-rate and efficiency numbers are attributed to the reporting source.

  • AlphaFold2 (Google DeepMind), prediction — CASP14 breakthrough on sequence-to-structure, ~90%-accuracy class; 2M+ users across 190 countries (as of 2024-10, Nobel material). Jumper et al., Nature 2021 (peer-reviewed).
  • AlphaFold3 (DeepMind/Isomorphic), prediction — co-folding: unified prediction of protein plus ligand plus nucleic acid plus ion complexes. Abramson et al., Nature 2024.
  • ESMFold / ESM-2 (Meta to EvolutionaryScale), prediction — ~15B-parameter protein language model, fast alignment-free prediction; ESM Metagenomic Atlas of 600M+ proteins. Lin et al., Science 2023.
  • RFdiffusion (Baker Lab/IPD), design — RoseTTAFold retrained as a denoising diffusion model to generate backbones; an influenza-HA binder matched the design model closely by cryo-EM. Watson et al., Nature 2023 (this is the base, non-all-atom RFdiffusion).
  • ProteinMPNN (Baker Lab), design — deep-learning inverse folding (structure to sequence), higher sequence recovery and lower compute than physics-based methods. Dauparas et al., Science 2022.
  • Chroma (Generate Biomedicines), design — programmable generative model conditioned on symmetry, shape, semantics and natural language; near-linear all-atom scaling. Ingraham et al., Nature 2023 (vol 623).
  • De novo binder (Baker Lab), function — binders designed from target structure alone, hundreds characterized experimentally; AF2/RF filtering reported ~10x higher success. Cao et al., Nature 2022 (605:551-560).
  • De novo serine hydrolase (Baker Lab), function — enzyme designed with RFdiffusion plus an active-site ensemble, kcat/Km up to 2.2×10^5 M^-1 s^-1 (a fold distinct from natural enzymes; activity still on the low side). Science 2025.
  • RFdiffusion-antibody (Baker Lab), function — first structural validation of de novo antibodies; the paper states success rates are “currently low.” Bennett et al., Nature 2025.

Layer character: prediction (1-3) is consistent with bio-FM Part 1 — AF3 leads most tasks but loses its baseline advantage on antibodies and nucleic acids, amid a data-leakage debate; this series inherits that conclusion and extends it to design. Design engines (4-6) — diffusion (backbone) plus inverse folding (sequence) is effectively the standard pipeline, and all are open code/weights on GitHub (a reproducibility strength). Function (7-9) reaches wet-lab validation (crystallography/cryo-EM), so it is not a demo — but success rate and activity magnitude are the bottleneck (the serine hydrolase is on the low-activity side versus natural enzymes; the antibody paper labels its success rate “low”). These tools were reported on different benchmarks, targets and metrics — cross-tool ranking is not warranted without a same-benchmark head-to-head (a discipline inherited from bio-FM).

3. The central science question — discovery engine, or low-hit-rate hypothesis generator? (falsifiable)

Generative de novo design scores impressively on in-silico metrics (pLDDT, ipTM, AF2 self-consistency, pae). The falsifiable question at the center of this series: does experimental success rate clear the bar to call it a discovery engine, or is it still a low-wet-lab-hit-rate hypothesis generator? Three hypotheses, with the observation that would falsify each:

  • (a) Discovery engine — in-silico filters lifted hit rate to a practical level. Support: Cao 2022 (~10x improvement), RFdiffusion binder near-matching the design model by cryo-EM, serine hydrolase yielding an active enzyme on a novel fold. Falsified if independent (non-originator) labs reproduce low hit rates across many novel targets under the same protocol (Parts 3-4).
  • (b) Low-hit-rate hypothesis generator — in-silico scores do not translate to the wet lab. Support: an independent evaluation (bioRxiv 2025.02.07.636769) found many RFdiffusion binders across six targets failing (low expression, non-specific, undetectable affinity); the RFdiffusion-antibody paper states success is “currently low.” Falsified if reproducible double-digit-percent hit rates appear across many novel targets, from several independent groups, under leakage-controlled prospective design (Part 4).
  • (c) Metric-reality gap — surrogate metrics such as self-consistency are themselves flawed. Support: this is isomorphic to bio-FM Part 1’s leakage/memorization signals; the AF2 self-consistency used as a design filter may predict “folding” but not “function” (binding, catalysis). Falsified if the in-silico-metric-to-wet-lab correlation holds strongly (predictively) on novel targets; strengthened if the correlation is confined to targets near the training set (Part 4).

Current provisional position: none of the three is excluded. The prediction axis leans toward (a) and is close to settled; the design axis cannot yet exclude (b) or (c), and the post-series work (below) shifts weight toward (c), the metric-reality gap. The decider is independent-group prospective, leakage-controlled experimental success rates and whether the in-silico-to-wet-lab correlation generalizes to novel targets. Part 0 does not assert any one and keeps all three falsifiable.

4. In-silico metric vs wet-lab-validated function — anatomy of the bottleneck

Firm discipline (analysis-standards Section 2): separate the surrogate metric from validated function. The mapping is direct.

In-silico surrogate (design stage) Wet-lab-validated function (outcome layer)
pLDDT / pTM / ipTM (prediction confidence) Recombinant expression success
AF2 self-consistency (re-fold RMSD) Target-specific binding and affinity (Kd)
pae (interface error estimate) Catalytic activity (kcat/Km) and functional assay
Design diversity / novelty score Structural confirmation (X-ray / cryo-EM)
The bottleneck’s substance: pipelines generate thousands to tens of thousands of candidates in-silico, filter, then test only a small subset (dozens to hundreds). Among that subset, many drop out on (i) expression failure, (ii) non-specific binding or (iii) undetectable affinity (the six-target observation in bioRxiv 2025.02.07.636769). The reported “success rate” is relative to candidates tested, not candidates generated, and swings widely by target, antigen and epitope. Even the RFdiffusion-antibody Nature 2025 paper concedes low success — antibodies stack CDR-loop, humanization and expression problems on top of the binder bottleneck. This is structurally isomorphic to the neuro series’ surrogate-marker (amyloid clearance) to hard-outcome (cognition) translation bottleneck: here the surrogate is the in-silico score, and the hard outcome is wet-lab-validated function (ultimately clinical benefit).

5. Commercialization and competitive context (TRL, related companies)

  • Maturity (TRL frame): prediction is settled and validated; generative design is wet-lab-validated (not a demo) but rate-limited on experimental success, activity magnitude and translation. The gating layers are success rate and metric-to-function translation, not scientific novelty.
  • Listed: Alphabet (Google DeepMind; Isomorphic Labs parent), Meta (EvolutionaryScale origin), NVIDIA, Recursion (RXRX), Schrodinger (SDGR), Absci (ABSI). Absci’s ABS-101 (TL1A, IBD) is at Phase 1, with ABS-201 entering — the clinical front is AI-optimized antibodies, not fully de novo backbones.
  • Private: Generate Biomedicines, Xaira Therapeutics (2024 launch, over $1B; Baker/Tessier-Lavigne), EvolutionaryScale (launch $142M, ESM3), Profluent (Series B $106M, $150M+ cumulative), Chai Discovery, Cradle. Chai’s Series B (~$130M), Cradle funding, and Generate’s IPO size/timing are [unverified] and deferred to Part 5.
  • The commercial reality: roughly $10^9 of capital (Xaira over $1B; Isomorphic $600M-plus with milestones) versus zero approved de novo AI-designed protein drugs at end of 2025. The clinical front is AI-optimized antibodies (Generate Phase 3, Absci Phase 1), not fully de novo backbones — an illustration at commercial scale of the “AI as hypothesis generator” thesis (convergence L-S01).
  • Company statements are limited to neutral, source-attributed description; vendor blog/preprint success-rate and funding figures are attributed and await independent verification. These are not buy/sell signals for any security.

6. The skeptic’s bottom line

  • In-silico metric is not wet-lab hit rate: an independent six-target evaluation reported many failures; high self-consistency does not guarantee expression, binding or catalysis.
  • Denominator caution: reported success rates are relative to candidates tested, not candidates generated — the framing inflates apparent performance.
  • No cross-tool ranking: tools were reported on different benchmarks and targets; without a same-benchmark head-to-head, ranking is not warranted.
  • Zero approved products: there are no marketed de novo AI-designed protein drugs as of late 2025 — the final outcome-layer signal.
  • Refuted overstatement: the claim that “generative de novo design is already a high-success, validated discovery engine” is an overstatement. The independent evaluation reports low wet-lab success, the antibody paper concedes “currently low,” and approved products are zero. In-silico scores must not be read as experimental success.
  • Neutral-framing note: vendor claims of large improvements (“16-20% hit rate,” “100x improvement,” “20 candidates per target”) arriving through non-peer-reviewed channels are treated as benchmark-hacking / data-leakage suspects and attributed, not adopted.

7. What to watch (falsifiable)

  • P1: if independent (non-originator) groups report reproducible double-digit-percent experimental success across many novel targets under prospective, leakage-controlled design, hypothesis (a) discovery engine is strengthened; if low expression / non-specific / undetectable failures reproduce, (b) is strengthened. (Parts 3-4)
  • P2: if the in-silico-metric-to-wet-lab correlation (AF2 self-consistency and similar) holds predictively beyond near-training-set targets to novel targets, the (c) metric-reality gap weakens; if it is confined to near targets, (c) strengthens. (Part 4, inheriting bio-FM Part 4 leakage methodology.)
  • P3: if a de novo AI-designed protein (binder/enzyme/antibody) enters the clinic (Phase 1 to efficacy readout) and shows a benefit signal, that is the decisive case for a “hypothesis generator to discovery engine” shift. With approved products at zero, the first efficacy readout (Absci/Xaira/Generate pipelines) is the touchstone. (Part 5, linked to the ai-drug-clinical-readout axis.)
  • Post-series update (Part 1-5 integration): the central question shifts weight toward (c) metric-reality gap — independent six-target reproduction failure (bioRxiv 636769), 3,766-binder precision spanning 0.1-1.0 by target, and a peer-reviewed finding (Protein Science 2026) that refolding metrics carry false positives together indicate self-consistency is “a surrogate for folding, not for function.” Separately, BindCraft (Nature 2025) reported one-shot experimental success of 10-100% (target-dependent) — the top-line generative-binder result, still target-dependent and not yet independently reproduced across many sites.

References

Disclosure

This post is for information only and is not investment advice, and not medical advice. Treatment decisions should always be made with your own clinician.

COI note: this post describes listed companies (Alphabet [Google DeepMind; Isomorphic Labs parent], Meta [EvolutionaryScale origin], NVIDIA, Recursion RXRX, Schrodinger SDGR, Absci ABSI) and private companies (Xaira Therapeutics, Generate Biomedicines, Profluent, Chai Discovery, Cradle) in a descriptive, neutral context. Every success-rate, hit-rate and efficiency figure is attributed to the reporting paper or company, and publication venue is labeled peer-reviewed (Nature/Science) vs preprint (bioRxiv) vs company blog. In-silico design metrics (pLDDT, ipTM, AF2 self-consistency) are stated as surrogates and kept separate from wet-lab-validated function. Vendor blog and preprint claims (for example “16-20% hit rate” or “100x improvement”) are attributed and flagged as awaiting independent verification. There are zero approved de novo AI-designed protein drugs as of late 2025. Competitive and capability statements are factual, neutral descriptions and are not buy/sell implications for any security. The author holds no position in, and has no financial interest in, the companies named.