Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice, not medical advice. This is the bottleneck layer of the AI-protein-design series, where its central question resolves. Self-reported in-silico success rates and improvement factors (from original authors and vendors) are separated throughout from independent, non-author wet-lab replication; every quantitative claim is attributed to its source, and some summary-only figures are marked unverified. Quotes are kept under 150 characters.
The 30-second version
- What. Generative de novo design pipelines score candidates with in-silico metrics (pLDDT, scRMSD, AF2 self-consistency, ipAE/ipTM), then test a small subset in the wet lab. This layer asks how well those scores translate into real function (expression, specific binding, affinity). The answer is: only target-dependently. An independent, non-Baker 6-target replication (bioRxiv 2025.02.07.636769) found many eukaryotic-target binders fail — low expression, non-specific binding, undetectable affinity — and a 3,766-binder meta-analysis (bioRxiv 670059) found in-silico predictivity swinging with the target (precision 0.1–1.0). These are preprints, not peer-reviewed.
- So what. The self-consistency / refolding metric that filters candidates is, in a peer-reviewed critique (Protein Science 2026), “a proxy for folding, not for function” — it produces false positives from evolutionary information and false negatives from flexibility. So a high in-silico score is not a wet-lab hit. The reported “16–90% success” versus a fraction of a percent depends almost entirely on one hidden choice: the denominator (candidates tested vs candidates generated).
- Now what. The decisive missing piece is an honest success-rate benchmark — prospective, leakage-controlled, with the denominator specified, failures fully reported, and replicated by independent labs. Such benchmarks are rare by construction (publication bias, denominator flexibility, IP incentives, target variability). Until one exists, “AI has solved protein design / it is a discovery engine” is an overstatement — though the independent failure data does not disprove the design paradigm either. There are zero marketed, de novo AI-designed protein drugs as of late 2025.
The five-minute read
How a “high in-silico score” hides a “low wet-lab hit rate”
A generative pipeline typically generates thousands to tens of thousands of candidates, filters them with in-silico metrics, and tests only a few dozen to a few hundred. The headline success rate is almost always reported relative to the candidates that reached the bench (the small set that expressed and purified), not relative to the candidates that were generated (the raw designs). That single choice of denominator is what separates a reported “16–90% success” from a fraction of a percent. The firm’s recurring lens — “the headline is the starting point; the real bottleneck is in the outcome layer” — bites hardest here.
The claim of this layer is fourfold: generative-design in-silico metrics (1) predict wet-lab success only in a target-dependent way, (2) rest on a refolding metric that generates false positives from evolutionary information and false negatives from structural flexibility, (3) fail on independent (non-author) replication through low expression, non-specific binding and undetectable affinity, and (4) are rarely accompanied by an honest success-rate benchmark (prospective, leakage-controlled, denominator = generated candidates, failures reported). These four levers overlap to bias the “high in-silico score = discovery engine” narrative upward. This is the design-axis re-run of the benchmark-versus-deployment gap documented for predictive foundation models.
The separation rule: self-report versus independent replication
Throughout, self-reported success rates from original authors and vendors — for example Cao et al. 2022 (“AF2 filter ~10x”), the 2023 Nature Communications deep-learning binder paper (“AF2 filter 8–30x over physically-based filtering”), and BindCraft’s one-shot design claim — are kept in a separate evidence tier from independent wet-lab replication such as the 6-target study below. The former are labelled [self-report / original author]; the latter [independent]. They cannot be placed on the same scale, because the self-report numbers may be the product of the originating lab’s optimized hands, easy targets, and a bench-tested denominator.
| What the in-silico filter sees | What it does not see (the outcome-layer physical wall) |
|---|---|
| pLDDT, scRMSD (fold confidence) | Recombinant expression success / yield — many of the 6-target binders failed here (low expression) |
| ipAE, ipTM, ipSAE (interface confidence) | Solubility / aggregation — inclusion bodies, precipitation |
| Rosetta ΔG (computed binding energy) | Non-specific binding — observed in the independent replication |
| Design novelty / diversity | Measured affinity (Kd) — undetectable for most eukaryotic targets in the replication |
Deep dive
1. Independent (non-author) replication — RFdiffusion’s low hit rate across 6 targets
This is the series’ central disconfirming dataset. Jiang, Li, Guo, Wei and Wu, “RFdiffusion Exhibits Low Success Rate in De Novo Design of Functional Protein Binders for Biochemical Detection,” bioRxiv 2025.02.07.636769 (posted 2025-02-08, a preprint from an independent, non-Baker group). [independent / medium confidence — abstract layer and PDF search snippets confirmed; full text not read behind the auth wall.]
- Six design targets: Strep-Tag II (a peptide tag) plus five eukaryotic proteins (STAT3, FGF4, EGF, PDGF-BB, CD4). Five binders were designed per target and experimentally characterized — so the denominator is 30 designs, a prospective small-batch design in which every candidate is tested, not a fraction relative to a pre-filtered subset.
- Result: two Strep-Tag II binders outperformed streptavidin on Western blot but did not reach the sensitivity of an anti-Strep-Tag II antibody. The binders for the other five eukaryotic targets failed through low expression, non-specific binding and undetectable affinity.
- Authors’ conclusion (paraphrased, in place of a quote): the pipeline produced structurally diverse candidates, but low-affinity designs and unstable recombinant expression constrained the success rate.
Firm read: this is not a demo but an independent, prospective validation, and it belongs to a separate evidence tier from the originating lab’s (Baker Lab’s) self-report. Two caveats: n=5 per target is a small sample, and this group’s design/expression protocol is not guaranteed to match the Baker Lab’s optimization (operator and facility effects cannot be excluded). So the precise statement is not “RFdiffusion fails” but “in non-originating hands, in the practical context of detection binders, small-batch success was low.” This strengthens hypothesis (b) — low-hit-rate hypothesis generator — without fully excluding (a). Cross-tool superiority claims are out of bounds: this is not a same-benchmark head-to-head, so it cannot be read as “RFdiffusion < BindCraft / other tools.”
2. How well do in-silico metrics predict the wet lab? Target-dependent precision, 0.1–1.0
The quantitative anchor for the metric–reality gap. Overath, Rygaard, Jacobsen et al., “Predicting Experimental Success in De Novo Binder Design: A Meta-Analysis of 3,766 Experimentally Characterised Binders,” bioRxiv 2025.08.14.670059 (posted 2025-08, preprint). [independent / medium confidence — abstract layer confirmed; some quantitative detail behind the auth wall (403), not read.]
- Scale: 3,766 experimentally characterized binders pooled across 15 structurally diverse targets. Each binder–target complex was re-predicted with AF2 (initial-guess / ColabFold), AF3 and Boltz-1, extracting 200+ structural, energetic and confidence features per design.
- Finding 1 — best predictor: an AF3-derived metric, ipSAE (interaction Score from Aligned Errors), outperformed the commonly used ipAE and ipTM, with roughly 1.4x the average precision of ipAE. Combining orthogonal physicochemical metrics (Rosetta ΔG / ΔSASA, interface shape complementarity) improved prediction further.
- Finding 2 (the core of the firm’s lens) — large target-to-target variation: predictive performance (precision) swung from 0.1 to 1.0 depending on the target. That is, no in-silico metric predicts wet-lab success uniformly across new targets.
Implication: this quantifies the metric–reality gap. Self-consistency / interface-confidence metrics are predictive on training-adjacent, easy targets but collapse toward precision 0.1 on hard or novel targets. The lesson mirrors the predictive-FM finding that leakage must be measured per benchmark rather than assumed: a metric’s predictive power is something to measure per target, not to assume. When a vendor says “our metric predicts success,” the firm’s question is: on which target distribution was that precision measured?
3. The self-consistency (refolding) metric’s own defects — false positives and false negatives
The AF2 self-consistency / refolding pipeline (does the redesigned sequence refold to the original backbone? — metric = pLDDT + scRMSD) is dissected in the most recent peer-reviewed evidence: Korbeld, Viliuga and Fürst, “Limitations of the refolding pipeline for de novo protein design,” bioRxiv 2025.12.09.693122 (posted 2025-12-11), published in Protein Science (Wiley) 2026, peer-reviewed, DOI 10.1002/pro.70613, with code and data released on GitHub (kt-korbeld/Limitations-refolding-pipeline-data). [confirmed / high — peer-reviewed publication plus open code/data; some quantitative detail per summary-layer snippets.]
- False-positive axis — evolutionary-information contamination: evolutionary information (MSA / homology) blurs the folding model’s sequence–structure fit judgment. A design that shares homology with natural sequences can pass the refolding metric regardless of its actual design quality (AF2 tends to be over-confident on poor designs in MSA mode) — biasing the in-silico success rate upward.
- False-negative axis — flexibility sensitivity: the refolding metric is sensitive to structural flexibility, so it can reject functionally sound but flexible designs on scRMSD.
- Arbitrariness of the designability threshold: even at a lenient threshold (pLDDT ≥ 70 and scRMSD ≤ 2.0 Å), only about 40% of the shortest natural sequences (≤ 50 aa) pass — the boundary of “designable” wobbles strongly with length and flexibility. [Quantitative detail per summary-layer snippet; re-check against full text advised — marked unverified.]
Firm read: this is the strongest support for the metric–reality-gap hypothesis, because it is peer-reviewed with open code. In the authors’ framing, self-consistency is “a proxy for folding, not for function” — and even folding itself is subject to false positives from evolutionary contamination. It is the design-filter re-run of the predictive-FM signal that a confidence threshold shifts once contamination is removed. Passing the filter is not, in itself, a wet-lab success.
4. Expression, solubility, aggregation — the physical wall behind the filter
The physical wall of the outcome layer that in-silico metrics do not touch — and it is exactly where the top failure modes in the 6-target study sit (see the five-minute-read table). Most vendor/original pipelines report a success rate with the small set that survived expression and purification as the denominator. But if, as in the independent replication, expression-stage failure is a leading cause of dropout, then “success rate relative to bench-tested candidates” systematically overestimates “success rate relative to generated candidates.” Expression, solubility and aggregation are governed by sequence physicochemistry (hydrophobic patches, pI, disulfides); inverse folding (ProteinMPNN) partly improves them, but along a different axis from the in-silico fold score.
5. Benchmark contamination and train–test leakage — the PDB-proximity problem
How the predictive-FM benchmark levers re-appear on the design axis.
- PDB overlap / near-neighbour leakage: RFdiffusion, ProteinMPNN and AF2 are all trained on the PDB. If the design/evaluation backbone is evolutionarily and structurally close to a training structure, self-consistency “success” can be a reflection of memorization — the evolutionary-information false positive of section 3 is the microscopic mechanism of this leakage. In the predictive-FM analog, random PPI splits showed leakage up to 86%, and a large fraction of genomics papers gained AUROC from feature-selection leakage. [independent / high, inherited.]
- Absent temporal split: design benchmarks rarely hold out on structures newer than a fixed cutoff. If the re-prediction models (AF2/AF3/Boltz-1) in the section-2 meta-analysis may already have seen the binder sequences/targets, the upper end of precision (near 1.0 on easy targets) may be a near-neighbour reflection.
- Missing strong baseline: the honest control on the design axis is the specialized baseline an expert can build in half a day — physically-based filters (Rosetta), or the classical physical metrics (Rosetta ΔG / ΔSASA, shape complementarity) that section 2 showed. The finding that these improve prediction when combined orthogonally with deep-learning metrics suggests any deep-learning-alone superiority may be overstated.
- Metric cherry-picking: measuring the success rate on easy targets (Strep-Tag, helical peptides) and headlining it hides the section-2 “target-to-target precision 0.1–1.0” variation behind an average.
6. What an honest wet-lab success-rate benchmark should look like — and why it is rare
Definition (firm standard): to judge falsifiably whether generative design is a discovery engine, a success-rate benchmark should satisfy the following seven points (the design-axis specialization of the predictive-FM checklist).
- State the denominator: is the denominator the generated (raw) candidates or the small set that reached the bench? Report both (pass rates at each funnel stage: generation → in-silico filter → expression → purification → binding → affinity).
- Prospective, target-diverse: pre-registered, multiple novel targets (not skewed to easy tags — include eukaryotic PPI, membrane proteins, antigens), not post-hoc cherry-picking.
- Leakage control: design/evaluation backbones controlled for similarity / temporal split against the training PDB, with the leakage rate measured and reported.
- Independent, multi-site: ≥ 2 independent labs outside the originating group reproduce the same protocol (separating operator and facility effects).
- Report all failures: low expression, non-specific binding and undetectable affinity, as numbers (blocking publication bias — negatives included).
- Tiered functional validation: not just binding (SPR/BLI Kd) but structural confirmation (X-ray / cryo-EM) and a function assay (kcat/Km for enzymes). Passing self-consistency does not count as success.
- Uncertainty: confidence intervals, multiple batches, statistical tests (no single point estimate).
Why such benchmarks are rare (the structure of publication bias):
- Failure under-reporting: the negatives of a design campaign (expression failures, undetectable binders) have low publication incentive — an independent low-hit-rate report like the section-1 study is itself scarce. When only successes reach the literature, the field-wide success rate is overstated.
- Denominator flexibility: vendors quote “relative to bench-tested candidates,” skeptics quote “relative to generated candidates” — there is no standard denominator.
- Competition / IP incentives: commercial platforms treat their success rate as a marketing asset and have weak incentive to disclose all failures (the non-peer-reviewed the channel, the lower the tier). This is the design-axis re-run of “the benchmark as a marketing asset.”
- Target variability: the section-2 precision of 0.1–1.0 means a single success-rate number is itself misleading — an honest report needs a per-target-distribution decomposition, which is costly and slow.
Contrasting self-report claims (separately labelled): BindCraft (“One-shot design of functional protein binders,” Nature 2025) self-reports a high experimental success rate from one-shot design; “Improving de novo protein binder design with deep learning” (Nature Communications 2023) reports SC50 < 4 µM as success, 1–17 successes per target, and an AF2 filter 8–30x over physically-based filtering; Cao et al. 2022 (Nature 605) reports an AF2/RF filter ~10x improvement. [all self-report / original author — kept separate from the independent-replication tier; BindCraft’s quantitative success rate is marked unverified.] These numbers may be the product of the originating lab’s optimized hands, easy targets and a bench-tested denominator, so they cannot be placed on the same scale as section 1 (independent, low hit rate) or section 2 (target variation). The firm files the former as [claim] and the latter as [evidence].
7. Verdict — where the three hypotheses sit on the data
| Hypothesis | Data at this layer | Current position | Evidence tier |
|---|---|---|---|
| (a) Discovery engine | Cao ~10x; Nat Commun 8–30x; BindCraft self-report | Strong in the originating lab, on easy targets — but self-report | self-report / original author |
| (b) Low-hit-rate hypothesis generator | Independent 6-target (636769): many failures via low expression / non-specific / undetectable | Strengthened in the independent, practical context (small-sample caveat) | independent / medium |
| (c) Metric–reality gap | Meta-analysis precision 0.1–1.0 (670059); refolding false positive/negative (693122, peer-reviewed) | Quantified and confirmed (target-dependent; evolutionary-info false positives) | independent + peer-reviewed / high |
| Marketed, approved de novo AI-designed protein drug | Inherited from Part 0 | Zero (end of 2025) — the final outcome-layer signal | confirmed |
8. The skeptic’s bottom line (six caveats)
- In-silico success rate ≠ wet-lab hit rate — many of the independent 6-target designs failed.
- Reported success rate is denominator-driven (bench-tested vs generated candidates); there is no standard denominator.
- Self-consistency / refolding is a proxy for folding, not for function — evolutionary-information false positives, flexibility false negatives (peer-reviewed).
- Predictive precision is target-dependent (0.1–1.0) — a single success-rate number is itself misleading.
- Separate self-report from independent replication (Cao / Nat Commun / BindCraft vs 636769), and state the failure under-reporting (publication bias).
- Zero marketed AI-designed protein drugs — the final outcome-layer signal. These six caveats are surfaced at Tier-1 level.
9. What to watch (falsifiable)
- P1: if a standard success-rate benchmark meeting the seven points of section 6 (denominator stated, prospective, leakage-controlled, independent multi-site, all failures reported) becomes common, a substantial share of current self-report success rates will fall significantly (especially on hard targets — eukaryotic PPI, membrane proteins). Conversely, if many independent groups produce reproducible double-digit-percent success across many novel targets, (a) is strengthened.
- P2: if the correlation between in-silico metrics (ipSAE / AF2 self-consistency) and the wet lab holds precision on novel (out-of-distribution) targets, (c) the metric–reality gap weakens; if it is confined to easy / near-neighbour targets (precision collapsing to 0.1), (c) is strengthened. Testable via a target-decomposed follow-up to the section-2 meta-analysis.
- P3: if the refolding false positives (evolutionary info) and false negatives (flexibility) reproduce in later independent groups, confidence in the self-consistency-alone filter falls further; and if an orthogonal combined filter (Rosetta ΔG, shape complementarity) that corrects for them significantly raises the success rate on novel targets, a path to metric improvement opens. Testable via follow-ups to 693122 and the 670059 feature combinations.
References
- Jiang, Li, Guo, Wei, Wu. 2025. “RFdiffusion Exhibits Low Success Rate in De Novo Design of Functional Protein Binders for Biochemical Detection.” bioRxiv 2025.02.07.636769 (preprint, non-peer-reviewed; independent replication). https://www.biorxiv.org/content/10.1101/2025.02.07.636769v1
- Overath, Rygaard, Jacobsen et al. 2025. “Predicting Experimental Success in De Novo Binder Design: A Meta-Analysis of 3,766 Experimentally Characterised Binders.” bioRxiv 2025.08.14.670059 (preprint). https://www.biorxiv.org/content/10.1101/2025.08.14.670059v2.full
- Korbeld, Viliuga, Fürst. 2025. “Limitations of the refolding pipeline for de novo protein design.” bioRxiv 2025.12.09.693122 (preprint version). https://www.biorxiv.org/content/10.64898/2025.12.09.693122v1
- Korbeld, Viliuga, Fürst. 2026. “Limitations of the refolding pipeline for de novo protein design.” Protein Science (Wiley), peer-reviewed. DOI 10.1002/pro.70613. https://onlinelibrary.wiley.com/doi/10.1002/pro.70613
- Korbeld et al. Code and data for the refolding-limitations study (open). GitHub: kt-korbeld/Limitations-refolding-pipeline-data. https://github.com/kt-korbeld/Limitations-refolding-pipeline-data
- “Improving de novo protein binder design with deep learning.” 2023. Nature Communications (self-report / original author). s41467-023-38328-5. https://www.nature.com/articles/s41467-023-38328-5
- “One-shot design of functional protein binders” (BindCraft). 2025. Nature (self-report / original author). s41586-025-09429-6. https://www.nature.com/articles/s41586-025-09429-6
- Bennett et al. (RFdiffusion antibody design). 2025. Nature. s41586-025-09721-5 (exact success-rate figures not read — marked unverified in the source). https://www.nature.com/articles/s41586-025-09721-5
- Cao et al. 2022. “Design of protein-binding proteins from the target structure alone.” Nature 605 (self-report / original author; AF2/RF filter ~10x). s41586-022-04654-9. https://www.nature.com/articles/s41586-022-04654-9
Disclosure
This post is for information only and is not investment advice, and not medical advice.
COI note: this post describes listed companies (Alphabet / Google DeepMind and Isomorphic Labs, Meta / EvolutionaryScale, NVIDIA, Recursion RXRX, Schrödinger SDGR, Absci ABSI), private platforms (Generate Biomedicines, Xaira Therapeutics, Profluent, Chai, Cradle) and academic groups (Baker Lab / IPD, Fürst Lab) in a descriptive, neutral context. Every success rate and improvement factor self-reported by original authors or vendors is attributed as such and kept in a separate tier from independent (non-author) wet-lab replication. Peer-reviewed results (the Protein Science 2026 refolding critique) are distinguished from preprints (bioRxiv 636769, 670059, 693122). Figures are attributed to the specific studies; quantitative claims are attributed to the vendor, author or preprint. Competitive and method statements are factual, neutral descriptions and are not buy/sell implications for any security. The author holds no position in, and has no financial interest in, the companies named.
Leave a comment