Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice. Every generalization figure below is attributed to the model’s arXiv paper or to a DeepMind / Nvidia / Physical Intelligence / Figure announcement; peer-reviewed and arXiv results are separated from company demos and blog claims. Crucially, all published numbers are in-protocol (benchmark / demonstration) success rates, not real-deployment autonomous-task success rates — and they are not head-to-head unless drawn from the same benchmark.
The 30-second version
- What. Vision-language-action (VLA) models train “image + language instruction → robot action” as a single foundation model. Starting with RT-2 (2023) and Open X-Embodiment / RT-X (2023), and continuing through π0 / π0.5, Nvidia GR00T N1 / N1.5, Google DeepMind Gemini Robotics and Figure Helix (2024–2025), they have demonstrated three things in benchmarks and demos: web-knowledge transfer, cross-embodiment transfer, and few-demo adaptation. RT-2 reported emergent semantic reasoning over 3× a top baseline (attributed); Open X-Embodiment reported RT-2-X roughly 3× on out-of-distribution objects/skills (attributed).
- So what. The improvements are real, but they are all in-protocol: they come from controlled evaluation sets and staged demonstrations, not from autonomous task success in unstructured real environments. No VLA model has published a reproducible, independent-protocol real-deployment autonomous success rate or intervention frequency. This is the exact gap the “ChatGPT moment for robotics” claim glosses over — benchmark generalization ≠ deployable generalization.
- Now what. Read the numbers for what they measure. The widely quoted “Gemini Robotics ~60% / ~80%” figures are generalization-axis scores (0–1) or a post-fine-tune 7-task average (~0.68) — not an “autonomous task success rate.” That semantic misread is the hype. Applying the firm’s bio-foundation-models lens (data leakage, strong baselines, subset cherry-pick, reproducibility), eval-contamination control across VLA benchmarks is mostly unreported. Company claims such as Helix’s “8-hour autonomous shift,” Skild’s “general robot brain,” and π0.5’s real-world autonomous success remain unverified.
The five-minute read
The VLA lineage — from web-knowledge transfer to cross-embodiment and dual-system control
A VLA model learns to map an image plus a language instruction to a robot action (an action token, or continuous control) inside one foundation model. The lineage advances along four axes. Web-knowledge transfer (RT-2): co-fine-tuning an internet-scale vision-language model (PaLI-X / PaLM-E) on robot actions, so the robot acquires semantic reasoning that was never in the robot data — the VLA prototype. Cross-embodiment (Open X-Embodiment / RT-X): pooling 22 robot types, 21 institutions, 527 skills, 160,266 tasks and over 1M trajectories to show positive transfer, where one robot’s data lifts another’s ability — a “robot ImageNet” attempt. Continuous control via flow/diffusion (π0 / π0.5, Helix): outputting high-frequency continuous control through a flow-matching / diffusion action expert instead of discrete action tokens. Edge deployment and few-demo adaptation (Gemini Robotics On-Device, GR00T N1.5): on-robot real-time inference and adaptation to a new task from 50–100 demonstrations.
All four share one hypothesis — port the scaling of language foundation models to robots. But the decisive asymmetry is that language and protein models train on internet-scale data, whereas robots have no internet-scale corpus of physical interaction. Whether a robot-FM scaling law holds despite that gap is the open question this series returns to.
The claim is “ChatGPT moment”; what is measured is in-protocol generalization
The firm’s recurring discipline — sharpened in the bio-foundation-models work — applies cleanly here: benchmark advantage ≠ deployable capability. RT-2’s “unseen scenario 32% → 62%,” Open X-Embodiment RT-2-X’s “~3×,” and Gemini Robotics On-Device’s “generalization-axis ~0.5–0.75” are all numbers inside controlled evaluation protocols. This is the same structural gap as in bio-FM, where “AlphaFold3 PoseBusters 84%” was not a drug and “in-silico design success” was not wet-lab function. The “ChatGPT moment for robotics” claim stays unestablished until three conditions are met: (1) benchmark numbers translate to real-deployment autonomous reliability; (2) the absence of internet-scale real-robot data (Goldberg’s “100,000-year data gap”) is somehow bridged; and (3) VLA benchmarks demonstrably control data leakage and eval contamination.
| Model (owner) | Key generalization figure (as attributed) | Nature of evaluation | Source tier |
|---|---|---|---|
| RT-2 (Google DeepMind) | Unseen scenario success RT-1 32% → RT-2 62%; emergent (symbol / semantic / human-recognition) over 3× the best baseline; broad generalization ~2×; ~6,000 real-robot eval trials | Benchmark / real-robot eval — in-protocol | arXiv 2307.15818 + DeepMind blog (primary + summary) |
| Open X-Embodiment / RT-X | RT-1-X ~50% higher success vs per-robot baselines in low-data regime; RT-2-X ~3× RT-2 on objects/skills absent from its robot data (OOD/emergent); 22 robots, 21 institutions, 527 skills, 160,266 tasks, >1M trajectories | Research eval — cross-embodiment (in-protocol) | arXiv 2310.08864 + DeepMind blog |
| π0 (Physical Intelligence) | Flow-matching VLA on top of a VLM; multi-platform (single-arm, bi-arm, mobile) data; laundry folding, table bussing, box assembly; zero-shot + fine-tuning | Research demonstration — per-task success in paper | arXiv 2410.24164 (RSS 2025) |
| π0.5 (Physical Intelligence) | Open-world generalization claim; kitchen cleaning / bedroom tidying (long-horizon) in untrained homes; tested in 3 SF rental houses | Research demonstration — real-world autonomous rate not independently confirmed | arXiv 2504.16054 + pi.website (primary + company) |
| GR00T N1 (Nvidia) | Open humanoid FM; human video + real + sim training; claims of exceeding SOTA imitation learning on sim benchmarks | Open model — sim-benchmark centric | arXiv 2503.14734 / Nvidia (primary) |
| GR00T N1.5 (Nvidia) | Over N1: frozen / reinforced VLM (language grounding), FLARE (human video), DreamGen synthetic actions; claims of exceeding N1 in sim + on real GR-1 robot; single-arm EEF / gripper support | Open model — sim + real robot (claim) | Nvidia GEAR blog (primary, company) |
| Gemini Robotics On-Device (Google DeepMind) | On-device model ~0.52–0.74 on generalization axes (visual / semantic / behavioral), off-device flagship ~0.60–0.75; adapts to a new task from 50–100 demos, then ~0.68 average across 7 dexterous tasks; ALOHA-trained → transfer to Franka FR3 / Apollo humanoid | Benchmark — generalization-axis scores / post-fine-tune (in-protocol) | DeepMind blog 2025-06-24 (primary) |
| Figure Helix (Figure AI, private) | Dual-system: S2 (internet-pretrained VLM 7B, 7–9Hz) + S1 (reactive visuomotor 80M, 200Hz); 35-DoF upper-body continuous control; split across on-robot embedded GPU | Company demo — architecture disclosed, autonomous rate not disclosed | figure.ai/news/helix (company claim) |
| Skild AI (private) | Foundation model aimed at a “general-purpose robot brain” | Company claim — no independent benchmark disclosed | Company statement (unverified) |
Deep dive
1. Background — what VLA is, and why “no internet-scale robot data” frames everything
The VLA narrative is clear: transfer the web knowledge of language / vision foundation models into physical control, and a robot should generalize to new tasks from little data, the way a language FM does. RT-2 opened the narrative in 2023; Open X-Embodiment added cross-embodiment data; π0 / π0.5, GR00T N1 / N1.5, Gemini Robotics and Figure Helix followed. The progress is real — in benchmarks and demonstrations, web-knowledge transfer, cross-embodiment transfer and few-demo adaptation are confirmed in arXiv / official primary sources.
But the decisive asymmetry sits underneath all of it. Language and protein foundation models are trained on internet-scale corpora; robots are not, because internet-scale real-world physical-interaction data does not exist (Goldberg’s “100,000-year data gap”). Strategies to fill the gap — cross-embodiment pooling, simulation and synthetic behavior (GR00T’s DreamGen), and human video (FLARE) — are exactly the attempts to substitute compute for missing data, which is where this thread meets the firm’s computing-power axis. Whether those substitutes genuinely stand in for real-robot teleoperation is unresolved.
2. What this landscape establishes — measured generalization, attributed and in-protocol
Principle: every generalization figure is reported as in the source, arXiv/peer-reviewed results are separated from company demos, and figures from different models/robots/protocols are not head-to-head unless drawn from the same benchmark.
- RT-2 (Google DeepMind) — web-knowledge transfer via co-fine-tuning a VLM on robot actions. Unseen-scenario success rose RT-1 32% → RT-2 62%; emergent capabilities (symbol understanding, semantic reasoning, human recognition) reached over 3× the best baseline; broad generalization averaged ~2×; across ~6,000 real-robot eval trials. All in-protocol (arXiv 2307.15818 + DeepMind blog). Note the internal spread: the “>3×” is a specific emergent category, while the broad average is ~2× — a subset versus a mean.
- Open X-Embodiment / RT-X — 22 robot types, 21 institutions, 527 skills, 160,266 tasks, over 1M trajectories pooled into one training set. RT-1-X showed ~50% higher success than per-robot baselines in the low-data regime; RT-2-X showed ~3× RT-2 on objects/skills absent from its original robot data (positive transfer). Research eval, in-protocol (arXiv 2310.08864 + DeepMind blog).
- π0 (Physical Intelligence) — a flow-matching VLA on top of a VLM, trained on multi-platform data (single-arm, bi-arm, mobile); demonstrated laundry folding, table bussing and box assembly with zero-shot plus fine-tuning. Quantitative rates are per-task inside the paper (arXiv 2410.24164, RSS 2025).
- π0.5 (Physical Intelligence) — an open-world generalization claim, showing long-horizon kitchen cleaning and bedroom tidying in homes not seen in training, tested across 3 San Francisco rental houses. The direction (open-world generalization) is confirmed by arXiv / company, but a quantitative autonomous success rate and intervention frequency are not independently confirmed (arXiv 2504.16054 + pi.website).
- GR00T N1 (Nvidia) — an open humanoid foundation model trained on human video + real + sim, with claims of exceeding SOTA imitation learning on simulation benchmarks (arXiv 2503.14734 / Nvidia). Sim-benchmark centric.
- GR00T N1.5 (Nvidia) — over N1, adds a frozen / reinforced VLM for language grounding, FLARE (human video) and DreamGen synthetic actions, with claims of exceeding N1 in sim and on the real GR-1 robot, plus expanded single-arm EEF / gripper support. The “exceeds N1 on real GR-1” magnitude is a company blog claim with no independent reproduction (Nvidia GEAR blog).
- Gemini Robotics On-Device (Google DeepMind) — the on-device model scored ~0.52–0.74 on visual / semantic / behavioral generalization axes, the off-device flagship ~0.60–0.75; adaptation from 50–100 demos yielded a ~0.68 average across 7 dexterous tasks; models trained on ALOHA transferred to Franka FR3 and the Apollo humanoid. A DeepMind-internal benchmark (not independent third-party), in-protocol (DeepMind blog 2025-06-24).
- Figure Helix (Figure AI, private) — a dual-system architecture, S2 (internet-pretrained 7B VLM, 7–9Hz) plus S1 (reactive visuomotor, 80M, 200Hz), driving 35-DoF upper-body continuous control split across an on-robot embedded GPU. The architecture is disclosed; an autonomous success rate is not (figure.ai/news/helix, company claim). The “8-hour fully autonomous shift” figure is a company/secondary claim, unverified.
- Skild AI (private) — a foundation model aimed at a “general-purpose robot brain”; no independent benchmark or quantitative success rate has been disclosed (company statement, unverified).
3. The core commercial question — is this a “ChatGPT moment for robotics”?
Narrow the series’ three hypotheses — (a) a genuine ChatGPT moment, (b) demo / teleop-bound stagnation, (c) narrow vertical commercialization — to what the VLA layer alone shows. Observations that support the “moment” claim: RT-2’s emergent semantic reasoning (>3×), whose shape resembles language-FM emergence; Open X-Embodiment’s positive transfer (RT-2-X ~3×), a precursor to a scaling law; Gemini Robotics On-Device’s 50–100-demo adaptation plus ALOHA→Franka/Apollo cross-hardware transfer, the core few-data property of an FM; and π0.5’s untrained-home long-horizon demonstration, in an open-world direction.
Observations that keep the claim from being established (firm lens): (1) benchmark ≠ real-world — every figure above is in-protocol; no VLA model has published an autonomous success rate in unstructured real environments (specular metal, wet surfaces, occlusion, lighting change) under a reproducible independent protocol. Just as “PoseBusters 84% ≠ a drug” in bio-FM, “RT-2 62% ≠ deployable autonomy.” (2) Eval-contamination uncontrolled — VLA benchmarks mostly do not report whether train/eval environments and objects are near-duplicates (cold split), or whether a strong subset (a tidy lab, specific objects) was promoted to the headline; if unmeasured, it is [unverified]. (3) Strong baselines missing — the robot analogue of bio-FM’s “specialist baseline an expert writes in half a day” is scripted control, task-specific behavior cloning and classical motion planning; whether these beat a VLA on structured, repetitive tasks is the fair contrast, but most VLA papers compare against a previous-generation learned policy instead. (4) Scaling law unverified — language-FM capability jumps came from internet-scale data that robots lack; whether RT-2-X’s ~3× keeps rising predictably with data scale, and whether cross-embodiment / sim-synthetic (GR00T DreamGen) / human video substitute for real-robot teleop, is untested.
Tentative position: VLA is the only layer to have demonstrated generalization improvement in benchmarks and demos, and the signals one would call “precursors to a ChatGPT moment” (emergence, transfer, few-demo adaptation) are real. But the conclusion “the moment has arrived” is unestablished — there is no independent-protocol real-deployment autonomous success rate, eval-contamination control is unreported, and it is untested whether a scaling law holds past the data bottleneck. The current evidence is a coexistence of (a)-type precursors with (b)/(c)-type real-world constraints. Framing “RT-2 / π0.5 / Gemini are the GPT of robots” is hype unless it states plainly: benchmark generalization ≠ deployable generalization.
4. Where the firm lens bites — five traps in VLA benchmarks (ported from bio-FM)
Port the bio-foundation-models “honest-evaluation checklist” onto VLA as a gate for any “VLA generalizes” claim. Failing one item lowers the tier and requires a caveat.
- In-protocol vs deployment. Is the number a controlled eval set / demo, or an autonomous success rate in an unstructured real environment? Is intervention frequency (interventions/hour) reported? If the former, label it “benchmark generalization” only — do not promote to “autonomous capability.”
- Eval contamination (leakage). Are eval scenes/objects near-duplicates of the training distribution? Is it reported how far success falls on held-out environments/objects (bio-FM: cold split, temporal split)? VLA usually does not report this → [unverified].
- Strong baseline. Is the control a task-specific script / behavior cloning / classical motion planning, or only a previous-generation learned policy (bio-FM: linear baseline)? Beating a specialist baseline on structured tasks is the fair contrast.
- Subset cherry-pick. Was a strong subset (tidy lab, specific objects, short horizon) promoted to the overall headline? What is the variance on long-horizon, contact-rich tasks (wet soap, foam packaging)? (RT-2’s emergent >3× is a specific category; broad generalization is ~2× — the variance behind the mean.)
- Reproducibility tier. Are code, weights, eval scripts and protocol public? Open (GR00T N1.5, some π0 code) vs closed (Gemini internal, Helix, Skild). A closed demo is structurally a lower verification tier → load it as a company claim only.
Summary: just as in bio-FM, when (a) favorable-split / near-duplicate eval, (b) a weak baseline, (c) a strong subset, and (d) a closed demo overlap, “SOTA / ChatGPT moment” is nearly guaranteed. The firm separates a vendor demonstration (a claim) from an independently reproducible protocol result (evidence) — and in the VLA layer the latter is scarce.
5. Commercialization and competitive context
- Maturity (TRL frame): benchmark/demonstration generalization is demonstrated, but real-deployment autonomous reliability is early — no VLA has published an independent-protocol autonomous success rate. The gating layer is the outcome layer (deployable autonomy), not benchmark generalization.
- Nvidia (NVDA): markets GR00T and Isaac as a “physical-AI foundation-model” growth axis, selling hardware, models and simulation vertically (the DreamGen / FLARE / Isaac Sim strategy of substituting sim/synthetic compute for scarce data). GR00T N1.5’s real-robot claims are company-blog, not independently reproduced.
- Google DeepMind / Alphabet (GOOGL): narrates robot-FM leadership through Gemini Robotics and Gemini Robotics On-Device; its generalization-axis scores are DeepMind-internal benchmarks.
- Physical Intelligence (private): π0 / π0.5, the flow-matching VLA line, with an open-world home demonstration whose quantitative autonomous rate is not independently confirmed.
- Figure AI (private): Helix, a dual-system on-robot architecture; the “8-hour autonomous shift” is a company/secondary claim, unverified.
- Skild AI (private): a “general robot brain” foundation model with no disclosed independent benchmark.
- Company statements here are limited to neutral, source-attributed description; leadership or capability rankings are not asserted and are not buy/sell signals. Cross-vendor comparisons are not head-to-head unless from the same benchmark.
6. The skeptic’s bottom line
- Benchmark ≠ real-world: every VLA generalization figure is in-protocol; the count of independent-protocol real-deployment autonomous success rates is zero.
- Read what the number measures: “Gemini Robotics ~60% / ~80%” are generalization-axis scores (0–1) or a post-fine-tune 7-task average (~0.68), not an autonomous task success rate — the semantic misread is the hype.
- Eval contamination uncontrolled: VLA benchmarks mostly do not report cold split, held-out generalization drop, strong-baseline contrast or subset cherry-pick (bio-FM Part 4 lens). Where unmeasured, treat as unverified.
- Company claims stay claims: Helix’s “8-hour autonomous shift,” Skild’s “general brain,” and π0.5’s real-world autonomous success rate are company statements with no independent primary verification.
- Scaling law unverified: whether a robot-FM scaling law holds is blocked by the absence of internet-scale real-robot data; substituting sim/synthetic/human-video for teleop is unproven.
- Neutral-framing note: to prevent misreading listed (Alphabet GOOGL, Nvidia NVDA) or private (Physical Intelligence, Figure, Skild) roadmap statements as security or capability signals.
7. What to watch (falsifiable)
- P1: if any VLA (π0.5, GR00T, Gemini, Helix) publishes, under a reproducible independent protocol, a teleop-free autonomous success rate plus intervention frequency in an unstructured real environment, and that number approaches its benchmark, hypothesis (a) ChatGPT-moment strengthens; if the deployment rate falls far short of the benchmark, or if no such disclosure ever appears, hypothesis (b) stagnation strengthens.
- P2: if a bio-FM-style standard protocol (cold split / new-environment holdout + strong baseline such as a specialist script / behavior cloning) spreads to VLA, a large share of the current generalization advantage will shrink significantly (with cases of more than a halving on specific benchmarks) — the robot analogue of the bio-FM leakage-control prediction.
- P3: if cross-embodiment / sim-synthetic (GR00T DreamGen) / human video genuinely substitute for real-robot teleop and data scale translates predictably into higher real-deployment autonomous success, a robot-FM scaling law holds and (a) strengthens — while simultaneously pushing sim/inference compute demand up the computing-power axis. If the sim-to-real gap persists, (b) strengthens.
References
- Brohan et al. 2023. “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.” arXiv 2307.15818. https://arxiv.org/abs/2307.15818
- Google DeepMind. “RT-2: New model translates vision and language into action.” https://deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/
- Open X-Embodiment Collaboration. 2023. “Open X-Embodiment: Robotic Learning Datasets and RT-X Models.” arXiv 2310.08864. https://arxiv.org/abs/2310.08864
- Google DeepMind. “Scaling up learning across many different robot types.” https://deepmind.google/blog/scaling-up-learning-across-many-different-robot-types/
- Physical Intelligence et al. 2024. “π0: A Vision-Language-Action Flow Model for General Robot Control.” arXiv 2410.24164 (RSS 2025). https://arxiv.org/abs/2410.24164
- Physical Intelligence et al. 2025. “π0.5: a VLA with Open-World Generalization.” arXiv 2504.16054. https://arxiv.org/abs/2504.16054
- Physical Intelligence. “π0.5” (company blog). https://www.pi.website/blog/pi05
- Nvidia. 2025. “GR00T N1: An Open Foundation Model for Humanoid Robots.” https://research.nvidia.com/publication/2025-03_nvidia-isaac-gr00t-n1-open-foundation-model-humanoid-robots
- Nvidia GEAR. “GR00T N1.5” (company blog). https://research.nvidia.com/labs/gear/gr00t-n1_5/
- Google DeepMind. 2025-06-24. “Gemini Robotics On-Device brings AI to local robotic devices.” https://deepmind.google/blog/gemini-robotics-on-device-brings-ai-to-local-robotic-devices/
- Figure AI. “Helix: A Vision-Language-Action Model for Generalist Humanoid Control” (company). https://www.figure.ai/news/helix
- Bio-foundation-models evaluation lens (leakage / strong baseline / subset). Nature Methods 2025, s41592-025-02772-6. https://www.nature.com/articles/s41592-025-02772-6
Disclosure
This post is for information only and is not investment advice.
COI note: this post describes listed (Alphabet / Google DeepMind, GOOGL; Nvidia, NVDA) and private (Physical Intelligence, Skild AI, Figure AI) robot-foundation-model efforts in a descriptive, neutral context. Every generalization figure is attributed to the model’s arXiv paper or to a DeepMind / Nvidia / Physical Intelligence / Figure announcement, and peer-reviewed / arXiv (academic) results are separated from company demos and blog claims (company claim). Benchmark / in-protocol success is not real-world autonomous task success, and figures from different models, robots or protocols are not head-to-head unless drawn from the same benchmark. Quantitative claims are attributed to the vendor, author or preprint. Company statements are factual, neutral descriptions and are not buy/sell implications for any security. The author holds no position in, and has no financial interest in, the companies named.
Leave a comment