Embodied AI Part 3 — the data bottleneck: robot foundation models have no internet-scale corpus, and teleop, sim and human video have not closed Goldberg’s 100,000-year gap

Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice. All dataset sizes, teleoperation costs, simulation-fidelity claims and synthetic-data figures are attributed to the relevant arXiv paper, journal, NVIDIA blog or company announcement. Peer-reviewed sources (Science Robotics, arXiv academic work), company/vendor claims (NVIDIA, 1X) and teleoperated demos are labeled separately. Teleop demo is not autonomous learning; simulation success is not real-world success.

The 30-second version

  • What. The deepest bottleneck in embodied AI is not the hardware and not the model architecture — it is the absence of an internet-scale corpus of real-robot interaction data. Language and protein foundation models scale on data that already exists at internet scale: Berkeley’s Ken Goldberg calculates that a modern vision-language model is trained on roughly 100,000 human-years of reading (Science Robotics, 2025-08-27, DOI 10.1126/scirobotics.aea7390). Robot interaction data does not exist that way — it has to be physically generated, and Goldberg names the deficit the “100,000-year data gap.”
  • So what. The four sources of robot data are all flawed. Teleoperation is linear — Goldberg’s framing is that “8 hours of work gives you 8 hours of data” — expensive and 1:1. Real-robot datasets are tiny next to internet scale (Open X-Embodiment >1M trajectories, DROID 350 hours, π0 ~10,000 hours, versus 100,000 years). Simulation/synthetic data hits sim ≠ real (NVIDIA’s GR00T-Dreams reports +40%, but that is within-benchmark, and domain randomization breaks on contact-rich manipulation). Human video has an embodiment mismatch and no action labels (Ego4D 3,670 hours, but no joint/gripper commands, and 2D→3D is hard).
  • Now what. No source shows a reproducible, deployed autonomous scaling result. Two claims are refuted outright: that “robot foundation models have the same internet-scale scaling corpus as LLMs,” and that “synthetic data plus video already substitutes for real-robot data and has closed the gap.” Data is the binding rate-limiter, which strengthens Part 0’s hypothesis (b) — a teleop-bound plateau — while hypothesis (a), a genuine ChatGPT moment, is not disproven but remains unproven in deployment. This is isomorphic to bio-foundation-models (“internet-scale text, but no wet-lab function data”) and to the space economy (“supply glut, but demand is the rate-limiter”).

The five-minute read

Why robot data is structurally different from language data

Language and protein foundation models could run scaling laws for one reason only — an internet-scale corpus already existed. Robots have no equivalent. Goldberg converts the text and image tokens used to train a modern large VLM into human reading time and arrives at roughly 100,000 years (Science Robotics, 2025-08-27). Dexterous manipulation is arguably harder than language, yet the corresponding real-world interaction data is not on the internet at all. It cannot be scraped; it must be physically produced — by a human teleoperating a robot, by synthesizing it in simulation, or by inferring it from human video. All three routes are flawed, and Goldberg names the deficit the “100,000-year data gap.”

The firm’s recurring lens — “the headline is the starting point; the real bottleneck is elsewhere” — applies cleanly. The visible headline layer (humanoid hardware) is maturing, so the bottleneck moves to the less visible outcome layer (the learning fuel, data). This is the same structural move seen in the space series, where launch-cost collapse floods supply (imagery, bandwidth) but paying downstream demand becomes the rate-limiter.

Four data sources, four different flaws

The four routes cannot be read head-to-head — their units differ (trajectories, hours, people) — and the numbers below are within-source values attributed to each paper, blog or company. A teleop demo is not autonomous learning; a simulation-benchmark gain is not a deployed autonomous success rate; human video is a pretraining aid, not autonomous manipulation data.

Data source Representative (company/research) Scale (source-attributed) Core flaw
1. Teleoperation Mobile ALOHA (Stanford); 1X Neo Expert Mode Mobile ALOHA system ~$32k (onboard power/compute); 1X Neo VR-headset teleop Linear and expensive (Goldberg: 8 h of work = 8 h of data; 1:1 human)
2. Real-robot datasets Open X-Embodiment; DROID; π0 OXE >1M trajectories / 22 embodiments; DROID 76k trajectories / 350 h / 564 scenes / 84 tasks; π0 ~10,000 h / 7 robots / 68 tasks Tiny vs internet scale (100,000 years vs hundreds–thousands of hours)
3. Simulation / synthetic NVIDIA Isaac Sim/Lab; MuJoCo; GR00T-Dreams; domain randomization GR00T-Dreams 780K synthetic trajectories (~6,500 h of human demo) in 11 h; GR00T N1 +40% (synthetic+real vs real-only, NVIDIA) Sim success ≠ real success (breaks on contact-rich manipulation)
4. Human video Meta Ego4D; GR00T egocentric video Ego4D 3,670 h / 931 participants / 74 locations (Meta 2022) Embodiment mismatch, no action labels, 2D→3D (pretraining aid only)
The figures are within-source values, not head-to-head comparisons (units differ). The +40% for GR00T N1 is a simulation-benchmark improvement, not a deployed autonomous success rate (sim success ≠ real success). 1X Neo Expert Mode is teleoperation, not autonomy. Human video is a pretraining signal, not autonomous manipulation data. Collection effort is not an internet-scale corpus.

Deep dive

1. Background — the deepest bottleneck is data, not hardware or the model

Part 0 framed embodied AI as a three-layer stack — HARDWARE (humanoid platform) × FOUNDATION MODEL (VLA) × DATA (real-robot data, sim-to-real) — and provisionally concluded that hardware is no longer the bottleneck (low-cost mass production, in-house manufacturing) and that the DATA layer is the deepest rate-limiter. Part 3 narrows that lowest layer to the empirical stage.

The counterintuitive point is that the bottleneck is neither “building the robot” nor “the model architecture” but the learning fuel (data). Language and protein foundation models train on an internet-scale corpus; robots have none. Robot data cannot be scraped from the internet and must be physically generated — by teleoperation, by simulation, or by inference from human video — and all three are flawed. The falsifiable central claim of this Part: the deepest bottleneck in embodied AI is the absence of internet-scale real-robot data (Goldberg’s 100,000-year gap), and none of the four data sources has closed that gap with an LLM-style scaling law in deployment.

2. The asymmetry — why robot foundation models have no scaling-law corpus

This is where the series’ “foundation model + data asymmetry” thesis is sharpest. Goldberg converts the text and image tokens used to train a modern large VLM into human reading time and reaches roughly 100,000 years (Science Robotics, 2025-08-27, DOI 10.1126/scirobotics.aea7390; techxplore 2025-08). GPT/VLM-class models are trained on the equivalent of “100,000 years of reading.” Dexterous manipulation is more complex than language, yet the matching real-world interaction data does not exist on the internet — at the current collection rate, Goldberg’s summary is that a general-purpose robot trained on a ChatGPT-scale robot dataset is roughly 100,000 years away.

The real-robot datasets are genuine progress but tiny. Open X-Embodiment (arXiv 2310.08864) is >1M trajectories / 22 embodiments / 60 datasets — the largest collaborative robotics dataset ever — yet a million trajectories is negligible next to trillions of language tokens. DROID (arXiv 2403.12945) is 76k trajectories / 350 hours / 564 scenes / 84 tasks; its “in-the-wild” diversity is real, but 350 hours is on the order of 0.00004% of 100,000 years (~8.76 billion hours). π0 (Physical Intelligence, pi.website) is ~10,000 hours / 7 robots / 68 tasks and fine-tunes on 1–20 hours per task — real generalization, but ~10,000 hours of pretraining still falls short of internet scale. Collection effort is not an internet-scale corpus. This is the same asymmetry the firm’s bio-foundation-models thread identified — internet-scale text and sequence exist, but phenotype and wet-lab function data are generated only experimentally and remain scarce — and for robots the asymmetry is even more extreme, because real-robot data is not on the internet at all.

3. Teleoperation — the main source, but linear so it does not scale

Teleoperation is the current workhorse for generating real-robot data, and it is the core of the data bottleneck because the physical act of making data is a linear rate-limiter. Goldberg’s formalization: operators drive robots “like puppets” in a warehouse, and the decisive limit is that “8 hours of work gives you 8 hours of data” (quote ≤150 chars, techxplore 2025-08). To double the data you must double the people and the time — the exact opposite of language data, which is already accumulated on the internet.

Low-cost teleop hardware lowers the entry barrier but not the linearity. Mobile ALOHA (Stanford, arXiv 2401.02117) is a bimanual whole-body teleop system at ~$32k (onboard power/compute included), and co-training with static ALOHA data raised success rates by up to +90% on tasks such as storing a pot, calling an elevator, pushing in a chair, and wiping spilled wine; ALOHA 2 (arXiv 2405.02292) improved the hardware and is fully open-source. But ALOHA lowers the cost of teleop without removing its linearity — it is still 1:1, so per-unit throughput is unchanged. 1X converts its Neo home robot into a teleop data engine: for chores the robot has not learned, a user schedules an “Expert Mode” session, a 1X operator drives the robot via a VR headset, and the trajectory is labeled and fed to the Redwood AI model (Humanoids Daily / hiverlab 2025). 1X states an autonomy target of ~30% (2026) rising to >90% (2029) — a company projection, not an achieved result, and the Redwood model’s deployed autonomous success rate has not been independently disclosed. A home teleop fleet also raises privacy questions (cameras and sensors granting external access inside the home, with blur and no-go-zone safeguards). Teleop is the main source, but its linearity is not overcome, and teleop data is not autonomous learning — the material basis for hypothesis (b), a teleop-bound plateau.

4. Simulation / synthetic data — sim success is not real success

The strongest route around teleop linearity is “buy data with compute” — synthesize it in simulation. Domain randomization (Tobin et al. 2017, arXiv 1703.06907) randomizes the simulator’s physics parameters (friction, mass, sensor noise, visual attributes) every episode for zero-shot transfer to reality; OpenAI’s in-hand manipulation (arXiv 1808.00177, 2018) is the canonical case, and it works well for locomotion and acrobatics. NVIDIA Isaac Sim/Lab, GR00T-Dreams and Cosmos (vendor blog) push further: the GR00T-Dreams blueprint generates synthetic trajectories from a single image plus a language prompt, producing 780K synthetic trajectories (~6,500 hours of human demo) in just 11 hours, and NVIDIA claims that combining synthetic and real data improves GR00T N1 by +40% versus real data only.

The decisive caveats. The +40% is a within-benchmark simulation improvement, not a deployed autonomous success rate — there is no independent evidence it translates to autonomous task success in unstructured real environments (demo/benchmark is not deployed). Domain randomization breaks on contact-rich manipulation: Goldberg notes it handles backflips and acrobatics but fails at practical dexterity such as plumbing and construction, because contact dynamics (wet surfaces, foam packaging, specular metal, variable friction) are hard for simulators to reproduce faithfully. The often-cited “~8 sim ≈ 1 teleop sample” ratio is a single secondary citation and is unverified. So synthetic data is the fastest route around teleop linearity and is isomorphic to the firm’s computing-power axis (GPU simulation/world-model demand surges to fill the data deficit, which is why NVIDIA sells hardware, model and simulator vertically) — but whether compute actually closes the data bottleneck or only lifts the benchmark is the unresolved gate for robot scaling laws. “Synthetic data solved the robot data problem” is a vendor narrative, not a deployed result.

5. Human video pretraining — embodiment mismatch, no action labels

The third route infers manipulation knowledge from human video — the internet holds vast footage of humans handling objects. Meta Ego4D (2022) is 3,670 hours of egocentric video, 931 participants, 74 locations, across 13 universities and labs, with hand-object contact labels, gaze tracking and 3D scans — the largest source of hand-object interaction for robotics — and GR00T N1 combines egocentric human video with real-robot and simulation data as a pretraining signal.

The decisive caveats on why human video does not translate to robot action. There are no action labels: Goldberg’s point is that video does not reveal the actual hand movements — footage shows the outcome (“what was done”) but no joint, force or gripper commands. 2D→3D is hard: recovering 3D hand and object pose from 2D video is extremely difficult, and manipulation needs accurate 3D contact and force. There is an embodiment gap: a human hand (five fingers, tactile) is not a robot gripper (two-finger, parallel), so the human method does not transfer directly. Some preprints claim egocentric human video can outperform real-robot pretraining, but that is a within-benchmark claim, not a deployed autonomous result (unverified). Human video is a useful pretraining aid, but it is the most indirect route — observation without action — and “teach robots from internet video” is a vision, not a solved path.

6. Data-engine strategies — will teleop fleets, RL or synthetic close the gap?

Goldberg names three routes to close the 100,000-year gap: simulation, spatial analysis of internet video, and real-robot data collection that works in real settings. Each player’s “data engine” is a combination of these. Teleop fleets (1X, Tesla) convert deployed robots into data-collection fleets (1X Neo via Expert Mode; Tesla Optimus positioned as “for training and data collection,” per Musk in Part 0) — parallelizing to soften linearity (N robots = N× data), but each unit still needs human teleop or supervision, and teleop is not autonomy. Synthetic scale-up (NVIDIA, Figure) mass-generates data on GPUs, where sim success ≠ real success is the gate. Reinforcement learning self-generates data inside simulation (MuJoCo, Isaac Lab), strong for locomotion but unresolved for contact-rich manipulation and real-world transfer. Co-training / transfer (π0, OXE) combines a little real data with a lot of diverse data (Mobile ALOHA +90%; π0 1–20 h per task), improving data efficiency but not changing the absence of an internet-scale corpus.

Does any of this close the gap? There is currently no deployed evidence that it does. All three routes show progress — teleop is linear (human-limited even when parallelized), sim is sim ≠ real (within-benchmark gains), video has an embodiment mismatch (pretraining aid) — but none has run a robot foundation model on an LLM-style scaling law with a reproducible, deployed result. Notably, Goldberg’s own prescription — “Good Old-Fashioned Engineering” (GOFE) — is peer skepticism of pure scaling optimism: because pure data scaling does not solve the problem, classical engineering (structure, constraints, modeling) must run alongside it. The data-engine strategies are all real progress, but each carries a fundamental flaw, and no deployed evidence closes the 100,000-year gap — which makes hypothesis (b), a teleop-bound plateau, the current best fit; hypothesis (a), a genuine ChatGPT moment, is not disproven but remains unproven in deployment, and the decider is whether synthetic/fleet data scale translates to deployed autonomous success rates (Parts 4–5).

7. Cross-domain links and what to watch (falsifiable)

This Part crosses three neighboring firm axes. Bio-foundation-models: the robot “no internet-scale real-robot data” is the isomorphic twin of bio-FM’s “internet-scale text and sequence, but scarce phenotype/wet-lab function data” — in both, the physical act of making data (wet-lab experiment / teleop) is the linear rate-limiter, and both try to route around it with synthesis (bio: simulated docking / robot: Isaac simulation) but hit sim ≠ real (bio: in-silico ≠ wet-lab / robot: sim success ≠ real success). Computing-power: the leading workaround is “buy data with compute” (GR00T-Dreams 780K trajectories in 11 hours, Isaac Sim GPU simulation), isomorphic to the computing-power thesis, with the unresolved gate being whether compute actually closes the data bottleneck or only lifts the benchmark. Space economy: structurally identical to “launch-cost collapse floods supply, but paying downstream demand is the rate-limiter” — the third instance of the firm’s lens that “the headline is the starting point; the real bottleneck is the outcome layer.”

  • P1 (synthetic/fleet data → deployed autonomy): in 2026–2028, if any platform reproducibly translates simulation/synthetic or teleop-fleet data scale-up into higher deployed autonomous task success (e.g. a GR00T-Dreams-style +40% carrying to unstructured real-world autonomy), that strengthens hypothesis (a) plus “compute substitutes for data.” If synthetic/fleet data grows but deployed autonomy keeps missing the benchmark (persistent sim-to-real gap), that strengthens hypothesis (b). (Verified in Part 4.)
  • P2 (overcoming teleop linearity): if new-task autonomous generalization is demonstrated across multiple platforms with a reproducible protocol using little or no teleop, and 1X-style autonomy ratios (30→90%) are actually achieved, that signals linearity overcome. If teleop-fleet / Expert-Mode dependence persists and autonomy stays a projection, that strengthens hypothesis (b). (Verified in Parts 4–5.)
  • P3 (data bottleneck vs compute substitution, cross-domain): if simulation/synthetic (Isaac, Cosmos, GR00T-Dreams) substantively replaces real-robot teleop, the data bottleneck eases and converges with the computing-power axis, dissolving the bio-FM asymmetry too. If the sim-to-real gap persists, data remains an irreducible physical rate-limiter, confirming the firm’s through-line that “the physical act of making data (wet-lab / teleop) is the linear rate-limiter.” (Verified in Part 5, with the bio-FM and computing-power series.)

Verdict: proceed-with-caveats. The data bottleneck is confirmed as the deepest rate-limiter — internet-scale real-robot data does not exist (Goldberg’s 100,000-year gap), teleop is linear (8 h = 8 h), real-robot datasets are tiny, sim success is not real success (GR00T-Dreams +40% is within-benchmark), and human video has an embodiment mismatch and no action labels. Two claims are refuted: “robot foundation models have the same internet-scale scaling corpus as LLMs,” and “synthetic plus video already substitutes for real-robot data and has closed the gap.” Demonstrated data-scaling is not a claimed scaling-law path.


References

Disclosure

This post is for information only and is not investment advice.

COI note: this post describes listed companies (Tesla TSLA, NVIDIA NVDA, Alphabet/Google GOOGL [DeepMind], Meta META) and private companies (Physical Intelligence, 1X Technologies, Figure AI, Skild AI) in a descriptive, neutral context. Every dataset size, teleoperation cost, simulation-fidelity claim and synthetic-data figure is attributed to the relevant arXiv paper, journal, NVIDIA blog or company announcement, with peer-reviewed sources (Science Robotics, arXiv academic work), company/vendor claims (NVIDIA, 1X) and teleoperated demos labeled separately. In particular, NVIDIA markets “solve the robot data problem with synthetic data” (Isaac Sim / GR00T-Dreams / Cosmos) as a growth narrative in which “substitute data with compute” is entangled with valuation logic; Tesla (Optimus fleet data) and 1X (Neo home teleop fleet) treat the “data engine” as central to their autonomy roadmaps. Teleop demo is not autonomous learning; simulation success is not real-world success; a within-benchmark gain is not a deployed autonomous success rate; demonstrated data-scaling is not a claimed scaling-law path. Quantitative claims are attributed to the vendor, author or preprint. Competitive and technology-narrative statements are factual, neutral descriptions and are not buy/sell implications for any security. The author holds no position in, and has no financial interest in, the companies named.