Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice. All deployment, pilot, uptime and cost figures are attributed to their source, and separated into three tiers: peer-reviewed / primary reporting (IEEE Spectrum), company press release or company page (company claim), and secondary aggregators (promotional). Demo is not deployment, a completed pilot is not ongoing production, and a pilot is not lights-out autonomous operation.
The 30-second version
- What. After Parts 1–3 (hardware, VLA models, data), Part 4 looks at the bottom of the stack — the deployment layer. Humanoid hardware is maturing and the demos and pilots are real, but every headline “deployment” resolves to a small, controlled, narrow pilot. BMW Spartanburg = one Figure 02 robot at a single station (sheet-metal insertion), described by BMW itself as a “successfully completed pilot” — not ongoing production. IEEE Spectrum (citing Agility’s former CPO Melonee Wise) summarizes the field as “a small handful of robots in carefully controlled pilot projects.”
- So what. The binding constraint is no longer the robot (hardware); it is the deployment layer’s two ceilings — lights-out autonomous reliability and labor-replacement unit economics. Industrial customers expect roughly 99.99% reliability; a stopped line costs tens of thousands of dollars per minute; “currently AI is not robust enough to meet the requirements of the market” (Wise). Battery is a physical sub-constraint: Digit’s roughly 10-to-1 charge ratio leaves about 30 minutes of effective uptime. Critically, no vendor discloses a real autonomous task-success rate, interventions-per-hour, or MTBF — the central transparency gap.
- Now what. The labor-replacement economics remain unproven. The circulating $/hr total-cost-of-ownership figures ($7.55; Digit $10–12/hr; sub-6-month payback) are promotional secondary aggregators that omit integration cost, hidden intervention labor and downtime; there is no audited, positive-ROI general-purpose deployment, and the humanoid-specific safety standard (ISO 25785-1) is unpublished. Two claims are refuted: that humanoids are “already deployed as lights-out autonomous workers at scale,” and that “labor-replacement ROI already holds.” An impressive demo or a completed pilot is not a deployable autonomous worker at scale.
The five-minute read
The deployment spectrum — demo is not pilot is not lights-out autonomous
To judge labor replacement you first have to split the spectrum that vendors tend to blur. A demo is a staged, controlled show (lighting, objects and timing controlled, cherry-picking possible); teleop / human-supervised means remote operation or constant human oversight and intervention; a pilot is a narrow real-environment trial (few units, supervised, time-limited); lights-out autonomous means continuous autonomous operation with minimal human intervention; and working-at-scale means hundreds to thousands of units repeating at a positive ROI. Almost every “deployment” headline today sits at the pilot rung. The decisive undisclosed metrics — interventions-per-hour, MTBF and real-world autonomous task-success rate — are exactly what would let an outsider tell a lights-out deployment from a supervised, teleop-assisted pilot, and no vendor publishes them under an independent protocol.
The bottleneck moved from the robot to reliability and labor economics
The firm’s recurring lens — “the headline is the entry point; the real bottleneck is elsewhere” — bites hardest here. The headline is the demo and the completed pilot, and those are real. But at the deployment layer the constraint moved off the robot itself onto four different ceilings: (1) lights-out autonomous reliability, where industrial customers expect roughly 99.99% and a stopped line costs tens of thousands of dollars per minute; (2) labor-replacement unit economics, where robot price plus integration plus intervention labor plus failure rate is weighed against wages, and no audited positive-ROI general-purpose deployment exists; (3) a safety/certification gap, where ISO 10218:2025 assumes statically stable robots and the dynamically stable legged/humanoid standard ISO 25785-1 is unpublished; and (4) demand, where, per Wise, no application has yet been found that would require several thousand robots per facility.
| Deployment stage | Observable metric | Where humanoids are today |
|---|---|---|
| Demo (staged, controlled) | Video; no reproduction protocol | Many (Optimus, Figure, NEO and others) |
| Teleop / human-supervised | Interventions/hour (undisclosed) | Widespread (1X NEO “Expert Mode”; teleop in demos, per Part 0) |
| Pilot (narrow, few units, time-limited) | Duration, unit count, success rate | Where most “deployment” headlines actually sit |
| Lights-out autonomous | Uptime, MTBF, interventions/hour, autonomous success rate | No independently demonstrated case (the central unknown) |
| Working-at-scale | Unit economics, scale, demand | Not demonstrated |
Deep dive
1. Background — the deployment spectrum and the no-fabrication distinctions
Parts 1–3 documented, at each layer, the gap of demo is not teleop is not autonomous is not deployed is not reliable. Part 4 takes the most downstream layer — deployment — and asks whether a matured robot with benchmark-generalizing VLA control actually repeats a task in the real world, lights-out, autonomously, reliably, and at a profit.
Four distinctions carry the argument. (1) Demo is not deployed: a staged show is not real-world autonomy; IEEE Spectrum summarizes the field as “even the most successful companies in this space have deployed only a small handful of robots in carefully controlled pilot projects.” (2) A completed pilot is not ongoing production: a finished pilot is a success signal, but it is not standing production deployment (see BMW below). (3) A pilot is not lights-out autonomous: a few-unit, supervised, time-limited pilot can include human intervention and teleop assistance. (4) Announced is not working-at-scale: announced deployment plans and volume targets (Hyundai 30,000/yr, Figure 100,000/yr, per Part 0) are not actual operation. The central undisclosed metrics — interventions-per-hour, MTBF and real-world autonomous task-success rate — are the transparency gap of the whole layer: without them, no deployment can be classified from the outside as lights-out versus supervised.
2. Lights-out autonomous reliability — the 99.99% gate (IEEE Spectrum, primary reporting)
The first binding constraint is reliability/uptime. For labor replacement to hold, a robot must repeat a task as reliably as a person, without intervention. IEEE Spectrum’s “Humanoid Robots: The Scaling Challenge” (2025-09-11), citing Agility Robotics’ former CPO Melonee Wise, quantifies the gate: industrial customers expect roughly 99.99% reliability; even 99% reliability implies about five hours of downtime per month, and a stopped production line costs tens of thousands of dollars per minute. A demo’s “nine out of ten” (90%) is therefore catastrophic in industrial deployment — the jump from 99% to 99.99% is a qualitatively different axis from a demo success rate. Industrial robot arms (FANUC, ABB, KUKA) reach 95–99% uptime, but on fixed, structured, repetitive tasks; humanoids target unstructured, variable, multi-task work, and that generalization is exactly what breaks reliability. In Wise’s words, “currently AI is not robust enough to meet the requirements of the market” — an industry insider naming autonomous reliability as the binding constraint.
Reliability has a physical sub-constraint: battery and runtime. Per IEEE Spectrum, Agility’s Digit runs a “10 to 1” charge ratio — about 90 minutes of operation for 9 minutes of charge — yet in real deployment it often runs only about 30 minutes, because the remaining 60 minutes is effectively held in reserve for unexpected floor events. This is a candidate exception to the Part 0/1 thesis that “hardware is not the bottleneck”: the “ten-hour shift” narrative runs into battery, thermal management and charging logistics as a deployment-layer uptime ceiling.
| Vendor / customer | Deployment status (attributed) | Autonomy / scale | Source tier |
|---|---|---|---|
| Figure @ BMW Spartanburg | “Successfully completed pilot” (2025); one Figure 02, single station (sheet-metal insertion); 10-hour shift M–F, 90,000 components, ~1,250 operating hours, 30,000 X3 supported; IT/safety teams involved early | One robot, narrow repetitive task, completed pilot (not ongoing production) | BMW press release (primary) |
| Figure 03 @ Spartanburg (logistics) | Announced 2026-06: pick parts from unsorted containers to a sequencing trolley | Unit count, autonomy scope, interventions undisclosed | Company/press (company claim) |
| Agility Digit @ GXO | “One-year anniversary of Digit’s full-time deployment at GXO” (company claim, 2025-10-02); AMR unloading, tote onto conveyor | Narrow warehouse task; interventions, success rate, uptime undisclosed | Agility company page (company claim) |
| Agility Digit @ Amazon | Agility’s current page does not mention Amazon; only secondary aggregators claim “still running” | Current status not confirmed by primary source (unverified) | Secondary aggregator |
| Apptronik Apollo @ Mercedes-Benz | “Commercial agreement that will pilot Apollo” (PR, 2024-03); intralogistics (kit delivery, tote transport, initial inspection) | Pilot; commercial quantities 2027 is a plan | PR (primary) / secondary |
3. Labor-replacement unit economics — ROI unproven, $/hr TCO is promotional
The second binding constraint is the labor-replacement economics: robot price plus integration plus intervention labor plus failure rate versus wages — does a positive ROI hold? The core of the labor-replacement narrative is “robot $/hr << human wage,” but the circulating figures are unverified. The robot TCO of about $7.55/hr (depreciation, power, maintenance, software), human manufacturing labor at about $25–30/hr (taxes, insurance, training, management), Digit at $10–12/hr versus a human at $30/hr, Digit “98% task success,” and a sub-6-month payback at a $20K target price are all secondary aggregators (promotional) that the firm could not independently confirm at the primary level; under skeptic discipline they are treated as unverified.
What those numbers structurally omit is (a) integration — cell design, safety fencing, MES integration and operating staff, where the per-robot integration and maintenance labor often exceeds the robot’s own price; (b) hidden intervention labor — the human teleop/supervision that fills the autonomous-reliability gap (see §2 and the cross-domain “hidden teleop labor” hook); (c) failure and downtime cost — tens of thousands of dollars per minute of stopped line when reliability falls short of 99.99%; and (d) battery uptime loss (the ~30-minute effective uptime of §2). Add these four back in and the promotional $/hr rises sharply. In short, there is no audited, positive-ROI general-purpose deployment.
On market size, Goldman Sachs (2024-01-08, “The global market for humanoid robots could reach $38 billion by 2035”) projects a humanoid TAM of $38bn by 2035 (a sixfold upgrade from a prior $6bn estimate), 250,000+ shipments by 2030 (mostly industrial), and manufacturing cost falling roughly 40% per year; a separate Goldman note gives a $6bn base and a $154bn blue-sky by 2035. These are model projections, not real deployments or revenue (isomorphic to the “announced is not operational” pattern of the space-economy series). Goldman itself conditions success on overcoming “product design, use case, technology, affordability, and wide public acceptance,” and names “significant bottlenecks in the development of AI and software for robot manipulation and interaction” — that is, even the analyst forecast locates the bottleneck at the results layer (AI/software, acceptance, economics).
The deepest constraint on the labor economics is effective demand. Per Wise, the core problem is demand, not technology: “I don’t think anyone has found an application for humanoids that would require several thousand robots per facility.” Most real tasks are either done more cheaply by existing fixed automation (conveyors, AMRs, robot arms) or still require human fingertip dexterity. The humanoid form factor’s “generality” becomes a dilemma in which it is more expensive and less reliable than dedicated automation on any single task. This mirrors the space-economy result exactly: falling launch cost ($/kg) did not guarantee a profitable space business, and matured, cheaply mass-produced robot hardware does not guarantee autonomous workers or labor-replacement ROI.
4. The safety and certification gap — the humanoid standard is unpublished
The third binding constraint is safety and certification. A humanoid moving autonomously in shared human space needs a dedicated safety frame that does not yet exist. ISO 10218:2025 (published 2025-02; -1 for the robot, -2 for the cell/application) is the reference for industrial-robot safety and integrates ISO/TS 15066 (collaborative-robot contact limits) — but it assumes statically stable robots. Dynamically stable legged/humanoid robots are addressed by the in-development ISO/WD 25785-1, which is unpublished; there is, in effect, no humanoid-specific safety standard, so deployment vendors must borrow and adapt ISO 10218:2025, ISO 13849-1 and IEC 62061 to make a safety case (standards commentary, secondary).
The implication is that the certification frame for the human-robot-collaboration risks (shared space, mobility, falls, impact) is incomplete, so the regulatory, insurance and liability path for large-scale lights-out deployment is uncertain. This constrains the speed at which a pilot (with safety-team involvement, fencing and supervision) translates into standing autonomous deployment — consistent with BMW Spartanburg’s early “IT/safety team involvement.” Safety and certification is thus the regulatory ceiling of the deployment layer, isomorphic to the space-economy series’ regulatory ceiling (EPFD, landing rights).
5. Long-horizon and edge-case failure, and the synthesis of the four ceilings
VLA benchmark success rates (RT-2 roughly 3x, Gemini roughly 60–80%, per Parts 0/2) are numbers inside short, structured protocols. Real deployment breaks on long-horizon (multi-step chains) and edge cases (specular metal, wet or deformable objects, lighting change, exception events): even at 90% per-step success, a ten-step chain collapses to about 35% (0.9 to the tenth). This long-horizon reliability collapse, combined with the 99.99% gate of §2, is what blocks general-purpose autonomy at the deployment layer; Digit’s “60-minute reserve” (§2) is itself a hedge against unexpected edge cases.
Synthesizing the four ceilings: the combination of autonomous-reliability shortfall (§2) and unproven positive ROI (§3) is precisely why matured hardware does not translate into labor replacement — when human intervention fills the reliability gap, that labor cost breaks the ROI (the “hidden teleop labor” of the computing-power hook). The combination of the certification gap (§4) and absent demand (“no application requiring several thousand robots,” §3) is the regulatory-and-demand ceiling on filling a factory the way one might “fill the sky.” Pilots may multiply, but the path to scale autonomous deployment stays narrow. This is isomorphic to the firm’s recurring thesis across space-economy (launch cost vs unit economics/orbit/demand), energy-storage (lab Wh/L vs GWh $/kWh), GLP-1 (mechanism vs hard outcome) and computing-power (chip performance vs power/data): the headline metric is the entry point; the contest is decided by a different ceiling at the results layer. In robotics that ceiling is not the robot (hardware) but lights-out autonomous reliability, labor economics, safety certification and demand.
6. The skeptic’s bottom line
- Deployments are small, controlled, narrow pilots. BMW Spartanburg is one robot at a single station, a “completed pilot” (not ongoing production); the field is “a small handful of robots in carefully controlled pilot projects.” Demo and pilot are not lights-out autonomous.
- Autonomous success rate, interventions-per-hour and MTBF are undisclosed by every vendor — the central transparency gap. Without them, one cannot tell lights-out from supervised, teleop-assisted operation.
- The labor-replacement $/hr TCO is promotional secondary — omitting integration, intervention labor, failure/downtime and battery uptime. No audited positive-ROI general-purpose deployment exists; Goldman’s TAM is a model projection that itself conditions on AI/software, affordability and acceptance.
- Reliability falls short of the 99.99% gate and battery caps uptime — a demo’s 90% is catastrophic industrially; Digit’s effective uptime is about 30 minutes.
- Safety and certification is a gap — the humanoid-specific ISO 25785-1 is unpublished, leaving the regulatory, insurance and liability path for autonomous deployment uncertain.
- Two claims are refuted: that humanoids are “already deployed as lights-out autonomous workers at scale,” and that “labor-replacement ROI already holds.” Both fail against small, controlled pilots, undisclosed autonomy metrics and promotional-only economics.
- Neutral-framing note: listed (TSLA, AMZN, MBG, BMW, NVDA) and private (Figure, Apptronik, Agility, 1X) references are neutral, source-attributed descriptions — not vendor rankings, not “labor replacement is imminent,” and not “humanoids have failed.” Neither reading is a security signal.
7. What to watch (falsifiable)
- P1: if, in 2026–2028, any vendor discloses real-world autonomous task-success rate, interventions-per-hour, MTBF and uptime under an independent protocol and those numbers converge on 99%+ with minimal human intervention, the bottleneck eases; if only company claims (pilots, narrow tasks) keep appearing while the autonomy metrics stay undisclosed, the “demo/teleop-bound stall” hypothesis strengthens. (Verify via Figure/Agility/Apptronik follow-on disclosures and audits.)
- P2: if an audited, positive-ROI general-purpose deployment emerges (integration, intervention labor and failure cost included), or a single facility profitably runs hundreds to thousands of units, the “narrow-vertical commercialization” hypothesis extends; if only promotional $/hr repeats while actual deployment stays at a few pilots, the “scale not reached” reading strengthens. (Verify via audited financials, unit counts, demand applications.)
- P3: if autonomous reliability converges to 99.99% and ISO 25785-1 is published, the deployment ceiling eases — while the open questions become whether that reliability translates into on-robot inference / simulation compute (the computing-power axis) and whether interventions-per-hour (hidden teleop labor) actually vanishes. If reliability stalls at demo level and interventions persist, the “demo/teleop-bound stall” hypothesis strengthens and the results-layer isomorphism with space-economy Part 4 (“falling launch cost is not profitability”) is confirmed. (Verify via IEEE/standards follow-ups, interventions disclosures, Part 5.)
References
- BMW Group. 2025. “BMW Group to Deploy Humanoid Robots in Production in Germany for the First Time” (Spartanburg Figure 02 completed pilot; Leipzig AEON pilot summer 2026). Press release T0455864EN. press.bmwgroup.com/…/T0455864EN
- IEEE Spectrum. 2025. “Humanoid Robots: The Scaling Challenge” (Melonee Wise; 99.99% reliability gate; Digit 10-to-1 charge; demand bottleneck). 11 September 2025. https://spectrum.ieee.org/humanoid-robot-scaling
- Agility Robotics. 2025. “Digit’s Next Steps” (one-year full-time deployment at GXO, company claim). 2 October 2025. https://www.agilityrobotics.com/content/digits-next-steps
- Apptronik. 2024. “Apptronik and Mercedes-Benz Enter Commercial Agreement That Will Pilot Apptronik’s Apollo Humanoid Robot in Mercedes-Benz Manufacturing Facilities.” PR Newswire, March 2024. prnewswire.com/…/apptronik-mercedes-benz
- Apptronik. “Apptronik Closes Over $935 Million Series A” (company). apptronik.com/…/series-a
- Goldman Sachs. 2024. “The Global Market for Robots Could Reach $38 Billion by 2035” (250,000 shipments by 2030; ~40% annual cost decline; AI/software bottleneck). 8 January 2024. goldmansachs.com/…/38-billion-by-2035
- Goldman Sachs. “Humanoid Robot: The AI Accelerant” ($6bn base / $154bn blue-sky; product design, use case, affordability, acceptance as obstacles). goldmansachs.com/…/humanoid-robot-the-ai-accelerant
- SICK / standards commentary. “What You Need to Know About the ISO 10218:2025 Standard” (ISO 10218:2025 for statically stable robots, ISO/TS 15066 integration; ISO/WD 25785-1 for dynamically stable legged/humanoid robots, unpublished). sickconnect.com/…/iso-10218-2025-standard
Disclosure
This post is for information only and is not investment advice.
COI note: this post describes listed and private robotics companies and their customers (Tesla TSLA, Amazon, BMW / Mercedes-Benz, Agility Robotics, Apptronik, Figure AI, 1X, Nvidia NVDA, GXO, Jabil) in a descriptive, neutral context. Every deployment, pilot, uptime and cost figure is attributed to its source and separated into tiers: primary reporting (IEEE Spectrum), company press release or company page (company claim), and secondary aggregators (promotional). Demo is not deployment, a completed pilot is not ongoing production, and a pilot is not lights-out autonomous operation. The circulating $/hr total-cost-of-ownership figures (about $7.55, Digit $10–12/hr, sub-6-month payback) are promotional and unverified: they omit integration cost, hidden intervention labor, failure/downtime and battery uptime loss, and no audited positive-ROI general-purpose deployment exists. Quantitative claims are attributed to the vendor, author or press release. Competitive and capability statements are factual, neutral descriptions and are not buy/sell implications for any security. The author holds no position in, and has no financial interest in, the companies named.
Leave a comment