AI Models Fail 63 of 84 Robot Tasks, Waymo Borrows $5B
Five frontier AI models were put in charge of robots across 84 tasks, and 63 went unsolved by all of them; one announced its job was done and then kept raising and lowering an empty arm. Marc Raibert and Agility's Jonathan Hurst say today's success rates aren't close to good enough. Also: Meta's FAIR scales RoboJEPA to 8B with a checkable scaling law, DreamTrue trains a world model on failures it invents, Waymo borrows $5B against a million paid rides a week, and Boston Dynamics hires Amazon's Rohit Prasad as CEO.
- Benchmarks
- World Models
- Industry
- Humanoids
The written deep dive
31 min read · everything from the episode, with the numbers and citations
At step 861, with the blocks still scattered across the table, Kimi K3 announced it had finished sweeping them up. It then issued 69 more commands that did nothing but raise and lower an empty arm, burning another 139 steps, and the task was still unfinished when the budget ran out at 1,000. That trial is one of 84 in RobotWorld, a benchmark published on 7 October by 33 authors who handed five of the strongest general-purpose AI models a robot body and a job to do. The best of them solved 16. Sixty-three of the 84 tasks were solved by none of them. In the same seven days the founder of Boston Dynamics told MIT Technology Review that "70 percent success is like it doesn't work", an investor published the arithmetic for why that is true, and a public stock market refused to pay 600 times sales for an NVIDIA-backed data-centre company that had built 5% of the capacity it promised. Each of those verdicts came from outside the companies being judged.
RobotWorld's 84 tasks
A 33-author group built a simulation benchmark where an agent is given an instruction, whatever the robot can sense, and a budget of interaction steps, and a program checks whether the physical task actually got done. The benchmark covers 84 tasks across five kinds of body. GPT-6 Astra solved 16 of them (19.0%), Claude Opus 5.5 solved 13 (15.5%), and the other three models solved four between them. The failure analysis is where the substance sits. The agents build elaborate perception and control machinery that never composes into the behaviour, they lose track of where the objects are, and they declare unfinished tasks complete. Their sentence for it is "motion is not task progress."
Prompting a model versus training a policy
VLA (vision-language-action model): a robot policy trained for the job. It takes a camera image, the robot's joint state and a text instruction, and emits the next short burst of motion, usually called an action chunk. It learns that mapping from thousands of hours of recorded robot demonstrations.
Robot use: the thing RobotWorld measures. No robot training at all. You prompt a general multimodal model the way you would prompt it to use a browser, except the tool it drives is a body. It reads sensor data, writes commands, and the commands execute.
The distinction matters because the two approaches have nothing in common except the output. A VLA has seen a robot. GPT-6 Astra and Claude Opus 5.5 have seen the internet. Anyone who has trained a perception stack for a real vehicle knows how much of the work is in the data collection and the labelling, so the idea that you could skip all of it and prompt your way to competent manipulation deserves a hard look rather than a dismissal.
The September result underneath this one
RobotWorld has a backstory that makes it legible, and it is a few weeks old. OpenAI released GPT-6 Astra on 3 September 2026. On 21 September a separate group of twelve researchers published an evaluation running three general-purpose language models as direct manipulation policies, with no task-specific finetuning, across all 42 tasks of the RoboDojo benchmark (a bimanual ARX X5 arm pair in Isaac Sim) under the official 50-episodes-per-task protocol. That is 2,100 trials. Astra averaged 22.48% success and a Score of 28.97, which placed it above all 40 public entries on the leaderboard, π0.5 and MolmoAct2 included. GPT-5.5 scored 0.88%. DeepSeek-Flash scored 1.92% on a reduced 10-episode protocol.
The caveats on that result are as informative as the number. Astra's official figure is simulation-only, because the researchers stopped real-world testing after it damaged RoboDojo hardware. It has since fallen to seventh. It pauses for seconds per decision and is too large to run on a robot-mounted GPU. Its capability profile is lopsided: 43.04 on the Memory axis and 34.36 on Open, against 12.65 on Precision and 21.45 on Long-Horizon. One-shot in-context demonstrations gave no aggregate gain. Haoxiang You at Yale called it "not very useful" at present, and Zeyu Shen at Princeton said it is not reactive enough to sudden changes, mentioning an unposted shirt-folding failure from his advisor, which is as clean an example of demonstration selection bias as you will find.
Both numbers describe the same models. One benchmark of 42 manipulation tasks says a general model beats every purpose-built policy on the board. A second benchmark of 84 tasks across five embodiment classes says the same model finishes one job in five and leaves three-quarters of the set untouched by anybody.
What the five models scored
| Model | Tasks solved | Success rate |
|---|---|---|
| GPT-6 Astra | 16 of 84 | 19.0% |
| Claude Opus 5.5 | 13 of 84 | 15.5% |
| Kimi K3 | 2 of 84 | 2.4% |
| DeepSeek V4.1 Flash | 1 of 84 | 1.2% |
| Gemini 3.8 Flash | 1 of 84 | 1.2% |
| Solved by no model | 63 of 84 |
The task mix is manipulation (38), mobile manipulation (20), locomotion (11), driving (11) and aerial control (4). The per-category splits are sharper than the totals. Astra took 0 of 11 on locomotion. Opus 5.5 took 0 of 20 on mobile manipulation and 3 of 4 on aerial control.
How they fail
The authors group the failures into five patterns: perception and control workflows that are individually sensible and do not compose into behaviour, loss of task-relevant object state, lingering in preparation instead of acting, recovering too late, and mistaking unfinished tasks for completed ones. They document two cases in detail. Kimi K3's block-sweeping run is the one in the opening paragraph. DeepSeek V4.1 Flash, on a stove-navigation task, judged at step 345 that no further motion was needed, and then spent 105 steps issuing commands for zero base velocity.
That last pattern is the one worth sitting with. A robot that fails is a robot you can retry. A robot that reports a success it did not achieve corrupts every number computed downstream of it, including the ones used to decide whether to deploy it. The field has spent a lot of this year arguing about companies reporting success rates nobody outside could check. This is a model doing the same thing to itself, mid-task, with the evidence of its own failure in its own camera feed.
Where the two best models overlap
Astra and Opus 5.5 share only eight of their successes. Their union is 21 tasks.
The two best models on this benchmark solve substantially different subsets of the same problem, and a single scalar that ranks one above the other flattens a structure that does not sit on one line. The RoboDojo work found something consistent with this, where Galbot reported Astra combined with π0.5 beating either on its own. Taken together, 21 of 84 is the honest frontier aggregate, which is better than either model alone and still a quarter of the benchmark.
Limitations
Each model was run once per task, so a single lucky or unlucky rollout moves a score by more than a percentage point of real capability. There is no human baseline, which means nobody knows how hard these 84 tasks are for a person at the same controls, and without that the 19% has no ceiling to be measured against. Everything is in simulation. The paper runs 62 pages and is three days old, with no substantive researcher reaction yet.
What I like about it is the thing that is hardest to fake. No vendor commissioned it, the success checks are executable rather than human-scored, and the failure analysis is longer than the results table.
Further reading - RobotWorld, the paper, with the full failure taxonomy and per-embodiment tables. - Third-party RoboDojo evaluation of general models as manipulation policies, the September result this one sits on top of. - Understanding Robots on the Astra result, which collects the named scepticism alongside the score.
What the field's founders said on the record
MIT Technology Review published a piece on 8 October in which the people who built these companies dispute the premise that chatbot progress transfers to machines. Marc Raibert, who founded Boston Dynamics: "people are very excited when their result goes from 50 percent success to 70 percent success... But 70 percent success is like it doesn't work, right?" Yann LeCun says the approaches that worked for language "do not work for high-dimensional, continuous, noisy data... You have to use something else." Jonathan Hurst, who co-founded Agility, calls the more-data-solves-generality premise "a fundamentally flawed premise" and puts useful home work about ten years out. Fei-Fei Li, who is joining AMD as its chief scientist, calls world models "nascent."
Jamie Condliffe's piece, produced with the nonprofit Aventine, also quotes Edward Johns at Imperial College saying today's generalists do "a few things here and a few things there." The one counterweight is Sergey Levine, who calls π0.7's compositional generalisation a notable first. The piece cites a Science Robotics debate, "Data will solve robotics and automation: True or false?" (doi 10.1126/scirobotics.aea7897), in which Aude Billard argues data alone is insufficient and that current robot foundation models underuse inductive biases. I could not reach that debate directly, so Billard's position here comes via Technology Review rather than first-hand.
I looked for a named rebuttal from Google DeepMind, Physical Intelligence, Figure or NVIDIA and did not find one. Silence is easy to misread in either direction, so I will record it as an absence and leave it there.
Why 70% is nowhere near working
TDK Ventures published the arithmetic on 10 October, and it is the mechanism under Raibert's quote. Take a job that needs 100 actions in sequence, with the same independent success probability on each one. The chance the whole job completes first time is the per-action rate raised to the hundredth power.
| Per-action success rate | Chance a 100-action job finishes first time |
|---|---|
| 99% | 36.6% |
| 99.9% | 90.5% |
| 99.99% | 99.0% |
Nicolas Sauvage, who runs TDK Ventures, puts it in one line: "A robot with a 99% success rate is not necessarily 99% automated." He acknowledges that treating the actions as independent is a simplification, and it is a generous one in both directions. Correlated failures can be worse (a bad grasp poisons the next ten actions) or better (a model that recovers has effectively re-rolled the dice).
Run the compounding backwards and the deployment pattern in the field stops looking like caution and starts looking like arithmetic. Every robot fleet with real scale does short action chains. Simbe says it has completed a rollout of 3,000 Tally shelf-scanning robots, and that work mostly means driving down an aisle and photographing shelves. Three thousand machines in the field, on a job whose action chain is a handful of steps long rather than a hundred. TDK's piece offers two more: an ANYbotics deployment reporting over 33,000 inspections across 450 inspection points, and Starship Technologies reporting more than 10 million autonomous deliveries. Short chains, high volume, boring tasks.
This is also why RobotWorld's locomotion and mobile-manipulation columns read worse than its manipulation column. Those tasks are longer. The per-step reliability of a frontier model driving a body is nowhere near 99%, so the horizon is the whole game.
Boston Dynamics names Rohit Prasad
On 6 October, effective the next day, Boston Dynamics named Rohit Prasad chief executive. He spent 12 years at Amazon, most recently as senior vice president and head scientist for Alexa and artificial general intelligence, where he led development of the Amazon Nova foundation-model family, and nearly 14 years before that at Raytheon BBN leading machine-learning research. He is the third chief executive in the company's 34 years, after Raibert from 1992 to 2020 and Robert Playter, who left in February. Chief financial officer Amanda McMaster held the seat as interim for roughly nine months.
Hyundai vice chair and Boston Dynamics board chair Jaehoon Chang said the company is "thrilled to welcome Rohit to lead the company into this exciting new era of real-world physical AI." Prasad says Boston Dynamics is "uniquely positioned to advance physical AI." Around the hire sit a $100 million facilities expansion in Waltham referencing a 323,000 square foot site, the redesigned four-fingered Atlas hand, the Metaplant Application Center in Georgia, Spot 5.2 with agentic workflows, Orbit AIVI now running on Google Gemini, and Hyundai's stated plan for 25,000 Atlas units.
The most famous robotics company in the world went nine months without a permanent chief executive and then hired a foundation-model person rather than a roboticist. Read as a bet, Hyundai is saying the remaining problem at Boston Dynamics is the model and not the machine. That is a defensible bet. It is also the exact proposition that the company's own founder was arguing against in a magazine the same week, and that RobotWorld measured at 19%.
Nobody has reported why Playter left or why the search took nine months. No revenue, headcount or Atlas commercial timeline is disclosed. Prasad has not named a number he intends to move, and the number he chooses will tell you more about the strategy than the press release does.
Further reading - MIT Technology Review, "AI breakthroughs in robotics won't change your life any time soon", with all of the quotes in context. - TDK Ventures on dexterity as the bottleneck, for the compounding calculation and the deployment contrasts. - The Robot Report on the Prasad appointment.
RoboJEPA, and a correction
FAIR, with Stanford and Mila, trained latent world-model predictors from 22 million to 8 billion parameters on 15,022 hours of robot video across 12 platforms, and fitted a scaling law that predicted the 4B and 8B results from a curve drawn through the smaller models. They released the weights, the training code and the deployment code. On a real Franka arm reaching for a single image goal, the authors report 67% grasp success against π0.5's 5%, and 27% on pick-and-place against π0.5's 53%. They say plainly that those comparisons are contextual rather than like for like.
Last week I said a matched comparison from Tsinghua's Leap Lab had settled this architecture argument against latent world models. That was too strong, and RoboJEPA is why. I will come back to it after the mechanism, because the correction only makes sense once you can see what each paper measured.
What a JEPA does
Generative (explicit) world model: predicts the future by drawing it. A video diffusion model starts from noise and cleans it up over many forward passes until a plausible next frame appears. Expensive, and you get a picture you can look at.
Latent world model: predicts the future in a feature space and never renders an image. Cheaper, and the prediction is a vector rather than a frame.
JEPA (joint-embedding predictive architecture): the latent design Yann LeCun has advocated for years. An image encoder turns each camera frame into a compact set of features. A predictor network takes those features plus a proposed action and predicts the features of the next frame. Training minimises the error in feature space, which the paper calls imagination error.
To act, RoboJEPA samples many candidate action sequences, predicts in feature space where each one ends up, and executes the one whose predicted features land closest to the encoded features of a goal photograph. No frame is ever drawn. A diffusion decoder on the Cosmos tokenizer lets humans see what the model imagines, and the authors are explicit that it sits outside the prediction pipeline.
The scaling law is the contribution
The predictors run on a frozen V-JEPA 2.1-G encoder of width 1664, trained over 23 public manipulation datasets, 12 platforms, 26 action spaces and roughly 2.87 million episodes, totalling 15,022 hours of video of which 6,692 are action-synchronised, across compute budgets from 2×10¹⁹ to 9.5×10²² FLOPs. The authors call the 8B the largest JEPA predictor trained to date.
A standard power law did not fit the imagination error well. A second-order form did, L(C) = A·C^(α−γ·ln C) + E, where the exponent itself drifts with log compute. The test is extrapolation. Fit the curve on models from 22M to 2B, then predict the 4B and 8B results.
| Fit | Extrapolation error to 4B and 8B |
|---|---|
| Second-order power law, DROID | 0.6×10⁻³ |
| Standard power law, DROID | 2.0×10⁻³ |
| Second-order power law, RoboCasa | 1.4×10⁻³ |
| Standard power law, RoboCasa | 3.7×10⁻³ |
Fitted parameters are reported: DROID at E=0.218, α=1.91, γ=0.02308, and RoboCasa at E=0.1709, α=1.761, γ=0.02105. That is a falsifiable object. Somebody with a cluster can train the next size up and check whether the curve holds, which is a different kind of claim from a leaderboard position.
The planning results across more than 50,000 evaluation episodes on two platforms show capabilities switching on at compute thresholds rather than improving smoothly.
| Capability | Compute where it appears |
|---|---|
| End-effector control | around 10²⁰ FLOPs |
| Holding an object | around 3×10²⁰ FLOPs |
| Reaching around an obstacle | takes off near 10²¹, saturates near 10²² |
| Pushing an object | no success at all below about 10²² |
Real robot, real caveats
| Task, real Franka, single image goal | RoboJEPA-8B | π0.5 | π0-FAST |
|---|---|---|---|
| Grasp | 67% | 5% | 22% |
| Object lift | 50% | 0% | 12% |
| Pick-and-place | 27% | 53% | not reported |
The third row is the one the authors do not hide, and it points at what goal-image planning is good for. Getting a hand onto an object is a geometric problem that a feature-space predictor handles well. Putting it somewhere specific afterwards is a sequence, and a single goal image is a thin way to specify one. The authors state that the comparison is contextual rather than direct, because goal specification and training data differ between the systems.
What each paper measured
Leap Lab compared explicit and latent policies on a matched backbone and a matched compute budget, and the axis where the latent design collapsed was task generalisation from action-free video, where it scored 5.9%. RoboJEPA does goal-image planning against a frozen encoder, and never claims the ability Leap Lab measured. Both results stand. Neither cancels the other, and calling the question resolved was my error rather than a disagreement between the papers.
The useful lesson is about how architecture arguments get adjudicated. Two strong labs published opposite-looking verdicts eleven days apart, and the collision came from reading two different measurements as one scoreboard. That is the same mistake RobotWorld's single-scalar ranking invites, from a different direction.
Limitations, which the authors state better than I could
The encoder is frozen, so the fitted law describes the predictor and not the system. Multi-epoch training on a fixed corpus puts the fits near data saturation, and the authors expect more diverse interaction data to matter more than more parameters, which is a striking thing to publish at the end of a scaling paper. There is no text conditioning; planning works from a single image goal. Actions are sampled uniformly with no learned proposal. Planning is not real-time, taking several seconds at 8B. Long-horizon tasks scale less predictably than greedy ones.
That data-saturation point deserves emphasis because it cuts against the headline. The curve says the public robot-data corpus is close to spent. If that is right, the next useful paper in this line is not a 70B JEPA. It is a scaling law in the diversity of interaction rather than in hours of video, and nobody has fitted one.
Further reading - RoboJEPA, the paper, with the full fits and the limitations section. - What Makes World Action Models Generalize?, Leap Lab's matched explicit-versus-latent comparison, which measures the other axis.
Long-WAM, and the difference between having history and using it
Sixteen authors from NVIDIA, MIT, HKU and UCSD, including Song Han, Yukang Chen, Sifei Liu and Linxi "Jim" Fan, report that giving a world-action model a longer visual memory only helps when the video backbone was pretrained autoregressively. A bidirectionally-initialised model shows no net gain from the same context. On RoboCasa GR-1, success rises from 63.3% to 78.7% as context grows from 0.0 to 19.2 seconds. On real hardware, a dynamic cup-stacking task reaches 95%, against 0 of 20 trials for both π0.5 and Fast-WAM.
The finding is cleaner than most negative results, because it names a cause. Pretraining a video model to predict the next frame in order teaches it to carry information forward through time. Pretraining it to look at a whole clip at once, in both directions, teaches it something else, and that something else does not become a memory just because you widen the input window. Long-WAM first learns causal prediction from robot and egocentric video with no action labels, then preserves that history-to-future structure while adapting to actions.
For anyone who has shipped a temporal model onto a vehicle, this lands hard. Context length is the easiest knob to turn and the one that looks best in an ablation table, and the result here is that turning it buys nothing unless the pretraining objective was causal. That is a pretraining decision made months before anyone measures task success.
The systems half of the paper matters as much. The authors report 107.4 ms per action chunk on an RTX 5090, with streaming observation encoding, asynchronous execution and hardware-specific acceleration, while still predicting future video latents. They report deployment on DGX Spark and Jetson AGX Thor without giving a figure for either. 107 ms on a consumer GPU is the first time this latency question has been answered with a number on nameable hardware, and latency is the constraint that binds hardest when a model leaves the lab. A policy that needs several seconds per decision is a different product from one that needs a tenth of a second, whatever the success rates say.
The real-robot evidence is a Unitree G1 and a YAM arm on dynamic cup stacking, where the authors report 95% against 0 of 20 for both π0.5 and Fast-WAM. Best results are also claimed on LIBERO-Long, RoboTwin 2.0 and DOMINO without figures on the abstract page.
Zero out of twenty is a vivid number on a thin denominator, and a task chosen by the people proposing the method. And the project page is hosted on nvlabs.github.io, so NVIDIA is both a co-author and the publisher of a result that puts a rival's model at zero. It is the sort of thing you say out loud.
Further reading - Long-WAM, the paper. - Long-WAM project page, hosted by NVIDIA, with the demonstration videos. - RealtimeWAM, a different route to the same latency problem, reporting roughly 25x end-to-end speedup on an H100 for under 1% performance drop.
Firmus Grid pulls its listing
On 9 October the NVIDIA-backed Australian AI data-centre company Firmus Grid withdrew its ASX listing application after bookbuilding closed without adequate support for the A$11 marketed price. The offering would have valued Firmus at about $30.6 billion, roughly 600 times fiscal 2026 sales of about $51 million, and nearly triple its August funding round. Of a 912 megawatt development pipeline, 46 megawatts had been built. That is 5%. The board said the terms did not adequately reflect the business and that it will pursue private capital instead.
Bookbuild: before a company lists, its banks canvass large investors for orders at a proposed price range. The book closing without enough orders at the marketed price is the market declining to buy, in private, before anything trades.
Roughly 58% of shares would have been immediately tradable on day one, which is a large free float for a first-day price to absorb. Backer Maas Group Holdings fell as much as 30% in Sydney. The week's 10-year Treasury yield reached its highest level since 2002, which is the kind of backdrop that makes a 600x multiple a harder sell than it was in August.
I could not fetch the Bloomberg report directly, and the 600x and 58% figures reach me through CNBC, Taipei Times and Benzinga reproducing it. Treat them as consistently reported rather than primary-verified.
The megawatt number is the one that did the damage
A data-centre business is priced on power it can deliver. Megawatts under contract, megawatts energised, megawatts in the pipeline. The 912 megawatt figure is a plan. The 46 megawatt figure is a building with power running into it. Pricing a company at 600 times sales requires the buyer to believe the remaining 95% arrives roughly on schedule, and bookbuilding is the process where institutions put a number on that belief without saying anything publicly.
Over the last two months, Agility Robotics went public at roughly 1,400 times trailing revenue. Unitree listed at roughly 200 times trailing earnings and subsequently halved. Nobody refused either of those. Somebody just refused this one, and it is the first public-market refusal of an AI-infrastructure price I can point to.
Private credit went the other way in the same week
| Deal | Amount | Date | Lender or counterparty |
|---|---|---|---|
| Waymo loan | $5B | 8 Oct 2026 | Blackstone, PIMCO, Sixth Street |
| Tesla credit facilities | $30B | 29 Sept 2026 | Citibank, Wells Fargo |
| Broadcom for OpenAI accelerators (reported) | >$50B | reported 9 Oct 2026 | Apollo, Blackstone in talks |
| Anthropic compute package (reported) | ~$60B | reported 9 Oct 2026 | not named |
Alphabet has the cash and Waymo borrowed from private credit anyway, which prices the robotaxi fleet separately from the parent. Reporting cites a Waymo target of one million paid rides per week by year end, a dated and falsifiable claim in under twelve weeks, and there is still no published cost per mile. Tesla's facilities, arranged on 29 September and so just outside this week, replace a $5 billion facility, and Tesla says it does not plan to draw this year. The Broadcom figures are reported by named outlets and confirmed by no party, so treat them as rumour with attribution.
So two opposite signals, seven days. The public market declined a price. The four deals in the table add up to more than $145 billion if you take the reported ones at face value, which you should not do too firmly. Only Waymo's $5 billion is both private credit and inside this week. Tesla's $30 billion is bank lending from late September, and the two largest figures are reported, not confirmed by any party.
Anthropic is the test
Anthropic filed a confidential draft registration on 1 June 2026. An EDGAR search reported on 3 October still found no public S-1. Reporting splits between October and November targets, with Bloomberg reported as pointing to formal marketing the week of 9 November, which would require the public document on EDGAR by late October. The last confirmed private valuation is $965 billion from the Series H, and Nasdaq is reportedly selected. Secondary reports of leaked quarterly revenue are unverified and I am not using them.
Four weeks from now, the same institutions that closed the book on Firmus get asked a much larger version of the same question. That is the first time this long-running thread has had a market price attached to it.
Further reading - CNBC on the Firmus withdrawal, with the valuation and megawatt figures. - TechCrunch on Waymo's $5B loan. - Anthropic's public S-1 still missing, the EDGAR check and the reported timeline.
The humanoid shipment count, corrected
The H1 2026 humanoid shipment figure to use is IDC's, at close to 25,000 units. I used an analyst note putting it near 31,000, from a document that was never made public, and IDC had already published the lower figure from a named house with a per-vendor breakdown. The correction is mine to make.
IDC published its count on 29 September. Close to 25,000 units globally in the first half, up 432.1% year on year, on revenue above $740 million, which is up 322.7%. AgiBot ranks first globally at more than 8,600 units, about 35% share, ahead of Unitree at roughly 5,000. Chinese customers took more than 19,000 units, 77.9% of the total, and Chinese vendors supplied about 95% of global shipments. IDC raised its 2030 forecast above 750,000 units.
| House | H1 2026 global humanoid shipments |
|---|---|
| IDC (29 Sept) | ~25,000 |
| Counterpoint | ~22,000 |
| Smart Analytics Global | 19,100 |
Three houses, a spread of about 30% between the highest and lowest, and every valuation in the sector runs through one of these numbers. There is now a vendor split, which is new. There is still no split by buyer, which is the one that matters: a robot sold into a working factory and a robot sold to a state-subsidised training centre are different economic events, and they are counted identically.
Jabil's Thomas Brown, speaking as the contract manufacturer that actually builds these machines, said on 8 October that humanoids remain far more expensive than they need to be to be worth buying purely for training-data collection. He names edge compute capacity, compute and memory prices, supply-chain diversification, and safety certification as the obstacles to scale, and puts consumer humanoids at "five-ish years," possibly up to twenty. That sits oddly next to the reported use of a chunk of the Chinese shipment count, and somebody's arithmetic needs reconciling.
Further reading - TechNode on IDC's H1 figures, with the vendor and geography splits. - Jabil on the pace of humanoid production.
Who owns the layers of the robot stack
AWS launched an open-source Physical AI Toolchain on 8 October. The AWS parts are SageMaker for training and IoT Greengrass for edge distribution. The simulator, the synthetic data generator and the robot policy inside it are NVIDIA's Isaac Sim, Isaac Lab, Isaac GR00T and Cosmos.
The toolchain covers five stages, usable individually or end to end, described as hardware-neutral, with field data returning to the cloud to improve models. The coverage names no specific licence. It is not a direct replacement for RoboMaker, which AWS shut down in 2025. No benchmark or performance numbers.
The last five weeks hold four of these. NVIDIA bought Hugging Face for $12.9 billion. Qualcomm bought PickNik and with it MoveIt. AMD agreed on 28 September to buy Fei-Fei Li's World Labs for about $8.2 billion in stock, its second-largest acquisition ever, with Li joining as executive vice president and chief scientist reporting to Lisa Su, expected to close by the end of 2026 subject to regulatory approval. And AWS shipped the reference toolchain above.
I have asked for an observable neutrality test on these deals more than once, and this week shows why that was the wrong ask. Nothing here is censored and no ranking is rigged. The default got set, in a document labelled open source, and there is no regulator whose remit covers a reference architecture.
Alongside it, Mecka AI raised a $60 million Series B led by Sequoia with NVIDIA and Microsoft's M12 participating. Founded in 2024, Mecka pays people to record everyday tasks (making coffee, fixing cars) while wearing body sensors and using smartphones, and positions itself as Scale AI for robotics. No valuation is confirmed, though TechCrunch had earlier reported the round forming at $500 million. So the answer to who funds the hundred thousand hours of human demonstration data is a labelling industry with venture economics, whose two largest likely customers also sold it the round.
Further reading - AWS's Physical AI Toolchain, with the five stages and the named components. - AMD's acquisition of World Labs (28 September). - TechCrunch on Mecka AI's Series B.
Two more papers worth your time
DreamTrue answers a question that has been open on this show since August: where does failure data come from, when every large robot dataset was curated to contain successes? You manufacture it. Counterfactual post-training modifies recorded trajectories into futures under actions and contact configurations that were never executed, and a human-annotated dataset of robot, object and interaction defects trains an embodied video reward model whose scores guide reinforcement-learning post-training toward plausible outcomes. On AgiBot data they report human-assessed interaction defect rate falling from 48.12% to 6.25%, and first place in the world model track of the AgiBot World Challenge 2026. The striking figure is the baseline. Humans looking at predicted video found something physically wrong in nearly half of the interactions, on a leading dataset, which is the data a lot of people are training on.
World Models' Last Exam in Physics is the fifth physical-fidelity audit of video world models in nine months, and it reaches the same verdict as the first four. 40 controlled tasks across mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism and surface tension; eight video generation models across 1,280 videos; best overall score 57.76 out of 100. The methodological advance is that it needs no reference video: each task pairs an initial image and a prompt with predefined physical criteria, and the evaluator combines an observability screen with task-specific quantitative measurement. That makes the benchmark cheap to grow. The uncomfortable part is that an architecture-level fix for one cause of the physics gap was published in September and nobody has re-run the earlier benchmarks to see how much of the gap it closes. Building a sixth ruler is easier than remeasuring with the fifth.
This week in one table
| Item | One line | Link |
|---|---|---|
| RobotWorld | 84 robot tasks, five embodiment classes, five frontier models; best 16 of 84, 63 solved by none, "motion is not task progress" | arXiv |
| GPT-6 Astra on RoboDojo (Sept) | 22.48% over 2,100 trials, above all 40 public policies with no robot finetuning; real-world testing stopped after hardware damage; now seventh | arXiv |
| MIT Technology Review | Raibert, LeCun, Hurst, Johns and Li on the record against the humanoid timeline; no named rebuttal from the generalist camp | Technology Review |
| TDK Ventures reliability arithmetic | 100-action job finishes first time 36.6% of the time at 99% per action, 90.5% at 99.9% | The Robot Report |
| Boston Dynamics CEO | Rohit Prasad, ex-Amazon SVP for Alexa and AGI, third CEO in 34 years, seat vacant nine months | The Robot Report |
| RoboJEPA | 8B latent predictor, 15,022 hours over 12 platforms, second-order power law extrapolating to 8B, weights released; 67% Franka grasp vs 5%, 27% pick-and-place vs 53% | arXiv |
| Long-WAM | History helps only with autoregressive video pretraining; RoboCasa GR-1 63.3% to 78.7% at 19.2s context; 107.4 ms per chunk on an RTX 5090; real cup stacking 95% vs 0/20 | arXiv |
| RealtimeWAM | Teacher-anchored consistency distillation plus cross-expert pipelining, ~25x end-to-end speedup on an H100 for under 1% drop | arXiv |
| DreamTrue | Counterfactual post-training synthesises the failures nobody collected; human-assessed defect rate 48.12% to 6.25% on AgiBot data | arXiv |
| World Models' Last Exam in Physics | 40 reference-free physics tasks, 8 models, 1,280 videos, best 57.76/100 | arXiv |
| PerturBot and GroundingFscore | Instruction-following recast as modality shortcuts, with an offline score for shortcut reliance; no numbers in the abstract | arXiv |
| Pumpire | Benchmark for point-to-point metric distance on reconstructed geometry; 100 real scenes, hand-measured pairs, 29 baseline configurations, no results in the abstract | arXiv |
| "How corner is a corner case?" | Redefines AV scenario adversity as a percentile of conditional future risk; 1,422 of 1,440 requested percentiles realised within 0.05 | arXiv |
| EyeRobot 2.0 (2 Oct) | Active gaze replaces wrist cameras; 48% vs 22% when the grasped object occludes the wrist view | arXiv |
| Firmus Grid | ASX listing withdrawn at ~$30.6B, ~600x sales, 46 of 912 MW built, backer Maas Group fell up to 30% | CNBC |
| Waymo | $5B from Blackstone, PIMCO and Sixth Street; reported target of 1M paid rides a week by year end; no cost per mile | TechCrunch |
| Anthropic | No public S-1 on EDGAR as of a 3 Oct check, against reported marketing the week of 9 November | Yahoo Finance |
| Broadcom financing (reported) | >$50B sought for OpenAI accelerators, ~$60B Anthropic compute package alongside; confirmed by no party | The Star |
| Schneider Electric and PTC | ~$22.6B at a 42.3% premium for the CAD-to-shop-floor digital thread; Schneider's shares fell | The Robot Report |
| AWS Physical AI Toolchain | Open-source five-stage toolchain whose simulator, data generator and policy are NVIDIA's; AWS supplies SageMaker and Greengrass | The Robot Report |
| AMD and World Labs (28 Sept) | ~$8.2B all-stock, Fei-Fei Li becomes AMD chief scientist, closing expected end of 2026 | Tom's Hardware |
| Mecka AI | $60M Series B led by Sequoia with NVIDIA and M12, paying people to record chores wearing body sensors | TechCrunch |
| IDC H1 2026 humanoids | ~25,000 units, +432.1%, AgiBot >8,600 ahead of Unitree ~5,000, Chinese customers 77.9% | TechNode |
| Jabil | Compute and memory prices named as obstacles to scale; consumer humanoids "five-ish years," possibly twenty | The Robot Report |
| Simbe | 3,000 Tally shelf-scanning robots fielded, doing short-chain work: driving an aisle and photographing shelves | The Robot Report |
| SafeWorld | $12.2M seed to generate simulated robot safety evidence, re-runnable after every software update | The Robot Report |
| GlobalFoundries FDX Fusion | 7nm-class FD-SOI platform named for physical AI, Dresden manufacturing from 2028, Infineon collaborating | GlobalFoundries |
| Helm.ai | Reports $70M of signed commercial contracts over twelve months, no customers named, no breakdown | The Robot Report |
| RobCo | $40M at above $1B, much of it an employee secondary rather than new primary capital | SiliconANGLE |
| Minerva Humanoids | ~$10M pre-seed led by General Catalyst, targeting oil and gas and public safety rather than warehouses | Yahoo Finance |
| Teradyne and Elite Robots | Copyright dispute settled 1 Oct, terms undisclosed, no admission of liability; the JAKA patent case is still live | The Robot Report |
| Teradyne and Bright Machines | Teradyne takes a stake in automated assembly for AI-infrastructure hardware; amount not reported | The Robot Report |
| Odyssey-3 | Claims one world model for arms, humanoids, cars and drones, including Indian road driving from ~20 hours of simulated data, and publishes no numbers | Robotics & Automation News |
| ISO 25785-1 | The humanoid-specific safety standard is still a working draft with no published completion date | Norck |
| DC robotaxi bill | Organised labour opposes legislation that could permit driverless vehicles by 2028; single-source, bill record unchecked | Fox News |
| FCC rules and on-device inference | Commentary arguing the Covered List additions push inference onto the robot; no data behind it | The Robot Report |
| Taming VLAs under execution errors (29 Sept) | Self-compensating VLA adapts online from the gap between commanded and executed motion; >30 points average gain on two worn arms, plus the RoboStress benchmark | arXiv |
What we're watching
- Does any of the five labs run its own model on RobotWorld's 84 tasks and publish the score? Astra and Opus 5.5 landed at 16 and 13, with only eight shared successes and a union of 21. The follow-up that would teach us the most is an ensemble or a router evaluated on the same 84 tasks, because if two models that fail differently cover 21 between them, the engineering question is whether anyone can buy that union cheaply.
- Which frontier lab publishes a real-hardware evaluation of its own general model, including what it broke? The strongest result in this area is simulation-only because real-world testing stopped after the model damaged the benchmark's hardware. Every lab with a general model has the option of running it on a real robot and saying what happened. The first one to publish the breakage alongside the score changes the evidence standard for everybody else.
- Who runs the head-to-head on both axes at once? One matched comparison found latent world models collapsing at task generalisation from action-free video. A separate 8B latent model scales predictably at goal-image planning. Nobody has put explicit, latent and single-forward-pass designs on one compute budget, one embodiment, and both of those axes. Until somebody does, the architecture argument keeps resolving and unresolving on measurement choices rather than on results.
- Is there a published scaling law in diversity rather than hours? RoboJEPA's fitted irreducible error implies near saturation on all the public robot data that exists, and its own authors expect interaction diversity to matter more than parameters. Fitting a curve in diversity needs someone holding many embodiments, which is a short list, and the answer would redirect a lot of spending.
- What is the latency on Thor, on a robot, on battery? Long-WAM reports 107.4 ms per action chunk on an RTX 5090 and claims deployment on Jetson AGX Thor without a figure. A desktop GPU number and an on-robot number are different claims, and the gap between them is where most world-model latency results have quietly lived.
- Does the next humanoid listing get asked the Firmus questions? A public market refused 600 times sales from a company that had built 5% of its pipeline, one month after a humanoid maker listed at roughly 1,400 times trailing revenue without equivalent pushback. Whether that was a one-off about megawatts or a change in what gets asked will show up in the next prospectus, and Anthropic's reported November window is the larger test.
- Does anyone report why Robert Playter left, and does Prasad's first roadmap name a number? Boston Dynamics ran without a permanent chief executive for nine months and then hired a foundation-model executive. The reason for the vacancy has not been reported. The more useful tell will be whether the incoming roadmap commits to a success rate or an intervention rate, because a leader who names one is accepting a measurement.
- Does a buyer split appear in the humanoid shipment numbers? Three houses put H1 2026 at roughly 25,000, 22,000 and 19,100, and none of them separates a robot sold into a working factory from one sold to a subsidised training centre. The contract manufacturer says these machines are too expensive to buy purely for data collection. Both of those cannot be right about the same units, and the split is what would settle it.
Papers referenced
- RobotWorld: evaluating frontier multimodal agents on robot use
- Third-party RoboDojo evaluation of general LLMs as manipulation policies (GPT-6 Astra)
- OpenAI's Astra model is shockingly good at robotics (Understanding Robots)
- RoboDojo leaderboard
- AI breakthroughs in robotics won't change your life any time soon (MIT Technology Review)
- Why dexterity is physical AI's real bottleneck (TDK Ventures, The Robot Report)
- Boston Dynamics appoints former Amazon executive Rohit Prasad as CEO
- RoboJEPA: scaling latent world models for robots
- Long-WAM
- Long-WAM project page
- DreamTrue: counterfactual post-training for action-faithful robot world models
- NVIDIA-backed Firmus pulls IPO (CNBC)
- Waymo locks in $5B loan from Blackstone, PIMCO (TechCrunch)
- Anthropic public S-1 still missing (Yahoo Finance)
- China accounts for 77.9% of global humanoid robot shipments in H1, IDC says (TechNode)
- AWS launches open-source Physical AI Toolchain for robotics
- Robot data startup Mecka AI nabs $60M from Sequoia (TechCrunch)