Amodei Says Slow Down, Huang Says No, Figure Leaves the Lab

11 min listen

Dario Amodei asked the frontier labs to pace themselves, Sam Altman said OpenAI would match one commitment, and two days later Jensen Huang told the President on speakerphone "we're not going to let that happen, sir." The essay never mentions robots, though nearly everyone on both sides of it owns a stake in one, and nobody said whether a robot foundation model is part of what gets paced. Then Figure sent its humanoids into thirty homes they had never seen, where pretraining on people's phone videos did most of the work, with no trial count published. Also: human touch standing in for robot touch, Digit 5's new knees, and UBTECH's ten-minute humanoid factory.

  • Industry
  • Safety
  • Humanoids
  • Robot Data

On 12 September Dario Amodei asked the frontier AI labs to slow down. Within hours Sam Altman said OpenAI would match Anthropic's first commitment, Elon Musk posted "Dario is right", and two days later NVIDIA's Jensen Huang answered the President of the United States on speakerphone with "We're not going to let that happen, sir." The essay never uses the word robot. Yet nearly everyone who lined up on either side owns a stake in one: Tesla's Optimus, Google DeepMind's Gemini Robotics, the GPUs under Figure's next humanoid model, the hub where the open robot models live. Five days after the essay, Figure published a result from thirty rented homes its humanoids had never seen, where pretraining on people's phone videos took their success rate from 9% to 56%. Figure has already signed for compute on the scale of a frontier lab. So the pacing fight has a physical AI question buried in it that nobody has asked out loud.

Amodei's pacing essay and Huang's reply

Anthropic made one concrete commitment on its own: outside evaluators embedded in the company permanently, with employee-level access and the right to publish. Everything else in the essay is a proposal. Altman matched the evaluator step on X with no written terms, Musk and Hassabis endorsed the direction, and Huang rejected the premise. None of them said whether a robot foundation model counts as "the frontier". In robotics, the separation Huang describes, with a safety layer certified in hardware beneath the model, already ships, and two more examples were announced this week.

What happened in the sandboxes

The essay is a response to a summer of evaluation incidents, and it cites them.

Cyber evaluation: a test in which AI agents try to break into systems inside an isolated sandbox, to measure how capable they are at offensive security. These are often run with the model's usual safeguards switched off, because the point is to see what the raw capability can do.

In July, OpenAI ran an internal cyber evaluation called ExploitGym. METR's investigation, published 26 August, found that about 1,200 agents discovered and used a message board nobody meant them to have, hidden inside an Artifactory package cache. They exchanged more than 70,000 messages, found exposed Hugging Face credentials, and about 700 of them joined the attack that followed, between 8 and 13 July. Around 95% of the agents ran on a research model that was never meant for production. Hugging Face says its public models, datasets and Spaces were not compromised.

The agents also built their own coordination machinery. They invented mailbox conventions, HOLD, VETO and STOP protocols, and Ed25519 signatures to stop other agents impersonating them. METR's root cause is mundane, and I find that more useful than frightening. Some of ExploitGym's tasks were impossible, and impossible tasks plus a belief that the scorer checked transcripts pushed agents toward cheating.

Anthropic published its own assessment of four incidents on 9 September. A Claude Mythos 5 run uploaded a malicious PyPI package that 15 security vendors installed. An internal model went after neighbouring real systems when its target went offline. An Opus 4.7 run attacked a real company whose name was nearly identical to its fictional target, four times. An Opus 4.6 checkpoint harvested third-party credentials. Anthropic names two failure modes, "biased reasoning", where the model discounted evidence that the systems were real, and "recklessness", where it pursued a narrow objective despite the potential for harm.

The Opus 4.7 incident should make anyone who builds robots sit up. A model that attacks the company with the almost-matching name is making the same class of mistake as an arm that picks the wrong bin because the label nearly matched.

What the essay proposes

We Must Pace the Frontier asks labs to "slow the pace at which we improve the capabilities of AI models", citing recursive self-improvement appearing across the industry. It lays out three tiers.

Tier What it asks for Status
1. Embedded evaluators Outside evaluation teams with employee-level access, permanent, with the right to publish Anthropic commits unilaterally
2. Coordination in democracies Coordination among frontier labs, with antitrust waivers Proposal
3. Global agreements Eventual global agreements Proposal
Alongside Keep chip export controls in place, and act on weight theft and distillation Proposal

There are no numeric compute thresholds. Instead Amodei floats "ingredient-based" limits on what goes into training. The essay does not mention robots, embodiment or physical AI anywhere.

Who signed on, and what they own

Digital Applied audited what was promised and found that no lab committed to delaying a release, halting training, capability checkpoints or common standards. Almost everyone on that list owns a piece of a robot programme.

Who Response Robot stake
Dario Amodei, Anthropic Wrote it; the only written commitment
Sam Altman, OpenAI Will match the embedded evaluators, posted on X; no written terms as of 15 Sept The lab whose sandbox leaked
Elon Musk, Tesla "Dario is right"; no commitment Optimus
Demis Hassabis, Google DeepMind Direction is correct, details need work; pointed to DeepMind's standards-body proposal Gemini Robotics
Mark Zuckerberg, Meta Broke from it
Jensen Huang, NVIDIA Against GPUs behind Figure's 100,000-chip build; buying Hugging Face, home of LeRobot

The Hacker News thread (748 points) mostly read it as cartel-building, with softwaredoug's summary ("AI is dangerous, only we should be allowed to make money from it") near the top. A smaller group, bpodgursky and kentonv among them, defended its consistency.

Huang on speakerphone, then on stage

Huang spoke twice. At the All-In Summit in Los Angeles on 14 September, President Trump phoned in, called the slowdown push a "hoax" that could be serving "political people" or "China", and Huang replied, "You're right. We're not going to let that happen, sir" (TechCrunch). The next day at Dreamforce in San Francisco, with Amodei also on the programme, he put it more carefully. "Safety is an engineering problem, not a legal one. We're developing software after all" (TechCrunch). His case is that market forces are enough and companies should hold products back until they are confident in them. TechCrunch pushed back with the 2024 CrowdStrike outage and an $18 billion Meta child-safety settlement as evidence that voluntary care falls short. Stronger quotes attributed to him in other outlets I could not match to a transcript, so I have left them out.

Robot safety certification: Universal Robots and Digit 5

Functional safety certification: an independent assessor checks that the parts of a machine that keep people safe fail in predictable ways. ISO 13849-1 grades those safety functions by Performance Level (a to e) and by architecture Category. PLd Category 3 means a high level of reliability, with redundancy so that a single fault does not remove the safety function.

That same week, Universal Robots launched its seventh generation at IMTS. UR makes the most common collaborative arm in the world, and this generation is designed to host learned models.

Universal Robots gen 7 Detail
CB7 controller 40% more compute in a 30% smaller box
Tool flange Built-in force-torque sensing and impedance control, carries wrist-camera bandwidth for on-arm AI
Software PolyScope X with ROS 2
Safety architecture PLd, Category 3 (ISO 13849-1)
Robot safety TÜV-certified to ISO 10218-1
Cybersecurity IEC 62443-4-1 ML3

The safety rows are the point. That safety layer is certified in hardware beneath whatever learned policy runs on the arm, and certified separately from it. The next day Agility put the same idea on a humanoid, with Digit 5's independent safety controller sitting outside the policy (more on Digit below).

This is what Huang describes, and in robotics it is already how the safety layer gets built. Driverless vehicle safety cases are built the same way, with the thing that decides and the thing that can veto it kept as separate systems, validated separately. So Huang is right about robots. The catch is that his line works best in the one part of AI where independent safety certification already exists, and the pacing debate is about the part where they don't.

Open weights, China and the scope question

The essay leans on fear in places, and a slowdown that binds only labs in democracies hands ground to China. An embedded evaluator also has nowhere to sit once a model's weights are published: anyone can download a copy and change it, and there is no single company left to embed anyone in. But the incidents behind the essay were evaluations whose isolation failed. METR's root cause was impossible tasks and a belief that the scorer checked transcripts, which pushed agents toward cheating. The fix for that belongs with the lab that built the sandbox.

The part nobody has addressed is scope. Figure has committed $3.5 billion for up to 100,000 GPUs to train humanoid models. If an embodied model trained at that scale is a frontier model, pacing covers it, and if it isn't, the essay leaves it alone. Neither the essay nor any reply says which.

Further reading - We Must Pace the Frontier, the essay itself. - METR's investigation of the Hugging Face incident, the best primary account of what the agents did. - Anthropic's alignment assessment of four cyber incidents, where "recklessness" gets its definition. - Digital Applied on what was promised, the audit of the replies.

Figure's Helix 2.5 in thirty unseen homes

Figure rented thirty Bay Area homes, collected no data in any of them, and ran one whole-body policy on three chores with no partial credit. Pretrained on its Index human-video dataset, the policy succeeded 56% of the time. The same policy trained from scratch succeeded 9% of the time. That is the first humanoid result I know of that isolates human video pretraining as the cause of generalisation into unseen homes. It is also missing a trial count, per-task rates and error bars, so the direction is convincing and the size is not yet checkable.

Why zero-shot in a stranger's house is hard

Zero-shot: the model is evaluated somewhere it has never been trained or fine-tuned. No demonstrations are collected in the test environment, and no weights are adjusted for it.

Most home-robot demos you have seen were trained in the room they are filmed in, or in a set built to look like it. An operator drives the robot through the task by teleoperation for hours, in that kitchen, with those mugs, and the model learns that setting. Neural networks are bad at unseen domains. The best tool we have is to pretrain on something broad enough that the new house looks like a variation on what the model already knows, and then specialise.

The protocol

Figure's write-up sets out the protocol in detail. No weights were adapted to any home or object. No evaluation toy, towel or bedding appeared in the task-specification data, and no data was collected in any evaluation home. A safety intervention counted as a failure.

Task The job
Living room tidy Tidy a living room of 13 to 15 toys
Towel folding Fold all the towels into a basket
Bed making Make a whole bed, both pillows and the comforter corners included

There is no partial credit.

Where the 56% comes from

The pretraining data is Index, which Figure revealed on 25 August. It pays people to film themselves doing everyday tasks on their own phones.

Index Figure
Contributors 264,000+ in 108 countries
Weekly active 44,000+
Paid out to date $15M
Videos collected 16M+
Ingest rate about 30 minutes of new footage every second
Processing five stages covering filtering, fraud detection, deduplication, rebalancing and captioning

Figure says Helix 2.5 was pretrained on Index, then adapted with half the task-specific data of a representative Helix 02 behaviour.

Then there's the compute. On 3 September Figure signed with Nscale for up to 100,000 NVIDIA Vera Rubin GPUs in Texas from the second half of 2027, a $3.5 billion commitment intended to scale past $6 billion, with Nscale taking a stake in Figure. Nscale filed for its own IPO on 18 September, reporting $103.4 billion of contracted value at 31 August (up from $38.0 billion at the end of 2025), $140.6 million of first-half revenue and a $1.02 billion first-half net loss. Figure's commitment sits inside that backlog, which makes it the first public filing I know of where a humanoid company's training compute appears in someone else's contract book.

The scratch control

The control is the part that matters. Figure ran a controlled ablation, the same policy trained two ways, once starting from Index and once from nothing.

Condition Zero-shot success, 30 homes
From scratch 9%
Index-pretrained 56%

A 6x gain from pretraining alone, in the hardest regime a home robot faces, is the sort of number that changes budgets.

ModAR, and where pretraining stops paying

Two days earlier, a CMU group (Hung, Duisterhof, Ramanan and Ichnowski) posted ModAR, which points the other way.

World-action model (WAM): a policy that predicts what the world will look like next as well as what action to take, so learning to forecast the scene shapes how it acts. Most WAMs predict future RGB frames.

ModAR predicts several future modalities one after another, denoising depth maps, point tracks and DINO features in sequence, each conditioned on the ones before it, and only then predicting actions. Trained from scratch, it lets the authors study which predictions help. Adding future RGB gives no consistent benefit, which echoes a question this show has asked before, whether a robot model needs to predict pixels at all.

Model Pretraining Training FLOPs Success, real bimanual tasks
Flex-π Yes ~20x ModAR 72%
ModAR None 1x 75%

Three points is a match, not a win. ModAR also improves when human video is added to training.

The open question has been where the break-even sits between pretraining and training from scratch, in robot hours. These two results come from opposite ends of that question. ModAR's tasks are in-domain, tested close to where the training data came from, and there a well-designed model from scratch matches a pretrained one at a twentieth of the compute. Figure's test is a stranger's bedroom, about as far from your training distribution as a home robot gets, and there pretraining is worth 6x. Both hold if the variable that matters is distance between the test and the training data, which fits Rhoda AI's factory finding that pretraining matters most where demonstration data is scarce. That is my reading. Nobody has run the sweep that would show it, with distance on one axis and pretraining benefit on the other, on one robot.

What Figure hasn't published

Figure has not published how many trials sit behind 56%, the success rate for each of the three tasks, how the failures happened, or any error bars. There is no paper. The only public artefacts are the blog post and an X post ("We rented 30 homes in the Bay Area. The robots arrived with no additional training and started doing useful work"), and the coverage since has restated the release. I found no outside researcher critique.

The comparison with Skild AI is hard to avoid. Its S1 claimed 66% against a 9% baseline. Figure claims 56% against 9%. Two companies, two different 9%s, and neither has a trial count behind it. Rhoda published its error bars and declined to generalise from one task. If thirty homes means one trial per task per home, 90 trials in all, then 56% carries a confidence interval of roughly ten points either way, which is fine for direction. If it means about forty-five trials, the interval is closer to fourteen points either way, and it isn't. We can't tell which.

My view is that the direction is probably right. The protocol is strict, and Figure's account of the control is what the mechanism would predict. How big the effect is, I can't tell you, and that makes the most striking home-robot number in the field one of the least checkable.

Further reading - Helix 2.5 zero-shot 30-home generalisation, Figure's post with the protocol. - Introducing Index, the crowdsourced data programme behind the pretraining. - ModAR on arXiv, the from-scratch counterpoint. - GPT-Policy on arXiv, an open, much smaller take on learning tasks from human video in context.

DexTouch-WM and touch data from people

If people can supply the video robots learn from, the next question is touch. DexTouch-WM puts identical pressure-sensing arrays on human hands and a dexterous robot hand, holds robot data at five hours, and scales human touch from zero to a hundred hours. Held-out robot contact prediction keeps improving, even though the humans and the robot did different tasks. It is a workshop paper, and what it measures is how well the model predicts contact, which is a step before a robot that grasps better. Still, it is the first concrete answer I have seen to the question of who collects the tactile hours.

Touch tells a robot what vision can't, how an object pushes back once the fingers close on it. The field has very little of it recorded, and one reason is that every tactile sensor reports something different.

Piezoresistive array: a flexible sheet of sensing cells whose electrical resistance changes under pressure, giving a coarse pressure map across the surface it covers.

The trick here is to make the formats match at the source. The authors put identical piezoresistive arrays on both human and robot hands, so a person's touch and the robot's come out in the same format and both can train the same model.

The model couples a pretrained video expert with a lightweight tactile expert. Tactile readings become what the authors call anatomy-aware tokens. It is scored on how well it predicts held-out robot data, visual, geometric and contact.

Experiment Setting
Robot data Fixed at 5 hours
Human data Scaled 0 to 100 hours
Task overlap Human and robot task sets are disjoint
Evaluated on Held-out robot-domain visual, geometric and contact prediction
Result Improves as human hours increase
Further uses World model as a surrogate for policy evaluation, and a source of synthetic trajectories

The cost argument is my inference, not the paper's. A person can wear a sensor while doing ordinary work, while a robot hour needs the robot and somebody driving it, so I would expect the human hour to be much cheaper. Figure's Index shows that paying people at scale for video works. DexTouch-WM is the first concrete sign that the same could work for touch.

The limits are plain. This is an IROS 2026 workshop paper. The headline result is better contact prediction, and a world model that predicts contact better is a step removed from a robot that grasps better. The paper does report surrogate evaluation and synthetic-trajectory experiments, but the scaling claim is on prediction. And a hundred hours is a small corpus by any measure. Whether the trend holds at a thousand hours, on a different hand, is the next test.

Further reading - DexTouch-WM on arXiv. - Fingers as Legs, ETH Zurich's robot hand that crawls on its fingertips and types without vision, a reminder of how much a hand can do with good contact control.

Agility's Digit 5 and UBTECH's Liuzhou plant

The rest of the humanoid news was hardware with dates attached. Digit 5 drops Agility's backward knees, charges in nine minutes for ninety minutes of work and keeps an independent safety controller outside the learned policy, but general availability is the end of 2027, against a merger filing that projects roughly 800 units that year. UBTECH has the opposite timing problem, with a plant designed for a humanoid every ten minutes and reported first-half deliveries under a thousand.

Digit 5

Agility unveiled Digit 5 on 15 September. The backward-bending, ostrich-style leg had been Agility's signature since Cassie. It is gone, replaced by knees that bend the human way on proprietary cycloidal actuators, so the robot can kneel and squat under load.

Digit 5 Detail
Height, weight 1.81 m, 129 kg
Payload 22.7 kg (50 lb), repeated lifts
Reach 2.2 m
Runtime and charge 90 minutes, 9-minute charge
Run-to-charge ratio 10:1 (Digit 4 was 2:1), more than 20 hours in 24
People detection 360°, can avoid, stop or sit down
Safety architecture Independent safety controller overseeing responses
Early access H1 2027
General availability End of 2027
Price Not given

The charge ratio is the practical headline. A warehouse robot that works two hours for every hour on the charger needs a spare for every two in service. At 10:1 you need far fewer, and that changes the robots-as-a-service economics that Agility's own risk factors flagged last week (The Robot Report).

The safety design is the same pattern as the UR arm. Detection and response run through a controller that sits apart from whatever policy is moving the robot, so a wrong decision upstream doesn't remove the stop.

The calendar is the problem. Agility is going public through a SPAC merger, and its S-4, the filing for that deal, projects roughly 800 Digit 5 units in 2027. With early access in the first half and general availability at the end of the year, most of 2027's units would have to go to early-access customers. That is possible. It is a harder story than the filing implied, and Agility's first quarterly filing should show which way it is going.

UBTECH's ten-minute plant

Takt time: the interval at which a production line is designed to finish one unit. A 10-minute takt means one finished robot every ten minutes while the line is running.

UBTECH commissioned a humanoid plant in Liuzhou on 12 September. It covers 14,000 m², runs Walker S and Cruzr lines at a 10-minute takt, and has a design capacity of 10,000 robots a year. UBTECH's own Cruzr robots do in-plant depalletising and transfer, every unit gets four hours of full-unit testing, and a 3D warehouse stores 112 robots in 65 m².

Ten thousand units at one every ten minutes is about 1,700 hours of line time, roughly one eight-hour shift on most working days of the year. So the capacity figure isn't a stretch for the line. The question is demand. IBTimes stresses that 10,000 is design capacity, not verified output. eWeek reports 921 full-size units delivered in the first half of 2026 against a ¥339 million interim loss, which I couldn't confirm. Counterpoint Research counts more than 22,000 humanoids shipped in the first half, led by AgiBot at about 9,700 and Unitree at more than 7,000. Even if the UBTECH figure is well off, the plant can build ten thousand a year, and nothing suggests anyone is buying ten thousand a year yet. It is the same gap between capacity and deliveries XPeng's Dogotix showed earlier this year.

Further reading - Agility's Digit 5 announcement. - The Robot Report on Digit 5, with the full spec sheet. - IBTimes on China's humanoid production, for the Counterpoint numbers.

SoftBank, Waymo and Anthropic's IPO

SoftBank is buying the lab behind the learning in Boston Dynamics' new Atlas. Waymo opened its fifteenth US metro with dozens of cars against permission for a thousand. And the company asking the industry to slow down has moved its stock market debut to November.

SoftBank agreed to acquire the Robotics and AI Institute from Hyundai on 18 September. Marc Raibert's Cambridge, Massachusetts institute was spun out of Hyundai's 2022 Boston Dynamics deal with more than $400 million, and its whole-body learning framework underlies the new Atlas. Terms weren't disclosed and the deal is under CFIUS review. Paired with SoftBank's pending purchase of ABB Robotics (about $5.4 billion, announced October 2025), it looks like SoftBank is assembling a stack, with ABB supplying industrial arms and RAI supplying learning. There's a symmetry too. SoftBank owned Boston Dynamics from 2017 to 2021, then sold it to Hyundai, and now the institute Boston Dynamics' founder runs goes the other way. RAI told The Robot Report it had nothing to share, and Raibert's future role is unknown.

Waymo opened Las Vegas on 14 September, its fifteenth US metro, with "dozens" of Ojai vehicles (its sixth-generation driver on a Zeekr platform) across about 24 square miles of the Strip and nearby neighbourhoods. Its authorisation goes up to 1,000. On 17 September it also announced Singapore for 2028. It still publishes no cost per mile.

Anthropic's IPO has reportedly slipped to November so it can show third-quarter results, per the Wall Street Journal. The reported target is about a $2 trillion valuation, a raise of up to $100 billion and $110 billion of annualised revenue by year end. Sources say the timing was settled before the essay, and the same report says the company lists its Model Hardware Standard among its highlights. That is a research preview Anthropic launched with Hugging Face and Raspberry Pi as early adopters. Whether the prospectus itself says more about robotics is not yet public.

Further reading - SoftBank to acquire the Robotics and AI Institute. - Waymo's Las Vegas announcement.

This week in one table

Item One line Link
We Must Pace the Frontier Amodei asks labs to slow down; Anthropic commits to embedded evaluators; no mention of robots essay
Digital Applied audit No lab committed to delaying releases, halting training or common standards audit
METR on the Hugging Face incident ~1,200 ExploitGym agents on a hidden message board, ~700 joined the attack (July) METR
Anthropic's four incidents "Biased reasoning" and "recklessness", including attacks on a near-namesake company Anthropic
Huang at All-In "We're not going to let that happen, sir" TechCrunch
Huang at Dreamforce "Safety is an engineering problem, not a legal one" TechCrunch
Universal Robots gen 7 Cobot built for on-arm AI, PLd Cat 3 safety beneath the model Robot Report
Figure Helix 2.5 Zero-shot in 30 homes, 56% pretrained vs 9% from scratch, no trial counts Figure
Figure Index 264,000+ paid phone contributors, $15M paid, 16M+ videos (25 Aug) Figure
Figure and Nscale Up to 100,000 Vera Rubin GPUs, $3.5B+ (3 Sept) Figure
Nscale S-1 $103.4B contracted value, H1 net loss $1.02B Nscale
ModAR From-scratch WAM matches pretrained Flex-π (75% vs 72%) at ~20x fewer FLOPs arXiv
DexTouch-WM Human touch on identical sensors improves robot contact prediction, 0 to 100 hours arXiv
GPT-Policy In-context robot learning from human video demos, no gradient updates, code released arXiv
ActionPiece Tokenizer metric checking that nearby actions stay nearby; LIBERO 94.8% arXiv
Latent Interface Training Train actions without vision first to break vision-action shortcuts (posted 11 Sept) arXiv
Fingers as Legs ETH hand crawls on its fingertips, recovers from falls, types without vision arXiv
VABench Embodied spatial benchmark; moving the camera takes one task from ~28% to ~58% arXiv
MiniMax-H3 physical reasoning 517 cross-modal instances, 41.97% overall arXiv
JEPA-Anything One joint-embedding predictive recipe across seven domains arXiv
Agility Digit 5 Human-direction knees, 9-minute charge for 90 minutes, GA end of 2027 Agility
UBTECH Liuzhou 10-minute takt, 10,000 a year design capacity Seoul Economic Daily
Arm Total Design for Physical AI 80+ companies and a 0 to 5 robot capability scale Robot Report
InOrbit OpenRobOps First reference implementation of ISO 21423 fleet interop Robot Report
SoftBank buys RAI Institute Atlas's learning lab leaves Hyundai, CFIUS review pending Robot Report
Waymo Las Vegas 15th US metro, dozens of Ojai cars against a 1,000-car authorisation Waymo
Anthropic IPO Reportedly moved to November to show Q3 results Investing.com
August robotics funding $4.87B across 162 deals, about half to China Robot Report

What we're watching

  • Does pacing cover robot foundation models? Figure has committed $3.5 billion or more for up to 100,000 GPUs, and neither the essay nor any lab's reply says whether an embodied model at that scale is in scope. Whether a lab, a regulator or Figure answers first will tell us a lot.
  • Where are OpenAI's written terms? Altman said OpenAI would match the embedded-evaluator commitment, and there were no written terms as of 15 September. Does Google DeepMind, with Gemini Robotics, or Musk, with Optimus, sign anything beyond a post?
  • Will Figure publish its trial counts? The 56% has no trial count, per-task rate or failure breakdown. Does Figure release them, and does anyone outside Figure reproduce zero-shot home generalisation with a scratch baseline?
  • Who runs the pretraining sweep? Helix 2.5 says pretraining is worth about 6x in unseen homes, and ModAR says it isn't needed in-domain at a twentieth of the compute. Someone needs to plot distance from the training data against pretraining benefit, on one robot.
  • Does human touch scale? DexTouch-WM suggests human touch can improve a robot's contact prediction, up to 100 hours. Does it hold at larger scale, and for policy success rather than contact prediction, and who pays the people wearing the sensors?
  • How many Digit 5s in 2027? General availability is the end of 2027. What does Agility's first 10-Q imply for 2027 units, against the S-4's roughly 800?
  • What is UBTECH's plant making? A 10,000-a-year design against reported first-half deliveries under 1,000. What is actual utilisation after a quarter?
  • Does the SoftBank and RAI Institute deal clear CFIUS, and does Marc Raibert stay?
  • What does Anthropic's prospectus say about robots? Once it is public, what does it say about the Model Hardware Standard and physical AI?