Agility $1.8M of Robots, Skild $100M ARR, Depth for Free

11 min listen

A humanoid maker's books went public for the first time under securities law, and they show $1.8 million of robot sales against a $2.5 billion merger valuation, while the company selling robot software rather than robots reported fifty-five times that in recurring revenue. In the same seven days, three groups published three incompatible answers to where a robot's imagination should come from, and none of them cites the others. Also: a model trained to edit pictures turns out to be better at depth than the models built for depth, and four tactile datasets landed at once, with the largest still short of what its own authors say the field needs.

  • Industry
  • World Models
  • Robot Data
  • Perception

One point eight million dollars. That is what a Western humanoid robot company sold last year, written down in a securities filing where being wrong is actionable, by a company merging at a valuation of two and a half billion. In the same seven days, the company that sells robot software instead of robots said it had reached $100 million of recurring revenue, and three research groups published three incompatible answers to where a robot's model of the world should come from. So the week priced two things at once: the body, in dollars, on file, and the knowledge you put inside it, in hours of data and dollars of compute. This week put a real number on both sides of that: what the hardware is worth, and what it costs to teach it anything. And the hardware came out looking worse.

Agility Robotics files a humanoid P&L

Agility's S-4 is the first time a Western humanoid maker's profit and loss statement has gone on file under US securities law. It reports $1.8 million of 2025 net sales and a $140 million operating loss, against a merger valuing the company at $2.5 billion, about 1,400 times trailing revenue. The operating record underneath is much better than the sales line: nine customer sites, 65,000-plus logged hours, over $300 million of booked orders. And the filing's own risk factors say the rental model has not been shown to make money at scale, which I respect a document for writing down.

Every humanoid valuation in this sector, the private ones included, has been marked against numbers nobody outside the companies had seen. That changed on 7 September, when Agility filed to go public by merging with Churchill Capital Corp. XI, a listed shell.

What a SPAC has to disclose

SPAC: a listed company holding cash and no operating business, which takes a private company public by merging with it. The shell already trades, so the private company inherits a listing without running a conventional IPO.

The route matters for the paperwork it forces. A merger needs a registration statement and a proxy explaining the deal to the shell's shareholders, and that document carries management's multi-year projections alongside the audited past. A conventional IPO prospectus rarely includes projections at all.

The numbers in the filing

Item Figure
2025 net sales $1.8M
2025 operating loss $140M
Operating expenses $111M, up from $71M in 2024
Cash burn ~$100M
Merger valuation $2.5B, roughly 1,400x trailing revenue
Gross proceeds expected $620M+
From the Churchill trust ~$420M
PIPE, led by Foxconn $200M
Customer sites 9
Cumulative Digit operating hours 65,000+
Booked multi-year Digit v5 orders $300M+
Projected Digit v5 units ~800 in 2027, 7,000 by 2030, 25,000 by 2035

Named customers include Amazon, GXO, Toyota Motor Manufacturing Canada, Schaeffler and Mercado Libre. Nine paying sites and sixty-five thousand hours of machine time are numbers you cannot fake with a demo video.

How Digit is sold

Robots as a service (RaaS): the robot is rented rather than sold. The customer pays monthly, and the vendor keeps ownership, maintenance and the risk that the machine wears out sooner than the model assumed.

Digit rents at a stated $8,500 per month. Under RaaS a signed order is a promise to pay over several years, and revenue only appears as the robot runs. At that rate a three-year contract is worth about $306,000 per robot, which is why the filing describes $300 million of backlog as roughly a thousand robots under three-year contracts, or something near $100 million a year if every one deploys and stays deployed. The 25,000 robots projected for 2035 would be about $2.55 billion of annual subscription revenue, which makes a $2.5 billion valuation arithmetically defensible if you believe the ladder. The near rung is the one to watch, and it is 800 units in 2027 against nine sites today, with quarterly filings from here.

Agility's counterargument, made in press interviews, is that the backlog is the number that matters and 2025 sales are an artifact of a business that only just started shipping v5. Fair enough, and it is the position the filing's own risk language declines to guarantee. The risk factors say RaaS has not been shown to make money at scale, and that the economics rest on unproven assumptions about Digit's useful life, repair costs, renewals and servicing. GeekWire flags a cash warning too. Those are hardware reliability questions wearing accounting clothes: useful life sets the depreciation schedule, repair cost sets the margin, and nobody has five years of data on either.

Skild's $100 million run rate

Three days later, Bloomberg reported that Skild AI had reached a $100 million annual recurring revenue run rate ten months after its first commercial deployment, with software on hundreds of robots at more than 60 companies, up from eight earlier in the year. Reported work spans NVIDIA and Foxconn server assembly, Sumitomo wire harnesses and Mitsui kitchen pilots.

So the company selling the brain reports roughly fifty-five times the revenue of the company selling the body. The comparison is not apples to apples. Agility's $1.8 million is audited net sales for a completed year, filed under securities law, while Skild's is an annualised run rate given to a journalist by a private company with no filing obligation. Cut Skild's figure hard and the software company is still an order of magnitude ahead, with a fraction of the capital equipment and none of the actuator wear. The field spent five years assuming hardware was the hard part and therefore the valuable part. The first public numbers say the business is in the software.

Something uncomfortable sits next to that number. There is still no paper, no benchmark and no third-party evaluation of S1, the model those 60 customers are paying for. Episode 5 called that a research problem. At $100 million of recurring revenue it becomes a procurement one.

Unitree, three weeks later

Item Figure
IPO price 150.80 yuan
First-day close, 19 August 845 yuan
First-day high 1,100 yuan
Close, 9 September 513.93 yuan (~$72.10)
Drawdown from first-day close ~39%
Drawdown from first-day high ~53%
Still above IPO price +240%
Market value now, versus peak ~$30B, against a ~$66B peak
2025 revenue 1.70B yuan (~$252M)
Humanoid share of 2025 revenue 51.78%
Humanoids shipped in 2025 5,500+

Nothing in the business broke. Every one of those fundamentals was in the prospectus before the stock halved. On revenue, Unitree's 1.70 billion yuan is about 140 times Agility's $1.8 million, which is the gap between the company shipping humanoids in volume and the company just starting to. What moved in three weeks is not that ratio. It is the multiple investors will pay for a humanoid manufacturer, and Unitree is the sector's only real listed comparable, so it is the multiple everything private gets marked against.

Limitations

A revenue multiple means little for a company whose product started shipping mid-year, and the 1,400x figure is trade-press arithmetic. I found no named researcher or investor publicly critiquing the deployment forecast, so the scepticism here is mine and the press's, not sourced analysis. And a filing is not an audit: nobody has verified anyone's shipment counts, or separated a robot sold to a working factory from one sold to a subsidised training centre. What we got is a dollar floor under the word "deployment," which is less than I wanted and more than we had.

Further reading: - The Robot Report on the S-4, the revenue line, the RaaS rate and the projection ladder - GeekWire's read of the filings, with the cash warning in the risk factors - Bloomberg on Skild's $100M run rate, 60+ customers, ten months from first deployment - The Robot Report on Unitree's drawdown, the price history and the regulator's tightened bar

Where a robot's world knowledge comes from

Three groups published three incompatible answers in one week and none cites the others. OpenWAM ran controlled ablations and concluded you should inherit from a capable video generator. AgiBot's GE-Act 2.0, out three days earlier, initialises every trainable component from scratch and scales co-training from 300 to 30,000 hours instead. Rhoda AI measured the question on a factory task with 200-plus hours of robot evaluation and found pretraining helps most exactly where demonstration data is scarce. That last clause reconciles the other two: both may be right about their own data situation and wrong as a general rule.

World action models

Policy: the model that decides what the robot's joints do next. Camera images and an instruction go in, joint commands come out.

World action model (WAM): a policy with an imagination attached. Alongside the action it predicts what the camera is about to see, and uses that prediction to choose. Predicting the next frame forces it to learn how a cup tips and how a cloth folds, which a model that only copies joint angles never has to learn.

The default recipe is to inherit. You take a video generation model, one of the big diffusion transformers trained on millions of internet clips, use those weights as your initialisation, attach an action head, and fine-tune on robot data. The bet is that internet video already taught the network how objects fall and how a hand occludes what it grasps, and that you cannot teach that from robot demonstrations because there are not enough of them. Almost nobody has tested that bet directly.

OpenWAM

OpenWAM is a seven-institution consortium (NUS, Tsinghua, Peking, HKU, Zhejiang, CUHK and SJTU), 24 authors, submitted 7 September. What I like most is that they did not release a model and claim a leaderboard. They built infrastructure to run controlled ablations across eight simulation benchmarks, ran them, and let the model fall out of the findings. That is unglamorous work nobody gets promoted for.

OpenWAM-α itself is a 518M-parameter dual-system design — one half models the world, one half produces actions — with an 80-dimensional unified action space. Pretraining is 518.5 million frames, about 6,400 hours, split 70% robot and 30% egocentric human video.

The three findings, paraphrased:

Finding Claim
What to inherit Upstream world knowledge transfers through a sufficiently capable generative backbone plus a compact, information-rich latent space — a small internal code rather than raw pixels.
How world and action interact Synergy requires dedicated action capacity, explicit world-to-action information flow during training, and synchronised joint denoising at inference.
How pretraining helps Embodied pretraining mainly buys out-of-domain generalisation. Robot trajectories preserve action grounding, egocentric human video broadens transfer, and one-stage co-training integrates both best.
RoboDojo-Real, bimanual Score Success rate
OpenWAM-α 37.6 24.4
Next best (π₀.₅) 22.9 12.8

Roughly double the success rate of π₀.₅, the policy sitting next best on that leaderboard. Forty-six checkpoints went out with the infra code, the evaluation protocols and the data recipes, released CC-BY-4.0.

"Sufficiently capable" is doing real work in finding one, because their reading is that inheriting from a weak generative model buys you very little. They also report that representation encoders with dimension compression are a viable alternative to reconstructive ones, a direct hit on episode 2's JEPA-versus-pixels argument.

GE-Act 2.0

GE-Act 2.0 comes from AgiBot Research, 45 authors, submitted 4 September, one day before this week's window opens. It is the paper OpenWAM is arguing with, and it was published first. Inheriting pretrained video generators, they say, has left the pretraining and scaling of world action models unexplored, so they initialise every trainable generative and action component from scratch, on manipulation data and nothing else.

The architecture has three pieces and the mechanism is clean. A control-oriented autoencoder squeezes the camera image into a compact code built for control rather than for looking good. A single-step visual planner predicts one future image, what the scene should look like after the next chunk of motion. And an inverse dynamics model takes where you are and where you want to be and works out the motion in between. So the robot imagines a frame ahead, then asks what movement would produce that picture. Evaluation is 100 atomic tasks across 20 skill groups, with no per-task fine-tuning anywhere.

Co-training data G1-OP success G2-90D success
300 hours 17.1% 13.4%
30,000 hours 44.1% 31.1%
Skill groups improved 19 of 20 18 of 20

The G2-90D column is the one I would point at. A second embodiment gains 17.7 points from the same scaling, which is what cross-embodiment transfer looks like if it is real. I could not establish what share of the corpus that embodiment accounts for, so read it as suggestive rather than as a measurement of transfer. The stated scope limit is honest: it says this is a low-level manipulation policy, not a high-level planner.

So one paper says inherit, and an older one says build it yourself, from a team with roughly a hundred times more manipulation data.

Rhoda AI's factory study

The third answer came from a company rather than a lab, and it is the one I keep coming back to. Tongzhou Mu and the Rhoda AI team ran the scaling question on an industrial task on 10 September, bearing unpacking, with 200-plus hours of real-robot evaluation behind the numbers.

Direct Video-Action model size At-speed completion
XS 4%
S 65%
M 75%
L 85%

Same data, same robot, same task, 4% to 85% on model size alone. They ran the other axis too: holding the model at M and scaling pretraining compute across four increasing budgets runs 58% → 67% → 74% → 75%. The compute sweep is the more interesting half, because of where the gains concentrate. Rhoda report the largest gains exactly where demonstration data is scarce, and the top of the sweep is nearly flat, 74% to 75%. Pretraining buys you the most when you have the least robot data, and its value flattens as demonstrations accumulate.

DINO FD

Buried in the same post is the most operationally useful result of the week. Rhoda scored their models on held-out web-video prediction, with no robot involved, and that score ranked their checkpoints in the same order real robot performance did, across every model size and compute budget, independently of which produced the improvement.

Picking between checkpoints today means putting each one on a robot and running hundreds of trials, which is why this study needed 200 hours of hardware time. A held-out video metric needs no robot at all. On any perception flywheel the expensive step is the same one, the part where you go and measure the thing out in the world.

Reading the three together

If the value of inherited world knowledge falls as your robot data grows, the answer depends on how many hours you hold. AgiBot has 30,000, so a video generator buys them close to nothing. The consortium had about 6,400, so inheriting carries them a long way. That is my reading and not a result, and nobody has run the crossover experiment.

Limitations

There is no head-to-head anywhere. The two papers use different benchmarks, embodiments and data, so "double π₀.₅" and "17.1% to 44.1%" are not comparable quantities, and neither team has acknowledged the other. Every number here is self-reported and I found no third-party reproduction. Rhoda's caveats are the most complete: one task, one embodiment, one setup, a single training run per condition with no seed variation, a model-size sweep that is not compute-matched, and a correlational pretraining-to-performance relationship. That is what a serious limitations section looks like, published by the outfit with the most incentive to skip it. OpenWAM, meanwhile, took 540 upvotes on Hugging Face and not one technical thread pulling it apart.

Further reading: - OpenWAM and its project page, the ablations, the dual-system architecture and the 46 released checkpoints - GE-Act 2.0, from-scratch initialisation and the 300 to 30,000 hour scaling sweep - Rhoda AI on web-video pretraining, the factory task, both scaling sweeps and DINO FD - SyncWorld, from the same week, arguing that action vectors are not a universal language in pixel space

Marigold V2 gets depth out of an image editor

A team at EPFL, Huawei's Bayer Lab and the University of Bologna did not train a depth model. They took a pretrained image editing diffusion transformer, quantised it to 4-bit, attached a rank-128 LoRA, and fine-tuned it into a single-step dense predictor. It beats the previous best by 16% to 26% in AbsRel on KITTI and ETH3D, and the same recipe delivers state-of-the-art normals, albedo and see-through depth, fine-tuned on a single consumer GPU in under a week. Set that against a week where everybody else was buying world knowledge in thousands of hours of robot time.

Monocular depth estimation: working out how far away everything is from a single ordinary photograph. With a stereo rig or a lidar you get range from geometry, and you pay for it in money, weight, power and calibration. With one camera you have to infer it.

For years that meant training a dedicated depth network on data collected for the purpose. The Marigold line asks whether a model trained to generate and edit images already contains the geometry, and whether you can just decode it.

The recipe

The base model is Qwen-Image-Edit-2509, a large pretrained diffusion transformer, quantised to 4-bit and frozen. Trainable capacity arrives as a rank-128 QLoRA on the DiT. Inference is a single step rather than an iterative sampling loop, which is the difference between something a robot can run and something that makes nice figures. Four modalities come out of the same recipe: depth, see-through depth, camera-space normals and linear-RGB albedo. The code is out, with weights and a demo on Hugging Face.

The loss, and why the ground truth is the problem

The part I found most interesting is why they needed a new loss at all. These models fail on fur, foliage and hair-thin edges, and the authors' diagnosis is that the problem lives in the ground truth rather than the model. The depth maps everyone trains against are not resolved finely enough to capture a strand of hair, so a pixel-wise loss has nothing correct to push toward at those pixels. It is supervised by a label that is wrong at the scale you care about.

Both fixes follow from that. Instead of only comparing pixels, they align the model's internal representations with semantic features extracted from the ground truth, so the supervision carries information the depth map's resolution destroyed. And they adopt a two-stage protocol built around a new Sinkhorn-based loss. Kwang Moo Yi, a computer vision faculty member, posted that his favourite part of the work is the reminder to pay attention to your data, because pixel-wise losses have limits when the ground truth is the constraint.

Results

On zero-shot depth, the headline is a 16% to 26% AbsRel improvement over the previous best on KITTI and ETH3D, with visibly correct fur, foliage and hair-thin edges — the places pixel-wise losses fail. The same recipe transfers to depth completion, see-through depth, surface normals and intrinsic image decomposition, state of the art on each, which matters more than any single benchmark line, because it is one adapter recipe producing four output modalities with no per-task architecture underneath.

The cost of all that is the other headline. Fine-tuning takes under a week on a single consumer GPU, and the code and weights are released.

Limitations

The depth outputs are affine-invariant, so they carry relative structure and no metric scale. A robot that needs to know a shelf edge is 1.4 metres away has to get that scale from a stereo pair, a lidar return, wheel odometry or a known object size. I also did not find inference latency or throughput figures, which is the number that decides whether this runs in a perception stack. A 4-bit quantised DiT is still a large model, and "one step" is not the same as "fast enough at 30 Hz on embedded hardware."

DriveZero, from the same week, needs four frozen vision foundation models in one backbone for comparable coverage, so a single adapter doing geometry, normals and materials weakens the case for specialists. Deployed stacks tend to keep dedicated heads anyway, because you can characterise their failure modes, bound their latency and defend them in a safety case. A frozen backbone with an adapter on top is harder to certify, and that gap is where a lot of very good research goes to sit.

Further reading: - Marigold V2, the recipe, the Sinkhorn loss and the cross-task results, accepted to ACM TOG and SIGGRAPH Asia 2026 - The code and training configs, including the single-GPU setup - DriveZero, the four-frozen-backbone alternative, plus a driving policy that beats log-replay by training under augmented intents

Four tactile datasets in one week

The Marigold trick does not work twice, because there is no web-scale archive of what things feel like. Four groups attacked that from the data side, from 100 curated hours at Berkeley to 30,000 synchronised hours out of Shanghai, and one skipped the sensor and inferred touch from vision. The builder of the largest corpus says the field needs 100,000 hours, so the biggest tactile dataset in existence is under a third of what its own author thinks is required. A sensor vendor published the upstream reason in the same week, and the reason, in the end, is that the skin wears out.

Why force is missing from VLAs

If you have ever seated a ribbon cable or picked a ripe strawberry without bruising it, you know the last centimetre of that job is not done with your eyes. Once the parts are touching, the contact patch is hidden by the parts themselves, and what you go on is how the thing pushes back. The motivation Trevor Darrell's group gives in IEEE Spectrum's roundup is that force, slip and precise grasping are not quantities an ordinary vision sensor reports well. Almost every vision-language-action model this show covers ignores force entirely, which is why fine-grained hand control on deformables and small objects is still broken.

Episode 3's SoftVTBench found recorded successes where the policy crushed what it was holding, because a success label is one bit and a tolerance violation is a continuous quantity nobody recorded. Episode 5's Facet-0 went from 15% to 82% on sub-millimetre assembly by supervising on expected force.

The four efforts

Effort Group Scale Reported result
T-Rex Trevor Darrell's group, UC Berkeley 100 hours, high quality, single hardware instance 65% average success across 12 complex tasks, close to double the best VLA models
Tactile aggregation Chengbo Yuan's group, Tsinghua 3,000+ hours from public datasets, 21 sensor types, multiple embodiments Generalisation to unseen tactile hardware
Vision-touch corpus NeoteAI and Fudan University, Shanghai 30,000+ hours of synchronised visual and tactile data Significant gains from tactile-prediction models
Vision to touch USC 2,700+ demonstrations A model inferring the tactile signal from vision alone; no success figures reported in the roundup

USC is the one I want numbers for, and the roundup does not give any. Inferring the touch signal from vision alone would mean either that part of what the sensor provides was already in the camera and models were not extracting it, or that 2,700 demonstrations in one setup is narrow enough for the mapping to memorise. Which of those it is decides whether this is a result or a curiosity, and nothing in the coverage settles it.

Twenty-one sensor types that disagree

The quiet story is fragmentation. Vision has a de facto standard, a rectangular array of RGB samples, which is why a model trained on one camera transfers to another and why an image dataset from 2012 is still usable. Touch has nothing of the kind. Capacitive, piezoresistive, optical and acoustic sensors report different quantities at different rates through different mechanical stacks, so half the work in building a tactile dataset is reconciling hardware that disagrees. Chengbo Yuan's group trains across the diverse setups anyway and reports generalisation to tactile hardware the model never saw, which is a workaround for the missing standard rather than a standard.

UltraSense's durability argument

Mo Maghsoudnia and Hao-Yen Tang of UltraSense Systems argue that surface-coupled electronic skin, whether capacitive, piezoresistive, piezoelectric, triboelectric or impedance-based, suffers wear, hysteresis, creep, delamination, baseline drift and a recalibration burden, because the sensing layer sits in the harshest mechanical environment on the robot. Their proposal puts ultrasound transducers beneath the contact surface and reads contact from acoustic echoes, since pressing an object into a compliant surface changes the acoustic path length, and the echo timing with it.

Claim Figure
Spatial resolution through elastomer 500 µm
Localised compressive force precision ~1.25 mN
Shear inference noise floor ~5 mN
Units shipped into automotive 4M+

This is a vendor arguing for its own product and I would discount it accordingly. The durability problem it names is real, though, and rarely stated with numbers. The part that wears in their design is a cheap layer of rubber rather than the sensor, which is the right instinct for anything that has to survive a production floor for years. The two stories fit together in an unhappy way: one set of groups says the tactile data does not exist, and the other says the sensors do not survive long enough to produce it.

Limitations

Every figure here comes from a trade-press roundup rather than the papers, which I have not read. The success rates are not comparable to each other: different tasks, different hardware, different baselines, all self-reported. "Close to double the best VLA models" is the roundup's framing and the baselines are unnamed. The 100,000-hour figure is Shunlin Lu's estimate in an interview, and his group has a commercial interest in tactile data being valuable. UltraSense's numbers are unaudited vendor specs.

Further reading: - IEEE Spectrum on the four tactile datasets, Edd Gent's roundup, with the researcher quotes and the 100,000-hour estimate - UltraSense on ultrasound under the skin, the durability argument and the sensor specifications - Facet-0, from episode 5, the constructive answer on the policy side: predict the wrench you expect to feel and score on whether contact matched

Also this week

Swarmer is buying Ratel Robotics for up to $224 million, folding the most field-proven multi-agent autonomy on earth, 100,000-plus combat missions since 2024, into one company no robotics paper cites. MassRobotics polled 14 members on the July FCC sourcing rules and got a 43/43 split tracking manufacturing footprint almost exactly, which is what a policy looks like when it redistributes rather than restricts. And Teradyne is suing JAKA over cobot patents with March 2006 priority dates, one covering a joint brake built from a solenoid and a ratchet. The FCC rule and the patent suit are different instruments aimed at the same contest.

This week in one table

Item What it is Link
Agility Robotics S-4 $1.8M 2025 net sales, $140M operating loss, ~$100M cash burn, $2.5B merger valuation (~1,400x revenue), 9 sites, 65,000+ operating hours, $300M+ booked orders, 800 units projected for 2027 The Robot Report
Agility filing detail RaaS at $8,500/month, roughly 1,000 robots behind the $300M backlog, and the cash warning in the risk factors GeekWire
Skild AI $100M annual recurring revenue run rate, 60+ customers, hundreds of robots, ten months after first deployment, still no paper Bloomberg
Unitree drawdown 513.93 yuan on 9 September, ~53% below the 1,100-yuan debut high, ~$30B against a ~$66B peak, on 1.70B yuan of 2025 revenue and 5,500+ humanoids shipped The Robot Report
Maven Robotics Out of stealth with a $100M Series A on eight wheeled dual-arm palletisers at 99%+ uptime, 30 kg via vacuum, 16-hour shifts TechCrunch
Chinese humanoid listing scrutiny CSRC reported to have informally raised the bar on recurring revenue, profitability path and genuine innovation technologies.org
OpenWAM Seven-university consortium, controlled WAM ablations, 518M-parameter dual-system model with an 80-D unified action space, 518.5M frames / ~6,400 hours at 70% robot and 30% egocentric human, 37.6 / 24.4 versus π₀.₅'s 22.9 / 12.8 on RoboDojo-Real, 46 checkpoints released arXiv
OpenWAM project page Architecture detail, the eight simulation benchmarks and the three study findings openwam-official.github.io
GE-Act 2.0 AgiBot, every trainable component from scratch, 300 to 30,000 hours takes G1-OP from 17.1% to 44.1% and improves 19 of 20 skill groups, submitted 4 September arXiv
Rhoda AI scaling study Bearing unpacking, 200+ hours of real robot evaluation, 4% to 85% by model size, 58% to 75% by pretraining compute, DINO FD predicts robot performance, full limitations published Rhoda AI
Marigold V2 Qwen-Image-Edit-2509 at 4-bit with a rank-128 LoRA, single-step depth, 16% to 26% AbsRel gain on KITTI and ETH3D, plus normals, albedo and see-through depth, under a week on a single consumer GPU arXiv
Marigold V2 weights and code Code and weights released, four output modalities, per-stage training configs Hugging Face
Marigold V2 demo Run it on your own image in the browser HF Spaces
Four tactile datasets T-Rex at 100 hours and 65% across 12 tasks, Tsinghua's 3,000+ hours across 21 sensor types, NeoteAI/Fudan's 30,000+ hours, USC inferring touch from vision alone on 2,700+ demonstrations, against a stated need for 100,000 hours IEEE Spectrum
UltraSense ultrasound tactile Transducers under the contact surface, 500 µm resolution, ~1.25 mN compressive precision, ~5 mN shear noise floor, 4M+ automotive units shipped The Robot Report
SyncWorld Visual calibration episodes teach a world model what your action numbers mean in your room, enabling zero-shot simulation in unseen setups arXiv
Show-Harness A semantic action interface plus embodiment-specific interpreters lets frontier VLMs drive robots zero-shot, no retraining; comparative claims unverified arXiv
MaP-WAM Episodic memory compiled into segment-level plans, 83.3% on RMBench and 78.0% on real robot tasks at roughly constant executor latency as history grows arXiv
Programmable World Model Language compiles into executable state programs so off-screen entities persist; 94% count accuracy and 98% state accuracy on CombatStateBench arXiv
Recursive Code World Models Single images reconstructed as executable recursive scene programs, with deeper recursion improving fine detail arXiv
DriveZero Four frozen vision foundation models consolidated into one driving backbone, plus a PPO teacher queried under augmented intents; 93.57 mean on nuPlan, claims it exceeds log-replay arXiv
SpatialBlock 15,000 synthetic block-stacking problems targeting projection, viewpoint transformation and structural combination, claiming transfer to real spatial tasks arXiv
SwingBot Whole-body humanoid brachiation with passive wrist hooks, biomimetic keyframes and recurrent privileged-state estimation; no success rates on the abstract page arXiv
ETH Zürich brachiation (August, context) Jump-up, brachiation and jump-down off raw head-mounted lidar, 14 of 15 hardware trials, 2 cm lateral clearance, up to 0.5 m/s arXiv
Swarmer and Ratel Robotics Up to $224M for a Ukrainian UGV maker, folding NATO-certified hardware into swarm software with 100,000+ combat missions behind it The Robot Report
MassRobotics FCC survey 14 companies, 43% beneficial and 43% detrimental, split tracking manufacturing footprint, 29% had not started onshoring The Robot Report
Teradyne versus JAKA Three Universal Robots patents with March 2006 priority dates asserted in the Unified Patent Court in Copenhagen The Robot Report
Chinese accelerator prices Huawei, Cambricon, MetaX and Iluvatar CoreX raise prices 20% to 50%, blaming HBM costs technologies.org
Vention Physical AI Lab Montreal lab for unstructured manufacturing manipulation, harvesting data from 28,000+ already-deployed machines, GRIIP pipeline slated for open source The Robot Report
"AI can't outrun a humanoid's hardware" Supplier op-ed on the 12V to 48V migration: 4x less current, 16x lower wiring loss. One real number, useful framing The Robot Report
ARM Institute awards $90M across 10 projects to modernise military manufacturing The Robot Report
Enovis and eCential $180M acquisition, surgical robotics consolidation The Robot Report
Figure "Index" (25 August, catch-up) Paid crowdsourced human video at population scale: 16M+ uploads, 44,000+ weekly contributors, 30 minutes ingested per second, $15M paid out, $1B+ committed to data and compute Figure
Figure's compute commitment (3 September) Up to 100,000 GPUs from Nscale on NVIDIA's Vera Rubin platform, reported as ~$3.5B against ~$1.9B raised Forbes
AMD Ryzen AI Embedded X100 (23 July, catch-up) Strix Halo-class APU for robots, up to 128 GB unified memory. AMD's Jetson Thor comparison ran on a mini PC "configured to reflect" X100 specs, not on real X100 hardware Tom's Hardware

What we're watching

Where is the crossover in hours? If inherited pretraining is worth most when robot data is scarce, there is a break-even point where building from scratch starts to win. Nobody has run that experiment, and until somebody does, OpenWAM and GE-Act 2.0 are both right about their own situation and neither is a general rule.

Does DINO FD survive a second task? Rhoda's held-out video score ranked checkpoints in the same order as real robot performance on one task, one embodiment, one architecture family. The follow-up worth doing is adversarial: go find the checkpoint pair where the metric and the robot disagree.

What does Agility's first 10-Q say? The S-4 projects roughly 800 Digit v5 units in 2027 against $1.8 million of 2025 revenue and nine customer sites. The unit count and the revenue recognition schedule are the lines to read, and this is the first time a humanoid maker's claims come with a filing deadline attached.

At what revenue does a missing evaluation become a procurement problem? Skild has $100 million of recurring revenue and no paper, no benchmark and no third-party evaluation of S1. Sixty companies have made a purchasing decision the field cannot check.

Does anyone stop training dedicated perception backbones? Marigold V2 says a generative image model already holds the geometry, the materials and the lighting. My guess is the field keeps paying for specialists anyway, and I would like to be wrong about that.

Who pays for 100,000 hours of touch, and in which format? The largest tactile dataset is 30,000 hours and its own builder puts the requirement at 100,000, collected in the wild. A standard would be worth more than another dataset right now.

Anthropic's prospectus still does not exist. Reporting said after Labor Day, then an investor day in mid-September, then a listing in late September or early October at roughly $1.5 to $2 trillion. As of 12 September there is no public S-1, only a confirmed confidential submission from June. Nothing physical-AI-specific can be said about a frontier lab's exposure under securities law until the document exists. Open since episode 3, and it either resolves in the next few weeks or comes off the list.