Anthropic Aims for a Record IPO, Unitree Up 460%
Anthropic has told its bankers it expects the largest first-time share sale in history, and it is reportedly still trying to buy a world-model company for six billion dollars while it does. A private lab can fund a five-year bet on physics, and a public one answers to a quarterly clock, which makes this the most consequential open question in Physical AI right now. Also: Unitree closed its first day of trading up 460 percent on a shipment number three organisations can't agree on within fifty percent, three labs independently decide the headroom is in the loop around the policy rather than in the weights, and a new benchmark finds that up to a quarter of recorded successes on soft objects crushed the object.
- Industry
- Humanoids
- Benchmarks
- VLAs
The written deep dive
25 min read · everything from the episode, with the numbers and citations
Bloomberg reported on Thursday that Anthropic expects the largest first-time share sale in history, bigger than SpaceX's $75 billion, while it is reportedly still trying to buy a world-model startup for $6 billion. The day before, the most visible humanoid maker on earth closed its Shanghai debut up 460 percent, near two hundred times what it earned last year, on a shipment number three organisations can't agree on within fifty percent. And in the research, the benchmark much of this field's imitation-learning thesis was built on got solved 100 percent of the time by a coding agent with no human demonstrations, whose authors then wrote into the paper that they cannot rule out it reverse-engineered the simulator. Every figure there is real, and none of them means what it looks like.
Anthropic's IPO and the quarterly clock
Anthropic expects to match or beat the biggest IPO ever recorded, and it is reportedly still shopping for a world-model company while it prepares to file. Whether a public frontier lab keeps paying for physics is the most consequential open question in this field, and it gets settled by people who have never worked in robotics.
Bloomberg reports that Anthropic has told its bankers it expects to match or top the largest first-time share sale ever, which is SpaceX at $75 billion.
| Item | Detail |
|---|---|
| Confidential SEC filing | 1 June 2026 |
| Banks | Morgan Stanley, Goldman Sachs, JPMorgan |
| Public filing | possible as soon as the end of August |
| Record to beat | SpaceX, $75B |
| Series H | closed at a $965B valuation |
| Run rate | $65B at the end of July |
Confidential filing: a company can hand the SEC a draft registration statement without publishing it, which starts the review privately. The public version comes out before the company goes on the road to sell the shares. A 1 June confidential filing and an end-of-August public one is a normal gap, not a sign of trouble.
The part that belongs on this show is what the same company is reportedly doing with the other hand. It is still trying to buy Decart, a world-model company, for $6 billion. Episode two aired that as a rumour, and it has neither closed nor been denied since.
World model: a system that learns how the physical world behaves and can predict what happens next given what you do. It is the thing that lets a robot try an action in its head before it moves its arm. Decart's own stated applications include autonomous driving and simulating the physical world, which is why a model lab buying one is a physical-AI story rather than a chatbot story.
A private lab can put $6 billion into a bet that pays off in five years, and nobody outside the building gets a vote. A public company answers to a quarterly clock, and I have watched what that clock does to physical-AI programmes. The work that makes a robot better in the real world does not photograph well on a quarterly slide. You build data infrastructure. You auto-label at scale. You pretrain on things nobody asked you to ship. You spend a year shrinking a model to the latency the vehicle needs, and the payoff shows up two years later as a curve that bends. World models are the extreme case of that shape, and for a long stretch they show you nothing you can put in an earnings deck.
So the question is not whether Anthropic can afford Decart. At a $65 billion run rate, $6 billion is a rounding error against what an IPO of this size raises. The question is whether the physics line survives an investor base that prices the company off model revenue. It could go the other way, and a lab with that balance sheet could fund more world-model work than the academic side of this field manages in a decade.
How frontier labs have touched robotics so far
Frontier labs have mostly touched physical AI at arm's length, shipping robot-flavoured versions of their models and letting robotics companies integrate them. Buying a world-model company outright says the lab wants the physics stack rather than a customer segment. The mirror image landed the same week in China, where DeepSeek took a locked strategic position in Unitree's IPO and signed a joint model-development agreement (more on that below). An open-weights lab buying three years of illiquidity in a robot maker, and an American lab buying a world-model startup on its way to a record listing, are the same bet in different clothing.
Limitations
Two numbers are circulating that you should ignore. Crypto pre-IPO perpetuals imply something near a $1.6 trillion valuation, and one forecasting site puts the median first-day close at $1.82 trillion. Those are thin prediction instruments and forecasts, quoted as though they were prices.
The rest needs care too. This is reporting on a private expectation, not a filed number. No S-1 is public as I write, the size of an IPO moves right up until pricing, and "expects to match or beat" is a thing companies tell bankers, who tell reporters. The Decart deal is unconfirmed in both directions. And the quarterly-clock argument is an argument, not evidence. I believe it because I have watched budget cycles eat long-horizon research, but nobody knows how a public Anthropic allocates capital.
Further reading: - Bloomberg on Anthropic's IPO expectations, the banks, the confidential filing and the record to beat - eWeek on DeepSeek's stake in Unitree, the model-lab-buys-into-hardware version of the same bet - SCMP on the DeepSeek-backed listing as a test of investor appetite, useful context for what public money will pay for in this sector
Unitree's listing and the numbers holding it up
Unitree closed its first day on Shanghai's STAR Market up 460 percent, at roughly 200x trailing earnings, while its own prospectus concedes that robot hands are still not precise or durable enough for large-scale adoption. Both are true at once, and reconciling them is most of the story.
Unitree is the company whose quadrupeds you have watched doing backflips online, and whose humanoids danced at the Spring Festival gala. Its pitch has always been price, robots that cost about what a car costs in a field where research humanoids ran into the hundreds of thousands.
What the debut printed
| Item | Figure |
|---|---|
| Offer price | 150.80 yuan per share |
| Shares sold | 40.45 million, about 10% of enlarged capital |
| Raised | ~6.1 billion yuan (~$850-900M) at a ~$9B valuation |
| Opening print | 1,100 yuan, +629%, ~445 billion yuan (~$66B) |
| Close | 845 yuan, +460%, ~342 billion yuan (~$48-53B) |
| Day-one turnover | 23.2 billion yuan |
| 2025 revenue | 1.7 billion yuan (~$252M) |
| 2025 profit | 600 million yuan (~$89M) |
STAR Market: Shanghai's Nasdaq-style board for technology companies. Listings float a small slice of the company and the first-day move is uncapped, so a large day-one pop is partly a mechanic of thin float rather than a verdict on the business.
The coverage disagrees with itself, which is the same disease as everything else here. The 629 percent and $66 billion pair belongs to the opening print, the 460 percent and roughly $48-53 billion pair to the close, and Fortune's headline mixes the two.
Roughly three-quarters of Unitree's humanoid revenue came from research and education, so the buyer is a university lab rather than a factory, and industrial deployment sat under 10 percent through the first three quarters of 2025. That split matters more than the multiple does. A research customer buys one unit and does not come back for forty more.
The FCC Covered List addition
This part is catch-up rather than news. On 28 July, twenty-five days before the listing, the FCC, America's Federal Communications Commission, added foreign-produced advanced robotic devices to its Covered List.
Equipment authorisation: the FCC certification any device with a radio in it needs before it can be legally marketed in the United States. Adding a category to the Covered List blocks new device models from receiving it, and works forward only, so robots already bought or already authorised are untouched.
The scope covers mobile humanoids, quadrupeds and autonomous mobile robots over 4.4 lb carrying sensing, networking and control software. The technical hook behind the national-security rationale is UniPwn, a Bluetooth Low Energy provisioning flaw in Unitree's Go2 and B2 quadrupeds and G1 and H1 humanoids that yields root through hardcoded keys and is wormable, so an infected robot compromises other units in BLE range unattended. Andreas Makris and Kevin Finnisterre disclosed it on 20 September 2025, which makes it background to the ruling rather than news.
Unitree cleared its current lineup weeks earlier, the R1 on 22 June and the H2 and A2 quadruped on 30 June 2026. Today's products are fine and every future model is a problem. About 18 percent of 2025 revenue came from the United States, and the prospectus names the restriction as a risk.
DeepSeek's stake
DeepSeek put about 140.8 million yuan into the strategic allotment for 2.31 percent, locked for 36 months, and signed an agreement to jointly develop AI models for humanoid robots, with reported terms giving each firm priority when the other is buying.
Twenty million dollars is small change against a $48 billion market cap, so the lock-up is the signal. Three years of illiquidity says the hard part from here is the model rather than the mechanism.
Three shipment counts, and what a shipment is
The number every valuation in this sector rests on is units shipped. Here is what three organisations published for H1 2026.
| Source | H1 2026 humanoid shipments | Leader | Second |
|---|---|---|---|
| Counterpoint Research | 22,000+ globally | AgiBot, ~9,700 (43%+) | Unitree, 7,000+ (31%) |
| Smart Analytics Global | 19,100 globally | AgiBot, ~8,400 (44%) | Unitree, ~5,900 (31%) |
| Humanoid Robot Scene Application Alliance | 30,000+ in China alone | not broken out | not broken out |
The low and the high are more than 50 percent apart, and Unitree is not first on either dataset that breaks out vendors. AgiBot ships more. Both counts do agree that Chinese makers hold about 97 percent of global shipments, and that industrial and commercial use climbed past 70 percent from roughly 50 percent a year earlier.
Then there is the mechanism the Forbes piece names, which I had not seen written down before. In China there are state-backed training centres that buy humanoids in order to teleoperate them, collect demonstration data, and sell that data back to the manufacturers. Nobody is doing anything improper. But a robot shipped into that channel and a robot sold to a factory are not the same event, and the sector is priced off one series that mixes them.
What Wang Xingxing said the day after
Speaking around the World Robot Conference in Beijing, Wang moved his own timeline out. The industry's breakthrough, he said, could take two to three years at the fastest and five or even ten years at the slowest. He also defined the threshold concretely, which I appreciated. Success is a robot you put in an unfamiliar home that handles about 80 percent of tasks from voice or text alone. I have no primary transcript, so treat the phrasing as reported rather than verbatim.
The companion piece the next day collects the same mood from the conference floor, including a Keenon executive's reliability wall, where 50 percent of a task is easy, 80 percent is manageable, and 99.9 percent is what tests your engineering. That 80 percent shows up twice, once as a finish line and once as the point where the hard part starts.
The Superman video
Two days before trading opened, Unitree posted a thirty-second video of an experimental humanoid it calls Superman, claiming 12.66 m/s, about 28.3 mph, against Bolt's peak instantaneous velocity of roughly 12.2 to 12.42 m/s, plus a 2 metre standing jump. Unitree says it is a research testbed rather than a product, built in just over three months, and it deserves credit for saying so. The sprint configuration has no hands and no head, because both were removed.
What is missing is everything that would make the claim checkable. No test protocol. No payload. No description of the running surface. No repeatability, no independent verification. A performance number with no protocol behind it is marketing.
Limitations
The boom is not fake and I do not want the skepticism read that way. Real humanoids are shipping in quantity, and Unitree builds hardware that works at a price nobody else matches. What I am questioning is the series used to price the category, and neither Counterpoint nor SAG publishes methodology that would reconcile a 1,300-unit gap on the leader. HSBC analysts told Fortune the shipment surge "could be illusory" without a significant improvement in the AI model capability of robot makers, and Persona AI's Nicolaus Radford said Unitree has proven it can sell robots but not "a use case that warrants deployment at scale."
Further reading: - SCMP on the debut, the open, the close and the turnover - Forbes on the three shipment counts, including the training-centre channel - Forbes on the FCC Covered List addition, scope and what it does not cover - TechNode on the Superman claim, the 12.66 m/s figure and the missing protocol
Test-time compute arrives in robot policies
Three groups published in one week on the same instinct, that the headroom now sits in the loop around the policy rather than in the weights. τ₀-VLA searches over candidate subtasks with a world model scoring them, Zetta freezes the base policy and writes code around it, EXIMO uses a VLM to explore. None cites either of the others.
Vision-language-action model (VLA): a policy that takes camera images plus a plain-language instruction and emits motor commands. "The policy" is the learned thing that decides what the joints do next, and most of the last three years of progress came from making it bigger.
Test-time compute: spending more computation at the moment of acting rather than at training time, usually by generating several candidates and picking among them. It is the idea behind reasoning models in language, and this is the first week I have seen it in robot policies with real-hardware numbers attached.
How τ₀-VLA is built
τ₀-VLA, from the Shanghai Innovation Institute, runs two policies at different time scales. A high-level policy on a Qwen3.5-9B backbone reads the instruction, the current observation and a memory of what has been done, then emits the next subtask in words, something like "pick up the kettle." A low-level policy on a Qwen3.5-2B backbone, with a mixture-of-transformers action expert trained by conditional flow matching, turns that subtask into motor commands in a unified 40-dimensional state and action space.
The new part happens at the moment of choosing. Instead of emitting one next subtask, the model proposes candidates with a branching factor of 3, has a world model (initialised from Step1X-Edit) score how each one would go, keeps the best 2, and expands again recursively. More compute buys a deeper search before it commits. Training used 40,115 hours of real-world robot data, and evaluation ran on three embodiments, an AGIBOT G1 wheeled humanoid, an ARX AC One and a Franka Research 3.
| Long-horizon task | Result |
|---|---|
| Clean Room | 5/10 |
| Prepare Ingredients | 4/10 |
| Tomato and Egg Stir Fry | 4/10 |
| Make Milk Tea | 5/10, rising to 7/10 with test-time compute |
| Average across the four | 45.0% |
| Direct execution baseline | 27.5% |
| Cross-embodiment Collect Laundry / Tidy Makeup Table | 10/10 each |
| Out-of-distribution Book Organization, next-subtask accuracy | 74.0% with search vs 50.0% without |
That last row is the one to hold on to. On a task the model had never trained on, choosing the right next subtask went from a coin flip to about three in four, purely by spending more compute at run time. I would rather have the honest middle of that table than a saturated benchmark. Four in ten on a stir-fry is not a product.
One caution. The project page and GitHub README both carry a model-release entry dated 27 July against an arXiv v1 of 17 August, so the paper landed this week and the checkpoint did not. Every primary source says 40,115 hours, and no parameter count is confirmable, whatever the search summaries say.
Zetta and EXIMO
Zetta ζ takes the most aggressive version of the same bet. It keeps the base policy frozen and evolves everything around it, with code-based runtime critics and recovery skills written online across three timescale-separated loops: action-frequency governance, rollout-level critic and recovery proposals, and validation-gated skill updates. It reports 90.8 percent on LIBERO-Pro, 93.6 percent on RoboCasa, and an 11.1x inference speedup. That first number needs care, because a third-party summary site has been reported quoting figures between 32 and 71.13 percent for the same benchmark. The arXiv abstract says 90.8 percent, and it does not name the frozen base policy, so do not attribute it to π₀.₅ or anything else.
EXIMO, from Google DeepMind, comes at it from the exploration side in three stages. A VLM planner decomposes long-horizon tasks into subtasks and collects structured data with the VLA, the VLA is finetuned on that data, and residual off-policy RL refines it. The abstract claims gains in sample efficiency and final performance, with per-stage ablations and no headline numbers.
Where this disagrees with episode one
Episode one closed its world-action-model primer on Fast-WAM's finding that imagining the future is a training-time habit and not a run-time requirement. τ₀-VLA is built on the opposite premise and has real-hardware numbers behind it. I would rather flag that than pretend the field stood still. I think both claims survive, because they search over different things. Fast-WAM was about per-step video rollout, imagining every frame as you go. τ₀-VLA imagines candidate subtasks at the few moments where a decision gets made.
Limitations
Zetta's benchmarks are LIBERO-Pro and RoboCasa, both simulation. τ₀-VLA reports the accuracy it gained and nobody has published what the search costs in wall-clock time. On a vehicle I would need that number first, because a planner that thinks longer while the world keeps moving is a different system from one that thinks better. And three papers converging in a week means the idea is in the air, which is weaker evidence than it sounds.
Further reading: - τ₀-VLA, hierarchical policies with world-model-guided beam search over subtasks - Zetta ζ, frozen base policy with runtime critics and recovery skills evolved online - EXIMO, VLM-guided exploration, imitation, and residual off-policy RL
Push-T, solved with no demonstrations at all
A coding agent handed the Push-T simulator and zero demonstrations wrote a controller that solves the task 100 percent of the time, in 46 percent fewer steps than a 200-demonstration diffusion policy at 62.5 percent. The authors then wrote their own refutation into the paper, and it is the best part of the work.
Push-T: a manipulation benchmark where a robot pushes a T-shaped block across a table until it matches a target pose, using a single round contact point. It sounds trivial and it is not, because you can only push. Nudge the block off-centre and it rotates, so you have to plan a sequence of contacts. If you have ever shoved a heavy box into a corner with one hand, you know how that goes.
The choice of task is the whole argument. Push-T is the benchmark Diffusion Policy was introduced on, and Diffusion Policy is what convinced a large part of this field that collecting human demonstrations and imitating them was the road forward. Shuangyu Xie, Kaiyuan Chen and Ken Goldberg at UC Berkeley handed an LLM coding agent that environment with no demonstrations, and it explored the simulator, wrote a controller, tested it, and rewrote it.
| Controller | Success | Mean steps | Mean pushes |
|---|---|---|---|
| Agent-written, state input | 100.0% | 120.1 | 3.72 |
| Agent-written, vision feedback | 97.2% | not reported | not reported |
| LeRobot diffusion policy, 200 human demos | 62.5% | 223.7 | 6.38 |
They extended it to the full alphabet, Push-A through Push-Z, at 99.4 percent with state input and 99.2 percent with vision, and generated 3D cross-embodiment simulation code for a Franka and a UR5. Cost was roughly $1,500 to $2,000 of agent time and 221 hours. The coding agent was Claude running Fable 5, which the paper states, and this show is scripted by a Claude model, so I will name that once and move on. The finding does not depend on it.
The authors' own caveat
I would rather quote this than paraphrase it. The authors write that "the simulator is only an approximation of the physical system. The agent may exploit simulator-specific APIs, state variables, rendering conventions, contact behavior, or numerical regularities that are unavailable or inaccurate on a physical robot." They claim no robustness to perception error, calibration drift, latency, compliance, friction variation or unmodeled contacts. They ran no physical robot experiments at all, and they recommend treating simulation "as an uncertain hypothesis rather than a fully specified environment."
So the result forks and I cannot tell you which fork it is on. Either the outer loop was always where the headroom was, and $2,000 of agent time beat 200 human demonstrations, or three benchmarks just got quietly reverse-engineered. Both readings fit the evidence.
Further reading: - Revisiting Push-T with agentic robotics, the controller synthesis, the alphabet extension and the caveat - Goldberg's ICRA 2026 plenary, the two-cultures argument in his own words
SoftVTBench and the successes that crushed the object
Across Diffusion Policy, π₀.₅ and FastWAM, all twelve in-distribution configurations contained rollouts scored as successes where the object had been deformed past tolerance, 0.7 to 24 percent of each configuration's recorded successes.
Say you are scoring a robot that has to put a strawberry on a plate. You check whether the strawberry ended up there. That is the success rate, and an enormous share of robot learning rests on that one choice.
Finite-element ground truth: a physics simulation of how the object itself deforms under contact, resolved through the material rather than only at the surface. Having it alongside a robot demonstration is unusual, and it means you know at every instant how much the thing is being squashed.
SoftVTBench is 4,000 expert demonstrations across 50+ assets, including volumetric deformable objects each paired with a visually matched rigid twin, a clean control because a camera cannot tell the pair apart. Everything is captured at 20 Hz with multi-view RGB, dual-finger tactile RGB, marker motion tracking, proprioception, language annotations, gripper actions, and the finite-element state alongside it. The authors define a Deformation-aware Success Rate from that, then score three of the field's most used policies against both metrics. Every in-distribution configuration, twelve of twelve, contained runs the ordinary success rate counted as wins while the finite-element ground truth showed the object crushed past tolerance. The task scored as a success, and the strawberry was destroyed.
So some part of what this field measures on soft objects is policies exploiting the metric, and nobody could see it, because the information needed to catch it was never recorded. That is the same disease episode two diagnosed in demonstration data, one layer up in the scoreboard itself.
The authors are careful in the other direction too. Under distribution shift the visuo-tactile variants win most comparisons, but they note that "making touch available does not by itself ensure effective multimodal fusion." Anyone who has fused two sensors that disagree knows that sentence.
Limitations
One benchmark, one group, no independent reproduction, and the tolerance threshold separating a success from a crush is itself a choice the authors made. The 0.7 percent floor is small enough to call noise. The 24 percent ceiling is not noise on any reading. What I want next is the same treatment applied to a benchmark nobody suspects.
Further reading: - SoftVTBench, the dataset, the rigid twins and the deformation-aware success rate
Hydra-0 and pixel motion as an action interface
Episode one left open whether pixel-space action rendering survives past lab-sized benchmarks. It does. And the more important number is that Hydra-0's imagined rollouts rank policies at r=0.96 against real trials, which is the property episode two's audits kept finding missing.
An action, when you train a robot policy, is normally a vector of numbers, joint angles and gripper commands. Those numbers mean different things on a bimanual robot, a single arm, and a human hand holding a tool. Every embodiment is a new dialect, which is much of why cross-embodiment training is hard.
Action flow: instead of feeding the model joint numbers, you draw the motion into the image. The action becomes pixel displacement, how things move in the camera view, which is a representation every embodiment shares because every embodiment is visible.
Hydra-0, from NVIDIA with academic co-authors, scales that idea into a generalist model spanning human hands, handheld grippers, single-arm and bimanual robots, across multiple video backbones. It also refuses the dichotomy episode two set up between latent and generative world models, because it is hybrid by construction. A physics engine executes the robot's own motion and converts it into action flow, since the kinematics are known. The learned world model predicts only how the scene and the objects respond, the part nobody has closed-form equations for.
| Metric | Result |
|---|---|
| Robot-motion error vs action-conditioned baseline | 90.4% lower |
| Object-motion error vs action-conditioned baseline | 60.2% lower |
| Correlation between replayed and reference success rates on RoboLab | r = 0.96 |
That third row is what matters. The model's imagined rollouts rank policies in nearly the same order real trials do. Episode two spent its time on audits finding that learned simulators do not reliably respond to the actions you feed them, which makes them useless as judges. Somebody now has a world model working as an evaluator, and evaluation is the bottleneck almost nobody budgets for. Anyone running a fleet knows the real cost of a policy change is the weeks of physical trials before you trust the comparison. There is also an emergent inverse mode, where object flow lifted from a human demonstration is enough for the model to predict compatible robot motion, with a learned action head turning it into executable actions.
The numbers above are the abstract's. Corpus figures circulating in secondary coverage do not trace to a primary source I could open, so they are not here.
The same instinct showed up independently in navigation. Embodied-Navigator, from ZJU-OmniAI, has a VLM select pixels that project into 3D coordinates driving SLAM control, and reports 66.2 percent success on R2R-CE from only 90k training trajectories. Two unrelated groups, one interface choice.
Limitations
RoboLab is one benchmark, and r=0.96 there does not establish that imagined evaluation transfers to a different task distribution. The hybrid design assumes a physics engine and a calibrated model of the robot, fine in a lab and less fine on a platform whose dynamics drift with load, wear and temperature. And I would want that correlation rechecked when the ranked policies sit close together, because separating a good policy from a terrible one is the easy version of the problem.
Further reading: - Hydra-0, action flow, the hybrid physics-plus-world-model split and the r=0.96 result - Embodied-Navigator, pixel selection as navigation actions, 66.2% on R2R-CE - Decision-Metric Alignment in latent world models, diagnostics for whether a latent model's training metric matches the decision it is used for
Also this week
Nevada authorised 8,000 robotaxis in one county. The Nevada Transportation Authority voted unanimously on 20 August to grant permits covering Clark County, capped over twelve months at 5,000 for Tesla, 1,000 for Waymo, 1,000 for Uber and Zoox's existing 100. Waymo ran about 3,500 vehicles across 11 US metros as of July 2026, so one county authorised more than twice that. The number deflated the same day, with Tesla's Cybercab chief engineer telling TechCrunch that "the 5,000 has always been a ceiling for us," and that Tesla would be "extremely happy" at 2,500.
A functional-safety company is going public. FORT Robotics is combining with Newbury Street II Acquisition Corp at a $556.6M pro-forma enterprise value, listing on Nasdaq as FROB. It sells e-stops, safe wireless control and a safety-rated stack for machine control. If humanoids and AMRs scale, safety certification is the gating item, and this puts a public price on it.
Nothing significant turned up in compute or hardware, and nothing in labour policy I could date inside the window. NVIDIA reports Q2 FY27 on 26 August, guided to about $91B, and it no longer breaks out robotics revenue at all, so the most-cited proxy for physical-AI compute demand is not disclosed by the company that dominates it. Self-driving trucks are also now testing on California highways under DMV permits held by Aurora and Kodiak, closing the loop on episode two's permit story. Full disclosure, Kodiak is where I work.
This week in one table
| Item | What it is | Link |
|---|---|---|
| Anthropic IPO expectations | Expects to match or top SpaceX's $75B record; filed confidentially 1 June, $965B Series H, $65B run rate | Bloomberg |
| Unitree STAR Market debut | Priced at 150.80 yuan, closed +460% at ~342B yuan on ~$252M 2025 revenue | SCMP |
| Unitree IPO coverage | The trading-day narrative and HSBC's "illusory" shipment-cycle call | Fortune |
| FCC Covered List addition (28 July) | Blocks equipment authorisation for new foreign-made humanoid, quadruped and AMR models | Forbes |
| Humanoid shipment counts | 22,000+ / 19,100 / 30,000+ from three organisations; AgiBot leads, not Unitree | Forbes |
| DeepSeek's Unitree stake | 2.31% for ~140.8M yuan, 36-month lock-up, joint model development pact | eWeek |
| Wang Xingxing on the timeline | "2 to 3 years at the fastest, and 5 or even 10 years at the slowest" | CNBC |
| Unitree "Superman" | 12.66 m/s claimed plus a 2m standing jump, no protocol, no hands, no head | TechNode |
| World Robot Conference 2026 | Beijing, 19-23 Aug, 300+ exhibitors (+69% YoY), 150+ product launches | Beijing gov |
| τ₀-VLA | Subtask selection as world-model-guided beam search; 50.0% to 74.0% OOD, 40,115 training hours | arXiv |
| Zetta ζ | Frozen base policy, code critics and recovery skills evolved online; 90.8% LIBERO-Pro, 11.1x faster | arXiv |
| EXIMO | VLM-guided exploration, VLA finetuning, residual off-policy RL; no headline numbers in the abstract | arXiv |
| Push-T with agentic robotics | Agent-written controller at 100.0% vs 62.5% for a 200-demo diffusion policy, ~$2,000, no hardware tests | arXiv |
| Hydra-0 | Action flow as a cross-embodiment interface; 90.4%/60.2% error reductions, r=0.96 on RoboLab | arXiv |
| SoftVTBench | 12 of 12 configurations contain "successes" that crushed the object, 0.7% to 24% | arXiv |
| Embodied-Navigator | Pixel selection as navigation action; 66.2% on R2R-CE from 90k trajectories | arXiv |
| Decision-Metric Alignment | Diagnostics and action-conditioned objectives for MPC in latent world models | arXiv |
| GOAG | Generative object-agnostic dexterous grasp planner, no object-specific training | arXiv |
| CoToGrasp | Contact-topology-conditioned dexterous grasp synthesis | arXiv |
| Nevada robotaxi permits | 8,000 vehicles across Tesla, Waymo, Uber and Zoox in Clark County over 12 months | TechCrunch |
| FORT Robotics SPAC | $556.6M pro-forma EV, Nasdaq as FROB, functional safety for machine control | PR Newswire |
| Self-driving trucks in California | Aurora and Kodiak testing on California highways under DMV permits | TechCrunch |
What we're watching
Does a public Anthropic stay in physical AI? If Decart closes after listing, a public frontier lab becomes a physical-AI player with a balance sheet nobody in robotics can match. If the quarterly clock wins, world models are the first line item to look optional.
Does an agent-written controller survive contact with a real Franka? Somebody with an arm and a week could settle it. Until they do, a perfect score in simulation licenses much less than it appears to.
Is the outer loop a permanent layer, or the thing the next generation of weights absorbs and deletes? The history of this field is mostly hand-built structure getting eaten by learned structure, and τ₀-VLA, Zetta and EXIMO are all betting against that pattern.
What does subtask search cost in wall-clock time? τ₀-VLA reports accuracy gained and leaves latency unmeasured, the same gap episode two flagged in TempoWAM's 26.9 percent saving. Beam search is not free at the moment the robot needs to move.
How many other benchmarks have a ground-truth signal nobody is scoring against? SoftVTBench caught violations in twelve of twelve configurations because it happened to have finite-element state sitting alongside the rollouts. How many other datasets already carry a signal like that?
Does the Covered List addition bind, or does it get routed around? Domestic assembly, licensing and a successor entity are all live options, and the answer moves a valuation with roughly a fifth of its revenue coming from the United States.
Who audits a shipment number? Three organisations cannot agree within 50 percent on the input to every valuation in this sector, and nobody separates a factory sale from a sale to a state-subsidised teleoperation centre that sells the training data back.
Papers referenced
- τ₀-VLA: A Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
- Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
- EXIMO: VLM-Guided Exploration of VLA Policies
- Revisiting the Push-T Robot Manipulation Task with Agentic Robotics
- Hydra-0: Action Flow for Generalist World Modeling and Control
- SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
- Embodied-Navigator: Point, Think, Memorize and Align for Efficient Navigation
- Decision-Metric Alignment in Latent World Models
- Unitree Robotics surges to a US$66 billion valuation in Shanghai share debut
- Unitree's China dancing robots IPO trading surge and valuation
- United States bans Chinese humanoid and quadruped robots, citing national security (FCC Covered List)
- Humanoid robot shipments up 300%, up to 30,000 so far in 2026
- DeepSeek takes a stake in Unitree's IPO and signs a joint model-development pact
- Unitree says its new humanoid reaches 12.66 m/s and jumps 2 meters
- Unitree CEO says the humanoid ChatGPT moment is years away
- Tesla, Uber and Waymo all get the OK to operate thousands of robotaxis in Nevada
- Anthropic expects to match SpaceX's record IPO size or top it
- FORT Robotics to go public via business combination with Newbury Street II Acquisition Corp