For decades, robots in factories succeeded by doing one thing, perfectly, forever. The new generation inverts that: instead of programming a robot for a task, train a general model on thousands of hours of humans doing tasks — teleoperation, motion capture, video — and let the model generalize to variations the programmers never anticipated.

The architectures are converging: a vision backbone that sees the scene, a language interface for instructions, and an action head that outputs joint commands at control frequency. The differences between labs are increasingly in data pipeline and training infrastructure, not in the idea.

Why it matters

Humanoid form factors are a bet on the existing world: stairs, doors, pallets and workstations were built for humans, so a machine shaped like one can, in principle, slot into environments that were never designed for robots. That is why the same companies talk about warehouses today and elder care or household work later.

The economic test is stark. A humanoid platform costs somewhere between a car and a small aircraft, and the labor it targets costs somewhere between minimum wage and skilled-trade rates. The technology only matters if uptime, reliability and task quality clear the bar across a multi-year depreciation window.

How the models actually learn

The data problem dominates. Demonstration teleoperation is expensive and slow — a human operator drives the robot through tasks while the system records joint states, camera views and gripper actions. Labs supplement this with simulation, video pretraining from internet-scale footage, and increasingly, model-generated synthetic tasks.

Generalization is the frontier. Models that perform a task in one kitchen or one warehouse bay may fail under different lighting, a different object placement, or a novel distractor. Benchmark suites for robot generalization are young and contested, which makes vendor claims hard to compare — a familiar pattern from the language-model era, arriving again with motors attached.

Evidence

Commercial signals are real but early: logistics operators and manufacturers have signed pilot agreements with multiple humanoid companies, and several labs have published model weights or research papers showing multi-task performance from single models. Funding has flowed at valuations that assume the economics work.

Independent verification remains thin. Most published results come from the vendors themselves; standardized, third-party evaluations of reliability, mean time between failures and task completion rates in production settings are still rare — the DailyTech pattern-check applies here as in any early market.

The competing read

Skeptics note that specialized automation — fixed arms, conveyors, AMRs — already does most economically valuable physical tasks better and cheaper than a bipedal generalist, and that the humanoid bet only pays off if the model layer truly generalizes across environments no integrator ever programmed.

Proponents respond that generality is precisely the product: a fleet of robots that can be retasked with a sentence rather than a re-engineering project. History is not encouraging — general-purpose robots have been five years away for fifty years — but the foundation-model lever is genuinely new, and it is the first mechanism that has ever plausibly addressed the generality gap at scale.

What happens next

Watch for the first published, third-party audited deployment metrics; watch which labs open their model weights (openness accelerates the field the way it did for language models); and watch the cost curve of the actuators, hands and batteries that dominate the bill of materials. The models are improving quickly. The physics of business cases is slower.