Data point
Joint actuators as share of humanoid robot bill-of-materials cost, 2025
Reported share of a humanoid robot's total bill-of-materials cost attributable to joint actuators, by robot configuration, 2025 estimates
Source: Interact Analysis, "Joint actuators: the fundamental component for humanoid robots' power and dexterity" - Figures are reported as thresholds (>30%, >50%), not point estimates. Chart values use the reported lower bound; actual shares may run higher.
Marco’s framing is that physical AI progress will mainly come from software (foundation models, simulation, sim-to-real transfer) rather than from hardware innovation (actuators, sensors, manipulators). This is a specific and testable claim, and 2025 gave it real evidence on both sides.
The software case
The strongest support for the software thesis is the sudden appearance of general-purpose vision-language-action (VLA) models that control humanoid bodies without body-specific retraining. Figure AI’s Helix, released February 20, 2025, was the first VLA to run high-rate continuous control of an entire humanoid upper body, including wrists, torso, head and individual fingers, and the first VLA to run two robots simultaneously on a shared manipulation task with objects neither had seen before. NVIDIA followed in March 2025 with GR00T N1, and Google DeepMind extended its Gemini 2.0 backbone into Gemini Robotics the same year. Three separate labs shipped generalist control models within about a month of each other, each claiming transfer across tasks and, in some cases, across robot bodies. If that cadence holds, it argues that the constraint on physical AI is increasingly data and model architecture, not the actuator underneath.
The hardware constraint that doesn’t move as fast
The software narrative has to clear a harder bar: does the physical body stop being the bottleneck. On current evidence, no. Joint actuators alone account for more than 30 percent of the bill-of-materials cost in a high-configuration humanoid, and can exceed 50 percent in basic versions without dexterous hands or high-end sensors, according to Interact Analysis. Tesla’s Optimus alone carries more than 28 rotary and linear actuators, each one a mechanical part with its own reliability, weight and cost curve. This is a manufacturing and materials problem, and those tend to improve linearly, not exponentially.
Rodney Brooks, the iRobot co-founder who has spent decades building physical robots, made this argument directly in a TechCrunch-covered essay from September 2025: human hands carry about 17,000 specialized touch receptors that no current robot hand approaches, and teaching dexterity purely from video, the approach behind most current VLA training, is what he calls pure fantasy thinking. His broader point is structural: robots that move physical mass do not get twice as capable every year the way a software model can get twice as good on a benchmark. Better software can route around some hardware limits, faster grasp planning can compensate for imprecise fingers, but it cannot substitute for touch sensing that isn’t there or torque an actuator can’t sustain over a factory shift.
What this means for the thesis
The honest read is that Marco’s framing is half right and probably premature as stated. Software is clearly the faster-moving layer right now, and the VLA releases of 2025 are a genuine inflection, because they show transfer across robot bodies rather than just improved benchmarks on one body. But the claim that physical AI will be mainly a software affair implicitly treats hardware as a solved or shrinking cost, and the bill-of-materials data says otherwise: actuators and hands remain the largest cost line in the system, and that share has not compressed the way GPU-driven software costs have. A more defensible version of the thesis is that software determines which tasks a given generation of hardware can attempt, while hardware determines which tasks are economically viable at all. Both curves matter, and betting only on the software curve assumes a hardware cost decline that has not yet been demonstrated at scale.
Countercase
The cleanest challenge to Marco’s framing is that the VLA evidence is a selection effect. The models getting attention (Helix, GR00T N1, Gemini Robotics) are demonstrations from well-funded labs on curated tasks, not audited production deployments at scale. Running on two robot bodies in a demo is a different claim from hardware no longer constraining commercial deployment. There is also a narrative fallacy risk: a story where software eats the hardware problem is a compelling pitch to investors because it makes today’s high hardware costs look temporary, and a similar story has been told about robotics before without paying off. On the hardware side, Brooks’ touch-receptor comparison and the bill-of-materials data are strong directional evidence but do not by themselves prove hardware costs will fail to fall. Actuator and sensor costs could still compress with manufacturing scale the way batteries and solar panels have, just on a longer timeline than software. The Interact Analysis figures are a 2025 snapshot, not a trend line, so this note cannot show whether the hardware cost share is falling, flat or rising over time. Evidence for that trajectory is not yet public in a form this note can cite responsibly.
