The Bridge
On defense procurement, the manipulation gap, and the data the country has not generated
In 1948, John Parsons, a manufacturer in Traverse City, Michigan, won an Air Force contract to produce helicopter rotor blade templates whose curves were too complex for manual machining to replicate at the precision the military demanded. A year later, the Air Force expanded the project and brought in MIT’s Servomechanisms Laboratory for the servo control expertise Parsons did not have. By 1952, the lab had built the first numerically controlled machine tool — a retrofitted Cincinnati milling machine that read instructions from punched paper tape and cut metal parts to specifications no human operator could match consistently.
The Air Force did not set out to transform manufacturing. It needed helicopter blades. But the capability that emerged from that need — the automation of machining itself — propagated outward. By the late 1950s, commercial machine tool builders were licensing MIT’s innovations. Within two decades, numerically controlled machining had become the foundation of precision manufacturing worldwide. Modern NC/CNC traces its lineage to a military contract that was trying to solve a problem with rotor blades.
The pattern is not unique to machining. Around 1963, the United States government was purchasing roughly 95 percent of all integrated circuits produced in the country — not as a subsidy but because the military needed chips and the civilian market did not yet exist at the volumes required. That demand floor gave Fairchild and Texas Instruments the revenue to justify production no civilian customer could underwrite, and drove cost reductions that eventually made civilian applications viable. Boeing and Douglas built their commercial aviation businesses on production capacity and engineering discipline funded by military contracts. The DARPA Grand Challenge did not produce autonomous vehicles, but it did produce the human capital pipeline that produced autonomous vehicles a decade later, for roughly $5.5 million in prize money awarded across the 2005 and 2007 events.
It is a simple pattern: defense procurement built much of the American technology base. But the standard account misses what the military actually did. Beyond providing demand, it imposed specifications the commercial market would never have set for itself, calibrated against operational realities the commercial market did not face. The capabilities that later moved into civilian industry were forged against demands the civilian market would never have generated on its own.
As I was re-reading the Dune novels last week, I realized that Frank Herbert understood something about this. The Fremen of Dune are made by the desert; the environment is not the obstacle but the mechanism that produces capability. History is more specific than Herbert’s fiction — the Pentagon imposed not raw harshness but demand-as-discipline — but the underlying logic holds. The Air Force’s precision requirements for numerical control were tighter than commercial manufacturing needed. The DARPA Grand Challenge course was more demanding than any commercial driving scenario. Each time, the military demanded capability against standards the commercial market would not have imposed on itself. Each time, the capability that emerged is what made the commercial market possible.
This is where I part ways with what has become the standard account of America’s robotics challenge.
Martin Casado and Anne Neuberger have written, in my view, the sharpest articulation of the competitive stakes: China is running away with the physical AI race, the data flywheel gives it a compounding advantage, and the United States is not set up to win. I think they are right about all of this. I also think their own argument contains a stronger version of itself that they did not pursue. Their call is to go “from permission-first to permissionless” — strip away regulatory accretion, align with allies, let entrepreneurial firms compete. Necessary, and I will trace several of these regulatory obstacles later. But in a single sentence, they note that the United States “won the key industries of the twentieth century” through “a dynamic and open market, sometimes combined with strategic government support — from Boeing and Lockheed to IBM and Intel.” That sentence deserves more weight than it gets.
Boeing and Lockheed, IBM and Intel — these are companies whose commercial businesses were built on production capacity and engineering discipline funded by military contracts. The government was the principal customer for the integrated circuit industry for over a decade, underwriting production at quality standards the civilian market had no reason to demand. The capacity it funded is what made the eventual commercial market possible. The history Casado and Neuberger invoke in passing is not context for the deregulation argument. It is a different mechanism — and for robotics, I think it is the more important one.
Deregulation is necessary. Permitting delays, fragmented state rules, and adversarial labor arrangements are real obstacles, and I plan on tracing several of them (hopefully) next week. But deregulation is a theory about removing friction. The historical mechanism that built American technology industries is a theory about creating demand. These are not the same thing, and the distinction matters more for robotics than it did for anything before it.
Removing friction will get you more robots in warehouses — structured environments, known objects, predictable workflows. That is valuable and insufficient. It will not get you the environments that close the manipulation gap. And because robotics is a technology whose capabilities are shaped by the environments in which it is deployed, where the commercial market deploys is not a secondary consideration. It is, in my mind, the strategic question.
For robotics, the demand mechanism is more consequential than for any previous technology. The reason is structural, and I want to be precise about it.
For the industries the Pentagon built in the twentieth century, specification and operating condition were more separable. Once a supplier met the military requirement, the underlying capability could transfer into civilian markets even if packaging, screening, and reliability qualification differed. A machine tool cut to Air Force precision is the same machine tool on a factory floor. The military imposed harder requirements, the supplier met them, the capability transferred. The product was fixed at the point of sale. It was cheaper and more reliable over time; it was not a different product.
A robot governed by a vision-language-action model is not fixed at the point of sale. Every deployment generates training data — every grasp, every collision, every failed insertion, every successful recovery — and that data feeds the models governing the next generation of capability. The robot that has spent three years navigating the interior of a destroyer’s engine compartment has encountered spatial configurations, material conditions, and failure modes that no simulation and no warehouse deployment could replicate. That experience lives in the model. The robot that has handled surgical instruments under vibration, inconsistent lighting, and unpredictable patient movement in a field hospital has generated data that shapes what subsequent robots can do in conditions a laboratory setting cannot reproduce. The specification the robot meets in year three was not written by the procurement officer at contract award. It was generated by where the robot was sent.
This is the data flywheel: deployment generates data, the data improves models, and better models enable more deployment. The critical variable is not volume but diversity. A thousand robots in a thousand identical warehouses generate a thousand copies of similar data. A hundred robots across ship maintenance, field repair, disaster response, medical evacuation, base logistics, and contested resupply generate data across conditions varied enough that the resulting models generalize in ways narrow deployment cannot produce.
For chips, the Pentagon’s contribution was demand carrying discipline: it imposed a standard, and when the price came down the standard transferred. For VLA-driven robotics, demand determines the training distribution itself. What the technology becomes is determined by where it is sent.
The Pentagon operates across some of the harshest and most operationally varied environments any single institution controls. It is the one institution that routinely sends people and equipment into confined spaces, extreme temperatures, unpredictable terrain, degraded infrastructure, and contested conditions where the tolerance for failure is not a warranty claim but a life. That operational diversity is not a side benefit of defense procurement for robotics. It is the asset the mechanism generates. A deregulated commercial market produces models trained on commercial environments. Anything that generalizes beyond them has to be trained somewhere, by someone, under conditions the commercial market will not pay for.
The mechanism is not entirely absent. In March, the Navy and the General Services Administration awarded Gecko Robotics — a Pittsburgh company (Let’s Go Pens!) that builds wall-climbing inspection robots — a five-year IDIQ contract with a $71 million ceiling to assess the health of warships across the Pacific Fleet. Eighteen ships. Destroyers, amphibious warships, littoral combat ships. Gecko says its robots identify structural problems up to fifty times faster and more accurately than manual methods. The Navy’s Chief Technology Officer, Justin Fanelli, framed the problem with unusual clarity: “Cracking the cost equation is just as important as cracking the physics equation.”
The contract shows the mechanism in miniature. The Navy’s binding constraint is maintenance backlog; the bottleneck within maintenance is inspection; the Chief of Naval Operations has set an aspirational target of 80 percent fleet combat-surge readiness by 2027, against a 2025 baseline in which fewer than half of ships completed scheduled repairs on time. Gecko’s robots crawl hulls, carry ultrasonic sensors, map structural conditions at resolutions human inspectors cannot match, and do it in spaces that are unsafe or inaccessible for human workers. And the data they generate — eighteen ships’ worth of hull conditions, weld geometries, corrosion patterns, defect morphologies under real operating wear — is training data that improves the next generation of Gecko’s inspection models. The flywheel is turning. For inspection.
Gecko’s robots look at ships. They do not repair them. And the distance between those two sentences is the distance between what physical AI has solved and what it has not.
Perception is far more tractable than manipulation. Vision models generalize; anomaly detection inherits from a decade of deep learning on natural images; navigation in many structured spaces is increasingly mature. A robot that crawls a hull, identifies a corroded weld, and produces a defect map is doing a hard job at the frontier of what’s currently deployable, but the underlying capabilities — see, classify, localize — are the ones the field is most confident in.
The leap from “identify the corroded section” to “remove the corroded section and install a replacement” is a different problem. It is the leap from perception to contact-rich manipulation — managing force interactions with objects the robot did not know about, under conditions it cannot fully see, using tolerances it has to feel rather than measure. Unfamiliar fasteners. Corroded fittings. Tools the robot has never handled. This is the frontier problem in physical AI, and it is the reason every serious company is racing to collect manipulation data at scale. Without it, the models do not generalize. With it, they might.
Simulation is the obvious response, and simulation is doing real work. NVIDIA’s Isaac Lab and DeepMind’s Genie are training policies in environments that did not exist before. For rigid-body, well-characterized contact at scale, simulation will get you a long way. But the tasks that matter most for the kinds of environments the Pentagon operates in — contact with surfaces whose friction, compliance, and material history cannot be specified in advance — are exactly the tasks where the sim-to-real gap is widest. You can simulate the mating of two clean fasteners. You cannot simulate the mating of a fastener that has been saltwater-corroded for six years in a compartment whose geometry does not match any CAD model. The diversity of physical conditions the Pentagon operates across is not a side benefit of its footprint. It is the training distribution nothing else can generate.
That is the gap. A robot that can navigate a destroyer’s bilge and diagnose a structural deficiency is valuable — Gecko has demonstrated it. A robot that can navigate the same bilge and fix the deficiency would be transformative. And the training data that takes the second robot from impossible to reliable can only be generated by deploying robots into the environments where the first one is already working. Not at scale, and not yet reliable, but deployed. Under conditions the commercial market will not create.
The desert makes the Fremen. The question is whether anyone is willing to enter the desert.
The mechanism is still idle. The authorities exist. The obstacle is simpler: procurement for a technology whose specification is generated by deployment breaks an assumption the system has never had to question.
Every existing mechanism for managing change in procured systems works by specifying the envelope of allowed change in advance. The FDA’s Predetermined Change Control Plan tells a sponsor to enumerate, at submission, the kinds of updates an AI/ML medical device will undergo and how each class of change will be validated. DoD’s continuous ATO lets a software system be re-authorized on a compressed timeline against a defined posture. The Software Acquisition Pathway accepts that capabilities will be delivered iteratively, against requirements that evolve between drops. All three depend on the same premise: the space of things the product might become is knowable at the outset, even if the specific instances are not. That is what lets the oversight framework travel with the product.
A VLA-driven robot in an operationally demanding environment does not have that property. The envelope is itself generated by deployment. The robot that operates in year three of an IDIQ is not what the manufacturer delivered, and the difference is not a planned release against an enumerated change-control plan — it is the accumulated effect of three years of training data from conditions nobody specified in advance because nobody could. It perceives differently. It handles failure modes it could not handle at delivery. Its manipulation repertoire has expanded in ways the acceptance test could not have measured. I have not found anyone in the acquisition system who is tracking this. The system assumes that whatever else a procured capability does, it operates within an envelope that was bounded at contract award. The envelope is what the system oversees. Robotics that learns from deployment has no envelope — or rather, its envelope is whatever the training data generates. An F-35 does not learn to fly better the longer it is in service. A robot does. The procurement system has partial language for the first case. It has no language for the second.
The integration problem extends the same claim. A robot is not a standalone deliverable; it is a system that has to be woven into existing workflows, physical spaces, and human teams. Gecko works in part because the robot supplements the existing inspection process — sailors are still there, the robot is a tool they use, the output feeds a maintenance decision a human makes. Harder cases — a robot that replaces a maintenance task, that works alongside a sailor in a confined space, that must hand off a partially completed repair — require more than a product purchase. They require the human team to train the robot, correct its failures, expose it to conditions that shape what it becomes next. The workforce doing the deploying is, in effect, writing the specification. The FAR was not written for this. The OTA structures that might accommodate it are still feeling their way toward what the terms should look like. And the workforce question — who operates the robot, who corrects it, whose judgments shape its learning — is the question the procurement system treats as someone else’s problem.
These obstacles are real, and they are solvable. The Office of Strategic Capital was designed for exactly this kind of market-shaping investment. Other Transaction Authorities provide the contracting flexibility FAR-based vehicles lack. Price-Anderson’s indemnification framework demonstrates that the government can underwrite risk in emerging technologies to enable deployment the private market would not otherwise attempt. The authorities and tools already exist. What is missing is the recognition that robotics procurement is a different category from anything the system has handled before — and that the difference is not that the product changes over time, which is old news. It is that the envelope of change is itself written by deployment.
China’s civil-military fusion doctrine does not solve the envelope problem. It sidesteps it. The institutional boundary the American system polices between commercial robotics and military capability does not exist on the other side in the same form — Unitree is partly owned by the state, PLA systems draw on commercially developed platforms, and under civil-military fusion the institutional barriers that would normally slow or block the movement of deployment knowledge, model improvements, and operational lessons into military capability development are much weaker than in the American system. The boundary has not been abolished. It has been made porous in the direction the state needs.
The standard framing understates the problem. China’s industrial advantage is well-documented. The informational advantage is less remarked on and more consequential for the argument I have been making. Chinese deployment at scale across electronics assembly, automotive manufacturing, logistics, and increasingly service and military environments generates training data at a pace and diversity no other country matches. Under civil-military fusion, that data is available for military applications on a compressed institutional timeline. The training distribution grows faster. The envelope — in the sense I used it in the last section — is written by a system designed for the data to flow.
The United States cannot build that system, and it should not try. The friction between Silicon Valley and the Pentagon, the ability of a company to negotiate terms, the existence of institutional boundaries between civilian and military technology: these are features of a state that does not own its technology companies. They are worth preserving, and preserving them is not free. Every institutional friction that protects the independence of a private firm is also a delay in the training data reaching the systems that most need it. I am not confident the American system can generate the deployment diversity it needs at the speed it needs without institutional changes that nobody in Washington is currently proposing. But I am confident about where the work starts.
There are two kinds of friction in the American system. The friction that protects a free society is not negotiable. The friction that prevents the United States from pointing defense procurement at the environments where the manipulation gap closes is procedural, and it is the kind that the authorities described a moment ago — OSC, OTA, Price-Anderson — were built to address. Separating the two is the work. Every month the work is not done is a month of training distribution the competition fills and the United States does not.
Every bridge the United States has built between its commercial technology sector and its defense establishment was built for a product whose specification was fixed at the point of sale. The procurement apparatus, the oversight framework, the contracting vehicles, the cost models — all of them assume that what was bought is what was deployed, and that whatever changes afterward happens inside an envelope bounded at acquisition. That assumption has held for every industry the historical mechanism built. It does not hold for physical AI.
The bridge robotics requires is for a category the procurement system has not encountered before: a system whose capabilities are shaped by the environments it operates in, whose year-three specification is written by deployment in years one and two, and whose training distribution is, in effect, what the government is buying, even though nobody in the acquisition process is tracking it that way.
The mechanism that built Boeing and Fairchild has to be pointed at a different kind of target. Not cheaper robots in warehouses. Not faster approvals for commercial deployment. Manipulation in the tail of the distribution, in the environments no commercial market will pay for, on a timeline short enough that the training data the competition is already generating does not become a lead the United States cannot close. The authorities exist. The technology is far enough along that perception is deploying and manipulation is a tractable research frontier. The gap is recognizing that what is being procured is the training distribution — and acting accordingly.
That gap is closable. It is also, at the moment, not being closed.



The real pivot in this piece isn't the policy argument. It's the ontological one you slipped past everyone. You're saying a VLA-driven robot is not a product but a trajectory: "its envelope is whatever the training data generates," which means what it *is* at year three is a function of everywhere it's been sent. The specification is retrospective, not prospective. That's a strange object for any procurement system to hold, because the system assumes the thing it's buying has a fixed identity at contract award. You've shown it doesn't, and that the diversity of the traversal (not its volume) is what determines generalization. The Fremen analogy undersells your own insight. Herbert's desert is uniform harshness. What you're actually describing is that the *heterogeneity* of the terrain writes the capability. A thousand identical warehouses are a basin. The Pentagon's value is that it breaks the basin open.
— Iman and Darja