Who Owns What Robots Learn? The Next Industrial Automation Contract Fight
Key Highlights
- Operational data in robotics behaves like a commodity, blending into shared models with no straightforward way to measure or unlearn individual site contributions.
- Contracts must explicitly allocate rights to source data, training, site-specific models and runtime use to prevent vendor overreach and protect deployers' assets.
- Perpetual licenses and clear documentation rights are essential for maintaining control over site-specific configurations and artifacts after contract termination.
- Market trends show large buyers now negotiate model ownership and training rights directly, recognizing the value of controlling the learning process and associated assets.
- Legal frameworks like the EU Data Act and FTC settlements influence access rights but often leave the model layer under vendor control, emphasizing the need for contractual clarity.
If you run robots in a warehouse or factory, your contract likely gives you the title to your operational data while the vendor earns the unrestricted right to train their models on your facilities.
There is no commercially accepted way to measure what the model gains from your site, and even where a vendor can remove your influence, removal returns nothing you can use. No statute and no standard contract that industrial automation deployers rely on reaches the model layer; only a negotiated contract does.
Before physical AI, I ran a commodities trading fund in electricity and gas. The distinction the industrial automation space keeps missing is one every energy trader lives with. Gas is a commodity you can own: storable, deliverable, severable, title moving with the molecule. Burn it to produce power, and the output enters a grid where it commingles beyond recovery. No meter on earth hands you back your electrons.
Data behaves like gas. The trained model, everything the robotics fleet has learned that makes it better every quarter, behaves like grid power.
Learning cannot be separated out
Once an industrial operator's data has been used to train a shared model, its influence is no longer separable from the model. It has become part of how the model behaves, blended with whatever every other site contributed. There is no copy of “your part” to hand back.
Removing an operator's data from a trained model generally means retraining without it, which is rarely practical at fleet scale: exact unlearning works only under particular architectures. In any case, removal is the wrong test: even if a system successfully unlearns a customer's data, it cannot hand that customer's contribution back as a standalone asset.
Measurement is no better. There is no commercially accepted method that can say what one industrial site's data contributed to fleet performance. A before-and-after benchmark shows that training helped, not whose data helped. Data is severable; learning that has merged into a shared model is not, unless the system was built for removal before training began.
Even the intuitive compromise (“Your fine-tune is yours.”) runs into the same limitation. A site fine-tune is often a small set of adjustments that adapts the vendor's general model to one site. Unlike the shared learning, it can be stored and handed over as a separate file. But it runs only on a compatible base model and runtime. Technical separability does not make it usable on its own, and no contract changes that. What a contract can do is license the customer to keep running it after exit.
The law stops one layer short
Now reread the legal record with that physics in mind.
The proposed FTC John Deere settlement would be a landmark repair-access outcome. Access to repair resources expressly “need not include title to or ownership of any of Deere's intellectual property rights.” Farmers secured diagnostic tools, not the model. The order never reaches the model behind See & Spray, Deere's machine-learning spraying system, which stays with Deere not by carve-out but by silence.
The EU Data Act, in application since September 2025, grants deployers statutory access to the data their machines generate. It reaches raw and pre-processed data. However, Recital 15 excludes information derived through proprietary complex algorithms “unless otherwise agreed.” These three words hand the entire model layer back to contract, which in practice means back to the vendor's paper.
Two instruments, two jurisdictions, one line: statutes carve the model out. Deployers keep winning access to the feedstock while the compounding derivative sits outside every instrument they won it with. Exhibit 1 lays the four instruments side by side against the stack from raw data to fleet model.
The largest buyers are catching on
Until recently, model rights were rarely addressed explicitly. At the top of the market, that has changed.
The largest industrial automation buyers now write the model layer directly into their draft physical AI agreements. I have seen the drafts and negotiated against them. Their claims climb a ladder: ownership of raw data, then of derived data, then of customer-specific weights, then perpetual licenses to any underlying model required to run those weights. The last rung is architecture itself: royalty triggers on systems that merely resemble what was built together. Buyers have understood that a fine-tune runs only on the vendor's base, so the claims climb until they reach the platform.
Two details stand out. First, the same documents that demand perpetual economics on anything that benefits from a robotics deployment treat hardware as a cost-plus line item. When the most sophisticated automation buyers allocate negotiating capital, they publish a depreciation schedule. Hardware depreciates. Learning compounds.
Second, the hardest-fought clause is rarely data ownership. It is pool membership: whether one customer's live production data may train the fleet’s shared brain. Spend a quarter teaching the fleet to handle your most fragile SKU, and that skill may ship to a competitor working with the same vendor in the next update. While a single site's deployment data might be a fraction of the total training data, those production edge cases carry disproportionate value.
Today fleet models remain stubbornly site-specific, and a fine-tune from one facility often degrades in the next building over. But the perpetual licenses being signed now will govern the models of five years from now, if cross-site transfer works. The asset is being allocated before it fully exists, which is exactly when it is cheapest to win by default.
Automation buyers demand consent gates on training, and the gate is technically real: site-specific fine-tuning can increasingly run on a single on-premises GPU, operational data never leaving the facility. But commingling is not inevitable. It is a pipeline decision made upstream, in the contract. Vendors defend the flywheel, base models trained on anonymized data across all sites, improving all customers. Some drafts now gesture at the honest resolution: a priced data license for training rights. Others impose model provenance obligations, documented inputs and decision logic, years ahead of any regulator.
Meanwhile the same buyers ask for source code: repositories, embedded engineering teams, escrow with hair-trigger release. Source-code access is an incomplete remedy; it transfers a snapshot, not the learning loop that improves it.
Vendors are already claiming the learning
Richtech Robotics' May 2026 8-K identifies robotics data collection and World Action Model training as intended uses of its newly acquired Las Vegas facility. Shanghai Seer's Hong Kong Public Offering prospectus describes a dual flywheel in which customer operational data may enter training when voluntarily shared with consent, while Seer retains its core technologies.
Read the published robotics terms side by side and the same structure recurs. The customer keeps title to raw data and facility specifics. The vendor takes a perpetual, royalty-free license to operational telemetry through improvements and aggregated-data clauses. The learning follows the license.
Where a no-train commitment exists at all, it is often scoped to personal data rather than operational data, derived data or customer artifacts. Terms of sale for Boston Dynamics' Spot, the most explicit I have found, define “Robot Technical Data” to include the geometric, terrain and image information the robot observes on site, permit its indefinite retention, and grant a perpetual right to use performance and application data for product improvement. So the market splits: the head negotiates the ladder clause by clause, and the long tail signs boilerplate that settles the same question by default.
What the contract must allocate
The constructive answer follows the energy precedent. Power markets could not return anyone's electrons, so they unbundled the attributes they could not separate physically, metered them and traded them as certificates. Robotics has the blending but no meter for learning, so it can copy only the contractual half, and must do it up front: define the derivative, allocate it, price it. The clause is not exotic: life-sciences agreements have long used unblocking licenses to stop follow-on IP from stranding the other party's platform.
At minimum, the agreement must allocate four things separately: (1) source data; (2) training rights; (3) site-specific models and site memory; and (4) the runtime rights required to use those artifacts after termination.
The third, site-specific models and site memory, is broader than a trained model. A robotics fleet now also improves by remembering: it keeps a record of what worked and what failed at a site and reuses it on the next job, without retraining the model at all. That site memory alone is an asset. Who owns it, who may reuse it and whether it leaves with the customer at exit all need settling, because a right to train is no longer the only way a vendor learns from a site. Exhibit 2 sets out the four allocations and the settlement they add up to.
The workable settlement is visible in the negotiations now underway at the head of the market. The vendor owns the base model and the cross-customer flywheel. The customer owns its site-specific fine-tunes, configurations and site memory, with a runtime license that makes them usable. Training rights are priced separately rather than smuggled through the service boilerplate. Documentation rights make the whole thing auditable.
Industrial automation operators should also know how the other side of the table will evaluate success. Data-unencumbered ARR is the metric I would put next to net revenue retention. This is ARR generated under contracts giving the vendor an unrestricted right to train shared models on deployment data, meaning no customer-specific contractual restriction on cross-customer training. It measures contractual permission, not the value of the data or what the training achieved. Two companies with identical ARR can hold entirely different assets. One holds the right to compound a moat. The other is running a service on whatever data it can learn from inside its own boundaries. Exhibit 3 sets out the formula, the four contract classes and an illustrative comparison.
What industrial automation operators and boards should do now
Two developments would weaken this thesis: a law giving automation deployers access to trained models, or fleets that train on site by default. Each exists in early form; neither exists at commercial scale today.
Certified removal of one customer's data would strengthen a customer's hand without returning anything usable. What would undermine the economics rather than the thesis is different: base models becoming a commodity or cross-site transfer never working, which would leave the fleet flywheel worth less than vendors claim. Meanwhile, the contracts being negotiated and signed today allocate tomorrow's asset: what the robot learned.
Robotics agreements should require four provisions before boards approve the deployment:
- Securing the source data. Title to raw and pre-processed data must be explicitly severable and returnable.
- Pricing the training right. Decide pool membership deliberately. If you join the pool, price the training right now on its own commercial terms, not inside the service boilerplate.
- Defining customer artifacts. Fine-tunes, configurations and site memory, named in the contract, not implied. Read the no-train clause closely; check whether it covers operational data, derived data and customer artifacts, or only personal data.
- Demanding a usable exit. Secure a runtime license and documentation rights that keep those artifacts working after termination. And ask the vendor what share of its revenue carries cross-customer training rights; the answer tells you what permissions it already holds over customers like you.
Don’t be the farmer who won repair access and still rents the model back by the acre. You can repatriate your data. You can’t repatriate what it taught the fleet.
About the Author
Thomas Stapp Thomas Stapp
Thomas Stapp is an investor, operator and member of the executive leadership team at GreyOrange, a leader in AI-powered orchestration software for supply chain and retail operations, where he leads strategy and corporate development. He is also an investor and advisor to Feather Robotics and Mbodi and serves on the board of Cartesian Kinetics.
Leaders LogoLeaders relevant to this article:



