Physical AI training data does not exist at the scale the robotics industry needs, and a startup called Encord is building a business around producing it. Inside a warehouse in San Leandro, California, a worker named Andrew Ceja plays Jenga while wearing a headset that films what he sees and measures his brain waves. It is an unlikely assembly line for one of technology’s most urgent supply-chain problems.
Brain Waves, Muscle Signals, and the Search for Better Robot Data
The brain wave headset was built by Zander Labs, a German neuroscience startup with offices in Berlin, Munich, Cottbus, and Delft. Zander is developing a wearable EEG system called the Zypher Suite for real-time brain activity monitoring, with local data processing intended to address privacy concerns. The theory, explained by Lucas Gehrke, a Zander neuroscientist supervising the Encord trial, is that the volume of brain activity at any point during a task offers clues about when AI models need to deploy their highest-effort processing.
Encord’s arrangement with Zander is currently a trial run. The goal is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it improves performance before deciding whether to scale. Vineeth Velmurugan, Encord’s head of robot learning and a veteran of OpenAI’s robot lab and the warehouse automation firm Berkshire Grey, describes it as the ‘bleeding edge’ of the effort to solve the robotics data bottleneck.
Brain waves are not the only new modality Encord is pursuing. A set of forearm sensors detects electrical signals in muscles, helping build a 3D picture of hand position during manipulation tasks, an angle that standard egocentric video typically misses. At other stations in the San Leandro facility, pilots use leader-follower rigs, paired robotic arms where one human-operated arm and one mimic arm generate training data around tasks like pouring coffee and stacking poker chips. Another pilot, Sofia Infante, practises plugging and unplugging ethernet cables from a server rack, the kind of fine-motor work data centre operators want automated but that robotic pincers, with fewer degrees of freedom than human fingers, still cannot reliably perform.
The Physical AI Training Data Problem Is an Economics Problem
Velmurugan estimates that densely annotated physical AI training data, tagged with physical descriptions such as ‘right hand tightens bolt’, is worth 100 times as much as raw egocentric footage for training specific tasks. It costs roughly 20 times more to produce. On paper, that is a sound trade. In practice, it exposes the fundamental difference between physical AI and the large language models that inspired it.
LLM builders scraped text from Stack Overflow and the wider web at negligible cost. Physical training data has to be manufactured. Velmurugan believes it will take a data set something like five times the size of YouTube’s entire video corpus to break through the current ceiling. That figure helps explain why data generation has become a commercial category rather than a research side-project.
Encord’s funding reflects that conviction. Incorporated as Cord Technologies Inc., according to SiliconANGLE, the company launched during Y Combinator’s winter 2021 batch as a two-person team. It raised $30 million in a Series B before closing a $60 million Series C led by Wellington Management, with existing backers Y Combinator, CRV, N47, Crane Venture Partners, and Harpoon Ventures joined by Bright Pixel Capital and Isomer Capital, bringing total funding to $110 million. The company reported that its physical AI revenue grew 10x in the twelve months preceding that announcement. Its platform is now used by more than 300 AI teams across data modalities ranging from video and audio to 3D point clouds and sensor data.
The competitive pressure is building. XDOF, founded in October 2024 by researchers from UC Berkeley, emerged from stealth in June 2026 with $70 million in backing from Thrive Capital, Spark Capital, a16z, Lux, and WndrCo. The company already counts approximately 20 customers and has released a 130,000-trajectory manipulation dataset developed in collaboration with Berkeley’s AI research lab. The data-factory model is attracting serious capital on multiple fronts.
Encord’s pitch rests on its position across multiple customers simultaneously. By working with many robotics firms at once, Velmurugan argues, the company can see which data techniques are gaining traction industry-wide before any single customer can. Ceja and Infante, both formerly at Scale AI, are part of what Velmurugan calls a manufacturing workforce for neural networks, a dozen or so pilots keeping the San Leandro facility running.
‘Every humanoid company has asked us for these pieces,’ Velmurugan says. The question hanging over all of it is whether the economics of manufacturing physical AI training data can compress fast enough to keep pace with the ambitions of the robotics labs buying it.
