A Jenga tower in a San Leandro, California warehouse has become a testbed for physical AI. The warehouse is operated by Encord, a company that builds data tooling and annotation services used to train machine learning models. Andrew Ceja — one of Encord’s so‑called pilots, the workers who generate robotic training data — wears a headset with a camera while carefully pulling wooden blocks from the unstable tower. The headset records what he sees and, uniquely in this trial, also measures his brain activity while he disassembles the tower.
The brain‑wave headset was developed by Zander Labs, a German neuroscience startup that aims to infer mental states such as error, intent and surprise from neural signals. Encord’s collaboration with Zander is a pilot: Encord plans to produce an initial dataset tagged with brain‑wave signals, run it through customer robotics models, and evaluate whether these labels improve performance before deciding on a broader rollout.
Lucas Gehrke, a neuroscientist at Zander supervising the work, says the amount of brain activity at particular moments during a task can inform model developers about when to deploy their highest‑effort models.
Why the approach matters
Encord is among a small but growing set of startups that believe the next real bottleneck for humanoid and warehouse robotics is scarcity of real‑world physical training data rather than model architecture. Encord began by helping companies that build machine‑vision applications annotate data and evaluate models. As customers started applying end‑to‑end learning to manipulation tasks, Encord’s leadership realized those customers often needed to produce training data themselves. “The data simply does not exist,” said Vineeth Velmurugan, Encord’s head of robot learning.
The comparison to LLMs only goes so far: large language models were trained on massive corpora scraped from the internet, but obtaining comparable volumes of physical manipulation data is far harder. Self‑driving car companies collect their own data, but that approach is difficult to scale. Video training helps but lacks the fidelity of real‑world multimodal data. Velmurugan estimates a breakthrough may require a dataset on the order of multiple times YouTube’s corpus, which helps explain why data generation has become a commercial activity as well as a research problem.
Two main data sources: egocentric video and remotely operated robots
Robotics teams today primarily rely on two kinds of data: egocentric video recorded by workers wearing head‑mounted cameras (often augmented with additional camera angles and other measurements), and data captured from robots operated remotely. Encord uses both: it pulls egocentric data from factories worldwide and uses its San Leandro facility to experiment with new modalities such as brain waves and to collect task‑specific datasets for fine‑tuning.
During a TechCrunch visit, pilots used leader‑follower rigs — pairs of robotic arms where one is directly controlled by a human and the other mirrors its movements — to produce data for tasks like pouring coffee from a pot into mugs (notoriously difficult because of splashing) and stacking poker chips. “Every humanoid company has asked us for these pieces,” Velmurugan said.
Storage racks in the warehouse hold cartons of fake flowers in vases, books, plastic vegetables, cat litter trays and scoops, bags and bundles of wires — a set of props intended to train manipulators for household tasks.
At one station, pilot Sofia Infante controlled robotic arms to plug and unplug Ethernet cables from the back of a server, the type of work data center operators would like to automate if robots could meet the necessary precision. Hands‑on experience highlights why this is still difficult: pincers lack the dexterity and degrees of freedom of human fingers and arms.
EMG, dense annotations and economics
Another sensor modality Encord is developing uses forearm sensors to detect electrical muscle signals (EMG). Video of human hands manipulating objects often fails to capture the whole hand, but Velmurugan hopes arm sensors can support a 3D reconstruction of hand position over time, producing a more robust input for models.
Encord’s datasets are annotated with physical descriptions of actions — for example, “right hand tightens bolt” — to help LLM‑based systems interpret what occurs. Velmurugan estimates that this kind of dense annotation can be 100 times more valuable for training a specific task than lower‑quality egocentric data, and that it costs roughly 20 times more to produce — a tradeoff that looks reasonable on paper.
The catch, however, is cost: scraping text from the internet to build LLMs was comparatively cheap, while manufacturing physical training data is expensive. That difference changes the economics of building robot models because such data must be created, not just collected.
Industry visibility and the workforce
Velmurugan says progress is happening and that Encord’s visibility across many programs allows the company to spot which data techniques work. That cross‑customer vantage point is part of Encord’s value proposition: noticing industry‑wide trends before a single customer does.
Roughly a dozen pilots work at Encord’s San Leandro facility producing the building blocks for neural networks. Both Infante and Ceja previously worked at Scale, another AI data annotation firm, before joining Encord. Ceja previously maintained a robotic waste sorter at a waste management company and now says he enjoys solving daily new challenges for robot training: “It’s something new every day!”



