⚡ AI Snapshot
- Brain-wave headsets tag robot training data
- Physical data must be manufactured, not scraped
- Dense annotation worth 100x, costs 20x
The update
Encord, a data-tooling startup with a warehouse in San Leandro, is trialing a headset from German neuroscience firm Zander Labs that measures a human pilot's brain waves while they perform manipulation tasks like disassembling a Jenga tower. The goal is to infer mental states—error, intent, surprise—and tag training data with them. It's an early trial: Encord plans to build an initial brain-wave-tagged data set, run it through customer robotics models, and check whether it actually improves performance before scaling.
By the numbers
Under the hood
Encord was founded to help machine-vision companies annotate data and evaluate models. As customers moved to end-to-end learning for robotic manipulation, it realized it had to manufacture data, not just manage it—"the data simply does not exist," says Vineeth Velmurugan, its head of robot learning and an alum of OpenAI's robot lab and Berkshire Grey. It now pulls egocentric video from factories worldwide and runs experimental modalities in San Leandro, where pilots use leader-follower rigs—paired arms, one human-controlled and one that mimics it—to generate data for tasks like pouring coffee and stacking poker chips. Beyond brain waves, Encord is testing forearm sensors that read muscle electrical signals to reconstruct a 3D hand pose that video alone misses. Its data comes densely annotated with physical descriptions like "right hand tightens bolt" to help LLM-based models understand the scene.
The signal
The interesting bet here isn't the headset—it's the premise. If brain activity reveals when a task is hard, model builders get a signal for when to deploy their highest-effort models, and robots could learn not just what people do but roughly why. That's a genuinely new axis for training data. It also drags physical AI into the same privacy questions self-driving already faces, except now the sensor is pointed at your head.
The backstory
This is the same wall the whole field keeps hitting: generative AI did for chatbots what it can't yet do for robots, because text was free to scrape off Stack Overflow and the web, and physical data isn't. Self-driving companies collect real-world data themselves, but it doesn't scale; training from video scales but lacks fidelity. Velmurugan estimates it will take a data set roughly five times the size of YouTube's video corpus to break through—which is why data generation has become a business, not just a research problem.
Who it's for
- Founders
- a robotics bottleneck that's now a business
- Researchers
- new training signals: intent, error, muscle pose
- Enterprises
- why warehouse automation is still costly
The catch
This is a trial run, explicitly framed as the "bleeding edge," with no results yet showing brain-wave tags improve model performance. The economics are the real catch: dense annotation is worth about 100x junky ego data but costs 20x more to produce—a good trade on paper, but 20x is still real money that LLM makers never had to spend. And the hardware humbles you fast—TechCrunch's writer found robotic pincers far less dexterous than human fingers, which is why tasks like unplugging ethernet cables from a server remain out of reach.
What to watch
Watch whether Encord's brain-wave data set clears the bar and moves from trial to scaled product. The bigger open question is economic: whether manufactured physical data can ever get cheap enough to close a YouTube-times-five gap, and which modalities—brain waves, muscle sensors, dense annotation—actually earn their cost.
Source-backed · official sources first, ecosystem reporting labelled