If you are an engineer, this is the one to stay for. Prerna Dhareshwar of Voxel51 returns with Nathan Maroney to go a layer deeper than Episode 12, into what actually runs inside a physical AI system. The short version: these models predict the next action the way a language model predicts the next token, and that one architectural fact quietly dissolved a problem the industry spent years solving in hardware.
Prerna works on the product side at Voxel51, where she led the platform that physical AI teams use to explore, visualize, curate and query their data. She came to it from the other end of the problem: an engineering degree from IIT Madras, a Master's from Stanford, research at India's National Aerospace Laboratories, predictive analytics at Pure Storage, and vision based anomaly detection for manufacturing at Instrumental. She has seen this from the model side, the data side, and the factory floor.
Start with VLMs. They are the large language models you already know, trained on images and video alongside text, so they carry an understanding of what they are looking at. Then VLA models, Vision-Language-Action. They take every sensor input, take a language prompt describing the goal, and output action tokens continuously, each one conditioned on what the sensors are saying at that instant.
Prerna walks it through with a robotic arm unloading a dishwasher. At timestamp zero it sees the dishes and decides its next move is to reach for a plate. That changes the inputs. Now the next action is to grasp. And so on. The model was never trained on your dishwasher, or Nathan's, and it does not need to be.
The consequence is the most contested claim in the episode. Because the model conditions on whatever it has at each moment, the sensor streams do not need to be time synchronized. In her words, alignment of sensors is not really something people are too worried about anymore.
Then trust. Titto puts the noise problem to her using his own house. A spotless kitchen is one thing. A dish sitting on the roof is another. Real environments are not controlled, and mathematically the difference is just noise. Her answer runs through post-training, the same human alignment step that makes language models sycophantic, applied instead to a human critiquing each action a robot takes. In autonomous driving that is the gap between a car that is safe and a car that behaves the way other drivers expect. Early Waymos followed the road rules exactly and got rear ended by humans who do not.
Nathan names the failure mode nobody wants to discuss. Industrial pilots that succeed technically and fail commercially. Heavy industry generates enormous volumes of sensor data, but identifying the small slice worth training on takes a specialist team, and operational leaders have KPIs tied to throughput rather than to technical change. Resistance is not ignorance, it is incentives.
The episode closes on Voxel51's platform, why edge cases and the long tail decide the last fraction of a percent, and why Prerna is bullish on physical AI while still calling it early.
CHAPTERS
00:00 What is a VLM, and why it matters
01:36 Multimodality and where the models are heading
03:52 Sensor coverage across a mine the size of a city
04:55 Why operational KPIs block adoption
06:29 Selling a model to a board
07:55 What happens when sensors are not time synced
08:31 VLA models: next action prediction explained
09:29 The dishwasher, step by step
11:23 Why sensor fusion stopped being the problem
13:10 The world's first fully autonomous rig
13:39 Noise, uncontrolled environments, and the dish on the roof
16:51 Waymo, road rules, and getting rear ended
17:20 Industry 5.0 and human in the loop
18:12 Post-training, sycophancy, and human alignment
20:20 Manufacturing: from defect detection to assembly
22:25 Pilots that succeed technically and fail commercially
24:44 Inside Voxel51's platform
26:10 Bullish, but early
27:05 Defense, swarms, and a higher bar
30:22 Edge cases, the long tail, and the last 0.9%
GUESTS
Prerna Dhareshwar, Voxel51
https://www.linkedin.com/in/prernamd/
Nathan Maroney, Director, Tryfecta Group
Host: Titto Thomas, Managing Partner and Co-Founder, Tryfecta Group
LINKS
Watch on YouTube: https://youtu.be/AFuHb-vLbyo
Voxel51: https://voxel51.com
Tryfecta: https://tryfecta.biz