If you are an engineer, this is the one to stay for. Prerna Dhareshwar of Voxel51 returns with Nathan Maroney to go a layer deeper than Episode 12, into what actually runs inside a physical AI system. The short version: these models predict the next action the way a language model predicts the next token, and that one architectural fact quietly dissolved a problem the industry spent years solving in hardware.

Prerna works on the product side at Voxel51, where she led the platform that physical AI teams use to explore, visualize, curate and query their data. She came to it from the other end of the problem: an engineering degree from IIT Madras, a Master's from Stanford, research at India's National Aerospace Laboratories, predictive analytics at Pure Storage, and vision based anomaly detection for manufacturing at Instrumental. She has seen this from the model side, the data side, and the factory floor.

Start with VLMs. They are the large language models you already know, trained on images and video alongside text, so they carry an understanding of what they are looking at. Then VLA models, Vision-Language-Action. They take every sensor input, take a language prompt describing the goal, and output action tokens continuously, each one conditioned on what the sensors are saying at that instant.

Prerna walks it through with a robotic arm unloading a dishwasher. At timestamp zero it sees the dishes and decides its next move is to reach for a plate. That changes the inputs. Now the next action is to grasp. And so on. The model was never trained on your dishwasher, or Nathan's, and it does not need to be.

The consequence is the most contested claim in the episode. Because the model conditions on whatever it has at each moment, the sensor streams do not need to be time synchronized. In her words, alignment of sensors is not really something people are too worried about anymore.

Then trust. Titto puts the noise problem to her using his own house. A spotless kitchen is one thing. A dish sitting on the roof is another. Real environments are not controlled, and mathematically the difference is just noise. Her answer runs through post-training, the same human alignment step that makes language models sycophantic, applied instead to a human critiquing each action a robot takes. In autonomous driving that is the gap between a car that is safe and a car that behaves the way other drivers expect. Early Waymos followed the road rules exactly and got rear ended by humans who do not.

Nathan names the failure mode nobody wants to discuss. Industrial pilots that succeed technically and fail commercially. Heavy industry generates enormous volumes of sensor data, but identifying the small slice worth training on takes a specialist team, and operational leaders have KPIs tied to throughput rather than to technical change. Resistance is not ignorance, it is incentives.

The episode closes on Voxel51's platform, why edge cases and the long tail decide the last fraction of a percent, and why Prerna is bullish on physical AI while still calling it early.

CHAPTERS

00:00 What is a VLM, and why it matters
 01:36 Multimodality and where the models are heading
 03:52 Sensor coverage across a mine the size of a city
 04:55 Why operational KPIs block adoption
 06:29 Selling a model to a board
 07:55 What happens when sensors are not time synced
 08:31 VLA models: next action prediction explained
 09:29 The dishwasher, step by step
 11:23 Why sensor fusion stopped being the problem
 13:10 The world's first fully autonomous rig
 13:39 Noise, uncontrolled environments, and the dish on the roof
 16:51 Waymo, road rules, and getting rear ended
 17:20 Industry 5.0 and human in the loop
 18:12 Post-training, sycophancy, and human alignment
 20:20 Manufacturing: from defect detection to assembly
 22:25 Pilots that succeed technically and fail commercially
 24:44 Inside Voxel51's platform
 26:10 Bullish, but early
 27:05 Defense, swarms, and a higher bar
 30:22 Edge cases, the long tail, and the last 0.9%

GUESTS

Prerna Dhareshwar, Voxel51
 https://www.linkedin.com/in/prernamd/

Nathan Maroney, Director, Tryfecta Group

Host: Titto Thomas, Managing Partner and Co-Founder, Tryfecta Group

LINKS

Watch on YouTube: https://youtu.be/AFuHb-vLbyo
Voxel51: https://voxel51.com
Tryfecta: https://tryfecta.biz

Podden och tillhörande omslagsbild på den här sidan tillhör Titto Thomas. Innehållet i podden är skapat av Titto Thomas och inte av, eller tillsammans med, Poddtoppen.