Computer vision has been around for decades. So why is physical AI suddenly everywhere, with a new humanoid robotics company funded almost every week? Prerna Dhareshwar of Voxel51 and Nathan Maroney of Tryfecta Group join Titto Thomas to explain what changed, why vision alone will never be enough, and why the real differentiator is the software and data layer that almost nobody is talking about.
Prerna traces the shift back to what large language models taught the field about generalization. Language modeling is next token prediction. Physical AI is next action prediction. Once it became clear the same architectures could work in both, capital and attention followed.
Prerna is on the product team at Voxel51, where she led a product that helps physical AI teams explore, visualize, curate, and query their data. Before that she was a machine learning engineer building vision based anomaly detection for manufacturing at Instrumental. Nathan Maroney, Director at Tryfecta Group and the group's mining lead, brings the operator view from mining and heavy industry, where you are processing hundreds of thousands of tons and a small consistency gain in concentration is worth real money.
The conversation gets practical fast. Nathan asks the question every operator is actually asking: how does a mine site get itself ready for this, and how is physical AI different from the conventional automation, machine vision, and predictive maintenance they already have? Prerna's answer is to look sideways. Auto manufacturing OEMs are already using humanoid robots to build and assemble parts, in exactly the complex three dimensional work that robotic arms could never automate away. Find what worked in an adjacent industry, then lift and generalize it. That approach also de risks the first move.
On whether physical AI is just self driving cars, Prerna points out that Waymo has had more than a ten year head start collecting data, which is why that sector looks a few steps ahead of everyone else rather than being the whole story.
There is an honest detour into risk. In San Francisco, Prerna keeps seeing ads for humanoid robots that will clean your house for a flat fee of $150 regardless of size. Great deal, until it leaves you with a pile of broken dishes. Titto's line for how heavy industry thinks about that: it is not the early bird that gets the worm, it is the second mouse that gets the cheese. Prerna's view is that the calculus has genuinely changed, and that she would have answered differently a year or two ago.
Then the technical core. Physical AI data is hard because the sensors are not time synchronized. Each one records at its own frame rate and frequency, across long time ranges, and all of it has to be aligned, curated, labeled, and fed downstream before a model ever sees it. That tooling problem is where a lot of the real work lives.
On why multimodal beats vision only, Prerna uses the human analogy. We do not perceive the world through sight alone. Two eyes give us depth. Touch tells us how much pressure a delicate object can take. Restrict a system to a single camera and you have severely limited what it can know about the world it is operating in. She takes a position on the vision only versus lidar debate, citing an edge case where a Waymo could not see a person crossing from behind a parked truck, and the lidar caught what the cameras could not.
Titto brings a story from the Apache gunship program, where the sensor lens was cut from pure sapphire because the software of the 1980s could not correct for chromatic aberration. Everything had to be fixed at the physical layer to hand the software a clean image. Today a ten dollar sensor does the same job, which is exactly the hardware to software shift the field is living through. With VLMs, the sensors are fixed and the goal is fixed as a text prompt. What the system controls is the sequence of actions it takes to get there.
Nathan closes on the practical blocker. Simulating a controlled separation process is achievable low hanging fruit, and most of the sensors are already installed, but a site typically has to wait two years to accumulate enough data before it can deploy. He wants systems that are self learning, self healing, and self governing from day one on a greenfield site. Prerna's answer is that it is never too early to start collecting data, and that even if you are years away from deploying anything, making sure it is clean, structured, and stored properly is the move available to you right now. She connects the self learning ambition to meta learning and few shot generalization, and to what today's models already do in a limited way through context.
Prerna is coming back for a deeper technical episode on VLMs and how to deal with noise.
CHAPTERS
00:00 Welcome and introducing Prerna Dhareshwar
01:06 Computer vision is not new, so what actually changed
01:22 Next token prediction to next action prediction
03:55 Nathan on mining: where the margin really sits
04:57 Lifting proven applications from adjacent industries
07:37 Is physical AI just self driving cars
09:08 The $150 humanoid and the broken dishes problem
10:33 Are we there yet, and how much should we fear it
11:56 Convincing conservative operators to move
13:46 Why software is the real differentiator
15:53 Multimodal versus vision alone
17:51 Vision only or lidar, and the Waymo edge case
18:55 Trust, and the Apache gunship sapphire lens
21:09 From hardware to software: what VLMs changed
23:10 Process simulation and the two year data wait
25:31 Start collecting data now, and the path to self learning
28:19 Next time: VLMs and filtering out noise
GUESTS
Prerna Dhareshwar, Product, Voxel51
https://www.linkedin.com/in/prernamd/
Nathan Maroney, Director, Tryfecta Group
Host: Titto Thomas, Managing Partner and Co-Founder, Tryfecta Group
LINKS
Watch on YouTube: https://www.youtube.com/watch?v=o7Rx3sNpgWM
Voxel51: https://voxel51.com
Tryfecta: https://tryfecta.biz