Back to Blog
Deployment
2026-09-04
6 min read

The models are good enough. The plumbing is the product.

Dev Sanghvi
Founder & CEO, DHI

An integrator, not a lab

On 3 September 2026, Ambarella, an edge AI silicon vendor that has shipped more than 50 million AI SoC units, announced a joint venture with Capgemini: a joint "Edge and Physical AI" Centre of Excellence. Capgemini is not a research house. It is a systems integrator with 420,000 employees and 22.5 billion euros in 2025 revenue, and its role in this partnership is to get physical AI installed across smart infrastructure, retail, logistics, industrial automation, healthcare and automotive sites.

Ambarella's chief executive, Fermi Wang, described the opportunity plainly: "Edge and Physical AI represent an important opportunity to extend AI from the data center into cameras, robots, vehicles, machines." That is a sentence about extending AI outward, not about making it smarter. Nobody in that announcement is promising a better model. They are promising to get the existing ones onto more machines.

That distinction is worth sitting with. If the binding constraint on physical AI adoption were model accuracy, the natural move would be to hire more machine learning researchers and chase a better benchmark score. Ambarella did not do that. It partnered with a company whose core competency is showing up at hundreds of thousands of client sites and making complicated systems talk to each other. That is a tell. When a hardware vendor's growth plan runs through a 420,000 person integrator rather than through its own research team, the constraint it is solving for is not model quality. It is deployment.

What actually breaks at a site

We say this from the inside, because it is the same wall we run into on every DHI installation, and it rarely has anything to do with detection accuracy.

It is an undocumented VLAN that nobody wrote down when the network was built, so the camera subnet cannot reach the box that needs it without someone tracing cable runs by hand. It is the network engineer who set that VLAN up leaving the company eighteen months ago, taking the mental map of the site with them. It is a mounting point with no spare power circuit, so a camera that could go up in an afternoon waits three weeks for an electrician. It is a change control window that opens one Saturday every quarter, so a five minute firewall rule change becomes a scheduling problem. It is an RTSP stream that only behaves on one specific codec path, discovered only after the second or third failed connection attempt. It is a camera installed years ago whose admin credentials exist nowhere anyone can find them.

None of this is exotic. It is the ordinary condition of physical infrastructure that was built for a different purpose, by a different set of contractors, documented to a different standard, over a different number of years. A detection model does not touch any of it. A model can be capable and still sit unused because nobody could get a clean video stream to it in the first place.

Why DHI starts with one camera

Our answer to this is not a magic integration layer that makes all of the above disappear. We do not have one, and we are not aware of anyone who does. Our answer is to shrink the problem: install on one camera, run it for a week, and only then talk about the second camera.

We want to be honest about what that is and is not. It is not a solution to the integration problem described above. The undocumented VLAN is still undocumented. The missing credentials are still missing. Starting small does not make a site's infrastructure debt go away. What it does is bound the damage from any one piece of that debt. If the first camera's RTSP stream turns out to only work on a codec nobody expected, we find that out on one camera, in one week, not across a twelve camera rollout planned around an assumption that turned out to be wrong. If the person who understood the VLAN is gone, we discover that gap while the cost of discovering it is one delayed camera, not a stalled programme.

This is a containment strategy, not a fix. Every site we have not yet touched still has its own undocumented VLAN and its own missing credentials waiting to be found. Scaling from one camera to ten does not average that risk away, it just means finding out about problems one at a time instead of all at once. We would rather say that plainly than imply we have solved deployment, because the moment a customer believes their integration risk is handled, they stop budgeting time for it, and that is when a rollout stalls.

The tell in who gets hired

Back to Ambarella and Capgemini. A Centre of Excellence aimed at smart infrastructure, retail, logistics, industrial automation, healthcare and automotive is, by its own description, a deployment programme spanning wildly different network topologies, procurement cycles, compliance regimes and physical environments. That breadth is the point. Capgemini's value is not that its engineers will make Ambarella's silicon detect objects better. It is that Capgemini has done enough site visits, across enough industries, to have a playbook for the boring failure modes: the VLAN, the missing spare circuit, the change window, the lost credentials.

We take that as a confirmation of something we already believed from running our own deployments: the hard part of physical AI was never mostly the model. Detection and classification models crossed a usability threshold for most straightforward warehouse and facility use cases some time ago. What has not crossed any threshold is the average site's network documentation, spare capacity, change management process and institutional memory. Those do not improve with a better checkpoint. They improve with someone doing the unglamorous work of finding out, camera by camera, site by site, what is actually plugged into what.

So when a vendor leads with an accuracy number, a benchmark score or a headline detection rate, we read that as a claim about a part of the stack that stopped being the bottleneck a while ago. The number that would actually tell you something is how many cameras a vendor got live on the last ten real sites, how long it took, and what broke along the way. That number is harder to make impressive, which is probably why it gets published less often. We would rather be judged on it anyway.

DeploymentSystems IntegrationEdge AINetwork Infrastructure