AIARTICLE

Standard Bots separates AI from deterministic logic to gain reliability

Industrial robot maker Standard Bots treats the AI model as one piece of the system, not the whole system. For those building software agents, the architecture lesson matters as much as the data one.

Standard Bots, an American maker of industrial robotic arms used by NASA, Amazon, and Lockheed Martin, recently raised $200 million in a Series C round led by General Catalyst and RoboStrategy, reaching a $1 billion valuation. The company describes itself as the largest AI-native industrial robot maker in the United States. In an interview with the Latent Space podcast, CEO Evan Beard and Head of AI Leif Jentoft detailed how the company's AI stack works under the hood when tackling tasks like machine tending, welding, and assembly.

What's interesting for developers isn't the robot itself but the architecture: Standard Bots doesn't let one giant model decide everything. It uses a relatively small model (according to Beard, "in the low billions of parameters," far from frontier models) for a well-defined function, and surrounds that function with conventional code. It's the kind of design decision that any team building AI agents in production in Brazil should be discussing right now, instead of stacking prompt on prompt until the system turns into an unpredictable black box.

The model doesn't need to do everything

For part swapping (machine tending) in high-mix manufacturing, Standard Bots built what it calls a zero-shot perception system: the user tells the robot which parts to look for in a task, and the model locates and identifies those parts. But arm movement and work-cell logic remain with traditional code, with no learning involved at all.

This division of labor is the core of the lesson. Jentoft was direct about why:

Despite claims elsewhere, no model today is truly hardware-agnostic, and co-optimizing low-level control and higher-level functions is a major advantage for both performance and iteration speed.

Leif Jentoft, Head of AI at Standard Bots

Translating for those building software agents: a generalist LLM calling tools with no control structure around it tends to be more fragile than a pipeline where the model handles only the part that requires judgment (understanding intent, extracting entities, classifying) and the rest runs on deterministic, testable, versioned code. This isn't a robotics fad, it's the same argument behind using state graphs (LangGraph, explicit state machines) instead of letting the agent decide its own flow at every step.

Good data beats more data

Standard Bots' base model is trained on more than a billion images, which, according to Jentoft, is what lets the robot tell apart lighting conditions, material types, and an object from the background itself. But the company doesn't rely on volume alone to solve edge cases: it uses corrections made in real-world deployments.

We believe data quality matters far more than raw volume, and we focus on getting the most out of targeted data rather than chasing the largest possible dataset.

Leif Jentoft, Head of AI at Standard Bots

Jentoft gave a concrete number: in-situ interventions can fix an edge case with just a few dozen examples. For those doing fine-tuning or curating small example sets for RAG and agents, this is one more data point in favor of building a fast-correction pipeline (few examples, high precision) instead of waiting for the next giant dataset to solve the problem on its own. The Latent Space piece also notes that this principle echoes what Diogo Almeida, creator of Jev, recently called the "bitterer lesson": the right task and the right data outweigh raw compute.

Why inference runs at the edge, not in the cloud

Training for Standard Bots' models happens in the cloud, but inference runs locally, inside the factory. Beard explained that this is a deliberate decision, not a technical limitation: most factories and warehouses don't have reliable internet, and for mobile robots that problem is even bigger. Keeping the decision loop local guarantees uptime, which is what determines whether the customer accepts the system or not.

For engineers building AI agents into a SaaS product, internet dependency isn't literally the same problem. But the design principle carries over: every point where the agent depends on an external network call (a model API, a third-party service, a distributed queue) is a failure point that needs fallback, timeout, and graceful degradation defined before going to production, not discovered during an incident. Standard Bots solved this by pushing the model to the edge; those building software agents solve it with circuit breakers, response caching, and deterministic fallback paths for when the model doesn't respond in time.

Production corrections become data, not just error

Not every task is easy to simulate: Beard cites liquids, suction, and cutting flexible material as cases that are hard to reproduce in current simulators. That's why the company relies on real-world demonstrations and on failure signals captured in the field.

Where that data ends up varies by customer. Defense customers, according to Jentoft, usually operate in air-gapped environments, where nothing flows back to the central model. Customers outside that regime, on the other hand, tend to accept contributing data in exchange for the performance gains of what's called "fleet learning" (shared learning across the robot fleet). For AI agent teams, the direct parallel is the feedback loop: every human correction made to an agent's output (a user editing a response, a reviewer flagging an error) is potential training data, and it's worth instrumenting this from the first deployment, not after the complaints backlog has already turned into noise.

What this changes for those building in Brazil

Standard Bots has also opened part of its stack to outside developers through StandardOS, a set of APIs and SDKs that lets teams assemble robotics applications using whichever pieces they want, including bringing their own model. Beard cited NVIDIA Cosmos, a family of world models for physical AI that launched version 3 in late May 2026, as an example of a compatible external model. Today this still requires manually written integration code; the company plans to simplify data collection and the deployment of custom models in the future.

It's worth contrasting this stance with that of other robotics companies, such as Skild, which pursues generalization across different hardware as a central goal. Jentoft's claim that "no model is truly hardware-agnostic" is, in practice, a direct disagreement with that thesis. For those evaluating agent frameworks, the warning applies: be wary of any tool that promises to be pluggable into any model, any stack, and any use case with no performance loss whatsoever.

In short, the lesson to take from Standard Bots isn't about robots: it's about reliability engineering. Keep the model restricted to a short-horizon, well-defined task, surround it with testable deterministic code, prioritize data quality over volume, and treat every production correction as input, not as noise to be ignored.

This approach doesn't fit every project: for a disposable prototype or a low-volume agent, the cost of building that correction pipeline outweighs the gain. But for any agent that's going to run in real production, deciding things that cost money when they go wrong, it's exactly this kind of discipline that separates a reliable system from a pretty demo.

Translated from the Brazilian Portuguese original · Read the original

View profile →