Dev & EngARTICLE

Your Agent Isn't as Autonomous as It Seems

Your Agent Isn't as Autonomous as It Seems
Image: Cezar Taurion

The more I work on AI projects and discuss agents with companies, the more I question a simplification that is starting to gain ground: attributing to the model all the autonomy we see in an agentic system.

The agent is not the LLM. And confusing the two can lead to poor decisions in architecture, security, and governance.

When I see an impressive demonstration of an “autonomous agent,” my attention today goes less to the model in isolation and more to everything that allows it to operate over time.

Persistent memory, state management, retrieval, tools, identity, permissions, policies, observability, checks, and interruption mechanisms are part of the system's architecture. The model plays a central role. It can interpret context, propose plans, select tools, and dynamically influence next steps. This is precisely what differentiates many agents from rigidly predefined workflows.

But this doesn't mean that all agency lies within the model. In a typical architecture, some infrastructure maintains or retrieves state, provides context, makes tools available, executes actions, returns their results to the model, and decides when the cycle should continue, be interrupted, or require approval.

So when we add persistent memory, it doesn't necessarily mean we're changing the model's weights or that it has “learned to remember” like a person. We're often creating external mechanisms to store, select, summarize, and reinsert information into the context.

When we connect tools, we greatly expand what the system can do. But concrete action still depends on software, APIs, credentials, permissions, and execution environments.

And an agentic loop should not automatically be confused with a kind of persistent mental reflection. We have successive inferences, actions, and observations coordinated by an architecture.

This is exactly where I find the interesting paradox: the more operational autonomy we grant an agent, the more control engineering we need to build around that autonomy.

We add retrieval to search for information. Memory to preserve relevant context. Verifiers because outputs can be wrong. Guardrails because certain actions must not occur. Checkpoints and approvals because some mistakes have bigger consequences. Observability to reconstruct what happened. Execution limits and interruption mechanisms because loops need to end.

We are not eliminating engineering by creating agents. We are creating a new layer of engineering to make probabilistic systems capable of acting in a manageable way.

That's why, in the projects I'm involved in, I've insisted on not governing only the model. We need to govern the entire system. Model, instructions, context, memory, data, tools, identity, credentials, permissions, orchestration, verifiers, policies, and execution limits are all part of the risk surface.

It's in this chain that a probabilistic output can turn into an action with real consequences.

I also try to avoid excessive anthropomorphizing. Saying that the agent “remembered,” “decided,” “noticed,” or “wanted to continue” can be a convenient linguistic shortcut. The problem starts when the shortcut begins to replace the explanation of the architecture.

My sense is that we are rediscovering, with much more powerful components, an old idea from computing and from AI itself: sophisticated capabilities can emerge from the composition of specialized components, and not necessarily from a monolithic intelligence.

The LLM is extraordinarily important. But it is not the entire system. That's why, when faced with an “autonomous agent,” I want to understand much more than just the model being used.

I want to know who maintains the state, which tools can be triggered, with which credentials, who grants the permissions, what is verified before an action, where approvals exist, how we reconstruct an execution, and who can interrupt it.

To me, it's clear that the more autonomous an agent seems on the surface, the more control engineering may be needed underneath it. And maturity in AI begins precisely when we stop asking only “which model are you using?” and start examining the system that turns probabilistic inferences into real-world actions.

Translated from the Brazilian Portuguese original · Read the original

More from Cezar Taurion
View profile →