AI agents require more than good models. They require authority engineering.

I see that we are attributing to the model capabilities that, in practice, belong to the entire architecture of the agent. This distinction becomes particularly important when AI systems stop merely responding and start querying data, using tools, and executing actions.
To understand this architecture, we need to separate model, workflow, agent, harness, tools, and execution environment. These boundaries are not universally standardized and vary between implementations, but the conceptual separation helps us understand where capability, execution, and control reside.
The model interprets the available context and produces inferences. It can generate text, classify information, propose next steps, and select or request tools. On its own, this does not necessarily turn it into an agent.
The difference between workflow and agent lies mainly in where the logic that determines the next steps resides. In a workflow, the path is predominantly defined by the software. There can be conditions, branches, and decisions based on an LLM's responses, but the overall structure was established by the developers.
In more agentic systems, part of that choice is delegated to the model. Based on the available state, it can select tools, interpret results, and determine next steps within the limits defined by the system.
These are not mutually exclusive categories. A workflow can incorporate agents, and an agent can operate within larger processes, with deterministic steps and predefined limits. In practice, many enterprise architectures will be hybrid.
The harness is a useful way to describe the software around the model that connects it to the other components. It can organize context, state, memory when necessary, tools, and execution cycles. It can also implement or trigger policies, checks, observability, and controls. But there is no universal definition determining that all these functions necessarily belong to the harness.
Imagine asking an agent to analyze a contract, identify risks, and compare its clauses with internal policies. The model may request the contract. The system forwards the request to the appropriate tool, within existing permissions. The document comes back as context. The model can then request another source, compare information, and propose the next step.
A cycle forms: inference, proposed action, execution, observation, and new inference. This is where a fundamental distinction emerges. Knowing how to do something, being able to do it, and being authorized to do it are different things.
The model can produce a request to delete a record. A tool may be technically capable of executing it. That does not mean the operation should be allowed.
Authorization needs to depend on the identity used, on who delegated that authority, on the permissions granted, on the resource involved, on the context of the operation, and on the applicable policies.
This raises a question frequently neglected in agentic architectures. On whose behalf is the agent acting?
Agents that execute actions need manageable identities and credentials, minimal permissions, and, in certain processes, segregation of duties. The technical capacity to execute an operation should not define the authority to execute it.
There is also the execution environment, or runtime, in which the software operates. It is there, together with other layers of infrastructure and control, that many concrete boundaries for access to systems, networks, and resources take shape.
When agents execute real actions, classic engineering problems reappear. Imagine a tool processing a refund, but the connection drops before confirming the result. If the system simply retries the operation, it may end up issuing two refunds.
The architecture needs to know how to check prior state, handle failures and retries, set time limits, recover operations, and, when possible, compensate for or reverse actions. Agents have not eliminated the classic problems of distributed systems. They have added probabilistic components to them.
Security also changes in scale. A contract, email, page, or retrieved document may contain malicious instructions intended to influence the model. External content should be treated as potentially untrusted data, not as an automatic source of authority.
Information is not instruction. And instruction is not authorization. That is why telling the model what it should not do is not the same as technically preventing a given action from being executed.
Critical controls also need to exist at the execution boundaries, through identity, permissions, deterministic validations, isolation, operational limits, approvals, and interruption mechanisms proportional to the risk.
There is yet another essential component. Observability. When an agent acts, we need to be able to reconstruct what happened. What goal it received, what relevant context it used, which tools it called, which actions it requested, under which identity it operated, which policies were applied, which approvals occurred, and what the outcome was. Without this, investigating errors and assigning responsibility becomes much harder.
It is also not enough to evaluate the system just once. Changes to the model, prompts, tools, data, policies, or integrations can alter its behavior. Enterprise agents need evaluations before deployment, monitoring in production, and clear criteria for expanding, reducing, or withdrawing autonomy. This explains why improving the model does not automatically mean improving the agent.
An enterprise system also depends on reliable data, integrations, identity, state, observability, failure handling, policies, and controls. Even swapping just the model requires reassessing the behavior of the whole, especially tool selection and use. An excellent model does not, by itself, make up for a deficient architecture.
The final distinction can be summarized like this. The model provides inference capability. The agent is the system configured to execute goal-oriented tasks using that capability. The harness coordinates part of the interaction between the model and the other components. The tools offer query or action capabilities. The runtime provides the environment in which the software executes.
Authority needs to be defined and enforced by mechanisms external to the mere intent expressed in the prompt. When the system only responds, much of the evaluation lies in the quality of the response. When it starts to act, a different discipline emerges. We need to determine who can do what, on whose behalf, over which resources, under which circumstances, and within which limits.
The greater the operational autonomy granted, the greater must be the capacity to observe, limit, interrupt, and audit its actions. The challenge, therefore, is not just building more capable agents. It is developing a true engineering of authority so that systems with probabilistic components can act in the real world without turning technical capability into unrestricted authority.
Translated from the Brazilian Portuguese original · Read the original
Does your site exist for AI? Start with the terminal
A site can work normally in the browser and still be inaccessible to AI crawlers. This article presents terminal tests to identify blocks by firewall or CDN and to check whether the content is in the HTML sent by the server. It also explains the order of Technical GEO checks: access, reading, interpretation, and measurement.



