AIARTICLE

OpenAI Agents SDK: how to build handoffs and guardrails between agents before going to production

The OpenAI Agents SDK offers three building blocks (agents, handoffs, and guardrails) to orchestrate teams of LLMs in Python. The tutorial shows how to put them together and where the architecture tends to break before production.

The OpenAI Agents SDK is the Python library that OpenAI maintains as the successor to Swarm, the experiment the company used internally to test multi-agent orchestration patterns. The difference stated in the project's official documentation is one of positioning: Swarm was experimental, while the Agents SDK is described as "production-ready." In practice, this means a small set of primitives (agents, handoffs, guardrails) rather than a stack of abstractions to learn before writing the first line of code.

The package is installed with pip install openai-agents and, by default, uses OpenAI's Responses API under the hood, but wrapped in a runtime that handles turns, tool calls, and safety checks without the developer having to write that loop by hand. It is this runtime, not the model itself, that the documentation calls the "agent loop": it keeps execution going until the task is considered complete, exchanging messages, calling tools, and, if configured, handing off to another agent along the way.

The minimal agent and what the Runner does under the hood

The documentation's "hello world" example fits in four lines and already exposes the central idea: an Agent is just an LLM with instructions and, optionally, tools; the one that runs it is the Runner.

python
from agents import Agent, Runner

agent = Agent(name="Assistant", instructions="You are a helpful assistant")
result = Runner.run_sync(agent, "Write a haiku about recursion in programming.")
print(result.final_output)

With OPENAI_API_KEY set in the environment, this code already runs locally, with no server and no message queue. The point is that Runner.run_sync hides the tool-calling loop: if the agent had a function tool attached, the SDK would decide on its own when to call it, read the return value, and decide whether the task is already done or needs another turn. This is convenient until the moment something goes wrong inside that loop and the developer has no visibility into what happened, which pushes the SDK's built-in tracing from "nice-to-have feature" to a debugging prerequisite.

Handoffs: when one agent passes the ball to another

The building block that lends its name to this piece's "team of agents" is the handoff. The documentation describes handoffs as agents treated as tools of other agents, a mechanism to "coordinate and delegate work across multiple agents." In practice, this solves a real problem: a single giant system prompt trying to cover triage, technical support, and billing tends to confuse the model about which set of instructions applies to each message.

The recommended pattern separates this into specialized agents with a triage agent in front:

python
from agents import Agent, Runner

billing_agent = Agent(
    name="Billing agent",
    instructions="Resolve billing questions. Be precise about amounts and dates.",
)

tech_agent = Agent(
    name="Tech support agent",
    instructions="Diagnose technical issues. Ask for logs when relevant.",
)

triage_agent = Agent(
    name="Triage agent",
    instructions="Decide if the user needs billing or technical help, and hand off accordingly.",
    handoffs=[billing_agent, tech_agent],
)

result = Runner.run_sync(triage_agent, "Minha fatura veio duplicada este mês")
print(result.final_output)

The triage agent doesn't resolve the duplicate invoice: based on its instructions, it decides that the billing_agent should take over the conversation, and the SDK transfers the execution context to that second agent. This is the practical difference between a handoff and simply calling another agent as a tool: in a handoff, ownership of the turn changes hands; in a regular tool call, the original agent stays in command and just uses the other agent's response as input.

In short: a handoff is for fully delegating responsibility; an agent-as-tool is for a one-off query without changing who's at the wheel. Picking the wrong model is one of the most common ways an orchestration gets stuck: a pipeline that should have delegated and only queries ends up with a generic agent trying to do everything poorly, and a pipeline that should have queried but hands off instead loses the thread of the original conversation.

Guardrails: the check that runs in parallel, not afterward

The documentation is specific about how guardrails should behave: "run input validation and safety checks in parallel with agent execution, and fail fast when checks do not pass." This is an architectural choice, not an implementation detail. A guardrail isn't a filter that runs after the response is ready, checking the final text; it runs alongside the main agent, and if it fails, it stops execution before spending more model calls.

A typical input guardrail blocks, for instance, requests that clearly don't belong to the agent's domain (someone trying to use the billing bot to ask for legal advice), and an output guardrail validates whether the model's response complies with a format or policy before it reaches the user. The gain from running in parallel is latency: the guardrail doesn't wait for the agent to finish before it starts checking, and the SDK aborts as soon as the first failure shows up, before the user receives a response that should never have been generated.

Sessions: memory across turns, and why the default option isn't fit for production

The third piece in this topic is context persistence. The documentation defines Sessions as "a persistent memory layer for maintaining working context within an agent loop," and here's the point that separates a local prototype from something that survives a service restart: there's more than one session implementation, and the difference between them is exactly where the orchestration breaks before going to production.

The SDK documents, among others, SQLAlchemySession, Async SQLite session, RedisSession, MongoDBSession, DaprSession, EncryptedSession, and AdvancedSQLiteSession. The existence of this entire list is the signal: an in-process memory session (the simplest option for running locally) disappears every time the service restarts, and any orchestration with multi-step handoffs that depends on remembering what the triage_agent decided two turns ago loses that history on the first deploy that restarts the container.

OpenAI Agents SDK documentation page shows the summary with the available session types, such as SQLite, Redis, SQLAlchemy, Dapr, and MongoDB, along with a code example using SQLiteSession
OpenAI Agents SDK documentation page shows the summary with the available session types, such as SQLite, Redis, SQLAlchemy, Dapr, and MongoDB, along with a code example using SQLiteSession. Screenshot: openai.github.io.

For anyone planning to take this kind of agent team to production, the choice of session backend isn't a late-stage configuration detail, it's a decision that needs to enter the architecture design from the prototype stage, especially if the flow involves compliance (where EncryptedSession comes into play) or multiple service instances running behind a load balancer (where an in-memory session simply isn't an option).

Where things get stuck before reaching production

Putting the three pieces together, the most likely points of friction in a real agent team show up in specific places:

  • Handoff loops with no stopping condition: if agent A hands off to B and B's instructions allow handing back to A, nothing in the SDK prevents an indefinite ping-pong without an explicitly defined turn limit.
  • A guardrail that fails silently: a misconfigured guardrail that never triggers gives a false sense of security; the SDK's built-in tracing is the only reliable way to confirm it's being evaluated on every run.
  • An in-memory session forgotten in production: it works perfectly on the developer's laptop and loses the entire history as soon as two workers of the service handle the same user in different requests.
  • Mistaking a handoff for a tool call: using a handoff for a simple one-off query makes the main agent needlessly lose control of the conversation.

Agents SDK or the Responses API directly: the question that comes before the code

The documentation itself recommends not treating this as an either-or choice: "many applications use the SDK for managed workflows and call the Responses API directly for low-level paths." Plain Responses API is worth it when the flow is short, a single turn, and the team wants to explicitly control tool dispatch and state. The Agents SDK is worth it when the runtime needs to manage turns, guardrails, handoffs, or sessions on its own, or when the agent needs to operate across multiple coordinated steps, like the triage, billing, and technical support team described above.

For anyone deciding whether to migrate a hand-built flow of chained prompts to the SDK, the safest path is to start with the single agent from the quickstart, add a handoff only once there are two clearly distinct instruction domains, and only then introduce a session with a persistent backend, in that order and not the other way around.

Translated from the Brazilian Portuguese original · Read the original

View profile →