Spotify details multi-agent architecture that already generates most of its ad creatives
At QCon AI, a Spotify Ads engineer showed how the Ads AI platform uses specialized agents, built with Google ADK Java, to generate scripts and campaign targeting for ads in production.
Spotify Ads Manager today runs a multi-agent system in production that already accounts for most of the platform's ad creative generation. Pratik Rasam, a senior engineer at Spotify Ads, presented the architecture at the QCon AI conference, in a talk recorded by InfoQ. The angle matters to anyone designing systems with multiple LLM agents: it is not a talk about the technology's potential, it is an account of architecture decisions tested with real traffic.
What Ads AI is, in numbers
The platform, called Ads AI, launched in 2025. According to Rasam, more than 70% of the ads currently served through Spotify Ads Manager use one of the platform's AI tools. It has already generated around 20,000 creative pieces for more than 7,000 advertisers.
In short: this is not a research prototype. It is real money, real audience and campaigns that actually run on Spotify's ad network, with most of the creative-creation pipeline passing through the agents described in the talk.
One agent, one package, one owner
The first structural pattern Spotify adopted is organizational before it is technical: each agent lives in its own package, with a single owning team. In practice, this means that even before writing the first line of instruction for the agent, a Bazel package already exists with:
lib-info, which defines which team is responsible for the agent and where pull request reviews go when an attribute changes;monitoring-info.yaml, with dashboards and PagerDuty alerts already configured;- dependency-injection scaffolding classes (
AgentFactory,AgentAdapter,AgentService,AgentModule) that make up the package.
The configuration of the language model used by the agent is stored as front matter inside the package itself, which keeps the agent's definition auditable alongside the code.
Boundaries enforced at build time
The most unusual part of the pattern is how Spotify prevents coupling between agents owned by different teams: by using Bazel visibility as the control mechanism. A script-generation agent can only call the audience-recommendation agent if that second agent explicitly allows the import. Without that permission, the build breaks, with no need for manual review or informal convention between teams.
Where the agent ends and deterministic code begins
The central question Rasam poses to anyone designing multi-agent systems is: who owns what? Spotify's answer splits responsibilities into three layers:
| Layer | Responsibility | Example in Ads AI |
|---|---|---|
| Agent (reasoning) | LLM-based judgment, natural-language interpretation | Extracting audience intent, identifying sensitive themes that violate policy |
| Tool/API | Verified data, no hallucination | Interest segments, geographic targeting (DMA) coming from Spotify's own APIs |
| Application code | Anything that can be deterministic | Business-rule validation, checks that don't depend on judgment |
"Anything that can be coded deterministically should always be coded deterministically," Rasam summarized in the presentation. The execution architecture follows the same logic of separation: client applications call a gRPC service responsible for session management, which triggers the AI agent layer, which in turn relies on a model-and-tools layer (LLM runtime, plugins and observability). Orchestration and safety guardrails sit outside the agent, in a separate layer, precisely to ensure consistent behavior across all of the platform's agents.
Why multi-agent became architecture, not just a research demo
Rasam places the turning point in four shifts, three on the technology-supply side and one on the demand side. On the supply side: language models became consistent enough to return structured JSON without improvised parsing; frameworks such as Google ADK Java, LangChain4j and CrewAI started offering ready-made primitives for parallel execution and session management, which each team previously had to write on its own; and OTel GenAI semantic conventions standardized observability for generative AI workloads, making it possible to capture and replay production traces instead of relying only on synthetic tests.
On the demand side, ad campaigns started having multiple entry points (an advertiser can start in the mobile app and continue in Ads Manager or via API), each with its own context that needs to be stitched together, on top of the requirement for closed-loop performance measurement feeding back into the system.
Until early 2026, multi-agent was a research demo, now it is an architecture.
Pratik Rasam, senior engineer at Spotify Ads
The cost of not having shared context
The talk uses a direct comparison between a function-style interface and an agent-style interface to show why this matters in practice. In the function style, each step of the ad-creation flow (brand briefing, script generation, targeting recommendation) extracts and uses only its own signals: the script-generation step, for example, receives the brand name and call-to-action, but loses the tone of voice that had been defined in the previous step.
In the agent style that Spotify adopted, each agent reads from and writes to a shared session state, so context accumulates at each step. By the end of the flow, the agent that generates the script has access to the brand tone, the call-to-action, the themes and the already-resolved audience, all together, instead of rebuilding that context from scratch on every call.
What remains open
Not everything on the platform is in full production: the Audience Recommendation Agent, which recommends audience targeting, remains in pilot. Spotify also maintains its own moderation layer, internally called kutest, for guardrail policy management, separate from the agents' business logic.
For those building multi-agent systems, the patterns of greatest practical value here are not about prompt engineering, but about platform engineering: explicit per-package ownership, import boundaries enforced at build time, and a clear separation between what is LLM judgment, what is data coming from an API, and what is simply deterministic code hidden behind a tool.
Translated from the Brazilian Portuguese original · Read the original
Cloudflare releases open-weight decision models for AI agents at the edge
The Clef and Clef-Flash models, announced by Cloudflare during its Birthday Week, trade text generation for typed probabilities, with open weights and latency designed to run inside the critical path of AI agents.