Dev & EngARTICLE

Rogo takes code written by AI agents to production in 5 minutes on Vercel

Financial startup Rogo reorganized its deploy pipeline on Vercel to put code generated by AI agents into production in five minutes, while another group of agents handles incidents without manual intervention.

Vercel published a case study about Rogo, a startup that builds Felix, an AI agent used by a large share of the world's largest financial institutions to generate decks, financial models, and reports. The material is short, but the central data point stands out: according to Vercel, Rogo's team restructured its internal applications so that code written by AI agents reaches production in five minutes. It's a concrete number, but it's worth separating what it proves from what it suggests.

One point that's easy to miss: the case study talks about "internal applications," not Felix itself. In other words, what's being automated in five minutes is Rogo's internal tooling pipeline (dashboards, integrations, operational workflows), not necessarily the engine that generates financial models for bank clients. This distinction matters because the appetite for letting an agent publish on its own changes a lot depending on who receives the result.

What fits inside the "5 minutes"

Vercel doesn't detail what counts within that window: just the build and deploy, or also the agent's code generation before the push? That ambiguity matters because five minutes for a build and deploy is already a good, but normal, number for anyone using Vercel's atomic deploys and instant rollback with any code, human-written or not. The differentiator Rogo is actually selling is something else: the confidence (or the guardrail structure) to let this pipeline accept commits from an agent without someone reviewing line by line before promoting to production.

This is the kind of change front-end feels first. Preview deployments per pull request, immutable aliases, and instant promotion have existed on Vercel for years; what changes here is who writes the diff. If an agent can open a PR, pass through the pipeline, and reach production in five minutes, the open question is: what gate remained between the agent's commit and the end user? The case study doesn't list automated tests, visual regression checks, or accessibility validation as part of this flow, and that absence says as much as the five-minute number.

Six agents in production, from churn to the deal desk

Rogo currently runs six AI agents in production doing work that ranges from churn analysis to support for the deal desk (the sales team that closes complex deals). These are operational tasks, not interface generation: the kind of automation that historically would have been a cron script or a spreadsheet with a heavy formula, now rewritten as an agent that makes decisions and, apparently, also writes and publishes its own supporting code.

In short: Rogo doesn't have a generalist agent writing just anything; it has agents specialized by business function, each with a defined scope. This is a more defensible architectural choice than "letting AI code the product," because it reduces the error surface to known domains.

Swarms of agents handling their own incidents

The boldest point in the case study is something else: when an incident happens in production, swarms of agents built on Vercel's AI SDK (the open source ai package, used to orchestrate calls to models and tools) handle triage and remediation without manual intervention. Vercel sums this up as "zero manual triage."

The text doesn't detail what "remediate" means in practice: rolling back to the previous deploy, applying a hotfix generated by another agent, or just isolating the problem and opening a ticket. These three things carry very different risks. Rolling back is safe and reversible; generating and publishing a hotfix on its own, in production, with no human in the loop, is the part I'd like to see with more transparency before replicating it on a team that serves banks.

The volume behind the narrative: 73,000 deploys per month

The other number in the case study is 73,000+ deployments in a single month. Divided across 30 days, that's more than 2,400 deploys per day. This is only possible with atomic deploys, cached builds, and, almost certainly, a large share of that volume coming from preview deployments (every PR, every agent iteration, generates a new environment), not just from production.

This volume is consistent with the rest of the case study: if agents write code all the time and every attempt generates a preview deploy for validation before promotion, the number of deploys explodes quickly compared to a purely human team. It's not an isolated indicator of production speed; it's an indicator of how many times the entire pipeline (preview + production) is exercised per month.

What remains open for front-end builders

Vercel's case study is, above all, customer marketing material, and that explains why it lists achievements without listing the brakes. For anyone deciding to adopt a similar model, the questions the text doesn't answer are the ones that matter most:

  • What tests (unit, E2E, visual) run before an agent's commit is promoted?
  • Is there an accessibility check, or is it only functional?
  • Who reviews the agent's code before deploy, or is the review only after the fact, via observability?
  • What happens when the remediation agent gets the diagnosis of its own incident wrong?

None of these questions invalidate what Rogo did. Vercel's atomic deploy and instant rollback infrastructure is real and already helps any team, with or without an agent writing code. But the leap from "fast deploy" to "agent publishes and also puts out its own fire without a human" requires guardrails that the case study doesn't show, and that's where any front-end team should spend their time before copying the model, especially if the final product reaches an end user and not just an internal dashboard.

Translated from the Brazilian Portuguese original · Read the original

Read also
↳

Threads