AIARTICLE

Asana Cuts Browsing Agent Cost by 76x With GPT-6.1 Sol

An internal study with 144 runs showed the bottleneck wasn't the chosen model, but how the agent managed cache and browsing history. OpenAI published the case as an example of optimization driven by a coding agent.

Asana, in partnership with OpenAI, published a study on how it cut the cost of a web browsing agent used by StackAI, a platform the company acquired to automate tasks on websites and no-code forms, by 76 times. The case, described by OpenAI, isn't about switching models: it's about fixing how the agent handled its own browsing history.

The technical protagonist of the story is Frank Hidalgo, CTO of StackAI at Asana, who used GPT-6 Astra running on Codex to investigate the agent, propose fixes, and compare results. According to Hidalgo, work that would have taken one to two months by hand was done in about a week.

This would have taken me one to two months by hand. With GPT-6 Astra in Codex, it took about a week: I'd set a /goal before going to bed and review the results in the morning.

Frank Hidalgo, PhD, CTO of StackAI at Asana

Where the Money Was Being Burned

The investigation began with GPT-6 Astra mapping the codebase to explain how the agent assembled each request to the model. The finding: the agent cached the fixed instructions and tool definitions, a standard practice in LLM APIs, but it didn't cache the growing history of page text and screenshots it accumulated during browsing.

In practice, each call resent that entire history at full price, even when much of it had already been processed in the previous call. To make matters worse, the agent discarded old screenshots and trimmed text on almost every step, which invalidated any attempt at caching, because each edit altered the history that came before it. It was a common pattern in agent prototypes: the context-pruning policy, meant to save tokens, ended up sabotaging the cache and multiplying the cost.

An Experiment With 144 Runs and Four Models

With the diagnosis in hand, Hidalgo selected three hypotheses to test: extending the cache to the browsing history, increasing the amount of retained text, and removing screenshots in batches instead of on every step. Since the original code wasn't designed for controlled experiments, GPT-6 Astra refactored the application so a single frontend and backend could support multiple configurations in parallel.

The study's final design tested two history budgets (120,000 and 480,000 characters) and six cache and screenshot-discard policies, with each combination run three times across four different models: GPT-6.1 Sol plus three models from other labs, referred to in the publication as Model A, B, and C. The benchmark task was always the same: collect six fields for 32 books from a public demo catalog, representative of what StackAI customers run in practice.

ModelOriginRelative Price
Model ACompeting lab, released in late 2025Half the price of GPT-6.1 Sol
Model BSame lab as A, originally used in production, released in mid-2026Same price as GPT-6.1 Sol
Model CUpdated version of B, released in late 2026Same price as GPT-6.1 Sol
GPT-6.1 SolOpenAIReference

All sessions, requests, and results were logged in Command, Asana's own software delivery platform, which made it possible to review the entire study afterward and turn the findings into tickets and pull requests all the way to production.

From US$36 to US$0.47 per Run

On Model B, which originally ran in production, the optimization cut the estimated cost from at least US$36.21 per run (some original runs hit the step limit before finishing) to US$1.24, a 29x drop. The optimized flow running on GPT-6.1 Sol was 2.6 times cheaper still, landing at US$0.47 per run, the figure that gives the study its name: 76 times cheaper than the original production setup.

Isolating just the effect of the cache policy on GPT-6.1 Sol, with the larger history budget, cost dropped from US$1.97 to US$0.47 per run, a 4x decrease. The explanation lies in the cache hit rate: 89% of the calls' input came from cache, billed at 5% of the price of a non-cached token. Execution time also dropped, from at least 22.5 minutes in the original Model B setup to about four minutes in the optimized flow, a 5x difference.

Perhaps the most relevant finding for agent developers isn't the cost, but the reliability: with the smaller history budget (120,000 characters), only 3 of the 18 runs on GPT-6.1 Sol produced a complete answer. With the larger budget (480,000 characters), all 18 runs finished, all with the correct answer. Aggressively cutting context didn't just fail to save money, it made the agent fail the task.

What This Changes for Agent Builders

In short: the gain didn't come from a cheaper model, it came from three engineering decisions about how to assemble the request's context. These are techniques replicable in any agent stack that accumulates state across multiple calls, not an OpenAI-specific trick.

The logic, simplified, looks something like this:

python
# naive pattern: only the system prompt is cached
request = system_prompt + tools + full_running_history  # recalculated on every call

# pattern observed in the study: cache extended to the history,
# with batch pruning instead of pruning on every step
if len(screenshots) > 20:
    screenshots = screenshots[-1:]
# page text remains intact across more calls,
# preserving the prefix that the prompt cache reuses

Three points are worth carrying into your own project:

  • Caching isn't just for the static prompt. If the agent accumulates history (page text, tool output, messages), that history can also be cached, as long as it isn't rewritten on every step.
  • Pruning context on every turn is the cache's enemy. Editing the history frequently invalidates the cached prefix; pruning in larger batches (every 20 steps, in the study's case) preserves the cache for longer.
  • A larger context can be cheaper, not more expensive, if it keeps the agent from having to revisit pages it already read. In the study, the more generous history budget was what made both the cache and the success rate possible.

The process used to get there is also notable: instead of an engineer manually testing hypotheses, GPT-6 Astra running on Codex executed the entire battery of 144 combinations and examined requests, usage logs, and outputs, while separate model sessions reviewed the work and produced a structured report. It's a use of a coding agent for empirical experimentation, not just for writing code, something that's still uncommon at most teams.

This is what teams of humans and agents look like in practice. An engineer set the direction, GPT-6 Astra ran the experiments, and the results went through Command to production.

Arnab Bose, CPO of Asana

What the Study Doesn't Resolve

It's worth weighing what's left out. The cost figures are specific to the tested task (32 books, six fields each) and to the prevailing pricing of the four compared models; the 76x ratio isn't a universal constant of agent engineering, it's the result of a specific workload with a specific cache bottleneck.

The study also doesn't account for how much it cost to run GPT-6 Astra itself to generate and review the 144 runs, nor the engineering effort to refactor the code to the point of supporting controlled parallel tests, something the publication acknowledges as a prerequisite for the experiment to work. For smaller teams, replicating the full methodology (history budgets, six cache policies, four models, three repetitions) may cost more in engineering time than the savings justify, unless the volume of production runs is high enough to offset it.

Asana says it has already rolled out the browsing changes to StackAI and plans to incorporate this kind of test into the platform's own evaluations, so customers can compare cost, execution time, and response quality when configuring agents. This suggests the methodology, today a one-off study, could become a product feature, which is the most interesting point for anyone tracking where agent tools are headed: from internal experiment to native platform capability.

Translated from the Brazilian Portuguese original · Read the original

View profile →