# How we cut a browser agent's cost 76x and made it 5x faster by keeping its cache intact

> How Asana cut a browser agent’s cost 76x and made it 5x faster — by caching the full conversation history and pruning screenshots in batches instead of one at a time.

Source: https://asana.com/inside-asana/cut-browsers-agent-cost

## How we cut a browser agent's cost 76x and made it 5x faster by keeping its cache intact

_A study with GPT-6 Astra in Codex across four models, from investigating the code to measuring the results_

_**Figure 1. Humans direct GPT-6 Astra in Codex, which logs its work in Command.**_

## **Why we looked**

Asana is bringing agent workflow automation to its platform through[StackAI](https://www.stackai.com/), and at Asana's scale, small inefficiencies add up. Our aim was to make one browser agent cheaper and faster without reducing answer quality.

_“This is what teams of humans and agents look like in practice. An engineer set the direction, an agent ran the experiments, and the results went through Command to production. This demonstrates how Asana brings human+agent teams to life.” -_**Arnab Bose, CPO at Asana**

## **What was going wrong**

A browser agent resends its tools, system prompt and growing history of page text and screenshots on every call. Prompt caching lowers the cost of repeated input: on the models we tested, cache reads cost 0.05x to 0.1x the standard input price. But a cache only reuses the longest unchanged prefix of a request. Our agent cached its tools and system prompt but not its history, and caching the history alone would not have helped: the agent removed the previous screenshot at every step and trimmed older text to fit a history budget. Each edit altered an earlier part of the request, so reuse broke on almost every call.

## **The fix**

First, cache the history too, with a cache marker on the latest tool result. Second, stop editing it on every call. Batch pruning keeps screenshots and removes them in batches: at 20:1, the agent keeps up to 20 and then cuts back to 1, so about 19 consecutive calls reuse the cached history. We also increased the history budget from 120,000 to 480,000 characters, so old text was no longer trimmed.

**Figure 2. Per-step screenshot removal breaks reuse on every call; batch pruning keeps the history unchanged between pruning steps.**

## **How we ran it**

GPT-6 Astra, working in Codex, did most of the work: it audited the code, instrumented every request, ran quick tests to identify the variables that mattered, refactored the code to run workflows in parallel, launched the runs and analyzed the traces. Humans set the goal and the standards and reviewed the conclusions.[Command](https://asana.com/product/command), Asana's software delivery platform, was the system of record: every session's requests, traces and results were recorded there, so we could analyze the full study afterward, and the findings became tickets, pull requests and reviewed changes that shipped.

_**Figure 3. Roles in the study: humans decide, GPT-6 Astra in Codex does the work, and Command keeps the record, from traces to shipped changes.**_

**I never thought to do it by hand as it would have likely taken me months. With Codex, it took about a week: I'd set a /goal before going to bed and review the results in the morning.**

**Now, every night our agents aren't running feels like a wasted night.**

We tested six caching and history policies at two budgets on GPT-6.1 Sol and three other frontier models, Models A, B and C (table below), with three runs per condition: 144 runs, plus a 12-run follow-up. We calculated costs from each provider's token counters, and every answer was scored against an independently prepared reference.

**Model**

**What it is**

**Price**

**Model A**

A smaller, lower-cost model from another frontier lab, released Fall 2025

Half of GPT-6.1 Sol's price

**Model B**

The model originally used in production, from the same lab as Model A, released Summer 2026

Same as GPT-6.1 Sol

**Model C**

A newer version of Model B, released Fall 2026

Same as GPT-6.1 Sol

**GPT-6.1 Sol**

OpenAI's model

Reference

## **Results**

_**Figure 4. Cost and time per run: original setup on Model B against the optimized agent on Models B and C and GPT-6.1 Sol. Means of three runs; capped baseline runs make these fold reductions lower bounds.**_

On Model B, the best condition cut cost per run 29x and ran 4x faster than the original production setup. On GPT-6.1 Sol, the same agent cost 76x less and ran 5x faster, reading 89% of its input from cache. On every model, every best-condition run cost less than every baseline run and encountered all 192 facts.

## **Learnings**
- **The budget has to fit the model**. Newer models used up the smaller budget faster: Model C first trimmed its history at call 10, Model A at call 64. At 120,000 characters, Model C answered in none of 18 runs and Sol in 3, most hitting the step limit; at 480,000, every run on both answered.
- **Caching alone is not enough**. Without batch pruning, caching the history at the larger budget cost more than not caching it on three of four models: the cache was continually rewritten and rarely read.

### **Do you need to prune at all?**

The per-call data suggested that pruning might not be needed here, so a follow-up kept every screenshot. Cost per call was 1.2x lower than the best condition on Model B and Sol, and about 5% lower on Model C. Pruning still matters for long tasks, small context windows and more expensive cache reads.

With three or four runs per condition and call counts that vary between runs, this study shows broad patterns rather than distinguishing conditions only a few percent apart.

### **Guardrails still matter**

No run reached the 480,000-character budget, so it did not constrain these runs. Limits still matter: a drifting agent grows toward the context limit, and if the cache breaks, every call pays full price. Caps on steps, tokens and cost per run limit the cost of a bad run. Compaction is another option, beyond the scope of this study.

## **From findings to production**

Some findings emerged later, when other agents reviewed every trace logged to Command. That is hard to do from a single session, where you can't be sure all the data was stored. Keeping everything in Command, which our agents accessed through MCP, meant we never had to repeat an experiment to recover missing data, which sped up the work. Findings became tickets for humans and coding agents, pull requests were reviewed there, and the changes shipped to StackAI. Since the study, we have also run the same task with Codex, which compared its approach with the StackAI agent's and suggested a second round of improvements, now tickets in Command.

## **What we took away**
- Cache the growing history, not just the system prompt.
- Keep the history append-only; when pruning is necessary, prune in large batches.
- Set the history budget to suit the model.
- Measure cache reads per call using the provider's own counters.
- Keep a full record of every session, so the analysis and the follow-up work use the same record.

Asana is building tools to make experiments like this routine. The [full study](https://assets.asana.biz/asset/fb86cc8e-6611-4404-b7c6-172692a724bb/How-Asana-used-Codex-to-optimize-browser-agent-costs-and-runtime.pdf) covers methods, results and limitations.

**Shipping speed is no longer the bottleneck; human attention is. I think we're close to a world where every engineer is a PM leading a fleet of agents.**

- [Inside Asana's Database Architecture: Sharding, Data Flow, and Scaling Tradeoffs](/inside-asana/database-architecture-sharding-scaling)

Engineering

At Asana, changes to work appear in real time across the product by default. That responsiveness increases database traffic and pushes our databases toward their scaling limits as ...

- [Microframeworks in the Admin Console](/inside-asana/microframeworks-admin-console)

Engineering

Every Asana deployment has an admin console. It's where IT admins configure how their company uses Asana, such as adjusting password requirements, roles and permissions, whether f ...

- [Spec-driven development: The Good Parts - and what we learned after three months](/inside-asana/spec-driven-development)

Engineering

#### Staff Software Engineer

After three months, we had a clearer picture of when the added structure helped and when it got in the way.One of our engineers was preparing a data migration and decided to use s ...

- [We migrated off Enzyme in 2 weeks. It should have taken five years.](/inside-asana/migrating-off-enzyme-2-weeks)

Engineering

We recently used AI to complete years of engineering work in about one sprint. Here's how, and why it's changed how we think about what's possible.The five-year problemBack in 202 ...

- [How Asana used Astra, Codex and Command to reduce browser agent costs](/inside-asana/cut-browsers-agent-cost)

Engineering

Artificial Intelligence (AI)

- [CTO of StackAI](/author/frank-hidalgo)

A study with GPT-6 Astra in Codex across four models, from investigating the code to measuring the resultsFigure 1. Humans direct GPT-6 Astra in Codex, which logs its work in Comm ...

- [Engineering](/inside-asana/engineering-spotlight)
