Evidence at a glance
First, identify what the 76-fold figure compares
Asana published a case study on optimizing StackAI, its browser agent for no-code business workflows that navigate websites, fill out forms, and gather information. StackAI CTO Frank Hidalgo asked GPT-6 Astra in Codex to inspect the agent’s code, design experiments, and compare results, with the goal of reducing runtime and cost without lowering answer quality. The optimized workflow using GPT-6.1 Sol averaged an estimated $0.47 in model costs per run and took about four minutes.
The headline claim of being 76 times cheaper and five times faster compares optimized Sol with the original production setup using Model B. It is not a before-and-after result for one model. Broken out, Model B fell from at least $36.21 to $1.24 per run, a reduction of about 29 times, while Sol fell from $1.97 to $0.47, about four times. The report notes that some baseline runs hit the step limit, so the cost and speed comparison is approximate, with the baseline cost and duration treated as lower bounds.
How the agent edits context determines whether it can reuse the cache
The problem lay in how the agent organized its request history. It already cached fixed instructions and tool definitions, but not the growing record of webpage text and screenshots, so those contents were sent again with each call. At the same time, the agent trimmed old text and removed screenshots at nearly every step. That kept changing the request history and breaking the stable prefix needed for cache reuse.
The fix was not simply to turn on more caching. It was to preserve previously collected history for longer and extend caching to that history. Caches rely on a continuous, unchanged prefix, so repeatedly rewriting old content undermines reuse. For Sol, cached input cost 5% of uncached input, and 89% of input in the best configuration came from the cache. The way the agent retained its history therefore became part of the cost architecture.
Batching screenshot cleanup keeps history stable for longer
The engineer chose three changes to test: caching the browsing history, retaining more text, and stopping the agent from deleting screenshots at every step. The best Sol configuration increased the history budget from 120,000 to 480,000 characters and let screenshots accumulate to as many as 20 before trimming them back to the most recent one. The goal was not to keep every screenshot forever, but to reduce how often the existing history changed.
A larger history budget also affected whether the task could be completed. With a 120,000-character limit, Sol produced an answer in only 3 of 18 runs. At 480,000 characters, it completed all 18 runs and returned correct answers. This is a useful warning for engineering leads: shrinking context may reduce input, but discarding information needed for the task can force the agent to revisit pages or prevent it from delivering an answer at all.
The 144-run study made optimization comparable
The main study was a comparison matrix, not a single demonstration: four models, six caching and screenshot policies, two history budgets, and three runs per condition, for 144 runs in total. Every configuration performed the same task: extracting six fields from each of 32 books in a public demo catalog, or 192 facts altogether. The team also ran 12 follow-up tests without deleting screenshots to examine that policy further.
The experiment framework itself had to change because the existing code was not designed for controlled comparisons. GPT-6 Astra helped refactor the frontend and backend so multiple workflows could run in parallel with separate settings, and examined requests, usage records, and outputs. The engineer reviewed the proposed changes and conclusions. Asana estimated that the work would have taken one to two months by hand but took about a week. The useful point is not to treat that time saving as a general productivity benchmark; it is that the agent helped execute and analyze the experiments while a human set the direction and reviewed the findings.
Transfer the diagnostic method, not the 76-fold promise
For teams running browser agents, the case suggests a practical diagnostic order: first measure how much history is resent with each request, then track how cache hits, cost, runtime, and completion rates change under different history policies. Small, controlled comparisons on the same task help distinguish savings from caching, history budgets, screenshot cleanup, or the model itself. Workflow improvements may create room to use a more capable model, but that does not mean a stronger model is automatically cheaper.
The limits of the result are just as important. The study covered one book-catalog demo task, with three runs per condition, so it cannot establish that other customer workflows will see the same reduction. The original production baseline was also affected by the step limit. The 76-fold figure is best treated as a reason to inspect duplicated context and history edits, not as a budget forecast. Before changing history limits or model settings, teams should repeat the comparison on their own tasks and check answer quality, completion rates, and cache costs together.