Cost Became an Engineering Problem After Access Expanded
LegalOn Technologies provides professional AI globally and also brought Codex into its own development process, gradually extending its use into everyday work across the organization. The company initially gave developers unlimited access to GPT-5.5 in Fast mode, encouraging teams to learn through practice where AI could contribute to design, implementation, and routine tasks. That approach addressed adoption first, but the scope of use and the resulting spend grew along with experimentation.
The conflict LegalOn then faced was concrete: leaving high-performance models available without limits could push spending beyond the annual budget, while a blanket ban or blunt cap could undermine productivity gains developers had already begun to realize. The company reports that after changing its usage practices, estimated daily Codex costs fell by 65% while development speed was maintained. The headline describes the result as “halving” costs, but the material does not explain how that summary relates to the 65% figure, so the two should not be treated as metrics with an established, identical basis.
Model Assignment Is an Escalation Path, Not a Leaderboard
The central change was to stop making the most capable model the default for every task. Teams start with a lighter model and move to a stronger one as complexity increases. GPT-6 Luna handles code implementation when requirements are clear, everyday automation, and relatively simple analysis. It can also act as a subagent for simpler work. GPT-6.1 Sol is used for standard design, data analysis, and document preparation, and the material also assigns it tasks expected to finish faster than they would with Luna.
GPT-6 Astra is reserved for complex analysis, architecture design, and coordinating multiple agents, where more judgment is required. This division shows that model selection is not just a contest over which model is more capable. It is a decision about when stronger capability is worth the extra resources. A lightweight model may be enough when requirements are clear and task boundaries are stable. Escalating for architectural judgment or complex coordination gives the additional capability a better chance of matching real work value.
Turning Selection Experience into Team Practice
This routing is not described as a fully automated dispatch system. LegalOn’s AI-powered Development CoE, or AID CoE, tests and monitors models, then turns observed use cases into guidance that managers share with their teams. Engineers use that guidance to choose a model for a specific task, while internal testing helps make the selection criteria more concrete over time.
The organizational value is that model selection need not depend only on the experience of individual developers. Teams can develop shared language around whether requirements are clear, whether a task needs complex judgment, and whether it involves coordinating agents. But the guidance will remain useful only if it is tested and updated as models and business needs change. The material describes monitoring and adjustment responsibilities, but does not provide data on routing accuracy, escalation frequency, or comparative output quality across models.
Speed and Spending Are Governed at Two Levels
Alongside task-based model selection, LegalOn changed Fast mode from a default setting to an option used when needed. Teams were concerned that this might slow development. The company says they used approaches such as running tasks in parallel to maintain performance and make the transition smoothly. This does not prove that a slower mode has no effect on work. Rather, it makes response speed an option that can be adjusted per request, while giving teams a way to compensate for waiting through how they organize execution.
The other layer of control is budget authority. Administrators set monthly usage limits for departments and individuals, while the AID CoE monitors consumption and can adjust allowances as business needs change. Model guidance answers which level of capability fits a task. Monthly limits answer who can use how much of the available resource. Together, the two controls reduce reliance on individual restraint alone while leaving managers room to recalibrate budgets against actual needs.
Different Business Stages Call for Different Budget Goals
LegalOn did not apply the same cost target uniformly across every business. Mature businesses were asked to improve cost efficiency, with a target of around 20%, while new businesses in their launch phase received more generous budgets to encourage active AI use. This is not simply a matter of giving newer teams more and established teams less. It puts spending back in the context of business stage: mature operations need to focus on efficiency, while newer ventures may need room to test and explore.
That differentiation has costs of its own. Tightening budgets for mature businesses too aggressively could obstruct valuable development work. Leaving launch-stage budgets generous indefinitely without examining outcomes could also detach spending from practical value. The material does not disclose specific allocations by business or explain how each stage is determined, so the transferable lesson is to adjust resources to business needs, not to copy a particular budget ratio. Uniform cuts may look easy to administer, but they do not necessarily preserve spending for the tasks where it matters most.
The 65% Figure Is a Reported Result, Not a Universal Formula
LegalOn describes a set of management actions that work together: assign models by complexity, limit Fast mode by default, set monthly caps for departments and individuals, and allocate budgets according to business stage. The company reports a 65% reduction in estimated daily costs while maintaining development speed. But the material does not explain how estimated daily cost was calculated, the comparison period, or provide comparative data on task completion time, quality, rework, or actual bills. This is a company-reported experience, not evidence that other teams will achieve the same reduction.
For technical leaders considering the approach, a safer first step is to establish a task-level baseline rather than copying the Luna, Sol, and Astra assignments. Track the model used, cost, time, and output quality for each task type, while recording model versions and operating modes. Then examine whether model escalation reduces rework, whether parallel execution offsets changes in speed, and whether budget differences across businesses correspond to real needs. Without continuous measurement, a cost reduction may reflect a changed estimate, while apparently stable speed may conceal shifts in quality or engineer workload. The decision criterion should be whether each increase in capability can be justified by task value, not whether the organization hits a fixed savings percentage.