


Evidence at a glance
The Four-Minute Result Is Not Just Model Speed
Chatham Financial advises clients on complex capital markets decisions. On October 2, 2026, OpenAI described how Chatham uses Codex to build internal and client-facing tools, with GPT-5.6 powering AI features across employee-built applications, the Chatham Onyx capital markets operating system, and a workflow reengineering service called Process Zero. The problem is concrete: transaction records must match what a client authorized and what was actually executed, while the validation process must remain traceable and auditable.
The important point is not that a financial firm has connected another large model. The efficiency gain comes from redesigning the workflow. Chatham says its trade validation application reduced a review that previously took about 30 minutes to under four, but the application did not gain authority to approve transactions on its own. It gathers evidence, compares key terms, flags discrepancies, and leaves judgment-intensive cases to professionals.
Process Zero Defines the Evidence Before the Automation
The core method behind Process Zero is to start with the desired outcome rather than with what a model happens to be able to do. For each workflow, Chatham identifies the minimum inputs and evidence required, separates the steps that require human judgment, and only then decides where AI and AI-built tools should participate. That sequence changes the starting point for system design and prevents a conversational interface from being mistaken for a complete business process.
Trade validation shows what this decomposition looks like. The Controls and Data Integrity team is responsible for the accuracy of transaction data, while the application handles evidence gathering, key-term comparison, and discrepancy flagging. It is not presented as a black box that replaces senior reviewers. It is a control layer that structures the mechanical work before expert review. For technology leaders, the system therefore needs more than an input-and-answer interface. It needs source evidence, comparison results, reasons for exceptions, and a record of the subsequent review.
The Validation Gate Defines What Four Minutes Means
Chatham explicitly says that the early measurements are being compared with real transactions and experienced reviewers before automation is expanded. That qualifies what “under four minutes” means. It is currently a measure of how quickly the application can perform a round of processing or preliminary review, not a promise of end-to-end unattended execution. The source also describes greater accuracy as an objective to validate, not as a fully established production result.
This is a crucial difference between financial workflows and ordinary office automation. Speed can be measured in a single run, but accuracy has to be tested repeatedly against real transactions, edge cases, and human baselines. If a system shows only a faster average while failing to preserve evidence sources, explanations for discrepancies, and the identity of the final decision-maker, it may simply spread errors more quickly. Chatham plans to extend the application to more trade types while retaining appropriate controls and professional oversight, making validation the gate for broader automation.
From Employee-Built Apps to Onyx, Two Paths Emerge
Chatham has not placed every AI capability inside a single top-down platform. Employees use ChatGPT and Codex for research, analysis, drafting, and software development, and they use the internal Chatham Vibes platform to create applications for specific jobs. Those applications support reviews of maturing-cap trades, pricing workbooks, client communications, fixed-income rate sheets, hedging dashboards, and trade confirmations. AI features in Vibes use GPT-5.6 Terra by default, with GPT-5.6 Sol available through application configuration.
Chatham Onyx has a different role. It brings assets, debt, and derivatives into a connected, governed data environment where clients and advisors can use AI while tracing outputs back to their underlying sources. Codex is used across planning, development, testing, documentation, and code review, and the platform uses GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.4, and GPT-4.1. The combination is operationally significant: employee-built applications provide speed close to the business, while Onyx turns selected capabilities into more durable and governable product functions.
Technology Leaders Should Copy the Method, Not the Headline Metric
The actionable lesson for financial technology teams is to begin with workflows that have stable evidence structures and exceptions that experts can handle. The first phase does not need to pursue full autonomy. It should establish the inputs, evidence, model outputs, exception categories, and human conclusions, then build a comparison baseline using real transactions and experienced reviewers. Only when those records support audit and reconstruction is there an engineering basis for expanding trade types or reducing manual steps.
Chatham's approach also exposes an organizational cost. Employee-built applications can move quickly and stay close to operational needs, but they create variation in model configuration, permissions, data sources, and quality standards. A centralized platform can provide governance and traceability, but requires more product and maintenance investment. Four minutes is therefore not a natural consequence of buying a model. It is the product of workflow boundaries, evidence management, professional review, and software engineering. If a team cannot answer what evidence triggered a system's flag, who confirmed the result, and how an error can be traced back, it is not ready to extend the same automation to higher-risk transaction workflows.