Evidence at a glance

700Evidence
7 AkashML, BoundlessEvidence
500 use 20Evidence
20% useEvidence
10% useEvidence
OpenRouter 500Evidence

More Than a Model Catalog: Procurement at Request Level

Architect Financial Technologies has launched Liquid Inference, a bidding router for large language model inference. It is aimed at development teams that want to keep using OpenAI- or Anthropic-format clients while choosing among inference providers. A client sends a request, and the platform selects a provider according to the buyer's rules. Architect describes the service as a spot market for inference, replacing some fixed-price or bilateral purchasing with provider bids on individual requests.

This is more than a larger model catalog. A conventional router can distribute traffic across available services, while Liquid Inference puts provider offers and buyer constraints into the same selection process: rules remove ineligible options, and the remaining offers compete. Technical leaders should therefore ask more than whether they can change an API base URL. They should also ask how much of the purchasing judgment once shared across model choice, vendor contracts, and runtime policy is now delegated to the platform.

Rules Come Before Bidding; the Lowest Price Does Not Always Win

In Architect's published flow, providers register offers for models. When a buyer's request arrives, the platform filters offers against rules such as a per-job cost cap, time to first token, minimum throughput, permitted regions, zero data retention, and provider or model allowlists. The lowest-priced offer that passes those checks wins. The system does not unconditionally choose the cheapest offer on the platform; it compares prices within the buyer's eligible set.

The maximum price is locked before the first token is generated, billing is based on metered usage, and the platform provides a per-job record that includes the winning provider, price cap, and final charge. This makes some procurement constraints configurable at request time rather than leaving them in purchasing documents. But rules can establish that a service meets a threshold; they cannot establish that outputs from different providers are equivalent or that latency will remain stable over time. Teams that treat model quality as a hard requirement beyond cost still need their own evaluation methods and a way to bring those results into routing or vendor management.

Market Transparency Is Not the Same as Market Depth

One distinctive feature of Liquid Inference is that account holders can view live order books, quotes broken down by provider and model, and cleared transactions. For procurement and platform teams, this data could extend the question from what a particular call cost to how offers and completed transactions change over time. That may inform budgets, routing policies, and vendor negotiations. Per-job records should also make it easier to trace where a charge came from than a single consolidated bill would.

But visible quotes do not guarantee a deep pool of eligible offers. Architect says its testing phase included hundreds of completed tasks across more than 700 models and names seven initial provider partners. Those are company-reported testing and partnership figures, not independent evidence of live competition. Public materials do not explain how often quotes update, how long bidding remains open, how ties are resolved, or exactly how a registered offer maps to an individual incoming request. An order book can improve visibility, but its existence alone does not prove that prices are formed through meaningful competition.

Easy Integration Does Not Remove Responsibility

For application teams, compatibility with OpenAI and Anthropic formats may lower integration effort, and Architect lists tools such as Claude Code and Cursor as compatible. Providers can register models and offers through REST or WebSocket interfaces and receive payment through Stripe. If an existing codebase already uses compatible APIs, changing the routing layer may be easier than rewriting inference calls. But “change the base URL” describes technical integration, not the absence of operational or governance costs.

A comparison with OpenRouter illustrates the trade-off. Liquid Inference emphasizes per-request bidding, a locked price cap, and quote and transaction data. The supplied material describes OpenRouter as supporting more than 500 models and 80-plus providers, with price-weighted load balancing. The former puts constraints and offers into each selection; the latter is presented with a more explicit scale of model and provider coverage. The available information does not establish which is better on latency, quality, or total cost, and bidding should not be treated as synonymous with lower prices.

There is also a boundary that the exchange metaphor can obscure. Architect's product terms say that buyers purchase inference from Architect, which in turn purchases it from independent providers; those providers are subcontractors and have no direct contract with the buyer. Before a trial, technical teams should clarify who handles data processing, incident escalation, refunds, and provider replacement, then record quality and cost under different routing rules. Liqu