Evidence at a glance
The mechanism in one line
Compress the visual or contextual input before the main reasoning path.
Route or verify the expensive step instead of repeating the full path.
Translate the mechanism into a bounded deployment or evaluation check.
Reframing the Task from Generation to Selection
Liquid AI’s d1 is a model for structured decisions, not a general-purpose language model for writing text. The caller submits a state and a set of predefined, typed questions. In one request, d1 returns probabilities over fixed outcomes rather than producing an explanation, label string, or JSON document. Every response reports usage.output_tokens as 0. The point is not that it writes less; the result is not organized as a generated token sequence at all.
The target is the large set of narrow tasks that teams often hand to general LLMs: ticket classification, routing, scoring, moderation, reranking, and LLM-as-judge checks. Liquid’s migration rule is straightforward: if the answer is one of N known options, consider a decision model; if the system must compose a new string, keep the LLM. That distinction determines where d1 belongs in a workflow, rather than treating it as a cheaper conversational model.
Three Primitives Turn Classification into a Probability Interface
The interface is built around three question primitives. Noul handles binary questions and returns a probability between 0 and 1, such as whether a message is a customer complaint. Choice selects from named options and returns the top pick, the full distribution, and a confidence value. Score places an input on an ordered rubric and returns a probability-weighted position. Because levels are indexed from 0, a four-level urgency rubric spans 0 through 3.
All three can be combined in one request, allowing the model to evaluate several questions against the same state in one round trip. In the published examples, a duplicate-charge ticket receives 0.9997 for billing, a production outage scores 2.9995 on a four-level urgency rubric, and a complaint check returns 0.999. These should be treated as probability outputs exposed by the interface, not as a model verbally reporting its confidence. For engineers, the practical difference is that the values can feed thresholds and routing logic without first parsing generated text.
The Architectural Gain Lies Beyond Decoding
A d1 call has three parts: the model, the state, and the questions. Requests go to Liquid API’s systemone endpoint, with the published example using d1:free through TypeSafe AI’s Python or TypeScript SDK. The state may be plain text or a JSON object, which makes input representation part of the system design rather than forcing every workflow to place one large prompt into context.
In its comparison with an LLM using structured output, Liquid highlights several engineering trade-offs: no billed output tokens, no decoding loop that grows with output length, and answers that always match the declared question type. That removes malformed-JSON and retry paths. Three sequential classification calls can also become one request. For low-latency routing and moderation pipelines, these changes matter more than sounding human because they reduce uncertainty in the protocol and scheduling layers.
Probabilities Become Useful Only When They Enter Policy
The value of d1 is not that it returns a precise-looking decimal. It is that uncertainty can be written into policy. Liquid’s moderation example blocks probabilities above 0.8, allows those below 0.2, and sends the middle band to human review. Its routing example falls back to a more capable model tier when confidence drops below 0.5. The model becomes a thresholded decision node inside a policy engine rather than a one-shot verdict generator.
That changes what teams need to evaluate. The question is not only whether the model can consistently produce a label, but whether its probabilities support thresholds, whether repeated evaluations reduce verdict flips, and whether the human-review band remains manageable. The material also says repeated evaluations of the same input are more consistent, but it provides no independent benchmark, latency measurements, or calibration error. The interface lowers integration friction; it does not remove the need to validate calibration and business thresholds.
Deployment Boundaries Keep It from Being a General Replacement
d1 is deployable today, but through Liquid API as a hosted service under the model name d1:free. Liquid’s model library marks it as API only and not trainable, so there are no GGUF, MLX, or ONNX weights for self-hosting. For teams that must keep data local, run at the edge, or fine-tune a model, this is not a minor limitation. It is a prerequisite in the architecture decision.
The Road Decider demo shows the intended shape of the system. A Node.js 18+ application uses a Choice question to select the left, center, or right lane on each decision tick, roughly 2 to 5 times per second, while a Vite proxy keeps the API key server-side. The useful lesson is not the game itself but state design: per-lane summaries with distance to the first obstacle produced more confident decisions than a raw road grid. Teams adopting d1 should first express state in a form that can be judged, then define thresholds and fallbacks. For summarization, drafting, multi-turn chat, code generation, or complex multi-step reasoning, keeping an LLM remains the safer boundary.