

From Writing an Answer to Answering a Defined Question
Liquid AI has released Open d1, two open-weight multimodal models in its d1 decision-model family. d1-3B accepts text and images, while d1-omni-600M accepts text with an image or text with audio. Neither produces a natural-language answer. Instead, in one forward pass, each takes an input state and a set of questions and returns typed probability results. The response reports output_tokens as 0, indicating that the model does not deliver its result by writing an answer token by token.
What changes is the interface between the model and the application, not merely the length of the output. A typical generative call first returns text, after which code extracts a label, score, or judgment. Parsing the format can itself fail. d1 puts the permitted answer space into the call and returns probabilities over those allowed outcomes directly. It still requires computation and can still make mistakes. The work shifts from interpreting a piece of prose to checking a structured decision.
The Business Defines the Answer Space First
d1 offers three question types. A noul question is a yes-or-no question and returns P(yes). A choice question selects among named options and returns the full probability distribution. A score question returns a probability-weighted position on an ordered rubric, which can contain between two and ten levels. Multiple questions can be submitted about the same input in one call, so a customer message could be assessed for eligibility, assigned to a team, and scored for urgency at once.
This design suits workflows whose answers can be enumerated in advance, and makes it easier for the calling application to see what it asked the model to decide. But a structured result does not repair a poorly framed question. If a refund rule omits an exception, or the routing options lack an appropriate destination, the model can only answer within an incomplete set. A precise-looking probability does not mean that it is calibrated on the business’s data. Technical leads need to treat options, rubric anchors, and thresholds as part of system design, not as incidental parameters filled in during an API call.
The Two Checkpoints Do Not Have Identical Modalities
d1-3B has 3.12 billion parameters and is based on the decoder-only LFM2.5-VL-3B. It uses a 400-million-parameter SigLIP2 NaFlex vision encoder and has a context length of 32,768 tokens. Liquid AI describes a build process that averages model weights, fine-tunes multiple checkpoints with different random seeds and data mixtures, and then merges those checkpoints. The material also highlights long inputs, shuffled answer options, and fixing data shortcuts. This points to a training challenge: preventing a model from treating superficial cues as evidence for its decision.
d1-omni-600M is an early research release with 587 million parameters in total. Its shared trunk and decision head account for 381 million, its vision encoder for 94 million, and its audio encoder for 112 million. It uses the bidirectional LFM2.5-Encoder-350M trunk and a 17-layer FastConformer audio encoder. Its context is 16,384 tokens, and audio clips are capped at 30 seconds. A request can contain an image or audio, but not both. Audio training covered English speaker-to-assistant requests. These constraints mean that “supports audio” should not be taken to mean suitability for every language, acoustic environment, or audio task.
Zero Output Tokens Are Only One Part of the Latency Picture
Liquid AI presents Open d1 for deployment across the NVIDIA stack, from DGX servers and RTX workstations to Jetson edge boards. Both checkpoints are available on Hugging Face, load through Transformers, and had day-one llama.cpp support. These entry points lower the barrier to integration experiments. Whether a model is production-ready still depends on the target hardware, inference stack, and request shape, not on open weights or runtime compatibility alone.
The d1-3B model card reports a fastest decision time of 8 milliseconds on an RTX 4090. Without model.compile(mode="reduce-overhead"), the figure on that card is 16 milliseconds per question. The published GPU numbers were measured in bf16 with one warm request at a time, and report the median of 20 runs. NVIDIA measured the Jetson figures. The material uses 33.3 milliseconds as the duration of one frame at 30 frames per second, but that reference does not replace an end-to-end latency budget for a particular service. In particular, d1-omni-600M has no published latency figures, so d1-3B measurements should not be extrapolated to it.
Business Risk and Evaluation Set the Boundary
Liquid AI lists routing, content moderation, intent classification, reranking, LLM-as-a-judge scoring, agent guardrails, and visual inspection as possible uses. These tasks share a useful property: a system can often define candidate labels or a scoring scale in advance. They may therefore benefit from direct probability outputs instead of parsing free text. For example, customer-service triage could separately ask whether a refund is eligible, which team should handle the case, and how urgent it is, then let downstream application code act according to policy.
But returning probabilities does not mean those probabilities can be used as thresholds without checking. Before deployment, teams still need to evaluate classification quality, calibration, the cost of errors, and boundary cases on business data, as well as end-to-end performance on the target device. Open d1 uses the LFM Open License v1.0, which permits free commercial use below $10 million in annual revenue. Deployers still need to confirm that the terms apply to their own circumstances. A practical approach is to begin with a narrow task whose answer space is stable and whose failures can be monitored, compare it against the existing approach, and only then decide whether to replace a generative call with d1. Open weights, zero output tokens, and a single forward pass do not substitute for that evaluation.