source preview
source preview Open source material ↗
Source figure
来自一手来源:model_card Open source material ↗
Source figure
来自一手来源:model_card Open source material ↗

From Generating Answers to Making Decisions

Liquid AI has released two open decision models, d1-3B and the experimental d1-omni-600M, for tasks that need fast, structured results from text or multimodal inputs. Their main difference from conventional generative models is how they produce an answer: rather than emitting tokens one by one, d1 answers in a single forward pass. For questions such as whether a customer is requesting a refund, which team should handle a ticket, or how urgent it is, the system can return a predefined decision instead of generating prose that another program must parse.

This changes the model’s role in an application. A generative model is often used for open-ended expression; a decision model is closer to a callable judgment component. The caller supplies a piece of state and a question with options or scoring criteria, then receives a structured result. Liquid AI’s examples ask about refund status, team assignment, and urgency from the same customer message, and also show an interface for processing multiple tickets. This design suits workflows whose possible outcomes can be bounded in advance, not every task that involves language.

Two Backbones, Two Input Paths

d1-3B is built on LFM2.5-VL-3B, a decoder-only vision-language model, so it accepts text and images. d1-omni-600M starts from LFM2.5-Encoder-350M, a bidirectional encoder, and adds vision and audio encoders. It accepts either text with images or text with audio; the release does not say that it handles image and audio together in a single input.

The trade-off is not simply about parameter count. The 3B model inherits a vision-language backbone, and Liquid AI says it retains that backbone’s performance on standard vision benchmarks. The 600M model aims for a smaller footprint and broader modality support, but remains an early research release. Liquid AI reports neither speed figures for d1-omni-600M nor decision benchmark scores for its vision or audio capabilities. Support for an input modality describes an interface, not yet publicly demonstrated decision quality or latency.

Rankings and Latency Answer Different Questions

On Decision Index 0.2.1, Liquid AI reports a score of 48.57 for d1-3B. The company calls it the best decision model under 10B parameters, ahead of every tested 4B and 9B model, and above Decider 35B-A3B at 47.11. A separate evaluation spans seven public datasets covering reading comprehension, toxicity detection, intent classification, medical question answering, and cross-lingual understanding. d1-3B has a reported mean score of 82.9; d1-omni-600M scores 78.4, above Decider 2B’s 77.1 with roughly one quarter of the parameters.

These figures demonstrate performance on particular public tasks, not general decision-making ability. The two sets of results are reported in different evaluation contexts and should not be treated as scores on one interchangeable scale. Liquid AI also says Decision Index v0.3 has only a private vision split and that audio decision benchmarks remain an open problem, so it publishes no vision or audio decision scores for d1. For technical leads, the practical question is not whether a leaderboard position transfers to their product, but whether their task resembles the evaluated datasets and how the model performs against their own labels and error costs.

Edge Speed Is Promising, but Deployment Still Needs Validation

The published deployment figures for d1-3B are single-question inference latencies. In evaluations with NVIDIA, Liquid AI reports 16 milliseconds on Jetson AGX Thor, 26 milliseconds on Jetson AGX Orin, and 50 milliseconds on Jetson Orin Nano. It says the model answers one question in under 50 milliseconds on every tested device. On Thor, three questions take 20 milliseconds compared with 16 milliseconds for one, suggesting that multiple judgments can share some processing overhead. On GPUs, the release reports under 10 milliseconds to answer a question and under 18 milliseconds to process a 384-pixel image on the tested platforms.

These are useful deployment signals, not a complete production capacity guarantee. The release does not provide figures for different concurrency levels, input lengths, power consumption, or long-term stability, and it gives no speed results for d1-omni-600M. The example code requires Transformers 5.14 or later and loads model-provided code with trust_remote_code=True, so teams should review the code and runtime environment before integration. If the workload is fixed-label classification, routing, or constrained-choice decisions, d1-3B’s latency and multi-question interface merit local testing. For open-ended generation, or workflows that depend on reliable audio decisions, neither leaderboard rankings nor modality support alone is sufficient evidence.