AI News Feed
Market watch
Companies

Liquid AI Releases Open-Weight d1-3B and d1-omni-600M Decision Models

Liquid AI has released Open d1, two open-weight multimodal decision models that return calibrated typed answers with zero output tokens. d1-3B and d1-omni-600M are on Hugging Face and target real-time decisions on NVIDIA hardware.

Both checkpoints are available on Hugging Face, load through Transformers, and have day-one llama.cpp support. The LFM Open License v1.0 permits free commercial use for organizations with less than $10 million in annual revenue. Liquid AI describes d1-omni-600M as an early research release and has not published latency figures for it.

A decision model takes a state and named questions, reads them once and returns a probability for every allowed answer. Liquid AI defines three question types: noul, a yes-or-no question returned as P(yes); choice, one label from named options with a full probability distribution; and score, a probability-weighted position on an ordered rubric of 2 to 10 levels. Several questions can share one state in a single call. Every response reports output_tokens: 0. Recommended uses include routing, moderation, intent classification, reranking, LLM-as-a-judge scoring, agent guardrails and visual inspection. Neither checkpoint is a chat model.

d1-3B has 3.12 billion parameters and starts from LFM2.5-VL-3B, a decoder-only vision-language model. Liquid AI averaged the weights of LFM2.5-2.6B with that model's text backbone, then fine-tuned several checkpoints with different seeds and data mixtures and merged them again. It uses a 400 million parameter SigLIP2 NaFlex vision encoder and a 32,768-token context. The company says long inputs, shuffled answer options and fixing data shortcuts mattered more than advanced techniques.

d1-omni-600M has 587 million parameters and starts from LFM2.5-Encoder-350M, a bidirectional encoder. It comprises a 381 million parameter shared trunk and decision head, a 94 million parameter vision encoder and a 112 million parameter audio encoder. The audio encoder is a 17-layer FastConformer. Context is 16,384 tokens, and audio clips are capped at 30 seconds. A request carries images or audio, never both. Audio training covered English speaker-to-assistant requests only.

Liquid AI presents support and ticket triage as a fit: one d1-3B call can answer a yes-or-no refund check, pick the owning team and rate urgency over the same message with no output tokens to parse. For real-time visual inspection and moderation at the edge, d1-3B reads a 384-pixel image in 35 milliseconds on Jetson AGX Thor, and Liquid AI's Open d1 Arcade runs it frame by frame on live camera input for content moderation and gesture control. For voice-command routing on small devices, d1-omni-600M takes up to 30 seconds of speech alongside text and returns the speaker's intent or topic directly. Its card lists voice-command routing and agent guardrails among recommended uses.

On Decision Index v0.2.1, d1-3B scores 48.57, beating every model under 10 billion parameters and edging Decider 35B-A3B at 47.11. Only Winnow-12B scores higher, at 50.02. Liquid AI ran the official scorer itself, so the d1 scores are not leaderboard submissions. d1-3B leads the Tools category at 74.5 and Arts at 36.3 but trails on Knowledge at 23.8. Across seven public text benchmarks, d1-3B averages 82.9, ahead of Decider 4B at 81.1. d1-omni-600M averages 78.4 and records the top Civil Comments score at 95.8 and PAWS-X at 79.5. On 11 image benchmarks, d1-3B averages 74.1 against 73.9 for its base model. Liquid AI calls audio decision benchmarks an open problem.

Liquid AI measured end-to-end latency with one warm request at a time. One question takes 8 milliseconds on an RTX 4090 and 9 milliseconds on an AMD MI325X. On Jetson, AGX Thor takes 16 milliseconds, AGX Orin takes 26 milliseconds and Orin Nano takes 50 milliseconds. On Jetson AGX Thor, three questions over one state take 20 milliseconds, compared with 16 milliseconds for one question. The RTX 4090 figure uses model.compile in reduce-overhead mode; without it, one question takes 16 milliseconds. NVIDIA's Jetson AI Lab also hosts d1-3B guides.

In the comparison published with the release, d1-3B uses LFM2.5-VL-3B as its base model, while d1-omni-600M uses LFM2.5-Encoder-350M. Decider 4B and Decider 35B-A3B come from Mapika and use Qwen3.5 base models, while Winnow-12B comes from EldanRing and uses Gemma 4 12B IT. Decision Index v0.2.1 scores are 48.57 for d1-3B, 15.95 for d1-omni-600M, 40.70 for Decider 4B, 47.11 for Decider 35B-A3B and 50.02 for Winnow-12B. d1-3B and d1-omni-600M use the LFM Open v1.0 license; Decider and Winnow models use Apache 2.0. Published latency is 8 milliseconds per question on RTX 4090 for d1-3B and not published for d1-omni-600M. Decider 4B is listed at 5.2 milliseconds per three-question request on B300, Decider 35B-A3B at 47 milliseconds per three-question request on B300, and Winnow-12B at 143 milliseconds for a cached four-question request near a 64K context on RTX 5070 Ti.

Editor's Summary Liquid AI has released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return calibrated typed answers with zero output tokens and are available on Hugging Face. The models target real-time routing, moderation, inspection and voice-command tasks on NVIDIA hardware, with d1-3B scoring 48.57 on Decision Index v0.2.1. The release expands the market for non-generative decision models while leaving audio decision benchmarks and d1-omni-600M latency as open questions.