AI News Feed
Market watch
Products & Applications

Liquid AI's d1-3B Tops Decision Index Under 10B as Open Multimodal Decision Models Target Edge

Liquid AI released open d1 decision models for the edge. d1-3B leads the Decision Index under 10B and answers in 16 ms on Nvidia Jetson AGX Thor, while d1-omni-600M adds text, image and audio support.

The models are built on Liquid Foundation Models, or LFMs. Unlike generative models, the d1 decision models do not produce tokens but answer in a single forward pass, the blog said. d1-3B is trained from LFM2.5-VL-3B, described as Liquid AI's latest vision-language model and a decoder-only architecture that accepts text and images. d1-omni-600M is trained from LFM2.5-Encoder-350M, a bidirectional encoder, and adds vision and audio encoders to handle text and images or text and audio. Liquid AI described d1-omni-600M as an early research release that is undergoing further development.

On seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding, d1-3B achieved a mean score of 82.9, the highest in the comparison and above Decider 4B at 81.1. d1-omni-600M scored 78.4, surpassing Decider 2B's 77.1 with one-quarter of the parameters, according to the blog. On individual benchmarks, d1-3B scored 83.3 on SQuAD 2.0, 93.3 on Civil Comments, 86.9 on MASSIVE intent, 68.3 on PubMedQA, 86.3 on BoolQ, 85.6 on XNLI, and 76.4 on PAWS-X. d1-omni-600M scored 74.0, 95.8, 86.1, 61.3, 77.7, 74.7, and 79.5 on the same benchmarks, respectively.

Liquid AI said it validated that d1-3B retains the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks and that d1-omni-600M handles all three modalities. It did not report vision or audio benchmarks, citing that Decision Index v0.3 includes only a private vision split and that audio decision benchmarks are currently an open problem.

In collaboration with NVIDIA, Liquid AI evaluated d1-3B on the NVIDIA stack across the GeForce RTX 4090, Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. On edge devices, d1-3B answered a single question in under 50 ms on every measured device, the blog said. Three questions took only 1.3 times the time of one, with the AGX Thor going from 16 ms to 20 ms. The edge table also listed the Apple M5 Pro at 30 ms for one question and 41 ms for three; Jetson AGX Orin 64 GB at 26 ms and 35 ms; and Jetson Orin Nano at 50 ms and 73 ms. For a 3.4K-token state, times ranged from 220 ms on the AGX Thor to 1,640 ms on the Orin Nano, and for a 384-pixel image, from 35 ms on the AGX Thor to 202 ms on the Orin Nano. Packed throughput for 64 states was 262 per second on the AGX Thor and 38 per second on the Orin Nano.

On GPU, d1-3B answered a question in under 10 ms and processed a 384-pixel image in under 18 ms on both platforms, according to the blog. The RTX 4090 took 8 ms for one question, 21 ms for three, 102 ms for a 3.4K-token state, and 17 ms for a 384-pixel image, with 475 packed states per second. AMD MI325X took 9 ms, 14 ms, 44 ms, and 18 ms, respectively, with 1,106 packed states per second. Liquid AI did not report speed numbers for d1-omni-600M because it is an early research release.

The blog said users should reach for d1 decision models when they need fast, structured decisions, including multimodal inputs. It said d1-3B delivers the highest decision quality at its size, while d1-omni-600M fits where footprint matters. The models require transformers 5.14 or later, along with torch, torchvision, and pillow, and ship their own code, so they should be loaded with trust_remote_code=True. The example code shows several named questions over one text state and an image as the whole state.