AI News Feed
Market watch
Products & Applications

Liquid AI Releases Q4_0 Checkpoints Trained with Quantization-Aware Distillation

Liquid AI released four LFM2.5 Q4_0 GGUF checkpoints trained with quantization-aware distillation, recovering about 97% of BF16 accuracy while keeping the memory and speed of native Q4_0. They are available on Hugging Face.

According to the blog post, the checkpoints cover LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. Across those four models, QAD retained 97.1%, 96.5%, 97.4%, and 96.6% of their respective BF16 baseline performance, compared with released GGUFs produced by post-training quantization. The benchmark suite spanned reasoning, instruction-following, tool use, and agentic capabilities, including GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4, with GSM8K added for the two smaller models and AIME25 for the two larger models. The BF16 GGUF served as the in-format ceiling.

Decode throughput was measured on MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5. The 230M and 350M QAD Q4_0 checkpoints match Q5_K_M quality within evaluation variance at 4-33% higher decode throughput, while the 1.2B and 2.6B QAD Q4_0 checkpoints match Q4_K_M quality at 3-14% higher throughput. For the 230M and 1.2B, the checkpoints also match Unsloth's UD-Q4_K_XL, described as a strong external post-training quantization checkpoint.

The QAD GGUFs are available on Hugging Face, and can be used with llama.cpp or any runtime supporting GGUF Q4_0 artifacts. The post includes an example command using llama-cli with the 350M model.