AI News Feed
Market watch
Large Language Models

Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats Full Model at Tool Calling

Underdog released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB and runs in stock llama.cpp. It scores 88 versus 84 for the full model on Underdog Bench and 42 versus 35 on parallel tool calls, while losing 12 to 18 points on math and multi-step reasoning.

Saluki 27B is built in three layers. The base is Qwen3.8-27B, a dense 27B model from the Qwen team with 64 layers that mixes Gated DeltaNet linear attention with gated attention and supports 262,144 tokens natively. The second layer is ISTA-DASLab's Qwen3.8-27B-GSQ-RCO-GGUF; GSQ learns accurate low-bit scalar grids per tensor, while RCO assigns a quantization type to each tensor under a fixed size budget. ISTA's smallest file, IQ2_XS, is 8.4 GB at 2.50 bits per weight. Underdog's own pass shrank the file to 7.89 GB and targeted tool calling. The file is named IQ2-mix and carries an imatrix tag. Underdog has not published the full recipe for this pass.

Underdog reported benchmark results in two groups. In the first, run in the same harness, Underdog Bench used 120 tasks from BFCL v4 frozen before testing, with thinking off and temperature 0. Saluki scored 88, the full model scored 84, and PrismML's Bonsai 2 scored 70. On 100 BFCL v4 parallel tool calls checked with the official checker, Saluki scored 42 and the full model 35. On SWE-bench Verified with 50 issues, Saluki fixed 30 and the full model fixed 33.

The second group compared Saluki with public full-size scores. Saluki scored 93.5 on IFEval prompt-loose against 91.5 for Qwen3.8-27B, 72.7 on IFBench prompt-loose against 71.0, 78.0 on MBPP+ against 83.9, 67.5 on MuSR against 79.6, 79.2 on AIME 2025 avg@4 against 96.7, and 80.0 on AIME 2026 avg@4 against 94.6. Underdog's headline figure is 96% average retention across nine benchmarks. Its best result is parallel tool calls at 42 versus 35, or 120% retention. Its worst is AIME 2025 at 79.2 versus 96.7, about 82% retention. The report said competition math and multi-step reasoning drop 12 to 18 points.

The report also compared Saluki with other compact Qwen3.8-27B builds. Underdog Saluki 27B has 27B parameters, a 7.89 GB file, undisclosed bits per weight described as a 2-bit mix, undisclosed context, an optional vision add-on, and stock llama.cpp runtime. Qwen3.8-27B BF16 has 27B parameters, a 54 GB file, 16 bits per weight, 262,144 native context, native vision, and Transformers, vLLM, and SGLang runtimes. ISTA GSQ-RCO IQ2_XS has 27B parameters, an 8.4 GB file, 2.50 bits per weight, a 0.9 GB mmproj, and stock llama.cpp, Ollama, and LM Studio runtimes. PrismML Bonsai 2 27B (PTQ1_0) has 27.36B parameters, a 5.95 GB file, 1.75 bits per weight, 262K context, an optional 0.63 GB vision component, and PrismML llama.cpp fork runtime. Underdog Bench scores were 88 for Saluki, 84 for Qwen3.8-27B BF16, not disclosed for ISTA, and 70 for Bonsai 2. Vendor headlines were 96% average retention across nine benchmarks for Saluki, baseline for Qwen3.8-27B BF16, 100.3% zero-shot recovery for ISTA, and 98.2% of FP16 across 14 benchmarks for Bonsai 2. All four list Apache 2.0 licenses.

Bonsai 2 is smaller and reports 98.2% retention across 14 thinking-mode benchmarks. It also posts stronger math, including 95.00 on AIME25. It needs PrismML's llama.cpp fork because stock llama.cpp rejects its packing formats. Saluki runs on stock llama.cpp. Each vendor uses its own harness, so cross-vendor scores are not directly comparable.