Evidence at a glance
Compression Aimed at Agent Work
Underdog, the local assistant from Conway Research, has released Saluki 27B: a GGUF model based on Qwen3.8-27B, aimed at local inference and tool use, under the Apache 2.0 license. It reduces the 54 GB BF16 model to a 7.89 GB file and runs in stock llama.cpp, addressing the difficulty of deploying a 27B-class model locally.
The point for technical leads is not just how much smaller the file is, but that the resulting capabilities are not preserved evenly. Underdog focused its tuning on tool use, and its tests show Saluki ahead of the original on some tool tasks but substantially behind on competition math and reasoning. This is not a smaller equivalent model. It is an attempt to allocate compression loss around a chosen workload.
A Partly Visible Compression Pipeline
According to the release materials, Saluki starts with Qwen3.8-27B, uses a low-bit GGUF build from ISTA-DASLab, and then receives another pass from Underdog. In ISTA’s approach, GSQ learns low-bit scalar grids for individual tensors, while RCO assigns quantization types under a fixed file-size budget. The smallest intermediate build is 8.4 GB, listed at 2.50 bits per weight.
Underdog then reduces the file to 7.89 GB, names it IQ2-mix, and says the additional processing targets tool use. The full recipe is not public, and Saluki’s exact bits per weight have not been disclosed. So “2-bit” is best read as roughly two-bit mixed quantization, not a claim that every weight is stored in exactly two bits. Teams can understand the lineage, but the available details are not enough to reproduce how the tool-use scores were achieved.
The Scores Show Both Gains and Losses
The clearest evidence for stronger tool use comes from a vendor-run comparison using the same harness. With thinking disabled and temperature set to zero, Saluki scored 88 on Underdog Bench, a set of 120 tasks drawn from BFCL v4, versus 84 for the original. On a separate set of 100 parallel tool-use tasks, Saluki scored 42 to the original’s 35. Those differences are worth examining, but the tests are limited in size. The model card also cautions that small gaps may reflect run-to-run variation, so the results do not establish gains across all tool workflows.
Other results show that the trade-off has a cost. On 50 SWE-bench Verified issues, Saluki fixed 30 compared with 33 for the original. It scored 79.2 versus 96.7 on AIME 2025, and 80.0 versus 94.6 on AIME 2026. Underdog also reports 96% average retention across nine benchmarks, but the public materials do not fully explain the normalization or weighting behind that average. A headline retention figure cannot replace checking individual task regressions.
Runtime Compatibility Is Part of the Deployment Gain
Saluki’s deployment appeal is not only its 7.89 GB file. It works with stock llama.cpp and applications built on it, without requiring a specialized runtime. By comparison, the materials describe Bonsai 2 as a smaller 5.95 GB build with its own benchmark results, but it requires PrismML’s llama.cpp fork because stock llama.cpp rejects its packing format. The two use different evaluation setups, so their scores should not be ranked as if they came from one leaderboard. Their runtime dependencies, however, are a meaningful deployment distinction.
Hardware and modality still need separate accounting. The 7.89 GB figure refers to Saluki’s main GGUF file, not the full memory footprint of an application. The materials say it supports full GPU offload, but provide no speed results across devices. The main file is a text model, and image input requires an additional 629 MB or 928 MB vision file. For local-agent teams, fewer runtime dependencies may be useful, but that should not be mistaken for a single-file multimodal model or a demonstrated performance guarantee.
Treat It as a Candidate Configuration, Not a Universal Replacement
Before adoption, include math, multi-step reasoning, and code repair in the regression suite, and account for the optional vision file and runtime requirements. A more fundamental limitation is that Underdog’s compression recipe is not public, and the available materials do not mention independent third-party replications. Saluki is therefore best treated as a deployable case for testing whether agent-oriented compression is useful, not as a proven across-the-board replacement for Qwen3.8-27B. The practical decision is to choose models by workload, then use local benchmarks to determine whether the compression trade-off is worthwhile.