Swift-1.5-Qwen3.8-27B-GGUF is a set of GGUF quantizations of Swift 1.5 Qwen3.8-27B, a text-generation and reasoning model derived from Qwen3.8-27B. UkisAI (maintainer profile) trained the Swift model to reduce pathological overthinking: its reported evaluation uses 58.5% fewer mean thinking tokens than the base on GPQA-Diamond, with a 0.31 percentage-point higher score, and the release reports a 9.18× speed-up on several tasks. That speed-up is a reported result, not a general guarantee for every prompt, quantization, or machine. The model has 27B parameters according to its name; the supplied material does not specify its architecture details, hardware requirements, or a context limit for this GGUF release. Run it with a current llama.cpp-compatible runtime such as llama-server. The key decision is whether lower reasoning-token use and local GGUF deployment suit your workload: Swift improves some coding and agent benchmarks, but trails the base on several reasoning and math scores.
Swift does not win every benchmark. It trails Qwen3.8-27B on IFBench (72.07% vs. 73.53%), ERQA (65.40% vs. 67.45%), AIME 2026 (96.00% vs. 98.67%), and HMMT November 2025 (97.33% vs. 99.33%). If your workload prioritizes those tasks, compare both models on representative prompts rather than assuming fewer tokens preserve accuracy.
The reported 9.18× speed-up applies to “several tasks”; the README does not provide a general throughput figure, tokens per second, latency distribution, or hardware details for that claim. The game-building demo is one example: the base took 104.6 minutes and Swift 11.39 minutes. It is not a controlled estimate of performance for other tasks or systems.
The GGUF file sizes range from 8.9 GB to 29.0 GB, but the README gives no VRAM or RAM requirements. File size alone does not establish the memory needed to run a model, especially with a long context or KV cache. The evaluation used context 262,144 for the main BF16 serving setup and 131,072 for Terminal-Bench; those are evaluation settings, not a stated maximum context for this GGUF release.
Quantization can change behavior. The supplied quantized benchmark is a single-seed evaluation on three datasets, and its authors caution that it does not establish quality parity or replace the broader multi-seed results. At the lowest GGUF tiers, the reported divergence from BF16 rises: IQ2_XXS has KLD 0.2769 at 32k and a 78.48% top-p figure, compared with Q8_0 at 0.0006 and 98.85%. These are distribution-comparison metrics, not direct task-accuracy scores.
The model card lists the license as “other” but does not provide the terms in the supplied material. Commercial-use rights therefore cannot be confirmed here. The material also does not specify safety evaluations, known bias findings, fine-tuning instructions, or a maintenance schedule.
Swift-1.5-Qwen3.8-27B-GGUF is derived directly from Swift 1.5 Qwen3.8-27B, itself a reasoning-efficient derivative of Qwen3.8-27B. The training approach identifies tokens associated with pathological overthinking, penalizes them without directly targeting reasoning length, then uses reinforcement learning (RL) and outcome-based preference optimization (OPD) to regain accuracy. Swift 1.5 scales up post-training methods used for Swift 1.0, with emphasis on long-horizon, agentic, and coding tasks. The linked training dataset is described as a source for resampling and constructing RL environments; it is not used out of the box. No dataset size, training-step count, or compute budget is supplied.
The main evaluation reports final aggregate percentages from five repeats under matched protocols. Serving used BF16, vLLM 0.27.1, a Qwen3 parser, context 262,144, and xhigh thinking. Sampling used temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 0, and repetition penalty 1. Benchmarks averaged five seeds (04) per model; Terminal-Bench used five trials per task, with both models served at context 131,072 on the same Harbor build. Output caps were 100,000 tokens for GPQA-Diamond, 16,384 for C-Eval, 81,920 for IFBench, 100,000 for ERQA, 250,000 for AIME 2026, 250,000 for HMMT November 2025, and 32,768 for LiveCodeBench v6. The supplied table gives no numeric output cap for Terminal-Bench.
Reported base-to-Swift scores and mean-token counts:
| Benchmark | Qwen3.8-27B | Swift 1.5 | Mean tokens: base → Swift |
|---|---|---|---|
| GPQA-Diamond | 88.28% | 88.59% | 15,014 → 8,717 |
| C-Eval | 90.00% | 90.92% | 1,492 → 819 |
| IFBench | 73.53% | 72.07% | 8,052 → 4,955 |
| ERQA | 67.45% | 65.40% | 4,137 → 1,906 |
| AIME 2026 | 98.67% | 96.00% | 22,014 → 13,203 |
| HMMT November 2025 | 99.33% | 97.33% | 22,032 → 14,957 |
| LiveCodeBench v6 | 76.76% | 81.71% | 11,184 → 8,448 |
| Terminal-Bench 2.1 | 69.21% | 72.13% | 52,265 → 43,733 |
The README also reports reasoning-effort results: at xhigh, GPQA-Diamond scores were 88.28% for the base and 88.59% for Swift, with 41.9% mean thinking-token reduction; at medium, scores were 84.14% and 82.22%, with 24.8% reduction; at low, scores were 84.04% and 84.85%, with 28.7% reduction.
Available quantization formats include GGUF for llama.cpp; GSQ-RCO GGUF, described as compact 23-bit; AWQ INT4 (W4A16) and AWQ + GPTQ INT4 (W4A16) for vLLM compressed-tensors; AutoRound INT4 (W4A16) for vLLM auto-round; NVFP4 for NVIDIA Blackwell; AMD Quark FP8 (W8A8); and MLX 5-bit, 4-bit, and 3-bit text-only variants for Apple MLX. The GGUF file-size and divergence table is:
| GGUF tier | Size | KLD wikitext @512 | KLD wikitext @32k | 99% KLD @32k | Top-p @32k |
|---|---|---|---|---|---|
| Q8_0 | 29.0 GB | 0.0008 | 0.0006 | 0.005 | 98.85% |
| Q6_K_L | 25.0 GB | 0.0015 | 0.0014 | 0.010 | 98.16% |
| Q6_K | 23.9 GB | 0.0018 | 0.0016 | 0.014 | 98.30% |
| Q6_K_S | 22.9 GB | 0.0020 | 0.0016 | 0.014 | 98.24% |
| Q5_K_M | 20.9 GB | 0.0052 | 0.0061 | 0.050 | 96.92% |
| Q5_K_S | 19.6 GB | 0.0060 | 0.0069 | 0.058 | 96.91% |
| Q4_K_L | 18.8 GB | 0.0106 | 0.0103 | 0.105 | 95.79% |
| Q4_K_M | 17.4 GB | 0.0137 | 0.0134 | 0.163 | 95.03% |
| IQ4_NL | 17.4 GB | 0.0152 | 0.0140 | 0.175 | 95.39% |
| Q4_K_S | 16.4 GB | 0.0164 | 0.0154 | 0.175 | 94.84% |
| IQ4_XS | 15.5 GB | 0.0179 | 0.0173 | 0.187 | 94.96% |
| IQ3_M | 14.9 GB | 0.0410 | 0.0380 | 0.409 | 91.83% |
| Q3_K_L | 14.1 GB | 0.0442 | 0.0410 | 0.415 | 91.63% |
| Q3_K_M | 13.4 GB | 0.0570 | 0.0562 | 0.614 | 90.30% |
| IQ3_XS | 12.8 GB | 0.0583 | 0.0885 | 1.130 | 88.54% |
| Q3_K_S | 12.7 GB | 0.0648 | 0.0658 | 0.712 | 89.53% |
| IQ3_XXS | 12.3 GB | 0.0742 | 0.0844 | 0.996 | 88.70% |
| Q2_K | 10.8 GB | 0.1655 | 0.1546 | 1.728 | 84.00% |
| IQ2_M | 10.5 GB | 0.1523 | 0.1493 | 1.568 | 84.17% |
| IQ2_S | 9.7 GB | 0.2095 | 0.2589 | 3.024 | 80.58% |
| IQ2_XS | 9.1 GB | 0.2433 | 0.2622 | 2.951 | 79.99% |
| IQ2_XXS | 8.9 GB | 0.2866 | 0.2769 | 2.902 | 78.48% |
The KLD figures compare each quantization with the Swift 1.5 BF16 source. The wikitext @512 measure uses 100 windows of 512 tokens from the wikitext-2 test set. The supplied description begins to define wikitext @32k but does not include the rest of that definition.
A separate quantized evaluation used vLLM 0.29.0, tensor parallelism 1, eager execution, BF16 activations, context 131,072, template-default thinking, and seed 0. Sampling used temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0, and repetition penalty 1. It tested 198 GPQA-Diamond questions, 300 IFBench prompts, and 30 AIME 2026 problems, with one sample per prompt and zero request errors. This was a single-seed evaluation, not a replacement for the five-repeat BF16 results. The AMD Quark INT4 and FP8 exports have separate sanity evaluations; completed results on these three reasoning benchmarks are not available.
The model is tagged Text-to-Text and has an image-text-to-text pipeline tag, but the supplied README does not document image input handling for this GGUF package. Do not assume multimodal support from the tag alone. The model card lists the license as other. The repository reports 21,524 downloads.
image-text-to-text pipeline tag, but the README does not document image formats, image preprocessing, or image support in the GGUF runtime.The README recommends a current llama.cpp-compatible runtime such as llama-server, but it does not provide a model-loading command, Python API example, or inference code. A Python snippet would require assumptions about the runtime and chat template, so no copy-pasteable Python example can be confirmed from the