OrcaSAQ-2 27B
AvailableOrcaRouter's first published open-weights model, released 2026-09-28 under Apache 2.0. OrcaSAQ-2 27B is a heavily compressed build of Alibaba's open-weight Qwen3.8-27B using OrcaRouter's proprietary sensitivity-aware mixed-precision quantization ("SAQ"): it stores weights at an average 3.21 bits, shrinking the 54 GB BF16 checkpoint to 12.3 GB (~4.4x smaller) so the 27.8B-parameter model fits and serves on a single 16 GB GPU. The vision tower is dropped, so it is text-only (text + code); it keeps the base's 262,144-token context and Qwen3_5ForCausalLM hybrid-attention design (48 Gated DeltaNet + 16 full-attention layers, hidden size 5120, 248,320-token vocab). OrcaRouter reports near-lossless fidelity — WikiText-2 perplexity 5.6482 vs 5.6468 at BF16 (+0.02%), 93.2% top-1 token agreement, 0.031 mean KL divergence — and agentic benchmarks of SWE-bench Verified 70.0% and Terminal-Bench 2.1 58.4%. Throughput is ~65 tok/s single-stream (rising to ~90 tok/s with MTP speculative decoding) and 330+ tok/s across 8-16 concurrent streams on a 16 GB GPU, served via vLLM with an OpenAI-compatible API. Weights are free on Hugging Face (orcarouter/OrcaSAQ-2-27B); the OrcaSAQ2 kernel is on GitHub (Continuum-AI-Corp). Benchmark and fidelity figures are vendor / self-reported.
Specifications
- License
- Open source · Apache 2.0
- Weights
- Downloadable
- Architecture
- hybrid
- Parameters
- 27.8B
- Context window
- 262K tokens
- Max output
- —
- Knowledge cutoff
- —
- Price (in / out, $/M)
- —
- Modalities
- TextCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.