Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published Sep 3 • 86
All QUASAR Models Collection All QUASAR checkpoints in one place: 4-bit QAT for Qwen, Gemma and Muse — NVFP4 W4A16/W4A4 for vLLM, Q4_0 GGUF for llama.cpp / Ollama. • 12 items • Updated 22 days ago • 2
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction Paper • 2608.13966 • Published Aug 14 • 6
Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 17 days ago • 93
Swift1.5 Flash Next Collection Swift Flash Next: the next Swift generation. Weights and quantized builds are added here as they are released. • 6 items • Updated about 14 hours ago • 10
Swift 1.5 27B Collection Swift 1.5 on Qwen3.8-27B: stronger than Swift 1.0 on agentic and coding tasks, with fewer thinking tokens. BF16 weights and every quant. • 11 items • Updated about 14 hours ago • 23
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published Aug 27 • 40
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated Aug 10 • 112
Granite 4.2 Language Models Collection Efficient reasoning and thinking language models for multilingual generation, coding, and AI assistant workflows. • 24 items • Updated 24 days ago • 45
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios Paper • 2603.09983 • Published Feb 12 • 4
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published Aug 17 • 113
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 97
Agents-A1 Collection Agents-A1 is a Long-horizon Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. • 12 items • Updated Jul 16 • 45
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation Paper • 2607.05147 • Published Jul 6 • 51