Model Card
- Base model:
Qwen/Qwen3-32B
- Quantization method: SqueezeLLM
- Target bit-width: 4
- Backend kernel: Any-Precision-LLM kernel (
ap-gemv
)
- Calibration data: RedPajama (1024 sentences / 4096 tokens)
- Calibration objective: Next-token prediction
How to run
References
Model tree for jusjinuk/Qwen3-32B-4bit-SqueezeLLM
Collection including
jusjinuk/Qwen3-32B-4bit-SqueezeLLM