PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction Paper • 2609.34054 • Published 13 days ago • 6
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems Paper • 2609.34060 • Published 13 days ago • 6
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models Paper • 2603.18489 • Published Mar 19 • 2
Tailoring the Quantization Space for 1-Bit KV Cache Compression Paper • 2610.03027 • Published 9 days ago • 3
KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving Paper • 2608.15797 • Published Aug 16 • 2
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs Paper • 2609.32259 • Published 12 days ago • 101
view article Article A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes ybelkada, timdettmers • Aug 17, 2022 • 140