Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Paper • 2608.04074 • Published 22 days ago • 2
L2G-Net: Local to Global Spectral Graph Neural Networks via Cauchy Factorizations Paper • 2602.18837 • Published May 29
Wrapper-Aware Rate-Distortion Optimization in Feature Coding for Machines Paper • 2601.22070 • Published Jan 29 • 1
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Paper • 2608.04074 • Published 22 days ago • 2