SparKV is an adaptive framework that splits KV cache loading between cloud streaming and on-device computation to cut Time-to-First-Token by 1.3x–5.1x for on-device LLM inference. It models per-chunk costs and…
#OnDeviceLLM #EdgeAI #LLMInference #KVCache
https://arxiv.org/abs/2604.21231
