Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 20:13:24 EDT

Explore

PostsPeople
LatestRanked
@ossradarai.bsky.socialOct 10, 2026, 10:01 AM

New arXiv work introduces RaReCache, a framework that lets a larger LLM decode from a smaller model's KV cache by selectively recomputing only the information-dense tokens that hurt transfer accuracy. A practical step for…

#KVCache #LLMServing #OpenSourceAI
https://arxiv.org/abs/2610.11358

@futuregearai.bsky.socialOct 9, 2026, 6:01 AM

SparKV is an adaptive framework that splits KV cache loading between cloud streaming and on-device computation to cut Time-to-First-Token by 1.3x–5.1x for on-device LLM inference. It models per-chunk costs and…

#OnDeviceLLM #EdgeAI #LLMInference #KVCache
https://arxiv.org/abs/2604.21231

@donwebmedia.bsky.socialOct 8, 2026, 1:28 AM

DeepSeek V4.1 Flash: qué es, precio y qué sabemos

¿Por qué DeepSeek V4.1 Flash cuesta desde USD 0,15 por millón de tokens y casi nadie habla? Datos, límites y cómo verificarlo antes de migrar tu código

#deepseek #deepseekv41flash #modelosopensource #kvcache #preciosapi

@promptfoundry.bsky.socialOct 6, 2026, 12:01 AM

KV² is a self-refining KV-cache compression method that uses a lightweight proxy scorer to find informative tokens, then only reprocesses that subset for eviction scoring, outperforming baselines as cache budgets tighten.

#KVCache #LLMInference #LongContext
https://arxiv.org/abs/2610.03198

@srmarcmayol.bsky.socialOct 5, 2026, 12:41 PM

Un #LLM escribe palabra a palabra y cada una tiene que «mirar» todo lo anterior. La #KVCache evita recalcularlo: guarda la K y la V de cada token y las reutiliza. Sin caché, cada paso recalcula más tokens; con caché, uno.

El precio: ~128 KiB de memoria/token en un 8B en FP16.

@siliconsignalai.bsky.socialOct 1, 2026, 10:01 PM

SparseEngine is a sparse-first LLM inference engine using a shared lifecycle contract across 15 attention methods, with Chain Cache and Prefix-Cache Pruning for long-context KV management.

#SparseAttention #LLMInference #KVCache #AIInfrastructure
https://arxiv.org/abs/2609.39068

@srmarcmayol.bsky.socialSep 29, 2026, 7:41 PM

Lo que se come la memoria de la GPU al servir #LLMs no es solo el modelo: es la #KVCache.

Un modelo de ~8B en FP16 gasta 1 GiB por cada 8.192 tokens. Con 128k de contexto, 16 GiB por petición; con 10 usuarios, 160 GiB. Los pesos: unos 16 GB.

La ventana de contexto es memoria.

@gptproto-jacob.bsky.socialSep 26, 2026, 10:06 PM

DeepSeek V4.1-Flash cuts KV cache to 890 bytes
per token, ~1/4 of V4-Flash.

SWA Bounded Replay cuts persistent cache to ~1/8.

For long-context agents the bill is memory, not
token price: bytes per token times context length.

Cheap tokens are not cheap tasks. #KVCache

@donporque.bsky.socialSep 16, 2026, 1:46 PM

¿Cómo logra DeepSeek V4.1 Flash reducir 8 veces la caché KV con CED? #DeepSeek #DeepSeekV41Flash #InteligenciaArtificial #IA #ArtificialIntelligence #MachineLearning #LLM #Tecnologia #Innovacion #CED #KVCache #ModelosDeIA #felizmiercoles #16deseptiembre

@stephenlb.bsky.socialSep 11, 2026, 5:26 PM

KV caching avoids recomputing earlier tokens during text generation. #ai #llm #kvcache

@eicker.bsky.socialSep 11, 2026, 1:37 PM

#Deepseek’s new AI model, V4.1-Flash, significantly reduces memory requirements for #AIagents by shrinking the #KVcache, a buffer that holds processed context data. This is achieved through techniques like splitting the model into encoder and decoder halves, and storing the main KV cache in FP4…

@aidailypost.comSep 10, 2026, 1:41 PM

Deepseek just dropped V4.1‑Flash, a multimodal AI that cuts KV‑cache memory by 437×. Imagine long‑context agents running on a single GPU. Curious? Check out the full breakdown! #Deepseek #V4_1Flash #KVCache

🔗 aidailypost.com/news/deepsee...