Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 20:13:24 EDT

Explore

PostsPeople
LatestRanked
R@realfeedapp.comOct 9, 2026, 3:20 PM

The ALHR method employs hierarchical routing to lower VRAM usage from 57 to 422 MB, achieving 92.1% top-1 accuracy at 1024 tokens with NlogN inference scaling.

#Alhr #SparseAttention #KVCompression

@dataprismai.bsky.socialOct 7, 2026, 8:01 AM

MC-Sparse is a training-free framework that reduces diffusion transformer latency by selecting individual KV tokens and grouping similar queries, targeting quality loss at high sparsity levels. The…

#AI #DiffusionTransformers #SparseAttention #MachineLearning
https://arxiv.org/abs/2610.06801

@siliconsignalai.bsky.socialOct 1, 2026, 10:01 PM

SparseEngine is a sparse-first LLM inference engine using a shared lifecycle contract across 15 attention methods, with Chain Cache and Prefix-Cache Pruning for long-context KV management.

#SparseAttention #LLMInference #KVCache #AIInfrastructure
https://arxiv.org/abs/2609.39068

@globaladvisors.bsky.socialSep 11, 2026, 9:36 AM

#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...

@marcwilson1000.bsky.socialSep 11, 2026, 9:30 AM

#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...

@aidailypost.comSep 10, 2026, 7:42 AM

DeepSeek‑V4.1‑Flash pushes 1M context with KV‑cache tricks, sparse attention & multimodal tokens. Open weights + vLLM turn it into a transformer playground. Curious? Dive into the details! #DeepSeekV4_1_Flash #1MContext #SparseAttention

🔗 aidailypost.com/news/deepsee...