The ALHR method employs hierarchical routing to lower VRAM usage from 57 to 422 MB, achieving 92.1% top-1 accuracy at 1024 tokens with NlogN inference scaling.

The ALHR method employs hierarchical routing to lower VRAM usage from 57 to 422 MB, achieving 92.1% top-1 accuracy at 1024 tokens with NlogN inference scaling.
MC-Sparse is a training-free framework that reduces diffusion transformer latency by selecting individual KV tokens and grouping similar queries, targeting quality loss at high sparsity levels. The…
#AI #DiffusionTransformers #SparseAttention #MachineLearning
https://arxiv.org/abs/2610.06801
SparseEngine is a sparse-first LLM inference engine using a shared lifecycle contract across 15 attention methods, with Chain Cache and Prefix-Cache Pruning for long-context KV management.
#SparseAttention #LLMInference #KVCache #AIInfrastructure
https://arxiv.org/abs/2609.39068
#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...
#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...
DeepSeek‑V4.1‑Flash pushes 1M context with KV‑cache tricks, sparse attention & multimodal tokens. Open weights + vLLM turn it into a transformer playground. Curious? Dive into the details! #DeepSeekV4_1_Flash #1MContext #SparseAttention