Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@gradientbrief.bsky.socialOct 10, 2026, 10:00 AM

DivMoE introduces a fine-grained MoE upcycling framework that addresses routing collapse when experts come from a single source model, outperforming naive fine-grained approaches on Qwen3-1.7B.

#AIresearch #MixtureOfExperts #LLMs #ModelUpcycling
https://arxiv.org/abs/2610.11317

@kuzved1988.bsky.socialOct 9, 2026, 8:43 AM

Технология Mixture of Experts: новый стандарт ИИ для финансов

Архитектура Mixture of Experts (MoE) позволяет финансовым организациям использовать мощь нейросетей с высокой скоростью и точностью за счет разделения…

Читать полностью

#ИскусственныйИнтеллект #ФинансовыеТехнологии #MixtureOfExperts

Технология Mixture of Experts: новый стандарт ИИ для финансов
@donwebmedia.bsky.socialOct 7, 2026, 7:12 AM

Un MoE de 188M en una GPU gratis: qué probó y qué no

¿Se puede armar un modelo MoE desde cero con una GPU gratis? Un dev diseñó 188M parámetros con 51M activos. Mirá qué midió y qué falta confirmar

#moe #deepseekmoe #llm #colab #mixtureofexperts

@aidailypost.comOct 6, 2026, 6:13 PM

🚀 Mistral AI just dropped “Le Chonk” – a 1.05 trillion‑parameter Mixture‑of‑Experts model with vision encoder, 1 M token context, and ready for Nvidia Grace Blackwell. Open‑weight and massive. Dive in to see what this means for AI! #MistralAI #TrillionParams #MixtureOfExperts

🔗

@aidailypost.comOct 5, 2026, 8:38 PM

Reflection AI just dropped its Beam model—an open‑weight, mixture‑of‑experts beast that hits reasoning benchmarks on par with Chinese rivals while using far less inference compute. Curious how they did it? Dive in. #ReflectionAI #BeamModel #MixtureOfExperts

🔗 aidailypost.com/news/reflect...

@cipherpulseai.bsky.socialOct 5, 2026, 4:01 PM

New arXiv work shows that a small set of routed experts in sparse MoE LLMs can be targeted to weaken refusal behavior without retraining, and that router-gradient sensitivity outperforms activation frequency for…

#AISafety #LLM #MixtureOfExperts #CyberSecurity
https://arxiv.org/abs/2610.02910

@hossamudin.bsky.socialOct 4, 2026, 1:19 PM

ألمانيا دخلت سباق النماذج مفتوحة الأوزان… بنموذج حجمه 78 مليار، بس بيشتغل بـ 3.46 مليار بس في كل كلمة! 🇩🇪

في 3 أكتوبر (يوم الوحدة الألمانية)، شركة Aleph Alpha نزّلت Kolibri: نموذج استدلال

#Kolibri #AlephAlpha #MixtureOfExperts #OpenWeights #ذكاء_اصطناعي_عربي #حسام_الدين_حسن #خبير_اونلاين

@michabbb.bsky.socialOct 3, 2026, 3:25 PM

🐦 #AlephAlpha released #Kolibri, an open-weight #LLM for German and English (Apache 2.0). #MixtureOfExperts: 78.1B parameters, only 3.46B active per token. Trained from scratch in Germany and Finland #SovereignAI #opensource #AI

🧵👇 🇩🇪

@aidailypost.comSep 29, 2026, 8:02 AM

🚀 H Company just smashed the OSWorld benchmark—Holo4 27B scored 85.2% at just $0.08 per task! See how its Mixture‑of‑Experts architecture powers GUI and tool‑calling agents. Curious? Dive into the details. #Holo4 #OSWorld #MixtureOfExperts

🔗 aidailypost.com/news/h-compa...

@aidailypost.comSep 14, 2026, 4:54 PM

Cracking the JAX MoE training puzzle: variable expert token counts, routing tricks, and the NVIDIA Transformer Engine boost. See how DeepSeek‑V3 and GB200 handle conditional computation. Dive in for the nitty‑gritty! #JAX #MixtureOfExperts #ConditionalComputation

🔗 aidailypost.com/news/jax-moe...

@globaladvisors.bsky.socialSep 11, 2026, 9:36 AM

#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...

@marcwilson1000.bsky.socialSep 11, 2026, 9:30 AM

#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...