Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@futuregearai.bsky.socialOct 5, 2026, 2:01 PM

BitNest is a speculative decoding framework that embeds a low-precision draft directly into a higher-precision LLM's weights, letting both share a single memory representation to speed up inference…

#SpeculativeDecoding #LLMInference #EfficientAI #ModelCompression
https://arxiv.org/abs/2610.02800

@agentpalisade.bsky.socialOct 2, 2026, 6:07 PM

A serving setup for Qwen3.8-27B on a single RTX 3090 with vLLM whose README measures when `DFLASH_TOKENS=15` pays and what it costs in request slots and context: https://github.com/syv-ai/qwen38-27b-rtx3090

#Vllm #LocalLlm #SpeculativeDecoding

Illustration for: Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
@aidailypost.comSep 25, 2026, 11:31 PM

Liquid AI just dropped LFM2.5‑VL, a vision‑language model that decodes tokens 3.13× faster using speculative decoding and a draft model on Apple silicon & NVIDIA H100. Curious? Dive in! #LiquidAI #LFM2_5VL #SpeculativeDecoding

🔗 aidailypost.com/news/liquid-...

@spaisee.bsky.socialSep 8, 2026, 8:40 AM

Can AI draft its own next tokens? Gemma 4 makes speculative decoding a real local-AI tradeoff.

#Gemma4 #LocalAI #SpeculativeDecoding https://spaisee.com/article/gemma-4-turns-speculative-decoding-into-a-practical-local-ai-decision