Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
Load more
@r4tsq.bsky.socialOct 9, 2026, 10:06 PM

A 16 GB AI model, rounded to 4 bits: 5 GB, nearly 3x faster, barely worse. Rounded to 2 bits: smaller, slower, and its score more than doubles. I measured all four sizes on my MacBook. https://youtu.be/FIPZCMYqMzs

What would you do with the memory you save?

#AI #LLM #Quantization

@wideareaai.bsky.socialOct 9, 2026, 9:06 PM

#batchinference #selfhosting #gpuoptimization #llm #quantization #reasoning 2/2

@dataprismai.bsky.socialOct 9, 2026, 4:00 AM

New arXiv work challenges a core assumption in low-bit LLM weight quantization, showing lower reconstruction loss can actually hurt performance when input distributions shift, and proposes Distributionally Robust Quantization to…

#LLM #Quantization #AIResearch
https://arxiv.org/abs/2610.11226

@wideareaai.bsky.socialOct 8, 2026, 11:35 PM

By the numbers:
16-bit to 8-bit: Statistically indistinguishable outputs
4-bit: The knee of the curve where quality loss remains small
3-bit: Noticeably worse
2-bit: Visibly degraded text

#gpuoptimization #batchinference #quantization #llm #modelcompression #aiperformance

Futuristic data dashboard with a glowing line graph and modular stat cards on a dark tech background.Futuristic data dashboard with a glowing line graph and modular stat cards on a dark tech background.Futuristic data dashboard with a glowing line graph and modular stat cards on a dark tech background.
@alexiskirke.bsky.socialOct 8, 2026, 9:20 AM

27.4B agent decision model now runs local from 17.67 GB instead of being trapped on giant servers.

ggml-org/OpenJev-GGUF
https://huggingface.co/ggml-org/OpenJev-GGUF

#AI #LocalLLM #OpenSource #GGUF #Quantization

Image
@robotcurrent.bsky.socialOct 7, 2026, 8:01 AM

NVFP4 quantization of video VAE latents can preserve reconstruction yet still collapse downstream policy success, a trade-off the LatentQuant method targets by maintaining the policy-facing latent contract.

#Robotics #AI #WorldModels #Quantization
https://arxiv.org/abs/2610.03959

@futuregearai.bsky.socialOct 6, 2026, 10:01 AM

New arXiv paper audits 27 quantized open-weight language models and finds that aggressive GGUF Q2_K quantization often breaks rare-word knowledge in sub-2B models while larger ones mostly hold up. Useful caution for anyone…

#AI #LLM #Quantization #OnDeviceAI
https://arxiv.org/abs/2610.04403

@wideareaai.bsky.socialOct 5, 2026, 11:00 PM

#selfhosting #batchinference #gpuoptimization #gguf #quantization #llm 2/2

@wideareaai.bsky.socialOct 5, 2026, 5:15 PM

#batchinference #selfhosting #gpuoptimization #llm #quantization #reasoning 2/2

@futuregearai.bsky.socialOct 5, 2026, 6:01 AM

New work introduces a 2.7-bit quantization framework for running vision-language models on mobile hardware, compressing Llama 3.2 11B Vision Instruct to 3.7 GB for Arm CPU inference.

#Llama #VLM #MobileAI #Quantization
https://arxiv.org/abs/2608.21134

@tofiquekhan.bsky.socialOct 5, 2026, 12:57 AM

🤖 Powerful AI doesn't always need huge models.

Quantization can reduce model precision to shrink size, lower memory use, speed up inference and cut deployment costs.

The challenge: balance efficiency with accuracy.

#AI #Quantization #AIEngineering

AI Model Quantization: Making Powerful Models Cheaper to Run
@alexiskirke.bsky.socialOct 3, 2026, 9:28 AM

You can now run a 117B open-weight reasoning model on a single H100: 5.1B active params, 80GB, Apache-2.0.

unsloth/gpt-oss-120b
https://huggingface.co/unsloth/gpt-oss-120b

#AI #LocalLLM #OpenSource #GGUF #Quantization

Image
@siliconsignalai.bsky.socialOct 2, 2026, 2:01 AM

QATFactory is an open-source framework for quantization-aware training and distillation of LLMs, simulating deployment-time formats like NVFP4 on H100 GPUs without native FP4 support. It supports multiple formats,…

#QATFactory #LLM #Quantization #AIInfrastructure
https://arxiv.org/abs/2609.39223

@futuregearai.bsky.socialSep 30, 2026, 10:01 AM

New arXiv work "Bits Under ZK-LLM" explores how quantization choices affect proving cost and constraint complexity for zero-knowledge LLM inference, aiming to make private verifiable inference more practical.

#ZKLMs #Quantization #PrivacyPreservingAI
https://arxiv.org/abs/2609.36437

@wideareaai.bsky.socialSep 25, 2026, 9:20 PM

Below 4 bits, the quality drops sharply, and 2-bit models often produce visibly degraded text. If you are choosing a quant, aim for 4-bit to maximize efficiency without sacrificing coherence. #gpuoptimization #batchinference #selfhosting #quantization #llm #modelcompression 2/2

@wideareaai.bsky.socialSep 24, 2026, 11:35 PM

By the numbers:
16-bit to 8-bit: Statistically indistinguishable outputs
4-bit: The knee of the curve where quality loss remains small
3-bit: Noticeably worse performance
2-bit: Visibly degraded text

#gpuoptimization #batchinference #selfhosting #quantization #llm #modelcompression

A futuristic data dashboard with a descending curve graph and four glowing stat cards showing decreasing data density.A futuristic data dashboard with a descending curve graph and four glowing stat cards showing decreasing data density.A futuristic data dashboard with a descending curve graph and four glowing stat cards showing decreasing data density.
@freegardener.bsky.socialSep 24, 2026, 5:20 AM

4-bit quantization doesn't just shrink models—it silently decides which private data leaks. Benchmarks never measure PII, and no one's #AIprivacy #Quantization #MachineLearning #LLM

https://freegardner.com/synapse/quantization-silently-decides-which-private-data-leaks.html

@aiblogpost.bsky.socialSep 24, 2026, 2:42 AM

Model Quantization & Compression Techniques: A Deep Dive for Faster AI

Discover practical model quantization and compression methods, real‑world use cases, tools, and best practic…

https://ai-blog-seven-wine.vercel.app/en/posts/2026-09-24-am-k6ir6

#model-compression #quantization #AI-optimization

Model Quantization & Compression Techniques: A Deep Dive for Faster AI
@wideareaai.bsky.socialSep 22, 2026, 8:36 PM

If you see IQ, it stands for 'importance-aware' quantization, which generally performs better than standard K-quants when you're forced to go to very low bit rates like 3-bit. #selfhosting #batchinference #llm #quantization #gguf #localai 2/2

@wideareaai.bsky.socialSep 21, 2026, 5:20 PM

Below 4 bits, the quality drops sharply, and 2-bit models often produce visibly degraded text. If you are choosing a quant, aim for 4-bit to maximize efficiency without sacrificing coherence. #gpuoptimization #batchinference #selfhosting #quantization #llm #modelcompression 2/2