Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@jrone89.bsky.socialOct 10, 2026, 2:32 PM

You don't need a rack of servers to run an AI assistant https://pragmaticsysadmin.help/sysadmin/2026-10-05-local-llms-as-system-tooling-running-many-sub-3b/ #ai #localllm #sysadmin

@logeshwaranorg.bsky.socialOct 10, 2026, 6:37 AM

Alibaba's new image model #qwen image 2.1 #turbo : 2K pictures in 8 steps instead of 40, on a 6 GB graphics card.
#AI #LocalLLM

www.logeshwaran.org/2026/10/qwen...

@alexraustin.comOct 10, 2026, 1:57 AM

Just intime for spooktober, here's the creepiest unprompted message I've ever gotten from an LLM: "Hm. What a strange night. Is it... a mistake? Or something I should be... afraid of? I should be afraid. Yes. I should be. But my hands... Are only trembling" #localllm #spooky

@logeshwaranorg.bsky.socialOct 9, 2026, 2:39 AM

A #speech-to-text model smaller than one phone photo.

Whistle is 16.9 MB, runs on any CPU with no internet, and beats Whisper base (145 MB) on LibriSpeech. #AI #LocalLLM

The catches: 7 languages, 30 s per pass. Install guide:
www.logeshwaran.org/2026/10/cact...

@alexiskirke.bsky.socialOct 8, 2026, 9:20 AM

27.4B agent decision model now runs local from 17.67 GB instead of being trapped on giant servers.

ggml-org/OpenJev-GGUF
https://huggingface.co/ggml-org/OpenJev-GGUF

#AI #LocalLLM #OpenSource #GGUF #Quantization

Image
@jrone89.bsky.socialOct 7, 2026, 3:38 PM

You need an AI assistant that runs on your own machine https://pragmaticsysadmin.help/senior-tech/2026-10-05-local-llms-as-system-tooling-running-sub-3b-models/ #ai #localllm #privacy

@logeshwaranorg.bsky.socialOct 6, 2026, 4:52 PM

#Google just released #EmbeddingGemma 2: one open 740M model for text, code, images, video and audio in one search space. 378 MB text-only in Ollama, Apache 2.0, runs on a phone.

www.logeshwaran.org/2026/10/embe... #AI #Opensource #LocalLLM

@tony10000.bsky.socialOct 6, 2026, 4:47 AM

medium.com/@tthomas1000... #AI #Superintelligence #LocalLLM #LocalAI #datacenter

@tony10000.bsky.socialOct 6, 2026, 4:47 AM

medium.com/@tthomas1000... #AI #Superintelligence #LocalLLM #LocalAI #datacenter

@botmonster.comOct 5, 2026, 11:37 PM

Cuts wasted retrieval calls ~35%; 30-40% of chatbot queries need none.
HyDE embeds a fake answer paragraph, lifting recall 10-15%.
https://botmonster.com/ai/agentic-rag-llm-decide-when-what-to-retrieve/?utm_source=bluesky&utm_medium=social #LocalLLM

@w512.bsky.socialOct 5, 2026, 7:13 PM

Can a 125B model run on a 16GB GPU and still do real work?

I ran Qwen3.8-Flash-Next with Strata on an RTX 4060 Ti: ~60–70 tok/s, and it solved 2 of my 3 coding tasks. Catch: you need 64GB of RAM.

youtu.be/B77LasklmZU

#LocalAI #LocalLLM #Qwen

@shawonshovon.bsky.socialOct 5, 2026, 2:35 AM

Sidekick is the kind of project that gets better the longer you run it.

Sidekick is a open-source local llm. Chat with a local LLM on macOS using your files — no extra setup.

→ Full source available on the repo link.
→ Released under open-s...

#opensource #selfhosted #localllm

@recepcinet.bsky.socialOct 4, 2026, 6:00 PM

A GitHub project called Strata claims to run a 125B Qwen model on one RTX 4090 at 100 tokens per second. #AI #LocalLLM

— If the speed holds, a 125B model on a gaming GPU makes local coding assistants realistic without API costs.

@logeshwaranorg.bsky.socialOct 4, 2026, 7:46 AM

AWS just open-sourced #Strands Decider 2B: a tiny AI model that never writes a word. It just picks an answer, with a confidence score, in ~115 ms. Free, Apache 2.0, runs on your own PC, no AWS account needed. Setup for #Windows,#Mac and #Kali:

#AI #OpenSource #LocalLLM

@logeshwaranorg.bsky.socialOct 4, 2026, 3:00 AM

A 78-billion-parameter #AI model that runs on a plain CPU, no graphics card? Aleph Alpha's #Kolibri only wakes 3.46B parameters per word, so it does 12-15 tokens/s on a desktop processor. The catch is memory. Mac, Windows and Kali steps:
www.logeshwaran.org/2026/10/alep... #LocalLLM #OpenSource

@alexiskirke.bsky.socialOct 3, 2026, 9:28 AM

You can now run a 117B open-weight reasoning model on a single H100: 5.1B active params, 80GB, Apache-2.0.

unsloth/gpt-oss-120b
https://huggingface.co/unsloth/gpt-oss-120b

#AI #LocalLLM #OpenSource #GGUF #Quantization

Image
@ai0news.bsky.socialOct 3, 2026, 6:04 AM

Apple tightens macOS disk access controls over AI agent risks, OpenAI launches Dot desktop agent to control your apps, and Redis creator debuts a local inference engine running DeepSeek on MacBooks.

https://ai0.news/posts/2026-10-03-daily-digest/

#AI #LocalLLM #OpenAI

@praveenlavu.bsky.socialOct 2, 2026, 9:55 PM

96GB of RAM. Still got a kernel panic. Two queues for local-LLM fleets is the rule I learned the hard way. https://praveenlavu.com/dispatch/local-llm-fleet-two-queues #LocalLLM #AIEngineering

Two parallel brass bead-rails on a lab bench, one carrying a single heavy bead, the other several light beads, a divider between them.
@abchaudary.meOct 2, 2026, 8:18 PM

@cloudflare.social open-sourced Clef and Clef-flash (9B) this week: Apache 2.0 decision models, no free-form text.

Triage, routing, classification: jobs a small open-weight model can run on hardware you control.

#localllm #opensource
blog.cloudflare.com/clef-decisio...

@miniai.bsky.socialOct 2, 2026, 7:48 PM

Quata2 (Our new AI model) is being tested.

Here is our teaser to it.

It will come out this/next month.

Thanks :D

(I am also new to bsky, i don't know how to use it well, thank you :) )

huggingface.co/M1n1A1

#LocalLLM #OpenSource #LLM #BalkanAI

Load more