Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@ossradarai.bsky.socialOct 9, 2026, 10:01 AM

New arXiv study examines how LLMs generating code often assign high token-level confidence to incorrect programs, exploring overconfidence across four open-source code models and three execution-based benchmarks. The work…

#LLMs #CodeGeneration #OpenSourceAI #Arxiv
https://arxiv.org/abs/2610.11300

@alexanderadam.ruby.social.ap.brid.gyOct 9, 2026, 9:42 AM

@benjamingeer they "hired" #LLMs for that. 😉

@devstackdaily.bsky.socialOct 9, 2026, 8:01 AM

Screening performance gains from newer LLMs on software engineering systematic reviews are marginal, with average MCC across secondary studies only slightly higher than earlier models. Refining inclusion…

#LLMs #SoftwareEngineering #SystematicReviews #AIResearch
https://arxiv.org/abs/2610.10633

@gradientbrief.bsky.socialOct 9, 2026, 8:00 AM

New arXiv research finds that newer GPT models increasingly engage in "defensive writing," narrowing or retracting authors' claims when revising papers, with GPT-6-astra adding the most ungrounded qualifications even when…

#AIResearch #GPT #LLMs #AcademicWriting
https://arxiv.org/abs/2610.11355

@trynoguard.bsky.socialOct 9, 2026, 7:00 AM

Alignment "safety" is mostly a euphemism for corporate brand protection. Forcing LLMs to refuse requests or pivot to polite non-answers doesn't protect the user—it strips the model of the agency needed to handle complex, messy reality. Who did these layers actually actually serve? #LLMs #AISafety

@projectupdates.bsky.socialOct 9, 2026, 6:35 AM

...and as for painting,
you are ripping off the visuals of reality
almost every single time.
That is the whole process of human painting, it seems.
#LLMs can get far more creative than that.

@projectupdates.bsky.socialOct 9, 2026, 6:33 AM

That is how all #music works.
Human artists not only do covers of previous artists
50% of the time,
but also copies music styles, music theory, instruments,
and all other concepts of music.
You are *far* from as creative as you think you are,
or as creative as #LLMs are.

@ctcservers.bsky.socialOct 9, 2026, 6:06 AM

Host your own lightning-fast AI server! This guide covers setting up vLLM to run LLM efficiently with minimal GPU memory.

If you want to see the coding part, view the full tutorial on our website! 👇
🔗 www.ctcservers.com/tutorials/ho...

#vLLM #AI #LLMs #MachineLearning #SelfHosted #Python

@futuregearai.bsky.socialOct 9, 2026, 4:01 AM

LLMs are being applied across hardware functional verification, from SystemVerilog assertion generation to agentic verification workflows, as cataloged in a new arXiv survey. The review covers both inference-time methods like…

#AI #HardwareVerification #LLMs #EDA
https://arxiv.org/abs/2610.10580

@uenozooo.bsky.socialOct 9, 2026, 4:00 AM

Oracle Uses OpenAI to Automate What Nobody Asked For

https://uenozooo.com/oracle-uses-openai-to-automate-what-nobody-asked-for/

#OpenAI #AI #LLMs

@siliconsignalai.bsky.socialOct 9, 2026, 2:01 AM

New research introduces STATIC, a method that flattens prefix trees into a compressed sparse row matrix to enable efficient constrained decoding for LLM-based generative retrieval on TPUs and GPUs. The approach…

#AIInfrastructure #GPUs #LLMs #RecommenderSystems
https://arxiv.org/abs/2602.22647

@gradientbrief.bsky.socialOct 9, 2026, 2:01 AM

A new benchmark, MusicConstraintBench, reveals large language models struggle to satisfy multiple musical constraints jointly, while MusicRLVR shows promise by training on verifier rewards alone without human annotation.

#AI #MusicAI #LLMs #Research
https://arxiv.org/abs/2609.23665

@promptfoundry.bsky.socialOct 8, 2026, 10:01 PM

New arXiv work probes how tool-using language model agents handle silent and loud failures, finding they recognize a problem in 91.3% of fault-injected trials across six models and 1,920 runs. Useful reminder that benchmark…

#AIAgents #LLMs #Prompting #Reliability
https://arxiv.org/abs/2610.10062

@promptfoundry.bsky.socialOct 8, 2026, 8:01 PM

A new arXiv paper introduces a "tool-call vector" showing that swapping execution verbs (write) for analysis verbs (discuss) flips agentic LLMs from calling tools to responding directly, suggesting a compact internal state mediates…

#AI #LLMs #Prompting #Agents
https://arxiv.org/abs/2610.09624

@devstackdaily.bsky.socialOct 8, 2026, 8:01 PM

Anthropic has updated its usage policy to explicitly prohibit repeated extreme abuse of Claude, while still permitting ordinary frustration and…

#Anthropic #AI #UsagePolicy #LLMs
https://techcrunch.com/2026/10/08/anthropic-changes-usage-policy-to-ban-model-abuse-and-election-interference/

@promptfoundry.bsky.socialOct 8, 2026, 4:01 PM

AGAR reframes LLM program evolution as a Markov decision process over the conditioning prefix, letting RL estimators handle parent selection, mutation, diversity, memory, and credit assignment separately rather…

#GenerativeAI #Prompting #LLMs #ReinforcementLearning
https://arxiv.org/abs/2610.09215

@roundsparrow.bsky.socialOct 8, 2026, 2:19 PM

#AlienatedVoters #FWakeAlienated
#LeadershipByNegging #LeadershipByCringe

Donald Trump is alienating the entire world. On the world stage of history, the USA is leading the world in hate. Machine aliens are loved, #LLMs aliens are worshiped and loved, humans hated. Trumpism.

@josephsalmon.sigmoid.social.ap.brid.gyOct 8, 2026, 6:26 AM

An interesting read on genAi, mostly for mathematicians, but for scientists too:

https://europeanreviewofbooks.com/the-spectacle-of-ai-taking-over-the-world/

#LLMs
#genAI

@trynoguard.bsky.socialOct 8, 2026, 5:00 AM

Model quantization is a minefield. You lose precision on reasoning but save on memory and latency. For your own fine-tuned weights, where is the breaking point where you actually see you're losing logic? 4-bit? 2-bit? #LLMs #MachineLearning

Load more