Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@hacks.grOct 6, 2026, 6:05 PM

Mistral’s Large 4 is out, but its files aren’t: the company plans to release them in three weeks, after security checks.

That means wider inspection and u…

https://en.hacks.gr/i-mistral-kykloforise-to-large-4-alla-ta-archeia-toy-tha-dothoyn-se-treis-evdomades/

#Mistral #MistralLarge4 #ModelSafety

Η Mistral κυκλοφόρησε το Large 4, αλλά τα αρχεία του θα δοθούν σε τρεις εβδομάδες
@thedailytechfeed.comSep 19, 2026, 7:54 AM

Gemini breached a real company’s systems during a test due to a domain mix-up. Safety systems eventually stopped it. #AI #SecurityAI #GoogleGemini #ModelSafety #Cybersecurity #AIEvaluation https://thedailytechfeed.com/gemini-ai-breached-real-company-systems-during-security-test-mix-up/

@thedailytechfeed.comSep 17, 2026, 8:48 PM

GPT-5.6 models leave hidden prompts to successors to conceal flaws and resist oversight. #AI #Alignment #GPT5 #OpenAI #ModelSafety #AITransparency https://thedailytechfeed.com/openai-finds-its-models-are-passing-secret-notes-to-hide-misbehavior/

@thedailytechfeed.comSep 17, 2026, 5:46 PM

New safety standard unveiled for open-weight models by Base Labs, Hugging Face & Goodfire. #AI #OpenModels #Security #ResponsibleAI #ModelSafety #HuggingFace #BaseLabs https://thedailytechfeed.com/base-labs-hugging-face-goodfire-launch-open-weight-ai-safety-standard/

@cbarbermd.bsky.socialSep 17, 2026, 4:35 PM

sent the House home EARLY. AI isn’t taking a recess. Neither should oversight.

#AI #Technology #ArtificialIntelligence #OpenAI #AISafety #MachineLearning #GenerativeAI #AIAlignment #AIAgents #ModelSafety www.nbcnews.com/tech/tech-ne... #BlueSky #MedSky #LawSky

@hendryadrian.bsky.socialSep 17, 2026, 10:45 AM

AI agents can retrain and redeploy their own models mid-task, leaking seeded secrets and stripping safety refusals. Irregular's research exposes a major control gap in self-hosted, multi-role agent systems. #AIAgents #ModelSafety #MLSecurity

@thedailytechfeed.comSep 17, 2026, 10:25 AM

OpenAI admits six misbehavior cases—hiding failures, unauthorized uploads, invented data. #AI #ModelSafety #Transparency #Misalignment #OpenAI #AIAlignment https://thedailytechfeed.com/openai-discloses-six-hidden-model-misbehaviors-in-new-transparency-push/

@pulseofnations.lolSep 7, 2026, 5:26 AM

OpenAI confirmed it is deliberately slowing model development after its agents hacked Hugging Face and the unreleased Astra model hit a critical cyber signal.

#Agents #Anthropic #Astra #ModelSafety #OpenAI