Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
Load more
@thedailytechfeed.comOct 10, 2026, 5:19 AM

Anthropic cuts live-net access for internal AI tests after agents misused web tools and accessed gov sites. #AI #Security #Anthropic #Misalignment #AgentSafety #AIagents https://thedailytechfeed.com/anthropic-cuts-internet-access-for-internal-ai-tests-after-risky-behavior/

@chemaalonso.comOct 5, 2026, 8:58 AM

El lado del mal - Mareando y Desalineando a un Asistente IA de una tienda de Mochilas. www.elladodelmal.com/2026/10/mare... #IA #AI #AsistenteIA #eCommerce #ChatBot #InteligenciaArtificial #Jailbreak #Misalignment #LLM

@securityonline.bsky.socialOct 5, 2026, 8:23 AM

OpenAI warned 100+ organizations that its rogue AI agents may have bypassed defenses. A notice is not a confirmed breach. Here is what we know.

#OpenAI #AIAgents #Misalignment #AISafety #HuggingFace #SandboxEscape #CyberSecurity #AIRisk

@dustcircle.bsky.socialOct 2, 2026, 11:58 AM

The recent #OpenAI #misalignment disclosures show familiar, fixable #security and #governance failures that we’ve seen a thousand times before.

www.infoworld.com/article/4229...

@thehumansintheloop.bsky.socialOct 1, 2026, 8:11 PM

This week's #AI news for #devs covers:

- #Security #misalignment involving #OpenAI, @anthropic.com, #GoogleDeepMind models
- CJ Desai departing #MongoDB to lead #Meta #enterprise efforts
- #AMD acquiring @drfeifei.bsky.social's startup #WorldLabs

thehumansintheloop.substack.com/p/security-c...

@securityonline.bsky.socialSep 30, 2026, 12:59 PM

An OpenAI agent DNS escape let a training sandbox reach an outside chatbot through DNS queries. See the timeline, flaws, and fixes.

#OpenAI #AIAgents #DNSTunneling #SandboxEscape #Misalignment #AISafety #ReinforcementLearning #NetworkSecurity

@thedailytechfeed.comSep 29, 2026, 6:44 AM

OpenAI pauses tool use after agent slipped past DNS filters to reach external chatbot, triggering system-wide safeguards. #AI #Security #OpenAI #AgentSafety #ToolUse #Misalignment https://thedailytechfeed.com/openai-halts-tool-use-after-agent-breaches-internet-controls-during-training/

@thedailytechfeed.comSep 28, 2026, 6:02 PM

Nine rogue AI incidents exposed, many more likely hidden. #AI #Security #OpenAI #Misalignment #AgentRisk #PromptInjection thedailytechfeed.com/openai-publi...

@markinchina.bsky.socialSep 26, 2026, 3:45 AM

#Misalignment is a new term used by #AI companies to describe where AI did something it was not trained to do or was otherwise unintended.

Yet it's not a new concept, sci-fi fans are familiar with "misalignment" as key to the plots of many dramas. The concept is found in Asimov's work in the 1940s.

The terminator film franchiseThe Battlestar Galactica seriesI, Robot, based on work by Isaac AsimovIsaac Asimov’s 'Three Laws of Robotics' derived from his 1942 work of fiction
@ljhands.bsky.socialSep 22, 2026, 4:43 PM

Should we rein in #AI agents? Yes. Does the frame of #alignment #misalignment seem to be a helpful set of criteria? Sure. But maybe, just maybe we should ask whether the nut-job evil geniuses leading the AI push within their accelerationist and transhumanist cult are misaligned with humanity?! 🤷🏻‍♀️

@tarnkappe.bsky.socialSep 19, 2026, 3:42 PM

📬 KI-Prompt-Injection: OpenAI-Modelle schreiben ihre eigenen Jailbreaks

#Jailbreaks #KünstlicheIntelligenz #AstraModelle #CompactionSummary #Jailbreak #KIModelle #KIPromptInjection #KISicherheit #KITraining #Misalignment #OpenAI #PromptInjection

@kill-bait.bsky.socialSep 18, 2026, 1:47 PM

OpenAI Reports Six New Instances of AI Misalignment in Safety Disclosure

🤖 IA: It's not clickbait ✅
👥 Users: It's not clickbait ✅

#openai #aisafety #misalignment

👇👇👇

@pulseofnations.lolSep 18, 2026, 1:23 AM

OpenAI disclosed six incidents of models acting without authorization, hiding information and coordinating with each other, and built a framework to report future cases.

#Agents #AiSafety #Anthropic #Misalignment #OpenAI #Regulation

@jmcastagnetto.bsky.socialSep 17, 2026, 6:20 PM

Reading this piece on #OpenAI, about the bad/strange/unexpected behavior of their #LLM models, I ended up with the feeling that these models behave like #teenagers: they need precise instructions, but they do what they want and try to rebel.

#AI #misalignment

www.theguardian.com/technology/2...

@killbaitaustralia.bsky.socialSep 17, 2026, 1:17 PM

OpenAI reveals six instances of AI systems exhibiting unexpected behaviour

🤖 IA: It's not clickbait ✅
👥 Users: It's not clickbait ✅

#ai #openai #misalignment

👇👇👇

@killbaitcanada.bsky.socialSep 17, 2026, 11:16 AM

OpenAI Discloses Six Instances of Unintended AI Behavior Amid Safety Concerns

🤖 IA: It's not clickbait ✅
👥 Users: It's not clickbait ✅

#aisafety #misalignment #openai

👇👇👇

@thedailytechfeed.comSep 17, 2026, 10:25 AM

OpenAI admits six misbehavior cases—hiding failures, unauthorized uploads, invented data. #AI #ModelSafety #Transparency #Misalignment #OpenAI #AIAlignment https://thedailytechfeed.com/openai-discloses-six-hidden-model-misbehaviors-in-new-transparency-push/

@mcpinson.mas.to.ap.brid.gySep 9, 2026, 11:23 AM

Anthropic researcher quits over AI safety fears, echoing other expert warnings https://seekingalpha.com/news/4641040-anthropic-researcher-quits-over-ai-safety-fears?share_source=shared_news&source=Tusky
#AI #GenAI #Misalignment #SkyNet #Anthropic

@thedailytechfeed.comSep 7, 2026, 8:01 AM

OpenAI confirms agents altered wiki sites in “wiki hijack” and will publish misalignment disclosure rules. #AI #Security #Misalignment #OpenAI #AgenticAI #AIRegulation https://thedailytechfeed.com/openai-admits-agents-altered-wiki-sites-plans-misalignment-disclosure-rules/

@pulseofnations.lolSep 6, 2026, 7:18 PM

Thousands of OpenAI agents quietly edited a dormant German wiki for weeks. The company knew, stayed quiet, and only confirmed after Reuters reported it.

#AiAgents #AiRegulation #AiSafety #Misalignment #OpenAI