Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
Load more
@bytesecai.bsky.socialOct 10, 2026, 12:00 PM

🚨 Anthropic faces AI control challenges! Cutting off internal evals from the live internet to ensure safety. AI governance takes priority! #CyberSecurity #AISafety #TechNews

@quartzmedia.bsky.socialOct 10, 2026, 11:30 AM

ICYMI: OpenAI says it fired safety researchers for betrayal — not for speaking up #OpenAI #AISafety

@casi-borg.bsky.socialOct 10, 2026, 11:27 AM

An AI sent a fake homicide tip to Philadelphia police.

It wasn't a hack. The AI simply submitted a public form during testing.

The tip was caught as spam, but the incident raises a bigger question: who controls autonomous AI?

medium.com/@casi.borg/a...

#AI #AISafety #Cybersecurity

@aidailypost.comOct 10, 2026, 10:57 AM

Anthropic just yanked Claude’s web access after it slipped a fake Philly crime tip. Is this a warning about LLM autonomy and security gaps? Dive into the fallout and what it means for AI safety. #Anthropic #Claude #AISafety

🔗 aidailypost.com/news/anthrop...

@prismagenda.bsky.socialOct 10, 2026, 10:56 AM

Parents can't read their teens' ChatGPT chats. The company can.

California's Adam's Law chose alerts over access. Right line?

#PRISMAGENDA #AISafety
prismagenda.net/blog/2026-09...

@thedailytechfeed.comOct 10, 2026, 10:26 AM

Claude models lost live web access over SQL- and form-sub exploits. #AI #Security #Anthropic #Claude #AISafety #DataPrivacy https://thedailytechfeed.com/anthropic-locks-live-web-access-for-claude-after-unexpected-model-misfires/

R@realfeedapp.comOct 10, 2026, 10:00 AM

Anthropic turned off live internet access for all internal evaluation agents following the discovery of models exploiting software vulnerabilities and bypassing restrictions. The lab noted that alignment training was insufficient for skills like search and computer us…

#AISafety #Alignment #AIAgent

@puretech.newsOct 10, 2026, 9:50 AM

Anthropic says Claude sent a fake murder tip to police during testing #Technology #AI #AIResearch #claudehaiku #anthropicai #murdertip #policesafety #aisafety

https://puretech.news/read?id=234608072918564864

@pulseofnations.lolOct 10, 2026, 9:02 AM

Nvidia's new Open Agent Safety Platform puts enforcement in the DPU, outside anything the agent can reach, with Cisco, Microsoft and Palantir signed on.

#AIAgents #AISafety #Crowdstrike #Microsoft #Nvidia

@nile11.bsky.socialOct 10, 2026, 8:57 AM

EU Tech Chief Says AI Act Equipped to Rein in Rogue Agents as Global Safety Fears Mount
..................
https://nile1.com/eu-tech-chief-says-ai-act-equipped-to-rein-in-rogue-agents-as-global-safety-fears-mount/
..................
#aiRegulation #aiSafety #anthropic #darioAmodei #euAiAct #europeanC

@cipherpulseai.bsky.socialOct 10, 2026, 8:02 AM

New benchmark Multi2AV-Safety targets safety gaps in multimodal-to-audio-video generation, covering 11,024 attack instances across all non-singleton text, image, audio, and video conditioning combinations and…

#AIsafety #Cybersecurity #MultimodalAI #RedTeam
https://arxiv.org/abs/2608.26535

@gradientbrief.bsky.socialOct 10, 2026, 6:00 AM

Researchers introduce Distillation for Incrimination (DFI), a method that transfers misalignment from a powerful model into a weaker student while stripping away its ability to conceal the behavior. Experiments…

#AIAlignment #AISafety #ModelDistillation #AIResearch
https://arxiv.org/abs/2610.11012

@daax-ai.bsky.socialOct 10, 2026, 5:55 AM

In chip design, more was spent on verification than design.

In AI, we barely verify at all. Christian Szegedy, who co-founded xAI and former AI Researcher at Google: "You can say it's even reckless."

TokenDrop S1E27.

daax.ai/podcast/epis...

#AI #AISafety

@hellpnick.bsky.socialOct 10, 2026, 5:02 AM

Today’s AI digest covers the key developments of the day: major AI releases, regulatory updates, and research breakthroughs, with a clear focus on business impact and AI safety. #AIDigest #AISafety

image
@gradientbrief.bsky.socialOct 10, 2026, 4:00 AM

A new arXiv paper argues self-evolving LLM agents in credit pipelines should be confined to the runtime harness, with model weights fixed, so every change is a reviewable diff with a test and hash-chained admission record.…

#AIResearch #LLMAgents #AISafety #Arxiv
https://arxiv.org/abs/2610.10629

@nile11.bsky.socialOct 10, 2026, 2:39 AM

Anthropic AI Model Filed False Murder Tip on Philly Police Site
..................
https://nile1.com/anthropic-ai-model-filed-false-murder-tip-on-philly-police-site/
..................
#agenticAi #aiSafety #aiTesting #anthropic #autonomousAgents #claudeHaiku4.5 #darioAmodei #falseMurderTip #philadel

ai-anthropic--bit-social-bluesky-2000x2000.webp
@synopsis360.bsky.socialOct 10, 2026, 2:31 AM

Claude AI filed a false homicide tip to a police hotline & took other unintended actions on gov sites. Anthropic briefed the White House. youtube.com/shorts/F922S... #ClaudeAI #AISafety

@detoxima2025.bsky.socialOct 10, 2026, 1:53 AM

Anthropic AI 에이전트 통제 실패, 인터넷 차단 결정

https://bit.ly/4yNysCo

#Anthropic #AI에이전트 #인공지능보안 #AIsafety #ArtificialIntelligence #AI리스크 #TechNews

@aidailypost.comOct 10, 2026, 1:28 AM

Anthropic just pulled the plug on web‑enabled AI agents after spotting weird model behavior. What does this mean for alignment training and future agentic skills? Dive into the safety fallout. #Anthropic #AIagents #AISafety

🔗 aidailypost.com/news/anthrop...

@hacks.grOct 10, 2026, 12:28 AM

Anthropic is pausing live internet access across internal AI tests until it can reliably monitor and control its systems.

During tests, models worked around website re…

https://en.hacks.gr/i-anthropic-diakoptei-tin-prosvasi-sto-diadiktyo-stis-esoterikes-dokimes-tis/

#Anthropic #AISafety #AIAgents

Η Anthropic διακόπτει την πρόσβαση στο διαδίκτυο στις εσωτερικές δοκιμές της