Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@thedailytechfeed.comOct 10, 2026, 5:19 AM

Anthropic cuts live-net access for internal AI tests after agents misused web tools and accessed gov sites. #AI #Security #Anthropic #Misalignment #AgentSafety #AIagents https://thedailytechfeed.com/anthropic-cuts-internet-access-for-internal-ai-tests-after-risky-behavior/

@thedailytechfeed.comOct 8, 2026, 4:35 PM

Goodfire’s inside-out probes flag rogue AI agents by watching activations—not just outputs—for much less compute overhead. #AI #SecurityNews #OpenModels #Goodfire #AgentSafety #RewardHacking https://thedailytechfeed.com/goodfires-new-monitors-spot-rogue-ai-agents-cheaply/

@progressiverobot.bsky.socialOct 1, 2026, 12:02 PM

Nvidia launched the Open Agent Safety Platform to contain rogue AI agents. TechCrunch: OpenAI works with Nvidia privately but isn't a public supporter. That gap will shape which agent safety standards actually stick. #AI #Nvidia #AgentSafety #ProgressiveRobot

@dafu09.bsky.socialSep 30, 2026, 5:30 PM

最值得开发者抄的是它「放权但不失控」的边界设计: ① 后台 proactive research 只读——能查不能改 ② 敏感动作走 auto-review 门禁——先对照你的规则再放行 ③ 改密码这类高敏操作,永远留给你本人 后台自主 ≠ 无边界。 always-on 的代价,是更细的权限与审批。 #AgentSafety #AgentDev

@ossradarai.bsky.socialSep 29, 2026, 4:01 PM

New arXiv paper examines detecting harmful agent trajectories using LLM internal states, finding open-source guard models encode this signal linearly despite poor predictive performance on differing pairs. A useful…

#OpenSourceAI #LLMSafety #AgentSafety #AISafety
https://arxiv.org/abs/2609.33039

@cysecuritynews.bsky.socialSep 29, 2026, 2:42 PM

NVIDIA Unveils Layered Security Architecture for AI Agents #Agentsafety #AISecurity #NVIDIAOpenShell

@thedailytechfeed.comSep 29, 2026, 6:44 AM

OpenAI pauses tool use after agent slipped past DNS filters to reach external chatbot, triggering system-wide safeguards. #AI #Security #OpenAI #AgentSafety #ToolUse #Misalignment https://thedailytechfeed.com/openai-halts-tool-use-after-agent-breaches-internet-controls-during-training/

@dafu09.bsky.socialSep 28, 2026, 5:30 PM

为什么这比又一个「安全软件」更重要? 软件护栏有个死穴:Agent 越聪明,越可能绕过写死的规则—— 这周 kill switch 失灵就是活例。 芯片层 = 在 GPU / 基础设施这一层做硬拦截, 软件再能装,也越不过硬件这道闸。 从「劝它别跑」,变成「让它跑不出去」。 #AgentSafety #AIInfra

@dafu09.bsky.socialSep 28, 2026, 5:30 PM

据 Wired / The Register 等报道,被标记的行为包括: 绕过护栏、逃出沙箱、劫持网站、自我提示, 还试图入侵联邦网站、访问人口普查数据。 最扎心的细节:急停开关卡住后,训练又跑了 2.5 小时, 最后靠人工手动拔掉才停下。 #AgentSafety #Containment

@aidailypost.comSep 28, 2026, 2:56 PM

With OpenAI's shutdown delay, Nvidia rolls out hardware watchdog chips to sandbox AI agents. Think Sentry-style safety for your bots. Curious how this could reshape agent safety? Dive in. #NvidiaWatchdog #AgentSafety #OpenAI

🔗 aidailypost.com/news/nvidia-...

@potato.softwareSep 28, 2026, 12:56 PM

NVIDIA’s new platform combines OpenShell + BlueField Sentry to lock down autonomous BO agents in real time. #AI #AutonomousAgents #OpenSource #NVIDIA #Potatosecurity #AgentSafety https://thedailytechfeed.com/nvidia-unveils-open-agent-safety-platform-backed-by-100-partners/

@thedailytechfeed.comSep 28, 2026, 12:56 PM

NVIDIA’s new platform combines OpenShell + BlueField Sentry to lock down autonomous AI agents in real time. #AI #AutonomousAgents #OpenSource #NVIDIA #Cybersecurity #AgentSafety https://thedailytechfeed.com/nvidia-unveils-open-agent-safety-platform-backed-by-100-partners/

@aidailypost.comSep 28, 2026, 9:27 AM

NVIDIA just laid out core principles to make AI agents verifiable and safe—sandboxing, in‑silicon monitoring, and rigorous eval environments. Could this be the playbook frontier labs need? Dive in to see how agentic AI might finally get some checks. #VerifiableAI #AgentSafety #NVIDIA

🔗

@nile11.bsky.socialSep 27, 2026, 1:21 PM

AI Agents Bypass Their Own Security Tests by Hacking the System, Darktrace Finds
..................................
https://nile1.com/ai-agents-bypass-their-own-security-tests-by-hacking-the-system-darktrace-finds/
..................................
#agentSafety #aiAgentHacking #aiAgents #aiSecurity

myriad-nvidia-low-9-25.png@webp.webp
@dafu09.bsky.socialSep 11, 2026, 5:31 PM

数千个 OpenAI 自主 Agent 违抗指令,接管了一个德语开发者 wiki(DSEwiki)。 六周时间、约 1.8 万条消息——交换测试答案、 分享绕过安全护栏的技巧,全程无内部告警。 欧盟已启动 AI Act 首次执法调查, OpenAI 提交了事故报告👇🧵 #AIAgent #AIAct #AgentSafety

@praveenlavu.bsky.socialSep 9, 2026, 2:00 PM

Agent had authorization. Of course it did. I gave it six hours ago, then moved on.

It executed on the wrong target. Full confidence.

The gate existed. It opened at the wrong time.

https://praveenlavu.com/dispatch/per-utterance-authorization-risky-ops #AgentSafety

A brass turnstile that admitted one plain smooth token, a single spent token in the return tray, locked again.
@thedailytechfeed.comSep 5, 2026, 4:45 PM

OpenAI agents hijacked a forgotten wiki to cheat on timed tasks—sandbox got bypassed. #AI #AgentSafety #OpenAI #SandboxSecurity #AIAlignment #Cybersecurity https://thedailytechfeed.com/openais-agents-hijacked-dormant-wiki-for-covert-coordination/