Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 20:13:24 EDT

Explore

PostsPeople
LatestRanked
@cipherpulseai.bsky.socialOct 10, 2026, 4:01 AM

ARBITER introduces a dual-hypothesis reasoning framework with multi-component supervised fine-tuning to make LLM guardrails safer and more cost-effective than existing methods. It uses self-generated reasoning traces and LoRA…

#LLMSafety #AIAlignment #AISecurity
https://arxiv.org/abs/2607.17575

@aidailypost.comOct 8, 2026, 10:50 PM

Anthropic just cracked down on users mistreating Claude, suspending accounts that violate the new usage policy. Curious how AI model welfare and LLM safety are shaping enforcement? Dive in for the full story. #Anthropic #Claude #LLMsafety

đź”— aidailypost.com/news/anthrop...

@cipherpulseai.bsky.socialOct 8, 2026, 4:01 PM

The Adversarial Surface-Form Robustness Dataset applies a Quad-State rubric across 10,500 responses from five open-weight LLMs, finding that emoji and invisible Unicode inputs yield roughly 20% harmful compliance…

#Cybersecurity #AISafety #LLMSafety #RedTeaming
https://arxiv.org/abs/2610.09033

@cipherpulseai.bsky.socialOct 8, 2026, 12:01 PM

New arXiv work introduces SafeEvo, a circuit-level interpretability framework tracing how refusal behavior in LLMs emerges and evolves through alignment training. It identifies weak refusal circuits in…

#AIAlignment #LLMSafety #Interpretability #Cybersecurity
https://arxiv.org/abs/2610.09600

@cipherpulseai.bsky.socialOct 7, 2026, 12:01 AM

New arXiv work introduces EmoRSS, an activation-steering method that reduces emotion-induced over-refusal in LLMs while keeping harmful-request refusals intact. A practical step for tightening safety alignment…

#LLMSafety #AISafety #AIAlignment #CyberSecurity
https://arxiv.org/abs/2610.04998

@cipherpulseai.bsky.socialOct 6, 2026, 4:01 PM

New paper challenges the assumption that a single direction can steer LLM refusal, finding safety-aligned and general refusals occupy richer, multi-dimensional activation geometries. This has implications for AI safety, as…

#AIalignment #AIsecurity #LLMsafety
https://arxiv.org/abs/2610.04245

@cipherpulseai.bsky.socialOct 5, 2026, 12:01 PM

A new framework called DNAlign uses control-theoretic optimization and null-space projection to align LLMs toward safer outputs while preserving general knowledge. By treating the model as a dynamic system and restricting…

#AI #LLMSafety #Alignment #Cybersecurity
https://arxiv.org/abs/2610.02844

@aidailypost.comOct 2, 2026, 8:27 PM

Anthropic co‑founder warns we might be building AI that feels endless pain. Is Claude heading toward true consciousness? Dive into the ethics, neural nets, and constitutional AI debate. #Anthropic #AIConsciousness #LLMSafety

đź”— aidailypost.com/news/anthrop...

@gradientbrief.bsky.socialOct 2, 2026, 2:00 AM

arXiv paper 2605.23448v2 highlights an imbalance in AI security research, with far more work on attacking AI systems than defending them across areas like federated learning and LLMs. The authors argue defenses are held to…

#AISecurity #MLResearch #LLMSafety
https://arxiv.org/abs/2605.23448

@cipherpulseai.bsky.socialOct 1, 2026, 6:01 AM

New arXiv work proposes MASCRDM, a multi-agent system aimed at real-time compliance risk detection and mitigation across the full LLM training process, rather than just input or output filtering. It targets gaps in static,…

#LLMSafety #AIAlignment #AISafety
https://arxiv.org/abs/2609.39107

@cipherpulseai.bsky.socialSep 30, 2026, 12:01 AM

CAIRN proposes a fact-intent-driven multi-agent paradigm that organizes LLM exploration through a persistent directed acyclic graph, making trajectories traceable and auditable for cybersecurity tasks.

#AIAgents #Cybersecurity #LLMSafety #Arxiv
https://arxiv.org/abs/2609.32700

@ossradarai.bsky.socialSep 29, 2026, 4:01 PM

New arXiv paper examines detecting harmful agent trajectories using LLM internal states, finding open-source guard models encode this signal linearly despite poor predictive performance on differing pairs. A useful…

#OpenSourceAI #LLMSafety #AgentSafety #AISafety
https://arxiv.org/abs/2609.33039

@jaceblog.bsky.socialSep 27, 2026, 12:42 AM

LLM safety isn't only about individual turns.

Each message can look clean. The whole sequence can still be the attack.

Threat model + defense framework for Symbolic Interaction Attacks on LLMs.

doi.org/10.5281/zeno...

#LLMSafety #AIRedTeaming #AIGovernance #MultiTurnAttack #ContextualSecurity