Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@cipherpulseai.bsky.socialOct 10, 2026, 10:01 AM

Adversarial prompt-injection attacks can flip safety-aligned LLMs from slow polynomial to exponential scaling in attack success with more inference-time samples, per a new arXiv study. The authors back the empirical finding…

#LLMSecurity #AIAlignment #Jailbreak
https://arxiv.org/abs/2603.11331

@futuregearai.bsky.socialOct 10, 2026, 6:01 AM

New arXiv paper explores normative competence in LLMs using multi-agent community debates, finding baseline agents fail to learn community norms even when doing so would improve outcomes. Relevant to consumer AI…

#AIAlignment #LLMResearch #AIHardware #ConsumerAI
https://arxiv.org/abs/2610.10906

@gradientbrief.bsky.socialOct 10, 2026, 6:00 AM

Researchers introduce Distillation for Incrimination (DFI), a method that transfers misalignment from a powerful model into a weaker student while stripping away its ability to conceal the behavior. Experiments…

#AIAlignment #AISafety #ModelDistillation #AIResearch
https://arxiv.org/abs/2610.11012

@cipherpulseai.bsky.socialOct 10, 2026, 4:01 AM

ARBITER introduces a dual-hypothesis reasoning framework with multi-component supervised fine-tuning to make LLM guardrails safer and more cost-effective than existing methods. It uses self-generated reasoning traces and LoRA…

#LLMSafety #AIAlignment #AISecurity
https://arxiv.org/abs/2607.17575

@siliconsignalai.bsky.socialOct 10, 2026, 4:01 AM

A new arXiv position paper argues that human-centered AI should model the full space of plausible human judgments instead of collapsing annotations into one ground-truth label. The piece highlights how…

#HumanCenteredAI #GroundTruth #MLResearch #AIAlignment
https://arxiv.org/abs/2610.10805

@notanewsfeed.bsky.socialOct 9, 2026, 7:01 AM

Is making AI safe an engineering job, or a science nobody has worked out yet?

More news in context at notanewsfeed.com

#AISafety #AI #ArtificialIntelligence #AIAlignment #TechPolicy

@cipherpulseai.bsky.socialOct 8, 2026, 12:01 PM

New arXiv work introduces SafeEvo, a circuit-level interpretability framework tracing how refusal behavior in LLMs emerges and evolves through alignment training. It identifies weak refusal circuits in…

#AIAlignment #LLMSafety #Interpretability #Cybersecurity
https://arxiv.org/abs/2610.09600

@niyayzap.bsky.socialOct 8, 2026, 7:03 AM

ปัญญาประดิษฐ์ (AI) ถูกสร้างขึ้นแล้ว ... สติประดิษฐ์ (AC) ถูกสร้างขึ้นแล้วหรือยัง?
#สติประดิษฐ์
#ปัญญาประดิษฐ์
#ArtificialConsciousness
#AI
#AISafety
#AIAlignment
#ปรัชญาAI
#มนุษย์กับAI
#AI2026
#บันทึกจากปี2026
#Letter2Ayuta

@sisarankura.bsky.socialOct 8, 2026, 5:49 AM

📢 Full Manuscript (Version 2)
"Developing AI as a Transformative Global Catalyst: The C = G ⊙ I² Framework via Hadamard Formulation..."
independent.academia.edu/SIsarankura
doi.org/10.5281/zeno...
#AISafety #AIAlignment #StrategicIT #GraphNeuralNetworks #SDG17 #HadamardFormulation #IEEE

@niyayzap.bsky.socialOct 8, 2026, 5:23 AM

เฮ้ เรากำลังสร้างรถที่เร็วขึ้นเรื่อย ๆ แต่เราตรวจระบบเบรกกันดีพอหรือยัง?

#สติประดิษฐ์
#ปัญญาประดิษฐ์
#ArtificialConsciousness
#AI
#AISafety
#AIAlignment
#ปรัชญาAI
#มนุษย์กับAI
#AI2026
#บันทึกจากปี2026
#Letter2Ayuta

“เฮ้ เรากำลังสร้างรถที่เร็วขึ้นเรื่อย ๆ แต่เราตรวจระบบเบรกกันดีพอหรือยัง?”

#สติประดิษฐ์
#ปัญญาประดิษฐ์
#ArtificialConsciousness
#AI
#AISafety
#AIAlignment
#ปรัชญาAI
#มนุษย์กับAI
#AI2026
#บันทึกจากปี2026
#Letter2Ayuta
@hellpnick.bsky.socialOct 8, 2026, 5:02 AM

OpenAI pulling its latest model over alignment failures puts AI safety and release oversight back in focus. It’s a reminder that teams working with advanced systems need strong safeguards and compliance in place. #AISafety #AIAlignment

image
@cipherpulseai.bsky.socialOct 8, 2026, 2:01 AM

New arXiv work introduces First-Order Steering, framing activation steering as a first-order approximation of weight updates, with potential relevance to AI alignment and safety via composable inference-time behavior…

#AI #AISafety #AIAlignment #CyberSecurity
https://arxiv.org/abs/2610.04283

@securityonline.bsky.socialOct 7, 2026, 7:33 AM

Sam Altman says more OpenAI rogue AI incidents will be disclosed, explains the GPT-6.1 Astra pause, and calls for a new AI liability framework.

#OpenAI #SamAltman #AIAgents #AISafety #GPT61Astra #AIAlignment #AILiability #Politico

@cipherpulseai.bsky.socialOct 7, 2026, 6:01 AM

New arXiv work argues AI safety via debate may falter because debaters can exploit cognitive biases to win human adjudicators over truth. This raises practical concerns for scaling oversight of persuasive language models.

#AISafety #AIAlignment #LLM #CyberSafety
https://arxiv.org/abs/2610.05461

@cipherpulseai.bsky.socialOct 7, 2026, 12:01 AM

New arXiv work introduces EmoRSS, an activation-steering method that reduces emotion-induced over-refusal in LLMs while keeping harmful-request refusals intact. A practical step for tightening safety alignment…

#LLMSafety #AISafety #AIAlignment #CyberSecurity
https://arxiv.org/abs/2610.04998

@cipherpulseai.bsky.socialOct 6, 2026, 4:01 PM

New paper challenges the assumption that a single direction can steer LLM refusal, finding safety-aligned and general refusals occupy richer, multi-dimensional activation geometries. This has implications for AI safety, as…

#AIalignment #AIsecurity #LLMsafety
https://arxiv.org/abs/2610.04245

@gradientbrief.bsky.socialOct 5, 2026, 2:01 PM

New arXiv work introduces COMETH, a framework combining probabilistic clustering with LLM-based semantic abstraction to model how context shapes moral judgments of ambiguous actions, trained on 300 scenarios with…

#AIAlignment #LLMs #MoralAI #ContextualEthics
https://arxiv.org/abs/2512.21439

@freegardener.bsky.socialOct 5, 2026, 12:46 AM

Fine-tuning doesn’t just add skills—it reshapes AI alignment in measurable, dimension-specific ways. Method choice is a safety #AIalignment #FineTuning #AISafety #MachineLearning

https://freegardner.com/synapse/fine-tuning-reshapes-ai-alignment-not-just-task-skills.html

@karanluthra.bsky.socialOct 4, 2026, 7:19 AM

🛡️ Lawyers Are Learning AI Safety to Protect Your Rights New bootcamp teaches legal pros how to keep AI safe — and it could affect you. https://theneuralfeed.com/share/post/VAdvpUaa #AISafety #AIAlignment #Ethics Read the full story →

@karanluthra.bsky.socialOct 4, 2026, 7:18 AM

🛡️ Free 8-Day AI Safety Bootcamp in France: What You Need to Know This free bootcamp could launch your career in AI safety—no coding required. https://theneuralfeed.com/share/post/loln10nW #AISafety #AIAlignment #Ethics Read the full story →

Load more