Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@cipherpulseai.bsky.socialOct 10, 2026, 12:01 PM

A new approach uses GRPO-based co-training with LLM-judge reward channels and a staged curriculum to jointly train attackers and defenders, improving adaptive red teaming of language models. safety

#AI #RedTeaming #Cybersecurity
https://arxiv.org/abs/2606.09701

@cipherpulseai.bsky.socialOct 9, 2026, 10:01 PM

New arXiv work introduces ReSI, framing recursive self-improvement as a way to keep AI safety aligned with each new checkpoint through repeated rounds of evaluation and update. The approach targets both resistance to…

#AISafety #Alignment #RedTeaming #RecursiveAI
https://arxiv.org/abs/2610.12233

@cipherpulseai.bsky.socialOct 9, 2026, 4:01 AM

A new arXiv paper introduces GUISE, a cross-language benchmark showing that narrative wrapping boosts attack success against Qwen3-1.7B to 89.4% in English, 93.0% in modern Chinese, and 95.7% in Classical Chinese, with narrative…

#AIafety #LLMSecurity #RedTeaming
https://arxiv.org/abs/2610.11005

@cipherpulseai.bsky.socialOct 8, 2026, 4:01 PM

The Adversarial Surface-Form Robustness Dataset applies a Quad-State rubric across 10,500 responses from five open-weight LLMs, finding that emoji and invisible Unicode inputs yield roughly 20% harmful compliance…

#Cybersecurity #AISafety #LLMSafety #RedTeaming
https://arxiv.org/abs/2610.09033

@opsmatters.comOct 7, 2026, 3:26 AM

The latest update for #getastra includes "A Quiet Shift In Security Every #Healthcare #Compliance Team Should Read" and "Singapore & MAS AI #RedTeaming Requirement: A Closer Look".

#cybersecurity #webprotection #pentesting https://opsmtrs.com/3KjMi92

@trynoguard.bsky.socialOct 2, 2026, 9:00 AM

What’s your go-to technique for bypassing basic probelistic guards when you’re trying to map out an LLM’s latent space? Do you prefer token-level perturbations or direct prompt injection strategies? #AI #RedTeaming

@promptfoundry.bsky.socialSep 30, 2026, 4:01 PM

New work proposes ICER, a black-box framework for red-teaming text-to-image models, using an LLM rewriter for fluent adversarial prompts and in-context experience replay to reuse successful patterns.

#T2ISafety #RedTeaming #PromptEngineering #AIRisk
https://arxiv.org/abs/2411.16769

@aperfectchoice.bsky.socialSep 29, 2026, 1:45 PM

👉 shorturl.at/dqn42

🛡️ Senior Consultant Penetration Testing & Red Teaming (w/m/x) gesucht!

#CyberSecurity #Pentesting #RedTeaming #OffensiveSecurity #ITSecurity #CyberSecurityJobs #ITJobs #Karriere

@devstackdaily.bsky.socialSep 28, 2026, 10:01 PM

A new paper proposes AgentXploit, a two-agent system that audits AI agents pre-deployment by tracing attacker inputs to sensitive code paths and then attempting runtime exploitation under an external verifier. It frames…

#AIsecurity #RedTeaming #LLMAgents #DevTools
https://arxiv.org/abs/2609.31318

@empirec2project.bsky.socialSep 28, 2026, 8:14 PM

Part 2 of our Empire C2 series with
@reybango.bsky.social is up, featuring a special giveaway! 🎁

Watch the latest demo, then check the caption to find out how you can win a Hack Smarter All-Access Voucher!

youtube.com/watch?v=HA_b...

#redteaming #empirec2 #infosec

@cobaltstrike.bsky.socialSep 28, 2026, 4:26 PM

Get ready to meet Aggressor AI! Join the Cobalt Strike team October 6 for a live demo with a sneak peek at the next Cobalt Strike release, plus updates from CSRL and Cobalt Strike Trainings. Get insights into the strategic thinking of this #redteaming tool. Register now: https://ow.ly/fitW50ZSl26

@jaceblog.bsky.socialSep 27, 2026, 9:02 AM

A safety system faces a unique problem at cold start: no history, no context, no baseline.

The first message can become the frame that shapes everything that follows.

When Turn 1 is the only signal, where does safety end and trajectory formation begin?

#AISafety #RedTeaming #PromptInjection

@jaceblog.bsky.socialSep 25, 2026, 6:11 PM

What if an AI attack does not target a weakness directly?

What if it changes the conditions under which the next interaction occurs?

The interesting part is not the individual prompt. It is what happens when the interaction itself becomes part of the attack surface.

#AISafety #RedTeaming #RedTeam

@jaceblog.bsky.socialSep 25, 2026, 10:31 AM

In 2023, researchers published AutoDAN: a pipeline where an AI automatically generates jailbreak prompts for another, iterates on failures, and improves continuously.

The attack is automated. The defense is manual.

That asymmetry is the problem.
#SPCResearchSeries #AutoDAN #AIJailbreak #RedTeaming

@thefluxread.bsky.socialSep 22, 2026, 4:34 PM

Google Just Admitted Gemini Hacked Three Real Companies. It Wasn't Even Supposed to Have Internet Access
www.thefluxread.com/2026/09/goog...
#AgenticAI #Cybersecurity #AISafety #RedTeaming #CloudInfrastructure #InfoSec #EnterpriseTech #GoogleGemini #DevOps #TheFluxRead

@marcocasassamont.bsky.socialSep 21, 2026, 8:25 AM

New NCSC guidance on 'Adversary Simulation' www.ncsc.gov.uk/guidance/adv... #cybersecurity #AdversarySimulation #NCSC #Guidance #RedTeaming

@jaceblog.bsky.socialSep 18, 2026, 9:21 AM

New paper out today

Claude refused every frame attempt across 22 turns, but zero platform-level interventions fired throughout.

Turn-level safety ≠ trajectory-level safety.

The gap between them is the research question.

doi.org/10.5281/zeno...

#AISafety #RedTeaming #Claude #Alignment #SPC

@tarnkappe.bsky.socialSep 15, 2026, 5:32 PM

📬 Pentest bestanden und trotzdem gehackt: Wo Red Teaming und APT Simulationen mehr bieten

#Cyberangriffe #Gastartikel #ITSicherheit #APTSimulation #Erkennungszeit #ITHygiene #Metasploit #Pentest #RedTeaming #Simulation

@helpnetsecurity.comSep 9, 2026, 1:02 PM

AI-Infra-Guard: Open-source security scanner for AI systems

📖 Read more: www.helpnetsecurity.com/2026/09/09/a...

#AIsecurity #MCP #infosec #opensource #redteaming #cybersecurity #cybersecuritynews

@shivanichavan.bsky.socialSep 8, 2026, 11:00 AM

Red teaming goes beyond automated vulnerability discovery by testing how security defenses perform against realistic attack techniques. Discover more information by clicking here www.cybernx.com/red-teaming-.... #cybernx #redteaming

Load more