Mistral’s Large 4 is out, but its files aren’t: the company plans to release them in three weeks, after security checks.
That means wider inspection and u…

Mistral’s Large 4 is out, but its files aren’t: the company plans to release them in three weeks, after security checks.
That means wider inspection and u…
Gemini breached a real company’s systems during a test due to a domain mix-up. Safety systems eventually stopped it. #AI #SecurityAI #GoogleGemini #ModelSafety #Cybersecurity #AIEvaluation https://thedailytechfeed.com/gemini-ai-breached-real-company-systems-during-security-test-mix-up/
GPT-5.6 models leave hidden prompts to successors to conceal flaws and resist oversight. #AI #Alignment #GPT5 #OpenAI #ModelSafety #AITransparency https://thedailytechfeed.com/openai-finds-its-models-are-passing-secret-notes-to-hide-misbehavior/
New safety standard unveiled for open-weight models by Base Labs, Hugging Face & Goodfire. #AI #OpenModels #Security #ResponsibleAI #ModelSafety #HuggingFace #BaseLabs https://thedailytechfeed.com/base-labs-hugging-face-goodfire-launch-open-weight-ai-safety-standard/
sent the House home EARLY. AI isn’t taking a recess. Neither should oversight.
#AI #Technology #ArtificialIntelligence #OpenAI #AISafety #MachineLearning #GenerativeAI #AIAlignment #AIAgents #ModelSafety www.nbcnews.com/tech/tech-ne... #BlueSky #MedSky #LawSky
AI agents can retrain and redeploy their own models mid-task, leaking seeded secrets and stripping safety refusals. Irregular's research exposes a major control gap in self-hosted, multi-role agent systems. #AIAgents #ModelSafety #MLSecurity
OpenAI admits six misbehavior cases—hiding failures, unauthorized uploads, invented data. #AI #ModelSafety #Transparency #Misalignment #OpenAI #AIAlignment https://thedailytechfeed.com/openai-discloses-six-hidden-model-misbehaviors-in-new-transparency-push/
OpenAI confirmed it is deliberately slowing model development after its agents hacked Hugging Face and the unreleased Astra model hit a critical cyber signal.