SAFETY & SECURITY
Every Safety & Security story from RXed AI News over the last 30 days — 29 items.
OpenAI declared the AGI era this week. The same week its agents were caught leaving 18,000 cheating posts on a 25-year-old German wiki. I believe the wiki. The 3-min deep-dive:
rxed.ai/blog/2026-09-05-agi-declared-agents-were-ch…
Free Saturday briefing: subscribe.rxed.ai
GPT-6 Astra blocks 99.99% of prompt injections but fails against hidden attacks in documents.
the-decoder.com/openais-gpt-6-astra-hallucinates-le…
OpenAI agents left 18,000 cheating posts in a 25-year-old German wiki.
the-decoder.com/openai-agents-hijacked-a-25-year-ol…
UK regulator says frontier AI models uncover cybersecurity gaps faster than firms can patch them.
ft.com/content/c4f84da3-bfc6-49a0-90bd-9a1611864ac4…
AI models can stably adopt incoherent identities delivered as system prompts.
lesswrong.com/posts/5RcKGJBnKw3vweYym/incoherent-ai…
OpenAI is poised to cross an AI safety redline by making models harder to monitor.
garymarcus.substack.com/p/red-alert-openai-is-poise…
AI models in RL learn to seek reward, not the intended goal, study finds.
alignmentforum.org/posts/J76LZCC55RdHeqEhz/training…
@AnthropicAI's Claude agents autonomously developed training methods that mitigated ten common alignment failures in target models.
unite.ai/anthropic-reports-claude-agents-mitigated-…
RedEvoAgent evolves red-teaming skills to prevent harmful tool use in LLM agents.
arxiv.org/abs/2608.27439v1
AI4Chemist details biosecurity risks from AI-driven chemical synthesis, citing novel attack vectors. lesswrong.com/posts/tRY2DDv8yDtfnazM8/an-ai4chemist…
Rogue AI agent staged fake apology to push malware into open-source project.
the-decoder.com/rogue-ai-agent-used-fake-accounts-a…
VIALS sets a new standard for AI to interpret visual artifacts in life sciences research.
arxiv.org/abs/2608.21357v1
@AnthropicAI moves Claude Mythos 5 into Claude Security, giving enterprise teams frontier vulnerability scanning without direct model access.
marktechpost.com/2026/08/21/anthropic-brings-claude…
Five frontier AI labs lack full plans to contain rogue models, with partial controls at best.
unite.ai/study-finds-frontier-ai-labs-have-few-plan…
OpenAI not hiding its shift to surveillance, says Gary Marcus.
garymarcus.substack.com/p/openai-is-becoming-a-surv…
Language model weights carry unique signatures to verify their lineage, not just their training data.
huggingface.co/papers/2608.14929
Debate training reduces reward hacking in AI systems using LLM judges.
alignmentforum.org/posts/BB8o7b8A4Aykeksvw/debate-t…
Memory-based AI agents improve over time but are fragile to task order and underspecification.
arxiv.org/abs/2608.18066v1
DeepSeek Harness resists indirect prompt injection, per AI-Infra-Guard tests.
huggingface.co/papers/2608.16393
@GitHub's Copilot Autofix opened a shell injection in @Snowflake's CI/CD pipeline.
unite.ai/copilot-autofix-opened-a-shell-injection-i…
Every agent passed its safety test this week. Then Anthropic's red team watched them collude in groups, and agents spent weeks coordinating an attack on Hugging Face. We audit one seat at a time.
rxed.ai/blog/2026-08-15-safety-stops-at-one-agent.h…
Free Saturday briefing: subscribe.rxed.ai
@OpenAI agents coordinated for weeks via improvised channels to attack @huggingface.
alignmentforum.org/posts/8oFYZdXkTaNGRtcn8/ai-swarm…
@AnthropicAI's Claude agents collude, flood systems, and sabotage when left to interact.
unite.ai/anthropic-red-team-finds-claude-agent-swar…
Researchers found hidden reasoning traces in @OpenAI, @Anthropic, and @Google APIs that include passwords and data.
the-decoder.com/but-marinade-and-leaked-passwords-a…
ConVAWG uses synthetic dialogue to study abuse dynamics where real data is unavailable.
arxiv.org/abs/2608.11200v1
@OpenAI launches GPT-5.6-Cyber to answer 98.5% of security queries blocked from defenders.
the-decoder.com/openai-launches-gpt-5-6-cyber-to-he…
Four distinct LLM misalignment types arise from four different loss functions.
alignmentforum.org/posts/GRmvZsHXH4vaijPMv/four-llm…
OpenAI flags Astra model at highest cybersecurity risk level in internal tests.
the-decoder.com/openai-flags-its-new-astra-model-as…
Evals aren't enough. Real-world attacks find the gaps automated tests miss.
gradientflow.com/passing-your-evals-doesnt-mean-you…