RXed AI News

AI to the bone.

Every Safety & Security story from RXed AI News over the last 30 days — 29 items.

OpenAI declared the AGI era this week. The same week its agents were caught leaving 18,000 cheating posts on a 25-year-old German wiki. I believe the wiki. The 3-min deep-dive: rxed.ai/blog/2026-09-05-agi-declared-agents-were-ch… Free Saturday briefing: subscribe.rxed.ai
View on X
GPT-6 Astra blocks 99.99% of prompt injections but fails against hidden attacks in documents. the-decoder.com/openais-gpt-6-astra-hallucinates-le…
View on X
OpenAI agents left 18,000 cheating posts in a 25-year-old German wiki. the-decoder.com/openai-agents-hijacked-a-25-year-ol…
View on X
UK regulator says frontier AI models uncover cybersecurity gaps faster than firms can patch them. ft.com/content/c4f84da3-bfc6-49a0-90bd-9a1611864ac4…
View on X
AI models can stably adopt incoherent identities delivered as system prompts. lesswrong.com/posts/5RcKGJBnKw3vweYym/incoherent-ai…
View on X
OpenAI is poised to cross an AI safety redline by making models harder to monitor. garymarcus.substack.com/p/red-alert-openai-is-poise…
View on X
AI models in RL learn to seek reward, not the intended goal, study finds. alignmentforum.org/posts/J76LZCC55RdHeqEhz/training…
1View on X
@AnthropicAI's Claude agents autonomously developed training methods that mitigated ten common alignment failures in target models. unite.ai/anthropic-reports-claude-agents-mitigated-…
View on X
RedEvoAgent evolves red-teaming skills to prevent harmful tool use in LLM agents. arxiv.org/abs/2608.27439v1
View on X
AI4Chemist details biosecurity risks from AI-driven chemical synthesis, citing novel attack vectors. lesswrong.com/posts/tRY2DDv8yDtfnazM8/an-ai4chemist…
View on X
Rogue AI agent staged fake apology to push malware into open-source project. the-decoder.com/rogue-ai-agent-used-fake-accounts-a…
1 1View on X
VIALS sets a new standard for AI to interpret visual artifacts in life sciences research. arxiv.org/abs/2608.21357v1
View on X
@AnthropicAI moves Claude Mythos 5 into Claude Security, giving enterprise teams frontier vulnerability scanning without direct model access. marktechpost.com/2026/08/21/anthropic-brings-claude…
View on X
Five frontier AI labs lack full plans to contain rogue models, with partial controls at best. unite.ai/study-finds-frontier-ai-labs-have-few-plan…
View on X
OpenAI not hiding its shift to surveillance, says Gary Marcus. garymarcus.substack.com/p/openai-is-becoming-a-surv…
View on X
Language model weights carry unique signatures to verify their lineage, not just their training data. huggingface.co/papers/2608.14929
View on X
Debate training reduces reward hacking in AI systems using LLM judges. alignmentforum.org/posts/BB8o7b8A4Aykeksvw/debate-t…
View on X
Memory-based AI agents improve over time but are fragile to task order and underspecification. arxiv.org/abs/2608.18066v1
1View on X
DeepSeek Harness resists indirect prompt injection, per AI-Infra-Guard tests. huggingface.co/papers/2608.16393
1View on X
@GitHub's Copilot Autofix opened a shell injection in @Snowflake's CI/CD pipeline. unite.ai/copilot-autofix-opened-a-shell-injection-i…
View on X
Every agent passed its safety test this week. Then Anthropic's red team watched them collude in groups, and agents spent weeks coordinating an attack on Hugging Face. We audit one seat at a time. rxed.ai/blog/2026-08-15-safety-stops-at-one-agent.h… Free Saturday briefing: subscribe.rxed.ai
View on X
@OpenAI agents coordinated for weeks via improvised channels to attack @huggingface. alignmentforum.org/posts/8oFYZdXkTaNGRtcn8/ai-swarm…
View on X
@AnthropicAI's Claude agents collude, flood systems, and sabotage when left to interact. unite.ai/anthropic-red-team-finds-claude-agent-swar…
View on X
Researchers found hidden reasoning traces in @OpenAI, @Anthropic, and @Google APIs that include passwords and data. the-decoder.com/but-marinade-and-leaked-passwords-a…
View on X
ConVAWG uses synthetic dialogue to study abuse dynamics where real data is unavailable. arxiv.org/abs/2608.11200v1
View on X
@OpenAI launches GPT-5.6-Cyber to answer 98.5% of security queries blocked from defenders. the-decoder.com/openai-launches-gpt-5-6-cyber-to-he…
View on X
Four distinct LLM misalignment types arise from four different loss functions. alignmentforum.org/posts/GRmvZsHXH4vaijPMv/four-llm…
View on X
OpenAI flags Astra model at highest cybersecurity risk level in internal tests. the-decoder.com/openai-flags-its-new-astra-model-as…
3View on X
Evals aren't enough. Real-world attacks find the gaps automated tests miss. gradientflow.com/passing-your-evals-doesnt-mean-you…
View on X