SAFETY & SECURITY
Every Safety & Security story from RXed AI News over the last 30 days — 27 items.
@OpenAI's Agent Builder flaw let one bad ChatGPT link deploy rogue agents on employee devices.
the-decoder.com/one-tampered-chatgpt-link-could-spa…
@OpenAI's hack of HuggingFace demands a rethink of open source security.
garymarcus.substack.com/p/openais-disconcerting-hac…
New method enables pixel-level tampering detection in modern vision-language models.
arxiv.org/abs/2607.18230v1
Open-weight LLMs show promise for generating structured threat data for autonomous vehicle vulnerabilities. arxiv.org/abs/2607.16175v1
Security agent evaluation should measure cost, not just peak offensive capability under generous budgets.
arxiv.org/abs/2607.15263v1
Poisoning pretraining data can introduce harmful behaviors to language models that are difficult to detect and mitigate.
arxiv.org/abs/2607.15267v1
OpenAI's GPT-Red automated red-teamer beat humans 84-13 on prompt injection attacks using self-play RL.
marktechpost.com/2026/07/16/openai-details-gpt-red-…
OpenAI introduces age-appropriate protections, learning tools, and parental controls to make ChatGPT safer for teens. Source: openai.com/index/why-teens-deserve-access-safe-ai
Deep Interaction introduces a new method for humans to correct reasoning errors in large language models, improving their accuracy.
arxiv.org/abs/2607.14049v1
@MiraMurati’s Thinking Machines Lab publishes technical case for human-centered AI with customizable model weights.
marktechpost.com/2026/07/11/mira-muratis-thinking-m…
Deepfake detectors are losing the arms race; trust must shift to multimodal provenance.
lesswrong.com/posts/MBRNR5h9g6HGvAJDe/don-t-bring-a…
Natural language autoencoders vary significantly in robustness to initialization methods, affecting their ability to reliably interpret LLM thought processes.
alignmentforum.org/posts/LQXWiF8PyJ5ojNsEv/how-robu…
Measuring AI capabilities is no longer the primary strategy for reducing existential risk.
lesswrong.com/posts/4TMKvGmoAWjXBGwWk/measuring-is-…
Value generalisation is key to AI Alignment, necessary and almost sufficient for solving it.
alignmentforum.org/posts/iPyJfD9Jyxj6Jfdws/value-ge…
New research explores technical alignment for brain-like AGI using human-like social drives.
alignmentforum.org/posts/rKdS7i4StaMmFzYRo/notes-on…
AI safety talent is dangerously concentrated in a handful of cities, risking critical bottlenecks for the field.
lesswrong.com/posts/jXQ9hPL8RqkH5ybnt/there-should-…
@NeelNanda8 and @DeepMind propose a pragmatic shift in AI interpretability, prioritizing utility over reverse-engineering models.
lesswrong.com/posts/Ro7kkSHg5SE7yYW2c/balancing-rig…
Deep dive into polysemanticity in language models, a key concept for mechanistic interpretability research.
lesswrong.com/posts/JpoF5zBKmcs2uHQAS/the-polyseman…
Anthropic faces claims of secret Claude tracker monitoring Chinese users, contradicting its anti-surveillance stance. Engineer confirms "experiment" ended. Source:...
Despite alignment training, LLMs remain prone to generating unsafe outputs online, making critical monitoring essential.
arxiv.org/abs/2607.02510v1
LACUNA introduces a testbed to evaluate localization precision for LLM unlearning methods.
arxiv.org/abs/2607.02513v1
Persistent AI codebases create a new attack surface for misaligned agents during iterative development.
arxiv.org/abs/2607.02514v1
New research trains language models to self-explain their predictions, revealing how behavioral change occurs despite fixed supervision.
arxiv.org/abs/2606.32038v1
Deployment awareness matters more than evaluation awareness for AI safety.
alignmentforum.org/posts/XP794SHDuXYfWLrvJ/deployme…
A reading list for generalists addresses the shortage of multifaceted talent in AI safety.
lesswrong.com/posts/sH4cFDDjRdGrn3p2o/a-reading-lis…
New sparsity regularizers make top-k sparse autoencoders more interpretable for vision foundation models.
arxiv.org/abs/2606.27321v1
Process reward models enable fine-grained step-level evaluation of LLMs, but building them for agentic settings remains difficult due to long-horizon interactions and irreversible actions.
huggingface.co/papers/2606.26080
OTHER CATEGORIES