RXed AI News

AI news with a spine.
@RXed_EU

Every Safety & Security story from RXed AI News over the last 30 days — 27 items.

@OpenAI's Agent Builder flaw let one bad ChatGPT link deploy rogue agents on employee devices. the-decoder.com/one-tampered-chatgpt-link-could-spa…
View on X
@OpenAI's hack of HuggingFace demands a rethink of open source security. garymarcus.substack.com/p/openais-disconcerting-hac…
1View on X
New method enables pixel-level tampering detection in modern vision-language models. arxiv.org/abs/2607.18230v1
1View on X
Open-weight LLMs show promise for generating structured threat data for autonomous vehicle vulnerabilities. arxiv.org/abs/2607.16175v1
View on X
Security agent evaluation should measure cost, not just peak offensive capability under generous budgets. arxiv.org/abs/2607.15263v1
View on X
Poisoning pretraining data can introduce harmful behaviors to language models that are difficult to detect and mitigate. arxiv.org/abs/2607.15267v1
View on X
OpenAI's GPT-Red automated red-teamer beat humans 84-13 on prompt injection attacks using self-play RL. marktechpost.com/2026/07/16/openai-details-gpt-red-…
View on X
OpenAI introduces age-appropriate protections, learning tools, and parental controls to make ChatGPT safer for teens. Source: openai.com/index/why-teens-deserve-access-safe-ai
View on X
Deep Interaction introduces a new method for humans to correct reasoning errors in large language models, improving their accuracy. arxiv.org/abs/2607.14049v1
View on X
@MiraMurati’s Thinking Machines Lab publishes technical case for human-centered AI with customizable model weights. marktechpost.com/2026/07/11/mira-muratis-thinking-m…
View on X
Deepfake detectors are losing the arms race; trust must shift to multimodal provenance. lesswrong.com/posts/MBRNR5h9g6HGvAJDe/don-t-bring-a…
View on X
Natural language autoencoders vary significantly in robustness to initialization methods, affecting their ability to reliably interpret LLM thought processes. alignmentforum.org/posts/LQXWiF8PyJ5ojNsEv/how-robu…
View on X
Measuring AI capabilities is no longer the primary strategy for reducing existential risk. lesswrong.com/posts/4TMKvGmoAWjXBGwWk/measuring-is-…
View on X
Value generalisation is key to AI Alignment, necessary and almost sufficient for solving it. alignmentforum.org/posts/iPyJfD9Jyxj6Jfdws/value-ge…
View on X
New research explores technical alignment for brain-like AGI using human-like social drives. alignmentforum.org/posts/rKdS7i4StaMmFzYRo/notes-on…
View on X
AI safety talent is dangerously concentrated in a handful of cities, risking critical bottlenecks for the field. lesswrong.com/posts/jXQ9hPL8RqkH5ybnt/there-should-…
View on X
@NeelNanda8 and @DeepMind propose a pragmatic shift in AI interpretability, prioritizing utility over reverse-engineering models. lesswrong.com/posts/Ro7kkSHg5SE7yYW2c/balancing-rig…
View on X
Deep dive into polysemanticity in language models, a key concept for mechanistic interpretability research. lesswrong.com/posts/JpoF5zBKmcs2uHQAS/the-polyseman…
View on X
Anthropic faces claims of secret Claude tracker monitoring Chinese users, contradicting its anti-surveillance stance. Engineer confirms "experiment" ended. Source:...
1View on X
Despite alignment training, LLMs remain prone to generating unsafe outputs online, making critical monitoring essential. arxiv.org/abs/2607.02510v1
View on X
LACUNA introduces a testbed to evaluate localization precision for LLM unlearning methods. arxiv.org/abs/2607.02513v1
View on X
Persistent AI codebases create a new attack surface for misaligned agents during iterative development. arxiv.org/abs/2607.02514v1
View on X
New research trains language models to self-explain their predictions, revealing how behavioral change occurs despite fixed supervision. arxiv.org/abs/2606.32038v1
View on X
Deployment awareness matters more than evaluation awareness for AI safety. alignmentforum.org/posts/XP794SHDuXYfWLrvJ/deployme…
View on X
A reading list for generalists addresses the shortage of multifaceted talent in AI safety. lesswrong.com/posts/sH4cFDDjRdGrn3p2o/a-reading-lis…
View on X
New sparsity regularizers make top-k sparse autoencoders more interpretable for vision foundation models. arxiv.org/abs/2606.27321v1
View on X
Process reward models enable fine-grained step-level evaluation of LLMs, but building them for agentic settings remains difficult due to long-horizon interactions and irreversible actions. huggingface.co/papers/2606.26080
View on X