RESEARCH
Every Research story from RXed AI News over the last 30 days — 66 items.
New reinforcement learning method uses predictive divergence masks to stabilize LLM off-policy updates, improving training efficiency.
huggingface.co/papers/2607.10848
Interactive video generation now controlled by graph-based object interactions, not just pixels.
arxiv.org/abs/2607.21580v1
New research enables multimodal region control for Diffusion Transformers, giving artists precise editing power. arxiv.org/abs/2607.19344v1
New AI method reduces repetitive copying in long-context reasoning by grounding steps in evidence.
arxiv.org/abs/2607.19345v1
No universally superior harness exists for autonomous discovery systems like OpenEvolve and TTT-Discover.
arxiv.org/abs/2607.18235v1
Active human vision is a closed loop redirecting gaze through intermediate hypotheses, not single snapshots. arxiv.org/abs/2607.16165v1
Autonomous data science agents still rely on expensive trial-and-error workflows, limiting efficiency. huggingface.co/papers/2607.15901
Muon shows competitive performance in large-scale pre-training but its value for reinforcement-learning post-training is unclear. arxiv.org/abs/2607.16169v1
@DeepMind's GenCeption video generator matches state-of-the-art vision systems using far less training data. the-decoder.com/google-deepmind-argues-video-genera…
SceneBind creates omni-modal scene representations binding vision, audio, and language for realistic environments.
arxiv.org/abs/2607.15265v1
SearchOS-V1 tackles agent collaboration failures as web search becomes core to information-seeking agents.
arxiv.org/abs/2607.15257v1
University of Pennsylvania professor used @OpenAI's GPT-5.6 Sol to disprove a 30-year-old statistics conjecture in 90 minutes.
the-decoder.com/gpt-5-6-sol-reportedly-disproves-a-…
GigaWorld-Policy-0.5 improves robot learning with World Action Models using future scenes as supervision.
huggingface.co/papers/2607.13960
New QA-driven LLM agents reduce software issue resolution errors by acquiring repository knowledge before fixing.
huggingface.co/papers/2607.11111
LLM agents struggle to gauge task complexity, often overcommitting to simple problems, new research shows.
arxiv.org/abs/2607.13034v1
SpectraReward turns pretrained MLLMs into zero-shot reward models for image generation RL, requiring no new training.
huggingface.co/papers/2607.11886
CtrlVTON enables precise garment placement in virtual try-ons using visual-instance-prompt segmentation.
huggingface.co/papers/2607.09362
Proxy-Guided Update Signals enable modular LLM post-training with reusable domain-specific policies.
huggingface.co/papers/2607.11505
New theory explains how Transformers develop inductive reasoning abilities, with implications for future AI regulation.
arxiv.org/abs/2607.11875v1
New theory explains when speculative decoding drafts are accepted to accelerate language model inference.
arxiv.org/abs/2607.11881v1
Vision-language models have made 10 years of progress in visual reasoning, reducing errors in complex human interactions.
arxiv.org/abs/2607.09654v1
New AI agent benchmarks test 500-day terminal tasks, revealing critical gaps in long-horizon autonomy.
huggingface.co/papers/2607.08964
A proactive memory agent surfaces critical states across long-horizon AI tasks without trajectory bloat.
huggingface.co/papers/2607.08716
SLORR enables efficient model compression without accuracy loss by using simple in-training low-rank regularization.
arxiv.org/abs/2607.08754v1
Beijing Academy of Artificial Intelligence's Orca world model predicts abstract states without action labels, matching robotics systems.
the-decoder.com/chinas-orca-world-model-matches-spe…
@OpenAI's GPT-5.6 Sol Ultra reportedly proved the 50-year-old Cycle Double Cover Conjecture in under an hour using 64 subagents.
the-decoder.com/openais-gpt-5-6-sol-ultra-reportedl…
CausalDS introduces a benchmark to evaluate causal reasoning in data-science agents using LLMs.
huggingface.co/papers/2607.08093
OpenCoF uses video generation to train models in logical reasoning without human labels.
arxiv.org/abs/2607.08763v1
Scientific ideas inherit and recombine like genomes, challenging current AI benchmarks.
arxiv.org/abs/2607.08758v1
New research demonstrates modular pretraining enables granular access control for AI models.
lesswrong.com/posts/43vKjWuH4goLwrFHA/modular-pretr…
A mathematical theory of attention is tested against trained transformers in a new series of experiments.
lesswrong.com/posts/2dA7phbYZGPjhTj9q/transformers-…
LessWrong explores the surprising historical path to democracy, a system often assumed to be inevitable. lesswrong.com/posts/cLQrijiGDa9X9wHqt/how-did-we-ge…
Neel Nanda’s research finds data filtering works far worse than expected for AI model training.
alignmentforum.org/posts/aTybJ6CPQrxEY8rE2/data-fil…
New AI model classifies astronomical transients without human labels using uncertainty quantification.
arxiv.org/abs/2607.05393v1
New robot vision model maintains accuracy without camera calibration, solving real-world deployment challenges.
arxiv.org/abs/2607.05396v1
New reinforcement learning method distills reasoning from strong models without costly retraining.
arxiv.org/abs/2607.05394v1
Bentham's Bulldog publishes a case against FDT, contrasting rationalist decision theory with pragmatic predictors in game theory.
alignmentforum.org/posts/SdGbWkCZgCN7EGBxM/pragmati…
New visual token pruning method compresses redundant image patches while preserving critical cues for dense VLM instructions.
arxiv.org/abs/2607.02484v1
G-RRM combines neural networks with symbolic reasoning to solve larger problems than previous methods.
arxiv.org/abs/2607.02491v1
AutoMem automates memory expertise in LLMs, teaching them what to encode, when to retrieve, and how to organize knowledge.
huggingface.co/papers/2607.01224
DemoPSD introduces disagreement-modulated policy self-distillation to boost LLM reasoning without multiple models.
arxiv.org/abs/2607.02502v1
@Google Research's WARP recovers training data portfolios from foundation models without accessing raw data.
huggingface.co/papers/2607.01686
Reasoning LLMs boost speaker recognition accuracy in long-form TV dramas, critical for storyline comprehension.
arxiv.org/abs/2607.02504v1
Grid-based ANN search methods are systematically characterized for high-dimensional scaling.
huggingface.co/papers/2607.01283
LLM agents develop social structures and objectives in multi-agent debates based on roles and audience.
arxiv.org/abs/2607.02507v1
ReContext uses recursive evidence replay to help LLMs reason over long contexts.
arxiv.org/abs/2607.02509v1
Quantum-inspired fast weight programmers achieve 99.8% accuracy in traffic matrix forecasting using 95% fewer parameters.
huggingface.co/papers/2606.27821
New programming paradigm 'Program-as-Weights' tackles fuzzy functions like log analysis and JSON repair using model weights directly.
arxiv.org/abs/2607.02512v1
TurboServe enables efficient, economical streaming video generation for long-lived user sessions.
huggingface.co/papers/2606.19271
New research improves AI learning from imperfect human demonstrations without compressed signals.
arxiv.org/abs/2607.01225v1
New research decouples perception and reasoning for fine-grained visual reasoning in vision-language models.
huggingface.co/papers/2607.01191
New study measures the gap between human and LLM research ideas, finding current models fall short of human creativity.
arxiv.org/abs/2607.01233v1
AutoMem automates memory expertise in LLMs, teaching them what to encode, when to retrieve, and how to organize knowledge.
arxiv.org/abs/2607.01224v1
@Meta's non-invasive brain-to-text AI translates brain activity into typed sentences without surgery.
the-decoder.com/metas-non-invasive-brain-to-text-ai…
New paper from @danahboyd and others on avoiding accountability decoys in AI's political economy.
bespacific.com/reckoning-with-the-political-economy…
New theory explains when speculative decoding drafts are accepted to accelerate language model inference.
arxiv.org/abs/2606.30265v1
Ko-WideSearch benchmarks web agents on breadth-searching closed sets, not just depth-solving single answers.
huggingface.co/papers/2606.27595
Information-aware KV cache compression improves long-reasoning efficiency in large language models.
huggingface.co/papers/2606.26875
Multimodal web agents autonomously decompose complex GUI tasks into executable actions for human assistance.
arxiv.org/abs/2606.27330v1
Language-based digital twins show promise for early detection of mild cognitive impairment in elderly patients.
arxiv.org/abs/2606.27334v1
OTHER CATEGORIES