#reinforcement-learning
16 briefs tagged #reinforcement-learning.
Agent Lightning v1.0 targets harnessed agentic RL
The framework connects arbitrary agent harnesses to RL training through an LLM endpoint proxy.
OpenAI slows AI training after Hugging Face breach
The company paused some reinforcement learning work while tightening model security and safeguards.
SA-MRPO targets unsolved rewards in RL post-training
The method reweights multi-objective RL rewards by saturation instead of fixed scalar weights.
Researchers target RL training for agent harnesses
LEGO-RL and ClawGym II propose proxy-and-sandbox approaches for long-horizon agent optimization.
New papers target VLMs' spatial reasoning gap
Researchers propose runtime memory, RL training, benchmarks and 3D generation methods for spatial AI.
Researchers target VLM spatial reasoning gaps
New papers propose training, memory and benchmarks for multimodal spatial reasoning.
Researchers target LLM exploration beyond temperature
DORA, 3PO and RISE-RL propose new ways to improve exploration in LLM agents and reinforcement learning.
UserIDA separates intent control in user simulation
The arXiv paper proposes directive-conditioned generation for controllable assistant evaluation.
New papers target scalable training for LLM agents
EnvACE and State2State reduce reliance on manual environments, expert trajectories and task design.
ABSeeker targets credit assignment for search agents
The arXiv paper proposes dense step-level rewards for training long-horizon search agents.
New papers probe self-distillation for LLM reinforcement learning
ArXiv reports propose ICE and OCSD while warning that privileged-information teachers can fail.
DASH targets overthinking in reasoning language models
The method uses segment-level credit assignment to reduce unproductive self-reflection in reasoning traces.
WorldCycle targets drift in video world models
The arXiv paper proposes self-verifiable RL using reversible action cycles.
Researchers propose Skill Entropy for LLM reasoning
Skill^2-Bench tests how models switch across 558 skills in long-horizon tasks.
New papers test self-distillation for LLM reinforcement learning
Studies propose ICE and OCSD while warning that privileged-information teachers can fail on harder tasks.
Wnuan paper tests staged post-training for enterprise QA
A 32B route lifted acceptable-answer rate on WnuanBench while reducing general-benchmark scores.