AAI News Hub

#reinforcement-learning

16 briefs tagged #reinforcement-learning.

Research

Agent Lightning v1.0 targets harnessed agentic RL

The framework connects arbitrary agent harnesses to RL training through an LLM endpoint proxy.

arXiv cs.AI+1 outlet·15h ago
Policy

OpenAI slows AI training after Hugging Face breach

The company paused some reinforcement learning work while tightening model security and safeguards.

TechCrunch AI+1 outlet·1d ago
Research

SA-MRPO targets unsolved rewards in RL post-training

The method reweights multi-objective RL rewards by saturation instead of fixed scalar weights.

arXiv cs.AI+1 outlet·1d ago
Research

Researchers target RL training for agent harnesses

LEGO-RL and ClawGym II propose proxy-and-sandbox approaches for long-horizon agent optimization.

HF Daily Papers+1 outlet·2d ago
Research

New papers target VLMs' spatial reasoning gap

Researchers propose runtime memory, RL training, benchmarks and 3D generation methods for spatial AI.

arXiv cs.AI+1 outlet·4d ago
Research

Researchers target VLM spatial reasoning gaps

New papers propose training, memory and benchmarks for multimodal spatial reasoning.

arXiv cs.AI+1 outlet·6d ago
Research

Researchers target LLM exploration beyond temperature

DORA, 3PO and RISE-RL propose new ways to improve exploration in LLM agents and reinforcement learning.

arXiv cs.AI+1 outlet·6d ago
Research

UserIDA separates intent control in user simulation

The arXiv paper proposes directive-conditioned generation for controllable assistant evaluation.

HF Daily Papers+1 outlet·Aug 10
Research

New papers target scalable training for LLM agents

EnvACE and State2State reduce reliance on manual environments, expert trajectories and task design.

arXiv cs.AI+1 outlet·Aug 7
Research

ABSeeker targets credit assignment for search agents

The arXiv paper proposes dense step-level rewards for training long-horizon search agents.

arXiv cs.AI+1 outlet·Aug 6
Research

New papers probe self-distillation for LLM reinforcement learning

ArXiv reports propose ICE and OCSD while warning that privileged-information teachers can fail.

arXiv cs.CL+1 outlet·Aug 6
Research

DASH targets overthinking in reasoning language models

The method uses segment-level credit assignment to reduce unproductive self-reflection in reasoning traces.

arXiv cs.CL+1 outlet·Aug 5
Research

WorldCycle targets drift in video world models

The arXiv paper proposes self-verifiable RL using reversible action cycles.

HF Daily Papers+1 outlet·Aug 5
Research

Researchers propose Skill Entropy for LLM reasoning

Skill^2-Bench tests how models switch across 558 skills in long-horizon tasks.

HF Daily Papers+1 outlet·Aug 5
Research

New papers test self-distillation for LLM reinforcement learning

Studies propose ICE and OCSD while warning that privileged-information teachers can fail on harder tasks.

arXiv cs.AI+1 outlet·Aug 4
Research

Wnuan paper tests staged post-training for enterprise QA

A 32B route lifted acceptable-answer rate on WnuanBench while reducing general-benchmark scores.

arXiv cs.AI+1 outlet·Aug 4