Forscher nehmen Fehler bei On-Policy-Distillation in Agenten ins Visier
Aktuelle arXiv-Papers schlagen turn-, schritt- und präfixbewusste Trainingskorrekturen für tool-nutzende LLM-Agenten vor.
Warum es wichtig ist
Die Arbeiten spiegeln eine breitere Verschiebung wider: weg von ausschließlich ergebnisorientiertem Reinforcement Learning hin zu dichterer, zustandsbewusster Supervision für das Agententraining. Das ist wichtig, weil tool-nutzende Agenten oft über mehrere Turns hinweg scheitern, während grobe Trajektorien-Rewards übersehen können, welcher Schritt oder welcher Tool Call den Fehler ausgelöst hat.
Die Kernpunkte
- 1.Neue Methoden gewichten Distillation nach Turn-, Schritt-, Präfix- oder Sample-Qualität neu.
- 2.Die Papers zielen auf sparse Rewards und unzuverlässige Teacher-Signale in langen Agenten-Trajektorien.
- 3.Tool-Call-Fehler können sich fortpflanzen und dadurch tokenbasierte Teacher-Supervision schwächen.
Eine Reihe aktueller arXiv-Papers schlägt neue Verfahren vor, um tool-integrierte und mehrzügige Sprachmodell-Agenten mit On-Policy-Distillation und Reinforcement Learning zu trainieren. Dazu gehören TurnSight, SOD, ATOD, PG-OPD und RSTG, die Zuverlässigkeitsprobleme wie sparse Rewards, sich fortpflanzende Tool-Call-Fehler, Teacher-Student-Divergenz, ineffiziente lange Rollouts und verlorene Gradienten in Reward-Gruppen mit Nullvarianz adressieren. Die Papers berichten Experimente über Long-Horizon-Reasoning, Mathematik, Naturwissenschaften, Code und interaktive Agentenaufgaben hinweg, doch die vorliegenden Abstracts belegen keinen einzelnen Benchmark-Spitzenreiter über alle Methoden hinweg.
⚡ Heute ausprobieren
Wer tool-nutzende Agenten trainiert, sollte prüfen, ob Teacher-Supervision nach Schritt- oder Turn-Qualität gewichtet wird, bevor OPD mit RL kombiniert wird.
Quellen & Originalberichte
Dieser Brief fasst die Berichterstattung der folgenden Medien zusammen und verlinkt sie.
- arXiv cs.AIInstruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned PolicyAug 6, 12:00 PM↗
- arXiv cs.AIAgentic Reinforcement Learning with Observation-Calibrated Self-DistillationAug 6, 12:00 PM↗
- arXiv cs.AIPrivileged, but Biased: How PI-Conditioned Teachers Break Self-DistillationAug 6, 12:00 PM↗
- arXiv cs.LGInstruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned PolicyAug 6, 12:00 PM↗
- arXiv cs.LGPrivileged, but Biased: How PI-Conditioned Teachers Break Self-DistillationAug 6, 12:00 PM↗
- arXiv cs.LGReward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement LearningAug 6, 12:00 PM↗
- arXiv cs.LGAgentic Reinforcement Learning with Observation-Calibrated Self-DistillationAug 6, 12:00 PM↗
- arXiv cs.AITurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated ReasoningAug 5, 12:00 PM↗
- arXiv cs.AIAgentic Reinforcement Learning with Self-Distilled Reward ShapingAug 5, 12:00 PM↗
- arXiv cs.AIRubrics as Privileged Information for Open-Ended GenerationAug 5, 12:00 PM↗
- arXiv cs.CLAgentic Reinforcement Learning with Self-Distilled Reward ShapingAug 5, 12:00 PM↗
- arXiv cs.CLTurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated ReasoningAug 5, 12:00 PM↗
- arXiv cs.LGAgentic Reinforcement Learning with Self-Distilled Reward ShapingAug 5, 12:00 PM↗
- arXiv cs.LGRubrics as Privileged Information for Open-Ended GenerationAug 5, 12:00 PM↗
- arXiv cs.LGRoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility StatesAug 4, 12:00 PM↗
- arXiv cs.CLRoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility StatesAug 4, 12:00 PM↗
- arXiv cs.AIDRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer TrainingAug 4, 12:00 PM↗
- arXiv cs.AIGroup-Reflective Self-Distillation for Agentic Reinforcement LearningAug 4, 12:00 PM↗
- arXiv cs.AIInstruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-DistillationAug 4, 12:00 PM↗
- arXiv cs.AIPCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement LearningAug 4, 12:00 PM↗
- arXiv cs.AIDAPD: Dual-Anchored Policy DistillationAug 4, 12:00 PM↗
- arXiv cs.AIIs More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled ReasoningAug 4, 12:00 PM↗
- arXiv cs.LGGroup-Reflective Self-Distillation for Agentic Reinforcement LearningAug 4, 12:00 PM↗
- arXiv cs.LGDRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer TrainingAug 4, 12:00 PM↗
- arXiv cs.LGInstruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-DistillationAug 4, 12:00 PM↗
- arXiv cs.LGRoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility StatesAug 4, 12:00 PM↗
- arXiv cs.CLRoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility StatesAug 4, 12:00 PM↗
- arXiv cs.CLInstruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-DistillationAug 4, 12:00 PM↗
- HF Daily PapersTurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated ReasoningAug 4, 4:00 AM↗
Hat dir dieses Briefing gefallen? Erhalte das nächste per Mail.
Mehr in Forschung
InternLM schlägt Mobius-Modellarchitektur vor
Das arXiv-Paper trennt Gedächtnis und Schlussfolgern, um Kompression und Inferenz effizienter zu machen.
KI-Entwicklertools und Arbeitsweisen lösen Debatte auf Hacker News aus
Beiträge zu MathCode, KI-Coding-Gewohnheiten, Cloudflare und Googles HEIR prägten die Diskussion.
Google treibt private KI mit homomorpher Verschlüsselung voran
Google will private KI mithilfe homomorpher Verschlüsselung praxistauglicher machen.