1 briefs tagged #self-rewarding.
The paper proposes multi-agent RL to reduce reliance on ground-truth supervision for reasoning models.