All news

July 2026

Going to ICML 2026 to co-present Emergence of Exploration

I will be at ICML 2026 to co-present Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying, our work on retry-based objectives and exploration emerging from policy-gradient optimization.

Lab news · arXiv · Code

June 2026

OrderGrad featured by alphaXiv on X

The OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation preprint was featured on alphaXiv's X channel.

alphaXiv X post · alphaXiv page · arXiv