跳转到内容

博客与开源

  • Lilian Weng 博客:Policy Gradient Algorithms、Exploration Strategies、RLHF 综述
  • Huang et al. “The 37 Implementation Details of PPO”(ICLR Blog Track)——第四、十一卷的必读
  • Nathan Lambert《RLHF Book》(在线)与 Interconnects 通讯(→ 第九卷)
  • Andy Jones “Debugging RL, Without the Agonizing Pain”(→ 第十一卷)
  • 库:CleanRL、Stable-Baselines3、RLlib、TorchRL、Tianshou、Acme;LLM 后训练:TRL、OpenRLHF、veRL、TRLX
  • 环境:Gymnasium、PettingZoo、ProcGen、Isaac Lab、Brax、MiniGrid、Atari ALE、D4RL/Minari