Energy Based Policies

Using energy-based models to represent policies in control and the benefit compared to other commonly used models.

Figure 1: Visualization of different types of policies.

Figure 1: Visualization of different types of policies. (source)

\[\pi(a \mid s) \propto \exp(-E_\theta(s, a))\]

Code

Code for this blog is available here: github.com/mht3/ebp

References

  1. Pete Florence, Corey Lynch, Andy Zeng, Oscar Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson. “Implicit Behavioral Cloning.” arXiv:2109.00137, 2021. [arXiv]
  2. Sumeet Singh, Stephen Tu, and Vikas Sindhwani. “Revisiting Energy Based Models as Policies: Ranking Noise Contrastive Estimation and Interpolating Energy Models.” arXiv:2309.05803, 2023. [arXiv]
  3. Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. “Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.” arXiv:2303.04137, 2024. [arXiv]
  4. Kevin Zakka. “A PyTorch Implementation of Implicit Behavioral Cloning.” Version 0.0.1, 2021. [GitHub]

© 2026 Matt Taylor. All rights reserved.

Powered by Hydejack v9.2.1