ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
arXiv:2412.03800v2 Announce Type: replace-cross Abstract: Reinforcement learning agents depend on reward signals whose density is rarely under the designer's control, and when such signals are absent, an agent must generate its own drive to explore. State entropy maximization offers a principled…