Reward Functions and Horizon Planning
Mikko Lauri explains how reward functions encode tasks by assigning high rewards to useful actions and low rewards to poor actions. He also discusses the importance of planning for long-term rewards over a finite horizon and the computational advantages of this approach.In this clip
From this podcast

Data Skeptic
Decentralized Information Gathering
Related Questions