Evaluating New Policies
Thorsten discusses the importance of random assignment in evaluating new policies, emphasizing that even non-uniform randomness can yield unbiased performance estimates. He explains the concept of inverse propensity weighting, which allows for the reuse of existing log data to assess new policies, provided that the actions evaluated had non-zero probabilities in the past. Understanding these conditions is crucial for effective policy evaluation and learning.In this clip
From this podcast

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Unbiased Learning from Biased User Feedback with Thorsten Joachims - TWiML Talk #207
Related Questions