Thorsten discusses the challenges of evaluating policies in machine learning, particularly the issues of bias when dealing with zero probabilities and high variance from small probabilities. He emphasizes the simplicity of the underlying mathematics for offline A/B testing while highlighting the complexities of ensuring random assignment and accurate propensity logging. By leveraging log data instead of hand-labeled data, existing learning algorithms can be repurposed for more effective training.