Dealing with Noisy Feedback
Ramya explains how noise in the form of contradictory signals from the person can impact the agent's decision-making process. She also discusses the challenges of distinguishing between safe and unsafe states when multiple safe labels are received. The team uses an Em style expectation maximization approach to learn the prior distribution and confusion matrix, allowing them to estimate aggregated labels for each state.In this clip
From this podcast

Data Skeptic
Blind Spots in Reinforcement Learning
Related Questions