Evaluating Text Generation

Pradeep and Alexis discuss the challenges of human evaluation in text generation systems, highlighting the subjectivity of human judgments. They delve into the importance of Inter Annotator Agreement in assessing the quality of collected data and share strategies to mitigate issues of bias and noise in human-labeled data.