Unveiling Model Interpretability

Kyle and Max discuss the use of Bert vectors as features and the challenge of interpretability in the model. They explore how logistic regression coefficients reveal the predictive words and phrases for compliance or refusal in AI language models like Chat GPT. The study also highlights the significance of negative generalizations and controversial figures in predicting refusals.