Unveiling Model Interpretability
Kyle and Max discuss the use of Bert vectors as features and the challenge of interpretability in the model. They explore how logistic regression coefficients reveal the predictive words and phrases for compliance or refusal in AI language models like Chat GPT. The study also highlights the significance of negative generalizations and controversial figures in predicting refusals.In this clip
From this podcast

Data Skeptic
Prompt Refusal
Related Questions