Transformer Relevance Patterns

Exploring the extension of LRP methods to transformer models reveals intriguing insights into how relevance is distributed across input tokens. As the number of target tokens increases, the model tends to hallucinate more, indicating a shift in attention away from the source. Notably, the entropy of contributions exhibits a unique curve, suggesting varying levels of confidence in token relevance at different stages.