Published Sep 9, 2021

Emily M. Bender — Language Models and Linguistics

Emily M. Bender delves into the intricacies of language models, discussing their limitations, biases, and ethical challenges while stressing the importance of more inclusive benchmarks, environmental responsibility, and ethical implementation strategies to harness their potential responsibly.
Episode Highlights
Gradient Dissent - A Machine Learning Podcast logo

Popular Clips

Episode Highlights

  • Limits of Understanding

    Emily M. Bender highlights the inherent limitations of language models in truly understanding language and concepts. She explains that while algorithms can grasp linguistic form, they lack the ability to comprehend meaning without additional inputs like visual grounding or knowledge bases 1. Emily uses the analogy of being dropped into a library with books in an unfamiliar language to illustrate the challenge of learning language without context or external references 1.

    Language modeling is not natural language understanding.

    ---

    She also emphasizes the importance of acknowledging language diversity in computational linguistics, advocating for research beyond English to include other languages, which often lack extensive data 2.

       

    Bias Concerns

    Bias in language models is a significant concern, as Emily illustrates with examples of word embeddings that reflect societal prejudices. She describes how word embeddings can inadvertently associate negative sentiments with certain terms due to biased training data, such as the case where Mexican restaurants were unfairly rated lower 3. This bias is not just a future concern but a present reality, as demonstrated by Safia Noble's work on search engine biases, where identity terms are manipulated for profit 4.

    Models also contribute to bias, not just the data.

    ---

    Emily stresses the need for more curated data and debiasing efforts to mitigate these issues, acknowledging that while a completely bias-free system is unattainable, improvements are possible 3.

       

    Ethical Issues

    The ethical challenges surrounding large language models are complex and controversial. Emily discusses the backlash from Google's ethical AI team following the publication of the paper on stochastic parrots, which questioned the unchecked growth of language models 5. Despite the paper being a survey of existing perspectives, it sparked significant controversy, highlighting the tension between technological advancement and ethical considerations 6.

    We thought we'd be ruffling some feathers with this paper.

    ---

    Emily emphasizes the importance of considering the societal impact of NLP technologies and encourages a nuanced discussion about ethics in AI, focusing on systemic issues rather than labeling individuals as ethical or unethical 7.

Related Episodes