Published Nov 27, 2019

97 - Automated Analysis Of Historical Printed Documents, With Taylor Berg-Kirkpatrick

Taylor Berg-Kirkpatrick delves into the cutting-edge intersection of machine learning and historical document analysis, highlighting the challenges and innovations in OCR for historical texts, uncovering printing biases, and using novel unsupervised learning techniques to reveal hidden insights from the past.
Episode Highlights
NLP Highlights logo

Popular Clips

Episode Highlights

  • Interpretability

    Interpretability is crucial in machine learning models for historical document analysis, as explains. He emphasizes that without interpretability, machine learning predictions are useless because they lack context and understanding 1. The challenge lies in balancing the power of modern neural networks with the need for models to be based on understood assumptions. This balance is essential for creating useful predictions from historical data, which often involves complex distributions and noise 2.

    What you're actually looking for is the model's interpretation, and you want that interpretation to be based in understood assumptions.

    ---

    The goal is to create models that can provide insights into historical processes while maintaining a level of interpretability that makes their predictions valuable.

       

    Model Capacity

    Balancing model capacity and interpretability in historical OCR models presents unique challenges. discusses the need for models that can handle nonlinear transformations while embedding prior knowledge 3. This balance is crucial in machine learning, where high-capacity models require vast data, but historical data is often limited. He highlights the importance of using unsupervised learning to bridge the gap between known processes and data-driven insights 4.

    It's this kind of interplay of what we know in advance that we embed in the system through modeling assumptions and what we don't know, which we try to learn with high capacity model.

    ---

    By leveraging both prior knowledge and data, researchers can develop models that are both powerful and interpretable, offering valuable insights into historical documents.