Published Sep 11, 2024

Futures of the data foundry business model

Nathan Lambert delves into the intricate challenges and future trends of the data foundry business model, focusing on data acquisition, market dynamics, and the critical role of high-quality data for AI training and utilization.
Episode Highlights
Interconnects Audio logo

Popular Clips

Questions from this episode

Episode Highlights

  • Annotation

    The methods and intricacies of data annotation for AI training are complex and costly. explains that basic instructions are cheap, but scientific reports are extremely expensive, contributing to the $10 million-plus data acquisition costs for fine-tuning 1. The process involves iterative training, evolving data instructions, and navigating bureaucratic challenges, making it difficult for new organizations to add human data to their pipelines 1.

    The process is difficult for new organizations trying to add human data to their pipelines, given the sensitivity, processes that work and improve the models are extracted until the performance runs out.

    ---

    Acquiring data from vendors is a high-stakes endeavor, often requiring legal or financial action to ensure delivery, and the data is delivered in weekly batches, with the goal of improving models over time 1.

       

    Quality

    Maintaining high-quality data for AI training presents significant challenges. Nathan notes that while synthetic data is becoming more prevalent due to its cost-effectiveness, human data remains crucial for certain tasks 1 2. The academic community is exploring how language models can replace humans for labeling preference data, but there are concerns about the quality and authenticity of data obtained from vendors 2.

    Waning data quality through AI slop is increasingly hard to guarantee that data obtained from a data vendor is not already partially or fully coming from a language model.

    ---

    The process of acquiring and utilizing human data involves iterative experimentation and high effort, with millions of dollars potentially wasted on datasets not used in final models 1.

       

    Utilization

    Companies strategically use data for AI training through innovative techniques and business strategies. Nathan highlights that data annotation companies are evolving to handle data collection and training tasks, with subsidiaries specializing in different areas 3. The drive for more specialized data is forcing companies like Scale AI to hire more labelers in the US and other western nations, which helps mitigate reputational challenges 4.

    Growth vectors scaling RLHF is in its early days RLHF has been touted as the leading contributor to model improvements since the original launches of Claud three, Gemini and GPT four.

    ---

    As demand grows, aggregating the supply allows companies to raise margins, and access to niche data prevents companies from bringing data sourcing in-house 4.

Related Episodes