Futures of the data foundry business model

Topics covered
Popular Clips
Questions from this episode
- Asked by 100 people
- Asked by 36 people
- Asked by 26 people
- Asked by 13 people
- Asked by 11 people
- Asked by 4 people
- Asked by 2 people
- Asked by 2 people
- Asked by 1 person
Episode Highlights
Annotation
The methods and intricacies of data annotation for AI training are complex and costly. explains that basic instructions are cheap, but scientific reports are extremely expensive, contributing to the $10 million-plus data acquisition costs for fine-tuning 1. The process involves iterative training, evolving data instructions, and navigating bureaucratic challenges, making it difficult for new organizations to add human data to their pipelines 1.
The process is difficult for new organizations trying to add human data to their pipelines, given the sensitivity, processes that work and improve the models are extracted until the performance runs out.
---
Acquiring data from vendors is a high-stakes endeavor, often requiring legal or financial action to ensure delivery, and the data is delivered in weekly batches, with the goal of improving models over time 1.
  Â
Quality
Maintaining high-quality data for AI training presents significant challenges. Nathan notes that while synthetic data is becoming more prevalent due to its cost-effectiveness, human data remains crucial for certain tasks 1 2. The academic community is exploring how language models can replace humans for labeling preference data, but there are concerns about the quality and authenticity of data obtained from vendors 2.
Waning data quality through AI slop is increasingly hard to guarantee that data obtained from a data vendor is not already partially or fully coming from a language model.
---
The process of acquiring and utilizing human data involves iterative experimentation and high effort, with millions of dollars potentially wasted on datasets not used in final models 1.
  Â
Utilization
Companies strategically use data for AI training through innovative techniques and business strategies. Nathan highlights that data annotation companies are evolving to handle data collection and training tasks, with subsidiaries specializing in different areas 3. The drive for more specialized data is forcing companies like Scale AI to hire more labelers in the US and other western nations, which helps mitigate reputational challenges 4.
Growth vectors scaling RLHF is in its early days RLHF has been touted as the leading contributor to model improvements since the original launches of Claud three, Gemini and GPT four.
---
As demand grows, aggregating the supply allows companies to raise margins, and access to niche data prevents companies from bringing data sourcing in-house 4.
Related Episodes


A recipe for frontier model post-training
Answers 383 questions
We aren't running out of training data, we are running out of open training data
Answers 383 questions
A realistic path to robotic foundation models
Answers 383 questions
Frontiers in synthetic data
Answers 383 questions

On the current definitions of open-source AI and the state of the data commons
Answers 383 questions
Alignment-as-a-Service: Scale AI vs. the new guys
Answers 383 questions
Interconnects year in review: 2023
Answers 383 questions
Llama 3.2 Vision and Molmo: Foundations for the multimodal open-source ecosystem
Answers 383 questions
DBRX: The new best open LLM and Databricks' ML strategy
Answers 383 questions
OpenAI's Model (behavior) Spec, RLHF transparency, and personalization questions
Answers 383 questions
Model merging lessons in The Waifu Research Department
Answers 383 questions
OLMoE and the hidden simplicity in training better foundation models
Answers 383 questions

A post-training approach to AI regulation with Model Specs
Answers 383 questions
