Model Training Insights
Nathan shares a lightbulb moment regarding the need for extensive generation and labeling from strong SFT models, highlighting a gap in current practices. Ross emphasizes the surprising effectiveness of sampling in reasoning tasks, referencing the outdated Galactica model and its impressive performance on benchmarks like GSMA K. The discussion reveals the complexities of model evaluation, particularly when training sets influence test results.In this clip
From this podcast

Interconnects Audio
Interviewing Ross Taylor on LLM reasoning, Llama fine-tuning, Galactica, agents
Related Questions