Compression and Replication
Bowen emphasizes that the primary focus of their research should be on compression rather than loss differential, especially as models scale. Jeffrey shares their journey of striving for replicability, detailing how they restarted their pre-training process from scratch to ensure accurate results. The team is committed to transparency, promising to publish their code and data in the upcoming distro paper.In this clip
From this podcast

AI + a16z
DisTrO and the Quest for Community-Trained AI Models
Related Questions