A significant effort has been made to collect a vast dataset of over 800GB, showcasing a diverse range of languages, including low-resource African languages. This initiative highlights the importance of community collaboration and expert curation, raising intriguing questions about data governance and model distribution. Observers note the fascinating evolution of this project and its potential impact on the field.