Transformative Audio Synthesis

Nvidia has launched Fugato, a revolutionary generative model that blends music, voices, and sounds in unprecedented ways. This innovative technology creates unique audio experiences, such as saxophones barking or voices singing underwater. The development process involved complex relationships between audio and language, utilizing large language models to enhance sound synthesis capabilities. Although not publicly available yet, a sample website showcases its impressive range of functionalities.