I have this image saved from 2023 when Elevenlabs first released. These were taken from their blog posts. It was also trained on only 32x3090s which is a surprisingly small amount of compute for a model that (imo) has been #1 at TTS for 2 years now. The key difference to me is that alternatives, like Kokoro, jump way too eagerly into synthetic data rather than using high-quality datasets: https://huggingface.co/posts/hexgrad/418806998707773 Training your AI on flawed TTS outputs will only get it as good as those flawed outputs. Elevenlabs trained on actual audiobook data and other high-quality voice sources. Elevenlabs early 2023 model is still a leap ahead of everyone else for voice cloning: https://youtu.be/pP35DxuAcac https://youtu.be/-gGLvg0n-uY https://youtu.be/kNipoNLC6Eg Start training on actual high quality data and you'll get there. More on reddit.com
No discussion yet. Be the first to share your thoughts!