Section 5 of 7 — 14 models

Audio, Voice & Music

Song and instrumental generation, voice cloning and text-to-speech, and speech-to-text transcription.

Suno
  • Best-in-class vocals, capturing whispers, vibrato, and emotional nuance
  • Full song structure with proper verse, chorus, and bridge arrangement
  • Rap and spoken word still sound noticeably synthetic
  • No official API; the workflow is largely web-only
Udio
  • Inpainting lets you regenerate one section without redoing the whole track
  • Stem separation and an official API for paid tiers
  • Smaller credit allowances than Suno at comparable price points
  • API access requires a Pro-tier subscription
ElevenLabs
  • Trained on licensed catalogs, giving strong legal safety for commercial use
  • Realistic voice cloning and text-to-speech with broad API access
  • Music composition quality trails Suno and Udio
  • Generation is slower and pricier than most competitors
Google
  • Generates vocals with auto-written lyrics from text, image, or video prompts
  • High output quality for short-form music
  • Currently locked to the Gemini app with no public API
  • Limited to 30-second maximum clips
MiniMax
  • Most affordable API-based music option available
  • Handles niche genre details well, with full commercial rights
  • Third-party API routes often cap clip length well below native limits
  • Much smaller brand recognition than Suno or Udio
Stability AI
  • Generates both music and sound effects, useful for production work
  • Audio inpainting for fine-tuning specific sections
  • No vocal generation
  • Commercial use is restricted by revenue thresholds
OpenAI
  • Robust, widely-used open speech-to-text across many languages and accents
  • Free and self-hostable, with a large surrounding tool ecosystem
  • No built-in speaker diarization out of the box
  • Architecture is aging relative to newer transcription models
Murf
  • Granular voice control, including emphasis, pitch, pacing, and pronunciation via IPA
  • Bundles voice cloning, dubbing, and translation alongside core text-to-speech
  • Voice library size and language coverage are less clearly documented than rivals
  • Free-tier limits and character caps aren't fully transparent
Play.ht
  • Large voice library with strong multilingual coverage
  • Good API access for developers building voice into products
  • Voice realism can vary noticeably across less common languages
  • Pricing tiers can get expensive at high usage volumes
WellSaid Labs
  • Studio-quality, natural-sounding voices favored for corporate narration
  • Strong focus on brand-safe, licensed voice talent
  • Smaller voice selection than mass-market competitors
  • Positioned mainly at enterprise budgets, less accessible for casual users
Descript
  • Voice cloning built directly into a full audio and video editing workflow
  • Lets you edit spoken audio like text, moving words to reshape a recording
  • Requires recording your own voice samples to train a usable clone
  • Best value comes from Descript's editor, not the voice model alone
AIVA Technologies
  • Focused on instrumental composition for film, games, and content creators
  • Lets users guide style and structure with more compositional control than most
  • No vocal generation
  • Less mainstream brand recognition than Suno or Udio
Boomy
  • Extremely fast, one-click song creation aimed at total beginners
  • Built-in path to distribute finished tracks to streaming platforms
  • Creative control is shallow compared to Suno or Udio
  • Output quality trails the leading music generators
Amazon
  • Deep integration with AWS, useful for developers already on that stack
  • Reliable, low-cost text-to-speech at scale for IVR and accessibility use
  • Voice realism trails newer generative voice platforms like ElevenLabs
  • Best suited to utilitarian use cases rather than expressive narration