Help CentreFeaturesAudio & TTS Models
Features

Audio / TTS Models (17 total)

  • WhisperSpeech-to-text transcription
  • TTS-1Text-to-speech optimized for speed
  • TTS-1 HDText-to-speech optimized for quality
  • Gemini 3.1 Flash TTSPowerful, low-latency speech generation with expressive audio tags for precise narration control — 70+ languages
  • Google TTS StandardGoogle Cloud Text-to-Speech — standard voices, 40+ languages
  • Google TTS Neural2Google Neural2 voices — highly natural-sounding TTS using novel synthesis methods
  • Google Chirp3 HDGoogle's most expressive TTS — Chirp3 HD voices with studio-quality audio
  • Google TTS StudioGoogle Studio voices — highest quality, human-like expressiveness
  • ElevenLabs TTSElevenLabs Eleven v3 — ultra-realistic voice synthesis with 30+ languages and voice cloning
  • ElevenLabs FlashElevenLabs Flash v2.5 — lowest latency TTS for real-time applications, 32 languages
  • Qwen3 TTS FlashAlibaba's multilingual TTS with 49 voices, 10+ languages - ElevenLabs alternative
  • Qwen3 TTS Flash (Nov 2025)Snapshot version of Qwen3 TTS Flash with 49 voices
  • Qwen3 TTS Instruct FlashInstruction-controllable TTS - control speech style via text instructions, 10+ languages
  • Qwen3 TTS Voice DesignGenerate custom voices from text descriptions - design unique voices without audio samples
  • Qwen3 TTS Voice CloneClone voices from 10-20 second audio samples - highly natural voice replication
  • CosyVoice V3 PlusNext-gen generative TTS model - high-quality real-time streaming synthesis
  • CosyVoice V3 FlashFast CosyVoice TTS - cost-effective streaming synthesis
  • How to use: Go to the Audio page for text-to-speech, transcription, and voice features.

    Need more help with this topic?

    Ask Kunya AI