Text-to-Speech

Characters: 0/1000
Speed
Language
Voices
🎤 Audio generation cost: 1 credit/10s

Text-to-Speech Feature Introduction

Text-to-Speech (TTS) is an AI technology that converts written text into natural-sounding speech. Our system supports multiple languages and voices, using advanced neural network models to provide high-quality speech synthesis services.

Supported Languages and Voices
  • Multi-language support: Supports 80+ languages including Chinese, English, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, Portuguese, and more
  • Rich voices: Built-in 400+ high-quality voices, including male, female, and different age groups
  • Voice cloning: Support using your own trained custom voices (requires separate training)
  • Multiple models: Supports A2E, Cartesia, Minimax, ElevenLabs and other TTS engines
How to Use
Input Text Content
Enter the content you want to convert in the text box, supports up to 1000 characters
Select Language and Voice
Choose the appropriate language, then select your preferred voice from the corresponding language voice list
Adjust Speed (Optional)
Adjust speech playback speed as needed, supports 0.5-2.0x speed
Generate and Download
Click the generate button, the system will create audio files that can be played and downloaded in history
Pricing Information
  • A2E Model: Free for VIP/MAX users, 1 credit/10s for regular users
  • Cartesia/Minimax Models: 2 credits/10s for all users
  • ElevenLabs Model: 3 credits/10s for all users
  • Voice Clone Training: A2E free, Cartesia 100 credits, Minimax/ElevenLabs 200 credits
Usage Tips
  • Text optimization: Use punctuation properly to help generate more natural speech rhythm
  • Language matching: Ensure the selected language matches the text content to avoid pronunciation errors
  • Voice selection: Choose appropriate voices based on content type (e.g., formal voices for news broadcasting)
  • Long text handling: For very long texts, consider processing in segments with reasonable length per segment
  • Special characters: Numbers, English abbreviations, etc. will be automatically converted to corresponding pronunciations based on language

Turn written content into AI speech

Create speech for narration, explainers, product demos, accessibility, and draft voiceovers. Select the language and voice that fit the script, then tune the pace before generating.

How to convert text to speech

  1. Enter your text

    Add up to 1,000 characters and use punctuation to shape pauses and rhythm.

  2. Choose a voice

    Select a language, a built-in voice or trained voice clone, and a speaking speed.

  3. Generate and download

    Create the audio, preview it in your recent results, and download the file.

Languages, voices, and playback controls

Browse voices by language and region, choose from the available speech engines, or reuse a trained custom voice. Speech-rate controls support slower delivery for tutorials and faster pacing for short-form content.

Train an AI voice clone

Text to speech FAQ

How much text can I convert at once?

The current text-to-speech editor accepts up to 1,000 characters for each generation.

Can I change the speaking speed?

Yes. You can adjust the speech rate from 0.5× to 2× before generating the audio.

Can I use my own cloned voice?

Yes. A voice trained in Voice Clone can be selected alongside the available built-in voices.

Can I download the generated speech?

Yes. Generated audio can be previewed and downloaded from My Results. The page currently shows results from the last 24 hours.