Text-to-Speech
Text-to-Speech Feature Introduction
Text-to-Speech (TTS) is an AI technology that converts written text into natural-sounding speech. Our system supports multiple languages and voices, using advanced neural network models to provide high-quality speech synthesis services.
Supported Languages and Voices
- Multi-language support: Supports 80+ languages including Chinese, English, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, Portuguese, and more
- Rich voices: Built-in 400+ high-quality voices, including male, female, and different age groups
- Voice cloning: Support using your own trained custom voices (requires separate training)
- Multiple models: Supports A2E, Cartesia, Minimax, ElevenLabs and other TTS engines
How to Use
Pricing Information
- A2E Model: Free for VIP/MAX users, 1 credit/10s for regular users
- Cartesia/Minimax Models: 2 credits/10s for all users
- ElevenLabs Model: 3 credits/10s for all users
- Voice Clone Training: A2E free, Cartesia 100 credits, Minimax/ElevenLabs 200 credits
Usage Tips
- Text optimization: Use punctuation properly to help generate more natural speech rhythm
- Language matching: Ensure the selected language matches the text content to avoid pronunciation errors
- Voice selection: Choose appropriate voices based on content type (e.g., formal voices for news broadcasting)
- Long text handling: For very long texts, consider processing in segments with reasonable length per segment
- Special characters: Numbers, English abbreviations, etc. will be automatically converted to corresponding pronunciations based on language
Turn written content into AI speech
Create speech for narration, explainers, product demos, accessibility, and draft voiceovers. Select the language and voice that fit the script, then tune the pace before generating.
How to convert text to speech
Enter your text
Add up to 1,000 characters and use punctuation to shape pauses and rhythm.
Choose a voice
Select a language, a built-in voice or trained voice clone, and a speaking speed.
Generate and download
Create the audio, preview it in your recent results, and download the file.
Languages, voices, and playback controls
Browse voices by language and region, choose from the available speech engines, or reuse a trained custom voice. Speech-rate controls support slower delivery for tutorials and faster pacing for short-form content.
Train an AI voice cloneText to speech FAQ
How much text can I convert at once?
The current text-to-speech editor accepts up to 1,000 characters for each generation.
Can I change the speaking speed?
Yes. You can adjust the speech rate from 0.5× to 2× before generating the audio.
Can I use my own cloned voice?
Yes. A voice trained in Voice Clone can be selected alongside the available built-in voices.
Can I download the generated speech?
Yes. Generated audio can be previewed and downloaded from My Results. The page currently shows results from the last 24 hours.