All tools

aiquick/text-to-speech

text-to-audio

Speech synthesis on the audio inference worker. Pick one of the built-in speakers or let the model choose, and get back a WAV clip you can play or download. Synthesis is slow — expect around a minute or two for a short passage — because each request generates the waveform from scratch.

Inference/audio/tts

Input

Text *What the model should say.
VoiceLeave on Auto to let the model pick. Each language has ten built-in speakers.

Synthesis runs on a shared GPU and is slow — usually one to two minutes for a short passage. Keep this tab open while it works.

Result

Idle
Fill in the input and run the tool to see results here.