Voxmint › Blog

Using a Text-to-Speech API in Your App: A Practical Guide

To use a text-to-speech API in your app, create an API key, call the generate endpoint with your text and a voice, then fetch the resulting audio. The Voxmint developer API supports plain speech with ready-made voices, voice design, and cloning from reference audio, along with streaming and batch generation for larger workloads.

This guide walks through the practical decisions you will make when integrating it.

What you can build

A text-to-speech API turns text into audio on demand, so your own product can talk. Typical uses include:

  • Read-aloud buttons for articles and documentation.
  • Narration for videos, courses and onboarding flows.
  • Voice replies in chat or support tools.
  • Audiobook or podcast production pipelines.
  • Accessibility features for content that is otherwise text-only.

Voxmint offers 90 ready-made voices across 30 languages, so you can pick a voice per language or per use case. You can browse and preview voices on the text to speech page before choosing one in code.

Step 1: Create an API key

Create a key in the app under the Developer section. The secret is shown once, so store it right away in an environment variable or a secrets manager, never in client-side code or a public repository.

Keys have scopes. A read scope covers listing voices and downloading audio, while a write scope covers generating and changing things. Give each service the narrowest scope it needs. Key management itself requires an interactive login, so a leaked key cannot create more keys.

Every request carries the key as a bearer token in the Authorization header.

Step 2: Choose a voice

List the voice catalogue and filter by language, gender or accent. Each entry has an id you can use when generating speech. Voice ids are stable, so the voice you choose today keeps sounding the same later.

A sensible pattern is to pick voices once during development, store the ids in your configuration, and avoid searching the catalogue on every request.

Step 3: Generate your first audio

Generation is a single POST request. The key fields are:

  • text: what to say.
  • mode: speech for a ready-made voice, design to describe a voice in words, or clone modes when you have reference audio.
  • catalog_voice_id: the voice to use in speech mode.
  • style_prompt: optional guidance on emotion and pace, such as "slower, warmer".

The response describes the generation, including its status and duration. You then download the audio from the generation's audio endpoint as a WAV file. Always check that the status says it succeeded before fetching the audio.

Because the audio is a standard WAV file, it works in browsers, mobile players and editing software, and you can convert it to another format on your side if you need smaller files.

Step 4: Pick the right delivery method

Think about what your user experience needs:

  • Standard generation is simplest: send text, get a finished file. Good for pre-rendered content such as lessons or articles.
  • Streaming returns audio in chunks as it is produced, which gets sound to the user sooner. Good for interactive features where waiting for a whole file would feel slow.
  • Batch applies one set of settings to many lines, each producing its own clip, and a failing line does not stop the rest. Good for course libraries or lists of notifications.

Pick one per feature. A read-aloud button might stream, while a nightly job that renders new articles can use batch.

Step 5: Handle errors and limits

Production integrations spend most of their effort on failure cases, so build these in from the start:

  • 401 means the key is missing, wrong or revoked.
  • 403 means the key lacks the needed scope or the feature is not available to your account.
  • 422 means a field is invalid; the response says which one.
  • 429 means you are being rate limited. Wait for the time given in the Retry-After header, then retry.
  • 503 means the service is warming up; retry after the indicated delay.

Responses include a request id. Log it so that any problem report can be traced quickly. Use exponential backoff for retries rather than hammering the service, and make sure retries are safe for your use case, so a user does not end up with duplicate audio.

Step 6: Design for text length and cost

Generation takes time that grows with the amount of text, and requests are processed one at a time per model, so heavy bursts queue up. A few habits keep your app responsive:

  • Split long text into paragraphs or sentences and generate them separately.
  • Cache audio for content that does not change, such as an article, instead of regenerating it on every view.
  • Generate in the background and notify the user when audio is ready, if immediate playback is not required.
  • Preview your text before you generate if you want to check how descriptions and style prompts combine.

Check the current limits for your account in the product rather than hard-coding assumptions.

Step 7: Voices of your own

If you want your app to use a custom voice, you have two routes. Voice design lets you describe a voice in words and generate it. Voice cloning builds a voice from a few seconds of reference audio. Cloning reproduces a real person's voice, so only upload audio you have permission to use, and tell your users when speech is synthetic if they might otherwise be misled.

Security and privacy checklist

  • Keep API keys on your server and call the service from your backend.
  • Use separate keys per environment, and revoke any that leak.
  • Do not log full user text if it may contain personal information.
  • Get consent for every cloned voice and store proof of it.
  • Add a clear label where synthetic speech could be mistaken for a real person.

FAQ

Do I need a backend to use the API?

Yes, in almost all cases. Calling the API from the browser would expose your key. Put a small endpoint on your server that forwards requests.

Can I get word timings?

The API can return word timings when you request them, which is useful for highlighting text as it is read aloud.

Which languages can I use?

Voices cover 30 languages. See pages such as Spanish, French and Chinese for examples.

What audio format do I get back?

A WAV file. Convert it on your side if your app needs another format.

Ready to add a voice to your app? Create a key and try Voxmint today.