Skip to main content
POST
The request and response follow OpenAI’s audio/speech API, so OpenAI SDKs work unchanged. The model is the name of your text-to-speech deployment, and the deployment decides which provider and voice model synthesize the audio.

Headers

Body

Which voice is used

Checked in this order; the first one that applies wins. If none of these gives a voice, the request is refused with 400 and param: "voice". Voice IDs are specific to each provider. An ID from one provider — alloy on an ElevenLabs deployment, for example — is refused with a 400 that names the voice and says how to fix it.

Output formats

Not every provider produces every format. When the deployment’s provider cannot produce the format you asked for, the request is refused with 400 and param: "response_format", and the message lists the formats that deployment can return. For example, ElevenLabs deployments return mp3, opus, wav and pcm. If a provider returns a different container from the one requested, the audio is served as returned: Content-Type and x-audio-format describe what you actually received, and an x-bud-config-warning says what changed.

Response headers

Deployment settings

Everything on a deployment’s Settings → Audio page is a default for every request to that deployment. The page shows only the settings the deployment’s provider and model support. A saved setting that the provider cannot apply is skipped with an x-bud-config-warning, and the request still succeeds. A request can override most of them for itself, using the field in the last column. The rest are operational settings that only the deployment controls. See Overriding settings per request. speed has no saved default: the request’s value, or the provider’s own default, is used.

Overriding settings per request

The fields in the table above are Bud additions to OpenAI’s request. All of them are optional, and a request that leaves them out gets the deployment’s settings, so plain OpenAI clients work unchanged. Each field applies to its own request only:
With an OpenAI SDK, pass them through the SDK’s extra-fields option, for example extra_body={"stability": 0.3} in Python.
  • A value that is out of range or not in the vocabulary is refused with 400 naming the field. Examples: stability above 1, an unknown emotion, or a sample_rate not in the list above.
  • A field the deployment’s provider cannot use is skipped with an x-bud-config-warning, as a saved setting would be.
  • pronunciations add to the deployment’s list. On the same word, the request’s entry wins.
  • A field nothing recognises, such as a misspelling, is ignored and named in an x-bud-config-warning.

Errors

Errors use OpenAI’s error format; param names the field at fault when there is one.
400

Supported providers

Any text-to-speech deployment works with this endpoint. The model catalog currently offers text-to-speech models from ElevenLabs, Deepgram, Cartesia, OpenAI, Azure, Amazon Polly, Google Cloud Text-to-Speech and Speechmatics.