Skip to main content
POST
The request and response follow OpenAI’s audio/transcriptions API, so OpenAI SDKs work unchanged. The model is the name of your speech-to-text deployment, and the deployment decides which provider and model transcribe the audio.

Headers

Form Data

Audio files

Response formats

verbose_json fields:

Segments and subtitles

Providers return timings per word, so segments are built from those words. A segment ends:
  • after a word that ends a sentence (., ! or ?);
  • before a word that would make it longer than 84 characters (two 42-character subtitle lines) or 7 seconds;
  • before a word that follows a pause of 1 second or more;
  • before a word spoken by a different speaker, when the deployment has diarization on.
verbose_json, srt and vtt share these segments. If a provider returns text but no word timings, the transcript is one segment spanning the whole audio, and an x-bud-config-warning says so. Silent audio returns no segments.

Response headers

Deployment settings

Everything on a deployment’s Settings → Audio page is a default for every request to that deployment. The page shows only the settings the deployment’s provider and model support. A saved setting that the provider cannot apply is skipped with an x-bud-config-warning, and the request still succeeds. A request can override most of them for itself, using the form field in the last column. Settings that cost extra per request, and the noise filter, only the deployment controls. See Overriding settings per request.

Overriding settings per request

The form fields in the table above are Bud additions to OpenAI’s request. All of them are optional, and a request that leaves them out gets the deployment’s settings, so plain OpenAI clients work unchanged. Each field applies to its own request only:
With an OpenAI SDK, pass them through the SDK’s extra-fields option. In Python, extra_body={"diarization": True, "keyterms": ["Kubernetes", "Dapr"]} sends exactly the form fields above.
  • Switches take true or false (1 or 0 also work). Any other value is refused with 400 naming the field.
  • language_detection=true and language in the same request are refused with 400. A request that names its language on a deployment that detects the language uses that language instead, and says so in an x-bud-config-warning.
  • keyterms[] and redaction[] take one value per form field, and add to the deployment’s list.
  • The two compliance settings only tighten. A request can switch profanity_filter on and add redaction[] categories. It cannot switch off or remove what the deployment set; trying says so in an x-bud-config-warning.
  • A field the provider cannot use is skipped with an x-bud-config-warning.
  • A field nothing recognises, such as a misspelling, is ignored and named in an x-bud-config-warning.
  • Self-hosted deployments receive only OpenAI’s own fields, so they ignore these, with a warning.

Errors

Errors use OpenAI’s error format; param names the field at fault when there is one.
404

Supported providers

Any speech-to-text deployment works with this endpoint. The model catalog currently offers speech-to-text models from Deepgram, ElevenLabs, AssemblyAI, Speechmatics, Gladia, Rev AI, Groq, OpenAI, Azure, Cartesia, Amazon Transcribe and Google Cloud Speech-to-Text.