Gateway APIs
Transcribe Audio
Convert an audio file to text in its original language.
POST
audio/transcriptions API, so OpenAI SDKs work
unchanged. The model is the name of your speech-to-text deployment, and the deployment decides
which provider and model transcribe the audio.
Headers
Form Data
Audio files
Response formats
verbose_json fields:
Segments and subtitles
Providers return timings per word, so segments are built from those words. A segment ends:- after a word that ends a sentence (
.,!or?); - before a word that would make it longer than 84 characters (two 42-character subtitle lines) or 7 seconds;
- before a word that follows a pause of 1 second or more;
- before a word spoken by a different speaker, when the deployment has diarization on.
verbose_json, srt and vtt share these segments. If a provider returns text but no word
timings, the transcript is one segment spanning the whole audio, and an x-bud-config-warning
says so. Silent audio returns no segments.
Response headers
Deployment settings
Everything on a deployment’s Settings → Audio page is a default for every request to that deployment. The page shows only the settings the deployment’s provider and model support. A saved setting that the provider cannot apply is skipped with anx-bud-config-warning, and the
request still succeeds.
A request can override most of them for itself, using the form field in the last column.
Settings that cost extra per request, and the noise filter, only the deployment controls. See
Overriding settings per request.
Overriding settings per request
The form fields in the table above are Bud additions to OpenAI’s request. All of them are optional, and a request that leaves them out gets the deployment’s settings, so plain OpenAI clients work unchanged. Each field applies to its own request only:extra_body={"diarization": True, "keyterms": ["Kubernetes", "Dapr"]} sends exactly the form
fields above.
- Switches take
trueorfalse(1or0also work). Any other value is refused with400naming the field. language_detection=trueandlanguagein the same request are refused with400. A request that names itslanguageon a deployment that detects the language uses that language instead, and says so in anx-bud-config-warning.keyterms[]andredaction[]take one value per form field, and add to the deployment’s list.- The two compliance settings only tighten. A request can switch
profanity_filteron and addredaction[]categories. It cannot switch off or remove what the deployment set; trying says so in anx-bud-config-warning. - A field the provider cannot use is skipped with an
x-bud-config-warning. - A field nothing recognises, such as a misspelling, is ignored and named in an
x-bud-config-warning. - Self-hosted deployments receive only OpenAI’s own fields, so they ignore these, with a warning.
Errors
Errors use OpenAI’s error format;param names the field at fault when there is one.
404