Audio transcriptions
POST /v1/audio/transcriptions — the OpenAI Audio contract.
Transcribe an uploaded audio file. The public model must support audio.transcribe; available
formats, timestamps, and streaming depend on its profile.
Example request
Set GATEWAY to the gateway origin and API_KEY to your virtual key. Replace the example
public model with one from Model discovery.
curl -X POST "$GATEWAY/v1/audio/transcriptions" \
-H "Authorization: Bearer $API_KEY" \
-F "model=transcribe" \
-F "[email protected]"Multipart only (file field). The uploaded file is streamed to temporary disk for validation — never
buffered fully in memory — and removed after success, failure, timeout, or client cancellation.
AUDIO_MAX_MULTIPART_BYTES (default 30 MB) caps the aggregate upload size.
Fields
model, file, language, languages, keywords, prompt, temperature, response_format (json, text, srt,
verbose_json, vtt), timestamp_granularities (word/segment, verbose_json only), include,
stream. Which response formats and whether streaming/timestamp granularities are available is
declared per model in its audio.transcribe catalog profile.
Vocabulary and language hints
keywords[] supplies literal words or phrases to help recognize names and domain terms;
languages[] supplies possible input languages. They are independent of prompt, which carries
free-form context. For example, add these fields to the request above:
-F 'keywords[]=Bifrost' \
-F 'keywords[]=AC-42' \
-F 'languages[]=es' \
-F 'languages[]=en'Repeated bare names (keywords, languages) are also accepted. A request may contain up to 1,024
text fields, including repeated hints. Keywords cannot contain <, >, carriage returns, or line
feeds. Send either language or languages, not both.
The gpt-transcribe catalog entries for OpenAI and Azure OpenAI enable both hints. Custom compatible
models can opt in through supportsKeywords and supportsLanguageHints in their audio.transcribe
profile. Hints without declared support are dropped for each selected deployment, including
fallbacks; keywords are never rewritten into a prompt. For models with supportsLanguageHints, a
singular language is sent upstream as a one-element languages[] list.
See OpenAI's transcription context guide.
Response format
Returned as JSON, plain text, or SSE depending on the requested response_format and stream — not
always JSON like the other endpoints.
When the provider reports detected languages, JSON responses and the final transcript.text.done
SSE event preserve languages: [{ "code": "es" }]. An empty list means no language was reliably
detected; an absent field stays absent.
Azure OpenAI note
Azure uses the classic deployment-based transcription API by default because some resources return
DeploymentNotFound from the v1 preview route for deployments that work through the classic route.
The adapter constructs the classic URL from your deployment name and resource baseUrl; v1 remains an
explicit transport override. See Providers → Azure OpenAI and the
Azure-specific example in Creating deployments.
Next steps
- Creating deployments — registering a transcription model.
- Streaming — the SSE variant of this endpoint.