OpenAI: GPT Transcribe API
GPT Transcribe is a speech-to-text model for completed audio files, streamed file transcripts, and committed turns in Realtime sessions over WebSocket. It supports unstructured context, keyword hints, and multiple language hints to improve transcription of domain terms, multilingual audio, and code-switching.
- Context window: 16,000 tokens
- Max output: 2,000 tokens
- Input: audio, text
- Output: text
- Tool calling: Supported
- File input: Supported
- Released: 2026-07-28
- Knowledge cutoff: 2024-05-31
Frequently Asked Questions
What is the context window of GPT Transcribe?
GPT Transcribe supports a context window of up to 16,000 tokens.
Does GPT Transcribe support function calling?
Yes. GPT Transcribe supports tool / function calling.
What is the knowledge cutoff of GPT Transcribe?
The knowledge cutoff of GPT Transcribe is 2024-05-31.