“OpenAI Audio” is three buying decisions, not one product: text-to-speech, transcription and low-latency conversational voice. They share an API account but use different models, limits, transports and billing units. A team that needs narration should not cost-model Realtime tokens; a call agent should not be evaluated from a TTS demo; a transcription pipeline should not pay for speech-generation features it never uses.
> Distinctive strength: A developer can build either modular transcription/TTS or a low-latency speech-to-speech agent through one platform, choosing between simpler file endpoints and the Realtime transport rather than stitching together unrelated vendors first. > > Where it stops being an advantage: The routes use different models, transports, retention behavior and billing units, and OpenAI still does not provide the editorial, mastering and approval workflow around the generated audio.
OpenAI supplies infrastructure, not a complete voice-production studio. Script approval, pronunciation testing, retries, audio storage, mastering, reviewer workflow, audit logs and publishing remain your responsibility. That can be an advantage when audio belongs inside an existing application, and a poor fit when nontechnical producers expect a timeline editor and approval workspace.
OpenAI voice quality, transcription accuracy and Realtime latency remain unbenchmarked by BenPicks. Prices and capabilities below come from current official documentation; the evaluation protocol is designed to produce workload-specific evidence instead of turning model labels into a verdict.
Choose TTS, transcription or Realtime before comparing price
| Need | Correct surface | Billing unit | Material boundary |
|---|
| Generate narration from text | `/v1/audio/speech` | Characters for TTS-1/HD; model-specific for newer speech models | 4,096 characters per request; your application assembles long scripts |
| Transcribe files or streamed audio | `/v1/audio/transcriptions` | Published per-minute estimates vary by model | Diarization, timestamps and response options are model-specific |
| Build a live voice agent | Realtime API | Text and audio tokens, including cached-input rules | Requires WebRTC, WebSocket or SIP lifecycle and interruption handling |
| Clone a voice | Eligible custom-voice APIs | Account/contract dependent | Requires a consent recording; access is not automatic |
| Finished editor and approvals | Not the core product | Your tooling and labour | OpenAI Audio is an API layer, not a multitrack publishing suite |
Is OpenAI Audio one product?
No. Start by choosing the job. Speech generation converts text into audio files or streams. Transcription converts recorded or live speech into text. Realtime combines audio and text input/output for interactive sessions, adding transport, turn detection, conversation context and failure recovery.
This separation prevents the most expensive planning mistake: comparing prices with incompatible denominators. TTS-1 is priced per character, GPT-Transcribe is displayed per audio minute, and Realtime is priced in text and audio tokens. There is no honest single “OpenAI Audio price.”
How much does OpenAI text-to-speech cost?
Current model pages list:
| Workload | TTS-1 at $15/1M chars | TTS-1 HD at $30/1M chars |
|---|
| 100,000 input characters | $1.50 | $3.00 |
| 1,000,000 | $15.00 | $30.00 |
| 5,000,000 | $75.00 | $150.00 |
Those are generation charges, not finished-production costs. A correction-heavy 1M-character workload adds $3.75/$7.50 at 25% regeneration or $7.50/$15 at 50%. Add reviewer time, failed requests, storage, delivery, monitoring and mastering.
Track submitted characters per request and accepted audio per asset. Do not forecast from the final script alone: text may be sent repeatedly while pronunciation, pacing or instruction prompts are corrected.
What limits matter for long-form speech?
The current speech reference caps `input` at 4,096 characters. A long article or course therefore requires deterministic segmentation. Your application must decide where to break text, preserve pauses, retry failed segments, order results and avoid duplicating charged work.
The endpoint lists MP3, Opus, AAC, FLAC, WAV and PCM; MP3 is default. Speed ranges from 0.25 to 4.0. Newer supported speech models accept delivery instructions, but TTS-1 and TTS-1 HD do not. SSE streaming is likewise model-dependent and unavailable for those two models.
Test the exact model, voice and format. “Streaming available” does not prove time to first audio, full latency, reconnect behaviour or stable pacing across chunks.
Which voices and controls are available?
The current endpoint enumerates alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin and cedar, plus eligible custom-voice identifiers. A catalogue entry establishes availability, not pronunciation, identity consistency or target-language quality.
For each production language, test names, phone numbers, currencies, acronyms, dates, punctuation, borrowed words and code-switching. Blind-grade samples so the reviewer does not know the model. Record the model snapshot and full settings; otherwise a later rerun is not reproducible.
OpenAI's TTS guidance requires clear disclosure that the voice is AI-generated rather than human. Put this in the actual user experience before or alongside playback, not only in internal documentation.
Can you create a custom voice safely?
The API exposes separate voice-consent and custom-voice resources. Creating a voice requires an audio sample and a previously uploaded consent recording; custom voices are limited to eligible customers. A consent object has a name and BCP 47 language tag, and deletion endpoints exist for voice and consent objects.
That technical workflow does not replace your legal process. Obtain written, purpose-specific consent covering speaker identity, languages, channels, duration, client use, sublicensing, withdrawal and already distributed files. The developer material reviewed here did not establish a complete contract for backup deletion or withdrawal propagation. Keep those unresolved until your contract and deletion test answer them.
How much does transcription cost?
The current GPT-Transcribe model page shows an estimated $0.0045 per input audio minute; Whisper lists $0.006. At those rates:
| Input audio | GPT-Transcribe | Whisper |
|---|
| 100 minutes | $0.45 | $0.60 |
| 1,000 minutes | $4.50 | $6.00 |
| 10,000 minutes | $45.00 | $60.00 |
Human correction often dominates, so calculate cost per accepted transcript, not API minute. Log corrections by category: names, terminology, numbers, speakers, punctuation and language switches.
GPT-Transcribe documents context, keyword and multiple-language hints. A dedicated GPT-4o Transcribe Diarize model identifies who spoke when. Do not assume every transcription model supports the same prompts, timestamps, diarization or output formats merely because all appear under Audio.
Is OpenAI Realtime cheaper than TTS?
That comparison is usually invalid. Realtime handles interactive audio/text sessions over WebRTC, WebSocket or SIP; batch TTS turns predetermined text into speech. GPT-Realtime currently lists text at $4 input/$16 output per million tokens and audio at $32 input/$64 output, with separate cached-input rates. History, interruptions, tool use, silence and retries change consumption.
The same page lists a 32,000-token context and 4,096 maximum output tokens. Long calls need an intentional truncation/history strategy. Test whether dropped context changes instructions or customer state, and measure tokens per successful call instead of converting a demo into minutes.
Turn detection is configurable. Evaluate talk-over, false starts, long silence, background noise, accents, numbers and mid-sentence interruption. Transport recovery, duplicate actions and human handoff belong to your product, not the model description.
What happens to audio and transcripts?
OpenAI's API data-control documentation says API inputs and outputs are not used to train models unless the customer explicitly opts in. Retention differs by endpoint:
- `/v1/audio/transcriptions` and `/v1/audio/translations` list no abuse-monitoring or application-state retention.
- `/v1/audio/speech` and `/v1/realtime` list up to 30 days of abuse-monitoring retention and no application state.
- All four are listed as eligible for Zero Data Retention, but ZDR/MAM require approval and have limitations.
Regional controls are specific too. Documentation lists Audio/Voice storage and processing with service/model restrictions; Europe audio processing requires an approved retention-control mode such as MAM or ZDR. Verify endpoint, snapshot, project setting and regional hostname before sending regulated or confidential content.
The retained developer sources did not complete a procurement pack for certifications, DPA/subprocessors, SSO/SCIM, incident notification, audit rights or a universal Audio SLA. Request those separately. “ZDR eligible” is not evidence that your organization has ZDR enabled.
Is there a meaningful free trial?
Model pages expose a Free limits tier, but that does not prove every new organization receives monetary credits or enough allowance for a representative production test. Check the live account's credits and limits. Do not shrink the test until it stops representing production.
Limits vary by model and organization tier and may constrain requests, tokens or audio duration. Capture live limits before load testing, then test peak concurrency, retry headers, backoff and idempotency with a spending guard.
When one API platform is useful—and when a studio is simpler
Shortlist it when audio belongs inside a product your engineers already operate, API-level control matters and you can own editor, approval, monitoring and compliance layers. It is coherent when one application needs TTS, transcription or Realtime while treating them as separate cost and quality decisions.
Look elsewhere when producers need a ready-made timeline, project library, approval comments, pronunciation workspace and publishing flow without engineering. Pause when custom-voice consent, regional processing, enterprise assurance or an SLA is a hard gate not documented for your account.
Separate API, studio and localization products in the AI Voice directory, resolve rights questions with the commercial-use voice guide, then run the final speech API candidates through the TTS benchmark.
Three separate tests for three different audio jobs
- Choose one surface—TTS, transcription or Realtime—and freeze endpoint, model and snapshot.
- For TTS, use a rights-cleared multilingual script with names, numbers, currencies, acronyms and emotional direction.
- For transcription, use clean/noisy recordings, overlapping speakers, domain vocabulary and a human reference transcript.
- For Realtime, test the deployed transport, interruption, silence, reconnect and duplicate-action recovery.
- Log request size, character/minute/token use, time to first result, total latency, errors, retries and corrections.
- Blind-grade results with a second reviewer and record rejection reasons rather than a star rating.
- Re-run 25% and 50% correction scenarios and reconcile billed usage to request logs.
- Validate formats and chunk joins in the real player/editor; check loudness, silence, ordering and duplicates.
- For custom voices, complete consent before upload and test voice/consent deletion; retain the legal scope.
- Put AI disclosure in every playback entry point.
- Confirm training opt-in, endpoint retention, ZDR/MAM approval, region and model support in the real organization.
- Capture limits and load-test below a spending ceiling with backoff and idempotent retries.
- Obtain unresolved rights, assurance, support and SLA answers in writing before consequential production.
- Calculate cost per accepted narrated minute, corrected transcript or successful call, including engineering and reviewer time.
Pass only if the selected surface meets its own criteria. One narration says nothing about diarization; one accurate transcript says nothing about voice latency; one smooth demo says nothing about production recovery.
Official sources