Source-led profile · Evidence checked 2026-08-26

Verified essentials

Voice-generation studio

OpenAI Audio

Choose an OpenAI Audio path using model-specific prices, limits, formats, consent, retention and a reproducible production test.

Decision first. Use the compact answer below before opening the complete research record.

Decision summary

The answer in one scan.

Decision-critical facts remain separate from the deeper editorial analysis.

Best fitprogrammable text-to-speech and audio generation within an OpenAI application stack
Free evaluationFree evaluation has not been established.
PricingAudio pricing is model- and token-dependent. The current pricing page must be checked against the exact speech or realtime model selected before cost modelling.
Commercial useCommercial-use eligibility has not been established.
Main cautionResolve the open evidence fields before buying.

From evidence to action

Make the OpenAI Audio decision with the right unit and route.

Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.

Which route fits you?

Architecture selector

OpenAI Audio: Chained endpoints versus Realtime based on latency, control, transcript and moderation needs.

Choose the closest route

The complete recommendation remains readable without JavaScript.

api · speech endpoint
POST /v1/audio/speech generates audio from text and accepts model, voice, input, format, speed and model-dependent instructions/streaming.
tts · voices
Current built-in voices are alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin and cedar, plus eligible custom voice IDs.
voice · custom access
Custom voices require an audio sample and prior consent recording and are limited to eligible customers.
voice · consent
A dedicated consent resource requires a named recording and BCP 47 language tag before custom voice creation.
stt · models
Current choices include GPT-Transcribe, GPT-4o Transcribe, GPT-4o Mini Transcribe, GPT-4o Transcribe Diarize and Whisper, with model-specific capabilities.
Route selector

Chained endpoints versus Realtime based on latency, control, transcript and moderation needs.

Keep in mind: Use only the retained OpenAI Audio evidence; do not generalize this decision asset to another product.
Sources and verification date

Continue your research

Move from profile to a sharper decision.

These links are explicit editorial relationships, not keyword matches or sponsored placements.

Good fit if

OpenAI Audio matches the job you need done

  • programmable text-to-speech and audio generation within an OpenAI application stack
  • The documented workflow and controls cover your required production steps.
  • You can validate the output with a representative project before committing.

Look elsewhere if

You need certainty this profile cannot provide

  • You need independently tested output quality rather than documented capabilities.
  • Your purchase depends on one of the 3 facts still requiring confirmation.
  • A narrower product would complete the same job with less workflow overhead.

Price, plan and risks

Confirm before you buy

Unknown, conflicted and stale facts stay visible before checkout.

Unknown

Self-serve evaluation

The retained evidence does not establish this field yet.

Unknown

Commercial-use eligibility

The retained evidence does not establish this field yet.

Unknown

Platforms

The retained evidence does not establish this field yet.

Commercial context

Compare the closest documented workflows.

Alternatives stay within the same vertical and use current internal profile routes.

Open the complete OpenAI Audio buying analysisTTS vs STT vs Realtime · cost scenarios · request limits · custom-voice consent · endpoint retention

“OpenAI Audio” is three buying decisions, not one product: text-to-speech, transcription and low-latency conversational voice. They share an API account but use different models, limits, transports and billing units. A team that needs narration should not cost-model Realtime tokens; a call agent should not be evaluated from a TTS demo; a transcription pipeline should not pay for speech-generation features it never uses.

> Distinctive strength: A developer can build either modular transcription/TTS or a low-latency speech-to-speech agent through one platform, choosing between simpler file endpoints and the Realtime transport rather than stitching together unrelated vendors first. > > Where it stops being an advantage: The routes use different models, transports, retention behavior and billing units, and OpenAI still does not provide the editorial, mastering and approval workflow around the generated audio.

OpenAI supplies infrastructure, not a complete voice-production studio. Script approval, pronunciation testing, retries, audio storage, mastering, reviewer workflow, audit logs and publishing remain your responsibility. That can be an advantage when audio belongs inside an existing application, and a poor fit when nontechnical producers expect a timeline editor and approval workspace.

OpenAI voice quality, transcription accuracy and Realtime latency remain unbenchmarked by BenPicks. Prices and capabilities below come from current official documentation; the evaluation protocol is designed to produce workload-specific evidence instead of turning model labels into a verdict.

Choose TTS, transcription or Realtime before comparing price

NeedCorrect surfaceBilling unitMaterial boundary
Generate narration from text`/v1/audio/speech`Characters for TTS-1/HD; model-specific for newer speech models4,096 characters per request; your application assembles long scripts
Transcribe files or streamed audio`/v1/audio/transcriptions`Published per-minute estimates vary by modelDiarization, timestamps and response options are model-specific
Build a live voice agentRealtime APIText and audio tokens, including cached-input rulesRequires WebRTC, WebSocket or SIP lifecycle and interruption handling
Clone a voiceEligible custom-voice APIsAccount/contract dependentRequires a consent recording; access is not automatic
Finished editor and approvalsNot the core productYour tooling and labourOpenAI Audio is an API layer, not a multitrack publishing suite

Is OpenAI Audio one product?

No. Start by choosing the job. Speech generation converts text into audio files or streams. Transcription converts recorded or live speech into text. Realtime combines audio and text input/output for interactive sessions, adding transport, turn detection, conversation context and failure recovery.

This separation prevents the most expensive planning mistake: comparing prices with incompatible denominators. TTS-1 is priced per character, GPT-Transcribe is displayed per audio minute, and Realtime is priced in text and audio tokens. There is no honest single “OpenAI Audio price.”

How much does OpenAI text-to-speech cost?

Current model pages list:

WorkloadTTS-1 at $15/1M charsTTS-1 HD at $30/1M chars
100,000 input characters$1.50$3.00
1,000,000$15.00$30.00
5,000,000$75.00$150.00

Those are generation charges, not finished-production costs. A correction-heavy 1M-character workload adds $3.75/$7.50 at 25% regeneration or $7.50/$15 at 50%. Add reviewer time, failed requests, storage, delivery, monitoring and mastering.

Track submitted characters per request and accepted audio per asset. Do not forecast from the final script alone: text may be sent repeatedly while pronunciation, pacing or instruction prompts are corrected.

What limits matter for long-form speech?

The current speech reference caps `input` at 4,096 characters. A long article or course therefore requires deterministic segmentation. Your application must decide where to break text, preserve pauses, retry failed segments, order results and avoid duplicating charged work.

The endpoint lists MP3, Opus, AAC, FLAC, WAV and PCM; MP3 is default. Speed ranges from 0.25 to 4.0. Newer supported speech models accept delivery instructions, but TTS-1 and TTS-1 HD do not. SSE streaming is likewise model-dependent and unavailable for those two models.

Test the exact model, voice and format. “Streaming available” does not prove time to first audio, full latency, reconnect behaviour or stable pacing across chunks.

Which voices and controls are available?

The current endpoint enumerates alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin and cedar, plus eligible custom-voice identifiers. A catalogue entry establishes availability, not pronunciation, identity consistency or target-language quality.

For each production language, test names, phone numbers, currencies, acronyms, dates, punctuation, borrowed words and code-switching. Blind-grade samples so the reviewer does not know the model. Record the model snapshot and full settings; otherwise a later rerun is not reproducible.

OpenAI's TTS guidance requires clear disclosure that the voice is AI-generated rather than human. Put this in the actual user experience before or alongside playback, not only in internal documentation.

Can you create a custom voice safely?

The API exposes separate voice-consent and custom-voice resources. Creating a voice requires an audio sample and a previously uploaded consent recording; custom voices are limited to eligible customers. A consent object has a name and BCP 47 language tag, and deletion endpoints exist for voice and consent objects.

That technical workflow does not replace your legal process. Obtain written, purpose-specific consent covering speaker identity, languages, channels, duration, client use, sublicensing, withdrawal and already distributed files. The developer material reviewed here did not establish a complete contract for backup deletion or withdrawal propagation. Keep those unresolved until your contract and deletion test answer them.

How much does transcription cost?

The current GPT-Transcribe model page shows an estimated $0.0045 per input audio minute; Whisper lists $0.006. At those rates:

Input audioGPT-TranscribeWhisper
100 minutes$0.45$0.60
1,000 minutes$4.50$6.00
10,000 minutes$45.00$60.00

Human correction often dominates, so calculate cost per accepted transcript, not API minute. Log corrections by category: names, terminology, numbers, speakers, punctuation and language switches.

GPT-Transcribe documents context, keyword and multiple-language hints. A dedicated GPT-4o Transcribe Diarize model identifies who spoke when. Do not assume every transcription model supports the same prompts, timestamps, diarization or output formats merely because all appear under Audio.

Is OpenAI Realtime cheaper than TTS?

That comparison is usually invalid. Realtime handles interactive audio/text sessions over WebRTC, WebSocket or SIP; batch TTS turns predetermined text into speech. GPT-Realtime currently lists text at $4 input/$16 output per million tokens and audio at $32 input/$64 output, with separate cached-input rates. History, interruptions, tool use, silence and retries change consumption.

The same page lists a 32,000-token context and 4,096 maximum output tokens. Long calls need an intentional truncation/history strategy. Test whether dropped context changes instructions or customer state, and measure tokens per successful call instead of converting a demo into minutes.

Turn detection is configurable. Evaluate talk-over, false starts, long silence, background noise, accents, numbers and mid-sentence interruption. Transport recovery, duplicate actions and human handoff belong to your product, not the model description.

What happens to audio and transcripts?

OpenAI's API data-control documentation says API inputs and outputs are not used to train models unless the customer explicitly opts in. Retention differs by endpoint:

  • `/v1/audio/transcriptions` and `/v1/audio/translations` list no abuse-monitoring or application-state retention.
  • `/v1/audio/speech` and `/v1/realtime` list up to 30 days of abuse-monitoring retention and no application state.
  • All four are listed as eligible for Zero Data Retention, but ZDR/MAM require approval and have limitations.

Regional controls are specific too. Documentation lists Audio/Voice storage and processing with service/model restrictions; Europe audio processing requires an approved retention-control mode such as MAM or ZDR. Verify endpoint, snapshot, project setting and regional hostname before sending regulated or confidential content.

The retained developer sources did not complete a procurement pack for certifications, DPA/subprocessors, SSO/SCIM, incident notification, audit rights or a universal Audio SLA. Request those separately. “ZDR eligible” is not evidence that your organization has ZDR enabled.

Is there a meaningful free trial?

Model pages expose a Free limits tier, but that does not prove every new organization receives monetary credits or enough allowance for a representative production test. Check the live account's credits and limits. Do not shrink the test until it stops representing production.

Limits vary by model and organization tier and may constrain requests, tokens or audio duration. Capture live limits before load testing, then test peak concurrency, retry headers, backoff and idempotency with a spending guard.

When one API platform is useful—and when a studio is simpler

Shortlist it when audio belongs inside a product your engineers already operate, API-level control matters and you can own editor, approval, monitoring and compliance layers. It is coherent when one application needs TTS, transcription or Realtime while treating them as separate cost and quality decisions.

Look elsewhere when producers need a ready-made timeline, project library, approval comments, pronunciation workspace and publishing flow without engineering. Pause when custom-voice consent, regional processing, enterprise assurance or an SLA is a hard gate not documented for your account.

Separate API, studio and localization products in the AI Voice directory, resolve rights questions with the commercial-use voice guide, then run the final speech API candidates through the TTS benchmark.

Three separate tests for three different audio jobs

  1. Choose one surface—TTS, transcription or Realtime—and freeze endpoint, model and snapshot.
  2. For TTS, use a rights-cleared multilingual script with names, numbers, currencies, acronyms and emotional direction.
  3. For transcription, use clean/noisy recordings, overlapping speakers, domain vocabulary and a human reference transcript.
  4. For Realtime, test the deployed transport, interruption, silence, reconnect and duplicate-action recovery.
  5. Log request size, character/minute/token use, time to first result, total latency, errors, retries and corrections.
  6. Blind-grade results with a second reviewer and record rejection reasons rather than a star rating.
  7. Re-run 25% and 50% correction scenarios and reconcile billed usage to request logs.
  8. Validate formats and chunk joins in the real player/editor; check loudness, silence, ordering and duplicates.
  9. For custom voices, complete consent before upload and test voice/consent deletion; retain the legal scope.
  10. Put AI disclosure in every playback entry point.
  11. Confirm training opt-in, endpoint retention, ZDR/MAM approval, region and model support in the real organization.
  12. Capture limits and load-test below a spending ceiling with backoff and idempotent retries.
  13. Obtain unresolved rights, assurance, support and SLA answers in writing before consequential production.
  14. Calculate cost per accepted narrated minute, corrected transcript or successful call, including engineering and reviewer time.

Pass only if the selected surface meets its own criteria. One narration says nothing about diarization; one accurate transcript says nothing about voice latency; one smooth demo says nothing about production recovery.

Official sources

Full evidence record7 fields · official links · dates · states

Documented limitations

Vendor claimChecked 2026-08-26

the speech endpoint documents a 4,096-character input limit; voice consent and disclosure requirements must be designed into the product workflow; latency, pronunciation and cost were not independently measured

Official sources (3)

Pricing context

Vendor claimChecked 2026-08-26

Audio pricing is model- and token-dependent. The current pricing page must be checked against the exact speech or realtime model selected before cost modelling.

Official sources (1)

Self-serve evaluation

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.

Commercial-use eligibility

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.

Platforms

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.