Source-led profile · Evidence checked 2026-08-26

Verified essentials

Voice-generation studio

Aura

Compare Deepgram Aura-1, Aura-2 and Flux using current character pricing, languages, concurrency, MIP data controls and a reproducible TTS pilot.

Decision first. Use the compact answer below before opening the complete research record.

Decision summary

The answer in one scan.

Decision-critical facts remain separate from the deeper editorial analysis.

Best fitlow-latency speech for agents, IVR and other developer-controlled applications
Free evaluationFree evaluation has not been established.
PricingThe retained official documentation establishes model and request limits, but the exact current per-character rate remains an account/pricing-page check before purchase.
Commercial useCommercial-use eligibility has not been established.
Main cautionResolve the open evidence fields before buying.

From evidence to action

Make the Aura decision with the right unit and route.

Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.

The number that changes the decision

Agent traffic calculator

Aura: Calls, turns, duration, concurrency and transport form a real proof-of-concept plan.

pricing · aura1
Pay As You Go is USD 0.015 per 1,000 input characters; Growth is USD 0.0135 per 1,000.
pricing · aura2
Pay As You Go is USD 0.030 per 1,000 input characters; Growth is USD 0.027 per 1,000.
pricing · scenarios
PayGo costs for 100k/1M/5M characters are USD 1.50/15/75 on Aura-1 and USD 3/30/150 on Aura-2, before rework and infrastructure.
pricing · growth commitment
Growth starts at a USD 4,000 annual commitment and advertises lower rates, higher concurrency and priority support.
pricing · mip
Public rates are stated for requests opted into the Model Improvement Program; `mip_opt_out` defaults false and opt-out price impact must be calculated for the account.
input · character limit
Aura-1 and Aura-2 accept at most 2,000 input characters per request; longer payloads can return HTTP 413 without audio.
Decision-changing number

Calls, turns, duration, concurrency and transport form a real proof-of-concept plan.

Keep in mind: Use only the retained Aura evidence; do not generalize this decision asset to another product.
Sources and verification date

Continue your research

Move from profile to a sharper decision.

These links are explicit editorial relationships, not keyword matches or sponsored placements.

Good fit if

Aura matches the job you need done

  • low-latency speech for agents, IVR and other developer-controlled applications
  • The documented workflow and controls cover your required production steps.
  • You can validate the output with a representative project before committing.

Look elsewhere if

You need certainty this profile cannot provide

  • You need independently tested output quality rather than documented capabilities.
  • Your purchase depends on one of the 3 facts still requiring confirmation.
  • A narrower product would complete the same job with less workflow overhead.

Price, plan and risks

Confirm before you buy

Unknown, conflicted and stale facts stay visible before checkout.

Unknown

Self-serve evaluation

The retained evidence does not establish this field yet.

Unknown

Commercial-use eligibility

The retained evidence does not establish this field yet.

Unknown

Platforms

The retained evidence does not establish this field yet.

Commercial context

Compare the closest documented workflows.

Alternatives stay within the same vertical and use current internal profile routes.

Open the complete Aura buying analysisAura-to-Flux model choice · character economics · streaming limits · MIP privacy · regional deployment

“Deepgram Aura” is no longer one simple TTS choice. Deepgram currently positions Flux TTS for English, Aura‑2 for the widest language coverage, and first-generation Aura‑1 as the lower-cost English family. A buyer must choose model, endpoint, transport, format, region and Model Improvement Program setting before price or latency comparisons mean anything.

> Distinctive strength: Deepgram lets developers choose among model families and transports for low-latency agents and programmatic speech rather than one monolithic TTS tier. > > Where it stops being an advantage: Model, endpoint, format, region and improvement setting must be fixed before latency or price comparisons are meaningful.

Deepgram's voice quality, latency, uptime and listener-preference claims have not been reproduced by BenPicks under production traffic. Current pricing and developer documentation are used here to construct the test that could prove—or reject—them.

The Deepgram TTS decision in 60 seconds

QuestionSource-led answer
Best fitDevelopers building low-latency agents, IVR or programmatic speech who can govern an API and test voice quality themselves.
Which model for English?Current docs recommend Flux; verify its endpoint and raw streaming-format constraints.
Which model for more languages?Aura‑2 covers seven documented languages; selected Spanish voices code-switch with English.
Public priceAura‑1 $0.015/1k characters; Aura‑2 $0.030/1k on PayGo.
Main cost/privacy interactionPublished rates opt into MIP; `mip_opt_out` defaults false and may affect pricing.
Hard technical limitsAura REST has 2,000 characters/request; PayGo documents 15 REST or 45 streaming concurrent Aura requests/project.

Should you choose Aura‑1, Aura‑2 or Flux?

Use the operating requirement, not the newest name:

FamilyCurrent roleBuyer-critical boundary
Aura‑1Cheaper first-generation English TTSLowest public character rate; older family and English-only.
Aura‑2Widest-language familySeven languages and broader voice catalog; twice Aura‑1 PayGo character price.
FluxRecommended English TTSStreaming-first turn control; its WebSocket uses raw formats and `/v2/speak`.

Existing Aura integrations should not swap model strings blindly. Flux rejects Aura model names on `/v2/speak`; its streaming transport accepts raw `linear16`, `mulaw` and `alaw`, while compressed/containerized output belongs to batch REST. Aura remains appropriate when an existing English pipeline needs compressed/containerized streaming compatibility or when a supported non-English voice is required.

Run the same fixtures through all eligible families. A migration passes only if pronunciation, preference, interruption behavior, output format, latency and cost all meet the application gate.

How much does Deepgram Aura cost?

Current public character rates are:

ModelPay As You GoGrowth
Aura‑1$0.015 / 1,000 chars$0.0135 / 1,000
Aura‑2$0.030 / 1,000 chars$0.027 / 1,000

Growth starts at a $4,000 annual commitment and advertises higher concurrency and priority support. It is not automatically economical for TTS alone: divide the commitment by the actual blended savings across every Deepgram service used.

Reproducible PayGo scenarios

Input charactersAura‑1Aura‑2
100,000$1.50$3.00
1,000,000$15.00$30.00
5,000,000$75.00$150.00

Character cost is not finished-audio cost. Add normalization, segmentation, failed/retried requests, regenerated lines, storage/CDN, telephony, monitoring and human review. At one million Aura‑2 characters, 25% regeneration makes synthesis $37.50; 50% makes it $45.00.

The public price table says listed rates opt into the Model Improvement Program. The streaming API exposes `mip_opt_out`, default `false`, and warns of pricing effects. Request an exact quote for both settings. A privacy-required opt-out can change the apparent model economics.

Is there a free Deepgram TTS trial?

Deepgram advertises $200 in credit for new accounts. That is credit, not a permanent free tier. Confirm current expiration, eligible APIs, automatic billing safeguards and whether MIP opt-out changes consumption before entering a card or sensitive text.

Use the credit for variance, not one polished sentence: names, addresses, dates, money, account codes, acronyms, medical/legal terms, emotional turns, punctuation, long paragraphs and every target language/accent.

Which languages and voices does Aura‑2 support?

Current documentation lists English, Spanish, German, French, Dutch, Italian and Japanese. English includes several accents; selected Spanish voices—Aquila, Carina, Diana, Javier and Selena—support English–Spanish code switching. Flux is currently English-only.

“Language supported” is not a quality verdict. For each market, test native reviewers, regional names/numbers, code-switch boundaries, abbreviations and unacceptable substitutions. Keep model IDs pinned and retain the audio because catalog/model behavior can change.

Deepgram's retained public Aura/Flux catalog does not document a self-serve cloning workflow. Buyers needing a branded custom voice should not infer cloning from the size of the stock catalog; request a separate written enterprise proposal and consent/deletion controls.

What are the input and concurrency limits?

Aura‑1 and Aura‑2 REST requests accept at most 2,000 input characters. Oversize payloads can return HTTP 413 without audio. PayGo documentation lists:

  • up to 15 concurrent Aura/Aura‑2 REST requests per project in North America, Europe and Australia;
  • up to 45 concurrent Aura/Aura‑2 streaming requests per project in those regions;
  • HTTP 429 when concurrency is exhausted.

Splitting text is not neutral. A poor boundary can change pauses, emphasis or pronunciation. Test sentence/paragraph segmentation, preserve punctuation, and compare the join against a short single-request control. Record retry count and duplicated billing/output behavior.

Do not size from average concurrency. Replay peak bursts with stepped 5/10/15 REST or larger streaming concurrency, record P50/P95/P99 time-to-first-byte and completion, and verify backoff under 429.

How do REST and WebSocket synthesis differ?

Aura REST `/v1/speak` streams audio in the response, so playback can begin at the first byte. Headers expose request ID, model name/UUID and input character count—useful for invoice reconciliation and support.

Aura WebSocket accepts `Speak`, `Flush`, `Clear` and `Close`. `Clear` can discard buffered text; graceful Close finishes/terminates the connection. Flux adds explicit turn-oriented `Interrupt` and `Configure`, assigns speech IDs and reports timing/billing metadata. These controls matter for user barge-in: measure how much audio still plays after an interrupt and whether the next turn preserves natural context.

Test abnormal closure, connection retry, duplicated frames, timeouts, empty flushes and cancellation. A demo that works on a quiet connection does not prove resilient conversation behavior.

Which audio formats and controls are available?

Aura supports telephony-friendly raw encodings and broader compressed/containerized outputs such as MP3, Opus, FLAC, AAC and WAV-compatible combinations. Valid sample-rate, encoding and container combinations differ; verify the exact downstream player/telephony path.

Aura‑2 speaking speed runs from 0.7 to 1.5 where the language supports it. Speed changes do not alter character billing. IPA pronunciation overrides allow up to 500 entries/request and 128 characters per IPA string; Deepgram bills the underlying word rather than markup.

Treat those controls as testable configuration, not automatic quality. A pronunciation dictionary can fix a drug or surname while damaging inflection in another context. Maintain fixture-based regression tests for every override and version them with the application.

Is Aura‑2 really sub‑200ms and better quality?

Deepgram markets sub‑200ms time to first byte and reports a blinded Aura‑2 preference study. Both are vendor evidence. They are useful hypotheses, not BenPicks results or universal rankings.

Latency depends on client location, selected endpoint, payload length, output format, connection reuse, concurrency and network. Listening preference depends on language, domain, voice, prompt text and evaluator mix. Reproduce both:

  1. warm and cold REST plus persistent WebSocket;
  2. EU/US/AU clients against appropriate endpoints;
  3. short acknowledgements, long sentences and 2,000-character payloads;
  4. normal and peak concurrency;
  5. native blind listeners with randomized vendor/model labels;
  6. pronunciation error, preference, intelligibility and unacceptable-error rate.

Report distributions, not the best sample.

What does MIP mean for sensitive text?

The API defaults `mip_opt_out` to false. That means the buyer must deliberately decide whether requests participate in Deepgram's Model Improvement Program, then verify both data handling and price for the chosen mode.

Public sources inspected here did not establish a numeric TTS input/output/metadata retention schedule or backup deletion period. Before sending health, financial, legal or customer conversation text, obtain written answers on retention, storage region, human access, subprocessors, training/improvement use, opt-out enforcement, logs and deletion.

Do not rely on a code comment. Add an automated check that every request/session uses the approved MIP setting and retain non-sensitive metadata proving it.

What security and deployment controls are documented?

Deepgram states SOC 2 Type I and II and GDPR readiness; the reports require request and scope review. It offers an EU endpoint (`api.eu.deepgram.com`) and Australia endpoint, plus enterprise Dedicated and self-hosted/VPC/on-prem deployments. It also describes a conditional HIPAA business-associate path for qualifying customers.

Those are useful procurement signals, not inherited compliance. Verify:

  • certificate period and services/endpoints in scope;
  • signed DPA/BAA where needed;
  • which telemetry leaves a regional or self-hosted deployment;
  • encryption/key/access/logging controls;
  • vulnerability and incident-notice obligations;
  • model/update rollback for self-hosted runtime;
  • infrastructure/GPU sizing and operational ownership;
  • SLA, RTO, RPO and support response.

PayGo priority/support and recovery commitments were not established in the public sources reviewed.

Can you use Deepgram TTS output commercially?

Do not infer an unconditional right merely because the API is paid. The public terms establish service use and customer responsibility for Content, but the retained text did not yield a simple unconditional TTS-output ownership warranty comparable across every use case.

Before commercial deployment, confirm output-use rights, input rights, voice/personality restrictions, prohibited impersonation, third-party claims, indemnity and survival after termination. Stock voices reduce cloning-consent complexity but do not remove content, publicity or deceptive-use duties.

Who should shortlist Deepgram TTS?

Shortlist it when:

  • low-latency agents/IVR are the primary job;
  • character pricing fits predictable text volume;
  • seven Aura‑2 languages or English Flux meet the catalog need;
  • regional/dedicated/self-hosted deployment matters;
  • engineers can govern streaming state, retries, formats and MIP;
  • the team will run native blind listening and load tests.

Look elsewhere or combine providers when you need self-serve cloning, languages outside the current catalog, a no-code media studio, theatrical expressiveness already proven for your material, or procurement that cannot resolve data/rights/SLA terms.

Compare API economics with Amazon Polly and inspect the wider AI Voice category. Use the TTS API benchmark guide to keep fixtures and metrics vendor-neutral.

A fair Deepgram evaluation protocol

  1. Choose Aura‑1, Aura‑2 and/or Flux according to language, format and endpoint—not novelty.
  2. Build fixed fixtures for every domain/language/accent, including codes, names, dates, money and interruptions.
  3. Blind native listeners and record preference, word/pronunciation errors and hard failures.
  4. Measure TTFB and completion P50/P95/P99 at stepped concurrency by region and transport.
  5. Exercise 2,000-character boundaries, safe splitting, all required formats/rates and speed/pronunciation controls.
  6. Test 429/backoff, disconnect/reconnect, Clear/Interrupt/Close and duplicate-output protection.
  7. Correlate `dg-request-id`, model UUID and `dg-char-count` with your usage ledger.
  8. Calculate baseline, +25% and +50% regeneration for PayGo/Growth and MIP opt-in/out.
  9. Verify regional routing, assurance documents, data retention/deletion, rights, SLA and full teardown in writing.
  10. Freeze the selected model/config/dictionary and rerun regression fixtures before every change.

The pass condition is not the fastest demo clip. It is acceptable listener quality and pronunciation at peak latency/cost, with resilient interruption behavior and a data/rights contract your organization can actually enforce.

Official sources checked

Sources checked 2026-08-28. Preserve the selected model UUID, endpoint, MIP setting and pricing page with every benchmark result.

Full evidence record7 fields · official links · dates · states

Documented capabilities

Vendor claimChecked 2026-08-26

REST and streaming synthesis; Aura-2 coverage across seven languages; speed and IPA pronunciation controls for supported Aura-2 languages; regional and self-hosted deployment options

Official sources (3)

Documented limitations

Vendor claimChecked 2026-08-26

Aura requests are limited to 2,000 input characters; Flux and Aura have different language, endpoint and output-format boundaries; voice quality and production latency were not independently tested

Official sources (3)

Pricing context

Vendor claimChecked 2026-08-26

The retained official documentation establishes model and request limits, but the exact current per-character rate remains an account/pricing-page check before purchase.

Self-serve evaluation

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.

Commercial-use eligibility

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.

Platforms

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.