Comparison · Evidence checked 2026-08-28

Deepgram vs Cartesia vs Rime for Realtime Text-to-Speech

A practical, evidence-led decision guide. Product capabilities and limits are separated from anything that would require hands-on testing.

# Deepgram vs Cartesia vs Rime: three ways to make an agent speak quickly

Deepgram Aura/Flux, Cartesia Sonic and Rime Mist/Coda are infrastructure choices for applications that must begin speaking before a long audio file has finished rendering. All three support streaming. All three publish language, model or latency claims. That shared vocabulary hides three different buying advantages.

The best product is not the one with the lowest vendor latency number. It is the one that speaks acceptably at the required concurrency, survives interruption and retry, fits the target language and has a data/rights contract the team can operate.

Start with the failure you cannot tolerate

Primary failureBest first testWhy
Existing Deepgram speech stack needs a TTS counterpartDeepgram Aura/FluxShared vendor operations and multiple model/transport choices may reduce integration work.
Context-rich WebSocket streaming and voice-asset lifecycle matterCartesia SonicContext IDs, timestamp events and cloning paths are central product surfaces.
The team must choose expressiveness versus minimum latencyRime Coda and Mist side by sideThe two model families are designed around different trade-offs.
Seven supported non-English languages are sufficientDeepgram Aura-2Verify exact voice and locale; Flux is an English route.
Broad public language catalogue is neededCartesia Sonic testSonic 3.5 documents 42 languages, subject to model/version reconciliation.
Private VPC or on-prem TTS is a hard gateRime or Deepgram enterpriseCompare actual deployment package, support, GPU/capacity and licence traffic.
Self-serve instant cloning is requiredCartesiaPro and higher document instant cloning; consent and deletion still require governance.
Public output-rights language must be complete before salesPause all three if unresolvedA working endpoint is not a commercial licence.

Model names matter more than brand names

Deepgram currently positions Flux as its recommended English TTS, Aura-2 as the seven-language family and Aura-1 as the lower-cost English generation. Aura's REST route accepts compressed/containerized combinations and limits input to 2,000 characters. Flux uses `/v2/speak`, a streaming-first turn-control design and raw `linear16`, μ-law or A-law output. Swapping a model string without changing endpoint and format assumptions can break an integration.

Cartesia's retained public pages need reconciliation: the technical model documentation details Sonic 3.5 and a dated snapshot, while the pricing calculator labels its included minutes as Sonic-3.6. That does not prove either page is false. It means pricing and behavior cannot be joined until the vendor identifies the production model ID, language/control set and credit rate.

Rime's choice is Coda versus Mist, but its model/language documentation also has drift. Coda is positioned for expressive multilingual conversation and word timestamps. Mist v3 focuses on pronunciation and low latency. Some current pages describe different language inventories and even speed parameters whose direction differs across generations. Pin the model image or API ID rather than purchasing “Rime TTS.”

Costs: characters, credits and privacy settings

Deepgram publishes straightforward PayGo character prices for Aura:

VolumeAura-1 at $0.015/1KAura-2 at $0.030/1K
100,000 characters$1.50$3
1 million$15$30
5 million$75$150

At one million Aura-2 characters, 25% regeneration makes synthesis $37.50; 50% makes it $45. The listed rates are tied to participation in Deepgram's Model Improvement Program. The streaming API exposes `mip_opt_out`; opting out can affect pricing, so a sensitive-data comparison needs a quote for the same privacy setting.

Cartesia sells credits through plans. The checked table shows Free $0/20,000 credits, Pro $5/100,000, Startup $49/1.25 million and Scale $299/8 million, with displayed Sonic capacities around 27, 133, 1,667 and 10,667 minutes. That display implies roughly 750 credits per minute, but the invoice and application operate on credits/characters and special features can use different rules. A Startup baseline consuming 1.25 million credits needs 1.5625 million at 25% regeneration and 1.875 million at 50%—already beyond the included pool.

Rime prices Mist at $0.03 and Coda at $0.05 per 1,000 characters in the checked public table. Ten million generated characters are therefore $300 or $500 before rejected output; 25% regeneration raises them to $375 or $625. The advertised free allowance is conflicted on the same public surface—approximately 800 minutes in one place and 3,000 in another—so it must not anchor the pilot until the account shows the balance and expiry.

These figures still omit the speech-to-text model, reasoning model, telephony, storage, monitoring and failed calls. Compare cost per completed acceptable turn, not cost per generated character alone.

Concurrency is an application property

Deepgram's PayGo documentation lists up to 15 concurrent Aura REST and 45 streaming requests per project in covered regions. Cartesia lists TTS context limits of 2/3/5/15 for Free/Pro/Startup/Scale. Rime Starter advertises 20 simultaneous generations.

None of those numbers equals simultaneous callers. A caller does not generate audio continuously; WebSocket contexts may serialize work; connection limits can differ from active synthesis; payload size and network geography change occupancy. Cartesia suggests a conversational application might support roughly four times its active TTS context count because speech is intermittent, but that is planning guidance, not guaranteed capacity.

Build a queueing test using the actual turn pattern. Step through 1, 5, 10, 20 and higher simultaneous sessions. Record active contexts, sockets, first playable audio, completion, 429s, retries, abandoned turns and cleanup after interruption. A leaked socket pool can fail while the generation counter looks healthy.

Voice controls and cloning separate the products further

Deepgram Aura-2 documents speed control and IPA pronunciation overrides for supported languages. The API can emit telephony encodings and broader file/container outputs depending on endpoint. Flux adds turn-oriented streaming behavior but a narrower raw-format route.

Cartesia supports byte, SSE and WebSocket endpoints, with timestamp events on eligible paths. Its instant clone can start from a short recording; guidance recommends a clean sample closer to ten seconds for stronger similarity. Professional Voice Cloning adds a separate training charge, 1.5 credits per generated character and retraining implications when the base model changes. It is a model-lifecycle commitment, not a premium toggle.

Rime Mist exposes pronunciation-oriented controls and text normalization choices. Disabling normalization may reduce processing time but can damage dates, currency, abbreviations and addresses. Coda emphasizes expressive conversation and timestamps. Enterprise advertises custom clones, but the public evidence retained for this review did not establish a complete consent, revocation and clone-deletion lifecycle.

If cloning is required, use one consenting internal speaker and test authorization, data access, old IDs and deletion. Do not infer safe governance from a five-second upload or an enterprise feature label.

Data posture can reverse a price result

Deepgram's public rate can depend on Model Improvement Program participation. Public evidence here did not establish one numeric retention/backup-deletion schedule for all TTS text, audio and metadata. Regional endpoints and Dedicated/self-hosted routes exist, but their controls and price belong in the actual quote.

Cartesia's ordinary terms permit use of inputs, outputs and interactions to improve models unless otherwise agreed. Its opt-out affects future training use and does not reverse prior use. Enterprise Zero Data Retention has eligibility and voice-asset exceptions that must be mapped to inference, clone samples, PVC datasets and operational logs.

Rime states zero content retention by default, no training by default and explicit opt-in through `trainableUtterance=true`; minimal connection/health logs may remain for 90 days. VPC and on-prem options can strengthen the boundary, but procurement must still review licence calls, images, updates, incident handling and responsibility for the deployment.

The cheapest public generation rate is irrelevant if the required privacy mode changes it or cannot be contracted.

A 1,000-turn benchmark for all three

Prepare 100 phrases covering acknowledgements, names, account codes, dates, currency, street addresses, abbreviations, code-switching, long sentences and interruption points. Replay ten times across a controlled distribution to produce 1,000 turns.

For every candidate:

  1. Pin endpoint, model/version, voice, region and audio format.
  2. Use connection reuse and payload sizes that match production.
  3. Measure first byte, first playable audio and completion separately.
  4. Blind-score pronunciation, intelligibility, suitability and take consistency.
  5. Interrupt at early/middle/late playback and confirm cancellation plus next-turn state.
  6. Step concurrency until the planned peak and one level above it.
  7. Inject disconnects, 429s and retries; reject duplicated/stale audio.
  8. Record generated, rejected and regenerated characters/credits in native units.
  9. Test required languages and telephony/media codecs independently.
  10. Run default and approved privacy configurations.
  11. Verify usage headers/logs reconcile with the account ledger.
  12. Calculate cost per accepted turn and per peak concurrent session.

Set thresholds before the run: p95 first audio, pronunciation error rate, interruption recovery, maximum 429/retry rate and cost ceiling. Keep listening scores separate from latency so a fast but unacceptable voice cannot hide behind one composite number.

The shortlist decision

Choose Deepgram when its model/transport choices fit an existing agent stack, seven-language Aura-2 or English Flux covers the need, and regional/MIP terms can be pinned. Aura-1 can be economical for suitable English workloads; do not treat it as equivalent to Aura-2 or Flux.

Choose Cartesia when context-based streaming, broader documented language coverage and a self-serve path from stock to instant/professional voices create real product leverage. Resolve Sonic 3.5/3.6 pricing identity and clone-data lifecycle before committing.

Choose Rime when the team values an explicit expressiveness-versus-latency model decision or genuinely needs VPC/on-prem deployment. Reconcile the free allowance, language/model matrix and output/clone rights in the signed service terms.

Choose more than one when fallback diversity is worth the integration cost. A second provider only improves resilience if the application has preapproved voices, normalized formats, tested retry/failover and a way to avoid speaking the same turn twice.

Which real-time stack earns the production pilot?

Deepgram, Cartesia and Rime are not three interchangeable low-latency voices. Deepgram offers model and transport choice, Cartesia organizes streaming around contexts and voice assets, and Rime makes model/deployment trade-offs explicit.

The correct winner appears only after the same 1,000 turns run at the same concurrency and privacy boundary. Anything less compares marketing surfaces rather than production systems.

Official sources checked

Sources checked 2026-08-28. Models, weights, rates, limits, language coverage and data terms can change; capture the exact production configuration and signed terms.