Cartesia Sonic is designed for streamed speech inside conversational agents and interactive applications. Its buying value is not a polished demo voice. It is the combination of pinned model behavior, first-byte latency, concurrency, predictable billing and a defensible data/voice lifecycle inside your own product.
> Distinctive strength: Sonic is designed around streaming speech for interactive agents, where first-byte behavior, concurrency and pinned model operation matter. > > Where it stops being an advantage: A fast demo does not establish production latency, reliability, accepted-output cost or voice-data governance.
Cartesia also sells adjacent speech-to-text and voice-agent services. This profile isolates Sonic text-to-speech and the voice assets it depends on. Production latency, naturalness, reliability and clone behavior remain unscored until the product-specific pilot is run.
The Cartesia Sonic decision in 60 seconds
| Buyer question | Source-led answer |
|---|
| Best fit | Developers building real-time agents, support flows, games or interactive products. |
| Public entry price | Free $0; Pro $5; Startup $49; Scale $299 per month; Enterprise custom. |
| Commercial use | Pro and higher according to pricing and plan-gated Terms. Free is not the commercial tier. |
| TTS concurrency | 2 / 3 / 5 / 15 concurrent contexts by Free / Pro / Startup / Scale. |
| Voice cloning | Instant from Pro; professional from Startup, with separate creation/use economics. |
| Main caution | Pricing says Sonic-3.6 while public model docs currently detail Sonic 3.5. Confirm the actual model ID and credit rate. |
| Data caution | Enterprise Zero Data Retention covers TTS/STT inference, not cloning/PVC/voice creation. |
What exactly is Cartesia Sonic?
Sonic is Cartesia's streaming TTS model family. Developers submit text, a model and a voice, then receive audio through one of three endpoint styles:
- bytes: straightforward raw/file stream when timestamps are unnecessary;
- SSE: streamed JSON events with timing information;
- WebSocket: reusable bidirectional connection and context IDs for conversational flows.
The surrounding product—script approval, turn taking, interruption, transcript storage, moderation, mastering and rollback—belongs to your application. Cartesia is infrastructure, not an audiobook or video-production workspace.
Which Sonic model will production actually use?
Current public pages need reconciliation. The technical model page details Sonic 3.5, 42 languages and a dated stable snapshot. The pricing calculator labels included TTS minutes as Sonic-3.6. That does not prove either page is wrong, but it prevents an honest assumption that the documented 3.5 behavior and displayed 3.6 economics are identical.
Cartesia documents three version styles:
- base ID routes to the newest stable snapshot;
- dated snapshot remains fixed;
- latest/preview track may be hot-swapped.
Pin the dated model and `Cartesia-Version` used in acceptance testing. Before purchase, obtain written confirmation of model ID, languages, controls, credits per character and retirement notice. A “latest” alias is convenient for exploration and inappropriate as the only production reproducibility control.
What languages, formats and timing data are available?
Sonic 3.5 documents 42 languages, including English, Romanian, Arabic, Hebrew and multiple Indic languages. Coverage does not establish equal pronunciation, emotion or latency in each one.
The API supports raw, WAV and MP3. Documented PCM formats include floating-point/signed 16-bit, μ-law and A-law with 8, 16, 22.05, 24, 44.1 and 48 kHz sample rates. SSE/WebSocket can provide word or segment timestamps subject to model/language support.
Test the exact production codec. A telephony agent may need 8 kHz μ-law; media output may use WAV at 44.1/48 kHz. Transcoding after generation adds latency and potential quality differences that do not appear in a playground comparison.
How much does Cartesia Sonic cost?
The public monthly pricing page currently shows:
| Plan | Price | Credits | Displayed Sonic minutes | TTS concurrency |
|---|
| Free | $0 | 20,000 | ~27 | 2 |
| Pro | $5 | 100,000 | ~133 | 3 |
| Startup | $49 | 1,250,000 | ~1,667 | 5 |
| Scale | $299 | 8,000,000 | ~10,667 | 15 |
| Enterprise | Custom | Custom | Custom | Custom |
These figures imply roughly 750 credits per displayed minute, but do not budget real operations. Calculate from generated characters and preserve usage records. Include:
- greeting, prompts and dynamic replies;
- retries after disconnects/errors;
- pronunciations and editorial corrections;
- test/staging traffic;
- model migrations and A/B tests;
- infill and cloning charges;
- idle/failed request behavior;
- seasonal peak concurrency.
If a baseline consumes 1.25M credits, 25% regeneration needs about 1.5625M and 50% needs 1.875M before special-feature charges. That already exceeds Startup's included credits even though the initial workload fits it.
Infill has a separate published rule: one credit per inserted character plus 300 credits each operation. Frequent small edits can therefore have a very different cost curve.
How many simultaneous conversations can it support?
Cartesia counts TTS concurrency by active context. HTTP requests are individual contexts. On WebSocket, requests sharing a `context_id` process sequentially and do not add concurrency. Exceeding the plan limit returns 429.
The vendor suggests conversational agents may support roughly four times the TTS concurrency because speech generation is intermittent. Treat that as capacity-planning guidance, not a guarantee. For batch generation, parallel jobs map much more directly to the limit.
WebSocket connections are capped at ten times TTS concurrency. Idle TTS sockets close after five minutes. A leaked connection pool can hit socket limits even when generation is low.
Load-test with real turn lengths and silence distributions. Record p50/p95/p99:
- time to first audio byte;
- time to final audio;
- active contexts/connections;
- 429 and reconnect rate;
- truncated/duplicated audio;
- credit usage per completed turn.
Do not infer production capacity from a single vendor latency figure.
What voice controls should be tested?
Cartesia documents expressive/pacing controls, contextual pronunciation and pronunciation dictionaries, but controls are model/version dependent. Build a retained fixture containing:
- names and brand terms;
- email, telephone, order and confirmation codes;
- currencies, dates and abbreviations;
- heteronyms and punctuation;
- language switching;
- short acknowledgements and long explanations;
- interruption at multiple audio positions.
Ask native reviewers to blind-score intelligibility, appropriateness, consistency and correction effort. Preserve the exact model, voice, controls and output format for every sample.
How does instant voice cloning work?
Pro and higher include instant cloning. The API reference suggests about five seconds; guidance says ten seconds is the high-similarity sweet spot and warns that background noise and pauses can be reproduced. Longer is not automatically better.
The source clip therefore becomes an output-control surface. Use a quiet, trimmed, target-language recording that reflects desired pace and energy. Store:
- speaker identity and signed authorization;
- source provenance and permitted uses;
- recording date/languages;
- hash of the submitted clip;
- returned voice ID and access visibility;
- approved applications/territories/term;
- revocation and deletion actions.
Terms prohibit using another person's voice without express permission, including a deceased person or political candidate. The public material inspected does not establish a complete identity/consent-verification workflow inside the product, so your own gate remains necessary.
Is Professional Voice Cloning economically different?
Yes. Startup and above expose Professional Voice Cloning (PVC). Current docs list:
- 1,000,000 credits after successful training;
- 1.5 credits per generated character;
- fine-tuning tied to a base model;
- retraining and another 1M-credit charge for a newer base model.
Audio guidance needs clarification: one table says three hours, while the detailed workflow says minimum 30 minutes and recommends two hours. Training is described as about three hours and produces four voices from source clips.
PVC is a model-lifecycle commitment, not merely a better voice button. Budget recording/editing, consent, failed training, evaluation, 50% higher inference credits and future retraining. Test whether an instant clone is sufficient before accepting that burden.
Who owns output and can it be used commercially?
Cartesia says it does not claim ownership of customer inputs or outputs. Outputs may be non-unique, and the user remains responsible for legality, accuracy and third-party rights.
Commercial use is expressly tier-gated: pricing includes it from Pro, while Terms permit commercial use only when the subscription tier allows it. This does not clear trademarks, scripts, performances, publicity rights or the cloned speaker's consent. Attach the intended use to the plan and obtain legal review for consequential deployments.
Does Cartesia use customer data to train models?
The public Terms allow inputs, outputs and interactions to train/improve models unless otherwise agreed. A form can opt selected categories out of future training use, but prior uses and improvements are unaffected.
Enterprise Zero Data Retention changes inference handling:
| Data under ZDR | Retained? |
|---|
| TTS text input | No |
| TTS audio output | No |
| STT audio input | No |
| STT transcript output | No |
| Request IDs, usage, account/service metadata | Yes |
ZDR explicitly excludes voice cloning, PVC and voice creation because those workflows need retained source material or derived artifacts. It can also remove Playground history and limit support recovery.
Do not describe the account as “zero retention” without that qualification. Separate ordinary TTS/STT requests from voice-asset workflows in the data inventory.
Can a cloned voice be fully deleted?
The API provides `DELETE /voices/{id}`. That establishes deletion of a voice resource, not the timing/scope for uploaded source audio, derived embeddings/fine-tuned models, reusable PVC datasets, logs and backups. ZDR's explicit cloning exclusion makes this a procurement question.
Require deletion documentation and test it with identifiers. If a speaker revokes consent, the runbook should locate every clone/PVC/dataset and dependent application—not only hide a voice in the Playground.
What security controls matter?
Cartesia's ZDR documentation states GDPR, SOC 2 Type II, PCI-DSS service-provider and HIPAA compliance. Review the current assurance artifacts and scope rather than repeating badges.
For browser/mobile apps, Cartesia documents short-lived access tokens and warns against embedding API keys. Keep key/token creation server-side, scope permissions, set expiry, rotate/revoke and prevent tokens from appearing in logs.
The 2026 changelog says requests route by origin to US/EU/APAC. Routing is not the same as contractual data residency. Confirm regions, subprocessors, transfer mechanism, incident notice, encryption, access logs and Enterprise SLA.
Who should shortlist Cartesia Sonic?
Shortlist Sonic when your product needs streamed speech, API-level timing/control, moderate-to-high concurrency and engineers who can own versioning, sockets, retries, consent and data governance.
Look elsewhere when:
- you need an all-in-one audiobook/video editor;
- commercial use must remain free;
- long-form project collaboration is the core workflow;
- no one can run load and failover tests;
- cloned-voice deletion/consent cannot be governed;
- Enterprise data terms cannot be obtained for sensitive speech.
Continue with the AI Voice directory, the commercial AI voice guide and the canonical software comparison workspace.
A product-specific Cartesia pilot
- Pin Cartesia API and dated Sonic model versions; record the billed model label.
- Build multilingual fixtures with names, codes, dates, emotions and interruptions.
- Generate bytes, SSE and WebSocket outputs in production codecs.
- Measure p50/p95/p99 latency, 429s, reconnects and completed-turn credits.
- Blind-review audio and calculate correction/regeneration cost.
- Saturate concurrency and socket limits; prove pool cleanup and backoff.
- Create one consented instant clone; test visibility, authorization and deletion.
- Compare PVC only if instant cloning fails; include retraining lifecycle cost.
- Verify training opt-out or Enterprise ZDR on eligible inference; document exclusions.
- Simulate snapshot retirement, timeout, malformed input and regional failure.
Pass only if the application—not just the model—meets latency, intelligibility, legal, cost and recovery thresholds.
Main unknowns before purchase
- whether production pricing currently maps to Sonic-3.5 or Sonic-3.6 and its exact credit rate;
- deletion timing/scope for clone source data, PVC models/datasets and backups;
- contractual availability/support response SLA;
- fixed regional processing/residency and subprocessor scope;
- Enterprise overage, ZDR, DPA/BAA and incident terms.
Official evidence boundary
Eighteen official Cartesia pricing, API, model, cloning, security and legal sources were checked on 28 August 2026. No hands-on latency, naturalness, cloning, deletion or uptime result is claimed. Cartesia Sonic remains `verified_essentials` and `noindex, follow`; the model-label conflict and two unresolved lifecycle/service fields remain explicit.