Source-led profile · Evidence checked 2026-08-26

Verified essentials

Voice-generation studio

Cartesia Sonic

Evaluate Cartesia Sonic pricing, streaming latency, concurrency, model versions, commercial rights, voice cloning and zero-data-retention boundaries.

Decision first. Use the compact answer below before opening the complete research record.

Decision summary

The answer in one scan.

Decision-critical facts remain separate from the deeper editorial analysis.

Best fitstreaming speech for conversational agents and real-time products
Free evaluationFree evaluation has not been established.
PricingThe official page lists Free at USD 0, Pro at USD 5, Startup at USD 49 and Scale at USD 299 per month, with credits and concurrency varying by tier.
Commercial useCommercial-use eligibility has not been established.
Main cautionResolve the open evidence fields before buying.

From evidence to action

Make the Cartesia Sonic decision with the right unit and route.

Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.

The number that changes the decision

Concurrency-cost estimator

Cartesia Sonic: Sessions, turns, duration and peak concurrency produce a load-test and pricing brief.

pricing · plans
Monthly page lists Free $0, Pro $5, Startup $49, Scale $299 and custom Enterprise, with unlimited workspace seats.
pricing · credits
Free/Pro/Startup/Scale include 20K/100K/1.25M/8M credits; page estimates about 27/133/1,667/10,667 Sonic minutes respectively.
pricing · regeneration
At one credit per ordinary TTS character, corrections consume the same unit again; model/feature exceptions such as PVC and infill require separate calculation.
limits · concurrency
Free/Pro/Startup/Scale allow 2/3/5/15 simultaneous TTS contexts; excess returns 429. Enterprise is custom.
limits · websocket
Parallel TTS WebSockets are capped at 10× TTS concurrency and idle TTS sockets close after five minutes; connection pooling/cleanup is required.
voice · clone · professional cost
Documentation lists one million credits on successful creation and 1.5 credits per TTS character; page gives inconsistent audio guidance (table says three hours, detailed section minimum 30 minutes/recommends two hours), so confirm current requirement.
Decision-changing number

Sessions, turns, duration and peak concurrency produce a load-test and pricing brief.

Keep in mind: Use only the retained Cartesia Sonic evidence; do not generalize this decision asset to another product.
Sources and verification date

Continue your research

Move from profile to a sharper decision.

These links are explicit editorial relationships, not keyword matches or sponsored placements.

Good fit if

Cartesia Sonic matches the job you need done

  • streaming speech for conversational agents and real-time products
  • The documented workflow and controls cover your required production steps.
  • You can validate the output with a representative project before committing.

Look elsewhere if

You need certainty this profile cannot provide

  • You need independently tested output quality rather than documented capabilities.
  • Your purchase depends on one of the 3 facts still requiring confirmation.
  • A narrower product would complete the same job with less workflow overhead.

Price, plan and risks

Confirm before you buy

Unknown, conflicted and stale facts stay visible before checkout.

Unknown

Self-serve evaluation

The retained evidence does not establish this field yet.

Unknown

Commercial-use eligibility

The retained evidence does not establish this field yet.

Unknown

Platforms

The retained evidence does not establish this field yet.

Commercial context

Compare the closest documented workflows.

Alternatives stay within the same vertical and use current internal profile routes.

Open the complete Cartesia Sonic buying analysisreal-time TTS · model pinning · concurrency economics · voice-clone lifecycle · ZDR boundary

Cartesia Sonic is designed for streamed speech inside conversational agents and interactive applications. Its buying value is not a polished demo voice. It is the combination of pinned model behavior, first-byte latency, concurrency, predictable billing and a defensible data/voice lifecycle inside your own product.

> Distinctive strength: Sonic is designed around streaming speech for interactive agents, where first-byte behavior, concurrency and pinned model operation matter. > > Where it stops being an advantage: A fast demo does not establish production latency, reliability, accepted-output cost or voice-data governance.

Cartesia also sells adjacent speech-to-text and voice-agent services. This profile isolates Sonic text-to-speech and the voice assets it depends on. Production latency, naturalness, reliability and clone behavior remain unscored until the product-specific pilot is run.

The Cartesia Sonic decision in 60 seconds

Buyer questionSource-led answer
Best fitDevelopers building real-time agents, support flows, games or interactive products.
Public entry priceFree $0; Pro $5; Startup $49; Scale $299 per month; Enterprise custom.
Commercial usePro and higher according to pricing and plan-gated Terms. Free is not the commercial tier.
TTS concurrency2 / 3 / 5 / 15 concurrent contexts by Free / Pro / Startup / Scale.
Voice cloningInstant from Pro; professional from Startup, with separate creation/use economics.
Main cautionPricing says Sonic-3.6 while public model docs currently detail Sonic 3.5. Confirm the actual model ID and credit rate.
Data cautionEnterprise Zero Data Retention covers TTS/STT inference, not cloning/PVC/voice creation.

What exactly is Cartesia Sonic?

Sonic is Cartesia's streaming TTS model family. Developers submit text, a model and a voice, then receive audio through one of three endpoint styles:

  • bytes: straightforward raw/file stream when timestamps are unnecessary;
  • SSE: streamed JSON events with timing information;
  • WebSocket: reusable bidirectional connection and context IDs for conversational flows.

The surrounding product—script approval, turn taking, interruption, transcript storage, moderation, mastering and rollback—belongs to your application. Cartesia is infrastructure, not an audiobook or video-production workspace.

Which Sonic model will production actually use?

Current public pages need reconciliation. The technical model page details Sonic 3.5, 42 languages and a dated stable snapshot. The pricing calculator labels included TTS minutes as Sonic-3.6. That does not prove either page is wrong, but it prevents an honest assumption that the documented 3.5 behavior and displayed 3.6 economics are identical.

Cartesia documents three version styles:

  • base ID routes to the newest stable snapshot;
  • dated snapshot remains fixed;
  • latest/preview track may be hot-swapped.

Pin the dated model and `Cartesia-Version` used in acceptance testing. Before purchase, obtain written confirmation of model ID, languages, controls, credits per character and retirement notice. A “latest” alias is convenient for exploration and inappropriate as the only production reproducibility control.

What languages, formats and timing data are available?

Sonic 3.5 documents 42 languages, including English, Romanian, Arabic, Hebrew and multiple Indic languages. Coverage does not establish equal pronunciation, emotion or latency in each one.

The API supports raw, WAV and MP3. Documented PCM formats include floating-point/signed 16-bit, μ-law and A-law with 8, 16, 22.05, 24, 44.1 and 48 kHz sample rates. SSE/WebSocket can provide word or segment timestamps subject to model/language support.

Test the exact production codec. A telephony agent may need 8 kHz μ-law; media output may use WAV at 44.1/48 kHz. Transcoding after generation adds latency and potential quality differences that do not appear in a playground comparison.

How much does Cartesia Sonic cost?

The public monthly pricing page currently shows:

PlanPriceCreditsDisplayed Sonic minutesTTS concurrency
Free$020,000~272
Pro$5100,000~1333
Startup$491,250,000~1,6675
Scale$2998,000,000~10,66715
EnterpriseCustomCustomCustomCustom

These figures imply roughly 750 credits per displayed minute, but do not budget real operations. Calculate from generated characters and preserve usage records. Include:

  • greeting, prompts and dynamic replies;
  • retries after disconnects/errors;
  • pronunciations and editorial corrections;
  • test/staging traffic;
  • model migrations and A/B tests;
  • infill and cloning charges;
  • idle/failed request behavior;
  • seasonal peak concurrency.

If a baseline consumes 1.25M credits, 25% regeneration needs about 1.5625M and 50% needs 1.875M before special-feature charges. That already exceeds Startup's included credits even though the initial workload fits it.

Infill has a separate published rule: one credit per inserted character plus 300 credits each operation. Frequent small edits can therefore have a very different cost curve.

How many simultaneous conversations can it support?

Cartesia counts TTS concurrency by active context. HTTP requests are individual contexts. On WebSocket, requests sharing a `context_id` process sequentially and do not add concurrency. Exceeding the plan limit returns 429.

The vendor suggests conversational agents may support roughly four times the TTS concurrency because speech generation is intermittent. Treat that as capacity-planning guidance, not a guarantee. For batch generation, parallel jobs map much more directly to the limit.

WebSocket connections are capped at ten times TTS concurrency. Idle TTS sockets close after five minutes. A leaked connection pool can hit socket limits even when generation is low.

Load-test with real turn lengths and silence distributions. Record p50/p95/p99:

  • time to first audio byte;
  • time to final audio;
  • active contexts/connections;
  • 429 and reconnect rate;
  • truncated/duplicated audio;
  • credit usage per completed turn.

Do not infer production capacity from a single vendor latency figure.

What voice controls should be tested?

Cartesia documents expressive/pacing controls, contextual pronunciation and pronunciation dictionaries, but controls are model/version dependent. Build a retained fixture containing:

  • names and brand terms;
  • email, telephone, order and confirmation codes;
  • currencies, dates and abbreviations;
  • heteronyms and punctuation;
  • language switching;
  • short acknowledgements and long explanations;
  • interruption at multiple audio positions.

Ask native reviewers to blind-score intelligibility, appropriateness, consistency and correction effort. Preserve the exact model, voice, controls and output format for every sample.

How does instant voice cloning work?

Pro and higher include instant cloning. The API reference suggests about five seconds; guidance says ten seconds is the high-similarity sweet spot and warns that background noise and pauses can be reproduced. Longer is not automatically better.

The source clip therefore becomes an output-control surface. Use a quiet, trimmed, target-language recording that reflects desired pace and energy. Store:

  • speaker identity and signed authorization;
  • source provenance and permitted uses;
  • recording date/languages;
  • hash of the submitted clip;
  • returned voice ID and access visibility;
  • approved applications/territories/term;
  • revocation and deletion actions.

Terms prohibit using another person's voice without express permission, including a deceased person or political candidate. The public material inspected does not establish a complete identity/consent-verification workflow inside the product, so your own gate remains necessary.

Is Professional Voice Cloning economically different?

Yes. Startup and above expose Professional Voice Cloning (PVC). Current docs list:

  • 1,000,000 credits after successful training;
  • 1.5 credits per generated character;
  • fine-tuning tied to a base model;
  • retraining and another 1M-credit charge for a newer base model.

Audio guidance needs clarification: one table says three hours, while the detailed workflow says minimum 30 minutes and recommends two hours. Training is described as about three hours and produces four voices from source clips.

PVC is a model-lifecycle commitment, not merely a better voice button. Budget recording/editing, consent, failed training, evaluation, 50% higher inference credits and future retraining. Test whether an instant clone is sufficient before accepting that burden.

Who owns output and can it be used commercially?

Cartesia says it does not claim ownership of customer inputs or outputs. Outputs may be non-unique, and the user remains responsible for legality, accuracy and third-party rights.

Commercial use is expressly tier-gated: pricing includes it from Pro, while Terms permit commercial use only when the subscription tier allows it. This does not clear trademarks, scripts, performances, publicity rights or the cloned speaker's consent. Attach the intended use to the plan and obtain legal review for consequential deployments.

Does Cartesia use customer data to train models?

The public Terms allow inputs, outputs and interactions to train/improve models unless otherwise agreed. A form can opt selected categories out of future training use, but prior uses and improvements are unaffected.

Enterprise Zero Data Retention changes inference handling:

Data under ZDRRetained?
TTS text inputNo
TTS audio outputNo
STT audio inputNo
STT transcript outputNo
Request IDs, usage, account/service metadataYes

ZDR explicitly excludes voice cloning, PVC and voice creation because those workflows need retained source material or derived artifacts. It can also remove Playground history and limit support recovery.

Do not describe the account as “zero retention” without that qualification. Separate ordinary TTS/STT requests from voice-asset workflows in the data inventory.

Can a cloned voice be fully deleted?

The API provides `DELETE /voices/{id}`. That establishes deletion of a voice resource, not the timing/scope for uploaded source audio, derived embeddings/fine-tuned models, reusable PVC datasets, logs and backups. ZDR's explicit cloning exclusion makes this a procurement question.

Require deletion documentation and test it with identifiers. If a speaker revokes consent, the runbook should locate every clone/PVC/dataset and dependent application—not only hide a voice in the Playground.

What security controls matter?

Cartesia's ZDR documentation states GDPR, SOC 2 Type II, PCI-DSS service-provider and HIPAA compliance. Review the current assurance artifacts and scope rather than repeating badges.

For browser/mobile apps, Cartesia documents short-lived access tokens and warns against embedding API keys. Keep key/token creation server-side, scope permissions, set expiry, rotate/revoke and prevent tokens from appearing in logs.

The 2026 changelog says requests route by origin to US/EU/APAC. Routing is not the same as contractual data residency. Confirm regions, subprocessors, transfer mechanism, incident notice, encryption, access logs and Enterprise SLA.

Who should shortlist Cartesia Sonic?

Shortlist Sonic when your product needs streamed speech, API-level timing/control, moderate-to-high concurrency and engineers who can own versioning, sockets, retries, consent and data governance.

Look elsewhere when:

  • you need an all-in-one audiobook/video editor;
  • commercial use must remain free;
  • long-form project collaboration is the core workflow;
  • no one can run load and failover tests;
  • cloned-voice deletion/consent cannot be governed;
  • Enterprise data terms cannot be obtained for sensitive speech.

Continue with the AI Voice directory, the commercial AI voice guide and the canonical software comparison workspace.

A product-specific Cartesia pilot

  1. Pin Cartesia API and dated Sonic model versions; record the billed model label.
  2. Build multilingual fixtures with names, codes, dates, emotions and interruptions.
  3. Generate bytes, SSE and WebSocket outputs in production codecs.
  4. Measure p50/p95/p99 latency, 429s, reconnects and completed-turn credits.
  5. Blind-review audio and calculate correction/regeneration cost.
  6. Saturate concurrency and socket limits; prove pool cleanup and backoff.
  7. Create one consented instant clone; test visibility, authorization and deletion.
  8. Compare PVC only if instant cloning fails; include retraining lifecycle cost.
  9. Verify training opt-out or Enterprise ZDR on eligible inference; document exclusions.
  10. Simulate snapshot retirement, timeout, malformed input and regional failure.

Pass only if the application—not just the model—meets latency, intelligibility, legal, cost and recovery thresholds.

Main unknowns before purchase

  • whether production pricing currently maps to Sonic-3.5 or Sonic-3.6 and its exact credit rate;
  • deletion timing/scope for clone source data, PVC models/datasets and backups;
  • contractual availability/support response SLA;
  • fixed regional processing/residency and subprocessor scope;
  • Enterprise overage, ZDR, DPA/BAA and incident terms.

Official evidence boundary

Eighteen official Cartesia pricing, API, model, cloning, security and legal sources were checked on 28 August 2026. No hands-on latency, naturalness, cloning, deletion or uptime result is claimed. Cartesia Sonic remains `verified_essentials` and `noindex, follow`; the model-label conflict and two unresolved lifecycle/service fields remain explicit.

Full evidence record7 fields · official links · dates · states

Documented capabilities

Vendor claimChecked 2026-08-26

streaming byte, SSE and WebSocket endpoints; Sonic text-to-speech models; instant and professional voice-cloning options by plan; pronunciation, accent and model-scoped delivery controls

Official sources (3)

Documented limitations

Vendor claimChecked 2026-08-26

credits must be translated into the buyer's real minutes and model; some controls are model-version dependent; the vendor's latency and quality claims were not independently reproduced

Official sources (3)

Pricing context

Vendor claimChecked 2026-08-26

The official page lists Free at USD 0, Pro at USD 5, Startup at USD 49 and Scale at USD 299 per month, with credits and concurrency varying by tier.

Official sources (1)

Self-serve evaluation

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.

Commercial-use eligibility

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.

Platforms

UnknownChecked 2026-08-26

The retained official evidence does not answer this yet.