Source-led profile · Evidence state visible

Listed · research pending

Voice-generation studio

Gradium

Evaluate Gradium's streaming TTS/STT API, shared credits, concurrency, five-language scope, voice cloning, deployment and a production voice-agent test.

Decision first. Use the compact answer below before opening the complete research record.

Decision summary

The answer in one scan.

Decision-critical facts remain separate from the deeper editorial analysis.

Best fitDeveloper speech infrastructure
Free evaluationFree evaluation has not been established.
PricingPricing has not been established.
Commercial useCommercial-use eligibility has not been established.
Main cautionTest output quality for your own language and workflow.

From evidence to action

Make the Gradium decision with the right unit and route.

Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.

The number that changes the decision

Real-time workload calculator

Gradium: Simultaneous sessions and audio duration generate a concurrency/cost test.

pricing · free
$0/month with 45,000 credits, approximately one TTS hour, API/Studio access, TTS concurrency 2, STT 3, S2S 2, five instant clones and no commercial use.
pricing · xs
$13/month with 225,000 credits, TTS concurrency 5, STT 20, S2S 5, commercial use and up to 1,000 instant clones.
pricing · s
$43/month with 900,000 credits and the same listed 5/20/5 concurrency; additional 100k credits cost $5.00.
pricing · m
$340/month with 9M credits, 10/40/10 concurrency, five Pro clones and additional 100k credits at $4.00.
pricing · l
$1,615/month with 45M credits, 15/60/15 concurrency, 20 Pro clones and additional 100k credits at $3.80.
pricing · credit conversion
TTS costs 1 credit/character; STT 3 credits/second; STT translation 4/second; speech-to-speech translation 30/second (launch display also shows a temporary 50% discount).
Decision-changing number

Simultaneous sessions and audio duration generate a concurrency/cost test.

Keep in mind: Use only the retained Gradium evidence; do not generalize this decision asset to another product.
Sources and verification date

Good fit if

Gradium matches the job you need done

  • Developer speech infrastructure
  • The documented workflow and controls cover your required production steps.
  • You can validate the output with a representative project before committing.

Look elsewhere if

You need certainty this profile cannot provide

  • You need independently tested output quality rather than documented capabilities.
  • Output quality still needs hands-on validation.
  • A narrower product would complete the same job with less workflow overhead.

Commercial context

Compare the closest documented workflows.

Alternatives stay within the same vertical and use current internal profile routes.

Open the complete Gradium buying analysisstreaming TTS/STT stack · shared metering · public concurrency · five languages · instant/pro cloning · retention/SLA evidence gates

# Gradium review: price the complete real-time voice loop, not TTS alone

Gradium is a developer voice-AI platform combining streaming text-to-speech, speech-to-text, translation and voice cloning. Its strongest advantage is infrastructure cohesion: REST and bidirectional WebSocket APIs, pronunciation dictionaries, custom-voice lifecycle and metering sit behind one API key and one shared credit pool.

> Distinctive strength: Gradium gives real-time voice-agent teams one metered API for TTS, STT, translation, pronunciation and custom voices, with plan-level concurrency published. > > Where it stops being an advantage: the public language scope is five languages, and claimed latency, zero retention, enterprise SLA and deployment assurances still require testing or contract evidence.

The shared pool simplifies procurement but complicates comparison. One million TTS characters and one hour of speech translation consume very different credit amounts. A plan that looks inexpensive in a TTS-only spreadsheet can become unsuitable when an agent listens, translates and speaks.

Gradium's latency, accuracy, voice and deployment claims have not been reproduced by BenPicks. They remain inputs to the production test below—not borrowed benchmark results.

The Gradium decision in 60 seconds

QuestionSource-led answer
Strongest reason to consider itUnified real-time TTS, STT, translation, cloning, pronunciation and metering API.
Current languagesEnglish, French, Spanish, Portuguese and German.
Stock catalogue237 documented voices across those five languages.
Free route45k credits; API/Studio; approximately one TTS hour; no commercial use.
Paid entryXS $13/month for 225k credits and commercial use.
Operational limits300-second sessions; Free adds 1,500 characters/session; concurrency is plan-specific.
Main unknownsQuantitative public SLA, complete retention/training scope, consent enforcement and scoped security assurance.

What is Gradium particularly good at?

Gradium is designed around a live conversational loop. The API inventory covers WebSocket and POST TTS/STT, voices, pronunciation dictionaries and metering. A team can stream input and output rather than stitching together unrelated batch services.

This matters when an agent must listen, decide and speak with interruption handling. Gradium also publishes Gradbot, a first-party open-source framework connecting STT, an OpenAI-compatible LLM and TTS with turn-taking, barge-in and tool calls. That does not guarantee production quality, but it gives engineers a more concrete starting point than a marketing demo.

The limit is scope. Five languages may be enough for a European deployment and inadequate for a global catalogue. Content creators seeking dozens of languages or deep long-form editing should not choose an agent API merely because its low-latency demo sounds good.

How does Gradium pricing work?

The same credits fund several workloads:

  • TTS: 1 credit per character;
  • STT: 3 credits per second;
  • speech-to-text translation: 4 credits per second;
  • speech-to-speech translation: 30 credits per second.

The page currently displays a limited 50% S2S launch offer. Treat the 30-credit contract and any discounted display separately; do not build permanent economics on a temporary promotion.

PlanMonthlyCreditsNominal TTS hoursTTS/STT/S2S concurrency
Free$045k~12 / 3 / 2
XS$13225k~55 / 20 / 5
S$43900k~205 / 20 / 5
M$3409M~20010 / 40 / 10
L$1,61545M~1,00015 / 60 / 15

Enterprise is custom. Free is marked noncommercial; paid plans permit commercial use. Additional 100k credits are $6.90 XS, $5 S, $4 M and $3.80 L.

What is the effective TTS cost?

If every bundled credit is consumed by TTS:

PlanApprox. bundled cost per 1M characters
XS$57.78
S$47.78
M$37.78
L$35.89

These figures divide plan price by included credits. They are not marginal overage rates and they overstate value if credits expire unused. They also stop describing reality when STT or translation shares the pool.

Model a real agent session: caller seconds × 3 for STT, translated seconds × applicable rate, and response characters × 1 for TTS. Add retries, silence handling, failed calls and test traffic. Compare monthly credit consumption with the plan's concurrency, not just its headline hours.

Is the free plan useful?

Free includes 45,000 credits, Studio/API access, two concurrent TTS streams, three STT streams, two S2S streams and five instant clones. Gradium approximates 45,000 characters as one TTS hour and says no credit card is required.

Free also has a 1,500-character session limit, while all sessions are capped at 300 seconds. That is enough for a prototype, not evidence that a production call centre fits. Free does not include commercial use.

Before upgrading, capture credit changes through the metering endpoint for every TTS, STT and translation action. Confirm credit rollover, billing date, upgrade/downgrade behavior and failed-request charging inside the account; the complete mechanics were not established in this evidence pass.

What languages and voices are available?

The FAQ lists English, French, Spanish, Portuguese and German. The documented catalogue currently contains 237 voices: 100 English, 43 French, 50 German, 20 Spanish and 24 Portuguese.

Counts do not prove accent suitability. Build a language-specific corpus containing names, addresses, dates, amounts, abbreviations, alphanumeric identifiers and code-switches. Use Gradium's pronunciation dictionaries for recurring domain vocabulary, then test whether entries persist across voices and API sessions.

Native reviewers should score both transcript accuracy and synthesized pronunciation. A voice agent can sound natural and still fail at the customer numbers that matter.

How should latency and concurrency be tested?

Gradium markets first audio in roughly 200ms and flat latency under load. These are vendor claims. The buyer should measure:

  • time to first accepted audio, not first byte only;
  • end-to-end listen-think-speak delay;
  • P50/P95/P99 at 1, 5 and plan-limit concurrency;
  • interruption and barge-in response;
  • reconnect, timeout and retry behavior;
  • WER on domain data in every target language;
  • credit charges for failed/retried sessions.

Run long enough to observe queueing. Compare the same telephony input, network region and LLM path across providers. Do not mix Gradium's benchmark setup with a competitor's marketing number.

How does Gradium voice cloning work?

Instant cloning uses about ten seconds of clear audio. Free includes five; XS through M list up to 1,000, while L/Enterprise list unlimited. The custom-voice API accepts an audio file plus metadata and returns a voice identifier for TTS.

Pro cloning requires at least 30 minutes of clean audio, with two hours recommended for emotional range and stability. M includes five Pro clones and L twenty.

Gradium explicitly requires voice-owner consent. Public docs reviewed here did not establish how that consent is technically verified. Before upload, require signed authorization, speaker identity, permitted uses/languages/territories, operator access, revocation and incident procedures.

The API documents permanent custom-voice deletion, which is useful. Still obtain the scope: source upload, derived embeddings/weights, logs, backups and downstream copies. An API `DELETE` response alone does not prove every layer is erased.

What is known about retention and deployment?

Gradium's homepage claims zero data retention and presents cloud, dedicated, self-hosted/on-prem and marketplace options. It also describes Phonon, an approximately 100M-parameter CPU model that runs offline/on-device.

These options could be decisive for data-sovereign or offline products. They require separate commercial and technical confirmation: which models/features are available, hardware requirements, update path, telemetry, licensing, support, residency and security responsibilities.

“Zero retention” must be scoped across text, audio, transcripts, request logs, abuse/security logs, custom-voice samples and models. This review did not establish a public DPA, scoped SOC/ISO report, complete training-use opt-out or numerical enterprise SLA.

Is Gradium appropriate for commercial use?

The pricing table marks paid plans commercial and Free noncommercial. That establishes a product route, not ownership of every input. The buyer remains responsible for scripts, caller consent, recorded conversations and cloned speakers.

For regulated or sensitive calls, obtain recording disclosures, retention policies, subprocessors, regional transfer terms, DPA/security exhibits and incident commitments. If self-hosting/on-device is proposed, map which data still reaches Gradium for licensing, updates, diagnostics or support.

Who should shortlist Gradium?

Shortlist it when a team is building real-time agents in the five supported languages and wants TTS, STT, translation, cloning and pronunciation controls under one developer contract. Published concurrency and metering support a disciplined proof.

Look elsewhere when broad language coverage, long-form editorial control or a fully public compliance package is required. A shared pool is not automatically economical when translation dominates usage.

Compare the AI voice software category, use the AI voice API buying guide, and model character/accepted-output economics through how to calculate AI voice generation cost.

A fail-closed Gradium production test

  1. Capture plan, credits, overage, renewal, rollover and commercial terms.
  2. Build one duplex agent through the documented WebSocket API.
  3. Test all five languages with domain-specific entities and native reviewers.
  4. Create pronunciation dictionaries and verify persistence/reproducibility.
  5. Measure P50/P95/P99 latency and WER at 1, 5 and plan-limit concurrency.
  6. Exercise interruptions, silence, reconnects, timeouts and retries.
  7. Reconcile every request with the credit endpoint across TTS/STT/translation.
  8. Compare instant and Pro clones using an authorized speaker and blind scoring.
  9. Delete a test clone and obtain deletion-scope evidence.
  10. Evaluate cloud versus dedicated/self-hosted/on-device for the actual constraint.
  11. Obtain retention, training, DPA, security, residency and SLA documents.
  12. Price the full listen-translate-speak loop, including failures and idle capacity.

Pass only when latency remains within budget under realistic concurrency, recognition/pronunciation survive native review, metering reconciles, cloning is governed and the deployment contract satisfies the data boundary.

Official sources

Full evidence record0 fields · official links · dates · states