# IBM Watson Text to Speech review: choose it for control, not a famous logo
IBM Watson Text to Speech is an API service available through IBM Cloud and installed IBM Cloud Pak for Data/Software Hub deployments. Its strongest advantage is governed deployment and data handling: a buyer can combine IAM, encrypted transport/storage, request-level no-logging, customer-linked deletion and large pronunciation dictionaries—or move the service behind its firewall.
> Distinctive strength: IBM combines cloud or installed deployment with explicit request-logging controls and pronunciation dictionaries of up to 20,000 entries. > > Where it stops being an advantage: the service has no production creator GUI, cannot switch languages inside one request, and Premium/Deploy Anywhere pricing, custom voice and higher assurance depend on negotiated contracts.
This is an enterprise/API choice. It should not win because IBM is familiar, and it should not lose because its public voice catalogue is smaller than a creator platform. The relevant questions are whether vocabulary, data boundary, output and SLA fit the application.
BenPicks has not used IBM Watson TTS, measured its latency or voices, deployed Cloud Pak, inspected an IBM contract or tested deletion. The following separates documented controls from the evidence still required.
IBM makes sense when governance matters as much as synthesis
| Question | Source-led answer |
|---|
| Strongest reason to choose it | IAM/no-log/deletion controls plus IBM Cloud or behind-firewall deployment. |
| Free route | 10,000 characters/month. |
| Standard public rate | As low as $0.02 per 1,000 characters. |
| Pronunciation | SSML plus language-specific dictionaries up to 20,000 entries. |
| Interfaces | HTTP GET/POST, WebSocket and Watson SDKs. |
| Language limitation | One voice/language per synthesis request; no multilingual synthesis. |
| Enterprise route | Premium 99.9% claim/custom voice; Deploy Anywhere behind firewall/on any cloud. |
Pronunciation governance is IBM's most defensible edge
The product is strongest when governance and domain vocabulary matter together. IBM documents account/request-level logging opt-out; when enabled, no user data from HTTP/WebSocket requests is written to disk. A customer ID can be attached through `X-Watson-Metadata`, then associated data can be removed with `DELETE /v1/user_data`.
At the same time, a custom pronunciation model can hold 20,000 terms. A hospital, insurer or financial platform can maintain product names, codes and specialist vocabulary while using IAM and plan-specific data controls. Installed Cloud Pak deployment can place the service behind a firewall when cloud processing is not acceptable.
That advantage is irrelevant for a creator who wants a visual timeline, instant self-serve voice cloning or dozens of languages. There is no graphical production interface; normal use is API/SDK.
Lite is easy to test; enterprise deployment is not easy to price
IBM's product page presents four routes:
| Route | Public price/allowance | Key boundary |
|---|
| Lite | free, 10,000 characters/month | evaluation |
| Standard | as low as $0.02/1,000 characters | cloud usage and high-value features |
| Premium | contact sales | custom-branded voice, stated 99.9% uptime |
| Deploy Anywhere | contact sales | behind firewall/any cloud, unlimited chars, stated 35 voices/16 language-dialect slots |
At the public starting Standard rate:
| Characters | Cost |
|---|
| 100,000 | $2 |
| 1,000,000 | $20 |
| 5,000,000 | $100 |
“As low as” is not a universal quote. During the 30 August check, IBM's United States Cloud catalogue displayed $0.0212 per thousand characters, while the product page retained the rounded $0.02 starting claim. The table above therefore illustrates the marketing floor, not a guaranteed invoice. Confirm tier, location, currency, tax, support, data separation and overage in IBM Cloud or the negotiated order. Whitespace counts toward request-size limits but IBM says it is excluded from billing.
Which languages and voices are supported?
Current documentation covers Dutch; English variants; Canadian/France French; German; Italian; Japanese; Korean; Brazilian Portuguese; and several Spanish variants. Voice availability and type differ between IBM Cloud and installed versions.
IBM distinguishes Natural, Expressive neural and Enhanced neural voices. Query the live voices endpoint for the current list and each voice's customization features rather than relying on a static total.
The service does not support multilingual synthesis in one request. The selected voice determines the input language. Split mixed-language scripts and test boundaries explicitly; `xml:lang` is not a workaround for changing voices/languages within a request.
SSML depth varies by voice
IBM supports most SSML 1.1 elements for pronunciation, volume, pitch, speed and related behavior, but support varies by voice and interface. Newer voices may omit specific elements. IBM also exposes global speaking-rate and pitch-percentage parameters.
Create a compatibility table for every selected voice and required element. A valid SSML document can still use an element that the chosen voice ignores. Store the exact voice identifier, model version and markup with the accepted output.
Input size is constrained: GET permits 8KB total request; POST/WebSocket allow 5KB of text including SSML (with POST headers/URL separately limited). Split long-form content and measure whether joins, pauses and loudness remain consistent.
Why do IBM custom pronunciation models matter?
Customization creates a language-specific dictionary of word/translation pairs. Entries can use sounds-like spelling or phonetic notation (IPA or IBM SPR). The model can apply across voices in its language.
Limits are explicit: 20,000 entries, 49 characters per word and 499 per translation. Lite cannot use customization; Standard/Premium can.
Test a representative 1,000-term set containing medicine/product names, acronyms, alphanumeric IDs, addresses, amounts and ambiguous words. Verify creation/update/deletion, latency impact and behavior when an entry conflicts with SSML. The dictionary's scale is useful only if operators can govern it.
The API returns Ogg/Opus by default
Default output is Ogg with Opus. IBM supports additional formats through its audio-format reference. The complete codec/sample-rate matrix should be captured for the intended destination.
Test telephony and media paths independently: inspect headers and file properties, stream through the real client, check transcoding and re-import into downstream tools. Do not assume the default fits PSTN, browser playback or archival masters equally.
What are IBM's data controls?
IBM Cloud authenticates through IAM API keys/bearer tokens. Documentation states TLS 1.2 in transit and AES-256/SHA-256 at rest. Credentials isolate customer data, while Standard and Premium offer different separation/encryption levels.
By default, Watson services may log request/response data to improve services. IBM documents opt-out at account or individual-request level; after opt-out, no user data from HTTP/WebSocket requests is written to disk. Test and document the header/account setting in every environment.
The `X-Watson-Metadata`/`DELETE /v1/user_data` path supports customer-linked deletion. Verify its scope across logged requests, custom dictionaries, custom voices, backups and installed deployments. No-data-written-to-disk applies to opted-out requests, not automatically to every other resource.
What does Premium or installed deployment add?
Premium states a 99.9% high-availability/service-level uptime guarantee, custom-branded voice and stronger capacity/data protection. HIPAA readiness is documented for Premium in `us-east` and `us-south`.
Deploy Anywhere uses Cloud Pak for Data behind the buyer's firewall or on its cloud. IBM lists unlimited characters, 35 neural voices and 16 languages/dialects. Installed deployments use access-token authentication and support FIPS-enabled clusters in documented versions.
Obtain the real order form: hardware sizing, license metric, HA/disaster recovery, update cadence, telemetry, support, model availability, residency and incident commitments. “On-prem” transfers operational responsibility; it does not eliminate it.
Can IBM train a custom voice?
Premium customers can request a custom-branded neural voice, and IBM says training can begin with as little as one hour of data. This is a sales-assisted programme, not instant self-serve cloning.
Public material reviewed here did not establish the full speaker identity/consent, data preparation, retention, revocation, output-rights and deletion process. Require those terms before recording anyone. Also review the controlling IBM agreement for commercial output rights; a priced API alone is not a legal grant.
IBM belongs on infrastructure shortlists, not creator shortlists
Shortlist it when an API team needs IAM, explicit no-log requests, customer-linked deletion, deep pronunciation customization or installed deployment. It is particularly relevant where domain vocabulary and data governance are both hard requirements.
Look elsewhere when a creator needs a visual studio, multilingual switching within a single request, self-serve cloning or broad language coverage. A cloud-native commodity API may also be simpler when on-prem and IBM enterprise controls add no value.
Explore AI voice software, compare API requirements in the AI voice API buying guide, and normalize character costs with how to calculate AI voice generation cost.
The IBM proof should include logging, deletion and vocabulary
- Capture the exact plan, region, rate, SLA, support and data-separation terms.
- Query the current voice endpoint and select target-language voices/features.
- Run identical HTTP and WebSocket corpora and measure P50/P95 latency/errors.
- Test every required SSML element on every chosen voice.
- Build a governed 1,000-term custom pronunciation dictionary.
- Reconcile characters, whitespace and bills at 100k/1M scale.
- Inspect Ogg/Opus and target output formats through real downstream paths.
- Enable account/request logging opt-out and preserve configuration evidence.
- Attach a synthetic customer ID, delete it and verify the documented scope.
- Compare IBM Cloud with installed Cloud Pak sizing/operations if relevant.
- Obtain custom-voice consent, retention and output-rights terms if needed.
- Validate DPA, HIPAA/region, encryption, incident and 99.9% remedy scope.
Pass only when pronunciation, latency and audio survive the real application, billing reconciles, no-log/deletion controls are demonstrated and the selected deployment contract meets the buyer's security boundary.
Start with Lite—or take IBM a security-and-throughput brief
If cloud API access is sufficient, use IBM's free Lite route ↗ to prove the voice, SSML and dictionary workflow before discussing a larger contract. If behind-firewall deployment or a custom voice is the reason IBM made the shortlist, skip the generic demo: take the exact regions, throughput, voice, vocabulary, deletion and recovery requirements to sales. IBM earns the next step only when those controls are requirements, not when its logo is simply familiar.
Official sources