Standard voices are listed at USD 4 per million characters, Neural at USD 16, Generative at USD 30 and Long-Form at USD 100 outside applicable free allowances.
Commercial use
Commercial-use eligibility has not been established.
Main caution
Resolve the open evidence fields before buying.
From evidence to action
Make the Amazon Polly decision with the right unit and route.
Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.
The number that changes the decision
One million characters costs $4, $16, $30 or $100 before regeneration
Which engine rate still fits after corrections and regeneration?
Standard per million characters
4
Neural per million characters
16
Generative per million characters
30
Long-Form per million characters
100
Calculation boundary
Normalize the selected engine, processed characters and regeneration before comparing the synthesis line item.
Decision-changing number
The engine choice changes the published synthesis charge by 25× between Standard and Long-Form at equal processed-character volume.
Keep in mind: Equal character volume does not make the engines equivalent in voice availability, quotas, region support or output behavior.
Open the complete Amazon Polly buying analysisDecision verdict · total-cost model · quality evidence · SLA · data governance · architecture · procurement and evaluation protocol
Amazon Polly belongs on a shortlist when speech synthesis is one component inside an AWS application and the team can own everything around it. It is a poor substitute for a creator studio: Polly generates audio, but it does not supply the editorial review, timeline, mastering, publishing or client-approval workflow.
> Distinctive strength: Polly offers usage-priced speech infrastructure inside AWS with unusually explicit model prices, quotas, security controls and operational documentation. > > Where it stops being an advantage: It supplies synthesis, not the editorial, mastering, approval or publishing workflow a creator studio provides.
The useful buying question is not “Does Polly have text to speech?” It is: does usage-priced AWS speech infrastructure remove enough engineering work for your volume, region and output requirements? BenPicks has not rated Polly's naturalness, latency, reliability or usability. This review converts current AWS documentation into a decision model you can reproduce.
The decision in 60 seconds
Your requirement
What it means
Speech inside an existing AWS application
Strong structural fit: API/SDK access, IAM, CloudTrail, CloudWatch, S3 tasks and PrivateLink can join an existing AWS operating model.
Transparent usage pricing
Strong evidence, but count regeneration and adjacent AWS services—not only first-pass characters.
A no-code production studio
Weak fit. Your team still owns script approval, mastering, storage and publishing.
Self-service voice cloning
Not documented. Amazon Brand Voice is a separate sales-assisted engagement.
Translation or automatic dubbing
Not provided by Polly; supply text in the target language and build the localization workflow.
A voice that must remain identical over a long series
Test carefully: AWS warns generative-model updates can cause slight variation over time.
A contractual availability commitment
Polly is covered by AWS's regional 99.9% monthly-uptime SLA, but credits—not damages—are the documented remedy.
Confidential or regulated scripts
Potential fit only after the team resolves region, opt-out, S3 retention, IAM, encryption and compliance obligations in its own architecture.
What you are actually buying
Polly accepts plain text or SSML through the console, API, CLI or SDKs and returns audio. Real-time synthesis suits shorter requests. Asynchronous synthesis accepts larger documents, writes output to a required S3 bucket and can notify SNS when complete.
That leaves the buyer responsible for text approval, engine/voice/region selection, pronunciation dictionaries, retries, storage, retention, listening checks, corrections, loudness, mixing and publishing. This is the central product boundary—not a footnote.
Real cost scenarios
AWS lists prices per million processed characters: $4 Standard, $16 Neural, $30 Generative and $100 Long-Form, outside applicable free allowances. AWS also permits generated speech to be cached and replayed without another Polly charge.
Using AWS's published approximation that one million characters creates about 23 hours and 8 minutes of speech:
Monthly input
Standard
Neural
Generative
Long-Form
100,000 characters (~2h 19m)
$0.40
$1.60
$3.00
$10.00
1 million characters (~23h 8m)
$4.00
$16.00
$30.00
$100.00
5 million characters (~115h 40m)
$20.00
$80.00
$150.00
$500.00
These are transparent arithmetic scenarios, not invoice guarantees. Taxes, regional pricing, S3, transfer, monitoring, support and adjacent services can change the total.
For a one-million-character Neural workload, 25% regeneration raises processed volume to 1.25 million and synthesis cost from $16 to $20. At 50% regeneration it becomes $24. Script approval and pronunciation testing can therefore matter more than the headline unit price.
AWS currently documents first-year monthly allowances of 5 million Standard, 1 million Neural, 500,000 Long-Form and 100,000 Generative characters. It also describes a newer account-credit model for customers joining from July 15, 2025. Check the billing console for your account rather than assuming the programs stack.
The invoice model the headline price omits
Use this procurement formula rather than multiplying only first-pass script length:
total monthly cost = synthesized characters × regeneration multiplier × engine rate + S3 + transfer + monitoring + support + engineering/QA
The last term can dominate at low volume. A creator tool may charge more per generated minute yet cost less overall if it replaces review, collaboration and mastering work. Conversely, Polly can be compelling at high repeat volume because AWS permits cached audio to be replayed without another synthesis charge.
GovCloud is a separate buying case: AWS publishes $4.80 per million Standard characters and $19.20 per million Neural characters there, rather than the commercial-region rates above. Generative and Long-Form availability must be checked by region before budgeting.
Four engines, four different decisions
Standard has the lowest published price and the highest documented real-time quota.
Neural costs four times Standard and has lower default throughput.
Long-Form costs $100 per million characters and AWS documents availability only in US East (N. Virginia).
Generative costs $30 per million characters and is available in a defined region set. AWS warns that model updates can alter voice output slightly. Speech Marks are unavailable, and AWS documents rare safety behavior that may cut a word or fail to prevent all hallucinated continuation.
For episodic or regulated content, record engine, voice ID and region and preserve a reference sample. The same voice ID does not prove future acoustic identity.
What AWS's own quality evidence does—and does not—show
AWS's 2026 Polly AI Service Card supplies unusually useful vendor evidence. It says each released voice is tested with 1,000–2,000 prompts and that CSAT uses at least 50 native listeners. AWS reports these engine-level results:
Engine
Voice QA critical-error rate
CSAT (1–7)
Neural
0.77%
5.75
Long-Form
1.53%
6.09
Generative
1.79%
6.15
The documented critical-error categories include inconsistent speaker identity, cutoff, audio glitch, hallucination, inappropriate persona, intonation, pausing, pronunciation and text normalization. These figures are stronger evidence than a marketing adjective, but they remain AWS-run aggregate tests—not a BenPicks hands-on result, not a voice-by-voice guarantee and not proof for your language or script.
This changes how Polly should be tested. A buyer should not ask only “Which sample sounds nicest?” Measure pronunciation failures, omissions, added words, pacing errors and listener preference separately. AWS itself warns that mixed-language text, specialist terminology and dynamic content can fail, and that Polly does not filter harmful or sensitive input.
Limits that may redesign your application
Real-time `SynthesizeSpeech` accepts up to 3,000 billed characters (6,000 total), up to five lexicons and at most ten minutes of audio. Asynchronous synthesis accepts 100,000 billed characters (200,000 total) and requires writable S3 storage.
Default real-time throughput is 80 TPS for Standard but 8 TPS for Neural, Long-Form and Generative. AWS recommends backoff and jitter. Test the intended concurrency in the intended region rather than extrapolating from one console request.
Common outputs include MP3, Ogg Vorbis and raw PCM; telephony can use Mu-law or A-law. Speech Marks return separate JSON timing metadata for sentences, words, visemes or SSML where supported. Generative voices currently do not support them.
Region and resilience implications
Polly exposes regional endpoints rather than one global speech endpoint. AWS currently lists 25 commercial/GovCloud regions for the service, but engines and voices are not uniform across them. Generative speech is documented in a smaller ten-region set; Long-Form is more restricted. PrivateLink is available in supported Polly regions, while FIPS endpoints are documented only in six regions.
The AWS ML Language Services SLA covers Polly separately per account and region at 99.9% monthly uptime. If uptime falls below 99.9% but stays at least 99.0%, the published credit is 10%; below 99.0% but at least 95.0%, 25%; below 95.0%, 100%. Claims require request logs and must be submitted within the stated window. This is useful contractual evidence, but it does not create automatic failover. A production design still needs retry policy, monitoring, cached/fallback audio where appropriate and an explicit response to regional failure.
Pronunciation and language controls
SSML can control pronunciation, volume, pitch, rate and pauses, but tag support differs by engine. Unsupported tags can fail instead of degrading gracefully. Polly also supports regional pronunciation lexicons: up to 100 per account, 40,000 characters each, with up to five applied per request.
Polly is not translation software. Voice and engine availability varies by language and region. A valid evaluation script must include names, abbreviations, dates, numbers, currency, domain terminology and every required language—not only fluent English marketing copy.
Rights, privacy and governance
AWS states that, as between the customer and AWS, Polly output belongs to the customer; the customer must hold rights to third-party input. That supports commercial use but does not clear copyright, trademark, publicity, consent or sector obligations.
AWS Service Terms also permit Polly content to be stored and used to improve the service and related AWS AI technologies, potentially outside the service region for that purpose. AWS documents an organization-level AI-services opt-out policy. Teams with confidential scripts should decide and record that setting before production use.
There is an important nuance rather than a simple contradiction. AWS's 2026 Service Card says Polly does not store synthesis inputs or outputs for the TTS function and does not use them to train Polly; pronunciation lexicons are the customer content stored by the service, and asynchronous outputs go to the customer's S3 bucket. The broader AWS terms and Organizations documentation address copies that may be retained for service improvement. AWS says applying the AI-services opt-out deletes historical improvement copies that are not required to provide the service. A serious buyer should preserve the effective Organizations policy, not rely on an assumption about the account default.
IAM, CloudTrail, CloudWatch and interface VPC endpoints provide real operational controls. PrivateLink can keep synthesis traffic off the public internet. These controls do not replace least privilege, retention rules, input classification or review of the AI-service terms. Asynchronous output inherits the security configuration of the destination S3 bucket.
AWS documents Polly within multiple compliance programs, including SOC, ISO, PCI, FedRAMP and HIPAA-eligible scope. That does not make the customer's application compliant automatically. For PHI, the account needs the appropriate AWS BAA and the complete architecture—including S3, logs, networking, access and deletion—must meet the customer's obligations.
CloudWatch exposes request-character counts, response latency and HTTP outcome metrics; CloudTrail records API activity. Those are the minimum observability inputs for cost attribution and an SLA claim. Neither service evaluates whether the audio is correct, so quality monitoring still requires human or application-level checks.
Brand Voice is procurement, not a clone button
AWS describes Brand Voice as a custom engagement with Polly scientists and linguists to create an exclusive voice for a brand persona. Public examples include KFC Canada and National Australia Bank. It is not documented as instant or self-service cloning, and public pricing, delivery time, training-data requirements and the engagement's consent evidence are not sufficiently specified.
Before signing, ask AWS to put these answers in writing:
Who may provide or authorize the source voice and what proof is retained?
What recordings, releases and usage territories are required?
Who can synthesize with the resulting voice, and how is access revoked?
Is the voice exclusive, portable or usable after contract termination?
What model/data artifacts are retained, where and for how long?
What happens after talent consent expires or is withdrawn?
What price, delivery schedule, support and change-control terms apply?
This is the profile's one material unresolved field. Leaving it unknown is more useful than implying that general AWS input-rights language proves a Brand Voice consent workflow.
What official sources do not prove
They do not establish that Polly sounds better than ElevenLabs, Azure or Google; that latency meets your target; that a voice stays acoustically identical after model updates; or that the complete AWS architecture costs less than a creator subscription. The SLA establishes a regional availability commitment, not the response time or expertise of the support plan you buy. Brand Voice exists, but its public procurement and consent details remain incomplete.
Product-specific test protocol
Use a versioned test corpus—not one polished demo. Include ordinary narration, proper nouns, acronyms, dates, currency, phone numbers, addresses, specialist terminology, mixed-language phrases, SSML, dynamic variables and every required locale.
Run it through every eligible engine in the production region.
Record voice ID, engine, region, format, sample rate, SSML, lexicons and request time.
Count every pronunciation error and manual correction.
Repeat the corrected script and calculate the regeneration multiplier.
Run at least 100 representative prompts per shortlisted voice and classify mispronunciation, omission, insertion, cutoff, glitch, pacing and persona mismatch separately.
Run sequential and bounded concurrent requests; record time-to-first-audio, completion latency, throttles, retries and HTTP outcomes.
Verify the exact output format and Speech Marks if required.
Run the asynchronous S3 path; verify bucket policy, encryption, notification, retention and cleanup.
Ask multiple native listeners to compare blinded outputs with at least one alternative using a predeclared rubric.
Recalculate cost using measured characters, regeneration, cache hit rate and adjacent AWS services.
Repeat a reference set after a model/service update to detect drift.
Leave unresolved procurement gates explicitly unknown; do not replace them with a star rating.
A pass/fail scorecard worth keeping
Record thresholds before listening: maximum critical-error rate, acceptable time-to-first-audio percentile, maximum regeneration multiplier, required region, required formats/Speech Marks, monthly total-cost ceiling, retention/opt-out approval, recovery behavior and Brand Voice consent evidence if applicable. Reject the engine or voice when a hard gate fails even if the demo sounds impressive.
Where alternatives may fit better
Priority
Compare first
Reason
Creator studio and self-service cloning
ElevenLabs
Different product shape; compare workflow burden, not only voice demos.
Existing Google Cloud estate
Google Cloud Text-to-Speech
Compare region, models, markup and surrounding cloud operations.
Microsoft enterprise workflow
Azure AI Speech
Compare geography, customisation and governance in the existing estate.
Conversational streaming
Deepgram Aura
Benchmark measured latency and concurrency rather than positioning.
Amazon Polly is strongest as AWS speech infrastructure. Its evidence is unusually concrete: transparent unit prices, explicit quotas, multiple output formats, asynchronous S3 delivery, IAM/CloudTrail/CloudWatch/PrivateLink controls and documented ownership of output. Its risks are equally concrete: the buyer owns the workflow, engine and region compatibility vary, generative speech has documented consistency caveats, and data-use terms require a governance decision.
That is enough to justify a serious shortlist—not enough to declare a quality winner without a workload-specific pilot.
API speech synthesis and streaming; standard, neural, long-form and generative voice families; SSML controls and Speech Marks; permission to cache and replay generated speech without an additional Polly charge
voice and engine availability varies by AWS Region; the editor, approval workflow and mastering chain must be supplied by the buyer; quality and pronunciation were not independently tested
Standard voices are listed at USD 4 per million characters, Neural at USD 16, Generative at USD 30 and Long-Form at USD 100 outside applicable free allowances.