Decision-critical facts remain separate from the deeper editorial analysis.
Best fit
Creator-facing voice production
Free evaluation
Free tier documented with usage limits
Pricing
Paid plans and prepaid options displayed; verify current locale pricing
Commercial use
Commercial-use eligibility has not been established.
Main caution
Resolve the open evidence fields before buying.
From evidence to action
Make the Voicemaker decision with the right unit and route.
Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.
The number that changes the decision
Credit/service calculator
Voicemaker: TTS/STT/conversion/cloning volumes and model family estimate the mixed allowance.
pricing · individual
Monthly Starter $5/200k credits, Creator $10/500k and Pro $20/1M; per-conversion limits 3k/5k/10k characters.
pricing · teams
Teams $30/month and Business $50/month; Enterprise custom with SSO and DPA/SLA options.
pricing · model multiplier
Default/Pro1/FlashX 1×, Pro2/Turbo 2×, Expressive/High-Res 4×; CJK can cost 2× on 1× families.
pricing · billing event
Credits are deducted on Convert, not download; unchanged text/voice may allow free setting-only reconverts per changelog.
pricing · rollover
Monthly Pro rolls one billing cycle; Starter/Creator and yearly individual plans do not; Teams/Business can roll up to three months.
pricing · topup
$20 buys 1M credits valid for one year on a paid plan; extra clone slot costs $2/year.
Decision-changing number
TTS/STT/conversion/cloning volumes and model family estimate the mixed allowance.
Keep in mind: Use only the retained Voicemaker evidence; do not generalize this decision asset to another product.
Open the complete Voicemaker buying analysisWorkflow · pricing · evidence boundaries · evaluation · official sources
# Voicemaker review: precise controls, models and true cost
> Distinctive strength: Voicemaker exposes unusually granular speech controls—SSML, pronunciation-oriented say-as rules, model families, effects, formats and streaming—while using one credit system across TTS, STT, speech-to-speech, cloning and audio effects. > > Where it stops being an advantage: The shared balance is not a shared unit: characters, model multipliers and seconds consume it differently, while clone governance and a public numerical SLA remain incomplete.
Voicemaker can serve a browser creator, audiobook producer or developer. Its model ladder ranges from economical default voices to multilingual, low-latency, high-resolution and expressive families. The buyer can control pauses, emphasis, dates, phone numbers, spelling, pitch, speed, volume and supported emotions.
The challenge is cost interpretation. One displayed credit is not always one generated character. Premium models can cost two or four times more, CJK can double some families, and STT/speech-to-speech use seconds rather than characters. A plan's nominal credit count is only the start of the calculation.
The model multiplier changes the buying answer
Question
Decision evidence
Strongest fit
Creator or developer needing granular SSML/output control and choice between economical and premium model families
Weak fit
Buyer needing a simple flat allowance, proven clone-consent enforcement or a public numerical SLA
Current monthly plans
Official surfaces agree on Starter $5/200k but conflict on Creator capacity and Pro price; preserve checkout
Model cost
1× default/Pro1/FlashX; 2× Pro2/Turbo; 4× Expressive/High-Res; CJK can be 2×
API
REST and WebSocket; 10k characters/request; WebSocket up to 20 concurrent requests
Rights
Live pricing says commercial rights on paid plans, but an older official surface says Creator onward—preserve checkout terms
Main unknown
Complete retention schedule and mandatory clone-consent enforcement
Voicemaker earns its keep through control and format breadth
Its strength is controllability without forcing every buyer into the most expensive model. Default voices can handle bulk narration at 1×, while Pro families trade credits for expression, clarity or low latency. SSML lets the editor solve predictable pronunciation problems rather than repeatedly regenerating a whole paragraph.
That is valuable in training, IVR, accessibility, audiobook and product-video work where dates, phone numbers, acronyms, pauses and emphasis must be repeatable. Catalogue size matters less than whether the required voice supports the exact control.
Use the AI voice software category to compare this production-control approach with simpler creator tools and cloud infrastructure.
Plan price hides the model-dependent capacity
Voicemaker's current official pricing response is internally inconsistent. Its comparison matrix lists the following individual plans:
Plan
Price
Monthly credits
Per conversion
Free
$0
Limited generations
250 characters
Starter
$5
200,000
3,000 characters
Creator
$10
500,000
5,000 characters
Pro
$20
1,000,000
10,000 characters
Another official card rendering captured during the same 30 August check showed Creator at 400,000 credits and Pro at $24 for one million. That is a first-party conflict, not a regional amount BenPicks can safely reconcile. Preserve the signed-in checkout before buying and treat the comparison matrix above as one published surface—not a guaranteed offer.
The comparison matrix also shows Teams at $30/month and Business at $50/month; Enterprise is custom. A $20 top-up adds 1M credits valid for one year. An additional eligible clone slot costs $2/year.
The dedicated audiobook plan is advertised at $25/year with 1M credits, an approximate 20 hours, 100,000 characters per conversion and 10GB storage. Treat the crossed-out/promo presentation as checkout-dependent.
Voicemaker bills conversions, not downloads. Therefore preview discipline matters: edit the full script before pressing Convert, and test whether a settings-only reconvert remains free under the current changelog rule.
At one million credits, a default English workload can process roughly one million characters. A 4× model can process about 250,000. If CJK and a multiplier both apply, confirm the exact arithmetic in usage history rather than extrapolating.
The plan page's hour estimates depend on language, speed and model. Characters are the safer planning unit; approved finished minutes are the safer business metric.
Which controls does Voicemaker provide?
The editor/help centre documents:
pause/break duration;
emphasis strength;
say-as for dates, time, addresses, telephone, spelling, digits, fractions and units;
global or selection-level speed, pitch and volume;
voice effects including happy, calm, sad, angry and shouting on supported voices;
accent/language control for compatible multilingual models.
Effects are not universal. Some default voice names carry an E suffix to signal compatibility. Test control support per voice and language; do not assume an effect button means every model honors it.
Catalogue breadth matters only after filtering by model and format
Pricing advertises 1,000+ default voices and 500+ Pro voices. Language scope reaches 130+ for default families, while higher-end families vary from 30+ to 90+.
Outputs include MP3 and OGG up to 192kbps, WAV 16-bit PCM up to 48kHz, OPUS, AAC and 8kHz telephony formats. This breadth is useful: a studio master, podcast download and IVR payload have different requirements.
Catalogue totals and “ultra-realistic” quality remain vendor claims. Test the actual language, model and format required.
Is the Voicemaker API suitable for production?
The standalone developer platform provides bearer-authenticated HTTPS REST endpoints for TTS, STT, speech-to-speech and clone operations. TTS requests accept up to 10,000 characters; documentation warns that inputs over 3,000 may take longer.
The WebSocket endpoint streams ordered chunks, supports up to 20 concurrent requests and closes after one minute of inactivity. Voice-list retrieval is not counted or billed. These are useful concrete implementation details.
A public numerical availability or latency SLA was not captured; Enterprise lists contractual DPA/SLA options. Before production, test timeout, retry, idempotency, credit restoration, stream assembly and version migration. The old endpoint is deprecated in favor of `/api/v1/`.
Speech-to-speech preserves source timing/pacing and uses ProPlus or cloned voices. Its documented rate is 100 credits per second. At that rate, a ten-minute file costs 60,000 credits before retries.
STT supports 90+ languages, optional speaker/audio-event handling and SRT export. Files up to three minutes process synchronously; longer files become asynchronous. The endpoint page says five credits/second, while the live pricing table says ten. This conflict materially doubles cost and should be confirmed before purchase.
Paid output has a commercial grant with boundaries
The live pricing FAQ says every paid plan includes personal and commercial rights and that subscribers retain copyright ownership of generated audio. Current terms exclude reselling Voicemaker's service itself.
An older official alpha pricing page says commercial rights begin at Creator rather than Starter. Because both surfaces exist, save the live checkout, Terms and invoice for the selected account. Inputs, third-party content and clone identity rights remain the user's responsibility even when output use is licensed.
Does Voicemaker train on customer text or audio?
The live pricing FAQ says input text and generated audio are not used to train models, are processed only to produce requested output and are not shared with third parties. That is a useful explicit statement.
However, a complete retention schedule for text, generated audio, histories, clones, API logs and backups was not captured. For sensitive material, request the current Privacy/GDPR terms, DPA, subprocessor list, data location and deletion schedule. Enterprise advertises DPA/SLA terms and SSO; Creator/Pro list 2FA.
How should voice cloning be evaluated?
Voicemaker offers plan-dependent clone slots and a professional cloning path based on about 30 minutes of source audio, with SSML, effects and API use. The public material does not document a mandatory speaker identity/consent ceremony.
Only clone a speaker with explicit authorization covering training, permitted outputs, duration, revocation and deletion. Follow the voice-cloning consent guide and obtain retention details before uploading biometric-sensitive recordings.
Choose Voicemaker when fine control pays for the complexity
Shortlist it when SSML precision, model choice, formats and API/browser parity solve real production problems. It can be especially attractive for long-form, training, IVR and multilingual workflows that benefit from pronunciation control.
Look elsewhere when a single predictable per-minute price matters more than model choice, public SLA/security evidence is mandatory, or clone governance cannot remain unknown.
Run the same script through three model multipliers
Save the live pricing, Terms and API version.
Prepare one fixed script with dates, telephone, acronym, currency, foreign phrase, emphasis and pauses.
Run it through one 1×, 2× and 4× model; record credits and accepted quality.
Change settings without text/voice changes and confirm billing behavior.
Export MP3, WAV and one telephony format; validate duration, sample rate and joins.
Stream the same input over WebSocket and test ordered concatenation, idle closure and concurrency.
Trigger one invalid request and document retry/credit handling.
If using STT, reconcile the five-versus-ten-credit rate in writing.
Confirm commercial rights and retention/training terms for the selected plan.
Test cloning only with an authorized speaker and document deletion.
Final assessment
Voicemaker's strength is not simply “many voices.” It gives a technical/editorial buyer unusually fine control over how speech is generated and delivered, with a model ladder that supports both economical bulk work and premium output.
That flexibility creates its own complexity. Capacity changes with model multipliers; STT documentation conflicts on price; commercial-rights wording differs across official surfaces; and retention/clone-consent evidence remains incomplete.
Choose it when the controls and formats reduce real correction work. Do not buy from the headline credit count alone.
Buy the smallest plan that proves the model
Start with the current Voicemaker editor and pricing ↗, then run the same production script through one 1×, one 2× and one 4× voice. The right upgrade is the cheapest plan that preserves the required controls and produces acceptable output after corrections. If the ledger cannot explain the model, language and tool debits, do not solve the uncertainty by buying a larger balance.
The evidence set was captured on 28 August 2026; current pricing, credit multipliers, usage rights and source availability were rechecked on 30 August 2026. BenPicks has not run hands-on voice-quality, latency, clone, security, deletion or billing tests.
Full evidence record7 fields · official links · dates · states
Best-fit workflow
Vendor claimChecked 2026-08-24
Browser TTS with granular voice controls and downloadable audio