Source-led profile · Evidence checked 2026-08-24

Verified essentials

Voice-generation studio

Voicemaker

Evaluate Voicemaker plans, model multipliers, SSML, API streaming, formats, rights, privacy gaps and a reproducible production test.

Decision first. Use the compact answer below before opening the complete research record.

Decision summary

The answer in one scan.

Decision-critical facts remain separate from the deeper editorial analysis.

Best fitCreator-facing voice production
Free evaluationFree tier documented with usage limits
PricingPaid plans and prepaid options displayed; verify current locale pricing
Commercial useCommercial-use eligibility has not been established.
Main cautionResolve the open evidence fields before buying.

From evidence to action

Make the Voicemaker decision with the right unit and route.

Each module separates documented facts, calculations and editorial conclusions. Missing or incompatible evidence stays visible instead of becoming a guess.

The number that changes the decision

Credit/service calculator

Voicemaker: TTS/STT/conversion/cloning volumes and model family estimate the mixed allowance.

pricing · individual
Monthly Starter $5/200k credits, Creator $10/500k and Pro $20/1M; per-conversion limits 3k/5k/10k characters.
pricing · teams
Teams $30/month and Business $50/month; Enterprise custom with SSO and DPA/SLA options.
pricing · model multiplier
Default/Pro1/FlashX 1×, Pro2/Turbo 2×, Expressive/High-Res 4×; CJK can cost 2× on 1× families.
pricing · billing event
Credits are deducted on Convert, not download; unchanged text/voice may allow free setting-only reconverts per changelog.
pricing · rollover
Monthly Pro rolls one billing cycle; Starter/Creator and yearly individual plans do not; Teams/Business can roll up to three months.
pricing · topup
$20 buys 1M credits valid for one year on a paid plan; extra clone slot costs $2/year.
Decision-changing number

TTS/STT/conversion/cloning volumes and model family estimate the mixed allowance.

Keep in mind: Use only the retained Voicemaker evidence; do not generalize this decision asset to another product.
Sources and verification date

Good fit if

Voicemaker matches the job you need done

  • Creator-facing voice production
  • The documented workflow and controls cover your required production steps.
  • You can validate the output with a representative project before committing.

Look elsewhere if

You need certainty this profile cannot provide

  • You need independently tested output quality rather than documented capabilities.
  • Your purchase depends on one of the 3 facts still requiring confirmation.
  • A narrower product would complete the same job with less workflow overhead.

Price, plan and risks

Confirm before you buy

Unknown, conflicted and stale facts stay visible before checkout.

Unknown

Languages

The retained evidence does not establish this field yet.

Unknown

Commercial-use eligibility

The retained evidence does not establish this field yet.

Unknown

Voice cloning

The retained evidence does not establish this field yet.

Commercial context

Compare the closest documented workflows.

Alternatives stay within the same vertical and use current internal profile routes.

Open the complete Voicemaker buying analysisWorkflow · pricing · evidence boundaries · evaluation · official sources

# Voicemaker review: precise controls, models and true cost

> Distinctive strength: Voicemaker exposes unusually granular speech controls—SSML, pronunciation-oriented say-as rules, model families, effects, formats and streaming—while using one credit system across TTS, STT, speech-to-speech, cloning and audio effects. > > Where it stops being an advantage: The shared balance is not a shared unit: characters, model multipliers and seconds consume it differently, while clone governance and a public numerical SLA remain incomplete.

Voicemaker can serve a browser creator, audiobook producer or developer. Its model ladder ranges from economical default voices to multilingual, low-latency, high-resolution and expressive families. The buyer can control pauses, emphasis, dates, phone numbers, spelling, pitch, speed, volume and supported emotions.

The challenge is cost interpretation. One displayed credit is not always one generated character. Premium models can cost two or four times more, CJK can double some families, and STT/speech-to-speech use seconds rather than characters. A plan's nominal credit count is only the start of the calculation.

An editorial illustration of a script flowing through SSML controls, voice models and multiple audio formats

The model multiplier changes the buying answer

QuestionDecision evidence
Strongest fitCreator or developer needing granular SSML/output control and choice between economical and premium model families
Weak fitBuyer needing a simple flat allowance, proven clone-consent enforcement or a public numerical SLA
Current monthly plansOfficial surfaces agree on Starter $5/200k but conflict on Creator capacity and Pro price; preserve checkout
Model cost1× default/Pro1/FlashX; 2× Pro2/Turbo; 4× Expressive/High-Res; CJK can be 2×
APIREST and WebSocket; 10k characters/request; WebSocket up to 20 concurrent requests
RightsLive pricing says commercial rights on paid plans, but an older official surface says Creator onward—preserve checkout terms
Main unknownComplete retention schedule and mandatory clone-consent enforcement

Voicemaker earns its keep through control and format breadth

Its strength is controllability without forcing every buyer into the most expensive model. Default voices can handle bulk narration at 1×, while Pro families trade credits for expression, clarity or low latency. SSML lets the editor solve predictable pronunciation problems rather than repeatedly regenerating a whole paragraph.

That is valuable in training, IVR, accessibility, audiobook and product-video work where dates, phone numbers, acronyms, pauses and emphasis must be repeatable. Catalogue size matters less than whether the required voice supports the exact control.

Use the AI voice software category to compare this production-control approach with simpler creator tools and cloud infrastructure.

Plan price hides the model-dependent capacity

Voicemaker's current official pricing response is internally inconsistent. Its comparison matrix lists the following individual plans:

PlanPriceMonthly creditsPer conversion
Free$0Limited generations250 characters
Starter$5200,0003,000 characters
Creator$10500,0005,000 characters
Pro$201,000,00010,000 characters

Another official card rendering captured during the same 30 August check showed Creator at 400,000 credits and Pro at $24 for one million. That is a first-party conflict, not a regional amount BenPicks can safely reconcile. Preserve the signed-in checkout before buying and treat the comparison matrix above as one published surface—not a guaranteed offer.

The comparison matrix also shows Teams at $30/month and Business at $50/month; Enterprise is custom. A $20 top-up adds 1M credits valid for one year. An additional eligible clone slot costs $2/year.

The dedicated audiobook plan is advertised at $25/year with 1M credits, an approximate 20 hours, 100,000 characters per conversion and 10GB storage. Treat the crossed-out/promo presentation as checkout-dependent.

Voicemaker bills conversions, not downloads. Therefore preview discipline matters: edit the full script before pressing Convert, and test whether a settings-only reconvert remains free under the current changelog rule.

Use the AI voice generation cost guide to calculate accepted minutes after multipliers and corrections.

Model multipliers shrink the headline allowance

Model choice changes the meter:

  • default AI1–AI6/HashCode: 1 credit/character;
  • Pro1: 1×, with CJK at 2×;
  • FlashX: 1×;
  • Pro2 and ProPlus Turbo: 2×;
  • ProPlus Expressive and High-Res: 4×.

At one million credits, a default English workload can process roughly one million characters. A 4× model can process about 250,000. If CJK and a multiplier both apply, confirm the exact arithmetic in usage history rather than extrapolating.

The plan page's hour estimates depend on language, speed and model. Characters are the safer planning unit; approved finished minutes are the safer business metric.

Which controls does Voicemaker provide?

The editor/help centre documents:

  • pause/break duration;
  • emphasis strength;
  • say-as for dates, time, addresses, telephone, spelling, digits, fractions and units;
  • global or selection-level speed, pitch and volume;
  • voice effects including happy, calm, sad, angry and shouting on supported voices;
  • accent/language control for compatible multilingual models.

Effects are not universal. Some default voice names carry an E suffix to signal compatibility. Test control support per voice and language; do not assume an effect button means every model honors it.

Catalogue breadth matters only after filtering by model and format

Pricing advertises 1,000+ default voices and 500+ Pro voices. Language scope reaches 130+ for default families, while higher-end families vary from 30+ to 90+.

Outputs include MP3 and OGG up to 192kbps, WAV 16-bit PCM up to 48kHz, OPUS, AAC and 8kHz telephony formats. This breadth is useful: a studio master, podcast download and IVR payload have different requirements.

Catalogue totals and “ultra-realistic” quality remain vendor claims. Test the actual language, model and format required.

Is the Voicemaker API suitable for production?

The standalone developer platform provides bearer-authenticated HTTPS REST endpoints for TTS, STT, speech-to-speech and clone operations. TTS requests accept up to 10,000 characters; documentation warns that inputs over 3,000 may take longer.

The WebSocket endpoint streams ordered chunks, supports up to 20 concurrent requests and closes after one minute of inactivity. Voice-list retrieval is not counted or billed. These are useful concrete implementation details.

A public numerical availability or latency SLA was not captured; Enterprise lists contractual DPA/SLA options. Before production, test timeout, retry, idempotency, credit restoration, stream assembly and version migration. The old endpoint is deprecated in favor of `/api/v1/`.

Use the AI voice API buying guide to compare operational contracts rather than voice demos alone.

What do speech-to-text and speech-to-speech cost?

Speech-to-speech preserves source timing/pacing and uses ProPlus or cloned voices. Its documented rate is 100 credits per second. At that rate, a ten-minute file costs 60,000 credits before retries.

STT supports 90+ languages, optional speaker/audio-event handling and SRT export. Files up to three minutes process synchronously; longer files become asynchronous. The endpoint page says five credits/second, while the live pricing table says ten. This conflict materially doubles cost and should be confirmed before purchase.

Paid output has a commercial grant with boundaries

The live pricing FAQ says every paid plan includes personal and commercial rights and that subscribers retain copyright ownership of generated audio. Current terms exclude reselling Voicemaker's service itself.

An older official alpha pricing page says commercial rights begin at Creator rather than Starter. Because both surfaces exist, save the live checkout, Terms and invoice for the selected account. Inputs, third-party content and clone identity rights remain the user's responsibility even when output use is licensed.

Does Voicemaker train on customer text or audio?

The live pricing FAQ says input text and generated audio are not used to train models, are processed only to produce requested output and are not shared with third parties. That is a useful explicit statement.

However, a complete retention schedule for text, generated audio, histories, clones, API logs and backups was not captured. For sensitive material, request the current Privacy/GDPR terms, DPA, subprocessor list, data location and deletion schedule. Enterprise advertises DPA/SLA terms and SSO; Creator/Pro list 2FA.

How should voice cloning be evaluated?

Voicemaker offers plan-dependent clone slots and a professional cloning path based on about 30 minutes of source audio, with SSML, effects and API use. The public material does not document a mandatory speaker identity/consent ceremony.

Only clone a speaker with explicit authorization covering training, permitted outputs, duration, revocation and deletion. Follow the voice-cloning consent guide and obtain retention details before uploading biometric-sensitive recordings.

Choose Voicemaker when fine control pays for the complexity

Shortlist it when SSML precision, model choice, formats and API/browser parity solve real production problems. It can be especially attractive for long-form, training, IVR and multilingual workflows that benefit from pronunciation control.

Look elsewhere when a single predictable per-minute price matters more than model choice, public SLA/security evidence is mandatory, or clone governance cannot remain unknown.

Run the same script through three model multipliers

  1. Save the live pricing, Terms and API version.
  2. Prepare one fixed script with dates, telephone, acronym, currency, foreign phrase, emphasis and pauses.
  3. Run it through one 1×, 2× and 4× model; record credits and accepted quality.
  4. Change settings without text/voice changes and confirm billing behavior.
  5. Export MP3, WAV and one telephony format; validate duration, sample rate and joins.
  6. Stream the same input over WebSocket and test ordered concatenation, idle closure and concurrency.
  7. Trigger one invalid request and document retry/credit handling.
  8. If using STT, reconcile the five-versus-ten-credit rate in writing.
  9. Confirm commercial rights and retention/training terms for the selected plan.
  10. Test cloning only with an authorized speaker and document deletion.

Final assessment

Voicemaker's strength is not simply “many voices.” It gives a technical/editorial buyer unusually fine control over how speech is generated and delivered, with a model ladder that supports both economical bulk work and premium output.

That flexibility creates its own complexity. Capacity changes with model multipliers; STT documentation conflicts on price; commercial-rights wording differs across official surfaces; and retention/clone-consent evidence remains incomplete.

Choose it when the controls and formats reduce real correction work. Do not buy from the headline credit count alone.

Buy the smallest plan that proves the model

Start with the current Voicemaker editor and pricing ↗, then run the same production script through one 1×, one 2× and one 4× voice. The right upgrade is the cheapest plan that preserves the required controls and produces acceptable output after corrections. If the ledger cannot explain the model, language and tool debits, do not solve the uncertainty by buying a larger balance.

Official sources checked

The evidence set was captured on 28 August 2026; current pricing, credit multipliers, usage rights and source availability were rechecked on 30 August 2026. BenPicks has not run hands-on voice-quality, latency, clone, security, deletion or billing tests.

Full evidence record7 fields · official links · dates · states

Best-fit workflow

Vendor claimChecked 2026-08-24

Browser TTS with granular voice controls and downloadable audio

Official sources (1)

Free evaluation

Vendor claimChecked 2026-08-24

Free tier documented with usage limits

Official sources (1)

Pricing observation

Vendor claimChecked 2026-08-24

Paid plans and prepaid options displayed; verify current locale pricing

Official sources (1)

Languages

UnknownChecked 2026-08-24

The retained official evidence does not answer this yet.

Official sources (1)

Commercial-use eligibility

UnknownChecked 2026-08-24

The retained official evidence does not answer this yet.

Official sources (1)

Voice cloning

UnknownChecked 2026-08-24

The retained official evidence does not answer this yet.

Official sources (1)

API

Vendor claimChecked 2026-08-24

API access available on qualifying plans

Official sources (1)