Guide · Evidence checked 2026-08-28
How to Calculate Multilingual Dubbing and Lip-Sync Cost
A practical, evidence-led decision guide. Product capabilities and limits are separated from anything that would require hands-on testing.
A 30-minute video translated into four languages is not a 30-minute dubbing job. It creates at least 120 target-language minutes before transcript cleanup, regenerated lines, lip-sync, subtitle correction, reviewer time and exports. A pricing page that says “minutes included” is incomplete until the buyer knows which of those minutes it counts.
This calculator keeps incompatible units separate and ends with the cost of an approved localized deliverable. It does not turn a credit multiplier into a price claim or assume every product sells the same workflow.
Draw the production map first
Write the source duration (`D`) and number of target languages (`L`). Then record which stages the vendor meters:
- source transcription;
- translation;
- generated target speech;
- voice cloning or speaker assignment;
- lip-sync;
- subtitles/captions;
- regenerated segments;
- export/render;
- storage, seats or projects;
- human language and production review.
The minimum localized workload is:
`target-language minutes = source duration × target languages`
For a 30-minute source in four languages:
`30 × 4 = 120 target-language minutes`
This number is a workload dimension, not a guaranteed billable unit. Some platforms charge source video minutes, some generated minutes, some credits and some separate lip-sync allowances.
Add regeneration without hiding it
Let `R` be the share of target-language audio regenerated after review:
`generated speech minutes = D × L × (1 + R)`
At 25% regeneration:
`30 × 4 × 1.25 = 150 generated minutes`
At 50% regeneration:
`30 × 4 × 1.50 = 180 generated minutes`
The regeneration rate should come from a pilot that includes real names, numbers, timing and brand language. A vendor’s ideal demo cannot supply it.
Keep correction types separate. A translation rewrite may trigger new speech and lip-sync. A subtitle punctuation correction may not. A pronunciation dictionary update can regenerate one line rather than the full video if the product supports segment-level work.
Lip-sync is its own production route
Lip-sync can be included, sold as a separate minute, or consume a multiplier of another allowance. Never describe a 2× or 4× credit rule as “two/four times the price” unless the plan price, included credits, overage, source-minute definition and complete workload are normalized.
For a multiplier `M` applied to target-language minutes:
`lip-sync allowance consumed = D × L × M`
The 30-minute/four-language project consumes 240 allowance-minutes at 2× or 480 at 4×. If the retained vendor evidence conflicts between 2× and 4×, publish both scenarios and stop before claiming a total.
This is the current evidence boundary for Rask AI: retained official surfaces do not reconcile one lip-sync multiplier. The conflict is a checkout/procurement question, not a value to average.
Build a unit ledger by product
Rask AI
Rask’s relevant workflow is editable localization of existing video: transcription/translation, speaker handling, generated speech and lip-sync. Model the base localization minutes separately from any lip-sync multiplier. Preserve the current conflict until the account shows the actual debit.
Its strength is the direct localization workflow. It stops being an advantage if disputed metering makes the full language basket unpredictable or native review reveals extensive semantic/timing correction.
HeyGen
HeyGen combines avatar/presenter production and video translation. A buyer translating an existing presenter video should not price it as if it were only TTS. Plan/credit rules, source duration, target languages, presenter/Avatar workflow and lip-sync route all matter.
HeyGen is stronger when the visual presenter is central to the deliverable. It can be unnecessary when the job is audio dubbing over footage that does not need an avatar system.
Murf
Murf separates Studio, API and Dubbing products. Their meters should never be pooled. Studio-generation economics do not automatically describe Dubbing, and API character billing does not price a video-localization job.
Choose the product route first, then record the entitlement and units for that route only. A single “Murf price” is not a useful input.
Dubverse, Maestra and Kapwing
Dubverse, Maestra and Kapwing AI Dubbing combine different mixes of transcription, translation, dubbing, subtitles and editor/export capabilities. Their apparent minute allowances can refer to different deliverables.
Use them to test whether an integrated editor reduces review and handoff cost, not merely whether the headline minute is lower. Record speaker limits, languages, export restrictions, subtitle workflows and what happens when one segment is regenerated.
Calculate the complete platform bill
For each candidate, keep a row for every native meter:
| Meter | Required quantity | Included | Overage/upgrade | Evidence state |
|---|---|---|---|---|
| Source minutes | `D` | plan-specific | plan-specific | verified/vendor claim/conflicted/unknown |
| Target-language minutes | `D × L` | plan-specific | plan-specific | same taxonomy |
| Generated speech | `D × L × (1+R)` | plan-specific | plan-specific | same taxonomy |
| Lip-sync | documented unit/multiplier | plan-specific | plan-specific | same taxonomy |
| Transcription/subtitles | product-specific | plan-specific | plan-specific | same taxonomy |
| Seats/projects/storage | actual requirement | plan-specific | plan-specific | same taxonomy |
| Exports/renders | accepted deliverables | plan-specific | plan-specific | same taxonomy |
Then calculate:
`platform cost = base plan + required upgrades + overage + add-ons`
Do not fill unknown overage with zero. Use a range if every bound is sourced; otherwise request a worked quote using the exact project.
Add human production cost
The accepted localized video needs work the platform may not price:
- source-transcript cleanup;
- translation and terminology review;
- cultural adaptation;
- speaker assignment and consent checks;
- pronunciation correction;
- timing and scene-fit edits;
- lip-sync artifact review;
- subtitle readability and line breaks;
- audio mixing and loudness;
- legal/brand approval;
- rendering, upload and final QA.
Use reviewer rates and measured hours:
`human cost = sum(role hours × loaded hourly rate)`
Then calculate the useful denominator:
`cost per accepted localized minute = (platform cost + human cost) ÷ approved target-language minutes`
For the 30-minute/four-language project, the denominator is 120 approved minutes only if all four versions pass. If one language is rejected, do not count its 30 minutes as accepted output.
A worked comparison template
Assume:
- 30-minute source;
- four target languages;
- 25% measured regeneration;
- 12 reviewer hours at $50/hour;
- platform subtotal `P` after native-meter calculation.
The workload is 120 target-language minutes and 150 generated minutes. Human review is $600. Complete accepted-minute cost, if all versions pass, is:
`(P + $600) ÷ 120`
If the platform subtotal is $300, complete cost is $7.50 per accepted localized minute. If only three languages pass, the accepted denominator becomes 90 minutes while some rejected-language cost remains, raising the effective cost to $10.00 before additional correction.
The illustration is not a vendor quote. It shows why approval and rejection belong in the denominator.
Model the language basket, not an average language
Use at least three difficulty classes:
- Near-timing language: translated speech commonly fits similar scene timing.
- Expansion language: translated text tends to require more words or faster delivery.
- Different script/segmentation: subtitles, names, line breaks and visual timing need distinct review.
Record regeneration and reviewer time by language. One difficult market can set the plan and schedule even when the average looks acceptable.
Also classify speakers: narrator, interview, overlapping dialogue, off-screen speech and emotionally expressive performance. A five-speaker documentary is not equivalent to a one-speaker training video of the same duration.
The ten-minute pilot that predicts the full job
Choose a ten-minute source containing:
- at least three speakers;
- names, numbers and product terminology;
- one overlapping exchange;
- visible mouth close-ups;
- on-screen text and subtitles;
- emotional and neutral passages;
- pauses, music and background noise.
Localize it into two contrasting languages. Preserve the original transcript, edited translation, speaker map, voice consent, generated segments, retries, credit ledger and exports.
Have native reviewers grade semantic accuracy, omissions/additions, pronunciation, voice consistency, timing, lip-sync, subtitle readability and cultural fit. Record every edit and debit.
Use pilot observations to calculate:
- regeneration rate by language;
- reviewer minutes per finished minute;
- lip-sync rejection rate;
- platform units per accepted minute;
- complete cost at baseline, +25% and +50% project growth.
Do not scale a vendor-provided sample or a clean single-speaker clip.
Rights and data can stop the cheapest route
For every cloned or assigned voice, preserve the speaker’s authority, permitted use, territories, duration, revocation process and model-deletion procedure. Commercial output rights do not prove voice consent.
Map source video, transcripts, translations, voice samples, face data, generated audio/video, logs and exports. Obtain retention, training/improvement, subprocessors, regions, deletion/backups and access controls. Lip-sync and avatar workflows can involve biometric or highly sensitive media; price does not override policy.
Confirm whether output and project files remain usable after cancellation, which assets can be exported, and whether a vendor watermark or plan restriction affects delivery.
Capacity cliffs to test before purchase
Run scenarios for:
- one additional language;
- 25% and 50% regeneration;
- a second reviewer pass;
- lip-sync on all languages versus only hero markets;
- peak monthly source duration;
- extra seats and projects;
- 4K or premium export where relevant;
- one rejected language requiring rework.
The capacity cliff is the smallest change that forces a higher plan or add-on. That number is more useful than average cost.
Fail closed on incompatible or disputed units
Stop the calculation when:
- “minute” is not defined as source, generated, downloaded or finished;
- lip-sync multiplier or allowance conflicts across official pages;
- credits cannot be mapped to the tested workflow;
- retries/regeneration treatment is unknown;
- languages or voices require undisclosed tiers;
- monthly and annual prices are mixed;
- output/export rights or consent are unresolved;
- a quote bundles services that cannot be separated.
The honest result can be “total unavailable pending checkout.” False precision does not help a buyer.
The decision this model should produce
Choose the platform that produces approved, rights-cleared localized video at the lowest complete cost while meeting language, timing, lip-sync, data and workflow gates. A lower headline minute is not enough. A more expensive integrated editor can win when it removes handoffs and correction work; a specialist can win when its native localization workflow reduces rejected output.
Compare Rask AI and HeyGen directly and use our AI voice localization shortlist to choose a small pilot set.