Guide · Evidence checked 2026-08-28

How to Calculate Multilingual Dubbing and Lip-Sync Cost

A practical, evidence-led decision guide. Product capabilities and limits are separated from anything that would require hands-on testing.

A 30-minute video translated into four languages is not a 30-minute dubbing job. It creates at least 120 target-language minutes before transcript cleanup, regenerated lines, lip-sync, subtitle correction, reviewer time and exports. A pricing page that says “minutes included” is incomplete until the buyer knows which of those minutes it counts.

This calculator keeps incompatible units separate and ends with the cost of an approved localized deliverable. It does not turn a credit multiplier into a price claim or assume every product sells the same workflow.

Draw the production map first

Write the source duration (`D`) and number of target languages (`L`). Then record which stages the vendor meters:

  1. source transcription;
  2. translation;
  3. generated target speech;
  4. voice cloning or speaker assignment;
  5. lip-sync;
  6. subtitles/captions;
  7. regenerated segments;
  8. export/render;
  9. storage, seats or projects;
  10. human language and production review.

The minimum localized workload is:

`target-language minutes = source duration × target languages`

For a 30-minute source in four languages:

`30 × 4 = 120 target-language minutes`

This number is a workload dimension, not a guaranteed billable unit. Some platforms charge source video minutes, some generated minutes, some credits and some separate lip-sync allowances.

Add regeneration without hiding it

Let `R` be the share of target-language audio regenerated after review:

`generated speech minutes = D × L × (1 + R)`

At 25% regeneration:

`30 × 4 × 1.25 = 150 generated minutes`

At 50% regeneration:

`30 × 4 × 1.50 = 180 generated minutes`

The regeneration rate should come from a pilot that includes real names, numbers, timing and brand language. A vendor’s ideal demo cannot supply it.

Keep correction types separate. A translation rewrite may trigger new speech and lip-sync. A subtitle punctuation correction may not. A pronunciation dictionary update can regenerate one line rather than the full video if the product supports segment-level work.

Lip-sync is its own production route

Lip-sync can be included, sold as a separate minute, or consume a multiplier of another allowance. Never describe a 2× or 4× credit rule as “two/four times the price” unless the plan price, included credits, overage, source-minute definition and complete workload are normalized.

For a multiplier `M` applied to target-language minutes:

`lip-sync allowance consumed = D × L × M`

The 30-minute/four-language project consumes 240 allowance-minutes at 2× or 480 at 4×. If the retained vendor evidence conflicts between 2× and 4×, publish both scenarios and stop before claiming a total.

This is the current evidence boundary for Rask AI: retained official surfaces do not reconcile one lip-sync multiplier. The conflict is a checkout/procurement question, not a value to average.

Build a unit ledger by product

Rask AI

Rask’s relevant workflow is editable localization of existing video: transcription/translation, speaker handling, generated speech and lip-sync. Model the base localization minutes separately from any lip-sync multiplier. Preserve the current conflict until the account shows the actual debit.

Its strength is the direct localization workflow. It stops being an advantage if disputed metering makes the full language basket unpredictable or native review reveals extensive semantic/timing correction.

HeyGen

HeyGen combines avatar/presenter production and video translation. A buyer translating an existing presenter video should not price it as if it were only TTS. Plan/credit rules, source duration, target languages, presenter/Avatar workflow and lip-sync route all matter.

HeyGen is stronger when the visual presenter is central to the deliverable. It can be unnecessary when the job is audio dubbing over footage that does not need an avatar system.

Murf

Murf separates Studio, API and Dubbing products. Their meters should never be pooled. Studio-generation economics do not automatically describe Dubbing, and API character billing does not price a video-localization job.

Choose the product route first, then record the entitlement and units for that route only. A single “Murf price” is not a useful input.

Dubverse, Maestra and Kapwing

Dubverse, Maestra and Kapwing AI Dubbing combine different mixes of transcription, translation, dubbing, subtitles and editor/export capabilities. Their apparent minute allowances can refer to different deliverables.

Use them to test whether an integrated editor reduces review and handoff cost, not merely whether the headline minute is lower. Record speaker limits, languages, export restrictions, subtitle workflows and what happens when one segment is regenerated.

Calculate the complete platform bill

For each candidate, keep a row for every native meter:

MeterRequired quantityIncludedOverage/upgradeEvidence state
Source minutes`D`plan-specificplan-specificverified/vendor claim/conflicted/unknown
Target-language minutes`D × L`plan-specificplan-specificsame taxonomy
Generated speech`D × L × (1+R)`plan-specificplan-specificsame taxonomy
Lip-syncdocumented unit/multiplierplan-specificplan-specificsame taxonomy
Transcription/subtitlesproduct-specificplan-specificplan-specificsame taxonomy
Seats/projects/storageactual requirementplan-specificplan-specificsame taxonomy
Exports/rendersaccepted deliverablesplan-specificplan-specificsame taxonomy

Then calculate:

`platform cost = base plan + required upgrades + overage + add-ons`

Do not fill unknown overage with zero. Use a range if every bound is sourced; otherwise request a worked quote using the exact project.

Add human production cost

The accepted localized video needs work the platform may not price:

Use reviewer rates and measured hours:

`human cost = sum(role hours × loaded hourly rate)`

Then calculate the useful denominator:

`cost per accepted localized minute = (platform cost + human cost) ÷ approved target-language minutes`

For the 30-minute/four-language project, the denominator is 120 approved minutes only if all four versions pass. If one language is rejected, do not count its 30 minutes as accepted output.

A worked comparison template

Assume:

The workload is 120 target-language minutes and 150 generated minutes. Human review is $600. Complete accepted-minute cost, if all versions pass, is:

`(P + $600) ÷ 120`

If the platform subtotal is $300, complete cost is $7.50 per accepted localized minute. If only three languages pass, the accepted denominator becomes 90 minutes while some rejected-language cost remains, raising the effective cost to $10.00 before additional correction.

The illustration is not a vendor quote. It shows why approval and rejection belong in the denominator.

Model the language basket, not an average language

Use at least three difficulty classes:

Record regeneration and reviewer time by language. One difficult market can set the plan and schedule even when the average looks acceptable.

Also classify speakers: narrator, interview, overlapping dialogue, off-screen speech and emotionally expressive performance. A five-speaker documentary is not equivalent to a one-speaker training video of the same duration.

The ten-minute pilot that predicts the full job

Choose a ten-minute source containing:

Localize it into two contrasting languages. Preserve the original transcript, edited translation, speaker map, voice consent, generated segments, retries, credit ledger and exports.

Have native reviewers grade semantic accuracy, omissions/additions, pronunciation, voice consistency, timing, lip-sync, subtitle readability and cultural fit. Record every edit and debit.

Use pilot observations to calculate:

Do not scale a vendor-provided sample or a clean single-speaker clip.

Rights and data can stop the cheapest route

For every cloned or assigned voice, preserve the speaker’s authority, permitted use, territories, duration, revocation process and model-deletion procedure. Commercial output rights do not prove voice consent.

Map source video, transcripts, translations, voice samples, face data, generated audio/video, logs and exports. Obtain retention, training/improvement, subprocessors, regions, deletion/backups and access controls. Lip-sync and avatar workflows can involve biometric or highly sensitive media; price does not override policy.

Confirm whether output and project files remain usable after cancellation, which assets can be exported, and whether a vendor watermark or plan restriction affects delivery.

Capacity cliffs to test before purchase

Run scenarios for:

The capacity cliff is the smallest change that forces a higher plan or add-on. That number is more useful than average cost.

Fail closed on incompatible or disputed units

Stop the calculation when:

The honest result can be “total unavailable pending checkout.” False precision does not help a buyer.

The decision this model should produce

Choose the platform that produces approved, rights-cleared localized video at the lowest complete cost while meeting language, timing, lip-sync, data and workflow gates. A lower headline minute is not enough. A more expensive integrated editor can win when it removes handoffs and correction work; a specialist can win when its native localization workflow reduces rejected output.

Compare Rask AI and HeyGen directly and use our AI voice localization shortlist to choose a small pilot set.

Official sources checked