# Altered Studio review: test the performance, not just the target voice
> Distinctive strength: Speech-to-speech morphing can preserve an actor's timing and performance while changing the apparent voice, which plain text-to-speech cannot reproduce. > > Where it stops being an advantage: Recording quality, model choice, repeated generation and rights clearance can erase the workflow advantage.
Altered Studio is not primarily a conventional text-to-speech generator. Its distinctive job is speech-to-speech morphing: an actor records a performance, then the software changes the apparent voice while preserving—or deliberately reshaping—timing, energy, accent and prosody. That can be more useful than TTS for character dialogue, pickups and localized performance, but it also creates more variables to review.
The buying question is whether Altered reduces total production effort after recording, repeated generation, sample selection, correction, rights clearance and export. A convincing voice-library demo does not answer that question because the final morph depends on the source performance, model family and settings.
BenPicks has not operated Altered Studio for this review. All feature, quota and licence statements below come from current official documentation. No hands-on quality ranking is implied.
Altered is a performance-conversion decision
| Buyer question | Source-led answer |
|---|
| Best fit | Performance-led character or voiceover post-production where an actor's timing and delivery should survive a voice change. |
| Weak fit | Simple narration, transparent per-character buying, one-click video dubbing or workflows that cannot clear source-voice/data rights. |
| Core workflow | Morph a whole file, transcript block or waveform selection using voice/model-specific controls. |
| Model tradeoff | Timbre, Performance, Flexi, Clone, Narration and Fast variants prioritize different aspects of accent, likeness and dynamics. |
| Regeneration | Most morph models are non-deterministic; Fast variants are documented as deterministic. |
| Rapid clone | Recommended 4–8 seconds, maximum 20 seconds; positioned for short-form pickups rather than a full nuanced custom model. |
| Main unknowns | Current Studio prices/allowances, full deletion lifecycle, security assurance and team governance. |
Morphing starts where text-to-speech stops
Text-to-speech starts from written language and asks a model to create a performance. Altered's speech-to-speech workflow starts with a human performance. The actor can supply pacing, emotion, breaths and scene timing, then the target model changes vocal identity or characteristics.
Official documentation lets an editor apply a Morph effect to an entire file, a transcript block or a selected waveform region. Different layers can use different voices, which makes one actor capable of roughing out multiple characters. The practical advantage is control: if a line needs a particular pause or emphasis, the actor can perform it rather than trying to prompt it into existence.
The practical risk is correction work. Background noise, room reverb and effects can interfere with synthesis. Overlapping morph layers can accidentally process already-morphed audio. Non-deterministic models can produce slightly different emphasis on repeat runs. A useful trial therefore needs finished scenes, not isolated one-line samples.
Each morph family sacrifices something different
Altered documents several model families:
| Model family | Intended distinction | What to test |
|---|
| Timbre | Cross-lingual use; preserves accent, sounds and emotes from input. | Code-switching, accent retention and intelligibility. |
| Performance | English, closer reproduction of input dynamics with Prosody control. | Emotional timing and actor identity leakage. |
| Flexi | English with age, gender and loudness shifts. | Whether transformation remains believable across a full scene. |
| Clone | Favors target-voice likeness more than input-speaker influence. | Target similarity versus loss of performance detail. |
| Narration | More constrained dynamics for narration. | Long-passage consistency and listener fatigue. |
| Fast | Lower-fidelity, deterministic versions for prototyping. | Whether faster iteration offsets the quality reduction. |
Availability depends on subscription and voice. Do not select one winner from marketing labels. Freeze the same input performance and generate every candidate under named settings. Blind the files, then ask reviewers to score intelligibility, performance preservation, target fit, artefacts and production readiness separately.
Variation creates review cost even when quota does not move
Altered says most morph models can vary slightly between runs; Fast models produce consistent output. That means the cost of one accepted line includes listening to alternatives and preserving the selected sample.
Quota accounting is unusual and potentially favorable: speech-to-speech usage is charged on the first morph of a selected audio region, while later morphs of that same selection are documented as not consuming additional quota. But a production team must establish what counts as “the same selection.” Changing boundaries, editing source audio or reopening a project may affect the operational result.
Track this during the trial:
| Metric | Why it matters |
|---|
| Source minutes recorded | Actor cost exists before morphing. |
| Unique selections first morphed | Closest documented quota unit. |
| Samples generated per accepted line | Captures non-deterministic selection work. |
| Lines requiring source retake | Reveals whether performance/noise constraints dominate. |
| Lines requiring transcript reinforcement | Measures intelligibility correction. |
| Editor review minutes | Can exceed synthesis time. |
| Final delivered minutes | Denominator for comparing alternatives. |
Do not convert subscription quota directly into delivered hours until this ratio is known.
What controls can shape the result?
Compatible models expose a meaningful set of controls. Target Prosody moves weighting between the input performance and the target voice's natural characteristics. Pitch Shift changes output pitch; official guidance says a range around plus or minus two semitones tends to work best even though the control extends further. Timbre and Flexi variants can expose age, gender and loudness shifts.
Text Reinforcement uses the transcript to correct minor mispronunciations on certain plans/models. The transcript must be corrected before the morph. Decreak, Power Envelope and post-processing can address fry, inconsistent dynamics or artefacts; post-processing is not available on Fast models. Output can be generated at 24 kHz by default or 48 kHz when selected.
These settings should become presets only after a scene passes review. Altered supports reusable global and file presets, but morph and TTS presets are separate. Preserve the model version, voice, settings and sample decision next to each approved line; otherwise later pickups may not reproduce the original production choice.
How does Rapid Cloning work, and what does it not promise?
Rapid Cloning creates what Altered calls a “vocal footprint.” Official guidance recommends 4–8 seconds of clean audio and caps total training input at 20 seconds. The source should be dry, close-mic audio without effects or room reverb. Trial users can preview a watermarked/noisy result, while saving the voice into the library requires paid access.
Altered explicitly distinguishes this from its separate full custom-voice service. Rapid Cloning is intended for short-form voiceovers and pickups, not a fully nuanced model. That boundary is useful: a buyer should not interpret a successful ten-second sample as evidence that the clone will sustain long-form narration, shouting, whispering or multiple languages.
The creator must hold rights to the voice and recordings or obtain the speaker's consent. Terms also reserve a private record of source files to enforce compliance. Maintain your own evidence package:
- speaker identity and authority;
- exact permitted projects, channels and duration;
- source recording provenance;
- whether model creation, sharing and future reuse are permitted;
- withdrawal and deletion process;
- final disposition when the subscription ends.
Do not upload a third party merely because a clip is publicly available.
Commercial permission depends on company and licence class
Commercial permission is plan- and buyer-dependent. Current terms describe a Creator licence for individuals or companies with no more than five employees and less than $100,000 in relevant revenue or funding over the previous 12 months. A Commercial licence provides broader commercial use under the terms. Attribution, Streaming and Real-Time Pro licences have different limits.
The output restrictions matter as much as the grant. Terms prohibit claiming stock voices through systems such as YouTube Content ID, using Outputs to train another AI or speech system, and using Altered's voice names in the Project. Altered retains ownership of the Voices while stating it does not claim ownership of Content or Output beyond the rights granted in the agreement.
Before production, record the legal entity, headcount, relevant revenue/funding, subscription and licence shown at checkout. If an agency creates work for clients, obtain written confirmation that client delivery, paid advertising, broadcast, game distribution and long-lived archive use fit the chosen licence. A Creator licence that qualified at signup may stop fitting as the business grows.
What happens to uploaded performances and generated audio?
This is a material procurement question. The terms grant Altered an irrevocable worldwide licence to use Content and Outputs to maintain, develop and enhance its software. Discounted Real-Time Pro plans add an explicit AI Improvement Program licence over voice recordings, prompts, metadata and outputs, with already-incorporated data potentially retained after withdrawal.
Studio workflows can also involve third parties. Official transcription documentation says desktop transcription uses Microsoft Azure, online transcription may use Microsoft or Google, and translation sends text to Amazon. The desktop Professional/Enterprise path advertises unlimited in-house transcription, but the exact routing for each operation should be verified.
Create a data-flow table before using unreleased dialogue, personal speech or regulated data:
| Operation | Potential processor named in documentation | Evidence to obtain |
|---|
| Speech-to-speech morph | Altered | Region, retention, improvement use, encryption and deletion. |
| Online transcription | Microsoft or Google | Selected provider, account terms and file retention. |
| Desktop transcription | Altered in-house | Whether audio stays local and what telemetry leaves. |
| Translation | Amazon | Region, text retention and provider terms. |
| Rapid Clone | Altered | Source-file record, clone lifecycle and deletion. |
The public privacy policy uses general necessity-based retention language but does not provide a complete project/audio/backup schedule for this workflow. Keep that unknown open until the controlling agreement answers it.
Is Altered Studio an automatic dubbing platform?
Not on the evidence reviewed. The editor can import video, transcribe speech, translate text, morph audio cross-lingually and edit the audio stream. That is a useful localization toolkit, but it is not the same contract as an automated end-to-end system that identifies speakers, translates, generates every voice, lip-syncs and delivers a mastered video.
If localization is the job, test the entire chain: transcription accuracy, translation review, actor performance in the target language, timing changes, speaker consistency, mix, captions and video export. Budget human linguistic and audio review. Do not infer lip sync or one-click project automation from the presence of translation and video import.
What does Altered Studio cost?
The current public pricing response reviewed did not expose a reliable Studio price/allowance table, so this profile does not invent one. The product describes Creator, Professional and Enterprise tiers, trials and different voice/model access, but the buyer must preserve the actual checkout.
Text-to-speech and translation use Text Tokens, generally one per character. Higher-quality third-party voices can consume more tokens per character, with the multiplier shown on voice cards. Speech-to-speech uses the first-morph selection accounting described earlier. Cost therefore depends on both plan access and the chosen workflow.
At evaluation time, capture:
- billing cadence, tax and renewal amount;
- morph minutes or quota and reset behavior;
- Text Tokens and multiplier for every shortlisted TTS voice;
- included Professional/Common voice counts and model families;
- Rapid Custom Voice capacity;
- desktop entitlement and hardware requirements;
- API, team and support entitlements;
- overage or upgrade path.
Do not publish structured Offer data until these amounts can be represented without simplifying different plans or units into one misleading price.
Altered fits performance-led post-production
Shortlist it when a human performance is the valuable input and the team wants to audition different identities without re-recording every target voice. It is particularly relevant for character dialogue, pickups, prototyping and post-production teams willing to manage settings, samples and rights.
Look elsewhere for bulk narration with transparent character pricing, a lightweight creator interface, or an end-to-end dubbing/lip-sync workflow. Also pause if sensitive voice recordings cannot be used under the content/output improvement licence or third-party routing.
Review the AI voice software category, use the voice cloning consent guide, and follow the AI voice API buying guide when automation rather than desktop post-production is required.
Benchmark a scene, not a demo line
- Preserve the exact plan, licence, quota, voice/model access and renewal terms.
- Obtain written speaker authority for every source voice and intended use.
- Record matched clean scenes with dialogue, whisper, intensity, names and code-switching.
- Run the same scenes through Timbre, Performance, Flexi, Clone, Narration and Fast candidates.
- Generate several samples for non-deterministic models and blind the review.
- Count source retakes, samples, transcript fixes, setting changes and editor minutes.
- Reconcile first-morph quota and Text Token balances after every action.
- Save/reopen the project, load presets, export at required sample rate and recreate a pickup.
- Test Rapid Clone preview, consent evidence, library retention and subscription exit.
- Map Altered, Microsoft, Google and Amazon processing for each enabled feature.
- Resolve improvement use, retention/deletion, security assurance and team governance.
- Compare accepted minutes and total labor with actor re-recording and competing workflows.
Pass only when blind reviewers prefer or accept the finished scenes, correction and selection effort is predictable, quota maps to delivered work, the licence covers the organization and channels, and voice/data handling passes procurement. Fail or extend the trial when convincing samples cannot be reproduced, rights are ambiguous, or data routes are unacceptable.
Test whether the performance survives the morph
Use Altered's trial on a short scene whose acting already works. Compare the original with multiple morph models and count the samples, corrections and editor time needed to keep timing and intent. Move to a paid licence only when the transformed performance is repeatable and the selected licence covers the organisation's revenue and distribution. If the source acting is not the valuable input, a conventional TTS or dubbing workflow is likely the cleaner purchase.
Official sources