Hume Octave is an expressive speech model, but a buying decision must separate Octave 1 from Octave 2 Preview and the API from Hume's Playground/Creator Studio. These surfaces differ in languages, controls, maturity and data handling. A demo that sounds expressive does not establish cost, repeatability or governance.
> Distinctive strength: Octave is built around expressive, context-sensitive speech rather than only neutral narration, giving buyers a distinct emotional-performance test. > > Where it stops being an advantage: Octave versions and API/Studio surfaces differ in maturity, controls, languages and data handling.
Hume's claims about naturalness, emotion and responsiveness remain hypotheses until a paid, blind listening panel reproduces them. The evidence below is used to design that buyer-controlled test, not to award Octave a quality win from its demo.
Octave 1 or Octave 2: the decision in 60 seconds
| Question | Source-led answer |
|---|
| Best fit | Teams building expressive agents, characters or narration that can test model-version behavior and operate an API workflow. |
| Octave 1 | English/Spanish, acting descriptions and voice design; vendor-reported model latency around 200ms. |
| Octave 2 | Preview, 11 languages, timestamps and reported ~100ms latency, but incomplete control/design parity. |
| Cost unit | Subscription includes characters; higher self-serve tiers add $0.15–$0.05 per 1,000 overage characters. |
| Input boundary | 5,000 text characters and 1,000 acting-description characters per utterance. |
| Cloning gate | User attests rights/consent; independent speaker verification and full revocation procedure were not established. |
| Data trap | API data policy is stricter than Playground/consumer handling; do not treat them as interchangeable. |
| Main procurement gap | Public TTS SLA and numeric TTS retention were not established. |
Should you use Octave 1 or Octave 2?
The newer version is not automatically the safer production choice.
| Decision surface | Octave 1 | Octave 2 |
|---|
| Launch state | Established API version | Preview |
| Languages | English, Spanish | English plus Japanese, Korean, French, Portuguese, Italian, German, Russian, Hindi and Arabic |
| Acting description | Available | Coming soon in retained docs |
| Voice design | Must be created with v1 | Can use a voice designed with v1 |
| Speed / trailing silence | Supported | Supported |
| Word/phoneme timestamps | Not documented | Supported when requested |
| Claimed model latency | ~200ms | ~100ms, network excluded |
Octave 1 and 2 generations cannot be used interchangeably for continuation. Pin `version` in every test and production call. If omitted, Hume can route to what it considers the appropriate model—convenient for experimentation, weak for reproducibility.
Price expressive speech by accepted character
The current public plan table combines a monthly subscription, included characters, an overage rate and a requests-per-minute ceiling:
| Plan | Monthly price | Included characters | Overage per 1,000 | TTS RPM |
|---|
| Free | $0 | 10,000 | Not publicly listed | 15 |
| Starter | $3 | 30,000 | Not publicly listed | 15 |
| Creator | $14 | 140,000 | $0.15 | 75 |
| Pro | $70 | 1,000,000 | $0.12 | 75 |
| Scale | $200 | 3,300,000 | $0.10 | 150 |
| Business | $500 | 10,000,000 | $0.05 | 225 |
| Enterprise | Custom | Custom | Custom | Custom |
The site currently shows a first-month Creator promotion; do not use that temporary figure for recurring economics. It estimates roughly one minute per 1,000 characters, but actual duration depends on language, punctuation and delivery.
If the included allowance is exhausted, 100k/1M/5M characters cost $15/$150/$750 Creator, $12/$120/$600 Pro, $10/$100/$500 Scale and $5/$50/$250 Business. Add the subscription first and subtract remaining allowance before applying overage.
Rework changes the result. At Pro's rate, regenerating 25% of a one-million-character workload adds $30; 50% adds $60 after allowance. Asking for multiple generations can create variants faster but must be metered. Record every rejected take and corrected passage, not just delivered audio.
Is Hume Octave free, and can free output be commercial?
Free includes 10,000 monthly characters and 15 RPM. That is enough for an integration/listening fixture, not proof of production rights or throughput. The pricing page includes a “Commercial license” row, but the retained rendered table did not expose reliable positive/negative values per plan. Confirm the exact tier in checkout or contract before publishing client/monetized work.
Hume's TTS documentation states users retain full ownership of generated Octave audio, subject to its Terms and lawful rights in source/voice content. Ownership and license entitlement are different questions; verify both.
How expressive are the acting controls?
Octave 1 accepts natural-language acting descriptions for emotion, delivery and context. Both versions support speed from 0.5 to 2.0 on a non-linear scale and trailing silence. Hume says that without instructions, Octave infers delivery from the voice description and text.
That flexibility can reduce deterministic repeatability. Test the same line many times, score whether required emphasis/pronunciation stays inside tolerance, and preserve the exact text, description, voice ID, version and context. Do not turn Hume's “industry-leading” language into a BenPicks quality claim.
For long-form work, continuation can chain adjacent utterances or reference a prior generation/context. It only follows immediate context, adds latency because context must be processed, and cannot bridge v1/v2 output. Test scene boundaries, speaker changes, homographs and re-generation halfway through a chapter—not just a clean first paragraph.
How do streaming, timestamps and formats work?
Hume offers:
- HTTP output streaming as JSON or file;
- WebSocket incremental text input with continuous audio output;
- non-streaming JSON/file responses;
- MP3, WAV and PCM synthesis outputs.
Instant mode is streaming-only, requires a selected voice and exactly one generation. Dynamic voice creation and multiple candidates require it to be disabled. Octave 2 can interleave word- or phoneme-level timestamps with audio, useful for captions, avatars and editing.
Measure time to first audio separately from complete-response time. Hume's ~100ms/~200ms figures exclude network transit and are vendor claims. Test at your plan's RPM, from the target region, with real payload lengths and simultaneous sessions.
What can voice design and cloning do?
Hume advertises more than 100 library voices. Octave 1 can design a reusable voice from a natural-language identity/style prompt and sample line; that voice can later be selected in Octave 2. Cloning can use guided microphone recording or uploaded speech, with Hume advertising a minimum around 15 seconds.
The upload workflow requires the user to confirm necessary rights/consent. That is a meaningful gate, but retained public docs did not establish independent identity/liveness verification. Nor did they establish a full speaker-led process for disputing a clone, withdrawing consent, revoking every cached key/session and certifying model/backup deletion.
Treat voice deletion in the UI/API as a feature to test, not proof of the whole lifecycle. Use a consenting internal speaker, delete the voice, then test old IDs, open sessions, Creator Studio assets and account export. Require the real speaker's consent record outside Hume.
What is voice conversion, and where can it fail?
Voice conversion applies a library/custom identity while preserving original timing and emotional delivery. Inputs include MP3, WAV, M4A and OGG, at least 12 seconds and under three minutes, with 44.1kHz recommended.
Test whether source identity leaks, noise/artifacts transfer, timing stays aligned and the target speaker remains recognizable across accent/emotion changes. Consent is required for both the source recording and cloned target. A technically successful conversion is not sufficient if either party lacks authorization.
Does Hume train on your data?
Hume's API data policy says customer-submitted API data is not used to train or improve its models. Consumer products such as Playground are different: the privacy documentation says content may be used to improve services unless opted out, can be stored in the US and elsewhere, and may be viewed by limited authorized personnel for abuse, support or legal purposes.
That distinction should change the test plan. Do not paste sensitive scripts or talent audio into the web Playground merely because production will use the API. Retained public sources did not give a numeric TTS API retention/deletion schedule for text, audio, metadata, logs and backups. Obtain one through a DPA/enterprise response if it matters.
The published subprocessor list includes Google Cloud, Datadog, Deepgram, SambaNova and operational/support providers. Confirm which apply to the exact TTS path and whether the list/notification terms meet policy.
What security and reliability commitments exist?
Server integrations use an API key header. For client TTS, a backend can issue a 30-minute access token so the permanent key is not exposed. Rotate keys and test invalidation before rollout; separate personal and organization credentials.
Pricing lists SOC 2 Type II, GDPR and HIPAA for Enterprise, and the privacy page directs buyers to request a DPA/BAA. This is not proof that every self-serve surface/workflow is in scope. Review the actual report, exceptions, BAA and architecture.
No public TTS uptime target, credit schedule, response-time target or disaster-recovery commitment was established in retained sources. Enterprise lists Slack support; lower tiers list Discord. A latency claim is not an SLA. Production buyers should request uptime, support severity targets, incident notification, RTO/RPO and status-history commitments.
When expressive acting control matters enough to shortlist Hume
Shortlist Octave when expressive direction and custom voice identity are more important than a broad GA-only language catalog, your team can benchmark Preview separately, and API operation plus human listening review are already planned.
Look elsewhere when every feature must be GA, a public SLA/numeric retention rule is mandatory, you need independently verified clone consent/revocation, or your required language/control exists on only one incompatible version.
Browse the AI Voice category, compare model economics in the AI Voice API buying guide, and use the TTS API benchmark guide for a neutral fixture.
A blind expression-and-continuation test for Octave
- Pin Octave version, voice ID, endpoint, output format and instant-mode setting.
- Use consented scripts with names, numbers, homographs, emotion shifts and each target language.
- Blind native listeners for pronunciation, meaning, identity, expression and fatigue.
- Generate repeat takes to quantify variability; record every rejected/corrected character.
- Measure first audio and complete response at low, expected and plan-RPM load.
- Test 5,000-character segmentation, continuation, speaker changes and cross-version rejection.
- Verify MP3/WAV/PCM and requested word/phoneme timestamps in the downstream stack.
- Test WebSocket interruption/reconnect and instant-mode restrictions.
- For cloning/conversion, record consent, then test deletion, old IDs and open sessions.
- Calculate subscription, allowance, overage, +25%/+50% rework and multiple generations.
- Compare API and Playground data settings; require retention, subprocessors and SLA in writing.
- Have a second reviewer reproduce the chosen configuration from preserved requests.
Official sources checked
Sources checked 2026-08-28. Octave 2 remains a preview surface; pin the version used in every result.
Pass only when listeners prefer or accept the output for the actual job, repeat runs stay within tolerance, version/transport survives peak load, cost remains predictable under correction, and consent/data/deletion terms are enforceable. Expressiveness in a demo is the beginning of the evaluation, not its conclusion.