Model Comparisons

Top Voice Cloning APIs of 2026: Speaker Fidelity, Consent Verification, and Cost per Million Characters

A hands-on 2026 comparison of seven voice cloning APIs tested with identical 10-second speaker samples. The guide covers reference audio requirements, consent verification methods, language support, and pricing per million characters. Fish Audio s2-pro led Hume's speaker similarity leaderboard at 4.03, ahead of Cartesia sonic-3.5 (3.70) and ElevenLabs Multilingual v2 (3.68), while ElevenLabs enforces the strictest consent checks via Voice Captcha. Inworld offered the lowest pricing at $12.50–$25 per million characters.

If you're evaluating a voice cloning API for a product, this guide is for you. A generic synthetic voice only has to sound pleasant; a cloned voice has to sound like one specific individual. That is a much tougher bar, and the seven providers tested here approach it in very different ways.

How we ran the 10-second test

We recorded 10 seconds of a single speaker in a quiet room, saved as a WAV file, and used that same clip on every platform. Each service read the same two sentences: one conversational, one loaded with numbers and a proper name. We used each vendor's instant cloning path with default settings.

Two vendors bent the rules slightly. Hume documents 15 seconds as its minimum, so its clone used the same speaker's clip extended to 15 seconds. ElevenLabs recommends 1 to 2 minutes for Instant Voice Cloning, meaning its 10-second clone ran below the vendor's own guidance.

Comparison at a glance

| Provider | Reference audio (instant / pro) | Consent verification | Commercial license from | Languages | List price per 1M chars |
|---|---|---|---|---|---|
| ElevenLabs | 1–2 min recommended / 30 min minimum | Rights attestation for IVC; Voice Captcha for PVC | Starter, $6/mo | 70+ (Eleven v3) | $50 (Flash) to $100 (v3) |
| Cartesia | 10 s, up to 60 s on Sonic 3.6+ / 30 min | Permission required by terms | Pro, $5/mo | 44 | About $37 to $50 |
| Inworld | 3–30 s / 10 min minimum (beta, English only) | Rights confirmation | Free On-Demand tier | 200+ languages and locales | $12.50 to $25 (TTS-2) |
| Gradium | 10 s / 30 min, 2 h recommended | Owner consent required by policy | XS, $13/mo | 5 | About $36 to $58 |
| Fish Audio | About 10 s / 10 to 180 min | Live ownership check for pro clones | Plus, $15/mo | 83 | $15 per 1M UTF-8 bytes |
| Resemble AI | 10 s / 10 to 25+ min | Verifiable consent for Professional Clone | Business plan for cloning API | 23 | About $30 (estimate) |
| Hume | 15 s / not documented | Upload from a consenting speaker | Creator, $14/mo | 11 | $50 to $150 by plan |

Credit-plan rates equal the plan price divided by included characters. Resemble AI's pricing page lists no TTS rate; trackers report $0.0005 per second in mid-2026, converted here at roughly 1,000 characters per minute. Prices were checked September 20, 2026.

What public data says about speaker similarity

Hume published its Voice Replication Leaderboard on September 10, 2026. It evaluated 11 models across 25 reference voices with 7 prompts, and 3 blind raters scored each clip from 1 to 5 on how much it sounded like the original speaker.

Among the vendors covered here, Fish Audio s2-pro led at 4.03. Cartesia sonic-3.5 scored 3.70, with ElevenLabs Multilingual v2 close behind at 3.68. Cartesia sonic-3.6-beta came in at 3.63 and Inworld TTS-2 at 3.62. ElevenLabs Eleven v3 ranked last of all 11 models at 2.91.

The board also shows why a single number can mislead. Cartesia's sonic-3.6-beta topped naturalness at 4.36 yet placed 8th on identity, while Inworld TTS-2 led audio quality at 4.61. Full per-category results are available on the Hugging Face leaderboard.

Provider breakdown

ElevenLabs offers Instant Voice Cloning (IVC) and Professional Voice Cloning (PVC). IVC docs recommend 1 to 2 minutes of clean audio; PVC requires at least 30 minutes, ideally 2 to 3 hours, on the Creator plan or above. It also has the strictest gate in this group: the voice owner must read on-screen text aloud via Voice Captcha. API rates are $0.10 per 1K characters on Eleven v3 and Multilingual v2, while Flash, Turbo, and v3 Conversational cost $0.05. Eleven v3 supports 70+ languages.

Cartesia clones from 10 seconds, and Sonic 3.6 and newer can use up to 60 seconds to hold accents better. Pro Voice Clones train on 30+ minutes and start on the $49 Startup plan. Sonic bills 1 credit per character, with plans from $5 (100K credits) to $299 (8M), and supports 44 languages.

Inworld instant clones work from 3 seconds, and samples up to 30 seconds improve similarity. Cloning itself is free. Realtime TTS-2 lists $25 per 1M characters on demand, dropping to $15 at $300 a month and $12.50 at $1,500; TTS-2 Flash starts at $15 and falls to $7. Professional cloning is in beta, needs 10 minutes of audio, and is English only. Inworld covers 200+ languages and locales, with 15 in its top quality tier.

Gradium, founded by Kyutai co-founders, clones from 10 seconds. The free tier includes 5 clones for non-commercial use, and paid plans start at $13 for 225K credits billed at 1 credit per character. Pro Voice Clone needs a 30-minute minimum and starts on the $340 M plan. Language coverage is narrow: English, French, German, Spanish, and Portuguese.

Fish Audio bills its API at $15 per 1M UTF-8 bytes. English runs roughly 1 byte per character, but Chinese, Japanese, and Korean run about 3, which nearly triples the effective price. S2.1 Pro covers 83 languages. Instant clones work from about 10 seconds; professional clones take 10 to 180 minutes and require live ownership verification. Commercial use starts on the $15 Plus plan and covers verified voices you own.

Resemble AI Rapid Clone needs 10 seconds and finishes in under a minute. Professional Clone needs 10 to 25+ minutes and about 40 minutes of training, with explicit, verifiable consent from the voice talent. PerTh watermarking is applied to every output. Chatterbox Multilingual covers 23 languages. The catch: Resemble now leads with deepfake detection, its Voice Cloning API requires a Business plan or higher, and its pricing page no longer lists TTS rates.

Hume Octave clones from as little as 15 seconds, recorded live or uploaded from a consenting speaker. Octave 2 covers 11 languages. Commercial use starts on Creator at $14 a month, since Free and Starter tiers are non-commercial. API access to voice cloning is listed only on Enterprise; other tiers create clones within the platform. Extra characters cost $0.15 per 1K on Creator, down to $0.05 on Business.

Which API fits which job

  • Budget voice agents at scale: Inworld or Fish Audio
  • Low latency with broad language coverage: Cartesia
  • Brand voices with strict consent: ElevenLabs PVC
  • European languages with a free test tier: Gradium
  • Cloning plus watermarking and deepfake detection: Resemble AI
  • Emotion-directed delivery: Hume

Key takeaways

  • ElevenLabs is the category default, but its models trail on Hume's new same-speaker leaderboard.
  • Only ElevenLabs, Fish Audio, and Resemble AI document technical consent checks, and only for pro clones.
  • Inworld and Fish Audio list $25 or less per 1M characters; ElevenLabs v3 lists $100.
  • Hume needs 15 seconds of audio, so a 10-second clip falls short.
  • Language coverage ranges from 5 (Gradium) to 200+ languages and locales (Inworld).

Who should use it / who should skip it

Use ElevenLabs if consent rigor and ecosystem maturity matter most, especially for professional brand voices. Use Inworld or Fish Audio if you need volume pricing for agents. Use Cartesia for low latency across 44 languages. Skip Gradium if you need anything beyond five European languages; skip Hume if your reference clips are under 15 seconds; and skip Resemble AI if you want a self-serve low-volume plan, since cloning requires a Business-tier commitment. No single vendor wins every axis, so match the tool to your deployment priorities.

Meta description: Compare 7 voice cloning APIs in 2026 on speaker similarity, consent checks, pricing per 1M characters, and language coverage before you commit.

Tags: voice cloning, text to speech, ElevenLabs, API comparison, AI audio

Featured image: Abstract friendly illustration of seven stylized soundwave shapes in different colors merging into one waveform, no people or logos.

Comments (0)

  1. No comments yet. Be the first to share what worked for you.

Leave a comment

Comments are reviewed before they appear. Your email address is not published.