The Same Ten-Second Clip Does Not Measure the Same Capability
MarkTechpost used the same ten-second WAV recording of one speaker in a quiet room as the reference for ElevenLabs, Cartesia, Inworld, Gradium, Fish Audio, Resemble AI, and Hume. Each provider ran its default instant-cloning path and read the same two sentences: one conversational, the other containing numbers and a proper name. The outputs were first shown side by side without labels, allowing listeners to judge them before seeing the provider names. This gets at the core difference between ordinary text-to-speech and voice cloning: a stock voice mainly needs to sound good, while a clone needs to sound like the same person.
The comparison is not perfectly symmetrical, however. Hume sets a fifteen-second floor for its reference audio, so its sample was extended from the same speaker. ElevenLabs recommends one to two minutes of clean audio for Instant Voice Cloning, meaning its ten-second result falls below vendor guidance. The test is therefore useful for comparing low-friction instant paths, but it should not be read as a ranking of the seven providers’ professional cloning capabilities.
Instant and Professional Cloning Make Different Bets
In the material, instant cloning means conditioning an existing model with reference audio at request time. Its advantages are a small data requirement and short turnaround, which make it suitable for interactive products. Its identity performance is also more exposed to the quality and duration of the reference clip and to the model’s ability to generalize. Professional cloning fine-tunes a model on speaker data, usually requiring much more audio in exchange for more stable identity preservation and a more explicit training process.
The providers’ stated thresholds make this distinction concrete. ElevenLabs recommends one to two minutes for its instant path, while its professional clone requires at least thirty minutes and ideally two to three hours. Other professional paths generally require ten to thirty minutes. Gradium lists thirty minutes and recommends two hours, while Inworld’s tested path is marked beta, English-only, and requires at least ten minutes. A technical leader should not ask only whether a vendor has a voice cloning API. The first question is whether the product needs a one-off character audition, low-latency personalization, or a production-grade replica of a specific speaker.
The Most Natural Voice May Not Be the Right Speaker
The public data also breaks the intuition that the most pleasant voice must be the most similar one. Hume’s Voice Replication Leaderboard, published on September 10, 2026, tested eleven models across twenty-five reference voices and seven prompts. Three blind raters scored each clip from one to five for resemblance to the reference. Among the providers covered here, Fish Audio s2-pro led on identity similarity at 4.03, followed by Cartesia sonic-3.5 at 3.70 and ElevenLabs Multilingual v2 at 3.68.
The category breakdown is more useful for procurement. Cartesia sonic-3.6-beta led naturalness at 4.36 but ranked eighth on identity. Inworld TTS-2 led audio quality at 4.61, while ElevenLabs Eleven v3 ranked last among the eleven models on identity similarity at 2.91. This does not mean that one model is universally worse. It means that naturalness, audio quality, and speaker identity are separate properties. Reducing them to a single listening impression can lead a team to choose a system that sounds polished but does not convincingly sound like the intended speaker.
Consent Is Part of the Product Surface
The comparison places reference-audio requirements, consent verification, commercial licensing, and language coverage in the same table. That reveals another shift: authorization is now part of whether a voice API can enter production. ElevenLabs requires a rights attestation for instant cloning and uses Voice Captcha for professional cloning, where the voice owner reads on-screen text. Cartesia requires permission under its terms, Inworld requires rights confirmation, Gradium’s policy requires owner consent, Fish Audio includes a live ownership check for professional clones, and Resemble AI requires verifiable consent for a Professional Clone. Hume describes uploads from a consenting speaker.
These controls are not equivalent, and the material also shows that many providers still rely mainly on declarations or contractual terms rather than universal technical verification. An enterprise integration therefore cannot treat the word “consent” on a vendor page as the end of the risk assessment. It should define who may upload and approve a voice, which languages and commercial uses are covered, how a voice can be withdrawn, and how long the audit trail is retained. For brand voices, customer-service identities, or public-figure-related use cases, stronger cloning makes the authorization chain more important, not less.
A Price Table Cannot Replace a Deployment Model
Published list prices appear to offer a straightforward answer, but there is no single cheapest provider in practice. The self-serve plans in the material range from five to fifteen dollars per month, while language coverage ranges from five languages to more than two hundred languages and locales. Fish Audio charges fifteen dollars per million UTF-8 bytes. Inworld TTS-2 is listed at roughly $12.50 to $25 per million characters, Gradium at about $36 to $58, and ElevenLabs at roughly $50 to $100 depending on the model and plan. Resemble AI’s pricing page does not list a TTS rate; the material uses a mid-2026 third-party estimate of about $0.0005 per second, converted at approximately 1,000 characters per minute, so it is not directly comparable with explicit vendor rates.
More importantly, unit price covers synthesis, not the full cost of obtaining a usable voice. Professional cloning may require a higher plan, more training audio, and a stricter authorization process. Multilingual products are also affected by the choice between byte-based, character-based, and time-based billing. A better procurement process is to score identity fidelity, naturalness, consent strength, language coverage, and unit economics, then calculate total cost at the expected monthly volume. Price and demo quality should be inputs to that decision, not substitutes for it.