Gen+ stays stable through long-form generation on the clean sample, with no voice breaks or instability observed.
What was measured
Long-Form Consistency
Whether voice quality, pacing, and pronunciation stay consistent over longer passages instead of degrading after a few sentences.
decisive for this rankingtransformation
For text-to-voiceover work, the voice must stay stable across longer scripts; degradation means the output is not reliable. (3 of 3 judges)
What was given, what came back
Test input: High-Quality Voice Sample · mixed · group: voice-cloning
Input — what we sent
Input, verbatim
Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.
0:00 / 0:00
Loading audio...
Research media topmediai highquality input.wav
0:00 / 0:00
Loading audio...
High-Quality Voice Sample
A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.
Why this input is hard
- · Maximum voice-cloning accuracy
- · Naturalness with optimal source quality
- · Long-form consistency
- · Pronunciation stability
- · Voice preservation under ideal conditions
Output — unretouched
0:00 / 0:00
Loading audio...
Also checked on this input — same tool, 9 other criteria
Naturalness & Human Quality⚠ StruggledGen still sounds robotic on the clean source sample, making it less convincing than HD.Naturalness & Human Quality✓ WorkedHD is the most human-like and realistic output on the clean sample, with better emotional delivery and conversational flow.Naturalness & Human Quality◐ MixedGen+ sounds more natural than Gen, but the weak identity preservation keeps the output from feeling fully convincing.Pronunciation Accuracy✓ WorkedHD has no misread or garbled words on the clean sample and is the cleanest pronunciation result across all nine English generations tested.Pronunciation Accuracy✓ WorkedGen keeps the script intelligible on the clean sample; the robotic character is a delivery issue, not misread words.Pronunciation Accuracy✓ WorkedGen+ keeps the script intelligible on the clean sample despite the gender shift, so pronunciation remains intact.Voice Match Accuracy✓ WorkedHD has the highest similarity to the clean source voice and preserves speaker identity best among the high-quality outputs.Voice Match Accuracy◐ MixedGen keeps some similarity to the clean source voice, but speaker identity remains only moderately accurate and still sounds robotic.Voice Match Accuracy✗ FailedGen+ reduces similarity to the clean source voice by shifting toward a feminine tone, so speaker representation is inaccurate.
Provenance
- Observation
- 0e63740f-6710-4e59-896e-d26271ddc32f
- Evidence run
- 46222c41-0046-41cc-bfaa-5f7ba6aa4933
- Study
- Clone Your Voice and Generate Voiceover from Text
- Research task
- 86ba42bx1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "topmediai-voice-cloning",
scenario: "voice-cloning"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on Long-Form Consistency
ElevenLabs✓ WorkedKeeps voice stability across extended scripts and is described as good for long-form narration, with no reported degradation over longer passages.Heygen⚠ StruggledThe longer-script high-quality generation had flow interruptions, word mispronunciations, inconsistent delivery, and the report says multiple regenerations may be required for production-ready results.Inworld⚠ StruggledIn the ~1:04 high-quality output, matching accuracy degraded further as the clip progressed, so consistency weakened over the longer passage.MiniMax✓ WorkedThe complete ~19-second output held up fine at that length, with no visible degradation reported within the generated clip.Uberduck✓ WorkedThe output stayed poor throughout the passage rather than degrading over length, so no later-stage dropoff was observed.VocalAI✓ WorkedConsistency was maintained throughout the script, though the speech still ran faster than expected.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com