Quality stayed uniformly poor from start to finish rather than degrading later in the passage, so no long-form drift was observed.
What was measured
Long-Form Consistency
Whether voice quality, pacing, and pronunciation stay consistent over longer passages instead of degrading after a few sentences.
decisive for this rankingtransformation
For text-to-voiceover work, the voice must stay stable across longer scripts; degradation means the output is not reliable. (3 of 3 judges)
What was given, what came back
Test input: Low-Quality Voice Sample · mixed · group: voice-cloning
Input — what we sent
Input, verbatim
Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.
Research media low quality input recording.mp4
0:00 / 0:00
Loading audio...
Low-Quality Voice Sample
A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.
Why this input is hard
- · Cloning accuracy from degraded audio
- · Noise and ambience robustness
- · Speaker identity preservation under poor recording conditions
- · Distinguishing enhancement from true cloning
Output — unretouched
0:00 / 0:00
Loading audio...
Also checked on this input — same tool, 3 other criteria
Naturalness & Human Quality✗ FailedThe output sounded heavily robotic, with frequent unnatural pauses that made it immediately identifiable as AI-generated.Pronunciation Accuracy✓ WorkedStandard-script English words stayed intelligible; the observed problem was delivery/prosody, not misread or garbled words.Voice Match Accuracy✗ FailedThe noisy-source clone barely resembled the original speaker and was described as the weakest voice-match result in the round.
Provenance
- Observation
- aab42c38-9b8f-47e6-aea3-f731d0fffffb
- Evidence run
- 46222c41-0046-41cc-bfaa-5f7ba6aa4933
- Study
- Clone Your Voice and Generate Voiceover from Text
- Research task
- 86ba42bx1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "uberduck",
scenario: "voice-cloning"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on Long-Form Consistency
ElevenLabs✓ WorkedHolds voice quality and pronunciation stable across extended narration, with no major degradation reported during long-form generation.Heygen⚠ StruggledIn the longer-script generation, the voice stayed human-like but broke conversational flow, mispronounced certain words, and became inconsistent across longer passages.Inworld⚠ StruggledThe ~53-second low-quality output drifted further from the source voice by the end of the clip, making long-form consistency a clear weak point.MiniMax✓ WorkedOn the complete ~18-second output, pacing stayed steady and natural pauses held up through the clip.TopMediai Voice Cloning✓ WorkedHD stays consistent through long-form narration on the noisy sample, with no abrupt changes in voice quality or pronunciation.VocalAI✓ WorkedThe output stayed consistent throughout the generated script, with no major pronunciation issues, although the pacing was noticeably fast.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com