-
Stefy Lanza (nextime / spora ) authored
Everything OpenVoice does that we didn't, built from models we already run rather than adopting theirs — plus MeloTTS itself so their exact V2 pairing is available. Engines on /v1/audio/clone (`engine`, default `auto`): f5 today's in-context clone — timbre AND prosody, needs the transcript xtts XTTS-v2 zero-shot from the clip, 17 languages, no transcript chain base TTS -> seed-vc timbre transfer: OpenVoice's architecture with seed-vc in place of its tone-colour converter melotts the same chain with MeloTTS as the base voice `auto` picks a transcript-free engine when the profile has no transcript or the target language differs from the reference — the case where F5 carries the wrong accent across. GET /v1/audio/clone/engines reports what is installed. The transcript is no longer a hard requirement: it is now checked per engine instead of rejecting the request outright. Multi-clip profiles: PATCH /v1/audio/voices/{name} takes add_clips/remove_clips, and a clone prompts on the cleanest take. Scored on clipping, level and crest factor — explicitly NOT loudest-wins, since a clipped take has the highest RMS of the set and is the worst possible prompt. Single-clip profiles keep working. Watermarking with AudioSeal (MIT), on by default, per-request or global override, detection at POST /v1/audio/watermark/detect. Missing audioseal degrades to unmarked audio with one warning, never an error. MeloTTS runs in an isolated venv worker like parler-tts — its pins conflict with this stack the same way. _RemoteParlerBackend generalised to _RemoteSpeechBackend since both speak the same POST /speak contract. Co-Authored-By:Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0182DhR1eFNPQyGmDrpHbedY
d4218c35