-
Stefy Lanza (nextime / spora ) authored
Transcription (/v1/audio/transcriptions): - Route multipart uploads by peeked model (front _peek_model) so they land on the model's assigned engine instead of scattering by capability. - New pluggable STT backends (codai/api/stt_backends.py): wav2vec2 (HF), vosk (CPU dirs), NVIDIA NeMo-Canary (isolated venv worker) + faster-whisper/whisper. - Per-model language selection + validation; Canary source->target translation. - OpenAI timestamp_granularities[]=word: top-level words[] in verbose_json (vosk/wav2vec2/faster-whisper/canary). - VRAM-eviction tracking (acquire_stt_backend) + keep_resident co-residency for the small set; /v1/models surfaces languages + backends. Diarization (/v1/audio/diarization + diarize=true): - pyannote.audio isolated venv worker (codai/api/pyannote_worker.py, tools/pyannote_service.py) on the 3090; speaker-labeled segments; graceful degradation. Pinned torch/pyannote/hf_hub/matplotlib in requirements-pyannote.txt. Speaker embeddings + recognition: - /v1/audio/speaker-embeddings (ecapa via speechbrain, pyannote/wespeaker). - Enroll/list/delete + identify/verify (codai/api/speaker_registry.py); persistent voiceprint store; diarization identify=true relabels turns with enrolled names. New isolated-venv requirements: requirements-nemo.txt, requirements-pyannote.txt (+ vosk in requirements.txt). Co-Authored-By:
Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R6iugoNeEgt9PkajKrFyq4
7e5a925c