• Stefy Lanza (nextime / spora )'s avatar
    audio: STT backends, diarization, speaker embeddings/ID, word timestamps · 7e5a925c
    Stefy Lanza (nextime / spora ) authored
    Transcription (/v1/audio/transcriptions):
    - Route multipart uploads by peeked model (front _peek_model) so they land on
      the model's assigned engine instead of scattering by capability.
    - New pluggable STT backends (codai/api/stt_backends.py): wav2vec2 (HF), vosk
      (CPU dirs), NVIDIA NeMo-Canary (isolated venv worker) + faster-whisper/whisper.
    - Per-model language selection + validation; Canary source->target translation.
    - OpenAI timestamp_granularities[]=word: top-level words[] in verbose_json
      (vosk/wav2vec2/faster-whisper/canary).
    - VRAM-eviction tracking (acquire_stt_backend) + keep_resident co-residency for
      the small set; /v1/models surfaces languages + backends.
    
    Diarization (/v1/audio/diarization + diarize=true):
    - pyannote.audio isolated venv worker (codai/api/pyannote_worker.py,
      tools/pyannote_service.py) on the 3090; speaker-labeled segments; graceful
      degradation. Pinned torch/pyannote/hf_hub/matplotlib in requirements-pyannote.txt.
    
    Speaker embeddings + recognition:
    - /v1/audio/speaker-embeddings (ecapa via speechbrain, pyannote/wespeaker).
    - Enroll/list/delete + identify/verify (codai/api/speaker_registry.py); persistent
      voiceprint store; diarization identify=true relabels turns with enrolled names.
    
    New isolated-venv requirements: requirements-nemo.txt, requirements-pyannote.txt
    (+ vosk in requirements.txt).
    Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_01R6iugoNeEgt9PkajKrFyq4
    7e5a925c