• Stefy Lanza (nextime / spora )'s avatar
    audio: CrisperWhisper support (isolated venv, verbatim + word timestamps) · b069137b
    Stefy Lanza (nextime / spora ) authored
    CrisperWhisper does verbatim ASR with precise word timestamps, but its custom
    generation config crashes on the server's transformers 5.x. Run it in its own
    venv worker (pinned torch 2.5.1+cu124 + transformers 4.46.3 — last pre-5.x with
    working Whisper word timestamps and a py3.13 tokenizers wheel):
    - codai/api/crisperwhisper_worker.py + tools/crisperwhisper_service.py
      (return_timestamps="word" + CrisperWhisper pause redistribution).
    - requirements-crisperwhisper.txt.
    - stt_backends.py: `crisperwhisper` is its own family -> _RemoteCrisperWhisperBackend
      (HTTP to the worker); generic `whisper-hf` still runs in the shared venv
      (text-only default + graceful timestamp fallback); HF-ASR path now converts any
      container (mp4/m4a/webm) to 16k mono wav first and uses model-correct
      return_timestamps (Whisper True/"word", CTC "chunk"/"word").
    - transcriptions.py: crisperwhisper dispatched as a GPU worker.
    - manager skip-download + admin backend option for whisper-hf/crisperwhisper.
    
    Verified: mp4 -> verbatim text + per-word timestamps, HTTP 200, on the 3090.
    Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_01R6iugoNeEgt9PkajKrFyq4
    b069137b