-
Stefy Lanza (nextime / spora ) authored
CrisperWhisper does verbatim ASR with precise word timestamps, but its custom generation config crashes on the server's transformers 5.x. Run it in its own venv worker (pinned torch 2.5.1+cu124 + transformers 4.46.3 — last pre-5.x with working Whisper word timestamps and a py3.13 tokenizers wheel): - codai/api/crisperwhisper_worker.py + tools/crisperwhisper_service.py (return_timestamps="word" + CrisperWhisper pause redistribution). - requirements-crisperwhisper.txt. - stt_backends.py: `crisperwhisper` is its own family -> _RemoteCrisperWhisperBackend (HTTP to the worker); generic `whisper-hf` still runs in the shared venv (text-only default + graceful timestamp fallback); HF-ASR path now converts any container (mp4/m4a/webm) to 16k mono wav first and uses model-correct return_timestamps (Whisper True/"word", CTC "chunk"/"word"). - transcriptions.py: crisperwhisper dispatched as a GPU worker. - manager skip-download + admin backend option for whisper-hf/crisperwhisper. Verified: mp4 -> verbatim text + per-word timestamps, HTTP 200, on the 3090. Co-Authored-By:
Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R6iugoNeEgt9PkajKrFyq4
b069137b