ocr(surya): parse Surya2 (>=0.20) result.blocks[].html + reading order
Surya2 returns PageOCRResult.blocks[] (html/polygon/confidence/label/reading_order), not .text_lines like classic surya. Worker now extracts text from block HTML (tag-stripped) ordered by reading_order, with polygon->bbox. Verified end-to-end: latest Surya2 served via the vLLM backend transcribes a scanned IT legal page with all codici fiscali correct. Co-Authored-By:Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mw2KQiswmD69T45fTfjKwW
Showing
Please
register
or
sign in
to comment