JA EN

#ocr

1 articles

01 ·VLMs & Multimodal·★ MEMBER·PAPER·10 min read Document AI and OCR Today — How an LLM Ends Up Reading Your Invoices How machines came to read invoices and scanned PDFs, from first principles: text detection and CTC, how errors are measured (CER), LayoutLM's trick of embedding coordinates alongside words, the OCR-free Donut line, and today's habit of handing the page straight to a VLM — plus the walls that matter in production: tables, handwriting, and hallucination.