JA EN

#document-ai

2 articles

01 ·★ MEMBER·PAPER·9 min read Paper Walkthrough: The Design Fundamentals of Pixel Text Representation Learning An encoder that reads meaning straight off the pixels, never converting glyphs to character codes. This EMNLP 2026 paper argues that what decides its quality is not data volume but four design choices — explained from zero. 02 ·VLMs & Multimodal·★ MEMBER·PAPER·10 min read Document AI and OCR Today — How an LLM Ends Up Reading Your Invoices How machines came to read invoices and scanned PDFs, from first principles: text detection and CTC, how errors are measured (CER), LayoutLM's trick of embedding coordinates alongside words, the OCR-free Donut line, and today's habit of handing the page straight to a VLM — plus the walls that matter in production: tables, handwriting, and hallucination.