Driven by Community.
Powered by AI.
PaddleOCR is more than just an OCR engine — it is a global collaborative ecosystem created by Baidu's PaddlePaddle team to make the world's printed knowledge accessible to machines and LLMs alike.
Our Mission
We believe that high-accuracy, production-ready document parsing and text recognition should be a universal public good. Our mission is to provide an industry-leading, 100% free, and private OCR solution that runs seamlessly on CPUs, GPUs, NPUs, and edge devices.
Baidu PaddlePaddle Origins
Originally developed at Baidu's Deep Learning Institute, PaddleOCR was built on the open-source PaddlePaddle framework. Today, it stands as the #1 most starred OCR repository on GitHub with over 86,000 stars and 6,000+ dependent repositories worldwide.
The PaddleOCR Timeline
From a research toolkit to the world's most popular open-source document AI engine.
Baidu PaddlePaddle Launch
PaddleOCR was open-sourced by Baidu. PP-OCRv1 and v2 introduced ultra-lightweight mobile models under 10MB, quickly reaching 20,000+ GitHub stars.
Global Multilingual Expansion
Expanded to 100+ native languages. Released PP-StructureV2 for complex table and layout parsing, becoming the default engine in RAG and AI agent frameworks.
Vision-Language & PP-OCRv6 Era
Released PP-OCRv6 and PaddleOCR-VL. OmniDocBench v1.6 precision reached 96.33%, surpassing 86,000+ stars as the global standard for document AI.
Why "PaddleOCR"?
The name stems from Baidu's PaddlePaddle (PArallel Distributed Deep LEarning) framework. PaddleOCR combines lightweight model compression with end-to-end parallel inference, allowing high-speed document processing on low-spec CPU machines without needing expensive GPU infrastructure.
Join the Community
PaddleOCR is maintained by Baidu engineers and thousands of open-source contributors worldwide. Whether you want to fix bugs, contribute multilingual training data, or build integrations, everyone is welcome.
Get Involved on GitHub →