Skip to main content
Features Install Cheatsheet Comparison Documentation About GitHub
Open Source Heritage

Driven by Community.
Powered by AI.

PaddleOCR is more than just an OCR engine — it is a global collaborative ecosystem created by Baidu's PaddlePaddle team to make the world's printed knowledge accessible to machines and LLMs alike.

Mission & Values

Our Mission

We believe that high-accuracy, production-ready document parsing and text recognition should be a universal public good. Our mission is to provide an industry-leading, 100% free, and private OCR solution that runs seamlessly on CPUs, GPUs, NPUs, and edge devices.

🛡️
100% Offline & Private No data ever leaves your machine. PaddleOCR operates strictly offline by default.
🌐
Universal Multilingual Access Native recognition models for 100+ languages and dozens of scripts.

Baidu PaddlePaddle Origins

Originally developed at Baidu's Deep Learning Institute, PaddleOCR was built on the open-source PaddlePaddle framework. Today, it stands as the #1 most starred OCR repository on GitHub with over 86,000 stars and 6,000+ dependent repositories worldwide.

Milestones

The PaddleOCR Timeline

From a research toolkit to the world's most popular open-source document AI engine.

2020 — 2022

Baidu PaddlePaddle Launch

PaddleOCR was open-sourced by Baidu. PP-OCRv1 and v2 introduced ultra-lightweight mobile models under 10MB, quickly reaching 20,000+ GitHub stars.

2023 — 2025

Global Multilingual Expansion

Expanded to 100+ native languages. Released PP-StructureV2 for complex table and layout parsing, becoming the default engine in RAG and AI agent frameworks.

2026 — Present

Vision-Language & PP-OCRv6 Era

Released PP-OCRv6 and PaddleOCR-VL. OmniDocBench v1.6 precision reached 96.33%, surpassing 86,000+ stars as the global standard for document AI.

Why "PaddleOCR"?

The name stems from Baidu's PaddlePaddle (PArallel Distributed Deep LEarning) framework. PaddleOCR combines lightweight model compression with end-to-end parallel inference, allowing high-speed document processing on low-spec CPU machines without needing expensive GPU infrastructure.

Join the Community

PaddleOCR is maintained by Baidu engineers and thousands of open-source contributors worldwide. Whether you want to fix bugs, contribute multilingual training data, or build integrations, everyone is welcome.

Get Involved on GitHub →