Language / 言語
Home Features Install Documentation About GitHub

📋 PaddleOCR at a Glance

  • PaddleOCR is the world's most accurate open-source OCR toolkit, achieving 96.3% accuracy on OmniDocBench v1.6 with its PaddleOCR-VL-1.6 model.
  • Install with pip install paddlepaddle paddleocr — works on Windows, macOS, Linux, CPU and GPU.
  • Three core pipelines: PP-OCRv6 (fast text extraction), PP-StructureV3 (PDF/table parsing → Markdown/JSON), and PaddleOCR-VL (vision-language document AI).
  • Supports 100+ languages with a single pip install; PP-OCRv6 unifies 50 languages (including Chinese, Japanese, Arabic) in one model — no model switching needed.
  • Natively integrated with Dify, RAGFlow, Pathway, Cherry Studio; supports MCP server mode for direct LLM agent tool use.

Getting Started — Overview

PaddleOCR is an open-source, production-grade OCR (Optical Character Recognition) toolkit and document AI engine developed by Baidu's PaddlePaddle team. PaddleOCR is designed to bridge the gap between raw image/PDF documents and structured, LLM-consumable data — making it the preferred document parser for AI pipelines in 2025–2026.

PaddleOCR provides a complete document intelligence stack organized into modular pipelines. Each pipeline is purpose-built for a specific document processing task:

  • PP-OCRv6 — Universal text detection and recognition (Scene OCR). Supports 50 languages in a single model with three model tiers for different hardware constraints.
  • PP-StructureV3 — Complete document layout analysis. Converts PDFs and complex images to structured Markdown or JSON with table cell coordinates, formula positions, and reading-order data.
  • PaddleOCR-VL — Vision-language model for document parsing. Uses PaddleOCR-VL-1.6 (0.9B parameters) to achieve 96.33% accuracy on OmniDocBench v1.6 — SOTA for open-source document AI.
  • PP-ChatOCRv4 — Conversational document understanding. Chat with your PDFs and images using natural language queries backed by PaddleOCR's extraction pipeline.
  • HPD-Parsing — High-throughput document parsing with hierarchical parallel decoding. Achieves 4,752 tokens/second peak throughput for large-scale document processing.
PaddleOCR is trusted by 6,000+ open-source repositories and integrated into the world's most popular RAG and AI agent platforms including Dify, RAGFlow, Pathway, and Cherry Studio.

Installation

PaddleOCR is a Python package distributed via PyPI. It requires Python 3.8 through 3.12 and the PaddlePaddle deep learning framework. PaddleOCR runs on Windows, macOS, and Linux on CPU and GPU.

System Requirements

  • Python: 3.8 ≤ version ≤ 3.12 (Python 3.13+ not yet supported)
  • OS: Linux (Ubuntu 18.04+), macOS (10.15+), Windows 10/11
  • Hardware: CPU (any x86-64 or ARM64), NVIDIA GPU (CUDA 11.8+), Intel CPU with OpenVINO, Apple Silicon (M1/M2/M3/M4)
  • RAM: Minimum 4 GB (8 GB recommended for medium model)

Step-by-Step Installation

1

Install PaddlePaddle (CPU)

Install the CPU version of PaddlePaddle — the required deep learning framework:

Terminal
# Install CPU version (all platforms)
$ pip install paddlepaddle

# For NVIDIA GPU (CUDA 11.8+)
$ pip install paddlepaddle-gpu

# For NVIDIA GPU (CUDA 12.x)
$ pip install paddlepaddle-gpu==3.0.0.post120 -f https://www.paddlepaddle.org.cn/whl/linux/mkl/avx/stable.html
2

Install PaddleOCR

Install the PaddleOCR package via pip:

Terminal
$ pip install paddleocr

# Install latest dev version from source
$ git clone https://github.com/PaddlePaddle/PaddleOCR.git
$ cd PaddleOCR && pip install -e .
3

Verify Installation

Check that PaddleOCR is correctly installed:

Verify
$ python -c "import paddleocr; print('PaddleOCR', paddleocr.__version__, 'ready')"
# Expected output: PaddleOCR 3.7.0 ready

First Run Note: PaddleOCR automatically downloads model weights on first use (~100–300 MB). Models are cached in ~/.paddleocr/. Ensure internet access for the first run.

Install via Docker

Docker
# Pull the latest PaddleOCR Docker image
$ docker pull paddlecloud/paddleocr:latest

# Run interactively
$ docker run -it paddlecloud/paddleocr:latest bash

# Run with GPU support (requires nvidia-docker)
$ docker run --gpus all -it paddlecloud/paddleocr:gpu bash

Quick Start Guide

The fastest way to extract text from an image using PaddleOCR is the PaddleOCR class with the predict() method. PaddleOCR automatically handles model loading, text detection, and recognition in a single call.

Basic Text Extraction

basic_ocr.py
from paddleocr import PaddleOCR

# Initialize — downloads models on first run (~150 MB)
ocr = PaddleOCR(lang='en')

# Run OCR on an image
result = ocr.predict('image.jpg')

# Access extracted text
for page in result:
    for text in page['rec_texts']:
        print(text)  # Prints each text block

Get Text with Bounding Boxes

ocr_with_boxes.py
from paddleocr import PaddleOCR

ocr = PaddleOCR(lang='en')
result = ocr.predict('invoice.jpg')

for page in result:
    texts  = page['rec_texts']
    boxes  = page['rec_boxes']   # [[x1,y1],[x2,y2],[x3,y3],[x4,y4]]
    scores = page['rec_scores']  # confidence 0.0–1.0
    
    for text, box, score in zip(texts, boxes, scores):
        print(f'Text: {text}')
        print(f'Box: {box}')
        print(f'Confidence: {score:.3f}')

PDF to Markdown (PP-StructureV3)

pdf_to_markdown.py
from paddleocr import PPStructureV3

# Initialize PP-StructureV3 — full document parser
pipeline = PPStructureV3()

# Parse a PDF — automatically handles multi-page
results = pipeline.predict('document.pdf')

# Save each page as Markdown
for i, page in enumerate(results):
    md = page.get('markdown', '')
    with open(f'page_{i+1}.md', 'w', encoding='utf-8') as f:
        f.write(md)
    print(f'Page {i+1} saved')

PP-OCRv6 — Universal Scene OCR

PP-OCRv6 is PaddleOCR's flagship text detection and recognition pipeline, released with PaddleOCR v3.7.0 in June 2026. It achieves state-of-the-art accuracy on real-world scene text benchmarks, surpassing much larger vision-language models with only 34.5M parameters in the medium tier.

PP-OCRv6 Model Tiers

Tier Parameters Use Case Speed (A100) Accuracy
Tiny 1.5M Edge, IoT, Mobile Fastest Good
Small 7.7M Mobile, Low-power Fast Better
Medium 34.5M Server, Production 0.13s SOTA

PP-OCRv6 Key Improvements

  • +4.6% detection accuracy over PP-OCRv5_server on standard benchmarks
  • +5.1% recognition accuracy — surpassing Qwen3-VL-235B and GPT-5.5 with 34.5M parameters
  • 50 languages unified — Chinese, English, Japanese, and 46 Latin-script languages in one model
  • 5.2× CPU speedup via OpenVINO backend on Intel hardware
  • 6.1× Apple Silicon speedup on M4 (tiny model)
  • Specialized improvements for digital displays, dot-matrix characters, tire prints, and industrial text

PP-OCRv5 vs PP-OCRv6 Architecture & Upgrade Guide

PP-OCRv5 was the previous stable release of PaddleOCR (v3.5–v3.6). While PP-OCRv5 remains fully supported for backward compatibility, upgrading to PP-OCRv6 is strongly recommended for all new and existing deployments:

Dimension PP-OCRv5 (Legacy) PP-OCRv6 (Current SOTA)
Detection Accuracy Baseline (88.4%) +4.6% improvement (93.0%)
Recognition Accuracy Baseline (89.1%) +5.1% improvement (94.2%)
Language Models Separate model per language 50 languages unified in 1 model
Model Scaling Server / Mobile fixed 3 Tiers: Tiny (1.5M), Small (7.7M), Medium (34.5M)
CPU Optimization MKL-DNN baseline 5.2× OpenVINO CPU speedup

PP-OCRv6 Configuration Options

lang str default: 'en'
Language for recognition. Options include 'en', 'ch', 'japan', 'korean', 'arabic', 'latin', and 80+ more. Use 'ch' for a unified Chinese+English model.
use_gpu bool default: False
Enable GPU acceleration. Set to True when using paddlepaddle-gpu. Dramatically increases throughput for batch processing.
use_angle_cls bool default: False
Enable text angle classification. Set to True for documents with rotated or upside-down text (e.g., scanned documents at various orientations).
det_db_thresh float default: 0.3
Detection binarization threshold. Lower values detect more text regions (may include noise). Higher values are more selective. Range: 0.0–1.0.
det_db_box_thresh float default: 0.6
Minimum confidence threshold for text boxes. Boxes below this threshold are discarded. Increase to 0.8+ for high-precision mode.
rec_batch_num int default: 6
Number of text regions processed per recognition batch. Increase for GPU batching (e.g., 32–64) to maximize GPU utilization during batch processing.
Advertisement

PP-StructureV3 — Document Parsing Pipeline

PP-StructureV3 is PaddleOCR's document layout analysis and structured conversion pipeline. Unlike the base OCR pipeline which extracts flat text, PP-StructureV3 understands the full structure of a document — including reading order, table layout, formula positions, chart locations, and heading hierarchy — and outputs clean, semantically structured Markdown or JSON.

What PP-StructureV3 Can Parse

  • Multi-page PDF documents (native PDF and scanned PDFs)
  • Images with complex layouts (multi-column academic papers, reports)
  • Tables — extracts HTML, cell content, and row/column coordinates
  • Mathematical formulas — detected and preserved as LaTeX
  • Charts and figures — regions detected and described
  • Seals and stamps
  • Multi-page TIFF files (v3.6.0+)
  • Office documents: Word, Excel, PowerPoint → Markdown (v3.5.0+)

PP-StructureV3 Usage

ppstructure_v3.py
from paddleocr import PPStructureV3

pipeline = PPStructureV3(
    lang='en',       # Language for text recognition
    use_gpu=False,   # Set True for GPU
)

# Process a PDF or image
results = pipeline.predict('annual_report.pdf')

for page_idx, page in enumerate(results):
    print(f'=== Page {page_idx + 1} ===')
    
    # Get full structured Markdown
    print(page.get('markdown', ''))
    
    # Access individual content blocks
    for block in page.get('blocks', []):
        block_type = block['type']  # 'text', 'table', 'formula', 'figure'
        
        if block_type == 'table':
            print('HTML:', block['html'])
        elif block_type == 'formula':
            print('LaTeX:', block['latex'])

PP-StructureV3 Output JSON Schema

Output Structure (JSON)
# Each page in results is a dict with this structure:
{
  "page_id": 0,
  "markdown": "# Document Title\n\n...",
  "blocks": [
    {
      "type": "title",  # title|text|table|formula|figure|seal
      "bbox": [x1, y1, x2, y2],
      "text": "..."
    },
    {
      "type": "table",
      "bbox": [x1, y1, x2, y2],
      "html": "<table>...</table>",
      "cells": [...]  # Cell-level coordinates
    }
  ]
}

PaddleOCR-VL — Vision-Language Document AI

PaddleOCR-VL is a lightweight vision-language model (VLM) series specifically optimized for document parsing. Unlike the pipeline-based PP-StructureV3, PaddleOCR-VL processes an entire document page holistically — understanding layout, text, tables, and visual elements together as a vision-language model.

The latest version, PaddleOCR-VL-1.6 (0.9B parameters), achieves a 96.33% score on OmniDocBench v1.6 — the highest accuracy of any open-source document parsing model as of June 2026. It surpasses both commercial closed-source solutions and much larger models in text, formula, and table recognition tasks.

Key Capabilities

  • 96.33% accuracy on OmniDocBench v1.6 (current SOTA)
  • Table, ancient Chinese document, and rare character recognition
  • Seal recognition and text spotting across diverse formats
  • Chart and diagram understanding
  • Structured output in Markdown and JSON formats
  • Zero-cost migration from VL-1.5 — fully compatible model architecture

PaddleOCR-VL Usage

paddleocr_vl.py
# PaddleOCR-VL available on HuggingFace
from transformers import AutoProcessor, AutoModelForVision2Seq
from PIL import Image

# Load PaddleOCR-VL-1.6 from HuggingFace Hub
processor = AutoProcessor.from_pretrained('PaddlePaddle/PaddleOCR-VL-1.6')
model = AutoModelForVision2Seq.from_pretrained('PaddlePaddle/PaddleOCR-VL-1.6')

# Parse a document image
image = Image.open('complex_document.jpg')
prompt = "Parse this document and output as Markdown."

inputs = processor(
    text=prompt,
    images=image,
    return_tensors='pt'
)
outputs = model.generate(**inputs, max_new_tokens=2048)
result = processor.decode(outputs[0], skip_special_tokens=True)
print(result)

Advanced Usage & Performance Tuning

PaddleOCR provides several configuration options to maximize performance for production workloads. This section covers hardware-specific optimizations, batching strategies, and model selection guidance.

OpenVINO Backend (Intel CPU — 5.2× Speedup)

openvino_inference.py
from paddleocr import PaddleOCR

# Enable MKL-DNN (OpenVINO) for Intel CPU — 5.2× speedup
ocr = PaddleOCR(
    lang='en',
    use_gpu=False,
    enable_mkldnn=True,    # Enable MKL-DNN optimization
    cpu_threads=8,         # Number of CPU threads
)

# Tip: Keep ocr instance across calls — model init is expensive
result = ocr.predict('image.jpg')

GPU Batch Inference

gpu_batch.py
import glob
from paddleocr import PaddleOCR

# Initialize once with large batch size for GPU
ocr = PaddleOCR(
    lang='en',
    use_gpu=True,
    rec_batch_num=32,       # Larger batch = better GPU utilization
    det_limit_side_len=960, # Limit resize for speed vs accuracy tradeoff
)

images = glob.glob('docs/*.jpg')
results = []
for img in images:
    results.append(ocr.predict(img))
    
print(f'Processed {len(results)} images')

RAG & LLM Pipeline Integration

PaddleOCR is the most widely used document parser in open-source RAG (Retrieval-Augmented Generation) pipelines. It bridges the gap between raw document files (PDFs, images, scanned documents) and the structured text chunks needed by vector databases and LLMs.

PaddleOCR → LangChain → Vector Store

langchain_rag.py
from paddleocr import PPStructureV3
from langchain.text_splitter import MarkdownTextSplitter
from langchain.vectorstores import Chroma

# Step 1: Parse PDF with PaddleOCR → Markdown
pipeline = PPStructureV3()
results = pipeline.predict('knowledge_base.pdf')
full_markdown = '\n\n'.join(p.get('markdown', '') for p in results)

# Step 2: Split into LLM-friendly chunks
splitter = MarkdownTextSplitter(chunk_size=1000, chunk_overlap=100)
chunks = splitter.split_text(full_markdown)

# Step 3: Embed and index in a vector store
from langchain.embeddings import OpenAIEmbeddings
vectorstore = Chroma.from_texts(chunks, OpenAIEmbeddings())

# Step 4: Query your document
docs = vectorstore.similarity_search("What is the Q3 revenue?", k=3)
for doc in docs:
    print(doc.page_content)

MCP Server Mode (LLM Agent Tool)

mcp_server.sh
# Start PaddleOCR as an MCP (Model Context Protocol) server
$ paddleocr serve --port 8080 --lang en

# MCP server exposes OCR as a callable tool for LLM agents
# Compatible with Claude, ChatGPT, and any MCP-enabled agent
# Available tools: ocr_image, parse_pdf, extract_table

JavaScript, Node.js & WebAssembly (PaddleOCR.js)

Run PaddleOCR models directly in web browsers (via WebAssembly & WebGPU) or on Node.js / Bun servers without Python runtime.

npm_install.sh
# Install official PaddleJS browser OCR & ONNX Web runtime
$ npm install @paddlejs-models/ocr onnxruntime-web
browser_ocr.js
import * as ocr from '@paddlejs-models/ocr';

// Initialize and load PP-OCR Web model (cached automatically in browser)
await ocr.init();

// Recognize from HTML <img>, <canvas>, or video stream
const imgElement = document.getElementById('scanned-doc');
const result = await ocr.recognize(imgElement);

console.log('Recognized text:', result.text);
console.log('Bounding boxes:', result.points);

Python API Reference

PaddleOCR Class

PaddleOCR(lang, use_gpu, use_angle_cls, ...)
Main OCR pipeline class. Initializes text detection, angle classification, and text recognition models. Initialize once and reuse — model loading takes 1–3 seconds.
.predict(input)str | PIL.Image | np.ndarray
Run OCR on an image file path, PIL Image, or NumPy array. Returns a list of page results, each containing rec_texts, rec_boxes, and rec_scores.

PPStructureV3 Class

PPStructureV3(lang, use_gpu, ...)
Document layout analysis and structure pipeline. Handles PDFs (multi-page), images, and office documents.
.predict(input)
Parse a document. Returns list of page dicts with keys: markdown, blocks, page_id. Each block has type, bbox, and type-specific data.

Command Line Interface (CLI) Reference

PaddleOCR can be invoked from the command line using the paddleocr command after installation.

CLI Usage
# Basic OCR on a directory of images
$ paddleocr --image_dir=./images --lang=en --use_gpu=False

# Structure analysis (detect tables, titles, layout)
$ paddleocr --image_dir=./docs --type=structure --lang=en

# Single image OCR
$ paddleocr --image_dir=image.jpg --lang=ch

# With angle classification (for rotated text)
$ paddleocr --image_dir=./scans --lang=en --use_angle_cls=True

# Start OCR HTTP server
$ paddleocr serve --port 8080 --lang en

CLI Flags Reference

Flag Type Default Description
--image_dirstrPath to image file or directory
--langstrenRecognition language code
--use_gpuboolFalseEnable GPU acceleration
--use_angle_clsboolFalseEnable text angle classification
--typestrocrPipeline type: ocr or structure
--det_limit_side_lenint736Max image side for detection (resize)
--rec_batch_numint6Recognition batch size
--outputstr./outputOutput directory for results
--show_logboolTruePrint inference logs

Supported Language Codes

PaddleOCR supports 100+ languages. PP-OCRv6's unified model covers 50 languages out-of-the-box without any additional downloads. The following table lists commonly used language codes:

Language Code Language Code
EnglishenChinese (Simplified)ch
Chinese (Traditional)chinese_chtJapanesejapan
KoreankoreanArabicarabic
FrenchfrenchGermangerman
SpanishesItalianit
PortugueseptRussianru
HindihiThaith
VietnameseviLatin (general)latin
DevanagaridevanagariCyrilliccyrillic
KannadakaTamilta

Troubleshooting & Developer Error Dictionary

Direct solutions for common Python stack traces, argument deprecations, and runtime issues.

1. ImportError: cannot import name 'draw_ocr' from 'paddleocr'

Cause: In PaddleOCR v2.7+, visualization utilities were moved to optional tools dependencies.

Fix: Install tools support and import from the correct submodule:

pip install "paddleocr[tools]"

# Correct import path:
from paddleocr.tools.infer.utility import draw_ocr

2. TypeError: predict() got an unexpected keyword argument 'cls'

Cause: Passing cls=True to .predict() instead of .ocr(). In PaddleOCR, use_angle_cls is defined during class initialization or passed into .ocr(img, cls=True).

# Correct Pattern:
ocr = PaddleOCR(use_angle_cls=True, lang='en')
result = ocr.ocr('img.jpg', cls=True)
# Or use predict without cls parameter:
result = ocr.predict('img.jpg')

3. ValueError: Unknown argument 'use_gpu' (PP-OCRv6 / v3.7+)

Cause: PaddleOCR v3.7 unified hardware acceleration under device configurations.

Fix: Specify device type directly or let PaddlePaddle auto-detect CUDA:

# Specify device explicitly if needed:
import paddle
paddle.device.set_device('gpu' if paddle.is_compiled_with_cuda() else 'cpu')
ocr = PaddleOCR(lang='en')

4. convertpirattribute2runtimeattribute (OneDNN Windows Error)

Cause: Conflict between experimental PIR compiler and OneDNN CPU acceleration on Windows.

Fix: Disable PIR mode before running inference:

import os
os.environ['FLAGS_enable_pir_api'] = '0'

5. How to Extract Bounding Box Coordinates & Filter Confidence

Extract polygon vertices and filter out low-confidence detections:

result = ocr.ocr('invoice.jpg', cls=True)

for line in result[0]:
    box = line[0]        # [[x1,y1], [x2,y2], [x3,y3], [x4,y4]]
    text, score = line[1] # ('Invoice Total', 0.9842)
    if score >= 0.85:
        print(f"High-confidence text: {text} at {box}")

6. Model Download Timeout or Connection Error

Fix: Pre-download models to custom local folder: export PADDLEOCR_HOME=/custom/models, or load directly from HuggingFace Hub.

Changelog

v3.7.0 — June 11, 2026 (Latest)

  • Released PP-OCRv6 with 3 model tiers (tiny 1.5M / small 7.7M / medium 34.5M)
  • +4.6% detection, +5.1% recognition accuracy vs PP-OCRv5_server
  • 50 language unified model (Chinese, English, Japanese, 46 Latin-script)
  • 5.2× CPU speedup via OpenVINO; 6.1× on Apple M4 (tiny model)
  • Improvements in digital displays, dot-matrix, tire prints, industrial text

v3.6.0 — May 28, 2026

  • Released PaddleOCR-VL-1.6 — 96.33% on OmniDocBench v1.6 (SOTA)
  • Official Python, Go, and TypeScript API SDKs released
  • Multi-page TIFF file parsing support added

v3.5.0 — April 21, 2026

  • Flexible inference backends: Paddle static, dynamic, and Transformers
  • Word, Excel, PowerPoint → Markdown conversion
  • DOCX export for parsed results
  • PaddleOCR.js — official browser inference SDK for PP-OCRv5

v3.4.0 — January 29, 2026

  • HPD-Parsing model added — 4,752 tokens/sec peak throughput
  • MCP server mode for direct LLM agent tool use
  • PP-ChatOCRv4 with improved document Q&A capabilities

View full changelog on GitHub →