A C++ inference implementation of PaddleOCR using onnxruntime and opencv, runnable on Windows x64.
Two OCR capabilities
- Full-image detection (text location) + recognition (text content)
- Recognition on a selected ROI
api_server is a single long-running exe that provides both:
- An HTTP API (
POST /ocr_detect,POST /ocr_recognize) for other services to call - A built-in web UI (see Frontend below) — open it in a browser to try OCR on an image
2026-07-26.022931_24fps.1.mp4
src/
core/ text_det / text_rec shared OCR logic
api_server/ main.cpp - HTTP API server (POST /ocr_detect, POST /ocr_recognize)
frontend/ Vue 3 + Vite web UI; once built, it's served by api_server itself (see below)
weights/ onnx models & dictionaries
images/ sample images, handy for testing the API manually with curl/Postman (see "Using api_server")
third_party/ httplib.h (single-header HTTP library used by api_server)
vs2022/ Visual Studio 2022 solution (.slnx + api_server's .vcxproj)
cmake/ CMakeLists.txt, for building via VSCode (CMake Tools extension) or any other IDE
packages/ opencv / onnxruntime third-party SDKs (download yourself, not version-controlled, see below)
frontend/ is a Vue 3 + Vite web page: pick an image, hit Detect/Recognize, and the detected text boxes get drawn on top of the image. It isn't a standalone web app — the built static files are served by api_server itself via httplib's set_mount_point(), so the page and the API share the same port, with no separate server and no CORS to deal with.
While developing (frontend and backend run separately, so UI edits show up instantly):
cd frontend
npm install
npm run devvite.config.js already proxies /health, /ocr_detect, /ocr_recognize to http://127.0.0.1:8080 (start api_server.exe separately). Open the URL Vite prints (default http://localhost:5173).
To have api_server.exe serve the page itself (for regular use / packaging):
cd frontend
npm run buildThis produces frontend/dist. When api_server is then built (VS2022 or CMake), the build script copies frontend/dist next to the output exe automatically; open http://127.0.0.1:8080/ in a browser and the web UI is right there — no npm run dev needed.
⚠️ Do this before the "Building" section below — both build methods only copy whateverfrontend/distalready exists; neither one builds the frontend for you.
Before you start, make sure
frontend/distis already built (see Frontend above) — otherwise the Post-Build Event won't find it to copy, and the web UI won't end up in the output directory.
Open vs2022/PaddleOCR-cpp.slnx, right-click the api_server project in Solution Explorer → Rebuild (or just Build).
Building it automatically triggers (configured in api_server.vcxproj):
- Pre-Build Event:
taskkills any still-runningapi_server.exe(otherwise the linker fails because the exe file is locked) - Compilation
- Post-Build Event: runs
copy_bin.bat, which copies the opencv/onnxruntime DLLs andweights/(andfrontend/dist, if it exists) into the output directoryvs2022/x64/<Debug|Release>/, thenstarts the freshly builtapi_server.exe
So whether you Rebuild or hit F5, api_server.exe pops up running on its own after the build — no need to hunt down the exe and launch it manually.
Before you start, make sure
frontend/distis already built (see Frontend above). This matters more for CMake:frontend/distneeds to exist before the firstcmakeconfigure for the copy step to be added at all; if you build the frontend after already configuring, run CMake: Delete Cache and Reconfigure once for it to take effect.
- Install the CMake Tools extension (and the C/C++ extension if you want to debug), then open this repo folder —
.vscode/settings.jsonalready setscmake.sourceDirectorytocmake/, so it picks upcmake/CMakeLists.txtautomatically. - Pick a build variant (Debug/Release):
Ctrl+Shift+P→ CMake: Select Variant and choose Debug (the status bar doesn't always show a variant button, so this command is the reliable way — it also tells you which one is currently active). Pick Debug if you plan to debug —launch.jsonpoints at the Debug build's exe specifically. - Run CMake: Build (or the build button in the status bar). This produces
cmake/build/Debug/api_server.exeand automatically copies the opencv/onnxruntime DLLs,weights/(andfrontend/dist, if present) into the same folder. The opencv/onnxruntime paths default to the folders underpackages/, and can be overridden with-DOPENCV_DIR=.../-DONNXRUNTIME_DIR=.... - Debugging: hit F5 — it uses the project's
.vscode/launch.jsondirectly (cppvsdbg, matching the PDBs MSVC produces), so you won't see the "Select debugger" picker, and breakpoints work normally.
api_server.exe [host] [port] # defaults: host 0.0.0.0 (listen on all interfaces), port 8080
Example: api_server.exe 127.0.0.1 9000 listens only on localhost, port 9000.
GET /health- health checkPOST /ocr_detect- body is the raw bytes of a full image (jpg/png/bmp...); detection only, returns the array of detected text boxes:[[[x,y],[x,y],[x,y],[x,y]], ...]POST /ocr_recognize- body is the raw bytes of an already-cropped single text line image; recognition only, returns the recognized text:{"text":"..."}
The body for both endpoints must be the raw binary content of an image file — not form-data, not base64, not wrapped in JSON.
-
curl: use
--data-binary @path/to/file(not-d/--data, which treats the content as text), and explicitly set-H "Content-Type: application/octet-stream"— curl defaults toapplication/x-www-form-urlencodedwhen unspecified, and httplib caps that content type at 8KB, so anything bigger gets rejected with 413.# detection (test_detection.bmp is a full image) curl -X POST -H "Content-Type: application/octet-stream" --data-binary @images/test_detection.bmp http://127.0.0.1:8080/ocr_detect # recognition (test_recognition.bmp is an already-cropped single text line) curl -X POST -H "Content-Type: application/octet-stream" --data-binary @images/test_recognition.bmp http://127.0.0.1:8080/ocr_recognize
-
Postman / Insomnia: on the Body tab, choose binary (not form-data or raw), then pick the file
-
PowerShell:
# detection $bytes = [System.IO.File]::ReadAllBytes("images/test_detection.bmp") Invoke-RestMethod -Uri "http://127.0.0.1:8080/ocr_detect" -Method Post -Body $bytes -ContentType "application/octet-stream" # recognition $bytes = [System.IO.File]::ReadAllBytes("images/test_recognition.bmp") Invoke-RestMethod -Uri "http://127.0.0.1:8080/ocr_recognize" -Method Post -Body $bytes -ContentType "application/octet-stream"
-
Python:
import requests # detection with open("images/test_detection.bmp", "rb") as f: r = requests.post("http://127.0.0.1:8080/ocr_detect", data=f.read()) print(r.json()) # recognition with open("images/test_recognition.bmp", "rb") as f: r = requests.post("http://127.0.0.1:8080/ocr_recognize", data=f.read()) print(r.json())
When the body isn't in the right format (e.g. JSON or form-data), the server responds with HTTP 400 and a hint field telling you to send raw binary:
{
"error": "could not decode image",
"hint": "body must be the raw image bytes (jpg/png/bmp/...), not form-data or base64, e.g. curl -H \"Content-Type: application/octet-stream\" --data-binary @file.jpg"
}- onnxruntime-win-x64-gpu-1.21.0 [ https://github.com/microsoft/onnxruntime/releases ]
- opencv 4.5.0 windows [ https://opencv.org/releases/ ]
Placement path (same directory level as vs2022/ and cmake/, at the project root):
./packages
├─ onnxruntime-win-x64-gpu-1.21.0
└─ opencv
List of PP-OCR models
Python environment setup
- python 3.10.10
- pip install paddle2onnx-2.0.2rc3
- Download the inference model and unzip it
- Run the following command to convert the model to onnx and place it at the path below (
./weights/), edit the paths yourself: - C:\Python31010-OCR-2onnx\Scripts\paddle2onnx.exe --model_dir "C:\Users\Users\Downloads\en_PP-OCRv5_mobile_rec_infer\en_PP-OCRv5_mobile_rec_infer" --model_filename inference.json --params_filename inference.pdiparams --save_file "C:\Users\Users\Downloads\en_PP-OCRv5_mobile_rec_infer\en_PP-OCRv5_mobile_rec_infer\model.onnx"
Code to edit
src/api_server/main.cpp-- TextDetector detect_model("det onnx model path");
- TextRecognizer rec_model("rec onnx model path", "rec dict.txt path");
Example: PP-OCRv6_medium_rec
-
Download the inference model and go through the .onnx conversion steps [https://www.paddleocr.ai/latest/en/version3.x/module_usage/text_detection.html]
Model page 
-
Find
character_dict_pathin the .yml file and download the dict.txt needed for recognition [https://github.com/PaddlePaddle/PaddleOCR/blob/main/configs/rec/PP-OCRv6/PP-OCRv6_medium_rec.yml]File page 