- 306 Ethiopic characters
- 356-character Ethiopic vocabulary
Context
Ethiopic script, used for Amharic, Tigrinya and Ge'ez, is poorly served by commercial OCR. Archives, records and teaching material sit on paper or in scanned images that no search engine can read. I built an OCR system for Ethiopic script in two parts: a working full-stack application on top of the pre-trained HHD-Ethiopic recognizer, and a research line for training better detection and recognition models with MMOCR.
Constraints
- Ethiopic is a syllabary with a large character set and many visually close glyphs. The application's recognizer covers 306 Ethiopic characters: the syllabary, punctuation and numerals.
- Documents arrive as photos, scans and multi-page PDFs, with skew, uneven lighting and mixed quality.
- People need to see what the model saw: region-level results with a confidence for each, not a block of text to take on trust.
- Personal hardware and public data only.
Architecture
The backend is a FastAPI service with endpoints for image OCR, base64 input, batch processing, detection only, detect-and-recognize, recognition of a chosen region, and PDF information and OCR. The pipeline has four stages. Preprocessing converts to grayscale, binarises with adaptive thresholding and deskews. Hierarchical detection finds lines with horizontal projection profiles, words with vertical projection and characters with connected-component analysis, returning bounding boxes at all three levels. Recognition feeds each region through the HHD-Ethiopic TensorFlow model and decodes the output with CTC, producing text and a confidence score per region. PDFs are rendered page by page with PyMuPDF and processed as images.
The frontend is a Next.js application with a detection canvas that overlays the boxes on the image; clicking a box shows its recognised text, and a results panel shows the line, word and character tree. Upload is drag and drop.
The research line uses MMOCR: a SATRN recognizer over a 356-character Ethiopic vocabulary, DBNet++ for detection on Amharic documents, a synthetic data generator for Amharic word images, and ONNX export so trained models can be served without the training stack.
Decisions and tradeoffs
- Ship on a pre-trained recognizer first. The application exists and can be tried today. The cost is domain: HHD-Ethiopic comes from historical handwriting, so modern printed documents differ from what it learned and results vary with the input.
- Projection-profile detection in the application rather than a learned detector. No training, fast, and every box is explainable. It struggles with strong skew, multi-column layouts and dense handwriting, which is why DBNet++ is in the research line.
- Confidence exposed per region. It lets a reviewer decide what to trust. CTC confidences are not calibrated, so a score is a ranking signal, not a probability.
- Two repositories. Training dependencies (MMOCR, PyTorch) and serving dependencies (TensorFlow, FastAPI) do not belong in one environment, and ONNX is the bridge between them.
Outcome
A working application: upload an image or a PDF, see detected lines, words and characters on the canvas, read the recognised text with its confidence, and call the same pipeline through the API, including in batch. The research line has training configurations for the recognizer and the detector, a synthetic data generator and an export path. No accuracy figures are published yet: the research README states a target error rate, and a target is not a result.
What I would change
- Evaluate first: character and word error rates on a held-out set of real documents, published with the split.
- Replace projection detection with DBNet++ in the application, and move the recognizer to SATRN through ONNX once both are evaluated.
- Fine-tune for modern printed text, where most practical demand is, using the synthetic generator for fonts.
- Package the backend with the model fetched at build time, so the demo runs with one command.
Stack and links
FastAPI, TensorFlow, OpenCV and PyMuPDF on the backend; Next.js, TypeScript and Tailwind CSS on the frontend; MMOCR, PyTorch and ONNX in the research line. The application repository is private for now.