Research · Ongoing
Tamil OCR Research Platform
An ongoing research initiative focused on improving handwritten
Tamil optical character recognition using Tesseract OCR and
AI-assisted post-processing.
- Large-scale dataset preparation
- Ground truth generation
- OCR model training
- Character Error Rate (CER) optimisation
- Word Error Rate (WER) evaluation
- AI-assisted correction pipelines
- Tesseract
- Python
- OCR
- Datasets
Developer tooling
GT Builder
A Python tool that generates ground truth data for Tesseract OCR
in Tamil — normalising raw text, rendering it across multiple
fonts into TIFF/.gt.txt pairs, verifying the result,
and scoring models with CER and WER. It produced the published
synthetic OCR dataset.
- Python
- Tesseract
- OCR
- Datasets
Open knowledge
FreeTamilEBooks Android
An Android application developed to improve access to Tamil digital
books, supporting open knowledge and community-driven publishing.
Personal lab
Home Automation Lab
A personal engineering lab built around Raspberry Pi, Home
Assistant, Linux, Docker, and self-hosted services — used to
experiment with automation and embedded technologies.
- Raspberry Pi
- Home Assistant
- Docker
- Linux