Research

Applying research to solve real-world problems — improving handwritten Tamil recognition, and the evaluation work that tells me whether it actually improved.

My research bridges software engineering and language technology, building practical solutions that improve accessibility, document digitisation, and regional language computing.

I approach research as an industry engineer, with a strong emphasis on reproducibility, practical implementation, and measurable improvement. The work combines traditional Optical Character Recognition with modern AI to better recognise handwritten and printed Tamil documents.

The goal is technology that preserves knowledge, improves accessibility, and supports the digital transformation of regional language content.

Research interests

  • Optical Character Recognition
  • Tamil Language Technologies
  • Artificial Intelligence
  • Large Language Models
  • Machine Translation
  • Document Digitisation
  • Accessibility Technologies
  • Computer Vision
  • Natural Language Processing
  • Human-Computer Interaction
  • Engineering Automation

Current research areas

Tamil Language Technologies

Technologies that support the digital ecosystem of the Tamil language — from encoding through to preservation.

  • Tamil OCR
  • Unicode processing
  • Text normalisation
  • Language resources
  • Digital preservation
  • Corpus generation

AI-enhanced OCR

Integrating large language models into OCR pipelines to correct errors, restore missing characters, recover formatting, and improve readability. AI serves as a post-processing layer that improves extracted text — it does not replace the OCR engine, and it must not invent text that was never on the page.

  • LLMs
  • Post-processing
  • Document understanding

Accessibility technologies

Improving access to printed and handwritten Tamil content for visually impaired users through OCR-powered assistive applications.

  • OCR
  • Android
  • Text-to-Speech
  • Mobile computing

Tamil OCR Research Platform

Improving handwritten Tamil optical character recognition using Tesseract OCR and AI-assisted post-processing.

Dataset & ground truth

Large-scale dataset preparation and ground truth generation for handwritten Tamil, including the tooling to make that work repeatable.

Model training

Training and fine-tuning OCR models, then optimising Character Error Rate against a held-out evaluation set.

Evaluation

Character Error Rate (CER) and Word Error Rate (WER) analysis, to measure whether a change actually helped.

AI-assisted correction

Post-processing pipelines that use language models to clean up OCR output without inventing text that was never on the page.

This research feeds directly into Vizhi Tamil, the accessibility app that puts Tamil OCR in users' hands.

How I work

An engineering-driven methodology, start to finish.

  1. Step 01

    Problem identification

    • Identify practical challenges affecting OCR accuracy, usability, or accessibility.
  2. Step 02

    Dataset preparation

    • Collect and prepare datasets from real handwritten documents, supplemented with synthetic data generation.
  3. Step 03

    Model development

    • Train and optimise OCR models using Tesseract and language-specific improvements.
  4. Step 04

    Evaluation

    • Measure with Character Error Rate, Word Error Rate, recognition accuracy, and performance benchmarks — before claiming anything improved.
  5. Step 05

    AI enhancement

    • Use large language models to improve OCR output through intelligent post-processing.
  6. Step 06

    Deployment

    • Apply the outcome in real-world software and accessibility solutions, where it either helps someone or it does not.

Research tools

OCR

  • Tesseract OCR

Programming

  • Python
  • Kotlin
  • Java

Artificial Intelligence

  • Large Language Models
  • Prompt Engineering
  • AI-assisted Text Processing

Data processing

  • OpenCV
  • ImageMagick
  • Unicode Processing

Development

  • Git
  • Linux
  • Docker

Publications

Published work, datasets, and experimental results.

Dataset · 2025 · Zenodo

Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations

An open benchmark for Tamil OCR covering handwritten and printed text — roughly 769,396 file pairs (TIFF images, ground truth, and box files) totalling 3.6 GB. Source text is drawn from Wikipedia, Wikisource, Maattru, and Theekkathir, rendered across 27 Unicode fonts: 9 handwritten and 18 printed, from Google Fonts and Tamil Virtual University. Released CC0, with data collection by the Kaniyam Foundation.

  • Sole author
  • CC0
  • Tesseract
  • Ground truth
Journal article · 2025 · Panamerican Mathematical Journal

A Comparative Analysis of Large Language Models for English-to-Tamil Machine Translation with Performance Evaluation

Machine translation matters most for low-resource languages, and Tamil is hard: complex morphology, syntax, and cultural nuance. This paper evaluates Claude, ChatGPT, and Gemini on English-to-Tamil translation across culturally sensitive, political, technical, and idiomatic content, scored with BLEU, BERT-based metrics, METEOR, and TER. Gemini ranked highest on accuracy and precision for complex syntax and idioms, ChatGPT led on semantic alignment, and Claude was strongest on document-level fluency.

  • First author
  • LLMs
  • Machine Translation
  • Evaluation

SyedKhaleel Jageer, Priyaradhikadevi T., Madhan K., Prasanna S. Panamerican Mathematical Journal, Vol. 35, No. 2s (2025).

My published research covers OCR evaluation methodologies, dataset generation, Tesseract optimisation, AI-enhanced OCR, machine translation, Tamil language computing, and accessibility. It is indexed on ResearchGate and Zenodo.

Research timeline

  1. 2023

    Publishing and education

    • Published Essentials of Computing.
    • Expanded my interest in technical education and language technologies.
  2. 2024

    Building the OCR foundation

    • Began focused research in Tamil OCR.
    • Built OCR datasets and trained custom Tesseract models.
    • Developed the OCR evaluation pipelines.
  3. 2025

    Publication and AI integration

    • Published OCR-related research.
    • Explored AI-assisted OCR workflows and improved accuracy through LLM integration.
  4. 2026

    Current exploration

    • AI-enhanced OCR, Tamil language technologies, vision-language models, large language models, accessibility technologies, and intelligent document processing.

Where this is going

Research should solve real problems. I believe meaningful research is practical, reproducible, measurable, open, accessible, and useful to people — measured not only by publications, but by whether the solutions built from it are any good.

Future directions

  • Intelligent Document Processing
  • Vision-Language Models
  • Multimodal AI
  • OCR for Indian Languages
  • Digital Preservation
  • AI-assisted Accessibility
  • Open Datasets
  • Open Research

Collaboration

I welcome opportunities to work with researchers, academic institutions, open-source communities, and industry partners across OCR, artificial intelligence, Tamil language technologies, accessibility, Android development, and open source software.

Research becomes meaningful when it transforms ideas into technologies that improve people's lives.