Dataset · 2025 · Zenodo
Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27
Diverse Fonts with Corresponding Ground Truth Annotations
An open benchmark for Tamil OCR covering handwritten and printed
text — roughly 769,396 file pairs (TIFF images, ground truth, and
box files) totalling 3.6 GB. Source text is drawn from Wikipedia,
Wikisource, Maattru, and Theekkathir, rendered across 27 Unicode
fonts: 9 handwritten and 18 printed, from Google Fonts and Tamil
Virtual University. Released CC0, with data collection by the
Kaniyam Foundation.
- Sole author
- CC0
- Tesseract
- Ground truth
Journal article · 2025 · Panamerican Mathematical Journal
A Comparative Analysis of Large Language Models for
English-to-Tamil Machine Translation with Performance Evaluation
Machine translation matters most for low-resource languages, and
Tamil is hard: complex morphology, syntax, and cultural nuance.
This paper evaluates Claude, ChatGPT, and Gemini on
English-to-Tamil translation across culturally sensitive,
political, technical, and idiomatic content, scored with BLEU,
BERT-based metrics, METEOR, and TER. Gemini ranked highest on
accuracy and precision for complex syntax and idioms, ChatGPT led
on semantic alignment, and Claude was strongest on document-level
fluency.
- First author
- LLMs
- Machine Translation
- Evaluation
SyedKhaleel Jageer, Priyaradhikadevi T., Madhan K., Prasanna S.
Panamerican Mathematical Journal, Vol. 35, No. 2s (2025).