தமிழ் வட்டார மொழிகளை AI மொழி மாதிரிகள் கையாளும் திறன்: கொங்கு வட்டாரத் தமிழை மையமாகக் கொண்ட ஆய்வு
The Capability of AI Language Models in Handling Tamil Regional Dialects: A Study Centred on Kongu Tamil
Keywords:
Kongu Tamil, Dialect, Large Language Model, LLM, Zero-shot learning, Few-shot learningAbstract
While Large Language Models (LLMs) demonstrate significant capabilities in standard Tamil, the extent to which they accurately understand and translate Kongu regional Tamil—spoken in the Kongu region including Coimbatore, Tiruppur, Erode, Salem, and Karur—remains an insufficiently explored area. This article examines the dialect-handling capabilities of AI language models, focusing on Kongu Tamil. For this purpose, a small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary. Using this dataset, zero-shot and few-shot experiments were conducted with a publicly accessible Large Language Model (Claude), and the responses were evaluated based on the author's linguistic knowledge, with errors categorized accordingly. Additionally, citing results from previously published studies comparing various LLMs (Gemini, ChatGPT, Claude), this paper proposes a comprehensive methodological plan for a complete multi-model comparative study for Kongu Tamil, along with a human expert evaluation protocol. Preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions. This study underscores the need to build large-scale, human-expert-verified benchmark datasets for Tamil dialects.
Downloads
References
1. Alam, M. M. I., & Anastasopoulos, A. (2024). CODET: A benchmark for contrastive dialectal evaluation of machine translation. In Findings of the Association for Computational Linguistics: EACL 2024 (pp. 1790–1859). Association for Computational Linguistics.
2. BhashaSutra: A task-centric unified survey of Indian NLP datasets, corpora, and resources. (2025). arXiv preprint.
3. Chakravarthi, B. R., Priyadharshini, R., Muralidaran, V., Suryawanshi, S., Jose, N., Sherly, E., & McCrae, J. P. (2021). Findings of the shared task on Dravidian language technology (DravidianLangTech). Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages.
4. From phonemes to meaning: Evaluating large language models on Tamil (ILAKKANAM). (2025). arXiv:2511.12387.
5. G.T.N. Arts College, Dindigul. (n.d.). The impact of artificial intelligence and natural language processing on the Tamil language preservation and cultural promotion. Tamil Computing Journal.
6. Jageer, S. K., & Priyaradhikadevi. (2025). A comparative analysis of large language models for English-to-Tamil machine translation with performance evaluation. Panamerican Mathematical Journal, 35(2s).
7. Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6282–6293).
8. Rajendran, S., Anand Kumar, M., Rajalakshmi, R., Dhanalakshmi, V., Balasubramanian, P., & Soman, K. P. (2022). Tamil NLP technologies: Challenges, state of the art, trends and future scope. In International Conference on Speech and Language Technologies for Low-Resource Languages (pp. 73–98). Springer.
9. Shaffi, N., & Hajamohideen, F. (2021). UTHCD: A new benchmarking for Tamil handwritten OCR. IEEE Access, 9, 101469–101493.
10. Tamil Wikipedia. (2024). தமிழ் வட்டார மொழி வழக்குகள் [Tamil regional dialects]. தமிழ் விக்கிப்பீடியா.
11. Tamil Wikipedia. (2025). கொங்கு வட்டார வழக்கு அகராதி [Kongu dialect dictionary]. தமிழ் விக்கிப்பீடியா.
12. Tamil Wikipedia. (2026). கொங்குத் தமிழ் [Kongu Tamil]. தமிழ் விக்கிப்பீடியா.
13. Tamilmanam International Research Journal of Tamil Studies. (2025a). Artificial intelligence technology: A boon in writing Tamil essays. Tamilmanam International Research Journal of Tamil Studies, 1(4), 221–228. https://doi.org/10.63300/dn796x92
14. Tamilmanam International Research Journal of Tamil Studies. (2025b). Artificial intelligence for Tamil literature: An overview. Tamilmanam International Research Journal of Tamil Studies.
15. Verma, S. S. U. R., Khan, M. S. U. R., Kumar, V., Murthy, R., & Sen, J. (2025). MILU: A multi-task Indic language understanding benchmark. In Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 10076–10132).
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Our journal adopts CC BY License Creative Commons Attribution 4.0 International License http://Creativecommons.org//license/by/4.0/ . It allows using, reusing, distributing and reproducing of the original work with proper citation.