தமிழ் வட்டார மொழிகளை AI மொழி மாதிரிகள் கையாளும் திறன்: கொங்கு வட்டாரத் தமிழை மையமாகக் கொண்ட ஆய்வு

The Capability of AI Language Models in Handling Tamil Regional Dialects: A Study Centred on Kongu Tamil

Authors

  • N.Manjula Assistant professor, Department of Tamil, A.V.P.College of Arts and Science,Thirumurugan poondi post,Tirupur-641652 Author
  • Dr. S. Padma Author
  • Dr.C.Sakthimurugan  Assistant professor in Tamil, Nandha Arts and Science College (Autonomous), Erode -638052 Author
  • Dr.V.C.Srinivasan Administrative Officer, Nandha Arts and Science College (Autonomous), Erode. Author

Keywords:

Kongu Tamil, Dialect, Large Language Model, LLM, Zero-shot learning, Few-shot learning

Abstract

While Large Language Models (LLMs) demonstrate significant capabilities in standard Tamil, the extent to which they accurately understand and translate Kongu regional Tamil—spoken in the Kongu region including Coimbatore, Tiruppur, Erode, Salem, and Karur—remains an insufficiently explored area. This article examines the dialect-handling capabilities of AI language models, focusing on Kongu Tamil. For this purpose, a small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary. Using this dataset, zero-shot and few-shot experiments were conducted with a publicly accessible Large Language Model (Claude), and the responses were evaluated based on the author's linguistic knowledge, with errors categorized accordingly. Additionally, citing results from previously published studies comparing various LLMs (Gemini, ChatGPT, Claude), this paper proposes a comprehensive methodological plan for a complete multi-model comparative study for Kongu Tamil, along with a human expert evaluation protocol. Preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions. This study underscores the need to build large-scale, human-expert-verified benchmark datasets for Tamil dialects.

Downloads

Download data is not yet available.

Author Biographies

References

1. Alam, M. M. I., & Anastasopoulos, A. (2024). CODET: A benchmark for contrastive dialectal evaluation of machine translation. In Findings of the Association for Computational Linguistics: EACL 2024 (pp. 1790–1859). Association for Computational Linguistics.

2. BhashaSutra: A task-centric unified survey of Indian NLP datasets, corpora, and resources. (2025). arXiv preprint.

3. Chakravarthi, B. R., Priyadharshini, R., Muralidaran, V., Suryawanshi, S., Jose, N., Sherly, E., & McCrae, J. P. (2021). Findings of the shared task on Dravidian language technology (DravidianLangTech). Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages.

4. From phonemes to meaning: Evaluating large language models on Tamil (ILAKKANAM). (2025). arXiv:2511.12387.

5. G.T.N. Arts College, Dindigul. (n.d.). The impact of artificial intelligence and natural language processing on the Tamil language preservation and cultural promotion. Tamil Computing Journal.

6. Jageer, S. K., & Priyaradhikadevi. (2025). A comparative analysis of large language models for English-to-Tamil machine translation with performance evaluation. Panamerican Mathematical Journal, 35(2s).

7. Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6282–6293).

8. Rajendran, S., Anand Kumar, M., Rajalakshmi, R., Dhanalakshmi, V., Balasubramanian, P., & Soman, K. P. (2022). Tamil NLP technologies: Challenges, state of the art, trends and future scope. In International Conference on Speech and Language Technologies for Low-Resource Languages (pp. 73–98). Springer.

9. Shaffi, N., & Hajamohideen, F. (2021). UTHCD: A new benchmarking for Tamil handwritten OCR. IEEE Access, 9, 101469–101493.

10. Tamil Wikipedia. (2024). தமிழ் வட்டார மொழி வழக்குகள் [Tamil regional dialects]. தமிழ் விக்கிப்பீடியா.

11. Tamil Wikipedia. (2025). கொங்கு வட்டார வழக்கு அகராதி [Kongu dialect dictionary]. தமிழ் விக்கிப்பீடியா.

12. Tamil Wikipedia. (2026). கொங்குத் தமிழ் [Kongu Tamil]. தமிழ் விக்கிப்பீடியா.

13. Tamilmanam International Research Journal of Tamil Studies. (2025a). Artificial intelligence technology: A boon in writing Tamil essays. Tamilmanam International Research Journal of Tamil Studies, 1(4), 221–228. https://doi.org/10.63300/dn796x92

14. Tamilmanam International Research Journal of Tamil Studies. (2025b). Artificial intelligence for Tamil literature: An overview. Tamilmanam International Research Journal of Tamil Studies.

15. Verma, S. S. U. R., Khan, M. S. U. R., Kumar, V., Murthy, R., & Sen, J. (2025). MILU: A multi-task Indic language understanding benchmark. In Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 10076–10132).

Downloads

Published

08/18/2026

How to Cite

தமிழ் வட்டார மொழிகளை AI மொழி மாதிரிகள் கையாளும் திறன்: கொங்கு வட்டாரத் தமிழை மையமாகக் கொண்ட ஆய்வு: The Capability of AI Language Models in Handling Tamil Regional Dialects: A Study Centred on Kongu Tamil. (2026). Tamilmanam International Research Journal of Tamil Studies, 12(03), 609-627. https://tamilmanam.in/journal/index.php/issue/article/view/645

Similar Articles

71-80 of 383

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)