Research Article Open Access

Context Tamil Llama-QA: A Question and Answering Framework for Semantic Information Retrieval Using Tamil Large Language Model

Jeevit Davidson S.1 and M. Murali1
  • 1 Department of Computing Technologies, SRM Institute of Science and Technology, India

Abstract

The Question Answering System (QA) is essential for effectively retrieving and providing relevant information, and its performance is heavily reliant on the efficiency of information retrieval mechanisms. As the volume of information within the corpus continues to grow, the complexity of generating accurate and pertinent answers also increases substantially. This challenge is particularly significant when designing QA systems for regional languages such as Tamil, which present unique linguistic nuances. Currently, the primary retrieval mechanism depends largely on token matching, which identifies relevant words within user queries. While this approach may be effective for straightforward questions, it tends to falter when faced with inquiries that require deeper contextual understanding. As a result, the reliance on token matching can lead to inconsistencies or irrelevant answers, ultimately undermining the quality of the QA system. To overcome these challenges and enhance the consistency and relevance of answers, we propose a novel methodology that combines generative models with context-based retrieval techniques. This approach integrates an ensemble classifier with the Tamil Llama model, significantly improving both the speed and accuracy of information retrieval from the vector store. By leveraging the strengths of both models, we ensure that the retrieval process is context-driven rather than solely reliant on superficial word matches. Additionally, the CHAII dataset, specifically crafted for handling inquiries in Tamil, undergoes classification and labeling using advanced real-time classifier models. This rigorous classification enhances the system's ability to grasp the nuances of user queries, thereby improving response accuracy. Our evaluations demonstrate that this approach not only elevates the quality of the retrieved answers but also substantially reduces the response time for generating accurate answers, leading to a more efficient and user-friendly QA experience.

Journal of Computer Science
Volume 22 No. 10, 2026, 3197-3203

DOI: https://doi.org/10.3844/jcssp.2026.3197.3203

Submitted On: 30 September 2024 Published On: 7 October 2026

How to Cite: S., J. D. & Murali, M. (2026). Context Tamil Llama-QA: A Question and Answering Framework for Semantic Information Retrieval Using Tamil Large Language Model. Journal of Computer Science, 22(10), 3197-3203. https://doi.org/10.3844/jcssp.2026.3197.3203

  • 43 Views
  • 10 Downloads
  • 0 Citations

Download

Keywords

  • Large Language Models (LLMs)
  • Retrieval Augmented Generation (RAG)