Evaluation of Gen-AI (LLM) Results on Prompts in the Batak Language

Authors

  • Ade Siahaan
  • Angelina Nadeak
  • Jeremi Samosir
  • Jesica Siburian Del Institute of Technology
  • Boy Hutajulu

DOI:

https://doi.org/10.54074/jicsa.v2i1.45

Abstract

The rapid advancement of Large Language Models (LLMs) has
significantly enhanced natural language processing capabilities;
however, their performance on low-resource regional languages
remains underexplored. The Toba Batak language, a regional
language spoken in North Sumatera, Indonesia, faces limited
representation in global AI training datasets, raising concerns about
linguistic accuracy and cultural preservation in AI-generated
content. This study evaluates the performance of four Large
Language Models—ChatGPT, Gemini, Perplexity, and Toba-LLM—
in generating responses to prompts written in the Toba Batak
language. The evaluation follows the CRISP-DM methodology and
employs a structured dataset consisting of categorized prompts and
corresponding ground truth answers. Model outputs are assessed
using six quantitative metrics: lexical similarity, keyword overlap,
length ratio, cosine similarity, cultural relevance, and fluency. The
results indicate that general-purpose models such as ChatGPT and
Gemini outperform the local Toba-LLM in terms of semantic
similarity and fluency, while Toba-LLM demonstrates higher lexical
alignment and response conciseness. However, cultural relevance
remains a challenge across all models, highlighting the difficulty of
capturing culturally grounded knowledge in low-resource language
settings. These findings underscore the need for culturally aware
evaluation metrics and enriched local language datasets to improve
the applicability of LLMs for regional language preservation and AI-
based educational applications.

References

[1] U. Janssens, “Functional hemodynamic monitoring,” 2024. doi: 10.1007/s00063-024-01190-4.

[2] J. Yang et al., “Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond,” ACM

Trans Knowl Discov Data, vol. 18, no. 6, pp. 1–24, 2024, doi: 10.1145/3649506.

[3] “Toba Batak language.” [Online]. Available: https://en.wikipedia.org/wiki/Toba_Batak_language

[4] D. F. Purba, F. Malau, M. L. A. Siahaan, and S. Napitupulu, “a Contrastive Analysis Between English

and Batak Toba Language in Verbal Affixes,” Journal of Humanities, Social Sciences and Business (Jhssb),

vol. 1, no. 4, pp. 121–128, 2022, doi: 10.55047/jhssb.v1i4.288.

[5] G. I. Winata et al., “NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local

Languages,” EACL 2023 - 17th Conference of the European Chapter of the Association for Computational

Linguistics, Proceedings of the Conference, pp. 815–834, 2023, doi: 10.18653/v1/2023.eacl-main.57.

[6] M. A. Hedderich, L. Lange, H. Adel, J. Strötgen, and D. Klakow, “A Survey on Recent Approaches

for Natural Language Processing in Low-Resource Scenarios,” NAACL-HLT 2021 - 2021 Conference of

the North American Chapter of the Association for Computational Linguistics: Human Language

Technologies, Proceedings of the Conference, pp. 2545–2568, 2021, doi: 10.18653/v1/2021.naacl-main.201.

[7] S. Kendre, A. Xu, H. Zhou, M. Ryoo, S. Joty, and J. C. Niebles, “SMILE: A Composite Lexical-

Semantic Metric for Question-Answering Evaluation,” pp. 1–23, 2025, [Online]. Available:

http://arxiv.org/abs/2511.17432

[8] S. Joshi, “Evaluation of Large Language Models: Review of Metrics, Applications, and

Methodologies,” Preprints (Basel), pp. 0–22, 2025, doi: 10.20944/preprints202504.0369.v1.

[9] Z. Ji et al., “Survey of Hallucination in Natural Language Generation,” ACM Comput Surv, vol. 55, no.

12, 2023, doi: 10.1145/3571730.

Downloads

Published

2026-09-30