Preview

Herald of the Kazakh-British Technical University

Advanced search

BENCHMARKING DEEP LEARNING MODELS FOR FEW-SHOT CLASSIFICATION OF KAZAKH NEWS TEXTS UNDER LIMITED ANNOTATION

https://doi.org/10.55452/1998-6688-2026-23-3-336-346

Abstract

Deep learning methods and transformer architectures have fundamentally reshaped the field of natural language processing; however, these advancements remain unevenly distributed. While high-resource languages like English benefit from large-scale benchmarks, robust text classification for low-resource languages, such as Kazakh, continues to pose a significant challenge. This paper presents a comparative analysis of the classical Word2Vec embedding, BiLSTM recurrent models, BERT and XLM-R transformers, the generative mT5 model, alongside classical and statistical baselines for few-shot Kazakh news text classification under limited annotation constraints. Experiments were conducted on the KazNews dataset, curated for few-shot classification from tengrinews.kz and inform.kz articles, as well as parallel HuffPost corpora. The datasets comprise five categories (sports, politics, business, travel, etc.). The experimental findings in 1-shot and 5-shot settings demonstrate that no single model consistently dominates across all datasets and regimes: Word2Vec achieved the best performance on the HuffPost corpora in the 5-shot regime (0.48 and 0.53); mT5-Seq2Seq outperformed others on KazNews under 5-shot (0.71); while BERT remained competitive on the long-text KazNews dataset in the 5-shot setup. The authors also observed a sharp performance gain for the mT5-Seq2Seq model in the 5-shot setup. Conversely, none of the evaluated models yielded acceptable quality metrics in the 1-shot regime.

About the Authors

D. Marlambekov
Al-Farabi Kazakh National University
Kazakhstan

PhD-student

Almaty



A. Akhmetova
Al-Farabi Kazakh National University
Kazakhstan

PhD-student

Almaty



S. Torekul
Al-Farabi Kazakh National University
Kazakhstan

PhD-student
Almaty



M. Ualiyeva
Al-Farabi Kazakh National University
Kazakhstan

Cand. Phys.-Math. Sc., Associate Professor

Almaty



References

1. Pakray, P., Gelbukh, A., & Bandyopadhyay, S. (2025). Natural language processing applications for low-resource languages. Natural Language Processing, 31(2), 183–197. https://doi.org/10.1017/nlp.2024.33

2. Kowsari, K., Jafari Meimandi, K., Heidarysafa, M., Mendu, S., Barnes, L., & Brown, D. (2019). Text classification algorithms: A survey. Information, 10(4), 150. https://doi.org/10.3390/info10040150

3. McCallum, A., & Nigam, K. (1998). A comparison of event models for naive Bayes text classification. AAAI-98 Workshop on Learning for Text Categorization. https://aaai.org/papers/041-ws98-05-007/

4. Joachims, T. (1998). Text categorization with support vector machines: Learning with many relevant features. In C. Nédellec & C. Rouveirol (Eds.), Machine Learning: ECML-98 (pp. 137–142). Springer. https://doi.org/10.1007/BFb0026683

5. Cox, D. R. (1958). The regression analysis of binary sequences. Journal of the Royal Statistical Society: Series B (Methodological), 20(2), 215–242.

6. Salton, G., Wong, A., & Yang, C.S. (1975). A vector space model for automatic indexing. Communications of the ACM, 18(11), 613–620. https://doi.org/10.1145/361219.361220

7. Spärck Jones, K. (2004). A statistical interpretation of term specificity in retrieval. Journal of Documentation, 60(5), 493–502. https://doi.org/10.1108/00220410410560573

8. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735– 1780. https://doi.org/10.1162/neco.1997.9.8.1735

9. Graves, A., & Schmidhuber, J. (2005). Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Networks, 18(5), 602–610. https://doi.org/10.1016/j.neunet.2005.06.042

10. Joulin, A., Grave, E., Bojanowski, P., & Mikolov, T. (2016). Bag of tricks for efficient text classification. arXiv. https://doi.org/10.48550/arXiv.1607.01759

11. Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv. https://doi.org/10.48550/arXiv.1301.3781

12. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171– 4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

13. Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., ... & Zettlemoyer, L. (2019). Unsupervised cross-lingual representation learning at scale. arXiv. https://arxiv.org/abs/1911.02116

14. Xue, L., Constant, N., Roberts, A., Kale, K., Al-Rfou, R., Siddhant, A., ... & Raffel, C. (2021). mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 483–498). https://doi.org/10.18653/v1/2021.naacl-main.41

15. Latief, A. D., Jarin, A., Yuyun, Hidayati, N. N., Afra, D. I. N., & Riza, H. (2025). A systematic review of few-shot and zero-shot learning for NLP in low-resource languages: Insights and challenges. IEEE Xplore. https://ieeexplore.ieee.org/abstract/document/11325582

16. Latief, A. D., Jarin, A., Yuyun, Hidayati, N. N., Afra, D. I. N., & Riza, H. (2025). A systematic review of few-shot and zero-shot learning for NLP in low-resource languages: Insights and challenges. In 2025 International Conference on Computer, Control, Informatics and its Applications (IC3INA) (pp. 358–363). IEEE. https://doi.org/10.1109/IC3INA68387.2025.11325582

17. Toleu, A., Tolegen, G., & Ualiyeva, I. (2025). Fine-Tuning Large Language Models for Kazakh Text Simplification. Applied Sciences, 15(15), 8344. https://doi.org/10.3390/app15158344


Review

For citations:


Marlambekov D., Akhmetova A., Torekul S., Ualiyeva M. BENCHMARKING DEEP LEARNING MODELS FOR FEW-SHOT CLASSIFICATION OF KAZAKH NEWS TEXTS UNDER LIMITED ANNOTATION. Herald of the Kazakh-British Technical University. 2026;23(3):336-346. (In Russ.) https://doi.org/10.55452/1998-6688-2026-23-3-336-346

Views: 5

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 1998-6688 (Print)
ISSN 2959-8109 (Online)