Preview

Herald of the Kazakh-British Technical University

Advanced search

MORPHOLOGICAL DISAMBIGUATION FOR THE KAZAKH LANGUAGE USING TRANSFORMER-BASED MODELS

https://doi.org/10.55452/1998-6688-2026-23-3-233-242

Abstract

Morphological ambiguity constitutes a significant challenge for natural language processing in agglutinative languages, as a single word form might include many grammatical categories. The Kazakh language features productive suffixation, vowel harmony, and intricate morphophonological patterns, which considerably hinder automatic morphological analysis. This work presents a transformer-based methodology for morphological disambiguation in Kazakh texts, with the objective of identifying the appropriate morphological interpretation of word forms within context. A contextual language model tailored for Kazakh is refined for token-level morphological tagging utilizing a manually validated annotated corpus of news articles. The suggested method utilizes self-attention mechanisms to capture long-range contextual dependencies that are challenging to represent with conventional rule­based or recurrent neural techniques. The experimental assessment reveals that the transformer-based model attains superior accuracy and F1-score relative to rule-based morphological analyzers and BiLSTM-based benchmarks. The findings demonstrate that contextualized embeddings significantly enhance the resolution of morphological ambiguity, especially with homonymous suffixes and infrequent grammatical structures. The results validate the efficacy of transformer topologies for low-resource agglutinative languages and establish a feasible basis for incorporating morphology-aware models into comprehensive Kazakh natural language processing frameworks.

About the Author

A. К. Aitim
International Information Technology University
Kazakhstan

PhD, associate professor

Almaty



References

1. Bach, M.P., Topalovic, A., Krstic, Z., and Ivec, A. Predictive maintenance in industry 4.0 for the SMEs: A decision support system case study using open-source software. Designs, 7, 98 (2023). https://doi.org/10.3390/designs7040098

2. Aitim, A., and Abdulla, M. Data Processing and Analysing Techniques in UX Research. Procedia Computer Science, 251, 591–596 (2024). https://doi.org/10.1016/j.procs.2024.11.154

3. Aitim, A., Sattarkhuzhayeva, D., and Khairullayeva, A. Development of a hybrid CNN-RNN model for enhanced recognition of dynamic gestures in Kazakh Sign Language. Eastern-European Journal of Enterprise Technologies, 2 (2 (134)), 58–67 (2025). https://doi.org/10.15587/1729-4061.2025.315834

4. Aitim, A. Building a high-quality annotated corpus for Kazakh NLP: a pipeline approach. Bulletin KazUTB, 4 (29) (2025). https://doi.org/10.58805/kazutb.v.4.29-1092

5. Aitim, A., and Satybaldiyeva, R. A comparison of Kazakh language processing models for improving semantic search results. Eastern-European Journal of Enterprise Technologies, 1 (2 (133)), 66–75 (2025). https://doi.org/10.15587/1729-4061.2025.315954

6. Aitim, A. Developing methods for automatic processing systems of Kazakh language. KazATC Bulletin, 133 (4), 254–265 (2024). https://doi.org/10.52167/1609-1817-2024-133-4-254-265

7. QNLP – Full Kazakh NLP Suite GitHub repository. https://github.com/Aigerimhub/qnlp

8. Ali, A., and Gravino, C. Improving software effort estimation using bio-inspired algorithms to select relevant features: an empirical study. Science of Computer Programming, 205, 102621 (2021). https://doi.org/10.1016/j.scico.2021.102621

9. Singh, K., and Gupta, P. Explainable artificial intelligence for software effort estimation: a survey and future directions. Information and Software Technology, 140, 106748 (2021). https://doi.org/10.1016/j.infsof.2021.106748

10. Khan, J.A., and Khan, S.U.R. Empirical investigation about the factors affecting the cost estimation in global software development context. IEEE Access, 9, 22274–22294 (2021). https://doi.org/10.1109/ACCESS.2021.3055858

11. Srivastava, D.K., Sharma, A.K., and Choudhary, D. Software development effort estimation using machine learning techniques: multi-linear regression versus random forest. 2021 International Conference on Computing, Communication and Green Engineering (CCGE) (2021), pp. 1–5. https://doi.org/10.1109/CCGE50943.2021.9776394

12. Alsaadi, M., and Saeedi, K. Agile effort estimation based on user stories: a systematic literature review. Artificial Intelligence Review, 55 (7), 5485–5516 (2022). https://doi.org/10.1007/s10462-021-10132-x

13. Matsubara, P.G.F. SEXTAMT: a systematic map to navigate the wide seas of factors affecting expert judgment software estimates. Journal of Systems and Software, 185, 111148 (2022). https://doi.org/10.1016/j.jss.2021.111148

14. Fávero, E.M.D.B. SE3M: a model for software effort estimation using pre-trained embedding models. Information and Software Technology, 147, 106886 (2022). https://doi.org/10.1016/j.infsof.2022.106886

15. Li, X., Zhao, H., and Yu, M. Hybrid deep learning models for software cost prediction using CNN and LSTM. Journal of Systems and Software, 188, 111282 (2022). https://doi.org/10.1016/j.jss.2022.111282

16. Rosa, C.C., and Jardine, D.A. Data-driven agile software cost estimation models for DHS and DoD. Journal of Systems and Software, 203, 111739 (2023). https://doi.org/10.1016/j.jss.2023.111739


Review

For citations:


Aitim A.К. MORPHOLOGICAL DISAMBIGUATION FOR THE KAZAKH LANGUAGE USING TRANSFORMER-BASED MODELS. Herald of the Kazakh-British Technical University. 2026;23(3):233-242. (In Russ.) https://doi.org/10.55452/1998-6688-2026-23-3-233-242

Views: 3

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 1998-6688 (Print)
ISSN 2959-8109 (Online)