MORPHOLOGICAL DISAMBIGUATION FOR THE KAZAKH LANGUAGE USING TRANSFORMER-BASED MODELS
https://doi.org/10.55452/1998-6688-2026-23-3-233-242
Abstract
Morphological ambiguity constitutes a significant challenge for natural language processing in agglutinative languages, as a single word form might include many grammatical categories. The Kazakh language features productive suffixation, vowel harmony, and intricate morphophonological patterns, which considerably hinder automatic morphological analysis. This work presents a transformer-based methodology for morphological disambiguation in Kazakh texts, with the objective of identifying the appropriate morphological interpretation of word forms within context. A contextual language model tailored for Kazakh is refined for token-level morphological tagging utilizing a manually validated annotated corpus of news articles. The suggested method utilizes self-attention mechanisms to capture long-range contextual dependencies that are challenging to represent with conventional rulebased or recurrent neural techniques. The experimental assessment reveals that the transformer-based model attains superior accuracy and F1-score relative to rule-based morphological analyzers and BiLSTM-based benchmarks. The findings demonstrate that contextualized embeddings significantly enhance the resolution of morphological ambiguity, especially with homonymous suffixes and infrequent grammatical structures. The results validate the efficacy of transformer topologies for low-resource agglutinative languages and establish a feasible basis for incorporating morphology-aware models into comprehensive Kazakh natural language processing frameworks.
Keywords
About the Author
A. К. AitimKazakhstan
PhD, associate professor
Almaty
References
1. Bach, M.P., Topalovic, A., Krstic, Z., and Ivec, A. Predictive maintenance in industry 4.0 for the SMEs: A decision support system case study using open-source software. Designs, 7, 98 (2023). https://doi.org/10.3390/designs7040098
2. Aitim, A., and Abdulla, M. Data Processing and Analysing Techniques in UX Research. Procedia Computer Science, 251, 591–596 (2024). https://doi.org/10.1016/j.procs.2024.11.154
3. Aitim, A., Sattarkhuzhayeva, D., and Khairullayeva, A. Development of a hybrid CNN-RNN model for enhanced recognition of dynamic gestures in Kazakh Sign Language. Eastern-European Journal of Enterprise Technologies, 2 (2 (134)), 58–67 (2025). https://doi.org/10.15587/1729-4061.2025.315834
4. Aitim, A. Building a high-quality annotated corpus for Kazakh NLP: a pipeline approach. Bulletin KazUTB, 4 (29) (2025). https://doi.org/10.58805/kazutb.v.4.29-1092
5. Aitim, A., and Satybaldiyeva, R. A comparison of Kazakh language processing models for improving semantic search results. Eastern-European Journal of Enterprise Technologies, 1 (2 (133)), 66–75 (2025). https://doi.org/10.15587/1729-4061.2025.315954
6. Aitim, A. Developing methods for automatic processing systems of Kazakh language. KazATC Bulletin, 133 (4), 254–265 (2024). https://doi.org/10.52167/1609-1817-2024-133-4-254-265
7. QNLP – Full Kazakh NLP Suite GitHub repository. https://github.com/Aigerimhub/qnlp
8. Ali, A., and Gravino, C. Improving software effort estimation using bio-inspired algorithms to select relevant features: an empirical study. Science of Computer Programming, 205, 102621 (2021). https://doi.org/10.1016/j.scico.2021.102621
9. Singh, K., and Gupta, P. Explainable artificial intelligence for software effort estimation: a survey and future directions. Information and Software Technology, 140, 106748 (2021). https://doi.org/10.1016/j.infsof.2021.106748
10. Khan, J.A., and Khan, S.U.R. Empirical investigation about the factors affecting the cost estimation in global software development context. IEEE Access, 9, 22274–22294 (2021). https://doi.org/10.1109/ACCESS.2021.3055858
11. Srivastava, D.K., Sharma, A.K., and Choudhary, D. Software development effort estimation using machine learning techniques: multi-linear regression versus random forest. 2021 International Conference on Computing, Communication and Green Engineering (CCGE) (2021), pp. 1–5. https://doi.org/10.1109/CCGE50943.2021.9776394
12. Alsaadi, M., and Saeedi, K. Agile effort estimation based on user stories: a systematic literature review. Artificial Intelligence Review, 55 (7), 5485–5516 (2022). https://doi.org/10.1007/s10462-021-10132-x
13. Matsubara, P.G.F. SEXTAMT: a systematic map to navigate the wide seas of factors affecting expert judgment software estimates. Journal of Systems and Software, 185, 111148 (2022). https://doi.org/10.1016/j.jss.2021.111148
14. Fávero, E.M.D.B. SE3M: a model for software effort estimation using pre-trained embedding models. Information and Software Technology, 147, 106886 (2022). https://doi.org/10.1016/j.infsof.2022.106886
15. Li, X., Zhao, H., and Yu, M. Hybrid deep learning models for software cost prediction using CNN and LSTM. Journal of Systems and Software, 188, 111282 (2022). https://doi.org/10.1016/j.jss.2022.111282
16. Rosa, C.C., and Jardine, D.A. Data-driven agile software cost estimation models for DHS and DoD. Journal of Systems and Software, 203, 111739 (2023). https://doi.org/10.1016/j.jss.2023.111739
Review
For citations:
Aitim A.К. MORPHOLOGICAL DISAMBIGUATION FOR THE KAZAKH LANGUAGE USING TRANSFORMER-BASED MODELS. Herald of the Kazakh-British Technical University. 2026;23(3):233-242. (In Russ.) https://doi.org/10.55452/1998-6688-2026-23-3-233-242
JATS XML






