МОРФОЛОГИЧЕСКАЯ ДИЗАМБИГУАЦИЯ КАЗАХСКОГО ЯЗЫКА С ИСПОЛЬЗОВАНИЕМ ТРАНСФОРМЕРНЫХ МОДЕЛЕЙ
https://doi.org/10.55452/1998-6688-2026-23-3-233-242
Аннотация
Морфологическая неоднозначность представляет собой значимую проблему для обработки естественного языка в агглютинативных языках, поскольку одна и та же словоформа может включать множество грамматических категорий. Казахский язык характеризуется продуктивной суффиксацией, сингармонизмом и сложными морфофонологическими закономерностями, что существенно затрудняет автоматический морфологический анализ. В данной работе представлена трансформер-ориентированная методология морфологической дизамбигуации казахских текстов, направленная на выбор корректной морфологической интерпретации словоформы в контексте. Контекстная языковая модель, адаптированная для казахского языка, дообучается для токен-уровневой морфологической разметки с использованием вручную верифицированного аннотированного корпуса новостных статей. Предложенный метод применяет механизмы самовнимания (self-attention) для захвата дальних контекстных зависимостей, которые сложно эффективно представить традиционными правилами или рекуррентными нейронными подходами. Экспериментальная оценка показывает, что трансформерная модель достигает более высокой точности и Fl-меры по сравнению с правилами морфологическими анализаторами и базовыми моделями на основе BiLSTM. Результаты демонстрируют, что контекстуализированные эмбеддинги заметно улучшают разрешение морфологической неоднозначности, особенно в случаях омонимичных суффиксов и редких грамматических структур. Полученные выводы подтверждают эффективность трансформерных архитектур для низкоресурсных агглютинативных языков и формируют практическую основу для интеграции морфологически чувствительных моделей в комплексные системы обработки казахского языка.
Об авторе
Э. Х. ЭйпмКазахстан
PhD, ассоциированный профессор
Алматы
Список литературы
1. Bach, M.P., Topalovic, A., Krstic, Z., and Ivec, A. Predictive maintenance in industry 4.0 for the SMEs: A decision support system case study using open-source software. Designs, 7, 98 (2023). https://doi.org/10.3390/designs7040098
2. Aitim, A., and Abdulla, M. Data Processing and Analysing Techniques in UX Research. Procedia Computer Science, 251, 591–596 (2024). https://doi.org/10.1016/j.procs.2024.11.154
3. Aitim, A., Sattarkhuzhayeva, D., and Khairullayeva, A. Development of a hybrid CNN-RNN model for enhanced recognition of dynamic gestures in Kazakh Sign Language. Eastern-European Journal of Enterprise Technologies, 2 (2 (134)), 58–67 (2025). https://doi.org/10.15587/1729-4061.2025.315834
4. Aitim, A. Building a high-quality annotated corpus for Kazakh NLP: a pipeline approach. Bulletin KazUTB, 4 (29) (2025). https://doi.org/10.58805/kazutb.v.4.29-1092
5. Aitim, A., and Satybaldiyeva, R. A comparison of Kazakh language processing models for improving semantic search results. Eastern-European Journal of Enterprise Technologies, 1 (2 (133)), 66–75 (2025). https://doi.org/10.15587/1729-4061.2025.315954
6. Aitim, A. Developing methods for automatic processing systems of Kazakh language. KazATC Bulletin, 133 (4), 254–265 (2024). https://doi.org/10.52167/1609-1817-2024-133-4-254-265
7. QNLP – Full Kazakh NLP Suite GitHub repository. https://github.com/Aigerimhub/qnlp
8. Ali, A., and Gravino, C. Improving software effort estimation using bio-inspired algorithms to select relevant features: an empirical study. Science of Computer Programming, 205, 102621 (2021). https://doi.org/10.1016/j.scico.2021.102621
9. Singh, K., and Gupta, P. Explainable artificial intelligence for software effort estimation: a survey and future directions. Information and Software Technology, 140, 106748 (2021). https://doi.org/10.1016/j.infsof.2021.106748
10. Khan, J.A., and Khan, S.U.R. Empirical investigation about the factors affecting the cost estimation in global software development context. IEEE Access, 9, 22274–22294 (2021). https://doi.org/10.1109/ACCESS.2021.3055858
11. Srivastava, D.K., Sharma, A.K., and Choudhary, D. Software development effort estimation using machine learning techniques: multi-linear regression versus random forest. 2021 International Conference on Computing, Communication and Green Engineering (CCGE) (2021), pp. 1–5. https://doi.org/10.1109/CCGE50943.2021.9776394
12. Alsaadi, M., and Saeedi, K. Agile effort estimation based on user stories: a systematic literature review. Artificial Intelligence Review, 55 (7), 5485–5516 (2022). https://doi.org/10.1007/s10462-021-10132-x
13. Matsubara, P.G.F. SEXTAMT: a systematic map to navigate the wide seas of factors affecting expert judgment software estimates. Journal of Systems and Software, 185, 111148 (2022). https://doi.org/10.1016/j.jss.2021.111148
14. Fávero, E.M.D.B. SE3M: a model for software effort estimation using pre-trained embedding models. Information and Software Technology, 147, 106886 (2022). https://doi.org/10.1016/j.infsof.2022.106886
15. Li, X., Zhao, H., and Yu, M. Hybrid deep learning models for software cost prediction using CNN and LSTM. Journal of Systems and Software, 188, 111282 (2022). https://doi.org/10.1016/j.jss.2022.111282
16. Rosa, C.C., and Jardine, D.A. Data-driven agile software cost estimation models for DHS and DoD. Journal of Systems and Software, 203, 111739 (2023). https://doi.org/10.1016/j.jss.2023.111739
Рецензия
Для цитирования:
Эйпм Э.Х. МОРФОЛОГИЧЕСКАЯ ДИЗАМБИГУАЦИЯ КАЗАХСКОГО ЯЗЫКА С ИСПОЛЬЗОВАНИЕМ ТРАНСФОРМЕРНЫХ МОДЕЛЕЙ. Вестник Казахстанско-Британского технического университета. 2026;23(3):233-242. https://doi.org/10.55452/1998-6688-2026-23-3-233-242
For citation:
Aitim A.К. MORPHOLOGICAL DISAMBIGUATION FOR THE KAZAKH LANGUAGE USING TRANSFORMER-BASED MODELS. Herald of the Kazakh-British Technical University. 2026;23(3):233-242. (In Russ.) https://doi.org/10.55452/1998-6688-2026-23-3-233-242
JATS XML






