<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">kaz44</journal-id><journal-title-group><journal-title xml:lang="ru">Вестник Университета Шакарима. Серия технические науки</journal-title><trans-title-group xml:lang="en"><trans-title>Bulletin of Shakarim University. Technical Sciences</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2788-7995</issn><issn pub-type="epub">3006-0524</issn><publisher><publisher-name>«Шәкәрім университеті» КеАҚ</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.53360/2788-7995-2026-2(22)-14</article-id><article-id custom-type="elpub" pub-id-type="custom">kaz44-2571</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>АВТОМАТИЗАЦИЯ И ИНФОРМАЦИОННЫЕ ТЕХНОЛОГИИ</subject></subj-group></article-categories><title-group><article-title>ФОРМАЛЬНЫЕ И НЕЙРОННЫЕ ПОДХОДЫ К МОРФОЛОГИЧЕСКОМУ АНАЛИЗУ В АГГЛЮТИНАТИВНЫХ ЯЗЫКАХ: ДАННЫЕ КАЗАХСКОГО ЯЗЫКА</article-title><trans-title-group xml:lang="en"><trans-title>FORMAL AND NEURAL APPROACHES TO MORPHOLOGICAL ANALYSIS IN AGGLUTINATIVE LANGUAGES: EVIDENCE FROM KAZAKH</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2982-214X</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Әйтім</surname><given-names>Ә. Қ.</given-names></name><name name-style="western" xml:lang="en"><surname>Aitim</surname><given-names>A. K.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Әйгерім Әйтім – PhD, ассоциированный профессор кафедры «Информационные системы», </p><p>050040, Алматы, ул.Манас, 34/1</p></bio><bio xml:lang="en"><p>Aigerim Aitim – PhD, associate-professor of Information Systems Department, </p><p>050040, Almaty, Manas street, 34/1</p></bio><email xlink:type="simple">a.aitim@iitu.edu.kz</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Международный Университет Информационных Технологий</institution><country>Казахстан</country></aff><aff xml:lang="en"><institution>International Information Technology University</institution><country>Kazakhstan</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>29</day><month>07</month><year>2026</year></pub-date><volume>0</volume><issue>2(22)</issue><fpage>132</fpage><lpage>141</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Әйтім Ә.Қ., 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Әйтім Ә.Қ.</copyright-holder><copyright-holder xml:lang="en">Aitim A.K.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://tech.vestnik.shakarim.kz/jour/article/view/2571">https://tech.vestnik.shakarim.kz/jour/article/view/2571</self-uri><abstract><p>Автоматический морфологический анализ традиционно представляет значительную сложность для агглютинативных языков вследствие их богатой словоизменительной системы, продуктивного словообразования и сложных морфофонологических правил. Современные нейронные модели, особенно архитектуры на основе трансформеров, демонстрируют высокие эмпирические показатели. Однако они нередко характеризуются недостаточной лингвистической прозрачностью и сталкиваются с трудностями систематического обобщения в условиях ограниченных ресурсов. В данной работе предлагается сравнительное и интегративное исследование формальных (правил-ориентированных и конечных автоматов) и нейронных подходов к морфологическому анализу на примере казахского языка как агглютинативного языка с низкими ресурсами. Вначале представляется формальная морфологическая модель, явно описывающая структуру «корень–аффикс», сингармонизм и морфотактические ограничения. Далее исследуются различные нейронные архитектуры для морфологической дизамбигуации и разметки, включая KazBERT в сочетании с декодированием на основе условных случайных полей (CRF). Помимо стандартных метрик точности проводится детальная типология ошибок и лингвистический анализ, в рамках которых изучается, как различные классы моделей справляются с неоднозначностью, редкими словоформами и длинными цепочками аффиксов. Результаты показывают, что, хотя нейронные модели превосходят исключительно правил-ориентированные системы по поверхностной точности, они демонстрируют устойчивые недостатки при обработке морфологически сложных и редких конструкций. Формальные модели, в свою очередь, обеспечивают более надёжное обобщение за счёт языковых ограничений. На основе полученных результатов предлагается гибридная морфологически осведомлённая архитектура, интегрирующая символические ограничения в нейронный вывод. Данный подход обеспечивает устойчивое улучшение результатов в различных условиях оценки. Работа демонстрирует, что эффективный морфологический анализ агглютинативных языков требует сочетания нейронного обучения представлений и явного лингвистического моделирования. Полученные выводы не ограничиваются одним языком и имеют более широкие последствия для морфологически чувствительного NLP в низкоресурсных условиях.</p></abstract><trans-abstract xml:lang="en"><p>Automatic morphological analysis remains a challenging task for agglutinative languages because of their rich inflectional systems, productive derivation, and complex morphophonological rules. Recent neural models, especially transformer-based architectures, have demonstrated impressive empirical performance. However, they frequently exhibit a deficiency in linguistic transparency and encounter challenges in systematic generalization inside low-resource environments. This research offers a comparative and integrative examination of formal (rule-based and finite-state) and neural (KazBERT-based) methodologies for morphological analysis, utilizing the Kazakh language as a case study of low-resource agglutinative morphology. Initially present a formal morphological model that distinctly represents root-affix structure, vowel harmony, and morphotactic restrictions. We next test many neural architectures for morphological disambiguation and tagging, such as KazBERT coupled with CRF-based decoding. In addition to typical accuracy measurements, we do a comprehensive error taxonomy and linguistic analysis, investigating how various model classes manage ambiguity, infrequent forms, and extended affix chains. The findings indicate that whereas neural models excel in surface-level accuracy compared to exclusively rule-based systems, they demonstrate consistent deficiencies in morphologically intricate and infrequent constructs. On the other hand, formal models show better generalization based on language limitations. Based on these results, we suggest a hybrid morphology-aware framework that adds symbolic restrictions to neural inference. This framework consistently improves results in a variety of assessment contexts. The study demonstrates that effective morphological analysis of agglutinative languages requires the integration of neural representation learning with explicit linguistic structure. The results are not tied to any one language and have wider implications for morphology-sensitive NLP in low-resource settings.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>вычислительная морфология</kwd><kwd>агглютинативные языки</kwd><kwd>морфологический анализ</kwd><kwd>языки с ограниченными ресурсами</kwd><kwd>формальные лингвистические модели</kwd><kwd>нейронные языковые модели</kwd></kwd-group><kwd-group xml:lang="en"><kwd>computational morphology</kwd><kwd>agglutinative languages</kwd><kwd>morphological analysis</kwd><kwd>low-resource languages</kwd><kwd>formal linguistic models</kwd><kwd>neural language models</kwd></kwd-group><funding-group><funding-statement xml:lang="en">This work was supported by the Ministry of Culture and Information of the Republic of Kazakhstan of grant "Tauelsizdik Urpaktary-2025", project named by “QazNLP is an open-source scientific system for intelligent processing of Kazakh-language text”.</funding-statement></funding-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Marquard C. Neural morphological tagging for Nguni languages. In Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025) / C. Marquard, S. Mawere, F. Meyer // Association for Computational Linguistics. – 2025. – Р. 210-220. https://aclanthology.org/2025.africanlp-1.31.</mixed-citation><mixed-citation xml:lang="en">Marquard C. Neural morphological tagging for Nguni languages. In Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025) / C. Marquard, S. Mawere, F. Meyer // Association for Computational Linguistics. – 2025. – Р. 210-220. https://aclanthology.org/2025.africanlp-1.31.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Stenlund M. Surface-level morphological segmentation of low-resource Inuktitut using pre-trained large language models. In Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025) / M. Stenlund, H. Myneni, M. Riedel // University of Tartu Library. – 2025. – Р. 688-696. https://aclanthology.org/2025.nodalida-1.69.</mixed-citation><mixed-citation xml:lang="en">Stenlund M. Surface-level morphological segmentation of low-resource Inuktitut using pre-trained large language models. In Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025) / M. Stenlund, H. Myneni, M. Riedel // University of Tartu Library. – 2025. – Р. 688-696. https://aclanthology.org/2025.nodalida-1.69.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Research on morphological knowledge-guided low-resource agglutinative languages-Chinese translation / G. Abudouwaili et al // Complex &amp; Intelligent Systems – 2025. https://doi.org/10.1007/s40747-025-01780-5.</mixed-citation><mixed-citation xml:lang="en">Research on morphological knowledge-guided low-resource agglutinative languages-Chinese translation / G. Abudouwaili et al // Complex &amp; Intelligent Systems – 2025. https://doi.org/10.1007/s40747-025-01780-5.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Developing a hybrid morphological analyzer for low-resource languages / M. Supriya et al // Applied Sciences. – 2025. – № 15(10). – Р. 5682. https://doi.org/10.3390/app15105682.</mixed-citation><mixed-citation xml:lang="en">Developing a hybrid morphological analyzer for low-resource languages / M. Supriya et al // Applied Sciences. – 2025. – № 15(10). – Р. 5682. https://doi.org/10.3390/app15105682.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Baitenova L. Hybrid artificial intelligence architectures for automatic text analysis and computational morphology / L. Baitenova // Frontiers in Artificial Intelligence. – 2025. – № 8. – Р. 1708566. https://doi.org/10.3389/frai.2025.1708566.</mixed-citation><mixed-citation xml:lang="en">Baitenova L. Hybrid artificial intelligence architectures for automatic text analysis and computational morphology / L. Baitenova // Frontiers in Artificial Intelligence. – 2025. – № 8. – Р. 1708566. https://doi.org/10.3389/frai.2025.1708566.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Yazar B.K. Improving low-resource Kazakh-English and Turkish-English neural machine translation using transfer learning and part-of-speech tags / B.K. Yazar, E. Kiliç, // IEEE Access. – 2025. – № 13. – Р. 32341-32356. https://doi.org/10.1109/ACCESS.2025.3542491.</mixed-citation><mixed-citation xml:lang="en">Yazar B.K. Improving low-resource Kazakh-English and Turkish-English neural machine translation using transfer learning and part-of-speech tags / B.K. Yazar, E. Kiliç, // IEEE Access. – 2025. – № 13. – Р. 32341-32356. https://doi.org/10.1109/ACCESS.2025.3542491.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Belth C. Meaning-informed low-resource segmentation of agglutinative morphology. In Proceedings of SCiL 2024. ACL Anthology.</mixed-citation><mixed-citation xml:lang="en">Belth C. Meaning-informed low-resource segmentation of agglutinative morphology. In Proceedings of SCiL 2024. ACL Anthology.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Aitim A. Developing methods for automatic processing systems of Kazakh language / A. Aitim // KazATC Bulletin. – 2024. – № 133(4). – Р. 254-265. https://doi.org/10.52167/1609-1817-2024-133-4-254-265.</mixed-citation><mixed-citation xml:lang="en">Aitim A. Developing methods for automatic processing systems of Kazakh language / A. Aitim // KazATC Bulletin. – 2024. – № 133(4). – Р. 254-265. https://doi.org/10.52167/1609-1817-2024-133-4-254-265.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Aitim A. Building a high-quality annotated corpus for Kazakh NLP: a pipeline approach / A. Aitim // Vestnik KazUTB. – 2025. – vol. 4, № 29. https://doi.org/10.58805/kazutb.v.4.29-1092.</mixed-citation><mixed-citation xml:lang="en">Aitim A. Building a high-quality annotated corpus for Kazakh NLP: a pipeline approach / A. Aitim // Vestnik KazUTB. – 2025. – vol. 4, № 29. https://doi.org/10.58805/kazutb.v.4.29-1092.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">QNLP – Full Kazakh NLP Suite GitHub repository. Retrieved from https://github.com/Aigerimhub/qnlp.</mixed-citation><mixed-citation xml:lang="en">QNLP – Full Kazakh NLP Suite GitHub repository. Retrieved from https://github.com/Aigerimhub/qnlp.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">A. Aitim QazNLP: Constraint-Aware Multi-Task Sequence Labeling for Morphologically Rich Low-Resource Languages / A. Aitim // in IEEE Access. – 2026. – vol. 14. – Р. 70955-70974. https://doi.org/10.1109/ACCESS.2026.3691193.</mixed-citation><mixed-citation xml:lang="en">A. Aitim QazNLP: Constraint-Aware Multi-Task Sequence Labeling for Morphologically Rich Low-Resource Languages / A. Aitim // in IEEE Access. – 2026. – vol. 14. – Р. 70955-70974. https://doi.org/10.1109/ACCESS.2026.3691193.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Aitim A. LLM-Assisted Weak Supervision for Low-Resource Kazakh Sequence Labeling: Synthetic Annotation and CRF-Refined NER/POS Models / A. Aitim // Applied Sciences. – 2026. – № 16(8). – Р. 3632. https://doi.org/10.3390/app16083632.</mixed-citation><mixed-citation xml:lang="en">Aitim A. LLM-Assisted Weak Supervision for Low-Resource Kazakh Sequence Labeling: Synthetic Annotation and CRF-Refined NER/POS Models / A. Aitim // Applied Sciences. – 2026. – № 16(8). – Р. 3632. https://doi.org/10.3390/app16083632.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
