<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">kaz44</journal-id><journal-title-group><journal-title xml:lang="ru">Вестник Университета Шакарима. Серия технические науки</journal-title><trans-title-group xml:lang="en"><trans-title>Bulletin of Shakarim University. Technical Sciences</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2788-7995</issn><issn pub-type="epub">3006-0524</issn><publisher><publisher-name>«Шәкәрім университеті» КеАҚ</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.53360/2788-7995-2026-1(21)-4</article-id><article-id custom-type="elpub" pub-id-type="custom">kaz44-2291</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>АВТОМАТИЗАЦИЯ И ИНФОРМАЦИОННЫЕ ТЕХНОЛОГИИ (ОРИГИНАЛЬНАЯ СТАТЬЯ)</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>AUTOMATION AND INFORMATION TECHNOLOGY (ORIGINAL ARTICLE)</subject></subj-group></article-categories><title-group><article-title>УНИВЕРСАЛЬНЫЙ ТРАНСФОРМЕРНЫЙ ФРЕЙМВОРК ДЛЯ АНАЛИЗА И ГЕНЕРАЦИИ ТЕКСТА</article-title><trans-title-group xml:lang="en"><trans-title>UNIVERSAL TRANSFORMER FRAMEWORK FOR TEXT ANALYSIS AND GENERATION</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Ахмедиярова</surname><given-names>А. Т.</given-names></name><name name-style="western" xml:lang="en"><surname>Akhmediarova</surname><given-names>A. T.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Айнур Танатаровна Ахмедиярова – PhD, профессор Института автоматики и информационных технологий, кафедра «Кибербезопасности, обработки и хранения информации» </p><p>050013, Алматы қаласы, Сатпаев көшесі, 22</p></bio><bio xml:lang="en"><p>Ainur Tanatarovna Akhmediarova – PhD, professor at the Institute of Automation and Information Technology, Department of Cybersecurity, Information Processing and Storage </p><p>050013, Almaty, Satpayev 22</p></bio><email xlink:type="simple">a.akhmediyarova@satbayev.university</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-6388-9458</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Аяпбергенова</surname><given-names>А. Т.</given-names></name><name name-style="western" xml:lang="en"><surname>Ayapbergenova</surname><given-names>A. T.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Асем Тултановна Аяпбергенова– магистр техники и технологии, старший преподаватель Института автоматики и информационных технологий, кафедра «Программной инженерии» </p><p>050013, Алматы қаласы, Сатпаев көшесі, 22</p></bio><bio xml:lang="en"><p>Asem Tultanovna Ayapbergenova– master of Engineering and Technology, Senior Lecturer at the Institute of Automation and Information Technology, Department of Software Engineering </p><p>050013, Almaty, Satpayev 22</p></bio><email xlink:type="simple">asem_007800@inbox.ru</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-9565-5621</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Алибиева</surname><given-names>Ж. М.</given-names></name><name name-style="western" xml:lang="en"><surname>Alibiyeva</surname><given-names>Zh. M.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Жибек Мейрамбековна Алибиева – PhD, ассоциированный профессор, Института автоматики и информационных технологий, кафедра «Программной инженерии» </p><p>050013, Алматы қаласы, Сатпаев көшесі, 22</p></bio><bio xml:lang="en"><p>Zhibek Meirambekovna Alibieva – PhD, associate professor, Institute of Automation and Information Technology, Department of Software Engineering </p><p>050013, Almaty, Satpayev 22</p></bio><email xlink:type="simple">zh.alibiyeva@satbayev.university</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-4835-5751</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Мукажанов</surname><given-names>Н. К.</given-names></name><name name-style="western" xml:lang="en"><surname>Мukazhanov</surname><given-names>N. K.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Нуржан Какенович Мукажанов – PhD., ассоциированный профессор, Института автоматики и информационных технологий, кафедра «Программной инженерии» </p><p>050013, Алматы қаласы, Сатпаев көшесі, 22</p></bio><bio xml:lang="en"><p>Nurzhan Kakenovich Mukazhanov – PhD, Associate Professor, Institute of Automation and Information Technology, Department of Software Engineering </p><p>050013, Almaty, Satpayev 22</p></bio><email xlink:type="simple">n.mukazhanov@satbayev.university</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2900-8025</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Исабеков</surname><given-names>Ж. Н.</given-names></name><name name-style="western" xml:lang="en"><surname>Issabekov</surname><given-names>Zh. N.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Жанибек Назарбекулы Исабеков – PhD, ассоциированный профессор кафедры «Робототехники и технических средств автоматики» </p><p>050013, Алматы қаласы, Сатпаев көшесі, 22</p></bio><bio xml:lang="en"><p>Zhanibek Issabekov – PhD, Associate Professor of the Department of Robotics and Engineering Tools of Automation </p><p>050013, Almaty, Satpayev 22</p></bio><email xlink:type="simple">z.issabekov@satbayev.university</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Satbayev University</institution><country>Казахстан</country></aff><aff xml:lang="en"><institution>Satbayev University</institution><country>Kazakhstan</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>25</day><month>05</month><year>2026</year></pub-date><volume>1</volume><issue>1(21)</issue><fpage>36</fpage><lpage>46</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Ахмедиярова А.Т., Аяпбергенова А.Т., Алибиева Ж.М., Мукажанов Н.К., Исабеков Ж.Н., 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Ахмедиярова А.Т., Аяпбергенова А.Т., Алибиева Ж.М., Мукажанов Н.К., Исабеков Ж.Н.</copyright-holder><copyright-holder xml:lang="en">Akhmediarova A.T., Ayapbergenova A.T., Alibiyeva Z.M., Мukazhanov N.K., Issabekov Z.N.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://tech.vestnik.shakarim.kz/jour/article/view/2291">https://tech.vestnik.shakarim.kz/jour/article/view/2291</self-uri><abstract><p>В статье предложен универсальный мультимодальный фреймворк для решения трёх задач обработки естественного языка: распознавания сарказма, сентиментноориентированной генерации описаний изображений и контекстного нейросетевого машинного перевода с использованием механизма reranking. Архитектура включает специализированные энкодеры (BERT, RoBERTa, ViT), кросс-модальные механизмы внимания и контрастивное обучение, что обеспечивает адаптацию к различным типам входных данных и семантических задач. Эксперименты продемонстрировали улучшение метрик BLEU, METEOR, F1 и NER по сравнению с базовыми моделями. Отдельное внимание уделено устойчивости модели при работе с редкими словами и именованными сущностями. Полученные результаты подтверждают эффективность предложенного подхода в условиях ограниченных данных и мультимодальной сложности. В условиях стремительного роста мультимедийного контента и экспрессии пользователей в социальных сетях, мультимодальные модели – объединяющие текст, визуальные и стилистические признаки – становятся ключевым направлением в развитии когнитивных ИИ-систем. В частности, трансформерные архитектуры открывают новые горизонты в интеграции многоканальной информации, позволяя учитывать не только прямой смысл текста, но и эмоциональный подтекст, визуальный контекст и стилистические особенности подачи. Современные интеллектуальные системы обработки языка всё чаще сталкиваются с необходимостью учитывать не только семантический и синтаксический контекст, но и эмоциональную окраску высказываний. Традиционные NLP-модели, ориентированные на буквальное понимание текста, демонстрируют ограниченную эффективность в задачах, где ключевым фактором становится скрытый смысл – как, например, в распознавании сарказма, генерации описаний изображений с учётом настроения, или переводе эмоционально окрашенных конструкций между языками.</p></abstract><trans-abstract xml:lang="en"><p>The article proposes a universal multimodal framework for solving three tasks of natural language processing: sarcasm recognition, sentiment-oriented generation of image descriptions, and contextual neural network machine translation using the reranking mechanism. The architecture includes specialized encoders (BERT, RoBERTa, ViT), cross-modal attention mechanisms, and contrastive learning, which provides adaptation to various types of input data and semantic tasks. The experiments demonstrated improvements in the BLEU, METEOR, F1, and NER metrics compared to the basic models. Special attention is paid to the stability of the model when working with rare words and named entities. The results obtained confirm the effectiveness of the proposed approach in conditions of limited data and multimodal complexity. In the context of the rapid growth of multimedia content and user expression on social networks, multimodal models combining text, visual and stylistic features are becoming a key direction in the development of cognitive AI systems. In particular, transformer architectures open up new horizons in the integration of multi-channel information, allowing us to take into account not only the direct meaning of the text, but also the emotional subtext, visual context and stylistic features of the presentation.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>мультимодальность</kwd><kwd>сарказм</kwd><kwd>сентимент-анализ</kwd><kwd>контрастивное обучение</kwd><kwd>трансформеры</kwd><kwd>ViT</kwd><kwd>BERT</kwd><kwd>низкоресурсные языки</kwd><kwd>кросс-модальный attention</kwd></kwd-group><kwd-group xml:lang="en"><kwd>multimodality</kwd><kwd>sarcasm</kwd><kwd>sentiment analysis</kwd><kwd>contrastive learning</kwd><kwd>transformers</kwd><kwd>ViT</kwd><kwd>BERT</kwd><kwd>low-resource languages</kwd><kwd>cross-modal attention</kwd></kwd-group><funding-group><funding-statement xml:lang="ru">исследование финансировалось Комитетом науки Министерства науки и высшего образования Республики Казахстан (ПЦФ № BR24993166 «Разработка комплексной инновационной онлайн-платформы, автоматизированной системы юридической помощи и единой системы автоматизации работы юристов»).</funding-statement></funding-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">A Contrastive Multimodal Representation Learning for Sarcasm / Alanoud Al Mazroa et al // Expert Systems With Applications. – 2025. – Vol. 298.</mixed-citation><mixed-citation xml:lang="en">A Contrastive Multimodal Representation Learning for Sarcasm / Alanoud Al Mazroa et al // Expert Systems With Applications. – 2025. – Vol. 298.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Optimizing Sentiment Integration in Image Captioning Using Transformer-Based Fusion Strategies / Komal Rani Narejo et al // Computers, Materials &amp; Continua. – 2025. – Vol. 84, №2. https://doi.org/10.32604/cmc.2025.065872.</mixed-citation><mixed-citation xml:lang="en">Optimizing Sentiment Integration in Image Captioning Using Transformer-Based Fusion Strategies / Komal Rani Narejo et al // Computers, Materials &amp; Continua. – 2025. – Vol. 84, №2. https://doi.org/10.32604/cmc.2025.065872.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Transformer-Based Re-Ranking Model for Enhancing Contextual and Syntactic Translation in Low-Resource Neural Machine Translation / A. Javed et al // Electronics. – 2025. – № 14(2). – Р. 243. https://doi.org/10.3390/electronics14020243.</mixed-citation><mixed-citation xml:lang="en">Transformer-Based Re-Ranking Model for Enhancing Contextual and Syntactic Translation in Low-Resource Neural Machine Translation / A. Javed et al // Electronics. – 2025. – № 14(2). – Р. 243. https://doi.org/10.3390/electronics14020243.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Zhang X. Multimodal sarcasm detection in Twitter with hierarchical fusion model / X. Zhang, H. Wang // Proceedings of the 58th ACL. – 2020. – Р. 2500-2505. https://doi.org/10.18653/v1/2020.aclmain.227.</mixed-citation><mixed-citation xml:lang="en">Zhang X. Multimodal sarcasm detection in Twitter with hierarchical fusion model / X. Zhang, H. Wang // Proceedings of the 58th ACL. – 2020. – Р. 2500-2505. https://doi.org/10.18653/v1/2020.aclmain.227.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Attention is all you need / А. Vaswani et al // Advances in Neural Information Processing Systems. – 2017. – Vol. 30. – Р. 5998-6008. https://doi.org/10.48550/arXiv.1706.03762.</mixed-citation><mixed-citation xml:lang="en">Attention is all you need / А. Vaswani et al // Advances in Neural Information Processing Systems. – 2017. – Vol. 30. – Р. 5998-6008. https://doi.org/10.48550/arXiv.1706.03762.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Tan H. LXMERT: Learning cross-modality encoder representations from transformers / H. Tan, M. Bansal // EMNLP. – 2019. – Р. 5103-5114. https://doi.org/10.18653/v1/D19-1514.</mixed-citation><mixed-citation xml:lang="en">Tan H. LXMERT: Learning cross-modality encoder representations from transformers / H. Tan, M. Bansal // EMNLP. – 2019. – Р. 5103-5114. https://doi.org/10.18653/v1/D19-1514.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Understanding back-translation at scale / S. Edunov et al // EMNLP. – 2018. – Р. 489-500. https://doi.org/10.18653/v1/D18-1150.</mixed-citation><mixed-citation xml:lang="en">Understanding back-translation at scale / S. Edunov et al // EMNLP. – 2018. – Р. 489-500. https://doi.org/10.18653/v1/D18-1150.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Hossain M.Z. Multimodal machine learning for emotion recognition: A review / M.Z. Hossain, G. Muhammad, M. Alsulaiman // Information Fusion. – 2021. – vol. 68. – Р. 21-39. https://doi.org/10.1016/j.inffus.2020.10.008.</mixed-citation><mixed-citation xml:lang="en">Hossain M.Z. Multimodal machine learning for emotion recognition: A review / M.Z. Hossain, G. Muhammad, M. Alsulaiman // Information Fusion. – 2021. – vol. 68. – Р. 21-39. https://doi.org/10.1016/j.inffus.2020.10.008.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Dabre R. Enabling multilingual neural machine translation with knowledge distillation / R. Dabre, C. Chu, S. Kurohashi // ACL. – 2017. – Р. 2662-2673. https://doi.org/10.18653/v1/P17-1243.</mixed-citation><mixed-citation xml:lang="en">Dabre R. Enabling multilingual neural machine translation with knowledge distillation / R. Dabre, C. Chu, S. Kurohashi // ACL. – 2017. – Р. 2662-2673. https://doi.org/10.18653/v1/P17-1243.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Wang R. Sentiment-aware image captioning with context disentangling / R. Wang, X. Wan, W. Li // CVPR. – 2021. – Р. 17586-17595. https://doi.org/10.1109/CVPR46437.2021.01732.</mixed-citation><mixed-citation xml:lang="en">Wang R. Sentiment-aware image captioning with context disentangling / R. Wang, X. Wan, W. Li // CVPR. – 2021. – Р. 17586-17595. https://doi.org/10.1109/CVPR46437.2021.01732.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">UNITER: Learning universal image-text representations / Y.-C. Chen et al // ECCV. – 2020. – Р. 104-120. https://doi.org/10.1007/978-3-030-58523-5_43.</mixed-citation><mixed-citation xml:lang="en">UNITER: Learning universal image-text representations / Y.-C. Chen et al // ECCV. – 2020. – Р. 104-120. https://doi.org/10.1007/978-3-030-58523-5_43.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">BERT: Pre-training of deep bidirectional transformers for language understanding / J. Devlin et al // NAACL-HLT. – 2019. – Р. 4171-4186. https://doi.org/10.18653/v1/N19-1423.</mixed-citation><mixed-citation xml:lang="en">BERT: Pre-training of deep bidirectional transformers for language understanding / J. Devlin et al // NAACL-HLT. – 2019. – Р. 4171-4186. https://doi.org/10.18653/v1/N19-1423.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
