UNIVERSAL TRANSFORMER FRAMEWORK FOR TEXT ANALYSIS AND GENERATION
https://doi.org/10.53360/2788-7995-2026-1(21)-4
Abstract
The article proposes a universal multimodal framework for solving three tasks of natural language processing: sarcasm recognition, sentiment-oriented generation of image descriptions, and contextual neural network machine translation using the reranking mechanism. The architecture includes specialized encoders (BERT, RoBERTa, ViT), cross-modal attention mechanisms, and contrastive learning, which provides adaptation to various types of input data and semantic tasks. The experiments demonstrated improvements in the BLEU, METEOR, F1, and NER metrics compared to the basic models. Special attention is paid to the stability of the model when working with rare words and named entities. The results obtained confirm the effectiveness of the proposed approach in conditions of limited data and multimodal complexity. In the context of the rapid growth of multimedia content and user expression on social networks, multimodal models combining text, visual and stylistic features are becoming a key direction in the development of cognitive AI systems. In particular, transformer architectures open up new horizons in the integration of multi-channel information, allowing us to take into account not only the direct meaning of the text, but also the emotional subtext, visual context and stylistic features of the presentation.
About the Authors
A. T. AkhmediarovaKazakhstan
Ainur Tanatarovna Akhmediarova – PhD, professor at the Institute of Automation and Information Technology, Department of Cybersecurity, Information Processing and Storage
050013, Almaty, Satpayev 22
A. T. Ayapbergenova
Kazakhstan
Asem Tultanovna Ayapbergenova– master of Engineering and Technology, Senior Lecturer at the Institute of Automation and Information Technology, Department of Software Engineering
050013, Almaty, Satpayev 22
Zh. M. Alibiyeva
Kazakhstan
Zhibek Meirambekovna Alibieva – PhD, associate professor, Institute of Automation and Information Technology, Department of Software Engineering
050013, Almaty, Satpayev 22
N. K. Мukazhanov
Kazakhstan
Nurzhan Kakenovich Mukazhanov – PhD, Associate Professor, Institute of Automation and Information Technology, Department of Software Engineering
050013, Almaty, Satpayev 22
Zh. N. Issabekov
Kazakhstan
Zhanibek Issabekov – PhD, Associate Professor of the Department of Robotics and Engineering Tools of Automation
050013, Almaty, Satpayev 22
References
1. A Contrastive Multimodal Representation Learning for Sarcasm / Alanoud Al Mazroa et al // Expert Systems With Applications. – 2025. – Vol. 298.
2. Optimizing Sentiment Integration in Image Captioning Using Transformer-Based Fusion Strategies / Komal Rani Narejo et al // Computers, Materials & Continua. – 2025. – Vol. 84, №2. https://doi.org/10.32604/cmc.2025.065872.
3. Transformer-Based Re-Ranking Model for Enhancing Contextual and Syntactic Translation in Low-Resource Neural Machine Translation / A. Javed et al // Electronics. – 2025. – № 14(2). – Р. 243. https://doi.org/10.3390/electronics14020243.
4. Zhang X. Multimodal sarcasm detection in Twitter with hierarchical fusion model / X. Zhang, H. Wang // Proceedings of the 58th ACL. – 2020. – Р. 2500-2505. https://doi.org/10.18653/v1/2020.aclmain.227.
5. Attention is all you need / А. Vaswani et al // Advances in Neural Information Processing Systems. – 2017. – Vol. 30. – Р. 5998-6008. https://doi.org/10.48550/arXiv.1706.03762.
6. Tan H. LXMERT: Learning cross-modality encoder representations from transformers / H. Tan, M. Bansal // EMNLP. – 2019. – Р. 5103-5114. https://doi.org/10.18653/v1/D19-1514.
7. Understanding back-translation at scale / S. Edunov et al // EMNLP. – 2018. – Р. 489-500. https://doi.org/10.18653/v1/D18-1150.
8. Hossain M.Z. Multimodal machine learning for emotion recognition: A review / M.Z. Hossain, G. Muhammad, M. Alsulaiman // Information Fusion. – 2021. – vol. 68. – Р. 21-39. https://doi.org/10.1016/j.inffus.2020.10.008.
9. Dabre R. Enabling multilingual neural machine translation with knowledge distillation / R. Dabre, C. Chu, S. Kurohashi // ACL. – 2017. – Р. 2662-2673. https://doi.org/10.18653/v1/P17-1243.
10. Wang R. Sentiment-aware image captioning with context disentangling / R. Wang, X. Wan, W. Li // CVPR. – 2021. – Р. 17586-17595. https://doi.org/10.1109/CVPR46437.2021.01732.
11. UNITER: Learning universal image-text representations / Y.-C. Chen et al // ECCV. – 2020. – Р. 104-120. https://doi.org/10.1007/978-3-030-58523-5_43.
12. BERT: Pre-training of deep bidirectional transformers for language understanding / J. Devlin et al // NAACL-HLT. – 2019. – Р. 4171-4186. https://doi.org/10.18653/v1/N19-1423.
Review
For citations:
Akhmediarova A.T., Ayapbergenova A.T., Alibiyeva Zh.M., Мukazhanov N.K., Issabekov Zh.N. UNIVERSAL TRANSFORMER FRAMEWORK FOR TEXT ANALYSIS AND GENERATION. Bulletin of Shakarim University. Technical Sciences. 2026;1(1(21)):36-46. (In Russ.) https://doi.org/10.53360/2788-7995-2026-1(21)-4
JATS XML















