<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">kaz44</journal-id><journal-title-group><journal-title xml:lang="ru">Вестник Университета Шакарима. Серия технические науки</journal-title><trans-title-group xml:lang="en"><trans-title>Bulletin of Shakarim University. Technical Sciences</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2788-7995</issn><issn pub-type="epub">3006-0524</issn><publisher><publisher-name>«Шәкәрім университеті» КеАҚ</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.53360/2788-7995-2026-1(21)-9</article-id><article-id custom-type="elpub" pub-id-type="custom">kaz44-2438</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>АВТОМАТИЗАЦИЯ И ИНФОРМАЦИОННЫЕ ТЕХНОЛОГИИ (ОРИГИНАЛЬНАЯ СТАТЬЯ)</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>AUTOMATION AND INFORMATION TECHNOLOGY (ORIGINAL ARTICLE)</subject></subj-group></article-categories><title-group><article-title>ОПТИМИЗАЦИЯ МОДЕЛЕЙ МАШИННОГО ОБУЧЕНИЯ ДЛЯ ДЕТЕКЦИИ ФЕЙКОВЫХ НОВОСТЕЙ: ПОДБОР ГИПЕРПАРАМЕТРОВ И АНАЛИЗ ROC-КРИВЫХ</article-title><trans-title-group xml:lang="en"><trans-title>OPTIMIZING MACHINE LEARNING MODELS FOR FAKE NEWS DETECTION: HYPERPARAMETER SELECTION AND ROC CURVES ANALYSIS</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0006-5319-7742</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Тюлемисова</surname><given-names>Д. Т.</given-names></name><name name-style="western" xml:lang="en"><surname>Tyulemissova</surname><given-names>D.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Дана Болатовна Тюлемисова – магистр технических наук, докторант специальности 8D06306 – «Системы информационной безопасности» </p><p>10000, г.Астана, ул. Сатпаева, 2</p></bio><bio xml:lang="en"><p>Dana Tyulemissova – Master of Science (Engineering), PhD candidate in the specialty 8D06306 – Information Security Systems </p><p>10000, Astana, Satpayev Street, 2</p></bio><email xlink:type="simple">tyulemissova_db_3@enu.kz</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-6006-4813</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Шайханова</surname><given-names>А. К.</given-names></name><name name-style="western" xml:lang="en"><surname>Shaikhanova</surname><given-names>A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Айгуль Кайрулаевна Шайханова – доктор философии, профессор кафедры информационной безопасности </p><p>10000, г.Астана, ул. Сатпаева, 2</p></bio><bio xml:lang="en"><p>Aigul Shaikhanova – PhD, Professor, Department of Information Security </p><p>10000, Astana, Satpayev Street, 2</p></bio><email xlink:type="simple">shaikhanova_ak@enu.kz</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-5622-1038</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Марценюк</surname><given-names>В. П.</given-names></name><name name-style="western" xml:lang="en"><surname>Martsenyuk</surname><given-names>V.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Васыл Марценюк – доктор технических наук, профессор кафедры информатики и автоматизации </p><p>ul.Willowa 2, 43-300, Бельско-Бяла </p></bio><bio xml:lang="en"><p>Vasyl Martsenyuk – Doctor of Engineering Sciences, Professor, Department of Computer Science and Automation </p><p>ul.Willowa 2, 43-300, Bielsko-Biała</p></bio><email xlink:type="simple">vmartsenyuk@ath.bielsko.pl</email><xref ref-type="aff" rid="aff-2"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-1635-4693</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Бекешова</surname><given-names>Г. Б.</given-names></name><name name-style="western" xml:lang="en"><surname>Bekeshova</surname><given-names>G.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Гульвира Бауржановна Бекешова – магистр технических наук, ст. преподаватель кафедры информационной безопасности </p><p>10000, г.Астана, ул. Сатпаева, 2</p></bio><bio xml:lang="en"><p>Gulvira Baurzhanovna Bekeshova – Master of Technical Sciences, Senior Lecturer at the Department of Information Security </p><p>10000, Astana, Satpayev Street, 2</p></bio><email xlink:type="simple">bekeshova_gb@enu.kz</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0005-4932-1774</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Смаилова</surname><given-names>Б. Т.</given-names></name><name name-style="western" xml:lang="en"><surname>Smailova</surname><given-names>B.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Балжан Темірболатқызы Смаилова – магистр, заведущая кафедрой математики </p><p>071412, г. Семей, ул. Глинки 20 А</p></bio><bio xml:lang="en"><p>Balzhan Smailova – Master of Science, Head of Mathematics Department </p><p>071412, Semey, Glinka Street 20 A</p></bio><email xlink:type="simple">st.balzhan@gmail.com</email><xref ref-type="aff" rid="aff-3"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Евразийский национальный университет имени Л.Н. Гумилева</institution><country>Казахстан</country></aff><aff xml:lang="en"><institution>L.N. Gumilyov Eurasian National University</institution><country>Kazakhstan</country></aff></aff-alternatives><aff-alternatives id="aff-2"><aff xml:lang="ru"><institution>Университет Бельско-Бяла</institution><country>Польша</country></aff><aff xml:lang="en"><institution>University of Bielsko-Biała</institution><country>Poland</country></aff></aff-alternatives><aff-alternatives id="aff-3"><aff xml:lang="ru"><institution>Шәкәрім университет</institution><country>Казахстан</country></aff><aff xml:lang="en"><institution>Shakarym University</institution><country>Kazakhstan</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>25</day><month>05</month><year>2026</year></pub-date><volume>1</volume><issue>1(21)</issue><fpage>83</fpage><lpage>92</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Тюлемисова Д.Т., Шайханова А.К., Марценюк В.П., Бекешова Г.Б., Смаилова Б.Т., 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Тюлемисова Д.Т., Шайханова А.К., Марценюк В.П., Бекешова Г.Б., Смаилова Б.Т.</copyright-holder><copyright-holder xml:lang="en">Tyulemissova D., Shaikhanova A., Martsenyuk V., Bekeshova G., Smailova B.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://tech.vestnik.shakarim.kz/jour/article/view/2438">https://tech.vestnik.shakarim.kz/jour/article/view/2438</self-uri><abstract><p>В условиях активного и неконтролируемого распространения информации в цифровой среде, особенно в социальных медиа и новостных агрегаторах, особую актуальность приобретает задача автоматического выявления фейковых новостей. Рост объёмов пользовательского контента и высокая скорость его распространения существенно усложняют ручную верификацию информации, что обуславливает необходимость применения методов машинного обучения. Целью настоящего исследования является анализ, сравнительная оценка и оптимизация базовых моделей машинного обучения для детекции фейковых новостей с акцентом на подбор гиперпараметров, вычислительную эффективность и интерпретацию результатов классификации. В работе использован набор текстовых данных ISOT Fake News Dataset, содержащий новостные сообщения на английском языке с бинарными метками «правда» и «ложь», прошедшие этапы очистки и векторизации с применением метода TF-IDF. В рамках исследования реализованы и проанализированы модели логистической регрессии, дерева решений, случайного леса и градиентного бустинга. Оценка качества классификации проводилась на основе метрик Accuracy, Precision, Recall, F1-score, а также ROC-анализа и значения площади под кривой (AUC). Показано, что модель градиентного бустинга обеспечивает наивысшую точность классификации, тогда как логистическая регрессия демонстрирует сопоставимое качество при значительно меньших вычислительных затратах и времени обучения. Дополнительно в работе выполнен контроль утечки данных и анализ причин завышенных метрик качества, что позволило обеспечить корректную интерпретацию результатов. Полученные выводы подтверждают, что оптимизированные классические модели машинного обучения способны обеспечивать высокое качество детекции фейковых новостей и могут рассматриваться как ресурсоэффективная альтернатива более сложным нейросетевым подходам в прикладных задачах кибербезопасности и мониторинга информационных потоков.</p></abstract><trans-abstract xml:lang="en"><p>In the context of the active and uncontrolled dissemination of information in the digital environment, especially on social media and news aggregators, the task of automatically identifying fake news has become especially relevant. The growing volume of user-generated content and the high speed of its distribution significantly complicate manual information verification, necessitating the use of machine learning methods. The aim of this study is to analyze, comparatively evaluate, and optimize baseline machine learning models for fake news detection, focusing on hyperparameter selection, computational efficiency, and interpretation of classification results. This study utilizes the ISOT Fake News Dataset, a text dataset containing Englishlanguage news items with binary labels labeled «true» and «false», cleaned and vectorized using the TF-IDF method. The study implemented and analyzed logistic regression, decision tree, random forest, and gradient boosting models. Classification quality was assessed using the following metrics: Accuracy, Precision, Recall, F1-score, as well as ROC analysis and the area under the curve (AUC). It was shown that the gradient boosting model provides the highest classification accuracy, while logistic regression demonstrates comparable quality with significantly lower computational costs and training time. Additionally, the study included data leakage monitoring and an analysis of the causes of inflated quality metrics, ensuring the correct interpretation of the results. The findings confirm that optimized classical machine learning models are capable of providing highquality fake news detection and can be considered a resource-efficient alternative to more complex neural network approaches in applied cybersecurity and information flow monitoring.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>фейковые новости</kwd><kwd>машинное обучение</kwd><kwd>оптимизация гиперпараметров</kwd><kwd>классификация текстов</kwd><kwd>ROC-анализ</kwd><kwd>AUC</kwd><kwd>логистическая регрессия</kwd></kwd-group><kwd-group xml:lang="en"><kwd>fake news</kwd><kwd>machine learning</kwd><kwd>hyperparameter optimization</kwd><kwd>text classification</kwd><kwd>ROC analysis</kwd><kwd>AUC</kwd><kwd>logistic regression</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Fake news detection on social media: a data mining perspective / K. Shu et al // ACM SIGKDD Explorations Newsletter. – 2017. – Vol. 19, № 1. – P. 22-36.</mixed-citation><mixed-citation xml:lang="en">Fake news detection on social media: a data mining perspective / K. Shu et al // ACM SIGKDD Explorations Newsletter. – 2017. – Vol. 19, № 1. – P. 22-36.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Zhou X. Fake news: a survey of research, detection methods, and opportunities / X. Zhou, R. Zafarani // ACM Computing Surveys. – 2020. – Vol. 53, № 5. – Article 109.</mixed-citation><mixed-citation xml:lang="en">Zhou X. Fake news: a survey of research, detection methods, and opportunities / X. Zhou, R. Zafarani // ACM Computing Surveys. – 2020. – Vol. 53, № 5. – Article 109.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Ahmed H. Detection of online fake news using n-gram analysis and machine learning techniques / H. Ahmed, I. Traore, S. Saad // Proceedings of the International Conference on Intelligent, Secure, and Dependable Systems. – 2018. – P. 127-138.</mixed-citation><mixed-citation xml:lang="en">Ahmed H. Detection of online fake news using n-gram analysis and machine learning techniques / H. Ahmed, I. Traore, S. Saad // Proceedings of the International Conference on Intelligent, Secure, and Dependable Systems. – 2018. – P. 127-138.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Ruchansky N. CSI: a hybrid deep model for fake news detection / N. Ruchansky, S. Seo, Y. Liu // Proceedings of the ACM International Conference on Information and Knowledge Management (CIKM). – 2017. – P. 797-806.</mixed-citation><mixed-citation xml:lang="en">Ruchansky N. CSI: a hybrid deep model for fake news detection / N. Ruchansky, S. Seo, Y. Liu // Proceedings of the ACM International Conference on Information and Knowledge Management (CIKM). – 2017. – P. 797-806.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">BERT: pre-training of deep bidirectional transformers for language understanding / J. Devlin et al // Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT). – 2019. – P. 4171-4186.</mixed-citation><mixed-citation xml:lang="en">BERT: pre-training of deep bidirectional transformers for language understanding / J. Devlin et al // Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT). – 2019. – P. 4171-4186.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">exBAKE: automatic fake news detection model based on bidirectional encoder representations from transformers / H. Jwa et al // Applied Sciences. – 2019. – Vol. 9, № 19. – P. 1-14.</mixed-citation><mixed-citation xml:lang="en">exBAKE: automatic fake news detection model based on bidirectional encoder representations from transformers / H. Jwa et al // Applied Sciences. – 2019. – Vol. 9, № 19. – P. 1-14.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Kaliyar R.K. FakeBERT: fake news detection in social media with a BERT-based deep learning approach / R.K. Kaliyar, A. Goswami, P. Narang // Multimedia Tools and Applications. – 2021. – Vol. 80, № 8. – P. 11765-11788.</mixed-citation><mixed-citation xml:lang="en">Kaliyar R.K. FakeBERT: fake news detection in social media with a BERT-based deep learning approach / R.K. Kaliyar, A. Goswami, P. Narang // Multimedia Tools and Applications. – 2021. – Vol. 80, № 8. – P. 11765-11788.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Strubell E., Ganesh A., McCallum A. Energy and policy considerations for deep learning in NLP / Strubell E., Ganesh A., McCallum A. // Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). – 2019. – P. 3645-3650.</mixed-citation><mixed-citation xml:lang="en">Strubell E., Ganesh A., McCallum A. Energy and policy considerations for deep learning in NLP / Strubell E., Ganesh A., McCallum A. // Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). – 2019. – P. 3645-3650.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Horne B.D. This just in: fake news packs a lot in title, uses simpler, repetitive content / B.D. Horne, S. Adalı // Proceedings of the International AAAI Conference on Web and Social Media (ICWSM). – 2017. – P. 759-766.</mixed-citation><mixed-citation xml:lang="en">Horne B.D. This just in: fake news packs a lot in title, uses simpler, repetitive content / B.D. Horne, S. Adalı // Proceedings of the International AAAI Conference on Web and Social Media (ICWSM). – 2017. – P. 759-766.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">FakeNewsNet: a data repository with news content, social context and dynamic information / K. Shu et al // Big Data. – 2020. – Vol. 8, № 3. – P. 171-188.</mixed-citation><mixed-citation xml:lang="en">FakeNewsNet: a data repository with news content, social context and dynamic information / K. Shu et al // Big Data. – 2020. – Vol. 8, № 3. – P. 171-188.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
