Ilya Makarov

dblp:161/1790 · DBLP profile ↗
← Back
41ranked-venue papers
3as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 20 · 20 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Framework GNN-AID: Graph Neural Network Analysis, Interpretation and Defense
abstract
The rising demand for Trusted AI (TAI) underscores the need for interpretable and robust models, yet existing tools rarely support graph-structured data or integrate interpretability with security. At the same time, Graph Neural Networks (GNNs) deliver state-of-the-art performance on numerous graph tasks. We present GNN-AID (Graph Neural Network Analysis, Interpretation, and Defense), an open-source Python framework for analyzing, interpreting, and defending GNNs, addressing this critical gap. Built on PyTorch-Geometric, GNN-AID offers preloaded datasets, model libraries, flexible APIs, and a web interface for visualization and no-code model design. MLOps features further support reproducibility and experiment tracking.
Kirill Lukianov, Mikhail Drobyshevskiy, Georgii V. Sazonov, Mikhail Soloviov, Ilya Makarov
AAAI5
2026 RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents
abstract
Climate change is driving more frequent and severe disasters, putting people and infrastructure at risk. Protecting communities requires models that capture both natural disasters dynamics and how people behave under extreme conditions. This demo presents RESPOND, a multi-agent LLM-enhanced platform that jointly simulates natural hazards and human response. RESPOND couples high-fidelity flood AI forecasting an agent-based model of human behavior. LLM modules improve each agent decision-making, enabling context-aware reasoning over alerts, road closures, social signals, and changing water levels. The system simulates evacuation flows, resource seeking, and communication patterns producing actionable outputs for emergency management, urban planning, and policy. In the live demo one can run what-if or predicted scenarios, adjust assumptions, and observe emergent population behavior and risk hot spots in real time. By tightly coupling dynamic hazards with LLM-driven multi-agent behavior, RESPOND moves beyond fragmented tools and offers a practical, integrated platform for disaster preparedness and response.
Roman Sultimov, Mikhail Mozikov, Dmitrii Abramov, Mariia Kovalchuk, Maksim Malykh, Ilya Makarov, Andrei Osiptsov, Aleksandr Volkov, Yury Maximov
AAAI6
2026 RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract)
abstract
Climate change is driving more frequent and severe disasters, putting people and infrastructure at risk. Protecting communities requires models that capture both natural disasters dynamics and how people behave under extreme conditions. This demo presents RESPOND, a multi-agent LLM-enhanced platform that jointly simulates natural hazards and human response. RESPOND couples high-fidelity flood AI forecasting an agent-based model of human behavior. LLM modules improve each agent decision-making, enabling context-aware reasoning over alerts, road closures, social signals, and changing water levels. The system simulates evacuation flows, resource seeking, and communication patterns producing actionable outputs for emergency management, urban planning, and policy. In the live demo one can run what-if or predicted scenarios, adjust assumptions, and observe emergent population behavior and risk hot spots in real time. By tightly coupling dynamic hazards with LLM-driven multi-agent behavior, RESPOND moves beyond fragmented tools and offers a practical, integrated platform for disaster preparedness and response.
Roman Sultimov, Mikhail Mozikov, Dmitrii Abramov, Mariia Kovalchuk, Maksim Malykh, Aleksandr Volkov, Ilya Makarov, Andrei Osiptsov, Yury Maximov
AAAI7
2026 Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
Nikita Severin, Danil Kartushov, Vladislav Urzhumov, Vladislav Kulikov, Oksana Konovalova, Alexey Grishanov, Anton Klenitskiy, Artem Fatkulin, Alexey Vasilev, Andrey V. Savchenko, Ilya Makarov
ECIR (2)11
2026 SOMIN: Agentic AI for Automating Professional Visibility and Countering Content Homogenization
Aleksandr Farseev, Kirill Lepikhin, Kevin Manuel, Maksim Gorodilov, Zhao Kui, Ilya Makarov, Jaime Francisco Maldonado, Ian Cassidy, Ronan Byrne
SIGIR6
2026 Beyond Isolated Clients: Integrating Graph-Based Embeddings into Event Sequence Models
Harry Proshian, Nikita Severin, Sergey I. Nikolenko, Ivan Kireev, Andrey V. Savchenko, Ivan Sergeev, Maria Postnova, Ilya Makarov
WWW8
2025 MEH: A Multi-Style Dataset and Toolkit for Advancing Egyptian Hieroglyph Recognition
Maksim Golyadkin, Valeria Rubanova, Aleksandr Utkov, Dmitry Nikolotov, Ilya Makarov
ICCV5
2025 Search Swarm: Multiagent Large Language Models Framework for E-commerce Product Search
abstract
Search engines are vital for online e-commerce but often struggle with long, detailed queries. We introduce Search Swarm, a novel multi-agent system designed to improve search engine navigation on platforms like Amazon by accurately locating relevant products based on user instructions. Search Swarm employs multiple large language model (LLM) agents, each with a specific role: query planner, searcher, critic, and attribute selector. These agents collaborate to generate search queries, evaluate results, and identify the best product options tailored to users' needs. Our framework outperforms existing methods like ReAct and Reflexion in the WebShop environment, achieving a reward score of 62.64, compared to scores of 54.1, 59.8, 61.5, and 58.2 for other approaches. Furthermore, in a comparison with a basic rule-based method on Amazon, Search Swarm achieved a score 38.71 points higher and a 41\% greater success rate, demonstrating its superior ability to provide relevant product matches over traditional search engines.
Nagim Isyanbaev, Ilya Makarov
IJCAI2
2025 Machine Learning Driven Optimization of Fe-Based TMCs for Photodynamic Therapy
abstract
Noble metal-based photoactive complexes have applications in photodynamic therapy (PDT), but their toxicity and high cost drive interest in sustainable and cheaper alternatives like iron-based compounds. In this paper, quantum chemistry and classical molecular dynamics were employed to characterize the photophysical properties and non-covalent interactions with DNA of two Fe(III) complexes. We explained the absorption of IR wavelength by bright ligand-to-metal transitions and showed that the complexes exhibit persistent, albeit modest, interaction with DNA. Building on these traditional simulation methods, we propose a conceptual ML-driven optimization module designed to refine the structure of iron complexes and enhance their photophysical features. While the framework is not yet implemented, we demonstrate that key properties relevant for PDT can be computationally evaluated, providing a foundation for future iterative optimization. The ML module integrates 3D molecular structures, simulation results, and quantum chemical insights to suggest modifications aimed at shifting the absorption spectrum more favorably into the visible range, improving their suitability for phototherapies.
Vladimir Manuilov, Antonio Francés-Monerris, Abdelazim M. A. Abdelgawwad, Daniel Escudero 0004, Ilya Makarov
IJCAI5
2025 FaceCluster: Interactive Photo Organization with Enhanced Face Recognition
Alexander Filonenko, Ilya Makarov, Andrey V. Savchenko
ACM Multimedia2
2025 MuMMy: Multimodal Dataset supporting VLM-based Egyptology Research Assistant
abstract
We present the first multimodal dataset MuMMy, for developing research assistants that can interpret Egyptian hieroglyphic texts. It pairs images with Gardiner codes, transliteration, and English translation at two levels of granularity. We also evaluate several deep learning pipelines across OCR, transliteration, and translation tasks, revealing the complexity of the domain and the challenges posed by error accumulation.
Maksim Golyadkin, Innokentiy Humonen, Valeria Rubanova, Danil Kalin, Yanis Plevokas, Dmitry Nikolotov, Aleksandr Utkov, Nikita Sidelnikov, Petr Ivanov, Ekaterina Bureeva, Ekaterina Alexandrova, Ilya Makarov
ACM Multimedia12
2025 Evaluation of Egyptian Hieroglyph Classification Across Diverse Writing Styles
abstract
The classification of Egyptian hieroglyphs remains a challenging problem due to the vast variability in writing styles across time periods, regions, and individual scribes. In this work, we present a comprehensive evaluation of hieroglyph classification performance across diverse stylistic domains, highlighting the limitations of current models in generalizing beyond a single style. We introduce a dataset that spans multiple writing styles, ranging from monumental inscriptions to handwritten manuscripts, and assess several near state-of-the-art recognition models. Our analysis reveals significant discrepancies in model performance when exposed to unseen styles, underscoring the need for style-aware learning strategies. This study provides a framework for future research on hieroglyph recognition with a focus on stylistic diversity and serves as a first step toward building vision-language systems capable of analyzing Egyptian hieroglyphic writings.
Maksim Golyadkin, Valeria Rubanova, Aleksandr Utkov, Dmitry Nikolotov, Ilya Makarov
ACM Multimedia5
2025 Real-Time SSL Sperm Whale Click Detector: Interactive Web Demo
abstract
Reliable, real-time detection of sperm-whale clicks is essential yet difficult in noisy ocean audio streams. We present the first browser-based pipeline that combines self-supervised embeddings with a BiLSTM to label clicks at millisecond resolution. The system attains 99% F1 on the Watkins benchmark, reduces false alarms by 60% against the best published baseline, and analyses one-second segments in ~40 ms on a consumer GPU. An interactive UI overlays multi-algorithm detections on waveform and spectrogram views with drag-zoom and live streaming. Code, pretrained weights and the public demo are released to advance bioacoustic event detection.
Anvar Iskhakov, Viktor Kovalev, Vladislav Naumov, Ilya Makarov
ACM Multimedia4
2025 HL-EAI: A Multimodal Framework Enabling Emotional Reciprocity in Human-AI Strategic Decision-Making
Mikhail Mozikov, Daniil Orekhov, Ivan Nasonov, Konstantin Baltsat, Vladislav Pedashenko, Dmitrii Abramov, Nikita Severin, Yury Maximov, Andrey V. Savchenko, Ilya Makarov
ACM Multimedia10
2025 Poster: Car2Vec: Task-Agnostic Latent Embeddings from CAN Bus for Efficient V2X Communication
abstract
The high bandwidth demand of raw sensor data transmission remains a critical bottleneck for scalable Vehicle-to-Everything (V2X) networks. We propose Car2Vec - a self-supervised contrastive learning framework that generates compact latent embeddings from vehicle CAN bus data. Our initial experiments demonstrate the framework's ability to create visually distinguishable embeddings in latent space, successfully separating clusters for fuel from different suppliers and identifying varying levels of motor oil viscosity. These early results suggest the framework's potential to encode semantic state information including vehicle dynamics and driver context, and serve as efficient communication primitives for V2V/V2I applications, potentially reducing wireless overhead by orders of magnitude compared to raw telemetry. By moving computation to vehicle edge devices, Car2Vec aligns with the AI-RAN co-design vision, offering a promising direction for semantic-aware resource optimization in next-gen transportation networks.
Aleksandr Kovalenko, Ahmad Ahmad, Alexander Karandeev, Alexey Maslov, Dmitry Zhevnenko, Ilya Makarov
MobiCom6
2025 Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
abstract
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In this paper, we introduce a method for evaluating the ability of LLM-based agents to cooperate in zero-shot, mixed-motive environments using Concordia, a natural language multi-agent simulation environment. Our method measures general cooperative intelligence by testing an agent's ability to identify and exploit opportunities for mutual gain across diverse partners and contexts. We present empirical results from the NeurIPS 2024 Concordia Contest, where agents were evaluated on their ability to achieve mutual gains across a suite of diverse scenarios ranging from negotiation to collective action problems. Our findings reveal significant gaps between current agent capabilities and the robust generalization required for reliable cooperation, particularly in scenarios demanding persuasion and norm enforcement.
Chandler Smith, Marwa Abdulhai, Manfred Diaz, Marko Tesic, Rakshit S. Trivedi, Alexander Vezhnevets, Lewis Hammond, Jesse Clifton, Minsuk Chang, Edgar A. Duéñez-Guzmán, John P. Agapiou, Jayd Matyas, Danny Karmon, Beining Zhang, Jim Dilkes, Akash Kundu, Emanuel Tewolde, Jebish Purbey, Ram Mohan Rao Kadiyala, Siddhant Gupta, Aliaksei Korshuk, Buyantuev Alexander, Ilya Makarov, Rolando Fernandez, Zhihan Wang, Caroline Wang, Jiaxun Cui, Lingyun Xiao, Yoonchang Sung, Muhammad Arrasy Rahman, Peter Stone 0001, Yipeng Kang, Hyeonggeun Yun, Ananya, Taehun Cha, Elizaveta Tennant, Olivia Macmillan-Scott, Marta Segura, Diana Riazi, Fuyang Cui, Sriram Ganapathi, Toryn Q. Klassen, Nico Schiavone, Mogtaba Alim, Sheila A. McIlraith, Manuel Ríos, Oswaldo Peña, Manuela Chacon-Chamorro, Rubén Manrique, Luis Felipe Giraldo, Nicanor Quijano, Fangwei Zhong, Wenming Tu, Zhaowei Zhang 0001, Zixia Jia, Zilong Zheng, Chichen Lin, Weijian Fan, Chenao Liu, Sneheel Sarangi, Shuqing Shi, Yali Du 0001, Avinaash Anand Kulandaivel, Yang Liu 0266, Ruiyang Wu 0007, Chetan Talele, Sunjia Lu, Gema Parreno, Shamika Dhuri, Bain McHale, Tim Baarslag, Dylan Hadfield-Menell, Natasha Jaques, José Hernández-Orallo, Joel Z. Leibo
NeurIPS24
2025 Poster Abstract: Exploring the Autoencoder Sequence Pooling
abstract
Sequence embeddings are essential for tasks like time series analysis and natural language processing, yet pooling techniques to create compact sequence representations remain underexplored. Poor pooling methods can lead to significant information loss, diminishing the effectiveness of strong feature extractors. In this early results paper, we propose an autoencoder-based sequence pooling approach that leverages autoencoders' ability to compress information and supports pretraining during self-supervised learning. Evaluated on time series data with a transformer-based encoder, our method outperforms traditional pooling techniques, such as mathematical functions and learnable weighted sums.
Petr Ivanov, Alexander Kozhevnikov, Maria Shtark, Ilya Makarov
SenSys4
2025 Poster Abstract: Autonomous AI-Driven Grid Protection: Sub-Cycle Fault Response via NPU-Optimized Neural Networks
abstract
Relay Protection and Automation (RPA) devices maintain power system stability by isolating faults using predefined or adaptive thresholds. However, even flexible configurations require manual pre-tuning, limiting their adaptability in dynamic grid environments, such as those with renewable energy integration. Artificial Intelligence (AI) augments RPA functionality by generalizing to new, unseen data, improving fault detection accuracy in transient or noisy scenarios. Our early studies demonstrate that embedded AI systems outperform static thresholds while operating in real time. These systems, implemented via neural networks on System-on-Chip (SoC) platforms with dedicated Neural Processing Units (NPUs), avoid cloud-dependent latency.
Aleksandr Kovalenko, Aleksey Evdakov, Galina Filatova, Andrey Yablokov, Aleksandr Bulashov, Ilya Makarov
SenSys6
2025 Poster Abstract: Minimizing Labeling Efforts for Fault Detection and Diagnosis
abstract
We present a semi-supervised fault detection method that combines contrastive and active learning techniques to efficiently analyze sensor data in industrial systems. It uses a transformer-based encoder with rotary positional embeddings for self-supervised pre-training, which allows it to build structured representations that can be used for density-based clustering using DBSCAN. This method significantly reduces the need for manual labeling, as it only requires 20% of the data to achieve high accuracy in clustering. It outperforms unsupervised alternatives and is scalable and adaptable to evolving fault patterns. Its reduced manual intervention makes it a valuable tool for real-world industrial health monitoring.
Maria Shtark, Alexander Kozhevnikov, Petr Ivanov, Ilya Makarov
SenSys4
2025 Closing the Domain Gap in Manga Colorization via Aligned Paired Dataset
abstract
This paper addresses the challenge of artwork colorization by proposing a benchmark for manga colorization using real black-and-white and colorized image pairs. Color images are widely recognized for their ability to capture attention and improve memory retention, yet the manual process of colorization is labor-intensive. Deep learning methods for supervised image-to-image translation offer a promising solution, relying on aligned pairs of black-and-white and color images for training. However, these pairs are often generated synthetically, introducing a domain gap that limits model performance. To address this, we explore the use of real data, proposing a method for creating such datasets. Our benchmarks reveal that models trained on real data significantly outperform those trained on synthetic pairs. Furthermore, we present a pipeline for text removal and panel segmentation, streamlining the comic colorization process. These contributions aim to enhance the generalization and applicability of deep learning models for artwork colorization.
Maksim Golyadkin, Yanis Plevokas, Ilya Makarov
WACV3
2024 Time Series Generation with GANs for Momentum Effect Simulation on Moscow Stock Exchange
abstract
The ability to accurately simulate financial markets is crucial, as it allows researchers and practitioners to rigorously test and refine trading strategies without the high risks associ-ated with real-world experimentation. By leveraging Generative Adversarial Networks (GANs), this research aims to enhance the robustness and effectiveness of trading strategies by providing a controlled environment to assess potential outcomes and strategy resilience under varied market conditions. In this work, we propose the application of GAN s for simu-1ating multidimensional time series in the context of developing and testing trading strategies. We conduct an experimental study to jointly simulate the log-returns of several stocks on Moscow Exchange for Momentum effect evaluation. Compared to traditional methods such as bootstrapping, GANs can better model and interpolate the non-parametric complex nature of the data, providing an increased diverse sample size. This methodology could be beneficial for investors seeking new opportunities to test and tune hyperparameters of other trading strategies.
Maksim Kazadaev, Vitaliy Pozdnyakov, Ilya Makarov
CIFEr3
2024 Weak-to-Strong 3D Object Detection with X-Ray Distillation
abstract
This paper addresses the critical challenges of sparsity and occlusion in LiDAR-based 3D object detection. Current methods often rely on supplementary modules or specific architectural designs, potentially limiting their applicability to new and evolving architectures. To our knowledge, we are the first to propose a versatile technique that seamlessly integrates into any existing framework for 3D Object Detection, marking the first instance of Weak-to-Strong generalization in 3D computer vision. We introduce a novel framework, X-Ray Distillation with Object-Complete Frames, suitable for both supervised and semi-supervised settings, that leverages the temporal aspect of point cloud sequences. This method extracts crucial information from both previous and subsequent LiDAR frames, creating Object-Complete frames that represent objects from multiple viewpoints, thus addressing occlusion and sparsity. Given the limitation of not being able to generate Object-Complete frames during online inference, we utilize Knowledge Distillation within a Teacher-Student framework. This technique encourages the strong Student model to emulate the behavior of the weaker Teacher, which processes simple and informative Object-Complete frames, effectively offering a comprehensive view of objects as if seen through X-ray vision. Our proposed methods surpass state-of-the-art in semi-supervised learning by 1-1.5 mAP and enhance the performance of five established supervised models by 1–2 mAP on standard autonomous driving datasets, even with default hyperparameters. Code for Object-Complete frames is available here: https://github.com/sakharok13/X-Ray-Teacher-Patching-Tools.
Alexander Gambashidze, Aleksandr Dadukin, Maksim Golyadkin, Maria Razzhivina, Ilya Makarov
CVPR5
2024 CA-SER: Cross-Attention Feature Fusion for Speech Emotion Recognition
abstract
In this paper, we introduce a novel tool for speech emotion recognition, CA-SER, that borrows self-supervised learning to extract semantic speech representations from a pre-trained wav2vec 2.0 model and combine them with spectral audio features to improve speech emotion recognition. Our approach involves a self-attention encoder on MFCC features to capture meaningful patterns in audio sequences. These MFCC features are combined with high-level representations using a multi-head cross-attention mechanism. Evaluation of speech emotion recognition on the IEMOCAP dataset shows that our system achieves a weighted accuracy of 74.6%, outperforming most existing techniques.
Bashar M. Deeb, Andrey V. Savchenko, Ilya Makarov
ECAI3
2024 InsideOut: Unifying Emotional LLMs to Foster Empathy
abstract
This paper introduces InsideOut, an original innovative framework that augments the emotional intelligence of Large Language Models (LLMs). Motivated by the cartoon, InsideOut is designed around a net of specialized agents, each dedicated to one of Ekman’s fundamental emotions. These agents collaboratively refine responses sensitive to the emotional context of interactions. Our assessments, conducted using EmpatheticDialogues and involving models like GPT-4 and GigaChat, indicate substantial improvements in identifying human emotions and generating empathetic responses. These improvements are most evident in situations with apparent valence-arousal differences. InsideOut offers a promising avenue for evolving AI into more perceptive and human-centric communicators.
Mikhail Mozikov, Nikita Severin, Maria Glushanina, Mikhail Baklashkin, Andrey V. Savchenko, Ilya Makarov
ECAI6
2024 Device-Specific Facial Descriptors: Winning a Lottery with a SuperNet
abstract
We address the challenge of devising neural network architectures to extract facial descriptors across diverse mobile and edge devices. Employing neural architecture search, we introduce a novel framework that selects optimal subnetworks from a SuperNet using an evolutionary search. Using a surrogate gradient boosting classifier to avoid direct accuracy estimation of subnetworks on validation sets, our approach swiftly delivers the most efficient and accurate models tailored to specific devices within minutes. Demonstrating versatility through an Android demo app, our framework excels in tasks like face recognition and emotion understanding across various devices, achieving real-time processing and superior accuracy compared to existing mobile models.
Andrey V. Savchenko, Dmitry Maslov, Ilya Makarov
ECAI3
2024 From Data to Decisions: Streamlining Geospatial Operations with Multimodal GlobeFlowGPT
abstract
As machine learning increasingly becomes a crucial tool for geospatial data analysis, finding and deploying a suitable model presents significant challenges, including the need for expertise in both programming and geospatial analysis, organizing data flow, and accurately assessing the results. To address these challenges, this paper introduces GlobeFlowGPT, a multimodal, chat-based framework designed to meet these demands by integrating domain-specific tools, machine learning models, Multimodal Large Language Models, and essential operational data. It leverages a Large Language Model orchestrator, facilitating complex geospatial tasks through a conversational interface. GlobeFlowGPT's flexible, containerized architecture allows for the rapid integration of cutting-edge models tailored for geospatial data, ensuring that the framework remains scalable and relevant amid ongoing technological advancements. We demonstrate the ability of our framework to streamline the analysis of geospatial data and expand the capabilities of modern MLLMs with complex geospatial machine learning models.
Danil Kononykhin, Mikhail Mozikov, Kirill Mishtal, Pavel Kuznetsov, Dmitrii Abramov, Nazar Sotiriadi, Yury Maximov, Andrey V. Savchenko, Ilya Makarov
SIGSPATIAL/GIS9
2024 Do You Remember the Future? Weak-to-Strong Generalization in 3D Object Detection
Alexander Gambashidze, Aleksandr Dadukin, Maksim Golyadkin, Maria Razzhivina, Ilya Makarov
IJCAI5
2024 SensorSCAN: Self-supervised learning and deep clustering for fault diagnosis in chemical processes (Abstract Reprint)
Maksim Golyadkin, Vitaliy Pozdnyakov, Leonid Zhukov, Ilya Makarov
IJCAI4
2024 Plug-and-Play Unsupervised Fault Detection and Diagnosis for Complex Industrial Monitoring
Maksim Golyadkin, Maria Shtark, Petr Ivanov, Alexander Kozhevnikov, Leonid Zhukov, Ilya Makarov
IJCAI6
2024 AADMIP: Adversarial Attacks and Defenses Modeling in Industrial Processes
Vitaliy Pozdnyakov, Aleksandr Kovalenko, Ilya Makarov, Mikhail Drobyshevskiy, Kirill S. Lukyanov
IJCAI3
2024 EAI: Emotional Decision-Making of LLMs in Strategic Games and Ethical Dilemmas
abstract
One of the urgent tasks of artificial intelligence is to assess the safety and alignment of large language models (LLMs) with human behavior. Conventional verification only in pure natural language processing benchmarks can be insufficient. Since emotions often influence human decisions, this paper examines LLM alignment in complex strategic and ethical environments, providing an in-depth analysis of the drawbacks of our psychology and the emotional impact on decision-making in humans and LLMs. We introduce the novel EAI framework for integrating emotion modeling into LLMs to examine the emotional impact on ethics and LLM-based decision-making in various strategic games, including bargaining and repeated games. Our experimental study with various LLMs demonstrated that emotions can significantly alter the ethical decision-making landscape of LLMs, highlighting the need for robust mechanisms to ensure consistent ethical standards. Our game-theoretic analysis revealed that LLMs are susceptible to emotional biases influenced by model size, alignment strategies, and primary pretraining language. Notably, these biases often diverge from typical human emotional responses, occasionally leading to unexpected drops in cooperation rates, even under positive emotional influence. Such behavior complicates the alignment of multiagent systems, emphasizing the need for benchmarks that can rigorously evaluate the degree of emotional alignment. Our framework provides a foundational basis for developing such benchmarks.
Mikhail Mozikov, Nikita Severin, Valeria Bodishtianu, Maria Glushanina, Ivan Nasonov, Daniil Orekhov, Pekhotin Vladislav, Ivan Makovetskiy, Mikhail Baklashkin, Vasily Lavrentyev, Akim Tsvigun, Denis Turdakov, Tatiana Shavrina, Andrey V. Savchenko, Ilya Makarov
NeurIPS15
2024 GEEF: A neural network model for automatic essay feedback generation by integrating writing skills assessment
Yuanchao Liu, Jiawei Han 0011, Alexander G. Sboev, Ilya Makarov
Expert Syst. Appl.4
2023 MonoVAN: Visual Attention for Self-Supervised Monocular Depth Estimation
abstract
Depth estimation is crucial in various computer vision applications, including autonomous driving, robotics, and virtual and augmented reality. An accurate scene depth map is beneficial for localization, spatial registration, and tracking. It converts 2D images into precise 3D coordinates for accurate positioning, seamlessly aligns virtual and real objects in applications like AR, and enhances object tracking by distinguishing distances. The self-supervised monocular approach is particularly promising as it eliminates the need for complex and expensive data acquisition setups relying solely on a standard RGB camera. Recently, transformer-based architectures have become popular to solve this problem, but at high quality, they suffer from high computational cost and poor perception of small details as they focus more on global information. In this paper, we propose a novel fully convolutional network for monocular depth estimation, called MonoVAN, which incorporates the visual attention mechanism and applies super-resolution techniques in decoder to better capture fine-grained details in depth maps. To the best of our knowledge, this work pioneers the use of a convolutional visual attention in the context of depth estimation. Our experiments on outdoor KITTI benchmark and the indoor NYUv2 dataset show that our approach outperforms the most advanced self-supervised methods, including such state-of-the-art models as transformer-based VTDepth from ISMAR’22 and hybrid convolutional-transformer MonoFormer from AAAI’23, while having a comparable or even fewer number of parameters in our model than competitors. We also validate the impact of each proposed improvement in isolation, providing evidence of its significant contribution. Code and weights are available at https://github.com/IlyaInd/MonoVAN.
Ilia Indyk, Ilya Makarov
ISMAR2
2023 Ti-DC-GNN: Incorporating Time-Interval Dual Graphs for Recommender Systems
abstract
Recommender systems are essential for personalized content delivery and have become increasingly popular recently. However, traditional recommender systems are limited in their ability to capture complex relationships between users and items. Dynamic graph neural networks (DGNNs) have recently emerged as a promising solution for improving recommender systems by incorporating temporal and sequential information in dynamic graphs. In this paper, we propose a novel method, "Ti-DC-GNN" (Time-Interval Dual Causal Graph Neural Networks), based on an intermediate representation of graph evolution as a sequence of time-interval graphs. The main parts of the method are the novel forms of interval graphs: graph of causality and graph of consequence that explicitly preserve inter-relationships between edges (user-items interactions). The local and global message passing are developed based on edge memory to identify short-term and long-term dependencies. Experiments on several well-known datasets show that our method consistently outperforms modern temporal GNNs with node memory alone in dynamic edge prediction tasks.
Nikita Severin, Andrey V. Savchenko, Dmitrii Kiselev, Maria Ivanova, Ivan Kireev, Ilya Makarov
RecSys6
2023 SensorSCAN: Self-supervised learning and deep clustering for fault diagnosis in chemical processes
Maksim Golyadkin, Vitaliy Pozdnyakov, Leonid Zhukov, Ilya Makarov
Artif. Intell.4
2022 Exploring Efficiency of Vision Transformers for Self-Supervised Monocular Depth Estimation
abstract
Depth estimation is a crucial task for the creation of depth maps, one of the most important components for augmented reality (AR) and other applications. However, the most widely used hardware for AR and smartphones has only sparse depth sensors with different ground truth depth acquisition methods. Thus, depth estimation models that are robust for downstream AR tasks performance can only be trained reliably using self-supervised learning based on camera information. Previous works in the field mostly focus on self-supervised models with pure convolutional architectures, without taking global spatial context into account.In this paper, we utilize vision transformer architectures for self-supervised monocular depth estimation and propose VTDepth, a vision transformer-based model, which provides a solution to the problem of the global spatial context. We compare various combinations of convolutional and transformer architectures for self-supervised depth estimation and show that the best combination of models is an encoder with a transformer basis and convolutional decoder. Our experiments demonstrate the efficiency of VTDepth for self-supervised depth estimation. Our set of models achieves state-of-the-art performance for self-supervised learning on NYUv2 and KITTI datasets. Our code is available at https://github.com/ahbpp/VTDepth.
Aleksei Karpov, Ilya Makarov
ISMAR2
2022 Classifying Emotions and Engagement in Online Learning Based on a Single Facial Expression Recognition Neural Network
abstract
In this article, behaviour of students in the e-learning environment is analyzed. The novel pipeline is proposed based on video facial processing. At first, face detection, tracking and clustering techniques are applied to extract the sequences of faces of each student. Next, a single efficient neural network is used to extract emotional features in each frame. This network is pre-trained on face identification and fine-tuned for facial expression recognition on static images from AffectNet using a specially developed robust optimization technique. It is shown that the resulting facial features can be used for fast simultaneous prediction of students’ engagement levels (from disengaged to highly engaged), individual emotions (happy, sad, etc.,) and group-level affect (positive, neutral or negative). This model can be used for real-time video processing even on a mobile device of each student without the need for sending their facial video to the remote server or teacher's PC. In addition, the possibility to prepare a summary of a lesson is demonstrated by saving short clips of different emotions and engagement of all students. The experimental study on the datasets from EmotiW (Emotion Recognition in the Wild) challenges showed that the proposed network significantly outperforms existing single models.
Andrey V. Savchenko, Lyudmila V. Savchenko, Ilya Makarov
IEEE Trans. Affect. Comput.3
2019 Deep Reinforcement Learning in Match-3 Game
abstract
An increasing number of algorithms in deep reinforcement learning area creates new challenges for environments, particularly, for their comprehensive analysis and searching application areas. The key purpose of this article is to provide an extensible environment for researches. We consider a Match-3 game, which has simple gameplay, but challenging game design for engaging players. The article provides metrics for evaluation of agents and corresponding baselines in different scenarios.
Ildar Kamaldinov, Ilya Makarov
CoG2
2019 On Reproducing Semi-dense Depth Map Reconstruction using Deep Convolutional Neural Networks with Perceptual Loss
abstract
In our recent papers, we proposed a new family of residual convolutional neural networks trained for semi-dense and sparse depth reconstruction without use of RGB channel. The proposed models can be used in low-resolution depth sensors or SLAM methods estimating partial depth with certain distributions. We proposed using perceptual loss for training depth reconstruction in order to better preserve edge structure and reduce over-smoothness of models trained on MSE loss alone. This paper contains reproducibility companion guide on training, running and evaluating suggested methods, while also presenting links on further studies in view of reviewers comments and related problems of depth reconstruction.
Ilya Makarov, Dmitrii Maslov, Olga Gerasimova, Vladimir Aliev, Alisa Korinevskaya, Ujjwal Sharma 0001
ACM Multimedia1
2017 Semi-Dense Depth Interpolation using Deep Convolutional Neural Networks
abstract
With advances of recent technologies, augmented reality systems and autonomous vehicles gained a lot of interest from academics and industry. Both these areas rely on scene geometry understanding, which usually requires depth map estimation. However, in case of systems with limited computational resources, such as smartphones or autonomous robots, high resolution dense depth map estimation may be challenging. In this paper, we study the problem of semi-dense depth map interpolation along with low resolution depth map upsampling. We present an end-to-end learnable residual convolutional neural network architecture that achieves fast interpolation of semi-dense depth maps with different sparse depth distributions: uniform, sparse grid and along intensity image gradient. We also propose a loss function combining classical mean squared error with perceptual loss widely used in intensity image super-resolution and style transfer tasks. We show that with some modifications, this architecture can be used for depth map super-resolution. Finally, we evaluate our results on both synthetic and real data, and consider applications for autonomous vehicles and creating AR/MR video games.
Ilya Makarov, Vladimir Aliev, Olga Gerasimova
ACM Multimedia1
2016 First-Person Shooter Game for Virtual Reality Headset with Advanced Multi-Agent Intelligent System
abstract
We present a multiplayer first-person shooter (FPS) game with advanced intelligent non-playable characters (NPC) under computer control. The game is specially adapted for playing in VR headset so the simulator sickness symptoms are significantly reduced. The demo allows users to play with the other human and NPC players in a shooter game made in Unreal Engine 4. User can verify his/her game skills versus evolving human-like NPCs with a level adjusting model. The humanness of NPC was verified with Alan Turing game test beating 52% record from BotPrize'12 competition.
Ilya Makarov, Mikhail Tokmakov, Pavel Polyakov, Peter Zyuzin, Maxim Martynov, Oleg Konoplya, George Kuznetsov, Ivan Guschenko-Cheverda, Maxim Uriev, Ivan Mokeev, Olga Gerasimova, Lada Tokmakova, Alexey Kosmachev
ACM Multimedia1