Egor Bondarev

dblp:52/2450 · also Egor R. V. Bondarev · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
14since 2021 · last 2025
0009-0005-2452-7389ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unmasking Performance Gaps: A Comparative Study of Human Anonymization and Its Effects on Video Anomaly Detection
Sara Abdulaziz, Egor Bondarev
ACIVS2
2025 Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
abstract
Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such as shoplifting from visually similar events. Therefore, towards robust complex real-world VAD, it is essential to augment RGB spatio-temporal features by additional modalities. Motivated by this, we introduce the Poly-modal Induced framework for VAD: "PI-VAD" (or π-VAD), a novel approach that augments RGB representations by five additional modalities. Specifically, the modalities include sensitivity to fine-grained motion (Pose), three dimensional scene and entity representation (Depth), surrounding objects (Panoptic masks), global motion (optical flow), as well as language cues (VLM). Each modality represents an axis of a polygon, streamlined to add salient cues to RGB. π-VAD includes two plug-in modules, namely Pseudo-modality Generation module and Cross Modal Induction module, which generate modality-specific prototypical representation and, thereby, induce multi-modal information into RGB cues. These modules operate by performing anomaly-aware auxiliary tasks and necessitate five modality backbones – only during training. Notably, π-VAD achieves state-of-the-art accuracy on three prominent VAD datasets encompassing real-world scenarios, without requiring the computational overhead of five modality backbones at inference.
Snehashis Majhi, Giacomo D'Amicantonio, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, Egor Bondarev, François Brémond
CVPR7
2025 MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks
abstract
Deep Neural Networks (DNNs) are accurate but compute-intensive, leading to substantial energy consumption during inference. Exploiting temporal redundancy through ∆-Σ convolution [26] in video processing has proven to greatly enhance computation efficiency. However, temporal ∆-Σ DNNs typically require substantial memory for storing neuron states to compute inter-frame differences, hindering their on-chip deployment. To mitigate this memory cost, directly compressing the states can disrupt the linearity of temporal ∆-Σ convolution, causing accumulated errors in long-term ∆-Σ processing. Thus, we propose MEET, an optimization framework for MEmory-Efficient Temporal ∆-Σ DNNs. MEET transfers the state compression challenge to a well-established weight compression problem by trading fewer activations for more weights and introduces a co-design of network architecture and suppression method to optimize for mixed spatial-temporal execution. Evaluations on three vision applications demonstrate a reduction of 5.1∼13.3 × in total memory compared to the most computation-efficient temporal DNNs, while preserving the computation efficiency and model accuracy in long-term ∆-Σ processing. MEET facilitates the deployment of temporal ∆-Σ DNNs within on-chip memory of embedded event-driven platforms, empowering low-power edge processing.
Zeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
CVPR4
2025 Mixture of Experts Guided by Gaussian Splatters Matters: A New Approach to Weakly-Supervised Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are provided during training, while predictions are made at the frame level. Although state-of-the-art models perform well on simple anomalies (e.g., explosions), they struggle with complex real-world events (e.g., shoplifting). This difficulty stems from two key issues: (1) the inability of current models to address the diversity of anomaly types, as they process all categories with a shared model, overlooking category-specific features; and (2) the weak supervision signal, which lacks precise temporal information, limiting the ability to capture nuanced anomalous patterns blended with normal events. To address these challenges, we propose Gaussian Splatting-guided Mixture of Experts (GS-MoE), a novel framework that employs a set of expert models, each specialized in capturing specific anomaly types. These experts are guided by a temporal Gaussian splatting loss, enabling the model to leverage temporal consistency and enhance weak supervision. The Gaussian splatting approach encourages a more precise and comprehensive representation of anomalies by focusing on temporal segments most likely to contain abnormal events. The predictions from these specialized experts are integrated through a mixture-of-experts mechanism to model complex relationships across diverse anomaly patterns. Our approach achieves state-of-the-art performance, with a 91.58% AUC on the UCF-Crime dataset, and demonstrates superior results on XD-Violence and MSAD datasets. By leveraging category-specific expertise and temporal guidance, GS-MoE sets a new benchmark for VAD under weak supervision.
Giacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond, Egor Bondarev
ICCV7
2024 ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration
Zeqi Zhu, Alberto García Ortiz, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
ECCV (11)4
2024 TeG: Temporal-Granularity Method for Anomaly Detection with Attention in Smart City Surveillance
abstract
Anomaly detection in video surveillance has recently gained interest from the research community. Temporal duration of anomalies vary within video streams, leading to complications in learning the temporal dynamics of specific events. This paper presents a temporal-granularity method for an anomaly detection model (TeG) in real-world surveillance, combining spatio-temporal features at different time-scales. The TeG model employs multi-head cross-attention (MCA) blocks and multi-head self-attention (MSA) blocks for this purpose. Additionally, we extend the UCF-Crime dataset with new anomaly types relevant to Smart City research project. The TeG model is deployed and validated in a city surveillance system, achieving successful real-time results in industrial settings.
Erkut Akdag, Egor Bondarev, Peter H. N. de With
VCIP2
2024 Automated Camera Calibration via Homography Estimation with GNNs
abstract
Over the past few decades, a significant rise of camera-based applications for traffic monitoring has occurred. Governments and local administrations are increasingly relying on the data collected from these cameras to enhance road safety and optimize traffic conditions. However, for effective data utilization, it is imperative to ensure accurate and automated calibration of the involved cameras. This paper proposes a novel approach to address this challenge by leveraging the topological structure of intersections.We propose a framework involving the generation of a set of synthetic intersection viewpoint images from a bird’seye-view image, framed as a graph of virtual cameras to model these images. Using the capabilities of Graph Neural Networks, we effectively learn the relationships within this graph, thereby facilitating the estimation of a homography matrix. This estimation leverages the neighbourhood representation for any real-world camera and is enhanced by exploiting multiple images instead of a single match. In turn, the homography matrix allows the retrieval of extrinsic calibration parameters. As a result, the proposed framework demonstrates superior performance on both synthetic datasets and real-world cameras, setting a new state-of-the-art benchmark.
Giacomo D'Amicantonio, Egor Bondarev, Peter H. N. de With
WACV2
2024 CATS: Combined Activation and Temporal Suppression for Efficient Network Inference
abstract
Brain-inspired event-driven processors execute deep neural networks (DNNs) in a sparsity-aware manner, leading to superior performance compared to conventional platforms. In the pursuit of higher event sparsity, prior studies suppress non-zero events by either eliminating the intra-frame activations (spatially) or leveraging the redundancy in the inter-frame differences for a video (temporally). However, we have empirically observed that simultaneously enhancing activation and temporal sparsity can lead to a synergistic suppression outcome. To this end, we propose an end-to-end event suppression training approach CATS −− Combined Activation and Temporal Suppression for efficient network inference. It utilizes a gradient-based method to search for the optimal temporal thresholds per layer while penalizing the presence of events in both spatial and temporal domains. Our experimental results show that CATS achieves 2 ∼ 6× higher event suppression compared to the inherent ReLU suppression across a wide range of vision applications, consistently outperforming the state-of-the-art (SOTA) methods by a significant margin at all accuracy levels. Furthermore, a case study on the commercial event-driven processor GrAI-VIP highlights that the induced event sparsity in SSD on the EgoHands dataset can be efficiently translated into a performance enhancement of 2.5× in FPS, 2.1× in latency, and 3.8× in energy consumption, while maintaining the model accuracy.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Ibrahim Batuhan Akkaya, Egor Bondarev, Orlando Moreira
WACV5
2022 ARTS: An adaptive regularization training schedule for activation sparsity exploration
abstract
Brain-inspired event-based processors have attracted considerable attention for edge deployment because of their ability to efficiently process Convolutional Neural Networks (CNNs) by exploiting sparsity. On such processors, one critical feature is that the speed and energy consumption of CNN inference are approximately proportional to the number of non-zero values in the activation maps. Thus, to achieve top performance, an efficient training algorithm is required to largely suppress the activations in CNNs. We propose a novel training method, called Adaptive-Regularization Training Schedule (ARTS), which dramatically decreases the non-zero activations in a model by adaptively altering the regularization coefficient through training. We evaluate our method across an extensive range of computer vision applications, including image classification, object recognition, depth estimation, and semantic segmentation. The results show that our technique can achieve 1.41 × to 6.00 × more activation suppression on top of ReLU activation across various networks and applications, and outperforms the state-of-the-art methods in terms of training time, activation suppression gains, and accuracy. A case study for a commercially-available event-based processor, Neuronflow, shows that the activation suppression achieved by ARTS effectively reduces CNN inference latency by up to 8.4 × and energy consumption by up to 14.1 ×.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Lennart Bamberg, Egor Bondarev, Orlando Moreira
DSD5
2022 Detection and correction of image distortions induced by parasitic light sensitivity using deep learning methods
abstract
Images captured by mobile camera systems are subject to distortions that can be irreversible. Sources of these distortions vary and can be attributed to sensor imperfections, lens defects, or shutter inefficiency. One form of image distortion is associated with high Parasitic-Light-Sensitivity (PLS) in CMOS Image Sensors when combined with Global Shutters (GS-CIS) in a moving camera system. The resulting distortion appears as widespread semi-transparent purple artifacts, or a complex purple fringe, covering a large area in the scene around high-intensity regions. Most of the earlier approaches addressing the purple fringing problems have been directed towards the simplest forms of this distortion and rely on heuristic image processing algorithms. Recently, machine learning methods have shown remarkable success in many image restoration and object detection problems. Nevertheless, they have not been applied for the complex purple fringing detection or correction. In this paper, we present our exploration and deployment of deep learning algorithms in a pipeline for the detection and correction of the purple fringing induced by high-PLS GS-CIS sensors. Experiments show that the proposed methods outperform state-of-the-art approaches for both problems of detection and color restoration. We achieve a final MS-SSIM of 0.966 on synthetic data, and a distortion classification accuracy of 96.97%. We further discuss the limitations and possible improvements over the proposed methods.
Sara Esam Azmi, Chris Bartels, Zill-e Hussnain, J. A. P. Guelen, Egor Bondarev
ICMV5
2022 Floor-plan generation from noisy point clouds
abstract
This paper proposes a growing-based floor-plan generation method that creates the global layout of buildings from noisy point clouds obtained by a stereo camera. We introduce a PCA-based line-growing concept with a subsequent filtering step, which is able to robustly handle the high noise levels in input point clouds. Experimental results show that this method outperforms the state-of-the-art techniques in floor-plan generation. The average F1 score for building layouts has increased from 0.38 to 0.66 on our test dataset, compared to the previous best floor-plan generation method. Furthermore, the resulting floor plans are multiple thousands of times smaller in memory size than the input point clouds, while still preserving the main building structures.
Egor Bondarev, Peter H. N. de With
ICMV2
2022 Critical Vehicle Detection for Intelligent Transportation Systems
abstract
An intelligent transportation system (ITS) is one of the core elements of smart cities, enhancing public safety and relieving traffic congestion. Detection and classification of critical vehicles, such as police cars and ambulances, passing through roadways form crucial use cases for ITS. This paper proposes a solution for detecting and classifying safety-critical vehicles on urban roadways using deep learning models. At present, a large-scale dataset for critical vehicles is not publicly available. The appearance scarcity of emergency vehicles and different coloring standards in various countries are significant challenges. To cope with the mentioned drawbacks and to address the unique requirements of our smart city project, we first generate a large-scale critical vehicle dataset, combining images retrieved from various sources with the support of the YOLO vehicle detection model. The classes of the generated dataset are: fire truck, police car, ambulance, military police car, dang erous truck, and standard vehicle. Second, we compare the performance of the Vision in Transformer (ViT) network against the traditional convolutional neural networks (CNNs) for the task of critical vehicle classification. Experimental results on our dataset reveal that the ViT-based solution reaches an average accuracy and recall of 99.39% and 99.34%, respectively.
Erkut Akdag, Egor Bondarev, Peter H. N. de With
VEHITS2
2022 Toward Multilabel Image Retrieval for Remote Sensing
abstract
The availability of large-scale remote sensing (RS) data facilitates a wide range of applications, such as disaster management and urban planning. An approach for such problems is image retrieval, where, given a query image, the goal is to find the most relevant match from a database. Most RS literature has been focused on single-label retrieval, where we assume an image has a single label. The primary challenge in single-label RS retrieval is that performance in most datasets is saturated, and it has become difficult to compare the performance of different methods. In this work, we extend the major multilabel classification datasets to the multilabel retrieval problem. We also define protocols, provide evaluation metrics, and study the impact of commonly used loss functions and reranking methods for multilabel retrieval. To this end, a novel multilabel loss function and a reranking technique are proposed, which circumvent the challenges present in conventional single-label image retrieval. The developed loss function considers both class and feature similarity. The proposed reranking technique achieves high performance with computation cost that is well-suited for fast online retrieval.
Raffaele Imbriaco, Clint Sebastian, Egor Bondarev, Peter H. N. de With
IEEE Trans. Geosci. Remote. Sens.3
2021 Maritime vessel re-identification: novel VR-VCA dataset and a multi-branch architecture MVR-net
abstract
Abstract Maritime vessel re-identification (re-ID) is a computer vision task of vessel identity matching across disjoint camera views. Prominent applications of vessel re-ID exist in the fields of surveillance and maritime traffic flow analysis. However, the field suffers from the absence of a large-scale dataset that enables training of deep learning models. In this study, we present a new dataset that includes 4614 images of 729 vessels along with 5-bin orientation and 8-class vessel-type annotations to promote further research. A second contribution of this study is the baseline re-ID analysis of our new dataset. Performances of 10 recent deep learning architectures are quantitatively compared to reveal the best practices. Lastly, we propose a novel multi-branch deep learning architecture, Maritime Vessel Re-ID network (MVR-net), to address the challenging problem of vessel re-ID. Evaluation of our approach on the new dataset yields 74.5% mAP and 77.9% Rank-1 score, providing a performance increase of 5.7% mAP and 5.0% Rank-1 over the best-performing baseline. MVR-net also outperforms the PRN (a pioneering vehicle re-ID network), by 2.9% and 4.3% higher mAP and Rank-1, respectively.
Amir Ghahremani, Tunç Alkanat, Egor Bondarev, Peter H. N. de With
Mach. Vis. Appl.3
2019 Improving open-set person re-identification by statistics-driven gallery refinement
abstract
Person re-identification (re-ID) is a valuable tool for multi-camera tracking of persons. Up till now, research on person re-ID has mainly focused on the closed-set case, where a given query is assumed to always have a correct match in the gallery set, which does not hold for practical scenarios. In this study, we explore the open-set person re-ID problem with queries not always included in the gallery set. First, we convert the popular closed-set person re-ID datasets into the open-set scenario. Second, we compare the performances of six state-of-the-art closed-set person re-ID methods under open-set conditions. Third, we investigate the impact of a simple and fast statistics-driven gallery refinement approach on the open-set person re-ID performance. Extensive experimental evaluations show that, gallery refinement increases the performance of existing methods in the low false-accept rate (FAR) region, while simultaneously reducing the computational demands of retrieval. Results show an average detection and identification rate (DIR) increase of 7.91% and 3.31% on the DukeMTMC-reID and Market1501 datasets, respectively, for an FAR of 1%.
Tunç Alkanat, Egor Bondarev, Peter H. N. de With
ICMV2
2018 Towards multi-class detection: a self-learning approach to reduce inter-class noise from training dataset
abstract
This paper proposes a novel self-learning framework, which converts a noisy, pre-labeled multi-class object dataset into a purified multi-class object dataset with object bounding-box annotations, by iteratively removing noise samples from the low-quality dataset, which may contain a high level of inter-class noise samples. The framework iteratively purifies the noisy training datasets for each class and updates the classification model for multiple classes. The procedure starts with a generic single-class object model which changes to a multi-class model in an iterative procedure of which the F-1 score is evaluated to reach a sufficiently high score. The proposed framework is based on learning the used models with CNNs. As a result, we obtain a purified multi-class dataset and as a spin-off, the updated multi-class object model. The proposed framework is evaluated on maritime surveillance, where vessels need to be classified into eight different types. The experimental results on the evaluation dataset show that the proposed framework improves the F-1 score approximately by 5% and 25% at the end of the third iteration, while the initial training datasets contain 40% and 60% inter-class noise samples (erroneously classified labels of vessels and without annotations), respectively. Additionally, the recall rate increases nearly by 38% (for the more challenging 60% inter-class noise case), while the mean Average Precision (mAP) rate remains stable.
Amir Ghahremani, Egor Bondarev, Peter H. N. de With
ICMV2
2018 Towards accurate camera geopositioning by image matching
abstract
In this work, we present a camera geopositioning system based on matching a query image against a database with panoramic images. For matching, our system uses memory vectors aggregated from global image descriptors based on convolutional features to facilitate fast searching in the database. To speed up searching, a clustering algorithm is used to balance geographical positioning and computation time. We refine the obtained position from the query image using a new outlier removal algorithm. The matching of the query image is obtained with a recall@5 larger than 90% for panorama-to-panorama matching. We cluster available panoramas from geographically adjacent locations into a single compact representation and observe computational gains of approximately 50% at the cost of only a small (approximately 3%) recall loss. Finally, we present a coordinate estimation algorithm that reduces the median geopositioning error by up to 20%.
Raffaele Imbriaco, Clint Sebastian, Egor Bondarev, Peter H. N. de With
ICMV3
2016 R3P: Real-time RGB-D Registration Pipeline
Hani Javan Hemmat, Egor Bondarev, Peter H. N. de With
ACIVS2
2016 ProMARTES: Accurate network and computation delay prediction for component-based distributed systems
Konstantinos Triantafyllidis, Waqar Aslam, Egor Bondarev, Johan J. Lukkien, Peter H. N. de With
J. Syst. Softw.3
2015 Solidarity Filter for Noise Reduction of 3D Edges in Depth Images
Hani Javan Hemmat, Egor Bondarev, Peter H. N. de With
ACIVS2
2014 Exploring Distance-Aware Weighting Strategies for Accurate Reconstruction of Voxel-Based 3D Synthetic Models
Hani Javan Hemmat, Egor Bondarev, Peter H. N. de With
MMM (1)2
2013 On photo-realistic 3D reconstruction of large-scale and arbitrary-shaped environments
abstract
This paper presents a system architecture for reconstructing photorealistic and accurate 3D models of indoor environments. The system specifically targets large-scale and arbitrary-shaped environments and enables processing of data obtained with an arbitrary-chosen capturing path. The system extends the baseline Kinect Fusion algorithm with a buffering algorithm to remove scene-size capturing limitations. Beside this, the paper presents the complete chain of advanced algorithms for point cloud segmentation/decimation, camera pose correction and texture mapping with post-processing filters. The presented architecture features memory- and processor-efficient processing, such that it can be executed on a conventional PC with a mainstream GPU card at the consumer premises.
Egor Bondarev, Francisco Heredia, Rafael Favier, Lingni Ma, Peter H. N. de With
CCNC1
2013 Plane segmentation and decimation of point clouds for 3D environment reconstruction
abstract
Three-dimensional (3D) models of environments are a promising technique for serious gaming and professional engineering applications. In this paper, we introduce a fast and memory-efficient system for the reconstruction of large-scale environments based on point clouds. Our main contribution is the emphasis on the data processing of large planes, for which two algorithms have been designed to improve the overall performance of the 3D reconstruction. First, a flatness-based segmentation algorithm is presented for plane detection in point clouds. Second, a quadtree-based algorithm is proposed for decimating the point cloud involved with the segmented plane and consequently improving the efficiency of triangulation. Our experimental results have shown that the proposed system and algorithms have a high efficiency in speed and memory for environment reconstruction. Depending on the amount of planes in the scene, the obtained efficiency gain varies between 20% and 50%.
Lingni Ma, Raphael Favier, Luat Do, Egor Bondarev, Peter H. N. de With
CCNC4
2013 Performance analysis method for RT systems: Promartes for autonomous robot
Konstantinos Triantafyllidis, Egor Bondarev, Peter H. N. de With
FDL2
2007 CARAT: a toolkit for design and performance analysis of component-based embedded systems
abstract
Solid frameworks and toolkits for design and analysis of embedded systems are of high importance, since they enable early reasoning about critical properties of a system. This paper presents a software toolkit that supports the design and performance analysis of real-time component-based software architectures deployed on heterogeneous multiprocessor platforms. The tooling environment contains a set of integrated tools for (a) component storage and retrieval, (b) graphics-based design of software and hardware architectures, (c) performance analysis of the designed architectures and, (d) automated code generation. The cornerstone of the toolkit is a performance analysis framework that automates composition of the individual component models into a system executable model, allows simulation of the system model and gives design-time predictions of key performance properties like response time, data throughput, and usage of hardware resources. The efficiency of this toolkit was illustrated on a car radio navigation benchmark system
Egor Bondarev, Michel R. V. Chaudron, Peter H. N. de With
DATE1
2006 A Toolkit for Design and Performance Analysis of Real-Time Component-Based Software Systems
abstract
Software tools supporting the design and analysis of complex software-intensive systems are highly desirable, since they enable earlier decision making about system realization. This paper presents a tooling environment that supports the design and performance analysis of time-critical component-based software architectures deployed on complex multiprocessor platforms. The tooling environment contains a set of integrated tools for (a) component storage and retrieval, (b) graphics-based design of software and hardware architectures, (c) performance analysis of the defined architectures and, (d) automated code generation. The cornerstone of the toolkit is a performance analyzer that provides efficient simulation of the designed architectures and enables design-time prediction of key performance properties like response time, data throughout, and usage of hardware resources (processor, memory and bus). For every architecture alternative, the performance predictions can be quickly obtained, thereby enabling a fast and yet broad design space exploration. We demonstrate the efficiency and robustness of this toolkit on a Car Radio Navigation benchmark case.
Egor Bondarev, Michel R. V. Chaudron, Heorhiy Byelas, Peter H. N. de With
ICSEA1