Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Stanislav Panev

dblp:15/10477 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-2193-5847ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Transfer learning and domain adaptation · 58% Image recognition and object detection · 29% 3D vision · 13%
Computer networks
1 paper
Wireless sensing and localization · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.912025
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision · ICCV 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.912025
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision · ICCV 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation
weakly-supervised domain adaptation
0.912025
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision · ICCV 2025
Computer vision › 3D vision
pose estimation
0.412019
Person-in-WiFi: Fine-Grained Person Perception Using WiFi · ICCV 2019
Wireless sensing and localization
wifi sensing
0.412019
Person-in-WiFi: Fine-Grained Person Perception Using WiFi · ICCV 2019

Methods — techniques the papers use, named apart from their topics

multi-modal knowledge transfer · 0.9latent diffusion model · 0.9generative data augmentation · 0.9wifi signal processing · 0.8deep learning · 0.8
YearPublicationVenuePosition
2025 Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
abstract
Detecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/AGenDA
Minhyek Jeon, Shuowen Hu, Zheyang Qin, Shayok Chakraborty, Stanislav Panev, Celso de Melo, Fernando De la Torre
ICCV6
2025 Texture- and Shape-Based Adversarial Attacks for Overhead Image Vehicle Detection
abstract
Detecting vehicles in aerial images is difficult due to complex backgrounds, small object sizes, shadows, and occlusions. Although recent deep learning advancements have improved object detection, these models remain susceptible to adversarial attacks (AAs), challenging their reliability. Traditional AA strategies often ignore practical implementation constraints. Our work proposes realistic and practical constraints on texture (lowering resolution, limiting modified areas, and color ranges) and analyzes the impact of shape modifications on attack performance. We conducted extensive experiments with three object detector architectures, demonstrating the performance-practicality trade-off: more practical modifications tend to be less effective, and vice versa. We release both code and data to support reproducibility at https://github.com/humansensinglab/texture-shape-adversarial-attacks.
Mikael Yeghiazaryan, Sai Abhishek Si Namburu, Emily Kim, Stanislav Panev, Celso de Melo, Fernando De la Torre, Jessica K. Hodgins
ICIP4
2024 Exploring the Impact of Rendering Method and Motion Quality on Model Performance when Using Multi-view Synthetic Data for Action Recognition
abstract
This paper explores the use of synthetic data in a human action recognition (HAR) task to avoid the challenges of obtaining and labeling real-world datasets. We introduce a new dataset suite comprising five datasets, eleven common human activities, three synchronized camera views (aerial and ground) in three outdoor environments, and three visual domains (real and two synthetic). For the synthetic data, two rendering methods (standard computer graphics and neural rendering) and two sources of human motions (motion capture and video-based motion reconstruction) were employed. We evaluated each dataset type by training popular activity recognition models and comparing the performance on the real test data. Our results show that synthetic data achieve slightly lower accuracy (4–8 %) than real data. On the other hand, a model pre-trained on synthetic data and fine-tuned on limited real data surpasses the performance of either domain alone. Standard computer graphics (CG)-rendered data delivers better performance than the data generated from the neural-based rendering method. The results suggest that the quality of the human motions in the training data also affects the test results: motion capture delivers higher test accuracy. Additionally, a model trained on CG aerial view synthetic data exhibits greater robustness against camera viewpoint changes than one trained on real data. See the project page: http://humansensinglab.github.io/REMAG/
Stanislav Panev, Emily Kim, Sai Abhishek Si Namburu, Desislava Nikolova, Celso de Melo, Fernando De la Torre, Jessica K. Hodgins
WACV1
2019 Person-in-WiFi: Fine-Grained Person Perception Using WiFi
abstract
Fine-grained person perception such as body segmentation and pose estimation has been achieved with many 2D and 3D sensors such as RGB/depth cameras, radars (e.g. RF-Pose), and LiDARs. These solutions require 2D images, depth maps or 3D point clouds of person bodies as input. In this paper, we take one step forward to show that fine-grained person perception is possible even with 1D sensors: WiFi antennas. Specifically, we used two sets of WiFi antennas to acquire signals, i.e., one transmitter set and one receiver set. Each set contains three antennas horizontally lined-up as a regular household WiFi router. The WiFi signal generated by a transmitter antenna, penetrates through and reflects on human bodies, furniture, and walls, and then superposes at a receiver antenna as 1D signal samples. We developed a deep learning approach that uses annotations on 2D images, takes the received 1D WiFi signals as input, and performs body segmentation and pose estimation in an end-to-end manner. To our knowledge, our solution is the first work based on off-the-shelf WiFi antennas and standard IEEE 802.11n WiFi signals. Demonstrating comparable results to image-based solutions, our WiFi-based person perception solution is cheaper and more ubiquitous than radars and LiDARs, while invariant to illumination and has little privacy concern comparing to cameras.
Fei Wang 0037, Sanping Zhou, Stanislav Panev, Jinsong Han
ICCV3
2019 Road Curb Detection and Localization With Monocular Forward-View Vehicle Camera
abstract
We propose a robust method for estimating road curb 3-D parameters (size, location, and orientation) using a calibrated monocular camera equipped with a fisheye lens. Automatic curb detection and localization is particularly important in the context of an advanced driver assistance system, i.e., to prevent possible collision and damage to the vehicle's bumper during perpendicular and diagonal parking maneuvers. Combining 3-D geometric reasoning with advanced vision-based detection methods, our approach is able to estimate the vehicle to curb distance in real time with a mean accuracy of more than 90%, as well as its orientation, height, and depth. Our approach consists of two distinct components-curb detection in each individual video frame and temporal analysis. The first part is comprised of sophisticated curb edges extraction and parameterized 3-D curb template fitting. Using a few assumptions regarding the real-world geometry, we can thus retrieve the curb's height and its relative position with respect to the moving vehicle on which the camera is mounted. Support vector machine classifier fed with histograms of oriented gradients is used for appearance-based filtering out outliers. In the second part, the detected curb regions are tracked in the temporal domain, so as to perform a second pass of false positives rejection. We have validated our approach on a newly collected database of 11 videos under different conditions. We have used point-wise LIDAR measurements and manual exhaustive labels as a ground truth.
Stanislav Panev, Francisco Vicente 0001, Fernando De la Torre, Véronique Prinet
IEEE Trans. Intell. Transp. Syst.1