VLDB 2026 Research / reviewers in the wild / expert
Reza Bahmanyar
dblp:141/9864
· DBLP profile ↗
15ranked-venue papers
5as first author
1since 2021 · last 2024
0000-0002-6999-714XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Video understanding and tracking · 87% 3D vision · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
multi-object tracking |
0.8 | 1 | 2024 | VETRA: A Dataset for Vehicle Tracking in Aerial Imagery - New Challenges for Multi-Object Tracking · ECCV (85) 2024 |
Computer vision › Video understanding and tracking › object tracking
UAV tracking |
0.8 | 1 | 2024 | VETRA: A Dataset for Vehicle Tracking in Aerial Imagery - New Challenges for Multi-Object Tracking · ECCV (85) 2024 |
Computer vision › 3D vision › remote sensing
aerial imagery |
0.2 | 1 | 2024 | VETRA: A Dataset for Vehicle Tracking in Aerial Imagery - New Challenges for Multi-Object Tracking · ECCV (85) 2024 |
Methods — techniques the papers use, named apart from their topics
multi-object tracking · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | VETRA: A Dataset for Vehicle Tracking in Aerial Imagery - New Challenges for Multi-Object Tracking
Jens Hellekes, Manuel Muehlhaus, Reza Bahmanyar, Seyed Majid Azimi, Franz Kurz |
ECCV (85) | 3 |
| 2020 | EAGLE: Large-Scale Vehicle Detection Dataset in Real-World Scenarios using Aerial ImageryabstractMulti-class vehicle detection from airborne imagery with orientation estimation is an important task in the near and remote vision domains with applications in traffic monitoring and disaster management. In the last decade, we have witnessed significant progress in object detection in ground imagery, but it is still in its infancy in airborne imagery, mostly due to the scarcity of diverse and large-scale datasets. Despite being a useful tool for different applications, current airborne datasets only partially reflect the challenges of real-world scenarios. To address this issue, we introduce EAGLE (oriEnted vehicle detection using Aerial imaGery in real-worLd scEnarios), a large-scale dataset for multi-class vehicle detection with object orientation information in aerial imagery. It features high-resolution aerial images composed of different real-world situations with a wide variety of camera sensor, resolution, flight altitude, weather, illumination, haze, shadow, time, city, country, occlusion, and camera angle. The annotation was done by airborne imagery experts with small-and large-vehicle classes. EAGLE contains 215,986 instances annotated with oriented bounding boxes defined by four points and orientation, making it by far the largest dataset to date in this task. It also supports researches on the haze and shadow removal as well as super-resolution and in-painting applications. We define three tasks: detection by (1) horizontal bounding boxes, (2) rotated bounding boxes, and (3) oriented bounding boxes. We carried out several experiments to evaluate several state-of-the-art methods in object detection on our dataset to form a baseline. Experiments show that the EAGLE dataset accurately reflects real-world situations and correspondingly challenging applications. The dataset will be made publicly available. Seyed Majid Azimi, Reza Bahmanyar, Corentin Henry, Franz Kurz |
ICPR | 2 |
| 2020 | AerialMPTNet: Multi-Pedestrian Tracking in Aerial Imagery Using Temporal and Graphical FeaturesabstractMulti-pedestrian tracking in aerial imagery has several applications such as large-scale event monitoring, disaster management, search-and-rescue missions, and as input into predictive crowd dynamic models. Due to the challenges such as the large number and the tiny size of the pedestrians (e.g., 4 × 4 pixels) with their similar appearances as well as different scales and atmospheric conditions of the images with their extremely low frame rates (e.g., 2 fps), current state-of-the-art algorithms including the deep learning-based ones are unable to perform well. In this paper, we propose AerialMPTNet, a novel approach for multi-pedestrian tracking in geo-referenced aerial imagery by fusing appearance features from a Siamese Neural Network, movement predictions from a Long Short-Term Memory, and pedestrian interconnections from a GraphCNN. In addition, to address the lack of diverse aerial pedestrian tracking datasets, we introduce the Aerial Multi-Pedestrian Tracking (AerialMPT) dataset consisting of 307 frames and 44,740 pedestrians annotated. We believe that AerialMPT is the largest and most diverse dataset to this date and will be released publicly. We evaluate AerialMPTNet on AerialMPT and KIT AIS, and benchmark with several state-of-the-art tracking methods. Results indicate that AerialMPTNet significantly outperforms other methods on accuracy and time-efficiency. Maximilian Kraus, Seyed Majid Azimi, Emec Ercelik, Reza Bahmanyar, Peter Reinartz, Alois C. Knoll |
ICPR | 4 |
| 2020 | Stepwise Refinement Of Low Resolution Labels For Earth Observation Data: Part 1abstractThis paper describes the contribution of the DLR team ranking 3rdin Track 1 of the 2020 IEEE GRSS Data Fusion Contest, with results ranking 2ndin Track 2 of the same contest being reported in a companion paper. The classifications are based on refinements of low-resolution MODIS labeling using available higher resolution Sentinel-1 and Sentinel-2 data. Results are initialized with a handcrafted decision tree integrating output from a random forest classifier, and subsequently boosted by detectors for specific classes. Daniele Cerra, Nina Merkle, Corentin Henry, Kevin Alonso 0001, Pablo d'Angelo, Stefan Auer, Reza Bahmanyar, Xiangtian Yuan, Ksenia Bittner, Maximilian Langheinrich, Guichen Zhang, Miguel Pato, Jiaojiao Tian, Peter Reinartz |
IGARSS | 7 |
| 2020 | Stepwise Refinement Of Low Resolution Labels For Earth Observation Data: Part 2abstractThis paper describes the contribution of the DLR team ranking 2ndin Track 2 of the 2020 IEEE GRSS Data Fusion Contest. The semantic classification of multimodal earth observation data proposed is based on the refinement of low-resolution MODIS labels, using as auxiliary training data higher resolution labels available for a validation data set. The classification is initialized with a handcrafted decision tree integrating output from a random forest classifier, and subsequently boosted by detectors for specific classes. The results of the team ranking 3rdin Track 1 of the same contest are reported in a companion paper. Daniele Cerra, Nina Merkle, Corentin Henry, Kevin Alonso 0001, Pablo d'Angelo, Stefan Auer, Reza Bahmanyar, Xiangtian Yuan, Ksenia Bittner, Maximilian Langheinrich, Guichen Zhang, Miguel Pato, Jiaojiao Tian, Peter Reinartz |
IGARSS | 7 |
| 2018 | Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
Seyed Majid Azimi, Eleonora Vig, Reza Bahmanyar, Marco Körner 0001, Peter Reinartz |
ACCV (3) | 3 |
| 2018 | Combining Deep and Shallow Neural Networks with Ad Hoc Detectors for the Classification of Complex Multi-Modal Urban ScenesabstractThis article describes the workflow of the classification algorithm which ranked at 2ndplace in the 2018 GRSS Data Fusion Contest. The objective of the contest was to provide a classification map with 20 classes on a complex urban scenario. The available multi-modal data were acquired from hyperspectral, LiDAR and very high-resolution RGB sensors flown on the same platform over the city of Houston, TX, USA. The classification was obtained by merging deep convolutional and shallow fully-connected neural networks on a simplified set of classes, complemented by a series of specific detectors and ad hoc classifiers. Daniele Cerra, Miguel Pato, Emiliano Carmona, Seyed Majid Azimi, Jiaojiao Tian, Reza Bahmanyar, Franz Kurz, Eleonora Vig, Ksenia Bittner, Corentin Henry, Pablo d'Angelo, Rupert Müller, Kevin Alonso 0001, Peter Fischer 0002, Peter Reinartz |
IGARSS | 6 |
| 2018 | Multisensor Earth Observation Image Classification Based on a Multimodal Latent Dirichlet Allocation ModelabstractMany previous researches have already shown the advantages of multisensor land-cover classification. Here, we propose an innovative land-cover classification approach based on learning a joint latent model of synthetic aperture radar (SAR) and multispectral satellite images using multimodal latent Dirichlet allocation (mmLDA), a probabilistic generative model. It has already been successfully applied to various other problems dealing with multimodal data. For our experiments, we chose overlapping SAR and multispectral images of two regions of interest. The images were tiled into patches and their local primitive features were extracted. Then each image patch is represented by SAR and multispectral bag-of-words (BoW) models. The BoW values are both fed to the mmLDA, resulting in a joint latent data model. A qualitative and quantitative validation of the topics based on ground-truth data demonstrate that the land-cover categories of the regions are correctly classified, outperforming the topics obtained using individual single modality data. Reza Bahmanyar, Daniela Espinoza-Molina, Mihai Datcu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Land-cover change detection using local feature descriptors extracted from spectral indicesabstractAn effective monitoring and analysis of ecosystems requires developing new tools and knowledge. In this paper, we propose an approach for detecting land-cover changes using satellite Image Time Series. This approach represents each image by spectral indices and then extracts local features of these representations. Next, a clustering technique (e.g., k-means) is applied to the extracted features, where the resulting clusters are assumed to refer to land-cover classes. The land-cover change is then obtained by counting the number of times an assigned class to each point changes along the time series. For our experiments, we use a collection of Landsat-5 images captured every second month from October 2009 to August 2010 over the protected area of the Doñana National Park in southwestern Spain, which is the largest sanctuary for migratory birds in western Europe. Results demonstrate that the proposed approach can detect the occurring changes in the main land-cover categories along the assessed time series. Daniela Espinoza-Molina, Reza Bahmanyar, Ricardo Díaz-Delgado, Javier Bustamante, Mihai Datcu |
IGARSS | 2 |
| 2017 | Discovery of Semantic Relationships in PolSAR Images Using Latent Dirichlet AllocationabstractWe propose a multilevel semantics discovery approach for bridging the semantic gap when mining high-resolution polarimetric synthetic aperture radar (PolSAR) remote sensing images. First, an Entropy/Anisotropy/Alpha-Wishart classifier is employed to discover low-level semantics as classes representing the physical scattering properties of targets (e.g., low-entropy/surface scattering/high anisotropy). Then, the images are tiled into patches and each patch is modeled as a bag-of-words, a histogram of the class labels. Next, latent Dirichlet allocation is applied to discover their higher level semantics as a set of topics. Our results demonstrate that topic semantics are close to human semantics used for basic land-cover types (e.g., grassland). Therefore, using the topic description (bag-of-topics) of PolSAR images leads to a narrower semantic gap in image mining. In addition, a visual exploration of the topic descriptions helps to find semantic relationships, which can be used for defining new semantic categories (e.g., mixed land-cover types) and designing rule-based categorization schemes. Radu Tanase, Reza Bahmanyar, Gottfried Schwarz, Mihai Datcu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Evaluating the sensory gap for earth observation images using human perception and an LDA-based computational modelabstractHigh resolution Earth Observation (EO) images contain detailed information, making it possible to recognize objects. However, issues such as the sensory gap (the difference between a real life scene and its sensory interpretation) cause difficulties for object recognition. In EO, this gap is rather wide due to sensor resolution, image perspective, scale and field of view (FOV). In this work, human perceptual and computational evaluations of the sensory gap are presented. For the human perceptual evaluation, user labels describing image patch content are gathered and analyzed. Results highlight issues caused by the sensory gap, e.g., FOV (image patch size) limits the contextual clues which can be used to disambiguate objects. The effect of FOV is then computationally analyzed as the difference between the scene context discovered by Latent Dirichlet Allocation from content within a certain FOV and the ground truth. Results indicate that increasing the FOV decreases the sensory gap. Reza Bahmanyar, Ambar Murillo |
ICIP | 1 |
| 2015 | A Comparative Study of Bag-of-Words and Bag-of-Topics Models of EO Image PatchesabstractThe large volume of detailed land cover features, provided by high resolution Earth observation (EO) images, has attracted considerable interest in the discovery of these features by learning systems. In this letter, we perform latent Dirichlet allocation on the bag of words (BoW) representation of collections of EO image patches to discover their semantic-level features, the so-called topics. To assess the discovered topics, the images are represented based on the occurrence of different topics, called bag of topics (BoT). The value added by BoT to the BoW model of image patches is then measured based on existing human annotations of the data. In our experiments, we compare the classification accuracy results of BoT and BoW representations of two different remote sensing image data sets, a multispectral optical data set and a synthetic-aperture-radar data set. Experimental results demonstrate that BoT can provide a compact and semantically meaningful representation of data; it either causes no significant reduction in the classification accuracy or increases the accuracy by a sufficient number of topics. Reza Bahmanyar, Shiyong Cui, Mihai Datcu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | The Semantic Gap: An Exploration of User and Computer Perspectives in Earth Observation ImagesabstractResearch on the semantic gap has considered differences between user and computer image interpretations and proposed methods to bridge it. These methods have been verified by comparing results to reference data or by measuring the degree of user acceptance. Although these methods result in a narrower semantic gap between computers and users, the resulting model for a specific user and search goal may still not be satisfactory to other users. Through an image annotation task with users, we find that this discrepancy is caused by the subjective biases present in the bridging methods, which we refer to as the “linguistic semantic gap.” Based on our findings, efforts to bridge the semantic gap should include different user perspectives to compensate the individual subjective biases, by increasing the diversity of data sets used in the domain. Moreover, models derived from proposed bridging methods could be stored and further used by other systems. Reza Bahmanyar, Ambar Murillo, Mihai Datcu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Farness preserving Non-negative matrix factorizationabstractDramatic growth in the volume of data made a compact and informative representation of the data highly demanded in computer vision, information retrieval, and pattern recognition. Non-negative Matrix Factorization (NMF) is used widely to provide parts-based representations by factorizing the data matrix into non-negative matrix factors. Since non-negativity constraint is not sufficient to achieve robust results, variants of NMF have been introduced to exploit the geometry of the data space. While these variants considered the local invariance based on the manifold assumption, we propose Farness preserving Non-negative Matrix Factorization (FNMF) to exploits the geometry of the data space by considering non-local invariance which is applicable to any data structure. FNMF adds a new constraint to enforce the far points (i.e., non-neighbors) in original space to stay far in the new space. Experiments on different kinds of data (e.g., Multimedia, Earth Observation) demonstrate that FNMF outperforms the other variants of NMF. Mohammadreza Babaee, Reza Bahmanyar, Gerhard Rigoll, Mihai Datcu |
ICIP | 2 |
| 2013 | Measuring the semantic gap based on a communication channel modelabstractThe collected Earth Observation (EO) data volumes are increasing immensely. In the meantime, the need for retrieval of focused information for decision making is increasing. Due to the particular nature of EO sensors, recording signals very differently than humans perceptual system, the challenges raised by the semantic and sensory gaps are immensely amplified in designing retrieval methods for EO images. This article introduces a method based on communication channel model to quantify and measure the semantic gap, used to assess various feature descriptors for semantic annotation purposes. The approach uses Latent Dirichlet Allocation (LDA), considering images as the source and the semantic topics as the receiver. The parameters of LDA are estimated for computing the Mutual Information to assess latent semantics of feature space. We further introduce a method to measure the distance between humans' and computer's semantics. The results are validated using an SVM-based classifier for an annotated dataset. Reza Bahmanyar, Mihai Datcu |
ICIP | 1 |