Adnan Firoze

dblp:84/10221 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0002-2751-009XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 73% 3D vision · 27%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
instance segmentation
1.722026
Building Instance Segmentation for Dense Urban Settlements · AAAI 2026
Tree Instance Segmentation with Temporal Contour Graph · CVPR 2023
Computer vision › Segmentation and scene understanding › building extraction
building instance segmentation
1.012026
Building Instance Segmentation for Dense Urban Settlements · AAAI 2026
Computer vision › 3D vision › remote sensing
remote sensing image analysis
1.012026
Building Instance Segmentation for Dense Urban Settlements · AAAI 2026
Multimedia analysis and retrieval
video analysis
0.212023
Tree Instance Segmentation with Temporal Contour Graph · CVPR 2023

Methods — techniques the papers use, named apart from their topics

temporal aggregation · 1.3graph convolutional network · 1.3contour graph · 1.3segment anything model 2 · 1.0graph neural network · 1.0
YearPublicationVenuePosition
2026 Building Instance Segmentation for Dense Urban Settlements
abstract
About 25% of the world’s population live in informal urban settlements containing densely packed buildings (approximately 8,000 houses per square-km) which do not lend themselves favorably to state-of-the-art satellite-based building segmentation methods due to, for example, occlusion, vegetation, shadows and low resolution. To address these challenges, we introduce a novel instance segmentation and counting approach for dense buildings. Our system first extracts a conservative set of tentative building center points using a deep network for jumpstarting a Segment Anything Model 2 (SAM2) module to produce an initial over-segmentation. Second, we use a graph neural network to refine the over-segmented regions into polygons representing accurate building masks. Experiments show that our approach achieves higher accuracy in instance segmentation and counting especially in challenging densely packed building areas in Brazil, Mexico, India, Pakistan, and Kenya, for instance.
Adnan Firoze, Raymond A. Yeh, Daniel G. Aliaga
AAAI1
2023 Tree Instance Segmentation with Temporal Contour Graph
abstract
We present a novel approach to perform instance segmentation and counting for densely packed self-similar trees using a top-view RGB image sequence. We propose a solution that leverages pixel content, shape, and self-occlusion. First, we perform an initial over-segmentation of the image sequence and aggregate structural characteristics into a contour graph with temporal information incorporated. Second, using a graph convolutional network and its inherent local messaging passing abilities, we merge adjacent tree crown patches into a final set of tree crowns. Per various studies and comparisons, our method is superior to all prior methods and results in high-accuracy instance segmentation and counting despite the trees being tightly packed. Finally, we provide various forest image sequence datasets suitable for subsequent benchmarking and evaluation captured at different altitudes and leaf conditions.
Adnan Firoze, Cameron Wingren, Raymond A. Yeh, Bedrich Benes, Daniel G. Aliaga
CVPR1
2022 Urban tree generator: spatio-temporal and generative deep learning for urban tree localization and modeling
Adnan Firoze, Bedrich Benes, Daniel G. Aliaga
Vis. Comput.1
2020 CAE: Towards Crowd Anarchism Exploration
abstract
Towards effective violence detection in densely crowded scenes, we introduce two novel computationally inexpensive real time pipelines. One is for automatic crowd violence detection and another for capturing the region with highest crowd concentration. The proposed violence detection architecture uses dense Histogram of Oriented Gradients (HOG) and dense Histogram of Motion Gradients (HMG) for feature extraction and Radial Bias Function Support Vector machine (RBF SVM) for classification. We further contribute by introducing a benchmark dataset, Dense Crowd Turbulence (DCT), having 120 videos of crowd violence and 120 for non-violence. DCT achieves an accuracy of 100% when evaluated with deep learning based violence detection frameworks. The violence detection architecture achieved a near the state of art accuracy of 87.3% on DCT.
Mayamin Hamid Raha, Tonmoay Deb, Mahieyin Rahmun, Shahriar Ali Bijoy, Adnan Firoze, Mohammad Ashrafuzzaman Khan
ICMLA5
2018 Scoring Photographic Rule of Thirds in a Large MIRFLICKR Dataset: A Showdown Between Machine Perception and Human Perception of Image Aesthetics
Adnan Firoze, Tousif Osman, Shahreen Shahjahan Psyche, Rashedur M. Rahman
ACIIDS (1)1
2018 Machine Cognition of Violence in Videos Using Novel Outlier-Resistant VLAD
abstract
Understanding highly accurate and real-time violent actions from surveillance videos is a demanding challenge. Our primary contribution of this work is divided into two parts. Firstly, we propose a computationally efficient Bag-of-Words (BoW) pipeline along with improved accuracy of violent videos classification. The novel pipeline's feature extraction stage is implemented with densely sampled Histogram of Oriented Gradients (HOG) and Histogram of Optical Flow (HOF) descriptors rather than Space-Time Interest Point (STIP) based extraction. Secondly, in encoding stage, we propose Outlier-Resistant VLAD (OR-VLAD), a novel higher order statistics-based feature encoding, to improve the original VLAD performance. In classification, efficient Linear Support Vector Machine (LSVM) is employed. The performance of the proposed pipeline is evaluated with three popular violent action datasets. On comparison, our pipeline achieved near perfect classification accuracies over three standard video datasets, outperforming most state-of-the-art approaches and having very low number of vocabulary size compared to previous BoW Models.
Tonmoay Deb, Aziz Arman, Adnan Firoze
ICMLA3
2015 Mining ICDDR, B Hospital Surveillance Data Using Locally Linear Embedding Based SMOTE Algorithm and Multilayer Perceptron
Adnan Firoze, Rashedur M. Rahman
ACIIDS (1)1