Saeed Anwar

dblp:176/1488 · DBLP profile ↗
← Back
58ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0002-0692-8411ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 8 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 20 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PointCaM: Cut-and-Mix for open-set point cloud learning
Shi Qiu 0001, Weihao Li 0005, Saeed Anwar, Mehrtash Harandi, Nick Barnes, Lars Petersson
Comput. Vis. Image Underst.4
2026 A lightweight model for perceptual image compression via implicit priors
Hao Wei 0005, Yiwen Jia, Chenyang Ge, Saeed Anwar, Ajmal Mian
Neural Networks5
2026 Correction: Vehicle and license plate recognition with novel dataset for toll collection
Hafeez Anwar, Abbas Anwar, Saeed Anwar
Pattern Anal. Appl.4
2026 Auto-Feedback Semantic Interface Learning for Text-Based Person Retrieval
Jiayi Li 0003, Min Jiang 0008, Jun Kong 0001, Saeed Anwar, Ajmal Mian
Pattern Recognit.5
2026 Representation-centric survey of supervised skeletal action recognition and the new benchmark
abstract
3D skeletal action recognition has emerged as a powerful alternative to traditional RGB and depth-based approaches, offering robustness to environmental variations, computational efficiency, and enhanced privacy. Despite remarkable progress, current research remains fragmented across diverse input representations and lacks evaluation under scenarios that reflect real-world challenges. This paper presents a representation-centric review of supervised skeletal action recognition, systematically categorizing state-of-the-art methods by their input feature types: joint coordinates, bone vectors, motion flows, and extended representations, and analyzing how these choices influence spatiotemporal modeling strategies. Building on the insights from this review, we introduce ANUBIS, a large-scale, challenging dataset designed to address critical gaps in existing benchmarks. ANUBIS incorporates multi-view recordings with back-view perspectives, complex multi-person interactions, fine-grained and violent actions, and contemporary social behaviors. We benchmark a diverse set of state-of-the-art models on ANUBIS and conduct an in-depth analysis of how different feature types affect recognition performance across 102 action categories. Our results show strong action-feature dependencies, highlight the limitations of naïve multi-representational fusion, and point toward the need for task-aware, semantically aligned integration strategies. This work offers both a comprehensive foundation and a practical benchmarking resource, aiming to guide the next generation of robust, generalizable skeleton-based action recognition systems for complex real-world scenarios. The dataset, benchmarking framework, and code are available at https://yliu1082.github.io/ANUBIS/ .
Yang Liu 0249, Jiyao Yang, Madhawa Perera, Pan Ji, Dongwoo Kim 0002, Min Xu 0009, Tianyang Wang 0004, Saeed Anwar, Tom Gedeon, Lei Wang 0108, Zhenyue Qin
Pattern Recognit.8
2026 DiffCom: Decoupled Sparse Priors Guided Diffusion Compression for Point Clouds
abstract
While conventional lossy compression methods predominantly depend on autoencoders to map point clouds into latent representations, they often neglect the intrinsic redundancy within these latent points. To address this limitation, this paper presents a diffusion-based architecture steered by sparse priors, designed to minimize latent redundancy while securing superior reconstruction fidelity, particularly in low-bitrate scenarios. A key feature of the framework is an efficient dual-density data flow that alleviates the stringent size constraints imposed on latent points. By integrating a Probabilistic Attention-based Conditional Denoiser (PACD), the method effectively encapsulates critical reconstruction details within sparse priors, which are hierarchically decoupled into intra- and inter-point components. Specifically, separate encoders are utilized to transform the source point cloud into latent points and decoupled sparse priors, respectively. To dynamically exploit geometric and semantic information, an attention-driven latent denoiser, conditioned on these decoupled priors, is applied across the encoding and decoding layers. Furthermore, inter-point distributions are incorporated into the arithmetic codec to refine local context modeling for sparse points, with the final point cloud recovered via a point decoder. Comprehensive experiments conducted on ShapeNet and standard MPEG PCC datasets demonstrate that the proposed method outperforms state-of-the-art techniques, achieving a superior rate-distortion trade-off.
Xiaoge Zhang 0003, Mingtao Feng, Mehwish Nasim, Saeed Anwar, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.5
2026 SPEGNet: Synergistic Perception-Guided Network for Camouflaged Object Detection
abstract
Camouflaged object detection segments objects with intrinsic similarity and edge disruption. Current detection methods rely on accumulated complex components. Each approach adds components such as boundary modules, attention mechanisms, and multi-scale processors independently. This accumulation creates a computational burden without proportional gains. To manage this complexity, they process at reduced resolutions, eliminating fine details essential for camouflage. We present SPEGNet, addressing fragmentation through synergistic design. The architecture integrates multi-scale features via channel calibration and spatial enhancement. Boundaries emerge directly from context-rich representations, maintaining semantic-spatial alignment. Progressive refinement implements scale-adaptive edge modulation with peak influence at intermediate resolutions. This design strikes a balance between boundary precision and regional consistency. SPEGNet achieves $0.887~S_{\alpha } $ on CAMO, 0.890 on COD10K, and 0.895 on NC4K, with real-time inference speed. Our approach excels across scales, from tiny, intricate objects to large, pattern-similar ones, while handling occlusion and ambiguous boundaries. Code, model weights, and results are available at https://github.com/Baber-Jan/SPEGNet.
Baber Jan, Saeed Anwar, Aiman H. El-Maleh, Abdul Jabbar Siddiqui, Abdul Bais
IEEE Trans. Image Process.2
2025 IARD: Intruder Activity Recognition Dataset for Threat Detection
abstract
Home security and surveillance systems are rapidly evolving, with Artificial Intelligence (AI) playing a transformative role in enhancing safety and threat detection. While several AI methods and datasets for intruder-related risk assessment exist, they predominantly focus on face detection and recognition, leaving a significant gap in addressing high-risk scenarios involving malicious intent, such as theft or harm. The lack of dedicated datasets for recognizing complex intruder activities, such as carrying weapons or engaging in destructive actions like kicking doors or breaking locks, limits the development of robust solutions. This work bridges this gap by introducing the Intruder Activity Recognition Dataset (IARD), a video dataset specifically designed to recognize four critical intruder activities: Armed Intruder, Door Kick, Intruder Inside and Lock Breaking. Leveraging IARD, we thoroughly benchmark various state-of-the-art methods, among which a Vision Transformer is found to achieve an impressive 93.3% accuracy in recognizing intruder actions. Our contribution highlights the potential of IARD in advancing AI-driven surveillance systems, providing a foundational dataset and benchmark for recognizing complex intruder activities.
Shehzad Ali, Md Tanvir Islam, Ikhyun Lee, Saeed Anwar, Javier Del Ser, Khan Muhammad 0001
CIKM4
2025 Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes
abstract
Fusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and RGB input has started gaining traction. However, effective integration of these modalities for precise object detection tasks still remains a largely open problem. To address that, we propose a MultiStream Detection (MuStD) network, which meticulously extracts task-relevant information from both data modalities. The network follows a three-stream structure. Its LiDAR-PillarNet stream extracts sparse 2D pillar features from the LiDAR input while the LiDAR-Height Compression stream computes Bird’s-Eye View features. An additional 3D Multimodal stream combines RGB and LiDAR features using UV mapping and polar coordinate indexing. Eventually, the features containing comprehensive spatial, textural, and geometric information are carefully fused and fed to a detection head for 3D object detection. We evaluate our method on the challenging KITTI Object Detection Benchmark, with results available on the official evaluation server.1. Our approach achieves strong performance, with an average precision (AP) of 85.39% in 3D detection, 91.34% in Bird’s Eye View (BEV) detection, and 96.39% in 2D detection. These results match or surpass existing state-of-the-art methods. In the difficult "Hard" category, our method attains 80.78% AP in 3D detection and 94.04% AP in 2D detection, highlighting its robustness in challenging scenarios. Furthermore, our method runs at 67 ms, demonstrating efficiency and real-time capability. Our code will be released through the MuStD GitHub repository at https://github.com/IbrahimUWA/MuStD.
Muhammad Ibrahim 0001, Naveed Akhtar, Haitian Wang 0002, Saeed Anwar, Ajmal Mian
IROS4
2025 Towards Hazardous Activity Recognition for A Novel Real-World Dataset
abstract
Detecting hazardous activities is essential for ensuring safety. However, existing datasets often lack coverage of the nuanced and diverse hazards present in indoor environments, which hinders the development of a specialized model. To address this, we introduce the Real-World Hazardous Activities Dataset (RHAD), a novel and diverse video dataset specifically curated for recognizing hazardous activities in real-world indoor settings. Leveraging RHAD, we introduce HazardNet, a hybrid deep-learning architecture designed for hazardous activity recognition. HazardNet integrates local and global spatial-temporal representation modules to effectively capture complex patterns, enabling a robust understanding of the activity. We perform comprehensive evaluations by benchmarking against a range of state-of-the-art activity recognition models. Experimental results show that our proposed model performs significantly better, surpassing the latest model, VideoMamba, with a 9.2% accuracy gain. Moreover, by providing the dataset and an effective recognition model, our work lays the foundation for further research, paving the way for enhanced safety measures and preventive interventions. The dataset and code are available at https://github.com/ShehzadCS18/RHAD.
Shehzad Ali, Md Tanvir Islam, Ikhyun Lee, Mingfu Xiong, Minh-Son Dao, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia6
2025 Bird eye-view to street-view: A survey
Khawlah Bajbaa, Muhammad Usman 0010, Saeed Anwar, Ibrahim Radwan, Abdul Bais
Appl. Intell.3
2025 FSBI: Deepfake detection with frequency enhanced self-blended images
Ahmed Abul Hasanaath, Hamzah Luqman, Raed Katib, Saeed Anwar
Image Vis. Comput.4
2025 Vehicle and license plate recognition with novel dataset for toll collection
Hafeez Anwar, Saeed Anwar
Pattern Anal. Appl.3
2025 A Comprehensive Overview of Large Language Models
abstract
Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multimodal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to provide not only a systematic survey but also a quick, comprehensive reference for the researchers and practitioners to draw insights from extensive, informative summaries of the existing works to advance the LLM research.
Humza Naveed, Asad Ullah Khan, Shi Qiu 0001, Saeed Anwar, Muhammad Usman 0010, Naveed Akhtar, Nick Barnes, Ajmal Mian
ACM Trans. Intell. Syst. Technol.5
2025 Attention-Based Real Image Restoration
abstract
Deep convolutional neural networks perform better on images containing spatially invariant degradations, also known as synthetic degradations; however, their performance is limited on real-degraded photographs and requires multiple-stage network modeling. To advance the practicability of restoration algorithms, this article proposes a novel single-stage blind real image restoration network ( Net) by employing a modular architecture. We use a residual on the residual structure to ease low-frequency information flow and apply feature attention to exploit the channel dependencies. Furthermore, the evaluation in terms of quantitative metrics and visual quality for four restoration tasks, i.e., denoising, super-resolution, raindrop removal, and JPEG compression on 11 real degraded datasets against more than 30 state-of-the-art algorithms, demonstrates the superiority of our Net. We also present the comparison on three synthetically generated degraded datasets for denoising to showcase our method's capability on synthetics denoising. The codes, trained models, and results are available on https://github.com/saeed-anwar/R2Net.
Saeed Anwar, Nick Barnes, Lars Petersson
IEEE Trans. Neural Networks Learn. Syst.1
2025 Position-Sensing Graph Neural Networks: Proactively Learning Nodes Relative Positions
abstract
Most existing graph neural networks (GNNs) learn node embeddings using the framework of message passing and aggregation. Such GNNs are incapable of learning relative positions between graph nodes within a graph. To empower GNNs with the awareness of node positions, some nodes are set as anchors. Then, using the distances from a node to the anchors, GNNs can infer relative positions between nodes. However, position-aware GNNs (P-GNNs) arbitrarily select anchors, leading to compromising position awareness and feature extraction. To eliminate this compromise, we demonstrate that selecting evenly distributed and asymmetric anchors is essential. On the other hand, we show that choosing anchors that can aggregate embeddings of all the nodes within a graph is NP-complete. Therefore, devising efficient optimal algorithms in a deterministic approach is practically not feasible. To ensure position awareness and bypass NP-completeness, we propose position-sensing GNNs (PSGNNs), learning how to choose anchors in a backpropagatable fashion. Experiments verify the effectiveness of PSGNNs against state-of-the-art GNNs, substantially improving performance on various synthetic and real-world graph datasets while enjoying stable scalability. Specifically, PSGNNs on average boost area under the curve (AUC) more than 14% for pairwise node classification and 18% for link prediction over the existing state-of-the-art position-aware methods. Our source code is publicly available at: https://github.com/ZhenyueQin/PSGNN.
Zhenyue Qin, Saeed Anwar, Dongwoo Kim 0002, Yang Liu 0249, Pan Ji, Tom Gedeon
IEEE Trans. Neural Networks Learn. Syst.3
2024 LoLI-Street: Benchmarking Low-Light Image Enhancement and Beyond
Md Tanvir Islam, Inzamamul Alam, Simon S. Woo, Saeed Anwar, Ikhyun Lee, Khan Muhammad 0001
ACCV (5)4
2024 Dual Deep Learning Network for Abnormal Action Detection
abstract
Neural networks have demonstrated remarkable effectiveness in solving distinct real-world vision problems pertaining to activity recognition and violence detection in surveillance scenarios. The broad reliance on practicing a single network for spatial and motion information collection has made them less effective for long-term dependency analysis in video snippets. Our work solves this issue through a multi-network fusion strategy suitable for real-world surveillance. Initially, the spatial information is accessed from a compound coefficient strategy inspired by a robust convolutional neural network (ConvNet). Next, the pyramidal convolutional features from two consecutive frames are obtained through LiteFlowNet. The output from both the networks (ConvNet and LiteFlowNet) is separately passed into a deep-gated recurrent Unit (GRU) that is assembled for a skip connection. The latter obtained from each GRU is fused and further propagated to the dense layer for final decision. The results on the datasets and the ablation study confirm our method’s efficiency, outperforming the state-of-the-art methods. (Code: GitHub)
Fath U Min Ullah, Zulfiqar Ahmad Khan 0002, Sung Wook Baik, Estefanía Talavera, Saeed Anwar, Khan Muhammad 0001
AVSS5
2024 Attention Down-Sampling Transformer, Relative Ranking and Self-Consistency For Blind Image Quality Assessment
abstract
The no-reference image quality assessment is a challenging domain that addresses estimating image quality without the original reference. We introduce an improved mechanism to extract local and non-local information from images via different transformer encoders and CNNs. The utilization of Transformer encoders aims to mitigate locality bias and generate a non-local representation by sequentially processing CNN features, which inherently capture local visual structures. Establishing a stronger connection between subjective and objective assessments is achieved through sorting within batches of images based on relative distance information. A self-consistency approach to self-supervision is presented, explicitly addressing the degradation of no-reference image quality assessment (NR-IQA) models under equivariant transformations. Our approach ensures model robustness by maintaining consistency between an image and its horizontally flipped equivalent. Through empirical evaluation of five popular image quality assessment datasets, the proposed model outperforms alternative algorithms in the context of no-reference image quality assessment datasets, especially on smaller datasets. Codes are available at https://github.com/mas94/ADTRS
Mohammed Alsaafin, Musab Alsheikh, Saeed Anwar
ICIP3
2024 AllWeather-Net: Unified Image Enhancement for Autonomous Driving Under Adverse Weather and Low-Light Conditions
Chenghao Qian, Mahdi Rezaei 0001, Saeed Anwar, Wenjing Li 0005, Tanveer Hussain 0001, Mohsen Azarmi, Wei Wang 0335
ICPR (30)3
2024 HazeSpace2M: A Dataset for Haze Aware Single Image Dehazing
abstract
Reducing the atmospheric haze and enhancing image clarity is crucial for computer vision applications. The lack of real-life hazy ground truth images necessitates synthetic datasets, which often lack diverse haze types, impeding effective haze type classification and dehazing algorithm selection. This research introduces the HazeSpace2M dataset, a collection of over 2 million images designed to enhance dehazing through haze type classification. HazeSpace2M includes diverse scenes with 10 haze intensity levels, featuring Fog, Cloud, and Environmental Haze (EH). Using the dataset, we introduce a technique of haze type classification followed by specialized dehazers to clear hazy images. Unlike conventional methods, our approach classifies haze types before applying type-specific dehazing, improving clarity in real-life hazy images. Benchmarking with state-of-the-art (SOTA) models, ResNet50 and AlexNet achieve 92.75\% and 92.50\% accuracy, respectively, against existing synthetic datasets. However, these models achieve only 80% and 70% accuracy, respectively, against our Real Hazy Testset (RHT), highlighting the challenging nature of our HazeSpace2M dataset. Additional experiments show that haze type classification followed by specialized dehazing improves results by 2.41% in PSNR, 17.14% in SSIM, and 10.2\% in MSE over general dehazers. Moreover, when testing with SOTA dehazing models, we found that applying our proposed framework significantly improves their performance. These results underscore the significance of HazeSpace2M and our proposed framework in addressing atmospheric haze in multimedia processing. Complete code and dataset is available on \href{https://github.com/tanvirnwu/HazeSpace2M} {\textcolor{blue}{\textbf{GitHub}}}.
Md Tanvir Islam, Nasir Rahim, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia3
2024 Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action Detection
abstract
Action detection and understanding provide the foundation for the generation and interaction of multimedia content. However, existing methods mainly focus on constructing complex relational inference networks, overlooking the judgment of detection effectiveness. Moreover, these methods frequently generate detection results with cognitive abnormalities. To solve the above problems, this study proposes a cognitive effectiveness network based on fuzzy inference (Cefdet), which introduces the concept of 'cognition--based detection' to simulate human cognition. First, a fuzzy-driven cognitive effectiveness evaluation module (FCM) is established to introduce fuzzy inference into action detection. FCM is combined with human action features to simulate the cognition-based detection process, which clearly locates the position of frames with cognitive abnormalities. Then, a fuzzy cognitive update strategy (FCS) is proposed based on the FCM, which utilizes fuzzy logic to re-detect the cognition-based detection results and effectively update the results with cognitive abnormalities. Experimental results demonstrate that Cefdet exhibits superior performance against several mainstream algorithms on the public datasets, validating its effectiveness and superiority.
Weina Fu, Shuai Liu 0002, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia4
2024 Survey: Image mixing and deleting for data augmentation
Humza Naveed, Saeed Anwar, Munawar Hayat, Kashif Javed, Ajmal Mian
Eng. Appl. Artif. Intell.2
2024 Fusing Higher-Order Features in Graph Neural Networks for Skeleton-Based Action Recognition
abstract
Skeleton sequences are lightweight and compact and thus are ideal candidates for action recognition on edge devices. Recent skeleton-based action recognition methods extract features from 3-D joint coordinates as spatial-temporal cues, using these representations in a graph neural network for feature fusion to boost recognition performance. The use of first- and second-order features, that is, joint and bone representations, has led to high accuracy. Nonetheless, many models are still confused by actions that have similar motion trajectories. To address these issues, we propose fusing higher-order features in the form of angular encoding (AGE) into modern architectures to robustly capture the relationships between joints and body parts. This simple fusion with popular spatial-temporal graph neural networks achieves new state-of-the-art accuracy in two large benchmarks, including NTU60 and NTU120, while employing fewer parameters and reduced run time. Our source code is publicly available at: https://github.com/ZhenyueQin/Angular-Skeleton-Encoding.
Zhenyue Qin, Yang Liu 0249, Pan Ji, Dongwoo Kim 0002, Lei Wang 0108, Robert I. McKay, Saeed Anwar, Tom Gedeon
IEEE Trans. Neural Networks Learn. Syst.7
2023 P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
abstract
Point cloud completion aims to recover the complete shape based on a partial observation. Existing methods require either complete point clouds or multiple partial observations of the same object for learning. In contrast to previous approaches, we present Partial2Complete (P2C), the first self-supervised framework that completes point cloud objects using training samples consisting of only a single incomplete point cloud per object. Specifically, our framework groups incomplete point clouds into local patches as input and predicts masked patches by learning prior information from different partial objects. We also propose Region-Aware Chamfer Distance to regularize shape mismatch without limiting completion capability, and devise the Normal Consistency Constraint to incorporate a local planarity assumption, encouraging the recovered shape surface to be continuous and complete. In this way, P2C no longer needs multiple observations or complete point clouds as ground truth. Instead, structural cues are learned from a category-specific dataset to complete partial point clouds of objects. We demonstrate the effectiveness of our approach on both synthetic ShapeNet data and real-world ScanNet data, showing that P2C produces comparable results to methods trained with complete shapes, and outperforms methods learned with multiple partial observations. Code is available at https://github.com/CuiRuikai/Partial2Complete.
Ruikai Cui, Shi Qiu 0001, Saeed Anwar, Jiawei Liu 0005, Chaoyue Xing, Jing Zhang 0052, Nick Barnes
ICCV3
2023 Slice Transformer and Self-supervised Learning for 6DoF Localization in 3D Point Cloud Maps
abstract
Precise localization is critical for autonomous vehicles. We present a self-supervised learning method that employs transformers for the first time for the task of outdoor localization using LiDAR data. We propose a pre-text task that reorganizes the slices of a 360° LiDAR scan to leverage its axial properties. Our model, called Slice Transformer, employs multi-head attention while systematically processing the slices. To the best of our knowledge, this is the first instance of leveraging multi-head attention for outdoor point clouds. We additionally introduce the Perth-Wadataset, which provides a large-scale LiDAR map of Perth city in Western Australia, covering ~4km2area. Localization annotations are provided for Perth - Wa.The proposed localization method is thoroughly evaluated on Perth-WA and Appollo-SouthBay datasets. We also establish the efficacy of our self-supervised learning approach for the common downstream task of object classification using ModelNet40 and ScanNN datasets. The code and Perth-WA data will be publicly released.
Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Michael J. Wise, Ajmal Mian
ICRA3
2023 UnLoc: A Universal Localization Method for Autonomous Vehicles using LiDAR, Radar and/or Camera Input
abstract
Localization is a fundamental task in robotics for autonomous navigation. Existing localization methods rely on a single input data modality or train several computational models to process different modalities. This leads to stringent computational requirements and sub-optimal results that fail to capitalize on the complementary information in other data streams. This paper proposes UnLoc, a novel unified neural modeling approach for localization with multi-sensor input in all weather conditions. Our multi-stream network can handle LiDAR, Camera and RADAR inputs for localization on demand, i.e., it can work with one or more input sensors, making it robust to sensor failure. UnLoc uses 3D sparse convolutions and cylindrical partitioning of the space to process LiDAR frames and implements ResNet blocks with a slot attention-based feature filtering module for the Radar and image modalities. We introduce a unique learnable modality encoding scheme to distinguish between the input sensor data. Our method is extensively evaluated on Oxford Radar RobotCar, ApolloSouthBay and Perth-WA datasets. The results ascertain the efficacy of our technique. The dataset, results, and codes are available at https://github.com/IbrahimUWA/UnLoc
Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Ajmal Mian
IROS3
2023 PnP-3D: A Plug-and-Play for 3D Point Clouds
abstract
With the help of the deep learning paradigm, many point cloud networks have been invented for visual analysis. However, there is great potential for development of these networks since the given information of point cloud data has not been fully exploited. To improve the effectiveness of existing networks in analyzing point cloud data, we propose a plug-and-play module, PnP-3D, aiming to refine the fundamental point cloud feature representations by involving more local context and global bilinear response from explicit 3D space and implicit feature space. To thoroughly evaluate our approach, we conduct experiments on three standard point cloud analysis tasks, including classification, semantic segmentation, and object detection, where we select three state-of-the-art networks from each task for evaluation. Serving as a plug-and-play module, PnP-3D can significantly boost the performances of established networks. In addition to achieving state-of-the-art results on four widely used point cloud benchmarks, we present comprehensive ablation studies and visualizations to demonstrate our approach's advantages. The code will be available at https://github.com/ShiQiu0419/pnp-3d.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 SAT3D: Slot Attention Transformer for 3D Point Cloud Semantic Segmentation
abstract
Semantic segmentation of 3D point cloud is a key task in numerous intelligent transportation system applications, e.g., self-driving vehicles, traffic monitoring. Due to the sparsity and varying density of points in the outdoor point clouds, it becomes particularly challenging to extract object-centric features from data. This leads to poor semantic segmentation, especially for the rare object classes. To address that, we introduce the first-ever Slot Attention Transformer based technique to effectively model object-centric features in point cloud data. Our method uses cylindrical splits of space for voxelization and computes channel-wise positional embeddings before repetitively encoding the point cloud with slot attentions. Our second major contribution is a Large-Scale Outdoor Point Cloud dataset (SWAN), collected in a dense urban environment, driving 150km distance. It provides 16 billion points in more than 200K frames. The dataset also provides annotations for 10K frames for 24 classes. We also contribute a data augmentation scheme to handle rare object classes in real-world point clouds. Besides benchmarking popular existing methods on SWAN for the first time, we thoroughly evaluate our technique on the existing large-scale datasets, Semantic KITTI and nuScenes. Our results demonstrate a consistent performance gain for our technique, and verify the need of the more challenging SWAN dataset.
Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Ajmal Mian
IEEE Trans. Intell. Transp. Syst.3
2022 PU-Transformer: Point Cloud Upsampling Transformer
Shi Qiu 0001, Saeed Anwar, Nick Barnes
ACCV (1)2
2022 Energy-Based Residual Latent Transport for Unsupervised Point Cloud Completion
Ruikai Cui, Shi Qiu 0001, Saeed Anwar, Jing Zhang 0052, Nick Barnes
BMVC3
2022 Image Dehazing Transformer with Transmission-Aware 3D Position Embedding
abstract
Despite single image dehazing has been made promising progress with Convolutional Neural Networks (CNNs), the inherent equivariance and locality of convolution still bottleneck deharing performance. Though Transformer has occupied various computer vision tasks, directly leveraging Transformer for image dehazing is challenging: 1) it tends to result in ambiguous and coarse details that are undesired for image reconstruction; 2) previous position embedding of Transformer is provided in logic or spatial position order that neglects the variational haze densities, which results in the sub-optimal dehazlng performance. The key insight of this study is to investigate how to combine CNN and Transformer for image dehazing. To solve the feature inconsistency issue between Transformer and CNN, we propose to modulate CNN features via learning modulation matrices (i.e., coefficient matrix and bias matrix) conditioned on Transformer features instead of simple feature addition or concatenation. The feature modulation naturally inherits the global context modeling capability of Transformer and the local representation capability of CNN. We bring a haze density-related prior into Trans-former via a novel transmission-aware 3D position embedding module, which not only provides the relative position but also suggests the haze density of different spatial regions. Extensive experiments demonstrate that our method, DeHamer, attains state-of-the-art performance on several image dehazing benchmarks.
Chunle Guo, Qixin Yan, Saeed Anwar, Runmin Cong, Wenqi Ren, Chongyi Li
CVPR3
2022 Deep localization of subcellular protein structures from fluorescence microscopy images
Muhammad Tahir 0006, Saeed Anwar, Ajmal Mian, Abdul Wahab Muzaffar
Neural Comput. Appl.2
2022 Densely Residual Laplacian Super-Resolution
abstract
Super-Resolution convolutional neural networks have recently demonstrated high-quality restoration for single images. However, existing algorithms often require very deep architectures and long training times. Furthermore, current convolutional neural networks for super-resolution are unable to exploit features at multiple scales and weigh them equally or at only static scale only, limiting their learning capability. In this exposition, we present a compact and accurate super-resolution algorithm, namely, densely residual laplacian network (DRLN). The proposed network employs cascading residual on the residual structure to allow the flow of low-frequency information to focus on learning high and mid-level features. In addition, deep supervision is achieved via the densely concatenated residual blocks settings, which also helps in learning from high-level complex features. Moreover, we propose Laplacian attention to model the crucial features to learn the inter and intra-level dependencies between the feature maps. Furthermore, comprehensive quantitative and qualitative evaluations on low-resolution, noisy low-resolution, and real historical image benchmark datasets illustrate that our DRLN algorithm performs favorably against the state-of-the-art methods visually and accurately.
Saeed Anwar, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Uncertainty Inspired RGB-D Saliency Detection
abstract
We propose the first stochastic framework to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection models treat this task as a point estimation problem by predicting a single saliency map following a deterministic learning pipeline. We argue that, however, the deterministic solution is relatively ill-posed. Inspired by the saliency data labeling process, we propose a generative architecture to achieve probabilistic RGB-D saliency detection which utilizes a latent variable to model the labeling variations. Our framework includes two main models: 1) a generator model, which maps the input image and latent variable to stochastic saliency prediction, and 2) an inference model, which gradually updates the latent variable by sampling it from the true or approximate posterior distribution. The generator model is an encoder-decoder saliency network. To infer the latent variable, we introduce two different solutions: i) a Conditional Variational Auto-encoder with an extra encoder to approximate the posterior distribution of the latent variable; and ii) an Alternating Back-Propagation technique, which directly samples the latent variable from the true posterior distribution. Qualitative and quantitative results on six challenging RGB-D benchmark datasets show our approach's superior performance in learning the distribution of saliency maps. The source code is publicly available via our project page: https://github.com/JingZhang617/UCNet.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Geometric Back-Projection Network for Point Cloud Classification
abstract
As the basic task of point cloud analysis, classification is fundamental but always challenging. To address some unsolved problems of existing methods, we propose a network that captures geometric features of point clouds for better representations. To achieve this, on the one hand, we enrich the geometric information of points in low-level 3D space explicitly. On the other hand, we apply CNN-based structures in high-level feature spaces to learn local geometric context implicitly. Specifically, we leverage an idea of error-correcting feedback structure to capture the local features of point clouds comprehensively. Furthermore, an attention module based on channel affinity assists the feature map to avoid possible redundancy by emphasizing its distinct channels. The performance on both synthetic and real-world point clouds datasets demonstrate the superiority and applicability of our network. Comparing with other state-of-the-art methods, our approach balances accuracy and efficiency.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
IEEE Trans. Multim.2
2021 Investigating Attention Mechanism in 3D Point Cloud Object Detection
abstract
Object detection in three-dimensional (3D) space attracts much interest from academia and industry since it is an essential task in AI-driven applications such as robotics, autonomous driving, and augmented reality. As the basic format of 3D data, the point cloud can provide detailed geometric information about the objects in the original 3D space. However, due to 3D data s sparsity and unorderedness, specially designed networks and modules are needed to process this type of data. Attention mechanism has achieved impressive performance in diverse computer vision tasks; however, it is unclear how attention modules would affect the performance of 3D point cloud object detection and what sort of attention modules could fit with the inherent properties of 3D data. This work investigates the role of the attention mechanism in 3D point cloud object detection and provides insights into the potential of different attention modules. To achieve that, we comprehensively investigate classical 2D attentions, novel 3D attentions, including the latest point cloud transformers on SUN RGB-D and ScanNetV2 datasets. Based on the detailed experiments and analysis, we conclude the effects of different attention modules. This paper is expected to serve as a reference source for benefiting attention-embedded 3D point cloud object detection. The code and trained models are available at: https://qithub.com/SkiQiu0419/attentions_in_3D_detection.
Shi Qiu 0001, Saeed Anwar, Chongyi Li
3DV3
2021 Invertible Denoising Network: A Light Solution for Real Noise Removal
abstract
Invertible networks have various benefits for image de-noising since they are lightweight, information-lossless, and memory-saving during back-propagation. However, applying invertible models to remove noise is challenging because the input is noisy, and the reversed output is clean, following two different distributions. We propose an invertible denoising network, InvDN, to address this challenge. InvDN transforms the noisy input into a low-resolution clean image and a latent representation containing noise. To discard noise and restore the clean image, InvDN replaces the noisy latent representation with another one sampled from a prior distribution during reversion. The de-noising performance of InvDN is better than all the existing competitive models, achieving a new state-of-the-art result for the SIDD dataset while enjoying less run time. Moreover, the size of InvDN is far smaller, only having 4.2% of the number of parameters compared to the most recently proposed DANet. Further, via manipulating the noisy latent representation, InvDN is also able to generate noise more similar to the original one. Our code is available at: https://github.com/Yang-Liu1082/InvDN.git.
Yang Liu 0249, Zhenyue Qin, Saeed Anwar, Pan Ji, Dongwoo Kim 0002, Sabrina B. Caldwell, Tom Gedeon
CVPR3
2021 Semantic Segmentation for Real Point Cloud Scenes via Bilateral Augmentation and Adaptive Fusion
abstract
Given the prominence of current 3D sensors, a fine-grained analysis on the basic point cloud data is worthy of further investigation. Particularly, real point cloud scenes can intuitively capture complex surroundings in the real world, but due to 3D data’s raw nature, it is very challenging for machine perception. In this work, we concentrate on the essential visual task, semantic segmentation, for large-scale point cloud data collected in reality. On the one hand, to reduce the ambiguity in nearby points, we augment their local context by fully utilizing both geometric and semantic features in a bilateral structure. On the other hand, we comprehensively interpret the distinctness of the points from multiple resolutions and represent the feature map following an adaptive fusion method at point-level for accurate semantic segmentation. Further, we provide specific ablation studies and intuitive visualizations to validate our key modules. By comparing with state-of-the-art networks on three different benchmarks, we demonstrate the effectiveness of our network.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
CVPR2
2021 Single Underwater Image Restoration by Contrastive Learning
abstract
Underwater image restoration attracts significant attention due to its importance in unveiling the underwater world. This paper elaborates on a novel method that achieves state-of-the-art results for underwater image restoration based on the unsupervised image-to-image translation framework. We design our method by leveraging from contrastive learning and generative adversarial networks to maximize mutual information between raw and restored images. Additionally, we release a large-scale real underwater image dataset to support both paired and unpaired training modules. Extensive experiments with comparisons to recent approaches further demonstrate the superiority of our proposed method.
Junlin Han, Mehrdad Shoeiby, Timothy J. Malthus, Elizabeth J. Botha, Janet M. Anstee, Saeed Anwar, Lars Petersson, Mohammad Ali Armin
IGARSS6
2021 Dense-Resolution Network for Point Cloud Classification and Segmentation
abstract
Point cloud analysis is attracting attention from Artificial Intelligence research since it can be widely used in applications such as robotics, Augmented Reality, self-driving. However, it is always challenging due to irregularities, unorderedness, and sparsity. In this article, we propose a novel network named Dense-Resolution Network (DRNet) for point cloud analysis. Our DRNet is designed to learn local point features from the point cloud in different resolutions. In order to learn local point groups more effectively, we present a novel grouping method for local neighborhood searching and an error-minimizing module for capturing local features. In addition to validating the network on widely used point cloud segmentation and classification benchmarks, we also test and visualize the performance of the components. Comparing with other state-of-the-art methods, our network shows superiority on ModelNet40, ShapeNet synthetic and ScanObjectNN real point cloud datasets.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
WACV2
2021 Deblur and deep depth from single defocus image
Saeed Anwar, Zeeshan Hayder, Fatih Porikli
Mach. Vis. Appl.1
2021 Multi-FAN: multi-spectral mosaic super-resolution via multi-scale feature aggregation network
Mehrdad Shoeiby, Mohammad Sadegh Ali Akbarian, Saeed Anwar, Lars Petersson
Mach. Vis. Appl.3
2021 Deep ancient Roman Republican coin classification via feature fusion and attention
Hafeez Anwar, Saeed Anwar, Sebastian Zambanini, Fatih Porikli
Pattern Recognit.2
2021 Underwater Image Enhancement via Medium Transmission-Guided Multi-Color Space Embedding
abstract
Underwater images suffer from color casts and low contrast due to wavelength- and distance-dependent attenuation and scattering. To solve these two degradation issues, we present an underwater image enhancement network via medium transmission-guided multi-color space embedding, called Ucolor. Concretely, we first propose a multi-color space encoder network, which enriches the diversity of feature representations by incorporating the characteristics of different color spaces into a unified structure. Coupled with an attention mechanism, the most discriminative features extracted from multiple color spaces are adaptively integrated and highlighted. Inspired by underwater imaging physical models, we design a medium transmission (indicating the percentage of the scene radiance reaching the camera)-guided decoder network to enhance the response of network towards quality-degraded regions. As a result, our network can effectively improve the visual quality of underwater images by exploiting multiple color spaces embedding and the advantages of both physical model-based and learning-based methods. Extensive experiments demonstrate that our Ucolor achieves superior performance against state-of-the-art methods in terms of both visual quality and quantitative metrics. The code is publicly available at: https://li-chongyi.github.io/Proj_Ucolor.html.
Chongyi Li, Saeed Anwar, Junhui Hou, Runmin Cong, Chunle Guo, Wenqi Ren
IEEE Trans. Image Process.2
2021 Light-DehazeNet: A Novel Lightweight CNN Architecture for Single Image Dehazing
abstract
Due to the rapid development of artificial intelligence technology, industrial sectors are revolutionizing in automation, reliability, and robustness, thereby significantly increasing quality and productivity. Most of the surveillance and industrial sectors are monitored by visual sensor networks capturing different surrounding environment images. However, during tempestuous weather conditions, the visual quality of the images is reduced due to contaminated suspended atmospheric particles that affect the overall surveillance systems. To tackle these challenges, this article presents a computationally efficient lightweight convolutional neural network referred to as Light-DehazeNet (LD-Net) for the reconstruction of hazy images. Unlike other learning-based approaches, which separately measure the transmission map and the atmospheric light, our proposed LD-Net jointly estimates both the transmission map and the atmospheric light using a transformed atmospheric scattering model. Furthermore, a color visibility restoration method is proposed to evade the color distortion in the dehaze image. Finally, we conduct extensive experiments using synthetic and natural hazy images. The quantitative and qualitative evaluation on different benchmark hazy datasets verify the superiority of the proposed method over other state-of-the-art image dehazing techniques. Moreover, additional experimentation validates the applicability of the proposed method in the object detection tasks. Considering the lightweight architecture with minimal computational cost, the proposed system is encouraged to be incorporated as an integral part of the vision-based monitoring systems to improve the overall performance.
Hayat Ullah, Khan Muhammad 0001, Saeed Anwar, Ali Shariq Imran, Victor Hugo C. de Albuquerque
IEEE Trans. Image Process.4
2020 From Depth What Can You See? Depth Completion via Auxiliary Image Reconstruction
abstract
Depth completion recovers dense depth from sparse measurements, e.g., LiDAR. Existing depth-only methods use sparse depth as the only input. However, these methods may fail to recover semantics consistent boundaries, or small/thin objects due to 1) the sparse nature of depth points and 2) the lack of images to provide semantic cues. This paper continues this line of research and aims to overcome the above shortcomings. The unique design of our depth completion model is that it simultaneously outputs a reconstructed image and a dense depth map. Specifically, we formulate image reconstruction from sparse depth as an auxiliary task during training that is supervised by the unlabelled gray-scale images. During testing, our system accepts sparse depth as the only input, i.e., the image is not required. Our design allows the depth completion network to learn complementary image features that help to better understand object structures. The extra supervision incurred by image reconstruction is minimal, because no annotations other than the image are needed. We evaluate our method on the KITTI depth completion benchmark and show that depth completion can be significantly improved via the auxiliary supervision of image reconstruction. Our algorithm consistently outperforms depth-only methods and is also effective for indoor scenes like NYUv2.
Kaiyue Lu, Nick Barnes, Saeed Anwar, Liang Zheng 0001
CVPR3
2020 UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders
abstract
In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following a deterministic learning pipeline. Inspired by the saliency data labeling process, we propose probabilistic RGB-D saliency detection network via conditional variational autoencoders to model human annotation uncertainty and generate multiple saliency maps for each input image by sampling in the latent space. With the proposed saliency consensus process, we are able to generate an accurate saliency map based on these multiple predictions. Quantitative and qualitative evaluations on six challenging benchmark datasets against 18 competing algorithms demonstrate the effectiveness of our approach in learning the distribution of saliency maps, leading to a new state-of-the-art in RGB-D saliency detection.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Tong Zhang 0023, Nick Barnes
CVPR4
2020 Are Deep Neural Architectures Losing Information? Invertibility is Indispensable
Yang Liu 0249, Zhenyue Qin, Saeed Anwar, Sabrina B. Caldwell, Tom Gedeon
ICONIP (3)3
2020 OfGAN: Realistic Rendition of Synthetic Colonoscopy Videos
Jiabo Xu, Saeed Anwar, Nick Barnes, Florian Grimpen, Olivier Salvado, Stuart Anderson 0004, Mohammad Ali Armin
MICCAI (3)2
2020 Underwater scene prior inspired deep underwater image and video enhancement
Chongyi Li, Saeed Anwar, Fatih Porikli
Pattern Recognit.2
2020 Diving deeper into underwater image enhancement: A survey
Saeed Anwar, Chongyi Li
Signal Process. Image Commun.1
2019 Real Image Denoising With Feature Attention
abstract
Deep convolutional neural networks perform better on images containing spatially invariant noise (synthetic noise); however, its performance is limited on real-noisy photographs and requires multiple stage network modeling. To advance the practicability of the denoising algorithms, this paper proposes a novel single-stage blind real image denoising network (RIDNet) by employing a modular architecture. We use residual on the residual structure to ease the flow of low-frequency information and apply feature attention to exploit the channel dependencies. Furthermore, the evaluation in terms of quantitative metrics and visual quality on three synthetic and four real noisy datasets against 19 state-of-the-art algorithms demonstrate the superiority of our RIDNet.
Saeed Anwar, Nick Barnes
ICCV1
2019 Image Deblurring with a Class-Specific Prior
abstract
A fundamental problem in image deblurring is to recover reliably distinct spatial frequencies that have been suppressed by the blur kernel. To tackle this issue, existing image deblurring techniques often rely on generic image priors such as the sparsity of salient features including image gradients and edges. However, these priors only help recover part of the frequency spectrum, such as the frequencies near the high-end. To this end, we pose the following specific questions: (i) Does any image class information offer an advantage over existing generic priors for image quality restoration? (ii) If a class-specific prior exists, how should it be encoded into a deblurring framework to recover attenuated image frequencies? Throughout this work, we devise a class-specific prior based on the band-pass filter responses and incorporate it into a deblurring strategy. More specifically, we show that the subspace of band-pass filtered images and their intensity distributions serve as useful priors for recovering image frequencies that are difficult to recover by generic image priors. We demonstrate that our image deblurring framework, when equipped with the above priors, significantly outperforms many state-of-the-art methods using generic image priors or class-specific exemplars.
Saeed Anwar, Cong Phuoc Huynh, Fatih Porikli
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Combined Internal and External Category-Specific Image Denoising
Saeed Anwar, Cong Phuoc Huynh, Fatih Porikli
BMVC1
2017 Depth Estimation and Blur Removal from a Single Out-of-focus Image
Saeed Anwar, Zeeshan Hayder, Fatih Porikli
BMVC1
2017 Category-Specific Object Image Denoising
abstract
We present a novel image denoising algorithm that uses external, category specific image database. In contrast to existing noisy image restoration algorithms that search patches either from a generic database or noisy image itself, our method first selects clean images similar to the noisy image from a database that consists of images of the same class. Then, within the spatial locality of each noisy patch, it assembles a set of "support patches" from the selected images. These noisy-free support samples resemble the noisy patch and correspond principally to the identical part of the depicted object. In addition, we employ a content adaptive distribution model for each patch, where we derive the parameters of the distribution from the support patches. We formulate noise removal task as an optimization problem in the transform domain. Our objective function composed of a Gaussian fidelity term that imposes category specific information, and a low-rank term that encourages the similarity between the noisy and the support patches in a robust manner. The denoising process is driven by an iterative selection of support patches and optimization of the objective function. Our extensive experiments on five different object categories confirm the benefit of incorporating category-specific information to noise removal and demonstrate the superior performance of our method over the state-of-the-art alternatives.
Saeed Anwar, Fatih Porikli, Cong Phuoc Huynh
IEEE Trans. Image Process.1
2015 Class-Specific Image Deblurring
abstract
In image deblurring, a fundamental problem is that the blur kernel suppresses a number of spatial frequencies that are difficult to recover reliably. In this paper, we explore the potential of a class-specific image prior for recovering spatial frequencies attenuated by the blurring process. Specifically, we devise a prior based on the class-specific subspace of image intensity responses to band-pass filters. We learn that the aggregation of these subspaces across all frequency bands serves as a good class-specific prior for the restoration of frequencies that cannot be recovered with generic image priors. In an extensive validation, our method, equipped with the above prior, yields greater image quality than many state-of-the-art methods by up to 5 dB in terms of image PSNR, across various image categories including portraits, cars, cats, pedestrians and household objects.
Saeed Anwar, Cong Phuoc Huynh, Fatih Porikli
ICCV1