VLDB 2026 Research / reviewers in the wild / expert
Naveed Akhtar
dblp:147/2745
· DBLP profile ↗
83ranked-venue papers
16as first author
60since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 13 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 8 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 10 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | SFA-DiffNet: Spatial-frequency aware diffusion network for robust medical image segmentation
Tongtong Xie, Hongshan Yu, Yan Zheng 0003, Yong He 0012, Zhengeng Yang, Zechuan Li, Naveed Akhtar |
Expert Syst. Appl. | 7 |
| 2026 | GenMix: Effective data augmentation with generative diffusion model image editingabstractData augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board. Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar, Naveed Akhtar |
Expert Syst. Appl. | 5 |
| 2026 | Efficient Point Cloud Processing With High-Dimensional Positional Encoding and Non-Local MLPsabstractMulti-Layer Perceptron (MLP) models are the foundation of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength and limit the application of these models. In this article, we develop a two-stage abstraction and refinement (ABS-REF) view for modular feature extraction in point cloud processing. This view elucidates that whereas the early models focused on ABS stages, the more recent techniques devise sophisticated REF stages to attain performance advantages. Then, we propose a High-dimensional Positional Encoding (HPE) module to explicitly utilize intrinsic positional information, extending the "positional encoding" concept from Transformer literature. HPE can be readily deployed in MLP-based architectures and is compatible with transformer-based methods. Within our ABS-REF view, we rethink local aggregation in MLP-based methods and propose replacing time-consuming local MLP operations, which are used to capture local relationships among neighbors. Instead, we use non-local MLPs for efficient non-local information updates, combined with the proposed HPE for effective local information representation. We leverage our modules to develop HPENets, a suite of MLP networks that follow the ABS-REF paradigm, incorporating a scalable HPE-based REF stage. Extensive experiments on seven public datasets across four different tasks show that HPENets deliver a strong balance between efficiency and effectiveness. Notably, HPENet surpasses PointNeXt, a strong MLP-based counterpart, by 1.1% mAcc, 4.0% mIoU, 1.8% mIoU and 0.2% Cls. mIoU, with only 50.0%, 21.5%, 23.1%, 44.4% of FLOPs on ScanObjectNN, S3DIS, ScanNet, and ShapeNetPart, respectively. Yanmei Zou, Hongshan Yu, Yaonan Wang 0001, Zhengeng Yang, Xieyuanli Chen, Kailun Yang 0001, Naveed Akhtar |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Prompt-guided selective frequency network for real-world scene text image super-Resolutionabstract• We introduce PGSFNet for real-world scene text image super-resolution. • Adaptive Frequency Modulator is proposed to extract informative frequency components. • Text Information Enhancement module is designed to incorporate text priors. • We develop a Sobel loss to guide optimization towards sharper text details. • PGSFNet is shown to achieve superior performance on public text image datasets. Real-world scene text image super-resolution is challenging due to complex writing strokes, random text distribution, and diverse scene degradations. Existing text super-resolution methods focus on pure text images or fixed-size single-line text, which limits their practical utility. To address that, we propose a Prompt-Guided Selective Frequency super-resolution Network (PGSFNet). Our unique bicephalous neural model comprises a super-resolution branch and a prompt guidance branch. The latter specifically helps in leveraging text content-aware information priors. To that end, we propose a Text Information Enhancement module. To exploit selective frequency information present in the image, PGSFNet employs a proposed Adaptive Frequency Modulator fused with multi-attention structures. Considering the criticality of text edges in our task, we also propose a tailored text edge perception loss. Extensive experiments on the standard open real-world scene text image datasets demonstrate remarkable performance of our method, achieving up to 8.75% PNSR gain for × 2 and 2.28% SSIM gain for × 4 super-resolution on the Real-CE dataset. Our code will be made public at https://github.com/holastq/PGSFNet . Tianqi Shan, Hanlin Qin, Naveed Akhtar, Hossein Rahmani 0001, Ajmal Mian |
Pattern Recognit. | 4 |
| 2026 | Multi-Modal Shape Encoding for 3-D Object Detection
Zechuan Li, Hongshan Yu, Niu Zhang, Jinhao Qiao, Wei Sun 0028, Naveed Akhtar |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | LPL3D: LVLM-Driven Pseudo-Labeling for 3D Object DetectionabstractEffective 3D object detection requires large-scale annotated datasets, which are expensive and time-consuming to produce - especially in indoor environments containing dense object arrangements. To address this, we propose an Large Vision-Language Model (LVLM)-driven automatic high-quality pseudo-label generation technique for 3D object detection in single- and multi-view scenarios. We propose an Auto3DLabeler that introduces the first-ever text-to-3D Bounding Box transformation. Its pipeline employs a text-based detector, a segmenter and an LVLM to generate annotation estimates, which are further refined by our IoU-guided iterative Box Aggregator and layout-aware prompt Class Refiner modules. We also introduce a semantic-enhanced multi-modal fusion module that integrates image-level semantic information into point cloud representations for precise detections. Collectively, our contributions provide a remarkable boost to the 3D object detection state-of-the-art. Extensive experiments on SUN RGB-D and ScanNet datasets show our unsupervised detector variant outperforming existing semi-supervised detectors, and our semi-supervised variant achieving up to 28.2% absolute gain in challenging scenarios - all this while maintaining considerable compute advantage over existing label-efficient methods. Our code and models will be made public for the community. Our code and model will be made public after acceptance. Zechuan Li, Hongshan Yu, Yihao Ding, Shuai Yuan 0013, Naveed Akhtar |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Monocular Multi-Object 3D Visual Language TrackingabstractVisual Language Tracking (VLT) enables machines to perform tracking in real world through human-like language descriptions. However, existing VLT methods are limited to 2D spatial tracking or single-object 3D tracking and do not support multi-object 3D tracking within monocular video. This limitation arises because advancements in 3D multi-object tracking have predominantly relied on sensor-based data (e.g., point clouds, depth sensors) that lacks corresponding language descriptions. Moreover, natural language descriptions in existing VLT literature often suffer from redundancy, impeding the efficient and precise localization of multiple objects. We present the first technique to extend VLT to multi-object 3D tracking using monocular video. We introduce a comprehensive framework that includes (i) a Monocular Multi-object 3D Visual Language Tracking (MoMo-3DVLT) task, (ii) a large-scale dataset, MoMo-3DRoVLT, tailored for this task, and (iii) a custom neural model. Our dataset, generated with the aid of Large Language Models (LLMs) and manual verification, contains 8,216 video sequences annotated with both 2D and 3D bounding boxes, with each sequence accompanied by three freely generated, human-level textual descriptions. We propose MoMo-3DVLTracker, the first neural model specifically designed for MoMo-3DVLT. This model integrates a multimodal feature extractor, a visual language encoder-decoder, and modules for detection and tracking, setting a strong baseline for MoMo-3DVLT. Beyond existing paradigms, it introduces a task-specific structural coupling that integrates a differentiable linked-memory mechanism with depth-guided and language-conditioned reasoning for robust monocular 3D multi-object tracking. Experimental results demonstrate that our approach outperforms existing methods on the MoMo-3DRoVLT dataset. Our dataset and code are available at https://github.com/hongkai-wei/MoMo-3DVLT. Hongkai Wei, Haixiang Hu, Shijie Sun 0001, Mingtao Feng, Keyu Guo, Yongle Huang, Naveed Akhtar |
IEEE Trans. Image Process. | 10 |
| 2025 | Skip Mamba Diffusion for Monocular 3D Semantic Scene Completionabstract3D semantic scene completion is critical for multiple downstream tasks in autonomous systems. It estimates missing geometric and semantic information in the acquired scene data. Due to the challenging real-world conditions, this task usually demands complex models that process multi-modal data to achieve acceptable performance. We propose a unique neural model, leveraging advances from the state space and diffusion generative modeling to achieve remarkable 3D semantic scene completion performance with monocular image input. Our technique processes the data in the conditioned latent space of a variational autoencoder where diffusion modeling is carried out with an innovative state space technique. A key component of our neural network is the proposed Skimba (Skip Mamba) denoiser, which is adept at efficiently processing long-sequence data. The Skimba diffusion model is integral to our 3D scene completion network, incorporating a triple Mamba structure, dimensional decomposition residuals and varying dilations along three directions. We also adopt a variant of this network for the subsequent semantic segmentation stage of our method. Extensive evaluation on the standard SemanticKITTI and SSCBench-KITTI360 datasets show that our approach not only outperforms other monocular techniques by a large margin, it also achieves competitive performance against stereo methods. Li Liang 0006, Naveed Akhtar, Jordan Vice, Xiangrui Kong, Ajmal Mian |
AAAI | 2 |
| 2025 | Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept ControlabstractEthical issues around text-to-image (T2I) models demand a comprehensive control over the generative content. Existing techniques addressing these issues for responsible T2I models aim for the generated content to be fair and safe (non-violent/explicit). However, these methods remain bounded to handling the facets of responsibility concepts individually, while also lacking in interpretability. Moreover, they often require alteration to the original model, which compromises the model performance. In this work, we propose a unique technique to enable responsible T2I generation by simultaneously accounting for an extensive range of concepts for fair and safe content generation in a scalable manner. The key idea is to distill the target T2I pipeline with an external plug-and-play mechanism that learns an interpretable composite responsible space for the desired concepts, conditioned on the target T2I pipeline. We use knowledge distillation and concept whitening to enable this. At inference, the learned space is utilized to modulate the generative content. A typical T2I pipeline presents two plug-in points for our approach, namely; the text embedding space and the diffusion model latent space. We develop modules for both points and show the effectiveness of our approach with a range of strong results. Our code can be accessed at https://basim-azam.github.io/responsiblediffusion/ Basim Azam, Naveed Akhtar |
CVPR | 2 |
| 2025 | Beyond Human Perception: Understanding Multi-Object World from Monocular ViewabstractLanguage and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG Keyu Guo, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001, Naveed Akhtar, Ajmal Mian |
CVPR | 10 |
| 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object DetectorabstractWe propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is challenging. Addressing that, we introduce a unique 3D positional information embedded voxel optimization mechanism to fuse multi-view features. To prioritize neural field reconstruction in object regions, we also devise a double importance sampling scheme for the NeRF branch of our detector. We additionally propose an opacity optimization module for precise voxel opacity prediction by enforcing multi-view consistency constraints. Moreover, to further improve voxel density consistency across multiple perspectives, we incorporate ray distance as a weighting factor to minimize cumulative ray errors. Our unique modules synergetically form an end-to-end neural model that establishes new state-of-the-art in NeRF-based multi-view 3D detection, verified with extensive experiments on ScanNet and ARKITScenes. Code will be available at https://github.com/ZechuanLi/GO-N3RDet. Zechuan Li, Hongshan Yu, Yihao Ding, Jinhao Qiao, Basim Azam, Naveed Akhtar |
CVPR | 6 |
| 2025 | Mono3DVLT: Monocular-Video-Based 3D Visual Language TrackingabstractVisual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal with 3D tracking in the confines of monocular video. Unfortunately, advances in 3D tracking mainly rely on expensive sensor inputs, e.g., point clouds, depth measurements, radar. Absence of language counterpart for the outputs of these mildly democratized sensors in the literature also hinders VLT expansion to 3D tracking. Addressing that, we make the first attempt towards extending VLT to 3D tracking based on monocular video. We present a comprehensive framework, introducing (i) the Monocular-Video-based 3D Visual Language Tracking (Mono3DVLT) task, (ii) a large-scale dataset for the task, called Mono3DVLT-V2X, and (iii) a customized neural model for the task. Our dataset is carefully curated, leveraging a Large Langauge Model (LLM) followed by human verification, composing natural language descriptions for 79,158 video sequences aiming at single object tracking, providing 2D and 3D bounding box annotations. Our neural model, termed Mono3DVLT-MT, is the first targeted approach for the Mono3DVLT task. Comprising the pipeline of multi-modal feature extractor, visual-language encoder, tracking decoder and a tracking head, our model sets a strong baseline for the task on Mono3DVLT-V2X. Experimental results show that our method significantly outperforms existing techniques on the Mono3DVLT-V2X dataset. Our dataset and code are available in https://github.com/hongkai-wei/Mono3DVLT. Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Hongli Hu, Huansheng Song, Naveed Akhtar, Ajmal Mian |
CVPR | 10 |
| 2025 | DDB: Diffusion Driven Balancing to Address Spurious Correlations
Aryan Yazdan Parast, Basim Azam, Naveed Akhtar |
ICCV | 3 |
| 2025 | Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor ScenesabstractFusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and RGB input has started gaining traction. However, effective integration of these modalities for precise object detection tasks still remains a largely open problem. To address that, we propose a MultiStream Detection (MuStD) network, which meticulously extracts task-relevant information from both data modalities. The network follows a three-stream structure. Its LiDAR-PillarNet stream extracts sparse 2D pillar features from the LiDAR input while the LiDAR-Height Compression stream computes Bird’s-Eye View features. An additional 3D Multimodal stream combines RGB and LiDAR features using UV mapping and polar coordinate indexing. Eventually, the features containing comprehensive spatial, textural, and geometric information are carefully fused and fed to a detection head for 3D object detection. We evaluate our method on the challenging KITTI Object Detection Benchmark, with results available on the official evaluation server.1. Our approach achieves strong performance, with an average precision (AP) of 85.39% in 3D detection, 91.34% in Bird’s Eye View (BEV) detection, and 96.39% in 2D detection. These results match or surpass existing state-of-the-art methods. In the difficult "Hard" category, our method attains 80.78% AP in 3D detection and 94.04% AP in 2D detection, highlighting its robustness in challenging scenarios. Furthermore, our method runs at 67 ms, demonstrating efficiency and real-time capability. Our code will be released through the MuStD GitHub repository at https://github.com/IbrahimUWA/MuStD. Muhammad Ibrahim 0001, Naveed Akhtar, Haitian Wang 0002, Saeed Anwar, Ajmal Mian |
IROS | 2 |
| 2025 | Attention-Guided Vector Quantized Variational Autoencoder for Brain Tumor Segmentation
Ajmal Mian, Naveed Akhtar, Ghulam M. Hassan |
MICCAI (1) | 3 |
| 2025 | CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene GenerationabstractOutdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large‑scale benchmark for generating 3D outdoor semantic scenes from abstract freehand sketches and pseudo‑labeled annotations of satellite images. SketchSem3D includes two subsets, Sketch-based SemanticKITTI and Sketch-based KITTI-360 (containing LiDAR voxels along with their corresponding sketches and annotated satellite images), to enable standardized, rigorous, and diverse evaluations. We also propose Cylinder Mamba Diffusion (CymbaDiff) that significantly enhances spatial coherence in outdoor 3D scene generation. CymbaDiff imposes structured spatial ordering, explicitly captures cylindrical continuity and vertical hierarchy, and preserves both physical neighborhood relationships and global context within the generated scenes. Extensive experiments on SketchSem3D demonstrate that CymbaDiff achieves superior semantic consistency, spatial realism, and cross-dataset generalization. The code and dataset will be available at here. Li Liang 0006, Bo Miao, Naveed Akhtar, Jordan Vice, Ajmal Mian |
NeurIPS | 4 |
| 2025 | HDTCNet: A hybrid-dimensional convolutional network for multivariate time series classification
Yongli Gu, Hanlin Qin, Naveed Akhtar, Shuai Yuan 0013, Honghao Fu, Shuowen Yang, Ajmal Mian |
Pattern Recognit. | 4 |
| 2025 | Quantifying Bias in Text-to-Image Generative ModelsabstractBias in text-to-image (T2I) generation can propagate unfair social representations and may be exploited to push ulterior agendas. These biases raise concerns on the dependability and fairness of models that have become widely popular and readily available for public consumption. Existing works in T2I bias analysis typically focus on social biases. We look beyond that and instead propose an evaluation methodology to quantify general bias in T2I generative models without any preconceived notion. We introduce a suite of three metrics; namely, distribution bias, Jaccard hallucination and generative miss-rate, to extensively appraise general model bias. To validate the efficacy of these metrics, we also introduce a backdoor-inspired strategy, which provides a convenient handle over the extent of bias in a model for controlled analysis. We assess T2I models implementing six widely used pipelines in this domain. Our extensive analysis covers both general and task-oriented scenarios, employing over 105 K generated images. For prior art comparison, it also encompasses social bias analysis. Moreover, we also extend our technique to analyze bias in seven popular captioned image datasets. Our experiments establish that our approach is objective, domain-agnostic and it consistently measures different forms of T2I model biases. To further research efforts into T2I model biases, we have developed an open-source web application and practical implementation of this work, which is available onhttps://huggingface.co/spaces/JVice/try-before-you-biasHuggingFace. All relevant code is also publicly available onhttps://github.com/JJ-Vice/TryBeforeYouBiasGitHub. Jordan Vice, Naveed Akhtar, Richard I. Hartley, Ajmal Mian |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | ASCNet: Asymmetric Sampling Correction Network for Infrared Image DestripingabstractIn a real-world infrared (IR) imaging system, effectively learning a consistent stripe noise removal model is essential. Most existing destriping methods cannot precisely reconstruct images due to cross-level semantic gaps and insufficient characterization of the global column features. To tackle this problem, we propose a novel IR image destriping method, called asymmetric sampling correction network (ASCNet), that can effectively capture global column relationships and embed them into a U-shaped framework, providing comprehensive discriminative representation and seamless semantic connectivity. Our ASCNet consists of three core elements: residual Haar discrete wavelet transform (RHDWT), pixel shuffle (PS), and column nonuniformity correction module (CNCM). Specifically, RHDWT is a novel downsampler that employs double-branch modeling to effectively integrate stripe-directional prior knowledge and data-driven semantic interaction to enrich the feature representation. Observing the semantic patterns crosstalk of stripe noise, PS is introduced as an upsampler to prevent excessive a priori decoding and performing semantic-bias-free image reconstruction. After each sampling, CNCM captures the column relationships in long-range dependencies. By incorporating column, spatial, and self-dependence information, CNCM well establishes a global context to distinguish stripes from the scene’s vertical structures. Extensive experiments on synthetic data, real data, and IR small target detection (IRSTD) tasks demonstrate that the proposed method outperforms state-of-the-art single-image destriping methods both visually and quantitatively. The code is available athttps://github.com/xdFai/ASCNet. Shuai Yuan 0013, Hanlin Qin, Shiqi Yang 0001, Shuowen Yang, Naveed Akhtar, Huixin Zhou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | RSVMamba for Tree Species Classification Using UAV RGB Remote Sensing ImagesabstractEffective forest tree species (TS) classification is critical for various application domains such as forest management, biodiversity conservation, and ecological research. However, existing studies on TS classification predominantly rely on high-cost and processing-intensive hyperspectral data, which limits practical applications on large scales. In this work, we focus on investigating the potential of cost-effective unmanned aerial vehicle (UAV) RGB images for TS classification in heterogeneous forests and propose a method that fully leverages the rich spatial, semantic, and visible spectral information of UAV RGB images. We propose an RSVMamba model, which incorporates improved visual state-space (VSS) blocks and an AutoDownsampling module to enhance accuracy and stability while paying particular attention to small objects in sparse spatial locations. The model achieves linear computational complexity while retaining the global receptive field, making it particularly suitable for processing high spatial-resolution images. Additionally, we collected UAV RGB images covering$40~\text {km}^{2}$of subtropical forest in southern China. A meticulous evaluation of this data shows that our method achieves an overall accuracy (OA) of 84.28% for eight TS, dead trees, and other broadleaves. We verify the superiority of our method through a series of comparative experiments on the collected and benchmark datasets. Our results affirm the usefulness of single-temporal UAV RGB images for TS classification in heterogeneous forest environments. Furthermore, the proposed method bridges the gap between data accessibility and precision in TS classification, broadening the boundaries of single-temporal UAV RGB images for practical forestry applications and providing a more cost-effective and time-flexible solution for this problem. Juntao Gu, Basim Azam, Moule Lin, Chao Li 0066, Weipeng Jing 0001, Naveed Akhtar |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | NPC-SPU: Nonlinear Phase Coding-Based Stereo Phase Unwrapping for Efficient 3D Measurementabstract3D imaging based on phase-shifting structured light is widely used in industrial measurement due to its non-contact nature. However, it typically requires a large number of additional images (multi-frequency heterodyne (M-FH) method) or introduces intensity features that compromise accuracy (space domain modulation phase-shifting (SDM-PS) method) for phase unwrapping, and it remains sensitive to motion. To overcome these issues, this article proposes a nonlinear phase coding-based stereo phase unwrapping (NPC-SPU) method that requires no additional patterns while maintaining measurement accuracy. In the encoding stage, a novel nonlinear distortion feature is introduced, while the signal-to-noise ratio of the phase codeword is preserved. In the decoding stage, a local phase unwrapping method that does not require additional auxiliary information is first proposed, closely associating the distortion information in the local wrapped phase. Then, a pre-calibrated stereo constraint system is used to filter potential matching phases, significantly reducing phase ambiguity and computational costs. Finally, to avoid the time-consuming and complex intensity kernel matching used in traditional methods, we propose a local phase correlation matching (LPCM) technique that enables lightweight and robust phase unwrapping. Experimental results demonstrate that this algorithm significantly enhances 3D reconstruction performance in scenarios with large depth, large disparity, complex colored structures, and dynamic scenes. Specifically, in dynamic environments (20mm/s), the proposed method achieves a lower measurement error rate (0.7829% vs. 6.4962%) with only 3 patterns, compared to the traditional three-frequency heterodyne (T-FH) method (using 9 patterns). Additionally, its measurement accuracy outperforms the advanced SDM-PS method, which also uses 3 patterns (0.1102 mm vs. 0.3232 mm). Ruiming Yu, Hongshan Yu, Wei Sun 0028, Yaonan Wang 0001, Naveed Akhtar, Kemao Qian |
IEEE Trans. Image Process. | 5 |
| 2025 | A Comprehensive Overview of Large Language ModelsabstractLarge Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multimodal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to provide not only a systematic survey but also a quick, comprehensive reference for the researchers and practitioners to draw insights from extensive, informative summaries of the existing works to advance the LLM research. Humza Naveed, Asad Ullah Khan, Shi Qiu 0001, Saeed Anwar, Muhammad Usman 0010, Naveed Akhtar, Nick Barnes, Ajmal Mian |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2024 | Improved MLP Point Cloud Processing with High-Dimensional Positional EncodingabstractMulti-Layer Perceptron (MLP) models are the bedrock of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength. We first develop an “abstraction and refinement” (ABS-REF) view for the neural modeling of point clouds. This view elucidates that whereas the early models focused on the ABS stage, the more recent techniques devise sophisticated REF stages to attain performance advantage in point cloud processing. We then borrow the concept of “positional encoding” from transformer literature, and propose a High-dimensional Positional Encoding (HPE) module, which can be readily deployed to MLP based architectures. We leverage our module to develop a suite of HPENet, which are MLP networks that follow ABS-REF paradigm, albeit with a sophisticated HPE based REF stage. The developed technique is extensively evaluated for 3D object classification, object part segmentation, semantic segmentation and object detection. We establish new state-of-the-art results of 87.6 mAcc on ScanObjectNN for object classification, and 85.5 class mIoU on ShapeNetPart for object part segmentation, and 72.7 and 78.7 mIoU on Area-5 and 6-fold experiments with S3DIS for semantic segmentation. The source code for this work is available at https://github.com/zouyanmei/HPENet. Yanmei Zou, Hongshan Yu, Zhengeng Yang, Zechuan Li, Naveed Akhtar |
AAAI | 5 |
| 2024 | Regulating Model Reliance on Non-robust Features by Smoothing Input Marginal Density
Peiyu Yang, Naveed Akhtar, Mubarak Shah, Ajmal Mian |
ECCV (57) | 2 |
| 2024 | A Statistical Image Realism Score For Deepfake DetectionabstractRecent advances in generative visual content have led to a quantum leap in the quality of artificially generated Deepfake content. Especially, diffusion models are causing growing concerns among communities due to their ever-increasing realism. However, quantifying the realism of generated content is still challenging. Existing evaluation metrics, such as Inception Score and Fréchet inception distance, fall short on benchmarking diffusion models due to the versatility of the generated images. To address this, we propose the Image Realism Score (IRS) evaluation metric, computed from five statistical measures of a given image. This non-learning-based metric not only efficiently quantifies the realism of generated images, but it is also a viable tool for detecting if an image is real or fake. We experimentally establish the model- and data-agnostic nature of the proposed IRS by successfully detecting fake images generated by Stable Diffusion Model (SDM), Dalle2, Dalle3, Deepfloyd, Kandinsky, Midjourney and BigGAN. Yunzhuo Chen, Naveed Akhtar, Nur Al Hasan Haldar, Jordan Vice, Ajmal Mian |
ICIP | 2 |
| 2024 | EFRNet-VL: An end-to-end feature refinement network for monocular visual localization in dynamic environments
Jingwen Wang 0009, Hongshan Yu, Xuefei Lin, Zechuan Li, Wei Sun 0028, Naveed Akhtar |
Expert Syst. Appl. | 6 |
| 2024 | LPL-VIO: monocular visual-inertial odometry with deep learning-based point and line features
Changxiang Liu, Qinhan Yang, Hongshan Yu, Qiang Fu 0013, Naveed Akhtar |
Neural Comput. Appl. | 5 |
| 2024 | Multimodal fusion for audio-image and video action recognitionabstractAbstract Multimodal Human Action Recognition (MHAR) is an important research topic in computer vision and event recognition fields. In this work, we address the problem of MHAR by developing a novel audio-image and video fusion-based deep learning framework that we call Multimodal Audio-Image and Video Action Recognizer (MAiVAR). We extract temporal information using image representations of audio signals and spatial information from video modality with the help of Convolutional Neutral Networks (CNN)-based feature extractors and fuse these features to recognize respective action classes. We apply a high-level weights assignment algorithm for improving audio-visual interaction and convergence. This proposed fusion-based framework utilizes the influence of audio and video feature maps and uses them to classify an action. Compared with state-of-the-art audio-visual MHAR techniques, the proposed approach features a simpler yet more accurate and more generalizable architecture, one that performs better with different audio-image representations. The system achieves an accuracy 87.9% and 79.0% on UCF51 and Kinetics Sounds datasets, respectively. All code and models for this paper will be available at https://tinyurl.com/4ps2ux6n . Muhammad Bilal Shaikh, Douglas Chai, Syed M. S. Islam, Naveed Akhtar |
Neural Comput. Appl. | 4 |
| 2024 | SCTransNet: Spatial-Channel Cross Transformer Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) has recently benefitted greatly from U-shaped neural models. However, largely overlooking effective global information modeling, existing techniques struggle when the target has high similarities with the background. We present aSpatial-channelCrossTransformerNetwork (SCTransNet) that leverages spatial-channel cross transformer blocks (SCTBs) on top of long-range skip connections to address the aforementioned challenge. In the proposed SCTBs, the outputs of all encoders are interacted with cross transformer to generate mixed features, which are redistributed to all decoders to effectively reinforce semantic differences between the target and clutter at full levels. Specifically, SCTB contains the following two key elements: (a) spatial-embedded single-head channel-cross attention (SSCA) for exchanging local spatial features and full-level global channel information to eliminate ambiguity among the encoders and facilitate high-level semantic associations of the images, and (b) a complementary feed-forward network (CFN) for enhancing the feature discriminability via a multi-scale strategy and cross-spatial-channel information interaction to promote beneficial information transfer. Our SCTransNet effectively encodes the semantic differences between targets and backgrounds to boost its internal representation for detecting small infrared targets accurately. Extensive experiments on three public datasets, NUDT-SIRST, NUAA-SIRST, and IRSTD-1K, demonstrate that the proposed SCTransNet outperforms existing IRSTD methods. Our code will be made public at https://github.com/xdFai/SCTransNet. Shuai Yuan 0013, Hanlin Qin, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | IRSTDID-800: A Benchmark Analysis of Infrared Small Target Detection-Oriented Image DestripingabstractDeep learning-based single-image infrared (IR) destriping has made significant advances. However, these methods are typically evaluated using “synthetic” images with specific stripe noise, making it unclear how well they handle “real” IR images. In fact, a clear and fair benchmarking of the existing destriping methods on real images, especially for the downstream IR small target detection (IRSTD) task, is currently an open gap. To tackle this problem, we introduce a novel benchmark, called IRSTD-oriented image destriping (IRSTDID-800), which thoroughly showcases the real distribution of IR small targets under stripe noise perturbation for the first time. Concretely, it consists of two subsets. (1) IRSTDID-SKY: composed of 500 real-world images afflicted with stripe noise, including unmanned aerial vehicles (UAVs) of various shapes, sizes, and contrasts. Moreover, these images are annotated with precise pixel levels for objective evaluation of IRSTD. (2) IRSTDID-GND: comprising 300 real-world images featuring common daily life objects, providing a richer scene under stripe noise. Based on the proposed IRSTDID-800, we comprehensively assess the performance of ten state-of-the-art (SOTA) destriping methods across eleven metrics, including full-reference, no-reference, and task-driven metrics with six advanced IRSTD methods. Furthermore, inspired by the correlation between image quality assessment and IRSTD, we proposed a task-oriented destriping optimization strategy. A loss function is introduced for IR image destriping, leveraging the structural properties of noise as a penalty term to strengthen image destriping and IRSTD. Overall, our analysis reveals interesting observations to guide future research in destriping and IRSTD tasks. Our dataset is available athttps://github.com/xdFai/IRSTDID-800. Shuai Yuan 0013, Hanlin Qin, Naveed Akhtar, Shiqi Yang 0001, Shuowen Yang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | BAGM: A Backdoor Attack for Manipulating Text-to-Image Generative ModelsabstractThe rise in popularity of text-to-image generative artificial intelligence (AI) has attracted widespread public interest. We demonstrate that this technology can be attacked to generate content that subtly manipulates its users. We propose a Backdoor Attack on text-to-image Generative Models (BAGM), which upon triggering, infuses the generated images with manipulative details that are naturally blended in the content. Our attack is the first to target three popular text-to-image generative models across three stages of the generative process by modifying the behaviour of the embedded tokenizer, the language model or the image generative model. Based on the penetration level, BAGM takes the form of a suite of attacks that are referred to assurface,shallowanddeepattacks in this article. Given the existing gap within this domain, we also contribute a comprehensive set of quantitative metrics designed specifically for assessing the effectiveness of backdoor attacks on text-to-image models. The efficacy of BAGM is established by attacking state-of-the-art generative models, using a marketing scenario as the target domain. To that end, we contribute a dataset of branded product images. Our embedded backdoors increase the bias towards the target outputs by more than five times the usual, without compromising the model robustness or the generated content utility. By exposing generative AI’s vulnerabilities, we encourage researchers to tackle these challenges and practitioners to exercise caution when using pre-trained models. Relevant code and input prompts can be found at https://github.com/JJ-Vice/BAGM, and the dataset is available at: https://ieee-dataport.org/documents/marketable-foods-mf-dataset. Jordan Vice, Naveed Akhtar, Richard I. Hartley, Ajmal Mian |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Enhanced Deep Predictive Modeling of Wastewater Plants With Limited DataabstractDeep learning is being widely utilized in industrial process monitoring, control, and optimization. However, in the wastewater industry, its applications are still underexplored. This is because deep learning requires a large amount of labeled training data to induce effective predictive models. Owing to the high cost of sensors and frequency and delay in sampling and laboratory analytics, wastewater treatment process data can be sparse with varying frequencies. One option to address training data limitations is to use transfer learning. However, owing to the large covariate shift between the commonly adopted source domains for transfer learning and the target domain of wastewater processes, this approach leads to unacceptable performance. We address this issue by proposing a novel synthetic data generation method for deep predictive modeling of wastewater plants. Employing a Markov process that utilizes random walk, our technique enables the generation of abundant annotated data for our target domain. The method preserves the temporal dynamics and distribution of the original data, thereby closely mimicking the potential original samples of the domain. We extensively evaluate our method over two different high-rate-algae-based treatment datasets, demonstrating considerable performance gains over existing transfer learning. Our proposed algorithm can assist plant operators to deploy responsive supportive models with limited data. Maira Alvi, Tim French 0002, Rachel Cardell-Oliver, Damien J. Batstone, Naveed Akhtar |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Yolo-3DMM for Simultaneous Multiple Object Detection and Tracking in Traffic ScenariosabstractVideo-based multiple object tracking (MOT) is a fundamental task in intelligent transportation with applications ranging from automated traffic surveillance to autonomous driving. MOT methods commonly follow a tracking-by-detection paradigm, tracking objects by associating their detections across video frames. However, insofar, these methods have not used the entire vehicle trajectory motion characteristics to perform tracking, which converts the vehicle localization problem into a motion parameter estimation problem. Moreover, MOT methods mainly rely on off-the-shelf detectors. An independently trained detector is sub-optimal for the tracking-by-detection paradigm and adversely affects the overall system performance. In this article, we address these issues by proposing a novel MOT method for moving vehicles in traffic scenarios. Our tracker treats the vehicle tracks as unified 3D spatio-temporal trajectory instances and leverages the power of deep learning to extract vehicle motion from the 3D instances. We propose a new simultaneous detection and tracking network, called YOLO-3D Motion Model Network (Yolo-3DMM) that employs spatio-temporal features of traffic videos for simultaneous vehicle detection and tracking in an end-to-end manner. We adopt a variety of different vehicle tracking datasets to evaluate our method. Moreover, we also propose a tunnel MOT dataset from real highway tunnel surveillance in Guangdong, China to expand the experimental scenarios. To establish the efficacy of our method, we evaluate it on 100 different roadside traffic scenarios. Our method shows excellent performance on UA-DETRAC and Omni-MOT datasets. It achieves a PR-MOTA score of 29.40% on UA-DETRAC and gets a 69.7% MOTA score on the Omni-MOT dataset. Lichen Liu, Huansheng Song, Shijie Sun 0001, Xian-Feng Han, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Co-Engaged Location Group Search in Location-Based Social NetworksabstractSearching for well-connected user communities in a Location-based Social Network (LBSN) has been extensively investigated. However, very few studies focus on finding a group of locations in an LBSN which are significantly engaged with socially cohesive user groups. In this work, we investigate the problem ofCo-engagedLocation groupSearch (CLS) from LBSNs where the selected locations are visited frequently by the members of the socially cohesive user groups, and the locations are reachable within a given distance threshold. To the best of our knowledge, this is the first work to search for socially co-engaged location groups in LBSNs. We devise a score function to measure the co-engagement of the location groups by combining social connectivity of the cohesive user groups and check-in density of the users to the selected locations. To solve theCLSproblem, we propose aFilter-and-Verifyalgorithm that effectively filters out ineligible locations, and their corresponding check-in users. Further, we derive a lower bound on the number of check-ins to prune the insignificant locations and develop a novel greedy forward expansion algorithm (GFA). To accelerate the computation ofCLS, we propose a ranking function and devise an incremental algorithm,GIA, that can filter the unqualified location groups. We establish the effectiveness of our solutions by conducting extensive experiments on three real-world datasets. Nur Al Hasan Haldar, Jianxin Li 0001, Naveed Akhtar, Yan Jia 0001, Ajmal Mian |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Mesh Convolution With Continuous Filters for 3-D Surface ParsingabstractGeometric feature learning for 3-D surfaces is critical for many applications in computer graphics and 3-D vision. However, deep learning currently lags in hierarchical modeling of 3-D surfaces due to the lack of required operations and/or their efficient implementations. In this article, we propose a series of modular operations for effective geometric feature learning from 3-D triangle meshes. These operations include novel mesh convolutions, efficient mesh decimation, and associated mesh (un)poolings. Our mesh convolutions exploit spherical harmonics as orthonormal bases to create continuous convolutional filters. The mesh decimation module is graphics processing unit (GPU)-accelerated and able to process batched meshes on-the-fly, while the (un)pooling operations compute features for upsampled/downsampled meshes. We provide an open-source implementation of these operations, collectively termed Picasso. Picasso supports heterogeneous mesh batching and processing. Leveraging its modular operations, we further contribute a novel hierarchical neural network for perceptual parsing of 3-D surfaces, named PicassoNet++. It achieves highly competitive performance for shape analysis and scene segmentation on prominent 3-D benchmarks. The code, data, and trained models are available at https://github.com/EnyaHermite/Picasso. Huan Lei, Naveed Akhtar, Mubarak Shah, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | From CNNs to Transformers in Multimodal Human Action Recognition: A SurveyabstractDue to its widespread applications, human action recognition is one of the most widely studied research problems in Computer Vision. Recent studies have shown that addressing it using multimodal data leads to superior performance as compared to relying on a single data modality. During the adoption of deep learning for visual modelling in the past decade, action recognition approaches have mainly relied on Convolutional Neural Networks (CNNs). However, the recent rise of Transformers in visual modelling is now also causing a paradigm shift for the action recognition task. This survey captures this transition while focusing on Multimodal Human Action Recognition (MHAR). Unique to the induction of multimodal computational models is the process of ‘fusing’ the features of the individual data modalities. Hence, we specifically focus on the fusion design aspects of the MHAR approaches. We analyze the classic and emerging techniques in this regard, while also highlighting the popular trends in the adaption of CNN and Transformer building blocks for the overall problem. In particular, we emphasize on recent design choices that have led to more efficient MHAR models. Unlike existing reviews, which discuss Human Action Recognition from a broad perspective, this survey is specifically aimed at pushing the boundaries of MHAR research by identifying promising architectural and fusion design choices to train practicable models. We also provide an outlook of the multimodal datasets from their scale and evaluation viewpoint. Finally, building on the reviewed literature, we discuss the challenges and future avenues for MHAR. Muhammad Bilal Shaikh, Douglas Chai, Syed M. S. Islam, Naveed Akhtar |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Contrastive Self-Supervised Learning Leads to Higher Adversarial SusceptibilityabstractContrastive self-supervised learning (CSL) has managed to match or surpass the performance of supervised learning in image and video classification. However, it is still largely unknown if the nature of the representations induced by the two learning paradigms is similar. We investigate this under the lens of adversarial robustness. Our analysis of the problem reveals that CSL has intrinsically higher sensitivity to perturbations over supervised learning. We identify the uniform distribution of data representation over a unit hypersphere in the CSL representation space as the key contributor to this phenomenon. We establish that this is a result of the presence of false negative pairs in the training process, which increases model sensitivity to input perturbations. Our finding is supported by extensive experiments for image and video classification using adversarial perturbations and other input corruptions. We devise a strategy to detect and remove false negative pairs that is simple, yet effective in improving model robustness with CSL training. We close up to 68% of the robustness gap between CSL and its supervised counterpart. Finally, we contribute to adversarial learning by incorporating our method in CSL. We demonstrate an average gain of about 5% over two different state-of-the-art methods in this domain. Rohit Gupta 0012, Naveed Akhtar, Ajmal Mian, Mubarak Shah |
AAAI | 2 |
| 2023 | Rethinking Interpretation: Input-Agnostic Saliency Mapping of Deep Visual ClassifiersabstractSaliency methods provide post-hoc model interpretation by attributing input features to the model outputs. Current methods mainly achieve this using a single input sample, thereby failing to answer input-independent inquiries about the model. We also show that input-specific saliency mapping is intrinsically susceptible to misleading feature attribution. Current attempts to use `general' input features for model interpretation assume access to a dataset containing those features, which biases the interpretation. Addressing the gap, we introduce a new perspective of input-agnostic saliency mapping that computationally estimates the high-level features attributed by the model to its outputs. These features are geometrically correlated, and are computed by accumulating model's gradient information with respect to an unrestricted data distribution. To compute these features, we nudge independent data points over the model loss surface towards the local minima associated by a human-understandable concept, e.g., class label for classifiers. With a systematic projection, scaling and refinement process, this information is transformed into an interpretable visualization without compromising its model-fidelity. The visualization serves as a stand-alone qualitative interpretation. With an extensive evaluation, we not only demonstrate successful visualizations for a variety of concepts for large-scale models, but also showcase an interesting utility of this new form of saliency mapping by identifying backdoor signatures in compromised classifiers. Naveed Akhtar, Mohammad A. A. K. Jalwana |
AAAI | 1 |
| 2023 | Local Path Integration for AttributionabstractPath attribution methods are a popular tool to interpret a visual model's prediction on an input. They integrate model gradients for the input features over a path defined between the input and a reference, thereby satisfying certain desirable theoretical properties. However, their reliability hinges on the choice of the reference. Moreover, they do not exhibit weak dependence on the input, which leads to counter-intuitive feature attribution mapping. We show that path-based attribution can account for the weak dependence property by choosing the reference from the local distribution of the input. We devise a method to identify the local input distribution and propose a technique to stochastically integrate the model gradients over the paths defined by the references sampled from that distribution. Our local path integration (LPI) method is found to consistently outperform existing path attribution techniques when evaluated on deep visual models. Contributing to the ongoing search of reliable evaluation metrics for the interpretation methods, we also introduce DiffID metric that uses the relative difference between insertion and deletion games to alleviate the distribution shift problem faced by existing metrics. Our code is available at https://github.com/ypeiyu/LPI. Peiyu Yang, Naveed Akhtar, Zeyi Wen, Ajmal Mian |
AAAI | 2 |
| 2023 | AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformersabstract3D object detection techniques commonly follow a pipeline that aggregates predicted object central point features to compute candidate points. However, these candidate points contain only positional information, largely ignoring the object-level shape information. This eventually leads to sub-optimal 3D object detection. In this work, we propose AShapeFormer, a semantics-guided object-level shape encoding module for 3D object detection. This is a plug-n-play module that leverages multi-head attention to encode object shape information. We also propose shape tokens and object-scene positional encoding to ensure that the shape information is fully exploited. Moreover, we introduce a semantic guidance sub-module to sample more foreground points and suppress the influence of background points for a better object shape perception. We demonstrate a straightforward enhancement of multiple existing methods with our AShapeFormer. Through extensive experiments on the popular SUN RGB-D and ScanNetV2 dataset, we show that our enhanced models are able to outperform the baselines by a considerable absolute margin of up to 8.1%. Code will be available at https://github.com/ZechuanLi/AShapeFormer Zechuan Li, Hongshan Yu, Zhengeng Yang, Tom Tongjia Chen, Naveed Akhtar |
CVPR | 5 |
| 2023 | Re-calibrating Feature Attributions for Model Interpretation
Peiyu Yang, Naveed Akhtar, Zeyi Wen, Mubarak Shah, Ajmal Mian |
ICLR | 2 |
| 2023 | Towards credible visual model interpretation with path attributionabstractWith its inspirational roots in game-theory, path attribution framework stands out among the post-hoc model interpretation techniques due to its axiomatic nature. However, recent developments show that despite being axiomatic, path attribution methods can compute counter-intuitive feature attributions. Not only that, for deep visual models, the methods may also not conform to the original game-theoretic intuitions that are the basis of their axiomatic nature. To address these issues, we perform a systematic investigation of the path attribution framework. We first pinpoint the conditions in which the counter-intuitive attributions of deep visual models can be avoided under this framework. Then, we identify a mechanism of integrating the attributions over the paths such that they computationally conform to the original insights of game-theory. These insights are eventually combined into a method, which provides intuitive and reliable feature attributions. We also establish the findings empirically by evaluating the method on multiple datasets, models and evaluation metrics. Extensive experiments show a consistent quantitative and qualitative gain in the results over the baselines. Naveed Akhtar, Mohammad A. A. K. Jalwana |
ICML | 1 |
| 2023 | Slice Transformer and Self-supervised Learning for 6DoF Localization in 3D Point Cloud MapsabstractPrecise localization is critical for autonomous vehicles. We present a self-supervised learning method that employs transformers for the first time for the task of outdoor localization using LiDAR data. We propose a pre-text task that reorganizes the slices of a 360° LiDAR scan to leverage its axial properties. Our model, called Slice Transformer, employs multi-head attention while systematically processing the slices. To the best of our knowledge, this is the first instance of leveraging multi-head attention for outdoor point clouds. We additionally introduce the Perth-Wadataset, which provides a large-scale LiDAR map of Perth city in Western Australia, covering ~4km2area. Localization annotations are provided for Perth - Wa.The proposed localization method is thoroughly evaluated on Perth-WA and Appollo-SouthBay datasets. We also establish the efficacy of our self-supervised learning approach for the common downstream task of object classification using ModelNet40 and ScanNN datasets. The code and Perth-WA data will be publicly released. Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Michael J. Wise, Ajmal Mian |
ICRA | 2 |
| 2023 | UnLoc: A Universal Localization Method for Autonomous Vehicles using LiDAR, Radar and/or Camera InputabstractLocalization is a fundamental task in robotics for autonomous navigation. Existing localization methods rely on a single input data modality or train several computational models to process different modalities. This leads to stringent computational requirements and sub-optimal results that fail to capitalize on the complementary information in other data streams. This paper proposes UnLoc, a novel unified neural modeling approach for localization with multi-sensor input in all weather conditions. Our multi-stream network can handle LiDAR, Camera and RADAR inputs for localization on demand, i.e., it can work with one or more input sensors, making it robust to sensor failure. UnLoc uses 3D sparse convolutions and cylindrical partitioning of the space to process LiDAR frames and implements ResNet blocks with a slot attention-based feature filtering module for the Radar and image modalities. We introduce a unique learnable modality encoding scheme to distinguish between the input sensor data. Our method is extensively evaluated on Oxford Radar RobotCar, ApolloSouthBay and Perth-WA datasets. The results ascertain the efficacy of our technique. The dataset, results, and codes are available at https://github.com/IbrahimUWA/UnLoc Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Ajmal Mian |
IROS | 2 |
| 2023 | Language Model Agnostic Gray-Box Adversarial Attack on Image CaptioningabstractAdversarial susceptibility of neural image captioning is still under-explored due to the complex multi-model nature of the task. We introduce a GAN-based adversarial attack to effectively fool encoder-decoder based image captioning frameworks. Unique to our attack is the systematic disruption of the internal representation of an image at the encoder stage which allows control over the captions generated at the decoder stage. We cause the desired disruption with an input perturbation that promotes similarity between the features of the input image with a target image of our choice. The target image provides a convenient handle to control the incorrect captions in our method. We do not assume any knowledge of the decoder module, which makes our attack ‘gray-box’. Moreover, our attack remains agnostic to the type of decoder module, thereby proving effective for RNNs as well as Transformers as the language models. This makes our attack highly pragmatic. Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Mubarak Shah, Ajmal Mian |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | SAT3D: Slot Attention Transformer for 3D Point Cloud Semantic SegmentationabstractSemantic segmentation of 3D point cloud is a key task in numerous intelligent transportation system applications, e.g., self-driving vehicles, traffic monitoring. Due to the sparsity and varying density of points in the outdoor point clouds, it becomes particularly challenging to extract object-centric features from data. This leads to poor semantic segmentation, especially for the rare object classes. To address that, we introduce the first-ever Slot Attention Transformer based technique to effectively model object-centric features in point cloud data. Our method uses cylindrical splits of space for voxelization and computes channel-wise positional embeddings before repetitively encoding the point cloud with slot attentions. Our second major contribution is a Large-Scale Outdoor Point Cloud dataset (SWAN), collected in a dense urban environment, driving 150km distance. It provides 16 billion points in more than 200K frames. The dataset also provides annotations for 10K frames for 24 classes. We also contribute a data augmentation scheme to handle rare object classes in real-world point clouds. Besides benchmarking popular existing methods on SWAN for the first time, we thoroughly evaluate our technique on the existing large-scale datasets, Semantic KITTI and nuScenes. Our results demonstrate a consistent performance gain for our technique, and verify the need of the more challenging SWAN dataset. Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Dense Video Captioning With Early Linguistic Information FusionabstractDense captioning methods generally detect events in videos first and then generate captions for the individual events. Events are localized solely based on the visual cues while ignoring the associated linguistic information and context. Whereas end-to-end learning may implicitly take guidance from language, these methods still fall short of the power of explicit modeling. In this paper, we propose aVisual-Semantic Embedding (ViSE) Frameworkthat models the word(s)-context distributional properties over the entire semantic space and computes weights for all then-gramssuch that higher weights are assigned to the more informativen-grams. The weights are accounted for in learning distributed representations of all the captions to construct a semantic space. To perform the contextualization of visual information and the constructed semantic space in a supervised manner, we designVisual-Semantic Joint Modeling Network (VSJM-Net). The learnedViSEembeddings are then temporally encoded with aHierarchical Descriptor Transformer (HDT). For caption generation, we exploit a transformer architecture to decode the input embeddings into natural language descriptions. Experiments on the large-scale ActivityNet Captions dataset and YouCook-II dataset demonstrate the efficacy of our method. Nayyer Aafaq, Ajmal Mian, Naveed Akhtar, Wei Liu 0006, Mubarak Shah |
IEEE Trans. Multim. | 3 |
| 2023 | GoMIC: Multi-view image clustering via self-supervised contrastive heterogeneous graph co-learningabstractAbstract Graph learning is being increasingly applied to image clustering to reveal intra-class and inter-class relationships in data. However, existing graph learning-based image clustering focuses on grouping images under a single view, which under-utilises the information provided by the data. To address that, we propose a self-supervised multi-view image clustering technique under contrastive heterogeneous graph learning. Our method computes a heterogeneous affinity graph for multi-view image data. It conducts Local Feature Propagation (LFP) for reasoning over the local neighbourhood of each node and executes an Influence-aware Feature Propagation (IFP) from each node to its influential node for learning the clustering intention. The proposed framework pioneeringly employs two contrastive objectives. The first targets to contrast and fuse multiple views for the overall LFP embedding, and the second maximises the mutual information between LFP and IFP representations. We conduct extensive experiments on the benchmark datasets for the problem, i.e. COIL-20, Caltech7 and CASIA-WebFace. Our evaluation shows that our method outperforms the state-of-the-art methods, including the popular techniques MVGL, MCGC and HeCo. Uno Fang, Jianxin Li 0001, Naveed Akhtar, Yan Jia 0001 |
World Wide Web (WWW) | 3 |
| 2022 | Deformation and Correspondence Aware Unsupervised Synthetic-to-Real Scene Flow Estimation for Point CloudsabstractPoint cloud scene flow estimation is of practical importance for dynamic scene navigation in autonomous driving. Since scene flow labels are hard to obtain, current methods train their models on synthetic data and transfer them to real scenes. However, large disparities between existing synthetic datasets and real scenes lead to poor model transfer. We make two major contributions to address that. First, we develop a point cloud collector and scene flow annotator for GTA-V engine to automatically obtain diverse realistic training samples without human intervention. With that, we develop a large-scale synthetic scene flow dataset GTA-SF. Second, we propose a mean-teacher-based domain adaptation framework that leverages self-generated pseudo-labels of the target domain. It also explicitly incorporates shape deformation regularization and surface correspondence refinement to address distortions and misalignments in domain transfer. Through extensive experiments, we show that our GTA-SF dataset leads to a consistent boost in model generalization to three real datasets (i.e., Waymo, Lyft and KITTI) as compared to the most widely used FT3D dataset. Moreover, our framework achieves superior adaptation performance on six source-target dataset pairs, remarkably closing the average domain gap by 60%. Data and codes are available at https://github.com/leolyj/DCA-SRSFE Yinjie Lei, Naveed Akhtar, Haifeng Li 0007, Munawar Hayat |
CVPR | 3 |
| 2022 | MAiVAR: Multimodal Audio-Image and Video Action RecognizerabstractCurrently, action recognition is predominately performed on video data as processed by CNNs. We investigate if the representation process of CNN s can also be leveraged for multimodal action recognition by incorporating image-based audio representations of actions in a task. To this end, we propose Multimodal Audio-Image and Video Action Recognizer (MAiVAR), a CNN-based audio-image to video fusion model that accounts for video and audio modalities to achieve superior action recognition performance. MAiVAR extracts meaningful image representations of audio and fuses it with video representation to achieve better performance as compared to both modalities individually on a large-scale action recognition dataset. Muhammad Bilal Shaikh, Douglas Chai, Syed M. S. Islam, Naveed Akhtar |
VCIP | 4 |
| 2022 | Transferable 3D Adversarial Textures using End-to-end OptimizationabstractDeep visual models are known to be vulnerable to adversarial attacks. The last few years have seen numerous techniques to compute adversarial inputs for these models. However, there are still under-explored avenues in this critical research direction. Among those is the estimation of adversarial textures for 3D models in an end-to-end optimization scheme. In this paper, we propose such a scheme to generate adversarial textures for 3D models that are highly transferable and invariant to different camera views and lighting conditions. Our method makes use of neural rendering with explicit control over the model texture and background. We ensure transferability of the adversarial textures by employing an ensemble of robust and non-robust models. Our technique utilizes 3D models as a proxy to simulate closer to real-life conditions, in contrast to conventional use of 2D images for adversarial attacks. We show the efficacy of our method with extensive experiments. Camilo Pestana, Naveed Akhtar, Nazanin Rahnavard, Mubarak Shah, Ajmal Mian |
WACV | 2 |
| 2022 | Attack to Fool and Explain Deep NetworksabstractDeep visual models are susceptible to adversarial perturbations to inputs. Although these signals are carefully crafted, they still appear noise-like patterns to humans. This observation has led to the argument that deep visual representation is misaligned with human perception. We counter-argue by providing evidence of human-meaningful patterns in adversarial perturbations. We first propose an attack that fools a network to confuse a whole category of objects (source class) with a target label. Our attack also limits the unintended fooling by samples from non-sources classes, thereby circumscribing human-defined semantic notions for network fooling. We show that the proposed attack not only leads to the emergence of regular geometric patterns in the perturbations, but also reveals insightful information about the decision boundaries of deep models. Exploring this phenomenon further, we alter the 'adversarial' objective of our attack to use it as a tool to 'explain' deep visual representation. We show that by careful channeling and projection of the perturbations computed by our method, we can visualize a model's understanding of human-defined semantic notions. Finally, we exploit the explanability properties of our perturbations to perform image generation, inpainting and interactive image manipulation by attacking adversarialy robust 'classifiers'. In all, our major contribution is a novel pragmatic adversarial attack that is subsequently transformed into a tool to interpret the visual models. The article also makes secondary contributions in terms of establishing the utility of our attack beyond the adversarial objective with multiple interesting applications. Naveed Akhtar, Mohammad A. A. K. Jalwana, Mohammed Bennamoun, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Adversarial Attack on Skeleton-Based Human Action RecognitionabstractDeep learning models achieve impressive performance for skeleton-based human action recognition. Graph convolutional networks (GCNs) are particularly suitable for this task due to the graph-structured nature of skeleton data. However, the robustness of these models to adversarial attacks remains largely unexplored due to their complex spatiotemporal nature that must represent sparse and discrete skeleton joints. This work presents the first adversarial attack on skeleton-based action recognition with GCNs. The proposed targeted attack, termed constrained iterative attack for skeleton actions (CIASA), perturbs joint locations in an action sequence such that the resulting adversarial sequence preserves the temporal coherence, spatial integrity, and the anthropomorphic plausibility of the skeletons. CIASA achieves this feat by satisfying multiple physical constraints and employing spatial skeleton realignments for the perturbed skeletons along with regularization of the adversarial skeletons with generative networks. We also explore the possibility of semantically imperceptible localized attacks with CIASA and succeed in fooling the state-of-the-art skeleton action recognition models with high confidence. CIASA perturbations show high transferability in black-box settings. We also show that the perturbed skeleton sequences are able to induce adversarial behavior in the RGB videos created with computer graphics. A comprehensive evaluation with NTU and Kinetics data sets ascertains the effectiveness of CIASA for graph-based skeleton action recognition and reveals the imminent threat to the spatiotemporal deep learning tasks in general. Jian Liu 0014, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image SaliencyabstractBackpropagation image saliency aims at explaining model predictions by estimating model-centric importance of individual pixels in the input. However, classinsensitivity of the earlier layers in a network only allows saliency computation with low resolution activation maps of the deeper layers, resulting in compromised image saliency. Remedifying this can lead to sanity failures. We propose CAMERAS, a technique to compute high-fidelity backpropagation saliency maps without requiring any external priors and preserving the map sanity. Our method systematically performs multi-scale accumulation and fusion of the activation maps and backpropagated gradients to compute precise saliency maps. From accurate image saliency to articulation of relative importance of input features for different models, and precise discrimination between model perception of visually similar objects, our high-resolution mapping offers multiple novel insights into the black-box deep visual models, which are presented in the paper. We also demonstrate the utility of our saliency maps in adversarial setup by drastically reducing the norm of attack signals by focusing them on the precise regions identified by our maps. Our method also inspires new evaluation metrics and a sanity check for this developing research direction. Mohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal Mian |
CVPR | 2 |
| 2021 | Picasso: A CUDA-Based Library for Deep Learning Over 3D MeshesabstractWe present Picasso, a CUDA-based library comprising novel modules for deep learning over complex real-world 3D meshes. Hierarchical neural architectures have proved effective in multi-scale feature extraction which signifies the need for fast mesh decimation. However, existing methods rely on CPU-based implementations to obtain multi-resolution meshes. We design GPU-accelerated mesh decimation to facilitate network resolution reduction efficiently on-the-fly. Pooling and unpooling modules are defined on the vertex clusters gathered during decimation. For feature learning over meshes, Picasso contains three types of novel convolutions namely, facet2vertex, vertex2facet, and facet2facet convolution. Hence, it treats a mesh as a geometric structure comprising vertices and facets, rather than a spatial graph with edges as previous methods do. Picasso also incorporates a fuzzy mechanism in its filters for robustness to mesh sampling (vertex density). It exploits Gaussian mixtures to define fuzzy coefficients for the facet2vertex convolution, and barycentric interpolation to define the coefficients for the remaining two convolutions. In this release, we demonstrate the effectiveness of the proposed modules with competitive segmentation results on S3DIS. The library will be made public through github. Huan Lei, Naveed Akhtar, Ajmal Mian |
CVPR | 2 |
| 2021 | Boosting Deep Transfer Learning For Covid-19 ClassificationabstractCOVID-19 classification using chest Computed Tomography (CT) has been found pragmatically useful by several studies. Due to the lack of annotated samples, these studies recommend transfer learning and explore the choices of pre-trained models and data augmentation. However, it is still unknown if there are better strategies than vanilla transfer learning for more accurate COVID-19 classification with limited CT data. This paper provides an affirmative answer, devising a novel ‘model’ augmentation technique that allows a considerable performance boost to transfer learning for the task. Our method systematically reduces the distributional shift between the source and target domains and considers augmenting deep learning with complementary representation learning techniques. We establish the efficacy of our method with publicly available datasets and models, along with identifying contrasting observations in the previous studies. Fouzia Altaf, Syed M. S. Islam, Naeem Janjua, Naveed Akhtar |
ICIP | 4 |
| 2021 | Adversarial Attacks and Defense on Deep Learning Classification Models using YCbCr Color ImagesabstractDeep neural network models are vulnerable to adversarial perturbations that are subtle but change the model predictions. Adversarial perturbations are generally computed for RGB images and are, hence, equally distributed among the RGB channels. We show, for the first time, that adversarial perturbations prevail in the Y-channel of the$\mathbf{YC}_{b}\mathbf{C}_{r}$> color space and exploit this finding to propose a defense mechanism. Our defense ResUpNet, which is end-to-end trainable, removes perturbations only from the Y-channel by exploiting ResNet features in a bottleneck free up-sampling framework. The refined Y-channel is combined with the untouched$\mathbf{C}_{b}\mathbf{C}_{r}$-channels to restore the clean image. We compare ResUpNet to existing defenses in the input transformation category and show that it achieves the best balance between maintaining the original accuracies on clean images and defense against adversarial attacks. Finally, we show that for the same attack and fixed perturbation magnitude, learning perturbations only in the Y-channel results in higher fooling rates. For example, with a very small perturbation magnitude$\epsilon=0.002$) the fooling rates of FGSM and PGD attacks on the ResNet50 model increase by 11.1% and 15.6% respectively, when the perturbations are learned only for the Y-channel. Camilo Pestana, Naveed Akhtar, Wei Liu 0006, David G. Glance, Ajmal Mian |
IJCNN | 2 |
| 2021 | Neural computing and applications (NCAA) special issue on best of DICTA 2019 papers
Ajmal Mian, Lei Wang 0108, Ruiping Wang 0001, Hamid Laga, Naveed Akhtar |
Neural Comput. Appl. | 5 |
| 2021 | Spherical Kernel for Efficient Graph Convolution on 3D Point CloudsabstractWe propose a spherical kernel for efficient graph convolution of 3D point clouds. Our metric-based kernels systematically quantize the local 3D space to identify distinctive geometric relationships in the data. Similar to the regular grid CNN kernels, the spherical kernel maintains translation-invariance and asymmetry properties, where the former guarantees weight sharing among similar local structures in the data and the latter facilitates fine geometric learning. The proposed kernel is applied to graph neural networks without edge-dependent filter generation, making it computationally attractive for large point clouds. In our graph networks, each vertex is associated with a single point location and edges connect the neighborhood points within a defined range. The graph gets coarsened in the network with farthest point sampling. Analogous to the standard CNNs, we define pooling and unpooling operations for our network. We demonstrate the effectiveness of the proposed spherical kernel with graph neural networks for point cloud classification and semantic segmentation using ModelNet, ShapeNet, RueMonge2014, ScanNet and S3DIS datasets. The source code and the trained models can be downloaded from https://github.com/hlei-ziyan/SPH3D-GCN. Huan Lei, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Deep Affinity Network for Multiple Object TrackingabstractMultiple Object Tracking (MOT) plays an important role in solving many fundamental problems in video analysis and computer vision. Most MOT methods employ two steps: Object Detection and Data Association. The first step detects objects of interest in every frame of a video, and the second establishes correspondence between the detected objects in different frames to obtain their tracks. Object detection has made tremendous progress in the last few years due to deep learning. However, data association for tracking still relies on hand crafted constraints such as appearance, motion, spatial proximity, grouping etc. to compute affinities between the objects in different frames. In this paper, we harness the power of deep learning for data association in tracking by jointly modeling object appearances and their affinities between different frames in an end-to-end fashion. The proposed Deep Affinity Network (DAN) learns compact, yet comprehensive features of pre-detected objects at several levels of abstraction, and performs exhaustive pairing permutations of those features in any two frames to infer object affinities. DAN also accounts for multiple objects appearing and disappearing between video frames. We exploit the resulting efficient affinity computations to associate objects in the current frame deep into the previous frames for reliable on-line tracking. Our technique is evaluated on popular multiple object tracking challenges MOT15, MOT17 and UA-DETRAC. Comprehensive benchmarking under twelve evaluation metrics demonstrates that our approach is among the best performing techniques on the leader board for these challenges. The open source implementation of our work is available at https://github.com/shijieS/SST.git. Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Attack to Explain Deep RepresentationabstractDeep visual models are susceptible to extremely low magnitude perturbations to input images. Though carefully crafted, the perturbation patterns generally appear noisy, yet they are able to perform controlled manipulation of model predictions. This observation is used to argue that deep representation is misaligned with human perception. This paper counter-argues and proposes the first attack on deep learning that aims at explaining the learned representation instead of fooling it. By extending the input domain of the manipulative signal and employing a model faithful channelling, we iteratively accumulate adversarial perturbations for a deep model. The accumulated signal gradually manifests itself as a collection of visually salient features of the target label (in model fooling), casting adversarial perturbations as primitive features of the target label. Our attack provides the first demonstration of systematically computing perturbations for adversarially non-robust classifiers that comprise salient visual features of objects. We leverage the model explaining character of our algorithm to perform image generation, inpainting and interactive image manipulation by attacking adversarially robust classifiers. The visually appealing results across these applications demonstrate the utility of our attack (and perturbations in general) beyond model fooling. Mohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal Mian |
CVPR | 2 |
| 2020 | SegGCN: Efficient 3D Point Cloud Segmentation With Fuzzy Spherical KernelabstractFuzzy clustering is known to perform well in real-world applications. Inspired by this observation, we incorporate a fuzzy mechanism into discrete convolutional kernels for 3D point clouds as our first major contribution. The proposed fuzzy kernel is defined over a spherical volume that uses discrete bins. Discrete volumetric division can normally make a kernel vulnerable to boundary effects during learning as well as point density during inference. However, the proposed kernel remains robust to boundary conditions and point density due to the fuzzy mechanism. Our second major contribution comes as the proposal of an efficient graph convolutional network, SegGCN for segmenting point clouds. The proposed network exploits ResNet like blocks in the encoder and 1 × 1 convolutions in the decoder. SegGCN capitalizes on the separable convolution operation of the proposed fuzzy kernel for efficiency. We establish the effectiveness of the SegGCN with the proposed kernel on the challenging S3DIS and ScanNet real-world datasets. Our experiments demonstrate that the proposed network can segment over one million points per second with highly competitive performance. Huan Lei, Naveed Akhtar, Ajmal Mian |
CVPR | 2 |
| 2020 | Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking
Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah |
ECCV (24) | 2 |
| 2020 | Efficient Detection of Pixel-Level Adversarial AttacksabstractDeep learning has achieved unprecedented performance in object recognition and scene understanding. However, deep models are also found vulnerable to adversarial attacks. Of particular relevance to robotics systems are pixel-level attacks that can completely fool a neural network by altering very few pixels (e.g. 1-5) in an image. We present the first technique to detect the presence of adversarial pixels in images for the robotic systems, employing an Adversarial Detection Network (ADNet). The proposed network efficiently recognize an input as adversarial or clean by discriminating the peculiar activation signals of the adversarial samples from the clean ones. It acts as a defense mechanism for the robotic vision system by detecting and rejecting the adversarial samples. We thoroughly evaluate our technique on three benchmark datasets including CIFAR-10, CIFAR-100 and Fashion MNIST. Results demonstrate effective detection of adversarial samples by ADNet. Syed Afaq Ali Shah, Moise Bougre, Naveed Akhtar, Mohammed Bennamoun, Liang Zhang 0010 |
ICIP | 3 |
| 2020 | Hyperspectral Recovery from RGB Images using Gaussian ProcessesabstractWe propose to recover spectral details from RGB images of known spectral quantization by modeling natural spectra under Gaussian Processes and combining them with the RGB images. Our technique exploits Process Kernels to model the relative smoothness of reflectance spectra, and encourages non-negativity in the resulting signals for better estimation of the reflectance values. The Gaussian Processes are inferred in sets using clusters of spatio-spectrally correlated hyperspectral training patches. Each set is transformed to match the spectral quantization of the test RGB image. We extract overlapping patches from the RGB image and match them to the hyperspectral training patches by spectrally transforming the latter. The RGB patches are encoded over the transformed Gaussian Processes related to those hyperspectral patches and the resulting image is constructed by combining the codes with the original processes. Our approach infers the desired Gaussian Processes under a fully Bayesian model inspired by Beta-Bernoulli Process, for which we also present the inference procedure. A thorough evaluation using three hyperspectral datasets demonstrates the effective extraction of spectral details from RGB images by the proposed technique. Naveed Akhtar, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Steganographic universal adversarial perturbations
Naveed Akhtar, Muhammad Shahzad Younis, Faisal Shafait, Atif Bin Mansoor, Muhammad Shafique 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video CaptioningabstractAutomatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These methods mainly focus on tailoring sequence learning through RNNs for better caption generation, whereas off-the-shelf visual features are borrowed from CNNs. We argue that careful designing of visual features for this task is equally important, and present a visual feature encoding technique to generate semantically rich captions using Gated Recurrent Units (GRUs). Our method embeds rich temporal dynamics in visual features by hierarchically applying Short Fourier Transform to CNN features of the whole video. It additionally derives high level semantics from an object detector to enrich the representation with spatial dynamics of the detected objects. The final representation is projected to a compact space and fed to a language model. By learning a relatively simple language model comprising two GRU layers, we establish new state-of-the-art on MSVD and MSR-VTT datasets for METEOR and ROUGELmetrics. Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Syed Zulqarnain Gilani, Ajmal Mian |
CVPR | 2 |
| 2019 | Octree Guided CNN With Spherical Kernels for 3D Point CloudsabstractWe propose an octree guided neural network architecture and spherical convolutional kernel for machine learning from arbitrary 3D point clouds. The network architecture capitalizes on the sparse nature of irregular point clouds,and hierarchically coarsens the data representation with space partitioning. At the same time, the proposed spherical kernels systematically quantize point neighborhoods to identify local geometric structures in the data, while maintaining the properties of translation-invariance and asymmetry. We specify spherical kernels with the help of network neurons that in turn are associated with spatial locations.We exploit this association to avert dynamic kernel generation during network training that enables efficient learning with high resolution point clouds. The effectiveness of the proposed technique is established on the benchmark tasks of 3D object classification and segmentation, achieving competitive performance on ShapeNet and RueMonge2014 datasets. Huan Lei, Naveed Akhtar, Ajmal Mian |
CVPR | 2 |
| 2019 | Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition
Jian Liu 0014, Hossein Rahmani 0001, Naveed Akhtar, Ajmal Mian |
Int. J. Comput. Vis. | 3 |
| 2019 | Benchmark Data and Method for Real-Time People Counting in Cluttered Scenes Using Depth SensorsabstractVision-based automatic counting of people has widespread applications in intelligent transportation systems, security, and logistics. However, there is currently no large-scale public dataset for benchmarking approaches on this problem. This paper fills this gap by introducing the first real-world RGB-D people counting dataset (PCDS) containing over 4500 videos recorded at the entrance doors of buses in normal and cluttered conditions. It also proposes an efficient method for counting people in real-world cluttered scenes related to public transportations using depth videos. The proposed method computes a point cloud from the depth video frame and re-projects it onto the ground plane to normalize the depth information. The resulting depth image is analyzed for identifying potential human heads. The human head proposals are meticulously refined using a 3D human model. The proposals in each frame of the continuous video stream are tracked to trace their trajectories. The trajectories are again refined to ascertain reliable counting. People are eventually counted by accumulating the head trajectories leaving the scene. To enable effective head and trajectory identification, we also propose two different compound features. A thorough evaluation on PCDS demonstrates that our technique is able to count people in cluttered scenes with high accuracy at 45 fps on a 1.7-GHz processor, and hence it can be deployed for effective real-time people counting for intelligent transportation systems. Shijie Sun 0001, Naveed Akhtar, Huansheng Song, ChaoYang Zhang, Jianxin Li 0001, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Defense Against Universal Adversarial PerturbationsabstractRecent advances in Deep Learning show the existence of image-agnostic quasi-imperceptible perturbations that when applied to 'any' image can fool a state-of-the-art network classifier to change its prediction about the image label. These 'Universal Adversarial Perturbations' pose a serious threat to the success of Deep Learning in practice. We present the first dedicated framework to effectively defend the networks against such perturbations. Our approach learns a Perturbation Rectifying Network (PRN) as 'pre-input' layers to a targeted model, such that the targeted model needs no modification. The PRN is learned from real and synthetic image-agnostic perturbations, where an efficient method to compute the latter is also proposed. A perturbation detector is separately trained on the Discrete Cosine Transform of the input-output difference of the PRN. A query image is first passed through the PRN and verified by the detector. If a perturbation is detected, the output of the PRN is used for label prediction instead of the actual image. A rigorous evaluation shows that our framework can defend the network classifiers against unseen adversarial perturbations in the real-world scenarios with up to 97.5% success rate. The PRN also generalizes well in the sense that training for one targeted network defends another network with a comparable success rate. Naveed Akhtar, Jian Liu 0014, Ajmal Mian |
CVPR | 1 |
| 2018 | Nonparametric Coupled Bayesian Dictionary and Classifier Learning for Hyperspectral ClassificationabstractWe present a principled approach to learn a discriminative dictionary along a linear classifier for hyperspectral classification. Our approach places Gaussian Process priors over the dictionary to account for the relative smoothness of the natural spectra, whereas the classifier parameters are sampled from multivariate Gaussians. We employ two Beta-Bernoulli processes to jointly infer the dictionary and the classifier. These processes are coupled under the same sets of Bernoulli distributions. In our approach, these distributions signify the frequency of the dictionary atom usage in representing class-specific training spectra, which also makes the dictionary discriminative. Due to the coupling between the dictionary and the classifier, the popularity of the atoms for representing different classes gets encoded into the classifier. This helps in predicting the class labels of test spectra that are first represented over the dictionary by solving a simultaneous sparse optimization problem. The labels of the spectra are predicted by feeding the resulting representations to the classifier. Our approach exploits the nonparametric Bayesian framework to automatically infer the dictionary size-the key parameter in discriminative dictionary learning. Moreover, it also has the desirable property of adaptively learning the association between the dictionary atoms and the class labels by itself. We use Gibbs sampling to infer the posterior probability distributions over the dictionary and the classifier under the proposed model, for which, we derive analytical expressions. To establish the effectiveness of our approach, we test it on benchmark hyperspectral images. The classification performance is compared with the state-of-the-art dictionary learning-based classification methods. Naveed Akhtar, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Joint Discriminative Bayesian Dictionary and Classifier LearningabstractWe propose to jointly learn a Discriminative Bayesian dictionary along a linear classifier using coupled Beta-Bernoulli Processes. Our representation model uses separate base measures for the dictionary and the classifier, but associates them to the class-specific training data using the same Bernoulli distributions. The Bernoulli distributions control the frequency with which the factors (e.g. dictionary atoms) are used in data representations, and they are inferred while accounting for the class labels in our approach. To further encourage discrimination in the dictionary, our model uses separate (sets of) Bernoulli distributions to represent data from different classes. Our approach adaptively learns the association between the dictionary atoms and the class labels while tailoring the classifier to this relation with a joint inference over the dictionary and the classifier. Once a test sample is represented over the dictionary, its representation is accurately labelled by the classifier due to the strong coupling between the dictionary and the classifier. We derive the Gibbs Sampling equations for our joint representation model and test our approach for face, object, scene and action recognition to establish its effectiveness. Naveed Akhtar, Ajmal Mian, Fatih Porikli |
CVPR | 1 |
| 2017 | Efficient classification with sparsity augmented collaborative representation
Naveed Akhtar, Faisal Shafait, Ajmal Mian |
Pattern Recognit. | 1 |
| 2017 | RCMF: Robust Constrained Matrix Factorization for Hyperspectral UnmixingabstractWe propose a constrained matrix factorization approach for linear unmixing of hyperspectral data. Our approach factorizes a hyperspectral cube into its constituent endmembers and their fractional abundances such that the endmembers are sparse nonnegative linear combinations of the observed spectra themselves. The association between the extracted endmembers and the observed spectra is explicitly noted for physical interpretability. To ensure reliable unmixing, we make the matrix factorization procedure robust to outliers in the observed spectra. Our approach simultaneously computes the endmembers and their abundances in an efficient and unsupervised manner. The extracted endmembers are nonnegative quantities, whereas their abundances additionally follow the sum-to-one constraint. We thoroughly evaluate our approach using synthetic data with white and correlated noise as well as real hyperspectral data. Experimental results establish the effectiveness of our approach. Naveed Akhtar, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Hierarchical Beta Process with Gaussian Process Prior for Hyperspectral Image Super Resolution
Naveed Akhtar, Faisal Shafait, Ajmal Mian |
ECCV (3) | 1 |
| 2016 | Discriminative Bayesian Dictionary Learning for ClassificationabstractWe propose a Bayesian approach to learn discriminative dictionaries for sparse representation of data. The proposed approach infers probability distributions over the atoms of a discriminative dictionary using a finite approximation of Beta Process. It also computes sets of Bernoulli distributions that associate class labels to the learned dictionary atoms. This association signifies the selection probabilities of the dictionary atoms in the expansion of class-specific data. Furthermore, the non-parametric character of the proposed approach allows it to infer the correct size of the dictionary. We exploit the aforementioned Bernoulli distributions in separately learning a linear classifier. The classifier uses the same hierarchical Bayesian model as the dictionary, which we present along the analytical inference solution for Gibbs sampling. For classification, a test instance is first sparsely encoded over the learned dictionary and the codes are fed to the classifier. We performed experiments for face and action recognition; and object and scene-category classification using five public datasets and compared the results with state-of-the-art discriminative sparse representation approaches. Experiments show that the proposed Bayesian approach consistently outperforms the existing approaches. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Bayesian sparse representation for hyperspectral image super resolutionabstractDespite the proven efficacy of hyperspectral imaging in many computer vision tasks, its widespread use is hindered by its low spatial resolution, resulting from hardware limitations. We propose a hyperspectral image super resolution approach that fuses a high resolution image with the low resolution hyperspectral image using non-parametric Bayesian sparse representation. The proposed approach first infers probability distributions for the material spectra in the scene and their proportions. The distributions are then used to compute sparse codes of the high resolution image. To that end, we propose a generic Bayesian sparse coding strategy to be used with Bayesian dictionaries learned with the Beta process. We theoretically analyze the proposed strategy for its accurate performance. The computed codes are used with the estimated scene spectra to construct the super resolution hyperspectral image. Exhaustive experiments on two public databases of ground based hyperspectral images and a remotely sensed image show that the proposed approach outperforms the existing state of the art. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
CVPR | 1 |
| 2015 | Futuristic Greedy Approach to Sparse Unmixing of Hyperspectral DataabstractSpectra measured at a single pixel of a remotely sensed hyperspectral image is usually a mixture of multiple spectral signatures (endmembers) corresponding to different materials on the ground. Sparse unmixing assumes that a mixed pixel is a sparse linear combination of different spectra already available in a spectral library. It uses sparse approximation (SA) techniques to solve the hyperspectral unmixing problem. Among these techniques, greedy algorithms suite well to sparse unmixing. However, their accuracy is immensely compromised by the high correlation of the spectra of different materials. This paper proposes a novel greedy algorithm, called OMP-Star, that shows robustness against the high correlation of spectral signatures. We preprocess the signals with spectral derivatives before they are used by the algorithm. To approximate the mixed pixel spectra, the algorithm employs a futuristic greedy approach that, if necessary, considers its future iterations before identifying an endmember. We also extend OMP-Star to exploit the nonnegativity of spectral mixing. Experiments on simulated and real hyperspectral data show that the proposed algorithms outperform the state-of-the-art greedy algorithms. Moreover, the proposed approach achieves results comparable to convex relaxation-based SA techniques, while maintaining the advantages of greedy approaches. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Sparse Spatio-spectral Representation for Hyperspectral Image Super-resolution
Naveed Akhtar, Faisal Shafait, Ajmal Mian |
ECCV (7) | 1 |
| 2014 | SUnGP: A Greedy Sparse Approximation Algorithm for Hyperspectral UnmixingabstractSpectra measured at a pixel of a remote sensing hyper spectral sensor is usually a mixture of multiple spectra (end-members) of different materials on the ground. Hyper spectral unmixing aims at identifying the end members and their proportions (fractional abundances) in the mixed pixels. Hyper spectral unmixing has recently been casted into a sparse approximation problem and greedy sparse approximation approaches are considered desirable for solving it. However, the high correlation among the spectra of different materials seriously affects the accuracy of the greedy algorithms. We propose a greedy sparse approximation algorithm, called SUnGP, for unmixing of hyper spectral data. SUnGP shows high robustness against the correlation of the spectra of materials. The algorithm employees a subspace pruning strategy for the identification of the end members. Experiments show that the proposed algorithm not only outperforms the state of the art greedy algorithms, its accuracy is comparable to the algorithms based on the convex relaxation of the problem, but with a considerable computational advantage. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
ICPR | 1 |
| 2014 | Repeated constrained sparse coding with partial dictionaries for hyperspectral unmixingabstractHyperspectral images obtained from remote sensing platforms have limited spatial resolution. Thus, each spectra measured at a pixel is usually a mixture of many pure spectral signatures (endmembers) corresponding to different materials on the ground. Hyperspectral unmixing aims at separating these mixed spectra into its constituent end-members. We formulate hyperspectral unmixing as a constrained sparse coding (CSC) problem where unmixing is performed with the help of a library of pure spectral signatures under positivity and summation constraints. We propose two different methods that perform CSC repeatedly over the hyperspectral data. However, the first method, Repeated-CSC (RCSC), systematically neglects a few spectral bands of the data each time it performs the sparse coding. Whereas the second method, Repeated Spectral Derivative (RSD), takes the spectral derivative of the data before the sparse coding stage. The spectral derivative is taken such that it is not operated on a few selected bands. Experiments on simulated and real hyperspectral data and comparison with existing state of the art show that the proposed methods achieve significantly higher accuracy. Our results demonstrate the overall robustness of RCSC to noise and better performance of RSD at high signal to noise ratio. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
WACV | 1 |
| 2013 | Unexpected Situations in Service Robot Environment: Classification and Reasoning Using Naive Physics
Anastassia Küstenmacher, Naveed Akhtar, Paul-Gerhard Plöger, Gerhard Lakemeyer |
RoboCup | 2 |