VLDB 2026 Research / reviewers in the wild / expert
Ajmal Mian
dblp:63/807 · also Ajmal S. Mian, Ajmal Saeed Mian
· DBLP profile ↗
222ranked-venue papers
16as first author
119since 2021 · last 2026
0000-0002-5206-3842ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 138 · 11 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 108 · 9 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 14 since 2021Systems, architecture and hardware · 14 · 13 since 2021Security and privacy · 7 · 7 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene GenerationabstractMost 3D scene generation methods are limited to only generating object bounding box parameters while newer diffusion methods also generate class labels and latent features. Using object size or latent feature, they then retrieve objects from a predefined database. For complex scenes of varied, multi-categorical objects, diffusion-based latents cannot be effectively decoded by current autoencoders into the correct point cloud objects which agree with target classes. We introduce a Class-Partitioned Vector Quantized Variational Autoencoder (CPVQ-VAE) that is trained to effectively decode object latent features, by employing a pioneering class-partitioned codebook where codevectors are labeled by class. To address the problem of codebook collapse, we propose a class-aware running average update which reinitializes dead codevectors within each partition. During inference, object features and class labels, both generated by a Latent-space Flow Matching Model (LFMM) designed specifically for scene generation, are consumed by the CPVQ-VAE. The CPVQ-VAE's class-aware inverse look-up then maps generated latents to codebook entries that are decoded to class-specific point cloud shapes. Thereby, we achieve pure point cloud generation without relying on an external objects database for retrieval. Extensive experiments reveal that our method reliably recovers plausible point cloud scenes, with up to 70.4% and 72.3% reduction in Chamfer and Point2Mesh errors on complex living room scenes. Dasith de Silva Edirimuni, Ajmal Mian |
AAAI | 2 |
| 2026 | Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Hossein Rahmani 0001, Nicu Sebe, Ajmal Mian |
Int. J. Comput. Vis. | 10 |
| 2026 | P-MLP: Language-conditioned task planning with multimodal lexical priors over labels
Wei Sun 0028, Yan Zheng 0003, Jian Liu 0014, Genwei Zhang, Hongshan Yu, Ajmal Mian |
Knowl. Based Syst. | 9 |
| 2026 | A lightweight model for perceptual image compression via implicit priors
Hao Wei 0005, Yiwen Jia, Chenyang Ge, Saeed Anwar, Ajmal Mian |
Neural Networks | 6 |
| 2026 | Diffusion-Driven Self-Supervised Learning for Shape Reconstruction and Pose EstimationabstractFully-supervised category-level pose estimation aims to determine the 6-DoF poses of unseen instances from known categories, requiring expensive manual labeling costs. Recently, various self-supervised category-level pose estimation methods have been proposed to reduce the requirement of the annotated datasets. However, most methods rely on synthetic data or 3D CAD model, and they are typically limited to addressing single-object pose problems without considering multi-objective tasks or shape reconstruction. To overcome these challenges and limitations, we introduce a diffusion-driven self-supervised network for multi-object shape reconstruction and categorical pose estimation, only leveraging the shape priors. Specifically, to capture the SE(3)-equivariant pose features and 3D scale-invariant shape information, we present a Prior-Aware Pyramid 3D Point Transformer. This module adopts a point convolutional layer with radial-kernels for pose-aware learning and a 3D scale-invariant graph convolution layer for object-level shape representation. Furthermore, we introduce a Pretrain-to-Refine Self-Supervised Training Paradigm to train our network. It enables proposed network to capture the associations between shape priors and observations, addressing the challenge of intra-class shape variations by utilising the diffusion mechanism. Extensive experiments conducted on four public datasets and a self-built dataset demonstrate that our method significantly outperforms state-of-the-art self-supervised category-level baselines and even surpasses some fully-supervised instance-level and category-level methods. The project page is released at Self-SRPE. Yaonan Wang 0001, Mingtao Feng, Chao Ding 0006, Zheng Shou 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Auto-Feedback Semantic Interface Learning for Text-Based Person Retrieval
Jiayi Li 0003, Min Jiang 0008, Jun Kong 0001, Saeed Anwar, Ajmal Mian |
Pattern Recognit. | 6 |
| 2026 | Prompt-guided selective frequency network for real-world scene text image super-Resolutionabstract• We introduce PGSFNet for real-world scene text image super-resolution. • Adaptive Frequency Modulator is proposed to extract informative frequency components. • Text Information Enhancement module is designed to incorporate text priors. • We develop a Sobel loss to guide optimization towards sharper text details. • PGSFNet is shown to achieve superior performance on public text image datasets. Real-world scene text image super-resolution is challenging due to complex writing strokes, random text distribution, and diverse scene degradations. Existing text super-resolution methods focus on pure text images or fixed-size single-line text, which limits their practical utility. To address that, we propose a Prompt-Guided Selective Frequency super-resolution Network (PGSFNet). Our unique bicephalous neural model comprises a super-resolution branch and a prompt guidance branch. The latter specifically helps in leveraging text content-aware information priors. To that end, we propose a Text Information Enhancement module. To exploit selective frequency information present in the image, PGSFNet employs a proposed Adaptive Frequency Modulator fused with multi-attention structures. Considering the criticality of text edges in our task, we also propose a tailored text edge perception loss. Extensive experiments on the standard open real-world scene text image datasets demonstrate remarkable performance of our method, achieving up to 8.75% PNSR gain for × 2 and 2.28% SSIM gain for × 4 super-resolution on the Real-CE dataset. Our code will be made public at https://github.com/holastq/PGSFNet . Tianqi Shan, Hanlin Qin, Naveed Akhtar, Hossein Rahmani 0001, Ajmal Mian |
Pattern Recognit. | 7 |
| 2026 | Soft-Masked Transformer for Point Cloud Processing With Skip Attention-Based UpsamplingabstractPoint cloud processing methods leverage local and global point features to cater to downstream tasks, yet they often overlook the task-level context inherent in point clouds during the encoding stage. We argue that integrating task-level information into the encoding stage significantly enhances performance. To that end, we propose SMTransformer which incorporates task-level information into a vector-based transformer by utilizing a soft mask generated from task-level queries and keys to learn the attention weights. Additionally, to facilitate effective communication between features from the encoding and decoding layers in high-level tasks such as segmentation, we introduce a skip-attention-based up-sampling block. This block dynamically fuses features from various resolution points across the encoding and decoding layers. To mitigate the increase in network parameters and training time resulting from the complexity of the aforementioned blocks, we propose a novel shared point position encoding strategy. This strategy allows various transformer blocks to share the same position information over the same resolution points, thereby reducing network parameters and training time without compromising accuracy. Experimental comparisons with existing methods on multiple datasets demonstrate the efficacy of SMTransformer and skip-attention-based up-sampling for semantic segmentation task. In particular, we achieve state-of-the-art semantic segmentation results of 73.9% mIoU on S3DIS Area 5 and 62.4% mIoU on SWAN dataset. Note to Practitioners—Point cloud processing underpins automation tasks such as robotic perception, navigation, and inspection, where accurate 3D understanding is essential. Existing methods often prioritize vision benchmarks while overlooking automation needs like efficiency on limited hardware and robustness in real-world environments. The proposed SMTransformer embeds task-level guidance into feature learning and employs skip-attention up-sampling to improve segmentation accuracy with practical efficiency. It is well-suited for robotic manipulation, autonomous driving, and inspection applications. Current limitations include reliance on GPUs and sensitivity to extreme density variations. Future work will target edge-device deployment and multi-task extensions. Yong He 0012, Hongshan Yu, Chaoxu Mu, Mingtao Feng, Tongjia Chen, Zechuan Li, Anwaar Ulhaq, Ajmal Mian |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2026 | Shape and Prototype-Guided Diffusion Model for Grape Amodal Completion in Vineyard Phenotyping SystemsabstractOcclusions often lead to underestimated grape phenotypes, thereby impeding accurate vineyard management and yield prediction. In grape de-occlusion tasks, existing amodal completion methods perform poorly due to complex cluster structures and fine-grained local textures. To this end, we propose a shape and prototype guided diffusion model for high-fidelity amodal completion, by developing a variational shape embedding learning strategy and a dynamic prototype prior extraction module. We model the shape embedding as a mixture of von Mises–Fisher distributions, indicating the potential complete shape of occluded grape clusters. Subsequently, we formulate prototype priors as discrete embeddings that are consistent with patch features, which represent local texture characteristics of grape berries. The shape and prototype embeddings are integrated into the reverse diffusion process via cross-attention mechanisms, stabilizing structural predictions and mitigating common artifacts such as deformation and adhesion. We construct and release a grape amodal completion dataset collected from real-world vineyard environments. Experimental results on the grape dataset and public KINS dataset demonstrate the superiority of our method in terms of perceptual fidelity and phenotypic accuracy, highlighting its effectiveness for vision-based phenotyping in practical vineyard applications. Yihan Wang 0006, Jianqiao Luo, Bailin Li, Mingtao Feng, Ajmal Mian |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Spatial Multimodal Knowledge-Driven 3D Scene Graph Prediction With Vision-Language ModelabstractIn-depth understanding of 3D environments not only involves locating and recognizing individual objects but also requires inferring the relationships and interactions among them. However, most existing methods heavily rely on scene-specific contents, which leads to poor performance due to the noisy, cluttered, and partial nature of real-world 3D scenes. In this work, we find that the inherently hierarchical structures of 3D environments, derived from support relationships, aid in the automatic association of semantic and spatial arrangements of objects and provide rich geometric and topological information independent of specific scenarios. To this end, we propose a 3D scene graph generation model that leverages the hierarchical structures of 3D environments as spatial multimodal knowledge to enhance 3D scene graph generation. Specifically, we first devise a cross-modal tuning approach, where a visually-prompted vision language model is learned to infer the support relationships between objects in a low-resource way. Subsequently, we build a hierarchical visual graph and hierarchical symbolic knowledge graph using the fine-tuned vision language model to extract contextualized visual contents and relevant textual facts, respectively. Finally, we progressively accumulate 3D spatial multimodal knowledge about the hierarchical structures by correlating contextualized visual contents and textual facts using a novel graph reasoning network. In addition, to better evaluate the performance of 3D scene graph generation models, we propose a new benchmark 3DSSG-M by reorganizing the widely-used 3D scene graph generation dataset 3DSSG. This reorganization balances the predicate distribution of 3DSSG and reduces the influence of frequency bias. Extensive results and ablations attest to the effectiveness of the hierarchical structures in 3D environments and demonstrate the superiority of our proposed method over current state-of-the-art competitors. Haoran Hou, Mingtao Feng, Yulan Guo, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Unsupervised Person Re-Identification With Diffusion Model via Semantic-Aware Disentanglement Representation LearningabstractUnsupervised person re-identification (Re-ID) requires learning semantic representation without identity labels. Existing methods entangle identity-related person features with camera-related background features, hindering discriminative feature learning. Also, these methods often disrupt the semantic structure of the person, weakening the semantic representation. In this paper, we propose the Semantic-Aware Disentanglement Representation Learning (SDRL) framework with diffusion models for unsupervised person Re-ID. Firstly, to enhance feature learning, we propose the Disentanglement Aggregation Model (DAM). This model disentangles identity-related features from camera-related features to generate multi-view features. Secondly, to promote the consistency of multi-view features, we design the multi-view similarity consistency (MSC) loss to constrain intra-camera and cross-camera similarity distributions. Thirdly, to generate semantically meaningful patches, we propose the Semantic Spatial Diffusion Model (SSDM). This model operates on identity-related features to perform the denoising diffusion process over spatial transformer parameters. Finally, to further enhance the semantic representation of generated patches, we design the Semantic Decoupled Contrastive (SDC) loss to perceive the inherent semantic structure. Numerous experiments on three demanding datasets prove that our approach is superior to the current unsupervised Re-ID approaches. The source code will be publicly available at https://github.com/taoxuefong/SDRL-reid. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Jiayi Li 0003, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | DiffCom: Decoupled Sparse Priors Guided Diffusion Compression for Point CloudsabstractWhile conventional lossy compression methods predominantly depend on autoencoders to map point clouds into latent representations, they often neglect the intrinsic redundancy within these latent points. To address this limitation, this paper presents a diffusion-based architecture steered by sparse priors, designed to minimize latent redundancy while securing superior reconstruction fidelity, particularly in low-bitrate scenarios. A key feature of the framework is an efficient dual-density data flow that alleviates the stringent size constraints imposed on latent points. By integrating a Probabilistic Attention-based Conditional Denoiser (PACD), the method effectively encapsulates critical reconstruction details within sparse priors, which are hierarchically decoupled into intra- and inter-point components. Specifically, separate encoders are utilized to transform the source point cloud into latent points and decoupled sparse priors, respectively. To dynamically exploit geometric and semantic information, an attention-driven latent denoiser, conditioned on these decoupled priors, is applied across the encoding and decoding layers. Furthermore, inter-point distributions are incorporated into the arithmetic codec to refine local context modeling for sparse points, with the final point cloud recovered via a point decoder. Comprehensive experiments conducted on ShapeNet and standard MPEG PCC datasets demonstrate that the proposed method outperforms state-of-the-art techniques, achieving a superior rate-distortion trade-off. Xiaoge Zhang 0003, Mingtao Feng, Mehwish Nasim, Saeed Anwar, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Second-Order Robust Iterative Pose Optimization for Fine-Grained Cross-View LocalizationabstractFine-grained cross-view localization seeks to estimate precise camera poses by matching ground images with GPS-tagged aerial imagery. Existing methods typically employ first-order iterative optimization to progressively update the camera pose based on cross-view feature correspondences. However, they rely on local features and neglect global and complementary contextual information, making them prone to local optima and slow convergence under large initial errors or strong disturbances. To overcome these limitations, we propose a second-order robust iterative pose estimation framework for fine-grained cross-view localization. Firstly, we devise a second-order deep iterative optimization module to capture complementary forward and backward motion cues, leading to a bidirectional correlation volume. A motion aggregator uses the volume to approximate the dynamics of second-order iterators, substantially facilitating convergence and robustness. In addition, a bidirectional motion-aware robust regularization module mitigates geometric distortions and outlier interference by leveraging bidirectional motion cues to generate fine-grained confidence maps, adaptively suppressing unreliable regions and enhancing the stability of iterative optimization and pose estimation accuracy. Extensive experiments demonstrate that the proposed framework achieves faster convergence and higher pose estimation accuracy than state-of-the-art methods, particularly under large initial errors and challenging conditions. Mingtao Feng, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Image Process. | 6 |
| 2026 | Scalable Unseen Objects 6-DoF Absolute Pose Estimation With Robotic Integration
Jian Liu 0014, Wei Sun 0028, Hossein Rahmani 0001, Ajmal Mian, Lin Wang 0025 |
IEEE Trans. Robotics | 7 |
| 2025 | Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-LocalizationabstractVisual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing. This makes it challenging to automatically control the relative importance of the two tasks, potentially compromising the retrieval task in favor of the offset regression. Consequently, the model may encounter conflicting dominating gradients during joint training. To address this, we propose to model the semantic ambiguity during the offset regression process by integrating associated uncertainty scores, represented as 2D Gaussian distributions, to mitigate negative transfer effects within the joint tasks. We further introduce an uncertainty-aware similarity metric to enhance similarity assessment between query and reference images, accounting for their semantic ambiguities. This metric propagates uncertainty scores into the retrieval task, focusing on certain samples and learning discriminative feature embeddings, allowing the model to adaptively handle conflicting dominating gradients during joint training. Extensive experiments demonstrate that our method improves the overall performance of the joint tasks, achieving state-of-the-art results on the VIGOR and CVACT datasets. Mingtao Feng, Fenghao Tian, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
AAAI | 7 |
| 2025 | Auto-Regressive Diffusion for Generating 3D Human-Object InteractionsabstractText-driven Human-Object Interaction (Text-to-HOI) generation is an emerging field with applications in animation, video games, virtual reality, and robotics. A key challenge in HOI generation is maintaining interaction consistency in long sequences. Existing Text-to-Motion-based approaches, such as discrete motion tokenization, cannot be directly applied to HOI generation due to limited data in this domain and the complexity of the modality. To address the problem of interaction consistency in long sequences, we propose an autoregressive diffusion model (ARDHOI) that predicts the next continuous token. Specifically, we introduce a Contrastive Variational Autoencoder (cVAE) to learn a physically plausible space of continuous HOI tokens, thereby ensuring that generated human-object motions are realistic and natural. For generating sequences autoregressively, we develop a Mamba-based context encoder to capture and maintain consistent sequential actions. Additionally, we implement an MLP-based denoiser to generate the subsequent token conditioned on the encoded context. Our model has been evaluated on the OMOMO and BEHAVE datasets, where it outperforms existing state-of-the-art methods in terms of both performance and inference speed. This makes ARDHOI a robust and efficient solution for text-driven HOI tasks. Zichen Geng, Zeeshan Hayder, Wei Liu 0006, Ajmal Mian |
AAAI | 4 |
| 2025 | Skip Mamba Diffusion for Monocular 3D Semantic Scene Completionabstract3D semantic scene completion is critical for multiple downstream tasks in autonomous systems. It estimates missing geometric and semantic information in the acquired scene data. Due to the challenging real-world conditions, this task usually demands complex models that process multi-modal data to achieve acceptable performance. We propose a unique neural model, leveraging advances from the state space and diffusion generative modeling to achieve remarkable 3D semantic scene completion performance with monocular image input. Our technique processes the data in the conditioned latent space of a variational autoencoder where diffusion modeling is carried out with an innovative state space technique. A key component of our neural network is the proposed Skimba (Skip Mamba) denoiser, which is adept at efficiently processing long-sequence data. The Skimba diffusion model is integral to our 3D scene completion network, incorporating a triple Mamba structure, dimensional decomposition residuals and varying dilations along three directions. We also adopt a variant of this network for the subsequent semantic segmentation stage of our method. Extensive evaluation on the standard SemanticKITTI and SSCBench-KITTI360 datasets show that our approach not only outperforms other monocular techniques by a large margin, it also achieves competitive performance against stereo methods. Li Liang 0006, Naveed Akhtar, Jordan Vice, Xiangrui Kong, Ajmal Mian |
AAAI | 5 |
| 2025 | Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelabstractWe introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the input question, and hence needs implicit spatiotemporal reasoning and grounding. It extends existing spatiotemporal grounding work focusing on explicit action/motion grounding, to a more general format by enabling implicit reasoning via questions. To facilitate the development of the new task, we collect a large-scale dataset called GroundMoRe, which comprises 1,715 video clips, 249K object masks that are deliberately designed with 4 question types for benchmarking deep and comprehensive motion reasoning abilities. GroundMoRe uniquely requires models to generate visual answers, providing a more concrete and visually interpretable response than plain texts. It evaluates models on both spatiotemporal grounding and reasoning, fostering to address complex challenges in motion-related video reasoning, temporal perception, and pixel-level understanding. Furthermore, we introduce a novel baseline model named MoRA, which achieves respectable performance on GroundMoRe outperforming the best existing visual grounding baseline model by an average of 21.5% relatively. We hope this novel and challenging task will pave the way for future advancements in robust and general motion understanding via video reasoning segmentation. Project available at: https://groundmore.github.io/ Andong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang, Lincoln Spencer, Yapeng Tian, Ajmal Mian, Mohit Bansal, Chen Chen 0001 |
CVPR | 7 |
| 2025 | Beyond Human Perception: Understanding Multi-Object World from Monocular ViewabstractLanguage and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG Keyu Guo, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001, Naveed Akhtar, Ajmal Mian |
CVPR | 11 |
| 2025 | Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object RecognitionabstractRecent open-world representation learning approaches have leveraged CLIP to enable zero-shot 3D object recognition. However, performance on real point clouds with occlusions still falls short due to unrealistic pretraining settings. Additionally, these methods incur high inference costs because they rely on Transformer’s attention modules. In this paper, we make two contributions to address these limitations. First, we propose occlusion-aware text-image-point cloud pretraining to reduce the training-testing domain gap. From 52K synthetic 3D objects, our framework generates nearly 630K partial point clouds for pretraining, consistently improving real-world recognition performances of existing popular 3D networks. Second, to reduce computational requirements, we introduce DuoMamba, a two-stream linear state space model tailored for point clouds. By integrating two spacefilling curves with 1D convolutions, DuoMamba effectively models spatial dependencies between point tokens, offering a powerful alternative to Transformer. When pretrained with our framework, DuoMamba surpasses current state-of-the-art methods while reducing latency and FLOPs, highlighting the potential of our approach for realworld applications. Our code and data are available at ndkhanh360.github.io/project-occtip. Ghulam M. Hassan, Ajmal Mian |
CVPR | 3 |
| 2025 | Mono3DVLT: Monocular-Video-Based 3D Visual Language TrackingabstractVisual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal with 3D tracking in the confines of monocular video. Unfortunately, advances in 3D tracking mainly rely on expensive sensor inputs, e.g., point clouds, depth measurements, radar. Absence of language counterpart for the outputs of these mildly democratized sensors in the literature also hinders VLT expansion to 3D tracking. Addressing that, we make the first attempt towards extending VLT to 3D tracking based on monocular video. We present a comprehensive framework, introducing (i) the Monocular-Video-based 3D Visual Language Tracking (Mono3DVLT) task, (ii) a large-scale dataset for the task, called Mono3DVLT-V2X, and (iii) a customized neural model for the task. Our dataset is carefully curated, leveraging a Large Langauge Model (LLM) followed by human verification, composing natural language descriptions for 79,158 video sequences aiming at single object tracking, providing 2D and 3D bounding box annotations. Our neural model, termed Mono3DVLT-MT, is the first targeted approach for the Mono3DVLT task. Comprising the pipeline of multi-modal feature extractor, visual-language encoder, tracking decoder and a tracking head, our model sets a strong baseline for the task on Mono3DVLT-V2X. Experimental results show that our method significantly outperforms existing techniques on the Mono3DVLT-V2X dataset. Our dataset and code are available in https://github.com/hongkai-wei/Mono3DVLT. Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Hongli Hu, Huansheng Song, Naveed Akhtar, Ajmal Mian |
CVPR | 11 |
| 2025 | Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generationabstract3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework designed to efficiently and controllably generate high-resolution 3D models from text prompts or images. Our key insights are three-fold: 1) Hierarchical Gaussian Mixture Model Splatting: We propose a hybrid hierarchical representation to extract fixed number of fine-grained Gaussians with multiscale details from textured object, also establish part-level representation of Gaussians primitives. 2) Mamba with adaptive tree topology: We present a diffusion mamba with tree-topology to adaptively generate Gaussians with disordered spatial structures, without the need for complex preprocessing and maintain linear complexity generation. 3) Controllable Generation: Building on the HGMM tree, we introduce a cascaded diffusion framework combining controllable implicit latent generation, which progressively generates condition-driven latents, and explicit splatting generation, which transforms latents into high-quality Gaussian primitives. Extensive experiments demonstrate the high fidelity and efficiency of our approach. Qitong Yang, Mingtao Feng, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
CVPR | 7 |
| 2025 | MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion ModelabstractObject pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on shape priors or CAD models of the intra-class known objects. We propose a diffusion-based monocular category-level 9D object pose generation method, MonoDiff9D. Our motivation is to leverage the probabilistic nature of diffusion models to alleviate the need for shape priors, CAD models, or depth sensors for intra-class unknown object pose estimation. We first estimate coarse depth via DINOv2 from the monocular image in a zero-shot manner and convert it into a point cloud. We then fuse the global features of the point cloud with the input image and use the fused features along with the encoded time step to condition MonoDiff9D. Finally, we design a transformer-based denoiser to recover the object pose from Gaussian noise. Extensive experiments on two popular benchmark datasets show that MonoDiff9D achieves state-of-the-art monocular category-level 9D object pose estimation accuracy without the need for shape priors or CAD models at any stage. Our code will be made public at https://github.com/CNJianLiu/MonoDiff9D. Jian Liu 0014, Wei Sun 0028, Zichen Geng, Hossein Rahmani 0001, Ajmal Mian |
ICRA | 7 |
| 2025 | Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment RetrievalabstractCurrent text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieves the target moment using purified multimodal representations. Following this paradigm, we introduce the Denoise-then-Retrieve Network (DRNet), comprising Text-Conditioned Denoising (TCD) and Text-Reconstruction Feedback (TRF) modules. TCD integrates cross-attention and structured state space blocks to dynamically identify noisy clips and produce a noise mask to purify multimodal video representations. TRF further distills a single query embedding from purified video representations and aligns it with the text embedding, serving as auxiliary supervision for denoising during training. Finally, we perform conditional retrieval using text embeddings on purified video representations for accurate VMR. Experiments on Charades-STA and QVHighlights demonstrate that our approach surpasses state-of-the-art methods on all metrics. Furthermore, our denoise-then-retrieve paradigm is adaptable and can be seamlessly integrated into advanced VMR models to boost performance. Jiuxin Cao, Bo Miao, Zhiheng Fu, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IJCAI | 9 |
| 2025 | Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor ScenesabstractFusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and RGB input has started gaining traction. However, effective integration of these modalities for precise object detection tasks still remains a largely open problem. To address that, we propose a MultiStream Detection (MuStD) network, which meticulously extracts task-relevant information from both data modalities. The network follows a three-stream structure. Its LiDAR-PillarNet stream extracts sparse 2D pillar features from the LiDAR input while the LiDAR-Height Compression stream computes Bird’s-Eye View features. An additional 3D Multimodal stream combines RGB and LiDAR features using UV mapping and polar coordinate indexing. Eventually, the features containing comprehensive spatial, textural, and geometric information are carefully fused and fed to a detection head for 3D object detection. We evaluate our method on the challenging KITTI Object Detection Benchmark, with results available on the official evaluation server.1. Our approach achieves strong performance, with an average precision (AP) of 85.39% in 3D detection, 91.34% in Bird’s Eye View (BEV) detection, and 96.39% in 2D detection. These results match or surpass existing state-of-the-art methods. In the difficult "Hard" category, our method attains 80.78% AP in 3D detection and 94.04% AP in 2D detection, highlighting its robustness in challenging scenarios. Furthermore, our method runs at 67 ms, demonstrating efficiency and real-time capability. Our code will be released through the MuStD GitHub repository at https://github.com/IbrahimUWA/MuStD. Muhammad Ibrahim 0001, Naveed Akhtar, Haitian Wang 0002, Saeed Anwar, Ajmal Mian |
IROS | 5 |
| 2025 | Attention-Guided Vector Quantized Variational Autoencoder for Brain Tumor Segmentation
Ajmal Mian, Naveed Akhtar, Ghulam M. Hassan |
MICCAI (1) | 2 |
| 2025 | CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene GenerationabstractOutdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large‑scale benchmark for generating 3D outdoor semantic scenes from abstract freehand sketches and pseudo‑labeled annotations of satellite images. SketchSem3D includes two subsets, Sketch-based SemanticKITTI and Sketch-based KITTI-360 (containing LiDAR voxels along with their corresponding sketches and annotated satellite images), to enable standardized, rigorous, and diverse evaluations. We also propose Cylinder Mamba Diffusion (CymbaDiff) that significantly enhances spatial coherence in outdoor 3D scene generation. CymbaDiff imposes structured spatial ordering, explicitly captures cylindrical continuity and vertical hierarchy, and preserves both physical neighborhood relationships and global context within the generated scenes. Extensive experiments on SketchSem3D demonstrate that CymbaDiff achieves superior semantic consistency, spatial realism, and cross-dataset generalization. The code and dataset will be available at here. Li Liang 0006, Bo Miao, Naveed Akhtar, Jordan Vice, Ajmal Mian |
NeurIPS | 6 |
| 2025 | UWB-IMU-Odometer Fusion for Simultaneous Calibration and LocalizationabstractThe location accuracy of fixed anchors plays a pivotal role in ultrawideband (UWB) positioning. However, existing calibration methods for calculating the anchor positions require anchor-to-anchor communication or known initial values for anchor locations. Moreover, the calibration accuracy is adversely affected by non line-of-sight (NLOS) conditions. We propose a two-stage calibration scheme to conduct simultaneous calibration and localization (SCAL) based on UWB-inertial measurement unit (IMU)–odometer sensor fusion without the need for anchor-to-anchor communication or manual intervention. Our method performs IMU-odometer-aided multidimensional scaling (IO-MDS) to provide the initial calibration value without anchor-to-anchor ranging measurements. This is followed by a novel factor-graph-based framework to achieve coarse-to-fine calibration based on UWB, inertial, and odometer measurements. In existing works, a single UWB range measurement is regarded as a weak constraint, as it often leads to incorrect estimates. Our method uses the derived radial velocity (DRV) and IO-MDS factors as additional strong constraints for reliable estimates. To minimize the NLOS influence on the calibration process, we introduce an improved least square-support vector machine (ILS-SVM) based on adaptive weight parameter and multikernel function. Experimental results on two field collected datasets show enhancements in NLOS identification and SCAL. In the lab dataset, identification accuracy increased by 5.5%, with improvements of 0.124 m and 0.309 m in root mean square error (RMSE) for robot and anchor locations, respectively. In the parking lot dataset, identification accuracy improved by 5.0%, with RMSE improvements of 0.223 m and 0.317 m for robot and anchor locations, respectively. Wei Sun 0028, Jian Liu 0014, Ajmal Mian |
IEEE Internet Things J. | 6 |
| 2025 | Hyperrectangle Embedding for Debiased 3D Scene Graph Prediction From RGB Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment and is regarded as a prerequisite for long-term autonomous robotic operations. A practical research problem here is to predict the 3D scene graph from sequentially captured data. However, existing methods neglect the polysemy of semantic roles that coarse feature vectors are insufficient to represent entities in different relationship semantics. This extremely limits their capability to predict relationships. We propose an approach to tackle the aforementioned challenge by introducing a novel representation, the hyperrectangle embedding, which represents entity using distinctive geometry for more effective scene understanding, rather than learning within vector-based feature with blindly increasing dimensions. By incorporating an entity within two affine-transformed embeddings, each representing either the subject or object and characterized by separate learnable transformations, we achieve the polysemy of semantic roles. The intersections of affine-transformed hyperrectangle embeddings represent the bidirectional relationship between two entities. We identify bias and reliability as two challenges impeding the model learning process. In response to the bias, that arises from long-tailed distributions in the data, we propose a history-guided debiasing strategy that utilizes a confusion history block comprised of previous hyperrectangle embeddings. This strategy mitigates inherent biases by extracting pertinent information and facilitating knowledge transfer from dominant categories to rare ones. To enhance the reliability of predictions, we introduce predictive uncertainty into the 3D scene graph prediction task. We develop a post-hoc reliability enhancement strategy to identify potentially unreliable predictions and subsequently enhance the model's predictive accuracy. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming existing state-of-the-art. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose EstimationabstractNine-degrees-of-freedom (9-DoF) object pose and size estimation is crucial for enabling augmented reality and robotic manipulation. Category-level methods have received extensive research attention due to their potential for generalization to intra-class unknown objects. However, these methods require manual collection and labeling of large-scale real-world training data. To address this problem, we introduce a diffusion-based paradigm for domain-generalized category-level 9-DoF object pose estimation. Our motivation is to leverage the latent generalization ability of the diffusion model to address the domain generalization challenge in object pose estimation. This entails training the model exclusively on rendered synthetic data to achieve generalization to real-world scenes. We propose an effective diffusion model to redefine 9-DoF object pose estimation from a generative perspective. Our model does not require any 3D shape priors during training or inference. By employing the Denoising Diffusion Implicit Model, we demonstrate that the reverse diffusion process can be executed in as few as 3 steps, achieving near real-time performance. Finally, we design a robotic grasping system comprising both hardware and software components. Through comprehensive experiments on two benchmark datasets and the real-world robotic system, we show that our method achieves state-of-the-art domain generalization performance. Jian Liu 0014, Wei Sun 0028, Pengchao Deng, Chongpei Liu, Nicu Sebe, Hossein Rahmani 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | HDTCNet: A hybrid-dimensional convolutional network for multivariate time series classification
Yongli Gu, Hanlin Qin, Naveed Akhtar, Shuai Yuan 0013, Honghao Fu, Shuowen Yang, Ajmal Mian |
Pattern Recognit. | 8 |
| 2025 | History-Enhanced 3D Scene Graph Reasoning From RGB-D Sequencesabstract3D scene graph has emerged as a powerful high-level representation of the environment, and is considered a prerequisite for long-term autonomous robotic operations. However, building rich representations from RGB-D sequences remains a challenging problem. Existing methods ignore the semantic gap between linguistic and geometric feature spaces or neglect the importance of historical context in incrementally captured data. This limits the learning of visual-textual correspondence and the capability of relationship prediction. To address these problems, we propose a history-enhanced 3D scene graph reasoning framework that incrementally builds a consistent 3D semantic scene graph from an RGB-D image sequence. Specifically, we first introduce a cross-domain unified feature representation module to describe the object instances and their relationships distinctly. Next, we build a one-hot candidate matrix-enabled recurrent mechanism to reason the 3D scene graph, combining the perceived global and local history information. Finally, we design history-aware supervised semantics contrastive learning to optimize the scene-specific global history features. Extensive experiments on the 3DSSG dataset show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be available athttps://github.com/cbyan1003/HE-3DSGR. Mingtao Feng, Chenbo Yan, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | RDEIC: Accelerating Diffusion-Based Extreme Image Compression With Relay Residual DiffusionabstractDiffusion-based extreme image compression methods have achieved impressive performance at extremely low bitrates. However, constrained by the iterative denoising process that starts from pure noise, these methods are limited in both fidelity and efficiency. To address these two issues, we present Relay Residual Diffusion Extreme Image Compression (RDEIC), which leverages compressed feature initialization and residual diffusion. Specifically, we first use the compressed latent features of the image with added noise, instead of pure noise, as the starting point to eliminate the unnecessary initial stages of the denoising process. Second, we directly derive a novel residual diffusion equation from Stable Diffusion’s original diffusion equation that reconstructs the raw image by iteratively removing the added noise and the residual between the compressed and target latent features. In this way, we effectively combine the efficiency of residual diffusion with the powerful generative capability of Stable Diffusion. Third, we propose a fixed-step fine-tuning strategy to eliminate the discrepancy between the training and inference phases, thereby further improving the reconstruction quality. Extensive experiments demonstrate that the proposed RDEIC achieves state-of-the-art visual quality and outperforms existing diffusion-based extreme image compression methods in both fidelity and efficiency. The source code and pre-trained models are available at https://github.com/huai-chang/RDEIC. Hao Wei 0005, Chenyang Ge, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Discriminative Correspondence Estimation for Unsupervised RGB-D Point Cloud RegistrationabstractPoint cloud registration is a fundamental task for estimating the rigid transformation matrix between two point clouds, and is regarded as a prerequisite for downstream vision tasks. Recent works have sought to address the registration problem using the obtainable RGB-D sequence, rather than relying solely on point clouds, which may not always be available. However, most existing unsupervised RGB-D point cloud registration works struggle to obtain fine-grained, robust, discriminative correspondences due to the simple concatenation of multimodal features and the increase in vector dimensions. These methods typically follow a common paradigm: extracting features from the input data, estimating correspondences, and obtaining the transformation matrix through geometric fitting. In this work, we design a generative feature extraction module to fully leverage multimodal information, and seek a novel perspective for correspondence estimation which expands the points in the source and target point clouds into hyperrectangle-based embeddings and considers their inner relationships, based on intersections in n-dimensional space, as the basis for estimating correspondences. Each hyperrectangle-based embedding is built upon the natural and discriminative semantics from the proposed generative feature extraction module, which involves a diffusion branch, a geometric branch, and point-pixel fusion. We harness the capability of the generative model to fully leverage the information from both complementary modalities in RGB-D frames. Furthermore, this distinctive geometry space allows for efficient calculation of intersection volumes and model conditional probabilistics for estimating correspondences. Extensive experiments on the 3DMatch and ScanNet datasets show the effectiveness of the proposed method in this challenging task, outperforming state-of-the-art approaches. Our code will be released at:https://github.com/cbyan1003/DCE. Chenbo Yan, Mingtao Feng, Yulan Guo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | A State Space Model for Multiobject Full 3-D Information Estimation From RGB-D ImagesabstractVisual understanding of 3-D objects is essential for robotic manipulation, autonomous navigation, and augmented reality. However, existing methods struggle to perform this task efficiently and accurately in an end-to-end manner. We propose a single-shot method based on the state space model (SSM) to predict the full 3-D information (pose, size, shape) of multiple 3-D objects from a single RGB-D image in an end-to-end manner. Our method first encodes long-range semantic information from RGB and depth images separately and then combines them into an integrated latent representation that is processed by a modified SSM to infer the full 3-D information in two separate task heads within a unified model. A heatmap/detection head predicts object centers, and a 3-D information head predicts a matrix detailing the pose, size and latent code of shape for each detected object. We also propose a shape autoencoder based on the SSM, which learns canonical shape codes derived from a large database of 3-D point cloud shapes. The end-to-end framework, modified SSM block and SSM-based shape autoencoder form major contributions of this work. Our design includes different scan strategies tailored to different input data representations, such as RGB-D images and point clouds. Extensive evaluations on the REAL275, CAMERA25, and Wild6D datasets show that our method achieves state-of-the-art performance. On the large-scale Wild6D dataset, our model significantly outperforms the nearest competitor, achieving 2.6% and 5.1% improvements on the IOU-50 and 5°10 cm metrics, respectively. Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Jian Liu 0014, Jianan Huang 0002, Ajmal Mian |
IEEE Trans. Cybern. | 7 |
| 2025 | Quantifying Bias in Text-to-Image Generative ModelsabstractBias in text-to-image (T2I) generation can propagate unfair social representations and may be exploited to push ulterior agendas. These biases raise concerns on the dependability and fairness of models that have become widely popular and readily available for public consumption. Existing works in T2I bias analysis typically focus on social biases. We look beyond that and instead propose an evaluation methodology to quantify general bias in T2I generative models without any preconceived notion. We introduce a suite of three metrics; namely, distribution bias, Jaccard hallucination and generative miss-rate, to extensively appraise general model bias. To validate the efficacy of these metrics, we also introduce a backdoor-inspired strategy, which provides a convenient handle over the extent of bias in a model for controlled analysis. We assess T2I models implementing six widely used pipelines in this domain. Our extensive analysis covers both general and task-oriented scenarios, employing over 105 K generated images. For prior art comparison, it also encompasses social bias analysis. Moreover, we also extend our technique to analyze bias in seven popular captioned image datasets. Our experiments establish that our approach is objective, domain-agnostic and it consistently measures different forms of T2I model biases. To further research efforts into T2I model biases, we have developed an open-source web application and practical implementation of this work, which is available onhttps://huggingface.co/spaces/JVice/try-before-you-biasHuggingFace. All relevant code is also publicly available onhttps://github.com/JJ-Vice/TryBeforeYouBiasGitHub. Jordan Vice, Naveed Akhtar, Richard I. Hartley, Ajmal Mian |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Hypergraph BiFormer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractWhile transformers are powerful neural network architectures for feature learning, current Transformer-based approaches for semantic segmentation of high-resolution remote sensing images (HRRSIs) struggle with the extraction of local semantic features. To address this issue, we incorporate a hypergraph into the Transformer. Hypergraph-based methods are proficient at discovering high-order correlations within limited-scale data, extracting pertinent representations to enhance the Transformer’s learning capabilities. We also propose dual pooling and feature aggregation modules (FAMs), inspired by the adaptive pooling’s potent local modeling capabilities, to additionally extract fine-grained features from HRRSIs. In particular, we conceive a hypergraph BiFormer (HGBT) based on these three proposed modules along with a BiFormer backbone. HGBT has the potential to learn general latent features as well as generate high-order representations of HRRSIs by modeling correlations of multiscale features and local topology within an entirely nonlinear space, leading to the aggregation of features in a compact and localized manner, enhancing the model’s ability to capture detailed variations within small areas. We validate our approach through extensive experiments on ISPRS Vaihingen and Potsdam datasets, where HGBT attains mean intersection over union (mIoU) of 83.71% and 87.88%, respectively. Both quantitative and qualitative assessments underscore the dominance of HGBT. Our code will be accessible at:https://github.com/ZhangIceNight/HGBFormer. Weipeng Jing 0001, Donglin Di, Chao Li 0066, Mahmoud Emam, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Division and Union: Latent Model WatermarkingabstractModel watermarking is a widely adopted mechanism for protecting deep learning (DL) model intellectual property (IP). Black-box verifiable watermarking typically involves injecting backdoors that cause the model to produce predetermined outputs for specific inputs. In contrast, white-box verifiable watermarking uses steganographic techniques to embed watermarks into weight parameters or activation values. However, the former poses new security risks, while the latter often lacks robustness against removal techniques. In this paper, we propose a latent model watermarking, constructing upon the model Division and Union operating concept, dubbed as DUO, leveraging the strengths of two watermarking methods above while eliminating each shortcoming. Once the model owner or provider embeds a watermark into the model using watermark data, the watermarked model is divided into two parts: the main model, which corresponds to the primary task and is made publicly available, and a small sub-network privately reserved by the owner. The watermark resides latently within the main model and can only be activated through the private sub-network (the reserved parameters) when they are united. Consequently, DUO does not adversely affect the performance of the main model on its primary task and does not induce any security risks, even in the presence of watermark data. We extensively validate DUO on four benchmark datasets (CIFAR-10, ImageNette, CIFAR-100, and Tiny-ImageNet) using various model architectures, including standardized ResNet and VGG. The results affirm its capability to accurately verify model ownership without compromising model accuracy. It exhibits a 100% detection accuracy on pirated/positive testing models (96 models are tested) with a 0% false positive rate on normal/negative testing models (64 models are tested). Due to its latent nature, DUO is both effective and robust, capable of withstanding a wide range of state-of-the-art watermark laundering including severe model fine-tuning and pruning. We further evaluate and demonstrate that DUO remains robust against adaptive attacks, even when both the watermark data and the reserved parameters are known to the adversary. Zhiyang Dai, Yansong Gao 0001, Boyu Kuang, Yifeng Zheng 0001, Ajmal Mian, Anmin Fu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Comprehensive Overview of Large Language ModelsabstractLarge Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multimodal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to provide not only a systematic survey but also a quick, comprehensive reference for the researchers and practitioners to draw insights from extensive, informative summaries of the existing works to advance the LLM research. Humza Naveed, Asad Ullah Khan, Shi Qiu 0001, Saeed Anwar, Muhammad Usman 0010, Naveed Akhtar, Nick Barnes, Ajmal Mian |
ACM Trans. Intell. Syst. Technol. | 9 |
| 2025 | Exploring Hierarchical Spatial Layout Cues for 3D Point Cloud Based Scene Graph Predictionabstract3D scene graph prediction is important for intelligent agents to gather information and perceive semantics of their environments. However, constructing an effective graph is nontrivial given the complexity of natural scenes. Existing solutions for graph representation of 3D scenes still distinguish each detailed discrepancy among all the relationships as flat thinking, ignoring the mechanism used by humans to perform this task. Inspired by the role of the prefrontal cortex in hierarchical reasoning, we analyze this problem from a novel perspective: exploring hierarchical spatial layout cues in 3D space and navigating that hierarchy to make the 3D scene graph more accurate in a vertical division to horizontal propagation strategy. To this end, we first encode the contextual object features for fine-gained object category classification. Next, we build a bottom-up hierarchical graph to predict remarkably diverse support relationships in a single concept regardless of numerous irrelevant relationships. Finally, equipped with the spatially-true and semantically-meaningful support relationships, we focus on the local region layout to propagate the semantic features to predict the additional non-support relationships under the guidance of the given referred hierarchical graph nodes. Experiments on the challenging 3DSSG benchmark show that our algorithm outperforms existing state-of-the-art, and can also alleviate the impact of the long-tailed distribution of training data. Our code is available athttps://github.com/HHrEtvP/HSLC-3DSG/. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Multim. | 7 |
| 2025 | Context-Enhanced Video Moment Retrieval With Large Language ModelsabstractCurrent methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives. To tackle this issue, we propose a Large Language Model-guided Moment Retrieval (LMR) approach that employs the extensive knowledge of Large Language Models (LLMs) to improve video context representation as well as cross-modal alignment, facilitating accurate localization of target moments. Specifically, LMR introduces a context enhancement technique with LLMs to generate crucial target-related context semantics. These semantics are integrated with visual features for producing discriminative video representations. Finally, a language-conditioned transformer is designed to decode free-form language queries, on the fly, using aligned video representations for moment retrieval. Extensive experiments demonstrate that LMR achieves state-of-the-art results, outperforming the nearest competitor by up to 3.28% and 4.06% on the challenging QVHighlights and Charades-STA benchmarks, respectively. More importantly, the performance gains are significantly higher for localization of complex queries. Bo Miao, Jiuxin Cao, Xuelin Zhu, Jiawei Ge 0002, Bo Liu 0004, Mehwish Nasim, Ajmal Mian |
IEEE Trans. Multim. | 8 |
| 2025 | Full Point Encoding for Local Feature Aggregation in 3-D Point CloudsabstractPoint cloud processing methods exploit local point features and global context through aggregation which does not explicitly model the internal correlations between local and global features. To address this problem, we propose full point encoding which is applicable to convolution and transformer architectures. Specifically, we propose full point convolution (FuPConv) and full point transformer (FPTransformer) architectures. The key idea is to adaptively learn the weights from local and global geometric connections, where the connections are established through local and global correlation functions, respectively. FuPConv and FPTransformer simultaneously model the local and global geometric relationships as well as their internal correlations, demonstrating strong generalization ability and high performance. FuPConv is incorporated in classical hierarchical network architectures to achieve local and global shape-aware learning. In FPTransformer, we introduce full point position encoding in self-attention, that hierarchically encodes each point position in the global and local receptive field. We also propose a shape-aware downsampling block that takes into account the local shape and the global context. Experimental comparison to existing methods on benchmark datasets shows the efficacy of FuPConv and FPTransformer for semantic segmentation, object detection, classification, and normal estimation tasks. In particular, we achieve state-of-the-art semantic segmentation results of 76.8% mIoU on S3DIS sixfold and 73.1% on S3DIS Area 5. Our code is available at https://github.com/hnuhyuwa/FullPointTransformer. Yong He 0012, Hongshan Yu, Zhengeng Yang, Wei Sun 0028, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Pixel-Level Noise Mining for Weakly Supervised Salient Object DetectionabstractTraining a deep model for visual saliency detection requires the collection and labor-intensive annotation of overwhelmingly large data. We propose to learn saliency detection in a weakly supervised manner from single noisy label, which is easy to obtain from unsupervised handcrafted feature-based methods. However, deep networks tend to overfit such noises leading to a dramatic drop in accuracy. Given our goal, we address a natural question: can we identify outliers during network prediction and rectify the label noises? To this end, we propose a pixel-level noise mining framework for robust salient object detection (SOD) by exploiting its own knowledge, and without the need for external models. Specifically, during the early training stage, we progressively identify the outliers from a novel perspective during saliency detection, before the network overfits to the noisy labels, and generate a selection matrix in each iteration. Next, we adaptively rectify the label noises under the guidance of the selection matrix for better supervision in the later training stage. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our method showing its ability to learn saliency detection comparable to state-of-the-art fully supervised methods. Furthermore, our approach outperforms existing weakly supervised methods utilizing single noisy label and surpasses the half of existing weakly supervised methods employing multiple noisy labels. Our approach, which trains with multiple noisy labels, outperforms all other methods employing multiple noisy labels across four major datasets. Furthermore, we also evaluate the generalization ability of our method on the multiclass semantic segmentation (SS) task. Our code is available at https://github.com/kendongdong/NoiseMining. Kendong Liu, Mingtao Feng, Wei Zhao 0019, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | MH6D: Multi-Hypothesis Consistency Learning for Category-Level 6-D Object Pose EstimationabstractSix-degree-of-freedom (6DoF) object pose estimation is a crucial task for virtual reality and accurate robotic manipulation. Category-level 6DoF pose estimation has recently become popular as it improves generalization to a complete category of objects. However, current methods focus on data-driven differential learning, which makes them highly dependent on the quality of the real-world labeled data and limits their ability to generalize to unseen objects. To address this problem, we propose multi-hypothesis (MH) consistency learning (MH6D) for category-level 6-D object pose estimation without using real-world training data. MH6D uses a parallel consistency learning structure, alleviating the uncertainty problem of single-shot feature extraction and promoting self-adaptation of domain to reduce the synthetic-to-real domain gap. Specifically, three randomly sampled pose transformations are first performed in parallel on the input point cloud. An attention-guided category-level 6-D pose estimation network with channel attention (CA) and global feature cross-attention (GFCA) modules is then proposed to estimate the three hypothesized 6-D object poses by extracting and fusing the global and local features effectively. Finally, we propose a novel loss function that considers both the process and the final result information allowing MH6D to perform robust consistency learning. We conduct experiments under two different training data settings (i.e., only synthetic data and synthetic and real-world data) to verify the generalization ability of MH6D. Extensive experiments on benchmark datasets demonstrate that MH6D achieves state-of-the-art (SOTA) performance, outperforming most data-driven methods even without using any real-world data. The code is available at https://github.com/CNJianLiu/MH6D. Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | TriCI: Triple Cross-Intra Branch Contrastive Learning for Point Cloud AnalysisabstractWhereas contrastive learning eliminates the need for labeled data, existing methods may suffer from inadequate features due to the conventional single shared encoder structure and struggle to fully harness the rich spectrum of 3D augmentations. In this paper, we propose TriCI, a self-supervised method that designs a triple-branch contrastive learning architecture. During contrastive pre-training, we generate three augmented versions of each input point cloud sample and pair each augmented sample with the original one, resulting in three unique positive pairs. We subsequently feed the pairs into three distinct encoders, each of which extracts features from its corresponding input positive pair. We design a novel cross-branch contrastive loss and use it along with the intra-branch contrastive loss to jointly train our network. The proposed cross-branch loss effectively aligns the output features from different perspectives for pre-training and facilitates their integration for downstream tasks, particularly in object-level scenarios. The intra-branch loss helps maximize the feature correspondences within positive pairs. Extensive experiments demonstrate the superiority of our TriCI in self-supervised learning, and show its strong ability in enhancing the performance of downstream object classification and part segmentation tasks. Interestingly, our TriCI achieves a 92.9% accuracy for linear SVM evaluation on ModelNet40, exceeding its closest competitor by 1.7% and even exceeding some supervised methods. Di Shao, Xuequan Lu, Xiao Liu 0004, Ajmal Mian |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | 3D Face Recognition with Contrastive Learning Network on Low-Quality Data
Yaping Jing, Ajmal Mian, Leo Zhang, Shang Gao 0003, Xuequan Lu |
CGI (1) | 2 |
| 2024 | Efficient Hyperparameter Optimization with Adaptive Fidelity IdentificationabstractHyperparameter Optimization and Neural Architecture Search are powerful in attaining state-of-the-art machine learning models, with Bayesian Optimization (BO) standing out as a mainstream method. Extending BO into the multi-fidelity setting has been an emerging research topic in this field, but faces the challenge of determining an appropriate fidelity for each hyperparameter configuration to fit the surrogate model. To tackle the challenge, we propose a multi-fidelity BO method named FastBO, which excels in adaptively deciding the fidelity for each configuration and providing strong performance while ensuring efficient resource usage. These advantages are achieved through our proposed techniques based on the concepts of efficient point and saturation point for each configuration, which can be obtained from the empirical learning curve of the configuration, estimated from early observations. Extensive experiments demonstrate FastBO's superior anytime performance and efficiency in identifying high-quality configurations and architectures. We also show that our method provides a way to extend any single-fidelity method to the multi-fidelity setting, highlighting the wide applicability of our approach. Jiantong Jiang, Zeyi Wen, Atif Bin Mansoor, Ajmal Mian |
CVPR | 4 |
| 2024 | L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Streamabstract3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However, the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper, we aim to handle a real-time multi-task for 6-DoF pose tracking of unknown objects, leveraging 3D-language pre-training scheme from a series of 3D point cloud video streams, while simultaneously performing 3D shape reconstruction in current observation. To this end, we present a generic Language-to-4D modeling paradigm termed L4D-Track, that tackles zero-shot 6-DoF Tracking and shape reconstruction by learning pairwise implicit 3D representation and multi-level multi-modal alignment. Our method constitutes two core parts. 1) Pairwise Implicit 3D Space Representation, that establishes spatial-temporal to language coherence descriptions across continuous 3D point cloud video. 2) Language-to-4D Association and Contrastive Alignment, enables multi-modality semantic connections between 3D point cloud video and language. Our method trained exclusively on public NOCS-REAL275 dataset, achieves promising results on both two publicly benchmarks. This not only shows powerful generalization performance, but also proves its remarkable capability in zero-shot inference. The project is released at L4D- Track. Yaonan Wang 0001, Mingtao Feng, Yulan Guo, Ajmal Mian, Zheng Shou 0001 |
CVPR | 5 |
| 2024 | External Knowledge Enhanced 3D Scene Generation from Sketch
Mingtao Feng, Yaonan Wang 0001, He Xie, Weisheng Dong, Bo Miao, Ajmal Mian |
ECCV (6) | 7 |
| 2024 | Regulating Model Reliance on Non-robust Features by Smoothing Input Marginal Density
Peiyu Yang, Naveed Akhtar, Mubarak Shah, Ajmal Mian |
ECCV (57) | 4 |
| 2024 | A Statistical Image Realism Score For Deepfake DetectionabstractRecent advances in generative visual content have led to a quantum leap in the quality of artificially generated Deepfake content. Especially, diffusion models are causing growing concerns among communities due to their ever-increasing realism. However, quantifying the realism of generated content is still challenging. Existing evaluation metrics, such as Inception Score and Fréchet inception distance, fall short on benchmarking diffusion models due to the versatility of the generated images. To address this, we propose the Image Realism Score (IRS) evaluation metric, computed from five statistical measures of a given image. This non-learning-based metric not only efficiently quantifies the realism of generated images, but it is also a viable tool for detecting if an image is real or fake. We experimentally establish the model- and data-agnostic nature of the proposed IRS by successfully detecting fake images generated by Stable Diffusion Model (SDM), Dalle2, Dalle3, Deepfloyd, Kandinsky, Midjourney and BigGAN. Yunzhuo Chen, Naveed Akhtar, Nur Al Hasan Haldar, Jordan Vice, Ajmal Mian |
ICIP | 5 |
| 2024 | Sparse Points to Dense Clouds: Enhancing 3D Detection with Limited LiDAR Dataabstract3D detection is a critical task that enables machines to identify and locate objects in three-dimensional space. It has a broad range of applications in several fields, including autonomous driving, robotics and augmented reality. Monocular 3D detection is attractive as it requires only a single camera, however, it lacks the accuracy and robustness required for real world applications. High resolution LiDAR on the other hand, can be expensive and lead to interference problems in heavy traffic given their active transmissions. We propose a balanced approach that combines the advantages of monocular and point cloud-based 3D detection. Our method requires only a small number of 3D points, that can be obtained from a low-cost, low-resolution sensor. Specifically, we use only 512 points, which is just 1% of a full LiDAR frame in the KITTI dataset. Our method reconstructs a complete 3D point cloud from this limited 3D information combined with a single image. The reconstructed 3D point cloud and corresponding image can be used by any multi-modal off-the-shelf detector for 3D object detection. By using the proposed network architecture with an off-the-shelf multi-modal 3D detector, the accuracy of 3D detection improves by 20% compared to the state-of-theart monocular detection methods and 6% to 9% compare to the baseline multi-modal methods on KITTI and JackRabbot datasets. Aakash Kumar, Chen Chen 0001, Ajmal Mian, Neils Lobo, Mubarak Shah |
IROS | 3 |
| 2024 | Referring Human Pose and Mask Estimation In the WildabstractWe introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Bo Miao, Mingtao Feng, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
NeurIPS | 6 |
| 2024 | Fast Inference for Probabilistic Graphical Models
Jiantong Jiang, Zeyi Wen, Atif Bin Mansoor, Ajmal Mian |
USENIX ATC | 4 |
| 2024 | Survey: Image mixing and deleting for data augmentation
Humza Naveed, Saeed Anwar, Munawar Hayat, Kashif Javed, Ajmal Mian |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | 3D Object Detection From Point Cloud via Voting Step Diffusionabstract3D object detection is a fundamental task in scene understanding. Numerous research efforts have been dedicated to better incorporate Hough voting into the 3D object detection pipeline. However, due to the noisy, cluttered, and partial nature of real 3D scans, existing voting-based methods often receive votes from the partial surfaces of individual objects together with severe noises, leading to sub-optimal detection performance. In this work, we focus on the distributional properties of point clouds and formulate the voting process as generating new points in the high-density region of the distribution of object centers. To achieve this, we propose a new method to move random 3D points toward the high-density region of the distribution by estimating the score function of the distribution with a noise conditioned score network. Specifically, we first generate a set of object center proposals to coarsely identify the high-density region of the object center distribution. To estimate the score function, we perturb the generated object center proposals by adding normalized Gaussian noise, and then jointly estimate the score function of all perturbed distributions. Finally, we generate new votes by moving random 3D points to the high-density region of the object center distribution according to the estimated score function. Extensive experiments on two large scale indoor 3D scene datasets, SUN RGB-D and ScanNet V2, demonstrate the superiority of our proposed method. The code will be released athttps://github.com/HHrEtvP/DiffVote. Haoran Hou, Mingtao Feng, Weisheng Dong, Qing Zhu 0003, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Temporally Consistent Referring Video Object Segmentation With Hybrid MemoryabstractReferring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects. We propose an end-to-end R-VOS paradigm that explicitly models temporal instance consistency alongside the referring segmentation. Specifically, we introduce a novel hybrid memory that facilitates inter-frame collaboration for robust spatio-temporal matching and propagation. Features of frames with automatically generated high-quality reference masks are propagated to segment the remaining frames based on multi-granularity association to achieve temporally consistent R-VOS. Furthermore, we propose a new Mask Consistency Score (MCS) metric to evaluate the temporal consistency of video segmentation. Extensive experiments demonstrate that our approach enhances temporal consistency by a significant margin, leading to top-ranked performance on popular R-VOS benchmarks, i.e., Ref-YouTube-VOS (67.1%) and Ref-DAVIS17 (65.6%). The code is available athttps://github.com/bo-miao/HTR. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Mubarak Shah, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Domain-Invariant Prototypes for Semantic SegmentationabstractDeep learning has greatly advanced the performance of semantic segmentation, however, its success relies on the availability of large amounts of annotated data for training. Hence, many efforts have been devoted to domain adaptive semantic segmentation that focuses on transferring semantic knowledge from a labeled source domain to an unlabeled target domain. Existing self-training methods typically require multiple rounds of training, while another popular framework based on adversarial training is known to be sensitive to hyper-parameters. We propose an easy-to-train framework that learns domain-invariant prototypes for domain adaptive semantic segmentation. In particular, we show that domain adaptation shares a common character with few-shot learning in that both aim to recognize some types of unseen data with knowledge learned from large amounts of seen data. Thus, we propose a unified framework for domain adaptation and few-shot learning. The core idea is to use the class prototypes extracted from few-shot annotated target images to classify pixels of both source images and target images. Our method involves only one-stage training and does not need to be trained on large-scale un-annotated target images. Moreover, our method can be extended to variants of both domain adaptation and few-shot learning. Competitive performances achieved on GTA5-to-Cityscapes and SYNTHIA-to-Cityscapes adaptation tasks show the effectiveness of the proposed novel while simple domain adaptation framework. The source code used in this paper is available at https://github.com/zgyang-hnu/DIP-hunnu. Zhengeng Yang, Hongshan Yu, Wei Sun 0028, Li Cheng 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | SCTransNet: Spatial-Channel Cross Transformer Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) has recently benefitted greatly from U-shaped neural models. However, largely overlooking effective global information modeling, existing techniques struggle when the target has high similarities with the background. We present aSpatial-channelCrossTransformerNetwork (SCTransNet) that leverages spatial-channel cross transformer blocks (SCTBs) on top of long-range skip connections to address the aforementioned challenge. In the proposed SCTBs, the outputs of all encoders are interacted with cross transformer to generate mixed features, which are redistributed to all decoders to effectively reinforce semantic differences between the target and clutter at full levels. Specifically, SCTB contains the following two key elements: (a) spatial-embedded single-head channel-cross attention (SSCA) for exchanging local spatial features and full-level global channel information to eliminate ambiguity among the encoders and facilitate high-level semantic associations of the images, and (b) a complementary feed-forward network (CFN) for enhancing the feature discriminability via a multi-scale strategy and cross-spatial-channel information interaction to promote beneficial information transfer. Our SCTransNet effectively encodes the semantic differences between targets and backgrounds to boost its internal representation for detecting small infrared targets accurately. Extensive experiments on three public datasets, NUDT-SIRST, NUAA-SIRST, and IRSTD-1K, demonstrate that the proposed SCTransNet outperforms existing IRSTD methods. Our code will be made public at https://github.com/xdFai/SCTransNet. Shuai Yuan 0013, Hanlin Qin, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Low-Rank and Sparse Decomposition for Low-Query Decision-Based Adversarial AttacksabstractDeep learning models are susceptible to contrived adversarial examples, even in the decision-based black-box setting where the attacker has access to the model’s decisions only. Developing more efficient and practical attacks help in better understanding the limitations of deep models. It is important that attacks are crafted with limited queries to avoid suspicion. Since the required number of queries increase with dimensions, low-dimensional embeddings are attractive. This low query budget constraint is a bottleneck for learning-based and data-driven attacks which rely heavily on querying the model. We propose LSDAT, an image-agnostic non-data-driven decision-based black-box attack that exploits low-rank and sparse decomposition (LSD) of images to dramatically reduce the queries and improve fooling rates compared to existing methods. LSDAT crafts perturbations in the low-dimensional subspace formed by the sparse component of the input image and that of a target adversarial image to obtain query-efficiency. A viable perturbation is obtained by traversing the path between the input and adversarial sparse components. Theoretical analyses are provided to justify the functionality of LSDAT. Unlike other competitors (e.g., FFT), LSD works directly in the image domain to guarantee that non-$\ell _{2}$constraints, such as sparsity, are satisfied. LSDAT offers better control over the number of queries and is computationally efficient as it performs sparse decomposition of the input and adversarial images only once to generate all queries. Four variants of LSDAT are presented for different scenarios including a pure black-box attack where no queries are allowed. We demonstrate$\ell _{0}$,$\ell _{2}$and$\ell _{\infty} $bounded attacks with LSDAT to evince its efficiency compared to baseline attacks in diverse low-query budget scenarios. LSDAT obtains 15 to 20% improvement in fooling ResNet-50 while using far fewer queries than competing methods in a similar setting. Ashkan Esmaeili, Marzieh Edraki, Nazanin Rahnavard, Ajmal Mian, Mubarak Shah |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | BAGM: A Backdoor Attack for Manipulating Text-to-Image Generative ModelsabstractThe rise in popularity of text-to-image generative artificial intelligence (AI) has attracted widespread public interest. We demonstrate that this technology can be attacked to generate content that subtly manipulates its users. We propose a Backdoor Attack on text-to-image Generative Models (BAGM), which upon triggering, infuses the generated images with manipulative details that are naturally blended in the content. Our attack is the first to target three popular text-to-image generative models across three stages of the generative process by modifying the behaviour of the embedded tokenizer, the language model or the image generative model. Based on the penetration level, BAGM takes the form of a suite of attacks that are referred to assurface,shallowanddeepattacks in this article. Given the existing gap within this domain, we also contribute a comprehensive set of quantitative metrics designed specifically for assessing the effectiveness of backdoor attacks on text-to-image models. The efficacy of BAGM is established by attacking state-of-the-art generative models, using a marketing scenario as the target domain. To that end, we contribute a dataset of branded product images. Our embedded backdoors increase the bias towards the target outputs by more than five times the usual, without compromising the model robustness or the generated content utility. By exposing generative AI’s vulnerabilities, we encourage researchers to tackle these challenges and practitioners to exercise caution when using pre-trained models. Relevant code and input prompts can be found at https://github.com/JJ-Vice/BAGM, and the dataset is available at: https://ieee-dataport.org/documents/marketable-foods-mf-dataset. Jordan Vice, Naveed Akhtar, Richard I. Hartley, Ajmal Mian |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Domain-Generalized Robotic Picking via Contrastive Learning-Based 6-D Pose EstimationabstractVision-guided robotic picking in 3-D space is a key technology for industrial automation and intelligent manufacturing. However, existing methods rely on labeled real-world data for learning, significantly limiting their ability to generalize to novel objects and robustness to challenging scenes containing occlusions and clutter. To address these problems, we propose a domain-generalized robotic picking method (DGPF6D) that builds on contrastive learning-based 6-D pose estimation. DGPF6D generalizes to real-world scenes by training only on synthetic data and without using shape priors. Specifically, we first perform continuous data augmentations on the synthetic RGB and point cloud images such that they can better simulate real-world scenes with occlusions and clutter. We then feed the augmented images in parallel to a two-stage (i.e., 3-D shape reconstruction and 6-D pose estimation) contrastive learning framework, thereby enhancing the domain-generalization ability and robustness of DGPF6D. Moreover, we propose a point cloud cross attention-guided intracategory unknown object 3-D shape reconstruction network, which can effectively fuse the observed and the unit random point clouds and explicitly highlight their differences, thus avoiding the dependence of DGPF6D on shape priors. Finally, we build a robotic picking system employing DGPF6D to realize domain-generalized robotic picking in 3-D space. Extensive experiments on two benchmarks and real-world scenes show that DGPF6D achieves state-of-the-art performance, and can be effectively applied for domain-generalized robotic picking. Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Ajmal Mian |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | A Systematic Point Cloud Edge Detection Framework for Automatic Aircraft Skin MillingabstractThe edge detection technique is an essential step for aircraft skin milling in aviation manufacturing. Most of the current detection methods focus on traditionally defined edge extraction tasks but disregard the crucial systematic requirement of edge milling. In this article, we proposed a novel edge detection framework for automatic edge milling of aircraft skins. First, an edge probability detector is proposed by the spatial tangent continuity to provide the essential reference. Second, we propose a hierarchical branch searching method to hierarchically strip the desired milling edges from the raw point cloud, which consists of the following three graded progressive steps: branch backbone generation, branch extension, and branch pruning. We demonstrate the performance of the proposed method on both synthetic models and aircraft skin workpieces. The proposed method outperforms the other baselines and shows accurate edges for the edge milling task. Yaonan Wang 0001, He Xie, Mingtao Feng, Haotian Wu 0002, Chao Ding 0006, Ajmal Mian |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | PoseDiffusion: A Coarse-to-Fine Framework for Unseen Object 6-DoF Pose EstimationabstractAccurately estimating the six-degrees of freedom (DoF) pose of unseen objects is crucial for successful robotic manipulation in industrial automation. Some existing methods for this task rely on prior knowledge of individual objects, i.e., the model must be trained on the exact object instance or object category. Others perform unseen object pose estimation but are limited in their feature learning and pose refinement ability. To address these problems, we propose an unseen object pose estimation method that follows a coarse-to-fine framework and leverages the powerful learning ability of diffusion models. We introduce a diffusion model for generating object poses, and conduct a comparison between the generated poses and the original pose to determine the optimal one. We design a novel pose estimation module to provide coarse poses for the PoseDiffusion. This module comprises two feature extraction modules that extract global and masked features. In addition, we propose a strategy to estimate the pose by comparing the similarity between rendered and query poses. The renderings of an unseen object from various viewpoints are generated from its computer-aided design (CAD) model. Our method requires a CAD model of the unseen object only during inference, a scenario well suited to industrial applications. Experimental evaluation on benchmark datasets demonstrates that the proposed framework outperforms existing approaches, achieving state-of-the-art performance in six-DoF object pose estimation. Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Chengzhong Wu, Xuebing Liu, Jianan Huang 0002, Ajmal Mian |
IEEE Trans. Ind. Informatics | 8 |
| 2024 | Region Aware Video Object Segmentation With Deep Motion ModelingabstractCurrent semi-supervised video object segmentation (VOS) methods often employ the entire features of one frame to predict object masks and update memory. This introduces significant redundant computations. To reduce redundancy, we introduce a Region Aware Video Object Segmentation (RAVOS) approach, which predicts regions of interest (ROIs) for efficient object segmentation and memory storage. RAVOS includes a fast object motion tracker to predict object ROIs in the next frame. For efficient segmentation, object features are extracted based on the ROIs, and an object decoder is designed for object-level segmentation. For efficient memory storage, we propose motion path memory to filter out redundant context by memorizing the features within the motion path of objects. In addition to RAVOS, we also propose a large-scale occluded VOS dataset, dubbed OVOS, to benchmark the performance of VOS models under occlusions. Evaluation on DAVIS and YouTube-VOS benchmarks and our new OVOS dataset show that our method achieves state-of-the-art performance with significantly faster inference time, e.g., 86.1 J & F at 42 FPS on DAVIS and 84.4 J & F at 23 FPS on YouTube-VOS. Project page: ravos.netlify.app. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
IEEE Trans. Image Process. | 4 |
| 2024 | Unsupervised Learning of Intrinsic Semantics With Diffusion Model for Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to learn semantic representations for person retrieval without using identity labels. Most existing methods generate fine-grained patch features to reduce noise in global feature clustering. However, these methods often compromise the discriminative semantic structure and overlook the semantic consistency between the patch and global features. To address these problems, we propose a Person Intrinsic Semantic Learning (PISL) framework with diffusion model for unsupervised person Re-ID. First, we design the Spatial Diffusion Model (SDM), which performs a denoising diffusion process from noisy spatial transformer parameters to semantic parameters, enabling the sampling of patches with intrinsic semantic structure. Second, we propose the Semantic Controlled Diffusion (SCD) loss to guide the denoising direction of the diffusion model, facilitating the generation of semantic patches. Third, we propose the Patch Semantic Consistency (PSC) loss to capture semantic consistency between the patch and global features, refining the pseudo-labels of global features. Comprehensive experiments on three challenging datasets show that our method surpasses current unsupervised Re-ID methods. The source code will be publicly available at https://github.com/taoxuefong/Diffusion-reid. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Ming Lu 0008, Ajmal Mian |
IEEE Trans. Image Process. | 5 |
| 2024 | Yolo-3DMM for Simultaneous Multiple Object Detection and Tracking in Traffic ScenariosabstractVideo-based multiple object tracking (MOT) is a fundamental task in intelligent transportation with applications ranging from automated traffic surveillance to autonomous driving. MOT methods commonly follow a tracking-by-detection paradigm, tracking objects by associating their detections across video frames. However, insofar, these methods have not used the entire vehicle trajectory motion characteristics to perform tracking, which converts the vehicle localization problem into a motion parameter estimation problem. Moreover, MOT methods mainly rely on off-the-shelf detectors. An independently trained detector is sub-optimal for the tracking-by-detection paradigm and adversely affects the overall system performance. In this article, we address these issues by proposing a novel MOT method for moving vehicles in traffic scenarios. Our tracker treats the vehicle tracks as unified 3D spatio-temporal trajectory instances and leverages the power of deep learning to extract vehicle motion from the 3D instances. We propose a new simultaneous detection and tracking network, called YOLO-3D Motion Model Network (Yolo-3DMM) that employs spatio-temporal features of traffic videos for simultaneous vehicle detection and tracking in an end-to-end manner. We adopt a variety of different vehicle tracking datasets to evaluate our method. Moreover, we also propose a tunnel MOT dataset from real highway tunnel surveillance in Guangdong, China to expand the experimental scenarios. To establish the efficacy of our method, we evaluate it on 100 different roadside traffic scenarios. Our method shows excellent performance on UA-DETRAC and Omni-MOT datasets. It achieves a PR-MOTA score of 29.40% on UA-DETRAC and gets a 69.7% MOTA score on the Omni-MOT dataset. Lichen Liu, Huansheng Song, Shijie Sun 0001, Xian-Feng Han, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Co-Engaged Location Group Search in Location-Based Social NetworksabstractSearching for well-connected user communities in a Location-based Social Network (LBSN) has been extensively investigated. However, very few studies focus on finding a group of locations in an LBSN which are significantly engaged with socially cohesive user groups. In this work, we investigate the problem ofCo-engagedLocation groupSearch (CLS) from LBSNs where the selected locations are visited frequently by the members of the socially cohesive user groups, and the locations are reachable within a given distance threshold. To the best of our knowledge, this is the first work to search for socially co-engaged location groups in LBSNs. We devise a score function to measure the co-engagement of the location groups by combining social connectivity of the cohesive user groups and check-in density of the users to the selected locations. To solve theCLSproblem, we propose aFilter-and-Verifyalgorithm that effectively filters out ineligible locations, and their corresponding check-in users. Further, we derive a lower bound on the number of check-ins to prune the insignificant locations and develop a novel greedy forward expansion algorithm (GFA). To accelerate the computation ofCLS, we propose a ranking function and devise an incremental algorithm,GIA, that can filter the unqualified location groups. We establish the effectiveness of our solutions by conducting extensive experiments on three real-world datasets. Nur Al Hasan Haldar, Jianxin Li 0001, Naveed Akhtar, Yan Jia 0001, Ajmal Mian |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Mesh Convolution With Continuous Filters for 3-D Surface ParsingabstractGeometric feature learning for 3-D surfaces is critical for many applications in computer graphics and 3-D vision. However, deep learning currently lags in hierarchical modeling of 3-D surfaces due to the lack of required operations and/or their efficient implementations. In this article, we propose a series of modular operations for effective geometric feature learning from 3-D triangle meshes. These operations include novel mesh convolutions, efficient mesh decimation, and associated mesh (un)poolings. Our mesh convolutions exploit spherical harmonics as orthonormal bases to create continuous convolutional filters. The mesh decimation module is graphics processing unit (GPU)-accelerated and able to process batched meshes on-the-fly, while the (un)pooling operations compute features for upsampled/downsampled meshes. We provide an open-source implementation of these operations, collectively termed Picasso. Picasso supports heterogeneous mesh batching and processing. Leveraging its modular operations, we further contribute a novel hierarchical neural network for perceptual parsing of 3-D surfaces, named PicassoNet++. It achieves highly competitive performance for shape analysis and scene segmentation on prominent 3-D benchmarks. The code, data, and trained models are available at https://github.com/EnyaHermite/Picasso. Huan Lei, Naveed Akhtar, Mubarak Shah, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Fully Convolutional Network-Based Self-Supervised Learning for Semantic SegmentationabstractAlthough deep learning has achieved great success in many computer vision tasks, its performance relies on the availability of large datasets with densely annotated samples. Such datasets are difficult and expensive to obtain. In this article, we focus on the problem of learning representation from unlabeled data for semantic segmentation. Inspired by two patch-based methods, we develop a novel self-supervised learning framework by formulating the jigsaw puzzle problem as a patch-wise classification problem and solving it with a fully convolutional network. By learning to solve a jigsaw puzzle comprising 25 patches and transferring the learned features to semantic segmentation task, we achieve a 5.8% point improvement on the Cityscapes dataset over the baseline model initialized from random values. It is noted that we use only about 1/6 training images of Cityscapes in our experiment, which is designed to imitate the real cases where fully annotated images are usually limited to a small number. We also show that our self-supervised learning method can be applied to different datasets and models. In particular, we achieved competitive performance with the state-of-the-art methods on the PASCAL VOC2012 dataset using significantly fewer time costs on pretraining. Zhengeng Yang, Hongshan Yu, Yong He 0012, Wei Sun 0028, Zhi-Hong Mao, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Faster-BNI: Fast Parallel Exact Inference on Bayesian NetworksabstractBayesian networks (BNs) have recently attracted more attention, because they are interpretable machine learning models and enable a direct representation of causal relations between variables. However, exact inference on BNs is time-consuming, especially for complex problems, which hinders the widespread adoption of BNs. To improve the efficiency, we propose a fast BN exact inference named Faster-BNI on multi-core CPUs. Faster-BNI enhances the efficiency of a well-known BN exact inference algorithm, namely the junction tree algorithm, through hybrid parallelism that tightly integrates coarse- and fine-grained parallelism. Moreover, we identify that the bottleneck of BN exact inference methods lies in recursively updating the potential tables of the network. To reduce the table update cost, Faster-BNI employs novel optimizations, including the reduction of potential tables and re-organizing the potential table storage, to avoid unnecessary memory consumption and simplify potential table operations. Comprehensive experiments on real-world BNs show that the sequential version of Faster-BNI outperforms existing sequential implementation by 9 to 22 times, and the parallel version of Faster-BNI achieves up to 11 times faster inference than its parallel counterparts. Jiantong Jiang, Zeyi Wen, Atif Bin Mansoor, Ajmal Mian |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Parallel and Distributed Bayesian Network Structure LearningabstractBayesian networks (BNs) are graphical models representing uncertainty in causal discovery, and have been widely used in medical diagnosis and gene analysis due to their effectiveness and good interpretability. However, mainstream BN structure learning methods are computationally expensive, as they must perform numerous conditional independence (CI) tests to decide the existence of edges. Some researchers attempt to accelerate the learning process by parallelism, but face issues including load unbalancing, costly dominant parallelism overhead. We propose a multi-thread method, namely Fast-BNS version 1 (Fast-BNS-v1 for short), on multi-core CPUs to enhance the efficiency of the BN structure learning. Fast-BNS-v1 incorporates a series of efficiency optimizations, including a dynamic work pool for better scheduling, grouping CI tests to avoid unnecessary operations, a cache-friendly data storage to improve memory efficiency, and on-the-fly conditioning sets generation to avoid extra memory consumption. To further boost learning performance, we develop a two-level parallel method Fast-BNS-v2 by integrating edge-level parallelism with multi-processes and CI-level parallelism with multi-threads. Fast-BNS-v2 is equipped with careful optimizations including dynamic work stealing for load balancing, SIMD edge list deletion for list updating, and effective communication policies for synchronization. Comprehensive experiments show that our Fast-BNS achieves 9 to 235 times speedup over the state-of-the-art multi-threaded method on a single machine. When running on multi-machines, it further reduces the execution time of the single-machine implementation by 80%. Jiantong Jiang, Zeyi Wen, Ajmal Mian |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Contrastive Self-Supervised Learning Leads to Higher Adversarial SusceptibilityabstractContrastive self-supervised learning (CSL) has managed to match or surpass the performance of supervised learning in image and video classification. However, it is still largely unknown if the nature of the representations induced by the two learning paradigms is similar. We investigate this under the lens of adversarial robustness. Our analysis of the problem reveals that CSL has intrinsically higher sensitivity to perturbations over supervised learning. We identify the uniform distribution of data representation over a unit hypersphere in the CSL representation space as the key contributor to this phenomenon. We establish that this is a result of the presence of false negative pairs in the training process, which increases model sensitivity to input perturbations. Our finding is supported by extensive experiments for image and video classification using adversarial perturbations and other input corruptions. We devise a strategy to detect and remove false negative pairs that is simple, yet effective in improving model robustness with CSL training. We close up to 68% of the robustness gap between CSL and its supervised counterpart. Finally, we contribute to adversarial learning by incorporating our method in CSL. We demonstrate an average gain of about 5% over two different state-of-the-art methods in this domain. Rohit Gupta 0012, Naveed Akhtar, Ajmal Mian, Mubarak Shah |
AAAI | 3 |
| 2023 | Local Path Integration for AttributionabstractPath attribution methods are a popular tool to interpret a visual model's prediction on an input. They integrate model gradients for the input features over a path defined between the input and a reference, thereby satisfying certain desirable theoretical properties. However, their reliability hinges on the choice of the reference. Moreover, they do not exhibit weak dependence on the input, which leads to counter-intuitive feature attribution mapping. We show that path-based attribution can account for the weak dependence property by choosing the reference from the local distribution of the input. We devise a method to identify the local input distribution and propose a technique to stochastically integrate the model gradients over the paths defined by the references sampled from that distribution. Our local path integration (LPI) method is found to consistently outperform existing path attribution techniques when evaluated on deep visual models. Contributing to the ongoing search of reliable evaluation metrics for the interpretation methods, we also introduce DiffID metric that uses the relative difference between insertion and deletion games to alleviate the distribution shift problem faced by existing metrics. Our code is available at https://github.com/ypeiyu/LPI. Peiyu Yang, Naveed Akhtar, Zeyi Wen, Ajmal Mian |
AAAI | 4 |
| 2023 | 3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point CloudabstractIn-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide variety of challenging relationships, existing methods perform quite poorly with limited training samples. In this work, we find that the inherently hierarchical structures of physical space in 3D scenes aid in the automatic association of semantic and spatial arrangements, specifying clear patterns and leading to less ambiguous predictions. Thus, they well meet the challenges due to the rich variations within scene categories. To achieve this, we explicitly unify these structural cues of 3D physical spaces into deep neural networks to facilitate scene graph prediction. Specifically, we exploit an external knowledge base as a baseline to accumulate both contextualized visual content and textual facts to form a 3D spatial multimodal knowledge graph. Moreover, we propose a knowledge-enabled scene graph prediction module benefiting from the 3D spatial knowledge to effectively regularize semantic space of relationships. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art competitors. Our code is available at https://github.com/HHrEtvP/SMKA. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Ajmal Mian |
CVPR | 6 |
| 2023 | Spectrum-guided Multi-granularity Referring Video Object SegmentationabstractCurrent referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perceive during the forward computation. This negatively affects the ability of segmentation kernels. To address the drift problem, we propose a Spectrum-guided Multi-granularity (SgMg) approach, which performs direct segmentation on the encoded features and employs visual details to further optimize the masks. In addition, we propose Spectrum-guided Cross-modal Fusion (SCF) to perform intra-frame global interactions in the spectral domain for effective multimodal representation. Finally, we extend SgMg to perform multi-object R-VOS, a new paradigm that enables simultaneous segmentation of multiple referred objects in a video. This not only makes R-VOS faster, but also more practical. Extensive experiments show that SgMg achieves state-of-the-art performance on four video benchmark datasets, outperforming the nearest competitor by 2.8% points on Ref-YouTube-VOS. Our extended SgMg enables multi-object R-VOS, runs about 3 faster while maintaining satisfactory performance. Code×is available at https://github.com/bo-miao/SgMg. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
ICCV | 4 |
| 2023 | Sketch and Text Guided Diffusion Model for Colored Point Cloud GenerationabstractDiffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently ambiguous and lack details. In this paper, we propose a sketch and text guided probabilistic diffusion model for colored point cloud generation that conditions the denoising process jointly with a hand drawn sketch of the object and its textual description. We incrementally diffuse the point coordinates and color values in a joint diffusion process to reach a Gaussian distribution. Colored point cloud generation thus amounts to learning the reverse diffusion process, conditioned by the sketch and text, to iteratively recover the desired shape and color. Specifically, to learn effective sketch-text embedding, our model adaptively aggregates the joint embedding of text prompt and the sketch based on a capsule attention network. Our model uses staged diffusion to generate the shape and then assign colors to different parts conditioned on the appearance prompt while preserving precise shapes from the first stage. This gives our model the flexibility to extend to multiple tasks, such as appearance re-editing and part segmentation. Experimental results demonstrate that our model outperforms recent state-of-the-art in point cloud generation. Yaonan Wang 0001, Mingtao Feng, He Xie, Ajmal Mian |
ICCV | 5 |
| 2023 | Dual Student Networks for Data-Free Model Stealing
James Beetham, Navid Kardan, Ajmal Mian, Mubarak Shah |
ICLR | 3 |
| 2023 | Re-calibrating Feature Attributions for Model Interpretation
Peiyu Yang, Naveed Akhtar, Zeyi Wen, Mubarak Shah, Ajmal Mian |
ICLR | 5 |
| 2023 | Slice Transformer and Self-supervised Learning for 6DoF Localization in 3D Point Cloud MapsabstractPrecise localization is critical for autonomous vehicles. We present a self-supervised learning method that employs transformers for the first time for the task of outdoor localization using LiDAR data. We propose a pre-text task that reorganizes the slices of a 360° LiDAR scan to leverage its axial properties. Our model, called Slice Transformer, employs multi-head attention while systematically processing the slices. To the best of our knowledge, this is the first instance of leveraging multi-head attention for outdoor point clouds. We additionally introduce the Perth-Wadataset, which provides a large-scale LiDAR map of Perth city in Western Australia, covering ~4km2area. Localization annotations are provided for Perth - Wa.The proposed localization method is thoroughly evaluated on Perth-WA and Appollo-SouthBay datasets. We also establish the efficacy of our self-supervised learning approach for the common downstream task of object classification using ModelNet40 and ScanNN datasets. The code and Perth-WA data will be publicly released. Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Michael J. Wise, Ajmal Mian |
ICRA | 5 |
| 2023 | 3DMODT: Attention-Guided Affinities for Joint Detection & Tracking in 3D Point CloudsabstractWe propose a method for joint detection and tracking of multiple objects in 3D point clouds, a task conventionally treated as a two-step process comprising object detection followed by data association. Our method embeds both steps into a single end-to-end trainable network eliminating the dependency on external object detectors. Our model exploits temporal information employing multiple frames to detect objects and track them in a single network, thereby making it a utilitarian formulation for real-world scenarios. Computing affinity matrix by employing features similarity across consecutive point cloud scans forms an integral part of visual tracking. We propose an attention-based refinement module to refine the affinity matrix by suppressing erroneous correspondences. The module is designed to capture the global context in affinity matrix by employing self-attention within each affinity matrix and cross-attention across a pair of affinity matrices. Unlike competing approaches, our network does not require complex post-processing algorithms, and directly processes raw LiDAR frames to output tracking results. We demonstrate the effectiveness of our method on three tracking benchmarks: JRDB, Waymo, and KITTI. Experimental evaluations indicate the ability of our model to generalize well across datasets. Jyoti Kini, Ajmal Mian, Mubarak Shah |
ICRA | 2 |
| 2023 | UnLoc: A Universal Localization Method for Autonomous Vehicles using LiDAR, Radar and/or Camera InputabstractLocalization is a fundamental task in robotics for autonomous navigation. Existing localization methods rely on a single input data modality or train several computational models to process different modalities. This leads to stringent computational requirements and sub-optimal results that fail to capitalize on the complementary information in other data streams. This paper proposes UnLoc, a novel unified neural modeling approach for localization with multi-sensor input in all weather conditions. Our multi-stream network can handle LiDAR, Camera and RADAR inputs for localization on demand, i.e., it can work with one or more input sensors, making it robust to sensor failure. UnLoc uses 3D sparse convolutions and cylindrical partitioning of the space to process LiDAR frames and implements ResNet blocks with a slot attention-based feature filtering module for the Radar and image modalities. We introduce a unique learnable modality encoding scheme to distinguish between the input sensor data. Our method is extensively evaluated on Oxford Radar RobotCar, ApolloSouthBay and Perth-WA datasets. The results ascertain the efficacy of our technique. The dataset, results, and codes are available at https://github.com/IbrahimUWA/UnLoc Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Ajmal Mian |
IROS | 4 |
| 2023 | Fast Parallel Exact Inference on Bayesian NetworksabstractBayesian networks (BNs) are attractive, because they are graphical and interpretable machine learning models. However, exact inference on BNs is time-consuming, especially for complex problems. To improve the efficiency, we propose a fast BN exact inference solution named Fast-BNI on multi-core CPUs. Fast-BNI enhances the efficiency of exact inference through hybrid parallelism that tightly integrates coarse- and fine-grained parallelism. We also propose techniques to further simplify the bottleneck operations of BN exact inference. Fast-BNI source code is freely available at https://github.com/jjiantong/FastBN. Jiantong Jiang, Zeyi Wen, Atif Bin Mansoor, Ajmal Mian |
PPoPP | 4 |
| 2023 | Entropy-based Selective Homomorphic Encryption for Smart Metering SystemsabstractSmart metering systems (SMS) are popular in industrial and residential areas but can risk privacy by revealing user behaviors. Homomorphic encryption (HE) is a technique that protects data privacy by enabling calculations on encrypted data. However, the high computational costs of HE can hinder real-time or resource-constrained applications. Our paper presents a framework to encrypt only selected SMS data for improved efficiency without compromising privacy significantly. By encrypting data blocks with higher entropy values, we can mitigate the leakage of key information to adversaries who may conduct privacy attacks, such as membership inference attacks (MIA). We evaluate our framework using two real-world datasets (i.e., electricity and water) to assess privacy and performance trade-offs. The results indicate a 47% performance increase while still providing a sufficient level of privacy when adopting an encryption ratio of 0.4. They demonstrate the effectiveness of the proposed framework, considering the trade-off between privacy and performance, where the user can determine the appropriate security level for privacy protection and enhance the performance using HE in practical SMS settings. Weiyan Xu, Rachel Cardell-Oliver, Ajmal Mian, Jin B. Hong |
PRDC | 4 |
| 2023 | Robust image clustering via context-aware contrastive graph learning
Uno Fang, Jianxin Li 0001, Xuequan Lu, Ajmal Mian, Zhaoquan Gu |
Pattern Recognit. | 4 |
| 2023 | Mesh-Based DGCNN: Semantic Segmentation of Textured 3-D Urban ScenesabstractTextured 3D mesh is one of the final user products in photogrammetry and remote sensing. However, research on the semantic segmentation of complex urban scenes represented by textured 3D meshes is in its infancy. We present a mesh-based dynamic graph CNN (DGCNN) for the semantic segmentation of textured 3D meshes. To represent each mesh facet, composite input feature vectors are constructed by concatenating the face-inherent features, i.e., XYZ coordinates of the center of gravity (CoG), texture values, and normal vectors. A texture fusion module is embedded into the proposed mesh-based DGCNN to generate high-level semantic features of the high-resolution texture information, which is useful for semantic segmentation. We achieve competitive accuracies when the proposed method is applied to the SUM mesh datasets. The overall accuracy (OA), Kappa coefficient (Kap), mean precision (mP), mean recall (mR), mean F1 score (mF1), and mean intersection over union (mIoU) are 93.3%, 88.7%, 79.6%, 83.0%, 80.7%, and 69.6%, respectively. In particular, the OA, mean class accuracy (mAcc), mIoU, and mF1 increase by 0.3%, 12.4%, 3.4%, and 6.9%, respectively, compared to the state-of-the-art method. Guangyun Zhang, Jihao Yin, Xiuping Jia, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Language Model Agnostic Gray-Box Adversarial Attack on Image CaptioningabstractAdversarial susceptibility of neural image captioning is still under-explored due to the complex multi-model nature of the task. We introduce a GAN-based adversarial attack to effectively fool encoder-decoder based image captioning frameworks. Unique to our attack is the systematic disruption of the internal representation of an image at the encoder stage which allows control over the captions generated at the decoder stage. We cause the desired disruption with an input perturbation that promotes similarity between the features of the input image with a target image of our choice. The target image provides a convenient handle to control the incorrect captions in our method. We do not assume any knowledge of the decoder module, which makes our attack ‘gray-box’. Moreover, our attack remains agnostic to the type of decoder module, thereby proving effective for RNNs as well as Transformers as the language models. This makes our attack highly pragmatic. Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Mubarak Shah, Ajmal Mian |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | SDFA: Structure-Aware Discriminative Feature Aggregation for Efficient Human Fall Detection in VideoabstractOlder people are susceptible to fall due to instability in posture and deteriorating health. Immediate access to medical support can greatly reduce repercussions. Hence, there is an increasing interest in automated fall detection, often incorporated into a smart health-care system to provide better monitoring. Existing systems focus on wearable devices that are inconvenient or video monitoring that has privacy concerns. Moreover, these systems provide a limited perspective of their generalization ability as they are tested on datasets containing few activities that have wide disparity in the action space and are easy to differentiate. Complex daily life scenarios pose much greater challenges with activities that overlap in action spaces due to similar posture or motion. To overcome these limitations, we propose a fall detection model, called structure-aware discriminative feature aggregation, based on human skeletons extracted from low-resolution videos. The use of skeleton data ensures privacy and low-resolution videos ensures low hardware and computational cost. Our model captures discriminative structural displacements and motion trends using unified joint and motion features projected onto a shared high-dimensional space. Particularly, the use of separable convolution combined with a powerful graph convolutional network architecture provides improved performance. Extensive experiments on five large-scale datasets with a wide range of evaluation settings show that our model achieves competitive performance with extremely low computational complexity and runs faster than existing models. Sania Zahan, Ghulam M. Hassan, Ajmal Mian |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | SAT3D: Slot Attention Transformer for 3D Point Cloud Semantic SegmentationabstractSemantic segmentation of 3D point cloud is a key task in numerous intelligent transportation system applications, e.g., self-driving vehicles, traffic monitoring. Due to the sparsity and varying density of points in the outdoor point clouds, it becomes particularly challenging to extract object-centric features from data. This leads to poor semantic segmentation, especially for the rare object classes. To address that, we introduce the first-ever Slot Attention Transformer based technique to effectively model object-centric features in point cloud data. Our method uses cylindrical splits of space for voxelization and computes channel-wise positional embeddings before repetitively encoding the point cloud with slot attentions. Our second major contribution is a Large-Scale Outdoor Point Cloud dataset (SWAN), collected in a dense urban environment, driving 150km distance. It provides 16 billion points in more than 200K frames. The dataset also provides annotations for 10K frames for 24 classes. We also contribute a data augmentation scheme to handle rare object classes in real-world point clouds. Besides benchmarking popular existing methods on SWAN for the first time, we thoroughly evaluate our technique on the existing large-scale datasets, Semantic KITTI and nuScenes. Our results demonstrate a consistent performance gain for our technique, and verify the need of the more challenging SWAN dataset. Muhammad Ibrahim 0001, Naveed Akhtar, Saeed Anwar, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Dense Video Captioning With Early Linguistic Information FusionabstractDense captioning methods generally detect events in videos first and then generate captions for the individual events. Events are localized solely based on the visual cues while ignoring the associated linguistic information and context. Whereas end-to-end learning may implicitly take guidance from language, these methods still fall short of the power of explicit modeling. In this paper, we propose aVisual-Semantic Embedding (ViSE) Frameworkthat models the word(s)-context distributional properties over the entire semantic space and computes weights for all then-gramssuch that higher weights are assigned to the more informativen-grams. The weights are accounted for in learning distributed representations of all the captions to construct a semantic space. To perform the contextualization of visual information and the constructed semantic space in a supervised manner, we designVisual-Semantic Joint Modeling Network (VSJM-Net). The learnedViSEembeddings are then temporally encoded with aHierarchical Descriptor Transformer (HDT). For caption generation, we exploit a transformer architecture to decode the input embeddings into natural language descriptions. Experiments on the large-scale ActivityNet Captions dataset and YouCook-II dataset demonstrate the efficacy of our method. Nayyer Aafaq, Ajmal Mian, Naveed Akhtar, Wei Liu 0006, Mubarak Shah |
IEEE Trans. Multim. | 2 |
| 2023 | DualConv: Dual Convolutional Kernels for Lightweight Deep Neural NetworksabstractConvolutional neural network (CNN) architectures are generally heavy on memory and computational requirements which make them infeasible for embedded systems with limited hardware resources. We propose dual convolutional kernels (DualConv) for constructing lightweight deep neural networks. DualConv combines 3×3 and 1×1 convolutional kernels to process the same input feature map channels simultaneously and exploits the group convolution technique to efficiently arrange convolutional filters. DualConv can be employed in any CNN model such as VGG-16 and ResNet-50 for image classification, you only look once (YOLO) and R-CNN for object detection, or fully convolutional network (FCN) for semantic segmentation. In this work, we extensively test DualConv for classification since these network architectures form the backbone for many other tasks. We also test DualConv for image detection on YOLO-V3. Experimental results show that, combined with our structural innovations, DualConv significantly reduces the computational cost and number of parameters of deep neural networks while surprisingly achieving slightly higher accuracy than the original models in some cases. We use DualConv to further reduce the number of parameters of the lightweight MobileNetV2 by 54% with only 0.68% drop in accuracy on CIFAR-100 dataset. When the number of parameters is not an issue, DualConv increases the accuracy of MobileNetV1 by 4.11% on the same dataset. Furthermore, DualConv significantly improves the YOLO-V3 object detection speed and improves its accuracy by 4.4% on PASCAL visual object classes (VOC) dataset. Jiachen Zhong, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | UTB180: A High-Quality Benchmark for Underwater Tracking
Basit Alawode, Mehnaz Ummar, Naoufel Werghi, Jorge Dias 0001, Ajmal Mian, Sajid Javed |
ACCV (5) | 6 |
| 2022 | Learning from Pixel-Level Noisy Label : A New Perspective for Light Field Saliency DetectionabstractSaliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obtained from unsupervised hand crafted featured-based saliency methods. Given this goal, a natural question is: can we efficiently incorporate the relationships among light field cues while identifying clean labels in a unified framework? We address this question by formulating the learning as a joint optimization of intra light field features fusion stream and inter scenes correlation stream to generate the predictions. Specially, we first introduce a pixel forgetting guided fusion module to mutually enhance the light field features and exploit pixel consistency across iterations to identify noisy pixels. Next, we introduce a cross scene noise penalty loss for better reflecting latent structures of training data and enabling the learning to be invariant to noise. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our framework showing that it learns saliency prediction comparable to state-of-the-art fully supervised light field saliency methods. Our code is available at h t tps://github.com/ OLobbCode/NoiseLF. Mingtao Feng, Kendong Liu, Liang Zhang 0010, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
CVPR | 6 |
| 2022 | UNICON: Combating Label Noise Through Uniform Selection and Contrastive LearningabstractSupervised deep learning methods require a large repository of annotated data; hence, label noise is inevitable. Training with such noisy data negatively impacts the generalization performance of deep neural networks. To combat label noise, recent state-of-the-art methods employ some sort of sample selection mechanism to select a possibly clean subset of data. Next, an off-the-shelf semi-supervised learning method is used for training where rejected samples are treated as unlabeled data. Our comprehensive analysis shows that current selection methods disproportionately select samples from easy (fast learnable) classes while rejecting those from relatively harder ones. This creates class imbalance in the selected clean set and in turn, deteriorates performance under high label noise. In this work, we propose UNICON, a simple yet effective sample selection method which is robust to high label noise. To address the disproportionate selection of easy and hard samples, we introduce a Jensen-Shannon divergence based uniform selection mechanism which does not require any probabilistic modeling and hyperparameter tuning. We complement our selection method with contrastive learning to further combat the memorization of noisy labels. Extensive experimentation on multiple benchmark datasets demonstrates the effectiveness of UNICON; we obtain an 11.4% improvement over the current state-of-the-art on CIFAR100 dataset with a 90% noise rate. Our code is publicly available.11https://github.com/nazmul-karim170/UNICON-Noisy-Label Nazmul Karim, Mamshad Nayeem Rizve, Nazanin Rahnavard, Ajmal Mian, Mubarak Shah |
CVPR | 4 |
| 2022 | Self-Supervised Video Object Segmentation by Motion-Aware Mask PropagationabstractWe propose a self-supervised spatio-temporal matching method, coined Motion-Aware Mask Propagation (MAMP), for video object segmentation. MAMP leverages the frame reconstruction task for training without the need for annotations. During inference, MAMP builds a dynamic memory bank and propagates masks according to our proposed motion-aware spatio-temporal matching module, which is able to handle fast motion and long-term matching scenarios. Evaluation on DAVIS-2017 and YouTube-VOS datasets show that MAMP achieves state-of-the-art performance with stronger generalization ability compared to existing self-supervised methods, i.e., 4.2% higher mean$\mathcal{J}$&$\mathcal{F}$on DAVIS-2017 and 4.85% higher mean$\mathcal{J}$&$\mathcal{F}$on the unseen categories of YouTube-VOS than the nearest competitor. Moreover, MAMP performs at par with many supervised video object segmentation methods. Our code is available at: https://github.com/bo-miao/MAMP. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
ICME | 4 |
| 2022 | Detecting Compromised Architecture/Weights of a Deep ModelabstractAdversarial attacks perturb data to modify a model’s prediction. These perturbations can be crafted in a white-box or black-box setting, depending on whether the target model architecture/weights are known or unknown. Compromised architecture and weights of a model makes it vulnerable to the more powerful white-box attacks. In this work, we determine if a deep model is compromised by distinguishing white-box from black-box adversarial attacks. The proposed method utilizes the internal representations of the target model and a proxy model to increase the detector efficacy. Additionally, it employs a spatial smoothing module to control the strength of white-box attacks relative to black-box attacks, and a proxy module to aid in measuring the transferability of the attack. Both modules work in tandem to increase the contrast of the internal representations between white-box and black-box attacks for better discrimination. We perform a detailed ablation of our method to showcase the importance of the different modules, and show that the spatial smoothing and proxy defense techniques enable our framework to significantly outperform the simple classification baseline on common vision datasets. James Beetham, Navid Kardan, Ajmal Mian, Mubarak Shah |
ICPR | 3 |
| 2022 | Multi-Grained Interpre table Network for Image RecognitionabstractGiven a classification problem with a large number of classes, humans often compare features at different granularities from coarse to fine to gradually recognize an object. However, current deep models are generally trained to directly make the final prediction, focusing on improving the ability of the network to extract features without considering the interpretability of the model. In this paper, we propose a multi-grained interpretable network to imitate the reasoning process of humans. The proposed network is equipped with techniques to assign images with multi-grained labels, so as to train a tree-structured classifier that learns features at different levels of granularity. The proposed method can hierarchically classify objects in images at different granularities, while providing a decision pathway with multi-grained explanations for practitioners. Experimental results demonstrate that our method achieves competitive prediction accuracy on CUB-200-2011 and Stanford Cars datasets, and simultaneously produces high-quality explanations of its decisions. Moreover, our method shows higher robustness of the learned features to adversarial examples generated by the FGSM and PGD attacks. Peiyu Yang, Zeyi Wen, Ajmal Mian |
ICPR | 3 |
| 2022 | Fast Parallel Bayesian Network Structure LearningabstractBayesian networks (BNs) are a widely used graphical model in machine learning for representing knowledge with uncertainty. The mainstream BN structure learning methods require performing a large number of conditional independence (CI) tests. The learning process is very time-consuming, especially for high-dimensional problems, which hinders the adoption of BNs to more applications. Existing works attempt to accelerate the learning process with parallelism, but face issues including load unbalancing, costly atomic operations and dominant parallel overhead. In this paper, we propose a fast solution named Fast-BNS on multi-core CPUs to enhance the efficiency of the BN structure learning. Fast-Bns is powered by a series of efficiency optimizations including (i) designing a dynamic work pool to monitor the processing of edges and to better schedule the workloads among threads, (ii) grouping the CI tests of the edges with the same endpoints to reduce the number of unnecessary CI tests, (iii) using a cache-friendly data storage to improve the memory efficiency, and (iv) generating the conditioning sets on-the-fly to avoid extra memory consumption. A comprehensive experimental study shows that the sequential version of Fast-BNS is up to 50 times faster than its counterpart, and the parallel version of Fast-Bns achieves 4.8 to 24.5 times speedup over the state-of-the-art multi-threaded solution. Moreover, Fast-BNS has a good scalability to the network size as well as sample size. Jiantong Jiang, Zeyi Wen, Ajmal Mian |
IPDPS | 3 |
| 2022 | Self Supervised Learning for Multiple Object Tracking in 3D Point CloudsabstractMultiple object tracking in 3D point clouds has applications in mobile robots and autonomous driving. This is a challenging problem due to the sparse nature of the point clouds and the added difficulty of annotation in 3D for supervised learning. To overcome these challenges, we propose a neural network architecture that learns effective object features and their affinities in a self supervised fashion for multiple object tracking in 3D point clouds captured with LiDAR sensors. For self supervision, we use two approaches. First, we generate two augmented LiDAR frames from a single real frame by applying translation, rotation and cutout to the objects. Second, we synthesize a LiDAR frame using CAD models or primitive geometric shapes and then apply the above three augmentations to them. Hence, the ground truth object locations and associations are known in both frames for self supervision. This removes the need to annotate object associations in real data, and additionally the need for training data collection and annotation for object detection in synthetic data. To the best of our knowledge, this is the first self supervised multiple object tracking method for 3D data. Our model achieves state of the art results. Aakash Kumar, Jyoti Kini, Ajmal Mian, Mubarak Shah |
IROS | 3 |
| 2022 | Spatial Consistency and Feature Diversity Regularization in Transfer Learning for Fine-Grained Visual CategorizationabstractFine-grained visual categorization is challenged by limited training data by localizing discriminative regions and learning diverse features. We propose an effective regularization method that simultaneously imposes spatial consistency and feature diversity on CNN feature maps from a unified perspective. The former guides different feature map channels to concentrate collaboratively on the discriminative areas while the latter ensures that the feature maps are diverse. The proposed method does not require additional supervision, and leverages the covariance matrix of multi-channel feature maps to regularize the loss at the last convolutional layer where the semantic information is the richest. This allows the influence to be backpropagated to update all convolutional layers. We perform experiments using four network architectures for transfer learning from two source domains to three target domains, and demonstrate that our regularization method improves accuracy in all different settings. The proposed regularization method achieves state-of-the-art performance on CUB-200-2011, Stanford-Cars, and Stanford-Dogs datasets with 89.8%, 94.6%, and 88.5% accuracy, respectively. Zhigang Dai, Ajmal Mian |
SMC | 3 |
| 2022 | Transferable 3D Adversarial Textures using End-to-end OptimizationabstractDeep visual models are known to be vulnerable to adversarial attacks. The last few years have seen numerous techniques to compute adversarial inputs for these models. However, there are still under-explored avenues in this critical research direction. Among those is the estimation of adversarial textures for 3D models in an end-to-end optimization scheme. In this paper, we propose such a scheme to generate adversarial textures for 3D models that are highly transferable and invariant to different camera views and lighting conditions. Our method makes use of neural rendering with explicit control over the model texture and background. We ensure transferability of the adversarial textures by employing an ensemble of robust and non-robust models. Our technique utilizes 3D models as a proxy to simulate closer to real-life conditions, in contrast to conventional use of 2D images for adversarial attacks. We show the efficacy of our method with extensive experiments. Camilo Pestana, Naveed Akhtar, Nazanin Rahnavard, Mubarak Shah, Ajmal Mian |
WACV | 5 |
| 2022 | Bi-CLKT: Bi-Graph Contrastive Learning based Knowledge Tracing
Jianxin Li 0001, Wei Zhao 0019, Yunliang Chen 0002, Ajmal Mian |
Knowl. Based Syst. | 6 |
| 2022 | Deep localization of subcellular protein structures from fluorescence microscopy images
Muhammad Tahir 0006, Saeed Anwar, Ajmal Mian, Abdul Wahab Muzaffar |
Neural Comput. Appl. | 3 |
| 2022 | Attack to Fool and Explain Deep NetworksabstractDeep visual models are susceptible to adversarial perturbations to inputs. Although these signals are carefully crafted, they still appear noise-like patterns to humans. This observation has led to the argument that deep visual representation is misaligned with human perception. We counter-argue by providing evidence of human-meaningful patterns in adversarial perturbations. We first propose an attack that fools a network to confuse a whole category of objects (source class) with a target label. Our attack also limits the unintended fooling by samples from non-sources classes, thereby circumscribing human-defined semantic notions for network fooling. We show that the proposed attack not only leads to the emergence of regular geometric patterns in the perturbations, but also reveals insightful information about the decision boundaries of deep models. Exploring this phenomenon further, we alter the 'adversarial' objective of our attack to use it as a tool to 'explain' deep visual representation. We show that by careful channeling and projection of the perturbations computed by our method, we can visualize a model's understanding of human-defined semantic notions. Finally, we exploit the explanability properties of our perturbations to perform image generation, inpainting and interactive image manipulation by attacking adversarialy robust 'classifiers'. In all, our major contribution is a novel pragmatic adversarial attack that is subsequently transformed into a tool to interpret the visual models. The article also makes secondary contributions in terms of establishing the utility of our attack beyond the adversarial objective with multiple interesting applications. Naveed Akhtar, Mohammad A. A. K. Jalwana, Mohammed Bennamoun, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Fast ORB-SLAM Without Keypoint DescriptorsabstractIndirect methods for visual SLAM are gaining popularity due to their robustness to environmental variations. ORB-SLAM2 (Mur-Artal and Tardós, 2017) is a benchmark method in this domain, however, it consumes significant time for computing descriptors that never get reused unless a frame is selected as a keyframe. To overcome these problems, we present FastORB-SLAM which is light-weight and efficient as it tracks keypoints between adjacent frames without computing descriptors. To achieve this, a two stage descriptor-independent keypoint matching method is proposed based on sparse optical flow. In the first stage, we predict initial keypoint correspondences via a simple but effective motion model and then robustly establish the correspondences via pyramid-based sparse optical flow tracking. In the second stage, we leverage the constraints of the motion smoothness and epipolar geometry to refine the correspondences. In particular, our method computes descriptors only for keyframes. We test FastORB-SLAM on TUM and ICL-NUIM RGB-D datasets and compare its accuracy and efficiency to nine existing RGB-D SLAM methods. Qualitative and quantitative results show that our method achieves state-of-the-art accuracy and is about twice as fast as the ORB-SLAM2. Qiang Fu 0013, Hongshan Yu, Xiaolong Wang 0005, Zhengeng Yang, Yong He 0012, Hong Zhang 0013, Ajmal Mian |
IEEE Trans. Image Process. | 7 |
| 2022 | Target-Aware Holistic Influence Maximization in Spatial Social NetworksabstractInfluence maximization has recently received significant attention for scheduling online campaigns or advertisements on social network platforms. However, most studies only focus on user influence via cyber interactions while ignoring their physical interactions which are also essential to gauge influence propagation. Additionally, targeted campaigns or advertisements have not received sufficient attention. To address these issues, we first devise a novel holistic influence diffusion model that takes into account both cyber and physical user interactions in an effective and practical way. Based on the new diffusion model, we formulate a new problem ofholistic influence maximization, denoted asHIMquery, for targeted advertisements in a spatial social network. TheHIMquery problem aims to find a minimum set of users whose holistic influence can cover all target users in the network, which belongs to a set covering problem. Since theHIMquery problem is NP-hard, we develop a greedy baseline algorithm and then improve on this algorithm to reduce the computational cost. To deal with large networks, we also design a spatial-social index to maintain the social, spatial and textual information of users, as well as developing an index-based efficient solution. Finally, we conduct extensive experiments using one synthetic and three real-world datasets to validate the efficiency and effectiveness of the proposed holistic influence diffusion model and our developed algorithms. Taotao Cai, Jianxin Li 0001, Ajmal Mian, Rong-Hua Li 0001, Timos K. Sellis, Jeffrey Xu Yu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Adversarial Attack on Skeleton-Based Human Action RecognitionabstractDeep learning models achieve impressive performance for skeleton-based human action recognition. Graph convolutional networks (GCNs) are particularly suitable for this task due to the graph-structured nature of skeleton data. However, the robustness of these models to adversarial attacks remains largely unexplored due to their complex spatiotemporal nature that must represent sparse and discrete skeleton joints. This work presents the first adversarial attack on skeleton-based action recognition with GCNs. The proposed targeted attack, termed constrained iterative attack for skeleton actions (CIASA), perturbs joint locations in an action sequence such that the resulting adversarial sequence preserves the temporal coherence, spatial integrity, and the anthropomorphic plausibility of the skeletons. CIASA achieves this feat by satisfying multiple physical constraints and employing spatial skeleton realignments for the perturbed skeletons along with regularization of the adversarial skeletons with generative networks. We also explore the possibility of semantically imperceptible localized attacks with CIASA and succeed in fooling the state-of-the-art skeleton action recognition models with high confidence. CIASA perturbations show high transferability in black-box settings. We also show that the perturbed skeleton sequences are able to induce adversarial behavior in the RGB videos created with computer graphics. A comprehensive evaluation with NTU and Kinetics data sets ascertains the effectiveness of CIASA for graph-based skeleton action recognition and reveals the imminent threat to the spatiotemporal deep learning tasks in general. Jian Liu 0014, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Attacking Image Classifiers To Generate 3D Textures
Camilo Pestana, Ajmal Mian |
3DV | 2 |
| 2021 | CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image SaliencyabstractBackpropagation image saliency aims at explaining model predictions by estimating model-centric importance of individual pixels in the input. However, classinsensitivity of the earlier layers in a network only allows saliency computation with low resolution activation maps of the deeper layers, resulting in compromised image saliency. Remedifying this can lead to sanity failures. We propose CAMERAS, a technique to compute high-fidelity backpropagation saliency maps without requiring any external priors and preserving the map sanity. Our method systematically performs multi-scale accumulation and fusion of the activation maps and backpropagated gradients to compute precise saliency maps. From accurate image saliency to articulation of relative importance of input features for different models, and precise discrimination between model perception of visually similar objects, our high-resolution mapping offers multiple novel insights into the black-box deep visual models, which are presented in the paper. We also demonstrate the utility of our saliency maps in adversarial setup by drastically reducing the norm of attack signals by focusing them on the precise regions identified by our maps. Our method also inspires new evaluation metrics and a sanity check for this developing research direction. Mohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal Mian |
CVPR | 4 |
| 2021 | Picasso: A CUDA-Based Library for Deep Learning Over 3D MeshesabstractWe present Picasso, a CUDA-based library comprising novel modules for deep learning over complex real-world 3D meshes. Hierarchical neural architectures have proved effective in multi-scale feature extraction which signifies the need for fast mesh decimation. However, existing methods rely on CPU-based implementations to obtain multi-resolution meshes. We design GPU-accelerated mesh decimation to facilitate network resolution reduction efficiently on-the-fly. Pooling and unpooling modules are defined on the vertex clusters gathered during decimation. For feature learning over meshes, Picasso contains three types of novel convolutions namely, facet2vertex, vertex2facet, and facet2facet convolution. Hence, it treats a mesh as a geometric structure comprising vertices and facets, rather than a spatial graph with edges as previous methods do. Picasso also incorporates a fuzzy mechanism in its filters for robustness to mesh sampling (vertex density). It exploits Gaussian mixtures to define fuzzy coefficients for the facet2vertex convolution, and barycentric interpolation to define the coefficients for the remaining two convolutions. In this release, we demonstrate the effectiveness of the proposed modules with competitive segmentation results on S3DIS. The library will be made public through github. Huan Lei, Naveed Akhtar, Ajmal Mian |
CVPR | 3 |
| 2021 | Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloudabstract3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a freeform language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of point clouds. There are three main challenges in 3D object grounding: to find the main focus in the complex and diverse description; to understand the point cloud scene; and to locate the target object. In this paper, we address all three challenges. Firstly, we propose a language scene graph module to capture the rich structure and long-distance phrase correlations. Secondly, we introduce a multi-level 3D proposal relation graph module to extract the object-object and object-scene co-occurrence relationships, and strengthen the visual features of the initial proposals. Lastly, we develop a description guided 3D visual graph module to encode global contexts of phrases and proposals by a nodes matching strategy. Extensive experiments on challenging benchmark datasets (ScanRefer [3] and Nr3D [42]) show that our algorithm outperforms existing state-of-the-art. Our code is available at https://github.com/PNXD/FFL-3DOG. Mingtao Feng, Liang Zhang 0010, Guangming Zhu 0001, Hui Zhang 0023, Yaonan Wang 0001, Ajmal Mian |
ICCV | 9 |
| 2021 | Adversarial Attacks and Defense on Deep Learning Classification Models using YCbCr Color ImagesabstractDeep neural network models are vulnerable to adversarial perturbations that are subtle but change the model predictions. Adversarial perturbations are generally computed for RGB images and are, hence, equally distributed among the RGB channels. We show, for the first time, that adversarial perturbations prevail in the Y-channel of the$\mathbf{YC}_{b}\mathbf{C}_{r}$> color space and exploit this finding to propose a defense mechanism. Our defense ResUpNet, which is end-to-end trainable, removes perturbations only from the Y-channel by exploiting ResNet features in a bottleneck free up-sampling framework. The refined Y-channel is combined with the untouched$\mathbf{C}_{b}\mathbf{C}_{r}$-channels to restore the clean image. We compare ResUpNet to existing defenses in the input transformation category and show that it achieves the best balance between maintaining the original accuracies on clean images and defense against adversarial attacks. Finally, we show that for the same attack and fixed perturbation magnitude, learning perturbations only in the Y-channel results in higher fooling rates. For example, with a very small perturbation magnitude$\epsilon=0.002$) the fooling rates of FGSM and PGD attacks on the ResNet50 model increase by 11.1% and 15.6% respectively, when the perturbations are learned only for the Y-channel. Camilo Pestana, Naveed Akhtar, Wei Liu 0006, David G. Glance, Ajmal Mian |
IJCNN | 5 |
| 2021 | Object-to-Scene: Learning to Transfer Object Knowledge to Indoor Scene RecognitionabstractAccurate perception of the surrounding scene is helpful for robots to make reasonable judgments and behaviours. Therefore, developing effective scene representation and recognition methods are of significant importance in robotics. Currently, a large body of research focuses on developing novel auxiliary features and networks to improve indoor scene recognition ability. However, few of them focus on directly constructing object features and relations for indoor scene recognition. In this paper, we analyze the weaknesses of current methods and propose an Object-to-Scene (OTS) method, which extracts object features and learns object relations to recognize indoor scenes. The proposed OTS first extracts object features based on the segmentation network and the proposed object feature aggregation module (OFAM). Afterwards, the object relations are calculated and the scene representation is constructed based on the proposed object attention module (OAM) and global relation aggregation module (GRAM). The final results in this work show that OTS successfully extracts object features and learns object relations from the segmentation network. Moreover, OTS outperforms the state-of-the-art methods by more than 2% on indoor scene recognition without using any additional streams. Code is publicly available at: https://github.com/FreeformRobotics/OTS. Bo Miao, Liguang Zhou, Ajmal Mian, Tin Lun Lam, Yangsheng Xu |
IROS | 3 |
| 2021 | Defense-friendly Images in Adversarial Attacks: Dataset and Metrics for Perturbation DifficultyabstractDataset bias is a problem in adversarial machine learning, especially in the evaluation of defenses. An adversarial attack or defense algorithm may show better results on the reported dataset than can be replicated on other datasets. Even when two algorithms are compared, their relative performance can vary depending on the dataset. Deep learning offers state-of-the-art solutions for image recognition, but deep models are vulnerable even to small perturbations. Research in this area focuses primarily on adversarial attacks and defense algorithms. In this paper, we report for the first time, a class of robust images that are both resilient to attacks and that recover better than random images under adversarial attacks using simple defense techniques. Thus, a test dataset with a high proportion of robust images gives a misleading impression about the performance of an adversarial attack or defense. We propose three metrics to determine the proportion of robust images in a dataset and provide scoring to determine the dataset bias. We also provide an ImageNet-R dataset of 15000+ robust images to facilitate further research on this intriguing phenomenon of image strength under attack. Our dataset, combined with the proposed metrics, is valuable for unbiased benchmarking of adversarial attack and defense algorithms. Camilo Pestana, Wei Liu 0006, David G. Glance, Ajmal Mian |
WACV | 4 |
| 2021 | Neural computing and applications (NCAA) special issue on best of DICTA 2019 papers
Ajmal Mian, Lei Wang 0108, Ruiping Wang 0001, Hamid Laga, Naveed Akhtar |
Neural Comput. Appl. | 1 |
| 2021 | Spherical Kernel for Efficient Graph Convolution on 3D Point CloudsabstractWe propose a spherical kernel for efficient graph convolution of 3D point clouds. Our metric-based kernels systematically quantize the local 3D space to identify distinctive geometric relationships in the data. Similar to the regular grid CNN kernels, the spherical kernel maintains translation-invariance and asymmetry properties, where the former guarantees weight sharing among similar local structures in the data and the latter facilitates fine geometric learning. The proposed kernel is applied to graph neural networks without edge-dependent filter generation, making it computationally attractive for large point clouds. In our graph networks, each vertex is associated with a single point location and edges connect the neighborhood points within a defined range. The graph gets coarsened in the network with farthest point sampling. Analogous to the standard CNNs, we define pooling and unpooling operations for our network. We demonstrate the effectiveness of the proposed spherical kernel with graph neural networks for point cloud classification and semantic segmentation using ModelNet, ShapeNet, RueMonge2014, ScanNet and S3DIS datasets. The source code and the trained models can be downloaded from https://github.com/hlei-ziyan/SPH3D-GCN. Huan Lei, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Deep Affinity Network for Multiple Object TrackingabstractMultiple Object Tracking (MOT) plays an important role in solving many fundamental problems in video analysis and computer vision. Most MOT methods employ two steps: Object Detection and Data Association. The first step detects objects of interest in every frame of a video, and the second establishes correspondence between the detected objects in different frames to obtain their tracks. Object detection has made tremendous progress in the last few years due to deep learning. However, data association for tracking still relies on hand crafted constraints such as appearance, motion, spatial proximity, grouping etc. to compute affinities between the objects in different frames. In this paper, we harness the power of deep learning for data association in tracking by jointly modeling object appearances and their affinities between different frames in an end-to-end fashion. The proposed Deep Affinity Network (DAN) learns compact, yet comprehensive features of pre-detected objects at several levels of abstraction, and performs exhaustive pairing permutations of those features in any two frames to infer object affinities. DAN also accounts for multiple objects appearing and disappearing between video frames. We exploit the resulting efficient affinity computations to associate objects in the current frame deep into the previous frames for reliable on-line tracking. Our technique is evaluated on popular multiple object tracking challenges MOT15, MOT17 and UA-DETRAC. Comprehensive benchmarking under twelve evaluation metrics demonstrates that our approach is among the best performing techniques on the leader board for these challenges. The open source implementation of our work is available at https://github.com/shijieS/SST.git. Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Odyssey: Creation, Analysis and Detection of Trojan ModelsabstractAlong with the success of deep neural network (DNN) models, rise the threats to the integrity of these models. A recent threat is the Trojan attack where an attacker interferes with the training pipeline by inserting triggers into some of the training samples and trains the model to act maliciously only for samples that contain the trigger. Since the knowledge of triggers is privy to the attacker, detection of Trojan networks is challenging. Existing Trojan detectors make strong assumptions about the types of triggers and attacks. We propose a detector that is based on the analysis of the intrinsic DNN properties; that are affected due to the Trojan insertion process. For a comprehensive analysis, we develop Odyssey, the most diverse dataset to date with over 3,000 clean and Trojan models. Odyssey covers a large spectrum of attacks; generated by leveraging the versatility in trigger designs and source to target class mappings. Our analysis results show that Trojan attacks affect the classifier margin and shape of decision boundary around the manifold of clean data. Exploiting these two factors, we propose an efficient Trojan detector that operates without any knowledge of the attack and significantly outperforms existing methods. Through a comprehensive set of experiments we demonstrate the efficacy of the detector on cross model architectures, unseen Triggers and regularized models. Marzieh Edraki, Nazmul Karim, Nazanin Rahnavard, Ajmal Mian, Mubarak Shah |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Relation Graph Network for 3D Object Detection in Point CloudsabstractConvolutional Neural Networks (CNNs) have emerged as a powerful tool for object detection in 2D images. However, their power has not been fully realised for detecting 3D objects directly in point clouds without conversion to regular grids. Moreover, existing state-of-the-art 3D object detection methods aim to recognize objects individually without exploiting their relationships during learning or inference. In this article, we first propose a strategy that associates the predictions of direction vectors with pseudo geometric centers, leading to a win-win solution for 3D bounding box candidates regression. Secondly, we propose point attention pooling to extract uniform appearance features for each 3D object proposal, benefiting from the learned direction features, semantic features and spatial coordinates of the object points. Finally, the appearance features are used together with the position features to build 3D object-object relationship graphs for all proposals to model their co-existence. We explore the effect of relation graphs on proposals' appearance feature enhancement under supervised and unsupervised settings. The proposed relation graph network comprises a 3D object proposal generation module and a 3D relation module, making it an end-to-end trainable network for detecting 3D objects in point clouds. Experiments on challenging benchmark point cloud datasets (SunRGB-D, ScanNet and KITTI) show that our algorithm performs better than existing state-of-the-art. Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Liang Zhang 0010, Ajmal Mian |
IEEE Trans. Image Process. | 5 |
| 2020 | Attack to Explain Deep RepresentationabstractDeep visual models are susceptible to extremely low magnitude perturbations to input images. Though carefully crafted, the perturbation patterns generally appear noisy, yet they are able to perform controlled manipulation of model predictions. This observation is used to argue that deep representation is misaligned with human perception. This paper counter-argues and proposes the first attack on deep learning that aims at explaining the learned representation instead of fooling it. By extending the input domain of the manipulative signal and employing a model faithful channelling, we iteratively accumulate adversarial perturbations for a deep model. The accumulated signal gradually manifests itself as a collection of visually salient features of the target label (in model fooling), casting adversarial perturbations as primitive features of the target label. Our attack provides the first demonstration of systematically computing perturbations for adversarially non-robust classifiers that comprise salient visual features of objects. We leverage the model explaining character of our algorithm to perform image generation, inpainting and interactive image manipulation by attacking adversarially robust classifiers. The visually appealing results across these applications demonstrate the utility of our attack (and perturbations in general) beyond model fooling. Mohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal Mian |
CVPR | 4 |
| 2020 | SegGCN: Efficient 3D Point Cloud Segmentation With Fuzzy Spherical KernelabstractFuzzy clustering is known to perform well in real-world applications. Inspired by this observation, we incorporate a fuzzy mechanism into discrete convolutional kernels for 3D point clouds as our first major contribution. The proposed fuzzy kernel is defined over a spherical volume that uses discrete bins. Discrete volumetric division can normally make a kernel vulnerable to boundary effects during learning as well as point density during inference. However, the proposed kernel remains robust to boundary conditions and point density due to the fuzzy mechanism. Our second major contribution comes as the proposal of an efficient graph convolutional network, SegGCN for segmenting point clouds. The proposed network exploits ResNet like blocks in the encoder and 1 × 1 convolutions in the decoder. SegGCN capitalizes on the separable convolution operation of the proposed fuzzy kernel for efficiency. We establish the effectiveness of the SegGCN with the proposed kernel on the challenging S3DIS and ScanNet real-world datasets. Our experiments demonstrate that the proposed network can segment over one million points per second with highly competitive performance. Huan Lei, Naveed Akhtar, Ajmal Mian |
CVPR | 3 |
| 2020 | Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking
Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah |
ECCV (24) | 5 |
| 2020 | Anchored Vertex Exploration for Community Engagement in Social NetworksabstractUser engagement has recently received significant attention in understanding decay and expansion of communities in social networks. However, the problem of user engagement hasn't been fully explored in terms of users' specific interests and structural cohesiveness altogether. Therefore, we fill the gap by investigating the problem of community engagement from the perspective of attributed communities. Given a set of keywords W, a structure cohesive parameter k, and a budget parameter l, our objective is to find l number of users who can induce a maximal expanded community. Meanwhile, every community member must contain the given keywords in W and the community should meet the specified structure cohesiveness constraint k. We introduce this problem as best-Anchored Vertex set Exploration (AVE).To solve the AVE problem, we develop a Filter-Verify framework by maintaining the intermediate results using multiway tree, and probe the best anchored users in a best search way. To accelerate the efficiency, we further design a keyword-aware anchored and follower index, and also develop an index-based efficient algorithm. The proposed algorithm can greatly reduce the cost of computing anchored users and their followers. Additionally, we present two bound properties that can guarantee the correctness of our solution. Finally, we demonstrate the efficiency of our proposed algorithms and index. We measure the effectiveness of attributed community-based community engagement model by conducting extensive experiments on five real-world datasets. Taotao Cai, Jianxin Li 0001, Nur Al Hasan Haldar, Ajmal Mian, John Yearwood, Timos K. Sellis |
ICDE | 4 |
| 2020 | Special issue on Advanced Machine Vision
Steven Puttemans, Toon Goedemé, Ajmal Mian, Thomas B. Moeslund, Rikke Gade |
Mach. Vis. Appl. | 3 |
| 2020 | Hyperspectral Recovery from RGB Images using Gaussian ProcessesabstractWe propose to recover spectral details from RGB images of known spectral quantization by modeling natural spectra under Gaussian Processes and combining them with the RGB images. Our technique exploits Process Kernels to model the relative smoothness of reflectance spectra, and encourages non-negativity in the resulting signals for better estimation of the reflectance values. The Gaussian Processes are inferred in sets using clusters of spatio-spectrally correlated hyperspectral training patches. Each set is transformed to match the spectral quantization of the test RGB image. We extract overlapping patches from the RGB image and match them to the hyperspectral training patches by spectrally transforming the latter. The RGB patches are encoded over the transformed Gaussian Processes related to those hyperspectral patches and the resulting image is constructed by combining the codes with the original processes. Our approach infers the desired Gaussian Processes under a fully Bayesian model inspired by Beta-Bernoulli Process, for which we also present the inference procedure. A thorough evaluation using three hyperspectral datasets demonstrates the effective extraction of spectral details from RGB images by the proposed technique. Naveed Akhtar, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Point attention network for semantic segmentation of 3D point clouds
Mingtao Feng, Liang Zhang 0010, Xuefei Lin, Syed Zulqarnain Gilani, Ajmal Mian |
Pattern Recognit. | 5 |
| 2020 | Small Object Augmentation of Urban Scenes for Real-Time Semantic SegmentationabstractSemantic segmentation is a key step in scene understanding for autonomous driving. Although deep learning has significantly improved the segmentation accuracy, current highquality models such as PSPNet and DeepLabV3 are inefficient given their complex architectures and reliance on multi-scale inputs. Thus, it is difficult to apply them to real-time or practical applications. On the other hand, existing real-time methods cannot yet produce satisfactory results on small objects such as traffic lights, which are imperative to safe autonomous driving. In this paper, we improve the performance of real-time semantic segmentation from two perspectives, methodology and data. Specifically, we propose a real-time segmentation model coined Narrow Deep Network (NDNet) and build a synthetic dataset by inserting additional small objects into the training images. The proposed method achieves 65.7% mean intersection over union (mIoU) on the Cityscapes test set with only 8.4G floatingpoint operations (FLOPs) on 1024×2048 inputs. Furthermore, by re-training the existing PSPNet and DeepLabV3 models on our synthetic dataset, we obtained an average 2% mIoU improvement on small objects. Zhengeng Yang, Hongshan Yu, Mingtao Feng, Wei Sun 0028, Xuefei Lin, Mingui Sun, Zhi-Hong Mao, Ajmal Mian |
IEEE Trans. Image Process. | 8 |
| 2019 | Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video CaptioningabstractAutomatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These methods mainly focus on tailoring sequence learning through RNNs for better caption generation, whereas off-the-shelf visual features are borrowed from CNNs. We argue that careful designing of visual features for this task is equally important, and present a visual feature encoding technique to generate semantically rich captions using Gated Recurrent Units (GRUs). Our method embeds rich temporal dynamics in visual features by hierarchically applying Short Fourier Transform to CNN features of the whole video. It additionally derives high level semantics from an object detector to enrich the representation with spatial dynamics of the detected objects. The final representation is projected to a compact space and fed to a language model. By learning a relatively simple language model comprising two GRU layers, we establish new state-of-the-art on MSVD and MSR-VTT datasets for METEOR and ROUGELmetrics. Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Syed Zulqarnain Gilani, Ajmal Mian |
CVPR | 5 |
| 2019 | Octree Guided CNN With Spherical Kernels for 3D Point CloudsabstractWe propose an octree guided neural network architecture and spherical convolutional kernel for machine learning from arbitrary 3D point clouds. The network architecture capitalizes on the sparse nature of irregular point clouds,and hierarchically coarsens the data representation with space partitioning. At the same time, the proposed spherical kernels systematically quantize point neighborhoods to identify local geometric structures in the data, while maintaining the properties of translation-invariance and asymmetry. We specify spherical kernels with the help of network neurons that in turn are associated with spatial locations.We exploit this association to avert dynamic kernel generation during network training that enables efficient learning with high resolution point clouds. The effectiveness of the proposed technique is established on the benchmark tasks of 3D object classification and segmentation, achieving competitive performance on ShapeNet and RueMonge2014 datasets. Huan Lei, Naveed Akhtar, Ajmal Mian |
CVPR | 3 |
| 2019 | Learning Interpretable Expression-sensitive Features for 3D Dynamic Facial Expression RecognitionabstractDifferent facial components carry different amount of information being conveyed for 3D dynamic expression recognition. Hence, identifying facial components that are highly relevant to specific expression changes is crucial for discriminative facial expression recognition. This work aims to learn expression-sensitive features, which are expected to not only yield comparable recognition performance with the state-of-the-art ones, but also can be interpreted by human. Firstly, spatio-temporal features (HOG3D) are extracted from local depth patch-sequences to represent facial expression dynamics. A two-phase feature selection process is then proposed to determine the facial components that can best distinguish the expressions. In order to verify the effectiveness of the resulting facial components, the expression-sensitive features from the corresponding area are fed into a hierarchical classifier for facial expression recognition. The proposed method is evaluated on the BU-4DFE benchmark database, and results show that learned expression-sensitive features can achieve a comparable recognition performance with existing methods. Additionally, the resulting HOG3D features after feature selection can be used to generate semantic interpretation of the expression dynamics. Mingliang Xue, Ajmal Mian, Xiaodong Duan, Wanquan Liu |
FG | 2 |
| 2019 | Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition
Jian Liu 0014, Hossein Rahmani 0001, Naveed Akhtar, Ajmal Mian |
Int. J. Comput. Vis. | 4 |
| 2019 | Benchmark Data and Method for Real-Time People Counting in Cluttered Scenes Using Depth SensorsabstractVision-based automatic counting of people has widespread applications in intelligent transportation systems, security, and logistics. However, there is currently no large-scale public dataset for benchmarking approaches on this problem. This paper fills this gap by introducing the first real-world RGB-D people counting dataset (PCDS) containing over 4500 videos recorded at the entrance doors of buses in normal and cluttered conditions. It also proposes an efficient method for counting people in real-world cluttered scenes related to public transportations using depth videos. The proposed method computes a point cloud from the depth video frame and re-projects it onto the ground plane to normalize the depth information. The resulting depth image is analyzed for identifying potential human heads. The human head proposals are meticulously refined using a 3D human model. The proposals in each frame of the continuous video stream are tracked to trace their trajectories. The trajectories are again refined to ascertain reliable counting. People are eventually counted by accumulating the head trajectories leaving the scene. To enable effective head and trajectory identification, we also propose two different compound features. A thorough evaluation on PCDS demonstrates that our technique is able to count people in cluttered scenes with high accuracy at 45 fps on a 1.7-GHz processor, and hence it can be deployed for effective real-time people counting for intelligent transportation systems. Shijie Sun 0001, Naveed Akhtar, Huansheng Song, ChaoYang Zhang, Jianxin Li 0001, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2018 | Defense Against Universal Adversarial PerturbationsabstractRecent advances in Deep Learning show the existence of image-agnostic quasi-imperceptible perturbations that when applied to 'any' image can fool a state-of-the-art network classifier to change its prediction about the image label. These 'Universal Adversarial Perturbations' pose a serious threat to the success of Deep Learning in practice. We present the first dedicated framework to effectively defend the networks against such perturbations. Our approach learns a Perturbation Rectifying Network (PRN) as 'pre-input' layers to a targeted model, such that the targeted model needs no modification. The PRN is learned from real and synthetic image-agnostic perturbations, where an efficient method to compute the latter is also proposed. A perturbation detector is separately trained on the Discrete Cosine Transform of the input-output difference of the PRN. A query image is first passed through the PRN and verified by the detector. If a perturbation is detected, the output of the PRN is used for label prediction instead of the actual image. A rigorous evaluation shows that our framework can defend the network classifiers against unseen adversarial perturbations in the real-world scenarios with up to 97.5% success rate. The PRN also generalizes well in the sense that training for one targeted network defends another network with a comparable success rate. Naveed Akhtar, Jian Liu 0014, Ajmal Mian |
CVPR | 3 |
| 2018 | Learning From Millions of 3D Scans for Large-Scale 3D Face RecognitionabstractDeep networks trained on millions of facial images are believed to be closely approaching human-level performance in face recognition. However, open world face recognition still remains a challenge. Although, 3D face recognition has an inherent edge over its 2D counterpart, it has not benefited from the recent developments in deep learning due to the unavailability of large training as well as large test datasets. Recognition accuracies have already saturated on existing 3D face datasets due to their small gallery sizes. Unlike 2D photographs, 3D facial scans cannot be sourced from the web causing a bottleneck in the development of deep 3D face recognition networks and datasets. In this backdrop, we propose a method for generating a large corpus of labeled 3D face identities and their multiple instances for training and a protocol for merging the most challenging existing 3D datasets for testing. We also propose the first deep CNN model designed specifically for 3D face recognition and trained on 3.1 Million 3D facial scans of 100K identities. Our test dataset comprises 1,853 identities with a single 3D scan in the gallery and another 31K scans as probes, which is several orders of magnitude larger than existing ones. Without fine tuning on this dataset, our network already outperforms state of the art face recognition by over 10%. We fine tune our network on the gallery set to perform end-to-end large scale 3D face recognition which further improves accuracy. Finally, we show the efficacy of our method for the open world face recognition problem. Syed Zulqarnain Gilani, Ajmal Mian |
CVPR | 2 |
| 2018 | 3D Face Reconstruction from Light Field Images: A Model-Free Approach
Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Ajmal Mian |
ECCV (10) | 4 |
| 2018 | Holistic Influence Maximization for Targeted Advertisements in Spatial Social NetworksabstractThe problem of influence maximization has recently received significant attention. However, most studies focused on user influence via cyber interactions while ignoring their physical interactions which are important to gauge influence propagation. Additionally, targeted campaigns or advertisements have not received sufficient attention. To do this, we first devise a novel holistic influence diffusion model and then formulate a new holistic influence maximization query problem and develop three algorithms. Finally, we conduct extensive experiments to evaluate the effectiveness and efficiency of the proposed solutions. Jianxin Li 0001, Taotao Cai, Ajmal Mian, Rong-Hua Li 0001, Timos K. Sellis, Jeffrey Xu Yu |
ICDE | 3 |
| 2018 | Representation learning with deep extreme learning machines for efficient image set classification
Faisal Shafait, Bernard Ghanem, Ajmal Mian |
Neural Comput. Appl. | 4 |
| 2018 | Dense 3D Face CorrespondenceabstractWe present an algorithm that automatically establishes dense correspondences between a large number of 3D faces. Starting from automatically detected sparse correspondences on the outer boundary of 3D faces, the algorithm triangulates existing correspondences and expands them iteratively by matching points of distinctive surface curvature along the triangle edges. After exhausting keypoint matches, further correspondences are established by generating evenly distributed points within triangles by evolving level set geodesic curves from the centroids of large triangles. A deformable model (K3DM) is constructed from the dense corresponded faces and an algorithm is proposed for morphing the K3DM to fit unseen faces. This algorithm iterates between rigid alignment of an unseen face followed by regularized morphing of the deformable model. We have extensively evaluated the proposed algorithms on synthetic data and real 3D faces from the FRGCv2, Bosphorus, BU3DFE and UND Ear databases using quantitative and qualitative benchmarks. Our algorithm achieved dense correspondences with a mean localisation error of 1.28 mm on synthetic faces and detected 14 anthropometric landmarks on unseen real faces from the FRGCv2 database with 3 mm precision. Furthermore, our deformable model fitting algorithm achieved 98.5 percent face recognition accuracy on the FRGCv2 and 98.6 percent on Bosphorus database. Our dense model is also able to generalize to unseen datasets. Syed Zulqarnain Gilani, Ajmal Mian, Faisal Shafait, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Learning a Deep Model for Human Action Recognition from Novel ViewpointsabstractRecognizing human actions from unknown and unseen (novel) views is a challenging problem. We propose a Robust Non-Linear Knowledge Transfer Model (R-NKTM) for human action recognition from novel views. The proposed R-NKTM is a deep fully-connected neural network that transfers knowledge of human actions from any unknown view to a shared high-level virtual view by finding a set of non-linear transformations that connects the views. The R-NKTM is learned from 2D projections of dense trajectories of synthetic 3D human models fitted to real motion capture data and generalizes to real videos of human actions. The strength of our technique is that we learn a single R-NKTM for all actions and all viewpoints for knowledge transfer of any real human action video without the need for re-training or fine-tuning the model. Thus, R-NKTM can efficiently scale to incorporate new action classes. R-NKTM is learned with dummy labels and does not require knowledge of the camera viewpoint at any stage. Experiments on three benchmark cross-view human action datasets show that our method outperforms existing state-of-the-art. Hossein Rahmani 0001, Ajmal Mian, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Distance metric learning for pattern recognition
Jiwen Lu, Ruiping Wang 0001, Ajmal Mian, Sudeep Sarkar |
Pattern Recognit. | 3 |
| 2018 | Benchmark Data Set and Method for Depth Estimation From Light Field ImagesabstractConvolutional Neural Networks (CNN) have performed extremely well for many image analysis tasks. However, supervised training of deep CNN architectures requires huge amounts of labelled data which is unavailable for light field images. In this paper, we leverage on synthetic light field images and propose a two stream CNN network that learns to estimate the disparities of multiple correlated neighbourhood pixels from their Epipolar Plane Images (EPI). Since the EPIs are unrelated except at their intersection, a two stream network is proposed to learn convolution weights individually for the EPIs and then combine the outputs of the two streams for disparity estimation. The CNN estimated disparity map is then refined using the central RGB light field image as a prior in a variational technique. We also propose a new real world dataset comprising light field images of 19 objects captured with the Lytro Illum camera in outdoor scenes and their corresponding 3D pointclouds, as ground truth, captured with the 3dMD scanner. This dataset will be made public to allow more precise 3D pointcloud level comparison of algorithms in the future which is currently not possible. Experiments on the synthetic and real world datasets show that our algorithm outperforms existing state-of-the-art for depth estimation from light field images. Mingtao Feng, Yaonan Wang 0001, Jian Liu 0014, Liang Zhang 0010, Hasan Firdaus M. Zaki, Ajmal Mian |
IEEE Trans. Image Process. | 6 |
| 2018 | Nonparametric Coupled Bayesian Dictionary and Classifier Learning for Hyperspectral ClassificationabstractWe present a principled approach to learn a discriminative dictionary along a linear classifier for hyperspectral classification. Our approach places Gaussian Process priors over the dictionary to account for the relative smoothness of the natural spectra, whereas the classifier parameters are sampled from multivariate Gaussians. We employ two Beta-Bernoulli processes to jointly infer the dictionary and the classifier. These processes are coupled under the same sets of Bernoulli distributions. In our approach, these distributions signify the frequency of the dictionary atom usage in representing class-specific training spectra, which also makes the dictionary discriminative. Due to the coupling between the dictionary and the classifier, the popularity of the atoms for representing different classes gets encoded into the classifier. This helps in predicting the class labels of test spectra that are first represented over the dictionary by solving a simultaneous sparse optimization problem. The labels of the spectra are predicted by feeding the resulting representations to the classifier. Our approach exploits the nonparametric Bayesian framework to automatically infer the dictionary size-the key parameter in discriminative dictionary learning. Moreover, it also has the desirable property of adaptively learning the association between the dictionary atoms and the class labels by itself. We use Gibbs sampling to infer the posterior probability distributions over the dictionary and the classifier under the proposed model, for which, we derive analytical expressions. To establish the effectiveness of our approach, we test it on benchmark hyperspectral images. The classification performance is compared with the state-of-the-art dictionary learning-based classification methods. Naveed Akhtar, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Joint Discriminative Bayesian Dictionary and Classifier LearningabstractWe propose to jointly learn a Discriminative Bayesian dictionary along a linear classifier using coupled Beta-Bernoulli Processes. Our representation model uses separate base measures for the dictionary and the classifier, but associates them to the class-specific training data using the same Bernoulli distributions. The Bernoulli distributions control the frequency with which the factors (e.g. dictionary atoms) are used in data representations, and they are inferred while accounting for the class labels in our approach. To further encourage discrimination in the dictionary, our model uses separate (sets of) Bernoulli distributions to represent data from different classes. Our approach adaptively learns the association between the dictionary atoms and the class labels while tailoring the classifier to this relation with a joint inference over the dictionary and the classifier. Once a test sample is represented over the dictionary, its representation is accurately labelled by the classifier due to the strong coupling between the dictionary and the classifier. We derive the Gibbs Sampling equations for our joint representation model and test our approach for face, object, scene and action recognition to establish its effectiveness. Naveed Akhtar, Ajmal Mian, Fatih Porikli |
CVPR | 2 |
| 2017 | Modeling Sub-Event Dynamics in First-Person Action RecognitionabstractFirst-person videos have unique characteristics such as heavy egocentric motion, strong preceding events, salient transitional activities and post-event impacts. Action recognition methods designed for third person videos may not optimally represent actions captured by first-person videos. We propose a method to represent the high level dynamics of sub-events in first-person videos by dynamically pooling features of sub-intervals of time series using a temporal feature pooling function. The sub-event dynamics are then temporally aligned to make a new series. To keep track of how the sub-event dynamics evolve over time, we recursively employ the Fast Fourier Transform on a pyramidal temporal structure. The Fourier coefficients of the segment define the overall video representation. We perform experiments on two existing benchmark first-person video datasets which have been captured in a controlled environment. Addressing this gap, we introduce a new dataset collected from YouTube which has a larger number of classes and a greater diversity of capture conditions thereby more closely depicting real-world challenges in first-person video analysis. We compare our method to state-of-the-art first person and generic video recognition algorithms. Our method consistently outperforms the nearest competitors by 10.3%, 3.3% and 11.7% respectively on the three datasets. Hasan Firdaus M. Zaki, Faisal Shafait, Ajmal Mian |
CVPR | 3 |
| 2017 | Guest Editorial: Language in Vision
Yan Yan 0002, Jiwen Lu, Ajmal Mian, Arun Ross, Vittorio Murino, Radu Horaud |
Comput. Vis. Image Underst. | 3 |
| 2017 | Regularization techniques for high-dimensional data analysis
Jiwen Lu, Xi Peng 0001, Weihong Deng, Ajmal Mian |
Image Vis. Comput. | 4 |
| 2017 | Efficient classification with sparsity augmented collaborative representation
Naveed Akhtar, Faisal Shafait, Ajmal Mian |
Pattern Recognit. | 3 |
| 2017 | Deep, dense and accurate 3D face correspondence for generating population specific deformable models
Syed Zulqarnain Gilani, Ajmal Mian, Peter R. Eastwood |
Pattern Recognit. | 2 |
| 2017 | Surface geodesic pattern for 3D deformable texture matching
Farshid Hajati, Ali Cheraghian, Soheila Gheisari, Yongsheng Gao 0001, Ajmal Mian |
Pattern Recognit. | 5 |
| 2017 | Blind Domain Adaptation With Augmented Extreme Learning Machine FeaturesabstractIn practical applications, the test data often have different distribution from the training data leading to suboptimal visual classification performance. Domain adaptation (DA) addresses this problem by designing classifiers that are robust to mismatched distributions. Existing DA algorithms use the unlabeled test data from target domain during training time in addition to the source domain data. However, target domain data may not always be available for training. We propose a blind DA algorithm that does not require target domain samples for training. For this purpose, we learn a global nonlinear extreme learning machine (ELM) model from the source domain data in an unsupervised fashion. The global ELM model is then used to initialize and learn class specific ELM models from the source domain data. During testing, the target domain features are augmented with the reconstructed features from the global ELM model. The resulting enriched features are then classified using the class specific ELM models based on minimum reconstruction error. Extensive experiments on 16 standard datasets show that despite blind learning, our method outperforms six existing state-of-the-art methods in cross domain visual recognition. Ajmal Mian |
IEEE Trans. Cybern. | 2 |
| 2017 | RCMF: Robust Constrained Matrix Factorization for Hyperspectral UnmixingabstractWe propose a constrained matrix factorization approach for linear unmixing of hyperspectral data. Our approach factorizes a hyperspectral cube into its constituent endmembers and their fractional abundances such that the endmembers are sparse nonnegative linear combinations of the observed spectra themselves. The association between the extracted endmembers and the observed spectra is explicitly noted for physical interpretability. To ensure reliable unmixing, we make the matrix factorization procedure robust to outliers in the observed spectra. Our approach simultaneously computes the endmembers and their abundances in an efficient and unsupervised manner. The extracted endmembers are nonnegative quantities, whereas their abundances additionally follow the sum-to-one constraint. We thoroughly evaluate our approach using synthetic data with white and correlated noise as well as real hyperspectral data. Experimental results establish the effectiveness of our approach. Naveed Akhtar, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Dynamic Texture Comparison Using Derivative Sparse Representation: Application to Video-Based Face RecognitionabstractVideo-based face, expression, and scene recognition are fundamental problems in human-machine interaction, especially when there is a short-length video. In this paper, we present a new derivative sparse representation approach for face and texture recognition using short-length videos. First, it builds local linear subspaces of dynamic texture segments by computing spatiotemporal directional derivatives in a cylinder neighborhood within dynamic textures. Unlike traditional methods, a nonbinary texture coding technique is proposed to extract high-order derivatives using continuous circular and cylinder regions to avoid aliasing effects. Then, these local linear subspaces of texture segments are mapped onto a Grassmann manifold via sparse representation. A new joint sparse representation algorithm is developed to establish the correspondences of subspace points on the manifold for measuring the similarity between two dynamic textures. Extensive experiments on the Honda/UCSD, the CMU motion of body, the YouTube, and the DynTex datasets show that the proposed method consistently outperforms the state-of-the-art methods in dynamic texture recognition, and achieved the encouraging highest accuracy reported to date on the challenging YouTube face dataset. The encouraging experimental results show the effectiveness of the proposed method in video-based face recognition in human-machine system applications. Farshid Hajati, Mohammad Tavakolian, Soheila Gheisari, Yongsheng Gao 0001, Ajmal Mian |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2016 | 3D Action Recognition from Novel ViewpointsabstractWe propose a human pose representation model that transfers human poses acquired from different unknown views to a view-invariant high-level space. The model is a deep convolutional neural network and requires a large corpus of multiview training data which is very expensive to acquire. Therefore, we propose a method to generate this data by fitting synthetic 3D human models to real motion capture data and rendering the human poses from numerous viewpoints. While learning the CNN model, we do not use action labels but only the pose labels after clustering all training poses into k clusters. The proposed model is able to generalize to real depth images of unseen poses without the need for re-training or fine-tuning. Real depth videos are passed through the model frame-wise to extract view-invariant features. For spatio-temporal representation, we propose group sparse Fourier Temporal Pyramid which robustly encodes the action specific most discriminative output features of the proposed human pose model. Experiments on two multiview and three single-view benchmark datasets show that the proposed method dramatically outperforms existing state-of-the-art in action recognition. Hossein Rahmani 0001, Ajmal Mian |
CVPR | 2 |
| 2016 | Hierarchical Beta Process with Gaussian Process Prior for Hyperspectral Image Super Resolution
Naveed Akhtar, Faisal Shafait, Ajmal Mian |
ECCV (3) | 3 |
| 2016 | Automatic Signature Segmentation Using Hyper-Spectral ImagingabstractIn this paper, we propose a method for automatic signature segmentation using hyper-spectral imaging. The proposed method first uses the connected component analysis and local features to segment the printed text and signatures. Secondly, it uses spectral response of text, signature, and background to extract signature pixels. The proposed method is robust, and remains unaffected by color and intensity of the ink, and by any structural information of the text, as the classification relies exclusively on the spectral response of the document. The proposed method can extract signature pixels either overlapping or non-overlapping from different backgrounds like, logos, tables, stamps, and printed text. We used high-resolution hyper-spectral imaging to study and classify 300 documents with varying backgrounds. We evaluated the proposed classification method and compared results with the state-of-the art system. The proposed method outperformed the state-of-the-art system and achieved 100% precision and 84% recall. Umair Muneer Butt, Sheraz Ahmed, Faisal Shafait, Christian Nansen, Ajmal Mian, Muhammad Imran Malik |
ICFHR | 5 |
| 2016 | Convolutional hypercube pyramid for accurate RGB-D object category and instance recognitionabstractDeep learning based methods have achieved unprecedented success in solving several computer vision problems involving RGB images. However, this level of success is yet to be seen on RGB-D images owing to two major challenges in this domain: training data deficiency and multi-modality input dissimilarity. We present an RGB-D object recognition framework that addresses these two key challenges by effectively embedding depth and point cloud data into the RGB domain. We employ a convolutional neural network (CNN) pre-trained on RGB data as a feature extractor for both color and depth channels and propose a rich coarse-to-fine feature representation scheme, coined Hypercube Pyramid, that is able to capture discriminatory information at different levels of detail. Finally, we present a novel fusion scheme to combine the Hypercube Pyramid features with the activations of fully connected neurons to construct a compact representation prior to classification. By employing Extreme Learning Machines (ELM) as non-linear classifiers, we show that the proposed method outperforms ten state-of-the-art algorithms for several tasks in terms of recognition accuracy on the benchmark Washington RGB-D and 2D3D object datasets by a large margin (upto 50% reduction in error rate). Hasan Firdaus M. Zaki, Faisal Shafait, Ajmal Mian |
ICRA | 3 |
| 2016 | Robust RGB-D face recognition using Kinect sensor
Billy Y. L. Li, Mingliang Xue, Ajmal Mian, Wanquan Liu, Aneesh Krishna |
Neurocomputing | 3 |
| 2016 | Unsupervised manifold alignment using soft-assign technique
Ajmal Mian, Wanquan Liu, Ling Li 0006 |
Mach. Vis. Appl. | 2 |
| 2016 | Face recognition based on Kinect
Billy Y. L. Li, Ajmal Mian, Wanquan Liu, Aneesh Krishna |
Pattern Anal. Appl. | 2 |
| 2016 | Discriminative Bayesian Dictionary Learning for ClassificationabstractWe propose a Bayesian approach to learn discriminative dictionaries for sparse representation of data. The proposed approach infers probability distributions over the atoms of a discriminative dictionary using a finite approximation of Beta Process. It also computes sets of Bernoulli distributions that associate class labels to the learned dictionary atoms. This association signifies the selection probabilities of the dictionary atoms in the expansion of class-specific data. Furthermore, the non-parametric character of the proposed approach allows it to infer the correct size of the dictionary. We exploit the aforementioned Bernoulli distributions in separately learning a linear classifier. The classifier uses the same hierarchical Bayesian model as the dictionary, which we present along the analytical inference solution for Gibbs sampling. For classification, a test instance is first sparsely encoded over the learned dictionary and the codes are fed to the classifier. We performed experiments for face and action recognition; and object and scene-category classification using five public datasets and compared the results with state-of-the-art discriminative sparse representation approaches. Experiments show that the proposed Bayesian approach consistently outperforms the existing approaches. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Histogram of Oriented Principal Components for Cross-View Action RecognitionabstractExisting techniques for 3D action recognition are sensitive to viewpoint variations because they extract features from depth images which are viewpoint dependent. In contrast, we directly process pointclouds for cross-view action recognition from unknown and unseen views. We propose the histogram of oriented principal components (HOPC) descriptor that is robust to noise, viewpoint, scale and action speed variations. At a 3D point, HOPC is computed by projecting the three scaled eigenvectors of the pointcloud within its local spatio-temporal support volume onto the vertices of a regular dodecahedron. HOPC is also used for the detection of spatio-temporal keypoints (STK) in 3D pointcloud sequences so that view-invariant STK descriptors (or Local HOPC descriptors) at these key locations only are used for action recognition. We also propose a global descriptor computed from the normalized spatio-temporal distribution of STKs in 4-D, which we refer to as STK-D. We have evaluated the performance of our proposed descriptors against nine existing techniques on two cross-view and three single-view human action recognition datasets. The experimental results show that our techniques provide significant improvement over state-of-the-art methods. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Discriminative human action classification using locality-constrained linear coding
Hossein Rahmani 0001, Du Q. Huynh, Arif Mahmood, Ajmal Mian |
Pattern Recognit. Lett. | 4 |
| 2015 | Bayesian sparse representation for hyperspectral image super resolutionabstractDespite the proven efficacy of hyperspectral imaging in many computer vision tasks, its widespread use is hindered by its low spatial resolution, resulting from hardware limitations. We propose a hyperspectral image super resolution approach that fuses a high resolution image with the low resolution hyperspectral image using non-parametric Bayesian sparse representation. The proposed approach first infers probability distributions for the material spectra in the scene and their proportions. The distributions are then used to compute sparse codes of the high resolution image. To that end, we propose a generic Bayesian sparse coding strategy to be used with Bayesian dictionaries learned with the Beta process. We theoretically analyze the proposed strategy for its accurate performance. The computed codes are used with the estimated scene spectra to construct the super resolution hyperspectral image. Exhaustive experiments on two public databases of ground based hyperspectral images and a remotely sensed image show that the proposed approach outperforms the existing state of the art. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
CVPR | 3 |
| 2015 | Shape-based automatic detection of a large number of 3D facial landmarksabstractWe present an algorithm for automatic detection of a large number of anthropometric landmarks on 3D faces. Our approach does not use texture and is completely shape based in order to detect landmarks that are morphologically significant. The proposed algorithm evolves level set curves with adaptive geometric speed functions to automatically extract effective seed points for dense correspondence. Correspondences are established by minimizing the bending energy between patches around seed points of given faces to those of a reference face. Given its hierarchical structure, our algorithm is capable of establishing thousands of correspondences between a large number of faces. Finally, a morphable model based on the dense corresponding points is fitted to an unseen query face for transfer of correspondences and hence automatic detection of landmarks. The proposed algorithm can detect any number of pre-defined landmarks including subtle landmarks that are even difficult to detect manually. Extensive experimental comparison on two benchmark databases containing 6, 507 scans shows that our algorithm outperforms six state of the art algorithms. Syed Zulqarnain Gilani, Faisal Shafait, Ajmal Mian |
CVPR | 3 |
| 2015 | Learning a non-linear knowledge transfer model for cross-view action recognitionabstractThis paper concerns action recognition from unseen and unknown views. We propose unsupervised learning of a non-linear model that transfers knowledge from multiple views to a canonical view. The proposed Non-linear Knowledge Transfer Model (NKTM) is a deep network, with weight decay and sparsity constraints, which finds a shared high-level virtual path from videos captured from different unknown viewpoints to the same canonical view. The strength of our technique is that we learn a single NKTM for all actions and all camera viewing directions. Thus, NKTM does not require action labels during learning and knowledge of the camera viewpoints during training or testing. NKTM is learned once only from dense trajectories of synthetic points fitted to mocap data and then applied to real video data. Trajectories are coded with a general codebook learned from the same mocap data. NKTM is scalable to new action classes and training data as it does not require re-learning. Experiments on the IXMAS and N-UCLA datasets show that NKTM outperforms existing state-of-the-art methods for cross-view action recognition. Hossein Rahmani 0001, Ajmal Mian |
CVPR | 2 |
| 2015 | Localized forgery detection in hyperspectral document imagesabstractHyperspectral imaging is emerging as a promising technology to discover patterns that are otherwise hard to identify with regular cameras. Recent research has shown the potential of hyperspectral image analysis to automatically distinguish visually similar inks. However, a major limitation of prior work is that automatic distinction only works when the number of inks to be distinguished is known a priori and their relative proportions in the inspected image are roughly equal. This research work aims at addressing these two problems. We show how anomaly detection combined with unsupervised clustering can be used to handle cases where the proportions of pixels belonging to the two inks are highly unbalanced. We have performed experiments on the publicly available UWA Hyperspectral Documents dataset. Our results show that INFLO anomaly detection algorithm is able to best distinguish inks for highly unbalanced ink proportions. Zhipei Luo, Faisal Shafait, Ajmal Mian |
ICDAR | 3 |
| 2015 | Automatic 4D Facial Expression Recognition Using DCT FeaturesabstractThis paper addresses the problem of person-independent 4D facial expression recognition. Unlike the majority of existing works, we propose to extract spatio-temporal features in 4D data (3D expression sequences changing over time) to represent 3D facial expression dynamics sufficiently, rather than extracting features frame-by-frame. First, the proposed method extracts local depth patch-sequences from consecutive expression frames based on the automatically detected facial landmarks. Three dimension discrete cosine transform (3D-DCT) is then applied on these patch-sequences to extract spatio-temporal features for facial expression dynamic representation. Finally, the extracted compact features (3D-DCT coefficients) are fed to nearest-neighbor classifier to finish expression recognition after feature selection and dimension reduction, in which the redundant features are filtered out. Experiments on the benchmark BU-4DFE database show that the proposed method achieves the best average recognition rate 78.8% among the existing automatic approaches, and outperforms the existing techniques in the recognition of those easily confused expressions (anger and sadness) significantly. Mingliang Xue, Ajmal Mian, Wanquan Liu, Ling Li 0006 |
WACV | 2 |
| 2015 | Periocular region-based person identification in the visible, infrared and hyperspectral imagery
Arif Mahmood, Ajmal Mian, Chris McDonald |
Neurocomputing | 3 |
| 2015 | Automatic ink mismatch detection for forensic document analysis
Zohaib Khan 0001, Faisal Shafait, Ajmal Mian |
Pattern Recognit. | 3 |
| 2015 | Futuristic Greedy Approach to Sparse Unmixing of Hyperspectral DataabstractSpectra measured at a single pixel of a remotely sensed hyperspectral image is usually a mixture of multiple spectral signatures (endmembers) corresponding to different materials on the ground. Sparse unmixing assumes that a mixed pixel is a sparse linear combination of different spectra already available in a spectral library. It uses sparse approximation (SA) techniques to solve the hyperspectral unmixing problem. Among these techniques, greedy algorithms suite well to sparse unmixing. However, their accuracy is immensely compromised by the high correlation of the spectra of different materials. This paper proposes a novel greedy algorithm, called OMP-Star, that shows robustness against the high correlation of spectral signatures. We preprocess the signals with spectral derivatives before they are used by the algorithm. To approximate the mixed pixel spectra, the algorithm employs a futuristic greedy approach that, if necessary, considers its future iterations before identifying an endmember. We also extend OMP-Star to exploit the nonnegativity of spectral mixing. Experiments on simulated and real hyperspectral data show that the proposed algorithms outperform the state-of-the-art greedy algorithms. Moreover, the proposed approach achieves results comparable to convex relaxation-based SA techniques, while maintaining the advantages of greedy approaches. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Joint Group Sparse PCA for Compressed Hyperspectral ImagingabstractA sparse principal component analysis (PCA) seeks a sparse linear combination of input features (variables), so that the derived features still explain most of the variations in the data. A group sparse PCA introduces structural constraints on the features in seeking such a linear combination. Collectively, the derived principal components may still require measuring all the input features. We present a joint group sparse PCA (JGSPCA) algorithm, which forces the basic coefficients corresponding to a group of features to be jointly sparse. Joint sparsity ensures that the complete basis involves only a sparse set of input features, whereas the group sparsity ensures that the structural integrity of the features is maximally preserved. We evaluate the JGSPCA algorithm on the problems of compressed hyperspectral imaging and face recognition. Compressed sensing results show that the proposed method consistently outperforms sparse PCA and group sparse PCA in reconstructing the hyperspectral scenes of natural and man-made objects. The efficacy of the proposed compressed sensing method is further demonstrated in band selection for face recognition. Zohaib Khan 0001, Faisal Shafait, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2015 | Hyperspectral Face Recognition With Spatiospectral Information Fusion and PLS RegressionabstractHyperspectral imaging offers new opportunities for face recognition via improved discrimination along the spectral dimension. However, it poses new challenges, including low signal-to-noise ratio, interband misalignment, and high data dimensionality. Due to these challenges, the literature on hyperspectral face recognition is not only sparse but is limited to ad hoc dimensionality reduction techniques and lacks comprehensive evaluation. We propose a hyperspectral face recognition algorithm using a spatiospectral covariance for band fusion and partial least square regression for classification. Moreover, we extend 13 existing face recognition techniques, for the first time, to perform hyperspectral face recognition.We formulate hyperspectral face recognition as an image-set classification problem and evaluate the performance of seven state-of-the-art image-set classification techniques. We also test six state-of-the-art grayscale and RGB (color) face recognition algorithms after applying fusion techniques on hyperspectral images. Comparison with the 13 extended and five existing hyperspectral face recognition techniques on three standard data sets show that the proposed algorithm outperforms all by a significant margin. Finally, we perform band selection experiments to find the most discriminative bands in the visible and near infrared response spectrum. Arif Mahmood, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2014 | Spatiotemporal Derivative Pattern: A Dynamic Texture Descriptor for Video Matching
Farshid Hajati, Mohammad Tavakolian, Soheila Gheisari, Ajmal Mian |
ACCV (5) | 4 |
| 2014 | Sparse Kernel Learning for Image Set Classification
Arif Mahmood, Ajmal Mian |
ACCV (2) | 3 |
| 2014 | Semi-supervised Spectral Clustering for Image Set ClassificationabstractWe present an image set classification algorithm based on unsupervised clustering of labeled training and unlabeled test data where labels are only used in the stopping criterion. The probability distribution of each class over the set of clusters is used to define a true set based similarity measure. To this end, we propose an iterative sparse spectral clustering algorithm. In each iteration, a proximity matrix is efficiently recomputed to better represent the local subspace structure. Initial clusters capture the global data structure and finer clusters at the later stages capture the subtle class differences not visible at the global scale. Image sets are compactly represented with multiple Grassmannian manifolds which are subsequently embedded in Euclidean space with the proposed spectral clustering algorithm. We also propose an efficient eigenvector solver which not only reduces the computational cost of spectral clustering by many folds but also improves the clustering quality and final classification results. Experiments on five standard datasets and comparison with seven existing techniques show the efficacy of our algorithm. Arif Mahmood, Ajmal Mian, Robyn A. Owens |
CVPR | 2 |
| 2014 | Sparse Spatio-spectral Representation for Hyperspectral Image Super-resolution
Naveed Akhtar, Faisal Shafait, Ajmal Mian |
ECCV (7) | 3 |
| 2014 | HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
ECCV (2) | 4 |
| 2014 | Adaptive spectral reflectance recovery using spatio-spectral support from hyperspectral imagesabstractAccurate knowledge of spectral reflectance is crucial for hyperspectral image analysis. We propose a novel spectral reflectance recovery method by adaptive spatio-spectral support. The proposed technique is evaluated in both simulated and real illumination scenarios. A multi-illuminant hyper-spectral scene database has been collected and made publicly available. Experiments show that the adaptive illuminant estimation reduces the mean angular error of recovered spectra by 13%. Zohaib Khan 0001, Faisal Shafait, Ajmal Mian |
ICIP | 3 |
| 2014 | SUnGP: A Greedy Sparse Approximation Algorithm for Hyperspectral UnmixingabstractSpectra measured at a pixel of a remote sensing hyper spectral sensor is usually a mixture of multiple spectra (end-members) of different materials on the ground. Hyper spectral unmixing aims at identifying the end members and their proportions (fractional abundances) in the mixed pixels. Hyper spectral unmixing has recently been casted into a sparse approximation problem and greedy sparse approximation approaches are considered desirable for solving it. However, the high correlation among the spectra of different materials seriously affects the accuracy of the greedy algorithms. We propose a greedy sparse approximation algorithm, called SUnGP, for unmixing of hyper spectral data. SUnGP shows high robustness against the correlation of the spectra of materials. The algorithm employees a subspace pruning strategy for the identification of the end members. Experiments show that the proposed algorithm not only outperforms the state of the art greedy algorithms, its accuracy is comparable to the algorithms based on the convex relaxation of the problem, but with a considerable computational advantage. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
ICPR | 3 |
| 2014 | Perceptual Differences between Men and Women: A 3D Facial Morphometric PerspectiveabstractUnderstanding the features employed by the human visual system in gender classification is considered a critical step towards improving machine based gender classification systems. We propose the use of 3D Euclidean and geodesic distances between biologically significant facial landmarks to classify gender. We perform five different experiments on the BU-3DFE face database to look for more representative features that can replicate our visual system. Based on our experiments we suggest that the human visual system looks at the ratio of 3D Euclidean and geodesic distance as these features can classify facial gender with an accuracy of 99.32%. The features selected by our proposed gender classification experiment are robust to ethnicity and moderate changes in expression. They also replicate the perceptual gender bias towards certain features and hence become good candidates for being a more representative feature set. Syed Zulqarnain Gilani, Ajmal Mian |
ICPR | 2 |
| 2014 | Action Classification with Locality-Constrained Linear CodingabstractWe propose an action classification algorithm which uses Locality-constrained Linear Coding (LLC) to capture discriminative information of human body variations in each spatio-temporal subsequence of a video sequence. Our proposed method divides the input video into equally spaced overlapping spatio-temporal sub sequences, each of which is decomposed into blocks and then cells. We use the Histogram of Oriented Gradient (HOG3D) feature to encode the information in each cell. We justify the use of LLC for encoding the block descriptor by demonstrating its superiority over Sparse Coding (SC). Our sequence descriptor is obtained via a logistic regression classifier with L2 regularization. We evaluate and compare our algorithm with ten state-of-the-art algorithms on five benchmark datasets. Experimental results show that, on average, our algorithm gives better accuracy than these ten algorithms. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
ICPR | 4 |
| 2014 | Repeated constrained sparse coding with partial dictionaries for hyperspectral unmixingabstractHyperspectral images obtained from remote sensing platforms have limited spatial resolution. Thus, each spectra measured at a pixel is usually a mixture of many pure spectral signatures (endmembers) corresponding to different materials on the ground. Hyperspectral unmixing aims at separating these mixed spectra into its constituent end-members. We formulate hyperspectral unmixing as a constrained sparse coding (CSC) problem where unmixing is performed with the help of a library of pure spectral signatures under positivity and summation constraints. We propose two different methods that perform CSC repeatedly over the hyperspectral data. However, the first method, Repeated-CSC (RCSC), systematically neglects a few spectral bands of the data each time it performs the sparse coding. Whereas the second method, Repeated Spectral Derivative (RSD), takes the spectral derivative of the data before the sparse coding stage. The spectral derivative is taken such that it is not operated on a few selected bands. Experiments on simulated and real hyperspectral data and comparison with existing state of the art show that the proposed methods achieve significantly higher accuracy. Our results demonstrate the overall robustness of RCSC to noise and better performance of RSD at high signal to noise ratio. Naveed Akhtar, Faisal Shafait, Ajmal Mian |
WACV | 3 |
| 2014 | Unsupervised iterative manifold alignment via local feature histogramsabstractWe propose a new unsupervised algorithm for the automatic alignment of two manifolds of different datasets with possibly different dimensionalities. Alignment is performed automatically without any assumptions on the correspondences between the two manifolds. The proposed algorithm automatically establishes an initial set of sparse correspondences between the two datasets by matching their underlying manifold structures. Local feature histograms are extracted at each point of the manifolds and matched using a robust algorithm to find the initial correspondences. Based on these sparse correspondences, an embedding space is estimated where the distance between the two manifolds is minimized while maximally retaining the original structure of the manifolds. The problem is formulated as a generalized eigenvalue problem and solved efficiently. Dense correspondences are then established between the two manifolds and the process is iteratively implemented until the two manifolds are correctly aligned consequently revealing their joint structure. We demonstrate the effectiveness of our algorithm on aligning protein structures, facial images of different subjects under pose variations and RGB and Depth data from Kinect. Comparison with an state-of-the-art algorithm shows the superiority of the proposed manifold alignment algorithm in terms of accuracy and computational time. Ajmal Mian, Wanquan Liu |
WACV | 2 |
| 2014 | Gradient based efficient feature selectionabstractSelecting a reduced set of relevant and non-redundant features for supervised classification problems is a challenging task. We propose a gradient based feature selection method which can search the feature space efficiently and select a reduced set of representative features. We test our proposed algorithm on five small and medium sized pattern classification datasets as well as two large 3D face datasets for computer vision applications. Comparison with the state of the art wrapper and filter methods shows that our proposed technique yields better classification results in lesser number of evaluations of the target classifier. The feature subset selected by our algorithm is representative of the classes in the data and has the least variation in classification accuracy. Syed Zulqarnain Gilani, Faisal Shafait, Ajmal Mian |
WACV | 3 |
| 2014 | Real time action recognition using histograms of depth gradients and random decision forestsabstractWe propose an algorithm which combines the discriminative information from depth images as well as from 3D joint positions to achieve high action recognition accuracy. To avoid the suppression of subtle discriminative information and also to handle local occlusions, we compute a vector of many independent local features. Each feature encodes spatiotemporal variations of depth and depth gradients at a specific space-time location in the action volume. Moreover, we encode the dominant skeleton movements by computing a local 3D joint position difference histogram. For each joint, we compute a 3D space-time motion volume which we use as an importance indicator and incorporate in the feature vector for improved action discrimination. To retain only the discriminant features, we train a random decision forest (RDF). The proposed algorithm is evaluated on three standard datasets and compared with nine state-of-the-art algorithms. Experimental results show that, on the average, the proposed algorithm outperform all other algorithms in accuracy and have a processing speed of over 112 frames/second. Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian |
WACV | 4 |
| 2014 | Fully automatic 3D facial expression recognition using local depth featuresabstractFacial expressions form a significant part of our nonverbal communications and understanding them is essential for effective human computer interaction. Due to the diversity of facial geometry and expressions, automatic expression recognition is a challenging task. This paper deals with the problem of person-independent facial expression recognition from a single 3D scan. We consider only the 3D shape because facial expressions are mostly encoded in facial geometry deformations rather than textures. Unlike the majority of existing works, our method is fully automatic including the detection of landmarks. We detect the four eye corners and nose tip in real time on the depth image and its gradients using Haar-like features and AdaBoost classifier. From these five points, another 25 heuristic points are defined to extract local depth features for representing facial expressions. The depth features are projected to a lower dimensional linear subspace where feature selection is performed by maximizing their relevance and minimizing their redundancy. The selected features are then used to train a multi-class SVM for the final classification. Experiments on the benchmark BU-3DFE database show that the proposed method outperforms existing automatic techniques, and is comparable even to the approaches using manual landmarks. Mingliang Xue, Ajmal Mian, Wanquan Liu, Ling Li 0006 |
WACV | 2 |
| 2014 | Recent Advances on Singlemodal and Multimodal Face Recognition: A SurveyabstractHigh performance for face recognition systems occurs in controlled environments and degrades with variations in illumination, facial expression, and pose. Efforts have been made to explore alternate face modalities such as infrared (IR) and 3-D for face recognition. Studies also demonstrate that fusion of multiple face modalities improve performance as compared with singlemodal face recognition. This paper categorizes these algorithms into singlemodal and multimodal face recognition and evaluates methods within each category via detailed descriptions of representative work and summarizations in tables. Advantages and disadvantages of each modality for face recognition are analyzed. In addition, face databases and system evaluations are also covered. Hailing Zhou, Ajmal Mian, Lei Wei 0002, Douglas C. Creighton, Mohammed Hossny, Saeid Nahavandi |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2013 | Hyperspectral Face Recognition using 3D-DCT and Partial Least SquaresabstractHyperspectral imaging offers new opportunities for inter-person facial discrimina-tion. However, compact and discriminative feature extraction from high dimensional hyperspectral image cubes is a challenging task. We propose a spatio-spectral feature extraction method based on the 3D Discrete Cosine Transform (3D-DCT). The 3D-DCT optimally compacts information in the low frequency coefficients. Therefore, we rep-resent each hyperspectral facial cube by a small number of low frequency DCT coef-ficients and formulate Partial Least Square (PLS) regression for accurate classification. The proposed algorithm is evaluated on three standard hyperspectral face databases. Ex-perimental results show that the proposed algorithm outperforms five current state of the art hyperspectral face recognition algorithms by a significant margin. 1 Arif Mahmood, Ajmal Mian |
BMVC | 3 |
| 2013 | Hyperspectral Imaging for Ink Mismatch DetectionabstractInk mismatch detection provides important clues to forensic document examiners by identifying whether a particular handwritten note was written with a specific pen, or to show that some part (e.g. signature) of a note is written with a different ink as compared to the rest of the note. In this paper, we show that a hyper spectral image (HSI) of handwritten notes can discriminate between inks that are visually similar in appearance. For this purpose, we develop the first ever hyper spectral image database of handwritten notes in various blue and black inks, comprising a total of 70 hyper spectral images each in 33 bands of the visible spectrum. In an unsupervised clustering scheme, the spectral responses of inks fall into separate clusters to allow segmentation of two different inks in a questioned document. The same method fails to segment inks correctly when applied to RGB scans of these documents, since the inks are very hard to distinguish in the visible spectral range. HSI overcomes the shortcomings of RGB and allows better discrimination between inks. We further evaluate which subset of bands from HSI is most useful for the purpose of ink mismatch detection. We hope that these findings will stimulate the use of HSI in document analysis research, especially for questioned document examination. Zohaib Khan 0001, Faisal Shafait, Ajmal Mian |
ICDAR | 3 |
| 2013 | 3D face recognition using topographic high-order derivativesabstractThis paper presents a novel feature, Topographic High-order Derivatives (THD) for 3D face recognition. THD is based on the high-order micro-pattern information extracted from face topography maps. Face topography maps are partitioned into polar sectors, and THDs are computed using directional highorder derivatives within the sectors. Local features are extracted by encoding directional high-order derivatives within polar neighborhoods. To evaluate the proposed method, we use Bosphorus and FRGC 3D face databases which include pose and expression changes. The performance of the proposed method is higher compared to the state-of-the-art benchmark approaches in 3D face recognition. Ali Cheraghian, Farshid Hajati, Ajmal Mian, Yongsheng Gao 0001, Soheila Gheisari |
ICIP | 3 |
| 2013 | Using Kinect for face recognition under varying poses, expressions, illumination and disguiseabstractWe present an algorithm that uses a low resolution 3D sensor for robust face recognition under challenging conditions. A preprocessing algorithm is proposed which exploits the facial symmetry at the 3D point cloud level to obtain a canonical frontal view, shape and texture, of the faces irrespective of their initial pose. This algorithm also fills holes and smooths the noisy depth data produced by the low resolution sensor. The canonical depth map and texture of a query face are then sparse approximated from separate dictionaries learned from training data. The texture is transformed from the RGB to Discriminant Color Space before sparse coding and the reconstruction errors from the two sparse coding steps are added for individual identities in the dictionary. The query face is assigned the identity with the smallest reconstruction error. Experiments are performed using a publicly available database containing over 5000 facial images (RGB-D) with varying poses, expressions, illumination and disguise, acquired using the Kinect sensor. Recognition rates are 96.7% for the RGB-D data and 88.7% for the noisy depth data alone. Our results justify the feasibility of low resolution 3D sensors for robust face recognition. Billy Y. L. Li, Ajmal Mian, Wanquan Liu, Aneesh Krishna |
WACV | 2 |
| 2013 | Periocular biometric recognition using image setsabstractHuman identification based on iris biometrics requires high resolution iris images of a cooperative subject. Such images cannot be obtained in non-intrusive applications such as surveillance. However, the full region around the eye, known as the periocular region, can be acquired non-intrusively and used as a biometric. In this paper we investigate the use of periocular region for person identification. Current techniques have focused on choosing a single best frame, mostly manually, for matching. In contrast, we formulate, for the first time, person identification based on periocular regions as an image set classification problem. We generate periocular region image sets from the Multi Bio-metric Grand Challenge (MBGC) NIR videos. Periocular regions of the right eyes are mirrored and combined with those of the left eyes to form an image set. Each image set contains periocular regions of a single subject. For imageset classification, we use six state-of-the-art techniques and report their comparative recognition and verification performances. Our results show that image sets of periocular regions achieve significantly higher recognition rates than currently reported in the literature for the same database. Arif Mahmood, Ajmal Mian, Chris McDonald |
WACV | 3 |
| 2013 | Multibiometric human recognition using 3D ear and face features
Syed M. S. Islam, Rowan Davies, Mohammed Bennamoun, Robyn A. Owens, Ajmal Mian |
Pattern Recognit. | 5 |
| 2013 | Image Set Based Face Recognition Using Self-Regularized Non-Negative Coding and Adaptive Distance Metric LearningabstractSimple nearest neighbor classification fails to exploit the additional information in image sets. We propose self-regularized nonnegative coding to define between set distance for robust face recognition. Set distance is measured between the nearest set points (samples) that can be approximated from their orthogonal basis vectors as well as from the set samples under the respective constraints of self-regularization and nonnegativity. Self-regularization constrains the orthogonal basis vectors to be similar to the approximated nearest point. The nonnegativity constraint ensures that each nearest point is approximated from a positive linear combination of the set samples. Both constraints are formulated as a single convex optimization problem and the accelerated proximal gradient method with linear-time Euclidean projection is adapted to efficiently find the optimal nearest points between two image sets. Using the nearest points between a query set and all the gallery sets as well as the active samples used to approximate them, we learn a more discriminative Mahalanobis distance for robust face recognition. The proposed algorithm works independently of the chosen features and has been tested on gray pixel values and local binary patterns. Experiments on three standard data sets show that the proposed method consistently outperforms existing state-of-the-art methods. Ajmal Mian, Yiqun Hu, Richard I. Hartley, Robyn A. Owens |
IEEE Trans. Image Process. | 1 |
| 2012 | Hierarchical Sparse Spectral Clustering For Image Set ClassificationabstractWe present a structural matching technique for robust classification based on image sets. In set based classification, a probe set is matched with a number of gallery sets and assigned the label of the most similar set. We represent each image set by a sparse dictionary and compute a similarity matrix by matching all the dictionary atoms of the gallery and probe sets. The similarity matrix comprises the sparse coding coefficients and forms a fully connected directed graph. The nodes of the graph are the dictionary atoms and the edges are the sparse coefficients. The graph is converted to an undirected graph with positive edge weights and spectral clustering is used to cut the graph into two balanced partitions using the normalized cut algorithm. This process is repeated until the graph reduces to critical and non-critical partitions. A critical partition contains atoms with the same gallery label along with one or more probe atoms whereas a non-critical partition either consists of only probe atoms or atoms with multiple gallery labels with no probe atom. Using the critical partitions, we define a novel set based similarity measure and assign the probe set the label of the gallery set with maximum similarity. The proposed algorithm is applied to image set based face recognition using two standard databases. Comparison with existing techniques shows the validity and robustness of our algorithm in the presence of outlier images. Arif Mahmood, Ajmal Mian |
BMVC | 2 |
| 2012 | Face Recognition Using Sparse Approximated Nearest Points between Image SetsabstractWe propose an efficient and robust solution for image set classification. A joint representation of an image set is proposed which includes the image samples of the set and their affine hull model. The model accounts for unseen appearances in the form of affine combinations of sample images. To calculate the between-set distance, we introduce the Sparse Approximated Nearest Point (SANP). SANPs are the nearest points of two image sets such that each point can be sparsely approximated by the image samples of its respective set. This novel sparse formulation enforces sparsity on the sample coefficients and jointly optimizes the nearest points as well as their sparse approximations. Unlike standard sparse coding, the data to be sparsely approximated are not fixed. A convex formulation is proposed to find the optimal SANPs between two sets and the accelerated proximal gradient method is adapted to efficiently solve this optimization. We also derive the kernel extension of the SANP and propose an algorithm for dynamically tuning the RBF kernel parameter while matching each pair of image sets. Comprehensive experiments on the UCSD/Honda, CMU MoBo, and YouTube Celebrities face datasets show that our method consistently outperforms the state of the art. Yiqun Hu, Ajmal Mian, Robyn A. Owens |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Spatially Optimized Data-Level Fusion of Texture and Shape for Face RecognitionabstractData-level fusion is believed to have the potential for enhancing human face recognition. However, due to a number of challenges, current techniques have failed to achieve its full potential. We propose spatially optimized data/pixel-level fusion of 3-D shape and texture for face recognition. Fusion functions are objectively optimized to model expression and illumination variations in linear subspaces for invariant face recognition. Parameters of adjacent functions are constrained to smoothly vary for effective numerical regularization. In addition to spatial optimization, multiple nonlinear fusion models are combined to enhance their learning capabilities. Experiments on the FRGC v2 data set show that spatial optimization, higher order fusion functions, and the combination of multiple such functions systematically improve performance, which is, for the first time, higher than score-level fusion in a similar experimental setup. Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
IEEE Trans. Image Process. | 3 |
| 2011 | Sparse approximated nearest points for image set classificationabstractClassification based on image sets has recently attracted great research interest as it holds more promise than single image based classification. In this paper, we propose an efficient and robust algorithm for image set classification. An image set is represented as a triplet: a number of image samples, their mean and an affine hull model. The affine hull model is used to account for unseen appearances in the form of affine combinations of sample images. We introduce a novel between-set distance called Sparse Approximated Nearest Point (SANP) distance. Unlike existing methods, the dissimilarity of two sets is measured as the distance between their nearest points, which can be sparsely approximated from the image samples of their respective set. Different from standard sparse modeling of a single image, this novel sparse formulation for the image set enforces sparsity on the sample coefficients rather than the model coefficients and jointly optimizes the nearest points as well as their sparse approximations. A convex formulation for searching the optimal SANP between two sets is proposed and the accelerated proximal gradient method is adapted to efficiently solve this optimization. Experimental evaluation was performed on the Honda, MoBo and Youtube datasets. Comparison with existing techniques shows that our method consistently achieves better results. Yiqun Hu, Ajmal Mian, Robyn A. Owens |
CVPR | 2 |
| 2011 | Contour Code: Robust and efficient multispectral palmprint encoding for human recognitionabstractWe propose `Contour Code', a novel representation and binary hash table encoding for multispectral palmprint recognition. We first present a reliable technique for the extraction of a region of interest (ROI) from palm images acquired with non-contact sensors. The Contour Code representation is then derived from the Nonsubsampled Contourlet Transform. A uniscale pyramidal filter is convolved with the ROI followed by the application of a directional filter bank. The dominant directional subband establishes the orientation at each pixel and the index corresponding to this subband is encoded in the Contour Code representation. Unlike existing representations which extract orientation features directly from the palm images, the Contour Code uses a two stage filtering to extract robust orientation features. The Contour Code is binarized into an efficient hash table structure that only requires indexing and summation operations for simultaneous one-to-many matching with an embedded score level fusion of multiple bands. We quantitatively evaluate the accuracy of the ROI extraction by comparison with a manually produced ground truth. Multispectral palmprint verification results on the PolyU and CASIA databases show that the Contour Code achieves an EER reduction upto 50%, compared to state-of-the-art methods. Zohaib Khan 0001, Ajmal Mian, Yiqun Hu |
ICCV | 2 |
| 2011 | Robust realtime feature detection in raw 3D face imagesabstract3D face data contains holes, spikes and significant noise which must be removed before any further operations such as feature detection or face recognition can be performed. Removing these anomalies from the complete data is expensive as it also contains non-facial regions. We present a realtime algorithm that can detect the eyes and the nose tip in raw 3D face images in about 210 msecs. With three points, the data can be aligned to a canonical pose or registered to a reference face allowing the face area to be accurately cropped. The more expensive preprocessing steps can then be applied to the cropped region of the face only. We calculate the x and y gradients from the range image and train separate feature detectors in the three representations. Each detector is trained using the AdaBoost algorithm and Haar-like features. Haar features detect higher order discontinuities in the gradient images which form the core of the proposed algorithm. Multiple feature detections in the three images are clustered and anthropometric ratios are used to eliminate outliers. The centroids of the remaining candidates are used as feature points. Experimental results on the FRGC v2 database gave over 99% detection rates. Detailed quantitative analysis and comparison with the ground truth feature locations is provided. Ajmal Mian |
WACV | 1 |
| 2011 | Efficient Detection and Recognition of 3D Ears
Syed M. S. Islam, Rowan Davies, Mohammed Bennamoun, Ajmal Mian |
Int. J. Comput. Vis. | 4 |
| 2011 | Illumination normalization of facial images by reversing the process of image formation
Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
Mach. Vis. Appl. | 3 |
| 2011 | Online learning from local features for video-based face recognition
Ajmal Mian |
Pattern Recognit. | 1 |
| 2011 | A training-free nose tip detection method from face range images
Xiaoming Peng, Mohammed Bennamoun, Ajmal Mian |
Pattern Recognit. | 3 |
| 2011 | Correlation based speech-video synchronization
Amar A. El-Sallam, Ajmal Mian |
Pattern Recognit. Lett. | 2 |
| 2010 | Face Recognition Using Contourlet Transform and Multidirectional Illumination from a Computer Screen
Ajmal Mian |
ACIVS (2) | 1 |
| 2010 | A Novel Analytical Approach for Lip Synchronization
Amar A. El-Sallam, Ajmal Mian |
ICIP | 2 |
| 2010 | On the Repeatability and Quality of Keypoints for Local Feature-based 3D Object Retrieval from Cluttered Scenes
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
Int. J. Comput. Vis. | 1 |
| 2009 | Comparison of Visible, Thermal Infra-Red and Range Images for Face Recognition
Ajmal Mian |
PSIVT | 1 |
| 2009 | An Expression Deformation Approach to Non-rigid 3D Face Recognition
Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
Int. J. Comput. Vis. | 3 |
| 2008 | A Fast and Fully Automatic Ear Recognition Approach Based on 3D Local Surface Features
Syed M. S. Islam, Rowan Davies, Ajmal Mian, Mohammed Bennamoun |
ACIVS | 3 |
| 2008 | Unsupervised learning from local features for video-based face recognition
Ajmal Mian |
FG | 1 |
| 2008 | Keypoint Detection and Local Feature Matching for Textured 3D Face Recognition
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
Int. J. Comput. Vis. | 1 |
| 2008 | Integration of local and global geometrical cues for 3D face recognition
Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
Pattern Recognit. | 3 |
| 2007 | Interest-point Based Face Recognition from Range ImagesabstractWe present a novel approach to interest-point detection tailored to range images. A range image is represented by two images with blob-like patterns that have easily detectable peaks and can be efficiently extracted using convolution kernels. These kernels were designed to produce repeatable and independent blob-like patterns when convolved with the range image. The interest-points correspond to peaks of the patterns after dropping the unstable ones and performing Non-Maximal Suppression (NMS) on their union. The approach was applied to facial range images from the FRGC V2.0 dataset and about 88% repeatability was achieved. Face recognition was also performed by matching the local range regions around the interest-points. An approach based on three levels of matching combined with RAN SAC algorithm was used to increase the correct matches and reduce the false ones. Preliminary recognition results for a database of 466 subjects and 1765 probes were 96.33% identification rate and 90% verification rate at 0.1% False Accept Rate (FAR) for faces under neutral expression. Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
BMVC | 3 |
| 2007 | An Efficient Multimodal 2D-3D Hybrid Approach to Automatic Face RecognitionabstractWe present a fully automatic face recognition algorithm and demonstrate its performance on the FRGC v2.0 data. Our algorithm is multimodal (2D and 3D) and performs hybrid (feature-based and holistic) matching in order to achieve efficiency and robustness to facial expressions. The pose of a 3D face along with its texture is automatically corrected using a novel approach based on a single automatically detected point and the Hotelling transform. A novel 3D Spherical Face Representation (SFR) is used in conjunction with the SIFT descriptor to form a rejection classifier which quickly eliminates a large number of candidate faces at an early stage for efficient recognition in case of large galleries. The remaining faces are then verified using a novel region-based matching approach which is robust to facial expressions. This approach automatically segments the eyes-forehead and the nose regions, which are relatively less sensitive to expressions, and matches them separately using a modified ICP algorithm. The results of all the matching engines are fused at the metric level to achieve higher accuracy. We use the FRGC benchmark to compare our results to other algorithms which used the same database. Our multimodal hybrid algorithm performed better than others by achieving 99.74% and 98.31% verification rates at 0.001 FAR and identification rates of 99.02% and 95.37% for probes with neutral and non-neutral expression respectively. Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | 2D and 3D Multimodal Hybrid Face Recognition
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
ECCV (3) | 1 |
| 2006 | A Novel Representation and Feature Matching Algorithm for Automatic Pairwise Registration of Range Images
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
Int. J. Comput. Vis. | 1 |
| 2006 | Three-Dimensional Model-Based Object Recognition and Segmentation in Cluttered ScenesabstractViewpoint independent recognition of free-form objects and their segmentation in the presence of clutter and occlusions is a challenging task. We present a novel 3D model-based algorithm which performs this task automatically and efficiently. A 3D model of an object is automatically constructed offline from its multiple unordered range images (views). These views are converted into multidimensional table representations (which we refer to as tensors). Correspondences are automatically established between these views by simultaneously matching the tensors of a view with those of the remaining views using a hash table-based voting scheme. This results in a graph of relative transformations used to register the views before they are integrated into a seamless 3D model. These models and their tensor representations constitute the model library. During online recognition, a tensor from the scene is simultaneously matched with those in the library by casting votes. Similarity measures are calculated for the model tensors which receive the most votes. The model with the highest similarity is transformed to the scene and, if it aligns accurately with an object in the scene, that object is declared as recognized and is segmented. This process is repeated until the scene is completely segmented. Experiments were performed on real and synthetic data comprised of 55 models and 610 scenes and an overall recognition rate of 95 percent was achieved. Comparison with the spin images revealed that our algorithm is superior in terms of recognition rate and efficiency. Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Region-based Matching for Robust 3D Face RecognitionabstractWe present a novel region-based matching approach for automatic 3D face recognition which is robust to facial expressions, facial hair, illumination changes and large occlusions. Each 3D face in the gallery is segmented offline into three disjoint regions, namely eyes-forehead, nose and cheeks. Recognition is performed on the basis of only the eyes-forehead and nose regions to avoid the effects of expressions and artifacts that occur in 3D faces due to a mustache or beard. These two regions of the gallery are matched with a probe using a modified version of the ICP algorithm and their matching scores are fused. The identity of the gallery face which gets the highest score is declared as the identity of the probe. Experiments were performed on the UND Biometrics Database which is so far the largest known database of 3D faces. We achieved a combined identification rate of 100% and a maximum verification rate of 99.42%. Our results also show that the eyes-forehead is the most significant region for 3D face recognition with individual identification and verification rates of 97.32% and 97.25% respectively. Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
BMVC | 1 |
| 2004 | Matching Tensors for Automatic Correspondence and Registration
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
ECCV (2) | 1 |
| 2004 | Performance analysis of an improved tensor based correspondence algorithm for automatic 3d modeling
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
ICIP | 1 |