VLDB 2026 Research / reviewers in the wild / expert
Mohammed Bennamoun
dblp:00/3214
· DBLP profile ↗
306ranked-venue papers
9as first author
112since 2021 · last 2026
0000-0002-6603-3257ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 185 · 3 first-author · 74 since 2021Graphics, computer vision, multimedia, augmented reality and games · 133 · 4 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 2 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Security and privacy · 7 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional SportsabstractAbstract Reasoning over sports videos for question answering is an important task with numerous applications, such as player training and information retrieval. However, this task has not been explored due to the lack of relevant datasets and the challenging nature it presents. Most datasets for video question answering (VideoQA) focus mainly on general and coarse-grained understanding of daily-life videos, which is not applicable to sports scenarios requiring professional action understanding and fine-grained motion analysis. In this paper, we introduce the first dataset, named Sports-QA, specifically designed for the sports VideoQA task. The Sports-QA dataset includes various types of questions, such as descriptions, chronologies, causalities, and counterfactual conditions, covering multiple sports. Furthermore, to address the characteristics of the sports VideoQA task, we propose a new Auto-Focus Transformer (AFT) capable of automatically focusing on particular scales of temporal information for question answering. We conduct extensive experiments on Sports-QA, including baseline studies and the evaluation of different methods. The results demonstrate that our AFT achieves state-of-the-art performance. Haopeng Li 0001, Andong Deng, Jun Liu 0036, Hossein Rahmani 0001, Yulan Guo, Bernt Schiele, Mohammed Bennamoun, Qiuhong Ke |
Int. J. Comput. Vis. | 7 |
| 2026 | Semantically-aware Neural Radiance Fields for Visual Scene Understanding: A Comprehensive Review
Thang-Anh-Quan Nguyen, Amine Bourki, Mátyás Macudzinski, Anthony Brunel, Mohammed Bennamoun |
Int. J. Comput. Vis. | 5 |
| 2026 | Unleashing the Power of Text-to-Image Diffusion Models for Category-Agnostic Pose EstimationabstractCategory-Agnostic Pose Estimation (CAPE) aims to detect keypoints of unseen object categories in a few-shot setting, where the scarcity of labeled data poses significant challenges to generalization. In this work, we propose Prompt Pose Matching (PPM), a novel framework that unleashes the power of off-the-shelf text-to-image diffusion models for CAPE. PPM learns pseudo prompts from few-shot examples via the text-to-image diffusion model. These learned pseudo prompts capture semantic information of keypoints, which can then be used to locate the same type of keypoints from images. To provide prompts with representative initialization, we introduce a category-agnostic pre-training strategy to capture the foreground prior shared across categories and keypoints. To support the reliable prompt pre-training, we propose a Foreground-Aware Region Aggregation (FARA) module to provide robust and consistent supervision signal. Based on the foreground prior, a Foreground-Guided Attention Refinement (FGAR) module is further proposed to reinforce cross-attention responses for accurate keypoint localization. For efficiency, a Prompt Ensemble Inference (PEI) scheme enables joint keypoint prediction. Unlike previous methods that highly rely on base-category annotated data, our PPM framework can operate in a base-category-free setting while retaining strong performance. Code will be available at: https://github.com/DuoPeng-CVer/Prompt-Pose-Matching. Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, De Wen Soh, Mohammed Bennamoun, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Triple Spectral Fusion for Sensor-Based Human Activity RecognitionabstractThe field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identify daily activities. Despite the advancements in learning-based methods, it is challenging to perform information fusion from the temporal perspective due to the complexities in fusing heterogeneous sensor data and establishing long-term context correlations. This paper proposes a novel triple spectral fusion framework tailored for HAR. First, we develop an adaptive complementary filtering technique for noise suppression and organize each IMU's sensors into posture and motion modality nodes. Given that IMU nodes form a dynamic heterogeneous graph, we then apply adaptive filtering within the graph Fourier domain to merge both homogeneous and heterogeneous node information. Furthermore, an adaptive wavelet frequency selection approach is implemented to suppress context redundancy and shorten the length of features. This approach enhances both timestamp-based graph aggregation and the correlation of long-term contexts. Our framework uses adaptive filtering in the Fourier, graph Fourier, and wavelet domains, enabling effective multi-sensor fusion and context correlation. Extensive experiments on ten benchmark datasets demonstrate the superior performance of our framework. Ye Zhang 0037, Longguang Wang, Qing Gao 0002, Chaocan Xiang, Mohammed Bennamoun, Yulan Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action RecognitionabstractZero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton features with static, class-level semantics. This coarse-grained alignment fails to bridge the domain shift between seen, unseen classes, thereby impeding the effective transfer of fine-grained visual knowledge. To address these limitations, we introduce DynaPURLS, a unified framework that establishes robust, multi-scale visual-semantic correspondences, dynamically refines them at inference time to enhance generalization. Our framework leverages a large language model to generate hierarchical textual descriptions that encompass both global movements, local body-part dynamics. Concurrently, an adaptive partitioning module produces fine-grained visual representations by semantically grouping skeleton joints. To fortify this fine-grained alignment against the train-test domain shift, DynaPURLS incorporates a dynamic refinement module. During inference, this module adapts textual features to the incoming visual stream via a lightweight learnable projection. This refinement process is stabilized by a confidence-aware, class-balanced memory bank, which mitigates error propagation from noisy pseudo-labels. Extensive experiments on three large-scale benchmark datasets, including NTU RGB+D 60/120, PKU-MMD, demonstrate that DynaPURLS significantly outperforms prior art, setting new state-of-the-art records. Jingmin Zhu, James Bailey 0001, Jun Liu 0036, Hossein Rahmani 0001, Mohammed Bennamoun, Farid Boussaïd, Qiuhong Ke |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | HAViG: Hierarchical adaptive visual grounding framework for video question answering
Lei Zhu 0005, Lingmin Pan, Siqiao Tan, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Farid Boussaïd, Mohammed Bennamoun |
Pattern Recognit. | 8 |
| 2026 | Pattern classification in unseen environments with incomplete multi-source domain generalization
Zhunga Liu, Mohammed Bennamoun |
Signal Process. | 4 |
| 2026 | SPA: Stable and Precise Alignment for Efficient Cross-Domain Palmprint RecognitionabstractPalmprint recognition has been extensively studied as an effective biometric technique for personal identification. With the rapid development of deep neural networks (DNNs), palmprint recognition methods have achieved remarkable progress. However, their performance often deteriorates significantly under domain shifts. Moreover, existing unsupervised domain adaptation approaches for palmprint recognition typically suffer from unstable training and imprecise feature alignment, thereby limiting their effectiveness. To address these challenges, we propose SPA, a Stable and Precise Alignment framework for cross-domain palmprint recognition. Specifically, we design a lightweight yet robust Style Transformation Module (STM) to mitigate variations in style, color, and illumination. With the aid of STM, we further align joint feature distributions across all high-level layers, achieving more accurate feature alignment and enhancing recognition robustness. We conduct extensive experiments on two public multi-domain palmprint databases encompassing 42 cross-domain scenarios. The results demonstrate that SPA consistently delivers superior performance across both databases, achieving higher recognition accuracy with lower computational overhead compared to existing methods. In particular, SPA improves the average identification accuracies to 94.21% and 81.93%, while reducing the average equal error rates (EER) to 1.36% and 3.62% on the two databases, respectively. Song Ruan, Yantao Li 0001, Huafeng Qin, Naeha Sharif, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | InfoARD: Enhancing Adversarial Robustness Distillation With Attack-Strength Adaptation and Mutual-Information MaximizationabstractAdversarial distillation (AD) aims to mitigate deep neural networks' inherent vulnerability to adversarial attacks, thereby providing robust protection for compact models through teacher-student interactions. Despite advancements, existing AD studies still suffer from insufficient robustness due to the limitations of fixed attack strength and attention region shifts. To address these challenges, we propose a strength-adaptive Info-maximizing Adversarial Robustness Distillation paradigm, namely "InfoARD", which strategically incorporates the Attack-Strength Adaptation (ASA) and Mutual-Information Maximization (MIM) to enhance adversarial robustness against adversarial attacks and perturbations. Unlike previous adversarial training (AT) methods that utilize fixed attack strength, the ASA mechanism is designed to capture smoother and generalized classification boundaries by dynamically tailoring the attack strength based on the characteristics of individual instances. Benefiting from mutual information constraints, our MIM strategy ensures the student model effectively learns from various levels of feature representations and attention patterns, thereby deepening the student model's understanding of the teacher model's decision-making processes. Furthermore, a comprehensive multi-granularity distillation is conducted to capture knowledge across multiple dimensions, enabling a more effective transfer of knowledge from the teacher model to the student model. Note that our InfoARD can be seamlessly integrated into existing AD frameworks, further boosting the adversarial robustness of deep learning models. Extensive experiments on various challenging datasets consistently demonstrate the effectiveness and robustness of our InfoARD, surpassing previous state-of-the-art methods. Ruihan Liu, Jieyi Cai, Yishu Liu 0001, Sudong Cai, Bingzhi Chen, Yulan Guo, Mohammed Bennamoun |
IEEE Trans. Image Process. | 7 |
| 2026 | RLAD: A Reliable Hippo-Guided Multi-Task Model for Alzheimer's Disease DiagnosisabstractEarly diagnosis of Alzheimer's disease (AD) is crucial for its prevention, and hippocampal atrophy is a significant lesion for early diagnosis. The current DL-based AD diagnosis methods only focus on either AD classification or hippocampus segmentation independently, neglecting the correlation between the two tasks and lacking pathological interpretability. To address this issue, we propose a Reliable Hippo-guided Learning model for Alzheimer's Disease diagnosis (RLAD), which employs multi-task learning for AD classification as a main task supplemented by hippocampus segmentation. More specifically, our model consists of 1) a hybrid shared features encoder that encodes local and global information in MRI to enhance the model's ability to learn discriminative features; 2) Task Specific Decoders to accomplish AD classification and hippocampus segmentation; and 3) Task Coordination module to correlate the two tasks and guide the classification task to focus on the hippocampus area. Our proposed RLAD model is evaluated on MRI scans of 1631 subjects from three independent datasets, including ADNI-1, ADNI-2, and HarP. Our extensive experimental results demonstrate that the proposed model significantly improves the performance of AD classification and hippocampus segmentation with strong generalization capabilities. Zhenxin Lei, Cong Hua, Johann Li, Syed Afaq Ali Shah, Liang Zhang 0010, Mohammed Bennamoun, Cuiping Mao |
IEEE J. Biomed. Health Informatics | 8 |
| 2026 | Toward Bidirectional Adaptability for Few-Shot Class-Incremental Learning With Forward-Backward Knowledge TransferabstractThe development of Deep Neural Networks (DNNs) has enabled AI-driven models to excel in recognizing a limited set of classes within static environments. As AI systems progress, few-shot class-incremental learning (FSCIL) aims to expand their understanding of novel classes from minimal samples while retaining knowledge of previously encountered ones. However, most existing FSCIL models face significant challenges, includinginadequate adaptabilityandcatastrophic forgetting, which hinder their ability to maintain robust forward and backward learning capabilities. To address these issues, this paper proposes a novel Forward-Backward Knowledge Transfer (FBKT) paradigm, which strategically integrates forward distribution adaptation (FDA) and backward semantic alignment (BSA) mechanisms to achieve bidirectional adaptability in knowledge transfer. The FDA mechanism enhances forward adaptability by expanding and reserving the embedding space for new classes using semantic-irrelevant masked images as virtual negative classes, thereby mitigating data overfitting. It also employs self-supervised representation learning to utilize semantic-relevant local embeddings as additional positive samples, fostering class separation and generalization. Meanwhile, the BSA mechanism ensures the semantic consistency of previously learned classes across sessions during class-incremental learning, promoting smoother backward adaptability and reducing model degradation. Extensive experiments conducted on multiple benchmark datasets consistently highlight the superior performance and effectiveness of our FBKT compared to state-of-the-art methods. Bingzhi Chen, Sudong Cai, Xiaozhao Fang, Mohammed Bennamoun, Shengli Xie 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Admitting Ignorance Helps the Video Question Answering Models to AnswerabstractSignificant progress has been made in the field of video question answering (VideoQA) thanks to deep learning and large-scale pretraining. Despite the presence of sophisticated model structures and powerful video-text foundation models, most existing methods focus solely on maximizing the correlation between answers and video-question pairs during training. We argue that these models often establish shortcuts, resulting in spurious correlations between questions and answers, especially when the alignment between video and text data is suboptimal. To address these spurious correlations, we propose a novel training framework in which the model is compelled to acknowledge its ignorance when presented with an intervened question, rather than making guesses solely based on superficial question-answer correlations. We introduce methodologies for intervening in questions, utilizing techniques such as displacement and perturbation, and design frameworks for the model to admit its lack of knowledge in both multi-choice VideoQA and open-ended settings. In practice, we integrate a state-of-the-art model into our framework to validate its effectiveness. The results clearly demonstrate that our framework can significantly enhance the performance of VideoQA models with minimal structural modifications. Haopeng Li 0001, Tom Drummond, Mingming Gong, Mohammed Bennamoun, Qiuhong Ke |
IEEE Trans. Multim. | 4 |
| 2026 | AquaticCLIP: A Vision-Language Foundation Model and Dataset for Underwater Scene AnalysisabstractThe preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role in aiding marine scientists in their decision-making processes. In this article, we introduce AquaticCLIP, a novel contrastive language-image pretraining (CLIP) model tailored for aquatic scene understanding. AquaticCLIP presents an underwater domain-specific learning framework that aligns images and texts in aquatic environments, enabling tasks such as segmentation, classification, detection, and object counting. By leveraging our large-scale underwater image-text paired dataset without the need for ground-truth (GT) annotations, our model enriches existing vision-language models (VLMs) in the aquatic domain. For this purpose, we construct a 2-million underwater image-text paired dataset using heterogeneous resources, including YouTube, Netflix, National Geographic (NatGeo), etc. To fine-tune AquaticCLIP, we propose a prompt-guided vision encoder (PGVE) that progressively aggregates patch features via learnable prompts, while a vision-guided mechanism enhances the language encoder by incorporating visual context. The model is optimized through a contrastive pretraining loss to align visual and textual modalities. AquaticCLIP achieves notable performance improvements in zero-shot settings across multiple underwater computer vision tasks, outperforming existing methods in both accuracy and robustness. Our model sets a new benchmark for vision-language applications in underwater environments. The code and dataset for AquaticCLIP are publicly available on GitHub at: https://github.com/BasitAlawode/AquaticCLIP. Basit Alawode, Iyyakutti Iyappan Ganapathi, Sajid Javed, Mohammed Bennamoun, Arif Mahmood |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual RepresentationabstractIn Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tissue phenotyping, and survival analysis due to the limited level of detail that a single-resolution image can provide. Addressing this, we propose a novel multi-resolution paradigm leveraging Whole Slide Images (WSIs) to extract histology patches at multiple resolutions and generate corresponding textual descriptions through advanced CPath VLM. We introduce visual-textual alignment at multiple resolutions as well as cross-resolution alignment to establish more effective text-guided visual representations. Cross-resolution alignment using a multi-modal encoder enhances the model’s ability to capture context from multiple resolutions in histology images. Our model aims to capture a broader range of information, supported by novel loss functions, enriches feature representation, improves discriminative ability, and enhances generalization across different resolutions. Pre-trained on a comprehensive TCGA dataset with 34 million image-language pairs at various resolutions, our fine-tuned model outperforms State-Of-The-Art (SOTA) counterparts across multiple datasets and tasks, demonstrating its effectiveness in CPath. The code is available on GitHub at: https://github.com/BasitAlawode/MR-PLIP. Shahad Albastaki, Anabia Sohail, Iyyakutti Iyappan Ganapathi, Basit Alawode, Asim Khan, Sajid Javed, Naoufel Werghi, Mohammed Bennamoun, Arif Mahmood |
CVPR | 8 |
| 2025 | Dynamic Neural Surfaces for Elastic 4D Shape Representation and AnalysisabstractWe propose a novel framework for the statistical analysis of genus-zero 4D surfaces, i.e., 3D surfaces that deform and evolve over time. This problem is particularly challenging due to the arbitrary parameterizations of these surfaces and their varying deformation speeds, necessitating effective spatiotemporal registration. Traditionally, 4D surfaces are discretized, in space and time, before computing their spatiotemporal registrations, geodesics, and statistics. However, this approach may result in suboptimal solutions and, as we demonstrate in this paper, is not necessary. In contrast, we treat 4D surfaces as continuous functions in both space and time. We introduce Dynamic Spherical Neural Surfaces (D-SNS), an efficient smooth and continuous spatiotemporal representation for genus-0 4D surfaces. We then demonstrate how to perform core 4D shape analysis tasks such as spatiotemporal registration, geodesics computation, and mean 4D shape estimation, directly on these continuous representations without upfront discretization and meshing. By integrating neural representations with classical Riemannian geometry and statistical shape analysis techniques, we provide the building blocks for enabling full functional shape analysis. We demonstrate the efficiency of the framework on 4D human and face datasets. The source code and additional results are available at https://4d-dsns.github.io/DSNS/. Awais Nizamani, Hamid Laga, Guanjin Wang, Farid Boussaïd, Mohammed Bennamoun, Anuj Srivastava |
CVPR | 5 |
| 2025 | STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security InspectionabstractAdvancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a closed-set paradigm with predefined labels. To address these challenges, we introduce STCray, the first multimodal X-ray baggage security dataset, comprising 46,642 image-caption paired scans across 21 threat categories, generated using an X-ray scanner for airport security. STCray is meticulously developed with our specialized protocol that ensures domain-aware, coherent captions, that lead to the multi-modal instruction following data in X-ray baggage security. This allows us to train a domain-aware visual AI assistant named STING-BEE that supports a range of vision-language tasks, including scene comprehension, referring threat localization, visual grounding, and visual question answering (VQA), establishing novel baselines for multi-modal learning in X-ray baggage security. Further, STING-BEE shows state-of-the-art generalization in cross-domain settings. Code, data, and models are available at https://divs1159.github.io/STING-BEE/. Divya Velayudhan, Abdelfatah Hassan Ahmed, Mohamad Alansari, Neha Gour, Abderaouf Behouch, Taimur Hassan, Syed Talal Wasim, Nabil Maalej, Muzammal Naseer, Juergen Gall, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
CVPR | 11 |
| 2025 | GeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingabstractRecent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS). The distinct overhead viewpoint, scale variation, and presence of small objects in high-resolution RS imagery present a unique challenge in region-level comprehension. Moreover, the development of the grounding conversation capability of LMMs within RS is hindered by the lack of granular, RS domain-specific grounded data. Addressing these limitations, we propose GeoPixel - the first end-to-end high-resolution RS-LMM that supports pixel-level grounding. This capability allows fine-grained visual perception by generating interleaved masks in conversation. GeoPixel supports up to 4K HD resolution in any aspect ratio, ideal for high-precision RS image analysis. To support the grounded conversation generation (GCG) in RS imagery, we curate a visually grounded dataset GeoPixelD through a semi-automated pipeline that utilizes set-of-marks prompting and spatial priors tailored for RS data to methodically control the data generation process. GeoPixel demonstrates superior performance in pixel-level comprehension, surpassing existing LMMs in both single-target and multi-target segmentation tasks. Our methodological ablation studies validate the effectiveness of each component in the overall architecture. Our code and data will be publicly released. Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad Shahbaz Khan, Salman Khan 0001 |
ICML | 3 |
| 2025 | Advancing RFI-Detection in Radio Astronomy with Liquid State MachinesabstractRadio Frequency Interference (RFI) from anthropogenic radio sources poses significant challenges to current and future radio telescopes. Contemporary approaches to detecting RFI treat the task as a semantic segmentation problem on radio telescope spectrograms. Typically, complex heuristic algorithms handle this task of ‘flagging’ in combination with manual labeling (in the most difficult cases). While recent machine-learning approaches have demonstrated high accuracy, they often fail to meet the stringent operational requirements of modern radio observatories. Owing to their inherently time-varying nature, spiking neural networks (SNNs) are a promising alternative method to RFI-detection by utilizing the time-varying nature of the spectrographic source data. In this work, we apply Liquid State Machines (LSMs), a class of spiking neural networks, to RFI-detection. We employ second-order Leaky Integrate-And-Fire (LiF) neurons, marking the first use of this architecture and neuron type for RFI-detection. We test three encoding methods and three increasingly complex readout layers, including a transformer decoder head, providing a hybrid of SNN and ANN techniques. Our methods extend LSMs beyond conventional classification tasks to fine-grained spatio-temporal segmentation. We train LSMs on simulated data derived from the Hydrogen Epoch of Reionization Array (HERA), a known benchmark for RFI-detection. Our model achieves a per-pixel accuracy of 98% and an F1-score of 0.743, demonstrating competitive performance on this highly challenging task. This work expands the sophistication of SNN techniques and architectures applied to RFI-detection, and highlights the effectiveness of LSMs in handling fine-grained, complex, spatio-temporal signal-processing tasks. Nicholas J. Pritchard, Andreas Wicenec, Mohammed Bennamoun, Richard Dodson |
IJCNN | 3 |
| 2025 | PMIL: A Topology Module to Improve MIL-based WSI ClassificationabstractDeep learning models have achieved remarkable success in pathology image analysis. However, they still face challenges in effectively modeling fine-grained, object-level features. Topological Data Analysis (TDA) has shown promise for addressing these issues but remains underexplored, particularly for whole-slide pathology applications. Additionally, the effectiveness of TDA has yet to be firmly established, as current studies largely use small-scale datasets. In this work, we address these gaps by introducing Persistent Homology in Multiple Instance Learning (PMIL), the first adaptable TDA-based module within the MIL framework. We validate our approach on a large-scale classification dataset, benchmarking against multiple state-of-the-art methods. Ahmad Obeid 0001, Anabia Sohail, Said Boumaraf, Xiabi Liu, Sajid Javed, Hasan Almarzouqi, Jorge Dias 0001, Mohammed Bennamoun, Naoufel Werghi, Ibrahim M. Elfadel |
ISCAS | 8 |
| 2025 | Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMabstractHumans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like “A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the audience nodding and applauding” requires simultaneous processing of visual, audio, and speech signals. However, existing models often struggle to effectively fuse and interpret audio information, limiting their capacity for comprehensive video temporal understanding. To address this, we present TriSense, a triple-modality large language model designed for holistic video temporal understanding through the integration of visual, audio, and speech modalities. Central to TriSense is a Query-Based Connector that adaptively reweights modality contributions based on the input query, enabling robust performance under modality dropout and allowing flexible combinations of available inputs. To support TriSense's multimodal capabilities, we introduce TriSense-2M, a high-quality dataset of over 2 million curated samples generated via an automated pipeline powered by fine-tuned LLMs. TriSense-2M includes long-form videos and diverse modality combinations, facilitating broad generalization. Extensive experiments across multiple benchmarks demonstrate the effectiveness of TriSense and its potential to advance multimodal video analysis. Zinuo Li, Yongxin Guo 0001, Mohammed Bennamoun, Farid Boussaïd, Girish Dwivedi, Luqi Gong, Qiuhong Ke |
NeurIPS | 4 |
| 2025 | Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time AdaptationabstractWe introduce Skeleton-Cache, the first training-free test-time adaptation framework for skeleton-based zero-shot action recognition (SZAR), aimed at improving model generalization to unseen actions during inference. Skeleton-Cache reformulates inference as a lightweight retrieval process over a non-parametric cache that stores structured skeleton representations, combining both global and fine-grained local descriptors. To guide the fusion of descriptor-wise predictions, we leverage the semantic reasoning capabilities of large language models (LLMs) to assign class-specific importance weights. By integrating these structured descriptors with LLM-guided semantic priors, Skeleton-Cache dynamically adapts to unseen actions without any additional training or access to training data. Extensive experiments on NTU RGB+D 60/120 and PKU-MMD II demonstrate that Skeleton-Cache consistently boosts the performance of various SZAR backbones under both zero-shot and generalized zero-shot settings. The code is publicly available at https://github.com/Alchemist0754/Skeleton-Cache. Jingmin Zhu, Hossein Rahmani 0001, Jun Liu 0036, Mohammed Bennamoun, Qiuhong Ke |
NeurIPS | 5 |
| 2025 | HOPE: A Memory-Based and Composition-Aware Framework for Zero-Shot Learning with Hopfield Network and Soft Mixture of ExpertsabstractCompositional Zero-Shot Learning (CZSL) has emerged as an essential paradigm in machine learning, aiming to overcome the constraints of traditional zero-shot learning by incorporating compositional thinking into its method-ology. Conventional zero-shot learning has difficulty managing unfamiliar combinations of seen and unseen classes because it depends on pre-defined class embeddings. In contrast, Compositional Zero-Shot Learning leverages the inherent hierarchies and structural connections among classes, creating new class representations by combining at-tributes, components, or other semantic elements. In our paper, we propose a novel framework that for the first time combines the Modern Hopfield Network with a Mixture of Experts (HOPE) to classify the compositions of previously unseen objects. Specifically, the Modern Hopfield Network creates a memory that stores label prototypes and identifies relevant labels for a given input image. Subsequently, the Mixture of Expert models integrates the image with the appropriate prototype to produce the final composition classi-fication. Our approach achieves SOTA performance on sev-eral benchmarks, including MIT-States and UT-Zappos. We also examine how each component contributes to improved generalization. Do Huu Dat, Po Yuan Mao, Tien Hoang Nguyen, Wray L. Buntine, Mohammed Bennamoun |
WACV | 5 |
| 2025 | Memory guided representation learning for cross-domain face anti-spoofing
Pengchao Deng, Zhiheng Fu, Shengjun Xu, Chenyang Ge, Farid Boussaïd, Mohammed Bennamoun |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | Generalized Closed-Form Formulae for Feature-Based Subpixel Alignment in Patch-Based MatchingabstractAbstract Patch-based matching is a technique meant to measure the disparity between pixels in a source and target image and is at the core of various methods in computer vision. When the subpixel disparity between the source and target images is required, the cost function or the target image has to be interpolated. While cost-based interpolation is easier to implement, multiple works have shown that image-based interpolation can increase the accuracy of the disparity estimate. In this paper we review closed-form formulae for subpixel disparity computation for one dimensional matching, e.g., rectified stereo matching, for the standard cost functions used in patch-based matching. We then propose new formulae to generalize to high-dimensional search spaces, which is necessary for unrectified stereo matching and optical flow. We also compare the image-based interpolation formulae with traditional cost-based formulae, and show that image-based interpolation brings a significant improvement over the cost-based interpolation methods for two dimensional search spaces, and small improvement in the case of one dimensional search spaces. The zero-mean normalized cross correlation cost function is found to be preferable for subpixel alignment. A new error model, based on very broad assumptions is outlined in the Supplementary Material to demonstrate why these image-based interpolation formulae outperform their cost-based counterparts and why the zero-mean normalized cross correlation function is preferable for subpixel alignement. Laurent Valentin Jospin, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun |
Int. J. Comput. Vis. | 4 |
| 2025 | Conditional plane-based multi-scene representation for novel view synthesisabstractThe method overview. The explicit representation on the right side represents the shared canonical space and the view space using 12 feature planes. The gray arrows indicate the feature projection from the canonical representation. The left side shows the deformation between the canonical space and the view space. Pairwise features (e.g., X Y − Z T ) from the canonical and view representations are aggregated to obtain the final feature vector. ⨂ indicates feature aggregation. Density and appearance decoders, which estimate the geometry and color, are conditioned on the scene’s latent s i . Existing explicit and implicit-explicit hybrid neural representations for novel view synthesis are scene-specific. In other words, they represent only a single scene and require retraining for every novel scene. Implicit scene-agnostic methods rely on large multilayer perception (MLP) networks conditioned on learned features. They are computationally expensive during training and rendering times. In contrast, we propose a novel plane-based representation that learns to represent multiple static and dynamic scenes during training and renders per-scene novel views during inference. The method consists of a deformation network, explicit feature planes, and a conditional decoder. Explicit feature planes are used to represent a time-stamped view space volume and a shared canonical volume across multiple scenes. The deformation network learns the deformations across shared canonical object space and time-stamped view space. The conditional decoder estimates the color and density of each scene constrained by a scene-specific latent code. We evaluated and compared the performance of the proposed representation on static (NeRF) and dynamic (Plenoptic videos) datasets. The results show that explicit planes combined with tiny MLPs can efficiently train multiple scenes simultaneously. The project page: https://anonpubcv.github.io/cplanes/ . • We present a novel multi-scene representation that uses twelve explicit feature planes. • The method uses encoder-less generalization to learn discriminative features for each scene. • We represent multiple dynamic scenes without relying on optical flow estimation. • The auto-decoded latent (the scene dimension) can interpolate between scenes. • It achieves state-of-the-art rendering results for both static and dynamic scenes. Uchitha Rajapaksha, Hamid Laga, Dean Diepeveen, Mohammed Bennamoun, Ferdous Sohel |
Neurocomputing | 4 |
| 2025 | Enhancing object recognition: The role of object knowledge decomposition and component-labeled datasets
Nuoye Xiong, Ning Wang 0047, Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 7 |
| 2025 | Information bottleneck-guided KNN contrastive hashing for unsupervised cross-modal retrievalabstractUnsupervised cross-modal hashing (UCMH) has emerged as a promising solution for scalable multi-modal retrieval without costly annotations. However, existing methods often rely on rigid pairwise contrastive learning and fixed-size neighborhood selection, which suffer from false negatives and semantic noise, respectively—limiting their ability to model complex semantic structures in open-world scenarios. In this paper, we propose a novel framework, I nformation B ottleneck-guided K NN C ontrastive H ashing ( IBKCH ), which introduces a flexible and semantically adaptive contrastive paradigm for UCMH. Specifically, we design an information-aware neighbor sampling strategy that integrates: (1) a Hard-negative and Soft-positive (HN-SP) mechanism to adaptively distinguish informative negatives and softly aggregate latent positives; (2) an information bottleneck loss to retain task-relevant semantics while suppressing redundancy; and (3) an entropy sparsity regularizer to mitigate noisy neighbor interference. Furthermore, we develop an adaptive KNN contrastive learning scheme that unifies intra-modal and inter-modal alignment, enabling robust and discriminative hash code learning. Extensive experiments on three benchmark datasets demonstrate that IBKCH consistently outperforms state-of-the-art methods, especially under noisy or semantically diverse conditions—highlighting its effectiveness and generalizability in real-world UCMH applications. Lei Zhu 0005, Zhengchang Yuan, Zeqian Yi, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Farid Boussaïd, Mohammed Bennamoun, Shichao Zhang 0001 |
Knowl. Based Syst. | 8 |
| 2025 | Non-Rigid Point Cloud Registration via Anisotropic Hybrid Field HarmonizationabstractCurrent point cloud registration algorithms struggle to effectively handle both deformations and occlusions simultaneously. Our manifold analysis reveals this limitation arises from the inaccurate modeling of the shape's underlying manifold and the lack of an effective optimization strategy for fragmented manifold structures. In this paper, we present AniSym-Net, a novel non-rigid registration framework designed to address near-isometric deformation registration in the presence of occlusions. To encode object's coarse topological properties and local geometric information, AniSym-Net introduces a novel anisotropic hybrid shape-motion deformation field. The effectiveness of the anisotropic hybrid shape-motion fields relies on both the holonomic constraints from the symplectic structure modeling in AniSym-Net and the motion-conditional cross-attention during fusion, which calibrates geometric features using velocity-boundary constrained point motion patterns. The harmonization of correspondences derived from anisotropic hybrid fields and those from motion-shape fields significantly mitigates registration errors and occlusions. This is achieved through the optimization of loop closures of cotangent bundles within the symplectic manifold framework. We conduct comprehensive evaluation across five popular benchmarks, namely CAPE, DT4D, SAPIEN, FAUST, and DeepDeform, to demonstrate our AniSym-Net's superior performance compared to the state-of-the-art methods. Code will be publicly available. Xuequan Lu, Mohammed Bennamoun, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Dual-Phase Framework for Few-Shot Hyperspectral Image Classification With Spatiospectral Masked Autoencoder and Episode TrainingabstractThis article introduces a two-phase learning approach for hyperspectral image (HSI) classification using few-shot learning (FSL). For the first phase, we present a novel spatiospectral masked autoencoder (ssMAE)—an advanced self-supervised learner. For the ssMAE backbone network, we designed a transformer encoder-decoder network, where we replaced the linear layer that is used as the initial feature embedding with a 3-D convolutional layer to better extract local spectral-spatial features from 3-D visible sub-patches. By tapping into vast unlabeled data, the ssMAE learns general HSI features. In the second phase, the ssMAE encoder is fine-tuned to extract discriminative features for classification using the few-shot labeled training samples. This is achieved through a unique hybrid episode learning method that integrates the ssMAE encoder in a prototypical network (PN). We innovate with a mix of global and local prototypes (combined global-local (CGL) prototype) to refine label predictions. This technique maximizes data usage, focuses on specific samples, and mitigates issues from subpar episodes. Tested on three HSI datasets, our approach outperforms alternative few-shot methods. The code will be made publicly available athttps://github.com/Weejaa04/SSMAE. Wijayanti Nurul Khotimah, Mohammed Bennamoun, Farid Boussaïd, Lian Xu, Ferdous Sohel |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | WSSIC-Net: Weakly-Supervised Semantic Instance Completion of 3D Point Cloud ScenesabstractSemantic instance completion aims to recover the complete 3D shapes of foreground objects together with their labels from a partial 2.5D scan of a scene. Previous works have relied on full supervision, which requires ground-truth annotations, in the form of bounding boxes and complete 3D objects. This has greatly limited their real-world application because the acquisition of ground-truth data is very costly and time-consuming. To address this bottleneck, we propose a Weakly-Supervised Semantic Instance Completion Network (WSSIC-Net), which learns real-world partial point cloud object completion without requiring the ground truth of complete 3D objects. Instead, WSSIC-Net leverages 3D ground-truth bounding boxes, partial objects of a raw scene, and unpaired synthetic 3D point clouds. More specifically, a 3D detector is used to encode partial point clouds into proposal features, which are then fed into two branches. The first branch uses fully supervised box prediction based on proposal features. The second branch, hereinafter called instance completion, leverages the proposal features as partial object features to achieve weakly-supervised instance completion. A Generative Adversarial Network (GAN) completes the partial features of the 2.5D foreground objects of real-world scenes using only unpaired but semantically-consistent complete synthetic point clouds. In our experiments, we demonstrate that the fully-supervised 3D detection and the weakly-supervised instance completion complement one another. The qualitative and quantitative evaluations on the ScanNet v2 dataset demonstrate that the proposed "weakly-supervised" approach consistently achieves comparable performance to the state-of-the-art "fully supervised" methods. Zhiheng Fu, Yulan Guo, Minglin Chen, Qingyong Hu, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Image Process. | 7 |
| 2025 | CompletionMamba: Taming State Space Model for Point Cloud CompletionabstractPoint cloud completion aims to reconstruct complete 3D shapes from partial scans. The long-range dependencies between points and shape perception are crucial for this task. While Transformers are effective due to their global processing ability, the quadratic complexity of their attention mechanism makes them unsuitable for long sequences when computational resources are constrained. As an alternative, State Space Models (SSMs) provide a memory-efficient solution for handling long-range dependencies, yet applying them directly to unordered point clouds presents challenges because of their intrinsic causality requirements. Existing methods attempt to address this by sorting points along a single axis. This, however, often overlooks complex causal relationships in 3D space since adjacency relationships based on Euclidean distance between points in the 3D space may not be preserved by this linear arrangement. To overcome this issue, we introduce CompletionMamba, a novel SSM-based network designed to harness SSMs for capturing both global and local dependencies within a point cloud. Initially, the input point cloud is causally structured by rearranging its coordinates. Then, a local SSM framework is proposed that defines neighborhood spaces around each point based on Euclidean distance, enhancing the causal structure. Although local SSM enhances relationships in short and long distance sequences, it still lacks full shape modeling of point cloud. To address this, we propose a novel shape-aware Mamba by integrating the shape code of each 3D shape into the model, enabling shape information propagation to all points. Our experiments show that CompletionMamba achieves state-of-the-art performance on both the MVP and PCN datasets. Zhiheng Fu, Longguang Wang, Lian Xu, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Image Process. | 8 |
| 2025 | A Guide to Image- and Video-Based Small Object Detection Using Deep Learning: Case Study of Maritime SurveillanceabstractDetecting small objects in optical images and videos is a significant challenge in numerous intelligent transportation and autonomous systems. State-of-the-art generic object detection methods fail to accurately localize and identify such small objects (e.g., pedestrians, small vehicles, obstacles). Because small objects occupy only a small area in the input image (e.g.,$32 \times 32$pixels or less), the information extracted from such a small area is not always rich enough to support decision-making. Multidisciplinary strategies are being developed by researchers working at the interface of deep learning and computer vision to enhance the performance of Small Object Detection (SOD). In this paper, we provide a comprehensive review of over 160 research papers published between 2017 and 2022 in order to survey this growing subject. This paper summarizes the existing literature and provides a taxonomy that illustrates the broad picture of current research. We further explore methods to boost the performance of small object detection in maritime settings, where enhanced performance is crucial for ensuring safety and managing traffic. Detecting small objects in the maritime environment requires additional considerations and the current survey aims to review the advanced techniques addressing those aspects. In addition, the popular SOD datasets for generic and maritime applications are discussed, and also well-known evaluation metrics for the state-of-the-art methods on some of the datasets are provided. The link to these datasets appears inhttps://github.com/arekavandi/Datasets_SOD. Aref Miri Rekavandi, Lian Xu, Farid Boussaïd, Abd-Krim Seghouane, Stephen Hoefs, Mohammed Bennamoun |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Box It to Bind It: Unified Layout Control and Attribute Binding in Text-to-Image Diffusion ModelsabstractWhile latent diffusion models (LDMs) excel at creating imaginative images, they often lack precision in semantic fidelity and spatial control over where objects are generated. To address these deficiencies, we introduce the Box-it-to-Bind-it (B2B) module—a novel, training-free approach for improving spatial control and semantic accuracy in text-to-image (T2I) diffusion models. B2B targets three key challenges in T2I: catastrophic neglect, attribute binding, and layout guidance. The process encompasses two main steps: (i)Object generation, which adjusts the latent encoding to guarantee object generation and directs it within specified bounding boxes, and (ii)Attribute binding, ensuring that generated objects adhere to their specified attributes in the prompt. B2B is designed as a compatible plug-and-play module for existing T2I models like Stable Diffusion and Gligen, markedly enhancing models’ performance in addressing these key challenges. We assess our technique on the well-established CompBench and TIFA score benchmarks, and HRS dataset where B2B not only surpasses methods specialized in either attribute binding or layout guidance but also uniquely excels by integrating these capabilities to deliver enhanced overall performance. Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun, Aref Miri Rekavandi, Hamid Laga, Farid Boussaïd |
IEEE Trans. Multim. | 3 |
| 2025 | Auxiliary Tasks Enhanced Dual-Affinity Learning for Weakly Supervised Semantic SegmentationabstractMost existing weakly supervised semantic segmentation (WSSS) methods rely on class activation mapping (CAM) to extract coarse class-specific localization maps using image-level labels. Prior works have commonly used an off-line heuristic thresholding process that combines the CAM maps with off-the-shelf saliency maps produced by a general pretrained saliency model to produce more accurate pseudo-segmentation labels. We propose AuxSegNet+, a weakly supervised auxiliary learning framework to explore the rich information from these saliency maps and the significant intertask correlation between saliency detection and semantic segmentation. In the proposed AuxSegNet+, saliency detection and multilabel image classification are used as auxiliary tasks to improve the primary task of semantic segmentation with only image-level ground-truth labels. We also propose a cross-task affinity learning mechanism to learn pixel-level affinities from the saliency and segmentation feature maps. In particular, we propose a cross-task dual-affinity learning module to learn both pairwise and unary affinities, which are used to enhance the task-specific features and predictions by aggregating both query-dependent and query-independent global context for both saliency detection and semantic segmentation. The learned cross-task pairwise affinity can also be used to refine and propagate CAM maps to provide better pseudo labels for both tasks. Iterative improvement of segmentation performance is enabled by cross-task affinity learning and pseudo-label updating. Extensive experiments demonstrate the effectiveness of the proposed approach with new state-of-the-art WSSS results on the challenging PASCAL VOC and MS COCO benchmarks. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Wanli Ouyang, Ferdous Sohel, Dan Xu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentabstractThis paper proposes Comprehensive Pathology Language Image Pretraining (CPLIP), a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision-language models by leveraging extensive data without needing ground truth annotations. CPLIP involves constructing a pathology-specific dictionary, generating textual descriptions for images using language models, and retrieving relevant images for each text snippet via a pretrained model. The model is then fine-tuned using a many-to-many contrastive learning method to align complex interrelated concepts across both modalities. Evaluated across multiple histopathology tasks, CPLIP shows notable improvements in zero-shot learning scenarios, outperforming existing methods in both interpretability and robustness and setting a higher benchmark for the application of vision-language models in the field. To encourage further research and replication, the code for CPLIP is available on GitHub at https://cplip.github.io/ Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, Mohammed Bennamoun |
CVPR | 6 |
| 2024 | Language Model Guided Interpretable Video Action ReasoningabstractWhile neural networks have excelled in video action recognition tasks, their “black-box” nature often obscures the understanding of their decision-making processes. Re-cent approaches used inherently interpretable models to an-alyze video actions in a manner akin to human reasoning. These models, however, usually fall short in performance compared to their “black-box” counterparts. In this work, we present a new framework named Language-guided Interpretable Action Recognition framework (La-IAR). LaIAR leverages knowledge from language models to enhance both the recognition capabilities and the inter-pretability of video models. In essence, we redefine the problem of understanding video model decisions as a task of aligning video and language models. Using the logical reasoning captured by the language model, we steer the training of the video model. This integrated approach not only improves the video model's adaptability to different domains but also boosts its overall performance. Extensive experiments on two complex video action datasets, Charades & CAD-120, validates the improved performance and inter-pretability of our LaIAR framework. The code of LaIAR is available at https://github.com/NingWang2049/LaIAR. Ning Wang 0047, Guangming Zhu 0001, HS Li, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
CVPR | 6 |
| 2024 | AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun |
ECCV (11) | 8 |
| 2024 | A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-Shaped Structures
Tahmina Khanam, Hamid Laga, Mohammed Bennamoun, Guanjin Wang, Ferdous Sohel, Farid Boussaïd, Anuj Srivastava |
ECCV (67) | 3 |
| 2024 | DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition
Qi Wang 0189, Yuming Lin 0007, Jingtao Ye, Hongsheng Li 0003, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Liang Zhang 0010 |
ECCV (84) | 8 |
| 2024 | CLIFS: Clip-Driven Few-Shot Learning for Baggage Threat ClassificationabstractBaggage screening in airports is a cornerstone in airport security measures. The advent of computer vision technologies in recent years has led to the development of several automated systems for identifying security threats in baggage scans. However, existing methods struggle to adapt to new threat categories when faced with a scarcity of data samples, and the rapid emergence of new threats. Hence, in this paper, we propose a novel CLIP-driven few-shot framework (CLIFS) to explore the potential of multi-modality using text-image fusion through contrastive learning to learn relevant contextual features for recognizing security threats with limited samples. By integrating features from GPT-4 generated captions with image features, CLIFS leverages both visual and textual data to significantly improve threat classification performance with limited samples in a few-shot learning context. Our proposed CLIFS was rigorously tested on the SIXray public available baggage X-ray dataset, where it outperformed state-of-the-art by 31.3% in accuracy and 28.40% in F1-score for the challenging 5-shots scenario, demonstrating its robustness and effectiveness in classifying threats from limited data samples. Abdelfatah Hassan Ahmed, Divya Velayudhan, Mahmoud Elmezain, Muaz Al Radi, Abderrahmene Boudiaf, Taimur Hassan, Mohamed Deriche 0001, Mohammed Bennamoun, Naoufel Werghi |
ICIP | 8 |
| 2024 | Model Predictive Control-Based Reinforcement LearningabstractReinforcement Learning (RL) has garnered much attention in the field of control due to its capacity to learn from interactions and adapt to complex and dynamic environments. However, RL is challenging because it needs to balance exploration, seeking new strategies, and exploitation, leveraging known strategies for maximum gain. To address these challenges, this paper proposes a Model Predictive Control (MPC) based RL approach, where the state value function in RL is utilized as the cost function in MPC, and the system dynamic model is represented by neural networks (NNs). This eliminates the need for human intervention and addresses inaccuracies in the system model. Additionally, MPC-guided RL accelerates convergence during RL training, thereby enhancing sample efficiency. Reported results demonstrate that the proposed method outperforms traditional RL algorithms and does not require prior knowledge of the system. Farid Boussaïd, Mohammed Bennamoun |
ISCAS | 3 |
| 2024 | Referring Human Pose and Mask Estimation In the WildabstractWe introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Bo Miao, Mingtao Feng, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
NeurIPS | 4 |
| 2024 | Enhancing security in X-ray baggage scans: A contour-driven learning approach for abnormality classification and instance segmentation
Abdelfatah Hassan Ahmed, Divya Velayudhan, Taimur Hassan, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Scene Graph Generation: A comprehensive surveyabstractDeep learning techniques have led to remarkable breakthroughs in the field of object detection and have spawned a lot of scene-understanding tasks in recent years. Scene graph has been the focus of research because of its powerful semantic representation and applications to scene understanding. Scene Graph Generation (SGG) refers to the task of automatically mapping an image or a video into a semantic structural scene graph, which requires the correct labeling of detected objects and their relationships. In this paper, a comprehensive survey of recent achievements is provided. This survey attempts to connect and systematize the existing visual relationship detection methods, to summarize, and interpret the mechanisms and the strategies of SGG in a comprehensive way. Deep discussions about current existing problems and future research directions are given at last. This survey will help readers to develop a better understanding of the current researches. Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Youliang Jiang, Yixuan Dang, Haoran Hou, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 10 |
| 2024 | MCTformer+: Multi-Class Token Transformer for Weakly Supervised Semantic SegmentationabstractThis paper proposes a novel transformer-based framework to generate accurate class-specific object localization maps for weakly supervised semantic segmentation (WSSS). Leveraging the insight that the attended regions of the one-class token in the standard vision transformer can generate class-agnostic localization maps, we investigate the transformer's capacity to capture class-specific attention for class-discriminative object localization by learning multiple class tokens. We present the Multi-Class Token transformer, which incorporates multiple class tokens to enable class-aware interactions with patch tokens. This is facilitated by a class-aware training strategy that establishes a one-to-one correspondence between output class tokens and ground-truth class labels. We also introduce a Contrastive-Class-Token (CCT) module to enhance the learning of discriminative class tokens, enabling the model to better capture the unique characteristics of each class. Consequently, the proposed framework effectively generates class-discriminative object localization maps from the class-to-patch attentions associated with different class tokens. To refine these localization maps, we propose the utilization of patch-level pairwise affinity derived from the patch-to-patch transformer attention. Furthermore, the proposed framework seamlessly complements the Class Activation Mapping (CAM) method, yielding significant improvements in WSSS performance on PASCAL VOC 2012 and MS COCO 2014. These results underline the importance of the class token for WSSS. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Hamid Laga, Wanli Ouyang, Dan Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Jacobian norm with Selective Input Gradient Regularization for interpretable adversarial defenseabstractDeep neural networks (DNNs) can be easily deceived by imperceptible alterations known as adversarial examples. These examples can lead to misclassification , posing a significant threat to the reliability of deep learning systems in real-world applications. Adversarial training (AT) is a popular technique used to enhance robustness by training models on a combination of corrupted and clean data. However, existing AT-based methods often struggle to handle transferred adversarial examples that can fool multiple defense models, thereby falling short of meeting the generalization requirements for real-world scenarios. Furthermore, AT typically fails to provide interpretable predictions, which are crucial for domain experts seeking to understand the behavior of DNNs. To overcome these challenges, we present a novel approach called Jacobian norm and Selective Input Gradient Regularization (J-SIGR). Our method leverages Jacobian normalization to improve robustness and introduces regularization of perturbation-based saliency maps, enabling interpretable predictions. By adopting J-SIGR, we achieve enhanced defense capabilities and promote high interpretability of DNNs. We evaluate the effectiveness of J-SIGR across various architectures by subjecting it to powerful adversarial attacks. Our experimental evaluations provide compelling evidence of the efficacy of J-SIGR against transferred adversarial attacks, while preserving interpretability. The project code can be found at https://github.com/Lywu-github/jJ-SIGR.git . Deyin Liu, Lin Wu 0001, Bo Li 0090, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie, Chengwu Liang |
Pattern Recognit. | 5 |
| 2024 | Temporally Consistent Referring Video Object Segmentation With Hybrid MemoryabstractReferring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects. We propose an end-to-end R-VOS paradigm that explicitly models temporal instance consistency alongside the referring segmentation. Specifically, we introduce a novel hybrid memory that facilitates inter-frame collaboration for robust spatio-temporal matching and propagation. Features of frames with automatically generated high-quality reference masks are propagated to segment the remaining frames based on multi-granularity association to achieve temporally consistent R-VOS. Furthermore, we propose a new Mask Consistency Score (MCS) metric to evaluate the temporal consistency of video segmentation. Extensive experiments demonstrate that our approach enhances temporal consistency by a significant margin, leading to top-ranked performance on popular R-VOS benchmarks, i.e., Ref-YouTube-VOS (67.1%) and Ref-DAVIS17 (65.6%). The code is available athttps://github.com/bo-miao/HTR. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Mubarak Shah, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Autonomous Localization of X-Ray Baggage Threats via Weakly Supervised LearningabstractAutonomous X-ray baggage security screening has shown significant strides recently, proving itself a viable solution to the flaws in manual screening, thanks to advancements in deep learning. However, these data-hungry techniques feed on extensively annotated data involving strenuous labor, impeding their advances in baggage screening. Consequently, we present a context-aware transformer for weakly supervised localization to relieve the annotation burden and provide visual interpretability that aids screeners in threat recognition and researchers in identifying the pitfalls of existing systems. The proposed approach can generalize and localize different types of contraband with only cost-effective binary labels without explicit training on item detection. Context extraction block, integrated into the dual-token framework, generates threat-aware context maps, while the token scoring block focuses on minimizing partial activations. Experimental results surpass state of the art (SOTA) methods in terms of classification and localization accuracies. Furthermore, we analyze failures to determine current vulnerabilities and provide new insights for future research. Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Neha Gour, Muhammad Owais, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Region Aware Video Object Segmentation With Deep Motion ModelingabstractCurrent semi-supervised video object segmentation (VOS) methods often employ the entire features of one frame to predict object masks and update memory. This introduces significant redundant computations. To reduce redundancy, we introduce a Region Aware Video Object Segmentation (RAVOS) approach, which predicts regions of interest (ROIs) for efficient object segmentation and memory storage. RAVOS includes a fast object motion tracker to predict object ROIs in the next frame. For efficient segmentation, object features are extracted based on the ROIs, and an object decoder is designed for object-level segmentation. For efficient memory storage, we propose motion path memory to filter out redundant context by memorizing the features within the motion path of objects. In addition to RAVOS, we also propose a large-scale occluded VOS dataset, dubbed OVOS, to benchmark the performance of VOS models under occlusions. Evaluation on DAVIS and YouTube-VOS benchmarks and our new OVOS dataset show that our method achieves state-of-the-art performance with significantly faster inference time, e.g., 86.1 J & F at 42 FPS on DAVIS and 84.4 J & F at 23 FPS on YouTube-VOS. Project page: ravos.netlify.app. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
IEEE Trans. Image Process. | 2 |
| 2024 | Spatio-Temporal Graph Representation Learning for Fraudster Group DetectionabstractMotivated by potential financial gain, companies may hire fraudster groups to write fake reviews to either demote competitors or promote their own businesses. Such groups are considerably more successful in misleading customers, as people are more likely to be influenced by the opinion of a large group. To detect such groups, a common model is to represent fraudster groups' static networks, consequently overlooking the longitudinal behavior of a reviewer, thus, the dynamics of coreview relations among reviewers in a group. Hence, these approaches are incapable of excluding outlier reviewers, which are fraudsters intentionally camouflaging themselves in a group and genuine reviewers happen to coreview in fraudster groups. To address this issue, we propose "FGDT," a framework for "fraudster group detection through temporal relations." FGDT first capitalizes on the effectiveness of the HIN-recurrent neural network (RNN) in both reviewers' representation learning while capturing the collaboration between reviewers. The HIN-RNN models the coreview relations of reviewers in a group in a fixed time window of 28 days. We refer to this as spatial relation learning representation to signify the generalizability of this work to other networked scenarios. Then, we use an RNN on the spatial relations to predict the spatio-temporal relations of reviewers in the group. In the third step, a graph convolution network (GCN) refines the reviewers' vector representations using these predicted relations. These refined representations are then used to remove outlier reviewers. The average of the remaining reviewers' representation is then fed to a simple fully connected layer to predict if the group is a fraudster group or not. Exhaustive experiments of FGDT showed a 5% (4%), 12% (5%), and 12% (5%) improvement over three of the most recent approaches on precision, recall, and F1-value over the Yelp (Amazon) dataset, respectively. Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Higher Order Polynomial Transformer for Fine-Grained Freezing of Gait DetectionabstractFreezing of Gait (FoG) is a common symptom of Parkinson's disease (PD), manifesting as a brief, episodic absence, or marked reduction in walking, despite a patient's intention to move. Clinical assessment of FoG events from manual observations by experts is both time-consuming and highly subjective. Therefore, machine learning-based FoG identification methods would be desirable. In this article, we address this task as a fine-grained human action recognition problem based on vision inputs. A novel deep learning architecture, namely, higher order polynomial transformer (HP-Transformer), is proposed to incorporate pose and appearance feature sequences to formulate fine-grained FoG patterns. In particular, a higher order self-attention mechanism is proposed based on higher order polynomials. To this end, linear, bilinear, and trilinear transformers are formulated in pursuit of discriminative fine-grained representations. These representations are treated as multiple streams and further fused by a cross-order fusion strategy for FoG detection. Comprehensive experiments on a large in-house dataset collected during clinical assessments demonstrate the effectiveness of the proposed method, and an area under the receiver operating characteristic (ROC) curve (AUC) of 0.92 is achieved for detecting FoG. Renfei Sun, Kun Hu 0008, Kaylena A. Ehgoetz Martens, Markus Hagenbuchner, Ah Chung Tsoi, Mohammed Bennamoun, Simon J. G. Lewis, Zhiyong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Learning Multi-Modal Class-Specific Tokens for Weakly Supervised Dense Object LocalizationabstractWeakly supervised dense object localization (WSDOL) relies generally on Class Activation Mapping (CAM), which exploits the correlation between the class weights of the image classifier and the pixel-level features. Due to the limited ability to address intra-class variations, the image classifier cannot properly associate the pixel features, leading to inaccurate dense localization maps. In this paper, we propose to explicitly construct multi-modal class representations by leveraging the Contrastive Language-Image Pre-training (CLIP), to guide dense localization. More specifically, we propose a unified transformer framework to learn two-modalities of class-specific tokens, i.e., class-specific visual and textual tokens. The former captures semantics from the target visual data while the latter exploits the class-related language priors from CLIP, providing complementary information to better perceive the intra-class diversities. In addition, we propose to enrich the multi-modal class-specific tokens with sample-specific contexts comprising visual context and image-language context. This enables more adaptive class representation learning, which further facilitates dense localization. Extensive experiments show the superiority of the proposed method for WSDOL on two multi-label datasets, i.e., PASCAL VOC and MS COCO, and one single-label dataset, i.e., OpenImages. Our dense localization maps also lead to the state-of-the-art weakly supervised semantic segmentation (WSSS) results on PASCAL VOC and MS COCO.11https://github.com/xulianuwa/MMCST Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Dan Xu 0002 |
CVPR | 3 |
| 2023 | Extended Expectation Maximization for Under-Fitted ModelsabstractIn this paper, we generalize the well-known Expectation Maximization (EM) algorithm using the α−divergence for Gaussian Mixture Model (GMM). This approach is used in robust subspace detection when the number of parameters is kept small to avoid overfitting and large estimation variances. The level of robustness can be tuned by the parameter α. When α → 1, our method is equivalent to the standard EM approach and for α < 1 the method is robust against potential outliers. Simulation results show that the method outperforms the standard EM when it comes to mismatches between noise models and their realizations. In addition, we use the proposed method to detect active brain areas using collected functional Magnetic Resonance Imaging (fMRI) data during task-related experiments. Aref Miri Rekavandi, Abd-Krim Seghouane, Farid Boussaïd, Mohammed Bennamoun |
ICASSP | 4 |
| 2023 | VAPCNet: Viewpoint-Aware 3D Point Cloud CompletionabstractMost existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating the viewpoint of each incomplete object is usually time-consuming and leads to huge annotation cost. In this paper, we thus propose an unsupervised viewpoint representation learning scheme for 3D point cloud completion without explicit viewpoint estimation. To be specific, we learn abstract representations of partial scans to distinguish various viewpoints in the representation space rather than the explicit estimation in the 3D space. We also introduce a Viewpoint-Aware Point cloud Completion Network (VAPCNet) with flexible adaption to various viewpoints based on the learned representations. The proposed viewpoint representation learning scheme can extract discriminative representations to obtain accurate viewpoint information. Reported experiments on two popular public datasets show that our VAPCNet achieves state-of-the-art performance for the point cloud completion task. Source code is available at https://github.com/FZH92128/VAPCNet. Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun |
ICCV | 8 |
| 2023 | Spectrum-guided Multi-granularity Referring Video Object SegmentationabstractCurrent referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perceive during the forward computation. This negatively affects the ability of segmentation kernels. To address the drift problem, we propose a Spectrum-guided Multi-granularity (SgMg) approach, which performs direct segmentation on the encoded features and employs visual details to further optimize the masks. In addition, we propose Spectrum-guided Cross-modal Fusion (SCF) to perform intra-frame global interactions in the spectral domain for effective multimodal representation. Finally, we extend SgMg to perform multi-object R-VOS, a new paradigm that enables simultaneous segmentation of multiple referred objects in a video. This not only makes R-VOS faster, but also more practical. Extensive experiments show that SgMg achieves state-of-the-art performance on four video benchmark datasets, outperforming the nearest competitor by 2.8% points on Ref-YouTube-VOS. Our extended SgMg enables multi-object R-VOS, runs about 3 faster while maintaining satisfactory performance. Code×is available at https://github.com/bo-miao/SgMg. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
ICCV | 2 |
| 2023 | Context-Aware Transformers for Weakly Supervised Baggage Threat LocalizationabstractRecent advances in deep learning have facilitated significant progress in the autonomous detection of concealed security threats from baggage X-ray scans, a plausible solution to overcome the pitfalls of manual screening. However, these data-hungry schemes rely on extensive instance-level annotations that involve strenuous skilled labor. Hence, this paper proposes a context-aware transformer for weakly supervised baggage threat localization, exploiting their inherent capacity to learn long-range semantic relations to capture the object-level context of the illegal items. Unlike the conventional single-class token transformers, the proposed dual-token architecture can generalize well to different threat categories by learning the threat-specific semantics from the token-wise attention to generate context maps. The framework has been evaluated on two public datasets, Compass-XP and SIXray, and surpassed other SOTA approaches. Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi |
ICIP | 4 |
| 2023 | Reinforced Learning for Label-Efficient 3D Face Reconstructionabstract3D face reconstruction plays a major role in many human-robot interaction systems, from automatic face authentication to human-computer interface-based entertainment. To improve robustness against occlusions and noise, 3D face reconstruction networks are often trained on a set of in-the-wild face images preferably captured along different viewpoints of the subject. However, collecting the required large amounts of 3D annotated face data is expensive and time-consuming. To address the high annotation cost and due to the importance of training on a useful set, we propose an Active Learning (AL) framework that actively selects the most informative and representative samples to be labeled. To the best of our knowledge, this paper is the first work on tackling active learning for 3D face reconstruction to enable a label-efficient training strategy. In particular, we propose a Reinforcement Active Learning approach in conjunction with a clustering-based pooling strategy to select informative view-points of the subjects. Experimental results on 300W-LP and AFLW2000 datasets demonstrate that our proposed method is able to 1) efficiently select the most influencing view-points for labeling and outperforms several baseline AL techniques and 2) further improve the performance of a 3D Face Reconstruction network trained on the full dataset. Hoda Mohaghegh, Hossein Rahmani 0001, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun |
ICRA | 5 |
| 2023 | UE4-NeRF: Neural Radiance Field for Real-Time Rendering of Large-Scale SceneabstractNeural Radiance Fields (NeRF) is a novel implicit 3D reconstruction method that shows immense potential and has been gaining increasing attention. It enables the reconstruction of 3D scenes solely from a set of photographs. However, its real-time rendering capability, especially for interactive real-time rendering of large-scale scenes, still has significant limitations. To address these challenges, in this paper, we propose a novel neural rendering system called UE4-NeRF, specifically designed for real-time rendering of large-scale scenes. We partitioned each large scene into different sub-NeRFs. In order to represent the partitioned independent scene, we initialize polygonal meshes by constructing multiple regular octahedra within the scene and the vertices of the polygonal faces are continuously optimized during the training process. Drawing inspiration from Level of Detail (LOD) techniques, we trained meshes of varying levels of detail for different observation levels. Our approach combines with the rasterization pipeline in Unreal Engine 4 (UE4), achieving real-time rendering of large-scale scenes at 4K resolution with a frame rate of up to 43 FPS. Rendering within UE4 also facilitates scene editing in subsequent stages. Furthermore, through experiments, we have demonstrated that our method achieves rendering quality comparable to state-of-the-art approaches. Project page: https://jamchaos.github.io/UE4-NeRF/. Jiaming Gu, Minchao Jiang, Hongsheng Li 0003, Xiaoyuan Lu, Guangming Zhu 0001, Syed Afaq Ali Shah, Liang Zhang 0010, Mohammed Bennamoun |
NeurIPS | 8 |
| 2023 | A novel quantum calculus-based complex least mean square algorithm (q-CLMS)
Alishba Sadiq, Imran Naseem, Shujaat Khan, Muhammad Moinuddin, Roberto Togneri, Mohammed Bennamoun |
Appl. Intell. | 6 |
| 2023 | Spiking neural networks for frame-based and event-based single object localization
Sami Barchid, José Mennesson, Jason Kamran Eshraghian, Chaabane Djeraba, Mohammed Bennamoun |
Neurocomputing | 5 |
| 2023 | Robust monocular 3D face reconstruction under challenging viewing conditions
Hoda Mohaghegh, Farid Boussaïd, Hamid Laga, Hossein Rahmani 0001, Mohammed Bennamoun |
Neurocomputing | 5 |
| 2023 | Position and structure-aware graph learning
Guoqiang Ye, Juan Song, Mingtao Feng, Guangming Zhu 0001, Peiyi Shen, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 8 |
| 2023 | Cascaded structure tensor for robust baggage threat detection
Taimur Hassan, Samet Akcay, Bilal Hassan, Mohammed Bennamoun, Salman Khan 0001, Jorge Dias 0001, Naoufel Werghi |
Neural Comput. Appl. | 4 |
| 2023 | Learning class-agnostic masks with cross-task refinement for weakly supervised semantic segmentationabstractAbstract Weakly supervised semantic segmentation (WSSS) commonly relies on Class Activation Mapping (CAM) to produce pseudo semantic labels using image-level annotations. However, because CAM maps often form sparse object regions with poor boundaries, they cannot provide sufficient segmentation supervision. Because off-the-shelf saliency maps can provide rich object boundaries that can be leveraged to improve semantic segmentation, we propose to jointly learn semantic segmentation and class-agnostic masks by using image-level annotations and off-the-shelf saliency maps as supervision. We also propose a cross-task label refinement mechanism, which takes advantage of the learned class-agnostic masks and semantic segmentation masks, to refine the pseudo labels and provide more accurate supervision to both tasks. Moreover, we introduce a new normalization method for CAM to generate more complete class-specific localization maps. The improved CAM maps complement our learned class-agnostic masks, leading to high-quality pseudo semantic segmentation labels. Extensive experiments demonstrate the effectiveness of the proposed approach, with state-of-the-art WSSS results established on PASCAL VOC 2012 and MS COCO. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Wanli Ouyang, Dan Xu 0002 |
Neural Comput. Appl. | 2 |
| 2023 | Multi-Kernel Fusion for RBF Neural NetworksabstractAbstract A simple yet effective architectural design of radial basis function neural networks (RBFNN) makes them amongst the most popular conventional neural networks. The current generation of radial basis function neural network is equipped with multiple kernels which provide significant performance benefits compared to the previous generation using only a single kernel. In existing multi-kernel RBF algorithms, multi-kernel is formed by the convex combination of the base/primary kernels. In this paper, we propose a novel multi-kernel RBFNN in which every base kernel has its own (local) weight. This novel flexibility in the network provides better performance such as faster convergence rate, better local minima and resilience against stucking in poor local minima. These performance gains are achieved at a competitive computational complexity compared to the contemporary multi-kernel RBF algorithms. The proposed algorithm is thoroughly analysed for performance gain using mathematical and graphical illustrations and also evaluated on three different types of problems namely: (i) pattern classification, (ii) system identification and (iii) function approximation. Empirical results clearly show the superiority of the proposed algorithm compared to the existing state-of-the-art multi-kernel approaches. Atif Muhammad Syed, Shujaat Khan, Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
Neural Process. Lett. | 5 |
| 2023 | 4D Atlas: Statistical Analysis of the Spatiotemporal Variability in Longitudinal 3D Shape DataabstractWe propose a novel framework to learn the spatiotemporal variability in longitudinal 3D shape data sets, which contain observations of objects that evolve and deform over time. This problem is challenging since surfaces come with arbitrary parameterizations and thus, they need to be spatially registered. Also, different deforming objects, hereinafter referred to as 4D surfaces, evolve at different speeds and thus they need to be temporally aligned. We solve this spatiotemporal registration problem using a Riemannian approach. We treat a 3D surface as a point in a shape space equipped with an elastic Riemannian metric that measures the amount of bending and stretching that the surfaces undergo. A 4D surface can then be seen as a trajectory in this space. With this formulation, the statistical analysis of 4D surfaces can be cast as the problem of analyzing trajectories embedded in a nonlinear Riemannian manifold. However, performing the spatiotemporal registration, and subsequently computing statistics, on such nonlinear spaces is not straightforward as they rely on complex nonlinear optimizations. Our core contribution is the mapping of the surfaces to the space of Square-Root Normal Fields (SRNF) where the [Formula: see text] metric is equivalent to the partial elastic metric in the space of surfaces. Thus, by solving the spatial registration in the SRNF space, the problem of analyzing 4D surfaces becomes the problem of analyzing trajectories embedded in the SRNF space, which has a euclidean structure. In this paper, we develop the building blocks that enable such analysis. These include: (1) the spatiotemporal registration of arbitrarily parameterized 4D surfaces even in the presence of large elastic deformations and large variations in their execution rates; (2) the computation of geodesics between 4D surfaces; (3) the computation of statistical summaries, such as means and modes of variation, of collections of 4D surfaces; and (4) the synthesis of random 4D surfaces. We demonstrate the performance of the proposed framework using 4D facial surfaces and 4D human body shapes. Hamid Laga, Marcel Padilla, Ian H. Jermyn, Sebastian Kurtek, Mohammed Bennamoun, Anuj Srivastava |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Untrained Neural Network Priors for Inverse Imaging Problems: A SurveyabstractIn recent years, advancements in machine learning (ML) techniques, in particular, deep learning (DL) methods have gained a lot of momentum in solving inverse imaging problems, often surpassing the performance provided by hand-crafted approaches. Traditionally, analytical methods have been used to solve inverse imaging problems such as image restoration, inpainting, and superresolution. Unlike analytical methods for which the problem is explicitly defined and the domain knowledge is carefully engineered into the solution, DL models do not benefit from such prior knowledge and instead make use of large datasets to predict an unknown solution to the inverse problem. Recently, a new paradigm of training deep models using a single image, named untrained neural network prior (UNNP) has been proposed to solve a variety of inverse tasks, e.g., restoration and inpainting. Since then, many researchers have proposed various applications and variants of UNNP. In this paper, we present a comprehensive review of such studies and various UNNP applications for different tasks and highlight various open research problems which require further research. Adnan Qayyum, Inaam Ilahi, Fahad Shamshad, Farid Boussaïd, Mohammed Bennamoun, Junaid Qadir 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Human Action Recognition From Various Data Modalities: A ReviewabstractHuman Action Recognition (HAR) aims to understand human behavior and assign a label to each action. It has a wide range of applications, and therefore has been attracting increasing attention in the field of computer vision. Human actions can be represented using various data modalities, such as RGB, skeleton, depth, infrared, point cloud, event stream, audio, acceleration, radar, and WiFi signal, which encode different sources of useful yet distinct information and have various advantages depending on the application scenarios. Consequently, lots of existing works have attempted to investigate different types of approaches for HAR using various modalities. In this article, we present a comprehensive survey of recent progress in deep learning methods for HAR based on the type of input data modality. Specifically, we review the current mainstream deep learning methods for single data modalities and multiple data modalities, including the fusion-based and the co-learning-based frameworks. We also present comparative results on several benchmark datasets for HAR, together with insightful observations and inspiring future research directions. Zehua Sun, Qiuhong Ke, Hossein Rahmani 0001, Mohammed Bennamoun, Gang Wang 0012, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Training Spiking Neural Networks Using Lessons From Deep LearningabstractThe brain is the perfect place to look for inspiration to develop more efficient neural networks. The inner workings of our synapses and neurons provide a glimpse at what the future of deep learning might look like. This article serves as a tutorial and perspective showing how to apply the lessons learned from several decades of research in deep learning, gradient descent, backpropagation, and neuroscience to biologically plausible spiking neural networks (SNNs). We also explore the delicate interplay between encoding data as spikes and the learning process; the challenges and solutions of applying gradient-based learning to SNNs; the subtle link between temporal backpropagation and spike timing-dependent plasticity; and how deep learning might move toward biologically plausible online learning. Some ideas are well accepted and commonly used among the neuromorphic engineering community, while others are presented or justified for the first time here. A series of companion interactive tutorials complementary to this article using our Python package,snnTorch, are also made available: https://snntorch.readthedocs.io/en/latest/tutorials/index.html. Jason Kamran Eshraghian, Max Ward 0001, Emre Neftci, Xinxin Wang 0002, Gregor Lenz, Girish Dwivedi, Mohammed Bennamoun, Doo Seok Jeong, Wei Lu 0003 |
Proc. IEEE | 7 |
| 2023 | Multi-stage information diffusion for joint depth and surface normal estimation
Zhiheng Fu, Siyu Hong, Hamid Laga, Mohammed Bennamoun, Farid Boussaïd, Yulan Guo |
Pattern Recognit. | 5 |
| 2023 | Cross domain 2D-3D descriptor matching for unconstrained 6-DOF pose estimationabstractThis paper presents a novel approach for cross-domain descriptor matching between 2D and 3D modalities. The 2D-3D matching is applied to localize 2D images in 3D point clouds. Direct cross-domain matching allows our technique to localize images in any type of 3D point cloud without any constraints on the nature or mechanism by which it is obtained. We propose a learning based framework, called Desc-Matcher, to directly match features between the two modalities. A dataset of 2D and 3D features with corresponding locations in images and point clouds is generated to train the Desc-Matcher. To estimate the pose of an image in any 3D cloud, keypoints and feature descriptors are extracted from the query image and the point cloud. The trained Desc-Matcher is then used to match the features from the image and the point cloud. A robust pose estimator is used to predict the location and orientation of the query image from the corresponding positions of the matched 2D and 3D features. We carried out an extensive evaluation of the proposed method for indoor and outdoor scenarios and with different types of point clouds to verify the feasibility of our approach. Experimental results show that the proposed approach can reliably estimate the 6-DOF poses of query cameras in any type of 3D point cloud with high precision. We achieved average median errors of 1.09cm/0.27∘ and 19cm/0.39∘ on the Stanford and Cambridge datasets, respectively. Uzair Nadeem, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel, Aref Miri Rekavandi, Farid Boussaïd |
Pattern Recognit. | 2 |
| 2023 | A Lie algebra representation for efficient 2D shape classification
Xiaohan Yu 0001, Yongsheng Gao 0001, Mohammed Bennamoun, Shengwu Xiong 0001 |
Pattern Recognit. | 3 |
| 2023 | Learning Resolution-Adaptive Representations for Cross-Resolution Person Re-IdentificationabstractCross-resolution person re-identification (CRReID) is a challenging and practical problem that involves matching low-resolution (LR) query identity images against high-resolution (HR) gallery images. Query images often suffer from resolution degradation due to the different capturing conditions from real-world cameras. State-of-the-art solutions for CRReID either learn a resolution-invariant representation or adopt a super-resolution (SR) module to recover the missing information from the LR query. In this paper, we propose an alternative SR-free paradigm to directly compare HR and LR images via a dynamic metric that is adaptive to the resolution of a query image. We realize this idea by learning resolution-adaptive representations for cross-resolution comparison. We propose two resolution-adaptive mechanisms to achieve this. The first mechanism encodes the resolution specifics into different subvectors in the penultimate layer of the deep neural network, creating a varying-length representation. To better extract resolution-dependent information, we further propose to learn resolution-adaptive masks for intermediate residual feature blocks. A novel progressive learning strategy is proposed to train those masks properly. These two mechanisms are combined to boost the performance of CRReID. Experimental results show that the proposed method outperforms existing approaches and achieves state-of-the-art performance on multiple CRReID benchmarks. Lin Wu 0001, Lingqiao Liu, Yang Wang 0023, Zheng Zhang 0006, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie |
IEEE Trans. Image Process. | 6 |
| 2023 | Graph Fusion Network-Based Multimodal Learning for Freezing of Gait DetectionabstractFreezing of gait (FoG) is identified as a sudden and brief episode of movement cessation despite the intention to continue walking. It is one of the most disabling symptoms of Parkinson's disease (PD) and often leads to falls and injuries. Many computer-aided FoG detection methods have been proposed to use data collected from unimodal sources, such as motion sensors, pressure sensors, and video cameras. However, there are limited efforts of multimodal-based methods to maximize the value of all the information collected from different modalities in clinical assessments and improve the FoG detection performance. Therefore, in this study, a novel end-to-end deep architecture, namely graph fusion neural network (GFN), is proposed for multimodal learning-based FoG detection by combining footstep pressure maps and video recordings. GFN constructs multimodal graphs by treating the encoded features of each modality as vertex-level inputs and measures their adjacency patterns to construct complementary FoG representations, thus reducing the representation redundancy among different modalities. In addition, since GFN is devised to process multimodal graphs of arbitrary structures, it is expected to achieve superior performance with inputs containing missing modalities, compared to the alternative unimodal methods. A multimodal FoG dataset was collected, which included clinical assessment videos and footstep pressure sequences of 340 trials from 20 PD patients. Our proposed GFN demonstrates a great promise of multimodal FoG detection with an area under the curve (AUC) of 0.882. To the best of our knowledge, this is one of the first studies to utilize multimodal learning for automated FoG detection, which offers significant opportunities for better patient assessments and clinical trials in the future. Kun Hu 0008, Zhiyong Wang 0001, Kaylena A. Ehgoetz Martens, Markus Hagenbuchner, Mohammed Bennamoun, Ah Chung Tsoi, Simon J. G. Lewis |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | HIN-RNN: A Graph Representation Learning Neural Network for Fraudster Group Detection With No Handcrafted FeaturesabstractSocial reviews are indispensable resources for modern consumers' decision making. For financial gain, companies pay fraudsters preferably in groups to demote or promote products and services since consumers are more likely to be misled by a large number of similar reviews from groups. Recent approaches on fraudster group detection employed handcrafted features of group behaviors without considering the semantic relation between reviews from the reviewers in a group. In this paper, we propose the first neural approach, HIN-RNN, a Heterogeneous Information Network (HIN) Compatible RNN for fraudster group detection that requires no handcrafted features. HIN-RNN provides a unifying architecture for representation learning of each reviewer, with the initial vector as the sum of word embeddings of all review text written by the same reviewer, concatenated by the ratio of negative reviews. Given a co-review network representing reviewers who have reviewed the same items with the same ratings and the reviewers' vector representation, a collaboration matrix is acquired through HIN-RNN training. The proposed approach is confirmed to be effective with marked improvement over state-of-the-art approaches on both the Yelp (22% and 12% in terms of recall and F1-value, respectively) and Amazon (4% and 2% in terms of recall and F1-value, respectively) datasets. Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Generative Metric Learning for Adversarially Robust Open-world Person Re-IdentificationabstractThe vulnerability of re-identification (re-ID) models under adversarial attacks is of significant concern as criminals may use adversarial perturbations to evade surveillance systems. Unlike a closed-world re-ID setting (i.e., a fixed number of training categories), a reliable re-ID system in the open world raises the concern of training a robust yet discriminative classifier, which still shows robustness in the context of unknown examples of an identity. In this work, we improve the robustness of open-world re-ID models by proposing a generative metric learning approach to generate adversarial examples that are regularized to produce robust distance metric. The proposed approach leverages the expressive capability of generative adversarial networks to defend the re-ID models against feature disturbance attacks. By generating the target people variants and sampling the triplet units for metric learning, our learned distance metrics are regulated to produce accurate predictions in the feature metric space. Experimental results on the three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17 demonstrate the robustness of our method. Deyin Liu, Lin Wu 0001, Richang Hong, ZongYuan Ge, Jialie Shen 0001, Farid Boussaïd, Mohammed Bennamoun |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | Multi-class Token Transformer for Weakly Supervised Semantic SegmentationabstractThis paper proposes a new transformer-based framework to learn class-specific object localization maps as pseudo labels for weakly supervised semantic segmentation (WSSS). Inspired by the fact that the attended regions of the one-class token in the standard vision transformer can be leveraged to form a class-agnostic localization map, we investigate if the transformer model can also effectively capture class-specific attention for more discriminative object localization by learning multiple class tokens within the transformer. To this end, we propose a Multi-class Token Transformer, termed as MCTformer, which uses multiple class tokens to learn interactions between the class tokens and the patch tokens. The proposed MCTformer can successfully produce class-discriminative object localization maps from the class-to-patch attentions corresponding to different class tokens. We also propose to use a patch-level pairwise affinity, which is extracted from the patch-to-patch transformer attention, to further refine the localization maps. Moreover, the proposed framework is shown to fully complement the Class Activation Mapping (CAM) method, leading to remarkably superior WSSS results on the PASCAL VOC and MS COCO datasets. These results underline the importance of the class token for WSSS.11https://github.com/xulianuwa/MCTformer Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Dan Xu 0002 |
CVPR | 3 |
| 2022 | Adversary Distillation for One-Shot Attacks on 3D Target TrackingabstractConsidering the vulnerability of existing deep models in the adversarial scenario, the robustness of 3D target tracking is not guaranteed. In this paper, we present an efficient generation based adversarial attack, termed Adversary Distillation Network (AD-Net), which is able to distract a victim tracker in a single shot. In contrast to existing adversarial attacks derived from point perturbations, the proposed method designs a generative network to distill an adversarial example from a tracking template through point-wise filtration. A binary distribution encoding layer is specialized to filter points, which is modeled as a Bernoulli distribution and approximated in a differentiable formulation. To boost the performance of adversarial example generation, a feature extraction module is deployed, which leverages the PointNet++ architecture to learn hierarchical features for the template points as well as similarities with the search areas. Experimental results on the KITTI vision benchmark show that the proposed adversarial attack can effectively mislead popular deep 3D trackers. Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun |
ICASSP | 4 |
| 2022 | Self-Supervised Video Object Segmentation by Motion-Aware Mask PropagationabstractWe propose a self-supervised spatio-temporal matching method, coined Motion-Aware Mask Propagation (MAMP), for video object segmentation. MAMP leverages the frame reconstruction task for training without the need for annotations. During inference, MAMP builds a dynamic memory bank and propagates masks according to our proposed motion-aware spatio-temporal matching module, which is able to handle fast motion and long-term matching scenarios. Evaluation on DAVIS-2017 and YouTube-VOS datasets show that MAMP achieves state-of-the-art performance with stronger generalization ability compared to existing self-supervised methods, i.e., 4.2% higher mean$\mathcal{J}$&$\mathcal{F}$on DAVIS-2017 and 4.85% higher mean$\mathcal{J}$&$\mathcal{F}$on the unseen categories of YouTube-VOS than the nearest competitor. Moreover, MAMP performs at par with many supervised video object segmentation methods. Our code is available at: https://github.com/bo-miao/MAMP. Bo Miao, Mohammed Bennamoun, Yongsheng Gao 0001, Ajmal Mian |
ICME | 2 |
| 2022 | Active-Passive SimStereo - Benchmarking the Cross-Generalization Capabilities of Deep Learning-based Stereo MethodsabstractIn stereo vision, self-similar or bland regions can make it difficult to match patches between two images. Active stereo-based methods mitigate this problem by projecting a pseudo-random pattern on the scene so that each patch of an image pair can be identified without ambiguity. However, the projected pattern significantly alters the appearance of the image. If this pattern acts as a form of adversarial noise, it could negatively impact the performance of deep learning-based methods, which are now the de-facto standard for dense stereo vision. In this paper, we propose the Active-Passive SimStereo dataset and a corresponding benchmark to evaluate the performance gap between passive and active stereo images for stereo matching algorithms. Using the proposed benchmark and an additional ablation study, we show that the feature extraction and matching modules of a selection of twenty selected deep learning-based stereo matching methods generalize to active stereo without a problem. However, the disparity refinement modules of three of the twenty architectures (ACVNet, CascadeStereo, and StereoNet) are negatively affected by the active stereo patterns due to their reliance on the appearance of the input images. Laurent Valentin Jospin, Allen Antony, Lian Xu, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun |
NeurIPS | 6 |
| 2022 | Sign Language Translation with Hierarchical Spatio-Temporal Graph Neural NetworkabstractSign language translation (SLT), which generates text in a spoken language from visual content in a sign language, is important to assist the hard-of-hearing community for their communications. Inspired by neural machine translation (NMT), most existing SLT studies adopted a general sequence to sequence learning strategy. However, SLT is significantly different from general NMT tasks since sign languages convey messages through multiple visual-manual aspects. Therefore, in this paper, these unique characteristics of sign languages are formulated as hierarchical spatio-temporal graph representations, including high-level and fine-level graphs of which a vertex characterizes a specified body part and an edge represents their interactions. Particularly, high-level graphs represent the patterns in the regions such as hands and face, and fine-level graphs consider the joints of hands and landmarks of facial regions. To learn these graph patterns, a novel deep learning architecture, namely hierarchical spatio-temporal graph neural network (HST-GNN), is proposed. Graph convolutions and graph self-attentions with neighborhood context are proposed to characterize both the local and the global graph properties. Experimental results on benchmark datasets demonstrated the effectiveness of the proposed method. Jichao Kan, Kun Hu 0008, Markus Hagenbuchner, Ah Chung Tsoi, Mohammed Bennamoun, Zhiyong Wang 0001 |
WACV | 5 |
| 2022 | Robust real-world point cloud registration by inlier detection
Xiaoshui Huang, Yangfu Wang, Sheng Li 0020, Guofeng Mei, Zongyi Xu, Yucheng Wang 0003, Jian Zhang 0002, Mohammed Bennamoun |
Comput. Vis. Image Underst. | 8 |
| 2022 | Deep-learning-based solution for data deficient satellite image segmentation
Henry Wing Fung Yeung, Vera Chung, Grant Moule, Wayne Thompson, Wanli Ouyang, Tom Weidong Cai, Mohammed Bennamoun |
Expert Syst. Appl. | 8 |
| 2022 | Deep learning for 3D visionabstractWith the rapid development of 3D imaging sensors, such as depth cameras and laser scanning systems, 3D data has become increasingly accessible. Meanwhile, the boost of various deep learning algorithms, such as convolutional neural networks and transformers, further increases the usability of 3D vision systems. Driven by these factors, 3D vision has become an emerging and core component for numerous applications, such as autonomous driving, augmented reality, virtual reality and robotics. Although remarkable progress has been achieved in this area during the last few years, there are still several challenges that need to be addressed, such as the noisy, sparse, and irregular nature of point clouds, the high cost to label 3D data and the necessity to integrate geometry-based and learning-based techniques. Besides, 3D data produced by different 3D imaging sensors (e.g. structured light, stereo, LiDAR and time-of-flight) can be highly different. It is, therefore, necessary to investigate general algorithms that can mitigate the domain gap between different types of 3D data. This special issue aims to collect and present the latest research development in learning-based 3D vision theories and their applications and to inspire future research in this area. In total, there are eight papers accepted for publication in this special issue through careful peer reviews and revisions. These accepted papers are broadly categorised into three topics, and the summary of each topic is given below. TOPIC A—OPTICAL FLOW AND DEPTH ESTIMATION Han et al., in their paper ‘DEMVSNet: Denoising and Depth Inference for Unstructured Multi-View Stereo on Noised Images’, proposed a DEMVSNet to simultaneously address the depth estimation and image denoising problems for unstructured multi-view stereo. The multi-scales feature maps for each image are wrapped to construct cost volumes containing both the depth and RGB information through differentiable homography and Gaussian probability mapping. The cost volume regularisation module is then adopted to predict the probability of depth and RGB. To avoid overfitting in multi-task learning, the gradient normalisation algorithm is utilised to dynamically fine-tune the weights between the depth prediction task and the denoising task. To evaluate the performance of proposed DEMVSNet, a noisy Technical University of Denmark dataset is generated by adding Gaussian-Poisson noise to each image, and the experimental results demonstrate the superiority of DEMVSNet on both the denoising and multi-view stereo reconstruction tasks. Lin et al., in their paper ‘EAGAN: Event-Based Attention Generative Adversarial Networks for Optical Flow and Depth Estimation’, proposed an event-based attention generative adversarial network named EAGAN to simultaneously deal with optical flow and depth estimation based on monocular event camera. The generator of EAGAN is similar to U-net except that a transformer structure is introduced between the encoder and decoder. The position-coding features learnt from the transformer is added to features learnt from the encoding layer, which helps to capture the correlation between sequence information. The discriminator of EAGAN is based on a fully convolutional network and aims to distinguish whether the depth image or the optical flow image is generated by the generator. Experimental results conducted on the multi-vehicle stereo event camera dataset demonstrate the effectiveness of EAGAN on both the depth and optical flow estimation tasks. TOPIC B—POSE ESTIMATION Gao et al., in their paper ‘Efficient 6D Object Pose Estimation based on Attentive Multi-Scale Contextual Information’, proposed an end-to-end 6D pose estimation network to utilise multi-scale contextual features learnt from two heterogeneous data. First, interesting objects are detected from an RGB-D image using an existing semantic segmentation method. Then, pixel-wise geometric and colour features are learnt from 3D point clouds and 2D images respectively. Next, three pixelwise feature attention mechanism modules are utilised to exploit the inter-channel relationship of multimodal features. Finally, multi-scale features are extracted at three different scales and 6D pose is estimated through a dense regression module. Experimental results conducted on the LineMOD and YCB-Video datasets demonstrate that the proposed method achieves state-of-the-art performances in terms of average point distance and average closest point distance. Liu et al., in their paper ‘Auto Calibration of Multi-Camera System for Human Pose Estimation’, proposed an iterative joint estimation of intrinsic and extrinsic parameters for a multi-camera system. Specifically, keypoints are detected with high confidence to estimate the essential matrix between two cameras, and the valid extrinsic parameters are estimated by assuming that the intrinsic parameters are known a priori. Then, the reconstructed 3D human body coordinates are projected into the pixel coordinate system, and the intrinsic parameters are estimated by minimising the projection errors. The experimental results show that the proposed method achieves better performance than commonly used calibration tools. TOPIC C—POINT CLOUD PROCESSING AND UNDERSTANDING Liu et al., in their paper ‘Point Cloud Completion by Dynamic Transformer with Adaptive Neighbourhood Feature Fusion’, utilised the adaptive neighbourhood feature extraction (ANE) module and genetic hierarchical point generation (GHG) module to accomplish the point cloud completion task. The ANE module selects k nearest points both in the spatial and feature spaces adaptively according to different target shapes. The GHG module generates finer point clouds hierarchically according to the local shape characteristics, and the shape information of current points is transferred to the next stage through a dynamic transformer structure. The experimental results conducted on the Point Completion Network and Completion3D datasets demonstrate the superiority of the proposed method. Wang et al., in their paper ‘PCCN-RE: Point Cloud Colourisation Network Based on Relevance Embedding’, proposed a highly authentic point cloud colorisation network based on conditional generative adversarial (cGANs) networks. The generator network predicts the colours from the coordinates of each point, while the discriminator utilises the coordinates and the generated colours to determine the reality of input colourised point clouds. Three key components are contained in the generator. Specifically, the relevance embedding structure captures the most related local information, the weighted pooling structure aggregates the local features based on the correlation values of the covariance matrix, and the enhanced spatial transform network keeps the point clouds invariant to the geometric transformations based on weighted pooling and maximal pooling. The experimental results show that the proposed method achieves the highest Peak Signal to Noise Ratio and Structural Similarity Index on the ShapeNetCore dataset. Fang et al., in their paper ‘Sparse Point-Voxel Aggregation Network for Efficient Point Cloud Semantic Segmentation’, proposed a sparse point-voxel aggregation network to overcome high computational costs in the point cloud semantic segmentation task. In the encoding layer, the local context features are learnt through a sparse convolutional network performed on the voxelised point cloud, and the individual point features are learnt through multi-layer perceptron (MLP)-based network performed on the original point cloud. In the decoding layer, these two kinds of features are aggregated at different encoding layers through simple MLP layers. The experimental results show that the proposed method achieves state-of-the-art performance on the SemanticKITTI and S3DIS datasets. Wang et al., in their paper ‘Scale Robust Point Matching-Net: End-to-End Scale Point Matching Using Lie Group’, proposed an end-to-end scale point cloud matching network named SRPM-Net based on Lie Group. The extracted pointwise features are composed of point absolute coordinates, relative coordinates and point pair features of neighbouring points, and the local context features are aggregated through an attentive pooling layer. The matching matrix is computed via the exponential map of Lie group, which represents the feature similarity of points in two point clouds. The final transformation estimation problem is transferred as estimating the coefficients of the Lie algebra optimisation problem and is optimised through an iterative linear optimisation approach. The experimental results show that SRPM-Net achieves the best performance on the ModelNet40 and Stanford 3D scanning datasets. SUMMARY/CONCLUSION The papers published in this Special Issue show that traditional topics, such as optical flow and depth estimation, pose estimation, and point cloud processing have developed very fast in recent years. In addition, many topics have emerged in deep learning-based 3D vision, such as multi-task joint learning and multimodality intelligence. Future research in this field is expected to boost the theoretical development and potential applications of 3D vision. Yulan Guo, Hanyun Wang, Ronald Clark, Stefano Berretti, Mohammed Bennamoun |
IET Comput. Vis. | 5 |
| 2022 | A comparative study on optical flow for facial expression analysis
Benjamin Allaert, Isaac Ronald Ward, Ioan Marius Bilasco, Chaabane Djeraba, Mohammed Bennamoun |
Neurocomputing | 5 |
| 2022 | Erratum to "Progressive conditional GAN-based augmentation for 3D object recognition" [Neurocomputing 460 (2021) 20-30]
A. A. M. Muzahid, Ferdous Sohel, Mohammed Bennamoun, Hidayat Ullah |
Neurocomputing | 4 |
| 2022 | A reinforcement learning-based approach for imputing missing dataabstractAbstract Missing data is a major problem in real-world datasets, which hinders the performance of data analytics. Conventional data imputation schemes such as univariate single imputation replace missing values in each column with the same approximated value. These univariate single imputation techniques underestimate the variance of the imputed values. On the other hand, multivariate imputation explores the relationships between different columns of data, to impute the missing values. Reinforcement Learning (RL) is a machine learning paradigm where the agent learns by taking actions and receiving rewards in response, to achieve its goal. In this work, we propose an RL-based approach to impute missing data by learning a policy to impute data through an action-reward-based experience. Our approach imputes missing values in a column by working only on the same column (similar to univariate single imputation) but imputes the missing values in the column with different values thus keeping the variance in the imputed values. We report superior performance of our approach, compared with other imputation techniques, on a number of datasets. Saqib Ejaz Awan, Mohammed Bennamoun, Ferdous Sohel, Frank M. Sanfilippo, Girish Dwivedi |
Neural Comput. Appl. | 2 |
| 2022 | Tensor pooling-driven instance segmentation framework for baggage threat recognition
Taimur Hassan, Samet Akcay, Mohammed Bennamoun, Salman Khan 0001, Naoufel Werghi |
Neural Comput. Appl. | 3 |
| 2022 | MEDAS: an open-source platform as a service to help break the walls between medicine and informatics
Liang Zhang 0010, Johann Li, Ping Li 0030, Xiaoyuan Lu, Maoguo Gong, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Kun Qian 0003, Björn W. Schuller |
Neural Comput. Appl. | 9 |
| 2022 | Attack to Fool and Explain Deep NetworksabstractDeep visual models are susceptible to adversarial perturbations to inputs. Although these signals are carefully crafted, they still appear noise-like patterns to humans. This observation has led to the argument that deep visual representation is misaligned with human perception. We counter-argue by providing evidence of human-meaningful patterns in adversarial perturbations. We first propose an attack that fools a network to confuse a whole category of objects (source class) with a target label. Our attack also limits the unintended fooling by samples from non-sources classes, thereby circumscribing human-defined semantic notions for network fooling. We show that the proposed attack not only leads to the emergence of regular geometric patterns in the perturbations, but also reveals insightful information about the decision boundaries of deep models. Exploring this phenomenon further, we alter the 'adversarial' objective of our attack to use it as a tool to 'explain' deep visual representation. We show that by careful channeling and projection of the perturbations computed by our method, we can visualize a model's understanding of human-defined semantic notions. Finally, we exploit the explanability properties of our perturbations to perform image generation, inpainting and interactive image manipulation by attacking adversarialy robust 'classifiers'. In all, our major contribution is a novel pragmatic adversarial attack that is subsequently transformed into a tool to interpret the visual models. The article also makes secondary contributions in terms of establishing the utility of our attack beyond the adversarial objective with multiple interesting applications. Naveed Akhtar, Mohammad A. A. K. Jalwana, Mohammed Bennamoun, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | A Survey on Deep Learning Techniques for Stereo-Based Depth EstimationabstractEstimating depth from RGB images is a long-standing ill-posed problem, which has been explored for decades by the computer vision, graphics, and machine learning communities. Among the existing techniques, stereo matching remains one of the most widely used in the literature due to its strong connection to the human binocular system. Traditionally, stereo-based depth estimation has been addressed through matching hand-crafted features across multiple images. Despite the extensive amount of research, these traditional techniques still suffer in the presence of highly textured areas, large uniform regions, and occlusions. Motivated by their growing success in solving various 2D and 3D vision problems, deep learning for stereo-based depth estimation has attracted a growing interest from the community, with more than 150 papers published in this area between 2014 and 2019. This new generation of methods has demonstrated a significant leap in performance, enabling applications such as autonomous driving and augmented reality. In this paper, we provide a comprehensive survey of this new and continuously growing field of research, summarize the most commonly used pipelines, and discuss their benefits and limitations. In retrospect of what has been achieved so far, we also conjecture what the future may hold for deep learning-based stereo for depth estimation research. Hamid Laga, Laurent Valentin Jospin, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | ScoreGAN: A Fraud Review Detector Based on Regulated GAN With Data AugmentationabstractThe promising performance of Deep Neural Networks (DNNs) in text classification has attracted researchers to use them for fraud review detection. However, the lack of trusted labeled data has limited the performance of the current solutions in detecting fraud reviews. The Generative Adversarial Network (GAN) as a semi-supervised method has been demonstrated to be effective for data augmentation purposes. The state-of-the-art solutions utilize GANs to overcome the data scarcity problem. However, they fail to incorporate the behavioral clues in fraud generation. Additionally, state-of-the-art approaches overlook the possible bot-generated reviews in the dataset. Finally, they also suffer from a common limitation in the generalization and stability of the GAN, slowing down the training procedure. In this work, we propose ScoreGAN for fraud review detection that makes use of both review text and review rating scores in the generation and detection process. Scores are incorporated through Information Gain Maximization (IGM) into the loss function for three reasons. One is to generate score-correlated reviews based on the scores given to the generator. Second, the generated reviews are employed to train the discriminator, allowing the discriminator to correctly label the possible bot-generated reviews through joint representations learned from the concatenation of GLobal Vector for Word representation (GLoVe) extracted from the text and the score. Finally, it can be used to improve the stability and generalization of the GAN. Results show that the proposed framework outperformed the existing state-of-the-art FakeGAN framework, in terms of AP by 7%, and 5% on the Yelp and TripAdvisor datasets, respectively. Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Soft Exemplar Highlighting for Cross-View Image-Based Geo-LocalizationabstractThe goal of ground-to-aerial image geo-localization is to determine the location of a ground query image by matching it against a reference database consisting of aerial/satellite images. This task is highly challenging due to the large appearance difference caused by extreme changes in viewpoint and orientation. In this work, we show that the training difficulty is an important cue that can be leveraged to improve metric learning on cross-view images. More specifically, we propose a new Soft Exemplar Highlighting (SEH) loss to achieve online soft selection of exemplars. Adaptive weights are generated for exemplars by measuring their associated training difficulty using distance rectified logistic regression. These weights are then constrained to remove simple exemplars from training and truncate the large weights of extremely hard exemplars to escape from the trap with a local optimal solution. We further use the proposed SEH loss to train two mainstream convolutional neural networks for ground-to-aerial image-based geo-localization. Experimental results on two benchmark cross-view image datasets demonstrate that the proposed method achieves significant improvements in feature discriminativeness and outperforms the state-of-the-art image-based geo-localization methods. Yulan Guo, Kunhong Li 0001, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Image Process. | 5 |
| 2022 | Dynamic Facial Expression Recognition Under Partial Occlusion With Optical Flow ReconstructionabstractVideo facial expression recognition is useful for many applications and received much interest lately. Although some methods give good results in controlled environments (no occlusion), recognition in the presence of partial facial occlusion remains a challenging task. To handle facial occlusions, methods based on the reconstruction of the occluded part of the face have been proposed. These methods are mainly based on the texture or the geometry of the face. However, the similarity of the face movement between different persons doing the same expression seems to be a real asset for the reconstruction. In this paper we exploit this asset and propose a new method based on an auto-encoder with skip connections to reconstruct the occluded part of the face in the optical flow domain. To the best of our knowledge, this is the first work that directly reconstructs the movement for facial expression recognition. We validated our approach in the controlled CK+ datasets on which different occlusions were generated. Our experiments show that the proposed method reduces the gap in the recognition accuracy between occluded and unoccluded situations. We also compare our approach with existing state-of-the-art approaches. In order to lay the basis of a reproducible and fair comparison in the future, we also propose a new experimental protocol that includes occlusion generation and reconstruction evaluation. Delphine Poux, Benjamin Allaert, Nacim Ihaddadene, Ioan Marius Bilasco, Chaabane Djeraba, Mohammed Bennamoun |
IEEE Trans. Image Process. | 6 |
| 2022 | Pseudo-Pair Based Self-Similarity Learning for Unsupervised Person Re-IdentificationabstractPerson re-identification (re-ID) is of great importance to video surveillance systems by estimating the similarity between a pair of cross-camera person shorts. Current methods for estimating such similarity require a large number of labeled samples for supervised training. In this paper, we present a pseudo-pair based self-similarity learning approach for unsupervised person re-ID without human annotations. Unlike conventional unsupervised re-ID methods that use pseudo labels based on global clustering, we construct patch surrogate classes as initial supervision, and propose to assign pseudo labels to images through the pairwise gradient-guided similarity separation. This can cluster images in pseudo pairs, and the pseudos can be updated during training. Based on pseudo pairs, we propose to improve the generalization of similarity function via a novel self-similarity learning:it learns local discriminative features from individual images via intra-similarity, and discovers the patch correspondence across images via inter-similarity. The intra-similarity learning is based on channel attention to detect diverse local features from an image. The inter-similarity learning employs a deformable convolution with a non-local block to align patches for cross-image similarity. Experimental results on several re-ID benchmark datasets demonstrate the superiority of the proposed method over the state-of-the-arts. Lin Wu 0001, Deyin Liu, Dapeng Chen, ZongYuan Ge, Farid Boussaïd, Mohammed Bennamoun, Jialie Shen 0001 |
IEEE Trans. Image Process. | 7 |
| 2022 | Probability-Based Framework to Fuse Temporal Consistency and Semantic Information for Background SegmentationabstractThe fusion of temporal consistency and semantic information with limited foreground information for background segmentation using deep learning is an underinvestigated problem. In this paper, we explore the relation between temporal consistency and semantic information based on the law of total probability. A highly concise framework is proposed to fuse these two types of information. A theoretical proof is given to show that the proposed framework is more accurate than either the temporal consistency-based model or the semantic information-based model and that each model is a special case of the proposed framework. The proposed framework is a white-box framework that can easily be embedded into a deep neural network as a merging layer. In the proposed model, only a few parameters must be learned, which substantially reduces the need for a large dataset. In addition, these interpretable parameters reflect our understanding of the background and can be applied to a wide range of environments. Extensive evaluations indicate the promising performance of the proposed method. Our code and trained weights for the experiments are available at GitHub.11https://github.com/zengzhi2015/SS_TC_BS(We encourage the reader to run the program for a better understanding of the proposed method). Ting Wang 0026, Fulei Ma, Liang Zhang 0010, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Multim. | 7 |
| 2022 | A Novel Incremental Learning Driven Instance Segmentation Framework to Recognize Highly Cluttered Instances of the Contraband ItemsabstractScreening cluttered and occluded contraband items from baggage X-ray scans is a cumbersome task even for the expert security staff. This article presents a novel strategy that extends a conventional encoder-decoder architecture to perform instance-aware segmentation and extract merged instances of contraband items without using any additional subnetwork or an object detector. The encoder-decoder network first performs conventional semantic segmentation and retrieves cluttered baggage items. The model then incrementally evolves during training to recognize individual instances using significantly reduced training batches. To avoid catastrophic forgetting, a novel objective function minimizes the network loss in each iteration by retaining the previously acquired knowledge while learning new class representations and resolving their complex structural interdependencies through Bayesian inference. A thorough evaluation of our framework on two publicly available X-ray datasets shows that it outperforms state-of-the-art methods, especially within the challenging cluttered scenarios, while achieving an optimal tradeoff between detection accuracy and efficiency. Taimur Hassan, Samet Akcay, Mohammed Bennamoun, Salman Khan 0001, Naoufel Werghi |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Analysis and Variants of Broad Learning SystemabstractThe broad learning system (BLS) is designed based on the technology of compressed sensing and pseudo-inverse theory, and consists of feature nodes and enhancement nodes, has been proposed recently. Compared with the popular deep learning structures, such as deep neural networks, BLS has the ability of rapid incremental learning and can remodel the system without the usual tedious retraining process. However, given that BLS is still in its infancy, it still needs analysis, improvements, and verification. In this article, we first analyze the principle of fast incremental learning ability of BLS in depth. Second, in order to provide an in-depth analysis of the BLS structure, according to the novel structure design concept of deep neural networks, we present four brand-new BLS variant networks and their incremental realizations. Third, based on our analysis of the effect of feature nodes and enhancement nodes, a new BLS structure with a semantic feature extraction layer has been proposed, which is called SFEBLS. The experimental results show that SFEBLS and its variants can increase the accuracy rate on the NORB dataset 6.18%, Fashion-MNIST dataset by 3.15%, ORL data by 5.00%, street view house number dataset by 12.88%, and CIFAR-10 dataset by 18.42%, respectively, and the four brand-new BLS variant networks also obviously outperform the original BLS. Liang Zhang 0010, Guoqing Lu, Peiyi Shen, Mohammed Bennamoun, Syed Afaq Ali Shah, Qiguang Miao, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image SaliencyabstractBackpropagation image saliency aims at explaining model predictions by estimating model-centric importance of individual pixels in the input. However, classinsensitivity of the earlier layers in a network only allows saliency computation with low resolution activation maps of the deeper layers, resulting in compromised image saliency. Remedifying this can lead to sanity failures. We propose CAMERAS, a technique to compute high-fidelity backpropagation saliency maps without requiring any external priors and preserving the map sanity. Our method systematically performs multi-scale accumulation and fusion of the activation maps and backpropagated gradients to compute precise saliency maps. From accurate image saliency to articulation of relative importance of input features for different models, and precise discrimination between model perception of visually similar objects, our high-resolution mapping offers multiple novel insights into the black-box deep visual models, which are presented in the paper. We also demonstrate the utility of our saliency maps in adversarial setup by drastically reducing the norm of attack signals by focusing them on the precise regions identified by our maps. Our method also inspires new evaluation metrics and a sanity check for this developing research direction. Mohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal Mian |
CVPR | 3 |
| 2021 | Leveraging Auxiliary Tasks with Affinity Learning for Weakly Supervised Semantic SegmentationabstractSemantic segmentation is a challenging task in the absence of densely labelled data. Only relying on class activation maps (CAM) with image-level labels provides deficient segmentation supervision. Prior works thus consider pre-trained models to produce coarse saliency maps to guide the generation of pseudo segmentation labels. However, the commonly used off-line heuristic generation process cannot fully exploit the benefits of these coarse saliency maps. Motivated by the significant inter-task correlation, we propose a novel weakly supervised multi-task framework termed as AuxSegNet, to leverage saliency detection and multi-label image classification as auxiliary tasks to improve the primary task of semantic segmentation using only image-level ground-truth labels. Inspired by their similar structured semantics, we also propose to learn a cross-task global pixellevel affinity map from the saliency and segmentation representations. The learned cross-task affinity can be used to refine saliency predictions and propagate CAM maps to provide improved pseudo labels for both tasks. The mutual boost between pseudo label updating and cross-task affinity learning enables iterative improvements on segmentation performance. Extensive experiments demonstrate the effectiveness of the proposed auxiliary learning network structure and the cross-task affinity learning method. The proposed approach achieves state-of-the-art weakly supervised segmentation performance on the challenging PASCAL VOC 2012 and MS COCO benchmarks.1 Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel, Dan Xu 0002 |
ICCV | 3 |
| 2021 | SubICap: Towards Subword-informed Image Captioning
Naeha Sharif, Mohammed Bennamoun, Wei Liu 0006, Syed Afaq Ali Shah |
WACV | 2 |
| 2021 | Imputation of missing data with class imbalance using conditional generative adversarial networks
Saqib Ejaz Awan, Mohammed Bennamoun, Ferdous Sohel, Frank M. Sanfilippo, Girish Dwivedi |
Neurocomputing | 2 |
| 2021 | Progressive conditional GAN-based augmentation for 3D object recognition
A. A. M. Muzahid, Ferdous Sohel, Mohammed Bennamoun, Hidayat Ullah |
Neurocomputing | 4 |
| 2021 | Atrous convolutional feature network for weakly supervised semantic segmentation
Lian Xu, Hao Xue 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel |
Neurocomputing | 3 |
| 2021 | Real time surveillance for low resolution and limited data scenarios: An image set classification approach
Uzair Nadeem, Syed Afaq Ali Shah, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
Inf. Sci. | 3 |
| 2021 | U-net based analysis of MRI for Alzheimer's disease diagnosis
Zhonghao Fan, Johann Li, Liang Zhang 0010, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun, Tao Hua, Wei Wei 0006 |
Neural Comput. Appl. | 9 |
| 2021 | Deep Learning for 3D Point Clouds: A SurveyabstractPoint cloud learning has lately attracted increasing attention due to its wide applications in many areas, such as computer vision, autonomous driving, and robotics. As a dominating technique in AI, deep learning has been successfully used to solve various 2D vision problems. However, deep learning on point clouds is still in its infancy due to the unique challenges faced by the processing of point clouds with deep neural networks. Recently, deep learning on point clouds has become even thriving, with numerous methods being proposed to address different problems in this area. To stimulate future research, this paper presents a comprehensive review of recent progress in deep learning methods for point clouds. It covers three major tasks, including 3D shape classification, 3D object detection and tracking, and 3D point cloud segmentation. It also presents comparative results on several publicly available datasets, together with insightful observations and inspiring future research directions. Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu 0061, Li Liu 0002, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Image-Based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Eraabstract3D reconstruction is a longstanding ill-posed problem, which has been explored for decades by the computer vision, computer graphics, and machine learning communities. Since 2015, image-based 3D reconstruction using convolutional neural networks (CNN) has attracted increasing interest and demonstrated an impressive performance. Given this new era of rapid evolution, this article provides a comprehensive survey of the recent developments in this field. We focus on the works which use deep learning techniques to estimate the 3D shape of generic objects either from a single or multiple RGB images. We organize the literature based on the shape representations, the network architectures, and the training mechanisms they use. While this survey is intended for methods which reconstruct generic objects, we also review some of the recent works which focus on specific object classes such as human body shapes and faces. We provide an analysis and comparison of the performance of some key papers, summarize some of the open problems in this field, and discuss promising directions for future research. Xian-Feng Han, Hamid Laga, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Quantitative performance evaluation of object detectors in hazy environments
Cameron Hodges, Mohammed Bennamoun, Farid Boussaïd |
Pattern Recognit. Lett. | 2 |
| 2021 | Similarity Based Block Sparse Subset Selection for Video SummarizationabstractVideo summarization (VS) is generally formulated as a subset selection problem where a set of representative keyframes or key segments is selected from an entire video frame set. Though many sparse subset selection based VS algorithms have been proposed in the past decade, most of them adopt linear sparse formulation in the explicit feature vector space of video frames, and don’t consider the local or global relationships among frames. In this paper, we first extend the conventional sparse subset selection for VS into kernel block sparse subset selection (KBS3) to utilize the advantage of kernel sparse coding and introduce a local inter-frame relationship through packing of frame blocks. Going a step further, we propose a similarity based block sparse subset selection (SB2S3) model by applying a specially designed transformation matrix on the KBS3 model in order to introduce a kind of global inter-frame relationship through the similarity. Finally, a greedy pursuit based algorithm is devised for the proposed NP-hard model optimization. The proposed SB2S3 has the following advantages: 1) through the similarity between each frame and any other frame, the global relationship among all frames can be considered; 2) through block sparse coding, the local relationship of adjacent frames is further considered; and 3) it has a wider application, since features can derive similarity, but not vice versa. It is believed that the effect of modeling such global and local relationships among frames in this paper, is similar to that of modeling the long-range and short-range dependencies among frames in deep learning based methods. Experimental results on three benchmark datasets have demonstrated that the proposed approach is superior to not only other sparse subset selection based VS methods but also most unsupervised deep-learning based VS methods. Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng, Mohammed Bennamoun |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | DFraud³: Multi-Component Fraud Detection Free of Cold-StartabstractFraud review detection is a hot research topic in recent years. The Cold-start is a particularly new but significant problem referring to the failure of a detection system to recognize the authenticity of a new user. State-of-the-art solutions employ a translational knowledge graph embedding approach (TransE) to model the interaction of the components of a review system. However, these approaches suffer from the limitation of TransE in handling N-1 relations and the narrow scope of a single classification task, i.e., detecting fraudsters only. In this paper, we model a review system as a Heterogeneous Information Network (HIN) which enables a unique representation to every component and performs graph inductive learning on the review data through aggregating features of nearby nodes. HIN with graph induction helps to address the camouflage issue (fraudsters with genuine reviews) which has shown to be more severe when it is coupled with cold-start, i.e., new fraudsters with genuine first reviews. In this research, instead of focusing only on one component, detecting either fraud reviews or fraud users (fraudsters), vector representations are learned for each component, enabling multi-component classification. In other words, we can detect fraud reviews, fraudsters, and fraud-targeted items, thus the name of our approach DFraud3. DFraud3demonstrates a significant accuracy increase of 13% over the state of the art on Yelp. Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Multi-Modal Co-Learning for Liver Lesion Segmentation on PET-CT ImagesabstractLiver lesion segmentation is an essential process to assist doctors in hepatocellular carcinoma diagnosis and treatment planning. Multi-modal positron emission tomography and computed tomography (PET-CT) scans are widely utilized due to their complementary feature information for this purpose. However, current methods ignore the interaction of information across the two modalities during feature extraction, omit the co-learning of the feature maps of different resolutions, and do not ensure that shallow and deep features complement each others sufficiently. In this paper, our proposed model can achieve feature interaction across multi-modal channels by sharing the down-sampling blocks between two encoding branches to eliminate misleading features. Furthermore, we combine feature maps of different resolutions to derive spatially varying fusion maps and enhance the lesions information. In addition, we introduce a similarity loss function for consistency constraint in case that predictions of separated refactoring branches for the same regions vary a lot. We evaluate our model for liver tumor segmentation using a PET-CT scans dataset, compare our method with the baseline techniques for multi-modal (multi-branches, multi-channels and cascaded networks) and then demonstrate that our method has a significantly higher accuracy ( ) than the baseline models. Zhongliang Xue, Ping Li 0030, Liang Zhang 0010, Xiaoyuan Lu, Guangming Zhu 0001, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Attack to Explain Deep RepresentationabstractDeep visual models are susceptible to extremely low magnitude perturbations to input images. Though carefully crafted, the perturbation patterns generally appear noisy, yet they are able to perform controlled manipulation of model predictions. This observation is used to argue that deep representation is misaligned with human perception. This paper counter-argues and proposes the first attack on deep learning that aims at explaining the learned representation instead of fooling it. By extending the input domain of the manipulative signal and employing a model faithful channelling, we iteratively accumulate adversarial perturbations for a deep model. The accumulated signal gradually manifests itself as a collection of visually salient features of the target label (in model fooling), casting adversarial perturbations as primitive features of the target label. Our attack provides the first demonstration of systematically computing perturbations for adversarially non-robust classifiers that comprise salient visual features of objects. We leverage the model explaining character of our algorithm to perform image generation, inpainting and interactive image manipulation by attacking adversarially robust classifiers. The visually appealing results across these applications demonstrate the utility of our attack (and perturbations in general) beyond model fooling. Mohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal Mian |
CVPR | 3 |
| 2020 | Efficient Scene Text Detection with Textual Attention TowerabstractScene text detection has received attention for years and achieved an impressive performance across various benchmarks. In this work, we propose an efficient and accurate approach to detect multi-oriented text in scene images. The proposed feature fusion mechanism allows us to use a shallower network to reduce the computational complexity. A self-attention mechanism is adopted to suppress false positive detections. Experiments on public benchmarks including ICDAR 2013, ICDAR 2015 and MSRA-TD500 show that our proposed approach can achieve better or comparable performances with fewer parameters and less computational cost. Liang Zhang 0010, Lu Yang 0019, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Peiyi Shen |
ICASSP | 7 |
| 2020 | Detecting Prohibited Items in X-Ray Images: a Contour Proposal Learning ApproachabstractX-ray baggage screening plays a vital role in aviation security. Manual inspection of potentially anomalous items is challenging due to the clutter and occlusion within Xray scans. Here, we address this issue by presenting an object-boundaries driven framework for the automated detection of suspicious items from X-ray baggage scans. Rather than recognizing objects directly from the X-ray images, our two-stage detection approach first extracts contour-based proposals using a novel cascaded structure tensor technique and subsequently passes the candidate proposals to a single feed-forward convolutional neural network for recognition. Thorough experimentation on GDXray and SIXray datasets demonstrates that the proposed model achieves a mean area under the curve of 0.9878, outperforming the existing renown state-of-the-art object detection frameworks. Taimur Hassan, Meriem Bettayeb, Samet Akcay, Salman Khan 0001, Mohammed Bennamoun, Naoufel Werghi |
ICIP | 5 |
| 2020 | Efficient Detection of Pixel-Level Adversarial AttacksabstractDeep learning has achieved unprecedented performance in object recognition and scene understanding. However, deep models are also found vulnerable to adversarial attacks. Of particular relevance to robotics systems are pixel-level attacks that can completely fool a neural network by altering very few pixels (e.g. 1-5) in an image. We present the first technique to detect the presence of adversarial pixels in images for the robotic systems, employing an Adversarial Detection Network (ADNet). The proposed network efficiently recognize an input as adversarial or clean by discriminating the peculiar activation signals of the adversarial samples from the clean ones. It acts as a defense mechanism for the robotic vision system by detecting and rejecting the adversarial samples. We thoroughly evaluate our technique on three benchmark datasets including CIFAR-10, CIFAR-100 and Fashion MNIST. Results demonstrate effective detection of adversarial samples by ADNet. Syed Afaq Ali Shah, Moise Bougre, Naveed Akhtar, Mohammed Bennamoun, Liang Zhang 0010 |
ICIP | 4 |
| 2020 | Structure-Feature based Graph Self-adaptive PoolingabstractVarious methods to deal with graph data have been proposed in recent years. However, most of these methods focus on graph feature aggregation rather than graph pooling. Besides, the existing top-k selection graph pooling methods have a few problems. First, to construct the pooled graph topology, current top-k selection methods evaluate the importance of the node from a single perspective only, which is simplistic and unobjective. Second, the feature information of unselected nodes is directly lost during the pooling process, which inevitably leads to a massive loss of graph feature information. To solve these problems mentioned above, we propose a novel graph self-adaptive pooling method with the following objectives: (1) to construct a reasonable pooled graph topology, structure and feature information of the graph are considered simultaneously, which provide additional veracity and objectivity in node selection; and (2) to make the pooled nodes contain sufficiently effective graph information, node feature information is aggregated before discarding the unimportant nodes; thus, the selected nodes contain information from neighbor nodes, which can enhance the use of features of the unselected nodes. Experimental results on four different datasets demonstrate that our method is effective in graph classification and outperforms state-of-the-art graph pooling methods. Liang Zhang 0010, Hongsheng Li 0003, Guangming Zhu 0001, Peiyi Shen, Ping Li 0030, Xiaoyuan Lu, Syed Afaq Ali Shah, Mohammed Bennamoun |
WWW | 9 |
| 2020 | ResFeats: Residual network based features for underwater image classification
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
Image Vis. Comput. | 2 |
| 2020 | Color vision deficiency datasets & recoloring evaluation using GANs
Hongsheng Li 0003, Liang Zhang 0010, Meili Zhang, Guangming Zhu 0001, Peiyi Shen, Ping Li 0030, Mohammed Bennamoun, Syed Afaq Ali Shah |
Multim. Tools Appl. | 8 |
| 2020 | Guest Editors' Introduction to the Special Issue on RGB-D Vision: Methods and ApplicationsabstractThe twenty-six papers in this special issue focus on Red Blue Green (RBG)-D vision, an emerging research topic in computer vision, with a number of applications in robotics, entertainment, biometrics and multimedia. Compared to 2D images and 3D data (including depth images, point clouds and meshes), RGB-D images represent both the photometric and geometric information of a scene. Moreover, low-cost consumer depth cameras (e.g., Microsoft Kinect v2, Intel Realsense, Orbbec Astra) can enable realtime applications due to their high acquisition frame-rate. In the last few years, a large number of RGB-D datasets have also been publicly released to tackle various vision tasks. Although remarkable progress has been achieved, several critical problems still remain open. The aim of this special issue is to stimulate researchers from different fields to present their state-of-the-art work, and to provide a cross-fertilization ground for discussions on the next steps in this important research area. Mohammed Bennamoun, Yulan Guo, Federico Tombari, Kamal Youcef-Toumi, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Topology-learnable graph convolution for skeleton-based action recognition
Guangming Zhu 0001, Liang Zhang 0010, Hongsheng Li 0003, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
Pattern Recognit. Lett. | 6 |
| 2020 | Learning Latent Global Network for Skeleton-Based Action PredictionabstractHuman actions represented with 3D skeleton sequences are robust to clustered backgrounds and illumination changes. In this paper, we investigate skeleton-based action prediction, which aims to recognize an action from a partial skeleton sequence that contains incomplete action information. We propose a new Latent Global Network based on adversarial learning for action prediction. We demonstrate that the proposed network provides latent long-term global information that is complementary to the local action information of the partial sequences and helps improve action prediction. We show that action prediction can be improved by combining the latent global information with the local action information. We test the proposed method on three challenging skeleton datasets and report state-of-the-art performance. Qiuhong Ke, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Image Process. | 2 |
| 2020 | Block Level Skip Connections Across Cascaded V-Net for Multi-Organ SegmentationabstractMulti-organ segmentation is a challenging task due to the label imbalance and structural differences between different organs. In this work, we propose an efficient cascaded V-Net model to improve the performance of multi-organ segmentation by establishing dense Block Level Skip Connections (BLSC) across cascaded V-Net. Our model can take full advantage of features from the first stage network and make the cascaded structure more efficient. We also combine stacked small and large kernels with an inception-like structure to help our model to learn more patterns, which produces superior results for multi-organ segmentation. In addition, some small organs are commonly occluded by large organs and have unclear boundaries with other surrounding tissues, which makes them hard to be segmented. We therefore first locate the small organs through a multi-class network and crop them randomly with the surrounding region, then segment them with a single-class network. We evaluated our model on SegTHOR 2019 challenge unseen testing set and Multi-Atlas Labeling Beyond the Cranial Vault challenge validation set. Our model has achieved an average dice score gain of 1.62 percents and 3.90 percents compared to traditional cascaded networks on these two datasets, respectively. For hard-to-segment small organs, such as the esophagus in SegTHOR 2019 challenge, our technique has achieved a gain of 5.63 percents on dice score, and four organs in Multi-Atlas Labeling Beyond the Cranial Vault challenge have achieved a gain of 5.27 percents on average dice score. Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Medical Imaging | 9 |
| 2020 | Redundancy and Attention in Convolutional LSTM for Gesture RecognitionabstractConvolutional long short-term memory (ConvLSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into ConvLSTM networks. This paper explores the redundancy of spatial convolutions and the effects of the attention mechanism in ConvLSTM, based on our previous gesture recognition architectures that combine the 3-D convolutional neural network (CNN) and ConvLSTM. Depthwise separable, group, and shuffle convolutions are used to replace the convolutional structures in ConvLSTM for the redundancy analysis. In addition, four ConvLSTM variants are derived for attention analysis: 1) by removing the convolutional structures of the three gates in ConvLSTM; 2) by applying the attention mechanism on the ConvLSTM input; and 3) by reconstructing the input and 4) output gates with the modified channelwise attention mechanism. Evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion and that the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn long-term spatiotemporal features when taking spatial or spatiotemporal features as input. A new LSTM variant is derived on this basis in which the convolutional structures are embedded only into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available.\footnotehttps://github.com/GuangmingZhu/ConvLSTMForGR. Guangming Zhu 0001, Liang Zhang 0010, Lu Yang 0019, Lin Mei 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Peiyi Shen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2019 | An Improved Approach to Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with image-level labels is of great significance since it alleviates the dependency on dense annotations. However, it is a challenging task as it aims to achieve a mapping from high-level semantics to low-level features. In this work, we propose a three-step method to bridge this gap. First, we rely on the interpretable ability of deep neural networks to generate attention maps with class localization information by back-propagating gradients. Secondly, we employ an off-the-shelf object saliency detector with an iterative erasing strategy to obtain saliency maps with spatial extent information of objects. Finally, we combine these two complementary maps to generate pseudo ground-truth images for the training of the segmentation network. With the help of the pre-trained model on the MS-COCO dataset and a multi-scale fusion method, we obtained mIoU of 62.1% and 63.3% on PASCAL VOC 2012 val and test sets, respectively, achieving new state-of-the-art results for the weakly supervised semantic segmentation task. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel |
ICASSP | 2 |
| 2019 | Attention-Based Image Captioning Using DenseNet Features
Ferdous Sohel, Mohd Fairuz Shiratuddin, Hamid Laga, Mohammed Bennamoun |
ICONIP (5) | 5 |
| 2019 | Learning-Based Confidence Estimation for Multi-modal Classifier Fusion
Uzair Nadeem, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
ICONIP (2) | 2 |
| 2019 | Direct Image to Point Cloud Descriptors Matching for 6-DOF Camera Localization in Dense 3D Point Clouds
Uzair Nadeem, Mohammad A. A. K. Jalwana, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
ICONIP (2) | 3 |
| 2019 | Coral Classification Using DenseNet and Cross-modality Transfer LearningabstractCoral classification is a challenging task due to the complex morphology and ambiguous boundaries of corals. This paper investigates the benefits of Densely connected convolutional network (DenseNet) and multi-modal image translation techniques in boosting image classification performance by synthesizing missing fluorescence information. To this end, an imageconditional Generative Adversarial Network (GAN) based image translator is trained to model the relationship between reflectance and fluorescence images. Through this image translator, fluorescence images can be generated from the available reflectance images to provide complementary information. During the classification phase, reflectance and translated fluorescence images are combined to obtain more discriminative representations and produce improved classification performance. We present results on the EFC and MLC datasets and report state-of-the-art coral classification performance. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel |
IJCNN | 2 |
| 2019 | Relationship Detection Based on Object Semantic Inference and Attention MechanismsabstractDetecting relations among objects is a crucial task for image understanding. However, each relationship involves different objects pair combinations, and different objects pair combinations express diverse interactions. This makes the relationships, based just on visual features, a challenging task. In this paper, we propose a simple yet effective relationship detection model, which is based on object semantic inference and attention mechanisms. Our model is trained to detect relation triples, such as , . To overcome the high diversity of visual appearances, the semantic inference module and the visual features are combined to complement each others. We also introduce two different attention mechanisms for object feature refinement and phrase feature refinement. In order to derive a more detailed and comprehensive representation for each object, the object feature refinement module refines the representation of each object by querying over all the other objects in the image. The phrase feature refinement module is proposed in order to make the phrase feature more effective, and to automatically focus on relative parts, to improve the visual relationship detection task. We validate our model on Visual Genome Relationship dataset. Our proposed model achieves competitive results compared to the state-of-the-art method MOTIFNET. Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun |
ICMR | 6 |
| 2019 | Deep learning-based 3D local feature descriptor from Mercator projections
Masoumeh Rezaei, Mehdi Rezaeian, Vali Derhami, Ferdous Sohel, Mohammed Bennamoun |
Comput. Aided Geom. Des. | 5 |
| 2019 | LCEval: Learned Composite Metric for Caption Evaluation
Naeha Sharif, Lyndon White, Mohammed Bennamoun, Wei Liu 0006, Syed Afaq Ali Shah |
Int. J. Comput. Vis. | 3 |
| 2019 | NormalNet: A voxel-based CNN for 3D object classification and retrieval
Cheng Wang 0003, Ming Cheng 0002, Ferdous Sohel, Mohammed Bennamoun, Jonathan Li 0001 |
Neurocomputing | 4 |
| 2019 | Single image dehazing using deep neural networks
Cameron Hodges, Mohammed Bennamoun, Hossein Rahmani 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | Continuous Gesture Segmentation and Recognition Using 3DCNN and Convolutional LSTMabstractContinuous gesture recognition aims at recognizing the ongoing gestures from continuous gesture sequences and is more meaningful for the scenarios, where the start and end frames of each gesture instance are generally unknown in practical applications. This paper presents an effective deep architecture for continuous gesture recognition. First, continuous gesture sequences are segmented into isolated gesture instances using the proposed temporal dilated Res3D network. A balanced squared hinge loss function is proposed to deal with the imbalance between boundaries and nonboundaries. Temporal dilation can preserve the temporal information for the dense detection of the boundaries at fine granularity, and the large temporal receptive field makes the segmentation results more reasonable and effective. Then, the recognition network is constructed based on the 3-D convolutional neural network (3DCNN), the convolutional long-short-term-memory network (ConvLSTM), and the 2-D convolutional neural network (2DCNN) for isolated gesture recognition. The “3DCNN-ConvLSTM-2DCNN” architecture is more effective to learn long-term and deep spatiotemporal features. The proposed segmentation and recognition networks obtain the Jaccard index of 0.7163 on the Chalearn LAP ConGD dataset, which is 0.106 higher than the winner of 2017 ChaLearn LAP Large-Scale Continuous Gesture Recognition Challenge. Guangming Zhu 0001, Liang Zhang 0010, Peiyi Shen, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Multim. | 6 |
| 2018 | Global Regularizer and Temporal-Aware Cross-Entropy for Skeleton-Based Early Action Recognition
Qiuhong Ke, Jun Liu 0036, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd |
ACCV (4) | 3 |
| 2018 | Finding Word Sense Embeddings of Known Meaning
Lyndon White, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun |
CICLing (2) | 4 |
| 2018 | NNEval: Neural Network Based Evaluation Metric for Image Captioning
Naeha Sharif, Lyndon White, Mohammed Bennamoun, Syed Afaq Ali Shah |
ECCV (8) | 3 |
| 2018 | Classification of Corals in Reflectance and Fluorescence Images Using Convolutional Neural Network RepresentationsabstractCoral species, with complex morphology and ambiguous boundaries, pose a great challenge for automated classification. CNN activations, which are extracted from fully connected layers of deep networks (FC features), have been successfully used as powerful universal representations in many visual tasks. In this paper, we investigate the transferability and combined performance of FC features and CONY features (extracted from convolutional layers) in the coral classification of two image modalities (reflectance and fluorescence), using a typical deep network (e.g. VGGNet). We exploit vector of locally aggregated descriptors (VLAD) encoding and principal component analysis (PCA) to compress dense CONY features into a compact representation. Experimental results demonstrate that encoded CONV3 features achieve superior performances on reflectance and fluorescence coral images, compared to FC features. The combination of these two features further improves the overall accuracy and achieves state-of-the-art performance on the challenging EFC dataset. Lian Xu, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
ICASSP | 2 |
| 2018 | Reflective Field for Pixel-Level TasksabstractPixelNet has achieved great success in dense prediction problems with a pure pixel-level architecture, but there is still much room for improvement. In this paper, we start from PixelNet and discuss the pixel-level architecture called hypercol-umn and its limitations in building feature representation with rich semantic information. To achieve this goal, we propose a concept in the context of neural networks called reflective field, representing the area reflected by the origin input. Furthermore, the proposed reflective field is used to solve the limitations of the hypercolumn architecture. Specifically, we give the method of calculating the size of the reflective field and analyze the effective reflective field in the calculated area. Then, we use the reflective field to build a new hypercolumn architecture, which has a more rational construction. The results on PASCAL VOC segmentation dataset with our new architecture are improved. Liang Zhang 0010, Xiangwen Kong, Peiyi Shen, Guangming Zhu 0001, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
ICPR | 7 |
| 2018 | Attention in Convolutional LSTM for Gesture RecognitionabstractConvolutional long short-term memory (LSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into the LSTM or the convolutional LSTM (ConvLSTM) networks. Based on the previous gesture recognition architectures which combine the three-dimensional convolution neural network (3DCNN) and ConvLSTM, this paper explores the effects of attention mechanism in ConvLSTM. Several variants of ConvLSTM are evaluated: (a) Removing the convolutional structures of the three gates in ConvLSTM, (b) Applying the attention mechanism on the input of ConvLSTM, (c) Reconstructing the input and (d) output gates respectively with the modified channel-wise attention mechanism. The evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion, and the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn the long-term spatiotemporal features, when taking as input the spatial or spatiotemporal features. On this basis, a new variant of LSTM is derived, in which the convolutional structures are only embedded into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available. Liang Zhang 0010, Guangming Zhu 0001, Lin Mei 0001, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
NeurIPS | 6 |
| 2018 | Efficient finer-grained incremental processing with MapReduce for big data
Liang Zhang 0010, Yuanyuan Feng, Peiyi Shen, Guangming Zhu 0001, Wei Wei 0006, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
Future Gener. Comput. Syst. | 8 |
| 2018 | Improved colour-to-grey method using image segmentation and colour difference model for colour vision deficiencyabstractColour vision deficiency (CVD) is a genetic condition that has troubled people for a long time. This study proposes an improved colour‐to‐grey method for CVD using image segmentation and a colour difference model. In this method, the colour image is first segmented using a region growing method so that each region corresponds to one colour. Next, the colour difference is computed between arbitrary segmented region pairs. Finally, the greyscale image is obtained by minimising a target function. Experimental results show that compared with state‐of‐the‐art colour‐to‐grey methods, the proposed algorithm can improve the E ‐score by about 10.99%. Liang Zhang 0010, Guangming Zhu 0001, Juan Song, Peiyi Shen, Wei Wei 0006, Syed Afaq Ali Shah, Mohammed Bennamoun |
IET Image Process. | 9 |
| 2018 | Semantic scene completion with dense CRF from a single depth image
Liang Zhang 0010, Peiyi Shen, Mohammed Bennamoun, Guangming Zhu 0001, Syed Afaq Ali Shah, Juan Song |
Neurocomputing | 5 |
| 2018 | Exploiting layerwise convexity of rectifier networks with sign constrained weights
Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel |
Neural Networks | 3 |
| 2018 | A Multi-Modal, Discriminative and Spatially Invariant CNN for RGB-D Object LabelingabstractWhile deep convolutional neural networks have shown a remarkable success in image classification, the problems of inter-class similarities, intra-class variances, the effective combination of multi-modal data, and the spatial variability in images of objects remain to be major challenges. To address these problems, this paper proposes a novel framework to learn a discriminative and spatially invariant classification model for object and indoor scene recognition using multi-modal RGB-D imagery. This is achieved through three postulates: 1) spatial invariance $-$ this is achieved by combining a spatial transformer network with a deep convolutional neural network to learn features which are invariant to spatial translations, rotations, and scale changes, 2) high discriminative capability $-$ this is achieved by introducing Fisher encoding within the CNN architecture to learn features which have small inter-class similarities and large intra-class compactness, and 3) multi-modal hierarchical fusion$-$ this is achieved through the regularization of semantic segmentation to a multi-modal CNN architecture, where class probabilities are estimated at different hierarchical levels (i.e., image- and pixel-levels), and fused into a Conditional Random Field (CRF)-based inference hypothesis, the optimization of which produces consistent class labels in RGB-D images. Extensive experimental evaluations on RGB-D object and scene datasets, and live video streams (acquired from Kinect) show that our framework produces superior object and scene classification results compared to the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Response to "Ghost Numbers"abstractThis note clarifies the experimental settings of [1] and shows that the issue raised by [2] is due to a lack of details in [1] which resulted in a misinterpretation of the experimental settings. Munawar Hayat, Mohammed Bennamoun, Senjian An |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | RAFP-Pred: Robust Prediction of Antifreeze Proteins Using Localized Analysis of n-Peptide CompositionsabstractIn extreme cold weather, living organisms produce Antifreeze Proteins (AFPs) to counter the otherwise lethal intracellular formation of ice. Structures and sequences of various AFPs exhibit a high degree of heterogeneity, consequently the prediction of the AFPs is considered to be a challenging task. In this research, we propose to handle this arduous manifold learning task using the notion of localized processing. In particular, an AFP sequence is segmented into two sub-segments each of which is analyzed for amino acid and di-peptide compositions. We propose to use only the most significant features using the concept of information gain (IG) followed by a random forest classification approach. The proposed RAFP-Pred achieved an excellent performance on a number of standard datasets. We report a high Youden's index (sensitivity+specificity-1) value of 0.75 on the standard independent test data set outperforming the AFP-PseAAC, AFP_PSSM, AFP-Pred, and iAFP by a margin of 0.05, 0.06, 0.14, and 0.68, respectively. The verification rate on the UniProKB dataset is found to be 83.19 percent which is substantially superior to the 57.18 percent reported for the iAFP method. Shujaat Khan, Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | Learning Clip Representations for Skeleton-Based 3D Action RecognitionabstractThis paper presents a new representation of skeleton sequences for 3D action recognition. Existing methods based on hand-crafted features or recurrent neural networks cannot adequately capture the complex spatial structures and the long-term temporal dynamics of the skeleton sequences, which are very important to recognize the actions. In this paper, we propose to transform each channel of the 3D coordinates of a skeleton sequence into a clip. Each frame of the generated clip represents the temporal information of the entire skeleton sequence and one particular spatial relationship between the skeleton joints. The entire clip incorporates multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We also propose a multitask convolutional neural network (MTCNN) to learn the generated clips for action recognition. The proposed MTCNN processes all the frames of the generated clips in parallel to explore the spatial and temporal information of the skeleton sequences. The proposed method has been extensively tested on six challenging benchmark datasets. Experimental results consistently demonstrate the superiority of the proposed clip representation and the feature learning method for 3D action recognition compared to the existing techniques. Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Image Process. | 2 |
| 2018 | Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction PredictionabstractPredicting an interaction before it is fully executed is very important in applications, such as human-robot interaction and video surveillance. In a two-human interaction scenario, there are often contextual dependency structures between the global interaction context of the two humans and the local context of the different body parts of each human. In this paper, we propose to learn the structure of the interaction contexts and combine it with the spatial and temporal information of a video sequence to better predict the interaction class. The structural models, including the spatial and the temporal models, are learned with long short term memory (LSTM) networks to capture the dependency of the global and local contexts of each RGB frame and each optical flow image, respectively. LSTM networks are also capable of detecting the key information from global and local interaction contexts. Moreover, to effectively combine the structural models with the spatial and temporal models for interaction prediction, a ranking score fusion method is introduced to automatically compute the optimal weight of each model for score fusion. Experimental results on the BIT-Interaction Dataset and the UT-Interaction Dataset clearly demonstrate the benefits of the proposed method. Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Multim. | 2 |
| 2018 | Cost-Sensitive Learning of Deep Feature Representations From Imbalanced DataabstractClass imbalance is a common problem in the case of real-world object detection and classification tasks. Data of some classes are abundant, making them an overrepresented majority, and data of other classes are scarce, making them an underrepresented minority. This imbalance makes it challenging for a classifier to appropriately learn the discriminating boundaries of the majority and minority classes. In this paper, we propose a cost-sensitive (CoSen) deep neural network, which can automatically learn robust feature representations for both the majority and minority classes. During training, our learning procedure jointly optimizes the class-dependent costs and the neural network parameters. The proposed approach is applicable to both binary and multiclass problems without any modification. Moreover, as opposed to data-level approaches, we do not alter the original data distribution, which results in a lower computational cost during the training process. We report the results of our experiments on six major image classification data sets and show that the proposed approach significantly outperforms the baseline algorithms. Comparisons with popular data sampling techniques and CoSen classifiers demonstrate the superior performance of our proposed method. Salman Khan 0001, Munawar Hayat, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Deep Learning on Underwater Marine Object Detection: A Survey
Md. Moniruzzaman 0001, Syed M. S. Islam, Mohammed Bennamoun, Paul Lavery |
ACIVS | 3 |
| 2017 | A New Representation of Skeleton Sequences for 3D Action RecognitionabstractThis paper presents a new method for 3D action recognition with skeleton sequences (i.e., 3D trajectories of human skeleton joints). The proposed method first transforms each skeleton sequence into three clips each consisting of several frames for spatial temporal feature learning using deep neural networks. Each clip is generated from one channel of the cylindrical coordinates of the skeleton sequence. Each frame of the generated clips represents the temporal information of the entire skeleton sequence, and incorporates one particular spatial relationship between the joints. The entire clips include multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We propose to use deep convolutional neural networks to learn long-term temporal information of the skeleton sequence from the frames of the generated clips, and then use a Multi-Task Learning Network (MTLN) to jointly process all frames of the clips in parallel to incorporate spatial structural information for action recognition. Experimental results clearly show the effectiveness of the proposed new representation and feature learning method for 3D action recognition. Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
CVPR | 2 |
| 2017 | Learning Action Recognition Model from Depth and Skeleton VideosabstractDepth sensors open up possibilities of dealing with the human action recognition problem by providing 3D human skeleton data and depth images of the scene. Analysis of human actions based on 3D skeleton data has become popular recently, due to its robustness and view-invariant representation. However, the skeleton alone is insufficient to distinguish actions which involve human-object interactions. In this paper, we propose a deep model which efficiently models human-object interactions and intra-class variations under viewpoint changes. First, a human body-part model is introduced to transfer the depth appearances of body-parts to a shared view-invariant space. Second, an end-to-end learning framework is proposed which is able to effectively combine the view-invariant body-part representation from skeletal and depth images, and learn the relations between the human body-parts and the environmental objects, the interactions between different human body-parts, and the temporal structure of human actions. We have evaluated the performance of our proposed model against 15 existing techniques on two large benchmark human action recognition datasets including NTU RGB+D and UWA3DII. The Experimental results show that our technique provides a significant improvement over state-of-the-art methods. Hossein Rahmani 0001, Mohammed Bennamoun |
ICCV | 2 |
| 2017 | Resfeats: Residual network based features for image classificationabstractDeep residual networks have recently emerged as the state-of-the-art architecture in image classification and object detection. In this paper, we propose new image features (called ResFeats) extracted from the last convolutional layer of the deep residual networks pre-trained on ImageNet. We propose to use ResFeats for diverse image classification tasks namely, object classification, scene classification and coral classification and show that ResFeats consistently perform better than their CNN counterparts on these classification tasks. Since the ResFeats are large feature vectors, we explore dimensionality reduction methods. Experimental results are provided to show the effectiveness of ResFeats with state-of-the-art classification accuracies on Caltech-101, Caltech-256 and MLC datasets and a significant performance improvement on MIT-67 dataset compared to the widely used CNN features. Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel |
ICIP | 2 |
| 2017 | Learning deep structured network for weakly supervised change detectionabstractConventional change detection methods require a large number of images to learn background models or depend on tedious pixel-level labeling by humans. In this paper, we present a weakly supervised approach that needs only image-level labels to simultaneously detect and localize changes in a pair of images. To this end, we employ a deep neural network with DAG topology to learn patterns of change from image-level labeled training data. On top of the initial CNN activations, we define a CRF model to incorporate the local differences and context with the dense connections between individual pixels. We apply a constrained mean-field algorithm to estimate the pixel-level labels, and use the estimated labels to update the parameters of the CNN in an iterative EM framework. This enables imposing global constraints on the observed foreground probability mass function. Our evaluations on four benchmark datasets demonstrate superior detection and localization performance. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IJCAI | 4 |
| 2017 | Empowering Simple Binary Classifiers for Image Set Based Face Recognition
Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun |
Int. J. Comput. Vis. | 3 |
| 2017 | Discriminative feature learning and region consistency activation for robust scene labeling
Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
Neurocomputing | 3 |
| 2017 | Keypoints-based surface representation for 3D modeling and 3D object recognition
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Pattern Recognit. | 2 |
| 2017 | Scale space clustering evolution for salient region detection on 3D deformable shapes
Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun, Yulan Guo |
Pattern Recognit. | 3 |
| 2017 | Deep feature learning for dummies: A simple auto-encoder training method using Particle Swarm Optimisation
Chao Sui, Mohammed Bennamoun, Roberto Togneri |
Pattern Recognit. Lett. | 2 |
| 2017 | A cascade gray-stereo visual feature extraction method for visual and audio-visual speech recognition
Chao Sui, Roberto Togneri, Mohammed Bennamoun |
Speech Commun. | 3 |
| 2017 | SkeletonNet: Mining Deep Part Features for 3-D Action RecognitionabstractThis letter presents SkeletonNet, a deep learning framework for skeleton-based 3-D action recognition. Given a skeleton sequence, the spatial structure of the skeleton joints in each frame and the temporal information between multiple frames are two important factors for action recognition. We first extract body-part-based features from each frame of the skeleton sequence. Compared to the original coordinates of the skeleton joints, the proposed features are translation, rotation, and scale invariant. To learn robust temporal information, instead of treating the features of all frames as a time series, we transform the features into images and feed them to the proposed deep learning network, which contains two parts: one to extract general features from the input images, while the other to generate a discriminative and compact representation for action recognition. The proposed method is tested on the SBU kinect interaction dataset, the CMU dataset, and the large-scale NTU RGB+D dataset and achieves state-of-the-art performance. Qiuhong Ke, Senjian An, Mohammed Bennamoun, Ferdous Sohel, Farid Boussaïd |
IEEE Signal Process. Lett. | 3 |
| 2017 | Forest Change Detection in Incomplete Satellite Images With Deep Neural NetworksabstractLand cover change monitoring is an important task from the perspective of regional resource monitoring, disaster management, land development, and environmental planning. In this paper, we analyze imagery data from remote sensing satellites to detect forest cover changes over a period of 29 years (1987-2015). Since the original data are severely incomplete and contaminated with artifacts, we first devise a spatiotemporal inpainting mechanism to recover the missing surface reflectance information. The spatial filling process makes use of the available data of the nearby temporal instances followed by a sparse encoding-based reconstruction. We formulate the change detection task as a region classification problem. We build a multiresolution profile (MRP) of the target area and generate a candidate set of bounding-box proposals that enclose potential change regions. In contrast to existing methods that use handcrafted features, we automatically learn region representations using a deep neural network in a data-driven fashion. Based on these highly discriminative representations, we determine forest changes and predict their onset and offset timings by labeling the candidate set of proposals. Our approach achieves the state-of-the-art average patch classification rate of 91.6% (an improvement of ~16%) and the mean onset/offset prediction error of 4.9 months (an error reduction of five months) compared with a strong baseline. We also qualitatively analyze the detected changes in the unlabeled image regions, which demonstrate that the proposed forest change detection approach is scalable to new regions. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | A Joint Deep Boltzmann Machine (jDBM) Model for Person Identification Using Mobile Phone DataabstractWe propose an audio-visual person identification approach based on a joint deep Boltzmann machine (jDBM) model. The proposed jDBM model is trained in three steps: 1) learning the unimodal DBM models corresponding to the speech and facial image modalities, 2) learning the shared layer parameters using a joint restricted Boltzmann machine (jRBM) model, and 3) the fine-tuning of the jDBM model after the initialization with the parameters of the unimodal DBMs and the shared layer. The activation probabilities of the units of the shared layer are used as the joint features and a logistic regression classifier is used for the combined speech and facial image recognition. We show that by learning the shared layer parameters using a jRBM, a higher accuracy can be achieved compared to the greedy layer-wise initialization. The performance of our proposed model is also compared with a state-of-the art support vector machine (SVM), deep belief network (DBN), and the deep auto-encoder (DAE) models. In addition, our experimental results show that the joint representations obtained from the proposed jDBM model are robust to noise and missing information. Experiments were carried out on the challenging MOBIO database, which includes audio-visual data captured using mobile phones. Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
IEEE Trans. Multim. | 2 |
| 2017 | RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded ForestsabstractThis paper presents an efficient framework to perform recognition and grasp detection of objects from RGB-D images of real scenes. The framework uses a novel architecture of hierarchical cascaded forests, in which object-class and grasp-pose probabilities are computed at different levels of an image hierarchy (e.g., patch and object levels) and fused to infer the class and the grasp of unseen objects. We introduce a novel training objective function that minimizes the uncertainties of the class labels and the grasp ground truths at the leaves of the forests, thereby enabling the framework to perform the recognition and grasp detection of objects. Our objective function is learned from features that are extracted from RGB-D point clouds of the objects. For that, we propose a novel method to encode an RGB-D point cloud into a representation that facilitates the use of large convolution neural networks to extract discriminative features from RGB-D images. We evaluate our framework on challenging object datasets, where we demonstrate that our framework outperforms the state-of-the-art methods in terms of object-recognition and grasp-detection accuracies. We also show experiments by using live video streams from a Kinect mounted on our in-house robotic platform. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IEEE Trans. Robotics | 2 |
| 2016 | Generating Bags of Words from the Sums of Their Word Embeddings
Lyndon White, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun |
CICLing (1) | 4 |
| 2016 | Coral classification with hybrid feature representationsabstractCoral reefs exhibit significant within-class variations, complex between-class boundaries and inconsistent image clarity. This makes coral classification a challenging task. In this paper, we report the application of generic CNN representations combined with hand-crafted features for coral reef classification to take advantage of the complementary strengths of these representation types. We extract CNN based features from patches centred at labelled pixels at multiple scales. We use texture and color based hand-crafted features extracted from the same patches to complement the CNN features. Our proposed method achieves a classification accuracy that is higher than the state-of-art methods on the MLC benchmark dataset for corals. Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd, Renae Hovey, Gary A. Kendrick, Robert B. Fisher |
ICIP | 2 |
| 2016 | Audio-visual biometric recognition via joint sparse representationsabstractIn this paper we present a novel audio-visual (AV) person identification system based on joint sparse representation. Video features used were vectorized raw pixel values, while i-vectors were used as the audio features. Classification is performed by solving the joint sparsity optimization problem, and fusion is carried out by using the quality (confidence) assigned to each matcher. Our experimental results on the challenging MOBIO database using 100 subjects show that the system based on joint sparse representation outperforms the system based on separate sparse representations for each modality. Furthermore, we show that our newly introduced quality measure improves the system's performance, when compared to conventionally used quality measures for sparse representation - based systems. Rudi Primorac, Roberto Togneri, Mohammed Bennamoun, Ferdous Sohel |
ICPR | 3 |
| 2016 | Simultaneous dense scene reconstruction and object labelingabstractThis paper presents an efficient system for simultaneous dense scene reconstruction and object labeling in real-world environments (captured with an RGB-D sensor). The proposed system starts with the generation of object proposals in the scene. It then tracks spatio-temporally consistent object proposals across multiple frames and produces a dense reconstruction of the scene. In parallel, the proposed system uses an efficient inference algorithm, where object class probabilities are computed at an object-level and fused into a voxel-based prediction hypothesis modeled on the voxels of the reconstructed scene. Our extensive experiments using challenging RGB-D object and scene datasets, and live video streams from Microsoft Kinect show that the proposed system achieved competitive 3D scene reconstruction and object labeling results compared to the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
ICRA | 2 |
| 2016 | Heat propagation contours for 3D non-rigid shape analysisabstractWe present a novel local shape descriptor by means of General Adaptive Neighborhoods (GANs) based on the properties of the heat diffusion process on a Riemannian manifold. The GAN is a spatial region, surrounding the feature point and fitting its local shape structure, which is isometric. Our signature, called the Heat Propagation Contours (HPCs), is obtained by analysing the well-known heat kernel and extracting contours automatically within the GAN as heat dissipates from the feature point onto the rest of the shape. HPCs capture geometric information around the feature point by investigating the heat propagation process both in the temporal and spatial domain. HPCs share many useful characteristics with the heat based methods. Particularly, it captures the intrinsic geometry of a shape and is suitable for non-rigid shape analysis. In addition, our signature provides an elegant and efficient way to describe the neighborhood of the feature point in a multi-scale approach. The proposed descriptor is evaluated on several datasets to demonstrate its effectiveness. Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun |
WACV | 3 |
| 2016 | Partial fingerprint indexing: a combination of local and reconstructed global featuresabstractSummary Existing work on partial fingerprint indexing attempts to make full use of the extracted features from the partial segments, such as singular points, minutiae, orientation field, and ridge count. However, singular points may not exist in partial fingerprints, and none of these features can form a complete set of feature vectors that can be used for matching with those derived from the corresponding full fingerprints for indexing. Our former work on fingerprint orientation model based on two‐dimensional Fourier expansion (FOMFE) coefficients‐based fingerprint indexing and global orientation field reconstruction has demonstrated the possibility of reconstructing a global feature vector for partial fingerprint indexing. In this paper, we design some novel features of minutiae triplets in addition to some commonly used features to constitute the local minutiae triplet features. Experiments carried out on fingerprint verification competition (FVC) 2000 DB2a, FVC 2002 DB1a, and National Institute of Standards and Technology (NIST) SD 14 demonstrate the performance improvement after adding the new features to minutiae triplet feature set. We then propose to combine the reconstructed global feature and local minutiae triplet features to improve the performance of partial fingerprint indexing. Specifically, the minutiae triplet‐based indexing scheme and the FOMFE coefficients‐based indexing scheme are applied separately to generate two candidate lists; then, a fuzzy‐based fusion scheme is designed to generate the final candidate list for matching. Experiments carried out on the public database NIST SD 14 show that the proposed approach can improve the performance that has been achieved by individual partial fingerprint indexing algorithms before fusion. Copyright © 2015 John Wiley & Sons, Ltd. Jiankun Hu, Song Wang 0003, Ian R. Petersen, Mohammed Bennamoun |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | A Comprehensive Performance Evaluation of 3D Local Feature Descriptors
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan, Ngai Ming Kwok |
Int. J. Comput. Vis. | 2 |
| 2016 | Integrating Geometrical Context for Semantic Labeling of Indoor Scenes using RGBD Images
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri, Imran Naseem |
Int. J. Comput. Vis. | 2 |
| 2016 | A semantic RBM-based model for image set classification
Said Elaiwat, Mohammed Bennamoun, Farid Boussaïd |
Neurocomputing | 2 |
| 2016 | An RGB-D based image set classification for robust face recognition from Kinect data
Munawar Hayat, Mohammed Bennamoun, Amar A. El-Sallam |
Neurocomputing | 2 |
| 2016 | Iterative deep learning for image set based face and object recognition
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Neurocomputing | 2 |
| 2016 | A novel feature representation for automatic 3D object recognition in cluttered scenes
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Neurocomputing | 2 |
| 2016 | Automatic Shadow Detection and Removal from a Single ImageabstractWe present a framework to automatically detect and remove shadows in real world scenes from a single image. Previous works on shadow detection put a lot of effort in designing shadow variant and invariant hand-crafted features. In contrast, our framework automatically learns the most relevant features in a supervised manner using multiple convolutional deep neural networks (ConvNets). The features are learned at the super-pixel level and along the dominant boundaries in the image. The predicted posteriors based on the learned features are fed to a conditional random field model to generate smooth shadow masks. Using the detected shadow masks, we propose a Bayesian formulation to accurately extract shadow matte and subsequently remove shadows. The Bayesian formulation is based on a novel model which accurately models the shadow generation process in the umbra and penumbra regions. The model parameters are efficiently estimated using an iterative optimization procedure. Our proposed framework consistently performed better than the state-of-the-art on all major shadow databases collected under a variety of conditions. Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | A spatio-temporal RBM-based model for facial expression recognition
Said Elaiwat, Mohammed Bennamoun, Farid Boussaïd |
Pattern Recognit. | 2 |
| 2016 | A Two-Phase Weighted Collaborative Representation for 3D partial face recognition with single sample
Yinjie Lei, Yulan Guo, Munawar Hayat, Mohammed Bennamoun, Xinzhi Zhou |
Pattern Recognit. | 4 |
| 2016 | EI3D: Expression-invariant 3D face recognition based on feature and shape matching
Yulan Guo, Yinjie Lei, Li Liu 0002, Yan Wang 0059, Mohammed Bennamoun, Ferdous Sohel |
Pattern Recognit. Lett. | 5 |
| 2016 | A Spatial Layout and Scale Invariant Feature Representation for Indoor Scene ClassificationabstractUnlike standard object classification, where the image to be classified contains one or multiple instances of the same object, indoor scene classification is quite different since the image consists of multiple distinct objects. Furthermore, these objects can be of varying sizes and are present across numerous spatial locations in different layouts. For automatic indoor scene categorization, large-scale spatial layout deformations and scale variations are therefore two major challenges and the design of rich feature descriptors which are robust to these challenges is still an open problem. This paper introduces a new learnable feature descriptor called “spatial layout and scale invariant convolutional activations” to deal with these challenges. For this purpose, a new convolutional neural network architecture is designed which incorporates a novel “spatially unstructured” layer to introduce robustness against spatial layout deformations. To achieve scale invariance, we present a pyramidal image representation. For feasible training of the proposed network for images of indoor scenes, this paper proposes a methodology, which efficiently adapts a trained network model (on a large-scale data) for our task with only a limited amount of available training data. The efficacy of the proposed approach is demonstrated through extensive experiments on a number of data sets, including MIT-67, Scene-15, Sports-8, Graz-02, and NYU data sets. Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun, Senjian An |
IEEE Trans. Image Process. | 3 |
| 2016 | A Discriminative Representation of Convolutional Features for Indoor Scene RecognitionabstractIndoor scene recognition is a multi-faceted and challenging problem due to the diverse intra-class variations and the confusing inter-class similarities. This paper presents a novel approach which exploits rich mid-level convolutional features to categorize indoor scenes. Traditionally used convolutional features preserve the global spatial structure, which is a desirable property for general object recognition. However, we argue that this structuredness is not much helpful when we have large variations in scene layouts, e.g., in indoor scenes. We propose to transform the structured convolutional activations to another highly discriminative feature space. The representation in the transformed space not only incorporates the discriminative aspects of the target dataset, but it also encodes the features in terms of the general object categories that are present in indoor scenes. To this end, we introduce a new large-scale dataset of 1300 object categories which are commonly present in indoor scenes. Our proposed approach achieves a significant performance boost over previous state of the art approaches on five major scene classification datasets. Salman Khan 0001, Munawar Hayat, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
IEEE Trans. Image Process. | 3 |
| 2015 | Separating objects and clutter in indoor scenesabstractObjects' spatial layout estimation and clutter identification are two important tasks to understand indoor scenes. We propose to solve both of these problems in a joint framework using RGBD images of indoor scenes. In contrast to recent approaches which focus on either one of these two problems, we perform ‘fine grained structure categorization’ by predicting all the major objects and simultaneously labeling the cluttered regions. A conditional random field model is proposed to incorporate a rich set of local appearance, geometric features and interactions between the scene elements. We take a structural learning approach with a loss of 3D localisation to estimate the model parameters from a large annotated RGBD dataset, and a mixed integer linear programming formulation for inference. We demonstrate that our approach is able to detect cuboids and estimate cluttered regions across many different object and scene categories in the presence of occlusion, illumination and appearance variations. Salman Khan 0001, Xuming He 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
CVPR | 3 |
| 2015 | Extracting deep bottleneck features for visual speech recognitionabstractMotivated by the recent progresses in the use of deep learning techniques for acoustic speech recognition, we present in this paper a visual deep bottleneck feature (DBNF) learning scheme using a stacked auto-encoder combined with other techniques. Experimental results show that our proposed deep feature learning scheme yields approximately 24% relative improvement for visual speech accuracy. To the best of our knowledge, this is the first study which uses deep bottleneck feature on visual speech recognition. Our work firstly shows that the deep bottleneck visual feature is able to achieve a significant accuracy improvement on visual speech recognition. Chao Sui, Roberto Togneri, Mohammed Bennamoun |
ICASSP | 3 |
| 2015 | Contractive Rectifier Networks for Nonlinear Maximum Margin ClassificationabstractTo find the optimal nonlinear separating boundary with maximum margin in the input data space, this paper proposes Contractive Rectifier Networks (CRNs), wherein the hidden-layer transformations are restricted to be contraction mappings. The contractive constraints ensure that the achieved separating margin in the input space is larger than or equal to the separating margin in the output layer. The training of the proposed CRNs is formulated as a linear support vector machine (SVM) in the output layer, combined with two or more contractive hidden layers. Effective algorithms have been proposed to address the optimization challenges arising from contraction constraints. Experimental results on MNIST, CIFAR-10, CIFAR-100 and MIT-67 datasets demonstrate that the proposed contractive rectifier networks consistently outperform their conventional unconstrained rectifier network counterparts. Senjian An, Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel |
ICCV | 4 |
| 2015 | Listening with Your Eyes: Towards a Practical Visual Speech Recognition System Using Deep Boltzmann MachinesabstractThis paper presents a novel feature learning method for visual speech recognition using Deep Boltzmann Machines (DBM). Unlike all existing visual feature extraction techniques which solely extracts features from video sequences, our method is able to explore both acoustic information and visual information to learn a better visual feature representation in the training stage. During the test stage, instead of using both audio and visual signals, only the videos are used for generating the missing audio feature, and both the given visual and given audio features are used to obtain a joint representation. We carried out our experiments on a large scale audio-visual data corpus, and experimental results show that our proposed techniques outperforms the performance of the hadncrafted features and features learned by other commonly used deep learning techniques. Chao Sui, Mohammed Bennamoun, Roberto Togneri |
ICCV | 2 |
| 2015 | Outdoor scene labelling with learned features and region consistency activationabstractThis paper presents a learned feature based method for scene labelling. This method is combined with a novel strategy to improve global label consistency. We first follow a traditional way to investigate trained features from convolutional neural networks (ConvNets) for scene labelling. Then, motivated by the recent successful use of general features extracted from ConvNets for various applications, we extend the use of the general features to scene labelling (for the first time). We further propose an algorithm called Region Consistency Activation (RCA) to improve the global label consistency. RCA is based on a novel transformation between Ultrametric Contour Map (UCM) and the Probability of Regions Consistency (PRC). Our algorithms were rigorously tested on the popular Stanford Background and SIFT Flow datasets. We achieved superior performances compared with the state-of-the-art methods on both of these datasets. Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
ICIP | 3 |
| 2015 | How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances?abstractThis paper investigates how hidden layers of deep rectifier networks are capable of transforming two or more pattern sets to be linearly separable while preserving the distances with a guaranteed degree, and proves the universal classification power of such distance preserving rectifier networks. Through the nearly isometric nonlinear transformation in the hidden layers, the margin of the linear separating plane in the output layer and the margin of the nonlinear separating boundary in the original data space can be closely related so that the maximum margin classification in the input data space can be achieved approximately via the maximum margin linear classifiers in the output layer. The generalization performance of such distance preserving deep rectifier neural networks can be well justified by the distance-preserving properties of their hidden layers and the maximum margin property of the linear classifiers in the output layer. Senjian An, Farid Boussaïd, Mohammed Bennamoun |
ICML | 3 |
| 2015 | Efficient RGB-D object categorization using cascaded ensembles of randomized decision treesabstractThis paper presents an efficient framework for the categorization of objects in real-world scenes (captured with an RGB-D sensor). The proposed framework uses ensembles of randomized decision trees in a hierarchical cascaded architecture to compute consistent object-class inferences of unseen objects. Specifically, the proposed framework computes object-class probabilities at three levels of an image hierarchy (i.e., pixel-, surfel-, and object-levels) using Random Forest classifiers. Next, these probabilities are fused together to compute a cumulative probabilistic output which is used to infer object categories. This fusion results in an improved object categorization performance compared with the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
ICRA | 2 |
| 2015 | Discriminative feature learning for efficient RGB-D object recognitionabstractThis paper presents an efficient approach to recognize objects captured with an RGB-D sensor. The proposed approach uses a Bag-of-Words (BOW) model to learn feature representations from raw RGB-D point clouds in a weakly supervised manner. To this end, we introduce a novel method based on randomized clustering trees to learn visual vocabularies which are fast to compute and more discriminative compared to the vocabularies generated by classical methods such as k-means. We show that, when combined with standard spatial pooling strategies, our proposed approach yields a powerful feature representation for RGB-D object recognition. Our extensive experimental evaluation on two challenging RGB-D object datasets and live video streams from Kinect shows that our learned features result in superior object recognition accuracies compared with the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IROS | 2 |
| 2015 | Sign Constrained Rectifier Networks with Applications to Pattern Decompositions
Senjian An, Qiuhong Ke, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel |
ECML/PKDD (1) | 3 |
| 2015 | Deep Boltzmann Machines for i-Vector Based Audio-Visual Person Identification
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
PSIVT | 2 |
| 2015 | Binary Descriptor Based on Heat Diffusion for Non-rigid Shape Analysis
Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun |
PSIVT | 3 |
| 2015 | Heterogeneous Multi-column ConvNets with a Fusion Framework for Object RecognitionabstractThe purpose of this paper is to investigate heterogeneous multi-column ConvNets (MCCNN) and fusion methods for them. We first construct heterogeneous MCCNN by combining ConvNets with different structures. We then use different fusion methods to check their performances to find out the effect of fusion methods for MCCNN. We also propose a novel sliding window based fusion framework which defines a specific subset of columns to be picked up from MCCNN for fusion. Two different strategies (exhaustive sliding window and sliding window from training) are investigated to determine the best performance of the fusion process. We tested the heterogeneous MCCNN and sliding window fusion on the MNIST dataset for optical character recognition. Experiments show that MCCNN improved the accuracy of recognition compared with a single column of ConvNets. Moreover, sliding window fusion is a more generalized fusion method and consistently achieves better results compared with the traditional fusion methods. We also tested the MCCNN and sliding window fusion on CIFAR-10 and Caltech-256 datasets. We achieved superior results compared to existing state-of-the-art techniques. Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
WACV | 3 |
| 2015 | A novel local surface feature for 3D object recognition under clutter and occlusion
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, Min Lu 0001 |
Inf. Sci. | 3 |
| 2015 | Deep Reconstruction Models for Image Set ClassificationabstractImage set classification finds its applications in a number of real-life scenarios such as classification from surveillance videos, multi-view camera networks and personal albums. Compared with single image based classification, it offers more promises and has therefore attracted significant research attention in recent years. Unlike many existing methods which assume images of a set to lie on a certain geometric surface, this paper introduces a deep learning framework which makes no such prior assumptions and can automatically discover the underlying geometric structure. Specifically, a Template Deep Reconstruction Model (TDRM) is defined whose parameters are initialized by performing unsupervised pre-training in a layer-wise fashion using Gaussian Restricted Boltzmann Machines (GRBMs). The initialized TDRM is then separately trained for images of each class and class-specific DRMs are learnt. Based on the minimum reconstruction errors from the learnt class-specific models, three different voting strategies are devised for classification. Extensive experiments are performed to demonstrate the efficacy of the proposed framework for the tasks of face and object recognition from image sets. Experimental results show that the proposed method consistently outperforms the existing state of the art methods. Munawar Hayat, Mohammed Bennamoun, Senjian An |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | A Curvelet-based approach for textured 3D face recognition
Said Elaiwat, Mohammed Bennamoun, Farid Boussaïd, Amar A. El-Sallam |
Pattern Recognit. | 2 |
| 2015 | A novel 3D vorticity based approach for automatic registration of low resolution range images
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Pattern Recognit. | 2 |
| 2015 | A confidence-based late fusion framework for audio-visual biometric identification
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
Pattern Recognit. Lett. | 2 |
| 2015 | Quantitative Error Analysis of Bilateral FilteringabstractOne of the fastest acceleration techniques for bilateral image filtering is the real time O(1) quantization method proposed by Yang 2009, which first computes some Principal Bilateral Filtered Image Components (PBFICs) and then applies linear interpolation to estimate the filtered output images. There is a trade-off between accuracy and efficiency in selecting the number of PBFICs: the more PBFICs are used, the higher the accuracy, and the higher the computational cost. A question arises: how many PBFICs are required to achieve a certain level of accuracy? In this letter, we address this question by investigating the properties of bilateral filtering and deriving the linear interpolation error bounds when only a subset of PBFICs is used. The provided theoretical analysis indicates that the necessary number of PBFICs for user-provided precision depends on the range kernel and, for typical Gaussian range kernels, a small percentage (typically less than 4%) of the PBFICs are enough for good approximations. Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel |
IEEE Signal Process. Lett. | 3 |
| 2015 | A Low-Cost Implementation of a 360° Vision Distributed Aperture SystemabstractVisible light cameras commonly used for surveillance applications usually have a limited field of view (FoV). To acquire a broader FoV, distributed aperture systems (DASs) combine views from multiple cameras. Although some commercial or proprietary systems already exist, the open literature in this field reports solely the performance and/or hardware architecture of the respective systems, and omits the required details for a reimplementation. In this paper we present a low-cost, personal-computer-based 360° DAS, with a full description of the hardware architecture and the software implementation details. In particular, we describe in detail two problems, the offline estimation of the orientations of the cameras of the proposed DAS and the online synthesis of a virtual view in real time. For the first problem, we propose an area-based bundle adjuster by combining the forward additive Lucas-Kanade algorithm with a bundle adjustment strategy. The color virtual view of up to 2048 × 1024 pixels is synthesized online at 39 frames/s. A large majority of the workload of the online phase is implemented on the graphic processing unit. The techniques and methods proposed in this paper are generic and independent on the arrangement of cameras. Xiaoming Peng, Mohammed Bennamoun, Qingbo Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Performance Evaluation of 3D Local Feature Descriptors
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan, Jun Zhang 0044 |
ACCV (2) | 2 |
| 2014 | Learning Non-linear Reconstruction Models for Image Set ClassificationabstractWe propose a deep learning framework for image set classification with application to face recognition. An Adaptive Deep Network Template (ADNT) is defined whose parameters are initialized by performing unsupervised pre-training in a layer-wise fashion using Gaussian Restricted Boltzmann Machines (GRBMs). The pre-initialized ADNT is then separately trained for images of each class and class-specific models are learnt. Based on the minimum reconstruction error from the learnt class-specific models, a majority voting strategy is used for classification. The proposed framework is extensively evaluated for the task of image set classification based face recognition on Honda/UCSD, CMU Mobo, YouTube Celebrities and a Kinect dataset. Our experimental results and comparisons with existing state-of-the-art methods show that the proposed method consistently achieves the best performance on all these datasets. Munawar Hayat, Mohammed Bennamoun, Senjian An |
CVPR | 2 |
| 2014 | Automatic Feature Learning for Robust Shadow DetectionabstractWe present a practical framework to automatically detect shadows in real world scenes from a single photograph. Previous works on shadow detection put a lot of effort in designing shadow variant and invariant hand-crafted features. In contrast, our framework automatically learns the most relevant features in a supervised manner using multiple convolutional deep neural networks (ConvNets). The 7-layer network architecture of each ConvNet consists of alternating convolution and sub-sampling layers. The proposed framework learns features at the super-pixel level and along the object boundaries. In both cases, features are extracted using a context aware window centered at interest points. The predicted posteriors based on the learned features are fed to a conditional random field model to generate smooth shadow contours. Our proposed framework consistently performed better than the state-of-the-art on all major shadow databases collected under a variety of conditions. Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
CVPR | 2 |
| 2014 | Model-Free Segmentation and Grasp Selection of Unknown Stacked Objects
Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
ECCV (5) | 2 |
| 2014 | Reverse Training: An Efficient Approach for Image Set Classification
Munawar Hayat, Mohammed Bennamoun, Senjian An |
ECCV (6) | 2 |
| 2014 | Geometry Driven Semantic Labeling of Indoor Scenes
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
ECCV (1) | 2 |
| 2014 | Confidence-based Rank-level Fusion for Audio-visual Person Identification SystemabstractA multibiometric identification system establishes the identity of a person based on the input biometric data presented to its sub-systems. Each sub-system compares the features extracted from the input against the templates of all identities stored in gallery. The best matched identity is ranked highest in the ranked list. In rank-level fusion, the ranked lists from different sub-systems are combined to reach a final decision. However, the state-of-the-art rank-level fusion methods consider that all sub-systems are equally reliable in terms of classifying the probe data. In practice, the probe data may be affected by different sources of degradation (e.g., illumination and pose variation on the face image, environmental noise) and thus affecting the overall recognition accuracy. In this paper, robust rank-level fusion methods (e.g., confidence based highest rank and Borda count) are proposed by using confidence measures for each sub-system in the decision making process. Experimental results show that the proposed confidence based rank-level fusion achieved higher recognition rates than state-of-the-art rank-level fusion methods. Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
ICPRAM | 2 |
| 2014 | A model-free approach for the segmentation of unknown objectsabstractWe address the problem of object segmentation from depth images of highly complex indoor scenes. We propose a model-free segmentation approach, which robustly separates unknown stacked objects in real-world scenes. Our approach constructs geometrically constrained 3D clusters known as salient-regions, which are subsequently merged into high-level object hypotheses by analyzing the local geometrical characteristics (such as local shape and homogeneity) of the area of their shared boundaries. We tested our approach using depth images from live Kinect video streams and publicly available RGB-D datasets. Our approach is highly efficient and achieves superior performance compared to state-of-the-art techniques. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IROS | 2 |
| 2014 | Fingerprint Indexing Based on Combination of Novel Minutiae Triplet Features
Jiankun Hu, Song Wang 0003, Ian R. Petersen, Mohammed Bennamoun |
NSS | 5 |
| 2014 | 3D Object Recognition in Cluttered Scenes with Local Surface Features: A Surveyabstract3D object recognition in cluttered scenes is a rapidly growing research area. Based on the used types of features, 3D object recognition methods can broadly be divided into two categories-global or local feature based methods. Intensive research has been done on local surface feature based methods as they are more robust to occlusion and clutter which are frequently present in a real-world scene. This paper presents a comprehensive survey of existing local surface feature based 3D object recognition methods. These methods generally comprise three phases: 3D keypoint detection, local surface feature description, and surface matching. This paper covers an extensive literature survey of each phase of the process. It also enlists a number of popular and contemporary databases together with their relevant attributes. Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | An efficient 3D face recognition approach using local geometrical signatures
Yinjie Lei, Mohammed Bennamoun, Munawar Hayat, Yulan Guo |
Pattern Recognit. | 2 |
| 2014 | An Automatic Framework for Textured 3D Video-Based Facial Expression RecognitionabstractMost of the existing research on 3D facial expression recognition has been done using static 3D meshes. 3D videos of a face are believed to contain more information in terms of the facial dynamics which are very critical for expression recognition. This paper presents a fully automatic framework which exploits the dynamics of textured 3D videos for recognition of six discrete facial expressions. Local video-patches of variable lengths are extracted from numerous locations of the training videos and represented as points on the Grassmannian manifold. An efficient graph-based spectral clustering algorithm is used to separately cluster these points for every expression class. Using a valid Grassmannian kernel function, the resulting cluster centers are embedded into a Reproducing Kernel Hilbert Space (RKHS) where six binary SVM models are learnt. Given a query video, we extract video-patches from it, represent them as points on the manifold and match these points with the learnt SVM models followed by a voting based strategy to decide about the class of the query video. The proposed framework is also implemented in parallel on 2D videos and a score level fusion of 2D & 3D videos is performed for performance improvement of the system. The experimental results on BU4DFE data set show that the system achieves a very high classification accuracy for facial expression recognition from 3D videos. Munawar Hayat, Mohammed Bennamoun |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | An Accurate and Robust Range Image Registration Algorithm for 3D Object ModelingabstractRange image registration is a fundamental research topic for 3D object modeling and recognition. In this paper, we propose an accurate and robust algorithm for pairwise and multi-view range image registration. We first extract a set of Rotational Projection Statistics (RoPS) features from a pair of range images, and perform feature matching between them. The two range images are then registered using a transformation estimation method and a variant of the Iterative Closest Point (ICP) algorithm. Based on the pairwise registration algorithm, we propose a shape growing based multi-view registration algorithm. The seed shape is initialized with a selected range image and then sequentially updated by performing pairwise registration between itself and the input range images. All input range images are iteratively registered during the shape growing process. Extensive experiments were conducted to test the performance of our algorithm. The proposed pairwise registration algorithm is accurate, and robust to small overlaps, noise and varying mesh resolutions. The proposed multi-view registration algorithm is also very accurate. Rigorous comparisons with the state-of-the-art show the superiority of our algorithm. Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, Min Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | 3D-Div: A novel local surface descriptor for feature matching and pairwise range image registrationabstractThis paper presents a novel local surface descriptor, called 3D-Div. The proposed descriptor is based on the concept of 3D vector fields divergence, extensively used in electromagnetic theory. To generate a 3D-Div descriptor of a 3D surface, a keypoint is first extracted on the 3D surface, then a local patch of a certain size is selected around that keypoint. A Local Reference Frame (LRF) is then constructed at the keypoint using all points forming the patch. A normalized 3D vector field is then computed at each point in the patch and referenced with LRF vectors. The 3D-Div descriptors are finally generated as the divergence of the reoriented 3D vector field. We tested our proposed descriptor on the low resolution Washington RGB-D (Kinect) object dataset. Performance was evaluated for the tasks of feature matching and pairwise range image registration. Experimental results showed that the proposed 3D-Div is 88% more computationally efficient and 47% more accurate than commonly used Spin Image (SI) descriptors. Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd, Amar A. El-Sallam |
ICIP | 2 |
| 2013 | Partial Fingerprint Reconstruction with Improved Smooth Extension
Jiankun Hu, Ian R. Petersen, Mohammed Bennamoun |
NSS | 4 |
| 2013 | A low cost 3D markerless system for the reconstruction of athletic techniquesabstractWe present a low cost markerless system for the optimization of athlete performance in sports such as pole vault, jumping and javelin throw. The system uses a number of calibrated cameras to capture a video of an athlete from different viewpoints. The athlete's body is then segmented from the background in each video frame. The silhouettes of the segmented body are then reprojected to reconstruct an estimate of the 3D body shape of the athlete, known as the visual hull (VH). The VH is tracked over a number of frames in real testing trials. A template combining a high resolution 3D scan and a 2D mass scan is then aligned with the VH in each frame. A set of motion analysis parameters such as the take-off data are finally estimated from the aligned template and compared with the ones obtained using a gold standard marker-based system, namely the Vicon. The proposed system was tested in real-time trials and was able to provide comparable results to the Vicon system. Amar A. El-Sallam, Mohammed Bennamoun, Ferdous Sohel, Jacqueline A. Alderson, Andrew Lyttle, M. Rossi |
WACV | 2 |
| 2013 | 3D free form object recognition using rotational projection statisticsabstractRecognizing 3D objects in the presence of clutter and occlusion is a challenging task. This paper presents a 3D free form object recognition system based on a novel local surface feature descriptor. For a randomly selected feature point, a local reference frame (LRF) is defined by calculating the eigenvectors of the covariance matrix of a local surface, and a feature descriptor called rotational projection statistics (RoPS) is constructed by calculating the statistics of the point distribution on 2D planes defined from the LRF. It finally proposes a 3D object recognition algorithm based on RoPS features. Candidate models and transformation hypotheses are generated by matching the scene features against the model features in the library, these hypotheses are then tested and verified by aligning the model to the scene. Comparative experiments were performed on two publicly available datasets and an overall recognition rate of 98.8% was achieved. Experimental results show that our method is robust to noise, mesh resolution variations and occlusion. Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Jianwei Wan, Min Lu 0001 |
WACV | 2 |
| 2013 | Clustering of video-patches on Grassmannian manifold for facial expression recognition from 3D videosabstractThis paper presents a fully automatic system which exploits the dynamics of 3D videos and is capable of recognizing six basic facial expressions. Local video-patches of variable lengths are extracted from different locations of the training videos and represented as points on the Grass-mannian manifold. An efficient spectral clustering based algorithm is used to separately cluster points for each of the six expression classes. The resulting cluster centers are matched with the points of a test video and a voting based strategy is used to decide about the expression class of the test video. The proposed system is tested on the largest publicly available 3D video database, BU4DFE. The experimental results show that the system achieves a very high classification accuracy and outperforms the current state of the art algorithms for facial expression recognition from 3D videos. Munawar Hayat, Mohammed Bennamoun, Amar A. El-Sallam |
WACV | 2 |
| 2013 | A lip extraction algorithm using region-based ACM with automatic contour initializationabstractIn a lipreading system, lip extraction is a fundamental method that directly affects the final speech recognition results. However, most existing systems need to detect some facial features as prior-knowledge to construct the initial contour, and any erroneous feature detection will lead to an incorrect lip extraction. In order to solve this problem, this paper presents a new framework which integrates both global region-based Active Contour Model (ACM) and localized region-based ACM. With the utilization of the proposed framework, the initial contour does not need to be specified according to the speaker facial features before extracting the lip, so that any erroneous extraction introduced by an incorrect initial contour is effectively eliminated. Experimental results show the efficiency of the proposed method in comparison with the existing methods. Chao Sui, Mohammed Bennamoun, Roberto Togneri, Serajul Haque |
WACV | 2 |
| 2013 | Rotational Projection Statistics for 3D Local Surface Description and Object Recognition
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Min Lu 0001, Jianwei Wan |
Int. J. Comput. Vis. | 3 |
| 2013 | Multibiometric human recognition using 3D ear and face features
Syed M. S. Islam, Rowan Davies, Mohammed Bennamoun, Robyn A. Owens, Ajmal Mian |
Pattern Recognit. | 3 |
| 2013 | An efficient 3D face recognition approach based on the fusion of novel local low-level features
Yinjie Lei, Mohammed Bennamoun, Amar A. El-Sallam |
Pattern Recognit. | 2 |
| 2013 | Discriminative fusion of shape and appearance features for human pose estimation
Suman Sedai, Mohammed Bennamoun, Du Q. Huynh |
Pattern Recognit. | 2 |
| 2013 | A Gaussian Process Guided Particle Filter for Tracking 3D Human Pose in VideoabstractIn this paper, we propose a hybrid method that combines Gaussian process learning, a particle filter, and annealing to track the 3D pose of a human subject in video sequences. Our approach, which we refer to as annealed Gaussian process guided particle filter, comprises two steps. In the training step, we use a supervised learning method to train a Gaussian process regressor that takes the silhouette descriptor as an input and produces multiple output poses modeled by a mixture of Gaussian distributions. In the tracking step, the output pose distributions from the Gaussian process regression are combined with the annealed particle filter to track the 3D pose in each frame of the video sequence. Our experiments show that the proposed method does not require initialization and does not lose tracking of the pose. We compare our approach with a standard annealed particle filter using the HumanEva-I dataset and with other state of the art approaches using the HumanEva-II dataset. The evaluation results show that our approach can successfully track the 3D human pose over long video sequences and give more accurate pose tracking results than the annealed particle filter. Suman Sedai, Mohammed Bennamoun, Du Q. Huynh |
IEEE Trans. Image Process. | 2 |
| 2012 | Evaluation of Spatiotemporal Detectors and Descriptors for Facial Expression RecognitionabstractLocal spatiotemporal detectors and descriptors have recently become very popular for video analysis in many applications. They do not require any preprocessing steps and are invariant to spatial and temporal scales. Despite their computational simplicity, they have not been evaluated and tested for video analysis of facial data. This paper considers two space-time detectors and four descriptors and uses bag of features framework for human facial expression recognition on BU_4DFE data set. A comparison of local spatiotemporal features with other non-spatiotemporal published techniques on the same data set is also given. Unlike spatiotemporal features, these techniques involve time consuming and computationally intensive preprocessing steps like manual initialization and tracking of facial points. Our results show that despite being totally automatic and not requiring any user intervention, local spacetime features provide promising and comparable performance for facial expression recognition on BU_4DFE data set. Munawar Hayat, Mohammed Bennamoun, Amar A. El-Sallam |
HSI | 2 |
| 2012 | Novel low level local features for 3D expression invariant face recognitionabstractIn this paper, we present a system based on novel low level local features to recognize 3D faces under varying facial expressions. Our local features are obtained by combinatorially selecting two points from expression insensitive semi-rigid portions of the face. The curve length between the two points is computed and the distribution of such curve lengths is used as a feature vector to model the geometric shape distribution of the face. Our proposed features are very simple to compute yet highly distinctive and discriminating. Kernel Fisher discriminant analysis is used for feature optimization, followed by a linear support vector machine classifier for recognition. The system is extensively tested on 2500 facial scans of BU 3DFE dataset. Our experimental results show that the proposed system achieves a very high average classification rate of 99.17% and verification rates of 99.0% and above for a false acceptance rate of 0.001. Munawar Hayat, Mohammed Bennamoun, Yinjie Lei, Amar A. El-Sallam |
ICARCV | 2 |
| 2012 | Fully automatic face recognition from 3D videos
Munawar Hayat, Mohammed Bennamoun, Amar A. El-Sallam |
ICPR | 2 |
| 2012 | Robust regression for face recognition
Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
Pattern Recognit. | 3 |
| 2012 | Spatially Optimized Data-Level Fusion of Texture and Shape for Face RecognitionabstractData-level fusion is believed to have the potential for enhancing human face recognition. However, due to a number of challenges, current techniques have failed to achieve its full potential. We propose spatially optimized data/pixel-level fusion of 3-D shape and texture for face recognition. Fusion functions are objectively optimized to model expression and illumination variations in linear subspaces for invariant face recognition. Parameters of adjacent functions are constrained to smoothly vary for effective numerical regularization. In addition to spatial optimization, multiple nonlinear fusion models are combined to enhance their learning capabilities. Experiments on the FRGC v2 data set show that spatial optimization, higher order fusion functions, and the combination of multiple such functions systematically improve performance, which is, for the first time, higher than score-level fusion in a similar experimental setup. Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
IEEE Trans. Image Process. | 2 |
| 2012 | Sliding-Window Designs for Vertex-Based Shape CodingabstractTraditionally the sliding window (SW) has been employed in vertex-based operational rate distortion (ORD) optimal shape coding algorithms to ensure consistent distortion (quality) measurement and improve computational efficiency. It also regulates the memory requirements for an encoder design enabling regular, symmetrical hardware implementations. This paper presents a series of new enhancements to existing techniques for determining the best SW-length within a rate-distortion (RD) framework, and analyses the nexus between SW-length and storage for ORD hardware realizations. In addition, it presents an efficient bit-allocation strategy for managing multiple shapes together with a generalized adaptive SW scheme which integrates localized curvature information (cornerity) on contour points with a bi-directional spatial distance, to afford a superior and more pragmatic SW design compared with existing adaptive SW solutions which are based on only cornerity values. Experimental results consistently corroborate the effectiveness of these new strategies. Ferdous Sohel, Gour C. Karmakar, Laurence Dooley, Mohammed Bennamoun |
IEEE Trans. Multim. | 4 |
| 2012 | Nature-Inspired Techniques in the Context of Fraud DetectionabstractElectronic fraud is highly lucrative, with estimates suggesting these crimes to be worth millions of dollars annually. Because of its complex nature, electronic fraud detection is typically impractical to solve without automation. However, the creation of automated systems to detect fraud is very difficult as adversaries readily adapt and change their fraudulent activities which are often lost in the magnitude of legitimate transactions. This study reviews the most popular types of electronic fraud and the existing nature-inspired detection methods that are used for them. The common characteristics of electronic fraud are examined in detail along with the difficulties and challenges that these present to computational intelligence systems. Finally, open questions and opportunities for further work, including a discussion of emerging types of electronic fraud, are presented to provide a context for ongoing research. Mohammad Behdad, Luigi Barone, Mohammed Bennamoun, Tim French 0002 |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2011 | Supervised particle filter for tracking 2D human pose in monocular videoabstractIn this paper, we propose a hybrid method that combines supervised learning and particle filtering to track the 2D pose of a human subject in monocular video sequences. Our approach, which we call a supervised particle filter method, consists of two steps: the training step and the tracking step. In the training step, we use a supervised learning method to train the regressors that take the silhouette descriptors as input and produce the 2D poses as output. In the tracking step, the output pose estimated from the regressors is combined with the particle filter to track the 2D pose in each video frame. Unlike the particle filter, our method does not require any manual initialization. We have tested our approach using the HumanEva video datasets and compared it with the standard particle filter and 2D pose estimation on individual frames. Our experimental results show that our approach can successfully track the pose over long video sequences and that it gives more accurate 2D human pose tracking than the particle filter and 2D pose estimation. Suman Sedai, Du Q. Huynh, Mohammed Bennamoun |
WACV | 3 |
| 2011 | A pitfall in fingerprint bio-cryptographic key generation
Peng Zhang 0063, Jiankun Hu, Cai Li 0001, Mohammed Bennamoun, B. V. K. Vijaya Kumar |
Comput. Secur. | 4 |
| 2011 | Efficient Detection and Recognition of 3D Ears
Syed M. S. Islam, Rowan Davies, Mohammed Bennamoun, Ajmal Mian |
Int. J. Comput. Vis. | 3 |
| 2011 | Illumination normalization of facial images by reversing the process of image formation
Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
Mach. Vis. Appl. | 2 |
| 2011 | A training-free nose tip detection method from face range images
Xiaoming Peng, Mohammed Bennamoun, Ajmal Mian |
Pattern Recognit. | 2 |
| 2011 | Biometric security for mobile computingabstractThis paper provides an editorial statement for the special issue on biometric security in mobile computing environment. Jiankun Hu, B. V. K. Vijaya Kumar, Mohammed Bennamoun, Kar-Ann Toh |
Secur. Commun. Networks | 3 |
| 2010 | An HMM-SVM-Based Automatic Image Annotation Approach
Yinjie Lei, Wilson Wong, Wei Liu 0006, Mohammed Bennamoun |
ACCV (4) | 4 |
| 2010 | Semi-supervised Neighborhood Preserving Discriminant Embedding: A Semi-supervised Subspace Learning Algorithm
Maryam Mehdizadeh, Cara MacNish, R. Nazim Khan, Mohammed Bennamoun |
ACCV (3) | 4 |
| 2010 | Localized fusion of Shape and Appearance features for 3D Human Pose Estimation
Suman Sedai, Mohammed Bennamoun, Du Q. Huynh |
BMVC | 2 |
| 2010 | On the problems of using learning classifier systems for fraud detectionabstractFraud detection problems have some uniquely challenging properties which make them difficult. In this paper, we investigate the fraud detection problem by describing the common properties of electronic fraud and examining how learning classifier systems (LCSs) can be applied to it. Also, we introduce "random Boolean function" (RBF); an abstract problem with high level of controllability which can be tuned to exhibit those characteristics individually, and report the results of using XCSR (a continuous variant of LCS) on RBF problem and also on a real-world problem. Results from our experiments demonstrate that XCSR can overcome most of the difficulties inherent to the fraud detection problem and can achieve good performance in case of the real-world problem. Mohammad Behdad, Tim French 0002, Luigi Barone, Mohammed Bennamoun |
GECCO | 4 |
| 2010 | Probabilistic human pose recovery from 2D imagesabstractImage based human pose recovery has many applications in different industries such as games, entertainment, physiological rehabilitation and biometrics. This paper presents a new pose estimation algorithm from monocular images based on a nonlinear mapping of human silhouettes, coded using a collection of local image moments, to the pose space using a mixture of Neural Networks (NN) regressors. All parameters are estimated automatically. Experiments and comparative results show a superior performance of the proposed method. Farid Flitti, Mohammed Bennamoun, Du Q. Huynh, Robyn A. Owens |
ICIP | 2 |
| 2010 | Robust Regression for Face RecognitionabstractIn this paper we address the problem of illumination invariant face recognition. Using a fundamental concept that in general, patterns from a single object class lie on a linear subspace, we develop a linear model representing a probe image as a linear combination of class-specific galleries. In the presence of noise, the well-conditioned inverse problem is solved using the robust Huber estimation and the decision is ruled in favor of the class with the minimum reconstruction error. The proposed Robust Linear Regression Classification (RLRC) algorithm is extensively evaluated for two standard databases and has shown good performance index compared to the state-of-art robust approaches. Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
ICPR | 3 |
| 2010 | Sparse Representation for Speaker IdentificationabstractWe address the closed-set problem of speaker identification by presenting a novel sparse representation classification algorithm. We propose to develop an over complete dictionary using the GMM mean super vector kernel for all the training utterances. A given test utterance corresponds to only a small fraction of the whole training database. We therefore propose to represent a given test utterance as a linear combination of all the training utterances, thereby generating a naturally sparse representation. Using this sparsity, the unknown vector of coefficients is computed via l1-minimization which is also the sparsest solution. Ideally, the vector of coefficients so obtained has nonzero entries representing the class index of the given test utterance. Experiments have been conducted on the standard TIMIT database and a comparison with the state-of-art speaker identification algorithms yields a favorable performance index for the proposed algorithm. Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
ICPR | 3 |
| 2010 | On the Repeatability and Quality of Keypoints for Local Feature-based 3D Object Retrieval from Cluttered Scenes
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
Int. J. Comput. Vis. | 2 |
| 2010 | Drift-correcting template update strategy for precision feature point tracking
Xiaoming Peng, Mohammed Bennamoun, Qiheng Zhang, Wufan Chen |
Image Vis. Comput. | 2 |
| 2010 | Linear Regression for Face RecognitionabstractIn this paper, we present a novel approach of face identification by formulating the pattern recognition problem in terms of linear regression. Using a fundamental concept that patterns from a single-object class lie on a linear subspace, we develop a linear model representing a probe image as a linear combination of class-specific galleries. The inverse problem is solved using the least-squares method and the decision is ruled in favor of the class with the minimum reconstruction error. The proposed Linear Regression Classification (LRC) algorithm falls in the category of nearest subspace classification. The algorithm is extensively evaluated on several standard databases under a number of exemplary evaluation protocols reported in the face recognition literature. A comparative study with state-of-the-art algorithms clearly reflects the efficacy of the proposed approach. For the problem of contiguous occlusion, we propose a Modular LRC approach, introducing a novel Distance-based Evidence Fusion (DEF) algorithm. The proposed methodology achieves the best results ever reported for the challenging problem of scarf occlusion. Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Probabilistic Satellite Image Fusion
Farid Flitti, Mohammed Bennamoun, Du Q. Huynh, Amine Bermak, Christophe Collet 0001 |
CAIP | 2 |
| 2009 | Face identification using linear regressionabstractIn this paper we present a novel approach of face identification by formulating the pattern recognition problem in terms of linear regression. Using a fundamental concept that patterns from a single object class lie on a linear subspace, we develop a linear model representing a probe image as a linear combination of class specific galleries. The inverse problem is solved using the least squares method and the decision is ruled in favor of the class with the minimum reconstruction error. The algorithm is extensively evaluated using two standard databases, a comparative study with the benchmark algorithms clearly reflects the efficacy of the proposed approach. Imran Naseem, Roberto Togneri, Mohammed Bennamoun |
ICIP | 3 |
| 2009 | Acquiring Semantic Relations Using the Web for Constructing Lightweight Ontologies
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun |
PAKDD | 3 |
| 2009 | A probabilistic framework for automatic term recognitionabstractTerm recognition identifies domain-relevant terms which are essential for discovering domain concepts and for the construction of terminologies required by a wide range of natural language applications. Many techniques have been developed in an attem Wilson Wong, Wei Liu 0006, Mohammed Bennamoun |
Intell. Data Anal. | 3 |
| 2009 | An Expression Deformation Approach to Non-rigid 3D Face Recognition
Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
Int. J. Comput. Vis. | 2 |
| 2009 | An extension of min/max flow framework
Hongchuan Yu, Mohammed Bennamoun, Chin-Seng Chua |
Image Vis. Comput. | 2 |
| 2008 | A Fast and Fully Automatic Ear Recognition Approach Based on 3D Local Surface Features
Syed M. S. Islam, Rowan Davies, Ajmal Mian, Mohammed Bennamoun |
ACIVS | 4 |
| 2008 | Determining the Unithood of Word Sequences Using a Probabilistic Approach
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun |
IJCNLP | 3 |
| 2008 | Fast and Fully Automatic Ear Detection Using Cascaded AdaBoostabstractEar detection from a profile face image is an important step in many applications including biometric recognition. But accurate and rapid detection of the ear for real-time applications is a challenging task, particularly in the presence of occlusions. In this work, a cascaded AdaBoost based ear detection approach is proposed. In an experiment with a test set of 203 profile face images, all the ears were accurately detected by the proposed detector with a very low (5 × 10-6) false positive rate. It is also very fast and relatively robust to the presence of occlusions and degradation of the ear images (e.g. motion blur). The detection process is fully automatic and does not require any manual intervention. Syed M. S. Islam, Mohammed Bennamoun, Rowan Davies |
WACV | 2 |
| 2008 | Keypoint Detection and Local Feature Matching for Textured 3D Face Recognition
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
Int. J. Comput. Vis. | 2 |
| 2008 | Integration of local and global geometrical cues for 3D face recognition
Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
Pattern Recognit. | 2 |
| 2007 | Interest-point Based Face Recognition from Range ImagesabstractWe present a novel approach to interest-point detection tailored to range images. A range image is represented by two images with blob-like patterns that have easily detectable peaks and can be efficiently extracted using convolution kernels. These kernels were designed to produce repeatable and independent blob-like patterns when convolved with the range image. The interest-points correspond to peaks of the patterns after dropping the unstable ones and performing Non-Maximal Suppression (NMS) on their union. The approach was applied to facial range images from the FRGC V2.0 dataset and about 88% repeatability was achieved. Face recognition was also performed by matching the local range regions around the interest-points. An approach based on three levels of matching combined with RAN SAC algorithm was used to increase the correct matches and reduce the false ones. Preliminary recognition results for a database of 466 subjects and 1765 probes were 96.33% identification rate and 90% verification rate at 0.1% False Accept Rate (FAR) for faces under neutral expression. Faisal R. Al-Osaimi, Mohammed Bennamoun, Ajmal Mian |
BMVC | 2 |
| 2007 | Tree-Traversing Ant Algorithm for term clustering based on featureless similarities
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun |
Data Min. Knowl. Discov. | 3 |
| 2007 | An Efficient Multimodal 2D-3D Hybrid Approach to Automatic Face RecognitionabstractWe present a fully automatic face recognition algorithm and demonstrate its performance on the FRGC v2.0 data. Our algorithm is multimodal (2D and 3D) and performs hybrid (feature-based and holistic) matching in order to achieve efficiency and robustness to facial expressions. The pose of a 3D face along with its texture is automatically corrected using a novel approach based on a single automatically detected point and the Hotelling transform. A novel 3D Spherical Face Representation (SFR) is used in conjunction with the SIFT descriptor to form a rejection classifier which quickly eliminates a large number of candidate faces at an early stage for efficient recognition in case of large galleries. The remaining faces are then verified using a novel region-based matching approach which is robust to facial expressions. This approach automatically segments the eyes-forehead and the nose regions, which are relatively less sensitive to expressions, and matches them separately using a modified ICP algorithm. The results of all the matching engines are fused at the metric level to achieve higher accuracy. We use the FRGC benchmark to compare our results to other algorithms which used the same database. Our multimodal hybrid algorithm performed better than others by achieving 99.74% and 98.31% verification rates at 0.001 FAR and identification rates of 99.02% and 95.37% for probes with neutral and non-neutral expression respectively. Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Complete invariants for robust face recognition
Hongchuan Yu, Mohammed Bennamoun |
Pattern Recognit. | 2 |
| 2006 | 2D and 3D Multimodal Hybrid Face Recognition
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
ECCV (3) | 2 |
| 2006 | Text-independent speaker identification in birds
E. J. S. Fox, J. D. Roberts, Mohammed Bennamoun |
INTERSPEECH | 3 |
| 2006 | A Novel Representation and Feature Matching Algorithm for Automatic Pairwise Registration of Range Images
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
Int. J. Comput. Vis. | 2 |
| 2006 | Three-Dimensional Model-Based Object Recognition and Segmentation in Cluttered ScenesabstractViewpoint independent recognition of free-form objects and their segmentation in the presence of clutter and occlusions is a challenging task. We present a novel 3D model-based algorithm which performs this task automatically and efficiently. A 3D model of an object is automatically constructed offline from its multiple unordered range images (views). These views are converted into multidimensional table representations (which we refer to as tensors). Correspondences are automatically established between these views by simultaneously matching the tensors of a view with those of the remaining views using a hash table-based voting scheme. This results in a graph of relative transformations used to register the views before they are integrated into a seamless 3D model. These models and their tensor representations constitute the model library. During online recognition, a tensor from the scene is simultaneously matched with those in the library by casting votes. Similarity measures are calculated for the model tensors which receive the most votes. The model with the highest similarity is transformed to the scene and, if it aligns accurately with an object in the scene, that object is declared as recognized and is segmented. This process is repeated until the scene is completely segmented. Experiments were performed on real and synthetic data comprised of 55 models and 610 scenes and an overall recognition rate of 95 percent was achieved. Comparison with the spin images revealed that our algorithm is superior in terms of recognition rate and efficiency. Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Region-based Matching for Robust 3D Face RecognitionabstractWe present a novel region-based matching approach for automatic 3D face recognition which is robust to facial expressions, facial hair, illumination changes and large occlusions. Each 3D face in the gallery is segmented offline into three disjoint regions, namely eyes-forehead, nose and cheeks. Recognition is performed on the basis of only the eyes-forehead and nose regions to avoid the effects of expressions and artifacts that occur in 3D faces due to a mustache or beard. These two regions of the gallery are matched with a probe using a modified version of the ICP algorithm and their matching scores are fused. The identity of the gallery face which gets the highest score is declared as the identity of the probe. Experiments were performed on the UND Biometrics Database which is so far the largest known database of 3D faces. We achieved a combined identification rate of 100% and a maximum verification rate of 99.42%. Our results also show that the eyes-forehead is the most significant region for 3D face recognition with individual identification and verification rates of 97.32% and 97.25% respectively. Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
BMVC | 2 |
| 2005 | A Phase Correlation Approach to Active Vision
Hongchuan Yu, Mohammed Bennamoun |
CAIP | 2 |
| 2004 | Matching Tensors for Automatic Correspondence and Registration
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
ECCV (2) | 2 |
| 2004 | Performance analysis of an improved tensor based correspondence algorithm for automatic 3d modeling
Ajmal Mian, Mohammed Bennamoun, Robyn A. Owens |
ICIP | 2 |
| 2003 | Editorial: Correspondence And Registration Techniques
Mohammed Bennamoun |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | Rigid Medical Image Registration And Its Association With Mutual InformationabstractImage registration plays a crucial role in the computer vision and medical imaging field where it is used to develop a spatial mapping between different sets of data. These transformations can range from simple rigid registrations to complex nonrigid deformations. Mutual information (MI) is a popular entropy-based similarity measure which has recently experienced a prolific expansion in a number of image registration applications. Stemming from information theory, this measure generally outperforms most other intensity-based measures in multimodal applications as it only assumes a statistical dependence between images. This paper provides a thorough introduction to the MI measure and its use in rigid medical image registration. A look at the extensions proposed to the original measure will also be provided. These were developed to improve the robustness of the measure and to avoid certain cases when maximizing MI does not lead to the correct spatial alignment. Clinton Fookes, Mohammed Bennamoun |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2002 | Optimal Gabor filters for textile flaw detection
Adriana Bodnarova, Mohammed Bennamoun, Shane J. Latham |
Pattern Recognit. | 2 |
| 2001 | A 3D acquisition and modelling systemabstractThis paper presents our implementation of an integrated 3D acquisition and modelling system. The proposed system consists of an acquisition module, a registration algorithm proposed by J.A. Williams and M. Bennamoun (see IEEE Int. Conf. on Acoustics, Speech and Sig. Proc.-ICASSP2000, p.2049-52, 2000; Computer Vision and Image Understanding, January 2001), and a reconstruction module. Techniques for addressing the implementation of the modules are first briefly described, then a more detailed discussion of the techniques implemented in the system is given. Results are also presented to demonstrate the system operation. Suhail Mahadevan, Haris Pandzo, Mohammed Bennamoun, John A. Williams 0001 |
ICASSP | 3 |
| 2001 | Automatic Bayesian knot placement for spline fittingabstractWe propose a Bayesian model for automatically determining knot placement in spline modelling. The random variables of the model are the number of knots and their locations, which we seek to estimate via a simulated annealing form of the reversible jump Markov chain Monte Carlo sampler. This novel technique has the ability to maximise the joint posterior distribution of the number of knots and their locations, without becoming stranded on local maxima. We provide results which verify the effectiveness of the proposed technique, in accurately fitting a non-uniform, cubic spline to data, whilst maintaining a relatively small number of knots. George J. Mamic, Mohammed Bennamoun |
ICIP (1) | 2 |
| 2001 | Simultaneous Registration of Multiple Corresponding Point Sets
John A. Williams 0001, Mohammed Bennamoun |
Comput. Vis. Image Underst. | 2 |
| 2001 | An Arabic optical character recognition system using recognition-based segmentation
A. Cheung, Mohammed Bennamoun, Neil W. Bergmann |
Pattern Recognit. | 2 |
| 2001 | Reliability analysis of the rank transform for stereo matchingabstractThe rank transform is a nonparametric technique which has been recently proposed for the stereo matching problem. The motivation behind its application to this problem is its invariance to certain types of image distortion and noise, as well as its amenability to real-time implementation. This paper derives one constraint which must be satisfied for a correct match. This has been termed the rank constraint. Experimental work has shown that this constraint is capable of resolving ambiguous matches, thereby improving matching reliability. A novel matching algorithm incorporating the rank constraint has also been proposed. This modified algorithm consistently resulted in an increased percentage of correct matches, for all test imagery used. Furthermore, the rank constraint has been used to devise a method of identifying regions of an image where the rank transform, and hence matching outcome, is more susceptible to noise. Experimental results have shown that the errors predicted using this technique are consistent with the actual errors which result when images are corrupted with noise. Such a method could be used to identify matches which are likely to be incorrect and/or provide a measure of confidence in a match. Jasmine Banks, Mohammed Bennamoun |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2000 | A constrained minimisation approach to optimise Gabor filters for detecting flaws in woven textilesabstractGabor filters have proved to be an effective segmentation and flaw detection tool. This study addresses the issue of an optimal 2-D Gabor filter design for automatically detecting defects in homogeneously textured woven fabrics. The parameters of these filters are derived through an optimisation process performing the minimisation of a Fisher cost function. By constraining some of the Gabor filter parameters to specific values the aim is to optimise the filter to detect a certain type of flaw as it appears in a particular textile background. To account for the potentially large variety of flaw types, the optimal parameters for multiple sets of constraints are computed. The detection outcomes from each set of optimal filters are combined to produce a final classification result. Successful detection results (with low false alarm rates) suggest that this optimal Gabor filter approach is a promising method for automated detection of flaws in homogenous textiles. Adriana Bodnarova, Mohammed Bennamoun, Shane J. Latham |
ICASSP | 2 |
| 2000 | Simultaneous registration of multiple point sets using orthonormal matricesabstractWe present a novel solution to the problem of simultaneously registering multiple corresponding point sets with rigid transformations. The proposed technique involves the pre-computation of a constant matrix which completely encodes the registration problem, followed by an iterative procedure to estimate the optimal rotations. From the estimated rotations, the optimal translations are then computed directly. Experimental results on both synthetic and real data demonstrate the algorithm's performance in terms of accuracy and efficiency. John A. Williams 0001, Mohammed Bennamoun |
ICASSP | 2 |
| 2000 | Global 3D Rigid Registration of Medical ImagesabstractWe present in this paper an iterative algorithm for the simultaneous registration of multiple 3D medical images. The proposed algorithm is a point-based registration method and is based on global registration techniques rather than the traditional pair-wise registration methods. Corresponding feature points, known as extremal points, are first automatically extracted from the 3D images and are used as the matching features in the registration process. These extremal points are stable landmarks, as the relative positions of these points are known to be invariant according to 3D rigid transformations. The registration algorithm is based on a novel weighted least squares formulation and it also incorporates 3D noise models on the extracted feature points. Results are presented for the 3D rigid registration of three successive MR images of the same patient taken at different periods of time. Clinton Fookes, John A. Williams 0001, Mohammed Bennamoun |
ICIP | 3 |
| 2000 | Multiple View Surface Registration with Error Modeling and AnalysisabstractWe describe a new 3D surface rigid registration system which allows the assignment of individual measurement error models to each surface point. Any number of surfaces may be registered simultaneously. The error models are used in the registration to compute statistically optimal registration parameters, as well as to estimate the error covariance of those parameters. Previous surface registration techniques have implicitly assumed an isotropic Gaussian measurement error distribution for all points. However, that does not reflect the nature of non-contact 3D surface measurements, whose error characteristics are often strongly directional. John A. Williams 0001, Mohammed Bennamoun |
ICIP | 2 |
| 2000 | Textile Flaw Detection Using Optimal Gabor FiltersabstractThis study presents a new automatic and fast approach to design optimised Gabor filters for textile flaw detection applications. The defect detection problem is solved by using a semi-supervised approach. The aim is to automatically discriminate between "known" nondefective background textures and "unknown" defective textures. The parameters of the optimal 2D Gabor filters are derived by constrained minimisation of a Fisher cost function. Such optimised Gabor filters are capable of detecting both structural and tonal defects. This adaptable approach can detect a large variety of flaw types, while at the same time, according for their changing appearance in different texture backgrounds. When applied to a large database of textile fabrics, accurate detection with a low false alarm rate was achieved. Adriana Bodnarova, Mohammed Bennamoun, Shane J. Latham |
ICPR | 2 |
| 2000 | Accurate Localization of Edges in Noisy Volume ImagesabstractAdvances in medical imaging modalities have made it possible to acquire volume images. One of the key steps used to split the raw volume image into meaningful sub-volumes is 3D edge detection. This paper describes a new approach for 3D edge detection. It is based on the 2D hybrid edge detector, which consists of a combination of the first and second order differential edge detectors. It was shown in the 2D case that the combination of the two differential edge detectors gave an accurate edge localisation whilst maintaining immunity to the noise in the image. Results using the 3D hybrid edge detector based on synthetic and real images are presented. They are also compared to the results using the 3D gradient of the Gaussian detector and the 3D Laplacian of the Gaussian detector. Pi-chi Chou, Mohammed Bennamoun |
ICPR | 2 |
| 2000 | Automatic Flaw Detection in Textiles Using a Neyman-Pearson DetectorabstractA system for the automated visual inspection of textiles is discussed. The system consists of two main components, (1) the extraction of the texture features utilising the Karhunen-Loeve (KL) transform which provides optimal compression of the image data into a feature vector and (2) the detection of the flaw patterns using a Neyman-Pearson detector, which maximises the rate of detection for a specified false alarm rate. The performance of the system was evaluated on various fabrics and different types of textile flaws. The results indicate that the system can detect flaws which vary drastically in physical dimension and nature with a very low false alarm rate. Experimental results in the paper demonstrate the performance of the detector on some typical textile flaws. George J. Mamic, Mohammed Bennamoun |
ICPR | 2 |
| 2000 | Evaluation of a Novel Multiple Point Set Registration AlgorithmabstractWe briefly describe a novel solution to the simultaneous multiple view point registration problem and evaluate its performance in a practical machine vision application. The algorithm involves computation of a constant matrix which encodes the point set registration, from which the optimal rigid transformations are found using a simple and efficient iterative algorithm. We describe this algorithm's incorporation into a multiple view surface registration system, and present results from experiments on two different sets of real surface data. These results demonstrate that the proposed algorithm is accurate and efficient. John A. Williams 0001, Mohammed Bennamoun |
ICPR | 2 |
| 2000 | Three-dimensional hybrid edge detection
Mohammed Bennamoun, Pi-chi Chou, Espen Norheim, Michael O'Loan |
VCIP | 1 |
| 2000 | Review of 3D object representation techniques for automatic object recognition
George J. Mamic, Mohammed Bennamoun |
VCIP | 2 |
| 2000 | Suitability Analysis of Techniques for Flaw Detection in Textiles using Texture Analysis
Adriana Bodnarova, Mohammed Bennamoun, Kurt Kubik |
Pattern Anal. Appl. | 2 |
| 1999 | A constraint to improve the reliability of stereo matching using the rank transformabstractThe rank transform is a non-parametric technique which has been previously proposed for the stereo matching problem. The motivation behind its application to the matching problem is its invariance to certain types of image distortion and noise, as well as its amenability to real-time implementation. This paper derives an analytic expression for the process of matching using the rank transform, and then goes on to derive one constraint which must be satisfied for a correct match. This has been dubbed the rank order constraint or simply the rank constraint. Experimental work has shown that this constraint is capable of resolving ambiguous matches, thereby improving matching reliability. This constraint was incorporated into a new algorithm for matching using the rank transform. This modified algorithm resulted in an increased proportion of correct matches, for all test imagery used. Jasmine Banks, Mohammed Bennamoun, Kurt Kubik, Peter I. Corke |
ICASSP | 2 |
| 1998 | An Extended Kalman Filtering Approach to High Precision Stereo Image MatchingabstractWe present a novel approach to stereo image matching for high precision applications which is based upon a non-linear filtering technique called the extended Kalman filter (EKF). The matching algorithm has three components-the matching process, false match rejection, and disparity prediction, which are all derived within the Kalman filtering framework. We first present the stereo matching model used, and then derive the matching equations and processes. We then present results of tests performed on a synthetic stereo pair, which allows comparison with ground truth data. The results indicate that the method is capable of very robust and high precision matching performance. John A. Williams 0001, Mohammed Bennamoun |
ICIP (2) | 2 |
| 1998 | Image segmentation and image matching for 3D terrain reconstructionabstractA 3D terrain reconstruction method using compound techniques is proposed. Normal matching results only supply a DSM (digital surface model). This means that the matching results may be on the top of objects such as houses or trees. This kind of matching results could not supply an accurate DEM (digital elevation model). The proposed method is a more efficient method for determining elevations from overlapping digital aerial images and satellite images. It combines image analysis and image matching methods and supplies a more accurate DEM. Yihui Lu, Kurt Kubik, Mohammed Bennamoun |
ICPR | 3 |
| 1998 | A non-linear filtering approach to image matchingabstractIn this paper we present a novel approach to high precision stereo image matching which employs an iterated extended Kalman filter (IEKF). The matching problem is formulated as a nonlinear filtering problem. We present a state-space representation of the matching problem, and show how nonlinear filtering techniques may be applied to determine optimal transformation parameters. To demonstrate the efficacy of this approach we use an IEKF for matching one-dimensional pixel intensity profiles. The results indicate that the IEKF technique is capable of high accuracy and precision. John A. Williams 0001, Mohammed Bennamoun |
ICPR | 2 |
| 1998 | Automatic visual inspection and flaw detection in textile materials: past, present and futureabstractThis paper provides a synthesis of a number of textile flaw detection techniques. The review highlights the issues pertaining to textile visual quality inspection and outlines its objectives and problems. Furthermore it looks at the general taxonomy of texture analysis approaches and their suitability for the task of textile quality inspection. The paper also details and provides a comparison of some of our most recent solutions to this problem and addresses the issues of their specifications for the real-time implementation. Mohammed Bennamoun, Adriana Bodnarova |
SMC | 1 |
| 1998 | Defect detection in textile materials based on aspects of the HVSabstractThe problem we address in this paper is that of detecting flaws in woven textile fabrics. This task is currently performed by human inspectors with a maximum detection rate of only about 80%. An Automated Visual Inspection System (AVIS) is potentially a more reliable, objective, time and cost effective solution. Using computer vision it detects irregularities in homogeneously structured images of textile fabrics. However, its successful performance is largely dependent of the choice of an accurate and robust texture analysis algorithm. The technique of blob detection described in this paper is representative of a structural texture analysis approach and it accounts for aspects of the human visual system (HVS) in detecting large varieties of textile flaws. Measures of texture discrimination based on psychophysical experiments are used to indicate the levels of perceivable differences in blob attributes indicating the presence of defects. Defects are detected by observing a threshold of acceptable differences in the properties of blobs based on human perception. Adriana Bodnarova, Mohammed Bennamoun, Kurt Kubik |
SMC | 2 |
| 1998 | A recognition-based Arabic optical character recognition systemabstractOptical character recognition systems improve human-machine interaction and are widely used in many government and commercial departments. After forty years of intensive research, OCR systems for most scripts are well developed. However, not for Arabic script. Since Arabic is a popular script, Arabic OCR systems should have great commercial value. Thus a recognition-based Arabic OCR system is proposed in this paper. It consists of the image acquisition, preprocessing, segmentation, character fragmentation, combination of character fragments, feature extraction, and classification. A signal is fed back to improve and determine the segmentation/recognition result. The system has been implemented and it has 90% recognition accuracy with a 20 chars/sec recognition rate. A. Cheung, Mohammed Bennamoun, Neil W. Bergmann |
SMC | 2 |
| 1998 | Handwritten character recognition by contour sequence moments and neural networkabstractContour sequence moments (CSM) have been used in the classification of four closed planar shapes. Gupta et al. described a neural network approach for the classification of four closed planar shapes using a contour sequence. In this paper, a backpropagation neural network is used in the recognition of handwritten numerals (from 0 to 9) using contour sequence moments. Experimental results indicate that the neural network approach gives better recognition accuracy when compared with the two conventional statistical classifiers, namely the nearest neighbour and minimum-mean-distance. This CSM technique was compared with geometrical moment (GM) invariants. We found that the recognition accuracy for handwritten character using GSM and neural network is over 95% while GM invariants and neural network can only give 82%. Vera Chung, Man To Wong, Mohammed Bennamoun |
SMC | 3 |
| 1998 | Implementing neural network in custom computersabstractThis paper describes the implementation of a partially connected neural network using FPGAs (field programmable gate arrays) based custom computers. Starting from the training data, a decision tree is generated using the classifier program C4.5. The tree is then used to initialise the architecture of the neural network to a nearly optimum configuration. This initialised partially connected network is then trained using training data. The trained neural network is then implemented by fine-grain Xilinx XC6200 series FPGAs. This implementation requires fewer connections and can provide a very high speed classification for many real-time image recognition applications. Vera Chung, Man To Wong, Neil W. Bergmann, Mohammed Bennamoun |
SMC | 4 |
| 1997 | Application of Time-Frequency Signal Analysis to Motion EstimationabstractIn this paper, a variant of a motion estimation technique for multiple velocities in an image neighbourhood is described. The proposed method ties in the concepts of motion estimation, instantaneous frequency and time-frequency distributions. A very brief overview about time-frequency signal analysis (TFSA) aimed to the reader who is not familiar with TFSA is given. Results are reported with the important factors to be considered when applying this technique. Mohammed Bennamoun |
ICIP (2) | 1 |
| 1997 | A structural-description-based vision system for automatic object recognitionabstractThis paper presents the results of the integration of a proposed part-segmentation-based vision system. The first stage of this system extracts the contour of the object using a hybrid first- and second-order differential edge detector. The object defined by its contour is then decomposed into its constituent parts using the part segmentation algorithm given by Bennamoun (1994). These parts are then isolated and modeled with 2D superquadrics. The parameters of the models are obtained by the minimization of a best-fit cost function. The object is then represented by its structural description which is a set of data structures whose predicates represent the constituent parts of the object and whose arguments represent the spatial relationship between these parts. This representation allows the recognition of objects independently of their positions, orientations, or sizes. It is also insensitive to objects with partially missing parts. In this paper, examples illustrating the acquired images of objects, the extraction of their contours, the isolation of the parts, and their fitting with 2D superquadrics are reported. The reconstruction of objects from their structural description is illustrated and improvements are suggested. Mohammed Bennamoun, Boualem Boashash |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1994 | A contour-based part segmentation algorithmabstractWithin the framework of a suggested vision system, a new part segmentation algorithm based on the contour of an object, assumed to be obtained from an edge detector is presented. A new approach is presented to extract the convex dominant points (CDPs) of the contour which are used to decompose the object into convex parts. Among all the points of the contour only the CDPs are moved along their normal until they touch another moving CDP or a point on the contour. The algorithm has been tested on many contours of objects with and without noise. The results are compared to other techniques based on a graph theoretic approach. The results appear to give a closer decomposition to the one performed by humans.> Mohammed Bennamoun |
ICASSP (5) | 1 |
| 1994 | Integrataion of a Part Segmentation Based Vision SystemabstractThis paper presents the results of the integration of a proposed part-segmentation based vision system. The first stage of this system extracts the contour of the object using a hybrid first and second-order differential edge detector. The object defined by its contour is then decomposed into its constituent parts using the part segmentation algorithm proposed by Bennamoun (see Proc. of the IEEE ICASSP'94, p.41-44, Adelaide, Australia, April 1994). These parts are then isolated and modelled with 2-D superquadrics. The parameters of the models are obtained by the minimization of a best-fit cost function. The object is then represented by its structural description which is a set of data structures whose predicates represent the constituent parts of the object and whose arguments represent the spatial relationship between these parts. This representation allows the recognition of objects independently of their positions, orientations or sizes. It is also insensitive to objects with partially missing parts. In this paper, examples illustrating acquired images of objects, the extraction of their contours, the isolation of the parts, and their fitting with 2-D superquadrics are reported. The reconstruction of objects from their structural description is illustrated and improvements are suggested.> Mohammed Bennamoun, Boualem Boashash |
ICIP (3) | 1 |
| 1991 | Avoidance of unknown obstacles using proximity fieldsabstractPresents a novel real-time obstacle avoidance approach based on proximity sensing. The scheme is designed for the two dimensional navigation of a point object in a totally unknown environment. Navigation is performed by utilizing two proximity sensors of different settings in conjunction with a simple memory less rule for motion control. Simulation results are given for different types of obstacles. The algorithm has also been implemented on an Adept-1 arm manipulator.> Mohammed Bennamoun, Ahmad A. Masoud, M. A. Ramsay, Mohamed M. Bayoumi |
IROS | 1 |