Farid Boussaïd

dblp:04/6354 · DBLP profile ↗
← Back
75ranked-venue papers
0as first author
38since 2021 · last 2026
0000-0001-7250-7407ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 15 since 2021Systems, architecture and hardware · 14 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
abstract
Zero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton features with static, class-level semantics. This coarse-grained alignment fails to bridge the domain shift between seen, unseen classes, thereby impeding the effective transfer of fine-grained visual knowledge. To address these limitations, we introduce DynaPURLS, a unified framework that establishes robust, multi-scale visual-semantic correspondences, dynamically refines them at inference time to enhance generalization. Our framework leverages a large language model to generate hierarchical textual descriptions that encompass both global movements, local body-part dynamics. Concurrently, an adaptive partitioning module produces fine-grained visual representations by semantically grouping skeleton joints. To fortify this fine-grained alignment against the train-test domain shift, DynaPURLS incorporates a dynamic refinement module. During inference, this module adapts textual features to the incoming visual stream via a lightweight learnable projection. This refinement process is stabilized by a confidence-aware, class-balanced memory bank, which mitigates error propagation from noisy pseudo-labels. Extensive experiments on three large-scale benchmark datasets, including NTU RGB+D 60/120, PKU-MMD, demonstrate that DynaPURLS significantly outperforms prior art, setting new state-of-the-art records.
Jingmin Zhu, James Bailey 0001, Jun Liu 0036, Hossein Rahmani 0001, Mohammed Bennamoun, Farid Boussaïd, Qiuhong Ke
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 HAViG: Hierarchical adaptive visual grounding framework for video question answering
Lei Zhu 0005, Lingmin Pan, Siqiao Tan, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Farid Boussaïd, Mohammed Bennamoun
Pattern Recognit.7
2026 SPA: Stable and Precise Alignment for Efficient Cross-Domain Palmprint Recognition
abstract
Palmprint recognition has been extensively studied as an effective biometric technique for personal identification. With the rapid development of deep neural networks (DNNs), palmprint recognition methods have achieved remarkable progress. However, their performance often deteriorates significantly under domain shifts. Moreover, existing unsupervised domain adaptation approaches for palmprint recognition typically suffer from unstable training and imprecise feature alignment, thereby limiting their effectiveness. To address these challenges, we propose SPA, a Stable and Precise Alignment framework for cross-domain palmprint recognition. Specifically, we design a lightweight yet robust Style Transformation Module (STM) to mitigate variations in style, color, and illumination. With the aid of STM, we further align joint feature distributions across all high-level layers, achieving more accurate feature alignment and enhancing recognition robustness. We conduct extensive experiments on two public multi-domain palmprint databases encompassing 42 cross-domain scenarios. The results demonstrate that SPA consistently delivers superior performance across both databases, achieving higher recognition accuracy with lower computational overhead compared to existing methods. In particular, SPA improves the average identification accuracies to 94.21% and 81.93%, while reducing the average equal error rates (EER) to 1.36% and 3.62% on the two databases, respectively.
Song Ruan, Yantao Li 0001, Huafeng Qin, Naeha Sharif, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Inf. Forensics Secur.5
2025 Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis
abstract
We propose a novel framework for the statistical analysis of genus-zero 4D surfaces, i.e., 3D surfaces that deform and evolve over time. This problem is particularly challenging due to the arbitrary parameterizations of these surfaces and their varying deformation speeds, necessitating effective spatiotemporal registration. Traditionally, 4D surfaces are discretized, in space and time, before computing their spatiotemporal registrations, geodesics, and statistics. However, this approach may result in suboptimal solutions and, as we demonstrate in this paper, is not necessary. In contrast, we treat 4D surfaces as continuous functions in both space and time. We introduce Dynamic Spherical Neural Surfaces (D-SNS), an efficient smooth and continuous spatiotemporal representation for genus-0 4D surfaces. We then demonstrate how to perform core 4D shape analysis tasks such as spatiotemporal registration, geodesics computation, and mean 4D shape estimation, directly on these continuous representations without upfront discretization and meshing. By integrating neural representations with classical Riemannian geometry and statistical shape analysis techniques, we provide the building blocks for enabling full functional shape analysis. We demonstrate the efficiency of the framework on 4D human and face datasets. The source code and additional results are available at https://4d-dsns.github.io/DSNS/.
Awais Nizamani, Hamid Laga, Guanjin Wang, Farid Boussaïd, Mohammed Bennamoun, Anuj Srivastava
CVPR4
2025 Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
abstract
Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like “A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the audience nodding and applauding” requires simultaneous processing of visual, audio, and speech signals. However, existing models often struggle to effectively fuse and interpret audio information, limiting their capacity for comprehensive video temporal understanding. To address this, we present TriSense, a triple-modality large language model designed for holistic video temporal understanding through the integration of visual, audio, and speech modalities. Central to TriSense is a Query-Based Connector that adaptively reweights modality contributions based on the input query, enabling robust performance under modality dropout and allowing flexible combinations of available inputs. To support TriSense's multimodal capabilities, we introduce TriSense-2M, a high-quality dataset of over 2 million curated samples generated via an automated pipeline powered by fine-tuned LLMs. TriSense-2M includes long-form videos and diverse modality combinations, facilitating broad generalization. Extensive experiments across multiple benchmarks demonstrate the effectiveness of TriSense and its potential to advance multimodal video analysis.
Zinuo Li, Yongxin Guo 0001, Mohammed Bennamoun, Farid Boussaïd, Girish Dwivedi, Luqi Gong, Qiuhong Ke
NeurIPS5
2025 Memory guided representation learning for cross-domain face anti-spoofing
Pengchao Deng, Zhiheng Fu, Shengjun Xu, Chenyang Ge, Farid Boussaïd, Mohammed Bennamoun
Eng. Appl. Artif. Intell.7
2025 Generalized Closed-Form Formulae for Feature-Based Subpixel Alignment in Patch-Based Matching
abstract
Abstract Patch-based matching is a technique meant to measure the disparity between pixels in a source and target image and is at the core of various methods in computer vision. When the subpixel disparity between the source and target images is required, the cost function or the target image has to be interpolated. While cost-based interpolation is easier to implement, multiple works have shown that image-based interpolation can increase the accuracy of the disparity estimate. In this paper we review closed-form formulae for subpixel disparity computation for one dimensional matching, e.g., rectified stereo matching, for the standard cost functions used in patch-based matching. We then propose new formulae to generalize to high-dimensional search spaces, which is necessary for unrectified stereo matching and optical flow. We also compare the image-based interpolation formulae with traditional cost-based formulae, and show that image-based interpolation brings a significant improvement over the cost-based interpolation methods for two dimensional search spaces, and small improvement in the case of one dimensional search spaces. The zero-mean normalized cross correlation cost function is found to be preferable for subpixel alignment. A new error model, based on very broad assumptions is outlined in the Supplementary Material to demonstrate why these image-based interpolation formulae outperform their cost-based counterparts and why the zero-mean normalized cross correlation function is preferable for subpixel alignement.
Laurent Valentin Jospin, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun
Int. J. Comput. Vis.3
2025 Information bottleneck-guided KNN contrastive hashing for unsupervised cross-modal retrieval
abstract
Unsupervised cross-modal hashing (UCMH) has emerged as a promising solution for scalable multi-modal retrieval without costly annotations. However, existing methods often rely on rigid pairwise contrastive learning and fixed-size neighborhood selection, which suffer from false negatives and semantic noise, respectively—limiting their ability to model complex semantic structures in open-world scenarios. In this paper, we propose a novel framework, I nformation B ottleneck-guided K NN C ontrastive H ashing ( IBKCH ), which introduces a flexible and semantically adaptive contrastive paradigm for UCMH. Specifically, we design an information-aware neighbor sampling strategy that integrates: (1) a Hard-negative and Soft-positive (HN-SP) mechanism to adaptively distinguish informative negatives and softly aggregate latent positives; (2) an information bottleneck loss to retain task-relevant semantics while suppressing redundancy; and (3) an entropy sparsity regularizer to mitigate noisy neighbor interference. Furthermore, we develop an adaptive KNN contrastive learning scheme that unifies intra-modal and inter-modal alignment, enabling robust and discriminative hash code learning. Extensive experiments on three benchmark datasets demonstrate that IBKCH consistently outperforms state-of-the-art methods, especially under noisy or semantically diverse conditions—highlighting its effectiveness and generalizability in real-world UCMH applications.
Lei Zhu 0005, Zhengchang Yuan, Zeqian Yi, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Farid Boussaïd, Mohammed Bennamoun, Shichao Zhang 0001
Knowl. Based Syst.7
2025 Dual-Phase Framework for Few-Shot Hyperspectral Image Classification With Spatiospectral Masked Autoencoder and Episode Training
abstract
This article introduces a two-phase learning approach for hyperspectral image (HSI) classification using few-shot learning (FSL). For the first phase, we present a novel spatiospectral masked autoencoder (ssMAE)—an advanced self-supervised learner. For the ssMAE backbone network, we designed a transformer encoder-decoder network, where we replaced the linear layer that is used as the initial feature embedding with a 3-D convolutional layer to better extract local spectral-spatial features from 3-D visible sub-patches. By tapping into vast unlabeled data, the ssMAE learns general HSI features. In the second phase, the ssMAE encoder is fine-tuned to extract discriminative features for classification using the few-shot labeled training samples. This is achieved through a unique hybrid episode learning method that integrates the ssMAE encoder in a prototypical network (PN). We innovate with a mix of global and local prototypes (combined global-local (CGL) prototype) to refine label predictions. This technique maximizes data usage, focuses on specific samples, and mitigates issues from subpar episodes. Tested on three HSI datasets, our approach outperforms alternative few-shot methods. The code will be made publicly available athttps://github.com/Weejaa04/SSMAE.
Wijayanti Nurul Khotimah, Mohammed Bennamoun, Farid Boussaïd, Lian Xu, Ferdous Sohel
IEEE Trans. Geosci. Remote. Sens.3
2025 WSSIC-Net: Weakly-Supervised Semantic Instance Completion of 3D Point Cloud Scenes
abstract
Semantic instance completion aims to recover the complete 3D shapes of foreground objects together with their labels from a partial 2.5D scan of a scene. Previous works have relied on full supervision, which requires ground-truth annotations, in the form of bounding boxes and complete 3D objects. This has greatly limited their real-world application because the acquisition of ground-truth data is very costly and time-consuming. To address this bottleneck, we propose a Weakly-Supervised Semantic Instance Completion Network (WSSIC-Net), which learns real-world partial point cloud object completion without requiring the ground truth of complete 3D objects. Instead, WSSIC-Net leverages 3D ground-truth bounding boxes, partial objects of a raw scene, and unpaired synthetic 3D point clouds. More specifically, a 3D detector is used to encode partial point clouds into proposal features, which are then fed into two branches. The first branch uses fully supervised box prediction based on proposal features. The second branch, hereinafter called instance completion, leverages the proposal features as partial object features to achieve weakly-supervised instance completion. A Generative Adversarial Network (GAN) completes the partial features of the 2.5D foreground objects of real-world scenes using only unpaired but semantically-consistent complete synthetic point clouds. In our experiments, we demonstrate that the fully-supervised 3D detection and the weakly-supervised instance completion complement one another. The qualitative and quantitative evaluations on the ScanNet v2 dataset demonstrate that the proposed "weakly-supervised" approach consistently achieves comparable performance to the state-of-the-art "fully supervised" methods.
Zhiheng Fu, Yulan Guo, Minglin Chen, Qingyong Hu, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Image Process.6
2025 CompletionMamba: Taming State Space Model for Point Cloud Completion
abstract
Point cloud completion aims to reconstruct complete 3D shapes from partial scans. The long-range dependencies between points and shape perception are crucial for this task. While Transformers are effective due to their global processing ability, the quadratic complexity of their attention mechanism makes them unsuitable for long sequences when computational resources are constrained. As an alternative, State Space Models (SSMs) provide a memory-efficient solution for handling long-range dependencies, yet applying them directly to unordered point clouds presents challenges because of their intrinsic causality requirements. Existing methods attempt to address this by sorting points along a single axis. This, however, often overlooks complex causal relationships in 3D space since adjacency relationships based on Euclidean distance between points in the 3D space may not be preserved by this linear arrangement. To overcome this issue, we introduce CompletionMamba, a novel SSM-based network designed to harness SSMs for capturing both global and local dependencies within a point cloud. Initially, the input point cloud is causally structured by rearranging its coordinates. Then, a local SSM framework is proposed that defines neighborhood spaces around each point based on Euclidean distance, enhancing the causal structure. Although local SSM enhances relationships in short and long distance sequences, it still lacks full shape modeling of point cloud. To address this, we propose a novel shape-aware Mamba by integrating the shape code of each 3D shape into the model, enabling shape information propagation to all points. Our experiments show that CompletionMamba achieves state-of-the-art performance on both the MVP and PCN datasets.
Zhiheng Fu, Longguang Wang, Lian Xu, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Image Process.7
2025 A Guide to Image- and Video-Based Small Object Detection Using Deep Learning: Case Study of Maritime Surveillance
abstract
Detecting small objects in optical images and videos is a significant challenge in numerous intelligent transportation and autonomous systems. State-of-the-art generic object detection methods fail to accurately localize and identify such small objects (e.g., pedestrians, small vehicles, obstacles). Because small objects occupy only a small area in the input image (e.g.,$32 \times 32$pixels or less), the information extracted from such a small area is not always rich enough to support decision-making. Multidisciplinary strategies are being developed by researchers working at the interface of deep learning and computer vision to enhance the performance of Small Object Detection (SOD). In this paper, we provide a comprehensive review of over 160 research papers published between 2017 and 2022 in order to survey this growing subject. This paper summarizes the existing literature and provides a taxonomy that illustrates the broad picture of current research. We further explore methods to boost the performance of small object detection in maritime settings, where enhanced performance is crucial for ensuring safety and managing traffic. Detecting small objects in the maritime environment requires additional considerations and the current survey aims to review the advanced techniques addressing those aspects. In addition, the popular SOD datasets for generic and maritime applications are discussed, and also well-known evaluation metrics for the state-of-the-art methods on some of the datasets are provided. The link to these datasets appears inhttps://github.com/arekavandi/Datasets_SOD.
Aref Miri Rekavandi, Lian Xu, Farid Boussaïd, Abd-Krim Seghouane, Stephen Hoefs, Mohammed Bennamoun
IEEE Trans. Intell. Transp. Syst.3
2025 Box It to Bind It: Unified Layout Control and Attribute Binding in Text-to-Image Diffusion Models
abstract
While latent diffusion models (LDMs) excel at creating imaginative images, they often lack precision in semantic fidelity and spatial control over where objects are generated. To address these deficiencies, we introduce the Box-it-to-Bind-it (B2B) module—a novel, training-free approach for improving spatial control and semantic accuracy in text-to-image (T2I) diffusion models. B2B targets three key challenges in T2I: catastrophic neglect, attribute binding, and layout guidance. The process encompasses two main steps: (i)Object generation, which adjusts the latent encoding to guarantee object generation and directs it within specified bounding boxes, and (ii)Attribute binding, ensuring that generated objects adhere to their specified attributes in the prompt. B2B is designed as a compatible plug-and-play module for existing T2I models like Stable Diffusion and Gligen, markedly enhancing models’ performance in addressing these key challenges. We assess our technique on the well-established CompBench and TIFA score benchmarks, and HRS dataset where B2B not only surpasses methods specialized in either attribute binding or layout guidance but also uniquely excels by integrating these capabilities to deliver enhanced overall performance.
Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun, Aref Miri Rekavandi, Hamid Laga, Farid Boussaïd
IEEE Trans. Multim.6
2025 Auxiliary Tasks Enhanced Dual-Affinity Learning for Weakly Supervised Semantic Segmentation
abstract
Most existing weakly supervised semantic segmentation (WSSS) methods rely on class activation mapping (CAM) to extract coarse class-specific localization maps using image-level labels. Prior works have commonly used an off-line heuristic thresholding process that combines the CAM maps with off-the-shelf saliency maps produced by a general pretrained saliency model to produce more accurate pseudo-segmentation labels. We propose AuxSegNet+, a weakly supervised auxiliary learning framework to explore the rich information from these saliency maps and the significant intertask correlation between saliency detection and semantic segmentation. In the proposed AuxSegNet+, saliency detection and multilabel image classification are used as auxiliary tasks to improve the primary task of semantic segmentation with only image-level ground-truth labels. We also propose a cross-task affinity learning mechanism to learn pixel-level affinities from the saliency and segmentation feature maps. In particular, we propose a cross-task dual-affinity learning module to learn both pairwise and unary affinities, which are used to enhance the task-specific features and predictions by aggregating both query-dependent and query-independent global context for both saliency detection and semantic segmentation. The learned cross-task pairwise affinity can also be used to refine and propagate CAM maps to provide better pseudo labels for both tasks. Iterative improvement of segmentation performance is enabled by cross-task affinity learning and pseudo-label updating. Extensive experiments demonstrate the effectiveness of the proposed approach with new state-of-the-art WSSS results on the challenging PASCAL VOC and MS COCO benchmarks.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Wanli Ouyang, Ferdous Sohel, Dan Xu 0002
IEEE Trans. Neural Networks Learn. Syst.3
2024 AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
ECCV (11)7
2024 A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-Shaped Structures
Tahmina Khanam, Hamid Laga, Mohammed Bennamoun, Guanjin Wang, Ferdous Sohel, Farid Boussaïd, Anuj Srivastava
ECCV (67)6
2024 Model Predictive Control-Based Reinforcement Learning
abstract
Reinforcement Learning (RL) has garnered much attention in the field of control due to its capacity to learn from interactions and adapt to complex and dynamic environments. However, RL is challenging because it needs to balance exploration, seeking new strategies, and exploitation, leveraging known strategies for maximum gain. To address these challenges, this paper proposes a Model Predictive Control (MPC) based RL approach, where the state value function in RL is utilized as the cost function in MPC, and the system dynamic model is represented by neural networks (NNs). This eliminates the need for human intervention and addresses inaccuracies in the system model. Additionally, MPC-guided RL accelerates convergence during RL training, thereby enhancing sample efficiency. Reported results demonstrate that the proposed method outperforms traditional RL algorithms and does not require prior knowledge of the system.
Farid Boussaïd, Mohammed Bennamoun
ISCAS2
2024 MCTformer+: Multi-Class Token Transformer for Weakly Supervised Semantic Segmentation
abstract
This paper proposes a novel transformer-based framework to generate accurate class-specific object localization maps for weakly supervised semantic segmentation (WSSS). Leveraging the insight that the attended regions of the one-class token in the standard vision transformer can generate class-agnostic localization maps, we investigate the transformer's capacity to capture class-specific attention for class-discriminative object localization by learning multiple class tokens. We present the Multi-Class Token transformer, which incorporates multiple class tokens to enable class-aware interactions with patch tokens. This is facilitated by a class-aware training strategy that establishes a one-to-one correspondence between output class tokens and ground-truth class labels. We also introduce a Contrastive-Class-Token (CCT) module to enhance the learning of discriminative class tokens, enabling the model to better capture the unique characteristics of each class. Consequently, the proposed framework effectively generates class-discriminative object localization maps from the class-to-patch attentions associated with different class tokens. To refine these localization maps, we propose the utilization of patch-level pairwise affinity derived from the patch-to-patch transformer attention. Furthermore, the proposed framework seamlessly complements the Class Activation Mapping (CAM) method, yielding significant improvements in WSSS performance on PASCAL VOC 2012 and MS COCO 2014. These results underline the importance of the class token for WSSS.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Hamid Laga, Wanli Ouyang, Dan Xu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Jacobian norm with Selective Input Gradient Regularization for interpretable adversarial defense
abstract
Deep neural networks (DNNs) can be easily deceived by imperceptible alterations known as adversarial examples. These examples can lead to misclassification , posing a significant threat to the reliability of deep learning systems in real-world applications. Adversarial training (AT) is a popular technique used to enhance robustness by training models on a combination of corrupted and clean data. However, existing AT-based methods often struggle to handle transferred adversarial examples that can fool multiple defense models, thereby falling short of meeting the generalization requirements for real-world scenarios. Furthermore, AT typically fails to provide interpretable predictions, which are crucial for domain experts seeking to understand the behavior of DNNs. To overcome these challenges, we present a novel approach called Jacobian norm and Selective Input Gradient Regularization (J-SIGR). Our method leverages Jacobian normalization to improve robustness and introduces regularization of perturbation-based saliency maps, enabling interpretable predictions. By adopting J-SIGR, we achieve enhanced defense capabilities and promote high interpretability of DNNs. We evaluate the effectiveness of J-SIGR across various architectures by subjecting it to powerful adversarial attacks. Our experimental evaluations provide compelling evidence of the efficacy of J-SIGR against transferred adversarial attacks, while preserving interpretability. The project code can be found at https://github.com/Lywu-github/jJ-SIGR.git .
Deyin Liu, Lin Wu 0001, Bo Li 0090, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie, Chengwu Liang
Pattern Recognit.4
2023 Learning Multi-Modal Class-Specific Tokens for Weakly Supervised Dense Object Localization
abstract
Weakly supervised dense object localization (WSDOL) relies generally on Class Activation Mapping (CAM), which exploits the correlation between the class weights of the image classifier and the pixel-level features. Due to the limited ability to address intra-class variations, the image classifier cannot properly associate the pixel features, leading to inaccurate dense localization maps. In this paper, we propose to explicitly construct multi-modal class representations by leveraging the Contrastive Language-Image Pre-training (CLIP), to guide dense localization. More specifically, we propose a unified transformer framework to learn two-modalities of class-specific tokens, i.e., class-specific visual and textual tokens. The former captures semantics from the target visual data while the latter exploits the class-related language priors from CLIP, providing complementary information to better perceive the intra-class diversities. In addition, we propose to enrich the multi-modal class-specific tokens with sample-specific contexts comprising visual context and image-language context. This enables more adaptive class representation learning, which further facilitates dense localization. Extensive experiments show the superiority of the proposed method for WSDOL on two multi-label datasets, i.e., PASCAL VOC and MS COCO, and one single-label dataset, i.e., OpenImages. Our dense localization maps also lead to the state-of-the-art weakly supervised semantic segmentation (WSSS) results on PASCAL VOC and MS COCO.11https://github.com/xulianuwa/MMCST
Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Dan Xu 0002
CVPR4
2023 Extended Expectation Maximization for Under-Fitted Models
abstract
In this paper, we generalize the well-known Expectation Maximization (EM) algorithm using the α−divergence for Gaussian Mixture Model (GMM). This approach is used in robust subspace detection when the number of parameters is kept small to avoid overfitting and large estimation variances. The level of robustness can be tuned by the parameter α. When α → 1, our method is equivalent to the standard EM approach and for α < 1 the method is robust against potential outliers. Simulation results show that the method outperforms the standard EM when it comes to mismatches between noise models and their realizations. In addition, we use the proposed method to detect active brain areas using collected functional Magnetic Resonance Imaging (fMRI) data during task-related experiments.
Aref Miri Rekavandi, Abd-Krim Seghouane, Farid Boussaïd, Mohammed Bennamoun
ICASSP3
2023 VAPCNet: Viewpoint-Aware 3D Point Cloud Completion
abstract
Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating the viewpoint of each incomplete object is usually time-consuming and leads to huge annotation cost. In this paper, we thus propose an unsupervised viewpoint representation learning scheme for 3D point cloud completion without explicit viewpoint estimation. To be specific, we learn abstract representations of partial scans to distinguish various viewpoints in the representation space rather than the explicit estimation in the 3D space. We also introduce a Viewpoint-Aware Point cloud Completion Network (VAPCNet) with flexible adaption to various viewpoints based on the learned representations. The proposed viewpoint representation learning scheme can extract discriminative representations to obtain accurate viewpoint information. Reported experiments on two popular public datasets show that our VAPCNet achieves state-of-the-art performance for the point cloud completion task. Source code is available at https://github.com/FZH92128/VAPCNet.
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
ICCV7
2023 Reinforced Learning for Label-Efficient 3D Face Reconstruction
abstract
3D face reconstruction plays a major role in many human-robot interaction systems, from automatic face authentication to human-computer interface-based entertainment. To improve robustness against occlusions and noise, 3D face reconstruction networks are often trained on a set of in-the-wild face images preferably captured along different viewpoints of the subject. However, collecting the required large amounts of 3D annotated face data is expensive and time-consuming. To address the high annotation cost and due to the importance of training on a useful set, we propose an Active Learning (AL) framework that actively selects the most informative and representative samples to be labeled. To the best of our knowledge, this paper is the first work on tackling active learning for 3D face reconstruction to enable a label-efficient training strategy. In particular, we propose a Reinforcement Active Learning approach in conjunction with a clustering-based pooling strategy to select informative view-points of the subjects. Experimental results on 300W-LP and AFLW2000 datasets demonstrate that our proposed method is able to 1) efficiently select the most influencing view-points for labeling and outperforms several baseline AL techniques and 2) further improve the performance of a 3D Face Reconstruction network trained on the full dataset.
Hoda Mohaghegh, Hossein Rahmani 0001, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun
ICRA4
2023 Robust monocular 3D face reconstruction under challenging viewing conditions
Hoda Mohaghegh, Farid Boussaïd, Hamid Laga, Hossein Rahmani 0001, Mohammed Bennamoun
Neurocomputing2
2023 Learning class-agnostic masks with cross-task refinement for weakly supervised semantic segmentation
abstract
Abstract Weakly supervised semantic segmentation (WSSS) commonly relies on Class Activation Mapping (CAM) to produce pseudo semantic labels using image-level annotations. However, because CAM maps often form sparse object regions with poor boundaries, they cannot provide sufficient segmentation supervision. Because off-the-shelf saliency maps can provide rich object boundaries that can be leveraged to improve semantic segmentation, we propose to jointly learn semantic segmentation and class-agnostic masks by using image-level annotations and off-the-shelf saliency maps as supervision. We also propose a cross-task label refinement mechanism, which takes advantage of the learned class-agnostic masks and semantic segmentation masks, to refine the pseudo labels and provide more accurate supervision to both tasks. Moreover, we introduce a new normalization method for CAM to generate more complete class-specific localization maps. The improved CAM maps complement our learned class-agnostic masks, leading to high-quality pseudo semantic segmentation labels. Extensive experiments demonstrate the effectiveness of the proposed approach, with state-of-the-art WSSS results established on PASCAL VOC 2012 and MS COCO.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Wanli Ouyang, Dan Xu 0002
Neural Comput. Appl.3
2023 Untrained Neural Network Priors for Inverse Imaging Problems: A Survey
abstract
In recent years, advancements in machine learning (ML) techniques, in particular, deep learning (DL) methods have gained a lot of momentum in solving inverse imaging problems, often surpassing the performance provided by hand-crafted approaches. Traditionally, analytical methods have been used to solve inverse imaging problems such as image restoration, inpainting, and superresolution. Unlike analytical methods for which the problem is explicitly defined and the domain knowledge is carefully engineered into the solution, DL models do not benefit from such prior knowledge and instead make use of large datasets to predict an unknown solution to the inverse problem. Recently, a new paradigm of training deep models using a single image, named untrained neural network prior (UNNP) has been proposed to solve a variety of inverse tasks, e.g., restoration and inpainting. Since then, many researchers have proposed various applications and variants of UNNP. In this paper, we present a comprehensive review of such studies and various UNNP applications for different tasks and highlight various open research problems which require further research.
Adnan Qayyum, Inaam Ilahi, Fahad Shamshad, Farid Boussaïd, Mohammed Bennamoun, Junaid Qadir 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Multi-stage information diffusion for joint depth and surface normal estimation
Zhiheng Fu, Siyu Hong, Hamid Laga, Mohammed Bennamoun, Farid Boussaïd, Yulan Guo
Pattern Recognit.6
2023 Cross domain 2D-3D descriptor matching for unconstrained 6-DOF pose estimation
abstract
This paper presents a novel approach for cross-domain descriptor matching between 2D and 3D modalities. The 2D-3D matching is applied to localize 2D images in 3D point clouds. Direct cross-domain matching allows our technique to localize images in any type of 3D point cloud without any constraints on the nature or mechanism by which it is obtained. We propose a learning based framework, called Desc-Matcher, to directly match features between the two modalities. A dataset of 2D and 3D features with corresponding locations in images and point clouds is generated to train the Desc-Matcher. To estimate the pose of an image in any 3D cloud, keypoints and feature descriptors are extracted from the query image and the point cloud. The trained Desc-Matcher is then used to match the features from the image and the point cloud. A robust pose estimator is used to predict the location and orientation of the query image from the corresponding positions of the matched 2D and 3D features. We carried out an extensive evaluation of the proposed method for indoor and outdoor scenarios and with different types of point clouds to verify the feasibility of our approach. Experimental results show that the proposed approach can reliably estimate the 6-DOF poses of query cameras in any type of 3D point cloud with high precision. We achieved average median errors of 1.09cm/0.27∘ and 19cm/0.39∘ on the Stanford and Cambridge datasets, respectively.
Uzair Nadeem, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel, Aref Miri Rekavandi, Farid Boussaïd
Pattern Recognit.6
2023 Learning Resolution-Adaptive Representations for Cross-Resolution Person Re-Identification
abstract
Cross-resolution person re-identification (CRReID) is a challenging and practical problem that involves matching low-resolution (LR) query identity images against high-resolution (HR) gallery images. Query images often suffer from resolution degradation due to the different capturing conditions from real-world cameras. State-of-the-art solutions for CRReID either learn a resolution-invariant representation or adopt a super-resolution (SR) module to recover the missing information from the LR query. In this paper, we propose an alternative SR-free paradigm to directly compare HR and LR images via a dynamic metric that is adaptive to the resolution of a query image. We realize this idea by learning resolution-adaptive representations for cross-resolution comparison. We propose two resolution-adaptive mechanisms to achieve this. The first mechanism encodes the resolution specifics into different subvectors in the penultimate layer of the deep neural network, creating a varying-length representation. To better extract resolution-dependent information, we further propose to learn resolution-adaptive masks for intermediate residual feature blocks. A novel progressive learning strategy is proposed to train those masks properly. These two mechanisms are combined to boost the performance of CRReID. Experimental results show that the proposed method outperforms existing approaches and achieves state-of-the-art performance on multiple CRReID benchmarks.
Lin Wu 0001, Lingqiao Liu, Yang Wang 0023, Zheng Zhang 0006, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie
IEEE Trans. Image Process.5
2023 Generative Metric Learning for Adversarially Robust Open-world Person Re-Identification
abstract
The vulnerability of re-identification (re-ID) models under adversarial attacks is of significant concern as criminals may use adversarial perturbations to evade surveillance systems. Unlike a closed-world re-ID setting (i.e., a fixed number of training categories), a reliable re-ID system in the open world raises the concern of training a robust yet discriminative classifier, which still shows robustness in the context of unknown examples of an identity. In this work, we improve the robustness of open-world re-ID models by proposing a generative metric learning approach to generate adversarial examples that are regularized to produce robust distance metric. The proposed approach leverages the expressive capability of generative adversarial networks to defend the re-ID models against feature disturbance attacks. By generating the target people variants and sampling the triplet units for metric learning, our learned distance metrics are regulated to produce accurate predictions in the feature metric space. Experimental results on the three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17 demonstrate the robustness of our method.
Deyin Liu, Lin Wu 0001, Richang Hong, ZongYuan Ge, Jialie Shen 0001, Farid Boussaïd, Mohammed Bennamoun
ACM Trans. Multim. Comput. Commun. Appl.6
2022 Multi-class Token Transformer for Weakly Supervised Semantic Segmentation
abstract
This paper proposes a new transformer-based framework to learn class-specific object localization maps as pseudo labels for weakly supervised semantic segmentation (WSSS). Inspired by the fact that the attended regions of the one-class token in the standard vision transformer can be leveraged to form a class-agnostic localization map, we investigate if the transformer model can also effectively capture class-specific attention for more discriminative object localization by learning multiple class tokens within the transformer. To this end, we propose a Multi-class Token Transformer, termed as MCTformer, which uses multiple class tokens to learn interactions between the class tokens and the patch tokens. The proposed MCTformer can successfully produce class-discriminative object localization maps from the class-to-patch attentions corresponding to different class tokens. We also propose to use a patch-level pairwise affinity, which is extracted from the patch-to-patch transformer attention, to further refine the localization maps. Moreover, the proposed framework is shown to fully complement the Class Activation Mapping (CAM) method, leading to remarkably superior WSSS results on the PASCAL VOC and MS COCO datasets. These results underline the importance of the class token for WSSS.11https://github.com/xulianuwa/MCTformer
Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Dan Xu 0002
CVPR4
2022 Active-Passive SimStereo - Benchmarking the Cross-Generalization Capabilities of Deep Learning-based Stereo Methods
abstract
In stereo vision, self-similar or bland regions can make it difficult to match patches between two images. Active stereo-based methods mitigate this problem by projecting a pseudo-random pattern on the scene so that each patch of an image pair can be identified without ambiguity. However, the projected pattern significantly alters the appearance of the image. If this pattern acts as a form of adversarial noise, it could negatively impact the performance of deep learning-based methods, which are now the de-facto standard for dense stereo vision. In this paper, we propose the Active-Passive SimStereo dataset and a corresponding benchmark to evaluate the performance gap between passive and active stereo images for stereo matching algorithms. Using the proposed benchmark and an additional ablation study, we show that the feature extraction and matching modules of a selection of twenty selected deep learning-based stereo matching methods generalize to active stereo without a problem. However, the disparity refinement modules of three of the twenty architectures (ACVNet, CascadeStereo, and StereoNet) are negatively affected by the active stereo patterns due to their reliance on the appearance of the input images.
Laurent Valentin Jospin, Allen Antony, Lian Xu, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun
NeurIPS5
2022 A Survey on Deep Learning Techniques for Stereo-Based Depth Estimation
abstract
Estimating depth from RGB images is a long-standing ill-posed problem, which has been explored for decades by the computer vision, graphics, and machine learning communities. Among the existing techniques, stereo matching remains one of the most widely used in the literature due to its strong connection to the human binocular system. Traditionally, stereo-based depth estimation has been addressed through matching hand-crafted features across multiple images. Despite the extensive amount of research, these traditional techniques still suffer in the presence of highly textured areas, large uniform regions, and occlusions. Motivated by their growing success in solving various 2D and 3D vision problems, deep learning for stereo-based depth estimation has attracted a growing interest from the community, with more than 150 papers published in this area between 2014 and 2019. This new generation of methods has demonstrated a significant leap in performance, enabling applications such as autonomous driving and augmented reality. In this paper, we provide a comprehensive survey of this new and continuously growing field of research, summarize the most commonly used pipelines, and discuss their benefits and limitations. In retrospect of what has been achieved so far, we also conjecture what the future may hold for deep learning-based stereo for depth estimation research.
Hamid Laga, Laurent Valentin Jospin, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Soft Exemplar Highlighting for Cross-View Image-Based Geo-Localization
abstract
The goal of ground-to-aerial image geo-localization is to determine the location of a ground query image by matching it against a reference database consisting of aerial/satellite images. This task is highly challenging due to the large appearance difference caused by extreme changes in viewpoint and orientation. In this work, we show that the training difficulty is an important cue that can be leveraged to improve metric learning on cross-view images. More specifically, we propose a new Soft Exemplar Highlighting (SEH) loss to achieve online soft selection of exemplars. Adaptive weights are generated for exemplars by measuring their associated training difficulty using distance rectified logistic regression. These weights are then constrained to remove simple exemplars from training and truncate the large weights of extremely hard exemplars to escape from the trap with a local optimal solution. We further use the proposed SEH loss to train two mainstream convolutional neural networks for ground-to-aerial image-based geo-localization. Experimental results on two benchmark cross-view image datasets demonstrate that the proposed method achieves significant improvements in feature discriminativeness and outperforms the state-of-the-art image-based geo-localization methods.
Yulan Guo, Kunhong Li 0001, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Image Process.4
2022 Pseudo-Pair Based Self-Similarity Learning for Unsupervised Person Re-Identification
abstract
Person re-identification (re-ID) is of great importance to video surveillance systems by estimating the similarity between a pair of cross-camera person shorts. Current methods for estimating such similarity require a large number of labeled samples for supervised training. In this paper, we present a pseudo-pair based self-similarity learning approach for unsupervised person re-ID without human annotations. Unlike conventional unsupervised re-ID methods that use pseudo labels based on global clustering, we construct patch surrogate classes as initial supervision, and propose to assign pseudo labels to images through the pairwise gradient-guided similarity separation. This can cluster images in pseudo pairs, and the pseudos can be updated during training. Based on pseudo pairs, we propose to improve the generalization of similarity function via a novel self-similarity learning:it learns local discriminative features from individual images via intra-similarity, and discovers the patch correspondence across images via inter-similarity. The intra-similarity learning is based on channel attention to detect diverse local features from an image. The inter-similarity learning employs a deformable convolution with a non-local block to align patches for cross-image similarity. Experimental results on several re-ID benchmark datasets demonstrate the superiority of the proposed method over the state-of-the-arts.
Lin Wu 0001, Deyin Liu, Dapeng Chen, ZongYuan Ge, Farid Boussaïd, Mohammed Bennamoun, Jialie Shen 0001
IEEE Trans. Image Process.6
2021 Leveraging Auxiliary Tasks with Affinity Learning for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation is a challenging task in the absence of densely labelled data. Only relying on class activation maps (CAM) with image-level labels provides deficient segmentation supervision. Prior works thus consider pre-trained models to produce coarse saliency maps to guide the generation of pseudo segmentation labels. However, the commonly used off-line heuristic generation process cannot fully exploit the benefits of these coarse saliency maps. Motivated by the significant inter-task correlation, we propose a novel weakly supervised multi-task framework termed as AuxSegNet, to leverage saliency detection and multi-label image classification as auxiliary tasks to improve the primary task of semantic segmentation using only image-level ground-truth labels. Inspired by their similar structured semantics, we also propose to learn a cross-task global pixellevel affinity map from the saliency and segmentation representations. The learned cross-task affinity can be used to refine saliency predictions and propagate CAM maps to provide improved pseudo labels for both tasks. The mutual boost between pseudo label updating and cross-task affinity learning enables iterative improvements on segmentation performance. Extensive experiments demonstrate the effectiveness of the proposed auxiliary learning network structure and the cross-task affinity learning method. The proposed approach achieves state-of-the-art weakly supervised segmentation performance on the challenging PASCAL VOC 2012 and MS COCO benchmarks.1
Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel, Dan Xu 0002
ICCV4
2021 Atrous convolutional feature network for weakly supervised semantic segmentation
Lian Xu, Hao Xue 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel
Neurocomputing4
2021 Quantitative performance evaluation of object detectors in hazy environments
Cameron Hodges, Mohammed Bennamoun, Farid Boussaïd
Pattern Recognit. Lett.3
2020 ResFeats: Residual network based features for underwater image classification
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
Image Vis. Comput.5
2020 Learning Latent Global Network for Skeleton-Based Action Prediction
abstract
Human actions represented with 3D skeleton sequences are robust to clustered backgrounds and illumination changes. In this paper, we investigate skeleton-based action prediction, which aims to recognize an action from a partial skeleton sequence that contains incomplete action information. We propose a new Latent Global Network based on adversarial learning for action prediction. We demonstrate that the proposed network provides latent long-term global information that is complementary to the local action information of the partial sequences and helps improve action prediction. We show that action prediction can be improved by combining the latent global information with the local action information. We test the proposed method on three challenging skeleton datasets and report state-of-the-art performance.
Qiuhong Ke, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd
IEEE Trans. Image Process.6
2019 An Improved Approach to Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with image-level labels is of great significance since it alleviates the dependency on dense annotations. However, it is a challenging task as it aims to achieve a mapping from high-level semantics to low-level features. In this work, we propose a three-step method to bridge this gap. First, we rely on the interpretable ability of deep neural networks to generate attention maps with class localization information by back-propagating gradients. Secondly, we employ an off-the-shelf object saliency detector with an iterative erasing strategy to obtain saliency maps with spatial extent information of objects. Finally, we combine these two complementary maps to generate pseudo ground-truth images for the training of the segmentation network. With the help of the pre-trained model on the MS-COCO dataset and a multi-scale fusion method, we obtained mIoU of 62.1% and 63.3% on PASCAL VOC 2012 val and test sets, respectively, achieving new state-of-the-art results for the weakly supervised semantic segmentation task.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel
ICASSP3
2019 Coral Classification Using DenseNet and Cross-modality Transfer Learning
abstract
Coral classification is a challenging task due to the complex morphology and ambiguous boundaries of corals. This paper investigates the benefits of Densely connected convolutional network (DenseNet) and multi-modal image translation techniques in boosting image classification performance by synthesizing missing fluorescence information. To this end, an imageconditional Generative Adversarial Network (GAN) based image translator is trained to model the relationship between reflectance and fluorescence images. Through this image translator, fluorescence images can be generated from the available reflectance images to provide complementary information. During the classification phase, reflectance and translated fluorescence images are combined to obtain more discriminative representations and produce improved classification performance. We present results on the EFC and MLC datasets and report state-of-the-art coral classification performance.
Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel
IJCNN3
2018 Global Regularizer and Temporal-Aware Cross-Entropy for Skeleton-Based Early Action Recognition
Qiuhong Ke, Jun Liu 0036, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd
ACCV (4)7
2018 Classification of Corals in Reflectance and Fluorescence Images Using Convolutional Neural Network Representations
abstract
Coral species, with complex morphology and ambiguous boundaries, pose a great challenge for automated classification. CNN activations, which are extracted from fully connected layers of deep networks (FC features), have been successfully used as powerful universal representations in many visual tasks. In this paper, we investigate the transferability and combined performance of FC features and CONY features (extracted from convolutional layers) in the coral classification of two image modalities (reflectance and fluorescence), using a typical deep network (e.g. VGGNet). We exploit vector of locally aggregated descriptors (VLAD) encoding and principal component analysis (PCA) to compress dense CONY features into a compact representation. Experimental results demonstrate that encoded CONV3 features achieve superior performances on reflectance and fluorescence coral images, compared to FC features. The combination of these two features further improves the overall accuracy and achieves state-of-the-art performance on the challenging EFC dataset.
Lian Xu, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
ICASSP5
2018 Room-Temperature Dual-mode CMOS Gas-FET Sensor for Diabetes Detection
abstract
A CMOS gas-sensitive field-effect transistor (Gas-FET) is proposed for noninvasive diabetes detection. The Gas-FET was fabricated in the GlobalFoundries 0.18μm 1P6M process with a lateral control gate and a floating gate to set operating point and gas sensing sensitivity, respectively. ZnO nanorods were used as the sensing material and deposited on top of the chip using a hydrothermal process at 80°C. Room-temperature acetone sensing down to sub-ppm level is demonstrated to enable noninvasive diagnosis of diabetes in exhaled breath. A dual-mode integrated readout circuit is also proposed to improve the sensor gas discrimination ability through the acquisition of 2-dimensional information by every single Gas-FET sensor.
Farid Boussaïd, Amine Bermak, Chi-Ying Tsui
ISCAS2
2018 Exploiting layerwise convexity of rectifier networks with sign constrained weights
Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel
Neural Networks2
2018 Learning Clip Representations for Skeleton-Based 3D Action Recognition
abstract
This paper presents a new representation of skeleton sequences for 3D action recognition. Existing methods based on hand-crafted features or recurrent neural networks cannot adequately capture the complex spatial structures and the long-term temporal dynamics of the skeleton sequences, which are very important to recognize the actions. In this paper, we propose to transform each channel of the 3D coordinates of a skeleton sequence into a clip. Each frame of the generated clip represents the temporal information of the entire skeleton sequence and one particular spatial relationship between the skeleton joints. The entire clip incorporates multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We also propose a multitask convolutional neural network (MTCNN) to learn the generated clips for action recognition. The proposed MTCNN processes all the frames of the generated clips in parallel to explore the spatial and temporal information of the skeleton sequences. The proposed method has been extensively tested on six challenging benchmark datasets. Experimental results consistently demonstrate the superiority of the proposed clip representation and the feature learning method for 3D action recognition compared to the existing techniques.
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
IEEE Trans. Image Process.5
2018 Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction Prediction
abstract
Predicting an interaction before it is fully executed is very important in applications, such as human-robot interaction and video surveillance. In a two-human interaction scenario, there are often contextual dependency structures between the global interaction context of the two humans and the local context of the different body parts of each human. In this paper, we propose to learn the structure of the interaction contexts and combine it with the spatial and temporal information of a video sequence to better predict the interaction class. The structural models, including the spatial and the temporal models, are learned with long short term memory (LSTM) networks to capture the dependency of the global and local contexts of each RGB frame and each optical flow image, respectively. LSTM networks are also capable of detecting the key information from global and local interaction contexts. Moreover, to effectively combine the structural models with the spatial and temporal models for interaction prediction, a ranking score fusion method is introduced to automatically compute the optimal weight of each model for score fusion. Experimental results on the BIT-Interaction Dataset and the UT-Interaction Dataset clearly demonstrate the benefits of the proposed method.
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
IEEE Trans. Multim.5
2017 A New Representation of Skeleton Sequences for 3D Action Recognition
abstract
This paper presents a new method for 3D action recognition with skeleton sequences (i.e., 3D trajectories of human skeleton joints). The proposed method first transforms each skeleton sequence into three clips each consisting of several frames for spatial temporal feature learning using deep neural networks. Each clip is generated from one channel of the cylindrical coordinates of the skeleton sequence. Each frame of the generated clips represents the temporal information of the entire skeleton sequence, and incorporates one particular spatial relationship between the joints. The entire clips include multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We propose to use deep convolutional neural networks to learn long-term temporal information of the skeleton sequence from the frames of the generated clips, and then use a Multi-Task Learning Network (MTLN) to jointly process all frames of the clips in parallel to incorporate spatial structural information for action recognition. Experimental results clearly show the effectiveness of the proposed new representation and feature learning method for 3D action recognition.
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd
CVPR5
2017 An ultra low-power capacitively-coupled chopper instrumentation amplifier for wheatstone-bridge readout circuits
abstract
This paper presents an ultra-low-power low-noise Capacitively-coupled Chopper Instrumentation Amplifier (CCIA). A current-reuse telescopic topology in the first stage along with a recycling folded-cascode topology in the second stage consumes net bias current of 26μA with enhanced efficiency and achieves an input-referred noise power-spectral-density of 12.77nV/√Hz. The proposed CCIA is chopped at 50kHz to bring the input-referred offset around 6μV and flicker-noise corner around 400mHz. Implemented in chartered 0.18μm CMOS process and designed for Thermoresistive Micro Calorimetric Flow (TMCF) sensors, the reported work achieves an excellent Noise Efficiency Factor (NEF) of 2.5 which is the lowest ever reported NEF for such applications.
Moaaz Ahmed, Farid Boussaïd, Amine Bermak
ISCAS2
2017 Dual transduction Gas sensor based on a surface acoustic wave resonator
abstract
This paper presents a novel dual transduction gas sensor providing both resistance and mass modalities for single sensor gas identification. The proposed sensor relies on a configurable dual mode frequency/resistance readout circuit, which enables the use of a single conventional surface acoustic wave (SAW) device. Unlike prior works which all rely on custom-made gas sensors, the proposed sensor is based on off-the-shelf SAW devices, making it low cost and easy to implement. Reported results validate the functionality of the proposed dual transduction gas sensor. The introduction of control switches for the dual mode readout is shown to only deteriorate the phase noise performance of the SAW oscillator by 4 dBc/Hz at 10 MHz offset and not affect the low offset part. This demonstrates that the mass sensing resolution of the SAW device is not reduced while including the resistive sensing feature.
Amine Bermak, Chi-Ying Tsui, Farid Boussaïd
ISCAS4
2017 Keypoints-based surface representation for 3D modeling and 3D object recognition
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd
Pattern Recognit.3
2017 SkeletonNet: Mining Deep Part Features for 3-D Action Recognition
abstract
This letter presents SkeletonNet, a deep learning framework for skeleton-based 3-D action recognition. Given a skeleton sequence, the spatial structure of the skeleton joints in each frame and the temporal information between multiple frames are two important factors for action recognition. We first extract body-part-based features from each frame of the skeleton sequence. Compared to the original coordinates of the skeleton joints, the proposed features are translation, rotation, and scale invariant. To learn robust temporal information, instead of treating the features of all frames as a time series, we transform the features into images and feed them to the proposed deep learning network, which contains two parts: one to extract general features from the input images, while the other to generate a discriminative and compact representation for action recognition. The proposed method is tested on the SBU kinect interaction dataset, the CMU dataset, and the large-scale NTU RGB+D dataset and achieves state-of-the-art performance.
Qiuhong Ke, Senjian An, Mohammed Bennamoun, Ferdous Sohel, Farid Boussaïd
IEEE Signal Process. Lett.5
2016 Coral classification with hybrid feature representations
abstract
Coral reefs exhibit significant within-class variations, complex between-class boundaries and inconsistent image clarity. This makes coral classification a challenging task. In this paper, we report the application of generic CNN representations combined with hand-crafted features for coral reef classification to take advantage of the complementary strengths of these representation types. We extract CNN based features from patches centred at labelled pixels at multiple scales. We use texture and color based hand-crafted features extracted from the same patches to complement the CNN features. Our proposed method achieves a classification accuracy that is higher than the state-of-art methods on the MLC benchmark dataset for corals.
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd, Renae Hovey, Gary A. Kendrick, Robert B. Fisher
ICIP5
2016 A hierarchical ZnO nanostructure gas sensor for human breath-level acetone detection
abstract
Analyzing the concentration of acetone in human breath constitutes a promising non-invasive means to diagnose the onset of diabetes, with acetone levels of at least 1.8ppm typically associated to individuals suffering from diabetes. In this paper, we report the performance of a hierarchical ZnO nanostructure gas sensor for acetone detection. The fabricated gas sensor can detect concentrations as low as 1ppm while operating at a comparatively lower temperature of 200°C. In addition, the proposed gas sensor can be fabricated on a silicon wafer using a MEMS process, making it thereby possible to fully integrate gas sensing and electronic circuitry on a single silicon chip.
Xiaofang Pan, Farid Boussaïd, Amine Bermak, Zhiyong Fan
ISCAS3
2016 A semantic RBM-based model for image set classification
Said Elaiwat, Mohammed Bennamoun, Farid Boussaïd
Neurocomputing3
2016 Iterative deep learning for image set based face and object recognition
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd
Neurocomputing3
2016 A novel feature representation for automatic 3D object recognition in cluttered scenes
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd
Neurocomputing3
2016 A spatio-temporal RBM-based model for facial expression recognition
Said Elaiwat, Mohammed Bennamoun, Farid Boussaïd
Pattern Recognit.3
2015 Contractive Rectifier Networks for Nonlinear Maximum Margin Classification
abstract
To find the optimal nonlinear separating boundary with maximum margin in the input data space, this paper proposes Contractive Rectifier Networks (CRNs), wherein the hidden-layer transformations are restricted to be contraction mappings. The contractive constraints ensure that the achieved separating margin in the input space is larger than or equal to the separating margin in the output layer. The training of the proposed CRNs is formulated as a linear support vector machine (SVM) in the output layer, combined with two or more contractive hidden layers. Effective algorithms have been proposed to address the optimization challenges arising from contraction constraints. Experimental results on MNIST, CIFAR-10, CIFAR-100 and MIT-67 datasets demonstrate that the proposed contractive rectifier networks consistently outperform their conventional unconstrained rectifier network counterparts.
Senjian An, Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel
ICCV5
2015 How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances?
abstract
This paper investigates how hidden layers of deep rectifier networks are capable of transforming two or more pattern sets to be linearly separable while preserving the distances with a guaranteed degree, and proves the universal classification power of such distance preserving rectifier networks. Through the nearly isometric nonlinear transformation in the hidden layers, the margin of the linear separating plane in the output layer and the margin of the nonlinear separating boundary in the original data space can be closely related so that the maximum margin classification in the input data space can be achieved approximately via the maximum margin linear classifiers in the output layer. The generalization performance of such distance preserving deep rectifier neural networks can be well justified by the distance-preserving properties of their hidden layers and the maximum margin property of the linear classifiers in the output layer.
Senjian An, Farid Boussaïd, Mohammed Bennamoun
ICML2
2015 Sign Constrained Rectifier Networks with Applications to Pattern Decompositions
Senjian An, Qiuhong Ke, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel
ECML/PKDD (1)4
2015 A Curvelet-based approach for textured 3D face recognition
Said Elaiwat, Mohammed Bennamoun, Farid Boussaïd, Amar A. El-Sallam
Pattern Recognit.3
2015 A novel 3D vorticity based approach for automatic registration of low resolution range images
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd
Pattern Recognit.3
2015 Quantitative Error Analysis of Bilateral Filtering
abstract
One of the fastest acceleration techniques for bilateral image filtering is the real time O(1) quantization method proposed by Yang 2009, which first computes some Principal Bilateral Filtered Image Components (PBFICs) and then applies linear interpolation to estimate the filtered output images. There is a trade-off between accuracy and efficiency in selecting the number of PBFICs: the more PBFICs are used, the higher the accuracy, and the higher the computational cost. A question arises: how many PBFICs are required to achieve a certain level of accuracy? In this letter, we address this question by investigating the properties of bilateral filtering and deriving the linear interpolation error bounds when only a subset of PBFICs is used. The provided theoretical analysis indicates that the necessary number of PBFICs for user-provided precision depends on the range kernel and, for typical Gaussian range kernels, a small percentage (typically less than 4%) of the PBFICs are enough for good approximations.
Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel
IEEE Signal Process. Lett.2
2014 A high speed configurable FPGA architecture for bilateral filtering
abstract
This paper presents a high speed configurable FPGA architecture for bilateral filtering. The proposed architecture is highly pipelined, parallel and fully configurable. It can achieve an operating frequency of 450 MHz and a throughput of one pixel value per clock cycle. This is almost three times faster than any reported FPGA architecture with such a throughput. Line Buffering was implemented using a novel BRAM architecture that ensures access to all pixels of the filter window in a single clock cycle. The proposed BRAM architecture also addresses the high speed and throughput requirements of convolution functions in image processing algorithms.
Jithin Sankar Sankaran Kutty, Farid Boussaïd, Abbes Amira
ICIP2
2014 32 Bit ×32 Bit Multiprecision Razor-Based Dynamic Voltage Scaling Multiplier With Operands Scheduler
abstract
In this paper, we present a multiprecision (MP) reconfigurable multiplier that incorporates variable precision, parallel processing (PP), razor-based dynamic voltage scaling (DVS), and dedicated MP operands scheduling to provide optimum performance for a variety of operating conditions. All of the building blocks of the proposed reconfigurable multiplier can either work as independent smaller-precision multipliers or work in parallel to perform higher-precision multiplications. Given the user's requirements (e.g., throughput), a dynamic voltage/frequency scaling management unit configures the multiplier to operate at the proper precision and frequency. Adapting to the run-time workload of the targeted application, razor flip-flops together with a dithering voltage unit then configure the multiplier to achieve the lowest power consumption. The single-switch dithering voltage unit and razor flip-flops help to reduce the voltage safety margins and overhead typically associated to DVS to the lowest level. The large silicon area and power overhead typically associated to reconfigurability features are removed. Finally, the proposed novel MP multiplier can further benefit from an operands scheduler that rearranges the input data, hence to determine the optimum voltage and frequency operating conditions for minimum power consumption. This low-power MP multiplier is fabricated in AMIS 0.35- μm technology. Experimental results show that the proposed MP design features a 28.2% and 15.8% reduction in circuit area and power consumption compared with conventional fixed-width multiplier. When combining this MP design with error-tolerant razor-based DVS, PP, and the proposed novel operands scheduler, 77.7%-86.3% total power reduction is achieved with a total silicon area overhead as low as 11.1%. This paper successfully demonstrates that a MP architecture can allow more aggressive frequency/supply voltage scaling for improved power efficiency.
Farid Boussaïd, Amine Bermak
IEEE Trans. Very Large Scale Integr. Syst.2
2013 3D-Div: A novel local surface descriptor for feature matching and pairwise range image registration
abstract
This paper presents a novel local surface descriptor, called 3D-Div. The proposed descriptor is based on the concept of 3D vector fields divergence, extensively used in electromagnetic theory. To generate a 3D-Div descriptor of a 3D surface, a keypoint is first extracted on the 3D surface, then a local patch of a certain size is selected around that keypoint. A Local Reference Frame (LRF) is then constructed at the keypoint using all points forming the patch. A normalized 3D vector field is then computed at each point in the patch and referenced with LRF vectors. The 3D-Div descriptors are finally generated as the divergence of the reoriented 3D vector field. We tested our proposed descriptor on the low resolution Washington RGB-D (Kinect) object dataset. Performance was evaluated for the tasks of feature matching and pairwise range image registration. Experimental results showed that the proposed 3D-Div is 88% more computationally efficient and 47% more accurate than commonly used Spin Image (SI) descriptors.
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd, Amar A. El-Sallam
ICIP3
2013 A high speed configurable FPGA architecture for k-mean clustering
abstract
This paper presents a high speed configurable FPGA architecture for k-means clustering. The proposed architecture is highly pipelined, parallel and fully configurable. It can achieve an operating frequency of 400 MHz, which is at least three times faster than prior works. The proposed architecture addresses the high speed and throughput requirements of machine vision, multi-media and data mining applications.
Jithin Sankar Sankaran Kutty, Farid Boussaïd, Abbes Amira
ISCAS2
2012 Bio-inspired gas recognition based on the organization of the olfactory pathway
abstract
Existing gas recognition techniques rely on complex signal processing techniques. This paper presents a simple bio-inspired gas recognition technique, exploiting fundamental characteristics of the organization of the olfactory pathway. The technique was validated using an-in house custom-fabricated tin gas sensor array together with three target gases: ethanol, methane, and carbon monoxide. Experimental results show that the proposed approach provides high accuracy, enabling the concept of a fully integrated electronic nose.
Jaber Hassan J. Al Yamani, Farid Boussaïd, Amine Bermak, Dominique Martinez
ISCAS2
2011 A low cost CMOS polarimetric ophthalmoscope scheme for cerebral malaria diagnostics
abstract
In this paper, we present a low cost CMOS polarimetric ophthalmoscope scheme enabling the capture of the retinal abnormalities that are unique to cerebral malaria. The proposed technology, which can be integrated into cellphones, offers the basis for quick and non-invasive screening of cerebral malaria. In addition, we report a high quality micropolarizer array for the proposed polarimetric ophthalmoscope, exploiting “guest-host” interactions in liquid crystals. With dichroic azodye-1 (AD1) molecules as the “guest” and nematic liquid crystal (NLC) molecules as the “host”, we demonstrate a better control of the molecular orientation of the “guest”, which in turn results in a ~25% increase of the major principal transmittance and a 139% increase of the peak extinction ratio. The proposed micropolarizer fabrication technology is simple and cost-effective, requiring only selective photo-patterning of a “guest-host” polymer spincoated over the image sensor.
Xiaojin Zhao, Amine Bermak, Farid Boussaïd
VLSI-SoC3
2010 A frequency-based signature gas identification circuit for SnO2 gas sensors
abstract
This paper presents a gas identification circuit for tin oxide (SnO2) gas sensors. The proposed circuit uses 2 gas sensors with different characteristics to achieve gas identification. A spike train is generated during operation, with the frequency of spike occurrence being gas dependent but concentration invariant. As a result, the spike firing frequency can be used to achieve gas identification. The calibration of this readout technique requires only a single exposure to the target gases to extract the sensor resistances. The low complexity processing is suitable for on-chip implementation. The functionality of this circuit has been validated with real data from our in-house fabricated sensors.
Kwan Ting Ng, Farid Boussaïd, Amine Bermak
ISCAS2
2010 Dynamic voltage and frequency scaling for low-power multi-precision reconfigurable multiplier
abstract
In this paper, a 32×32-bit low power multi-precision multiplier is described, in which each building block can be either an independent smaller-precision multiplier or work in parallel to perform higher-precision operations. The proposed multi-precision multiplier enables voltage and frequency scaling for low power operation, while still maintaining full throughput. According to user's arbitrary throughput requirements, the highly dynamic voltage and frequency scaling circuits can autonomously configure the multiplier to operate with the lowest possible voltage and frequency to achieve the lowest power consumption. By carrying out optimizations at the algorithmic and architectural levels, we have completely removed silicon area and power overheads which is always associated with the reconfigurability features. The 32×32-bit low power multi-precision multiplier has been implemented in TSMC 0.18 μm technology. Compared with fixed-width multipliers, the proposed design features around 13.8% and 30% reduction in circuit area and power, respectively. Multi-precision processing featured in this paper accordingly enables voltage and frequency scaling resulting in up to 68% reduction in power consumption.
Amine Bermak, Farid Boussaïd
ISCAS3
2010 Liquid-crystal micropolarimeter array for visible linear and circular polarization imaging
abstract
In this paper, we propose a liquid-crystal mi-cropolarimeter (LCMP) array with high spatial resolution for real-time linear and circular polarization imaging in visible spectrum. LCMPs for extracting 0°, 90° linearly and right-handed circularly polarized components of incident light are implemented by micro-patterning a liquid crystal (LC) layer on top of a 45° oriented ultra-thin metal-wire-grid polarizer (MWGP). A compact LCMP pitch of 5μm × 5μm is achieved with sulfonic-dye-1 (SD1) as the LC alignment material. In addition, these micron-scale LCMPs feature ~5μm overall thickness and ~1100 extinction ratio. Reported experimental results validate the concept of real-time linear and circular polarization image sensing and processing with targets illuminated by collimated artificial light.
Xiaojin Zhao, Amine Bermak, Farid Boussaïd, Vladimir G. Chigrinov
ISCAS3
2009 A Robust Spike-based Gas Identification Technique for SnO2 Gas Sensors
abstract
This paper presents a robust gas identification technique for tin oxide (SnO2) gas sensors. The proposed technique generates a unique spike pattern or signature for each sensed gas, irrespective of its concentration. The proposed gas identification technique is insensitive to drift in the sensor baseline resistance. Furthermore, its calibration requires a single measurement to be made for each targeted gas. The proposed spike-based gas identification technique has been implemented in TSMC 0.18 mum CMOS technology and validated using experimental data from a fabricated in-house 4 times 4 SnO2gas sensor array. Reported results reveal a 10% increase in correct gas detection rate.
Kwan Ting Ng, Hung Tat Chen, Farid Boussaïd, Amine Bermak, Dominique Martinez
ISCAS3