Naoufel Werghi

dblp:36/5115 · DBLP profile ↗
← Back
128ranked-venue papers
19as first author
74since 2021 · last 2026
0000-0002-5542-448XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 10 first-author · 26 since 2021Artificial intelligence and machine learning · 43 · 9 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 13 · 4 first-author · 9 since 2021Systems, architecture and hardware · 6 · 5 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021
YearPublicationVenuePosition
2026 X-ThreatDet: Enhancing X-Ray Threat Detection with Self-Supervised and Multi-Modal Learning
Yonathan Michael, Mohamad Alansari, Maregu Assefa, Naoufel Werghi, Andreas Henschel
MMM (4)4
2026 PRISM-X: Progressive semi-supervised threat detection in X-ray scans with self-guided multimodal refinement
Abdelfatah Hassan Ahmed, Mohammad Irshaid, Mohamad Alansari, Divya Velayudhan, Mohammed Tarnini, Mohammed El-Amine Azz, Naser A. Abou-Elheggag, Taimur Hassan, Ernesto Damiani, Naoufel Werghi
Inf. Process. Manag.10
2026 Emotion and noise-robust speaker identification via filter-free self-supervised learning
abstract
• The proposed SWF tokenization method, dynamically capturing emotional subtleties, noise robustness, and contextual information to robustly preserve speaker identity. • A transformative deep-self-supervised spectrogram transformer back-end, outperforming conventional approaches by effectively addressing their inherent limitations in preserving local features without the need for extensive labeled training data. • Complete elimination of dependency on additional speech enhancement, enabling seamless, efficient, and robust end-to-end learning tailored for real-world deployments. Identifying speakers in noisy and emotional conditions remains a significant challenge due to the distortion of spectral cues. This study proposes the Speech Without Filter (SWF) framework, a novel self-supervised learning paradigm that operates directly on raw spectrograms. Theoretically, this research introduces a progressive tokenization mechanism that acts as a structural inductive bias, mimicking the contracting path of a U-Net to preserve local spectro-temporal continuity. Unlike standard fixed-patch Transformers that often smooth over speaker-specific micro-textures, the SWF architecture integrates denoising and feature extraction into a single stage, challenging the traditional decoupled paradigm of speech enhancement and recognition. Using a sample of 1.58 million pre-training instances, the model was evaluated across English (RAVDESS), Arabic (ESD), and stressful (SUSAS) datasets. Results demonstrate significant improvements, with the SWF model achieving 91.01% accuracy in clean conditions and maintaining 88.5% in high-noise cocktail party environments, outperforming state-of-the-art models like WavLM and HuBERT. These findings suggest that architectural innovation in tokenization is as critical as pre-training scale for robust speech processing.
Shibani Hamsa, Youssef Iraqi, Ismail Shahin, Ernesto Damiani, Kinda Khalaf, Herbert F. Jelinek, Naoufel Werghi
Inf. Process. Manag.7
2026 X-SSL: Self-supervised X-ray threat detection with zero-shot and multi-modal learning
abstract
Automated X-ray threat detection is challenged by cluttered baggage scans, severe object occlusions, and the scarcity of annotated datasets. Traditional supervised approaches are impractical, as they require large amounts of labeled data, which is difficult to obtain given the rarity of threat objects. Unsupervised methods, on the other hand, often fail to differentiate between threat and nonthreat items due to the complex grayscale nature and high object overlap inherent in X-ray imagery. To overcome these limitations, we propose X-SSL a novel self-supervised learning framework that eliminates manual annotations designed to perform threat localization. Our approach integrates spatial region extraction using MaskCut for zero-shot object proposal generation, contrastive multi-modal clustering that leverages both image and text encoders to cluster and label proposals into threat and nonthreat categories, and self-supervised knowledge distillation where a teacher–student model refines multiscale features from global and local image crops for improved representation learning. We evaluated X-SSL on two benchmark datasets: PIDray (39,000+ images) and CLCXray (14,000+ images), demonstrating significant improvements over previous state-of-the-art (SOTA) methods. In the hidden PIDray subset, X-SSL improves detection AP to 18.16 (+4.41 AP) and segmentation AP to 13.76 (+1.01 AP) over previous methods. On CLCXray, it achieves 39.26 detection AP (+5.93 AP) and 38.67 segmentation AP (+10.20 AP), significantly surpassing previous approaches. For classification, X-SSL achieves an accuracy of 43% on PIDray Hidden and 65% on CLCXray, further highlighting its superior performance compared to existing weakly supervised and unsupervised methods. Code will be available here: https://github.com/yonathan-kiflom/X-SSL .
Yonathan Michael, Mohamad Alansari, Abdelfatah Hassan Ahmed, Naoufel Werghi, Andreas Henschel
Inf. Process. Manag.4
2026 Data poisoning-based backdoor attacks against supervised learning rules of Spiking Neural Networks
Lingxin Jin, Wei Jiang 0016, Jinyu Zhan, Meiyu Lin, Letian Chen, Boran Quan, Lin Zuo, Xingzhi Zhou 0001, Maregu Assefa, Naoufel Werghi
J. Syst. Archit.10
2026 TriGAN-SiaMT: A triple-segmentor adversarial network with bounding box priors for semi-supervised brain lesion segmentation
Mohammad Alshurbaji, Maregu Assefa, Ahmad Obeid 0001, Mohamed L. Seghier, Taimur Hassan, Kamal Taha, Naoufel Werghi
Pattern Recognit. Lett.7
2026 DUCore: Dual Uncertainty-Guided Consistency and Regional Contrastive Learning for Semi-Supervised Medical Image Segmentation
abstract
Uncertainty-aware consistency learning is one of the reliable approaches in semi-supervised medical image segmentation, enforcing robust model predictions under various perturbations. However, existing methods often rely on multiple stochastic predictions or dual-network/decoder discrepancies to estimate uncertainty, which increases computational cost and discards uncertain regions, potentially missing complex structures such as ambiguous lesion boundaries. To address these challenges, we introduce a Dual Uncertainty-Guided Consistency and Regional Contrastive Learning (DUCore) framework. DUCore improves segmentation robustness by integrating two complementary loss functions within consistency learning. The dual uncertainty-guided consistency loss (DuCL) adaptively calibrates the prediction alignment by prioritizing uncertain regions. DuCL uses deterministic single-pass uncertainty estimation, employing entropy-based calibration for aleatoric uncertainty and Proxy Dirichlet calibration for epistemic uncertainty. These uncertainty measures are computed directly from network output, and moderately uncertain regions are weighted instead of being discarded, which preserves valuable learning signals. The Regional Contrastive Loss (ReCL) further refines feature separability using boundary- and gradient-based hard negative mining in the encoded representation space. By explicitly targeting structural ambiguities, ReCL distinguishes lesion and organ edges from visually similar boundary-adjacent regions and mitigates intensity overlaps in gradient-rich transitions. As a result, DUCore is able to delineate fine structures and complex boundaries with higher precision. Extensive experiments on various medical segmentation benchmarks reveal that DUCore outperforms existing consistency methods.
Maregu Assefa, Muzammal Naseer, Kumie Gedamu, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Mohamed L. Seghier, Ernesto Damiani, Naoufel Werghi
IEEE J. Biomed. Health Informatics8
2025 Structured Comprehensive Textual Representations for Medical Vision-Language Pretraining
abstract
Medical vision-language learning faces a persistent challenge: limited paired image-text data, especially in expert domains like radiology. To address this, we propose a data-efficient framework that enhances dual-encoder vision-language models through structured comprehensive textual representations. Using GPT-4 and curated clinical knowledge, we generate semantically comprehensive representations for each disease class. These representations capture detailed visual descriptions of the disease, the major causes, and the major symptoms related to chest X-rays. These representations are paired with medical images to pre-train BiomedCLIP using contrastive learning, effectively aligning vision and language modalities without the need for large-scale manual annotations. We then fine-tune the vision encoder for chest X-ray disease classification. Our results show that this approach achieves competitive performance, even with limited supervision. This work highlights the potential of structured generative supervision to scale vision-language learning in data-constrained medical domains.
Youssef Ibrahim, Anabia Sohail, Sajid Javed, Hasan Al Marzouqi, Naoufel Werghi
AICCSA5
2025 Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
abstract
In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tissue phenotyping, and survival analysis due to the limited level of detail that a single-resolution image can provide. Addressing this, we propose a novel multi-resolution paradigm leveraging Whole Slide Images (WSIs) to extract histology patches at multiple resolutions and generate corresponding textual descriptions through advanced CPath VLM. We introduce visual-textual alignment at multiple resolutions as well as cross-resolution alignment to establish more effective text-guided visual representations. Cross-resolution alignment using a multi-modal encoder enhances the model’s ability to capture context from multiple resolutions in histology images. Our model aims to capture a broader range of information, supported by novel loss functions, enriches feature representation, improves discriminative ability, and enhances generalization across different resolutions. Pre-trained on a comprehensive TCGA dataset with 34 million image-language pairs at various resolutions, our fine-tuned model outperforms State-Of-The-Art (SOTA) counterparts across multiple datasets and tasks, demonstrating its effectiveness in CPath. The code is available on GitHub at: https://github.com/BasitAlawode/MR-PLIP.
Shahad Albastaki, Anabia Sohail, Iyyakutti Iyappan Ganapathi, Basit Alawode, Asim Khan, Sajid Javed, Naoufel Werghi, Mohammed Bennamoun, Arif Mahmood
CVPR7
2025 DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
abstract
Semi-supervised learning in medical image segmentation leverages unlabeled data to reduce annotation burdens through consistency learning. However, current methods struggle with class imbalance and high uncertainty from pathology variations, leading to inaccurate segmentation in 3D medical images. To address these challenges, we present DyCON, a Dynamic Uncertainty-aware Consistency and Contrastive Learning framework that enhances the generalization of consistency methods with two complementary losses: Uncertainty-aware Consistency Loss (UnCL) and Focal Entropy-aware Contrastive Loss (FeCL). UnCL enforces global consistency by dynamically weighting the contribution of each voxel to the consistency loss based on its uncertainty, preserving high-uncertainty regions instead of filtering them out. Initially, UnCL prioritizes learning from uncertain voxels with lower penalties, encouraging the model to explore challenging regions. As training progress, the penalty shift towards confident voxels to refine predictions and ensure global consistency. Meanwhile, FeCL enhances local feature discrimination in imbalanced regions by introducing dual focal mechanisms and adaptive confidence adjustments into the contrastive principle. These mechanisms jointly prioritizes hard positives and negatives while focusing on uncertain sample pairs, effectively capturing subtle lesion variations under class imbalance. Extensive evaluations on four diverse medical image segmentation datasets (ISLES’22, BraTS’19, LA, Pancreas) show DyCON’s superior performance against SOTA methods1.
Maregu Assefa, Muzammal Naseer, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Mohamed L. Seghier, Naoufel Werghi
CVPR6
2025 STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
abstract
Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a closed-set paradigm with predefined labels. To address these challenges, we introduce STCray, the first multimodal X-ray baggage security dataset, comprising 46,642 image-caption paired scans across 21 threat categories, generated using an X-ray scanner for airport security. STCray is meticulously developed with our specialized protocol that ensures domain-aware, coherent captions, that lead to the multi-modal instruction following data in X-ray baggage security. This allows us to train a domain-aware visual AI assistant named STING-BEE that supports a range of vision-language tasks, including scene comprehension, referring threat localization, visual grounding, and visual question answering (VQA), establishing novel baselines for multi-modal learning in X-ray baggage security. Further, STING-BEE shows state-of-the-art generalization in cross-domain settings. Code, data, and models are available at https://divs1159.github.io/STING-BEE/.
Divya Velayudhan, Abdelfatah Hassan Ahmed, Mohamad Alansari, Neha Gour, Abderaouf Behouch, Taimur Hassan, Syed Talal Wasim, Nabil Maalej, Muzammal Naseer, Juergen Gall, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi
CVPR13
2025 Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
abstract
This paper presents a summary of the 2025 Sclera Segmentation Benchmarking Competition (SSBC), which focused on the development of privacy-preserving sclera-segmentation models trained using synthetically generated ocular images. The goal of the competition was to evaluate how well models trained on synthetic data perform in comparison to those trained on real-world datasets. The competition featured two tracks: (i) one relying solely on synthetic data for model development, and (ii) one combining/mixing synthetic with (a limited amount of) real-world data. A total of nine research groups submitted diverse segmentation models, employing a variety of architectural designs, including transformer-based solutions, lightweight models, and segmentation networks guided by generative frameworks. Experiments were conducted across three evaluation datasets containing both synthetic and real-world images, collected under diverse conditions. Results show that models trained entirely on synthetic data can achieve competitive performance, particularly when dedicated training strategies are employed, as evidenced by the top performing models that achieved F1scores of over 0.8 in the synthetic data track. Moreover, performance gains in the mixed track were often driven more by methodological choices rather than by the inclusion of real data, highlighting the promise of synthetic data for privacy-aware biometric development. The code and data for the competition is available at: https://github.com/dariant/SSBC_2025.
Matej Vitek, Darian Tomasevic, Abhijit Das 0001, Sabari Nathan, Gökhan Özbulak, G. A. T. Özbulak, Jean-Paul Calbimonte, André Anjos, Hariohm Hemant Bhatt, Dhruv Dhirendra Premani, Jay Chaudhari, Caiyong Wang, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Divya Velayudan, Maregu Assefa, Naoufel Werghi, Zachary A. Daniels, Leeon John, Ritesh Vyas, Jalil Nourmohammadi Khiarak, Taher Akbari Saeed, Mahsa Nasehi, Ali Kianfar, Mobina Pashazadeh Panahi, Geetanjali Sharma, Pushp Raj Panth, Ramachandra Raghavendra, Aditya Nigam, Umapada Pal 0001, Peter Peer, Vitomir Struc
IJCB20
2025 Vision-Language Neural Graph Featurization for Extracting Retinal Lesions
Taimur Hassan, Anabia Sohail, Muzammal Naseer, Naoufel Werghi
ICCV4
2025 Enhancing Medical Vision-Language Models with Rich Textual Descriptions and Multiple Alignments for Chest X-Ray Diagnosis
abstract
Vision-Language models (VLMs) integrate natural language understanding with visual data interpretation, crucial in diverse applications such as medical imaging. However, training VLMs on limited data, especially in radiology, remains a challenge. We propose a strategy to improve dual encoder performance under data constraints. Using contrastive learning to align visual and textual embeddings effectively, we generated a bag of rich textual descriptions using GPT-4 to augment merged information from esteemed medical resources and pre-trained BiomedCLIP. These rich textual descriptions provide in-depth information on disease visual description, major causes, and major symptoms, enhancing the model’s contextual understanding and classification accuracy. Unlike previous methods relying on a single alignment, our multiple alignment strategy associates multiple images with multiple textual descriptions per disease class while capping descriptors to maintain computational efficiency. Adapting the vision encoder for chest X-ray classification, our approach achieves competitive accuracy with fewer training pairs, highlighting its potential for data-limited domains.
Youssef Ibrahim, Anabia Sohail, Sajid Javed, Hasan Almarzouqi, Mohamed Deriche 0001, Naoufel Werghi
ICIP6
2025 PMIL: A Topology Module to Improve MIL-based WSI Classification
abstract
Deep learning models have achieved remarkable success in pathology image analysis. However, they still face challenges in effectively modeling fine-grained, object-level features. Topological Data Analysis (TDA) has shown promise for addressing these issues but remains underexplored, particularly for whole-slide pathology applications. Additionally, the effectiveness of TDA has yet to be firmly established, as current studies largely use small-scale datasets. In this work, we address these gaps by introducing Persistent Homology in Multiple Instance Learning (PMIL), the first adaptable TDA-based module within the MIL framework. We validate our approach on a large-scale classification dataset, benchmarking against multiple state-of-the-art methods.
Ahmad Obeid 0001, Anabia Sohail, Said Boumaraf, Xiabi Liu, Sajid Javed, Hasan Almarzouqi, Jorge Dias 0001, Mohammed Bennamoun, Naoufel Werghi, Ibrahim M. Elfadel
ISCAS9
2025 A statistical 3D watermarking method based on logistic chaotic scrambling system
abstract
Digital watermarking is recognized as an efficient solution in the field of information security. In 3D mesh watermarking, vertex positions are often subtly adjusted to maintain a predefined relationship with their neighboring vertices. In this paper, we propose a robust blind 3D mesh watermarking method based on mesh saliency and a scrambling technique. Before embedding, the watermark image is scrambled using the logistic chaos scrambling method. The proposed algorithm segments the mesh surface into distinct regions based on detected salient points, after which the watermark is statistically embedded within each region. Experimental results demonstrate that the proposed method effectively resists various attacks such as additive binary noise, Laplacian smoothing, simplification, and cropping while maintaining high imperceptibility.
Nassima Medimegh, Hela Haj Mohamed, Samir Belaid, Naoufel Werghi
KES4
2025 A Brain-Inspired Dual-Stream Neural Network for Tumor Classification in Ultrasound Images
abstract
Early and accurate tumor classification in ultrasound images plays a pivotal role in improving cancer diagnosis and patient outcomes. Existing computer-aided diagnostic (CAD) algorithms often rely on cropping-based single feedforward pathways, which can result in the loss of crucial contextual information around the tumor. The surrounding ultrasound data, including relative intensity, plays a significant role in tumor diagnosis, and incorrect cropping or positioning may lead to unreliable results. To overcome these limitations, we propose a novel Brain-inspired Dual-stream Network (BidsNet), aiming to emulate the functional mechanisms of the dorsal and ventral streams in human visual processing. BidsNet processes the entire ultrasound image as input, preventing errors or loss of contextual details from cropping. The dorsal stream in BidsNet specializes in extracting spatial features, such as shape and texture, while the ventral stream focuses on object recognition and classification. A cross-stream communication mechanism is introduced to facilitate dynamic information sharing between the streams: spatial attention generated in the dorsal stream informs the ventral stream to improve feature localization, while channel attention derived from the ventral stream refines spatial feature representation in the dorsal stream. This collaborative interplay boosts both the interpretability and performance of the network. Extensive experiments on multiple ultrasound datasets demonstrate that BidsNet delivers superior accuracy and interpretability, validating the effectiveness of its dual-stream design and cross-stream communication mechanism.
Chaochao Lin, Said Boumaraf, Xiabi Liu, Qianglin Liu, Lijuan Niu, Naoufel Werghi
SMC6
2025 EfficientFaceV2S: A lightweight model and a benchmarking approach for drone-captured face recognition
Mohamad Alansari, Khaled Alnuaimi, Iyyakutti Iyappan Ganapathi, Sara Alansari, Sajid Javed, Abdulhadi Shoufan, Yahya Zweiri, Naoufel Werghi
Expert Syst. Appl.8
2025 Vision-based air-flow monitoring in an industrial flare system design using deep convolutional neural networks
Said Boumaraf, Muaz Al Radi, Fares Oussama Abdelhafez, Khalid Yousef Al Awadhi, Hamad Karki, Sahraoui Dhelim, Naoufel Werghi
Expert Syst. Appl.8
2025 Video anomaly detection in 10 years: a survey and outlook
Moshira Abdalla, Sajid Javed, Muaz Al Radi, Anwaar Ulhaq, Naoufel Werghi
Neural Comput. Appl.5
2025 A Vision Language Correlation Framework for Screening Disabled Retina
abstract
Retinopathy is a group of retinal disabilities that causes severe visual impairments or complete blindness. Due to the capability of optical coherence tomography to reveal early retinal abnormalities, many researchers have utilized it to develop autonomous retinal screening systems. However, to the best of our knowledge, most of these systems rely only on mathematical features, which might not be helpful to clinicians since they do not encompass the clinical manifestations of screening the underlying diseases. Such clinical manifestations are critically important to be considered within the autonomous screening systems to match the grading of ophthalmologists within the clinical settings. To overcome these limitations, we present a novel framework that exploits the fusion of vision language correlation between the retinal imagery and the set of clinical prompts to recognize the different types of retinal disabilities. The proposed framework is rigorously tested on six public datasets, where, across each dataset, the proposed framework outperformed state-of-the-art methods in various metrics. Moreover, the clinical significance of the proposed framework is also tested under strict blind testing experiments, where the proposed system achieved a statistically significant correlation coefficient of 0.9185 and 0.9529 with the two expert clinicians. These blind test experiments highlight the potential of the proposed framework to be deployed in the real world for accurate screening of retinal diseases.
Taimur Hassan, Hina Raja, Kais Belwafi, Samet Akcay, Mohamed Jleli, Bessem Samet, Naoufel Werghi, Jawad Yousaf, Mohammed Ghazal
IEEE J. Biomed. Health Informatics7
2025 RobMOT: 3D Multi-Object Tracking Enhancement Through Observational Noise and State Estimation Drift Mitigation in LiDAR Point Clouds
abstract
This paper addresses key limitations in recent 3D tracking-by-detection methods, focusing on the challenges of identifying legitimate trajectories and mitigating state estimation drift in the Kalman filter. Current methods rely heavily on threshold-based detection score filtering approaches to reduce false positives and prevent ghost trajectories. However, these approaches fail for distant and partially occluded objects, where detection scores drop, and false positives surpass that threshold. Additionally, many existing methods assume that detections provide precise localization, overlooking the inherent noise that affects localization accuracy and causes state drift for occluded objects, as demonstrated in this work. To this end, a novel track validity mechanism, combined with a multi-stage observational gating process, is proposed that significantly reduces ghost tracks and improves tracking performance. Our method achieves 29.47% enhancement in Multi-Object tracking accuracy (MOTA) on the KITTI validation dataset with the Second detector. Furthermore, a refined Kalman filter term mitigates localization noise, ensuring robust state estimation for objects that are occluded and superior recovery during prolonged occlusions. This results in higher-order tracking accuracy (HOTA) improving by 4.8% on the KITTI validation dataset with the PV-RCNN detector. The proposed online framework, RobMOT, outperforms state-of-the-art methods, including deep learning approaches, across multiple detectors, with HOTA improvements of up to 3.92% on the KITTI testing dataset and 8.7% on the KITTI validation dataset while achieving the lowest identity switch (IDSW) scores of 7 and 0, respectively. RobMOT excels under challenging scenarios, such as tracking distant objects and handling prolonged occlusions, surpassing state-of-the-art methods on the Waymo Open testing dataset with a 1.77% improvement in MOTA for objects at distances exceeding 50 meters. RobMOT achieves a groundbreaking runtime of 3221 FPS using a single CPU, establishing itself as a highly efficient and scalable solution for real-time multi-object tracking.
Mohamed Nagy, Naoufel Werghi, Bilal Hassan, Jorge Dias 0001, Majid Khonji
IEEE Trans. Intell. Transp. Syst.2
2025 Unsupervised Dual Transformer Learning for 3-D Textured Surface Segmentation
abstract
Analysis of the 3-D texture is indispensable for various tasks, such as retrieval, segmentation, classification, and inspection of sculptures, knit fabrics, and biological tissues. A 3-D texture represents a locally repeated surface variation (SV) that is independent of the overall shape of the surface and can be determined using the local neighborhood and its characteristics. Existing methods mostly employ computer vision techniques that analyze a 3-D mesh globally, derive features, and then utilize them for classification or retrieval tasks. While several traditional and learning-based methods have been proposed in the literature, only a few have addressed 3-D texture analysis, and none have considered unsupervised schemes so far. This article proposes an original framework for the unsupervised segmentation of 3-D texture on the mesh manifold. The problem is approached as a binary surface segmentation task, where the mesh surface is partitioned into textured and nontextured regions without prior annotation. The proposed method comprises a mutual transformer-based system consisting of a label generator (LG) and a label cleaner (LC). Both models take geometric image representations of the surface mesh facets and label them as texture or nontexture using an iterative mutual learning scheme. Extensive experiments on three publicly available datasets with diverse texture patterns demonstrate that the proposed framework outperforms standard and state-of-the-art unsupervised techniques and performs reasonably well compared to supervised methods.
Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Sajid Javed, Syed Sadaf Ali, Naoufel Werghi
IEEE Trans. Neural Networks Learn. Syst.5
2025 Continuous Wavelet Network for Efficient and Transferable Collision Detection in Collaborative Robots
abstract
This article addresses the crucial aspect of safety in collaborative robotics by introducing a new continuous wavelet transform-convolutional neural network (CWT-CNN) for efficient robot collision detection. Unlike conventional methods, CWT-CNN exhibits superior data efficiency, requiring minimal collision data for robust training without relying on a dynamic model. The network’s adaptability extends to varying internal stiffness levels, offering robustness to changes in robotic system characteristics. Through comprehensive experimental studies, we investigate the impact of input signal types, wavelet types, wavelet scale ranges, and time-moving window sizes on collision detection performance, offering critical insights for optimal CWT parameter selection. Additionally, our transferability analysis demonstrates that the CWT-CNN can seamlessly adapt from one joint to another, requiring only minimal free-motion data from the new joint. This adaptability is validated through extensive experiments on an industrial robot and the robot equipped with variable stiffness actuators. In conclusion, the CWT-CNN is highly generalizable and data-efficient, making it a reliable solution for real-time collision detection in human-robot interactions, addressing a key aspect of safety in collaborative environments.
Zhenwei Niu, Taimur Hassan, Mohamed Nassim Boushaki, Naoufel Werghi, Irfan Hussain
IEEE Trans. Syst. Man Cybern. Syst.4
2024 3D-TexSeg: Unsupervised Segmentation of 3D Texture Using Mutual Transformer Learning
abstract
Analysis of the 3D Texture is indispensable for various tasks, such as retrieval, segmentation, classification, and inspection of sculptures, knitted fabrics, and biological tissues. A 3D texture is a locally repeated surface variation independent of the surface’s overall shape and can be determined using the local neighborhood and its characteristics. Existing techniques typically employ computer vision techniques that analyze a 3D mesh globally, derive features, and then utilize the obtained features for retrieval or classification. Several traditional and learning-based methods exist in the literature; however, only a few are on 3D texture, and nothing yet, to the best of our knowledge, on the unsupervised schemes. This paper presents an original framework for the unsupervised segmentation of the 3D texture on the mesh manifold. We approach this problem as binary surface segmentation, partitioning the mesh surface into textured and non-textured regions without prior annotation. We devise a mutual transformer-based system comprising a label generator and a cleaner. The two models take geometric image representations of the surface mesh facets and label them as texture or non-texture across an iterative mutual learning scheme. Extensive experiments on three publicly available datasets with diverse texture patterns demonstrate that the proposed framework outperforms standard and SOTA unsupervised techniques and competes reasonably with supervised methods.
Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Sajid Javed, Syed Sadaf Ali, Naoufel Werghi
3DV5
2024 CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
abstract
This paper proposes Comprehensive Pathology Language Image Pretraining (CPLIP), a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision-language models by leveraging extensive data without needing ground truth annotations. CPLIP involves constructing a pathology-specific dictionary, generating textual descriptions for images using language models, and retrieving relevant images for each text snippet via a pretrained model. The model is then fine-tuned using a many-to-many contrastive learning method to align complex interrelated concepts across both modalities. Evaluated across multiple histopathology tasks, CPLIP shows notable improvements in zero-shot learning scenarios, outperforming existing methods in both interpretability and robustness and setting a higher benchmark for the application of vision-language models in the field. To encourage further research and replication, the code for CPLIP is available on GitHub at https://cplip.github.io/
Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, Mohammed Bennamoun
CVPR5
2024 Human Action Recognition with Multi-Level Granularity and Pair-Wise Hyper GCN
abstract
Lately, there has been a surge in interest in utilizing Graph Convolutional Networks (GCNs) for the purpose of action recognition using skeletal data. In order to achieve optimal results, it is crucial to generate high-quality representations of the skeletal graph. Graph Convolutional Networks (GCNs) often employ the Message-Passing Mechanism (MPM) to acquire knowledge about various components of the skeleton by iteratively computing new features at each step. However, the interconnections between joints in the skeletal structure are intricate and extend beyond mere proximity. In order to address this issue, we propose the implementation of our Disassembled Hyper-Graph (DH-Graph), which draws inspiration from hyper-graph edges. The process of constructing the DH-network entails a few steps: partitioning the skeleton network into clusters of hyper-edges according to their semantic significance and relevance to action recognition, arranging these clusters in a hierarchical structure to enhance granularity, and establishing connections between joints within these clusters to discover hidden relationships. The DH-Graph employs a spatial domain GCN technique to construct the Pair-wise Hyper Hierarchical GCN (PH-GCN). In addition, we incorporate the HyperAttention module, which employs Multi-scale Representative Spatial Average Pooling and Edge Convolution techniques to emphasize significant sets of hyper-hierarchical information. Extensive experiments demonstrate that PH-GCN achieves remarkable performance on challenging NTU RGB+D and Northwestern UCLA datasets.
Tamam Alsarhan, Syed Sadaf Ali, Ayoub Alsarhan, Iyyakutti Iyappan Ganapathi, Naoufel Werghi
FG5
2024 CLIFS: Clip-Driven Few-Shot Learning for Baggage Threat Classification
abstract
Baggage screening in airports is a cornerstone in airport security measures. The advent of computer vision technologies in recent years has led to the development of several automated systems for identifying security threats in baggage scans. However, existing methods struggle to adapt to new threat categories when faced with a scarcity of data samples, and the rapid emergence of new threats. Hence, in this paper, we propose a novel CLIP-driven few-shot framework (CLIFS) to explore the potential of multi-modality using text-image fusion through contrastive learning to learn relevant contextual features for recognizing security threats with limited samples. By integrating features from GPT-4 generated captions with image features, CLIFS leverages both visual and textual data to significantly improve threat classification performance with limited samples in a few-shot learning context. Our proposed CLIFS was rigorously tested on the SIXray public available baggage X-ray dataset, where it outperformed state-of-the-art by 31.3% in accuracy and 28.40% in F1-score for the challenging 5-shots scenario, demonstrating its robustness and effectiveness in classifying threats from limited data samples.
Abdelfatah Hassan Ahmed, Divya Velayudhan, Mahmoud Elmezain, Muaz Al Radi, Abderrahmene Boudiaf, Taimur Hassan, Mohamed Deriche 0001, Mohammed Bennamoun, Naoufel Werghi
ICIP9
2024 Integrating Vision-Language Supervision for Uniform Appearance Tracking
abstract
Integrating detailed Natural Language (NL) descriptions with modern tracking technologies represents a significant and emerging field within Uniform Appearance (UA) crowd-tracking research, demonstrating substantial potential for future developments. A prominent challenge in this area is the lack of NL descriptions tailored for UA crowd tracking datasets. Existing datasets for Drone-Person Tracking in Uniform Appearance Crowd (D-PTUAC) lack essential textual annotations. Our study aims to bridge this gap by innovatively introducing comprehensive natural language descriptions for the D-PTUAC dataset, specifically designed for Uniform Appearance crowd tracking using drones. This enhancement aims to provide a richer understanding of the dataset and facilitate more effective utilization in research and applications related to drone-based crowd tracking. These descriptions are meticulously designed to include extensive information about the target entities, thereby significantly augmenting the dataset’s depth and applicability. Our evaluations utilizing the latest state-of-the-art (SOTA) NL-based tracking algorithms showed us a remarkable competitive performance in tracking when juxtaposed against SOTA visual trackers benchmarked on the D-PTUAC dataset. This outcome highlights the critical role and efficacy of integrated language descriptions in enhancing the methodologies employed in UA crowd tracking.
Mohamad Alansari, Ahmed Abughali, Obadah Habash, Khaled Alnuaimi, Sajid Javed, Naoufel Werghi
ICIP6
2024 SMO-CLIP: Enhancing Anomalous Smoke Density Assessment Using A Hybrid LLM-VLM Approach
abstract
Flare stacks are among the crucial components in the safety and emission control of petrochemical plants. However, due to the imperceptibility of smoke and contaminants, analyzing these released particles during flare stack operation is one of the top challenges. To stress the problem, our work presents a novel solution called SMO-CLIP that can hybridize knowledge from Vision-Language Models (VLMs), specifically the Contrastive Language Image Pretraining (CLIP) model, with extra insights derived from GPT-4 Large Language Model (LLM). Furthermore, two new tasks, Finegrained Smoke Density Recognition (FSDR) and Coarsegrained Smoke Density Recognition (CSDR) are investigated in this paper to accurately detect and evaluate varying smoke intensities. Notable advancements over current approaches are observed through extensive experiments, demonstrating the superior performance of the proposed approach against state-of-the-art models.
Muaz Al Radi, Mahmoud Said Elmezain, Abdelfatah Hassan Ahmed, Abderrahmene Boudiaf, Said Boumaraf, Jorge Dias 0001, Hamad Karki, Sajid Javed, Khalid Yousef Al Awadhi, Naoufel Werghi
ICIP11
2024 Recurrent 3-D Multi-Level Visual Transformer For Joint Classification of Heterogeneous 2-d AND 3-D Radiographic Data
abstract
Recent advancements in artificial intelligence algorithms for medical imaging show significant potential in automating the detection of lung infections from chest radiograph scans. However, current approaches often focus solely on either 2-D or 3-D scans, failing to leverage the combined advantages of both modalities. Moreover, conventional slice-based methods place a manual burden on radiologists for slice selection. To overcome these challenges, we propose the Recurrent 3-D Multi-level Vision Transformer (R3DM-ViT) model, capable of handling multimodal data to enhance diagnostic accuracy. Our quantitative evaluations demonstrate that R3DM-ViT surpasses existing methods, achieving an impressive accuracy of $96.67 \%$, F1-score of $96.88 \%$, mean average precision of $96.75 \%$, and mean average recall of $97.02 \%$. This research signifies a significant stride forward in the automated detection of lung infections through multimodal imaging.
Muhammad Owais, Taimur Hassan, Divya Velayudhan, Irfan Hussain, Naoufel Werghi
ICIP6
2024 A Fusion-Based Approach for Blind Contrast-Enhanced Image Ranking
abstract
Cameras are now available at extremely low prices due to ongoing advancements in image acquisition hardware. However, the quality of images can be compromised by various distortions that occur throughout the entire process, from acquisition to processing and delivery. Over the past few decades, researchers have primarily focused on developing algorithms to assess the quality of distorted images. Unfortunately, certain distortions can also result from enhancement processes, such as over-enhancement and color saturation. Although there are metrics available for measuring contrast levels in images, there is currently no standard metric for evaluating the extent and effects of contrast enhancement. In this paper, we propose a new framework that expands the evaluation of contrast levels to ranking contrast-enhanced images. Our technique involves extracting a new set of features that accurately describe the effects of contrast enhancement. Furthermore, we integrate additional statistical indicators, such as skewness and kurtosis, which describe the degree of visual satisfaction linked to human perception. These identified characteristics are subsequently use with a simple classification module to determine the rank order for a given collection of contrast enhanced images. The results show excellent accuracy in correct ranking which outperforms state-of-the-art by more than $15 \%$.
Wael Suliman, Mohamed Deriche 0001, Naoufel Werghi, Azeddine Beghdadi
ICIP3
2024 An interpretable dual attention network for diabetic retinopathy grading: IDANet
Amit Bhati, Neha Gour, Pritee Khanna, Aparajita Ojha, Naoufel Werghi
Artif. Intell. Medicine5
2024 FinTem: A secure and non-invertible technique for fingerprint template protection
Amber Hayat, Syed Sadaf Ali, Ashok Kumar Bhateja, Naoufel Werghi
Comput. Secur.4
2024 Enhancing security in X-ray baggage scans: A contour-driven learning approach for abnormality classification and instance segmentation
Abdelfatah Hassan Ahmed, Divya Velayudhan, Taimur Hassan, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi
Eng. Appl. Artif. Intell.6
2024 B3D-EAR: Binarized 3D descriptors for ear-based human recognition
Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Surya Prakash 0001, Sambit Bakshi, Naoufel Werghi
Expert Syst. Appl.5
2024 Efficient and lightweight in-memory computing architecture for hardware security
Hala Ajmi, Fakhreddine Zayer, Amira Hadj Fredj, Belgacem Hamdi, Baker Mohammad, Naoufel Werghi, Jorge Dias 0001
J. Parallel Distributed Comput.6
2024 Unsupervised mutual transformer learning for multi-gigapixel Whole Slide Image classification
Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi, Nasir M. Rajpoot
Medical Image Anal.4
2024 Programmable broad learning system for baggage threat recognition
Muhammad Shafay, Abdelfatah Hassan Ahmed, Taimur Hassan, Jorge Dias 0001, Naoufel Werghi
Multim. Tools Appl.5
2024 ViT-LSTM synergy: a multi-feature approach for speaker identification and mask detection
Ali Bou Nassif, Ismail Shahin, Mohamed Bader, Abdelfatah Hassan Ahmed, Naoufel Werghi
Neural Comput. Appl.5
2024 Incremental convolutional transformer for baggage threat detection
Taimur Hassan, Bilal Hassan, Muhammad Owais, Divya Velayudhan, Jorge Dias 0001, Mohammed Ghazal, Naoufel Werghi
Pattern Recognit.7
2024 Autonomous Localization of X-Ray Baggage Threats via Weakly Supervised Learning
abstract
Autonomous X-ray baggage security screening has shown significant strides recently, proving itself a viable solution to the flaws in manual screening, thanks to advancements in deep learning. However, these data-hungry techniques feed on extensively annotated data involving strenuous labor, impeding their advances in baggage screening. Consequently, we present a context-aware transformer for weakly supervised localization to relieve the annotation burden and provide visual interpretability that aids screeners in threat recognition and researchers in identifying the pitfalls of existing systems. The proposed approach can generalize and localize different types of contraband with only cost-effective binary labels without explicit training on item detection. Context extraction block, integrated into the dual-token framework, generates threat-aware context maps, while the token scoring block focuses on minimizing partial activations. Experimental results surpass state of the art (SOTA) methods in terms of classification and localization accuracies. Furthermore, we analyze failures to determine current vulnerabilities and provide new insights for future research.
Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Neha Gour, Muhammad Owais, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi
IEEE Trans. Ind. Informatics8
2024 Neural Graph Refinement for Robust Recognition of Nuclei Communities in Histopathological Landscape
abstract
Accurate classification of nuclei communities is an important step towards timely treating the cancer spread. Graph theory provides an elegant way to represent and analyze nuclei communities within the histopathological landscape in order to perform tissue phenotyping and tumor profiling tasks. Many researchers have worked on recognizing nuclei regions within the histology images in order to grade cancerous progression. However, due to the high structural similarities between nuclei communities, defining a model that can accurately differentiate between nuclei pathological patterns still needs to be solved. To surmount this challenge, we present a novel approach, dubbed neural graph refinement, that enhances the capabilities of existing models to perform nuclei recognition tasks by employing graph representational learning and broadcasting processes. Based on the physical interaction of the nuclei, we first construct a fully connected graph in which nodes represent nuclei and adjacent nodes are connected to each other via an undirected edge. For each edge and node pair, appearance and geometric features are computed and are then utilized for generating the neural graph embeddings. These embeddings are used for diffusing contextual information to the neighboring nodes, all along a path traversing the whole graph to infer global information over an entire nuclei network and predict pathologically meaningful communities. Through rigorous evaluation of the proposed scheme across four public datasets, we showcase that learning such communities through neural graph refinement produces better results that outperform state-of-the-art methods.
Taimur Hassan, Zhu Li 0001, Sajid Javed, Jorge Dias 0001, Naoufel Werghi
IEEE Trans. Image Process.5
2024 Center-Focused Affinity Loss for Class Imbalance Histology Image Classification
abstract
Early-stage cancer diagnosis potentially improves the chances of survival for many cancer patients worldwide. Manual examination of Whole Slide Images (WSIs) is a time-consuming task for analyzing tumor-microenvironment. To overcome this limitation, the conjunction of deep learning with computational pathology has been proposed to assist pathologists in efficiently prognosing the cancerous spread. Nevertheless, the existing deep learning methods are ill-equipped to handle fine-grained histopathology datasets. This is because these models are constrained via conventional softmax loss function, which cannot expose them to learn distinct representational embeddings of the similarly textured WSIs containing an imbalanced data distribution. To address this problem, we propose a novel center-focused affinity loss (CFAL) function that exhibits 1) constructing uniformly distributed class prototypes in the feature space, 2) penalizing difficult samples, 3) minimizing intra-class variations, and 4) placing greater emphasis on learning minority class features. We evaluated the performance of the proposed CFAL loss function on two publicly available breast and colon cancer datasets having varying levels of imbalanced classes. The proposed CFAL function shows better discrimination abilities as compared to the popular loss functions such as ArcFace, CosFace, and Focal loss. Moreover, it outperforms several SOTA methods for histology image classification across both datasets.
Taslim Mahbub, Ahmad Obeid 0001, Sajid Javed, Jorge Dias 0001, Taimur Hassan, Naoufel Werghi
IEEE J. Biomed. Health Informatics6
2023 The Unconstrained Ear Recognition Challenge 2023: Maximizing Performance and Minimizing Bias
abstract
The paper provides a summary of the 2023 Unconstrained Ear Recognition Challenge (UERC), a benchmarking effort focused on ear recognition from images acquired in uncontrolled environments. The objective of the challenge was to evaluate the effectiveness of current ear recognition techniques on a challenging ear dataset while analyzing the techniques from two distinct aspects, i.e., verification performance and bias with respect to specific demographic factors, i.e., gender and ethnicity. Seven research groups participated in the challenge and submitted a seven distinct recognition approaches that ranged from descriptor-based methods and deep-learning models to ensemble techniques that relied on multiple data representations to maximize performance and minimize bias. A comprehensive investigation into the performance of the submitted models is presented, as well as an in-depth analysis of bias and associated performance differentials due to differences in gender and ethnicity. The results of the challenge suggest that a wide variety of models (e.g., transformers, convolutional neural networks, ensemble models) is capable of achieving competitive recognition results, but also that all of the models still exhibit considerable performance differentials with respect to both gender and ethnicity. To promote further development of unbiased and effective ear recognition models, the starter kit of UERC 2023 together with the baseline model, and training and test data is made available from: http://ears.fri.uni-lj.si/
Ziga Emersic, Tetsushi Ohki, Muku Akasaka, Takahiko Arakawa, Soshi Maeda, Masora Okano, Yuya Sato, Anjith George, Sébastien Marcel, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Sajid Javed, Naoufel Werghi, S. G. Isik, Erdi Saritas, Hazim Kemal Ekenel, V. Hudovernik, Jan Niklas Kolf, Fadi Boutros, Naser Damer, G. Sharma, Aman Kamboj, Aditya Nigam, Deepak Kumar Jain 0001, G. Cámara-Chávez, Peter Peer, Vitomir Struc
IJCB13
2023 EFaR 2023: Efficient Face Recognition Competition
abstract
This paper presents the summary of the Efficient Face Recognition Competition (EFaR) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition received 17 submissions from 6 different teams. To drive further development of efficient face recognition models, the submitted solutions are ranked based on a weighted score of the achieved verification accuracies on a diverse set of benchmarks, as well as the deployability given by the number of floating-point operations and model size. The evaluation of submissions is extended to bias, cross-quality, and large-scale recognition benchmarks. Overall, the paper gives an overview of the achieved performance values of the submitted solutions as well as a diverse set of baselines. The submitted solutions use small, efficient network architectures to reduce the computational cost, some solutions apply model quantization. An outlook on possible techniques that are underrepresented in current solutions is given as well.
Jan Niklas Kolf, Fadi Boutros, Jurek Elliesen, Markus Theuerkauf, Naser Damer, Mohamad Alansari, Oussama Abdul Hay, Sara Alansari, Sajid Javed, Naoufel Werghi, Klemen Grm, Vitomir Struc, Fernando Alonso-Fernandez, Kevin Hernandez-Diaz, Josef Bigün, Anjith George, Christophe Ecabert, Hatef Otroshi-Shahreza, Ketan Kotwal, Sébastien Marcel, Iurii Medvedev, Bo Jin 0018, Diogo Nunes, Ahmad Hassanpour, Pankaj Khatiwada, Aafan Ahmad Toor, Bian Yang
IJCB10
2023 A Shallow U-Net with Split-Fused Attention Mechanism for Retinal Vessel Segmentation
abstract
Extraction of retinal vascular parts is an important task in retinal disease diagnosis. Precise segmentation of the retinal vascular pattern is challenging due to its complex structure, overlapping with other anatomical structures, and crucial thin vascular structures. In recent years, complex and heavy deep learning networks have been proposed to segment retinal blood vessels accurately. However, these methods fail to detect the thin vascular structure among different patterns of thick vessels. An attention-based novel architecture is proposed to segment the thin vasculature to address this limitation. The proposed model comprises a shallow U-Net based encoder-decoder architecture with split-fuse attention (SFA) block. The proposed SFA block enables the network to identify the placement of pixels for the tree-shaped vessel patterns at their relative position during the reconstruction phase in the decoder. The attention block aggregates low-level and high-level semantic information, improving the vessel segmentation performance. Experimentation performed on publicly available fundus datasets, DRIVE, HRF, CHASE-DB1, and STARE show that the proposed method performs better than the current state-of-the-art methods. The results demonstrate the adaptability of the proposed model for clinical applications due to its low memory footprint and better performance.
Amit Bhati, Samir Jain, Neha Gour, Pritee Khanna, Aparajita Ojha, Naoufel Werghi
ICIP6
2023 Facet-Level Segmentation of 3d Textures on Cultural Heritage Objects
abstract
Textures in 3D meshes exhibit intrinsic surface variations and are indispensable for various applications, such as retrieval, segmentation, and classification of sculptures, artifacts, and paintings. A 3D texture pattern is a locally repeated surface variation independent of the overall surface geometry and can be determined using the local neighborhood and its characteristics. Texture analysis typically employs computer vision techniques that analyze the entire 3D mesh, derive hand-crafted features, and then utilize the derived features for retrieval or classification. Several traditional and learning-based techniques exist in the literature on surface variations; however, textures are the subject of limited works. We propose a binary classification framework at the facet level for classifying texture and non-texture regions on 3D surfaces. An image sequence is generated at each facet, which serves as input to a deep vision transformer. To generate images at each facet, we construct a grid where each cell is filled with the geometric properties of its neighboring facets. We evaluated the proposed method using two datasets with diverse texture patterns, and the results are encouraging.
Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Muhammad Owais, Neha Gour, Sajid Javed, Naoufel Werghi
ICIP6
2023 Robust Nucleus Classification with Iterative Graph Representational Learning
abstract
Classifying nuclei communities in histology images is vital for early cancer treatment, but it remains challenging due to the similar structure of nuclei communities. To address this, we propose an iterative neural graph improvement and broadcasting approach. A fully connected graph is constructed with nuclei as nodes starting with a baseline classification. Node and edge features are updated and exchanged along a Hamiltonian path, removing weak connections. This process filters communities by disconnecting weakly connected nodes and iterates until stability is reached. Loose nodes from this refining stage are then assigned to their closest community clusters. Experimental results on two public datasets demonstrate the superiority of the proposed approach over state-of-the-art methods.
Taimur Hassan, Moshira Abdalla, Hina Raja, Muhammad Owais, Naoufel Werghi
ICIP5
2023 Strengthening Deep Learning Model for Robust Screening of Volumetric Chest Radiographic Scans
abstract
The emerging deep learning algorithms have shown significant potential in the development of efficient computer-aided diagnosis tools for automated detection of lung infections using chest radiographs. However, many existing methods are slice-based and require manual selection of appropriate slices from the entire CT scan, which is tedious and requires expert radiologists. To overcome these limitations, we propose a recurrent 3D Inception network (R3DI-Net) that sequentially exploits spatial and 3D structural features of the entire CT scan, ultimately leading to improved diagnostic performance. Additionally, the proposed method flexibly handles input CT scans with a variable number of slices without incurring performance degradation. A quantitative evaluation of R3DI-Net was made using a combined collection of three publicly accessible datasets containing a sufficient number of data samples. Our method outperforms various existing methods by achieving remarkable performances of 98.39%, 98.36%, 98.1%, and 98.64% in terms of accuracy, F1-score, sensitivity, and average precision, respectively.
Muhammad Owais, Taimur Hassan, Neha Gour, Iyyakutti Iyappan Ganapathi, Naoufel Werghi
ICIP5
2023 Context-Aware Transformers for Weakly Supervised Baggage Threat Localization
abstract
Recent advances in deep learning have facilitated significant progress in the autonomous detection of concealed security threats from baggage X-ray scans, a plausible solution to overcome the pitfalls of manual screening. However, these data-hungry schemes rely on extensive instance-level annotations that involve strenuous skilled labor. Hence, this paper proposes a context-aware transformer for weakly supervised baggage threat localization, exploiting their inherent capacity to learn long-range semantic relations to capture the object-level context of the illegal items. Unlike the conventional single-class token transformers, the proposed dual-token architecture can generalize well to different threat categories by learning the threat-specific semantics from the token-wise attention to generate context maps. The framework has been evaluated on two public datasets, Compass-XP and SIXray, and surpassed other SOTA approaches.
Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi
ICIP6
2023 Multi-view Inspection of Flare Stacks Operation Using a Vision-controlled Autonomous UAV
abstract
Flare stacks are crucial safety control components in petrochemical plants that required efficient monitoring and inspection. In this work, an Unmanned Aerial Vehicle (UAV)-based multi-view operation inspection system for monitoring and assessing the operation of flare stacks is proposed. Image-Based Visual Servoing (IBVS) control is used to guide the autonomous UAV for multi-view visual data collection. Afterwards, the collected visual data is analyzed using a new Multi-View Convolutional Neural Network (MV-CNN) deep learning model to obtain useful conclusions on the system's operation and classify the current state of the observed system. The proposed system's performance was validated in a simulated petrochemical plant environment with operational flare stacks and the results showed superior performance of the proposed MV-CNN model compared to a conventional single-view CNN model.
Muaz Al Radi, Hamad Karki, Naoufel Werghi, Sajid Javed, Jorge Dias 0001
IECON4
2023 Multimodal hybrid features in 3D ear recognition
Karthika Ganesan, Iyyakutti Iyappan Ganapathi, Sajid Javed, Naoufel Werghi
Appl. Intell.5
2023 RHEMAT: Robust human ear based multimodal authentication technique
Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Uttam Sharma, Pradeep Tomar, Muhammad Owais, Naoufel Werghi
Comput. Secur.6
2023 Speaker identification from emotional and noisy speech using learned voice segregation and speech VGG
Shibani Hamsa, Ismail Shahin, Youssef Iraqi, Ernesto Damiani, Ali Bou Nassif, Naoufel Werghi
Expert Syst. Appl.6
2023 Cascaded structure tensor for robust baggage threat detection
Taimur Hassan, Samet Akcay, Bilal Hassan, Mohammed Bennamoun, Salman Khan 0001, Jorge Dias 0001, Naoufel Werghi
Neural Comput. Appl.7
2023 Robot-Person Tracking in Uniform Appearance Scenarios: A New Dataset and Challenges
abstract
Person-tracking robots have many applications including security, surveillance, and autonomous driving. Despite the abundance of uniform appearance in many contexts and the challenges they exhibit, there is a lack of video datasets dedicated to benchmarking tracking algorithms in such contexts. In this article, we propose a new high-quality RGB-D benchmark called PTUA for robot–person tracking in uniform appearance scenarios. PTUA is recorded using an RGB-D sensor on top of a moving robot and consists of 45 sequences containing more than 85 K frames. Each frame is manually annotated with a bounding box and attributes, making PTUA the largest and the most challenging person tracking RGB-D dataset. To the best of our knowledge, such a densely annotated and properly synchronized RGB-D tracking benchmark does not exist in the literature. Each sequence comprises various challenges deriving from real-life scenarios where the target person appears highly similar to the background or distractors. By releasing PTUA, we expect to provide the community with a large-scale challenging RGB-D benchmark with high quality for the robust evaluation of trackers on uniform appearance scenarios for autonomous robots. We also present a rigorous experimental evaluation of the state-of-the-art trackers on the PTUA dataset with a comprehensive analysis. The findings evidence the challenges of person tracking in a uniform appearance scenario for both target tracking and robot–person tracking, and the need to bridge the performance gap. In addition, we propose a new RGB-D tracker that extracts features from RGB-D frames and it achieves the best performance on each challenging scenario of PTUA.
Xiaoxiong Zhang 0001, Adarsh Ghimire, Sajid Javed, Jorge Dias 0001, Naoufel Werghi
IEEE Trans. Hum. Mach. Syst.5
2023 Knowledge Distillation in Histology Landscape by Multi-Layer Features Supervision
abstract
Automatic tissue classification is a fundamental task in computational pathology for profiling tumor micro-environments. Deep learning has advanced tissue classification performance at the cost of significant computational power. Shallow networks have also been end-to-end trained using direct supervision however their performance degrades because of the lack of capturing robust tissue heterogeneity. Knowledge distillation has recently been employed to improve the performance of the shallow networks used as student networks by using additional supervision from deep neural networks used as teacher networks. In the current work, we propose a novel knowledge distillation algorithm to improve the performance of shallow networks for tissue phenotyping in histology images. For this purpose, we propose multi-layer feature distillation such that a single layer in the student network gets supervision from multiple teacher layers. In the proposed algorithm, the size of the feature map of two layers is matched by using a learnable multi-layer perceptron. The distance between the feature maps of the two layers is then minimized during the training of the student network. The overall objective function is computed by summation of the loss over multiple layers combination weighted with a learnable attention-based parameter. The proposed algorithm is named as Knowledge Distillation for Tissue Phenotyping (KDTP). Experiments are performed on five different publicly available histology image classification datasets using several teacher-student network combinations within the KDTP algorithm. Our results demonstrate a significant performance increase in the student networks by using the proposed KDTP algorithm compared to direct supervision-based training methods.
Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi
IEEE J. Biomed. Health Informatics4
2022 UTB180: A High-Quality Benchmark for Underwater Tracking
Basit Alawode, Mehnaz Ummar, Naoufel Werghi, Jorge Dias 0001, Ajmal Mian, Sajid Javed
ACCV (5)4
2022 Balanced Affinity Loss for Highly Imbalanced Baggage Threat Contour-Driven Instance Segmentation
abstract
Autonomous detection of threat items from baggage X-ray imagery is one of the most vital and challenging tasks. Manual detection of these items is a cumbersome, slow, and error-ridden process which is also limited by the examination capacity of the security inspector. To overcome these limitations, many researchers have proposed deep learning-driven approaches to recognize suspicious objects from the baggage X-ray scans. However, threat items are rarely seen in the real world compared to innocuous baggage content. Therefore, when trained with imbalanced data, the performance of the conventional threat detection models drastically decreases. This paper addresses these issues with a contour-driven instance segmentation model optimized with a novel combined loss function, dubbed balanced affinity loss function. In addition to mitigating the class imbalance, this function best handles the fine-grained classification aspect inferred by contours and the instance segmentation. We validated the proposed system on three public baggage X-ray datasets, where it outperformed state-of-the-art methods by 7.76%, 25.81%, and 8.78% in terms of intersection-over-union score.
Abdelfatah Hassan Ahmed, Ahmad Obeid 0001, Divya Velayudhan, Taimur Hassan, Ernesto Damiani, Naoufel Werghi
ICIP6
2022 Vision-based Inspection of Flare Stacks Operation Using a Visual Servoing Controlled Autonomous Unmanned Aerial Vehicle (UAV)
abstract
The inspection of flare stacks’ operation is a challenging task that requires technical expertise and human effort. Flare stack systems undergo various types of faults that need to be monitored in a timely manner to avoid costly and dangerous accidents. Automating this process via the application of autonomous robotic systems for collecting comprehensive data of the flare stack’s operation is a promising solution for minimizing the involved hazards and costs. In this work, a novel Unmanned Aerial Vehicle (UAV)-based autonomous inspection system for flare stacks performance monitoring is proposed. The system employs a deep learning detection network that was trained for detection of flame and smoke for vision-based flaring performance analysis. A visual servoing control technique was used for guiding the UAV’s movement throughout the inspection mission for collecting comprehensive visual inspection data. Simulations in a simulated petrochemical plant environment with flare stacks were performed for validating the performance of the proposed system. The proposed UAV system was able to collect the required data successfully and analysis of the obtained data returned useful information about the flare stack’s operation.
Muaz Al Radi, Hamad Karki, Naoufel Werghi, Sajid Javed, Jorge Dias 0001
IECON3
2022 Learning to localize image forgery using end-to-end attention network
Iyyakutti Iyappan Ganapathi, Sajid Javed, Syed Sadaf Ali, Arif Mahmood, Ngoc-Son Vu, Naoufel Werghi
Neurocomputing6
2022 Unsupervised Land-Cover Segmentation Using Accelerated Balanced Deep Embedded Clustering
abstract
In this letter, we address the issue of the automatic labeling of remote sensing datasets using a novel deep learning clustering algorithm. The proposed algorithm addresses the inherent susceptibility of the deep embedded clustering (DEC) algorithm to data imbalance using additional search and extraction steps. Furthermore, the proposed algorithm is highly parallelizable. A graphics processing unit (GPU) implementation is shown to achieve 40X to 2600X of performance speedup and improved clustering accuracy with respect to DEC and other clustering approaches.
Ahmad Obeid 0001, Ibrahim M. Elfadel, Naoufel Werghi
IEEE Geosci. Remote. Sens. Lett.3
2022 Nucleus classification in histology images using message passing network
Taimur Hassan, Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi, Nasir M. Rajpoot
Medical Image Anal.5
2022 Tensor pooling-driven instance segmentation framework for baggage threat recognition
Taimur Hassan, Samet Akcay, Mohammed Bennamoun, Salman Khan 0001, Naoufel Werghi
Neural Comput. Appl.5
2022 Hierarchical Spatiotemporal Graph Regularized Discriminative Correlation Filter for Visual Object Tracking
abstract
Visual object tracking is a fundamental and challenging task in many high-level vision and robotics applications. It is typically formulated by estimating the target appearance model between consecutive frames. Discriminative correlation filters (DCFs) and their variants have achieved promising speed and accuracy for visual tracking in many challenging scenarios. However, because of the unwanted boundary effects and lack of geometric constraints, these methods suffer from performance degradation. In the current work, we propose hierarchical spatiotemporal graph-regularized correlation filters for robust object tracking. The target sample is decomposed into a large number of deep channels, which are then used to construct a spatial graph such that each graph node corresponds to a particular target location across all channels. Such a graph effectively captures the spatial structure of the target object. In order to capture the temporal structure of the target object, the information in the deep channels obtained from a temporal window is compressed using the principal component analysis, and then, a temporal graph is constructed such that each graph node corresponds to a particular target location in the temporal dimension. Both spatial and temporal graphs span different subspaces such that the target and the background become linearly separable. The learned correlation filter is constrained to act as an eigenvector of the Laplacian of these spatiotemporal graphs. We propose a novel objective function that incorporates these spatiotemporal constraints into the DCFs framework. We solve the objective function using alternating direction methods of multipliers such that each subproblem has a closed-form solution. We evaluate our proposed algorithm on six challenging benchmark datasets and compare it with 33 existing state-of-the art trackers. Our results demonstrate an excellent performance of the proposed algorithm compared to the existing trackers.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Lakmal D. Seneviratne, Naoufel Werghi
IEEE Trans. Cybern.5
2022 A Novel Incremental Learning Driven Instance Segmentation Framework to Recognize Highly Cluttered Instances of the Contraband Items
abstract
Screening cluttered and occluded contraband items from baggage X-ray scans is a cumbersome task even for the expert security staff. This article presents a novel strategy that extends a conventional encoder-decoder architecture to perform instance-aware segmentation and extract merged instances of contraband items without using any additional subnetwork or an object detector. The encoder-decoder network first performs conventional semantic segmentation and retrieves cluttered baggage items. The model then incrementally evolves during training to recognize individual instances using significantly reduced training batches. To avoid catastrophic forgetting, a novel objective function minimizes the network loss in each iteration by retaining the previously acquired knowledge while learning new class representations and resolving their complex structural interdependencies through Bayesian inference. A thorough evaluation of our framework on two publicly available X-ray datasets shows that it outperforms state-of-the-art methods, especially within the challenging cluttered scenarios, while achieving an optimal tradeoff between detection accuracy and efficiency.
Taimur Hassan, Samet Akcay, Mohammed Bennamoun, Salman Khan 0001, Naoufel Werghi
IEEE Trans. Syst. Man Cybern. Syst.5
2021 Temporal Fusion Based Mutli-scale Semantic Segmentation for Detecting Concealed Baggage Threats
abstract
Detection of illegal and threatening items in bag-gage is one of the utmost security concern nowadays. Even for experienced security personnel, manual detection is a time-consuming and stressful task.Many academics have created automated frameworks for detecting suspicious and contraband data from X-ray scans of luggage. However, to our knowledge, no framework exists that utilizes temporal baggage X-ray imagery to effectively screen highly concealed and occluded objects which are barely visible even to the naked eye. To address this, we present a novel temporal fusion driven multi-scale residual fashioned encoder-decoder that takes series of consecutive scans as input and fuses them to generate distinct feature representations of the suspicious and non-suspicious baggage content, leading towards a more accurate extraction of the contraband data. The proposed methodology has been thoroughly tested using the publicly accessible GDXray dataset, which is the only dataset containing temporally linked grayscale X-ray scans showcasing extremely concealed contraband data. The proposed framework outperforms its competitors on the GDXray dataset on various metrics.
Muhammad Shafay, Taimur Hassan, Ernesto Damiani, Naoufel Werghi
SMC4
2021 Spatially Constrained Context-Aware Hierarchical Deep Correlation Filters for Nucleus Detection in Histology Images
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi, Nasir M. Rajpoot
Medical Image Anal.4
2021 Representing and analyzing relief patterns using LBP variants on mesh manifold
Claudio Tortorici, Naoufel Werghi, Stefano Berretti
Pattern Anal. Appl.2
2021 Convolution operations for relief-pattern retrieval, segmentation and classification on mesh manifolds
Claudio Tortorici, Stefano Berretti, Ahmad Obeid 0001, Naoufel Werghi
Pattern Recognit. Lett.4
2021 Estimating ambient visibility in the presence of fog: a deep convolutional neural network approach
Fatma Outay, Bilal Taha, Hazar Chaabani, Faouzi Kamoun, Naoufel Werghi, Ansar-Ul-Haque Yasar
Pers. Ubiquitous Comput.5
2021 Estimating Left Ventricle Ejection Fraction Levels Using Circadian Heart Rate Variability Features and Support Vector Regression Models
abstract
OBJECTIVES: The purpose of this study was to set an optimal fit of the estimated LVEF at hourly intervals from 24-hour ECG recordings and compare it with the fit based on two gold-standard guidelines. METHODS: Support vector regression (SVR) models were applied to estimate LVEF from ECG derived heart rate variability (HRV) data in one-hour intervals from 24-hour ECG recordings of patients with either preserved, mid-range, or reduced LVEF, obtained from the Intercity Digital ECG Alliance (IDEAL) study. A step-wise feature selection approach was used to ensure the best possible estimations of LVEF levels. RESULTS: The experimental results have shown that the lowest Root Mean Square Error (RMSE) between the original and estimated LVEF levels was during 3-4 am, 5-6 am and 6-7 pm. CONCLUSION: The observations suggest these hours as possible times for intervention and optimal treatment outcomes. In addition, LVEF classifications following the ACCF/AHA guidelines leads to a more accurate assessment of mid-range LVEF. SIGNIFICANCE: This study paves the way to explore the use of HRV features in the prediction of LVEF percentages as an indicator of disease progression, which may lead to an automated classification process for CAD patients.
Mohanad Alkhodari, Herbert F. Jelinek, Naoufel Werghi, Leontios J. Hadjileontiadis, Ahsan H. Khandoker
IEEE J. Biomed. Health Informatics3
2021 RAG-FW: A Hybrid Convolutional Framework for the Automated Extraction of Retinal Lesions and Lesion-Influenced Grading of Human Retinal Pathology
abstract
The identification of retinal lesions plays a vital role in accurately classifying and grading retinopathy. Many researchers have presented studies on optical coherence tomography (OCT) based retinal image analysis over the past. However, to the best of our knowledge, there is no framework yet available that can extract retinal lesions from multi-vendor OCT scans and utilize them for the intuitive severity grading of the human retina. To cater this lack, we propose a deep retinal analysis and grading framework (RAG-FW). RAG-FW is a hybrid convolutional framework that extracts multiple retinal lesions from OCT scans and utilizes them for lesion-influenced grading of retinopathy as per the clinical standards. RAG-FW has been rigorously tested on 43,613 scans from five highly complex publicly available datasets, containing multi-vendor scans, where it achieved the mean intersection-over-union score of 0.8055 for extracting the retinal lesions and the accuracy of 98.70% for the correct severity grading of retinopathy.
Taimur Hassan, M. Usman Akram, Naoufel Werghi, Muhammad Noman Nazir
IEEE J. Biomed. Health Informatics3
2020 Trainable Structure Tensors for Autonomous Baggage Threat Detection Under Extreme Occlusion
Taimur Hassan, Naoufel Werghi
ACCV (6)2
2020 Exploiting the Transferability of Deep Learning Systems Across Multi-modal Retinal Scans for Extracting Retinopathy Lesions
abstract
Retinal lesions play a vital role in the accurate classification of retinal abnormalities. Many researchers have proposed deep lesion-aware screening systems that analyze and grade the progression of retinopathy. However, to the best of our knowledge, no literature exploits the tendency of these systems to generalize across multiple scanner specifications and multi-modal imagery. Towards this end, this paper presents a detailed evaluation of semantic segmentation, scene parsing and hybrid deep learning systems for extracting the retinal lesions such as intra-retinal fluid, sub-retinal fluid, hard exudates, drusen, and other chorioretinal anomalies from fused fundus and optical coherence tomography (OCT) imagery. Furthermore, we present a novel strategy exploiting the transferability of these models across multiple retinal scanner specifications. A total of 363 fundus and 173,915 OCT scans from seven publicly available datasets were used in this research (from which 297 fundus and 59,593 OCT scans were used for testing purposes). Overall, a hybrid retinal analysis and grading network (RAGNet), backboned through ResNet50, stood first for extracting the retinal lesions, achieving a mean dice coefficient score of 0.822. Moreover, the complete source code and its documentation are released at http://biomisa.org/index.php/downloads/.
Taimur Hassan, M. Usman Akram, Naoufel Werghi
BIBE3
2020 Deep Bidirectional Correlation Filters for Visual Object Tracking
abstract
Visual Object Tracking (VOT) is an essential task for many computer vision applications. VOT becomes challenging when a target object faces severe occlusion, drastic illumination changes, and scale variation problems. In the literature, Discriminative Correlation Filters (DCFs)-based tracking methods have achieved promising results in terms of accuracy and efficiency in many complex VOT scenarios. A plethora of DCFs trackers have been proposed which exploit information observed in past frames to create and update DCFs for VOT. To adapt to target appearance variations, the DCFs are enhanced by incorporating spatial and temporal consistency constraints. Nevertheless, the performance degradation is observed for these methods because of the aforementioned limitations. To address these issues, we propose a novel algorithm based on bidirectional DCFs for VOT. In this algorithm, we propose the original idea of leveraging information from both past and future frames. The proposed algorithm first tracks the target object forward in the video sequence and then its uses the predicted location of the last window frame and track the target object backward towards the current frame. We design an appearance consistency loss function by taking the$L_{2}$norm between the regression target of the forward tracking and response map of the backward tracking to obtain the resulting response map. Our proposed algorithm realizes a highly accurate DCFs because forward and backward tracking information are fused together for consistent VOT. Although, a result will be output with some small delay because information is taken from a future to the present period, our proposed algorithm has the merit of addressing the drastic appearance variations VOT challenges. We evaluate our proposed tracker using deep features on three publicly available challenging datasets. Our results demonstrate the superior performance of the proposed tracker compared to the existing state-of-the-art trackers.
Sajid Javed, Xiaoxiong Zhang 0001, Lakmal D. Seneviratne, Jorge Dias 0001, Naoufel Werghi
FUSION5
2020 Localizing Firearm Carriers By Identifying Human-Object Pairs
abstract
Visual identification of gunmen in a crowd is a challenging problem, that requires resolving the association of a person with an object (firearm). We present a novel approach to address this problem, by defining human-object interaction (and non-interaction) bounding boxes. In a given image, human and firearms are separately detected. Each detected human is paired with each detected firearm, allowing us to create a paired bounding box that contains both object and the human. A network is trained to classify these paired-bounding-boxes into human carrying the identified firearm or not. Extensive experiments were performed to evaluate the effectiveness of the algorithm, including exploiting full pose of the human, hand-keypoints, and their association with the firearm. The knowledge of spatially localized features is key to the success of our method by using multi-size proposals with adaptive average pooling. We have also extended a previously existing firearm detection dataset, by adding more images and tagging in the extended dataset the human-firearm pairs (including bounding boxes for firearms and gunmen). The experimental results $({78.5 AP}_{hold})$ demonstrate effectiveness of the proposed method.
Abdul Basit 0019, Muhammad Akhtar Munir, Mohsen Ali, Naoufel Werghi, Arif Mahmood
ICIP4
2020 Detecting Prohibited Items in X-Ray Images: a Contour Proposal Learning Approach
abstract
X-ray baggage screening plays a vital role in aviation security. Manual inspection of potentially anomalous items is challenging due to the clutter and occlusion within Xray scans. Here, we address this issue by presenting an object-boundaries driven framework for the automated detection of suspicious items from X-ray baggage scans. Rather than recognizing objects directly from the X-ray images, our two-stage detection approach first extracts contour-based proposals using a novel cascaded structure tensor technique and subsequently passes the candidate proposals to a single feed-forward convolutional neural network for recognition. Thorough experimentation on GDXray and SIXray datasets demonstrates that the proposed model achieves a mean area under the curve of 0.9878, outperforming the existing renown state-of-the-art object detection frameworks.
Taimur Hassan, Meriem Bettayeb, Samet Akcay, Salman Khan 0001, Mohammed Bennamoun, Naoufel Werghi
ICIP6
2020 CS-RPCA: Clustered Sparse RPCA for Moving Object Detection
abstract
Moving object detection (MOD) is an important step for many computer vision applications. In the last decade, it is evident that RPCA has shown to be a potential solution for MOD and achieved a promising performance under various challenging background scenes. However, because of the lack of different types of features, RPCA still shows degraded performance in many complicated background scenes such as dynamic backgrounds, cluttered foreground objects, and camouflage. To address these problems, this paper presents a Clustered Sparse RPCA (CS-RPCA) for MOD under challenging environments. The proposed algorithm extracts multiple features from video sequences and then employs RPCA to get the low-rank and sparse component from each representation. The sparse subspaces are then emerged into a common sparse component using Grassmann manifold. We proposed a novel objective function which computes the composite sparse component from multiple representations and it is solved using non-negative matrix factorization method. The proposed algorithm is evaluated on two challenging datasets for MOD. Results demonstrate excellent performance of the proposed algorithm as compared to existing state-of-the-art methods.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi
ICIP4
2020 Diagnosing Autism Using T1-W MRI With Multi-Kernel Learning and Hypergraph Neural Network
abstract
The field of network neuroscience provided unprecedented insights into how brain connectivity gets altered by autism spectrum disorder (ASD) on functional, structural, and morphological levels. However, a few studies have looked to design a framework that captures the complex network structure of the brain and disentangles the heterogeneity of ASD. In this paper, we leverage multi-kernel unsupervised learning in the construction of multiview hypergraph neural networks (HGNN), each capturing a particular view of the brain connectome, to eventually distinguish between ASD and normal control (NC) subjects. Additionally, we tested and measured how our proposed framework compares to other variants based on previous baseline methods. Our classification results outperformed comparison methods and agreed with the literature in the sense that the right hemisphere connectivity was more discriminative in ASD diagnosis than the left hemisphere.
Mohammad Moussa Madine, Islem Rekik, Naoufel Werghi
ICIP3
2020 Fused Geometry Augmented Images For Analyzing Textured Mesh
abstract
In this paper, we propose a multi-modal mesh surface representation by fusing texture and geometric data. Our approach defines an inverse mapping between different geometric descriptors computed on the mesh surface, and the corresponding 2D texture image of the mesh, allowing the construction of fused geometrically augmented images. This new fused modality enables us to learn feature representations from 3D data in a highly efficient manner by employing standard convolutional neural networks in a transfer-learning mode. In contrast to existing methods, the proposed approach is both computationally and memory efficient, preserves intrinsic geometric information and learns highly discriminative feature representations by effectively fusing shape and texture information at the data level. The efficacy is demonstrated on the task of facial expression classification, showing competitive performance with state-of-the-art methods.
Bilal Taha, Munawar Hayat, Stefano Berretti, Naoufel Werghi
ICIP4
2020 CSIOR: An Algorithm For Ordered Triangular Mesh Regularization
abstract
3D scanners generate irregularly distributed cloud of points in most of the cases. Dealing with such data, often in the form of triangular meshes, requires a pre-processing step to regularize the triangle facets shape and size. In this paper, we propose CSIOR, a novel mesh regularization technique which is capable of producing quasi-equilateral triangles, and distinguished by two novel features, namely, its intrinsic ordered aspect and its preservation of the geometric texture of the surface (relief patterns). We evidence the superiority of our technique over current methods through a series of experiments performed on a variety of geometric textured surfaces.
Claudio Tortorici, Naoufel Werghi, Stefano Berretti
ICIP2
2020 Gender Recognition on RGB-D Image
abstract
In this paper, we propose a deep-learning approach for human gender classification on RGB-D images. Unlike most of the existing methods, which use hand-crafted features from the human face, we exploit local information from the head and global information from the whole body to classify people's gender. A head detector is fine-tuned on YOLO to detect the head regions on the images automatically. Two gender classifiers are trained using head images and whole-body images separately. The final prediction is made by fusing the two classifiers' results. The presented method outperforms the state-of-art with an improvement in the accuracy of 2.6%, 7.6%, and 8.4% on three different test data of a challenging gender dataset which includes human standing, walking, and interacting scenarios.
Xiaoxiong Zhang 0001, Sajid Javed, Ahmad Obeid 0001, Jorge Dias 0001, Naoufel Werghi
ICIP5
2020 Long-Range Visual UAV Detection and Tracking System with Threat Level Assessment
abstract
Unmanned aerial vehicles (UAVs) can pose a serious threat to critical infrastructure which has motivated researchers to develop solutions for early detection. Nevertheless, the problem remains unsolved due to the limitations of the current detection techniques. In this paper, a vision-based approach using deep learning and a pan-tilt-zoom camera is proposed. In addition to detecting and tracking UAVs at long distances, the approach also assesses the threat level of the intruder UAVs based on their orientation. The proposed system offers long-range coverage while being cheap and practically feasible.
Abdel Gafoor Haddad, Muhammad Ahmed Humais, Naoufel Werghi, Abdulhadi Shoufan
IECON3
2020 CSIOR: Circle-Surface Intersection Ordered Resampling
Claudio Tortorici, Mohamed Kamel Riahi, Stefano Berretti, Naoufel Werghi
Comput. Aided Geom. Des.4
2020 SHREC 2020: Retrieval of digital surfaces with similar geometric reliefs
Elia Moscoso Thompson, Silvia Biasotti, Andrea Giachetti 0001, Claudio Tortorici, Naoufel Werghi, Ahmad Obeid 0001, Stefano Berretti, Hoang-Phuc Nguyen-Dinh, Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran, Leonardo Gigli, Santiago Velasco-Forero, Beatriz Marcotegui, Ivan Sipiran, Benjamin Bustos, Ioannis Romanelis, Vlassis Fotis, Ramamoorthy Luxman
Comput. Graph.5
2020 Statistical 3D watermarking algorithm using non negative matrix factorization
Nassima Medimegh, Samir Belaid, Mohamed Atri, Naoufel Werghi
Multim. Tools Appl.4
2020 Learned 3D Shape Representations Using Fused Geometrically Augmented Images: Application to Facial Expression and Action Unit Detection
abstract
In this paper, we propose an approach to learn generic multi-modal mesh surface representations using a novel scheme for fusing texture and geometric data. Our approach defines an inverse mapping between different geometric descriptors computed on the mesh surface or its down-sampled version, and the corresponding 2D texture image of the mesh, allowing the construction of fused geometrically augmented images (FGAI). This new fused modality enables us to learn feature representations from 3D data in a highly efficient manner by simply employing standard CNNs in a transfer-learning mode. The proposed approach is both computationally and memory efficient, preserves intrinsic geometric information and learns highly discriminative feature representations by effectively fusing shape and texture information at data level. The efficacy of our approach is demonstrated for the tasks of facial action unit detection and expression classification. The extensive experiments conducted on the Bosphorus and BU-4DFE datasets show that our method produces a significant boost in the performance when compared to state-of-the-art solutions.
Bilal Taha, Munawar Hayat, Stefano Berretti, Dimitrios Hatzinakos, Naoufel Werghi
IEEE Trans. Circuits Syst. Video Technol.5
2020 Robust Structural Low-Rank Tracking
abstract
Visual object tracking is an essential task for many computer vision applications. It becomes very challenging when the target appearance changes especially in the presence of occlusion, background clutter, and sudden illumination variations. Methods, that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when facing the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking in complex scenarios. In the proposed algorithm, we consider spatial and temporal appearance consistency constraints, among the particles in the low-rank subspace, embedded in four different graphs. The resulting objective function encoding these constraints is novel and it is solved using linearized alternating direction method with adaptive penalty both in batch fashion as well as in online fashion. Our proposed objective function jointly learns the spatial and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on four challenging datasets demonstrate excellent performance of the proposed algorithm as compared to current state-of-the-art methods.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi
IEEE Trans. Image Process.4
2020 Multiplex Cellular Communities in Multi-Gigapixel Colorectal Cancer Histology Images for Tissue Phenotyping
abstract
In computational pathology, automated tissue phenotyping in cancer histology images is a fundamental tool for profiling tumor microenvironments. Current tissue phenotyping methods use features derived from image patches which may not carry biological significance. In this work, we propose a novel multiplex cellular community-based algorithm for tissue phenotyping integrating cell-level features within a graph-based hierarchical framework. We demonstrate that such integration offers better performance compared to prior deep learning and texture-based methods as well as to cellular community based methods using uniplex networks. To this end, we construct celllevel graphs using texture, alpha diversity and multi-resolution deep features. Using these graphs, we compute cellular connectivity features which are then employed for the construction of a patch-level multiplex network. Over this network, we compute multiplex cellular communities using a novel objective function. The proposed objective function computes a low-dimensional subspace from each cellular network and subsequently seeks a common low-dimensional subspace using the Grassmann manifold. We evaluate our proposed algorithm on three publicly available datasets for tissue phenotyping, demonstrating a significant improvement over existing state-of-the-art methods.
Sajid Javed, Arif Mahmood, Naoufel Werghi, Ksenija Benes, Nasir M. Rajpoot
IEEE Trans. Image Process.3
2019 Detecting Drivers Smartphone: A Learned Features Approach using Aggregated Scalogram Images
abstract
In this paper, we propose an image representation approach for detecting driver mobile phone from the accelerometer signals produced by a set of smartphones in a vehicle. Rather than following the classic paradigm of classifying the signal as driver or non-driver, we propose an original paradigm whereby we aggregate the signals together and train a classifier to detect the driver signal in that aggregation. We do so by stacking-up the Scalograms images of the smartphone signals and training a CNN classifier to identify the driver's Scalograms instance in the Scalograms stack image. To the best our knowledge, this is the first time such an image-fusion and classification scheme is proposed for detecting driver's smartphone. Experiments performed with an in-house dataset confirms the potential and the merit of our approach.
Mohammad Moussa Madine, Ammar Battah, Aaminah Khan, Naoufel Werghi
AICCSA4
2019 Structural Low-Rank Tracking
abstract
Visual object tracking is an important step for many computer vision applications. The task becomes very challenging when the target undergoes heavy occlusion, background clutters, and sudden illumination variations. Methods that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when an object faces the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking. In the proposed algorithm, we enforce local spatial, global spatial and temporal appearance consistency among the particles in the low-rank subspace by constructing three graphs. The Laplacian matrices of these graphs are incorporated into the novel low-rank objective function which is solved using linearized alternating direction method with an adaptive penalty. Our proposed objective function jointly learns the spatial, global, and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on two challenging benchmark datasets show the superiority of the proposed algorithm as compared to current state-of-the-art methods.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi
AVSS4
2019 Extending LBP and Convolution-Like Operations on the Mesh
abstract
Extending the concept of texture to the geometry of a mesh manifold surface is an emerging topic in image processing. This concept is different from gluing images to the surface, but rather indicates the presence of relief patterns that locally change the surface geometry, showing some regular and repetitive pattern. In this paper, we propose an efficient and effective framework to address this novel task, which encompasses the convolution operation and the casting of a variety of Local Binary Pattern on the mesh manifold. Results show that our technique outperforms the existing state-of-the-art methods in the challenging task of relief patterns classification.
Claudio Tortorici, Naoufel Werghi, Stefano Berretti
ICIP2
2018 Balancing Incident and Ambient Light for Illumination Compensation in Video Applications
abstract
Hand-gesture interaction with a smart TV using a conventional camera under restricted ambient illumination is complicated by the varying incident illumination coming from the screen. In this paper, we propose a method for compensating for the poor ambient illumination in the scene captured by a smart TV camera by balancing it against this kind of incident illumination. The proposed framework models the variations in the ambient illumination as a factor of the varying incident illumination caused by the changing multimedia content displayed on the smart TV. Further, by estimating the average incident illumination and measuring its deviation during the scene transition, an illumination intensity correction factor is deduced, and thus ambient illumination is corrected iteratively. To validate the effectiveness of the proposed method, we benchmark it against multiple video sequences depicting the human gesture interaction with a smart device under both low-lighting (typically in the night) and natural-lighting conditions. Furthermore, experiments on such challenging video sequences have demonstrated improved accuracy and robustness of visual target trackers by preprocessing the sequence using our method.
Buti Al Delail, Harish Bhaskar, Mohamed Jamal Zemerly, Naoufel Werghi
ICIP4
2018 3D mesh watermarking using salient points
Nassima Medimegh, Samir Belaid, Mohamed Atri, Naoufel Werghi
Multim. Tools Appl.4
2017 Joint Registration and Representation Learning for Unconstrained Face Identification
abstract
Recent advances in deep learning have resulted in human-level performances on popular unconstrained face datasets including Labeled Faces in the Wild and YouTube Faces. To further advance research, IJB-A benchmark was recently introduced with more challenges especially in the form of extreme head poses. Registration of such faces is quite demanding and often requires laborious procedures like facial landmark localization. In this paper, we propose a Convolutional Neural Networks based data-driven approach which learns to simultaneously register and represent faces. We validate the proposed scheme on template based unconstrained face identification. Here, a template contains multiple media in the form of images and video frames. Unlike existing methods which synthesize all template media information at feature level, we propose to keep the template media intact. Instead, we represent gallery templates by their trained one-vs-rest discriminative models and then employ a Bayesian strategy which optimally fuses decisions of all medias in a query template. We demonstrate the efficacy of the proposed scheme on IJB-A, YouTube Celebrities and COX datasets where our approach achieves significant relative performance boosts of 3.6%, 21.6% and 12.8% respectively.
Munawar Hayat, Salman Khan 0001, Naoufel Werghi, Roland Göcke
CVPR3
2017 Convolutional neural networkasa feature extractor for automatic polyp detection
abstract
Colorectal cancer is one of the major causes of cancer deaths worldwide. To achieve early cancer screening, detecting the presence of polyps in the colon tract is the preferred technique. In this paper, a deep learning approach for identifying polyps in colonoscopy images is proposed. The novelty of our technique stems from the fact that it fully employs a pre-trained Convolutional Neural Network (CNN) architecture as a feature extractor. Contrary to the conventional methods which either perform fine-tuning or train the CNN from scratch, we utilize the CNN output features as an input to train the Support Vector Machine (SVM) Classifier. The efficiency of the presented framework is demonstrated on the public CVC ColonDB, in which the experimental results indicate that our methodology significantly outperforms other competitive paradigms.
Bilal Taha, Jorge Dias 0001, Naoufel Werghi
ICIP3
2016 3D constrained local model with independent component analysis and non-Gaussian shape prior distribution: Application to 3D facial landmark detection
abstract
We present a novel statistical shape model and fitting process for the 3D Constrained Local Models (CLM), exploiting the properties of Independent Component Analysis (ICA), instead of the classic use of Principal Component Analysis (PCA), and adopting a non-Gaussian distribution of the shape prior information. Using ICA permits to exploit the real distribution of shape priors by adopting a Generalised Gaussian Distribution (GGD) model. Consequently, we derive a modified approach of the mean shift optimizer by using the Expectation-Maximization algorithms. We apply this novel method for the localization of face landmarks on 3D facial mesh models, which, to the best of our knowledge, is the first employment of the CLM variant on this kind of modality. Experiments conduced on the Bosphorus face database demonstrated that our approach outperforms state-of-the-art methods.
Marwa Chendeb, Claudio Tortorici, Hassan Al-Muhairi, Marius George Linguraru, Naoufel Werghi
ICIP5
2016 Computer-aided diagnostic tool for early detection of prostate cancer
abstract
In this paper, we propose a novel non-invasive framework for the early diagnosis of prostate cancer from diffusion-weighted magnetic resonance imaging (DW-MRI). The proposed approach consists of three main steps. In the first step, the prostate is localized and segmented based on a new level-set model. In the second step, the apparent diffusion coefficient (ADC) of the segmented prostate volume is mathematically calculated for different b-values. To preserve continuity, the calculated ADC values are normalized and refined using a Generalized Gauss-Markov Random Field (GGMRF) image model. The cumulative distribution function (CDF) of refined ADC for the prostate tissues at different b-values are then constructed. These CDFs are considered as global features describing water diffusion which can be used to distinguish between benign and malignant tumors. Finally, a deep learning auto-encoder network, trained by a stacked non-negativity constraint algorithm (SNCAE), is used to classify the prostate tumor as benign or malignant based on the CDFs extracted from the previous step. Preliminary experiments on 53 clinical DW-MRI data sets resulted in 100% correct classification, indicating the high accuracy of the proposed framework and holding promise of the proposed CAD system as a reliable non-invasive diagnostic tool.
Islam Reda, Ahmed Shalaby 0002, Fahmi Khalifa, Mohammed M. Elmogy, Ahmed Abou El-Fetouh, Mohamed Abou El-Ghar, Ehsan Hosseini-Asl, Naoufel Werghi, Robert Keynton, Ayman El-Baz
ICIP8
2016 Boosting 3D LBP-Based Face Recognition by Fusing Shape and Texture Descriptors on the Mesh
abstract
In this paper, we present a novel approach for fusing shape and texture local binary patterns (LBPs) on a mesh for 3D face recognition. Using a recently proposed framework, we compute LBP directly on the face mesh surface, then we construct a grid of the regions on the facial surface that can accommodate global and partial descriptions. Compared with its depth-image counterpart, our approach is distinguished by the following features: 1) inherits the intrinsic advantages of mesh surface (e.g., preservation of the full geometry); 2) does not require normalization; and 3) can accommodate partial matching. In addition, it allows early level fusion of texture and shape modalities. Through experiments conducted on the BU-3DFE and Bosphorus databases, we assess different variants of our approach with regard to facial expressions and missing data, also in comparison to the state-of-the-art solutions.
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Inf. Forensics Secur.1
2015 Representing 3D texture on mesh manifolds for retrieval and recognition applications
abstract
In this paper, we present and experiment a novel approach for representing texture of 3D mesh manifolds using local binary patterns (LBP). Using a recently proposed framework [37], we compute LBP directly on the mesh surface, either using geometric or photometric appearance. Compared to its depth-image counterpart, our approach is distinguished by the following features: a) inherits the intrinsic advantages of mesh surface (e.g., preservation of the full geometry); b) does not require normalization; c) can accommodate partial matching. In addition, it allows early-level fusion of the geometry and photometric texture modalities. Through experiments conducted on two application scenarios, namely, 3D texture retrieval and 3D face recognition, we assess the effectiveness of the proposed solution with respect to state of the art approaches.
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo
CVPR1
2015 Boosting 3D LBP-based face recognition by fusing shape and texture descriptors on the mesh
abstract
In this paper, we present a novel approach for fusing shape and texture local binary patterns (LBP) for 3D face recognition. Using the framework proposed in [1], we compute LBP directly on the face mesh surface, then we construct a grid of the regions on the facial surface that can accommodate global and partial descriptions. Compared to its depth-image counterpart, our approach is distinguished by the following features: a) inherits the intrinsic advantages of mesh surface; b) does not require normalization; c) can accommodate partial matching. In addition, it allows early-level fusion of texture and shape modalities. Through experiments conducted on the BU-3DFE and Bosphorus databases, we assess different variants of our approach with regard to facial expressions and missing data.
Claudio Tortorici, Naoufel Werghi, Stefano Berretti
ICIP2
2015 Local binary patterns on triangular meshes: Concept and applications
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo
Comput. Vis. Image Underst.1
2015 The Mesh-LBP: A Framework for Extracting Local Binary Patterns From Discrete Manifolds
abstract
In this paper, we present a novel and original framework, which we dubbed mesh-local binary pattern (LBP), for computing local binary-like-patterns on a triangular-mesh manifold. This framework can be adapted to all the LBP variants employed in 2D image analysis. As such, it allows extending the related techniques to mesh surfaces. After describing the foundations, the construction and the main features of the mesh-LBP, we derive its possible variants and show how they can extend most of the 2D-LBP variants to the mesh manifold. In the experiments, we give evidence of the presence of the uniformity aspect in the mesh-LBP, similar to the one noticed in the 2D-LBP. We also report repeatability experiments that confirm, in particular, the rotation-invariance of mesh-LBP descriptors. Furthermore, we analyze the potential of mesh-LBP for the task of 3D texture classification of triangular-mesh surfaces collected from public data sets. Comparison with state-of-the-art surface descriptors, as well as with 2D-LBP counterparts applied on depth images, also evidences the effectiveness of the proposed framework. Finally, we illustrate the robustness of the mesh-LBP with respect to the class of mesh irregularity typical to 3D surface-digitizer scans.
Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Image Process.1
2014 Computing Local Binary Patterns on Discrete Manifolds
abstract
In this paper, we present a novel and original framework for computing Local Binary Pattern (LBP)-like patterns on a triangular mesh manifold. This framework, that we called mesh-LBP, can be adapted to all the LBP variants employed in 2D image analysis. As such, it allows extending the related techniques to mesh surfaces. First, we describe the foundations, the construction and the features of the mesh-LBP. In the experiments, we first show evidence of the presence of the "uniformity" aspect in the mesh-LBP patterns. Then, we report about the application of mesh-LBP to the problem of 3D texture-classification in comparison to standard 3D surface descriptors and show the mesh-LBP robustness to mesh irregularities.
Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo
ICPR1
2014 Multiscale Roughness Approach for Assessing Posterior Capsule Opacification
abstract
Posterior capsule opacification (PCO) is a common complication in patients who have undergone cataract surgery, occurring in up to 50% of patients by two to three years after the operation. Assessment of PCO has been mainly subjective, making it difficult to understand its progression over time or assess the effectiveness of strategies used for the prevention of PCO. Fully automated PCO assessment systems developed so far offer objective grades. However, they do not provide morphological PCO data useful for an effective analysis of scores. This paper proposes a novel method based on multiscale roughness estimation to detect and quantify the PCO areas. This method is also characterized by its robustness against monotonic illumination variations. Extensive experimentation showcases a distinctive analysis and assessment power of our method compared to other competitive methods. The results show a high correlation of 84.6% with respect to clinical scores.
Aruna Vivekanand, Naoufel Werghi
IEEE J. Biomed. Health Informatics2
2014 Selecting stable keypoints and local descriptors for person identification using 3D face scans
Stefano Berretti, Naoufel Werghi, Alberto Del Bimbo, Pietro Pala
Vis. Comput.2
2013 Robust and Fragile Watermarking Scheme Based on DCT and Hash Function for Color Satellite Images
Sultan Al-Shehhi, Mohammed Al-Muhairi, Dalia Abuoeida, Naoufel Werghi, Alavi Kunhu
DeSE5
2013 Local descriptors matching for 3D face recognition
abstract
An original solution to 3D face recognition, which supports face matching also in the case of probes with varying expressions and missing parts is proposed in this work. Distinguishing traits of the face are captured by first extracting 3D keypoints of the face scan, then measuring how the face surface changes in the neighborhood of the keypoints using a local descriptor. To this end, an adaptation of the meshDOG detector to the case of 3D faces is proposed, together with a multi-ring geometric histogram descriptor. Face similarity is then evaluated by comparing local keypoint descriptors across inlier pairs of matching keypoints between probe and gallery scans. Experiments have been performed on the Bosphorus database, showing competitive results with respect to existing solutions for 3D face biometrics.
Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICIP1
2013 Matching 3D face scans using interest points and local histogram descriptors
abstract
In this work, we propose and experiment an original solution to 3D face recognition that supports face matching also in the case of probe scans with missing parts. In the proposed approach, distinguishing traits of the face are captured by first extracting 3D keypoints of the scan and then measuring how the face surface changes in the keypoints neighborhood using local shape descriptors. In particular: 3D keypoints detection relies on the adaptation to the case of 3D faces of the meshDOG algorithm that has been demonstrated to be effective for 3D keypoints extraction from generic objects; as 3D local descriptors we used the HOG descriptor and also proposed two alternative solutions that develop, respectively, on the histogram of orientations and the geometric histogram descriptors. Face similarity is evaluated by comparing local shape descriptors across inlier pairs of matching keypoints between probe and gallery scans. The face recognition accuracy of the approach has been first experimented on the difficult probes included in the new 2D/3D Florence face dataset that has been recently collected and released at the University of Firenze, and on the Binghamton University 3D facial expression dataset. Then, a comprehensive comparative evaluation has been performed on the Bosphorus, Gavab and UND/FRGC v2.0 databases, where competitive results with respect to existing solutions for 3D face biometrics have been obtained.
Stefano Berretti, Naoufel Werghi, Alberto Del Bimbo, Pietro Pala
Comput. Graph.2
2012 Detection and segmentation of sputum cell for early lung cancer detection
abstract
Lung cancer has been the largest cause of cancer deaths worldwide with an overall 5-year survival rate of only 15%. Its early detection significantly increases the chances of an effective treatment. For this purpose, a computer-aided design system using images of sputum stained smears is a practical, low-cost, and totally non invasive solution. In this paper, we present a framework for the detection and segmentation of sputum cells in sputum images using respectively, a Bayesian classification and mean shift segmentation. Our methods are validated and compared with other competitive approaches via a series of experiments conducted with a data set of 88 images.
Naoufel Werghi, Christian Donner, Fatma Taher
ICIP1
2011 A novel surface pattern for 3D facial surface encoding and alignment
abstract
In this paper, a novel topological structured pattern, dubbed a spiral facet, is proposed and applied for the encoding and alignment of 3D facial triangular mesh surfaces. After describing the foundation of this representation, we present two direct face alignment methods, that exploit the ordered structure of the spiral facet in addressing the correspondence problem. We study the behavior and the performance of these two methods through experimentation conducted with a subset of a 3D face database. We also showcase the potential of our alignment framework for face recognition.
Naoufel Werghi, Mohamed K. Naqbi
SMC1
2010 Combined spatial and transform domain analysis for rectangle detection
Harish Bhaskar, Naoufel Werghi, Saeed Al-Mansoori
FUSION2
2010 The 3D facial kernel: Application to facial surface spherical mapping and alignment
abstract
This paper presents a framework for the computation of a novel 3D facial attribute, namely,the face kernel. The first part of the paper exposes the theoretical background related to the kernel computation and demonstrates some properties of the surface kernel. The second part describes two applications of the kernel concept, namely, spherical facial surface mapping and facial surface alignment. This framework has been tested and experimented on real 3D face surfaces. The usefulness of the face kernel is manifested in two aspects: 1) It provides a simple scheme for an optimal spherical embedding in the sense of minimizing the related distortion error. 2) It can be used for an efficient facial surface alignment.
Naoufel Werghi
SMC1
2010 An unsupervised learning approach based on a Hopfield-like network for assessing posterior capsule opacification
Naoufel Werghi, Rachid Sammouda, Fatma AlKirbi
Pattern Anal. Appl.1
2007 Segmentation and Modeling of Full Human Body Shape From 3-D Scan Data: A Survey
abstract
The recent advances in full human body (HB) imaging technology illustrated by the 3D human body scanner (HBS), a device delivering full HB shape data, opened up large perspectives for the deployment of this technology in various fields such as the clothing industry, anthropology, and entertainment. However, these advances also brought challenges on how to process and interpret the data delivered by the HBS in order to bridge the gap between this technology and potential applications. This paper presents a literature survey of research work on HBS data segmentation and modeling aiming at overcoming these challenges, and discusses and evaluates different approaches with respect to several requirements.
Naoufel Werghi
IEEE Trans. Syst. Man Cybern. Part C1
2006 A robust approach for constructing a graph representation of articulated and tubular-like objects from 3D scattered data
Naoufel Werghi
Pattern Recognit. Lett.1
2006 A functional-based segmentation of human body scans in arbitrary postures
abstract
This paper presents a general framework that aims to address the task of segmenting three-dimensional (3-D) scan data representing the human form into subsets which correspond to functional human body parts. Such a task is challenging due to the articulated and deformable nature of the human body. A salient feature of this framework is that it is able to cope with various body postures and is in addition robust to noise, holes, irregular sampling and rigid transformations. Although whole human body scanners are now capable of routinely capturing the shape of the whole body in machine readable format, they have not yet realized their potential to provide automatic extraction of key body measurements. Automated production of anthropometric databases is a prerequisite to satisfying the needs of certain industrial sectors (e.g., the clothing industry). This implies that in order to extract specific measurements of interest, whole body 3-D scan data must be segmented by machine into subsets corresponding to functional human body parts. However, previously reported attempts at automating the segmentation process suffer from various limitations, such as being restricted to a standard specific posture and being vulnerable to scan data artifacts. Our human body segmentation algorithm advances the state of the art to overcome the above limitations and we present experimental results obtained using both real and synthetic data that confirm the validity, effectiveness, and robustness of our approach.
Naoufel Werghi, Yijun Xiao, J. Paul Siebert
IEEE Trans. Syst. Man Cybern. Part B1
2005 A discriminative 3D wavelet-based descriptors: Application to the recognition of human body postures
Naoufel Werghi
Pattern Recognit. Lett.1
1999 Construction of Articulated Models from Range Data
abstract
In this paper we present an algorithm for automatically building models of articulated objects from range data. These models not only describe the surface shape of the object but also describe the kinematics that constrain the movement of one object component in relation to another. This is more difficult than building models of rigid objects because the association of surface measurements to object components must be determined. The algorithm is demonstrated on a difficult object with free-form surfaces. 1 Introduction The ability to automatically acquire geometric models from example objects is useful in a growing number of application areas. In the field of computer graphics, the need for improvements in realism requires more complex models, but manual model construction is time-consuming and difficult. Users of Computer-Aided Design technology would like to be able to make improvements to a manufactured part and then update their CAD model to reflect this. This provides a very eff...
Anthony Ashbrook, Robert B. Fisher, Naoufel Werghi, Craig Robertson
BMVC3
1999 Improving Second-order Surface Estimation
abstract
The paper proposes a reliable method for estimating second-order surfaces from 3D range data in the framework of object recognition and localization or object modelling. Instead of estimating such surface individually the approach ts all the surfaces captured in the scene together, taking into account the geometric relationships between them and their speci c characteristics. The technique is compared with other methods through experiments performed on real objects and demonstrates that the use of constrained relationships improves shape estimates.
Naoufel Werghi, Robert B. Fisher, Anthony Ashbrook, Craig Robertson
BMVC1
1999 Object reconstruction by incorporating geometric constraints in reverse engineering
Naoufel Werghi, Robert B. Fisher, Craig Robertson, Anthony Ashbrook
Comput. Aided Des.1
1998 Finding Surface Correspondance for Object Recognition and Registration Using Pairwise Geometric Histograms
Anthony Ashbrook, Robert B. Fisher, Craig Robertson, Naoufel Werghi
ECCV (2)4
1998 Modelling Objects having Quadric Surfaces Incorporating Geometric cCnstraints
Naoufel Werghi, Robert B. Fisher, Craig Robertson, Anthony Ashbrook
ECCV (2)1
1997 Segmentation of Range Data into Rigid Subsets using Planar Surface Patches
Anthony Ashbrook, Robert B. Fisher, Craig Robertson, Naoufel Werghi
BMVC4
1997 Improving model shape acquisition by incorporating geometric constraints
Naoufel Werghi, Robert B. Fisher, Anthony Ashbrook, Craig Robertson
BMVC1
1996 Ellipse Fitting And Three-Dimensional Localization Of Objects Based On Elliptic Features
Naoufel Werghi, Christophe Doignon, Gabriel Abba
ICIP (1)1