Sajid Javed

dblp:143/2961 · DBLP profile ↗
← Back
58ranked-venue papers
17as first author
42since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 10 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Underwater visual tracking with a large scale dataset and image enhancement
abstract
This paper presents a new dataset and general tracker enhancement method for Underwater Visual Object Tracking (UVOT). Despite its significance, underwater tracking has remained unexplored due to data inaccessibility. It poses distinct challenges; the underwater environment exhibits non-uniform lighting conditions, low visibility, lack of sharpness, low contrast, camouflage, and reflections from suspended particles. Performance of traditional tracking methods designed primarily for terrestrial or open-air scenarios drops in such conditions. We address the problem by proposing a novel underwater image enhancement algorithm designed specifically to boost tracking quality. The method has resulted in a significant performance improvement, of up to 5.0% AUC, of state-of-the-art (SOTA) visual trackers. To develop robust and accurate UVOT methods, large-scale datasets are required. To this end, we introduce a large-scale UVOT benchmark dataset consisting of 400 video segments and 275,000 manually annotated frames enabling underwater training and evaluation of deep trackers. The videos are labelled with several underwater-specific tracking attributes including watercolor variation, target distractors, camouflage, target relative size, and low visibility conditions. The UVOT400 dataset, tracking results, and the code are publicly available on: https://github.com/BasitAlawode/UVOT400 .
Basit Alawode, Sajid Javed, Fayaz Ali Dharejo, Mehnaz Ummar, Arif Mahmoud, Fahad Shahbaz Khan, Jiri Matas
Neurocomputing2
2026 Hierarchical vision-language model with comprehensive language description for video anomaly detection
abstract
Video Anomaly Detection (VAD) is a crucial task in computer vision, with applications in surveillance, transportation, and industrial monitoring. Recent advancements in Vision-Language Models (VLMs) have shown promising direction toward VAD in weakly supervised and unsupervised settings by leveraging visual and textual modalities. However, existing VLM-based methods often overlook coarse-to-fine temporal information, limiting their ability to handle complex anomalies. To address this issue, we propose a hierarchical VLM that enhances visual-textual feature representation by capturing video content at multiple levels of abstraction. Our algorithm generates a hierarchical view of the video, dividing it into short and long views. We extract hierarchical visual features and construct a bag containing comprehensive textual descriptions of anomalies using existing VLMs without relying on ground truth data. Our model fuses these modalities and is fine-tuned for anomaly score prediction in weakly supervised, unsupervised, and one-class settings. We also introduce a training-free VAD framework based on similarity scores. By aligning complex concepts across hierarchical views, our model captures both fine-grained details and high-level contextual information, leading to robust feature representations. Extensive experiments on UCF-Crime, ShanghaiTech, and XD-Violence datasets demonstrate the superior performance of our method compared to State-Of-The-Art VAD methods. The code is available on GitHub at: https://github.com/MR81224/HVLMCLD-VAD .
Muaz Al Radi, Sajid Javed
Knowl. Based Syst.2
2026 Unsupervised Background Subtraction Using Generator-Discriminator Learning
abstract
Background subtraction is a core problem in computer vision, widely used in video surveillance to segment moving foreground objects from video sequences. While deep learning approaches have shown strong performance—especially under dynamic backgrounds and sudden illumination changes—they typically rely on large-scale, high-quality labeled video datasets. Acquiring such data is time-consuming and expensive, making existing supervised or weakly supervised methods less suitable for real-time applications. Moreover, many of these methods suffer from performance degradation when applied to unseen video sequences. To address these challenges, we present UTGMP-BS algorithm: an Unsupervised Transformer-based pseudo-label Generator with a Message-Passing network for the Background Subtraction task. UTGMP-BS is a fully unsupervised framework designed to learn directly from unlabeled video sequences. It comprises two key components: a transformer-based pseudo-label generator, which produces initial pixel-level foreground and background labels using an encoder-decoder architecture and anL1loss, and a message-passing network, which acts as a label-cleaner discriminator to refine the pseudo labels and enforce spatial consistency. These two branches engage in mutual learning through consecutive iterations, enhancing one another’s performance without any ground-truth supervision. The framework is trained using an alternating iterative learning strategy with binary cross-entropy loss, achieving robust background subtraction across varied scenes. Extensive experiments on six publicly available benchmark datasets demonstrate that UTGMP-BS achieves competitive results compared to existing State-of-The-Art (SOTA) methods.
Basit Alawode, Sajid Javed
IEEE Trans. Circuits Syst. Video Technol.2
2026 AquaticCLIP: A Vision-Language Foundation Model and Dataset for Underwater Scene Analysis
abstract
The preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role in aiding marine scientists in their decision-making processes. In this article, we introduce AquaticCLIP, a novel contrastive language-image pretraining (CLIP) model tailored for aquatic scene understanding. AquaticCLIP presents an underwater domain-specific learning framework that aligns images and texts in aquatic environments, enabling tasks such as segmentation, classification, detection, and object counting. By leveraging our large-scale underwater image-text paired dataset without the need for ground-truth (GT) annotations, our model enriches existing vision-language models (VLMs) in the aquatic domain. For this purpose, we construct a 2-million underwater image-text paired dataset using heterogeneous resources, including YouTube, Netflix, National Geographic (NatGeo), etc. To fine-tune AquaticCLIP, we propose a prompt-guided vision encoder (PGVE) that progressively aggregates patch features via learnable prompts, while a vision-guided mechanism enhances the language encoder by incorporating visual context. The model is optimized through a contrastive pretraining loss to align visual and textual modalities. AquaticCLIP achieves notable performance improvements in zero-shot settings across multiple underwater computer vision tasks, outperforming existing methods in both accuracy and robustness. Our model sets a new benchmark for vision-language applications in underwater environments. The code and dataset for AquaticCLIP are publicly available on GitHub at: https://github.com/BasitAlawode/AquaticCLIP.
Basit Alawode, Iyyakutti Iyappan Ganapathi, Sajid Javed, Mohammed Bennamoun, Arif Mahmood
IEEE Trans. Neural Networks Learn. Syst.3
2025 Structured Comprehensive Textual Representations for Medical Vision-Language Pretraining
abstract
Medical vision-language learning faces a persistent challenge: limited paired image-text data, especially in expert domains like radiology. To address this, we propose a data-efficient framework that enhances dual-encoder vision-language models through structured comprehensive textual representations. Using GPT-4 and curated clinical knowledge, we generate semantically comprehensive representations for each disease class. These representations capture detailed visual descriptions of the disease, the major causes, and the major symptoms related to chest X-rays. These representations are paired with medical images to pre-train BiomedCLIP using contrastive learning, effectively aligning vision and language modalities without the need for large-scale manual annotations. We then fine-tune the vision encoder for chest X-ray disease classification. Our results show that this approach achieves competitive performance, even with limited supervision. This work highlights the potential of structured generative supervision to scale vision-language learning in data-constrained medical domains.
Youssef Ibrahim, Anabia Sohail, Sajid Javed, Hasan Al Marzouqi, Naoufel Werghi
AICCSA3
2025 Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
abstract
In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tissue phenotyping, and survival analysis due to the limited level of detail that a single-resolution image can provide. Addressing this, we propose a novel multi-resolution paradigm leveraging Whole Slide Images (WSIs) to extract histology patches at multiple resolutions and generate corresponding textual descriptions through advanced CPath VLM. We introduce visual-textual alignment at multiple resolutions as well as cross-resolution alignment to establish more effective text-guided visual representations. Cross-resolution alignment using a multi-modal encoder enhances the model’s ability to capture context from multiple resolutions in histology images. Our model aims to capture a broader range of information, supported by novel loss functions, enriches feature representation, improves discriminative ability, and enhances generalization across different resolutions. Pre-trained on a comprehensive TCGA dataset with 34 million image-language pairs at various resolutions, our fine-tuned model outperforms State-Of-The-Art (SOTA) counterparts across multiple datasets and tasks, demonstrating its effectiveness in CPath. The code is available on GitHub at: https://github.com/BasitAlawode/MR-PLIP.
Shahad Albastaki, Anabia Sohail, Iyyakutti Iyappan Ganapathi, Basit Alawode, Asim Khan, Sajid Javed, Naoufel Werghi, Mohammed Bennamoun, Arif Mahmood
CVPR6
2025 Enhancing Medical Vision-Language Models with Rich Textual Descriptions and Multiple Alignments for Chest X-Ray Diagnosis
abstract
Vision-Language models (VLMs) integrate natural language understanding with visual data interpretation, crucial in diverse applications such as medical imaging. However, training VLMs on limited data, especially in radiology, remains a challenge. We propose a strategy to improve dual encoder performance under data constraints. Using contrastive learning to align visual and textual embeddings effectively, we generated a bag of rich textual descriptions using GPT-4 to augment merged information from esteemed medical resources and pre-trained BiomedCLIP. These rich textual descriptions provide in-depth information on disease visual description, major causes, and major symptoms, enhancing the model’s contextual understanding and classification accuracy. Unlike previous methods relying on a single alignment, our multiple alignment strategy associates multiple images with multiple textual descriptions per disease class while capping descriptors to maintain computational efficiency. Adapting the vision encoder for chest X-ray classification, our approach achieves competitive accuracy with fewer training pairs, highlighting its potential for data-limited domains.
Youssef Ibrahim, Anabia Sohail, Sajid Javed, Hasan Almarzouqi, Mohamed Deriche 0001, Naoufel Werghi
ICIP3
2025 PMIL: A Topology Module to Improve MIL-based WSI Classification
abstract
Deep learning models have achieved remarkable success in pathology image analysis. However, they still face challenges in effectively modeling fine-grained, object-level features. Topological Data Analysis (TDA) has shown promise for addressing these issues but remains underexplored, particularly for whole-slide pathology applications. Additionally, the effectiveness of TDA has yet to be firmly established, as current studies largely use small-scale datasets. In this work, we address these gaps by introducing Persistent Homology in Multiple Instance Learning (PMIL), the first adaptable TDA-based module within the MIL framework. We validate our approach on a large-scale classification dataset, benchmarking against multiple state-of-the-art methods.
Ahmad Obeid 0001, Anabia Sohail, Said Boumaraf, Xiabi Liu, Sajid Javed, Hasan Almarzouqi, Jorge Dias 0001, Mohammed Bennamoun, Naoufel Werghi, Ibrahim M. Elfadel
ISCAS5
2025 EfficientFaceV2S: A lightweight model and a benchmarking approach for drone-captured face recognition
Mohamad Alansari, Khaled Alnuaimi, Iyyakutti Iyappan Ganapathi, Sara Alansari, Sajid Javed, Abdulhadi Shoufan, Yahya Zweiri, Naoufel Werghi
Expert Syst. Appl.5
2025 Video anomaly detection in 10 years: a survey and outlook
Moshira Abdalla, Sajid Javed, Muaz Al Radi, Anwaar Ulhaq, Naoufel Werghi
Neural Comput. Appl.2
2025 Neuromorphic Vision-Based Motion Segmentation With Graph Transformer Neural Network
abstract
Moving object segmentation is critical to interpret scene dynamics for robotic navigation systems in challenging environments. Neuromorphic vision sensors are tailored for motion perception due to their asynchronous nature, high temporal resolution, and reduced power consumption. However, their unconventional output requires novel perception paradigms to leverage their spatially sparse and temporally dense nature. In this work, we propose a novel event-based motion segmentation algorithm using a Graph Transformer Neural Network, dubbed GTNN. Our proposed algorithm processes event streams as 3D graphs by a series of nonlinear transformations to unveil local and global spatiotemporal correlations between events. Based on these correlations, events belonging to moving objects are segmented from the background without prior knowledge of the dynamic scene geometry. The algorithm is trained on publicly available datasets including MOD, EV-IMO, and EV-IMO2 using the proposed training scheme to facilitate efficient training on extensive datasets. Moreover, we introduce the Dynamic Object Mask-aware Event Labeling (DOMEL) approach for generating approximate ground-truth labels for event-based motion segmentation datasets. We use DOMEL to label our own recorded Event dataset for Motion Segmentation (EMS-DOMEL), which we release to the public for further research and benchmarking. Rigorous experiments are conducted on several unseen publicly-available datasets where the results revealed that GTNN outperforms state-of-the-art methods in the presence of dynamic background variations, motion patterns, and multiple dynamic objects with varying sizes and velocities. GTNN achieves significant performance gains with an average increase of 9.4% and 4.5% in terms of motion segmentation accuracy (IoU%) and detection rate (DR%), respectively.
Yusra Alkendi, Rana Azzam, Sajid Javed, Lakmal D. Seneviratne, Yahya Zweiri
IEEE Trans. Multim.3
2025 Learning Spatial-Temporal Regularized Tensor Sparse RPCA for Background Subtraction
abstract
Background subtraction in videos is a core challenge in computer vision, aiming to accurately identify moving objects. Robust principal component analysis (RPCA) has emerged as a promising unsupervised (US) paradigm for this task, showing strong performance on various benchmark datasets. Building on RPCA, tensor RPCA (TRPCA) variants have further enhanced background subtraction performance. However, current TRPCA methods often treat moving object pixels independently, lacking spatial-temporal structured-sparsity constraints. This limitation leads to performance degradation in scenarios with dynamic backgrounds, camouflage, and camera jitter. In this work, we introduce a novel spatial-temporal regularized tensor sparse RPCA algorithm to address these issues. By incorporating normalized graph-Laplacian matrices into the sparse component, we enforce spatial-temporal regularization. We construct two graphs-one across spatial locations and another across temporal slices-to guide regularization. By maximizing our objective function, we ensure that the tensor sparse component aligns with the spatiotemporal eigenvectors of the graph-Laplacian matrices, preserving disconnected moving object pixels. We formulate a new objective function and employ batch and online-based optimization methods to jointly optimize background-foreground separation and spatial-temporal regularization. Experimental evaluation on six publicly available datasets demonstrates the superior performance of our algorithm compared to existing methods.
Basit Alawode, Sajid Javed
IEEE Trans. Neural Networks Learn. Syst.2
2025 Unsupervised Dual Transformer Learning for 3-D Textured Surface Segmentation
abstract
Analysis of the 3-D texture is indispensable for various tasks, such as retrieval, segmentation, classification, and inspection of sculptures, knit fabrics, and biological tissues. A 3-D texture represents a locally repeated surface variation (SV) that is independent of the overall shape of the surface and can be determined using the local neighborhood and its characteristics. Existing methods mostly employ computer vision techniques that analyze a 3-D mesh globally, derive features, and then utilize them for classification or retrieval tasks. While several traditional and learning-based methods have been proposed in the literature, only a few have addressed 3-D texture analysis, and none have considered unsupervised schemes so far. This article proposes an original framework for the unsupervised segmentation of 3-D texture on the mesh manifold. The problem is approached as a binary surface segmentation task, where the mesh surface is partitioned into textured and nontextured regions without prior annotation. The proposed method comprises a mutual transformer-based system consisting of a label generator (LG) and a label cleaner (LC). Both models take geometric image representations of the surface mesh facets and label them as texture or nontexture using an iterative mutual learning scheme. Extensive experiments on three publicly available datasets with diverse texture patterns demonstrate that the proposed framework outperforms standard and state-of-the-art unsupervised techniques and performs reasonably well compared to supervised methods.
Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Sajid Javed, Syed Sadaf Ali, Naoufel Werghi
IEEE Trans. Neural Networks Learn. Syst.3
2024 3D-TexSeg: Unsupervised Segmentation of 3D Texture Using Mutual Transformer Learning
abstract
Analysis of the 3D Texture is indispensable for various tasks, such as retrieval, segmentation, classification, and inspection of sculptures, knitted fabrics, and biological tissues. A 3D texture is a locally repeated surface variation independent of the surface’s overall shape and can be determined using the local neighborhood and its characteristics. Existing techniques typically employ computer vision techniques that analyze a 3D mesh globally, derive features, and then utilize the obtained features for retrieval or classification. Several traditional and learning-based methods exist in the literature; however, only a few are on 3D texture, and nothing yet, to the best of our knowledge, on the unsupervised schemes. This paper presents an original framework for the unsupervised segmentation of the 3D texture on the mesh manifold. We approach this problem as binary surface segmentation, partitioning the mesh surface into textured and non-textured regions without prior annotation. We devise a mutual transformer-based system comprising a label generator and a cleaner. The two models take geometric image representations of the surface mesh facets and label them as texture or non-texture across an iterative mutual learning scheme. Extensive experiments on three publicly available datasets with diverse texture patterns demonstrate that the proposed framework outperforms standard and SOTA unsupervised techniques and competes reasonably with supervised methods.
Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Sajid Javed, Syed Sadaf Ali, Naoufel Werghi
3DV3
2024 Text-Guided Multi-Modal Fusion for Underwater Visual Tracking
abstract
The integration of Natural Language (NL) descriptions with contemporary tracking algorithms constitutes a new and dynamic field, exhibiting no indications of deceleration in the near future. Nevertheless, the absence of comprehensive language descriptions for tracking datasets, particularly in the domain of underwater tracking datasets, presents a substantial impediment to the advancement of this field. Typically, the textual descriptions accompanying these datasets are brief, inadequately informative, lack details regarding relative location or directional movement, and occasionally deviate from the manner in which a human would naturally describe the target in ordinary conversation. In response to this challenge, we propose the development of vividly descriptive NL descriptions tailored for the UVOT400 dataset, which focuses on underwater tracking. These descriptions aim to encapsulate a myriad of factors in order to furnish as comprehensive an understanding as possible regarding the target fish. Subsequent evaluations of these descriptions, conducted in conjunction with contemporary language-based tracking systems, have revealed superior performance in comparison to the best-performing visual-only trackers employed for benchmarking purposes with the aforementioned dataset.
Yonathan Michael, Mohamad Alansari, Sajid Javed
AVSS3
2024 CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment
abstract
This paper proposes Comprehensive Pathology Language Image Pretraining (CPLIP), a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision-language models by leveraging extensive data without needing ground truth annotations. CPLIP involves constructing a pathology-specific dictionary, generating textual descriptions for images using language models, and retrieving relevant images for each text snippet via a pretrained model. The model is then fine-tuned using a many-to-many contrastive learning method to align complex interrelated concepts across both modalities. Evaluated across multiple histopathology tasks, CPLIP shows notable improvements in zero-shot learning scenarios, outperforming existing methods in both interpretability and robustness and setting a higher benchmark for the application of vision-language models in the field. To encourage further research and replication, the code for CPLIP is available on GitHub at https://cplip.github.io/
Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, Mohammed Bennamoun
CVPR1
2024 Integrating Vision-Language Supervision for Uniform Appearance Tracking
abstract
Integrating detailed Natural Language (NL) descriptions with modern tracking technologies represents a significant and emerging field within Uniform Appearance (UA) crowd-tracking research, demonstrating substantial potential for future developments. A prominent challenge in this area is the lack of NL descriptions tailored for UA crowd tracking datasets. Existing datasets for Drone-Person Tracking in Uniform Appearance Crowd (D-PTUAC) lack essential textual annotations. Our study aims to bridge this gap by innovatively introducing comprehensive natural language descriptions for the D-PTUAC dataset, specifically designed for Uniform Appearance crowd tracking using drones. This enhancement aims to provide a richer understanding of the dataset and facilitate more effective utilization in research and applications related to drone-based crowd tracking. These descriptions are meticulously designed to include extensive information about the target entities, thereby significantly augmenting the dataset’s depth and applicability. Our evaluations utilizing the latest state-of-the-art (SOTA) NL-based tracking algorithms showed us a remarkable competitive performance in tracking when juxtaposed against SOTA visual trackers benchmarked on the D-PTUAC dataset. This outcome highlights the critical role and efficacy of integrated language descriptions in enhancing the methodologies employed in UA crowd tracking.
Mohamad Alansari, Ahmed Abughali, Obadah Habash, Khaled Alnuaimi, Sajid Javed, Naoufel Werghi
ICIP5
2024 SMO-CLIP: Enhancing Anomalous Smoke Density Assessment Using A Hybrid LLM-VLM Approach
abstract
Flare stacks are among the crucial components in the safety and emission control of petrochemical plants. However, due to the imperceptibility of smoke and contaminants, analyzing these released particles during flare stack operation is one of the top challenges. To stress the problem, our work presents a novel solution called SMO-CLIP that can hybridize knowledge from Vision-Language Models (VLMs), specifically the Contrastive Language Image Pretraining (CLIP) model, with extra insights derived from GPT-4 Large Language Model (LLM). Furthermore, two new tasks, Finegrained Smoke Density Recognition (FSDR) and Coarsegrained Smoke Density Recognition (CSDR) are investigated in this paper to accurately detect and evaluate varying smoke intensities. Notable advancements over current approaches are observed through extensive experiments, demonstrating the superior performance of the proposed approach against state-of-the-art models.
Muaz Al Radi, Mahmoud Said Elmezain, Abdelfatah Hassan Ahmed, Abderrahmene Boudiaf, Said Boumaraf, Jorge Dias 0001, Hamad Karki, Sajid Javed, Khalid Yousef Al Awadhi, Naoufel Werghi
ICIP9
2024 Multi-scale feature reconstruction network for industrial anomaly detection
abstract
Unsupervised anomaly detection techniques, which operate without prior knowledge of anomalies, have garnered significant attention in industrial inspection due to their adaptability and generalization. Therefore, knowledge-based computer vision techniques have been broadly applied to identify unusual image patterns. However, real-time industrial applications present challenges such as limited anomalous samples, inadequate defect knowledge, and complex background textures. These factors lead to difficulties in accurately identifying defect regions, and conventional auto-encoder networks often struggle to overcome these issues.To address these limitations, we propose a multi-scale feature reconstruction (MSFR) network specifically designed for domain shift scenarios. Our approach employs a pyramidal vision transformer network (PVTN) to reconstruct multi-scale feature maps, capturing discriminative features at various scales. Additionally, a pre-trained module extracts multi-level features at the same scale, and a dedicated feature matching module enhances accuracy by improving the alignment probability between features. The MSFR strategy surpasses conventional auto-encoders by filtering pixel-level information at multiple depths. Empirical evaluations were conducted using benchmark datasets such as MVTec AD and AeBAD-S. Furthermore, an extensive ablation study demonstrates the effectiveness and viability of the proposed MSFR approach for industrial anomaly detection tasks. The experimental results show that the proposed model significantly outperforms recent approaches, making it highly suitable for real-world industrial applications, particularly in manufacturing. • MSFR is introduced for industrial anomaly detection to handle domain shifts generalization. • A multi-scale features network is developed to efficiently detect defects. • A pyramidal network is proposed to cover the limitations of traditional autoencoders. • Comprehensive experiments are generated over AeBAD-S and MVTec IAD benchmark datasets.
Ehtesham Iqbal 0002, Samee Ullah Khan, Sajid Javed, Brain Moyo, Yahya Zweiri, Yusra Abdulrahman
Knowl. Based Syst.3
2024 Unsupervised mutual transformer learning for multi-gigapixel Whole Slide Image classification
Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi, Nasir M. Rajpoot
Medical Image Anal.1
2024 Spatial deep feature augmentation technique for FER using genetic algorithm
Nudrat Nida, Muhammad Haroon Yousaf, Aun Irtaza, Sajid Javed, Sergio A. Velastin
Neural Comput. Appl.4
2024 Neural Graph Refinement for Robust Recognition of Nuclei Communities in Histopathological Landscape
abstract
Accurate classification of nuclei communities is an important step towards timely treating the cancer spread. Graph theory provides an elegant way to represent and analyze nuclei communities within the histopathological landscape in order to perform tissue phenotyping and tumor profiling tasks. Many researchers have worked on recognizing nuclei regions within the histology images in order to grade cancerous progression. However, due to the high structural similarities between nuclei communities, defining a model that can accurately differentiate between nuclei pathological patterns still needs to be solved. To surmount this challenge, we present a novel approach, dubbed neural graph refinement, that enhances the capabilities of existing models to perform nuclei recognition tasks by employing graph representational learning and broadcasting processes. Based on the physical interaction of the nuclei, we first construct a fully connected graph in which nodes represent nuclei and adjacent nodes are connected to each other via an undirected edge. For each edge and node pair, appearance and geometric features are computed and are then utilized for generating the neural graph embeddings. These embeddings are used for diffusing contextual information to the neighboring nodes, all along a path traversing the whole graph to infer global information over an entire nuclei network and predict pathologically meaningful communities. Through rigorous evaluation of the proposed scheme across four public datasets, we showcase that learning such communities through neural graph refinement produces better results that outperform state-of-the-art methods.
Taimur Hassan, Zhu Li 0001, Sajid Javed, Jorge Dias 0001, Naoufel Werghi
IEEE Trans. Image Process.3
2024 Center-Focused Affinity Loss for Class Imbalance Histology Image Classification
abstract
Early-stage cancer diagnosis potentially improves the chances of survival for many cancer patients worldwide. Manual examination of Whole Slide Images (WSIs) is a time-consuming task for analyzing tumor-microenvironment. To overcome this limitation, the conjunction of deep learning with computational pathology has been proposed to assist pathologists in efficiently prognosing the cancerous spread. Nevertheless, the existing deep learning methods are ill-equipped to handle fine-grained histopathology datasets. This is because these models are constrained via conventional softmax loss function, which cannot expose them to learn distinct representational embeddings of the similarly textured WSIs containing an imbalanced data distribution. To address this problem, we propose a novel center-focused affinity loss (CFAL) function that exhibits 1) constructing uniformly distributed class prototypes in the feature space, 2) penalizing difficult samples, 3) minimizing intra-class variations, and 4) placing greater emphasis on learning minority class features. We evaluated the performance of the proposed CFAL loss function on two publicly available breast and colon cancer datasets having varying levels of imbalanced classes. The proposed CFAL function shows better discrimination abilities as compared to the popular loss functions such as ArcFace, CosFace, and Focal loss. Moreover, it outperforms several SOTA methods for histology image classification across both datasets.
Taslim Mahbub, Ahmad Obeid 0001, Sajid Javed, Jorge Dias 0001, Taimur Hassan, Naoufel Werghi
IEEE J. Biomed. Health Informatics3
2024 Neuromorphic Camera Denoising Using Graph Neural Network-Driven Transformers
abstract
Neuromorphic vision is a bio-inspired technology that has triggered a paradigm shift in the computer vision community and is serving as a key enabler for a wide range of applications. This technology has offered significant advantages, including reduced power consumption, reduced processing needs, and communication speedups. However, neuromorphic cameras suffer from significant amounts of measurement noise. This noise deteriorates the performance of neuromorphic event-based perception and navigation algorithms. In this article, we propose a novel noise filtration algorithm to eliminate events that do not represent real log-intensity variations in the observed scene. We employ a graph neural network (GNN)-driven transformer algorithm, called GNN-Transformer, to classify every active event pixel in the raw stream into real log-intensity variation or noise. Within the GNN, a message-passing framework, referred to as EventConv, is carried out to reflect the spatiotemporal correlation among the events while preserving their asynchronous nature. We also introduce the known-object ground-truth labeling (KoGTL) approach for generating approximate ground-truth labels of event streams under various illumination conditions. KoGTL is used to generate labeled datasets, from experiments recorded in challenging lighting conditions, including moon light. These datasets are used to train and extensively test our proposed algorithm. When tested on unseen datasets, the proposed algorithm outperforms state-of-the-art methods by at least 8.8% in terms of filtration accuracy. Additional tests are also conducted on publicly available datasets (ETH Zürich Color-DAVIS346 datasets) to demonstrate the generalization capabilities of the proposed algorithm in the presence of illumination variations and different motion dynamics. Compared to state-of-the-art solutions, qualitative results verified the superior capability of the proposed algorithm to eliminate noise while preserving meaningful events in the scene.
Yusra Alkendi, Rana Azzam, Abdulla Ayyad, Sajid Javed, Lakmal D. Seneviratne, Yahya Zweiri
IEEE Trans. Neural Networks Learn. Syst.4
2023 Higher-Order Sparse Convolutions in Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have been applied to many problems in computer sciences. Capturing higher-order relationships between nodes is crucial to increase the expressive power of GNNs. However, existing methods to capture these relationships could be infeasible for large-scale graphs. In this work, we introduce a new higher-order sparse convolution based on the Sobolev norm of graph signals. Our Sparse Sobolev GNN (S-SobGNN) computes a cascade of filters on each layer with increasing Hadamard powers to get a more diverse set of functions, and then a linear combination layer weights the embeddings of each filter. We evaluate S-SobGNN in several applications of semi-supervised learning. S-SobGNN shows competitive performance in all applications as compared to several state-of-the-art methods.
Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Arif Mahmood, Fragkiskos D. Malliaros, Thierry Bouwmans
ICASSP2
2023 The Unconstrained Ear Recognition Challenge 2023: Maximizing Performance and Minimizing Bias
abstract
The paper provides a summary of the 2023 Unconstrained Ear Recognition Challenge (UERC), a benchmarking effort focused on ear recognition from images acquired in uncontrolled environments. The objective of the challenge was to evaluate the effectiveness of current ear recognition techniques on a challenging ear dataset while analyzing the techniques from two distinct aspects, i.e., verification performance and bias with respect to specific demographic factors, i.e., gender and ethnicity. Seven research groups participated in the challenge and submitted a seven distinct recognition approaches that ranged from descriptor-based methods and deep-learning models to ensemble techniques that relied on multiple data representations to maximize performance and minimize bias. A comprehensive investigation into the performance of the submitted models is presented, as well as an in-depth analysis of bias and associated performance differentials due to differences in gender and ethnicity. The results of the challenge suggest that a wide variety of models (e.g., transformers, convolutional neural networks, ensemble models) is capable of achieving competitive recognition results, but also that all of the models still exhibit considerable performance differentials with respect to both gender and ethnicity. To promote further development of unbiased and effective ear recognition models, the starter kit of UERC 2023 together with the baseline model, and training and test data is made available from: http://ears.fri.uni-lj.si/
Ziga Emersic, Tetsushi Ohki, Muku Akasaka, Takahiko Arakawa, Soshi Maeda, Masora Okano, Yuya Sato, Anjith George, Sébastien Marcel, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Sajid Javed, Naoufel Werghi, S. G. Isik, Erdi Saritas, Hazim Kemal Ekenel, V. Hudovernik, Jan Niklas Kolf, Fadi Boutros, Naser Damer, G. Sharma, Aman Kamboj, Aditya Nigam, Deepak Kumar Jain 0001, G. Cámara-Chávez, Peter Peer, Vitomir Struc
IJCB12
2023 EFaR 2023: Efficient Face Recognition Competition
abstract
This paper presents the summary of the Efficient Face Recognition Competition (EFaR) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition received 17 submissions from 6 different teams. To drive further development of efficient face recognition models, the submitted solutions are ranked based on a weighted score of the achieved verification accuracies on a diverse set of benchmarks, as well as the deployability given by the number of floating-point operations and model size. The evaluation of submissions is extended to bias, cross-quality, and large-scale recognition benchmarks. Overall, the paper gives an overview of the achieved performance values of the submitted solutions as well as a diverse set of baselines. The submitted solutions use small, efficient network architectures to reduce the computational cost, some solutions apply model quantization. An outlook on possible techniques that are underrepresented in current solutions is given as well.
Jan Niklas Kolf, Fadi Boutros, Jurek Elliesen, Markus Theuerkauf, Naser Damer, Mohamad Alansari, Oussama Abdul Hay, Sara Alansari, Sajid Javed, Naoufel Werghi, Klemen Grm, Vitomir Struc, Fernando Alonso-Fernandez, Kevin Hernandez-Diaz, Josef Bigün, Anjith George, Christophe Ecabert, Hatef Otroshi-Shahreza, Ketan Kotwal, Sébastien Marcel, Iurii Medvedev, Bo Jin 0018, Diogo Nunes, Ahmad Hassanpour, Pankaj Khatiwada, Aafan Ahmad Toor, Bian Yang
IJCB9
2023 Facet-Level Segmentation of 3d Textures on Cultural Heritage Objects
abstract
Textures in 3D meshes exhibit intrinsic surface variations and are indispensable for various applications, such as retrieval, segmentation, and classification of sculptures, artifacts, and paintings. A 3D texture pattern is a locally repeated surface variation independent of the overall surface geometry and can be determined using the local neighborhood and its characteristics. Texture analysis typically employs computer vision techniques that analyze the entire 3D mesh, derive hand-crafted features, and then utilize the derived features for retrieval or classification. Several traditional and learning-based techniques exist in the literature on surface variations; however, textures are the subject of limited works. We propose a binary classification framework at the facet level for classifying texture and non-texture regions on 3D surfaces. An image sequence is generated at each facet, which serves as input to a deep vision transformer. To generate images at each facet, we construct a grid where each cell is filled with the geometric properties of its neighboring facets. We evaluated the proposed method using two datasets with diverse texture patterns, and the results are encouraging.
Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Muhammad Owais, Neha Gour, Sajid Javed, Naoufel Werghi
ICIP5
2023 DFR-FastMOT: Detection Failure Resistant Tracker for Fast Multi-Object Tracking Based on Sensor Fusion
abstract
Persistent multi-object tracking (MOT) allows autonomous vehicles to navigate safely in highly dynamic environments. One of the well-known challenges in MOT is object occlusion when an object becomes unobservant for subsequent frames. The current MOT methods store objects information, such as trajectories, in internal memory to recover the objects after occlusions. However, they retain short-term memory to save computational time and avoid slowing down the MOT method. As a result, they lose track of objects in some occlusion scenarios, particularly long ones. In this paper, we propose DFR-FastMOT, a light MOT method that uses data from a camera and LiDAR sensors and relies on an algebraic formulation for object association and fusion. The formulation boosts the computational time and permits long-term memory that tackles more occlusion scenarios. Our method shows outstanding tracking performance over recent learning and non-learning benchmarks with about 3% and 4% margin in MOTA, respectively. Also, we conduct extensive experiments that simulate occlusion phenomena by employing detectors with various distortion levels. The proposed solution enables superior performance under various distortion levels in detection over current state-of-art methods. Our framework processes about 7,763 frames in 1.48 seconds, which is seven times faster than recent benchmarks. The framework will be available at https://github.com/MohamedNagyMostafa/DFR-FastMOT.
Mohamed Nagy, Majid Khonji, Jorge Dias 0001, Sajid Javed
ICRA4
2023 Multi-view Inspection of Flare Stacks Operation Using a Vision-controlled Autonomous UAV
abstract
Flare stacks are crucial safety control components in petrochemical plants that required efficient monitoring and inspection. In this work, an Unmanned Aerial Vehicle (UAV)-based multi-view operation inspection system for monitoring and assessing the operation of flare stacks is proposed. Image-Based Visual Servoing (IBVS) control is used to guide the autonomous UAV for multi-view visual data collection. Afterwards, the collected visual data is analyzed using a new Multi-View Convolutional Neural Network (MV-CNN) deep learning model to obtain useful conclusions on the system's operation and classify the current state of the observed system. The proposed system's performance was validated in a simulated petrochemical plant environment with operational flare stacks and the results showed superior performance of the proposed MV-CNN model compared to a conventional single-view CNN model.
Muaz Al Radi, Hamad Karki, Naoufel Werghi, Sajid Javed, Jorge Dias 0001
IECON5
2023 Multimodal hybrid features in 3D ear recognition
Karthika Ganesan, Iyyakutti Iyappan Ganapathi, Sajid Javed, Naoufel Werghi
Appl. Intell.4
2023 Window-based transformer generative adversarial network for autonomous underwater image enhancement
Mehnaz Ummar, Fayaz Ali Dharejo, Basit Alawode, Taslim Mahbub, Mohammad Jalil Piran, Sajid Javed
Eng. Appl. Artif. Intell.6
2023 Visual Object Tracking With Discriminative Filters and Siamese Networks: A Survey and Outlook
abstract
Accurate and robust visual object tracking is one of the most challenging and fundamental computer vision problems. It entails estimating the trajectory of the target in an image sequence, given only its initial location, and segmentation, or its rough approximation in the form of a bounding box. Discriminative Correlation Filters (DCFs) and deep Siamese Networks (SNs) have emerged as dominating tracking paradigms, which have led to significant progress. Following the rapid evolution of visual object tracking in the last decade, this survey presents a systematic and thorough review of more than 90 DCFs and Siamese trackers, based on results in nine tracking benchmarks. First, we present the background theory of both the DCF and Siamese tracking core formulations. Then, we distinguish and comprehensively review the shared as well as specific open research challenges in both these tracking paradigms. Furthermore, we thoroughly analyze the performance of DCF and Siamese trackers on nine benchmarks, covering different experimental aspects of visual tracking: datasets, evaluation metrics, performance, and speed comparisons. We finish the survey by presenting recommendations and suggestions for distinguished open challenges based on our analysis.
Sajid Javed, Martin Danelljan, Fahad Shahbaz Khan, Muhammad Haris Khan, Michael Felsberg, Jiri Matas
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Robot-Person Tracking in Uniform Appearance Scenarios: A New Dataset and Challenges
abstract
Person-tracking robots have many applications including security, surveillance, and autonomous driving. Despite the abundance of uniform appearance in many contexts and the challenges they exhibit, there is a lack of video datasets dedicated to benchmarking tracking algorithms in such contexts. In this article, we propose a new high-quality RGB-D benchmark called PTUA for robot–person tracking in uniform appearance scenarios. PTUA is recorded using an RGB-D sensor on top of a moving robot and consists of 45 sequences containing more than 85 K frames. Each frame is manually annotated with a bounding box and attributes, making PTUA the largest and the most challenging person tracking RGB-D dataset. To the best of our knowledge, such a densely annotated and properly synchronized RGB-D tracking benchmark does not exist in the literature. Each sequence comprises various challenges deriving from real-life scenarios where the target person appears highly similar to the background or distractors. By releasing PTUA, we expect to provide the community with a large-scale challenging RGB-D benchmark with high quality for the robust evaluation of trackers on uniform appearance scenarios for autonomous robots. We also present a rigorous experimental evaluation of the state-of-the-art trackers on the PTUA dataset with a comprehensive analysis. The findings evidence the challenges of person tracking in a uniform appearance scenario for both target tracking and robot–person tracking, and the need to bridge the performance gap. In addition, we propose a new RGB-D tracker that extracts features from RGB-D frames and it achieves the best performance on each challenging scenario of PTUA.
Xiaoxiong Zhang 0001, Adarsh Ghimire, Sajid Javed, Jorge Dias 0001, Naoufel Werghi
IEEE Trans. Hum. Mach. Syst.3
2023 Knowledge Distillation in Histology Landscape by Multi-Layer Features Supervision
abstract
Automatic tissue classification is a fundamental task in computational pathology for profiling tumor micro-environments. Deep learning has advanced tissue classification performance at the cost of significant computational power. Shallow networks have also been end-to-end trained using direct supervision however their performance degrades because of the lack of capturing robust tissue heterogeneity. Knowledge distillation has recently been employed to improve the performance of the shallow networks used as student networks by using additional supervision from deep neural networks used as teacher networks. In the current work, we propose a novel knowledge distillation algorithm to improve the performance of shallow networks for tissue phenotyping in histology images. For this purpose, we propose multi-layer feature distillation such that a single layer in the student network gets supervision from multiple teacher layers. In the proposed algorithm, the size of the feature map of two layers is matched by using a learnable multi-layer perceptron. The distance between the feature maps of the two layers is then minimized during the training of the student network. The overall objective function is computed by summation of the loss over multiple layers combination weighted with a learnable attention-based parameter. The proposed algorithm is named as Knowledge Distillation for Tissue Phenotyping (KDTP). Experiments are performed on five different publicly available histology image classification datasets using several teacher-student network combinations within the KDTP algorithm. Our results demonstrate a significant performance increase in the student networks by using the proposed KDTP algorithm compared to direct supervision-based training methods.
Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi
IEEE J. Biomed. Health Informatics1
2022 UTB180: A High-Quality Benchmark for Underwater Tracking
Basit Alawode, Mehnaz Ummar, Naoufel Werghi, Jorge Dias 0001, Ajmal Mian, Sajid Javed
ACCV (5)7
2022 Vision-based Inspection of Flare Stacks Operation Using a Visual Servoing Controlled Autonomous Unmanned Aerial Vehicle (UAV)
abstract
The inspection of flare stacks’ operation is a challenging task that requires technical expertise and human effort. Flare stack systems undergo various types of faults that need to be monitored in a timely manner to avoid costly and dangerous accidents. Automating this process via the application of autonomous robotic systems for collecting comprehensive data of the flare stack’s operation is a promising solution for minimizing the involved hazards and costs. In this work, a novel Unmanned Aerial Vehicle (UAV)-based autonomous inspection system for flare stacks performance monitoring is proposed. The system employs a deep learning detection network that was trained for detection of flame and smoke for vision-based flaring performance analysis. A visual servoing control technique was used for guiding the UAV’s movement throughout the inspection mission for collecting comprehensive visual inspection data. Simulations in a simulated petrochemical plant environment with flare stacks were performed for validating the performance of the proposed system. The proposed UAV system was able to collect the required data successfully and analysis of the obtained data returned useful information about the flare stack’s operation.
Muaz Al Radi, Hamad Karki, Naoufel Werghi, Sajid Javed, Jorge Dias 0001
IECON4
2022 Learning to localize image forgery using end-to-end attention network
Iyyakutti Iyappan Ganapathi, Sajid Javed, Syed Sadaf Ali, Arif Mahmood, Ngoc-Son Vu, Naoufel Werghi
Neurocomputing2
2022 Nucleus classification in histology images using message passing network
Taimur Hassan, Sajid Javed, Arif Mahmood, Talha Qaiser, Naoufel Werghi, Nasir M. Rajpoot
Medical Image Anal.2
2022 Graph Moving Object Segmentation
abstract
Moving Object Segmentation (MOS) is a fundamental task in computer vision. Due to undesirable variations in the background scene, MOS becomes very challenging for static and moving camera sequences. Several deep learning methods have been proposed for MOS with impressive performance. However, these methods show performance degradation in the presence of unseen videos; and usually, deep learning models require large amounts of data to avoid overfitting. Recently, graph learning has attracted significant attention in many computer vision applications since they provide tools to exploit the geometrical structure of data. In this work, concepts of graph signal processing are introduced for MOS. First, we propose a new algorithm that is composed of segmentation, background initialization, graph construction, unseen sampling, and a semi-supervised learning method inspired by the theory of recovery of graph signals. Second, theoretical developments are introduced, showing one bound for the sample complexity in semi-supervised learning, and two bounds for the condition number of the Sobolev norm. Our algorithm has the advantage of requiring less labeled data than deep learning methods while having competitive results on both static and moving camera videos. Our algorithm is also adapted for Video Object Segmentation (VOS) tasks and is evaluated on six publicly available datasets outperforming several state-of-the-art methods in challenging conditions.
Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Thierry Bouwmans
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Hierarchical Spatiotemporal Graph Regularized Discriminative Correlation Filter for Visual Object Tracking
abstract
Visual object tracking is a fundamental and challenging task in many high-level vision and robotics applications. It is typically formulated by estimating the target appearance model between consecutive frames. Discriminative correlation filters (DCFs) and their variants have achieved promising speed and accuracy for visual tracking in many challenging scenarios. However, because of the unwanted boundary effects and lack of geometric constraints, these methods suffer from performance degradation. In the current work, we propose hierarchical spatiotemporal graph-regularized correlation filters for robust object tracking. The target sample is decomposed into a large number of deep channels, which are then used to construct a spatial graph such that each graph node corresponds to a particular target location across all channels. Such a graph effectively captures the spatial structure of the target object. In order to capture the temporal structure of the target object, the information in the deep channels obtained from a temporal window is compressed using the principal component analysis, and then, a temporal graph is constructed such that each graph node corresponds to a particular target location in the temporal dimension. Both spatial and temporal graphs span different subspaces such that the target and the background become linearly separable. The learned correlation filter is constrained to act as an eigenvector of the Laplacian of these spatiotemporal graphs. We propose a novel objective function that incorporates these spatiotemporal constraints into the DCFs framework. We solve the objective function using alternating direction methods of multipliers such that each subproblem has a closed-form solution. We evaluate our proposed algorithm on six challenging benchmark datasets and compare it with 33 existing state-of-the art trackers. Our results demonstrate an excellent performance of the proposed algorithm compared to the existing trackers.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Lakmal D. Seneviratne, Naoufel Werghi
IEEE Trans. Cybern.1
2021 Spatially Constrained Context-Aware Hierarchical Deep Correlation Filters for Nucleus Detection in Histology Images
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi, Nasir M. Rajpoot
Medical Image Anal.1
2020 Deep Bidirectional Correlation Filters for Visual Object Tracking
abstract
Visual Object Tracking (VOT) is an essential task for many computer vision applications. VOT becomes challenging when a target object faces severe occlusion, drastic illumination changes, and scale variation problems. In the literature, Discriminative Correlation Filters (DCFs)-based tracking methods have achieved promising results in terms of accuracy and efficiency in many complex VOT scenarios. A plethora of DCFs trackers have been proposed which exploit information observed in past frames to create and update DCFs for VOT. To adapt to target appearance variations, the DCFs are enhanced by incorporating spatial and temporal consistency constraints. Nevertheless, the performance degradation is observed for these methods because of the aforementioned limitations. To address these issues, we propose a novel algorithm based on bidirectional DCFs for VOT. In this algorithm, we propose the original idea of leveraging information from both past and future frames. The proposed algorithm first tracks the target object forward in the video sequence and then its uses the predicted location of the last window frame and track the target object backward towards the current frame. We design an appearance consistency loss function by taking the$L_{2}$norm between the regression target of the forward tracking and response map of the backward tracking to obtain the resulting response map. Our proposed algorithm realizes a highly accurate DCFs because forward and backward tracking information are fused together for consistent VOT. Although, a result will be output with some small delay because information is taken from a future to the present period, our proposed algorithm has the merit of addressing the drastic appearance variations VOT challenges. We evaluate our proposed tracker using deep features on three publicly available challenging datasets. Our results demonstrate the superior performance of the proposed tracker compared to the existing state-of-the-art trackers.
Sajid Javed, Xiaoxiong Zhang 0001, Lakmal D. Seneviratne, Jorge Dias 0001, Naoufel Werghi
FUSION1
2020 CS-RPCA: Clustered Sparse RPCA for Moving Object Detection
abstract
Moving object detection (MOD) is an important step for many computer vision applications. In the last decade, it is evident that RPCA has shown to be a potential solution for MOD and achieved a promising performance under various challenging background scenes. However, because of the lack of different types of features, RPCA still shows degraded performance in many complicated background scenes such as dynamic backgrounds, cluttered foreground objects, and camouflage. To address these problems, this paper presents a Clustered Sparse RPCA (CS-RPCA) for MOD under challenging environments. The proposed algorithm extracts multiple features from video sequences and then employs RPCA to get the low-rank and sparse component from each representation. The sparse subspaces are then emerged into a common sparse component using Grassmann manifold. We proposed a novel objective function which computes the composite sparse component from multiple representations and it is solved using non-negative matrix factorization method. The proposed algorithm is evaluated on two challenging datasets for MOD. Results demonstrate excellent performance of the proposed algorithm as compared to existing state-of-the-art methods.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi
ICIP1
2020 Gender Recognition on RGB-D Image
abstract
In this paper, we propose a deep-learning approach for human gender classification on RGB-D images. Unlike most of the existing methods, which use hand-crafted features from the human face, we exploit local information from the head and global information from the whole body to classify people's gender. A head detector is fine-tuned on YOLO to detect the head regions on the images automatically. Two gender classifiers are trained using head images and whole-body images separately. The final prediction is made by fusing the two classifiers' results. The presented method outperforms the state-of-art with an improvement in the accuracy of 2.6%, 7.6%, and 8.4% on three different test data of a challenging gender dataset which includes human standing, walking, and interacting scenarios.
Xiaoxiong Zhang 0001, Sajid Javed, Ahmad Obeid 0001, Jorge Dias 0001, Naoufel Werghi
ICIP2
2020 Cellular community detection for tissue phenotyping in colorectal cancer histology images
Sajid Javed, Arif Mahmood, Muhammad Moazam Fraz, Navid Alemi Koohbanani, Ksenija Benes, Yee-Wah Tsang, Katherine Hewitt, David B. A. Epstein, David R. J. Snead, Nasir M. Rajpoot
Medical Image Anal.1
2020 Robust Structural Low-Rank Tracking
abstract
Visual object tracking is an essential task for many computer vision applications. It becomes very challenging when the target appearance changes especially in the presence of occlusion, background clutter, and sudden illumination variations. Methods, that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when facing the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking in complex scenarios. In the proposed algorithm, we consider spatial and temporal appearance consistency constraints, among the particles in the low-rank subspace, embedded in four different graphs. The resulting objective function encoding these constraints is novel and it is solved using linearized alternating direction method with adaptive penalty both in batch fashion as well as in online fashion. Our proposed objective function jointly learns the spatial and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on four challenging datasets demonstrate excellent performance of the proposed algorithm as compared to current state-of-the-art methods.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi
IEEE Trans. Image Process.1
2020 Multiplex Cellular Communities in Multi-Gigapixel Colorectal Cancer Histology Images for Tissue Phenotyping
abstract
In computational pathology, automated tissue phenotyping in cancer histology images is a fundamental tool for profiling tumor microenvironments. Current tissue phenotyping methods use features derived from image patches which may not carry biological significance. In this work, we propose a novel multiplex cellular community-based algorithm for tissue phenotyping integrating cell-level features within a graph-based hierarchical framework. We demonstrate that such integration offers better performance compared to prior deep learning and texture-based methods as well as to cellular community based methods using uniplex networks. To this end, we construct celllevel graphs using texture, alpha diversity and multi-resolution deep features. Using these graphs, we compute cellular connectivity features which are then employed for the construction of a patch-level multiplex network. Over this network, we compute multiplex cellular communities using a novel objective function. The proposed objective function computes a low-dimensional subspace from each cellular network and subsequently seeks a common low-dimensional subspace using the Grassmann manifold. We evaluate our proposed algorithm on three publicly available datasets for tissue phenotyping, demonstrating a significant improvement over existing state-of-the-art methods.
Sajid Javed, Arif Mahmood, Naoufel Werghi, Ksenija Benes, Nasir M. Rajpoot
IEEE Trans. Image Process.1
2019 Structural Low-Rank Tracking
abstract
Visual object tracking is an important step for many computer vision applications. The task becomes very challenging when the target undergoes heavy occlusion, background clutters, and sudden illumination variations. Methods that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when an object faces the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking. In the proposed algorithm, we enforce local spatial, global spatial and temporal appearance consistency among the particles in the low-rank subspace by constructing three graphs. The Laplacian matrices of these graphs are incorporated into the novel low-rank objective function which is solved using linearized alternating direction method with an adaptive penalty. Our proposed objective function jointly learns the spatial, global, and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on two challenging benchmark datasets show the superiority of the proposed algorithm as compared to current state-of-the-art methods.
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi
AVSS1
2019 Unsupervised deep context prediction for background estimation and foreground segmentation
Maryam Sultana, Arif Mahmood, Sajid Javed, Soon Ki Jung
Mach. Vis. Appl.3
2019 Deep neural network concepts for background subtraction: A systematic review and comparative evaluation
Thierry Bouwmans, Sajid Javed, Maryam Sultana, Soon Ki Jung
Neural Networks2
2019 Moving Object Detection in Complex Scene Using Spatiotemporal Structured-Sparse RPCA
abstract
Moving object detection is a fundamental step in various computer vision applications. Robust Principal Component Analysis (RPCA) based methods have often been employed for this task. However, the performance of these methods deteriorates in the presence of dynamic background scenes, camera jitter, camouflaged moving objects, and/or variations in illumination. It is because of an underlying assumption that the elements in the sparse component are mutually independent, and thus the spatiotemporal structure of the moving objects is lost. To address this issue, we propose a spatiotemporal structured sparse RPCA algorithm for moving objects detection, where we impose spatial and temporal regularization on the sparse component in the form of graph Laplacians. Each Laplacian corresponds to a multi-feature graph constructed over superpixels in the input matrix. We enforce the sparse component to act as eigenvectors of the spatial and temporal graph Laplacians while minimizing the RPCA objective function. These constraints incorporate a spatiotemporal subspace structure within the sparse component. Thus, we obtain a novel objective function for separating moving objects in the presence of complex backgrounds. The proposed objective function is solved using a linearized alternating direction method of multipliers based batch optimization. Moreover, we also propose an online optimization algorithm for real-time applications. We evaluated both the batch and online solutions using six publicly available datasets that included most of the aforementioned challenges. Our experiments demonstrated the superior performance of the proposed algorithms compared with the current state-of-the-art methods.
Sajid Javed, Arif Mahmood, Somaya Al-Máadeed, Thierry Bouwmans, Soon Ki Jung
IEEE Trans. Image Process.1
2018 Two Stream Deep CNN-RNN Attentive Pooling Architecture for Video-Based Person Re-identification
Wajeeha Ansar, Muhammad Moazam Fraz, Muhammad Shahzad 0002, Imad Gohar, Sajid Javed, Soon Ki Jung
CIARP5
2018 On the Applications of Robust PCA in Image and Video Processing
abstract
Robust principal component analysis (RPCA) via decomposition into low-rank plus sparse matrices offers a powerful framework for a large variety of applications such as image processing, video processing, and 3-D computer vision. Indeed, most of the time these applications require to detect sparse outliers from the observed imagery data that can be approximated by a low-rank matrix. Moreover, most of the time experiments show that RPCA with additional spatial and/or temporal constraints often outperforms the state-of-the-art algorithms in these applications. Thus, the aim of this paper is to survey the applications of RPCA in computer vision. In the first part of this paper, we review representative image processing applications as follows: 1) low-level imaging such as image recovery and denoising, image composition, image colorization, image alignment and rectification, multifocus image, and face recognition; 2) medical imaging such as dynamic magnetic resonance imaging (MRI) for acceleration of data acquisition, background suppression, and learning of interframe motion fields; and 3) imaging for 3-D computer vision with additional depth information such as in structure from motion (SfM) and 3-D motion recovery. In the second part, we present the applications of RPCA in video processing which utilize additional spatial and temporal information compared to image processing. Specifically, we investigate video denoising and restoration, hyperspectral video, and background/foreground separation. Finally, we provide perspectives on possible future research directions and algorithmic frameworks that are suitable for these applications.
Thierry Bouwmans, Sajid Javed, Hongyang Zhang 0001, Zhouchen Lin, Ricardo Otazo
Proc. IEEE2
2018 Spatiotemporal Low-Rank Modeling for Complex Scene Background Initialization
abstract
Background modeling constitutes the building block of many computer-vision tasks. Traditional schemes model the background as a low rank matrix with corrupted entries. These schemes operate in batch mode and do not scale well with the data size. Moreover, without enforcing spatiotemporal information in the low-rank component, and because of occlusions by foreground objects and redundancy in video data, the design of a background initialization method robust against outliers is very challenging. To overcome these limitations, this paper presents a spatiotemporal low-rank modeling method on dynamic video clips for estimating the robust background model. The proposed method encodes spatiotemporal constraints by regularizing spectral graphs. Initially, a motion-compensated binary matrix is generated using optical flow information to remove redundant data and to create a set of dynamic frames from the input video sequence. Then two graphs are constructed, one between frames for temporal consistency and the other between features for spatial consistency, to encode the local structure for continuously promoting the intrinsic behavior of the low-rank model against outliers. These two terms are then incorporated in the iterative Matrix Completion framework for improved segmentation of background. Rigorous evaluation on severely occluded and dynamic background sequences demonstrates the superior performance of the proposed method over state-of-the-art approaches.
Sajid Javed, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung
IEEE Trans. Circuits Syst. Video Technol.1
2017 Background-Foreground Modeling Based on Spatiotemporal Sparse Subspace Clustering
abstract
Background estimation and foreground segmentation are important steps in many high-level vision tasks. Many existing methods estimate background as a low-rank component and foreground as a sparse matrix without incorporating the structural information. Therefore, these algorithms exhibit degraded performance in the presence of dynamic backgrounds, photometric variations, jitter, shadows, and large occlusions. We observe that these backgrounds often span multiple manifolds. Therefore, constraints that ensure continuity on those manifolds will result in better background estimation. Hence, we propose to incorporate the spatial and temporal sparse subspace clustering into the robust principal component analysis (RPCA) framework. To that end, we compute a spatial and temporal graph for a given sequence using motion-aware correlation coefficient. The information captured by both graphs is utilized by estimating the proximity matrices using both the normalized Euclidean and geodesic distances. The low-rank component must be able to efficiently partition the spatiotemporal graphs using these Laplacian matrices. Embedded with the RPCA objective function, these Laplacian matrices constrain the background model to be spatially and temporally consistent, both on linear and nonlinear manifolds. The solution of the proposed objective function is computed by using the linearized alternating direction method with adaptive penalty optimization scheme. Experiments are performed on challenging sequences from five publicly available datasets and are compared with the 23 existing state-of-the-art methods. The results demonstrate excellent performance of the proposed algorithm for both the background estimation and foreground segmentation.
Sajid Javed, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung
IEEE Trans. Image Process.1
2016 Motion-Aware Graph Regularized RPCA for background modeling of complex scenes
abstract
Computing a background model from a given sequence of video frames is a prerequisite for many computer vision applications. Recently, this problem has been posed as learning a low-dimensional subspace from high dimensional data. Many contemporary subspace segmentation methods have been proposed to overcome the limitations of the methods developed for simple background scenes. Unfortunately, because of the absence of motion information and without preserving intrinsic geometric structure of video data, most existing algorithms do not provide promising nature of the low-rank component for complex scenes. Such as largely occluded background by foreground objects, superfluity in video frames in order to cope with intermittent motion of foreground objects, sudden lighting condition variation, and camera jitter sequences. To overcome these difficulties, we propose a motion-aware regularization of graphs on low-rank component for video background modeling. We compute optical flow and use this information to make a motion-aware matrix. In order to learn the locality and similarity information within a video we compute inter-frame and intra-frame graphs which we use to preserve geometric information in the low-rank component. Finally, we use linearized alternating direction method with parallel splitting and adaptive penalty to incorporate the preceding steps to recover the model of the background. Experimental evaluations on challenging sequences demonstrate promising results over state-of-the-art methods.
Sajid Javed, Soon Ki Jung, Arif Mahmood, Thierry Bouwmans
ICPR1
2014 OR-PCA with MRF for Robust Foreground Detection in Highly Dynamic Backgrounds
Sajid Javed, Seon Ho Oh, Andrews Sobral, Thierry Bouwmans, Soon Ki Jung
ACCV (3)1