EDBT 2026 Demo / reviewers in the wild / expert
Ferdous Sohel
dblp:87/1736 · also Ferdous Ahmed Sohel
· DBLP profile ↗
125ranked-venue papers
9as first author
41since 2021 · last 2026
0000-0003-1557-4907ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 67 · 3 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 6 first-author · 14 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A computational learning pipeline for glaucoma progression detection based on the prediction of visual field changes from fundus photographs
Md. Reduanul Haque, Andrew Mehnert, William H. Morgan, Graham Mann, Ferdous Sohel |
Expert Syst. Appl. | 5 |
| 2026 | 3D-CDNeT: Cross-domain learning with enhanced speed and robustness for point cloud recognitionabstractDespite progress in 3D object recognition using deep learning (DL), challenges such as domain shift, occlusion, and viewpoint variations hinder robust performance. Additionally, the high computational cost and lack of labeled data limit real-time deployment in applications such as autonomous driving and robotic manipulation. To address these challenges, we propose 3D-CDNeT, a novel cross-domain deep learning network designed for unsupervised learning, enabling efficient and robust point cloud recognition. At the core of our model is a lightweight graph-infused attention encoder (GIAE) that enables effective feature interaction between the source and target domains. It not only improves recognition accuracy but also reduces inference time, which is essential for real-time applications. To enhance robustness and adaptability, we introduce a feature invariance learning module (FILM) using contrastive loss for learning invariant features. In addition, we adopt a Generative Decoder (GD) based on a Variational Auto-Encoder (VAE) to model diverse latent spaces and reconstruct meaningful 3D structures from the point cloud. This reconstruction process acts as a self-supervised generative objective that complements the discriminative recognition task, guiding the encoder to learn structure-preserving and domain-invariant features that improve recognition under occlusion and cross-domain conditions. Our proposed model unifies generative and discriminative tasks by using self-attention on the object covariance matrix to facilitate efficient information exchange, enabling the extraction of both local and global features. We further develop a self-supervised pretraining strategy that learns both global and local object invariances through GIAE and GD, respectively. A new loss function, combining contrastive loss and Chamfer distance, is proposed to strengthen cross-domain feature alignment. Experimental results on three benchmark datasets demonstrate that 3D-CDNeT outperforms existing state-of-the-art (SOTA) methods in recognition accuracy and inference speed, offering a practical solution for real-time 3D perception tasks. It achieves accuracies of 90.6 % on ModelNet40, 95.2 % on ModelNet10, and 76.4 % on the ScanObjectNN dataset in linear evaluation tasks, all while reducing runtime by 45 % without compromising performance. Detailed qualitative comparisons and ablation studies are provided to validate the effectiveness of each component and demonstrate the superior performance of our proposed method. Abu Bakor Hayat Arnob, A. A. M. Muzahid, Hua Han 0002, Ferdous Sohel |
Neurocomputing | 5 |
| 2026 | QTopic: A novel quantum perspective on learning topics from textabstractTopic modeling is an unsupervised technique in natural language processing (NLP) used to identify hidden topic structures within large text datasets. Among traditional approaches of topic modeling, latent Dirichlet allocation, BERTopic, and Top2Vec, are widely adopted to uncover hidden topics in text data. However, these methods often struggle with poor performance in scenarios involving limited data availability or high-dimensional textual features. In this research, we propose QTopic, a novel hybrid quantum-classical topic modeling architecture that leverages quantum properties through parameterized quantum circuits. By integrating quantum-enhanced sampling into the inference pipeline, the proposed model captures richer topic distributions by mapping textual data into a higher-dimensional space. Benchmark experiments demonstrate that QTopic consistently outperforms classical approaches in terms of coherence, diversity, and topic distinctiveness, particularly when modeling a small number of topics. This study demonstrates the promise of quantum techniques in advancing unsupervised NLP, while also highlighting hardware limitations that present challenges for future research. Monika Kabir, Mohammed Kaosar, Ferdous Sohel |
Neurocomputing | 3 |
| 2026 | StructFuse: Harmonizing multiple structural cues for diffusion-driven image compositingabstractCompositing visually coherent foreground and background elements remains a pivotal challenge in generative image editing. Prominent diffusion-based approaches either rely on rigid image-level inputs, which limit scene diversity, or focus on localized mask conditioning (e.g., edges, depth), while struggling to effectively combine multiple structural cues. To address these limitations, we propose an advanced diffusion-based framework that achieves precise foreground-background integration through the use of structured visual cues. Our approach introduces three key innovations: (1) a learnable adaptive gating strategy that dynamically fuses multiple structural cues (edge, contour, and depth) while effectively capturing richer and balanced conditional representations; (2) a bidirectional feature modulation method that injects multi-resolution control signals across the diffusion U-Net encoder–decoder hierarchy, significantly enhancing compositional quality; and (3) a mask-routed cross-attention mechanism that isolates text influence to user-specified regions, eliminating semantic bleeding. Additionally, we propose a novel pipeline for constructing a composite image dataset tailored to the image compositing task. Experimental results demonstrate that our framework outperforms state-of-the-art baseline methods across benchmark metrics, including Harmony Score, FID, SSIM, CLIP Score, and MUSIQ, reflecting strong compositional integrity and visual coherence. This approach offers a scalable and extensible solution for compositional image generation, with broad applications in synthetic dataset generation, creative content design, and augmented reality. Dean Diepeveen, Ferdous Sohel |
Pattern Recognit. | 3 |
| 2026 | Joint Adversarial Attack: An Effective Approach to Evaluate Robustness of 3D Object Tracking
Riran Cheng, Xupeng Wang 0001, Ferdous Sohel |
Pattern Recognit. | 3 |
| 2025 | LiDAR-SPD: Improving Adversarial Robustness of 3D Object Detection via Spherical Projection and DiffusionabstractThe advancements in light detection and ranging (LiDAR) sensors and 3D object detection techniques have boosted their deployment in a wide range of applications, autonomous driving, in particular. However, it has been demonstrated that 3D object detection models based on deep neural networks exhibit vulnerabilities and tend to be susceptible to adversarial attacks. Nonetheless, there exists a scarcity of defensive strategies explicitly tailored for mitigating adversarial attacks on 3D object detection. In this paper, we introduce LiDAR-SPD, a novel approach to defend against adversarial attacks targeting LiDAR-based 3D object detectors. Specifically, a spherical purification unit is designed, which encompasses two pivotal processes: spherical projection and spherical diffusion. The former leverages a spatial projection strategy to eliminate adversarial point clouds inserted in occluded regions, while the latter employs a diffusion model to regenerate points, rendering it closer to a pristine LiDAR scene. Comprehensive experiments conducted on the KITTI dataset demonstrate that our proposed LiDAR-SPD method effectively thwarts various types of adversarial attacks, decreasing the attack success rates against 3D object detectors by 60%. Mumuxin Cai, Xupeng Wang 0001, Ferdous Sohel |
ICASSP | 3 |
| 2025 | Superpoints Guided Local Explanation For Deep 3D Trackersabstract3D object tracking has become a popular research topic because of its broad application prospects. However, it remains a challenging task to advance the trustworthiness of deep trackers, caused by the complex network structure of black-box models. In this paper, a local explanation method for 3D object tracking is proposed, which trains an interpretable surrogate model to reveal the contribution of superpoints in the search area. Specifically, local points of the search area with comparable geometric features are aggregated as superpoints, which serve as the fundamentals of the explanation. In contrast to the commonly used voxels, superpoints capture semantic information of the search area, and facilitate an intuitive understanding of the predictions. In addition, a distance-aware masking strategy is proposed for generating the sample set to train a surrogate model, which corresponds to latent contributions of superpoints to predictions and improves the efficacy of the explanation. Experiments have demonstrated that the proposed explainability approach can effectively provide explanations to deep tracking models. Riran Cheng, Xupeng Wang 0001, Ferdous Sohel |
ICASSP | 3 |
| 2025 | RSM: Refined Saliency Map For Explainable 3D Object TrackingabstractSaliency maps play a major role in understanding the decision-making process of 3D models by illustrating the importance of individual points from the input to model predictions. However, saliency maps typically suffer from inaccuracies due to not considering the potential classification of contributions made by a point. In this paper, a two-stage explainability method for 3D object tracking is proposed to generate a refined saliency map (RSM), which refines the contributions of points to positive and negative based on their actual effects on tracking performances. Specifically, in stage I, a point-wise growing downsampling algorithm is developed to generate subsets of the search area, under which the model’s behavior is evaluated to precisely identify the points with negative contributions. Subsequently, a voxel-wise downsampling algorithm is performed along with the deviation metric to select points with positive contributions in stage II. Experiments demonstrate that RSM can generate high-quality explanations to popular 3D trackers. Riran Cheng, Xupeng Wang 0001, Ferdous Sohel |
ICASSP | 3 |
| 2025 | GNF: Gaussian Neural Fields for Multidimensional Signal Representation and ReconstructionabstractAbstract Neural fields have emerged as a powerful framework for representing continuous multidimensional signals such as images and videos, 3D and 4D objects and scenes, and radiance fields. While efficient, achieving high‐quality representation requires the use of wide and deep neural networks. These, however, are slow to train and evaluate. Although several acceleration techniques have been proposed, they either trade memory for faster training and/or inference, rely on thousands of fitted primitives with considerable optimization time, or compromise the smooth, continuous nature of neural fields. In this paper, we introduce Gaussian Neural Fields (GNF), a novel compact neural decoder that maps learned feature grids into continuous non‐linear signals, such as RGB images, Signed Distance Functions (SDFs), and radiance fields, using a single compact layer of Gaussian kernels defined in a high‐dimensional feature space. Our key observation is that neurons in traditional MLPs perform simple computations, usually a dot product followed by an activation function, necessitating wide and deep MLPs or high‐resolution feature grids to model complex functions. In this paper, we show that replacing MLP‐based decoders with Gaussian kernels whose centers are learned features yields highly accurate representations of 2D (RGB), 3D (geometry), and 5D (radiance fields) signals with just a single layer of such kernels. This representation is highly parallelizable, operates on low‐resolution grids, and trains in under 15 seconds for 3D geometry and under 11 minutes for view synthesis. GNF matches the accuracy of deep MLP‐based decoders with far fewer parameters and significantly higher inference throughput. The source code is publicly available at https://grbfnet.github.io/ . Abelaziz Bouzidi, Hamid Laga, Hazem Wannous, Ferdous Sohel |
Comput. Graph. Forum | 4 |
| 2025 | Enhanced concrete crack segmentation with MSMC-U-Net: integrating multiscale features and contextual analysis for infrastructure safety
Sachal Pervaiz, Changqing Cai, Rawal Javed, Shuangyin Liu, Ferdous Sohel, Shahbaz Gul Hassan, Huakun Huang |
Expert Syst. Appl. | 5 |
| 2025 | Conditional plane-based multi-scene representation for novel view synthesisabstractThe method overview. The explicit representation on the right side represents the shared canonical space and the view space using 12 feature planes. The gray arrows indicate the feature projection from the canonical representation. The left side shows the deformation between the canonical space and the view space. Pairwise features (e.g., X Y − Z T ) from the canonical and view representations are aggregated to obtain the final feature vector. ⨂ indicates feature aggregation. Density and appearance decoders, which estimate the geometry and color, are conditioned on the scene’s latent s i . Existing explicit and implicit-explicit hybrid neural representations for novel view synthesis are scene-specific. In other words, they represent only a single scene and require retraining for every novel scene. Implicit scene-agnostic methods rely on large multilayer perception (MLP) networks conditioned on learned features. They are computationally expensive during training and rendering times. In contrast, we propose a novel plane-based representation that learns to represent multiple static and dynamic scenes during training and renders per-scene novel views during inference. The method consists of a deformation network, explicit feature planes, and a conditional decoder. Explicit feature planes are used to represent a time-stamped view space volume and a shared canonical volume across multiple scenes. The deformation network learns the deformations across shared canonical object space and time-stamped view space. The conditional decoder estimates the color and density of each scene constrained by a scene-specific latent code. We evaluated and compared the performance of the proposed representation on static (NeRF) and dynamic (Plenoptic videos) datasets. The results show that explicit planes combined with tiny MLPs can efficiently train multiple scenes simultaneously. The project page: https://anonpubcv.github.io/cplanes/ . • We present a novel multi-scene representation that uses twelve explicit feature planes. • The method uses encoder-less generalization to learn discriminative features for each scene. • We represent multiple dynamic scenes without relying on optical flow estimation. • The auto-decoded latent (the scene dimension) can interpolate between scenes. • It achieves state-of-the-art rendering results for both static and dynamic scenes. Uchitha Rajapaksha, Hamid Laga, Dean Diepeveen, Mohammed Bennamoun, Ferdous Sohel |
Neurocomputing | 5 |
| 2025 | Transferable universal adversarial attack against 3D object detection with latent feature disruption
Mumuxin Cai, Xupeng Wang 0001, Ferdous Sohel, Dian Xiao |
J. Syst. Archit. | 3 |
| 2025 | Black-Box Explainability-Guided Adversarial Attack for 3D Object TrackingabstractWith the development of deep 3D tracking models and their broad prospects for safety-critical applications, adversarial robustness, i.e., the ability of deep models to resist malicious adversarial attacks, has become an important research topic. Previous works generate adversarial examples by tampering with points of the input point cloud indiscriminately. Consequently, they suffer from high computing costs and limited attack performance caused by the trade-off between imperceptibility and adversarial strength. In this paper, we propose a novel adversarial attack against 3D object tracking, which is guided by an occlusion-based explainability method to target points crucial for the predictions in the search area and results in a significant deviation between the predictions and the ground truth. Specifically, an attribution map is generated to reveal the importance of points to the model decision, which is achieved by measuring the variations of tracking performance under subsets generated by the downsampling strategy. To facilitate the generation of attribution maps, the downsampling strategy considers prior knowledge of 3D trackers, which assigns higher sampling probabilities to points with potentially higher contributions enclosed by bounding boxes. Multi-scale fusion is also leveraged to integrate the sensitivity of the model to local regions of varying sizes. Considering the requirement of imperceptibility on adversarial attacks, a hard geometric constraint is imposed on the targeted critical points, which produces perturbations with the property of surface invariance. Furthermore, in contrast to existing works devoted to spatial information manipulation only, multiple loss functions are developed to guide the perturbation generation, where the predicted motions of the tracking target representing the spatial-temporal information unique to the tracking task are distorted to deceive 3D trackers. Extensive experiments conducted on public benchmarks and 3D trackers demonstrate that our method can generate effective and imperceptible adversarial examples with tiny perturbations. Riran Cheng, Xupeng Wang 0001, Ferdous Sohel |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Dual-Phase Framework for Few-Shot Hyperspectral Image Classification With Spatiospectral Masked Autoencoder and Episode TrainingabstractThis article introduces a two-phase learning approach for hyperspectral image (HSI) classification using few-shot learning (FSL). For the first phase, we present a novel spatiospectral masked autoencoder (ssMAE)—an advanced self-supervised learner. For the ssMAE backbone network, we designed a transformer encoder-decoder network, where we replaced the linear layer that is used as the initial feature embedding with a 3-D convolutional layer to better extract local spectral-spatial features from 3-D visible sub-patches. By tapping into vast unlabeled data, the ssMAE learns general HSI features. In the second phase, the ssMAE encoder is fine-tuned to extract discriminative features for classification using the few-shot labeled training samples. This is achieved through a unique hybrid episode learning method that integrates the ssMAE encoder in a prototypical network (PN). We innovate with a mix of global and local prototypes (combined global-local (CGL) prototype) to refine label predictions. This technique maximizes data usage, focuses on specific samples, and mitigates issues from subpar episodes. Tested on three HSI datasets, our approach outperforms alternative few-shot methods. The code will be made publicly available athttps://github.com/Weejaa04/SSMAE. Wijayanti Nurul Khotimah, Mohammed Bennamoun, Farid Boussaïd, Lian Xu, Ferdous Sohel |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | RSUIA: Dynamic No-Reference Underwater Image Assessment via Reinforcement SequencesabstractUnderwater image quality assessment (UIQA) is a challenging task due to the complexities of underwater environments. Traditional UIQA methods primarily rely on fitting mean opinion scores (MOS), which are limited by human visual biases. To address the above limitation, we propose a no-reference underwater image quality assessment paradigm using reinforcement sequences. Our paradigm leverages reinforcement learning to iteratively merge the input image with the corresponding ground truth, generating an optimized sequence of images. A classifier generates probability arrays for the optimized sequence, which are converted into objective scores by a regression model. Unlike existing methods that focus solely on the final quality score, our paradigm emphasizes dynamic quality changes throughout the image-enhancement process. By employing objective mixing ratio labels, our reinforcement sequence dataset reduces subjective bias. The multiscale classifier captures local and global information differences between the input and ground truth images, effectively preserving the contrast and detail in diverse lighting conditions. Our paradigm combines multi-source data classification with support vector regression, optimizing the mapping of feature vectors to quality scores through fine-tuning libsvm kernel parameters. Experimental results on multiple benchmark datasets demonstrate that our paradigm outperforms the state-of-the-art UIQA methods, providing an effective solution for Underwater Image quality Assessment via Reinforcement Sequences (RSUIA). Jingchun Zhou, Chunjiang Liu, Dehuan Zhang, Zongxin He, Ferdous Sohel, Qiuping Jiang |
IEEE Trans. Multim. | 5 |
| 2025 | Auxiliary Tasks Enhanced Dual-Affinity Learning for Weakly Supervised Semantic SegmentationabstractMost existing weakly supervised semantic segmentation (WSSS) methods rely on class activation mapping (CAM) to extract coarse class-specific localization maps using image-level labels. Prior works have commonly used an off-line heuristic thresholding process that combines the CAM maps with off-the-shelf saliency maps produced by a general pretrained saliency model to produce more accurate pseudo-segmentation labels. We propose AuxSegNet+, a weakly supervised auxiliary learning framework to explore the rich information from these saliency maps and the significant intertask correlation between saliency detection and semantic segmentation. In the proposed AuxSegNet+, saliency detection and multilabel image classification are used as auxiliary tasks to improve the primary task of semantic segmentation with only image-level ground-truth labels. We also propose a cross-task affinity learning mechanism to learn pixel-level affinities from the saliency and segmentation feature maps. In particular, we propose a cross-task dual-affinity learning module to learn both pairwise and unary affinities, which are used to enhance the task-specific features and predictions by aggregating both query-dependent and query-independent global context for both saliency detection and semantic segmentation. The learned cross-task pairwise affinity can also be used to refine and propagate CAM maps to provide better pseudo labels for both tasks. Iterative improvement of segmentation performance is enabled by cross-task affinity learning and pseudo-label updating. Extensive experiments demonstrate the effectiveness of the proposed approach with new state-of-the-art WSSS results on the challenging PASCAL VOC and MS COCO benchmarks. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Wanli Ouyang, Ferdous Sohel, Dan Xu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-Shaped Structures
Tahmina Khanam, Hamid Laga, Mohammed Bennamoun, Guanjin Wang, Ferdous Sohel, Farid Boussaïd, Anuj Srivastava |
ECCV (67) | 5 |
| 2024 | LoRa localisation using single mobile gatewayabstractEffective use of GPS and mobile networks for localisation in rangeland areas is constrained by their high power consumption and high deployment costs. Long-range (LoRa), a low-power wide area network (LPWAN) technology, can be employed to mitigate these challenges. In contrast to prior research where the prevalent approaches entail multiple gateways. This work proposes a valuable methodology focused on a single mobile LoRa gateway for localisation. A particle filtering and machine learning-based pipeline is employed to map the distance between a target node and the gateway from the received signal strength indicator (RSSI). Particle filtering is used to reduce the impact of noise on the RSSI values. Then, several machine learning techniques, such as support vector machines, random forest, and k-nearest neighbour, are used on the RSSI values to estimate the distance. The estimated distance is then used for tracking using a centroid pseudo-trilateration method. The proposed method was tested in a real-world semi-line-of-sight setting, using three datasets generated by LoRaWAN-specified hardware components and a server. Two forms of experiments were performed: active searching and passive monitoring. We propose an iterative estimation process to address the dilution of precision caused by the initial positions of the gateway required for active searching applications. The results show that active searching typically requires 2 to 3 hops to reach a target node. The accuracy of passive monitoring depends on the proximity of the gateway, which varies from 20 m to 170 m. This proposed approach has the potential to open the way for localising, tracking, or monitoring target objects within sparsely populated rangeland areas, even when resources are severely constrained. Khondoker Ziaul Islam, David Murray, Dean Diepeveen, Michael G. K. Jones, Ferdous Sohel |
Comput. Commun. | 5 |
| 2024 | Deep learning for 3D object recognition: A survey
A. A. M. Muzahid, Hua Han 0002, Junaid Jamshid, Ferdous Sohel |
Neurocomputing | 7 |
| 2024 | DRC: Chromatic aberration intensity priors for underwater image enhancement
Zongxin He, Dehuan Zhang, Weishi Zhang, Zifan Lin, Ferdous Sohel |
J. Vis. Commun. Image Represent. | 6 |
| 2024 | A Pixel Distribution Remapping and Multi-Prior Retinex Variational Model for Underwater Image EnhancementabstractHigh-quality underwater imaging is crucial for underwater exploration. However, particle scattering and light absorption by seawater significantly degrade image clarity. To address these issues, we propose a novel underwater image enhancement (UIE) method that combines pixel distribution remapping (PDR) with a multi-priority Retinex variational model. We design a pre-compensation method for severely attenuated channels that effectively prevents new color artifacts during color correction. By combining the inter-channel coupling relationships, we compute a limiting factor to remap pixel distribution curves to improve image contrast. In addition, considering the significant noise interference, we introduce the prior knowledge, including underwater noise and texture priors, while constructing the variational model, and design penalty terms that match the underwater characteristics to remove excessive noise in the reflectance component. Our approach efficiently decouples the illumination and reflectance components using a rapid solver. Subsequently, gamma correction adjusts the illumination component, and the corrected illumination and reflectance components are fused to reconstruct the final natural output image. Comprehensive evaluations across various datasets reveal that our approach significantly surpasses current state-of-the-art (SOTA) methods. These results demonstrate the effectiveness of our method in correcting color bias and compensating for luminance losses in underwater imagery. Our code is available at:https://github.com/zhoujingchun03/PDRMRV. Jingchun Zhou, Shiyin Wang, Zifan Lin, Qiuping Jiang, Ferdous Sohel |
IEEE Trans. Multim. | 5 |
| 2023 | An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
Linping Xu, Dejun Zhang, Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Sixing Yin, Ferdous Sohel |
INTERSPEECH | 10 |
| 2023 | Anti-aliasing deep image classifiers using novel depth adaptive blurring and activation function
Md Tahmid Hossain, Shyh Wei Teng, Guojun Lu, Mohammad Arifur Rahman, Ferdous Sohel |
Neurocomputing | 5 |
| 2023 | An AUC-maximizing classifier for skewed and partially labeled data with an application in clinical prediction modelingabstractPartially labeled and skewed datasets are common in many applications including healthcare, due to the high costs and time constraints of data collection and annotation. However, training machine learning classifiers on such data can undermine their prediction performances. In this paper, we propose a novel classifier to address this problem by focusing on the Area Under the Curve (AUC), which is widely recognized as a more robust performance metric for skewed datasets than other metrics such as accuracy and error rate. We introduce a new classifier called PSVM-AUC Maximizer (PSVM-AUCMax) which is based on Proximal Support Vector Machines (PSVM) and directly maximizes a new AUC-based metric in its learning objective. PSVM-AUCMax has several merits. First, by directly integrating the maximization of the proposed AUC-based metric, PSVM-AUCMax can be proved to have the enhanced generalization capability on the partially labeled and skewed dataset. Second, it simplifies the model selection process with fewer tuning hyperparameters. Third, PSVM-AUCMax’s analytical solution remains the same form as traditional PSVM, preserving its advantages such as fast incremental updating in incremental learning scenarios. The efficacy of PSVM-AUCMax has been demonstrated through extensive experiments on several public datasets and a healthcare case study using data collected at the US Mayo Clinic. In the healthcare case study, we utilized PSVM-AUCMax to develop a clinical prediction model for forecasting composite outcomes in hospitalized COVID-19 patients which yielded promising results. Guanjin Wang, Stephen Wai Hang Kwok, Daniel Axford, Muhammed Yousufuddin, Ferdous Sohel |
Knowl. Based Syst. | 5 |
| 2023 | Cross domain 2D-3D descriptor matching for unconstrained 6-DOF pose estimationabstractThis paper presents a novel approach for cross-domain descriptor matching between 2D and 3D modalities. The 2D-3D matching is applied to localize 2D images in 3D point clouds. Direct cross-domain matching allows our technique to localize images in any type of 3D point cloud without any constraints on the nature or mechanism by which it is obtained. We propose a learning based framework, called Desc-Matcher, to directly match features between the two modalities. A dataset of 2D and 3D features with corresponding locations in images and point clouds is generated to train the Desc-Matcher. To estimate the pose of an image in any 3D cloud, keypoints and feature descriptors are extracted from the query image and the point cloud. The trained Desc-Matcher is then used to match the features from the image and the point cloud. A robust pose estimator is used to predict the location and orientation of the query image from the corresponding positions of the matched 2D and 3D features. We carried out an extensive evaluation of the proposed method for indoor and outdoor scenarios and with different types of point clouds to verify the feasibility of our approach. Experimental results show that the proposed approach can reliably estimate the 6-DOF poses of query cameras in any type of 3D point cloud with high precision. We achieved average median errors of 1.09cm/0.27∘ and 19cm/0.39∘ on the Stanford and Cambridge datasets, respectively. Uzair Nadeem, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel, Aref Miri Rekavandi, Farid Boussaïd |
Pattern Recognit. | 4 |
| 2023 | A Novel AUC Maximization Imbalanced Learning Approach for Predicting Composite Outcomes in COVID-19 Hospitalized PatientsabstractThe COVID-19 patient data for composite outcome prediction often comes with class imbalance issues, i.e., only a small group of patients develop severe composite events after hospital admission, while the rest do not. An ideal COVID-19 composite outcome prediction model should possess strong imbalanced learning capability. The model also should have fewer tuning hyperparameters to ensure good usability and exhibit potential for fast incremental learning. Towards this goal, this study proposes a novel imbalanced learning approach called Imbalanced maximizing-Area Under the Curve (AUC) Proximal Support Vector Machine (ImAUC-PSVM) by the means of classical PSVM to predict the composite outcomes of hospitalized COVID-19 patients within 30 days of hospitalization. ImAUC-PSVM offers the following merits: (1) it incorporates straightforward AUC maximization into the objective function, resulting in fewer parameters to tune. This makes it suitable for handling imbalanced COVID-19 data with a simplified training process. (2) Theoretical derivations reveal that ImAUC-PSVM has the same analytical solution form as PSVM, thus inheriting the advantages of PSVM for handling incremental COVID-19 cases through fast incremental updating. We built and internally and externally validated our proposed classifier using real COVID-19 patient data obtained from three separate sites of Mayo Clinic in the United States. Additionally, we validated it on public datasets using various performance metrics. Experimental results demonstrate that ImAUC-PSVM outperforms other methods in most cases, showcasing its potential to assist clinicians in triaging COVID-19 patients at an early stage in hospital settings, as well as in other prediction applications. Guanjin Wang, Stephen Wai Hang Kwok, Muhammed Yousufuddin, Ferdous Sohel |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Adversary Distillation for One-Shot Attacks on 3D Target TrackingabstractConsidering the vulnerability of existing deep models in the adversarial scenario, the robustness of 3D target tracking is not guaranteed. In this paper, we present an efficient generation based adversarial attack, termed Adversary Distillation Network (AD-Net), which is able to distract a victim tracker in a single shot. In contrast to existing adversarial attacks derived from point perturbations, the proposed method designs a generative network to distill an adversarial example from a tracking template through point-wise filtration. A binary distribution encoding layer is specialized to filter points, which is modeled as a Bernoulli distribution and approximated in a differentiable formulation. To boost the performance of adversarial example generation, a feature extraction module is deployed, which leverages the PointNet++ architecture to learn hierarchical features for the template points as well as similarities with the search areas. Experimental results on the KITTI vision benchmark show that the proposed adversarial attack can effectively mislead popular deep 3D trackers. Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun |
ICASSP | 3 |
| 2022 | Erratum to "Progressive conditional GAN-based augmentation for 3D object recognition" [Neurocomputing 460 (2021) 20-30]
A. A. M. Muzahid, Ferdous Sohel, Mohammed Bennamoun, Hidayat Ullah |
Neurocomputing | 3 |
| 2022 | A reinforcement learning-based approach for imputing missing dataabstractAbstract Missing data is a major problem in real-world datasets, which hinders the performance of data analytics. Conventional data imputation schemes such as univariate single imputation replace missing values in each column with the same approximated value. These univariate single imputation techniques underestimate the variance of the imputed values. On the other hand, multivariate imputation explores the relationships between different columns of data, to impute the missing values. Reinforcement Learning (RL) is a machine learning paradigm where the agent learns by taking actions and receiving rewards in response, to achieve its goal. In this work, we propose an RL-based approach to impute missing data by learning a policy to impute data through an action-reward-based experience. Our approach imputes missing values in a column by working only on the same column (similar to univariate single imputation) but imputes the missing values in the column with different values thus keeping the variance in the imputed values. We report superior performance of our approach, compared with other imputation techniques, on a number of datasets. Saqib Ejaz Awan, Mohammed Bennamoun, Ferdous Sohel, Frank M. Sanfilippo, Girish Dwivedi |
Neural Comput. Appl. | 3 |
| 2022 | Integrated generalized zero-shot learning for fine-grained classification
Tasfia Shermin, Shyh Wei Teng, Ferdous Sohel, M. Manzur Murshed, Guojun Lu |
Pattern Recognit. | 3 |
| 2022 | Bidirectional Mapping Coupled GAN for Generalized Zero-Shot LearningabstractBidirectional mapping-based generalized zero-shot learning (GZSL) methods rely on the quality of synthesized features to recognize seen and unseen data. Therefore, learning a joint distribution of seen-unseen classes and preserving the distinction between seen-unseen classes is crucial for GZSL methods. However, existing methods only learn the underlying distribution of seen data, although unseen class semantics are available in the GZSL problem setting. Most methods neglect retaining seen-unseen classes distinction and use the learned distribution to recognize seen and unseen data. Consequently, they do not perform well. In this work, we utilize the available unseen class semantics alongside seen class semantics and learn joint distribution through a strong visual-semantic coupling. We propose a bidirectional mapping coupled generative adversarial network (BMCoGAN) by extending the concept of the coupled generative adversarial network into a bidirectional mapping model. We further integrate a Wasserstein generative adversarial optimization to supervise the joint distribution learning. We design a loss optimization for retaining distinctive information of seen-unseen classes in the synthesized features and reducing bias towards seen classes, which pushes synthesized seen features towards real seen features and pulls synthesized unseen features away from real seen features. We evaluate BMCoGAN on benchmark datasets and demonstrate its superior performance against contemporary methods. Tasfia Shermin, Shyh Wei Teng, Ferdous Sohel, M. Manzur Murshed, Guojun Lu |
IEEE Trans. Image Process. | 3 |
| 2021 | Leveraging Auxiliary Tasks with Affinity Learning for Weakly Supervised Semantic SegmentationabstractSemantic segmentation is a challenging task in the absence of densely labelled data. Only relying on class activation maps (CAM) with image-level labels provides deficient segmentation supervision. Prior works thus consider pre-trained models to produce coarse saliency maps to guide the generation of pseudo segmentation labels. However, the commonly used off-line heuristic generation process cannot fully exploit the benefits of these coarse saliency maps. Motivated by the significant inter-task correlation, we propose a novel weakly supervised multi-task framework termed as AuxSegNet, to leverage saliency detection and multi-label image classification as auxiliary tasks to improve the primary task of semantic segmentation using only image-level ground-truth labels. Inspired by their similar structured semantics, we also propose to learn a cross-task global pixellevel affinity map from the saliency and segmentation representations. The learned cross-task affinity can be used to refine saliency predictions and propagate CAM maps to provide improved pseudo labels for both tasks. The mutual boost between pseudo label updating and cross-task affinity learning enables iterative improvements on segmentation performance. Extensive experiments demonstrate the effectiveness of the proposed auxiliary learning network structure and the cross-task affinity learning method. The proposed approach achieves state-of-the-art weakly supervised segmentation performance on the challenging PASCAL VOC 2012 and MS COCO benchmarks.1 Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel, Dan Xu 0002 |
ICCV | 5 |
| 2021 | PD-Net: Point Dropping Network for Flexible Adversarial Example Generation with $L_{0}$ RegularizationabstractIt is a challenging task to generate adversarial point clouds, considering the irregular structure of a point cloud, the large search space, and the requirement of imperception to humans. In this paper, a flexible adversarial point cloud generation method, named Point Dropping Network (PD-Net), is proposed, which can be trained to craft adversarial examples in a single forward pass. The network is designed to launch untargeted black-box attacks to deep 3D models through point dropping regularized by the$L_{0}$norm, in contrast to the widely adopted point perturbation methods. To enable incorporation into a deep neural network, the probability of a point to be dropped, which can be described by a Bernoulli distribution, is approximated by a hard concrete distribution. The network of PD-Net consists of an encoder and a decoder, where the former encodes geometric information of each point and the latter learns to drop points from their local features in an unsupervised way. Experiments on two popular deep 3D models (including PointNet and PointNet++) show that the proposed PD-Net degrades the recognition accuracy to a large extent and achieves a high flexibility at the same time. Xupeng Wang 0001, Ferdous Sohel |
IJCNN | 3 |
| 2021 | Imputation of missing data with class imbalance using conditional generative adversarial networks
Saqib Ejaz Awan, Mohammed Bennamoun, Ferdous Sohel, Frank M. Sanfilippo, Girish Dwivedi |
Neurocomputing | 3 |
| 2021 | Progressive conditional GAN-based augmentation for 3D object recognition
A. A. M. Muzahid, Ferdous Sohel, Mohammed Bennamoun, Hidayat Ullah |
Neurocomputing | 3 |
| 2021 | Adversarial point cloud perturbations against 3D object detection in autonomous driving systems
Xupeng Wang 0001, Mumuxin Cai, Ferdous Sohel, Nan Sang, Zhengwei Chang |
Neurocomputing | 3 |
| 2021 | Atrous convolutional feature network for weakly supervised semantic segmentation
Lian Xu, Hao Xue 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel |
Neurocomputing | 5 |
| 2021 | Real time surveillance for low resolution and limited data scenarios: An image set classification approach
Uzair Nadeem, Syed Afaq Ali Shah, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
Inf. Sci. | 5 |
| 2021 | Anomaly detection in a forensic timeline with deep autoencoders
Hudan Studiawan, Ferdous Sohel |
J. Inf. Secur. Appl. | 2 |
| 2021 | Anomaly Detection in Operating System Logs with Deep Learning-Based Sentiment AnalysisabstractThe purpose of sentiment analysis is to detect an opinion or polarity in text data. We can apply such an analysis to detect negative sentiment, which represents the anomalous activities in operating system (OS) logs. Existing methods involve manual searching, predefined rules, or traditional machine learning techniques to detect such suspicious events. In this article, we propose a novel deep learning-based sentiment analysis technique to check whether there are anomalous activities in OS logs. Log messages are modeled as sentences and we identify the sentiments using the gated recurrent unit (GRU) networks. OS log datasets inherently have a class imbalance in the sense that the number of negative sentiment is much lower than that of the number of positive ones. In order to address the class imbalance, we build a GRU layer on top of a class imbalance solver using the Tomek link method. Experimental results demonstrate that the proposed method can detect anomalous events in OS logs with an overall F1 and accuracy of 99.84 and 99.93 percent, respectively. Hudan Studiawan, Ferdous Sohel, Christian Payne |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Adversarial Network With Multiple Classifiers for Open Set Domain AdaptationabstractDomain adaptation aims to transfer knowledge from a domain with adequate labeled samples to a domain with scarce labeled samples. Prior research has introduced various open set domain adaptation settings in the literature to extend the applications of domain adaptation methods in real-world scenarios. This paper focuses on the type of open set domain adaptation setting where the target domain has both private (‘unknown classes’) label space and the shared (‘known classes’) label space. However, the source domain only has the ‘known classes’ label space. Prevalent distribution-matching domain adaptation methods are inadequate in such a setting that demands adaptation from a smaller source domain to a larger and diverse target domain with more classes. For addressing this specific open set domain adaptation setting, prior research introduces a domain adversarial model that uses a fixed threshold for distinguishing known from unknown target samples and lacks at handling negative transfers. We extend their adversarial model and propose a novel adversarial domain adaptation model with multiple auxiliary classifiers. The proposed multi-classifier structure introduces a weighting module that evaluates distinctive domain characteristics for assigning the target samples with weights which are more representative to whether they are likely to belong to the known and unknown classes to encourage positive transfers during adversarial training and simultaneously reduces the domain gap between the shared classes of the source and target domains. A thorough experimental investigation shows that our proposed method outperforms existing domain adaptation methods on a number of domain adaptation datasets. Tasfia Shermin, Guojun Lu, Shyh Wei Teng, M. Manzur Murshed, Ferdous Sohel |
IEEE Trans. Multim. | 5 |
| 2020 | RCNN for Region of Interest Detection in Whole Slide Images
Anupiya Nugaliyadde, Kevin Kok Wai Wong, Jeremy Parry, Ferdous Sohel, Hamid Laga, Upeka Somaratne, Chris Yeomans, Orchid Foster |
ICONIP (5) | 4 |
| 2020 | ResFeats: Residual network based features for underwater image classification
Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
Image Vis. Comput. | 4 |
| 2020 | Sound Event Detection Using Multiple Optimized KernelsabstractSound event detection (SED) has been widely applied in real world applications. Convolutional recurrent neural network based SED approaches have achieved state-of-the-art performance. However, the convolution process is typically performed by using a fixed sized kernel, which adversely affects the detection accuracy especially when the acoustic features of different event classes are characterized by high variations. To deal with this, this article proposes a sound event detection technique using a convolutional recurrent neural network framework with multiple convolutional kernels of different sizes. The top performing kernels are selected from a kernel pool based on the unsupervised clustering errors and the accuracies of the temporarily trained models. Afterwards, the selected kernels are fed to multiple convolution layers to deal with the acoustic feature variations. Experimental results on different subsets of AudioSet, namely the DCASE Challenge 2017 Task 4 and DCASE Challenge 2018 Task 4, demonstrate the performance of the proposed approach compared to state-of-the-art systems. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Learning Latent Global Network for Skeleton-Based Action PredictionabstractHuman actions represented with 3D skeleton sequences are robust to clustered backgrounds and illumination changes. In this paper, we investigate skeleton-based action prediction, which aims to recognize an action from a partial skeleton sequence that contains incomplete action information. We propose a new Latent Global Network based on adversarial learning for action prediction. We demonstrate that the proposed network provides latent long-term global information that is complementary to the local action information of the partial sequences and helps improve action prediction. We show that action prediction can be improved by combining the latent global information with the local action information. We test the proposed method on three challenging skeleton datasets and report state-of-the-art performance. Qiuhong Ke, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Image Process. | 5 |
| 2020 | Multi-Task Learning for Acoustic Event Detection Using Event and Frame Position InformationabstractAcoustic event detection deals with the acoustic signals to determine the sound type and to estimate the audio event boundaries. Multi-label classification based approaches are commonly used to detect the frame wise event types with a median filter applied to determine the happening acoustic events. However, the multi-label classifiers are trained only on the acoustic event types ignoring the frame position within the audio events. To deal with this, this paper proposes to construct a joint learning based multi-task system. The first task performs the acoustic event type detection and the second task is to predict the frame position information. By sharing representations between the two tasks, we can enable the acoustic models to generalize better than the original classifier by averaging respective noise patterns to be implicitly regularized. Experimental results on the monophonic UPC-TALP and the polyphonic TUT Sound Event datasets demonstrate the superior performance of the joint learning method by achieving lower error rate and higher F-score compared to the baseline AED system. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
IEEE Trans. Multim. | 3 |
| 2019 | Automatic Graph-Based Clustering for Security Logs
Hudan Studiawan, Christian Payne, Ferdous Sohel |
AINA | 3 |
| 2019 | An Improved Approach to Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with image-level labels is of great significance since it alleviates the dependency on dense annotations. However, it is a challenging task as it aims to achieve a mapping from high-level semantics to low-level features. In this work, we propose a three-step method to bridge this gap. First, we rely on the interpretable ability of deep neural networks to generate attention maps with class localization information by back-propagating gradients. Secondly, we employ an off-the-shelf object saliency detector with an iterative erasing strategy to obtain saliency maps with spatial extent information of objects. Finally, we combine these two complementary maps to generate pseudo ground-truth images for the training of the segmentation network. With the help of the pre-trained model on the MS-COCO dataset and a multi-scale fusion method, we obtained mIoU of 62.1% and 63.3% on PASCAL VOC 2012 val and test sets, respectively, achieving new state-of-the-art results for the weakly supervised semantic segmentation task. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel |
ICASSP | 5 |
| 2019 | Attention-Based Image Captioning Using DenseNet Features
Ferdous Sohel, Mohd Fairuz Shiratuddin, Hamid Laga, Mohammed Bennamoun |
ICONIP (5) | 2 |
| 2019 | MON: Multiple Output Neurons
Yasir Jan, Ferdous Sohel, Mohd Fairuz Shiratuddin, Kevin Kok Wai Wong |
ICONIP (5) | 2 |
| 2019 | Learning-Based Confidence Estimation for Multi-modal Classifier Fusion
Uzair Nadeem, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
ICONIP (2) | 3 |
| 2019 | Direct Image to Point Cloud Descriptors Matching for 6-DOF Camera Localization in Dense 3D Point Clouds
Uzair Nadeem, Mohammad A. A. K. Jalwana, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
ICONIP (2) | 5 |
| 2019 | Language Modeling through Long-Term Memory NetworkabstractRecurrent Neural Networks (RNN), Long Short-Term Memory Networks (LSTM), and Memory Networks which contain memory are popularly used to learn patterns in sequential data. Sequential data has long sequences that hold relationships. RNN can handle long sequences but suffers from the vanishing and exploding gradient problems. While LSTM and other memory networks address this problem, they are not capable of handling long sequences (50 or more data points long sequence patterns). Language modelling requiring learning from longer sequences are affected by the need for more information in memory. This paper introduces Long Term Memory network (LTM), which can tackle the exploding and vanishing gradient problems and handles long sequences without forgetting. LTM is designed to scale data in the memory and gives a higher weight to the input in the sequence. LTM avoid overfitting by scaling the cell state after achieving the optimal results. The LTM is tested on Penn treebank dataset, and Text8 dataset and LTM achieves test perplexities of 83 and 82 respectively. 650 LTM cells achieved a test perplexity of 67 for Penn treebank, and 600 cells achieved a test perplexity of 77 for Text8. LTM achieves state of the art results by only using ten hidden LTM cells for both datasets. Anupiya Nugaliyadde, Ferdous Sohel, Kevin Kok Wai Wong, Hong Xie 0003 |
IJCNN | 2 |
| 2019 | Coral Classification Using DenseNet and Cross-modality Transfer LearningabstractCoral classification is a challenging task due to the complex morphology and ambiguous boundaries of corals. This paper investigates the benefits of Densely connected convolutional network (DenseNet) and multi-modal image translation techniques in boosting image classification performance by synthesizing missing fluorescence information. To this end, an imageconditional Generative Adversarial Network (GAN) based image translator is trained to model the relationship between reflectance and fluorescence images. Through this image translator, fluorescence images can be generated from the available reflectance images to provide complementary information. During the classification phase, reflectance and translated fluorescence images are combined to obtain more discriminative representations and produce improved classification performance. We present results on the EFC and MLC datasets and report state-of-the-art coral classification performance. Lian Xu, Mohammed Bennamoun, Farid Boussaïd, Senjian An, Ferdous Sohel |
IJCNN | 5 |
| 2019 | Body Detection in Spectator Crowd Images Using Partial Heads
Yasir Jan, Ferdous Sohel, Mohd Fairuz Shiratuddin, Kevin Kok Wai Wong |
PSIVT | 2 |
| 2019 | Enhanced Transfer Learning with ImageNet Trained Classification Layer
Tasfia Shermin, Shyh Wei Teng, M. Manzur Murshed, Guojun Lu, Ferdous Sohel, Manoranjan Paul |
PSIVT | 5 |
| 2019 | Deep learning-based 3D local feature descriptor from Mercator projections
Masoumeh Rezaei, Mehdi Rezaeian, Vali Derhami, Ferdous Sohel, Mohammed Bennamoun |
Comput. Aided Geom. Des. | 4 |
| 2019 | NormalNet: A voxel-based CNN for 3D object classification and retrieval
Cheng Wang 0003, Ming Cheng 0002, Ferdous Sohel, Mohammed Bennamoun, Jonathan Li 0001 |
Neurocomputing | 3 |
| 2019 | Auxiliary Classifier Generative Adversarial Network With Soft Labels in Imbalanced Acoustic Event DetectionabstractIn acoustic event detection, the training data size of some acoustic events is often small and imbalanced. To deal with this, this paper proposes generating the virtual training data categorically using the auxiliary classifier generative adversarial networks. Soft labels of acoustic events are first calculated to represent the acoustic event localization information. The closer the current frame is to the middle of the manually labeled acoustic event, the higher the soft label will be, which makes the soft labels positively correlated with the acoustic event localization. Then, the acoustic event class and the quantized soft labels are used as the input condition to the auxiliary classifier generative adversarial networks to generate an arbitrary number of training samples. Experimental results on the TUT Sound Event 2016 under the home environment and TUT Sound Event 2017 under the street environment demonstrate the improved performance of the proposed technique compared to existing acoustic event detection systems. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
IEEE Trans. Multim. | 3 |
| 2018 | Global Regularizer and Temporal-Aware Cross-Entropy for Skeleton-Based Early Action Recognition
Qiuhong Ke, Jun Liu 0036, Mohammed Bennamoun, Hossein Rahmani 0001, Senjian An, Ferdous Sohel, Farid Boussaïd |
ACCV (4) | 6 |
| 2018 | Acoustic Scene Classification Using Joint Time-Frequency Image-Based Feature RepresentationsabstractThe classification of acoustic scenes is important in emerging applications such as automatic audio surveillance, machine listening and multimedia content analysis. In this paper, we present an approach for acoustic scene classification by using joint time-frequency image-based feature representations. In acoustic scene classification, joint time-frequency representation (TFR) is shown to better represent important information across a wide range of low and middle frequencies in the audio signal. The audio signal is converted to Constant-Q Transform (CQT) and Mel-spectrum TFRs and local binary patterns (LBP) are used to extract the features from these TFRs. To ensure localized spectral information is not lost, the TFRs are divided into a number of zones. Then, we perform score level fusion to further improve the classification performance accuracy. Our technique achieves a competitive performance with a classification accuracy of 83.4% on the DCASE 2016 development dataset compared to the existing current state of the art. Shamsiah Abidin, Roberto Togneri, Ferdous Sohel |
AVSS | 3 |
| 2018 | Confidence Based Acoustic Event DetectionabstractAcoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt the multi-label classification technique to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the manually labeled boundaries are error-prone and cannot always be accurate, especially when the frame length is too short to be accurately labeled by human annotators. To deal with this, a confidence is assigned to each frame and acoustic event detection is performed using a multi-variable regression approach in this paper. Experimental results on the latest TUT sound event 2017 database of polyphonic events demonstrate the superior performance of the proposed approach compared to the multi-label classification based AED method. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
ICASSP | 3 |
| 2018 | Classification of Corals in Reflectance and Fluorescence Images Using Convolutional Neural Network RepresentationsabstractCoral species, with complex morphology and ambiguous boundaries, pose a great challenge for automated classification. CNN activations, which are extracted from fully connected layers of deep networks (FC features), have been successfully used as powerful universal representations in many visual tasks. In this paper, we investigate the transferability and combined performance of FC features and CONY features (extracted from convolutional layers) in the coral classification of two image modalities (reflectance and fluorescence), using a typical deep network (e.g. VGGNet). We exploit vector of locally aggregated descriptors (VLAD) encoding and principal component analysis (PCA) to compress dense CONY features into a compact representation. Experimental results demonstrate that encoded CONV3 features achieve superior performances on reflectance and fluorescence coral images, compared to FC features. The combination of these two features further improves the overall accuracy and achieves state-of-the-art performance on the challenging EFC dataset. Lian Xu, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
ICASSP | 4 |
| 2018 | Local Binary Pattern with Random Forest for Acoustic Scene ClassificationabstractThis paper presents an approach for acoustic scene classification using the local binary pattern (LBP) and random forest (RF). The audio signal is converted to a Constant-Q transform (CQT) representation and LBP is used to extract the features from this time-frequency representation. The CQT representations are divided into a number of sub-bands to obtain more localized features relevant to the spectral information. We then use random forest to select the most important features for each band of extracted LBP features. For further performance enhancement, we use feature level fusion of LBP and HOG features. The proposed system has achieved an accuracy of 85% on the DCASE 2016 dataset. Shamsiah Abidin, Xianjun Xia, Roberto Togneri, Ferdous Sohel |
ICME | 4 |
| 2018 | Exploiting layerwise convexity of rectifier networks with sign constrained weights
Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel |
Neural Networks | 4 |
| 2018 | A Multi-Modal, Discriminative and Spatially Invariant CNN for RGB-D Object LabelingabstractWhile deep convolutional neural networks have shown a remarkable success in image classification, the problems of inter-class similarities, intra-class variances, the effective combination of multi-modal data, and the spatial variability in images of objects remain to be major challenges. To address these problems, this paper proposes a novel framework to learn a discriminative and spatially invariant classification model for object and indoor scene recognition using multi-modal RGB-D imagery. This is achieved through three postulates: 1) spatial invariance $-$ this is achieved by combining a spatial transformer network with a deep convolutional neural network to learn features which are invariant to spatial translations, rotations, and scale changes, 2) high discriminative capability $-$ this is achieved by introducing Fisher encoding within the CNN architecture to learn features which have small inter-class similarities and large intra-class compactness, and 3) multi-modal hierarchical fusion$-$ this is achieved through the regularization of semantic segmentation to a multi-modal CNN architecture, where class probabilities are estimated at different hierarchical levels (i.e., image- and pixel-levels), and fused into a Conditional Random Field (CRF)-based inference hypothesis, the optimization of which produces consistent class labels in RGB-D images. Extensive experimental evaluations on RGB-D object and scene datasets, and live video streams (acquired from Kinect) show that our framework produces superior object and scene classification results compared to the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Random forest classification based acoustic event detection utilizing contextual-information and bottleneck features
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
Pattern Recognit. | 3 |
| 2018 | Spectrotemporal Analysis Using Local Binary Pattern Variants for Acoustic Scene ClassificationabstractIn this paper, we present an approach for acoustic scene classification, which aggregates spectral and temporal features. We do this by proposing the first use of the variable-Q transform (VQT) to generate the time-frequency representation for acoustic scene classification. The VQT provides finer control over the resolution compared to the constant-Q transform (CQT) or short time fourier transform and can be tuned to better capture acoustic scene information. We then adopt a variant of the local binary pattern (LBP), the adjacent evaluation completed LBP (AECLBP), which is better suited to extracting features from acoustic time-frequency images. Our results yield a 5.2% improvement on the DCASE 2016 dataset compared to the application of standard CQT with LBP. Fusing our proposed AECLBP with HOG features, we achieve a classification accuracy of 85.5%, which outperforms one of the top performing systems. Shamsiah Abidin, Roberto Togneri, Ferdous Sohel |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Learning Clip Representations for Skeleton-Based 3D Action RecognitionabstractThis paper presents a new representation of skeleton sequences for 3D action recognition. Existing methods based on hand-crafted features or recurrent neural networks cannot adequately capture the complex spatial structures and the long-term temporal dynamics of the skeleton sequences, which are very important to recognize the actions. In this paper, we propose to transform each channel of the 3D coordinates of a skeleton sequence into a clip. Each frame of the generated clip represents the temporal information of the entire skeleton sequence and one particular spatial relationship between the skeleton joints. The entire clip incorporates multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We also propose a multitask convolutional neural network (MTCNN) to learn the generated clips for action recognition. The proposed MTCNN processes all the frames of the generated clips in parallel to explore the spatial and temporal information of the skeleton sequences. The proposed method has been extensively tested on six challenging benchmark datasets. Experimental results consistently demonstrate the superiority of the proposed clip representation and the feature learning method for 3D action recognition compared to the existing techniques. Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Image Process. | 4 |
| 2018 | Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction PredictionabstractPredicting an interaction before it is fully executed is very important in applications, such as human-robot interaction and video surveillance. In a two-human interaction scenario, there are often contextual dependency structures between the global interaction context of the two humans and the local context of the different body parts of each human. In this paper, we propose to learn the structure of the interaction contexts and combine it with the spatial and temporal information of a video sequence to better predict the interaction class. The structural models, including the spatial and the temporal models, are learned with long short term memory (LSTM) networks to capture the dependency of the global and local contexts of each RGB frame and each optical flow image, respectively. LSTM networks are also capable of detecting the key information from global and local interaction contexts. Moreover, to effectively combine the structural models with the spatial and temporal models for interaction prediction, a ranking score fusion method is introduced to automatically compute the optimal weight of each model for score fusion. Experimental results on the BIT-Interaction Dataset and the UT-Interaction Dataset clearly demonstrate the benefits of the proposed method. Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
IEEE Trans. Multim. | 4 |
| 2018 | Cost-Sensitive Learning of Deep Feature Representations From Imbalanced DataabstractClass imbalance is a common problem in the case of real-world object detection and classification tasks. Data of some classes are abundant, making them an overrepresented majority, and data of other classes are scarce, making them an underrepresented minority. This imbalance makes it challenging for a classifier to appropriately learn the discriminating boundaries of the majority and minority classes. In this paper, we propose a cost-sensitive (CoSen) deep neural network, which can automatically learn robust feature representations for both the majority and minority classes. During training, our learning procedure jointly optimizes the class-dependent costs and the neural network parameters. The proposed approach is applicable to both binary and multiclass problems without any modification. Moreover, as opposed to data-level approaches, we do not alter the original data distribution, which results in a lower computational cost during the training process. We report the results of our experiments on six major image classification data sets and show that the proposed approach significantly outperforms the baseline algorithms. Comparisons with popular data sampling techniques and CoSen classifiers demonstrate the superior performance of our proposed method. Salman Khan 0001, Munawar Hayat, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | A New Representation of Skeleton Sequences for 3D Action RecognitionabstractThis paper presents a new method for 3D action recognition with skeleton sequences (i.e., 3D trajectories of human skeleton joints). The proposed method first transforms each skeleton sequence into three clips each consisting of several frames for spatial temporal feature learning using deep neural networks. Each clip is generated from one channel of the cylindrical coordinates of the skeleton sequence. Each frame of the generated clips represents the temporal information of the entire skeleton sequence, and incorporates one particular spatial relationship between the joints. The entire clips include multiple frames with different spatial relationships, which provide useful spatial structural information of the human skeleton. We propose to use deep convolutional neural networks to learn long-term temporal information of the skeleton sequence from the frames of the generated clips, and then use a Multi-Task Learning Network (MTLN) to jointly process all frames of the clips in parallel to incorporate spatial structural information for action recognition. Experimental results clearly show the effectiveness of the proposed new representation and feature learning method for 3D action recognition. Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd |
CVPR | 4 |
| 2017 | Enhanced LBP texture features from time frequency representations for acoustic scene classificationabstractThis paper introduces the use of local binary patterns (LBP) extracted from a time-frequency representation (TFR) for acoustic scene classification. As LBP provides a description of the global TFR texture we propose a novel zoning mechanism that provides a simple solution to extract spectrally relevant local features which better characterize the audio TFRs. To further improve the classification performance, we perform feature and score level fusion of the proposed LBP (with zoning) with histogram of gradients (HOG) of the TFR images. Our technique demonstrates an improved performance by achieving a classification accuracy of 95.2% using a fusion of time-frequency derived features. Shamsiah Abidin, Roberto Togneri, Ferdous Sohel |
ICASSP | 3 |
| 2017 | Resfeats: Residual network based features for image classificationabstractDeep residual networks have recently emerged as the state-of-the-art architecture in image classification and object detection. In this paper, we propose new image features (called ResFeats) extracted from the last convolutional layer of the deep residual networks pre-trained on ImageNet. We propose to use ResFeats for diverse image classification tasks namely, object classification, scene classification and coral classification and show that ResFeats consistently perform better than their CNN counterparts on these classification tasks. Since the ResFeats are large feature vectors, we explore dimensionality reduction methods. Experimental results are provided to show the effectiveness of ResFeats with state-of-the-art classification accuracies on Caltech-101, Caltech-256 and MLC datasets and a significant performance improvement on MIT-67 dataset compared to the widely used CNN features. Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel |
ICIP | 4 |
| 2017 | Random forest regression based acoustic event detection with bottleneck featuresabstractThis paper deals with random forest regression based acoustic event detection (AED) by combining acoustic features with bottleneck features (BN). The bottleneck features have a good reputation of being inherently discriminative in acoustic signal processing. To deal with the unstructured and complex real-world acoustic events, an acoustic event detection system is constructed using bottleneck features combined with acoustic features. Evaluations were carried out on the UPC-TALP and ITC-Irst databases which consist of highly variable acoustic events. Experimental results demonstrate the usefulness of the low-dimensional and discriminative bottleneck features with relative 5.33% and 5.51% decreases in error rates respectively. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
ICME | 3 |
| 2017 | Random forest classification based acoustic event detectionabstractThis paper deals with the acoustic event detection (AED) to improve the detection accuracy of acoustic events. Acoustic event detection task is performed by a regression via classification (RvC) based approach along with the random forest technique. A discretization process is used to convert the continuous frame positions within acoustic events into event duration class labels. Outputs of the category-specific random forest classifiers are then reversed back to the event boundary information. Evaluations on the UPC-TALP database which consists of highly variable acoustic events demonstrate the efficiency of the proposed approaches with improvements in detection error rate compared to the best baseline system. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
ICME | 3 |
| 2017 | Reinforced Memory Network for Question Answering
Anupiya Nugaliyadde, Kevin Kok Wai Wong, Ferdous Sohel, Hong Xie 0003 |
ICONIP (2) | 3 |
| 2017 | Learning deep structured network for weakly supervised change detectionabstractConventional change detection methods require a large number of images to learn background models or depend on tedious pixel-level labeling by humans. In this paper, we present a weakly supervised approach that needs only image-level labels to simultaneously detect and localize changes in a pair of images. To this end, we employ a deep neural network with DAG topology to learn patterns of change from image-level labeled training data. On top of the initial CNN activations, we define a CRF model to incorporate the local differences and context with the dense connections between individual pixels. We apply a constrained mean-field algorithm to estimate the pixel-level labels, and use the estimated labels to update the parameters of the CNN in an iterative EM framework. This enables imposing global constraints on the observed foreground probability mass function. Our evaluations on four benchmark datasets demonstrate superior detection and localization performance. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IJCAI | 5 |
| 2017 | Frame-Wise Dynamic Threshold Based Polyphonic Acoustic Event DetectionabstractAcoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt multi-label classification techniques to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the global threshold has to be set manually and is highly dependent on the database being tested. To deal with this, we replaced the fixed threshold method with a frame-wise dynamic threshold approach in this paper. Two novel approaches, namely contour and regressor based dynamic threshold approaches are proposed in this work. Experimental results on the popular TUT Acoustic Scenes 2016 database of polyphonic events demonstrated the superior performance of the proposed approaches. Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang |
INTERSPEECH | 3 |
| 2017 | Discriminative feature learning and region consistency activation for robust scene labeling
Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
Neurocomputing | 2 |
| 2017 | Scale space clustering evolution for salient region detection on 3D deformable shapes
Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun, Yulan Guo |
Pattern Recognit. | 2 |
| 2017 | SkeletonNet: Mining Deep Part Features for 3-D Action RecognitionabstractThis letter presents SkeletonNet, a deep learning framework for skeleton-based 3-D action recognition. Given a skeleton sequence, the spatial structure of the skeleton joints in each frame and the temporal information between multiple frames are two important factors for action recognition. We first extract body-part-based features from each frame of the skeleton sequence. Compared to the original coordinates of the skeleton joints, the proposed features are translation, rotation, and scale invariant. To learn robust temporal information, instead of treating the features of all frames as a time series, we transform the features into images and feed them to the proposed deep learning network, which contains two parts: one to extract general features from the input images, while the other to generate a discriminative and compact representation for action recognition. The proposed method is tested on the SBU kinect interaction dataset, the CMU dataset, and the large-scale NTU RGB+D dataset and achieves state-of-the-art performance. Qiuhong Ke, Senjian An, Mohammed Bennamoun, Ferdous Sohel, Farid Boussaïd |
IEEE Signal Process. Lett. | 4 |
| 2017 | A Joint Deep Boltzmann Machine (jDBM) Model for Person Identification Using Mobile Phone DataabstractWe propose an audio-visual person identification approach based on a joint deep Boltzmann machine (jDBM) model. The proposed jDBM model is trained in three steps: 1) learning the unimodal DBM models corresponding to the speech and facial image modalities, 2) learning the shared layer parameters using a joint restricted Boltzmann machine (jRBM) model, and 3) the fine-tuning of the jDBM model after the initialization with the parameters of the unimodal DBMs and the shared layer. The activation probabilities of the units of the shared layer are used as the joint features and a logistic regression classifier is used for the combined speech and facial image recognition. We show that by learning the shared layer parameters using a jRBM, a higher accuracy can be achieved compared to the greedy layer-wise initialization. The performance of our proposed model is also compared with a state-of-the art support vector machine (SVM), deep belief network (DBN), and the deep auto-encoder (DAE) models. In addition, our experimental results show that the joint representations obtained from the proposed jDBM model are robust to noise and missing information. Experiments were carried out on the challenging MOBIO database, which includes audio-visual data captured using mobile phones. Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
IEEE Trans. Multim. | 4 |
| 2017 | RGB-D Object Recognition and Grasp Detection Using Hierarchical Cascaded ForestsabstractThis paper presents an efficient framework to perform recognition and grasp detection of objects from RGB-D images of real scenes. The framework uses a novel architecture of hierarchical cascaded forests, in which object-class and grasp-pose probabilities are computed at different levels of an image hierarchy (e.g., patch and object levels) and fused to infer the class and the grasp of unseen objects. We introduce a novel training objective function that minimizes the uncertainties of the class labels and the grasp ground truths at the leaves of the forests, thereby enabling the framework to perform the recognition and grasp detection of objects. Our objective function is learned from features that are extracted from RGB-D point clouds of the objects. For that, we propose a novel method to encode an RGB-D point cloud into a representation that facilitates the use of large convolution neural networks to extract discriminative features from RGB-D images. We evaluate our framework on challenging object datasets, where we demonstrate that our framework outperforms the state-of-the-art methods in terms of object-recognition and grasp-detection accuracies. We also show experiments by using live video streams from a Kinect mounted on our in-house robotic platform. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IEEE Trans. Robotics | 3 |
| 2016 | Coral classification with hybrid feature representationsabstractCoral reefs exhibit significant within-class variations, complex between-class boundaries and inconsistent image clarity. This makes coral classification a challenging task. In this paper, we report the application of generic CNN representations combined with hand-crafted features for coral reef classification to take advantage of the complementary strengths of these representation types. We extract CNN based features from patches centred at labelled pixels at multiple scales. We use texture and color based hand-crafted features extracted from the same patches to complement the CNN features. Our proposed method achieves a classification accuracy that is higher than the state-of-art methods on the MLC benchmark dataset for corals. Ammar Mahmood, Mohammed Bennamoun, Senjian An, Ferdous Sohel, Farid Boussaïd, Renae Hovey, Gary A. Kendrick, Robert B. Fisher |
ICIP | 4 |
| 2016 | Audio-visual biometric recognition via joint sparse representationsabstractIn this paper we present a novel audio-visual (AV) person identification system based on joint sparse representation. Video features used were vectorized raw pixel values, while i-vectors were used as the audio features. Classification is performed by solving the joint sparsity optimization problem, and fusion is carried out by using the quality (confidence) assigned to each matcher. Our experimental results on the challenging MOBIO database using 100 subjects show that the system based on joint sparse representation outperforms the system based on separate sparse representations for each modality. Furthermore, we show that our newly introduced quality measure improves the system's performance, when compared to conventionally used quality measures for sparse representation - based systems. Rudi Primorac, Roberto Togneri, Mohammed Bennamoun, Ferdous Sohel |
ICPR | 4 |
| 2016 | Simultaneous dense scene reconstruction and object labelingabstractThis paper presents an efficient system for simultaneous dense scene reconstruction and object labeling in real-world environments (captured with an RGB-D sensor). The proposed system starts with the generation of object proposals in the scene. It then tracks spatio-temporally consistent object proposals across multiple frames and produces a dense reconstruction of the scene. In parallel, the proposed system uses an efficient inference algorithm, where object class probabilities are computed at an object-level and fused into a voxel-based prediction hypothesis modeled on the voxels of the reconstructed scene. Our extensive experiments using challenging RGB-D object and scene datasets, and live video streams from Microsoft Kinect show that the proposed system achieved competitive 3D scene reconstruction and object labeling results compared to the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
ICRA | 3 |
| 2016 | Heat propagation contours for 3D non-rigid shape analysisabstractWe present a novel local shape descriptor by means of General Adaptive Neighborhoods (GANs) based on the properties of the heat diffusion process on a Riemannian manifold. The GAN is a spatial region, surrounding the feature point and fitting its local shape structure, which is isometric. Our signature, called the Heat Propagation Contours (HPCs), is obtained by analysing the well-known heat kernel and extracting contours automatically within the GAN as heat dissipates from the feature point onto the rest of the shape. HPCs capture geometric information around the feature point by investigating the heat propagation process both in the temporal and spatial domain. HPCs share many useful characteristics with the heat based methods. Particularly, it captures the intrinsic geometry of a shape and is suitable for non-rigid shape analysis. In addition, our signature provides an elegant and efficient way to describe the neighborhood of the feature point in a multi-scale approach. The proposed descriptor is evaluated on several datasets to demonstrate its effectiveness. Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun |
WACV | 2 |
| 2016 | A Comprehensive Performance Evaluation of 3D Local Feature Descriptors
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan, Ngai Ming Kwok |
Int. J. Comput. Vis. | 3 |
| 2016 | Integrating Geometrical Context for Semantic Labeling of Indoor Scenes using RGBD Images
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri, Imran Naseem |
Int. J. Comput. Vis. | 3 |
| 2016 | Automatic Shadow Detection and Removal from a Single ImageabstractWe present a framework to automatically detect and remove shadows in real world scenes from a single image. Previous works on shadow detection put a lot of effort in designing shadow variant and invariant hand-crafted features. In contrast, our framework automatically learns the most relevant features in a supervised manner using multiple convolutional deep neural networks (ConvNets). The features are learned at the super-pixel level and along the dominant boundaries in the image. The predicted posteriors based on the learned features are fed to a conditional random field model to generate smooth shadow masks. Using the detected shadow masks, we propose a Bayesian formulation to accurately extract shadow matte and subsequently remove shadows. The Bayesian formulation is based on a novel model which accurately models the shadow generation process in the umbra and penumbra regions. The model parameters are efficiently estimated using an iterative optimization procedure. Our proposed framework consistently performed better than the state-of-the-art on all major shadow databases collected under a variety of conditions. Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | EI3D: Expression-invariant 3D face recognition based on feature and shape matching
Yulan Guo, Yinjie Lei, Li Liu 0002, Yan Wang 0059, Mohammed Bennamoun, Ferdous Sohel |
Pattern Recognit. Lett. | 6 |
| 2016 | A Discriminative Representation of Convolutional Features for Indoor Scene RecognitionabstractIndoor scene recognition is a multi-faceted and challenging problem due to the diverse intra-class variations and the confusing inter-class similarities. This paper presents a novel approach which exploits rich mid-level convolutional features to categorize indoor scenes. Traditionally used convolutional features preserve the global spatial structure, which is a desirable property for general object recognition. However, we argue that this structuredness is not much helpful when we have large variations in scene layouts, e.g., in indoor scenes. We propose to transform the structured convolutional activations to another highly discriminative feature space. The representation in the transformed space not only incorporates the discriminative aspects of the target dataset, but it also encodes the features in terms of the general object categories that are present in indoor scenes. To this end, we introduce a new large-scale dataset of 1300 object categories which are commonly present in indoor scenes. Our proposed approach achieves a significant performance boost over previous state of the art approaches on five major scene classification datasets. Salman Khan 0001, Munawar Hayat, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
IEEE Trans. Image Process. | 5 |
| 2015 | Separating objects and clutter in indoor scenesabstractObjects' spatial layout estimation and clutter identification are two important tasks to understand indoor scenes. We propose to solve both of these problems in a joint framework using RGBD images of indoor scenes. In contrast to recent approaches which focus on either one of these two problems, we perform ‘fine grained structure categorization’ by predicting all the major objects and simultaneously labeling the cluttered regions. A conditional random field model is proposed to incorporate a rich set of local appearance, geometric features and interactions between the scene elements. We take a structural learning approach with a loss of 3D localisation to estimate the model parameters from a large annotated RGBD dataset, and a mixed integer linear programming formulation for inference. We demonstrate that our approach is able to detect cuboids and estimate cluttered regions across many different object and scene categories in the presence of occlusion, illumination and appearance variations. Salman Khan 0001, Xuming He 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
CVPR | 4 |
| 2015 | Contractive Rectifier Networks for Nonlinear Maximum Margin ClassificationabstractTo find the optimal nonlinear separating boundary with maximum margin in the input data space, this paper proposes Contractive Rectifier Networks (CRNs), wherein the hidden-layer transformations are restricted to be contraction mappings. The contractive constraints ensure that the achieved separating margin in the input space is larger than or equal to the separating margin in the output layer. The training of the proposed CRNs is formulated as a linear support vector machine (SVM) in the output layer, combined with two or more contractive hidden layers. Effective algorithms have been proposed to address the optimization challenges arising from contraction constraints. Experimental results on MNIST, CIFAR-10, CIFAR-100 and MIT-67 datasets demonstrate that the proposed contractive rectifier networks consistently outperform their conventional unconstrained rectifier network counterparts. Senjian An, Munawar Hayat, Salman Khan 0001, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel |
ICCV | 6 |
| 2015 | Outdoor scene labelling with learned features and region consistency activationabstractThis paper presents a learned feature based method for scene labelling. This method is combined with a novel strategy to improve global label consistency. We first follow a traditional way to investigate trained features from convolutional neural networks (ConvNets) for scene labelling. Then, motivated by the recent successful use of general features extracted from ConvNets for various applications, we extend the use of the general features to scene labelling (for the first time). We further propose an algorithm called Region Consistency Activation (RCA) to improve the global label consistency. RCA is based on a novel transformation between Ultrametric Contour Map (UCM) and the Probability of Regions Consistency (PRC). Our algorithms were rigorously tested on the popular Stanford Background and SIFT Flow datasets. We achieved superior performances compared with the state-of-the-art methods on both of these datasets. Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
ICIP | 2 |
| 2015 | Efficient RGB-D object categorization using cascaded ensembles of randomized decision treesabstractThis paper presents an efficient framework for the categorization of objects in real-world scenes (captured with an RGB-D sensor). The proposed framework uses ensembles of randomized decision trees in a hierarchical cascaded architecture to compute consistent object-class inferences of unseen objects. Specifically, the proposed framework computes object-class probabilities at three levels of an image hierarchy (i.e., pixel-, surfel-, and object-levels) using Random Forest classifiers. Next, these probabilities are fused together to compute a cumulative probabilistic output which is used to infer object categories. This fusion results in an improved object categorization performance compared with the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
ICRA | 3 |
| 2015 | Discriminative feature learning for efficient RGB-D object recognitionabstractThis paper presents an efficient approach to recognize objects captured with an RGB-D sensor. The proposed approach uses a Bag-of-Words (BOW) model to learn feature representations from raw RGB-D point clouds in a weakly supervised manner. To this end, we introduce a novel method based on randomized clustering trees to learn visual vocabularies which are fast to compute and more discriminative compared to the vocabularies generated by classical methods such as k-means. We show that, when combined with standard spatial pooling strategies, our proposed approach yields a powerful feature representation for RGB-D object recognition. Our extensive experimental evaluation on two challenging RGB-D object datasets and live video streams from Kinect shows that our learned features result in superior object recognition accuracies compared with the state-of-the-art methods. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IROS | 3 |
| 2015 | Sign Constrained Rectifier Networks with Applications to Pattern Decompositions
Senjian An, Qiuhong Ke, Mohammed Bennamoun, Farid Boussaïd, Ferdous Sohel |
ECML/PKDD (1) | 5 |
| 2015 | Deep Boltzmann Machines for i-Vector Based Audio-Visual Person Identification
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
PSIVT | 4 |
| 2015 | Binary Descriptor Based on Heat Diffusion for Non-rigid Shape Analysis
Xupeng Wang 0001, Ferdous Sohel, Mohammed Bennamoun |
PSIVT | 2 |
| 2015 | Heterogeneous Multi-column ConvNets with a Fusion Framework for Object RecognitionabstractThe purpose of this paper is to investigate heterogeneous multi-column ConvNets (MCCNN) and fusion methods for them. We first construct heterogeneous MCCNN by combining ConvNets with different structures. We then use different fusion methods to check their performances to find out the effect of fusion methods for MCCNN. We also propose a novel sliding window based fusion framework which defines a specific subset of columns to be picked up from MCCNN for fusion. Two different strategies (exhaustive sliding window and sliding window from training) are investigated to determine the best performance of the fusion process. We tested the heterogeneous MCCNN and sliding window fusion on the MNIST dataset for optical character recognition. Experiments show that MCCNN improved the accuracy of recognition compared with a single column of ConvNets. Moreover, sliding window fusion is a more generalized fusion method and consistently achieves better results compared with the traditional fusion methods. We also tested the MCCNN and sliding window fusion on CIFAR-10 and Caltech-256 datasets. We achieved superior results compared to existing state-of-the-art techniques. Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
WACV | 2 |
| 2015 | A novel local surface feature for 3D object recognition under clutter and occlusion
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, Min Lu 0001 |
Inf. Sci. | 2 |
| 2015 | A confidence-based late fusion framework for audio-visual biometric identification
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
Pattern Recognit. Lett. | 4 |
| 2015 | Quantitative Error Analysis of Bilateral FilteringabstractOne of the fastest acceleration techniques for bilateral image filtering is the real time O(1) quantization method proposed by Yang 2009, which first computes some Principal Bilateral Filtered Image Components (PBFICs) and then applies linear interpolation to estimate the filtered output images. There is a trade-off between accuracy and efficiency in selecting the number of PBFICs: the more PBFICs are used, the higher the accuracy, and the higher the computational cost. A question arises: how many PBFICs are required to achieve a certain level of accuracy? In this letter, we address this question by investigating the properties of bilateral filtering and deriving the linear interpolation error bounds when only a subset of PBFICs is used. The provided theoretical analysis indicates that the necessary number of PBFICs for user-provided precision depends on the range kernel and, for typical Gaussian range kernels, a small percentage (typically less than 4%) of the PBFICs are enough for good approximations. Senjian An, Farid Boussaïd, Mohammed Bennamoun, Ferdous Sohel |
IEEE Signal Process. Lett. | 4 |
| 2014 | Performance Evaluation of 3D Local Feature Descriptors
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan, Jun Zhang 0044 |
ACCV (2) | 3 |
| 2014 | Automatic Feature Learning for Robust Shadow DetectionabstractWe present a practical framework to automatically detect shadows in real world scenes from a single photograph. Previous works on shadow detection put a lot of effort in designing shadow variant and invariant hand-crafted features. In contrast, our framework automatically learns the most relevant features in a supervised manner using multiple convolutional deep neural networks (ConvNets). The 7-layer network architecture of each ConvNet consists of alternating convolution and sub-sampling layers. The proposed framework learns features at the super-pixel level and along the object boundaries. In both cases, features are extracted using a context aware window centered at interest points. The predicted posteriors based on the learned features are fed to a conditional random field model to generate smooth shadow contours. Our proposed framework consistently performed better than the state-of-the-art on all major shadow databases collected under a variety of conditions. Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
CVPR | 3 |
| 2014 | Model-Free Segmentation and Grasp Selection of Unknown Stacked Objects
Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
ECCV (5) | 3 |
| 2014 | Geometry Driven Semantic Labeling of Indoor Scenes
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
ECCV (1) | 3 |
| 2014 | Confidence-based Rank-level Fusion for Audio-visual Person Identification SystemabstractA multibiometric identification system establishes the identity of a person based on the input biometric data presented to its sub-systems. Each sub-system compares the features extracted from the input against the templates of all identities stored in gallery. The best matched identity is ranked highest in the ranked list. In rank-level fusion, the ranked lists from different sub-systems are combined to reach a final decision. However, the state-of-the-art rank-level fusion methods consider that all sub-systems are equally reliable in terms of classifying the probe data. In practice, the probe data may be affected by different sources of degradation (e.g., illumination and pose variation on the face image, environmental noise) and thus affecting the overall recognition accuracy. In this paper, robust rank-level fusion methods (e.g., confidence based highest rank and Borda count) are proposed by using confidence measures for each sub-system in the decision making process. Experimental results show that the proposed confidence based rank-level fusion achieved higher recognition rates than state-of-the-art rank-level fusion methods. Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
ICPRAM | 4 |
| 2014 | A model-free approach for the segmentation of unknown objectsabstractWe address the problem of object segmentation from depth images of highly complex indoor scenes. We propose a model-free segmentation approach, which robustly separates unknown stacked objects in real-world scenes. Our approach constructs geometrically constrained 3D clusters known as salient-regions, which are subsequently merged into high-level object hypotheses by analyzing the local geometrical characteristics (such as local shape and homogeneity) of the area of their shared boundaries. We tested our approach using depth images from live Kinect video streams and publicly available RGB-D datasets. Our approach is highly efficient and achieves superior performance compared to state-of-the-art techniques. Umar Asif, Mohammed Bennamoun, Ferdous Sohel |
IROS | 3 |
| 2014 | 3D Object Recognition in Cluttered Scenes with Local Surface Features: A Surveyabstract3D object recognition in cluttered scenes is a rapidly growing research area. Based on the used types of features, 3D object recognition methods can broadly be divided into two categories-global or local feature based methods. Intensive research has been done on local surface feature based methods as they are more robust to occlusion and clutter which are frequently present in a real-world scene. This paper presents a comprehensive survey of existing local surface feature based 3D object recognition methods. These methods generally comprise three phases: 3D keypoint detection, local surface feature description, and surface matching. This paper covers an extensive literature survey of each phase of the process. It also enlists a number of popular and contemporary databases together with their relevant attributes. Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | An Accurate and Robust Range Image Registration Algorithm for 3D Object ModelingabstractRange image registration is a fundamental research topic for 3D object modeling and recognition. In this paper, we propose an accurate and robust algorithm for pairwise and multi-view range image registration. We first extract a set of Rotational Projection Statistics (RoPS) features from a pair of range images, and perform feature matching between them. The two range images are then registered using a transformation estimation method and a variant of the Iterative Closest Point (ICP) algorithm. Based on the pairwise registration algorithm, we propose a shape growing based multi-view registration algorithm. The seed shape is initialized with a selected range image and then sequentially updated by performing pairwise registration between itself and the input range images. All input range images are iteratively registered during the shape growing process. Extensive experiments were conducted to test the performance of our algorithm. The proposed pairwise registration algorithm is accurate, and robust to small overlaps, noise and varying mesh resolutions. The proposed multi-view registration algorithm is also very accurate. Rigorous comparisons with the state-of-the-art show the superiority of our algorithm. Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, Min Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | A low cost 3D markerless system for the reconstruction of athletic techniquesabstractWe present a low cost markerless system for the optimization of athlete performance in sports such as pole vault, jumping and javelin throw. The system uses a number of calibrated cameras to capture a video of an athlete from different viewpoints. The athlete's body is then segmented from the background in each video frame. The silhouettes of the segmented body are then reprojected to reconstruct an estimate of the 3D body shape of the athlete, known as the visual hull (VH). The VH is tracked over a number of frames in real testing trials. A template combining a high resolution 3D scan and a 2D mass scan is then aligned with the VH in each frame. A set of motion analysis parameters such as the take-off data are finally estimated from the aligned template and compared with the ones obtained using a gold standard marker-based system, namely the Vicon. The proposed system was tested in real-time trials and was able to provide comparable results to the Vicon system. Amar A. El-Sallam, Mohammed Bennamoun, Ferdous Sohel, Jacqueline A. Alderson, Andrew Lyttle, M. Rossi |
WACV | 3 |
| 2013 | 3D free form object recognition using rotational projection statisticsabstractRecognizing 3D objects in the presence of clutter and occlusion is a challenging task. This paper presents a 3D free form object recognition system based on a novel local surface feature descriptor. For a randomly selected feature point, a local reference frame (LRF) is defined by calculating the eigenvectors of the covariance matrix of a local surface, and a feature descriptor called rotational projection statistics (RoPS) is constructed by calculating the statistics of the point distribution on 2D planes defined from the LRF. It finally proposes a 3D object recognition algorithm based on RoPS features. Candidate models and transformation hypotheses are generated by matching the scene features against the model features in the library, these hypotheses are then tested and verified by aligning the model to the scene. Comparative experiments were performed on two publicly available datasets and an overall recognition rate of 98.8% was achieved. Experimental results show that our method is robust to noise, mesh resolution variations and occlusion. Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Jianwei Wan, Min Lu 0001 |
WACV | 3 |
| 2013 | Rotational Projection Statistics for 3D Local Surface Description and Object Recognition
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Min Lu 0001, Jianwei Wan |
Int. J. Comput. Vis. | 2 |
| 2012 | Sliding-Window Designs for Vertex-Based Shape CodingabstractTraditionally the sliding window (SW) has been employed in vertex-based operational rate distortion (ORD) optimal shape coding algorithms to ensure consistent distortion (quality) measurement and improve computational efficiency. It also regulates the memory requirements for an encoder design enabling regular, symmetrical hardware implementations. This paper presents a series of new enhancements to existing techniques for determining the best SW-length within a rate-distortion (RD) framework, and analyses the nexus between SW-length and storage for ORD hardware realizations. In addition, it presents an efficient bit-allocation strategy for managing multiple shapes together with a generalized adaptive SW scheme which integrates localized curvature information (cornerity) on contour points with a bi-directional spatial distance, to afford a superior and more pragmatic SW design compared with existing adaptive SW solutions which are based on only cornerity values. Experimental results consistently corroborate the effectiveness of these new strategies. Ferdous Sohel, Gour C. Karmakar, Laurence Dooley, Mohammed Bennamoun |
IEEE Trans. Multim. | 1 |
| 2008 | Dynamic Bezier curves for variable rate-distortion
Ferdous Sohel, Gour C. Karmakar, Laurence Dooley |
Pattern Recognit. | 1 |
| 2008 | Quasi-Bezier curves integrating localised information
Ferdous Sohel, Gour C. Karmakar, Laurence Dooley, John R. Arkinstall |
Pattern Recognit. | 1 |
| 2007 | Fast Distortion Measurement Using Chord-Length Parameterization Within the Vertex-Based Rate-Distortion Optimal Shape Coding FrameworkabstractExisting vertex-based operational rate-distortion (ORD) optimal shape coding algorithms can use a number of different distortion measurement techniques, including the shortest absolute distance (SAD), the distortion band (DB), the tolerance band (TB), and the accurate distortion measurement technique for shape coding (ADMSC). From a computational time perspective, an N-point contour requires O(N2) time for DB and TB for both polygon and B-spline-based encoding, while SAD and ADMSC incur O(N) time for polygonal encoding but O(N2) for B-spline based encoding, thereby rendering the ORD optimal algorithms computationally inefficient. This letter presents a novel distortion measurement strategy based on chord-length parameterization (DMCLP) of a boundary that incurs order O(N) complexity for both polygon and B-spline-based encoding while preserving a comparable rate-distortion performance to the original ORD optimal shape coding algorithms Ferdous Sohel, Gour C. Karmakar, Laurence Dooley |
IEEE Signal Process. Lett. | 1 |
| 2007 | New Dynamic Enhancements to the Vertex-Based Rate-Distortion Optimal Shape Coding FrameworkabstractExisting vertex-based operational rate-distortion (ORD) optimal shape coding algorithms use a vertex band around the shape boundary as the source of candidate control points (CP) usually in combination with a tolerance band (TB) and sliding window (SW) arrangement, as their distortion measuring technique. These algorithms however, employ a fixed vertex-band width irrespective of the shape and admissible distortion (AD), so the full bit-rate reduction potential is not fulfilled. Moreover, despite the causal impact of the SW-length upon both the bit-rate and computational-speed, there is no formal mechanism for determining the most suitable SW-length. This paper introduces the concept of a variable width admissible CP band and new adaptive SW-length selection strategy to address these issues. The presented quantitative and qualitative results analysis endorses the superior performance achieved by integrating these enhancements into the existing vertex-based ORD optimal algorithms. Ferdous Sohel, Laurence Dooley, Gour C. Karmakar |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Variable Width Admissible Control Point Band for Vertex Based Operational-Rate-Distortion Optimal Shape Coding AlgorithmsabstractExisting vertex-based operational-rate-distortion (ORD) optimal shape coding algorithms use a fixed width admissible control point band (FCB) around the shape boundary as the search space for possible control points. The width of the band however, is fixed and arbitrarily chosen independent of the admissible distortion and shape contour, so it fails to fully exploit the admissible control point band to reduce the bit-rate. This paper proposes a variable width admissible control point band (VCB) where the width associated to each boundary point is dynamically determined from the admissible peak distortion and shape information. In addition, the paper uses an accurate distortion measurement method to overcome a key limitation of existing distortion and tolerance band based methods. Experimental results reveal that both the qualitative and quantitative performance of the existing ORD algorithms are improved by seamlessly integrating the VCB and accurate distortion measuring approach. Ferdous Sohel, Laurence Dooley, Gour C. Karmakar |
ICIP | 1 |
| 2006 | Accurate distortion measurement for generic shape coding
Ferdous Sohel, Laurence Dooley, Gour C. Karmakar |
Pattern Recognit. Lett. | 1 |
| 2005 | Enhanced Bezier curve models incorporating local informationabstractThe Bezier curve is fundamental to many challenging and practical applications, ranging from computer aided geometric design and postscript font representations through to generic object shape descriptors and surface representation. A drawback of the Bezier curve however, is that it only considers global information about the control points, so there is often a large gap between the curve and its control polygon, leading to considerable error in curve representations. To address this issue, this paper presents enhanced Bezier curve (EBC) models which seamlessly incorporate local information. The performance of the models is empirically evaluated upon a number of natural and synthetic objects having arbitrary shape and both qualitative and quantitative results confirm the superiority of both EBC models in comparison with the classical Bezier curve representation, with no increase in the order of computational complexity. Ferdous Sohel, Gour C. Karmakar, Laurence Dooley, John R. Arkinstall |
ICASSP (4) | 1 |
| 2005 | A dynamic Bezier curve modelabstractBezier curves (BC) are fundamental to a wide range of applications from computer-aided design through to object shape descriptions and surface mapping. Since BC only consider global information with respect to their control points, this can lead to erroneous shape representations, though integrating local control point information minimises this error. This paper presents a new dynamic Bezier curve (DBC) model which combines both localised and global shape information by making a parametric shift of the BC points in the gap between the curve and its control polygon. The value of the shifting parameter is dynamically determined for a prescribed maximum distortion. DBC retains the kernel properties of the BC without increasing computational complexity order. The model's performance has been empirically evaluated on a number of arbitrary-shaped objects from geometric modelling to shape coding. Both qualitative and quantitative results confirm the improvement achieved compared with the classical BC representation. Ferdous Sohel, Laurence Dooley, Gour C. Karmakar |
ICIP (2) | 1 |