EDBT 2026 Demo / reviewers in the wild / expert
Syed Afaq Ali Shah
dblp:141/9937
· DBLP profile ↗
44ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0003-2181-8445ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic context improvisational retrieval-augmented generation for empathic conversational AIabstractAbstract A fundamental limitation of modern conversational AI is its limited capacity to demonstrate sustained empathy in long-form interactions. We propose SCIRAG (Semantic Context Improvisational Retrieval-Augmented Generation), a feedback-driven retrieval framework for adaptive empathic dialogue. It employs a dual-loop retrieval framework, iteratively optimizing a static counseling dataset through user metadata and feedback memory refinement. To enhance contextual alignment, we deploy retrieval adaptation, enabling the model to retain and leverage past conversational cues based on user preferences. When integrated with Mixtral-8x7B, SCIRAG improves human-rated empathic understanding by +1.26 points and empathic response by +1.00 point on the RoPE scale, while increasing acceptability by +7.66 points compared to a fine-tuned non-RAG baseline. Automatic evaluation further shows gains in semantic alignment (BERTScore-F1 +0.11) and fluency (perplexity reduced from 19.1 to 12.3). We also present EMPATHIC , a dataset of unscripted, therapeutic conversations. Unlike conventional datasets that contain only 2-4 dialogue turns per conversation, EMPATHIC provides extended conversational trajectories (50+ turns per session), allowing models to learn long-range coherence and empathic listening. The proposed dataset will be publicly released. Sharjeel Tahir, Judith Johnson, Jumana M. Abu-Khalaf, Syed Afaq Ali Shah |
Neural Comput. Appl. | 4 |
| 2026 | RLAD: A Reliable Hippo-Guided Multi-Task Model for Alzheimer's Disease DiagnosisabstractEarly diagnosis of Alzheimer's disease (AD) is crucial for its prevention, and hippocampal atrophy is a significant lesion for early diagnosis. The current DL-based AD diagnosis methods only focus on either AD classification or hippocampus segmentation independently, neglecting the correlation between the two tasks and lacking pathological interpretability. To address this issue, we propose a Reliable Hippo-guided Learning model for Alzheimer's Disease diagnosis (RLAD), which employs multi-task learning for AD classification as a main task supplemented by hippocampus segmentation. More specifically, our model consists of 1) a hybrid shared features encoder that encodes local and global information in MRI to enhance the model's ability to learn discriminative features; 2) Task Specific Decoders to accomplish AD classification and hippocampus segmentation; and 3) Task Coordination module to correlate the two tasks and guide the classification task to focus on the hippocampus area. Our proposed RLAD model is evaluated on MRI scans of 1631 subjects from three independent datasets, including ADNI-1, ADNI-2, and HarP. Our extensive experimental results demonstrate that the proposed model significantly improves the performance of AD classification and hippocampus segmentation with strong generalization capabilities. Zhenxin Lei, Cong Hua, Johann Li, Syed Afaq Ali Shah, Liang Zhang 0010, Mohammed Bennamoun, Cuiping Mao |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | CubeSat Downlink Communications Enhanced by Movable Antennas
Zeynab Khodkar, Shihao Yan, Syed Afaq Ali Shah, Rajen Biswa, Paulo de Souza, Leshan Uggalla, Derrick Wing Kwan Ng |
ICC | 3 |
| 2025 | Enhancing object recognition: The role of object knowledge decomposition and component-labeled datasets
Nuoye Xiong, Ning Wang 0047, Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 6 |
| 2024 | Language Model Guided Interpretable Video Action ReasoningabstractWhile neural networks have excelled in video action recognition tasks, their “black-box” nature often obscures the understanding of their decision-making processes. Re-cent approaches used inherently interpretable models to an-alyze video actions in a manner akin to human reasoning. These models, however, usually fall short in performance compared to their “black-box” counterparts. In this work, we present a new framework named Language-guided Interpretable Action Recognition framework (La-IAR). LaIAR leverages knowledge from language models to enhance both the recognition capabilities and the inter-pretability of video models. In essence, we redefine the problem of understanding video model decisions as a task of aligning video and language models. Using the logical reasoning captured by the language model, we steer the training of the video model. This integrated approach not only improves the video model's adaptability to different domains but also boosts its overall performance. Extensive experiments on two complex video action datasets, Charades & CAD-120, validates the improved performance and inter-pretability of our LaIAR framework. The code of LaIAR is available at https://github.com/NingWang2049/LaIAR. Ning Wang 0047, Guangming Zhu 0001, HS Li, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
CVPR | 5 |
| 2024 | DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition
Qi Wang 0189, Yuming Lin 0007, Jingtao Ye, Hongsheng Li 0003, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Liang Zhang 0010 |
ECCV (84) | 7 |
| 2024 | GenAI in Rule-based Systems for IoMT Security: Testing and EvaluationabstractGenerative AI (GenAI) represents a significant advancement in Artificial intelligence research, offering numerous benefits and opening new avenues for innovation across various domains. In healthcare, Generative AI has shown promise in applications such as drug discovery, personalized medicine, and medical imaging. This paper examines the role of Generative AI in rule-based systems, where vulnerabilities are detected with the help of formal logic. In this context, the ruleset is generated and tested to evaluate the performance of rule-based systems with the aid of GenAI. The effectiveness of the GenAI tool was evaluated using a publicly available case study from a laboratory setting. The results show that using generative Artificial intelligence in rule-based systems leads to increased creativity, continuous learning, and robust performance. GenAI responded to each use case and provided the desired results compared to traditional rule-based systems. This integration of advanced AI techniques with traditional rule-based systems ensures that these hybrid systems perform reliably and effectively. Kulsoom Saima Bughio, David M. Cook, Syed Afaq Ali Shah |
KES | 3 |
| 2024 | Therapying Outside the Box: Innovating the Implementation and Evaulation of CBT in Therapeutic Artificial Agents
Sharjeel Tahir, Jumana M. Abu-Khalaf, Syed Afaq Ali Shah, Judith Johnson |
WISE (4) | 3 |
| 2024 | Scene Graph Generation: A comprehensive surveyabstractDeep learning techniques have led to remarkable breakthroughs in the field of object detection and have spawned a lot of scene-understanding tasks in recent years. Scene graph has been the focus of research because of its powerful semantic representation and applications to scene understanding. Scene Graph Generation (SGG) refers to the task of automatically mapping an image or a video into a semantic structural scene graph, which requires the correct labeling of detected objects and their relationships. In this paper, a comprehensive survey of recent achievements is provided. This survey attempts to connect and systematize the existing visual relationship detection methods, to summarize, and interpret the mechanisms and the strategies of SGG in a comprehensive way. Deep discussions about current existing problems and future research directions are given at last. This survey will help readers to develop a better understanding of the current researches. Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Youliang Jiang, Yixuan Dang, Haoran Hou, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 9 |
| 2024 | Text-to-text generative approach for enhanced complex word identificationabstractThis paper presents a novel approach for solving the Complex Word Identification (CWI) task using the text-to-text generative model. The CWI task involves identifying complex words in text, which is a challenging Natural Language Processing task. To our knowledge, it is a first attempt to address CWI problem into text-to-text context. In this work, we propose a new methodology that leverages the power of the Transformer model to evaluate complexity of words in binary and probabilistic settings. We also propose a novel CWI dataset, which consists of 62,200 phrases, both complex and simple. We train and fine-tune our proposed model on our CWI dataset. We also evaluate its performance on separate test sets across three different domains. Our experimental results demonstrate the effectiveness of our proposed approach compared to state-of-the-art methods. • The paper proposes a transformer based generative approach for complex word identification (CWI) • Our technique uses text generation and addresses the CWI task in the text-to-text context. • The paper also proposes and develops a new CWI dataset for solving the CWI task using the proposed method. • We fine-tune our model for CWI task in binary settings and it performs on par with state-of-the-art methods. • In addition, we also fine-tune the model using probabilistic settings and it achieves state-of-the-art results. • Our dataset and code is publicly available for the research community. Patrycja Sliwiak, Syed Afaq Ali Shah |
Neurocomputing | 2 |
| 2023 | A Large Scale Multi-View RGBD Visual Affordance Learning DatasetabstractThe physical and textural attributes of objects have been widely studied for recognition, detection and segmentation tasks in computer vision. A number of datasets, such as large scale ImageNet, have been proposed for feature learning using data hungry deep neural networks and for hand-crafted feature extraction. To intelligently interact with objects, robots and intelligent machines need the ability to infer beyond the traditional physical/textural attributes, and understand/learn visual cues, called visual affordances, for affordance recognition, detection and segmentation. To date there is no publicly available large dataset for visual affordance understanding and learning. In this paper, we introduce a large scale multi-view RGBD visual affordance learning dataset, a benchmark of 47210 RGBD images from 37 object categories, annotated with 15 visual affordance categories. To the best of our knowledge, this is the first ever and the largest multi-view RGBD visual affordance learning dataset. We benchmark the proposed dataset for affordance segmentation and recognition tasks using popular Vision Transformer and Convolutional Neural Networks. Several state-of-the-art deep learning networks are evaluated each for affordance recognition and segmentation tasks. Our experimental results showcase the challenging nature of the dataset and present definite prospects for new and robust affordance learning algorithms. The dataset is publicly available at https://sites.google.com/view/afaqshah/dataset. Zeyad Osama Khalifa, Syed Afaq Ali Shah |
ICIP | 2 |
| 2023 | 3D Brain Registration with Intensity Shift RobustnessabstractTechnological advances in medical imaging are enabling us to understand healthcare datasets in great detail. Machine Learning enabled methods, specifically, deep neural networks are continuously achieving benchmark performances in terms of accuracy and computational efficiency. However, the lack of agreed-upon standard procedures, variations in the devices by different vendors, and artifacts induced by the physical phenomenon in the sensors make the data inconsistent and noisy. These variations in the data are detrimental to the performance of learning-based methods. In this study, we analyze the behavior of traditional and deep learning-based image registration methods and explore strategies to handle the problem of intensity distributional shifts without compromising the performance. To achieve this, we propose an intensity-based loss function and demonstrate that the models trained with our proposed loss function are better at handling unseen data from different sites using machines from different vendors. In addition, our trained model is superior in preserving the boundaries of anatomical regions after registration. Hassan Mahmood, Asim Iqbal, Syed M. S. Islam, Syed Afaq Ali Shah |
ICIP | 4 |
| 2023 | Hierarchical Transformer for Visual Affordance Understanding using a Large-scale DatasetabstractRecognition, detection, and segmentation tasks in machine vision have focused on studying the physical and textural attributes of objects. However, robots and intelligent machines require the ability to understand visual cues, such as the visual affordances that objects offer, to interact intelligently with novel objects. In this paper, we present a large-scale multi-view RGBD visual affordance learning dataset a benchmark of 47,210 RGBD images from 37 object categories, annotated with 15 visual affordance categories and 35 cluttered/complex scenes. We deploy a Vision Transformer (ViT), called Visual Affordance Transformer (VAT), for the affordance segmentation task. Due to its hierarchical architecture, VAT can learn multiple affordances at various scales, making it suitable for objects of varying sizes. Our experimental results show the superior performance of VAT compared to state-of-the-art deep learning networks. In addition, the challenging nature of the proposed dataset highlights the potential for new and robust affordance learning algorithms. Our dataset is publicly available at https://sites.google.com/view/afaqshah/dataset. Syed Afaq Ali Shah, Zeyad Osama Khalifa |
IROS | 1 |
| 2023 | UE4-NeRF: Neural Radiance Field for Real-Time Rendering of Large-Scale SceneabstractNeural Radiance Fields (NeRF) is a novel implicit 3D reconstruction method that shows immense potential and has been gaining increasing attention. It enables the reconstruction of 3D scenes solely from a set of photographs. However, its real-time rendering capability, especially for interactive real-time rendering of large-scale scenes, still has significant limitations. To address these challenges, in this paper, we propose a novel neural rendering system called UE4-NeRF, specifically designed for real-time rendering of large-scale scenes. We partitioned each large scene into different sub-NeRFs. In order to represent the partitioned independent scene, we initialize polygonal meshes by constructing multiple regular octahedra within the scene and the vertices of the polygonal faces are continuously optimized during the training process. Drawing inspiration from Level of Detail (LOD) techniques, we trained meshes of varying levels of detail for different observation levels. Our approach combines with the rasterization pipeline in Unreal Engine 4 (UE4), achieving real-time rendering of large-scale scenes at 4K resolution with a frame rate of up to 43 FPS. Rendering within UE4 also facilitates scene editing in subsequent stages. Furthermore, through experiments, we have demonstrated that our method achieves rendering quality comparable to state-of-the-art approaches. Project page: https://jamchaos.github.io/UE4-NeRF/. Jiaming Gu, Minchao Jiang, Hongsheng Li 0003, Xiaoyuan Lu, Guangming Zhu 0001, Syed Afaq Ali Shah, Liang Zhang 0010, Mohammed Bennamoun |
NeurIPS | 6 |
| 2023 | Position and structure-aware graph learning
Guoqiang Ye, Juan Song, Mingtao Feng, Guangming Zhu 0001, Peiyi Shen, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 7 |
| 2023 | CommuNety: deep learning-based face recognition system for the prediction of cohesive communitiesabstractAbstract Effective mining of social media, which consists of a large number of users is a challenging task. Traditional approaches rely on the analysis of text data related to users to accomplish this task. However, text data lacks significant information about the social users and their associated groups. In this paper, we propose CommuNety, a deep learning system for the prediction of cohesive networks using face images from photo albums. The proposed deep learning model consists of hierarchical CNN architecture to learn descriptive features related to each cohesive network. The paper also proposes a novel Face Co-occurrence Frequency algorithm to quantify existence of people in images, and a novel photo ranking method to analyze the strength of relationship between different individuals in a predicted social network. We extensively evaluate the proposed technique on PIPA dataset and compare with state-of-the-art methods. Our experimental results demonstrate the superior performance of the proposed technique for the prediction of relationship between different individuals and the cohesiveness of communities. Syed Afaq Ali Shah, WeiFeng Deng, Muhammad Aamir Cheema, Abdul Bais |
Multim. Tools Appl. | 1 |
| 2022 | MEDAS: an open-source platform as a service to help break the walls between medicine and informatics
Liang Zhang 0010, Johann Li, Ping Li 0030, Xiaoyuan Lu, Maoguo Gong, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Kun Qian 0003, Björn W. Schuller |
Neural Comput. Appl. | 8 |
| 2022 | Probability-Based Framework to Fuse Temporal Consistency and Semantic Information for Background SegmentationabstractThe fusion of temporal consistency and semantic information with limited foreground information for background segmentation using deep learning is an underinvestigated problem. In this paper, we explore the relation between temporal consistency and semantic information based on the law of total probability. A highly concise framework is proposed to fuse these two types of information. A theoretical proof is given to show that the proposed framework is more accurate than either the temporal consistency-based model or the semantic information-based model and that each model is a special case of the proposed framework. The proposed framework is a white-box framework that can easily be embedded into a deep neural network as a merging layer. In the proposed model, only a few parameters must be learned, which substantially reduces the need for a large dataset. In addition, these interpretable parameters reflect our understanding of the background and can be applied to a wide range of environments. Extensive evaluations indicate the promising performance of the proposed method. Our code and trained weights for the experiments are available at GitHub.11https://github.com/zengzhi2015/SS_TC_BS(We encourage the reader to run the program for a better understanding of the proposed method). Ting Wang 0026, Fulei Ma, Liang Zhang 0010, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Multim. | 6 |
| 2022 | Analysis and Variants of Broad Learning SystemabstractThe broad learning system (BLS) is designed based on the technology of compressed sensing and pseudo-inverse theory, and consists of feature nodes and enhancement nodes, has been proposed recently. Compared with the popular deep learning structures, such as deep neural networks, BLS has the ability of rapid incremental learning and can remodel the system without the usual tedious retraining process. However, given that BLS is still in its infancy, it still needs analysis, improvements, and verification. In this article, we first analyze the principle of fast incremental learning ability of BLS in depth. Second, in order to provide an in-depth analysis of the BLS structure, according to the novel structure design concept of deep neural networks, we present four brand-new BLS variant networks and their incremental realizations. Third, based on our analysis of the effect of feature nodes and enhancement nodes, a new BLS structure with a semantic feature extraction layer has been proposed, which is called SFEBLS. The experimental results show that SFEBLS and its variants can increase the accuracy rate on the NORB dataset 6.18%, Fashion-MNIST dataset by 3.15%, ORL data by 5.00%, street view house number dataset by 12.88%, and CIFAR-10 dataset by 18.42%, respectively, and the four brand-new BLS variant networks also obviously outperform the original BLS. Liang Zhang 0010, Guoqing Lu, Peiyi Shen, Mohammed Bennamoun, Syed Afaq Ali Shah, Qiguang Miao, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2021 | SubICap: Towards Subword-informed Image Captioning
Naeha Sharif, Mohammed Bennamoun, Wei Liu 0006, Syed Afaq Ali Shah |
WACV | 4 |
| 2021 | Real time surveillance for low resolution and limited data scenarios: An image set classification approach
Uzair Nadeem, Syed Afaq Ali Shah, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel |
Inf. Sci. | 2 |
| 2021 | U-net based analysis of MRI for Alzheimer's disease diagnosis
Zhonghao Fan, Johann Li, Liang Zhang 0010, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun, Tao Hua, Wei Wei 0006 |
Neural Comput. Appl. | 8 |
| 2021 | Multi-Modal Co-Learning for Liver Lesion Segmentation on PET-CT ImagesabstractLiver lesion segmentation is an essential process to assist doctors in hepatocellular carcinoma diagnosis and treatment planning. Multi-modal positron emission tomography and computed tomography (PET-CT) scans are widely utilized due to their complementary feature information for this purpose. However, current methods ignore the interaction of information across the two modalities during feature extraction, omit the co-learning of the feature maps of different resolutions, and do not ensure that shallow and deep features complement each others sufficiently. In this paper, our proposed model can achieve feature interaction across multi-modal channels by sharing the down-sampling blocks between two encoding branches to eliminate misleading features. Furthermore, we combine feature maps of different resolutions to derive spatially varying fusion maps and enhance the lesions information. In addition, we introduce a similarity loss function for consistency constraint in case that predictions of separated refactoring branches for the same regions vary a lot. We evaluate our model for liver tumor segmentation using a PET-CT scans dataset, compare our method with the baseline techniques for multi-modal (multi-branches, multi-channels and cascaded networks) and then demonstrate that our method has a significantly higher accuracy ( ) than the baseline models. Zhongliang Xue, Ping Li 0030, Liang Zhang 0010, Xiaoyuan Lu, Guangming Zhu 0001, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Efficient Scene Text Detection with Textual Attention TowerabstractScene text detection has received attention for years and achieved an impressive performance across various benchmarks. In this work, we propose an efficient and accurate approach to detect multi-oriented text in scene images. The proposed feature fusion mechanism allows us to use a shallower network to reduce the computational complexity. A self-attention mechanism is adopted to suppress false positive detections. Experiments on public benchmarks including ICDAR 2013, ICDAR 2015 and MSRA-TD500 show that our proposed approach can achieve better or comparable performances with fewer parameters and less computational cost. Liang Zhang 0010, Lu Yang 0019, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Peiyi Shen |
ICASSP | 6 |
| 2020 | Efficient Detection of Pixel-Level Adversarial AttacksabstractDeep learning has achieved unprecedented performance in object recognition and scene understanding. However, deep models are also found vulnerable to adversarial attacks. Of particular relevance to robotics systems are pixel-level attacks that can completely fool a neural network by altering very few pixels (e.g. 1-5) in an image. We present the first technique to detect the presence of adversarial pixels in images for the robotic systems, employing an Adversarial Detection Network (ADNet). The proposed network efficiently recognize an input as adversarial or clean by discriminating the peculiar activation signals of the adversarial samples from the clean ones. It acts as a defense mechanism for the robotic vision system by detecting and rejecting the adversarial samples. We thoroughly evaluate our technique on three benchmark datasets including CIFAR-10, CIFAR-100 and Fashion MNIST. Results demonstrate effective detection of adversarial samples by ADNet. Syed Afaq Ali Shah, Moise Bougre, Naveed Akhtar, Mohammed Bennamoun, Liang Zhang 0010 |
ICIP | 1 |
| 2020 | Structure-Feature based Graph Self-adaptive PoolingabstractVarious methods to deal with graph data have been proposed in recent years. However, most of these methods focus on graph feature aggregation rather than graph pooling. Besides, the existing top-k selection graph pooling methods have a few problems. First, to construct the pooled graph topology, current top-k selection methods evaluate the importance of the node from a single perspective only, which is simplistic and unobjective. Second, the feature information of unselected nodes is directly lost during the pooling process, which inevitably leads to a massive loss of graph feature information. To solve these problems mentioned above, we propose a novel graph self-adaptive pooling method with the following objectives: (1) to construct a reasonable pooled graph topology, structure and feature information of the graph are considered simultaneously, which provide additional veracity and objectivity in node selection; and (2) to make the pooled nodes contain sufficiently effective graph information, node feature information is aggregated before discarding the unimportant nodes; thus, the selected nodes contain information from neighbor nodes, which can enhance the use of features of the unselected nodes. Experimental results on four different datasets demonstrate that our method is effective in graph classification and outperforms state-of-the-art graph pooling methods. Liang Zhang 0010, Hongsheng Li 0003, Guangming Zhu 0001, Peiyi Shen, Ping Li 0030, Xiaoyuan Lu, Syed Afaq Ali Shah, Mohammed Bennamoun |
WWW | 8 |
| 2020 | Color vision deficiency datasets & recoloring evaluation using GANs
Hongsheng Li 0003, Liang Zhang 0010, Meili Zhang, Guangming Zhu 0001, Peiyi Shen, Ping Li 0030, Mohammed Bennamoun, Syed Afaq Ali Shah |
Multim. Tools Appl. | 9 |
| 2020 | Topology-learnable graph convolution for skeleton-based action recognition
Guangming Zhu 0001, Liang Zhang 0010, Hongsheng Li 0003, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
Pattern Recognit. Lett. | 5 |
| 2020 | Block Level Skip Connections Across Cascaded V-Net for Multi-Organ SegmentationabstractMulti-organ segmentation is a challenging task due to the label imbalance and structural differences between different organs. In this work, we propose an efficient cascaded V-Net model to improve the performance of multi-organ segmentation by establishing dense Block Level Skip Connections (BLSC) across cascaded V-Net. Our model can take full advantage of features from the first stage network and make the cascaded structure more efficient. We also combine stacked small and large kernels with an inception-like structure to help our model to learn more patterns, which produces superior results for multi-organ segmentation. In addition, some small organs are commonly occluded by large organs and have unclear boundaries with other surrounding tissues, which makes them hard to be segmented. We therefore first locate the small organs through a multi-class network and crop them randomly with the surrounding region, then segment them with a single-class network. We evaluated our model on SegTHOR 2019 challenge unseen testing set and Multi-Atlas Labeling Beyond the Cranial Vault challenge validation set. Our model has achieved an average dice score gain of 1.62 percents and 3.90 percents compared to traditional cascaded networks on these two datasets, respectively. For hard-to-segment small organs, such as the esophagus in SegTHOR 2019 challenge, our technique has achieved a gain of 5.63 percents on dice score, and four organs in Multi-Atlas Labeling Beyond the Cranial Vault challenge have achieved a gain of 5.27 percents on average dice score. Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Redundancy and Attention in Convolutional LSTM for Gesture RecognitionabstractConvolutional long short-term memory (ConvLSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into ConvLSTM networks. This paper explores the redundancy of spatial convolutions and the effects of the attention mechanism in ConvLSTM, based on our previous gesture recognition architectures that combine the 3-D convolutional neural network (CNN) and ConvLSTM. Depthwise separable, group, and shuffle convolutions are used to replace the convolutional structures in ConvLSTM for the redundancy analysis. In addition, four ConvLSTM variants are derived for attention analysis: 1) by removing the convolutional structures of the three gates in ConvLSTM; 2) by applying the attention mechanism on the ConvLSTM input; and 3) by reconstructing the input and 4) output gates with the modified channelwise attention mechanism. Evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion and that the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn long-term spatiotemporal features when taking spatial or spatiotemporal features as input. A new LSTM variant is derived on this basis in which the convolutional structures are embedded only into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available.\footnotehttps://github.com/GuangmingZhu/ConvLSTMForGR. Guangming Zhu 0001, Liang Zhang 0010, Lu Yang 0019, Lin Mei 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Peiyi Shen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Relationship Detection Based on Object Semantic Inference and Attention MechanismsabstractDetecting relations among objects is a crucial task for image understanding. However, each relationship involves different objects pair combinations, and different objects pair combinations express diverse interactions. This makes the relationships, based just on visual features, a challenging task. In this paper, we propose a simple yet effective relationship detection model, which is based on object semantic inference and attention mechanisms. Our model is trained to detect relation triples, such as , . To overcome the high diversity of visual appearances, the semantic inference module and the visual features are combined to complement each others. We also introduce two different attention mechanisms for object feature refinement and phrase feature refinement. In order to derive a more detailed and comprehensive representation for each object, the object feature refinement module refines the representation of each object by querying over all the other objects in the image. The phrase feature refinement module is proposed in order to make the phrase feature more effective, and to automatically focus on relative parts, to improve the visual relationship detection task. We validate our model on Visual Genome Relationship dataset. Our proposed model achieves competitive results compared to the state-of-the-art method MOTIFNET. Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun |
ICMR | 5 |
| 2019 | LCEval: Learned Composite Metric for Caption Evaluation
Naeha Sharif, Lyndon White, Mohammed Bennamoun, Wei Liu 0006, Syed Afaq Ali Shah |
Int. J. Comput. Vis. | 5 |
| 2019 | Continuous Gesture Segmentation and Recognition Using 3DCNN and Convolutional LSTMabstractContinuous gesture recognition aims at recognizing the ongoing gestures from continuous gesture sequences and is more meaningful for the scenarios, where the start and end frames of each gesture instance are generally unknown in practical applications. This paper presents an effective deep architecture for continuous gesture recognition. First, continuous gesture sequences are segmented into isolated gesture instances using the proposed temporal dilated Res3D network. A balanced squared hinge loss function is proposed to deal with the imbalance between boundaries and nonboundaries. Temporal dilation can preserve the temporal information for the dense detection of the boundaries at fine granularity, and the large temporal receptive field makes the segmentation results more reasonable and effective. Then, the recognition network is constructed based on the 3-D convolutional neural network (3DCNN), the convolutional long-short-term-memory network (ConvLSTM), and the 2-D convolutional neural network (2DCNN) for isolated gesture recognition. The “3DCNN-ConvLSTM-2DCNN” architecture is more effective to learn long-term and deep spatiotemporal features. The proposed segmentation and recognition networks obtain the Jaccard index of 0.7163 on the Chalearn LAP ConGD dataset, which is 0.106 higher than the winner of 2017 ChaLearn LAP Large-Scale Continuous Gesture Recognition Challenge. Guangming Zhu 0001, Liang Zhang 0010, Peiyi Shen, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Multim. | 5 |
| 2018 | NNEval: Neural Network Based Evaluation Metric for Image Captioning
Naeha Sharif, Lyndon White, Mohammed Bennamoun, Syed Afaq Ali Shah |
ECCV (8) | 4 |
| 2018 | Reflective Field for Pixel-Level TasksabstractPixelNet has achieved great success in dense prediction problems with a pure pixel-level architecture, but there is still much room for improvement. In this paper, we start from PixelNet and discuss the pixel-level architecture called hypercol-umn and its limitations in building feature representation with rich semantic information. To achieve this goal, we propose a concept in the context of neural networks called reflective field, representing the area reflected by the origin input. Furthermore, the proposed reflective field is used to solve the limitations of the hypercolumn architecture. Specifically, we give the method of calculating the size of the reflective field and analyze the effective reflective field in the calculated area. Then, we use the reflective field to build a new hypercolumn architecture, which has a more rational construction. The results on PASCAL VOC segmentation dataset with our new architecture are improved. Liang Zhang 0010, Xiangwen Kong, Peiyi Shen, Guangming Zhu 0001, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
ICPR | 6 |
| 2018 | Attention in Convolutional LSTM for Gesture RecognitionabstractConvolutional long short-term memory (LSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into the LSTM or the convolutional LSTM (ConvLSTM) networks. Based on the previous gesture recognition architectures which combine the three-dimensional convolution neural network (3DCNN) and ConvLSTM, this paper explores the effects of attention mechanism in ConvLSTM. Several variants of ConvLSTM are evaluated: (a) Removing the convolutional structures of the three gates in ConvLSTM, (b) Applying the attention mechanism on the input of ConvLSTM, (c) Reconstructing the input and (d) output gates respectively with the modified channel-wise attention mechanism. The evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion, and the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn the long-term spatiotemporal features, when taking as input the spatial or spatiotemporal features. On this basis, a new variant of LSTM is derived, in which the convolutional structures are only embedded into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available. Liang Zhang 0010, Guangming Zhu 0001, Lin Mei 0001, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
NeurIPS | 5 |
| 2018 | Efficient finer-grained incremental processing with MapReduce for big data
Liang Zhang 0010, Yuanyuan Feng, Peiyi Shen, Guangming Zhu 0001, Wei Wei 0006, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
Future Gener. Comput. Syst. | 7 |
| 2018 | Improved colour-to-grey method using image segmentation and colour difference model for colour vision deficiencyabstractColour vision deficiency (CVD) is a genetic condition that has troubled people for a long time. This study proposes an improved colour‐to‐grey method for CVD using image segmentation and a colour difference model. In this method, the colour image is first segmented using a region growing method so that each region corresponds to one colour. Next, the colour difference is computed between arbitrary segmented region pairs. Finally, the greyscale image is obtained by minimising a target function. Experimental results show that compared with state‐of‐the‐art colour‐to‐grey methods, the proposed algorithm can improve the E ‐score by about 10.99%. Liang Zhang 0010, Guangming Zhu 0001, Juan Song, Peiyi Shen, Wei Wei 0006, Syed Afaq Ali Shah, Mohammed Bennamoun |
IET Image Process. | 8 |
| 2018 | Semantic scene completion with dense CRF from a single depth image
Liang Zhang 0010, Peiyi Shen, Mohammed Bennamoun, Guangming Zhu 0001, Syed Afaq Ali Shah, Juan Song |
Neurocomputing | 7 |
| 2017 | Keypoints-based surface representation for 3D modeling and 3D object recognition
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Pattern Recognit. | 1 |
| 2016 | Iterative deep learning for image set based face and object recognition
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Neurocomputing | 1 |
| 2016 | A novel feature representation for automatic 3D object recognition in cluttered scenes
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Neurocomputing | 1 |
| 2015 | A novel 3D vorticity based approach for automatic registration of low resolution range images
Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd |
Pattern Recognit. | 1 |
| 2013 | 3D-Div: A novel local surface descriptor for feature matching and pairwise range image registrationabstractThis paper presents a novel local surface descriptor, called 3D-Div. The proposed descriptor is based on the concept of 3D vector fields divergence, extensively used in electromagnetic theory. To generate a 3D-Div descriptor of a 3D surface, a keypoint is first extracted on the 3D surface, then a local patch of a certain size is selected around that keypoint. A Local Reference Frame (LRF) is then constructed at the keypoint using all points forming the patch. A normalized 3D vector field is then computed at each point in the patch and referenced with LRF vectors. The 3D-Div descriptors are finally generated as the divergence of the reoriented 3D vector field. We tested our proposed descriptor on the low resolution Washington RGB-D (Kinect) object dataset. Performance was evaluated for the tasks of feature matching and pairwise range image registration. Experimental results showed that the proposed 3D-Div is 88% more computationally efficient and 47% more accurate than commonly used Spin Image (SI) descriptors. Syed Afaq Ali Shah, Mohammed Bennamoun, Farid Boussaïd, Amar A. El-Sallam |
ICIP | 1 |