EDBT 2026 Demo / reviewers in the wild / expert
Liang Zhang 0010
dblp:50/6759-10
· DBLP profile ↗
59ranked-venue papers
15as first author
31since 2021 · last 2026
0000-0003-4331-5830ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 29 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OC-SemR: Ontology-Constrained Semantic Re-ranking for Knowledge Graph Completion
Tianci Wu, Guangming Zhu 0001, Jun Sheng, Siyuan Wang 0004, Liang Zhang 0010 |
ICIC (23) | 6 |
| 2026 | CIFE-FSOD: Few-shot object detection via joint extraction of common and instance-specific features
Shengye Zhou, Johann Li, Guangming Zhu 0001, Lin Mei 0001, Liang Zhang 0010 |
Neurocomputing | 5 |
| 2026 | RLAD: A Reliable Hippo-Guided Multi-Task Model for Alzheimer's Disease DiagnosisabstractEarly diagnosis of Alzheimer's disease (AD) is crucial for its prevention, and hippocampal atrophy is a significant lesion for early diagnosis. The current DL-based AD diagnosis methods only focus on either AD classification or hippocampus segmentation independently, neglecting the correlation between the two tasks and lacking pathological interpretability. To address this issue, we propose a Reliable Hippo-guided Learning model for Alzheimer's Disease diagnosis (RLAD), which employs multi-task learning for AD classification as a main task supplemented by hippocampus segmentation. More specifically, our model consists of 1) a hybrid shared features encoder that encodes local and global information in MRI to enhance the model's ability to learn discriminative features; 2) Task Specific Decoders to accomplish AD classification and hippocampus segmentation; and 3) Task Coordination module to correlate the two tasks and guide the classification task to focus on the hippocampus area. Our proposed RLAD model is evaluated on MRI scans of 1631 subjects from three independent datasets, including ADNI-1, ADNI-2, and HarP. Our extensive experimental results demonstrate that the proposed model significantly improves the performance of AD classification and hippocampus segmentation with strong generalization capabilities. Zhenxin Lei, Cong Hua, Johann Li, Syed Afaq Ali Shah, Liang Zhang 0010, Mohammed Bennamoun, Cuiping Mao |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Votesplat: Hough Voting Gaussian Splatting for 3D Scene Understandingabstract3D Gaussian Splatting (3DGS) has become horsepower in high-quality, real-time rendering for novel view synthesis of 3D scenes. However, existing methods focus primarily on geometric and appearance modeling, lacking deeper scene understanding while also incurring high training costs that complicate the originally streamlined differentiable rendering pipeline. To this end, we propose VoteSplat, a novel 3D scene understanding framework that integrates Hough voting with 3DGS. Specifically, Segment Anything Model (SAM) is utilized for instance segmentation, extracting objects, and generating 2D vote maps. We then embed spatial offset vectors into Gaussian primitives. These offsets construct 3D spatial votes by associating them with 2D image votes, while depth distortion constraints refine localization along the depth axis. For open-vocabulary object localization, VoteSplat maps 2D image semantics to 3D point clouds via voting points, reducing training costs associated with high-dimensional CLIP features while preserving semantic unambiguity. Extensive experiments demonstrate effectiveness of VoteSplat in open-vocabulary 3D instance localization, 3D point cloud understanding, click-based 3D object localization, hierarchical segmentation, and ablation studies. Our code is available at https://sy-ja.github.io/votesplat/ Minchao Jiang, Shunyu Jia, Jiaming Gu, Xiaoyuan Lu, Guangming Zhu 0001, Anqi Dong, Liang Zhang 0010 |
ICCV | 7 |
| 2025 | Intervening in Black Box: Concept Bottleneck Model for Enhancing Human Neural Network Mutual UnderstandingabstractRecent advances in deep learning have led to increasingly complex models with deeper layers and more parameters, reducing interpretability and making their decisions harder to understand. While many methods explain black-box reasoning, most lack effective interventions or only operate at sample-level without modifying the model itself. To address this, we propose the Concept Bottleneck Model for Enhancing Human-Neural Network Mutual Understanding (CBM-HNMU). CBM-HNMU leverages the Concept Bottleneck Model (CBM) as an interpretable framework to approximate black-box reasoning and communicate conceptual understanding. Detrimental concepts are automatically identified and refined (removed/replaced) based on global gradient contributions. The modified CBM then distills corrected knowledge back into the black-box model, enhancing both interpretability and accuracy. We evaluate CBM-HNMU on various CNN and transformer-based models across Flower-102, CIFAR-10, CIFAR-100, FGVC-Aircraft, and CUB-200, achieving a maximum accuracy improvement of 2.64% and a maximum increase in average accuracy across 1.03%. Source code is available at: https://github.com/XiGuaBo/CBM-HNMU. Nuoye Xiong, Anqi Dong, Ning Wang 0047, Cong Hua, Guangming Zhu 0001, Lin Mei 0001, Peiyi Shen, Liang Zhang 0010 |
ICCV | 8 |
| 2025 | Sprig: Low-Latency Startup for Knative Serverless Platforms with Async-CforkabstractServerless computing has emerged as a transformative paradigm in the cloud-native era, widely adopted across major public cloud platforms. However, its pay-as-use model necessitates frequent dynamic instance creation and destruction, making cold start latency a critical system bottleneck. While the industry predominantly employs caching techniques to mitigate cold start overheads, these approaches exhibit inherent limitations. Existing research prototypes often optimize container runtimes manually, resulting in systems incompatible with production environments that cannot leverage automated management capabilities or integrate optimizations effectively. Moreover, runtime optimizations fail to fit system features like asynchronous cold start. This paper presents Sprig, the first serverless system that seamlessly integrates cutting-edge academic research with production-grade platforms. We introduce a novel computational abstraction called Template Pod to bridge the gap between research-oriented flexible designs and minimal computational units in production systems. Sprig further incorporates Asynccfork, a low-latency startup technique optimized for large-scale serverless platforms, significantly reducing function cold start latency. Additionally, we design a multiplexed proxy queue with enhanced scheduling elasticity to resolve request accumulation issues in Knative's asynchronous startup mechanism. Evaluations show Sprig achieves a 70% average reduction in cold start latency compared with Knative, and decreases P99 end-to-end latency from 27s to 0.5s under high load while maintaining only 16.7% of baseline memory footprint for 20 instances. Dong Du 0003, Liang Zhang 0010, Yubin Xia |
JCC | 3 |
| 2025 | Prompt-guided Disentangled Representation for Action RecognitionabstractAction recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging to model the interactions between different objects in multi-action scenarios. To alleviate this issue, we explore disentangling any specified actions from complex scenes as an effective solution. In this paper, we propose Prompt-guided Disentangled Representation for Action Recognition (ProDA), a novel framework that disentangles any specified actions from a multi-action scene. ProDA leverages Spatio-temporal Scene Graphs (SSGs) and introduces Dynamic Prompt Module (DPM) to guide a Graph Parsing Neural Network (GPNN) in generating action-specific representations. Furthermore, we design a video-adapted GPNN that aggregates information using dynamic weights. Extensive experiments on two complex video action datasets, Charades and SportsHHI, demonstrate the effectiveness of our approach against state-of-the-art methods. Our code can be found in https://github.com/iamsnaping/ProDA.git. Tianci Wu, Guangming Zhu 0001, Jiang Lu, Siyuan Wang 0004, Ning Wang 0047, Nuoye Xiong, Liang Zhang 0010 |
NeurIPS | 7 |
| 2025 | PD-NeRF: a general pseudo-depth supervision method for neural radiance fields
Jiaming Gu, Minchao Jiang, Xiaoyuan Lu, Cong Hua, Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010 |
Sci. China Inf. Sci. | 7 |
| 2025 | Enhancing object recognition: The role of object knowledge decomposition and component-labeled datasets
Nuoye Xiong, Ning Wang 0047, Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 5 |
| 2025 | Exploring Hierarchical Spatial Layout Cues for 3D Point Cloud Based Scene Graph Predictionabstract3D scene graph prediction is important for intelligent agents to gather information and perceive semantics of their environments. However, constructing an effective graph is nontrivial given the complexity of natural scenes. Existing solutions for graph representation of 3D scenes still distinguish each detailed discrepancy among all the relationships as flat thinking, ignoring the mechanism used by humans to perform this task. Inspired by the role of the prefrontal cortex in hierarchical reasoning, we analyze this problem from a novel perspective: exploring hierarchical spatial layout cues in 3D space and navigating that hierarchy to make the 3D scene graph more accurate in a vertical division to horizontal propagation strategy. To this end, we first encode the contextual object features for fine-gained object category classification. Next, we build a bottom-up hierarchical graph to predict remarkably diverse support relationships in a single concept regardless of numerous irrelevant relationships. Finally, equipped with the spatially-true and semantically-meaningful support relationships, we focus on the local region layout to propagate the semantic features to predict the additional non-support relationships under the guidance of the given referred hierarchical graph nodes. Experiments on the challenging 3DSSG benchmark show that our algorithm outperforms existing state-of-the-art, and can also alleviate the impact of the long-tailed distribution of training data. Our code is available athttps://github.com/HHrEtvP/HSLC-3DSG/. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
IEEE Trans. Multim. | 3 |
| 2024 | Enhance Sketch Recognition's Explainability via Semantic Component-Level ParsingabstractFree-hand sketches are appealing for humans as a universal tool to depict the visual world. Humans can recognize varied sketches of a category easily by identifying the concurrence and layout of the intrinsic semantic components of the category, since humans draw free-hand sketches based a common consensus that which types of semantic components constitute each sketch category. For example, an airplane should at least have a fuselage and wings. Based on this analysis, a semantic component-level memory module is constructed and embedded in the proposed structured sketch recognition network in this paper. The memory keys representing semantic components of each sketch category can be self-learned and enhance the recognition network's explainability. Our proposed networks can deal with different situations of sketch recognition, i.e., with or without semantic components labels of strokes. Experiments on the SPG and SketchIME datasets demonstrate the memory module's flexibility and the recognition network's explainability. The code and data are available at https://github.com/GuangmingZhu/SketchESC. Guangming Zhu 0001, Siyuan Wang 0004, Tianci Wu, Liang Zhang 0010 |
AAAI | 4 |
| 2024 | Language Model Guided Interpretable Video Action ReasoningabstractWhile neural networks have excelled in video action recognition tasks, their “black-box” nature often obscures the understanding of their decision-making processes. Re-cent approaches used inherently interpretable models to an-alyze video actions in a manner akin to human reasoning. These models, however, usually fall short in performance compared to their “black-box” counterparts. In this work, we present a new framework named Language-guided Interpretable Action Recognition framework (La-IAR). LaIAR leverages knowledge from language models to enhance both the recognition capabilities and the inter-pretability of video models. In essence, we redefine the problem of understanding video model decisions as a task of aligning video and language models. Using the logical reasoning captured by the language model, we steer the training of the video model. This integrated approach not only improves the video model's adaptability to different domains but also boosts its overall performance. Extensive experiments on two complex video action datasets, Charades & CAD-120, validates the improved performance and inter-pretability of our LaIAR framework. The code of LaIAR is available at https://github.com/NingWang2049/LaIAR. Ning Wang 0047, Guangming Zhu 0001, HS Li, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
CVPR | 4 |
| 2024 | DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition
Qi Wang 0189, Yuming Lin 0007, Jingtao Ye, Hongsheng Li 0003, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Liang Zhang 0010 |
ECCV (84) | 9 |
| 2024 | Scene Graph Generation: A comprehensive surveyabstractDeep learning techniques have led to remarkable breakthroughs in the field of object detection and have spawned a lot of scene-understanding tasks in recent years. Scene graph has been the focus of research because of its powerful semantic representation and applications to scene understanding. Scene Graph Generation (SGG) refers to the task of automatically mapping an image or a video into a semantic structural scene graph, which requires the correct labeling of detected objects and their relationships. In this paper, a comprehensive survey of recent achievements is provided. This survey attempts to connect and systematize the existing visual relationship detection methods, to summarize, and interpret the mechanisms and the strategies of SGG in a comprehensive way. Deep discussions about current existing problems and future research directions are given at last. This survey will help readers to develop a better understanding of the current researches. Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Youliang Jiang, Yixuan Dang, Haoran Hou, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 3 |
| 2023 | 3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point CloudabstractIn-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide variety of challenging relationships, existing methods perform quite poorly with limited training samples. In this work, we find that the inherently hierarchical structures of physical space in 3D scenes aid in the automatic association of semantic and spatial arrangements, specifying clear patterns and leading to less ambiguous predictions. Thus, they well meet the challenges due to the rich variations within scene categories. To achieve this, we explicitly unify these structural cues of 3D physical spaces into deep neural networks to facilitate scene graph prediction. Specifically, we exploit an external knowledge base as a baseline to accumulate both contextualized visual content and textual facts to form a 3D spatial multimodal knowledge graph. Moreover, we propose a knowledge-enabled scene graph prediction module benefiting from the 3D spatial knowledge to effectively regularize semantic space of relationships. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art competitors. Our code is available at https://github.com/HHrEtvP/SMKA. Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Ajmal Mian |
CVPR | 3 |
| 2023 | Sketch Input Method Editor: A Comprehensive Dataset and Methodology for Systematic Input RecognitionabstractWith the recent surge in the use of touchscreen devices, free-hand sketching has emerged as a promising modality for human-computer interaction. While previous research has focused on tasks such as recognition, retrieval, and generation of familiar everyday objects, this study aims to create a Sketch Input Method Editor (SketchIME) specifically designed for a professional Command, Control, Communications, Computer, and Intelligence (C4I) system. Within this system, sketches are utilized as low-fidelity prototypes for recommending standardized symbols in the creation of comprehensive situation maps. This paper also presents a systematic dataset comprising 374 specialized sketch types, and proposes a simultaneous recognition and segmentation architecture with multilevel supervision between recognition and segmentation to improve performance and enhance interpretability. By incorporating few-shot domain adaptation and class-incremental learning, the network's ability to adapt to new users and extend to new task-specific classes is significantly enhanced. Results from experiments conducted on both the proposed dataset and the SPG dataset illustrate the superior performance of the proposed architecture. Our dataset and code are publicly available at https://github.com/GuangmingZhu/SketchIME. Guangming Zhu 0001, Siyuan Wang 0004, Qing Cheng 0004, Kelong Wu, Hao Li 0179, Liang Zhang 0010 |
ACM Multimedia | 6 |
| 2023 | UE4-NeRF: Neural Radiance Field for Real-Time Rendering of Large-Scale SceneabstractNeural Radiance Fields (NeRF) is a novel implicit 3D reconstruction method that shows immense potential and has been gaining increasing attention. It enables the reconstruction of 3D scenes solely from a set of photographs. However, its real-time rendering capability, especially for interactive real-time rendering of large-scale scenes, still has significant limitations. To address these challenges, in this paper, we propose a novel neural rendering system called UE4-NeRF, specifically designed for real-time rendering of large-scale scenes. We partitioned each large scene into different sub-NeRFs. In order to represent the partitioned independent scene, we initialize polygonal meshes by constructing multiple regular octahedra within the scene and the vertices of the polygonal faces are continuously optimized during the training process. Drawing inspiration from Level of Detail (LOD) techniques, we trained meshes of varying levels of detail for different observation levels. Our approach combines with the rasterization pipeline in Unreal Engine 4 (UE4), achieving real-time rendering of large-scale scenes at 4K resolution with a frame rate of up to 43 FPS. Rendering within UE4 also facilitates scene editing in subsequent stages. Furthermore, through experiments, we have demonstrated that our method achieves rendering quality comparable to state-of-the-art approaches. Project page: https://jamchaos.github.io/UE4-NeRF/. Jiaming Gu, Minchao Jiang, Hongsheng Li 0003, Xiaoyuan Lu, Guangming Zhu 0001, Syed Afaq Ali Shah, Liang Zhang 0010, Mohammed Bennamoun |
NeurIPS | 7 |
| 2023 | Position and structure-aware graph learning
Guoqiang Ye, Juan Song, Mingtao Feng, Guangming Zhu 0001, Peiyi Shen, Liang Zhang 0010, Syed Afaq Ali Shah, Mohammed Bennamoun |
Neurocomputing | 6 |
| 2023 | Exploring Spatio-Temporal Graph Convolution for Video-Based Human-Object Interaction RecognitionabstractVideo-based human-object interaction recognition is a challenging task since the state of objects as well as their correlations change constantly in the video. Existing methods mainly use 3DCNN or use separate components (e.g., GCN + RNN) to model the spatial correlation or the temporal correlation respectively, but ignore modeling spatio-temporal correlations simultaneously and long-term temporal dynamics of objects. In this paper, we propose a novel model, named Spatio-Temporal Interaction Graph Parsing Networks (STIGPN), for human-object interaction recognition in videos. STIGPN captures both spatial and temporal correlations simultaneously and thus can capture intra-frame and inter-frame dependencies efficiently and effectively. To model long-term temporal dynamics of objects, we introduce spatio-temporal feature enhancement, which can improve the detection of the salient human-object interaction pairs. We explore three types of spatio-temporal graph convolutions to simultaneously capture the spatio-temporal correlations and assess their effectiveness as the basic building block of STIGPN. Extensive experiments on CAD-120, Something-Else and Charades datasets show that our proposed solution leads to competitive results compared with the state-of-the-art methods. Code for STIGPN is available at:https://github.com/NingWang2049/STIGPN2 Ning Wang 0047, Guangming Zhu 0001, Hongsheng Li 0003, Mingtao Feng, Lan Ni, Peiyi Shen, Lin Mei 0001, Liang Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2022 | Learning from Pixel-Level Noisy Label : A New Perspective for Light Field Saliency DetectionabstractSaliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obtained from unsupervised hand crafted featured-based saliency methods. Given this goal, a natural question is: can we efficiently incorporate the relationships among light field cues while identifying clean labels in a unified framework? We address this question by formulating the learning as a joint optimization of intra light field features fusion stream and inter scenes correlation stream to generate the predictions. Specially, we first introduce a pixel forgetting guided fusion module to mutually enhance the light field features and exploit pixel consistency across iterations to identify noisy pixels. Next, we introduce a cross scene noise penalty loss for better reflecting latent structures of training data and enabling the learning to be invariant to noise. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our framework showing that it learns saliency prediction comparable to state-of-the-art fully supervised light field saliency methods. Our code is available at h t tps://github.com/ OLobbCode/NoiseLF. Mingtao Feng, Kendong Liu, Liang Zhang 0010, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian |
CVPR | 3 |
| 2022 | Spatial Parsing and Dynamic Temporal Pooling networks for Human-Object Interaction detectionabstractThe key of Human-Object Interaction(HOI) recognition is to infer the relationship between human and objects. Recently, the image's Human-Object Interaction(HOI) detection has made significant progress. However, there is still room for improvement in video HOI detection performance. Existing one-stage methods use well-designed end-to-end networks to detect a video segment and directly predict an interaction. A side effect of these approaches is that we have no way of knowing the human-object pair of the interaction or the keyframes in which the interaction took place. It makes the model learning and further optimization of the network more complex. This paper introduces the Spatial Parsing and Dynamic Temporal Pooling (SPDTP) network, which takes the entire video as a spatio-temporal graph with human and object nodes as input. Unlike existing methods, our proposed network predicts the difference between interactive and non-interactive pairs through explicit spatial parsing, and then performs interaction recognition. Moreover, we propose a learnable and differentiable Dynamic Temporal Module(DTM) to emphasize the keyframes of the video and suppress the redundant frame. Furthermore, the experimental results show that SPDTP can pay more attention to active human-object pairs and valid keyframes. Overall, we achieve state-of-the-art performance on CAD-120 dataset and Something- Else dataset. Hongsheng Li 0003, Guangming Zhu 0001, Wu Zhen, Lan Ni, Peiyi Shen, Liang Zhang 0010, Ning Wang 0047, Cong Hua |
IJCNN | 6 |
| 2022 | MEDAS: an open-source platform as a service to help break the walls between medicine and informatics
Liang Zhang 0010, Johann Li, Ping Li 0030, Xiaoyuan Lu, Maoguo Gong, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Kun Qian 0003, Björn W. Schuller |
Neural Comput. Appl. | 1 |
| 2022 | Probability-Based Framework to Fuse Temporal Consistency and Semantic Information for Background SegmentationabstractThe fusion of temporal consistency and semantic information with limited foreground information for background segmentation using deep learning is an underinvestigated problem. In this paper, we explore the relation between temporal consistency and semantic information based on the law of total probability. A highly concise framework is proposed to fuse these two types of information. A theoretical proof is given to show that the proposed framework is more accurate than either the temporal consistency-based model or the semantic information-based model and that each model is a special case of the proposed framework. The proposed framework is a white-box framework that can easily be embedded into a deep neural network as a merging layer. In the proposed model, only a few parameters must be learned, which substantially reduces the need for a large dataset. In addition, these interpretable parameters reflect our understanding of the background and can be applied to a wide range of environments. Extensive evaluations indicate the promising performance of the proposed method. Our code and trained weights for the experiments are available at GitHub.11https://github.com/zengzhi2015/SS_TC_BS(We encourage the reader to run the program for a better understanding of the proposed method). Ting Wang 0026, Fulei Ma, Liang Zhang 0010, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Multim. | 4 |
| 2022 | Analysis and Variants of Broad Learning SystemabstractThe broad learning system (BLS) is designed based on the technology of compressed sensing and pseudo-inverse theory, and consists of feature nodes and enhancement nodes, has been proposed recently. Compared with the popular deep learning structures, such as deep neural networks, BLS has the ability of rapid incremental learning and can remodel the system without the usual tedious retraining process. However, given that BLS is still in its infancy, it still needs analysis, improvements, and verification. In this article, we first analyze the principle of fast incremental learning ability of BLS in depth. Second, in order to provide an in-depth analysis of the BLS structure, according to the novel structure design concept of deep neural networks, we present four brand-new BLS variant networks and their incremental realizations. Third, based on our analysis of the effect of feature nodes and enhancement nodes, a new BLS structure with a semantic feature extraction layer has been proposed, which is called SFEBLS. The experimental results show that SFEBLS and its variants can increase the accuracy rate on the NORB dataset 6.18%, Fashion-MNIST dataset by 3.15%, ORL data by 5.00%, street view house number dataset by 12.88%, and CIFAR-10 dataset by 18.42%, respectively, and the four brand-new BLS variant networks also obviously outperform the original BLS. Liang Zhang 0010, Guoqing Lu, Peiyi Shen, Mohammed Bennamoun, Syed Afaq Ali Shah, Qiguang Miao, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloudabstract3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a freeform language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of point clouds. There are three main challenges in 3D object grounding: to find the main focus in the complex and diverse description; to understand the point cloud scene; and to locate the target object. In this paper, we address all three challenges. Firstly, we propose a language scene graph module to capture the rich structure and long-distance phrase correlations. Secondly, we introduce a multi-level 3D proposal relation graph module to extract the object-object and object-scene co-occurrence relationships, and strengthen the visual features of the initial proposals. Lastly, we develop a description guided 3D visual graph module to encode global contexts of phrases and proposals by a nodes matching strategy. Extensive experiments on challenging benchmark datasets (ScanRefer [3] and Nr3D [42]) show that our algorithm outperforms existing state-of-the-art. Our code is available at https://github.com/PNXD/FFL-3DOG. Mingtao Feng, Liang Zhang 0010, Guangming Zhu 0001, Hui Zhang 0023, Yaonan Wang 0001, Ajmal Mian |
ICCV | 4 |
| 2021 | Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction RecognitionabstractFor a given video-based Human-Object Interaction scene, modeling the spatio-temporal relationship between humans and objects is the important cue to understand the contextual information presented in the video. With the efficient spatio-temporal relationship modeling, it is possible not only to uncover contextual information in each frame, but to directly capture inter-frame dependencies as well. Capturing the position changes of human and objects over the spatio-temporal dimension is more critical when significant changes in the appearance features may not occur over time. When utilizing appearance features, the spatial location and the semantic information are also the key to improve the video-based Human-Object Interaction recognition performance. In this paper, Spatio-Temporal Interaction Graph Parsing Networks (STIGPN) are constructed, which encode the videos with a graph composed of human and object nodes. These nodes are connected by two types of relations: (i) intra-frame relations: modeling the interactions between human and the interacted objects within each frame. (ii) inter-frame relations: capturing the long range dependencies between human and the interacted objects across frame. With the graph, STIGPN learn spatio-temporal features directly from the whole video-based Human-Object Interaction scenes. Multi-modal features and a multi-stream fusion strategy are used to enhance the reasoning capability of STIGPN. Two Human-Object Interaction video datasets, including CAD-120 and Something-Else, are used to evaluate the proposed architectures, and the state-of-the-art performance demonstrates the superiority of STIGPN. Code for STIGPN is available at https://github.com/GuangmingZhu/STIGPN. Ning Wang 0047, Guangming Zhu 0001, Liang Zhang 0010, Peiyi Shen, Hongsheng Li 0003, Cong Hua |
ACM Multimedia | 3 |
| 2021 | U-net based analysis of MRI for Alzheimer's disease diagnosis
Zhonghao Fan, Johann Li, Liang Zhang 0010, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun, Tao Hua, Wei Wei 0006 |
Neural Comput. Appl. | 3 |
| 2021 | COVID-19 Chest CT Image Segmentation Network by Multi-Scale Fusion and Enhancement OperationsabstractA novel coronavirus disease 2019 (COVID-19) was detected and has spread rapidly across various countries around the world since the end of the year 2019. Computed Tomography (CT) images have been used as a crucial alternative to the time-consuming RT-PCR test. However, pure manual segmentation of CT images faces a serious challenge with the increase of suspected cases, resulting in urgent requirements for accurate and automatic segmentation of COVID-19 infections. Unfortunately, since the imaging characteristics of the COVID-19 infection are diverse and similar to the backgrounds, existing medical image segmentation methods cannot achieve satisfactory performance. In this article, we try to establish a new deep convolutional neural network tailored for segmenting the chest CT images with COVID-19 infections. We first maintain a large and new chest CT image dataset consisting of 165,667 annotated chest CT images from 861 patients with confirmed COVID-19. Inspired by the observation that the boundary of the infected lung can be enhanced by adjusting the global intensity, in the proposed deep CNN, we introduce a feature variation block which adaptively adjusts the global properties of the features for segmenting COVID-19 infection. The proposed FV block can enhance the capability of feature representation effectively and adaptively for diverse cases. We fuse features at different scales by proposing Progressive Atrous Spatial Pyramid Pooling to handle the sophisticated infection areas with diverse appearance and shapes. The proposed method achieves state-of-the-art performance. Dice similarity coefficients are 0.987 and 0.726 for lung and COVID-19 segmentation, respectively. We conducted experiments on the data collected in China and Germany and show that the proposed deep CNN can produce impressive performance effectively. The proposed network enhances the segmentation ability of the COVID-19 infection, makes the connection with other techniques and contributes to the development of remedying COVID-19 infection. Qingsen Yan, Bo Wang 0011, Dong Gong, Chuan Luo 0003, Jianhu Shen, Jingyang Ai, Qinfeng Shi, Yanning Zhang 0001, Liang Zhang 0010, Zheng You |
IEEE Trans. Big Data | 11 |
| 2021 | Relation Graph Network for 3D Object Detection in Point CloudsabstractConvolutional Neural Networks (CNNs) have emerged as a powerful tool for object detection in 2D images. However, their power has not been fully realised for detecting 3D objects directly in point clouds without conversion to regular grids. Moreover, existing state-of-the-art 3D object detection methods aim to recognize objects individually without exploiting their relationships during learning or inference. In this article, we first propose a strategy that associates the predictions of direction vectors with pseudo geometric centers, leading to a win-win solution for 3D bounding box candidates regression. Secondly, we propose point attention pooling to extract uniform appearance features for each 3D object proposal, benefiting from the learned direction features, semantic features and spatial coordinates of the object points. Finally, the appearance features are used together with the position features to build 3D object-object relationship graphs for all proposals to model their co-existence. We explore the effect of relation graphs on proposals' appearance feature enhancement under supervised and unsupervised settings. The proposed relation graph network comprises a 3D object proposal generation module and a 3D relation module, making it an end-to-end trainable network for detecting 3D objects in point clouds. Experiments on challenging benchmark point cloud datasets (SunRGB-D, ScanNet and KITTI) show that our algorithm performs better than existing state-of-the-art. Mingtao Feng, Syed Zulqarnain Gilani, Yaonan Wang 0001, Liang Zhang 0010, Ajmal Mian |
IEEE Trans. Image Process. | 4 |
| 2021 | Attention-Guided Deep Neural Network With Multi-Scale Feature Fusion for Liver Vessel SegmentationabstractLiver vessel segmentation is fast becoming a key instrument in the diagnosis and surgical planning of liver diseases. In clinical practice, liver vessels are normally manual annotated by clinicians on each slice of CT images, which is extremely laborious. Several deep learning methods exist for liver vessel segmentation, however, promoting the performance of segmentation remains a major challenge due to the large variations and complex structure of liver vessels. Previous methods mainly using existing UNet architecture, but not all features of the encoder are useful for segmentation and some even cause interferences. To overcome this problem, we propose a novel deep neural network for liver vessel segmentation, called LVSNet, which employs special designs to obtain the accurate structure of the liver vessel. Specifically, we design Attention-Guided Concatenation (AGC) module to adaptively select the useful context features from low-level features guided by high-level features. The proposed AGC module focuses on capturing rich complemented information to obtain more details. In addition, we introduce an innovative multi-scale fusion block by constructing hierarchical residual-like connections within one single residual block, which is of great importance for effectively linking the local blood vessel fragments together. Furthermore, we construct a new dataset containing 40 thin thickness cases (0.625 mm) which consist of CT volumes and annotated vessels. To evaluate the effectiveness of the method with minor vessels, we also propose an automatic stratification method to split major and minor liver vessels. Extensive experimental results demonstrate that the proposed LVSNet outperforms previous methods on liver vessel segmentation datasets. Additionally, we conduct a series of ablation studies that comprehensively support the superiority of the underlying concepts. Qingsen Yan, Bo Wang 0011, Wei Zhang 0098, Chuan Luo 0003, Wei Xu 0005, Zhengqing Xu, Yanning Zhang 0001, Qinfeng Shi, Liang Zhang 0010, Zheng You |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | Multi-Modal Co-Learning for Liver Lesion Segmentation on PET-CT ImagesabstractLiver lesion segmentation is an essential process to assist doctors in hepatocellular carcinoma diagnosis and treatment planning. Multi-modal positron emission tomography and computed tomography (PET-CT) scans are widely utilized due to their complementary feature information for this purpose. However, current methods ignore the interaction of information across the two modalities during feature extraction, omit the co-learning of the feature maps of different resolutions, and do not ensure that shallow and deep features complement each others sufficiently. In this paper, our proposed model can achieve feature interaction across multi-modal channels by sharing the down-sampling blocks between two encoding branches to eliminate misleading features. Furthermore, we combine feature maps of different resolutions to derive spatially varying fusion maps and enhance the lesions information. In addition, we introduce a similarity loss function for consistency constraint in case that predictions of separated refactoring branches for the same regions vary a lot. We evaluate our model for liver tumor segmentation using a PET-CT scans dataset, compare our method with the baseline techniques for multi-modal (multi-branches, multi-channels and cascaded networks) and then demonstrate that our method has a significantly higher accuracy ( ) than the baseline models. Zhongliang Xue, Ping Li 0030, Liang Zhang 0010, Xiaoyuan Lu, Guangming Zhu 0001, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Efficient Scene Text Detection with Textual Attention TowerabstractScene text detection has received attention for years and achieved an impressive performance across various benchmarks. In this work, we propose an efficient and accurate approach to detect multi-oriented text in scene images. The proposed feature fusion mechanism allows us to use a shallower network to reduce the computational complexity. A self-attention mechanism is adopted to suppress false positive detections. Experiments on public benchmarks including ICDAR 2013, ICDAR 2015 and MSRA-TD500 show that our proposed approach can achieve better or comparable performances with fewer parameters and less computational cost. Liang Zhang 0010, Lu Yang 0019, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Peiyi Shen |
ICASSP | 1 |
| 2020 | Efficient Detection of Pixel-Level Adversarial AttacksabstractDeep learning has achieved unprecedented performance in object recognition and scene understanding. However, deep models are also found vulnerable to adversarial attacks. Of particular relevance to robotics systems are pixel-level attacks that can completely fool a neural network by altering very few pixels (e.g. 1-5) in an image. We present the first technique to detect the presence of adversarial pixels in images for the robotic systems, employing an Adversarial Detection Network (ADNet). The proposed network efficiently recognize an input as adversarial or clean by discriminating the peculiar activation signals of the adversarial samples from the clean ones. It acts as a defense mechanism for the robotic vision system by detecting and rejecting the adversarial samples. We thoroughly evaluate our technique on three benchmark datasets including CIFAR-10, CIFAR-100 and Fashion MNIST. Results demonstrate effective detection of adversarial samples by ADNet. Syed Afaq Ali Shah, Moise Bougre, Naveed Akhtar, Mohammed Bennamoun, Liang Zhang 0010 |
ICIP | 5 |
| 2020 | Recurrent Graph Convolutional Networks for Skeleton-based Action RecognitionabstractHuman action recognition is one of the challenging and active research fields due to its wide range of applications. Recently, graph convolutions for skeleton-based action recognition have attracted much attention. Generally, the adjacency matrices of the graph are fixed to the hand-crafted physical connectivity of the human joints, or learned adaptively via deep learning. The hand-crafted or learned adjacency matrices are fixed when processing each frame of an action sequence. However, the interactions of different subsets of joints may play a core role at different phases of an action. Therefore, it is reasonable to evolve the graph topology with time. In this paper, a recurrent graph convolution is proposed, in which the graph topology is evolved via a long short-term memory (LSTM) network. The proposed recurrent graph convolutional network (R-GCN) can recurrently learn the data-dependent graph topologies for different layers, different time steps and different kinds of actions. Experimental results on the NTU RGB+D and Kinetics-Skeleton datasets demonstrate the advantages of the proposed R-GCN. Guangming Zhu 0001, Lu Yang 0019, Liang Zhang 0010, Peiyi Shen, Juan Song |
ICPR | 3 |
| 2020 | A Benchmark Dataset for Segmenting Liver, Vasculature and Lesions from Large-scale Computed Tomography DataabstractHow to build a high-performance liver-related computer assisted diagnosis system is an open question of great interest. However, the performance of the state-of-art algorithm is always limited by the amount of data and the quality of the label. To address this problem, we propose the biggest treatment-oriented liver cancer dataset for liver surgery and treatment planning. This dataset provides 216 cases (total about 268K frames) scanned images in contrast-enhanced computed tomography (CT). We labeled all the CT images with the liver, liver vasculature, and liver tumor segmentation ground truth for train and tune segmentation algorithms in advance. Based on that, we evaluate several recent and state-of-the-art segmentation algorithms, including 7 deep learning methods, on CT sequences. All results are compared to reference segmentations five error metrics that highlight different aspects of segmentation accuracy. In general, compared with previous datasets, our dataset is really a challenging dataset. To our knowledge, the proposed dataset and benchmark allow for the first time systematic exploration of such issues, and will be made available to allow for further research in this field. Bo Wang 0011, Qingsen Yan, Zhengqing Xu, Jingyang Ai, Wei Xu 0005, Liang Zhang 0010, Zheng You |
ICPR | 8 |
| 2020 | Graph-Temporal LSTM Networks for Skeleton-Based Action Recognition
Hongsheng Li 0003, Guangming Zhu 0001, Liang Zhang 0010, Juan Song, Peiyi Shen |
PRCV (2) | 3 |
| 2020 | Structure-Feature based Graph Self-adaptive PoolingabstractVarious methods to deal with graph data have been proposed in recent years. However, most of these methods focus on graph feature aggregation rather than graph pooling. Besides, the existing top-k selection graph pooling methods have a few problems. First, to construct the pooled graph topology, current top-k selection methods evaluate the importance of the node from a single perspective only, which is simplistic and unobjective. Second, the feature information of unselected nodes is directly lost during the pooling process, which inevitably leads to a massive loss of graph feature information. To solve these problems mentioned above, we propose a novel graph self-adaptive pooling method with the following objectives: (1) to construct a reasonable pooled graph topology, structure and feature information of the graph are considered simultaneously, which provide additional veracity and objectivity in node selection; and (2) to make the pooled nodes contain sufficiently effective graph information, node feature information is aggregated before discarding the unimportant nodes; thus, the selected nodes contain information from neighbor nodes, which can enhance the use of features of the unselected nodes. Experimental results on four different datasets demonstrate that our method is effective in graph classification and outperforms state-of-the-art graph pooling methods. Liang Zhang 0010, Hongsheng Li 0003, Guangming Zhu 0001, Peiyi Shen, Ping Li 0030, Xiaoyuan Lu, Syed Afaq Ali Shah, Mohammed Bennamoun |
WWW | 1 |
| 2020 | Color vision deficiency datasets & recoloring evaluation using GANs
Hongsheng Li 0003, Liang Zhang 0010, Meili Zhang, Guangming Zhu 0001, Peiyi Shen, Ping Li 0030, Mohammed Bennamoun, Syed Afaq Ali Shah |
Multim. Tools Appl. | 2 |
| 2020 | Point attention network for semantic segmentation of 3D point clouds
Mingtao Feng, Liang Zhang 0010, Xuefei Lin, Syed Zulqarnain Gilani, Ajmal Mian |
Pattern Recognit. | 2 |
| 2020 | Topology-learnable graph convolution for skeleton-based action recognition
Guangming Zhu 0001, Liang Zhang 0010, Hongsheng Li 0003, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
Pattern Recognit. Lett. | 2 |
| 2020 | Block Level Skip Connections Across Cascaded V-Net for Multi-Organ SegmentationabstractMulti-organ segmentation is a challenging task due to the label imbalance and structural differences between different organs. In this work, we propose an efficient cascaded V-Net model to improve the performance of multi-organ segmentation by establishing dense Block Level Skip Connections (BLSC) across cascaded V-Net. Our model can take full advantage of features from the first stage network and make the cascaded structure more efficient. We also combine stacked small and large kernels with an inception-like structure to help our model to learn more patterns, which produces superior results for multi-organ segmentation. In addition, some small organs are commonly occluded by large organs and have unclear boundaries with other surrounding tissues, which makes them hard to be segmented. We therefore first locate the small organs through a multi-class network and crop them randomly with the surrounding region, then segment them with a single-class network. We evaluated our model on SegTHOR 2019 challenge unseen testing set and Multi-Atlas Labeling Beyond the Cranial Vault challenge validation set. Our model has achieved an average dice score gain of 1.62 percents and 3.90 percents compared to traditional cascaded networks on these two datasets, respectively. For hard-to-segment small organs, such as the esophagus in SegTHOR 2019 challenge, our technique has achieved a gain of 5.63 percents on dice score, and four organs in Multi-Atlas Labeling Beyond the Cranial Vault challenge have achieved a gain of 5.27 percents on average dice score. Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Redundancy and Attention in Convolutional LSTM for Gesture RecognitionabstractConvolutional long short-term memory (ConvLSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into ConvLSTM networks. This paper explores the redundancy of spatial convolutions and the effects of the attention mechanism in ConvLSTM, based on our previous gesture recognition architectures that combine the 3-D convolutional neural network (CNN) and ConvLSTM. Depthwise separable, group, and shuffle convolutions are used to replace the convolutional structures in ConvLSTM for the redundancy analysis. In addition, four ConvLSTM variants are derived for attention analysis: 1) by removing the convolutional structures of the three gates in ConvLSTM; 2) by applying the attention mechanism on the ConvLSTM input; and 3) by reconstructing the input and 4) output gates with the modified channelwise attention mechanism. Evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion and that the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn long-term spatiotemporal features when taking spatial or spatiotemporal features as input. A new LSTM variant is derived on this basis in which the convolutional structures are embedded only into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available.\footnotehttps://github.com/GuangmingZhu/ConvLSTMForGR. Guangming Zhu 0001, Liang Zhang 0010, Lu Yang 0019, Lin Mei 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Peiyi Shen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Recoloring Image For Color Vision Deficiency By GANSabstractCompared to normal people who can recognize all colors in 3-D color space, people with Color Vision Deficiency (CVD) just could recognize colors in 2-D color space or only in 1-D color space. Therefore, color version deficient people cannot distinguish some colors under certain circumstances. In this paper, a new GAN-based architecture is constructed to improve the color identifiability for color vision deficient people. The improved framework is based on BicycleGAN [1]. First, a CVD simulation module is introduced and integrated into the BicycleGAN [1], allowing us to inspect the recoloring progress from the perspective of people with CVD. Second, a color-to-gray module is added to eliminate the distortion. Third, Gaussian Mixture Model (GMM) [2] and KullbackLeibler (KL) divergence [3], [4] calculation model are inserted as priori information of the conditional GANs [5] to guide the direction of recoloring. Experimental results show that our proposed architecture achieves impressive performance. Meili Zhang, Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Ping Li 0030 |
ICIP | 3 |
| 2019 | Relationship Detection Based on Object Semantic Inference and Attention MechanismsabstractDetecting relations among objects is a crucial task for image understanding. However, each relationship involves different objects pair combinations, and different objects pair combinations express diverse interactions. This makes the relationships, based just on visual features, a challenging task. In this paper, we propose a simple yet effective relationship detection model, which is based on object semantic inference and attention mechanisms. Our model is trained to detect relation triples, such as , . To overcome the high diversity of visual appearances, the semantic inference module and the visual features are combined to complement each others. We also introduce two different attention mechanisms for object feature refinement and phrase feature refinement. In order to derive a more detailed and comprehensive representation for each object, the object feature refinement module refines the representation of each object by querying over all the other objects in the image. The phrase feature refinement module is proposed in order to make the phrase feature more effective, and to automatically focus on relative parts, to improve the visual relationship detection task. We validate our model on Visual Genome Relationship dataset. Our proposed model achieves competitive results compared to the state-of-the-art method MOTIFNET. Liang Zhang 0010, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun |
ICMR | 1 |
| 2019 | Continuous Gesture Segmentation and Recognition Using 3DCNN and Convolutional LSTMabstractContinuous gesture recognition aims at recognizing the ongoing gestures from continuous gesture sequences and is more meaningful for the scenarios, where the start and end frames of each gesture instance are generally unknown in practical applications. This paper presents an effective deep architecture for continuous gesture recognition. First, continuous gesture sequences are segmented into isolated gesture instances using the proposed temporal dilated Res3D network. A balanced squared hinge loss function is proposed to deal with the imbalance between boundaries and nonboundaries. Temporal dilation can preserve the temporal information for the dense detection of the boundaries at fine granularity, and the large temporal receptive field makes the segmentation results more reasonable and effective. Then, the recognition network is constructed based on the 3-D convolutional neural network (3DCNN), the convolutional long-short-term-memory network (ConvLSTM), and the 2-D convolutional neural network (2DCNN) for isolated gesture recognition. The “3DCNN-ConvLSTM-2DCNN” architecture is more effective to learn long-term and deep spatiotemporal features. The proposed segmentation and recognition networks obtain the Jaccard index of 0.7163 on the Chalearn LAP ConGD dataset, which is 0.106 higher than the winner of 2017 ChaLearn LAP Large-Scale Continuous Gesture Recognition Challenge. Guangming Zhu 0001, Liang Zhang 0010, Peiyi Shen, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
IEEE Trans. Multim. | 2 |
| 2018 | Reflective Field for Pixel-Level TasksabstractPixelNet has achieved great success in dense prediction problems with a pure pixel-level architecture, but there is still much room for improvement. In this paper, we start from PixelNet and discuss the pixel-level architecture called hypercol-umn and its limitations in building feature representation with rich semantic information. To achieve this goal, we propose a concept in the context of neural networks called reflective field, representing the area reflected by the origin input. Furthermore, the proposed reflective field is used to solve the limitations of the hypercolumn architecture. Specifically, we give the method of calculating the size of the reflective field and analyze the effective reflective field in the calculated area. Then, we use the reflective field to build a new hypercolumn architecture, which has a more rational construction. The results on PASCAL VOC segmentation dataset with our new architecture are improved. Liang Zhang 0010, Xiangwen Kong, Peiyi Shen, Guangming Zhu 0001, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
ICPR | 1 |
| 2018 | Attention in Convolutional LSTM for Gesture RecognitionabstractConvolutional long short-term memory (LSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into the LSTM or the convolutional LSTM (ConvLSTM) networks. Based on the previous gesture recognition architectures which combine the three-dimensional convolution neural network (3DCNN) and ConvLSTM, this paper explores the effects of attention mechanism in ConvLSTM. Several variants of ConvLSTM are evaluated: (a) Removing the convolutional structures of the three gates in ConvLSTM, (b) Applying the attention mechanism on the input of ConvLSTM, (c) Reconstructing the input and (d) output gates respectively with the modified channel-wise attention mechanism. The evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion, and the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn the long-term spatiotemporal features, when taking as input the spatial or spatiotemporal features. On this basis, a new variant of LSTM is derived, in which the convolutional structures are only embedded into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available. Liang Zhang 0010, Guangming Zhu 0001, Lin Mei 0001, Peiyi Shen, Syed Afaq Ali Shah, Mohammed Bennamoun |
NeurIPS | 1 |
| 2018 | Efficient finer-grained incremental processing with MapReduce for big data
Liang Zhang 0010, Yuanyuan Feng, Peiyi Shen, Guangming Zhu 0001, Wei Wei 0006, Juan Song, Syed Afaq Ali Shah, Mohammed Bennamoun |
Future Gener. Comput. Syst. | 1 |
| 2018 | Improved colour-to-grey method using image segmentation and colour difference model for colour vision deficiencyabstractColour vision deficiency (CVD) is a genetic condition that has troubled people for a long time. This study proposes an improved colour‐to‐grey method for CVD using image segmentation and a colour difference model. In this method, the colour image is first segmented using a region growing method so that each region corresponds to one colour. Next, the colour difference is computed between arbitrary segmented region pairs. Finally, the greyscale image is obtained by minimising a target function. Experimental results show that compared with state‐of‐the‐art colour‐to‐grey methods, the proposed algorithm can improve the E ‐score by about 10.99%. Liang Zhang 0010, Guangming Zhu 0001, Juan Song, Peiyi Shen, Wei Wei 0006, Syed Afaq Ali Shah, Mohammed Bennamoun |
IET Image Process. | 1 |
| 2018 | Semantic scene completion with dense CRF from a single depth image
Liang Zhang 0010, Peiyi Shen, Mohammed Bennamoun, Guangming Zhu 0001, Syed Afaq Ali Shah, Juan Song |
Neurocomputing | 1 |
| 2018 | Benchmark Data Set and Method for Depth Estimation From Light Field ImagesabstractConvolutional Neural Networks (CNN) have performed extremely well for many image analysis tasks. However, supervised training of deep CNN architectures requires huge amounts of labelled data which is unavailable for light field images. In this paper, we leverage on synthetic light field images and propose a two stream CNN network that learns to estimate the disparities of multiple correlated neighbourhood pixels from their Epipolar Plane Images (EPI). Since the EPIs are unrelated except at their intersection, a two stream network is proposed to learn convolution weights individually for the EPIs and then combine the outputs of the two streams for disparity estimation. The CNN estimated disparity map is then refined using the central RGB light field image as a prior in a variational technique. We also propose a new real world dataset comprising light field images of 19 objects captured with the Lytro Illum camera in outdoor scenes and their corresponding 3D pointclouds, as ground truth, captured with the 3dMD scanner. This dataset will be made public to allow more precise 3D pointcloud level comparison of algorithms in the future which is currently not possible. Experiments on the synthetic and real world datasets show that our algorithm outperforms existing state-of-the-art for depth estimation from light field images. Mingtao Feng, Yaonan Wang 0001, Jian Liu 0014, Liang Zhang 0010, Hasan Firdaus M. Zaki, Ajmal Mian |
IEEE Trans. Image Process. | 4 |
| 2017 | Viewpoint calibration method based on point features for point cloud fusionabstractPoint cloud fusion is a significant process in many computer vision applications such as multi-robot maps building or multi-SLAM. When robots capture maps in different places, the maps suffer from large angle difference between respective viewpoints of the robots and spatial misalignment. These problems bring great challenges to point cloud fusion. In order to overcome such difficulty, in this study, we propose a 3D viewpoint calibration method. The method relies on the novel combination of 3D-SIFT (scale-invariant feature transform) keypoints and FPFH (fast point feature histogram) features, which are used for feature matching. Based on the feature correspondences, we compute the transformation to resolve the difference of the viewpoints. The experiments demonstrate that (1) the keypoints and features used in our method are distinctive and robust to camera viewpoint change; and (2) using viewpoint calibration can reduce the iterations of registration algorithms, and transforming different viewpoints to a close one could save computation time and improve accuracy. Liang Zhang 0010, Juan Song, Peiyi Shen, Guangming Zhu 0001, Shaokai Dong |
ICIP | 1 |
| 2016 | Depth enhancement with improved exemplar-based inpainting and joint trilateral guided filteringabstractHeavy noises and large amounts of holes exist in depth images captured by sensors, such as Kinect, which would severely hinder the application of depth information. In this paper, a novel depth enhancement algorithm with improved exemplar-based inpainting and joint trilateral guided filtering is proposed. The improved examplar-based inpainting method is applied to fill the holes in the depth images, in which the level set distance component is introduced in the priority evaluation function. Then a joint trilateral guided filter is adopted to denoise and smooth the inpainted results. Experimental results reveal that the proposed algorithm can achieve better enhancement results compared with the existing methods in terms of subjective and objective quality measurements. Liang Zhang 0010, Peiyi Shen, Shu'e Zhang, Juan Song, Guangming Zhu 0001 |
ICIP | 1 |
| 2016 | Large-scale Isolated Gesture Recognition using pyramidal 3D convolutional networksabstractHuman gesture recognition is one of the central research fields of computer vision, and effective gesture recognition is still challenging up to now. In this paper, we present a pyramidal 3D convolutional network framework for large-scale isolated human gesture recognition. 3D convolutional networks are utilized to learn the spatiotemporal features from gesture video files. Pyramid input is proposed to preserve the multi-scale contextual information of gestures, and each pyramid segment is uniformly sampled with temporal jitter. Pyramid fusion layers are inserted into the 3D convolutional networks to fuse the features of pyramid input. This strategy makes the networks recognize human gestures from the entire video files, not just from segmented clips independently. We present the experiment results on the 2016 ChaLearn LAP Large-scale Isolated Gesture Recognition Challenge, in which we placed third. Guangming Zhu 0001, Liang Zhang 0010, Lin Mei 0001, Jie Shao 0013, Juan Song, Peiyi Shen |
ICPR | 2 |
| 2016 | Human activity recognition based on weighted limb featuresabstractHuman activity recognition plays an important role in personal assistive robot, being able to recognize human activity and perform corresponding assistive action is a great challenges for personal assistive robot. Human body is an articulated system of rigid segments that can be divided into five parts, but many existing methods always identify actions based on the motion trajectories of whole body. In this paper, taking into account the fact that most actions can be performed by a few limbs and the other limbs should not impact on the action recognition, we proposed an activity recognition method based on limb weights. The weight of each limb is composed of consistency weight and uniqueness weight, which are learned according to the similarity degree among different sequences for each specific action. The covariance descriptor, which is the concatenation of eigenvalues extracted from covariance matrices, is adopted to represent the motion trajectory of each limb. In order to distinguish action instances from each other in the feature sequences, a simple annotation method is used. Experimental results on the Cornell activity dataset and the Lab dataset show that the proposed method not only can outperform the state-of-the-art algorithms, but also is appropriate to recognize the actions whose non-core limbs' trajectories are different from each other. Liang Zhang 0010, Wenhan Yang, Guangming Zhu 0001, Peiyi Shen, Juan Song |
IROS | 1 |
| 2016 | Simultaneous enhancement and noise reduction of a single low-light imageabstractImages obtained under low‐light conditions tend to have the characteristics of low‐grey levels, high‐noise levels, and indistinguishable details. Image degradation not only affects the recognition of images, but also influences the performance of the computer vision system. The low‐light image enhancement algorithm based on the dark channel prior de‐hazing technique can enhance the contrast of images effectively and can highlight the details of images. However, the dark channel prior de‐hazing technique ignores the effects of noise, which leads to significant noise amplification after the enhancement process. In this study, a de‐hazing‐based simultaneous enhancement and noise reduction algorithm of are proposed by analysing the essence of the dark channel prior de‐hazing technique and bilateral filter. First, the authors estimate the values of the initial parameters of the hazy image model by de‐hazing technique. Then, they correct the parameters of the hazy image model alternately with the iterative joint bilateral filter. Experimental results indicate that the proposed algorithm can simultaneously enhance the low‐light images and reduce noise effectively. The proposed algorithm could also perform quite well compared with the current common image enhancement and noise reduction algorithms in terms of the subjective visual effects and objective quality assessments. Liang Zhang 0010, Peiyi Shen, Xilu Peng, Guangming Zhu 0001, Juan Song, Wei Wei 0006, Houbing Song |
IET Image Process. | 1 |
| 2016 | An algorithm combined with color differential models for license-plate location
Yuanmei Tian, Juan Song, Peiyi Shen, Liang Zhang 0010, Weibin Gong, Wei Wei 0006, Guangming Zhu 0001 |
Neurocomputing | 5 |
| 2016 | Human action recognition using multi-layer codebooks of key poses and atomic motions
Guangming Zhu 0001, Liang Zhang 0010, Peiyi Shen, Juan Song |
Signal Process. Image Commun. | 2 |
| 2012 | Enhancement and noise reduction of very low light level images
Peiyi Shen, Lingli Luo, Liang Zhang 0010, Juan Song |
ICPR | 4 |