Sheng Liu 0002

dblp:03/5747-2 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0001-8082-0903ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SignDAGC: Dynamic axial graph structure for continuous sign language recognition and translation
Hong-Xiang Hu, Xuhua Yang 0001, Gang-Feng Ma, Sheng Liu 0002, Yuan Feng 0002
Pattern Recognit.5
2025 Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition
abstract
In continuous sign language recognition (CSLR), 2D-CNN-based extractors are often insufficiently trained for spatial capture and struggle with temporal modeling. This leads to incomplete spatial discrimination, hindering the understanding actions across frames. To address these limitations, we propose Hierarchical Spatial-Temporal Enhancement network (HSTE) through two key modules: Cross-scale Semantic Alignment (CSA) and Temporal Extension Shift (TES). CSA innovatively utilizes multi-scale features generated within the network, enriching feature representation through semantic alignment across scales. By integrating a novel temporal shift strategy with dilated convolutions, TES expands the receptive field and captures temporal changes between frames. These modules work independently and are hierarchically integrated into the network in a plug-and-play manner. Extensive experiments show that our method achieves state-of-the-art performance on the challenging CSLR benchmarks: PHOENIX14, PHOENIX14-T, and CSL-Daily. Code will be available at https://github.com/justlis/HSTENet.
Sheng Liu 0002, Yuan Feng 0002, Yiheng Yu, Zhelun Jin, Xuhua Yang 0001
ICASSP2
2025 Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces
abstract
Current continuous sign language recognition (CSLR) methods typically rely on single or adjacent frames for calculations, which can overlook broader contextual information and result in lower accuracy. To address this issue, we introduce CVSign, which constructs an extended contextual space frame by frame while enabling comprehensive cross-frame interaction. Specifically, we present two innovative modules: Contextual Correspondence Awareness (CCA) and Contextual Variability Awareness (CVA). CCA enhances the relevance of contextual features by utilizing cross-frame multi-head query attention to identify and prioritize related areas while suppressing irrelevant regions. CVA captures motion changes at varying speeds by employing difference calculations between multiple frames, effectively minimizing static redundancy. Remarkably, experimental results show that CVSign outperforms the previous state-of-the-art method by a clear margin on widely used datasets, including PHOENIX14, PHOENIX14-T, and CSL-Daily.
Yiheng Yu, Sheng Liu 0002, Yuan Feng 0002, Zhelun Jin, Xuhua Yang 0001
ICASSP2
2025 Fine-grained cross-modality consistency mining for Continuous Sign Language Recognition
Zhenghao Ke, Sheng Liu 0002, Yuan Feng 0002
Pattern Recognit. Lett.2
2025 Selective directed graph convolutional network for skeleton-based action recognition
Chengyuan Ke, Sheng Liu 0002, Yuan Feng 0002, Shengyong Chen
Pattern Recognit. Lett.2
2025 Cross-Modal Adaptive Prototype Learning for Continuous Sign Language Recognition
Xuhua Yang 0001, Yiyang Weng, Xuanyu Lin, Hong-Xiang Hu, Sheng Liu 0002
IEEE Trans. Circuits Syst. Video Technol.6
2025 Highly Condensed All-MLP Architecture for Long-Term Human Motion Prediction
abstract
In artificial intelligence (AI) scenarios where computational resources are constrained, such as in autonomous driving systems, it is challenging to construct a lightweight model that can accurately predict human motion overextended duration. To tackle this challenge, we introduce a highly condensed all-multilayer perceptron (HCMLP) architecture that is engineered for supreme lightweight efficiency. This design facilitates extended-range motion predictions while maintaining uncompromised performance. First, the spatiotemporal dynamic perception (STDP) block enhances operational efficiency while maintaining a simple structure. In STDP, the distinct but parallel spatial multilayer perceptron (SMLP) and temporal multilayer perceptron (TMLP) simultaneously capture the spatial correlations between pose joints and the temporal dynamics of each joint. The subsequent dynamic aggregation (DA), coupled with the channel multilayer perceptron (CMLP), dynamically consolidates and refines spatial and temporal features, leading to improved predictive accuracy. Second, the multiterm union prediction (MTUP) block directly delivers precise predictions for periods ranging from 0 to 4000 ms, eliminating the need for repetitive short-term (ST) prediction iterations. Our experimental results on the Human3.6M, AMASS, 3DPW, and CMU-Mocap datasets demonstrate that HCMLP outperforms existing state-of-the-art (SOTA) methods in ST prediction, long-term (LT) prediction, and especially in extended and extra extended LT (ELT) predictions, all while utilizing the fewest parameters.
Sheng Liu 0002, Shaobo Zhang 0005, Fei Gao 0014, Yuan Feng 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 POSE-HMR: Heuristic Transformer with Postural Prior Constraints for 3D Human Mesh Reconstruction
abstract
This paper proposes an efficient and lightweight model called PoseHMR to address the interference of irrelevant image features and the issues of model inefficiency in 3D human body mesh reconstruction. PoseHMR uses a transformer-decoder architecture and obtains holistic and regional prior constraints about human posture, which serve as signals for the model throughout the process of human mesh reconstruction. To filter out the useless features extracted from the image, the self-attention module is guided by holistic prior constraints to focus on the area where the human body is located. Likewise, the cross-attention module is guided by regional prior constraints to focus on key sampling points around vertices. Furthermore, after generating key query areas by regional prior constraints, a heuristic fine-tuning strategy is applied to refine the local human mesh effectively. Our model is evaluated on mainstream Human3.6M and 3DPW datasets and achieves a state-of-the-art result with fewer parameters. The codes are available at https://github.com/Sookiep/Pose-HMR.
Songqi Pan, Sheng Liu 0002, Yuan Feng 0002, Yineng Zhang, Xiaopeng Tian
ICASSP2
2024 Dynamic Mutual-Activated Transformer for Human Motion Prediction
abstract
Accurate human motion prediction is vital for diverse artificial intelligence applications, and recent research has yielded substantial advancements. Despite this, the prediction process often encounters abrupt discontinuities and accumulates errors over the long term due to insufficient modeling of spatial and temporal correlations, which significantly impacts predictive accuracy. To tackle these challenges, we introduce the Dynamic Mutual-Activated Transformer (DyMAT). This innovative approach learns spatial correlation among joints in pose and temporal correlation of each joint. It is achieved through separate yet concurrent Pose-wise Spatial Attention (PSA) and Joint-specific Temporal Attention (JTA). The dynamic mutual-activation block (DMA) adeptly combines spatio-temporal features, significantly enhancing DyMAT’s representational capacity. Moreover, we integrate a Temporal Self-Enhancement (TSE) block with JTA, serving as a supplement for refining temporal correlation learning. Our experiments conducted on Human3.6M and CMU Mocap datasets underscore that DyMAT consistently outperforms state-of-the-art methods in terms of prediction accuracy. Code is available at https://github.com/alanzhangv123/DyMAT.
Shaobo Zhang 0005, Sheng Liu 0002, Fei Gao 0014, Yuan Feng 0002
ICASSP2
2024 Cross-Modality Consistency Mining For Continuous Sign Language Recognition with Text-Domain Equivalents
abstract
Continuous Sign Language Recognition (CSLR) approaches share similarities with conventional NLP approaches in which language understanding is involved. However, CSLR approaches face the significant challenge of limited scale and vocabulary in existing datasets. Unlike language models that benefit from extensive training datasets, CSLR models often contend with data constraints, hindering their ability to generalize effectively and consistently capture the rich sign language expressions. To leverage the strong contextual and memorial capabilities of pre-trained language models, in this work, we propose Cross-Modality Consistency (XMC) loss to mine the alignment between the visual model and pre-trained language model, enabling the direct alignment at the gloss level. Towards this, we construct a small-scale gloss description corpus named DCSLG with rich descriptive text. Accompanied by CSL-Daily, text-domain equivalents are made for each video in the dataset, making fine-level alignment possible. The experiment results show that the proposed XMC Loss significantly improves the activations, producing more spatio-temporally accurate and relevant activations. Our approach achieves an average reduction in WER by 3.5%. The code and our gloss description corpus named DCSLG are made publicly available on GitHub1.
Zhenghao Ke, Sheng Liu 0002, Chengyuan Ke, Yuan Feng 0002, Shengyong Chen
ICME2
2024 DPS-Net: Dual-Path Stimulation Network for Continuous Sign Language Recognition
abstract
Continuous Sign Language Recognition (CSLR) is a current research hotspot. However, most existing CSLR methods are suffered from inadequate emphasis on motion information, leading to poor recognition accuracy. To tackle this issue, we propose DPS-Net, a Dual-Path Stimulation Network that captures human motion information for CSLR without relying on extra expensive supervision. The DPS consists of two path of stimulation, Local Volatility Stimulation (LVS) and Global Interpretation Stimulation (GIS). Specifically, LVS calculates stimulation based on motion features, capturing the fine-grained movements of sign language. By focusing on the volatility of these movements, LVS is able to extract and amplify critical motion cues that are often missed by other models. GIS stimulates the model with global features to focus on crucial features at a global level and enhance the model’s perception of the entire sequence. The results of our experiments indicate that our proposed approach achieves notable enhancement over the current methods on three large-scale datasets (PHOENIX14, PHOENIX14-T, and CSL-Daily). Moreover, visualizations demonstrate the effects of DPS-Net on accentuating and capturing the human body movements in successive frames. The code will be available soon.
Xiaopeng Tian, Sheng Liu 0002, Yuan Feng 0002, Yineng Zhang, Songqi Pan
IJCNN2
2024 ATCE: Adaptive Temporal Context Exploitation for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization aims to predict action categories and temporal boundaries in long untrimmed videos with only video-level labels. The insufficient utilization of temporal information is a key factor leading to limited results. Moreover, due to the varying durations of different actions, a uniform temporal sampling strategy struggles to accommodate these diverse temporal contexts, which has been overlooked by previous methods. To address this issue, we propose a novel framework called Adaptive Temporal Context Exploitation (ATCE) to adaptively exploit the temporal contexts of different actions for feature enhancement. Specifically, we introduce an Adaptive Temporal Context Capture (ATCC) module to capture diverse temporal contexts at adaptive sampling scales. This module is mainly implemented by temporal deformable convolution which has a learnable receptive field. To better exploit the learned temporal information, we further propose a Teacher-Guided Modality Consensus (TGMC) module. This module introduced a teacher model to accumulate temporal knowledge across training steps and guide the consistency training of RGB and optical flow modalities. Extensive experiments demonstrate that our proposed ATCE outperforms state-of-the-art methods on two popular benchmarks, THUMOS14 and ActivityNet1.2.
Sheng Liu 0002, Yuan Feng 0002, Xiaopeng Tian, Yineng Zhang, Songqi Pan
IJCNN2
2024 SE-FewDet: Semantic-Enhanced Feature Generation and Prediction Refinement for Few-Shot Object Detection
abstract
Few-Shot Object Detection (FSOD) entails learning from few examples. Due to the lack of data diversity, feature generation emerges as an effective method to improve performance. However, the process of generating diverse features for the novel classes, which introduces excessive intra-class variations of the base classes, resulting in blurring the boundaries between the novel and the base classes. To ensure the diversity and boundary clarity of the generated features, our SE-FewDet explores a new structure called SemVAE to integrate semantic and visual information. This structure allows the generator to strengthen the class-centred representation through cross-modal constraints, thus clarifying the boundaries of different classes while ensuring the enhancement of data diversity. Additionally, our SE-FewDet includes a semantic-enhanced prediction refinement module that accurately filters out potential false positives caused by bounding box offsets, ensuring that only the most reliable detections remain. We evaluate our approach on the PASCAL VOC and MS COCO datasets. With these improvements, SE-FewDet significantly improves detection performance on new classes compared to the baseline (VFA).
Yineng Zhang, Sheng Liu 0002, Yuan Feng 0002, Songqi Pan, Xiaopeng Tian
IJCNN2
2024 Consistency-Driven Cross-Modality Transferring for Continuous Sign Language Recognition
abstract
Sign language consists of a unique grammar and expression system. While Continuous Sign Language Recognition (CSLR) approaches share similarities with conventional NLP approaches in which language understanding is involved, the current works on CSLR usually focus on feature extraction and fusion, neglecting the language semantics, causing false positives and overfitting in gloss detection. In this paper, we propose a novel Consistency-Driven Cross-Modality Transferring (CDCM) mechanism to transfer the language modality to visual modality under a consistency-driven optimization. By progressively reducing the gap between text and visual modality, we are able to stably train CSLR networks. The experiments show the efficacy of our approach, with a notable relative reduction in Word Error Rate of 6.91% on average across multiple datasets. We also demonstrate that our approach contributes corrections to suppress the false peaks on highly related and visually similar glosses while training, making glosses in semantic space distinct, thereby achieving improved overall performance.
Zhenghao Ke, Sheng Liu 0002, Chengyuan Ke, Yuan Feng 0002
SMC2
2024 FP-GCN: A Novel Feature Pyramid Graph Convolutional Network For Skeleton-based Action Recognition
abstract
For skeleton-based action recognition, the aggregation of features among human skeletal joints is a critical factor, which influences recognition accuracy in graph convolutional networks. Existing methods often neglect the extraction of skeletal structure features at different scales, which limits the ability of the model to understand actions. To address this issue, we propose a novel Feature Pyramid Graph Convolutional Network(FP-GCN) that enhances the representational capability of the model by capturing the multi-scale spatial features of the skeleton sequence. In detail, we propose an attention-based graph pooling module that effectively contracts the skeleton to multiple lower-order sub-graphs, which serve as spatial representations of the skeleton at corresponding levels. The original skeleton and these sub-graphs are combined to form the feature pyramid, where joints of each level span in the same semantic space. Additionally, we introduce a graph unpooling module to restore the pooled sub-graphs to their original topology. Moreover, we adopt a multi-loss strategy across different spatial scales, encouraging the model to learn more comprehensive skeletal features. Finally, we validate our proposed model on three large-scale datasets, achieving the highest accuracy compared to state-of-the-art methods. We conduct numerous comparative experiments to verify the effectiveness of modules.
Chengyuan Ke, Sheng Liu 0002, Zhenghao Ke, Yuan Feng 0002
SMC2
2024 HCMLP: A Highly Condensed All-MLP Architecture for Extended Long-term Human Motion Prediction
abstract
Accurate human motion prediction has significant potential in various artificial intelligence applications. To accommodate the demands of applications such as autonomous driving on mobile devices, it is essential to utilize models that are both lightweight and capable of performing extended-duration predictions to ensure the system remains swift and reliable. To address these challenges, we present the HCMLP, a highly condensed all-MLP architecture designed for optimal lightweight efficiency, enabling extended long-term predictions without compromising performance. This pioneering method simultaneously captures the spatial correlations between pose joints and the temporal dynamics of each joint by employing distinct but parallel spatial and temporal MLPs. Then, Dynamic Aggregation component dynamically assimilates the spatial and temporal correlations. Finally, channel MLP synergizes and refines these spatio-temporal features for enhanced prediction accuracy. Our experiments on the Human3.6M, AMASS, and 3DPW datasets reveal that HCMLP surpasses the performance of current state-of-the-art methods in short-term, long-term, and particularly extended long-term predictions, while maintaining the least parameters. Code will be available at https://github.com/alanzhangv123/HCMLP.
Shaobo Zhang 0005, Sheng Liu 0002, Fei Gao 0014, Yuan Feng 0002
SMC2
2024 DFCNet +: Cross-modal dynamic feature contrast net for continuous sign language recognition
Yuan Feng 0002, Nuoyi Chen, Yumeng Wu, Caoyu Jiang, Sheng Liu 0002, Shengyong Chen
Image Vis. Comput.5
2024 Asymmetric Dual-Decoder U-Net for Joint Rain and Haze Removal
abstract
This work studies the multi-weather restoration problem. In real-life scenarios, rain and haze, two often co-occurring common weather phenomena, can greatly degrade the clarity and quality of the scene images, leading to a performance drop in the visual applications, such as autonomous driving. However, jointly removing the rain and haze in scene images is ill-posed and challenging, where the existence of haze and rain and the change of atmosphere light, can both degrade the scene information. Current methods focus on the contamination removal part, thus ignoring the restoration of the scene information affected by the change of atmospheric light. We propose a novel deep neural network, named Asymmetric Dual-decoder U-Net (ADU-Net), to address the aforementioned challenge. The ADU-Net produces both the contamination residual and the scene residual to efficiently remove the contamination while preserving the fidelity of the scene information. Extensive experiments show our work outperforms the existing state-of-the-art methods by a considerable margin in both synthetic data and real-world data benchmarks, including RainCityscapes, BID Rain, and SPA-Data. For instance, we improve the state-of-the-art PSNR value by 2.26/4.57 on the RainCityscapes/SPA-Data, respectively. Codes will be made available freely to the research community.
Yuan Feng 0002, Yaojun Hu, Pengfei Fang, Sheng Liu 0002, Yanhong Yang, Shengyong Chen
ACM Trans. Multim. Comput. Commun. Appl.4
2023 ICDT: Maintaining Interaction Consistency for Deformable Transformer with Multi-scale Features in HOI Detection
Bingnan Guo, Sheng Liu 0002, Ruixiang Chen
ICANN (6)2
2023 IAST: Instance Association Relying on Spatio-Temporal Features for Video Instance Segmentation
abstract
Most offline video instance segmentation (VIS) methods lack consideration for multi-scale spatio-temporal features, which leads to unstable instance association across frames. To address this problem, we propose IAST that builds Instance Association relying on Spatio-Temporal features for video instance segmentation. In detail, we design a novel Scale-to-Scale Attention Module in the encoder of IAST, which constructs stable cross-frame instance associations by completely leveraged multi-scale spatio-temporal features. In addition, we introduce a new data augmentation method called Sequential Copy-Paste, which effectively alleviates the overfitting problem caused by insufficient training data and enhances the robustness of the model. Empirically, IAST achieves the state-of-the-art VIS benchmarks with a ResNet-50 backbone: 47.4% AP, 41.6% AP on YouTube-VIS 2019 & 2021. Such achievements significantly outperform the previous state-of-the-art performance of 1.0% at the expense of fewer parameters. Code is available at https://github.com/clozureyez/IAST.
Sheng Liu 0002, Ruixiang Chen, Bingnan Guo
ICASSP2
2023 SQA: Strong Guidance Query with Self-Selected Attention for Human-Object Interaction Detection
abstract
The attention mechanism in Transformer-based HOI models plays important role in the comprehension of human and object interaction. However, most previous Transformer-based models ignore the guidance on the query and attention, which leads to a poor understanding of interaction behaviour. In this paper, we propose a strong guidance query model with self-selected attention called SQA. The model includes two novel modules, query feature extraction (QFE) and attention mask construction (AMC). QFE builds strong guidance query by concatenating guidance features. The strong guidance query effectively improves the ability to capture both human and object relationships. Meanwhile, AMC establishes distinctive attention masks for each query. The masks allow each query to contact self-selected particular attention regions. It facilitates directing query to obtain more accurate information during cross-attention even in the rare-sample case. We evaluate our SQA model on the mainstream HICO-DET and V-COCO datasets and it achieves a state-of-the-art result. The codes are available at https://github.com/nmbzdwss/SQA.
Sheng Liu 0002, Bingnan Guo, Ruixiang Chen
ICASSP2
2022 PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition
abstract
Due to overly relying on appearance information or adopting direct static feature fusion, most of the existing action recognition methods based on multi-modality have poor robustness and insufficient consideration of modality differences. To address these problems, we propose a two-stream adaptive weight integration network with a three-dimensional parallel attention module, PA-AWCNN. Firstly, a three-dimensional Parallel Attention (PA) module is proposed to effectively extract features of spatial, temporal and channel dimensions and reduce the cross-dimensional interference, to achieve better robustness. Secondly, a Common Feature-driven (CFD) feature integration module is proposed to dynamically integrate appearance and depth features with adaptive weights, utilizing modality differences to redeem the lack of each feature, thereby balance the influence of both. The proposed PA-AW CNN uses the representative integrated feature generated by attention enhancement and feature integration for action recognition; it can not only get higher recognition accuracy but also improve the performance of distinguishing similar actions. Experiments illustrate that the proposed method achieves com-parable performances to state-of-the-art methods and obtains the accuracy of 92.76% and 95.65% on NTU RGB+D Dataset and SBU Kinect Interaction Dataset, respectively. The code is publicly available at: https://github.com/Luu-Yao/PA-AWCNN.
Sheng Liu 0002, Chaonan Li, Siyu Zou, Shengyong Chen, Diyi Guan
ICRA2
2022 HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency
abstract
We present a transformer-based network, Human Mesh Deformer (HMD-former), to tackle the problem of 3D human mesh reconstruction from a single RGB image. HMD-former applies a pre-trained CNN to extract image grid features and a transformer decoder to gradually warp the template 3D mesh to the deformed mesh. On each decoder layer, the fine-grained local information of grid features is well utilized using cross-attention by softly and content-dependently transforming the grid features to vertex embeddings. Auxiliary losses and proposed bi-directional mapping layers inherently ensure semantic consistency throughout the whole decoder, which free the network from learning unnecessary embedding transformation between layers. This further induces each layer of the decoder to focus on refining vertex embeddings and makes the whole network work in a progressively refining manner. Experiments on different public datasets Human3.6M and 3DPW show better reconstruction accuracy and faster inference speed than previous state-of-the-art methods, demonstrating the effectiveness and generalizability of HMD-former. Code is publicly available at https://github.com/siyuzou/HMD-former.
Siyu Zou, Sheng Liu 0002, Chaonan Li, Shengyong Chen
ICRA2
2022 Human Interaction Recognition with Skeletal Attention and Shift Graph Convolution
abstract
Human interaction recognition has wide applications including intelligent surveillance, intelligent transportation and the analysis of sports videos. In recent years, benefiting from the development of action recognition based on deep learning, the performance of human interaction recognition has been boosted. This paper tackles two vital issues in recognizing human interactions, namely target missing and inadequate feature expression. To this end, we first design a data preprocessing method using skeleton estimation and multi-object tracking, which effectively reduces the chance of missing detection. Second, we propose a two-stream network composing of an appearance branch and a pose branch. The appearance branch extracts features enhanced via part affinity maps and part confidences maps, while the pose branch trains a customized Shift-GCN to extract skeletal features from people-pairs. Appearance and pose features are then fused to generate a more powerful representation of human interactions. Extensive experiments on two existing benchmarks, UT and BIT-Interaction, as well as a new dataset crafted by us, namely Campus-Interaction (CI), demonstrate the superior performance of the proposed approach over the state-of-the-arts.
Zhenhua Wang 0003, Jiajun Meng, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen
IJCNN4
2022 SMS-Net: Sparse multi-scale voxel feature aggregation network for LiDAR-based 3D object detection
Sheng Liu 0002, Yifeng Cao, Dingda Li, Shengyong Chen
Neurocomputing1
2021 MRAC-Net: Multi-resolution Anisotropic Convolutional Network for 3D Point Cloud Completion
Sheng Liu 0002, Dingda Li, Yifeng Cao, Shengyong Chen
PRICAI (3)1
2021 CT-UNet: Context-Transfer-UNet for Building Segmentation in Remote Sensing Images
Sheng Liu 0002, Huanran Ye, Haohao Cheng
Neural Process. Lett.1
2020 CT-UNet: An Improved Neural Network Based on U-Net for Building Segmentation in Remote Sensing Images
abstract
With the proliferation of remote sensing images, how to segment buildings more accurately in remote sensing images is a critical challenge. First, the high resolution leads to blurred boundaries in the extracted building maps. Second, the similarity between buildings and background results in intra-class inconsistency. To address these two problems, we propose an UNet-based network named Context-Transfer-UNet (CT-UNet). Specifically, we design Dense Boundary Block (DBB). Dense Block utilizes reuse mechanism to refine features and increase recognition capabilities. Boundary Block introduces the low-level spatial information to solve the fuzzy boundary problem. Then, to handle intra-class inconsistency, we construct Spatial Channel Attention Block (SCAB). It combines context space information and selects more distinguishable features from space and channel. Finally, we propose a novel loss function to enhance the purpose of loss by adding evaluation indicator. Based on our proposed CT-UNet, we achieve 85.33% mean IoU on the Inria dataset and 91.00% mean IoU on the WHU dataset, which outperforms our baseline (U-Net ResNet-34) by 3.76% and Web-Net by 2.24%.
Huanran Ye, Sheng Liu 0002, Haohao Cheng
ICPR2
2018 Cross Modal Multiscale Fusion Net for Real-time RGB-D Detection
abstract
This paper presents a novel multi-modal CNN architecture for object detection by exploiting complementary input cues in addition to sole color information. Our one-stage architecture fuses the multiscale mid-level features from two individual feature extractor, so that our end-to-end net can accept cross modal streams to obtain high-precision detection results. In comparison to other cross modal fusion neural networks, our solution successfully reduces runtime to meet the real-time requirement with still high-level accuracy. Experimental evaluation on challenging NYUD2 dataset shows that our network achieves 49.1% mAP, and processes images in real-time at 35.3 frames per second on one single Nvidia GTX1080 GPU. Compared to baseline one stage network SSD on RGB images which gets 39.2% mAP, our method has great accuracy improvement.
Kejie Yin, Sheng Liu 0002, Ruyu Liu, Yibin Chen
ICPR2
2018 Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao
Neurocomputing4
2017 Joint label-interaction learning for human action recognition
abstract
Human interactions and their action categories preserve strong correlations, and the identification of the interaction configuration is of significant importance to improve the action recognition result. However, interactions are typically estimated using heuristics or treated as latent variables. The former usually produces incorrect interaction configuration while the latter introduces challenging training problem. Hence we propose a framework to jointly learn interactions and actions by designing a potential function using both features learned via deep neural networks and human interaction context. We propose an iterative approach to solve the associated inference problem efficiently and approximately. Experimental results on real datasets demonstrate that the proposed approach outperforms baselines by a large margin, and is competitive compared with the state-of-the-arts.
Jiali Jin, Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan
ICIP3
2017 Fast initialization for feature-based monocular slam
abstract
Initial map determines the effect of followed slam tracking. Most feature-based monocular slam initialize their map according to key points matching in close frames. Nevertheless, it will consume lots of computational resources and time. And it is easy to fail in some far scene or close scene. In this paper, we present a fast initialization method to reduce runtime and improve success rate of initialization for feature-based monocular slam. First, vanishing points detection based on line segment detector [1] is adopted. Second, we extract orb key points. And the coordinates of every key points are undistorted and normalized. Third, we generate the corresponding depth for each key point by normalizing its distance to the existing vanishing points or gaussian random number. We compare our method with state-of-the-art on public data sets and ours. The experiments show that our method outperforms on runtime and accuracy.
Shaobo Zhang 0005, Sheng Liu 0002, Jianhua Zhang 0002, Zhenhua Wang 0003, Xiaoyan Wang 0007
ICIP2
2017 A Spatio-Temporal CRF for Human Interaction Understanding
abstract
A better understanding of human interactions in videos can be achieved by simultaneously considering the coarse interactions between people, the action of each individual, and the activity of all people as a whole. We divide the recognition task into two stages. The first stage discriminates interactions and noninteractions, actions and activities based on local image information, while during the second stage, actions and activities are recognized in a global manner based on the local recognition results. A conditional random field (CRF) is designed to model human interactions in the spatio-temporal space. Different from most existing global models which cover either action or activity variables only, our model covers them both by considering the interactions between different types of variables. The graph structure of the CRF is predicted by a model learned from training data, which is different from traditional graph construction methods that typically rely on human heuristics. We learn the parameters of the CRF via structured support vector machine. We propose an efficient inference algorithm to tackle the estimation of labels in long videos containing many people. Our model admits both semantic-level understanding of human interactions in videos and competitive action and activity recognition performance.
Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan
IEEE Trans. Circuits Syst. Video Technol.2
2013 Self-Adaptive Matching In Local Windows For Depth Estimation
abstract
This paper proposes a novel local stereo matching approach based on self-adapting matching window. We improve the accuracy of stereo matching in 3 steps. First, we integrate shape and size information, and construct robust minimum matching windows by applying a self-adapting method. Then, two matching cost optimization strategies are employed for handling both occlusion regions and image borders. Last, we perform a refinement algorithm for obtaining more accurate depth map. Experiment results on the Middlebury stereo image pairs prove that the proposed matching method performs equally well in comparison with other state-of-the-art local approaches.
Haiqiang Jin, Sheng Liu 0002, Xuhua Yang 0001, Shengyong Chen
ECMS2
2013 Image Super-Resolution Reconstruction Using Map Estimation
abstract
This paper presents a promising super-resolution (SR) approach using maximum a posteriori (MAP) estimation. We consider the high resolution (HR) estimation as a Markov Random Field (MRF), using a transformed gradient field prior to repair the image fuzzy problem caused by MRF. An improved Normalized Convolution method is proposed to obtain a first good estimation. We build a reasonable energy function and minimize the posterior energy by gradient descent algorithm. Experimental results on realistic image sequence and comparisons with several other SR techniques show that our approach gives the best results both qualitative and quantitative.
Xin-Long Lu, Shengyong Chen, Xin Wang 0204, Sheng Liu 0002, Chunyan Yao, Xianping Huan
ECMS4