Yuan Feng 0002

dblp:14/6701-2 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0003-2563-8724ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SignDAGC: Dynamic axial graph structure for continuous sign language recognition and translation
Hong-Xiang Hu, Xuhua Yang 0001, Gang-Feng Ma, Sheng Liu 0002, Yuan Feng 0002
Pattern Recognit.6
2025 Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition
abstract
In continuous sign language recognition (CSLR), 2D-CNN-based extractors are often insufficiently trained for spatial capture and struggle with temporal modeling. This leads to incomplete spatial discrimination, hindering the understanding actions across frames. To address these limitations, we propose Hierarchical Spatial-Temporal Enhancement network (HSTE) through two key modules: Cross-scale Semantic Alignment (CSA) and Temporal Extension Shift (TES). CSA innovatively utilizes multi-scale features generated within the network, enriching feature representation through semantic alignment across scales. By integrating a novel temporal shift strategy with dilated convolutions, TES expands the receptive field and captures temporal changes between frames. These modules work independently and are hierarchically integrated into the network in a plug-and-play manner. Extensive experiments show that our method achieves state-of-the-art performance on the challenging CSLR benchmarks: PHOENIX14, PHOENIX14-T, and CSL-Daily. Code will be available at https://github.com/justlis/HSTENet.
Sheng Liu 0002, Yuan Feng 0002, Yiheng Yu, Zhelun Jin, Xuhua Yang 0001
ICASSP3
2025 Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces
abstract
Current continuous sign language recognition (CSLR) methods typically rely on single or adjacent frames for calculations, which can overlook broader contextual information and result in lower accuracy. To address this issue, we introduce CVSign, which constructs an extended contextual space frame by frame while enabling comprehensive cross-frame interaction. Specifically, we present two innovative modules: Contextual Correspondence Awareness (CCA) and Contextual Variability Awareness (CVA). CCA enhances the relevance of contextual features by utilizing cross-frame multi-head query attention to identify and prioritize related areas while suppressing irrelevant regions. CVA captures motion changes at varying speeds by employing difference calculations between multiple frames, effectively minimizing static redundancy. Remarkably, experimental results show that CVSign outperforms the previous state-of-the-art method by a clear margin on widely used datasets, including PHOENIX14, PHOENIX14-T, and CSL-Daily.
Yiheng Yu, Sheng Liu 0002, Yuan Feng 0002, Zhelun Jin, Xuhua Yang 0001
ICASSP3
2025 Fine-grained cross-modality consistency mining for Continuous Sign Language Recognition
Zhenghao Ke, Sheng Liu 0002, Yuan Feng 0002
Pattern Recognit. Lett.3
2025 Selective directed graph convolutional network for skeleton-based action recognition
Chengyuan Ke, Sheng Liu 0002, Yuan Feng 0002, Shengyong Chen
Pattern Recognit. Lett.3
2025 Highly Condensed All-MLP Architecture for Long-Term Human Motion Prediction
abstract
In artificial intelligence (AI) scenarios where computational resources are constrained, such as in autonomous driving systems, it is challenging to construct a lightweight model that can accurately predict human motion overextended duration. To tackle this challenge, we introduce a highly condensed all-multilayer perceptron (HCMLP) architecture that is engineered for supreme lightweight efficiency. This design facilitates extended-range motion predictions while maintaining uncompromised performance. First, the spatiotemporal dynamic perception (STDP) block enhances operational efficiency while maintaining a simple structure. In STDP, the distinct but parallel spatial multilayer perceptron (SMLP) and temporal multilayer perceptron (TMLP) simultaneously capture the spatial correlations between pose joints and the temporal dynamics of each joint. The subsequent dynamic aggregation (DA), coupled with the channel multilayer perceptron (CMLP), dynamically consolidates and refines spatial and temporal features, leading to improved predictive accuracy. Second, the multiterm union prediction (MTUP) block directly delivers precise predictions for periods ranging from 0 to 4000 ms, eliminating the need for repetitive short-term (ST) prediction iterations. Our experimental results on the Human3.6M, AMASS, 3DPW, and CMU-Mocap datasets demonstrate that HCMLP outperforms existing state-of-the-art (SOTA) methods in ST prediction, long-term (LT) prediction, and especially in extended and extra extended LT (ELT) predictions, all while utilizing the fewest parameters.
Sheng Liu 0002, Shaobo Zhang 0005, Fei Gao 0014, Yuan Feng 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 POSE-HMR: Heuristic Transformer with Postural Prior Constraints for 3D Human Mesh Reconstruction
abstract
This paper proposes an efficient and lightweight model called PoseHMR to address the interference of irrelevant image features and the issues of model inefficiency in 3D human body mesh reconstruction. PoseHMR uses a transformer-decoder architecture and obtains holistic and regional prior constraints about human posture, which serve as signals for the model throughout the process of human mesh reconstruction. To filter out the useless features extracted from the image, the self-attention module is guided by holistic prior constraints to focus on the area where the human body is located. Likewise, the cross-attention module is guided by regional prior constraints to focus on key sampling points around vertices. Furthermore, after generating key query areas by regional prior constraints, a heuristic fine-tuning strategy is applied to refine the local human mesh effectively. Our model is evaluated on mainstream Human3.6M and 3DPW datasets and achieves a state-of-the-art result with fewer parameters. The codes are available at https://github.com/Sookiep/Pose-HMR.
Songqi Pan, Sheng Liu 0002, Yuan Feng 0002, Yineng Zhang, Xiaopeng Tian
ICASSP3
2024 Dynamic Mutual-Activated Transformer for Human Motion Prediction
abstract
Accurate human motion prediction is vital for diverse artificial intelligence applications, and recent research has yielded substantial advancements. Despite this, the prediction process often encounters abrupt discontinuities and accumulates errors over the long term due to insufficient modeling of spatial and temporal correlations, which significantly impacts predictive accuracy. To tackle these challenges, we introduce the Dynamic Mutual-Activated Transformer (DyMAT). This innovative approach learns spatial correlation among joints in pose and temporal correlation of each joint. It is achieved through separate yet concurrent Pose-wise Spatial Attention (PSA) and Joint-specific Temporal Attention (JTA). The dynamic mutual-activation block (DMA) adeptly combines spatio-temporal features, significantly enhancing DyMAT’s representational capacity. Moreover, we integrate a Temporal Self-Enhancement (TSE) block with JTA, serving as a supplement for refining temporal correlation learning. Our experiments conducted on Human3.6M and CMU Mocap datasets underscore that DyMAT consistently outperforms state-of-the-art methods in terms of prediction accuracy. Code is available at https://github.com/alanzhangv123/DyMAT.
Shaobo Zhang 0005, Sheng Liu 0002, Fei Gao 0014, Yuan Feng 0002
ICASSP4
2024 Cross-Modality Consistency Mining For Continuous Sign Language Recognition with Text-Domain Equivalents
abstract
Continuous Sign Language Recognition (CSLR) approaches share similarities with conventional NLP approaches in which language understanding is involved. However, CSLR approaches face the significant challenge of limited scale and vocabulary in existing datasets. Unlike language models that benefit from extensive training datasets, CSLR models often contend with data constraints, hindering their ability to generalize effectively and consistently capture the rich sign language expressions. To leverage the strong contextual and memorial capabilities of pre-trained language models, in this work, we propose Cross-Modality Consistency (XMC) loss to mine the alignment between the visual model and pre-trained language model, enabling the direct alignment at the gloss level. Towards this, we construct a small-scale gloss description corpus named DCSLG with rich descriptive text. Accompanied by CSL-Daily, text-domain equivalents are made for each video in the dataset, making fine-level alignment possible. The experiment results show that the proposed XMC Loss significantly improves the activations, producing more spatio-temporally accurate and relevant activations. Our approach achieves an average reduction in WER by 3.5%. The code and our gloss description corpus named DCSLG are made publicly available on GitHub1.
Zhenghao Ke, Sheng Liu 0002, Chengyuan Ke, Yuan Feng 0002, Shengyong Chen
ICME4
2024 DPS-Net: Dual-Path Stimulation Network for Continuous Sign Language Recognition
abstract
Continuous Sign Language Recognition (CSLR) is a current research hotspot. However, most existing CSLR methods are suffered from inadequate emphasis on motion information, leading to poor recognition accuracy. To tackle this issue, we propose DPS-Net, a Dual-Path Stimulation Network that captures human motion information for CSLR without relying on extra expensive supervision. The DPS consists of two path of stimulation, Local Volatility Stimulation (LVS) and Global Interpretation Stimulation (GIS). Specifically, LVS calculates stimulation based on motion features, capturing the fine-grained movements of sign language. By focusing on the volatility of these movements, LVS is able to extract and amplify critical motion cues that are often missed by other models. GIS stimulates the model with global features to focus on crucial features at a global level and enhance the model’s perception of the entire sequence. The results of our experiments indicate that our proposed approach achieves notable enhancement over the current methods on three large-scale datasets (PHOENIX14, PHOENIX14-T, and CSL-Daily). Moreover, visualizations demonstrate the effects of DPS-Net on accentuating and capturing the human body movements in successive frames. The code will be available soon.
Xiaopeng Tian, Sheng Liu 0002, Yuan Feng 0002, Yineng Zhang, Songqi Pan
IJCNN3
2024 ATCE: Adaptive Temporal Context Exploitation for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization aims to predict action categories and temporal boundaries in long untrimmed videos with only video-level labels. The insufficient utilization of temporal information is a key factor leading to limited results. Moreover, due to the varying durations of different actions, a uniform temporal sampling strategy struggles to accommodate these diverse temporal contexts, which has been overlooked by previous methods. To address this issue, we propose a novel framework called Adaptive Temporal Context Exploitation (ATCE) to adaptively exploit the temporal contexts of different actions for feature enhancement. Specifically, we introduce an Adaptive Temporal Context Capture (ATCC) module to capture diverse temporal contexts at adaptive sampling scales. This module is mainly implemented by temporal deformable convolution which has a learnable receptive field. To better exploit the learned temporal information, we further propose a Teacher-Guided Modality Consensus (TGMC) module. This module introduced a teacher model to accumulate temporal knowledge across training steps and guide the consistency training of RGB and optical flow modalities. Extensive experiments demonstrate that our proposed ATCE outperforms state-of-the-art methods on two popular benchmarks, THUMOS14 and ActivityNet1.2.
Sheng Liu 0002, Yuan Feng 0002, Xiaopeng Tian, Yineng Zhang, Songqi Pan
IJCNN3
2024 SE-FewDet: Semantic-Enhanced Feature Generation and Prediction Refinement for Few-Shot Object Detection
abstract
Few-Shot Object Detection (FSOD) entails learning from few examples. Due to the lack of data diversity, feature generation emerges as an effective method to improve performance. However, the process of generating diverse features for the novel classes, which introduces excessive intra-class variations of the base classes, resulting in blurring the boundaries between the novel and the base classes. To ensure the diversity and boundary clarity of the generated features, our SE-FewDet explores a new structure called SemVAE to integrate semantic and visual information. This structure allows the generator to strengthen the class-centred representation through cross-modal constraints, thus clarifying the boundaries of different classes while ensuring the enhancement of data diversity. Additionally, our SE-FewDet includes a semantic-enhanced prediction refinement module that accurately filters out potential false positives caused by bounding box offsets, ensuring that only the most reliable detections remain. We evaluate our approach on the PASCAL VOC and MS COCO datasets. With these improvements, SE-FewDet significantly improves detection performance on new classes compared to the baseline (VFA).
Yineng Zhang, Sheng Liu 0002, Yuan Feng 0002, Songqi Pan, Xiaopeng Tian
IJCNN3
2024 Consistency-Driven Cross-Modality Transferring for Continuous Sign Language Recognition
abstract
Sign language consists of a unique grammar and expression system. While Continuous Sign Language Recognition (CSLR) approaches share similarities with conventional NLP approaches in which language understanding is involved, the current works on CSLR usually focus on feature extraction and fusion, neglecting the language semantics, causing false positives and overfitting in gloss detection. In this paper, we propose a novel Consistency-Driven Cross-Modality Transferring (CDCM) mechanism to transfer the language modality to visual modality under a consistency-driven optimization. By progressively reducing the gap between text and visual modality, we are able to stably train CSLR networks. The experiments show the efficacy of our approach, with a notable relative reduction in Word Error Rate of 6.91% on average across multiple datasets. We also demonstrate that our approach contributes corrections to suppress the false peaks on highly related and visually similar glosses while training, making glosses in semantic space distinct, thereby achieving improved overall performance.
Zhenghao Ke, Sheng Liu 0002, Chengyuan Ke, Yuan Feng 0002
SMC4
2024 FP-GCN: A Novel Feature Pyramid Graph Convolutional Network For Skeleton-based Action Recognition
abstract
For skeleton-based action recognition, the aggregation of features among human skeletal joints is a critical factor, which influences recognition accuracy in graph convolutional networks. Existing methods often neglect the extraction of skeletal structure features at different scales, which limits the ability of the model to understand actions. To address this issue, we propose a novel Feature Pyramid Graph Convolutional Network(FP-GCN) that enhances the representational capability of the model by capturing the multi-scale spatial features of the skeleton sequence. In detail, we propose an attention-based graph pooling module that effectively contracts the skeleton to multiple lower-order sub-graphs, which serve as spatial representations of the skeleton at corresponding levels. The original skeleton and these sub-graphs are combined to form the feature pyramid, where joints of each level span in the same semantic space. Additionally, we introduce a graph unpooling module to restore the pooled sub-graphs to their original topology. Moreover, we adopt a multi-loss strategy across different spatial scales, encouraging the model to learn more comprehensive skeletal features. Finally, we validate our proposed model on three large-scale datasets, achieving the highest accuracy compared to state-of-the-art methods. We conduct numerous comparative experiments to verify the effectiveness of modules.
Chengyuan Ke, Sheng Liu 0002, Zhenghao Ke, Yuan Feng 0002
SMC4
2024 HCMLP: A Highly Condensed All-MLP Architecture for Extended Long-term Human Motion Prediction
abstract
Accurate human motion prediction has significant potential in various artificial intelligence applications. To accommodate the demands of applications such as autonomous driving on mobile devices, it is essential to utilize models that are both lightweight and capable of performing extended-duration predictions to ensure the system remains swift and reliable. To address these challenges, we present the HCMLP, a highly condensed all-MLP architecture designed for optimal lightweight efficiency, enabling extended long-term predictions without compromising performance. This pioneering method simultaneously captures the spatial correlations between pose joints and the temporal dynamics of each joint by employing distinct but parallel spatial and temporal MLPs. Then, Dynamic Aggregation component dynamically assimilates the spatial and temporal correlations. Finally, channel MLP synergizes and refines these spatio-temporal features for enhanced prediction accuracy. Our experiments on the Human3.6M, AMASS, and 3DPW datasets reveal that HCMLP surpasses the performance of current state-of-the-art methods in short-term, long-term, and particularly extended long-term predictions, while maintaining the least parameters. Code will be available at https://github.com/alanzhangv123/HCMLP.
Shaobo Zhang 0005, Sheng Liu 0002, Fei Gao 0014, Yuan Feng 0002
SMC4
2024 DFCNet +: Cross-modal dynamic feature contrast net for continuous sign language recognition
Yuan Feng 0002, Nuoyi Chen, Yumeng Wu, Caoyu Jiang, Sheng Liu 0002, Shengyong Chen
Image Vis. Comput.1
2024 Asymmetric Dual-Decoder U-Net for Joint Rain and Haze Removal
abstract
This work studies the multi-weather restoration problem. In real-life scenarios, rain and haze, two often co-occurring common weather phenomena, can greatly degrade the clarity and quality of the scene images, leading to a performance drop in the visual applications, such as autonomous driving. However, jointly removing the rain and haze in scene images is ill-posed and challenging, where the existence of haze and rain and the change of atmosphere light, can both degrade the scene information. Current methods focus on the contamination removal part, thus ignoring the restoration of the scene information affected by the change of atmospheric light. We propose a novel deep neural network, named Asymmetric Dual-decoder U-Net (ADU-Net), to address the aforementioned challenge. The ADU-Net produces both the contamination residual and the scene residual to efficiently remove the contamination while preserving the fidelity of the scene information. Extensive experiments show our work outperforms the existing state-of-the-art methods by a considerable margin in both synthetic data and real-world data benchmarks, including RainCityscapes, BID Rain, and SPA-Data. For instance, we improve the state-of-the-art PSNR value by 2.26/4.57 on the RainCityscapes/SPA-Data, respectively. Codes will be made available freely to the research community.
Yuan Feng 0002, Yaojun Hu, Pengfei Fang, Sheng Liu 0002, Yanhong Yang, Shengyong Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Hyperspectral Image Restoration via Subspace-Based Nonlocal Low-Rank Tensor Approximation
abstract
In this letter, we present a subspace-based nonlocal low-rank tensor approximation framework (SNLRTA) for hyperspectral image (HSI) restoration. The proposed method consists of a subspace learning method to achieve an accurate subspace characterization of HSI and a nonlocal low-rank tensor approximation to take spatial nonlocal self-similarity into consideration. Specifically, the HSI first exploits residual statistics on median filtered image to estimate a robust subspace. Laplacian scale mixture (LSM) modeling is then investigated to model tensor coefficients from overlapping cubes in low-rank subspace. Both the hidden scale parameters and the sparse coefficients therein are adaptively shrink, characterizing the sparsity of similar patches. Meanwhile, the$\ell _{1}$data fidelity facilitates the implicit detection of outliers after median filtering. Substantiated by extensive experimental results, the proposed method outperforms several state-of-the-art approaches on mixed noise removal, qualitatively and quantitatively.
Yanhong Yang, Yuan Feng 0002, Jianhua Zhang 0002, Shengyong Chen
IEEE Geosci. Remote. Sens. Lett.2