Lianyu Hu 0003

dblp:324/5573 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SSL-SSAW: Self-supervised Learning with Sigmoid Self-attention Weighting for question-based Sign Language Translation
Zekang Liu, Wei Feng 0005, Fanhua Shang, Lianyu Hu 0003, Jichao Feng, Liqing Gao
Pattern Recognit.4
2025 Greg: GEometry-Aware RegIon Refinement for Sign Language Video Generation
Tongkai Shi, Lianyu Hu 0003, Fanhua Shang, Liqing Gao, Wei Feng 0005
ICCV2
2025 Gloss-Free Sign Language Translation With Optical-Flow Guided Two-Stream Network
abstract
Sign Language Translation (SLT) is a challenging task, with existing approaches often constrained by the necessity of gloss annotations (sign language lexemes). Since gloss annotations are expensive and difficult to obtain, they severely limit the scalability of SLT in practical applications. To address this, we propose leveraging inter-frame optical flow as crucial prior information to enhance visual feature extraction, which naturally corresponds to signers’ semantic movements. We design a novel gloss-free optical-Flow Guided Two-Stream Network architecture (FGTSN): The Image Encoder Stream employs optical flow as a static prior to guide spatial attention while the Skeleton Encoder Stream integrates it to enrich the keypoint features and enhance the representation of local temporal dynamics and structural changes. To tackle the alignment challenge inherent in gloss-free SLT, we introduce a novel auxiliary block prediction loss that strengthens the local temporal relationships of the visual features. Experiments on three major datasets demonstrate that our FGTSN achieves state-of-the-art performance with significant improvements over existing methods, providing a novel and lightweight perspective for gloss-free sign language translation.
Lianyu Hu 0003, Tongkai Shi, Fanhua Shang, Jichao Feng, Wei Feng 0005
MMAsia2
2025 A structure-based disentangled network with contrastive regularization for sign language recognition
Liqing Gao, Lei Zhu 0003, Lianyu Hu 0003, Liang Wang 0001, Wei Feng 0005
Expert Syst. Appl.3
2025 Rethinking the temporal downsampling paradigm for continuous sign language recognition
Caifeng Liu, Lianyu Hu 0003
Multim. Syst.2
2025 A large-scale combinatorial benchmark for sign language recognition
Liqing Gao, Liang Wang 0001, Lianyu Hu 0003, Rui-Ze Han, Zekang Liu, Fanhua Shang, Wei Feng 0005
Pattern Recognit.3
2025 DBL-SC: background-independent sign language recognition based on spatial channel separation computation
Zekang Liu, Wei Feng 0005, Liqing Gao, Lianyu Hu 0003
Vis. Comput.4
2024 COMMA: Co-articulated Multi-Modal Learning
abstract
Pretrained large-scale vision-language models such as CLIP have demonstrated excellent generalizability over a series of downstream tasks. However, they are sensitive to the variation of input text prompts and need a selection of prompt templates to achieve satisfactory performance. Recently, various methods have been proposed to dynamically learn the prompts as the textual inputs to avoid the requirements of laboring hand-crafted prompt engineering in the fine-tuning process. We notice that these methods are suboptimal in two aspects. First, the prompts of the vision and language branches in these methods are usually separated or uni-directionally correlated. Thus, the prompts of both branches are not fully correlated and may not provide enough guidance to align the representations of both branches. Second, it's observed that most previous methods usually achieve better performance on seen classes but cause performance degeneration on unseen classes compared to CLIP. This is because the essential generic knowledge learned in the pretraining stage is partly forgotten in the fine-tuning process. In this paper, we propose Co-Articulated Multi-Modal Learning (COMMA) to handle the above limitations. Especially, our method considers prompts from both branches to generate the prompts to enhance the representation alignment of both branches. Besides, to alleviate forgetting about the essential knowledge, we minimize the feature discrepancy between the learned prompts and the embeddings of hand-crafted prompts in the pre-trained CLIP in the late transformer layers. We evaluate our method across three representative tasks of generalization to novel classes, new target datasets and unseen domain shifts. Experimental results demonstrate the superiority of our method by exhibiting a favorable performance boost upon all tasks with high efficiency. Code is available at https://github.com/hulianyuyy/COMMA.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005
AAAI1
2024 Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
abstract
Skeleton-aware sign language recognition (SLR) has gained popularity due to its ability to remain unaffected by background information and its lower computational requirements. Current methods utilize spatial graph modules and temporal modules to capture spatial and temporal features, respectively. However, their spatial graph modules are typically built on fixed graph structures such as graph convolutional networks or a single learnable graph, which only partially explore joint relationships. Additionally, a simple temporal convolution kernel is used to capture temporal information, which may not fully capture the complex movement patterns of different signers. To overcome these limitations, we propose a new spatial architecture consisting of two concurrent branches, which build input-sensitive joint relationships and incorporates specific domain knowledge for recognition, respectively. These two branches are followed by an aggregation process to distinguishe important joint connections. We then propose a new temporal module to model multi-scale temporal information to capture complex human dynamics. Our method achieves state-of-the-art accuracy compared to previous skeleton-aware methods on four large-scale SLR benchmarks. Moreover, our method demonstrates superior accuracy compared to RGB-based methods in most cases while requiring much fewer computational resources, bringing better accuracy-computation trade-off. Code is available at https://github.com/hulianyuyy/DSTA-SLR.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
LREC/COLING1
2024 Pose-Guided Fine-Grained Sign Language Video Generation
Tongkai Shi, Lianyu Hu 0003, Fanhua Shang, Jichao Feng, Wei Feng 0005
ECCV (77)2
2024 Deep Correlated Prompting for Visual Recognition with Missing Modalities
abstract
Large-scale multimodal models have shown excellent performance over a series of tasks powered by the large corpus of paired multimodal training data. Generally, they are always assumed to receive modality-complete inputs. However, this simple assumption may not always hold in the real world due to privacy constraints or collection difficulty, where models pretrained on modality-complete data easily demonstrate degraded performance on missing-modality cases. To handle this issue, we refer to prompt learning to adapt large pretrained multimodal models to handle missing-modality scenarios by regarding different missing cases as different types of input. Instead of only prepending independent prompts to the intermediate layers, we present to leverage the correlations between prompts and input features and excavate the relationships between different layers of prompts to carefully design the instructions. We also incorporate the complementary semantics of different modalities to guide the prompting design for each modality. Extensive experiments on three commonly-used datasets consistently demonstrate the superiority of our method compared to the previous approaches upon different missing scenarios. Plentiful ablations are further given to show the generalizability and reliability of our method upon different modality-missing ratios and types.
Lianyu Hu 0003, Tongkai Shi, Wei Feng 0005, Fanhua Shang
NeurIPS1
2024 Cross-modal knowledge distillation for continuous sign language recognition
Liqing Gao, Lianyu Hu 0003, Jichao Feng, Lei Zhu 0003, Liang Wang 0001, Wei Feng 0005
Neural Networks3
2024 Scalable frame resolution for efficient continuous sign language recognition
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
Pattern Recognit.1
2023 Self-Emphasizing Network for Continuous Sign Language Recognition
abstract
Hand and face play an important role in expressing sign language. Their features are usually especially leveraged to improve system performance. However, to effectively extract visual representations and capture trajectories for hands and face, previous methods always come at high computations with increased training complexity. They usually employ extra heavy pose-estimation networks to locate human body keypoints or rely on additional pre-extracted heatmaps for supervision. To relieve this problem, we propose a self-emphasizing network (SEN) to emphasize informative spatial regions in a self-motivated way, with few extra computations and without additional expensive supervision. Specifically, SEN first employs a lightweight subnetwork to incorporate local spatial-temporal features to identify informative regions, and then dynamically augment original features via attention maps. It's also observed that not all frames contribute equally to recognition. We present a temporal self-emphasizing module to adaptively emphasize those discriminative frames and suppress redundant ones. A comprehensive comparison with previous methods equipped with hand and face features demonstrates the superiority of our method, even though they always require huge computations and rely on expensive extra supervision. Remarkably, with few extra computations, SEN achieves new state-of-the-art accuracy on four large-scale datasets, PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. Visualizations verify the effects of SEN on emphasizing informative spatial and temporal features. Code is available at https://github.com/hulianyuyy/SEN_CSLR
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
AAAI1
2023 Continuous Sign Language Recognition with Correlation Network
abstract
Human body trajectories are a salient cue to identify actions in the video. Such body trajectories are mainly conveyed by hands and face across consecutive frames in sign language. However, current methods in continuous sign language recognition (CSLR) usually process frames independently, thus failing to capture cross-frame trajectories to effectively identify a sign. To handle this limitation, we propose correlation network (CorrNet) to explicitly capture and leverage body trajectories across frames to identify signs. In specific, a correlation module is first proposed to dynamically compute correlation maps between the current frame and adjacent frames to identify trajectories of all spatial patches. An identification module is then presented to dynamically emphasize the body trajectories within these correlation maps. As a result, the generated features are able to gain an overview of local temporal movements to identify a sign. Thanks to its special attention on body trajectories, CorrNet achieves new state-of-the-art accuracy on four largescale datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. A comprehensive comparison with previous spatial-temporal reasoning methods verifies the effectiveness of CorrNet. Visualizations demonstrate the effects of CorrNet on emphasizing human body trajectories across adjacent frames.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
CVPR1
2023 AdaBrowse: Adaptive Video Browser for Efficient Continuous Sign Language Recognition
abstract
Raw videos have been proven to own considerable feature redundancy where in many cases only a portion of frames can already meet the requirements for accurate recognition. In this paper, we are interested in whether such redundancy can be effectively leveraged to facilitate efficient inference in continuous sign language recognition (CSLR). We propose a novel adaptive model (AdaBrowse) to dynamically select a most informative subsequence from input video sequences by modelling this problem as a sequential decision task. In specific, we first utilize a lightweight network to quickly scan input videos to extract coarse features. Then these features are fed into a policy network to intelligently select a subsequence to process. The corresponding subsequence is finally inferred by a normal CSLR model for sentence prediction. As only a portion of frames are processed in this procedure, the total computations can be considerably saved. Besides temporal redundancy, we are also interested in whether the inherent spatial redundancy can be seamlessly integrated together to achieve further efficiency, i.e., dynamically selecting a lowest input resolution for each sample, whose model is referred to as AdaBrowse+. Extensive experimental results on four large-scale CSLR datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily and CSL, demonstrate the effectiveness of AdaBrowse and AdaBrowse+ by achieving comparable accuracy with state-of-the-art methods with 1.44X throughput and 2.12X fewer FLOPs. Comparisons with other commonly-used 2D CNNs and adaptive efficient methods verify the effectiveness of AdaBrowse. Code is available at https://github.com/hulianyuyy/AdaBrowse.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005
ACM Multimedia1
2023 Skeleton-based action recognition with local dynamic spatial-temporal aggregation
Lianyu Hu 0003, Shenglan Liu 0001, Wei Feng 0005
Expert Syst. Appl.1
2023 Difference-guided multi-scale spatial-temporal representation for sign language recognition
Liqing Gao, Lianyu Hu 0003, Fan Lyu, Lei Zhu 0003, Chi-Man Pun, Wei Feng 0005
Vis. Comput.2
2022 Temporal Lift Pooling for Continuous Sign Language Recognition
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
ECCV (35)1