EDBT 2026 Demo / reviewers in the wild / expert
Liqing Gao
dblp:24/11235
· DBLP profile ↗
28ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0003-4518-2154ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 12 since 2021Theory of computation · 3 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSL-SSAW: Self-supervised Learning with Sigmoid Self-attention Weighting for question-based Sign Language Translation
Zekang Liu, Wei Feng 0005, Fanhua Shang, Lianyu Hu 0003, Jichao Feng, Liqing Gao |
Pattern Recognit. | 6 |
| 2025 | Greg: GEometry-Aware RegIon Refinement for Sign Language Video Generation
Tongkai Shi, Lianyu Hu 0003, Fanhua Shang, Liqing Gao, Wei Feng 0005 |
ICCV | 4 |
| 2025 | Unsupervised 3D Coronary Angiography Segmentation Based on Generative Adversarial Networks
Yaoxian Yang, Liqing Gao, Jiabao Wen |
PRCV (13) | 2 |
| 2025 | A structure-based disentangled network with contrastive regularization for sign language recognition
Liqing Gao, Lei Zhu 0003, Lianyu Hu 0003, Liang Wang 0001, Wei Feng 0005 |
Expert Syst. Appl. | 1 |
| 2025 | A large-scale combinatorial benchmark for sign language recognition
Liqing Gao, Liang Wang 0001, Lianyu Hu 0003, Rui-Ze Han, Zekang Liu, Fanhua Shang, Wei Feng 0005 |
Pattern Recognit. | 1 |
| 2025 | DBL-SC: background-independent sign language recognition based on spatial channel separation computation
Zekang Liu, Wei Feng 0005, Liqing Gao, Lianyu Hu 0003 |
Vis. Comput. | 3 |
| 2024 | COMMA: Co-articulated Multi-Modal LearningabstractPretrained large-scale vision-language models such as CLIP have demonstrated excellent generalizability over a series of downstream tasks. However, they are sensitive to the variation of input text prompts and need a selection of prompt templates to achieve satisfactory performance. Recently, various methods have been proposed to dynamically learn the prompts as the textual inputs to avoid the requirements of laboring hand-crafted prompt engineering in the fine-tuning process. We notice that these methods are suboptimal in two aspects. First, the prompts of the vision and language branches in these methods are usually separated or uni-directionally correlated. Thus, the prompts of both branches are not fully correlated and may not provide enough guidance to align the representations of both branches. Second, it's observed that most previous methods usually achieve better performance on seen classes but cause performance degeneration on unseen classes compared to CLIP. This is because the essential generic knowledge learned in the pretraining stage is partly forgotten in the fine-tuning process. In this paper, we propose Co-Articulated Multi-Modal Learning (COMMA) to handle the above limitations. Especially, our method considers prompts from both branches to generate the prompts to enhance the representation alignment of both branches. Besides, to alleviate forgetting about the essential knowledge, we minimize the feature discrepancy between the learned prompts and the embeddings of hand-crafted prompts in the pre-trained CLIP in the late transformer layers. We evaluate our method across three representative tasks of generalization to novel classes, new target datasets and unseen domain shifts. Experimental results demonstrate the superiority of our method by exhibiting a favorable performance boost upon all tasks with high efficiency. Code is available at https://github.com/hulianyuyy/COMMA. Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005 |
AAAI | 2 |
| 2024 | Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language RecognitionabstractSkeleton-aware sign language recognition (SLR) has gained popularity due to its ability to remain unaffected by background information and its lower computational requirements. Current methods utilize spatial graph modules and temporal modules to capture spatial and temporal features, respectively. However, their spatial graph modules are typically built on fixed graph structures such as graph convolutional networks or a single learnable graph, which only partially explore joint relationships. Additionally, a simple temporal convolution kernel is used to capture temporal information, which may not fully capture the complex movement patterns of different signers. To overcome these limitations, we propose a new spatial architecture consisting of two concurrent branches, which build input-sensitive joint relationships and incorporates specific domain knowledge for recognition, respectively. These two branches are followed by an aggregation process to distinguishe important joint connections. We then propose a new temporal module to model multi-scale temporal information to capture complex human dynamics. Our method achieves state-of-the-art accuracy compared to previous skeleton-aware methods on four large-scale SLR benchmarks. Moreover, our method demonstrates superior accuracy compared to RGB-based methods in most cases while requiring much fewer computational resources, bringing better accuracy-computation trade-off. Code is available at https://github.com/hulianyuyy/DSTA-SLR. Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005 |
LREC/COLING | 2 |
| 2024 | Combinational sign language recognition
Liqing Gao, Wei Feng 0005, Fan Lyu, Liang Wang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | Intelligent Decision-Making Method for AUV Path Planning Against Ocean Current Disturbance via Reinforcement LearningabstractWith the development of society and the economy, low-carbon and low-energy means of exploiting marine resources are receiving increasing attention. Autonomous path planning is a fundamental capability for IoT Autonomous Underwater Vehicle (AUV) to carry out ocean exploration tasks. Currently, the main issue lies in the numerous disturbances and uncertainties present in the marine environment during practical applications, which can significantly impact path planning, leading to high energy consumption and carbon emissions. To address this challenge, this paper presents a sustainable reinforcement learning algorithm for handling time-varying current disturbances to achieve low-carbon AUV path planning, which is delineated into three steps. Firstly, a three-dimensional time-varying current environment is established as the environmental framework for reinforcement learning, and the dynamic model of the AUV is formulated. Secondly, to enhance training efficiency and reduce AUV’s energy consumption, this paper puts forth the OCDRP (Ocean Current Disturbance Rejection PPO) algorithm, which incorporates tidal current information to enhance the AUV’s resilience to time-varying currents. Lastly, expectile regression methods are introduced to facilitate the algorithm’s convergence. Experimental results confirm the efficacy of the proposed algorithm and its adaptability to time-varying currents, making it an efficient, adaptable, and low-carbon sustainable path planning approach. Jiabao Wen, Huiao Dai, Jingyi He 0001, Lijiao Sun, Liqing Gao |
IEEE Internet Things J. | 5 |
| 2024 | Sign language translation with hierarchical memorized context in question answering scenarios
Liqing Gao, Wei Feng 0005, Rui-Ze Han, Di Lin 0002, Liang Wang 0001 |
Neural Comput. Appl. | 1 |
| 2024 | Cross-modal knowledge distillation for continuous sign language recognition
Liqing Gao, Lianyu Hu 0003, Jichao Feng, Lei Zhu 0003, Liang Wang 0001, Wei Feng 0005 |
Neural Networks | 1 |
| 2024 | Scalable frame resolution for efficient continuous sign language recognition
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005 |
Pattern Recognit. | 2 |
| 2024 | Overcoming Modality Bias in Question-Driven Sign Language Video TranslationabstractQuestion-Driven Sign Language Translation (QSLT) addresses the challenge of translating sign language using pertinent questions in question-answering contexts. However, the pronounced modality complexity between question text and sign video poses a predicament: the model tends to overly depend on questions to generate translations, thereby neglecting the value of visual cues. To tackle this issue, the paper presents a Gloss-Bridged Translator (GBT), which introduces sign gloss as an intermediary conduit to establish semantic connections between questions and videos. By leveraging gloss, visual features are transformed into textual counterparts, mitigating the modality imbalance between these representations. Moreover, a cross-modal contrastive learning strategy is implemented, bolstering the global contextual relevance and local semantic alignment between questions and sign language. The proposed methodology is validated through extensive experiments on the proposed QSL dataset and other public sign language datasets. The results show the efficacy of integrating questions into sign language translation. The GBT yields remarkable improvements over prevailing SLT methods, attesting to its effectiveness and rationale. Our code and dataset is available athttps://github.com/glq-1992/QSL. Liqing Gao, Fan Lyu, Lei Zhu 0003, Junfu Pu, Liang Wang 0001, Wei Feng 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multi-scale context-aware network for continuous sign language recognitionabstractThe hands and face are the most important parts for expressing sign language morphemes in sign language videos. However, we find that existing Continuous Sign Language Recognition (CSLR) methods lack the mining of hand and face information in visual backbones or use expensive and time-consuming external extractors to explore this information. In addition, the signs have different lengths, whereas previous CSLR methods typically use a fixed-length window to segment the video to capture sequential features and then perform global temporal modeling, which disturbs the perception of complete signs. In this study, we propose a Multi-Scale Context-Aware network (MSCA-Net) to solve the aforementioned problems. Our MSCA-Net contains two main modules: (1) Multi-Scale Motion Attention (MSMA), which uses the differences among frames to perceive information of the hands and face in multiple spatial scales, replacing the heavy feature extractors; and (2) Multi-Scale Temporal Modeling (MSTM), which explores crucial temporal information in the sign language video from different temporal scales. We conduct extensive experiments using three widely used sign language datasets, i.e., RWTH-PHOENIX-Weather-2014, RWTH-PHOENIX-Weather-2014T, and CSL-Daily. The proposed MSCA-Net achieve state-of-the-art performance, demonstrating the effectiveness of our approach. Senhua Xue, Liqing Gao, Wei Feng 0005 |
Virtual Real. Intell. Hardw. | 2 |
| 2023 | Self-Emphasizing Network for Continuous Sign Language RecognitionabstractHand and face play an important role in expressing sign language. Their features are usually especially leveraged to improve system performance. However, to effectively extract visual representations and capture trajectories for hands and face, previous methods always come at high computations with increased training complexity. They usually employ extra heavy pose-estimation networks to locate human body keypoints or rely on additional pre-extracted heatmaps for supervision. To relieve this problem, we propose a self-emphasizing network (SEN) to emphasize informative spatial regions in a self-motivated way, with few extra computations and without additional expensive supervision. Specifically, SEN first employs a lightweight subnetwork to incorporate local spatial-temporal features to identify informative regions, and then dynamically augment original features via attention maps. It's also observed that not all frames contribute equally to recognition. We present a temporal self-emphasizing module to adaptively emphasize those discriminative frames and suppress redundant ones. A comprehensive comparison with previous methods equipped with hand and face features demonstrates the superiority of our method, even though they always require huge computations and rely on expensive extra supervision. Remarkably, with few extra computations, SEN achieves new state-of-the-art accuracy on four large-scale datasets, PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. Visualizations verify the effects of SEN on emphasizing informative spatial and temporal features. Code is available at https://github.com/hulianyuyy/SEN_CSLR Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005 |
AAAI | 2 |
| 2023 | Spatial-Temporal Consistency Constraints for Chinese Sign Language Synthesis
Liqing Gao, Wei Feng 0005 |
CAD/Graphics | 1 |
| 2023 | Continuous Sign Language Recognition with Correlation NetworkabstractHuman body trajectories are a salient cue to identify actions in the video. Such body trajectories are mainly conveyed by hands and face across consecutive frames in sign language. However, current methods in continuous sign language recognition (CSLR) usually process frames independently, thus failing to capture cross-frame trajectories to effectively identify a sign. To handle this limitation, we propose correlation network (CorrNet) to explicitly capture and leverage body trajectories across frames to identify signs. In specific, a correlation module is first proposed to dynamically compute correlation maps between the current frame and adjacent frames to identify trajectories of all spatial patches. An identification module is then presented to dynamically emphasize the body trajectories within these correlation maps. As a result, the generated features are able to gain an overview of local temporal movements to identify a sign. Thanks to its special attention on body trajectories, CorrNet achieves new state-of-the-art accuracy on four largescale datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. A comprehensive comparison with previous spatial-temporal reasoning methods verifies the effectiveness of CorrNet. Visualizations demonstrate the effects of CorrNet on emphasizing human body trajectories across adjacent frames. Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005 |
CVPR | 2 |
| 2023 | AdaBrowse: Adaptive Video Browser for Efficient Continuous Sign Language RecognitionabstractRaw videos have been proven to own considerable feature redundancy where in many cases only a portion of frames can already meet the requirements for accurate recognition. In this paper, we are interested in whether such redundancy can be effectively leveraged to facilitate efficient inference in continuous sign language recognition (CSLR). We propose a novel adaptive model (AdaBrowse) to dynamically select a most informative subsequence from input video sequences by modelling this problem as a sequential decision task. In specific, we first utilize a lightweight network to quickly scan input videos to extract coarse features. Then these features are fed into a policy network to intelligently select a subsequence to process. The corresponding subsequence is finally inferred by a normal CSLR model for sentence prediction. As only a portion of frames are processed in this procedure, the total computations can be considerably saved. Besides temporal redundancy, we are also interested in whether the inherent spatial redundancy can be seamlessly integrated together to achieve further efficiency, i.e., dynamically selecting a lowest input resolution for each sample, whose model is referred to as AdaBrowse+. Extensive experimental results on four large-scale CSLR datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily and CSL, demonstrate the effectiveness of AdaBrowse and AdaBrowse+ by achieving comparable accuracy with state-of-the-art methods with 1.44X throughput and 2.12X fewer FLOPs. Comparisons with other commonly-used 2D CNNs and adaptive efficient methods verify the effectiveness of AdaBrowse. Code is available at https://github.com/hulianyuyy/AdaBrowse. Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005 |
ACM Multimedia | 2 |
| 2023 | Difference-guided multi-scale spatial-temporal representation for sign language recognition
Liqing Gao, Lianyu Hu 0003, Fan Lyu, Lei Zhu 0003, Chi-Man Pun, Wei Feng 0005 |
Vis. Comput. | 1 |
| 2022 | Temporal Lift Pooling for Continuous Sign Language Recognition
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005 |
ECCV (35) | 2 |
| 2022 | Harmful algal bloom warning based on machine learning in maritime site monitoring
Jiabao Wen, Yang Li 0111, Liqing Gao |
Knowl. Based Syst. | 4 |
| 2021 | RNN-Transducer based Chinese Sign Language Recognition
Liqing Gao, Zekang Liu, Wei Feng 0005 |
Neurocomputing | 1 |
| 2020 | Key Action and Joint CTC-Attention based Sign Language RecognitionabstractSign Language Recognition (SLR) translates sign language video into natural language. In practice, sign language video, owning a large number of redundant frames, is necessary to be selected the essential. However, unlike common video that describes actions, sign language video is characterized as continuous and dense action sequence, which is difficult to capture key actions corresponding to meaningful sentence. In this paper, we propose to hierarchically search key actions by a pyramid BiLSTM. Specifically, we first construct three BiL-STMs to produce temporal relationships among input video sequence. Then, we associate these BiLSTMs by searching the salient responses in two groups of fixed-scale sliding window and capture key actions. Additionally, in order to balance the sequence alignment and dependency, we propose to jointly train Connectionist Temporal Classification (CTC) and Long Short-Term Memory (LSTM). Experimental results demonstrate the effectiveness of the proposed method. Liqing Gao, Rui-Ze Han, Wei Feng 0005 |
ICASSP | 2 |
| 2018 | Crowd counting considering network flow constraints in videosabstractThe growth of the number of people in the monitoring scene may increase the probability of security threat, which makes crowd counting more and more important. Most of the existing approaches estimate the number of pedestrians within one frame, which results in inconsistent predictions in terms of time. This study, for the first time, introduces a quadratic programming (QP) model with the network flow constraints to improve the accuracy of crowd counting. Firstly, the foreground of each frame is segmented into groups, each of which contains several pedestrians. Then, a regression‐based map is developed in accordance with the relationship between low‐level features of each group and the number of people in it. Secondly, a directed graph is constructed to simulate constraints on people's flow, whose vertices represent groups of each frame and arcs represent people moving from one group to another. Finally, by solving a QP problem with network flow constraints in the directed graph, the authors obtain consistency in people counting. The experimental results show that the proposed method can reduce the crowd counting errors and improve the accuracy. Moreover, this method can also be applied to any ultramodern group‐based regression counting approach to get improvements. Liqing Gao, Yanzhang Wang, Xin Ye 0004 |
IET Image Process. | 1 |
| 2015 | The decycling number of generalized Petersen graphs
Liqing Gao, Xirong Xu, Dejun Zhu, Yuansheng Yang |
Discret. Appl. Math. | 1 |
| 2015 | Decycling bubble sort graphs
Xirong Xu, Liqing Gao, Yuansheng Yang |
Discret. Appl. Math. | 3 |
| 2012 | On the bounds of feedback numbers of (n, k)-star graphs
Xirong Xu, Dejun Zhu, Liqing Gao, Jun-Ming Xu 0001 |
Inf. Process. Lett. | 4 |