VLDB 2026 Research / reviewers in the wild / expert
Tiantian Yuan
dblp:46/2025
· DBLP profile ↗
15ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ADASign: Adaptive deformable visual attention for continuous sign language recognition
Xuyan Zhang, Huayu Ma, Wanli Xue, Leming Guo, Tiantian Yuan, Shengyong Chen |
Neurocomputing | 7 |
| 2026 | Frequency transform attack: a transferable adversarial framework for continuous sign language recognition
Yachao Lin, Wanli Xue, Leming Guo, Tiantian Yuan |
Multim. Syst. | 6 |
| 2026 | Trustworthy Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) uses visual cues (e.g., hands, face, mouth, and body) to automatically recognize the sign language of deaf people, helping them to actively communicate with hearing people. The effects of these visual cues change dynamically with the demonstration of sign language. However, previous CSLR methods usually model visual information from the entire frame or simple fused visual cues, and thus do not well describe such dynamic change among visual cues. Therefore, we propose the Trustworthy Fusion Network ( TFN) of visual cues for CSLR, which comprises two fundamental modules: Intra-cue Cross-modality Feature Fusion module ( IntraCFF) and Inter-cue Trustworthy Fusion module (InterTF). IntraCFF uses the calibrated joint-belief method to dynamically fuse cross-modality features of RGB and keypoint information, to obtain a robust visual cue feature. InterTF innovatively employs the Dempster-Shafer Theory (DST) to evaluate the uncertainty of different cues in expressing sign movements. Then, the trustworthy fusion via DST is used to adaptively weigh and credibly fuse the visual cues based on uncertainty. In addition, to address the semantic gap when fusing different cues, we design consistency fusion constraints during the training stage. These constraints enhance the semantic consistency of different cues with global sign movements. Experiments on publicly CSLR datasets validate the effectiveness of our TFN. Yan Zhang 0154, Wanli Xue, Leming Guo, Yangcan Wu, Tiantian Yuan, Shengyong Chen |
IEEE Trans. Multim. | 7 |
| 2025 | Dynamic Multi-Scale Spatial-Motion Feature Learning for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) translates sign language videos into gloss sequences. Most CSLR models are trained using Connectionist Temporal Classification (CTC) loss, which may lead to insufficient feature extraction. To enhance the feature extractor, we propose a dynamic multi-scale spatial-motion feature learning method, introducing two key modules: the Multi-Scale Spatial Feature Learning Module (MSSFLM) and the Multi-Scale Motion Feature Learning Module (MSMFLM). The MSSFLM employs dilated convolution and Fast Fourier Transform (FFT) to extract spatial features and applies spatial attention to highlight key regions. The MSMFLM captures motion features via frame-by-frame and cross-N-frame differences, followed by self-attention to emphasize salient motions. Experiments on PHOENIX14, PHOENIX14-T, and CSL-Daily demonstrate the effectiveness of our method. Cui-Hong Xue, Tiantian Yuan, Longcan Yuan |
IJCNN | 3 |
| 2025 | Dual-Path Collaborative Optimization Network for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize gesture sequences in sign language videos. However, traditional CSLR methods face the challenges of the motion and long-term spatio-temporal features modeling. To address this problem, we propose a dual-path collaborative optimization network (DCO-Net), it balances long-term dependency and fast motion feature modeling by the dual-path CSLR. In DCO-Net, the Channel-Spatial Optimization (CSO) generates dynamic weights through a multi-path pooling strategy to focus on key regions; the Sequential Spatio-Temporal Attention (SSTA) module enhances slow motion and long-term dependency modeling; the Parallel Spatio-Temporal Attention (PSTA) module captures fast-changing features while preserving fine-grained spatial information; Horizontal Cross Connects (HCC) facilitate feature interaction, and the Fusion Soft Filter (FSF) module fuses features from both paths. Extensive experiments on public datasets (PHOENIX14, PHOENIX14-T and CSL-Daily) show that our method outperform the state-of-the-art performance in CSLR. Cui-Hong Xue, Longcan Yuan, Tiantian Yuan |
IJCNN | 3 |
| 2025 | CORE: Multi-link graph attention network with inter-regional collaboration for continuous sign language recognition
Yan Zhang 0154, Wanli Xue, Tiantian Yuan, Shengyong Chen |
Pattern Recognit. | 4 |
| 2025 | Dual-stage temporal perception network for continuous sign language recognition
Wanli Xue, Jinlu Sun, Yazhou Wu, Tiantian Yuan, Shengyong Chen |
Vis. Comput. | 6 |
| 2024 | CoSLR: Contrastive Chinese Sign Language Recognition with prior knowledge And Multi-Tasks Joint LearningabstractPerceiving by computer vision, Sign Language Recognition (SLR) obtains the advantage of transforming the posture video into a sentence, compared with the methods of sensors to collect signals. However, learning representative features from a multimodal perspective is challenging. To this end, this study proposes a multi-task joint learning framework termed Contrastive Learning-based Sign Language Recognition Network (CoSLR) for Chinese sign language, which embeds text representation into the general video-based framework of SLR. In virtue of the profound ability of the pre-trained multimodal encoders, they are employed as processing modules to extract features from the original input video and text. Then, a contrastive learning between the video representation and corresponding token embedding is utilized to the feature extractor training. Finally, the linear combination of contrast and cross-entropy loss functions drives the end-to-end network to converge. Experiments show that the 1.27% WER of CoSLR has outperformed the state-of-the-art works in the comparison. Tiantian Yuan |
ICASSP | 3 |
| 2024 | Gloss Prior Guided Visual Feature Learning for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) is to recognize the glosses in a sign language video. Enhancing the generalization ability of CSLR's visual feature extractor is a worthy area of investigation. In this paper, we model glosses as priors that help to learn more generalizable visual features. Specifically, the signer-invariant gloss feature is extracted by a pre-trained gloss BERT model. Then we design a gloss prior guidance network (GPGN). It contains a novel parallel densely-connected temporal feature extraction (PDC-TFE) module for multi-resolution visual feature extraction. The PDC-TFE captures the complex temporal patterns of the glosses. The pre-trained gloss feature guides the visual feature learning through a cross-modality matching loss. We propose to formulate the cross-modality feature matching into a regularized optimal transport problem, it can be efficiently solved by a variant of the Sinkhorn algorithm. The GPGN parameters are learned by optimizing a weighted sum of the cross-modality matching loss and CTC loss. The experiment results on German and Chinese sign language benchmarks demonstrate that the proposed GPGN achieves competitive performance. The ablation study verifies the effectiveness of several critical components of the GPGN. Furthermore, the proposed pre-trained gloss BERT model and cross-modality matching can be seamlessly integrated into other RGB-cue-based CSLR methods as plug-and-play formulations to enhance the generalization ability of the visual feature extractor. Leming Guo, Wanli Xue, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 5 |
| 2023 | Distilling Cross-Temporal Contexts for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the over-all model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thor-ough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a cross-temporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distil-lation learning objective to aggregate the two types of con-text and the linguistic prior. The knowledge distillation en-ables the resultant one-branch temporal aggregation mod-ule to perceive local-global temporal and semantic context. This shallow temporal perception module structure facili-tates spatial perception module learning. Extensive exper-iments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods. Leming Guo, Wanli Xue, Qing Guo 0005, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Shengyong Chen |
CVPR | 6 |
| 2022 | Multi-level Temporal Relation Graph for Continuous Sign Language Recognition
Wanli Xue, Leming Guo, Tiantian Yuan, Shengyong Chen |
PRCV (3) | 4 |
| 2021 | CRB-Net: A Sign Language Recognition Deep Learning Strategy Based on Multi-modal Fusion with Attention MechanismabstractAt present, sign language recognition (SLR) researchers are mainly committed to establishing a sign language recognition model based on single-mode data. Nevertheless, this manipulation often leads to a defective understanding of the sign language semantics and ignoring some visual information. In a nutshell, the challenges locate redundancy removing and the alignment of the sign language data with the given tag. To solve the conundrum, this paper proposes a deep learning strategy called CRB-Net, which has used a kind of multimodal fusion attention mechanism. We first extract the features from RGB video and depth video, respectively, then conduct multi-modal fusion. Finally, the fused feature information is fed into an encoder-decoder network to achieve the goal of end-to-end continuous SLR. We verify the effectiveness of our method on three datasets, including the German dataset RWTH-Phoenix-Weather-2014, the Chinese dataset USTC-CSL and the Chinese dataset TJUT-SLRT. As shown by experimental results, the accuracy of 98.5% of our framework CRB-Net has outperformed the state-of-the-art works in the comparison, both in accuracy and algorithm execution efficiency. Feng Xiao 0005, Tiantian Yuan, Shengyong Chen |
SMC | 3 |
| 2020 | A dissimilarity measure for mixed nominal and ordinal attribute data in k-Modes algorithm
Fang Yuan 0003, Youlong Yang, Tiantian Yuan |
Appl. Intell. | 3 |
| 2019 | Large Scale Sign Language InterpretationabstractSign language is the primary way of communication between deaf people, but the majority of hearing people do not know how to sign. The reliance of deaf people on interpreters is both inconvenient and cost inefficient. Many research groups have experimented with using machine learning to develop automatic translators. Largely, these efforts have been constrained to restrictive dictionaries or insufficiently small signers or signed content. We introduce the world's largest sign language dataset to date- a collection of 50,000 video snippets taken from a pool of 10,000 unique utterances signed by 50 signers. We further propose several sequence-to-sequence deep learning approaches to automatically translate from Chinese sign language to both English and Mandarin written text. These methods utilize body joint position, facial expression, as well as finger articulation. While models can overfit on training sets, generalization to unforeseen utterances remains challenging with real-world data. The introduced dataset and methods demonstrate how modern machine learning methods are able to close the communication gap between deaf and hearing people. Tiantian Yuan, Shagan Sah, Tejaswini Ananthanarayana, Chi Zhang 0023, Aneesh Bhat, Sahaj Gandhi, Raymond W. Ptucha |
FG | 1 |
| 2009 | Expression Recognition Based on Multi-scale Block Local Gabor Binary Patterns with Dichotomy-Dependent Weights
Tiantian Yuan |
ISNN (2) | 3 |