Yidong Chen 0001

dblp:11/1492-1 · DBLP profile ↗
← Back
66ranked-venue papers
3as first author
42since 2021 · last 2026
0000-0002-0243-7228ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 2 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 14 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding and generation by integrating visual and textual data. However, these models frequently exhibit object hallucination problems: generating outputs that are inconsistent with the input image. Existing improved methods for mitigating hallucinations still suffer from two key limitations: dynamic approaches based on logits or attention mechanisms risk suppressing valuable linguistic priors, whereas static methods that employ fixed intervention vectors lack the flexibility to adapt to diverse images and questions. To address these issues, we propose RFI (Rectified Flow Intervention), a novel approach that harnesses the linear trajectory design of rectified flow for input-specific adaptation and employs gradient correction to ensure coherent generation, effectively combining the adaptability of dynamic methods with the stability of static ones. RFI dynamically predicts latent-space intervention vectors while requiring only a single forward pass in LVLMs per question, achieving computational efficiency (1.09x latency overhead for 100 new tokens). Extensive experiments show RFI significantly reduces hallucinations, achieving superior performance compared to existing advanced methods, highlighting its effectiveness as a lightweight plug-and-play method for reducing LVLM's hallucination in practical applications.
Junyu Cheng, Zhibiao Liang, Yidong Chen 0001, Shuangyin Li
AAAI3
2026 Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
abstract
Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines.
Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li 0014, Xinjie Chen, Yidong Chen 0001, Kai Fan 0002, Xiaodong Shi
AAAI6
2026 PLaST: Towards Paralinguistic-aware Speech Translation
abstract
Speech translation (ST) aims to translate speech from a source language into text in the target language. Naturally, speech signals contain paralinguistic cues beyond linguistic content, which could influence or even alter the interpretation of a lexically identical sentence, thereby yielding distinct translations. However, existing ST models lack direct and sufficient modeling of paralinguistic information, which limits their ability to perceive paralinguistic cues and understand speech comprehensively, leading to degraded translation performance. In response, we propose Paralinguistic-aware Speech Translation (PLaST), a novel dual-branch framework which directly leverages paralinguistic cues beyond the linguistic content. Specifically, PLaST employs a speech encoder and a style extractor to independently generate linguistic and paralinguistic representations, respectively. To obtain a purified linguistic representation aligned with the text representation, a hierarchical Optimal Transport (OT) is applied on the layer-wise outputs from an LLM decoder. Then, the paralinguistic information is retrieved and refined with an Attention-based Retrieval (AR) module, with the linguistic representation serving as queries to enable joint guidance for semantic understanding and translation generation. PLaST outperforms the strong baseline with an average of 5.0 directional and 4.5 global contrastive likelihood scores on the paralinguistic-sensitive benchmark ContraProST, demonstrating its superior capability in paralinguistic perception. Further experiments on the standard speech translation benchmark CoVoST-2 show that PLaST generalizes well to typical ST scenarios.
Ruiquan Zhang, Jinsong Su, Daimeng Wei, Min Zhang 0042, Yidong Chen 0001
AAAI7
2026 Selective Contrastive Learning For Gloss Free Sign Language Translation
abstract
Sign language translation (SLT) converts continuous sign videos into spoken-language text, yet it remains challenging due to the intrinsic modality mismatch between visual signs and written text, particularly in gloss-free settings.Recent SLT systems increasingly adopt CLIP-like Vision-Language pretraining (VLP) for cross-modal alignment, but the random inbatch contrast provides few, batch-dependent negatives and may mislabel semantically similar (or even identical) pairs as negatives, introducing noisy and potentially inconsistent alignment supervision.In this work, we first conduct a preliminary trajectory-based analysis that tracks negative video-text similarity over training.The results show that only a small subset of negatives exhibits the desired behavior of being consistently pushed away, while the remaining negatives display heterogeneous and often non-decreasing similarity dynamics, suggesting that random in-batch negatives are frequently uninformative for effective alignment.Inspired by this, we propose Selective Contrastive Learning for SLT (SCL-SLT) with a Pair Selection (PS) strategy.PS scores candidate negatives using similarity dynamics from reference checkpoints and constructs mini-batches via a curriculum that progressively emphasizes more challenging negatives, thereby strengthening contrastive supervision while reducing the influence of noisy or semantically invalid negatives.
Changhao Lai, Xuewen Zhong, Jinsong Su, Yidong Chen 0001
ACL (1)5
2026 Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
abstract
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators’ ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-though traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32–60%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.
Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen 0001, Xiaodong Shi
ACL (1)4
2026 CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language
abstract
Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs to understand sign language, especially in multimodal contexts, remains underexplored. To address this limitation, we introduce \textbf{CNSL-bench}, the first comprehensive \textbf{C}hinese \textbf{N}ational \textbf{S}ign \textbf{L}anguage \textbf{bench}mark designed for evaluating multimodal large language models (MLLMs) in sign language understanding. The proposed CNSL-bench is characterized by: 1) Authoritative grounding, as it is anchored to the officially standardized \textit{National Common Sign Language Dictionary}, mitigating ambiguity from regional or non-canonical variants and ensuring consistent semantic definitions; 2) Multimodal coverage, providing aligned textual descriptions, illustrative images, and sign language videos; and 3) Articulatory diversity, supporting fine-grained analysis across key manual articulatory forms, including air-writing, finger-spelling, and the Chinese manual-alphabet. Using CNSL-bench, we extensively evaluate 21 open-source and proprietary up-to-date MLLMs. Our results reveal that, despite recent advances in multimodal modeling, current MLLMs remain substantially inferior to human performance, exhibiting systematic disparities across input modalities and manual articulatory forms. Additional diagnostic analyses suggest that several performance limitations persist beyond improvements in reasoning and that instruction-following robustness varies substantially across models.
Xuewen Zhong, Xiaoyun Zheng, Jinsong Su, Yidong Chen 0001
ACL (1)5
2025 Improving Multilingual Sign Language Translation with Automatically Clustered Language Family Information
abstract
Sign Language Translation (SLT) bridges the communication gap between deaf and hearing individuals by converting sign language videos into spoken language texts. While most SLT research has focused on bilingual translation models, the recent surge in interest has led to the exploration of Multilingual Sign Language Translation (MSLT). However, MSLT presents unique challenges due to the diversity of sign languages across nations. This diversity can lead to cross-linguistic conflicts and hinder translation accuracy. To use the similarity of actions and semantics between sign languages to alleviate conflict, we propose a novel approach that leverages sign language families to improve MSLT performance. Sign languages were clustered into families automatically based on their Language distribution in the MSLT network. We compare the results of our proposed family clustering method with the analysis conducted by sign language linguists and then train dedicated translation models for each family in the many-to-one translation scenario. Our experiments on the SP-10 dataset demonstrate that our approach can achieve a balance between translation accuracy and computational cost by regulating the number of language families.
Ruiquan Zhang, Pei Yu, Yidong Chen 0001
COLING4
2025 Representation Purification for End-to-End Speech Translation
abstract
Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowledge transfer from machine translation, exploring various methods to bridge the gap between speech and text modalities. Despite substantial progress made, factors in speech that are not relevant to translation content, such as timbre and rhythm, often limit the efficiency of knowledge transfer. In this paper, we conceptualize speech representation as a combination of content-agnostic and content-relevant factors. We examine the impact of content-agnostic factors on translation performance through preliminary experiments and observe a significant performance deterioration when content-agnostic perturbations are introduced to speech signals. To address this issue, we propose a Speech Representation Purification with Supervision Enhancement (SRPSE) framework, which excludes the content-agnostic components within speech representations to mitigate their negative impact on ST. Experiments on MuST-C and CoVoST-2 datasets demonstrate that SRPSE significantly improves translation performance across all translation directions in three settings and achieves preeminent performance under a transcript-free setting.
Yue Zhou 0012, Yidong Chen 0001, Xiaodong Shi
COLING4
2025 FedSign: Federated Learning for Enhancing Sign Language Recognition with Adaptive Model Selection
abstract
Sign language recognition (SLR) serves as a critical technological bridge facilitating seamless communication between the hearing-impaired population and the general public. While deep learning-based SLR systems have demonstrated remarkable progress, they remain constrained by three fundamental challenges: (i) the necessity for large-scale annotated datasets, (ii) heightened privacy concerns surrounding biometric data, and (iii) inherent heterogeneity in sign language data distributions. To address these limitations, we present FedSign, a novel model-agnostic federated learning (FL) framework designed for universal applicability across diverse SLR architectures. Our framework simultaneously preserves data privacy through decentralized training while enhancing model robustness in non-independent and identically distributed (Non-IID) environments via an innovative entropy-based pseudo-label selection mechanism. This adaptive approach dynamically optimizes knowledge distillation by selectively leveraging either global or local model outputs based on uncertainty quantification. Comprehensive evaluations on three benchmark datasets (CSL-Daily, Phoenix-2014, and Phoenix-2014T) demonstrate that FedSign significantly outperforms both centralized and federated baselines, particularly in scenarios with extreme data heterogeneity. Notably, our framework establishes new state-of-the-art (SOTA) performance metrics on two challenging German sign language corpora.
Ruiquan Zhang, Yidong Chen 0001
ECAI4
2025 TempParaphraser: "Heating Up" Text to Evade AI-Text Detection through Paraphrasing
abstract
The widespread adoption of large language models (LLMs) has increased the need for reliable AI-text detection.While current detectors perform well on benchmark datasets, we highlight a critical vulnerability: increasing the temperature parameter during inference significantly reduces detection accuracy.Based on this weakness, we propose Temp-Paraphraser, a simple yet effective paraphrasing framework that simulates high-temperature sampling effects through multiple normaltemperature generations, effectively evading detection.Experiments show that TempParaphraser reduces detector accuracy by an average of 82.5% while preserving high text quality.We also demonstrate that training on TempParaphraser-augmented data improves detector robustness.
Ruiquan Zhang, Jinsong Su, Yidong Chen 0001
EMNLP4
2025 KFM-SLR: A Novel Privacy-Aware Framework for Sign Language Recognition Using Keypoint-Filled Face Masking
Yidong Chen 0001
ICIC (11)2
2025 Corruption-Agnostic Sign Language Translation
abstract
Sign Language Translation (SLT) aims to translate sign languages into spoken languages, acting as a bridge between the hard-of-hearing community and the hearing world. However, existing SLT research predominantly depends on meticulously curated datasets (e.g., PHOENIX-2014T), whereas practical SLT systems may encounter various forms of unanticipated corruptions that can compromise input quality, such as weather effects, camera blurring, external noise, or intrusion of extraneous objects. Consequently, the challenges faced in real-world applications are not fully considered, nor is the robustness required for SLT systems operating in diverse and unpredictable environments adequately evaluated. In order to make existing SLT models applicable in real-world and improve their robustness against various unknown corruptions, we introduce a Corruption-Agnostic Robust Sign Language Translation (CAR-SLT) framework. CAR-SLT consists of two primary components: Decoupled Information Bottleneck(DIB) and Adversarial Perturbation Alignment(APA). DIB decouples SLT-relevant from SLT-irrelevant information within the input, ensuring that the features used for translation contain only necessary elements for SLT. By focusing on SLT-relevant information that is less susceptible to corruptions, DIB enhances the robustness of SLT model. APA employs adversarial strategies to generate gradient-perturbed visual features that simulate scenarios where inputs are corrupted. Subsequently, APA uses two alignment strategies to maintain consistent model performance regardless of input condition, thus improving robustness against unknown corruptions. To evaluate the robustness of SLT systems and simulate the corruptions that a SLT system might encounter in real-world applications, we introduce a challenging setting called conrruption-agnostic sign language translation, along with the PHOENIX-2014T-C dataset. In this setting, the SLT model is trained on clean data but tested on corrupted data, where both the types of corruptions and their locations within the videos are unknown. The PHOENIX-2014T-C incorporates various types of corruptions encountered in real-world scenarios. We evaluate previous state-of-the-art methods as well as our method in this setting. The experimental results demonstrate that previous SLT methods do not perform well under corruption-agnostic setting, while our method exhibits superior performance.
Honghao Fu, Yidong Chen 0001
IJCNN2
2025 CRS3D: Consistency Regularization for Sparsely-supervised 3D object detection
abstract
3D object detection is an indispensable component of autonomous driving. Due to the expensive and labor-intensive annotation required for full supervision, sparsely supervised 3D object detection is emerging as a promising alternative. Although some existing sparse supervision methods have achieved encouraging detection results, they do not perform well with distant or occluded objects. To address this issue, we propose a Consistency Regularization method for Sparsely-supervised 3D object detection(CRS3D). CRS3D consists of three modules: the Point Cloud Adjust module and the Consistency Loss module, which enhance the model’s ability to perceive distant objects by aligning predictions between the original and perturbed point clouds; and the Priori Ratio module, which optimizes the perception of occluded objects by imposing constraints based on priori information. In the KITTI benchmark, CRS3D achieves 87.4% 3D-mAP for the car class at easy difficulty using only 2% of the annotations, surpassing the accuracy of fully supervised methods.
Binghui Zeng, Zongyue Wang, Zhaoliang Liu, Yidong Chen 0001, Weiquan Liu
IJCNN4
2025 Hy2CRE: Hypernetworks with Hybrid Data Augmentation for Continual Relation Extraction
abstract
Continual Relation Extraction (CRE) aims to learn constantly emerging tasks while avoiding forgetting the learned tasks. Previous studies suggest that interference among similar relations is the primary factor contributing to the forgetting of learned tasks. To address this issue, robust training strategies have been employed to enhance the model’s ability for distinguishing between similar relations. However, these methods use the same model parameters for learning all tasks, which increases the risk of conflicts between similar relations within the unified feature space. To address this issue, we propose a model that utilizes a Hypernetwork with Hybrid data augmentation for CRE (Hy2CRE). Specifically, Hy2CRE employs a hypernetwork-based network generator to generate task-specific projection heads for each task during the continual learning process, thereby mitigating the risk of conflicts emerging between similar relations within the model’s feature space. Meanwhile, we introduce a hybrid data augmentation method to further enhance the model’s robustness to similar relations, which integrates data augmentation in both the discrete text space and the continuous semantic space. Experimental results on two benchmark datasets prove the effectiveness of our model.
Yang Zhang 0079, Yidong Chen 0001
IJCNN2
2025 Advancing Continuous Sign Language Recognition Through Denoising Diffusion Transformer-Based Spatial-Temporal Enhancement
abstract
ABSTRACT The intricate spatial‐temporal dynamics and variability of sign language gestures pose significant challenges for Continuous Sign Language Recognition (CSLR) systems. Existing models often fall short in accurately capturing these complexities, leading to performance issues and frequent misalignments. To address these shortcomings, we introduce a new approach that leverages Denoising Diffusion Models (DDMs) to improve feature representation in the visual‐sequential module of CSLR systems. Originally intended for generative tasks, DDMs have shown strong potential in representation learning through a denoising process akin to Denoising Autoencoders. Our method incorporates a denoising diffusion transformer into the CSLR framework to refine spatial‐temporal features, capitalizing on the ability of diffusion models to enhance representation quality. By conditionally denoising visual feature sequences, our approach increases the discriminative capability of the system. Additionally, we introduce an additional classifier, trained with Connectionist Temporal Classification (CTC) loss, to provide complementary supervision and further boost performance. Extensive experiments demonstrate that our method significantly improves CSLR accuracy by effectively capturing the subtle details of continuous sign language gestures and overcoming the representation limitations of current models.
Suhail Muhammad Kamal, Yidong Chen 0001, Shaozi Li
Concurr. Comput. Pract. Exp.2
2025 Towards Simultaneous Sign Language Production: A Future-Context-Aware Approach
abstract
Sign Language Production (SLP) has achieved promising progress in offline settings, where full input text is available before generation. However, such methods are unsuitable for real-time applications requiring low latency. In this work, we introduce Simultaneous Sign Language Production (SimulSLP), a new task that generates sign pose sequences incrementally from streaming text input. We first formalize the SimulSLP task and adapt the Average Token Delay metric to quantify latency. Then, we benchmark this task using three strong baselines from offline SLP—an end-to-end system and two cascaded pipelines with neural and dictionary-based Gloss-to-Pose modules—under a wait-k policy. However, all baselines suffer from a mismatch between full-sequence training and partial-input inference. To mitigate this, we propose a Future-Context-Aware Inference (FCAI) strategy. FCAI enhances partial input representations by predicting a small number of future tokens using a large language model. Before decoding, speculative features from the predicted tokens are discarded to ensure alignment with the observed input. Experiments on PHOENIX2014T show that FCAI significantly improves the quality-latency trade-off, especially in low-latency settings, offering a promising step toward SimulSLP.
Biao Fu, Xiaodong Shi, Yidong Chen 0001
IEEE Signal Process. Lett.4
2025 Data augmentation and debiasing for signers in signer-independent sign language translation
Honghao Fu, Yidong Chen 0001
J. Supercomput.2
2024 Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal Alignment
abstract
Sign language translation (SLT) aims to convert continuous sign language videos into textual sentences. As a typical multi-modal task, there exists an inherent modality gap between sign language videos and spoken language text, which makes the cross-modal alignment between visual and textual modalities crucial. However, previous studies tend to rely on an intermediate sign gloss representation to help alleviate the cross-modal problem thereby neglecting the alignment across modalities that may lead to compromised results. To address this issue, we propose a novel framework based on Conditional Variational autoencoder for SLT (CV-SLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text. Specifically, our CV-SLT consists of two paths with two Kullback-Leibler (KL) divergences to regularize the outputs of the encoder and decoder, respectively. In the prior path, the model solely relies on visual information to predict the target text; whereas in the posterior path, it simultaneously encodes visual information and textual knowledge to reconstruct the target text. The first KL divergence optimizes the conditional variational autoencoder and regularizes the encoder outputs, while the second KL divergence performs a self-distillation from the posterior path to the prior path, ensuring the consistency of decoder outputs.We further enhance the integration of textual information to the posterior path by employing a shared Attention Residual Gaussian Distribution (ARGD), which considers the textual information in the posterior path as a residual component relative to the prior path. Extensive experiments conducted on public datasets demonstrate the effectiveness of our framework, achieving new state-of-the-art results while significantly alleviating the cross-modal representation discrepancy. The code and models are available at https://github.com/rzhao-zhsq/CV-SLT.
Biao Fu, Jinsong Su, Yidong Chen 0001
AAAI6
2024 Layer-Wise Representation Fusion for Compositional Generalization
abstract
Existing neural models are demonstrated to struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components. A key reason for failure on CG is that the syntactic and semantic representations of sequences in both the uppermost layer of the encoder and decoder are entangled. However, previous work concentrates on separating the learning of syntax and semantics instead of exploring the reasons behind the representation entanglement (RE) problem to solve it. We explain why it exists by analyzing the representation evolving mechanism from the bottom to the top of the Transformer layers. We find that the ``shallow'' residual connections within each layer fail to fuse previous layers' information effectively, leading to information forgetting between layers and further the RE problems. Inspired by this, we propose LRF, a novel Layer-wise Representation Fusion framework for CG, which learns to fuse previous layers' information back into the encoding and decoding process effectively through introducing a fuse-attention module at each encoder and decoder layer. LRF achieves promising results on two realistic benchmarks, empirically demonstrating the effectiveness of our proposal. Codes are available at https://github.com/thinkaboutzero/LRF.
Yafang Zheng, Shuangtao Li, Zhaohong Lai, Biao Fu, Yidong Chen 0001, Xiaodong Shi
AAAI8
2024 Adaptive Simultaneous Sign Language Translation with Confident Translation Length Estimation
abstract
Traditional non-simultaneous Sign Language Translation (SLT) methods, while effective for pre-recorded videos, face challenges in real-time scenarios due to inherent inference delays. The emerging field of simultaneous SLT aims to address this issue by progressively translating incrementally received sign video. However, the sole existing work in simultaneous SLT adopts a fixed gloss-based policy, which suffer from limitations in boundary prediction and contextual comprehension. In this paper, we delve deeper into this area and propose an adaptive policy for simultaneous SLT. Our approach introduces the concept of “confident translation length”, denoting maximum accurate translation achievable from current input. An estimator measures this length for streaming sign video, enabling the model to make informed decisions on whether to wait for more input or proceed with translation. To train the estimator, we construct a training data of confident translation length based on the longest common prefix between translations of partial and complete inputs. Furthermore, we incorporate adaptive training, utilizing pseudo prefix pairs, to refine the offline translation model for optimal performance in simultaneous scenarios. Experimental results on PHOENIX2014T and CSL-Daily demonstrate the superiority of our adaptive policy over existing methods, particularly excelling in situations requiring extremely low latency.
Biao Fu, Ruiquan Zhang, Xiaodong Shi, Jinsong Su, Yidong Chen 0001
LREC/COLING8
2024 Improving Non-Autoregressive Sign Language Translation with Random Ordering Progressive Prediction Pretraining
abstract
Recently, the Non-AutoRegressive (NAR) decoding mechanism, effectively reducing the inference latency of text generation, has been applied to Sign Language Translation (SLT). Typically, the current best NAR SLT model using a Curriculum-based Non-autoregressive Decoder (CND) outperforms AutoRegressive (AR) baselines in speed and performance. Although it has been proven that AutoRegressive Pre-trained Language Models (AR-PLMs) further boost the performance of AR SLT models, combining NAR Pretrained Language Models (NAR-PLMs) with NAR SLT model remains challenge due to (1) existing NAR-PLMs’ inability to model token dependencies between decoder layers, crucial for NAR SLT models using CND; (2) the modality gap between the decoder’s inputs of the NAR-PLMs and NAR SLT models. To address these, we propose a Random Ordering Progressive Prediction Pre-training task for NAR SLT models using CND, enabling the decoder to predict target sequences in diverse orderings and enhancing the modeling of target token dependencies between layers. Moreover, we propose a CTC-enhanced Soft Copy method to incorporate target-side information in the decoder’s inputs, alleviating the modality gap. Experimental results on PHOENIX-2014T and CSL-Daily demonstrate that our model consistently outperforms all strong baselines and achieves competitive performance with AR SLT models equipped with AR-PLMs.
Pei Yu, Changhao Lai, Biao Fu, Yidong Chen 0001
ECAI7
2024 An Explicit Multi-Modal Fusion Method for Sign Language Translation
abstract
Sign Language Translation (SLT) aims to convert sign language videos into corresponding spoken text sequences. However, the inherent modality gap between sign language video and text hinders the development of SLT. Motivated by the linguistic consistency between gloss1and text, we propose EMF-SLT, an Explicit Multi-modal Fusion method for Sign Language Translation to mitigate the modality gap with the help of gloss. Specifically, EMF-SLT first leverages a vector quantizer and a fusion module to align and fuse sign language and gloss features, respectively, resulting in more informative multi-modal features for the decoder. Then, a multi-task mutual learning framework is introduced to regularize the output predictions from different modalities, which ensures the consistency of outputs across modalities and encourages different modalities to learn from each other. Experiments on two SLT benchmarks and further analyses show that our method achieves significant improvements over the baselines and effectively alleviates the modality gap.
Biao Fu, Pei Yu, Xiaodong Shi, Yidong Chen 0001
ICASSP6
2024 Adversarial autoencoder for continuous sign language recognition
abstract
Summary Sign language serves as a vital communication medium for the deaf community, encompassing a diverse array of signs conveyed through distinct hand shapes along with non‐manual gestures like facial expressions and body movements. Accurate recognition of sign language is crucial for bridging the communication gap between deaf and hearing individuals, yet the scarcity of large‐scale datasets poses a significant challenge in developing robust recognition technologies. Existing works address this challenge by employing various strategies, such as enhancing visual modules, incorporating pretrained visual models, and leveraging multiple modalities to improve performance and mitigate overfitting. However, the exploration of the contextual module, responsible for modeling long‐term dependencies, remains limited. This work introduces an Adversarial Autoencoder for Continuous Sign Language Recognition, AA‐CSLR, to address the constraints imposed by limited data availability, leveraging the capabilities of generative models. The integration of pretrained knowledge, coupled with cross‐modal alignment, enhances the representation of sign language by effectively aligning visual and textual features. Through extensive experiments on publicly available datasets (PHOENIX‐2014, PHOENIX‐2014T, and CSL‐Daily), we demonstrate the effectiveness of our proposed method in achieving competitive performance in continuous sign language recognition.
Suhail Muhammad Kamal, Yidong Chen 0001, Shaozi Li
Concurr. Comput. Pract. Exp.2
2024 Improving few-shot relation extraction through semantics-guided learning
Hui Wu 0008, Yidong Chen 0001, Xiaodong Shi
Neural Networks3
2024 Self-Growing Binary Activation Network: A Novel Deep Learning Model With Dynamic Architecture
abstract
For a deep learning model, the network architecture is crucial as a model with inappropriate architecture often suffers from performance degradation or parameter redundancy. However, it is experiential and difficult to find the appropriate architecture for a certain application. To tackle this problem, we propose a novel deep learning model with dynamic architecture, named self-growing binary activation network (SGBAN), which can extend the design of a fully connected network (FCN) progressively, resulting in a more compact architecture with higher performance on a certain task. This constructing process is more efficient than neural architecture search methods that train mass of networks to search for the optimal one. Concretely, the training technique of SGBAN is based on the function-preserving transformations that can expand the architecture and combine the information in the new data without neglecting the knowledge learned in the previous steps. The experimental results on four different classification tasks, i.e., Iris, MNIST, CIFAR-10, and CIFAR-100, demonstrate the effectiveness of SGBAN. On the one hand, SGBAN achieves competitive accuracy when compared with the FCN composed of the same architecture, which indicates that the new training technique has the equivalent optimization ability as the traditional optimization methods. On the other hand, the architecture generated by SGBAN achieves 0.59% improvements of accuracy, with only 33.44% parameters when compared with the FCNs composed of manual design architectures, i.e., 500 + 150 hidden units, on MNIST. Furthermore, we demonstrate that replacing the fully connected layers of the well-trained VGG-19 with SGBAN can gain a slightly improved performance with less than 1% parameters on all these tasks. Finally, we show that the proposed method can conduct the incremental learning tasks and outperform the three outstanding incremental learning methods, i.e., learning without forgetting, elastic weight consolidation, and gradient episodic memory, on both the incremental learning tasks on Disjoint MNIST and Disjoint CIFAR-10.
Yidong Chen 0001, Changle Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2023 Exploring Self-Distillation Based Relational Reasoning Training for Document-Level Relation Extraction
abstract
Document-level relation extraction (RE) aims to extract relational triples from a document. One of its primary challenges is to predict implicit relations between entities, which are not explicitly expressed in the document but can usually be extracted through relational reasoning. Previous methods mainly implicitly model relational reasoning through the interaction among entities or entity pairs. However, they suffer from two deficiencies: 1) they often consider only one reasoning pattern, of which coverage on relational triples is limited; 2) they do not explicitly model the process of relational reasoning. In this paper, to deal with the first problem, we propose a document-level RE model with a reasoning module that contains a core unit, the reasoning multi-head self-attention unit. This unit is a variant of the conventional multi-head self-attention and utilizes four attention heads to model four common reasoning patterns, respectively, which can cover more relational triples than previous methods. Then, to address the second issue, we propose a self-distillation training framework, which contains two branches sharing parameters. In the first branch, we first randomly mask some entity pair feature vectors in the document, and then train our reasoning module to infer their relations by exploiting the feature information of other related entity pairs. By doing so, we can explicitly model the process of relational reasoning. However, because the additional masking operation is not used during testing, it causes an input gap between training and testing scenarios, which would hurt the model performance. To reduce this gap, we perform conventional supervised training without masking operation in the second branch and utilize Kullback-Leibler divergence loss to minimize the difference between the predictions of the two branches. Finally, we conduct comprehensive experiments on three benchmark datasets, of which experimental results demonstrate that our model consistently outperforms all competitive baselines. Our source code is available at https://github.com/DeepLearnXMU/DocRE-SD
Jinsong Su, Zijun Min, Zhongjian Miao, Qingguo Hu, Biao Fu, Xiaodong Shi, Yidong Chen 0001
AAAI8
2023 CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational Alignment
abstract
Sign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modules and develop two mainstream solutions. The multi-stream architectures extend multi-cue visual features, yielding the current SOTA performances but requiring complex designs and might introduce potential noise. Alternatively, the advanced single-cue SLR frameworks using explicit cross-modal alignment between visual and textual modalities are simple and effective, potentially competitive with the multi-cue framework. In this work, we propose a novel contrastive visual-textual transformation for SLR, CVT-SLR, to fully explore the pretrained knowledge of both the visual and language modalities. Based on the single-cue cross-modal alignment framework, we propose a variational autoencoder (VAE) for pretrained contextual knowledge while introducing the complete pretrained language module. The VAE implicitly aligns visual and textual modalities while benefiting from pretrained contextual knowledge as the traditional contextual module. Meanwhile, a contrastive cross-modal alignment algorithm is designed to explicitly enhance the consistency constraints. Extensive experiments on public datasets (PHOENIX-2014 and PHOENIX-2014T) demonstrate that our proposed CVT-SLR consistently outperforms existing single-cue methods and even outperforms SOTA multi-cue methods. The source codes and models are available at https://github.com/binbinjiang/CVT-SLR.
Jiangbin Zheng 0002, Yile Wang 0001, Cheng Tan 0012, Siyuan Li 0002, Jun Xia 0001, Yidong Chen 0001, Stan Z. Li
CVPR7
2023 Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and Inference
abstract
A popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different latency constraints.However, there is a mismatch problem in using a model trained with complete utterances for streaming inference with partial input.We demonstrate that speech representations extracted at the end of a streaming input are significantly different from those extracted from a complete utterance.To address this issue, we propose a new approach called Future-Aware Streaming Translation (FAST) that adapts an offline ST model for streaming input.FAST includes a Future-Aware Inference (FAI) strategy that incorporates future context through a trainable masked embedding, and a Future-Aware Distillation (FAD) framework that transfers future context from an approximation of full speech to streaming input.Our experiments on the MuST-C EnDe, EnEs, and EnFr benchmarks show that FAST achieves better trade-offs between translation quality and latency than strong baselines.Extensive analyses suggest that our methods effectively alleviate the aforementioned mismatch problem between offline training and online inference.1
Biao Fu, Minpeng Liao, Kai Fan 0002, Zhongqiang Huang, Boxing Chen, Yidong Chen 0001, Xiaodong Shi
EMNLP6
2023 Exploring All-In-One Knowledge Distillation Framework for Neural Machine Translation
abstract
Conventional knowledge distillation (KD) approaches are commonly employed to compress neural machine translation (NMT) models.However, they only obtain one lightweight student each time.Consequently, we have to conduct KD multiple times when different students are required at the same time, which could be resource-intensive.Additionally, these students are individually optimized, and thus lack interactions with each other, leading to their potential not being fully exerted.In this work, we propose a novel All-In-One Knowledge Distillation (AIO-KD) framework for NMT, which generates multiple satisfactory students at once.Under AIO-KD, we first randomly extract fewer-layer subnetworks from the teacher as the sample students.Then, we jointly optimize the teacher and these students, where the students simultaneously learn the knowledge from the teacher and interact with other students via mutual learning.When utilized, we re-extract the candidate students, satisfying the specifications of various devices.Particularly, we adopt carefully-designed strategies for AIO-KD: 1) we dynamically detach gradients to prevent poorly-performed students from negatively affecting the teacher during the knowledge transfer, which could subsequently impact other students; 2) we design a twostage mutual learning strategy, which alleviates the negative impacts of poorly-performed students on the early-stage student interactions.Extensive experiments and in-depth analyses on three benchmarks demonstrate the effectiveness and eco-friendliness of AIO-KD.Our source code is available at https://github. com/DeepLearnXMU/AIO-KD.
Zhongjian Miao, Wen Zhang 0015, Jinsong Su, Xiang Li 0104, Jian Luan 0001, Yidong Chen 0001, Bin Wang 0004, Min Zhang 0005
EMNLP6
2023 HyperNetwork-based Decoupling to Improve Model Generalization for Few-Shot Relation Extraction
abstract
Few-shot relation extraction (FSRE) aims to train a model that can deal with new relations using only a few labeled examples.Most existing studies employ Prototypical Networks for FSRE, which usually overfits the relation classes in the training set and cannot generalize well to unseen relations.By investigating the class separation of an FSRE model, we find that model upper layers are prone to learn relation-specific knowledge.Therefore, in this paper, we propose a HyperNetworkbased Decoupling approach to improve the generalization of FSRE models.Specifically, our model consists of an encoder, a network generator (for producing relation classifiers) and the generated-then-finetuned classifiers for every N -way-K-shot episode.Meanwhile, we design a two-step training strategy along with a class-agnostic aligner, by which the generated classifiers focus on acquiring relation-specific knowledge and the encoder is encouraged to learn more general relation knowledge.In this way, the roles of upper and lower layers in our FSRE model are explicitly decoupled, thus enhancing its generalizing capability during testing.Experiments on two public datasets demonstrate the effectiveness of our method.Our source code is available at https: //github.com/DeepLearnXMU/FSRE-HDN.
Chulun Zhou, Fandong Meng, Jinsong Su, Yidong Chen 0001, Jie Zhou 0016
EMNLP5
2023 A Token-Level Contrastive Framework for Sign Language Translation
abstract
Sign Language Translation (SLT) is a promising technology to bridge the communication gap between the deaf and the hearing people. Recently, researchers have adopted Neural Machine Translation (NMT) methods, which usually require large-scale corpus for training, to achieve SLT. However, the publicly available SLT corpus is very limited, which causes the collapse of the token representations and the inaccuracy of the generated tokens. To alleviate this issue, we propose Con-SLT, a novel token-level Contrastive learning framework for Sign Language Translation , which learns effective token representations by incorporating token-level contrastive learning into the SLT decoding process. Concretely, ConSLT treats each token and its counterpart generated by different dropout masks as positive pairs during decoding, and then randomly samples K tokens in the vocabulary that are not in the current sentence to construct negative examples. We conduct comprehensive experiments on two benchmarks (PHOENIX14T and CSL-Daily) for both end-to-end and cascaded settings. The experimental results demonstrate that ConSLT can achieve better translation quality than the strong baselines1.
Biao Fu, Peigen Ye, Pei Yu, Xiaodong Shi, Yidong Chen 0001
ICASSP7
2023 Efficient Sign Language Translation with a Curriculum-based Non-autoregressive Decoder
abstract
Most existing studies on Sign Language Translation (SLT) employ AutoRegressive Decoding Mechanism (AR-DM) to generate target sentences. However, the main disadvantage of the AR-DM is high inference latency. To address this problem, we introduce Non-AutoRegressive Decoding Mechanism (NAR-DM) into SLT, which generates the whole sentence at once. Meanwhile, to improve its decoding ability, we integrate the advantages of curriculum learning and NAR-DM and propose a Curriculum-based NAR Decoder (CND). Specifically, the lower layers of the CND are expected to predict simple tokens that could be predicted correctly using source-side information solely. Meanwhile, the upper layers could predict complex tokens based on the lower layers' predictions. Therefore, our CND significantly reduces the model's inference latency while maintaining its competitive performance. Moreover, to further boost the performance of our CND, we propose a mutual learning framework, containing two decoders, i.e., an AR decoder and our CND. We jointly train the two decoders and minimize the KL divergence between their outputs, which enables our CND to learn the forward sequential knowledge from the strengthened AR decoder. Experimental results on PHOENIX2014T and CSL-Daily demonstrate that our model consistently outperforms all competitive baselines and achieves 7.92/8.02× speed-up compared to the AR SLT model respectively. Our source code is available at https://github.com/yp20000921/CND.
Pei Yu, Biao Fu, Yidong Chen 0001
IJCAI4
2023 Exploring Effective Inter-Encoder Semantic Interaction for Document-Level Relation Extraction
abstract
In document-level relation extraction (RE), the models are required to correctly predict implicit relations in documents via relational reasoning. To this end, many graph-based methods have been proposed for this task. Despite their success, these methods still suffer from several drawbacks: 1) their interaction between document encoder and graph encoder is usually unidirectional and insufficient; 2) their graph encoders often fail to capture the global context of nodes in document graph. In this paper, we propose a document-level RE model with a Graph-Transformer Network (GTN). The GTN includes two core sublayers: 1) the graph-attention sublayer that simultaneously models global and local contexts of nodes in the document graph; 2) the cross-attention sublayer, enabling GTN to capture the non-entity clue information from the document encoder. Furthermore, we introduce two auxiliary training tasks to enhance the bidirectional semantic interaction between the document encoder and GTN: 1) the graph node reconstruction that can effectively train our cross-attention sublayer to enhance the semantic transition from the document encoder to GTN; 2) the structure-aware adversarial knowledge distillation, by which we can effectively transfer the structural information of GTN to the document encoder. Experimental results on four benchmark datasets prove the effectiveness of our model. Our source code is available at https://github.com/DeepLearnXMU/DocRE-BSI.
Zijun Min, Jinsong Su, Pei Yu, Ante Wang, Yidong Chen 0001
IJCAI6
2023 A Novel POS-Guided Data Augmentation Method for Sign Language Gloss Translation
Yafang Zheng, Yidong Chen 0001, Xiaodong Shi
NLPCC (2)4
2023 MOPRD: A multidisciplinary open peer review dataset
Jialiang Lin 0001, Zhangping Zhou, Yidong Chen 0001, Xiaodong Shi
Neural Comput. Appl.4
2022 Towards Robust Neural Machine Translation with Iterative Scheduled Data-Switch Training
abstract
Most existing methods on robust neural machine translation (NMT) construct adversarial examples by injecting noise into authentic examples and indiscriminately exploit two types of examples. They require the model to translate both the authentic source sentence and its adversarial counterpart into the identical target sentence within the same training stage, which may be a suboptimal choice to achieve robust NMT. In this paper, we first conduct a preliminary study to confirm this claim and further propose an Iterative Scheduled Data-switch Training Framework to mitigate this problem. Specifically, we introduce two training stages, iteratively switching between authentic and adversarial examples. Compared with previous studies, our model focuses more on just one type of examples at each single stage, which can better exploit authentic and adversarial examples, and thus obtaining a better robust NMT model. Moreover, we introduce an improved curriculum learning method with a sampling strategy to better schedule the process of noise injection. Experimental results show that our model significantly surpasses several competitive baselines on four translation benchmarks. Our source code is available at https://github.com/DeepLearnXMU/RobustNMT-ISDST.
Zhongjian Miao, Xiang Li 0104, Liyan Kang, Wen Zhang 0015, Chulun Zhou, Yidong Chen 0001, Bin Wang 0004, Min Zhang 0005, Jinsong Su
COLING6
2022 Towards Better Document-level Relation Extraction via Iterative Inference
abstract
Document-level relation extraction (RE) aimsto extract the relations between entities from the input document that usually containing many difficultly-predicted entity pairs whose relations can only be predicted through relational inference.Existing methods usually directly predict the relations of all entity pairs of input document in a one-pass manner, ignoring the fact that predictions of some entity pairs heavily depend on the predicted results of other pairs.To deal with this issue, in this paper, we propose a novel document-level RE model with iterative inference.Our model is mainly composed of two modules: 1) a base module expected to provide preliminary relation predictions on entity pairs; 2) an inference module introduced to refine these preliminary predictions by iteratively dealing with difficultlypredicted entity pairs depending on other pairs in an easy-to-hard manner.Unlike previous methods which only consider feature information of entity pairs, our inference module is equipped with two Extended Cross Attention units, allowing it to exploit both feature information and previous predictions of entity pairs during relational inference.Furthermore, we adopt a two-stage strategy to train our model.At the first stage, we only train our base module.During the second stage, we train the whole model, where contrastive learning is introduced to enhance the training of inference module.Experimental results on three commonly-used datasets show that our model consistently outperforms other competitive baselines.Our source code is available at https://github. com/DeepLearnXMU/DocRE-II.
Jinsong Su, Yidong Chen 0001, Zhongjian Miao, Zijun Min, Qingguo Hu, Xiaodong Shi
EMNLP3
2022 Automatic Analysis of Available Source Code of Top Artificial Intelligence Conference Papers
abstract
Source code is essential for researchers to reproduce the methods and replicate the results of artificial intelligence (AI) papers. Some organizations and researchers manually collect AI papers with available source code to contribute to the AI community. However, manual collection is a labor-intensive and time-consuming task. To address this issue, we propose a method to automatically identify papers with available source code and extract their source code repository URLs. With this method, we find that 20.5% of regular papers of 10 top AI conferences published from 2010 to 2019 are identified as papers with available source code and that 8.1% of these source code repositories are no longer accessible. We also create the XMU NLP Lab README Dataset, the largest dataset of labeled README files for source code document research. Through this dataset, we have discovered that quite a few README files have no installation instructions or usage tutorials provided. Further, a large-scale comprehensive statistical analysis is made for a general picture of the source code of AI conference papers. The proposed solution can also go beyond AI conference papers to analyze other scientific papers from both journals and conferences to shed light on more domains.
Jialiang Lin 0001, Yingmin Wang, Yao Yu 0001, Yu Zhou 0007, Yidong Chen 0001, Xiaodong Shi
Int. J. Softw. Eng. Knowl. Eng.5
2021 CTRD: A Chinese Theme-Rheme Discourse Dataset
Biao Fu, Yiqi Tong, Dawei Tian, Yidong Chen 0001, Xiaodong Shi
NLPCC (1)4
2021 Aggregating Inter-viewpoint Relationships of User's Review for Accurate Recommendation
Xingchen He, Yidong Chen 0001, Guocheng Zhang, Xu-Ling Zheng
NLPCC (1)2
2021 Enhancing Neural Sign Language Translation by highlighting the facial expression information
Jiangbin Zheng 0002, Yidong Chen 0001, Chong Wu 0007, Xiaodong Shi, Suhail Muhammad Kamal
Neurocomputing2
2021 Knowledge Graph Embedding Based on Multi-View Clustering Framework
abstract
Knowledge representation is one of the critical problems in knowledge engineering and artificial intelligence, while knowledge embedding as a knowledge representation methodology indicates entities and relations in knowledge graph as low-dimensional, continuous vectors. In this way, knowledge graph is compatible with numerical machine learning models. Major knowledge embedding methods employ geometric translation to design score function, which is weak-semantic for natural language processing. To overcome this disadvantage, in this paper, we propose our model based on multi-view clustering framework, which could generate semantic representations of knowledge elements (i.e., entities/relations). With our semantic model, we also present an empowered solution to entity retrieval with entity description. Extensive experiments show that our model achieves substantial improvements against baselines on the task of knowledge graph completion, triple classification, entity classification, and entity retrieval.
Yidong Chen 0001, Xiaodong Shi
IEEE Trans. Knowl. Data Eng.2
2020 A Document-Level Neural Machine Translation Model with Dynamic Caching Guided by Theme-Rheme Information
abstract
Research on document-level Neural Machine Translation (NMT) models has attracted increasing attention in recent years.Although the proposed works have proved that the inter-sentence information is helpful for improving the performance of the NMT models, what information should be regarded as context remains ambiguous.To solve this problem, we proposed a novel cache-based document-level NMT model which conducts dynamic caching guided by theme-rheme information.The experiments on NIST evaluation sets demonstrate that our proposed model achieves substantial improvements over the state-of-the-art baseline NMT models.As far as we know, we are the first to introduce theme-rheme theory into the field of machine translation.
Yiqi Tong, Jiangbin Zheng 0002, Hongkang Zhu, Yidong Chen 0001, Xiaodong Shi
COLING4
2019 Boosting implicit discourse relation recognition with connective-based word embeddings
Changxing Wu, Jinsong Su, Yidong Chen 0001, Xiaodong Shi
Neurocomputing3
2019 Multi-perspective neural architecture for recommendation system
Yidong Chen 0001, Xiaodong Shi
Neural Networks2
2019 POS Tag-enhanced Coarse-to-fine Attention for Neural Machine Translation
abstract
Although neural machine translation (NMT) has certain capability to implicitly learn semantic information of sentences, we explore and show that Part-of-Speech (POS) tags can be explicitly incorporated into the attention mechanism of NMT effectively to yield further improvements. In this article, we propose an NMT model with tag-enhanced attention mechanism. In our model, NMT and POS tagging are jointly modeled via multi-task learning. Besides following common practice to enrich encoder annotations by introducing predicted source POS tags, we exploit predicted target POS tags to refine attention model in a coarse-to-fine manner. Specifically, we first implement a coarse attention operation solely on source annotations and target hidden state, where the produced context vector is applied to update target hidden state used for target POS tagging. Then, we perform a fine attention operation that extends the coarse one by further exploiting the predicted target POS tags. Finally, we facilitate word prediction by simultaneously utilizing the context vector from fine attention and the predicted target POS tags. Experimental results and further analyses on Chinese-English and Japanese-English translation tasks demonstrate the superiority of our proposed model over the conventional NMT models. We release our code at https://github.com/middlekisser/PEA-NMT.git.
Yongjing Yin, Jinsong Su, Huating Wen, Jiali Zeng, Yang Liu 0005, Yidong Chen 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2018 Deep Semantic Role Labeling With Self-Attention
abstract
Semantic Role Labeling (SRL) is believed to be a crucial step towards natural language understanding and has been widely studied. Recent years, end-to-end SRL with recurrent neural networks (RNN) has gained increasing attention. However, it remains a major challenge for RNNs to handle structural information and long range dependencies. In this paper, we present a simple and effective architecture for SRL which aims to address these problems. Our model is based on self-attention which can directly capture the relationships between two tokens regardless of their distance. Our single model achieves F1=83.4 on the CoNLL-2005 shared task dataset and F1=82.7 on the CoNLL-2012 shared task dataset, which outperforms the previous state-of-the-art results by 1.8 and 1.0 F1 score respectively. Besides, our model is computationally efficient, and the parsing speed is 50K tokens per second on a single Titan X GPU.
Zhixing Tan, Mingxuan Wang, Yidong Chen 0001, Xiaodong Shi
AAAI4
2018 Lattice-to-sequence attentional Neural Machine Translation models
Zhixing Tan, Jinsong Su, Boli Wang, Yidong Chen 0001, Xiaodong Shi
Neurocomputing4
2018 Exploring Implicit Semantic Constraints for Bilingual Word Embeddings
Jinsong Su, Zhenqiao Song, Yaojie Lu 0001, Mu Xu, Changxing Wu, Yidong Chen 0001
Neural Process. Lett.6
2018 Constructing and validating word similarity datasets by integrating methods from psychology, brain science and computational linguistics
Yu Wan 0004, Yidong Chen 0001, Xiaodong Shi, Changle Zhou
Soft Comput.2
2017 A synergetic semantic role labeling model with the introduction of fluctuating force accompanied with word sense information
abstract
Semantic role labeling (SRL) is a key problem in natural language processing which goal is to find a sentence-level semantic representation. Word sense information plays an important role on the determination of semantic roles. The introduction of word sense in the process of semantic role labeling will hopefully lead to achieve better result. But how to better reflect the relationship between word sense information and semantic role information is a key task. Synergetic neural network (SNN) provides an opportunity for us to study how to use word sense for semantic role labeling. The role labeling process can be seen as a competition process of many roles chain order parameters with word sense, of which order parameter with the largest support will win, thereby obtaining desired pattern. There are three main contributions in this article: firstly, we introduce synergetic theory to semantic analysis and propose a semantic analysis method based on synergetic neural network, which can effectively use semantic information and word sense information. Secondly, fluctuating force is introduced into potential evolution function which can effectively make use of prior semantic knowledge. Finally, we use artificial fish swarm algorithm (AFSA) to realize the optimization of network parameter which has both global and local search ability, and not easy to fall into local extremism. Experiment results show the proposed model in this paper can further improve the performance of semantic role labeling, and thus provides an important reference value to future research.
Zhehuang Huang, Yidong Chen 0001, Xiaodong Shi
Intell. Data Anal.2
2017 Leveraging bilingually-constrained synthetic data via multi-task neural networks for implicit discourse relation recognition
Changxing Wu, Xiaodong Shi, Yidong Chen 0001, Yanzhou Huang, Jinsong Su
Neurocomputing3
2017 Co-training for Implicit Discourse Relation Recognition Based on Manual and Distributed Features
Changxing Wu, Xiaodong Shi, Jinsong Su, Yidong Chen 0001, Yanzhou Huang
Neural Process. Lett.4
2016 Bilingually-constrained Synthetic Data for Implicit Discourse Relation Recognition
abstract
To alleviate the shortage of labeled data, we propose to use bilingually-constrained synthetic implicit data for implicit discourse relation recognition.These data are extracted from a bilingual sentence-aligned corpus according to the implicit/explicit mismatch between different languages.Incorporating these data via a multi-task neural network model achieves significant improvements over baselines, on both the English PDTB and Chinese CDTB data sets.
Changxing Wu, Xiaodong Shi, Yidong Chen 0001, Yanzhou Huang, Jinsong Su
EMNLP3
2016 Sentiment analysis via integrating distributed representations of variable-length word sequence
Zhijian Cui, Xiaodong Shi, Yidong Chen 0001
Neurocomputing3
2016 Adapted competitive learning on continuous semantic space for word sense induction
Yanzhou Huang, Deyi Xiong, Xiaodong Shi, Yidong Chen 0001, Changxing Wu, Guimin Huang
Neurocomputing4
2016 An SNN-Based Semantic Role Labeling Model with Its Network Parameters Optimized Using an Improved PSO Algorithm
Yidong Chen 0001, Zhehuang Huang, Xiaodong Shi
Neural Process. Lett.1
2015 Unsupervised word sense induction using rival penalized competitive learning
Yanzhou Huang, Xiaodong Shi, Jinsong Su, Yidong Chen 0001, Guimin Huang
Eng. Appl. Artif. Intell.4
2014 Topic-aware pivot language approach for statisticalmachine translation
abstract
The pivot language approach for statistical machine translation (SMT) is a good method to break the resource bottleneck for certain language pairs. However, in the implementation of conventional approaches, pivot-side context information is far from fully utilized, resulting in erroneous estimations of translation probabilities. In this study, we propose two topic-aware pivot language approaches to use different levels of pivot-side context. The first method takes advantage of document-level context by assuming that the bridged phrase pairs should be similar in the document-level topic distributions. The second method focuses on the effect of local context. Central to this approach are that the phrase sense can be reflected by local context in the form of probabilistic topics, and that bridged phrase pairs should be compatible in the latent sense distributions. Then, we build an interpolated model bringing the above methods together to further enhance the system performance. Experimental results on French-Spanish and French-German translations using English as the pivot language demonstrate the effectiveness of topic-based context in pivot-based SMT.
Jinsong Su, Xiaodong Shi, Yanzhou Huang, Yang Liu 0005, Qingqiang Wu 0001, Yidong Chen 0001, Huailin Dong
J. Zhejiang Univ. Sci. C6
2013 Improving Alignment of System Combination by Using Multi-objective Optimization
abstract
This paper proposes a multi-objective optimization framework which supports heterogeneous information sources to improve alignment in machine translation system combination techniques.In this area, most of techniques usually utilize confusion networks (CN) as their central data structure to compact an exponential number of an potential hypotheses, and because better hypothesis alignment may benefit constructing better quality confusion networks, it is natural to add more useful information to improve alignment results.However, these information may be heterogeneous, so the widely-used Viterbi algorithm for searching the best alignment may not apply here.In the multi-objective optimization framework, each information source is viewed as an independent objective, and a new goal of improving all objectives can be searched by mature algorithms.The solutions from this framework, termed Pareto optimal solutions, are then combined to construct confusion networks.Experiments on two Chinese-to-English translation datasets show significant improvements, 0.97 and 1.06 BLEU points over a strong Indirected Hidden Markov Model-based (IHMM) system, and 4.75 and 3.53 points over the best single machine translation systems.
Tian Xia 0004, Zongcheng Ji, Shaodan Zhai, Yidong Chen 0001, Qun Liu 0001
EMNLP4
2012 Translation Model Adaptation for Statistical Machine Translation with Monolingual Topic Information
Jinsong Su, Hua Wu 0003, Haifeng Wang 0001, Yidong Chen 0001, Xiaodong Shi, Huailin Dong, Qun Liu 0001
ACL (1)4
2011 Improving the Hierarchical Phrase-Based Translation Model
Xiaodong Shi, Yidong Chen 0001
MTSummit3
2007 Dependency-Based Chinese-English Statistical Machine Translation
Xiaodong Shi, Yidong Chen 0001, Jianfeng Jia
CICLing2
2007 Translation Memory Sharing Models in XMCAT
abstract
In this paper, two Translation Memory (TM) sharing models adopted in XMCAT, a Computer Assisted Translation tool (CAT) supporting cooperated work in machine translation, was described in detail. One is Center-based TM sharing model, which is only fit for users in a local area network (LAN) and the other is a novel model called P2P-based TM sharing model, which could be used through Internet by geographically distributed users. With the two TM sharing models, a user may share data with other users through network, so that he/she may reduce the repeated work further and cooperate with others more easily. Besides, the methods used in XMCA T to deal with the problem of multi-translations arose in the cooperated memory sharing models, were also proposed in this paper. XMCAT system has been adopted and approved by some translation companies.
Yidong Chen 0001, Xiaodong Shi, Changle Zhou, Tangqiu Li, Qingyang Hong
CSCWD1
2003 A Word Selection Model Based on Lexical Semantic Knowledge in English Generation
Yidong Chen 0001, Tangqiu Li, Xu-Ling Zheng
PACLIC1
2002 IP Multicasting Technique and its Application on Multimedia Network Teaching System
abstract
This paper presents the basic principles of IP multicasting. The design and application in a multimedia network teaching system are then given.
Shaozi Li, Tangqiu Li, Yidong Chen 0001
CSCWD3