VLDB 2026 Research / reviewers in the wild / expert
Yun Chen 0007
dblp:10/5680-7
· DBLP profile ↗
33ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0002-3563-7592ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal BeliefabstractLarge Language Models (LLMs) have achieved remarkable success across a wide range of natural language tasks, but often exhibit overconfidence and generate plausible yet incorrect answers. This overconfidence, especially in models undergone Reinforcement Learning from Human Feedback (RLHF), poses significant challenges for reliable uncertainty estimation and safe deployment. In this paper, we propose EAGLE (Expectation of AGgregated internaL bEief), a novel self-evaluation-based calibration method that leverages the internal hidden states of LLMs to derive more accurate confidence scores. Instead of relying on the model's final output, our approach extracts internal beliefs from multiple intermediate layers during self-evaluation. By aggregating these layer-wise beliefs and calculating the expectation over the resulting confidence score distribution, EAGLE produces a refined confidence score that more faithfully reflects the model's internal certainty. Extensive experiments on diverse datasets and LLMs demonstrate that EAGLE significantly improves calibration performance over existing baselines. We also provide an in-depth analysis of EAGLE, including a layer-wise examination of uncertainty patterns, a study of the impact of self-evaluation prompts, and an analysis of the effect of self-evaluation score range. Zeguan Xiao, Diyang Dou, Boya Xiong, Yun Chen 0007, Guanhua Chen 0001 |
AAAI | 4 |
| 2026 | GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language ModelsabstractZhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen, Peng Li, Jinsong Su, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen 0007, Peng Li 0030, Jinsong Su, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 5 |
| 2026 | InstructDiff: Domain-Adaptive Data Selection via Contrastive Entropy for Efficient LLM Fine-TuningabstractJunyou Su, He Zhu, Xiao Luo, Liyu Zhang, Hong-Yu Zhou, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junyou Su, Xiao Luo 0001, Liyu Zhang 0010, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 6 |
| 2026 | SPPO: Sequence-Level PPO for Long-Horizon Reasoning TasksabstractTianyi Wang, Yixia Li, Long Li, Yibiao Chen, Shaohan Huang, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixia Li, Yibiao Chen, Shaohan Huang, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 6 |
| 2026 | Modeling LLM Unlearning as an Asymmetric Two-Task Learning ProblemabstractMachine unlearning for large language models (LLMs) aims to remove targeted knowledge while preserving general capability.In this paper, we recast LLM unlearning as an asymmetric two-task problem: retention is the primary objective and forgetting is an auxiliary.From this perspective, we propose a retention-prioritized gradient synthesis framework that decouples task-specific gradient extraction from conflict-aware combination.Instantiating the framework, we adapt established PCGrad to resolve gradient conflicts, and introduce SAGO, a novel retention-prioritized gradient synthesis method.Theoretically, both variants ensure non-negative cosine similarity with the retain gradient, while SAGO achieves strictly tighter alignment through constructive sign-constrained synthesis.Empirically, on WMDP Bio/Cyber and RWKU benchmarks, SAGO consistently pushes the Pareto frontier: e.g., on WMDP Bio (SimNPO+GD), recovery of target model MMLU performance progresses from 44.6% (naive) to 94.0% (+PC-Grad) and further to 96.0% (+SAGO), while maintaining comparable forgetting strength.Our results show that re-shaping gradient geometry, rather than re-balancing losses, is the key to mitigating unlearning-retention trade-offs. Zeguan Xiao, Siqing Li, Yong Wang 0032, Xuetao Wei, Jian Yang 0003, Yun Chen 0007, Guanhua Chen 0001 |
ACL (1) | 6 |
| 2025 | ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMsabstractWith the proliferation of task-specific large language models, delta compression has emerged as a method to mitigate the resource challenges of deploying numerous such models by effectively compressing the delta model parameters. Previous delta-sparsification methods either remove parameters randomly or truncate singular vectors directly after singular value decomposition (SVD). However, these methods either disregard parameter importance entirely or evaluate it with too coarse a granularity. In this work, we introduce ImPart, a novel importance-aware delta sparsification approach. Leveraging SVD, it dynamically adjusts sparsity ratios of different singular vectors based on their importance, effectively retaining crucial task-specific knowledge even at high sparsity ratios. Experiments show that ImPart achieves state-of-the-art delta sparsification performance, demonstrating 2\times higher compression ratio than baselines at the same performance level. When integrated with existing methods, ImPart sets a new state-of-the-art on delta quantization and model merging. Yixia Li, Hongru Wang 0003, Xuetao Wei, James Jian Qiao Yu, Yun Chen 0007, Guanhua Chen 0001 |
ACL (1) | 6 |
| 2025 | G2: Guided Generation for Enhanced Output Diversity in LLMsabstractLarge Language Models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks.However, these models exhibit a critical limitation in output diversity, often generating highly similar content across multiple attempts.This limitation significantly affects tasks requiring diverse outputs, from creative writing to reasoning.Existing solutions, like temperature scaling, enhance diversity by modifying probability distributions but compromise output quality.We propose Guide-to-Generation (G2), a trainingfree plug-and-play method that enhances output diversity while preserving generation quality.G2 employs a base generator alongside dual Guides, which guide the generation process through decoding-based interventions to encourage more diverse outputs conditioned on the original query.Comprehensive experiments demonstrate that G2 effectively improves output diversity while maintaining an optimal balance between diversity and quality. Zhiwen Ruan, Yixia Li, Yefeng Liu, Yun Chen 0007, Weihua Luo, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
EMNLP | 4 |
| 2025 | MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM FinetuningabstractHanqing Wang, Yixia Li, Shuo Wang, Guanhua Chen, Yun Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hanqing Wang 0003, Yixia Li, Shuo Wang 0013, Guanhua Chen 0001, Yun Chen 0007 |
NAACL (Long Papers) | 5 |
| 2025 | SeqAR: Jailbreak LLMs with Sequential Auto-Generated CharactersabstractYan Yang, Zeguan Xiao, Xin Lu, Hongru Wang, Xuetao Wei, Hailiang Huang, Guanhua Chen, Yun Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zeguan Xiao, Hongru Wang 0003, Xuetao Wei, Guanhua Chen 0001, Yun Chen 0007 |
NAACL (Long Papers) | 8 |
| 2025 | Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsabstractThe growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as “LLM-as-a-judge”. However, improving its alignment with human preferences without complex prompts or fine-tuning remains challenging. Previous studies mainly optimize based on shallow outputs, overlooking rich cross-layer representations. In this work, motivated by preliminary findings that middle-to-upper layers encode semantically and task-relevant representations that are often more aligned with human judgments than the final layer, we propose LAGER, a post-hoc, plug-and-play framework for improving the alignment of LLM-as-a-Judge point-wise evaluations with human scores by leveraging internal representations. LAGER produces fine-grained judgment scores by aggregating cross-layer score-token logits and computing the expected score from a softmax-based distribution, while keeping the LLM backbone frozen and ensuring no impact on the inference process.
LAGER fully leverages the complementary information across different layers, overcoming the limitations of relying solely on the final layer.
We evaluate our method on the standard alignment benchmarks Flask, HelpSteer, and BIGGen using Spearman correlation, and find that LAGER achieves improvements of up to 7.5% over the best baseline across these benchmarks. Without reasoning steps, LAGER matches or outperforms reasoning-based methods. Experiments on downstream applications, such as data selection and emotional understanding, further show the generalization of LAGER. Peng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
NeurIPS | 4 |
| 2024 | LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative TasksabstractLoRA employs lightweight modules to customize large language models (LLMs) for each downstream task or domain, where different learned additional modules represent diverse skills.Combining existing LoRA modules to address new tasks can enhance the reusability of learned LoRA modules, particularly beneficial for tasks with limited annotated data.Most prior works on LoRA combination primarily rely on task-level weights for each involved LoRA, making different examples and tokens share the same LoRA weights.However, in generative tasks, different tokens may necessitate diverse skills to manage.Taking the Chinese math task as an example, understanding the problem description may depend more on the Chinese LoRA, while the calculation part may rely more on the math LoRA.To this end, we propose LoRA-Flow, which utilizes dynamic weights to adjust the impact of different LoRA modules.The weights at each step are determined by a fusion gate with extremely few parameters, which can be learned with only 200 training examples.Experiments across six generative tasks demonstrate that our method consistently outperforms baselines with tasklevel fusion weights.This underscores the necessity of introducing dynamic fusion weights for LoRA combination. 1 Hanqing Wang 0003, Bowen Ping, Shuo Wang 0013, Xu Han 0007, Yun Chen 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 5 |
| 2024 | Distract Large Language Models for Automatic Jailbreak AttackabstractExtensive efforts have been made before the public release of Large language models (LLMs) to align their behaviors with human values.However, even meticulously aligned LLMs remain vulnerable to malicious manipulations such as jailbreaking, leading to unintended behaviors.In this work, we propose a novel black-box jailbreak framework for automated red teaming of LLMs.We designed malicious content concealing and memory reframing with an iterative optimization algorithm to jailbreak LLMs, motivated by the research about the distractibility and over-confidence phenomenon of LLMs.Extensive experiments of jailbreaking both open-source and proprietary LLMs demonstrate the superiority of our framework in terms of effectiveness, scalability and transferability.We also evaluate the effectiveness of existing jailbreak defense methods against our attack and highlight the crucial need to develop more effective and practical defense strategies.Warning: This paper contains unfiltered content generated by LLMs that may be offensive to readers. Zeguan Xiao, Guanhua Chen 0001, Yun Chen 0007 |
EMNLP | 4 |
| 2024 | SeTAR: Out-of-Distribution Detection with Selective Low-Rank ApproximationabstractOut-of-distribution (OOD) detection is crucial for the safe deployment of neural networks. Existing CLIP-based approaches perform OOD detection by devising novel scoring functions or sophisticated fine-tuning methods. In this work, we propose SeTAR, a novel, training-free OOD detection method that leverages selective low-rank approximation of weight matrices in vision-language and vision-only models. SeTAR enhances OOD detection via post-hoc modification of the model's weight matrices using a simple greedy search algorithm. Based on SeTAR, we further propose SeTAR+FT, a fine-tuning extension optimizing model performance for OOD detection tasks. Extensive evaluations on ImageNet1K and Pascal-VOC benchmarks show SeTAR's superior performance, reducing the relatively false positive rate by up to 18.95\% and 36.80\% compared to zero-shot and fine-tuning baselines. Ablation studies further validate our approach's effectiveness, robustness, and generalizability across different model backbones. Our work offers a scalable, efficient solution for OOD detection, setting a new state-of-the-art in this area. Yixia Li, Boya Xiong, Guanhua Chen 0001, Yun Chen 0007 |
NeurIPS | 4 |
| 2024 | Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language ModelsabstractFine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability. Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 7 |
| 2023 | mCLIP: Multilingual CLIP via Cross-lingual TransferabstractGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai, Lifeng Shang, Xin Jiang, Qun Liu, Jia Pan, Wenping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Guanhua Chen 0001, Lu Hou 0002, Yun Chen 0007, Wenliang Dai, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Jia Pan 0001, Wenping Wang 0001 |
ACL (1) | 3 |
| 2023 | PrivateRec: Differentially Private Model Training and Online Serving for Federated News RecommendationabstractFederated recommendation can potentially alleviate the privacy concerns in collecting sensitive and personal data for training personalized recommendation systems. However, it suffers from a low recommendation quality when a local serving is inapplicable due to the local resource limitation and the data privacy of querying clients is required in online serving. Furthermore, a theoretically private solution in both the training and serving of federated recommendation is essential but still lacking. Naively applying differential privacy (DP) to the two stages in federated recommendation would fail to achieve a satisfactory trade-off between privacy and utility due to the high-dimensional characteristics of model gradients and hidden representations. In this work, we propose a federated news recommendation method for achieving better utility in model training and online serving under a DP guarantee. We first clarify the DP definition over behavior data for each round in the pipeline of federated recommendation systems. Next, we propose a privacy-preserving online serving mechanism under this definition based on the idea of decomposing user embeddings with public basic vectors and perturbing the lower-dimensional combination coefficients. We apply a random behavior padding mechanism to reduce the required noise intensity for better utility. Besides, we design a federated recommendation model training method, which can generate effective and public basic vectors for serving while providing DP for training participants. We avoid the dimension-dependent noise for large models via label permutation and differentially private attention modules. Experiments on real-world news recommendation datasets validate that our method achieves superior utility under a DP guarantee in both training and serving of federated news recommendations. Ruixuan Liu, Yang Cao 0011, Yanlin Wang 0001, Lingjuan Lyu, Yun Chen 0007, Hong Chen 0001 |
KDD | 5 |
| 2022 | Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine TranslationabstractGuanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang, Jia Pan, Wenping Wang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
ACL (1) | 3 |
| 2022 | Multitasking Framework for Unsupervised Simple Definition GenerationabstractThe definition generation task can help language learners by providing explanations for unfamiliar words.This task has attracted much attention in recent years.We propose a novel task of Simple Definition Generation (SDG) to help language learners and low literacy readers.A significant challenge of this task is the lack of learner's dictionaries in many languages, and therefore the lack of data for supervised training.We explore this task and propose a multitasking framework SimpDefiner that only requires a standard dictionary with complex definitions and a corpus containing arbitrary simple texts.We disentangle the complexity factors from the text by carefully designing a parameter sharing scheme between two decoders.By jointly training these components, the framework can generate both complex and simple definitions simultaneously.We demonstrate that the framework can generate relevant, simple definitions for the target words through automatic and manual evaluations on English and Chinese datasets.Our method outperforms the baseline model by a 1.77 SARI score on the English dataset, and raises the proportion of the low level (HSK level 1-3) words in Chinese definitions by 3.87% 1 . Cunliang Kong, Yun Chen 0007, Liner Yang, Erhong Yang |
ACL (1) | 2 |
| 2022 | XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine TranslationabstractPre-training language models have achieved thriving success in numerous natural language understanding and autoregressive generation tasks, but non-autoregressive generation in applications such as machine translation has not sufficiently benefited from the pre-training paradigm.In this work, we establish the connection between a pre-trained masked language model (MLM) and non-autoregressive generation on machine translation.From this perspective, we present XLM-D, which seamlessly transforms an off-the-shelf cross-lingual pre-training model into a non-autoregressive translation (NAT) model with a lightweight yet effective decorator.Specifically, the decorator ensures the representation consistency of the pre-trained model and brings only one additional trainable parameter.Extensive experiments on typical translation datasets show that our models obtain state-of-the-art performance while realizing the inference speedup by 19.9×.One striking result is that on WMT14 En⇒De, our XLM-D obtains 29.80 BLEU points with multiple iterations, which outperforms the previous mask-predict model by 2.77 points. Yong Wang 0032, Shilin He, Guanhua Chen 0001, Yun Chen 0007, Daxin Jiang |
EMNLP | 4 |
| 2022 | Controllable data synthesis method for grammatical error correction
Liner Yang, Yun Chen 0007, Yongping Du, Erhong Yang |
Frontiers Comput. Sci. | 3 |
| 2021 | Lexically Constrained Neural Machine Translation with Explicit Alignment GuidanceabstractLexically constrained neural machine translation (NMT), which leverages pre-specified translation to constrain NMT, has practical significance in interactive translation and NMT domain adaption. Previous work either modify the decoding algorithm or train the model on augmented dataset. These methods suffer from either high computational overheads or low copying success rates. In this paper, we investigate Att-Input and Att-Output, two alignment-based constrained decoding methods. These two methods revise the target tokens during decoding based on word alignments derived from encoder-decoder attention weights. Our study shows that Att-Input translates better while Att-Output is more computationally efficient. Capitalizing on both strengths, we further propose EAM-Output by introducing an explicit alignment module (EAM) to a pretrained Transformer. It decodes similarly as EAM-Output, except using alignments derived from the EAM. We leverage the word alignments induced from Att-Input as labels and train the EAM while keeping the parameters of the Transformer frozen. Experiments on WMT16 De-En and WMT16 Ro-En show the effectiveness of our approaches on constrained NMT. In particular, the proposed EAM-Output method consistently outperforms previous approaches in translation quality, with light computational overheads over unconstrained baseline. Guanhua Chen 0001, Yun Chen 0007, Victor O. K. Li |
AAAI | 2 |
| 2021 | Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersabstractPrevious work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT.However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transferability of NMT model.In this paper, we focus on a zero-shot cross-lingual transfer task in NMT.In this task, the NMT model is trained with parallel dataset of only one language pair and an off-the-shelf MPE, then it is directly tested on zero-shot language pairs.We propose SixT, a simple yet effective model for this task.SixT leverages the MPE with a two-stage training schedule and gets further improvement with a position disentangled encoder and a capacity-enhanced decoder.Using this method, SixT significantly outperforms mBART, a pretrained multilingual encoderdecoder model explicitly designed for NMT, with an average improvement of 7.1 BLEU on zero-shot any-to-English test sets across 14 source languages.Furthermore, with much less training computation cost and training data, our model achieves better performance on 15 any-to-English test sets than CRISS and m2m-100, two strong multilingual NMT baselines. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Li Dong 0004, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
EMNLP (1) | 3 |
| 2020 | Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTabstractBy introducing a small set of additional parameters, a probe learns to solve specific linguistic tasks (e.g., dependency parsing) in a supervised manner using feature representations (e.g., contextualized embeddings).The effectiveness of such probing tasks is taken as evidence that the pre-trained model encodes linguistic knowledge.However, this approach of evaluating a language model is undermined by the uncertainty of the amount of knowledge that is learned by the probe itself.Complementary to those works, we propose a parameter-free probing technique for analyzing pre-trained language models (e.g., BERT).Our method does not require direct supervision from the probing tasks, nor do we introduce additional parameters to the probing process.Our experiments on BERT show that syntactic trees recovered from BERT using our method are significantly better than linguistically-uninformed baselines.We further feed the empirically induced dependency structures into a downstream sentiment classification task and find its improvement compatible with or even superior to a human-designed dependency schema. Zhiyong Wu 0003, Yun Chen 0007, Ben Kao, Qun Liu 0001 |
ACL | 2 |
| 2020 | Accurate Word Alignment Induction from Neural Machine TranslationabstractDespite its original goal to jointly learn to align and translate, prior researches suggest that Transformer captures poor word alignments through its attention mechanism.In this paper, we show that attention weights DO capture accurate word alignments and propose two novel word alignment induction methods SHIFT-ATT and SHIFT-AET.The main idea is to induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work.SHIFT-ATT is an interpretation method that induces alignments from the attention weights of Transformer and does not require parameter update or architecture change.SHIFT-AET extracts alignments from an additional alignment module which is tightly integrated into Transformer and trained in isolation with supervision from symmetrized SHIFT-ATT alignments.Experiments on three publicly available datasets demonstrate that both methods perform better than their corresponding neural baselines and SHIFT-AET significantly outperforms GIZA++ by 1.4-4.8AER points. 1 Yun Chen 0007, Yang Liu 0005, Guanhua Chen 0001, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 1 |
| 2020 | Lexical-Constraint-Aware Neural Machine Translation via Data AugmentationabstractLeveraging lexical constraint is extremely significant in domain-specific machine translation and interactive machine translation. Previous studies mainly focus on extending beam search algorithm or augmenting the training corpus by replacing source phrases with the corresponding target translation. These methods either suffer from the heavy computation cost during inference or depend on the quality of the bilingual dictionary pre-specified by user or constructed with statistical machine translation. In response to these problems, we present a conceptually simple and empirically effective data augmentation approach in lexical constrained neural machine translation. Specifically, we make constraint-aware training data by first randomly sampling the phrases of the reference as constraints, and then packing them together into the source sentence with a separation symbol. Extensive experiments on several language pairs demonstrate that our approach achieves superior translation results over the existing systems, improving translation of constrained sentences without hurting the unconstrained ones. Guanhua Chen 0001, Yun Chen 0007, Yong Wang 0032, Victor O. K. Li |
IJCAI | 2 |
| 2020 | Reinforced Zero-Shot Cross-Lingual Neural Headline GenerationabstractCross-lingual neural headline generation (CNHG), which aims at training a single, large neural network that directly generates a target language headline given a source language news document, has received considerable attention in recent years. Unlike conventional neural headline generation, CNHG faces the problem that there are no large-scale parallel corpora of source language articles and target language headlines. Consequently, CNHG is a zero-shot scenario. To solve this problem, we propose zero resource CNHG with reinforcement learning. We develop a reinforcement learning framework that is composed of two modules: a neural machine translation (NMT) module and a CNHG module. The translation module translates an input document into a source language document, and the headline generation module takes the previous output as input to generate a target language headline. Then, both modules receive a reward for joint training. The experimental results reveal that our method significantly outperforms baseline models. Ayana, Yun Chen 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Incorporating Sememes into Chinese Definition ModelingabstractChinese definition modeling is a challenging task that generates a dictionary definition in Chinese for a given Chinese word. To accomplish this task, we built two novel datasets based on Chinese Concept Dictionary (CCD) and Chinese WordNet (CWN) respectively. Each dataset contains triples of a word, sememes, and a corresponding definition. We present two novel models to improve Chinese definition modeling: the Adaptive-Attention model (AAM) and the Self- and Adaptive-Attention Model (SAAM). AAM successfully incorporates sememes for generating the definition with an adaptive attention mechanism. It has the capability to decide which sememes to focus on and when to pay attention to sememes. SAAM further replaces recurrent connections in AAM with self-attention and relies entirely on the attention mechanism, reducing the path length between word, sememes and definition. Experiments on both datasets demonstrate that by incorporating sememes, our model can generate definitions with more concrete information. And the best model that we proposed outperforms the state-of-the-art method by a large margin on both datasets. Liner Yang, Cunliang Kong, Yun Chen 0007, Yang Liu 0005, Qinan Fan, Erhong Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Zero-Resource Neural Machine Translation with Multi-Agent Communication GameabstractWhile end-to-end neural machine translation (NMT) has achieved notable success in the past years in translating a handful of resource-rich language pairs, it still suffers from the data scarcity problem for low-resource language pairs and domains. To tackle this problem, we propose an interactive multimodal framework for zero-resource neural machine translation. Instead of being passively exposed to large amounts of parallel corpora, our learners (implemented as encoder-decoder architecture) engage in cooperative image description games, and thus develop their own image captioning or neural machine translation model from the need to communicate in order to succeed at the game. Experimental results on the IAPR-TC12 and Multi30K datasets show that the proposed learning mechanism significantly improves over the state-of-the-art methods. Yun Chen 0007, Yang Liu 0005, Victor O. K. Li |
AAAI | 1 |
| 2018 | A Stable and Effective Learning Strategy for Trainable Greedy DecodingabstractBeam search is a widely used approximate search strategy for neural network decoders, and it generally outperforms simple greedy decoding on tasks like machine translation.However, this improvement comes at substantial computational cost.In this paper, we propose a flexible new method that allows us to reap nearly the full benefits of beam search with nearly no additional computational cost.The method revolves around a small neural network actor that is trained to observe and manipulate the hidden state of a previouslytrained decoder.To train this actor network, we introduce the use of a pseudo-parallel corpus built using the output of beam search on a base model, ranked by a target quality metric like BLEU.Our method is inspired by earlier work on this problem, but requires no reinforcement learning, and can be trained reliably on a range of models.Experiments on three parallel corpora and three architectures show that the method yields substantial improvements in translation quality and speed over each base system. Yun Chen 0007, Victor O. K. Li, Kyunghyun Cho, Samuel R. Bowman |
EMNLP | 1 |
| 2018 | Meta-Learning for Low-Resource Neural Machine TranslationabstractIn this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT).We frame low-resource translation as a metalearning problem, and we learn to adapt to low-resource languages based on multilingual high-resource language tasks.We use the universal lexical representation (Gu et al., 2018b) to overcome the input-output mismatch across different languages.We evaluate the proposed meta-learning strategy using eighteen European languages (Bg, Cs, Da, De, El, Es, Et, Fr, Hu, It, Lt, Nl, Pl, Pt, Sk, Sl, Sv and Ru) as source tasks and five diverse languages (Ro, Lv, Fi, Tr and Ko) as target tasks.We show that the proposed approach significantly outperforms the multilingual, transfer learning based approach (Zoph et al., 2016) and enables us to train a competitive NMT system with only a fraction of training examples.For instance, the proposed approach can achieve as high as 22.04 BLEU on Romanian-English WMT'16 by seeing only 16,000 translated words (⇠ 600 parallel sentences). Jiatao Gu, Yong Wang 0032, Yun Chen 0007, Victor O. K. Li, Kyunghyun Cho |
EMNLP | 3 |
| 2018 | Zero-Shot Cross-Lingual Neural Headline GenerationabstractNeural headline generation (NHG) has been proven to be effective in generating a fully abstractive headline recently. Existing NHG systems are only capable of producing headline of the same language as the original document. Cross lingual headline generation is an important task since it provides an efficient way to understand the key point of a document in a different language. Due to the lack of those parallel corpora of direct source language articles and target language headlines, we propose to deal with the cross-lingual neural headline generation (CNHG) under the zero-shot scenario. A trivial solution is to translate and summarize the source document in a pipeline way. However, a pipeline solution will lead to error propagation in the translation and summarization phases. This challenge motivates us to build a direct source-to-target CNHG model based on existing parallel corpora of translation and monolingual headline generation. Specifically, we let a parameterized CNHG model (student model) mimic the output of a pretrained translation or headline generation model (teacher model). To the best of our knowledge, this is the first effort to address CNHG problem. Besides, we construct English-Chinese headline generation evaluation datasets by manual translation. Experimental results on English-to-Chinese cross-lingual headline generation demonstrate that our proposed method significantly outperforms the baseline models. Shiqi Shen, Yun Chen 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | A Teacher-Student Framework for Zero-Resource Neural Machine TranslationabstractWhile end-to-end neural machine translation (NMT) has made remarkable progress recently, it still suffers from the data scarcity problem for low-resource language pairs and domains.In this paper, we propose a method for zero-resource NMT by assuming that parallel sentences have close probabilities of generating a sentence in a third language.Based on the assumption, our method is able to train a source-to-target NMT model ("student") without parallel corpora available guided by an existing pivot-to-target NMT model ("teacher") on a source-pivot parallel corpus.Experimental results show that the proposed method significantly improves over a baseline pivot-based model by +3.0 BLEU points across various language pairs. Yun Chen 0007, Yang Liu 0005, Yong Cheng 0003, Victor O. K. Li |
ACL (1) | 1 |
| 2017 | Deep Learning Model to Estimate Air Pollution Using M-BP to Fill in Missing Proxy Urban DataabstractAir quality has deteriorated rapidly in Hong Kong and China in the past two decades, with NO2and PM2.5levels frequently exceeding WHO safety guidelines. While poor air quality has clear public health impacts, there are very limited air quality monitoring (AQM) stations, severely constraining evidence-based air quality decision-making, leading to severe criticisms about the utility of the current official Air Quality Health Index to the public. Since air pollution is highly location-dependent, a city-wide deployment of traditional, highly sophisticated air quality monitors would be prohibitively expensive. In this paper, we propose a deep learning model to estimate air pollution throughout the city, utilizing the readily available urban data as proxy data. As with many big data driven approaches, the proxy data may be sparse/missing. We propose the M-BP algorithm to recover/fill in such missing data. Our results show that the proposed model gives better estimates compared with existing big data approaches. Victor O. K. Li, Jacqueline C. K. Lam, Yun Chen 0007, Jiatao Gu |
GLOBECOM | 3 |