EDBT 2026 Demo / reviewers in the wild / expert
Baosong Yang
dblp:203/8245
· DBLP profile ↗
52ranked-venue papers
6as first author
40since 2021 · last 2026
0000-0001-5002-2409ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 6 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEabstractHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She, Linjuan Wu, Hao-Ran Wei, Baosong Yang, Jiajun Chen, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Zhou 0012, Shuaijie She, Linjuan Wu, Baosong Yang, Jiajun Chen 0001, Shujian Huang |
ACL (1) | 7 |
| 2026 | A Systematic Assessment of Language Models with Linguistic Minimal Pairs in ChineseabstractAbstract We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the Ba construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that Anaphor, Quantifiers, and Ellipsis in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs. Yikang Liu 0002, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Jialong Tang, Pei Zhang 0011, Baosong Yang, Rui Wang 0015, Hai Hu 0001 |
Trans. Assoc. Comput. Linguistics | 10 |
| 2025 | Unveiling Language-Specific Features in Large Language Models via Sparse AutoencodersabstractThe mechanisms behind multilingual capabilities in Large Language Models (LLMs) have been examined using neuron-based or internal-activation-based methods.However, these methods often face challenges such as superposition and layer-wise activation variance, which limit their reliability.Sparse Autoencoders (SAEs) offer a more nuanced analysis by decomposing the activations of LLMs into a sparse linear combination of SAE features.We introduce a novel metric to assess the monolinguality of features obtained from SAEs, discovering that some features are strongly related to specific languages.Additionally, we show that ablating these SAE features only significantly reduces abilities in one language of LLMs, leaving others almost unaffected.Interestingly, we find some languages have multiple synergistic SAE features, and ablating them together yields greater improvement than ablating individually.Moreover, we leverage these SAE-derived language-specific features to enhance steering vectors, achieving control over the language generated by LLMs.The code is publicly available at https://github.com/ Aatrox103/multilingual-llm-features. Boyi Deng, Yu Wan 0004, Baosong Yang, Yidan Zhang 0004, Fuli Feng |
ACL (1) | 3 |
| 2025 | Enhancing Machine Translation with Self-Supervised Preference DataabstractModel alignment methods like Direct Preference Optimization and Contrastive Preference Optimization have enhanced machine translation performance by leveraging preference data to enable models to reject suboptimal outputs. During preference data construction, previous approaches primarily rely on humans, strong models like GPT4 or model self-sampling. In this study, we first explain the shortcomings of this practice. Then, we propose Self-Supervised Preference Optimization (SSPO), a novel framework which efficiently constructs translation preference data for iterative DPO training. Applying SSPO to 14B parameters large language models (LLMs) achieves comparable or better performance than GPT-4o on FLORES and multi-domain test datasets. We release an augmented MQM dataset in https://github.com/sunny-sjtu/MQM-aug. Haoxiang Sun, Pei Zhang 0011, Baosong Yang, Rui Wang 0015 |
ACL (1) | 4 |
| 2025 | Locate-and-Focus: Enhancing Terminology Translation in Speech Language ModelsabstractDirect speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance. Suhang Wu, Jialong Tang, Pei Zhang 0011, Baosong Yang, Junhui Li 0001, Junfeng Yao, Min Zhang 0005, Jinsong Su |
ACL (1) | 5 |
| 2025 | From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction TuningabstractSupervised Fine-Tuning (SFT) with translated instruction data effectively adapts Large Language Models (LLMs) from English to non-English languages.We introduce Cross-Lingual Continued Instruction Tuning (X-CIT), which fully leverages translation-based parallel instruction data to enhance cross-lingual adaptability.X-CIT emulates the human process of second language acquisition and is guided by Chomsky's Principles and Parameters Theory.It first fine-tunes the LLM on English instruction data to establish foundational capabilities (i.e.Principles), then continues with target language translation and customized chatinstruction data to adjust "parameters" specific to the target language.This chat-instruction data captures alignment information in translated parallel data, guiding the model to initially think and respond in its native language before transitioning to the target language.To further mimic human learning progression, we incorporate Self-Paced Learning (SPL) during continued training, allowing the model to advance from simple to complex tasks.Implemented on Llama-2-7B across five languages, X-CIT was evaluated against three objective benchmarks and an LLM-as-a-judge benchmark, improving the strongest baseline by an average of 1.97% and 8.2% in these two benchmarks, respectively. Linjuan Wu, Baosong Yang, Weiming Lu 0001 |
ACL (1) | 3 |
| 2025 | Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseabstractYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang, Pei Zhang, Baosong Yang, Fei Huang, Rui Wang, Hai Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yikang Liu 0002, Wanyang Zhang, Yiming Wang 0011, Jialong Tang, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Rui Wang 0015, Hai Hu 0001 |
EMNLP | 6 |
| 2025 | Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-trainingabstractLarge language models (LLMs) exhibit remarkable multilingual capabilities despite Englishdominated pre-training, attributed to crosslingual mechanisms during pre-training.Existing methods for enhancing cross-lingual transfer remain constrained by parallel resources, suffering from limited linguistic and domain coverage.We propose Cross-lingual In-context Pre-training (CrossIC-PT), a simple and scalable approach that enhances cross-lingual transfer by leveraging semantically related bilingual texts via simple next-word prediction.We construct CrossIC-PT samples by interleaving semantic-related bilingual Wikipedia documents into a single context window.To access window size constraints, we implement a systematic segmentation policy to split long bilingual document pairs into chunks while adjusting the sliding window mechanism to preserve contextual coherence.We further extend data availability through a semantic retrieval framework to construct CrossIC-PT samples from web-crawled corpus.Experimental results demonstrate that CrossIC-PT improves multilingual performance on three models (Llama-3.1-8B,Qwen2.5-7B, and Qwen2.5-1.5B)across six target languages, yielding performance gains of 3.79%, 3.99%, and 1.95%, respectively, with additional improvements after data augmentation. Linjuan Wu, Baosong Yang, Fei Huang 0002, Weiming Lu 0001 |
EMNLP | 5 |
| 2025 | P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMsabstractYidan Zhang, Yu Wan, Boyi Deng, Baosong Yang, Hao-Ran Wei, Fei Huang, Bowen Yu, Dayiheng Liu, Junyang Lin, Fei Huang, Jingren Zhou. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yidan Zhang 0004, Yu Wan 0004, Boyi Deng, Baosong Yang, Fei Huang 0002, Bowen Yu 0002, Dayiheng Liu, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
EMNLP | 4 |
| 2025 | NOVA-63: Native Omni-lingual Versatile Assessments of 63 DisciplinesabstractThe multilingual capabilities of large language models (LLMs) have attracted considerable attention over the past decade. Assessing the accuracy with which LLMs provide answers in multilingual contexts is essential for determining their level of multilingual proficiency. Nevertheless, existing multilingual benchmarks generally reveal severe drawbacks, such as overly translated content (translationese), the absence of difficulty control, constrained diversity, and disciplinary imbalance, making the benchmarking process unreliable and showing low convincingness. To alleviate those shortcomings, we introduce NOVA-63 (Native Omni-lingual Versatile Assessments of 63 Disciplines), a comprehensive, difficult multilingual benchmark featuring 93,536 questions sourced from native speakers across 14 languages and 63 academic disciplines. Leveraging a robust pipeline that integrates LLM-assisted formatting, expert quality verification, and multi-level difficulty screening, NOVA-63 is balanced on disciplines with consistent difficulty standards while maintaining authentic linguistic elements. Extensive experimentation with current LLMs has shown significant insights into cross-lingual consistency among language families, and exposed notable disparities in models’ capabilities across various disciplines. This work provides valuable benchmarking data for the future development of multilingual models. Furthermore, our findings underscore the importance of moving beyond overall scores and instead conducting fine-grained analyses of model performance. Kexin Yang 0002, Yu Wan 0004, Muyang Ye, Baosong Yang, Junyang Lin, Dayiheng Liu |
EMNLP | 5 |
| 2025 | Latent Space Chain-of-Embedding Enables Output-free LLM Self-EvaluationabstractLLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability.
In this research track, we propose the Chain-of-Embedding (CoE) in the latent space to enable LLMs to perform output-free self-evaluation. CoE consists of all progressive hidden states produced during the inference time, which can be treated as the latent thinking path of LLMs. We find that when LLMs respond correctly and incorrectly, their CoE features differ, these discrepancies assist us in estimating LLM response correctness. Experiments in four diverse domains and seven LLMs fully demonstrate the effectiveness of our method. Meanwhile, its label-free design intent without any training and millisecond-level computational cost ensure real-time feedback in large-scale scenarios.
More importantly, we provide interesting insights into LLM response correctness from the perspective of hidden state changes inside LLMs. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Rui Wang 0015 |
ICLR | 3 |
| 2025 | ConText: Driving In-context Learning for Text Removal and SegmentationabstractThis paper presents the first study on adapting the visual in-context learning (V-ICL) paradigm to optical character recognition tasks, specifically focusing on text removal and segmentation. Most existing V-ICL generalists employ a reasoning-as-reconstruction approach: they turn to using a straightforward image-label compositor as the prompt and query input, and then masking the query label to generate the desired output. This direct prompt confines the model to a challenging single-step reasoning process. To address this, we propose a task-chaining compositor in the form of image-removal-segmentation, providing an enhanced prompt that elicits reasoning with enriched intermediates. Additionally, we introduce context-aware aggregation, integrating the chained prompt pattern into the latent query representation, thereby strengthening the model’s in-context reasoning. We also consider the issue of visual heterogeneity, which complicates the selection of homogeneous demonstrations in text recognition. Accordingly, this is effectively addressed through a simple self-prompting strategy, preventing the model’s in-context learnability from devolving into specialist-like, context-free inference. Collectively, these insights culminate in our ConText model, which achieves new state-of-the-art across both in- and out-of-domain benchmarks. The code is available at https://github.com/Ferenas/ConText. Fei Zhang 0016, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Yanfeng Wang 0001, Ya Zhang 0002 |
ICML | 3 |
| 2025 | Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingabstractTest-time scaling enhances large language model performance by allocating additional compute resources during decoding. Best-of-$N$ (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. However, its cost–performance trade-off is still underexplored. Two main challenges limit the efficiency of BoN sampling:
(1) Generating $N$ full samples consumes substantial GPU memory, reducing inference capacity under limited resources.
(2) Reward models add extra memory and latency overhead, and training strong reward models introduces potential training data costs.
Although some studies have explored efficiency improvements, none have addressed both challenges at once.
To address this gap, we propose **Self-Truncation Best-of-$N$ (ST-BoN)**, a decoding method that avoids fully generating all $N$ samples and eliminates the need for reward models. It leverages early sampling consistency in the model’s internal states to identify the most promising path and truncate suboptimal ones.
In terms of cost, ST-BoN reduces dynamic GPU memory usage by over 80% and inference latency by 50%.
In terms of cost–performance trade-off, ST-BoN achieves the same performance as Full-BoN while saving computational cost by 70%–80%, and under the same cost, it can improve accuracy by 3–4 points. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Zhuosheng Zhang 0001, Fei Huang 0002, Rui Wang 0015 |
NeurIPS | 4 |
| 2025 | PolyMath: Evaluating Mathematical Reasoning in Multilingual ContextsabstractIn this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs.We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level.From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning:(1) Reasoning performance varies widely across languages for current LLMs;(2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance;(3) The thinking length differs significantly by language for current LLMs.Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs. Yiming Wang 0011, Pei Zhang 0011, Jialong Tang, Baosong Yang, Rui Wang 0015, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang 0002, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
NeurIPS | 5 |
| 2024 | MoNMT: Modularly Leveraging Monolingual and Bilingual Knowledge for Neural Machine TranslationabstractThe effective use of monolingual and bilingual knowledge represents a critical challenge within the neural machine translation (NMT) community. In this paper, we propose a modular strategy that facilitates the cooperation of these two types of knowledge in translation tasks, while avoiding the issue of catastrophic forgetting and exhibiting superior model generalization and robustness. Our model is comprised of three functionally independent modules: an encoding module, a decoding module, and a transferring module. The former two acquire large-scale monolingual knowledge via self-supervised learning, while the latter is trained on parallel data and responsible for transferring latent features between the encoding and decoding modules. Extensive experiments in multi-domain translation tasks indicate our model yields remarkable performance, with up to 7 BLEU improvements in out-of-domain tests over the conventional pretrain-and-finetune approach. Our codes are available at https://github.com/NLP2CT/MoNMT. Jianhui Pang, Baosong Yang, Derek F. Wong, Dayiheng Liu, Xiangpeng Wei, Lidia S. Chao |
LREC/COLING | 2 |
| 2024 | Embedding Trajectory for Out-of-Distribution Detection in Mathematical ReasoningabstractReal-world data deviating from the independent and identically distributed (\textit{i.i.d.}) assumption of in-distribution training data poses security threats to deep networks, thus advancing out-of-distribution (OOD) detection algorithms. Detection methods in generative language models (GLMs) mainly focus on uncertainty estimation and embedding distance measurement, with the latter proven to be most effective in traditional linguistic tasks like summarization and translation. However, another complex generative scenario mathematical reasoning poses significant challenges to embedding-based methods due to its high-density feature of output spaces, but this feature causes larger discrepancies in the embedding shift trajectory between different samples in latent spaces. Hence, we propose a trajectory-based method TV score, which uses trajectory volatility for OOD detection in mathematical reasoning. Experiments show that our method outperforms all traditional algorithms on GLMs under mathematical reasoning scenarios and can be extended to more applications with high-density features in output spaces, such as multiple-choice questions. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Zhuosheng Zhang 0001, Rui Wang 0015 |
NeurIPS | 3 |
| 2024 | Rethinking the Exploitation of Monolingual Data for Low-Resource Neural Machine TranslationabstractAbstract The utilization of monolingual data has been shown to be a promising strategy for addressing low-resource machine translation problems. Previous studies have demonstrated the effectiveness of techniques such as back-translation and self-supervised objectives, including masked language modeling, causal language modeling, and denoise autoencoding, in improving the performance of machine translation models. However, the manner in which these methods contribute to the success of machine translation tasks and how they can be effectively combined remains an under-researched area. In this study, we carry out a systematic investigation of the effects of these techniques on linguistic properties through the use of probing tasks, including source language comprehension, bilingual word alignment, and translation fluency. We further evaluate the impact of pre-training, back-translation, and multi-task learning on bitexts of varying sizes. Our findings inform the design of more effective pipelines for leveraging monolingual data in extremely low-resource and low-resource machine translation tasks. Experiment results show consistent performance gains in seven translation directions, which provide further support for our conclusions and understanding of the role of monolingual data in machine translation. Jianhui Pang, Baosong Yang, Derek F. Wong, Yu Wan 0004, Dayiheng Liu, Lidia S. Chao |
Comput. Linguistics | 2 |
| 2023 | Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text GenerationabstractKexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Mingfeng Xue, Boxing Chen, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Mingfeng Xue, Boxing Chen |
ACL (1) | 4 |
| 2023 | Bridging the Domain Gaps in Context Representations for k-Nearest Neighbor Neural Machine TranslationabstractZhiwei Cao, Baosong Yang, Huan Lin, Suhang Wu, Xiangpeng Wei, Dayiheng Liu, Jun Xie, Min Zhang, Jinsong Su. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Baosong Yang, Suhang Wu, Xiangpeng Wei, Dayiheng Liu, Min Zhang 0005, Jinsong Su |
ACL (1) | 2 |
| 2023 | Fantastic Expressions and Where to Find Them: Chinese Simile Generation with Multiple ConstraintsabstractKexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu |
ACL (1) | 4 |
| 2023 | MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksabstractMixture-of-Experts (MoE) based sparse architectures can significantly increase model capacity with sublinear computational overhead, which are hence widely used in massively multilingual neural machine translation (MNMT).However, they are prone to overfitting on lowresource language translation.In this paper, we propose a modularized MNMT framework that is able to flexibly assemble dense and MoEbased sparse modules to achieve the best of both worlds.The training strategy of the modularized MNMT framework consists of three stages: (1) Pre-training basic MNMT models with different training objectives or model structures, (2) Initializing modules of the framework with pre-trained couterparts (e.g., encoder, decoder and embedding layers) from the basic models and (3) Fine-tuning the modularized MNMT framework to fit modules from different models together.We pre-train three basic MNMT models from scratch: a dense model, an MoE-based sparse model and a new MoE model, termed as MoE-LGR that explores multiple Language-Group-specifc Routers to incorporate language group knowledge into MNMT.The strengths of these pre-trained models are either on low-resource language translation, highresource language translation or zero-shot translation.Our modularized MNMT framework attempts to incorporate these advantages into a single model with reasonable initialization and fine-tuning.Experiments on widely-used benchmark datasets demonstrate that the proposed modularized MNMT framwork substantially outperforms both MoE and dense models on high-and low-resource language translation as well as zero-shot translation.Our framework facilitates the combination of different methods with their own strengths and recycling off-the-shelf models for multilingual neural machine translation.Codes are available at https://github.com/lishangjie1/MMNMT. Shangjie Li, Xiangpeng Wei, Shaolin Zhu, Baosong Yang, Deyi Xiong |
EMNLP | 5 |
| 2023 | Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationabstractMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu, Jian Lan, Mei Li, Baosong Yang, Jun Xie, Yidan Zhang, Dezhong Peng, Jiancheng Lv. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Mingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu 0001, Baosong Yang, Yidan Zhang 0004, Dezhong Peng, Jiancheng Lv 0001 |
EMNLP | 7 |
| 2023 | EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningabstractExpressing universal semantics common to all languages is helpful to understand the meanings of complex and culture-specific sentences. The research theme underlying this scenario focuses on learning universal representations across languages with the usage of massive parallel corpora. However, due to the sparsity and scarcity of parallel data, there is still a big challenge in learning authentic ``universals'' for any two languages. In this paper, we propose Emma-X: an EM-like Multilingual pre-training Algorithm, to learn Cross-lingual universals with the aid of excessive multilingual non-parallel data. Emma-X unifies the cross-lingual representation learning task and an extra semantic relation prediction task within an EM framework. Both the extra semantic classifier and the cross-lingual sentence encoder approximate the semantic relation of two sentences, and supervise each other until convergence. To evaluate Emma-X, we conduct experiments on xrete, a newly introduced benchmark containing 12 widely studied cross-lingual tasks that fully depend on sentence-level representations. Results reveal that Emma-X achieves state-of-the-art performance. Further geometric analysis of the built representation space with three requirements demonstrates the superiority of Emma-X over advanced models. Ping Guo 0002, Xiangpeng Wei, Yue Hu 0002, Baosong Yang, Dayiheng Liu, Fei Huang 0002 |
NeurIPS | 4 |
| 2023 | Imitation Attacks Can Steal More Than You Think from Machine Translation Systems
Tianxiang Hu, Pei Zhang 0011, Baosong Yang, Rui Wang 0015 |
NLPCC (1) | 3 |
| 2023 | From statistical methods to deep learning, automatic keyphrase prediction: A survey
Binbin Xie, Jia Song 0003, Liangying Shao, Suhang Wu, Xiangpeng Wei, Baosong Yang, Jinsong Su |
Inf. Process. Manag. | 6 |
| 2023 | Towards Energy-Preserving Natural Language Understanding With Spiking Neural NetworksabstractArtificial neural networks have shown promising results in a variety of natural language understanding (NLU) tasks. Despite their successes, conventional neural-based NLU models are criticized for high energy consumption, making them laborious to be widely applied in low-power electronics, such as smartphones and intelligent terminals. In this paper, we introduce a potential direction to alleviate this bottleneck by proposing a spiking encoder. The core of our model is bi-directional spiking neural network (SNN) which transforms numeric values into discrete spiking signals and replaces massive multiplications with much cheaper additive operations. We examine our model on sentiment classification and machine translation tasks. Experimental results reveal that our model achieves comparable classification and translation accuracy to advancedTransformerbaseline, whereas significantly reduces the required computational energy to 0.82%. Rong Xiao 0001, Yu Wan 0004, Baosong Yang, Haibo Zhang 0013, Huajin Tang, Derek F. Wong, Boxing Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | KGR4: Retrieval, Retrospect, Refine and Rethink for Commonsense GenerationabstractGenerative commonsense reasoning requires machines to generate sentences describing an everyday scenario given several concepts, which has attracted much attention recently. However, existing models cannot perform as well as humans, since sentences they produce are often implausible and grammatically incorrect. In this paper, inspired by the process of humans creating sentences, we propose a novel Knowledge-enhanced Commonsense Generation framework, termed KGR4, consisting of four stages: Retrieval, Retrospect, Refine, Rethink. Under this framework, we first perform retrieval to search for relevant sentences from external corpus as the prototypes. Then, we train the generator that either edits or copies these prototypes to generate candidate sentences, of which potential errors will be fixed by an autoencoder-based refiner. Finally, we select the output sentence from candidate sentences produced by generators with different hyper-parameters. Experimental results and in-depth analysis on the CommonGen benchmark strongly demonstrate the effectiveness of our framework. Particularly, KGR4 obtains 33.56 SPICE in the official leaderboard, outperforming the previously-reported best result by 2.49 SPICE and achieving state-of-the-art performance. We release the code at https://github.com/DeepLearnXMU/KGR-4. Xin Liu 0066, Dayiheng Liu, Baosong Yang, Haibo Zhang 0013, Junwei Ding, Wenqing Yao, Weihua Luo, Jinsong Su |
AAAI | 3 |
| 2022 | Frequency-Aware Contrastive Learning for Neural Machine TranslationabstractLow-frequency word prediction remains a challenge in modern neural machine translation (NMT) systems. Recent adaptive training methods promote the output of infrequent words by emphasizing their weights in the overall training objectives. Despite the improved recall of low-frequency words, their prediction precision is unexpectedly hindered by the adaptive objectives. Inspired by the observation that low-frequency words form a more compact embedding space, we tackle this challenge from a representation learning perspective. Specifically, we propose a frequency-aware token-level contrastive learning method, in which the hidden state of each decoding step is pushed away from the counterparts of other target words, in a soft contrastive way based on the corresponding word frequencies. We conduct experiments on widely used NIST Chinese-English and WMT14 English-German translation tasks. Empirical results show that our proposed methods can not only significantly improve the translation quality but also enhance lexical diversity and optimize word representation space. Further investigation reveals that, comparing with related adaptive training strategies, the superiority of our method on low-frequency word prediction lies in the robustness of token-level recall across different frequencies without sacrificing precision. Tong Zhang 0001, Wei Ye 0004, Baosong Yang, Long Zhang 0012, Xingzhang Ren, Dayiheng Liu, Jinan Sun, Shikun Zhang, Haibo Zhang 0013 |
AAAI | 3 |
| 2022 | UniTE: Unified Translation EvaluationabstractYu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek Wong, Lidia Chao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yu Wan 0004, Dayiheng Liu, Baosong Yang, Haibo Zhang 0013, Boxing Chen, Derek F. Wong, Lidia S. Chao |
ACL (1) | 3 |
| 2022 | Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?abstractNeural machine translation (NMT) is often criticized for failures that happen without awareness.The lack of competency awareness makes NMT untrustworthy.This is in sharp contrast to human translators who give feedback or conduct further investigations whenever they are in doubt about predictions.To fill this gap, we propose a novel competency-aware NMT by extending conventional NMT with a selfestimator, offering abilities to translate a source sentence and estimate its competency.The selfestimator encodes the information of the decoding procedure and then examines whether it can reconstruct the original semantics of the source sentence.Experimental results on four translation tasks demonstrate that the proposed method not only carries out translation tasks intact but also delivers outstanding performance on quality estimation.Without depending on any reference or annotated data typically required by state-of-the-art metric and quality estimation methods, our model yields an even higher correlation with human quality judgments than a variety of aforementioned methods, such as BLEURT, COMET, and BERTScore.Quantitative and qualitative analyses show better robustness of competency awareness in our model.1 Pei Zhang 0011, Baosong Yang, Dayiheng Liu, Kai Fan 0002, Luo Si |
EMNLP | 2 |
| 2022 | WR-One2Set: Towards Well-Calibrated Keyphrase GenerationabstractKeyphrase generation aims to automatically generate short phrases summarizing an input document.The recently emerged ONE2SET paradigm (Ye et al., 2021) generates keyphrases as a set and has achieved competitive performance.Nevertheless, we observe serious calibration errors outputted by ONE2SET, especially in the over-estimation of ∅ token (means "no corresponding keyphrase").In this paper, we deeply analyze this limitation and identify two main reasons behind: 1) the parallel generation has to introduce excessive ∅ as padding tokens into training instances; and 2) the training mechanism assigning target to each slot is unstable and further aggravates the ∅ token over-estimation.To make the model wellcalibrated, we propose WR-ONE2SET which extends ONE2SET with an adaptive instancelevel cost Weighting strategy and a target Reassignment mechanism.The former dynamically penalizes the over-estimated slots for different instances thus smoothing the uneven training distribution.The latter refines the original inappropriate assignment and reduces the supervisory signals of over-estimated slots.Experimental results on commonly-used datasets demonstrate the effectiveness and generality of our proposed paradigm. Binbin Xie, Xiangpeng Wei, Baosong Yang, Xiaoli Wang 0002, Min Zhang 0005, Jinsong Su |
EMNLP | 3 |
| 2022 | Should We Rely on Entity Mentions for Relation Extraction? Debiasing Relation Extraction with Counterfactual AnalysisabstractYiwei Wang, Muhao Chen, Wenxuan Zhou, Yujun Cai, Yuxuan Liang, Dayiheng Liu, Baosong Yang, Juncheng Liu, Bryan Hooi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yiwei Wang 0001, Muhao Chen 0001, Wenxuan Zhou 0002, Yujun Cai, Yuxuan Liang 0002, Dayiheng Liu, Baosong Yang, Bryan Hooi |
NAACL-HLT | 7 |
| 2022 | Cross-Lingual Product Retrieval in E-Commerce Search
Wenya Zhu, Xiaoyu Lv, Baosong Yang, Xu Yong, Linlong Xu, Yinfu Feng, Haibo Zhang 0013, Qing Da, Anxiang Zeng, Ronghua Chen |
PAKDD (2) | 3 |
| 2022 | Effective Approaches to Neural Query Language IdentificationabstractAbstract Query language identification (Q-LID) plays a crucial role in a cross-lingual search engine. There exist two main challenges in Q-LID: (1) insufficient contextual information in queries for disambiguation; and (2) the lack of query-style training examples for low-resource languages. In this article, we propose a neural Q-LID model by alleviating the above problems from both model architecture and data augmentation perspectives. Concretely, we build our model upon the advanced Transformer model. In order to enhance the discrimination of queries, a variety of external features (e.g., character, word, as well as script) are fed into the model and fused by a multi-scale attention mechanism. Moreover, to remedy the low resource challenge in this task, a novel machine translation–based strategy is proposed to automatically generate synthetic query-style data for low-resource languages. We contribute the first Q-LID test set called QID-21, which consists of search queries in 21 languages. Experimental results reveal that our model yields better classification accuracy than strong baselines and existing LID systems on both query and traditional LID tasks.1 Xingzhang Ren, Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Xiaoyu Lv |
Comput. Linguistics | 2 |
| 2022 | Challenges of Neural Machine Translation for Short TextsabstractAbstract Short texts (STs) present in a variety of scenarios, including query, dialog, and entity names. Most of the exciting studies in neural machine translation (NMT) are focused on tackling open problems concerning long sentences rather than short ones. The intuition behind is that, with respect to human learning and processing, short sequences are generally regarded as easy examples. In this article, we first dispel this speculation via conducting preliminary experiments, showing that the conventional state-of-the-art NMT approach, namely, Transformer (Vaswani et al. 2017), still suffers from over-translation and mistranslation errors over STs. After empirically investigating the rationale behind this, we summarize two challenges in NMT for STs associated with translation error types above, respectively: (1) the imbalanced length distribution in training set intensifies model inference calibration over STs, leading to more over-translation cases on STs; and (2) the lack of contextual information forces NMT to have higher data uncertainty on short sentences, and thus NMT model is troubled by considerable mistranslation errors. Some existing approaches, like balancing data distribution for training (e.g., data upsampling) and complementing contextual information (e.g., introducing translation memory) can alleviate the translation issues in NMT for STs. We encourage researchers to investigate other challenges in NMT for STs, thus reducing ST translation errors and enhancing translation quality. Yu Wan 0004, Baosong Yang, Derek F. Wong, Lidia S. Chao, Haibo Zhang 0013, Boxing Chen |
Comput. Linguistics | 2 |
| 2022 | Multi-view self-attention networks
Mingzhou Xu, Baosong Yang, Derek F. Wong, Lidia S. Chao |
Knowl. Based Syst. | 2 |
| 2021 | Towards User-Driven Neural Machine TranslationabstractHuan Lin, Liang Yao, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Degen Huang, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Weihua Luo, Degen Huang, Jinsong Su |
ACL/IJCNLP (1) | 3 |
| 2021 | Bridging Subword Gaps in Pretrain-Finetune Paradigm for Natural Language GenerationabstractXin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xin Liu 0066, Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Weihua Luo, Min Zhang 0005, Jinsong Su |
ACL/IJCNLP (1) | 2 |
| 2021 | Multi-Hop Transformer for Document-Level Machine TranslationabstractLong Zhang, Tong Zhang, Haibo Zhang, Baosong Yang, Wei Ye, Shikun Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Long Zhang 0012, Tong Zhang 0001, Haibo Zhang 0013, Baosong Yang, Wei Ye 0004, Shikun Zhang |
NAACL-HLT | 4 |
| 2021 | Context-aware Self-Attention Networks for Natural Language Processing
Baosong Yang, Longyue Wang, Derek F. Wong, Shuming Shi 0001, Zhaopeng Tu |
Neurocomputing | 1 |
| 2020 | Neuron Interaction Based Representation Composition for Neural Machine TranslationabstractRecent NLP studies reveal that substantial linguistic information can be attributed to single neurons, i.e., individual dimensions of the representation vectors. We hypothesize that modeling strong interactions among neurons helps to better capture complex information by composing the linguistic properties embedded in individual neurons. Starting from this intuition, we propose a novel approach to compose representations learned by different components in neural machine translation (e.g., multi-layer networks or multi-head attention), based on modeling strong interactions among neurons in the representation vectors. Specifically, we leverage bilinear pooling to model pairwise multiplicative interactions among individual neurons, and a low-rank approximation to make the model computationally feasible. We further propose extended bilinear pooling to incorporate first-order representations. Experiments on WMT14 English⇒German and English⇒French translation tasks show that our model consistently improves performances over the SOTA Transformer baseline. Further analyses demonstrate that our approach indeed captures more syntactic and semantic information as expected. Jian Li 0054, Xing Wang 0007, Baosong Yang, Shuming Shi 0001, Michael R. Lyu, Zhaopeng Tu |
AAAI | 3 |
| 2020 | Unsupervised Neural Dialect Translation with Commonality and Diversity ModelingabstractAs a special machine translation task, dialect translation has two main characteristics: 1) lack of parallel training corpus; and 2) possessing similar grammar between two sides of the translation. In this paper, we investigate how to exploit the commonality and diversity between dialects thus to build unsupervised translation models merely accessing to monolingual data. Specifically, we leverage pivot-private embedding, layer coordination, as well as parameter sharing to sufficiently model commonality and diversity among source and target, ranging from lexical, through syntactic, to semantic levels. In order to examine the effectiveness of the proposed models, we collect 20 million monolingual corpus for each of Mandarin and Cantonese, which are official language and the most widely used dialect in China. Experimental results reveal that our methods outperform rule-based simplified and traditional Chinese conversion and conventional unsupervised translation models over 12 BLEU scores. Yu Wan 0004, Baosong Yang, Derek F. Wong, Lidia S. Chao, Haihua Du, Ben C. H. Ao |
AAAI | 2 |
| 2020 | Uncertainty-Aware Curriculum Learning for Neural Machine TranslationabstractNeural machine translation (NMT) has proven to be facilitated by curriculum learning which presents examples in an easy-to-hard order at different training stages. The keys lie in the assessment of data difficulty and model competence. We propose uncertainty-aware curriculum learning, which is motivated by the intuition that: 1) the higher the uncertainty in a translation pair, the more complex and rarer the information it contains; and 2) the end of the decline in model uncertainty indicates the completeness of current training stage. Specifically, we serve cross-entropy of an example as its data difficulty and exploit the variance of distributions over the weights of the network to present the model uncertainty. Extensive experiments on various translation tasks reveal that our approach outperforms the strong baseline and related methods on both translation quality and convergence speed. Quantitative analyses reveal that the proposed strategy offers NMT the ability to automatically govern its learning schedule. Yikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan 0004, Lidia S. Chao |
ACL | 2 |
| 2020 | Domain Transfer based Data Augmentation for Neural Query TranslationabstractQuery translation (QT) serves as a critical factor in successful cross-lingual information retrieval (CLIR).Due to the lack of parallel query samples, neural-based QT models are usually optimized with synthetic data which are derived from large-scale monolingual queries.Nevertheless, such kind of pseudo corpus is mostly produced by a general-domain translation model, making it be insufficient to guide the learning of QT model.In this paper, we extend the data augmentation with a domain transfer procedure, thus to revise synthetic candidates to search-aware examples.Specifically, the domain transfer model is built upon advanced Transformer, in which layer coordination and mixed attention are exploited to speed up the refining process and leverage parameters from a pre-trained cross-lingual language model.In order to examine the effectiveness of the proposed method, we collected French-to-English and Spanish-to-English QT test sets, each of which consists of 10,000 parallel query pairs with careful manual-checking.Qualitative and quantitative analyses reveal that our model significantly outperforms strong baselines and the related domain transfer methods on both translation quality and retrieval accuracy.1 Baosong Yang, Haibo Zhang 0013, Boxing Chen, Weihua Luo |
COLING | 2 |
| 2020 | Self-Paced Learning for Neural Machine TranslationabstractRecent studies have proven that the training of neural machine translation (NMT) can be facilitated by mimicking the learning process of humans.Nevertheless, achievements of such kind of curriculum learning rely on the quality of artificial schedule drawn up with the handcrafted features, e.g.sentence length or word rarity.We ameliorate this procedure with a more flexible manner by proposing self-paced learning, where NMT model is allowed to 1) automatically quantify the learning confidence over training examples; and 2) flexibly govern its learning via regulating the loss in each iteration step.Experimental results over multiple translation tasks demonstrate that the proposed model yields better performance than strong baselines and those models trained with human-designed curricula on both translation quality and convergence speed. 1 Yu Wan 0004, Baosong Yang, Derek F. Wong, Yikai Zhou, Lidia S. Chao, Haibo Zhang 0013, Boxing Chen |
EMNLP (1) | 2 |
| 2020 | Improving tree-based neural machine translation with dynamic lexicalized dependency encoding
Baosong Yang, Derek F. Wong, Lidia S. Chao, Min Zhang 0005 |
Knowl. Based Syst. | 1 |
| 2019 | Context-Aware Self-Attention NetworksabstractSelf-attention model has shown its flexibility in parallel computation and the effectiveness on modeling both long- and short-term dependencies. However, it calculates the dependencies between representations without considering the contextual information, which has proven useful for modeling dependencies among neural representations in various natural language tasks. In this work, we focus on improving self-attention networks through capturing the richness of context. To maintain the simplicity and flexibility of the self-attention networks, we propose to contextualize the transformations of the query and key layers, which are used to calculate the relevance between elements. Specifically, we leverage the internal representations that embed both global and deep contexts, thus avoid relying on external resources. Experimental results on WMT14 English⇒German and WMT17 Chinese⇒English translation tasks demonstrate the effectiveness and universality of the proposed methods. Furthermore, we conducted extensive analyses to quantify how the context vectors participate in the self-attention model. Baosong Yang, Jian Li 0054, Derek F. Wong, Lidia S. Chao, Xing Wang 0007, Zhaopeng Tu |
AAAI | 1 |
| 2019 | Leveraging Local and Global Patterns for Self-Attention NetworksabstractSelf-attention networks have received increasing research attention.By default, the hidden states of each word are hierarchically calculated by attending to all words in the sentence, which assembles global information.However, several studies pointed out that taking all signals into account may lead to overlooking neighboring information (e.g.phrase pattern).To address this argument, we propose a hybrid attention mechanism to dynamically leverage both of the local and global information.Specifically, our approach uses a gating scalar for integrating both sources of the information, which is also convenient for quantifying their contributions.Experiments on various neural machine translation tasks demonstrate the effectiveness of the proposed method.The extensive analyses verify that the two types of contexts are complementary to each other, and our method gives highly effective improvements in their integration. Mingzhou Xu, Derek F. Wong, Baosong Yang, Yue Zhang 0004, Lidia S. Chao |
ACL (1) | 3 |
| 2019 | Assessing the Ability of Self-Attention Networks to Learn Word OrderabstractSelf-attention networks (SAN) have attracted a lot of interests due to their high parallelization and strong performance on a variety of NLP tasks, e.g. machine translation.Due to the lack of recurrence structure such as recurrent neural networks (RNN), SAN is ascribed to be weak at learning positional information of words for sequence modeling.However, neither this speculation has been empirically confirmed, nor explanations for their strong performances on machine translation tasks when "lacking positional information" have been explored.To this end, we propose a novel word reordering detection task to quantify how well the word order information learned by SAN and RNN.Specifically, we randomly move one word to another position, and examine whether a trained model can detect both the original and inserted positions.Experimental results reveal that: 1) SAN trained on word reordering detection indeed has difficulty learning the positional information even with the position embedding; and 2) SAN trained on machine translation learns better positional information than its RNN counterpart, in which position embedding plays a critical role.Although recurrence structure make the model more universally-effective on learning word order, learning objectives matter more in the downstream tasks such as machine translation. Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, Zhaopeng Tu |
ACL (1) | 1 |
| 2018 | Multi-Head Attention with Disagreement RegularizationabstractMulti-head attention is appealing for the ability to jointly attend to information from different representation subspaces at different positions.In this work, we introduce a disagreement regularization to explicitly encourage the diversity among multiple attention heads.Specifically, we propose three types of disagreement regularization, which respectively encourage the subspace, the attended positions, and the output representation associated with each attention head to be different from other heads.Experimental results on widely-used WMT14 English⇒German and WMT17 Chinese⇒English translation tasks demonstrate the effectiveness and universality of the proposed approach.* Zhaopeng Tu is the corresponding author of the paper.This work was mainly conducted when Jian Li and Baosong Yang were interning at Tencent AI Lab. Jian Li 0054, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, Tong Zhang 0001 |
EMNLP | 3 |
| 2018 | Modeling Localness for Self-Attention NetworksabstractSelf-attention networks have proven to be of profound value for its strength of capturing global dependencies.In this work, we propose to model localness for self-attention networks, which enhances the ability of capturing useful local context.We cast localness modeling as a learnable Gaussian bias, which indicates the central and scope of the local region to be paid more attention.The bias is then incorporated into the original attention distribution to form a revised distribution.To maintain the strength of capturing long distance dependencies and enhance the ability of capturing shortrange dependencies, we only apply localness modeling to lower layers of self-attention networks.Quantitative and qualitative analyses on Chinese⇒English and English⇒German translation tasks demonstrate the effectiveness and universality of the proposed approach. Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, Tong Zhang 0001 |
EMNLP | 1 |
| 2017 | Towards Bidirectional Hierarchical Representations for Attention-based Neural Machine TranslationabstractThis paper proposes a hierarchical attentional neural translation model which focuses on enhancing source-side hierarchical representations by covering both local and global semantic information using a bidirectional tree-based encoder.To maximize the predictive likelihood of target words, a weighted variant of an attention mechanism is used to balance the attentive information between lexical and phrase vectors.Using a tree-based rare word encoding, the proposed model is extended to sub-word level to alleviate the out-of-vocabulary (OOV) problem.Empirical results reveal that the proposed model significantly outperforms sequence-to-sequence attention-based and tree-based neural translation models in English-Chinese translation tasks. Baosong Yang, Derek F. Wong, Tong Xiao 0001, Lidia S. Chao |
EMNLP | 1 |