Xiangyu Duan

dblp:09/5412 · DBLP profile ↗
← Back
29ranked-venue papers
10as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
abstract
Recent advancements in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increased computational resources. In this work, we identify two key challenges contributing to this inefficiency: \textit{over-exploration} due to redundant states with semantically equivalent content, and \textit{under-exploration} caused by high variance in verifier scoring leading to frequent trajectory switching. To address these issues, we propose FETCH – an e{\bf f}fici{\bf e}nt {\bf t}ree sear{\bf ch} framework, which is a flexible, plug-and-play system compatible with various tree search algorithms.Our framework mitigates over-exploration by merging semantically similar states using agglomerative clustering of text embeddings obtained from a fine-tuned SimCSE model. To tackle under-exploration, we enhance verifiers by incorporating temporal difference learning with adjusted \lambda-returns during training to reduce variance, and employing a verifier ensemble to aggregate scores during inference. Experiments on GSM8K, GSM-Plus, and MATH datasets demonstrate that our methods significantly improve reasoning accuracy and computational efficiency across four different tree search algorithms, paving the way for more practical applications of LLM-based reasoning. The code is available at https://github.com/DeepLearnXMU/Fetch.
Ante Wang, Linfeng Song, Dian Yu 0001, Haitao Mi, Xiangyu Duan, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
ACL (1)6
2025 Basic Reading Distillation
abstract
Large language models (LLMs) have demonstrated remarkable abilities in various natural language processing areas, but they demand high computation resources which limits their deployment in real-world.Distillation is one technique to solve this problem through either knowledge distillation or task distillation.Both distillation approaches train small models to imitate specific features of LLMs, but they all neglect basic reading education for small models on generic texts that are unrelated to downstream tasks.In this paper, we propose basic reading distillation (BRD) which educates a small model to imitate LLMs basic reading behaviors, such as named entity recognition, question raising and answering, on each sentence.After such basic education, we apply the small model on various tasks including language inference benchmarks and BIG-bench tasks.It shows that the small model can outperform or perform comparable to over 20x bigger LLMs.Analysis reveals that BRD effectively influences the probability distribution of the small model, and has orthogonality to either knowledge distillation or task distillation.
Sirui Miao, Xiangyu Duan, Hao Yang 0006, Min Zhang 0005
ACL (1)3
2024 Multimodal Cross-lingual Phrase Retrieval
abstract
Cross-lingual phrase retrieval aims to retrieve parallel phrases among languages. Current approaches only deals with textual modality. There lacks multimodal data resources and explorations for multimodal cross-lingual phrase retrieval (MXPR). In this paper, we create the first MXPR data resource and propose a novel approach for MXPR to explore the effectiveness of multi-modality. The MXPR data resource is built by marrying the benchmark dataset for textual cross-lingual phrase retrieval with Wikimedia Commons, which is a media store containing tremendous texts and related images. In the built resource, the phrase pairs of the textual benchmark dataset are equipped with their related images. Based on this novel data resource, we introduce a strategy to bridge the gap between different modalities by multimodal relation generation with a large multimodal pre-trained model and consistency training. Experiments on benchmarked dataset covering eight language pairs show that our MXPR approach, which deals with multimodal phrases, performs significantly better than pure textual cross-lingual phrase retrieval.
Chuanqi Dong, Xiangyu Duan
LREC/COLING3
2024 Submodular-based In-context Example Selection for LLMs-based Machine Translation
abstract
Large Language Models (LLMs) have demonstrated impressive performances across various NLP tasks with just a few prompts via in-context learning. Previous studies have emphasized the pivotal role of well-chosen examples in in-context learning, as opposed to randomly selected instances that exhibits unstable results.A successful example selection scheme depends on multiple factors, while in the context of LLMs-based machine translation, the common selection algorithms only consider the single factor, i.e., the similarity between the example source sentence and the input sentence.In this paper, we introduce a novel approach to use multiple translational factors for in-context example selection by using monotone submodular function maximization.The factors include surface/semantic similarity between examples and inputs on both source and target sides, as well as the diversity within examples.Importantly, our framework mathematically guarantees the coordination between these factors, which are different and challenging to reconcile.Additionally, our research uncovers a previously unexamined dimension: unlike other NLP tasks, the translation part of an example is also crucial, a facet disregarded in prior studies.Experiments conducted on BLOOMZ-7.1B and LLAMA2-13B, demonstrate that our approach significantly outperforms random selection and robust single-factor baselines across various machine translation tasks.
Baijun Ji, Xiangyu Duan, Zhenyu Qiu, Junhui Li 0001, Hao Yang 0006, Min Zhang 0005
LREC/COLING2
2024 Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation
abstract
Multi-choice questions (MCQ) are a common method for assessing the world knowledge of large language models (LLMs), demonstrated by benchmarks such as MMLU and C-Eval. However, recent findings indicate that even top-tier LLMs, such as ChatGPT and GPT4, might display inconsistencies when faced with slightly varied inputs. This raises concerns about the credibility of MCQ-based evaluations. To address this issue, we introduced three knowledge-equivalent question variants: option position shuffle, option label replacement, and conversion to a True/False format. We rigorously tested a range of LLMs, varying in model size (from 6B to 70B) and types—pretrained language model (PLM), supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF). Our findings from MMLU and C-Eval revealed that accuracy for individual questions lacks robustness, particularly in smaller models (<30B) and PLMs. Consequently, we advocate that consistent accuracy may serve as a more reliable metric for evaluating and ranking LLMs.
Mingzhou Xu, Xiangyu Duan
LREC/COLING5
2024 Exploring and Improving Consistency in Large Language Models for Multiple-Choice Question Assessment
abstract
With the evolution of Large Language Models (LLMs), accurately evaluating their capabilities has become a critical focus. The Multi-choice Questions (MCQ) benchmark is widely adopted for its definitive answers and straightforward assessment approach. However, recent studies have found that there is inconsistency in the model, raising concerns about potential biases and their genuine comprehension abilities in MCQ contexts. To delve into and improve the consistency performance of LLMs in MCQ answering domain, this study first proposes new consistency metrics, including option position consistency and option symbol consistency. These metrics are designed to quantify and reveal the level of consistency in the model, thereby assessing the authenticity of its knowledge comprehension. Secondly, utilizing these metrics, we propose two novel improvement strategies: 1) An enhanced In-context Learning (ICL) prompt customization technique, which adaptively modifies prompts based on the model’s demonstrated capabilities, aligning the prompts with the model’s inherent abilities to filter out questions it deems consistently answerable; and 2) Consistency Supervised Fine-Tuning (CSFT), which enriches the training data set focused on consistency, followed by specialized fine-tuning to augment the model’s inherent capabilities. Our research is committed to exploring and improving consistency levels in models, with the goal of bolstering their integrity and reliability in the realm of MCQ answering.
Xiangyu Duan
IJCNN2
2022 Bilingual Terminology Extraction from Comparable E-Commerce Corpora
abstract
Bilingual terminologies are important machine translation resources in the field of e-commerce, which are usually either manually translated or automatically extracted from parallel data. The human translation is costly and e-commerce parallel corpora is very scarce. However, the comparable data in different languages in the same commodity field is abundant. In this paper, we propose a novel framework of extracting e-commercial bilingual terminologies from comparable data. Benefiting from the cross-lingual pre-training in e-commerce, our framework can make full use of the deep semantic relationship between source-side terminology and target-side sentence to extract corresponding target terminology. Experimental results on various language pairs show that our approaches achieve significantly better performance than various strong baselines.
Shuqin Gu, Xiangyu Duan
IJCNN4
2022 Random Concatenation: A Simple Data Augmentation Method for Neural Machine Translation
Nini Xiao, Huaao Zhang, Chang Jin, Xiangyu Duan
NLPCC (1)4
2021 Improving Context-Aware Neural Machine Translation with Source-side Monolingual Documents
abstract
Document context-aware machine translation remains challenging due to the lack of large-scale document parallel corpora. To make full use of source-side monolingual documents for context-aware NMT, we propose a Pre-training approach with Global Context (PGC). In particular, we first propose a novel self-supervised pre-training task, which contains two training objectives: (1) reconstructing the original sentence from a corrupted version; (2) generating a gap sentence from its left and right neighbouring sentences. Then we design a universal model for PGC which consists of a global context encoder, a sentence encoder and a decoder, with similar architecture to typical context-aware NMT models. We evaluate the effectiveness and generality of our pre-trained PGC model by adapting it to various downstream context-aware NMT models. Detailed experimentation on four different translation tasks demonstrates that our PGC approach significantly improves the translation performance of context-aware NMT. For example, based on the state-of-the-art SAN model, we achieve an averaged improvement of 1.85 BLEU scores and 1.59 Meteor scores on the four translation tasks.
Linqing Chen, Junhui Li 0001, Zhengxian Gong, Xiangyu Duan, Boxing Chen, Weihua Luo, Min Zhang 0005, Guodong Zhou 0001
IJCAI4
2020 Cross-Lingual Pre-Training Based Transfer for Zero-Shot Neural Machine Translation
abstract
Transfer learning between different language pairs has shown its effectiveness for Neural Machine Translation (NMT) in low-resource scenario. However, existing transfer methods involving a common target language are far from success in the extreme scenario of zero-shot translation, due to the language space mismatch problem between transferor (the parent model) and transferee (the child model) on the source side. To address this challenge, we propose an effective transfer learning approach based on cross-lingual pre-training. Our key idea is to make all source languages share the same feature space and thus enable a smooth transition for zero-shot translation. To this end, we introduce one monolingual pre-training method and two bilingual pre-training methods to obtain a universal encoder for different languages. Once the universal encoder is constructed, the parent model built on such encoder is trained with large-scale annotated data and then directly applied in zero-shot translation scenario. Experiments on two public datasets show that our approach significantly outperforms strong pivot-based baseline and various multilingual NMT approaches.
Baijun Ji, Zhirui Zhang, Xiangyu Duan, Min Zhang 0005, Boxing Chen, Weihua Luo
AAAI3
2020 Alignment-Enhanced Transformer for Constraining NMT with Pre-Specified Translations
abstract
We investigate the task of constraining NMT with pre-specified translations, which has practical significance for a number of research and industrial applications. Existing works impose pre-specified translations as lexical constraints during decoding, which are based on word alignments derived from target-to-source attention weights. However, multiple recent studies have found that word alignment derived from generic attention heads in the Transformer is unreliable. We address this problem by introducing a dedicated head in the multi-head Transformer architecture to capture external supervision signals. Results on five language pairs show that our method is highly effective in constraining NMT with pre-specified translations, consistently outperforming previous methods in translation quality.
Heng Yu 0006, Yue Zhang 0004, Zhongqiang Huang, Weihua Luo, Xiangyu Duan, Min Zhang 0005
AAAI7
2020 Bilingual Dictionary Based Neural Machine Translation without Using Parallel Sentences
abstract
International audience
Xiangyu Duan, Baijun Ji, Min Zhang 0005, Boxing Chen, Weihua Luo, Yue Zhang 0004
ACL1
2020 Token Drop mechanism for Neural Machine Translation
abstract
Neural machine translation with millions of parameters is vulnerable to unfamiliar inputs.We propose Token Drop to improve generalization and avoid overfitting for the NMT model.Similar to word dropout, whereas we replace dropped token with a special token instead of setting zero to words.We further introduce two self-supervised objectives: Replaced Token Detection and Dropped Token Prediction.Our method aims to force model generating target translation with less information, in this way the model can learn textual representation better.Experiments on Chinese-English and English-Romanian benchmark demonstrate the effectiveness of our approach and our model achieves significant improvements over a strong Transformer baseline 1 .
Huaao Zhang, Shigui Qiu, Xiangyu Duan, Min Zhang 0005
COLING3
2020 Layer-Wise De-Training and Re-Training for ConvS2S Machine Translation
abstract
The convolutional sequence-to-sequence (ConvS2S) machine translation system is one of the typical neural machine translation (NMT) systems. Training the ConvS2S model tends to get stuck in a local optimum in our pre-studies. To overcome this inferior behavior, we propose to de-train a trained ConvS2S model in a mild way and retrain to find a better solution globally. In particular, the trained parameters of one layer of the NMT network are abandoned by re-initialization while other layers’ parameters are kept at the same time to kick off re-optimization from a new start point and safeguard the new start point not too far from the previous optimum. This procedure is executed layer by layer until all layers of the ConvS2S model are explored. Experiments show that when compared to various measures for escaping from the local optimum, including initialization with random seeds, adding perturbations to the baseline parameters, and continuing training (con-training) with the baseline models, our method consistently improves the ConvS2S translation quality across various language pairs and achieves better performance.
Hongfei Yu, Xiaoqing Zhou, Xiangyu Duan, Min Zhang 0005
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2020 Towards Better Word Alignment in Transformer
abstract
While neural models based on the Transformer architecture achieve the State-of-the-Art translation performance, it is well known that the learned target-to-source attentions do not correlate well with word alignment. There is an increasing interest in inducing accurate word alignment in Transformer, due to its important role in practical applications such as dictionary-guided translation and interactive translation. In this article, we extend and improve the recent work on unsupervised learning of word alignment in Transformer on two dimensions: a) parameter initialization from a pre-trained cross-lingual language model to leverage large amounts of monolingual data for learning robust contextualized word representations, and b) regularization of the training objective to directly model characteristics of word alignments which results in favorable word alignments receiving more concentrated probabilities. Experiments on benchmark data sets of three language pairs show that the proposed methods can significantly reduce alignment error rate (AER) by at least 3.7 to 7.7 points on each language pair over two recent works on improving the Transformer's word alignment. Moreover, our methods can achieve better alignment results than GIZA++ on certain test sets.
Xiaoqing Zhou, Heng Yu 0006, Zhongqiang Huang, Yue Zhang 0004, Weihua Luo, Xiangyu Duan, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.7
2019 Zero-Shot Cross-Lingual Abstractive Sentence Summarization through Teaching Generation and Attention
abstract
Abstractive Sentence Summarization (AS-SUM) targets at grasping the core idea of the source sentence and presenting it as the summary.It is extensively studied using statistical models or neural models based on the large-scale monolingual source-summary parallel corpus.But there is no cross-lingual parallel corpus, whose source sentence language is different to the summary language, to directly train a cross-lingual ASSUM system.We propose to solve this zero-shot problem by using resource-rich monolingual AS-SUM system to teach zero-shot cross-lingual ASSUM system on both summary word generation and attention.This teaching process is along with a back-translation process which simulates source-summary pairs.Experiments on cross-lingual ASSUM task show that our proposed method is significantly better than pipeline baselines and previous works, and greatly enhances the cross-lingual performances closer to the monolingual performances.We release the code and data at https://github.com/KelleyYin/ Cross-lingual-Summarization.
Xiangyu Duan, Mingming Yin, Min Zhang 0005, Boxing Chen, Weihua Luo
ACL (1)1
2019 Contrastive Attention Mechanism for Abstractive Sentence Summarization
abstract
Xiangyu Duan, Hongfei Yu, Mingming Yin, Min Zhang, Weihua Luo, Yue Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiangyu Duan, Hongfei Yu, Mingming Yin, Min Zhang 0005, Weihua Luo, Yue Zhang 0004
EMNLP/IJCNLP (1)1
2019 Question Generation Based Product Information
Kang Xiao, Xiabing Zhou, Xiangyu Duan, Min Zhang 0005
NLPCC (2)4
2016 Exploiting meta features for dependency parsing and part-of-speech tagging
Wenliang Chen, Min Zhang 0005, Yue Zhang 0004, Xiangyu Duan
Artif. Intell.4
2014 Synchronous Constituent Context Model for Inducing Bilingual Synchronous Structures
Xiangyu Duan, Min Zhang 0005, Qiaoming Zhu
COLING1
2014 Bayesian Constituent Context Model for Grammar Induction
abstract
Constituent Context Model (CCM) is an effective generative model for grammar induction, the aim of which is to induce hierarchical syntactic structure from natural text. The CCM simply defines the Multinomial distribution over constituents, which leads to a severe data sparse problem because long constituents are unlikely to appear in unseen data sets. This paper proposes a Bayesian method for constituent smoothing by defining two kinds of prior distributions over constituents: the Dirichlet prior and the Pitman-Yor Process prior. The Dirichlet prior functions as an additive smoothing method, and the PYP prior functions as a back-off smoothing method. Furthermore, a modified CCM is proposed to differentiate left constituents and right constituents in binary branching trees. Experiments show that both the proposed Bayesian smoothing method and the modified CCM are effective, and combining them attains or significantly improves the state-of-the-art performance of grammar induction evaluated on standard treebanks of various languages.
Min Zhang 0005, Xiangyu Duan, Wenliang Chen
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Smoothing for Bracketing Induction
Xiangyu Duan, Min Zhang 0005, Wenliang Chen
IJCAI1
2013 Improving Graph-Based Dependency Parsing Models With Dependency Language Models
abstract
For graph-based dependency parsing, how to enrich high-order features without increasing decoding complexity is a very challenging problem. To solve this problem, this paper presents an approach to representing high-order features for graph-based dependency parsing models using a dependency language model and beam search. Firstly, we use a baseline parser to parse a large-amount of unannotated data. Then we build the dependency language model (DLM) on the auto-parsed data. A set of new features is represented based on the DLM. Finally, we integrate the DLM-based features into the parsing model during decoding by beam search. We also utilize the features in bilingual text (bitext) parsing models. The main advantages of our approach are: 1) we utilize rich high-order features defined over a view of large scope and additional large raw corpus; 2) our approach does not increase the decoding complexity. We evaluate the proposed approach on the monotext and bitext parsing tasks. In the monotext parsing task, we conduct the experiments on Chinese and English data. The experimental results show that our new parser achieves the best accuracy on the Chinese data and comparable accuracy with the best known systems on the English data. In the bitext parsing task, we conduct the experiments on a Chinese-English bilingual data and our score is the best reported so far.
Min Zhang 0005, Wenliang Chen, Xiangyu Duan, Rong Zhang 0002
IEEE Trans. Speech Audio Process.3
2011 Joint Alignment and Artificial Data Generation: An Empirical Study of Pivot-based Machine Transliteration
Min Zhang 0005, Xiangyu Duan, Yunqing Xia, Haizhou Li 0001
IJCNLP2
2010 Pseudo-Word for Phrase-Based Machine Translation
Xiangyu Duan, Min Zhang 0005, Haizhou Li 0001
ACL1
2007 Ungreedy Methods for Chinese Deterministic Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
AAAI1
2007 Probabilistic Models for Action-Based Chinese Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
ECML1
2007 Probabilistic Parsing Action Models for Multi-Lingual Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
EMNLP-CoNLL1
2007 Word Sense Disambiguation through Sememe Labeling
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
IJCAI1