Yixuan Su

dblp:262/3282 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-1472-7791ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SARE: Sketch-Aware Random Erasing for Transformer-based 3D Shape Retrieval
Yixuan Su, Pengxue Wu, Jinglan Tian, Jingliang Peng
ICIC (20)1
2025 500xCompressor: Generalized Prompt Compression for Large Language Models
abstract
Prompt compression is important for large language models (LLMs) to increase inference speed, reduce costs, and improve user experience.However, current methods face challenges such as low compression ratios and potential training-test overlap during evaluation.To address these issues, we propose 500xCompressor, a method that compresses natural language contexts into a minimum of one special token and demonstrates strong generalization ability.The 500xCompressor introduces approximately 0.3% additional parameters and achieves compression ratios ranging from 6x to 500x, achieving 27-90% reduction in calculations and 55-83% memory savings when generating 100-400 tokens for new and reused prompts at 500x compression, while retaining 70-74% (F1) and 77-84% (Exact Match) of the LLM capabilities compared to using non-compressed prompts.It is designed to compress any text, answer various types of questions, and can be utilized by the original LLM without requiring fine-tuning.Initially, 500xCompressor was pretrained on the Arx-ivCorpus, followed by fine-tuning on the Arx-ivQA dataset, and subsequently evaluated on strictly unseen and cross-domain question answering (QA) datasets.This study shows that KV values outperform embeddings in preserving information at high compression ratios.The highly compressive nature of natural language prompts, even for detailed information, suggests potential for future applications and the development of a new LLM language.1
Zongqian Li, Yixuan Su, Nigel Collier
ACL (1)2
2025 To Code or Not To Code? Exploring Impact of Code in Pre-training
abstract
Including code in the pre-training data mixture, even for models not specifically designed for code, has become a common practice in LLMs pre-training. While there has been anecdotal consensus among practitioners that code data plays a vital role in general LLMs' performance, there is only limited work analyzing the precise impact of code on non-code tasks. In this work, we systematically investigate the impact of code data on general performance. We ask “what is the impact of code data used in pre-training on a large variety of downstream tasks beyond code generation”. We conduct extensive ablations and evaluate across a broad range of natural language reasoning tasks, world knowledge tasks, code benchmarks, and LLM-as-a-judge win-rates for models with sizes ranging from 470M to 2.8B parameters. Across settings, we find a consistent results that code is a critical building block for generalization far beyond coding tasks and improvements to code quality have an outsized impact across all tasks. In particular, compared to text-only pre-training, the addition of code results in up to relative increase of 8.2% in natural language (NL) reasoning, 4.2% in world knowledge, 6.6% improvement in generative win-rates, and a 12x boost in code performance respectively. Our work suggests investments in code quality and preserving code during pre-training have positive impacts.
Viraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot, Ivan Zhang, Acyr Locatelli, Marzieh Fadaee, Ahmet Üstün, Sara Hooker
ICLR2
2025 Prompt Compression for Large Language Models: A Survey
abstract
Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier
NAACL (Long Papers)3
2025 PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
abstract
Parameter-efficient fine-tuning (PEFT) methods have shown promise in adapting large language models, yet existing approaches exhibit counter-intuitive phenomena: integrating either matrix decomposition or mixture-of-experts (MoE) individually decreases performance across tasks, though decomposition improves results on specific domains despite reducing parameters, while MoE increases parameter count without corresponding decrease in training efficiency. Motivated by these observations and the modular nature of PT, we propose PT-MoE, a novel framework that integrates matrix decomposition with MoE routing for efficient PT. Evaluation results across 17 datasets demonstrate that PT-MoE achieves state-of-the-art performance in both question answering (QA) and mathematical problem solving tasks, improving F1 score by 1.49 points over PT and 2.13 points over LoRA in QA tasks, while improving mathematical accuracy by 10.75 points over PT and 0.44 points over LoRA, all while using 25% fewer parameters than LoRA. Our analysis reveals that while PT methods generally excel in QA tasks and LoRA-based methods in math datasets, the integration of matrix decomposition and MoE in PT-MoE yields complementary benefits: decomposition enables efficient parameter sharing across experts while MoE provides dynamic adaptation, collectively enabling PT-MoE to demonstrate cross-task consistency and generalization abilities. These findings, along with ablation studies on routing mechanisms and architectural components, provide insights for future PEFT methods.
Zongqian Li, Yixuan Su, Nigel Collier
NeurIPS2
2024 Exploring Dense Retrieval for Dialogue Response Selection
abstract
Recent progress in deep learning has continuously improved the accuracy of dialogue response selection. However, in real-world scenarios, the high computation cost forces existing dialogue response selection models to rank only a small number of candidates, recalled by a coarse-grained model, precluding many high-quality candidates. To overcome this problem, we present a novel and efficient response selection model and a set of tailor-designed learning strategies to train it effectively. The proposed model consists of a dense retrieval module and an interaction layer, which could directly select the proper response from a large corpus. We conduct re-rank and full-rank evaluations on widely used benchmarks to evaluate our proposed model. Extensive experimental results demonstrate that our proposed model notably outperforms the state-of-the-art baselines on both re-rank and full-rank evaluations. Moreover, human evaluation results show that the response quality could be improved further by enlarging the candidate pool with nonparallel corpora. In addition, we also release high-quality benchmarks that are carefully annotated for more accurate dialogue response selection evaluation. All source codes, datasets, model parameters, and other related resources have been publicly available. 1
Tian Lan 0003, Deng Cai 0002, Yan Wang 0060, Yixuan Su, Heyan Huang, Xianling Mao
ACM Trans. Inf. Syst.4
2023 Biomedical Named Entity Recognition via Dictionary-based Synonym Generalization
abstract
Biomedical named entity recognition is one of the core tasks in biomedical natural language processing (BioNLP).To tackle this task, numerous supervised/distantly supervised approaches have been proposed.Despite their remarkable success, these approaches inescapably demand laborious human effort.To alleviate the need of human effort, dictionarybased approaches have been proposed to extract named entities simply based on a given dictionary.However, one downside of existing dictionary-based approaches is that they are challenged to identify concept synonyms that are not listed in the given dictionary, which we refer as the synonym generalization problem.In this study, we propose a novel Synonym Generalization (SynGen) framework that recognizes the biomedical concepts contained in the input text using span-based predictions.In particular, SynGen introduces two regularization terms, namely, (1) a synonym distance regularizer; and (2) a noise perturbation regularizer, to minimize the synonym generalization error.To demonstrate the effectiveness of our approach, we provide a theoretical analysis of the bound of synonym generalization error.We extensively evaluate our approach on a wide range of benchmarks and the results verify that SynGen outperforms previous dictionary-based models by notable margins.Lastly, we provide a detailed analysis to further reveal the merits and inner-workings of our approach.1
Yixuan Su, Zaiqiao Meng, Nigel Collier
EMNLP2
2023 Specialist or Generalist? Instruction Tuning for Specific NLP Tasks
abstract
The potential of large language models (LLMs) to simultaneously perform a wide range of natural language processing (NLP) tasks has been the subject of extensive research.Although instruction tuning has proven to be a data-efficient method for transforming LLMs into such generalist models, their performance still lags behind specialist models trained exclusively for specific tasks.In this paper, we investigate whether incorporating broadcoverage generalist instruction tuning can contribute to building a specialist model.We hypothesize that its efficacy depends on task specificity and skill requirements.Our experiments assess four target tasks with distinct coverage levels, revealing that integrating generalist instruction tuning consistently enhances model performance when the task coverage is broad.The effect is particularly pronounced when the amount of task-specific training data is limited.Further investigation into three target tasks focusing on different capabilities demonstrates that generalist instruction tuning improves understanding and reasoning abilities.However, for tasks requiring factual knowledge, generalist data containing hallucinatory information may negatively affect the model's performance.Overall, our work provides a systematic guide for developing specialist models with general instruction tuning.Our code and other related resources can be found at https://github.com/DavidFanzz/ Generalist_or_Specialist.
Chufan Shi, Yixuan Su, Cheng Yang 0002, Yujiu Yang 0001, Deng Cai 0002
EMNLP2
2023 Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data Perspective
abstract
There are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by presenting a straightforward and fundamental explanation from the data perspective. Our preliminary investigation reveals a strong correlation between the degeneration issue and the presence of repetitions in training data. Subsequent experiments also demonstrate that by selectively dropping out the attention to repetitive words in training data, degeneration can be significantly minimized. Furthermore, our empirical analysis illustrates that prior works addressing the degeneration issue from various standpoints, such as the high-inflow words, the likelihood objective, and the self-reinforcement phenomenon, can be interpreted by one simple explanation. That is, penalizing the repetitions in training data is a common and fundamental factor for their effectiveness. Moreover, our experiments reveal that penalizing the repetitions in training data remains critical even when considering larger model sizes and instruction tuning.
Tian Lan 0003, Deng Cai 0002, Lemao Liu, Nigel Collier, Taro Watanabe, Yixuan Su
NeurIPS8
2022 Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models
abstract
Zaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zaiqiao Meng, Fangyu Liu 0001, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier
ACL (1)4
2022 Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System
abstract
Pre-trained language models have been recently shown to benefit task-oriented dialogue (TOD) systems.Despite their success, existing methods often formulate this task as a cascaded generation problem which can lead to error accumulation across different sub-tasks and greater data annotation overhead.In this study, we present PPTOD, a unified plug-andplay model for task-oriented dialogue.In addition, we introduce a new dialogue multi-task pre-training strategy that allows the model to learn the primary TOD task completion skills from heterogeneous dialog corpora.We extensively test our model on three benchmark TOD tasks, including end-to-end dialogue modelling, dialogue state tracking, and intent classification.Experimental results show that PPTOD achieves new state of the art on all evaluated tasks in both high-resource and lowresource scenarios.Furthermore, comparisons against previous SOTA methods show that the responses generated by PPTOD are more factually correct and semantically coherent as judged by human annotators. 1
Yixuan Su, Lei Shu 0004, Elman Mansimov, Arshit Gupta, Deng Cai 0002, Yi-An Lai
ACL (1)1
2022 From Easy to Hard: A Dual Curriculum Learning Framework for Context-Aware Document Ranking
abstract
Contextual information in search sessions is important for capturing users' search intents. Various approaches have been proposed to model user behavior sequences to improve document ranking in a session. Typically, training samples of (search context, document) pairs are sampled randomly in each training epoch. In reality, the difficulty to understand user's search intent and to judge document's relevance varies greatly from one search context to another. Mixing up training samples of different difficulties may confuse the model's optimization process. In this work, we propose a curriculum learning framework for context-aware document ranking, in which the ranking model learns matching signals between the search context and the candidate document in an easy-to-hard manner. In so doing, we aim to guide the model gradually toward a global optimum. To leverage both positive and negative examples, two curricula are designed. Experiments on two real query log datasets show that our proposed framework can improve the performance of several existing methods significantly, demonstrating the effectiveness of curriculum learning for context-aware document ranking.
Yutao Zhu 0001, Jian-Yun Nie, Yixuan Su, Haonan Chen 0005, Xinyu Zhang 0019, Zhicheng Dou
CIKM3
2022 Measuring and Reducing Model Update Regression in Structured Prediction for NLP
abstract
Recent advance in deep learning has led to rapid adoption of machine learning based NLP models in a wide range of applications. Despite the continuous gain in accuracy, backward compatibility is also an important aspect for industrial applications, yet it received little research attention. Backward compatibility requires that the new model does not regress on cases that were correctly handled by its predecessor. This work studies model update regression in structured prediction tasks. We choose syntactic dependency parsing and conversational semantic parsing as representative examples of structured prediction tasks in NLP. First, we measure and analyze model update regression in different model update settings. Next, we explore and benchmark existing techniques for reducing model update regression including model ensemble and knowledge distillation. We further propose a simple and effective method, Backward-Congruent Re-ranking (BCR), by taking into account the characteristics of structured output. Experiments show that BCR can better mitigate model update regression than model ensemble and knowledge distillation approaches.
Deng Cai 0002, Elman Mansimov, Yi-An Lai, Yixuan Su, Lei Shu 0004
NeurIPS4
2022 A Contrastive Framework for Neural Text Generation
abstract
Text generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g., beam search) of neural language models often lead to degenerate solutions---the generated text is unnatural and contains undesirable repetitions. Existing approaches introduce stochasticity via sampling or modify training objectives to decrease the probabilities of certain tokens (e.g., unlikelihood training). However, they often lead to solutions that lack coherence. In this work, we show that an underlying reason for model degeneration is the anisotropic distribution of token representations. We present a contrastive solution: (i) SimCTG, a contrastive training objective to calibrate the model's representation space, and (ii) a decoding method---contrastive search---to encourage diversity while maintaining coherence in the generated text. Extensive experiments and analyses on three benchmarks from two languages demonstrate that our proposed approach outperforms state-of-the-art text generation methods as evaluated by both human and automatic metrics.
Yixuan Su, Tian Lan 0003, Yan Wang 0060, Dani Yogatama, Lingpeng Kong, Nigel Collier
NeurIPS1
2021 Dialogue Response Selection with Hierarchical Curriculum Learning
abstract
Yixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060
ACL/IJCNLP (1)1
2021 Non-Autoregressive Text Generation with Pre-trained Language Models
abstract
Yixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, Nigel Collier. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Yixuan Su, Deng Cai 0002, Yan Wang 0060, David Vandyke, Simon Baker, Piji Li, Nigel Collier
EACL1
2021 PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval Memory
abstract
The ability of dialogue systems to express pre-specified style during conversations has a direct, positive impact on their usability and user satisfaction. While it has attracted much research interest, existing methods often generate stylistic responses at the cost of content quality. In this work, we introduce a prototype-to-style (PS) framework to tackle the challenge of stylistic dialogue generation. The proposed framework first exploits an Information Retrieval (IR) system and extracts a response prototype from the retrieved response. A stylistic response generator then takes the response prototype and the desired style as input to produce a high-quality and stylistic response. To effectively train the proposed model and imitate the real testing environment, we introduce a new style-aware learning objective and a denoising learning strategy. Results on three benchmark datasets (gender, emotion, and sentiment) from two languages demonstrate that the proposed approach significantly outperforms existing baselines both in terms of in-domain and cross-domain evaluations.
Yixuan Su, Yan Wang 0060, Deng Cai 0002, Simon Baker, Anna Korhonen, Nigel Collier
IEEE ACM Trans. Audio Speech Lang. Process.1