Shuo Wang 0013

dblp:63/1591-13 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0001-5408-3145ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 5 first-author · 33 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM Simulation
abstract
Deploying Large Language Models (LLMs) in specialized domains introduces significant societal and compliance risks, including bias amplification, misinformation propagation, and privacy violations. These risks predominantly emerge from the dynamic interactions between LLMs and humans in specific contexts. Different domains face unique distribution of hazards, and varying interaction modalities introduce distinct levels of exposure and vulnerability. However, current risk assessment frameworks lack a systematic methodology to capture this dynamic interplay. In this work, we introduce the HEV Generative Sandbox, a novel risk evaluation framework that simulates human-LLM behavior to quantify domain-contextual risks across three interdependent dimensions: 1) Hazard (H): Domain-specific threats inherent to a given context; 2) Exposure (E): The extent to which the LLM and its users are subjected to hazardous scenarios; 3) Vulnerability (V): The susceptibility of the system to risk due to human interaction or model weaknesses. Our approach pioneers "domain-rooted scenario generation", wherein we sample contextual distributions from domain-specific corpora and simulate diverse inputs. By unifying dynamic scenario simulation, causal risk decomposition, and closed-loop evaluation, the HEV Generative Sandbox provides a scalable, domain-sensitive methodology for responsible LLM deployment. This work contributes to advancing the safe deployment of LLMs by providing a comprehensive and automated risk evaluation framework.
Zhiyi Hou, Xiaoang Xu, Shuo Wang 0013, Huijia Wu, Kaicheng Yu, Yang Yu 0011, ChengXiang Zhai
AAAI4
2026 Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
abstract
Shaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu, Xiaoyuan Yi, Yukun Yan, Shuo Wang, Yu Gu, Ge Yu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shaohua Duan, Pengcheng Huang 0004, Zhenghao Liu 0001, Xiaoyuan Yi, Yukun Yan, Shuo Wang 0013, Yu Gu 0002, Ge Yu 0001, Maosong Sun 0001
ACL (1)7
2026 Empirical Analysis of Decoding Biases in Masked Diffusion Models
abstract
Pengcheng Huang, Tianming Liu, Zhenghao Liu, Yukun Yan, Shuo Wang, Tong Xiao, Zulong Chen, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengcheng Huang 0004, Zhenghao Liu 0001, Yukun Yan, Shuo Wang 0013, Tong Xiao 0001, Zulong Chen, Maosong Sun 0001
ACL (1)5
2026 Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
abstract
Zhenghao Liu, Zhuoyang Wu, Xinze Li, Yukun Yan, Shuo Wang, Zulong Chen, Yu Gu, Ge Yu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhenghao Liu 0001, Zhuoyang Wu, Yukun Yan, Shuo Wang 0013, Zulong Chen, Yu Gu 0002, Ge Yu 0001, Maosong Sun 0001
ACL (1)5
2026 CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning
abstract
Dingling Xu, Ruobing Wang, Qingfei Zhao, Yukun Yan, Zhichun Wang, Daren Zha, Shi Yu, Zhenghao Liu, Shuo Wang, Xu Han, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingling Xu, Qingfei Zhao, Yukun Yan, Zhichun Wang, Daren Zha, Shi Yu 0001, Zhenghao Liu 0001, Shuo Wang 0013, Maosong Sun 0001
ACL (1)9
2026 AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
abstract
Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanle Zhao, Zilin Sang, Qi Shi 0002, Wei-Lun Zhao, Shuo Wang 0013, Duzhen Zhang, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)6
2026 LISRec: Modeling User Preferences with Learned Item Shortcuts for Sequential Recommendation
abstract
User-item interaction histories are pivotal for sequential recommendation systems but often include noise, such as unintended clicks or actions that fail to reflect genuine user preferences. To address this, we propose Learned Item Shortcuts for Sequential Recommendation (LISRec), a novel framework that explicitly captures stable preferences by extracting personalized semantic shortcuts from historical interactions. LISRec first learns task-agnostic semantic representations to assess item similarities, then constructs a personalized semantic graph over all user-interacted items. By identifying the maximal semantic connectivity subset within this graph, LISRec selects the most representative items as semantic shortcuts to guide user preference modeling. This focused representation filters out irrelevant actions while preserving the diversity of genuine interests. Experimental results on the Yelp and Amazon Product datasets illustrate that LISRec achieves a 13% improvement over baseline recommendation models, showing its effectiveness in capturing stable user interests. Further analysis indicates that shortcut-based histories better capture user preferences, making more accurate and relevant recommendations. All codes and datasets are available at https://github.com/NEUIR/LISRec.
Haidong Xin, Zhenghao Liu 0001, Sen Mei, Yukun Yan, Shi Yu 0001, Shuo Wang 0013, Zulong Chen, Yu Gu 0002, Ge Yu 0001, Chenyan Xiong
KDD (1)6
2026 Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
abstract
Multimodal Retrieval-Augmented Generation (MRAG) has shown promise in mitigating hallucinations in Multimodal Large Language Models (MLLMs) by incorporating external knowledge. However, existing methods typically adhere to rigid retrieval paradigms by mimicking fixed retrieval trajectories and thus fail to fully exploit the knowledge of different retrieval experts through dynamic interaction based on the model's knowledge needs or evolving reasoning states. To overcome this limitation, we introduce Mixture-of-Retrieval Experts (MoRE), a novel framework that enables MLLMs to collaboratively interact with diverse retrieval experts for more effective knowledge exploitation. Specifically, MoRE learns to dynamically determine which expert to engage with, conditioned on the evolving reasoning state. To effectively train this capability, we propose Stepwise Group Relative Policy Optimization (Step-GRPO), which goes beyond sparse outcome-based supervision by encouraging MLLMs to interact with multiple retrieval experts and synthesize fine-grained rewards, thereby teaching the MLLM to fully coordinate all experts when answering a given query. Experimental results on diverse open-domain QA benchmarks demonstrate the effectiveness of MoRE, achieving average performance gains of over 7% compared to competitive baselines. Notably, MoRE exhibits strong adaptability by dynamically coordinating heterogeneous experts to precisely locate relevant information, validating its capability for robust, reasoning-driven expert collaboration. All codes and data are released on https://github.com/OpenBMB/MoRE.
Zhenghao Liu 0001, Yishan Li, Yukun Yan, Shuo Wang 0013, Yu Gu 0002, Minghe Yu 0001, Ge Yu 0001, Maosong Sun 0001
SIGIR6
2026 ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment
Yifan Ji, Zhenghao Liu 0001, Yukun Yan, Zulong Chen, Shuo Wang 0013, Yu Gu 0002, Ge Yu 0001
SIGIR7
2025 ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks.However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts.In contrast, parsing charts into code provides lossless representations that can effectively contain all critical details.Although existing open-source MLLMs have achieved success in chart understanding tasks, they still face two major challenges when applied to chart-to-code tasks: (1) Low executability and poor restoration of chart details in the generated code and (2) Lack of large-scale and diverse training data.To address these challenges, we propose ChartCoder, the first dedicated chart-to-code MLLM, which leverages Code LLMs as the language backbone to enhance the executability of the generated code.Furthermore, we introduce Chart2Code-160k, the first large-scale and diverse dataset for chartto-code generation, and propose the Snippetof-Thought (SoT) method, which transforms direct chart-to-code generation data into stepby-step generation.Experiments demonstrate that ChartCoder, with only 7B parameters, surpasses existing open-source MLLMs on chartto-code benchmarks, achieving superior chart restoration and code excitability.Our code is available at https://github.com/thunlp/ ChartCoder.89 seed code with 27 chart types Available functions and parameters
Xuanle Zhao, Xianzhen Luo, Qi Shi 0002, Chi Chen 0005, Shuo Wang 0013, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)5
2025 LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models
abstract
Zihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Qi Shi, Zhixing Tan, Xu Han, Xiaodong Shi, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuo Wang 0013, Yu Chao, Zhili Li, Qi Shi 0002, Zhixing Tan, Xu Han 0007, Xiaodong Shi, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)4
2025 RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
abstract
Kunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Ruobing Wang, Shuo Wang, Yishan Li, Nan Zhang, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kunlun Zhu, Dingling Xu, Yukun Yan, Zhenghao Liu 0001, Shi Yu 0001, Shuo Wang 0013, Yishan Li, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)8
2025 LegalDuet: Learning Fine-Grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
Buqiang Xu, Zhenghao Liu 0001, Huiyuan Xie, Xiaoyuan Yi, Shuo Wang 0013, Yukun Yan, Liner Yang, Yu Gu 0002, Ge Yu 0001
ADMA (1)6
2025 On LLM-Based Scientific Inductive Reasoning Beyond Equations
abstract
Brian S. Lin, Jiaxin Yuan, Zihan Zhou, Shouli Wang, Shuo Wang, Cunliang Kong, Qi Shi, Yuxuan Li, Liner Yang, Zhiyuan Liu, Maosong Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Brian S. Lin, Shouli Wang, Shuo Wang 0013, Cunliang Kong, Qi Shi 0002, Liner Yang, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP5
2025 From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
abstract
Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages.However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics.In contrast, multi-way parallel data, where identical content is aligned across multiple languages, provides stronger cross-lingual consistency and offers greater potential for improving multilingual performance.In this paper, we introduce a large-scale, high-quality multiway parallel corpus, TED2025, based on TED Talks.The corpus spans 113 languages, with up to 50 languages aligned in parallel, ensuring extensive multilingual coverage.Using this dataset, we investigate best practices for leveraging multi-way parallel data to enhance LLMs, including strategies for continued pretraining, instruction tuning, and the analysis of key influencing factors.Experiments on six multilingual benchmarks show that models trained on multiway parallel data consistently outperform those trained on unaligned multilingual data.
Yingli Shen, Wen Lai, Shuo Wang 0013, Kangyang Luo, Alexander Fraser 0001, Maosong Sun 0001
EMNLP3
2025 Exploring the Impact of Personality Traits on LLM Bias and Toxicity
abstract
With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interest.While the "personification" enhances human experiences of interactivity and adaptability of LLMs, it gives rise to critical concerns about content safety, particularly regarding bias, sentiment, and toxicity of LLM generation.This study explores how assigning different personality traits to LLMs affects the toxicity and biases of their outputs.Leveraging the widely accepted HEXACO personality framework developed in social psychology, we design experimentally sound prompts to test three LLMs' performance on three toxic and bias benchmarks.The findings demonstrate the sensitivity of all three models to HEXACO personality traits and, more importantly, a consistent variation in the biases, negative sentiment, and toxicity of their output.In particular, adjusting the levels of several personality traits can effectively reduce bias and toxicity in model performance, similar to humans' correlations between personality traits and toxic behaviors.The findings highlight the additional need to examine content safety besides the efficiency of training or fine-tuning methods for LLM personification, they also suggest a potential for the adjustment of personalities to be a simple and low-cost method to conduct controlled text generation.
Shuo Wang 0013, Renhao Li, Yulin Yuan, Min Yang 0007, Derek F. Wong
EMNLP1
2025 Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors
abstract
LLMs are transforming software development, yet most code benchmarks still emphasize syntactic or functional correctness in simple, single-error cases.These settings miss the core difficulty of real-world data science debugging, where runtime errors propagate across multiple lines (multi-hop) and often appear in sets (multi-bug).We introduce DSDBench: Data Science Debugging Benchmark, the first benchmark to systematically evaluate LLMs on this challenge.Unlike general debugging benchmark suites such as SWE-bench, DSD-Bench targets non-expert, data-centric scripting, where practitioners rely heavily on blackbox libraries and write exploratory code that is error-prone and difficult to debug.Evaluations of state-of-the-art LLMs reveal large performance gaps: even frontier models that excel at code generation fail to reliably trace and resolve these errors, exposing a critical "generation versus understanding" gap.DSDBench provides a resource to drive progress toward more robust and trustworthy AI-assisted data science.
Zhiyu Yang 0001, Shuo Wang 0013, Yukun Yan, Yang Deng 0002
EMNLP2
2025 RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
abstract
Retrieval-Augmented Generation (RAG) has proven its effectiveness in mitigating hallucinations in Large Language Models (LLMs) by retrieving knowledge from external resources. To adapt LLMs for the RAG systems, current approaches use instruction tuning to optimize LLMs, improving their ability to utilize retrieved knowledge. This supervised fine-tuning (SFT) approach focuses on equipping LLMs to handle diverse RAG tasks using different instructions. However, it trains RAG modules to overfit training signals and overlooks the varying data preferences among agents within the RAG system. In this paper, we propose a Differentiable Data Rewards (DDR) method, which end-to-end trains RAG systems by aligning data preferences between different RAG modules. DDR works by collecting the rewards to optimize each agent in the RAG system with the rollout method, which prompts agents to sample some potential responses as perturbations, evaluates the impact of these perturbations on the whole RAG system, and subsequently optimizes the agent to produce outputs that improve the performance of the RAG system. Our experiments on various knowledge-intensive tasks demonstrate that DDR significantly outperforms the SFT method, particularly for LLMs with smaller-scale parameters that depend more on the retrieved knowledge. Additionally, DDR exhibits a stronger capability to align the data preference between RAG modules. The DDR method makes the generation module more effective in extracting key information from documents and mitigating conflicts between parametric memory and external knowledge. All codes are available at https://github.com/OpenMatch/RAG-DDR.
Sen Mei, Zhenghao Liu 0001, Yukun Yan, Shuo Wang 0013, Shi Yu 0001, Zheni Zeng, Ge Yu 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Chenyan Xiong
ICLR5
2025 VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
abstract
Retrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG systems are solely based on text, rendering it impossible to utilize vision information like layout and images that play crucial roles in real-world multi-modality documents. In this paper, we introduce VisRAG, which tackles this issue by establishing a vision-language model (VLM)-based RAG pipeline. In this pipeline, instead of first parsing the document to obtain text, the document is directly embedded using a VLM as an image and then retrieved to enhance the generation of a VLM. Compared to traditional text-based RAG, VisRAG maximizes the retention and utilization of the data information in the original documents, eliminating the information loss introduced during the parsing process. We collect both open-source and synthetic data to train the retriever in VisRAG and explore a variety of generation methods. Experiments demonstrate that VisRAG outperforms traditional RAG in both the retrieval and generation stages, achieving a 20–40% end-to-end performance gain over traditional text-based RAG pipeline. Further analysis reveals that VisRAG is efficient in utilizing training data and demonstrates strong generalization capability, positioning it as a promising solution for RAG on multi-modality documents. Our code and data are available at https://github.com/openbmb/visrag.
Shi Yu 0001, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu 0001, Shuo Wang 0013, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR8
2025 CITR: Efficient Long Video Understanding Needs Causal Importance
abstract
Long video understanding is essential for various practical applications including surveillance and film analysis. While recent Vision-Language Models (VLMs) have advanced performance in this domain, efficiency remains a key challenge, especially for hour-long videos. Existing methods commonly reduce visual tokens via compression in the vision encoder, but token count still grows linearly with video length. Alternative approaches apply importance-based token reduction in the language model, yet their non-causal design limits efficiency gains to offline, single-query settings. In this work, we emphasize the need for causal importance estimation-where a token's relevance is determined only from prior context-to enable efficient, real-time long video understanding. We propose ØurMethod, a Causal Importance-based Token Reduction framework to reduce visual token redundancy in long video understanding tasks, enabling practical memory control and enhanced computational efficiency. Experiments on both offline and streaming benchmarks show that ØurMethod reduces latency by 49% in offline multi-query scenarios and effectively controls chunked prefilling time in streaming, all within a 24GB memory footprint and with less than 1% performance drop. The code and appendix are available at https://github.com/Columbine21/CITR.
Yanghao Li, Yuxiang Huang 0001, Chi Chen 0005, Shuo Wang 0013, Zhinan Gou
ACM Multimedia6
2025 MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
abstract
Hanqing Wang, Yixia Li, Shuo Wang, Guanhua Chen, Yun Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hanqing Wang 0003, Yixia Li, Shuo Wang 0013, Guanhua Chen 0001, Yun Chen 0007
NAACL (Long Papers)3
2025 DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection
abstract
The rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a large-scale multilingual corpus constructed from newly extracted Common Crawl data and existing multilingual sources. DCAD-2000 covers 2,282 languages, 46.72TB of text, and 8.63 billion documents, spanning 155 high- and medium-resource languages and 159 writing scripts. To overcome the limitations of existing data cleaning approaches, which rely on manually designed heuristic thresholds, we reframe data cleaning as an anomaly detection problem. This dynamic filtering paradigm substantially improves data quality by automatically identifying and removing noisy or anomalous content. By fine-tuning LLMs on DCAD-2000, we demonstrate notable improvements in data quality, robustness of the cleaning pipeline, and downstream performance, particularly for low-resource languages across multiple multilingual benchmarks.
Yingli Shen, Wen Lai, Shuo Wang 0013, Xueren Zhang, Kangyang Luo, Alexander Fraser 0001, Maosong Sun 0001
NeurIPS3
2025 A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
abstract
Large Reasoning Models (LRMs) achieve superior performance by extending the thought length. However, a lengthy thinking trajectory leads to reduced efficiency. Most of the existing methods are stuck in the assumption of overthinking and attempt to reason efficiently by compressing the Chain-of-Thought, but this often leads to performance degradation. To address this problem, we introduce A*-Thought, an efficient tree search-based unified framework designed to identify and isolate the most essential thoughts from the extensive reasoning chains produced by these models. It formulates the reasoning process of LRMs as a search tree, where each node represents a reasoning span in the giant reasoning space. By combining the A* search algorithm with a cost function specific to the reasoning path, it can efficiently compress the chain of thought and determine a reasoning path with high information density and low cost. In addition, we also propose a bidirectional importance estimation mechanism, which further refines this search process and enhances its efficiency beyond uniform sampling. Extensive experiments on several advanced math tasks show that A*-Thought effectively balances performance and efficiency over a huge search space. Specifically, A*-Thought can improve the performance of QwQ-32B by 2.39$\times$ with low-budget and reduce the length of the output token by nearly 50\% with high-budget. The proposed method is also compatible with several other LRMs, demonstrating its generalization capability. The code can be accessed at: https://github.com/AI9Stars/AStar-Thought.
Xiaoang Xu, Shuo Wang 0013, Zhenghao Liu 0001, Huijia Wu, Peipei Li 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Zhaofeng He 0001
NeurIPS2
2025 Building a Coding Assistant via the Retrieval-Augmented Language Model
abstract
Pretrained language models have shown strong effectiveness in code-related tasks, such as code retrieval, code generation, code summarization, and code completion tasks. In this article, we propose COde assistaNt viA retrieval-augmeNted language model (CONAN), which aims to build a code assistant by mimicking the knowledge-seeking behaviors of humans during coding. Specifically, it consists of a code structure-aware retriever (CONAN-R) and a dual-view code representation-based retrieval-augmented generation model (CONAN-G). CONAN-R pretrains CodeT5 using Code-Documentation Alignment and Masked Entity Prediction tasks to make language models code structure-aware and learn effective representations for code snippets and documentation. Then CONAN-G designs a dual-view code representation mechanism for implementing a retrieval-augmented code generation model. CONAN-G regards the code documentation descriptions as prompts, which help language models better understand the code semantics. Our experiments show that CONAN achieves convincing performance on different code generation tasks and significantly outperforms previous retrieval augmented code generation models. Our further analyses show that CONAN learns tailored representations for both code snippets and documentation by aligning code-documentation data pairs and capturing structural semantics by masking and predicting entities in the code data. Additionally, the retrieved code snippets and documentation provide necessary information from both program language and natural language to assist the code generation process. CONAN can also be used as an assistant for Large Language Models (LLMs), providing LLMs with external knowledge in shorter code document lengths to improve their effectiveness on various code tasks. It shows the ability of CONAN to extract necessary information and help filter out the noise from retrieved code documents.
Hanbin Wang, Zhenghao Liu 0001, Shi Yu 0001, Shuo Wang 0013, Yukun Yan, Yu Gu 0002, Ge Yu 0001
ACM Trans. Inf. Syst.5
2024 LoRA-Flow: Dynamic LoRA Fusion for Large Language Models in Generative Tasks
abstract
LoRA employs lightweight modules to customize large language models (LLMs) for each downstream task or domain, where different learned additional modules represent diverse skills.Combining existing LoRA modules to address new tasks can enhance the reusability of learned LoRA modules, particularly beneficial for tasks with limited annotated data.Most prior works on LoRA combination primarily rely on task-level weights for each involved LoRA, making different examples and tokens share the same LoRA weights.However, in generative tasks, different tokens may necessitate diverse skills to manage.Taking the Chinese math task as an example, understanding the problem description may depend more on the Chinese LoRA, while the calculation part may rely more on the math LoRA.To this end, we propose LoRA-Flow, which utilizes dynamic weights to adjust the impact of different LoRA modules.The weights at each step are determined by a fusion gate with extremely few parameters, which can be learned with only 200 training examples.Experiments across six generative tasks demonstrate that our method consistently outperforms baselines with tasklevel fusion weights.This underscores the necessity of introducing dynamic fusion weights for LoRA combination. 1
Hanqing Wang 0003, Bowen Ping, Shuo Wang 0013, Xu Han 0007, Yun Chen 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)3
2024 UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
abstract
Haoyu Wang, Shuo Wang, Yukun Yan, Xujia Wang, Zhiyu Yang, Yuzhuang Xu, Zhenghao Liu, Liner Yang, Ning Ding, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shuo Wang 0013, Yukun Yan, Xujia Wang, Zhiyu Yang 0001, Yuzhuang Xu, Zhenghao Liu 0001, Liner Yang, Ning Ding 0002, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)2
2024 ınftyBench: Extending Long Context Evaluation Beyond 100K Tokens
abstract
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yingfa Chen, Shengding Hu, Zihang Xu, Moo Khai Hao, Xu Han 0007, Zhen Leng Thai, Shuo Wang 0013, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)9
2024 Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
abstract
Yuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang, Xiaolong Wang, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuanchi Zhang, Yile Wang 0001, Shuo Wang 0013, Xiaolong Wang 0014, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)4
2024 MCTS: A Multi-Reference Chinese Text Simplification Dataset
abstract
Text simplification aims to make the text easier to understand by applying rewriting transformations. There has been very little research on Chinese text simplification for a long time. The lack of generic evaluation data is an essential reason for this phenomenon. In this paper, we introduce MCTS, a multi-reference Chinese text simplification dataset. We describe the annotation process of the dataset and provide a detailed analysis. Furthermore, we evaluate the performance of several unsupervised methods and advanced large language models. We additionally provide Chinese text simplification parallel data that can be used for training, acquired by utilizing machine translation and English text simplification. We hope to build a basic understanding of Chinese text simplification through the foundational work and provide references for future research. All of the code and data are released at https://github.com/blcuicall/mcts/.
Ruining Chong, Luming Lu, Liner Yang, Jinran Nie, Zhenghao Liu 0001, Shuo Wang 0013, Shuhan Zhou, Yaoxin Li, Erhong Yang
LREC/COLING6
2024 Pluggable Neural Machine Translation Models via Memory-augmented Adapters
abstract
Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratch for each user requirement, we propose a memory-augmented adapter to steer pretrained NMT models in a pluggable manner. Specifically, we construct a multi-granular memory based on the user-provided text samples and propose a new adapter architecture to combine the model representations and the retrieved results. We also propose a training strategy using memory dropout to reduce spurious dependencies between the NMT model and the memory. We validate our approach on both style- and domain-specific experiments and the results indicate that our method can outperform several representative pluggable baselines.
Yuzhuang Xu, Shuo Wang 0013, Peng Li 0030, Xuebo Liu 0002, Xiaolong Wang 0014, Yang Liu 0005
LREC/COLING2
2024 Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
abstract
Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability.
Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS2
2024 OneBit: Towards Extremely Low-bit Large Language Models
abstract
Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs. However, current quantization methods suffer severe performance degradation when the bit-width is extremely reduced, and thus focus on utilizing 4-bit or 8-bit values to quantize models. This paper boldly quantizes the weight matrices of LLMs to 1-bit, paving the way for the extremely low bit-width deployment of LLMs. For this target, we introduce a 1-bit model compressing framework named OneBit, including a novel 1-bit parameter representation method to better quantize LLMs as well as an effective parameter initialization method based on matrix decomposition to improve the convergence speed of the quantization framework. Sufficient experimental results indicate that OneBit achieves good performance (at least 81% of the non-quantized performance on LLaMA models) with robust training processes when only using 1-bit weight matrices.
Yuzhuang Xu, Xu Han 0007, Zonghan Yang, Shuo Wang 0013, Qingfu Zhu, Zhiyuan Liu 0001, Wanxiang Che
NeurIPS4
2024 Understanding and Mitigating the Uncertainty in Zero-Shot Translation
abstract
Zero-shottranslation is a promising direction for building a comprehensive multilingual neural machine translation (MNMT) system. However, its quality is still not satisfactory due to off-target issues. In this paper, we aim to understand and alleviate the off-target issues from the perspective of uncertainty in zero-shot translation. By carefully examining the translation output and model confidence, we identify two uncertainties that are responsible for the off-target issues, namely, extrinsic data uncertainty and intrinsic model uncertainty. Based on the observations, we propose two lightweight and complementary approaches to denoise the training data for model training and explicitly penalize the off-target translations by unlikelihood training during model training. Extensive experiments on both balanced and imbalanced datasets show that our approaches significantly improve the performance of zero-shot translation over strong MNMT baselines.
Wenxuan Wang 0001, Wenxiang Jiao, Shuo Wang 0013, Zhaopeng Tu, Michael R. Lyu
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 TemplateGEC: Improving Grammatical Error Correction with Detection Template
abstract
Yinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong, Derek F. Wong, Yang Gao, Heyan Huang, Min Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xuebo Liu 0002, Shuo Wang 0013, Peiyuan Gong, Derek F. Wong, Yang Gao 0016, Heyan Huang, Min Zhang 0005
ACL (1)3
2022 MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators
abstract
Prompting has recently been shown as a promising approach for applying pre-trained language models to perform downstream tasks.We present Multi-Stage Prompting, a simple and automatic approach for leveraging pre-trained language models to translation tasks.To better mitigate the discrepancy between pre-training and translation, MSP divides the translation process via pre-trained language models into multiple separate stages: the encoding stage, the re-encoding stage, and the decoding stage.During each stage, we independently apply different continuous prompts for allowing pretrained language models better shift to translation tasks.We conduct extensive experiments on three translation tasks.Experiments show that our method can significantly improve the translation performance of pre-trained language models.
Zhixing Tan, Xiangwen Zhang, Shuo Wang 0013, Yang Liu 0005
ACL (1)3
2022 Integrating Vectorized Lexical Constraints for Neural Machine Translation
abstract
Lexically constrained neural machine translation (NMT), which controls the generation of NMT models with pre-specified constraints, is important in many practical scenarios.Due to the representation gap between discrete constraints and continuous vectors in NMT models, most existing works choose to construct synthetic data or modify the decoding algorithm to impose lexical constraints, treating the NMT model as a black box.In this work, we propose to open this black box by directly integrating the constraints into NMT models.Specifically, we vectorize source and target constraints into continuous keys and values, which can be utilized by the attention modules of NMT models.The proposed integration method is based on the assumption that the correspondence between keys and values in attention modules is naturally suitable for modeling constraint pairs.Experimental results show that our method consistently outperforms several representative baselines on four language pairs, demonstrating the superiority of integrating vectorized lexical constraints.
Shuo Wang 0013, Zhixing Tan, Yang Liu 0005
ACL (1)1
2022 A Template-based Method for Constrained Neural Machine Translation
abstract
Machine translation systems are expected to cope with various types of constraints in many practical scenarios.While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints into the translation process of NMT models.Although many approaches have been proposed to address this issue, most existing methods can not satisfy the following three desiderata at the same time: (1) high translation quality, (2) high match accuracy, and (3) low latency.In this work, we propose a template-based method that can yield results with high translation quality and match accuracy and the inference speed of our method is comparable with unconstrained NMT models.Our basic idea is to rearrange the generation of constrained and unconstrained tokens through a template.Our method does not require any changes in the model architecture and the decoding algorithm.Experimental results show that the proposed template-based approach can outperform several representative baselines in both lexically and structurally constrained translation tasks.
Shuo Wang 0013, Peng Li 0030, Zhixing Tan, Zhaopeng Tu, Maosong Sun 0001, Yang Liu 0005
EMNLP1
2020 On the Inference Calibration of Neural Machine Translation
abstract
Confidence calibration, which aims to make model predictions equal to the true correctness measures, is important for neural machine translation (NMT) because it is able to offer useful indicators of translation errors in the generated output.While prior studies have shown that NMT models trained with label smoothing are well-calibrated on the groundtruth training data, we find that miscalibration still remains a severe challenge for NMT during inference due to the discrepancy between training and inference.By carefully designing experiments on three language pairs, our work provides in-depth analyses of the correlation between calibration and translation performance as well as linguistic properties of miscalibration and reports a number of interesting findings that might help humans better analyze, understand and improve NMT models.Based on these observations, we further propose a new graduated label smoothing method that can improve both inference calibration and translation performance.1
Shuo Wang 0013, Zhaopeng Tu, Shuming Shi 0001, Yang Liu 0005
ACL1
2019 Improving Back-Translation with Uncertainty-based Confidence Estimation
abstract
Shuo Wang, Yang Liu, Chao Wang, Huanbo Luan, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shuo Wang 0013, Yang Liu 0005, Chao Wang 0049, Huan-Bo Luan, Maosong Sun 0001
EMNLP/IJCNLP (1)1