Shuzheng Si

dblp:324/3680 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
abstract
Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to construct high-quality and easily verifiable training data without human annotation. Also, we propose Dual-GRPO, a rule-based reinforcement learning method that includes three tailored rule-based rewards derived from synthesized short-form QA data, while simultaneously optimizing both short-form and long-form response generation. Notably, Dual-GRPO eliminates the need to manually label preference data to train reward models and avoids over-optimizing short-form generation when relying only on the synthesized short-form QA data. Experimental results show that CANOE greatly improves the faithfulness of LLMs across 11 different tasks, even outperforming the most advanced LLMs, e.g., GPT-4o and OpenAI o1.
Shuzheng Si, Haozhe Zhao, Yuzhuo Bai, Zhitong Wang, Bofei Gao, Kangyang Luo, Wenhao Li 0003, Yufei Huang 0008, Gang Chen 0039, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun 0001
AAAI1
2026 ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement
abstract
Kangyang Luo, Yuzhuo Bai, Shuzheng Si, Cheng Gao, Zhitong Wang, Yingli Shen, Wenhao Li, Zhu Liu, Yufeng Han, Jiayi Wu, Cunliang Kong, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kangyang Luo, Yuzhuo Bai, Shuzheng Si, Zhitong Wang, Yingli Shen, Wenhao Li 0003, Zhu Liu 0005, Yufeng Han, Cunliang Kong, Maosong Sun 0001
ACL (1)3
2026 A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Task
abstract
Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen 0039, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun 0001
ACL (1)1
2025 Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
abstract
Shuzheng Si, Haozhe Zhao, Gang Chen, Cheng Gao, Yuzhuo Bai, Zhitong Wang, Kaikai An, Kangyang Luo, Chen Qian, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yuzhuo Bai, Zhitong Wang, Kaikai An, Kangyang Luo, Fanchao Qi, Baobao Chang, Maosong Sun 0001
ACL (1)1
2025 UltraIF: Advancing Instruction Following from the Wild
abstract
Instruction-following made modern large language models (LLMs) helpful assistants.However, the key to taming LLMs on complex instructions remains mysterious, for that there are huge gaps between models trained by opensource community and those trained by leading companies.To bridge the gap, we propose a simple and scalable approach ULTRAIF for building LLMs that can follow complex instructions with open-source data.ULTRAIF first decomposes real-world user prompts into simpler queries, constraints, and corresponding evaluation questions for the constraints.Then, we train an UltraComposer to compose constraintassociated prompts with evaluation questions.This prompt composer allows us to synthesize complicated instructions as well as filter responses with evaluation questions.In our experiment, for the first time, we successfully align LLaMA-3.1-8B-Base to catch up with its instruct version on 5 instruction-following benchmarks without any benchmark information, using only 8B model as response generator and evaluator.The aligned model also achieved competitive scores on other benchmarks.Moreover, we also show that ULTRAIF could further improve LLaMA-3.1-8B-Instruct through self-alignment, motivating broader use cases for the method.Our code is available at https://github.com/kkk-an/UltraIF.
Kaikai An, Ganqu Cui, Shuzheng Si, Baobao Chang
EMNLP4
2025 Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation
abstract
Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang, Pu Zhao, Lele Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Baobao Chang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang 0029, Pu Zhao 0004, Le-le Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Baobao Chang
EMNLP6
2025 GATEAU: Selecting Influential Samples for Long Context Alignment
abstract
Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun 0001
EMNLP1
2025 Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
abstract
Large vision-language models (LVLMs) have achieved impressive results in vision-language tasks.However, LVLMs suffer from hallucinations caused by language bias, which neglects images while over-relying on text.We identify two reasons for the bias: 1).Different training scales between the LLM pretraining and LVLM alignment stage.2).The learned inference bias due to short-term dependency of text data.Therefore, we propose LACING, designed to address such bias with MuLtimodal DuAlattention MeChanIsm (MDA) aNd Soft-Image Guidance (SIG).Specifically, MDA adopts a parallel dual-attention mechanism that constructs separate attention for visual and text inputs to enhance integration of visual inputs across model.SIG uses a learnable soft visual prompt during training and inference to replace visual inputs, designed to compel LVLMs to prioritize text inputs during inference.Experiments across different model architectures and scales demonstrate that LACING effectively debiases LVLMs from their language bias, enhancing visual comprehension and reducing hallucinations without additional resources.
Haozhe Zhao, Shuzheng Si, Liang Chen 0024, Yichi Zhang 0010, Maosong Sun 0001, Baobao Chang, Minjia Zhang
EMNLP2
2025 SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
abstract
Resolution of complex SQL issues persists as a significant bottleneck in real-world database applications. Current Large Language Models (LLMs), while adept at text-to-SQL translation, have not been rigorously evaluated on the more challenging task of debugging on SQL issues. In order to address this gap, we introduce **BIRD-CRITIC**, a new SQL issue debugging benchmark comprising 530 carefully curated PostgreSQL tasks (**BIRD-CRITIC-PG**) and 570 multi-dialect tasks (**BIRD-CRITIC-Multi**), which are distilled from authentic user issues and replayed within new environments to facilitate rigorous and contamination-free evaluation. Baseline evaluations on BIRD-CRITIC underscore the task's complexity, with the leading reasoning model **O3-Mini** achieving only 38.87% success rate on **BIRD-CRITIC-PG** and 33.33% on **BIRD-CRITIC-Multi**. Meanwhile, realizing open-source models for database tasks is crucial which can empower local development while safeguarding data privacy. Therefore, we present **Six-Gym** (**S**ql-f**IX**-Gym), a training environment for elevating the capabilities of open-source models specifically for SQL issue debugging. This environment leverages **SQL-Rewind** strategy, which automatically generates executable issue-solution datasets by reverse-engineering issues from verified SQLs. However, popular trajectory-based fine-tuning methods do not explore substantial supervisory signals. We further propose *f*-Plan Boosting, which extracts high-level debugging plans automatically from SQL solutions, enabling the teacher LLMs to harvest and produce 73.7% more successful trajectories for training. We integrate these components into an open-source agent, **BIRD-Fixer**. Based on Qwen-2.5-Coder-14B, **BIRD-Fixer** raises its success rate to 38.11% on **BIRD-CRITIC-PG** and 29.65% on **BIRD-CRITIC-Multi**, surpassing many leading proprietary models such as Claude-3.7-Sonnet and GPT-4.1, marking a significant step toward democratizing sophisticated SQL-debugging capabilities for both research and industry.
Jinyang Li 0003, Ge Qu, Per Jacobsson, Bowen Qin, Binyuan Hui, Shuzheng Si, Nan Huo, Ziwei Tang, Yuanshuai Li, Florensia Widjaja, Xintong Zhu, Feige Zhou, Yannis Papakonstantinou, Fatma Özcan 0001, Chenhao Ma 0001, Reynold Cheng
NeurIPS7
2024 One-Shot Learning as Instruction Data Prospector for Large Language Models
abstract
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang, Min Yang, Lei Zhang, Shuzheng Si, Ling-Hao Chen, Junhao Liu, Tongliang Liu, Fei Huang, Yongbin Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang 0004, Min Yang 0007, Lei Zhang 0201, Shuzheng Si, Junhao Liu 0001, Tongliang Liu, Fei Huang 0002, Yongbin Li 0001
ACL (1)7
2024 UniPCM: Universal Pre-trained Conversation Model with Task-aware Automatic Prompt
abstract
Recent researches have shown that multi-task instruction tuning after pre-training greatly improves the model’s robustness and transfer ability, which is crucial for building a high-quality dialog system. However, most previous works on multi-task instruction tuning rely heavily on human-defined input format or prompt, which is not optimal in quality and quantity.In this work, we propose to use Task-aware Automatic Prompt generation (TAP) to automatically generate high-quality prompts. Using the high-quality prompts generated, we scale the corpus of the pre-trained conversation model to 122 datasets from 15 dialog-related tasks, resulting in Universal Pre-trained Conversation Model (UniPCM), a powerful foundation model for various conversational tasks and different dialog systems. Extensive experiments have shown that UniPCM is robust to input prompts and capable of various dialog-related tasks. Moreover, UniPCM has strong transfer ability and excels at low resource scenarios, achieving SOTA results on 9 different datasets ranging from task-oriented dialog to open-domain conversation. Furthermore, we are amazed to find that TAP can generate prompts on par with those collected with crowdsourcing.
Yucheng Cai, Yuchuan Wu, Shuzheng Si, Yuan Shao, Zhijian Ou
LREC/COLING4
2024 MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
abstract
Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks. In this paper, we address the limitation above by 1) introducing vision-language Model with **M**ulti-**M**odal **I**n-**C**ontext **L**earning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts. Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context. Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC.
Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 0001, Kaikai An, Liang Chen 0024, Zixuan Liu 0001, Sheng Wang 0012, Wenjuan Han, Baobao Chang
ICLR3
2024 Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
abstract
Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen, Yufeng He, Kaikai An, Baobao Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen 0024, Yufeng He, Kaikai An, Baobao Chang
NAACL-HLT3
2024 UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
abstract
This paper presents UltraEdit, a large-scale (~ 4M editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples: 1) UltraEdit includes more diverse editing instructions by combining LLM creativity and in-context editing examples by human raters; 2) UltraEdit is anchored on real images (photographs or artworks), which offers more diversity and less biases than those purely synthesized by text-to-image models; 3) UltraEdit supports region-based editing with high-quality, automatically produced region annotations. Our experiments show that canonical diffusion-based editing baselines trained on UltraEdit set new records on challenging MagicBrush and Emu-Edit benchmarks, respectively. Our analysis further confirms the crucial role of real image anchors and region-based editing data. The dataset, code, and models will be made public.
Haozhe Zhao, Xiaojian Ma 0001, Liang Chen 0024, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li 0003, Baobao Chang
NeurIPS4
2023 SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
abstract
Task-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and real-world spoken con- versation scenarios. While several small-scale spoken TOD datasets are proposed to address robustness issues such as ASR errors, they ignore the unique challenges in spoken conversation. To tackle the limitations, we introduce SpokenWOZ, a large-scale speech-text dataset for spoken TOD, containing 8 domains, 203k turns, 5.7k dialogues and 249 hours of audios from human-to-human spoken conversations. SpokenWOZ further incorporates common spoken characteristics such as word-by-word processing and reasoning in spoken language. Based on these characteristics, we present cross-turn slot and reasoning slot detection as new challenges. We conduct experiments on various baselines, including text-modal models, newly proposed dual-modal models, and LLMs, e.g., ChatGPT. The results show that the current models still have substantial room for improvement in spoken conversation, where the most advanced dialogue state tracker only achieves 25.65% in joint goal accuracy and the SOTA end-to-end model only correctly completes the user request in 52.1% of dialogues. Our dataset, code, and leaderboard are available at https://spokenwoz.github.io/SpokenWOZ-github.io/.
Shuzheng Si, Yuchuan Wu, Ting-En Lin, Yinpei Dai, Hangyu Li 0003, Rui Yan 0001, Fei Huang 0002, Yongbin Li 0001
NeurIPS1
2022 SCL-RAI: Span-based Contrastive Learning with Retrieval Augmented Inference for Unlabeled Entity Problem in NER
abstract
Unlabeled Entity Problem (UEP) in Named Entity Recognition (NER) datasets seriously hinders the improvement of NER performance. This paper proposes SCL-RAI to cope with this problem. Firstly, we decrease the distance of span representations with the same label while increasing it for different ones via span-based contrastive learning, which relieves the ambiguity among entities and improves the robustness of the model over unlabeled entities. Then we propose retrieval augmented inference to mitigate the decision boundary shifting problem. Our method significantly outperforms the previous SOTA method by 4.21% and 8.64% F1-score on two real-world datasets.
Shuzheng Si, Shuang Zeng, Jiaxing Lin, Baobao Chang
COLING1
2022 Mining Clues from Incomplete Utterance: A Query-enhanced Network for Incomplete Utterance Rewriting
abstract
Incomplete utterance rewriting has recently raised wide attention.However, previous works do not consider the semantic structural information between incomplete utterance and rewritten utterance or model the semantic structure implicitly and insufficiently.To address this problem, we propose a QUEry-Enhanced Network (QUEEN).Firstly, our proposed query template explicitly brings guided semantic structural knowledge between the incomplete utterance and the rewritten utterance making model perceive where to refer back to or recover omitted tokens.Then, we adopt a fast and effective edit operation scoring network to model the relation between two tokens.Benefiting from extra information and the well-designed network, QUEEN achieves state-of-the-art performance on several public datasets.
Shuzheng Si, Shuang Zeng, Baobao Chang
NAACL-HLT1