Shichao Sun

dblp:143/5725 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Personalized Large Language Model Assistant with Evolving Conditional Memory
abstract
With the rapid development of large language models, AI assistants like ChatGPT have become increasingly integrated into people’s works and lives but are limited in personalized services. In this paper, we present a plug-and-play framework that could facilitate personalized large language model assistants with evolving conditional memory. The personalized assistant focuses on intelligently preserving the knowledge and experience from the history dialogue with the user, which can be applied to future tailored responses that better align with the user’s preferences. Generally, the assistant generates a set of records from the dialogue, stores them in a memory bank, and retrieves related memory to improve the quality of the response. For the crucial memory design, we explore different ways of constructing the memory and propose a new memorizing mechanism named conditional memory to enhance the memory management of the framework. We also investigate the retrieval and usage of memory in the generation process. To better evaluate the personalized assistants’ abilities, we build the first evaluation benchmark from three critical aspects: continuing previous dialogue, learning personalized knowledge and learning from user feedback. The experimental results illustrate the effectiveness of our method.
Ruifeng Yuan, Shichao Sun, Yongqi Li 0001, Ziqiang Cao, Wenjie Li 0002
COLING2
2024 Dissecting Human and LLM Preferences
abstract
As a relative quality comparison of model responses, human and Large Language Model (LLM) preferences serve as common alignment goals in model fine-tuning and criteria in evaluation.Yet, these preferences merely reflect broad tendencies, resulting in less explainable and controllable models with potential safety risks.In this work, we dissect the preferences of human and 32 different LLMs to understand their quantitative composition, using annotations from real-world user-model conversations for a fine-grained, scenario-wise analysis.We find that humans are less sensitive to errors, favor responses that support their stances, and show clear dislike when models admit their limits.On the contrary, advanced LLMs like GPT-4-Turbo emphasize correctness, clarity, and harmlessness more.Additionally, LLMs of similar sizes tend to exhibit similar preferences, regardless of their training methods, and fine-tuning for alignment does not significantly alter the preferences of pretrained-only LLMs.Finally, we show that preference-based evaluation can be intentionally manipulated.In both training-free and training-based settings, aligning a model with the preferences of judges boosts scores, while injecting the least preferred properties lowers them.This results in notable score shifts: up to 0.59 on MT-Bench (1-10 scale) and 31.94 on AlpacaEval 2.0 (0-100 scale), highlighting the significant impact of this strategic adaptation.We have made all resources of this project publicly available.
Shichao Sun, Yikai Zhang 0003, Hai Zhao 0001, Pengfei Liu 0003
ACL (1)3
2024 Recovery Should Never Deviate from Ground Truth: Mitigating Exposure Bias in Neural Machine Translation
abstract
In Neural Machine Translation, models are often trained with teacher forcing and suffer from exposure bias due to the discrepancy between training and inference. Current token-level solutions, such as scheduled sampling, aim to maximize the model’s capability to recover from errors. Their loss functions have a side effect: a sequence with errors may have a larger probability than the ground truth. The consequence is that the generated sequences may recover too much and deviate from the ground truth. This side effect is verified in our experiments. To address this issue, we propose using token-level contrastive learning to coordinate three training objectives: the usual MLE objective, an objective for recovery from errors, and a new objective to explicitly constrain the recovery in a scope that does not impact the ground truth. Our empirical analysis shows that this method effectively achieves these objectives in training and reduces the frequency with which the third objective is violated. We conduct experiments on three language pairs: German-English, Russian-English, and English-Russian. Results show that our method outperforms the vanilla Transformer and other methods addressing the exposure bias.
Jianfei He, Shichao Sun, Xiaohua Jia, Wenjie Li 0002
EAMT (1)2
2024 FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in LLMs
abstract
Fuzzy reasoning is vital due to the frequent use of imprecise information in daily contexts.However, the ability of current large language models (LLMs) to handle such reasoning remains largely uncharted.In this paper, we introduce a new benchmark, FROG, for fuzzy reasoning, featuring real-world mathematical word problems that incorporate generalized quantifiers.Our experimental findings reveal that fuzzy reasoning continues to pose significant challenges for LLMs.Moreover, we find that existing methods designed to enhance reasoning do not consistently improve performance in tasks involving fuzzy logic.Additionally, our results show an inverse scaling effect in the performance of LLMs on FROG.Interestingly, we also demonstrate that strong mathematical reasoning skills are not necessarily indicative of success on our benchmark 1 .
Yiyuan Li, Shichao Sun, Pengfei Liu 0003
EMNLP2
2024 Generative Judge for Evaluating Alignment
abstract
The rapid development of Large Language Models (LLMs) has substantially expanded the range of tasks they can address. In the field of Natural Language Processing (NLP), researchers have shifted their focus from conventional NLP tasks (e.g., sequence tagging and parsing) towards tasks that revolve around aligning with human needs (e.g., brainstorming and email writing). This shift in task distribution imposes new requirements on evaluating these aligned models regarding *generality* (i.e., assessing performance across diverse scenarios), *flexibility* (i.e., examining under different protocols), and *interpretability* (i.e., scrutinizing models with explanations). In this paper, we propose a generative judge with 13B parameters, **Auto-J**, designed to address these challenges. Our model is trained on user queries and LLM-generated responses under massive real-world scenarios and accommodates diverse evaluation protocols (e.g., pairwise response comparison and single-response evaluation) with well-structured natural language critiques. To demonstrate the efficacy of our approach, we construct a new testbed covering 58 different scenarios. Experimentally, **Auto-J** outperforms a series of strong competitors, including both open-source and closed-source models, by a large margin. We also provide detailed analysis and case studies to further reveal the potential of our method and make a variety of resources public at https://github.com/GAIR-NLP/auto-j.
Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao 0001, Pengfei Liu 0003
ICLR2
2024 OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
abstract
The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11,163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39.97\% overall accuracy (28.67\% for mathematics and 29.71\% for physics), illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features.
Zengzhi Wang, Shijie Xia, Xuefeng Li 0003, Haoyang Zou, Ruijie Xu 0005, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, Yikai Zhang 0003, Yuqing Yang 0004, Binjie Wang, Shichao Sun, Yiyuan Li, Steffi Chern, Yiwei Qin, Jiadi Su, Yixiu Liu, Shaoting Zhang 0001, Dahua Lin, Yu Qiao 0001, Pengfei Liu 0003
NeurIPS15
2024 RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
abstract
Despite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems.
Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001
NeurIPS9
2024 Dialogue acts enhanced extract-abstract framework for meeting summarization
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Ziqiang Cao, Sujian Li
Inf. Process. Manag.1
2023 Empirical Analysis of Beam Search Curse and Search Errors with Model Errors in Neural Machine Translation
abstract
Beam search is the most popular decoding method for Neural Machine Translation (NMT) and is still a strong baseline compared with the newly proposed sampling-based methods. To better understand beam search, we investigate its two well-recognized issues, beam search curse and search errors, at the sentence level. We find that only less than 30% of sentences in the test set experience these issues. Meanwhile, there is a related phenomenon. For the majority of sentences, their gold references have lower probabilities than the predictions from beam search. We also test with different levels of model errors including a special test using training samples and models without regularization. We find that these phenomena still exist even for a model with an accuracy of 95% although they are mitigated. These findings show that it is not promising to improve beam search by seeking higher probabilities in searching and further reducing its search errors. The relationship between the quality and the probability of predictions at the sentence level in our results provides useful information to find new ways to improve NMT.
Jianfei He, Shichao Sun, Xiaohua Jia, Wenjie Li 0002
EAMT2
2023 Improving Sentence Similarity Estimation for Unsupervised Extractive Summarization
abstract
Unsupervised extractive summarization aims to extract salient sentences from a document as the summary without labeled data. Recent literatures mostly research how to leverage sentence similarity to rank sentences in the order of salience. However, sentence similarity estimation using pre-trained language models mostly takes little account of document-level information and has a weak correlation with sentence salience ranking. In this paper, we proposed two novel strategies to improve sentence similarity estimation for unsupervised extractive summarization. We use contrastive learning to optimize a document-level objective that sentences from the same document are more similar than those from different documents. Moreover, we use mutual learning to enhance the relationship between sentence similarity estimation and sentence salience ranking, where an extra signal amplifier is used to refine the pivotal information. Experimental results demonstrate the effectiveness of our strategies.1
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Sujian Li
ICASSP1
2023 Load Change Assessment-Based Feedforward Compensation for FCS-MPCC Used in PMSMs Considering Load Disturbances
abstract
This paper presents a load change assessment-based feedforward compensation method for finite control set model predictive current control (FCS-MPCC) in permanent magnet synchronous motors (PMSMs). The objective is to address the adverse effects of load disturbances on FCS-MPCC, which can lead to deteriorated control performance and system instability. To mitigate these effects, a novel feedforward compensation mechanism is proposed by integrating a load change assessment mechanism within the FCS-MPCC framework. The mechanism enables real-time estimation of load changes by accurately capturing their rate and direction. A sliding mode torque observer (SMTO) is developed to ensure accurate load estimation, characterized by fast response and strong robustness. The stability of the SMTO is analyzed using a Lyapunov function. Furthermore, a technique is proposed to generate feedforward compensation values based on the load change assessment, specifically applied to the q-axis reference current. Comparative simulation results verify the effectiveness of the proposed strategies.
Shichao Sun, Yaofei Han, Chao Gong 0001
IECON2
2023 An Adaptive Passivity-Based Controller for Boost Converter Supplying Constant Power Load
abstract
This paper presents a robust and adaptive passivity-based controller (PBC) to regulate the output voltage of a DC-DC boost converter supplying a constant power load (CPL) paralleling with a resistor in dc microgrids which exhibiting limit-cycle behavior when CPLs dominate the load. Based on Lyapunov's theory, the limit-cycle behavior is analyzed in detail. This instability effect caused by CPLs has been attempted to be removed by introducing PBC, but the previously proposed PBC is not robust. The aim of this paper is to propose a new PBC with a certain degree of robustness. To deal with the steady-state error problem caused by the input voltage variation, the nonlinear disturbance observer (NDO) is added to PBC for disturbance estimation. The proposed PBC strategy is robust when dealing with load variation. The steady-state error caused by source variation is eliminated by introducing the estimated value of NDO into the design of PBC. The simulation results are provided to verify the effectiveness and robustness of the proposed control strategy.
Shichao Sun, Xinrong Huang
IECON1
2023 Aligning Language Models with Human Preferences via a Bayesian Approach
abstract
In the quest to advance human-centric natural language generation (NLG) systems, ensuring alignment between NLG models and human preferences is crucial. For this alignment, current popular methods leverage a reinforcement learning (RL) approach with a reward model trained on feedback from humans. However, inherent disagreements due to the subjective nature of human preferences pose a significant challenge for training the reward model, resulting in a deterioration of the NLG performance. To tackle this issue, previous approaches typically rely on majority voting or averaging to consolidate multiple inconsistent preferences into a merged one. Although straightforward to understand and execute, such methods suffer from an inability to capture the nuanced degrees of disaggregation among humans and may only represent a specialized subset of individuals, thereby lacking the ability to quantitatively disclose the universality of human preferences. To address this challenge, this paper proposes a novel approach, which employs a Bayesian framework to account for the distribution of disagreements among human preferences as training a preference model, and names it as $\textbf{d-PM}$. Besides, considering the RL strategy's inefficient and complex training process over the training efficiency, we further propose utilizing the contrastive learning strategy to train the NLG model with the preference scores derived from the d-PM model. Extensive experiments on two human-centric NLG tasks, i.e., emotional support conversation and integrity ``Rule-of-Thumb'' generation, show that our method consistently exceeds previous SOTA models in both automatic and human evaluations.
Jiashuo Wang, Haozhao Wang, Shichao Sun, Wenjie Li 0002
NeurIPS3
2022 Rethinking the framework constructed by counterfactual functional model
Chao Wang 0095, Linfang Liu, Shichao Sun, Wei Wang 0009
Appl. Intell.3
2019 A Goal-Driven Tree-Structured Neural Model for Math Word Problems
abstract
Most existing neural models for math word problems exploit Seq2Seq model to generate solution expressions sequentially from left to right, whose results are far from satisfactory due to the lack of goal-driven mechanism commonly seen in human problem solving. This paper proposes a tree-structured neural model to generate expression tree in a goal-driven manner. Given a math word problem, the model first identifies and encodes its goal to achieve, and then the goal gets decomposed into sub-goals combined by an operator in a top-down recursive way. The whole process is repeated until the goal is simple enough to be realized by a known quantity as leaf node. During the process, two-layer gated-feedforward networks are designed to implement each step of goal decomposition, and a recursive neural network is used to encode fulfilled subtrees into subtree embeddings, which provides a better representation of subtrees than the simple goals of subtrees. Experimental results on the dataset Math23K have shown that our tree-structured model outperforms significantly several state-of-the-art models.
Zhipeng Xie, Shichao Sun
IJCAI2
2017 BiLSTM-Based Models for Metaphor Detection
Shichao Sun, Zhipeng Xie
NLPCC1
2013 Min-max regret approach to the optimal path finding problem in stochastic time-dependent networks
abstract
A theoretical study was conducted on finding optimal paths in transportation networks where link travel times were stochastic and time-dependent (STD). The methodology of robust optimization was applied as measures for comparing time-varying, random path travel times for a priori optimization. In accordance with the situation in real world, a stochastic consistent condition was provided for the STD networks and under this condition, a mathematical proof was given that the STD robust optimal path problem can be simplified into a minimum problem in a specific time-dependent network. Then a modified label setting algorithm was designed and tested to find travelers' robust optimal path in a sampled STD network with computation complexity of O(n2∗ e). The computational results confirmed the validity of the Min-Max regret approach and proved that the proposed algorithm can solve for the optimal path in STD networks with a polynomial-time computation complexity. Besides, some other advantages of the Min-Max regret approach are also discussed.
Shichao Sun, Zhengyu Duan, Dongyuan Yang
ISADS1