EDBT 2026 Demo / reviewers in the wild / expert
Xingyu Sui
dblp:389/7980
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 42% Trustworthy machine learning · 21% Efficient and distributed learning · 10% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.0 | 1 | 2026 | Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities · AAAI 2026 |
Natural language and speech › Question answering and dialogue systems › open-domain dialogue
emotional support conversation |
1.0 | 1 | 2026 | TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent · ACL (1) 2026 |
Natural language and speech › Language models and text generation
hallucination mitigation |
1.0 | 1 | 2026 | TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model
large reasoning model |
1.0 | 1 | 2026 | Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.9 | 1 | 2025 | When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
multilingual reasoning |
0.9 | 1 | 2025 | When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs · ACL (1) 2025 |
Security and privacy of machine learning › large language model safety
jailbreak defense |
0.9 | 1 | 2025 | AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender · EMNLP 2025 |
Natural language and speech › Language models and text generation › alignment
aligned large language models |
0.3 | 1 | 2025 | AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
representation steering · 1.7adversarial prompting · 1.7tool augmentation · 1.0supervised fine-tuning · 1.0empirical analysis · 1.0adaptive reasoning · 1.0safety risk measurement · 0.9representation ablation · 0.9mitigation · 0.9causal intervention · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational CapabilitiesabstractRecent advancements in Large Reasoning Models (LRMs), such as OpenAI's o1/o3 and DeepSeek-R1, have demonstrated remarkable performance in specialized reasoning tasks through human-like deliberative thinking and long chain-of-thought reasoning. However, our systematic evaluation across various model families (DeepSeek, Qwen, and LLaMA) and scales (7B to 32B) reveals that acquiring these deliberative reasoning capabilities significantly reduces the foundational capabilities of LRMs, including notable declines in helpfulness and harmlessness, alongside substantially increased inference costs. Importantly, we demonstrate that adaptive reasoning---employing modes like Zero-Thinking, Less-Thinking, and Summary-Thinking---can effectively alleviate these drawbacks. Our empirical insights underline the critical need for developing more versatile LRMs capable of dynamically allocating inference-time compute according to specific task characteristics. Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng 0002, Xuda Zhi, Yongbo Huang, Wanxiang Che, Ting Liu 0001, Bing Qin 0001 |
AAAI | 2 |
| 2026 | When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue AgentsabstractJiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Weixiang Zhao, Bing Qin 0001 |
ACL (1) | 5 |
| 2026 | TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue AgentabstractEmotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance.However, existing ESC systems and benchmarks largely focus on affective support in text-only settings, overlooking how external tools can enable factual grounding and reduce hallucination in multi-turn emotional support.We introduce TEA-Bench, the first interactive benchmark for evaluating tool-augmented agents in ESC, featuring realistic emotional scenarios, an MCP-style tool environment, and process-level metrics that jointly assess the quality and factual grounding of emotional support.Experiments on nine LLMs show that tool augmentation generally improves emotional support quality and reduces hallucination, but the gains are strongly capacity-dependent: stronger models use tools more selectively and effectively, while weaker models benefit only marginally.We further release TEA-Dialog, a dataset of toolenhanced ESC dialogues, and find that supervised fine-tuning improves in-distribution support but generalizes poorly.Our results underscore the importance of tool use in building reliable emotional support agents. 1 Xingyu Sui, Yulin Hu, Jiahe Guo, Weixiang Zhao, Bing Qin 0001 |
ACL (1) | 1 |
| 2026 | The gains do not make up for the losses: a comprehensive evaluation for safety alignment of large language models via machine unlearningabstractAbstract Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased , leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark M u B ench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful inputs and 10 types of jailbreak attacks. Furthermore, we examine whether MU introduces side effects, focusing on over-safety and utility-loss. Extensive experiments are performed on 3 popular LLMs with 7 recent MU methods. The results highlight a challenging trilemma in safety alignment without side effects, indicating that there is still considerable room for further exploration. M u B ench serves as a comprehensive benchmark, fostering future research on MU for safety alignment of LLMs. Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng 0002, Bing Qin 0001, Wanxiang Che |
Frontiers Comput. Sci. | 3 |
| 2025 | Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMsabstractWeixiang Zhao, Yulin Hu, Yang Deng, Jiahe Guo, Xingyu Sui, Xinyang Han, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weixiang Zhao, Yulin Hu, Yang Deng 0002, Jiahe Guo, Xingyu Sui, An Zhang 0003, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001 |
ACL (1) | 5 |
| 2025 | AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak DefenderabstractWeixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng 0002, An Zhang 0003, Xingyu Sui, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001 |
EMNLP | 6 |
| 2025 | When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual ReasonersabstractMultilingual reasoning remains a significant challenge for large language models (LLMs), with performance disproportionately favoring high-resource languages. Drawing inspiration from cognitive neuroscience, which suggests that human reasoning functions largely independently of language processing, we hypothesize that LLMs similarly encode reasoning and language as separable components that can be disentangled to enhance multilingual reasoning. To evaluate this, we perform a causal intervention by ablating language-specific representations at inference time. Experiments on 10 open-weight LLMs spanning 11 typologically diverse languages show that this language-specific ablation consistently boosts multilingual reasoning performance. Layer-wise analyses further confirm that language and reasoning representations can be effectively disentangled throughout the model, yielding improved multilingual reasoning capabilities, while preserving top-layer language features remains essential for maintaining linguistic fidelity. Compared to post-training methods such as supervised fine-tuning or reinforcement learning, our training-free language-reasoning disentanglement achieves comparable or superior results with minimal computational overhead. These findings shed light on the internal mechanisms underlying multilingual reasoning in LLMs and suggest a lightweight and interpretable strategy for improving cross-lingual generalization. Weixiang Zhao, Jiahe Guo, Yang Deng 0002, Tongtong Wu, Wenxuan Zhang 0001, Yulin Hu, Xingyu Sui, Wanxiang Che, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001 |
NeurIPS | 7 |
| 2025 | Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized AlignmentabstractPersonalized alignment is essential for enabling large language models (LLMs) to engage effectively in user-centric dialogue. While recent prompt-based and offline optimization methods offer preliminary solutions, they fall short in cold-start scenarios and long-term personalization due to their inherently static and shallow designs. In this work, we introduce the Reinforcement Learning for Personalized Alignment (RLPA) framework, in which an LLM interacts with a simulated user model to iteratively infer and refine user profiles through dialogue. The training process is guided by a dual-level reward structure: the Profile Reward encourages accurate construction of user representations, while the Response Reward incentivizes generation of responses consistent with the inferred profile. We instantiate RLPA by fine-tuning Qwen-2.5-3B-Instruct, resulting in Qwen-RLPA, which achieves state-of-the-art performance in personalized dialogue. Empirical evaluations demonstrate that Qwen-RLPA consistently outperforms prompting and offline fine-tuning baselines, and even surpasses advanced commercial models such as Claude-3.5 and GPT-4o. Further analysis highlights Qwen-RLPA's robustness in reconciling conflicting user preferences, sustaining long-term personalization and delivering more efficient inference compared to recent reasoning-focused LLMs. These results emphasize the potential of dynamic profile inference as a more effective paradigm for building personalized dialogue systems. Weixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo, Haixiao Liu, Biye Li, Bing Qin 0001, Ting Liu 0001 |
NeurIPS | 2 |