EDBT 2026 Demo / reviewers in the wild / expert
Shihan Dou
dblp:282/6213
· DBLP profile ↗
43ranked-venue papers
11as first author
42since 2021 · last 2026
0009-0002-6013-3035ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 9 first-author · 33 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningabstractOutcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B. Zhiheng Xi, Yiwen Ding, Senjie Jin, Shichun Liu, Jixuan Huang, Dingwen Yang, Jiafu Tang, Boyang Hong, Junjie Ye 0005, Shihan Dou, Ming Zhang 0030, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 12 |
| 2026 | OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic CodingabstractDeming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Pengyu Zhao, Chengjun Xiao, Qunhong Zeng, Qi Zhang, Xuanjing Huang, Qidi Xu, Tao Gui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Chengjun Xiao, Qunhong Zeng, Qi Zhang 0001, Xuanjing Huang 0001, Qidi Xu, Tao Gui |
ACL (1) | 6 |
| 2026 | Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-TrainingabstractChanghao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 6 |
| 2026 | DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency DataabstractShaofan Liu, Guoqiang Zhang, Shihan Dou, Huiyuan Zheng, Yiming Zhou, Junjie Ye, Shaowen Wang, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaofan Liu, Shihan Dou, Huiyuan Zheng, Junjie Ye 0005, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2026 | Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationabstractSearch agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and longtail information unavailable during pretraining.While reinforcement learning has been widely adopted for training such agents, existing approaches face key limitations: process supervision often suffers from unstable value estimation, whereas outcome supervision struggles with credit assignment due to sparse, trajectory-level rewards.To bridge this gap, we propose Contribution-Weighted GRPO (CW-GRPO), a framework that integrates process supervision into group relative policy optimization.Instead of directly optimizing process rewards, CW-GRPO employs an LLM judge to assess the retrieval utility and reasoning correctness at each search round, producing per-round contribution weights.These weights are used to rescale outcome-based advantages along the trajectory, enabling finegrained credit assignment without sacrificing optimization stability.Experiments on multiple knowledge-intensive benchmarks show that CW-GRPO outperforms standard GRPO by 5.0% on Qwen3-8B and 6.3% on Qwen3-1.7B,leading to more effective search behaviors.Additional analysis reveals that successful trajectories exhibit concentrated contributions in specific rounds, providing empirical insight into search agent tasks.Our code is available at https://github.com/zsxmwjz/CW-GRPO. Junzhe Wang 0001, Zhiheng Xi, Yajie Yang, Shihan Dou, Tao Gui, Qi Zhang 0001 |
ACL (1) | 5 |
| 2026 | JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal ConversationsabstractXinyi Xu, Bingguang Hao, Yongyi Xiong, Zimo Chen, Xinchen Liu, Hongxin Guo, Xuelong Wang, Silin Zhou, Shihan Dou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bingguang Hao, Yongyi Xiong, Zimo Chen, Xinchen Liu, Hongxin Guo, Xuelong Wang, Silin Zhou, Shihan Dou |
ACL (1) | 9 |
| 2026 | LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsabstractMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang, Junzhe Wang, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ming Zhang 0030, Yujiong Shen, Jingyi Deng, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang 0004, Junzhe Wang 0001, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang 0002, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 19 |
| 2026 | PRISM: Probabilistic Reward Model with Inherent Structural ModelingabstractYuhang Zhou, Yixin Cao, Yuchen Ni, Shihan Dou, Xutian Chen, Ge Zhang, Xiang Liu, Guangnan Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixin Cao 0002, Yuchen Ni, Shihan Dou, Xutian Chen, Ge Zhang 0009, Guangnan Ye |
ACL (1) | 4 |
| 2026 | VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-TrainingabstractDingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang, Chenhao Huang, Yunke Zhang, Yuran Wang, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiazheng Zhang, Junjie Ye 0005, Mingxu Chai, Enyu Zhou, Ming Zhang 0030, Caishuang Huang, Chenhao Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 2 |
| 2026 | What is wrong with your code generated by large language models? An extensive study
Shihan Dou, Haoxiang Jia, Shenxi Wu, Huiyuan Zheng, Muling Wu, Yunbo Tao, Ming Zhang 0030, Mingxu Chai, Jessica Fan, Zhiheng Xi, Yueming Wu 0001, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 1 |
| 2026 | SpikeBERT: A language spikformer learned from BERT with knowledge distillation
Changze Lv, Tianlong Li, Weiming Qiao, Muling Wu, Shihan Dou, Xiaoqing Zheng, Xuanjing Huang 0001 |
Neural Networks | 7 |
| 2026 | Eler: Ensemble Learning-Based Automated Verification of Code Clones
Shihan Dou, Siyue Feng, Yueming Wu 0001, Deqing Zou |
IEEE Trans. Software Eng. | 1 |
| 2025 | Alleviating Shifted Distribution in Human Preference Alignment through Meta-LearningabstractThe capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from the output distribution of the early policy model, gradually loses its ability to distinguish between responses from the newly shifted distribution. This issue is further compounded when the RM, trained on a specific data distribution, struggles to generalize to examples outside of that distribution. These two issues can be united as a challenge posed by the shifted distribution of the environment. To surmount this challenge, we introduce MetaRM, a novel method leveraging meta-learning to adapt the RM to the shifted environment distribution. MetaRM optimizes the RM in an alternating way, by preserving both the preferences of the original preference pairs, as well as maximizing discrimination power over new examples of the shifted distribution. Extensive experiments demonstrate that MetaRM can iteratively enhance the performance of human preference alignment by improving the RM's capacity to identify subtle differences in samples of shifted distributions. Shihan Dou, Yan Liu 0002, Enyu Zhou, Songyang Gao, Tianlong Li, Limao Xiong, Haoxiang Jia, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 1 |
| 2025 | Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference ModelingabstractShihan Dou, Jiayi Chen, Chenhao Huang, Feng Chen, Wei Chengzhi, Huiyuan Zheng, Shichun Liu, Yan Liu, Chenxiao Liu, Chao Xin, Lin Yan, Zongzhang Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shihan Dou, Chenhao Huang, Feng Chen 0042, Wei Chengzhi, Huiyuan Zheng, Shichun Liu, Yan Liu 0002, Chenxiao Liu, Chao Xin, Zongzhang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2025 | Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable MetricabstractData diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive fine-tuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level “novelty.” Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric. Yuming Yang 0001, Junjie Ye 0005, Shihan Dou, Xiao Wang 0042, Huijie Lv, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2025 | SpikeBERT: A Language Understanding Spiking Neural Network Learned from BERT with Knowledge Distillation
Changze Lv, Tianlong Li, Muling Wu, Shihan Dou, Xiaoqing Zheng, Xuanjing Huang 0001 |
CogSci | 6 |
| 2025 | Revisiting Jailbreaking for Large Language Models: A Representation Engineering PerspectiveabstractThe recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community. Tianlong Li, Zhenghua Wang, Muling Wu, Shihan Dou, Changze Lv, Xiaoqing Zheng, Xuanjing Huang 0001 |
COLING | 5 |
| 2025 | ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world ScenariosabstractExisting evaluations of tool learning primarily focus on validating the alignment of selected tools for large language models (LLMs) with expected outcomes. However, these approaches rely on a limited set of scenarios where answers can be pre-determined. Furthermore, a sole emphasis on outcomes disregards the complex capabilities required for LLMs to effectively use tools. To tackle this issue, we propose ToolEyes, a fine-grained system tailored for the evaluation of the LLMs’ tool learning capabilities in authentic scenarios. The system meticulously examines seven real-world scenarios, analyzing five dimensions crucial to LLMs in tool learning: format alignment, intent comprehension, behavior planning, tool selection, and answer organization. Additionally, ToolEyes incorporates a tool library boasting approximately 600 tools, serving as an intermediary between LLMs and the physical world. Evaluations involving ten LLMs across three categories reveal a preference for specific scenarios and limited cognitive abilities in tool learning. Intriguingly, expanding the model size even exacerbates the hindrance to tool learning. The code and data are available at https://github.com/Junjie-Ye/ToolEyes. Junjie Ye 0005, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
COLING | 8 |
| 2025 | Governance in Motion: Co-evolution of Constitutions and AI models for Scalable SafetyabstractChenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng, Jiazheng Zhang, Mingxu Chai, Ming Zhang, Shihan Dou, Fan Mo, Jie Shi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng, Jiazheng Zhang, Mingxu Chai, Ming Zhang 0030, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 8 |
| 2025 | RMB: Comprehensively benchmarking reward models in LLM alignmentabstractReward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization.
We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. We will release our evaluation code and datasets upon publication. Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Limao Xiong, Jessica Fan, Yurong Mou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 5 |
| 2025 | Pre-Trained Policy Discriminators are General Reward ModelsabstractWe offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance.
For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines.
POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks.
Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99.
The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models. Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026 |
NeurIPS | 1 |
| 2025 | EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem SolvingabstractWe introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3.7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available in the supplementary materials. Shihan Dou, Ming Zhang 0030, Chenhao Huang, Feng Chen 0042, Shichun Liu, Yan Liu 0002, Chenxiao Liu, Zongzhang Zhang, Tao Gui, Chao Xin, Wei Chengzhi, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 1 |
| 2025 | Improving RL Exploration for LLM Reasoning Through Retrospective Replay
Shihan Dou, Muling Wu, Tao Gui, Qi Zhang 0001 |
NLPCC (1) | 1 |
| 2025 | Visual Sketchbook: Enhancing Chart-to-Code Generation via Reflective Refinement
Junzhe Wang 0001, Zhiheng Xi, Wei He 0024, Dingwei Zhu, Shihan Dou, Tao Gui, Qi Zhang 0001 |
NLPCC (2) | 5 |
| 2025 | Fighting Fire with Fire: Continuous Attack for Adversarial Android Malware Detection
Yinyuan Zhang, Cuiying Gao, Yueming Wu 0001, Shihan Dou, Cong Wu 0003, Ying Zhang 0066, Wei Yuan 0001, Yang Liu 0003 |
USENIX Security Symposium | 4 |
| 2025 | The dual-edged sword: artificial intelligence's evolving role in academic peer review
Xuanjing Huang 0001, Shihan Dou, Zhangyue Yin |
Sci. China Inf. Sci. | 2 |
| 2025 | The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui |
Sci. China Inf. Sci. | 21 |
| 2024 | StepCoder: Improving Code Generation with Reinforcement Learning from Compiler FeedbackabstractShihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Yan Liu 0002, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang 0001, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2024 | LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style PluginabstractShihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Enyu Zhou, Yan Liu 0002, Songyang Gao, Limao Xiong, Yuhao Zhou 0005, Xiao Wang 0001, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2024 | TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesabstractMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong, Yujiong Shen, Shihan Dou, Jun Zhao, Junjie Ye, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ming Zhang 0030, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong 0001, Yujiong Shen, Shihan Dou, Jun Zhao 0019, Junjie Ye 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
EMNLP | 8 |
| 2024 | Improving Generalization of Alignment with Human Preferences through Group Invariant LearningabstractThe success of AI assistants based on language models (LLMs) hinges crucially on Reinforcement Learning from Human Feedback (RLHF), which enables the generation of responses more aligned with human preferences.
As universal AI assistants, there's a growing expectation for them to perform consistently across various domains.
However, previous work shows that Reinforcement Learning (RL) often exploits shortcuts to attain high rewards and overlooks challenging samples.
This focus on quick reward gains undermines both the stability in training and the model's ability to generalize to new, unseen data.
In this work, we propose a novel approach that can learn a consistent policy via RL across various data groups or domains.
Given the challenges associated with acquiring group annotations, our method automatically classifies data into different groups, deliberately maximizing performance variance.
Then, we optimize the policy to perform well on challenging groups.
Lastly, leveraging the established groups, our approach adaptively adjusts the exploration space, allocating more learning capacity to more challenging data and preventing the model from over-optimizing on simpler data. Experimental results indicate that our approach significantly enhances training stability and model generalization. Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou 0005, Zhiheng Xi, Xiao Wang 0001, Haoran Huang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 5 |
| 2024 | Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and FeedbackabstractThe success of AI assistants based on Language Models (LLMs) hinges on Reinforcement Learning from Human Feedback (RLHF) to comprehend and align with user intentions. However, traditional alignment algorithms, such as PPO, are hampered by complex annotation and training requirements. This reliance limits the applicability of RLHF and hinders the development of professional assistants tailored to diverse human preferences. In this work, we introduce Linear Alignment, a novel algorithm that aligns language models with human preferences in one single inference step, eliminating the reliance on data annotation and model training. Linear alignment incorporates a new parameterization for policy optimization under divergence constraints, which enables the extraction of optimal policy in a closed-form manner and facilitates the direct estimation of the aligned response. Extensive experiments on both general and personalized preference datasets demonstrate that linear alignment significantly enhances the performance and efficiency of LLM alignment across diverse scenarios. Songyang Gao, Qiming Ge, Shihan Dou, Junjie Ye 0005, Xiao Wang 0001, Yicheng Zou, Zhi Chen 0006, Hang Yan 0001, Qi Zhang 0001, Dahua Lin |
ICML | 4 |
| 2024 | Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningabstractIn this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$. Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICML | 15 |
| 2024 | CausalAPM: Generalizable Literal Disentanglement for NLU Debiasing
Shihan Dou, Songyang Gao, Tao Gui, Qi Zhang 0001 |
NLPCC (1) | 1 |
| 2024 | COCL: An Intelligent Framework for Enhancing Deep Learning-Based Vulnerability DetectionabstractDue to the powerful feature extraction capability ofdeep learning(DL), many recent studies have used it to conduct source code vulnerability analysis. However, although it has a good performance on artificial datasets, it does not perform satisfactorily on the real-world vulnerabilities with higher complexity. In this article, we introduce contrastive curriculum learning into DL-based vulnerability detection to find a suitable boundary to distinguish vulnerabilities from normal codes. Contrastive learning can be used to reduce the difference between different vulnerabilities while amplifying the difference between vulnerabilities and normal codes. To make the training phase of contrastive learning more intelligent, we apply curriculum learning to mimic the way humans acquire knowledge, which means that the model will learn simple samples first and then increase the difficulty of training samples. Specifically, we implement an intelligent framework (i.e.,contrastive curriculum learning (COCL)) that can enhance the detection effect of existing DL-based vulnerability detectors. To verify the capability ofCOCL, we select four state-of-the-art DL-based vulnerability detectors (i.e.,AutoVulTC,VulDeePecker,BenchSG, andDevign) as our base models. The experimental results show that usingCOCLcan bring an improvement of 8.1% to the F1 scores of these models on a real-world vulnerability dataset. Shihan Dou, Yueming Wu 0001, Yang Liu 0003 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | DSRM: Boost Textual Adversarial Training with Distribution Shift Risk MinimizationabstractSongYang Gao, Shihan Dou, Yan Liu, Xiao Wang, Qi Zhang, Zhongyu Wei, Jin Ma, Ying Shan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Songyang Gao, Shihan Dou, Yan Liu 0002, Xiao Wang 0001, Qi Zhang 0001, Zhongyu Wei, Jin Ma 0003, Ying Shan |
ACL (1) | 2 |
| 2023 | Gitor: Scalable Code Clone Detection by Building Global Sample GraphabstractCode clone detection is about finding out similar code fragments, which has drawn much attention in software engineering since it is important for software maintenance and evolution. Researchers have proposed many techniques and tools for source code clone detection, but current detection methods concentrate on analyzing or processing code samples individually without exploring the underlying connections among code samples. Junjie Shan, Shihan Dou, Yueming Wu 0001, Hairu Wu, Yang Liu 0003 |
ESEC/SIGSOFT FSE | 2 |
| 2022 | MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic PerspectiveabstractXiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang, Tao Gui, Liang Qiao, Zhanzhan Cheng, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xiao Wang 0001, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang 0001, Tao Gui, Liang Qiao 0001, Zhanzhan Cheng, Xuanjing Huang 0001 |
ACL (1) | 2 |
| 2022 | Decorrelate Irrelevant, Purify Relevant: Overcome Textual Spurious Correlations from a Feature PerspectiveabstractNatural language understanding (NLU) models tend to rely on spurious correlations (i.e., dataset bias) to achieve high performance on in-distribution datasets but poor performance on out-of-distribution ones. Most of the existing debiasing methods often identify and weaken these samples with biased features (i.e., superficial surface features that cause such spurious correlations). However, down-weighting these samples obstructs the model in learning from the non-biased parts of these samples. To tackle this challenge, in this paper, we propose to eliminate spurious correlations in a fine-grained manner from a feature space perspective. Specifically, we introduce Random Fourier Features and weighted re-sampling to decorrelate the dependencies between features to mitigate spurious correlations. After obtaining decorrelated features, we further design a mutual-information-based method to purify them, which forces the model to learn features that are more relevant to tasks. Extensive experiments on two well-studied NLU tasks demonstrate that our method is superior to other comparative approaches. Shihan Dou, Songyang Gao, Junjie Shan, Qi Zhang 0001, Yueming Wu 0001, Xuanjing Huang 0001 |
COLING | 1 |
| 2022 | Kernel-Whitening: Overcome Dataset Bias with Isotropic Sentence EmbeddingabstractDataset bias has attracted increasing attention recently for its detrimental effect on the generalization ability of fine-tuned models.The current mainstream solution is designing an additional shallow model to pre-identify biased instances.However, such two-stage methods scale up the computational complexity of training process and obstruct valid feature information while mitigating bias.To address this issue, we utilize the representation normalization method which aims at disentangling the correlations between features of encoded sentences.We find it also promising in eliminating the bias problem by providing isotropic data distribution.We further propose Kernel-Whitening, a Nyström kernel approximation method to achieve more thorough debiasing on nonlinear spurious correlations.Our framework is end-to-end with similar time consumption to fine-tuning.Experiments show that Kernel-Whitening significantly improves the performance of BERT on out-of-distribution datasets while maintaining in-distribution accuracy. Songyang Gao, Shihan Dou, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 2 |
| 2022 | VulCNN: An Image-inspired Scalable Vulnerability Detection SystemabstractSince deep learning (DL) can automatically learn features from source code, it has been widely used to detect source code vulnerability. To achieve scalable vulnerability scanning, some prior studies intend to process the source code directly by treating them as text. To achieve accurate vulnerability detection, other approaches consider distilling the program semantics into graph representations and using them to detect vulnerability. In practice, text-based techniques are scalable but not accurate due to the lack of program semantics. Graph-based methods are accurate but not scalable since graph analysis is typically time-consuming. Yueming Wu 0001, Deqing Zou, Shihan Dou, Wei Yang 0013, Hai Jin 0001 |
ICSE | 3 |
| 2021 | IntDroid: Android Malware Detection Based on API Intimacy AnalysisabstractAndroid, the most popular mobile operating system, has attracted millions of users around the world. Meanwhile, the number of new Android malware instances has grown exponentially in recent years. On the one hand, existing Android malware detection systems have shown that distilling the program semantics into a graph representation and detecting malicious programs by conducting graph matching are able to achieve high accuracy on detecting Android malware. However, these traditional graph-based approaches always perform expensive program analysis and suffer from low scalability on malware detection. On the other hand, because of the high scalability of social network analysis, it has been applied to complete large-scale malware detection. However, the social-network-analysis-based method only considers simple semantic information (i.e., centrality) for achieving market-wide mobile malware scanning, which may limit the detection effectiveness when benign apps show some similar behaviors as malware. In this article, we aim to combine the high accuracy of traditional graph-based method with the high scalability of social-network-analysis--based method for Android malware detection. Instead of using traditional heavyweight static analysis, we treat function call graphs of apps as complex social networks and apply social-network--based centrality analysis to unearth the central nodes within call graphs. After obtaining the central nodes, the average intimacies between sensitive API calls and central nodes are computed to represent the semantic features of the graphs. We implement our approach in a tool called IntDroid and evaluate it on a dataset of 3,988 benign samples and 4,265 malicious samples. Experimental results show that IntDroid is capable of detecting Android malware with an F-measure of 97.1% while maintaining a True-positive Rate of 99.1%. Although the scalability is not as fast as a social-network-analysis--based method (i.e., MalScan ), compared to a traditional graph-based method, IntDroid is more than six times faster than MaMaDroid . Moreover, in a corpus of apps collected from GooglePlay market, IntDroid is able to identify 28 zero-day malware that can evade detection of existing tools, one of which has been downloaded and installed by more than ten million users. This app has also been flagged as malware by six anti-virus scanners in VirusTotal, one of which is Symantec Mobile Insight . Deqing Zou, Yueming Wu 0001, Siru Yang, Anki Chauhan, Wei Yang 0013, Jiangying Zhong, Shihan Dou, Hai Jin 0001 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2020 | SCDetector: Software Functional Clone Detection Based on Semantic Tokens AnalysisabstractCode clone detection is to find out code fragments with similar functionalities, which has been more and more important in software engineering. Many approaches have been proposed to detect code clones, in which token-based methods are the most scalable but cannot handle semantic clones because of the lack of consideration of program semantics. To address the issue, researchers conduct program analysis to distill the program semantics into a graph representation and detect clones by matching the graphs. However, such approaches suffer from low scalability since graph matching is typically time-consuming. Yueming Wu 0001, Deqing Zou, Shihan Dou, Siru Yang, Wei Yang 0013, Hai Jin 0001 |
ASE | 3 |