Dawei Li 0008

dblp:13/5856-8 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0003-4139-7841ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or Harm
abstract
Agents based on Large Language Models (LLMs) have demonstrated strong capabilities across a wide range of tasks. However, deploying LLM-based agents in high-stakes domains comes with significant safety and ethical risks. Unethical behavior by these agents can directly result in serious real-world consequences, including physical harm and financial loss. To efficiently steer the ethical behavior of agents, we frame agent behavior steering as a model editing task, which we term Behavior Editing. Model editing is an emerging area of research that enables precise and efficient modifications to LLMs while preserving their overall capabilities. To systematically study and evaluate this approach, we introduce BehaviorBench, a multi-tier benchmark grounded in psychological moral theories. This benchmark supports both the evaluation and editing of agent behaviors across a variety of scenarios, with each tier introducing more complex and ambiguous scenarios. We first demonstrate that Behavior Editing can dynamically steer agents toward the target behavior within specific scenarios. Moreover, Behavior Editing enables not only scenario-specific local adjustments but also more extensive shifts in an agent’s global moral alignment. We demonstrate that Behavior Editing can be used to promote ethical and benevolent behavior or, conversely, to induce harmful or malicious behavior. Through extensive evaluations of agents built on frontier LLMs, BehaviorBench validates the effectiveness of behavior editing across a wide range of models and scenarios. Our findings offer key insights into a new paradigm for steering agent behavior, highlighting both the promise and perils of Behavior Editing.
Baixiang Huang, Zhen Tan 0001, Haoran Wang 0005, Dawei Li 0008, Ali Payani, Huan Liu 0001, Tianlong Chen 0001, Kai Shu
AAAI5
2026 Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
abstract
Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps.However, excessive or redundant reasoning -so-called overthinking -can increase inference costs and lead LLMs toward incorrect conclusions.In this paper, we present REFRAIN (REFlective-Redundancy for Adaptive INference), a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking.REFRAIN integrates a two-stage stop discriminator to identify reflective yet redundant reasoning and a sliding-window Upper Confidence Bound (SW-UCB) multi-armed bandit controller to dynamically adjust stopping thresholds according to problem difficulty without supervision or fine-tuning.Across four representative benchmarks and two model families, REFRAIN reduces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting.Extensive ablation and robustness analyses demonstrate its stability across models, scorers, and prompt variations.In summary, our findings highlight when-tostop as a new and practical axis of test-time scaling -enabling models to reason not just more, but just enough.
Renliang Sun, Wei Cheng 0002, Dawei Li 0008, Wei Wang 0010
ACL (1)3
2025 Can LLMs Improve Multimodal Fact-Checking by Asking Relevant Questions?
Alimohammad Beigi, Bohan Jiang, Dawei Li 0008, Zhen Tan 0001, Pouya Shaeri, Tharindu Kumarage, Amrita Bhattacharjee, Huan Liu 0001
IEEE Big Data3
2025 Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
Liangjie Huang, Dawei Li 0008, Huan Liu 0001, Lu Cheng 0001
IEEE Big Data2
2025 Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
Dawei Li 0008, Yue Huang 0001, Ming Li 0010, Tianyi Zhou 0001, Xiangliang Zhang 0001, Huan Liu 0001
CIKM1
2025 Building Safer Sites: A Large-Scale Multi-Level Dataset for Construction Safety Benchmark
abstract
Construction safety research is a critical field in civil engineering, aiming to mitigate risks and prevent injuries through the analysis of site conditions and human factors. However, the limited volume and lack of diversity in existing construction safety datasets pose significant challenges to conducting in-depth analyses. To address this research gap, this paper introduces the Construction Safety Dataset (CSDataset), a well-organized comprehensive multi-level dataset that encompasses incidents, inspections, and violations recorded sourced from the Occupational Safety and Health Administration (OSHA). This dataset uniquely integrates structured attributes with unstructured narratives, facilitating a wide range of approaches driven by machine learning and large language models. We also conduct a preliminary approach benchmarking and various cross-level analyses using our dataset, offering insights to inform and enhance future efforts in construction safety. For example, we found that complaint-driven inspections were associated with a 17.3% reduction in the likelihood of subsequent incidents. Our dataset and code are released at https://github.com/zhenhuiou/Construction-Safety-Dataset-CSDataset.
Zhenhui Ou, Dawei Li 0008, Zhen Tan 0001, Huan Liu 0001, Siyuan Song
CIKM2
2025 From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
abstract
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, Huan Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Dawei Li 0008, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan 0001, Amrita Bhattacharjee, Canyu Chen, Kai Shu, Lu Cheng 0001, Huan Liu 0001
EMNLP1
2025 BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
abstract
Sizhe Wang, Yongqi Tong, Hengyuan Zhang, Dawei Li, Xin Zhang, Tianlong Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yongqi Tong, Dawei Li 0008, Tianlong Chen 0001
NAACL (Long Papers)4
2025 SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents
Dawei Li 0008, Zhen Tan 0001, Peijia Qian, Kumar Satvik Chaudhary, Lijie Hu
PAKDD (3)1
2024 Large Language Models for Data Annotation and Synthesis: A Survey
abstract
Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, Huan Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhen Tan 0001, Dawei Li 0008, Song Wang 0013, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng 0001, Huan Liu 0001
EMNLP2
2023 Multi-level Contrastive Learning for Script-based Character Understanding
abstract
In this work, we tackle the scenario of understanding characters in scripts, which aims to learn the characters' personalities and identities from their utterances.We begin by analyzing several challenges in this scenario, and then propose a multi-level contrastive learning framework to capture characters' global information in a fine-grained manner.To validate the proposed framework, we conduct extensive experiments on three character understanding sub-tasks by comparing with strong pretrained language models, including SpanBERT, Longformer, BigBird and ChatGPT-3.5.Experimental results demonstrate that our method improves the performances by a considerable margin.Through further in-depth analysis, we show the effectiveness of our method in addressing the challenges and provide more hints on the scenario of character understanding.We will open-source our work in this URL....
Dawei Li 0008, Yanran Li
EMNLP1