Dacheng Wen

dblp:344/1944 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
abstract
State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the partial results to produce verdicts with justifications (explanations for the verdicts). The security of these systems is crucial, as compromised fact-checkers can amplify misinformation, but remains largely underexplored. To bridge this gap, this work introduces a novel threat model against such fact-checking systems and presents Fact2Fiction, the first poisoning attack framework targeting SOTA agentic fact-checking systems. Fact2Fiction employs LLMs to mimic the decomposition strategy and exploit system-generated justifications to craft tailored malicious evidences that compromise sub-claim verification. Extensive experiments demonstrate that Fact2Fiction achieves 8.9%-21.2% higher attack success rates than SOTA attacks across various poisoning budgets and exposes security weaknesses in existing fact-checking systems, highlighting the need for defensive countermeasures.
Haorui He, Yupeng Li 0001, Bin B. Zhu, Dacheng Wen, Reynold Cheng, Francis C. M. Lau 0001
AAAI4
2026 RLSeek: Evidence-Grounded Reasoning for RAG Hallucination Detection
abstract
Zhaoheng Huang, Dacheng Wen, Yutao Zhu, Xiaoying Lian, Yushi Liang, Kai Hao, Nan Li, Liangjie Zhang, Qi Zhang, Ji-Rong Wen, Zhicheng Dou, Fangzhao Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaoheng Huang, Dacheng Wen, Yutao Zhu 0001, Xiaoying Lian, Yushi Liang, Kai Hao, Liangjie Zhang, Qi Zhang 0001, Ji-Rong Wen, Zhicheng Dou, Fangzhao Wu
ACL (1)2
2026 Near-Optimal Online Learning with Non-Stochastic and Unbounded Erroneous Feedback
Dacheng Wen, Yupeng Li 0001, Francis C. M. Lau 0001, Tian Wang 0001, Yang Chen 0001
INFOCOM1
2026 Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents
abstract
State-of-the-art single-agent claim verification methods struggle with complex claims that require nuanced analysis of multifaceted evidence. Inspired by real-world professional fact-checkers, we propose DebateCV, the first debate-driven claim verification framework powered by multiple LLM agents. In DebateCV, two Debaters argue opposing stances to surface subtle errors in single-agent assessments. A decisive Moderator is then required to weigh the evidential strength of conflicting arguments to deliver an accurate verdict. Yet, zero-shot Moderators are biased toward neutral judgments, and no datasets exist for training them. To bridge this gap, we propose Debate-SFT, a post-training framework that leverages synthetic data to enhance agents' ability to effectively adjudicate debates for claim verification. Results show that our methods surpass state-of-the-art non-debate approaches in both accuracy (across various evidence conditions) and justification quality.
Haorui He, Yupeng Li 0001, Dacheng Wen, Yang Chen 0001, Reynold Cheng, Donald Donglong Chen, Francis C. M. Lau 0001
WWW3
2026 Robust Decentralized Online Learning Against Targeted and Untargeted Malicious Data Feature Manipulation
abstract
Motivated by real-world applications, we study the problem of decentralized online learning with dynamic feedback delays in the presence of malicious data generators under different threat models. In this problem, multiple agents collaborate to classify the features of streaming data samples generated online and receive dynamically delayed feedback on the ground-truth labels. While some data generators are benign, others—due to internal motives or external factors such as cyberattacks—may maliciously manipulate data features to compromise the classification performance. In this work, we first investigate the targeted attacks by malicious data generators, i.e., feature manipulation with aims to gain preferred classification outcomes from the agents. In response, we propose two robust algorithms,RDOC-TOandRDOC-TC, countering ordinary and clairvoyant adversaries that can access certain outdated and the latest classification models of the agents, respectively. Subsequently, we address the untargeted attacks by malicious data generators, which aim to disrupt the classification outcomes without targeting any particular class, by proposing another algorithm,RDOC-U. Our theoretical analysis establishes that all three proposed algorithms achieve sublinear regret bounds. The evaluations conducted in the application of network traffic classification with two real-world datasets demonstrate the competitiveness of the proposed algorithms compared to advanced baselines.
Yupeng Li 0001, Dacheng Wen, Mengjia Xia, Mingzhe Chen, Xiaoming Fu 0001
IEEE Trans. Mob. Comput.2
2026 Fairness-Aware Online Pricing for Profit Maximization in Ride-Sharing
abstract
Ride-sharing represents a sustainable transportation paradigm that is beneficial to human, society, and environment. Common ride-sharing pricing approaches determine prices for riders through optimizing one or more figures of merit, e.g., the profit or revenue. However, they overlook an important issue—fairness—which, when perceived by the riders, can critically affect their degree of satisfaction. In this work, we take the initiative to consider an intuitive and appropriate notion of individual fairness called fairness-in-hindsight for riders in ride-sharing pricing. We study the problem of online fair pricing of shared rides (which allow multiple riders to share one ride) with an aim to maximize the profit of the ride-sharing operator/platform. We design an online fair ride-sharing pricing algorithm called OnFairRP, which comprises phases oflearning, transition, and exploitation. We prove that OnFairRP has a sub-linear regret bound and can guarantee the fairness between riders. Our extensive performance evaluations using real-world data traces of ride-sharing demonstrate the advantages of OnFairRP over benchmarking schemes including commonly used methods with or without fairness guarantee.
Yupeng Li 0001, Mengjia Xia, Dacheng Wen, Francis C. M. Lau 0001, Shunbo Lei, Zhaocheng Huang
IEEE Trans. Netw.3
2024 Augment Decentralized Online Convex Optimization with Arbitrarily Bad Machine-Learned Predictions
abstract
Decentralized online convex optimization (DOCO), as a pivotal computational paradigm in machine learning, has been applied to many critical tasks. However, existing DOCO algorithms, due to their excessive emphasis on the worst-case theoretical performance, appear to be overly cautious in making decisions across all possible cases, especially in real-world applications where the worst cases actually hardly occur. Therefore, these existing approaches typically are limited in performance in practice. To avoid such pessimistic strategies, we propose to study the approach of augmenting DOCO with machine-learned predictions that can guide the decision-making process. We present an overview of the problem along with the preliminary results and outlook in this work.
Dacheng Wen, Yupeng Li 0001, Francis C. M. Lau 0001
ICDCS1
2024 Robust Decentralized Online Optimization Against Malicious Agents
abstract
Decentralized online optimization, a pivotal paradigm in machine learning, involves multiple agents making online decisions cooperatively in a decentralized network. Despite its outstanding capabilities in processing large-scale streaming data, the ubiquitous existence of malicious agents, capable of disseminating arbitrary information among their neighbors and undetectable a priori, poses a severe threat to the reliability and efficacy of existing decentralized online optimization solutions. In response to the above critical vulnerability in practice, we take the first step to properly address the threat posed by malicious agents. We propose ROOO, a novel robust decentralized online optimization algorithm, specifically designed to counteract the detrimental impact of malicious agents. Our theoretical analysis shows that the regret bound of ROOO is sub-linear, indicating that, over time, its performance progressively approximates that of an offline oracle operating with the benefit of hindsight. Empirical evaluations in two networking applications, including opportunistic channel selection and mobile crowdsensing, further validate our theoretical results and demonstrate the competitiveness of ROOO compared to several advanced baselines.
Dacheng Wen, Yupeng Li 0001, Xiaoxi Zhang 0001, Francis C. M. Lau 0001
ICDCS1
2024 Augment Online Linear Optimization with Arbitrarily Bad Machine-Learned Predictions
abstract
The online linear optimization paradigm is important to many real-world network applications as well as theoretical algorithmic studies. Recent studies have made attempts to augment online linear optimization with machine-learned predictions of the cost function that are meant to improve the performance of the algorithms. However, they fail to address the critical case in practical systems where the predictions can be arbitrarily bad. In this work, we take the first step to study the problem of online linear optimization with a dynamic number of arbitrarily bad machine-learned predictions per round and propose an algorithm termed OLOAP. Our theoretical analysis shows that, when the qualities of the predictions are satisfactory, OLOAP achieves a regret bound of O(logT), which circumvents the tight lower bound of Ω($\sqrt T $) for the vanilla problem of online linear optimization (i.e., the one without any predictions). Meanwhile, the regret of our algorithm is never worse than O($\sqrt T $) irrespective of the qualities of predictions. In addition, we further derive a lower bound for the regret of the studied problem, which demonstrates that OLOAP is near-optimal. We consider two important network applications and conduct extensive evaluations. Our results validate the superiority of our algorithm over state-of-the-art approaches.
Dacheng Wen, Yupeng Li 0001, Francis C. M. Lau 0001
INFOCOM1
2024 MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection
abstract
The prevalence of fake news across various online sources has had a significant influence on the public. Existing Chinese fake news detection datasets are limited to news sourced solely from Weibo. However, fake news originating from multiple sources exhibits diversity in various aspects, including its content and social context. Methods trained on purely one single news source can hardly be applicable to real-world scenarios. Our pilot experiment demonstrates that the F1 score of the state-of-the-art method that learns from a large Chinese fake news detection dataset, Weibo-21, drops significantly from 0.943 to 0.470 when the test data is changed to multi-source news data, failing to identify more than one-third of the multi-source fake news. To address this limitation, we constructed the first multi-source benchmark dataset for Chinese fake news detection, termed MCFEND, which is composed of news we collected from diverse sources such as social platforms, messaging apps, and traditional online news outlets. Notably, such news has been fact-checked by 14 authoritative fact-checking agencies worldwide. In addition, various existing Chinese fake news detection methods are thoroughly evaluated on our proposed dataset in cross-source, multi-source, and unseen source ways. MCFEND, as a benchmark dataset, aims to advance Chinese fake news detection approaches in real-world scenarios.
Yupeng Li 0001, Haorui He, Dacheng Wen
WWW4
2024 Message Injection Attack on Rumor Detection under the Black-Box Evasion Setting Using Large Language Model
abstract
Recent analyses have disclosed that existing rumor detection techniques, despite playing a pivotal role in countering the dissemination of misinformation on social media, are vulnerable to both white-box and surrogate-based black-box adversarial attacks. However, such attacks depend heavily on unrealistic assumptions, e.g., modifiable user data and white-box access to the rumor detection models, or appropriate selections of surrogate models, which are impractical in the real world. Thus, existing analyses fail to uncover the robustness of rumor detectors in practice. In this work, we take a further step towards the investigation about the robustness of existing rumor detection solutions. Specifically, we focus on the state-of-the-art rumor detectors, which leverage graph neural network based models to predict whether a post is rumor based on the Message Propagation Tree (MPT), a conversation tree with the post as its root and the replies to the post as the descendants of the root. We propose a novel black-box attack method, HMIA-LLM, against these rumor detectors, which uses the Large Language Model to generate malicious messages and inject them into the targeted MPTs. Our extensive evaluation conducted across three rumor detection datasets, four target rumor detectors, and three baselines for comparison demonstrates the effectiveness of our proposed attack method in compromising the performance of the state-of-the-art rumor detectors.
Yifeng Luo, Yupeng Li 0001, Dacheng Wen, Liang Lan
WWW3
2024 A Survey of Machine Learning-Based Ride-Hailing Planning
abstract
Ride-hailing is a sustainable transportation paradigm where riders access door-to-door traveling services through a mobile phone application, which has attracted a colossal amount of usage. There are two major planning tasks in a ride-hailing system: 1) matching, i.e., assigning available vehicles to pick up the riders; and 2) repositioning, i.e., proactively relocating vehicles to certain locations to balance the supply and demand of ride-hailing services. Recently, many studies of ride-hailing planning that leverage machine learning techniques have emerged. In this article, we present a comprehensive overview on latest developments of machine learning-based ride-hailing planning. To offer a clear and structured review, we introduce a taxonomy into which we carefully fit the different categories of related works according to the types of their planning tasks and solution schemes, which include collective matching, distributed matching, collective repositioning, distributed repositioning, and joint matching and repositioning. We further shed light on many real-world data sets and simulators that are indispensable for empirical studies on machine learning-based ride-hailing planning strategies. At last, we propose several promising research directions for this rapidly growing research and practical field.
Dacheng Wen, Yupeng Li 0001, Francis C. M. Lau 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Contextual Target-Specific Stance Detection on Twitter: Dataset and Method
abstract
To understand different aspects of online human behaviors, e.g., the public stances toward various social and political issues, contextual target-specific stance detection has become one of the most important studies on social media. Considering the lack of appropriate data for the studies of contextual target-specific stance detection on Twitter, which is one of the most popular online social platforms worldwide, we introduce CTSDT, a new dataset that consists of a large number of annotated target-specific conversations collected from Twitter. Furthermore, we propose a new contextual target-specific stance detection model called ConMulAttn, which is the first method that can learn both the contents of the posts and the concrete relationships between the posts in a conversation. We conduct extensive evaluation using CTSDT as well as another two popular datasets, CreateDebate and ConvinceMe, for contextual target-specific stance detection. The evaluation results validate the necessity of introducing our dataset CTSDT. Besides, according to the evaluation results, our proposed model ConMulAttn can outperform the state-of-the-art contextual target-specific stance detection method by up to 25% in F1score, indicating the effectiveness and superiority of our solution. Our study has the potential to assist policymakers in utilizing conversation data from online social platforms to efficiently gain real-time insights into public stances on target topics, such as vaccination.
Yupeng Li 0001, Dacheng Wen, Haorui He, Jianxiong Guo, Xuan Ning, Francis C. M. Lau 0001
ICDM2
2023 Robust Decentralized Online Learning against Malicious Data Generators and Dynamic Feedback Delays with Application to Traffic Classification
abstract
Motivated by the real-world application of traffic classification at the network edge, we study the problem of robust decentralized online learning against malicious data generators that can manipulate their data features with an aim to gain preferred classification outcomes. Multiple agents cooperatively learn classification models to make online decisions. They periodically exchange their models, e.g., traffic classification models, between neighbors in a decentralized network and update local model parameters on the fly based on the models they have access to and feedback on the observed local data samples that are dynamically delayed. In this work, we propose two decentralized online learning algorithms, RDOC-O and RDOC-C, respectively against ordinary malicious and clairvoyant malicious data generators. Our theoretical performance analysis shows that the two algorithms have provable sub-linear individual regret bounds under mild conditions. To validate our analysis, extensive performance evaluations are conducted in the application of network traffic classification using two real-world data traces. Our results show that the two proposed algorithms compare favorably with an optimal offline classification model in the presence of malicious data generators, and they can achieve a steady-state F1score of around 0.85, which validates their effectiveness and makes them appealing in practice.
Yupeng Li 0001, Dacheng Wen, Mengjia Xia
SECON2