Hao Wang 0193

dblp:181/2812-193 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0004-0199-0488ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
abstract
Hao Wang, Hao Gu, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue, Sirui Han, Yike Guo, Dapeng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Wang 0193, Hao Gu 0001, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue 0001, Sirui Han, Yike Guo, Dapeng Oliver Wu
ACL (1)1
2026 Multi-Task-Oriented Emergency-Aware UAV Crowdsensing: A Hierarchical Multi-Agent Deep Reinforcement Learning Approach
abstract
Integrated sensing and communication (ISAC) has emerged as a transformative paradigm, merging the capabilities of sensing and communication to enhance efficiency and enable advanced applications. Mobile crowdsensing (MCS), as a important example of ISAC, leverages unmanned vehicles such as UAVs to continuously gather and transmit environmental data, supporting critical applications like traffic monitoring, urban congestion management, and accident investigation. In this paper, we focus on multi-task-oriented UAV crowdsensing (UCS), where diverse tasks—such as surveillance and emergency response—each have distinct age-of-information (AoI) requirements. We introduce a novel metric, the “valid task handling index,” to evaluate the performance of handling multiple tasks effectively. Our proposed hierarchical multi-agent deep reinforcement learning (MADRL) framework, DRL-MTUCS, integrates seamlessly with multi-agent actor-critic reinforcement learning methods. It features dynamically weighted queues for UAV goal assignment, enabling efficient management of multiple emergency tasks, and a low-level UAV execution module with a self-balancing intrinsic reward mechanism. This ensures all tasks are completed within their individual AoI constraints. Extensive experiments and trajectory visualizations validate the superior performance and robustness of DRL-MTUCS compared to six baselines across varying conditions, including the number of UAVs, surveillance task AoI thresholds, and emergency task image blur requirements.
Chi Harold Liu, Hao Wang 0193, Guangpeng Qi, Zhongyi Liu 0002, Dapeng Oliver Wu
IEEE J. Sel. Areas Commun.3
2026 A Framework of Knowledge Graph-Enhanced Large Language Model Based on Global Planning
abstract
Knowledge graphs (KGs) can provide structured knowledge to assist large language models (LLMs) in interpretable reasoning. Knowledge graph question answering (KGQA) is a typical benchmark to evaluate KG-enhanced LLM methods. Previous methods of KG-enhanced LLMs for KGQA mainly include: 1) origin question-oriented methods, which perform KG retrieval based solely on the original question without explicitly analyzing multi-step reasoning logic; and 2) stepwise reasoning-oriented methods, which alternate between LLM generating the next reasoning step and targeted KG retrieval but lack systematic planning, leading to poor controllability. To tackle these limitations, we propose KELGoP, a framework of KG-enhanced LLM based on global planning. We propose fine-grained question categorization based on reasoning patterns and corresponding category-driven question decomposition for complex questions, enabling more controllable reasoning and atomic KG retrieval targeted to sub-questions. Furthermore, we propose an adaptive strategy that allows adjusting the reasoning pattern based on the performance of question answering, making the reasoning more flexible and robust. Finally, we introduce several efficient atomic KG retrieval strategies that operate on KG subgraphs to assist the LLM in answering atomic-level questions. A series of experiments on KGQA datasets demonstrate that our proposed framework achieves superior performance compared to existing baselines.
Yading Li, Dandan Song 0005, Yuhang Tian 0002, Hao Wang 0193, Changzhi Zhou, Shuhao Zhang 0001
IEEE Trans. Knowl. Data Eng.4
2026 Indoor Fingerprint Collection Under Environment Changes by Vehicular Crowdsensing: A Bayesian Reinforcement Learning Approach
abstract
Indoor localization is crucial for applications such as navigation, asset tracking, and emergency response. Fingerprint-based methods that use RSSI are widely adopted; however, they fail under large environmental changes. Unmanned Vehicles (UVs) equipped with high precision sensors are able to collect fingerprints, serving as a promising way by forming a Vehicular Crowdsensing (VCS) campaign. In this paper, we propose “BRAVE”, a Bayesian RL Approach for VCS under Environment changing, while introducing a new metric “Calibration Benefit” to explicitly quantify how effectively a learned trajectory updates those regions of the fingerprint database that have changed and matter most for localization. Specifically, we propose a spatial-temporal Bayesian Network(BN) for change detection, a region rearrangement method for fewer restarts, and an optimistic strategy to balance the exploration and exploitation trade-offs in optimizing calibration benefit. Extensive results on two real-world datasets from SML Center (Shanghai) and Haopu Fashion City (Shanghai) demonstrate that BRAVE outperforms eight baselines and the derived dataset has better localization accuracy compared with the original dataset.
Haoming Yang, Chi Harold Liu, Guozheng Li 0002, Hao Wang 0193, Jianxin Zhao 0001, Guangpeng Qi, Dapeng Oliver Wu
IEEE Trans. Mob. Comput.4
2025 A$^3$E: Towards Compositional Model Editing
abstract
Model editing has become a *de-facto* practice to address hallucinations and outdated knowledge of large language models (LLMs). However, existing methods are predominantly evaluated in isolation, i.e., one edit at a time, failing to consider a critical scenario of compositional model editing, where multiple edits must be integrated and jointly utilized to answer real-world multifaceted questions. For instance, in medical domains, if one edit informs LLMs that COVID-19 causes "fever" and another that it causes "loss of taste", a qualified compositional editor should enable LLMs to answer the question "What are the symptoms of COVID-19?" with both "fever" and "loss of taste" (and potentially more). In this work, we define and systematically benchmark this compositional model editing (CME) task, identifying three key undesirable issues that existing methods struggle with: *knowledge loss*, *incorrect preceding* and *knowledge sinking*. To overcome these issues, we propose A$^3$E, a novel compositional editor that (1) ***a**daptively combines and **a**daptively regularizes* pre-trained foundation knowledge in LLMs in the stage of edit training and (2) ***a**daptively merges* multiple edits to better meet compositional needs in the stage of edit composing. Extensive experiments demonstrate that A$^3$E improves the composability by at least 22.45\% without sacrificing the performance of non-compositional model editing.
Hongming Piao, Hao Wang 0193, Dapeng Oliver Wu, Ying Wei 0001
NeurIPS2
2024 Indoor Periodic Fingerprint Collections by Vehicular Crowdsensing via Primal-Dual Multi-Agent Deep Reinforcement Learning
abstract
Indoor localization is drawing more and more attentions due to the growing demand of various location-based services, where fingerprinting is a popular data driven techniques that does not rely on complex measurement equipment, yet it requires site surveys which is both labor-intensive and time-consuming. Vehicular crowdsensing (VCS) with unmanned vehicles (UVs) is a novel paradigm to navigate a group of UVs to collect sensory data from certain point-of-interests periodically (PoIs, i.e., coverage holes in localization scenarios). In this paper, we formulate the multi-floor indoor fingerprint collection task with periodical PoI coverage requirements as a constrained optimization problem. Then, we propose a multi-agent deep reinforcement learning (MADRL) based solution, “MADRL-PosVCS”, which consists of a primal-dual framework to transform the above optimization problem into the unconstrained duality, with adjustable Lagrangian multipliers to ensure periodic fingerprint collection. We also propose a novel intrinsic reward mechanism consists of the mutual information between a UV’s observations and environment transition probability parameterized by a Bayesian Neural Network (BNN) for exploration, and a elevator-based reward to allow UVs to go cross different floors for collaborative fingerprint collections. Extensive simulation results on three real-world datasets in SML Center (Shanghai), Joy City (Hangzhou) and Haopu Fashion City (Shanghai) show that MADRL-PosVCS achieves better results over four baselines on fingerprint collection ratio, PoI coverage ratio for collection intervals, geographic fairness and average moving distance.
Haoming Yang, Qiran Zhao, Hao Wang 0193, Chi Harold Liu, Guozheng Li 0002, Guoren Wang, Jian Tang 0008, Dapeng Oliver Wu
IEEE J. Sel. Areas Commun.3
2024 QoI-Aware Mobile Crowdsensing for Metaverse by Multi-Agent Deep Reinforcement Learning
abstract
Metaverse is expected to provide mobile users with emerging applications both in regular situation like intelligent transportation services and in emergencies like wireless search and disaster response. These applications are usually associated with stringent quality-of-information (QoI) requirements like throughput and age-of-information (AoI), which can be further guaranteed by using unmanned aerial vehicles (UAVs) as aerial base stations (BSs) to compensate the existing 5G infrastructures. In this paper, we consider a new QoI-aware mobile crowdsensing (MCS) campaign by UAVs which move around and collect data from mobile users wearing metaverse devices. Specifically, we propose “MetaCS”, a multi-agent deep reinforcement learning (MADRL) framework with improvements on a Transformer-based user mobility prediction module between regions and a relational graph learning mechanism to enable the selection of most informative partners to communicate for each UAV. Extensive results and trajectory visualizations on three real mobility datasets in NCSU, KAIST and Beijing show that MetaCS consistently outperforms six baselines in terms of overall QoI index, when varying different numbers of UAVs, throughput requirement, and AoI threshold.
Yuxiao Ye, Hao Wang 0193, Chi Harold Liu, Zipeng Dai, Guozheng Li 0002, Guoren Wang, Jian Tang 0008
IEEE J. Sel. Areas Commun.2
2024 HiBid: A Cross-Channel Constrained Bidding System With Budget Allocation by Hierarchical Offline Deep Reinforcement Learning
abstract
Online display advertising platforms service numerous advertisers by providing real-time bidding (RTB) for the scale of billions of ad requests every day. The bidding strategy handles ad requests cross multiple channels to maximize the number of clicks under the set financial constraints, i.e., total budget and cost-per-click (CPC), etc. Different from existing works mainly focusing on single channel bidding, we explicitly consider cross-channel constrained bidding with budget allocation. Specifically, we propose a hierarchical offline deep reinforcement learning (DRL) framework called “HiBid”, consisted of a high-level planner equipped with auxiliary loss for non-competitive budget allocation, and a data augmentation enhanced low-level executor for adaptive bidding strategy in response to allocated budgets. Additionally, a CPC-guided action selection mechanism is introduced to satisfy the cross-channel CPC constraint. Through extensive experiments on both the large-scale log data and online A/B testing, we confirm that HiBid outperforms six baselines in terms of the number of clicks, CPC satisfactory ratio, and return-on-investment (ROI). We also deploy HiBid on Meituan advertising platform to already service tens of thousands of advertisers every day.
Hao Wang 0193, Bo Tang 0018, Chi Harold Liu, Shangqin Mao, Jiahong Zhou, Zipeng Dai, Yaqi Sun, Qianlong Xie, Dong Wang 0022
IEEE Trans. Computers1
2024 Ensuring Threshold AoI for UAV-Assisted Mobile Crowdsensing by Multi-Agent Deep Reinforcement Learning With Transformer
abstract
Unmanned aerial vehicle (UAV) crowdsensing (UCS) is an emerging data collection paradigm to provide reliable and high quality urban sensing services, with age-of-information (AoI) requirement to measure data freshness in real-time applications. In this paper, we explicitly consider the case to ensure that the attained AoI always stay within a specific threshold. The goal is to maximize the total amount of collected data from diverse Point-of-Interests (PoIs) while minimizing AoI and AoI threshold violation ratio under limited energy supplement. To this end, we propose a decentralized multi-agent deep reinforcement learning framework called “DRL-UCS($\text {AoI}_{th}$)” for multi-UAV trajectory planning, which consists of a novel transformer-enhanced distributed architecture and an adaptive intrinsic reward mechanism for spatial cooperation and exploration. Extensive results and trajectory visualization on two real-world datasets in Beijing and San Francisco show that, DRL-UCS($\text {AoI}_{th}$) consistently outperforms all nine baselines when varying the number of UAVs, AoI threshold and generated data amount in a timeslot.
Hao Wang 0193, Chi Harold Liu, Haoming Yang, Guoren Wang, Kin K. Leung
IEEE/ACM Trans. Netw.1
2021 Constrained Route Planning over Large Multi-Modal Time-Dependent Networks
abstract
Constrained route planning (CRP) on transportation networks has been extensively studied because of its broad applications, such as route recommendation. However, the existing works on CRP neglect the time-dependent and multi-modal properties of transportation networks. This paper proposes an approach for CRP over multi-modal time-dependent networks. Specifically, we design two novel constrained route planning algorithms, function-dependent routing and labeling-index-based routing. While function-dependent routing generates an accurate route to CRP by traversing the network, labeling-index-based one ensures the fast response with the support of an efficient index and the compression scheme of networks. In order to demonstrate the efficiency and effectiveness of our proposed algorithms, experiments are performed over real datasets.
Yishu Wang 0001, Ye Yuan 0001, Hao Wang 0193, Xiangmin Zhou, Congcong Mu, Guoren Wang
ICDE3
2021 Mobile Crowdsensing for Data Freshness: A Deep Reinforcement Learning Approach
abstract
Data collection by mobile crowdsensing (MCS) is emerging as data sources for smart city applications, however how to ensure data freshness has sparse research exposure but quite important in practice. In this paper, we consider to use a group of mobile agents (MAs) like UAVs and driverless cars which are equipped with multiple antennas to move around in the task area to collect data from deployed sensor nodes (SNs). Our goal is to minimize the age of information (AoI) of all SNs and energy consumption of MAs during movement and data upload. To this end, we propose a centralized deep reinforcement learning (DRL)-based solution called "DRL-freshMCS" for controlling MA trajectory planning and SN scheduling. We further utilize implicit quantile networks to maintain the accurate value estimation and steady policies for MAs. Then, we design an exploration and exploitation mechanism by dynamic distributed prioritized experience replay. We also derive the theoretical lower bound for episodic AoI. Extensive simulation results show that DRL-freshMCS significantly reduces the episodic AoI per remaining energy, compared to five baselines when varying different number of antennas and data upload thresholds, and number of SNs. We also visualize their trajectories and AoI update process for clear illustrations.
Zipeng Dai, Hao Wang 0193, Chi Harold Liu, Rui Han 0001, Jian Tang 0008, Guoren Wang
INFOCOM2
2021 Energy-Efficient 3D Vehicular Crowdsourcing for Disaster Response by Distributed Deep Reinforcement Learning
abstract
Fast and efficient access to environmental and life data is key to the successful disaster response. Vehicular crowdsourcing (VC) by a group of unmanned vehicles (UVs) like drones and unmanned ground vehicles to collect these data from Point-of-Interests (PoIs) e.g., possible survivor spots and fire site, provides an efficient way to assist disaster rescue. In this paper, we explicitly consider to navigate a group of UVs in a 3-dimensional (3D) disaster workzone to maximize the amount of collected data, geographical fairness, energy efficiency, while minimizing data dropout due to limited transmission rate. We propose DRL-DisasterVC(3D), a distributed deep reinforcement learning framework, with a repetitive experience replay (RER) to improve learning efficiency, and a clipped target network to increase learning stability. We also use a 3D convolutional neural network (3D CNN) with multi-head-relational attention (MHRA) for spatial modeling, and add auxiliary pixel control (PC) for spatial exploration. We designed a novel disaster response simulator, called "DisasterSim", and conduct extensive experiments to show that DRL-DisasterVC(3D) outperforms all five baselines in terms of energy efficiency when varying the numbers of UVs, PoIs and SNR threshold.
Hao Wang 0193, Chi Harold Liu, Zipeng Dai, Jian Tang 0008, Guoren Wang
KDD1