VLDB 2026 Research / reviewers in the wild / expert
Hao Wang 0182
dblp:181/2812-182
· DBLP profile ↗
18ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-6959-7237ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PersuHSG: Adaptive Persuasion Strategy Planning for Dialogue Agents Based on Hierarchical Strategy GraphabstractPersuasion, a vital social skill, influences beliefs, attitudes, and behaviors through conversation. Yet, current dialogue agents either rely on scenario-specific strategies, restricting their cross-context adaptability, or neglect persuasion’s logical structure. They focus on isolated strategy classification, overlooking the significance of fine-grained sequential planning for real-world scenarios. To address these limitations, inspired by basic human mental activities, we present PersuHSG, an adaptive persuasion strategy planning framework. The core idea is to conceptualize persuasion as a tripartite framework comprising cognition, affection, and volition, with each stage represented as a graph layer and principle-based strategies for efficient multi-stage persuasion. Specifically, we first develop PersuInstruct, a fine-tuning dataset to improve dialogue agents’ strategic planning and response generation. Then, we propose a graph-aware planning algorithm for stage-strategy-response reasoning to generate persuasive responses for diverse scenarios. Extensive experiments confirm that PersuHSG significantly enhances the persuasiveness of Large Language Models (LLMs), allows smaller models (e.g., 9B, 13B) to achieve competitive performance, and demonstrates the efficacy of structured strategy planning in improving model efficiency and adaptability. Bin Guo 0001, Hao Wang 0182, Jingqi Liu, Yan Liu 0045, Yunji Liang, Yan Pan 0003, Zhiwen Yu 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2025 | HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-trainingabstractAdvances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set. Heli Sun, Danlei Huang, Xinyi Yin, Hao Wang 0182, Jia Zhang 0016, Fei Wang 0128, Peihao Guo, Suyu Xing, Junxiao Xue, Liang He 0006 |
ACM Multimedia | 6 |
| 2025 | The future of cognitive strategy-enhanced persuasive dialogue agents: new perspectives and trendsabstractAbstract Persuasion, as one of the crucial abilities in human communication, has garnered extensive attention from researchers within the field of intelligent dialogue systems. Developing dialogue agents that can persuade others to accept certain standpoints is essential to achieving truly intelligent and anthropomorphic dialogue systems. Benefiting from the substantial progress of Large Language Models (LLMs), dialogue agents have acquired an exceptional capability in context understanding and response generation. However, as a typical and complicated cognitive psychological system, persuasive dialogue agents also require knowledge from the domain of cognitive psychology to attain a level of human-like persuasion. Consequently, the cognitive strategy-enhanced persuasive dialogue agent (defined as CogAgent ), which incorporates cognitive strategies to achieve persuasive targets through conversation, has become a predominant research paradigm. To depict the research trends of CogAgent, in this paper, we first present several fundamental cognitive psychology theories and give the formalized definition of three typical cognitive strategies, including the persuasion strategy, the topic path planning strategy, and the argument structure prediction strategy. Then we propose a new system architecture by incorporating the formalized definition to lay the foundation of CogAgent. Representative works are detailed and investigated according to the combined cognitive strategy, followed by the summary of authoritative benchmarks and evaluation metrics. Finally, we summarize our insights on open issues and future directions of CogAgent for upcoming researchers. Bin Guo 0001, Hao Wang 0182, Jingqi Liu, Yasan Ding, Yan Pan 0003, Zhiwen Yu 0001 |
Frontiers Comput. Sci. | 3 |
| 2025 | Cascade context-oriented spatio-temporal attention network for efficient and fine-grained video-grounded dialogues
Hao Wang 0182, Bin Guo 0001, Qiuyun Zhang, Yasan Ding, Ying Zhang 0047, Zhiwen Yu 0001 |
Frontiers Comput. Sci. | 1 |
| 2025 | EvolveDetector: Towards an evolving fake news detector for emerging events with continual knowledge accumulation and transfer
Yasan Ding, Bin Guo 0001, Yan Liu 0045, Yao Jing, Maolong Yin, Hao Wang 0182, Zhiwen Yu 0001 |
Inf. Process. Manag. | 7 |
| 2025 | Design Graph Guided Element Importance-Aware Layout Generation With Multimodality Cascade TransformerabstractGraphic designs are pervasive in our daily lives and widely used to communicate information hierarchically to humans. To achieve this, the layout plays an essential role in guiding readers to understand the importance of different elements and comprehend the content. To deal with the rapidly increasing demands of graphic designs, recent studies attempt to automatically generate layouts based on category information and spatial relations, often resulting in layouts with poor communication quality. In this article, we make the first attempt to explore element importance-aware layout generation under the guidance of a novel design graph, which attracts readers’ attention to a layout by formulating aesthetic relations implicitly involved in graphic designs between element pairs. The core of our approach is a learning-based framework with a new multimodality cascade transformer (MCT) in a coarse-to-fine manner. A hierarchical multimodality fusion (HMF) mechanism and two new losses are introduced to guide the training process progressively. We further collect a new fine-grained advertisement poster layout dataset containing more than 30 K layouts labeled with 91 element labels. Both qualitative and quantitative experiments demonstrate the effectiveness of our approach against existing works. We also conduct user studies and cognitive experiments to evaluate the direct adaptability and attractiveness of generated layouts. Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Xiaotian Qiao, Hao Wang 0182, Ying Zhang 0047, Zhiwen Yu 0001 |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2025 | Enabling Harmonious Human-Machine Interaction with Visual-Context Augmented Dialogue System: A ReviewabstractThe intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually complex human-computer interaction requirements, it is difficult for traditional text-based dialogue system to meet the demands for more vivid and convenient interaction. Consequently, Visual-Context Augmented Dialogue (VAD) System, which has the potential to communicate with humans by perceiving and understanding multimodal information (i.e., visual context in images or videos, textual dialogue history), has become a predominant research paradigm. Benefiting from the consistency and complementarity between visual and textual context, VAD possesses the potential to generate engaging and context-aware responses. To depict the development of VAD, we first characterize the concept model of VAD and then present its generic system architecture to illustrate the system workflow, followed by a summary of multimodal fusion techniques. Subsequently, several research challenges and representative works are investigated, followed by the summary of authoritative benchmarks and real-world application of VAD. We conclude this article by putting forward some open issues and promising research trends for VAD, e.g., the cognitive mechanisms of human-machine dialogue under cross-modal dialogue context, mobile and lightweight deployment of VAD. Hao Wang 0182, Bin Guo 0001, Yating Zeng, Yasan Ding, Ying Zhang 0047, Lina Yao 0001, Zhiwen Yu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Memory-Enhanced Emotional Support Conversations with Motivation-Driven Strategy Inference
Hao Wang 0182, Bin Guo 0001, Yasan Ding, Qiuyun Zhang, Ying Zhang 0047, Zhiwen Yu 0001 |
ECML/PKDD (5) | 1 |
| 2024 | TD3-Based Collaborative Computation Offloading and Charging Scheduling in Multi-UAV-Assisted MEC NetworksabstractComputation offloading, resource allocation, and endurance issues in unmanned aerial vehicle (UAV)-aided mobile edge computing (MEC) networks have always been a research focus. UAV-aided MEC allows mobile users (MUs)' tasks to be offloaded to drones for processing in special scenarios, such as natural disasters or military attacks. However, as the number and size of offloaded tasks increase, a single UAV is difficult to meet all computational demands, result in the decline of QoS. To address this issue, this paper presents a collaborative computation offloading scheme where multiple UAVs can cooperate to handle massive tasks. Firstly, considering that battery-limited UAVs cannot complete all tasks and sustain flight without charging, we incorporate charging stations (CS) into multi-UAV-assisted MEC networks. Subsequently, we design a price-based incentive mechanism to maximize the total revenue obtained from UAVs' collaborative computation. Then, we formulate the joint optimization problem of computation offloading, resource allocation and charging scheduling as a Markov Decision Process (MDP), and propose a Twin Delayed Deep Deterministic policy gradient (TD3) algorithm to find optimal strategies. Finally, extensive simulations demonstrate that the proposed TD3 algorithm outperforms other benchmark methods, achieving the highest overall system utility under different scenarios. Liang Zhao 0014, Yujun Yao, Huan Zhou 0002, Hao Wang 0182, Victor C. M. Leung |
WCNC | 4 |
| 2024 | AdaMEC: Towards a Context-adaptive and Dynamically Combinable DNN Deployment Framework for Mobile Edge ComputingabstractWith the rapid development of deep learning, recent research on intelligent and interactive mobile applications (e.g., health monitoring, speech recognition) has attracted extensive attention. And these applications necessitate the mobile edge computing scheme, i.e., offloading partial computation from mobile devices to edge devices for inference acceleration and transmission load reduction. The current practices have relied on collaborative DNN partition and offloading to satisfy the predefined latency requirements, which is intractable to adapt to the dynamic deployment context at runtime. AdaMEC, a context-adaptive and dynamically combinable DNN deployment framework, is proposed to meet these requirements for mobile edge computing, which consists of three novel techniques. First, once-for-all DNN pre-partition divides DNN at the primitive operator level and stores partitioned modules into executable files, defined as pre-partitioned DNN atoms. Second, context-adaptive DNN atom combination and offloading introduces a graph-based decision algorithm to quickly search the suitable combination of atoms and adaptively make the offloading plan under dynamic deployment contexts. Third, runtime latency predictor provides timely latency feedback for DNN deployment considering both DNN configurations and dynamic contexts. Extensive experiments demonstrate that AdaMEC outperforms state-of-the-art baselines in terms of latency reduction by up to 62.14% and average memory saving by 55.21%. Sicong Liu 0005, Bin Guo 0001, Yuzhan Wang, Hao Wang 0182, Zhenli Sheng, Zhiwen Yu 0001 |
ACM Trans. Sens. Networks | 6 |
| 2024 | Federated Distributed Deep Reinforcement Learning for Recommendation-Enabled Edge CachingabstractRecently, in response to the low efficiency and high transmission latency of traditional centralized content delivery networks, especially in congested scenarios, edge caching has emerged as a promising method to bring content caching closer to the edge of the network. However, traditional content delivery methods might still lead to low utilization of cache resources. To tackle this challenge, this paper investigates a content recommendation-based edge caching method in multi-tier edge-cloud networks while considering content delivery and cache replacement decisions as well as bandwidth allocation strategies. First, we consider a multi-tier edge caching-enabled content delivery network architecture combined with a content recommendation system and formulate the optimization problem with the objective of minimizing long-term content delivery delay and maximizing cache hit rate. Second, considering time-varying system environments and uncertain content demands, we approximate the optimization process of content delivery and cache replacement for each agent as a Partially Observable Markov Decision Process (POMDP) and propose a single-agent Deep Deterministic Policy Gradient (DDPG)-based method. Subsequently, we extend the POMDP to a multi-agent scenario. To address the issue of agents converging to local optima and establish more personalized models, we propose a Federated Distributed DDPG-based method (FD3PG) to solve the corresponding problem in a multi-agent system. Finally, simulation results demonstrate that the proposed FD3PG achieves lower delivery delay and higher cache hit rate compared with other baselines in various scenarios. Specifically, compared with FADE, MADRL, and DDPG, FD3PG achieves a significant decrease in average delivery delay, approximately 10%, 11%, and 35% on the Synthetic dataset, and 12%, 14%, and 48% on the MovieLens Latest Small dataset, respectively. Huan Zhou 0002, Hao Wang 0182, Zhiwen Yu 0001, Bin Guo 0001, Mingjun Xiao, Jie Wu 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | PiercingEye: Identifying Both Faint and Distinct Clues for Explainable Fake News Detection with Progressive Dynamic Graph MiningabstractExplainability is crucial for the successful use of AI for fake news detection (FND). Researchers aim to improve the explainability of FND by highlighting important descriptions in crowd-contributed comments as clues. From the perspective of law and sociology, there are distinct clues that are easy to discover and understand, and faint clues that require careful observation and analysis. For example, in fake news related to COVID-Omicron showing increased pathogenicity and transmissibility, distinct clues might involve virologists’ opinions regarding the inverse correlation between pathogenicity and transmissibility. Meanwhile, faint clues might be reflected in an infected person’s claim that the symptoms are milder than a cold (indirectly indicating reduced pathogenicity). Occasionally, the statements of some ordinary eyewitnesses can decisively reveal the truth of the news, leading to the judgment of fake news. Existing methods generally use static networks to model the entire news life-cycle, which makes it fail to capture the subtle dynamic interactions between individual clues and news. Thereby faint clues, whose relations to the truth of news are challenging to be characterized and extracted directly, are more likely to be overshadowed by distinct clues. To address this issue, we propose an explainable FND method, dubbed as PiercingEye, which leverages dynamic interaction information to progressively mine valuable clues. PiercingEye models the news propagation topology as a dynamic graph, with interactive comments serving as nodes, and employs the time-semantic encoding mechanism to refine the modeling of temporal interaction information between comments and news to preserve faint clues. Subsequently, it utilizes the self-attention mechanism to aggregate distinct and faint clues for FND. Experimental results demonstrate that PiercingEye outperforms state-of-the-art methods and is capable of identifying both faint and distinct clues for humans to debunk fake news. Yasan Ding, Bin Guo 0001, Yan Liu 0045, Hao Wang 0182, Haocheng Shen, Zhiwen Yu 0001 |
ECAI | 4 |
| 2023 | Towards Informative and Diverse Dialogue Systems Over Hierarchical Crowd Intelligence Knowledge GraphabstractKnowledge-enhanced dialogue systems aim at generating factually correct and coherent responses by reasoning over knowledge sources, which is a promising research trend. The truly harmonious human-agent dialogue systems need to conduct engaging conversations from three aspects as humans, namely (1) stating factual contents (e.g., records in Wikipedia), (2) conveying subjective and informative opinions about objects (e.g., user discussions on Twitter), and (3) impressing interlocutors with diverse expression styles (e.g., personalized expression habits). The existing knowledge base is a standardized and unified coding for factual knowledge, which could not portray the other two kinds of knowledge to make responses more informative and expressive diverse. To address this, we present CrowdDialog , a crowd intelligence knowledge-enhanced dialogue system, which takes advantage of “crowd intelligence knowledge” extracted from social media (with rich subjective descriptions and diversified expression styles) to promote the performance of dialogue systems. Firstly, to thoroughly mine and organize the crowd intelligence knowledge underlying large-scale and unstructured online contents, we elaborately design the C rowd I ntelligence K nowledge G raph ( CIKG ) structure, including the domain commonsense subgraph, descriptive subgraph, and expressive subgraph. Secondly, to reasonably integrate heterogeneous crowd intelligence knowledge into responses while ensuring logicality and fluency, we propose the G ated F usion with D ynamic Knowledge- D ependent ( GFDD ) model, which generates responses from the semantic and syntactic perspective with the context-aware knowledge gate and dynamic knowledge decoding. Finally, extensive experiments over both Chinese and English dialogue datasets demonstrate that our approach GFDD outperforms competitive baselines in terms of both automatic evaluation and human judgments. Besides, ablation studies indicate that the proposed CIKG has the potential to promote dialogue systems to generate fluent, informative, and diverse dialogue responses. Hao Wang 0182, Bin Guo 0001, Jiaqi Liu 0002, Yasan Ding, Zhiwen Yu 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | Towards information-rich, logical dialogue systems with knowledge-enhanced neural models
Hao Wang 0182, Bin Guo 0001, Wei Wu 0014, Sicong Liu 0005, Zhiwen Yu 0001 |
Neurocomputing | 1 |
| 2021 | Conditional Text Generation for Harmonious Human-Machine InteractionabstractIn recent years, with the development of deep learning, text-generation technology has undergone great changes and provided many kinds of services for human beings, such as restaurant reservation and daily communication. The automatically generated text is becoming more and more fluent so researchers begin to consider more anthropomorphic text-generation technology, that is, the conditional text generation, including emotional text generation, personalized text generation, and so on. Conditional Text Generation (CTG) has thus become a research hotspot. As a promising research field, we find that much attention has been paid to exploring it. Therefore, we aim to give a comprehensive review of the new research trends of CTG. We first summarize several key techniques and illustrate the technical evolution route in the field of neural text generation, based on the concept model of CTG. We further make an investigation of existing CTG fields and propose several general learning models for CTG. Finally, we discuss the open issues and promising research directions of CTG. Bin Guo 0001, Hao Wang 0182, Yasan Ding, Wei Wu 0014, Shaoyang Hao, Yueqi Sun, Zhiwen Yu 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | DeepDepict: Enabling Information Rich, Personalized Product Description Generation With the Deep Multiple Pointer Generator NetworkabstractIn e-commerce platforms, the online descriptive information of products shows significant impacts on the purchase behaviors. To attract potential buyers for product promotion, numerous workers are employed to write the impressive product descriptions. The hand-crafted product descriptions are less-efficient with great labor costs and huge time consumption. Meanwhile, the generated product descriptions do not take consideration into the customization and the diversity to meet users’ interests. To address these problems, we propose one generic framework, namely DeepDepict, to automatically generate the information-rich and personalized product descriptive information. Specifically, DeepDepict leverages the graph attention to retrieve the product-related knowledge from external knowledge base to enrich the diversity of products, constructs the personalized lexicon to capture the linguistic traits of individuals for the personalization of product descriptions, and utilizes multiple pointer-generator network to fuse heterogeneous data from multi-sources to generate informative and personalized product descriptions. We conduct intensive experiments on one public dataset. The experimental results show that DeepDepict outperforms existing solutions in terms of description diversity, BLEU, and personalized degree with significant margin gain, and is able to generate product descriptions with comprehensive knowledge and personalized linguistic traits. Shaoyang Hao, Bin Guo 0001, Hao Wang 0182, Yunji Liang, Lina Yao 0001, Qianru Wang, Zhiwen Yu 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | MateBot: The Design of a Human-Like, Context-Sensitive Virtual Bot for Harmonious Human-Computer Interaction
Bin Guo 0001, Hao Wang 0182, Helei Cui, Zhiwen Yu 0001 |
GPC | 3 |
| 2015 | Lightweight Management of Resource-Constrained Sensor Devices in Internet of ThingsabstractIt is predicted that billions of intelligent devices and networks, such as wireless sensor networks (WSNs), will not be isolated but connected and integrated with computer networks in future Internet of Things (IoT). In order to well maintain those sensor devices, it is often necessary to evolve devices to function correctly by allowing device management (DM) entities to remotely monitor and control devices without consuming significant resources. In this paper, we propose a lightweight RESTful Web service (WS) approach to enable device management of wireless sensor devices. Specifically, motivated by the recent development of IPv6-based open standards for accessing wireless resource-constrained networks, we consider to implement IPv6 over low-power wireless personal area network (6LoWPAN)/routing protocol for low power and lossy network (RPL)/constrained application protocol (CoAP) protocols on sensor devices and propose a CoAP-based DM solution to allow easy access and management of IPv6 sensor devices. By developing a prototype cloud system, we successfully demonstrate the proposed solution in efficient and effective management of wireless sensor devices. Zhengguo Sheng, Hao Wang 0182, Changchuan Yin, Xiping Hu, Shusen Yang, Victor C. M. Leung |
IEEE Internet Things J. | 2 |