Shihan Wang 0001

dblp:173/9268-1 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0001-5971-7522ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite Individuals
abstract
Deep Reinforcement Learning (DRL) is widely used in task-oriented dialogue systems to optimize dialogue policy, but it struggles to balance exploration and exploitation due to the high dimensionality of state and action spaces.This challenge often results in local optima or poor convergence.Evolutionary Algorithms (EAs) have been proven to effectively explore the solution space of neural networks by maintaining population diversity.Inspired by this, we innovatively combine the global search capabilities of EA with the local optimization of DRL to achieve a balance between exploration and exploitation.Nevertheless, the inherent flexibility of natural language in dialogue tasks complicates this direct integration, leading to prolonged evolutionary times.Thus, we further propose an elite individual injection mechanism to enhance EA's search efficiency by adaptively introducing best-performing individuals into the population.Experiments across four datasets show that our approach significantly improves the balance between exploration and exploitation, boosting performance.Moreover, the effectiveness of the EII mechanism in reducing exploration time has been demonstrated, achieving an efficient integration of EA and DRL on task-oriented dialogue policy tasks.
Libo Qin 0001, Shihan Wang 0001
ACL (1)4
2025 Reducing Variance Caused by Communication in Decentralized Multi-agent Deep Reinforcement Learning
Changxi Zhu, Mehdi Dastani, Shihan Wang 0001
AAMAS3
2025 Sparse communication in multi-agent deep reinforcement learning
Mehdi Dastani, Shihan Wang 0001
Neurocomputing3
2025 Communication with factorized policy gradients in multi-agent deep reinforcement learning
abstract
Abstract In multi-agent deep reinforcement learning (MADRL), agents can learn to communicate to broaden their view and understanding of the environment and their teammates. Previous works on communication in MADRL mainly rely on centralized or independent value functions for learning communication, which cannot differentiate how communicating agents individually contribute to the overall learning process. Moreover, continuous environments that incorporate continuous state/action spaces have received limited attention in previous research. In this paper, we propose a novel architecture for communicating agents and apply centralized but factorized value functions to differentiate how each agent contributes to learning during communication, along with gradient backpropagation. Additionally, to address the complexity introduced by communication, we investigate the use of an attention mechanism that aggregates messages, enabling policies to maintain a fixed input length. We then present a new policy gradient method termed communication with factorized policy gradients (CFPG), featuring full backpropagation from factorized value functions to communicating agents’ architecture. We demonstrate that CFPG can enhance performance and accelerate learning in continuous predator–prey scenarios and multi-agent MuJoCo, when compared to other learning communication methods.
Changxi Zhu, Mehdi Dastani, Shihan Wang 0001
Neural Comput. Appl.3
2024 Learning Reward Structure with Subtasks in Reinforcement Learning
abstract
Improving sample efficiency of Reinforcement Learning (RL) in sparse-reward environments poses a significant challenge. In scenarios where the reward structure is complex, accurate action evaluation often relies heavily on precise information about past achieved subtasks and their order. Previous approaches have often failed or proved inefficient in constructing and leveraging such intricate reward structures. In this work, we propose an RL algorithm that can automatically structure the reward function for sample efficiency, given a set of labels that signify subtasks. Given such minimal knowledge about the task, we train a high-level policy that selects optimal subtasks in each state together with a low-level policy that efficiently learns to complete each sub-task. We evaluate our algorithm in a variety of sparse-reward environments. The experiment results show that our method significantly outperforms the state-of-art baselines as the difficulty of the task increases.
Mehdi Dastani, Shihan Wang 0001
ECAI3
2024 Bootstrapped Policy Learning for Task-oriented Dialogue through Goal Shaping
abstract
Reinforcement learning shows promise in optimizing dialogue policies, but addressing the challenge of reward sparsity remains crucial.While curriculum learning offers a practical solution by strategically training policies from simple to complex, it hinges on the assumption of a gradual increase in goal difficulty to ensure a smooth knowledge transition across varied complexities.In complex dialogue environments without intermediate goals, achieving seamless knowledge transitions becomes tricky.This paper proposes a novel Bootstrapped Policy Learning (BPL) framework, which adaptively tailors progressively challenging subgoal curriculum for each complex goal through goal shaping, ensuring a smooth knowledge transition.Goal shaping involves goal decomposition and evolution, decomposing complex goals into subgoals with solvable maximum difficulty and progressively increasing difficulty as the policy improves.Moreover, to enhance BPL's adaptability across various environments, we explore various combinations of goal decomposition and evolution within BPL, and identify two universal curriculum patterns that remain effective across different dialogue environments, independent of specific environmental constraints.By integrating the summarized curriculum patterns, our BPL has exhibited efficacy and versatility across four publicly available datasets with different difficulty levels.
Mehdi Dastani, Shihan Wang 0001
EMNLP4
2024 A survey of multi-agent deep reinforcement learning with communication
abstract
Abstract Communication is an effective mechanism for coordinating the behaviors of multiple agents, broadening their views of the environment, and to support their collaborations. In the field of multi-agent deep reinforcement learning (MADRL), agents can improve the overall learning performance and achieve their objectives by communication. Agents can communicate various types of messages, either to all agents or to specific agent groups, or conditioned on specific constraints. With the growing body of research work in MADRL with communication (Comm-MADRL), there is a lack of a systematic and structural approach to distinguish and classify existing Comm-MADRL approaches. In this paper, we survey recent works in the Comm-MADRL field and consider various aspects of communication that can play a role in designing and developing multi-agent reinforcement learning systems. With these aspects in mind, we propose 9 dimensions along which Comm-MADRL approaches can be analyzed, developed, and compared. By projecting existing works into the multi-dimensional space, we discover interesting trends. We also propose some novel directions for designing future Comm-MADRL systems through exploring possible combinations of the dimensions.
Changxi Zhu, Mehdi Dastani, Shihan Wang 0001
Auton. Agents Multi Agent Syst.3
2024 Correction: A survey of multi-agent deep reinforcement learning with communication
abstract
guideline of Comm-MADRL systems. The guideline positions dimensions where communication influences interaction with the environment and training phases.
Changxi Zhu, Mehdi Dastani, Shihan Wang 0001
Auton. Agents Multi Agent Syst.3
2024 Viewpoint: Hybrid Intelligence Supports Application Development for Diabetes Lifestyle Management
abstract
Type II diabetes is a complex health condition requiring patients to closely and continuously collaborate with healthcare professionals and other caretakers on lifestyle changes. While intelligent products have tremendous potential to support such Diabetes Lifestyle Management (DLM), existing products are typically conceived from a technology-centered perspective that insufficiently acknowledges the degree to which collaboration and inclusion of stakeholders is required. In this article, we argue that the emergent design philosophy of Hybrid Intelligence (HI) forms a suitable alternative lens for research and development. In particular, we (1) highlight a series of pragmatic challenges for effective AI-based DLM support based on results from an expert focus group, and (2) argue for HI’s potential to address these by outlining relevant research trajectories.
Bernd Dudzik, Jasper van der Waa, Roel Dobbe, Inago M. D. R. de Troya, Roos M. Bakker, Maaike de Boer, Quirine T. S. Smit, Davide Dell'Anna, Emre Erdogan, Pinar Yolum, Shihan Wang 0001, Selene Baez, Lea Krause, Bart Kamphorst
J. Artif. Intell. Res.12
2024 Rescue Conversations from Dead-ends: Efficient Exploration for Task-oriented Dialogue Policy Optimization
abstract
Abstract Training a task-oriented dialogue policy using deep reinforcement learning is promising but requires extensive environment exploration. The amount of wasted invalid exploration makes policy learning inefficient. In this paper, we define and argue that dead-end states are important reasons for invalid exploration. When a conversation enters a dead-end state, regardless of the actions taken afterward, it will continue in a dead-end trajectory until the agent reaches a termination state or maximum turn. We propose a Dead-end Detection and Resurrection (DDR) method that detects dead-end states in an efficient manner and provides a rescue action to guide and correct the exploration direction. To prevent dialogue policies from repeating errors, DDR also performs dialogue data augmentation by adding relevant experiences that include dead-end states and penalties into the experience pool. We first validate the dead-end detection reliability and then demonstrate the effectiveness and generality of the method across various domains through experiments on four public dialogue datasets.
Mehdi Dastani, Jinchuan Long, Zhenyu Wang 0001, Shihan Wang 0001
Trans. Assoc. Comput. Linguistics5
2024 Decomposed Deep Q-Network for Coherent Task-Oriented Dialogue Policy Learning
abstract
Reinforcement learning (RL) has emerged as a key technique for designing dialogue policies. However, action space inflation in dialogue tasks has led to a heavy decision burden and incoherence problems for dialogue policies. In this paper, we propose a novel decomposed deep Q-network (D2Q) that exploits the natural structure of dialogue actions to perform decomposition on Q-function, realizing efficient and coherent dialogue policy learning. Instead of directly evaluating the Q-function, it consists of two separate estimators, one for the abstract action-value functions and the other for the specific action-value functions, both sharing a common feature layer. The abstract action-value function determines the speech act of the system action, while the specific action-value function focuses on the concrete action. This structure establishes a logical relationship between the user and the system on speech actions, avoiding the problem of incoherence. Moreover, the abstract action-value function shields unreasonable specific actions in the inflated action space, reducing the decision complexity. Our results show that the problem of incoherence is prevalent in existing approaches, which significantly impacts the efficiency and quality of dialogue policy learning. Our D2Q architecture alleviates this problem and performs significantly better than competitive baselines in both evaluated and human experiments. Further experiments validate the generality of our method. It can be easily extended to other RL-based dialogue policy approaches.
Zhenyu Wang 0001, Mehdi Dastani, Shihan Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Extracting Stances on Pandemic Measures from Social Media Data
abstract
Support for national measures against the COVID-19 pandemic can be measured by collecting responses to questionnaires. In this paper we explore a less costly and less time-consuming method: by analyzing social media data. We compare stances regarding anti-pandemic measures extracted from tweets with the questionnaire results provided by the Dutch national health institute RIVM. We find similarities and differences and discuss the results.
Erik F. Tjong Kim Sang, Shihan Wang 0001, Marijn Schraagen, Mehdi Dastani
e-Science2
2021 Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory Policy
abstract
Deep reinforcement learning has shown great potential in training dialogue policies.However, its favorable performance comes at the cost of many rounds of interaction.Most of the existing dialogue policy methods rely on a single learning system, while the human brain has two specialized learning and memory systems, supporting to find good solutions without requiring copious examples.Inspired by the human brain, this paper proposes a novel complementary policy learning (CPL) framework, which exploits the complementary advantages of the episodic memory (EM) policy and the deep Q-network (DQN) policy to achieve fast and effective dialogue policy learning.In order to coordinate between the two policies, we proposed a confidence controller to control the complementary time according to their relative efficacy at different stages.Furthermore, memory connectivity and time pruning are proposed to guarantee the flexible and adaptive generalization of the EM policy in dialog tasks.Experimental results on three dialogue datasets show that our method significantly outperforms existing methods relying on a single learning system.
Zhenyu Wang 0001, Changxi Zhu, Shihan Wang 0001
EMNLP (1)4
2020 METNet: A Mutual Enhanced Transformation Network for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to determine the sentiment polarity of each specific aspect in a given sentence.Existing researches have realized the importance of the aspect for the ABSA task and have derived many interactive learning methods that model context based on specific aspect.However, current interaction mechanisms are ill-equipped to learn complex sentences with multiple aspects, and these methods underestimate the representation learning of the aspect.In order to solve the two problems, we propose a mutual enhanced transformation network (METNet) for the ABSA task.First, the aspect enhancement module in METNet improves the representation learning of the aspect with contextual semantic features, which gives the aspect more abundant information.Second, METNet designs and implements a hierarchical structure, which enhances the representations of aspect and context iteratively.Experimental results on SemEval 2014 Datasets demonstrate the effectiveness of METNet, and we further prove that METNet is outstanding in multi-aspect scenarios.
Bin Jiang 0006, Wanyue Zhou, Chao Yang 0015, Shihan Wang 0001, Liang Pang 0001
COLING5
2020 PEDNet: A Persona Enhanced Dual Alternating Learning Network for Conversational Response Generation
abstract
Endowing a chatbot with a personality is essential to deliver more realistic conversations.Various persona-based dialogue models have been proposed to generate personalized and diverse responses by utilizing predefined persona information.However, generating personalized responses is still a challenging task since the leverage of predefined persona information is often insufficient.To alleviate this problem, we propose a novel Persona Enhanced Dual Alternating Learning Network (PEDNet) aiming at producing more personalized responses in various opendomain conversation scenarios.PEDNet consists of a Context-Dominated Network (CDNet) and a Persona-Dominated Network (PDNet), which are built upon a common encoder-decoder backbone.CDNet learns to select a proper persona as well as ensure the contextual relevance of the predicted response, while PDNet learns to enhance the utilization of persona information when generating the response by weakening the disturbance of specific content in the conversation context.CDNet and PDNet are trained alternately using a multi-task training approach to equip PEDNet with the both capabilities they have learned.Both automatic and human evaluations on a newly released dialogue dataset Persona-chat demonstrate that our method could deliver more personalized responses than baseline methods.
Bin Jiang 0006, Wanyue Zhou, Jingxu Yang, Chao Yang 0015, Shihan Wang 0001, Liang Pang 0001
COLING5
2020 STBins: Visual Tracking and Comparison of Multiple Data Sequences Using Temporal Binning
abstract
While analyzing multiple data sequences, the following questions typically arise: how does a single sequence change over time, how do multiple sequences compare within a period, and how does such comparison change over time. This paper presents a visual technique named STBins to answer these questions. STBins is designed for visual tracking of individual data sequences and also for comparison of sequences. The latter is done by showing the similarity of sequences within temporal windows. A perception study is conducted to examine the readability of alternative visual designs based on sequence tracking and comparison tasks. Also, two case studies based on real-world datasets are presented in detail to demonstrate usage of our technique.
Vincent Bloemen, Shihan Wang 0001, Jarke J. van Wijk, Huub van de Wetering
IEEE Trans. Vis. Comput. Graph.3
2017 Early Signals of Trending Rumor Event in Streaming Social Media
abstract
In this study, we propose a mechanism for identifying early signals of trending rumor events (i.e. controversial emerging topics) in streaming social media. The pattern, combining features of both user's attitude and information diffusion, is applied in the sliding windows of social media data streams. By capturing and analyzing frequent patterns within early windows, we found signal patterns appearing at very early stages of trending rumor events (in average, months before their peak time). Our preliminary empirical analysis is applied in two different Twitter datasets. The obtained results indicate the potential of our approach to detect trending rumor event candidates (with high probability of being false) as early as possible in real-time environments.
Shihan Wang 0001, Izabela Moise, Dirk Helbing, Takao Terano
COMPSAC (2)1
2015 Detecting rumor patterns in streaming social media
abstract
Rumor detection in streaming social media is a significant but challenging problem. In this paper, we present a method to identify rumor patterns in the streaming social media environment. Patterns which combine both structural and behavioral properties of rumor are firstly proposed to distinguish false rumors from valid news. A novel graph-based pattern matching algorithm is also described to detect rumor patterns from streaming social media data. Compared within Twitter data of rumors and non-rumors, our selected rumor patterns contain distinct properties of rumors in short-term series.
Shihan Wang 0001, Takao Terano
IEEE BigData1