Rui Wang 0079

dblp:06/2293-79 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0001-5369-9116ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
abstract
Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely based on unlabeled test data may induce prompt optimization bias, ultimately leading to suboptimal performance on downstream tasks. In this work, we analyze the underlying causes of prompt optimization bias from both the model and data perspectives. In terms of the model, the entropy minimization objective typically focuses on reducing the entropy of model predictions while overlooking their correctness. This can result in overconfident yet incorrect outputs, thereby compromising the quality of prompt optimization. On the data side, prompts affected by optimization bias can introduce misalignment between visual and textual modalities, which further aggravates the prompt optimization bias. To this end, we propose a Doubly Debiased Test-Time Prompt Tuning method, abbreviated as D2TPT. Specifically, we first introduce a dynamic retrieval-augmented modulation module that retrieves high-confidence knowledge from a dynamic knowledge base using the test image feature as a query, and uses the retrieved knowledge to modulate the predictions. Guided by the refined predictions, we further develop a reliability-aware prompt optimization module that incorporates a confidence-based weighted ensemble and cross-modal consistency distillation to impose regularization constraints during prompt tuning. Extensive experiments across 15 benchmark datasets involving both natural distribution shifts and cross-datasets generalization demonstrate that D2TPT outperforms baselines, validating its effectiveness in mitigating prompt optimization bias.
Rui Wang 0079, Jiahuan Zhou, Changwen Zheng, Jiangmeng Li
AAAI3
2026 M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference
abstract
Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.
Chuxiong Sun, Qirui Ji, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Wei Wang 0353
AAAI6
2026 TMAE: Learning Targeted Multi-Agent Exploration via Causal Inference
abstract
Exploration in sparse-reward tasks remains a fundamental challenge in multi-agent reinforcement learning (MARL) due to complex inter-agent interactions and the expansive exploration space. To address this issue, we propose Targeted Multi-Agent Exploration (TMAE), a novel framework that uncovers the causal relationships between the state space and the reward function, thereby reducing the exploration space and enabling more targeted exploration. Specifically, we construct a structural causal model (SCM) to model the causality between sub-state variables and sparse rewards, providing a robust analytical foundation for subsequent causal inference. Through counterfactual causal intervention, TMAE identifies the most critical subspaces for discovering rare but pivotal events while filtering out confounders. By incorporating these causal insights into the exploration process, TMAE prioritizes subspaces with stronger causal effects on sparse rewards, significantly enhancing exploration efficiency. We evaluate TMAE on a range of MARL benchmarks featuring sparse rewards, consistently demonstrating superior exploration efficiency compared to state-of-the-art methods. Furthermore, visualized causal insights derived from TMAE reveal its ability to effectively capture intricate dependencies and priorities in targeted exploration, showcasing strong alignment with prior domain knowledge.
Chuxiong Sun, Dunqi Yao, Rui Wang 0079, Wenwen Qiang, Changwen Zheng, Jiangmeng Li
AAAI3
2026 AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
Jiangmeng Li, Rui Wang 0079, Changwen Zheng, Fanjiang Xu, Hui Xiong 0001
Int. J. Comput. Vis.4
2026 Visual reinforcement learning via sequential consistency preserved policy contrast from optimal transport view
Zehua Zang, Jiangmeng Li, Chuxiong Sun, Rui Wang 0079, Fuchun Sun 0001
Neural Networks4
2025 Mastering Visual Reinforcement Learning via Positive Unlabeled Policy-Guided Contrast
Zehua Zang, Qirui Ji, Rui Wang 0079, Kai Li 0047, Fuchun Sun 0001
ICIC (12)3
2025 A Game of Finding Delicate Data Augmentations: Reinforcement Contrastive Learning of Visual Representations
Zehua Zang, Rui Wang 0079, Fuchun Sun 0001
ICIC (12)2
2025 Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective
Chuxiong Sun, Rui Wang 0079, Changwen Zheng
AAMAS3
2025 Empowering Graph Contrastive Learning with Topological Rationale
Qirui Ji, Rui Wang 0079
PRICAI5
2024 Rethinking Dimensional Rationale in Graph Contrastive Learning from Causal Perspective
abstract
Graph contrastive learning is a general learning paradigm excelling at capturing invariant information from diverse perturbations in graphs. Recent works focus on exploring the structural rationale from graphs, thereby increasing the discriminability of the invariant information. However, such methods may incur in the mis-learning of graph models towards the interpretability of graphs, and thus the learned noisy and task-agnostic information interferes with the prediction of graphs. To this end, with the purpose of exploring the intrinsic rationale of graphs, we accordingly propose to capture the dimensional rationale from graphs, which has not received sufficient attention in the literature. The conducted exploratory experiments attest to the feasibility of the aforementioned roadmap. To elucidate the innate mechanism behind the performance improvement arising from the dimensional rationale, we rethink the dimensional rationale in graph contrastive learning from a causal perspective and further formalize the causality among the variables in the pre-training stage to build the corresponding structural causal model. On the basis of the understanding of the structural causal model, we propose the dimensional rationale-aware graph contrastive learning approach, which introduces a learnable dimensional rationale acquiring network and a redundancy reduction constraint. The learnable dimensional rationale acquiring network is updated by leveraging a bi-level meta-learning technique, and the redundancy reduction constraint disentangles the redundant features through a decorrelation process during learning. Empirically, compared with state-of-the-art methods, our method can yield significant performance boosts on various benchmarks with respect to discriminability and transferability. The code implementation of our method is available at https://github.com/ByronJi/DRGCL.
Qirui Ji, Jiangmeng Li, Jie Hu 0019, Rui Wang 0079, Changwen Zheng, Fanjiang Xu
AAAI4
2024 T2MAC: Targeted and Trusted Multi-Agent Communication through Selective Engagement and Evidence-Driven Integration
abstract
Communication stands as a potent mechanism to harmonize the behaviors of multiple agents. However, existing work primarily concentrates on broadcast communication, which not only lacks practicality, but also leads to information redundancy. This surplus, one-fits-all information could adversely impact the communication efficiency. Furthermore, existing works often resort to basic mechanisms to integrate observed and received information, impairing the learning process. To tackle these difficulties, we propose Targeted and Trusted Multi-Agent Communication (T2MAC), a straightforward yet effective method that enables agents to learn selective engagement and evidence-driven integration. With T2MAC, agents have the capability to craft individualized messages, pinpoint ideal communication windows, and engage with reliable partners, thereby refining communication efficiency. Following the reception of messages, the agents integrate information observed and received from different sources at an evidence level. This process enables agents to collectively use evidence garnered from multiple perspectives, fostering trusted and cooperative behaviors. We evaluate our method on a diverse set of cooperative multi-agent tasks, with varying difficulties, involving different scales and ranging from Hallway, MPE to SMAC. The experiments indicate that the proposed model not only surpasses the state-of-the-art methods in terms of cooperative performance and communication efficiency, but also exhibits impressive generalization.
Chuxiong Sun, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Changwen Zheng
AAAI6
2024 Delve into Cosine: Few-shot Object Detection via Adaptive Norm-cutting Scale
abstract
In this paper, we provide an analysis of the existing two-stage representation learning framework for the few-shot object detection from the perspective of normalization in the latent space, which is achieved by delving into cosine, thereby exploring the intrinsic reason behind its promotion of the model performance. Accordingly, the motivating experiments demonstrate that the cosine-similarity-based classifier enhances performance by alleviating the model’s bias towards the background class. However, the previously fixed scale applied in the classifier still leads to an overestimation of the probability that a novel instance belongs to the background class, thereby retaining the model’s bias towards the background. To address this issue, we establish the adaptive norm-cutting scale for instances that are predicted as the background instances, utilizing a learnable parameter to generate an appropriate scale. This optimization leverages our exploration, resulting in an improved Few-Shot Object Detection model. Experimental results validate the effectiveness of our approach.
Rui Wang 0079
IJCNN3
2024 CMIX: Causal Value Decomposition for Cooperative Multi-Agent Reinforcement Learning
abstract
Value decomposition plays a pivotal role in ensuring effective credit assignment within Multi-Agent Reinforcement Learning (MARL), particularly in cooperative multi-agent tasks where agents are limited to accessing team rewards only. However, existing methods treat the mixing network as a black box, implicitly assuming that neural networks can autonomously extract important information and achieve rational credit assignment during policy learning. This approach not only lacks interpretability but may also prove inefficient in complex scenarios. To enhance the interpretability and rationality of value decomposition, we propose an innovative approach called “Causal Value Decomposition”(CMIX). CMIX employs causal inference-based models, introducing a set of metrics beyond environmental rewards to enhance robustness and model interpretability. Specifically, CMIX establishes intricate relational structures among agents in complex environments and leverages causal relationships between agents and their surroundings to address the credit assignment challenge in MARL. By employing do-calculus, CMIX accurately measures the impact of each agent's actions on environmental states, precisely determining their contribution to the collective reward. This approach not only enhances the interpretability of existing black-box models but also improves the accuracy of credit assignment in multi-agent systems. Moreover, CMIX exhibits high scalability and complements existing value decomposition techniques. Its effectiveness and scalability have been rigorously tested across various settings, including MPE, LBF, and SMAC environments.
Dunqi Yao, Chuxiong Sun, Kai Li 0047, Kaijie Zhou, Rui Wang 0079
SMC6
2023 Adaptive Graph Augmentation for Graph Contrastive Learning
Zeming Wang, Rui Wang 0079, Changwen Zheng
ICIC (4)3
2023 Semi-Supervised Text Classification via Self-Paced Semantic-Level Contrast
Kaijie Zhou, Rui Wang 0079
PAKDD (2)4
2023 Physical Layer Security Against Passive Eavesdropper in Digital Twin-Enabler Power Grid: An IRS-Assisted Approach
abstract
The paper explores the issue of multiple-user fairness of intelligent reflecting surface (IRS)-assisted physical layer security (PLS) in the digital twin (DT)-enabler power grid. Previous research works have focused on achieving secrecy rate fairness through beamforming or phase shift optimization. However, in the DT-enabler power grid, the secrecy rate is not available as the instantaneous channel state information (CSI) of the passive eavesdropper is unknown. To address these challenges, we apply an expression for secrecy outage probability, measured based on the statistical CSI of the eavesdropper for the scenario where multiple DT users are present. Using zero-forcing (ZF) precoding at the transmitter, we formulate the problem of achieving fairness in secrecy outage probability, and then solve it by optimizing the phase shift matrices. Simulation results demonstrate that the proposed methods can achieve higher fairness among users in comparison to existing IRS-assisted PLS schemes.
Rui Wang 0079, Yiliang Liu, Donglan Liu, Fangzhe Zhang, Lili Sun, Tom H. Luan
PIMRC2
2023 Cross modification attention-based deliberation model for image captioning
Zheng Lian 0002, Haichang Li, Rui Wang 0079
Appl. Intell.4
2022 Context-Assisted Attention for Image Captioning
Zheng Lian 0002, Rui Wang 0079, Haichang Li
ICANN (1)2
2022 Deep Reinforcement Learning for Object Detection with the Updatable Target Network
abstract
We propose a reinforcement learning method for setting up the game of detecting objects within an image. Unlike some traditional image detection methods, which produce a large number of candidate boxes to detect objects, our approach allows an agent to use a small set of image locations to detect a visual object effectively. We train the agent to identify the useless information in the four edges of the image and discard them by representing predefined candidate areas in a tree-like hierarchy. Such a series of procedures that build the game environment is designed to detect locations of target objects. We also show that the updatable target network can make the agent reach stability faster and improve training results prominently, which utilizes good samples. Extensive comparison experiments on the benchmark dataset of Pascal VOC verify the outperformance of the proposed method.
Wenwu Yu, Rui Wang 0079
SMC2
2021 Multiple Fusion Adaptation: A Strong Framework for Unsupervised Semantic Segmentation Adaptation
Rui Wang 0079, Haichang Li
BMVC3
2021 LFMAC: Low-Frequency Multi-Agent Communication
Cong Cong 0004, Chuxiong Sun, Rui Wang 0079
ICONIP (5)4
2021 Reward Space Noise for Exploration in Deep Reinforcement Learning
abstract
A fundamental challenge for reinforcement learning (RL) is how to achieve efficient exploration in initially unknown environments. Most state-of-the-art RL algorithms leverage action space noise to drive exploration. The classical strategies are computationally efficient and straightforward to implement. However, these methods may fail to perform effectively in complex environments. To address this issue, we propose a novel strategy named reward space noise (RSN) for farsighted and consistent exploration in RL. By introducing the stochasticity from reward space, we are able to change agent’s understanding about environment and perturb its behaviors. We find that the simple RSN can achieve consistent exploration and scale to complex domains without intensive computational cost. To demonstrate the effectiveness and scalability of the proposed method, we implement a deep Q-learning agent with reward noise and evaluate its exploratory performance on a set of Atari games which are challenging for the naive [Formula: see text]-greedy strategy. The results show that reward noise outperforms action noise in most games and performs comparably in others. Concretely, we found that in the early training, the best exploratory performance of reward noise is obviously better than action noise, which demonstrates that the reward noise can quickly explore the valuable states and aid in finding the optimal policy. Moreover, the average scores and learning efficiency of reward noise are also higher than action noise through the whole training, which indicates that the reward noise can generate more stable and consistent performance.
Chuxiong Sun, Rui Wang 0079, Qian Li 0003
Int. J. Pattern Recognit. Artif. Intell.2
2021 Online Multiview Deep Forest for Remote Sensing Image Classification via Data Fusion
abstract
Remote sensing data can be sequentially acquired from different sources or feature spaces, which are regarded as multiple views. For the classification task where the training data arrive in a sequence, online learning (OL) methods are effective by learning new knowledge from incoming samples incrementally. However, it is known that shallow OL models usually have limited performance. In this letter, an online multiview deep forest (OMDF) architecture is proposed, which consists of multiple layers and employs a cascade structure. Each layer is an ensemble of multiple random forests, which process data from different views, respectively. For each view, the outputs of one layer concatenated with the original feature are fed into the next layer. The proposed method learns a deep forest model in an online manner from a stream of multiview data. The structure of every random forest and the weights adjusting the importance among different views will be updated dynamically. Experimental results on multifeature or multifrequency PolSAR data and the fusion of PolSAR and optical data demonstrate that the proposed method can achieve higher test accuracy and significantly improve the performance, especially on small-scale training data, compared with the other methods.
Xiangli Nie, Ruofei Gao, Rui Wang 0079, Deliang Xiang
IEEE Geosci. Remote. Sens. Lett.3
2020 Enhanced soft attention mechanism with an inception-like module for image captioning
abstract
Visual soft attention has been widely adopted in image captioning models. Traditional Soft Attention Mechanism (TSAM) assigns a weight to a certain region by using a multilayer perceptron with input from its own features. As image classification networks extract regional features based on spatial locations, TSAM fails to adequately consider the spatial contexts of regions, which leads to unreasonable weight distribution. In this paper, we introduce a flexible and universal attention framework with an inception-like module, named Enhanced Soft Attention Mechanism (ESAM), which can balance the attention levels of adjacent regions and alleviate the problem caused by local features with weak representational ability. Furthermore, we add an LSTM to the attention module so that it can take into account the previous attention distribution while generating the current word. Experimental results show that our ESAM significantly surpasses the TSAM by 4.1% on BLEU-4 and 2.7% on CIDEr, and achieves better results when verifying universality under the same experimental setups.
Zheng Lian 0002, Haichang Li, Rui Wang 0079
ICTAI3
2020 Learning Effective Value Function Factorization via Attentional Communication
abstract
How to achieve efficient cooperation among agents in partially observed environments remains an overarching problem in multi-agent reinforcement learning (MARL). Value function factorization learning is a promising way as it can efficiently address multi-agent credit assignment problem. However, existing value function factorization methods have been focusing on learning fully decentralized value functions, which are not effective for some complex tasks. To address this limitation, we propose a framework which enhances value function factorization by allowing communication during execution. Communication introduces extra information to help agents understand the complex environment and learn sophisticated factorization. Furthermore, the proposed mechanism of communication differs from existing methods since we additionally design a descriptive key along with the message. By the descriptive key, agents can dynamically measure the importance of different messages and achieve attentional communication. We evaluate our framework on a challenging set of StarCraft II micromanagement tasks, and show that it significantly outperforms existing value function factorization methods.
Xiaoya Yang, Chuxiong Sun, Rui Wang 0079
SMC4
2020 Detecting stealthy attacks on industrial control systems using a permutation entropy-based method
Hong Li 0004, Tom H. Luan, An Yang, Limin Sun 0001, Rui Wang 0079
Future Gener. Comput. Syst.7
2019 Efficient and Scalable Exploration via Estimation-Error
abstract
Exploring efficiently in complex environments is still a challenging problem in reinforcement learning. Recent exploration algorithms based on "optimism in the face of uncertainty" or intrinsic motivation achieved promising performance in sparse reward settings, but they often rely on additional structures which are hard to build in large scale problems. It renders them impractical and hinders the process of combining with reinforcement learning algorithms. Hence, the most state-of-the-art RL algorithms still use the naive action space noise as exploration strategy. In this paper, we model the uncertainty about environment through agent's ability to estimate the value across state and action space. Then, we parameterize the uncertainty by a neural network and regard it as a reward bonus signal to reward uncertain states. In this way, we generate an end-to-end bonus which can scale to complex environments with less computational cost. In order to prove the effectiveness of our method, we evaluate it on the challenging Atari 2600 games. We observed that our method achieves superior or comparable exploratory performance compared to action space noise in all environments, including environments whose rewards are sparse. The results demonstrate that our exploration method can motivate agent to explore effectively even in complex environments and it generally outperforms the naive action space noise.
Chuxiong Sun, Rui Wang 0079, Ruiying Li
IJCNN2
2018 Historical Best Q-Networks for Deep Reinforcement Learning
abstract
The popular DQN algorithm is known to have some instability and variability which make its performance poor sometimes. In prior work, there is only one target network, the network that is updated by the latest learned Q-value estimate. In this paper, we present multiple target networks which are the extension to the Deep Q-Networks (DQN). Based on the previously learned Q-value estimate networks, we choose several networks that perform best in all previous networks as our auxiliary networks. We show that in order to solve the problem of determining which network is better, we use the score of each episode as a measure of the quality of the network. The key behind our method is that each auxiliary network has some states that it is good at handling and guides the agent to make the right choices. We apply our method to the Atari 2600 games from the OpenAI Gym. We find that DQN with auxiliary networks significantly improves the performance and the stability of games.
Wenwu Yu, Rui Wang 0079, Ruiying Li
ICTAI2
2018 Multi-critic DDPG Method and Double Experience Replay
abstract
The remarkable Deep Deterministic Policy Gradient (DDPG) reinforcement learning method commonly consists of actor learning and critic learning. The actor learning highly relies on the critic learning, which makes the performance of DDPG method rather sensitive to critic learning and leads to stability issues. To further improve the stability and performance of DDPG method, the multi-critic DDPG method (MCDDPG) is proposed for a reliable critic learning. The average value of multiple critics is used to replace the single critic in DDPG method for better resistance when one critic performs badly, and multiple independent critics can learn knowledges from environment more widely. Besides, an extension of experience replay mechanism is revealed for accelerating the training process. All the methods are tested on simulated environments in OpenAI Gym platform, and convincing experiment results are obtained to support the proposed methods.
Rui Wang 0079, Ruiying Li, Hui Zhang 0061
SMC2
2015 Software architecture construction and collaboration based on service dependency
abstract
Cloud computing is a new software paradigm for resource integration and sharing in the open, dynamic and autonomous network environment. For large-scale complex software system on the dynamic evolution of the software architecture, taking service as the basic unit, the user can select service components and configuration online, can release of application form, and the ability to modify the published application. Current service dependencies including user application software architecture of logic dependency and dependency of evolution in the history of evolution, through the analysis and optimization of clustering in the service dependencies and reconstruction scheme of the user application software architecture is given.
Rui Wang 0079, Qimin Peng
CSCWD1