VLDB 2026 Research / reviewers in the wild / expert
Wenbing Tang 0001
dblp:237/7970-1
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0125-1939ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SimpleDiffusion: A Lightweight and Efficient Conditional Diffusion Model for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (MSOD), which integrates complementary modalities such as depth or thermal data, primarily faces two challenges: accurately preserving salient object details and effectively aligning cross-modal features. Recent advances in using Stable Diffusion to generate images with fine edge details have inspired researchers to reformulate MSOD as a conditional mask generation process guided by salient features, which has achieved excellent visual results. However, these approaches often overlook the high computational cost and large-scale architecture of Stable Diffusion, both of which render it unsuitable for real-world MSOD applications. Therefore, we propose SimpleDiffusion, the first lightweight and efficient conditional diffusion model for MSOD that does not rely on Stable Diffusion. Specifically, we propose an Adaptive Cross-Modal Fusion Conditional Network and a Latent Denoising Network to reduce the complexity of diffusion models. Furthermore, we design a Multi-modal Feature Rectification and Fusion Module to enhance the representational capacity of cross-modal salient features. Customized training and sampling strategies are also developed to improve inference efficiency and reduce erroneous object segmentations. Experiments on multiple MSOD datasets demonstrate that SimpleDiffusion reduces model size by over tenfold and improves inference speed by more than fivefold compared to other diffusion-based methods, while maintaining comparable or superior performance. Shuo Zhang 0013, Wenbing Tang 0001, Jing Liu 0012, Li Han 0001, Jiandun Li, Hongchun Yuan, Zizhu Fan |
AAAI | 3 |
| 2026 | Fuzzy-DDPG: Integrating fuzzy logic with continuous deep reinforcement learning for mobile robot motion planning
Fenghua Wu, Wenbing Tang 0001, Yuan Zhou 0005, Hesuan Hu, Yang Liu 0003, Zuohua Ding |
Fuzzy Sets Syst. | 2 |
| 2026 | Causality-Aware Safety Testing for Autonomous Driving SystemsabstractSimulation-based testing is essential for evaluating the safety of Autonomous Driving Systems (ADSs). Comprehensive evaluation requires testing across diverse scenarios that can trigger various types of violations under different conditions. While existing methods typically focus on individual diversity metrics, such as input scenarios, ADS-generated motion commands, and system violations, they often fail to capture the complex interrelationships among these elements. For instance, identical motion commands can produce different collision risks in varying scenes, and the same collision may result from different commands under different scenarios. This oversight leads to gaps in testing coverage, potentially missing critical issues in the ADS under evaluation. In this paper, we proposeCausal-Fuzzer, the first causality-aware fuzzing technique that enables efficient and comprehensive testing of ADSs by constructing causal graphs to model the interrelationships among scenarios, actions, and violations. Unlike existing methods that treat diversity metrics independently, we recognize these elements are causally interconnected and use their relationships to identify more diverse violations triggered by fundamentally different causal mechanisms. Specifically,Causal-Fuzzerproposes (1) a causality-based feedback mechanism that quantifies the combined diversity of test scenarios by assessing whether they activate new causal relationships, and (2) a causality-driven mutation strategy that prioritizes mutations on input scenario elements with higher causal impact on ego action changes and violation occurrence to enable interpretable and efficient test generation. We evaluatedCausal-Fuzzeron an industry-grade ADS Apollo, with a high-fidelity simulator LGSVL. Our empirical results demonstrate thatCausal-Fuzzersignificantly outperforms existing methods in (1) identifying a greater diversity of violations (96.5 violations on average, compared to 66.9 for the best baseline method), (2) providing enhanced testing sufficiency with improved coverage of causal relationships (13.6 unique sceneaction- violation patterns on average, compared to 8.6 for the best baseline method), and (3) achieving greater efficiency in detecting critical scenarios, strong robustness under noise conditions, and good generalizability across varying scenario complexities and violation types. Our source code and experimental results are available athttps://sites.google.com/view/causal-fuzzer. Wenbing Tang 0001, Mingfei Cheng, Yuan Zhou 0005, Yang Liu 0003, Zuohua Ding |
IEEE Trans. Software Eng. | 1 |
| 2025 | DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) images are tackled separately. However, without intricate cross-modal fusion strategies, such approaches struggle to effectively integrate multi-modal information, often resulting in poorly defined object edges or overconfident inaccurate predictions. Recent studies have shown that designing a unified end-to-end framework to handle all three types of SOD tasks simultaneously is both necessary and difficult. To address this need, we propose a novel approach that treats multi-modal SOD as a conditional mask generation task utilizing diffusion models. We introduce DiMSOD, which enables the concurrent use of local (depth maps, thermal maps) and global controls (original images) within a unified model for progressive denoising and refined prediction. DiMSOD is efficient, only requiring fine-tuning of our newly introduced modules on the existing stable diffusion, which not only reduces the fine-tuning cost, making it more viable for practical use, but also enhances the integration of multi-modal conditional controls. Specifically, we have developed modules including SOD-ControlNet, Feature Adaptive Network (FAN), and Feature Injection Attention Network (FIAN) to enhance the model's performance. Extensive experiments demonstrate that DiMSOD efficiently detects salient objects across RGB, RGB-D, and RGB-T datasets, achieving superior performance compared to previous well-established methods. Shuo Zhang 0013, Wenbing Tang 0001, Terrence Hu, Xiaogang Xu 0002, Jing Liu 0012 |
AAAI | 3 |
| 2025 | Multi-modal Salient Object Detection via a Unified Diffusion ModelabstractSalient Object Detection (SOD) aims to identify and segment the most striking elements within an image. Salient object detection methods can be differentiated into several types according to the input data, such as RGB-D (Depth) and RGB-T (Thermal). Previous research primarily focused on saliency detection for single data types. However, forcing an RGB-D SOD model to process RGB-T data will degrade its performance significantly. In addition, current methods still face challenges in detecting fine edge details of salient objects and achieving end-to-end training. To address these issues, we introduce diffSOD, which leverages stable diffusion and cross-modal feature rectification and fusion module for saliency detection by transforming salient object detection into a denoising process from a noisy mask to an object mask. It offers a unified solution for salient object detection that seamlessly spans both RGB-D SOD and RGB-T SOD. Extensive experiments validate the effectiveness of the proposed diffSOD, demonstrating its ability to efficiently detect salient objects across both RGB-D and RGB-T data, while achieving superior performance over state-of-the-art methods. Shuo Zhang 0013, Wenbing Tang 0001, Lili Tian, Yuang Wei, Jing Liu 0012 |
ICASSP | 3 |
| 2025 | Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVSS) is a challenging computer vision task that labels each pixel within an image based on text descriptions. Recent advancements in OVSS are largely attributed to the increased model capacity. However, these models often struggle with unfamiliar images or unseen text, as their visual language understanding is limited to training data. Text-to-image (T2I) diffusion models have demonstrated strong image generation with diverse open-vocabulary descriptions. It prompted us to explore whether the comprehensive priors in T2I diffusion models could enhance the zero-shot generalization of OVSS. In this study, we define OVSS as a denoising diffusion task from noisy to object mask and introduce Seg-diffusion, a novel method based on Stable Diffusion that utilizes its extensive visual and linguistic prior knowledge. Specifically, the object mask diffuses from ground-truth to a random distribution in latent space. The model learns to reverse this noisy process to reconstruct object mask to segment objectives using text embeddings with our proposed Content Attention Module (CAM). Extensive experiments on popular OVSS benchmarks show that Seg-diffusion outperforms previous well-established methods and achieves impressive zero-shot generalization to unseen datasets. Shuo Zhang 0013, Wenbing Tang 0001, Jing Liu 0012 |
ICASSP | 5 |
| 2025 | Empowering Embodied Agents with Semantic Intelligence
Wenbing Tang 0001, Meilin Zhu, Fenghua Wu, Xinfeng Li, Yang Liu 0003 |
ICECCS | 1 |
| 2025 | Formal Modeling and Quantitative Evaluation for Online Monitoring Systems in Nuclear FacilitiesabstractAdvanced online execution monitoring is an essential system for ensuring the safety of nuclear facilities. Formal modeling and quantitative evaluation of these systems offer a promising approach to verifying their behaviors and identifying potential vulnerabilities. However, existing modeling languages often lack the capability to represent the system’s control flow logic. Additionally, the absence of automated transformation rules hinders the verification of generated models using available verification tools. Hence, in this paper, we propose a novel synchronous modeling language, Hybrid SynLong, which integrates data flow and control flow to effectively describe the real-time dynamic behaviors of online monitoring systems. Additionally, we present a method for converting the Hybrid SynLong language model into a network of stochastic hybrid automata, enabling direct verification with existing statistical model checkers. In consequence, the performance of an online monitoring system can be quantitatively evaluated by executing well-defined queries. The experimental results illustrate the effectiveness and efficiency of the proposed modeling language and transformation algorithms, as demonstrated through their application in an online monitoring system for a nuclear plant. Letian Fang, Wenbing Tang 0001, Jing Liu 0012 |
SMC | 2 |
| 2025 | A Role-based Hierarchical Adaptive Routing Protocol for Clustered UAV SwarmsabstractFlying Ad Hoc Networks (FANETs) are self-organizing wireless networks composed of Unmanned Aerial Vehicles (UAVs), designed to enable collaborative tasks without relying on fixed infrastructure. Characterized by high mobility, dynamic topology, and three-dimensional spatial movement, FANETs pose critical challenges to routing protocols, particularly to the broadcast storm problem caused by uncontrolled flooding in large-scale deployments. To address the challenges of broadcast storms in FANETs, this paper proposes RHARP, a role-based hierarchical routing protocol for FANETs. RHARP innovatively integrates bio-inspired swarm communication principles with role-based network clustering, establishing functional correspondence between decentralized coordination in biological swarms and differentiated node roles in FANETs. Through functional decoupling, RHARP delegates intra-cluster routing decisions to gateway nodes to alleviate cluster head (CH) workload. It achieves dynamic inter-cluster communication optimization through hierarchical routing strategies and on-demand path discovery mechanisms. By systematically integrating role-based clustering with adaptive routing, RHARP provides a scalable and cost-effective communication framework for large-scale FANET deployments while maintaining protocol agility. The simulation results demonstrate the effectiveness of RHARP across various dynamic scenarios. Tianen Guan, Xinliang Wu, Zhongliang Zhao, Wenbing Tang 0001, Yang Liu 0003 |
VTC2025-Fall | 5 |
| 2024 | Causal deconfounding deep reinforcement learning for mobile robot motion planning
Wenbing Tang 0001, Fenghua Wu, Shang-wei Lin, Zuohua Ding, Jing Liu 0012, Yang Liu 0003, Jifeng He 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Robust Motion Planning for Multi-Robot Systems Against Position Deception AttacksabstractDeep reinforcement learning (DRL) is widely applied in motion planning for multi-robot systems as DRL leverages the offline training process to improve the real-time computation efficiency. In DRL-based methods, the DRL models compute an action for a robot based on the states of its surrounding obstacles, including other robots in the system. They always assume that the number of obstacles is fixed and the obtained obstacles’ states are reliable. However, in the real world, a multi-robot system may suffer from various attacks, such as remote control attacks and network attacks, that cause wrong positions of the surrounding obstacles received by a robot. In this paper, we propose a robust motion planning methodDAE-Crit-LSTM, integrating a denoising autoencoder (DAE) with DRL models, to mitigate such position deception attacks in environments with a different number of obstacles.DAE-Crit-LSTMshows the following two advantages. First,DAE-Crit-LSTMcan be applied in benign and attacked scenarios and thus does not require any detector. It learns an encoder and a decoder to approximate the accurate positions of the obstacles, no matter under attack or not. Second,DAE-Crit-LSTMapplies an LSTM (Long Short-Term Memory)-based DRL model to deal with a variable number of obstacles in the environment. It is worth noting thatDAE-Crit-LSTMis method-agnostic and can be easily implemented in state-of-the-art motion planning methods. Comprehensive experiments show thatDAE-Crit-LSTMcan mitigate position deception attacks and guarantee safe motion. We also demonstrate the effectiveness and generalization ofDAE-Crit-LSTM. Wenbing Tang 0001, Yuan Zhou 0005, Yang Liu 0003, Zuohua Ding, Jing Liu 0012 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Causality-Guided Counterfactual Debiasing for Anomaly Detection of Cyber-Physical SystemsabstractMachine learning has become a promising technology for anomaly detection of cyber-physical systems (CPSs). However, the trained anomaly detection models always suffer from bias due to the scarcity of anomaly data in CPSs and the biased data collection process, which may poison the models' generalization ability. Recent debiasing methods are proposed to deal with the bias via resampling the training dataset, reweighting during the training phase, or adjusting the classification threshold. However, they may lose valuable information, need extra knowledge of the models, or lead to overfitting. Especially, they lack a causal understanding of the debiasing process, so they cannot point out the source and propagation of the bias and, thus, cannot deal with it in an explainable way. In this article, we propose a counterfactual debiasing framework to mitigate the bias in a well-trained model. First, we formalize the model's training and inference processes using causal graphs. Thus, we can understand the source and propagation of the model's bias through causal inference. Then, we use counterfactual inference to estimate the bias's detrimental causal effect on the prediction and remove it from the total causal effect. Therefore, we can conduct unbiased inferences with a biased model. The proposed method can remove the bias in an explainable way by incorporating causal graphs. Comprehensive experiments are conducted on seven real-world CPS datasets, i.e., IDA, MFP, ACS, SPF, UNS, NSL, and ICS. The results demonstrate the effectiveness, compatibility, and unbiasedness of the proposed approach. Wenbing Tang 0001, Jing Liu 0012, Yuan Zhou 0005, Zuohua Ding |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Tracking a Ground Moving Target with UAV Based on Interval Type-2 Fuzzy LogicabstractIn recent years, tracking a ground moving target with an unmanned aerial vehicle (UAV) has played a key role in navigation, military reconnaissance, field rescue, and traffic monitoring tasks. However, due to the maneuvering uncertainty of the moving target and various uncertainties in the complex environment, it has brought many difficulties to the tracking task. In this paper, aiming at addressing the various uncertainties in the process of target tracking with UAV, we present a control strategy based on interval type-2 fuzzy logic. We design two interval type-2 fuzzy controllers to control the yaw angle and speed of the UAV, respectively. Images obtained from an embedded camera on the UAV platform are visually processed by the target recognition and tracking algorithms in the early stage to detect the pixel position of the target in every frame, which is used as the input of the controller we have designed. Through the analysis of the input and output data of controllers designed by us and the analysis of the tracking error, the reliability of control strategy proposed by us is proved. At last, a series of simulation experiments are implemented in ROS to demonstrate the performance of the target-tracking strategy we proposed. Yao Li 0011, Wenbing Tang 0001, Bochen Chen, Zuohua Ding |
TASE | 2 |