EDBT 2026 Demo / reviewers in the wild / expert
Shuo Zhang 0013
dblp:83/3714-13
· DBLP profile ↗
9ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0001-8523-4903ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SimpleDiffusion: A Lightweight and Efficient Conditional Diffusion Model for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (MSOD), which integrates complementary modalities such as depth or thermal data, primarily faces two challenges: accurately preserving salient object details and effectively aligning cross-modal features. Recent advances in using Stable Diffusion to generate images with fine edge details have inspired researchers to reformulate MSOD as a conditional mask generation process guided by salient features, which has achieved excellent visual results. However, these approaches often overlook the high computational cost and large-scale architecture of Stable Diffusion, both of which render it unsuitable for real-world MSOD applications. Therefore, we propose SimpleDiffusion, the first lightweight and efficient conditional diffusion model for MSOD that does not rely on Stable Diffusion. Specifically, we propose an Adaptive Cross-Modal Fusion Conditional Network and a Latent Denoising Network to reduce the complexity of diffusion models. Furthermore, we design a Multi-modal Feature Rectification and Fusion Module to enhance the representational capacity of cross-modal salient features. Customized training and sampling strategies are also developed to improve inference efficiency and reduce erroneous object segmentations. Experiments on multiple MSOD datasets demonstrate that SimpleDiffusion reduces model size by over tenfold and improves inference speed by more than fivefold compared to other diffusion-based methods, while maintaining comparable or superior performance. Shuo Zhang 0013, Wenbing Tang 0001, Jing Liu 0012, Li Han 0001, Jiandun Li, Hongchun Yuan, Zizhu Fan |
AAAI | 1 |
| 2025 | DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) images are tackled separately. However, without intricate cross-modal fusion strategies, such approaches struggle to effectively integrate multi-modal information, often resulting in poorly defined object edges or overconfident inaccurate predictions. Recent studies have shown that designing a unified end-to-end framework to handle all three types of SOD tasks simultaneously is both necessary and difficult. To address this need, we propose a novel approach that treats multi-modal SOD as a conditional mask generation task utilizing diffusion models. We introduce DiMSOD, which enables the concurrent use of local (depth maps, thermal maps) and global controls (original images) within a unified model for progressive denoising and refined prediction. DiMSOD is efficient, only requiring fine-tuning of our newly introduced modules on the existing stable diffusion, which not only reduces the fine-tuning cost, making it more viable for practical use, but also enhances the integration of multi-modal conditional controls. Specifically, we have developed modules including SOD-ControlNet, Feature Adaptive Network (FAN), and Feature Injection Attention Network (FIAN) to enhance the model's performance. Extensive experiments demonstrate that DiMSOD efficiently detects salient objects across RGB, RGB-D, and RGB-T datasets, achieving superior performance compared to previous well-established methods. Shuo Zhang 0013, Wenbing Tang 0001, Terrence Hu, Xiaogang Xu 0002, Jing Liu 0012 |
AAAI | 1 |
| 2025 | Multi-modal Salient Object Detection via a Unified Diffusion ModelabstractSalient Object Detection (SOD) aims to identify and segment the most striking elements within an image. Salient object detection methods can be differentiated into several types according to the input data, such as RGB-D (Depth) and RGB-T (Thermal). Previous research primarily focused on saliency detection for single data types. However, forcing an RGB-D SOD model to process RGB-T data will degrade its performance significantly. In addition, current methods still face challenges in detecting fine edge details of salient objects and achieving end-to-end training. To address these issues, we introduce diffSOD, which leverages stable diffusion and cross-modal feature rectification and fusion module for saliency detection by transforming salient object detection into a denoising process from a noisy mask to an object mask. It offers a unified solution for salient object detection that seamlessly spans both RGB-D SOD and RGB-T SOD. Extensive experiments validate the effectiveness of the proposed diffSOD, demonstrating its ability to efficiently detect salient objects across both RGB-D and RGB-T data, while achieving superior performance over state-of-the-art methods. Shuo Zhang 0013, Wenbing Tang 0001, Lili Tian, Yuang Wei, Jing Liu 0012 |
ICASSP | 1 |
| 2025 | Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVSS) is a challenging computer vision task that labels each pixel within an image based on text descriptions. Recent advancements in OVSS are largely attributed to the increased model capacity. However, these models often struggle with unfamiliar images or unseen text, as their visual language understanding is limited to training data. Text-to-image (T2I) diffusion models have demonstrated strong image generation with diverse open-vocabulary descriptions. It prompted us to explore whether the comprehensive priors in T2I diffusion models could enhance the zero-shot generalization of OVSS. In this study, we define OVSS as a denoising diffusion task from noisy to object mask and introduce Seg-diffusion, a novel method based on Stable Diffusion that utilizes its extensive visual and linguistic prior knowledge. Specifically, the object mask diffuses from ground-truth to a random distribution in latent space. The model learns to reverse this noisy process to reconstruct object mask to segment objectives using text embeddings with our proposed Content Attention Module (CAM). Extensive experiments on popular OVSS benchmarks show that Seg-diffusion outperforms previous well-established methods and achieves impressive zero-shot generalization to unseen datasets. Shuo Zhang 0013, Wenbing Tang 0001, Jing Liu 0012 |
ICASSP | 1 |
| 2024 | Feature-Constrained and Attention-Conditioned Distillation Learning for Visual Anomaly DetectionabstractVisual anomaly detection in computer vision is an essential one-class classification and segmentation problem. The student-teacher (S-T) approach has proven effective in addressing this challenge. However, previous studies based on S-T underutilize the feature representations learned by the teacher network, which restricts anomaly detection performance. In this study, we propose a novel feature-constrained and attention-conditioned distillation learning method for visual anomaly detection with localization, which fully uses the features of the teacher model and the local semantics of the critical structure to instruct the student model to detect anomalies efficiently. Specifically, we introduce the Vision Transformer (ViT) as the back-bone for anomaly detection tasks, and the central feature strategy and self-attention masking strategy are proposed to constrain the output features and impose agreement between multi-image views. It improves the ability of the student network to describe normal data features and widens the feature difference between the student and teacher networks for abnormal data. Experiments on the benchmark datasets demonstrate that the proposed method significantly improves the performance of visual anomaly detection compared with the competing methods. Shuo Zhang 0013, Jing Liu 0012 |
ICASSP | 1 |
| 2024 | TranBF: Deep Transformer Networks and Bayesian Filtering for Time Series Anomalous Signal Detection in Cyber-physical SystemsabstractEffective anomalous signal detection in time series multimedia data is imperative for safety-critical cyber-physical systems (CPS). Nevertheless, constructing a system for precise and rapid anomaly detection is challenging due to complex system dynamics, long-range dependencies, and unidentified sensor noise in modern CPS. This study proposes TranBF, an innovative time series anomalous signal detection method designed with a carefully engineered deep Transformer network and Bayesian Filtering. TranBF aims to capture the dynamics and broader temporal dependencies of CPS within a dynamic state-space and recursively track the uncertainty of system noise over time, thereby significantly improving the robustness and accuracy of anomaly detection. Extensive experiments on three real-world public datasets demonstrate that TranBF can significantly out-perform state-of-the-art baseline methods in terms of detection performance. Specifically, TranBF enhances F1 scores by a maximum of 16.5% while concurrently reducing training times by as much as 39.3% compared to the baseline models. Furthermore, the ablation study furnishes empirical evidence supporting the effectiveness of each component within TranBF. Shuo Zhang 0013, Xiongpeng Hu, Jing Liu 0012 |
ICME | 1 |
| 2024 | Causal Fusion of Convolutional Neural Network and Vision Transformer for Image Anomaly Detection and LocalizationabstractTo address the challenge of visual anomaly detection amidst complex background interference. First, we construct a structural causal model for anomaly detection under complex background interference and propose an intervention strategy to block background feature interference. Then, we build an anomaly feature-sensitive neural network (AFSNN) containing two feature extraction modules based on the causal intervention strategy. Given the limitations of convolutional neural networks in capturing global features associated with spatial location dependence, and the substantial data requirements of vision transformers, we opt for the enhanced Swin Transformer module and the deformable convolutional networks encoder module to extract global features and local details, respectively. We also designed the cross-attention to fuse these two scales of feature representation. Finally, we introduce a causality-sensitive learning module that differentiates the outputs of the two feature extraction modules and constructs a causality-sensitive loss function by maximizing the output differences. This approach blocks background features and enhances sensitivity to anomaly features during training. Experiments show that AFSNN can effectively attenuate the confusing interference of the background pattern. Shuo Zhang 0013, Xiongpeng Hu, Jing Liu 0012 |
ICME | 1 |
| 2024 | SOD-diffusion: Salient Object Detection via Diffusion-Based Image GeneratorsabstractAbstract Salient Object Detection (SOD) is a challenging task that aims to precisely identify and segment the salient objects. However, existing SOD methods still face challenges in making explicit predictions near the edges and often lack end‐to‐end training capabilities. To alleviate these problems, we propose SOD‐diffusion, a novel framework that formulates salient object detection as a denoising diffusion process from noisy masks to object masks. Specifically, object masks diffuse from ground‐truth masks to random distribution in latent space, and the model learns to reverse this noising process to reconstruct object masks. To enhance the denoising learning process, we design an attention feature interaction module (AFIM) and a specific fine‐tuning protocol to integrate conditional semantic features from the input image with diffusion noise embedding. Extensive experiments on five widely used SOD benchmark datasets demonstrate that our proposed SOD‐diffusion achieves favorable performance compared to previous well‐established methods. Furthermore, leveraging the outstanding generalization capability of SOD‐diffusion, we applied it to publicly available images, generating high‐quality masks that serve as an additional SOD benchmark testset. Shuo Zhang 0013, Shizhe Chen, Jing Liu 0012 |
Comput. Graph. Forum | 1 |
| 2023 | Anomalous Signal Detection for Cyber-Physical Systems Using Interpretable Causal Neural NetworkabstractAnomalous signal detection aims to detect unknown abnormal signals of machines from normal signals. However, building effective and interpretable anomaly detection models for safety-critical cyber-physical systems (CPS) is rather difficult due to the unidentified system noise and extremely intricate system dynamics of CPS and the neural network black box. This work proposes a novel time series anomalous signal detection model based on neural system identification and causal inference to track the dynamics of CPS in a dynamical state-space and avoid absorbing spurious correlation caused by confounding bias generated by system noise, which improves the stability, security and interpretability in detection of anomalous signals from CPS. Experiments on three real-world CPS datasets show that the proposed method achieved considerable improvements compared favorably to the state-of-the-art methods on anomalous signal detection in CPS. Moreover, the ablation study empirically demonstrates the efficiency of each component in our method. Shuo Zhang 0013, Jing Liu 0012 |
ICASSP | 1 |