EDBT 2026 Demo / reviewers in the wild / expert
Siao Liu
dblp:315/9254
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-4285-3573ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint-Guided Spatial and Semantic Sensitive Diffusion Policy for Robotic ManipulationabstractImitation learning has shown strong potential for enabling robots to acquire dexterous manipulation skills by integrating visual observations with proprioceptive states. However, common approaches typically use visual encoders pretrained in computer vision domains, which mainly aim to extract generic representations without emphasizing the precise spatial and semantic structures that are crucial for robotic manipulation. In this work, we propose the Joint-Guided Spatial and Semantic Sensitive Diffusion Policy (S3D), which effectively fuses structured and generic features by incorporating depth and semantic maps with RGB and proprioceptive inputs to strengthen spatial–semantic understanding in manipulation. However, naively incorporating these multimodal representations inevitably introduces additional computational overhead. Thus, we introduce a Joint-Guided Dynamic Attention module that generates joint-conditioned queries to extract behavior-specific representations with controlled complexity. Experiments across a variety of simulated and real-world robotic manipulation tasks demonstrate that S3D yields consistent performance gains over state-of-the-art methods. Hongda Zhang, Siao Liu, Yi Liu 0027, Chun Ouyang 0002, Zhongxue Gan 0001 |
ICMR | 2 |
| 2026 | Privacy-Preserving Video Anomaly Detection: A SurveyabstractThe video anomaly detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm, such as fighting, stealing, and car accidents. However, vision-based surveillance systems such as closed-circuit television (CCTV) often capture personally identifiable information. The lack of transparency and interpretability in video transmission and usage raises public concerns about privacy and ethics, limiting the real-world application of VAD. Recently, researchers have focused on privacy concerns in VAD by conducting systematic studies from various perspectives, including data, features, and systems, making privacy-preserving VAD (P2VAD) a hotspot in the AI community. However, the current research in P2VAD is fragmented, and prior reviews have mostly focused on methods using RGB sequences, overlooking privacy leakage and appearance bias considerations. To address this gap, this article is the first to systematically review the progress of P2VAD, defining its scope and providing an intuitive taxonomy. We outline the basic assumptions, learning frameworks, and optimization objectives of various approaches, analyzing their strengths, weaknesses, and potential correlations. In addition, we provide open access to research resources such as benchmark datasets and available code. Finally, we discuss key challenges and future opportunities from the perspectives of AI development and P2VAD deployment, aiming to the guide future work in the field. Yang Liu 0246, Siao Liu, Xiaoguang Zhu, Hao Yang 0055, Juncen Guo, Liangyu Teng, Dingkang Yang, Yan Wang 0068, Jing Liu 0050 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationabstractBuilding a lifelong robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in alleviating catastrophic forgetting problem, naively applying these methods causes a failure to leverage the shared primitives between skills. To tackle these issues, we propose Primitive Prompt Learning (PPL), to achieve lifelong robot manipulation via reusable and extensible primitives. Within our two stage learning scheme, we first learn a set of primitive prompts to represent shared primitives through multi-skills pre-training stage, where motion-aware prompts are learned to capture semantic and motion shared primitives across different skills. Secondly, when acquiring new skills in lifelong span, new prompts are concatenated and optimized with frozen pretrained prompts, boosting the learning via knowledge transfer from old skills to new ones. For evaluation, we construct a large-scale skill dataset and conduct extensive experiments in both simulation and real-world tasks, demonstrating PPL’s superior performance over state-of-the-art methods. Yuanqi Yao, Siao Liu, Haoming Song, Delin Qu, Yan Ding 0002, Bin Zhao 0001, Zhigang Wang 0002, Xuelong Li 0001, Dong Wang 0028 |
CVPR | 2 |
| 2025 | ACORN: Acyclic Coordination with Reachability Network to Reduce Communication Redundancy in Multi-Agent Systems
Ziqing Zhou, Chun Ouyang 0002, Siao Liu, Linqiang Hu, Zhongxue Gan 0001 |
AAMAS | 4 |
| 2025 | Heuristics-Assisted Experience Replay Strategy for Cooperative Multi-Agent Reinforcement Learning
Ziqing Zhou, Chun Ouyang 0002, Siao Liu, Linqiang Hu, Zhongxue Gan 0001 |
AAMAS | 4 |
| 2024 | De-Confounded Data-Free Knowledge Distillation for Handling Distribution ShiftsabstractData-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However, a long-overlooked issue is that the severe distribution shifts between their substitution and original data, which mani-fests as huge differences in the quality of images and class proportions. The harmful shifts are essentially the con-founder that significantly causes performance bottlenecks. To tackle the issue, this paper proposes a novel perspective with causal inference to disentangle the student models from the impact of such shifts. By designing a customized causal graph, we first reveal the causalities among the variables in the DFKD task. Subsequently, we propose a Knowledge Distillation Causal Intervention (KDCI) framework based on the backdoor adjustment to de-confound the confounder. KDCI can be flexibly combined with most existing state-of-the-art baselines. Experiments in combination with six representative DFKD methods demonstrate the effectiveness of our KDCI, which can obviously help existing methods under almost all settings, e.g., improving the base-line by up to 15.54% accuracy on the CIFAR-100 dataset. Dingkang Yang, Zhaoyu Chen 0001, Yang Liu 0246, Siao Liu, Lihua Zhang 0002, Lizhe Qi |
CVPR | 5 |
| 2024 | Improving Domain Generalization in Self-supervised Monocular Depth Estimation via Stabilized Adversarial Training
Yuanqi Yao, Gang Wu 0010, Kui Jiang, Siao Liu, Jian Kuai, Xianming Liu 0005, Junjun Jiang |
ECCV (24) | 4 |
| 2024 | Sampling to Distill: Knowledge Transfer from Open-World DataabstractData-Free Knowledge Distillation (DFKD) is a novel task that aims to train high-performance student models using only the pre-trained teacher network without original training data. Most of the existing DFKD methods rely heavily on additional generation modules to synthesize the substitution data resulting in high computational costs and ignoring the massive amounts of easily accessible, low-cost, unlabeled open-world data. Meanwhile, existing methods ignore the domain shift issue between the substitution data and the original data, resulting in knowledge from teachers not always trustworthy and structured knowledge from data becoming a crucial supplement. To tackle the issue, we propose a novel Open-world Data Sampling Distillation (ODSD) method for the DFKD task without the redundant generation process. First, we try to sample open-world data close to the original data's distribution by an adaptive sampling module and introduce a low-noise representation to alleviate the domain shift issue. Then, we build structured relationships of multiple data examples to exploit data knowledge through the student model itself and the teacher's structured representation. Extensive experiments on CIFAR-10, CIFAR-100, NYUv2, and ImageNet show that our ODSD method achieves state-of-the-art performance with lower FLOPs and parameters. Especially, we improve 1.50%-9.59% accuracy on the ImageNet dataset and avoid training the separate generator for each class. Zhaoyu Chen 0001, Jie Zhang 0107, Dingkang Yang, Zuhao Ge, Yang Liu 0246, Siao Liu, Yunquan Sun, Lizhe Qi |
ACM Multimedia | 7 |
| 2024 | Heterogeneous Robot Swarms with an Attention Mechanism for Dynamic Target TrackingabstractMultirobot collaboration offers significant potential for diverse applications, including tracking and surveillance. In this paper, we introduce an attention mechanism tailored for heterogeneous robot swarms characterized by varied sensing ranges. This mechanism effectively utilizes the swarm's intrinsic characteristics, enabling rapid information transmission and ensuring consistent collective responses to external stimuli. Additionally, we introduce a pigeon-inspired navigation strategy that effectively replaces the traditional obstacle repulsion term by preventing the swarm from becoming trapped in local min-ima and reducing oscillatory behaviors. To validate the efficacy of our algorithm, we have developed an autonomously designed PlusBot swarm platform, which consists of agile vibration-driven miniature robots. Each of them is equipped with its own computing and communication system and is capable of precise closed-loop motion control. This setup meets the requirements for conducting heterogeneous swarm movement experiments in indoor environments. Through comprehensive numerical simulations and real-world experiments, our method has demonstrated exceptional precision and adaptability in tracking dynamic targets. The comparative analysis under-scores the superiority of our approach, particularly in minimizing swarm collisions and ensuring safe navigation in dynamic target-tracking scenarios involving obstacles. Ziqing Zhou, Chun Ouyang 0002, Xinyang Dong, Siao Liu, Linqiang Hu, Zhile Zhao, Zhongxue Gan 0001 |
SMC | 5 |
| 2024 | DiffSkill: Improving Reinforcement Learning through diffusion-based skill denoiser for robotic manipulation
Siao Liu, Yang Liu 0246, Linqiang Hu, Ziqing Zhou, Zhile Zhao, Wei Li 0055, Zhongxue Gan 0001 |
Knowl. Based Syst. | 1 |
| 2024 | AMP-Net: Appearance-Motion Prototype Network Assisted Automatic Video Anomaly Detection SystemabstractAs essential tools for industry safety protection, automatic video anomaly detection systems (AVADS) are designed to detect anomalous events of concern in surveillance videos. Existing VAD methods lack effective exploration of the prototypical appearance and motion features leading to poor performance in realistic scenarios. Specifically, they either misreport regular events as anomalies due to insufficient representation power, or lead to missed detections with over-power generalization. In this regard, we propose an appearance-motion prototype network (AMP-net) that uses external memories to record prototype features and augments the appearance-motion prototype with a spatial-temporal fusion. In addition, AMP-net sequentially fuses appearance features from deep to shallow to utilize multiscale spatial context. Additionally, we introduce temporal attention to capture important dynamics and enhance AMP-net for representing regular motion. The proposed method achieves a delicate balance of effective representation of normal events and limited generalization to anomalies. Experiments on three benchmark datasets demonstrate that our method can accurately detect anomalous events, achieving performance comparable to state-of-the-art methods with frame-level AUCs of 98.7%, 92.4%, and 78.8% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. Moreover, we conducted a case study on the self-collected industrial dataset, and the results indicate that our AMP-net can cope with complex industrial scenarios and outperform existing methods. Yang Liu 0246, Jing Liu 0050, Kun Yang 0010, Bobo Ju, Siao Liu, Dingkang Yang, Peng Sun 0007 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Context De-Confounded Emotion RecognitionabstractContext-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representations from subjects and contexts. However, a long-overlooked issue is that a context bias in existing datasets leads to a significantly unbalanced distribution of emotional states among different context scenarios. Concretely, the harmful bias is a confounder that misleads existing models to learn spurious correlations based on conventional likelihood estimation, significantly limiting the models' performance. To tackle the issue, this paper provides a causality-based perspective to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a tailored causal graph. Then, we propose a Contextual Causal Intervention Module (CCIM) based on the backdoor adjustment to de-confound the confounder and exploit the true causal effect for model training. CCIM is plug-in and model-agnostic, which improves diverse state-of-the-art approaches by considerable margins. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our CCIM and the significance of causal insight. Dingkang Yang, Zhaoyu Chen 0001, Shunli Wang 0001, Mingcheng Li, Siao Liu, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002 |
CVPR | 6 |
| 2023 | Adversarial Contrastive Distillation with Adaptive DenoisingabstractAdversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not always make correct predictions, interfering with the student’s robust performance. Besides, in the previous ARD methods, the robustness comes entirely from one-to-one imitation, ignoring the relationship between examples. To this end, we propose a novel structured ARD method called Contrastive Relationship DeNoise Distillation (CRDND). We design an adaptive compensation module to model the instability of the teacher. Moreover, we utilize the contrastive relationship to explore implicit robustness knowledge among multiple examples. Experimental results on multiple attack benchmarks show CRDND can transfer robust knowledge efficiently and achieves state-of-the-art performance. Zhaoyu Chen 0001, Dingkang Yang, Yang Liu 0246, Siao Liu, Lizhe Qi |
ICASSP | 5 |
| 2023 | Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement AugmentationabstractLearning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficiency, suffering from serve performance degradation. In this paper, we first conduct qualitative analysis and illuminate the main causes: (i) high-variance gradient magnitudes and (ii) gradient conflicts existed in various augmentation methods. To alleviate these issues, we propose a general policy gradient optimization framework, named Conflict-aware Gradient Agreement Augmentation (CG2A), and better integrate augmentation combination into visual RL algorithms to address the generalization bias. In particular, CG2A develops a Gradient Agreement Solver to adaptively balance the varying gradient magnitudes, and introduces a Soft Gradient Surgery strategy to alleviate the gradient conflicts. Extensive experiments demonstrate that CG2A significantly improves the generalization performance and sample efficiency of visual RL algorithms. Siao Liu, Zhaoyu Chen 0001, Yang Liu 0246, Dingkang Yang, Zhile Zhao, Ziqing Zhou, Xie Yi, Wei Li 0055, Zhongxue Gan 0001 |
ICCV | 1 |
| 2023 | Learning Causality-inspired Representation Consistency for Video Anomaly DetectionabstractVideo anomaly detection is an essential yet challenging task in the multimedia community, with promising applications in smart cities and secure communities. Existing methods attempt to learn abstract representations of regular events with statistical dependence to model the endogenous normality, which discriminates anomalies by measuring the deviations to the learned distribution. However, conventional representation learning is only a crude description of video normality and lacks an exploration of its underlying causality. The learned statistical dependence is unreliable for diverse regular events in the real world and may cause high false alarms due to over generalization. Inspired by causal representation learning, we think that there exists a causal variable capable of adequately representing the general patterns of regular events in which anomalies will present significant variations. Therefore, we design a causality-inspired representation consistency (CRC) framework to implicitly learn the unobservable causal variables of normality directly from available normal videos and detect abnormal events with the learned representation consistency. Extensive experiments show that the causality-inspired normality is robust to regular events with label-independent shifts, and the proposed CRC framework can quickly and accurately detect various complicated anomalies from real-world surveillance videos. Yang Liu 0246, Zhaoyang Xia, Mengyang Zhao 0002, Donglai Wei 0002, Siao Liu, Bobo Ju, Gaoyun Fang, Jing Liu 0050 |
ACM Multimedia | 6 |
| 2022 | Efficient Universal Shuffle Attack for Visual Object TrackingabstractRecently, adversarial attacks have been applied in visual object tracking to deceive deep trackers by injecting imperceptible perturbations into video frames. However, previous work only generates the video-specific perturbations, which restricts its application scenarios. In addition, existing attacks are difficult to implement in reality due to the real-time of tracking and the re-initialization mechanism. To address these issues, we propose an offline universal adversarial attack called Efficient Universal Shuffle Attack. It takes only one perturbation to cause the tracker malfunction on all videos. To improve the computational efficiency and attack performance, we propose a greedy gradient strategy and a triple loss to efficiently capture and attack model-specific feature representations through the gradients. Experimental results show that EUSA can significantly reduce the performance of state-of-the-art trackers on OTB2015 and VOT2018. Siao Liu, Zhaoyu Chen 0001, Wei Li 0055, Jiwei Zhu, Zhongxue Gan 0001 |
ICASSP | 1 |