Shuguang Dou

dblp:224/4114 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-3231-8817ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Visual Bridge: Universal Visual Perception Representations Generating
abstract
Recent advances in diffusion models have achieved remarkable success in isolated computer vision tasks such as text-to-image generation, depth estimation, and optical flow. However, these models are often restricted by a ``single-task-single-model'' paradigm, severely limiting their generalizability and scalability in multi-task scenarios. Motivated by the cross-domain generalization ability of large language models, we propose a universal visual perception framework based on flow matching that can generate diverse visual representations across multiple tasks. Our approach formulates the process as a universal flow-matching problem from image patch tokens to task-specific representations rather than an independent generation or regression problem. By leveraging a strong self-supervised foundation model as the anchor and introducing a multi-scale, circular task embedding mechanism, our method learns a universal velocity field to bridge the gap between heterogeneous tasks, supporting efficient and flexible representation transfer. Extensive experiments on classification, detection, segmentation, depth estimation, and image-text retrieval demonstrate that our model achieves competitive performance in both zero-shot and fine-tuned settings, outperforming prior generalist and several specialist models. Ablation studies further validate the robustness, scalability, and generalization of our framework. Our work marks a significant step towards general-purpose visual perception, providing a solid foundation for future research in universal vision modeling.
Shuguang Dou, Junzhou Li, Zhiheng Yu, Dongsheng Jiang, Shugong Xu
AAAI2
2026 Person identity shift for privacy-preserving person re-identification
Shuguang Dou, Xinyang Jiang, Yansen Wang, Dongsheng Li 0002, Cairong Zhao
Sci. China Inf. Sci.1
2026 Active Dataset Distillation via Dual-Space Informative Matching
abstract
Dataset distillation improves neural network training efficiency by compressing large real datasets into compact synthetic datasets. Existing methods typically optimize matching objectives, such as aligning gradients, features, and trajectories between the synthetic and original datasets to ensure the distilled data retains essential properties for model training. However, many of these approaches rely on predefined distillation pools to streamline the process or treat all real data points equally, overlooking the dynamic nature of the synthetic dataset's training requirements during optimization. To address these limitations, we propose Active Dataset Distillation via Dual-Space Informative Matching (ACDD), an active learning-based algorithm that dynamically selects the most informative real data subset to align with the synthetic dataset's evolving needs. By adaptively refining the distillation pool, ACDD enhances training efficiency and generalization while ensuring the synthetic dataset effectively captures the original data's key characteristics. ACDD operates through two interconnected loops: the dual-space active loop (DAL) and the distillation loop. DAL plays a key role by dynamically selecting samples that balance diversity and uncertainty, adding them to the target distillation pool to meet the evolving informational needs of the current distillation loop. As a result, ACDD enables the synthetic dataset to achieve superior performance compared to SOTA methods across multiple benchmarks, including SVHN, CIFAR-10, CIFAR-100, TinyImageNet, and ImageNet subset. Moreover, ACDD reduces the required real dataset to just 20%-40% of the original, demonstrating its efficiency and effectiveness in data distillation.
Ding Qi, Jian Li 0062, Shuguang Dou, Junyao Gao 0002, Yabiao Wang, Bo Zhao 0015, Cairong Zhao
IEEE Trans. Image Process.3
2025 Towards Universal Dataset Distillation via Task-Driven Diffusion
abstract
Dataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, most recent research has primarily focused on image classification tasks, with limited exploration in detection and segmentation. Two key challenges remain: (i) Task Optimization Heterogeneity, where existing methods focus on class-level information but fail to address the diverse needs of detection and segmentation, and (ii) Inflexible Image Generation, where current generation methods rely on global updates for single-class targets and lack localized optimization for specific object regions. To address these challenges, we propose UniDD, a universal dataset distillation framework built on a task-driven diffusion model for diverse DD tasks, as shown in Fig. 1. Our approach operates in two stages: Universal Task Knowledge Mining, which captures task-relevant information through task-specific proxy model training, and Universal Task-Driven Diffusion, where these proxies guide the diffusion process to generate task-specific synthetic images. Extensive experiments across ImageNet-1K, Pascal VOC, and MS COCO demonstrate that UniDD consistently outperforms state-of-the-art methods. In particular, on ImageNet-1K with IPC-10, UniDD surpasses previous diffusion-based methods by 6.1%, while also reducing deployment costs.
Ding Qi, Jian Li 0062, Junyao Gao 0002, Shuguang Dou, Ying Tai, Jianlong Hu, Bo Zhao 0015, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
CVPR4
2025 DPL++: Advancing the Network Performance via Image and Label Perturbations
abstract
Recent advances in supervised learning have predominantly focused on regularizations, optimizers, and architectures, yet the potential of simultaneously optimizing data distributions and supervisory signals for training samples remains underexplored. In this paper, we propose a novel paradigm that leverages the benefits of image perturbations for rectifying data distributions. Our method, called DPL (Deep Perturbation Learning), introduces new insights into utilizing image perturbations and focuses on improving generalizability on normal samples, rather than resisting adversarial attacks. DPL formulates a differentiable function w.r.t. image perturbations and implements an alternative optimization process that seamlessly integrates with downstream tasks. However, the limitations of DPL stem from the inefficiency in employing differentiable targets caused by the exclusive optimization of image perturbations, while neglecting the critical role of supervisory signals in training effectiveness. These lead to the excessive necessity of DPL iterations and yield inferior performance-cost trade-off. To track this, we extend DPL to DPL++ with synchronous optimization for image perturbations and label perturbations. In our DPL++ paradigm, the post-hoc application of perturbations to images and labels endows amendments toward both data distributions and supervisory signals, significantly furthering the generalizability of models over various benchmarks. Crucially, the proposed synchronous optimization process shares key differentiable objectives to reduce computational complexity, thereby achieving enhanced effectiveness within fewer optimization iterations. Theoretically, as a generic and flexible approach, DPL++ can be applied to a variety of backbone architectures (e.g., ResNet, DenseNet, and ViT) and downstream tasks (e.g., image classification and object detection). To validate the efficacy of DPL++, we conduct extensive performance experiments and in-depth analytical studies on 2 visual tasks over 5 mainstream benchmarks across 13 backbone networks. The comprehensive results verify the superiority of DPL++ over DPL and demonstrate its promising capabilities for advancing decision-making capacity, risk minimization, class distinguishability, and training convergence.
Zifan Song, Guosheng Hu, Shuguang Dou, Cairong Zhao
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 EA-HAS-Bench and Language-Enhanced Shrinkage Search for Energy-Aware NAS
abstract
This paper takes a crucial step in the development of energy-aware (EA) NAS methods by offering a benchmark that enhances the reproducibility and accessibility of EA-NAS research. Specifically, we introduce EA-HAS-Bench, the first large-scale energy-aware benchmark designed to enable the study of AutoML methods in achieving improved trade-offs between performance and search energy consumption. EA-HAS-Bench offers a vast architecture/hyperparameter joint search space, encompassing diverse configurations relevant to energy consumption, and proposes a novel surrogate model based on Bézier curves for predicting learning curves with versatile shapes and lengths. On the other hand, recent studies have started integrating large language models (LLMs) into AutoML frameworks to enhance model search efficiency and configuration prediction, yet challenges remain in adapting these methods for energy-efficient searches across vast configuration spaces, as they often neglect energy consumption metrics. As a result, we introduce the Language-Enhanced Shrinkage Search (LESS), a plug-and-play method that utilizes the analytical capabilities of LLMs to enhance the energy efficiency of existing hyperparameter optimization techniques. Moreover, we adapt existing AutoML algorithms to construct baselines. Our experiments demonstrate that these modified energy-aware AutoML methods and LESS achieve an improved balance between energy consumption and model performance.
Cairong Zhao, Shuguang Dou, Xinyang Jiang, Junyao Gao 0002, Yuge Zhang, Bo Li 0080, Dongsheng Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Reviving Static Charts Into Live Charts
abstract
Data charts are prevalent across various fields due to their efficacy in conveying complex data relationships. However, static charts may sometimes struggle to engage readers and efficiently present intricate information, potentially resulting in limited understanding. We introduce "Live Charts," a new format of presentation that decomposes complex information within a chart and explains the information pieces sequentially through rich animations and accompanying audio narration. We propose an automated approach to revive static charts into Live Charts. Our method integrates GNN-based techniques to analyze the chart components and extract data from charts. Then we adopt large natural language models to generate appropriate animated visuals along with a voice-over to produce Live Charts from static ones. We conducted a thorough evaluation of our approach, which involved the model performance, use cases, a crowd-sourced user study, and expert interviews. The results demonstrate Live Charts offer a multi-sensory experience where readers can follow the information and understand the data insights better. We analyze the benefits and drawbacks of Live Charts over static charts as a new information consumption experience.
Lu Ying, Yun Wang 0012, Haotian Li 0001, Shuguang Dou, Xinyang Jiang, Huamin Qu, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.4
2024 Fetch and Forge: Efficient Dataset Condensation for Object Detection
abstract
Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses on image classification, with less exploration of object detection. This is primarily due to two challenges: (i) the multitasking nature of object detection complicates the condensation process, and (ii) Object detection datasets are characterized by large-scale and high-resolution data, which are difficult for existing DC methods to handle. As a remedy, we propose DCOD, the first dataset condensation framework for object detection. It operates in two stages: Fetch and Forge, initially storing key localization and classification information into model parameters, and then reconstructing synthetic images via model inversion. For the complex of multiple objects in an image, we propose Foreground Background Decoupling to centrally update the foreground of multiple instances and Incremental PatchExpand to further enhance the diversity of foregrounds. Extensive experiments on various detection datasets demonstrate the superiority of DCOD. Even at an extremely low compression rate of 1\%, we achieve 46.4\% and 24.7\% $\text{AP}_{50}$ on the VOC and COCO, respectively, significantly reducing detector training duration.
Ding Qi, Jian Li 0062, Jinlong Peng, Bo Zhao 0015, Shuguang Dou, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
NeurIPS5
2024 Re-ID-leak: Membership Inference Attacks Against Person Re-identification
Junyao Gao 0002, Xinyang Jiang, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cairong Zhao
Int. J. Comput. Vis.3
2024 Adaptive Discriminative Regularization for Visual Classification
Yi Wang 0033, Shuguang Dou, Cairong Zhao
Int. J. Comput. Vis.3
2024 Hierarchically Recognizing Vector Graphics and A New Chart-Based Vector Graphics Dataset
abstract
The conventional approach to image recognition has been based on raster graphics, which can suffer from aliasing and information loss when scaled up or down. In this paper, we propose a novel approach that leverages the benefits of vector graphics for object localization and classification. Our method, called YOLaT (You Only Look at Text), takes the textual document of vector graphics as input, rather than rendering it into pixels. YOLaT builds multi-graphs to model the structural and spatial information in vector graphics and utilizes a dual-stream graph neural network (GNN) to detect objects from the graph. However, for real-world vector graphics, YOLaT only models in flat GNN with vertexes as nodes ignore higher-level information of vector data. Therefore, we propose YOLaT++ to learn Multi-level Abstraction Feature Learning from a new perspective: Primitive Shapes to Curves and Points. On the other hand, given few public datasets focus on vector graphics, data-driven learning cannot exert its full power on this format. We provide a large-scale and challenging dataset for Chart-based Vector Graphics Detection and Chart Understanding, termed VG-DCU, with vector graphics, raster graphics, annotations, and raw data drawn for creating these vector charts. Experiments show that the YOLaT series outperforms both vector graphics and raster graphics-based object detection methods on both subsets of VG-DCU in terms of both accuracy and efficiency, showcasing the potential of vector graphics for image recognition tasks.
Shuguang Dou, Xinyang Jiang, Lu Liu 0019, Lu Ying, Yifei Shen 0004, Xuanyi Dong, Yun Wang 0012, Dongsheng Li 0002, Cairong Zhao
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Invisible Backdoor Attack With Dynamic Triggers Against Person Re-Identification
abstract
In recent years, person Re-IDentification (ReID) has rapidly progressed with wide real-world applications but is also susceptible to various forms of attack, including proven vulnerability to adversarial attacks. In this paper, we focus on the backdoor attack on deep ReID models. Existing backdoor attack methods follow an all-to-one or all-to-all attack scenario, where all the target classes in the test set have already been seen in the training set. However, ReID is a much more complex fine-grained open-set recognition problem, where the identities in the test set are not contained in the training set. Thus, previous backdoor attack methods for classification are not applicable to ReID. To ameliorate this issue, we propose a novel backdoor attack on deep ReID under a new all-to-unknown scenario, called Dynamic Triggers Invisible Backdoor Attack (DT-IBA). Instead of learning fixed triggers for the target classes from the training set, DT-IBA can dynamically generate new triggers for any unknown identities. Specifically, an identity hashing network is proposed to first extract target identity information from a reference image, which is then injected into the benign images by image steganography. We extensively validate the effectiveness and stealthiness of the proposed attack on benchmark datasets and evaluate the effectiveness of several defense methods against our attack.
Wenli Sun, Xinyang Jiang, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cheng Deng 0002, Cairong Zhao
IEEE Trans. Inf. Forensics Secur.3
2023 Similarity Distribution Based Membership Inference Attack on Person Re-identification
abstract
While person Re-identification (Re-ID) has progressed rapidly due to its wide real-world applications, it also causes severe risks of leaking personal information from training data. Thus, this paper focuses on quantifying this risk by membership inference (MI) attack. Most of the existing MI attack algorithms focus on classification models, while Re-ID follows a totally different training and inference paradigm. Re-ID is a fine-grained recognition task with complex feature embedding, and model outputs commonly used by existing MI like logits and losses are not accessible during inference. Since Re-ID focuses on modelling the relative relationship between image pairs instead of individual semantics, we conduct a formal and empirical analysis which validates that the distribution shift of the inter-sample similarity between training and test set is a critical criterion for Re-ID membership inference. As a result, we propose a novel membership inference attack method based on the inter-sample similarity distribution. Specifically, a set of anchor images are sampled to represent the similarity distribution conditioned on a target image, and a neural network with a novel anchor selection module is proposed to predict the membership of the target image. Our experiments validate the effectiveness of the proposed approach on both the Re-ID task and conventional classification task.
Junyao Gao 0002, Xinyang Jiang, Huishuai Zhang, Yifan Yang 0004, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cheng Deng 0002, Cairong Zhao
AAAI5
2023 EA-HAS-Bench: Energy-aware Hyperparameter and Architecture Search Benchmark
Shuguang Dou, Xinyang Jiang, Cairong Zhao, Dongsheng Li 0002
ICLR1
2023 Human Co-Parsing Guided Alignment for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) is a challenging task due to more background noises and incomplete foreground information. Although existing human parsing-based ReID methods can tackle this problem with semantic alignment at the finest pixel level, their performance is heavily affected by the human parsing model. Most supervised methods propose to train an extra human parsing model aside from the ReID model with cross-domain human parts annotation, suffering from expensive annotation cost and domain gap; Unsupervised methods integrate a feature clustering-based human parsing process into the ReID model, but lacking supervision signals brings less satisfactory segmentation results. In this paper, we argue that the pre-existing information in the ReID training dataset can be directly used as supervision signals to train the human parsing model without any extra annotation. By integrating a weakly supervised human co-parsing network into the ReID network, we propose a novel framework that exploits shared information across different images of the same pedestrian, called the Human Co-parsing Guided Alignment (HCGA) framework. Specifically, the human co-parsing network is weakly supervised by three consistency criteria, namely global semantics, local space, and background. By feeding the semantic information and deep features from the person ReID network into the guided alignment module, features of the foreground and human parts can then be obtained for effective occluded person ReID. Experiment results on two occluded and two holistic datasets demonstrate the superiority of our method. Especially on Occluded-DukeMTMC, it achieves 70.2% Rank-1 accuracy and 57.5% mAP.
Shuguang Dou, Cairong Zhao, Xinyang Jiang, Shanshan Zhang 0001, Wei-Shi Zheng 0001, Wangmeng Zuo
IEEE Trans. Image Process.1
2022 Context-Aware Feature Learning for Noise Robust Person Search
abstract
Person search aims to localize and identify specific pedestrians from numerous surveillance scene images. In this work, we focus on the noise in person search. We categorize the noise into scene-inherent noise and human-introduced noise. Scene-inherent noise comes from congestion, occlusion, and illumination changes. Human-introduced noise originates from the labeling process. For scene-inherent noise, we propose a novel context contrastive loss to take advantage of the latent contextual information from scene images. Features from context regions are utilized to construct contrastive pairs to constrain the feature discrimination among pedestrians in scene images while maintaining the feature consistency of the same identity. The network can thus learn to distinguish congested and overlapped pedestrians and more robust features can be obtained. For human-introduced noise, we propose a noise-discovery and noise-suppression training process for mislabeling robust person search. After the first training pass, the relation between feature prototypes of different identities is analyzed and the mislabeled pedestrians are discovered. During the second training pass, the label noise is suppressed to reduce the negative influence of mislabeled data. Experiments show that the proposed context-aware noise-robust (CANR) person search can achieve competitive performance. Further ablation studies confirm the effectiveness of CANR.
Cairong Zhao, Shuguang Dou, Zefan Qu, Jiawei Yao, Jun Wu 0006, Duoqian Miao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Detecting Overlapped Objects in X-Ray Security Imagery by a Label-Aware Mechanism
abstract
One of the key challenges to the X-ray security check is to detect the overlapped items in backpacks or suitcases in the X-ray images. Most existing methods improve the robustness of models to the object overlapping problem by enhancing the underlying visual information such as colors and edges. However, this strategy ignores the situations that the objects have similar visual clues as to the background, and objects overlapping each other. Since the two cases rarely appear in existing datasets, we contribute a novel dataset – Cutters and Liquid Containers X-ray Dataset (CLCXray) to complete the related research. Furthermore, we propose a novel Label-aware Mechanism (LA) to tackle the object overlapping problem. Particularly, LA establishes the associations between feature channels and different labels and adjusts the features according to the assigned labels (or pseudo labels) to help improve the prediction results. Extensive experiments demonstrate that the LA is accurate and robust to detect overlapped objects, and also validate the effectiveness and the good generalization of the LA for arbitrary state-of-the-art (SOTA) methods. Furthermore, experimental results show that the network constructed by the LA is superior to the SOTA models on OPIXray and CLCXray, especially solving the challenges of the subset of the highly overlapped objects.
Cairong Zhao, Shuguang Dou, Weihong Deng, Liang Wang 0001
IEEE Trans. Inf. Forensics Secur.3
2021 Incremental Generative Occlusion Adversarial Suppression Network for Person ReID
abstract
Person re-identification (re-id) suffers from the significant challenge of occlusion, where an image contains occlusions and less discriminative pedestrian information. However, certain work consistently attempts to design complex modules to capture implicit information (including human pose landmarks, mask maps, and spatial information). The network, consequently, focuses on discriminative features learning on human non-occluded body regions and realizes effective matching under spatial misalignment. Few studies have focused on data augmentation, given that existing single-based data augmentation methods bring limited performance improvement. To address the occlusion problem, we propose a novel Incremental Generative Occlusion Adversarial Suppression (IGOAS) network. It consists of 1) an incremental generative occlusion block, generating easy-to-hard occlusion data, that makes the network more robust to occlusion by gradually learning harder occlusion instead of hardest occlusion directly. And 2) a global-adversarial suppression (G&A) framework with a global branch and an adversarial suppression branch. The global branch extracts steady global features of the images. The adversarial suppression branch, embedded with two occlusion suppression module, minimizes the generated occlusion's response and strengthens attentive feature representation on human non-occluded body regions. Finally, we get a more discriminative pedestrian feature descriptor by concatenating two branches' features, which is robust to the occlusion problem. The experiments on the occluded dataset show the competitive performance of IGOAS. On Occluded-DukeMTMC, it achieves 60.1% Rank-1 accuracy and 49.4% mAP.
Cairong Zhao, Xinbi Lv, Shuguang Dou, Shanshan Zhang 0001, Jun Wu 0006, Liang Wang 0001
IEEE Trans. Image Process.3