VLDB 2026 Research / reviewers in the wild / expert
Shitong Shao
dblp:329/2735
· DBLP profile ↗
21ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-4689-6140ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and RefineabstractRecently, enhancing the generative capability of text-to-image (T2I) models has become a promising direction in both academia and industry. Prior studies often focused on either improving generative quality or reducing inference latency, but typically failed to improve both quality and speed simultaneously. Moreover, existing inference-enhancement methods do not achieve significant improvements simultaneously across both diffusion models (DMs) and autoregressive models (ARMs). In this paper, we introduce a general tuning-based inference-enhancement framework, named CoRe$^{2}$2, which is the first to simultaneously achieve significant generative quality and reduced inference overhead across DMs and ARMs, to the best of our knowledge. CoRe$^{2}$2 comprises three stages: Collect, Reflect, and Refine. During the Collect stage, classifier-free guidance (CFG) trajectories are collected and subsequently used in the Reflect stage to train a weak model capable of reflecting the "easy-to-learn" content. Finally, during the Refine stage, CoRe$^{2}$2 can utilize the trained weak model to achieve speedup and performance gain in inference. Specifically, in the early sampling steps, CoRe$^{2}$2 employs weak-to-strong guidance to refine the "difficult-to-learn" and realistic content, thereby improving generative quality. In the later sampling steps, CoRe$^{2}$2 can use the weak model to generate "easy-to-learn" content instead of CFG, dramatically reducing inference time. Experimental outcomes substantiates CoRe$^{2}$2 achieve significant performance improvements on HPD v2, Pick-of-Pic, Drawbench, GenEval, and T2I-Compbench across SDXL, SD3.5, FLUX and LlamaGen. Notably, for SD3.5, CoRe$^{2}$2 can be seamlessly integrated with the state-of-the-art inference-enhancement algorithm Z-Sampling, outperforming it even with less time. Shitong Shao, Zikai Zhou, Dian Xie, Yuetong Fang, Tian Ye 0001, Lichen Bai, Bo Han 0003, Zeke Xie |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | DELT: A Simple Diversity-driven EarlyLate Training for Dataset DistillationabstractRecent advances in dataset distillation have led to solutions in two main directions. The conventional batch-to-batch matching mechanism is ideal for small-scale datasets and includes bi-level optimization methods on models and syntheses, such as FRePo, RCIG, and RaT-BPTT, as well as other methods like distribution matching, gradient matching, and weight trajectory matching. Conversely, batch-to-global matching typifies decoupled methods, which are particularly advantageous for large-scale datasets. This approach has garnered substantial interest within the community, as seen in SRe2L, G-VBSM, WMDD, and CDA. A primary challenge with the second approach is the lack of diversity among syntheses within each class since samples are optimized independently and the same global supervision signals are reused across different synthetic images. In this study, we propose a new Diversity-driven EarlyLate Training (DELT) scheme to enhance the diversity of images in batch-to-global matching with less computation. Our approach is conceptually simple yet effective, it partitions predefined IPC samples into smaller subtasks and employs local optimizations to distill each subset into distributions from distinct phases, reducing the uniformity induced by the unified optimization process. These distilled images from the subtasks demonstrate effective generalization when applied to the entire task. We conduct extensive experiments on CIFAR, Tiny-ImageNet, ImageNet-1K, and its sub-datasets. Our approach outperforms the previous state-of-the-art by 2~5% on average across different datasets and IPCs (images per class), increasing diversity per class by more than 5% while reducing synthesis time by up to 39.3% for enhancing the training efficiency. Ammar Sherif, Zeyuan Yin 0001, Shitong Shao |
CVPR | 4 |
| 2025 | Golden Noise for Diffusion Models: A Learning FrameworkabstractText-to-image diffusion model is a popular paradigm that synthesizes personalized images by providing a text prompt and a random Gaussian noise. While people observe that some noises are ``golden noises'' that can achieve better text-image alignment and higher human preference than others, we still lack a machine learning framework to obtain those golden noises. To learn golden noises for diffusion sampling, we mainly make three contributions in this paper. First, we identify a new concept termed the \textit{noise prompt}, which aims at turning a random Gaussian noise into a golden noise by adding a small desirable perturbation derived from the text prompt. Following the concept, we first formulate the \textit{noise prompt learning} framework that systematically learns ``prompted'' golden noise associated with a text prompt for diffusion models. Second, we design a noise prompt data collection pipeline and collect a large-scale \textit{noise prompt dataset}~(NPD) that contains 100k pairs of random noises and golden noises with the associated text prompts. With the prepared NPD as the training dataset, we trained a small \textit{noise prompt network}~(NPNet) that can directly learn to transform a random noise into a golden noise. The learned golden noise perturbation can be considered as a kind of prompt for noise, as it is rich in semantic information and tailored to the given text prompt. Third, our extensive experiments demonstrate the impressive effectiveness and generalization of NPNet on improving the quality of synthesized images across various diffusion models, including SDXL, DreamShaper-xl-v2-turbo, and Hunyuan-DiT. Moreover, NPNet is a small and efficient controller that acts as a plug-and-play module with very limited additional inference and computational costs, as it just provides a golden noise instead of a random noise without accessing the original pipeline. Zikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang, Zhiqiang Xu 0003, Bo Han 0003, Zeke Xie |
ICCV | 2 |
| 2025 | Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-ReflectionabstractDiffusion models, the most popular generative paradigm so far, can inject conditional information into the generation path to guide the latent towards desired directions. However, existing text-to-image diffusion models often fail to maintain high image quality and high prompt-image alignment for those challenging prompts. To mitigate this issue and enhance existing pretrained diffusion models, we mainly made three contributions in this paper. First, we propose **diffusion self-reflection** that alternately performs denoising and inversion and demonstrate that such diffusion self-reflection can leverage the guidance gap between denoising and inversion to capture prompt-related semantic information with theoretical and empirical evidence. Second, motivated by theoretical analysis, we derive Zigzag Diffusion Sampling (Z-Sampling), a novel self-reflection-based diffusion sampling method that leverages the guidance gap between denosing and inversion to accumulate semantic information step by step along the sampling path, leading to improved sampling results. Moreover, as a plug-and-play method, Z-Sampling can be generally applied to various diffusion models (e.g., accelerated ones and Transformer-based ones) with very limited coding and computational costs. Third, our extensive experiments demonstrate that Z-Sampling can generally and significantly enhance generation quality across various benchmark datasets, diffusion models, and performance evaluation metrics. For example, DreamShaper with Z-Sampling can self-improve with the HPSv2 winning rate up to **94%** over the original results. Moreover, Z-Sampling can further enhance existing diffusion models combined with other orthogonal methods, including Diffusion-DPO. The code is publicly available at
[github.com/xie-lab-ml/Zigzag-Diffusion-Sampling](https://github.com/xie-lab-ml/Zigzag-Diffusion-Sampling). Lichen Bai, Shitong Shao, Zikai Zhou, Zipeng Qi, Zhiqiang Xu 0003, Haoyi Xiong, Zeke Xie |
ICLR | 2 |
| 2025 | IV-mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video SynthesisabstractExploring suitable solutions to improve performance by increasing the computational cost of inference in visual diffusion models is a highly promising direction. Sufficient prior studies have demonstrated that correctly scaling up computation in the sampling process can successfully lead to improved generation quality, enhanced image editing, and compositional generalization. While there have been rapid advancements in developing inference-heavy algorithms for improved image generation, relatively little work has explored inference scaling laws in video diffusion models (VDMs). Furthermore, existing research shows only minimal performance gains that are perceptible to the naked eye. To address this, we design a novel training-free algorithm IV-Mixed Sampler that leverages the strengths of image diffusion models (IDMs) to assist VDMs surpass their current capabilities. The core of IV-Mixed Sampler is to use IDMs to significantly enhance the quality of each video frame and VDMs ensure the temporal coherence of the video during the sampling process. Our experiments have demonstrated that IV-Mixed Sampler achieves state-of-the-art performance on 4 benchmarks including UCF-101-FVD, MSR-VTT-FVD, Chronomagic-Bench-150/1649, and VBench. For example, the open-source Animatediff with IV-Mixed Sampler reduces the UMT-FVD score from 275.2 to 228.6, closing to 223.1 from the closed-source Pika-2.0. Shitong Shao, Zikai Zhou, Bai Lichen, Haoyi Xiong, Zeke Xie |
ICLR | 1 |
| 2025 | Prompt-Enhanced: Leveraging language representation for prompt continual learning
Wei Li 0049, Shitong Shao, Kaizhu Huang, Zhen Lei 0001 |
Neural Networks | 3 |
| 2024 | Generalized Large-Scale Data Condensation via Various Backbone and Statistical MatchingabstractThe lightweight “local-match-global” matching introduced by SRe2L successfully creates a distilled dataset with comprehensive information on the full 224×224 ImageNetlk. However, this one-sided approach is limited to a particular backbone, layer, and statistics, which limits the improvement of the generalization of a distilled dataset. We suggest that sufficient and various “local-match-global” matching are more precise and effective than a single one and have the ability to create a distilled dataset with richer information and better generalization ability. We call this perspective “generalized matching” and propose Generalized Various Backbone and Statistical Matching (G-VBSM) in this work, which aims to create a synthetic dataset with densities, ensuring consistency with the complete dataset across various backbones, layers, and statistics. As experimentally demonstrated, G-VBSM is the first algorithm to obtain strong performance across both small-scale and large-scale datasets. Specifically, G-VBSM achieves performances of 38.7% on CIFAR-I00, 47.6% on Tiny-ImageNet, and 31.4% on the full 224×224 ImageNet1 k, respectively11Settings: CIFAR-I00 with 128-width ConvNet under 10 images per class (lPC), Tiny-ImageNet with ResNet18 under 50 IPC, and ImageNetlk with ResNet18 under 10 IPC.. These results surpass all SOTA methods by margins of 3.9%, 6.5%, and 10.1%, respectively. Shitong Shao, Zeyuan Yin 0001, Muxin Zhou |
CVPR | 1 |
| 2024 | Auto-DAS: Automated Proxy Discovery for Training-Free Distillation-Aware Architecture Search
Haosen Sun, Lujun Li 0001, Peijie Dong, Zimian Wei, Shitong Shao |
ECCV (5) | 5 |
| 2024 | Rethinking Centered Kernel Alignment in Knowledge Distillation
Zikai Zhou, Yunhang Shen, Shitong Shao, Linrui Gong, Shaohui Lin |
IJCAI | 3 |
| 2024 | Diffusion Models are Certifiably Robust ClassifiersabstractGenerative learning, recognized for its effective modeling of data distributions, offers inherent advantages in handling out-of-distribution instances, especially for enhancing robustness to adversarial attacks. Among these, diffusion classifiers, utilizing powerful diffusion models, have demonstrated superior empirical robustness. However, a comprehensive theoretical understanding of their robustness is still lacking, raising concerns about their vulnerability to stronger future attacks. In this study, we prove that diffusion classifiers possess $O(1)$ Lipschitzness, and establish their certified robustness, demonstrating their inherent resilience. To achieve non-constant Lipschitzness, thereby obtaining much tighter certified robustness, we generalize diffusion classifiers to classify Gaussian-corrupted data. This involves deriving the evidence lower bounds (ELBOs) for these distributions, approximating the likelihood using the ELBO, and calculating classification probabilities via Bayes' theorem. Experimental results show the superior certified robustness of these Noised Diffusion Classifiers (NDCs). Notably, we achieve over 80\% and 70\% certified robustness on CIFAR-10 under adversarial perturbations with \(\ell_2\) norms less than 0.25 and 0.5, respectively, using a single off-the-shelf diffusion model without any additional data. Huanran Chen, Yinpeng Dong, Shitong Shao, Zhongkai Hao, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001 |
NeurIPS | 3 |
| 2024 | Elucidating the Design Space of Dataset CondensationabstractDataset condensation, a concept within $\textit{data-centric learning}$, aims to efficiently transfer critical attributes from an original dataset to a synthetic version, meanwhile maintaining both diversity and realism of syntheses. This approach can significantly improve model training efficiency and is also adaptable for multiple application areas. Previous methods in dataset condensation have faced several challenges: some incur high computational costs which limit scalability to larger datasets ($\textit{e.g.,}$ MTT, DREAM, and TESLA), while others are restricted to less optimal design spaces, which could hinder potential improvements, especially in smaller datasets ($\textit{e.g.,}$ SRe$^2$L, G-VBSM, and RDED). To address these limitations, we propose a comprehensive designing-centric framework that includes specific, effective strategies like implementing soft category-aware matching, adjusting the learning rate schedule and applying small batch-size. These strategies are grounded in both empirical evidence and theoretical backing. Our resulting approach, $\textbf{E}$lucidate $\textbf{D}$ataset $\textbf{C}$ondensation ($\textbf{EDC}$), establishes a benchmark for both small and large-scale dataset condensation. In our testing, EDC achieves state-of-the-art accuracy, reaching 48.6% on ImageNet-1k with a ResNet-18 model at an IPC of 10, which corresponds to a compression ratio of 0.78\%. This performance surpasses those of SRe$^2$L, G-VBSM, and RDED by margins of 27.3%, 17.2%, and 6.6%, respectively. Code is available at: https://github.com/shaoshitong/EDC. Shitong Shao, Zikai Zhou, Huanran Chen |
NeurIPS | 1 |
| 2024 | Multi-perspective analysis on data augmentation in knowledge distillationabstractKnowledge distillation stands as a capable technique for transferring knowledge from a larger to a smaller model, thereby notably enhancing the smaller model’s performance. In the recent past, data augmentation has been employed in contrastive learning based knowledge distillation techniques yielding superior results. Despite the significant role of data augmentation, its value remains underappreciated within the domain of knowledge distillation, with no in-depth analysis in the literature thus far. To make up for this oversight, we conduct a multi-perspective theoretical and experimental analysis on the role that data augmentation can play in knowledge distillation. We summarize the properties of data augmentation and list the core findings as follows. (a) Our investigations validate that data augmentation significantly boosts the performance of knowledge distillation on the tasks of image classification and object detection. And this holds true even if the teacher model lacks comprehensive information about the augmented samples. Moreover, our novel J oint D ata A ugmentation (JDA) approach outperforms single data augmentation in knowledge distillation. (b) The pivotal role of data augmentation in knowledge distillation can be theoretically explained via Sharpness-Aware Minimization. (c) The compatibility of data augmentation with various knowledge distillation methods can enhance their performance. In light of these observations, we propose a new method called C osine C onfidence D istillation (CCD) for more reasonable knowledge transfer from augmented samples. Experimental results not only demonstrate that CCD becomes the state-of-the-art method with less storage requirement on CIFAR-100 and ImageNet-1k, but also validate the superiority of CCD over DIST on the object detection benchmark dataset, MS-COCO. Wei Li 0049, Shitong Shao, Ziming Qiu, Aiguo Song |
Neurocomputing | 2 |
| 2024 | Attention-Based Intrinsic Reward Mixing Network for Credit Assignment in Multiagent Reinforcement LearningabstractCredit assignment is a critical problem in cooperative Multi-Agent Reinforcement Learning (MARL). To address this problem, current studies mainly rely on the intrinsic reward, which is directly summed with the global reward to generate a total reward. However, such kinds of intrinsic reward functions ignore the dependence among agents and inevitably limit the adaptivity and effectiveness of MARL methods. In this paper, we propose a novel method, Attention-based Intrinsic Reward Mixing Network (AIRMN), for credit assignment in MARL. Specifically, we design a new intrinsic reward network on the basis of the attention mechanism, in order to enhance the effectiveness of teamwork. Besides, we devise a new mixing network that combines the intrinsic and extrinsic rewards in a nonlinear and dynamic manner, so as to adapt the total reward to the variation of the environment. Experimental results on the battle games of StarCraft II demonstrate that AIRMN outperforms the state-of-the-art methods in terms of the average test win rate, and also validate that AIRMN can dynamically return the precise intrinsic reward to each agent based on their contributions to the team cooperation, thereby better dealing with the credit assignment problem. Wei Li 0049, Weiyan Liu, Shitong Shao, Shiyi Huang, Aiguo Song |
IEEE Trans. Games | 3 |
| 2024 | MDDP: Making Decisions From Different Perspectives in Multiagent Reinforcement LearningabstractMultiagent reinforcement learning (MARL) has made remarkable progress in recent years. However, in most MARL methods, agents share a policy or value network, which is easy to result in similar behaviors of agents, and thus, limits the flexibility of the method to handle complex tasks. To enhance the diversity of agent behaviors, we propose a novel method, making decisions from different perspectives (MDDP). This method enables agents to switch flexibly between different policy roles and make decisions from different perspectives, which can improve the adaptability of policy learning in complex scenarios. Specifically, in MDDP, we design a new self-attention and gated recurrent unit (GRU)-based dueling architecture network (SG-DAN) to estimate the individual$Q$-values. SG-DAN contains two components: 1) the new self-attention-based role-switching network (SAR) and the capable GRU-based state value estimation network (GSE). SAR takes charge of action advantage estimation and GSE is responsible for state value estimation. Experimental results on the challengingStarCraftII micromanagement benchmark not only verify the modeling reasonability of MDDP but also demonstrate its performance superiority over the related advanced approaches. Wei Li 0049, Ziming Qiu, Shitong Shao, Aiguo Song |
IEEE Trans. Games | 3 |
| 2023 | Teaching What You Should Teach: A Data-Based Distillation MethodabstractIn real teaching scenarios, an excellent teacher always teaches what he (or she) is good at but the student is not. This gives the student the best assistance in making up for his (or her) weaknesses and becoming a good one overall. Enlightened by this, we introduce the "Teaching what you Should Teach" strategy into a knowledge distillation framework, and propose a data-based distillation method named "TST" that searches for desirable augmented samples to assist in distilling more efficiently and rationally. To be specific, we design a neural network-based data augmentation module with priori bias to find out what meets the teacher's strengths but the student's weaknesses, by learning magnitudes and probabilities to generate suitable data samples. By training the data augmentation module and the generalized distillation paradigm alternately, a student model is learned with excellent generalization ability. To verify the effectiveness of our method, we conducted extensive comparative experiments on object recognition, detection, and segmentation tasks. The results on the CIFAR-100, ImageNet-1k, MS-COCO, and Cityscapes datasets demonstrate that our method achieves state-of-the-art performance on almost all teacher-student pairs. Furthermore, we conduct visualization studies to explore what magnitudes and probabilities are needed for the distillation process. Shitong Shao, Huanran Chen, Zhen Huang 0007, Linrui Gong, Shuai Wang 0048, Xinxiao Wu |
IJCAI | 1 |
| 2023 | Spatial-Temporal Constraint Learning for Cross-Subject EEG-Based Emotion RecognitionabstractRecent researches combine domain adaptation methods with elaborate feature extractors to better learn domain-invariant and discriminative features for cross-subject Electroencephalogram(EEG)-based emotion recognition. Existing models only utilize domain adaptation to constrain spatial learning or temporary learning, though the domain shift will possibly appear in both spatial and temporal learning stages. And some models simply treat the different subjects in the source domain as a whole, ignoring the data structure of the source domain. Motivated by the above problems, we design a novel model, Spatial-Temporal Constraint Learning (STCL), which adopts the Multi-Layer Perceptron (MLP) and Transformer Encoder for spatial and temporal features learning, respectively. In the spatial learning stage, we design Multi-Subject Prototypes Alignment (MSPA), which treats different subjects as different domains. In the temporary learning stage, we utilize the adversarial training strategy which treats different subjects in the source domain as a whole to further narrow down the domain gap. In addition, to improve the representative ability of our model in the target domain, we put forward Target Samples Selective Strategy (TSSS) which selects the samples from the target domain with reliable pseudo-labels for training STCL. The cross-subject experiments on two benchmark datasets have demonstrated the effectiveness of our model. Our model achieves the accuracies of 83.43 %, 78.36 %, and 80.09 % for three sessions respectively on SEED, 60.51 % for valence classification, and 63.68 % for arousal classification on DEAP. Wei Li 0049, Shitong Shao, Wei Huan, Ye Tian 0034 |
IJCNN | 3 |
| 2023 | Hybrid knowledge distillation from intermediate layers for efficient Single Image Super-Resolution
Jiao Xie, Linrui Gong, Shitong Shao, Shaohui Lin, Linkai Luo |
Neurocomputing | 3 |
| 2023 | MS-FRAN: A Novel Multi-Source Domain Adaptation Method for EEG-Based Emotion RecognitionabstractElectroencephalogram (EEG)-based emotion recognition has gradually become a research hotspot. However, the large distribution differences of EEG signals across subjects make the current research stuck in a dilemma. To resolve this problem, in this article, we propose a novel and effective method, Multi-Source Feature Representation and Alignment Network (MS-FRAN). The effectiveness of proposed method mainly comes from three new modules: Wide Feature Extractor (WFE) for feature learning, Random Matching Operation (RMO) for model training, and Top- h ranked domain classifier selection (TOP) for emotion classification. MS-FRAN is not only effective in aligning the distributions of each pair of source and target domains, but also capable of reducing the distributional differences among the multiple source domains. Experimental results on the public benchmark datasets SEED and DEAP have demonstrated the advantage of our method over the related competitive approaches for cross-subject EEG-based emotion recognition. Wei Li 0049, Wei Huan, Shitong Shao, Aiguo Song |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | What Role Does Data Augmentation Play in Knowledge Distillation?
Wei Li 0049, Shitong Shao, Weiyan Liu, Ziming Qiu, Wei Huan |
ACCV (2) | 2 |
| 2022 | AIIR-MIX: Multi-Agent Reinforcement Learning Meets Attention Individual Intrinsic Reward Mixing Network
Wei Li 0049, Weiyan Liu, Shitong Shao, Shiyi Huang |
ACML | 3 |
| 2022 | BiSMSM: A Hybrid MLP-Based Model of Global Self-Attention Processes for EEG-Based Emotion Recognition
Wei Li 0049, Ye Tian 0034, Jianzhang Dong, Shitong Shao |
ICANN (1) | 5 |