VLDB 2026 Research / reviewers in the wild / expert
Kaiyu Song
dblp:321/8769
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0007-6443-7104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flow-Matching Posterior Sampling: A Training-Free Conditional Generation for Flow MatchingabstractTraining-free conditional generation based on flow matching aims to leverage pre-trained unconditional flow matching models to perform conditional generation without retraining. Recently, a successful training-free conditional generation approach incorporates conditions via posterior sampling, which relies on the availability of a score function in the unconditional diffusion model. However, flow matching models lack an explicit score function, rendering this strategy inapplicable. Approximate posterior sampling for flow matching has been explored, but it is limited to linear inverse problems. In this paper, we propose Flow Matching-based Posterior Sampling (FMPS) to broaden its scope of application. We introduce a correction term by steering the velocity field. This correction term can be reformulated to incorporate a surrogate score function, thereby bridging the gap between flow matching models and score-based posterior sampling. Hence, FMPS enables posterior sampling to be adjusted within the flow-matching framework. Furthermore, we propose two practical implementations of the correction mechanism: one to improve generation quality and the other to enhance computational efficiency. Experimental results on diverse conditional generation tasks demonstrate that our method achieves superior generation quality compared to existing state-of-the-art approaches, validating the effectiveness and generality of FMPS. Kaiyu Song, Hanjiang Lai, Yan Pan 0002, Kun Yue, Jian Yin 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Test-time Alignment-Enhanced Adapter for Vision-Language ModelsabstractTest-time adaptation with pre-trained vision-language models (VLMs) has attracted increasing attention for tackling the issue of distribution shift during the test phase. While prior methods have shown effectiveness in addressing distribution shift by adjusting classification logits, they are not optimal due to keeping text features unchanged. To address this issue, we introduce a new approach called Test-time Alignment-Enhanced Adapter (TAEA), which trains an adapter with test samples to adjust text features during the test phase. We can enhance the text-to-image alignment prediction by utilizing an adapter to adapt text features. Furthermore, we also propose to adopt the negative cache from TDA as enhancement module, which further improves the performance of TAEA. Our approach outperforms the state-of-the-art TTA method of pre-trained VLMs by an average of 0.75% on the out-of-distribution benchmark and 2.5% on the cross-domain benchmark, with an acceptable training time. Code will be available at https://github.com/BaoshunWq/clip-TAEA. Baoshun Tong, Kaiyu Song, Hanjiang Lai |
ICASSP | 2 |
| 2025 | Enhancing Few-Shot Out-of-Distribution Detection with Gradient Aligned Context OptimizationabstractFew-shot out-of-distribution (OOD) detection aims to detect OOD images from unseen classes with only a few labeled in-distribution (ID) images. To detect OOD images and classify ID samples, prior methods have been proposed by regarding the background regions of ID samples as the OOD knowledge and performing OOD regularization and ID classification optimization. However, the gradient conflict still exists between ID classification optimization and OOD regularization caused by biased recognition. To address this issue, we present Gradient Aligned Context Optimization (GaCoOp) to mitigate this gradient conflict. Specifically, we decompose the optimization gradient to identify the scenario when the conflict occurs. Then we alleviate the conflict in inner ID samples and optimize the prompts via leveraging gradient projection. Extensive experiments over the large-scale ImageNet OOD detection benchmark demonstrate that our GaCoOp can effectively mitigate the conflict and achieve great performance. Code will be available at https://github.com/BaoshunWq/ood-GaCoOp. Baoshun Tong, Kaiyu Song, Hanjiang Lai |
ICASSP | 2 |
| 2025 | CPMDiff: Classifier Probability Measurement for Out-of-Distribution Detection via Diffusion ModelsabstractRecent research has explored using diffusion models for out-of-distribution (OOD) detection, leveraging their ability to distinguish OOD samples by measuring reconstruction errors. Existing methods use visual feature metrics to measure reconstruction errors. However, the inherent stochasticity in diffusion models introduces variability, distorting reconstruction error measurements with visual feature metrics. To address this issue, we propose Classifier Probability Measurement via Diffusion Models (CPMDiff), a novel OOD detection method that measures the classifier probabilities’ discrepancy between input and reconstructed data. By leveraging classifier space measurement, CPMDiff alleviates the distortion in reconstruction error metrics caused by the stochasticity of diffusion models. Experiments on benchmark datasets demonstrate that CPMDiff significantly improves the performance of diffusion model-based methods in OOD detection, especially on challenging datasets. Yongheng Xu, Kaiyu Song, Hanjiang Lai |
ICME | 2 |
| 2024 | MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion ModelabstractDeep neural networks (DNNs) are vulnerable to adversarial perturbation, where an imperceptible perturbation is added to the image that can fool the DNNs. Diffusion-based adversarial purification uses the diffusion model to generate a clean image against such adversarial attacks. Un-fortunately, the generative process of the diffusion model is also inevitably affected by adversarial perturbation since the diffusion model is also a deep neural network where its input has adversarial perturbation. In this work, we propose MimicDiffusion, a new diffusion-based adversarial purification technique that directly approximates the generative process of the diffusion model with the clean image as input. Concretely, we analyze the differences between the guided terms using the clean image and the ad-versarial sample. After that, we first implement MimicD-iffusion based on Manhattan distance. Then, we propose two guidance to purify the adversarial perturbation and ap-proximate the clean diffusion model. Extensive experiments on three image datasets, including CIFAR-10, CIFAR-100, and ImageNet, with three classifier backbones including WideResNet-70-16, WideResNet-28-10, and ResNet-50 demonstrate that MimicDiffusion significantly performs better than the state-of-the-art baselines. On CIFAR -10, CIFAR-100, and ImageNet, it achieves 92.67%, 61.35%, and 61.53% average robust accuracy, which are 18.49%, 13.23%, and 17.64% higher, respectively. The code is available at https://github.com/psky1111/MimicDiffusion. Kaiyu Song, Hanjiang Lai, Yan Pan 0002, Jian Yin 0001 |
CVPR | 1 |
| 2024 | MALIP: Improving Few-Shot Image Classification with Multimodal Fusion EnhancementabstractWith the significant progress in pre-trained visionlanguage models like CLIP, recent CLIP-based methods have shown impressive performance in few-shot tasks. However, CLIPbased representations have a natural gap in downstream few-shot tasks due to the label-related multimodal information scarcity caused by limited data. We then question, whether the generative model trained in downstream tasks could be used to enhance label-related multimodal fusion. In this paper, we propose a generative model-based multimodal fusion enhancement method, MALIP, to improve the few-shot performance of CLIP via a Multimodal Adapter module. Specifically, we first leverage a variational autoencoder (VAE) that could be trained in the few-shot scenario to extend the data. Then we create adapter weights by a key-value cache model constructed from the image and text information based on the expanded data. In the end, through extensive experiments on 11 datasets, we demonstrate the effectiveness of MALIP to perform state-of-the-art few-shot image classification. Kaifen Cai, Kaiyu Song, Yan Pan 0002, Hanjiang Lai |
ICME | 2 |
| 2024 | DNAF: Diffusion with Noise-Aware Feature for Pose-Guided Person Image SynthesisabstractPose-guided person image synthesis aims at generating images based on the related pose skeleton and the appearance of a source image. As a popular generative model, the diffusion model shows its potential. However, there are two gaps to hinder the fusion between pose information and appearance: 1) Directly injecting pixel-level pose information into semantic features leads to the representation gap. 2) The timestep-dependent nature of the diffusion model introduces the noise-induced gap. To alleviate these, we propose Diffusion with Noise-Aware Feature(DNAF). Concretely, we leverage the T2I-Adapter-based pose adapter to achieve the mapping from the pixel level to the feature level. Then, we propose a lightweight trainable layer to infuse the multi-scale constant feature adaptively. In the end, we construct noise-aware features to more effectively guide the diffusion process. Experimental results show that DNAF achieves competitive results on DeepFashion and Market-1501 datasets. Liyan Guo, Kaiyu Song, Mengying Xu, Hanjiang Lai |
ICME | 2 |
| 2022 | Mutual information based Bayesian graph neural network for few-shot learningabstractIn the deep neural network based few-shot learning, the limited training data may make the neural network extract ineffective features, which leads to inaccurate results. By Bayesian graph neural network (BGNN), the probability distributions on hidden layers imply useful features, and the few-shot learning could improved by establishing the correlation among features. Thus, in this paper, we incorporate mutual information (MI) into BGNN to describe the correlation, and propose an innovative framework by adopting the Bayesian network with continuous variables (BNCV) for effective calculation of MI. First, we build the BNCV simultaneously when calculating the probability distributions of features from the Dropout in hidden layers of BGNN. Then, we approximate the MI values efficiently by probabilistic inferences over BNCV. Finally, we give the correlation based loss function and training algorithm of our BGNN model. Experimental results show that our MI based BGNN framework is effective for few-shot learning and outperforms some state-of-the-art competitors by large margins on accuracy. Kaiyu Song, Kun Yue, Liang Duan, Mingze Yang, Angsheng Li |
UAI | 1 |
| 2020 | An Efficient Approach for Parameters Learning of Bayesian Network with Multiple Latent Variables Using Neural Networks and P-EM
Kaiyu Song, Kun Yue, Jia Hao 0001 |
CollaborateCom (1) | 1 |