VLDB 2026 Research / reviewers in the wild / expert
Jiayu Xiao
dblp:255/7118
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exact: Exploring Space-Time Perceptive Clues for Weakly Supervised Satellite Image Time Series Semantic SegmentationabstractAutomated crop mapping through Satellite Image Time Series (SITS) has emerged as a crucial avenue for agricultural monitoring and management. However, due to the low resolution and unclear parcel boundaries, annotating pixel-level masks is exceptionally complex and time-consuming in SITS. This paper embraces the weakly supervised paradigm (i.e., only image-level categories available) to liberate the crop mapping task from the exhaustive annotation burden. The unique characteristics of SITS give rise to several challenges in weakly supervised learning: (1) noise perturbation from spatially neighboring regions, and (2) erroneous semantic bias from anomalous temporal periods. To address the above difficulties, we propose a novel method, termed exploring space-time perceptive clues (Exact). First, we introduce a set of spatial clues to explicitly capture the representative patterns of different crops from the most class-relative regions. Besides, we leverage the temporal-to-class interaction of the model to emphasize the contributions of pivotal clips, thereby enhancing the model perception for crop regions. Building upon the space-time perceptive clues, we derive the clue-based CAMs to effectively supervise the SITS segmentation network. Our method demonstrates impressive performance on various SITS benchmarks. Remarkably, the segmentation network trained on Exact-generated masks achieves 95% of its fully supervised performance, showing the bright promise of weakly supervised paradigm in crop mapping scenario. Our code will be publicly available here. Jiayu Xiao, Tianxiang Xiao, Yike Ma |
CVPR | 3 |
| 2024 | ADIFT: Zero-Shot Generative Model Adaption Via Adaptive Domain-Invariant Feature TransferabstractCLIP-guided zero-shot image generative model adaption methods only require textual domain labels without any target domain images, but there are some dilemmas remain unsolved, such as identity degradation and pattern overfitting. To address these issues, an adaptive domain-invariant feature transfer (ADIFT) method is proposed. It makes the target domain generator learn domain-invariant features from the source domain generator but learn domain-variant features from the CLIP space. We first introduce a local self-similarity map to represent and preserve the image identity features, and then add a parameter learnable point-wise gate module on the alignment path of the local self-similarity maps to transfer cross-domain features adaptively. Qualitative and quantitative experimental results validate that the proposed ADIFT solves the problems of identity degradation and pattern over-fitting effectively. Chaofei Wang, Xiangan Zhao, Jiayu Xiao, Guotong Geng |
ICASSP | 5 |
| 2024 | R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image GenerationabstractRecent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout instruction. In this work, we probe into zero-shot grounded T2I generation with diffusion models, that is, generating images corresponding to the input layout information without training auxiliary modules or finetuning diffusion models. We propose a **R**egion and **B**oundary (R&B) aware cross-attention guidance approach that gradually modulates the attention maps of diffusion model during generative process, and assists the model to synthesize images (1) with high fidelity, (2) highly compatible with textual input, and (3) interpreting layout instructions accurately. Specifically, we leverage the discrete sampling to bridge the gap between consecutive attention maps and discrete layout constraints, and design a region-aware loss to refine the generative layout during diffusion process. We further propose a boundary-aware loss to strengthen object discriminability within the corresponding regions. Experimental results show that our method outperforms existing state-of-the-art zero-shot grounded T2I generation methods by a large margin both qualitatively and quantitatively on several benchmarks.
Project page: https://sagileo.github.io/Region-and-Boundary. Jiayu Xiao, Henglei Lv, Liang Li 0003, Shuhui Wang, Qingming Huang |
ICLR | 1 |
| 2024 | MISA: MIning Saliency-Aware Semantic Prior for Box Supervised Instance Segmentation
Jiayu Xiao, Yike Ma |
IJCAI | 3 |
| 2024 | Towards Robustness and Diversity: Continual Learning in Dialog Generation with Text-Mixup and Batch Nuclear-Norm MaximizationabstractIn our dynamic world where data arrives in a continuous stream, continual learning enables us to incrementally add new tasks/domains without the need to retrain from scratch. A major challenge in continual learning of language model is catastrophic forgetting, the tendency of models to forget knowledge from previously trained tasks/domains when training on new ones. This paper studies dialog generation under the continual learning setting. We propose a novel method that 1) uses Text-Mixup as data augmentation to avoid model overfitting on replay memory and 2) leverages Batch-Nuclear Norm Maximization (BNNM) to alleviate the problem of mode collapse. Experiments on a 37-domain task-oriented dialog dataset and DailyDialog (a 10-domain chitchat dataset) demonstrate that our proposed approach outperforms the state-of-the-art in continual learning. Jiayu Xiao, Mengxiang Li, Zhongjiang He, Shuangyong Song |
IJCNN | 2 |
| 2024 | Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
Henglei Lv, Jiayu Xiao, Liang Li 0003 |
ACM Multimedia | 2 |
| 2023 | Text-Driven Generative Domain Adaptation with Spectral Consistency RegularizationabstractCombined with the generative prior of pre-trained models and the flexibility of text, text-driven generative domain adaptation can generate images from a wide range of target domains. However, current methods still suffer from overfitting and the mode collapse problem. In this paper, we analyze the mode collapse from the geometric point of view and reveal its relationship to the Hessian matrix of generator. To alleviate it, we propose the spectral consistency regularization to preserve the diversity of source domain without restricting the semantic adaptation to target domain. We also design granularity adaptive regularization to flexibly control the balance between diversity and stylization for target model. We conduct experiments for broad target domains compared with state-of-the-art methods and extensive ablation studies. The experiments demonstrate the effectiveness of our method to preserve the diversity of source domain and generate high fidelity target images. Source code has been released in https://github.com/Victarry/Adaptation-SCR. Zhenhuan Liu, Liang Li 0003, Jiayu Xiao, Zhengjun Zha, Qingming Huang |
ICCV | 3 |
| 2022 | Few Shot Generative Model Adaption via Relaxed Spatial Structural AlignmentabstractTraining a generative adversarial network (GAN) with limited data has been a challenging task. A feasible solution is to start with a GAN well-trained on a large scale source domain and adapt it to the target domain with a few samples, termed as few shot generative model adaption. However, existing methods are prone to model overfitting and collapse in extremely few shot setting (less than 10). To solve this problem, we propose a relaxed spatial structural alignment (RSSA) method to calibrate the target generative models during the adaption. We design a cross-domain spatial structural consistency loss comprising the self-correlation and disturbance correlation consistency loss. It helps align the spatial structural information between the synthesis image pairs of the source and target domains. To relax the cross-domain alignment, we compress the original latent space of generative models to a subspace. Image pairs generated from the subspace are pulled closer. Qualitative and quantitative experiments show that our method consistently surpasses the state-of-the-art methods in few shot setting. Our source code: https://github.com/StevenShaw1999/RSSA. Jiayu Xiao, Liang Li 0003, Chaofei Wang, Zhengjun Zha, Qingming Huang |
CVPR | 1 |
| 2022 | A fast neighborhood classifier based on hash bucket with application to medical diagnosis
Jiayu Xiao, Qinghua Zhang 0001, Zhihua Ai, Guoyin Wang 0001 |
Int. J. Approx. Reason. | 1 |
| 2021 | Towards Learning Spatially Discriminative Feature RepresentationsabstractThe backbone of traditional CNN classifier is generally considered as a feature extractor, followed by a linear layer which performs the classification. We propose a novel loss function, termed as CAM-loss, to constrain the embedded feature maps with the class activation maps (CAMs) which indicate the spatially discriminative regions of an image for particular categories. CAM-loss drives the backbone to express the features of target category and suppress the features of non-target categories or background, so as to obtain more discriminative feature representations. It can be simply applied in any CNN architecture with neglectable additional parameters and calculations. Experimental results show that CAM-loss is applicable to a variety of network structures and can be combined with mainstream regularization methods to improve the performance of image classification. The strong generalization ability of CAMloss is validated in the transfer learning and few shot learning tasks. Based on CAM-loss, we also propose a novel CAAM-CAM matching knowledge distillation method. This method directly uses the CAM generated by the teacher network to supervise the CAAM generated by the student network, which effectively improves the accuracy and convergence rate of the student network. Chaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang, Shiji Song, Gao Huang 0001 |
ICCV | 2 |