EDBT 2026 Demo / reviewers in the wild / expert
Xuechao Zou
dblp:341/8058
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-3743-7738ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial GenerationabstractWhile significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this challenge, we introduce MDiTFace—a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminating discrepancies between heterogeneous modality representations. The framework facilitates comprehensive multimodal feature interaction through stacked, newly designed multivariate transformer blocks that process all conditions synchronously. Additionally, we design a novel decoupled attention mechanism by dissociating implicit dependencies between mask tokens and temporal embeddings. This mechanism segregates internal computations into dynamic and static pathways, enabling caching and reuse of features computed in static pathways after initial calculation, thereby reducing additional computational overhead introduced by mask condition by over 94% while maintaining performance. Extensive experiments demonstrate that MDiTFace significantly outperforms other competing methods in terms of both facial fidelity and conditional consistency. Yushe Cao, Dian-xi Shi, Xuechao Zou, Haikuo Peng, Chun Yu, Junliang Xing |
AAAI | 4 |
| 2026 | Hierarchical fusion of local and global visual features with mixture-of-experts for remote sensing image scene classification
Yuanhao Tang, Xuechao Zou, Zhengpei Hu, Junliang Xing, Jianqiang Huang 0002 |
Neurocomputing | 2 |
| 2025 | UV-Mamba: A DCN-Enhanced State Space Model for Urban Village Boundary Identification in High-Resolution Remote Sensing ImagesabstractDue to the diverse geographical environments, intricate landscapes, and high-density settlements, the automatic identification of urban village boundaries using remote sensing images remains a highly challenging task. This paper proposes a novel and efficient neural network model called UV-Mamba for accurate boundary detection in high-resolution remote sensing images. UV-Mamba mitigates the memory loss problem in lengthy sequence modeling, which arises in state space models (SSM) with increasing image size, by incorporating deformable convolutions (DCN). Its architecture utilizes an encoder-decoder framework and includes an encoder with four deformable state space augmentation (DSSA) blocks for efficient multi-level semantic extraction and a decoder to integrate the extracted semantic information. We conducted experiments on two large datasets showing that UV-Mamba achieves state-of-the-art performance. Specifically, our model achieves 73.3% and 78.1% IoU on the Beijing and Xi’an datasets, respectively, representing improvements of 1.2% and 3.4% IoU over the previous best model while also being 6× faster in inference speed and 40× smaller in parameter count. Source code and pre-trained models are available at https://github.com/Devin-Egber/UV-Mamba. Lulin Li, Xuechao Zou, Junliang Xing, Pin Tao |
ICASSP | 3 |
| 2025 | Dynamic Dictionary Learning for Remote Sensing Image Segmentation
Xuechao Zou, Kai Li 0023, Pin Tao, Junliang Xing, Congyan Lang |
ICCV | 1 |
| 2025 | Knowledge Transfer and Domain Adaptation for Fine-Grained Remote Sensing Image SegmentationabstractFine-Grained remote sensing image segmentation is essential for accurately identifying detailed objects in remote sensing images. Recently, vision transformer models (VTMs) pretrained on large-scale datasets have demonstrated strong zero-shot generalization. However, directly applying them to specific tasks may lead to domain shift. We introduce a novel end-to-end learning paradigm combining knowledge guidance with domain refinement to enhance performance. We present two key components: the Feature Alignment Module (FAM) and the Feature Modulation Module (FMM). FAM aligns features from a CNN-based backbone with those from the pretrained VTM’s encoder using channel transformation and spatial interpolation, and transfers knowledge via KL divergence and L2 normalization constraint. FMM further adapts the knowledge to the specific domain to address domain shift. We also introduce a fine-grained grass segmentation dataset and demonstrate, through experiments on two datasets, that our method achieves a significant improvement of 2.57 mIoU on the grass dataset and 3.73 mIoU on the cloud dataset. The results highlight the potential of combining knowledge transfer and domain adaptation to overcome domain-related challenges and data limitations. The project page is available at https://xavierjiezou.github.io/KTDA/. Xuechao Zou, Kai Li 0023, Congyan Lang, Pin Tao |
ICME | 2 |
| 2025 | A Novel Two-Stream Algorithm for Spatio-Temporal Fusion in Remote SensingabstractThe spatio-temporal fusion technology can effectively solve the problem of missing time series data. However, existing algorithms often struggle to accurately capture surface feature changes and fine details. To overcome these limitations, this letter introduces FemSTF, a flexible two-stream algorithm for remote sensing spatio-temporal fusion, which consists of three main stages: image sharpening, image fusion, and edge optimization. In the first stage, techniques such as interpolation, classification, and blending are employed to address missing pixel values caused by minimal input data. In the second stage, a two-stream strategy is introduced to enhance fusion accuracy, enabling the algorithm to capture regional detail changes with high sensitivity to surface variations. The third stage focuses on edge optimization, generating high-resolution predictive images under minimal input conditions (three images) while retaining rich texture and edge details. Notably, FemSTF achieves a performance comparable to that of methods requiring five input images. Experimental results demonstrate that FemSTF outperforms mainstream algorithms across multiple datasets. Ablation experiments confirm the effectiveness of the two-stream strategy, highlighting its role in enhancing remote sensing data processing accuracy. This study offers an efficient solution for spatio-temporal fusion, demonstrating great potential in the field. Xuechao Zou, Pin Tao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing ImagesabstractCloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. Code and model checkpoints are available at https://xavierjiezou.github.io/Cloud-Adapter/. Xuechao Zou, Kai Li 0023, Junliang Xing, Lei Jin 0003, Congyan Lang, Pin Tao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | LEFormer: A Hybrid CNN-Transformer Architecture for Accurate Lake Extraction from Remote Sensing ImageryabstractLake extraction from remote sensing images is challenging due to the complex lake shapes and inherent data noises. Existing methods suffer from blurred segmentation boundaries and poor foreground modeling. This paper proposes a hybrid CNN-Transformer architecture, called LEFormer, for accurate lake extraction. LEFormer contains three main modules: CNN encoder, Transformer encoder, and cross-encoder fusion. The CNN encoder effectively recovers local spatial information and improves fine-scale details. Simultaneously, the Transformer encoder captures long-range dependencies between sequences of any length, allowing them to obtain global features and context information. The cross-encoder fusion module integrates the local and global features to improve mask prediction. Experimental results show that LEFormer consistently achieves state-of-the-art performance and efficiency on the Surface Water and the Qinghai-Tibet Plateau Lake datasets. Specifically, LEFormer achieves 90.86% and 97.42% mIoU on two datasets with a parameter count of 3.61M, respectively, while being 20× minor than the previous best lake extraction method. The source code is available at https://github.com/BastianChen/LEFormer. Xuechao Zou, Yu Zhang 0165, Jiayu Li 0008, Kai Li 0023, Junliang Xing, Pin Tao |
ICASSP | 2 |
| 2024 | ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal PredictionabstractSpatiotemporal prediction aims to generate future sequences by paradigms learned from historical contexts. It is essential in numerous domains, such as traffic flow prediction and weather forecasting. Recently, research in this field has been predominantly driven by deep neural networks based on autoencoder architectures. However, existing methods commonly adopt autoencoder architectures with identical receptive field sizes. To address this issue, we propose an Asymmetric Receptive Field Autoencoder (ARFA) model, which introduces corresponding sizes of receptive field modules tailored to the distinct functionalities of the encoder and decoder. In the encoder, we present a large kernel module for global spatiotemporal feature extraction. In the decoder, we develop a small kernel module for local spatiotemporal information reconstruction. Experimental results demonstrate that ARFA consistently achieves state-of-the-art performance on popular datasets. Additionally, we construct the RainBench, a large-scale radar echo dataset for precipitation prediction, to address the scarcity of meteorological data in the domain. Xuechao Zou, Xiaoying Wang 0002, Jianqiang Huang 0002, Junliang Xing |
ICASSP | 2 |
| 2024 | High-Fidelity Lake Extraction Via Two-Stage Prompt Enhancement: Establishing A Novel Baseline and BenchmarkabstractLake extraction from remote sensing imagery is a complex challenge due to the varied lake shapes and data noise. Current methods rely on multispectral image datasets, making it challenging to learn lake features accurately from pixel arrangements. This, in turn, affects model learning and the creation of accurate segmentation masks. This paper introduces a prompt-based dataset construction approach that provides approximate lake locations using point, box, and mask prompts. We also propose a two-stage prompt enhancement framework, LEPrompter, with prompt-based and prompt-free stages during training. The prompt-based stage employs a prompt encoder to extract prior information, integrating prompt tokens and image embedding through self- and cross-attention in the prompt decoder. Prompts are deactivated to ensure independence during inference, enabling automated lake extraction without introducing additional parameters and GFlops. Extensive experiments showcase performance improvements of our proposed approach compared to the previous state- of-the-art method. The source code is available at https://github.com/BastianChen/LEPrompter. Xuechao Zou, Yu Zhang 0165, Junliang Xing, Pin Tao |
ICME | 2 |
| 2024 | A Parallel Attention Network For Cattle Face RecognitionabstractCattle face recognition holds paramount significance in domains such as animal husbandry and behavioral research. Despite significant progress in confined environments, applying these accomplishments in wild settings remains challenging. Thus, we create the first large-scale cattle face recognition dataset, ICRWE, for wild environments. It encompasses 483 cattle and 9,816 high-resolution image samples. Each sample undergoes annotation for face features, light conditions, and face orientation. Furthermore, we introduce a novel parallel attention network, PANet. Comprising several cascaded Transformer modules, each module incorporates two parallel Position Attention Modules (PAM) and Feature Mapping Modules (FMM). PAM focuses on local and global features at each image position through parallel channel attention, and FMM captures intricate feature patterns through non-linear mappings. Experimental results indicate that PANet achieves a recognition accuracy of 88.03% on the ICRWE dataset, establishing itself as the current state-of-the-art approach. The source code is available at https://github.com/1jy-0124/PANet Jiayu Li 0008, Xuechao Zou, Junliang Xing, Pin Tao |
ICME | 2 |
| 2024 | GOP: A Group Object Perception Framework for Optical Remote Sensing
Lei Jin 0003, Xuechao Zou, Jian Zhao 0006, Junliang Xing |
PRCV (12) | 3 |
| 2024 | DiffCR: A Fast Conditional Diffusion Framework for Cloud Removal From Optical Satellite ImagesabstractOptical satellite images are a critical data source; however, cloud cover often compromises their quality, hindering image applications and analysis. Consequently, effectively removing clouds from optical satellite images has emerged as a prominent research direction. Recent advances in deep learning-based cloud removal methods have been significant, but image generation quality still needs improvement. Diffusion models have demonstrated remarkable success in diverse image-generation tasks, showcasing their potential in addressing this challenge. This paper presents a novel framework called DiffCR, which leverages conditional guided diffusion with deep convolutional networks for high-performance cloud removal for optical satellite imagery. Specifically, we introduce a decoupled encoder for conditional image feature extraction, providing a robust color representation to ensure the close similarity of appearance information between the conditional input and the synthesized output. Moreover, we propose a novel and efficient time and condition fusion block within the cloud removal model to accurately simulate the correspondence between the appearance in the conditional image and the target image at a low computational cost. Extensive experimental evaluations on three commonly used benchmark datasets demonstrate that DiffCR consistently achieves state-of-the-art performance on all metrics, with parameter and computational complexities amounting to only 5.1% and 5.4%, respectively, of those previous best methods. The source code, pre-trained models, and all the experimental results will be publicly available at https://github.com/XavierJiezou/DiffCR upon the paper’s acceptance of this work. Xuechao Zou, Kai Li 0023, Junliang Xing, Yu Zhang 0165, Lei Jin 0003, Pin Tao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | PMAA: A Progressive Multi-Scale Attention Autoencoder Model for High-Performance Cloud Removal from Multi-Temporal Satellite ImageryabstractSatellite imagery analysis plays a pivotal role in remote sensing; however, information loss due to cloud cover significantly impedes its application. Although existing deep cloud removal models have achieved notable outcomes, they scarcely consider contextual information. This study introduces a high-performance cloud removal architecture, termed Progressive Multi-scale Attention Autoencoder (PMAA), which concurrently harnesses global and local information to construct robust contextual dependencies using a novel Multi-scale Attention Module (MAM) and a novel Local Interaction Module (LIM). PMAA establishes long-range dependencies of multi-scale features using MAM and modulates the reconstruction of fine-grained details utilizing LIM, enabling simultaneous representation of fine- and coarse-grained features at the same level. With the help of diverse and multi-scale features, PMAA consistently outperforms the previous state-of-the-art model CTGAN on two benchmark datasets. Moreover, PMAA boasts considerable efficiency advantages, with only 0.5% and 14.6% of the parameters and computational complexity of CTGAN, respectively. These comprehensive results underscore PMAA’s potential as a lightweight cloud removal network suitable for deployment on edge devices to accomplish large-scale cloud removal tasks. Our source code and pre-trained models are available at https://github.com/XavierJiezou/PMAA. Xuechao Zou, Kai Li 0023, Junliang Xing, Pin Tao, Yachao Cui |
ECAI | 1 |
| 2023 | PlantDet: A Benchmark for Plant Detection in the Three-Rivers-Source Region
Huanhuan Li 0008, Yuan Zhang 0026, Xuechao Zou, Jiangcai Zhaba, Guomei Li, Lamao Yongga |
ICANN (9) | 3 |