EDBT 2026 Demo / reviewers in the wild / expert
Zipeng Qi
dblp:286/8776
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-8571-6942ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bayesian Neural Networks for One-to-Many Mapping in Image EnhancementabstractIn image enhancement tasks, such as low-light and underwater image enhancement, a degraded image can correspond to multiple plausible target images due to dynamic photography conditions. This naturally results in a one-to-many mapping problem. To address this, we propose a Bayesian Enhancement Model (BEM) that incorporates Bayesian Neural Networks (BNNs) to capture data uncertainty and produce diverse outputs. To enable fast inference, we introduce a BNN-DNN framework: a BNN is first employed to model the one-to-many mapping in a low-dimensional space, followed by a Deterministic Neural Network (DNN) that refines fine-grained image details. Extensive experiments on multiple low-light and underwater image enhancement benchmarks demonstrate the effectiveness of our method. Guoxi Huang, Ruirui Lin, Zipeng Qi, David Bull 0001, Nantheera Anantrasirichai |
AAAI | 4 |
| 2025 | Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-ReflectionabstractDiffusion models, the most popular generative paradigm so far, can inject conditional information into the generation path to guide the latent towards desired directions. However, existing text-to-image diffusion models often fail to maintain high image quality and high prompt-image alignment for those challenging prompts. To mitigate this issue and enhance existing pretrained diffusion models, we mainly made three contributions in this paper. First, we propose **diffusion self-reflection** that alternately performs denoising and inversion and demonstrate that such diffusion self-reflection can leverage the guidance gap between denoising and inversion to capture prompt-related semantic information with theoretical and empirical evidence. Second, motivated by theoretical analysis, we derive Zigzag Diffusion Sampling (Z-Sampling), a novel self-reflection-based diffusion sampling method that leverages the guidance gap between denosing and inversion to accumulate semantic information step by step along the sampling path, leading to improved sampling results. Moreover, as a plug-and-play method, Z-Sampling can be generally applied to various diffusion models (e.g., accelerated ones and Transformer-based ones) with very limited coding and computational costs. Third, our extensive experiments demonstrate that Z-Sampling can generally and significantly enhance generation quality across various benchmark datasets, diffusion models, and performance evaluation metrics. For example, DreamShaper with Z-Sampling can self-improve with the HPSv2 winning rate up to **94%** over the original results. Moreover, Z-Sampling can further enhance existing diffusion models combined with other orthogonal methods, including Diffusion-DPO. The code is publicly available at
[github.com/xie-lab-ml/Zigzag-Diffusion-Sampling](https://github.com/xie-lab-ml/Zigzag-Diffusion-Sampling). Lichen Bai, Shitong Shao, Zikai Zhou, Zipeng Qi, Zhiqiang Xu 0003, Haoyi Xiong, Zeke Xie |
ICLR | 4 |
| 2025 | Efficient Semantic Splatting for Remote Sensing Multiview SegmentationabstractRemote sensing multi-view image segmentation is essential for achieving accurate and consistent stereoscopic perception of target scenes. This task involves processing RGB images from multiple viewpoints to generate high-accuracy, view-consistent semantic segmentation across all views. Traditional training-based methods struggle with maintaining cross-view consistency, while optimization-driven approaches using implicit neural networks improve view consistency but suffer from slow parameter optimization and inference. To overcome these limitations, we propose a novel Gaussian Splatting-based semantic segmentation framework. Our method efficiently projects the color attributes and semantic features of 3D Gaussians onto the image plane, enabling the simultaneous generation of both RGB images and segmentation outputs. By leveraging explicit spatial structures and a splatting rendering strategy, our approach significantly enhances optimization efficiency and rendering speed. Additionally, we incorporate SAM2 to generate pseudo-labels for boundary regions, addressing the lack of supervision in sparsely labeled views (e.g., 3%). To further enforce cross-view consistency and feature coherence of 3D Gaussians, we introduce a two-level aggregation loss that operates at both the 2D feature map and 3D spatial levels. Extensive experiments across nine datasets demonstrate the superiority of our method, achieving competitive segmentation quality with limited supervisory views. Notably, our approach reduces rendering (inference) times by 90%, while improving the average mIoU by up to 3.5%. Zipeng Qi, Hao Chen 0045, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Layered Rendering Diffusion Model for Controllable Zero-Shot Image Synthesis
Zipeng Qi, Guoxi Huang |
ECCV (66) | 1 |
| 2024 | Learning to Detect Cloud and Snow in Remote Sensing Images from Noisy LabelsabstractDetecting clouds and snow in remote sensing images is an essential preprocessing task for remote sensing imagery. Previous works draw inspiration from semantic segmentation models in computer vision, with most research focusing on improving model architectures to enhance detection performance. However, unlike natural images, the complexity of scenes and the diversity of cloud types in remote sensing images result in many inaccurate labels in cloud and snow detection datasets, introducing unnecessary noises into the training and testing processes. By constructing a new dataset and proposing a novel training strategy with the curriculum learning paradigm, we guide the model in reducing overfitting to noisy labels. Additionally, we design a more appropriate model performance evaluation method, that alleviates the performance assessment bias caused by noisy labels. By conducting experiments on models with UNet and Segformer, we have validated the effectiveness of our proposed method. This paper is the first to consider the impact of label noise on the detection of clouds and snow in remote sensing images. Hao Chen 0045, Wenyuan Li 0002, Keyan Chen 0001, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001 |
IGARSS | 5 |
| 2024 | Pixel-Level Change Detection Pseudo-Label Learning For Remote Sensing Change CaptioningabstractThe existing Remote Sensing Image Change Captioning (RSICC) methods perform well in simple scenes but exhibit poorer performance in complex scenes. This limitation is primarily attributed to the model’s constrained visual ability to distinguish and locate changes. Acknowledging the inherent correlation between change detection (CD) and RSICC tasks, we believe pixel-level CD is significant for describing the differences between images through language. Regrettably, the current RSICC dataset lacks readily available pixel-level CD labels. To address this deficiency, we leverage a model trained on existing CD datasets to derive CD pseudo-labels. We propose an innovative network with an auxiliary CD branch, supervised by pseudo-labels. Furthermore, a semantic fusion augment (SFA) module is proposed to fuse the feature information extracted by the CD branch, thereby facilitating the nuanced description of changes. Experiments demonstrate that our method achieves state-of-the-art performance and validate that learning pixel-level CD pseudo-labels significantly contributes to change captioning. Keyan Chen 0001, Zipeng Qi, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001 |
IGARSS | 3 |
| 2024 | Multi-View Remote Sensing Image Segmentation with Sam PriorsabstractMulti-view segmentation in Remote Sensing (RS) seeks to segment images from diverse perspectives within a scene. Recent methods leverage 3D information extracted from an Implicit Neural Field (INF), bolstering result consistency across multiple views while using limited accounts of labels (even within 3-5 labels) to streamline labor. Nonetheless, achieving superior performance within the constraints of limited-view labels remains challenging due to inadequate scene-wide supervision and insufficient semantic features within the INF. To address these. we propose to inject the prior of the visual foundation model-Segment Anything(SAM), to the INF to obtain better results under the limited number of training data. Specifically, we contrast SAM features between testing and training views to derive pseudo labels for each testing view, augmenting scene-wide labeling information. Subsequently, we introduce SAM features via a transformer into the INF of the scene, supplementing the semantic information. The experimental results demonstrate that our method outperforms the mainstream method, confirming the efficacy of SAM as a supplement to the INF for this task. Zipeng Qi, Hao Chen 0045, Yongchang Wu, Zhengxia Zou, Zhenwei Shi 0001 |
IGARSS | 1 |
| 2024 | Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and AnalysisabstractMonitoring changes in the Earth’s surface is crucial for understanding natural processes and human impacts, necessitating precise and comprehensive interpretation methodologies. Remote sensing (RS) satellite imagery offers a unique perspective for monitoring these changes, leading to the emergence of RS image change interpretation (RSICI) as a significant research focus. Current RSICI technology encompasses change detection and change captioning, each with its limitations in providing comprehensive interpretation. To address this, we propose an interactive Change-Agent, which can follow user instructions to achieve comprehensive change interpretation and insightful analysis, such as change detection and change captioning, change object counting, and change cause analysis. The Change-Agent integrates a multilevel change interpretation (MCI) model as the eyes and a large language model (LLM) as the brain. The MCI model contains two branches of pixel-level change detection and semantic-level change captioning, in which the BI-temporal iterative interaction (BI3) layer is proposed to enhance the model’s discriminative feature representation capabilities. To support the training of the MCI model, we build the LEVIR-MCI dataset with a large number of change masks and captions of changes. Experiments demonstrate the state-of-the-art (SOTA) performance of the MCI model in achieving both change detection and change description simultaneously and highlight the promising application value of our Change-Agent in facilitating comprehensive interpretation of surface changes, which opens up a new avenue for intelligent RS applications. To facilitate future research, we will make our dataset and codebase publicly available athttps://github.com/Chen-Yang-Liu/Change-Agent. Keyan Chen 0001, Haotian Zhang 0010, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Progressive Scale-Aware Network for Remote Sensing Image Change CaptioningabstractRemote sensing (RS) images contain numerous objects of different scales, which poses significant challenges for the RS image change captioning (RSICC) task to identify visual changes of interest in complex scenes and describe them via language. However, current methods still have some weaknesses in sufficiently extracting and utilizing multi-scale information. In this paper, we propose a progressive scale-aware network (PSNet) to address the problem. PSNet is a pure Transformer-based model. To sufficiently extract multi-scale visual features, multiple progressive difference perception (PDP) layers are stacked to progressively exploit the differencing features of bitemporal features. To sufficiently utilize the extracted multi-scale features for captioning, we propose a scale-aware reinforcement (SR) module and combine it with the Transformer decoding layer to progressively utilize the features from different PDP layers. Experiments show that the PDP layer and SR module are effective and our PSNet outperforms previous methods. Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001 |
IGARSS | 3 |
| 2023 | A Decoupling Paradigm With Prompt Learning for Remote Sensing Image Change CaptioningabstractRemote sensing image change captioning (RSICC) is a novel task that aims to describe the differences between bi-temporal images by natural language. Previous methods ignore a significant specificity of the task: the difficulty of RSICC is different for unchanged and changed image pairs. They process the unchanged and changed image pairs in a coupled way, which usually causes confusion for change captioning. In this paper, we decouple the task into two issues to ease it: whether and what changes have occurred. An image-level classifier performs binary classification to address the first issue. A feature-level encoder contributes to extracting discriminative features to help the caption generation module address the second issue. Besides, for caption generation, we utilize prompt learning to introduce pre-trained large language models (LLMs) into the RSICC task. A multi-prompt learning strategy is proposed to generate a set of unified prompts and a class-specific prompt conditioned on the image-level classifier’s results. The strategy can prompt a pre-trained LLM to know whether changes exist and generate captions. Finally, the multiple prompts and the visual features of the feature-level encoder are fed into a frozen LLM for language generation. Compared with previous methods, our method can leverage the powerful abilities of the pre-trained LLM in language to generate plausible captions, which is free of training. Extensive experiments show that our method is effective and achieves state-of-the-art performance. Besides, an additional experiment demonstrates that our decoupling paradigm is more promising than the previous coupled paradigm for the RSICC task. We will make our codebase publicly available to facilitate future research at https://github.com/Chen-Yang-Liu/PromptCC. Rui Zhao 0019, Jianqi Chen, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Implicit Ray Transformers for Multiview Remote Sensing Image SegmentationabstractThe mainstream CNN-based remote sensing (RS) image semantic segmentation approaches typically rely on massively labeled training data. Such a paradigm struggles with the problem of RS multi-view scene segmentation with limited labeled views due to the lack of consideration of 3D information within the scene. In this paper, we propose “Implicit Ray-Transformer (IRT)” based on Implicit Neural Representation (INR) for RS scene semantic segmentation with sparse labels (5% of the images being labeled). We explore a new way of introducing the multi-view 3D structure priors to the task for accurate and view-consistent semantic segmentation. The proposed method includes a two-stage learning process. In the first stage, we optimize a neural field to encode the color and 3D structure of the remote sensing scene based on multi-view images. In the second stage, we design a Ray Transformer to leverage the relations between the neural field 3D features and 2D texture features for learning better semantic representations. Different from previous methods that only consider 3D priors or 2D features, we incorporate additional 2D texture information and 3D priors by broadcasting CNN features to different point features along the sampled ray. To verify the effectiveness of the proposed method, we construct a challenging dataset containing six synthetic sub-datasets collected from the Carla platform and three real sub-datasets from Google Maps. Experiments show that the proposed method outperforms the CNN-based methods and the state-of-the-art INR-based segmentation methods in quantitative and qualitative metrics. The ablation study shows that under a limited number of fully annotated images, the combination of the 3D structure priors and 2D texture can significantly improve the performance and effectively complete missing semantic information in novel views. Experiments also demonstrate the proposed method could yield geometry-consistent segmentation results against illumination changes and viewpoint changes. Our data and code will be public. Zipeng Qi, Hao Chen 0045, Zhenwei Shi 0001, Zhengxia Zou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Remote-Sensing Image Segmentation Based on Implicit 3-D Scene RepresentationabstractRemote sensing image segmentation, as a challenging but fundamental task, has drawn increasing attention in the remote sensing field. Recent advances in deep learning have greatly boosted research on this task. However, the existing deep learning-based segmentation methods heavily rely on a large amount of pixel-wise labeled training data, and the labeling process is time-consuming and labor-intensive. In this paper, we focus on the scenario that leverages the 3D structure of multi-view images and a limited number of annotations to generate accurate novel view segmentation. Under this scenario, we propose a novel method for remote sensing image segmentation based on implicit 3D scene representation, which generates arbitrary-view segmentation output from limited segmentation annotations. The proposed method employs a two-stage training strategy. In the first stage, we optimize the implicit neural representations of a 3D scene and encode their multi-view images into a neural radiance field. In the second stage, we transform the scene color attribute into semantic labels and propose a ray-convolution network to aggregate local 3D consistency cues across different locations. We also design a color-radiance network to help our method generalize to unseen views. Experiments on both synthetic and real-world data suggest that our method significantly outperforms deep convolutional networks (CNN)-based methods and other view synthesis-based methods. We also show that the proposed method can be applied as a novel data augmentation approach that benefits CNN-based segmentation methods. Zipeng Qi, Zhengxia Zou, Hao Chen 0045, Zhenwei Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Remote Sensing Image Change Detection With TransformersabstractModern change detection (CD) has achieved remarkable success by the powerful discriminative ability of deep convolutions. However, high-resolution remote sensing CD remains challenging due to the complexity of objects in the scene. Objects with the same semantic concept may show distinct spectral characteristics at different times and spatial locations. Most recent CD pipelines using pure convolutions are still struggling to relate long-range concepts in space-time. Nonlocal self-attention approaches show promising performance via modeling dense relationships among pixels, yet are computationally inefficient. Here, we propose a bitemporal image transformer (BIT) to efficiently and effectively model contexts within the spatial-temporal domain. Our intuition is that the high-level concepts of the change of interest can be represented by a few visual words, that is, semantic tokens. To achieve this, we express the bitemporal image into a few tokens and use a transformer encoder to model contexts in the compact token-based space-time. The learned context-rich tokens are then fed back to the pixel-space for refining the original features via a transformer decoder. We incorporate BIT in a deep feature differencing-based CD framework. Extensive experiments on three CD datasets demonstrate the effectiveness and efficiency of the proposed method. Notably, our BIT-based model significantly outperforms the purely convolutional baseline using only three times lower computational costs and model parameters. Based on a naive backbone (ResNet18) without sophisticated structures (e.g., feature pyramid network (FPN) and UNet), our model surpasses several state-of-the-art CD methods, including better than four recent attention-based methods in terms of efficiency and accuracy. Our code is available athttps://github.com/justchenhao/BIT_CD. Hao Chen 0045, Zipeng Qi, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |