VLDB 2026 Research / reviewers in the wild / expert
Jiang Qin
dblp:180/6605
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion ModelsabstractRecent advances in text-to-image (T2I) diffusion models have enabled remarkable control over various attributes, yet precise color specification remains a fundamental challenge. Existing approaches, such as ColorPeel, rely on model personalization, requiring additional optimization and limiting flexibility in specifying arbitrary colors. In this work, we introduce ColorWave, a novel training-free approach that achieves exact RGB-level color control in diffusion models without fine-tuning. By systematically analyzing the cross-attention mechanisms within IP-Adapter, we uncover an implicit binding between textual color descriptors and reference image features. Leveraging this insight, our method rewires these bindings to enforce precise color attribution while preserving the generative capabilities of pretrained models. Our approach maintains generation quality and diversity, outperforming prior methods in accuracy and applicability across diverse object categories. Through extensive evaluations, we demonstrate that ColorWave establishes a new paradigm for structured, color-consistent diffusion-based image synthesis. Héctor Laria Mantecon, Alexandra Gomez-Villa, Jiang Qin, Muhammad Atif Butt, Bogdan Raducanu, Javier Vazquez-Corral, Joost van de Weijer 0001, Kai Wang 0060 |
WACV | 3 |
| 2025 | Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationabstractRecent advances in Text-to-Image (T2I) diffusion models have transformed image generation, enabling significant progress in stylized generation using only a few style reference images. However, current diffusion-based methods struggle with \textit{fine-grained} style customization due to challenges in controlling multiple style attributes, such as color and texture. This paper introduces the first tuning-free approach to achieve free-lunch color-texture disentanglement in stylized T2I generation, addressing the need for independently controlled style elements for the Disentangled Stylized Image Generation (DisIG) problem. Our approach leverages the \textit{Image-Prompt Additivity} property in the CLIP image embedding space to develop techniques for separating and extracting Color-Texture Embeddings (CTE) from individual color and texture reference images. To ensure that the color palette of the generated image aligns closely with the color reference, we apply a whitening and coloring transformation to enhance color consistency. Additionally, to prevent texture loss due to the signal-leak bias inherent in diffusion training, we introduce a noise term that preserves textural fidelity during the Regularized Whitening and Coloring Transformation (RegWCT). Through these methods, our Style Attributes Disentanglement approach (SADis) delivers a more precise and customizable solution for stylized image generation. Experiments on images from the WikiArt and StyleDrop datasets demonstrate that, both qualitatively and quantitatively, SADis surpasses state-of-the-art stylization methods in the DisIG task. Jiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 0002, Yaxing Wang, Kai Wang 0060, Joost van de Weijer 0001 |
NeurIPS | 1 |
| 2025 | Efficient End-to-End Diffusion Model for One-Step SAR-to-Optical TranslationabstractThe undesirable distortions of synthetic aperture radar (SAR) images pose a challenge to intuitive SAR interpretation. SAR-to-optical (S2O) image translation provides a feasible solution for easier interpretation of SAR and supports multisensor analysis. Currently, diffusion-based S2O models are emerging and have achieved remarkable performance in terms of perceptual metrics and fidelity. However, the numerous iterative sampling steps and slow inference speed of these diffusion models (DMs) limit their potential for practical applications. In this letter, an efficient end-to-end diffusion model (E3Diff) is developed for real-time one-step S2O translation. E3Diff not only samples as fast as generative adversarial network (GAN) models, but also retains the powerful image synthesis performance of DMs to achieve high-quality S2O translation in an end-to-end manner. To be specific, SAR spatial priors are first incorporated to provide enriched conditional clues and achieve more precise control from the feature level to synthesize optical images. Then, E3Diff is accelerated by a hybrid refinement loss, which effectively integrates the advantages of both GAN and diffusion components to achieve efficient one-step sampling. Experiments show that E3Diff achieves real-time inference speed (0.17 s per image on an A6000 GPU) and demonstrates significant image-quality improvements (35% and 27% improvement in Frechet inception distance (FID) on the UNICORN and SEN12 dataset, respectively) compared to existing state-of-the-art (SOTA) diffusion S2O methods. This advancement of E3Diff highlights its potential to enhance SAR interpretation and cross-modal applications. The code is available athttps://github.com/DeepSARRS/E3Diff. Jiang Qin, Bin Zou 0001, Lamei Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Cross-Resolution SAR Target Detection Using Structural Hierarchy Adaptation and Reliable Adjacency AlignmentabstractIn recent years, continuous improvements in SAR resolution have significantly benefited applications such as urban monitoring and target detection. However, these improvements in resolution have also led to increased discrepancies in scattering characteristics, posing challenges to the generalization ability of target detection models. While domain adaptation technologies provide a potential solution, the inevitable discrepancies caused by resolution differences often result in blind feature adaptation and unreliable semantic propagation, ultimately degrading the domain adaptation performance. To address these challenges, this paper proposes a novel SAR target detection method, termedCR-Net, which incorporates structure priors and evidential learning theory into the detection model, enabling reliable domain adaptation for cross-resolution detection. To be specific,CR-Netintegrates Structure-induced Hierarchical Feature Adaptation (SHFA) and Reliable Structural Adjacency Alignment (RSAA). TheSHFAmodule is designed to establish structural correlations between targets and achieve structure-aware feature adaptation, thereby enhancing the interpretability of the adaptation process. Afterwards, theRSAAmodule is proposed to enhance reliable semantic alignment, by leveraging the secure adjacency set to transfer valuable discriminative knowledge from the source domain to the target domain. This further improves the discriminability of the detection model in the target domain. Based on experimental results from different-resolution datasets, the proposedCR-Netsignificantly enhances cross-resolution adaptation by preserving intra-domain structures and improving discriminability. It achieves state-of-the-art (SOTA) performance in cross-resolution SAR target detection. Jiang Qin, Bin Zou 0001, Lamei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Hierarchical Attribute Scattering Center Extraction Method for SAR ImageabstractThe attribute scattering center (ASC) is a common target characteristic in synthetic aperture radar (SAR) image. In addition to the optical features, the physical parameters extracted by ASC provide the specific target structure information related to the SAR system. However, the existing methods for ASC extraction usually extract all parameters at the same time, which cause many errors in the extracted results. In this paper, a hierarchical ASC extraction method is proposed, in which the parameters of each scattering center are extracted hierarchically. In the experimental section, a SAR image simulation algorithm is used to simulate SAR images of some typical structures and the effect of ASC extraction is analyzed. Besides, a set of measured data is used for validation as well. Experimental results show that the hierarchical ASC extraction method can extract the scattering centers more accurately. Lamei Zhang, Bin Zou 0001, Jiang Qin |
IGARSS | 4 |
| 2024 | Multi-Aspect Feature Enhancement Network for Aircraft Detection in High-Resolution SAR ImagesabstractAircraft detection using Synthetic Aperture Radar (SAR) images plays a crucial role in transportation and military applications. Nevertheless, the unique imaging characteristics of SAR often render aircraft targets as discrete points, and their complex geometric structures vary under different imaging conditions. Moreover, the complex background, especially with strong-scattering elements like buildings, significantly complicates detection. To overcome these challenges, this paper introduces a multi-aspect feature enhancement network called MAFEN for aircraft detection in high-resolution SAR images. MAFEN combines a Multi-scale Feature Enhancement Module (MSFEM) with a Key Structure Enhancement Module (KSEM), thereby enhancing detection accuracy and efficiency. Experimental results on public datasets demonstrate significant improvements in detecting aircraft targets in complex scenes. Bin Zou 0001, Jiang Qin, Lamei Zhang |
IGARSS | 3 |
| 2024 | Scientific and Technological Problems of Modeling Jointly Functioning Artificial and Bioenergy Agro-Ecological SystemsabstractWe live in a time that is complex and difficult to predict, even for events shortly. Along with the uniform development of modern technologies, wars, and artificial and natural disasters arise sporadically, which contemporary society has not yet learned to cope with effectively. Various world-class programs and projects exist to overcome the consequences of such phenomena, but these efforts are not enough to prevent the conditions for cataclysms. There is a need to consider the construction and development of new technologies as a particular component of a complex equivalent in the form of a system of concepts: “nature - man - society - state.” Considering the above circumstances, solving the problem of building effective bioenergy agricultural production in these conditions is becoming a promising area of scientific research in the coming decades. The efforts of scientists in many countries aim to find optimal solutions to ensure the effective operation of agricultural technologies associated with the production and distribution of energy resources available in the state. The work analyzes some conditions and theoretical prerequisites for creating new generation models using a virtual ecosystem model at the pace of technological energy production and consumption processes in bioenergy ecosystems. The shortcomings of existing approaches and technologies for modeling complex heterogeneous ecosystems are discussed. New innovative technology elements for constructing models and modeling the functioning modes of bioenergy agricultural technologies using artificial intelligence methods are proposed. Viktor Gurieiev, Jiang Qin |
TENCON | 2 |
| 2024 | Semi-Supervised SAR Image Change Detection via Structure-Optimized Complex-Valued Graph Contrastive LearningabstractSignificant progress has been achieved by using the graph convolutional network (GCN) in image change detection. However, the limited quantity of labeled data and the inherent speckle noise adversely impact the generalization ability of the existing GCN-based methods in practical synthetic aperture radar (SAR) image applications. To address these challenges, we introduce the structure-optimized complex-valued graph contrastive learning network (SCGCLN) for semi-supervised SAR image change detection. Specifically, we explore how to learn effective feature representations from complex-valued SAR data with limited supervised information using the GCN architecture. We present a structure-optimized graph reconstruction strategy based on optimizing node features and edge structures. By combining efficient spectral clustering with graph reconnection, our method learns high-quality graph structures that enable the network to capture long-range dependencies, thereby mitigating the impact of speckle noise. Moreover, we construct a complex-valued graph contrastive learning (GCL) network to train a graph feature representation model from unlabeled SAR data. Subsequently, the pretrained model is fine-tuned for the downstream limited labeled SAR change detection task. The effectiveness of SCGCLN is validated through experimental results on three SAR image datasets. Bin Zou 0001, Lamei Zhang, Jiang Qin |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | CausalCD: A Causal Graph Contrastive Learning Framework for Self-Supervised SAR Image Change DetectionabstractIn recent years, self-supervised synthetic aperture radar (SAR) image change detection methods have achieved remarkable results, particularly in reducing dependence on expensive supervised signals. However, some critical issues remain open: how to design the optimal unsupervised feature representation model, and how to utilize noisy pseudo-labels for self-training to improve the change detection model? In this article, efforts are made to find a principled and fundamental solution to the above issues from the new perspective of causal reasoning, proposing a self-supervised SAR change detection framework named CausalCD. Specifically, we first construct a structural causal model (SCM) to formalize the self-supervised SAR image change detection process, and carry out a principled analysis to assess the influence of the unsupervised feature representation module and the pseudo-label-driven self-training module on the performance of change detection model. On this basis, we present a feature representation module that employs graph contrastive learning (GCL) and leverages the causal invariant mechanism to extract the optimal augmented representation, thus improving the model’s generalizability and soundness. In addition, we employ a causal intervention strategy to mitigate the negative impact of confusing bias from noisy pseudo-labels by blocking the backdoor path. CausalCD’s strengths lie in reducing the model’s dependence on labeled samples through causal GCL and obtaining optimal feature representation. Concurrently, it helps disentangle the noisy pseudo-label’s impact, thereby improving the performance of the self-supervised SAR change detection model. Finally, comprehensive experiments on three bitemporal SAR scenes demonstrate that CausalCD significantly outperforms several mainstream change detection models, confirming the effectiveness of CausalCD. Bin Zou 0001, Lamei Zhang, Jiang Qin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Conditional Diffusion Model With Spatial-Frequency Refinement for SAR-to-Optical Image TranslationabstractThe presence of speckles and geometric distortions poses a serious challenge to the visual interpretation of synthetic aperture radar (SAR) images. SAR-to-optical (S2O) image translation technology provides a feasible solution and has attracted increasing attention. Restricted by substantial gaps between optical and SAR images, current S2O translation methods unavoidably result in geometric distortions, target missing, and generating low-fidelity images, thereby limiting subsequent cross-modal applications. In this article, we propose an augmented conditional denoising diffusion probabilistic model with spatial-frequency refinement (SFDiff) for high-fidelity S2O image translation. SFDiff progressively narrows the gap between synthesized and real images in both spatial and frequency perspectives, showcasing notable performance in terms of quality and consistency. Specifically, to incorporate rich spatial content priors provided by SAR images, we design an SAR context prior extractor (SCPE) with denoising enhancement to extract multiscale conditional representations, thereby aiding SFDiff in capturing more descriptive cues for S2O translation. In addition, a spatial-frequency complementary learning (SFCL) module is designed to learn spatial semantics and simultaneously enhances informative frequency components and global dependencies. Furthermore, SFDiff is optimized using the joint spatial-frequency refinement loss, facilitating iterative refinement in both spatial and frequency domains to enhance content consistency and fidelity in the synthesized images. Based on the experimental findings from the UNICORN dataset and the SEN12 dataset, SFDiff maintains a high level of content and structural consistency, resulting in visually appealing translation results that surpass the state-of-the-art (SOTA) methods. In particular, SFDiff exhibits excellent performance in preserving small targets and details, which is crucial in cross-modal detection applications. Jiang Qin, Kai Wang 0060, Bin Zou 0001, Lamei Zhang, Joost van de Weijer 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Vehicle Detection Based on Semantic-Context Enhancement for High-Resolution SAR Images in Complex BackgroundabstractSmall-scale target detection (such as vehicles) in complex synthetic aperture radar (SAR) image scenes has always been a pain point for the advanced convolutional neural network (CNN)-based target detectors because of the downsampling operations and the local receptive field characteristics of CNNs. To tackle these limitations, a vehicle detector named SCEDet for the small-scale vehicles in SAR images is proposed to improve the detection performance in this letter. SCEDet mainly consists of two parts: subaperture semantic feature extraction and subaperture semantic-context enhancement (SCE) with SCE module. First, ResNet34 with subaperture decomposition is used to efficiently exploit the latent subaperture semantic features. Then, the SCE module is proposed to balance the multiscale semantic information as well as aggregate the global context information for vehicle detection with a small number of parameters and computation costs. The experimental results on the FARAD dataset (0.1 m$\times0.1$m, Ka-band) demonstrate that both the detection performance and the speed are much better than other detection methods under the same hardware conditions. Bin Zou 0001, Jiang Qin, Lamei Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |