VLDB 2026 Research / reviewers in the wild / expert
Xiaoyuan Yang 0003
dblp:95/11478-3
· DBLP profile ↗
23ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-7367-0115ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Inconsistency Guidance for Super-resolution Video Quality AssessmentabstractAs super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches. Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
AAAI | 2 |
| 2025 | IcGAN4ColSAR: A Novel Multispectral Conditional Generative Adversarial Network Approach for SAR Image ColorizationabstractSAR colorization aims to enrich gray-scale SAR images with color while ensuring the preservation of original radiometric and spatial details. However, researchers often limit themselves to using only the red, green, and blue bands of a multispectral image as the source of color information, coupled with a single-polarization channel from the SAR image. This approach neglects the intrinsic characteristics of remote sensing data and thus fails to fully leverage available information. To overcome this limitation, this research attempts to explore inclusion of all available bands from multispectral images along with dual-polarization channels from SAR imagery in the colorization process. Furthermore, we present a new colorization method called improved conditional generative adversarial network for SAR colorization (IcGAN4ColSAR). This method tries to include the spectral angle mapper index within its loss function. Sufficient experiments show that our explorations in the number of data channels and the loss function are helpful in improving the colorization performance of the SAR image. Kangqing Shen, Gemine Vivone, Simone Lolli, Michael Schmitt 0003, Xiaoyuan Yang 0003, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | EAMNet: Efficient Adaptive Mamba Network for Infrared Small-Target DetectionabstractInfrared small target detection (ISTD) is essential for various fields. Recent approaches based on existing network structures including convolutional neural networks (CNNs), Transformers, and diffusion models, still face challenges in balancing accuracy and efficiency. To address this problem, this paper proposes an Efficient Adaptive Mamba Network (EAMNet) based on the advanced Mamba structure, which effectively models long-range dependencies while maintaining linear complexity, enabling EAMNet to achieve superior detection performance while significantly improving efficiency. First, a Mamba-based UNet architecture is introduced, which processes separated features in parallel, making it highly efficient with a low parameter count and computational cost. To better adapt the Mamba-based framework to the unique characteristics of infrared images, such as low contrast and small target sizes, we propose an adaptive filter module (AFM) that applies adaptive filtering by predicting filter parameters through an additional designed sub-network, enhancing the boundaries and visibility of infrared targets. To further enhance model performance and ensure efficient feature fusion, we propose a shared adaptive spatial attention module (SASAM), which enables a more compact and efficient feature representation in generating spatial attention maps, while minimizing additional computational overhead. Extensive experiments on public benchmarks demonstrate the effectiveness of the proposed EAMNet in both improving accuracy and efficiency compared to existing state-of-the-art methods. Besides, ablation experiments verify the effectiveness of each module. The code is available at https://github.com/jiangjin1246/EAMNet. Jin Jiang 0003, Shengcai Liao, Xiaoyuan Yang 0003, Kangqing Shen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Cross-Modal Contrastive Pansharpening via Uncertainty GuidanceabstractDeep learning (DL)-based pansharpening has been widely applied in high-resolution imaging. Yet, artifacts related to generalization and oversmoothing have continuously been the challenge, primarily due to the mismatch between the simulation dataset and the unseen real-world scenarios. Current approaches address these through unsupervised frameworks or generative models, while modal inconsistency is not fully considered, leading to suboptimal performance. In this article, we propose a contrastive cross-modal framework via uncertainty guidance (UGCC), which comprises three key modules: a contrast feature enhancement module (CFEM), a cross-modal compensation module (CMCM), and an uncertainty guidance module (UGM). First, to enhance generalization and reduce overfitting, CFEM is introduced. Robust contrast features are augmented and learned sparsely in latent space, where sample distributions are refined, and redundant information is filtered from highly similar sample pairs for enhanced training stability. Furthermore, CMCM mitigates modal inconsistency effectively by domain transfer and collaborative attention, achieving efficient modal separation and interaction. Finally, to adaptively balance the performance of CMCM and CFEM based on prediction confidence, a hybrid loss function is designed, where UGM adjusts the weights through quantifying statistical-versus-structural uncertainties. Extensive experiments on Quickbird, Gaofen-2, WorldView-2, and WorldView-3 demonstrate that the performance of the proposed method surpasses or matches the state of the arts. Furthermore, ablation studies validate the effectiveness of each component. The code is now available at:https://github.com/meimeizeng/UGCF. Haoying Zeng, Xiaoyuan Yang 0003, Kangqing Shen, Jin Jiang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality AssessmentabstractMany super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods. Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
IEEE Trans. Image Process. | 2 |
| 2024 | Deep Bi-directional Attention Network for Image Super-Resolution Quality AssessmentabstractThere has emerged a growing interest in exploring efficient quality assessment algorithms for image super-resolution (SR). However, employing deep learning techniques, especially dual-branch algorithms, to automatically evaluate the visual quality of SR images remains challenging. Existing SR image quality assessment (IQA) metrics based on two-stream networks lack interactions between branches. To address this, we propose a novel full-reference IQA (FR-IQA) method for SR images. Specifically, producing SR images and evaluating how close the SR images are to the corresponding HR references are separate processes. Based on this consideration, we construct a deep Bidirectional Attention Network (BiAtten-Net) that dynamically deepens visual attention to distortions in both processes, which aligns well with the human visual system (HVS). Experiments on public SR quality databases demonstrate the superiority of our proposed BiAtten-Net over state-of-the-art quality assessment methods. In addition, the visualization results and ablation study show the effectiveness of bi-directional attention. Xiaoyuan Yang 0003, Jun Fu 0007, Guanghui Yue 0001, Wei Zhou 0021 |
ICME | 2 |
| 2024 | A benchmarking protocol for SAR colorization: From regression to deep learning approaches
Kangqing Shen, Gemine Vivone, Xiaoyuan Yang 0003, Simone Lolli, Michael Schmitt 0003 |
Neural Networks | 3 |
| 2023 | IIANet: Information Interactivity Attention Network with adversarial learning for infrared small object detection
Jin Jiang 0003, Xiaoyuan Yang 0003 |
Comput. Vis. Image Underst. | 2 |
| 2023 | WeightFace: weight adaptive scaling loss for face recognition
Huwei Ren, Xiaoyuan Yang 0003 |
Multim. Tools Appl. | 2 |
| 2023 | Adversarial feature hybrid framework for steganography with shifted window local loss
Zhengze Li, Xiaoyuan Yang 0003, Kangqing Shen, Fazhen Jiang, Jin Jiang 0003, Huwei Ren |
Neural Networks | 2 |
| 2023 | DuaFace: Data uncertainty in angular based loss for face recognition
Fazhen Jiang, Xiaoyuan Yang 0003, Huwei Ren, Zhengze Li, Kangqing Shen, Jin Jiang 0003 |
Pattern Recognit. Lett. | 2 |
| 2022 | Dual branch parallel steganographic framework based on multi-scale distillation in framelet domain
Zhengze Li, Xiaoyuan Yang 0003, Kangqing Shen, Fazhen Jiang, Jin Jiang 0003, Huwei Ren |
Neurocomputing | 2 |
| 2022 | MultiBSP: multi-branch and multi-scale perception object tracking framework based on siamese CNN
Jin Jiang 0003, Xiaoyuan Yang 0003, Zhengze Li, Kangqing Shen, Fazhen Jiang, Huwei Ren |
Neural Comput. Appl. | 2 |
| 2021 | PSGU: Parametric self-circulation gating unit for deep neural networks
Zhengze Li, Xiaoyuan Yang 0003, Kangqing Shen, Fazhen Jiang, Jin Jiang 0003, Huwei Ren |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Real-time least-squares ensemble visual trackingabstractIn this study, the authors present a novel ensemble tracking system by formulating the tracking task in terms of a linear regression which is a least‐squares problem. A set of weak classifiers are trained using least squares which are solved efficiently using the Moore–Penrose inverse. Then, these weak classifiers are combined into a strong classifier using bagging. The strong classifier is used to recognise the target and locate its position, which is obtained efficiently in the Fourier domain. For obtaining a good ensemble, a novel sampling strategy is proposed to train accurate and diverse weak classifiers. By exploiting historical targets to monitor the training process, pose change and occlusion are well‐handled. The proposed method is extensively evaluated using a variety of evaluation protocols on the recent standard datasets including OTB50, OTB100 and VOT2016. Experimental results show that the proposed methodology performs favourably against state‐of‐the‐art methods in terms of efficiency, accuracy and robustness. Ridong Zhu, Xiaoyuan Yang 0003, Jingkai Wang 0001, Zhengze Li |
IET Image Process. | 2 |
| 2020 | Information encryption communication system based on the adversarial networks Foundation
Zhengze Li, Xiaoyuan Yang 0003, Kangqing Shen, Ridong Zhu, Jin Jiang 0003 |
Neurocomputing | 2 |
| 2019 | Random Walks for Pansharpening in Complex Tight Framelet DomainabstractIn this paper, a new random walk (RW) pansharpening method on the basis of the complex framelet domain is proposed. In the process of fusion, the hidden Markov tree model is first established based on the statistical properties of complex high-pass framelet coefficients. On this basis, a novel RW fusion algorithm is presented. Then, the probabilities of complex framelet coefficients being allotted original images are solved by the linear system of equations. Based on these probabilities, the spatial details of the panchromatic image are selectively injected into the multispectral (MS) image to get a space-enhanced MS image. In the end, the GeoGye-1, WorldView-3, and WorldView-2 remote sensing image data sets are used to evaluate the performance of the presented method quantitatively and qualitatively. The results of the experiment show that our method outperforms some state-of-the-art approaches. It can improve the spatial resolution of the MS image while keeping the spectral information. Jingkai Wang 0001, Xiaoyuan Yang 0003, Ridong Zhu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Random Walks for Synthetic Aperture Radar Image Fusion in Framelet DomainabstractA new framelet-based random walks (RWs) method is presented for synthetic aperture radar (SAR) image fusion, including SAR-visible images, SAR-infrared images, and Multi-band SAR images. In this method, we build a novel RWs model based on the statistical characteristics of framelet coefficients to fuse the high-frequency and low-frequency coefficients. This model converts the fusion problem to estimate the probability of each framelet coefficient being assigned each input image. Experimental results show that the proposed approach improves the contrast while preserves the edges simultaneously, and outperforms many traditional and state-of-the-art fusion techniques in both qualitative and quantitative analysis. Xiaoyuan Yang 0003, Jingkai Wang 0001, Ridong Zhu |
IEEE Trans. Image Process. | 1 |
| 2014 | Random Error Modeling and Analysis of Airborne Lidar SystemsabstractWhen airborne lidar is used to produce digital elevation models, the random error of airborne lidar systems is often the limiting factor. To improve the overall 3-D expected accuracy of airborne lidar systems, we develop random error models of airborne lidar systems. First, we present a footprint geolocation equation for the airborne lidar that describes an ideal system, followed by a modified footprint geolocation equation that contains systematic and random errors. To better constrain these uncertainties, we describe the mathematical characteristics of the random error in detail by modeling it as a stochastic process. The probabilistic random error models are developed by analyzing the differences between the errorless equation and the equation with errors. This paper also addresses the issue of recovering the errors and presents an error recovery model. Finally, based on a linear scanner and a new coordinate system, we simulate the geometric processes of point positioning of airborne lidar systems and study the impacts of the random measurement errors on the 3-D coordinate accuracy of airborne lidar systems. Junchen Dan, Xiaoyuan Yang 0003, Yan Shi 0012, Yuhua Guo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Translation Invariant Directional Framelet Transform Combined With Gabor Filters for Image DenoisingabstractThis paper is devoted to the study of a directional lifting transform for wavelet frames. A nonsubsampled lifting structure is developed to maintain the translation invariance as it is an important property in image denoising. Then, the directionality of the lifting-based tight frame is explicitly discussed, followed by a specific translation invariant directional framelet transform (TIDFT). The TIDFT has two framelets ψ1, ψ2 with vanishing moments of order two and one respectively, which are able to detect singularities in a given direction set. It provides an efficient and sparse representation for images containing rich textures along with properties of fast implementation and perfect reconstruction. In addition, an adaptive block-wise orientation estimation method based on Gabor filters is presented instead of the conventional minimization of residuals. Furthermore, the TIDFT is utilized to exploit the capability of image denoising, incorporating the MAP estimator for multivariate exponential distribution. Consequently, the TIDFT is able to eliminate the noise effectively while preserving the textures simultaneously. Experimental results show that the TIDFT outperforms some other frame-based denoising methods, such as contourlet and shearlet, and is competitive to the state-of-the-art denoising approaches. Yan Shi 0012, Xiaoyuan Yang 0003, Yuhua Guo |
IEEE Trans. Image Process. | 2 |
| 2011 | The Lifting Factorization and Construction of Wavelet Bi-Frames With Arbitrary Generators and ScalingabstractIn this paper, we present the lifting factorization and construction of wavelet bi-frames with arbitrary generators and scaling. We show that an arbitrary polyphase matrix of a wavelet bi-frame can be factorized into a series of lifting steps. Based on the proposed factorization, we present a general construction of bi-frames. Especially, we give an explicit formula to construct the bi-frames of two scaling and two generators with symmetry and one vanishing moment. This paper does not involve many theories of wavelet frames in mathematics, but focuses on the algebraic issues related to Laurent polynomials, which are efficient expressions in redundant filter banks associated with wavelet frames. As an extension of the classical two-channel filter bank, the redundant filter bank is more complicated but also more flexible. Furthermore, we present an algorithm to increase the number of vanishing moments to arbitrary order by lifting, which is iterated and is straightforward in implementation. Yan Shi 0012, Xiaoyuan Yang 0003 |
IEEE Trans. Image Process. | 2 |
| 2010 | The Lifting Scheme for Wavelet Bi-Frames: Theory, Structure, and AlgorithmabstractIn this paper, we present the lifting scheme of wavelet bi-frames along with theory analysis, structure, and algorithm. We show how any wavelet bi-frame can be decomposed into a finite sequence of simple filtering steps. This decomposition corresponds to a factorization of a polyphase matrix of a wavelet bi-frame. Based on this concept, we present a new idea for constructing wavelet bi-frames. For the construction of symmetric bi-frames, we use generalized Bernstein basis functions, which enable us to design symmetric prediction and update filters. The construction allows more efficient implementation and provides tools for custom design of wavelet bi-frames. By combining the different designed filters for the prediction and update steps, we can devise practically unlimited forms of wavelet bi-frames. Moreover, we present an algorithm of increasing the number of vanishing moments of bi-framelets to arbitrary order via the presented lifting scheme, which adopts an iterative algorithm and ensures the shortest lifting scheme. Several construction examples are given to illustrate the results. Xiaoyuan Yang 0003, Yan Shi 0012, Liuhe Chen, Zongfeng Quan |
IEEE Trans. Image Process. | 1 |
| 2008 | Embedded zerotree wavelets coding based on adaptive fuzzy clustering for image compression
Xiaoyuan Yang 0003, Bo Li 0006 |
Image Vis. Comput. | 1 |