VLDB 2026 Research / reviewers in the wild / expert
Xiaowan Hu
dblp:299/4760
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-7364-4489ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Continuous Spatiotemporal Implicit Neural Fields for Unsupervised Video DenoisingabstractVideo denoising is fundamental to low-level vision and real-world imaging, yet existing self-supervised methods remain fragile under severe noise and complex motion. Most approaches still rely on spatially and temporally discrete grid-based representations: blind-spot networks enforce J-invariance by masking center pixels with a limited receptive field, while recurrent models build temporal dependencies on discretized frame sequences and noise-sensitive optical flow, leading to error accumulation and motion artifacts. We address this model bottleneck by reformulating self-supervised video denoising as learning a continuous spatiotemporal implicit field. Building on coordinate-based implicit neural representations, we propose a unified video denoising model with a spatiotemporal implicit neural field (SINF). In the spatial domain, a blind-spot implicit spatial field maps coordinates directly to pixel-level representations, enabling globally informed texture recovery beyond receptive-field limits. In the temporal domain, an implicit temporal embedding with periodic activations encodes motion continuously over time, while a time-aware spatial graph module refines cross-frame alignment. Together, SINF remodels discretized video signals into a continuous spatiotemporal intensity field, enabling more robust pixel-wise associations than coarse optical flow. Extensive experiments on synthetic and real noisy video benchmarks demonstrate that our SINF achieves state-of-the-art performance on synthetic and real noisy video benchmarks. Xiaowan Hu, Henan Liu, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | A Novel Visible-Infrared Image Compression Framework for High-Value Target ProtectionabstractThe joint compression of visible-infrared images is crucial for military and surveillance applications. The challenge lies in the protection of high-value targets (HVT) while maintaining high compression efficiency. This paper proposes a novel dual-stream compression framework that effectively addresses this challenge. In our framework, the sensitive HVT infrared signatures are concealed within the visible image stream, while residual infrared image is encoded separately. This dual-stream compression framework introduces three key innovations. 1) HVT protection: The HVT information is physically isolated and hidden within public visible images through a dedicated concealment stream; 2) Key-conditioned reconstruction: A novel decoding mechanism enables active camouflage by replacing HVTs with plausible background content when unauthorized access is detected; 3) Unified optimization: The framework integrates compression efficiency and HVT protection within an endto- end trainable network. Extensive experiments demonstrate that our approach achieves state-of-the-art compression performance while providing superior HVT protection, significantly outperforming traditional encrypt-then-compress methods. The code and weights are open-source athttps://github.com/eecoder-dyf/rgbir-compress. Yufan Deng, Xin Deng 0002, Shengxi Li, Xiaowan Hu, Mai Xu |
IEEE Signal Process. Lett. | 4 |
| 2026 | Dual-Domain Visual Prompt Learning for Multi-Modal Medical Image Saliency PredictionabstractMedical image saliency prediction plays a pivotal role in emulating clinician visual attention to prioritize diagnostically critical regions. Current methods remain constrained by their spatial-domain dependency and limited cross-modality generalizability, neglecting frequency-domain patterns critical for subtle pathology detection while suffering from over-specialization in specific imaging modalities. Therefore, we propose a dual-domain visual prompt network (DVPNet) that integrates cross-modality generalization with spectral pattern awareness. On the one hand, DVPNet establishes a dataset prompt branch that dynamically modulates spatial feature encoding through modality-specific priors, allowing adaptive interpretation of heterogeneous medical imaging domains. On the other hand, a spatial-frequency hybrid prompt module employs learnable wavelet filters to decompose images into multi-scale spectral components, preserving low-frequency anatomical context while enhancing discriminative high-frequency biomarkers that are typically obscured in previous pixel-level analysis. By seamlessly integrating these complementary representations, DVPNet optimally synthesizes spatial and spectral evidence, enabling robust generalization across diverse medical imaging modalities while sustaining computational efficiency. Extensive experimental results on two distinct datasets demonstrate that the proposed method outperforms state-of-the-art approaches, showing superior saliency prediction performance and enhanced generalizability across medical contexts. Mai Xu, Xiaowan Hu, Lai Jiang 0004 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Spatiotemporal Blind-Spot Network with Calibrated Flow Alignment for Self-Supervised Video DenoisingabstractSelf-supervised video denoising aims to remove noise from videos without relying on ground truth data, leveraging the video itself to recover clean frames. Existing methods often rely on simplistic feature stacking or apply optical flow without thorough analysis. This results in suboptimal utilization of both inter-frame and intra-frame information, and it also neglects the potential of optical flow alignment under self-supervised conditions, leading to biased and insufficient denoising outcomes. To this end, we first explore the practicality of optical flow in the self-supervised setting and introduce a SpatioTemporal Blind-spot Network (STBN) for global frame feature utilization. In the temporal domain, we utilize bidirectional blind-spot feature propagation through the proposed blind-spot alignment block to ensure accurate temporal alignment and effectively capture long-range dependencies. In the spatial domain, we introduce the spatial receptive field expansion module, which enhances the receptive field and improves global perception capabilities. Additionally, to reduce the sensitivity of optical flow estimation to noise, we propose an unsupervised optical flow distillation mechanism that refines fine-grained inter-frame interactions during optical flow alignment. Our method demonstrates superior performance across both synthetic and real-world video denoising datasets. Xiaowan Hu, Huaqiu Li, Haoqian Wang |
AAAI | 3 |
| 2025 | Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single Image DenoisingabstractMany studies have concentrated on constructing supervised models utilizing paired datasets for image denoising, which proves to be expensive and time-consuming. Current self-supervised and unsupervised approaches typically rely on blind-spot networks or sub-image pairs sampling, resulting in pixel information loss and destruction of detailed structural information, thereby significantly constraining the efficacy of such methods. In this paper, we introduce Prompt-SID, a prompt-learning-based single image denoising framework that emphasizes the preservation of structural details. This approach is trained in a self-supervised manner using downsampled image pairs. It captures original-scale image information through structural encoding and integrates this prompt into the denoiser. To achieve this, we propose a structural representation generation model based on the latent diffusion process and design a structural attention module within the transformer-based denoiser architecture to decode the prompt. Additionally, we introduce a scale replay training mechanism, which effectively mitigates the scale gap from images of different resolutions. We conduct comprehensive experiments on synthetic, real-world, and fluorescence imaging datasets, showcasing the remarkable effectiveness of Prompt-SID. Huaqiu Li, Xiaowan Hu, Haoqian Wang |
AAAI | 3 |
| 2025 | Interpretable Unsupervised Joint Denoising and Enhancement for Real-World low-light ScenariosabstractReal-world low-light images often suffer from complex degradations such as local overexposure, low brightness, noise, and uneven illumination. Supervised methods tend to overfit to specific scenarios, while unsupervised methods, though better at generalization, struggle to model these degradations due to the lack of reference images. To address this issue, we propose an interpretable, zero-reference joint denoising and low-light enhancement framework tailored for real-world scenarios. Our method derives a training strategy based on paired sub-images with varying illumination and noise levels, grounded in physical imaging principles and retinex theory. Additionally, we leverage the Discrete Cosine Transform (DCT) to perform frequency domain decomposition in the sRGB space, and introduce an implicit-guided hybrid representation strategy that effectively separates intricate compounded degradations. In the backbone network design, we develop retinal decomposition network guided by implicit degradation representation mechanisms. Extensive experiments demonstrate the superiority of our method. Code will be available at https://github.com/huaqlili/unsupervised-light-enhance-ICLR2025. Huaqiu Li, Xiaowan Hu, Haoqian Wang |
ICLR | 2 |
| 2025 | Measuring and Controlling the Spectral Bias for Self-Supervised Image DenoisingabstractCurrent self-supervised denoising methods for paired noisy images typically involve mapping one noisy image through the network to the other noisy image. However, after measuring the spectral bias of such methods using our proposed Image Pair Frequency-Band Similarity, it suffers from two practical limitations. Firstly, the high-frequency structural details in images are not preserved well enough. Secondly, during the process of fitting high frequencies, the network learns high-frequency noise from the mapped noisy images. To address these challenges, we introduce a Spectral Controlling network (SCNet) to optimize self-supervised denoising of paired noisy images. First, we propose a selection strategy to choose frequency band components for noisy images, to accelerate the convergence speed of training. Next, we present a parameter optimization method that restricts the learning ability of convolutional kernels to high-frequency noise using the Lipschitz constant, without changing the network structure. Finally, we introduce the Spectral Separation and low-rank Reconstruction module (SSR module), which separates noise and high-frequency details through frequency domain separation and low-rank space reconstruction, to retain the high-frequency structural details of images. Experiments performed on synthetic and real-world datasets verify the effectiveness of SCNet. The code will be released soon. Huaqiu Li, Xiaowan Hu, Haoqian Wang |
ICME | 3 |
| 2024 | Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
Xiaowan Hu, Yan Li 0043, Minquan Wang, Haoqian Wang, Quan Chen 0006, Han Li 0005, Peng Jiang 0002 |
ACM Multimedia | 1 |
| 2022 | Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image ReconstructionabstractHyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactions is beneficial for HSI reconstruction. However, existing CNN-based methods show limitations in capturing spectral-wise similarity and long-range dependencies. Besides, the HSI information is modulated by a coded aperture (physical mask) in CASSI. Nonetheless, current algorithms have not fully explored the guidance effect of the mask for HSI restoration. In this paper, we propose a novel framework, Mask-guided Spectral-wise Transformer (MST), for HSI reconstruction. Specifically, we present a Spectral-wise Multi-head Self-Attention (S-MSA) that treats each spectral feature as a token and calculates self-attention along the spectral dimension. In addition, we customize a Mask-guided Mechanism (MM) that directs S- MSA to pay attention to spatial regions with high-fidelity spectral representations. Extensive experiments show that our MST significantly outperforms state-of-the-art (SOTA) methods on simulation and real HSI datasets while requiring dramatically cheaper computational and memory costs. https://github.com/caiyuanhao1998/MST/ Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
CVPR | 3 |
| 2022 | HDNet: High-resolution Dual-domain Learning for Spectral Compressive ImagingabstractThe rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance against complexity, losing fine-grained high-resolution (HR) features. Secondly, even if the optimization focusing on spatial-spectral domain learning (SDL) converges to the ideal solution, there is still a significant visual difference between the reconstructed HSI and the truth. So we propose a high-resolution dual-domain learning network (HDNet) for HSI reconstruction. On the one hand, the proposed HR spatial-spectral attention module with its efficient feature fusion provides continuous and fine pixel-level features. On the other hand, frequency domain learning (FDL) is introduced for HSI reconstruction to narrow the frequency domain discrepancy. Dynamic FDL supervision forces the model to reconstruct fine-grained frequencies and compensate for excessive smoothing and distortion caused by pixel-level losses. The HR pixel-level attention and frequency-level refinement in our HDNet mutually promote HSI perceptual quality. Extensive quantitative and qualitative experiments show that our method achieves SOTA performance on simulated and real HSI datasets. https://github.com/Huxiaowan/HDNet Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
CVPR | 1 |
| 2022 | Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
ECCV (17) | 3 |
| 2022 | Flow-Guided Sparse Transformer for Video DeblurringabstractExploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Transformer (FGST), for video deblurring. In FGST, we customize a self-attention module, Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA). For each $query$ element on the blurry reference frame, FGSW-MSA enjoys the guidance of the estimated optical flow to globally sample spatially sparse yet highly related $key$ elements corresponding to the same scene patch in neighboring frames. Besides, we present a Recurrent Embedding (RE) mechanism to transfer information from past frames and strengthen long-range temporal dependencies. Comprehensive experiments demonstrate that our proposed FGST outperforms state-of-the-art (SOTA) methods on both DVD and GOPRO datasets and yields visually pleasant results in real video deblurring. https://github.com/linjing7/VR-Baseline Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Youliang Yan, Xueyi Zou, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
ICML | 3 |
| 2022 | Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video RestorationabstractHow to properly model the inter-frame relation within the video sequence is an important but unsolved challenge for video restoration (VR). In this work, we propose an unsupervised flow-aligned sequence-to-sequence model (S2SVR) to address this problem. On the one hand, the sequence-to-sequence model, which has proven capable of sequence modeling in the field of natural language processing, is explored for the first time in VR. Optimized serialization modeling shows potential in capturing long-range dependencies among frames. On the other hand, we equip the sequence-to-sequence model with an unsupervised optical flow estimator to maximize its potential. The flow estimator is trained with our proposed unsupervised distillation loss, which can alleviate the data discrepancy and inaccurate degraded optical flow issues of previous flow-based methods. With reliable optical flow, we can establish accurate correspondence among multiple frames, narrowing the domain difference between 1D language and 2D misaligned frames and improving the potential of the sequence-to-sequence model. S2SVR shows superior performance in multiple VR tasks, including video deblurring, video super-resolution, and compressed video quality enhancement. https://github.com/linjing7/VR-Baseline Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Youliang Yan, Xueyi Zou, Yulun Zhang 0001, Luc Van Gool |
ICML | 2 |
| 2022 | Wide Weighted Attention Multi-Scale Network for Accurate MR Image Super-ResolutionabstractHigh-quality magnetic resonance (MR) images afford more detailed information for reliable diagnoses and quantitative image analyses. Given low-resolution (LR) images, the deep convolutional neural network (CNN) has shown its promising ability for image super-resolution (SR). The LR MR images usually share some visual characteristics: structural textures of different sizes, edges with high correlation, and less informative background. However, multi-scale structural features are informative for image reconstruction, while the background is more smooth. Most previous CNN-based SR methods use a single receptive field and equally treat the spatial pixels (including the background). It neglects to sense the entire space and get diversified features from the input, which is critical for high-quality MR image SR. We propose a wide weighted attention multi-scale network ($\text{W}^{2}$AMSN) for accurate MR image SR to address these problems. On the one hand, the features of varying sizes can be extracted by the wide multi-scale branches. On the other hand, we design a non-reduction attention mechanism to recalibrate feature responses adaptively. Such attention preserves continuous cross-channel interaction and focuses on more informative regions. Meanwhile, the learnable weighted factors fuse extracted features selectively. The encapsulated wide weighted attention multi-scale block ($\text{W}^{2}$AMSB) is integrated through a recurrent framework and global attention mechanism. Extensive experiments and diversified ablation studies show the effectiveness of our proposed$\text{W}^{2}$AMSN, which surpasses state-of-the-art methods on most popular MR image SR benchmarks quantitatively and qualitatively. And our method still offers superior accuracy and adaptability on real MR images. Haoqian Wang, Xiaowan Hu, Xiaole Zhao, Yulun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Pseudo 3D Auto-Correlation Network for Real Image DenoisingabstractThe extraction of auto-correlation in images has shown great potential in deep learning networks, such as the self-attention mechanism in the channel domain and the self-similarity mechanism in the spatial domain. However, the realization of the above mechanisms mostly requires complicated module stacking and a large number of convolution calculations, which inevitably increases model complexity and memory cost. Therefore, we propose a pseudo 3D auto-correlation network (P3AN) to explore a more efficient way of capturing contextual information in image de-noising. On the one hand, P3AN uses fast 1D convolution instead of dense connections to realize criss-cross interaction, which requires less computational resources. On the other hand, the operation does not change the feature size and makes it easy to expand. It means that only a simple adaptive fusion is needed to obtain contextual information that includes both the channel domain and the spatial domain. Our method built a pseudo 3D auto-correlation attention block through 1D convolutions and a lightweight 2D structure for more discriminative features. Extensive experiments have been conducted on three synthetic and four real noisy datasets. According to quantitative metrics and visual quality evaluation, the P3AN shows great superiority and surpasses state-of-the-art image denoising methods. Xiaowan Hu, Ruijun Ma 0001, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001, Haoqian Wang |
CVPR | 1 |
| 2021 | Margin Loss Based On Adaptive Metric For Image RecognitionabstractCross-entropy (CE) loss is one of the most commonly used supervision in image recognition. However, the features trained by CE loss are not discriminative enough. There are many methods using the angular-softmax CE loss to learn angularly discriminative features. But these methods still use manually designed metric, which cannot deal with complicated distribution of high-dimensional features, and there are no appropriate inter-class restraint. Therefore, we propose a novel method to restrain the distribution of features in high-dimensional. We construct trainable centers, and design adaptive metric to express the distance between features. Specifically, we design AdaMetricLoss which can manipulate inter-class distance and intra-class distance simultaneously. We evaluate the effectiveness of AdaMetricLoss on Cifar10 and Cifar100 datasets, and our method shows preferable classification performance. We also visualize the discriminative distribution of features on MNIST, proves the proposed method is more suitable for image recognition. Lei Song 0003, Xiaowan Hu, Haoqian Wang |
ICIP | 3 |
| 2021 | Pyramid Orthogonal Attention Network based on Dual Self-Similarity for Accurate Mr Image Super-ResolutionabstractFor magnetic resonance (MR) images sharing visual characteristics, the internal structure repetitions of different scales are considerable image-specific priors. Following the traditional algorithms, we try to combine external dataset-driven learning with the internal self-similarity for MR image super-resolution (SR). We propose a pyramid orthogonal attention network (POAN) based on dual self-similarity. On the one hand, by combining the point-similarity and the pyramid-similarity, sufficient spatial autocorrelation is explored to alleviate less training data limitation. On the other hand, the non-reduction channel attention mechanism maximizes inter-channel dependence. It increases the probability of the high-frequency region (e.g., structural textures and edges) being activated while suppresses low-frequency regions (e.g., background) adaptively. Out proposed POAN reconstructs the MR image under the guidance of pyramid orthogonal attention. Extensive experiments demonstrate that our method obtains the best results compared with state-of-the-art MR image SR methods quantitatively and visually. Xiaowan Hu, Haoqian Wang, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001 |
ICME | 1 |
| 2021 | Multi-Scale Selective Feedback Network with Dual Loss for Real Image DenoisingabstractThe feedback mechanism in the human visual system extracts high-level semantics from noisy scenes. It then guides low-level noise removal, which has not been fully explored in image denoising networks based on deep learning. The commonly used fully-supervised network optimizes parameters through paired training data. However, unpaired images without noise-free labels are ubiquitous in the real world. Therefore, we proposed a multi-scale selective feedback network (MSFN) with the dual loss. We allow shallow layers to access valuable contextual information from the following deep layers selectively between two adjacent time steps. Iterative refinement mechanism can remove complex noise from coarse to fine. The dual regression is designed to reconstruct noisy images to establish closed-loop supervision that is training-friendly for unpaired data. We use the dual loss to optimize the primary clean-to-noisy task and the dual noisy-to-clean task simultaneously. Extensive experiments prove that our method achieves state-of-the-art results and shows better adaptability on real-world images than the existing methods. Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Yulun Zhang 0001 |
IJCAI | 1 |
| 2021 | Learning to Generate Realistic Noisy Images via Pixel-level Noise-aware Adversarial TrainingabstractExisting deep learning real denoising methods require a large amount of noisy-clean image pairs for supervision. Nonetheless, capturing a real noisy-clean dataset is an unacceptable expensive and cumbersome procedure. To alleviate this problem, this work investigates how to generate realistic noisy images. Firstly, we formulate a simple yet reasonable noise model that treats each real noisy pixel as a random variable. This model splits the noisy image generation problem into two sub-problems: image domain alignment and noise domain alignment. Subsequently, we propose a novel framework, namely Pixel-level Noise-aware Generative Adversarial Network (PNGAN). PNGAN employs a pre-trained real denoiser to map the fake and real noisy images into a nearly noise-free solution space to perform image domain alignment. Simultaneously, PNGAN establishes a pixel-level adversarial training to conduct noise domain alignment. Additionally, for better noise fitting, we present an efficient architecture Simple Multi-scale Network (SMNet) as the generator. Qualitative validation shows that noise generated by PNGAN is highly similar to real noise in terms of intensity and distribution. Quantitative experiments demonstrate that a series of denoisers trained with the generated noisy images achieve state-of-the-art (SOTA) results on four real denoising benchmarks. Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Yulun Zhang 0001, Hanspeter Pfister, Donglai Wei 0001 |
NeurIPS | 2 |