EDBT 2026 Demo / reviewers in the wild / expert
Yang Zou 0004
dblp:03/7447-4
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0007-3004-7650ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark DatasetabstractPansharpening under thin cloudy conditions is a practically significant yet rarely addressed task, challenged by simultaneous spatial resolution degradation and cloud-induced spectral distortions. Existing methods often address cloud removal and pansharpening sequentially, leading to cumulative errors and suboptimal performance due to the lack of joint degradation modeling. To address these challenges, we propose a Unified Pansharpening Model with Thin Cloud Removal (Pan-TCR), an end-to-end framework that integrates physical priors. Motivated by theoretical analysis in the frequency domain, we design a frequency-decoupled restoration (FDR) block that disentangles the restoration of multispectral image (MSI) features into amplitude and phase components, each guided by complementary degradation-robust prompts: the near-infrared (NIR) band amplitude for cloud-resilient restoration, and the panchromatic (PAN) phase for high-resolution structural enhancement. To ensure coherence between the two components, we further introduce an interactive inter-frequency consistency (IFC) module, enabling cross-modal refinement that enforces consistency and robustness across frequency cues. Furthermore, we introduce the first real-world thin-cloud contaminated pansharpening dataset (PanTCR-GF2), comprising paired clean and cloudy PAN-MSI images, to enable robust benchmarking under realistic conditions. Extensive experiments on real-world and synthetic datasets demonstrate the superiority and robustness of Pan-TCR, establishing a new benchmark for pansharpening under realistic atmospheric degradations. Songcheng Du, Yang Zou 0004, Ying Li 0017, Changjing Shang, Qiang Shen 0001 |
AAAI | 2 |
| 2026 | HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-ResolutionabstractInfrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fail to restore turbulence-induced distortions. Directly cascading turbulence mitigation (TM) algorithms with VSR methods leads to error propagation and accumulation due to the decoupled modeling of degradation between turbulence and resolution. We introduce HATIR, a Heat-Aware Diffusion for Turbulent InfraRed Video Super-Resolution, which injects heat-aware deformation priors into the diffusion sampling path to jointly model the inverse process of turbulent degradation and structural detail loss. Specifically, HATIR constructs a Phasor-Guided Flow Estimator, rooted in the physical principle that thermally active regions exhibit consistent phasor responses over time, enabling reliable turbulence-aware flow to guide the reverse diffusion process. To ensure the fidelity of structural recovery under nonuniform distortions, a Turbulence-Aware Decoder is proposed to selectively suppress unstable temporal cues and enhance edge-aware feature aggregation via turbulence gating and structure-aware attention. We built FLIR-IVSR, the first dataset for turbulent infrared VSR, comprising paired LR-HR sequences from a FLIR T1050sc camera (1024 X 768) spanning 640 diverse scenes with varying camera and object motion conditions. This encourages future research in infrared VSR. Yang Zou 0004, Xingyue Zhu, Kaiqi Han, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
AAAI | 1 |
| 2026 | Unsupervised Hyperspectral Image Super-Resolution via Self-Supervised Modality Decoupling
Songcheng Du, Yang Zou 0004, Xingyuan Li 0005, Ying Li 0017, Changjing Shang, Qiang Shen 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | Contourlet Refinement Gate Framework for Thermal Spectrum Distribution Regularized Infrared Image Super-Resolution
Yang Zou 0004, Zhixin Chen, Xingyuan Li 0005, Long Ma 0002, Jinyuan Liu 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | Adversarially robust Fourier-aware multimodal medical image fusion for LSCI
Zhidong Jiao, Yuchun He, Xingyue Zhu, Weipeng Sun, Xingyuan Li 0005, Yang Zou 0004 |
Neurocomputing | 7 |
| 2026 | Illumination Refinement via Textual Cues: A Prompt-Driven Approach for Low-Light NeRF EnhancementabstractIn the realm of 3D scene modeling and rendering, the emergence of Neural Radiance Fields (NeRF) represents a significant leap forward. However, NeRF’s rendering performance suffers significantly when rendering images under low-light conditions. Existing approaches are optimized by enhancing low-light input images and combining NeRF models, but still fail to address the issues of multiview consistency and image quality. To address these challenges, our research introduces a textual constraint-prompted enhancement method that facilitates low-light image brightening and new view synthesis in an unsupervised manner. Specifically, we devise a semantic calibration strategy that employs positive and negative prompts to motivate and penalize the network towards attributes associated with high-quality images and exploits the capability of visual language models in semantic parsing to align the generated images with textual descriptors to improve image generation quality. In addition, to address the multiview consistency problem, we propose a two-layer optimization strategy, where the semantic cue optimization in the upper layer and the new view generation in the lower layer interact with each other to achieve a balance between luminance consistency and structural integrity by combining these improved images with text-driven semantic features. Comprehensive tests on two datasets with different resolutions, LOM and LLFF, show that our approach outperforms existing methods by significantly improving the brightness and clarity of low-light images to state-of-the-art while preserving the natural appearance and details. Xinrui Ju, Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | AxisPose: Model-Free Matching-Free Single-Shot 6D Object Pose Estimation via Axis GenerationabstractObject pose estimation is a fundamental task in computer vision and plays an important role in various applications such as robotics, augmented reality, and autonomous manipulation. Existing studies often demand complex inputs or depend on correspondence-based matching between 2D image features and 3D object representations. While effective, these methods rely strongly on explicit appearance matching, often requiring multi-view inputs, depth sensors, or CAD models, which limits their scalability and robustness. Building on top of the pioneering generative studies, we propose AxisPose, a model-free, matching-free, and single-view 6D pose estimation framework that departs from conventional correspondence-based paradigms. Unlike existing methods, AxisPose directly infers a pose representation by learning a latent distribution of object orientation axes through a diffusion model. Specifically, AxisPose introduces an Axis Generation Module (AGM) that progressively denoises tri-axial orientation fields guided by geometric consistency constraints, and a Triaxial Back-projection Module (TBM) to recover the final 6D pose from the generated orientation axes without relying on explicit 2D- 2D/3D correspondences. AxisPose achieves strong cross-instance generalization, enabling a single model to handle multiple object categories without retraining. Extensive experiments on LINEMOD and YCB-Video datasets demonstrate that AxisPose improves the Average Distance Deviation score from 0.733 to 0.814 over the strong baseline NOPE, using only a single RGB input. The code is available at https://github.com/pubyLu/AxisPose/tree/main. Yang Zou 0004, Zhaoshuai Qi, Weipeng Sun, Xingyuan Li 0005, Jiaqi Yang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-ResolutionabstractInfrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challenge imaging quality and subsequent visual tasks. Hence, infrared image super-resolution (IISR) has been developed to address this challenge. While recent developments in diffusion models have greatly advanced this field, current methods to solve it either ignore the unique modal characteristics of infrared imaging or overlook the machine perception requirements. To bridge these gaps, we propose DifIISR, an infrared image super-resolution diffusion model optimized for visual quality and perceptual performance. Our approach achieves task-based guidance for diffusion by injecting gradients derived from visual and perceptual priors into the noise during the reverse process. Specifically, we introduce an infrared thermal spectrum distribution regulation to preserve visual fidelity, ensuring that the reconstructed infrared images closely align with high-resolution images by matching their frequency components. Subsequently, we incorporate various visual foundational models as the perceptual guidance for downstream visual tasks, infusing generalizable perceptual features beneficial for detection and segmentation. As a result, our approach gains superior visual results while attaining State-Of-The-Art downstream task performance. Code is available at https://github.com/zirui0625/DifIISR Xingyuan Li 0005, Yang Zou 0004, Zhixin Chen, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001 |
CVPR | 3 |
| 2025 | DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image FusionabstractInfrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resulting in fused images that offer only marginal gains in task performance and fail to provide constructive feedback for optimizing the fusion process. To overcome these limitations, we propose a Discriminative Cross-Dimension Evolutionary Learning Framework, termed DCEvo, which simultaneously enhances visual quality and perception accuracy. Leveraging the robust search capabilities of Evolutionary Learning, our approach formulates the optimization of dual tasks as a multi-objective problem by employing an Evolutionary Algorithm (EA) to dynamically balance loss function parameters. Inspired by visual neuroscience, we integrate a Discriminative Enhancer (DE) within both the encoder and decoder, enabling the effective learning of complementary features from different modalities. Additionally, our Cross-Dimensional Embedding (CDE) block facilitates mutual enhancement between high-dimensional task features and low-dimensional fusion features, ensuring a cohesive and efficient feature integration process. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving an average improvement of 9.32% in visual quality while also enhancing subsequent high-level tasks. The code is available at https://github.com/Beate-Suy-Zhang/DCEvo. Jinyuan Liu 0001, Qingyun Mei, Xingyuan Li 0005, Yang Zou 0004, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
CVPR | 5 |
| 2025 | Toward a Training-Free Plug-and-Play Refinement Framework for Infrared and Visible Image Registration and FusionabstractInfrared and Visible Image Fusion (IVIF) under unregistered conditions has been of great interest in various visual tasks under challenging environments. While existing approaches often demonstrate promising results on specific benchmarks, they tend to exhibit performance drops in unseen scenarios and incur high computational overhead when retrained on new datasets. To address these challenges, we propose TRACE, a Training-free Reinforcement-based Alignment method for Cross-modality Enhancement, which incorporates Evaluator, a rewarding network, into an evaluation-driven Reinforcement Learning (RL) framework, enabling efficient and plug-and-play refinement of any existing registration approach. Specifically, TRACE constructs the Evaluator network to assess the alignment quality of the given registration model, generating confidence scores and adjustment masks via spatial and channel attention. Leveraging these cues as RL rewards, TRACE iteratively refines the registration network to mitigate misalignments until the accumulated improvement is satisfied. Due to its training-free and plug-and-play nature, TRACE notably enhances fusion results across diverse and unseen scenarios. TRACE achieves impressive improvements in different methods across diverse datasets with minimal computational cost. The project page is available at https://github.com/pubyLu/TRACE. Yang Zou 0004, Xingyuan Li 0005, Xingyue Zhu, Kaiqi Han, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001 |
ACM Multimedia | 2 |
| 2025 | Modeling detail feature connections for infrared image enhancement
Ziang Meng, Kaiqi Han, Yuchun He, Yuhan He, Xingyuan Li 0005, Yang Zou 0004 |
Neurocomputing | 6 |
| 2024 | Enhancing Neural Radiance Fields with Adaptive Multi-Exposure Fusion: A Bilevel Optimization Approach for Novel View SynthesisabstractNeural Radiance Fields (NeRF) have made significant strides in the modeling and rendering of 3D scenes. However, due to the complexity of luminance information, existing NeRF methods often struggle to produce satisfactory renderings when dealing with high and low exposure images. To address this issue, we propose an innovative approach capable of effectively modeling and rendering images under multiple exposure conditions. Our method adaptively learns the characteristics of images under different exposure conditions through an unsupervised evaluator-simulator structure for HDR (High Dynamic Range) fusion. This approach enhances NeRF's comprehension and handling of light variations, leading to the generation of images with appropriate brightness. Simultaneously, we present a bilevel optimization method tailored for novel view synthesis, aiming to harmonize the luminance information of input images while preserving their structural and content consistency. This approach facilitates the concurrent optimization of multi-exposure correction and novel view synthesis, in an unsupervised manner. Through comprehensive experiments conducted on the LOM and LOL datasets, our approach surpasses existing methods, markedly enhancing the task of novel view synthesis for multi-exposure environments and attaining state-of-the-art results. The source code can be found at https://github.com/Archer-204/AME-NeRF. Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
AAAI | 1 |
| 2024 | Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution
Xingyuan Li 0005, Jinyuan Liu 0001, Zhixin Chen, Yang Zou 0004, Long Ma 0002, Xin Fan 0001, Risheng Liu |
ECCV (3) | 4 |
| 2024 | Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders images across a spectrum of exposure conditions. Our approach utilizes an unsupervised classifier-generator structure for HDR fusion, significantly enhancing NeRF’s ability to comprehend and adjust to light variations, leading to the generation of images with appropriate brightness. Extensive evaluations on the LOM[1] and LOL[2] datasets underscore our method’s edge. Our approach significantly improves the task of novel view synthesis for multi-exposure images, attaining state-of-the-art results. Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Tiantian Yan, Jinyuan Liu 0001 |
ICASSP | 1 |