VLDB 2026 Research / reviewers in the wild / expert
Xingyuan Li 0005
dblp:51/6800-5
· DBLP profile ↗
18ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0001-9081-817XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain Adaptation Guided Infrared and Visible Image FusionabstractInfrared and Visible Image Fusion (IVIF) integrates complementary information from distinct modalities to enhance image quality. However, the effectiveness declines under unseen conditions such as novel weather or scenes, due to domain shifts primarily from variations of data distribution in the visible modality, while the infrared modality remains relatively stable. To overcome domain shifts caused by the imbalance between modalities during image fusion, we propose a Domain Adaptation Guided Infrared and Visible Image Fusion method, termed DAFusion, leveraging a dual-rank domain adapter to enable fast adaptation to diverse adverse conditions during image fusion. Specifically, trainable low-rank and high-rank embedding spaces are respectively used to capture knowledge common across domains (domain-shared) and those unique to target domains (domain-specific). To leverage the dual-rank adapter more effectively, we develop a homeostatic knowledge allotment strategy to integrate the distinct types of knowledge dynamically based on the uncertainty value of target domains. Since domain adaptation typically optimizes for feature alignment across domains and emphasizes invariance rather than preserving specific cues critical for image fusion, while the fusion objective requires retaining discriminative and complementary features, a conflict between the two modules appears. To reconcile this, we further adopt a bi-level optimization framework that structurally decouples the two objectives, enabling the fusion module to steer the adaptation process while benefiting in return from domain-aligned representations. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving both an enhancement in fusion quality and an improvement on subsequent high-level tasks. Tianwei Guan, Haozhen Wei, Zecheng Xu, Zhiying Jiang, Jinyuan Liu 0001, Xingyuan Li 0005 |
AAAI | 8 |
| 2026 | HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-ResolutionabstractInfrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fail to restore turbulence-induced distortions. Directly cascading turbulence mitigation (TM) algorithms with VSR methods leads to error propagation and accumulation due to the decoupled modeling of degradation between turbulence and resolution. We introduce HATIR, a Heat-Aware Diffusion for Turbulent InfraRed Video Super-Resolution, which injects heat-aware deformation priors into the diffusion sampling path to jointly model the inverse process of turbulent degradation and structural detail loss. Specifically, HATIR constructs a Phasor-Guided Flow Estimator, rooted in the physical principle that thermally active regions exhibit consistent phasor responses over time, enabling reliable turbulence-aware flow to guide the reverse diffusion process. To ensure the fidelity of structural recovery under nonuniform distortions, a Turbulence-Aware Decoder is proposed to selectively suppress unstable temporal cues and enhance edge-aware feature aggregation via turbulence gating and structure-aware attention. We built FLIR-IVSR, the first dataset for turbulent infrared VSR, comprising paired LR-HR sequences from a FLIR T1050sc camera (1024 X 768) spanning 640 diverse scenes with varying camera and object motion conditions. This encourages future research in infrared VSR. Yang Zou 0004, Xingyue Zhu, Kaiqi Han, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
AAAI | 5 |
| 2026 | Unsupervised Hyperspectral Image Super-Resolution via Self-Supervised Modality Decoupling
Songcheng Du, Yang Zou 0004, Xingyuan Li 0005, Ying Li 0017, Changjing Shang, Qiang Shen 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Contourlet Refinement Gate Framework for Thermal Spectrum Distribution Regularized Infrared Image Super-Resolution
Yang Zou 0004, Zhixin Chen, Xingyuan Li 0005, Long Ma 0002, Jinyuan Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Adversarially robust Fourier-aware multimodal medical image fusion for LSCI
Zhidong Jiao, Yuchun He, Xingyue Zhu, Weipeng Sun, Xingyuan Li 0005, Yang Zou 0004 |
Neurocomputing | 6 |
| 2026 | Structured Prompt-Guided Segmentation of Enhanced CT Images for Predicting Lateral Cervical Lymph Node Metastasis in Papillary Thyroid CarcinomaabstractAccurate segmentation of metastatic lymph nodes in head-and-neck CT is crucial for treatment planning in papillary thyroid carcinoma. Although deep networks have advanced medical image segmentation, most existing methods depend solely on visual features and neglect clinical semantic cues. In this work, we propose a text-guided segmentation framework that introduces structured clinical baseline feature into a YOLOv8-based segmentation network. To effectively fuse semantic and spatial features, we design a cross-modal fusion mechanism that aligns textual embeddings with image features at multiple scales. This allows the model to focus on diagnostically relevant regions and improves performance in challenging scenarios involving low contrast and ambiguous boundaries. Extensive experiments on a multi-center, pathology-confirmed CT dataset demonstrate that our method achieves superior results in AP@50, mAP@[.5:.95], and AR@[.5:.95], outperforming existing state-of-the-art instance segmentation models. Xingyuan Li 0005, Jinyuan Liu 0001, Yanhui Peng |
IEEE Signal Process. Lett. | 3 |
| 2026 | Illumination Refinement via Textual Cues: A Prompt-Driven Approach for Low-Light NeRF EnhancementabstractIn the realm of 3D scene modeling and rendering, the emergence of Neural Radiance Fields (NeRF) represents a significant leap forward. However, NeRF’s rendering performance suffers significantly when rendering images under low-light conditions. Existing approaches are optimized by enhancing low-light input images and combining NeRF models, but still fail to address the issues of multiview consistency and image quality. To address these challenges, our research introduces a textual constraint-prompted enhancement method that facilitates low-light image brightening and new view synthesis in an unsupervised manner. Specifically, we devise a semantic calibration strategy that employs positive and negative prompts to motivate and penalize the network towards attributes associated with high-quality images and exploits the capability of visual language models in semantic parsing to align the generated images with textual descriptors to improve image generation quality. In addition, to address the multiview consistency problem, we propose a two-layer optimization strategy, where the semantic cue optimization in the upper layer and the new view generation in the lower layer interact with each other to achieve a balance between luminance consistency and structural integrity by combining these improved images with text-driven semantic features. Comprehensive tests on two datasets with different resolutions, LOM and LLFF, show that our approach outperforms existing methods by significantly improving the brightness and clarity of low-light images to state-of-the-art while preserving the natural appearance and details. Xinrui Ju, Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | AxisPose: Model-Free Matching-Free Single-Shot 6D Object Pose Estimation via Axis GenerationabstractObject pose estimation is a fundamental task in computer vision and plays an important role in various applications such as robotics, augmented reality, and autonomous manipulation. Existing studies often demand complex inputs or depend on correspondence-based matching between 2D image features and 3D object representations. While effective, these methods rely strongly on explicit appearance matching, often requiring multi-view inputs, depth sensors, or CAD models, which limits their scalability and robustness. Building on top of the pioneering generative studies, we propose AxisPose, a model-free, matching-free, and single-view 6D pose estimation framework that departs from conventional correspondence-based paradigms. Unlike existing methods, AxisPose directly infers a pose representation by learning a latent distribution of object orientation axes through a diffusion model. Specifically, AxisPose introduces an Axis Generation Module (AGM) that progressively denoises tri-axial orientation fields guided by geometric consistency constraints, and a Triaxial Back-projection Module (TBM) to recover the final 6D pose from the generated orientation axes without relying on explicit 2D- 2D/3D correspondences. AxisPose achieves strong cross-instance generalization, enabling a single model to handle multiple object categories without retraining. Extensive experiments on LINEMOD and YCB-Video datasets demonstrate that AxisPose improves the Average Distance Deviation score from 0.733 to 0.814 over the strong baseline NOPE, using only a single RGB input. The code is available at https://github.com/pubyLu/AxisPose/tree/main. Yang Zou 0004, Zhaoshuai Qi, Weipeng Sun, Xingyuan Li 0005, Jiaqi Yang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-ResolutionabstractInfrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challenge imaging quality and subsequent visual tasks. Hence, infrared image super-resolution (IISR) has been developed to address this challenge. While recent developments in diffusion models have greatly advanced this field, current methods to solve it either ignore the unique modal characteristics of infrared imaging or overlook the machine perception requirements. To bridge these gaps, we propose DifIISR, an infrared image super-resolution diffusion model optimized for visual quality and perceptual performance. Our approach achieves task-based guidance for diffusion by injecting gradients derived from visual and perceptual priors into the noise during the reverse process. Specifically, we introduce an infrared thermal spectrum distribution regulation to preserve visual fidelity, ensuring that the reconstructed infrared images closely align with high-resolution images by matching their frequency components. Subsequently, we incorporate various visual foundational models as the perceptual guidance for downstream visual tasks, infusing generalizable perceptual features beneficial for detection and segmentation. As a result, our approach gains superior visual results while attaining State-Of-The-Art downstream task performance. Code is available at https://github.com/zirui0625/DifIISR Xingyuan Li 0005, Yang Zou 0004, Zhixin Chen, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001 |
CVPR | 1 |
| 2025 | DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image FusionabstractInfrared and visible image fusion integrates information from distinct spectral bands to enhance image quality by leveraging the strengths and mitigating the limitations of each modality. Existing approaches typically treat image fusion and subsequent high-level tasks as separate processes, resulting in fused images that offer only marginal gains in task performance and fail to provide constructive feedback for optimizing the fusion process. To overcome these limitations, we propose a Discriminative Cross-Dimension Evolutionary Learning Framework, termed DCEvo, which simultaneously enhances visual quality and perception accuracy. Leveraging the robust search capabilities of Evolutionary Learning, our approach formulates the optimization of dual tasks as a multi-objective problem by employing an Evolutionary Algorithm (EA) to dynamically balance loss function parameters. Inspired by visual neuroscience, we integrate a Discriminative Enhancer (DE) within both the encoder and decoder, enabling the effective learning of complementary features from different modalities. Additionally, our Cross-Dimensional Embedding (CDE) block facilitates mutual enhancement between high-dimensional task features and low-dimensional fusion features, ensuring a cohesive and efficient feature integration process. Experimental results on three benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, achieving an average improvement of 9.32% in visual quality while also enhancing subsequent high-level tasks. The code is available at https://github.com/Beate-Suy-Zhang/DCEvo. Jinyuan Liu 0001, Qingyun Mei, Xingyuan Li 0005, Yang Zou 0004, Zhiying Jiang, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
CVPR | 4 |
| 2025 | Toward a Training-Free Plug-and-Play Refinement Framework for Infrared and Visible Image Registration and FusionabstractInfrared and Visible Image Fusion (IVIF) under unregistered conditions has been of great interest in various visual tasks under challenging environments. While existing approaches often demonstrate promising results on specific benchmarks, they tend to exhibit performance drops in unseen scenarios and incur high computational overhead when retrained on new datasets. To address these challenges, we propose TRACE, a Training-free Reinforcement-based Alignment method for Cross-modality Enhancement, which incorporates Evaluator, a rewarding network, into an evaluation-driven Reinforcement Learning (RL) framework, enabling efficient and plug-and-play refinement of any existing registration approach. Specifically, TRACE constructs the Evaluator network to assess the alignment quality of the given registration model, generating confidence scores and adjustment masks via spatial and channel attention. Leveraging these cues as RL rewards, TRACE iteratively refines the registration network to mitigate misalignments until the accumulated improvement is satisfied. Due to its training-free and plug-and-play nature, TRACE notably enhances fusion results across diverse and unseen scenarios. TRACE achieves impressive improvements in different methods across diverse datasets with minimal computational cost. The project page is available at https://github.com/pubyLu/TRACE. Yang Zou 0004, Xingyuan Li 0005, Xingyue Zhu, Kaiqi Han, Zhiying Jiang, Long Ma 0002, Jinyuan Liu 0001 |
ACM Multimedia | 3 |
| 2025 | Efficient Rectified Flow for Image FusionabstractImage fusion is a fundamental and important task in computer vision, aiming to combine complementary information from different modalities to fuse images. In recent years, diffusion models have made significant developments in the field of image fusion. However, diffusion models often require complex computations and redundant inference time, which reduces the applicability of these methods. To address this issue, we propose RFfusion, an efficient one-step diffusion model for image fusion based on Rectified Flow. We incorporate Rectified Flow into the image fusion task to straighten the sampling path in the diffusion model, achieving one-step sampling without the need for additional training, while still maintaining high-quality fusion results. Furthermore, we propose a task-specific variational autoencoder (VAE) architecture tailored for image fusion, where the fusion operation is embedded within the latent space to further reduce computational complexity. To address the inherent discrepancy between conventional reconstruction-oriented VAE objectives and the requirements of image fusion, we introduce a two-stage training strategy. This approach facilitates the effective learning and integration of complementary information from multi-modal source images, thereby enabling the model to retain fine-grained structural details while significantly enhancing inference efficiency. Extensive experiments demonstrate that our method outperforms other state-of-the-art methods in terms of both inference speed and fusion quality. Tianwei Guan, Xingyuan Li 0005, Minjing Dong, Jinyuan Liu 0001 |
NeurIPS | 5 |
| 2025 | Modeling detail feature connections for infrared image enhancement
Ziang Meng, Kaiqi Han, Yuchun He, Yuhan He, Xingyuan Li 0005, Yang Zou 0004 |
Neurocomputing | 5 |
| 2024 | Towards Robust Image Stitching: An Adaptive Resistance Learning against Compatible AttacksabstractImage stitching seamlessly integrates images captured from varying perspectives into a single wide field-of-view image. Such integration not only broadens the captured scene but also augments holistic perception in computer vision applications. Given a pair of captured images, subtle perturbations and distortions which go unnoticed by the human visual system tend to attack the correspondence matching, impairing the performance of image stitching algorithms. In light of this challenge, this paper presents the first attempt to improve the robustness of image stitching against adversarial attacks. Specifically, we introduce a stitching-oriented attack (SoA), tailored to amplify the alignment loss within overlapping regions, thereby targeting the feature matching procedure. To establish an attack resistant model, we delve into the robustness of stitching architecture and develop an adaptive adversarial training (AAT) to balance attack resistance with stitching precision. In this way, we relieve the gap between the routine adversarial training and benign models, ensuring resilience without quality compromise. Comprehensive evaluation across real-world and synthetic datasets validate the deterioration of SoA on stitching performance. Furthermore, AAT emerges as a more robust solution against adversarial perturbations, delivering superior stitching results. Code is available at: https://github.com/Jzy2017/TRIS. Zhiying Jiang, Xingyuan Li 0005, Jinyuan Liu 0001, Xin Fan 0001, Risheng Liu |
AAAI | 2 |
| 2024 | Enhancing Neural Radiance Fields with Adaptive Multi-Exposure Fusion: A Bilevel Optimization Approach for Novel View SynthesisabstractNeural Radiance Fields (NeRF) have made significant strides in the modeling and rendering of 3D scenes. However, due to the complexity of luminance information, existing NeRF methods often struggle to produce satisfactory renderings when dealing with high and low exposure images. To address this issue, we propose an innovative approach capable of effectively modeling and rendering images under multiple exposure conditions. Our method adaptively learns the characteristics of images under different exposure conditions through an unsupervised evaluator-simulator structure for HDR (High Dynamic Range) fusion. This approach enhances NeRF's comprehension and handling of light variations, leading to the generation of images with appropriate brightness. Simultaneously, we present a bilevel optimization method tailored for novel view synthesis, aiming to harmonize the luminance information of input images while preserving their structural and content consistency. This approach facilitates the concurrent optimization of multi-exposure correction and novel view synthesis, in an unsupervised manner. Through comprehensive experiments conducted on the LOM and LOL datasets, our approach surpasses existing methods, markedly enhancing the task of novel view synthesis for multi-exposure environments and attaining state-of-the-art results. The source code can be found at https://github.com/Archer-204/AME-NeRF. Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Jinyuan Liu 0001 |
AAAI | 2 |
| 2024 | Contourlet Residual for Prompt Learning Enhanced Infrared Image Super-Resolution
Xingyuan Li 0005, Jinyuan Liu 0001, Zhixin Chen, Yang Zou 0004, Long Ma 0002, Xin Fan 0001, Risheng Liu |
ECCV (3) | 1 |
| 2024 | AEAM3D: Adverse Environment-Adaptive Monocular 3D Object Detection via Feature Extraction Regularizationabstract3D object detection plays a crucial role in intelligent vision systems. Detection in the open world inevitably encounters various adverse scenes while most of existing methods fail in these scenes. To address this issue, this paper proposes a monocular 3D detection model, termed AEAM3D, which effectively mitigates the degradation of detection performance in various harsh environments. Additionally, we assemble a new adverse 3D object detection dataset encompassing some challenging scenes, including rainy, foggy, and low light weather conditions. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches by an average of 3.12% in terms of APR40for car category across adverse environments. Yixin Lei, Xingyuan Li 0005, Zhiying Jiang, Xinrui Ju, Jinyuan Liu 0001 |
ICASSP | 2 |
| 2024 | Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders images across a spectrum of exposure conditions. Our approach utilizes an unsupervised classifier-generator structure for HDR fusion, significantly enhancing NeRF’s ability to comprehend and adjust to light variations, leading to the generation of images with appropriate brightness. Extensive evaluations on the LOM[1] and LOL[2] datasets underscore our method’s edge. Our approach significantly improves the task of novel view synthesis for multi-exposure images, attaining state-of-the-art results. Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Tiantian Yan, Jinyuan Liu 0001 |
ICASSP | 2 |