VLDB 2026 Research / reviewers in the wild / expert
Zeyu Xiao 0002
dblp:276/3139-2
· DBLP profile ↗
40ranked-venue papers
17as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 15 first-author · 35 since 2021Artificial intelligence and machine learning · 20 · 9 first-author · 20 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing the Unseen: Zooming in the Dark with Event CamerasabstractThis paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these challenges, we present RetinexEVSR, the first event-driven LVSR framework that leverages high-contrast event signals and Retinex-inspired priors to enhance video quality under low-light scenarios. Unlike previous approaches that directly fuse degraded signals, RetinexEVSR introduces a novel bidirectional cross-modal fusion strategy to extract and integrate meaningful cues from noisy event data and degraded RGB frames. Specifically, an illumination-guided event enhancement module is designed to progressively refine event features using illumination maps derived from the Retinex model, thereby suppressing low-light artifacts while preserving high-contrast details. Furthermore, we propose an event-guided reflectance enhancement module that utilizes the enhanced event features to dynamically recover reflectance details via a multi-scale fusion mechanism. Experimental results show that our RetinexEVSR achieves state-of-the-art performance on three datasets. Notably, on the SDSD benchmark, our method can get up to 2.95 dB gain while reducing runtime by 65% compared to prior event-based methods. Dachun Kai, Zeyu Xiao 0002, Huyue Zhu, Jiaxiao Wang, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
AAAI | 2 |
| 2026 | FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image GenerationabstractLayout-to-Image generation has significantly advanced content creation by enabling the rendering of visual text under predefined spatial layouts. Current approaches achieve training-free layout guidance by constructing attention-based energy functions to derive correction gradients. In this paper, we demonstrate that vanilla energy functions suffer from two limitations, resulting in imprecise layout control and visually unrealistic artifacts. First, the normalizing factor of the Boltzmann distribution defined by the energy functions is non-negligible when calculating correction gradients, yet current energy functions cannot compute this factor exactly. Furthermore, while attention varies over time during the denoising process, existing approaches employ a fixed formulation. To address these challenges, we introduce FreLay, a novel training-free approach equipped with a frequency-aware energy function. Our method first reformulates the energy function to handle the normalization factor, enabling accurate computation of correction gradients. Simultaneously, leveraging the prior knowledge that low-frequency information deteriorates slower during noise addition, we design a time-specific energy function for each timestep from a frequency-domain perspective. Experimental results demonstrate that FreLay consistently outperforms existing state-of-the-art training-free methods by a large margin both qualitatively and quantitatively across multiple datasets. Bonan Li, Yinhan Hu, Songhua Liu, Zeyu Xiao 0002, Xinchao Wang |
AAAI | 4 |
| 2026 | Event-Guided Scene Text Image Super-ResolutionabstractScene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods. Zihan Qi, Zeyu Xiao 0002, Haoyi Zhao, Yang Zhao 0002, Feng Xue 0002, Wei Jia 0001 |
AAAI | 2 |
| 2026 | Exploiting Blurry Representations for Event-guided Video Super-ResolutionabstractBlurry video super-resolution (BVSR) remains fundamentally ill-posed due to the simultaneous loss of high-frequency spatial details and reliable motion cues in blurry low-resolution frames. While cascade-based and joint BVSR methods struggle under severe blur, existing event-guided VSR approaches largely assume clean inputs and are ineffective against complex motion degradation. These methods fail to model blurry representations or leverage event signals for blur-aware motion cues, leading to sub-optimal performance. We propose BluR-EVSR, a unified framework that implicitly models Blurry Representations and leverages Event cameras to jointly address both blur and resolution degradation for VSR. The framework begins with a self-supervised degradation learning strategy guided by event streams and neighboring frames, enabling adaptive blur representation without requiring explicit supervision. A dynamic routing mechanism encodes spatially varying degradations, while a motion-saliency degradation-aware attention module injects motion saliency priors to facilitate efficient RGB-event fusion. Integrated into a bidirectional recurrent framework, BluR-EVSR enables temporally consistent and detail-preserving restoration with low computational cost. Extensive experiments across multiple benchmarks show that our method significantly outperforms prior BVSR and event-based approaches. Zeyu Xiao 0002, Xinchao Wang |
AAAI | 1 |
| 2026 | Event-Based Dynamic Turbulence MitigationabstractAtmospheric turbulence induces coupled spatio-temporal distortions, including blur, geometric deformation, and temporal jitter, which severely degrade image quality. We propose EvTurM, a practical framework leveraging event camera data for dynamic turbulence mitigation with precise motion cues and stable temporal modeling. Leveraging the high temporal resolution and dynamic range of events, EvTurM achieves robust restoration under diverse turbulence conditions. EvTurM comprises two key modules: (1) the event-aware modality enhancement module, which uses event-derived motion to enrich RGB features and recover structural details, and (2) the bidirectional modality calibration module, which jointly aligns RGB and event features in forward and backward propagation to reduce misalignment and enhance temporal consistency. Extensive experiments show EvTurM consistently surpasses existing methods and achieves superior performance. Haoyi Zhao, Zeyu Xiao 0002, Zihan Qi, Yang Zhao 0002, Wei Jia 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | Learning Implicit and Detail-Enhanced Network for Light Field Image Spatial-Angular Super-ResolutionabstractLight field (LF) imaging holds immense promise for applications such as post-capture refocusing and virtual reality. However, its inherent spatial-angular trade-off significantly limits both spatial and angular resolution, restricting its practicality in real-world scenarios. To address these limitations, spatial-angular super-resolution methods have been proposed to simultaneously enhance both dimensions. Yet, existing methods struggle to fully exploit the intertwined spatial-angular correlations and fail to effectively handle sparsely sampled LFs with low spatial resolution, often leading to cumulative errors during reconstruction. In this paper, we propose an Implicit and Detail-Enhanced Network (IDNet) to overcome these challenges. Our IDNet employs 3D convolution for the joint extraction of spatial and angular information, leveraging their interdependencies for more effective LF reconstruction. Additionally, we introduce an implicit detail restoration module that enhances features while encoding positional information to refine fine details. To overcome the limitations of sparse spatial and angular information on high-detail reconstruction and angular consistency in low-resolution LFs, we design a multi-representation enhancement block. This block enhances features by learning pixel differences across multiple directions in diverse representations, effectively capturing intricate details and complex correlations. Thanks to these designs, our IDNet reconstructs novel views with finer details, effectively learns occlusion relationships, and ensures geometric consistency. Experimental results on benchmark datasets demonstrate its superior quantitative and qualitative performance. The code is publicly available at https://github.com/ldyorchid/IDNet. Deyang Liu, Shizheng Li, Xiaofei Zhou 0003, Zeyu Xiao 0002, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Wiener-Deconvolution-Driven Event-Based Deblurring for Low-Light ImagingabstractWe address event-based deblurring for low-light imaging, where conventional frames suffer severe blur, noise and saturation, while events capture sharp high-frequency contrast changes with microsecond latency that can guide the recovery of lost structures. Existing event-based reconstruction methods neither explicitly model low-light noise and saturation nor enforce precise alignment between events and frames, which limits cross-modal fusion and deblurring quality. We propose the Wiener-Deconvolution-Driven Event-Based Deblurring Network (WiED-Net), which embeds the Wiener deconvolution into a deep architecture so that the physical imaging model and noise statistics are encoded in the frequency domain and high-frequency recovery is stabilized on noise dominated night data. WiED-Net adopts a two stage design. The first stage applies Wiener deconvolution in both image and feature spaces to suppress noise, recover saturated regions and reduce ringing, assisted by an eventguided cross-modal feature fusion (ECFF) module for accurate alignment. The second stage uses a multi-scale fusion module to integrate the complementary event and image branches. Training is constrained by a set of losses, including a tailored blur kernel loss that provides closed-loop regularization from physical priors. Together, these designs enable WiED-Net to recover fine details while robustly suppressing artifacts and noise, and to achieve superior quantitative and qualitative performance, achieving superior quantitative and qualitative performance with a notable improvement of 1.97 dB in PSNR and 5% in SSIM over the previous state-of-the-art methods in low-light deblurring. Code will be available at https://github.com/zhuzifeng38/WiED-Net. Zeyu Xiao 0002, Jianlong Jin, Feng Xue 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Event-Guided Online Video Super-ResolutionabstractEvent-guided video super-resolution (VSR) leverages high-temporal-resolution event streams to address motion blur, rapid dynamics, and poor illumination that challenge frame-only VSR methods. However, most existing approaches emphasize reconstruction quality while overlooking real-time performance and computational efficiency, limiting their deployment in latency-sensitive scenarios. To overcome these issues, we present E2VSR, a lightweight and Efficient Event-guided VSR framework tailored for real-time applications. Operating under a causal setting with only current and past observations, E2VSR is designed for low-latency event-guided VSR. We propose an event-confidence adaptive propagation strategy comprising two key modules: the Event-induced Feature Modulation (EvFM) block for robust cross-modal event-frame integration, and the Event-Confidence Feature Fusion (EvCFF) block, which exploits events as motion cues for adaptive inter-frame aggregation. This design improves motion-aware temporal aggregation in challenging dynamic conditions, where event cues may provide complementary temporal information. Furthermore, an Implicit Event Reconstruction (IER) technique leverages event information during training to enrich feature representations without adding inference-time cost, enhancing spatial and temporal fidelity. Experimental results demonstrate that E2VSR achieves superior quantitative and qualitative performance while maintaining a low parameter count and computational cost. Zeyu Xiao 0002, Xinchao Wang |
IEEE Trans. Image Process. | 1 |
| 2026 | Learning Dual Modality Interactions for Event-Based Motion DeblurringabstractEvent cameras hold great potential for motion deblurring because they capture motion information with microsecond precision, offering robustness to motion blur. However, the limited interaction between RGB frames and event streams presents a significant challenge, preventing the full utilization of the event cameras' unique advantages. To address this, we proposeDual frame-eventInteraction and introduce a multi-scaleNetwork structure, DuInt-Net. DuInt-Net aims to tackle two key challenges: (1) enhancing the representational and interaction capabilities between RGB frames and event streams, and (2) adaptively selecting richer visual features for improved motion deblurring. We introduce an event-frame joint interaction module that consists of three branches: a base branch, a global awareness attention branch, and a local enhancement attention branch. The base branch processes essential pixel-level features that retain the original structural information. The global branch integrates event data to improve large-scale motion understanding, while the local branch uses large-kernel convolutions to refine fine-grained details in RGB frames. For superior reconstruction performance, we also propose the event-guided multi-scale fusion attention module, which effectively combines local visual information and global frame-event relationships. Extensive experiments demonstrate that DuInt-Net achieves superior performance, both quantitatively and qualitatively, showcasing its superior motion deblurring capabilities. Zeyu Xiao 0002, Zhuoyuan Li 0001, Yang Zhao 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Event-Enhanced Blurry Video Super-ResolutionabstractIn this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insufficient motion information for deconvolution and the lack of high-frequency details in LR frames. To address these challenges, we introduce event signals into BVSR and propose a novel event-enhanced network, Ev-DeblurVSR. To effectively fuse information from frames and events for feature deblurring, we introduce a reciprocal feature deblurring module that leverages motion information from intra-frame events to deblur frame features while reciprocally using global scene context from the frames to enhance event features. Furthermore, to enhance temporal consistency, we propose a hybrid deformable alignment module that fully exploits the complementary motion information from inter-frame events and optical flow to improve motion estimation in the deformable alignment process. Extensive evaluations demonstrate that Ev-DeblurVSR establishes a new state-of-the-art performance on both synthetic and real-world datasets. Notably, on real data, our method is 2.59 dB more accurate and 7.28× faster than the recent best BVSR baseline FMA-Net. Dachun Kai, Yueyi Zhang 0001, Jin Wang 0023, Zeyu Xiao 0002, Zhiwei Xiong, Xiaoyan Sun 0001 |
AAAI | 4 |
| 2025 | Occlusion-Embedded Hybrid Transformer for Light Field Super-ResolutionabstractTransformer-based networks have set new benchmarks in light field super-resolution (SR), but adapting them to capture both global and local spatial-angular correlations efficiently remains challenging. Moreover, many methods fail to account for geometric details like occlusions, leading to performance drops. To tackle these issues, we introduce OHT. This hybrid network leverages occlusion maps through an occlusion-embedded mix layer. It combines the strengths of convolutional networks and Transformers via spatial-angular separable convolution (SASep-Conv) and angular self-attention (ASA). SASep-Conv offers a lightweight alternative to 3D convolution for capturing spatial-angular correlations, while the ASA mechanism applies 3D self-attention across the angular dimension. These designs allow OHT to capture global angular correlations effectively. Extensive experiments on multiple datasets demonstrate OHT's superior performance. Zeyu Xiao 0002, Zhuoyuan Li 0001, Wei Jia 0001 |
AAAI | 1 |
| 2025 | Event-based Video Super-Resolution via State Space ModelsabstractExploiting temporal correlations is crucial for video super-resolution (VSR). Recent approaches enhance this by incorporating event cameras. In this paper, we introduce MamEVSR, a Mamba-based network for event-based VSR that leverages the selective state space model, Mamba. MamEVSR stands out by offering global receptive field coverage with linear computational complexity, thus addressing the limitations of convolutional neural networks and Transformers. The key components of MamEVSR include: (1) The interleaved Mamba (iMamba) block, which interleaves tokens from adjacent frames and applies multidirectional selective state space modeling, enabling efficient feature fusion and propagation across bi-directional frames while maintaining linear complexity. (2) The cross-modality Mamba (cMamba) block facilitates further interaction and aggregation between event information and the output from the iMamba block. The cMamba block can leverage complementary spatio-temporal information from both modalities and allows MamEVSR to capture finer motion details. Experimental results show that the proposed MamEVSR achieves superior performance on various datasets quantitatively and qualitatively. Zeyu Xiao 0002, Xinchao Wang |
CVPR | 1 |
| 2025 | Asymmetric Dual-Lens Video DeblurringabstractModern smartphones often feature asymmetric dual-lens systems, capturing wide-angle and ultra-wide views with complementary perspectives and details. Motion and shake can blur the wide lens, while the ultra-wide lens, despite lower resolution, retains sharper details. This natural complementarity offers valuable cues for video deblurring. However, existing methods focus mainly on single-camera inputs or symmetric stereo pairs, neglecting the cross-lens redundancy in mobile dual-camera systems. In this paper, we propose a practical video deblurring method, AsLeD-Net, which recurrently aligns and propagates temporal reference features from ultra-wide views fused with features extracted from wide-angle blurry frames. AsLeD-Net consists of two key modules: the adaptive local matching (ALM) module, which refines blurry features using $K$-nearest neighbor reference features, and the difference compensation (DC) module, which ensures spatial consistency and reduces misalignment. Additionally, AsLeD-Net uses the reference-guided motion compensation (RMC) module for temporal alignment, further improving frame-to-frame consistency in the deblurring process. We validate the effectiveness of AsLeD-Net through extensive experiments, benchmarking it against potential solutions for asymmetric lens deblurring. Zeyu Xiao 0002, Xinchao Wang |
NeurIPS | 1 |
| 2025 | Incorporating degradation estimation in light field spatial super-resolutionabstractRecent advancements in light field super-resolution (SR) have yielded impressive results. In practice, however, many existing methods are limited by assuming fixed degradation models , such as bicubic downsampling, which hinders their robustness in real-world scenarios with complex degradations. To address this limitation, we present LF-DEST, an effective blind L ight F ield SR method that incorporates explicit D egradation Est imation to handle various degradation types. LF-DEST consists of two primary components: degradation estimation and light field restoration. The former concurrently estimates blur kernels and noise maps from low-resolution degraded light fields, while the latter generates super-resolved light fields based on the estimated degradations. Notably, we introduce a modulated and selective fusion module that intelligently combines degradation representations with image information, effectively handling diverse degradation types. We conduct extensive experiments on benchmark datasets, demonstrating that LF-DEST achieves superior performance across various degradation scenarios in light field SR. The implementation code is available at https://github.com/zeyuxiao1997/LF-DEST . Zeyu Xiao 0002, Zhiwei Xiong |
Comput. Vis. Image Underst. | 1 |
| 2025 | L3FMamba: Low-Light Light Field Image Enhancement With Prior-Injected State Space ModelsabstractIn this paper, we address the problem of low-light light field (LF) image enhancement, where spatial details and angular coherence are severely degraded due to noise and insufficient illumination. Existing methods often rely on local aggregation or naive view stacking, which fail to capture global illumination and long-range spatial-angular correlations. To overcome these limitations, we propose L3FMamba, a lightweight enhancement method that integrates Retinex and Atmospheric Scattering models with dark, bright, and average channel priors for robust illumination decomposition. Moreover, we incorporate a state space model to capture non-local spatial-angular dependencies, enabling effective propagation of global context across views. By combining physics-inspired priors with structured modeling, L3FMamba achieves accurate illumination correction and fine-detail preservation with minimal parameters. Experiments show that L3FMamba outperforms the state-of-the-art in quality. Deyang Liu, Shizheng Li, Zeyu Xiao 0002, Ping An 0001, Caifeng Shan |
IEEE Signal Process. Lett. | 3 |
| 2025 | Task-to-Instance Prompt Learning for Vision-Language Models at Test TimeabstractPrompt learning has been recently introduced into the adaption of pre-trained vision-language models (VLMs) by tuning a set of trainable tokens to replace hand-crafted text templates. Despite the encouraging results achieved, existing methods largely rely on extra annotated data for training. In this paper, we investigate a more realistic scenario, where only the unlabeled test data is available. Existing test-time prompt learning methods often separately learn a prompt for each test sample. However, relying solely on a single sample heavily limits the performance of the learned prompts, as it neglects the task-level knowledge that can be gained from multiple samples. To that end, we propose a novel test-time prompt learning method of VLMs, called Task-to-Instance PromPt LEarning (TIPPLE), which adopts a two-stage training strategy to leverage both task- and instance-level knowledge. Specifically, we reformulate the effective online pseudo-labeling paradigm along with two tailored components: an auxiliary text classification task and a diversity regularization term, to serve the task-oriented prompt learning. After that, the learned task-level prompt is further combined with a tunable residual for each test sample to integrate with instance-level knowledge. We demonstrate the superior performance of TIPPLE on 15 downstream datasets, e.g., the average improvement of 1.87% over the state-of-the-art method, using ViT-B/16 visual backbone. Our code is open-sourced at https://github.com/zhiheLu/TIPPLE. Zhihe Lu, Jiawang Bai, Xin Li 0082, Zeyu Xiao 0002, Xinchao Wang |
IEEE Trans. Image Process. | 4 |
| 2025 | Deep Sparse-to-Dense Inbetweening for Multi-View Light FieldsabstractLight field (LF) imaging, which captures both intensity and directional information of light rays, extends the capabilities of traditional imaging techniques. In this paper, we introduce a task in the field of LF imaging, sparse-to-dense inbetweening, which focuses on generating dense novel views from sparse multi-view LFs. By synthesizing intermediate views from sparse inputs, this task enhances LF view synthesis through filling in interperspective gaps within an expanded field of view and increasing data robustness by leveraging complementary information between light rays from different perspectives, which are limited by non-robust single-view synthesis and the inability to handle sparse inputs effectively. To address these challenges, we construct a high-quality multi-view LF dataset, consisting of 60 indoor scenes and 59 outdoor scenes. Building upon this dataset, we propose a baseline method. Specifically, we introduce an adaptive alignment module to dynamically align information by capturing relative displacements. Next, we explore angular consistency and hierarchical information using a multi-level feature decoupling module. Finally, a multi-level feature refinement module is applied to enhance features and facilitate reconstruction. Additionally, we introduce a universally applicable artifact-aware loss function to effectively suppress visual artifacts. Experimental results demonstrate that our method outperforms existing approaches, establishing a benchmark for sparse-to-dense inbetweening. The code is available at https://github.com/Starmao1/MutiLF. Zeyu Xiao 0002, Ping An 0001, Deyang Liu, Caifeng Shan |
IEEE Trans. Image Process. | 2 |
| 2025 | BVSR-EvD: Blurry Video Space-Time Super-Resolution With Events via Diffusion ModelsabstractVideo restoration from low-resolution and low-frame-rate blurry sources remains challenging due to insufficient data priors. In this paper, we propose BVSR-EvD, leveraging event cameras and diffusion models to boost blurry video space-time super-resolution. Specifically, we identify three distinct data priors from event-video dual modalities: motion prior from events, content prior from videos, and physical prior from their integration, contributing to temporal stability, content preservation, and detail enhancement respectively. To effectively utilize these data priors, BVSR-EvD creates the Trident Diffusion Model (Trident-DM), which decomposes each denoising step into trident decoupling and adaptive self-composition stages. The former employs single-modal and dual-modal meta-networks to extract the three unique data priors, while the latter dynamically integrates them through learned prior-aware weight maps. BVSR-EvD achieves up to $\times 8$ spatial super-resolution and $\times 64$ temporal super-resolution from blurry videos, surpassing existing methods on public video datasets. Wenming Weng, Yueyi Zhang 0001, Zeyu Xiao 0002, Zhiwei Xiong |
IEEE Trans. Image Process. | 4 |
| 2025 | ALOHA: Adapting Local Spatio-Temporal Context to Enhance the Audio-Visual Semantic SegmentationabstractAudio-Visual Semantic Segmentation (AVSS) plays a crucial role in pixel-level multi-modal perception for real-world applications such as robotic navigation and autonomous driving. Existing methods typically rely on global spatio-temporal modules to fuse audio and visual representations, which aids in generating pixel-level semantic masks. However, these approaches often overlook the importance of local spatio-temporal context in understanding semantics, leading to suboptimal performance. This limitation makes it difficult for models to accurately distinguish sound-emitting objects from irrelevant background noise, resulting in erroneous segmentation across the spatio-temporal dimension. To address this issue, we propose the ALOHA framework, which A dapts LO cal spatio-temporal context to en HA nce AVSS. The framework introduces two key components designed to leverage and enhance local spatio-temporal context information: the LOHA adapter and the Selective Context Enhancement (SCE) module. Specifically, the LOHA adapter adaptively captures essential modality information across spatio-temporal dimensions, while implicitly learning fine-grained local context through the local attention mechanism. Furthermore, the SCE module selectively enhances the local context related to the semantics, thereby facilitating the distinction between the sounding object and irrelevant background and improving segmentation accuracy. Moreover, to better adapt to embodied AI systems, our framework utilizes a parameter-shared encoder and applies the adapters in a staged manner. This design significantly reduces the number of trainable parameters, making it more parameter-efficient. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance on the AVSBench-Semantic benchmark dataset and shows competitive results on the AVSBench-Object benchmark, while exhibiting broad adaptability across different visual backbone networks. Yang-Hao Zhou, Heyan Huang, Cunhan Guo, Rongcheng Tu, Zeyu Xiao 0002, Bo Wang 0134, Xianling Mao |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Mamba-Based Light Field Super-Resolution with Efficient Subspace Scanning
Ruisheng Gao, Zeyu Xiao 0002, Zhiwei Xiong |
ACCV (5) | 2 |
| 2024 | Learning Complementary Maps for Light Field Salient Object Detection
Zeyu Xiao 0002, Jiateng Shou, Zhiwei Xiong |
ACCV (5) | 1 |
| 2024 | Learning Large-Factor EM Image Super-Resolution with Generative PriorsabstractAs the mainstream technique for capturing images of biological specimens at nanometer resolution, electron microscopy (EM) is extremely time-consuming for scanning wide field-of-view (FOV) specimens. In this paper, we investigate a challenging task of large-factor EM image super-resolution (EMSR), which holds great promise for reducing scanning time, relaxing acquisition conditions, and expanding imaging FOV. By exploiting the repetitive structures and volumetric coherence of EM images, we propose the first generative learning-based framework for large-factor EMSR. Specifically, motivated by the predictability ofrepetitive structures and textures in EM images, we first learn a discrete codebook in the latent space to represent highresolution (HR) cell-specific priors and a latent vector indexer to map low-resolution (LR) EM images to their corresponding latent vectors in a generative manner. By incorporating the generative cell-specific priors from HR EM images through a multi-scale prior fusion module, we then deploy multi-image feature alignment and fusion to further exploit the inter-section coherence in the volumetric EM data. Extensive experiments demonstrate that our proposed framework outperforms advanced single-image and video super-resolution methods for 8× and 16× EMSR (i.e., with 64 times and 256 times less data acquired, respectively), achieving superior visual reconstruction quality and down-stream segmentation accuracy on benchmark EM datasets. Code is available at https://github.com/jtshou/GPEMSR. Jiateng Shou, Zeyu Xiao 0002, Shiyu Deng, Wei Huang 0036, Peiyao Shi, Ruobing Zhang, Zhiwei Xiong, Feng Wu 0001 |
CVPR | 2 |
| 2024 | Event-Adapted Video Super-Resolution
Zeyu Xiao 0002, Dachun Kai, Yueyi Zhang 0001, Zhengjun Zha, Xiaoyan Sun 0001, Zhiwei Xiong |
ECCV (42) | 1 |
| 2024 | Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language ModelsabstractFine-tuning pre-trained vision-language models (VLMs), e.g., CLIP, for the open-world generalization has gained increasing popularity due to its practical value. However, performance advancements are limited when relying solely on intricate algorithmic designs for a single model, even one exhibiting strong performance, e.g., CLIP-ViT-B/16. This paper, for the first time, explores the collaborative potential of leveraging much weaker VLMs to enhance the generalization of a robust single model. The affirmative findings motivate us to address the generalization problem from a novel perspective, i.e., ensemble of pre-trained VLMs. We introduce three customized ensemble strategies, each tailored to one specific scenario. Firstly, we introduce the zero-shot ensemble, automatically adjusting the logits of different models based on their confidence when only pre-trained VLMs are available. Furthermore, for scenarios with extra few-shot samples, we propose the training-free and tuning ensemble, offering flexibility based on the availability of computing resources. The code is available at https://github.com/zhiheLu/Ensemble_VLM.git. Zhihe Lu, Jiawang Bai, Xin Li 0082, Zeyu Xiao 0002, Xinchao Wang |
ICML | 4 |
| 2024 | Asymmetric Event-Guided Video Super-ResolutionabstractEvent cameras are novel bio-inspired cameras that record asynchronous events with high temporal resolution and dynamic range. Leveraging the auxiliary temporal information recorded by event cameras holds great promise for the task of video super-resolution (VSR). However, existing event-guided VSR methods assume that the event and RGB cameras are strictly calibrated (e.g., pixel-level sensor designs in DAVIS 240/346). This assumption proves limiting in emerging high-resolution devices, such as dual-lens smartphones and unmanned aerial vehicles, where such precise calibration is typically unavailable. To unlock more event-guided application scenarios, we perform the task of asymmetric event-guided VSR for the first time, and we propose an Asymmetric Event-guided VSR Network (AsEVSRN) for this new task. AsEVSRN incorporates two specialized designs for leveraging the asymmetric event stream in VSR. Firstly, the content hallucination module dynamically enhances event and RGB information by exploiting their complementary nature, thereby adaptively boosting representational capacity. Secondly, the event-enhanced bidirectional recurrent cells align and propagate temporal features fused with features from content-hallucinated frames. Within the bidirectional recurrent cells, event-enhanced flow is employed to simultaneously utilize and fuse temporal information at both the feature and pixel levels. Comprehensive experimental results affirm that our method consistently generates superior quantitative and qualitative results. Zeyu Xiao 0002, Dachun Kai, Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ACM Multimedia | 1 |
| 2024 | Unraveling Motion Uncertainty for Local Motion DeblurringabstractIn real-world photography, local motion blur often arises from the interplay between moving objects and stationary backgrounds during exposure. Existing deblurring methods face challenges in addressing local motion deblurring due to (i) the presence of arbitrary localized blurs and uncertain blur extents; (ii) the limited ability to accurately identify specific blurs resulting from ambiguous motion boundaries. These limitations often lead to suboptimal solutions when estimating blur maps and generating final deblurred images. To that end, we propose a novel method named Motion-Uncertainty-Guided Network (MUGNet), which harnesses a probabilistic representational model to explicitly address the intricacies stemming from motion uncertainties. Specifically, MUGNet consists of two key components, i.e., motion-uncertainty quantification (MUQ) module and motion-masked separable attention (M2SA) module, serving for complementary purposes. Concretely, MUQ aims to learn a conditional distribution for accurate and reliable blur map estimation, while the M2SA module is to enhance the representation of regions influenced by local motion blur and static background, which is achieved by promoting the establishment of extensive global interactions. We demonstrate the superiority of our MUGNet with extensive experiments. The code is publicly available at: https://github.com/zeyuxiao1997/MUGNet. Zeyu Xiao 0002, Zhihe Lu, Michael Bi Mi, Zhiwei Xiong, Xinchao Wang |
ACM Multimedia | 1 |
| 2024 | P-BiC: Ultra-High-Definition Image Moiré Patterns Removal via Patch Bilateral CompensationabstractPeople nowadays use smartphones to capture photos from multimedia platforms. The presence of moire patterns resulting from spectral aliasing can significantly degrade the visual quality of images, particularly in ultra-high-definition (UHD) images. However, existing demoireing methods have mostly been designed for low-definition images, making them unsuitable for handling moire patterns in UHD images due to their substantial memory requirements. In this paper, we propose a novel patch bilateral compensation network (P-BiC) for the demoire pattern removal in UHD images, which is memory-efficient and prior-knowledge-based. Specifically, we divide the UHD images into small patches and perform patch-level demoireing to maintain the low memory cost even for ultra-large image sizes. Moreover, a pivotal insight, namely that the green channel of an image remains relatively less affected by moire patterns, while the tone information in moire images is still well-retained despite color shifts, is directly harnessed for the purpose of bilateral compensation. The bilateral compensation is achieved by two key components in our P-BiC, i.e., a green-guided detail transfer (G2DT) module that complements distorted features with the intact content, and a style-aware tone adjustment (STA) module for the color adjustment. We quantitatively and qualitatively evaluate the effectiveness of P-BiC with extensive experiments. The code is publicly available at: https://github.com/zeyuxiao1997/P-BiC. Zeyu Xiao 0002, Zhihe Lu, Xinchao Wang |
ACM Multimedia | 1 |
| 2024 | TSA2: Temporal Segment Adaptation and Aggregation for Video HarmonizationabstractVideo composition merges the foreground and background of different videos, presenting challenges due to variations in capture conditions (e.g., saturation, brightness, and contrast). Video harmonization is a vital process in achieving a realistic composite by seamlessly adjusting the foreground’s appearance to match the background. In this paper, we propose TSA2, a novel method for video harmonization that incorporates temporal segment adaptation and aggregation. TSA2divides the inharmonious input sequence into temporal segments, each corresponding to a different frame rate, allowing effective utilization of complementary information within each segment. The method includes the Temporal Segment Adaptation module, which learns and remaps the distribution difference between background and foreground regions, and the Temporal Segment Aggregation module, which emphasizes and aggregates cross-segment information through element-wise correlations. Experimental results demonstrate that TSA2outperforms advanced image and video harmonization methods quantitatively and qualitatively. Zeyu Xiao 0002, Yurui Zhu, Xueyang Fu, Zhiwei Xiong |
WACV | 1 |
| 2024 | Light Field Super-Resolution Using Decoupled Selective MatchingabstractNon-local self-similarity has been well exploited in the single image super-resolution task as an effective prior. However, due to the difficulty of modeling the 4D correspondence globally, the potential of the non-local prior is less revealed for light field (LF) super-resolution. Meanwhile, existing non-local models only utilize the global spatial correspondence, but largely neglect the global geometric correspondence. To address the aforementioned problems, we propose a Decoupled Selective Matching Network (DSMNet) for LF super-resolution, by designing a novel selective matching mechanism to flexibly extract non-local information from specific 4D positions in an LF. Such a mechanism matches the reference patch with several auxiliary patches dynamically searched from predefined windows, which promotes efficiency while improving performance compared to the existing non-local models. Specifically, our DSMNet decouples the whole LF into Sub-Aperture Images (SAIs) and Epipolar Plane Images (EPIs). For each SAI patch, we separately perform the selective matching inside the current SAI and cross different SAIs to exploit the global spatial correspondence efficiently. For each EPI patch, we separately perform the selective matching in EPIs of different orientations to embed robust LF geometric information into features by enhancing EPI textures, which exploits the global geometric correspondence in an efficient manner. Comprehensive experiments validate that DSMNet outperforms state-of-the-art LF super-resolution methods both quantitatively and qualitatively. Zhen Cheng 0002, Zeyu Xiao 0002, Zhiwei Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | CutMIB: Boosting Light Field Super-Resolution via Multi-View Image BlendingabstractData augmentation (DA) is an efficient strategy for improving the performance of deep neural networks. Recent DA strategies have demonstrated utility in single image super-resolution (SR). Little research has, however, focused on the DA strategy for light field SR, in which multi-view information utilization is required. For the first time in light field SR, we propose a potent DA strategy called CutMIB to improve the performance of existing light field SR networks while keeping their structures unchanged. Specifically, Cut-MIB first cuts low-resolution (LR) patches from each view at the same location. Then CutMIB blends all LR patches to generate the blended patch and finally pastes the blended patch to the corresponding regions of high-resolution light field views, and vice versa. By doing so, CutMIB enables light field SR networks to learn from implicit geometric information during the training stage. Experimental results demonstrate that CutMIB can improve the reconstruction performance and the angular consistency of existing light field SR networks. We further verify the effectiveness of CutMIB on real-world light field SR and light field denoising. The implementation code is available at https://github.com/zeyuxiao1997/CutMIB. Zeyu Xiao 0002, Ruisheng Gao, Zhiwei Xiong |
CVPR | 1 |
| 2023 | Space-Time Super-Resolution for Light Field VideosabstractLight field (LF) cameras suffer from a fundamental trade-off between spatial and angular resolutions. Additionally, due to the significant amount of data that needs to be recorded, the Lytro ILLUM, a modern LF camera, can only capture three frames per second. In this paper, we consider space-time super-resolution (SR) for LF videos, aiming at generating high-resolution and high-frame-rate LF videos from low-resolution and low-frame-rate observations. Extending existing space-time video SR methods to this task directly will meet two key challenges: 1) how to re-organize sub-aperture images (SAIs) efficiently and effectively given highly redundant LF videos, and 2) how to aggregate complementary information between multiple SAIs and frames considering the coherence in LF videos. To address the above challenges, we propose a novel framework for space-time super-resolving LF videos for the first time. First, we propose a novel Multi-Scale Dilated SAI Re-organization strategy for re-organizing SAIs into auxiliary view stacks with decreasing resolution as the Chebyshev distance in the angular dimension increases. In particular, the auxiliary view stack with original resolution preserves essential visual details, while the down-scaled view stacks capture long-range contextual information. Second, we propose the Multi-Scale Aggregated Feature extractor and the Angular-Assisted Feature Interpolation module to utilize and aggregate information from the spatial, angular, and temporal dimensions in LF videos. The former aggregates similar contents from different SAIs and frames for subsequent reconstruction in a disparity-free manner at the feature level, whereas the latter interpolates intermediate frames temporally by implicitly aggregating geometric information. Compared to other potential approaches, experimental results demonstrate that the reconstructed LF videos generated by our framework achieve higher reconstruction quality and better preserve the LF parallax structure and temporal consistency. The implementation code is available at https://github.com/zeyuxiao1997/LFSTVSR. Zeyu Xiao 0002, Zhen Cheng 0002, Zhiwei Xiong |
IEEE Trans. Image Process. | 1 |
| 2023 | Low-Light Stereo Image EnhancementabstractStereo cameras are now commonly used in more and more devices. Nevertheless, visually unpleasant images captured under low-light conditions hinder their practical application. As an initial attempt at low-light stereo image enhancement, we propose a novel Dual-View Enhancement Network (DVENet) based on the Retinex theory, which consists of two stages. The first stage estimates an illumination map to obtain a coarse enhancement result, which boosts the correlation of two views, while the second stage recovers details by integrating the information from two views to achieve fine image quality improvement with the guidance of the illumination map. To fully utilize the dual-view correlation, we further design a wavelet-based view transfer module to efficiently carry out multi-scale detail recovery. Then, we design an illumination-aware attention fusion module to exploit the complementarity between the fused features from two views and the single-view features. Experiments on both synthetic and real-world stereo datasets demonstrate the superiority of our proposed method over existing solutions. The code and model are publicly available at:https://github.com/KevinJ-Huang/Stereo-Low-Light. Jie Huang 0017, Xueyang Fu, Zeyu Xiao 0002, Feng Zhao 0004, Zhiwei Xiong |
IEEE Trans. Multim. | 3 |
| 2022 | Efficient Model-Driven Network for Shadow RemovalabstractDeep Convolutional Neural Networks (CNNs) based methods have achieved significant breakthroughs in the task of single image shadow removal. However, the performance of these methods remains limited for several reasons. First, the existing shadow illumination model ignores the spatially variant property of the shadow images, hindering their further performance. Second, most deep CNNs based methods directly estimate the shadow free results from the input shadow images like a black box, thus losing the desired interpretability. To address these issues, we first propose a new shadow illumination model for the shadow removal task. This new shadow illumination model ensures the identity mapping among unshaded regions, and adaptively performs fine grained spatial mapping between shadow regions and their references. Then, based on the shadow illumination model, we reformulate the shadow removal task as a variational optimization problem. To effectively solve the variational problem, we design an iterative algorithm and unfold it into a deep network, naturally increasing the interpretability of the deep model. Experiments show that our method could achieve SOTA performance with less than half parameters, one-fifth of floating-point of operations (FLOPs), and over seventeen times faster than SOTA method (DHAN). Yurui Zhu, Zeyu Xiao 0002, Yanchi Fang, Xueyang Fu, Zhiwei Xiong, Zhengjun Zha |
AAAI | 2 |
| 2022 | Propagating Difference Flows for Efficient Video Super-Resolution
Ruisheng Gao, Zeyu Xiao 0002, Zhiwei Xiong |
BMVC | 2 |
| 2022 | Frequency and Spatial Dual Guidance for Image Dehazing
Hu Yu 0001, Naishan Zheng, Man Zhou 0003, Jie Huang 0017, Zeyu Xiao 0002, Feng Zhao 0004 |
ECCV (19) | 5 |
| 2022 | Dast-Net: Depth-Aware Spatio-Temporal Network for Video DeblurringabstractVideo deblurring is a challenging task due to inevitable blurs caused by depth variation, object motion, and camera shake. Although several video deblurring methods resort to depth maps, they rarely produce visually appealing results since the information in the depth maps is used insufficiently. To address this issue, we propose a Depth-Aware Modulated (DAM) block for efficiently utilizing the depth map characteristics, in which the intensity and variation of depth are exploited according to the depth map value and edges. Based on the DAM block, we develop the Depth-Aware Spatio-Temporal Network (DAST-Net) tailored for video deblurring. Particularly, the Depth-Aware Temporal Alignment module uses the depth cues to guide the alignment of adjacent frames. The Depth-Modulated Spatial Fusion module then warps the aligned frames to maintain spatial invariance with the aligned features. The warped depth features are more effective in video deblurring, since they allow for the aggregation of multiple frames. Extensive quantitative and qualitative evaluations demonstrate that the proposed DAST-Net outperforms other state-of-the-art methods. Qi Zhu 0010, Zeyu Xiao 0002, Jie Huang 0017, Feng Zhao 0004 |
ICME | 2 |
| 2021 | Space-Time Distillation for Video Super-ResolutionabstractCompact video super-resolution (VSR) networks can be easily deployed on resource-limited devices, e.g., smartphones and wearable devices, but have considerable performance gaps compared with complicated VSR networks that require a large amount of computing resources. In this paper, we aim to improve the performance of compact VSR networks without changing their original architectures, through a knowledge distillation approach that transfers knowledge from a complicated VSR network to a compact one. Specifically, we propose a space-time distillation (STD) scheme to exploit both spatial and temporal knowledge in the VSR task. For space distillation, we extract spatial attention maps that hint the high-frequency video content from both networks, which are further used for transferring spatial modeling capabilities. For time distillation, we narrow the performance gap between compact models and complicated models by distilling the feature similarity of the temporal memory cells, which are encoded from the sequence of feature maps generated in the training clips using ConvLSTM. During the training process, STD can be easily incorporated into any network without changing the original network architecture. Experimental results on standard benchmarks demonstrate that, in resource-constrained situations, the proposed method notably improves the performance of existing VSR networks without increasing the inference time. Zeyu Xiao 0002, Xueyang Fu, Jie Huang 0017, Zhen Cheng 0002, Zhiwei Xiong |
CVPR | 1 |
| 2021 | Stereo Video Super-Resolution via Exploiting View-Temporal CorrelationsabstractStereo Video Super-Resolution (StereoVSR) aims to generate high-resolution video steams from two low-resolution videos under stereo settings. Existing video super-resolution and stereo image super-resolution techniques can be extended to tackle the StereoVSR task, yet they cannot make full use of the multi-view and temporal information to achieve satisfactory performance. In this paper, we propose a novel Stereo Video Super-Resolution Network (SVSRNet) to fulfill the StereoVSR task via exploiting view-temporal correlations. First, we devise a view-temporal attention module (VTAM) to integrate the information of cross-time-cross-view for constructing high-resolution stereo videos. Second, we propose a spatial-temporal fusion module (STFM), which aggregates the information across time in intra-view to emphasize important features for subsequent restoration. In addition, we design a view-temporal consistency loss function to enforce consistency constraint of superresolved stereo videos. Comprehensive experimental results demonstrate that our method generates superior results. Ruikang Xu, Zeyu Xiao 0002, Mingde Yao, Yueyi Zhang 0001, Zhiwei Xiong |
ACM Multimedia | 2 |
| 2021 | Unfolding Taylor's Approximations for Image RestorationabstractDeep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping networks without deepening into the rationality, and neglect the intrinsic prior knowledge of restoration task. To solve the above problems, inspired by Taylor’s Approximations, we unfold Taylor’s Formula to construct a novel framework for image restoration. We find the main part and the derivative part of Taylor’s Approximations take the same effect as the two competing goals of high-level contextualized information and spatial details of image restoration respectively. Specifically, our framework consists of two steps, which are correspondingly responsible for the mapping and derivative functions. The former first learns the high-level contextualized information and the later combines it with the degraded input to progressively recover local high-order spatial details. Our proposed framework is orthogonal to existing methods and thus can be easily integrated with them for further improvement, and extensive experiments demonstrate the effectiveness and scalability of our proposed framework. Man Zhou 0003, Xueyang Fu, Zeyu Xiao 0002, Aiping Liu, Zhiwei Xiong |
NeurIPS | 3 |
| 2020 | Space-Time Video Super-Resolution Using Temporal ProfilesabstractIn this paper, we propose a novel space-time video super-resolution method, which aims to recover a high-frame-rate and high-resolution video from its low-frame-rate and low-resolution observation. Existing solutions seldom consider the spatial-temporal correlation and the long-term temporal context simultaneously and thus are limited in the restoration performance. Inspired by the epipolar-plane image used in multi-view computer vision tasks, we first propose the concept of temporal-profile super-resolution to directly exploit the spatial-temporal correlation in the long-term temporal context. Then, we specifically design a feature shuffling module for spatial retargeting and spatial-temporal information fusion, which is followed by a refining module for artifacts alleviation and detail enhancement. Different from existing solutions, our method does not require any explicit or implicit motion estimation, making it lightweight and flexible to handle any number of input frames. Comprehensive experimental results demonstrate that our method not only generates superior space-time video super-resolution results but also retains competitive implementation efficiency. Zeyu Xiao 0002, Zhiwei Xiong, Xueyang Fu, Dong Liu 0002, Zhengjun Zha |
ACM Multimedia | 1 |