VLDB 2026 Research / reviewers in the wild / expert
Wenming Weng
dblp:294/0803
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-1042-8903ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E2SL: Efficient Depth Sensing from Event-Based Structured LightabstractStructured light (SL) is a popular approach for 3D reconstruction. Most SL techniques rely on frame-based cameras and are often not robust in high-speed dynamic scenes. Recently, event cameras have sparked growing interest in high-speed SL imaging, due to their high temporal resolution. The event-based SL enjoys the high-speed data acquisition, however, most existing methods tend to pursue the reconstruction accuracy but sacrificing the computational efficiency, limiting the applicability in real-world scenarios. To this end, we propose E2SL, an Efficient deep network tailored for monocular Event-based SL. Specifically, E2SL comprises three key components: binary embedding lookup table (BE-LUT), spatial context enhancement (SCE), and geometric-prior regression (GPR). Given the input event frame, BE-LUT, which is precomputed and stored, first retrieves the features efficiently. Then, SCE extends the receptive field of the features and captures the spatial context. Finally, GPR conducts the geometric-prior-based tree classification for fast and robust depth estimation. To support training and evaluation, we contribute an event-based SL simulator, which generates a large-scale and diverse synthetic dataset. Besides, we develop an event-based SL prototype and collect a dataset with accurate ground truth for real-world evaluation. Extensive experiments demonstrate that our method achieves state-of-the-art accuracy while maintaining a per-frame reconstruction time of 7.7 ms, meeting the demands of high-speed depth sensing. The code and dataset are available on the project page https://dongxin000.github.io/E2SL/. Jiacheng Fu, Wenming Weng, Yueyi Zhang 0001, Bingyao Huang, Zhiwei Xiong |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | S2D-LFE: Sparse-to-Dense Light Field Event GenerationabstractIn this paper, we present S2D-LFE, an innovative approach for sparse-to-dense light field event generation. For the first time to our knowledge, S2D-LFE enables controllable novel view synthesis only from sparse-view light field event (LFE) data, and addresses three critical challenges for the LFE generation task: simplicity, controllability, and consistency. The simplicity aspect eliminates the dependency on frame-based modality, which often suffers from motion blur and low frame-rate limitations. The controllability aspect enables precise view synthesis under sparse LFE conditions with view-related constraints. The consistency aspect ensures both cross-view and temporal coherence in the generated results. To realize S2D-LFE, we develop a novel diffusion-based generation network with two key components. First, we design an LFE-customized variational auto-encoder that effectively compresses and reconstructs LFE by integrating cross-view information. Second, we design an LFE-aware injection adaptor to extract comprehensive geometric and texture priors. Furthermore, we construct a large-scale synthetic LFE dataset containing 162 one-minute sequences using simulator, and capture a real-world testset using our custom-built sparse LFE acquisition system, covering diverse indoor and outdoor scenes. Extensive experiments demonstrate that S2D-LFE successfully generates up to 9 × 9 dense LFE from 2 × 2 sparse inputs and outperforms existing methods on both synthetic and real-world data. The datasets and code are available at https://github.com/Yutong2022/S2D-LFE. Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
CVPR | 2 |
| 2025 | GenFlow3D: Generative Scene Flow Estimation and Prediction on Point Cloud Sequences
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
ICCV | 2 |
| 2025 | Event-Boosted Deformable 3D Gaussians for Dynamic Scene Reconstruction
Wenming Weng, Yueyi Zhang 0001, Ruikang Xu, Zhiwei Xiong |
ICCV | 2 |
| 2025 | Event-based HDR Structured LightabstractEvent-based structured light (SL) systems have attracted increasing attention for their potential in high-performance 3D measurement. Despite the inherent HDR capability of event cameras, reflective and absorptive surfaces still cause event cluttering and absence, which produce overexposed and underexposed regions that degrade the reconstruction quality. In this work, we present the first HDR 3D measurement framework specifically designed for event-based SL systems. First, we introduce a multi-contrast HDR coding strategy that facilitates imaging of areas with different reflectance. Second, to alleviate inter-frame interference caused by overexposed and underexposed areas, we propose a universal confidence-driven stereo matching strategy. Specifically, we estimate a confidence map as the fusion weight for features via an energy-guided confidence estimation. Further, we propose the confidence propagation volume, an innovative cost volume that offers both effective suppression of inter-frame interference and strong representation capability. Third, we contribute an event-based SL simulator and propose the first event-based HDR SL dataset. We also collect a real-world benchmarking dataset with ground truth. We validate the effectiveness of our method with the proposed confidence-driven strategy on both synthetic and real-world datasets. Experimental results demonstrate that our proposed HDR framework enables accurate 3D measurement even under extreme conditions. Jiacheng Fu, Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
NeurIPS | 4 |
| 2025 | BVSR-EvD: Blurry Video Space-Time Super-Resolution With Events via Diffusion ModelsabstractVideo restoration from low-resolution and low-frame-rate blurry sources remains challenging due to insufficient data priors. In this paper, we propose BVSR-EvD, leveraging event cameras and diffusion models to boost blurry video space-time super-resolution. Specifically, we identify three distinct data priors from event-video dual modalities: motion prior from events, content prior from videos, and physical prior from their integration, contributing to temporal stability, content preservation, and detail enhancement respectively. To effectively utilize these data priors, BVSR-EvD creates the Trident Diffusion Model (Trident-DM), which decomposes each denoising step into trident decoupling and adaptive self-composition stages. The former employs single-modal and dual-modal meta-networks to extract the three unique data priors, while the latter dynamically integrates them through learned prior-aware weight maps. BVSR-EvD achieves up to $\times 8$ spatial super-resolution and $\times 64$ temporal super-resolution from blurry videos, surpassing existing methods on public video datasets. Wenming Weng, Yueyi Zhang 0001, Zeyu Xiao 0002, Zhiwei Xiong |
IEEE Trans. Image Process. | 1 |
| 2025 | EGVD: Event-Guided Video DerainingabstractRecent research has explored leveraging event cameras, known for their prowess in capturing scenes with nonuniform motion, for video deraining, leading to performance improvements. However, the existing event-based method still faces the challenge that the complex spatiotemporal distribution disrupts temporal information fusion and complicates feature separation. This article proposes a novel end-to-end learning framework for video deraining that effectively extracts the rich dynamic information provided by the event stream. Our framework incorporates two key modules: an event-aware motion detection (EAMD) module that adaptively aggregates multiframe motion information using event-driven masks and a pyramidal adaptive selection module that separates background and rain layers by leveraging contextual priors from both event and conventional camera data. To facilitate efficient training, we introduce a real-world dataset of synchronized rainy videos and event streams. Extensive evaluations on both synthetic and real-world datasets demonstrate the superiority of our proposed method compared to state-of-the-art approaches. The code is available at https://github.com/booker-max/EGVD. Yueyi Zhang 0001, Jin Wang 0023, Wenming Weng, Xiaoyan Sun 0001, Zhiwei Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | CCEdit: Creative and Controllable Video Editing via Diffusion ModelsabstractIn this paper, we present CCEdit, a versatile generative video editing framework based on diffusion models. Our approach employs a novel trident network structure that separates structure and appearance control, ensuring precise and creative editing capabilities. Utilizing the foundational ControlNet architecture, we maintain the structural integrity of the video during editing. The incorporation of an additional appearance branch enables users to exert fine-grained control over the edited key frame. These two side branches seamlessly integrate into the main branch, which is constructed upon existing text-to-image (T2I) generation models, through learnable temporal layers. The versatility of our framework is demonstrated through a diverse range of choices in both structure representations and personalized T2I models, as well as the option to provide the edited key frame. To facilitate comprehensive evaluation, we introduce the BalanceCC benchmark dataset, comprising 100 videos and 4 target prompts for each video. Our extensive user studies compare CCEdit with eight state-of-the-art video editing methods. The outcomes demonstrate CCEdit's substantial superiority over all other methods. Ruoyu Feng 0001, Wenming Weng, Yuhui Yuan, Jianmin Bao, Chong Luo 0001, Zhibo Chen 0001, Baining Guo |
CVPR | 2 |
| 2024 | Event-Assisted Low-Light Video Object SegmentationabstractIn the realm of video object segmentation (VOS), the challenge of operating under low-light conditions persists, resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras, characterized by their high dynamic range and ability to capture motion information of objects, offer promise in enhancing object visibility and aiding VOS methods under such low-light conditions. This paper introduces a pioneering framework tai-lored for low-light VOS, leveraging event camera data to elevate segmentation accuracy. Our approach hinges on two pivotal components: the Adaptive Cross-Modal Fusion (ACMF) module, aimed at extracting pertinent features while fusing image and event modalities to mitigate noise interference, and the Event-Guided Memory Matching (EGMM) module, designed to rectify the issue of in-accurate matching prevalent in low-light settings. Additionally, we present the creation of a synthetic LLE-DAVIS dataset and the curation of a real-world LLE-vas dataset, encompassing frames and events. Experimental evaluations corroborate the efficacy of our method across both datasets, affirming its effectiveness in low-light scenarios. The datasets are available at https://github.com/HebeiFast/EventLowLightVOS. Hebei Li, Jin Wang 0023, Jiahui Yuan, Wenming Weng, Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001 |
CVPR | 5 |
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 3 |
| 2024 | Depth from Asymmetric Frame-Event Stereo: A Divide-and-Conquer ApproachabstractEvent cameras asynchronously measure brightness changes in a scene without motion blur or saturation, while frame cameras capture images with dense intensity and fine details at a fixed rate. The exclusive advantages of the two modalities make depth estimation from Stereo Asymmetric Frame-Event (SAFE) systems appealing. However, due to the inevitable information absence of one modality in certain challenging regions, existing stereo matching methods lose efficacy for asymmetric inputs from SAFE systems. In this paper, we propose a divide-and-conquer approach that decomposes depth estimation from SAFE systems into three sub-tasks, i.e., frame-event stereo matching, frame-based Structure-from-Motion (SfM), and event-based SfM. In this way, the above challenging regions are addressed by monocular SfM, which estimates robust depth with two views belonging to the same functioning modality. Moreover, we propose a dual sampling strategy to construct cost volumes with identical spatial locations and depth hypotheses for different sub-tasks, which enables sub-task fusion at the cost volume level. To tackle the occlusion issue raised by the sampling strategy, we further introduce a temporal fusion scheme to utilize long-term sequential inputs with multi-view information. Experimental results validate the superior performance of our method over existing solutions. Xihao Chen, Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
WACV | 2 |
| 2023 | Event-based Blurry Frame Interpolation under Blind ExposureabstractRestoring sharp high frame-rate videos from low frame-rate blurry videos is a challenging problem. Existing blurry frame interpolation methods assume a predefined and known exposure time, which suffer from severe performance drop when applied to videos captured in the wild. In this paper, we study the problem of blurry frame interpolation under blind exposure with the assistance of an event camera. The high temporal resolution of the event camera is beneficial to obtain the exposure prior that is lost during the imaging process. Besides, sharp frames can be restored using event streams and blurry frames relying on the mutual constraint among them. Therefore, we first propose an exposure estimation strategy guided by event streams to estimate the lost exposure prior, transforming the blind exposure problem well-posed. Second, we propose to model the mutual constraint with a temporal-exposure control strategy through iterative residual learning. Our blurry frame interpolation method achieves a distinct performance boost over existing methods on both synthetic and self-collected real- world datasets under blind exposure. Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
CVPR | 1 |
| 2023 | Unsupervised Video Deraining with An Event CameraabstractCurrent unsupervised video deraining methods are inefficient in modeling the intricate spatio-temporal properties of rain, which leads to unsatisfactory results. In this paper, we propose a novel approach by integrating a bio-inspired event camera into the unsupervised video deraining pipeline, which enables us to capture high temporal resolution information and model complex rain characteristics. Specifically, we first design an end-to-end learning-based network consisting of two modules, the asymmetric separation module and the cross-modal fusion module. The two modules are responsible for segregating the features of the rain-background layer, and for positive enhancement and negative suppression from a cross-modal perspective, respectively. Second, to regularize the network training, we elaborately design a cross-modal contrastive learning method that leverages the complementary information from event cameras, exploring the mutual exclusion and similarity of rain-background layers in different domains. This encourages the deraining network to focus on the distinctive characteristics of each layer and learn a more discriminative representation. Moreover, we construct the first real-world dataset comprising rainy videos and events using a hybrid imaging system. Extensive experiments demonstrate the superior performance of our method on both synthetic and real-world datasets. Jin Wang 0023, Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
ICCV | 2 |
| 2022 | Boosting Event Stream Super-Resolution with a Recurrent Neural Network
Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
ECCV (6) | 1 |
| 2021 | Training Spiking Neural Networks with Accumulated Spiking FlowabstractThe fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demonstrate the potential of SNNs in supervised learning by introducing approximated backpropagation (BP) methods. To deal with the non-differentiable spike function in SNNs, these BP methods utilize information from the spatio-temporal domain to adjust the model parameters. With the increasing of time window and network size, the computational complexity of spatio-temporal backpropagation augments dramatically. In this paper, we propose a new backpropagation method for SNNs based on the accumulated spiking flow (ASF), i.e. ASF-BP. In the proposed ASF-BP method, updating parameters does not rely on the spike train of spiking neurons but leverage accumulated inputs and outputs of spiking neurons over the time window, which reduces the BP complexity significantly. We further present an adaptive linear estimation model to approach the dynamic characteristics of spiking neurons statistically. Experimental results demonstrate that with our proposed ASF-BP method, light-weight convolutional SNNs achieve superior performances compared with other spike-based BP methods on both non-neuromorphic (MNIST, CIFAR10) and neuromorphic (CIFAR10-DVS) datasets. The code is available at https://github.com/neural-lab/ASF-BP. Hao Wu 0042, Yueyi Zhang 0001, Wenming Weng, Yongting Zhang, Zhiwei Xiong, Zhengjun Zha, Xiaoyan Sun 0001, Feng Wu 0001 |
AAAI | 3 |
| 2021 | Event-based Video Reconstruction Using TransformerabstractEvent cameras, which output events by detecting spatio- temporal brightness changes, bring a novel paradigm to image sensors with high dynamic range and low latency. Previous works have achieved impressive performances on event-based video reconstruction by introducing convolutional neural networks (CNNs). However, intrinsic locality of convolutional operations is not capable of modeling long-range dependency, which is crucial to many vision tasks. In this paper, we present a hybrid CNN- Transformer network for event-based video reconstruction (ET-Net), which merits the fine local information from CNN and global contexts from Transformer In addition, we further propose a Token Pyramid Aggregation strategy to implement multi-scale token integration for relating internal and intersected semantic concepts in the token-space. Experimental results demonstrate that our proposed method achieves superior performance over state-of-the-art methods on multiple real-world event datasets. The code is available at https://github.com/WarranWeng/ET-Net. Wenming Weng, Yueyi Zhang 0001, Zhiwei Xiong |
ICCV | 1 |