Wei Zhang 0161

dblp:10/4661-161 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0002-2358-0779ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo Vision
abstract
High spatio‑temporal resolution novel‑view scene rendering is crucial for applications such as sports analysis and scientific experiments. However, existing Dynamic Scene Rendering (DSR) approaches typically rely on conventional RGB cameras with limited frame rates, making it difficult to achieve high spatio‑temporal resolution. In this paper, we present BulletTime4D, a high spatio‑temporal resolution DSR framework, which is the first trial to integrate a spike camera with binocular RGB cameras for dynamic scene reconstruction. Specifically, we first develop a hybrid camera prototype and build a real‑world dynamic scene reconstruction dataset. Then, BulletTime4D presents a multi‑timescale deformation representation by combining low‑frequency spatio‑temporal features with high‑frequency inter‑frame motion features. Finally, a rendering network is designed capable of projecting 4D Gaussians into the spike domain for spike rendering, and a cross‑domain supervision strategy is proposed to achieve high‑frame‑rate texture and color rendering. The results show that BulletTime4D outperforms state‑of‑the‑art methods on both simulated and real‑world datasets. In addition, BulletTime4D can synthesize 300 FPS novel‑view renderings using stereo RGB cameras at 30 FPS and a single spike camera.
Yiqian Chang, Haoran Xu 0004, Qinghong Ye, Jianing Li 0001, Xuan Wang 0002, Wei Zhang 0161, Peixi Peng
AAAI6
2025 Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset
abstract
Object detection in event streams has emerged as a cutting-edge research area, demonstrating superior performance in low-light conditions, scenarios with motion blur, and rapid movements. Current detectors leverage spiking neural networks, Transformers, or convolutional neural networks as their core architectures, each with its own set of limitations including restricted performance, high computational overhead, or limited local receptive fields. This paper introduces a novel MoE (Mixture of Experts) heat conduction-based object detection algorithm that strikingly balances accuracy and computational efficiency. Initially, we employ a stem network for event data embedding, followed by processing through our innovative MoE-HCO blocks. Each block integrates various expert modules to mimic heat conduction within event streams. Subsequently, an IoU-based query selection module is utilized for efficient token extraction, which is then channeled into a detection head for the final object detection process. Furthermore, we are pleased to introduce EvDET200K, a novel benchmark dataset for event-based object detection. Captured with a high-definition Prophesee EVK4-HD event camera, this dataset encompasses 10 distinct categories, 200,000 bounding boxes, and 10,054 samples, each spanning 2 to 5 seconds. We also provide comprehensive results from over 15 state-of-the-art detectors, offering a solid foundation for future research and comparison. The source code has been released on: https://github.com/Event-AHU/OpenEvDET
Xiao Wang 0014, Wei Zhang 0161, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001
CVPR4
2025 Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion
Quanmin Liang, Shuai Liu 0009, Xinzi Cao, Jinyi Lu, Feidiao Yang, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001
ICCV7
2025 ESOD: Event-Based Small Object Detection
abstract
Event-based object detection plays a crucial role in scenarios involving high-speed motion, extreme lighting conditions, and high-frequency detection. However, existing methods fail to address the challenges posed by small objects, including discriminative feature deficiency, the loss of critical information, and the inherent sparsity of event data. Moreover, the lack of benchmark datasets has significantly hindered progress in this field. To tackle these issues, we propose the Fully Deformable Detection Network (FDDNet), a lightweight framework that dynamically adapts to extract key features. First, we introduce a Long-Term Deformable Temporal Receptive Module (LDTR), which aligns critical features across consecutive event streams and leverages a State Space Model for long-range temporal modeling, enhancing the detection of high-speed small objects. Second, to address the sparsity of event data and the concentration of key features along object edges, we design a Sparse Feature Aggregation Block (SFAB) within the backbone and a coarse-to-fine deformable detection head, enabling hierarchical feature refinement from local to global, and improving the detection quality of sparse targets. Finally, to mitigate the lack of event-based small object datasets, we develop a high-quality, annotation-free data acquisition method and collect a real-world benchmark dataset for validation. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on event-based small object detection tasks, with a mAP of 37.4% (+2.4%) on our benchmark and runs at 88 FPS, showcasing both accuracy and real-time capability. Our code and Supplement are available at https://github.com/Lqm26/ESOD.
Quanmin Liang, Jinyi Lu, Shuai Liu 0009, Yinzheng Zhao, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001
ACM Multimedia7
2025 DSF-Net: Dynamic Sparse Fusion of Event-RGB via Spike-Triggered Attention for High-Speed Detection
Dongyang Ma, Zhengyu Ma, Wei Zhang 0161, Yonghong Tian 0001
ACM Multimedia3
2025 Spike4DGS: Towards High-Speed Dynamic Scene Rendering with 4D Gaussian Splatting via a Spike Camera Array
abstract
Spike camera with high temporal resolution offers a new perspective on high-speed dynamic scene rendering. Most existing rendering methods rely on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for static scenes using a monocular spike camera. However, these methods struggle with dynamic motion, while a single camera suffers from limited spatial coverage, making it challenging to reconstruct fine details in high-speed scenes. To address these problems, we propose Spike4DGS, the first high-speed dynamic scene rendering framework with 4D Gaussian Splatting using spike camera arrays. Technically, we first build a multi-view spike camera array to validate our solution, then establish both synthetic and real-world multi-view spike-based reconstruction datasets. Then, we design a multi-view spike-based dense initialization module that obtains dense point clouds and camera poses from continuous spike streams. Finally, we propose a spike-pixel synergy constraint supervision to optimize Spike4DGS, incorporating both rendered image quality loss and dynamic spatiotemporal spike loss. The results show that our Spike4DGS outperforms state-of-the-art methods in terms of novel view rendering quality on both synthetic and real-world datasets. More details are available at https://github.com/Qinghongye/Spike4DGS.
Qinghong Ye, Yiqian Chang, Jianing Li 0001, Haoran Xu 0004, Xuan Wang 0002, Wei Zhang 0161, Yonghong Tian 0001, Peixi Peng
NeurIPS6
2025 Towards Ultra High-Speed Hyperspectral Imaging by Integrating Compressive and Neuromorphic Sampling
Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001
Int. J. Comput. Vis.4
2025 High-Rate Monocular Depth Estimation via Cross Frame-Rate Collaboration of Frames and Events
Xu Liu 0006, Xiaopeng Fan 0001, Jianing Li 0001, Dianze Li, Wei Zhang 0161, Zhengyu Ma, Yonghong Tian 0001
Int. J. Comput. Vis.5
2025 Event-Enhanced Snapshot Mosaic Hyperspectral Frame Deblurring
abstract
Snapshot Mosaic Hyperspectral Cameras (SMHCs) are popular hyperspectral imaging devices for acquiring both color and motion details of scenes. However, the narrow-band spectral filters in SMHCs may negatively impact their motion perception ability, resulting in blurry SMHC frames. In this paper, we propose a hardware-software collaborative approach to address the blurring issue of SMHCs. Our approach involves integrating SMHCs with neuromorphic event cameras for efficient event-enhanced SMHC frame deblurring. To achieve spectral information recovery guided by event signals, we formulate a spectral-aware Event-based Double Integral (sEDI) model that links SMHC frames and events from a spectral perspective, providing principled model design insights. Then, we develop a Diffusion-guided Noise Awareness (DNA) training framework that utilizes diffusion models to learn noise-aware features and promote model robustness towards camera noise. Furthermore, we design an Event-enhanced Hyperspectral frame Deblurring Network (EvHDNet) based on sEDI, which is trained with DNA and features improved spatial-spectral learning and modality interaction for reliable SMHC frame deblurring. Experiments on both synthetic data and real data show that the proposed DNA + EvHDNet outperforms state-of-the-art methods on both spatial and spectral fidelity. The code and dataset will be made publicly available.
Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition
abstract
Current event stream-based pattern recognition models typically present the event stream as the point cloud, voxel, image, and the like, and formulate multiple deep neural networks to acquire their features. Although considerable results can be achieved in simple cases, however, the performance of the model might be restricted by monotonous modality expressions, sub-optimal fusion, and readout mechanisms. In this article, we put forward a novel dual-stream framework for event stream-based pattern recognition through differentiated fusion, which is called EFV++. It models two common event representations simultaneously, i.e., event images and event voxels. The spatial and three-dimensional stereo information can be separately learned by making use of Transformer and Graph Neural Network (GNN). We believe the features of each representation still contain both efficient and redundant features and a sub-optimal solution may be obtained if we directly fuse them without differentiation. Thus, we divide each feature into three levels and retain high-quality features, blend medium-quality features, and exchange low-quality features. The enhanced dual features will be provided to the fusion Transformer together with bottleneck features. In addition, we introduce a novel hybrid interaction readout mechanism to enhance the diversity of features as final representations. Comprehensive experiments validate that the framework we have proposed attains cutting-edge performance on a variety of extensively utilized event stream-based classification datasets. Particularly, we have realized a freshly pioneering performance on the Bullying10 k dataset, precisely 90.51%, and this outpaces the runner-up by$+2.21\%$.
Lan Chen 0003, Xiao Wang 0014, Pengpeng Shao, Wei Zhang 0161, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001
IEEE Trans. Multim.5
2024 Event-Based Visible and Infrared Fusion via Multi-Task Collaboration
abstract
Visible and Infrared image Fusion (VIF) offers a comprehensive scene description by combining thermal infrared images with the rich textures from visible cameras. However, conventional VIF systems may capture over/under exposure or blurry images in extreme lighting and high dynamic motion scenarios, leading to degraded fusion results. To address these problems, we propose a novel Event-based Visible and Infrared Fusion (EVIF) system that employs a visible event camera as an alternative to traditional frame-based cameras for the VIF task. With extremely low latency and high dynamic range, event cameras can effectively address blurriness and are robust against diverse luminous ranges. To produce high-quality fused images, we develop a multitask collaborative framework that simultaneously performs event-based visible texture reconstruction, event-guided infrared image deblurring, and visible-infrared fusion. Rather than independently learning these tasks, our framework capitalizes on their synergy, leveraging cross-task event enhancement for efficient deblurring and bi-level min-max mutual information optimization to achieve higher fusion quality. Experiments on both synthetic and real data show that EVIF achieves remarkable performance in dealing with extreme lighting conditions and high-dynamic scenes, ensuring high-quality fused images across a broad range of practical scenarios.
Mengyue Geng, Lin Zhu 0012, Lizhi Wang 0001, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001
CVPR4
2023 Content-Aware Subspace Low-Rank Tensor Recovery for Hyperspectral Image Restoration
abstract
The low-rank tensor model has made great progress for hyperspectral image (HSI) restoration. Recently, the low-rank tensor methods have further been boosted with subspace learning by transforming original HSI into a low-dimensional subspace with reduced computational burden and discriminative feature representation. However, existing subspace-based methods consistently employ a fixed subspace dimension for all patches, which may violate the intrinsic dimension discrepancy of different image content, leading to information loss or redundancy. In this work, our key observation is that the intrinsic subspace of different image patches along different dimensions is different, which should be adaptively modelled for compact feature extraction. Therefore, we propose a content-aware subspace low-rank tensor recovery (CSLRTR) methods by leveraging both deep network and low-rank tensor model. Specifically, we first analyze the intrinsic discrepancy of different HSI patches among both spatial and spectral dimensions, and design a simple network to adaptively learn the optimal subspace dimension. The adaptive subspace learning and low-rank tensor recovery are iteratively performed and mutually promote each other. On one hand, the learned subspace would contribute to more compact low-rank representation for better restoration; on the other hand, the low rank tensor recovery with less degradations would definitely ease the difficulty of the subspace estimation. Note that, the adaptive content-aware subspace strategy has been simultaneously employed on both spectral and nonlocal dimensions, where the spectral-spatial relationship has been further strengthened with better restoration. We have performed extensive experiments on different datasets and restoration tasks, and extended the content-aware subspace strategy to previous methods.
Xueyao Xiao, Wei Zhang 0161, Yi Chang 0002, Shuning Cao, Wei He 0003, Houzhang Fang, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.2
2022 Robust Blind Deblurring Under Stripe Noise for Remote Sensing Images
abstract
The blind image deblurring methods have achieved great progress for Gaussian random noise. Few works have paid attention to the image deblurring under the structural noise, which is a very common degradation in multi-detectors imaging systems. This paper considers the practical yet challenging problem of blind deblurring in presence of the line-pattern stripe noise for remote sensing images. To overcome this issue, we explicitly formulate the structural noise into a novel and robust blind image deblurring framework. We observe that the structural line-pattern stripe noise would deteriorate both the kernel estimation and non-blind deblurring, and propose a three-stage restoration framework to progressively estimate the blur kernel and clean image. Specifically, we first estimate an intermediate blur kernel by getting rid of the negative influence of the stripe noise in the unidirectional gradient domain. Next, a learning-based kernel refinement network is introduced to rectify the missing details of the inaccurate kernel. Finally, a low-rank decomposition-based non-blind deblurring model is proposed to simultaneously estimate the clean image and stripe noise. Experimental results on real and synthetic datasets demonstrate that the proposed RBDS method outperforms the state-of-the-art blind deblurring methods.
Shuning Cao, Houzhang Fang, Wei Zhang 0161, Yi Chang 0002, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.4
2022 A Distributed Algorithm for Task Offloading in Vehicular Networks With Hybrid Fog/Cloud Computing
abstract
Fog computing has been an effective paradigm of real-time applications in the IoT area, which enables task offloading at network edge devices. Particularly, many emerging vehicular applications require real-time interaction between the terminal users and computation servers, which can be implemented in fog-based architecture. However, it is still challenging to apply fog computing in vehicular networks due to high mobility of vehicles and uneven distribution of vehicle density, which may result in performance degradation, such as unbalanced workload and unexpected task failure. In this article, we investigate a new service scenario of task offloading under a three-layer service architecture, where the resources of vehicular fog (VF), fog server (FS), and central cloud (CC) are utilized in a cooperative way. On this basis, we formulate the probabilistic task offloading (PTO) problem by synthesizing task transmission, computation, and result retrieval, as well as characterizing the heterogeneity of computation servers. The objective of the PTO is to minimize the weighted sum of execution delay, energy consumption, and payment cost. To resolve the PTO problem, we propose a comprehensive task offloading algorithm by combining the alternating direction method of multipliers (ADMMs) and particle swarm optimization (PSO), called ADMM-PSO. The basic idea of the ADMM-PSO is to divide the PTO problem into multiple unconstrained subproblems and achieve the optimal solution in the form of an iterative coordination process. For each iteration, the solution is achieved by solving each subproblem with the PSO and updated based on a designed rule, which is able to converge to the optimal solution when the stop criterion is satisfied. Finally, we build the simulation model and implement the proposed algorithm for performance evaluation. The simulation results demonstrate the superiority of the proposed algorithm under a wide range of service scenarios.
Zongkai Liu, Penglin Dai, Huanlai Xing, Zhaofei Yu, Wei Zhang 0161
IEEE Trans. Syst. Man Cybern. Syst.5