EDBT 2026 Demo / reviewers in the wild / expert
Lin Zhu 0012
dblp:z/LinZhu12
· DBLP profile ↗
61ranked-venue papers
18as first author
57since 2021 · last 2026
0000-0001-6487-0441ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 11 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 13 first-author · 35 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms
Xiao Wang 0014, Shiao Wang, Pengpeng Shao, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | ESTR-CoT: Towards explainable and accurate event stream based scene text recognition with chain-of-thought reasoning
Xiao Wang 0014, Jingtao Jiang, Qiang Chen 0007, Lan Chen 0003, Lin Zhu 0012, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001 |
Neurocomputing | 5 |
| 2026 | Learning Physics-Informed Noise Models from Dark Frames for Low-Light Raw Image DenoisingabstractRecently, the mainstream practice for training low-light raw image denoising methods has shifted towards employing synthetic data. Noise modeling, which focuses on characterizing the noise distribution of real-world sensors, profoundly influences the effectiveness and practicality of synthetic data. Currently, physics-based noise modeling struggles to characterize the entire real noise distribution, while learning-based noise modeling impractically depends on paired real data. In this paper, we propose a novel strategy: learning the noise model from dark frames instead of paired real data, to break down the data dependency. Based on this strategy, we introduce an efficient physics-informed noise neural proxy (PNNP) to approximate the real-world sensor noise model. Specifically, we integrate physical priors into neural proxies and introduce three efficient techniques: physics-guided noise decoupling (PND), physics-aware proxy model (PPM), and differentiable distribution loss (DDL). PND decouples the dark frame into different components and handles different levels of noise flexibly, which reduces the complexity of noise modeling. PPM incorporates physical priors to constrain the synthetic noise, which promotes the accuracy of noise modeling. DDL provides explicit and reliable supervision for noise distribution, which promotes the precision of noise modeling. PNNP exhibits powerful potential in characterizing the real noise distribution. Extensive experiments on public datasets demonstrate superior performance in practical low-light raw image denoising. The source code will be publicly available at the https://fenghansen.github.io/publication/PNNP. Hansen Feng, Lizhi Wang 0001, Yiqi Huang, Yuzhi Wang, Lin Zhu 0012, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | DSNeRF: Dynamic View Synthesis for Ultra-Fast Scenes From Continuous Spike StreamsabstractSpike cameras generate binary spikes in response to light intensity changes, enabling high-speed visual perception with unprecedented temporal resolution. However, the unique characteristics of spike stream present significant challenges for reconstructing dense 3D scene representations, particularly in dynamic environments and under non-ideal lighting conditions. In this paper, we introduce DSNeRF, the first method to derive a NeRF-based volumetric scene representation from spike camera data. Our approach leverages NeRF's multi-view consistency to establish robust self-supervision, effectively eliminating erroneous measurements and uncovering coherent structures within exceedingly noisy input amidst diverse real-world illumination scenarios. We propose a novel mapping from pixel rays to the spike domain, integrating the spike generation process directly into NeRF training. Specifically, DSNeRF introduces an integrate-and-fire neuron layer that models non-idealities to capture intrinsic camera noise, including both random and fixed-pattern spike noise, thereby enhancing scene fidelity. Additionally, we propose a motion-guided spiking neuron layer and a long-term rendering photometric loss to better align dynamic spike streams, ensuring accurate scene geometry. Our method optimizes neural radiance fields to render photorealistic novel views from continuous spike streams, demonstrating advantages over other vision sensors in certain scenes. Empirical evaluations on both real and simulated sequences validate the effectiveness of our approach. Lin Zhu 0012, Kangmin Jia, Yifan Zhao 0002, Yunshan Qi, Lizhi Wang 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Revisiting color-event based tracking: A unified network, dataset, and metric
Chuanming Tang, Xiao Wang 0014, Ju Huang, Bo Jiang 0002, Lin Zhu 0012, Shifeng Chen, Jianlin Zhang 0001, Yaowei Wang 0001, Yonghong Tian 0001 |
Pattern Recognit. | 5 |
| 2026 | MambaEVT: Event Stream-Based Visual Object Tracking Using State Space ModelabstractEvent camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high dynamic range, and dense temporal resolution. Current event-based tracking algorithms are gradually hitting their performance bottlenecks, due to the utilization of vision Transformer and the static template for target object localization. In this paper, we propose a novel Mamba-based visual tracking framework that adopts the state space model with linear complexity as a backbone network. The search regions and target template are fed into the vision Mamba network for simultaneous feature extraction and interaction. The output tokens of search regions will be fed into the tracking head for target localization. More importantly, we consider introducing a dynamic template update strategy into the tracking framework using the Memory Mamba network. By considering the diversity of samples in the target template library and making appropriate adjustments to the template memory module, a more effective dynamic template can be integrated. The effective combination of dynamic and static templates allows our Mamba-based tracking algorithm to achieve a good balance between accuracy and computational cost on multiple large-scale datasets, including EventVOT, VisEvent, and FE240hz. The source code and checkpoint have been released on https://github.com/Event-AHU/MambaEVT. Xiao Wang 0014, Shiao Wang, Xixi Wang 0005, Zhicheng Zhao 0002, Lin Zhu 0012, Bo Jiang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Positive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image DenoisingabstractImage denoising enhances image quality, serving as a foundational technique across various computational photography applications. The obstacle to clean image acquisition in real scenarios necessitates the development of self-supervised image denoising methods only depending on noisy images, especially a single noisy image. Existing self-supervised image denoising paradigms (Noise2Noise and Noise2Void) rely heavily on information-lossy operations, such as downsampling and masking, culminating in low-quality denoising performance. In this paper, we propose a novel self-supervised single image denoising paradigm, Positive2Negative, to break the information-lossy barrier. Our paradigm involves two key steps: Renoised Data Construction (RDC) and Denoised Consistency Supervision (DCS). RDC renoises the predicted denoised image by the predicted noise to construct multiple noisy images, preserving all the information of the original image. DCS ensures consistency across the multiple denoised images, supervising the network to learn robust denoising. Our Positive2Negative paradigm achieves state-of-the-art performance in self-supervised single image denoising with significant speed improvements. The code is released to the public at https://github.com/Li-Tong-621/P2N. Tong Li 0016, Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 4 |
| 2025 | Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark DatasetabstractObject detection in event streams has emerged as a cutting-edge research area, demonstrating superior performance in low-light conditions, scenarios with motion blur, and rapid movements. Current detectors leverage spiking neural networks, Transformers, or convolutional neural networks as their core architectures, each with its own set of limitations including restricted performance, high computational overhead, or limited local receptive fields. This paper introduces a novel MoE (Mixture of Experts) heat conduction-based object detection algorithm that strikingly balances accuracy and computational efficiency. Initially, we employ a stem network for event data embedding, followed by processing through our innovative MoE-HCO blocks. Each block integrates various expert modules to mimic heat conduction within event streams. Subsequently, an IoU-based query selection module is utilized for efficient token extraction, which is then channeled into a detection head for the final object detection process. Furthermore, we are pleased to introduce EvDET200K, a novel benchmark dataset for event-based object detection. Captured with a high-definition Prophesee EVK4-HD event camera, this dataset encompasses 10 distinct categories, 200,000 bounding boxes, and 10,054 samples, each spanning 2 to 5 seconds. We also provide comprehensive results from over 15 state-of-the-art detectors, offering a solid foundation for future research and comparison. The source code has been released on: https://github.com/Event-AHU/OpenEvDET Xiao Wang 0014, Wei Zhang 0161, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001 |
CVPR | 5 |
| 2025 | Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image DenoisingabstractExisting single-image denoising algorithms often struggle to restore details when dealing with complex noisy images. The introduction of near-infrared (NIR) images offers new possibilities for RGB image denoising. However, due to the inconsistency between NIR and RGB images, the existing works still struggle to balance the contributions of two fields in the process of image fusion. In response to this, in this paper, we develop a cross-field Frequency Correlation Exploiting Network (FCENet) for NIR-assisted image denoising. We first propose the frequency correlation prior based on an in-depth statistical frequency analysis of NIR-RGB image pairs. The prior reveals the complementary correlation of NIR and RGB images in the frequency domain. Leveraging frequency correlation prior, we then establish a frequency learning framework composed of Frequency Dynamic Selection Mechanism (FDSM) and Frequency Exhaustive Fusion Mechanism (FEFM). FDSM dynamically selects complementary information from NIR and RGB images in the frequency domain, and FEFM strengthens the control of common and differential features during the fusion process of NIR and RGB features. Extensive experiments on simulated and real data validate that the proposed method outperforms other state-of-the-art methods. The code will be released at https://github.com/yuchenwang815/FCENet. Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 5 |
| 2025 | Noise-Modeled Diffusion Models for Low-Light Spike Image Restoration
Lin Zhu 0012, Xijie Xiang, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 2 |
| 2025 | EMatch: A Unified Framework for Event-Based Optical Flow and Stereo MatchingabstractEvent cameras have shown promise in vision applications like optical flow estimation and stereo matching, with many specialized architectures leveraging the asynchronous and sparse nature of event data. However, existing works only focus event data within the confines of task-specific domains, overlooking how tasks across the temporal and spatial domains can reinforce each other. In this paper, we reformulate event-based flow estimation and stereo matching as a unified dense correspondence matching problem, enabling us to solve both tasks within a single model by directly matching features in a shared representation space. Specifically, our method utilizes a Temporal Recurrent Network to aggregate event features across temporal or spatial domains, and a Spatial Contextual Attention to enhance knowledge transfer across event flows via temporal or spatial interactions. By utilizing a shared feature similarities module that integrates knowledge from event streams via temporal or spatial interactions, our network performs optical flow estimation from temporal event segment inputs and stereo matching from spatial event segment inputs simultaneously. We demonstrate that our unified model inherently supports multi-task fusion and cross-task transfer. Without the need for retraining for specific task, our model can effectively handle both optical flow and stereo estimation, achieving state-of-the-art performance on both tasks. Pengjie Zhang, Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 2 |
| 2025 | EvFocus: Learning to Reconstruct Sharp Images from Out-of-Focus Event StreamsabstractEvent cameras are innovative sensors that capture brightness changes as asynchronous events rather than traditional intensity frames. These cameras offer substantial advantages over conventional cameras, including high temporal resolution, high dynamic range, and the elimination of motion blur. However, defocus blur, a common image quality degradation resulting from out-of-focus lenses, complicates the challenge of event-based imaging. Due to the unique imaging mechanism of event cameras, existing focusing algorithms struggle to operate efficiently on sparse event data. In this work, we propose EvFocus, a novel architecture designed to reconstruct sharp images from defocus event streams for the first time. Our work includes the development of an event-based out-of-focus camera model and a simulator to generate realistic defocus event streams for robust training and testing. EvDefous integrates a temporal information encoder, a blur-aware two-branch decoder, and a reconstruction and re-defocus module to effectively learn and correct defocus blur. Extensive experiments on both simulated and real-world datasets demonstrate that EvFocus outperforms existing methods across varying lighting conditions and blur sizes, proving its robustness and practical applicability in event-based defocus imaging. Lin Zhu 0012, Xiantao Ma, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICML | 1 |
| 2025 | Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events
Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 1 |
| 2025 | Re-coding for Uncertainties: Edge-awareness Semantic Concordance for Resilient Event-RGB SegmentationabstractSemantic segmentation has achieved great success in ideal conditions. However, when facing extreme conditions (e.g., insufficient light, fierce camera motion), most existing methods suffer from significant information loss of RGB, severely damaging segmentation results. Several researches exploit the high-speed and high-dynamic event modality as a complement, but event and RGB are naturally heterogeneous, which leads to feature-level mismatch and inferior optimization of existing multi-modality methods. Different from these researches, we delve into the edge secret of both modalities for resilient fusion and propose a novel Edge-awareness Semantic Concordance framework to unify the multi-modality heterogeneous features with latent edge cues. In this framework, we first propose Edge-awareness Latent Re-coding, which obtains uncertainty indicators while realigning event-RGB features into unified semantic space guided by re-coded distribution, and transfers event-RGB distributions into re-coded features by utilizing a pre-established edge dictionary as clues. We then propose Re-coded Consolidation and Uncertainty Optimization, which utilize re-coded edge features and uncertainty indicators to solve the heterogeneous event-RGB fusion issues under extreme conditions. We establish two synthetic and one real-world event-RGB semantic segmentation datasets for extreme scenario comparisons. Experimental results show that our method outperforms the state-of-the-art by a 2.55% mIoU on our proposed DERS-XS, and possesses superior resilience under spatial occlusion. Our code and datasets are publicly available at https://github.com/iCVTEAM/ESC. Yifan Zhao 0002, Lin Zhu 0012, Jia Li 0003 |
NeurIPS | 3 |
| 2025 | Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionabstractEvent cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms. Code: https://github.com/BIT-Vision/SATE. Lin Zhu 0012, Tengyu Long, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
NeurIPS | 1 |
| 2025 | Towards Ultra High-Speed Hyperspectral Imaging by Integrating Compressive and Neuromorphic Sampling
Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Learning Refractive-Diffractive Optics with Unidirectional Transformer for Large Field-of-View Imaging
Xiangtian Ma, Lizhi Wang 0001, Qilin Sun 0001, Lin Zhu 0012, Hua Huang 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | Event-Enhanced Snapshot Mosaic Hyperspectral Frame DeblurringabstractSnapshot Mosaic Hyperspectral Cameras (SMHCs) are popular hyperspectral imaging devices for acquiring both color and motion details of scenes. However, the narrow-band spectral filters in SMHCs may negatively impact their motion perception ability, resulting in blurry SMHC frames. In this paper, we propose a hardware-software collaborative approach to address the blurring issue of SMHCs. Our approach involves integrating SMHCs with neuromorphic event cameras for efficient event-enhanced SMHC frame deblurring. To achieve spectral information recovery guided by event signals, we formulate a spectral-aware Event-based Double Integral (sEDI) model that links SMHC frames and events from a spectral perspective, providing principled model design insights. Then, we develop a Diffusion-guided Noise Awareness (DNA) training framework that utilizes diffusion models to learn noise-aware features and promote model robustness towards camera noise. Furthermore, we design an Event-enhanced Hyperspectral frame Deblurring Network (EvHDNet) based on sEDI, which is trained with DNA and features improved spatial-spectral learning and modality interaction for reliable SMHC frame deblurring. Experiments on both synthetic data and real data show that the proposed DNA + EvHDNet outperforms state-of-the-art methods on both spatial and spectral fidelity. The code and dataset will be made publicly available. Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Building Non-Uniform Degradation Model for Position-Aware Hyperspectral Image FusionabstractThe fusion of low-spatial-resolution hyperspectral image (LR-HSI) with high-spatial-resolution multispectral image (HR-MSI) has become an effective way to obtain the high-spatial-resolution hyperspectral image (HR-HSI). Currently, learning-based methods have emerged as the mainstream solution in this field. However, these methods typically rely on predefined or simplified degradation models during fusion training, resulting in inaccurate supervision of the fusion networks. Meanwhile, most methods overlook the degradation characteristics in designing the fusion networks, leading to a mismatch between the degradation and fusion processes. These limitations ultimately result in unsatisfactory fusion performance on real data. To enhance the practicality of learning-based methods, accurate degradation modeling and effective network design have become the critical priorities. We observe that, in practical scenarios, the degree of pixel degradation varies across different positions due to the unforeseen factors such as illumination variations and imaging system fluctuations. Considering this, we propose a non-uniform degradation model (NUD), which introduces non-uniformity into the degradation processes of LR-HSI and HR-MSI. In addition, we emphasize that the essence of fusion is to reverse the degradation process. Therefore, to align with the non-uniform degradation process, the fusion process should exhibit similar positional specificity. For this purpose, we propose a position-aware fusion network (PAF), which employs positional encoding to endow the fusion process with the position-aware attribute. Experimental results show that our proposed methods provide an effective solution for HSI fusion in practical scenarios. Lizhi Wang 0001, Lin Zhu 0012, Renwei Dian, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | NER-Net+: Seeing Motion at Nighttime With an Event CameraabstractWe focus on a very challenging task: imaging at nighttime dynamic scenes. Conventional RGB cameras struggle with the trade-off between long exposure for low-light imaging and short exposure for capturing dynamic scenes. Event cameras react to dynamic changes, with their high temporal resolution (microsecond) and dynamic range (120 dB), and thus offer a promising alternative. However, existing methods are mostly based on simulated datasets due to the lack of paired event-clean image data for nighttime conditions, where the domain gap leads to performance limitations in real-world scenarios. Moreover, most existing event reconstruction methods are tailored for daytime data, overlooking issues unique to low-light events at night, such as strong noise, temporal trailing, and spatial non-uniformity, resulting in unsatisfactory reconstruction results. To address these challenges, we construct the first real paired low-light event dataset (RLED) through a co-axial imaging system, comprising 80,400 spatially and temporally aligned image GTs and low-light events, which provides a unified training and evaluation dataset for existing methods. We further conduct a comprehensive analysis of the causes and characteristics of strong noise, temporal trailing, and spatial non-uniformity in nighttime events, and propose a nighttime event reconstruction network (NER-Net+). It includes a learnable event timestamps calibration module (LETC) to correct the temporal trailing events and a non-stationary spatio-temporal information enhancement module (NSIE) to suppress sensor noise and spatial non-uniformity. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods in visual quality and generalization on real-world nighttime datasets. Haoyue Liu 0001, Shihan Peng, Yi Chang 0002, Hanyu Zhou, Yuxing Duan, Lin Zhu 0012, Yonghong Tian 0001, Luxin Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Continuous-Time Object Segmentation Using High Temporal Resolution Event CameraabstractEvent cameras are novel bio-inspired sensors, where individual pixels operate independently and asynchronously, generating intensity changes as events. Leveraging the microsecond resolution (no motion blur) and high dynamic range (compatible with extreme light conditions) of events, there is considerable promise in directly segmenting objects from sparse and asynchronous event streams in various applications. However, different from the rich cues in video object segmentation, it is challenging to segment complete objects from the sparse event stream. In this paper, we present the first framework for continuous-time object segmentation from event stream. Given the object mask at the initial time, our task aims to segment the complete object at any subsequent time in event streams. Specifically, our framework consists of a Recurrent Temporal Embedding Extraction (RTEE) module based on a novel ResLSTM, a Cross-time Spatiotemporal Feature Modeling (CSFM) module which is a transformer architecture with long-term and short-term matching modules, and a segmentation head. The historical events and masks (reference sets) are recurrently fed into our framework along with current-time events. The temporal embedding is updated as new events are input, enabling our framework to continuously process the event stream. To train and test our model, we construct both real-world and simulated event-based object segmentation datasets, each comprising event streams, APS images, and object annotations. Extensive experiments on our datasets demonstrate the effectiveness of the proposed recurrent architecture. Lin Zhu 0012, Xianzhang Chen, Lizhi Wang 0001, Xiao Wang 0014, Yonghong Tian 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Siamese Visual Tracking With Multi-Parallel Interactive TransformersabstractIn recent years, Siamese network-based visual tracking methods have gained popularity and success in terms of efficiency and accuracy. However, typical Siamese trackers utilize two independent weight-sharing streams to describe the exemplar and search region without any interaction between the two streams. As a result, such trackers employ only shallow cross-correlation or correlation filters to obtain the final information association, which neglects the deep interaction between the exemplar and search region and may reduce the discriminative power of the trackers. To address this issue, we propose a novel multi-parallel interactive transformer-based (MPIT) tracking framework to introduce sufficient interaction so that the two streams can guide the prediction heads to focus on the target more easily. Unlike recent one-stream transformer-based trackers that directly concatenate template and search tokens to perform joint feature learning, our multi-parallel interactive framework introduces a transmission band module to deliver global information for both the exemplar and the search region with low computational cost. Moreover, to integrate dynamic information, we incorporate temporal level extraction into the tracking framework to increase the variety of the templates. The experimental results show that the proposed MPIT method achieves a remarkable tracking speed of 136 frames per second (FPS) while attaining performance better than or comparable to that of state-of-the-art trackers. Wuwei Wang, Meibo Lv, Lin Zhu 0012, Tuo Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Enhancing Real-Time Object Detection With Optical Flow-Guided Streaming PerceptionabstractReal-time object detection in Unmanned Aerial Vehicle (UAV) videos remains a significant challenge due to the fast motion and small scale of objects. Existing streaming perception models struggle to accurately capture fine-grained motion cues between consecutive frames, leading to suboptimal performance in dynamic UAV scenarios. To address these challenges, Stream-Flow is proposed to integrate optical flow information and enhance real-time object detection in UAV videos. StreamFlow incorporates Flow-Guided Dynamic Prediction (FGDP) to refine position predictions using local optical flow information and Optical Flow Guided Optimization (OFGO) to optimize model parameters considering both localization loss and optical flow reliability. Central to OFGO is the Adaptive Flow Weighting (AFW) module, which focuses on reliable flow samples during training. The proposed integration of optical flow and adaptive weighting scheme significantly enhances the ability of streaming perception models to handle fast-moving objects in dynamic UAV environments. Extensive experiments on four challenging UAV video datasets demonstrate the superior performance of StreamFlow compared to state-of-the-art methods in terms of accuracy. Tongbo Wang, Lin Zhu 0012, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Simultaneous Learning Intensity and Optical Flow From High-Speed Spike StreamabstractBio-inspired vision sensors, which emulate the human retina by recording light intensity as binary spikes, have gained increasing interest in recent years. Among them, the spike camera is capable of perceiving fine textures by simulating a small retinal region called the fovea and producing high temporal resolution (20,000 Hz) spatiotemporal spike streams. To bridge the gap between binary spike streams and human vision in high-speed scenes, reconstructing intensity and optical flow from high temporal resolution spikes is particularly important. In this paper, we present a hybrid SNN-ANN network designed for simultaneous intensity and optical flow learning from spike streams. To adaptively extract spatial and temporal features from continuous spike streams, we propose a spiking neuron module with dense connections that efficiently processes both short-term and long-term spike data, while maintaining low power consumption characteristics. Subsequently, we introduce two decoders for optical flow and intensity estimation that complement each other. A temporal-aware warping module, based on flow features, is specifically designed to align the temporal features of the intensity decoder, thereby reducing motion artifacts. Concurrently, improved intensity features contribute to more accurate flow feature predictions, resulting in a mutually beneficial relationship within our network. To evaluate the effectiveness of our proposed network, we conduct experiments on both simulated and real spike datasets. Our network outperforms existing state-of-the-art spike-based reconstruction and optical flow estimation methods, demonstrating its potential for advancing the field of bio-inspired vision sensors. Our code is available athttps://github.com/LinZhu111/SLIO. Lin Zhu 0012, Weiquan Yan, Yi Chang 0002, Yonghong Tian 0001, Hua Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event CamerasabstractExisting datasets for RGB-DVS tracking are collected with DVS346 camera and their resolution ($346 \times 260$) is low for practical applications. Actually, only visible cameras are deployed in many practical systems, and the newly designed neuromorphic cameras may have different resolutions. The latest neuromorphic sensors can output high-definition event streams, but it is very difficult to achieve strict alignment between events and frames on both spatial and temporal views. Therefore, how to achieve accurate tracking with unaligned neuromorphic and visible sensors is a valuable but unresearched problem. In this work, we formally propose the task of object tracking using unaligned neuromorphic and visible cameras. We build the first unaligned frame-event dataset CRSOT collected with a specially built data acquisition system, which contains 1,030 high-definition RGB-Event video pairs, 304,974 video frames. In addition, we propose a novel unaligned object tracking framework that can realize robust tracking even using the loosely aligned RGB-Event data. This proposed method utilizes uncertainty perception techniques, which can effectively reduce the negative impact of noise (especially noise in event data) on tracking performance. Specifically, we extract the template and search regions of RGB and Event data and feed them into a unified ViT backbone for feature embedding. Next, we propose uncertainty perception modules to encode the RGB and Event features, respectively, then, we propose a modality uncertainty fusion module to aggregate the two modalities. These three branches are jointly optimized in the training phase. Extensive experiments demonstrate that our tracker can collaborate the dual modalities for high-performance tracking even without strictly temporal and spatial alignment. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Lin Zhu 0012, Zhixiang Huang, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | HARDVS: Revisiting Human Activity Recognition with Dynamic Vision SensorsabstractThe main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS. Xiao Wang 0014, Zongzhen Wu, Bo Jiang 0002, Zhimin Bao, Lin Zhu 0012, Guoqi Li 0002, Yaowei Wang 0001, Yonghong Tian 0001 |
AAAI | 5 |
| 2024 | Finding Visual Saliency in Continuous Spike StreamabstractAs a bio-inspired vision sensor, the spike camera emulates the operational principles of the fovea, a compact retinal region, by employing spike discharges to encode the accumulation of per-pixel luminance intensity. Leveraging its high temporal resolution and bio-inspired neuromorphic design, the spike camera holds significant promise for advancing computer vision applications. Saliency detection mimic the behavior of human beings and capture the most salient region from the scenes. In this paper, we investigate the visual saliency in the continuous spike stream for the first time. To effectively process the binary spike stream, we propose a Recurrent Spiking Transformer (RST) framework, which is based on a full spiking neural network. Our framework enables the extraction of spatio-temporal features from the continuous spatio-temporal spike stream while maintaining low power consumption. To facilitate the training and validation of our proposed model, we build a comprehensive real-world spike-based visual saliency dataset, enriched with numerous light conditions. Extensive experiments demonstrate the superior performance of our Recurrent Spiking Transformer framework in comparison to other spike neural network-based methods. Our framework exhibits a substantial margin of improvement in capturing and highlighting visual saliency in the spike stream, which not only provides a new perspective for spike-based saliency segmentation but also shows a new paradigm for full SNN-based transformer models. The code and dataset are available at https://github.com/BIT-Vision/SVS. Lin Zhu 0012, Xianzhang Chen, Xiao Wang 0014, Hua Huang 0001 |
AAAI | 1 |
| 2024 | SpikeNeRF: Learning Neural Radiance Fields from Continuous Spike StreamabstractSpike cameras, leveraging spike-based integration sampling and high temporal resolution, offer distinct advan-tages over standard cameras. However, existing approaches reliant on spike cameras often assume optimal illumination, a condition frequently unmet in real-world scenarios. To address this, we introduce SpikeNeRF, the first work that derives a NeRF-based volumetric scene representation from spike camera data. Our approach leverages NeRF's multi-view consistency to establish robust self-supervision, effectively eliminating erroneous measurements and uncovering coherent structures within exceedingly noisy input amidst diverse real-world illumination scenarios. The framework comprises two core elements: a spike generation model incorporating an integrate-and-fire neuron layer and parameters accounting for non-idealities, such as threshold variation, and a spike rendering loss capable of general-izing across varying illumination conditions. We describe how to effectively optimize neural radiance fields to render photorealistic novel views from the novel continuous spike stream, demonstrating advantages over other vision sen-sors in certain scenes. Empirical evaluations conducted on both real and novel realistically simulated sequences affirm the efficacy of our methodology. The dataset and source code are released at https://github.com/BIT-Vision/SpikeNeRF. Lin Zhu 0012, Kangmin Jia, Yifan Zhao 0002, Yunshan Qi, Lizhi Wang 0001, Hua Huang 0001 |
CVPR | 1 |
| 2024 | Event-Based Visible and Infrared Fusion via Multi-Task CollaborationabstractVisible and Infrared image Fusion (VIF) offers a comprehensive scene description by combining thermal infrared images with the rich textures from visible cameras. However, conventional VIF systems may capture over/under exposure or blurry images in extreme lighting and high dynamic motion scenarios, leading to degraded fusion results. To address these problems, we propose a novel Event-based Visible and Infrared Fusion (EVIF) system that employs a visible event camera as an alternative to traditional frame-based cameras for the VIF task. With extremely low latency and high dynamic range, event cameras can effectively address blurriness and are robust against diverse luminous ranges. To produce high-quality fused images, we develop a multitask collaborative framework that simultaneously performs event-based visible texture reconstruction, event-guided infrared image deblurring, and visible-infrared fusion. Rather than independently learning these tasks, our framework capitalizes on their synergy, leveraging cross-task event enhancement for efficient deblurring and bi-level min-max mutual information optimization to achieve higher fusion quality. Experiments on both synthetic and real data show that EVIF achieves remarkable performance in dealing with extreme lighting conditions and high-dynamic scenes, ensuring high-quality fused images across a broad range of practical scenarios. Mengyue Geng, Lin Zhu 0012, Lizhi Wang 0001, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
CVPR | 2 |
| 2024 | Seeing Motion at Nighttime with an Event CameraabstractWe focus on a very challenging task: imaging at night-time dynamic scenes. Most previous methods rely on the low-light enhancement of a conventional RGB camera. However, they would inevitably face a dilemma between the long exposure time of nighttime and the motion blur of dynamic scenes. Event cameras react to dynamic changes with higher temporal resolution (microsecond) and higher dynamic range (120dB), offering an alternative solution. In this work, we present a novel nighttime dynamic imaging method with an event camera. Specifically, we discover that the event at nighttime exhibits temporal trailing characteristics and spatial non-stationary distribution. Consequently, we propose a nighttime event reconstruction network (NER-Net) which mainly includes a learnable event timestamps calibration module (LETC) to align the temporal trailing events and a non-uniform illumination aware module (NIAM) to stabilize the spatiotemporal distribution of events. Moreover, we construct a paired real low-light event dataset (RLED) through a co-axial imaging system, including 64,200 spatially and temporally aligned image GTs and low-light events. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art methods in terms of visual quality and generalization ability on real-world nighttime datasets. The project are available at: https://github.com/Liu-haoyue/NER-Net. Haoyue Liu 0001, Shihan Peng, Lin Zhu 0012, Yi Chang 0002, Hanyu Zhou, Luxin Yan |
CVPR | 3 |
| 2024 | In2SET: Intra-Inter Similarity Exploiting Transformer for Dual-Camera Compressive Hyperspectral ImagingabstractDual-camera compressive hyperspectral imaging (DC-CHI) offers the capability to reconstruct 3D hyperspectral image (HSI) by fusing compressive and panchromatic (PAN) image, which has shown great potential for snapshot hyperspectral imaging in practice. In this paper, we introduce a novel DCCHI reconstruction network, intra-inter similarity exploiting Transformer (In2SET). Our key insight is to make full use of the PAN image to assist the reconstruction. To this end, we propose to use the intra-similarity within the PAN image as a proxy for approximating the intra-similarity in the original HSI, thereby offering an enhanced content prior for more accurate HSI reconstruction. Furthermore, we propose to use the inter-similarity to align the features between HSI and PAN images, thereby maintaining semantic consistency between the two modalities during the reconstruction process. By integrating In2SET into a PAN-guided deep unrolling (PGDU)framework, our method substantially enhances the spatial-spectral fidelity and detail of the reconstructed images, providing a more comprehensive and accurate depiction of the scene. Experiments conducted on both real and simulated datasets demonstrate that our approach consistently outperforms existing state-of-the-art methods in terms of reconstruction quality and computational complexity. The code is available at https://github.com/2JONAS/In2SET. Lizhi Wang 0001, Xiangtian Ma, Maoqing Zhang, Lin Zhu 0012, Hua Huang 0001 |
CVPR | 5 |
| 2024 | Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel BaselineabstractTracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impact of noisy events or sparse spatial resolution. In this paper, we propose a novel hierarchical knowledge distillation framework that can fully utilize multimodal / multi-view information during training to facilitate knowledge transfer, enabling us to achieve high-speed and low-latency visual tracking during testing by using only event signals. Specifically, a teacher Transformer-based multimodal tracking framework is first trained by feeding the RGB frame and event stream simultaneously. Then, we design a new hierarchical knowledge distillation strategy which includes pairwise similarity, feature representation, and response maps-based knowledge distillation to guide the learning of the student Transformer network. In particular, since existing event-based tracking datasets are all low-resolution (346 × 260), we propose the first large-scale high-resolution (1280 × 720) dataset named EventVOT. It contains 1141 videos and covers a wide range of categories such as pedestrians, vehicles, UAVs, ping pong, etc. Ex-tensive experiments on both low-resolution (FE240hz, Vi-sEvent, COESOT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method. Xiao Wang 0014, Shiao Wang, Chuanming Tang, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001, Jin Tang 0001 |
CVPR | 4 |
| 2024 | Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
Lin Zhu 0012, Yunlong Zheng, Yijun Zhang 0003, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ECCV (40) | 1 |
| 2024 | Deblurring Neural Radiance Fields with Event-driven Bundle AdjustmentabstractNeural Radiance Fields (NeRF) achieves impressive 3D representation learning and novel view synthesis results with high-quality multi-view images as input. However, motion blur in images often occurs in low-light and high-speed motion scenes, which significantly degrades the reconstruction quality of NeRF. Previous deblurring NeRF methods struggle to estimate pose and lighting changes during the exposure time, making them unable to accurately model the motion blur. The bio-inspired event camera measuring intensity changes with high temporal resolution makes up this information deficiency. In this paper, we propose Event-driven Bundle Adjustment for Deblurring Neural Radiance Fields (EBAD-NeRF) to jointly optimize the learnable poses and NeRF parameters by leveraging the hybrid event-RGB data. An intensity-change-metric event loss and a photo-metric blur loss are introduced to strengthen the explicit modeling of camera motion blur. Experiments on both synthetic and real-captured data demonstrate that EBAD-NeRF can obtain accurate camera trajectory during the exposure time and learn a sharper 3D representations compared to prior works. Yunshan Qi, Lin Zhu 0012, Yifan Zhao 0002, Jia Li 0003 |
ACM Multimedia | 2 |
| 2024 | Mirror complementary transformer network for RGB-thermal salient object detectionabstractAbstract Conventional RGB‐T salient object detection treats RGB and thermal modalities equally to locate the common salient regions. However, the authors observed that the rich colour and texture information of the RGB modality makes the objects more prominent compared to the background; and the thermal modality records the temperature difference of the scene, so the objects usually contain clear and continuous edge information. In this work, a novel mirror‐complementary Transformer network (MCNet) is proposed for RGB‐T SOD, which supervise the two modalities separately with a complementary set of saliency labels under a symmetrical structure. Moreover, the attention‐based feature interaction and serial multiscale dilated convolution (SDC)‐based feature fusion modules are introduced to make the two modalities complement and adjust each other flexibly. When one modality fails, the proposed model can still accurately segment the salient regions. To demonstrate the robustness of the proposed model under challenging scenes in real world, the authors build a novel RGB‐T SOD dataset VT723 based on a large public semantic segmentation RGB‐T dataset used in the autonomous driving domain. Extensive experiments on benchmark and VT723 datasets show that the proposed method outperforms state‐of‐the‐art approaches, including CNN‐based and Transformer‐based methods. The code and dataset can be found at https://github.com/jxr326/SwinMCNet . Xiurong Jiang, Hui Tian 0003, Lin Zhu 0012 |
IET Comput. Vis. | 4 |
| 2024 | Learning a physics-based filter attachment for hyperspectral imaging with RGB cameras
Maoqing Zhang, Lizhi Wang 0001, Lin Zhu 0012, Hua Huang 0001 |
Neurocomputing | 3 |
| 2024 | Unsupervised Deraining: Where Asymmetric Contrastive Learning Meets Self-SimilarityabstractMost existing learning-based deraining methods are supervisedly trained on synthetic rainy-clean pairs. The domain gap between the synthetic and real rain makes them less generalized to complex real rainy scenes. Moreover, the existing methods mainly utilize the property of the image or rain layers independently, while few of them have considered their mutually exclusive relationship. To solve above dilemma, we explore the intrinsic intra-similarity within each layer and inter-exclusiveness between two layers and propose an unsupervised non-local contrastive learning (NLCL) deraining method. The non-local self-similarity image patches as the positives are tightly pulled together and rain patches as the negatives are remarkably pushed away, and vice versa. On one hand, the intrinsic self-similarity knowledge within positive/negative samples of each layer benefits us to discover more compact representation; on the other hand, the mutually exclusive property between the two layers enriches the discriminative decomposition. Thus, the internal self-similarity within each layer (similarity) and the external exclusive relationship of the two layers (dissimilarity) serving as a generic image prior jointly facilitate us to unsupervisedly differentiate the rain from clean image. We further discover that the intrinsic dimension of the non-local image patches is generally higher than that of the rain patches. This insight motivates us to design an asymmetric contrastive loss that precisely models the compactness discrepancy of the two layers, thereby improving the discriminative decomposition. In addition, recognizing the limited quality of existing real rain datasets, which are often small-scale or obtained from the internet, we collect a large-scale real dataset under various rainy weathers that contains high-resolution rainy images. Extensive experiments conducted on different real rainy datasets demonstrate that the proposed method obtains state-of-the-art performance in real deraining. Yi Chang 0002, Yun Guo, Yuntong Ye, Changfeng Yu, Lin Zhu 0012, Xi-Le Zhao, Luxin Yan, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Reliable Event Generation With Invertible Conditional Normalizing FlowabstractEvent streams provide a novel paradigm to describe visual scenes by capturing intensity variations above specific thresholds along with various types of noise. Existing event generation methods usually rely on one-way mappings using hand-crafted parameters and noise rates, which may not adequately suit diverse scenarios and event cameras. To address this limitation, we propose a novel approach to learn a bidirectional mapping between the feature space of event streams and their inherent parameters, enabling the generation of reliable event streams with enhanced generalization capabilities. We first randomly generate a vast number of parameters and synthesize massive event streams using an event simulator. Subsequently, an event-based normalizing flow network is proposed to learn the invertible mapping between the representation of a synthetic event stream and its parameters. The invertible mapping is implemented by incorporating an intensity-guided conditional affine simulation mechanism, facilitating better alignment between event features and parameter spaces. Additionally, we impose constraints on event sparsity, edge distribution, and noise distribution through novel event losses, further emphasizing event priors in the bidirectional mapping. Our framework surpasses state-of-the-art methods in video reconstruction, optical flow estimation, and parameter estimation tasks on synthetic and real-world datasets, exhibiting excellent generalization across diverse scenes and cameras. Daxin Gu, Jia Li 0003, Lin Zhu 0012, Yu Zhang 0035, Jimmy S. J. Ren |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Stimulating Diffusion Model for Image Denoising via Adaptive Embedding and EnsemblingabstractImage denoising is a fundamental problem in computational photography, where achieving high perception with low distortion is highly demanding. Current methods either struggle with perceptual quality or suffer from significant distortion. Recently, the emerging diffusion model has achieved state-of-the-art performance in various tasks and demonstrates great potential for image denoising. However, stimulating diffusion models for image denoising is not straightforward and requires solving several critical problems. For one thing, the input inconsistency hinders the connection between diffusion models and image denoising. For another, the content inconsistency between the generated image and the desired denoised image introduces distortion. To tackle these problems, we present a novel strategy called the Diffusion Model for Image Denoising (DMID) by understanding and rethinking the diffusion model from a denoising perspective. Our DMID strategy includes an adaptive embedding method that embeds the noisy image into a pre-trained unconditional diffusion model and an adaptive ensembling method that reduces distortion in the denoised image. Our DMID strategy achieves state-of-the-art performance on both distortion-based and perception-based metrics, for both Gaussian and real-world image denoising. Tong Li 0016, Hansen Feng, Lizhi Wang 0001, Lin Zhu 0012, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Learning Adaptive Parameter Representation for Event-Based Video ReconstructionabstractEvent-based video reconstruction aims to generate images from asynchronous event streams, which record the intensity changes exceeding specific contrast thresholds. However, the contrast thresholds are varied among pixels with manufacturing imperfections and circumstancing interference, which causes undesirable events. It may cause the existing works to output blurry frames with unpleasing artifacts. To address this, we propose a novel two-stage framework to reconstruct images with learnable parameter representations. The learnable representation of the contrast threshold is extracted with a transformer network from corresponding asynchronous events in the first stage. Then a UNet architecture is utilized in the second stage to fuse the representations with the event encoding features to refine the decoding features in spatiotemporal dimensions. The representation learned from asynchronous events can adapt to the variety of contrast thresholds when processing event data in diverse scenes, motivating the proposed framework to generate high-quality frames. Quantitative and qualitative experimental results on the four public datasets show that our approach achieves better performance. Daxin Gu, Jia Li 0003, Lin Zhu 0012 |
IEEE Signal Process. Lett. | 3 |
| 2024 | VisEvent: Reliable Object Tracking via Collaboration of Frame and Event FlowsabstractDifferent from visible cameras which record intensity images frame by frame, the biologically inspired event camera produces a stream of asynchronous and sparse events with much lower latency. In practice, visible cameras can better perceive texture details and slow motion, while event cameras can be free from motion blurs and have a larger dynamic range which enables them to work well under fast motion and low illumination (LI). Therefore, the two sensors can cooperate with each other to achieve more reliable object tracking. In this work, we propose a large-scale Visible-Event benchmark (termed VisEvent) due to the lack of a realistic and scaled dataset for this task. Our dataset consists of 820 video pairs captured under LI, high speed, and background clutter scenarios, and it is divided into a training and a testing subset, each of which contains 500 and 320 videos, respectively. Based on VisEvent, we transform the event flows into event images and construct more than 30 baseline methods by extending current single-modality trackers into dual-modality versions. More importantly, we further build a simple but effective tracking algorithm by proposing a cross-modality transformer, to achieve more effective feature fusion between visible and event data. Extensive experiments on the proposed VisEvent dataset, FE108, COESOT, and two simulated datasets (i.e., OTB-DVS and VOT-DVS), validated the effectiveness of our model. The dataset and source code have been released on: https://github.com/wangxiao5791509/VisEvent_SOT_Benchmark. Xiao Wang 0014, Jianing Li 0001, Lin Zhu 0012, Zhe Chen 0013, Xin Li 0034, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Prompt-Based Learning for Unpaired Image CaptioningabstractUnpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement learning. However, these existing works were only able to learn limited cross-domain information in vision and language domains, which restrains the captioning performance of UIC. Inspired by the success of Vision-Language Pre-Trained Models (VL-PTMs) in this research, we attempt to infer the cross-domain cue information about a given image from the large VL-PTMs for the UIC task. This research is also motivated by recent successes of prompt learning in many downstream multi-modal tasks, including image-text retrieval and vision question answering. In this work, a semantic prompt is introduced and aggregated with visual features for more accurate caption prediction under the adversarial learning framework. In addition, a metric prompt is designed to select high-quality pseudo image-caption samples obtained from the basic captioning model and refine the model in an iterative manner. Extensive experiments on the COCO and Flickr30 K datasets validate the promising captioning ability of the proposed model. We expect that the proposed prompt-based UIC model will stimulate a new line of research for the VL-PTMs based captioning. Peipei Zhu, Xiao Wang 0014, Lin Zhu 0012, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2023 | E2NeRF: Event Enhanced Neural Radiance Fields from Blurry ImagesabstractNeural Radiance Fields (NeRF) achieves impressive rendering performance by learning volumetric 3D representation from several images of different views. However, it is difficult to reconstruct a sharp NeRF from blurry input as often occurred in the wild. To solve this problem, we propose a novel Event-Enhanced NeRF (E2NeRF) by utilizing the combination data of a bio-inspired event camera and a standard RGB camera. To effectively introduce event stream into the learning process of neural volumetric representation, we propose a blur rendering loss and an event rendering loss, which guide the network via modelling real blur process and event generation process, respectively. Moreover, a camera pose estimation framework for real-world data is built with the guidance of event stream to generalize the method to practical applications. In contrast to previous image-based or event-based NeRF, our framework effectively utilizes the internal relationship between events and images. As a result, E2NeRF not only achieves image deblurring but also achieves high-quality novel view image generation. Extensive experiments on both synthetic data and real-world data demonstrate that E2NeRF can effectively learn a sharp NeRF from blurry images, especially in complex and low-light scenes. Our code and datasets are publicly available at https://github.com/iCVTEAM/E2NeRF. Yunshan Qi, Lin Zhu 0012, Yu Zhang 0035, Jia Li 0003 |
ICCV | 2 |
| 2023 | Recurrent Spike-based Image Restoration under General IlluminationabstractSpike camera is a new type of bio-inspired vision sensor that records light intensity in the form of a spike array with high temporal resolution (20,000 Hz). This new paradigm of vision sensor offers significant advantages for many vision tasks such as high speed image reconstruction. However, existing spike-based approaches typically assume that the scenes are with sufficient light intensity, which is usually unavailable in many real-world scenarios such as rainy days or dusk scenes. To unlock more spike-based application scenarios, we propose a Recurrent Spike-based Image Restoration (RSIR) network, which is the first work towards restoring clear images from spike arrays under general illumination. Specifically, to accurately describe the noise distribution under different illuminations, we build a physical-based spike noise model according to the sampling process of the spike camera. Based on the noise model, we design our RSIR network which consists of an adaptive spike transformation module, a recurrent temporal feature fusion module, and a frequency-based spike denoising module. Our RSIR can process the spike array in a recursive manner to ensure that the spike temporal information is well utilized. In the training process, we generate the simulated spike data based on our noise model to train our network. Extensive experiments on real-world datasets with different illuminations demonstrate the effectiveness of the proposed network. The code and dataset are released at https://github.com/BIT-Vision/RSIR. Lin Zhu 0012, Yunlong Zheng, Mengyue Geng, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 1 |
| 2023 | Ultra-High Temporal Resolution Visual Reconstruction From a Fovea-Like Spike Camera via Spiking Neuron ModelabstractNeuromorphic vision sensor is a new bio-inspired imaging paradigm emerged in recent years. It uses the asynchronous spike signals instead of the traditional frame-based manner to achieve ultra-high speed sampling. Unlike the dynamic vision sensor (DVS) that perceives movement by imitating the retinal periphery, the spike camera was developed recently to perceive fine textures by simulating a small retinal region called the fovea. For this new type of neuromorphic camera, how to reconstruct ultra-high speed visual images from spike data becomes an important yet challenging issue in visual scene perception, analysis, and recognition applications. In this paper, a bio-inspired visual reconstruction framework for the spike camera is proposed for the first time. Its core idea is to use the biologically inspired adaptive adjustment mechanisms, combined with the spatiotemporal spike information extracted by the proposed model, to reconstruct the full texture of natural scenes in an ultra-high temporal resolution. Specifically, the proposed model consists of a motion local excitation layer, a spike refining layer and a visual reconstruction layer motivated by the bio-realistic leaky integrate-and-fire (LIF) neurons and synapse connection with spike-timing dependent plasticity (STDP) rule. To evaluate the performance, a spike dataset was constructed for normal and high-speed scenes in real-world recorded by the spike camera. The experimental results show that the proposed approach can reconstruct the visual images with 40,000 frames per second in both normal and high-speed scenes, while achieving high dynamic range and high image quality. Lin Zhu 0012, Siwei Dong, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Learning Super-Resolution Reconstruction for High Temporal Resolution Spike StreamabstractSpike camera is a new type of bio-inspired vision sensor, each pixel of which perceives the brightness of the scene independently, and finally outputs 3-dimensional spatiotemporal spike streams. To bridge the spike camera and traditional frame-based vision, there is some works to reconstruct spike streams into regular images. However, the low spatial resolution ($400\times 250$) of the spike camera limits the quality of the reconstructed images. Thus, it is meaningful to explore a super-resolution reconstruction for spike streams. In this paper, we propose an end-to-end network to reconstruct high-resolution images from low-resolution spike streams. To utilize more spatiotemporal features of spike streams, our network adopts a multi-level features learning mechanism, including intra-stream feature extraction by spike encoder, inter-stream dependencies extraction based on optical flow module, and joint features learning via spike-based iterative projection. Experimental results demonstrate that our network is superior to the combination of state-of-the-art intensity image reconstruction methods and super-resolution networks on simulated and real datasets. Xijie Xiang, Lin Zhu 0012, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Retinomorphic Object Detection in Asynchronous Visual StreamsabstractDue to high-speed motion blur and challenging illumination, conventional frame-based cameras have encountered an important challenge in object detection tasks. Neuromorphic cameras that output asynchronous visual streams instead of intensity frames, by taking the advantage of high temporal resolution and high dynamic range, have brought a new perspective to address the challenge. In this paper, we propose a novel problem setting, retinomorphic object detection, which is the first trial that integrates foveal-like and peripheral-like visual streams. Technically, we first build a large-scale multimodal neuromorphic object detection dataset (i.e., PKU-Vidar-DVS) over 215.5k spatio-temporal synchronized labels. Then, we design temporal aggregation representations to preserve the spatio-temporal information from asynchronous visual streams. Finally, we present a novel bio-inspired unifying framework to fuse two sensing modalities via a dynamic interaction mechanism. Our experimental evaluation shows that our approach has significant improvements over the state-of-the-art methods with the single-modality, especially in high-speed motion and low-light scenarios. We hope that our work will attract further research into this newly identified, yet crucial research direction. Our dataset can be available at https://www.pkuml.org/resources/pku-vidar-dvs.html. Jianing Li 0001, Xiao Wang 0014, Lin Zhu 0012, Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001 |
AAAI | 3 |
| 2022 | Event-based Video Reconstruction via Potential-assisted Spiking Neural NetworkabstractNeuromorphic vision sensor is a new bio-inspired imaging paradigm that reports asynchronous, continuously perpixel brightness changes called ‘events’ with high temporal resolution and high dynamic range. So far, the event-based image reconstruction methods are based on artificial neural networks (ANN) or hand-crafted spatiotemporal smoothing techniques. In this paper, we first implement the image reconstruction work via deep spiking neural network (SNN) architecture. As the bio-inspired neural networks, SNNs operating with asynchronous binary spikes distributed over time, can potentially lead to greater computational efficiency on event-driven hardware. We propose a novel Event-based Video reconstruction framework based on a fully Spiking Neural Network (EVSNN), which utilizes Leaky-Integrate-and-Fire (LIF) neuron and Membrane Potential (MP) neuron. We find that the spiking neurons have the potential to store useful temporal information (memory) to complete such time-dependent tasks. Further-more, to better utilize the temporal information, we propose a hybrid potential-assisted framework (PAEVSNN) using the membrane potential of spiking neuron. The proposed neuron is referred as Adaptive Membrane Potential (AMP) neuron, which adaptively updates the membrane potential according to the input spikes. The experimental results demonstrate that our models achieve comparable performance to ANN-based models on IJRR, MVSEC, and HQF datasets. The energy consumptions of EVSNN and PAEVSNN are$19.36\times$and$7.75\times$more computationally ef-ficient than their ANN architectures, respectively. The code and pretrained model are available at https://sites.google.com/view/evsnn. Lin Zhu 0012, Xiao Wang 0014, Yi Chang 0002, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
CVPR | 1 |
| 2022 | Unsupervised Deraining: Where Contrastive Learning Meets Self-similarityabstractImage deraining is a typical low-level image restoration task, which aims at decomposing the rainy image into two distinguishable layers: clean image layer and rain layer. Most of the existing learning-based deraining methods are supervisedly trained on synthetic rainy-clean pairs. The domain gap between the synthetic and real rains makes them less generalized to different real rainy scenes. Moreover, the existing methods mainly utilize the property of the two layers independently, while few of them have considered the mutually exclusive relationship between the two layers. In this work, we propose a novel non-local contrastive learning (NLCL) method for unsupervised image deraining. Consequently, we not only utilize the intrinsic self-similarity property within samples, but also the mutually exclusive property between the two layers, so as to better differ the rain layer from the clean image. Specifically, the non-local self-similarity image layer patches as the positives are pulled together and similar rain layer patches as the negatives are pushed away. Thus the similar positive/negative samples that are close in the original space benefit us to enrich more discriminative representation. Apart from the self-similarity sampling strategy, we analyze how to choose an appropriate feature encoder in NLCL. Extensive experiments on different real rainy datasets demonstrate that the proposed method obtains state-of-the-art performance in real deraining. Yuntong Ye, Changfeng Yu, Yi Chang 0002, Lin Zhu 0012, Xi-Le Zhao, Luxin Yan, Yonghong Tian 0001 |
CVPR | 4 |
| 2022 | Learning Stereo Depth Estimation with Bio-Inspired Spike CamerasabstractBio-inspired spike cameras, offering high temporal resolution spike streams, have brought a new perspective to address common challenges (e.g.,high-speed motion blur) in depth estimation tasks. In this paper, we propose a novel problem setting, spike-based stereo depth estimation, which is the first trail that explores an end-to-end network to learn stereo depth estimation with transformers for spike cameras, named Spike-based Stereo Depth Estimation Transformer (SSDEFormer). We first build a hybrid camera platform and provide a new stereo depth estimation dataset (i.e.,PKU-Spike-Stereo) with spatiotemporal synchronized labels. Then, we propose a novel spike representation to effectively exploit spatiotemporal information from spike streams. Finally, a transformer-based network is designed to generate dense depth maps without a fixed-disparity cost volume. Empirically, it shows that our approach is extremely effective on both synthetic and real-world datasets. The results verify that spike cameras can perform robust depth estimation even in cases where conventional cameras and event cameras fail in fast motion scenarios. Jianing Li 0001, Lin Zhu 0012, Xijie Xiang, Tiejun Huang 0001, Yonghong Tian 0001 |
ICME | 3 |
| 2022 | Temporal Up-Sampling for Asynchronous EventsabstractThe event camera is a novel bio-inspired vision sensor. When the brightness change exceeds the preset threshold, the sensor generates events asynchronously. The number of valid events directly affects the performance of event-based tasks, such as reconstruction, detection, and recognition. However, when in low-brightness or slow-moving scenes, events are often sparse and accompanied by noise, which poses challenges for event-based tasks. To solve these challenges, we propose an event temporal up-sampling algorithm11Code: https://github.com/XIJIE-XIANG/Event-Temporal-Up-sampling to generate more effective and reliable events. The main idea of our algorithm is to generate up-sampling events on the event motion trajectory. First, we estimate the event motion trajectory by contrast maximization algorithm and then up-sampling the events by temporal point processes. Experimental results show that up-sampling events can provide more effective information and improve the performance of downstream tasks, such as improving the quality of reconstructed images and increasing the accuracy of object detection. Xijie Xiang, Lin Zhu 0012, Jianing Li 0001, Yonghong Tian 0001, Tiejun Huang 0001 |
ICME | 2 |
| 2022 | Asynchronous Spatio-Temporal Memory Network for Continuous Event-Based Object DetectionabstractEvent cameras, offering extremely high temporal resolution and high dynamic range, have brought a new perspective to addressing common object detection challenges (e.g., motion blur and low light). However, how to learn a better spatio-temporal representation and exploit rich temporal cues from asynchronous events for object detection still remains an open issue. To address this problem, we propose a novel asynchronous spatio-temporal memory network (ASTMNet) that directly consumes asynchronous events instead of event images prior to processing, which can well detect objects in a continuous manner. Technically, ASTMNet learns an asynchronous attention embedding from the continuous event stream by adopting an adaptive temporal sampling strategy and a temporal attention convolutional module. Besides, a spatio-temporal memory module is designed to exploit rich temporal cues via a lightweight yet efficient inter-weaved recurrent-convolutional architecture. Empirically, it shows that our approach outperforms the state-of-the-art methods using the feed-forward frame-based detectors on three datasets by a large margin (i.e., 7.6% in the KITTI Simulated Dataset, 10.8% in the Gen1 Automotive Dataset, and 10.5% in the 1Mpx Detection Dataset). The results demonstrate that event cameras can perform robust object detection even in cases where conventional cameras fail, e.g., fast motion and challenging light conditions. Jianing Li 0001, Jia Li 0003, Lin Zhu 0012, Xijie Xiang, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | NeuSpike-Net: High Speed Video Reconstruction via Bio-inspired Neuromorphic CamerasabstractNeuromorphic vision sensor is a new bio-inspired imaging paradigm that emerged in recent years, which continuously sensing luminance intensity and firing asynchronous spikes (events) with high temporal resolution. Typically, there are two types of neuromorphic vision sensors, namely dynamic vision sensor (DVS) and spike camera. From the perspective of bio-inspired sampling, DVS only perceives movement by imitating the retinal periphery, while the spike camera was developed to perceive fine textures by simulating the fovea. It is meaningful to explore how to combine two types of neuromorphic cameras to reconstruct high quality image like human vision. In this paper, we propose a NeuSpike-Net to learn both the high dynamic range and high motion sensitivity of DVS and the full texture sampling of spike camera to achieve high-speed and high dynamic image reconstruction. We propose a novel representation to effectively extract the temporal information of spike and event data. By introducing the feature fusion module, the two types of neuromorphic data achieve complementary to each other. The experimental results on the simulated and real datasets demonstrate that the proposed approach is effective to reconstruct high-speed and high dynamic range images via the combination of spike and event data. Lin Zhu 0012, Jianing Li 0001, Xiao Wang 0014, Tiejun Huang 0001, Yonghong Tian 0001 |
ICCV | 1 |
| 2021 | Retinomorphic Sensing: A Novel Paradigm for Future Multimedia ComputingabstractConventional frame-based cameras for multimedia computing have encountered important challenges in high-speed and extreme light scenarios. However, how to design a novel paradigm for visual perception that overcomes the disadvantages of conventional cameras still remains an open issue. In this paper, we propose a novel solution, namely retinomorphic sensing, which integrates fovea-like and peripheral-like sampling mechanisms to generate asynchronous visual streams using a unified representation as the retina does. Technically, our encoder incorporates an interaction controller to switch flexibly between dynamic and static sensing. Then, the decoder effectively extracts dynamic events for machine vision and reconstructs visual textures for human vision. The results show that our strategy enables it to sense dynamic events and visual textures meanwhile reduce data redundancy. We further build a prototype hybrid camera system to verify this strategy on vision tasks such as image reconstruction and object detection. We believe that this novel paradigm will provide insight into future multimedia computing. The code can be available at https://github.com/acmmm2021-bni-retinomorphic/retinomorphic-sensing. Zhaodong Kang, Jianing Li 0001, Lin Zhu 0012, Yonghong Tian 0001 |
ACM Multimedia | 3 |
| 2021 | Learning event guided network for salient object detection
Xiurong Jiang, Lin Zhu 0012, Hui Tian 0003 |
Pattern Recognit. Lett. | 2 |
| 2021 | Hybrid Coding of Spatiotemporal Spike Data for a Bio-Inspired CameraabstractRecently, a novel bio-inspired camera was developed by mimicking the retina fovea to continuously accumulate luminance intensity and then fire spikes once the dispatch threshold is reached. In contrast to the conventional frame-based cameras and the emerging dynamic vision sensors, this spike camera has shown remarkable advantages in capturing fast-moving scenes in a frame-free manner with full texture reconstruction capabilities. However, the ultra-high temporal resolution makes the transmission or storage of the output data of spike camera (referred to as spike data) quite difficult. To address the above challenges, we propose a unified lossy spike coding framework, which exploits the motion patterns hidden in the spike data distribution to design the motion-fidelity coding modes for the first time. We investigate the spatiotemporal distribution of spike data and propose an intensity-based measurement of the spike train distance. Then, the adaptive polyhedron partitioning is proposed to deal with the spike data with different motion characteristics. Finally, the intra-/inter-polyhedron prediction with spike-time and spike-rate modes, transform and multi-layer quantization are proposed and introduced into the codec. We also construct a PKU-Spike dataset captured by the spike camera to evaluate the compression performance. The experimental results on the dataset demonstrate that the proposed approach is effective in compressing such spike data while maintaining the visual fidelity especially for high-speed scenarios. Lin Zhu 0012, Siwei Dong, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Motion-Aware Structured Matrix Factorization for Foreground Detection in Complex ScenesabstractForeground detection is one of the key steps in computer vision applications. Many foreground and background models have been proposed and achieved promising performance in static scenes. However, due to challenges such as dynamic background, irregular movement, and noise, most algorithms degrade sharply in complex scenes. To address the problem, we propose a motion-aware structured matrix factorization approach (MSMF), which integrates the structural and spatiotemporal motion information into a unified sparse-low-rank matrix factorization framework. Technologically, it has three main contributions: First, a variant of structured sparsity-inducing norm is proposed to constrain both structure and sparsity of foreground. The model is robust to the statistical variability of the underlying foreground pixels in complex scenes. Second, to capture the ambiguous pixels, a spatiotemporal cube-based motion trajectory is extracted for assisting matrix factorization. Finally, to solve the optimization problem of structured matrix factorization, we develop an augmented Lagrange multiplier method with the alternating direction strategy and Douglas-Rachford monotone operator splitting algorithm. Experiments demonstrate that the proposed approach achieves impressive performance in separating irregular moving foreground while suppressing the dynamic background and the noise, and outperforms some state-of-the-art algorithms. Lin Zhu 0012, Xiurong Jiang, Jianing Li 0001, Yuanhong Hao, Yonghong Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Retina-Like Visual Image Reconstruction via Spiking Neural ModelabstractThe high-sensitivity vision of primates, including humans, is mediated by a small retinal region called the fovea. As a novel bio-inspired vision sensor, spike camera mimics the fovea to record the nature scenes by continuous-time spikes instead of frame-based manner. However, reconstructing visual images from the spikes remains to be a challenge. In this paper, we design a retina-like visual image reconstruction framework, which is flexible in reconstructing full texture of natural scenes from the totally new spike data. Specifically, the proposed architecture consists of motion local excitation layer, spike refining layer and visual reconstruction layer motivated by bio-realistic leaky integrate and fire (LIF) neurons and synapse connection with spike-timing-dependent plasticity (STDP) rules. This approach may represent a major shift from conventional frame-based vision to the continuous-time retina-like vision, owning to the advantages of high temporal resolution and low power consumption. To test the performance, a spike dataset is constructed which is recorded by the spike camera. The experimental results show that the proposed approach is extremely effective in reconstructing the visual image in both normal and high speed scenes, while achieving high dynamic range and high image quality. Lin Zhu 0012, Siwei Dong, Jianing Li 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
CVPR | 1 |
| 2019 | An Efficient Coding Method for Spike Camera Using Inter-Spike IntervalsabstractRecently, a novel bio-inspired spike camera has been proposed, which continuously accumulates luminance intensity and fires spikes once the dispatch threshold is reached. It has shown great advantages in capturing fast-moving scene in a frame-free manner with full texture reconstruction capabilities. However, it is difficult to transmit or store the large amount of spike data. By investigating the spatiotemporal distribution of the spikes, we propose an intensity-based measurement for spike train distance and design an efficient coding method to meet the challenge. First, the spike train is transformed into inter-spike intervals (ISIs), and ISIs are adaptively partitioned into multiple segments in temporal. Then, intra-and inter-pixel prediction are performed to find the best reference candidate. The prediction residuals are quantized to achieve lossy compression. Finally, the quantized residuals are fed into an adaptive context-based entropy coder. Overall, to achieve the best performance, each prediction mode will be tried and the one with minimum rate-distortion cost is chosen. Siwei Dong, Lin Zhu 0012, Daoyuan Xu, Yonghong Tian 0001, Tiejun Huang 0001 |
DCC | 2 |
| 2019 | A Retina-Inspired Sampling Method for Visual Texture ReconstructionabstractConventional frame-based camera is not able to meet the demand of rapid reaction for real-time applications, while the emerging dynamic vision sensor (DVS) can realize high speed capturing for moving objects. However, to achieve visual texture reconstruction, DVS need extra information apart from the output spikes. This paper introduces a fovea-like sampling method inspired by the neuron signal processing in retina, which aims at visual texture reconstruction only taking advantage of the properties of spikes. In the proposed method, the pixels independently respond to the luminance changes with temporal asynchronous spikes. Analyzing the arrivals of spikes makes it possible to restore the luminance information, enabling reconstructing the natural scene for visualization. Three decoding methods of spike stream for texture reconstruction are proposed for high-speed motion and stationary scenes. Compared to conventional frame-based camera and DVS, our model can achieve better image quality and higher flexibility, which is capable of changing the way that demanding machine vision applications are built. Lin Zhu 0012, Siwei Dong, Tiejun Huang 0001, Yonghong Tian 0001 |
ICME | 1 |
| 2018 | L1/2 Norm and Spatial Continuity Regularized Low-Rank Approximation for Moving Object Detection in Dynamic BackgroundabstractLow-rank modeling-based moving object detection approaches proposed so far use fixed l1-norm penalty to capture the sparse nature of foreground in video, and thus, hardly adapt readily to the statistical variability of underlying foreground pixels in dynamic background. Additionally, they ignore the spatial continuity prior among the neighbor foreground pixels. Consequently, they cannot offer a satisfactory performance in practical dynamic background. In this letter, we present a unified regularization framework, namely l1/2-norm and spatial continuity regularized low-rank approximation (SCLR-l1/2), to solve this problem. First, in order to promote accuracy, we introduce an l1/2constraint into the framework. Second, to guarantee the continuity among the neighbor foreground pixels, we introduce a spatial continuity regularization term, motivated by total variation. Finally, we generalize our framework to the lq-norm penalized case (SCLR-lq). By adjusting the shrinkage parameter q, the framework gets better flexibility to choose a reasonable sparse domain. To deal with the present constrained minimization problem, the augmented Lagrange multiplier method is employed and extended with the help of the alternating direction minimizing strategy. Experimental results show that the proposed method outperforms some state-of-the-art algorithms especially for the cases with dynamic backgrounds. Lin Zhu 0012, Yuanhong Hao, Yuejin Song |
IEEE Signal Process. Lett. | 1 |