Junwei Zhao 0003

dblp:70/1157-3 · DBLP profile ↗
← Back
8ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0002-0301-3638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Learning Robust Representation from Imbalanced Event Streams with Adaptive STDP
abstract
UAVs require robust visual perception under rapid motion and varying illumination, conditions for which event cameras are well-suited. Such cameras asynchronously capture brightness changes through ON and OFF events, providing an efficient modality for visual data processing. Bio-inspired Spiking Neural Networks (SNNs), utilizing spike-driven neuron models, are highly suitable for processing such data due to their ability to encode temporal dynamics. However, the inherent imbalance between ON and OFF events poses challenges, limiting the performance of event-based algorithms. This paper introduces an adaptive representation learning method that models event streams as temporal point processes and dynamically balances synaptic plasticity between ON and OFF events. The proposed method ensures that learning is equally driven by both event polarities, effectively mitigating the imbalance issue and enhancing the robustness of SNNs. Experimental results on benchmark datasets, including N-CARS, N-CALTECH101, DVS-CIFAR10, and CEP-DVS, demonstrate improvements in SNN classification accuracy. These findings underscore the effectiveness of the proposed adaptive bio-inspired approach, offering new possibilities for visual data analysis and neuromorphic computing.
Junwei Zhao 0003
ECAI1
2025 HDCFN: Haze Distribution-aware Cross-modal Fusion Network for Infrared-guided Dense Haze Removal in UAVs
abstract
In UAV applications, dense haze severely obscures small ground-level objects, hindering the recovery of fine details. Existing visible-only dehazing methods struggle with such dense occlusions, while infrared imaging lacks color and fine texture information. To address these limitations, we propose the Haze Distribution-aware Cross-modal Fusion Network (HDCFN). HDCFN features two key components: (i) an infrared-guided multiscale feature enhancement framework that integrates haze-resistant structural cues from infrared modality with visible features across coarse to fine, improving the recovery of small objects, and (ii) a haze distribution-aware cross-modal fusion module that adaptively prioritizes relevant information from each modality according to haze density. This framework effectively combines the complementary strengths of visible and infrared imaging for dense haze removal. Extensive experiments on multiple public datasets show that HDCFN outperforms state-of-the-art dehazing and fusion methods, yielding higher-quality and more detailed images.
Junwei Zhao 0003, Qianchun Luo, Shiliang Zhang, Shen Gao, Jie Wu 0001
ACM Multimedia1
2024 Recognizing Ultra-High-Speed Moving Objects with Bio-Inspired Spike Camera
abstract
Bio-inspired spike camera mimics the sampling principle of primate fovea. It presents high temporal resolution and dynamic range, showing great promise in fast-moving object recognition. However, the physical limit of CMOS technology in spike cameras still hinders their capability of recognizing ultra-high-speed moving objects, e.g., extremely fast motions cause blur during the imaging process of spike cameras. This paper presents the first theoretical analysis for the causes of spiking motion blur and proposes a robust representation that addresses this issue through temporal-spatial context learning. The proposed method leverages multi-span feature aggregation to capture temporal cues and employs residual deformable convolution to model spatial correlation among neighbouring pixels. Additionally, this paper contributes an original real-captured spiking recognition dataset consisting of 12,000 ultra-high-speed (equivalent speed > 500 km/h) moving objects. Experimental results show that the proposed method achieves 73.2% accuracy in recognizing 10 classes of ultra-high-speed moving objects, outperforming all existing spike-based recognition methods. Resources will be available at https://github.com/Evin-X/UHSR.
Junwei Zhao 0003, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001
AAAI1
2024 SpiReco: Fast and Efficient Recognition of High-Speed Moving Objects With Spike Camera
abstract
Benefited from the high temporal resolution and high dynamic range, spike cameras have shown great potential in recognizing high-speed moving objects. However, the computer vision community has not explored this task due to the lack of spike data and annotations of high-speed moving objects. This paper contributes a novel dataset, namedSpiReco(Spiking datasets forRecognition), by recording high-speed moving objects using a spike camera. To annotate the dataset, image labels from established datasets such as MNIST, CIFAR10, and CALTECH101 are utilized. Based on this new dataset, this paper proposes the first spike-based object recognition framework. The proposed framework includes a denoise module, which is designed to suppress spike noise by learning spatio-temporal correlation from neighbouring pixels. Additionally, a motion enhancement module is introduced to address high-speed and random motions. Afterward, binarized neural networks are adopted to save computation costs. These efforts result in a fast and efficient processing framework for spiking data. Experimental results demonstrate the effectiveness of the proposed methods. For example, the proposed spike-based recognition framework achieves 80.2% accuracy in recognizing 101 classes of high-speed moving objects using only 2.2ms of spike streams. The SpiReco is available at https://github.com/Evin-X/SpiReco.
Junwei Zhao 0003, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Recognizing High-Speed Moving Objects with Spike Camera
abstract
Spike camera is a novel bio-inspired vision sensor that mimics the sampling mechanism of the primate fovea. It presents high temporal resolution and dynamic range, showing great potentials in the high-speed moving object recognition task, which has not been fully explored in the Multimedia community due to the lack of data and annotations. This paper contributes the first large-scale High-Speed Spiking Recognition (HSSR) dataset, by recording high-speed moving objects using a spike camera. The HSSR dataset contains 135,000 indoor objects annotated using ImageNet labels and 3,100 outdoor objects collected from real-world scenarios. Furthermore, we propose an original spiking recognition framework, which employs long-term spike stream features to supervise the feature learning from short-term spike streams. This framework improves the recognition accuracy, meanwhile substantially decreasing the recognition latency, making our method can accurately recognize moving objects at an equivalent speed of 514 km/h, using only 1 ms of spike stream. Experimental results show that, the proposed method achieves 76.5% accuracy for recognizing 100 fine-grained indoor objects and 84.3% accuracy for recognizing 8 outdoor objects using 1 ms of spike streams. Resources will be available at https://github.com/Evin-X/HSSR.
Junwei Zhao 0003, Jianming Ye, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001
ACM Multimedia1
2022 Modeling The Detection Capability Of High-Speed Spiking Cameras
abstract
The novel working principle enables spiking cameras to capture high-speed moving objects. However, the applications of spiking cameras can be affected by many factors, such as brightness intensity, detectable distance, and the maximum speed of moving targets. Improper settings such as weak ambient brightness and too short object-camera distance, will lead to failure in the application of such cameras. To address the issue, this paper proposes a modeling algorithm that studies the detection capability of spiking cameras. The algorithm deduces the maximum detectable speed of spiking cameras corresponding to different scenario settings (e.g., brightness intensity, camera lens, and object-camera distance) based on the basic technical parameters of cameras (e.g., pixel size, spatial and temporal resolution). Thereby, the proper camera settings for various applications can be determined. Extensive experiments verify the effectiveness of the modeling algorithm. To our best knowledge, it is the first work to investigate the detection capability of spiking cameras.
Junwei Zhao 0003, Zhaofei Yu, Lei Ma 0008, Ziluo Ding, Shiliang Zhang, Yonghong Tian 0001, Tiejun Huang 0001
ICASSP1
2022 Transformer-Based Domain Adaptation for Event Data Classification
abstract
Event cameras encode the change of brightness into events, differing from conventional frame cameras. The novel working principle makes them to have stronger potential in high-speed applications. However, the lack of labeled event annotations limits the applications of such cameras in deep learning frameworks, making it appealing to study more efficient deep learning algorithms and architectures. This paper devises the Convolutional Transformer Network (CTN) for processing event data. The CTN enjoys the advantages of convolution networks and transformers, presenting stronger capability in event-based classification tasks compared with existing models. To address the insufficiency issue of annotated event data, we propose to train the CTN via the source-free Unsupervised Domain Adaptation (UDA) algorithm leveraging large-scale labeled image data. Extensive experiments verify the effectiveness of the UDA algorithm. And our CTN outperforms recent state-of-the-art methods on event-based classification tasks, suggesting that it is an effective model for this task. To our best acknowledge, it is an early attempt of employing vision transformers with the source-free UDA algorithm to process event data.
Junwei Zhao 0003, Shiliang Zhang, Tiejun Huang 0001
ICASSP1
2022 SpikingSIM: A Bio-Inspired Spiking Simulator
abstract
Large-scale neuromorphic dataset is costly to construct and difficult to annotate because of the unique high-speed asynchronous imaging principle of bio-inspired cameras. Lacking of large-scale annotated neuromorphic datasets has significantly hindered the applications of bio-inspired cameras in deep neural networks. Synthesizing neuromorphic data from annotated RGB images can be considered to alleviate this challenge. This paper proposes a simulator to generate simulated spiking data from images recorded by frame cameras. To minimize the deviations between synthetic data and real data, the proposed simulator named SpikingSIM considers the sensing principle of spiking cameras, and generates high-quality simulated spiking data, e.g., the noises in real data are also simulated. Experimental results show that, our simulator generates more realistic spiking data than existing methods. We hence train deep neural networks with synthesized spiking data. Experiments show that, the net- work trained by our simulated data generalizes well on real spiking data. The source code of SpikingSIM is available at http://github.com/Evin-X/SpikingSIM.
Junwei Zhao 0003, Shiliang Zhang, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001
ISCAS1