Yue Zhou 0010

dblp:78/6191-10 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-9323-4524ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 35% Video understanding and tracking · 32% Efficient and distributed learning · 14%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
1.522025
Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025
TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023
Computer vision › Video understanding and tracking › action recognition
event-based action recognition
1.522025
Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025
TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023
Computer vision › 3D vision › visual localization
camera relocalization
1.022025
A Simple and Effective Point-Based Network for Event Camera 6-DOFs Pose Relocalization · CVPR 2024
Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection › point set representation
point cloud representation
0.912025
Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training
spiking neural network
0.912025
ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring · ICCV 2025
Image and video processing
image restoration
0.912025
ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring · ICCV 2025
Image and video processing › image restoration › image deblurring
motion deblurring
0.912025
ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring · ICCV 2025
Machine learning › Efficient and distributed learning
model compression
0.712023
TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023
Computer vision › 3D vision › point cloud analysis
point cloud learning
0.712023
TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023
Computer vision › 3D vision › geometric deep learning
point cloud network
0.712023
TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition
0.712023
TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

visual attention · 1.7spiking neural network · 1.7cross-modal fusion · 1.7state space model · 0.9mamba · 0.9hierarchical feature extraction · 0.9point-based network · 0.8hierarchical feature abstraction · 0.8attentive bidirectional LSTM · 0.8downsampling · 0.7
YearPublicationVenuePosition
2026 Event-Based Motion Deblurring via Multi-Temporal Granularity Fusion
abstract
Conventional frame-based cameras inevitably produce blurry effects due to motion occurring during the exposure time. Event camera, a bio-inspired sensor offering continuous visual information could enhance the deblurring performance. Effectively utilizing the high-temporal-resolution event data is crucial for extracting precise motion information and enhancing deblurring performance. However, existing event-based image deblurring methods usually utilize voxel-based event representations, losing the fine-grained temporal details that are mathematically essential for fast motion deblurring. In this paper, we first introduce point cloud-based event representation into the image deblurring task and propose a Multi-Temporal Granularity Network (MTGNet). It combines the spatially dense but temporally coarse-grained voxel-based event representation and the temporally fine-grained but spatially sparse point cloud-based event. To seamlessly integrate such complementary representations, we design a Fine-grained Point Branch. An Aggregation and Mapping Module (AMM) is proposed to align the low-level point-based features with frame-based features and an Adaptive Feature Diffusion Module (AFDM) is designed to manage the resolution discrepancies between event data and image data by enriching the sparse point feature. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art approaches on both synthetic and real-world datasets. Our code is available at: https://github.com/xplin13/MTGNet.
Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Yue Zhou 0010, Haotian Fu, Biao Pan, Bojun Cheng
IEEE Trans. Circuits Syst. Video Technol.5
2025 ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring
abstract
Motion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of event streams, we employ Spiking Neural Networks (SNNs) for motion feature extraction and Artificial Neural Networks (ANNs) for color information processing. Due to the non-uniform distribution and inherent redundancy of event data, existing cross-modal feature fusion methods exhibit certain limitations. Inspired by the visual attention mechanism in the human visual system, this study introduces a bioinspired dual-drive hybrid network (BDHNet). Specifically, the Neuron Configurator Module (NCM) is designed to dynamically adjusts neuron configurations based on cross-modal features, thereby focusing the spikes in blurry regions and adapting to varying blurry scenarios dynamically. Additionally, the Region of Blurry Attention Module (RBAM) is introduced to generate a blurry mask in an unsupervised manner, effectively extracting motion clues from the event features and guiding more accurate cross-modal feature fusion. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art methods on both synthetic and real-world datasets.
Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Hongxiang Huang, Yue Zhou 0010, Haotian Fu, Bojun Cheng
ICCV6
2025 Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression
abstract
Event cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks.
Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 A Simple and Effective Point-Based Network for Event Camera 6-DOFs Pose Relocalization
abstract
Event cameras exhibit remarkable attributes such as high dynamic range, asynchronicity, and low latency, making them highly suitable for vision tasks that involve highspeed motion in challenging lighting conditions. These cameras implicitly capture movement and depth information in events, making them appealing sensors for Camera Pose Relocalization (CPR) tasks. Nevertheless, existing CPR networks based on events neglect the pivotal finegrained temporal information in events, resulting in unsatisfactory performance. Moreover, the energy-efficient features are further compromised by the use of excessively complex models, hindering efficient deployment on edge devices. In this paper, we introduce PEPNet, a simple and effective point-based network designed to regress six degrees of freedom (6-DOFs) event camera poses. We rethink the relationship between the event camera and CPR tasks, leveraging the raw Point Cloud directly as network input to harness the high-temporal resolution and inherent sparsity of events. PEPNet is adept at abstracting the spatial and implicit temporal features through hierarchical structure and explicit temporal features by Attentive Bidirectional Long Short-Term Memory (A-Bi-LSTM). Byemploying a carefully crafted lightweight design, PEPNet delivers state-of-the-art (SOTA) performance on both indoor and outdoor datasets with meager computational resources. Specifically, PEPNet attains a significant 38% and 33% performance improvement on the random split IJRR and M3ED datasets, respectively. Moreover, the lightweight design version PEPNettinyaccomplishes results comparable to the SOTA while employing a mere 0.5% of the parameters.
Jiadong Zhu, Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Bojun Cheng
CVPR3
2024 DS-CIM: A 40nm Asynchronous Dual-Spike Driven, MRAM Compute-In-Memory Macro for Spiking Neural Network
abstract
Compute-in-memory (CIM) based on emerging nonvolatile memory (eNVM) is an effective way to deploy neural networks to low-power edge devices for both storage and computation. NVMs such as ReRAM have been widely used in CIM. Meanwhile, MRAM has higher read and write cycles, lower device and cycle variation and a lower bit error rate, making it equally attractive for storage. However, the high read current and low on/off ratio result in large energy consumption in MRAM read limiting its large-scale application in CIM. The spiking neural network (SNN) represents the information as sparse spike sequences and facilitates hardware to achieve low-power computing by taking advantage of its spatial-temporal sparsity. To further increase the input sparsity of SNN and reduce the read energy consumption, this paper proposes ADC-free, dual-spike (DS) -CIM macro, a spiking MRAM CIM macro driven by asynchronous dual spikes. Compared to the conventional rate coding, our dual-spike coding method uses only 2 spikes to encode the information without losing accuracy. Moreover, the event-driven feature allows the macro to have sub-nW static power consumption. Our DS-CIM macro achieves comparable or higher accuracy while maintaining very low energy consumption. Specifically, it achieves accuracies of 96.99%, 82.87%, 90.00%, and 85.97% for digit classification, image classification, gesture recognition, and action recognition tasks, with energy consumption of only 8.07nJ, 71.26nJ, 729.3nJ, and 369.82nJ, respectively. These results emphasize the significance of DS-CIM and provide ideas for low-power inference on edge devices.
Haotian Fu, Yulong Huang 0001, Tingran Chen, Chenyi Fu, Yue Zhou 0010, Shouzhong Peng, Zhirui Zong, Biao Pan, Bojun Cheng
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras
abstract
Event cameras have gained popularity in computer vision due to their data sparsity, high dynamic range, and low latency. As a bio-inspired sensor, event cameras generate sparse and asynchronous data, which is inherently incompatible with the traditional frame-based method. Alternatively, the point-based method can avoid additional modality transformation and naturally adapt to the sparsity of events. Still, it typically cannot reach a comparable accuracy as the frame-based method. We propose a lightweight and generalized point cloud network called TTPOINT which achieves competitive results even compared to the state-of-the-art (SOTA) frame-based method in action recognition tasks while only using 1.5 % of the computational resources. The model is adept at abstracting local and global geometry by hierarchy structure. By leveraging tensor-train compressed feature extractors, TTPOINT can be designed with minimal parameters and computational complexity. Additionally, we developed a straightforward downsampling algorithm to maintain the spatio-temporal feature. In the experiment, TTPOINT emerged as the SOTA method on three datasets while also attaining SOTA among point cloud methods on all five datasets. Moreover, by using the tensor-train decomposition method, the accuracy of the proposed TTPOINT is almost unaffected while compressing the parameter size by 55% in all five datasets.
Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Renjing Xu, Bojun Cheng
ACM Multimedia2