EDBT 2026 Demo / reviewers in the wild / expert
Yulong Huang 0001
dblp:06/10848-1
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-4261-775XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 35% Video understanding and tracking · 32% Efficient and distributed learning · 14% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
1.5 | 2 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023 |
Computer vision › Video understanding and tracking › action recognition
event-based action recognition |
1.5 | 2 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023 |
Computer vision › 3D vision › visual localization
camera relocalization |
1.0 | 2 | 2025 | A Simple and Effective Point-Based Network for Event Camera 6-DOFs Pose Relocalization · CVPR 2024 Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Image recognition and object detection › point set representation
point cloud representation |
0.9 | 1 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Deep learning architectures and training
spiking neural network |
0.9 | 1 | 2025 | ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring · ICCV 2025 |
Image and video processing
image restoration |
0.9 | 1 | 2025 | ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring · ICCV 2025 |
Image and video processing › image restoration › image deblurring
motion deblurring |
0.9 | 1 | 2025 | ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring · ICCV 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023 |
Computer vision › 3D vision › point cloud analysis
point cloud learning |
0.7 | 1 | 2023 | TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023 |
Computer vision › 3D vision › geometric deep learning
point cloud network |
0.7 | 1 | 2023 | TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023 |
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition |
0.7 | 1 | 2023 | TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
visual attention · 1.7spiking neural network · 1.7cross-modal fusion · 1.7state space model · 0.9mamba · 0.9hierarchical feature extraction · 0.9point-based network · 0.8hierarchical feature abstraction · 0.8attentive bidirectional LSTM · 0.8downsampling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Based Motion Deblurring via Multi-Temporal Granularity FusionabstractConventional frame-based cameras inevitably produce blurry effects due to motion occurring during the exposure time. Event camera, a bio-inspired sensor offering continuous visual information could enhance the deblurring performance. Effectively utilizing the high-temporal-resolution event data is crucial for extracting precise motion information and enhancing deblurring performance. However, existing event-based image deblurring methods usually utilize voxel-based event representations, losing the fine-grained temporal details that are mathematically essential for fast motion deblurring. In this paper, we first introduce point cloud-based event representation into the image deblurring task and propose a Multi-Temporal Granularity Network (MTGNet). It combines the spatially dense but temporally coarse-grained voxel-based event representation and the temporally fine-grained but spatially sparse point cloud-based event. To seamlessly integrate such complementary representations, we design a Fine-grained Point Branch. An Aggregation and Mapping Module (AMM) is proposed to align the low-level point-based features with frame-based features and an Adaptive Feature Diffusion Module (AFDM) is designed to manage the resolution discrepancies between event data and image data by enriching the sparse point feature. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art approaches on both synthetic and real-world datasets. Our code is available at: https://github.com/xplin13/MTGNet. Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Yue Zhou 0010, Haotian Fu, Biao Pan, Bojun Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ClearSight: Human Vision-Inspired Solutions for Event-Based Motion DeblurringabstractMotion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of event streams, we employ Spiking Neural Networks (SNNs) for motion feature extraction and Artificial Neural Networks (ANNs) for color information processing. Due to the non-uniform distribution and inherent redundancy of event data, existing cross-modal feature fusion methods exhibit certain limitations. Inspired by the visual attention mechanism in the human visual system, this study introduces a bioinspired dual-drive hybrid network (BDHNet). Specifically, the Neuron Configurator Module (NCM) is designed to dynamically adjusts neuron configurations based on cross-modal features, thereby focusing the spikes in blurry regions and adapting to varying blurry scenarios dynamically. Additionally, the Region of Blurry Attention Module (RBAM) is introduced to generate a blurry mask in an unsupervised manner, effectively extracting motion clues from the event features and guiding more accurate cross-modal feature fusion. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art methods on both synthetic and real-world datasets. Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Hongxiang Huang, Yue Zhou 0010, Haotian Fu, Bojun Cheng |
ICCV | 2 |
| 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and RegressionabstractEvent cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks. Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | A Simple and Effective Point-Based Network for Event Camera 6-DOFs Pose RelocalizationabstractEvent cameras exhibit remarkable attributes such as high dynamic range, asynchronicity, and low latency, making them highly suitable for vision tasks that involve highspeed motion in challenging lighting conditions. These cameras implicitly capture movement and depth information in events, making them appealing sensors for Camera Pose Relocalization (CPR) tasks. Nevertheless, existing CPR networks based on events neglect the pivotal finegrained temporal information in events, resulting in unsatisfactory performance. Moreover, the energy-efficient features are further compromised by the use of excessively complex models, hindering efficient deployment on edge devices. In this paper, we introduce PEPNet, a simple and effective point-based network designed to regress six degrees of freedom (6-DOFs) event camera poses. We rethink the relationship between the event camera and CPR tasks, leveraging the raw Point Cloud directly as network input to harness the high-temporal resolution and inherent sparsity of events. PEPNet is adept at abstracting the spatial and implicit temporal features through hierarchical structure and explicit temporal features by Attentive Bidirectional Long Short-Term Memory (A-Bi-LSTM). Byemploying a carefully crafted lightweight design, PEPNet delivers state-of-the-art (SOTA) performance on both indoor and outdoor datasets with meager computational resources. Specifically, PEPNet attains a significant 38% and 33% performance improvement on the random split IJRR and M3ED datasets, respectively. Moreover, the lightweight design version PEPNettinyaccomplishes results comparable to the SOTA while employing a mere 0.5% of the parameters. Jiadong Zhu, Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Bojun Cheng |
CVPR | 5 |
| 2024 | DS-CIM: A 40nm Asynchronous Dual-Spike Driven, MRAM Compute-In-Memory Macro for Spiking Neural NetworkabstractCompute-in-memory (CIM) based on emerging nonvolatile memory (eNVM) is an effective way to deploy neural networks to low-power edge devices for both storage and computation. NVMs such as ReRAM have been widely used in CIM. Meanwhile, MRAM has higher read and write cycles, lower device and cycle variation and a lower bit error rate, making it equally attractive for storage. However, the high read current and low on/off ratio result in large energy consumption in MRAM read limiting its large-scale application in CIM. The spiking neural network (SNN) represents the information as sparse spike sequences and facilitates hardware to achieve low-power computing by taking advantage of its spatial-temporal sparsity. To further increase the input sparsity of SNN and reduce the read energy consumption, this paper proposes ADC-free, dual-spike (DS) -CIM macro, a spiking MRAM CIM macro driven by asynchronous dual spikes. Compared to the conventional rate coding, our dual-spike coding method uses only 2 spikes to encode the information without losing accuracy. Moreover, the event-driven feature allows the macro to have sub-nW static power consumption. Our DS-CIM macro achieves comparable or higher accuracy while maintaining very low energy consumption. Specifically, it achieves accuracies of 96.99%, 82.87%, 90.00%, and 85.97% for digit classification, image classification, gesture recognition, and action recognition tasks, with energy consumption of only 8.07nJ, 71.26nJ, 729.3nJ, and 369.82nJ, respectively. These results emphasize the significance of DS-CIM and provide ideas for low-power inference on edge devices. Haotian Fu, Yulong Huang 0001, Tingran Chen, Chenyi Fu, Yue Zhou 0010, Shouzhong Peng, Zhirui Zong, Biao Pan, Bojun Cheng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasabstractEvent cameras have gained popularity in computer vision due to their data sparsity, high dynamic range, and low latency. As a bio-inspired sensor, event cameras generate sparse and asynchronous data, which is inherently incompatible with the traditional frame-based method. Alternatively, the point-based method can avoid additional modality transformation and naturally adapt to the sparsity of events. Still, it typically cannot reach a comparable accuracy as the frame-based method. We propose a lightweight and generalized point cloud network called TTPOINT which achieves competitive results even compared to the state-of-the-art (SOTA) frame-based method in action recognition tasks while only using 1.5 % of the computational resources. The model is adept at abstracting local and global geometry by hierarchy structure. By leveraging tensor-train compressed feature extractors, TTPOINT can be designed with minimal parameters and computational complexity. Additionally, we developed a straightforward downsampling algorithm to maintain the spatio-temporal feature. In the experiment, TTPOINT emerged as the SOTA method on three datasets while also attaining SOTA among point cloud methods on all five datasets. Moreover, by using the tensor-train decomposition method, the accuracy of the proposed TTPOINT is almost unaffected while compressing the parameter size by 55% in all five datasets. Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Renjing Xu, Bojun Cheng |
ACM Multimedia | 4 |