Bojun Cheng

dblp:285/0564 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-5551-4242ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Partner Project: Scalable, Ferroelectric-based Accelerators for Energy Efficient Edge AI (Ferro4EdgeAI)
abstract
The Computing-In-Memory (CIM) paradigm offers a promising solution to the memory-wall bottleneck that limits conventional Von Neumann architectures. By performing data processing at the same physical location where the data are stored, CIM-based architectures minimize costly data movement and drastically improve energy efficiency. When implemented with Ferroelectric Field Effect Transistors (FeFETs), additional advantages from the non-volatility, fast switching, and low operating voltage of FeFETs are added. However, the widespread adoption of FeFETs is limited by their poor endurance, which is overcome by a Back End of the Line (BEoL) integration of FeFET-2, where a ferroelectric capacitor (FeCAP) is wired to the gate of a CMOS transistor providing high endurance compatible with low-power edge applications. These properties enable dense, low-power, and high-speed matrix operations essential for AI workloads. As a result, FeFET-2-based CIM accelerators offer a promising solution for energy-efficient, high-performance AI at the edge. The Ferro4EdgeAI project aims to develop an ultra low-power, scalable edge accelerator for AI, targeting a significant gain in energy efficiency with respect to state-of-the-art AI hardware accelerators. To attain this, our project focuses on innovation all along the value chain from materials, physic concepts, device architecture, integration technologies, and accelerators in a holistic design space exploration approach.
Theofilos Spyrou, Yashvardhan Biyani, Konstantinos Stavrakakis, Rajendra Bishnoi, Said Hamdioui, Joel Minguet Lopez, Louise Dumas, Jean Coignus, Denys Ly, Hugo Chazot-Ranquet, Laurent Grenouillet, Fabien Grimaud, Simon Martin 0006, Olivier Billoint, François Andrieu, Ruben Alcala, Stefan Slesazeck, Athira Sunil, Antoine Cauquil, Rosario Pronsat, Damien Deleruyelle, Cédric Marchand 0002, Alberto Bosio, Ian O'Connor, Giulio Urlini, Simon Jeannot, Mohammad Sajedi Alvar, Nima Akbari Moghaddam, Thilo Werner, Tony Schenk, Bojun Cheng, Mina Khoei, Lucía Pérez Ramírez, EunJin Koh, Somnath Kale, Nicholas Barrett
DATE32
2026 Event-Based Motion Deblurring via Multi-Temporal Granularity Fusion
abstract
Conventional frame-based cameras inevitably produce blurry effects due to motion occurring during the exposure time. Event camera, a bio-inspired sensor offering continuous visual information could enhance the deblurring performance. Effectively utilizing the high-temporal-resolution event data is crucial for extracting precise motion information and enhancing deblurring performance. However, existing event-based image deblurring methods usually utilize voxel-based event representations, losing the fine-grained temporal details that are mathematically essential for fast motion deblurring. In this paper, we first introduce point cloud-based event representation into the image deblurring task and propose a Multi-Temporal Granularity Network (MTGNet). It combines the spatially dense but temporally coarse-grained voxel-based event representation and the temporally fine-grained but spatially sparse point cloud-based event. To seamlessly integrate such complementary representations, we design a Fine-grained Point Branch. An Aggregation and Mapping Module (AMM) is proposed to align the low-level point-based features with frame-based features and an Adaptive Feature Diffusion Module (AFDM) is designed to manage the resolution discrepancies between event data and image data by enriching the sparse point feature. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art approaches on both synthetic and real-world datasets. Our code is available at: https://github.com/xplin13/MTGNet.
Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Yue Zhou 0010, Haotian Fu, Biao Pan, Bojun Cheng
IEEE Trans. Circuits Syst. Video Technol.8
2025 ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring
abstract
Motion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of event streams, we employ Spiking Neural Networks (SNNs) for motion feature extraction and Artificial Neural Networks (ANNs) for color information processing. Due to the non-uniform distribution and inherent redundancy of event data, existing cross-modal feature fusion methods exhibit certain limitations. Inspired by the visual attention mechanism in the human visual system, this study introduces a bioinspired dual-drive hybrid network (BDHNet). Specifically, the Neuron Configurator Module (NCM) is designed to dynamically adjusts neuron configurations based on cross-modal features, thereby focusing the spikes in blurry regions and adapting to varying blurry scenarios dynamically. Additionally, the Region of Blurry Attention Module (RBAM) is introduced to generate a blurry mask in an unsupervised manner, effectively extracting motion clues from the event features and guiding more accurate cross-modal feature fusion. Extensive subjective and objective evaluations demonstrate that our method outperforms current state-of-the-art methods on both synthetic and real-world datasets.
Xiaopeng Lin, Yulong Huang 0001, Zunchang Liu, Hongxiang Huang, Yue Zhou 0010, Haotian Fu, Bojun Cheng
ICCV8
2025 E2B: A Single Modality Point-Based Tracker with Event Cameras
abstract
High-speed object tracking holds significant relevance across robotic domains, such as drones and autonomous driving. Compared to conventional cameras, event cameras are equipped with the ability to capture object motion information at exceptionally high temporal resolution with relatively low power consumption and remain immune from motion-blurring effects. Regrettably, many existing methods adopt a framebased approach by stacking events into Event Frame, which overlooks the sparsity and high temporal resolution of events. This approach is also reliant on the huge pre-training backbone and reaches a performance plateau but demands unrealistically large networks and high power consumption, rendering it impractical for real-time applications in battery-constrained robotic scenarios. In this paper, we propose an efficient and effective single-modality tracker using Point Cloud representation named E2B (Event to Box). By directly handling the raw output of event cameras without dataformat transformation, E2B leverages events' coordinate guidance to accurately map Event Cloud features to 2D bounding boxes. Moreover, E2B incorporates the pyramid structure into the multi-stage feature extraction architecture to effectively track objects across diverse scales. In the experiments, E2B performs outstandingly on two large-scale and one synthetic event-based tracking datasets, covering both indoor and outdoor environments, as well as rigid and non-rigid objects.
Aiersi Tuerhong, Haobo Liu, Yongxiang Feng, Wenhui Wang 0001, Yaoyuan Wang, Weihua He, Bojun Cheng
ICRA11
2025 An FPGA Processor Combining Point Cloud and SNN for DVS-based ADAS Application
abstract
Automatic Emergency Braking (AEB) has become an important component in Advanced Driver Assistance Systems (ADAS) and a potential solution for AEB lies in the integration of Dynamic Vision Sensor (DVS) with Spiking Neural Network (SNN). A high-precision behavioural recognition algorithm called Spikepoint has been proposed by us, which combines Point Cloud with SNN to enable recognition of DVS event data. This work concentrates on the FPGA implementation of Spikepoint, aiming to improve real-time recognition capabilities. The deployment of Spikepoint on FPGA encounters two challenges: 1) Point Cloud processing introduces additional latency 2) Storing parameters that require to be accessed frequently from DDR introduces a significant time overhead. In order to address challenges aforementioned, a novel reference point-based filtering technique for Point Cloud is introduced. Meanwhile, a fine-grained quantization method and other optimization strategies are used on the neuron model. The Xilinx UltraScale+ is employed in the experiments conducted in this work. Our Point-based SNN Processor achieves a recognition frame rate of 92.08 FPS through the novel algorithm and corresponding hardware optimization, while achieving an accuracy of 94.3% on the DVS128 Gesture dataset.
Wente Yi, Kexun Cheng, Lehao Tan, Bojun Cheng, Biao Pan
ISCAS8
2025 Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression
abstract
Event cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks.
Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng
IEEE Trans. Pattern Anal. Mach. Intell.10
2024 A Simple and Effective Point-Based Network for Event Camera 6-DOFs Pose Relocalization
abstract
Event cameras exhibit remarkable attributes such as high dynamic range, asynchronicity, and low latency, making them highly suitable for vision tasks that involve highspeed motion in challenging lighting conditions. These cameras implicitly capture movement and depth information in events, making them appealing sensors for Camera Pose Relocalization (CPR) tasks. Nevertheless, existing CPR networks based on events neglect the pivotal finegrained temporal information in events, resulting in unsatisfactory performance. Moreover, the energy-efficient features are further compromised by the use of excessively complex models, hindering efficient deployment on edge devices. In this paper, we introduce PEPNet, a simple and effective point-based network designed to regress six degrees of freedom (6-DOFs) event camera poses. We rethink the relationship between the event camera and CPR tasks, leveraging the raw Point Cloud directly as network input to harness the high-temporal resolution and inherent sparsity of events. PEPNet is adept at abstracting the spatial and implicit temporal features through hierarchical structure and explicit temporal features by Attentive Bidirectional Long Short-Term Memory (A-Bi-LSTM). Byemploying a carefully crafted lightweight design, PEPNet delivers state-of-the-art (SOTA) performance on both indoor and outdoor datasets with meager computational resources. Specifically, PEPNet attains a significant 38% and 33% performance improvement on the random split IJRR and M3ED datasets, respectively. Moreover, the lightweight design version PEPNettinyaccomplishes results comparable to the SOTA while employing a mere 0.5% of the parameters.
Jiadong Zhu, Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Bojun Cheng
CVPR6
2024 SpikePoint: An Efficient Point-based Spiking Neural Network for Event Cameras Action Recognition
abstract
Event cameras are bio-inspired sensors that respond to local changes in light intensity and feature low latency, high energy efficiency, and high dynamic range. Meanwhile, Spiking Neural Networks (SNNs) have gained significant attention due to their remarkable efficiency and fault tolerance. By synergistically harnessing the energy efficiency inherent in event cameras and the spike-based processing capabilities of SNNs, their integration could enable ultra-low-power application scenarios, such as action recognition tasks. However, existing approaches often entail converting asynchronous events into conventional frames, leading to additional data mapping efforts and a loss of sparsity, contradicting the design concept of SNNs and event cameras. To address this challenge, we propose SpikePoint, a novel end-to-end point-based SNN architecture. SpikePoint excels at processing sparse event cloud data, effectively extracting both global and local features through a singular-stage structure. Leveraging the surrogate training method, SpikePoint achieves high accuracy with few parameters and maintains low power consumption, specifically employing the identity mapping feature extractor on diverse datasets. SpikePoint achieves state-of-the-art (SOTA) performance on four event-based action recognition datasets using only 16 timesteps, surpassing other SNN methods. Moreover, it also achieves SOTA performance across all methods on three datasets, utilizing approximately 0.3 % of the parameters and 0.5 % of power consumption employed by artificial neural networks (ANNs). These results emphasize the significance of Point Cloud and pave the way for many ultra-low-power event-based data processing applications.
Xiaopeng Lin, Haotian Fu, Bojun Cheng
ICLR7
2024 CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are promising brain-inspired energy-efficient models. Compared to conventional deep Artificial Neural Networks (ANNs), SNNs exhibit superior efficiency and capability to process temporal information. However, it remains a challenge to train SNNs due to their undifferentiable spiking mechanism. The surrogate gradients method is commonly used to train SNNs, but often comes with an accuracy disadvantage over ANNs counterpart. We link the degraded accuracy to the vanishing of gradient on the temporal dimension through the analytical and experimental study of the training process of Leaky Integrate-and-Fire (LIF) Neuron-based SNNs. Moreover, we propose the Complementary Leaky Integrate-and-Fire (CLIF) Neuron. CLIF creates extra paths to facilitate the backpropagation in computing temporal gradient while keeping binary output. CLIF is hyperparameter-free and features broad applicability. Extensive experiments on a variety of datasets demonstrate CLIF's clear performance advantage over other neuron models. Furthermore, the CLIF's performance even slightly surpasses superior ANNs with identical network structure and training conditions. The code is available at https://github.com/HuuYuLong/Complementary-LIF.
Xiaopeng Lin, Haotian Fu, Zunchang Liu, Biao Pan, Bojun Cheng
ICML8
2024 PipeCIM: A High-Throughput Computing-In-Memory Microprocessor With Nested Pipeline and RISC-V Extended Instructions
abstract
The large number of multiply accumulate (MAC) operations in Convolutional Neural Network (CNN) leads to substantial data migration and computation. Although computing-in-memory (CIM) proves to be a promising paradigm for MAC operations, high throughput CNN accelerator still confronts bottlenecks from: the low MAC utilization and the uncessary off-chip memory access. In this paper, we propose a high throughput CIM-based CNN accelerator PipeCIM with three hierarchies of pipelines: Intra-Macro, Near-Memory and Tile-Level. The Intra-Macro Pipeline parallelly executes data transfer and in-memory-computing (IMC) operations. The Near-Memory Pipeline alleviates memory access for pooling and data reshaping. The Tile-Level Pipeline establishes a layer-wise pipeline to further improve the throughput while reducing control complexity. PipeCIM introduces the nested scheme and a Unidirectional Divergent Connection Protocol (UDTCP) to simplify the control of data flow with the help of customized RISC-V instructions. To validate our design, PipeCIM was prototyped in 55 nm process node, achieving energy efficiency of 133.8 TOPS/W and peak throughput of 819 GOPS with a 16KB CIM array, which can accelerate VGG-16 to 128.56$\times$or Inception to 19.754$\times$compared to the baseline.
Tingran Chen, Wenjia Wang 0011, Haotian Fu, Wente Yi, Bojun Cheng, He Zhang 0011, Biao Pan
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 DS-CIM: A 40nm Asynchronous Dual-Spike Driven, MRAM Compute-In-Memory Macro for Spiking Neural Network
abstract
Compute-in-memory (CIM) based on emerging nonvolatile memory (eNVM) is an effective way to deploy neural networks to low-power edge devices for both storage and computation. NVMs such as ReRAM have been widely used in CIM. Meanwhile, MRAM has higher read and write cycles, lower device and cycle variation and a lower bit error rate, making it equally attractive for storage. However, the high read current and low on/off ratio result in large energy consumption in MRAM read limiting its large-scale application in CIM. The spiking neural network (SNN) represents the information as sparse spike sequences and facilitates hardware to achieve low-power computing by taking advantage of its spatial-temporal sparsity. To further increase the input sparsity of SNN and reduce the read energy consumption, this paper proposes ADC-free, dual-spike (DS) -CIM macro, a spiking MRAM CIM macro driven by asynchronous dual spikes. Compared to the conventional rate coding, our dual-spike coding method uses only 2 spikes to encode the information without losing accuracy. Moreover, the event-driven feature allows the macro to have sub-nW static power consumption. Our DS-CIM macro achieves comparable or higher accuracy while maintaining very low energy consumption. Specifically, it achieves accuracies of 96.99%, 82.87%, 90.00%, and 85.97% for digit classification, image classification, gesture recognition, and action recognition tasks, with energy consumption of only 8.07nJ, 71.26nJ, 729.3nJ, and 369.82nJ, respectively. These results emphasize the significance of DS-CIM and provide ideas for low-power inference on edge devices.
Haotian Fu, Yulong Huang 0001, Tingran Chen, Chenyi Fu, Yue Zhou 0010, Shouzhong Peng, Zhirui Zong, Biao Pan, Bojun Cheng
IEEE Trans. Circuits Syst. I Regul. Pap.10
2024 RDCIM: RISC-V Supported Full-Digital Computing-in-Memory Processor With High Energy Efficiency and Low Area Overhead
abstract
Digital computing-in-memory (DCIM) that merges computing logic into memory has been proven to be an efficient architecture for accelerating multiply-and-accumulates (MACs). However, low energy efficiency and high area overhead pose a primary restriction for integrating DCIM in re-configurable processors required for multi-functional workloads. To alleviate this dilemma, a novel RISC-V supported full-digital computing-in-memory processor (RDCIM) is designed and fabricated with 55nm CMOS technology. In RDCIM, an adding-on-memory-boundary (AOMB) scheme is adopted to improve the energy efficiency of DCIM. Meanwhile, a multi-precision adaptive accumulator (MPAA) and a serial-parallel conversion supported SRAM buffer (SPBUF) are employed to reduce the area overhead caused by the peripheral circuits and the intermediate buffer for multi-precision support. The results show that the energy efficiency in our design is 16.6 TOPS/W (8-bit) and 66.3 TOPS/W (4-bit). Compared to related works, the proposed RDCIM macro shows a maximum energy efficiency improvement of 1.22$\times$in a continuous computing scenario, an area saving of 1.22$\times$in the accumulator, and an area saving of 3.12$\times$in the input buffer. Moreover, in RDCIM, 5 fine-grained RISC-V extended instructions are designed to dynamically adjust the state of DCIM, reaching 1.2$\times$computation efficiency.
Wente Yi, Kefan Mo, Wenjia Wang 0011, Yejun Zeng, Zihan Yuan, Bojun Cheng, Biao Pan
IEEE Trans. Circuits Syst. I Regul. Pap.7
2023 TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event Cameras
abstract
Event cameras have gained popularity in computer vision due to their data sparsity, high dynamic range, and low latency. As a bio-inspired sensor, event cameras generate sparse and asynchronous data, which is inherently incompatible with the traditional frame-based method. Alternatively, the point-based method can avoid additional modality transformation and naturally adapt to the sparsity of events. Still, it typically cannot reach a comparable accuracy as the frame-based method. We propose a lightweight and generalized point cloud network called TTPOINT which achieves competitive results even compared to the state-of-the-art (SOTA) frame-based method in action recognition tasks while only using 1.5 % of the computational resources. The model is adept at abstracting local and global geometry by hierarchy structure. By leveraging tensor-train compressed feature extractors, TTPOINT can be designed with minimal parameters and computational complexity. Additionally, we developed a straightforward downsampling algorithm to maintain the spatio-temporal feature. In the experiment, TTPOINT emerged as the SOTA method on three datasets while also attaining SOTA among point cloud methods on all five datasets. Moreover, by using the tensor-train decomposition method, the accuracy of the proposed TTPOINT is almost unaffected while compressing the parameter size by 55% in all five datasets.
Yue Zhou 0010, Haotian Fu, Yulong Huang 0001, Renjing Xu, Bojun Cheng
ACM Multimedia6