EDBT 2026 Demo / reviewers in the wild / expert
De Ma
dblp:18/8568
· DBLP profile ↗
42ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0001-8700-938XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Systems, architecture and hardware · 13 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SNN-Driven Event-Based Flow and Rotation Estimation with SO(3) RefinementabstractSpiking Neural Networks (SNNs) offer a promising direction for energy-efficient event-based vision by leveraging sparse, temporally precise spikes. We propose a directly trained, fully spiking model for optical flow estimation, featuring a novel Spike GRU and membrane potential carryover for improved temporal modeling. On the DSEC-Flow benchmark, our model achieves competitive accuracy while reducing energy consumption by 42.88× over EV-FlowNet and 38× over TIDNet. Building on the predicted motion field, we infer camera rotation and, to the best of our knowledge, are the first to construct panoramic event images from SNN-based flow. We further introduce an optional unsupervised SO(3) refinement step that improves rotation accuracy by maximizing panorama consistency—without IMU or pose supervision. Our results achieve comparable visual quality to CMax-SLAM, showing that SNNs can enable fast and high-level spatial perception using only event-based input. Ruimin Sun, De Ma |
AAAI | 3 |
| 2026 | S³: Spiking Neurons as an Isolating Segmenter for Brain Signal DecodingabstractRecent brain decoding studies have primarily emphasized the development of brain decoders, while largely neglecting the segmentation step. Existing methods typically adopt fixed-length segmentation, which might overlook subject- or task-level variability and disrupt temporal patterns within brain signals. To address this gap, we propose S3, which leverages spiking neurons as an isolating segmenter for brain signal decoding. S3 segments brain signals adaptively, considering subject- and task-level variability while preserving intrinsic temporal patterns of brain signals. It exploits the unique reset mechanism of spiking neurons to isolate previous irrelevant temporal patterns during the generation of each segmentation point. To optimize S3 for enhancing task performance in the absence of segmentation labels, we develop an optimization method where segmentation pseudo-labels are created with a stochastic-greedy algorithm to optimize them, while circumventing gradient blockade between S3 and task performance. Experiments on 10 downstream tasks across 13 public datasets demonstrate that S3 consistently outperforms existing methods, validating its effectiveness, generalizability and interpretability. Sha Zhao, Shi Gu, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 6 |
| 2026 | SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language ModelsabstractMing Chen, Wenyao Li, Chao Liang, Shi Gu, Peng Lin, De Ma, Huajin Tang, Qian Zheng, Gang Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shi Gu, De Ma, Huajin Tang, Gang Pan 0001 |
ACL (1) | 6 |
| 2026 | A Lightweight Die-to-Die Interconnect with Co-Scheduling Framework for Scalable Multi-Chip Inference Accelerators
De Ma |
ISCAS | 3 |
| 2026 | DepAsync: An Asynchronous SNN Accelerator Based on Core-DependencyabstractSpiking Neural Networks (SNNs) are widely used in brain-inspired computing and neuroscience research. Several many-core accelerators have been built to improve the running speed and energy efficiency of SNNs. However, current accelerators generally need explicit synchronization among all cores after each timestep of SNNs, which poses a challenge to overall efficiency. This paper proposes DepAsync, an asynchronous architecture that eliminates inter-core synchronization, facilitating fast and energy-efficient SNN inference with commendable scalability. The main idea is to exploit the dependency of neuromorphic cores predetermined at compile time. We design a DepAsync scheduler for each core to trace the running state of its dependencies and control the core to safely forward to the next timestep without waiting for other cores to complete their tasks. This approach prevents the necessity for global synchronization, allowing DepAsync to minimize core waiting time facing inherent core and time imbalance in SNN workloads. The comprehensive evaluations using five SNN workloads show that DepAsync achieves 2.47x speedup and 1.55x energy efficiency compared to the state-of-the-art synchronization architectures. Zhuo Chen 0044, De Ma, Xiaofei Jin, Qinghui Xing, Ouwen Jin, Xin Du 0002, Shuibing He, Gang Pan 0001 |
IEEE Trans. Computers | 2 |
| 2026 | NPEva: An NoC-Based Neuromorphic Processors Performance Evaluation Framework for Benchmarking Spiking Neural NetworksabstractSpiking Neural Network (SNN) applications place diverse demands on neuromorphic processors’ computational and communication capabilities. Specifically, the sparse parallel computation of SNNs requires storage systems capable of extensive parallel data access and processing. Additionally, the time-dependent nature of SNN computations demands a communication framework to manage spike transmission and synchronization. Addressing these requires continuous early-stage evaluation of potential designs, which can be time-intensive. To this end, this paper introduces a performance evaluation model for NoC-based neuromorphic processors, named NPEva, which includes a Computation Model (CpMo) and a Communication Model (CoMo). This model facilitates rapid, high-dimensional exploration of design spaces across a wide range of microarchitectural parameters. CpMo quickly assesses computation-related latency, power consumption, and area by extracting relevant parameters from hierarchical storage organizations and neuron computation models. CoMo employs communication upper-bound latency analysis to evaluate packet latencies to reduce redundant synchronization time. Comparisons with TrueNorth’s publicly released data reveal that the model’s evaluation results are within an 8% error margin. The CoMo offers enhanced precision in estimating latency upper-bounds compared to other models like worst-contention delay (WCD) and worst-case traversal time (WCTT), especially when varying the number of VCs. With VC settings of 2, 4, 8, and 16, CoMo reduced latency by 0.55× to 2.53× compared to WCTT and by 0.83× to 25.57× compared to WCD. Additionally, using the NPEva model for exploring the design space of Liquid State Machines (LSM) networks has shown that optimized hardware can improve cost efficiency by 2.94× with only a 1% increase in latency. Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | EvHDR-GS: Event-guided HDR Video Reconstruction with 3D Gaussian SplattingabstractHigh Dynamic Range (HDR) video reconstruction seeks to accurately restore the extensive dynamic range present in real-world scenes and is widely employed in downstream applications. Existing methods typically operate on one or a small number of consecutive frames, which often leads to inconsistent brightness across the video due to their limited perspective on the video sequence. Moreover, supervised learning-based approaches are susceptible to data bias, resulting in reduced effectiveness when confronted with test inputs exhibiting a domain gap relative to the training data. To address these limitations, we present an event-guided HDR video reconstruction method through building 3D Gaussian Splatting (3DGS), to ensure consistent brightness imposed by 3D consistency. We introduce HDR 3D Gaussians capable of simultaneously representing HDR and low-dynamic-range (LDR) colors. Furthermore, we incorporate a learnable HDR-to-LDR transformation optimized by input event streams and LDR frames to eliminate the data bias. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method achieves state-of-the-art performance. Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001 |
AAAI | 3 |
| 2025 | EvHDR-NeRF: Building High Dynamic Range Radiance Fields with Single Exposure Images and EventsabstractWe present EvHDR-NeRF to recover a High Dynamic Range (HDR) radiance field from event streams and a set of Low Dynamic Range (LDR) views with single exposures. Using the EvHDR-NeRF, we can generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the new relationship between events streams and LDR images, which considers both the Camera Response Function (CRF) and exposure time. Based on this relationship, we categorize events into inter-frame events and intra-exposure. The former is utilized for building HDR radiance field and the latter is used to deblur potentially blurred images. Compared to existing methods, this method can effectively reconstruct the HDR radiance field even when the input images are degraded. Experimental results demonstrate that our method achieves state-of-the-art HDR reconstruction, providing a more adaptable and accurate solution for complex imaging applications. Zhanfeng Liao, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 3 |
| 2025 | EvSTVSR: Event Guided Space-Time Video Super-ResolutionabstractIn the domain of space-time video super-resolution, it is typically challenging to handle complex motions (including large and nonlinear motions) and varying illumination scenes due to the lack of inter-frame information. Leveraging the dense temporal information provided by event signals offers a promising solution. Traditional event-based methods typically rely on multiple images, using motion estimation and compensation, which can introduce errors. Accumulated errors from multiple frames often lead to artifacts and blurriness in the output. To mitigate these issues, we propose EvSTVSR, a method that uses fewer adjacent frames and integrates dense temporal information from events to guide alignment. Additionally, we introduce a coordinate-based feature fusion upsampling module to achieve spatial super-resolution. Experimental results demonstrate that our method not only outperforms existing RGB-based approaches but also excels in handling large motion scenarios. Haojie Yan, Zhan Lu, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 4 |
| 2025 | VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision MakingabstractRecent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm.We show that this paradigm fundamentally limits performance in multi-input multi-output (MIMO) scenarios, where parallel task execution is required.In MISO architectures, tasks compete for a shared output channel, creating mutual exclusion effects that cause unbalanced optimization and degraded performance.To address this gap, we introduce MIMO-VLA (VLASCD), a unified training framework that enables concurrent multi-task outputs, exemplified by simultaneous dialogue generation and decision-making.Inspired by human cognition, MIMO-VLA eliminates interference between tasks and supports efficient parallel processing.Experiments on the CARLA autonomous driving platform demonstrate that MIMO-VLA substantially outperforms stateof-the-art MISO-based LLMs, reinforcement learning models, and VLAs in MIMO settings, establishing a new direction for multimodal and multitask learning.Our code is available at: Zuojin Tang, De Ma, Gang Pan 0001 |
EMNLP | 4 |
| 2025 | E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic Scenes
Haojie Yan, De Ma, Huajin Tang, Gang Pan 0001 |
ICCV | 4 |
| 2025 | Training High Performance Spiking Neural Network by Temporal Model CalibrationabstractSpiking Neural Networks (SNNs) are considered promising energy-efficient models due to their dynamic capability to process spatial-temporal spike information. Existing work has demonstrated that SNNs exhibit temporal heterogeneity, which leads to diverse outputs of SNNs at different time steps and has the potential to enhance their performance. Although SNNs obtained by direct training methods achieve state-of-the-art performance, current methods introduce limited temporal heterogeneity through the dynamics of spiking neurons or network structures. They lack the improvement of temporal heterogeneity through the lens of the gradient. In this paper, we first conclude that the diversity of the temporal logit gradients in current methods is limited. This leads to insufficient temporal heterogeneity and results in temporally miscalibrated SNNs with degraded performance. Based on the above analysis, we propose a Temporal Model Calibration (TMC) method, which can be seen as a logit gradient rescaling mechanism across time steps. Experimental results show that our method can improve the temporal logit gradient diversity and generate temporally calibrated SNNs with enhanced performance. In particular, our method achieves state-of-the-art accuracy on ImageNet, DVSCIFAR10, and N-Caltech101. Codes are available at https://github.com/zju-bmi-lab/TMC. Changping Wang, De Ma, Huajin Tang, Gang Pan 0001 |
ICML | 3 |
| 2025 | HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot NavigationabstractReinforcement Learning (RL) has shown promise in robotic navigation tasks, yet applying it to real-world environments remains challenging due to dynamic complexities and the need for dynamically feasible actions. We propose a hierarchical control framework based on Spiking Deep Reinforcement Learning (SDRL) for robust robot navigation in real environments. Our approach utilizes a two-layer architecture: a high-level decision layer powered by a Spiking GRU network for handling partially observable environments, and a low-level executive layer employing Continuous Attractor Neural Networks (CANNs) to ensure precise and continuous actions. This hierarchical structure allows real-time decisionmaking that respects the physical constraints of the robot. Experimental results show that our method adapts effectively to new environments without fine-tuning and surpasses existing methods in performance. We also explore the implementation on the Darwin3 chip, paving the way for biologically inspired motion control in future robotic applications. Shibo Zhou, Chaohui Lin, Qingao Chai, Rui Yan 0005, De Ma, Gang Pan 0001, Huajin Tang |
ICRA | 6 |
| 2025 | Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
Lang Feng 0002, Dong Xing, Li Zhang 0045, De Ma, Gang Pan 0001 |
AAMAS | 5 |
| 2025 | EDyGS: Event Enhanced Dynamic 3D Radiance Fields from Blurry Monocular VideoabstractThe task of generating novel views in dynamic scenes plays a critical role in the 3D vision domain. Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) have shown great promise in this domain but struggle with motion blur, which often arises in real-world scenarios due to camera or object motion. Existing methods address camera motion blur but fall short in dynamic scenes, where the coupling of camera and object motion complicates multi-view consistency and temporal coherence. In this work, we propose EDyGS, a model designed to reconstruct sharp novel views from event streams and monocular videos of dynamic scenes with motion blur. Our approach introduces a motion-mask 3D Gaussian model that assigns each Gaussian an additional attribute to distinguish between static and dynamic regions. By leveraging this motion mask field, we separate and optimize the static and dynamic regions independently. A progressive learning strategy is adopted, where static regions are reconstructed by jointly optimizing camera poses and learnable 3D Gaussians, while dynamic regions are modeled using an implicit deformation field alongside learnable 3D Gaussians. We conduct both quantitative and qualitative experiments on synthetic and real-world data. Experimental results demonstrate that EDyGS effectively handles blurry inputs in dynamic scenes. Mengxu Lu, De Ma, Huajin Tang, Gang Pan 0001 |
IJCAI | 4 |
| 2025 | An SRAM-Based Digital Compute-in-Memory Macro with Dual-Bit Input Data Sparsification and Restructuring for Energy-Efficient MACabstractSRAM-based digital computing-in-memory (CIM) provides a low-power, high-precision method for multiply-and-accumulate (MAC) operations. However, input data with a high toggle rate and low sparsity will significantly increase the energy consumption of digital CIMs. This work presents an input-transforming and weight-precomputation (ITWP) SRAM-based digital CIM macro. Two strategies involving input sparsification and input restructuring are proposed to transform the dual-bit input data, resulting in higher sparsity and lower toggle rate. The sum of the weights is precomputed to correct the partial sum (PSUM) of the sparsified input data. A 13T-based hybrid SRAM CIM unit (HSCU) and a split adder tree (SAT) are proposed to compute on the restructured two input sets with different sparsity levels, leading to a higher throughput and energy efficiency. Finally, the ITWP CIM macro is implemented in a 28 nm process and achieves a peak throughput of 0.935 TOPS and an energy efficiency of 20.08-33.78 TOPS/W with signed 8b input and weight. Compared to the conventional bit-serial CIM macro, the ITWP CIM macro is demonstrated to increase the energy efficiency to 1.42x and up to 3.04x. De Ma, Zhiping Yu |
ISCAS | 2 |
| 2025 | An NoC-Based Latency Upper-Bound Model for Reducing Timestep Length in SNN CommunicationabstractSpiking Neural Networks (SNNs) with real-time demands are deployed on Network-on-Chip (NoC)-based neuromorphic processors for specific tasks. While meeting real-time constraints for spiking data streams is crucial, overly long timesteps (e.g., 1ms) result in idle neuron cores and routers, reducing efficiency. This paper proposes a communication performance model using a recursive calculation method to assess the worst-case upper-bound latency. Integrated with the NoC router microarchitecture, the model analyzes the behavior of spiking streams during communication. It effectively balances timestep length to ensure real-time communication while minimizing idle time. Empirical results show that our model outperforms worst-contention delay (WCD) and worst-case traversal time (WCTT) models, achieving latency reductions of 0.55× to 2.53× compared to WCTT and 0.83× to 25.57× compared to WCD across various spiking datasets and virtual channels. Encouragingly, with only a 1% accuracy reduction, latency was reduced by 18%. Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001 |
ISCAS | 4 |
| 2025 | Training multi-bit Spiking Neural Network with Virtual Neurons
Zonghua Gu 0001, Ruimin Sun, De Ma |
Neurocomputing | 4 |
| 2025 | Post-training quantization for efficient ANN-SNN conversion
Ruimin Sun, De Ma, Gang Pan 0001 |
Neural Networks | 2 |
| 2025 | Mapping Large-Scale Spiking Neural Network on Arbitrary Meshed Neuromorphic HardwareabstractNeuromorphic hardware systems—designed as 2D-mesh structures with parallel neurosynaptic cores—have proven highly efficient at executing large-scale spiking neural networks (SNNs). A critical challenge, however, lies in mapping neurons efficiently to these cores. While existing approaches work well with regular, fully functional mesh structures, they falter in real-world scenarios where hardware has irregular shapes or non-functional cores caused by defects or resource fragmentation. To address these limitations, we propose a novel mapping method based on an innovative space-filling curve: the Adaptive Locality-Preserving (ALP) curve. Using a unique divide-and-conquer construction algorithm, the ALP curve ensures adaptability to meshes of any shape while maintaining crucial locality properties—essential for efficient mapping. Our method demonstrates exceptional computational efficiency, making it ideal for large-scale deployments. These distinctive characteristics enable our approach to handle complex scenarios that challenge conventional methods. Experimental results show that our method matches state-of-the-art solutions in regular-shape mapping while achieving significant improvements in irregular scenarios, reducing communication overhead by up to 57.1%. Ouwen Jin, Qinghui Xing, Zhuo Chen 0044, Ming Zhang 0018, De Ma, Ying Li 0001, Xin Du 0002, Shuibing He, Shuiguang Deng, Gang Pan 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | HetSub: A Heterogeneous Multi-NoC With Reconfigurable Long-Range Links for Neuromorphic Systems
Youneng Hu, Xiaofei Jin, Ziyang Kang, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Learning to Manipulate Artistic ImagesabstractRecent advancement in computer vision has significantly lowered the barriers to artistic creation. Exemplar-based image translation methods have attracted much attention due to flexibility and controllability. However, these methods hold assumptions regarding semantics or require semantic information as the input, while accurate semantics is not easy to obtain in artistic images. Besides, these methods suffer from cross-domain artifacts due to training data prior and generate imprecise structure due to feature compression in the spatial domain. In this paper, we propose an arbitrary Style Image Manipulation Network (SIM-Net), which leverages semantic-free information as guidance and a region transportation strategy in a self-supervised manner for image generation. Our method balances computational efficiency and high resolution to a certain extent. Moreover, our method facilitates zero-shot style image manipulation. Both qualitative and quantitative experiments demonstrate the superiority of our method over state-of-the-art methods.Code is available at https://github.com/SnailForce/SIM-Net. De Ma |
AAAI | 3 |
| 2024 | EAS-SNN: End-to-End Adaptive Sampling and Representation for Event-Based Detection with Recurrent Spiking Neural Networks
Ziling Wang, Huaning Li, Runhao Jiang, De Ma, Huajin Tang |
ECCV (60) | 6 |
| 2024 | Event-ID: Intrinsic Decomposition Using an Event CameraabstractReconstructing 3D scenes from multi-view images is challenging, especially under extreme scenarios. We propose Event-ID, an event-based intrinsic decomposition framework that leverages events and images for stable decomposition under extreme scenarios. Our method is based on two observations: event cameras maintain good imaging quality under blurry or poorly exposed scenarios, and event signals from different viewpoints exhibit similarity in diffuse regions while varying in specular regions. We establish an event-based reflectance model and introduce an event-based warping method to extract specular clues. Our two-stage framework constructs a radiance field and decomposes the scene into normal, material, and lighting. Experimental results demonstrate superior performance compared to state-of-the-art methods. Our project can be found at https://zehaoc.github.io/EventID.github.io/ Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001 |
ACM Multimedia | 3 |
| 2024 | RSNN: Recurrent Spiking Neural Networks for Dynamic Spatial-Temporal Information Processing
Qi Xu 0008, Xuanye Fang, Jiangrong Shen, De Ma, Yi Xu 0008, Gang Pan 0001 |
ACM Multimedia | 5 |
| 2024 | FEEL-SNN: Robust Spiking Neural Networks with Frequency Encoding and Evolutionary Leak FactorabstractCurrently, researchers think that the inherent robustness of spiking neural networks (SNNs) stems from their biologically plausible spiking neurons, and are dedicated to developing more bio-inspired models to defend attacks. However, most work relies solely on experimental analysis and lacks theoretical support, and the direct-encoding method and fixed membrane potential leak factor they used in spiking neurons are simplified simulations of those in the biological nervous system, which makes it difficult to ensure generalizability across all datasets and networks. Contrarily, the biological nervous system can stay reliable even in a highly complex noise environment, one of the reasons is selective visual attention and non-fixed membrane potential leaks in biological neurons. This biological finding has inspired us to design a highly robust SNN model that closely mimics the biological nervous system. In our study, we first present a unified theoretical framework for SNN robustness constraint, which suggests that improving the encoding method and evolution of the membrane potential leak factor in spiking neurons can improve SNN robustness. Subsequently, we propose a robust SNN (FEEL-SNN) with Frequency Encoding (FE) and Evolutionary Leak factor (EL) to defend against different noises, mimicking the selective visual attention mechanism and non-fixed leak observed in biological systems. Experimental results confirm the efficacy of both our FE, EL, and FEEL methods, either in isolation or in conjunction with established robust enhancement algorithms, for enhancing the robustness of SNNs. Mengting Xu, De Ma, Huajin Tang, Gang Pan 0001 |
NeurIPS | 2 |
| 2024 | Learning improvement of spiking neural networks with dynamic adaptive hyperparameter neurons
Jiakai Liang, De Ma, Ruixue Li, Keqiang Yue |
Appl. Intell. | 3 |
| 2024 | Efficient spiking neural network design via neural architecture search
Qianhui Liu, Malu Zhang, Lang Feng 0002, De Ma, Haizhou Li 0001, Gang Pan 0001 |
Neural Networks | 5 |
| 2024 | LSM-Based Hotspot Prediction and Hotspot-Aware Routing in NoC-Based Neuromorphic ProcessorabstractThe traffic patterns of spiking neural networks (SNNs) exhibit high variability and stochastic, leading to the emergence of elevated traffic hotspots on the network-on-chip (NoC)-based neuromorphic processors. Predicting the occurrence of hotspots remains one of the most challenging issues in NoC design. This article presents the first attempt toward traffic hotspot prediction by utilizing liquid state machine (HP-LSM). The predictor extracts essential information reflecting the current state of the NoC to predict potential routing hotspots in the subsequent time step. Furthermore, we designed the hardware architecture for HP-LSM, which incorporates leaky-integrate-and-fire (LIF) neurons with configurable biological parameters. Meanwhile, we introduce a novel hotspot-aware path-based multicast (HaPM) routing algorithm that utilizes advanced knowledge acquired from HP-LSM to guide packet routing throughout the network, aiming to improve the performance of NoC. Results indicate that the HP-LSM can forecast hotspot formation with an accuracy up to 89.36% and 90.19% for two spiking-based datasets, respectively. The hardware experiment results demonstrate a 92.03% reduction in the average execution time of zero skipping compared with nonzero skipping. Moreover, the HP-LSM exhibits a reduction of up to 79.30% in the number of neurons compared with other related SNN predictor models. The experiments reveal a reduction of 73.67% and 53.42% in the average length of the multicast path when compared with dual-path (DP) or multipath (MP) multicast routing. The HaPM demonstrates improved performance in terms of average latency and throughput compared with DP, MP, and path-based multicast (PbM) multicast routing. Ziyang Kang, Xun Xiao, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | Spike-EFI: Spiking Neural Network for Event-Based Video Frame Interpolation
Dong-Sheng Wu, De Ma |
PSIVT | 2 |
| 2022 | Multi-Level Firing with Spiking DS-ResNet: Enabling Better and Deeper Directly-Trained Spiking Neural NetworksabstractSpiking neural networks (SNNs) are bio-inspired neural networks with asynchronous discrete and sparse characteristics, which have increasingly manifested their superiority in low energy consumption. Recent research is devoted to utilizing spatio-temporal information to directly train SNNs by backpropagation. However, the binary and non-differentiable properties of spike activities force directly trained SNNs to suffer from serious gradient vanishing and network degradation, which greatly limits the performance of directly trained SNNs and prevents them from going deeper. In this paper, we propose a multi-level firing (MLF) method based on the existing spatio-temporal back propagation (STBP) method, and spiking dormant-suppressed residual network (spiking DS-ResNet). MLF enables more efficient gradient propagation and the incremental expression ability of the neurons. Spiking DS-ResNet can efficiently perform identity mapping of discrete spikes, as well as provide a more suitable connection for gradient propagation in deep SNNs. With the proposed method, our model achieves superior performances on a non-neuromorphic dataset and two neuromorphic datasets with much fewer trainable parameters and demonstrates the great ability to combat the gradient vanishing and degradation problem in deep SNNs. Lang Feng 0002, Qianhui Liu, Huajin Tang, De Ma, Gang Pan 0001 |
IJCAI | 4 |
| 2021 | Event-based Action Recognition Using Motion Information and Spiking Neural NetworksabstractEvent-based cameras have attracted increasing attention due to their advantages of biologically inspired paradigm and low power consumption. Since event-based cameras record the visual input as asynchronous discrete events, they are inherently suitable to cooperate with the spiking neural network (SNN). Existing works of SNNs for processing events mainly focus on the task of object recognition. However, events from the event-based camera are triggered by dynamic changes, which makes it an ideal choice to capture actions in the visual scene. Inspired by the dorsal stream in visual cortex, we propose a hierarchical SNN architecture for event-based action recognition using motion information. Motion features are extracted and utilized from events to local and finally to global perception for action recognition. To the best of the authors’ knowledge, it is the first attempt of SNN to apply motion information to event-based action recognition. We evaluate our proposed SNN on three event-based action recognition datasets, including our newly published DailyAction-DVS dataset comprising 12 actions collected under diverse recording conditions. Extensive experimental results show the effectiveness of motion information and our proposed SNN architecture for event-based action recognition. Qianhui Liu, Dong Xing, Huajin Tang, De Ma, Gang Pan 0001 |
IJCAI | 4 |
| 2018 | Scalable NoC-based Neuromorphic Hardware Learning and InferenceabstractBio-inspired neuromorphic hardware is a research direction to approach brain's computational power and energy efficiency. Spiking neural networks (SNN) encode information as sparsely distributed spike trains and employ spike-timingdependent plasticity (STDP) mechanism for learning. Existing hardware implementations of SNN are limited in scale or do not have in-hardware learning capability. In this work, we propose a low-cost scalable Network-on-Chip (NoC) based SNN hardware architecture with fully distributed in-hardware STDP learning capability. All hardware neurons work in parallel and communicate through the NoC. This enables chip-level interconnection, scalability and reconfigurability necessary for deploying different applications. The hardware is applied to learn MNIST digits as an evaluation of its learning capability. We explore the design space to study the trade-offs between speed, area and energy. How to use this procedure to find optimal architecture configuration is also discussed. Haowen Fang, Amar Shrestha, De Ma, Qinru Qiu |
IJCNN | 3 |
| 2018 | Dating ancient paintings of Mogao Grottoes using deeply learnt visual codes
Qingquan Li 0001, Qin Zou 0001, De Ma, Qian Wang 0002, Song Wang 0002 |
Sci. China Inf. Sci. | 3 |
| 2018 | Relative ordering learning in spiking neural network for pattern recognition
Zhitao Lin, De Ma, Jian-Yi Meng, Linna Chen |
Neurocomputing | 2 |
| 2017 | Darwin: A neuromorphic hardware co-processor based on spiking neural networks
De Ma, Juncheng Shen, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001 |
J. Syst. Archit. | 1 |
| 2016 | Darwin: a neuromorphic hardware co-processor based on Spiking Neural Networks
Juncheng Shen, De Ma, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001 |
Sci. China Inf. Sci. | 2 |
| 2015 | Profiling and annotation combined method for multimedia application specific MPSoC performance estimationabstractAccurate and fast performance estimation is necessary to drive design space exploration and thus support important design decisions. Current techniques are either time consuming or not accurate enough. In this paper, we solve these problems by presenting a hybrid method for multimedia multiprocessor system-on-chip (MPSoC) performance estimation. A general coverage analysis tool GNU gcov is employed to profile the execution statistics during the native simulation. To tackle the complexity and keep the analysis and simulation manageable, the orthogonalization of communication and computation parts is adopted. The estimation result of the computation part is annotated to a transaction accurate model for further analysis, by which a gradual refinement of MPSoC performance estimation is supported. The implementation and its experimental results prove the feasibility and efficiency of the proposed method. Kai Huang 0002, Siwen Xiu, Dandan Zheng 0001, Min Yu 0006, De Ma, Kai Huang 0001, Gang Chen 0023, Xiaolang Yan |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2014 | Annotation and analysis combined cache modeling for native simulationabstractTo accelerate the speed of performance estimation and raise its accuracy for MPSoC, we propose a static analysis and dynamic annotation combined method to efficiently model cache mechanism in native simulation. We use a new cache model to statically analyze segmental profiling results to speed up simulation, and utilize a dynamic annotation technique to exactly trace the addresses of local variables. Experimental results show the efficiency of the proposed techniques for more accurate system performance estimation. Rongjie Yan, De Ma, Kai Huang 0002, Siwen Xiu |
ASP-DAC | 2 |
| 2013 | High throughput VLSI architecture for H.264/AVC context-based adaptive binary arithmetic coding (CABAC) decodingabstractContext-based adaptive binary arithmetic coding (CABAC) is the major entropy-coding algorithm employed in H.264/AVC. In this paper, we present a new VLSI architecture design for an H.264/AVC CABAC decoder, which optimizes both decode decision and decode bypass engines for high throughput, and improves context model allocation for efficient external memory access. Based on the fact that the most possible symbol (MPS) branch is much simpler than the least possible symbol (LPS) branch, a newly organized decode decision engine consisting of two serially concatenated MPS branches and one LPS branch is proposed to achieve better parallelism at lower timing path cost. A look-ahead context index (ctxIdx) calculation mechanism is designed to provide the context model for the second MPS branch. A head-zero detector is proposed to improve the performance of the decode bypass engine according to UEG k encoding features. In addition, to lower the frequency of memory access, we reorganize the context models in external memory and use three circular buffers to cache the context models, neighboring information, and bit stream, respectively. A pre-fetching mechanism with a prediction scheme is adopted to load the corresponding content to a circular buffer to hide external memory latency. Experimental results show that our design can operate at 250 MHz with a 20.71k gate count in SMIC18 silicon technology, and that it achieves an average data decoding rate of 1.5 bins/cycle. Kai Huang 0002, De Ma, Rongjie Yan, Haitong Ge, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 2 |
| 2013 | Performance Estimation Techniques With MPSoC Transaction-Accurate ModelsabstractEfficient design of multiprocessor system-on-chip (MPSoC) requires early, fast, and accurate performance estimation techniques. In this paper, we present new techniques based on fine-grained code analysis to estimate accurate performance during simulation of MPSoC transaction accurate models. First, a GCC profiling tool is applied in the native simulation process. Based on the profiling result, an instruction analyzer of the target CPU architecture is proposed to analyze the cycle cost of C code under estimation. In addition, a memory analyzer is used to further estimate memory access latency including both instruction/data cache time cost and global memory access cycles. Both data and instruction cache models are proposed to estimate cache miss penalty, and a segment-based strategy is adopted to update the cache models more efficiently. Furthermore, an equalized access model is presented to imitate the memory access behavior of processors for estimating global memory access latency caused by bus contention and memory bandwidth. We have applied these techniques on an H.264 decoder application with different hardware architectures. The experimental results show that applying these techniques can obviously improve estimation accuracy of transaction accurate models close to that of the virtual prototype models, with a tolerable overhead on simulation speed. De Ma, Rongjie Yan, Kai Huang 0002, Min Yu 0006, Siwen Xiu, Haitong Ge, Xiaolang Yan, Ahmed Amine Jerraya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | A high efficient memory architecture for H.264/AVC motion compensationabstractIn H.264/AVC decoding system, motion compensation operation occupies about 80% of the total memory access and becomes the system bottleneck. In this paper, a high efficient memory architecture for H.264/AVC motion compensation is proposed to extremely reduce external memory access bandwidth. A four-level hierarchical memory organization scheme is utilized to explore the reusability of neighboring blocks at an acceptable area cost. To improve the system processing throughput, five optimization techniques are adopted in motion compensation operation, which enable video decoder to achieve real-time decoding of HD 1080p video stream when operating at 110 MHz. Compared with the existing works, the proposed architecture is able to reduce the memory bandwidth requirement in motion compensation progress by 83.7% and performs better in the real-time application. Chunshu Li, Kai Huang 0002, Xiaolang Yan, Jiong Feng, De Ma, Haitong Ge |
ASAP | 5 |