EDBT 2026 Demo / reviewers in the wild / expert
Pengju Ren
dblp:99/2460
· DBLP profile ↗
57ranked-venue papers
6as first author
40since 2021 · last 2026
0000-0003-1163-2014ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Double Rounding: Nearly Lossless Adaptive Bit Switching for QATabstractModel quantization is widely applied for compressing and accelerating deep neural networks (DNNs). However, conventional Quantization-Aware Training (QAT) focuses on training DNNs with uniform bit-width. The bit-width settings vary across different hardware and transmission demands, which induces considerable training and storage costs. Hence, the scheme of one-shot joint training multiple precisions is proposed to address this issue. Previous works either store a larger FP32 model to switch between different precision models for higher accuracy or store a smaller INT8 model but compromise accuracy due to using shared quantization parameters. In this paper, we introduce the Double Rounding quantization method, which fully utilizes the quantized representation range to accomplish nearly lossless bit-switching while reducing storage by using the highest integer precision instead of full precision. Furthermore, we observe a competitive interference among different precisions during one-shot joint training, primarily due to inconsistent gradients of quantization scales during backward propagation. To tackle this problem, we propose an Adaptive Learning Rate Scaling (AdaScale) technique that dynamically adapts learning rates for various precisions to optimize the training process. Additionally, we extend our Double Rounding to one-shot mixed precision training and develop a Hessian-Aware Stochastic Bit-switching (HessBit) strategy. Experimental results on the ImageNet-1K classification demonstrate that our methods have enough advantages to state-of-the-art one-shot joint QAT in both multi-precision and mixed-precision. We validate the feasibility of our method on detection and segmentation tasks, as well as on LLMs task. Haiduo Huang, Tian Xia 0008, Pengju Ren |
AAAI | 4 |
| 2026 | PartialNet: Compute Less, Perform BetterabstractAchieving a balance between low parameter count, reduced FLOPs, and high accuracy and throughput remains a central challenge in neural network design. To address this, we propose the partial channel mechanism (PCM), which leverages the inherent redundancy in feature map channels. PCM divides feature map channels into multiple groups, each processed by distinct operations such as convolution, attention, pooling, or identity mapping. Building on this, we introduce partial attention convolution (PATConv), a novel module that efficiently fuses convolution and visual attention within a unified framework. Our results demonstrate that PATConv can fully replace both standard convolution and visual attention modules, leading to significant reductions in parameters and FLOPs. Furthermore, PATConv enables three efficient visual attention variants: Partial Channel Attention, Partial Spatial Attention, and Partial Self-Attention. To further optimize the allocation of channel splits, we propose dynamic {partial convolution (DPConv), which adaptively learns the optimal split ratio for each layer, achieving a better trade-off between speed and accuracy. By integrating PATConv and DPConv, we develop a new hybrid network family, PartialNet, which achieves superior top-1 accuracy and inference speed on ImageNet-1K, and demonstrates strong performance on COCO detection and segmentation tasks. Haiduo Huang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
AAAI | 4 |
| 2026 | MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT DistillationabstractJin Cui, Jiaqi Guo, Jiepeng Zhou, Ruixuan Yang, Jiayi Lu, Jiajun Xu, Jiangcheng Song, Boran Zhao, Pengju Ren. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiepeng Zhou, Ruixuan Yang, Jiangcheng Song, Boran Zhao, Pengju Ren |
ACL (1) | 9 |
| 2026 | Jakiro: Boosting Speculative Decoding via Decoupled MoEabstractSpeculative decoding has emerged as a promising technique to accelerate large language model inference by employing a smaller draft model to predict multiple tokens, which are then verified in parallel by the larger target model.However, existing approaches face a fundamental limitation: candidates at the same tree layer share identical feature representations, constraining diversity and diminishing overall effectiveness.We identify this as an intra-layer coupling problem that limits prediction accuracy.To address this challenge, we propose Jakiro, which introduces decoupled Mixture of Experts (MoE) into the draft model, enabling different experts to generate diverse candidate tokens from distinct feature spaces.We further propose Contrastive-Enhanced Parallel Decoding (CEPD) that combines autoregressive and parallel decoding with a contrastive mechanism to reduce inference steps while maintaining accuracy.Extensive experiments across diverse models and tasks demonstrate that Jakiro achieves significant speedups over strong baselines, with particularly notable improvements in non-greedy decoding scenarios where token diversity is crucial. Haiduo Huang, Fuwei Yang, Pengju Ren |
ACL (1) | 4 |
| 2026 | MoEA: A Mixed-Precision Edge Accelerator for CNN-MSA Models with Fine-Tuning Support
Qiwei Dang, Chengyu Ma, Zhiwang Huo, Guoming Yang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
ASP-DAC | 7 |
| 2026 | RWIP: Region-Level Write-Intensity Prediction for GPU L2 Cache Based on Hybrid-Retention STT-MRAM
Yujie Pu, Chen Zhao 0009, Jianfeng An, Wenzhe Zhao 0001, Pengju Ren |
Euro-Par (1) | 5 |
| 2026 | An Adaptive Neuro-fuzzy Framework for Stock Price Forecasting
Pengju Ren, Chengbin Peng 0001, Baisong Liu, Xiaoqin Fan |
ICIC (7) | 1 |
| 2026 | Hierarchical-ISA Supporting Row-Wise Operands for Efficient DNN ComputationabstractDeep neural networks (DNNs) have become a cornerstone in advancing artificial intelligence, but their complexity often leads to inefficient hardware utilization due to varying structure characteristics and excessive memory accesses. Domain-specific architectures (DSAs) offer a solution by optimizing data locality through data stationary, tiling, and layer fusion, which minimize memory access and energy consumption while boosting performance. However, current approaches lack flexibility for efficient memory management at the appropriate granularity, causing misaligned accesses and decreasing reuse potential for variable-sized tiles. To this end, we propose a hierarchical Instruction Set Architecture (hierarchical-ISA) combining a RISC-V ISA and a flexible CISC-style macro-ISA (mISA). Unlike byte-level RISC-V, mISA employs row-wise tiles as the fundamental operand, enabling efficient data reuse across adjacent iterations as well as residual connections. This mISA approach simplifies DNN programming, enhances data partitioning and manipulation efficiency, and enables a hardware-software co-designed Remapping mechanism that facilitates data reuse without physical data movement. Experiments show 31.8%–72.0% reductions in off-chip memory access across MobileNet, ResNet, Swin Transformer, MobileViT, along with speedups of 2.9× to 7.4× compared to previous DNN accelerators. We also conduct comparisons under the roofline model with NVIDIA RTX A6000 and Intel Core i7-10700K. The results show that our arithmetic intensity reaches up to 26.0× that of i7-10700K and 22.6× that of A6000. Zhiwang Huo, Wenzhe Zhao 0001, Qiwei Dang, Chengyu Ma, Guoming Yang, Gelin Fu, Tian Xia 0008, Pengju Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2026 | SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution StrategyabstractThe growing demand for sparse tensor algebra (SpTA) in machine learning and big data has driven the development of various sparse tensor accelerators. However, most existing manually designed accelerators are limited to specific scenarios, and it’s time-consuming and challenging to adjust a large number of design factors when scenarios change. Therefore, automating the design of SpTA accelerators is crucial. Nevertheless, previous works focus solely on eithermapping(i.e., tiling communication and computation in space and time) orsparse strategy(i.e., bypassing zero elements for efficiency), leading to suboptimal designs due to the lack of consideration of both. A unified framework that jointly optimizes both is urgently needed. However, integrating mapping and sparse strategies leads to a combinatorial explosion in the design space(e.g., as large asO(1041) for the workloadP32×64×Q64×48=Z32×48). This vast search space renders most conventional optimization methods (e.g., particle swarm optimization, reinforcement learning and Monte Carlo tree search) inefficient. To address this challenge, we propose an evolution strategy-based sparse tensor accelerator optimization framework, called SparseMap. SparseMap constructing a more essential-factor design space with the consideration of both mapping and sparse strategy. We introduce a series of enhancements to genetic encoding and evolutionary operators, enabling SparseMap to efficiently explore the vast and diverse design space. We quantitatively compare SparseMap with prior works and classical optimization methods, demonstrating that SparseMap consistently finds superior solutions. Boran Zhao, Haiming Zhai, Zihang Yuan, Hetian Liu, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | FP2: A 2-bit Floating-Point Format for Edge-AI Inference and Fine-TuningabstractThe increasing scale of Deep Neural Networks (DNNs) has made 2-bit quantization crucial for mitigating memory bottlenecks on edge devices. Low-bitwidth floating-point formats, offering larger dynamic ranges and avoiding quantization steps, have emerged as promising alternatives to fixed-point quantization. However, constructing viable floating-point representations with fewer than 3 bits remains challenging, as conventional formats require at least one sign bit, one exponent bit, and one mantissa bit. We address this challenge by introducing a novel data compression method that uses a 4-bit encoding space to represent two floating-point values, achieving an effective storage density of 2 bits per value. Depending on the bit width of the exponent and mantissa, we propose two different 2-bit floating-point encodings:fp2-e1m0andfp2-e0m1. Based onfp2, we introduce two computing architectures that simplify floating-point multiply-accumulate (MAC) operations into bitwise addition and logic operations, reducing floating-point computation by factors of$2\times $and$4\times $. As a result,fp2offers a practical solution for efficient inference using floating-point arithmetic on resource-constrained edge devices. Moreover, we analyze the error characteristics of thefp2data format from three perspectives. To validate the effectiveness of thefp2format, we conduct experiments on ResNet18/50 and ConvNeXt-Tiny using the CIFAR-10 and ImageNet-1K datasets. Compared tofp4, our approach reduces model size by 47%, with accuracy loss is less than 2 percentage points. Notably, on CIFAR-10, some results are close to those offp32. In contrast, when evaluated under 2-bit GPTQ,fp2demonstrates significant advantages over the baseline method on the LLAMA model. For hardware evaluation, we implement our design at the RTL level and evaluate it on both FPGA and ASIC platforms. Compared to computation architectures based onfp4, ourfp4$\times $fp2processing element (PE) array reduces area by 15% and power consumption by 8%. Furthermore, ourfp2$\times $fp2PE array achieves a remarkable 78% reduction in both area and power consumption. Qiwei Dang, Chengyu Ma, Haiduo Huang, Gelin Fu, Zhiwang Huo, Guoming Yang, Pengchen Zong, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2026 | AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN TrainingabstractTraining deep neural networks (DNNs) on edge devices faces challenges due to the large-scale datasets required, which are costly for edge devices, especially in large language model (LLM) tasks. To address this, a DNN-free method called Near-Memory Sampling (NMS) has been introduced. NMS reduces dimensionality and performs exemplar sampling in the reduced space, avoiding architectural bias and improving generalization. However, NMS has two limitations: 1) The mismatch between the search method and the non-monotonic property of the perplexity error function leads to the emergence of outliers; 2) Key parameter (i.e., target perplexity) is selected empirically, introducing arbitrariness and leading to uneven sampling. These two issues lead torepresentative biasof exemplars, resulting in degraded accuracy. To overcome these, we propose AdapSNE, which integrates the Fireworks Algorithm (FWA) for efficient non-monotonic search to avoid outliers and uses entropy-guided optimization for uniform sampling, ensuring representative training samples. To reduce the cost of iterative computations, we design an accelerator with custom dataflow and time-multiplexing mechanisms. Experimental results show that AdapSNE outperforms state-of-the-art methods, including both DNN-based (DQAS) and DNN-free (NMS) approaches, across small-scale image datasets, large-scale datasets, and the MMLU benchmark for LLM tasks. Boran Zhao, Hetian Liu, Zihang Yuan, Li Zhu 0003, Lina Xie, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2026 | A Groupwise Add-Multiply-Shift-Accumulate Datapath for Efficient DNN AcceleratorsabstractMultiply–accumulate(MAC) units account for a large fraction of the power and area in modern deep neural network (DNN) accelerators. Although low-bitwidth quantization reduces hardware overhead, the high cost of multipliers remains a fundamental bottleneck in modern accelerator datapaths. This article proposes add–multiply–shift–accumulate (AMC), a groupwise arithmetic datapath that reduces multiplier count by sharing base multiplications across groups of neighboring weights and generating residual products using lightweight shift operation. To support efficient deployment, we design a compact residual encoding and buffer organization that allows AMC arrays to be constructed with minimal decoding and control overheads. While AMC can be directly applied to existing quantized models, we further introduce a lightweight residual-aware fine-tuning (RAF) procedure to increase AMC compatibility. We implement AMC-based accelerators in SystemVerilog and synthesize them in TSMC 28-nm CMOS technology across operating frequencies from 500MHz to 1GHz. At the compute unit level, AMC reduces arithmetic area by 39.5%–62.8% and dynamic power by 32.2%–60.3% compared with optimized baseline multipliers. When integrated into CNN and Vision Transformer accelerators, AMC achieves$1.34\times $–$18.90\times $higher area efficiency and up to$10.16\times $higher energy efficiency than prior designs while preserving baseline inference accuracy. Zhiwang Huo, Wenzhe Zhao 0001, Yuanchang Gong, Tian Xia 0008, Zheng Wang 0001, Pengju Ren |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Rapid Dynamic Obstacle Avoidance for UAVs Enhanced by DVS and Neuromorphic ComputingabstractAchieving rapid and accurate dynamic obstacle avoidance is crucial for enhancing the survivability of unmanned aerial vehicles (UAVs) in hazardous conditions. To accomplish dynamic obstacle avoidance, sensors with high temporal resolution and efficient processing models are required. Dynamic vision sensors (DVS) fulfill the sensing requirements, while spiking neural networks (SNNs) address the processing demands. In this paper, we develop an end-to-end obstacle avoidance algorithm for UAVs using only a single monocular DVS as the sensor and further enhance accuracy and speed through our proposed mechanisms. The algorithm consists of three components: ego-motion compensation, an SNN model for movement analysis, and a force filter inspired by spiking neurons. In movement analysis, we propose the temporal potential pooling (TPP) and incremental event (EI) mechanisms to accelerate our SNN model. The real-flight experiments confirm that our algorithm achieves approximately 90% accuracy with a processing latency as low as 4ms on a GPU, surpassing state-of-the-art methods. Ablation studies show that the proposed method maintains high accuracy in movement detection while significantly reducing computational time. Our method operates in real-time, achieves high accuracy, and is feasible across a wide range of environments. Our code is available at https://github.com/AmperiaWang/oanet_s1 for reproducibility. Tingbang Liang, Yilin Shi, Pengju Ren |
ICRA | 6 |
| 2025 | Magellan: A High-Performance Loop-Guided Prefetcher for Indirect Memory AccessabstractGraph analytics and sparse linear algebra applications heavily rely on indirect memory access (IMA).IMAs are characterized by poor temporal and spatial locality, which causes frequent high-latency DRAM accesses.While dedicated hardware prefetchers for IMA have been explored, they target narrow access patterns and tend to introduce significant hardware complexity.Software prefetching offers a promising alternative, leveraging compiler analysis to prefetch indirection patterns.However, existing software prefetchers struggle with sparse applications due to limited loop iterations and complex IMA patterns across nested loops.We propose Magellan, a novel loop-guided software prefetcher designed to detect and schedule IMA prefetches efficiently.Magellan introduces two key innovations: (1) extracting dependence graphs across loop levels to detect complex IMA patterns and (2) capturing inner-outer loop semantics to prefetch for both current and future iterations.We evaluate Magellan on 14 memory-intensive benchmarks using real-world datasets from social networks and web graphs.Compared to the best existing IMA software prefetcher, Magellan reduces cache misses by 25% and dynamic instruction counts by 14% on average.This results in a 1.14× average speedup, with performance gains of up to 1.41×. Gelin Fu, Tian Xia 0008, Mingzhuo Yin, Prashant J. Nair, Mieszko Lis, Pengju Ren |
ISCA | 6 |
| 2025 | DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation TrainerabstractRecent advances in knowledge distillation have emphasized the importance of decoupling different knowledge components. While existing methods utilize momentum mechanisms to separate task-oriented and distillation gradients, they overlook the inherent conflict between target-class and non-target-class knowledge flows. Furthermore, low-confidence dark knowledge in non-target classes introduces noisy signals that hinder effective knowledge transfer. To address these limitations, we propose DeepKD, a novel training framework that integrates dual-level decoupling with adaptive denoising. First, through theoretical analysis of gradient signal-to-noise ratio (GSNR) characteristics in task-oriented and non-task-oriented knowledge distillation, we design independent momentum updaters for each component to prevent mutual interference. We observe that the optimal momentum coefficients for task-oriented gradient (TOG), target-class gradient (TCG), and non-target-class gradient (NCG) should be positively related to their GSNR. Second, we introduce a dynamic top-k mask (DTM) mechanism that gradually increases K from a small initial value to incorporate more non-target classes as training progresses, following curriculum learning principles. The DTM jointly filters low-confidence logits from both teacher and student models, effectively purifying dark knowledge during early training. Extensive experiments on CIFAR-100, ImageNet, and MS-COCO demonstrate DeepKD's effectiveness. Haiduo Huang, Jiangcheng Song, Pengju Ren |
NeurIPS | 4 |
| 2025 | GeGS-PCR: Fast and Robust Color 3D Point Cloud Registration with Two-Stage Geometric-3DGS FusionabstractWe address the challenge of point cloud registration using color information, where traditional methods relying solely on geometric features often struggle in low-overlap and incomplete scenarios. To overcome these limitations, we propose GeGS-PCR, a novel two-stage method that combines geometric, color, and Gaussian information for robust registration. Our approach incorporates a dedicated color encoder that enhances color features by extracting multi-level geometric and color data from the original point cloud. We introduce the Geometric-3DGS module, which encodes the local neighborhood information of colored superpoints to ensure a globally invariant geometric-color context. Leveraging LORA optimization, we maintain high performance while preserving the expressiveness of 3DGS. Additionally, fast differentiable rendering is utilized to refine the registration process, leading to improved convergence. To further enhance performance, we propose a joint photometric loss that exploits both geometric and color features. This enables strong performance in challenging conditions with extremely low point cloud overlap. We validate our method by colorizing the Kitti dataset as ColorKitti and testing on both Color3DMatch and Color3DLoMatch datasets. Our method achieves state-of-the-art performance with Registration Recall at 99.9%, Relative Rotation Error as low as 0.013, and Relative Translation Error as low as 0.024, improving precision by at least a factor of 2. Haiduo Huang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
NeurIPS | 5 |
| 2025 | Diffattack-X: An effective transferable adversarial attack based on diffusion models
Lisha Li, Yini Pu, Pengju Ren, Jiaxing Chu |
Appl. Intell. | 7 |
| 2025 | A self-supervised learning method for Raman spectroscopy based on masked autoencoders
Pengju Ren, Rigui Zhou, Yaochong Li |
Expert Syst. Appl. | 1 |
| 2024 | Differential-Matching Prefetcher for Indirect Memory AccessabstractIndirect memory access is a critical bottleneck for modern CPUs, especially for graph analysis and sparse linear algebra applications, where the values of one data array are used to generate the fetching addresses of another array. It often causes irregular data accesses with poor temporal and spatial locality that are difficult to be captured by conventional hardware prefetchers. For many complex workloads, such indirect access patterns may have different types and are nested in a multiplelevel form. Moreover, branch mispredictions would further disturb their patterns, making them even harder to detect. As a result, existing hardware prefetchers are unable to fully prefetch complex indirect patterns. This paper proposes DMP, a low-cost hardware prefetcher to improve the memory latency in several representative irregular workloads. DMP targets four types of indirect memory access patterns including single, range, multi-level, and multi-way indirect access. DMP uses differential matching to identify an indirect access pattern in pair with its corresponding index stream. Then DMP uses a flexible prefetching mechanism to dynamically adapt the prefetching degree to maintain prefetching coverage. We evaluate the performance, energy consumption, and transistor cost of DMP among various algorithms from GAP, NAS, and HPCG benchmarks. DMP improves performance by 1.8 × (up to 5.6 ×) on average against state-of-the-art hardware prefetchers and 1.2 × (up to 2.3 ×) speedup against state-of-the-art compiler-based prefetcher Prodigy. Besides, the proposed design is optimized to take only 0.9KB of storage, making it feasible to be integrated into current CPU designs. Gelin Fu, Tian Xia 0008, Zhongpei Luo, Wenzhe Zhao 0001, Pengju Ren |
HPCA | 6 |
| 2024 | Towards Optimal Lane-changing Coordination of CAVs in Multi-lane Mixed Traffic ScenariosabstractLane changing is a fundamental but challenging operation for moving vehicles. Connected and Automated Vehicles(CAVs) enable autonomous vehicles to cooperate with each other to accomplish the lane changing tasks, profiting from their communication ability. However, dispatching CAVs in mixed traffic remains difficult due to the stochastic behaviors and uncertain intentions of Human-Driven Vehicles(HDVs). To tackle this issue, this paper devises a coordination approach based on Conflict-Based Search(CBS) theory. Firstly, HDVs are accurately modeled as constraints to enable usage of CBS in the mixed traffic. Additionally, virtual goals are introduced to search CAVs’ priority and outlets along with path finding. Furthermore, we optimize the performance of CBS in dense traffic by defining the concept of following vehicles. Experiments show that performance is improved by utilizing new conflict prioritizing rules and a heuristic value calculation method that derived from following vehicles. Finally, we introduce grouping vehicles to extend the proposed method for solving extremely dense and large instances at a scale of more than one hundred without significant loss in efficiency. Yijun Mao, Chongshan Jiao, Pengju Ren |
ICRA | 4 |
| 2024 | Error Loss NetworksabstractA novel model called error loss network (ELN) is proposed to build an error loss function for supervised learning. The ELN is similar in structure to a radial basis function (RBF) neural network, but its input is an error sample and output is a loss corresponding to that error sample. That means the nonlinear input-output mapper of the ELN creates an error loss function. The proposed ELN provides a unified model for a large class of error loss functions, which includes some information-theoretic learning (ITL) loss functions as special cases. The activation function, weight parameters, and network size of the ELN can be predetermined or learned from the error samples. On this basis, we propose a new machine learning paradigm where the learning process is divided into two stages: first, learning a loss function using an ELN; second, using the learned loss function to continue to perform the learning. Experimental results are presented to demonstrate the desirable performance of the new method. Badong Chen, Pengju Ren |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | PrSpMV: An Efficient Predictable Kernel for SpMVabstractSparse Matrix-Vector Multiplication (SpMV) has been widely applied in scientific computation, industry simulation, and intelligent computation domains, which is the critical algorithm in all these applications. Due to the poor data locality, low cache usage, and extremely irregular branch patterns caused by the highly sparse and random distributions, SpMV optimization has become one of the most challenging problems for modern high-performance processors. In this paper, we study the bottlenecks of SpMV on current out-of-order CPUs and propose a novel SpMV kernel named PrSpMV to improve its performance by pursuing high predictability. Specifically, we improve the memory access regularity and locality by creating serialized access patterns so that the data prefetching efficiency and cache usage are optimized. We also improve pipeline efficiency by creating regular branch patterns to make branch prediction more accurate. Experiment results show that using the above optimization approaches, PrSpMV can eliminate nearly all branch mispredictions. Moreover, it can also significantly reduce the average L2 cache miss rate from 57% to 20% via efficiently leveraging hardware prefetchers. By using PrSpMV, stride prefetcher can be boosted with 1.31× speedup and dedicated irregular prefetcher can be improved with 1.40× speedup. Meanwhile, on commercial high-end Intel processors, it achieves 1.32× speedup against some state-of-the-art SpMV kernels. Gelin Fu, Tian Xia 0008, Shaoru Qu, Zhongpei Luo, Pengyu Cheng, Runfan Guo, Yitong Ding, Pengju Ren |
ICCD | 9 |
| 2023 | TAQ: Top-K Attention-Aware Quantization for Vision TransformersabstractModel quantization can reduce the memory footprint of the neural network and improve the computing efficiency. However, the sparse attention in Transformer models is difficult to quantize, the main challenge is that changing the order of attention values and shifting attention regions might lead to incorrect prediction results. To address this problem, we propose quantization method, termed TAQ, which uses the proposed TOP-K attention-aware loss to search the quantization parameters. Further, we combine the sequential and parallel quantization methods to optimize the procedure. We evaluate the generalization ability of TAQ on various vision Transformer variants, and its performance on image classification and object detection tasks. TAQ makes the TOP-K attention ranking more consistent before and after quantization, and significantly reduces the attention shifting rate, compared with PTQ4ViT, TAQ improves the performance by 0.66 and 0.45, respectively on ImageNet and COCO, achieves the state-of-the-art performance. Lili Shi, Haiduo Huang, Bowei Song, Meng Tan, Wenzhe Zhao 0001, Tian Xia 0008, Pengju Ren |
ICIP | 7 |
| 2023 | UCLF: An Uncertainty-Aware Cooperative Lane-Changing Framework for Connected Autonomous Vehicles in Mixed TrafficabstractHuman-driven vehicles (HDVs) will still exist for a long time as we move towards the era of connected autonomous vehicles (CAVs). It is challenging to ensure the safety of the system and improve the efficiency of convoys in mixed traffic environments due to the stochastic behaviors and uncertain intentions of HDVs. To address these issues, this paper develops an uncertainty-aware cooperative lane-changing framework, termed UCLF, for CAVs based on partially observable Markov decision process (POMDP). We extend POMDP to multi-agent cooperative lane-changing by prioritizing CAVs according to lane-changing urgency and planning for CAVs sequentially. Two novel cooperation mechanisms, namely cooperative implicit branching and cooperative explicit pruning, are proposed to promote efficiency and ensure safety. Numerical experiments are conducted to show the smooth and efficient lane-changing maneuvers under intention uncertainty. Compared to baseline, UCLF achieves up to 28.7% decrease in total travel time on average. We also validate UCLF in a real multi-AGV (Automated Guided Vehicle) system to demonstrate the usability and reliability of our study. Yijun Mao, Chongshan Jiao, Pengju Ren |
IV | 4 |
| 2023 | REMAP: A Spatiotemporal CNN Accelerator Optimization Methodology and Toolkit ThereofabstractDesigning convolutional neural network (CNN) accelerators is getting more difficult owing to the fast-increasing types of CNN models. Some approaches use constant dataflow and microarchitecture that have lower design complexity. However, these accelerators are difficult to adapt with the highly-diverse CNN models and often suffer from low process element utilization. Some other accelerators resort to reconfigurable devices, such as field-programmable gate array (FPGA) and coarse-grained reconfigurable array to support flexible dataflows in order to fit diverse CNN layers. However, layer-by-layer processing may require more energy for frequent reconfiguration and off-chip DDR access. In this work, we introduce a reconfigurable pipeline accelerator (RPA) that can reduce the latency and DDR access by pipelining the compuptation of CNN layers. Although there have been several researches that try to speedup the design process by automatically exploring subset of the accelerator design space, identifying an available automated design tool that can effectively find the complete and optimal design scheme remains a problem, especially for the novel RPA architecture type. Unfortunately, comprehensive exploration of the whole design space faces an excessive large searching space. To tackle this problem, we propose REMAP, a toolkit for designing CNN accelerators based on the Monte Carlo tree search (MCTS) method. To efficiently search the huge design space, we propose several methods to improve searching efficiency. Evaluations show that REMAP significantly outperforms some state-of-the-art approaches; compared with GAMMA, it achieves an average speed increase of$14.75\times $, and an energy reduction of 45.45%; it also achieves a speed increase of$32.6\times $against ConfuciuX on MobileNetV2 and ResNet50. We also show an FPGA accelerator implementation which is based on REMAP’s search result, and it achieves high performance in real-time CNN tasks. This indicates that REMAP can provide high-quality design exploration with valuable insights and useful architecture design guidances. Boran Zhao, Tian Xia 0008, Haiming Zhai, Fulun Ma, Hanzhi Chang, Wenzhe Zhao 0001, Pengju Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | An Energy-and-Area-Efficient CNN Accelerator for Universal Powers-of-Two QuantizationabstractCNN model computation on edge devices is tightly restricted to the limited resource and power budgets, which motivates the low-bit quantization technology to compress CNN models into 4-bit or lower format to reduce the model size and increase hardware efficiency. Most current low-bit quantization methods use uniform quantization that maps weight and activation values onto evenly-distributed levels, which usually results in accuracy loss due to distribution mismatch. Meanwhile, some non-uniform quantization methods propose specialized representation that can better match various distribution shapes but are usually difficult to be efficiently accelerated on hardware. In order to achieve low-bit quantization with high accuracy and hardware efficiency, this paper proposes Universal Power-of-Two (UPoT), a novel low-bit quantization method that represents values as the addition of multiple power-of-two values selected from a series of subsets. By updating the subset contents, UPoT can provide adaptive quantization levels for various distributions. For each CNN model layer, UPoT automatically searches for the optimized distribution that minimizes the quantization error. Moreover, we design an efficient accelerator system with specifically optimized power-of-two multipliers and requantization units. Evaluations show that the proposed architecture can provide high-performance CNN inference with reduced circuit area and energy, and outperforms several mainstream CNN accelerators with higher ($8\times $–$65\times $) area efficiency and ($2\times $–$19\times $) energy efficiency. Further experiments of 4/3/2-bit quantization on ResNet18/50, MobileNet_V2 and EfficientNet models show that our UPoT can achieve high model accuracy which greatly outperform other state-of-the-art low-bit quantization methods by 0.3%–6%. The results indicate that our approach provides a highly-efficient accelerator for low-bit CNN model quantization with low hardware overheads and good model accuracy. Tian Xia 0008, Boran Zhao, Gelin Fu, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | Optimizing FPGA-Based DNN Accelerator With Shared Exponential Floating-Point FormatabstractIn recent years, low-precision fixed-point computation has become a widely used technique for neural network inference on FPGAs. However, this approach has some limitations, as certain neural networks are difficult to quantify using fixed-point arithmetic, such as those involved in super-resolution scaling, image denoising, and other scenarios that lack sufficient conditions for fine-tuning. Furthermore, deploying a floating-point precision neural network directly on an FPGA would lead to significant hardware overhead and low computational efficiency. To address this issue, this paper proposes an FPGA-friendly floating-point data format that achieves the same storage density as int8 without sacrificing inference accuracy or requiring fine-tuning. Additionally, this paper presents an FPGA-based neural network accelerator that is compatible with the proposed format, utilizing DSP resources to increase the number of DSP cascading from 7 to 16, and solving the back-to-back accumulation issue of floating-point numbers. This design achieves comparable resource consumption and execution efficiency to those of 8-bit fixed-point accelerators. Experimental results demonstrate that the accelerator proposed in this study achieves the same accuracy as the native floating point on multiple neural networks without fine-tuning, and remains high computing performance. When deployed on the Xilinx ZU9P, the performance achieves 4.072 TFlops at 250 MHz, which outperforms the previous works, including the Xilinx official DPU. Wenzhe Zhao 0001, Qiwei Dang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | A Comprehensive Performance Model of Sparse Matrix-Vector Multiplication to Guide Kernel OptimizationabstractSparse Matrix-Vector Multiplication (SpMV) is important in scientific and industrial applications and remains a well-known challenge for modern CPUs due to high sparsity and irregularity. Many researchers try to improve SpMV performance by designing dedicated data formats and computation patterns. However, out-of-order superscalar CPUs have complex micro-architectures where exist complicated interactions and restrictions among software and hardware factors. It is hard to systematically study the effectiveness of optimization methods on the overall performance, as its benefits may be undermined by other factors. In this paper, we thoroughly study the execution of SpMV on modern CPUs and propose a comprehensive performance model to reveal the critical factors and their relationships. Specifically, we first study the coding characteristics of SpMV kernels to identify key factors worthy of attention. Then we model the execution of SpMV as two overlapped parts: CPU pipeline and memory latency. Both are carefully modeled with related hardware and software factors. We also model SIMD performance with the usage of specific SIMD instructions and vector registers. Experiments show that our model matches the actual execution of real-world processors. Guided by the model, we propose SpV8, a novel SpMV kernel that optimizes critical factors to improve computation efficiency and memory bandwidth. Experiments on Intel/AMD x86 and ARM AArch64 platforms show that SpV8 outperforms several state-of-the-art approaches with large margins, achieving average$3.4\times$over Intel Math Kernel Library and$1.4\times$over the best existing approach. Such results indicate that the proposed model is capable of valuable guidance for efficient SpMV optimizations. Tian Xia 0008, Gelin Fu, Zhongpei Luo, Lucheng Zhang, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2023 | HIPU: A Hybrid Intelligent Processing Unit With Fine-Grained ISA for Real-Time Deep Neural Network Inference ApplicationsabstractNeural network algorithms have shown superior performance over conventional algorithms, leading to the designation and deployment of dedicated accelerators in practical scenarios. Coarse-grained accelerators achieve high performance but can support only a limited number of predesigned operators, which cannot cover the flexible operators emerging in modern neural network algorithms. Therefore, fine-grained accelerators, such as instruction set architecture (ISA)-based accelerators, have become a hot research topic due to their sufficient flexibility to cover the unpredefined operators. The main challenges for fine-grained accelerators include the undesired long delays of single-image inference when performing multibatch inference, as well as the difficulty of meeting real-time constraints when processing multiple tasks simultaneously. This article proposes a hybrid intelligent processing unit (HIPU) to address the aforementioned problems. Specifically, we design a novel conversion-free data format, expanding the single-instruction multiple-data (SIMD) instruction set and optimizing the microarchitecture design to improve the performance. We also arrange the inference schedule to guarantee scalability on multicores. The experimental results show that the proposed accelerator maintains high multiply–accumulation (MAC) utilization for all common operators and achieves high performance with 4–$7\times $speedup against NVIDIA RTX2080Ti GPU. Finally, the proposed accelerator is manufactured using TSMC 28-nm technology, achieving 1 GHz for each core, with a peak performance of 13 TOPS. Wenzhe Zhao 0001, Guoming Yang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | AdaBin: Improving Binary Neural Networks with Adaptive Binary Sets
Zhijun Tu, Xinghao Chen 0001, Pengju Ren, Yunhe Wang 0001 |
ECCV (11) | 3 |
| 2022 | MI2D: Accelerating Matrix Inversion with 2-Dimensional Tile ManipulationsabstractMatrix inversion is critical in mathematics and scientific applications. Large-scale dense matrix inversion is especially challenging for modern computers due to its heavy dependency of matrix elements and the poor temporal data locality. In this paper, we propose a novel accelerator termed MI2D, which converts matrix inversion into regular matrix multiplications using 2-dimensional cross-tile operations and novel algorithms for efficient data reuse and computations. Our evaluations show that MI2D can be easily integrated with existing matrix engines in modern high-end CPU and NPU, and effectively improves matrix inversion with 2.7× speedup against Intel Skylake CPU, and 24× against NVIDIA RTX 2080 Ti. Lingfeng Chen, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | Multikernel Correntropy for Robust LearningabstractAs a novel similarity measure that is defined as the expectation of a kernel function between two random variables, correntropy has been successfully applied in robust machine learning and signal processing to combat large outliers. The kernel function in correntropy is usually a zero-mean Gaussian kernel. In a recent work, the concept of mixture correntropy (MC) was proposed to improve the learning performance, where the kernel function is a mixture Gaussian kernel, namely, a linear combination of several zero-mean Gaussian kernels with different widths. In both correntropy and MC, the center of the kernel function is, however, always located at zero. In the present work, to further improve the learning performance, we propose the concept of multikernel correntropy (MKC), in which each component of the mixture Gaussian kernel can be centered at a different location. The properties of the MKC are investigated and an efficient approach is proposed to determine the free parameters in MKC. Experimental results show that the learning algorithms under the maximum MKC criterion (MMKCC) can outperform those under the original maximum correntropy criterion (MCC) and the maximum MC criterion (MMCC). Badong Chen, Yuqing Xie 0002, Zejian Yuan, Pengju Ren, Harry Qin |
IEEE Trans. Cybern. | 5 |
| 2022 | Perturbation of Spike Timing Benefits Neural Network Performance on Similarity SearchabstractPerturbation has a positive effect, as it contributes to the stability of neural systems through adaptation and robustness. For example, deep reinforcement learning generally engages in exploratory behavior by injecting noise into the action space and network parameters. It can consistently increase the agent's exploration ability and lead to richer sets of behaviors. Evolutionary strategies also apply parameter perturbations, which makes network architecture robust and diverse. Our main concern is whether the notion of synaptic perturbation introduced in a spiking neural network (SNN) is biologically relevant or if novel frameworks and components are desired to account for the perturbation properties of artificial neural systems. In this work, we first review part of the locality-sensitive hashing (LSH) of similarity search, the FLY algorithm, as recently published in Science, and propose an improved architecture, time-shifted spiking LSH (TS-SLSH), with the consideration of temporal perturbations of the firing moments of spike pulses. Experiment results show promising performance of the proposed method and demonstrate its generality to various spiking neuron models. Therefore, we expect temporal perturbation to play an active role in SNN performance. Ziru Wang, Badong Chen, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | SpV8: Pursuing Optimal Vectorization and Regular Computation Pattern in SpMVabstractSparse Matrix-Vector Multiplication (SpMV) plays an important role in many scientific and industry applications, and remains a well-known challenge due to the high sparsity and irregularity. Most existing researches on SpMV try to pursue high vectorization efficiency. However, such approaches may suffer from non-negligible speculation penalty due to their irregular computation patterns. In this paper, we propose SpV8, a novel approach that optimizes both speculation and vectorization in SpMV. Specifically, SpV8 analyzes data distribution in different matrices and row panels, and accordingly applies optimization method that achieves the maximal vectorization with regular computation patterns. We evaluate SpV8 on Intel Xeon CPU and compare with multiple state-of-art SpMV algorithms using 71 sparse matrices. The results show that SpV8 achieves up to 10× speedup (average 2.8×) against the standard MKL SpMV routine, and up to 2.4× speedup (average 1.4×) against the best existing approach. Moreover, SpMV features very low preprocessing overhead in all compared approaches, which indicates SpV8 is highly-applicable in real-world applications. Tian Xia 0008, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
DAC | 5 |
| 2021 | ac2SLAM: FPGA Accelerated High-Accuracy SLAM with Heapsort and Parallel Keypoint ExtractorabstractIn order to fulfill the rich functions of the application layer, robust and accurate Simultaneous Localization and Mapping (SLAM) technique is very critical for robotics. However, due to the lack of sufficient computing power and storage capacity, it is challenging to delpoy high-accuracy SLAM in embedded devices efficiently. In this work, we propose a complete acceleration scheme, termed ac2SLAM, based on the ORB-SLAM2 algorithm including both front and back ends, and implement it on an FPGA platform. Specifically, the proposed ac2SLAM features with: 1) a scalable and parallel ORB extractor to extract sufficient keypoints and scores for throughput matching with 4% error, 2) a PingPong heapsort component (pp-heapsort) to select the significant keypoints, that could achieve single-cycle initiation interval to reduce the amount of data transfer between accelerator and the host CPU, and 3) the potential parallel acceleration strategies for the back-end optimization. Compared with running ORB-SLAM2 on the ARM processor, ac2SLAM achieves 2.1 × and 2.7 × faster in the TUM and KITTI datasets, while maintaining 10% error of SOTA eSLAM. In addition, the FPGA accelerated front-end achieves 4.55 × and 40 × faster than eSLAM and ARM. The ac2SLAM is fully open-sourced at https://github.com/SLAM-Hardware/acSLAM. Cheng Wang 0045, Yingkun Liu, Kedai Zuo, Jianming Tong, Pengju Ren |
FPT | 6 |
| 2021 | CAQ: Context-Aware Quantization via Reinforcement LearningabstractModel quantization is a crucial step for porting Deep Neural Networks (DNNs) on embedded devices to meet the limited computation and storage resources requirement. Traditional methods usually obtain the scaling factor and quantize the weights based on the information of single layer. However, our analysis indicate that these selection methods of scaling factor overlook the differences and dependencies among layers, leading to large truncation errors or zeroing errors, which is the main reason for the performance degradation. To this end, we propose a Context-Aware Quantization (CAQ) scheme, which formalizes the model quantization as a global optimization problem and leverages reinforcement learning to search for the optimal scaling factors based on the entire model. Further, we adopt shift-based scaling factors to narrow the search space to improve the search efficiency, additionally, it reduces the computational complexity during the inference phase, and also provides a simpler and more robust activation calibration solution. We extensively test our scheme on a wide range of Neural Networks, including ResNet 50/101/152, InceptionV3 and MobileNetV2 on ImageNet, the entire search process only takes about 1 hour on a single GeForce RTX 2080 Ti. Compared with the existed methods, Our scheme can get a better performance, which could maintain the post-quantization accuracy loss less than 0.25%, while reducing memory footprint by 5%-8% and multiply accumulate (MAC) operations by 2%-4%. Besides, we further show that the CAQ can be applied on other tasks, such as object detection and segmentation. Zhijun Tu, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren, Nanning Zheng 0001 |
IJCNN | 5 |
| 2021 | Joint Critics Mechanism: A Universal Framework for Multi-targets Visual NavigationabstractRegarding to target-driven visual navigation problem, training a universal value function or policy function approximator is considered to be a fairly difficult task, because there may exits potential conflicts among different targets. When modeling navigation as a goal-conditional reinforcement learning problem, the algorithm can only support a relatively small number of goals, which limits the universality of the reinforcement learning based methods. In this work, we proposed a framework for multi-targets visual navigation, termed Joint Critics Mechanism, to better train the universal policy function approximator. Recognizing that target-specific network has better convergence, we use the target-specific value network to estimate the advantage of the target-universal policy network for better convergence. In this way, we avoid complexity of learning competitive targets and achieve a better convergence with a larger number of targets. For evaluation, we conduct experiments in realistic simulation environments and the results prove the rationality and effectiveness of our proposed framework. Youzhuo Wang, Wenzhe Zhao 0001, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren |
IJCNN | 5 |
| 2021 | A Lightweight sequence-based Unsupervised Loop Closure DetectionabstractStable, effective and lightweight loop closure detection is an always pursued goal in real-time SLAM systems, that can be ported on embedded processors and deployed on autonomous robotics. Deep learning methods have extended the expressive ability and adaptability of the descriptor, and sequence-based methods can greatly improve the matching accuracy. However, the increased computation complexity and storage bandwidth requirements of matching calculations for high-dimensional descriptor make it infeasible for real-time deployment, especially for robots that navigate in relatively big maps. To address this challenge, we propose a lightweight sequence-based unsupervised loop closure detection scheme. To be specific, Principal Component Analysis (PCA) is applied to squeeze the descriptor dimensions while maintaining sufficient expressive ability. Additionally, with the consideration of the image sequence and combining linear query with fast approximate nearest neighbor search to further reduce the execution time and improve the efficiency of sequence matching. We implement our method on CALC, a state-of-the-art unsupervised solution, and conduct experiments on NVIDIA TX2, results demonstrate that the accuracy has been improved by 5%, while the execution speed is 2× faster. Source code is available at https://github.com/Mingrui-Yu/Seq-CALC. Fan Xiong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
IJCNN | 6 |
| 2021 | PIT: Processing-In-Transmission With Fine-Grained Data Manipulation NetworksabstractIn the domain of data parallel computation, most works focus on data flow optimization inside the PE array and favorable memory hierarchy to pursue the maximum parallelism and efficiency, while the importance of data contents has been overlooked for a long time. As we observe, for structured data, insights on the contents (i.e., their values and locations within a structured form) can greatly benefit the computation performance, as fine-grained data manipulation can be performed. In this paper, we claim that by providing a flexible and adaptive data path, an efficient architecture with capability of fine-grained data manipulation can be built. Specifically, we design SOM, a portable and highly-adaptive data transmission network, with the capability of operand sorting, non-blocking self-route ordering and multicasting. Based on SOM, we propose the processing-in-transmission architecture (PITA), which extends the traditional SIMD architecture to perform some fundamental data processing during its transmission, by embedding multiple levels of SOM networks on the data path. We evaluate the performance of PITA in two irregular computation problems. We first map the matrix inversion task onto PITA and show considerable performance gain can be achieved, resulting in 3x-20x speedup against Intel MKL, and 20x-40x against cuBLAS. Then we evaluate our PITA on sparse CNNs. The results indicate that PITA can greatly improve computation efficiency and reduce memory bandwidth pressure. We achieved 2x-9x speedup against several state-of-art accelerators on sparse CNN, where nearly 100 percent PE efficiency is maintained under high sparsity. We believe the concept of PIT is a promising computing paradigm that can enlarge the capability of traditional parallel architecture. Pengchen Zong, Tian Xia 0008, Jianming Tong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Computers | 8 |
| 2021 | Linear and Nonlinear Regression-Based Maximum Correntropy Extended Kalman FilteringabstractThe extended Kalman filter (EKF) is a method extensively applied in many areas, particularly, in nonlinear target tracking. The optimization criterion commonly used in EKF is the celebrated minimum mean square error (MMSE) criterion, which exhibits excellent performance under Gaussian noise assumption. However, its performance may degrade dramatically when the noises are heavy tailed. To cope with this problem, this paper proposes two new nonlinear filters, namely the linear regression maximum correntropy EKF (LRMCEKF) and nonlinear regression maximum correntropy EKF (NRMCEKF), by applying the maximum correntropy criterion (MCC) rather than the MMSE criterion to EKF. In both filters, a regression model is formulated, and a fixed-point iterative algorithm is utilized to obtain the posterior estimates. The effectiveness and robustness of the proposed algorithms in target tracking are confirmed by an illustrative example. Xi Liu 0006, Hongqiang Lyu, Zhihong Jiang, Pengju Ren, Badong Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object DetectionabstractKeypoint-based detectors have achieved pretty-well performance. However, incorrect keypoint matching is still widespread and greatly affects the performance of the detector. In this paper, we propose CentripetalNet which uses centripetal shift to pair corner keypoints from the same instance. CentripetalNet predicts the position and the centripetal shift of the corner points and matches corners whose shifted results are aligned. Combining position information, our approach matches corner points more accurately than the conventional embedding approaches do. Corner pooling extracts information inside the bounding boxes onto the border. To make this information more aware at the corners, we design a cross-star deformable convolution network to conduct feature adaption. Furthermore, we explore instance segmentation on anchor-free detectors by equipping our CentripetalNet with a mask prediction module. On COCO test-dev, our CentripetalNet not only outperforms all existing anchor-free detectors with an AP of 48.0% but also achieves comparable performance to the state-of-the-art instance segmentation approaches with a 40.2% Mask AP. Code is available at https: //github.com/KiveeDong/CentripetalNet. Zhiwei Dong, Guoxuan Li, Yue Liao, Fei Wang 0032, Pengju Ren, Chen Qian 0006 |
CVPR | 5 |
| 2020 | COCOA: Content-Oriented Configurable Architecture Based on Highly-Adaptive Data Transmission NetworksabstractIn domain of parallel computation, most works focus on optimizing PE organization or memory hierarchy to pursue the maximum efficiency, while the importance of data contents has been overlooked for a long time. Actually for structured data, insights on data contents (i.e. values and locations within a structured form) can greatly benefit the computation performance, as fine-grained data manipulation can be performed. In this paper, we claim that by providing a flexible and adaptive data path, an efficient architecture with capability of fine-grained data manipulation can be built. Specifically, we propose COCOA, a novel content-oriented configurable architecture, which integrates multi-functional data reorganization networks in traditional computing scheme to handle the contents of data during the transmission path, so that they can be processed more efficiently. We evaluate COCOA on various problems: complex matrix algorithm (matrix inversion) and sparse DNN. The results indicates that COCOA is versatile enough to achieve high computation efficiency in both cases. Tian Xia 0008, Pengchen Zong, Jianming Tong, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | Exploring Better Speculation and Data Locality in Sparse Matrix-Vector Multiplication on Intel XeonabstractSparse Matrix-Vector Multiplication (SpMV) is a fundamental workload of numerous applications. However, for today's high-end superscalar CPUs, such as Intel Xeon series, it is usually difficult to efficiently perform SpMV due to the irregular, matrix-dependent data access and computation pattern. While many researches focus on optimizing the memory bandwidth bound by improving data locality, this work dives into the execution of SpMV computation on Intel Xeon CPU and reveals that the bad-speculation penalty is significant in many sparse matrices and too expensive to be ignored. We study and characterize sparsity structure types that are more vulnerable to the cache miss penalty or the bad speculation penalty, respectively. Based on this insight, we proposed a fast preprocessing method, which divides the matrix into sub-matrices and determines the critical performance bound of sub-matrices according to the data distribution characteristics. On each submatrix, a combination of dedicated row reordering strategies is performed to efficiently alleviate its key performance bounds: bad speculation, cache miss, or both. Our matrix representation is based on standard Compressed Sparse Row (CSR) format, and can be easily adapted to existing SpMV libraries. Our approach is evaluated on Intel Xeon Gold 6146 Processor with a wide-range of matrices from the SuiteSparse benchmarks. The results demonstrate that the proposed approach achieves an average 1.8× speedup (up to 2.5×) on multi-threaded MKL Sparse Routines, with a quite low pre-processing cost. Additionally, when used in conjunction with MKL's original optimization method, our approach can further prompt the speedup, to average 3.6 × (up to 8.3 ×), This result indicates that our method can serve as a fast and wide-spectrum optimization method which is compatible with existing routines. Tian Xia 0008, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren |
ICCD | 6 |
| 2019 | Design Space Exploration of Neural Network Activation Function CircuitsabstractThe widespread application of artificial neural networks has prompted researchers to experiment with field-programmable gate array and customized ASIC designs to speed up their computation. These implementation efforts have generally focused on weight multiplication and signal summation operations, and less on activation functions used in these applications. Yet, efficient hardware implementations of nonlinear activation functions like exponential linear units (ELU), scaled ELU (SELU), and hyperbolic tangent (tanh), are central to designing effective neural network accelerators, since these functions require lots of resources. In this paper, we explore efficient hardware implementations of activation functions using purely combinational circuits, with a focus on two widely used nonlinear activation functions, i.e., SELU and tanh. Our experiments demonstrate that neural networks are generally insensitive to the precision of the activation function. The results also prove that the proposed combinational circuit-based approach is very efficient in terms of speed and area, with negligible accuracy loss on the MNIST, CIFAR-10, and IMAGE NET benchmarks. Synopsys design compiler synthesis results show that circuit designs for tanh and SELU can save between ${\times 3.13\sim \times 7.69}$ and ${ {\times 4.45\sim \times 8.45}}$ area compared to the look-up table/memory-based implementations, and can operate at 5.14 GHz and 4.52 GHz using the 28-nm SVT library, respectively. The implementation is available at: https://github.com/ThomasMrY/ActivationFunctionDemo. Tao Yang 0032, Yadong Wei, Zhijun Tu, Haolun Zeng, Michel A. Kinsy, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2018 | Spiking Locality-Sensitive Hash: Spiking Computation with Phase Encoding MethodabstractA novel similarity search method, named spiking locality sensitive hash (SLSH), a forward spiking neuron network(SNN) is proposed in this paper. The SLSH architecture is composed of successively connected encoding and fully connected layer. We optimize phase encoding to maximize the difference between corresponding pixels of any two different images. Then we test the performance of the encoding method and the SLSH model on graphic datasets. Experimental results prove that improved phase encoding method based on the difference exhibits the accuracy of 100%, 100% and 92%, which has superiority over previous phase encoding whose accuracies are 93%, 78% and 55% when the noise level is 5%, 20% and 40% respectively. Furthermore, experiments demonstrate that SLSH method is more capable than the traditional Locality-Sensitive Hash(LSH) and the FLY algorithm published in SCIENCE in similarity search. The mean average precision of SLSH is twice of FLY algorithm when the hash length is 5. In addition, the SLSH achieves a good recognition performance even under the influence of noise for MNIST, SVHN and SIFT datasets. Ziru Wang, Zhiwei Dong, Nanning Zheng 0001, Pengju Ren |
IJCNN | 5 |
| 2018 | A novel spiking neural network of receptive field encoding with groups of neurons decisionabstractHuman information processing depends mainly on billions of neurons which constitute a complex neural network, and the information is transmitted in the form of neural spikes. In this paper, we propose a spiking neural network (SNN), named MD-SNN, with three key features: (1) using receptive field to encode spike trains from images; (2) randomly selecting partial spikes as inputs for each neuron to approach the absolute refractory period of the neuron; (3) using groups of neurons to make decisions. We test MD-SNN on the MNIST data set of handwritten digits, and results demonstrate that: (1) Different sizes of receptive fields influence classification results significantly. (2) Considering the neuronal refractory period in the SNN model, increasing the number of neurons in the learning layer could greatly reduce the training time, effectively reduce the probability of over-fitting, and improve the accuracy by 8.77%. (3) Compared with other SNN methods, MD-SNN achieves a better classification; compared with the convolution neural network, MD-SNN maintains flip and rotation invariance (the accuracy can remain at 90.44% on the test set), and it is more suitable for small sample learning (the accuracy can reach 80.15% for 1000 training samples, which is 7.8 times that of CNN). Ziru Wang, Si-yu Yu, Badong Chen, Nanning Zheng 0001, Pengju Ren |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2017 | Hybrid-augmented intelligence: collaboration and cognitionabstractThe long-term goal of artificial intelligence (AI) is to make machines learn and think like human beings. Due to the high levels of uncertainty and vulnerability in human life and the open-ended nature of problems that humans are facing, no matter how intelligent machines are, they are unable to completely replace humans. Therefore, it is necessary to introduce human cognitive capabilities or human-like cognitive models into AI systems to develop a new form of AI, that is, hybrid-augmented intelligence. This form of AI or machine intelligence is a feasible and important developing model. Hybrid-augmented intelligence can be divided into two basic models: one is human-in-the-loop augmented intelligence with human-computer collaboration, and the other is cognitive computing based augmented intelligence, in which a cognitive model is embedded in the machine learning system. This survey describes a basic framework for human-computer collaborative hybrid-augmented intelligence, and the basic elements of hybrid-augmented intelligence based on cognitive computing. These elements include intuitive reasoning, causal models, evolution of memory and knowledge, especially the role and basic principles of intuitive reasoning for complex problem solving, and the cognitive learning framework for visual scene understanding based on memory and reasoning. Several typical applications of hybrid-augmented intelligence in related fields are given. Nanning Zheng 0001, Ziyi Liu 0001, Pengju Ren, Shi-tao Chen, Si-yu Yu, Jianru Xue, Badong Chen, Fei-Yue Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | A reconfigurable parallel FPGA accelerator for the adapt-then-combine diffusion LMS algorithmabstractThe combination of diffusion strategies and least-mean-square (LMS) algorithm provides many advantages for adaptive-filter to solve distributed optimization, estimation and inference problems. However, suffering from high computation complexity, software implementation of diffusion LMS algorithm is unsuitable for real-time and portable applications. In order to extend its availability, we design a reconfigurable parallel FPG accelerator by exploring multiple dimensions of parallelism, including: parallel execution of agents state updating, data combining, data training and multi-stages pipeline to speedup the execution time. The accelerator for networks with various number of agents and different input dimensions is implemented. Results demonstrate that, it can achieve a speedup of three orders of magnitude at 100Mhz compared with C implementation for a 32-nodes network with 16-dimensional input-data. Qihang Yu, Badong Chen, José C. Príncipe, Nanning Zheng 0001, Pengju Ren |
ISCAS | 6 |
| 2016 | Fault-Aware Load-Balancing Routing for 2D-Mesh and Torus On-Chip Network TopologiesabstractRouting algorithm design for on-chip networks (OCNs) has become increasingly challenging due to high levels of integration and complexity of modern systems-on-chip (SoCs). The inherent unreliability of components, embedded oversized IP blocks, and finegrained voltage-frequency islands (VFIs) management among others, raise several challenges in OCNs: (a) network topologies become irregular or asymmetric making circular route dependencies that lead to deadlock hard to detect; and (b) routing algorithms that lack strong load-balancing properties often saturate prematurely. In order to address the aforementioned deadlock and loadbalancing problems, we propose the traffic balancing oblivious routing (TBOR) algorithm. It is a two-phase routing algorithm consisting of: (1) construction of the weighted acyclic channel dependency graph (CDG) for the OCN to efficiently maximize available resource utilization; and (2) channel ordering across turn models to keep the underlying CDG cycle-free to guarantee deadlock-freedom using one or more turn-models. Channel bandwidth utilization and traffic balancing are achieved through static virtual channel allocation according to residual bandwidth of healthy links. In addition, we introduce in this work two schemes of different granularity of fault detection and analysis while guaranteeing in-order packet delivery by assigning a unique path to each flow. Extensive experiments demonstrate the proposed routing methodology outperforms previous algorithms. Pengju Ren, Michel A. Kinsy, Nanning Zheng 0001 |
IEEE Trans. Computers | 1 |
| 2016 | A Deadlock-Free and Connectivity-Guaranteed Methodology for Achieving Fault-Tolerance in On-Chip NetworksabstractTo improve the reliability of on-chip network based systems, we design a deadlock-free routing technique that is more resilient to component failures and guarantees a higher degree of node connectivity. The routing methodology consists of three key steps. First, we determine the maximal connected subgraph of the faulty network by checking whether the defective components happen to be the cut vertices and bridges of the network topology. A precise fault diagnosis mechanism is used to identify partial defective routers. Second, we construct an acyclic channel dependency graph that breaks all cycles and preserves connectivity of the maximal connected subgraph. This is done through the cycle-breaking and connectivity guaranteed (CBCG) algorithm. Finally, we introduce a fault-tolerant adaptive routing scheme that can be used with or without virtual channels for network congestion avoidance and high-throughput routing. The simulation results show both the effectiveness and robustness of the proposed approach. For an 8 × 8 2D-Mesh with 40 percent of link damage, full connectivity and deadlock freedom are still archived without disabling any faultless router in 98.18 percent of the simulations. In a 2D-Torus, the simulation percentage is even higher (99.93 percent). The hardware overhead for supporting the introduced features is minimal. An on-line implementation of CBCG using TSMC 65nm library has only 0.966 and 1.139 percent area overhead for the 8 × 8 and 16 × 16 2D-Meshes. Pengju Ren, Xiaowei Ren, Sudhanshu Sane, Michel A. Kinsy, Nanning Zheng 0001 |
IEEE Trans. Computers | 1 |
| 2015 | A 128-way FPGA platform for the acceleration of KLMS algorithmabstractThis paper proposes a 128-way parallel FPGA platform to accelerate the kernel least mean square (KLMS) algorithm. With the adoption of a quantized method and pipeline technology, this platform which works at 200MHz is 4827 times faster, on average, than the Matlab code running on a 3GHz Intel(R) Core(TM) i5-2320 CPU. Xiaowei Ren, Qihang Yu, Badong Chen, Nanning Zheng 0001, Pengju Ren |
ASP-DAC | 5 |
| 2015 | A real-time permutation entropy computation for EEG signalsabstractIn this paper, we implement a reconfigurable FPGA accelerator which could compute multiscale permutation entropy for 128 EEG signals simultaneously in real time. When it works at 150MHz and the window size is 256, compared with C code running on a 3GHz Intel(R) Core(TM) i5-2320 CPU, the average speedup is 3748. Xiaowei Ren, Qihang Yu, Badong Chen, Nanning Zheng 0001, Pengju Ren |
ASP-DAC | 5 |
| 2015 | A high efficient hardware architecture for multiview 3DTVabstractThere are three main challenges to design an efficient multiview 3DTV SoC:(1)how to organize DRAM address mapping to maximize off-chip bandwidth utilization;(2)how to design a parallel configurable image scaling engine to interpolate various viewpoints in real-time; (3)how to reduce computational complexity of float-point sub-pixel rearrangement with sufficient accuracy. To this end, we present a highly optimized hardware architecture, which saves 38.4% logic and 37.5% memory resources when implementing a multiview 1080P@60Hz 3DTV on the Xilinx XC5VLX330 FPGA. Pengju Ren |
ASP-DAC | 4 |
| 2014 | Fault-tolerant Routing for On-chip Network Without Using Virtual ChannelsabstractThanks to its less design complexity, less power consumption and service time, to avoid using virtual channel has became a very attractive approach to building future reliable and massively parallel many-core systems. Furthermore, less area of the light-weight router decrease the probability of failure. To this end, by constructing an acyclic channel dependency graph that breaks all cycles and preserves connectivity of the network, we propose a new deadlock-free fault-tolerant adaptive routing without virtual channel. Extensive experiments of 8x8 2D-mesh network demonstrate 99.73% and 97.56% reliability under uniform random traffic when 10% and 20% of the links are failed. Pengju Ren, Xiaowei Ren, Nanning Zheng 0001 |
DAC | 1 |
| 2014 | Hardware implementation of KLMS algorithm using FPGAabstractFast and accurate machine learning algorithms are needed in many physical applications. However, the learning efficiency is badly subjected to the intensive computation. Knowing that hardware implementation could speed up computation effectively, we use a FPGA hardware platform to implement an on-line kernel learning algorithm, namely the kernel least mean square (KLMS) which adopts the simple survival kernel as the Mercer kernel. By using an on-line quantization method and pipeline technology, the requirement of hardware resources and computation burden can be reduced significantly and the data processing speed can be accelerated apparently without losing accuracy. Finally, a 128-way parallel FPGA platform which works at 200MHz is implemented. It could achieve an average speedup of 6553 versus Matlab running on a 3GHz Intel(R) Core(TM) i5-2320 CPU. Xiaowei Ren, Pengju Ren, Badong Chen, Tai Min, Nanning Zheng 0001 |
IJCNN | 2 |
| 2012 | HORNET: A Cycle-Level Multicore SimulatorabstractWe present hornet, a parallel, highly configurable, cycle-level multicore simulator based on an ingress-queued wormhole router network-on-chip (NoC) architecture. The parallel simulation engine offers cycle-accurate as well as periodic synchronization; while preserving functional accuracy, this permits tradeoffs between perfect timing accuracy and high speed with very good accuracy. When run on six separate physical cores on a single die, speedups can exceed a factor of over 5, and when run on a two-die 12-core system with 2-way hyperthreading, speedups exceed$12\times$. Most hardware parameters are configurable, including memory hierarchy, interconnect geometry, bandwidth, crossbar dimensions, parameters driving power, and thermal effects. A highly parametrized table-based NoC design allows a variety of routing and virtual channel allocation algorithms out of the box, ranging from simple dimension-ordered routing to complex Valiant, ROMM, O1Turn or PROM schemes, BSOR, and adaptive routing. Hornet can run in network-only mode using synthetic traffic or traces, or directly emulate a MIPS-based multicore. Hornet is freely available under the open-source MIT license at http://csg.csail.mit.edu/hornet/. Pengju Ren, Mieszko Lis, Myong Hyon Cho, Keun Sup Shim, Christopher W. Fletcher, Omer Khan, Nanning Zheng 0001, Srini Devadas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Scalable, accurate multicore simulation in the 1000-core eraabstractWe present HORNET, a parallel, highly configurable, cycle-level multicore simulator based on an ingress-queued worm-hole router NoC architecture. The parallel simulation engine offers cycle-accurate as well as periodic synchronization; while preserving functional accuracy, this permits tradeoffs between perfect timing accuracy and high speed with very good accuracy. When run on 6 separate physical cores on a single die, speedups can exceed a factor of over 5, and when run on a two-die 12-core system with 2-way hyperthreading, speedups exceed 11 ×. Most hardware parameters are configurable, including memory hierarchy, interconnect geometry, bandwidth, crossbar dimensions, and parameters driving power and thermal effects. A highly parametrized table-based NoC design allows a variety of routing and virtual channel allocation algorithms out of the box, ranging from simple DOR routing to complex Valiant, ROMM, or PROM schemes, BSOR, and adaptive routing. HORNET can run in network-only mode using synthetic traffic or traces, directly emulate a MIPS-based multicore, or function as the memory subsystem for native applications executed under the Pin instrumentation tool. HORNET is freely available under the open-source MIT license at http://csg.csail.mit.edu/hornet/. Mieszko Lis, Pengju Ren, Myong Hyon Cho, Keun Sup Shim, Christopher W. Fletcher, Omer Khan, Srini Devadas |
ISPASS | 2 |