Man Shi

dblp:252/5454 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SunPar: An Analytical Design Space Exploration Framework Modeling Performance Uncertainty in Sparse AI Accelerators
Jiacong Sun, Man Shi, Mahesh Subedar, Georges Gielen, Marian Verhelst
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 FlexiGen: An Automated AI Accelerator Generation Framework With Decoupled-Access-Execute and Dynamic Dataflows
abstract
Modern tensor applications, especially artificial intelligence (AI) applications, are evolving rapidly, posing a significant demand for agile hardware design. While numerous hardware generators have been developed, they suffer from three significant limitations: 1) they are either limited to a single dataflow/data type generation, failing to cater to the computational requirements of diverse workloads; 2) or focus only on the array level optimization, omitting system-level effects, such as the influence of on-chip memory bandwidth and contention; and 3) customized workload mapping/configuration is needed, resulting in increased programming complexity. To address these challenges, we proposeFlexiGen, a flexible and extensible hardware generation framework, which targets diverse deep neural networks (DNN) tensor applications and can generate a complete synthesizable acceleration system at the RTL level with arbitrary dataflow and its combinations. Our key contributions are threefold: 1) we incorporate decoupled-access-execute architecture insideFlexiGen, enabling full system generation while maintaining flexibility and efficiency; 2) we propose a versatile spatial core generator that supports dynamic spatial dataflows and multiple data precisions in the same array and a compatible data streaming engine generator that can support arbitrary temporal dataflows and$N$-dimensional data access patterns; and 3) we leverage a uniform programming interface and provide a customized kernel library, enabling agile configuration programming. We conduct an intensive evaluation to demonstrate the versatility ofFlexiGenin dataflow accelerator generation and show the trade-offs of performance, area, and power across a wide range of dataflows and workloads at both the array level and system level. Our case study experiment showsFlexiGen’s usefulness as a hardware generator to rapidly generate desired dataflow acceleration systems. Compared with the state-of-the-art (SotA) hardware generation framework LEGO,FlexiGenachieves 36.79% and 57.16% less area and power when generating the same dual spatial dataflow design.FlexiGenis open-source and available athttps://github.com/KULeuven-MICAS/snax_cluster
Xiaoling Yi, Man Shi, Joren Dumoulin, Jiacong Sun, Yunhao Deng, Ryan Antonio, Fanchen Kong, Marian Verhelst
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
abstract
The widely-used, weight-only quantized large language models (LLMs), which leverage low-bit integer (INT) weights and retain floating-point (FP) activations, reduce storage requirements while maintaining accuracy. However, this shifts the energy and latency bottlenecks towards the FP activations that are associated with costly memory accesses and computations. Existing LLM accelerators focus primarily on computation optimizations, overlooking the potential of jointly optimizing FP computations and data movement, particularly for the dominant FP-INT GeMM operations in LLM inference. To address these challenges, we investigate the sensitivity of activation precision across various LLM modules and its impact on overall model accuracy. Based on our findings, we first propose the Anda data type: an adaptive data format with group-shared exponent bits and dynamic mantissa bit allocation. Secondly, we develop an iterative post-training adaptive precision search algorithm that optimizes the bit-width for different LLM modules to balance model accuracy, energy efficiency, and inference speed. Lastly, a suite of hardware optimization techniques is proposed to maximally exploit the benefits of the Anda format. These include a bit-plane-based data organization scheme, Anda-enhanced processing units with bit-serial computation, and a runtime bit-plane Anda compressor to simultaneously optimize storage, computation, and memory footprints. Our evaluations on FP-INT GeMM operations show that Anda achieves a $2.4 \times$ speedup, $4.0 \times$ area efficiency, and $3.1 \times$ energy efficiency improvement on average for popular LLMs including OPT, LLaMA, and LLaMA-2 series over the GPU-like FP-FP baseline. Anda demonstrates strong adaptability across various application scenarios, accuracy requirements, and system performance, enabling efficient LLM inference across a wide range of deployment scenarios.
Chao Fang 0005, Man Shi, Robin Geens, Arne Symons, Zhongfeng Wang 0001, Marian Verhelst
HPCA2
2024 BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Acceleration
abstract
Bit-serial computation facilitates bit-wise sequential data processing, offering numerous benefits, such as a reduced area footprint and dynamically-adaptive computational precision. It has emerged as a prominent approach, particularly in leveraging bit-level sparsity in Deep Neural Networks (DNNs). However, existing bit-serial accelerators exploit bit-level sparsity to reduce computations by skipping zero bits, but they suffer from inefficient memory accesses due to the irregular indices of the non-zero bits. As memory accesses typically are the dominant contributor to DNN accelerator performance, this paper introduces a novel computing approach called “bit-column-serial” and a compatible architecture design named “BitWave.” BitWave harnesses the advantages of the “bit-column-serial” approach, leveraging structured bit-level sparsity in combination with dynamic dataflow techniques. This achieves a reduction in computations and memory footprints through redundant computation skipping and weight compression. BitWave is able to mitigate the performance drop or the need for retraining that is typically associated with sparsity-enhancing techniques using a post-training optimization involving selected weight bit-flips. Empirical studies conducted on four deep-learning benchmarks demonstrate the achievements of BitWave: (1) Maximally realize 13.25x higher speedup, 7.71 x efficiency compared to state-of-the-art sparsity-aware accelerators. (2) Occupying 1.138 mm2area and consuming 17.56 mW power in 16nm FinFet process node.
Man Shi, Vikram Jain, Antony Joseph, Maurice Meijer, Marian Verhelst
HPCA1
2024 Learning Based Model Predictive Path Tracking Control for Autonomous Buses
abstract
In addressing the trade-off between prediction model accuracy and computational cost in the context of path tracking control, this paper proposes a learning-based model predictive control (LB-MPC) strategy for autonomous buses. A three-degree-of-freedom (DOF) single-track vehicle dynamic model is established, and an in-depth analysis is conducted on its step response error with respect to variations in vehicle speed, pedal position, and front wheel steering angle compared to the IPG TruckMaker model. Methods for constructing error datasets and receding horizon updates are designed, and a Gaussian process regression (GPR) is employed to establish an error fitting model for real-time error compensation and correction of the nominal single-track model. The error correction model is utilized as the prediction model, and a path tracking cost function is designed to formulate a quadratic programming (QP) optimization problem, proposing an LB-MPC path tracking control architecture. Through joint simulations using the IPG TruckMaker & Simulink platform and real bus experiment, the real-time performance and effectiveness of the proposed GPR error correction model and LB-MPC path tracking control strategy are verified. Results demonstrate that compared to traditional MPC path tracking control strategy, the proposed LB-MPC strategy reduces the average path tracking error by 79.00%.
Mo Han, Hongwen He, Jianfei Cao, Jingda Wu, Wei Liu 0058, Man Shi
IV6
2023 From Grayscale Image to Battery Aging Awareness - A New Battery Capacity Estimation Model With Computer Vision Approach
abstract
Accurate detection of capacity degradation is critical to the safe and efficient utilization of battery systems. Many data-driven capacity estimators were proposed based on emerging intelligent algorithms, but their accuracy depends on the data of complete charged/discharged process and complex algorithm structures. This article developed a computer vision (CV)-based method, constructing battery multidimensional aging features as the key image to estimate capacity using specific charging data segment. Specifically, the designed image-aging recognition method is used to extract multidimensional aging features from the partial charging current sequence and then establish map inputs for a computer vision model that recognizes the constructed feature maps. Consequently, the mapping relationship between the charging information and capacity degradation can be obtained as the 2-D grayscale images that contain massive extracted features in their small size hence greatly simplify the network structure in CV model so as to improve estimation accuracy and efficiency significantly. More importantly, since the model input is a specific charging current segment rather than the data of complete charging process, the model applicability to the random and incomplete charging process of electric vehicles can be greatly improved. Battery cycling data from different types of Li-ion cells were utilized for performance verification. Compared with the conventional estimation methods proposed previously, the proposed method demonstrates the great superiority in terms of the model applicability, estimation accuracy, and computational efficiency for online capacity estimation in actual battery usage.
Hongwen He, Jianwei Li 0001, Zhongbao Wei, Ruchen Huang, Man Shi
IEEE Trans. Ind. Informatics6
2023 COAC: Cross-Layer Optimization of Accelerator Configurability for Efficient CNN Processing
abstract
To achieve high accuracy, convolutional neural networks (CNNs) are increasingly growing in complexity and diversity in layer types and topologies. This makes it very challenging to efficiently deploy such networks on custom processor architectures for resource-scarce edge devices. Existing mapping exploration frameworks enable searching for the optimal execution schedules or hardware mappings of individual network layers, by optimizing each layer’s spatial (dataflow parallelization) and temporal unrolling (TU, execution order). However, these tools fail to take into account the overhead of supporting different unrolling schemes within a common hardware architecture. Using a fixed unrolling scheme across all layers is also not ideal, as this misses significant opportunities for energy and latency savings from optimizing the mapping of diverse layer types. A balanced approach assesses the right amount of mapping flexibility needed across target neural networks, while taking into account the overhead to support multiple unrollings. This article, therefore, presents cross-layer optimization of accelerator configurability (COAC), a cross-layer design space exploration and mapping framework to optimize the flexibility of neural processing architectures by balancing configurability overhead against resulting energy and latency savings for end-to-end inference. COAC does not only provide a systematical analysis of the architectural overhead in function of the supported spatial unrollings (SUs), but also builds an automated flow to find the best unrolling combination(s) for efficient end-to-end inference with limited hardware overhead. Results demonstrate that architectures with carefully optimized flexibility can achieve up to 38% energy-delay-product (EDP) savings for a set of six neural networks at the expense of a relative area increase of 9.5%.
Steven Colleman, Man Shi, Marian Verhelst
IEEE Trans. Very Large Scale Integr. Syst.2
2020 STC: Significance-aware Transform-based Codec Framework for External Memory Access Reduction
abstract
Deep convolutional neural networks (DCNNs), with extensive computation, require considerable external memory bandwidth and storage for intermediate feature maps. External memory accesses for feature maps become a significant energy bottleneck for DCNN accelerators. Many works have been done on quantizing feature maps into low precision to decrease the costs for computation and storage. There is an opportunity that the large amount of correlation among channels in feature maps can be exploited to further reduce external memory access. Towards this end, we propose a novel compression framework called Significance-aware Transform-based Codec (STC). In its compression process, significance-aware transform is introduced to obtain low-correlated feature maps in an orthogonal space, as the intrinsic representations of original feature maps. The transformed feature maps are quantized and encoded to compress external data transmission. For the next layer computation, the data will be reloaded with STC's reconstruction process. The STC framework can be supported with a small set of extensions to current DCNN accelerators. We implement STC extensions to the baseline TPU architecture for hardware evaluation. The strengthened TPU achieves average reduction of 2.57x in external memory access, 1.95x~2.78x improvement of system-level energy efficiency, with a negligible accuracy loss of only 0.5%.
Fengbin Tu, Man Shi, Yang Wang 0089, Leibo Liu, Shaojun Wei, Shouyi Yin
DAC3