Juin-Ming Lu

dblp:58/794 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-4219-1349ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Low-complexity and Reconfigurable Design for Nonlinear Function Approximation in Transformers
abstract
Nonlinear function approximation has been studied for decades. With the rise of transformer-based AI models, the need for efficient, low-complexity circuit implementations for functions like softmax, GELU, and layer normalization has intensified due to non-negligible hardware overhead. Existing methods reuse softmax hardware for GELU or reconfigure both, but the preprocessing required for their GELU approximation results in area inefficiency. To address this, we propose a novel successive approximation technique that reduces preprocessing complexity and compensates for errors in successive steps. Additionally, our reconfigurable design supports softmax, GELU, and square root functions, optimizing hardware area and flexibility. Experimental results show a 2.48x increase in throughput per area for softmax and a 4.96x increase for GELU, with only a 0.09% accuracy loss under BERT-base models in comparison to related work.
Qi-Xian Wu, Shu-Sian Teng, Ming-Der Shieh, Chih-Tsun Huang, Juin-Ming Lu
ISCAS5
2023 MultiFuse: Efficient Cross Layer Fusion for DNN Accelerators with Multi-level Memory Hierarchy
abstract
In order to facilitate the deployment of diverse deep learning models while maintaining scalability, modern DNN accelerators frequently employ reconfigurable structures such as Network-on-Chip (NoC) and multi-level on-chip memory hierarchy. To achieve high energy efficiency, it is imperative to store intermediate DNN-layer results within the on-chip memory hierarchy, thereby reducing the need for off-chip data transfers to/from the DRAM memory.Two well-established optimization techniques, node fusion and loop tiling, have proven effective in retaining temporary results within the on-chip buffers, commonly used to minimize off-chip DRAM accesses. In this paper, we introduce MultiFuse, an infrastructure designed to automatically explore multiple DNN layer node fusion techniques, enabling optimal utilization of the on-chip multi-level memory hierarchy.Experimental results demonstrate the effectiveness of our retargetable infrastructure, which outperforms Ansor’s algorithm. Our exploration algorithm achieves a remarkable 70% reduction in Energy-Delay Product (EDP) while gaining a 67x speedup in search time when executing the data-intensive MobileNet model on a single DNN accelerator.
Chia-Wei Chang, Jing-Jia Liou, Chih-Tsun Huang, Wei-Chung Hsu, Juin-Ming Lu
ICCD5
2023 Optimization of AI SoC with Compiler-assisted Virtual Design Platform
abstract
As deep learning keeps evolving dramatically with rapidly increasing complexity, the demand for efficient hardware accelerators has become vital. However, the lack of software/hardware co-development toolchains makes designing AI SoCs (artificial intelligent system-on-chips) considerably challenging. This paper presents a compiler-assisted virtual platform to facilitate the development of AI SoCs from the early design stage. The electronic system-level design platform provides rapid functional verification and performance/energy analysis. Cooperating with the neural network compiler, AI software and hardware can be co-optimized on the proposed virtual design platform. Our Deep Inference Processor is also utilized on the virtual design platform to demonstrate the effectiveness of the architectural evaluation and exploration methodology.
Chih-Tsun Huang, Juin-Ming Lu, Yao-Hua Chen, Ming-Chih Tung, Shih-Chieh Chang 0001
ISPD2
2022 Fault Modeling and Testing of Memristor-Based Spiking Neural Networks
abstract
The resistive random-access memory (RRAM), whose core is composed of memristor cell arrays, has recently been proposed for implementing the deep neural network (DNN) and spiking neural network (SNN), which potentially can improve the performance and energy efficiency of AI computing. In this paper, we address high-quality memristor-based SNNs. We propose two functional fault models for them, i.e., the Slow Integration Fault (SIF) and Fast Integration Fault (FIF), and show the circuit-level fault simulation results, taking process variation into account. The detailed simulation and analysis results show that the SIF and FIF are feasible functional fault models for memristor-based SNNs, as the interconnect open and short defects, as well as the memristor and transistor faults, can all be covered by the proposed SIF and FIF. We have also developed a test algorithm for the SNNs, i.e., the SNN-Test algorithm. Experimental results show that the SNN-Test covers 100% of the open and short defects and the proposed functional fault models.
Kuan-Wei Hou, Hsueh-Hung Cheng, Chi Tung, Cheng-Wen Wu, Juin-Ming Lu
ITC5
2021 An Improved STBP for Training High-Accuracy and Low-Spike-Count Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) that facilitate energy-efficient neuromorphic hardware are getting increasing attention. Directly training SNN with backpropagation has already shown competitive accuracy compared with Deep Neural Networks. Besides the accuracy, the number of spikes per inference has a direct impact on the processing time and energy once employed in the neuromorphic processors. However, previous direct-training algorithms do not put great emphasis on this metric. Therefore, this paper proposes four enhancing schemes for the existing direct-training algorithm, Spatio-Temporal Back-Propagation (STBP), to improve not only the accuracy but also the spike count per inference. We first modify the reset mechanism of the spiking neuron model to address the information loss issue, which enables the firing threshold to be a trainable variable. Then we propose two novel output spike decoding schemes to effectively utilize the spatio-temporal information. Finally, we reformulate the derivative approximation of the non-differentiable firing function to simplify the computation of STBP without accuracy loss. In this way, we can achieve higher accuracy and lower spike count per inference on image classification tasks. Moreover, the enhanced STBP is feasible for the on-line learning hardware implementation in the future.
Pai-Yu Tan, Cheng-Wen Wu, Juin-Ming Lu
DATE3
2021 Thermal-Aware Fixed-Outline Floorplanning Using Analytical Models With Thermal-Force Modulation
abstract
High temperature or temperature nonuniformity has become a serious threat to performance and reliability of high-performance integrated circuits (ICs), which makes the thermal effect turn into a nonignorable issue in the circuit design or the physical design. In order to estimate temperature accurately, the locations of modules have to be determined in advance, which makes an efficient and effective thermal-aware floorplanning play a more important role. Hence, this article proposes a differentiable nonlinear placement model that can optimize temperature and minimize wirelength at the same time without needing to construct a congestion map. In addition, to avoid inducing longer wirelength while optimizing temperature, we propose some techniques, such as thermal-aware clustering, shrink of hot modules, or thermal-force modulation in the multilevel framework. The experimental results demonstrate that temperature and wirelength are greatly improved by our method compared to Corblivar. More importantly, our runtime is quite fast and the fixed-outline constraint can also be satisfied.
Jai-Ming Lin, Tai-Ting Chen, Hao-Yuan Hsieh, Ya-Ting Shyu, Yeong-Jar Chang, Juin-Ming Lu
IEEE Trans. Very Large Scale Integr. Syst.6
2021 Thermal-Aware Floorplanning and TSV-Planning for Mixed-Type Modules in a Fixed-Outline 3-D IC
abstract
High temperature or temperature nonuniformity has become a serious threat to performance and reliability of high-performance integrated circuits (ICs). Since the temperature in a 3-D IC is mainly determined by a power distribution across tiers, this article proposes a novel method to allocate modules to tiers while exploring a better power distribution according to the hyperparameter optimization technique. In addition, we integrate the sub-3-D thermal mask into the analytical formulation to interleave high power consumption modules in contiguous tiers during distributing modules over placement regions so that it is possible to insert through-silicon vias (TSVs) around high power modules to further reduce the temperature at a later stage. Since the temperature in a tier will be changed every time the locations of TSVs in its lower tier are moved, we also propose a procedure to update the temperature map before refining locations of TSVs. Experimental results have demonstrated that the proposed methodology can effectively reduce the temperature of a 3-D IC with a slight increase in the wirelength. Moreover, its runtime is quite fast.
Jai-Ming Lin, Wei-Yi Chang, Hao-Yuan Hsieh, Ya-Ting Shyu, Yeong-Jar Chang, Juin-Ming Lu
IEEE Trans. Very Large Scale Integr. Syst.6
2020 A 90nm 103.14 TOPS/W Binary-Weight Spiking Neural Network CMOS ASIC for Real-Time Object Classification
abstract
This paper introduces a low-power 90nm CMOS binary weight spiking neural network (BW-SNN) ASIC for real-time image classification. The chip maximizes data reuse through systolic arrays that house the entire 5-layer BW-SNN, requiring a minimum off-chip bandwidth for data access. The chip achieves 97.57% accuracy for real-time bottled-drink recognition, consuming only 0.62uJ per inference. For comparison purpose, it achieves 98.73% accuracy for MNIST hand-written character recognition, consuming only 0.59uJ per inference. The bottled-drink recognition is demonstrated at 300 fps that is well enough for many other real-time applications. The peak efficiency point is 103.14TOPS/W at a voltage of 0.6V, which outperforms other designs so far as we know. By normalizing to the 28nm technology node, the proposed ASIC is about 5× more efficient and 7× lower hardware cost as compared with the state-of-the-art designs.
Po-Yao Chuang, Pai-Yu Tan, Cheng-Wen Wu, Juin-Ming Lu
DAC4
2018 A fast thermal-aware fixed-outline floorplanning methodology based on analytical models
abstract
Today, secure systems are built by identifying potential vulnerabilities and then adding protections to thwart the associated attacks. Unfortunately, the complexity of today's systems makes it impossible to prove that all attacks are stopped, so clever attackers find a way around even the most carefully designed protections. In this article, we take a sobering look at the state of secure system design, and ask ourselves why the “security arms race” never ends? The answer lies in our inability to develop adequate security verification technologies. We then examine an advanced defensive system in nature – the human immune system – and we discover that it does not remove vulnerabilities, rather it adds offensive measures to protect the body when its vulnerabilities are penetrated We close the article with brief speculation on how the human immune system could inspire more capable secure system designs.
Jai-Ming Lin, Tai-Ting Chen, Yen-Fu Chang, Wei-Yi Chang, Ya-Ting Shyu, Yeong-Jar Chang, Juin-Ming Lu
ICCAD7
2018 Full System Emulation of Embedded Heterogeneous Multicores Based on QEMU
abstract
The emerging edge computing is poised to move computing and intelligence to the network's edge so as to be close to the data sources for fast responses and reduced network traffic. In edge computing, edge devices need to encompass a wide variety of applications or services, from data preprocessing, intelligence inference, to multimedia human interface. Many such applications are well suited for special-purpose hardware accelerators. With the increasing number of accelerators on the edge devices, a promising architecture for edge devices is an asymmetric heterogeneous multicore that incorporates one or more microcontrollers to offload accelerator scheduling and interrupt handling from the main CPU, as exemplified in the NVIDIA Deep Learning Accelerator (NVDLA). To develop such computing systems, virtual platforms such as QEMU are often used. Unfortunately, QEMU only supports symmetric homogeneous multicore systems. In this paper, we tackle the challenging problem of supporting asymmetric heterogeneous multicore systems on QEMU by considering two possible implementation strategies: one-process and multi-process. The challenges are discussed and our implementations are presented. The two approaches are then compared qualitatively and quantitatively.
I-Hua Chen, Chung-Ta King, Yao-Hua Chen, Juin-Ming Lu
ICPADS4
2000 Cost and Benefit Models for Logic and Memory BIST
abstract
We present cost and benefit models and analyze the economics effects of built-in self-test (BIST) for logic and memory cores. In our cost and benefit models for BIST, we take into consideration the design verification time and test development time associated with testability. Experimental results for logic BIST and memory BIST examples show that a threshold volume exists when BIST is profitable for the logic core under consideration - it is not recommended for a higher volume. However, BIST is a good choice for memory cores in general.
Juin-Ming Lu, Cheng-Wen Wu
DATE1