VLDB 2026 Research / reviewers in the wild / expert
Kun-Chih Chen
dblp:16/7595 · also Kun-Chih Jimmy Chen
· DBLP profile ↗
27ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0002-8908-468XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 14 first-author · 12 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-speed Bufferless NoC-based LLM Accelerator with Distributed Ring-Shaped Buffer
Bo-Chun Chen, Kun-Chih Chen |
ISCAS | 2 |
| 2026 | Direction-based Route Integrity Validation for Eavesdropping Prevention in NoC Systems
Hsuan-Yu Huang, Kun-Chih Chen |
ISCAS | 2 |
| 2026 | Entropy-Based Thermal Sensor Placement and Temperature Reconstruction Based on Adaptive Compressive Sensing Theory
Kun-Chih Chen, Chia-Hsin Chen, Lei-Qi Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | High Reliable and Accurate Stochastic Computing-based Artificial Neural Network Architecture DesignabstractHardware reliability has emerged as a paramount consideration in the modern Artificial Neural Network (ANN) design in recent years. The un-reliable ANN leads to un-trustable inference results and disasters (e.g., finance system or transportation system crash). To ensure hardware reliability, the common way is to insert some error correctness blocks or fault-tolerant computing blocks, which bring considerable hardware overhead and are improper to the resource-limited edge AI designs. To design hardware-friendly and highly reliable hardware, the Stochastic Computing (SC) method has been proven to be an efficient way to achieve fault-tolerant computing goals. Consequently, many SC-based computing architectures have been introduced recently. However, because of the stochastic number representation, the computing accuracy issue is the design challenge to implement the SC-based computing architecture. To solve this problem, we propose a novel scaling-free adder and input data pre-processing method to achieve a reliable SC-based computing architecture and improve the accuracy of conventional SC-based ANN design. Compared with the traditional ANN design, the proposed SC-based ANN design maintains computing accuracy and enhances the performance by 32% to 55% while facing serious fault injection. In addition, the proposed SC-based ANN architecture reduces 48% to 81% power consumption and 51% to 92% area cost compared with the conventional ANN architecture. Kun-Chih Chen, Wei-Ren Syu |
ISCAS | 1 |
| 2024 | Q-learning Assisted LASSO-based Thermal Sensor Placement for Thermal-aware Multi-core SystemsabstractThermal problems become severe in contemporary multi-core systems because of the complicated workload and high power density. The problems impact the system performance and damage the system reliability. The practical way to monitor the system temperature is to place number-limited thermal sensors on multi-core systems. Unfortunately, finding proper locations for thermal sensor placement is an NP-hard problem. Many pieces of research proposed methods to allocate number-limited thermal sensors under different perspectives. However, the conventional methods still do not consider the time-varying temperature distribution or the correlation between different thermal hotspot points. To solve these problems, we apply the Q-learning method to assist with thermal sensor placements determined by the Least Absolute Shrinkage and Selection Operator (LASSO) theory. While LASSO primarily excels in feature selection, Q-learning is better equipped to handle dynamic environmental changes and interdependencies. Thus, we employ the Q-learning method to formulate a precise cost function for accurately estimating the outcomes following the placement of a thermal sensor at a specific location. Compared with the state-of-the-art, our proposed methods, using both the LASSO-based method and the Q-learning Assisted LASSO method, reduce the average errors by 67%-85% and 69%-87%, respectively, and the maximum errors by 80%-92% and 82%-93%. Kun-Chih Chen, Leiqi Wang |
ISCAS | 1 |
| 2024 | Hardware Accelerator for MobileViT Vision Transformer with Reconfigurable ComputationabstractWith the great success of the Transformer model in Natural Language Processing (NLP), Vision Transformer (ViT) was proposed achieving comparable performance to traditional Convolutional Neural Network (CNN) models in tasks such as image classification and object detection. This paper focuses on the acceleration of a new lightweight hybrid model, named MobileViT, which has less computation complexity and higher accuracy compared with ViT and other CNN-based lightweight models such as MobileNets. We introduce an adaptive systolic array (SA) design with a flexible shape size, called LEGO SA, that enhances the efficiency of hardware utilization and memory accesses during standard convolution, Depth-wise Separable Convolution (DWC), and self-attention operations. Furthermore, matrix transpose in self-attention is implemented efficiently with significantly reduced wastage of execution time, memory buffers, and power consumption. The proposed MobileViT hardware accelerator with 112KB on-chip buffers occupies an area of just 1.64mm^2 on the TSMC 40nm process, and achieves a performance of 1.2 TOPS at 600 MHz with energy efficiency of 5.34 TOPS/W. Shen-Fu Hsiao, Tzu-Hsien Chao, Yen-Che Yuan, Kun-Chih Chen |
ISCAS | 4 |
| 2024 | Neural Network Acceleration Using Digit-Plane Computation with Early TerminationabstractDeep neural network (DNN) hardware accelerator designs can be divided into two categories, bit-parallel and bit-serial, depending on whether the input of the multiplication is in bit-parallel or bit-serial style. Bit-serial DNN designs with only shift-add operations are more flexible for computation when the per-layer bit-widths are different in aggressive model quantization. Rectified Linear Unit (ReLU) is a common non-linear activation function after each layer of DNN computation where negative values are replaced by zeros. Thus, early-termination could be employed to reduce ineffective computation, resulting in improved speed performance and reduced energy consumption. In this paper, we propose a new bit-serial computation style by re-organizing data in bit-plane order. We observe that bit-plane computation allows more efficient early-termination. The proposed bit-plane DNN design is also extended to digit-plane through radix-4 Booth recoding, leading to a smaller area cost. Experimental results show that the speed performance of the proposed designs is 2.25x and 1.48x times that of the bit-parallel and bit-serial designs respectively when executing the VGG-16 model. Shen-Fu Hsiao, Hou-Chun Kuo, Yu Kuo, Kun-Chih Chen |
ISCAS | 4 |
| 2023 | Adaptive Machine Learning-Based Proactive Thermal Management for NoC SystemsabstractBecause of the high-complex interconnection in contemporary multicore systems, the network-on-chip (NoC) technology has been proven as an efficient way to solve the communication problem in multicore systems. However, the thermal problem becomes the main design challenge in the current NoC systems due to the high-diverse workload distribution and large power density. Therefore, proactive dynamic thermal management (PDTM) is employed as an efficient way to control the system temperature. Based on the predicted temperature information, the PDTM can control the system temperature in advance to reduce the performance impact during the temperature control period. However, conventional temperature prediction models are usually built based on specific physical parameters, which are usually temperature-sensitive. Consequently, the current temperature prediction models still result in significant temperature prediction errors. To solve this problem, a novel adaptive machine learning (ML)-based PDTM is proposed in this work. The adaptive ML-based PDTM first uses an adaptive single layer perceptron (ASLP), which is composed of a single-neuron operation and a least mean square (LMS) adaptive filter technology, to precisely predict the future temperature. Afterward, the proposed adaptive reinforcement learning (RL) is used to find the proper throttling ratio to control the system temperature. In this way, the proposed adaptive ML-based PDTM can adapt to the hyperplane of the temperature behavior of the NoC system and provide a proper temperature control strategy at runtime. Compared with related works, the proposed approach reduces average temperature prediction error by 0.2%–78.0% and improves the system performance by 2.4%–43.0% with smaller hardware overhead. Kun-Chih Chen, Yuan-Hao Liao, Cheng-Ting Chen, Leiqi Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Entropy-based Thermal Sensor Allocation for Temperature-aware Multi-core PlatformsabstractBecause of the high design complexity in multicore systems, contemporary multi-core systems usually suffer from serious thermal issues caused by the large variety of workloads. To monitor the heat phenomenon, the cost-efficient way is to allocate number-limited thermal sensors on the multicore system. Consequently, a lot of methods were proposed to find proper locations for thermal sensors allocation in the recent decade. However, due to the time-varying characteristic of the temperature behavior on the multi-core system, the temperature distribution usually changes along with time. Besides, the temperature distribution also depends on the target application on the multi-core system, which increases the difficulty to find the proper locations for the thermal sensor allocation. The improper locations for the thermal error while using the sensing information from the allocated error while using the sensing information from the allocated art, we propose an entropy-based thermal sensor allocation method, which aims to find locations to cover as many different temperature behaviors on the system as possible. In this way, we can apply the Restricted Isometry Property (RIP) to reconstruct the full-chip temperature distribution efficiently based on the number-limited thermal sensing information. Compared with the previous thermal sensor allocation methods, we can reduce the average full-chip temperature reconstruction error by 3% to 93%. In addition, the maximum error can be reduced by 3% to 96% as well. Kun-Chih Chen, Chia-Hsin Chen |
ISCAS | 1 |
| 2022 | Thermal Sensor Placement for Multicore Systems Based on Low-Complex Compressive Sensing TheoryabstractAs the complexity of the multicore system grows, the large workload diversity results in serious thermal problems. In a practical way, the number of placed thermal sensors is usually limited due to the manufacturing cost. In recent years, the compressive sensing (CS) theory is proven as an efficient way to reconstruct the original signal by using fewer sampling data. However, due to the high computational complexity during the signal reconstruction, the CS theory is not appropriate to apply to the real-time temperature monitoring in the current multicore system. In this article, we propose a grid-based sensor placement approach to placement the number-limited thermal sensors on the target multicore system. On the other hand, we adopt the matrix inversion bypass (MIB) property to reduce the computational complexity of two widely used signal reconstruction approaches in CS theory [i.e., the orthogonal matching pursuit (OMP) and stagewise OMP (StOMP)]. Due to the characteristic of random sampling in CS theory, the complexity of thermal sensor placement for multicore systems can be reduced significantly. In addition, the proposed MIB-based temperature reconstruction method helps to satisfy the requirement of real-time temperature estimation. The experimental results show that the proposed approach can reduce 57%–93% average full-system temperature reconstruction error compared with the previous non-CS-based approaches. Besides, we can also reduce 22%–41% computing latency compared with the current CS-based reconstruction algorithm. Due to the MIB-based operation, we can bypass the matrix inversion operation for temperature reconstruction. Therefore, the hardware overhead of the temperature reconstruction unit can be reduced significantly. Compared with the conventional approaches, we can reduce 24%–87% area overhead and improve 50%–220% hardware efficiency. Kun-Chih Chen, Hsueh-Wen Tang, Chi-Hsun Wu, Chia-Hsin Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | High-Accurate Stochastic Computing for Artificial Neural Network by Using Extended Stochastic LogicabstractThe Artificial Neural Network (ANN) already shows the superiority in many real-world applications. However, due to the high dense neuron computing, the power issue becomes the design challenge of the ANN hardware implementation. On the other hand, the Stochastic Computing (SC) method has been proven as an efficient way to substitute the high-power arithmetic unit through stochastic bit-stream- based computing. Therefore, many SC-based ANN designs were proposed in recent years. However, due to the stochastic bitstream computing, the conventional SC-based ANN designs suffer from low computing accuracy. In this work, we apply the Extended Stochastic Logic (ESL) method to solve the accuracy problem of the conventional SC-based ANN designs. Because the ESL method supports a wider input coding range for the SC process, the computing accuracy can be improved. With this design concept, we propose an ESL-based adder to substitute the accumulation process in ANN computing. Furthermore, an ESL-based ReLU function is proposed to be used as the involved activation function instead. Compared with the conventional SC-based approaches, the proposed ESL-based ANN approach can help to improve the system accuracy by 48%. In addition, compared with the non-SC-based ANN, the proposed ESL- based ANN can reduce 84% area cost and 60% power consumption. Kun-Chih Chen, Chi-Hsun Wu |
ISCAS | 1 |
| 2021 | A Fast ECG Diagnosis by Using Non-Uniform Spectral Analysis and the Artificial Neural NetworkabstractThe electrocardiogram (ECG) has been proven as an efficient diagnostic tool to monitor the electrical activity of the heart and has become a widely used clinical approach to diagnose heart diseases. In a practical way, the ECG signal can be decomposed into P, Q, R, S, and T waves. Based on the information of the features in these waves, such as the amplitude and the interval between each wave, many types of heart diseases can be detected by using the neural network (NN)-based ECG analysis approach. However, because of a large amount of computing to preprocess the raw ECG signal, it is time consuming to analyze the ECG signal in the time domain. In addition, the non-linear ECG signal analysis worsens the difficulty to diagnose the ECG signal. To solve the problem, we propose a fast ECG diagnosis approach based on spectral analysis and the artificial neural network. Compared with the conventional time-domain approaches, the proposed approach analyzes the ECG signal only in the frequency domain. However, because most of the noises in the raw ECG signal belong to high-frequency signals, it is necessary to acquire more features in the low-frequency spectrum and fewer features in the high-frequency spectrum. Hence, a non-uniform feature extraction approach is proposed in this article. According to less data preprocessing in the frequency domain than the one in the time domain, the proposed approach not only reduces the total diagnosis latency but also reduces the computing power consumption of the ECG diagnosis. To verify the proposed approach, the well-known MIT-BIH arrhythmia database is involved in this work. The experimental results show that the proposed approach can reduce ECG diagnosis latency by 47% to 52% compared with conventional ECG analysis methods under similar diagnostic accuracy of heart diseases. In addition, because of less data preprocessing, the proposed approach can achieve lower area overhead by 22% to 29% and lower computing power consumption by 29% to 34% compared with the related works, which is proper for applying this approach to portable medical devices. Kun-Chih Chen, Po-Chen Chien, Zi-Jie Gao, Chi-Hsun Wu |
ACM Trans. Comput. Heal. | 1 |
| 2021 | A Hierarchical K-Means-Assisted Scenario-Aware Reconfigurable Convolutional Neural NetworkabstractThe superiority of convolutional neural network (CNN) has been proven in various object recognition tasks and has received much attention. However, the modern CNN approaches usually that assume the testing data and the training data belong to an identical category. Therefore, the current CNN approaches are not efficient for some applications with multisource data (i.e., the input data come from different sources), such as remote sensing scene. To increase the adaptability of the involved CNN approach, we first propose a K-means-assisted scenario-aware reconfigurable convolutional neural network (KASR-CNN) mechanism. The KASR-CNN is composed of a fully convolutional autoencoder-based K-means clustering (FCA-KC) and a reconfigurable convolutional neural network (RCNN), which are used to perform the coarsegrained classification and fine-grained classification to the input data, respectively. Furthermore, a Lego-like architecture design methodology is proposed to reduce the design complexity and improve computing flexibility. To show the adaptability, the KASR-CNN mechanism has been applied to different CNN models. In addition, the KASR-CNN has been verified on the Xilinx Zynq-ZC706 field-programmable gate array (FPGA) and implemented with TSMC 40-nm technology. Compared with the conventional approaches, the proposed KASR-CNN can help the involved CNN model to improve 3.73%-36.65% classification accuracy with only 0.38%-0.74% area overhead. Kun-Chih Chen, Ya-Wei Huang, Geng-Ming Liu, Jing-Wen Liang, Yueh-Chi Yang, Yuan-Hao Liao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Adaptive Machine Learning-Based Temperature Prediction Scheme for Thermal-Aware NoC SystemabstractBecause of the high-complex interconnection in the contemporary manycore system, the Network-on-Chip (NoC) technology is proven as an efficient way to solve the communication problem in multicore systems. However, the thermal problem becomes the main design challenge in the current NoC systems due to the high-diverse workload distribution and large power density. Therefore, the Proactive Dynamic Thermal Management (PDTM) is employed as an efficient way to control the system temperature in modern multicore systems. Based on the predicted temperature information, the PDTM can control the system temperature in advance, which helps to reduce the performance impact during the temperature control period. However, the conventional temperature prediction model is usually built based on several specific physical parameters, which are usually temperature-sensitive as well. As a result, the current temperature prediction models still suffer from large prediction errors, which reduces the benefit of the PDTM. To solve this problem, we combine the artificial neural network and LMS adaptive filter theory to propose an adaptive machine learning-based temperature prediction model. Because the proposed model can adapt to the hyperplane of the temperature behavior of NoC system during the runtime, the proposed approach can reduce average error by 37.2% to 62.3%, which helps to improve the system performance by 9.16% to 38.37% and can bring smaller area overhead than the related works by 18.59% to 22.11%. Kun-Chih Chen, Yuan-Hao Liao |
ISCAS | 1 |
| 2020 | Special issue on energy-efficient many-core embedded systems and architectures (SI: NoCArc18)
Maurizio Palesi, Kun-Chih Chen, Midia Reshadi |
J. Syst. Archit. | 2 |
| 2019 | Dual-Precision Acceleration of Convolutional Neural Network Computation with Mixed Input and Output Data ReuseabstractMemory access dominates power consumption in hardware acceleration of deep neural networks (DNN) computation due to the movement of huge data and weights. This paper design a DNN accelerator using mixed input and output data reuse scheme to achieve balance between internal memory size and memory access amount, two contradictory design goals in resource limited embedded systems. First, analytical forms for memory size and accesses are derived for different data reuse methods in DNN convolution. After comparing the analysis results across different convolutional layers of the VGG-16 model with different levels of hardware parallelism, we implement a low-cost DNN hardware accelerator using mixed input and output data reuse scheme with 32 processing elements (PEs) operating in parallel. Furthermore, the design supports two precision modes (8-bit and 16-bit) allowing variable precision requirements across DNN layers, resulting in more efficient computation compared with single-precision designs through sharing of hardware resource. Shen-Fu Hsiao, Pei-Hsuan Wu, Jien-Min Chen, Kun-Chih Chen |
ISCAS | 4 |
| 2019 | NoC-based DNN accelerator: a future design paradigmabstractDeep Neural Networks (DNN) have shown significant advantages in many domains such as pattern recognition, prediction, and control optimization. The edge computing demand in the Internet-of-Things era has motivated many kinds of computing platforms to accelerate the DNN operations. The most common platforms are CPU, GPU, ASIC, and FPGA. However, these platforms suffer from low performance (i.e., CPU and GPU), large power consumption (i.e., CPU, GPU, ASIC, and FPGA), or low computational flexibility at runtime (i.e., FPGA and ASIC). In this paper, we suggest the NoC-based DNN platform as a new accelerator design paradigm. The NoC-based designs can reduce the off-chip memory accesses through a flexible interconnect that facilitates data exchange between processing elements on the chip. We first comprehensively investigate conventional platforms and methodologies used in DNN computing. Then we study and analyze different design parameters to implement the NoC-based DNN accelerator. The presented accelerator is based on mesh topology, neuron clustering, random mapping, and XY-routing. The experimental results on LeNet, MobileNet, and VGG-16 models show the benefits of the NoC-based DNN accelerator in reducing off-chip memory accesses and improving runtime computational flexibility. Kun-Chih Chen, Masoumeh Ebrahimi, Ting-Yi Wang, Yuch-Chi Yang |
NOCS | 1 |
| 2018 | Optimization of Lookup Table Size in Table-Bound Design of Function ComputationabstractComputation of function values is critical for designing datapath units in digital signal processors (DSP) and graphics processing units (GPU). Most function computation methods requires lookup tables (LUT) and simple arithmetic components. In table-bound methods, LUT size takes a significant portion of total hardware area, in particular for high-precision applications. This paper presents a new multi-level lossless table decomposition to further reduce total table size in a recently proposed hierarchical multipartite (HMP) table method which is a generation of the prior bipartite/multipartite table-addition methods. Furthermore, hardware design parameters are optimized by jointly considering all the error sources. Experimental results shows that the proposed design has the chance of further reducing the total table size of HMP, which already has significant table-size saving over previous similar designs. Shen-Fu Hsiao, Kun-Chih Chen, Yi-Hau Chen |
ISCAS | 2 |
| 2018 | Game-Based Thermal-Delay-Aware Adaptive Routing (GTDAR) for Temperature-Aware 3D Network-on-Chip SystemsabstractThe thermal problem of three-dimensional Network-on-Chip (3D NoC) is proven to be severer than 2D NoC due to the stacking dies and heterogeneous thermal conduction between each silicon layers. To control the system temperature under a certain thermal limit, the current Dynamic Thermal Managements (DTMs) can be classified into temporal approaches and spatial approaches. The temporal DTM approaches reduce the processing speed of those overheated NoC components. However, for emerging cooling consideration, the full throttling scheme is usually applied as the system temperature reaches the alarming level, which results in significant system performance overhead. On the other hand, the spatial DTM approaches migrate the traffic load away from the overheated components. Although the spatial DTM approaches can mitigate the performance impact during the temperature control, the cooling period is longer than the temporal approaches because of the asynchronous phenomenon of traffic and temperature behavior among the NoC components. To consider the advantages of the temporal and spatial DTM approaches, it is necessary to synchronize the information of traffic and temperature behavior in the NoC systems. In this paper, we apply the Game Theory to propose a Game-based Thermal-Delay-aware Adaptive Routing (GTDAR) scheme. The GTDAR first adopts the Thermal-Delay principle to transfer the long-term temperature information to short-term traffic information by allocating the input buffer length of each NoC routers, which can reduce the thermal problem into the traffic problem. Afterward, the GTDAR involves the Nash Equilibrium property to distribute the packet routing to mitigate the thermal problem by considering the traffic and temperature simultaneously. In our experiments, the proposed Game-based Thermal-Delay-aware Adaptive Routing (GTDAR) scheme can help to improve 8.7 percent to 130 percent system performance with only 2.4 percent area overhead compared with the previous works. Kun-Chih Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Enabling fast preemption via Dual-Kernel support on GPUsabstractTo consider QoS for resource-limited mobile systems, we introduce a fast preemption mechanism on GPUs. First, we involve a dual-kernel execution model to support fine-grained preemption, and a resource allocation policy to avoid resource fragmentation problem. Second, we propose a preemption victim selection scheme to reduce the throughput overhead while satisfying a required preemption latency. Evaluations show that we can reach very close to the ideal preemption scheme within 2% difference in terms of deadline violations. Furthermore, on average we improve GPU resource utilization by 2.93× over prior technique during preemption. Li-Wei Shieh, Kun-Chih Chen, Hsueh-Chun Fu, Po-Han Wang 0001, Chia-Lin Yang |
ASP-DAC | 2 |
| 2017 | Analyzing OpenCL 2.0 workloads using a heterogeneous CPU-GPU simulatorabstractHeterogeneous CPU-GPU systems have recently emerged as an energy-efficient computing platform. A robust integrated CPU-GPU simulator is essential to facilitate researches in this direction. While few integrated CPU-GPU simulators are available, similar tools that support OpenCL 2.0, a widely used new standard with promising heterogeneous computing features, are currently missing. In this paper, we extend the existing integrated CPU-GPU simulator, gem5-gpu, to support OpenCL 2.0. In addition, we conduct experiments on the extended simulator to see the impact of new features introduced by OpenCL 2.0. Our OpenCL 2.0 compatible simulator is successfully validated against a state-of-the-art commercial product, and is expected to help boost future studies in heterogeneous CPU-GPU systems. Ren-Wei Tsai, Shao-Chung Wang, Kun-Chih Chen, Po-Han Wang 0001, Hsiang-Yun Cheng, Yi-Chung Lee, Sheng-Jie Shu, Chun-Chieh Yang, Min-Yih Hsu, Li-Chen Kan, Chao-Lin Lee, Tzu-Chieh Yu, Rih-Ding Peng, Chia-Lin Yang, Yuan-Shin Hwang, Jenq Kuen Lee, Shiao-Li Tsao, Ouhyoung Ming |
ISPASS | 4 |
| 2017 | Path-Diversity-Aware Fault-Tolerant Routing Algorithm for Network-on-Chip SystemsabstractNetwork-on-Chip (NoC) is the regular and scalable design architecture for chip multiprocessor (CMP) systems. With the increasing number of cores and the scaling of network in deep submicron (DSM) technology, the NoC systems become subject to manufacturing defects and have low production yield. Due to the fault issues, the reduction in the number of available routing paths for packet delivery may cause severe traffic congestion and even to a system crash. Therefore, the fault-tolerant routing algorithm is desired to maintain the correctness of system functionality. To overcome fault problems, conventional fault-tolerant routing algorithms employ fault information and buffer occupancy information of the local regions. However, the information only provides a limited view of traffic in the network, which still results in heavy traffic congestion. To achieve fault-resilient packet delivery and traffic balancing, this work proposes a Path-Diversity-Aware Fault-Tolerant Routing (PDA-FTR) algorithm, which simultaneously considers path diversity information and buffer information. Compared with other fault-tolerant routing algorithms, the proposed work can improve average saturation throughput by 175 percent with only 8.9 percent average area overhead and 7.1 percent average power overhead. Yu-Yin Chen, En-Jui Chang, Hsien-Kai Hsin, Kun-Chih Chen, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | RC-Based Temperature Prediction Scheme for Proactive Dynamic Thermal Management in Throttle-Based 3D NoCsabstractThe three-dimensional Network-on-Chip (3D NoC) has been proposed to solve the complex on-chip communication issues in multicore systems using die stacking in recent days. Because of the larger power density and the heterogeneous thermal conductance in different silicon layers of 3D NoC, the thermal problems of 3D NoC become more exacerbated than that of 2D NoC and become a major design constraint for a high-performance system. To control the system temperature under a certain thermal limit, many Dynamic Thermal Managements (DTMs) have been proposed. Recently, for emergent cooling, the full throttling scheme is usually employed as the system temperature reaches the alarming level. Hence, the conventional reactiveDTMsuffers from significant performance impact because of the pessimistic reaction. In this paper, we propose a throttle-based proactiveDTM(T-PDTM) scheme to predict the future temperature through a newThermal RC-based temperature prediction (RCTP) model. TheRCTPmodel can precisely predict the temperature with heterogeneous workload assignment with low constant computational complexity. Based on the predictive temperature, the proposedT-PDTMscheme will assign the suitable clock frequency for each node of the NoC system to perform early temperature control through power budget distribution. Based on the experimental results, compared with the conventional reactive throttled-basedDTMs, theT-PDTMscheme can help to reduce 11.4∼80.3 percent fully throttled nodes and improves the network throughput by around 1.5∼211.8 percent. Kun-Chih Chen, En-Jui Chang, Huai-Ting Li, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Traffic- and Thermal-aware Adaptive Beltway Routing for three dimensional Network-on-Chip systemsabstractThe distribution of traffic and temperature in a high-performance three dimensional Network-on-Chip (3D NoC) system become more unbalanced because of chip stacking and applied minimal routing algorithms. To regulate the temperature under a certain thermal limit, the overheated nodes are usually throttled by run-time thermal management (RTM). Therefore, the network topology becomes a Non-Stationary Irregular Mesh (NSI-Mesh) and leads to heavy traffic congestion around the throttled nodes. Because of the traffic imbalance in the network, the system performance degrades sharply as temperature rises. In this paper, a Traffic- and Thermal-aware Adaptive Beltway Routing (TTABR) is proposed to balance both the distribution of the traffic and temperature in the network. The proposed TTABR can be applied to NSI-Mesh and regular mesh. The experimental results show that the proposed TTABR can achieve more balanced both traffic and temperature distribution, and the network throughput is improved by around 3.4~113% with less than 18% area overhead. Kun-Chih Chen, Che-Chuan Kuo, Hui-Shun Hung, An-Yeu Wu |
ISCAS | 1 |
| 2013 | Transport-layer-assisted routing for runtime thermal management of 3D NoC systems
Chih-Hao Chao, Kun-Chih Chen, Tsu-Chu Yin, Shu-Yen Lin, An-Yeu Wu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Topology-Aware Adaptive Routing for Nonstationary Irregular Mesh in Throttled 3D NoC SystemsabstractThree-dimensional network-on-chip (3D NoC) has been proposed to solve the complex on-chip communication issues in future 3D multicore systems. However, the thermal problems of 3D NoC are more serious than 2D NoC due to chip stacking. To keep the temperature below a certain thermal limit, the thermal emergent routers are usually throttled. Then, the topology of 3D NoC becomes a Nonstationary Irregular Mesh (NSI-Mesh). To ensure the successful packet delivery in the NSI-Mesh, some routing algorithms had been proposed in the previous works. However, the network still suffers from extremely traffic imbalance among lateral and vertical logic layer. In this paper, we propose a Topology Aware Adaptive Routing (TAAR) to balance the traffic load for NSI-Mesh in 3D NoC. TAAR has three routing modes, which can be dynamically adjusted based on the topology status of the routing path. In addition to increasing routing flexibility, the TAAR also increases both vertical and lateral path diversity to balance the traffic load. Compared with the related adaptive routing methods, the experimental results show that the proposed TAAR can reduce 19 to 295 percent traffic loads in the bottom logic layer and improve around 7.7 to 380 percent network throughput. According to our proposed VLSI architecture, the TAAR only needs less than 24.8 percent hardware overhead compared with the previous works. Kun-Chih Chen, Shu-Yen Lin, Hui-Shun Hung, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Routing-Based Traffic Migration and Buffer Allocation Schemes for 3-D Network-on-Chip Systems With Thermal LimitabstractThe 3-D network-on-chip (NoC) router is a major source of thermal hotspots, limiting the performance gain of 3-D integration. Due to the varying cooling efficiency of different silicon layers in 3-D NoC, the optimal criteria of traditional load balancing design (LBD) scheme and temperature balancing design (TBD) scheme may not be satisfied. To analyze the tradeoff between performance and temperature, we provide a new analytical model. The model shows that the LBD scheme and the TBD scheme can be considered as two corner cases in the design space, and design cases can be categorized by comparing the bandwidth bound and the thermal-limited bound. To find the optimal design criteria between the LBD and the TBD schemes in 3-D NoC, we propose a new routing-based traffic migration, vertical-downward lateral-adaptive proactive routing (VDLAPR), and buffer allocation methods, vertical buffer allocation (VBA). The VDLAPR algorithm enables to tradeoff between the LBD and the TBD schemes. The proposed VBA method mitigates the traffic congestion caused by traffic migration. To reach the optimal configuration, we propose a systematic design flow, which assists in finding the best design parameters in the expanded space between LBD and TBD. Based on the traffic-thermal co-simulation experiments, the achievable throughput can be improved from 2.7% to 45.2% using the proposed design scheme. Chih-Hao Chao, Kun-Chih Chen, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |