EDBT 2026 Demo / reviewers in the wild / expert
Chia-Lin Yang
dblp:03/711
· DBLP profile ↗
95ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0003-0091-5027ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 85 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 11 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | In-Storage Read-Centric Seed Location Filtering Using 3D-NAND Flash for Genome Sequence AnalysisabstractRead mapping is a critical bottleneck in genome sequence analysis, requiring costly approximate string matching to identify potential matches between reads and a reference genome. Pre-alignment filtering methods aim to mitigate this issue by filtering out unnecessary mapping locations, and implementing them with processing-in-memory (PIM) approaches offers potential benefits by offloading filtering from the computing unit. However, the sparse number of potential mapping locations for each read limits the utilization of PIM's parallel computing capabilities, thereby hindering the overlapping of filtering and sequence alignment to hide filtering latency overheads. In this paper, we propose a 3D NAND-based in-storage pre-alignment filtering approach. Leveraging the read depth property, we introduce a read-centric pre-alignment filtering method that enables parallel comparison of multiple reads. We co-design software and hardware for in-situ processing of read-centric pre-alignment filtering within the storage, capitalizing on 3D NAND Flash's approximate parallel search capability. When integrating with a representative read mapping accelerator, our design achieves an average 1.36x performance improvement with comparable energy consumption. Compared to the state-of-the-art (SOTA) PIM solution, our design results in 123.8x and 53.3x performance gain and energy efficiency improvement. You-Kai Zheng, Ming-Liang Wei, Hsiang-Yun Cheng, Chia-Lin Yang, Ming-Hsiang Tsai, Chia-Chun Chien, Yuan-Hao Zhong, Po-Hao Tseng, Hsiang-Pang Li |
ASP-DAC | 4 |
| 2025 | Filter-Based Adaptive Model Pruning for Efficient Incremental Learning on Edge DevicesabstractIncremental Learning (IL) enhances Machine Learning (ML) models over time with new data, ideal for edge devices at the forefront of data collection. However, executing IL on edges faces challenges due to limited resources. Common methods involve IL followed by model pruning or specialized IL methods for edges. However, the former increases training time due to fine-tuning and compromises accuracy for past classes due to limited retained samples or features. Meanwhile, existing edge-specific IL methods utilize weight pruning, which requires specialized hardware or compilers to speed up and cannot reduce computations on general embedded platforms. In this paper, we propose Filter-based Adaptive Model Pruning (FAMP), the first pruning method designed specifically for IL. FAMP prunes the model before the IL process, allowing fine-tuning to occur concurrently with IL, thereby avoiding extended training time. To maintain high accuracy for both new and past data classes, FAMP adapts the compressed model based on observed data classes and retains filter settings from the previous IL iteration to mitigate forgetting. Across all tests, FAMP achieves the best average accuracy, with only a 2.78% accuracy drop over full ML models with IL. Moreover, unlike the common methods that prolong training time, FAMP takes 35% shorter training time on average than using the full ML models for IL. Jing-Jia Hung, Yi-Jung Chen, Hsiang-Yun Cheng, Hsu Kao, Chia-Lin Yang |
DATE | 5 |
| 2025 | REAP-NVM: Resilient Endurance-Aware NVM-Based PUF Against Learning-Based AttacksabstractNVM-based PUFs offer secure authentication and cryptographic applications by exploiting NVMs' MLC to generate diverse, ML-attack-resistant responses. Yet, frequent writes degrade these PUFs, lowering reliability and lifespan. This paper presents a model to assess endurance effects on NVM PUFs, guiding the creation of more robust PUFs. Our novel NVM PUF design enhances endurance by evenly distributing writes, thus mitigating cell stress, achieving a 62x improvement over current solutions while preserving security against learning-based attacks. Hassan Nassar, Ming-Liang Wei, Chia-Lin Yang, Jörg Henkel, Kuan-Hsun Chen |
DATE | 3 |
| 2025 | Accelerating Genome Alignment Pipeline with In-NAND Search Technology and Group Testing TechniquesabstractGenomic sequence analysis deciphers and interprets an organism’s DNA, offering crucial insights into personalized medicine, disease diagnosis, evolutionary biology, and agricultural biotechnology. While Next-Generation Sequencing (NGS) has revolutionized genomics by providing a fast and cost-effective method for generating genomic sequences, the computational complexity of aligning short reads back to a reference genome remains a significant bottleneck. The exact-match-based preseeding filter has emerged as an effective and general methodology to address this issue, capable of removing 70% to 80% of exact-matched genomic reads at the source and applicable to a wide range of alignment tools. However, the state-of-the-art exact-match filter architecture, GenStore, encounters performance limitations due to the need to load reference sequences from NAND flash memory to the controller page by page.In this work, we propose a novel Solid-State Drive (SSD) architecture that leverages computing-in-NAND-flash techniques to perform match detection directly within memory. By harnessing the two-dimensional input capability of 3D NAND flash memory and integrating group testing methods, our design enables comparisons across hundreds of pages in a single read cycle and supports simultaneous multi-query searches. Combined with a Bloom filter for in-NAND search, our architecture significantly reduces data movement by 48% to 96%, achieves a speedup of 1.60× to 4.99× over GenStore, and delivers 30% higher energy efficiency with only a 4.5% circuit overhead. Ming-Hsiang Tsai, Ming-Liang Wei, Chia-Chun Chien, Po-Hao Tseng, Yung-Chun Lee, Hsiang-Pang Li, Chia-Lin Yang |
ICCAD | 7 |
| 2024 | Co-Designing NVM-based Systems for Machine Learning and In-memory Search ApplicationsabstractWith the rapid development of the Internet of Things, machine learning applications on edge devices with limited resources face challenges due to large data scales and irregular memory access patterns. Non-volatile memory (NVM) technologies provide promising solutions by offering larger capacity, low leakage power, and data persistence. In this paper, we discuss the potential of NVM technology in enhancing machine learning applications by improving energy efficiency and reducing latency through in-memory computation and different NVM write modes. The insights from this analysis provide valuable guidance to device researchers and system architects working to develop highperformance systems for machine learning and accelerators in large-scale search applications using NVMs. Jörg Henkel, Lokesh Siddhu, Hassan Nassar, Lars Bauer, Jian-Jia Chen, Christian Hakert, Tristan Taylan Seidl, Kuan-Hsun Chen, Xiaobo Sharon Hu, Mengyuan Li 0001, Chia-Lin Yang, Ming-Liang Wei |
ICCAD | 11 |
| 2024 | PointCIM: A Computing-in-Memory Architecture for Accelerating Deep Point Cloud AnalyticsabstractEfficient deep point cloud (PC) analytics is crucial for numerous emerging applications such as autonomous vehicles and augmented and virtual reality. Our roofline model analysis reveals that the “memory wall” bottleneck primarily constrains the execution efficiency of deep PC analytics, providing valuable insight into optimization opportunities. In contrast to previous works, which greatly rely on approximating the original algorithm to fit hardware limitations, the approach presented in this paper is analytical; that is, our approach does not require any modification to the original algorithm, thus preserving its integrity and accuracy. In this paper, we introduce PointCIM, the first deep PC analytics accelerator that leverages computing-in-memory (CIM) optimization opportunities to address memory inefficiency. We identify that existing in-memory methods cannot fully support the distance function required by PC network inference. To address the challenge, we propose computation optimizations, including the Base+Offset mapping and early stopping for bit-serial computation, not only to enable full support for PC network inference in memory, but also to significantly improve hardware efficiency. We design the CIM architecture support for the proposed computation optimizations, including the memristor crossbar architecture, custom peripheral logic, data layout, and pipelined execution. Evaluation results show that the designed accelerator provides an average speedup of 17.1× and an energy reduction of 9.6× compared to the baseline of a typical edge SoC. We also compare PointCIM with several state-of-the-art PC accelerators, yielding up to 10.7× speedup and 4.9× energy savings. Xuanjun Chen, Han-Ping Chen, Chia-Lin Yang |
MICRO | 3 |
| 2024 | RecTS: A Temporal-Aware Memory System Optimization for Training Deep Learning Recommendation Models
Jui-Nan Yen, You-Ru Lai, Yun-Ping Lin, Chia-Lin Yang |
SYSTOR | 5 |
| 2023 | Special Session - Non-Volatile Memories: Challenges and Opportunities for Embedded System Architectures with Focus on Machine Learning ApplicationsabstractThis paper explores the challenges and opportunities of integrating non-volatile memories (NVMs) into embedded systems for machine learning. NVMs offer advantages such as increased memory density, lower power consumption, non-volatility, and compute-in-memory capabilities. The paper focuses on integrating NVMs into embedded systems, particularly in intermittent computing, where systems operate during periods of available energy. NVM technologies bring persistence closer to the CPU core, enabling efficient designs for energy-constrained scenarios. Next, computation in resistive NVMs is explored, highlighting its potential for accelerating machine learning algorithms. However, challenges related to reliability and device non-idealities need to be addressed. The paper also discusses memory-centric machine learning, leveraging NVMs to overcome the memory wall challenge. By optimizing memory layouts and utilizing probabilistic decision tree execution and neural network sparsity, NVM-based systems can improve cache behavior and reduce unnecessary computations. In conclusion, the paper emphasizes the need for further research and optimization for the widespread adoption of NVMs in embedded systems presenting relevant challenges, especially for machine learning applications. Jörg Henkel, Lokesh Siddhu, Lars Bauer, Jürgen Teich, Stefan Wildermann, Mehdi Baradaran Tahoori, Mahta Mayahinia, Jerónimo Castrillón, Asif Ali Khan, Hamid Farzaneh, João Paulo C. de Lima, Jian-Jia Chen, Christian Hakert, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng |
CASES | 15 |
| 2023 | Unified Agile Accuracy Assessment in Computing-in-Memory Neural Accelerators by Layerwise Dynamical IsometryabstractDeploying neural networks (NN) on computing-in-memory (CIM) neural accelerators incurs additional hardware factors in the test accuracy, which add substantial extra evaluation overhead. This work takes the first step to quantitatively analyze how information propagates in CIM neural accelerators as well as how additional CIM factors influence that information propagation. From our analysis, we propose a new metric named Unified-QCN that is theoretically linked to the test accuracy according to layerwise dynamical isometry (LDI), providing us with a compass to avoid direct time-consuming simulations. Our method consistently delivers high correlations with the test accuracy for various NN backbones on different datasets. Xuanjun Chen, Cynthia Kuan, Chia-Lin Yang |
DAC | 3 |
| 2023 | Tensor Movement Orchestration in Multi-GPU Training SystemsabstractAs deep neural network (DNN) models grow deeper and wider, one of the main challenges for training large-scale neural networks is overcoming limited GPU memory capacity. One common solution is to utilize the host memory as the external memory for swapping tensors in and out of GPU memory. However, the effectiveness of such tensor swapping can be impaired in data-parallel training systems due to contention on the shared PCIe channel to the host. In this paper, we propose the first large-model support framework that coordinates tensor movements among GPUs to alleviate PCIe channel contention. We design two types of coordination mechanisms. In the first mechanism, PCIe channel accesses from different GPUs are interleaved by selecting disjoint swapped-out tensors for each GPU. In the second method, swap commands are orchestrated to avoid contention. The effectiveness of these two methods depends on the model size and how often the GPUs synchronize on gradients. Experimental results show that compared to large-model support that is oblivious to channel contention, the proposed solution achieves average speedups of 38.3% to 31.8% when the memory footprint size is 1.33 to 2 times the GPU memory size. Shao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin Yang |
HPCA | 4 |
| 2023 | Reliable Brain-inspired AI Accelerators using Classical and Emerging MemoriesabstractBy taking inspiration from the operation of biological brains, emerging brain-inspired hardware has the potential to revolutionize the way computations are performed. Brain-inspired computing can be realized using both classical CMOS and emerging beyond-CMOS technologies, whereas the latter holds the promise to provide substantial energy savings akin to the employment of non-volatile memories. One way to implement highly efficient brain-inspired AI applications is through analog computing schemes, such as Integrate-and-Fire (IF) Spiking Neural Networks (SNNs), which can be implemented using both CMOS and beyond-CMOS technologies as synaptic storage. However, managing the inherent degradation of computing accuracy in analog circuits and mitigating their effects on the predictive accuracy of AI systems remains a key challenge due to the inherent nature of analog computing.In this paper, we discuss how the aforementioned challenges can be addressed. In the first part, we present our SPICE-Torch, a framework that connects low-level SPICE simulations of circuits and memories performing analog computations with high-level accuracy evaluations of NN models based on PyTorch. Furthermore, we present an example of neuromorphic optimization using classical CMOS technology. In the second part, we introduce memristors as an emerging beyond-CMOS technology that can retain their state without any outside influence and are well-suited for brain-inspired neuromorphic hardware. We demonstrate that brain-inspired hardware, realized using classical CMOS or beyond-CMOS technologies, has the potential to revolutionize the way we process information and solve complex computation problems. Nevertheless, to harness its full potential, reliability issues have to be managed carefully and HW/SW codesign is key. Our presented framework SPICE-Torch, which connects low-level SPICE simulations of circuits performing analog computations with high-level accuracy evaluations of NN models based on PyTorch is available as open-source in https://github.com/myay/SPICE-Torch. Mikail Yayla, Simon Thomann, Md. Mazharul Islam 0006, Ming-Liang Wei, Shu-Yin Ho, Ahmedullah Aziz, Chia-Lin Yang, Jian-Jia Chen, Hussam Amrouch |
VTS | 7 |
| 2023 | Impact of Non-Volatile Memory Cells on Spiking Neural Network Annealing Machine With In-Situ Synapse ProcessingabstractSolving constraint satisfaction problems (CSPs) is in high demand for various applications. SNN serves as a competitive annealing machine that can solve the CSP more efficiently than well-known Metropolis sampling and Hopfield networks. NVM-based crossbars with analog Integrate and Fire (IF) neurons can evolve the state of SNN to solve CSP more efficiently. However, analog computations inherently suffer from imprecisions in NVM cells, e.g., current variation, OFF-state leakage, and temperature-induced drift. We are the first to analyze the impacts of various memory technologies, including 2T-NOR, FeFET, WOx ReRAM, and HfOx ReRAM, on solving the Ising model, Sudoku, and Traveling-salesman-problem (TSP). The results show that both 2T-NOR Flash and FeFET with normalized standard deviation( ${\sigma}/{u}$ ) $<$ $5\%$ and ON-OFF ratio $>$ $1000$ are both ideal candidates as synapse devices at room temperature, while other devices suffer from the effects of current variation and OFF-state leakage, which would require the neuron circuits to have infeasible membrane capacitance size. However, the drift of cell current and the reduction of the ON-OFF ratio drops the success rate as the temperature increases. The success rate of solving TSP drops by 60 $\%$ and 90 $\%$ while the temperature increases from 300K to 358K for 2T-NOR and FeFET, respectively. Throughout the simulation, we show that the transistor-based memory is suggested to be a synapse device. Yet, we also find that the tolerance of temperature is inevitable under limited capacitance. Exploration of temperature-tolerated design of circuit and memory design is still in demand for future works. Ming-Liang Wei, Mikail Yayla, Shu-Yin Ho, Jian-Jia Chen, Hussam Amrouch, Chia-Lin Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | PUMP: Profiling-free Unified Memory Prefetcher for Large DNN Model SupportabstractModern DNNs are going deeper and wider to achieve higher accuracy. However, existing deep learning frameworks require the whole DNN model to fit into the GPU memory when training with GPUs, which puts an unwanted limitation on training large models. Utilizing NVIDIA Unified Memory (UM) could inherently support training DNN models beyond GPU memory capacity. However, naively adopting UM would suffer a significant performance penalty due to the delay of data transfer. In this paper, we propose PUMP, a Profiling-free Unified Memory Prefetcher. PUMP exploits GPU asynchronous execution for prefetch; that is, there exists a delay between the time that CPU launches a kernel and the time the kernel executes in GPU. PUMP extracts memory blocks accessed by the kernel when launching and swaps these blocks into GPU memory. Experimental results show PUMP achieves about 2x speedup on the average compared to the baseline that naively enables UM. Chung-Hsiang Lin, Shao-Fu Lin, Yi-Jung Chen, En-Yu Jenp, Chia-Lin Yang |
ASP-DAC | 5 |
| 2022 | This is SPATEM! A Spatial-Temporal Optimization Framework for Efficient Inference on ReRAM-based CNN AcceleratorabstractResistive memory-based computing-in-memory (CIM) has been considered as a promising solution to accelerate convolutional neural networks (CNN) inference, which stores the weights in crossbar memory arrays and performs in-situ matrix-vector multiplications (MVMs) in an analog manner. Several techniques assume that a whole crossbar can operate concurrently and discuss how to efficiently map the weights onto crossbar arrays. However, in practice, the accumulated effect of per-cell current deviation and Analog-to-Digital-Converter overhead may greatly degrade inference accuracy, which motivates the concept of Operation Unit (OU), by which an operation per cycle in a crossbar only involve limited wordlines and bitlines to preserve satisfactory inference accuracy. With OU-based operations, the mapping of weights and scheduling strategy for parallelizing CNN convolution operations should take the cost of communication overhead and resource utilization into consideration to optimize the inference acceleration. In this work, we propose the first optimization framework named SPATEM, that efficiently executes MVMs with OU-based operations on ReRAM-based CIM accelerators. It decouples the design space into tractable steps, models the expected inference latency, and derives an optimized spatial-temporal-aware scheduling strategy. By comparing with state-of-the-arts, the experimental result shows that the derived scheduling strategy of SPATEM achieves on average 29.24% inference latency reduction with 31.28% less communication overhead by exploiting more originally unused crossbar cells. Yen-Ting Tsou, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng, Jian-Jia Chen, Der-Yu Tsai |
ASP-DAC | 3 |
| 2022 | RM-SSD: In-Storage Computing for Large-Scale Recommendation InferenceabstractTo meet the strict service level agreement requirements of recommendation systems, the entire set of embeddings in recommendation systems needs to be loaded into the memory. However, as the model and dataset for production-scale recommendation systems scale up, the size of the embeddings is approaching the limit of memory capacity. Limited physical memory constrains the algorithms that can be trained and deployed, posing a severe challenge for deploying advanced recommendation systems. Recent studies offload the embedding lookups into SSDs, which targets the embedding-dominated recommendation models. This paper takes it one step further and proposes to offload the entire recommendation system into SSD with in-storage computing capability. The proposed SSD-side FPGA solution leverages a low-end FPGA to speed up both the embedding-dominated and MLP-dominated models with high resource efficiency. We evaluate the performance of the proposed solution with a prototype SSD. Results show that we can achieve 20-100× throughput improvement compared with the baseline SSD and 1.5-15× improvement compared with the state-of-art. Xuan Sun 0003, Hu Wan 0001, Qiao Li 0001, Chia-Lin Yang, Tei-Wei Kuo, Chun Jason Xue |
HPCA | 4 |
| 2022 | Efficient Bad Block Management with Cluster SimilarityabstractProcess variation in the 3D flash memory architecture raises the difficulty of bad block management. Since the error characteristics vary among different blocks, it is difficult for the existing P/E cycle-based bad block management policies to decide a suitable cycle threshold. This increases the possibility of data loss and decreases the SSD’s lifetime. In this work, we characterize the 3D flash memory and observe spatial correlation among flash blocks in the aspect of error behaviors. This phenomenon is referred to as cluster similarity. A novel cluster-based bad block management policy is proposed, which treats the failure of a block as an indicator of near-future failures of its neighboring blocks. Moreover, we provide quantitative methods to enable judicious selection of the cluster size to meet the desired tradeoff between the SSD lifetime and reliability. Compared with the commonly-used cycle-based bad block management policy, our cluster-based management policy has a lifetime improvement of 2x with comparable failure rates. And with comparable lifetime, the failure rate of the cycle-based policy is 9x higher than our method. To alleviate the I/O performance impact caused by the cluster retirement, we proposes a critical-block first reallocation scheduling. Our experiments show up to two times improvement of the 95th percentile latency compared to the naive scheduling of cluster reallocation. Jui-Nan Yen, Tseng-Yi Chen, Chia-Lin Yang, Hsiang-Yun Cheng |
HPCA | 5 |
| 2022 | DL-RSIM: A Reliability and Deployment Strategy Simulation Framework for ReRAM-based CNN AcceleratorsabstractMemristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. In addition, due to the hardware constraints, the way to deploy neural network models on memristor crossbar arrays affects the computation parallelism and communication overheads. To enable reliable and energy-efficient memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit/device properties on the inference accuracy and the influence of different deployment strategies on performance and energy consumption. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. A rich set of reliability impact factors and deployment strategies are explored by DL-RSIM, and it can be incorporated with any deep learning neural networks implemented by TensorFlow. Using several representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and energy-efficient deployment strategies and develop optimization techniques accordingly. Hsiang-Yun Cheng, Chia-Lin Yang, Meng-Yao Lin, Kai Lien, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang, Yen-Ting Tsou, Chin-Fu Nien |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | Future Computing Platform Design: A Cross-Layer Design ApproachabstractFuture computing platforms are facing a paradigm shift with the emerging resistive memory technologies. First, they offer fast memory accesses and data persistence in a single large-capacity device deployed on the memory bus, blurring the boundary between memory and storage. Second, they enable computing-in-memory for neuromorphic computing to mitigate costly data movements. Due to the non-ideality of these resistive memory devices at the moment, we envision that cross-layer design is essential to bring such a system into practice. In this paper, we showcase a few examples to demonstrate how cross-layer design can be developed to fully exploit the potential of resistive memories and accelerate its adoption for future computing platforms. Hsiang-Yun Cheng, Chun-Feng Wu, Christian Hakert, Kuan-Hsun Chen, Yuan-Hao Chang 0001, Jian-Jia Chen, Chia-Lin Yang, Tei-Wei Kuo |
DATE | 7 |
| 2021 | Binarized SNNs: Efficient and Error-Resilient Spiking Neural Networks through BinarizationabstractSpiking Neural Networks (SNNs) are considered the third generation of NNs and can reach similar accuracy as conventional deep NNs, but with a considerable improvement in efficiency. However, to achieve high accuracy, state-of-the-art SNNs employ stochastic spike coding of the inputs, requiring multiple cycles of computation. Because of this and due to the nature of analog computing, it is required to accumulate and hold the charges of multiple cycles, necessitating a large membrane capacitor. This results in high energy, long latency, and expensive area costs, constituting one of the major bottlenecks in analog SNN implementations. Membrane capacitor size determines the precision of the firing time. Hence reducing the capacitor size considerably degrades the inference accuracy. To alleviate this, we focus on bridging the gap between binarized NNs (BNNs) and SNNs. BNNs are rapidly emerging as an attractive alternative for NNs due to their high efficiency and error tolerance. In this work, we evaluate the impact of deploying error-resilient BNNs, i.e. BNNs that have been proactively trained in the presence of errors, on analog implementation of SNNs. We show that for BNNs, the capacitor size and latency can be reduced significantly compared to state-of-the-art SNNs, which employ multi-bit models. Our experiments demonstrate that when error-resilient BNNs are deployed on analog-based SNN accelerator, the size of the membrane capacitor is reduced by 50%, the inference latency is decreased by two orders of magnitude, and energy is reduced by 57% compared to the baseline 4-bit SNN implementation, under minimal accuracy cost. Ming-Liang Wei, Mikail Yayla, Shu-Yin Ho, Jian-Jia Chen, Chia-Lin Yang, Hussam Amrouch |
ICCAD | 5 |
| 2021 | A Dense Tensor Accelerator with Data Exchange Mesh for DNN and Vision WorkloadsabstractWe propose a dense tensor accelerator called VectorMesh, a scalable, memory-efficient architecture that can support a wide variety of DNN and computer vision workloads. Its building block is a tile execution unit (TEU), which includes dozens of processing elements (PEs) and SRAM buffers connected through a butterfly network. A mesh of FIFOs between the TEUs facilitates data exchange between tiles and promote local data to global visibility. Our design performs better according to the roofline model for CNN, GEMM, and spatial matching algorithms compared to state-of-the-art architectures. It can reduce global buffer and DRAM fetches by 2-22 times and up to 5 times, respectively. Wei-Chao Chen, Chia-Lin Yang, Shao-Yi Chien |
ISCAS | 3 |
| 2021 | Analyzing the Interplay Between Random Shuffling and Storage Devices for Efficient Machine LearningabstractMachine learning algorithms, such as Support Vector Machine (SVM) and Deep Neural Network (DNN), have gained a lot of interest recently. When training a machine learning algorithm, randomly shuffling all the training data can improve the testing accuracy and boost the convergence rate. Nevertheless, realizing training data random shuffling in a real system is not straightforward due to the slow random accesses in hard disk drives (HDDs). Common random shuffling implementations assume that HDD is used as storage, so they sacrifice the random degree of shuffling to reduce random storage accesses. Different from conventional HDD, emerging solid-state drive (SSD) based storage devices, such as Intel Optane SSD, offer fast random accesses. In this paper, we explore the opportunities to take advantage of the fast random access property in SSD to perform full-range random shuffling without taking up precious CPU memory and study the interplay between different shuffling methods and various types of storage devices. We use a lightweight implementation of random shuffling (LIRS) as an example of the SSD-aware shuffling method to conduct performance analysis. Evaluations show that, compared to conventional shuffling methods, LIRS can improve convergence rate and reduce the total training time of SVM and DNN by 67.1% and 33.9% on average. Zhi-Lin Ke, Hsiang-Yun Cheng, Chia-Lin Yang, Han-Wei Huang |
ISPASS | 3 |
| 2021 | ezGeno: an automatic model selection package for genomic data analysisabstractMOTIVATION: To facilitate the process of tailor-making a deep neural network for exploring the dynamics of genomic DNA, we have developed a hands-on package called ezGeno. ezGeno automates the search process of various parameters and network structures and can be applied to any kind of 1D genomic data. Combinations of multiple abovementioned 1D features are also applicable. RESULTS: For the task of predicting TF binding using genomic sequences as the input, ezGeno can consistently return the best performing set of parameters and network structure, as well as highlight the important segments within the original sequences. For the task of predicting tissue-specific enhancer activity using both sequence and DNase feature data as the input, ezGeno also regularly outperforms the hand-designed models. Furthermore, we demonstrate that ezGeno is superior in efficiency and accuracy compared to the one-layer DeepBind model and AutoKeras, an open-source AutoML package. AVAILABILITY AND IMPLEMENTATION: The ezGeno package can be freely accessed at https://github.com/ailabstw/ezGeno. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun-Liang Lin, Tsung-Ting Hsieh, Yi-An Tung, Xuanjun Chen, Yu-Chun Hsiao, Chia-Lin Yang, Tyng-Luh Liu, Chien-Yu Chen 0001 |
Bioinform. | 6 |
| 2020 | Lattice: An ADC/DAC-less ReRAM-based Processing-In-Memory Architecture for Accelerating Deep Convolution Neural NetworksabstractNonvolatile Processing-In-Memory (NVPIM) has demonstrated its great potential in accelerating Deep Convolution Neural Networks (DCNN). However, most of existing NVPIM designs require costly analog-digital conversions and often rely on excessive data copies or writes to achieve performance speedup. In this paper, we propose a new NVPIM architecture, namely, Lattice, which calculates the partial sum of the dot products between the feature map and weights of network layers in a CMOS peripheral circuit to eliminate the analog-digital conversions. Lattice also naturally offers an efficient data mapping scheme to align the data of the feature maps and the weights and hence, avoiding the excessive data copies or writes in the previous NVPIM designs. Finally, we develop a zero-flag encoding scheme to save the energy of processing zero-values in sparse DCNNs. Our experimental results show that Lattice improves the system energy efficiency by 4× ~ 13.22× compared to three state-of-the-art NVPIM designs: ISAAC, PipeLayer, and FloatPIM. Qilin Zheng, Zongwei Wang 0001, Zishun Feng, Bonan Yan, Yimao Cai, Ru Huang 0001, Yiran Chen 0001, Chia-Lin Yang, Hai Li 0001 |
DAC | 8 |
| 2019 | The Impact of Emerging Technologies on Architectures and System-level Management: Invited PaperabstractThe goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management. Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang |
ICCAD | 12 |
| 2019 | Iotbench: A Benchmark Suite for Intelligent Internet of Things Edge DevicesabstractIoT devices must and will be more intelligent in the future. However, due to the lack of benchmarks representative to the diverse IoT applications, there are limited architecture performance studies on IoT devices. This paper presents IoTBench, the first benchmark suite targeting at IoT edge-device applications. This suite includes seven representative programs from three major IoT categories. We investigate the computational demand and execution efficiency of these benchmarks running on a popular IoT device platform. We also analyze and characterize the energy consumption of IoTBench using analytic approaches. Overall, IoTBench establishes a foundation for innovative architecture design of IoT edge devices. Chien-I Lee, Meng-Yao Lin, Chia-Lin Yang, Yen-Kuang Chen |
ICIP | 3 |
| 2019 | Sparse ReRAM engine: joint exploration of activation and weight sparsity in compressed neural networksabstractExploiting model sparsity to reduce ineffectual computation is a commonly used approach to achieve energy efficiency for DNN inference accelerators. However, due to the tightly coupled crossbar structure, exploiting sparsity for ReRAM-based NN accelerator is a less explored area. Existing architectural studies on ReRAM-based NN accelerators assume that an entire crossbar array can be activated in a single cycle. However, due to inference accuracy considerations, matrix-vector computation must be conducted in a smaller granularity in practice, called Operation Unit (OU). An OU-based architecture creates a new opportunity to exploit DNN sparsity. In this paper, we propose the first practical Sparse ReRAM Engine that exploits both weight and activation sparsity. Our evaluation shows that the proposed method is effective in eliminating ineffectual computation, and delivers significant performance improvement and energy savings. Tzu-Hsien Yang, Hsiang-Yun Cheng, Chia-Lin Yang, I-Ching Tseng, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li |
ISCA | 3 |
| 2018 | Active forwarding: eliminate IOMMU address translation for accelerator-rich architecturesabstractAccelerator-rich architectures employ IOMMUs to support unified virtual address, but researches show that they fail to meet the performance and energy requirements of accelerators. Instead of optimizing the speed/energy of IOMMU address translation, this work tackles the issue from a new perspective, eliminating the need for translation with an active forwarding (AF) mechanism that forwards input data of accelerators directly from the CPU cache to the scratchpad memory of the accelerator. Results show that on average, AF can provide 8% performance improvement compared to the state-of-the-art mechanism, hostPageWalk, and reduce 22.1% accelerator power. Hsueh-Chun Fu, Po-Han Wang 0001, Chia-Lin Yang |
DAC | 3 |
| 2018 | DL-RSIM: a simulation framework to enable reliable ReRAM-based accelerators for deep learningabstractMemristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. To enable reliable memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit and device properties on the inference accuracy. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. DL-RSIM simulates the error rates of every sum-of-products computation in the memristor-based accelerator and injects the errors in the targeted TensorFlow-based neural network model. A rich set of reliability impact factors are explored by DL-RSIM, and it can be incorporated with any deep learning neural network implemented by TensorFlow. Using three representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and develop reliability optimization techniques. Meng-Yao Lin, Hsiang-Yun Cheng, Tzu-Hsien Yang, I-Ching Tseng, Chia-Lin Yang, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang |
ICCAD | 6 |
| 2017 | Enabling fast preemption via Dual-Kernel support on GPUsabstractTo consider QoS for resource-limited mobile systems, we introduce a fast preemption mechanism on GPUs. First, we involve a dual-kernel execution model to support fine-grained preemption, and a resource allocation policy to avoid resource fragmentation problem. Second, we propose a preemption victim selection scheme to reduce the throughput overhead while satisfying a required preemption latency. Evaluations show that we can reach very close to the ideal preemption scheme within 2% difference in terms of deadline violations. Furthermore, on average we improve GPU resource utilization by 2.93× over prior technique during preemption. Li-Wei Shieh, Kun-Chih Chen, Hsueh-Chun Fu, Po-Han Wang 0001, Chia-Lin Yang |
ASP-DAC | 5 |
| 2017 | Leave the Cache Hierarchy Operation as It Is: A New Persistent Memory Accelerating ApproachabstractPersistent memory places NVRAM on the memory bus, offering fast access to persistent data. Yet maintaining NVRAM data persistence raises a host of challenges. Most proposed schemes either incur much performance overhead or require substantial modifications to existing architectures. Chun-Hao Lai, Jishen Zhao, Chia-Lin Yang |
DAC | 3 |
| 2017 | Message from the general co-chairsabstractOn behalf of the Organizing Committee, it is our pleasure to welcome you to the 22nd IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED 2017) in Taipei, Taiwan. The conference is being held on the beautiful campus of National Taiwan University (NTU) from July 24–26, 2017. Taiwan plays an enormous role in our field of semiconductors, with over 15% of worldwide semiconductors sales coming from an island just over 400km long. Our conference has a unique mix of topics attacking power and energy optimization at all levels of the design space from circuit design, to high-level architectures, as well as Software optimizations. The ideas here can help run your electronics on smaller batteries, or push the density of data centers. David Garrett, Chia-Lin Yang |
ISLPED | 2 |
| 2017 | Analyzing OpenCL 2.0 workloads using a heterogeneous CPU-GPU simulatorabstractHeterogeneous CPU-GPU systems have recently emerged as an energy-efficient computing platform. A robust integrated CPU-GPU simulator is essential to facilitate researches in this direction. While few integrated CPU-GPU simulators are available, similar tools that support OpenCL 2.0, a widely used new standard with promising heterogeneous computing features, are currently missing. In this paper, we extend the existing integrated CPU-GPU simulator, gem5-gpu, to support OpenCL 2.0. In addition, we conduct experiments on the extended simulator to see the impact of new features introduced by OpenCL 2.0. Our OpenCL 2.0 compatible simulator is successfully validated against a state-of-the-art commercial product, and is expected to help boost future studies in heterogeneous CPU-GPU systems. Ren-Wei Tsai, Shao-Chung Wang, Kun-Chih Chen, Po-Han Wang 0001, Hsiang-Yun Cheng, Yi-Chung Lee, Sheng-Jie Shu, Chun-Chieh Yang, Min-Yih Hsu, Li-Chen Kan, Chao-Lin Lee, Tzu-Chieh Yu, Rih-Ding Peng, Chia-Lin Yang, Yuan-Shin Hwang, Jenq Kuen Lee, Shiao-Li Tsao, Ouhyoung Ming |
ISPASS | 15 |
| 2017 | Exploiting Write Heterogeneity of Morphable MLC/SLC SSDs in Datacenters with Service-Level ObjectivesabstractGiven the needs of data-intensive web services and cloud computing applications, storage centers play an important role in serving the demanded data access while jointly considering low cost, qualified performance, and good scalability. To manage peak workloads with performance requirements for read/ write latencies, overprovisioning more storage nodes is common but also increases total cost as well as power consumption. Recently, due to the growing capacity and dropping price, NAND-flash-based Solid-State Drives (SSDs) have become an attractive storage solution in datacenters. In this work, we exploit the write heterogeneity in Multi-Level-Cell (MLC) NAND flash memory to meet Service-Level Objectives (SLOs) of applications and to avoid storage overprovision. In MLC NAND flash memory, a memory cell can be programmed as a Single-Level Cell (SLC) or a multi-level cell at runtime, and SLC writes take shorter latency with the cost of larger consumed capacity. The proposed SLO-aware morphable SSD design seeks to meet the SLO requirement by deciding the write mode of each write request while minimizing the number of SLC writes. Experimental results show that the proposed design meets the SLO requirement for all of the tested I/O traces with less than 2.8 percent extra erase counts in average, while conventional MLC SSDs require up to 2.375 times storage overprovision to meet the SLO requirement. Geng-You Chen, Yi-Jung Chen, Chia-Wei Yeh, Pei Yin Eng, Ana Cheung, Chia-Lin Yang |
IEEE Trans. Computers | 7 |
| 2017 | A Hybrid DRAM/PCM Buffer Cache Architecture for Smartphones with QoS ConsiderationabstractFlash memory is widely used in mobile phones to store contact information, application files, and other types of data. In an operating system, the buffer cache keeps the I/O blocks in dynamic random access memory (DRAM) to reduce the slow flash accesses. However, in smartphones, we observed two issues which reduce the benefits of the buffer cache. First, a large number of synchronous writes force writing the data from the buffer cache to flash frequently. Second, the large amount of I/O accesses from background applications diminishes the buffer cache efficiency of the foreground application, which degrades the quality-of-service (QoS). In this article, we propose a buffer cache architecture with hybrid DRAM and phase change memory (PCM) memory, which improves the I/O performance and QoS for smartphones. We use a DRAM first-level buffer cache to provide high buffer cache performance and a PCM last-level buffer cache to reduce the impact of frequent synchronous writes. Based on the proposed hierarchical buffer cache architecture, we propose a sub-block management and background flush to reduce the impact of the PCM write limitation and the dirty block write-back overhead, respectively. To improve the QoS, we propose a least-recently-activated first replacement policy (LRA) to keep the data from the applications that are most likely to become the foreground one. The experimental results show that with the proposed mechanisms, our hierarchical buffer cache can improve the I/O response time by 20% compared to the conventional buffer cache. The proposed LRA can improve the foreground application performance by 1.74x compared to the conventional CLOCK policy. Ye-Jyun Lin, Chia-Lin Yang, Hsiang-Pang Li, Cheng-Yuan Michael Wang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | MCSSim: A memory channel storage simulatorabstractRecently, NVDIMM (Non-Volatile Dual In-line Memory Module) is being widely supported by leading hardware design companies, such as IBM. Nevertheless, existing efforts largely focus on NVDIMM specification and fabrication issues, and the potential performance gains brought by NVDIMM are not fully investigated. In this paper, we present a NVDIMM-based simulator called MCSSim to help study the memory channel storage techniques. MCSSim is a cycle-accurate simulator that is elaborated with the consideration of differences between the memory channel interface and the NAND flash memory features. MCSSim is also implemented with the DRAMSim2 [31] simulator thus enabling the simulation of a variety of hybrid memory systems by combining of DRAM DIMM and NVDIMM. We have done some experiments with MCSSim, and the experimental results show the effectiveness of the proposed simulator. Renhai Chen, Zili Shao, Chia-Lin Yang, Tao Li 0006 |
ASP-DAC | 3 |
| 2016 | Latency sensitivity-based cache partitioning for heterogeneous multi-core architectureabstractShared last-level cache (LLC) management is a critical design issue for heterogeneous multi-cores. In this paper, we observe two major challenges: the contribution of LLC latency to overall performance varies among applications/cores and also across time; overlooking the off-chip latency factor often leads to adverse effects on overall performance. Hence, we propose a Latency Sensitivity-based Cache Partitioning (LSP) framework, including a lightweight runtime mechanism to quantify the latency-sensitivity and a new cost function to guide the LLC partitioning. Results show that LSP improves the overall throughput by 8% on average (27% at most), compared with the state-of-the-art partitioning mechanism, TAP. Po-Han Wang 0001, Cheng-Hsuan Li, Chia-Lin Yang |
DAC | 3 |
| 2016 | Improving Read Performance of NAND Flash SSDs by Exploiting Error LocalityabstractNAND flash-based solid-state drives (SSDs), which can serve as the caches of hard disk drives, have gained popularity in large-scale, high-performance storage. A type of advanced error correction code for SSDs, low-density parity-check (LDPC), is required to mitigate a considerable number of errors in the raw data of NAND flash. However, LDPC imposes read performanceoverhead due to the complex decoding procedure of LDPC. In this paper, we propose an error-correcting cache (EC-Cache) that exploits “error locality”, a characteristic of NAND flash memory, to improve the read performance of SSDs. We use the term “error locality” to refer to the property that the majority of errors in reads to the same flash page appear at the same positions until thepage is erased. By caching detected errors, we can correct a significant portion of errors in the requested flash page prior to the LDPC decoding process. This design significantly reduces LDPC decoding overhead because the latency of LDPC is correlated with thenumber of errors in the input data. We conduct experiments, including flash characterization, LDPC simulation, and SSD simulation,to evaluate EC-Cache. The experimental results demonstrate that EC-Cache can improve the read performance of LDPC-based SSDs by up to$2.6\times$. Ren-Shuo Liu, Meng-Yen Chuang, Chia-Lin Yang, Cheng-Hsuan Li, Kin-Chu Ho, Hsiang-Pang Li |
IEEE Trans. Computers | 3 |
| 2015 | Fine-grained write scheduling for PCM performance improvement under write power budgetabstractPhase-change memory (PCM) has gained much attention recently since it offers several advantages over DRAM, such as high cell density and low leakage power. PCM has similar read power and latency as DRAM; however, its write power and latency are significantly higher than DRAM. Therefore, one challenge with PCM is how to increase write throughput under write power budget constraints. To increase write concurrency, PCM often adopts division programming, where a write occurs in a series of divisions, so that writes to different banks proceed concurrently. In this study, we observe that since the write scheduling granularity in the memory controller differs from the actual write granularity in PCM chips, i.e., requests vs. divisions, the available power budget cannot be fully utilized. We therefore propose enhancing the interface between the memory controller and PCM chips to allow the memory controller to schedule writes in the division granularity. To further increase power budget utilization, we design a variable-length division mechanism to allow the division granularity to be adjusted at runtime according to the available write power budget. Our experimental results show that these techniques improve system performance by up to 65%. Chun-Hao Lai, Shun-Chih Yu, Chia-Lin Yang, Hsiang-Pang Li |
ISLPED | 3 |
| 2015 | Improving DRAM latency with dynamic asymmetric subarrayabstractThe evolution of DRAM technology has been driven by capacity and bandwidth during the last decade. In contrast, DRAM access latency stays relatively constant and is trending to increase. Much efforts have been devoted to tolerate memory access latency but these techniques have reached the point of diminishing returns. Having shorter bitline and wordline length in a DRAM device will reduce the access latency. However by doing so it will impact the array efficiency. In the mainstream market, manufacturers are not willing to trade capacity for latency. Prior works had proposed hybrid-bitline DRAM design to overcome this problem. However, those methods are either intrusive to the circuit and layout of the DRAM design, or there is no direct way to migrate data between the fast and slow levels. Shih-Lien Lu, Ying-Chen Lin, Chia-Lin Yang |
MICRO | 3 |
| 2015 | SECRET: A Selective Error Correction Framework for Refresh Energy Reduction in DRAMsabstractDRAMs are used as the main memory in most computing systems today. Studies show that DRAMs contribute to a significant part of overall system power consumption. One of the main challenges in low-power DRAM design is the inevitable refresh process. Due to process variation, memory cells exhibit retention time variations. Current DRAMs use a single refresh period determined by the cell with the largest leakage. Since prolonging refresh intervals introduces retention errors, a set of previous works adopt conventional error-correcting code (ECC) to correct retention errors. However, these approaches introduce significant area and energy overheads. In this article, we propose a novel error correction framework for retention errors in DRAMs, called SECRET (selective error correction for refresh energy reduction). The key observations we make are that retention errors are hard errors rather than soft errors, and only few DRAM cells have large leakage. Therefore, instead of equipping error correction capability for all memory cells as existing ECC schemes, we only allocate error correction information to leaky cells under a refresh interval. Our SECRET framework contains two parts: an offline phase to identify memory cells with retention errors given a target error rate and a low-overhead error correction mechanism. The experimental results show that among all test cases performed, the proposed SECRET framework can reduce refresh power by 87.2% and overall DRAM power up to 18.57% with negligible area and performance overheads. Chung-Hsiang Lin, De-Yu Shen, Yi-Jung Chen, Chia-Lin Yang, Cheng-Yuan Michael Wang |
ACM Trans. Archit. Code Optim. | 4 |
| 2015 | System-Level Performance and Power Optimization for MPSoC: A Memory Access-Aware ApproachabstractAs the number of IPs in a multimedia Multi-Processor System-on-Chip (MPSoC) continues to increase, concurrent memory accesses from different IPs increasingly stress memory systems, which presents both opportunities and challenges for future MPSoC design. The impact of such requirements on the system-level design for MPSoC is twofold. First, contention among IPs prolongs memory access time, which exacerbates the persisting memory wall problem. Second, longer memory accesses lead to longer IP stall time, which results in unnecessary leakage waste. In this article, we propose two memory access-aware system-level design approaches for performance and leakage optimization. To alleviate the memory wall problem, we propose a Hierarchical Memory Scheduling (HMS) policy that schedules memory requests from the same IP and application consecutively to reduce interference among memory accesses from different IPs with a fairness guarantee. To reduce IP leakage waste due to long memory access, we propose a memory access-aware power-gating policy. A straightforward power-gating approach is to power gate an IP when it needs to fetch data from memory. However, due to the response time variation among memory accesses, aggressively power gating an IP whenever a memory request occurs may result in incorrect power-gating decisions. The proposed memory access-aware power-gating policy makes these decisions judiciously, based on the predicted memory latency of an individual IP and its energy breakeven time. The experimental results show that the proposed HMS memory scheduling policy improves system throughput by 42% compared to First-Come-First-Serve (FCFS) and by 21% compared to First-Ready First-Come-First-Serve (FR-FCFS) on an MPSoC for mobile phones. For the improvement of fairness, HMS improves fairness by 1.52× compared to FCFS and by 1.23× compared to FRFCFS. In the aspect of leakage optimization, our memory access-aware power-gating mechanism improves energy savings by 3.88× and reduces the performance penalty by 70% compared to conventional timeout-based power gating. We further demonstrate that our HMS memory scheduler can regulate memory access orders, thereby reducing memory response time variation. This leads to more accurate power-down decisions for both conventional timeout power gating and the proposed memory access- aware power gating. Ye-Jyun Lin, Chia-Lin Yang, Jiao-Wei Huang, Tay-Jyi Lin, Chih-Wen Hsueh, Naehyuck Chang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2014 | NVM duet: unified working memory and persistent store architectureabstractEmerging non-volatile memory (NVM) technologies have gained a lot of attention recently. The byte-addressability and high density of NVM enable computer architects to build large-scale main memory systems. NVM has also been shown to be a promising alternative to conventional persistent store. With NVM, programmers can persistently retain in-memory data structures without writing them to disk. Therefore, one can envision that in the future, NVM will play the role of both working memory and persistent store at the same time. Ren-Shuo Liu, De-Yu Shen, Chia-Lin Yang, Shun-Chih Yu, Cheng-Yuan Michael Wang |
ASPLOS | 3 |
| 2014 | EC-Cache: Exploiting Error Locality to Optimize LDPC in NAND Flash-Based SSDsabstractLow-density parity-check (LDPC) is widely accepted as the baseline error-correction codes offering strong error-correcting capability for future NAND flash-based SSDs. However, LDPC incurs read performance overhead because of its complex decoding procedure. To mitigate such overhead, we propose the error-correcting cache (EC-Cache) that exploits the "error locality" of NAND flash. Error locality means that the majority of errors in reads to the same NAND flash page appear in the same positions until the page is erased. By caching detected errors, EC-Cache can correct a significant portion of errors present in a requested flash page before the associated LDPC decoding process begins. EC-Cache can greatly speed up LDPC decoding because LDPC's latency is directly correlated to the number of errors present in the input data. Experimental results show that EC-Cache achieves up to 2.6× SSD read performance gain. Ren-Shuo Liu, Meng-Yen Chuang, Chia-Lin Yang, Cheng-Hsuan Li, Kin-Chu Ho, Hsiang-Pang Li |
DAC | 3 |
| 2013 | DuraCache: a durable SSD cache using MLC NAND flashabstractAdopting SSDs as caches for HDD arrays has gained popularity in datacenters because SSDs are superior in handling random reads that HDDs cannot efficiently deal with. Two types of flash memory cells are available for building SSD caches, single-level cells (SLC) and multi-level cells (MLC). MLC is more appealing than SLC because it can achieve higher cache capacity at the same cost. However, we see a critical issue for SSD caches to adopt MLC NAND flash: the endurance of modern MLC NAND flash is too low to sustain datacenter workloads. In this paper, we propose DuraCache that addresses the durability issue of SSD caches. DuraCache exploits the fact that SSD caches are write-through caches in datacenters. Therefore, uncorrectable errors in SSD caches can be handled like cache misses which bring in correct data from HDD arrays. In addition, DuraCache gradually allocates more ECC parities associated with data when NAND flash reaches wearout thresholds. This allows SSD caches to continue operating by sacrificing available capacity. We conduct empirical experiments and demonstrate that DuraCache enables MLC SSD caches to achieve 4.1 years of service life assuming a TPC-C workload. Ren-Shuo Liu, Chia-Lin Yang, Cheng-Hsuan Li, Geng-You Chen |
DAC | 2 |
| 2013 | Exploring synergistic DVFS control of cores and DRAMs for thermal efficiency in CMPs with 3D-stacked DRAMsabstractThe three-dimensional (3D) integration technology that utilizes low-latency and high-density Through-Silicon Vias (TSVs) to integrate DRAMs and Chip-Multiprocessors(CMPs) in the third dimension has been demonstrated as a promising way to mitigate the memory wall problem in CMPs. In addition to the improved interconnection performance and heterogeneous integration, the 3D IC technology also provides the advantages of high packaging density and small chip area. However, the power density of 3D ICs increases with the number of active devices. Therefore, alleviating the thermal stress issue of 3D ICs is one of the major design challenges. Ping-Sheng Lin, Yi-Jung Chen, Chia-Lin Yang, Yi-Chang Lu |
ISLPED | 3 |
| 2012 | Memory access aware power gating for MPSoCsabstractAs technology continues to scale, reducing leakage is critical to achieve energy efficiency. Power gating can potentially save a significant part of leakage but it incurs both energy and performance penalties. Therefore, power gating decisions need to be made carefully. In the current low-power SoC design, an IP core is power gated when it is not operating. In this paper, we explore the IP idle time due to memory accesses for further leakage reduction. In MPSoCs, due to contention among concurrent memory accesses from different IP cores, memory stall cycles vary significantly, ranging from 10 to 600 cycles according to our experiments. We propose a run-time mechanism that predict the memory stall cycles of an individual IP, and make the power gating decision based on the predicted memory latency and its break-even time. With the predicted memory latency, a power-gated IP can be woken up in advance to avoid performance degradation. The experimental results show that our power management mechanism can achieve 25.3% leakage energy saving within 4% performance penalty. Ye-Jyun Lin, Chia-Lin Yang, Jiao-Wei Huang, Naehyuck Chang |
ASP-DAC | 2 |
| 2012 | Age-based PCM wear leveling with nearly zero search costabstractImproving the endurance of PCM is a fundamental issue when the technology is considered as an alternative to main memory usage. In the design of memory-based wear leveling approaches, a major challenge is how to efficiently determine the appropriate memory pages for allocation or swapping. In this paper, we present an efficient wear-leveling design that is compatible with existing virtual memory management. Two implementations, namely, bucket-based and array-based wear leveling, with nearly zero search cost are proposed to tradeoff time and space complexity. The results of experiments conducted based on popular benchmarks to evaluate the efficacy of the proposed design are very encouraging. Chi-Hao Chen, Pi-Cheng Hsiu, Tei-Wei Kuo, Chia-Lin Yang, Cheng-Yuan Michael Wang |
DAC | 4 |
| 2012 | Optimizing NAND flash-based SSDs via retention relaxation
Ren-Shuo Liu, Chia-Lin Yang |
FAST | 2 |
| 2012 | Distributed memory interface synthesis for Network-on-Chips with 3D-stacked DRAMsabstractStacking DRAMs on processing cores by Through-Silicon Vias (TSVs) provides abundant bandwidth and enables a distributed memory interface design. To achieve the best balance in performance and cost in an application-specific system, the distributed memory interface should be tailored for the target applications. In this paper, we propose the first distributed memory interface synthesis framework for application-specific Network-on-Chips (NoCs) with 3D-stacked DRAMs. To maximize the performance of a selected hardware configuration, the proposed framework co-synthesizes the hardware configuration of the distributed memory interface, and the software configuration, e.g. task mapping and data assignment. Since TSVs have adverse impact on chip costs and yields, the goal of the framework is minimizing the number of TSVs provided that the user-defined performance constraint is met. Yi-Jung Chen, Chia-Lin Yang, Jian-Jia Chen |
ICCAD | 2 |
| 2012 | SECRET: Selective error correction for refresh energy reduction in DRAMsabstractDRAMs are used as the main memory in most computing systems today. Studies show that DRAMs contribute to a significant part of overall system power consumption. Therefore, one of the main challenges in low-power DRAM design is the inevitable refresh process. Due to process variation, memory cells exhibit retention time variations. Current DRAMs use a single worst-case refresh period. Prolonging refresh intervals introduces retention errors. Previous works adopt conventional ECC (Error Correcting Code) to correct retention errors. These approaches introduce significant area and energy overheads. In this paper, we propose a novel error correction framework for retention errors in DRAMs, called SECRET (Selective Error Correction for Refresh Energy reducTion). The key observation we make is that retention errors can be treated as hard errors rather than soft errors, and only few DRAM cells have large leakage. Therefore, instead of equipping error correction capability in all memory cells as existing ECC schemes, we only allocate error correction information to leaky cells under a refresh interval. Our SECRET framework contains two parts, an off-line phase to identify memory cells with retention errors given a target error rate, and a low-overhead error correction mechanism. The experimental results show that the proposed SECRET framework can reduce refresh power by 87.2%, and overall DRAM power by 18.57% with negligible area and performance overheads. Chung-Hsiang Lin, De-Yu Shen, Yi-Jung Chen, Chia-Lin Yang, Cheng-Yuan Michael Wang |
ICCD | 4 |
| 2012 | A cycle-level SIMT-GPU simulation frameworkabstractThe massive parallelism provided by the modern graphics processing units (GPUs) makes them the attractive processors to accelerate the applications with high data-level parallelism. Therefore, the GPU architecture has recently gained a lot of attention in research community. However, the advance in the GPU architecture is impeded by the limited documents released from the major GPU vendors. Furthermore, current studies on GPUs often focus only on general-purpose (GPGPU) applications. The behaviors of the graphics applications, which are considered as the major GPU workloads, are often overlooked in these studies. A GPU design good for the GPGPU applications is not necessarily good for the graphics applications. Therefore, a simulation framework that is able to provide performance characterization for both applications is mandatory for the innovation of the GPU architecture. Po-Han Wang 0001, Chien-Wei Lo, Chia-Lin Yang, Yu-Jung Cheng |
ISPASS | 3 |
| 2011 | Power gating strategies on GPUsabstractAs technology continues to shrink, reducing leakage is critical to achieving energy efficiency. Previous studies on low-power GPUs (Graphics Processing Units) focused on techniques for dynamic power reduction, such as DVFS (Dynamic Voltage and Frequency Scaling) and clock gating. In this paper, we explore the potential of adopting architecture-level power gating techniques for leakage reduction on GPUs. We propose three strategies for applying power gating on different modules in GPUs. The Predictive Shader Shutdown technique exploits workload variation across frames to eliminate leakage in shader clusters. Deferred Geometry Pipeline seeks to minimize leakage in fixed-function geometry units by utilizing an imbalance between geometry and fragment computation across batches. Finally, the simple time-out power gating method is applied to nonshader execution units to exploit a finer granularity of the idle time. Our results indicate that Predictive Shader Shutdown eliminates up to 60% of the leakage in shader clusters, Deferred Geometry Pipeline removes up to 57% of the leakage in the fixed-function geometry units, and the simple time-out power gating mechanism eliminates 83.3% of the leakage in nonshader execution units on average. All three schemes incur negligible performance degradation, less than 1%. Po-Han Wang 0001, Chia-Lin Yang, Yen-Ming Chen, Yu-Jung Cheng |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | TACLC: Timing-Aware Cache Leakage Control for Hard Real-Time SystemsabstractLeakage energy consumption is an increasingly important issue as the technology continues to shrink. Existing leakage reduction techniques for hard real-time systems utilize slack to turn off a CPU completely. However, turning on/off a processor involves high performance and energy overheads. Hence, a hard real-time system is very likely to have unutilized slack if only the CPU shutdown technique is used to reduce leakage. Architectural-level shutdown techniques in all instances have a much lower overheads than turning off a CPU; therefore, they can be utilized in a hard real-time system to further reduce CPU leakage. However, existing architecture-level shutdown techniques cause unpredictable performance degradation thereby unsuitable for a hard real-time system that must meet the timing constraint in all cases. This paper is the first attempt to bridge this gap. This paper focuses on cache leakage reduction and proposes the first Timing-Aware Cache Leakage Control (TACLC) mechanism. TACLC exploits system slack to turn cache lines into low-leakage states provided that the timing constraint is met. The experimental results demonstrate that TACLC effectively utilizes system slack to reduce cache leakage. For systems with low CPU utilization, TACLC achieves comparable leakage reduction to the leakage control policy that aggressively turns cache lines into low-leakage modes while neglecting the timing constraint. Yi-Jung Chen, Chia-Lin Yang, Jaw-Wei Chi, Jian-Jia Chen |
IEEE Trans. Computers | 2 |
| 2011 | Thermal Modeling and Analysis for 3-D ICs With Integrated Microchannel CoolingabstractIntegrated microchannel liquid-cooling technology is envisioned as a viable solution to alleviate an increasing thermal stress imposed by 3-D stacked ICs. Thermal modeling for microchannel cooling is challenging due to its complicated thermal-wake effect, a localized temperature wake phenomenon downstream of a heated source in the flow. This paper presents a fast and accurate thermal-wake aware thermal model for integrated microchannel 3-D ICs. A combination of the microchannel thermal-wake function and the channel merging technique achieves more than 3300× speedup with less than 5% error in comparison with a commercial numerical finite volume simulation tool. With the proposed model, we characterize thermal behaviors of microchannel-cooled 3-D ICs and compare them with the case of conventional air-cooled 3-D ICs. We also demonstrate thermal-aware placements using our thermal model. It shows that the proposed model can be used to reduce peak temperatures, which is considered important for 3-D IC designs. Hitoshi Mizunuma, Yi-Chang Lu, Chia-Lin Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | PM-COSYN: PE and memory co-synthesis for MPSoCsabstractMulti-Processor System-on-Chips (MPSoCs) exploit task-level parallelism to achieve high computation throughput, but concurrent memory accesses from multiple PEs may cause memory bottleneck. Therefore, to maximize system performance, it is important to simultaneously consider the PE and on-chip memory architecture design. However, in a traditional MPSoC design flow, PE allocation and on-chip memory allocation are often considered independently. To tackle this problem, we propose the first PE and Memory Co-synthesis (PM-COSYN) framework for MPSoCs. One critical issue in such a memory-aware MPSoC design is how to utilize the available die area to achieve a balanced design between memory and computation subsystems. Therefore, the goal of PM-COSYN is to allocate PE and on-chip memory for MPSoCs with Network-on-Chip (NoC) architecture such that system performance is maximized and the area constraint is met. The experimental results show that, PM-COSYN can synthesize NoC resource allocation according to the needs of the target task set. When comparing to a Simulated-Annealing method, PM-COSYN generates a comparable solution with much shorter CPU time. Yi-Jung Chen, Chia-Lin Yang, Po-Han Wang 0001 |
DATE | 2 |
| 2010 | Hierarchical memory scheduling for multimedia MPSoCsabstractOptimizing memory system performance is critical for delivering high system performance for multimedia applications since they are usually memory intensive. As the number of IP cores in a multimedia MPSoC (Multi-Processor System-on-Chip) continues to increase, system performance will be eventually limited by the memory system. In this paper, we tackle the memory performance issue of multimedida MPSoCs through intelligent memory access scheduling. We observe that since memory resources are shared by all processing elements in an MPSoC, interferences among requests from different IP cores cause not only delay in memory accesses but also unfair DRAM accesses among IPs. Traditional memory scheduling policies that only emphasize on maximizing memory system throughput do not take into account these interferences. Therefore, in this paper, we propose a hierarchical memory scheduling policy to minimize interferences among requests. The experimental results show that the proposed scheduling policy improves system throughput by 21% compared to FR-FCFS (first-ready first-come-first-serve) on an MPSoC for mobile phones with QoS guarantee. Ye-Jyun Lin, Chia-Lin Yang, Tay-Jyi Lin, Jiao-Wei Huang, Naehyuck Chang |
ICCAD | 2 |
| 2010 | Dynamic thermal management for networked embedded systems under harsh ambient temperature variationabstractModern vehicle electronics control units (ECUs) are getting rapidly complicated because of active safety and semi-autonomous driving controls, such as electric stability program (ESP) and adaptive cruise control (ACC). Furthermore, the operational environment of ECUs is extremely harsh, especially in terms of an ambient temperature well exceeding 100°C, which causes a very small temperature headroom. Thus, ECUs require a careful temperature management and high performance at the same time. Sangyoung Park, Jian-Jia Chen, Donghwa Shin, Younghyun Kim 0001, Chia-Lin Yang, Naehyuck Chang |
ISLPED | 5 |
| 2010 | Memory Latency Reduction via Thread ThrottlingabstractMemory Wall is a well-known obstacle to processor performance improvement. The popularity of multi-core architecture will further exaggerate the problem since the memory resource is shared by all cores. Interferences among requests from different cores may prolong the latency of memory accesses thereby degrading the system performance. To tackle the problem, this paper proposes to decouple application threads into compute and memory tasks, and restrict the number of concurrent memory tasks to avoid the interference among memory requests. Yet with this scheduling restriction, a CPU core may unnecessarily stay idle, which incurs adverse impact on the overall performance. Therefore, we develop a memory thread throttling mechanism that tunes the allowable memory threads dynamically under workload variation to improve system performance. The proposed run-time mechanism monitors memory and computation ratios of a program for phase detection. It then decides the memory thread constraint for the next program phase based on an analytical model that can estimate system performance under different constraint values. To prove the concept, we prototype the mechanism in some real-world applications as well as synthetic workloads. We evaluate their performance on real machines. The experimental results demonstrate up to 20% speedup with a pool of synthetic workloads on an Intel i7 (Nehalem) machine and match with the speedup estimated by the proposed analytical model. Furthermore, the intelligent run-time scheduling leads to a geometric mean of 12% performance improvement for real-world applications on the same hardware. Hsiang-Yun Cheng, Chung-Hsiang Lin, Chia-Lin Yang |
MICRO | 4 |
| 2009 | Thermal modeling for 3D-ICs with integrated microchannel coolingabstractIntegrated microchannel liquid-cooling technology is envisioned as a viable solution to alleviate an increasing thermal stress imposed by 3D stacked ICs. Thermal modeling for microchannel cooling is challenging due to its complicated thermal-wake effect, a localized temperature wake phenomenon downstream of a heated source in the flow. This paper presents a fast and accurate thermal-wake aware thermal model for integrated microchannel 3D ICs. Validation results show the proposed thermal model achieves more than 400x speed up and only 2.0% error in comparison with a commercial numerical simulation tool. We also demonstrate the use of the proposed thermal model for thermal optimization during the IC placement stage. We find that due to the thermal-wake effect, tiles are placed in the descending order of power magnitude along the flow direction. We also find that modeling thermal-wakes is critical for generating a thermal-aware placement for integrated microchannel-cooled 3D IC. It could result in up to 25°C peak temperature difference according to our experiments. Hitoshi Mizunuma, Chia-Lin Yang, Yi-Chang Lu |
ICCAD | 2 |
| 2009 | PPT: joint performance/power/thermal management of DRAM memory for multi-core systemsabstractWith the popularity of multi-core architecture, to sustain the memory demands from different cores, the memory system is expected to grow significantly in both speed and capacity. This will lead to increasing power consumption in the memory system. Therefore, it is critical to address the power issue in the memory subsystem. In designing a power-aware memory system, due to the interplay among power, thermal and performance, all the three factors need to be taken into account. In this paper, we propose the first joint performance, power and thermal management framework (PPT) through orchestrating task execution and page allocation. The PPT framework adapts to system loading to maximize power saving and avoid memory hotspot at the same time whiling sustaining the system bandwidth demand. Chung-Hsiang Lin, Chia-Lin Yang, Ku-Jei King |
ISLPED | 2 |
| 2009 | An architectural co-synthesis algorithm for energy-aware Network-on-Chip design
Yi-Jung Chen, Chia-Lin Yang, Yen-Sheng Chang |
J. Syst. Archit. | 2 |
| 2009 | A Progressive-ILP-Based Routing Algorithm for the Synthesis of Cross-Referencing BiochipsabstractDue to recent advances in microfluidics technology, digital microfluidic biochips and their associated computer-aided-design problems have gained much attention, most of which has been devoted to direct-addressing biochips. In this paper, we solve the droplet routing problem under the more scalable cross-referencing biochip paradigm. We propose the first droplet routing algorithm that directly solves the problem of routing. We first present an optimal basic integer-linear-programming (ILP) formulation. Due to its complexity, we also propose a progressive-ILP scheme to determine the locations of droplets at each time step. Simulation results demonstrate the efficiency and effectiveness of our algorithm. Ping-Hung Yuh, Sachin S. Sapatnekar, Chia-Lin Yang, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | T-trees: A tree-based representation for temporal and three-dimensional floorplanningabstractImproving logic capacity by time-sharing, dynamically reconfigurable FPGAs are employed to handle designs of high complexity and functionality. In this article, we model each task as a 3D-box and deal with the temporal floorplanning/placement problem for dynamically reconfigurable FPGA architectures. We present a tree-based data structure, called T-trees , to represent the spatial and temporal relations among tasks. Each node in a T-tree has at most three children which represent the dimensional relationship among tasks. For the T-tree, we develop an efficient packing method and derive the condition to ensure the satisfaction of precedence constraints which model the temporal ordering among tasks induced by the execution of dynamically reconfigurable FPGAs. Experimental results show that our tree-based formulation can obtain significantly better solution quality with less execution time than the most recent state-of-the-art work. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | Leakage-aware task scheduling for partially dynamically reconfigurable FPGAsabstractAs technology continues to shrink, reducing leakage power of Field-Programmable Gate Arrays (FPGAs) becomes a critical issue for the practical use of FPGAs. In this article, we address the leakage issue of partially dynamically reconfigurable FPGA architectures with sleep transistors embedded into FPGA fabrics. In particular, we focus on eliminating leakage waste due to the delay between reconfiguration and execution time of a task. For partially dynamically reconfigurable FPGAs, the configuration prefetching technique is commonly used to hide runtime reconfiguration overhead. With prefetching, the configuration of a task is loaded into FPGAs as early as possible. Therefore, there is often a delay between reconfiguration and execution time of a task. In this period of time, the SRAM cells allocated to a task cannot be turned off even though they are not utilized. In this article, we propose a two-stage task scheduling methodology to reduce leakage waste due to the delay between reconfiguration and execution time of a task without sacrificing performance. In the first stage, a performance-driven task scheduler that targets at minimizing the schedule length is invoked to generate an initial placement. In the second stage, a postplacement leakage-aware task scheduling is applied to refine the initial placement such that leakage waste is minimized provided that the schedule length is not increased. To solve the postplacement leakage optimization problem, we propose two algorithms. The first one is an optimal algorithm based on Integer Linear Programming (ILP). The second algorithm is a heuristic approach that iteratively refines the placement to reduce leakage waste. Experimental results on real and synthetic designs show that the efficiency and effectiveness of the proposed postplacement leakage reduction techniques. Ping-Hung Yuh, Chia-Lin Yang, Chi-Feng Li, Chung-Hsiang Lin |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2008 | A progressive-ILP based routing algorithm for cross-referencing biochipsabstractDue to recent advances in microfluidics technology, digital microfluidic biochips and their associated CAD problems have gained much attention, most of which has been devoted to direct-addressing biochips. In this paper, we solve the droplet routing problem under the more scalable cross-referencing biochip paradigm, which uses row/column addressing scheme to activate electrodes. We propose the first droplet routing algorithm that directly solves the problem of routing in cross-referencing biochips. The main challenge of this type of biochips is the electrode interference which prevents simultaneous movement of multiple droplets. We first present a basic integer linear programming (ILP) formulation to optimally solve the droplet routing problem. Due to its complexity, we also propose a progressive ILP scheme to determine the locations of droplets at each time step. Experimental results demonstrate the efficiency and effectiveness of our progressive ILP scheme on a set of practical bioassays. Ping-Hung Yuh, Sachin S. Sapatnekar, Chia-Lin Yang, Yao-Wen Chang |
DAC | 3 |
| 2008 | Obstacle-Avoiding Rectilinear Steiner Tree Construction Based on Spanning GraphsabstractGiven a set of pins and a set of obstacles on a plane, an obstacle-avoiding rectilinear Steiner minimal tree (OARSMT) connects these pins, possibly through some additional points (called the Steiner points), and avoids running through any obstacle to construct a tree with a minimal total wirelength. The OARSMT problem becomes more important than ever for modern nanometer IC designs which need to consider numerous routing obstacles incurred from power networks, prerouted nets, IP blocks, feature patterns for manufacturability improvement, antenna jumpers for reliability enhancement, etc. Consequently, the OARSMT problem has received dramatically increasing attention recently. Nevertheless, considering obstacles significantly increases the problem complexity, and thus, most previous works suffer from either poor quality or expensive running time. Based on the obstacle-avoiding spanning graph, this paper presents an efficient algorithm with some theoretical optimality guarantees for the OARSMT construction. Unlike previous heuristics, our algorithm guarantees to find an optimal OARSMT for any two-pin net and many higher pin nets. Extensive experiments show that our algorithm results in significantly shorter wirelengths than all state-of-the-art works. Chung-Wei Lin, Szu-Yu Chen, Chi-Feng Li, Yao-Wen Chang, Chia-Lin Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2008 | BioRoute: A Network-Flow-Based Routing Algorithm for the Synthesis of Digital Microfluidic BiochipsabstractDue to recent advances in microfluidics, digital microfluidic biochips are expected to revolutionize laboratory procedures. One critical problem for biochip synthesis is the droplet routing problem. Unlike traditional very large scale integration routing problems, in addition to routing path selection, the biochip routing problem needs to address the issue of scheduling droplets under practical constraints imposed by the fluidic property and timing restriction of synthesis results. In this paper, we present the first network-flow-based routing algorithm that can concurrently route a set of noninterfering nets for the droplet routing problem on biochips. We adopt a two-stage technique of global routing followed by detailed routing. In global routing, we first identify a set of noninterfering nets and then adopt the network-flow approach to generate optimal global-routing paths for nets. In detailed routing, we present thefirstpolynomial-time algorithm for simultaneous routing and scheduling using the global-routing paths with a negotiation-based routing scheme. Our algorithm targets at both the minimization of cells used for routing for better fault tolerance and minimization of droplet transportation time for better reliability and faster bioassay execution. Experimental results show the robustness and efficiency of our algorithm. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Energy-Aware Flash Memory Management in Virtual Memory SystemabstractThe traditional virtual memory system is designed for decades assuming a magnetic disk as the secondary storage. Recently, flash memory becomes a popular storage alternative for many portable devices with the continuing improvements on its capacity, reliability and much lower power consumption than mechanical hard drives. The characteristics of flash memory are quite different from a magnetic disk. Therefore, in this paper, we revisit virtual memory system design considering limitations imposed by flash memory. In particular, we focus on the energy efficient aspect since power is the first-order design consideration for embedded systems. Due to the write-once feature of flash memory, frequent writes incur frequent garbage collection thereby introducing significant energy overhead. Therefore, in this paper, we propose three methods to reduce writes to flash memory. The HotCache scheme adds an SRAM cache to buffer frequent writes. The subpaging technique partitions a page into subunits, and only dirty subpages are written to flash memory. The duplication-aware garbage collection method exploits data redundancy between the main memory and flash memory to reduce writes incurred by garbage collection. We also identify one type of data locality that is inherent in accesses to flash memory in the virtual memory system, intrapage locality. Intrapage locality needs to be carefully maintained for data allocation in flash memory. Destroying intrapage locality causes noticeable increases in energy consumption. Experimental results show that the average energy reduction of combined subpaging, HotCache, and duplication-aware garbage collection techniques is 42.2%. Chia-Lin Yang, Hung-Wei Tseng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Cache leakage control mechanism for hard real-time systemsabstractLeakage energy consumption is an increasingly important issue as the technology continues to shrink. Since on-chip caches constitute a major portion of the processor’s transistor budget, several leakage control policies have been proposed to reduce cache leakage. However, these policies introduce performance unpredictability thereby not suitable for hard real-time applications that require the timing constraint is met in all cases. In this paper, we propose the first approach to apply existing low leakage circuit techniques on hard real-time applications. The proposed timing-aware cache leakage control mechanism exploits task slack time to turn cache lines into the low-leakage state provided that the timing constraint is met. The experimental results show that the proposed cache leakage control policy achieves comparable leakage reduction to the leakage control policy that aggressively turns cache lines into low-leakage modes without considering the timing constraint. Jaw-Wei Chi, Chia-Lin Yang, Yi-Jung Chen, Jian-Jia Chen |
CASES | 2 |
| 2007 | Energy-efficient real-time task scheduling with task rejection
Jian-Jia Chen, Tei-Wei Kuo, Chia-Lin Yang, Ku-Jei King |
DATE | 3 |
| 2007 | BioRoute: a network-flow based routing algorithm for digital microfluidic biochipsabstractDue to the recent advances in microfluidics, digital microfluidic biochips are expected to revolutionize laboratory procedures. One critical problem for biochip synthesis is the droplet routing problem. Unlike traditional VLSI routing problems, in addition to routing path selection, the biochip routing problem needs to address the issue of scheduling droplets under the practical constraints imposed by the fluidic property and the timing restriction of the synthesis result. In this paper, we present the first network-flow based routing algorithm that can concurrently route a set of non-interfering nets for the droplet routing problem on biochips. We adopt a two-stage technique of global routing followed by detailed routing. In global routing, we first identify a set of non-interfering nets and then adopt the network-flow approach to generate optimal global-routing paths for the nets. In detailed routing, we present the first polynomialtime algorithm for simultaneous routing and scheduling using the global-routing paths with a negotiation-based routing scheme. The experimental results show the robustness and efficiency of our algorithm. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
ICCAD | 2 |
| 2007 | 3D Video Applications and Intelligent Video Surveillance Camera and its VLSI DesignabstractIn this demonstration, the core processing engines of two video applications, 3D video and intelligent video surveillance, are demonstrated. The developed algorithms and its VLSI design results are shown with hardware prototypes processing input video on-the-fly. In addition to the processing engine design, the development tools for efficiently designing these chips are also demonstrated. Shao-Yi Chien, Chi-Sheng Shih 0001, Mong-Kai Ku, Chia-Lin Yang, Yao-Wen Chang, Tei-Wei Kuo, Liang-Gee Chen |
ICME | 4 |
| 2007 | Post-placement leakage optimization for partially dynamically reconfigurable FPGAsabstractAs technology continues to shrink, leakage power becomes animportant issue for modern FPGAs. In this paper, we address the leakage issue of partially dynamical reconfigurable FPGAs. We focus on eliminating leakage waste due to the delay between reconfiguration and task execution. We propose a post-placement leakage-aware scheduling algorithm that refines a placement generated by a performance-driven scheduler such that leakage waste is minimized and performance is not sacrificed. Experimental results on real and synthetic designs demonstrate the effectiveness and efficiency of our algorithm on leakage optimization. Chi-Feng Li, Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
ISLPED | 3 |
| 2007 | Efficient obstacle-avoiding rectilinear steiner tree constructionabstractGiven a set of pins and a set of obstacles on a plane, an obstacle-avoiding rectilinear Steiner minimal tree (OARSMT) connects these pins, possibly through some additional points (called Steiner points), and avoids running through any obstacle to construct a tree with a minimal total wirelength. The OARSMT problem becomes more important than ever for modern nanometer IC designs which need to consider numerous routing obstacles incurred from power networks, prerouted nets, IP blocks, feature patterns for manufacturability improvement, antenna jumpers for reliability enhancement, etc. Consequently, the OARSMT problem has received dramatically increasing attention recently. Nevertheless, considering obstacles significantly increases the problem complexity, and thus most previous works suffer from either poor quality or expensive running time. Based on the obstacle-avoiding spanning graph (OASG), this paper presents an efficient algorithm with some theoretical optimality guarantees for the OARSMT construction. Unlike previous heuristics, our algorithm guarantees to find an optimal OARSMT for any 2-pin net and many higher-pin nets. Extensive experiments show that our algorithm results in significantly shorter wirelengths than all state-of-the-art works. Chung-Wei Lin, Szu-Yu Chen, Chi-Feng Li, Yao-Wen Chang, Chia-Lin Yang |
ISPD | 5 |
| 2007 | Placement of defect-tolerant digital microfluidic biochips using the T-tree formulationabstractDroplet-based microfluidic biochips have recently gained much attention and are expected to revolutionize the biological laboratory procedures. As biochips are adopted for the complex procedures in molecular biology, its complexity is expected to increase due to the need of multiple and concurrent assays on a chip. In this article, we formulate the placement problem of digital microfluidic biochips with a tree-based topological representation, called T-tree . To the best knowledge of the authors, this is the first work that adopts a topological representation to solve the placement problem of digital microfluidic biochips. We also consider the defect tolerant issue to avoid to use defective cells due to fabrication. Experimental results demonstrate that our approach is more efficient and effective than the previous unified synthesis and placement framework. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2007 | Temporal floorplanning using the three-dimensional transitive closure subGraphabstractImproving logic capacity by time-sharing, dynamically reconfigurable Field Gate Programmable Arrays (FPGAs) are employed to handle designs of high complexity and functionality. In this paper, we use a novel graph-based topological floorplan representation, named 3D-subTCG (3-Dimensional Transitive Closure subGraph), to deal with the 3-dimensional (temporal) floorplanning/placement problem, arising from dynamically reconfigurable FPGAs. The 3D-subTCG uses three transitive closure graphs to model the temporal and spatial relations between modules. We derive the feasibility conditions for the precedence constraints induced by the execution of the dynamically reconfigurable FPGAs. Because the geometric relationship is transparent to the 3D-subTCG and its induced operations (i.e., we can directly detect the relationship between any two tasks from the representation), we can easily detect any violation of the temporal precedence constraints on 3D-subTCG. We also derive important properties of the 3D-subTCG to reduce the solution space and shorten the running time for 3D (temporal) foorplanning/placement. Experimental results show that our 3D-subTCG-based algorithm is very effective and efficient. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Placement of digital microfluidic biochips using the t-tree formulationabstractDroplet-based microfluidic biochips have recently gained much attention and are expected to revolutionize the biological laboratory procedure. As biochips are adopted for the complex procedures in molecular biology, its complexity is expected to increase due to the need of multiple and concurrent assays on a chip. In this paper, we formulate the placement problem of digital microfluidic biochips with a tree-based topological representation, called T-tree. To the best knowledge of the authors, this is the first work that adopts a topological representation to solve the placement problem of digital microfluidic biochips. Experimental results demonstrate that our approach is much more efficient and effective, compared with the previous unified synthesis and placement framework. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
DAC | 2 |
| 2006 | Hierarchical value cache encoding for off-chip data busabstractOff-chip data bus consumes a significant part of system power. Recent works use small caches (Value Cache) at each side of the off-chip data bus, and transmit cache indexes instead of data values to reduce bus switching activity. A larger VC has a higher VC hit rate, but it also incurs more switching activity on a VC hit. In this paper, we propose the hierarchical VC design concept that provides a good tradeoff between VC capacity and bus switching activity. Our experimental results show that the proposed hierarchical VC design reduces the off-chip data bus energy by 60.2%. Chung-Hsiang Lin, Chia-Lin Yang, Ku-Jei King |
ISLPED | 2 |
| 2006 | An energy-efficient virtual memory system with flash memory as the secondary storageabstractThe traditional virtualmemory system is designed for decades assuming a magnetic disk as the secondary storage. Recently, flash memory becomes a popular storage alternative formany portable devices with the continuing improvements on its capacity, reliability and much lower power consumption than mechanical hard drives. TheNAND flash memory is organized with blocks, and each block contains a set of pages. The characteristics of flash memory are quite different from a magnetic disk. Therefore, in this paper, we revisit virtual memory system desigin considering limitations imposed by flash memory. In particular, we study the effects of the subpaging technique and storage cache management. In the traditional virtual memory system, a full page is written back to the secondary storage on a page fault. We found that this could result in unnecessary writes thereby wasting energy. The subpaging technique that partitions a page into subunits, and only dirty subpages are written to flash memory is beneficial to the energy efficiency. For the storage cache management, unlike traditional disk cache maniagement, care needs to be taken to guarantee that the flash pages of a main memory page are replaced from the cache in sequence. Experimental results show that the average energy reduction of combined subpaging and caching techniques is 35.6%. Hung-Wei Tseng 0001, Chia-Lin Yang |
ISLPED | 3 |
| 2006 | A Space-Efficient Caching Mechanism for Flash-Memory Address TranslationabstractWhile flash memory has been widely adopted for various embedded systems, space efficiency with reasonable performance has become a critical issue for the design of the flash-memory translation layer. The target of this paper is to improve the performance of existing designs by proposing a search-tree-like caching mechanism for efficient address translation. A replacement strategy with a low time complexity is presented to monitor the access status of recently used LBA's. The proposed caching mechanism and replacement strategy were shown being highly effective in the reducing of the address translation time over popular translation layer designs, such as NAND, where realistic workloads were used for experiments. Chin-Hsien Wu, Tei-Wei Kuo, Chia-Lin Yang |
ISORC | 3 |
| 2005 | Joint exploration of architectural and physical design spaces with thermal considerationabstractHeat is a main concern for processors in deep sub-micron technologies. The chip temperature is affected by both the power consumption of processor components and the chip layout. Therefore, for thermal-aware design it is crucial to consider the thermal effects of different floorplans during micro-architectural design space exploration. In this paper, we propose a thermal-aware architectural floorplanning framework. With the aid of this framework, an architect can explore both physical and architectural design spaces simultaneously to find an architecture and the corresponding chip layout that maximizes performance under a thermal limitation Yen-Wei Wu, Chia-Lin Yang, Ping-Hung Yuh, Yao-Wen Chang |
ISLPED | 2 |
| 2005 | Reconfigurable Platform for Content Science ResearchabstractThe College of Electrical Engineering and Computer Science at the National Taiwan University has identified the area of content science for media-rich life, broadly construed, as one of core areas for the college's future directions. One major aspect of this project is to develop the enabling technology for reconfigurable platforms for multimedia applications. Specifically, the goal is to develop reconfigurable platforms of system-on-a-chip (SoC) components for the applications to support the needs of multi-modal multimedia contents and to provide rapid system prototyping. The faculties in the College of Electrical Engineering and Computer Science have formed a multi-discipline team to develop such technology. Our team includes seven faculties and more than thirty students from the college. This short report describes the reconfigurable platform for content science research activities currently underway by our team. Our current activities include to develop the technology to analyze the critical path for avoiding hardware contention, to minimize the use of logic components, to design the multimedia IPs, to optimally route the bus and place the logic units, to design energy efficient cache, to evaluate the performance and power consumption, to design the algorithm for temporal floor-planning/placement. Chi-Sheng Shih 0001, Chia-Lin Yang, Mong-Kai Ku, Tei-Wei Kuo, Shao-Yi Chien, Yao-Wen Chang, Liang-Gee Chen |
RTCSA | 2 |
| 2005 | Software-Controlled Cache Architecture for Energy EfficiencyabstractPower consumption is an important design issue of current multimedia embedded systems. Data caches consume a significant portion of total processor power for multimedia applications because they are data intensive. In an integrated multimedia system, the cache architecture cannot be tuned specifically for an application. Therefore, a significant amount of cache energy is actually wasted. In this paper, we propose the software-controlled cache architecture that improves the energy efficiency of the shared cache in an integrated multimedia system on an application-specific base. Data types in an application are allocated to different cache regions. On each access, only the allocated cache regions need to be activated. We test the effectiveness of the software-controlled cache of the MPEG-2 software decoder. The results show up to 40% of cache energy reduction on an ARM-like cache architecture without sacrificing performance. Chia-Lin Yang, Hong-Wei Tseng, Chia-Chiang Ho, Ja-Ling Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Temporal floorplanning using 3D-subTCG
Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang, Hsin-Lung Chen |
ASP-DAC | 2 |
| 2004 | Value-Conscious Cache: Simple Technique for Reducing Cache Access PowerabstractMost microprocessors employ the on-chip caches to bridge the performance gap between the processor and main memory. However, the cache accesses usually contribute significantly to the total power consumption of the chip. Based on the observation that an overwhelming majority of the cache access bits are '0', in this paper we propose a value-conscious (VC) cache to reduce the average cache power consumption during an access. Unlike the conventional cache with differential-bitline implementation, the VC cache is a single-bitline design. Depending on the access bit value, the VC cache can dynamically prevent the bitline from being discharged such that the power dissipated in accessing '0' is much less than the power dissipated in accessing '1'. The implementation of the VC cache is a circuit-level technique, which is software independent and orthogonal to other low power techniques at architecture-level. The experimental results based on the SPEC2000 and MediaBench traces show that without compromise of both performance and stability, by exploiting the prevalence of '0' bits in access data the VC cache can reduce the average cache read and write power by about 18%/spl sim/22% and 36%/spl sim/40%, respectively. Yen-Jen Chang, Chia-Lin Yang, Feipei Lai |
DATE | 2 |
| 2004 | Multiprocessor Energy-Efficient Scheduling with Task Migration Considerations
Jian-Jia Chen, Heng-Ruey Hsu, Kai-Hsiang Chuang, Chia-Lin Yang, Ai-Chun Pang, Tei-Wei Kuo |
ECRTS | 4 |
| 2004 | Temporal floorplanning using the T-tree formulationabstractImproving logic capacity by time-sharing, dynamically reconfigurable FPGAs are employed to handle designs of high complexity and functionality. We model each task as a 3D-box and deal with the temporal floorplanning/placement problem for dynamically reconfigurable FPGA architectures. We present a tree-based data structure, called T-trees, to represent the spatial and temporal relations among tasks. Each node in a T-tree has at most three children which represent the dimensional relationship among tasks. For the T-tree, we develop an efficient packing method and derive the condition to ensure the satisfaction of precedence constraints which model the temporal ordering among tasks induced by the execution of dynamically reconfigurable FPGAs. Experimental results show that our tree-based formulation can achieve significantly better solution quality with less execution time than the most recent state-of-the-art work. Ping-Hung Yuh, Chia-Lin Yang, Yao-Wen Chang |
ICCAD | 2 |
| 2004 | HotSpot cache: joint temporal and spatial locality exploitation for i-cache energy reductionabstractPower consumption is an important design issue of current embedded systems. It has been shown that the instruction cache accounts for a significant portion of the power dissipation of the whole chip. Several studies propose to add a cache (L0 cache) that is very small relative to the conventional L1 cache on chip for power optimization since a smaller cache has lower load capacitance. However, energy savings often come at the cost of performance degradation. In this paper, we propose a novel instruction cache architecture, the HotSpot cache, that achieves energy savings without sacrificing performance. The HotSpot cache identifies frequently accessed instructions dynamically and stores them in the L0 cache. Other instructions are placed only in the L1 cache. A steering mechanism is employed to direct an instruction to its allocated cache in the instruction fetch stage. The simulation results show that the HotSpot cache can achieve 52% instruction cache energy reduction on the average for a set of multimedia applications without performance degradation. Chia-Lin Yang, Chien-Hao Lee |
ISLPED | 1 |
| 2004 | Tolerating memory latency through push prefetching for pointer-intensive applicationsabstractPrefetching is often used to overlap memory latency with computation for array-based applications. However, prefetching for pointer-intensive applications remains a challenge because of the irregular memory access pattern and pointer-chasing problem. In this paper, we proposed a cooperative hardware/software prefetching framework, the push architecture, which is designed specifically for linked data structures. The push architecture exploits program structure for future address generation instead of relying on past address history. It identifies the load instructions that traverse a LDS and uses a prefetch engine to execute them ahead of the CPU execution. This allows the prefetch engine to successfully generate future addresses. To overcome the serial nature of LDS address generation, the push architecture employs a novel data movement model. It attaches the prefetch engine to each level of the memory hierarchy and pushes , rather than pulls , data to the CPU. This push model decouples the pointer dereference from the transfer of the current node up to the processor. Thus a series of pointer dereferences becomes a pipelined process rather than a serial process. Simulation results show that the push architecture can reduce up to 100% of memory stall time on a suite of pointer-intensive applications, reducing overall execution time by an average 15%. Chia-Lin Yang, Alvin R. Lebeck, Hung-Wei Tseng 0001, Chien-Hao Lee |
ACM Trans. Archit. Code Optim. | 1 |
| 2004 | Zero-aware asymmetric SRAM cell for reducing cache power in writing zeroabstractMost microprocessors employ the on-chip caches to bridge the performance gap between the processor and the main memory. However, the cache accesses usually contribute significantly to the total power consumption of the chip. Based on the observation that an overwhelming majority of the values written to the cache are "0", in this paper we propose a zero-aware SRAM cell with an asymmetric inverter pair, called ZA cell, to minimize the cache power consumption in writing "0". The ZA cell uses a circuit-level technique, which is software independent and orthogonal to other low-power techniques at architecture-level. Compared to the conventional SRAM cell, the experimental results based on the SPEC2000 and MediaBench traces show that without compromise of both performance and stability, the ZA cell can reduce the average cache write power consumption over 60% for both the baseline instruction and data caches. In particular, the ZA cell is attractive in the data caches, which reveal the high write-zero rate. Yen-Jen Chang, Feipei Lai, Chia-Lin Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | A power-aware SWDR cell for reducing cache write powerabstractLow power caches have become a critical component of both hand-held devices and high-performance processors. Based on the observation that an overwhelming majority of the data written to the cache are '0', in this paper we propose a power-aware SRAM cell with one single-bitline write port and one differential-bitlines read port, called SWDR cell, to minimize the cache power consumption in writing '0'. The SWDR cell uses a circuit-level technique, which is software independent and orthogonal to other low power techniques at architecture-level. Compared to the conventional SRAM cell, the experimental results show that without compromise of both performance and stability, the SWDR cell can result in 73%∼92% reduction in average cache write power dissipated in bitlines. Yen-Jen Chang, Chia-Lin Yang, Feipei Lai |
ISLPED | 2 |
| 2000 | Push vs. pull: data movement for linked data structuresabstractAs the performance gap between the CPU and main memory continues to grow, techniques to hide memory latency are essential to deliver a high performance computer system. Prefetching can often overlap memory latency with computation for array-based numeric applications. However, prefetching for pointer-intensive applications still remains a challenging problem. Prefetching linked data structures (LDS) is difficult because the address sequence of LDS traversal does not present the same arithmetic regularity as array-based applications and the data dependence of pointer dereferences can serialize the address generation process. Chia-Lin Yang, Alvin R. Lebeck |
ICS | 1 |
| 2000 | Exploiting Parallelism in Geometry Processing with General Purpose Processors and Floating-Point SIMD InstructionsabstractThree-dimensional (3D) graphics applications have become very important workloads running on today's computer systems. A cost-effective graphics solution is to perform geometry processing of 3D graphics on the host CPU and have specialized hardware handle the rendering task. In this paper, we analyze microarchitecture and SIMD instruction set enhancements to a RISC superscalar processor for exploiting parallelism in geometry processing for 3D computer graphics. Our results show that 3D geometry processing has inherent parallelism. Adding SIMD operations improves performance from 8 percent to 28 percent on a 4-issue dynamically scheduled processor that can issue at most two floating-point operations. In comparison, an 8-issue processor, ignoring cycle time effects, can achieve 20 to 60 percent performance improvement over a 4-issue. If processor cycle time scales with the number of ports to the register file, then doubling only the floating-point issue width of a 4-issue processor with SIMD instructions gives the best performance among the architectural configurations that we examine (the most aggressive configuration is an 8-issue processor with SIMD instructions). Chia-Lin Yang, Barton Sano, Alvin R. Lebeck |
IEEE Trans. Computers | 1 |
| 1999 | Annotated Memory References: A Mechanism for Informed Cache Management
Alvin R. Lebeck, David R. Raymond, Chia-Lin Yang, Mithuna Thottethodi |
Euro-Par | 3 |
| 1998 | Exploiting Instruction Level Parallelism in Geometry Processing for Three Dimensional Graphics ApplicationsabstractThree dimensional (3D) graphics applications have become very important workloads running on today's computer systems. A cost-effective graphics solution is to perform geometry processing of 3D graphics on the host CPU and have specialized hardware handle the rendering task. In this paper, we analyze microarchitecture and SIMD instruction set enhancements to a RISC superscalar processor for exploiting instruction level parallelism (ILP) in geometry processing for 3D computer graphics. Our results show that 3D geometry processing has inherent parallelism. When ignoring cycle time effects, an 8-issue processor can achieve up to 60% performance improvement over a 4-issue. However, certain application attributes can hinder the exploitation of ILP on a superscalar processor. Adding SIMD operations improves performance from 8% to 28% on a 4-issue processor that can issue at most 2 floating-point operations. If processor cycle time scales with the number of ports to the register file, doubling only the floating-point issue width of a 4-issue processor with SIMD instructions gives the best performance among the architecture configurations that we examine (the most aggressive configuration is an 8-issue processor with SIMD instructions). Chia-Lin Yang, Barton Sano, Alvin R. Lebeck |
MICRO | 1 |