EDBT 2026 Demo / reviewers in the wild / expert
Bo-Cheng Lai
dblp:l/BoChengLai · also Bo-Cheng Charles Lai
· DBLP profile ↗
47ranked-venue papers
14as first author
15since 2021 · last 2026
0000-0002-9729-5196ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 14 first-author · 12 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Switchable-Precision Universally Slimmable Networks
Chi-Jui Chen, Edwin Arkel Rios, Bo-Cheng Lai |
IEEE Signal Process. Lett. | 3 |
| 2026 | hls4ml: A Flexible, Open Source Platform for Deep Learning Acceleration on Reconfigurable HardwareabstractWe present hls4ml , a free and open source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this article, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results. Jan-Frederik Schulte, Benjamin Ramhorst, Jovan Mitrevski, Nicolò Ghielmetti, Enrico Lupi, Dimitrios Danopoulos, Vladimir Loncar, Javier M. Duarte, David Burnette, Lauri Laatu, Stylianos Tzelepis, Konstantinos Axiotis, Quentin Berthet, Haoyan Wang, Suleyman Demirsoy, Marco Colombo, Thea Aarrestad, Sioni Summers, Maurizio Pierini, Giuseppe Di Guglielmo, Jennifer Ngadiuba, Javier Campos, Benjamin Hawks, Abhijith Gandrakota, Farah Fahim, George A. Constantinides, Zhiqiang Que, Wayne Luk, Alexander D. Tapper, Duc Hoang, Noah Paladino, Philip C. Harris, Bo-Cheng Lai, Manuel Valentin, Ryan Forelli, Seda Ogrenci Memik, Lino Gerlach, Rian Brooks Flynn, Mia Liu, Daniel Diaz 0003, Elham E Khoda, Melissa Quinnan, Russell Solares, Santosh Parajuli, Mark S. Neubauer, Christian Herwig, Ho Fung Tsoi, Dylan S. Rankin, Shih-Chieh Hsu, Scott Hauck |
ACM Trans. Reconfigurable Technol. Syst. | 36 |
| 2025 | BAQET: BRAM-aware Quantization for Efficient Transformer Inference via Stream-based Architecture on an FPGAabstractFPGAs are a compelling substrate for supporting machine learning inference. Tools such as High-Level Synthesis and hls4ml can shorten the development cycle for deploying ML algorithms on FPGAs, but can struggle to handle the large on-chip storage needed for many of these models. In particular the high BRAM usage found in many of these flows can cause Place & Route failures during synthesis. In this paper we propose using a Simulated-Annealing based flow to perform BRAM-aware quantization. This approach trades off inference accuracy with BRAM usage, to provide a high-quality inference engine that still meets on-chip resource constraints. We demonstrate this flow for Transformer-based machine learning algorithms, which include Flash Attention in a Stream-based Dataflow architecture. Our system imposes minimal accuracy drops, yet can reduce BRAM usage by 20%-50%, and improve power efficiency by 264%-812% compared to existing Transformer-based accelerators on FPGAs LingChi Yang, Chi-Jui Chen, Trung Le 0002, Bo-Cheng Lai, Scott Hauck, Shih-Chieh Hsu |
FPGA | 4 |
| 2025 | HiGTR: High-Performance FPGA Implementation of Complete GNN-based Trajectory Reconstruction for HEPabstractCharged particle trajectory reconstruction is a critical task in high-energy physics (HEP), particularly for collision analysis in the Large Hadron Collider (LHC). In the LHC, the Level-1 Trigger (L1T) system must perform trajectory reconstruction with ultralow latency and very high throughput. Graph Neural Networks (GNNs)-based trajectory reconstruction on FPGAs has shown promising performance. However, the existing FPGA-based implementations are incomplete, supporting only the GNN processing stage and lacking the graph construction and track building parts of the task flow. Yun-Chen Yang, Hsuan-Wei Yu, Bo-Cheng Lai, Shih-Chieh Hsu, Mark S. Neubauer, Santosh Parajuli |
FPGA | 3 |
| 2025 | Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image RecognitionabstractUltra-fine-grained image recognition (UFGIR) is a challenging task that involves classifying images within a macro-category. While traditional FGIR deals with classifying different species, UFGIR goes beyond by classifying sub-categories within a species such as cultivars of a plant. In recent times the usage of Vision Transformer-based backbones has allowed methods to obtain outstanding recognition performances in this task but this comes at a significant cost in terms of computation specially since this task significantly benefits from incorporating higher resolution images. Therefore, techniques such as token reduction have emerged to reduce the computational cost. However, dropping tokens leads to loss of essential information for fine-grained categories, specially as the token keep rate is reduced. Therefore, to counteract the loss of information brought by the usage of token reduction we propose a novel Cross-Layer Aggregation Classification Head and a Cross-Layer Cache mechanism to recover and access information from previous layers in later locations. Extensive experiments covering more than 2000 runs across diverse settings including 5 datasets, 9 backbones, 7 token reduction methods, 5 keep rates, and 2 image sizes demonstrate the effectiveness of the proposed plug-and-play modules and allow us to push the boundaries of accuracy vs cost for UFGIR by reducing the kept tokens to extremely low ratios of up to 10% while maintaining a competitive accuracy to state-of-the-art models. Code is available at: https://github.com/arkel23/CLCA Edwin Arkel Rios, Jansen Christopher Yuanda, Vincent Leon Ghanz, Cheng-Wei Yu, Bo-Cheng Lai, Min-Chun Hu 0001 |
ICASSP | 5 |
| 2025 | Global-Local Similarity for Efficient Fine-Grained Image Recognition with Vision TransformersabstractFine-Grained recognition involves the classification of images from subordinate macro-categories, and it is challenging due to small inter-class differences. To overcome this, most methods perform discriminative feature selection enabled by a feature extraction backbone followed by a high-level feature refinement step. Recently, many studies have shown the potential behind vision transformers as a backbone for fine-grained recognition, but their usage of its attention mechanism to select discriminative tokens can be computationally expensive. In this work, we propose a novel and computationally inexpensive metric to identify discriminative regions in an image. We compare the similarity between the global representation of an image given by the CLS token, a learnable token used by transformers for classification, and the local representation of individual patches. We select the regions with the highest similarity to obtain crops, which are forwarded through the same transformer encoder. Finally, high-level features of the original and cropped representations are further refined together in order to make more robust predictions. We demonstrate the effectiveness of our proposed method through comprehensive experiments, obtaining superior accuracy across multiple datasets. Code and checkpoints are available at: https://github.com/arkel23/GLSim. Edwin Arkel Rios, Min-Chun Hu 0001, Bo-Cheng Lai |
ISCAS | 3 |
| 2025 | Multi-dimensional Range Joins on HBM-enabled FPGAsabstractRange join is an essential analysis in various big data applications such as genomic variant annotations, spatiotemporal databases and EDA design rule checking, often requiring searches on billion-scale records. However, its performance on multi-threaded CPUs/GPUs have been bottlenecked by both memory-access bandwidth and instruction/data dependencies. Furthermore, the limited intra-task parallelism in state-of-the-art renders SIMD platforms severely under-utilized. In this work, we present an efficient hardware-software co-design for multi-dimensional range joins on clusters of HBM-enabled FPGAs. Our highly-scalable processing system achieves up-to 2.6x/6.9x/7.1x speedup/energy improvements/memory access reductions compared to state-of-the-art CPU solution for one-dimensional range joins, while being easily extensible to larger data dimensionality. Shih-Chen Lo, Bo-Cheng Lai |
ISCAS | 3 |
| 2023 | Low Latency Edge Classification GNN for Particle Trajectory Tracking on FPGAsabstractIn-time particle trajectory reconstruction in the Large Hadron Collider is challenging due to the high collision rate and numerous particle hits. Using GNN (Graph Neural Network) on FPGA has enabled superior accuracy with flexible trajectory classification. However, existing GNN architectures have inefficient resource usage and insufficient parallelism for edge classification. This paper introduces a resource-efficient GNN architecture on FPGAs for low latency particle tracking. The modular architecture facilitates design scalability to support large graphs. Leveraging the geometric properties of hit detectors further reduces graph complexity and resource usage. Our results on Xilinx UltraScale+ VU9P demonstrate 1625x and 1574x performance improvement over CPU and GPU respectively. Shi-Yu Huang, Yun-Chen Yang, Yu-Ru Su, Bo-Cheng Lai, Javier M. Duarte, Scott Hauck, Shih-Chieh Hsu, Jin-Xuan Hu, Mark S. Neubauer |
FPL | 4 |
| 2023 | REGAL: Reprogrammable Engines for Genome Analysis on LPDDR4x-based Stacked DRAMabstractGenomic big data analysis pipelines are bottlenecked with massive data movement between CPU and memory hierarchy. Recently developed Stacked Embedded DRAM (SEDRAM) with high density hybrid bonding offers not only high bandwidth concurrent data accesses, but also highly parallel distributed data processing on the logic layer. However, the complex logical flow and data dependencies pose challenges to existing Near-DRAM Processing (NDP) solutions for analysis such as string pattern matching using FM-Index. In this work, we propose REGAL, a highly scalable and re-programmable solution on SEDRAM memory system. REGAL architecture maximizes intra-query parallelism and enhances occupancy of all components. The efficient data layout and mapping minimizes round-trip communications between processing engines. The programmability of REGAL further enables effective prefetching to support variations in data characteristics and involved algorithms. REGAL demonstrates up-to 17.3x and 70.6x speedup and energy reductions compared to multithreaded CPU implementation for FM-Index queries, while achieving performance comparable to state-of-the-art fixed-function NDP implementation. Yuhao Fang, Bo-Cheng Lai |
ISCAS | 3 |
| 2023 | A Bin-Based Indexing for Scalable Range Join on Genomic DataabstractRange-join is an operation for finding overlaps in interval-form genomic data. Range-join is widely used in various genome analysis processes such as annotation, filtering and comparison of variants in whole-genome and exome analysis pipelines. The quadratic complexity of current algorithms with sheer data volume has surged the design challenges. Existing tools have limitations on algorithm efficiency, parallelism, scalability and memory consumption. This paper proposes BIndex, a novel bin-based indexing algorithm and its distributed implementation to attain high throughput range-join processing. BIndex features near-constant search complexity while the inherently parallel data structure facilitates exploitation of parallel computing architectures. Balanced partitioning of dataset further enables scalability on distributed frameworks. The implementation on Message Passing Interface shows upto 933.5x speedup in comparison to state-of-the-art tools. Parallel nature of BIndex further enables GPU-based acceleration with 3.72x speedup than CPU implementations. The add-in modules for Apache Spark provides upto 4.65x speedup than the previously best available tool. BIndex supports wide variety of input and output formats prevalent in bioinformatics community and the algorithm is easily extendable to streaming data in recent Big Data solutions. Furthermore, the index data structure is memory-efficient and consumes upto two orders-of-magnitude lesser RAM, while having no adverse effect on speedup. Bo-Cheng Lai, Jhih-Yong Mai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | DASC: A DRAM Data Mapping Methodology for Sparse Convolutional Neural NetworksabstractThe data transferring of sheer model size of CNN (Convolution Neural Network) has become one of the main performance challenges in modern intelligent systems. Although pruning can trim down substantial amount of non-effective neurons, the excessive DRAM accesses of the non-zero data in a sparse network still dominate the overall system performance. Proper data mapping can enable efficient DRAM accesses for a CNN. However, previous DRAM mapping methods focus on dense CNN and become less effective when handling the compressed format and irregular accesses of sparse CNN. The extensive design space search for mapping parameters also results in a time-consuming process. This paper proposes DASC: a DRAM data mapping methodology for sparse CNNs. DASC is designed to handle the data access patterns and block schedule of sparse CNN to attain good spatial locality and efficient DRAM accesses. The bank-group feature in modern DDR is further exploited to enhance processing parallelism. DASC also introduces an analytical model to facilitate fast exploration and quick convergence of parameter search in minutes instead of days from previous work. When compared with the state-of-the-art, DASC decreases the total DRAM latencies and attains an average of 17.1x, 14.3x, and 23.3x better DRAM performance for sparse AlexNet, VGG-16, and ResNet-50 respectively. Bo-Cheng Lai, Tzu-Chieh Chiang, Po-Shen Kuo, Wan-Ching Wang, Yan-Lin Hung, Hung-Ming Chen, Chien-Nan Jimmy Liu, Shyh-Jye Jou |
DATE | 1 |
| 2022 | A Highly Parallel Fine-Grained Sort-Merge Join on Near Memory ComputingabstractIn-time processing of database system is imperative to reveal the hidden information. JOIN operation is critical in data analysis, as it occupies almost half of the average execution time in the standard TPC-H benchmark for database processing. In modern databases, transferring data between computing engines and system memory has become one of the main performance challenges. Previous works of Near Memory Computing (NMC) alleviated the costly data transfer, however, the designs still pose inefficiency in terms of processing flow and data management. In this paper, we propose FG-SMJ: a highly parallel fine-grained sort-merge join on near memory computing. The novel data layout allows us to access data from memory chips with fine-grained chip-level parallelism and exploit memory bandwidth. Compared with previous NMC designs, the proposed FG-SMJ attains 3.08x speedup. Po-Yen Lin, Yen-Shi Kuo, Bo-Cheng Lai |
ISCAS | 3 |
| 2022 | Anime Character Recognition using Intermediate Features AggregationabstractIn this work we study the problem of anime character recognition. Anime, refers to animation produced within Japan and work derived or inspired from it. We propose a novel Intermediate Features Aggregation classification head, which helps smooth the optimization landscape of Vision Transformers (ViTs) by adding skip connections between intermediate layers and the classification head, thereby improving relative classification accuracy by up to 28%. The proposed model, named as Animesion, is the first end-to-end framework for large-scale anime character recognition. We conduct extensive experiments using a variety of classification models, including CNNs and self-attention based ViTs. We also adapt its multimodal variation Vision-Language Transformer (ViLT), to incorporate external tag data for classification, without additional multimodal pre-training. Through our results we obtain new insights into the effects of how hyperparameters such as input sequence length, mini-batch size, and variations on the architecture, affect the transfer learning performance of Vi(L)Ts. Edwin Arkel Rios, Min-Chun Hu 0001, Bo-Cheng Lai |
ISCAS | 3 |
| 2022 | DLPrPPG: Development and Design of Deep Learning Platform for Remote PhotoplethysmographyabstractThis paper presents a comprehensive neural network-based development platform for remote photoplethysmography (rPPG). rPPG is a growing and popular research area, especially with the introduction of deep learning methods that can significantly improve its signal quality and heart rate prediction reliability. However, there are still many problems with the experimental methods in current studies, such as non-standardized and private data, different pre-processing methods, and incomplete or irreproducible experiment methodologies, among others. These problems prevent methods from being compared fairly and lead to lower reliability of the proposed experimental results, hindering progress in this area. For these reasons, we propose an open-source framework to facilitate the design and experimentation of deep learning-based rPPG development, and it’s made freely available on GitHub(DLPrPPG). Through our platform we provide ready-to-use implementations of CNN-AE, LSTM, GAN, and Transformer models, whose hyperparameters we can easily and quickly optimize, and efficiently compare in a fair fashion. From our experiments we show that if the parameters of different neural networks are optimized, the performance of older architectures can be on par or even outperform newer ones. Bo-Rong Yan, Edwin Arkel Rios, Wen-Hsien Lee, Bo-Cheng Lai |
ISCAS | 4 |
| 2021 | Parametric Study of Performance of Remote Photopletysmography SystemabstractRemote photoplethysmography (RPPG) is a technique in which we measure sub-cutaneous variations in blood flow, usually through a camera, to obtain physiological signals. Studies involving RPPG have increased in the past few years due to its numerous applications including remote healthcare, anti-spoofing, among others. While there have been many studies on how to increase RPPG's accuracy of bio-markers predictions in a variety of settings, most of them are usually done using workstation computers, yet some of the most promising applications of RPPG probably would be on limited resources, low-power embedded systems. Therefore, we did an extensive study on the effects of one of the most important design parameters in RPPG systems, sliding window (SW) size, for a variety of algorithms, in order to quantify the trade-off between computational cost in time and accuracy in root-mean-squared-error (RMSE), using a standardized public database. We also studied how different face detection and region-of-interest selection affected these results. Finally, based on these, we came up with a new and simple metric that takes into account both computation and accuracy, as a means to design dynamic systems which make the best out of the available resources With correct tuning, we can use this metric to reduce computational costs by up to 47%. Edwin Arkel Rios, Chih-Chieh Lai, Bo-Rong Yan, Bo-Cheng Lai |
ISCAS | 4 |
| 2020 | On EDA Solutions for Reconfigurable Memory-Centric AI Edge ApplicationsabstractMemory-centric designs deploy computation to storage and enable efficient in-memory computation while avoiding massive amount of data movement. The in-memory-computing schemes have shown distinct advantages and concerns when applying to different types of memory technologies, from conventional SRAM, DRAM to emerging ReRAM. Moreover, the next-generation smart edge systems are expected to support various intelligent applications by employing multi-task machine learning models which would be dynamically activated. To attain an efficient design within short design cycle, it is imperative to have an integrated design framework with automated tools to support hybrid memory systems and perform effective optimization across design stages. This work will introduce a unified framework which integrates EDA solutions to address the design and optimization challenges at different aspects of next-generation memory-centric designs, including fast reconfiguring in-memory/near-memory computing designs to provide optimized solutions (behavioral models and APR cell layouts) for designers to choose the best suitable architectures for their applications. Hung-Ming Chen, Chia-Lin Hu, Kang-Yu Chang, Alexandra Küster, Yu-Hsien Lin, Po-Shen Kuo, Wei-Tung Chao, Bo-Cheng Lai, Chien-Nan Jimmy Liu, Shyh-Jye Jou |
ICCAD | 8 |
| 2020 | Selective bypassing and mapping for heterogeneous applications on GPGPUs
Moustafa Emara, Bo-Cheng Lai |
J. Parallel Distributed Comput. | 2 |
| 2020 | REMAP+: An Efficient Banking Architecture for Multiple Writes of Algorithmic MemoryabstractSupporting multiple write ports is one of the main challenges when designing algorithmic multiported memory (AMM). AMM supports concurrent accesses by cooperating multiple, low-complexity memory modules together with logical operations. When scaling the number of write ports, the nontable-based approaches quadratically increase the number of memory modules, whereas the table-based approaches tend to introduce complex lookup tables and access handling logics. In this article, we introduce REMAP+, an efficient banking architecture to support multiple writes. We optimize the pipeline of REMAP+ to achieve high access bandwidth and more efficient table access. We also exploit the structured architecture of REMAP+ and propose a systematic design flow to automate the scaling of write ports and optimization of banking. Comprehensive analysis is presented to reveal the insight into design features and concerns. Based on extensive experiments, we have shown that REMAP+ outperforms the existing write schemes (XOR, live value table (LVT), and REMAP) with higher bandwidth (49%, 50%, 18%), lower energy (28%, 49%, 54%), and smaller area (43%, 37%, 35%). Bo-Cheng Lai, Bo-Ya Chen, Bo-En Chen, Yi-Da Hsin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Enhancing Utilization of SIMD-Like Accelerator for Sparse Convolutional Neural NetworksabstractAlthough the existing single-instruction-multiple-data-like (SIMD) accelerators can handle the compressed format of sparse convolutional neural networks, the sparse and irregular distributions of nonzero elements cause low utilization of multipliers in a processing engine (PE) and imbalanced computation between PEs. This brief addresses the above issues by proposing a data screening and task mapping (DSTM) accelerator which integrates a series of techniques, including software refinement and hardware modules. An efficient indexing module is introduced to identify the effectual computation pairs and skip unnecessary computation in a line-grained manner. The intra-PE load imbalance is alleviated with weight data rearrangement. An effective task sharing mechanism further balances the computation between PEs. When compared with the state-of-the-art SIMD-like accelerator, the proposed DSTM enhances the average PE utilization by 3.5×. The overall processing throughput is 59.7% higher than the previous design. Bo-Cheng Lai, Jyun-Wei Pan, Chien-Yu Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Supporting compressed-sparse activations and weights on SIMD-like accelerator for sparse convolutional neural networksabstractSparsity is widely observed in convolutional neural networks by zeroing a large portion of both activations and weights without impairing the result. By keeping the data in a compressed-sparse format, the energy consumption could be considerably cut down due to less memory traffic. However, the wide SIMD-like MAC engine adopted in many CNN accelerators can not support the compressed input due to the data misalignment. In this work, a novel Dual Indexing Module (DIM) is proposed to efficiently handle the alignment issue where activations and weights are both kept in compressed-sparse format. The DIM is implemented in a representative SIMD-like CNN accelerator, and able to exploit both compressed-sparse activations and weights. The synthesis results with 40nm technology have shown that DIM can enhance up to 46% of energy consumption and 55.4% Energy-Delay-Product (EDP). Chien-Yu Lin, Bo-Cheng Lai |
ASP-DAC | 2 |
| 2018 | Towards high performance data analytic on heterogeneous many-core systems: A study on Bayesian Sequential Partitioning
Bo-Cheng Lai, Tung-Yu Wu, Tsou-Han Chiu, Kun-Chun Li, Chia-Ying Lee, Wei-Chen Chien, Wing Hung Wong |
J. Parallel Distributed Comput. | 1 |
| 2017 | An Efficient Hierarchical Banking Structure for Algorithmic Multiported Memory on FPGAabstractAlgorithmic multiported memory supports concurrent accesses by cooperating block RAMs (BRAMs) with algorithmic operations, and demonstrates the better performance per resource usage on FPGA when compared with register-based designs. However, the current approaches still use significant amount of FPGA resources and pose great design challenges when increasing the access ports. This paper proposes HB-NTX with a resource efficient hierarchical banking structure for nontable-based multi-ported memory design on FPGA. The regular design style enables a systematic flow to scale both read and write ports. When compared with the previous approaches, HB-NTX can reduce 62.03% BRAMs when composing a 2R4W memory with 32K depth. This paper further extends the HB-NTX to alleviate the complexity of the table-based memory designs. When compared with the previous table-based TBLVT approach, the proposed design for a 2R4W memory with 8K depth attains the cost reduction of 39.9%, 14.3%, and 15.6%, for registers, lookup tables, and BRAMs, respectively. Bo-Cheng Lai, Kun-Hua Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Efficient Designs of Multiported Memory on FPGAabstractThe utilization of block RAMs (BRAMs) is a critical performance factor for multiported memory designs on field-programmable gate arrays (FPGAs). Not only does the excessive demand on BRAMs block the usage of BRAMs from other parts of a design, but the complex routing between BRAMs and logic also limits the operating frequency. This paper first introduces a brand new perspective and a more efficient way of using a conventional two reads one write (2R1W) memory as a 2R1W/4R memory. By exploiting the 2R1W/4R as the building block, this paper introduces a hierarchical design of 4R1W memory that requires 25% fewer BRAMs than the previous approach of duplicating the 2R1W module. Memories with more read/write ports can be extended from the proposed 2R1W/4R memory and the hierarchical 4R1W memory. Compared with previous xor-based and live value table-based approaches, the proposed designs can, respectively, reduce up to 53% and 69% of BRAM usage for 4R2W memory designs with 8K-depth. For complex multiported designs, the proposed BRAM-efficient approaches can achieve higher clock frequencies by alleviating the complex routing in an FPGA. For 4R3W memory with 8K-depth, the proposed design can save 53% of BRAMs and enhance the operating frequency by 20%. Bo-Cheng Lai, Jiun-Liang Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Enhancing Data Reuse in Cache Contention Aware Thread Scheduling on GPGPUabstractGPGPUs have been widely adopted as throughput processing platforms for modern big-data and cloud computing. Attaining a high performance design on a GPGPU requires careful tradeoffs among various design concerns. Data reuse, cache contention, and thread level parallelism, have been demonstrated as three imperative performance factors for a GPGPU. The correlated performance impacts of these factors pose non-trivial concerns when scheduling threads on GPGPUs. This paper proposes a three-staged scheduling scheme to coschedule the threads with consideration of the three factors. The experiment results on a set of irregular parallel applications, when compared with previous approaches, have demonstrated up to 70% execution time improvement. Chin-Fu Lu, Hsien-Kai Kuo, Bo-Cheng Lai |
CISIS | 3 |
| 2016 | Unified Designs for High Performance LDPC Decoding on GPGPUabstractModern GPGPU's have enabled massively parallel computing with programmability that can exploit the highly parallel nature of LDPC decoding. Previous works customized the design on a GPGPU towards specific execution attributes of a particular LDPC decoding matrix. Supporting different LDPC decoding matrices requires either substantial rework on the current program, or a brand new parallel design. This paper proposes two unified designs that can achieve high performance for both regular and irregular LDPC decoding on a GPGPU. The first design introduces a node-based scheme with a versatile translation array mechanism that can efficiently handle the complex data access patterns of different LDPC decoding matrices. The second design proposes an edge-based parallel paradigm that uses more intuitive data layout. More edges than nodes in a Tanner graph also give the edge-based design higher computation parallelism when there are limited concurrent codewords. With the proposed unified designs, designers can be ignorant of the types of LDPC matrices and achieve high performance LDPC decoding. The experiments on a GTX 470 GPGPU have demonstrated up to 134.56x runtime improvement, when compared with designs on a high-end CPU. The maximum throughput can reach 80.25 Mbps. When compared with the previous customized designs, the proposed systematic designs can reach better performance while relieving the effort of customization. Bo-Cheng Lai, Chia-Ying Lee, Tsou-Han Chiu, Hsien-Kai Kuo, Chun-Kai Chang |
IEEE Trans. Computers | 1 |
| 2015 | Self adaptable multithreaded object detection on embedded multicore systems
Bo-Cheng Lai, Kun-Chun Li, Guan-Ru Li, Chih-Hsuan Chiang |
J. Parallel Distributed Comput. | 1 |
| 2015 | A Cache Hierarchy Aware Thread Mapping Methodology for GPGPUsabstractThe recently proposed GPGPU architecture has added a multi-level hierarchy of shared cache to better exploit the data locality of general purpose applications. The GPGPU design philosophy allocates most of the chip area to processing cores, and thus results in a relatively small cache shared by a large number of cores when compared with conventional multi-core CPUs. Applying a proper thread mapping scheme is crucial for gaining from constructive cache sharing and avoiding resource contention among thousands of threads. However, due to the significant differences on architectures and programming models, the existing thread mapping approaches for multi-core CPUs do not perform as effective on GPGPUs. This paper proposes a formal model to capture both the characteristics of threads as well as the cache sharing behavior of multi-level shared cache. With appropriate proofs, the model forms a solid theoretical foundation beneath the proposed cache hierarchy aware thread mapping methodology for multi-level shared cache GPGPUs. The experiments reveal that the three-staged thread mapping methodology can successfully improve the data reuse on each cache level of GPGPUs and achieve an average of 2.3× to 4.3× runtime enhancement when compared with existing approaches. Bo-Cheng Lai, Hsien-Kai Kuo, Jing-Yang Jou |
IEEE Trans. Computers | 1 |
| 2015 | Scalable Global Power Management Policy Based on Combinatorial Optimization for MultiprocessorsabstractMultiprocessors have become the main architecture trend in modern systems due to the superior performance; nevertheless, the power consumption remains a critical challenge. Global power management (GPM) aims at dynamically finding the power state combination that satisfies the power budget constraint while maximizing the overall performance (or vice versa). Due to the increasing number of cores in a multiprocessor system, the scalability of GPM policies has become critical when searching satisfactory state combinations within acceptable time. This article proposes a highly scalable policy based on combinatorial optimization with theoretical proofs, whereas previous works take exhaustive search or heuristic methods. The proposed policy first applies an optimum algorithm to construct a state combination table in pseudo--polynomial time using dynamic programming. Then, the state combination is assigned to cores with minimum transition cost in linear time by mapping to the network flow problem. Simulation results show that the proposed policy achieves better system performance for any given power budget when compared to the state-of-the-art heuristic. Furthermore, the proposed policy demonstrates its prominent scalability with 125 times faster policy runtime for 512 cores. Gung-Yu Pan, Jed Yang, Jing-Yang Jou, Bo-Cheng Lai |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2015 | A High-Performance Double-Layer Counting Bloom Filter for Multicore SystemsabstractThe snoopy-based protocol is a widely used cache coherence mechanism for a symmetric multiprocessor (SMP) system. However, this broadcast-based protocol blindly disseminates data sharing information across the system, and introduces many unnecessary data operations. This paper proposes a novel architecture of double-layer counting Bloom filter (DLCBF) to reduce the unnecessary data lookups on the local cache and redundant data transactions on the shared interconnection of an SMP system. By adding an extra filtering layer, the DLCBF effectively exploits the data locality of applications. The two-layer hierarchy reduces the storage size of DLCBF by 18.75%, and achieves 81.99% and 31.36% better filtering rates when compared with a classic Bloom filter (BF) and original counting BF, respectively. When applied on the segmented shared bus of an SMP system, the DLCBF outperforms the previous work by 58% for In-filters and $1.86\times $ for Out-filters. This paper also comprehensively explores the key design parameters of DLCBF, including the sizes of top-layer, bottom-layer, and multilayer design. The results show that enlarging the layer filters enhance the filtering rates of DLCBF, while adding an extra filter layer only provides slight benefit. Bo-Cheng Lai, Kuan-Ting Chen, Ping-Ru Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | A learning-on-cloud power management policy for smart devicesabstractEnergy consumption poses severe limitations for smart devices, urging the development of effective and efficient power management policies. State-of-the-art learning-based policies are autonomous and adaptive to the environment, but they are subject to costly computational overhead and lengthy convergence time. As smart devices are connected to Internet, this paper proposes the Learning-on-Cloud (LoC) policy to exploit cloud computing for power management. Sophisticated learning engines are offloaded from local devices to the cloud with minimal communication data, thus the runtime overhead is reduced. The learning data are shared between many devices with the same model, hence the convergence rate is raised. With one thousand devices connecting to the cloud, the LoC agent is able to converge within a few iterations; the energy saving is better than both of the greedy and the learning-based policies with less latency penalty. By implementing the LoC policy as an Android App, the measured overhead is only 0.01% of the system time. Gung-Yu Pan, Bo-Cheng Lai, Sheng-Yen Chen, Jing-Yang Jou |
ICCAD | 2 |
| 2014 | Automatic Data Layout Transformation for Heterogeneous Many-Core Systems
Ying-Yu Tseng, Bo-Cheng Lai, Jiun-Liang Lin |
NPC | 3 |
| 2014 | Reducing Contention in Shared Last-Level Cache for Throughput ProcessorsabstractDeploying the Shared Last-Level Cache (SLLC) is an effective way to alleviate the memory bottleneck in modern throughput processors, such as GPGPUs. A commonly used scheduling policy of throughput processors is to render the maximum possible thread-level parallelism. However, this greedy policy usually causes serious cache contention on the SLLC and significantly degrades the system performance. It is therefore a critical performance factor that the thread scheduling of a throughput processor performs a careful trade-off between the thread-level parallelism and cache contention. This article characterizes and analyzes the performance impact of cache contention in the SLLC of throughput processors. Based on the analyses and findings of cache contention and its performance pitfalls, this article formally formulates the aggregate working-set-size-constrained thread scheduling problem that constrains the aggregate working-set size on concurrent threads. With a proof to be NP-hard, this article has integrated a series of algorithms to minimize the cache contention and enhance the overall system performance on GPGPUs. The simulation results on NVIDIA's Fermi architecture have shown that the proposed thread scheduling scheme achieves up to 61.6% execution time enhancement over a widely used thread clustering scheme. When compared to the state-of-the-art technique that exploits the data reuse of applications, the improvement on execution time can reach 47.4%. Notably, the execution time improvement of the proposed thread scheduling scheme is only 2.6% from an exhaustive searching scheme. Hsien-Kai Kuo, Bo-Cheng Lai, Jing-Yang Jou |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Scalable Power Management Using Multilevel Reinforcement Learning for MultiprocessorsabstractDynamic power management has become an imperative design factor to attain the energy efficiency in modern systems. Among various power management schemes, learning-based policies that are adaptive to different environments and applications have demonstrated superior performance to other approaches. However, they suffer the scalability problem for multiprocessors due to the increasing number of cores in a system. In this article, we propose a scalable and effective online policy called MultiLevel Reinforcement Learning (MLRL). By exploiting the hierarchical paradigm, the time complexity of MLRL is O ( n lg n ) for n cores and the convergence rate is greatly raised by compressing redundant searching space. Some advanced techniques, such as the function approximation and the action selection scheme, are included to enhance the generality and stability of the proposed policy. By simulating on the SPLASH-2 benchmarks, MLRL runs 53% faster and outperforms the state-of-the-art work with 13.6% energy saving and 2.7% latency penalty on average. The generality and the scalability of MLRL are also validated through extensive simulations. Gung-Yu Pan, Jing-Yang Jou, Bo-Cheng Lai |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2013 | Cache Capacity Aware Thread Scheduling for Irregular Memory Access on many-core GPGPUsabstractOn-chip shared cache is effective to alleviate the memory bottleneck in modern many-core systems, such as GPGPUs. However, when scheduling numerous concurrent threads on a GPGPU, a cache capacity agnostic scheduling scheme could lead to severe cache contention among threads and thus significant performance degradation. Moreover, the diverse working sets in irregular applications make the cache contention issue an even more serious problem. As a result, taking cache capacity into account has become a critical scheduling issue of GPGPUs. This paper formulates a Cache Capacity Aware Thread Scheduling Problem to capture the impact of cache capacity as well as different architectural considerations. With a proof to be NP-hard, this paper has proposed two algorithms to perform the cache capacity aware thread scheduling. The simulation results on Nvidia's Fermi configuration have shown that the proposed scheduling scheme can effectively avoid cache contention, and achieve an average of 44.7% cache miss reduction and 28.5% runtime enhancement. The paper also shows the runtime can be enhanced up to 62.5% for more complex applications. Hsien-Kai Kuo, Ta-Kan Yen, Bo-Cheng Lai, Jing-Yang Jou |
ASP-DAC | 3 |
| 2013 | A Locality-Aware Dynamic Thread Scheduler for GPGPUsabstractModern GPGPUs implement on-chip shared cache to better exploit the data reuse of various general purpose applications. Given the massive amount of concurrent threads in a GPGPU, striking the balance between Data Locality and Load Balance has become a critical design concern. To achieve the best performance, the trade-off between these two factors needs to be performed concurrently. This paper proposes a dynamic thread scheduler which co-optimizes both the data locality and load balance on a GPGPU. The proposed approach is evaluated using three applications with various input datasets. The results show that the proposed approach reduces the overall execution cycles by up to 16% when compared with other approaches concerning only one objective. Ying-Yu Tseng, Hsien-Kai Kuo, Ta-Kan Yen, Bo-Cheng Lai |
PDCAT | 5 |
| 2012 | Thread affinity mapping for irregular data access on shared Cache GPGPUabstractMemory Coalescing and on-chip shared Cache are two effective techniques to alleviate the memory bottleneck in modern GPGPUs. These two techniques are very useful on applications with regular memory accesses. However, they become ineffective on concurrent threads with large numbers of uncoordinated accesses and the potential performance benefit could be significantly degraded. This paper proposes a thread affinity mapping methodology to coordinate the irregular data accesses on shared cache GPGPUs. Based on the proposed affinity metrics, threads are congregated into execution groups which are able to fully exploit the memory coalescing and data sharing within an application. An average of 3.5x runtime speedup is achieved on a Fermi GPGPU. The speedup scales with the sizes of test cases, which makes the proposed methodology an effective and promising solution for the continually increasing complexities of applications in the future many-core systems. Hsien-Kai Kuo, Kuan-Ting Chen, Bo-Cheng Lai, Jing-Yang Jou |
ASP-DAC | 3 |
| 2011 | Classifier Grouping to Enhance Data Locality for a Multi-threaded Object Detection AlgorithmabstractObject detection has become an enabling function for modern smart embedded devices to perform intelligent applications and interact with the environment appropriately and promptly. However, the limited computation resource of embedded devices has become a barrier to execute the computation intensive object detection algorithm. Leveraging the multi-threading scheme on embedded multi-core systems provides an opportunity to boost the performance. However, the memory bottleneck limits the performance scalability. Improving data locality of applications and maximizing the data reuse for on-chip caches have therefore become critical design concerns. This paper comprehensively analyzes the memory behavior and data locality of a multi-threaded object detection algorithm. A novel Classifier-Grouping scheme is proposed to significantly enhance the data reuse for on-chip caches of embedded multicore systems. By executing a multi-threaded object detection algorithm on a cycle-accurate multi-core simulator, the proposed approach can achieve up to 62% better performance when compared with the original parallel program. Bo-Cheng Lai, Chih-Hsuan Chiang, Guan-Ru Li |
ICPADS | 1 |
| 2008 | A Cost-Effective Latency-Aware Memory Bus for Symmetric Multiprocessor SystemsabstractThis paper presents how a multi-core system can benefit from the use of a latency-aware memory bus capable of dual-concurrent data transfers on a single wire line: Source synchronous CDMA interconnect (SSCDMA-I) has been adopted to implement the memory bus of a shared-memory multi-core system. Two types of bus-based homogeneous and heterogeneous multi-core systems are modeled and simulated by a cycle-accurate simulation platform. Unlike the conventional time-division multiplexing (TDM) bus-based multi-core system that shows degradation in performance as the number of processing cores increases, the proposed SSCDMA bus-based multi-core shows higher performance up to 23.1% for 4 cores. The maximum latency of a heterogeneous multi-core system with a mix of traffic loads has been reduced up to 78%. These results demonstrate that the performance of multi-core systems can be improved with less cost and network complexity by reducing the bus contention interferences and by supporting higher concurrency in memory accesses that brings shorter critical word access latency. Jongsun Kim, Bo-Cheng Lai, Mau-Chung Frank Chang, Ingrid Verbauwhede |
IEEE Trans. Computers | 2 |
| 2006 | Cross Layer Design to Multi-thread a Data-Pipelining Application on a Multi-processor on ChipabstractData-Pipelining is a widely used model to represent streaming applications. Incremental decomposition and optimization of a data-pipelining application onto a multi-processor platform spans multiple design layers, including the application layer, the system software layer, the architecture layer and the micro-architecture layer. For best results, designers have to consider multiple design layers (vertical exploration) and multiple architecture options (horizontal exploration). By using a data-pipelining JPEG encoder as the application driver, this paper presents a comprehensive analysis of mapping a data-pipelined application through multiple design layers, to a shared-memory SMP (Symmetric Multi- Processing) system. It is shown that a single-layered optimization ends up with a 110% worse design if the system effects from other layers are not taken into account. Compared to the nominal case, with appropriate mapping of the application, we achieve 47.5% improvement for high performance design and 77.6% energy reduction for energy efficient design under constant performance. Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede |
ASAP | 1 |
| 2005 | Prototype IC with WDDL and Differential Routing - DPA Resistance Assessment
Kris Tiri, David Hwang 0001, Alireza Hodjat, Bo-Cheng Lai, Shenglin Yang, Patrick Schaumont, Ingrid Verbauwhede |
CHES | 4 |
| 2005 | Cooperative multithreading on 3mbedded multiprocessor architectures enables energy-scalable designabstractWe propose an embedded multiprocessor architecture and its associated thread-based programming model. Using a cycle-true simulation model of this architecture, we are able to estimate energy savings for a threaded C program. The savings are obtained by voltage- and frequency-scaling of the individual processors. We port a fingerprint minutiae detection application onto this architecture, and show the resulting performance on single-, dual-, and quad-processor configurations. The energy-scaled quadprocessor version results in a 77% energy reduction over the single-processor non-scaled implementation, at only a 2.2% degradation in cycle count. Patrick Schaumont, Bo-Cheng Lai, Ingrid Verbauwhede |
DAC | 2 |
| 2005 | A side-channel leakage free coprocessor IC in 0.18µm CMOS for embedded AES-based cryptographic and biometric processingabstractSecurity ICs are vulnerable to side-channel attacks (SCAs) that find the secret key by monitoring the power consumption and other information that is leaked by the switching behavior of digital CMOS gates. This paper describes a side-channel attack resistant coprocessor IC and its design techniques. The IC has been fabricated in 0.18µm CMOS. The coprocessor, which is used for embedded cryptographic and biometric processing, consists of four components: an Advanced Encryption Standard (AES) based cryptographic engine, a fingerprint-matching oracle, a template storage, and an interface unit. Two functionally identical coprocessors have been fabricated on the same die. The first, 'secure', coprocessor is implemented using a logic style called Wave Dynamic Digital Logic (WDDL) and a layout technique called differential routing. The second, 'insecure', coprocessor is implemented using regular standard cells and regular routing techniques. Measurement-based experimental results show that a differential power analysis (DPA) attack on the insecure coprocessor requires only 8,000 acquisitions to disclose the entire 128b secret key. The same attack on the secure coprocessor still does not disclose the entire secret key at 1,500,000 acquisitions. This improvement in DPA resistance of at least 2 orders of magnitude makes the attack de facto infeasible. The required number of measurements is larger than the lifetime of the secret key in most practical systems. Kris Tiri, David Hwang 0001, Alireza Hodjat, Bo-Cheng Lai, Shenglin Yang, Patrick Schaumont, Ingrid Verbauwhede |
DAC | 4 |
| 2005 | A 3.84 gbits/s AES crypto coprocessor with modes of operation in a 0.18-µm CMOS technologyabstractIn this paper an AES crypto coprocessor that is fabricated using a 0.18-μm CMOS technology is presented. This crypto coprocessor performs the AES-128 encryption in both feedback and non-feedback modes of operation. A maximum throughput of 3.84 Gbits/s is achieved at a 330 MHz clock frequency for ECB, OFB, and CBC modes of operation. This crypto coprocessor can be programmed using the memory-mapped interface of an embedded CPU core and is tested using a LEON 32-bit (SPARC V8) processor in the ThumbPod secure system-on-chip. Alireza Hodjat, David Hwang 0001, Bo-Cheng Lai, Kris Tiri, Ingrid Verbauwhede |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Energy and Performance Analysis of Mapping Parallel Multithreaded Tasks for An On-Chip Multi-Processor SystemabstractMultiprocessor systems offer superior performance and potentially better energy-reduction than single-processor systems. It all depends, however, on how well the application can be mapped onto the architecture. Indeed, a careful tradeoff of energy and performance requires a thorough understanding of the energy consumption pattern of the application across the architecture. We develop a simulation platform, MultiPo-Sim, which returns the cycle-accurate performance and energy consumption of a multiprocessor system, for both hardware components and software primitives. On the hardware level, energy scaling techniques can be modeled and each processing core can operate at different energy modes. MultiPo-Sim achieves 331K cycles per second simulation speed for a four-processor system on a 3GHz, 512MByte Fedora-2 PC. On the software level, data parallelizing and task parallelizing are two common models of multi-thread programming. By using MultiPo-Sim, we show that they show different energy and performance characteristics when mapping onto a multi-processor system. Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede |
ICCD | 1 |
| 2004 | Reducing radio energy consumption of key management protocols for wireless sensor networksabstractThe security of sensor networks is a challenging area. Key management is one of the crucial parts in constructing the security among sensor nodes. However, key management protocols require a great deal of energy consumption, particularly in the transmission of initial key negotiation messages. In this paper, we examine three previously published sensor network security schemes: SPINS and C&R for master-key-based schemes, and Eschenhaur-Gligor (EG) for distributed-key-based schemes. We then present two new low-power schemes, which we call BROSK and OKS as alternatives to master-key-based schemes and distributed-key-based schemes, respectively. Compared to SPINS and C&R protocols, BROSK can reduce energy consumption by up to 12X by reducing the number of data transmissions in the key negotiation process. Compared with EG, OKS reduces energy by up to 96% and reduces memory requirements by up to 78%. Bo-Cheng Lai, David Hwang 0001, Sungha Pete Kim, Ingrid Verbauwhede |
ISLPED | 1 |
| 2003 | Design flow for HW / SW acceleration transparency in the thumbpod secure embedded systemabstractThis paper describes a case study and design flow of a secure embedded system called ThumbPod, which uses cryptographic and biometric signal processing acceleration. It presents the concept of HW/SW acceleration transparency, a systematic method to accelerate Java functions in both software and hardware. An example of acceleration transparency for a Rijndael encryption function is presented. The embedded prototype hardware platform is also described. Acceleration transparency yields software and hardware performance gains of 333X. David Hwang 0001, Bo-Cheng Lai, Patrick Schaumont, Kazuo Sakiyama, Shenglin Yang, Alireza Hodjat, Ingrid Verbauwhede |
DAC | 2 |
| 2002 | A Security Protocol for Biometric Smart Cards
David Hwang 0001, Bo-Cheng Lai, Patrick Schaumont, Ingrid Verbauwhede |
CARDIS | 2 |