EDBT 2026 Demo / reviewers in the wild / expert
Yun-Chen Lo
dblp:220/9112
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-1324-7649ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | P-DAC: Power-Efficient Photonic Accelerators for LLM InferenceabstractAs traditional electronic hardware encounters the limitations of Moore’s Law, optical computing is emerging as a promising alternative, delivering high data transmission rates, especially beneficial for big data and AI applications. Photonic accelerators, such as the LighteningTransformer, utilize optical analog signals to accelerate Transformerbased models, achieving exceptional speed and low energy consumption. However, controlling modern optical intensity modulators (e.g., MachZehnder Modulators) requires using electrical analog signals (e.g., voltage values) to adjust the optical signal intensity for realizing optical-based vector inner product calculations. Managing this modulation consumes significant power, as it involves selecting optimal electrical values through an electrical controller and converting digital signals to analog using digital-to-analog converters (DACs). In this work, we introduce P-DAC, a solution designed to reduce DAC power consumption, significantly enhancing the energy efficiency of optical accelerators for Transformer models. Wen-Tse Chang, Chun-Feng Wu, Yun-Chen Lo |
DAC | 3 |
| 2024 | ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-MemoryabstractAmong several emerging architectures, computing in memory (CIM), which featuresin-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out thatmutually stationary vectors (MSVs), which can be maximized by introducingassociativityto CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. We have designed and realized an SA-CIM silicon prototype and corresponding architecture and acceleration schemes in the TSMC 28 nm process. More specifically, the contributions of this paper are fivefold: 1) We identify MSVs as new features that can be exploited to improve the current performance and energy challenges of the CIM-based hardware. 2) We propose SA-CIM to enhance MSVs (input-reordering flexibility) for skipping the zeros, small values, and sparse vectors. 3) We propose channel swapping to enhance the zero-skipping technique. 4) We propose a transposed systolic dataflow to efficiently conduct conv3×3 while being capable of exploiting input-skipping schemes. 5) We propose a design flow to search for optimal aggressive skipping scheme setups while satisfying the accuracy loss constraint. The proposed ISSA architecture improves the throughput by 1.91× to 2.97× speedup and the energy efficiency by 2.5× to 4.2×. Yun-Chen Lo, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Chih-Chen Yeh, Wen-Chien Ting, Ren-Shuo Liu |
IEEE Trans. Computers | 1 |
| 2023 | Bit-Serial Cache: Exploiting Input Bit Vector Repetition to Accelerate Bit-Serial InferenceabstractBit-serial computation has demonstrated superiority in processing precision-varying DNNs by slicing multi-bit vectors into multiple single-bit vectors and computing the inner product using multiple steps of shift-and-adds. In this paper, we identify that performing real-world DNNs inference with bit-serial computation exhibits high input bit vector locality, where up to 85.7% of non-zero input bit vectors, as well as their associated computation, are previously-seen and previously-done ones. We propose Bit-Serial Cache to transfer the identified locality into performance and energy efficiency gains. The key design strategy is to store recently-computed partial sums of input bit vectors to a cache and utilize cache accesses to replace redundant computations. In addition to the bit-serial computation architecture, we also present: 1) request clustering and 2) interleaved scheduling, to further enhance the performance and energy efficiency.Our experiments using six popular DNNs (in both 8-b and 4-b) show that Bit-Serial Cache speeds up DNN inference by up to 2.72×, 1.82×, and 4.03×, energy efficiency by 3.19×, 3.29×, and 2.82×, area efficiency by 1.35×, 1.24×, and 2.76× over state-of-the-art Loom, DPRed Loom, and Laconic. Yun-Chen Lo, Ren-Shuo Liu |
DAC | 1 |
| 2023 | Morphable CIM: Improving Operation Intensity and Depthwise Capability for SRAM-CIM ArchitectureabstractSRAM-based computing in memory (SRAM-CIM) supports weight-stationary dataflows and in-situ computation, successfully achieving lower weight SRAM traffic and hence better operation intensity (OI) than Von Neumann architecture. However, sticking to weight-stationary dataflows is sub-optimal since some layers of CNNs, e.g., depth-wise and input-dominant, suffer from large activation traffic and underutilization.We propose Morphable CIM to address the above challenges, in which the key contributions are: 1) We propose a dataflow-morphable SRAM-CIM architecture that adaptively switches between weight-stationary and input-stationary dataflow to enhance the overall operation intensity (OI) and computation utilization. 2) We propose an input-stationary systolic dataflow and a word-wise mapping to efficiently achieve dataflow reconfigurability for SRAM-CIM. 3) We propose a depthwise-capable CIM macro to improve the utilization of processing depth-wise layers.The experimental results on ten workloads show that our proposed SRAM-CIM architecture successfully outperforms the traffic-optimized and the performance-optimized baselines by up to 16.7× performance speedup, 3.6× higher energy efficiency, and 94.8% memory traffic reduction. Yun-Chen Lo, Ren-Shuo Liu |
DAC | 1 |
| 2023 | BICEP: Exploiting Bitline Inversion for Efficient Operation-Unit-Based Compute-in-Memory Architecture: No Retraining Needed!abstractCompute-in-memory (CIM) architecture is promising for its in-situ analog computing ability. However, one practical constraint for CIM architectures is the limited number of activated rows in an operation Unit (OU). OU-based CIM architecture only activates a subgroup of memory cells to ensure a large signal margin and enough consideration of non-ideal device/circuit effect, which pays the cost of lowered computing throughput. In short, the OU-based CIM architectures suffer from array underutilization to ensure high accuracy.This work proposes a novel architecture, BICEP, which exploits bitline inversion technique to enlarge the OU size without the need to prune, approximate, and retrain. More specifically, the key contributions of this work are threefold: 1) We propose a bitline inversion scheme, which guarantees more than 2× larger OU size without affecting the numerical results and the ADC resolution. The key insight is to selectively apply code inversion on heavy bitlines to constrain their MAC outputs and compensate using low-cost compensation units. We mathematically prove that the proposal can be applied to both single- and multi-level cells (SLC and MLC). 2) We propose an inversion-aware weight swapping scheme, which swaps the weight order to maximize the OU size exploiting bitline inversion. 3) We propose weight order propagation to enable inversion-aware weight swapping without storage overheads. The extensive experiments on ImageNet classification tasks demonstrate that this work outperforms state-of-the-art OU-based CIM architecture (DL-RSIM) by up to 2.06× speedup and 1.97× energy efficiency. Yun-Chen Lo, Chia-Chun Wang, Ren-Shuo Liu |
ICCD | 1 |
| 2023 | Exploiting and Enhancing Computation Latency Variability for High-Performance Time-Domain Computing-in-Memory Neural Network AcceleratorsabstractTo address the inefficiency resulting from data movement in Von Neumann architecture, computing-in-memory (CIM) is a promising solution due to its in-situ analog computation. Among the various types of CIMs, time-domain CIM stands out as a promising solution for achieving high energy efficiency and high readout resolution by employing time-to-digital converters (TDC) instead of analog-to-digital converters (ADC) to convert time-domain delays into digital values. However, the performance of the accelerator may be constrained by the maximum operating frequency of time-domain CIM, which is significantly lower than that of digital circuits.This paper proposes an architecture for a time-domain CIM-based neural network accelerator that leverages the varying output time of the TDC. The key contributions of this work are as follows: 1) We introduce an early-termination scheme for time-domain CIM, which dynamically determines the length of the CIM clock period by deriving the maximum possible multiply-accumulate (MAC) value based on the current input. This approach reduces computation time for low-MAC results. 2) We propose an input-inversion scheme to decrease the computation time for high-MAC results. By employing linear combination, we perform bit-inversion on large inputs and compensate for the results using a low-cost digital circuit. 3) We propose a hardware optimization on the compensation circuit by combining it with shift-adders in traditional neural network accelerators.Experiments show that our schemes could gain 2× ∼ 2.9× speedup under different clock period specifications with 5.82% area overhead compared to the CIM macro. Chia-Chun Wang, Yun-Chen Lo, Jun-Shen Wu, Yu-Chih Tsai, Chia-Cheng Chang, Tsen-Wei Hsu, Min-Wei Chu, Chuan-Yao Lai, Ren-Shuo Liu |
ICCD | 2 |
| 2023 | Block and Subword-Scaling Floating-Point (BSFP) : An Efficient Non-Uniform Quantization For Low Precision Inference
Yun-Chen Lo, Tse-Kuang Lee, Ren-Shuo Liu |
ICLR | 1 |
| 2023 | Bucket Getter: A Bucket-based Processing Engine for Low-bit Block Floating Point (BFP) DNNsabstractBlock floating point (BFP), an efficient numerical system for deep neural networks (DNNs), achieves a good trade-off between dynamic range and hardware costs. Specifically, prior works have demonstrated that BFP format with 3 ∼ 5-bit mantissa can achieve FP32-comparable accuracy for various DNN workloads. We find that the floating-point adder (FP-Acc), which contains modules for normalization, alignment, addition, and fixed-point-to-floating-point (FXP2FP) conversion, dominates the power and area overheads, hence hindering the hardware efficiency of state-of-the-art low-bit BFP processing engines (BFP-PE). Yun-Chen Lo, Ren-Shuo Liu |
MICRO | 1 |
| 2022 | ISSA: Input-Skippable, Set-Associative Computing-in-Memory (SA-CIM) Architecture for Neural Network AcceleratorsabstractAmong several emerging architectures, computing in memory (CIM), which features in-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out that mutually stationary vectors (MSVs), which can be maximized by introducing associativity to CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. Yun-Chen Lo, Chih-Chen Yeh, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Wen-Chien Ting, Ren-Shuo Liu |
ICCAD | 1 |
| 2021 | Interference-Free Design Methodology for Paper-Based Digital Microfluidic BiochipsabstractPaper-based digital microfluidic biochips (P-DMFBs) have recently attracted great attention for its low-cost, in-place, and fast fabrication. This technology is essential for agile bio-assay development and deployment. P-DMFBs print electrodes and associate control lines on paper to control droplets and complete bio-assays. However, P-DMFBs have following issues: 1) control line interference may cause unwanted droplet movements, 2) avoiding control interference degrades assay performance and routability, 3) single layer fabrication limits routability, and 4) expensive ink cost limits low-cost benefits of P-DMFBs. To solve above issues, this work proposes an interference-free design methodology to design P-DMFBs with fast assay speed, better routability, and compact printing area. The contributions are as follows: First, we categorize control interference into soft and hard. Second, we identify only soft interference happens and propose to remove soft control interference constraints. Third, we propose an interference-free design methodology. Finally, we propose a cost-efficient ILP-based fluidic design module. Experimental results show proposed method outperforms prior work [14] across all bio-assay benchmarks. Compared to previous work, our cost-optimized designs use only 47%~78% area, gain 3.6%~16.2% more routing resources, and achieve 0.97x~1.5x shorter assay completion time. Our performance-optimized designs can accelerate assay speed by 1.05x~1.65x using 81%~96% printed area. Yun-Chen Lo, Bing Li 0005, Sooyong Park, Kwanwoo Shin, Tsung-Yi Ho |
ASP-DAC | 1 |