VLDB 2026 Research / reviewers in the wild / expert
Longyang Lin
dblp:196/1722
· DBLP profile ↗
22ranked-venue papers
1as first author
19since 2021 · last 2026
0000-0002-4702-737XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 1 first-author · 18 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gundam: A Generalized Unified Design and Analysis Model for Matrix Multiplication on Edge
Weirong Dong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto |
ASP-DAC | 5 |
| 2026 | Ramen: Radiation-Aware Modeling Framework for PDK-Enabled Design and Library CharacterizationabstractRadiation-induced degradation poses a critical challenge to the reliability of space-grade integrated circuits (ICs). Existing radiation-aware models largely remain at the device level and lack direct integration with circuit or system design flows, limiting their practical use in radiation-aware IC design. To address this, this work proposes Ramen, a non-invasive radiation-aware device modeling framework that is fully compatible with commercial Process Design Kits (PDKs). Ramen accurately captures total ionizing dose (TID) and displacement damage dose (DDD), enabling early-stage evaluation at both circuit and system levels without requiring modifications to existing PDK structures. By seamlessly integrating with standard analog, mixed-signal, and digital flows, the radiation-aware models not only support SPICE-based circuit simulation but also feed into standard library characterization tools to generate radiation-aware Liberty libraries. These libraries encode dose-dependent timing, leakage, and power information, allowing radiation effects to be captured in synthesis, timing analysis, and back-end implementation. Experimental validation on a 180 nm CMOS imager under radiation stress shows that the proposed framework achieves <15% simulation errors for both analog and logic circuit, confirming the reliability of Ramen for radiation-aware IC design. Zhenzhe Chen, Wang Liao 0001, Jing-Jia Liou, Masanori Hashimoto, Longyang Lin |
DATE | 7 |
| 2026 | Gohan: A Golden-Copy-Aided Platform Enabling Online Hybrid-Interactive Reliability AnalysisabstractEnsuring reliable operation of modern silicon systems in safety-critical domains requires fault injection (FI) platforms that simultaneously achieve accuracy, observability, and efficiency. Traditional simulation-based FI provides full observability but is prohibitively slow, while hardware-based FI improves speed but struggles to provide cycle-level precision, cross-domain support, and comprehensive monitoring. To address this, this work presents Gohan, a golden-copy-aided platform that enables online, hybrid-interactive reliability analysis across multi-clock-domain systems. To preserve cycle-accurate state transitions, it introduces a per-domain golden copy that is generated independently for each domain through simulation. In addition, an FPGA-based host–DUT co-execution loop is used, incorporating clock domain-crossing (CDC)-aware pause-resume mechanisms and scan-chain-based FI. Experimental results on both lightweight RISC-V cores and complex AI processor demonstrate that Gohan achieves 100% consistency with simulation models under repeated pause–resume operations and fault campaigns, while providing 3 orders-of-magnitude speedup over pure simulation. By bridging simulation accuracy and hardware realism, Gohan offers a scalable, low-cost, and high-fidelity solution for reliability evaluation at pre-silicon stage. Wang Liao 0001, Longyang Lin, Masanori Hashimoto |
DATE | 5 |
| 2026 | A 40-nm Resilient MLC RRAM Macro with Self-Referenced Time-Based Readout and 3-Bit Interleaved ECC Achieving 0.22 pJ/bit Read Energy
Zhen Kong, Yida Liang, Humiao Li, Yida Li 0004, Jiamin Li 0008, Longyang Lin |
ISCAS | 7 |
| 2026 | Mitigating Conductance Drift via In-Situ Calibration for Reliable RRAM-Based CIM Edge Inference
Zhen Kong, Weirong Dong, Zhengke Yang, Yida Liang, Jiamin Li 0008, Yida Li 0004, Longyang Lin |
ISCAS | 9 |
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 12 |
| 2026 | A 3-D Connectivity CMOS Ising Machine With 12-Way Toroidal Hexagonal Close-Packed Supply-and-Bulk Injection Locking Oscillators for Combinatorial OptimizationabstractFinding optimal solutions for Combinatorial Optimization (CO) problems is challenging. Compared to power-hungry cryogenic quantum computer and time-consuming classical computer, quantum-inspired Ising machine solves CO problems at room temperature with fast optimization speed, low power consumption, and low cost. Nevertheless, the Ising machine still faces several challenges: digital CMOS Ising machines increase interaction freedom at the cost of greater area and larger processing time; in analog Ising machine, oscillator spins find it hard to differentiate spin states without the assistance of the post-processing algorithm, and latch spins suffer from mismatches. To address these issues, we propose an oscillator-based 3-D CMOS Analog Ising Machine (CAIM) which adopts the 12-way toroidal Hexagonal Close Packed (HCP) structure, Supply-And-Bulk Injection Locking (SABIL), and dual-mode tunable coupler. The toroidal HCP structure exhibits >2 & times; interactions compared to conventional 3-D Ising machine, while SABIL and dual-mode tunable coupler settle oscillators to a bistable ground state 2.2 & times; quicker with an 8.5 & times; wider lock range (within 5 cycles). The proposed CAIM successfully solves 3-D max-cut problems and achieves a normalized Hamiltonian energy of more than 0.98 with a maximum perfect accuracy prevalence (PAP) of 91.67%. Measurement results on a sample random Max-Cut instance demonstrate that CAIM converges to within 1.7% of the reference optimum in 5 cycles. Jiaer Chen, Yingna Huang, Zhong-Qi Li, Han Wu 0003, Longyang Lin, Jiamin Li 0008, Jerald Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2026 | A Low-Power Speech-Based Depression Recognition Processor With Hierarchical Local-Global NetworkabstractDepression is a critical public health concern characterized by underdiagnosis, often due to stigma, lack of awareness, and reluctance to seek help. Cases of delayed intervention could be alleviated by wearable solutions which enable continuous and unobtrusive monitoring of depression indicators. Compared to electroencephalogram (EEG)-based and video-based depression recognition, speech-based approaches can be performed without deliberate user attention. However, due to the limited accuracy of existing algorithms and constrained resources at edge, performing accurate speech-based depression recognition on wearable platforms remains a challenge. Therefore, to achieve unobtrusive, accurate, and efficient depression recognition at edge, this work presents a hierarchical local–global network (HLG-Net) and optimized processor design for speech-based depression recognition. The proposed HLG-Net integrates convolutional neural networks (CNNs) with multihead attention (MHA) mechanism to simultaneously capture local acoustic features and global utterance-level coherence, enhancing depression stage recognition. For efficient processor design, a cross-layer buffered dataflow is proposed for efficient data handling, reducing data storage by 98.77%. The computing unit (CU) employs layer fusion, operator optimization, and quantization techniques to further improve resource utilization and reduce power consumption while preserving recognition accuracy. System-level low-power techniques such as clock/input gating and near-threshold design for application specific integrated circuit (ASIC) further reduce power consumption. The proposed processor implemented on field-programmable gate array (FPGA) (XC7Z100-2FFG900) achieves the lowest reported mean absolute error (MAE) of 5.13 on AVEC 2014 database. The 180-nm ASIC implementation shows a simulated power consumption of$17.4~\mu $W at 0.4 V. The results demonstrate the feasibility of accurate and efficient speech-based depression recognition on wearables. Yuxing Zhi, Weirong Dong, Huaijun Wang, Longyang Lin, Jiamin Li 0008 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2026 | TriCIM: A General CIM-Capacity-Aware Framework for Optimizing Model, Layer, and Tile-Stationary Dataflows in CIM AcceleratorsabstractCompute-in-memory (CIM) has emerged as a promising paradigm for accelerating neural networks (NNs), offering high parallelism and superior energy efficiency. To fully leverage the potential of CIM accelerators for deploying NNs, it is essential to carefully optimize their mapping dataflows, which determine how computations and data are scheduled across CIM resources over space and time. However, the diversity of CIM architectures, with their varying hardware parameters and mapping constraints, makes it challenging to develop a unified framework that efficiently optimizes dataflows across different CIM systems. Most existing CIM dataflow frameworks are tailored to a limited range of architectures or dataflow types, and often neglect CIM-specific features such as weight-update scheduling, which limits their effectiveness and generality. In this work, we propose TriCIM, a general dataflow optimization framework designed to adapt to diverse CIM architectures. We begin by introducing a CIM-capacity-aware formulation that classifies the CIM dataflow space into three distinct regions, model-stationary, layer-stationary, and tile-stationary, based on the available CIM capacity and the model size. For each dataflow region, we identify specific bottlenecks and propose tailored optimization strategies, including load balancing, layer grouping, weight-update scheduling, tiling, and inter-tile ordering, to address region-specific bottlenecks and enhance dataflow efficiency. Our evaluation across various CIM accelerators and representative NN models demonstrates that TriCIM consistently finds optimal dataflows, achieving$1.1\times $to$13.2\times $runtime speedup compared to state-of-the-art frameworks. Jin Wang 0044, Yufu Zhang, Longyang Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | HachiFI: A Lightweight SoC Architecture-Independent Fault-Injection Framework for SEU Impact EvaluationabstractSingle-Event Upsets (SEUs), triggered by energetic particles, manifest as unexpected bit-flips in memory cells or registers, potentially causing significant anomalies in electronic devices. Driven by the needs of safety-critical applications, it is crucial to evaluate the reliability of these electronic devices before they are deployed. However, traditional reliability analysis techniques, such as irradiation experiments, are costly, while fault injection (FI) simulations often fail to provide full coverage and have limited effectiveness and accuracy. To address these issues, we introduce HachiFI, a lightweight, architecture-independent framework that automates fault injection with 100% coverage via memory and scan-chain accesses and simulates the behavior of SEUs based on specific cross-sections. HachiFI supports configurable fault injection patterns for both system-level and module-level reliability analysis. Using HachiFI, we demonstrate a low hardware overhead (2=0.984) between FI and irradiation experiments, verified on a 22nm edge-AI chip. Wang Liao 0001, Hao Yu 0001, Longyang Lin, Masanori Hashimoto |
DATE | 5 |
| 2025 | Post-Layout Automated Optimization for Capacitor Array in Digital-to-Time ConverterabstractThe integral nonlinearity (INL) of Digital-to-Time Converter (DTC) in fractional-N phase-locked loops introduces fractional spurs, especially at near-integer channels, resulting in increased jitter. To meet the strict jitter and spur performance requirements of high-performance wireless transceivers, minimizing the INL in DTC designs is crucial. This work presents a computer-aided, automated optimization methodology that focuses on addressing issues stemming from the uniform capacitor unit structure within the capacitor array in Variable-Slope DTC. These issues include parasitic resistance and capacitance, which distort the charging and discharging behavior of the capacitors, contributing to INL. By systematically optimizing the capacitor layout and mitigating parasitic effects, the methodology allows precise tuning of each capacitor unit in capacitor array to reduce INL, enhancing the overall performance of the DTC. Hefei Wang, Jianghao Su, Junhe Xue, Haoran Lyu, Longyang Lin, Shenghua Zhou |
DATE | 6 |
| 2025 | Tenpura: A General Transient Fault Evaluation and Scope Narrowing Platform for Ultra-fast Reliability AnalysisabstractFor reliability-critical silicon systems, transient errors caused by cosmic rays necessitate comprehensive and efficient reliability analysis before product deployment. Fault injection (FI) serves as a cost-effective alternative to expensive irradiation experiments for evaluating system robustness. However, simulation-based FI is constrained by the performance of the underlying hardware platform, making it impractical for large-scale designs, where achieving high fault coverage can take months or even years. Furthermore, most transient errors have no impact on system functionality, and filtering out these insignificant errors in advance can significantly enhance the efficiency of reliability analysis. To address these challenges, we propose Tenpura, a fault evaluation platform designed for ultra-fast reliability analysis. In Tenpura, a transient fault scope narrowing method is introduced to narrow the FI scope via the proposed scan-based activity tracing flow, further optimizing fault analysis and improving overall efficiency. By leveraging FPGA emulation and scan chain-based fault analysis at the pre-silicon stage, Tenpura achieves high-efficiency fault reduction (88.49–96.26% across three design under tests (DUTs) including RISC-V cores and NVDLA-based AI accelerator) within one month, delivering over an order of magnitude faster fault analysis compared to SOTA methods. Huizi Zhang, Chien-Hsing Liang, Jing-Jia Liou, Jinjun Xiong, Longyang Lin, Masanori Hashimoto |
ICCAD | 7 |
| 2025 | CIMWise: An IREE-based End-To-End AI Compiler with Auto-Tuning for CIM ProcessorsabstractComputing-In-memory (CIM) processors have emerged as a promising approach for accelerating deep neural networks (DNNs). To unleash the potential of diverse CIM architectures, it is essential to develop a tailored compiler that is aware of both the architectural parameters of CIM processors (e.g., CIM array size, number of CIM units, on-chip buffer capacity) and their dataflow characteristics (e.g., CIM parallelism, dataflow scheduling, buffer allocation). However, existing CIM compilers primarily focus on hardware parameters, while neglecting the diversity of dataflow patterns across different architectures. In this paper, we propose CIMWise, a general, IREE-based, end-to-end AI compiler for CIM processors. At its core, an auto-tuning framework has been developed, comprising a comprehensive search space (including both hardware parameters and dataflow characteristics), an analytical cost model (integrating an adaptable on-chip memory model), and a two-stage simulated annealing search algorithm, to explore optimal scheduling strategies for the target workload on CIM processors. Additionally, CIMWise leverages IREE’s frontend for graph-level optimization and implements a customized backend to perform operator-level optimization using the proposed auto-tuning framework, enabling seamless end-to-end deployment of neural networks on CIM processors. The proposed CIMWise is validated through actual hardware measurements. Experimental results show that CIMWise achieves up to a 58% reduction in energy consumption and a 19% decrease in latency compared to prior CIM compilers. Bo Mai, Jin Wang 0044, Zhen Zhai, Yufu Zhang, Longyang Lin |
ICCAD | 6 |
| 2025 | A Scalable External Memory Access and On-Chip Storage Architecture for Edge-AI Accelerators : - Multi-Path Rolling Data Refresh and Layer-Wise Bank Allocation -abstractFor resource-constrained AI accelerators applied in edge computing, achieving high power efficiency in neural network (NN) model computation is crucial. However, current designs often overlook the efficiency of off-chip/on-chip data interaction, leading to high latency, which in turn results in suboptimal power efficiency during computation. Additionally, inefficient memory bank allocation further exacerbates latency by causing underutilization of storage resources, thereby contributing to higher overall latency and energy consumption. To address these challenges, this paper proposes a scalable multi-path rolling data refresh and layer-wise bank allocation architecture. The rolling data refresh mechanism enables efficient data interaction between off-chip and on-chip storage, reducing latency and minimizing the area overhead of on-chip memories. The layer-wise bank allocation optimizes on-chip memory utilization according to specific application requirements, improving memory efficiency. A case study on a 28nm AI accelerator demonstrates a 30.6% reduction in area, achieves a power efficiency of 7.36–10.28 TOPS/W, and reduces external memory access by 2.63% to 37.24% on VGG16 and ViT-Small. Huizi Zhang, Qiufeng Li, Yuan Liang 0004, Zhenzhe Chen, Jinjun Xiong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto |
ISLPED | 10 |
| 2025 | Genshin: A Generalized Framework with Software-Hardware Co-design and Pruned Fault Injection for Reliability AnalysisabstractReliability-demanding devices often require numerous fault injections (FIs) for reliability analysis in the product cycle. However, software-based FI typically demonstrates extremely low efficiency due to low simulation throughput, especially for large-scale designs, while hardware-based FI presents challenges related to complexity of setup and limited scalability. Additionally, FIs often occur in intervals where errors do not affect the system’s outcome, e.g., after final read before next write, necessitating efficient pruning of non-impactful FIs. To address this, a general-purpose FI-specialized framework, Genshin, is proposed for rapid reliability analysis. On the hardware side, we provide an FI-specialized design, which works with Design Under Test (DUT) chips on PCB boards and supports FI control based on the scan chain (SC). An integrated programmable logic allows for flexible and custom FI pattern definitions. Furthermore, an architecturally correct execution (ACE) analysis generates pruned fault tables for DUTs. In Genshin, the SC logic achieves 3,802-65,388 cycles/FI across SC lengths ranging from 2,795 to 61,393 in different DUTs, while the programmable logic enables custom error patterns such as layout-aware multi-bit upset (MBU). Furthermore, the pruned fault tables achieve fault reduction rates from 45.80% to 83.21%. Hao-Yang Chi, Chien-Hsing Liang, Yu-Hong Chao, Huizi Zhang, Yuan Liang 0004, Wang Liao 0001, Jinjun Xiong, Jing-Jia Liou, Masanori Hashimoto, Longyang Lin |
ITC | 12 |
| 2024 | How accurately can soft error impact be estimated in black-box/white-box cases? - a case study with an edge AI SoC -abstractArtificial intelligence (AI) edge devices often feature numerous storage units and sequential logic circuits, making them vulnerable to soft errors. For reliable and critical edge AI applications, assessing System-on-Chip (SoC) reliability in advance is essential. Here, there are two cases: a self-designed SoC (white-box), or a commercial off-the-shelf (COTS) chip (black-box). This study uses alpha particle irradiation results on our 22nm AI SoC as a golden reference to estimate soft error impacts, injecting faults across the entire chip in the white-box case and into the accessible memory and registers in the black-box case. The results demonstrate a high degree of consistency between the white-box case and golden reference, meaning that pre-silicon reliability assessment is feasible. As for the black-box case, the proportion of memory in the SoC remains unchanged and is still significantly larger than that of registers, and hence the simulation results between black-box and white-box are not substantially different. Qiufeng Li, Longyang Lin, Wang Liao 0001, Liuyao Dai, Hao Yu 0001, Masanori Hashimoto |
DAC | 3 |
| 2024 | S3M: Static Semi-Segmented Multipliers for Energy-Efficient DNN Inference AcceleratorsabstractApproximate multipliers offer an efficient approach to reduce power consumption in compute-intensive applications, such as Deep Neural Networks (DNNs). However, current 8-bit approximate multipliers struggle to maintain high accuracy across various DNN applications. In this paper, we highlight challenges in 8-bit multiplier designs with body approximation strategies and evaluate the effectiveness of input approximation methods. Recognizing that exact multipliers with quantization bit-widths below 8 bits have demonstrated superior performance, we aim to explore whether alternative input approximation methods can provide an even better tradeoff between accuracy and energy consumption. To this end, by exploiting the fact that weight operand values are smaller than activations and prepared offline in DNNs, we simplify a static segmented multiplier (SSM) into a static semi-segmented multiplier$(\mathbf{S}^{3}\mathbf{M})$, achieving a 31.58% reduction in power-delay product (PDP) compared to the original SSM, with similar classification accuracy. Additionally, we propose Coded$\mathbf{S}^{3}\mathbf{M}$with optimized memory usage and im-plement various multipliers on a systolic array-based accelerator. Experimental results show that the proposed$\mathbf{S}^{3}\mathbf{M}$and Coded$\mathbf{S}^{3}\mathbf{M}$outperform existing 8-bit approximate multipliers in DNN applications, effectively bridging the PDP and inference accuracy tradeoff observed across exact commercial IP multipliers of varied bit-widths without requiring time-consuming retraining. Consequently, the proposed multiplier designs provide enhanced computational solutions for energy-efficient DNN inference ac-celerators. Hiromitsu Awano, Longyang Lin, Masanori Hashimoto |
ICCD | 4 |
| 2023 | NUTS-BSNN: A non-uniform time-step binarized spiking neural network with energy-efficient in-memory computing macro
Van-Ngoc Dinh, Ngoc-My Bui, Chacko John Deepu, Longyang Lin, Quang-Kien Trinh |
Neurocomputing | 5 |
| 2022 | A Configurable Floating-Point Multiple-Precision Processing Element for HPC and AI Converged ComputingabstractThere is an emerging need to design configurable accelerators for the high-performance computing (HPC) and artificial intelligence (AI) applications in different precisions. Thus, the floating-point (FP) processing element (PE), which is the key basic unit of the accelerators, is necessary to meet multiple-precision requirements with energy-efficient operations. However, the existing structures by using high-precision-split (HPS) and low-precision-combination (LPC) methods result in low utilization rate of the multiplication array and long multiterm processing period, respectively. In this article, a configurable FP multiple-precision PE design is proposed with the LPC structure. Half precision, single precision, and double precision are supported. The 100% multiplier utilization rate of the multiplication array for all precisions is achieved with improved speed in the comparison and summation process. The proposed design is realized in a 28-nm process with 1.429-GHz clock frequency. Compared with the existing multiple-precision FP methods, the proposed structure achieves 63% and 88% area-saving performance for FP16 and FP32 operations, respectively. The$4\times $and$20\times $maximum throughput rates are obtained when compared with fixed FP32 and FP64 operations. Compared with the previous multiple-precision PEs, the proposed one achieves the best energy-efficiency performance with 975.13 GFLOPS/W. Wei Mao 0002, Kai Li 0024, Liuyao Dai, Xinang Xie, He Li 0008, Longyang Lin, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2020 | Automated Design of Reconfigurable Microarchitectures for Accelerators under Wide-Voltage ScalingabstractThis article introduces a systematic methodology to design microarchitectures that are reconfigurable down to the pipeline stage. Reconfigurable microarchitectures were showed to provide significant energy improvements in accelerators under wide-voltage scaling. However, prior art is based on ad hoc techniques that limit their applicability, without addressing the challenge of enabling general design flows for reconfigurable microarchitectures. The proposed methodology introduces the unprecedented capability of translating a conventional fixed microarchitecture into a reconfigurable one. The methodology relies on commercial EDA tools, which are integrated into a design flow through the manipulation of the gate-level netlist via a set of graph algorithms. The proposed methodology is shown to be architecture-agnostic, fully automated, and applicable to designs that are either developed at the register transfer level (RTL), or provided by third-party soft IP vendors. Ultimately, the proposed methodology allows to add microarchitectural adjustment as a run-time knob to augment the energy benefits of wide-voltage scaling. Reconfiguration is shown to improve the energy efficiency by up to 35% beyond the conventional dynamic voltage frequency scaling (DVFS), through the analysis of various test vehicles. Longyang Lin, Massimo Alioto |
ISCAS | 2 |
| 2020 | Automated Design of Reconfigurable Microarchitectures for Accelerators Under Wide-Voltage ScalingabstractThis article introduces a systematic methodology to design microarchitectures that are reconfigurable down to the pipeline stage. Reconfigurable microarchitectures were showed to provide significant energy improvements in accelerators under wide-voltage scaling. However, prior art is based on ad hoc techniques that limit their applicability, without addressing the challenge of enabling general design flows for reconfigurable microarchitectures. The proposed methodology introduces the unprecedented capability of translating a conventional fixed microarchitecture into a reconfigurable one. The methodology relies on commercial EDA tools, which are integrated into a design flow through the manipulation of the gate-level netlist via a set of graph algorithms. The proposed methodology is shown to be architecture-agnostic, fully automated, and applicable to designs that are either developed at the register transfer level (RTL), or provided by third-party soft IP vendors. Ultimately, the proposed methodology allows to add microarchitectural adjustment as a run-time knob to augment the energy benefits of wide-voltage scaling. Reconfiguration is shown to improve the energy efficiency by up to 35% beyond the conventional dynamic voltage frequency scaling (DVFS), through the analysis of various test vehicles. Longyang Lin, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Transistor sizing strategy for simultaneous energy-delay optimization in CMOS buffersabstractIn this work, a systematic transistor sizing strategy is proposed to meet an arbitrary energy-delay target in CMOS buffers, as defined by the considered applications. This is particularly important in VLSI systems, as buffers driving large capacitive loads consume a very large energy compared to other logic gates. To this aim, an analytical and technology-independent model was first developed to find optimal circuit design parameters (e.g., sizing, number of stages). To assure true optimality, general Variable-stage effort Tapered Buffers (VTB) are considered, as opposed to Fixed-stage effort Tapered Buffers (FTB). Results show that optimized VTBs reduce energy by as much as 30%, compared to FTBs. Under balanced energy and delay, VTBs achieve 10-20% energy saving with nearly the same performance as FTBs. The adopted models and design guidelines are shown to agree well with circuit simulations in 28 and 65nm CMOS across the voltage range from 0.6 V to 1 V. This design strategy is a useful tool for circuit designers to systematically manage the energy-delay tradeoff of CMOS buffers in a simple and technology-agnostic manner. Longyang Lin, Quang-Kien Trinh, Massimo Alioto |
ISCAS | 1 |