EDBT 2026 Demo / reviewers in the wild / expert
Masanori Hashimoto
dblp:81/317
· DBLP profile ↗
115ranked-venue papers
17as first author
26since 2021 · last 2026
0000-0002-0377-2108ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 110 · 17 first-author · 25 since 2021Software engineering, systems software and programming languages · 12 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gundam: A Generalized Unified Design and Analysis Model for Matrix Multiplication on Edge
Weirong Dong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto |
ASP-DAC | 6 |
| 2026 | FIawase: A SET Fault Injection Framework Towards Exhaustive System-Level Impact EvaluationabstractSingle-event transients (SETs) threaten modern reliability-demanding SoCs equipped with error correction codes (ECC) for single event upset (SEU) mitigation. However, conventional gate-level SET fault injection (FI) remains prohibitively slow for practical reliability evaluation. This work presents Flawase, a high-throughput SET injection framework that enables comprehensive system-level evaluation of SET-induced soft errors. Flawase consists of two phases: a netlist-level inject-and-capture simulation, which systematically flips the output of every gate-cycle pair during program execution to record flipflop changes one cycle later, and a scan-chain-based replay-andmeasure emulation, which replays these patterns at hardware speed to quantify system-level impact. Implemented on an opensource RISC-V system, FIawase reduces a comprehensive SET injection campaign from decades of pure simulation to nearly a day, achieving over four orders of magnitude end-to-end speedup. Flawase takes a critical first step toward exhaustive, cycle-accurate SET analysis, enabling architectural and reliability research at previously infeasible scales. Masanori Hashimoto |
ASP-DAC | 3 |
| 2026 | Ramen: Radiation-Aware Modeling Framework for PDK-Enabled Design and Library CharacterizationabstractRadiation-induced degradation poses a critical challenge to the reliability of space-grade integrated circuits (ICs). Existing radiation-aware models largely remain at the device level and lack direct integration with circuit or system design flows, limiting their practical use in radiation-aware IC design. To address this, this work proposes Ramen, a non-invasive radiation-aware device modeling framework that is fully compatible with commercial Process Design Kits (PDKs). Ramen accurately captures total ionizing dose (TID) and displacement damage dose (DDD), enabling early-stage evaluation at both circuit and system levels without requiring modifications to existing PDK structures. By seamlessly integrating with standard analog, mixed-signal, and digital flows, the radiation-aware models not only support SPICE-based circuit simulation but also feed into standard library characterization tools to generate radiation-aware Liberty libraries. These libraries encode dose-dependent timing, leakage, and power information, allowing radiation effects to be captured in synthesis, timing analysis, and back-end implementation. Experimental validation on a 180 nm CMOS imager under radiation stress shows that the proposed framework achieves <15% simulation errors for both analog and logic circuit, confirming the reliability of Ramen for radiation-aware IC design. Zhenzhe Chen, Wang Liao 0001, Jing-Jia Liou, Masanori Hashimoto, Longyang Lin |
DATE | 6 |
| 2026 | Gohan: A Golden-Copy-Aided Platform Enabling Online Hybrid-Interactive Reliability AnalysisabstractEnsuring reliable operation of modern silicon systems in safety-critical domains requires fault injection (FI) platforms that simultaneously achieve accuracy, observability, and efficiency. Traditional simulation-based FI provides full observability but is prohibitively slow, while hardware-based FI improves speed but struggles to provide cycle-level precision, cross-domain support, and comprehensive monitoring. To address this, this work presents Gohan, a golden-copy-aided platform that enables online, hybrid-interactive reliability analysis across multi-clock-domain systems. To preserve cycle-accurate state transitions, it introduces a per-domain golden copy that is generated independently for each domain through simulation. In addition, an FPGA-based host–DUT co-execution loop is used, incorporating clock domain-crossing (CDC)-aware pause-resume mechanisms and scan-chain-based FI. Experimental results on both lightweight RISC-V cores and complex AI processor demonstrate that Gohan achieves 100% consistency with simulation models under repeated pause–resume operations and fault campaigns, while providing 3 orders-of-magnitude speedup over pure simulation. By bridging simulation accuracy and hardware realism, Gohan offers a scalable, low-cost, and high-fidelity solution for reliability evaluation at pre-silicon stage. Wang Liao 0001, Longyang Lin, Masanori Hashimoto |
DATE | 6 |
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 13 |
| 2025 | ML-assisted SRAM Soft Error Rate Characterization: Opportunities and ChallengesabstractSoft errors from cosmic rays are a significant concern for reliability-critical applications such as autonomous driving and supercomputers. In this paper, we review soft error rate (SER) estimation for SRAM, the most sensitive component in digital logic chips, and explore how machine learning can assist in SRAM SER characterization. We propose an efficient discriminator construction method for single-event upset (SEU) using active learning and adaptive hyperparameter tuning in the learning algorithm. This method iteratively labels samples through technology computer-aided design (TCAD) simulations, determining whether an upset occurs for an unlabeled sample with the lowest confidence in prediction. Our approach eliminates the need for empirical modeling based on tacit knowledge, systematically building a model while reducing the training data needed to achieve sufficient event-wise accuracy. Experiments with a 12-nm SRAM show that the training data required to achieve the same accuracy was reduced by 41% for 80% accuracy and by 31% for 85% accuracy. Finally, we discuss future directions and challenges in advanced nano-sheet and CFET transistors. Masanori Hashimoto, Ryuichi Yasuda, Kazusa Takami, Yuibi Gomi, Kozo Takeuchi |
ASP-DAC | 1 |
| 2025 | Hardware Error Detection with In-Situ Monitoring of Control Flow-Related SpecificationsabstractIn hardware accelerators used in data centers and safety-critical applications, soft errors and resultant silent data corruption significantly compromise reliability, particularly when upsets occur in control-flow operations, leading to severe failures. To address this, we introduce a method for monitoring control flow-related specifications using Petri nets. We validated our method across three designs: convolutional layers in LeNet-5, Gaussian blur in Canny edge detection, and AES encryption. Our fault injection campaign targeting the control registers and primary control inputs demonstrated high error detection rates in both datapath and control logic. Synthesis results show that a maximum detection rate is achieved with a few to around 10 % area overhead in most cases. The proposed detectors quickly detect 88.0% to 99.9% of failures resulting from upsets in internal control registers and perturbation in primary control inputs. Tomonari Tanaka, Takumi Uezono, Kohei Suenaga, Masanori Hashimoto |
ASP-DAC | 4 |
| 2025 | HachiFI: A Lightweight SoC Architecture-Independent Fault-Injection Framework for SEU Impact EvaluationabstractSingle-Event Upsets (SEUs), triggered by energetic particles, manifest as unexpected bit-flips in memory cells or registers, potentially causing significant anomalies in electronic devices. Driven by the needs of safety-critical applications, it is crucial to evaluate the reliability of these electronic devices before they are deployed. However, traditional reliability analysis techniques, such as irradiation experiments, are costly, while fault injection (FI) simulations often fail to provide full coverage and have limited effectiveness and accuracy. To address these issues, we introduce HachiFI, a lightweight, architecture-independent framework that automates fault injection with 100% coverage via memory and scan-chain accesses and simulates the behavior of SEUs based on specific cross-sections. HachiFI supports configurable fault injection patterns for both system-level and module-level reliability analysis. Using HachiFI, we demonstrate a low hardware overhead (2=0.984) between FI and irradiation experiments, verified on a 22nm edge-AI chip. Wang Liao 0001, Hao Yu 0001, Longyang Lin, Masanori Hashimoto |
DATE | 6 |
| 2025 | Tenpura: A General Transient Fault Evaluation and Scope Narrowing Platform for Ultra-fast Reliability AnalysisabstractFor reliability-critical silicon systems, transient errors caused by cosmic rays necessitate comprehensive and efficient reliability analysis before product deployment. Fault injection (FI) serves as a cost-effective alternative to expensive irradiation experiments for evaluating system robustness. However, simulation-based FI is constrained by the performance of the underlying hardware platform, making it impractical for large-scale designs, where achieving high fault coverage can take months or even years. Furthermore, most transient errors have no impact on system functionality, and filtering out these insignificant errors in advance can significantly enhance the efficiency of reliability analysis. To address these challenges, we propose Tenpura, a fault evaluation platform designed for ultra-fast reliability analysis. In Tenpura, a transient fault scope narrowing method is introduced to narrow the FI scope via the proposed scan-based activity tracing flow, further optimizing fault analysis and improving overall efficiency. By leveraging FPGA emulation and scan chain-based fault analysis at the pre-silicon stage, Tenpura achieves high-efficiency fault reduction (88.49–96.26% across three design under tests (DUTs) including RISC-V cores and NVDLA-based AI accelerator) within one month, delivering over an order of magnitude faster fault analysis compared to SOTA methods. Huizi Zhang, Chien-Hsing Liang, Jing-Jia Liou, Jinjun Xiong, Longyang Lin, Masanori Hashimoto |
ICCAD | 8 |
| 2025 | A Scalable External Memory Access and On-Chip Storage Architecture for Edge-AI Accelerators : - Multi-Path Rolling Data Refresh and Layer-Wise Bank Allocation -abstractFor resource-constrained AI accelerators applied in edge computing, achieving high power efficiency in neural network (NN) model computation is crucial. However, current designs often overlook the efficiency of off-chip/on-chip data interaction, leading to high latency, which in turn results in suboptimal power efficiency during computation. Additionally, inefficient memory bank allocation further exacerbates latency by causing underutilization of storage resources, thereby contributing to higher overall latency and energy consumption. To address these challenges, this paper proposes a scalable multi-path rolling data refresh and layer-wise bank allocation architecture. The rolling data refresh mechanism enables efficient data interaction between off-chip and on-chip storage, reducing latency and minimizing the area overhead of on-chip memories. The layer-wise bank allocation optimizes on-chip memory utilization according to specific application requirements, improving memory efficiency. A case study on a 28nm AI accelerator demonstrates a 30.6% reduction in area, achieves a power efficiency of 7.36–10.28 TOPS/W, and reduces external memory access by 2.63% to 37.24% on VGG16 and ViT-Small. Huizi Zhang, Qiufeng Li, Yuan Liang 0004, Zhenzhe Chen, Jinjun Xiong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto |
ISLPED | 11 |
| 2025 | Genshin: A Generalized Framework with Software-Hardware Co-design and Pruned Fault Injection for Reliability AnalysisabstractReliability-demanding devices often require numerous fault injections (FIs) for reliability analysis in the product cycle. However, software-based FI typically demonstrates extremely low efficiency due to low simulation throughput, especially for large-scale designs, while hardware-based FI presents challenges related to complexity of setup and limited scalability. Additionally, FIs often occur in intervals where errors do not affect the system’s outcome, e.g., after final read before next write, necessitating efficient pruning of non-impactful FIs. To address this, a general-purpose FI-specialized framework, Genshin, is proposed for rapid reliability analysis. On the hardware side, we provide an FI-specialized design, which works with Design Under Test (DUT) chips on PCB boards and supports FI control based on the scan chain (SC). An integrated programmable logic allows for flexible and custom FI pattern definitions. Furthermore, an architecturally correct execution (ACE) analysis generates pruned fault tables for DUTs. In Genshin, the SC logic achieves 3,802-65,388 cycles/FI across SC lengths ranging from 2,795 to 61,393 in different DUTs, while the programmable logic enables custom error patterns such as layout-aware multi-bit upset (MBU). Furthermore, the pruned fault tables achieve fault reduction rates from 45.80% to 83.21%. Hao-Yang Chi, Chien-Hsing Liang, Yu-Hong Chao, Huizi Zhang, Yuan Liang 0004, Wang Liao 0001, Jinjun Xiong, Jing-Jia Liou, Masanori Hashimoto, Longyang Lin |
ITC | 11 |
| 2025 | An 88.5 fsrms Integrated Jitter and -76.2 dBc Reference Spur mmW PLL Utilizing a Ripple Compensation Phase/Frequency DetectorabstractMillimeter-wave (mmW) phase-locked loops (PLLs) typically favor a wide loop bandwidth for stronger suppression of the out-of-band phase noise from a voltage-controlled oscillator (VCO). Unfortunately, doing so lowers the degree of attenuation to the PLL reference spurs. This paper proposes a ripple compensation phase detector (RCPD) for extending PLL loop bandwidth and phase noise suppression without sacrificing reference spur performance. The RCPD inherently consists of a pair of PDs that generate respective ripple simultaneously, with each PD’s ripple current compensating the other, resulting in a glitch-free RCPD output. A calibrator is also introduced to reduce device mismatches. With the proposed techniques, the proposed mmW PLL was implemented using 22 nm bulk CMOS technology. The mmW PLL operates from 32.7 to 39.4 GHz, achieving an integrated jitter and reference spur of 88.5 fsrms (1 kHz to 100 MHz) and –76.2 dBc, respectively, with a figure-of-merit (FoM) of –247.5 dB. Yuan Liang 0004, Zhongyuan Fang, Masanori Hashimoto |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Logic Locking over TFHE for Securing User Data and AlgorithmsabstractThis paper proposes the application of logic locking over TFHE to protect both user data and algorithms, such as input user data and models in machine learning inference applications. With the proposed secure computation protocol algorithm evaluation can be performed distributively on honest-but-curious user computers while keeping the algorithm secure. To achieve this, we combine conventional logic locking for untrusted foundries with TFHE to enable secure computation. By encrypting the logic locking key using TFHE, the key is secured with the degree of TFHE. We implemented the proposed secure protocols for combinational logic neural networks and decision trees using LUT-based obfuscation. Regarding the security analysis, we subjected them to the SAT attack and evaluated their resistance based on the execution time. We successfully configured the proposed secure protocol to be resistant to the SAT attack in all machine learning benchmarks. Also, the experimental result shows that the proposed secure computation involved almost no TFHE runtime overhead in a test case with thousands of gates. Kohei Suemitsu, Kotaro Matsuoka, Takashi Sato 0001, Masanori Hashimoto |
ASPDAC | 4 |
| 2024 | How accurately can soft error impact be estimated in black-box/white-box cases? - a case study with an edge AI SoC -abstractArtificial intelligence (AI) edge devices often feature numerous storage units and sequential logic circuits, making them vulnerable to soft errors. For reliable and critical edge AI applications, assessing System-on-Chip (SoC) reliability in advance is essential. Here, there are two cases: a self-designed SoC (white-box), or a commercial off-the-shelf (COTS) chip (black-box). This study uses alpha particle irradiation results on our 22nm AI SoC as a golden reference to estimate soft error impacts, injecting faults across the entire chip in the white-box case and into the accessible memory and registers in the black-box case. The results demonstrate a high degree of consistency between the white-box case and golden reference, meaning that pre-silicon reliability assessment is feasible. As for the black-box case, the proportion of memory in the SoC remains unchanged and is still significantly larger than that of registers, and hence the simulation results between black-box and white-box are not substantially different. Qiufeng Li, Longyang Lin, Wang Liao 0001, Liuyao Dai, Hao Yu 0001, Masanori Hashimoto |
DAC | 7 |
| 2024 | S3M: Static Semi-Segmented Multipliers for Energy-Efficient DNN Inference AcceleratorsabstractApproximate multipliers offer an efficient approach to reduce power consumption in compute-intensive applications, such as Deep Neural Networks (DNNs). However, current 8-bit approximate multipliers struggle to maintain high accuracy across various DNN applications. In this paper, we highlight challenges in 8-bit multiplier designs with body approximation strategies and evaluate the effectiveness of input approximation methods. Recognizing that exact multipliers with quantization bit-widths below 8 bits have demonstrated superior performance, we aim to explore whether alternative input approximation methods can provide an even better tradeoff between accuracy and energy consumption. To this end, by exploiting the fact that weight operand values are smaller than activations and prepared offline in DNNs, we simplify a static segmented multiplier (SSM) into a static semi-segmented multiplier$(\mathbf{S}^{3}\mathbf{M})$, achieving a 31.58% reduction in power-delay product (PDP) compared to the original SSM, with similar classification accuracy. Additionally, we propose Coded$\mathbf{S}^{3}\mathbf{M}$with optimized memory usage and im-plement various multipliers on a systolic array-based accelerator. Experimental results show that the proposed$\mathbf{S}^{3}\mathbf{M}$and Coded$\mathbf{S}^{3}\mathbf{M}$outperform existing 8-bit approximate multipliers in DNN applications, effectively bridging the PDP and inference accuracy tradeoff observed across exact commercial IP multipliers of varied bit-widths without requiring time-consuming retraining. Consequently, the proposed multiplier designs provide enhanced computational solutions for energy-efficient DNN inference ac-celerators. Hiromitsu Awano, Longyang Lin, Masanori Hashimoto |
ICCD | 5 |
| 2023 | Avoiding Soft Error-Induced Illegal Memory Accesses in GPU with Inter-Thread CommunicationabstractA soft error caused by terrestrial neutrons poses a threat to the reliability of safety-critical systems, such as self-driving applications. These applications, often comprised of neural networks, rely on graphic processing units (GPUs) due to their requirement for massive parallel computation. While neural networks inherently include redundant computation and possess a certain level of error tolerance, detectable unrecoverable errors (DUEs) can be more detrimental than silent data corruption (SDC), as they can result in temporary service unavailability. This study specifically focuses on addressing illegal memory access, a primary cause of DUEs, and proposes a programming method that can detect illegal addresses. In the single instruction, multiple threads (SIMT) scheme, the data address is regularly calculated based on the thread ID, and this regularity is exploited to identify illegal addresses through inter-thread communication. To evaluate the effectiveness of the proposed method, fault injection campaigns were conducted for matrix multiplication, vector addition, and transposition. The experimental results indicate that the proposed method resulted in a reduction of the DUE rate by 17.3%, 86.8%, and 87.1% for these respective operations. Riku Iwamoto, Masanori Hashimoto |
IOLTS | 2 |
| 2023 | B2N2: Resource efficient Bayesian neural network accelerator using Bernoulli sampler on FPGAabstractA resource efficient hardware accelerator for Bayesian neural network (BNN) named B 2 N 2 , B ernoulli random number based B ayesian n eural network accelerator, is proposed. As neural networks expand their application into risk sensitive domains where mispredictions may cause serious social and economic losses, evaluating the NN’s confidence on its prediction has emerged as a critical concern. Among many uncertainty evaluation methods, BNN provides a theoretically grounded way to evaluate the uncertainty of NN’s output by treating network parameters as random variables . By exploiting the central limit theorem , we propose to replace costly Gaussian random number generators (RNG) with Bernoulli RNG which can be efficiently implemented on hardware since the possible outcome from Bernoulli distribution is binary. We demonstrate that B 2 N 2 implemented on Xilinx ZCU104 FPGA board consumes only 465 DSPs and 81661 LUTs which corresponds to 50.9% and 14.3% reductions compared to Gaussian-BNN (Hirayama et al., 2020) implemented on the same FPGA board for fair comparison. We further compare B 2 N 2 with VIBNN (Cai et al., 2018), which shows that B 2 N 2 successfully reduced DSPs and LUTs usages by 50.9% and 57.9%, respectively. Owing to the reduced hardware resources, B 2 N 2 improved energy efficiency by 7.50% and 57.5% compared to Gaussian-BNN (Hirayama et al., 2020) and VIBNN (Cai et al., 2018), respectively. Hiromitsu Awano, Masanori Hashimoto |
Integr. | 2 |
| 2023 | Reliability Exploration of System-on-Chip With Multi-Bit-Width Accelerator for Multi-Precision Deep Neural NetworksabstractDeep neural networks (DNNs) in safety-critical applications demand high reliability even when running on edge-computing devices. Recent works on System-on-Chip (SoC) design with state-of-the-art (SOTA) hardware artificial intelligence (AI) accelerators and corresponding multi-bit-width (MBW) convolutional neural network (CNN) generation strategies show that MBW CNNs can effectively explore the trade-off between network accuracy and hardware efficiency. However, reliability has not been considered in such trade-off analysis, even though highly quantized CNNs may elevate the impact of bit flips in the hardware. Also, the reliability of the microcontroller and its interface operating with the AI accelerator are not studied. This work evaluates the reliability of DNN computation in an SoC that includes a processor, SOTA AI accelerator, and NN models highly optimized for computation efficiency using a neural architecture search (NAS) method. Focusing on neutron-induced soft error, which is the primary source of bit-flip errors in a terrestrial environment, we perform fault injection and neutron beam experiments. For these experiments, we prototype the SoC on a flash-based FPGA platform, in which the configuration memory is robust to neutron irradiation. Then, we analyze the experimental data and identify vulnerable components in the system. Furthermore, we evaluate how the SoC running different NAS-optimized MBW LeNet5 networks impact the performance, radiation sensitivity, failure rate of MBW accelerator, and crash rate of the system on the FPGAs. Our results show that instruction and data tightly coupled memory (I/DTCM) are the most vulnerable parts and the control status registers (CSRs) in our accelerator are the second most vulnerable component. Moreover, MBW networks have higher susceptibility to critical errors than single-precision networks, low-precision data are more likely to affect the classification results, and the high bits are more sensitive to faults. Mingqiang Huang, Changhai Man, Liuyao Dai, Hao Yu 0001, Masanori Hashimoto |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | Estimating Vulnerability of All Model Parameters in DNN with a Small Number of Fault InjectionsabstractThe reliability of deep neural networks (DNNs) against hardware errors is essential as DNNs are increasingly employed in safety-critical applications such as automatic driving. Transient errors in memory, such as radiation-induced soft error, may propagate through the inference computation, resulting in unexpected output, which can adversely trigger catastrophic system failures. As a first step to tackle this problem, this paper proposes constructing a vulnerability model (VM) with a small number of fault injections to identify vulnerable model parameters in DNN. We reduce the number of bit locations for fault injection significantly and develop a flow to incrementally collect the training data, i.e., the fault injection results, for VM accuracy improvement. Experimental results show that VM can estimate vulnerabilities of all DNN model parameters only with 1/3490 computations compared with traditional fault injection-based vulnerability estimation. Yangchao Zhang, Hiroaki Itsuji, Takumi Uezono, Tadanobu Toba, Masanori Hashimoto |
DATE | 5 |
| 2022 | VirtualSync+: Timing Optimization With Virtual SynchronizationabstractIn digital circuit designs, sequential components such as flip-flops are used to synchronize signal propagations. Logic computations are aligned at and thus isolated by flip-flop stages. Although this fully synchronous style can reduce design efforts significantly, it may affect circuit performance negatively, because sequential components can only introduce delays into signal propagations but never accelerate them. In this article, we propose a new timing model, VirtualSync+, in which signals, specially those along critical paths, are allowed to propagate through several sequential stages without flip-flops. Timing constraints are still satisfied at the boundary of the optimized circuit to maintain a consistent interface with existing designs. By removing clock-to-q delays and setup time requirements of flip-flops on critical paths, the performance of a circuit can be pushed even beyond the limit of traditional sequential designs. In addition, we further enhance the optimization with VirtualSync+ by fine-tuning with commercial design tools, e.g., design compiler from Synopsys, to achieve more accurate result. To achieve this fine-tuning, we first optimize the circuits by reallocating sequential components with sequential and combinational components as delay units. Afterward, the removal locations of flip-flops with respect to the circuits under optimization are extracted and the corresponding wave-pipelining timing constraints compatible with commercial design tools are established. These timing constraints are then incorporated into the optimization flow of commercial tools to generate the optimized circuits. The experimental results demonstrate that circuit performance can be improved by up to 4% (average 1.5%) compared with that after extreme retiming and sizing, while the increase of area is still negligible. This timing performance is enhanced beyond the limit of traditional sequential designs. It also demonstrates that compared with those after retiming and sizing, the circuits with VirtualSync+ can achieve better timing performance under the same area cost or smaller area cost under the same clock period, respectively. Grace Li Zhang, Bing Li 0005, Xing Huang 0001, Xunzhao Yin, Cheng Zhuo, Masanori Hashimoto, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Mode-wise Voltage-scalable Design with Activation-aware Slack Assignment for Energy MinimizationabstractThis paper proposes a design optimization methodology that can achieve a mode-wise voltage scalable (MWVS) design with applying the activation-aware slack assignment (ASA). Originally, ASA allocates the timing margin of critical paths with a stochastic treatment of timing errors, which limits its application. Instead, this work employs ASA with guaranteeing no timing errors. The MWVS design is formulated as an optimization problem that minimizes the overall power consumption considering each mode duration, achievable voltage reduction, and accompanied circuit overhead explicitly, and explores the solution space with the downhill simplex algorithm that does not require numerical derivation. For obtaining a solution, i.e., a design, in the optimization process, we exploit the multi-corner multi-mode design flow in a commercial tool for performing mode-wise ASA with sets of false paths dedicated to individual modes. Experimental results based on RISC-V design show that the proposed methodology saves 20% more power compared to the conventional voltage scaling approach and attains 15% gain from the single-mode ASA. Also, the cycle-by-cycle fine-grained false path identification reduced leakage power by 42%. TaiYu Cheng, Yukata Masuda, Jun Nagayama, Yoichi Momiyama, Masanori Hashimoto |
ASP-DAC | 6 |
| 2021 | Hidden-Fold Networks: Random Recurrent Residuals Using Sparse Supermasks
Ángel López García-Arias, Masanori Hashimoto, Masato Motomura, Jaehoon Yu |
BMVC | 2 |
| 2021 | MUX Granularity Oriented Iterative Technology Mapping for Implementing Compute-Intensive Applications on Via-Switch FPGAabstractThis paper proposes a technology mapping algorithm for implementing application circuits on via-switch FPGA (VS-FPGA). The via-switch is a novel non-volatile and rewritable memory element. Its small footprint and low parasitic RC are expected to improve the area- and energy-efficiency of an FPGA system. Some unique features of the VS-FPGA require a dedicated technology mapping strategy for implementing application circuits with maximum energy-efficiency. One of the features is the small ratio of logic blocks to arithmetic blocks (ABs). Given an application circuit, the proposed algorithm first detects word-wise circuit elements, such as MUXs. These elements are evaluated with an index of how resource utilization and fan-out change when the corresponding element is implemented with AB. All these elements are sorted in descending order based on this index. According to this order, each element is mapped to AB one by one, and synthesis and evaluation are repeated iteratively until satisfying given design constraints. The experimental results show that resource utilization and maximum fan-out can be reduced by about 30 % to 50 % and 12 % to 87 %, respectively. The proposed algorithm is not limited to the VS-FPGA and is expected to improve computation density and energy-efficiency of various FPGAs dedicated to compute-intensive signal processing applications. Takashi Imagawa, Jaehoon Yu, Masanori Hashimoto, Hiroyuki Ochi |
DATE | 3 |
| 2021 | BloomCA: A Memory Efficient Reservoir Computing Hardware Implementation Using Cellular Automata and Ensemble Bloom FilterabstractIn this work, we propose a BloomCA which utilizes cellular automata (CA) and ensemble Bloom filter to organize an RC system by using only binary operations, which is suitable for hardware implementation. The rich pattern dynamics created by CA can map the input into high-dimensional space and provide more features for the classifier. Utilizing the ensemble Bloom filter as the classifier, the features can be memorized effectively. Our experiment reveals that applying the ensemble mechanism to Bloom filter endues a significant reduction in inference memory cost. Comparing with the state-of-the-art reference, the BloomCA achieves a 43× reduction for memory cost without hurting the accuracy. Our hardware implementation also demonstrates that BloomCA achieves over 21× and 43.64% reduction in area and power, respectively. Dehua Liang, Masanori Hashimoto, Hiromitsu Awano |
DATE | 2 |
| 2021 | Critical Path Isolation and Bit-Width Scaling Are Highly Compatible for Voltage Over-Scalable DesignabstractThis work proposes a design methodology that saves the power under voltage over-scaling (VOS) operation. The key idea of the proposed design methodology is to combine critical path isolation (CPI) and bit-width scaling (BWS) under the constraint of computational quality, e.g., Peak Signal-to-Noise Ratio (PSNR). Conventional CPI inherently cannot reduce the delay of intrinsic critical paths (CPs), which may significantly restrict the power saving effect. On the other hand, the proposed methodology tries to reduce both intrinsic and non-intrinsic CPs. Therefore, our design dramatically reduces the supply voltage and power dissipation while satisfying the quality constraint. Moreover, for reducing co-design exploration space, the proposed methodology utilizes the exclusiveness of the paths targeted by CPI and BWS, where CPI aims at reducing the minimum supply voltage of non-intrinsic CP, and BWS focuses on intrinsic CPs in arithmetic units. From this key exclusiveness, the proposed design splits the simultaneous optimization problem into three sub-problems; (1) the determination of bit-width reduction, (2) the timing optimization for non-intrinsic CPs, and (3) investigating the minimum supply voltage of the BWS and CPI-applied circuit under quality constraint, for reducing power dissipation. Thanks to the problem splitting, the proposed methodology can efficiently find quality-constrained minimum-power design. Evaluation results show that CPI and BWS are highly compatible, and they significantly enhance the efficacy of VOS. In a case study of GPGPU processor, the proposed design saves the power dissipation by 42.7% for an image processing and by 51.2% for a neural network inference workload. Yutaka Masuda, Jun Nagayama, TaiYu Cheng, Tohru Ishihara, Yoichi Momiyama, Masanori Hashimoto |
DATE | 6 |
| 2021 | Minimizing Energy of DNN Training with Adaptive Bit-Width and Voltage ScalingabstractTraining DNN mostly relies on GPUs with FP32 format. While FP16 is acknowledged for its advantage of high computation and memory efficiencies for training DNN, the training must be accompanied with techniques dedicated for a particular dataset. Therefore, a hardware engine with a configurable bit-width feature is desirable for covering any datasets and applications. This work proposes an adaptive bit-width and voltage scaling (ABVS) scheme for DNN training. The key idea is to increase fraction bit-width (FB) gradually from a small value according to current training quality (e.g., accuracy, mAP). Since less FB achieves shorter hardware latency, this training scheme concurrently adapts bit-width and voltage scaling and intensify energy reduction. Experimental results show that the ABVS scheme achieves the comparable quality to FP32 with at most 0.5% accuracy drop, but up to 63% energy reduction. TaiYu Cheng, Masanori Hashimoto |
ISCAS | 2 |
| 2020 | Soft Error and Its Countermeasures in Terrestrial EnvironmentabstractThis paper discusses soft errors in digital chips consisting of SRAM, flip-flops, and combinational logic in the terrestrial environment. We review the effectiveness of error-correction coding (ECC) in processor systems and point out the importance of radiation-hardened flip-flops for further error mitigation. The discussion includes the difference between planar and FD-SOI transistors, and the type of secondary cosmic rays including neutron and muon, using irradiation test results. Also, the difficulty in characterizing SER of a commercial GPU chip is exemplified. Masanori Hashimoto, Wang Liao 0001 |
ASP-DAC | 1 |
| 2020 | When Single Event Upset Meets Deep Neural Networks: Observations, Explorations, and RemediesabstractDeep Neural Network has proved its potential in various perception tasks and hence become an appealing option for interpretation and data processing in security sensitive systems. However, security-sensitive systems demand not only high perception performance, but also design robustness under various circumstances. Unlike prior works that study network robustness from software level, we investigate from hardware perspective about the impact of Single Event Upset (SEU) induced parameter perturbation (SIPP) on neural networks. We systematically define the fault models of SEU and then provide the definition of sensitivity to SIPP as the robustness measure for the network. We are then able to analytically explore the weakness of a network and summarize the key findings for the impact of SIPP on different types of bits in a floating point parameter, layer-wise robustness within the same network and impact of network depth. Based on those findings, we propose two remedy solutions to protect DNNs from SIPPs, which can mitigate accuracy degradation from 28% to 0.27% for ResNet with merely 0.24-bit SRAM area overhead per parameter. Zheyu Yan, Yiyu Shi 0001, Wang Liao 0001, Masanori Hashimoto, Xichuan Zhou, Cheng Zhuo |
ASP-DAC | 4 |
| 2020 | BYNQNet: Bayesian Neural Network with Quadratic Activations for Sampling-Free Uncertainty Estimation on FPGAabstractAn efficient inference algorithm for Bayesian neural network (BNN) named BYNQNet, Bayesian neural network with quadratic activations, and its FPGA implementation are proposed. As neural networks find applications in mission critical systems, uncertainty estimations in network inference become increasingly important. BNN is a theoretically grounded solution to deal with uncertainty in neural network by treating network parameters as random variables. However, an inference in BNN involves Monte Carlo (MC) sampling, i.e., a stochastic forwarding is repeated N times with randomly sampled network parameters, which results in N times slower inference compared to non-Bayesian approach. Although recent papers proposed sampling-free algorithms for BNN inference, they still require evaluation of complex functions such as a cumulative distribution function (CDF) of Gaussian distribution for propagating uncertainties through nonlinear activation functions such as ReLU and Heaviside, which requires considerable amount of resources for hardware implementation. Contrary to conventional BNN, BYNQNet employs quadratic nonlinear activation functions and hence the uncertainty propagation can be achieved using only polynomial operations. Our numerical experiment reveals that BYNQNet has comparative accuracy with MC-based BNN which requires N=10 forwardings. We also demonstrate that BYNQNet implemented on Xilinx PYNQ-Z1 FPGA board achieves the throughput of 131×103images per second and the energy efficiency of 44.7×103images per joule, which corresponds to 4.07× and 8.99× improvements from the state-of-the-art MC-based BNN accelerator. Hiromitsu Awano, Masanori Hashimoto |
DATE | 2 |
| 2020 | Fault Diagnosis of Via-Switch Crossbar in Non-volatile FPGAabstractFPGA that exploits via-switches, which are a kind of non-volatile resistive RAMs, for crossbar implementation is attracting attention due to its high integration density and energy efficiency. Via-switch crossbar is responsible for the signal routing by changing on/off-states of via-switches. To verify the via-switch crossbar functionality after manufacturing, fault testing that checks whether we can turn on/off via-switches normally is essential. This paper confirms that a general differential pair comparator successfully discriminates on/off-states of via-switches, and clarifies fault modes of a via-switch by transistor-level SPICE simulation that injects stuck-on/off faults to atom switch and varistor, where a via-switch consists of two atom switches and two varistors. We then propose a fault diagnosis methodology that diagnoses the fault modes of each via-switch using the comparator response difference between normal and faulty via-switches. The proposed method achieves 100% fault detection by checking the comparator responses after turning on/off the via-switch. In case that the number of faulty components in a via-switch is one, the ratio of the fault diagnosis, which exactly identifies the faulty varistor and atom switch inside the faulty via-switch, is 100%, and in case of up to two faults, the fault diagnosis ratio is 79%. Ryutaro Doi, Toshitsugu Sakamoto, Masanori Hashimoto |
DATE | 4 |
| 2020 | Low-Cost Reservoir Computing using Cellular Automata and Random ForestsabstractHigh-performance image classification models involve massive computation and an energy cost that are unaffordable for resource-limited platforms. As a solution, reservoir computing based on cellular automata has been proposed, but there is still room for improvement in terms of classification cost. This research builds on the previous work introducing enhancements at both the algorithmic and architectural level. Using a random forest classifier with binary features completely eliminates multiplication operations and 97% of addition operations. Also, memory usage can be decreased by pruning 82% of the least relevant augmented features. An architecture with an increased level of parallelism which processes images in a single pass reduces memory accesses, and reduces 60% of logic by optimizing FPGA mapping. These speed, power, and memory optimizations come at an accuracy tradeoff of a mere 0.6%. Ángel López García-Arias, Jaehoon Yu, Masanori Hashimoto |
ISCAS | 3 |
| 2020 | Proactive Supply Noise Mitigation with Low-Latency Minor Voltage Regulator and Lightweight Current PredictionabstractPower supply noise induces extra timing delay or even malfunctions in modern power-demanding VLSI chips. Traditional reactive noise mitigation is often too late to suppress emergent supply noise due to the long latency of voltage boosting. This paper proposes a proactive method for mitigating emergent supply noises and avoiding unexpected failures in power-hungry VLSI designs with two contributions. First, a major-minor voltage regulator (MMVR) structure, which enables quick and widerange voltage scaling with small ripples, is proposed. Second, a lightweight current predictor consisting of a six-layer decision tree regressor achieves over 0.98 correlation for 50-cycle-ahead prediction in 25 RISC-V benchmark programs. Experimental results with a multi-core RISC-V design show that the proposed method mitigates the supply noise within 30 mV while the noise exceeds 70 mV with the conventional reactive mitigation. Also, the average supply voltage is compensated during the power-demanding operation. Masanori Hashimoto |
ITC | 2 |
| 2020 | Concurrent Detection of Failures in GPU Control Logic for Reliable Parallel ComputingabstractThe reliability of GPUs is becoming a major concern due to the increased probability of failures and the high vulnerability of GPUs compared to conventional CPUs in terms of tasks per failure. While there are extensive countermeasures against failures in GPU data units, there are fewer countermeasures for failures in GPU control logics. Currently, software-based techniques, such as inserting signature codes for detecting GPU control-logic failures by comparing the expected signature value with the current signature value, are being utilized. However, in the conventional software-based techniques, application calculations, signature calculations, and signature comparison calculations are executed in sequence, which degrades the application throughputs. We have developed a software-based technique that concurrently detects GPU control-logic failures in a running application while largely maintaining its throughput. Experimental results show that when our technique concurrently executed application calculations, signature calculations, and signature comparison calculations for a matrix multiplication application, the application throughput remains 78% of the original one, whereas 62% is reported in literature. We also developed fault injection simulators specialized for injecting GPU-specific control-logic faults into GPU intermediate codes and found that 100% of GPU-specific failures could be detected both during and after application execution. The proposed approach can be utilized for a wide variety of safety-and reliability-critical applications. Hiroaki Itsuji, Takumi Uezono, Tadanobu Toba, Kojiro Ito, Masanori Hashimoto |
ITC | 5 |
| 2020 | Logarithm-approximate floating-point multiplier is applicable to power-efficient neural network training
TaiYu Cheng, Yukata Masuda, Jaehoon Yu, Masanori Hashimoto |
Integr. | 5 |
| 2020 | Sneak Path Free Reconfiguration With Minimized Programming Steps for Via-Switch Crossbar-Based FPGAabstractField programmable gate array (FPGA) that utilizes via-switches, which are a kind of nonvolatile resistive RAMs, for crossbar implementation is attracting attention due to higher integration density and performance. However, programming via-switches arbitrarily in a crossbar is not trivial since a programming current must be provided through signal wires shared by multiple via-switches. Consequently, depending on the previous programming status in sequential programming, unintentional switch programming may occur due to signal detour, which is called the sneak path problem. This article identifies the circuit status that causes the sneak path problem and proposes a sneak path avoidance method that gives sneak path free programming order of via-switches in a crossbar. We prove that sneak path free programming order necessarily exists for arbitrary ON-OFF patterns in a crossbar as long as no loops exist. This article also proposes a partial reconfiguration method that achieves the minimum number of switch programming steps while avoiding the sneak path problem. This method contributes to the extension of via-switch lifetime and fast reconfiguration of the via-switch FPGA. Experimental results show that the proposed partial reconfiguration method reduces the number of programmed switches by 77.4% compared to the conventional approach. This 77.4% reduction improves the number of reconfigurations of the via-switch FPGA by 4.4× and reduces reconfiguration time by 77.4%. Ryutaro Doi, Jaehoon Yu, Masanori Hashimoto |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Characterizing SRAM and FF soft error rates with measurement and simulation
Masanori Hashimoto, Kazutoshi Kobayashi, Jun Furuta, Shin-ichiro Abe, Yukinobu Watanabe |
Integr. | 1 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 16 |
| 2018 | MTTF-aware design methodology of error prediction based adaptively voltage-scaled circuitsabstractAdaptive voltage scaling is a promising approach to overcome manufacturing variability, dynamic environmental fluctuation, and aging. This paper focuses on error prediction based adaptive voltage scaling (EP-AVS) and proposes an MTTF-aware design methodology for EP-AVS circuits. Main contributions of this work include (1) optimization of both voltage-scaled circuit and voltage control logic, and (2) quantitative evaluation of voltage reduction for practically long MTTF. Evaluation results show that the proposed EP-AVS design methodology achieves 20.8% voltage reduction while satisfying target MTTF. Yutaka Masuda, Masanori Hashimoto |
ASP-DAC | 2 |
| 2018 | Virtualsync: timing optimization by synchronizing logic waves with sequential and combinational components as delay unitsabstractIn digital circuit designs, sequential components such as flip-flops are used to synchronize signal propagations. Logic computations are aligned at and thus isolated by flip-flop stages. Although this fully synchronous style can reduce design efforts significantly, it may affect circuit performance negatively, because sequential components can only introduce delays into signal propagations instead of accelerating them. In this paper, we propose a new timing model, VirtualSync, in which signals, specially those along critical paths, are allowed to propagate through several sequential stages without flip-flops. Timing constraints are still satisfied at the boundary of the optimized circuit to maintain a consistent interface with existing designs. By removing clock-to-q delays and setup time requirements of lip-lops on critical paths, the performance of a circuit can be pushed even beyond the limit of traditional sequential designs. Experimental results demonstrate that circuit performance can be improved by up to 11.5% (average 3.1%) compared with that after thorough sizing and retiming, while the increase of area is still negligible. Grace Li Zhang, Bing Li 0005, Masanori Hashimoto, Ulf Schlichtmann |
DAC | 3 |
| 2018 | Sneak path free reconfiguration of via-switch crossbars based FPGAabstractFPGA that utilizes via-switches, which are a kind of nonvolatile resistive RAMs, for crossbar implementation is attracting attention due to higher integration density and performance. However, programming via-switches arbitrarily in a crossbar is not trivial since a programming current must be provided through signal wires that are shared by multiple via-switches. Consequently, depending on the previous programming status in sequential programming, unintentional switch programming may occur due to signal detour, which is called sneak path problem. This problem interferes the reconfiguration of via-switch FPGA, and hence countermeasures for sneak path problem are indispensable. This paper identifies the circuit status that causes sneak path problem and proposes a sneak path avoidance method that gives sneak path free programming order of via-switches in a crossbar. We prove that sneak path free programming order necessarily exists for arbitrary on-off patterns in a crossbar as long as no loops exist, and also validate the proof and the proposed method with simulation-based evaluation. Thanks to the proposed method, any practical configurations of via-switch FPGA can be successfully programmed without sneak path problem. Ryutaro Doi, Jaehoon Yu, Masanori Hashimoto |
ICCAD | 3 |
| 2018 | Comparing voltage adaptation performance between replica and in-situ timing monitorsabstractAdaptive voltage scaling (AVS) is a promising approach to overcome manufacturing variability, dynamic environmental fluctuation, and aging. This paper focuses on timing sensors necessary for AVS implementation and compares in-situ timing error predictive FF (TEP-FF) and critical path replica in terms of how much voltage margin can be reduced. For estimating the theoretical bound of ideal AVS, this work proposes linear programming based minimum supply voltage analysis and discusses the voltage adaptation performance quantitatively by investigating the gap between the lower bound and actual supply voltages. Experimental results show that TEP-FF based AVS and replica based AVS achieve up to 13.3% and 8.9% supply voltage reduction, respectively while satisfying the target MTTF. AVS with TEP-FF tracks the theoretical bound with 2.5 to 5.6% voltage margin while AVS with replica needs 7.2 to 9.9% margin. Yutaka Masuda, Jun Nagayama, Hirotaka Takeno, Yoshimasa Ogawa, Yoichi Momiyama, Masanori Hashimoto |
ICCAD | 6 |
| 2018 | Activation-Aware Slack Assignment for Time-to-Failure Extension and Power Saving
Yutaka Masuda, Takao Onoye, Masanori Hashimoto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Via-Switch FPGA: Highly Dense Mixed-Grained Reconfigurable Architecture With Overlay Via-Switch CrossbarsabstractThis paper proposes a highly dense reconfigurable architecture that introduces via-switch device, which is a nonvolatile resistive-change switch and is used in crossbar switches. Via-switch is implemented in back-end-of-line layers only, and hence the front-end-of-line (FEoL) layers under the crossbar can be fully exploited for highly dense logic blocks. The proposed architecture uses the FEoL layers for fine-grained lookup tables and coarse-grained arithmetic/memory units for improving performance and compatibility with various applications. A case study of application mapping shows the proposed architecture can reduce the array area by 21.7%, thanks to the bidirectional interconnection. Thanks to F2footprint and one order of magnitude lower resistivity of via-switch compared to MOS switch, the crossbar density is improved by up to 26× and the delay and energy in the interconnection are reduced by 90% and 94% at 0.5-V operation. Hiroyuki Ochi, Kosei Yamaguchi, Tetsuaki Fujimoto, Junshi Hotate, Takashi Kishimoto, Toshiki Higashi, Takashi Imagawa, Ryutaro Doi, Munehiro Tada, Tadahiko Sugibayashi, Wataru Takahashi 0002, Kazutoshi Wakabayashi, Hidetoshi Onodera, Yukio Mitsuyama, Jaehoon Yu, Masanori Hashimoto |
IEEE Trans. Very Large Scale Integr. Syst. | 16 |
| 2017 | Near-field dual-use antenna for magnetic-field based communication and electrical-field based distance sensing in mm3-class sensor nodeabstractThis paper proposes a mm3-class dual-use near-field antenna that can be used for both magnetic-field based communication and electrical-field based distance sensing. The proposed antenna consists of two spiral coils, and they are used as a coil antenna in communication mode and signal electrodes in distance sensing mode. We evaluated the performance of the communication mode with a prototype antenna. The measured S21 is −8.3 dB to −45.1 dB in the range from 6 mm to 24 mm, which is highly correlated to 3D full-wave electromagnetic simulation. With this antenna, we performed BER evaluation with commercial transceiver boards showing that the proposed antenna could be used for ASK/OOK signaling. We also confirmed that the proposed antenna could be used for cm-scale node-to-node distance sensing through capacitive coupling. Ryo Shirai, Jin Kono, Tetsuya Hirose, Masanori Hashimoto |
ISCAS | 4 |
| 2017 | GPGPU-based Highly Parallelized 3D Node Localization for Real-Time 3D Model ReproductionabstractThis paper proposes a highly parallelized 3D node localization method based on cross-entropy method for the 3D modeling system. Cross-entropy localization statistically estimates node positions from node-to-node distance information by sampling, and each sample evaluation and internal computation of objective function can be processed in parallel. Experimental results show our GPGPU-based implementation achieved 5,163x and 61.5x speed up compared to a single processor and 80-processor implementations. In addition, for enhancing model reproduction accuracy, this work introduces a penalty function to mitigate flip ambiguity. Kauzki Hirosue, Shohei Ukawa, Yuichi Itoh, Takao Onoye, Masanori Hashimoto |
IUI | 5 |
| 2017 | Near-future traffic evaluation based navigation for automated driving vehiclesabstractOnce vehicles start to be driven automatically, people expect the driving routing is automatically and optimally selected. Supposing all the vehicles are navigated by a single system in the future, the navigation system will be able to provide instructions to each vehicle based on the evaluated near-future traffic information while the current navigation system frequently updates the routing based on current traffic information. This paper proposes a navigation method that guides vehicles based on the evaluated near-future traffic information. Experimental results with actual city maps show the evaluated near-future traffic information is helpful to mitigate traffic jam and reduce driving time. Kuen-Wey Lin, Yih-Lang Li, Masanori Hashimoto |
Intelligent Vehicles Symposium | 3 |
| 2017 | Minimizing detection-to-boosting latency toward low-power error-resilient circuits
Chih-Cheng Hsu, Masanori Hashimoto, Mark Po-Hung Lin |
Integr. | 2 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 15 |
| 2016 | Reliability, adaptability and flexibility in timing: Buy a life insurance for your circuitsabstractAt nanometer manufacturing technology nodes, process variations affect circuit performance significantly. In addition, performance deterioration of circuits due to aging effects is also increasing. Consequently, a large timing margin is required to maintain yield. To combat the pessimism and the resulting overdesign, aging analysis with highlevel models, on-chip timing margin monitoring and tuning, and flexible delay models of flip-flops can be deployed. This paper gives an overview of the state of the art of applying these techniques to improve the health of circuits. Ulf Schlichtmann, Masanori Hashimoto, Iris Hui-Ru Jiang, Bing Li 0005 |
ASP-DAC | 2 |
| 2016 | A highly-dense mixed grained reconfigurable architecture with overlay crossbar interconnect using via-switchabstractThis paper proposes a highly-dense reconfigurable architecture that introduces via-switch device, which is a kind of resistive RAM and is used in crossbar switches. Since via-switch is implemented in BEOL layers only, the FEOL layer under the crossbar can be fully exploited for highly-dense logic blocks. The proposed architecture uses the FEOL layer for fine-grained look-up tables and coarse-grained arithmetic/memory units for better performance and highly wide applications. In a case study of application mapping, the proposed architecture reduces array area by 76% thanks to mixed grained logic structure and overlay bidirectional interconnection. Thanks to 18F2footprint and one order of magnitude lower resistivity of via-switch compared to MOS switch, the crossbar density is improved by 26× and the delay and energy in the interconnection are reduced by 90% and 93% at 0.5V operation. Junshi Hotate, Takashi Kishimoto, Toshiki Higashi, Hiroyuki Ochi, Ryutaro Doi, Munehiro Tada, Tadahiko Sugibayashi, Kazutoshi Wakabayashi, Hidetoshi Onodera, Yukio Mitsuyama, Masanori Hashimoto |
FPL | 11 |
| 2016 | Critical path isolation for time-to-failure extension and lower voltage operationabstractDevice miniaturization due to technology scaling has made manufacturing variability and aging more significant, and lower supply voltage makes circuits sensitive to dynamic environmental fluctuation. These may shorten the time to failure (TTF) of fabricated chips unexpectedly. This paper focuses on critical path isolation, which increases timing slack of non-intrinsic critical paths and decreases timing error occurrence probability in the circuit, and proposes a design methodology of isolated circuits for TTF extension and/or lower voltage operation. The proposed methodology selects a set of FFs for isolation using ILP so that it maximumly reduces the sum of gate-wise failure probabilities. We evaluated MTTF (Mean Time To Failure) of circuits with/without critical path isolation and examined how much supply voltage could be reduced without MTTF degradation. Evaluation results show that circuits with the proposed critical path isolation achieved 25% supply voltage reduction with 1.4% area overhead. With the same supply voltage, MTTF was improved by 14 orders of magnitude. Yutaka Masuda, Masanori Hashimoto, Takao Onoye |
ICCAD | 2 |
| 2016 | Hardware-simulation correlation of timing error detection performance of software-based error detection mechanismsabstractSoftware-based error detection techniques, which includes EDM (error detection mechanisms) transformation, are used for error localization in post-silicon validation. This paper evaluates the performance of EDM for timing error localization with 65-nm test chips assuming the following two EDM usage scenarios; (1) localizing a timing error occurred in the original program, and (2) localizing potential timing errors that vary execution results. Experimental results show that the EDM transformation customized for quick error detection detects 25% of timing errors in the original program in the first scenario and 56% of non-masked errors in the second scenario. However, these hardware measurement results are not consistent with the simulation results of our previous work. To investigate the reason, we focus on the following two differences between hardware and simulation; (1) design of power distribution network, and (2) definition of timing error occurrence frequency. We update the simulation setup for filling the difference and re-execute the simulation. We confirm that the simulation and the chip measurement results are consistent, which validates our simulation methodology. Yutaka Masuda, Masanori Hashimoto, Takao Onoye |
IOLTS | 2 |
| 2015 | An oscillator-based true random number generator with process and temperature toleranceabstractThis paper presents an oscillator-based true random number generator (TRNG) that automatically adjusts the duty cycle of a fast oscillator to 50 %, and generates unbiased random numbers tolerating process variation and dynamic temperature fluctuation. Measurement results with 65nm test chips show that the proposed TRNG adjusted the probability of `1' to within 50 ± 0.07 % in five chips in the temperature range of 0 °C to 75 °C. Consequently, the proposed TRNG passed the NIST and DIEHARD tests at 7.5 Mbps with 6,670 μm2area. Takehiko Amaki, Masanori Hashimoto, Takao Onoye |
ASP-DAC | 2 |
| 2015 | Reliability-configurable mixed-grained reconfigurable array compatible with high-level synthesisabstractThis paper presents a mixed-grained reconfigurable VLSI array architecture that can cover mission-critical applications to consumer products through C-to-array application mapping. A proof-of-concept VLSI chip was fabricated in a 65nm process. Measurement results show that applications on the chip can be working in a harsh radiation environment. Masanori Hashimoto, Dawood Alnajiar, Hiroaki Konoura, Yukio Mitsuyama, Hajime Shimada, Kazutoshi Kobayashi, Hiroyuki Kanbara, Hiroyuki Ochi, Takashi Imagawa, Kazutoshi Wakabayashi, Takao Onoye, Hidetoshi Onodera |
ASP-DAC | 1 |
| 2015 | Area efficient device-parameter estimation using sensitivity-configurable ring oscillatorabstractThis paper proposes an area efficient device parameter estimation method with sensitivity-configurable ring oscillator (RO). This sensitivity-configurable RO has a number of configurations and the proposed method exploits this property for reducing sensor area and/or improving estimation accuracy. The proposed method selects multiple sets of sensitivity configurations, obtains multiple estimates and computes the average of them for accuracy improvement exploiting an averaging effect. Experimental results with a 32-nm predictive technology model show that the proposed method can reduce the estimation error by 49% or reduce the sensor area by 75% while keeping the accuracy. Shoichi Iizuka, Yuma Higuchi, Masanori Hashimoto, Takao Onoye |
ASP-DAC | 3 |
| 2015 | Performance Evaluation of Software-based Error Detection Mechanisms for Localizing Electrical Timing Failures under Dynamic Supply NoiseabstractFor facilitating error localization, software-based error detection techniques have been proposed and EDM (error detection mechanisms) transformation is one of these techniques. To discuss the effectiveness of EDM for electrical bug localization, two scenarios are considered; (1) localizing an electrical bug occurred in the original program, and (2) localizing as many potential bugs as possible. We experimentally evaluated the error detection performance in these two scenarios under dynamic power supply noise. Experimental results show that the EDM transformation customized for quick error detection cannot locate electrical bugs in the original program in the firs scenario, but it is useful for findin potential bugs in the second scenario. Yutaka Masuda, Masanori Hashimoto, Takao Onoye |
ICCAD | 2 |
| 2015 | Real-time on-chip supply voltage sensor and its application to trace-based timing error localizationabstractThis paper presents an all-digital on-chip supply voltage sensor that captures one-shot voltage fluctuation every clock cycle. The proposed sensor was implemented on ASIC in 65nm process and FPGA. The obtained voltage resolution was 3.9mV and 29mV, respectively. This sensor is suitable for providing voltage information to trace-based error localization system. We experimentally show that the proposed sensor contributes to the facilitation of error localization. Miho Ueno, Masanori Hashimoto, Takao Onoye |
IOLTS | 2 |
| 2015 | Stochastic timing error rate estimation under process and temporal variationsabstractReducing design and operational margin is a key factor that makes fabricated chips competitive in terms of speed and power consumption. On the other hand, a smaller margin involves a higher risk that a timing error occurs in field. This paper proposes a stochastic framework that estimates timing error rate under static process variations and dynamic environmental variations for circuits with and without run-time adaptive speed control. The proposed framework extends the state assignment of the continuous-time Markov process used in the previous work so as to take into account within-die random variation, and speeds up the database construction for the transition rate matrix by combining logic simulation and statistical static timing analysis. This paper also demonstrates that the proposed framework can cope with transistor-by-transistor stochastic aging processes. Experimental results show that the within-die random variation deviates the MTTF with σ of 52%. The CPU time for the transition rate matrix computation is reduced to 1/30. Shoichi Iizuka, Yutaka Masuda, Masanori Hashimoto, Takao Onoye |
ITC | 3 |
| 2015 | 3D node localization from node-to-node distance information using cross-entropy methodabstractThis paper proposes a 3D node localization method that uses cross-entropy method for the 3D modeling system. The proposed localization method statistically estimates the most probable positions overcoming measurement errors through iterative sample generation and evaluation. The generated samples are evaluated in parallel, and then a significant speedup can be obtained. We also demonstrate that the iterative sample generation and evaluation performed in parallel are highly compatible with interactive node movement. Shohei Ukawa, Tatsuya Shinada, Masanori Hashimoto, Yuichi Itoh, Takao Onoye |
VR | 3 |
| 2014 | Opportunities and Verification Challenges of Run-Time Performance AdaptationabstractRun-time performance adaptation with field delay testing is a promising approach for minimizing design margin while sustaining necessary operational margin in the field. However, run-time performance adaptation has not been adopted in industrial designs since a serious concern on timing error occurrence exists. For putting the run-time performance adaptation in a practical use, we need to verify and optimize the run-time adaptation system in design time, but a straightforward verification with logic simulation could need billion years and is totally insufficient. For this problem, we have developed a stochastic framework for error rate estimation that models adaptive speed control as a continuous-time Markov process. This paper first exemplifies the power reduction thanks to run-time performance adaptation with a 65nm test chip. Then, the proposed stochastic framework is introduced. With this framework, we evaluate MTTF of an embedded processor whose performance is adaptively controlled with online testing and offline testing. This evaluation shows how design parameters affect MTTF as an example. Masanori Hashimoto |
ATS | 1 |
| 2013 | Extracting device-parameter variations using a single sensitivity-configurable ring oscillatorabstractThe RO(Ring-Oscillator)-based sensor is one of easily-implementable variation sensors, but for decomposing the observed variability into multiple unique device-parameter variations, a large number of ROs with different structures and sensitivities to device-parameters is required. This paper proposes a scheme for sensing multiple device-parameter variations with just a single reconfigurable RO. This sensitivity-configurable RO has a number of configurations available and this property can be exploited for reducing sensor area while improving estimation accuracy through iterative estimation. To minimize the prospective error, the proposed estimation iterates: (1) selecting the best configuration that minimizes the prospective estimation error around the current estimates; and (2) updating the estimates with the selected configuration. This experiment was carried out assuming a 32-nm predictive technology model. Experimental results show that device-parameter extraction with a single RO is feasible and the error of the extracted parameters is reduced by 35 to 53% with the improved objective function and iterative estimation. Yuma Higuchi, Kenichi Shinkai, Masanori Hashimoto, Rahul M. Rao, Sani R. Nassif |
ETS | 3 |
| 2013 | Stochastic error rate estimation for adaptive speed control with field delay testingabstractThis paper proposes a stochastic framework for error rate estimation that models adaptive speed control as a continuous-time Markov process and derives its transition rates using developed similarity database. The proposed framework is implemented for adaptive speed control systems based on timing error prediction and scan-test. Experimental results show that the proposed framework enabled 12 orders of magnitude faster MTTF estimation than ordinary logic simulation. The accuracy of MTTF estimation under random delay fluctuation is clarified through a comparison with logic simulation. The proposed estimation can contribute to design and validation of adaptive speed control systems with field delay testing. Shoichi Iizuka, Masafumi Mizuno, Dan Kuroda, Masanori Hashimoto, Takao Onoye |
ICCAD | 4 |
| 2013 | A gate-delay model focusing on current fluctuation over wide range of process-voltage-temperature variations
Kenichi Shinkai, Masanori Hashimoto, Takao Onoye |
Integr. | 2 |
| 2013 | A Worst-Case-Aware Design Methodology for Noise-Tolerant Oscillator-Based True Random Number Generator With Stochastic Behavior ModelingabstractThis paper presents a worst-case-aware design methodology for an oscillator-based true random number generator (TRNG) that produces highly random bit streams even under deterministic noise. We propose a stochastic behavior model to efficiently determine design parameters, and identify a class of deterministic noise under which the randomness gets the worst. They can be used to directly estimate the worst χ value of a poker test under deterministic noise without generating bit streams, which enables efficient exploration of design space and guarantees sufficient randomness in a hostile environment. The proposed model is validated by measuring prototype TRNGs fabricated with a 65-nm CMOS process. Takehiko Amaki, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Implementing Flexible Reliability in a Coarse-Grained Reconfigurable ArchitectureabstractThis paper proposes a coarse-grained dynamically reconfigurable architecture that offers flexible reliability to deal with soft errors and aging. The notion of a cluster is introduced as a basic architectural element; each cluster can select four operation modes with different levels of spatial redundancy and area efficiency. We evaluate the aging effect due to negative bias temperature instability and illustrate that periodically alternating active cells with resting ones will greatly mitigate the effects of the aging process with a negligible power overhead. The area of circuits that are added for immunity to soft errors and for mitigating aging effects is 29.3% of the proposed reconfigurable device. A fault-tolerance evaluation of a Viterbi decoder mapped on the architecture suggests that there is a considerable tradeoff between reliability and area overhead. Finally, we design and fabricate a test chip that contains a 4 × 8 cluster array in a 65-nm process and demonstrate its immunity to soft errors. Accelerated tests using an alpha particle foil showed that the mean time to failure and failure in time are well characterized with the number of sensitive bits and that our architecture can trade off soft error immunity with the area of implementation. Dawood Alnajiar, Hiroaki Konoura, Younghun Ko, Yukio Mitsuyama, Masanori Hashimoto, Takao Onoye |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2013 | Supply Noise Suppression by Triple-Well StructureabstractThis brief discusses the impact of twin- and triple-well structures on power supply noise, and a substrate model for simulating the power supply noise. We observedVssnoise reduction by the resistive network of the p-substrate andVddnoise reduction by the junction capacitance of a triple-well structure on a 90-nm test chip. Measurement results also showed that the total noise reduction of a triple-well structure is superior to that of a twin-well structure. The measurement results correlate well with the results obtained from the power supply noise simulation using a hierarchical resistive mesh model. Our simulation-based verification indicates that in common CMOS design, a triple-well structure can reduce the power supply drop by 10%-40% or the decoupling capacitance area by 5%-10%. We also verified that supply drop sensitivity to variation of the well junction capacitance is sufficiently small and that supply noise reduction using a triple-well structure is robust to process variation. Yasuhiro Ogasahara, Masanori Hashimoto, Toshiki Kanamoto, Takao Onoye |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Body bias clustering for low test-cost post-silicon tuningabstractPost-silicon tuning is attracting a lot of attention for coping with increasing process variation. However, its tuning cost via testing is still a crucial problem. In this paper, we propose tuning-friendly body bias clustering with multiple bias voltages. The proposed method provides a small set of compensation levels so that the speed and leakage current vary monotonically according to the level. Thanks to this monotonic leveling and limitation of the number of levels, the test-cost of post-silicon tuning is significantly reduced. During the body bias clustering, the proposed method explicitly estimates and minimizes the average leakage after the post-silicon tuning. Experimental results demonstrate that the proposed method reduces the average leakage by 25.3 to 51.9% compared to non clustering case. We reveal that two bias voltages are sufficient when only a small number of compensation levels are allowed for test-cost reduction. We also give an implication on how to synthesize a circuit to which post-silicon tuning will be applied. Shuta Kimura, Masanori Hashimoto, Takao Onoye |
ASP-DAC | 2 |
| 2012 | A predictive delay fault avoidance scheme for coarse-grained reconfigurable architectureabstractA scheme for avoiding delay faults with slack assessment during standby time is proposed in this paper. The proposed scheme performs path delay testing and checks if the slack is larger than a threshold value using selectable delay embedded in basic elements (BE) on a coarse-grained reconfigurable device. If the slack is smaller than the threshold, a pair of BEs to be replaced, which maximizes the path slack, is identified. Experimental results show that for aging-induced delay degradation a small threshold slack, which is less than 1 ps in a test case, is enough to ensure the delay fault prediction. Toshihiro Kameda, Hiroaki Konoura, Dawood Alnajiar, Yukio Mitsuyama, Masanori Hashimoto, Takao Onoye |
FPL | 5 |
| 2012 | Adaptive Performance Compensation With In-Situ Timing Error Predictive Sensors for Subthreshold CircuitsabstractWe present an adaptive technique for compensating manufacturing and environmental variability in subthreshold circuits using “canary flip-flop (FF),” which can predict timing errors. A 32-bit Kogge-Stone adder whose performance was controlled by body-biasing was fabricated in a 65-nm CMOS process. Measurement results show that the adaptive control can compensate process, supply voltage, and temperature variations and improve the energy efficiency of subthreshold circuits by up to 46% compared to worst-case design and operation with guardbanding. We also discuss how to determine design parameters, such as the inserted location and the buffer delay of the canary FF, supposing two approaches: configuration in the design phase and post-silicon tuning. Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Jitter amplifier for oscillator-based true random number generatorabstractThis paper presents a jitter amplifier for oscillator-based TRNG (true random number generator). The proposed jitter amplifier fabricated in a 65nm CMOS process occupying the area of 3,300 μm2archives 8.4× gain at 25°C and significantly improves the entropy enough to pass randomness test. Takehiko Amaki, Masanori Hashimoto, Takao Onoye |
ASP-DAC | 2 |
| 2011 | Run-time adaptive performance compensation using on-chip sensorsabstractThis paper discusses run-time adaptive performance control with on-chip sensors that predict timing errors. The sensors embedded into functional circuits capture delay variations due to not only die-to-die process variation but also random process variation, environmental fluctuation and aging. By compensating circuit performance according to the sensor outputs, we can overcome PVT worst-case design and reduce power dissipation while satisfying circuit performance. We applied the adaptive speed control to subthreshold circuits that are very sensitive to random variation and environmental fluctuation. Measurement results of a 65nm test chip show that the adaptive speed control can compensate PVT variations and improve energy efficiency by up to 46% compared to the worst-case design and operation with guardbanding. Masanori Hashimoto |
ASP-DAC | 1 |
| 2011 | Device-parameter estimation with on-chip variation sensors considering random variabilityabstractDevice-parameter monitoring sensors inside a chip are gaining its importance as the post-fabrication tuning is becoming of a practical use. In estimation of variational parameters using on-chip sensors, it is often assumed that the outputs of variation sensors are not affected by random variations. However, random variations can deteriorate the accuracy of the estimation result. In this paper, we propose a device-parameter estimation method with on-chip variation sensors explicitly considering random variability. The proposed method derives the global variation parameters and the standard deviation of the random variability using the maximum likelihood estimation. We experimentally verified that the proposed method can accurately estimate variations, whereas the estimation result deteriorates when neglecting random variations. We also demonstrate an application result of the proposed method to test chips fabricated in a 65-nm process technology. Kenichi Shinkai, Masanori Hashimoto |
ASP-DAC | 2 |
| 2011 | Implications of Reliability Enhancement Achieved by Fault Avoidance on Dynamically Reconfigurable ArchitecturesabstractFault avoidance methods on dynamically reconfigurable devices have been proposed to extend device life-time, while their quantitative comparison has not been sufficiently presented. This paper shows results of quantitative life-time evaluation by simulating fault avoidance procedures of representative five methods under the same conditions of wear-out scenario, application and device architecture. Experimental results reveal 1) MTTF is highly correlated with the number of avoided faults, 2) there is the efficiency difference of spare usage in five fault avoidance methods, and 3) spares should be prevented from wear-out not to spoil life-time enhancement. Hiroaki Konoura, Yukio Mitsuyama, Masanori Hashimoto, Takao Onoye |
FPL | 3 |
| 2011 | An oscillator-based true random number generator with jitter amplifierabstractThis paper presents an oscillator-based TRNG (true random number generator) with jitter amplifier. The proposed jitter amplifier fabricated in a 65nm CMOS process archives 8.4× gain at 25°C, and significantly improves randomness of output bitstream. The TRNG with the jitter amplifier enhances throughput per area by 94% compared to a TRNG with frequency dividers. The prototype TRNG occupies 6,300 μm2, generates 2 Mbps random bitstreams, and passes FIPS 140-2 randomness tests and 12 tests in NIST test suite. Takehiko Amaki, Masanori Hashimoto, Takao Onoye |
ISCAS | 2 |
| 2010 | Adaptive performance control with embedded timing error predictive sensors for subthreshold circuitsabstractThis paper presents an adaptive technique for compensating manufacturing and environmental variability in subthreshold circuits using ¿canary flip-flop¿ that can predict timing errors. A 32-bit Kogge-Stone adder whose performance was controlled by body-biasing was fabricated in a 65 nm CMOS process. Measurement results show that the adaptive control can reduce the power dissipation by 46% in comparison with the worst-case design with guardbanding. Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
ASP-DAC | 2 |
| 2010 | Gate delay estimation in STA under dynamic power supply noiseabstractThis paper presents a gate delay estimation method that takes into account dynamic power supply noise. We review STA based on static IR-drop analysis and a conventional method for dynamic noise waveform, and reveal their limitations and problems that originate from circuit structures and higher delay sensitivity to voltage in advanced technologies. We then propose a gate delay computation that overcomes the problems with iterative computations and consideration of input voltage drop. Evaluation results with various circuits and noise injection timings show that the proposed method estimates path delay fluctuation well within 2% error on average. Takaaki Okumura, Fumihiro Minami, Kenji Shimazaki, Kimihiko Kuwada, Masanori Hashimoto |
ASP-DAC | 5 |
| 2010 | Clock skew reduction by self-compensating manufacturing variability with on-chip sensorsabstractThis paper presents a self-compensation scheme of manufacturing variability for clock skew reduction. In the proposed scheme, a CDN with embedded variability sensors tunes variable clock drivers for canceling the clock skew induced by manufacturing variability. We apply the proposed scheme for a mesh-style CDN in a 65nm technology and evaluate the deskewing effect as a function of the sensor performance. Experimental results show that the skew can be reduced by over 70% and the correlation coefficient between estimated and actual variabilities, which represents the sensor performance, should be more than 0.3 for skew reduction. Shinya Abe, Kenichi Shinkai, Masanori Hashimoto, Takao Onoye |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | Modeling the Overshooting Effect for CMOS Inverter Delay Analysis in Nanometer TechnologiesabstractWith the scaling of complementary metal-oxide-semiconductor (CMOS) technology into the nanometer regime, the overshooting effect due to the input-to-output coupling capacitance has more significant influence on CMOS gate analysis, especially on CMOS gate static timing analysis. In this paper, the overshooting effect is modeled for CMOS inverter delay analysis in nanometer technologies. The results produced by the proposed model are close to simulation program with integrated circuit emphasis (SPICE). Moreover, the influence of the overshooting effect on CMOS inverter analysis is discussed. An analytical model is presented to calculate the CMOS inverter delay time based on the proposed overshooting effect model, which is verified to be in good agreement with SPICE results. Furthermore, the proposed model is used to improve the accuracy of the switch-resistor model for approximating the inverter output waveform. Zhangcai Huang, Atsushi Kurokawa, Masanori Hashimoto, Takashi Sato 0001, Minglu Jiang, Yasuaki Inoue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Transistor Variability Modeling and its Validation With Ring-Oscillation Frequencies for Body-Biased Subthreshold CircuitsabstractThis paper presents transistor variability modeling and its validation for body-biased subthreshold circuits based on measurements of a device-array circuit using a 90-nm technology. The device array consists of p/nMOS transistors and ring oscillators. We examine and confirm the correlation between the performance variation model extracted from measured I-V characteristics and fabricated oscillation frequencies. We demonstrate that delay variations in subthreshold circuits are well characterized with two parameters, i.e., threshold voltage and subthreshold swing parameter. We also reveal that threshold voltage shift by body biasing can be deterministically modeled and statistical modeling is less meaningful. Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Trade-off analysis between timing error rate and power dissipation for adaptive speed control with timing error predictionabstractTiming margin of a chip varies chip by chip due to manufacturing variability, and depends on operating environment and aging. Adaptive speed control with timing error prediction is a promising approach to mitigate the timing margin variation, whereas it inherently has a critical risk of timing error occurrence when a circuit is slowed down. This paper presents how to evaluate the relation between timing error rate and power dissipation in self-adaptive circuits with timing error prediction. The discussion is experimentally validated using a 32-bit ripple carry adder in subthreshold operation in a 90nm CMOS process. We show a trade-off between timing error rate and power dissipation, and reveal the dependency of the trade-off on design parameters. Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
ASP-DAC | 2 |
| 2009 | High performance on-chip differential signaling using passive compensation for global communicationabstractTo address the performance limitation brought by the scaling issues of on-chip global wires, a new configuration for global wiring using on-chip lossy transmission lines is proposed and optimized. We propose a signaling structure to compensate the distortion and attenuation of on-chip transmission lines, which uses passive compensation and inserts repeated transceivers composing sense amplifiers and inverter chains. An optimization flow for designing this scheme based on eye-diagram prediction and sequential quadratic programming (SQP) is devised. This flow is used to study the latency, power dissipation and throughput performance of the new global wiring scheme as the technology scales from 90 nm to 22 nm. Comparing to repeated RC wire, experimental results demonstrate that at 22 nm technology node, the new scheme can reduce the normalized delay by 80%-95%, the normalized energy consumption by 50%-94%. The normalized latency is 10 ps/mm, the energy per bit is 20 pJ/m, and the throughput is 15 Gbps/mum. All performance metrics are scalable with technology, which makes this approach a potential candidate to break the "interconnect wall" of digital system performance. Yulei Zhang 0002, Akira Tsuchiya, Masanori Hashimoto, Ernest S. Kuh, Chung-Kuan Cheng |
ASP-DAC | 4 |
| 2009 | Coarse-grained dynamically reconfigurable architecture with flexible reliabilityabstractThis paper proposes a coarse-grained dynamically reconfigurable architecture, which offers flexible reliability to soft errors and aging. A notion of cluster is introduced as a basic element of the proposed architecture, each of which can select four operation modes with different levels of spatial redundancy and area-efficiency. Evaluation of permanent error rates demonstrates that four different reliability levels can be achieved by the proposed architecture. We also evaluate aging effect due to NBTI, and illustrate that alternating active cells with resting ones periodically will greatly mitigate the aging process with negligible power overhead. The area of additional circuits to attain immunity to soft errors and reliability configuration is 26.6% of the proposed reconfigurable device. Finally, a fault-tolerance evaluation of Viterbi decoder mapped on the proposed architecture suggests that there is a considerable trade-off between reliability and area overhead. Dawood Alnajiar, Younghun Ko, Takashi Imagawa, Hiroaki Konoura, Masayuki Hiromoto, Yukio Mitsuyama, Masanori Hashimoto, Hiroyuki Ochi, Takao Onoye |
FPL | 7 |
| 2009 | Tuning-friendly body bias clustering for compensating random variability in subthreshold circuitsabstractPost-fabrication tuning for mitigating manufacturing variability is receiving a significant attention. To reduce leakage increase involved in performance compensation by body biasing, body bias clustering methods have been proposed. However, conventional methods suffer from a large test cost for tuning after fabrication, since there are a tremendous number of body bias assignments. We in this paper propose a low-cost tuning scheme after fabrication and present a layout aware body bias clustering method. The proposed method estimates average leakage power after post-fabrication tuning, and minimizes it. We applied the proposed method to ultralow voltage circuits for suppressing their high sensitivity to random Vth variability, and demonstrated the effectiveness of the proposed method. In the experiments, by just introducing two clusters, leakage power after post-fabrication tuning was reduced by up to 70% compared to a single cluster case. Koichi Hamamoto, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
ISLPED | 2 |
| 2009 | Statistical Timing Analysis Considering Spatially and Temporally Correlated Dynamic Power Supply NoiseabstractPower supply noise is having increasingly more influence on timing, even though noise-aware timing analysis has not yet been fully established, because of several difficulties such as its dependence on input vectors and dynamic behavior. This paper proposes static timing analysis that takes power supply noise into consideration where the dependence of noise on input vectors and spatial and temporal correlations are handled statistically. We construct a statistical model of power supply voltage that dynamically varies with spatial and temporal correlation, and represent it as a set of uncorrelated variables. We demonstrate that power-voltage variations are highly correlated and adopting principal component analysis as an orthogonalization technique can effectively reduce the number of variables. Experiments confirmed the validity of our model and the accuracy of timing analysis. We also discuss the accuracy and CPU time in association with the reduced number of variables. Takashi Enami, Shinyu Ninomiya, Masanori Hashimoto |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Dynamic supply noise measurement circuit composed of standard cells suitable for in-site SoC power integrity verificationabstractThis paper presents an all digital measurement circuit called “gated oscillator” for capturing waveforms of dynamic power supply noise. The gated oscillator is constructed with standard cells, and thus can be easily embedded in SoCs for design verification. The performance of the gated oscillator is verified with fabricated test chips in a 90nm process. Yasuhiro Ogasahara, Masanori Hashimoto, Takao Onoye |
ASP-DAC | 2 |
| 2008 | High performance current-mode differential logicabstractThis paper presents a new logic style, named Current-Mode Differential logic (CMDL), that achieves both high operating speed and low power consumption. Inspired by the low-voltage swing (LVS) logic, CMDL uses a shunt resistor at the differential output to obtain constant low swing signal without the need to reset low. Furthermore, conditional shunt transistors are used for the internal nodes to prevent high-voltage swing, thus entirely eliminate the power-hungry clocked reset network in LVS circuits. We show that the CMDL is suitable for high-end microprocessor integer core by providing three datapath modules implemented in CMDL. Our simulation results indicate that, operating at comparable speed with LVS logic, CMDL circuits can achieve up to 50% reduction of delay-power product compared to CMOS logic and LVS logic. In addition, CMDL reduces the power consumption of LVS by up to 40%. Haikun Zhu, Chung-Kuan Cheng, Masanori Hashimoto |
ASP-DAC | 5 |
| 2008 | Experimental study on body-biasing layout style-- negligible area overhead enables sufficient speed controllability --abstractBody-biasing is expected to be a common design technique, then area efficient implementation in layout has been demanded. Body-biasing outside standard cells is one of possible layouts. However in this case body-bias controllability, especially when forward bias is applied, is a concern. To investigate the controllability, we fabricated a ring oscillator in a 90nm technology, and measured the controllability. Our measurement result and evaluation of area efficiency reveal that body-biased circuits can be implemented with area overhead of less than 1% yet with sufficient speed controllability. Koichi Hamamoto, Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Decoupling capacitance allocation for timing with statistical noise model and timing analysisabstractThis paper presents an allocation method of decoupling capacitance that explicitly considers timing. We have found and focused that decap does not necessarily improve a gate delay at all the switching timing within a cycle, and devised an efficient sensitivity calculation of timing to decap for decap allocation. The proposed method, which is based on a statistical noise modeling and timing analysis, accelerates the sensitivity calculation with an approximation and adjoint sensitivity analysis. Experimental results show that the decap allocation based on the sensitivity analysis efficiently optimizes the worst-case circuit delay within a given decap budget. Compared to the maximum decap placement, the delay improvement due to decap increases by 5% even while the total amount of decap is reduced to 40%. Takashi Enami, Masanori Hashimoto, Takashi Sato 0001 |
ICCAD | 2 |
| 2008 | On-chip high performance signaling using passive compensationabstractTo address the performance limitation brought by the scaling issues of on-chip global wires, a new configuration for global wiring using on-chip lossy transmission lines(T-lines) is proposed and optimized in this paper. Firstly, we use passive compensation and repeated transceivers composed by sense amplifier and inverter chain to compensate the distortion and attenuation of on-chip T-lines. Secondly, an optimization flow for designing this scheme based on eye-diagram prediction and sequential quadratic programming (SQP) is proposed. This flow is employed to study the latency, power dissipation and throughput performance of the new global wiring scheme as the technology scales from 90nm to 22nm. Compared with conventional repeater insertion methods, our experimental results demonstrate that, at 22nm technology node, this new scheme reduces the normalized delay by 85.1%, the normalized energy consumption by 98.8%. Furthermore, all the performance metrics are scalable as the technology advances, which makes this new signaling scheme a potential candidate to break the “interconnect wall” of digital system performance. Yulei Zhang 0002, Akira Tsuchiya, Masanori Hashimoto, Chung-Kuan Cheng |
ICCD | 4 |
| 2008 | Correlation verification between transistor variability model with body biasing and ring oscillation frequency in 90nm subthreshold circuitsabstractThis paper presents modeling of manufacturing variability and body bias effect for subthreshold circuits based on measurement of a device array circuit in a 90nm technology. The device array consists of P/NMOS transistors and ring oscillators. This work verifies the correlation between the variation model extracted from IV measurement results and oscillation frequencies, which means the transistor-level variation model is examined and confirmed in terms of circuit performance. We demonstrate that delay variations of subthreshold circuits are well characterized with two parameters - threshold voltage and subthreshold swing parameter. We reveal that body bias effect is a less statistical phenomenon and threshold voltage shift by body biasing can be modeled deterministically. Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye |
ISLPED | 2 |
| 2008 | Statistical timing analysis considering spatially and temporally correlated dynamic power supply noiseabstractPower supply noise is becoming more and more influential on timing, though noise aware timing analysis has not been well established yet, because of several difficulties such as its dependency on input vectors and dynamic behavior. This paper proposes a static timing analysis considering power supply noise in which the dependency of noise on input vectors and spatial and temporal correlations are handled in a statistical manner. We construct a statistical model of power supply voltage that dynamically varies with spatial and temporal correlation, and represent it as a set of uncorrelated variables. We demonstrate that power voltage variation is highly correlated and adopting principal component analysis as an orthogonalization technique is effective in variable reduction. Experiments confirm the validity of our model and the accuracy of timing analysis. We also discuss the accuracy and CPU time in association with variable reduction Takashi Enami, Shinyu Ninomiya, Masanori Hashimoto |
ISPD | 3 |
| 2006 | Interconnect RL extraction at a single representative frequencyabstractThis paper proposes a method to determine a single frequency for interconnect RL extraction. Resistance and inductance of interconnects depend on frequency, and hence the extraction frequency strongly affects the modeling accuracy of interconnects. The proposed method determines an extraction frequency based on the transfer characteristic of interconnects. By choosing the frequency where the transfer characteristic becomes maximum, the extracted RL values achieve the accurate modeling of the waveform. We experimentally verify that the proposed method provides accurate transition waveforms over various interconnect topologies. Akira Tsuchiya, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 2 |
| 2006 | A gate delay model focusing on current fluctuation over wide-range of process and environmental variabilityabstractThis paper proposes a gate delay model that is suitable for timing analysis considering wide-range process and environmental variability. The proposed model focuses on current variation and its impact on delay is considered by replacing output load. The proposed model is applicable for large variability with current model constructed by DC analysis whose cost is small. The proposed model can also be used both in statistical static timing analysis and in conventional corner-based static timing analysis. Experimental results in a 90nm technology show that the gate delays of inverter, NAND and NOR are accurately estimated under gate length, threshold voltage, supply voltage and temperature fluctuation. We also verify that the proposed model can cope with slow input transition and RC output load. We demonstrate applicability to multiple-stage path delay and flip-flop delay, and show an application of sensitivity calculation for statistical timing analysis. Kenichi Shinkai, Masanori Hashimoto, Atsushi Kurokawa, Takao Onoye |
ICCAD | 2 |
| 2006 | Quantitative Prediction of On-chip Capacitive and Inductive Crosstalk Noise and Discussion on Wire Cross-Sectional Area Toward Inductive Crosstalk Free InterconnectsabstractCapacitive and inductive crosstalk noises are expected to be more serious in advanced technologies. However, capacitive and inductive crosstalk noises in the future have not been concurrently and sufficiently discussed quantitatively, though capacitive crosstalk noise has been intensively studied solely as a primary factor of interconnect delay variation. This paper quantitatively predicts the impact of capacitive and inductive crosstalk in prospective processes, and reveals that interconnect scaling strategies strongly affect relative dominance between capacitive and inductive coupling. Our prediction also makes the point that the interconnect resistance significantly influences both inductive coupling noise and propagation delay. We then evaluate a tradeoff between wire cross-sectional area and worst-case propagation delay focusing on inductive coupling noise, and show that an appropriate selection of wire cross-section can reduce delay uncertainty by the small sacrifice of propagation delay. Yasuhiro Ogasahara, Masanori Hashimoto, Takao Onoye |
ICCD | 2 |
| 2005 | Timing analysis considering temporal supply voltage fluctuationabstractThis paper proposes an approach to cope with temporal power/ground voltage fluctuation for static timing analysis. The proposed approach replaces temporal noise with an equivalent power/ground voltage. This replacement reduces complexity that comes from the variety in noise waveform shape, and improves compatibility of power/ground noise aware timing analysis with conventional timing analysis framework. Experimental results show that the proposed approach can compute gate propagation delay considering temporal noise within 10% error in maximum and 0.5% in average. Masanori Hashimoto, Junji Yamaguchi, Takashi Sato 0001, Hidetoshi Onodera |
ASP-DAC | 1 |
| 2005 | Successive pad assignment algorithm to optimize number and location of power supply pad using incremental matrix inversionabstractAn efficient pad assignment algorithm to minimize voltage drop on a power distribution network is proposed. Combination of the successive pad assignment (SPA) and the incremental matrix inversion (IMI) provides an efficient assignment for both location and number of power supply pads. The SPA creates equivalent resistance matrix which preserves both pad candidates and power consumption points as external ports so that topological modification due to connection or disconnection between voltage sources and candidate pads are consistently represented. By reusing sub-matrix of equivalent matrix, the SPA greedily searches next pad location that minimizes the worst drop voltage. Each time the candidate pad is added, the IMI reduces computational complexity significantly. Experimental results show that the proposed procedures efficiently enumerate pad order in practical time. Takashi Sato 0001, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 2 |
| 2005 | On-chip thermal gradient analysis and temperature flattening for SoC designabstractThis paper quantitatively analyzes thermal gradient of SoC and proposes a thermal flattening procedure. First, the impact of dominant parameters, such as area occupancy of memory/logic, power density, and floorplan on thermal gradient and clock skew are studied. Important results obtained here are 1) the maximum temperature difference increases with higher memory area occupancy and 2) the difference is very floorplan sensitive. Then, we propose a procedure to amend thermal gradient. A slight floorplan modification using the proposed procedure improves on-chip thermal gradient significantly. Takashi Sato 0001, Junji Ichimiya, Nobuto Ono, Koutaro Hachiya, Masanori Hashimoto |
ASP-DAC | 5 |
| 2005 | Design and measurement of 6.4 Gbps 8: 1 multiplexer in 0.18µm CMOS processabstractWe develop and measure a 8:1 multiplexer in a CMOS 0.18μm process. We design the hybrid multiplexer based on a prior detailed performance evaluation both of CMOS static and current mode logic circuits, and build a hybrid structure. The fabricated chip operates at up to 6.4 Gbps with power consumption of 84mW. Akinori Shinmyo, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 2 |
| 2005 | Return path selection for loop RL extractionabstractThis paper propose a systematic method to select power/ground wires that should be considered in interconnect RL extraction. The return current distribution affects loop characteristic of interconnects. To extract exact RL value, all of return paths have to be considered. However it is impossible because there are huge number of P/G wires in LSIs. As more wires are considered, the extraction accuracy improves but the extraction cost increases undesirably. The proposed method focuses the energy dissipated at P/G wires and utilizes it for screening return paths. Experimental results reveal that our method enables accurate and computationally efficient RL extraction with considering return current distribution. Akira Tsuchiya, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 2 |
| 2005 | Interconnect capacitance extraction for system LCD circuitsabstractThis paper discusses interconnect capacitance extraction for system LCD circuits, where coupling capacitance is much significant since a ground plane locates far away unlike LSI interconnects. We focus on a pattern matching method with interpolation to implement an accurate and efficient capacitance extraction system, and present good implementations that are suitable for system LCD circuits. To reduce computational cost, interconnect structures are spatially divided into several sub-regions considering capacitance coupling range, and analyzed in each sub-region using a capacitance database pre-characterized by a 3-D field solver. This paper evaluates tradeoff curves between characterization cost and extraction accuracy for four division methods in lattice structures that are basic and common structures in LCD driver circuits. Experimental results reveal efficient division methods for accurate capacitance extraction. Yoshihiro Uchida, Sadahiro Tani, Masanori Hashimoto, Shuji Tsukiyama, Isao Shirakawa |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Effects of on-chip inductance on power distribution gridabstractWith increase of clock frequency, on-chip wire inductance starts to play an important role in power/ground distribution analysis, although it has not been considered so far. We perform a case study work that evaluates relation between decoupling capacitance position and noise suppression effect, and we reveal that placing decoupling capacitance close to current load is necessary for noise reduction. We experimentally show that impact of on-chip inductance becomes small when on-chip decoupling capacitance is well placed according to local power consumption. We also examine influences of grid pitch, wire area, and spacing between paired power and ground wires on power supply noise. Minification of grid pitch is more efficient than increase in wire area, and small spacing reduces power noise as we expected. Atsushi Muramatsu, Masanori Hashimoto, Hidetoshi Onodera |
ISPD | 2 |
| 2004 | A performance comparison of PLLs for clock generation using ring oscillator VCO and LC oscillator in a digital CMOS process
Takahito Miyazaki, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 2 |
| 2004 | Representative frequency for interconnect R(f)L(f)C extraction
Akira Tsuchiya, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 2 |
| 2004 | Timing analysis considering spatial power/ground level variationabstractSpatial power/ground level variation causes power/ground level mismatch between driver and receiver, and the mismatch affects gate propagation delay. This work proposes a timing analysis method based on a concept called "PG level equalization" which is compatible with conventional STA frameworks. We equalize the power/ground levels of driver and receiver. The charging/discharging current variation due to equalization is compensated by replacing output load. We present an implementation method of the proposed concept, and demonstrate that the proposed method works well for multiple-input gates and RC load models. Masanori Hashimoto, Junji Yamaguchi, Hidetoshi Onodera |
ICCAD | 1 |
| 2004 | Equivalent waveform propagation for static timing analysisabstractThis paper proposes a scheme that captures diverse input waveforms of CMOS gates for static timing analysis (STA). Conventionally latest arrival and transition times are calculated from the timings when a transient waveform goes across predetermined reference voltages. However, this method cannot accurately consider the impact of waveform shape on gate delay when crosstalk-induced nonmonotonic waveforms or inductance-dominant stepwise waveforms are injected. We propose a new timing analysis scheme called "equivalent waveform propagation." The proposed scheme calculates the equivalent waveform that makes the output waveform close to the actual waveform, and uses the equivalent waveform for timing calculation. The proposed scheme can cope with various waveforms affected by resistive shielding, crosstalk noise, wire inductance, etc. In this paper, we devise a method to calculate the equivalent waveform. The proposed calculation method is compatible with conventional methods in gate delay library and characterization and, hence, our method is easily implemented with conventional STA tools. Masanori Hashimoto, Yuji Yamada, Hidetoshi Onodera |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2003 | Standard cell libraries with various driving strength cells for 0.13, 0.18 and 0.35 μm technologiesabstractWe developed standard cell libraries for three technologies(0.13, 0.18 and 0.35 μm) using an automatic layout generation tool that we have developed. The developed libraries are as competitive as manually-designed libraries in layout density and speed. We verify the functionalities of all cells and the speed of basic combinational cells on the fabricated chips. The libraries are currently public to educational organizations in Japan. Masanori Hashimoto, Kazunori Fujimori, Hidetoshi Onodera |
ASP-DAC | 1 |
| 2003 | Accurate prediction of the impact of on-chip inductance on interconnect delay using electrical and physical parameter-based RSFabstractThis paper proposes a new methodology to accurately predict the impact of inductance on on-chip wire delay using response surface functions (RSF). The proposed methodology consists of two stages which involves first calculating the delay difference between RC and RLC wire models for a set of parameter variations, then building RSFs using electrical parameters such as wire resistance, capacitance, etc., and physical parameters such as wire width, pitch, etc. as variables. The proposed methodology can help 1) to define design rules for avoiding inductance effects, 2) to point out wires that require RLC delay calculation, and 3) to estimate and correct the delay when using an RC model. An example design rule for limiting self inductance and accurate estimation of the delay difference for a 100 nm technology node is also presented. Takashi Sato 0001, Toshiki Kanamoto, Atsushi Kurokawa, Yoshiyuki Kawakami, Hiroki Oka, Tomoyasu Kitaura, Hiroyuki Kobayashi, Masanori Hashimoto |
ASP-DAC | 8 |
| 2003 | Equivalent Waveform Propagation for Static Timing AnalysisabstractThis paper proposes a scheme that captures diverse input waveforms of CMOS gates for static timing analysis. Conventionally the latest arrival time and transition time are calculated from the timings when a transient waveform goes across pre-determined reference voltages. However, this method cannot accurately consider the impact of waveform shape on gate delay, when crosstalk-induced non-monotonic waveforms or inductance-dominant step-wise waveforms are injected. We propose a new timing analysis scheme called "equivalent waveform propagation". The proposed scheme calculates the equivalent waveform that makes the output waveform close to the actual waveform, and uses the equivalent waveform for timing calculation. The proposed scheme can cope with various waveforms affected by resistive shielding, crosstalk noise, wire inductance etc. In this paper, we devise a method to calculate equivalent waveform. The proposed calculation method is compatible with conventional methods in gate delay library and characterization, and hence our method is easy to be implemented with conventional static timing analysis tools. Masanori Hashimoto, Yuji Yamada, Hidetoshi Onodera |
ICCAD | 1 |
| 2003 | Capturing crosstalk-induced waveform for accurate static timing analysisabstractWe propose a method to capture crosstalk-induced noisy waveform for crosstalk-aware static timing analysis. The effects of capacitive coupling noise on timing are conventionally measured as delay variation. On the other hand, the propose method derives an equivalent waveform to a crosstalk-induced noisy waveform. The crosstalk effects on timing are all included in the equivalent waveform. With the derived equivalent waveform, we can perform static timing analysis with consideration of dynamic delay variation due to crosstalk noise. The equivalent waveform is derived by our improved least square fitting with weighting coefficient. Our method can naturally consider the slew variation due to crosstalk noise as well as the delay variation. We experimentally verify that our method can estimate the delay variation at the output of the receiver gate accurately. The strength is that the proposed method requires no additional library characterization and is easy to be integrated into usual static timing analysis methods. Masanori Hashimoto, Yuji Yamada, Hidetoshi Onodera |
ISPD | 1 |
| 2002 | Crosstalk noise optimization by post-layout transistor sizingabstractThis paper proposes a post-layout transistor sizing method for crosstalk noise reduction. The proposed method downsizes the drivers of the aggressor wires for noise reduction, utilizing the precise interconnect information extracted from the detail-routed layouts. We develop a transistor sizing algorithm for crosstalk noise reduction under delay constraints, and construct a crosstalk noise optimization method utilizing a crosstalk noise estimation method and a transistor sizing framework which are previously developed. Our method exploits the transistor sizing framework that can vary the transistor widths inside cells with interconnects unchanged. Our optimization method therefore never cause a new crosstalk noise problem, and does not need iterative layout optimization. The effectiveness of the proposed method is experimentally examined using 2 circuits. The maximum noise voltage is reduced by more than 50% without delay increase. These results show that the risk of crosstalk noise problems can be considerably reduced after detail-routing. Masanori Hashimoto, Masao Takahashi, Hidetoshi Onodera |
ISPD | 1 |
| 2001 | Post-layout transistor sizing for power reduction in cell-based designabstractWe propose a transistor sizing method that down-sizes MOSFETs inside a cell to eliminate redundancy of cell-based circuits as much as possible. Our method reduces power dissipation of detail-routed circuits while preserving interconnects. The effectiveness of our method is experimentally evaluated using 5 circuits. The power dissipation is reduced by 77% maximum and 65% on average without delay increase. Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 1 |
| 2001 | Crosstalk Noise Estimation for Generic RC TreesabstractWe propose an estimation method of crosstalk noise for generic RC trees. The proposed method derives an analytic waveform of crosstalk noise in a 2-/spl pi/ equivalent circuit. The peak voltage is calculated from the closed-form expression, and the crosstalk induced delay is estimated using the derived noise waveform. We also develop a transformation method from generic RC trees with branches into the 2-/spl pi/ model circuit. The proposed method can hence estimate crosstalk noise for any RC trees. Our estimation method is evaluated in a 0.13 /spl mu/m technology. The peak noise of two partially-coupled interconnects is estimated with the average error of 13%. Our method transforms generic RC interconnects with branches into the 2-/spl pi/ model with 14% error on average. Masao Takahashi, Masanori Hashimoto, Hidetoshi Onodera |
ICCD | 2 |
| 2000 | A performance optimization method by gate sizing using statistical static timing analysisabstractWe propose a gate resizing method for delay and power optimization that is based on statistical static timing analysis.Our method focuses on the component of timing uncertainties due to local random fluctuation.Utilizing our method, over-design of a circuit can be eliminated and high-performance and high-reliability LSI design can be realized.The effectiveness of our method is examined by 6 benchmark circuits.We verify that our method can reduce delay and power dissipation from the circuits optimized without the consideration of fluctuation.otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific pennission and/or a fee. Masanori Hashimoto, Hidetoshi Onodera |
ISPD | 1 |
| 1999 | A Practical Gate Resizing Technique Considering Glitch Reduction for Low Power DesignabstractWe propose a method for power optimization that considers glitch reduction by gate sizing based on the statistical estimation of glitch transitions. Our method reduces not only the amount of capacitive and short-circuit power consumption but also the power dissipated by glitches which has not been exploited previously. The effect of our method is verified experimentally using 8 benchmark circuits with a 0.6 m standard cell library. Our method reduces the power dissipation from the minimum-sized circuits further by 9.8 % on average and 23.0 % maximum. We also verify that our method is effective under manufacturing variation. 1 Masanori Hashimoto, Hidetoshi Onodera, Keikichi Tamaru |
DAC | 1 |
| 1998 | A power optimization method considering glitch reduction by gate sizingabstractWe propose a power optimization method considering glitch reduction by gate sizing. Our method reduces not only the amount of capacitive and short-circuit power consumption but also the power dissipated by glitches which has not been exploited previously. In the optimization method, we improve the accuracy of statistical glitch estimation method and device a gate sizing algorithm that utilizes perturbations for escaping a bad local solution. The effect of our method is verified experimentally using 12 benchmark circuits with a 0.5 µm standard cell library. Gate sizing reduces the number of glitch transitions by 38.2 % on average and by 63.4 % maximum. This results in the reduction of total transitions by 12.8 % on average. When the circuits are optimized for power without delay constraints, the power dissipation is reduced by 7.4 % on average and by 15.7 % maximum further from the minimum-sized circuits. Masanori Hashimoto, Hidetoshi Onodera, Keikichi Tamaru |
ISLPED | 1 |