Takumi Uezono

dblp:23/5074 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-1804-2714ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Hardware Error Detection with In-Situ Monitoring of Control Flow-Related Specifications
abstract
In hardware accelerators used in data centers and safety-critical applications, soft errors and resultant silent data corruption significantly compromise reliability, particularly when upsets occur in control-flow operations, leading to severe failures. To address this, we introduce a method for monitoring control flow-related specifications using Petri nets. We validated our method across three designs: convolutional layers in LeNet-5, Gaussian blur in Canny edge detection, and AES encryption. Our fault injection campaign targeting the control registers and primary control inputs demonstrated high error detection rates in both datapath and control logic. Synthesis results show that a maximum detection rate is achieved with a few to around 10 % area overhead in most cases. The proposed detectors quickly detect 88.0% to 99.9% of failures resulting from upsets in internal control registers and perturbation in primary control inputs.
Tomonari Tanaka, Takumi Uezono, Kohei Suenaga, Masanori Hashimoto
ASP-DAC2
2022 Estimating Vulnerability of All Model Parameters in DNN with a Small Number of Fault Injections
abstract
The reliability of deep neural networks (DNNs) against hardware errors is essential as DNNs are increasingly employed in safety-critical applications such as automatic driving. Transient errors in memory, such as radiation-induced soft error, may propagate through the inference computation, resulting in unexpected output, which can adversely trigger catastrophic system failures. As a first step to tackle this problem, this paper proposes constructing a vulnerability model (VM) with a small number of fault injections to identify vulnerable model parameters in DNN. We reduce the number of bit locations for fault injection significantly and develop a flow to incrementally collect the training data, i.e., the fault injection results, for VM accuracy improvement. Experimental results show that VM can estimate vulnerabilities of all DNN model parameters only with 1/3490 computations compared with traditional fault injection-based vulnerability estimation.
Yangchao Zhang, Hiroaki Itsuji, Takumi Uezono, Tadanobu Toba, Masanori Hashimoto
DATE3
2022 Achieving Automotive Safety Requirements through Functional In-Field Self-Test for Deep Learning Accelerators
abstract
Deep learning (DL) accelerators are prominent in automotive systems, and it is essential to guarantee that these accelerators can meet the stringent automotive safety standard even in the presence of various hardware failures. In our previous work [1], we developed an efficient functional in-field self-test generation technique targeting DL accelerators, which achieves high (99.9%) stuck-at fault coverage. In this paper, we present an industry case study that extends our previous work to generate functional in-field self-tests with high transition test coverage, which is critical for screening timing degradation (e.g., caused by circuit aging). We will first present an overview of the general in-vehicle system architecture and discuss reliability/safety requirements and goals. Next, we will discuss the details of our functional in-field self-test generation technique for the transition fault model. Finally, through detailed evaluation on an industrial DL accelerator design, we will show that our approach is able to achieve extremely high transition test coverage (> 99.0%), thereby successfully achieving the reliability/safety requirements for DL accelerators in automotive applications. Moreover, the total in-field self-test time and test storage costs of our technique are low, within the required constraints.
Takumi Uezono, Yi He 0010, Yanjing Li
ITC1
2021 Efficient Functional In-Field Self-Test for Deep Learning Accelerators
abstract
We present a technique that generates high-quality functional in-field self-tests specifically targeting deep learning (DL) accelerators. These functional tests can be applied in the field during normal operation of a DL accelerator, which is crucial to ensure that the safety and/or reliability requirements are met for any given application, including safety-critical applications such as self-driving cars, robotics, and more.Our technique takes advantage of special architectural characteristics and application properties to achieve high functional test coverage while incurring minimal system-level costs. Moreover, we devise different strategies for the compute units (which support computation operations) and the control units (which control data movement) because these two types of units exhibit different properties. For the compute units of a DL accelerator, we first use combinational ATPG to generate test patterns with high test coverage, which is possible because these units do not contain complex sequential logic. Next, we map the ATPG patterns to one or more equivalent deep neural networks (DNNs) that can be directly executed on the accelerator, which is possible given the well-defined dataflow/reuse algorithm of a DL accelerator. For the control units, we leverage the property that typically only one or a few fixed DNNs are deployed at a time in many application domains (e.g., self-driving cars). Thus, it is sufficient to target only the faults that can directly affect the correctness of the DNNs that are currently deployed. This is done by executing different layers of each target DNN using carefully-crafted input and weight values to maximize test coverage while minimizing test time.We apply our technique using Nvidia’s open-source accelerator as a case study to demonstrate its efficacy. Our results show that our technique achieves high test coverage. For the compute units, 99.9% single stuck-at functional test coverage is achieved. For the control units, we are able to prove that, given any target DNN, 100% coverage can be achieved for a large class of single and multiple fault models. The in-field functional self-test time is also very low, < 17 ms for various representative DNNs. These functional tests can be applied during boot-up, reset, and even concurrently with normal operation by executing DNN test programs directly on the accelerator, without requiring any test support in the hardware.
Yi He 0010, Takumi Uezono, Yanjing Li
ITC2
2020 Concurrent Detection of Failures in GPU Control Logic for Reliable Parallel Computing
abstract
The reliability of GPUs is becoming a major concern due to the increased probability of failures and the high vulnerability of GPUs compared to conventional CPUs in terms of tasks per failure. While there are extensive countermeasures against failures in GPU data units, there are fewer countermeasures for failures in GPU control logics. Currently, software-based techniques, such as inserting signature codes for detecting GPU control-logic failures by comparing the expected signature value with the current signature value, are being utilized. However, in the conventional software-based techniques, application calculations, signature calculations, and signature comparison calculations are executed in sequence, which degrades the application throughputs. We have developed a software-based technique that concurrently detects GPU control-logic failures in a running application while largely maintaining its throughput. Experimental results show that when our technique concurrently executed application calculations, signature calculations, and signature comparison calculations for a matrix multiplication application, the application throughput remains 78% of the original one, whereas 62% is reported in literature. We also developed fault injection simulators specialized for injecting GPU-specific control-logic faults into GPU intermediate codes and found that 100% of GPU-specific failures could be detected both during and after application execution. The proposed approach can be utilized for a wide variety of safety-and reliability-critical applications.
Hiroaki Itsuji, Takumi Uezono, Tadanobu Toba, Kojiro Ito, Masanori Hashimoto
ITC2
2016 Evaluation Technique for Soft-Error Rate in Terrestrial Environment Utilizing Low-Energy Neutron Irradiation
abstract
Technology scaling of semiconductor devices improves circuit performance but at the same time degrades a radiation-induced soft-error tolerance. The measurement of soft-error tolerance of devices is one of key technologies to guarantee quality reliability. The conventional measurement methods require high-energy neutron beam. In this paper, we propose a method to measure soft-error rate in terrestrial environment irradiating with a low-energy neutron beam. Our proposed method requires low-energy neutron beam whose energy is less than 40MeV and the measurement cost can be reduced comparing the conventional methods. Our proposed method is applied to FPGAs fabricated in 90nm, 65nm, 40nm, and 28nm processes, and results of our proposed method are compared with those of conventional methods. The comparison results shows the accuracy of our proposed method is comparable with that of conventional ones for the CRAMs of FPGAs fabricated in processes less than 40nm. Therefore, measurement cost reduction can be achieved with our proposed low-energy neutron method.
Takumi Uezono, Tadanobu Toba, Ken-ichi Shimbo, Fumihiko Nagasaki, Kenji Kawamura
ATS1
2016 Path Clustering for Test Pattern Reduction of Variation-Aware Adaptive Path Delay Testing
Michihiro Shintani, Takumi Uezono, Kazumi Hatayama, Kazuya Masu, Takashi Sato 0001
J. Electron. Test.2
2014 A Variability-Aware Adaptive Test Flow for Test Quality Improvement
abstract
In this paper, we propose a process-variability-aware adaptive test flow that realizes efficient and comprehensive detection of parametric faults. A parametric fault is essentially a malfunction in a large-scale integration chip, which is caused by the variability in fabrication processes. In our adaptive test framework, test pattern sets are altered on individual chips in order to apply the optimal set of test patterns for each chip, and thus the test coverage is improved and the test time is reduced. The test pattern is chosen on the basis of parameter estimations measured using an on-chip sensor with respect to statistical timing information. We also propose a novel metric to quantize the test coverage suitable for evaluating the test quality of parametric faults. Our experimental results using an industrial design show that the proposed flow significantly improves the parametric fault coverage and test efficiency compared to conventional test flows.
Michihiro Shintani, Takumi Uezono, Tomoyuki Takahashi, Kazumi Hatayama, Takashi Aikyo, Kazuya Masu, Takashi Sato 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Decomposition of drain-current variation into gain-factor and threshold voltage variations
abstract
A predictable device models should correctly handle parameter variations. Good recognition of the variation of physical parameters, which are being the sources of current variations of modern devices, is thus significantly important. In this paper, we present a practical procedure for decomposing device current variation into physical parameter variations. Based on the I-V curve measurements, two variation components: threshold voltage variation and gain-factor variation are clearly separated. Cause of gain-factor variation is further discussed with measurement results of poly-Si resistor. The impact of the variation-sources on circuit performance is also evaluated using SRAM noise margin as an example.
Takashi Sato 0001, Takumi Uezono, Noriaki Nakayama, Kazuya Masu
ISCAS2
2010 Scan based process parameter estimation through path-delay inequalities
abstract
A novel technique that estimates on-chip process parameters, such as threshold voltages or channel length, is proposed. The proposed method is particularly useful as process condition estimator for reliability and yield enhancement techniques such as adaptive delay test or post-fabric performance compensation. Test paths consisting of a flip-flop and designated delay circuit, which is sensitive to individual process parameters, are inserted to obtain simultaneous delay inequalities. Then, the inequalities are solved for process parameters. The test path insertion is only on short paths to reduce delay and area overhead. Through numerical experiments, the proposed estimation flow using 150 paths achieve 10mV accuracy in estimating threshold voltages.
Takumi Uezono, Tomoyuki Takahashi, Michihiro Shintani, Kazumi Hatayama, Kazuya Masu, Hiroyuki Ochi, Takashi Sato 0001
ISCAS1
2010 Path clustering for adaptive test
abstract
Adaptive test is one of the most efficient techniques that practically ensure high yield and reliability of designed chips. In this paper, a novel path-clustering method suitable for the adaptive test, in which test paths are altered according to the monitored process-parameters, is proposed. Considering the probability function of the die-to-die systematic process variation, the proposed method clusters path sets so that the total number of test-paths are minimized. For quantitative evaluation of different clusterings, figure of merit for clustering, which represents the expected number of test-paths at a particular test coverage, is also proposed. The proposed clustering is experimentally evaluated by applying to an industrial circuit. With our clustering, the average test paths in the adaptive test have been reduced to less than 50% compared with the ones of the conventional test.
Takumi Uezono, Tomoyuki Takahashi, Michihiro Shintani, Kazumi Hatayama, Kazuya Masu, Hiroyuki Ochi, Takashi Sato 0001
VTS1
2009 An Adaptive Test for Parametric Faults Based on Statistical Timing Information
abstract
The continuing miniaturization of LSI dimension is causing the increase of process-related variations which significantly affects not only its design turn around time but also its manufacturing yield. Statistical static timing analysis (SSTA) is expected as a promising way to estimate the performance of circuits more accurately considering delay variations. However, LSIs designed using SSTA may have higher probability of parametric faults than the ones designed with deterministic timing analysis. In order to test these parametric faults, effective extraction techniques of critical paths are needed. In this paper, we discuss a general trend between the delay margin of LSIs designed by SSTA and their parametric fault ratio. Then we propose an adaptive test flow for parametric faults using statistical static timing information, and a concept of parametric fault coverage. Experimental results demonstrate the effectiveness of our approach.
Michihiro Shintani, Takumi Uezono, Tomoyuki Takahashi, Hiroyuki Ueyama, Takashi Sato 0001, Kazumi Hatayama, Takashi Aikyo, Kazuya Masu
Asian Test Symposium2
2007 Improvement of power distribution network using correlation-based regression analysis
abstract
Stochastic approaches for effective power supply network optimization are proposed. Considering node voltages obtained using dynamic voltage drop analysis as sample variables, multi-variate regression is conducted to optimize clock timing metrics, such as clock skew or jitter. Aggregate correlation coefficient (ACC) which quantifies the resistivity between different chip regions is defined in order to find apossible insufficiency in the wire connections of the powersupply network. Based on the ACC, we also propose a procedure using linear regression to find the most effective region for improving clock timing metrics. In our example, clockskew has been reduced by 20% through two iterations.
Shiho Hagiwara, Takumi Uezono, Takashi Sato 0001, Kazuya Masu
ACM Great Lakes Symposium on VLSI2
2005 Evaluation of on-chip transmission line interconnect using wire length distribution
abstract
On-chip transmission-line interconnect has been proposed to reduce delay time and power consumption. The transmission line is used to replace long RC interconnects. This paper proposes the methodology to replace RC lines with transmission lines, which are estimated with Wire Length Distribution (WLD). Advantages of on-chip transmission line are discussed from the view point of delay time and power consumption.
Junpei Inoue, Hiroyuki Ito, Shinichiro Gomi, Takanori Kyogoku, Takumi Uezono, Kenichi Okada 0001, Kazuya Masu
ASP-DAC5