EDBT 2026 Demo / reviewers in the wild / expert
Takashi Sato 0001
dblp:48/4595-1
· DBLP profile ↗
90ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0002-1577-8259ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 81 · 8 first-author · 27 since 2021Software engineering, systems software and programming languages · 8 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Security and privacy · 3 · 2 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LMESN: A Leakage-Driven MOSFET Reservoir for Scalable and Ultra-Low-Power Temporal InferenceabstractEdge-based temporal inference demands energy-efficient and scalable computing architectures, but existing analog reservoir computing models often face high energy costs and limited reconfigurability. We present LMESN, a leakage-current-driven, pulse-based reservoir computing architecture that exploits intrinsic threshold-voltage variation in standard CMOS to realize ultra-low-power stochastic dynamics. To overcome physical array size constraints, we propose a Shift-Multi-Mask (SMM) technique that emulates large virtual reservoirs through cyclic mask shifts, reducing update energy by over $100 \times$ and enabling single-cycle reconfiguration. To further boost task-level performance, we develop a hardware-software co-optimization framework that jointly tunes the ADC quantization range and reservoir mask structure via a discrete genetic algorithm. Post-layout simulations in 22 nm CMOS and evaluations on eight time-series datasets demonstrate up to 13.7% accuracy improvement and $5 \times$ variance reduction over unoptimized LMESN baselines. Compared to prior analog and neural network models, LMESN achieves 3–7 orders of magnitude lower energy consumption while delivering competitive or superior accuracy. Together, these innovations make LMESN a scalable, energy-efficient, and task-adaptive platform for edge temporal processing, setting a new direction in physical reservoir computing. Haoyuan, Masami Utsunomiya, Ryuko Seki, Weirong Dong, Feng Liang 0001, Takashi Sato 0001 |
ASP-DAC | 7 |
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 10 |
| 2026 | Prime Factorization Using Partially Constrained Multiple Quantum Annealing With Analytical and Pattern-Based Variable ReductionabstractFactorization of large semiprimes remains one of the most challenging problems for classical computers. Shor’s algorithm offers a quantum approach that reduces computational complexity, but its practical application is currently limited by hardware constraints. Meanwhile, as a provisional approach, quantum annealing (QA) has been explored through formulations of the quadratic unconstrained binary optimization (QUBO) problem. Among existing methods, the blockwise partial-product approach effectively reduced the QUBO variable count but was limited to semiprimes up to 21 bits. To extend factorization to larger semiprimes, this paper addresses key engineering challenges in constructing efficient QUBO formulations for prime factorization. We propose five techniques to reduce variable counts and improve scalability with current QA hardware: (1) dividing the problem into subproblems with partial constraints; (2) applying analytical reductions near the LSB; (3) applying analytical reductions near the MSB; (4) exploiting special patterns in semiprimes, with odd bit widths and long MSB-side zero sequences; and (5) balancing variable usage across both sides of the subproblem. Integrated into a QUBO converter, these methods enable stable factorization of semiprimes up to 47-bits within 20 seconds and can extend to special 2049-bit instances with 1001 consecutive MSB-side zeros. Geguang Miao, Shinichi Nishizawa, Shinji Kimura, Takashi Sato 0001 |
IEEE Trans. Computers | 5 |
| 2026 | Biologically Constrained DNA Encoding With Triplet Networks for Similarity Image RetrievalabstractAs the volume of digital data continues to grow exponentially, DNA has emerged as a promising medium for long-term data storage due to its high density and durability. For enabling data retrieval via DNA's biochemical reactions, the encoding strategy plays a critical role. This paper proposes a training framework for a DNA encoder that improves both accuracy and training efficiency in content-based image retrieval by incorporating deep metric learning. In addition, we introduce loss functions that enforce biological constraints, specifically homopolymer length and GC content, thereby improving the biochemical stability of the generated DNA sequences. To evaluate the effectiveness of the proposed method, we conduct quantitative assessments based on image classification performance. Simulations on the CIFAR-10 and CIFAR-100 datasets demonstrate that our method achieves classification accuracy comparable to CNN-based baselines and a 20-fold speedup over the training time of the existing method. Moreover, the generated DNA sequences enable strict control of homopolymer length and maintain GC content within the optimal 40-60% range, significantly improving biological feasibility compared to baseline methods. Takefumi Koike, Hiromitsu Awano, Takashi Sato 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Cryo-Compact Modeling Based on Sparse Gaussian ProcessabstractTo enhance the scalability of quantum computers, CMOS circuits operating at cryogenic temperatures (Cryo-CMOS) are being extensively studied for qubit control applications. Designing robust Cryo-CMOS integrated circuits necessitates a transistor model applicable at cryogenic temperatures. Recently, a sparse Gaussian process (SGP)-based modeling method was introduced to consistently model the current characteristics across the room to cryogenic temperatures. However, this method is limited to modeling the current characteristics alone and does not address behaviors in the subthreshold region. In this paper, we extend the SGP-based approach to accurately represent ultra-steep subthreshold slopes at cryogenic temperatures and integrate it into the BSIM-BULK model for the comprehensive representation of transistor behavior. Evaluations using transistors fabricated with 22 nm technology demonstrate that the proposed model effectively simulates the transitions between off-current and on-current at 4 K and that our model can be applicable to predict the DC and transient characteristics of an inverter. Tetsuro Iwasaki, Takashi Sato 0001, Michihiro Shintani |
ASP-DAC | 2 |
| 2025 | Random Telegraph Noise Observed on 65-nm Bulk pMOS Transistors at 3.8KabstractThis paper presents a detailed study on Random Telegraph Noise (RTN) behavior under cryogenic conditions. The study leverages a device array, BTIarray, to statistically measure RTN in a temperature range from room temperature down to 3.8 K. The measurement results indicate that while RTN's impact decreases in the low-temperature region at about 100 K, it becomes more pronounced at lower temperatures, especially in transistors with shorter channel lengths. This research advances the understanding of RTN in cryogenic environments, offering essential insights for future integrated circuit (IC) design. Takuma Kawakami, Takashi Sato 0001, Hiromitsu Awano |
ASP-DAC | 2 |
| 2025 | Physics-based Modeling to Extend a MOSFET Compact Model for Cryogenic OperationabstractThis paper extends the low-temperature modeling capabilities of an industry-standard compact metal-oxide-semiconductor field-effect transistor (MOSFET) model by incorporating physics-based representations of cryogenic effects in semiconductors. Specifically, the incomplete dopant ionization effect is integrated into the bulk Fermi potential calculation of the compact model and applied as a threshold voltage shift in the formulation of Poisson's equation. Temperature-related models for bandgap energy, saturation velocity, and contact resistance at the source/drain regions are also enhanced. Using transistors fabricated with 22 nm process technology, we demonstrate that this consistent modeling approach accurately reproduces current-voltage and threshold voltage-temperature characteristics across a temperature range from 300 K to 4 K. Dondee Navarro, Shin Taniguchi, Chika Tanaka, Kazutoshi Kobayashi, Takashi Sato 0001, Michihiro Shintani |
ASP-DAC | 5 |
| 2025 | Weighted Range-Constrained Ising-Model Decoder for Quantum Error CorrectionabstractIsing model-based Quantum Error Correction decoders reduce topological complexity compared to classical decoders. However, the SOTA Ising decoder has a higher time complexity than union-find (UF) and a lower threshold than minimum-weight perfect-matching (MWPM). We propose the Weighted Range-Constrained Ising Model-Based (WRIM) decoder. WRIM uses a polygonal region to enclose flipped syndromes, ensuring the coverage of all potential error chains while optimizing coupling and external field coefficients. WRIM reduces the variable count by $97.8 x$, achieves microsecondlevel decoding, and has a worst-case time complexity of $O(n)$, outperforming UF. WRIM exhibits threshold behavior up to 10.7$\mathbf{1 1. 0 \%}$, surpassing the MWPM’s highest reported threshold. Hiromitsu Awano, Takashi Sato 0001 |
DAC | 3 |
| 2025 | Lookup Table-based Multiplication-free All-digital DNN Accelerator Featuring Self-Synchronous Pipeline AccumulationabstractDeep neural networks (DNNs) have been widely applied in our society, yet reducing power consumption due to large-scale matrix computations remains a critical challenge. MADDNESS is a known approach to improving energy efficiency by substituting matrix multiplication with table lookup operations. Previous research has employed large analog computing circuits to convert inputs into LUT addresses, which presents challenges to area efficiency and computational accuracy. This paper proposes a novel MADDNESS-based all-digital accelerator featuring a self-synchronous pipeline accumulator, resulting in a compact, energy-efficient, and PVT-invariant computation. Post-layout simulation using a commercial 22nm process showed that 2.5 × higher energy efficiency (174 TOPS/W) and 5× higher area efficiency (2.01 TOPS/mm2) can be achieved compared to the conventional accelerator. Hiroto Tagata, Takashi Sato 0001, Hiromitsu Awano |
DAC | 2 |
| 2025 | SOME: Symmetric One-Hot Matching Elector - A Lightweight Microsecond Decoder for Quantum Error CorrectionabstractConventional quantum error correction (QEC) de-coders such as Minimum-Weight Perfect Matching (MWPM) and Union-Find (UF) offer high thresholds and fast decoding, respectively, but both suffer from high topological complexity. In contrast, Ising model-based decoders reduce topological complexity but demand considerable decoding time. We propose the Symmetric One-Hot Matching Elector (SOME), a novel decoder that reformulates the QEC decoding task as a Quadratic Unconstrained Binary Optimization (QUBO) problem—termed the One-Hot QUBO (OHQ). Each variable in the QUBO represents whether a given pair of flipped syndromes is matched, while the error probabilities between the pair are encoded as interaction coefficients (weight). Constraints ensure that each flipped syndrome is matched exactly once. Valid solutions of OHQ correspond to self-inverse permutation matrices, characterized by symmetric one-hot encoding. To solve the OHQ efficiently, SOME reformulates the decoding task as the construction of permutation matrices that minimize the total weight. It initializes each candidate matrix from one of the minimum-weight syndrome pairs, then iteratively appends additional pairs in ascending order of weight, and finally selects the permutation matrix with the lowest total energy. SOME achieves up to a 99.9x reduction in variable count and reduces decoding times from milliseconds to microseconds on a single-threaded commodity CPU. OHQ also maintains performance up to a 10.5% physical error rate, surpassing the highest known threshold of MWPM. Geguang Miao, Shinichi Nishizawa, Hiromitsu Awano, Shinji Kimura, Takashi Sato 0001 |
ICCAD | 6 |
| 2025 | Zero-Aware Regularization for Energy-Efficient Inference on Akida Neuromorphic ProcessorabstractSpiking Neural Networks (SNNs) and their hardware accelerators have emerged as promising systems for advanced cognitive processing with low power consumption. Although the development of SNN hardware accelerators is particularly active, research on the intelligent use of these accelerators remains limited. This study focuses on the SNN accelerator Akida, a commercially available neuromorphic processor, and presents a novel training method designed to reduce inference energy by leveraging the unique architecture of the hardware. Specifically, we apply sparse constraints on neuron activations and synaptic connection weights, aiming to minimize the number of firing neurons by considering Akida's batch spike processing feature. Our proposed method was applied to a network consisting of three convolutional layers and two fully connected layers. In the MNIST image classification task, the activations became 76.1% sparser, and the weights became 22.1% sparser, resulting in a 13.8% reduction in energy consumption per image. Takehiro Habara, Takashi Sato 0001, Hiromitsu Awano |
ISCAS | 2 |
| 2025 | Gaitcloud: Leveraging Spatial-Temporal Information for Lidar-Base Gait Recognition With a True-3D Gait Representation
Hiromitsu Awano, Takashi Sato 0001 |
WACV | 3 |
| 2025 | Online Training and Inference System on Edge FPGA Using Delayed Feedback ReservoirabstractA delayed feedback reservoir (DFR) is a hardware-friendly reservoir computing system. Implementing DFRs in embedded hardware requires efficient online training. However, two main challenges prevent this: 1) hyperparameter selection, which is typically done by offline grid search, and 2) training of the output linear layer, which is memory-intensive. This article introduces a fast and accurate parameter optimization method for the reservoir layer utilizing backpropagation and gradient descent by adopting a modular DFR model. A truncated backpropagation strategy is proposed to reduce memory consumption associated with the expansion of the recursive structure while maintaining accuracy. The computation time is significantly reduced compared to grid search. In addition, an in-place Ridge regression for the output layer via 1-D Cholesky decomposition is presented, reducing memory usage to be 1/4. These methods enable the realization of an online edge training and inference system of DFR on an FPGA, reducing computation time by about 1/13 and power consumption by about 1/27 compared to software implementation on the same board. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | HOGE: Homomorphic Gate on An FPGAabstractThis paper proposes HOGE, a new accelerator architecture on FPGA, for the Fully Homomorphic Encryption over the Torus (TFHE). TFHE is one of the FHE schemes that allows arbitrary logical circuits to be evaluated over encrypted ciphertexts. To the best of the our knowledge, HOGE is the first hardware architecture that achieves the evaluation of a complete homomorphic gate with bootstrapping in one commercial hardware device and the proven capability for working with the host machine to evaluate homomorphic logic circuits. HOGE is equipped with the carefully designed resource-efficient four-step radix-32 NTT/INTT architectures that achieve higher parallelism, resulting in an overall lower latency. HOGE is implemented on the Xilinx Alveo U280 platform and demonstrated that the actual performance can be 5 – 6 × faster than the-state-of-the-art CPU implementation of TFHE, carrying out a homomorphic gate within about 1.6ms. Kotaro Matsuoka, Song Bian 0001, Takashi Sato 0001 |
ASPDAC | 3 |
| 2024 | Logic Locking over TFHE for Securing User Data and AlgorithmsabstractThis paper proposes the application of logic locking over TFHE to protect both user data and algorithms, such as input user data and models in machine learning inference applications. With the proposed secure computation protocol algorithm evaluation can be performed distributively on honest-but-curious user computers while keeping the algorithm secure. To achieve this, we combine conventional logic locking for untrusted foundries with TFHE to enable secure computation. By encrypting the logic locking key using TFHE, the key is secured with the degree of TFHE. We implemented the proposed secure protocols for combinational logic neural networks and decision trees using LUT-based obfuscation. Regarding the security analysis, we subjected them to the SAT attack and evaluated their resistance based on the execution time. We successfully configured the proposed secure protocol to be resistant to the SAT attack in all machine learning benchmarks. Also, the experimental result shows that the proposed secure computation involved almost no TFHE runtime overhead in a test case with thousands of gates. Kohei Suemitsu, Kotaro Matsuoka, Takashi Sato 0001, Masanori Hashimoto |
ASPDAC | 3 |
| 2024 | Triplet Network-Based DNA Encoding for Enhanced Similarity Image RetrievalabstractWith the exponential growth of digital data, DNA is emerging as an attractive medium for storage and computing. Thus, design methods for encoding, storing, and searching digital data within DNA storage are of utmost importance. This paper introduces image classification as a measurable task for evaluating the performance of DNA encoders in similar image searches. Furthermore, we propose a novel triplet network-based DNA encoder to improve the accuracy and efficiency. The evaluation using the CIFAR-100 dataset demonstrates that the proposed encoder outperforms existing encoders in retrieving similar images, with an accuracy of 0.77, which is equivalent to 94% of the practical upper limit, and 16 times faster training time. Takefumi Koike, Hiromitsu Awano, Takashi Sato 0001 |
DAC | 3 |
| 2024 | Fast Parameter Optimization of Delayed Feedback Reservoir with Backpropagation and Gradient DescentabstractA delayed feedback reservoir (DFR) is a reservoir computing system well-suited for hardware implementations. However, achieving high accuracy in DFRs depends heavily on selecting appropriate hyperparameters. Conventionally, due to the presence of a non-linear circuit block in the DFR, the grid search has only been the preferred method, which is computationally intensive and time-consuming and thus performed offline. This paper presents a fast and accurate parameter optimization method for DFRs. To this end, we leverage the well-known backpropagation and gradient descent framework with the state-of-the-art DFR model for the first time to facilitate parameter optimization. We further propose a truncated backpropagation strategy applicable to the recursive dot-product reservoir representation to achieve the highest accuracy with reduced memory usage. With the proposed lightweight implementation, the computation time has been significantly reduced by up to 1/700 of the grid search. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
DATE | 3 |
| 2024 | DNA-Based Similar Image Retrieval via Triplet Network-Driven EncoderabstractWith the exponential growth of digital data, DNA has emerged as an attractive medium for storage and computing. Design methods for encoding, storing, and searching digital data within DNA storage are thus of utmost importance. This paper introduces image classification as a measurable task for evaluating the performance of DNA encoders in similar image searches. In addition, we propose a triplet network-based DNA encoder to improve accuracy and efficiency. The evaluation using the CIFAR-100 dataset demonstrates that the proposed encoder outperforms existing encoders in retrieving similar images, with an accuracy of 0.77, which is equivalent to 94 % of the practical upper limit, and achieves 16 times faster training time. Takefumi Koike, Hiromitsu Awano, Takashi Sato 0001 |
DATE | 3 |
| 2023 | Invited Paper: Overview of 2023 CAD Contest at ICCADabstractThe “CAD Contest at ICCAD” is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2023 CAD Contest has 210 teams from all over the world, which generates the highest participation record. Moreover, the problems of this year cover state-of-the-art EDA research trends such as circuit verification, hardware security, 3D-IC, and Machine Learning (ML) for EDA from well-known EDA/IC companies. We believe the contest keeps enhancing impact and boosting EDA researches. Takashi Sato 0001, Chun-Yao Wang, Yu-Guang Chen, Tsung-Wei Huang |
ICCAD | 1 |
| 2023 | Improving Efficiency and Robustness of Gaussian Process Based Outlier Detection via Ensemble LearningabstractAlthough automotive semiconductors must comply with the standard dynamic part average testing (DPAT) defined by the Automotive Electronics Council, it remains challenging to detect outliers that deviate from the spatial trend within a wafer. Outlier detection using Gaussian process (GP) regression has recently been proposed and outperformed DPAT. However, the detection performance degrades when faulty large-scale integrations are densely included in the regression. Furthermore, the applicable test items are limited because of the long computation time for regression. We propose an outlier detection method by applying ensemble learning to GP regression for simultaneously improving the detection performance and shortening the learning time. Experimental results on industrial production test data demonstrate that the proposed method improves the robustness against latent faulty chip detection by 15.6% while reducing the computation time by 98.6% compared with the conventional GP-based method. Makoto Eiki, Tomoki Nakamura, Masuo Kajiyama, Michiko Inoue, Takashi Sato 0001, Michihiro Shintani |
ITC | 5 |
| 2023 | Modular DFR: Digital Delayed Feedback Reservoir Model for Enhancing Design FlexibilityabstractA delayed feedback reservoir (DFR) is a type of reservoir computing system well-suited for hardware implementations owing to its simple structure. Most existing DFR implementations use analog circuits that require both digital-to-analog and analog-to-digital converters for interfacing. However, digital DFRs emulate analog nonlinear components in the digital domain, resulting in a lack of design flexibility and higher power consumption. In this paper, we propose a novel modular DFR model that is suitable for fully digital implementations. The proposed model reduces the number of hyperparameters and allows flexibility in the selection of the nonlinear function, which improves the accuracy while reducing the power consumption. We further present two DFR realizations with different nonlinear functions, achieving 10× power reduction and 5.3× throughput improvement while maintaining equal or better accuracy. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Respiratory Rate Estimation Based on WiFi Frame CaptureabstractThis paper presents a method that estimates the respiratory rate based on the frame capturing of wireless local area networks. The method uses beamforming feedback matrices (BFMs) contained in the captured frames, which is a rotation matrix of channel state information (CSI). BFMs are transmitted unencrypted and easily obtained using frame capturing, requiring no specific firmware or WiFi chipsets, unlike the methods that use CSI. Such properties of BFMs allow us to apply frame capturing to various sensing tasks, e.g., vital sensing. In the proposed method, principal component analysis is applied to BFMs to isolate the effect of the chest movement of the subject, and then, discrete Fourier transform is performed to extract respiratory rates in a frequency domain. Experimental evaluation results confirm that the frame-capture-based respiratory rate estimation can achieve estimation error lower than 3.5 breaths/minute. Takamochi Kanda, Takashi Sato 0001, Hiromitsu Awano, Sota Kondo, Koji Yamamoto 0001 |
CCNC | 2 |
| 2022 | Overview of 2022 CAD Contest at ICCADabstractThe "CAD Contest at ICCAD" is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2022 CAD Contest has 166 teams from all over the world. Moreover, the problems of this year cover state-of-the-art EDA research trends such as circuit security, 3D-IC, and design space exploration from well-known EDA/IC companies. We believe the contest keeps enhancing impact and boosting EDA researches. Yu-Guang Chen, Chun-Yao Wang, Tsung-Wei Huang, Takashi Sato 0001 |
ICCAD | 4 |
| 2022 | VisualNet: An End-to-End Human Visual System Inspired Framework to Reduce Inference Latency of Deep Neural NetworksabstractAcceleration of deep neural network (DNN) inference has gained increasing attention recently with the wide adoption of DNNs for practical applications. For computer vision tasks where inputs are images, existing works mostly focus on improving the throughput of inference for multiple images. However, in many real-time applications, it is critical to reduce the latency of a single image inference, which is more complicated than improving the throughput because of the inherent data dependencies. On the other hand, from human brain's perspective, the complexity in our visual surroundings is first encoded as a pattern of light on a two dimensional array of photoreceptors, with little direct resemblance to the original input or the ultimate percept. Within just a few hundred microns of retinal thickness, this initial signal encoded by our photoreceptors must be transformed into an adequate representation of the entire visual scene. Inspired by how the retina helps human brain incept new information efficiently, we present an end-to-end structured framework built using any existing convolutional neural network (CNN) as the backbone. The proposed framework, called VisualNet, can create task parallelism for the backbone during the inference of a single image. Experiments using a number of neural networks for the ImageNet classification task and the CIFAR-10 classification task on GPUs and CPUs show that the proposed VisualNet reduces the latency of the regular network it builds on by up to 80.6% when both are fully parallelized with state-of-the-art acceleration libraries. At the same time, VisualNet can achieve similar or slightly higher accuracy. Jinjun Xiong, Song Bian 0001, Zheyu Yan, Meiping Huang, Jian Zhuang, Takashi Sato 0001, Xiaowei Xu 0004, Yiyu Shi 0001 |
IEEE Trans. Computers | 8 |
| 2022 | Hardware-Friendly Delayed-Feedback Reservoir for Multivariate Time-Series ClassificationabstractReservoir computing (RC) is attracting attention as a machine-learning technique for edge computing. In time-series classification tasks, the number of features obtained using a reservoir depends on the length of the input series. Therefore, the features must be converted to a constant-length intermediate representation (IR), such that they can be processed by an output layer. Existing conversion methods involve computationally expensive matrix inversion that significantly increases the circuit size and requires processing power when implemented in hardware. In this article, we propose a simple but effective IR, namely, dot-product-based reservoir representation (DPRR), for RC based on the dot product of data features. Additionally, we propose a hardware-friendly delayed-feedback reservoir (DFR) consisting of a nonlinear element and delayed feedback loop with DPRR. The proposed DFR successfully classified multivariate time series data that has been considered particularly difficult to implement efficiently in hardware. In contrast to conventional DFR models that require analog circuits, the proposed model can be implemented in a fully digital manner suitable for high-level syntheses. A comparison with existing machine-learning methods via field-programmable gate array implementation using 12 multivariate time-series classification tasks confirmed the superior accuracy and small circuit size of the proposed method. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Efficient Analysis for Mitigation of Workload-Dependent Aging DegradationabstractThe effect of negative bias temperature instability (NBTI) varies significantly according to given workloads. Finding a feasible worst case workload is difficult due to logical correlation within the logic circuit under consideration. In this article, we propose an NBTI-aware subset simulation (SS) framework that efficiently and accurately finds the failure probability covering various input duty cycles determined by different workloads. In addition, the proposed method is incorporated with the NBTI mitigation technique to facilitate workload-aware mitigation. Through numerical experiments using benchmark circuits, the proposed method achieves up to 36 times speedup compared to a naive Monte Carlo method. The NBTI mitigation based on SS demonstrates$1.78\times $better mitigation for multiple input duty cycles compared to the conventional method. Shumpei Morita, Song Bian 0001, Michihiro Shintani, Takashi Sato 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Privacy-Preserving Medical Image Segmentation via Hybrid Trusted Execution EnvironmentabstractRecently, it is reported that the-state-of-the-art secure protocol is able to segment a three-dimensional heart CT scan in roughly 3,000 seconds, without revealing any sensitive information related to the parties involved in the computation. In this work, building upon the existing mix-protocol approach, we make use of the trusted execution environment (TEE) to implement a more efficient privacy-preserving medical image segmentation protocol. In the experiment, we show that by offloading the computations of single-party operators to trusted hardware, the latency for a round of privacy-preserving segmentation can be further reduced by 25×. Song Bian 0001, Weiwen Jiang, Takashi Sato 0001 |
DAC | 3 |
| 2021 | Overview of 2021 CAD Contest at ICCADabstractThe “CAD Contest at ICCAD” is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2021 CAD Contest has 137 teams from all over the world. The contest keeps enhancing impact and boosting EDA research. Tsung-Wei Huang, Yu-Guang Chen, Chun-Yao Wang, Takashi Sato 0001 |
ICCAD | 4 |
| 2021 | Automatic Parallelism Tuning for Module Learning with Errors Based Post-Quantum Key Exchanges on GPUsabstractThe module learning with errors (MLWE) problem is one of the most promising candidates for constructing quantum-resistant cryptosystems. In this work, we propose an open-source framework to automatically adjust the level of parallelism for MLWE-based key exchange protocols to maximize the protocol execution efficiency. We observed that the number of key exchanges handled by primitive functions in parallel, and the dimension of the grids in the GPUs have significant impacts on both the latencies and throughputs of MLWE key exchange protocols. By properly adjusting the related parameters, in the experiments, we show that performance of MLWE based key exchange protocols can be improved across GPU platforms. Tatsuki Ono, Song Bian 0001, Takashi Sato 0001 |
ISCAS | 3 |
| 2021 | Clonable PUF: on the Design of PUFs That Share Equivalent ResponsesabstractWhile numerous physically unclonable functions (PUFs) were proposed in recent years, the conventional PUF- based authentication model is centralized by the data of challenge-response pairs (CRPs), particularly when n-party authentication is required. In this work, we propose a novel concept of clonable PUF (CPUF), wherein two or more PUFs having equivalent responses are manufactured to facilitate decentralized authentication. By design, cloning is only possible in the fabrication period and the responses are determined based on the variability induced during the fabrication. We establish the usage model and the circuit design of CPUFs. Numerical experiments using a circuit simulator show an ideal matching rate of responses between the CPUFs. Takashi Sato 0001, Song Bian 0001 |
ISCAS | 1 |
| 2021 | Virtual Secure Platform: A Five-Stage Pipeline Processor over TFHE
Kotaro Matsuoka, Ryotaro Banno, Naoki Matsumoto, Takashi Sato 0001, Song Bian 0001 |
USENIX Security Symposium | 4 |
| 2021 | APAS: Application-Specific Accelerators for RLWE-Based Homomorphic Linear TransformationsabstractRecently, the application of multi-party secure computing schemes based on homomorphic encryption in the field of machine learning attracts attentions across the research fields. Previous studies have demonstrated that secure protocols adopting packed additive homomorphic encryption (PAHE) schemes based on the ring learning with errors (RLWE) problem exhibit significant practical merits, and are particularly promising in enabling efficient secure inference in machine-learning-as-a-service applications. In this work, we introduce a new technique for performing homomorphic linear transformation (HLT) over PAHE ciphertexts. Using the proposed HLT technique, homomorphic convolutions and inner products can be executed without the use of number theoretic transform and the rotate-and-add algorithms that were proposed in existing works. To maximize the efficiency of the HLT technique, we propose APAS, a hardware-software co-design framework consisting of approximate arithmetic units for the hardware acceleration of HLT. In the experiments, we use actual neural network architectures as benchmarks to show that APAS can improve the computational and communicational efficiency of homomorphic convolution by 8× and 3×, respectively, with an energy reduction of up to 26× as compared to the ASIC implementations of existing methods. Song Bian 0001, Dur-e-Shahwar Kundi, Kazuma Hirozawa, Weiqiang Liu 0001, Takashi Sato 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | A Tuning-Free Hardware Reservoir Based on MOSFET Crossbar Array for Practical Echo State Network ImplementationabstractEcho state network (ESN) is a class of recurrent neural network, and is known for drastically reducing the training time by the use of reservoir, a random and fixed network as the input and middle layers. In this paper, we propose a hardware implementation of ESN that uses practical MOSFET-based reservoir. As opposed to existing reservoirs that require additional tuning of network weights for improved stability, our ESN requires no post-training parameter tuning. To this end, we apply the circular law of random matrix to sparse reservoirs to determine a stable and fixed feedback gain. Through the evaluations using Mackey-Glass time-series dataset, the proposed ESN performs successful inference without post parameter tuning. Yuki Kume, Song Bian 0001, Takashi Sato 0001 |
ASP-DAC | 3 |
| 2020 | Influence of Device Parameter Variability on Current Sharing of Parallel-Connected SiC MOSFETsabstractIn this paper, impact of device parameter variation on current sharing between silicon carbide (SiC) metal-oxide-semiconductor field-effect transistors (MOSFETs) connected in parallel has been studied via Monte Carlo simulation. Paralleled MOSFETs driving an inductive load in a switching circuit, which are expressed by using a surface-potential-based SiC MOSFET model, have been analyzed. From the simulation results, dominant device parameters that affect the mismatch in the SiC MOSFET currents are identified. We also evaluate the effect of current mismatch on the energy loss imbalance. In our analysis, the model parameters related to flat-band voltage, channel length modulation, and current gain factor are found to be particularly important. Yohei Nakamura, Naotaka Kuroda, Atsushi Yamaguchi, Michihiro Shintani, Takashi Sato 0001 |
ATS | 6 |
| 2020 | Measurement of BTI-induced Threshold Voltage Shift for Power MOSFETs under Switching OperationabstractWhile silicon carbide (SiC) MOSFETs can tolerate high-voltage and high-temperature operations with low power loss, long-term reliability of SiC devices is of a concern due to the threshold voltage shift caused by bias temperature instability (BTI). Although the industry-standard BTI characterization method can measure long-term threshold voltage fluctuation, it is hard to quantify the fluctuation under the actual switching operation. In this paper, we propose a long-term BTI characterization method that reflects a realistic switching operation of power devices. We demonstrate the effectiveness of the proposed method using a commercial SiC MOSFET and then discuss the suitable model to represent BTI on the basis of the measurement data. Aoi Ueda, Michihiro Shintani, Michiko Inoue, Takashi Sato 0001 |
ATS | 4 |
| 2020 | ENSEI: Efficient Secure Inference via Frequency-Domain Homomorphic Convolution for Privacy-Preserving Visual RecognitionabstractIn this work, we propose ENSEI, a secure inference (SI) framework based on the frequency-domain secure convolution (FDSC) protocol for the efficient execution of image inference in the encrypted domain. Our observation is that, under the combination of homomorphic encryption and secret sharing, homomorphic convolution can be obliviously carried out in the frequency domain, significantly simplifying the related computations. We provide protocol designs and parameter derivations for number-theoretic transform (NTT) based FDSC. In the experiment, we thoroughly study the accuracy-efficiency trade-offs between time- and frequency-domain homomorphic convolution. With ENSEI, compared to the best known works, we achieve 5--11x online time reduction, up to 33x setup time reduction, and up to 10x reduction in the overall inference time. A further 33% of bandwidth reductions can be obtained on binary neural networks with only 3% of accuracy degradation on the CIFAR-10 dataset. Song Bian 0001, Masayuki Hiromoto, Yiyu Shi 0001, Takashi Sato 0001 |
CVPR | 5 |
| 2020 | Clustering Approach for Solving Traveling Salesman Problems via Ising Model Based SolverabstractIsing model based solver have gained increasing attention due to their efficiency in finding approximate solutions for combinatorial optimization problems. However, when solving doubly constrained problems, such as traveling salesman problem using the Ising model-based solver, both the execution speed and the quality of solutions deteriorate significantly due to the quadratically increasing spin counts and strong constraints placed on the spins. In this paper, we propose a recursive clustering approach that accelerates the calculations of the Ising model and that also helps to obtain high-quality solutions. Through evaluations using the TSP benchmarks, the qualities with the proposed method have been improved by up to 67.1% and runtime were reduced by 73.8x. Akira Dan, Riu Shimizu, Takeshi Nishikawa, Song Bian 0001, Takashi Sato 0001 |
DAC | 5 |
| 2020 | NASS: Optimizing Secure Inference via Neural Architecture SearchabstractDue to increasing privacy concerns, neural network (NN) based secure inference (SI) schemes that simultaneously hide the client inputs and server models attract major research interests. While existing works focused on developing secure protocols for NN-based SI, in this work, we take a different approach. We propose NASS, an integrated framework to search for tailored NN architectures designed specifically for SI. In particular, we propose to model cryptographic protocols as design elements with associated reward functions. The characterized models are then adopted in a joint optimization with predicted hyperparameters in identifying the best NN architectures that balance prediction accuracy and execution efficiency. In the experiment, it is demonstrated that we can achieve the best of both worlds by using NASS, where the prediction accuracy can be improved from 81.6% to 84.6%, while the inference runtime is reduced by 2x and communication bandwidth by 1.9x on the CIFAR-10 dataset. Song Bian 0001, Weiwen Jiang, Qing Lu 0001, Yiyu Shi 0001, Takashi Sato 0001 |
ECAI | 5 |
| 2020 | BUNET: Blind Medical Image Segmentation Based on Secure UNET
Song Bian 0001, Xiaowei Xu 0004, Weiwen Jiang, Yiyu Shi 0001, Takashi Sato 0001 |
MICCAI (2) | 5 |
| 2020 | Ed-PUF: Event-Driven Physical Unclonable Function for Camera Authentication in Reactive Monitoring SystemabstractAs surveillance footage plays an increasingly significant role in law enforcement, it is imperative to ensure the integrity of recorded video data and the authenticity of its originator, and instill situation awareness into these monitoring systems with a fidelity record of the incidents. Unfortunately, existing frame-based networked surveillance systems could only partially fulfill these requirements. The emerging Dynamic Vision Sensor (DVS) sheds new light on solving this problem with its completely different sensor design, i.e., DVS responds only to temporal intensity change and records only sparse asynchronous address-events with precise timing information. Motivated by the reduced data size of activities and the prevention of privacy intrusion of subjects under surveillance as well as other appealing attributes, this work introduces the first event-driven physical unclonable function (Ed-PUF) system to fill the forensic gap of simultaneously authenticating the event data integrity and source camera identity for reactive monitoring by DVS camera. New DVS sensor architecture is proposed with negligible modifications made to the original DVS pixel. The Ed-PUF response bit can only be triggered by and uniquely dependent on the asynchronous addressed event without being interfered by the simultaneous firing of other address events. Address event streams are securely transmitted with an event package tag created by a keyed hash-based message authentication code with the key being the Ed-PUF response. A secure protocol to authenticate the identity of DVS camera and the integrity of address events transmitted through cellular network is also proposed. A camera lock is embedded to protect against severing and splicing the inter-chip connectivity within the camera for raw PUF responses. The proposed system is evaluated using raw PUF data obtained by post-layout Monte Carlo simulation in UMC 180nm technology and real event stream captured by a DVS camera. The proposed Ed-PUF has been demonstrated to have excellent uniqueness, randomness and reliability. Collision test is also conducted to show that the quality of DVS imaging is not compromised. Besides keeping the hardware/power/timing overheads low, the proposed scheme is also analyzed to be resilient against multiple attack scenarios. Xiaojin Zhao, Takashi Sato 0001, Yuan Cao 0003, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Towards practical homomorphic email filtering: a hardware-accelerated secure naïve bayesian filterabstractA secure version of the naïve Bayesian filter (NBF) is proposed utilizing partially homomorphic encryption (PHE) scheme. SNBF can be implemented with only the additive homomorphism from the Paillier system, and we derive new techniques to reduce the computational cost of PHE-based SNBF. In the experiment, we implemented SNBF both in software and hardware. Compared to the best existing PHE scheme, we achieved 1,200x (resp., 398,840x) runtime reduction in the CPU (resp., ASIC) implementations, with additional 1,919x power reduction on the designated hardware multiplier. Our hardware implementation is able to classify an average-length email in 0.5 s, making it one of the most practical NBF schemes to date. Song Bian 0001, Masayuki Hiromoto, Takashi Sato 0001 |
ASP-DAC | 3 |
| 2019 | Filianore: Better Multiplier Architectures for LWE-based Post-Quantum Key ExchangeabstractThe (ring) learning with errors (RLWE/LWE) problem is one of the most promising candidates for constructing quantum-secure key exchange protocols. In this work, we design and implement specialized hardware multiplier units for both LWE and RLWE key exchange schemes to maximize their computational efficiency. By exploiting the algebraic structure with aggressive parameter sets, we show that the design and implementation of LWE key exchange on hardware is considerably easier and more flexible than RLWE. Using the proposed architectures, we show that client-side energy-efficiency of LWE-based key exchange can be on the same order, or even (slightly) better than RLWE-based schemes, making LWE an attractive option for designing post-quantum cryptographic suite. Song Bian 0001, Masayuki Hiromoto, Takashi Sato 0001 |
DAC | 3 |
| 2019 | DArL: Dynamic Parameter Adjustment for LWE-based Secure InferenceabstractPacked additive homomorphic encryption (PAHE) based secure neural network inference is attracting increasing attention in the field of applied cryptography. In this work, we seek to improve the practicality of LWE-based secure inference by dynamically changing the cryptographic parameters depending on the underlying architecture of the neural network. First, we develop and apply theoretical methods to closely examine the error behavior of secure inference, and propose parameters that can reduce as much as 67% of ciphertext size when smaller networks are used. Second, we use rare-event simulation techniques based on the sigma-scale sampling method to provide tight bounds on the size of cumulative errors drawn from (somewhat) arbitrary distributions. Finally, in the experiment, we instantiate an example PAHE scheme and show that we can further reduce the ciphertext size by 3.3x if we adopt a binarized neural network architecture, along with a computation speedup of 2x-3x. Song Bian 0001, Masayuki Hiromoto, Takashi Sato 0001 |
DATE | 3 |
| 2019 | GPU-based Ising computing for solving max-cut combinatorial optimization problems
Chase Cook, Hengyang Zhao, Takashi Sato 0001, Masayuki Hiromoto, Sheldon X.-D. Tan |
Integr. | 3 |
| 2018 | Efficient worst-case timing analysis of critical-path delay under workload-dependent aging degradationabstractThe effect of negative bias temperature instability (NBTI) varies significantly according to given workloads, and thus path delay degradation is strongly dependent on each use case. In this paper, we propose a subset simulation (SS) framework that efficiently and accurately finds the worst case workload and the failure probability covering various workloads. In the proposed method, workloads that yield worst aged delay are efficiently generated by the NBTI-aware Markov chain Monte Carlo method. Through numerical experiments using benchmark circuits, the proposed method achieves up to 36 times speedup compared to the naive Monte Carlo method. From the result of the SS, feasible workload that gives worst aged delay is obtained. Shumpei Morita, Song Bian 0001, Michihiro Shintani, Masayuki Hiromoto, Takashi Sato 0001 |
ASP-DAC | 5 |
| 2018 | DWE: decrypting learning with errors with errorsabstractThe Learning with Errors (LWE) problem is a novel foundation of a variety of cryptographic applications, including quantumly-secure public-key encryption, digital signature, and fully homomorphic encryption. In this work, we propose an approximate decryption technique for LWE-based cryptosystems. Based on the fact that the decryption process for such systems is inherently approximate, we apply hardware-based approximate computing techniques. Rigorous experiments have shown that the proposed technique simultaneously achieved 1.3x (resp., 2.5x) speed increase, 2.06x (resp., 7.89x) area reduction, 20.5% (resp., 4x) of power reduction, and an average of 27.1% (resp., 65.6%) ciphertext size reduction for public-key encryption scheme (resp., a state-of-the-art fully homomorphic encryption scheme). Song Bian 0001, Masayuki Hiromoto, Takashi Sato 0001 |
DAC | 3 |
| 2018 | Ising-PUF: A machine learning attack resistant PUF featuring lattice like arrangement of Arbiter-PUFsabstractA concept of Ising-PUF, a novel PUF structure that utilizes chaotic behavior of mutually interacting small PUFs, is proposed. Ising-PUF consists of a lattice like arrangement of small PUFs, each of which contains a spin register that stores the response of the small PUF, which also serves as a challenge of its neighbors. The spin patterns that develop along time determine the 1-bit response of the Ising-PUF. Utilizing state-memorizing nature of the spin registers, Ising-PUF attains a challenge hysteresis, i.e., allowing sequence of challenge inputs that continuously stimulate its chaotic behavior, which provides the drastically large challenge-to-response space. Experimental results demonstrate nearly ideal metrics; inter-chip Hamming distance (HD) of 50.1% and inter-environment HD of 2.26%. Further, Ising-PUF is remarkably tolerant to machine learning attacks, demonstrating that, even with a deep neural network using a 50k training cRPs, the prediction accuracy remains 50%, which is comparable to a random guess. Hiromitsu Awano, Takashi Sato 0001 |
DATE | 2 |
| 2018 | Enhancing the solution quality of hardware ising-model solver via parallel temperingabstractWe propose an efficient Ising processor with approximated parallel tempering (IPAPT) implemented on an FPGA. Hardware-friendly approximations of the components of parallel tempering (PT) are proposed to enhance solution quality with low hardware overhead. Multiple replicas of Ising states having different temperatures run in parallel by sharing a single network structure, and the replicas are exchanged based on the approximated energy evaluation. The application of PT substantially improves the quality of optimization solutions. The experimental results on the various max-cut problems have shown that utilization of PT significantly increases the probability of obtaining optimal solutions, and IPAPT obtains optimal solutions two orders magnitude faster than a software solver. Hidenori Gyoten, Masayuki Hiromoto, Takashi Sato 0001 |
ICCAD | 3 |
| 2017 | Efficient circuit failure probability calculation along product lifetime considering device agingabstractA device-aging simulation that efficiently estimates temporal degradation of failure probability of a circuit is proposed. As the size of transistors shrinks, consideration of device aging in addition to manufacturing variability has become an urgent issue for maintaining reliability of LSIs. Contrary to existing techniques that separately handle manufacturing variability and the device aging, we propose a simultaneous evaluation approach using an augmented reliability and subset simulation. By eliminating the repetitive failure-probability calculations at each device-age, the proposed method reduces the number of required circuit simulations to about 1/6 of that of the conventional method without compromising accuracy. Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
ASP-DAC | 3 |
| 2017 | Pattern based runtime voltage emergency prediction: An instruction-aware block sparse compressed sensing approachabstractThe relentless technology scaling calls for reduced supply voltage for dynamic power suppression. On the other hand, transistor threshold voltage cannot be scaled at the same pace to avoid excessive leakage power. Consequently, the noise margin is significantly reduced, leading to the deployment of various noise management systems that handle runtime voltage emergencies. Most of these systems rely on on-chip noise sensors, which are large in size and consume significant power. To tackle this issue, in this paper we propose a sensor-less voltage emergency estimation framework. It explores the relationship between switching activities and noise, and takes advantage of block sparse compressed sensing developed by the signal processing society. Experimental results on a few industrial designs show that by monitoring registers, voltage emergencies can be successfully predicted. Yu-Guang Chen, Michihiro Shintani, Takashi Sato 0001, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ASP-DAC | 3 |
| 2017 | LSTA: Learning-Based Static Timing Analysis for High-Dimensional Correlated On-Chip VariationsabstractAs the transistor process technology continues to scale, the aging effect posits new challenges to the already complex static timing analysis (STA) process. In this paper, we first observe that aging can be thought of a type of correlated dynamic on-chip variations (OCV), and identify the problem introduced by such type of OCV. In particular, we take the negative bias temperature instability (NBTI) as an example dynamic OCV mechanism. We then propose a learning-based STA (LSTA) library to "predict" the timing of gates by capturing the correlation between our designed predictors. In the experiment, we used a linear regressor, support vector regression, and a non-linear method, random forest, to create the prediction model. An ISCAS'89 benchmark circuit is used as a training sample to for the algorithms to learn the aging model of gates, and the accuracies of the model is then tested on two processor-scale designs using the library are evaluated, achieving a maximum absolute error of 3.42%. Song Bian 0001, Michihiro Shintani, Masayuki Hiromoto, Takashi Sato 0001 |
DAC | 4 |
| 2017 | SCAM: Secured content addressable memory based on homomorphic encryptionabstractWe propose an implementation of a secured content addressable memory (SCAM) based on homomorphic encryption (HE), where HE is used to compute the word matching function without the processor knowing what is being searched and the result of matching. By exploiting the shallow logic structure (XNOR followed by AND) of content addressable memory (CAM), we show that SCAM can be implemented with only additive homomorphism, greatly improving the efficiency of the HE algorithm. In the proposed method, the logic of homomorphic XNOR-AND is replaced with homomorphic XOR-OR, requiring only simple additions to be performed on the ciphertext. We also show that our scheme can be implemented by highly parallelizable and simple hardware architecture. Through experiment, we demonstrate that our software implementation is already 403x faster than the fastest known algorithm. With the help of hardware, we can achieve an energy reduction per word match by a factor of 477 million times, making our SCAM scheme much more practical. Song Bian 0001, Masayuki Hiromoto, Takashi Sato 0001 |
DATE | 3 |
| 2017 | Scalable Device Array for Statistical Characterization of BTI-Related ParametersabstractA device array circuit, scalable in terms of the number of transistors used, is proposed. The proposed array facilitates accurate and simultaneous bias voltage application to a large number of devices, making it suitable for the measurement-based statistical characterization of device degradation, known as bias temperature instability. Using the proposed array, the degradation measurement of thousands of transistors is made possible in a practical amount of time. The experimental results show that the defect-centric model can approximate the statistical variation in magnitudes of threshold voltage shifts (ΔVTH) and that the variance of ΔVTHbears an inverse relationship to the channel areas of transistors. The degradation variability under ac stress conditions is also presented for the first time. Hiromitsu Awano, Shumpei Morita, Takashi Sato 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | RTN in Scaled Transistors for On-Chip Random Seed GenerationabstractRandom numbers play a vital role in cryptography, where they are used to generate keys, nonce, one-time pads, and initialization vectors for symmetric encryption. The quality of random number generator (RNG) has significant implications on vulnerability and performance of these algorithms. A pseudo-RNG uses a deterministic algorithm to produce numbers with a distribution very similar to uniform. True RNGs (TRNGs), on the other hand, use some natural phenomenon/process to generate random bits. They are nondeterministic, because the next number to be generated cannot be determined in advance. In this paper, a novel on-chip noise source, random telegraph noise (RTN), is exploited for simple and reliable TRNG. RTN, a microscopic process of stochastic trapping/detrapping of charges, is usually considered as a noise and mitigated in design. Through physical modeling and silicon measurement, we demonstrate that RTN is appropriate for TRNG, especially in highly scaled MOSFETs. Due to the slow speed of RTN, we purpose the system for on-chip seed generation for random number. Our contributions are: 1) physical model calibration of RTN with comprehensive 65- and 180-nm transistor measurements; 2) the scaling trend of RTN, validated with silicon data down to 28 nm; 3) design principles to achieve 50% signal probability by using intrinsic RTN physical properties, without traditional postprocessing algorithms, the generated sequence passes the National Institute of Standards and Technology (NIST) tests; and 4) solutions to manage realistic issues in practice, including multilevel RTN signal, robustness to voltage and temperature fluctuations and the operation speed. Abinash Mohanty, Ketul Sutaria, Hiromitsu Awano, Takashi Sato 0001, Yu Cao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Runtime NBTI Mitigation for Processor Lifespan Extension via Selective Node ControlabstractNegative bias temperature instability (NBTI) has become one of the major reliability concerns for nanoscale CMOS technology. The NBTI effect degrades pMOS transistors by stressing them with negatively biased voltage, while the transistors heal themselves as the negative bias is removed. In this paper, we propose a cross-layer mitigation technique for NBTI-induced timing degradation in processors. The NOP (No Operation) instruction is replaced by a custom NOP instruction for healing purpose. Cells that are likely to be stressed under negative bias are classified and their upstream cell will be replaced by the internal node control (INC) logics. Upon encountering a custom NOP instruction, the INC logics will force the NBTI-stressed cell to be in its healing mode. The optimal INC logic insertion through genetic programming approach achieves much greater delay mitigation of 44.3% than prior works in a 10-year span with less than 4% of power and negligible area overhead. Song Bian 0001, Michihiro Shintani, Zheng Wang 0020, Masayuki Hiromoto, Anupam Chattopadhyay, Takashi Sato 0001 |
ATS | 6 |
| 2016 | Efficient transistor-level timing yield estimation via line samplingabstractYield estimation has been and will continue to be the integral part in design flow, particularly under large process variability in advanced technology nodes. This paper proposes an efficient method that accelerates transistor-level statistical timing simulations which are intensively conducted in various design stages, such as in final timing verification, while considering thousands of random variables. The proposed method utilizes line sampling (LS), in which integration of randomly generated lines, not the random points, are evaluated. Numerical experiments show that the proposed method achieves 14× to 300× speed-up compared to the fastest one that has ever reported. Hiromitsu Awano, Takashi Sato 0001 |
DAC | 2 |
| 2016 | Workload-Aware Worst Path Analysis of Processor-Scale NBTI DegradationabstractAs technology further scales semiconductor devices, aging-induced device degradation has become one of the major threats to device reliability. In addition, aging mechanisms like the negative bias temperature instability (NBTI) is known to be sensitive to workload (i.e., signal probability) that is hard to be assumed at design phase. In this work, we analyze the workload dependence of NBTI degradation using a processor, and propose a novel technique to estimate the worst-case paths. In our approach, with careful examination, we exploit the fact that the deterministic nature of circuit structure limits the amount of NBTI degradation on different paths, and proposes a two-stage path extraction algorithm to identify the invariable critical paths in the processor. Through numerical experiment on a MIPS32 processor, we performed a detailed signal probability analysis, and successfully extracted 85 invariable critical paths out of the 24,978 path candidates, achieving nearly 300x reduction in the sheer number of paths. Song Bian 0001, Michihiro Shintani, Shumpei Morita, Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2016 | Physically unclonable function using RTN-induced delay fluctuation in ring oscillatorsabstractThis paper proposes RTN-PUF, a novel PUF that utilizes random telegraph noise (RTN) of transistors as the physical uniqueness of individual devices. Our proposed RTN-PUF generates a response from a pair of ring oscillators (ROs) by comparing the numbers of frequency changes, which depend on the time constants of RTN. Due to the log-uniform distribution of the time constants, our RTN-PUF provides more stable responses than the existing manufacturing-variation-based PUFs. The numerical experiments show that the RTN-PUF reduces false negative errors by about 60 times compared to the conventional RO-based PUF. This facilitates to implement PUF into security purposes. Motoki Yoshinaga, Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
ISCAS | 4 |
| 2016 | Path Clustering for Test Pattern Reduction of Variation-Aware Adaptive Path Delay Testing
Michihiro Shintani, Takumi Uezono, Kazumi Hatayama, Kazuya Masu, Takashi Sato 0001 |
J. Electron. Test. | 5 |
| 2015 | ECRIPSE: an efficient method for calculating RTN-induced failure probability of an SRAM cell
Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
DATE | 3 |
| 2015 | Introduction to: Special Issue on Cross-Layer System DesignabstractNo abstract available. Yiyu Shi 0001, Takashi Sato 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2014 | Sensorless estimation of global device-parameters based on Fmax testingabstractPost-fabrication performance compensation and adaptive delay testing are indispensable means for improving yield and reliability of LSIs, The global parameter estimations, such as of threshold voltages, play a key role in maximizing their effectiveness. This paper proposes a novel technique that realizes an accurate device-parameter estimation through Fmaxtesting framework. In the proposed method, statistical path delay distributions of sensitized paths in Fmaxtesting are utilized to calculate device-parameters, such that they most likely explain the measurements in the Fmaxtesting. Two estimation procedures are proposed: one utilizes discrete Bayesian estimation and the other uses maximum likelihood estimation. Numerical experiments demonstrate that both methods achieve 2.5mV accuracy in estimating threshold voltages. Michihiro Shintani, Takashi Sato 0001 |
ICCAD | 2 |
| 2014 | A Variability-Aware Adaptive Test Flow for Test Quality ImprovementabstractIn this paper, we propose a process-variability-aware adaptive test flow that realizes efficient and comprehensive detection of parametric faults. A parametric fault is essentially a malfunction in a large-scale integration chip, which is caused by the variability in fabrication processes. In our adaptive test framework, test pattern sets are altered on individual chips in order to apply the optimal set of test patterns for each chip, and thus the test coverage is improved and the test time is reduced. The test pattern is chosen on the basis of parameter estimations measured using an on-chip sensor with respect to statistical timing information. We also propose a novel metric to quantize the test coverage suitable for evaluating the test quality of parametric faults. Our experimental results using an industrial design show that the proposed flow significantly improves the parametric fault coverage and test efficiency compared to conventional test flows. Michihiro Shintani, Takumi Uezono, Tomoyuki Takahashi, Kazumi Hatayama, Takashi Aikyo, Kazuya Masu, Takashi Sato 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2013 | Realization of frequency-domain circuit analysis through random walkabstractThis paper presents the realization of frequency-domain circuit analysis based on random walk framework for the first time. In conventional random walk based circuit analyses, the sample movement at a node is randomly chosen to follow the edge probabilities. The probabilities are determined by edge-admittances connecting to the node, which is impossible to apply for the frequency-domain analysis because the probabilities are imaginary numbers. By applying the idea of importance sampling, the intractable imaginary probabilities are converted into real numbers while maintaining the estimation correctness. Runtime acceleration through incremental analysis is also proposed. Tetsuro Miyakawa, Hiroshi Tsutsui, Hiroyuki Ochi, Takashi Sato 0001 |
ASP-DAC | 4 |
| 2013 | An adaptive current-threshold determination for IDDQ testing based on Bayesian process parameter estimationabstractApplication of IDDQ testing to LSIs fabricated using advanced process technology is becoming increasingly difficult due to large variability of scaled devices. In this paper, we propose a novel technique that adaptively determines per-chip current-threshold for IDDQ testing to enhance test accuracy. In the proposed technique, process condition of a chip and fault-sensitization vector are first estimated based on measured IDDQ currents through Bayesian inference. Then, using the estimated process condition, a statistical distribution of the leakage current for each test pattern is calculated and suitable current-threshold is determined by the distribution. Simulation experiments demonstrate that the proposed technique can successfully detect a very small leakage fault, down to 16% of the nominal IDDQ current with the test escape ratio of 3.1 %. Michihiro Shintani, Takashi Sato 0001 |
ASP-DAC | 2 |
| 2013 | A cost-effective selective TMR for heterogeneous coarse-grained reconfigurable architectures based on DFG-level vulnerability analysisabstractThis paper proposes a method to determine a priority for applying selective triple modular redundancy (selective TMR) against single event upset (SEU) to achieve cost-effective reliable implementation of an application circuit to a coarse-grained reconfigurable architecture (CGRA). The priority is determined by an estimation of the vulnerability of each node in the data flow graph (DFG) of the application circuit. The estimation is based on a weighted sum of the features and parameters of each node in the DFG which characterize impact of the SEU in the node to the output data. This method does not require time-consuming placement-and-routing processes, as well as extensive fault simulations for various triplicating patterns, which allows us to identify the set of nodes to be triplicated for minimizing the vulnerability under given area constraint at the early stage of design flow. Therefore, the proposed method enables us efficient design space exploration of reliability-oriented CGRAs and their applications. Takashi Imagawa, Hiroshi Tsutsui, Hiroyuki Ochi, Takashi Sato 0001 |
DATE | 4 |
| 2013 | Hot-swapping architecture with back-biased testing for mitigation of permanent faults in functional unit arrayabstractDue to latest advances in semiconductor integration, systems are becoming more susceptible to faults leading to temporary or permanent failures. We propose a new architecture extension suitable for arrays of functional units (FUs), that will provide testing and replacement of faulty units, without interrupting normal system operation. The extension relies on data-path switching realized by the proposed hot-swapping algorithm and structures, by use of which functional units are tested and replaced by spares, at lower overheads than traditional modular redundancy. For a case study architecture, hot-swapping support could be added with only 29% area overhead. In this paper we focus on experimental evaluation of the hot-swapping system from a fabricated chip in 65nm CMOS process. Autonomous testing of the hot-swapping system is enhanced with back-bias circuitry to attain an early fault detection and restoration system. Experimental measurements prove that the proposed concept works well, predicting fault occurrence with a configurable prediction interval, while power measurements reveal that with only 20% power overhead the proposed system can attain reliability levels similar to triple modular redundancy. Additionally, measurements reveal that manufacturing randomness across the die can significantly influence identical sub-circuit reliability located in different parts in the die, although identical layout has been employed. Zoltán Endre Rákossy, Masayuki Hiromoto, Hiroshi Tsutsui, Takashi Sato 0001, Yukihiro Nakamura, Hiroyuki Ochi |
DATE | 4 |
| 2013 | Fast and memory-efficient GPU implementations of krylov subspace methods for efficient power grid analysisabstractPower grid analysis for modern LSI is computationally challenging in terms of both runtime and memory usage. In this paper, we implement Krylov subspace based linear circuit solvers on a graphics processing unit (GPU) to realize fast power grid analysis. Efficiencies of memory space and access performance are pursued by improving a data structure that stores elements of large sparse matrices. Experimental results on benchmark circuits show that the proposed data structures are more suitable than widely used compressed sparse row (CSR) format and our GPU implementations can achieve up to 17x speedup over CPU implementations. Takumi Morishita, Hiroshi Tsutsui, Hiroyuki Ochi, Takashi Sato 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2012 | Physics matters: statistical aging prediction under trapping/detrappingabstractRandomness in Negative Bias Temperature Instability (NBTI) process poses a dramatic challenge on reliability prediction of digital circuits. Accurate statistical aging prediction is essential in order to develop robust guard banding and protection strategies during the design stage. Variations in device level and supply voltage due to Dynamic Voltage Scaling (DVS) need to be considered in aging analysis. The statistical device data collected from 65nm test chip shows that degradation behavior derived from trapping/detrapping mechanism is accurate under statistical variations compared to conventional Reaction Diffusion (RD) theory. The unique features of this work include (1) Aging model development as a function of technology parameters based on trapping/detrapping theory (2) Reliability prediction under device variations and DVS with solid validation with using 65nm statistical silicon data (3) Asymmetric aged timing analysis under NBTI and comprehensive evaluation of our framework in ISCAS89 sequential circuits. Further, we show that RD based NBTI model significantly overestimates the degradation and TD model correctly captures aging variability. These results provide design insights under statistical NBTI aging and enhance the prediction efficiency. Jyothi Velamala, Ketul Sutaria, Takashi Sato 0001, Yu Cao 0001 |
DAC | 3 |
| 2012 | A Bayesian-based process parameter estimation using IDDQ current signatureabstractPost-fabrication performance compensation and adaptive delay testing are effective means to improve yield and reliability of LSIs. In these methods, process parameter estimation plays a key role. In this paper, we propose a novel technique for accurate on-chip process parameter estimation. The proposed technique is based on Bayes' theorem, in which on-chip parameters, such as threshold voltages, are estimated by current signatures obtained within a regular IDDQ testing. No additional circuit and additional measurements are required for the purpose of estimation. Numerical experiments demonstrate that the proposed technique can achieve less than 10 mV accuracy in estimating threshold voltages. Michihiro Shintani, Takashi Sato 0001 |
VTS | 2 |
| 2011 | Acceleration of random-walk-based linear circuit analysis using importance samplingabstractThis paper proposes an importance sampling (IS) technique based on quasi-zero-variance estimation for accelerating convergence of random-walk-based power grid analysis. In our approach, the alternative probability for IS is incrementally updated after every Mr samples of random walk so that more recent and thus more accurate node voltages are utilized to asymptotically achieve ideal zero-variance estimation. We also propose a method to determine efficient Mr for the r-th probability update; although smaller Mr results more aggressive update of alternative probability, the alternative probability becomes inaccurate if Mr is too small. The estimation error of the proposed method decreases O((M/r)-r/2), which breaks O(M-1/2), the slow convergence-rate barrier of normal Monte Carlo analysis. Our trial implementation achieved 790x speedup compared with a conventional random-walk-based circuit analysis for analyzing IBM power grid benchmark circuits at 1mV accuracy. Tetsuro Miyakawa, Koh Yamanaga, Hiroshi Tsutsui, Hiroyuki Ochi, Takashi Sato 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2010 | Sequential importance sampling for low-probability and high-dimensional SRAM yield analysisabstractIn this paper, a significant acceleration of estimating low-failure rate in a high-dimensional SRAM yield analysis is achieved using sequential importance sampling. The proposed method systematically, autonomously, and adaptively explores failure region of interest, whereas all previous works needed to resort to brute-force search. Elimination of brute-force search and adaptive trial distribution significantly improves the efficiency of failure-rate estimation of hitherto unsolved high-dimensional cases wherein a lot of variation sources including threshold voltages, channel-length, carrier mobility, etc. are simultaneously considered. The proposed method is applicable to wide range of Monte Carlo simulation analyses dealing with high-dimensional problem of rare events. In SRAM yield estimation example, we achieved 106times acceleration compared to a standard Monte Carlo simulation for a failure probability of 3 × 10-9in a six-dimensional problem. The example of 24-dimensional analysis on which other methods are ineffective is also presented. Kentaro Katayama, Shiho Hagiwara, Hiroshi Tsutsui, Hiroyuki Ochi, Takashi Sato 0001 |
ICCAD | 5 |
| 2010 | Decomposition of drain-current variation into gain-factor and threshold voltage variationsabstractA predictable device models should correctly handle parameter variations. Good recognition of the variation of physical parameters, which are being the sources of current variations of modern devices, is thus significantly important. In this paper, we present a practical procedure for decomposing device current variation into physical parameter variations. Based on the I-V curve measurements, two variation components: threshold voltage variation and gain-factor variation are clearly separated. Cause of gain-factor variation is further discussed with measurement results of poly-Si resistor. The impact of the variation-sources on circuit performance is also evaluated using SRAM noise margin as an example. Takashi Sato 0001, Takumi Uezono, Noriaki Nakayama, Kazuya Masu |
ISCAS | 1 |
| 2010 | Scan based process parameter estimation through path-delay inequalitiesabstractA novel technique that estimates on-chip process parameters, such as threshold voltages or channel length, is proposed. The proposed method is particularly useful as process condition estimator for reliability and yield enhancement techniques such as adaptive delay test or post-fabric performance compensation. Test paths consisting of a flip-flop and designated delay circuit, which is sensitive to individual process parameters, are inserted to obtain simultaneous delay inequalities. Then, the inequalities are solved for process parameters. The test path insertion is only on short paths to reduce delay and area overhead. Through numerical experiments, the proposed estimation flow using 150 paths achieve 10mV accuracy in estimating threshold voltages. Takumi Uezono, Tomoyuki Takahashi, Michihiro Shintani, Kazumi Hatayama, Kazuya Masu, Hiroyuki Ochi, Takashi Sato 0001 |
ISCAS | 7 |
| 2010 | Path clustering for adaptive testabstractAdaptive test is one of the most efficient techniques that practically ensure high yield and reliability of designed chips. In this paper, a novel path-clustering method suitable for the adaptive test, in which test paths are altered according to the monitored process-parameters, is proposed. Considering the probability function of the die-to-die systematic process variation, the proposed method clusters path sets so that the total number of test-paths are minimized. For quantitative evaluation of different clusterings, figure of merit for clustering, which represents the expected number of test-paths at a particular test coverage, is also proposed. The proposed clustering is experimentally evaluated by applying to an industrial circuit. With our clustering, the average test paths in the adaptive test have been reduced to less than 50% compared with the ones of the conventional test. Takumi Uezono, Tomoyuki Takahashi, Michihiro Shintani, Kazumi Hatayama, Kazuya Masu, Hiroyuki Ochi, Takashi Sato 0001 |
VTS | 7 |
| 2010 | Modeling the Overshooting Effect for CMOS Inverter Delay Analysis in Nanometer TechnologiesabstractWith the scaling of complementary metal-oxide-semiconductor (CMOS) technology into the nanometer regime, the overshooting effect due to the input-to-output coupling capacitance has more significant influence on CMOS gate analysis, especially on CMOS gate static timing analysis. In this paper, the overshooting effect is modeled for CMOS inverter delay analysis in nanometer technologies. The results produced by the proposed model are close to simulation program with integrated circuit emphasis (SPICE). Moreover, the influence of the overshooting effect on CMOS inverter analysis is discussed. An analytical model is presented to calculate the CMOS inverter delay time based on the proposed overshooting effect model, which is verified to be in good agreement with SPICE results. Furthermore, the proposed model is used to improve the accuracy of the switch-resistor model for approximating the inverter output waveform. Zhangcai Huang, Atsushi Kurokawa, Masanori Hashimoto, Takashi Sato 0001, Minglu Jiang, Yasuaki Inoue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2009 | An Adaptive Test for Parametric Faults Based on Statistical Timing InformationabstractThe continuing miniaturization of LSI dimension is causing the increase of process-related variations which significantly affects not only its design turn around time but also its manufacturing yield. Statistical static timing analysis (SSTA) is expected as a promising way to estimate the performance of circuits more accurately considering delay variations. However, LSIs designed using SSTA may have higher probability of parametric faults than the ones designed with deterministic timing analysis. In order to test these parametric faults, effective extraction techniques of critical paths are needed. In this paper, we discuss a general trend between the delay margin of LSIs designed by SSTA and their parametric fault ratio. Then we propose an adaptive test flow for parametric faults using statistical static timing information, and a concept of parametric fault coverage. Experimental results demonstrate the effectiveness of our approach. Michihiro Shintani, Takumi Uezono, Tomoyuki Takahashi, Hiroyuki Ueyama, Takashi Sato 0001, Kazumi Hatayama, Takashi Aikyo, Kazuya Masu |
Asian Test Symposium | 5 |
| 2008 | Determination of optimal polynomial regression function to decompose on-die systematic and random variationsabstractA procedure that decomposes measured parametric device variation into systematic and random components is studied by considering the decomposition process as selecting the most suitable model for describing on-die spatial variation trend. In order to maximize model predictability, the log-likelihood estimate called corrected Akaike information criterion is adopted. Depending on on-die contours of underlying systematic variation, necessary and sufficient complexity of the systematic regression model is objectively and adaptively determined. The proposed procedure is applied to 90-nm threshold voltage data and found the low order polynomials describe systematic variation very well. Designing cost-effective variation monitoring circuits as well as appropriate model determination of on-die variation are hence facilitated. Takashi Sato 0001, Hiroyuki Ueyama, Noriaki Nakayama, Kazuya Masu |
ASP-DAC | 1 |
| 2008 | Non-parametric statistical static timing analysis: an SSTA framework for arbitrary distributionabstractWe present a new statistical STA framework based on Monte Carlo analysis that can deal with arbitrary statistical distribution and delay models. Order statistics (non-parametrics) is consistently adopted by which the timing analysis and criticality calculation become distribution-independent. To make Monte Carlo process computationally practical, delays are handled as vectors so that iterations are eliminated. The vector dimension or required number of Monte Carlo iterations which guarantees no timing violation at any user-specified probability is analytically determined. A path criticality metric using order statistics is also defined. Experimental results using various delay models show the validity and usefulness of our proposed algorithm. Masanori Imai, Takashi Sato 0001, Noriaki Nakayama, Kazuya Masu |
DAC | 2 |
| 2008 | Decoupling capacitance allocation for timing with statistical noise model and timing analysisabstractThis paper presents an allocation method of decoupling capacitance that explicitly considers timing. We have found and focused that decap does not necessarily improve a gate delay at all the switching timing within a cycle, and devised an efficient sensitivity calculation of timing to decap for decap allocation. The proposed method, which is based on a statistical noise modeling and timing analysis, accelerates the sensitivity calculation with an approximation and adjoint sensitivity analysis. Experimental results show that the decap allocation based on the sensitivity analysis efficiently optimizes the worst-case circuit delay within a given decap budget. Compared to the maximum decap placement, the delay improvement due to decap increases by 5% even while the total amount of decap is reduced to 40%. Takashi Enami, Masanori Hashimoto, Takashi Sato 0001 |
ICCAD | 3 |
| 2007 | A Multi-Drop Transmission-Line Interconnect in Si LSIabstractThis paper proposes a branching method for on-chip transmission line (TL) interconnects, which can reduce delay and power of global interconnects. A 6-mm-long TL interconnect with a branch is fabricated by using a 0.18 mum standard Si CMOS process, and the measurement result performs 4Gbps signal transmission. Junki Seita, Hiroyuki Ito, Kenichi Okada 0001, Takashi Sato 0001, Kazuya Masu |
ASP-DAC | 4 |
| 2007 | Improvement of power distribution network using correlation-based regression analysisabstractStochastic approaches for effective power supply network optimization are proposed. Considering node voltages obtained using dynamic voltage drop analysis as sample variables, multi-variate regression is conducted to optimize clock timing metrics, such as clock skew or jitter. Aggregate correlation coefficient (ACC) which quantifies the resistivity between different chip regions is defined in order to find apossible insufficiency in the wire connections of the powersupply network. Based on the ACC, we also propose a procedure using linear regression to find the most effective region for improving clock timing metrics. In our example, clockskew has been reduced by 20% through two iterations. Shiho Hagiwara, Takumi Uezono, Takashi Sato 0001, Kazuya Masu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Timing analysis considering temporal supply voltage fluctuationabstractThis paper proposes an approach to cope with temporal power/ground voltage fluctuation for static timing analysis. The proposed approach replaces temporal noise with an equivalent power/ground voltage. This replacement reduces complexity that comes from the variety in noise waveform shape, and improves compatibility of power/ground noise aware timing analysis with conventional timing analysis framework. Experimental results show that the proposed approach can compute gate propagation delay considering temporal noise within 10% error in maximum and 0.5% in average. Masanori Hashimoto, Junji Yamaguchi, Takashi Sato 0001, Hidetoshi Onodera |
ASP-DAC | 3 |
| 2005 | Successive pad assignment algorithm to optimize number and location of power supply pad using incremental matrix inversionabstractAn efficient pad assignment algorithm to minimize voltage drop on a power distribution network is proposed. Combination of the successive pad assignment (SPA) and the incremental matrix inversion (IMI) provides an efficient assignment for both location and number of power supply pads. The SPA creates equivalent resistance matrix which preserves both pad candidates and power consumption points as external ports so that topological modification due to connection or disconnection between voltage sources and candidate pads are consistently represented. By reusing sub-matrix of equivalent matrix, the SPA greedily searches next pad location that minimizes the worst drop voltage. Each time the candidate pad is added, the IMI reduces computational complexity significantly. Experimental results show that the proposed procedures efficiently enumerate pad order in practical time. Takashi Sato 0001, Masanori Hashimoto, Hidetoshi Onodera |
ASP-DAC | 1 |
| 2005 | On-chip thermal gradient analysis and temperature flattening for SoC designabstractThis paper quantitatively analyzes thermal gradient of SoC and proposes a thermal flattening procedure. First, the impact of dominant parameters, such as area occupancy of memory/logic, power density, and floorplan on thermal gradient and clock skew are studied. Important results obtained here are 1) the maximum temperature difference increases with higher memory area occupancy and 2) the difference is very floorplan sensitive. Then, we propose a procedure to amend thermal gradient. A slight floorplan modification using the proposed procedure improves on-chip thermal gradient significantly. Takashi Sato 0001, Junji Ichimiya, Nobuto Ono, Koutaro Hachiya, Masanori Hashimoto |
ASP-DAC | 1 |
| 2004 | Probabilistic crosstalk delay estimation for ASICsabstractThe crosstalk delay caused by capacitive coupling between wires on a chip is investigated by using a statistical approach and circuit simulations. Two metrics are introduced in order to evaluate an impact of the crosstalk delay on timing design in advance. The first is probabilistic coupling rate (CPR), which can be obtained by the short segment model of the aggressors. Then, the CPR roughly obeys normal distribution and its standard deviation is determined by the slew time of the victim along with the number of aggressor segments. The second is crosstalk delay normalized by the original delay without crosstalk, /spl Delta/t/sub pd//t/sub pd/. The /spl Delta/t/sub pd//t/sub pd/ is equal to 2*CPR at the maximum, and CPR on average, regardless of victim length. The two metrics in conjunction with empirical slew distribution allows us to set the appropriate crosstalk delay budget, at the prelayout stage, for reducing the possibility of the crosstalk violation found in the postlayout verification process. Kan Takeuchi, Kazumasa Yanagisawa, Takashi Sato 0001, Kazuko Sakamoto, Saburo Hojo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2003 | Approximate formulae approach for efficient inductance extractionabstractIn this paper, we present a new and effective approach to the extraction of on-chip inductance, in which we apply approximate formulae. The equations are based on the assumption of filaments or bars of finite width and zero thickness and are derived through Taylor's expansion of the exact formula for mutual inductance between filaments. Despite the assumption of uniform current density in each of the bars, the model is sufficiently accurate for the interconnections of current and future LSIs, in which most of the wires are not affected by the skin and proximity effects. Expression of the equations in polynomial form provides a balance between accuracy and computational complexity. These equations are mapped according to the geometric structures for which they are most suitable in minimizing runtime in the calculation of inductance while remaining accurate to within 3%. Within the geometrical constraints, the wires are of arbitrary specification.From a comprehensive evaluation on the ITRS-specified global wiring structure for 2003, the values for inductance extracted through the proposed approach are within 3% of the values obtained by commercial three-dimensional (3-D) field solvers. The efficiency of the proposed approach is also demonstrated by extraction from a real layout design that has 300-k interconnecting segments. Atsushi Kurokawa, Takashi Sato 0001, Hiroo Masuda |
ASP-DAC | 2 |
| 2003 | Accurate prediction of the impact of on-chip inductance on interconnect delay using electrical and physical parameter-based RSFabstractThis paper proposes a new methodology to accurately predict the impact of inductance on on-chip wire delay using response surface functions (RSF). The proposed methodology consists of two stages which involves first calculating the delay difference between RC and RLC wire models for a set of parameter variations, then building RSFs using electrical parameters such as wire resistance, capacitance, etc., and physical parameters such as wire width, pitch, etc. as variables. The proposed methodology can help 1) to define design rules for avoiding inductance effects, 2) to point out wires that require RLC delay calculation, and 3) to estimate and correct the delay when using an RC model. An example design rule for limiting self inductance and accurate estimation of the delay difference for a 100 nm technology node is also presented. Takashi Sato 0001, Toshiki Kanamoto, Atsushi Kurokawa, Yoshiyuki Kawakami, Hiroki Oka, Tomoyasu Kitaura, Hiroyuki Kobayashi, Masanori Hashimoto |
ASP-DAC | 1 |
| 2003 | Bidirectional closed-form transformation between on-chip coupling noise waveforms and interconnect delay-change curvesabstractA novel concept of bidirectional transformation between on-chip coupling noise waveform and delay-change curve (DCC) using closed-form equations is described in this paper. These equations are targeted for use in: 1) the efficient generation of DCCs and 2) accurate experimental determination of subnanosecond coupling noise. In particular, we explore the concept of using analytical models to efficiently generate DCCs that can then be used to characterize the impact of noise on any victim/aggressor configuration. The concept is model independent, although we investigate several common noise modeling choices and perform a sensitivity analysis to optimize the generation of DCCs. By extending existing noise models, arbitrary configurations can be considered including multiple aggressors in the timing-analysis framework. Simulation using the analytical approach closely matches time-consuming SPICE simulations, making noise-aware timing analysis using DCCs both efficient and accurate. A test chip using a 0.25-/spl mu/m CMOS process was designed and its measurement results also show good agreement with SPICE simulations. Takashi Sato 0001, Yu Cao 0001, Kanak Agarwal 0001, Dennis Sylvester, Chenming Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1989 | An efficient algorithm for layout compaction problem with symmetry constraintsabstractAn efficient algorithm is presented for the symbolic layout compaction problem with symmetry constraints. The symmetry constraint maintains the geometric symmetry of the circuit components during the layout compaction. It is indispensable to the symbolic layout for analog LSIs where the geometric symmetry between the components is important. However, it makes the compaction problem so complicated that no efficient algorithm has ever been shown except for the time-consuming linear programming algorithm. The proposed algorithm uses both the graph-based technique and the linear programming technique, and takes advantage of the high speed of the former and the generality of the latter. The authors implemented the proposed algorithm in a layout compaction program. The experimental results show that the proposed algorithm is fast enough for practical use.> R. Okuda, Takashi Sato 0001, Hidetoshi Onodera, K. Tamariu |
ICCAD | 2 |