Masahiro Fujita 0004

dblp:56/1768-4 · DBLP profile ↗
← Back
65ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-6516-4175ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 61 · 2 first-author · 17 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 A SAT-Hard Compound Logic Locking Scheme with Empirical Resistance to Known Structural Attacks
abstract
Logic Locking aims to hide the original functionality of the design using a secret key. It protects hardware intellectual properties (IPs) against IP piracy or IC overproduction. However, an attacker analyzes the structural traces and/or uses Boolean satisfiability based technique called SAT attack to break such logic locking schemes. This motivates us to find a logic locking technique that can work against both SAT and structural analysis attacks. Therefore, this paper introduces a novel multiplier-based logic locking scheme. Leveraging the inherent complexity of multiplier circuits, the proposed scheme exponentially increases the time required for each iteration of a SAT attack. Moreover, a heuristic is also proposed to identify appropriate locations for inserting multiplier instances to increase the number of iterations. The multiplier-based logic locking scheme is further combined with the Anti-SAT scheme to create a robust and effective defense mechanism against SAT attacks and the various other attacks exploiting structural traces. The proposed technique is resilient to the state-of-the-art attack dedicated to the existing compound logic locking schemes. Moreover, the proposed compound logic locking scheme, requires half the number of key inputs than the state-of-the-art logic locking scheme while providing the similar level of security.
Sonali Shukla, Govind Rajhans Jadhav, Durgesh Sardan, Suryakant Toraskar, Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
DDECS7
2026 A comparative study on formal verification techniques to verify large integer multiplier circuits
Asutosh Srivastava, Masahiro Fujita 0004
Integr.3
2026 LibSCAT: Library-Based Formal Verification of Heavily Optimized Multipliers via GNN-Guided Reference Selection
abstract
Formal verification of heavily optimized multipliers is a critical yet challenging problem in both industry and academia. Current approaches suffer from fundamental limitations: Symbolic Computer Algebra (SCA) techniques struggle with heavily optimized multipliers, Satisfiability (SAT)-based approaches require structurally similar reference designs, and hybrid methods fail to handle Booth multipliers. On the other hand, industrial design flows possess extensive libraries of verified multipliers for optimization workflows, creating an underutilized opportunity for library-based verification. Yet optimal reference selection becomes challenging due to large-scale libraries and optimization-obscured architectural relationships. To address these challenges, we propose LibSCAT, a verification framework that leverages large-scale reference libraries in a scalable manner. First, we propose a reference library-based methodology that adaptively combines SCA and SAT techniques through intelligent reference selection and predictive method choice. Second, we propose a Siamese Graph Neural Network model that captures multiplier structural relationships in latent space from reverse-engineered graphs, generating robust embeddings for efficient reference selection. Third, we propose a Random Forest-based predictor that leverages learned embeddings for accurate selection of verification strategies. Experimental results show our method achieves 88.2% success on heavily optimized simple partial product multipliers and 94.0% success on heavily optimized Booth multipliers, significantly outperforming state-of-the-art methods.
Rui Li 0095, Masahiro Fujita 0004, Heng Yu 0001, Guangyao Yan, Lin Li 0079, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Learned Image Codec on FPGA: Algorithm, Architecture and System Design
abstract
This paper describes our design for learned image codec (LIC) on FPGA, from the aspects of algorithm, architecture and system. For the algorithm, we build the neural network on the hyperprior structure. Besides, we present a quantization aware training scheme specifically adapted to LIC. For the architecture, we propose a fine-grained pipeline architecture. Channel parallelism constraint and neural network search are proposed to improve the DSP utilization and efficiency, respectively. For the system, we make a CPU-FPGA heterogeneous coding system in which a system-level pipeline is proposed to maximize the throughput. A 720P@30FPS demo and a cross-platform demo are provided in the websites.12
Heming Sun, Jing Wang 0181, Silu Liu, Shinji Kimura, Masahiro Fujita 0004
ASP-DAC5
2025 Weight-Aware Scan Chain Stitching for Shift Power Minimization under Routing Constraints
abstract
Scan chain architecture is a widely adopted design-for-testability (DFT) technique in modern VLSI circuits. However, with the increasing complexity of integrated circuits, excessive shift power during scan testing has become a critical concern, especially for power-constrained designs, as it can lead to increased thermal stress and reduced reliability. In this paper, we propose a weight-aware scan chain stitching methodology that effectively reduces both shift-in and shift-out power while maintaining optimized routing length. Experimental evaluations on ISCAS’89 and ITC’99 benchmark circuits demonstrate average reductions of 14.5% in scan shift power and a 68.5% improvement in routing efficiency compared to state-of-the-art techniques. Index Terms-scan-based testing, scan chain stitching, design for testability (DFT), routing constraint, low power testing.
Rohit Badjatya, Masahiro Fujita 0004, Virendra Singh
ATS4
2025 Breaking the Barriers of One-to-One Usage of Implicit Neural Representation in Image Compression: A Linear Combination Approach With Performance Guarantees
abstract
In an era, where the exponential growth of image data driven by the Internet of Things (IoT) is outpacing traditional storage solutions, this work explores and advances the potential of implicit neural representation (INR) as a transformative approach to image compression. INR leverages the function approximation capabilities of neural networks to represent various types of data. While previous research has employed INR to achieve compression by training small networks to reconstruct large images, no work has explored past the fundamental barrier of using one network per image. This work proposes a novel advancement by breaking this barrier and representing multiple images with a single network. By modifying the loss function during training, the proposed approach allows a small number of weights to represent a large number of images, even those significantly different from each other. A thorough analytical study of the convergence of this new training method is also carried out, establishing upper bounds that not only confirm the method’s validity but also offer insights into optimal hyperparameter design. The proposed method is evaluated on the Kodak, ImageNet, and CIFAR-10 datasets. Experimental results demonstrate that all 24 images in the Kodak dataset can be represented by linear combinations of two sets of weights, achieving a peak signal-to-noise ratio (PSNR) of 26.5 dB with as low as 0.2 bits per pixel (BPP). The proposed method matches the rate-distortion performance of state-of-the-art image codecs, such as BPG, on the CIFAR-10 dataset. Additionally, the proposed method maintains the fundamental properties of INR, such as arbitrary resolution reconstruction of images.
Sai Sanjeet, Seyyedali Hosseinalipour, Jinjun Xiong, Masahiro Fujita 0004, Bibhudatta Sahoo 0002
IEEE Internet Things J.4
2025 RefSCAT: Formal Verification of Logic-Optimized Multipliers via Automated Reference Multiplier Generation and SCA-SAT Synergy
abstract
Formally verifying logic-optimized integer multipliers remains a crucial yet insufficiently addressed problem in both industry and academia, presenting significant verification challenges, particularly when verifying the large-scale logic-optimized multipliers with diverse architectures. Satisfiability (SAT)-based methods require structurally similar and known correct reference multipliers, which may not always be readily accessible. Symbolic computer algebra (SCA) techniques can verify multipliers without references but encounter difficulties with optimized multipliers due to unclear adder boundaries. To enable effective formal verification of the optimized multipliers, we propose the RefSCAT framework, which contains a reference multiplier generator that produces references structurally similar to the optimized multiplier with clear adder boundaries, enabling a synergistic SCA-SAT verification flow. First, we propose a reverse engineering algorithm that extracts the essential adder tree from the optimized multiplier, ensuring similarity. Second, since only a partial netlist is extractable after optimization, we propose a constraint satisfaction algorithm to complete the generation using only adders while following the extracted netlist, ensuring both similarity and clear adder boundaries. Third, leveraging the generated reference, we propose a synergized SCA-SAT verification flow that verifies the generated reference using SCA and then uses it as a correct reference for the SAT-based verification. The experiments demonstrate that RefSCAT can successfully verify logic-optimized multipliers with diverse partial-product-based architectures up to 128 bits, outperforming the state-of-the-art methods by verifying at least 29% more benchmarks.
Rui Li 0095, Lin Li 0079, Heng Yu 0001, Masahiro Fujita 0004, Weixiong Jiang, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 RefSCAT-2.0: Formal Verification of Large-Scale Optimized Multipliers via Quantum-Inspired Ant Colony Optimization-Based Reference Generation
abstract
Formal verification of large-scale optimized integer multipliers remains a critical yet insufficiently addressed challenge in industry and academia. Current methods employ reference multiplier generators to automatically construct structurally similar reference multipliers, which are then used by Satisfiability (SAT)-based techniques to verify equivalence with optimized multipliers. However, these approaches face limitations when generating references for large-scale optimized multipliers within acceptable timeframes. To address these limitations, we introduce the RefSCAT-2.0 framework, designed to rapidly produce high-quality large-scale reference multipliers. Firstly, we generate the macro-architecture to determine the number of adders required for constructing the reference multiplier. We propose a novel Integer Linear Programming (ILP)-based macro-architecture generation algorithm that minimizes the number of allocated adders, thereby reducing the overall problem complexity. Secondly, we organize the allocated adders into groups to simplify the subsequent generation process. We present a multi-level scheduler that automatically decomposes adders into groups with minimized interdependencies, ensuring both the quality of generation and a reduction in overall generation complexity. Thirdly, we generate the micro-architecture for each scheduled group, wherein we finalize the connections between adders. We present a graph-based design space representation coupled with a quantum-inspired ant colony optimization (QACO)-based generation algorithm that can efficiently explores the micro-architectures of each scheduled group. Experimental results show that RefSCAT-2.0 successfully verifies all 124 cases in a 256-bit optimized multiplier benchmark suite, outperforming SCA-based tcad22revsca and hybrid RefSCATTCAD24 methods which solve only 24 cases each.
Rui Li 0095, Lin Li 0079, Heng Yu 0001, Masahiro Fujita 0004, Weixiong Jiang, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 LLM-aided Front-End Design Framework For Early Development of Verified RTLs
abstract
This work demonstrates the potential of a proposed Large language model (LLM) aided front-end design flow in the early development of verified Register Transfer Level (RTL) design. The proposed framework consists of three task-specific LLMs that generate RTL description, test-bench, and review the design to suggest required modifications depending on the simulation results feedback into the model. The proposed framework has been implemented twice with two different versions of OpenAI LLM viz. GPT-3.5-turbo and GPT-4o-mini. Each implementation has been used independently for developing ten distinct designs of different complexities. The results show that low-complexity designs get generated within a few minutes with fewer feedback or review iterations, moderate complexity designs require more time and iterations as compared to low-complexity designs. However, developing a highly complex design is a bit tricky task. Experimental results show that the proposed framework achieves a higher success rate compared to state-of-the-art methods, automatically fixing various types of bugs when simulation results are fed back into the model. Notably, the framework achieves a 90% success rate with GPT-4o-mini and 84% with GPT-3.5-turbo, demonstrating its robustness across different model configurations.
Vyom Kumar Gupta, Abhishek Yadav 0003, Masahiro Fujita 0004, Binod Kumar 0001
ATS3
2024 Bidirectional LSTM Model for Accurate and Real-Time Landslide Detection: A Case Study in Mawiongrim, Meghalaya, India
abstract
This article presents a bidirectional long short-term memory (LSTM) model for the detection of landslides. Previous uses of machine learning (ML) in this setting have demonstrated its general potential, which necessitates the implementation of a suitable algorithm. Landslides are natural disasters that can cause significant destruction and disruption in the affected areas. Early detection is the key to minimizing the impact of landslides, so it is important to develop accurate and efficient models. An area selected for this study is located in Mawiongrim, Meghalaya, India, which is an active landslide zone. The proposed model uses a bidirectional LSTM to capture the temporal patterns of the input data collected from a long-term real-time monitoring system set up in the area. To evaluate the effectiveness of the predictions, the model is trained using a data set composed of various landslide-related characteristics, such as topography, rainfall, hydrological, and soil properties. The results show that the suggested model is capable of detecting landslides with greater accuracy and the lowest error value relative to other models. Additionally, the model is also able to provide a real-time warning system, making it a viable tool for early landslide detection. The research also highlights the prediction models for matric suction and groundwater level, which are crucial in determining slope stability.
J. Sharailin Gidon, Jintu Borah, Smrutirekha Sahoo, Shubhankar Majumdar, Masahiro Fujita 0004
IEEE Internet Things J.5
2023 IIR Filter-Based Spiking Neural Network
abstract
Spiking Neural Networks (SNNs) are closely related to the dynamics of the human brain and use spatiotemporal encoding of information to generate spikes. Implementing various neuronal models in hardware is a popular field of research aiming to mimic biological behavior. The leaky integrate-and-fire model of the neuron is generally chosen for hardware implementation owing to its simplicity and accuracy in modeling the neuron. This paper proposes an infinite impulse response (IIR) filter-based neuron model and describes a backpropagation-based training algorithm for an SNN built using the proposed neurons. The trained network is implemented on an Ultra96-V2 FPGA to validate the design and demonstrate the power and resource efficiency. The implemented design achieves an accuracy of 98.91% on the MNIST dataset and classifies images at 13,021 frames-per-second (FPS) with a 200 MHz clock while consuming$\approx 7.5\times$higher resource efficiency than previous publications.
Sai Sanjeet, Rahul K. Meena, Bibhudatta Sahoo 0002, Keshab K. Parhi, Masahiro Fujita 0004
ISCAS5
2023 Formal Verification of Integer Multiplier Circuits Using Binary Decision Diagrams
abstract
Multiplier circuit covers a more extensive area of embedded system application in digital signal processing, cryptography, and multimedia. Nonstandard implementations and custom optimization are being done to reduce the size of multipliers. The circuit became prone to a buggy, and hence the demand for verification increased. Formal verification methods, such as satisfiability (SAT), symbolic computer algebra (SCA), and binary decision diagrams (BDDs) have made massive progress over the last few decades. However, these methods are insufficient to verify the optimized multipliers. SAT-based equivalence checking is computationally expensive. SCA-based backward rewriting is limited to algebraic-friendly multipliers. The complexity of BDDs is exponential with the input size. Although, by allowing an additional variable method, the size of the BDD is limited to 4th degree polynomial of the number of the inputs, this method is not explored to verify optimized multipliers. This article focus on verifying integer multipliers with diverse architectures. We propose an algorithm for the direct construction of BDDs without traversing circuits and generate BDDs up to 1024 bits. We utilize the additional variable method and constructing BDDs using high-to-low variable ordering. We reduce the complexity of BDD size to a 3rd degree polynomial. We generate BDDs and verify the multipliers with various architectures up to 64 bits. We propose a method to verify optimized multipliers by checking equivalence and verifying up to 32-bits optimized multipliers. We do the error tolerance analysis of our approach by inserting bugs in a circuit at various locations.
Yukio Miyasaka, Asutosh Srivastava, Masahiro Fujita 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Deep Learning-assisted Scan Chain Diagnosis with Different Fault Models during Manufacturing Test
abstract
Manufacturing of integrated circuits at the smaller technology nodes leads to several defects in them that must be screened and appropriately diagnosed for minimization of cost overruns. A substantial portion of the functional failures during the process of manufacturing test is often attributed to the defects inside the scan chains. With the advancements in the digital test technologies, almost every chip is manufactured with in-built pattern compression infrastructure. This exacerbates the problem of scan chain diagnosis from the collected failure traces. In this work, an automated methodology to perform this diagnosis in the presence of multiple faults is proposed. Deep learning is utilized to predict the probable candidate locations given the compressed scan chain response. Experiments have been performed on different fault models. Experimental results indicate that the proposed methodology is able to perform the diagnosis with a success rate of approximately 80-100%.
Utsav Jana, Sourav Banerjee, Binod Kumar 0001, Madhu B, Shankar Umapathi, Masahiro Fujita 0004
ATS6
2022 Aries: A Semiformal Technique for Fine-Grained Bug Localization in Hardware Designs
abstract
Effective bug localization during verification is a challenging step in the development cycle of complex hardware designs. While meeting different coverage goals is possible in the verification process, yet bug localization cannot be directly related to such goals. We propose a two-step methodology to achieve fine-grained design bug localization. First, we obtain multiple error traces based on a failing property. Starting from an initial error trace, we employ model checking to generate supportive error traces that are utilized to mine important assertions. In the second step, we utilize these assertions for fine-grained design bug localization. The mapping of the assertions leads to specific regions in register transfer level descriptions that are highly probable to be the root cause of the design bug. Specifically, we devise a binning methodology to categorize multiple suspects in different bins that need to be investigated by the design engineer for arriving at the correction for corresponding bugs. Experiments on multiple designs illustrate the efficacy of the proposed methodology in comparison to previous work and state-of-the-art industrial tool.
Binod Kumar 0001, Vineesh V. S., Puneet Nemade, Masahiro Fujita 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 High-Precision Sub-Nyquist Sampling System Based on Modulated Wideband Converter for Communication Device Testing
abstract
This paper proposes a new method of constructing a compensation filter for the modulated wideband converter (MWC) system. The proposed method can be directly used in the MWC circuit without disconnecting any components. Furthermore, the non-ideal transfer characteristics of the real mixer output and the real ADC input are taken into account in the proposed compensation filter. The proposed method can be utilized in the advanced MWC that enables more flexibility in the design parameters of the MWC. This paper also demonstrates the reconstruction of the Bluetooth signal as a practical example for the evaluation of the reconstruction performance. The MWC successfully reconstructs the time-domain waveform of the signal based on the proposed compensation filter. Even in the case that the total effective number of channels in MWC is small as close to the necessary condition of the MWC, the error vector magnitude (EVM) was still under 1%, which is acceptable for Bluetooth device testing applications.
Zolboo Byambadorj, Koji Asami, Takahiro J. Yamaguchi, Akio Higo, Masahiro Fujita 0004, Tetsuya Iizuka
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 BMC-Based Temperature-Aware SBST for Worst-Case Delay Fault Testing Under High Temperature
abstract
This article presents a bounded model checking (BMC)-based temperature-aware software-based self-testing (SBST) technique to test worst case delay faults within the highest temperature range. The BMC-based SBST method first defines the sequential constraint. It develops a sequentially constrained automatic test pattern generation (ATPG) to ensure that the generated delay test patterns can emerge in functional mode. It then uses the processor’s multiple-level information to reduce the model complexity, avoid aborts due to time-outs during the BMC process, and generate test programs automatically. A temperature-aware SBST method has then been developed to ensure that the test temperature is within the specified range and test the worst case delays under high temperature. Experimental results demonstrate that the proposed technique achieves an extremely high coverage for delay faults and effectively avoids yield loss caused by the overtesting problem. Its test quality also outperforms that of the existing methods. The generated SBST programs are successful and efficient in testing worst case delay faults under high temperature.
Ying Zhang 0040, Zebo Peng, Huawei Li 0001, Masahiro Fujita 0004, Jianhui Jiang
IEEE Trans. Very Large Scale Integr. Syst.5
2021 Logic Synthesis Meets Machine Learning: Trading Exactness for Generalization
abstract
Logic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning.
Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee
DATE7
2021 Enhanced Design Debugging With Assistance From Guidance-Based Model Checking
abstract
Design debugging is one of the most important steps in the modern integrated circuits (ICs) development cycle. Simulation-based verification is never sufficient for ensuring design correctness because of its incomplete nature. Formal techniques such as model checking promise to solve this issue through a complete state-space traversal approach. However, because of increasing design complexity, such methods suffer from scalability issues. Guidance-based state-space traversal techniques have been proposed in the past to assist the model checkers in overcoming the complexity bottleneck. Automatically identifying these guidance hints is relatively difficult and requires heuristic-based reasoning procedures. Additionally, to come up with quick fixes during the debug stage, an effective bug localization strategy is needed. In this article, we revisit the paradigm of guidance-based model checking and propose a methodology to improve these guidance generation mechanisms for achieving fine-grained bug localization. In particular, this work proposes a systematic methodology to localize the buggy RTL lines from the erroneous RTL simulation trace. The proposed technique involves the mining of invariant-like assertions from simulation traces. The mined assertions act as probable guidance candidates for the model checking exercise. To identify useful guidance hints from possible ones, we use the Bayesian networks that explore conditional dependence between the various hints at different levels and the target property. These guidance hints are utilized for obtaining possible buggy subregions, which are analyzed via an iterative model checking methodology for fine-grained bug localization. By using the proposed framework, bugs can be localized to within a few lines of RTL description.
Vineesh V. S., Binod Kumar 0001, Rushikesh Shinde, Masahiro Fujita 0004, Virendra Singh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 SAT-Based On-Track Bus Routing
abstract
In modern integrated circuit design, bus routing is a challenge because of complex design rules and wiring constraints. Despite extensive research, state-of-the-art in bus routing is not effective when nonuniform tracks, various obstacles, wire width constraints, and multiple spacing rules should be handled simultaneously. A new bus routing framework proposed in this article is based on maze routing and Boolean satisfiability. It produces high-quality results quickly and allows for additional optimizations, such as minimizing wire length on the critical paths. A number of challenging bus routing benchmarks appeared in 2018 ICCAD Contest. Experiments on these benchmarks not only show that the framework is faster than the winners of the competition and previous work but also produces better results, improving the overall cost by 12% while at the same time minimizing the number of spacing violations.
He-Teng Zhang, Masahiro Fujita 0004, Chung-Kuan Cheng, Jie-Hong Roland Jiang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 LUT-based Circuit Approximation with Targeted Error Guarantees
abstract
Approximate circuits are widely gaining popularity in various fields where error tolerance is applicable. However, striking the right balance between error tolerance and the output quality is a challenging step in the overall design of approximate systems. We propose a systematic approach utilizing Look-Up Table (LUT)-based netlist transformations to achieve approximation while targeting specific error guarantees. Specifically, we employ a SAT-based property checking technique to accommodate worst-case error constraints acting as error guarantees. The proposed methodology involves the formulation of templates to enable the reusability of the technique for different design choices. The analysis comprises of fitness function evaluation based on layout area or the considered error guarantees. We analyze the impact of different parameters on the quality of the output of the resulting approximation and the time taken to obtain them.
Vinod G. U, Vineesh V. S., Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
ATS4
2020 Synthesis and Optimization of Multiple Portions of Circuits for ECO based on Set-Covering and QBF Formulations
abstract
Engineering Change Order (ECO) and logic debugging problems where multiple locations in the circuit must be modified are formulated with Quantified Boolean Function (QBF) and set-sovering techniques. The formulation is based on the fanin selection method for each gate. Although the resulting formulation for single portion changes is basically equivalent to Sets of Pairs of Functions to be Distinguished (SPFD) [3], the way of its computations is quite different. Moreover, the simultaneous changes for multipl portions becomes Boolean Relation extension of SPFD. Experimental results and applications to various logic optimization problems are also shown.
Masahiro Fujita 0004, Yusuke Kimura, Xingming Le, Yukio Miyasaka, Amir Masoud Gharehbaghi
DATE1
2020 Post-Silicon Gate-Level Error Localization With Effective and Combined Trace Signal Selection
abstract
Incorporating on-chip trace buffers (TBs) helps to overcome the limited observability by tracing selected signals during post-silicon validation. The effectiveness of TB-based techniques largely relies on selection of appropriate trace signals. For processor-based systems, the selection becomes relatively easier because important signals can be identified. However, for a general digital block in a complex system-on-chip, recognizing necessary trace signals becomes extremely challenging and requires a systematic approach. Previous research on trace signal selection has mainly focused on improving reconstruction of unknown signal values with the help of traced signals. Even though it serves as a good selection principle, an effective signal selection must consider other important factors such as error detection (ED) with the traced signals, which in turn assist in localization and root-cause discovery. Additionally, from practical point of view, the signal selection algorithm needs to cater to factors like routing congestion and minimizing routing wire length. The proposed methodology of signal selection attempts to combine these three crucial factors of signal selection: restoration of untraced signal states, ED with traced signals and routing considerations. The concurrent maximization of all these three parameters is difficult as they have conflicting preference of the candidate trace signals. Hence, the proposed signal selection approach presents a methodology of judiciously mixing the choices of these three objectives. Furthermore, the restored and traced signal states are analyzed for the purpose of error localization at the gate level for several design error models.
Binod Kumar 0001, Kanad Basu, Masahiro Fujita 0004, Virendra Singh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 An Automatic Test Pattern Generation Method for Multiple Stuck-At Faults by Incrementally Extending the Test Patterns
abstract
As the number of transistors in the fabricated circuits becomes extremely larger, not only single stuck-at faults, but also multiple stuck-at faults (MSAFs) are likely to happen in the circuits, especially for the large-scale circuit. Multiple faults are difficult to be fully covered due to the exponentially enormous number of all possible faults. Although there are methods proposed to deal with the multiple faults, they fail to generate compact test patterns to detect all faults within an acceptable running time. In order to solve these problems, this article proposes an incremental automatic test pattern generation method to deal with MSAFs. Instead of traversing the entire n multiple fault list, the proposed method only selects the faults undetected by the existing test patterns for n - 1 faults, and then generates additional test patterns. Staring from a complete test set for single faults, the proposed ATPG method can be incrementally applied to handle all multiple faults. Moreover, since the number of undetected faults that are selected is extremely smaller comparing to the total number of the entire fault list, the proposed method can generate compact test patterns to cover all faults within an acceptable running time.
Peikun Wang, Amir Masoud Gharehbaghi, Masahiro Fujita 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 A Methodology to Capture Fine-Grained Internal Visibility During Multisession Silicon Debug
abstract
Silicon debugging is carried out in multiple sessions which are characterized by run-and-halt intervals. One of the important criteria for the success of this method is that the debugging infrastructure should capture only the erroneous data which can add important insights to the debugging process. However, identification of such suspect clock cycles is not a trivial exercise and requires an systematic approach. We propose a debugging architecture for enhancing the multisession procedure using the technique of on-chip debug data compression. The first session assists in identifying those erroneous clock cycles, and the useful debug data are collected in the second session with the help of markers called tag bits. At the cost of a minimal increase in area overhead, the proposed architecture achieves finer temporal visibility expansion because of the debug data collection in a segregated manner. During the offline analysis of the collected debug data, error localization can be achieved to a finer resolution. We evaluate our methodology on several designs for different kinds of error configurations. Experimental results show that the proposed methodology can achieve better on-chip storage utilization and the expansion in the temporal observation window compared to similar techniques in the literature.
Binod Kumar 0001, Jay Adhaduk, Kanad Basu, Masahiro Fujita 0004, Virendra Singh
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Validating Multi-Processor Cache Coherence Mechanisms under Diminished Observability
abstract
Modern chip multi-processors (CMP) inevitably require cache coherence mechanisms for their correct operation. However, exhaustive functional verification of a complex cache coherence mechanism is a challenging task. This leads to bugs escaping to the first silicon and necessitates validation at the post- silicon stage. In this work, an on-chip signal logging method is proposed which helps in bug detection in case of design errors and soft-errors arising out of reliability issues. The logged contents can then be further dumped off-line for fine-grained bug localization. The proposed methodology utilizes cache coherence protocol specifications to obtain the signal states of coherence transactions and the detector module flags an error once a mismatch is found between observed signal states and correct signal states. The proposed logging mechanism decreases the error detection latency at minimal area and power overheads. Experiments on a four core multiprocessor having a 7-stage MIPS pipeline implementing the widely utilized directory-based MESI protocol indicate that the proposed methodology succeeds in detecting design errors. Analysis of soft errors have also been performed and shorter error detection latency is achieved compared to a previously proposed technique in the literature.
Binod Kumar 0001, Atul Kumar Bhosale, Masahiro Fujita 0004, Virendra Singh
ATS3
2019 Securing Scan through Plain-text Restriction
abstract
Scan design-for-test (DfT) feature can be exploited as a side channel to break a cryptographic chip. The stringent test and diagnosis requirements of present-day complex system-on-chip (SoC) make use of the scan DfT feature unavoidable. However, being a threat to cryptographic chips, it needs to be secured against the scan-based side-channel attacks. In this paper, we propose a simple yet effective technique to prevent scan attack on Advanced Encryption Standard (AES) cryptographic chip. The proposed technique restricts the user from applying any random inputs at the plain-text inputs. To use the scan feature the plain-text inputs must be forced to a constant all-0 or all-1 value throughout the test session. Because of this feature, there is no possibility of mounting any differential scan attack. The proposed technique is simple to implement and does not have any impact on test coverage.
Satyadev Ahlawat, Kailash Ahirwar, Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
IOLTS4
2019 Signal Selection Methods for Efficient Multi-Target Correction
abstract
This paper proposes methods to get a minimally rectified logic circuit equivalent to a new specification. One of the proposed methods can deal with multiple target signals to be modified at the same time. Moreover, a graph-cut based method is proposed to reduce the search space. In the case the problem is difficult to be handled by the at-the-same-time method, an enhanced method is proposed to efficiently deal with the target signals one by one. The quality of results on ICCAD'17 contest benchmarks is equally good or better, average 16 times better score, than the previous work.
Yusuke Kimura, Amir Masoud Gharehbaghi, Masahiro Fujita 0004
ISCAS3
2019 Live Demonstration: Automatic Synthesis of Algorithms on Multi Chip/FPGA with Communication Constraints
abstract
Mapping of large systems/computations on multiple chips/multiple cores needs sophisticated compilation methods. In this demonstration, we present our compiler tools for multi-chip and multi-core systems that considers communication architecture and the related constraints for optimal mapping. Specifically, we demonstrate compilation methods for multi-chip connected with ring topology, and multi-core connected with mesh topology, assuming fine-grained reconfigurable cores, as well as generalization techniques for large problems size as convolutional neural networks. We will demonstrate our mappings methods starting from data-flow graphs (DFGs) and equations, specifically with applications to convolutional neural networks (CNNs) for convolution layers as well as fully connected layers.
Tomohiro Maruoka, Yukio Miyasaka, Akihiro Goda, Amir Masoud Gharehbaghi, Masahiro Fujita 0004
ISCAS5
2019 High-Level Engineering Change Through Programmable Datapath and SMT Solvers
abstract
In this paper, we present a technique to automatically adjust Register-transfer level (RTL) implementation for Engineering Change Order (ECO) in high level. Our method focuses on the datapath structure, where circuit topologies are mostly, and only partial portions are replaced by programmable datapath. Exploring the correct configuration of the datapath is formulated as Quantified Boolean Formula (QBF) problem, and can be solved by using Satisfiability (SAT) /Satisfiability Modulo Theories (SMT) solver in an incremental way automatically. The experimental results with several example cases have confirmed that effectiveness of the proposed method.
Qinhao Wang, Amir Masoud Gharehbaghi, Takeshi Matsumoto, Masahiro Fujita 0004
ISCAS4
2019 Approximate Arithmetic Circuit Design Using a Fast and Scalable Method
abstract
Approximate computing can be applied to error-tolerant applications, by trading off accuracy for lower power consumption, shorter delay and smaller area. In this paper, we focus on the approximate arithmetic circuit design especially targeting combinational multipliers and adders, which are essential computing components in machine learning such as neural network computation. We propose a novel approach to generate approximate circuits from the given correct circuits. The basic idea is to study the circuit's characteristics from the small instances of the circuits and then establish a common algorithm for larger circuits with the same architectures. We propose two different methods and apply them to adders and multipliers with different architectures. The experimental results demonstrate that our method outperforms the state-of-art methods in terms of the quality of the circuits with orders of magnitude shorter processing time.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
VLSI-SoC3
2019 An Incremental Automatic Test Pattern Generation Method for Multiple Stuck-at Faults
abstract
This paper proposes an incremental ATPG method to deal with multiple stuck-at faults. In order to generate the test set for n multiple faults, only the additional test patterns for the undetected faults by the existing test patterns for n - 1 multiple faults are generated. Moreover, by introducing an efficient fault selection method, the size of the fault list to be dealt with is reduced drastically compared to the entire fault list of n multiple faults. Our experimental results on ISCAS 89 and IWLS 2005 benchmarks up to triple faults indicate that the proposed method can generate a compact test set to cover all the faults within an acceptable runtime.
Peikun Wang, Amir Masoud Gharehbaghi, Masahiro Fujita 0004
VTS3
2019 SAT-based Silicon Debug of Electrical Errors under Restricted Observability Enhancement
Binod Kumar 0001, Masahiro Fujita 0004, Virendra Singh
J. Electron. Test.2
2018 On Securing Scan Design Through Test Vector Encryption
abstract
Scan-based side-channel attacks have been gaining a prominence among the malicious attackers. The unprotected scan chains are extremely vulnerable and could be exploited to extract the secret information from a security chip such as an Advanced Encryption Standard (AES) cryptochip. To protect the secret information from being hacked it is utmost necessary to redesign the scan chain with security features. In this paper, we propose a secure scan architecture aiming at the protection of AES cryptochips against scan-based attacks. The proposed idea is based on the principle of test pattern encryption. The major contribution of our architecture is area efficiency and its security features without hampering test, diagnose, and debug capability of the original scan chain. The experimental results and security analysis shows the efficacy of proposed design.
Darshit Vaghani, Satyadev Ahlawat, Jaynarayan T. Tudu, Masahiro Fujita 0004, Virendra Singh
ISCAS4
2018 Multiple Stuck-at Fault Testability Analysis of ROBDD Based Combinational Circuit Design
Toral Shah, Anzhela Yu. Matrosova, Masahiro Fujita 0004, Virendra Singh
J. Electron. Test.3
2017 Post Silicon Debugging of Electrical Bugs Using Trace Buffers
abstract
Post-silicon debug is the task of finding the bugs that could not be found before manufacturing. Electrical bugs are an important category of post-silicon bugs that are typically hard to debug due to complex interdependence of layout and netlist as well as their dependency to the running workload and environment. In this paper, we tackle the problem of debugging electrical bugs using trace buffers. Given an erroneous values captured in the trace buffer due to occurrence of an electrical bug, we try to identify the spatial and temporal location of the error. We have formulated the problem of debugging electrical bugs as a SAT problem and proposed a debugging method based on that. Moreover, we have shown that the existing signal selection methods for logical bugs that are typically trying to maximize signal restoration ratio (SRR) metric, do not perform better than random selection when debugging electrical bugs. On the other hands, utilizing a more sophisticated signal selection that is considering propagation of bit-flips due to electrical bugs is more effective.
Kentaro Iwata, Amir Masoud Gharehbaghi, Mehdi Baradaran Tahoori, Masahiro Fujita 0004
ATS4
2017 Combining Restorability and Error Detection Ability for Effective Trace Signal Selection
abstract
Persistent growth in design complexity has led to increased chances of bugs appearing during post-silicon validation. Debugging errors at this stage requires some arrangement for expanding the observability of internal signals of the design. Limited number of trace buffers help in increasing visibility of these internal states. However, appropriate trace signal selection is very difficult. Restorability of untraced states with the help of traced ones is a popular approach for signal selection although it fails to address the main issue of error detection. We propose a signal selection methodology which combines the restoration capability and error detection ability of these signals. Experimental evaluation of the proposed signal selection approach on benchmark circuits indicates improved error detection for various kind of design errors. Practical consideration like minimizing routing overhead has also been analyzed.
Binod Kumar 0001, Ankit Jindal, Masahiro Fujita 0004, Virendra Singh
ACM Great Lakes Symposium on VLSI3
2017 Instruction-based self-test for delay faults maximizing operating temperature
abstract
In today's technology, reliability is one of the major challenges. Process variations and increasing power density in advanced technology nodes make the condition even worse. Process variations during manufacturing induce delay variations and such variations may manifest into faults at the higher temperature during operation. Delay faults at higher functional temperature is a critical factor for the reliability of a system. Current DFT-based methods do not handle this properly. In this paper, we propose an instruction based self-test (IBST) technique to elevate the temperature of a chip near to functional temperature and test it for delay faults at that temperature. In the first step, integer linear programming (ILP) is used to find out power hungry instructions. These instructions cause maximum toggling in a unit. In the second step, delay test instructions are combined to create a program for testing functional unit of an out-of-order superscalar processor. Experimentation of the proposed technique is carried out on portable ISA (PISA) based superscalar processor. Higher operating temperature 95°C ± 2°C is achieved and it is maintained while applying delay test.
Nihar Hage, Rohini Gulve, Masahiro Fujita 0004, Virendra Singh
IOLTS3
2017 A new approach for diagnosing bridging faults in logic designs
abstract
As fabricated chips get larger and denser, more kinds of defects happen that cannot be explained by the traditional stuck-at-fault model. In this paper, we propose a new approach for diagnosis of bridging faults based on analysis of the logic design. The proposed approach is based on iteratively solving SAT problems until finding the internal signals that can explain the misbehavior of the circuit caused by bridging faults. Our experiments on ISCAS circuits shows efficiency and effectiveness of our method.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
ISCAS2
2017 Test pattern generation for multiple stuck-at faults not covered by test patterns for single faults
abstract
Previous works have shown that given an initial set of test patterns for single faults, relatively few additional tests are required to cover all multiple faults. In this paper, the exact situations in which the test patterns for single stuck-at faults do not detect multiple stuck-at faults are examined. We present proofs which show the conditions that is required to be met for ATPG on single faults to cover all multiple faults. Based on this analysis, we propose a new incremental ATPG method which first targets only single faults and then incrementally expands it to multiple faults with larger cardinalities, and present the experimental results for double faults.
Conrad J. Moore, Peikun Wang, Amir Masoud Gharehbaghi, Masahiro Fujita 0004
ISCAS4
2017 A low-cost approximate 32-point transform architecture
abstract
This paper presents an area-efficient approximate method for 32-point transform which is one of the most area-consuming parts in High Efficiency Video Coding (HEVC) applications. Compared to prior literatures, this work reduces the hardware cost of transform by 1) eliminating all the arithmetic operations of 6 least significant bits (LSB), 2) presenting a low-delay method for generating carry propagation from the remaining 5 LSBs and 3) truncating the most significant bits (MSB) according to the position of component. In the implementation of a 32-point forward transform, the experimental results show that 27% area consumption can be saved and the coding efficiency loss aroused by the approximation is only 0.044% compared with the origin.
Heming Sun, Zhengxue Cheng, Amir Masoud Gharehbaghi, Shinji Kimura, Masahiro Fujita 0004
ISCAS5
2017 Improving post-silicon error detection with topological selection of trace signals
abstract
Drastic growth in design complexity of VLSI circuits has increased the chances of bugs escaping to first released silicon. This has resulted in an increased emphasis on post-silicon validation and debug which is typically hindered by limited observability of internal signals. Trace buffers assist in curbing this bottleneck by storing selected signal states for limited clock cycles. For efficient use of these on-chip buffers, devising a proper selection criterion is of utmost importance. Maximization of restoration of untraced signals is a widely utilized signal selection metric. However, this approach has been seen to not be very effective for error detection. This paper proposes a trace signal selection technique based on error transmission, taking into account the topology of the design. The proposed signal selection methodology can be effectively applied to trace as well as a combination of trace and scan based observability techniques. Experimental evaluation of the proposed methodology on different design errors indicates improvement in error detection as compared to restorability based selection techniques.
Binod Kumar 0001, Kanad Basu, Ankit Jindal, Masahiro Fujita 0004, Virendra Singh
VLSI-SoC4
2017 A new approach for constructing logic functions after ECO
abstract
Engineering change orders (ECO) are small changes in the design due to last minute bug fixes or spec changes. In this paper, we focus on functional ECO in a logic design and try to construct the new logic function, reusing the existing logic as much as possible. Traditional approaches try to find appropriate locations in the original design that their modifications may result in the new functionality. However, those methods usually fail if additional inputs are required. We propose a new approach based on iterative SAT solving to find the inputs of the function for the given internal nodes, or the primary outputs that are the target of the ECO, out of all the internal signals and primary inputs such that it is guaranteed to be able to correct the functionality without explicitly generating the functions. Our experimental results on ITC99 benchmarks show the efficiency and effectiveness of our approach, specially for hard ECO cases.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
VLSI-SoC2
2017 Subthreshold Operation of CAAC-IGZO FPGA by Overdriving of Programmable Routing Switch and Programmable Power Switch
abstract
A field-programmable gate array (FPGA) using a crystalline oxide semiconductor of c-axis-aligned crystal indium-gallium-zinc oxide (CAAC-IGZO) has been developed, which is capable of subthreshold operation used for energy harvesting. To achieve subthreshold operation, the CAAC-IGZO FPGA has a structure designed as an extension of a boosting pass gate using a CAAC-IGZO FET and employs overdriving of a programmable routing switch and a programmable power switch for power gating (PG). A CAAC-IGZO FET is used to give an ideal floating gate with excellent charge retention. A chip fabricated using a 0.8-μm CAAC-IGZO/0.18-μm CMOS hybrid process achieves subthreshold operation while maintaining the features required for normally off computing proposed in our previous study. Specifically, these features are realized by fine-grained PG for individual programmable logic elements (PLEs), fast configuration switching between contexts, and load/store between a volatile register and a nonvolatile shadow register in the PLEs. The chip operation at a minimum operating voltage of 180 mV with a combinational circuit configuration is demonstrated. With a sequential circuit configuration, the chip operates at a minimum operating voltage of 190 mV with 12.5 kHz, and the minimum power-delay product is 3.40 pJ/operation at 330 mV.
Munehiro Kozuma, Yuki Okamoto, Takashi Nakagawa, Takeshi Aoki, Yoshiyuki Kurokawa, Takayuki Ikeda, Yoshinori Ieda, Naoto Yamade, Hidekazu Miyairi, Masahiro Fujita 0004, Shunpei Yamazaki
IEEE Trans. Very Large Scale Integr. Syst.11
2016 A New Approach for Debugging Logic Circuits without Explicitly Debugging Their Functionality
abstract
Traditional debugging methods try to find an appropriate function of the given inputs for the internal nodes such that the modified circuit becomes correct. In this paper, we propose a new approach based on iterative SAT solving to find the inputs of the function for the internal nodes out of all the internal signals and primary inputs such that it is guaranteed to be able to correct the functionality without explicitly generating the functions. Our experimental results on ITC'99 benchmarks shows the efficiency and effectiveness of our approach.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
ATS2
2015 Temperature-aware software-based self-testing for delay faults
Ying Zhang 0040, Zebo Peng, Jianhui Jiang, Huawei Li 0001, Masahiro Fujita 0004
DATE5
2015 Trace signal selection methods for post silicon debugging
abstract
In post-silicon debugging, only a limited number of states (flip-flops) can be traced, due to the area overhead that is introduced by trace buffers. Therefore, it is important to select the states which can restore most of the other states. There exist researches that try to heuristically select a set of flip-flops (FFs) which maximizes the number of restored FFs. We first show that those existing works are not so robust, as the cost functions used for selections do not work well in some cases. In this paper, we introduce a new signal selection that tries to improve the selection by swapping the FFs that are going to be traced. Furthermore, we introduce a hardware implementation of the method that is more than 3 orders of magnitude faster than software-based swapping. With the proposed methods, we can improve the signal selection and get consistent results even for large circuits.
Shridhar Choudhary, Amir Masoud Gharehbaghi, Takeshi Matsumoto, Masahiro Fujita 0004
VLSI-SoC4
2015 Efficient signature-based sub-circuit matching
abstract
We introduce a new approach for circuit matching using signatures. We have introduced a signature based on the topology of the fanin cone of each circuit element. First, we generate the signature for each circuit element. Then, we find all the circuit elements with unique signature among the given circuits. Finally, we expand the matching area by our expansion rules. Our experiments on IWLS2005 benchmark suite show that our method is able to find the perfect matching between two 160,000-gate IP in 43 minutes. In addition, our method is more than one order of magnitude faster than a state of the art graph matching based approach, while the size of the matched area is more than 30% larger.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
VLSI-SoC2
2013 Rectification of advanced microprocessors without changing routing on FPGAs (abstract only)
abstract
We propose a method for rectification of bugs in microprocessors that are implemented on FPGAs, by only changing the configuration of LUTs, without any modification to the routing. Therefore, correcting the bugs does not require resynthesis, which can be very long for complex microprocessors due to possible timing closure problems. As the structure of the circuit is preserved, correcting the bugs does not affect the timings of the circuit. In design phase, we may add additional LUTs to the original circuit, so that we can use them in the correction phase. After a bug is found, we perform the following two tasks. Fist, we find the candidate control signals as well as the required change to correct their behavior. This is done by using symbolic simulation and equivalency checking between the formal specification and the erroneous formal model of the processor. Then, we try to map the corrected functionality into the existing LUT structure. This is done by a novel method that formulates the problem as a QBF (Quantified Boolean Formula) problem, and solves it by repeatedly applying normal SAT solvers instead of QBF solvers under a CEGAR (Counter Example Guided Abstraction Refinement) paradigm. We show effectiveness of our method by correcting bugs in two complex out-of-order superscalar processors with two different timing error recovery mechanisms.
Satoshi Jo, Amir Masoud Gharehbaghi, Takeshi Matsumoto, Masahiro Fujita 0004
FPGA4
2013 Debugging processors with advanced features by reprogramming LUTs on FPGA
abstract
In this paper, we propose an automated method for debugging and rectification of logical bugs in processors that are implemented on FPGAs. Our method is based on preserving the current circuit topology, and debugging and rectifying bugs by only changing the contents of LUTs, without any modification to the wiring. As a result, correcting the bugs does not require re-synthesis, which can be very time consuming for complex processors due to possible timing closure problems. As the topology of the circuit is preserved, correcting the bugs does not affect the timings of the circuit. In the design phase, we may add additional LUTs or additional inputs to LUTs in the original circuit, so that we can use them in debugging and rectification phase. After a bug is found, first we try to identify the candidate signals as well as their required changes to correct their behavior. This is achieved by using symbolic simulation and equivalence checking between an instruction-set architecture model of the processor and its erroneous model at micro-architecture level. Then, we try to map the corrected functionality into the existing LUT topology. This is realized by a novel method that formulates the problem as a QBF (Quantified Boolean Formula) problem, and solves it by repeatedly applying normal SAT solvers incrementally instead of QBF solvers utilizing ideas from CEGAR (Counter Example Guided Abstraction Refinement) paradigm. We show effectiveness as well as efficiency of our method by correcting bugs in two complex out-of-order superscalar processors with a timing error recovery mechanism.
Satoshi Jo, Amir Masoud Gharehbaghi, Takeshi Matsumoto, Masahiro Fujita 0004
FPT4
2013 Fast simulation of Digital Spiking Silicon Neuron model employing reconfigurable dataflow computing
abstract
A new simulation scheme of the Digital Spiking Silicon Neuron (DSSN) model is proposed. This scheme is based on the reconfigurable dataflow computing paradigm and targets the Maxeler MaxWorkstation. Compared to the previous implementation of the DSSN network, the new scheme has the virtues of better flexibility and better programmability. More importantly, computing with dataflow cores takes good advantage of the intrinsic parallelism of the reconfigurable hardware and better pipelining is achievable. The proposed scheme has good potential of conducting large-scale and fast simulation of the DSSN-model-based network which is pivotal to future neuroscience research.
Xiangyu Li 0005, Shridhar Choudhary, Ray C. C. Cheung, Takeshi Matsumoto, Masahiro Fujita 0004
FPT5
2013 Special session 4B: Elevator talks
abstract
Start of the "Special session 4B: Elevator talks" section of the conference record.
Jennifer Dworak, R. D. (Shawn) Blanton, Masahiro Fujita 0004, Kazumi Hatayama, Naghmeh Karimi, Michail Maniatakos, Antonis M. Paschalis, Adit D. Singh
VTS3
2013 On the integration of model-driven design and dynamic assertion-based verification for embedded software
Giuseppe Di Guglielmo, Luigi Di Guglielmo, Andreas Foltinek, Masahiro Fujita 0004, Franco Fummi, Cristina Marconcini, Graziano Pravadelli
J. Syst. Softw.4
2012 Error Model Free Automatic Design Error Correction of Complex Processors Using Formal Methods
abstract
This paper presents a method for automatic diagnosis and correction of design bugs in processors. Given a golden sequential instruction-set architecture model of a processor and its erroneous detailed cycle-accurate model at micro-architecture level, we employ a symbolic simulator and a property checker in an iterative process to formally find the candidate buggy locations and their corresponding fixes, without requiring an error model. We have shown the effectiveness of our method on a complex out-of-order super scalar processors supporting atomic execution.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
Asian Test Symposium2
2012 Automatic rectification of design errors in complex processors with programmable hardware
abstract
In this paper, we address the problem of automatic correction of design errors in microprocessors, when the correction function is implemented in lookup tables. The formal specification of an erroneous processor and its reference instruction-set architecture (ISA) model are used in an iterative process, employing formal methods, to find and fix the bugs. Then, the correction function is optimized to reduce the lookup table size. We have shown the effectiveness of our method by correcting the bugs in two complex out-of-order superscalar processors with two different timing error recovery mechanisms.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
FPT2
2012 SEU tolerant robust memory cell design
abstract
The implementation of semiconductor circuits and systems in nano-technology makes it possible to achieve high speed, lower voltage level and smaller area. The unintended and undesirable result of this scaling is that it makes integrated circuits susceptible to soft errors normally caused by alpha particle or neutron hits. These events of radiation strike resulting into bit upsets referred to as single event upsets(SEU), become increasingly of concern for the reliable circuit operation in the field. Storage elements are worst hit by this phenomenon. As we further scale down, there is greater interest in reliability of the circuits and systems, apart from the performance, power and area aspects. In this paper we propose an improved 12T SEU tolerant SRAM cell design. The proposed SRAM cell is economical in terms of area overhead. It is easy to fabricate as compared to earlier designs. Simulation results show that the proposed cell is highly robust, as it does not flip even for a transient pulse with 62 times the Qcritof a standard 6T SRAM cell.
Mohammed Shayan, Virendra Singh, Adit D. Singh, Masahiro Fujita 0004
IOLTS4
2012 Time-Constraint-Aware Optimization of Assertions in Embedded Software
Viacheslav Izosimov, Giuseppe Di Guglielmo, Michele Lora, Graziano Pravadelli, Franco Fummi, Zebo Peng, Masahiro Fujita 0004
J. Electron. Test.7
2011 Optimization of Assertion Placement in Time-Constrained Embedded Systems
abstract
We present an approach for optimization of assertion placement in time-constrained HW/SW modules for detection of errors due to transient and intermittent faults. During the design phases, these assertions have to be inserted into the executable code and, hence, will always be executed with the corresponding code branches. As the result, they can significantly increase execution time of a module, in particular, contributing to a much longer execution of the worst case, and cause deadline misses. Assertions have different characteristics such as tightness (or "local error coverage") and execution latency. Taking into account these properties can increase efficiency of assertion checks in time-constrained embedded HW/SW modules. We have developed a design optimization framework, which (1) identifies candidate locations for assertions, (2) associates a candidate assertion to each location, and (3) selects a set of assertions in terms of performance degradation and assertion tightness. Experimental results have shown the efficiency of the proposed techniques.
Viacheslav Izosimov, Michele Lora, Graziano Pravadelli, Franco Fummi, Zebo Peng, Giuseppe Di Guglielmo, Masahiro Fujita 0004
ETS7
2011 EFSM-based model-driven approach to concolic testing of system-level design
abstract
State-of-the-art approaches for testing of system-level design of embedded systems generally work at source-code level, thus they require an implementation of the system to be tested. For this reason, they cannot be applied in the context of model-driven design, where code is available only at end of the design process. Moreover, traditional approaches based on combined concrete and symbolic execution (concolic) suffer two main drawbacks: they are limited in width and depth of the search and not corner-cases oriented. To address such limitations, this paper presents a concolic testing approach for model-driven design of embedded systems. It explores a model of the system, i.e., the extended finite state machine (EFSM), and it relies on weight-oriented analysis of the EFSM paths to achieve high controllability of EFSM transitions, by interleaving longrange concrete approach with symbolic multi-level backjumping strategy. The experimental evaluation on several case studies demonstrates the competitiveness of the proposed approach, which achieves higher transition and instruction coverage than other approaches in significantly reduced time.
Giuseppe Di Guglielmo, Masahiro Fujita 0004, Franco Fummi, Graziano Pravadelli, Stefano Soffia
MEMOCODE2
2010 Aggressive overclocking support using a novel timing error recovery technique on FPGAs (abstract only)
abstract
Clock period of pipelined designs are usually determined by the critical paths to avoid timing errors and guarantee reliable operations. The worst case delays of the slowest pipeline stages cause the clock frequency to be less than the average path delays. This may introduce enormous performance loss if the critical paths rarely happen, and the critical path delays are far larger than the average path delays, which is common for many pipelined circuits. In this paper we present a novel timing error recovery technique that guarantees reliable operation of pipelined designs in presence of any arbitrary number of timing errors in different pipeline stages. We allow the clock frequency to be higher than the worst case; hence increasing the performance. We demonstrate the usefulness of our technique by implementing a pipelined arithmetic circuit with the proposed technique on top of a FPGA board. Our experimental results show that we could successfully increase the clock frequency by 30% with the timing error rate of 13%, all of which are automatically corrected with negligible performance penalty. The timing error recovery circuits need extra flipflops for timing error detection and correction. Since typical FPGAs have LUT with flipflops, the extra area for additional flipflops is minimized as the experimental results have shown. With sophisticated synthesis algorithms for pipelined arithmetic circuits which balanced path delays in the circuits, significantly more performance improvements can be expected.
Amir Masoud Gharehbaghi, Bijan Alizadeh, Masahiro Fujita 0004
FPGA3
2010 SEU tolerant SRAM for FPGA applications
abstract
Modern integrated circuits require careful attention to the soft errors resulting into bit upsets, which are normally caused by alpha particle or neutron hits. These events, also referred to as single-event upsets (SEUs), will become more severe for future technologies. LUT-based FPGAs are heavily using SRAM and there is a growing concern on correct operations of such FPGAs. Although there have been researches on enhancing fault tolerance of such FPGAs, they are based on TMR (triple modular redundancy) and simply too costly for normal application. In this paper we propose a novel 10T SEU tolerant SRAM cell and discuss its modifications for storage of configuration bits in FPGA so that reasonable protection against soft errors can be achieved with small area increase.
Sudipta Sarkar, Anubhav Adak, Virendra Singh, Kewal K. Saluja, Masahiro Fujita 0004
FPT5
2009 Debugging from high level down to gate level
abstract
C-based hardware designs are now accepted as means to increase design productivity. Starting with rather algorithmic design descriptions, incremental refinements are applied to generate high-level synhesizable descriptions which are further processed by high-level and logic synthesis tools. C-based system level design descriptions, such as in SpecC [?] and SystemC [?], can give concise and global views on the behaviors of the designs as well as structures, and various types of dependencies, such as control, data, concurrency, and others, can be extracted quickly. These dependencies can be the bases for efficient and effective debugging for all levels of design descriptions. In this paper, graph representations for various dependencies which are extracted from C-based descriptions are introduced. Then techniques on their uses for debugging in various design levels are discussed. We present static and dynamic tracing methods for dependence analysis as well as techniques that try to establish mapping between implementations and C-based design descriptions.
Masahiro Fujita 0004, Yoshihisa Kojima, Amir Masoud Gharehbaghi
DAC1
2009 Transaction-based debugging of system-on-chips with patterns
abstract
This paper presents a debug method for system communications in post-silicon verification. First, we extract transaction sequences at run-time using on-chip circuits and store them in a trace buffer. Then, we read the stored transactions and analyze them with software. The analysis software tries to find certain patterns in the extracted transactions that are defined by our transaction debug pattern specification language (TDPSL). We have also defined a number of standard patterns for common communication problems such as race and deadlock in TDPSL. To show the feasibility of the method, it is applied to a number of on chip buses. It is shown that the area overhead of the method is very low. Also we have implemented the analysis software and shown that it is memory efficient, scalable and effective to find bugs. The proposed method can also be applied to fault analysis including transient faults.
Amir Masoud Gharehbaghi, Masahiro Fujita 0004
ICCD2
1996 Solving the net matching problem in high-performance chip design
abstract
In high-performance chip design, the problem of net matching is often critical for achieving correct circuit performance. We adopt a conservative design, to route all matched nets with identical topologies and equal wire lengths to achieve zero skew. The problem is formulated as a variant of the D-dimensional Steiner tree problem. We propose a two-stage solution. The first stage uses an iterative improvement strategy to generate the Steiner tree topology for all the nets. The second stage places the nodes using one of two methods. The first approach expresses the optimal Steiner node positions as a linear programming solution, with average computational complexity O(n/sup 2/m/sup 2/), where n is the number of nets and m is the number of pins. Improved efficiency is achieved under the other approach by transforming the Manhattan metric to an l/sub /spl infin// norm using a 45/spl deg/ rotation of the solution space. The norm is then approximated by either an l/sub /spl lambda// norm, for suitably large values of /spl lambda/, or an exponential "penalty" function. The solution space in both approaches becomes strictly convex, allowing us to apply a greedy approach which converges to an optimal solution with great efficiency, leading to a dramatic speed-up versus the linear programming approach.
Robert J. Carragher, Chung-Kuan Cheng, Xiao-Ming Xiong, Masahiro Fujita 0004, Ramamohan Paturi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1995 Simple tree-construction heuristics for the fanout problem
abstract
We address in this paper the fanout tree problem introduced by Berman, et. al., that is using buffer fanout trees to reduce the fanout delay in a technology mapped network. We construct two basic types of fanout trees and provide simple techniques to manipulate them for further delay reduction. These trees are inserted along critical paths throughout the network. We also perform gate-transformation, that is substitution of a gates of equivalent logical functions, if the technology permits. Experimental results show improvement over Touati's LT-tree construction technique.
Robert J. Carragher, Masahiro Fujita 0004, Chung-Kuan Cheng
ICCD2
1993 An efficient algorithm for the net matching problem
abstract
Net matching is often critical for the correct performance of a circuit. We adopt a conservative design to route all matched nets with identical topology and equal wire lengths in order to achieve zero skew. We formulate the problem as one special D-dimensional Steiner tree problem. An iterative improvement strategy is used to generate the Steiner tree topology. We provide a more efficient method of computing the optimal Steiner point positions over the linear programming method involving a 45-degree rotation of the domain, and approximation matrices.
Robert J. Carragher, Chung-Kuan Cheng, Masahiro Fujita 0004
ICCAD3