Hidetoshi Onodera

dblp:37/1497 · DBLP profile ↗
← Back
77ranked-venue papers
4as first author
3since 2021 · last 2022
0000-0001-5198-0668ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 77 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 NCTUcell: A DDA- and Delay-Aware Cell Library Generator for FinFET Structure With Implicitly Adjustable Grid Map
abstract
For the 7-nm technology node, cell placement with a drain-to-drain abutment (DDA) requires additional filler cells, increasing the placement area. This is the first work to fully automatically synthesize a DDA-aware cell library with the optimized number of drains on cell boundary based on ASAP 7-nm PDK. We propose a DDA-aware dynamic programming-based transistor placement. Previous works ignore the use of the M0 layer in cell routing. We first propose an ILP-based M0 routing planning. With M0 routing, the congestion of M1 routing can be reduced and the pin accessibility (PA) can be improved due to the diminished use of M2 routing. We also present a quadratic-programming based-coupling-capacitance-aware initial routing to optimize cell delay, cell area, and M2 usage. To improve the routing resource utilization, we propose an implicitly adjustable grid map, making the maze routing able to explore more routing solutions. The experimental results show that block placement using the DDA-aware cell library requires fewer filler cells than that using the traditional cell library by 25.1%, which achieves a block area reduction rate of 0.97%.
Yih-Lang Li, Shih-Ting Lin, Shinichi Nishizawa, Hong-Yan Su, Ming-Jie Fong, Oscar Chen, Hidetoshi Onodera
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 Tamper-Resistant Optical Logic Circuits Based on Integrated Nanophotonics
abstract
A tamper-resistant logical operation method based on integrated nanophotonics is proposed focusing on electromagnetic side-channel attacks. In the proposed method, only the phase of each optical signal is modulated depending on its logical state, which keeps the power of optical signals in optical logic circuits constant. This provides logic-gate-level tamper resistance which is difficult to achieve with CMOS circuits. An optical implementation method based on electronically-controlled phase shifters is then proposed. The electrical part of proposed circuits achieves 300 times less instantaneous current change, which is proportional to intensity of the leaked electromagnetic wave, than a CMOS logic gate.
Jun Shiomi, Shuya Kotsugi, Boyu Dong, Hidetoshi Onodera, Akihiko Shinya, Masaya Notomi
DAC4
2021 CDF Distance Based Statistical Parameter Extraction Using Nonlinear Delay Variation Models
abstract
This paper proposes a parameter extraction method by comparing the cumulative distribution functions (CDF) between measurement and model-based estimation. We propose a nonlinear delay variation model for fast Monte Carlo simulation to obtain CDFs. We demonstrate the validity of our method by extracting within-die and random telegraph noise induced threshold voltage variations using measured data obtained from a 65 nm test structure. Our proposed method can accurately extract the statistical parameters and can reproduce the measured delay variations by simulation.
Kensuke Murakami, Islam A. K. M. Mahfuzul, Hidetoshi Onodera
IOLTS3
2020 On-chip Memory Optimized CNN Accelerator with Efficient Partial-sum Accumulation
abstract
In convolutional neural networks (CNNs), data movement inside convolution layers between memory and PEs is most energy dominant. This paper proposes a convolution processing dataflow that reduces both the number of memory accesses and the on-chip buffer capacity for convolution operations. Based on the dataflow, we design an on-chip buffer-minimized CNN accelerator. Compared with the state-of-the-art CNN accelerator, the proposed CNN accelerator utilizes 2.30 times less on-chip buffer and 2.18 times energy efficiency to achieve the same data throughput under Alexnet. The proposed architecture is able to achieve higher data throughput with the almost constant on-chip buffer capacity.
Hongjie Xu, Jun Shiomi, Hidetoshi Onodera
ACM Great Lakes Symposium on VLSI3
2020 MCell: Multi-Row Cell Layout Synthesis with Resource Constrained MAX-SAT Based Detailed Routing
abstract
Multi-row cell structure has become popular for modern designs, especially for the multi-bit flip-flop (MBFF) cells, but has not been under full investigation in previous cell library synthesis researches. In this work, we propose an entire placement and routing flow for synthesizing multi-row cell layouts. The proposed new A*-based multi-row transistor placement algorithm can optimize the intra-row and inter-row connections. We also present the first MAX-SAT based detailed router to optimize the cross-row connections that also conform to primitive design rules, but not only to obtain a legal routing result in previous SAT-based detailed router. Experimental results show that the quality of synthesized cells is similar to that of a state-of-the-art cell library in [6], and better aspect ratios of multi-row cells also offer more flexible capability in assembling block designs under some aspect ratio constraints as compared to single-row cell library.
Yih-Lang Li, Shih-Ting Lin, Shinichi Nishizawa, Hidetoshi Onodera
ICCAD4
2019 BDD-based synthesis of optical logic circuits exploiting wavelength division multiplexing
abstract
Optical circuits using nanophotonic devices attract significant interest due to its ultra-high speed operation. As a consequence, the synthesis methods for the optical circuits also attract increasing attention. However, existing methods for synthesizing optical circuits mostly rely on straight-forward mappings from established data structures such as Binary Decision Diagram (BDD). The strategy of simply mapping a BDD to an optical circuit sometimes results in an explosion of size and involves significant power losses in branches and optical devices. To address these issues, this paper proposes a method for reducing the size of BDD-based optical logic circuits exploiting wavelength division multiplexing (WDM). The paper also proposes a method for reducing the number of branches in a BDD-based circuit, which reduces the power dissipation in laser sources. Experimental results obtained using a partial product accumulation circuit in parallel multipliers demonstrates significant advantages of our method over existing approaches in terms of area and power consumption.
Ryosuke Matsuo, Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera, Akihiko Shinya, Masaya Notomi
ASP-DAC4
2019 NCTUcell: A DDA-Aware Cell Library Generator for FinFET Structure with Implicitly Adjustable Grid Map
abstract
For 7nm technology node, cell placement with drain-to-drain abutment (DDA) requires additional filler cells, increasing placement area. This is the first work to fully automatically synthesize a DDA-aware cell library with optimized number of drains on cell boundary based on ASAP 7nm PDK. We propose a DDA-aware dynamic programming based transistor placement. Previous works ignore the use of M0 layer in cell routing. We firstly propose an ILP-based M0 routing planning. With M0 routing, the congestion of M1 routing can be reduced and the pin accessibility can be improved due to the diminished use of M2 routing. To improve the routing resource utilization, we propose an implicitly adjustable grid map, making the maze routing able to explore more routing solutions. Experimental results show that block placement using the DDA-aware cell library requires less filler cells than that using traditional cell library by 70.9%, which achieves a block area reduction rate of 5.7%.
Yih-Lang Li, Shih-Ting Lin, Shinichi Nishizawa, Hong-Yan Su, Ming-Jie Fong, Oscar Chen, Hidetoshi Onodera
DAC7
2019 Area-efficient fully digital memory using minimum height standard cells for near-threshold voltage computing
Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera
Integr.3
2018 PVT2: process, voltage, temperature and time-dependent variability in scaled CMOS process
abstract
In addition to the conventional PVT (Process, Voltage and Temperature) variation, time-dependent current fluctuation such as random telegraph noise (RTN) poses a new challenge on VLSI reliability. In this paper, we show that compared with the static random variation, RTN amplitude of a particular device is not constant across supply voltages and temperatures. A device may show large RTN amplitude at one operating condition and small amplitude at another operating condition. As a result, RTN amplitude distribution becomes uncorrelated across a wide range of voltage and temperature. The emergence of uncorrelated distribution causes significant degradation of worst-case values. Analysis results based on variability models from a 65 nm silicon-on-insulator process show that uncorrelated RTN degrades the worst-case threshold voltage value significantly compared with that where RTN is not considered. Delay variation analysis shows that consideration of RTN in the statistical analysis have little impact at high supply voltage. However, at low voltage operation, RTN can degrade the worst-case value by more than 5%.
Islam A. K. M. Mahfuzul, Hidetoshi Onodera
ICCAD2
2018 Independent N-Well And P-Well Biasing For Minimum Leakage Energy Operation
abstract
This paper proposes a method for minimizing leakage energy consumption under a specific supply voltage and a delay constraint by independently tuning threshold voltages of nMOSFETs and pMOSFETs with body-biasing. We first show a necessary and sufficient condition for the minimum leakage energy operation of a circuit under a delay constraint. We next show that the condition can be identified by a ratio of the leakage currents drawn through an nMOSFET and a pMOSFET in a leakage monitor circuit integrated with the targeting circuit. The leakage current ratio can be monitored at runtime using the leakage monitor. Assuming a constant supply voltage, it is thus possible to minimize the total energy consumption of the circuit by independent tuning of n-well and p-well bias voltages so that the leakage current ratio tracks the predetermined value while keeping the delay constraint. The proposed strategy is experimentally verified by measurements using a 32-bit RISC processor integrating the leakage monitor on the same die fabricated with a 65 nm CMOS process.
Yosuke Okamura, Tohru Ishihara, Hidetoshi Onodera
IOLTS3
2018 Via-Switch FPGA: Highly Dense Mixed-Grained Reconfigurable Architecture With Overlay Via-Switch Crossbars
abstract
This paper proposes a highly dense reconfigurable architecture that introduces via-switch device, which is a nonvolatile resistive-change switch and is used in crossbar switches. Via-switch is implemented in back-end-of-line layers only, and hence the front-end-of-line (FEoL) layers under the crossbar can be fully exploited for highly dense logic blocks. The proposed architecture uses the FEoL layers for fine-grained lookup tables and coarse-grained arithmetic/memory units for improving performance and compatibility with various applications. A case study of application mapping shows the proposed architecture can reduce the array area by 21.7%, thanks to the bidirectional interconnection. Thanks to F2footprint and one order of magnitude lower resistivity of via-switch compared to MOS switch, the crossbar density is improved by up to 26× and the delay and energy in the interconnection are reduced by 90% and 94% at 0.5-V operation.
Hiroyuki Ochi, Kosei Yamaguchi, Tetsuaki Fujimoto, Junshi Hotate, Takashi Kishimoto, Toshiki Higashi, Takashi Imagawa, Ryutaro Doi, Munehiro Tada, Tadahiko Sugibayashi, Wataru Takahashi 0002, Kazutoshi Wakabayashi, Hidetoshi Onodera, Yukio Mitsuyama, Jaehoon Yu, Masanori Hashimoto
IEEE Trans. Very Large Scale Integr. Syst.13
2016 A closed-form stability model for cross-coupled inverters operating in sub-threshold voltage region
abstract
A cross-coupled inverter which is an essential element of on-chip memory subsystems plays an important role in synchronous LSI circuits. In this paper, an analytical stability model for a cross-coupled inverter operating in a sub-threshold voltage region is proposed. The proposed model analytically shows that the minimum operating voltage of the cross-coupled inverter distributes normally in a high-s region if the distribution of the threshold voltage is Gaussian. The minimum supply voltage at which the yield of the cross-coupled inverter becomes a specific value can be accurately derived by a simple calculation using the model. Monte-Carlo simulation assuming a commercial 28 nm process technology demonstrates the accuracy and the validity of the proposed model. Based on the model, this paper shows strategies for variation tolerant memory design.
Tatsuya Kamakari, Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera
ASP-DAC4
2016 On-chip monitoring and compensation scheme with fine-grain body biasing for robust and energy-efficient operations
abstract
Aggressive technology scaling and strong demand for lowering supply voltage impose a serious challenge in achieving robust and energy-efficient circuit operation. This paper first overviews on device-circuit interactions to enable cross-layer resiliency, and energy optimization. We show that the ability to monitor and control device and circuit characteristics not only increase energy-efficiency by more than 20% but also relax the severe design constraints, which were required because of the uncertainties of variability. We then demonstrate two proof-of-concept circuits in a 65 nm process to show variability resiliency and energy optimization with local body biasing.
Islam A. K. M. Mahfuzul, Hidetoshi Onodera
ASP-DAC2
2016 A highly-dense mixed grained reconfigurable architecture with overlay crossbar interconnect using via-switch
abstract
This paper proposes a highly-dense reconfigurable architecture that introduces via-switch device, which is a kind of resistive RAM and is used in crossbar switches. Since via-switch is implemented in BEOL layers only, the FEOL layer under the crossbar can be fully exploited for highly-dense logic blocks. The proposed architecture uses the FEOL layer for fine-grained look-up tables and coarse-grained arithmetic/memory units for better performance and highly wide applications. In a case study of application mapping, the proposed architecture reduces array area by 76% thanks to mixed grained logic structure and overlay bidirectional interconnection. Thanks to 18F2footprint and one order of magnitude lower resistivity of via-switch compared to MOS switch, the crossbar density is improved by 26× and the delay and energy in the interconnection are reduced by 90% and 93% at 0.5V operation.
Junshi Hotate, Takashi Kishimoto, Toshiki Higashi, Hiroyuki Ochi, Ryutaro Doi, Munehiro Tada, Tadahiko Sugibayashi, Kazutoshi Wakabayashi, Hidetoshi Onodera, Yukio Mitsuyama, Masanori Hashimoto
FPL9
2015 Reliability-configurable mixed-grained reconfigurable array compatible with high-level synthesis
abstract
This paper presents a mixed-grained reconfigurable VLSI array architecture that can cover mission-critical applications to consumer products through C-to-array application mapping. A proof-of-concept VLSI chip was fabricated in a 65nm process. Measurement results show that applications on the chip can be working in a harsh radiation environment.
Masanori Hashimoto, Dawood Alnajiar, Hiroaki Konoura, Yukio Mitsuyama, Hajime Shimada, Kazutoshi Kobayashi, Hiroyuki Kanbara, Hiroyuki Ochi, Takashi Imagawa, Kazutoshi Wakabayashi, Takao Onoye, Hidetoshi Onodera
ASP-DAC12
2015 Microarchitectural-level statistical timing models for near-threshold circuit design
abstract
Near-threshold computing has emerged as a promising solution for drastically improving the energy efficiency of microprocessors. This paper proposes architectural-level statistical static timing analysis (SSTA) models for the near-threshold voltage computing where the path delay distribution is approximated as a lognormal distribution. First, we prove several important theorems that help consider architectural design strategies for high performance and energy efficient near-threshold computing. After that, we show the numerical experiments with Monte Carlo simulations using a commercial 28-nm process technology model and demonstrate that the properties presented in the theorems hold for the practical near-threshold logic circuits.
Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera
ASP-DAC3
2014 Variability and Soft-Error Resilience in Dependable VLSI Platform
abstract
Extreme scaling imposes enormous challenges, such as variability increase and soft-error vulnerability, on the resilience of VLSI circuits and systems. For coping with those threats, we have been developing a VLSI platform that can realize a dependable circuit with required level of reliability. In the platform, circuit-level resilience to variability is achieved by on-chip performance monitoring and variability compensation by localized body biasing. Architecture-level resilience to soft-errors is accommodated by a mixed-grained reconfigurable array in which functionality as well as reliability can be configured. Those properties have been experimentally verified by proof-of-concept chips in 65 nm process. Overview of the variability and soft-error resilience of the platform will be explained, followed by experimental demonstrations.
Yukio Mitsuyama, Hidetoshi Onodera
ATS2
2014 A 65-nm CMOS burst-mode CDR based on a GVCO with symmetric loops
abstract
A 12.5-Gb/s burst-mode clock and data recovery (BCDR) circuit based on a simple gated voltage-controlled oscillator (GVCO) is presented. A simple symmetric circuit topology makes the area for the GVCO smaller and leads to an easier timing design. The GVCO consists of two loops which operate complementarily. The same type of circuit configurations are adopted for AND and OR in the loops to reduce the difficulties in the timing alignment of the signals from the loops. To confirm the validity of the proposed topology, we fabricated a 12.5-Gb/s-BCDR IC with the 65-nm-MOSFET process. Without a circuit for precise timing adjustment for the signals in the two loops, the IC provides instantaneous phase locking of 1 bit for burst data input of 12.5 G/s. The measured jitter is lower than 2 ps rms. The area and the power consumption for the core GVCO are 0.03 mm2and 60 mW, respectively.
Keiji Kishine, Hiroshi Inoue, Hiromi Inaba, Makoto Nakamura, Akira Tsuchiya, Hidetoshi Onodera, Hiroaki Katsurai
ISCAS6
2014 Frequency-Independent Warning Detection Sequential for Dynamic Voltage and Frequency Scaling in ASICs
abstract
In this paper, a metastability immune warning flip-flop (FF) is proposed, which consists of an edge detector, a warning window generator, and a warning detector along with a traditional FF. The delayed data are monitored during the warning window to flag a warning signal before the data enter the erroneous zone. In this scheme, the warning window is independent of input clock frequency and hence is suitable for frequency scaling application. A 16-bit Kogge-stone adder is implemented in 65-nm technology, which uses warning FF for dynamic voltage and frequency scaling (DVFS). The warning FF-based DVFS allows elimination of safety margins and operates till the point of first warning of the adder without any erroneous results. The experiments were conducted with different supply voltages, phase-shifted clocks, and process conditions. The circuit is helpful to determine when to stop further reduction in supply voltage by producing the warning signal with predefined timing slacks in DVFS application. The test chip results demonstrate that the proposed circuit can track the critical path delay of 2.4-7.5 ns at warning voltage of 1.15-0.72 V, respectively. The measured results from 10 different chips show the effectiveness of the proposed concept across process variation.
Bishnu Prasad Das, Hidetoshi Onodera
IEEE Trans. Very Large Scale Integr. Syst.2
2013 A 25-Gb/s LD driver with area-effective inductor in a 0.18-µm CMOS
abstract
This paper presents high-speed and area-efficient laser-diode driver with interwoven inductor in a 0.18-μm CMOS. We interweave ten peaking inductors for area-effective implementation as well as performance enhancement. Interwoven inductor can not only achieve area-efficiency but also tune frequency characteristic. Mutual inductances of interwoven inductor enhance bandwidth and suppress group delay dispersion. The test chip area is 0.32 mm2and the maximum operating speed is 25 Gb/s.
Takeshi Kuboki, Yusuke Ohtomo, Akira Tsuchiya, Keiji Kishine, Hidetoshi Onodera
ASP-DAC5
2013 Dependable VLSI Platform using Robust Fabrics
abstract
Technology scaling and growing complexity have an increasing impact on the resilience of VLSI circuits and systems. Severe challenges have been emerging for the realization of dependable VLSI circuits and systems with necessary and sufficient amount of reliability and security. For coping with the increasing threats on manufacturability, variability, and transient (soft) errors, we have been working on the development of “Dependable VLSI Platform using Robust Fabrics.” The project tackles the challenges with collaborative researches on layout, circuit, architecture, and design automation. Overview of the project as well as key achievements on the component-level (Fabrics) and the architecture-level (reconfigurable architecture) will be explained, followed by a brief introduction of the platform SoC and its C-based design tools.
Hidetoshi Onodera
ASP-DAC1
2013 Perturbation-immune radiation-hardened PLL with a switchable DMR structure
abstract
This paper proposes a perturbation-immune radiation-hardened PLL with a switchable dual modular redundancy (DMR) structure. By a radiation-strike, a PLL has clock-perturbation for a while. Conventional RHPLLs are proposed to reduce recovery-time which is the time to recover from perturbation. However, recovery still needs tens of clock cycles. Our proposal is `detecting' and `switching' instead of `recovering' clock-perturbation. For robust perturbation-immunity, detecting speed is important. We identify types of clock-perturbation and - then propose a set of detectors to detect each type. With this, detectors guarantee high speed in detection.
SinNyoung Kim, Akira Tsuchiya, Hidetoshi Onodera
IOLTS3
2012 A 16Gb/s area-efficient LD driver with interwoven inductor in a 0.18µm CMOS
abstract
This paper presents the fastest laser-diode driver with interwoven peaking inductor in 0.18μm CMOS. Six and four inductors are interwoven into two sets of inductors for area-effective implementation as well as performance enhancement. The operation speed enhancement of the proposed circuit is achieved by tuning mutual inductances of interwoven inductors. The circuit area is 0.34-mm2and the maximum operating speed is 16-Gb/s.
Takeshi Kuboki, Yusuke Ohtomo, Akira Tsuchiya, Keiji Kishine, Hidetoshi Onodera
ASP-DAC5
2012 On-Chip Detection of Process Shift and Process Spread for Silicon Debugging and Model-Hardware Correlation
abstract
This paper proposes the use of ROs (Ring Oscillators) for process shift and process spread detection for silicon debugging and model-hardware correlation. ROs are designed to be sensitive to either nMOSFET orpMOSFET variation, thus the location of the chip in the process spacecan be detected directly from the RO measurements. Test chip measurements in a 65-nm process shows the validity of the proposed ROs. Amounts of process shift and process spread for key process parameters as threshold voltages and gate length are extracted from test chip measurements.
Islam A. K. M. Mahfuzul, Hidetoshi Onodera
Asian Test Symposium2
2012 A flexible structure of standard cell and its optimization method for near-threshold voltage operation
abstract
With ever growing demands of mobile devices, low power consumption has become essential for VLSI circuits. Since standard cell libraries are typically used in many parts of VLSI circuits, their performance has a strong impact on realizing high speed and low power VLSI circuits. One of the most promising approaches for reducing the power consumption of the circuit is lowering the supply voltage. However this causes an increase of imbalance between rise and fall delays especially for cells having transistor stacks. For mitigating this imbalance, this paper proposes a structure of standard cells where the P/N ratio of each cell can be independently customized for near-threshold operation in VLSI circuits. The structure cancels the imbalance between rise and fall delays at the expense of cell area. The experiments with ISCAS'85 benchmark circuits demonstrate that the standard cell library consisting of the proposed cells reduces the power consumption of the benchmark circuits by 16% on average without increasing the circuit area, compared to that of the same circuit synthesized with a library which is not optimized for the near-threshold operation.
Shinichi Nishizawa, Tohru Ishihara, Hidetoshi Onodera
ICCD3
2011 A 65nm flip-flop array to measure soft error resiliency against high-energy neutron and alpha particles
abstract
We fabricated a 65nm LSI including flip-flop array to measure soft error resiliency against high-energy neutron and alpha particles. It consists of two FF arrays as follows. One is an array composed of redundant FFs to confirm radiation hardness of the proposed and conventional redundant FFs. The other is an array composed of conventional D-FFs to measure SEU (Single Event Upset) and MCU(Multiple Cell Upset) by the distance from tap cells.
Jun Furuta, Chikara Hamanaka, Kazutoshi Kobayashi, Hidetoshi Onodera
ASP-DAC4
2011 Dependable VLSI Program in Japan: Program Overview and the Current Status of Dependable VLSI Platform Project
abstract
Technology scaling and growing complexity have an increasing impacton the resilience of VLSI circuits and systems. Severe challenges have been emerging for the realization of dependable VLSI circuits and systems with required reliability and security. For coping with the increasing threats to dependability of VLSIs, the Dependable VLSI (DVLSI) program has been established in 2007 and 11 projects are now in progress. This paper gives an overview of the DVLSI program, followed by a brief introduction of one of the 11 projects entitled ``Dependable VLSI Platform using Robust Fabrics''.
Hidetoshi Onodera
Asian Test Symposium1
2009 Dependable VLSI: device, design and architecture: how should they cooperate?
abstract
VLSI dependability is one of the most significant issues in the modern world. Here the panelists will discuss the key technologies for it as well as the cost optimization among device, design and architecture.
Shuichi Sakai, Hidetoshi Onodera, Hiroto Yasuura, James C. Hoe
ASP-DAC2
2008 Statistical gate delay model for Multiple Input Switching
abstract
In this paper, we propose a calculation method of gate delay for SSTA (Statistical Static Timing Analysis) considering MIS (Multiple Input Switching). Most SSTA approaches assume a single input switching model and ignore the effect of MIS on gate delay. MIS occurs when multiple inputs of a gate switch nearly simultaneously. Thus, ignoring MIS causes error in MAX operation in SSTA. We propose a statistical gate delay model considering MIS. We verify the proposed method by SPICE based Monte Carlo simulations and experimental results show that the proposed method improves the error due to ignoring MIS.
Takayuki Fukuoka, Akira Tsuchiya, Hidetoshi Onodera
ASP-DAC3
2008 Best ways to use billions of devices on a chip - Error predictive, defect tolerant and error recovery designs
abstract
Error rates on an LSI are increasing according to the Moore’s law. Now is the time to start incorporating error-tolerant design methodologies. This paper introduces sources of failures in semiconductor devices, levels of dependability according to applications of devices and some circuit-level techniques to detect or recover faults after shipping.
Kazutoshi Kobayashi, Hidetoshi Onodera
ASP-DAC2
2008 Speed and yield enhancement by track swapping on critical paths utilizing random variations for FPGAs
abstract
FPGAs in future deep submicron fabrication process will suffer from drastic speed and yield loss caused by device variations. We propose variation-aware reconfiguration which utilizes variations for performance enhancement. To utilize random variations for performance enhancement, optimizing each device from a common initial configuration is better than producing optimized configurations according to detailed measurement results because it is very hard to measure detailed variation maps chip by chip when random uncorrelated variations are dominant. In the critical path reconfiguration scheme, an initial configuration is gradually optimized chip by chip according to the delay variations. We apply the track swapping procedure to critical path reconfiguration which obtains an optimized configuration to repeat measurement and reconfiguration. First we configure all fabricated FPGAs with a common configuration data without considering variations. The configuration of each die is optimized to reroute the critical paths by choosing a faster path. To reroute a critical path we swap a wire track on a critical path with the adjacent track. It can be realized to use switch blocks with more flexibility. We implement the track swapping to VPR and experiment performance enhancement by applying the track swapping to LGSynth93 benchmark circuits. The average speed and yield enhancements are 2.57%, 26.01% respectively when the standard deviation of random variations is 10.0%
Yuuri Sugihara, Yohei Kume, Kazutoshi Kobayashi, Hidetoshi Onodera
FPGA4
2008 A variation-aware constant-order optimization scheme utilizing delay detectors to search for fastest paths on FPGAS
abstract
We propose a variation-aware post-fabrication optimization scheme on FPGAs. Variation-aware optimization usually takes huge measurement cost. The proposed scheme achieves a constant optimization cost for any circuit configuration. We utilize delay detectors embedded in clustered CLBs to choose fastest paths among multiple candidates. The delay detectors enable simultaneous measurement of critical path candidates to partition all critical paths into segments. The number of measurement to choose fastest paths on all critical paths does not depends on configurations but on FPGA architectures. We confirm that a simple heuristic algorithm can find the order of measurement near the lowest bound of the measurement cost and it is almost constant regardless of circuit configurations.
Kazutoshi Kobayashi, Yohei Kume, Cam Lai Ngo, Yuuri Sugihara, Hidetoshi Onodera
FPL5
2008 Performance optimization by track swapping on critical paths utilizing random variations for FPGAS
abstract
Since FPGAs in future deep sub-micron processes will suffer from drastic speed and yield losses caused by device variations, we propose variation-aware reconfiguration that utilizes these variations for performance enhancement. To utilize random variations on a current deep submicron process for performance enhancement, optimizing each device from a common configuration is better than producing optimized configurations based on detailed measurement results. In this paper we apply a track swapping procedure to critical path reconfiguration. First, we configure all fabricated FPGAs with common configuration data. The configuration of each die is optimized to reroute the critical paths that do not satisfy timing specifications. The rerouting of a critical path usually causes serious topology changes that may prolong other paths and create new critical paths. In the track swapping procedure, we swap a wire track on a critical path for the adjacent track without any topology changes by switching blocks with more flexibility. We experiment on performance enhancement by applying track swapping to LGSynth93 benchmark circuits. The average speed enhancement is 2.45%, and the average yield enhancement is 32.7% when the standard deviation of the random variations is 10.0%.
Yuuri Sugihara, Yohei Kume, Kazutoshi Kobayashi, Hidetoshi Onodera
FPL4
2007 A 10Gbps/channel On-Chip Signaling Circuit with an Impedance-Unmatched CML Driver in 90nm CMOS Technology
abstract
An on-chip signaling system consists of a CML driver, a differential transmission-line and a CML receiver is fabricated. We developed an impedance-unmatched driver for power reduction. The impedance-unmatched driver reduces the tail current of the CML buffer by tuning the load resistance. The designed circuit achieves 3mm, 10Gbps/channel on-chip signal transmission and the impedance-unmatched driver saves the energy per bit by 21% compared with a conventional impedance-matched driver.
Takeshi Kuboki, Akira Tsuchiya, Hidetoshi Onodera
ASP-DAC3
2007 A 90nm 8×16 FPGA Enhancing Speed and Yield Utilizing Within-Die Variations
abstract
We have fabricated an LUT-based FPGA device with functionalities measuring within-die variations in a 90nm process. Measured variations are used to configure each device to maximize the operating frequency by allocating critical paths in faster portions. Variations are measured using ring oscillators implemented as a configuration of the FPGA. Placement optimization using a simple model circuit reveals that performance of the circuit is enhanced by 4% in average, which is the same amount as the measured within-die variations. The yield is enhanced by 32% to the worst case.
Yuuri Sugihara, Manabu Kotani, Kazuya Katsuki, Kazutoshi Kobayashi, Hidetoshi Onodera
ASP-DAC5
2007 Worst-case delay analysis considering the variability of transistors and interconnects
abstract
This paper discusses the condition that gives the statistical worst-casedelay of a stage under the fluctuation of interconnect structure and transistor performance. The delay of a stage is a function of many parameters such as drive strength of the gate, interconnect length and interconnect structures (width, thickness, spacing, etc.), and therefore the condition for the worst-case delay also becomes a function of those parameters. We examine the worst-case condition using a simple equivalent circuit and show how the worst-case condition varies. It is shown that the worst-case condition of an interconnect structure for a certain range of interconnect length moves toward the best-case for other range of interconnect length, and hence it is important to locate the worst-case condition correctly for accurate estimation of the worst-case stage delay. We show a simple criteria for the direction of the worst-case condition of an interconnect structure.
Takayuki Fukuoka, Akira Tsuchiya, Hidetoshi Onodera
ISPD3
2006 Measurement results of within-die variations on a 90nm LUT array for speed and yield enhancement of reconfigurable devices
abstract
It is possible to enhance speed and yield of reconfigurable devices utilizing WID variations. An LUT array LSI is fabricated on a 90nm process to measure WID and D2D variations. Performance fluctuations are measured by counting the number of LUTs through which a signal is passing within a certain time. D2D and WID variations are clearly observed by the measurement
Kazuya Katsuki, Manabu Kotani, Kazutoshi Kobayashi, Hidetoshi Onodera
ASP-DAC4
2006 Interconnect RL extraction at a single representative frequency
abstract
This paper proposes a method to determine a single frequency for interconnect RL extraction. Resistance and inductance of interconnects depend on frequency, and hence the extraction frequency strongly affects the modeling accuracy of interconnects. The proposed method determines an extraction frequency based on the transfer characteristic of interconnects. By choosing the frequency where the transfer characteristic becomes maximum, the extracted RL values achieve the accurate modeling of the waveform. We experimentally verify that the proposed method provides accurate transition waveforms over various interconnect topologies.
Akira Tsuchiya, Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC3
2006 A Yield and Speed Enhancement Technique Using Reconfigurable Devices Against Within-Die Variations on the Nanometer Regime
abstract
A reconfigurable device can be utilized to enhance speed and yield on the sub-100nm device technologies, in which large within-die (WID) variations will degrade speed and cause huge yield loss in conventional fixed-structured ASICs. In the proposed scheme, configurations of all fabricated chips are optimized according to measured intra variations of LUTs and switch matrixes. Two LSIs are fabricated in a 90nm CMOS process. We successfully measured WID variations on the first LUT array LSI. The speed is enhanced by 4.1% in average on the second variation-aware FPGA LSIs to optimize configurations by the measured WID variations
Kazutoshi Kobayashi, Manabu Kotani, Kazuya Katsuki, Y. Takatsukasa, K. Ogata, Yuuri Sugihara, Hidetoshi Onodera
FPL7
2005 Timing analysis considering temporal supply voltage fluctuation
abstract
This paper proposes an approach to cope with temporal power/ground voltage fluctuation for static timing analysis. The proposed approach replaces temporal noise with an equivalent power/ground voltage. This replacement reduces complexity that comes from the variety in noise waveform shape, and improves compatibility of power/ground noise aware timing analysis with conventional timing analysis framework. Experimental results show that the proposed approach can compute gate propagation delay considering temporal noise within 10% error in maximum and 0.5% in average.
Masanori Hashimoto, Junji Yamaguchi, Takashi Sato 0001, Hidetoshi Onodera
ASP-DAC4
2005 A resource-shared VLIW processor architecture for area-efficient on-chip multiprocessing
abstract
We propose an area-efficient resource-shared VLIW processor (RSVP) for future leaky nm process technologies. It consists of several single-way independent processor units (IPUs) that share parallel processor resources. Each IPU works as a variable-way VLIW processor sharing the parallel resources according to priorities of given tasks. RSVP allocates shared parallel resources to the IPUs cycle by cycle. It can minimize the number of NOPs that waste power. The performance per power (P3) of a 4-parallel 4-way RSVP that corresponds to four 4way VLIWs is 3.7% better than a conventional 4-parallel 4-way VLIW multiprocessor in the current 90nm process. We estimate that the RSVP achieves 36% less leakage power and 28% better P3 in the future 25nm process. We have fabricated an RSVP test chip that contains two IPU and a shared resource equivalent to two 2way VLIWs in a 180nm process. It is functional at 100MHz clock speed and its power is 130mW.
Kazutoshi Kobayashi, Masao Aramoto, Yoichi Yuyama, Akihiko Higuchi, Hidetoshi Onodera
ASP-DAC5
2005 Successive pad assignment algorithm to optimize number and location of power supply pad using incremental matrix inversion
abstract
An efficient pad assignment algorithm to minimize voltage drop on a power distribution network is proposed. Combination of the successive pad assignment (SPA) and the incremental matrix inversion (IMI) provides an efficient assignment for both location and number of power supply pads. The SPA creates equivalent resistance matrix which preserves both pad candidates and power consumption points as external ports so that topological modification due to connection or disconnection between voltage sources and candidate pads are consistently represented. By reusing sub-matrix of equivalent matrix, the SPA greedily searches next pad location that minimizes the worst drop voltage. Each time the candidate pad is added, the IMI reduces computational complexity significantly. Experimental results show that the proposed procedures efficiently enumerate pad order in practical time.
Takashi Sato 0001, Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC3
2005 Design and measurement of 6.4 Gbps 8: 1 multiplexer in 0.18µm CMOS process
abstract
We develop and measure a 8:1 multiplexer in a CMOS 0.18μm process. We design the hybrid multiplexer based on a prior detailed performance evaluation both of CMOS static and current mode logic circuits, and build a hybrid structure. The fabricated chip operates at up to 6.4 Gbps with power consumption of 84mW.
Akinori Shinmyo, Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC3
2005 Return path selection for loop RL extraction
abstract
This paper propose a systematic method to select power/ground wires that should be considered in interconnect RL extraction. The return current distribution affects loop characteristic of interconnects. To extract exact RL value, all of return paths have to be considered. However it is impossible because there are huge number of P/G wires in LSIs. As more wires are considered, the extraction accuracy improves but the extraction cost increases undesirably. The proposed method focuses the energy dissipated at P/G wires and utilizes it for screening return paths. Experimental results reveal that our method enables accurate and computationally efficient RL extraction with considering return current distribution.
Akira Tsuchiya, Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC3
2005 Effects of on-chip inductance on power distribution grid
abstract
With increase of clock frequency, on-chip wire inductance starts to play an important role in power/ground distribution analysis, although it has not been considered so far. We perform a case study work that evaluates relation between decoupling capacitance position and noise suppression effect, and we reveal that placing decoupling capacitance close to current load is necessary for noise reduction. We experimentally show that impact of on-chip inductance becomes small when on-chip decoupling capacitance is well placed according to local power consumption. We also examine influences of grid pitch, wire area, and spacing between paired power and ground wires on power supply noise. Minification of grid pitch is more efficient than increase in wire area, and small spacing reduces power noise as we expected.
Atsushi Muramatsu, Masanori Hashimoto, Hidetoshi Onodera
ISPD3
2004 A performance comparison of PLLs for clock generation using ring oscillator VCO and LC oscillator in a digital CMOS process
Takahito Miyazaki, Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC3
2004 Representative frequency for interconnect R(f)L(f)C extraction
Akira Tsuchiya, Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC3
2004 An SoC architecture and its design methodology using unifunctional heterogeneous processor array
Yoichi Yuyama, Masao Aramoto, Kazutoshi Kobayashi, Hidetoshi Onodera
ASP-DAC4
2004 Timing analysis considering spatial power/ground level variation
abstract
Spatial power/ground level variation causes power/ground level mismatch between driver and receiver, and the mismatch affects gate propagation delay. This work proposes a timing analysis method based on a concept called "PG level equalization" which is compatible with conventional STA frameworks. We equalize the power/ground levels of driver and receiver. The charging/discharging current variation due to equalization is compensated by replacing output load. We present an implementation method of the proposed concept, and demonstrate that the proposed method works well for multiple-input gates and RC load models.
Masanori Hashimoto, Junji Yamaguchi, Hidetoshi Onodera
ICCAD3
2004 Equivalent waveform propagation for static timing analysis
abstract
This paper proposes a scheme that captures diverse input waveforms of CMOS gates for static timing analysis (STA). Conventionally latest arrival and transition times are calculated from the timings when a transient waveform goes across predetermined reference voltages. However, this method cannot accurately consider the impact of waveform shape on gate delay when crosstalk-induced nonmonotonic waveforms or inductance-dominant stepwise waveforms are injected. We propose a new timing analysis scheme called "equivalent waveform propagation." The proposed scheme calculates the equivalent waveform that makes the output waveform close to the actual waveform, and uses the equivalent waveform for timing calculation. The proposed scheme can cope with various waveforms affected by resistive shielding, crosstalk noise, wire inductance, etc. In this paper, we devise a method to calculate the equivalent waveform. The proposed calculation method is compatible with conventional methods in gate delay library and characterization and, hence, our method is easily implemented with conventional STA tools.
Masanori Hashimoto, Yuji Yamada, Hidetoshi Onodera
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2003 Standard cell libraries with various driving strength cells for 0.13, 0.18 and 0.35 μm technologies
abstract
We developed standard cell libraries for three technologies(0.13, 0.18 and 0.35 μm) using an automatic layout generation tool that we have developed. The developed libraries are as competitive as manually-designed libraries in layout density and speed. We verify the functionalities of all cells and the speed of basic combinational cells on the fabricated chips. The libraries are currently public to educational organizations in Japan.
Masanori Hashimoto, Kazunori Fujimori, Hidetoshi Onodera
ASP-DAC3
2003 A statistical gate delay model for intra-chip and inter-chip variabilities
abstract
This paper proposes a model to calculate statistical gate-delay variation caused by intra-chip and inter-chip variabilities. Our model consists of a statistical transistor model and a gate-delay model. We present a modeling and extracting method of transistor characteristics for the intra-chip variability and the inter-chip variability. In the modeling of the intra-chip variability, it is important to consider a gate-size dependence by which the amount of intra-chip variation is affected. This effect is not captured in a statistical delay analysis reported so far. Our gate-delay model characterizes a statistical gate delay variation using a response surface method (RSM) according to the intra-chip and inter-chip variability of each transistor in a gate. We evaluate the accuracy of our model, and we show some simulated results of a circuit delay variation characterized by the measured variances of transistor currents.
Kenichi Okada 0001, Kento Yamaoka, Hidetoshi Onodera
ASP-DAC3
2003 Equivalent Waveform Propagation for Static Timing Analysis
abstract
This paper proposes a scheme that captures diverse input waveforms of CMOS gates for static timing analysis. Conventionally the latest arrival time and transition time are calculated from the timings when a transient waveform goes across pre-determined reference voltages. However, this method cannot accurately consider the impact of waveform shape on gate delay, when crosstalk-induced non-monotonic waveforms or inductance-dominant step-wise waveforms are injected. We propose a new timing analysis scheme called "equivalent waveform propagation". The proposed scheme calculates the equivalent waveform that makes the output waveform close to the actual waveform, and uses the equivalent waveform for timing calculation. The proposed scheme can cope with various waveforms affected by resistive shielding, crosstalk noise, wire inductance etc. In this paper, we devise a method to calculate equivalent waveform. The proposed calculation method is compatible with conventional methods in gate delay library and characterization, and hence our method is easy to be implemented with conventional static timing analysis tools.
Masanori Hashimoto, Yuji Yamada, Hidetoshi Onodera
ICCAD3
2003 A Statistical Gate-Delay Model Considering Intra-Gate Variability
Kenichi Okada 0001, Kento Yamaoka, Hidetoshi Onodera
ICCAD3
2003 Capturing crosstalk-induced waveform for accurate static timing analysis
abstract
We propose a method to capture crosstalk-induced noisy waveform for crosstalk-aware static timing analysis. The effects of capacitive coupling noise on timing are conventionally measured as delay variation. On the other hand, the propose method derives an equivalent waveform to a crosstalk-induced noisy waveform. The crosstalk effects on timing are all included in the equivalent waveform. With the derived equivalent waveform, we can perform static timing analysis with consideration of dynamic delay variation due to crosstalk noise. The equivalent waveform is derived by our improved least square fitting with weighting coefficient. Our method can naturally consider the slew variation due to crosstalk noise as well as the delay variation. We experimentally verify that our method can estimate the delay variation at the output of the receiver gate accurately. The strength is that the proposed method requires no additional library characterization and is easy to be integrated into usual static timing analysis methods.
Masanori Hashimoto, Yuji Yamada, Hidetoshi Onodera
ISPD3
2002 Crosstalk noise optimization by post-layout transistor sizing
abstract
This paper proposes a post-layout transistor sizing method for crosstalk noise reduction. The proposed method downsizes the drivers of the aggressor wires for noise reduction, utilizing the precise interconnect information extracted from the detail-routed layouts. We develop a transistor sizing algorithm for crosstalk noise reduction under delay constraints, and construct a crosstalk noise optimization method utilizing a crosstalk noise estimation method and a transistor sizing framework which are previously developed. Our method exploits the transistor sizing framework that can vary the transistor widths inside cells with interconnects unchanged. Our optimization method therefore never cause a new crosstalk noise problem, and does not need iterative layout optimization. The effectiveness of the proposed method is experimentally examined using 2 circuits. The maximum noise voltage is reduced by more than 50% without delay increase. These results show that the risk of crosstalk noise problems can be considerably reduced after detail-routing.
Masanori Hashimoto, Masao Takahashi, Hidetoshi Onodera
ISPD3
2001 Post-layout transistor sizing for power reduction in cell-based design
abstract
We propose a transistor sizing method that down-sizes MOSFETs inside a cell to eliminate redundancy of cell-based circuits as much as possible. Our method reduces power dissipation of detail-routed circuits while preserving interconnects. The effectiveness of our method is experimentally evaluated using 5 circuits. The power dissipation is reduced by 77% maximum and 65% on average without delay increase.
Masanori Hashimoto, Hidetoshi Onodera
ASP-DAC2
2001 A vector-pipeline DSP for low-rate videophones
abstract
We propose a vector-pipeline processor VP-DSP for low-rate videophones, which can encode and decode 10 frames/sec. of QCIF through a 29.2kbps low-rate line. We have already fabricated a VP-DSP LSI by a 0.35 um CMOS process. The area of the VP-DSP core is 4.2mm. It works properly at 25MHz/1.6V with the power dissipation of 49mW. Its peak performance is up to 400MOPS, 8.2GOPS/W.
Kazutoshi Kobayashi, Makoto Eguchi, Takuya Iwahashi, Takehide Shibayama, Kousuke Takai, Hidetoshi Onodera
ASP-DAC7
2001 Beyond the red brick wall (panel): challenges and solutions in 50nm physical design
abstract
Aggressive technology scaling will push us into a 50nm regime within a decade. Most entries in current ITRS for the technology node are painted out in red, indicating "No know solutions". In the physical implementation domain, we are facing severe challenges in various aspects such as interconnect performance degradation, signal integrity, reliability, manufacturing variability, etc. These challenges will continue to grow for the future. In this panel, our panelists will present their own view of the most difficult challenges in the 50nm regime, and possible solutions to break through the red brick wall as well, followed by a live discussion on the approaches we should take for successful 50nm physical implementation.
Hidetoshi Onodera, Andrew B. Kahng, Wayne Wei-Ming Dai, Sani R. Nassif, Akira Tanabe, Toshihiro Hattori
ASP-DAC1
2001 A dynamically phase adjusting PLL with a variable delay
abstract
Phase locked loop (PLL) is widely used for many purposes. The lock-up performance is one of the most important target items in designing PLLs. In a digital PLL, it is difficult to control the frequency and phase independently, which makes it difficult to improve lock-up performance. A variable delay circuit which adjusts only the phase of the PLL is introduced here. A full loop model simulation with measured controllable delay shows the effectiveness of applying the phase adjust method with the variable delay to the PLL. Index Terms - PLL, phase adjust, variable delay, lock-up.
Takeo Yasuda, Hiroaki Fujita, Hidetoshi Onodera
ASP-DAC3
2001 Crosstalk Noise Estimation for Generic RC Trees
abstract
We propose an estimation method of crosstalk noise for generic RC trees. The proposed method derives an analytic waveform of crosstalk noise in a 2-/spl pi/ equivalent circuit. The peak voltage is calculated from the closed-form expression, and the crosstalk induced delay is estimated using the derived noise waveform. We also develop a transformation method from generic RC trees with branches into the 2-/spl pi/ model circuit. The proposed method can hence estimate crosstalk noise for any RC trees. Our estimation method is evaluated in a 0.13 /spl mu/m technology. The peak noise of two partially-coupled interconnects is estimated with the average error of 13%. Our method transforms generic RC interconnects with branches into the 2-/spl pi/ model with 14% error on average.
Masao Takahashi, Masanori Hashimoto, Hidetoshi Onodera
ICCD3
2000 A method for linking process-level variability to system performances
Tomohiro Fujita, Kenichi Okada 0001, Hiroaki Fujita, Hidetoshi Onodera, Keikichi Tamaru
ASP-DAC4
2000 Statistical delay calculation with vector synthesis model
abstract
A statistical delay model for CMOS digital circuits called the "vector synthesis model" is proposed. The model provides a relationship between process random variables and a digital circuit path delay. A first order coefficient vector (FOCV), which characterizes the drain current of a transistor, is introduced as a characteristic parameter of the cell delay. The circuit path delay is modeled by synthesizing a FOCV of the path using the FOCVs of the cells constituting the path. The simple structure of the vector synthesis model enables the reduction of simulation cost for a statistical analysis. The accuracy of the vector synthesis model has been verified experimentally. The deviation of the worst case delay from the result by SPICE Monte Carlo analysis is around 5%, whereas that of an usual corner (slow-slow and fast-fast) analysis is as high as 25%.
Tomohiro Fujita, Hidetoshi Onodera
ISCAS2
2000 Statistical modeling of device characteristics with systematic fluctuation
abstract
The fluctuations of device characteristics are usually regarded as a normal distribution. However, if we consider the fluctuation over the whole wafer, the fluctuation cannot be expressed as a normal distribution due to the existence of a systematic component. We propose a model, characterizing the systematic component according to the distance from the center die, which can express the fluctuation over the whole wafer statistically.
Kenichi Okada 0001, Hidetoshi Onodera
ISCAS2
2000 A performance optimization method by gate sizing using statistical static timing analysis
abstract
We propose a gate resizing method for delay and power optimization that is based on statistical static timing analysis.Our method focuses on the component of timing uncertainties due to local random fluctuation.Utilizing our method, over-design of a circuit can be eliminated and high-performance and high-reliability LSI design can be realized.The effectiveness of our method is examined by 6 benchmark circuits.We verify that our method can reduce delay and power dissipation from the circuits optimized without the consideration of fluctuation.otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific pennission and/or a fee.
Masanori Hashimoto, Hidetoshi Onodera
ISPD2
1999 A Practical Gate Resizing Technique Considering Glitch Reduction for Low Power Design
abstract
We propose a method for power optimization that considers glitch reduction by gate sizing based on the statistical estimation of glitch transitions. Our method reduces not only the amount of capacitive and short-circuit power consumption but also the power dissipated by glitches which has not been exploited previously. The effect of our method is verified experimentally using 8 benchmark circuits with a 0.6 m standard cell library. Our method reduces the power dissipation from the minimum-sized circuits further by 9.8 % on average and 23.0 % maximum. We also verify that our method is effective under manufacturing variation. 1
Masanori Hashimoto, Hidetoshi Onodera, Keikichi Tamaru
DAC2
1998 Proposal of a timing model for CMOS logic gates driving a CRC load
abstract
We present a gate delay model of CMOS logic gates driving a CRC n load for deep sub-micron technology.Our approach is to repIace series-parallel connected MOSFETS to an equivalent MOSFET and calculate the output waveform by an analytically derived formula.We present a MOSFET drain current model improved from the n-th power law hlOS~T model to represent the characteristic of the equivalent inverter accurately, The accuracy of our gate delay model is evaluated in several gates under various conditions of input transition time and CRC parameters.The maximum error is less than 10.3% in the experiments.Our approach will contribute to fast and accurate estimation of circuit speed under various supply voltage, which will enable us to optimize the circuit speed and power dissipation.
Akio Hirata, Hidetoshi Onodera, Keikichi Tamaru
ICCAD2
1998 A power optimization method considering glitch reduction by gate sizing
abstract
We propose a power optimization method considering glitch reduction by gate sizing. Our method reduces not only the amount of capacitive and short-circuit power consumption but also the power dissipated by glitches which has not been exploited previously. In the optimization method, we improve the accuracy of statistical glitch estimation method and device a gate sizing algorithm that utilizes perturbations for escaping a bad local solution. The effect of our method is verified experimentally using 12 benchmark circuits with a 0.5 µm standard cell library. Gate sizing reduces the number of glitch transitions by 38.2 % on average and by 63.4 % maximum. This results in the reduction of total transitions by 12.8 % on average. When the circuits are optimized for power without delay constraints, the power dissipation is reduced by 7.4 % on average and by 15.7 % maximum further from the minimum-sized circuits.
Masanori Hashimoto, Hidetoshi Onodera, Keikichi Tamaru
ISLPED2
1998 Model-adaptable MOSFET parameter-extraction method using an intermediate model
abstract
We present a parameter-extraction method that is applicable to many metal-oxide-semiconductor field-effect-transistor (MOSFET) models. A simple intermediate model is introduced to eliminate model dependency of parameter estimation for numerical optimization techniques. The process of the parameter estimation is decomposed into two parts: extraction of parameters of the intermediate model and transformation of the intermediate parameters into target model parameters. Only the latter part should be devised to accommodate new MOSFET models, which may be mostly identical with that for another model. We have integrated the method onto an extraction system, and verified that the method is effective for parameter extraction of major SPICE models.
Masaki Kondo, Hidetoshi Onodera, Keikichi Tamaru
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1997 A functional memory type parallel processor for vector quantization
abstract
We propose a memory-based parallel processor for vector quantization, called a functional memory type parallel processor for vector quantization (FMPP-VQ). It accelerates the nearest neighbour search of vector quantization. All distances between an input vector and reference vectors in a codebook are computed simultaneously in all PEs. The minimum value of all distances is searched in parallel. The nearest vector is obtained in O(k), where k stands for the dimension of vectors. An LSI including four PEs has been implemented. It operates at 25 MHz clock frequency.
Kazutoshi Kobayashi, Masayoshi Kinoshita, Masahiro Takeuchi, Hidetoshi Onodera, Keikichi Tamaru
ASP-DAC4
1997 A current mode cyclic A/D converter with a 0.8 μm CMOS process
abstract
We have developed a current mode cyclic analog-to-digital converter using a 0.8 /spl mu/m CMOS process. Our circuit structure makes it possible to construct the converter without any precise analog components, hence, it is well compatible with submicron processes. The fabricated circuit has an area of 0.014 mm/sup 2/ and performs 8-bit resolution at a sampling rate of 40 kHz and average power dissipation of 370 /spl mu/W at 4 V supply voltage.
Masaki Kondo, Hidetoshi Onodera, Keikichi Tamaru
ASP-DAC2
1996 Timing and Power Optimization by Gate Sizing Considering False Paths
abstract
This paper introduces a new gate sizing approach for area and power optimization considering path sensitization. The approach selects a set of long paths from a combinational circuit by means of a performance optimization oriented heuristic path selection approach. The longest sensitizable path delay of the circuit can be restricted within the specified delay limit if we set the specified delay limit on these paths in an LP based iterative gate sizing process. Since the approach get rid of unnecessary delay constraints on long false paths, results with smaller circuit area or power dissipation is expected. Experiments on benchmark circuits show that the proposed approach can substantially reduce the circuit area and power dissipation by considering path sensitization for some false path dominated circuits.
Guangqiu Chen, Hidetoshi Onodera, Keikichi Tamaru
Great Lakes Symposium on VLSI2
1995 A model-adaptable MOSFET parameter extraction system
abstract
No abstract available.
Masaki Kondo, Hidetoshi Onodera, Keikichi Tamaru
ASP-DAC2
1995 An iterative gate sizing approach with accurate delay evaluation
abstract
This paper introduces a new gate sizing approach with accurate delay evaluation. The approach solves gate sizing problems by iterating local sizing results from linear programming within mall variable ranges of gate sizes. In each iterative step, variable ranges of gate sizes are updated according to the result from a previous step. Solutions with accurate delay evaluation which consider input signal slopes and separately evaluate rising and falling delays are obtained after several iterative steps. A speedup technique is used to pick out gates actually involved in each local sizing step so as to reduce CPU time. Experiments on sample circuits show that our approach can provide solutions with smaller circuit area than conventional approaches for the same circuit delay or provide solutions under tight delay constraints where conventional approaches can nor reach. Moreover, our approach is faster than the conventional approaches for most circuits, especially under loose delay constraints.
Guangqiu Chen, Hidetoshi Onodera, Keikichi Tamaru
ICCAD2
1993 Layout-driven module selection for register-transfer synthesis of sub-micron ASIC's
abstract
As sub-micron design rules are utilized for IC fabrication, wiring is becoming an important issue in the register-transfer synthesis of high-speed application-specific integrated circuits. This paper proposes a new algorithm that incorporates performance-driven placement in module selection phase of the synthesis. The algorithm not only efficiently exploits multiple module implementations in the design library, but also finds the module placement which minimizes wiring delay. Experimental results on a practical size example show that considering both module and wiring issues, the algorithm is able to improve the design performance more than 20%.
Vasily G. Moshnyaga, Hiroshi Mori, Hidetoshi Onodera, Keikichi Tamaru
ICCAD3
1991 Branch-and-Bound Placement for Building Block Layout
abstract
We present a branch-and-bound placement technique for building block layout that effectively searches for an optimal placement in the whole solution space. We first describe a block placement problem and its solution space. Then we explain branching and bounding operations designed for the placement problem. Constraints on critical nets and/or the shape of a resulting chip can be taken into account in the search process. Experiments reveals that the number of blocks the method can manage is around six if the whole solution space is explored. For a problem which contains more blocks than the limit, we decompose the problem hierarchically and apply the method to each subproblem. The results for standard benchmark examples and a comparison with those of other systems are given to demonstrate the performance of the method.
Hidetoshi Onodera, Yo Taniguchi, Keikichi Tamaru
DAC1
1989 An efficient algorithm for layout compaction problem with symmetry constraints
abstract
An efficient algorithm is presented for the symbolic layout compaction problem with symmetry constraints. The symmetry constraint maintains the geometric symmetry of the circuit components during the layout compaction. It is indispensable to the symbolic layout for analog LSIs where the geometric symmetry between the components is important. However, it makes the compaction problem so complicated that no efficient algorithm has ever been shown except for the time-consuming linear programming algorithm. The proposed algorithm uses both the graph-based technique and the linear programming technique, and takes advantage of the high speed of the former and the generality of the latter. The authors implemented the proposed algorithm in a layout compaction program. The experimental results show that the proposed algorithm is fast enough for practical use.>
R. Okuda, Takashi Sato 0001, Hidetoshi Onodera, K. Tamariu
ICCAD3