Tohru Ishihara

dblp:29/4628 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
7since 2021 · last 2023
0000-0002-1650-9958ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 9 first-author · 6 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author
YearPublicationVenuePosition
2023 An Efficient Fault Injection Algorithm for Identifying Unimportant FFs in Approximate Computing Circuits
abstract
Approximate Computing (AC) saves energy and improves performance by introducing approximation into computation in error-torrent applications. This work focuses on an AC strategy that accurately performs important computations and approximates others. In order to determine which calculations are unimportant, we propose a novel importance evaluation algorithm, in which the key idea is a two-step fault injection to extract the near-optimal set of unimportant flip-flops in the circuit. The proposed algorithm reduces the complexity of architecture exploration from an exponential order to a linear order with-out understanding the functionality and behavior of the target application program.
Jiaxuan Lu, Yutaka Masuda, Tohru Ishihara
DATE3
2023 Feedback-Tuned Fuzzing for Accelerating Quality Verification of Approximate Computing Design
abstract
This paper proposes a novel quality verification (q-verification) technique for approximate computing (AC) design, which explores the input patterns that violate the constraint of computational quality, e.g., image quality. The key component of the proposed technique is coverage-based grey-box fuzzing (CGF). For enhancing the efficacy of CGF, this work proposes a feedback tuning algorithm that controls the feedback information to the queue by referring to the computational quality. In a case study of approximate image processing design where the bit-width scaling is applied, we experimentally confirmed that the proposed method reduced the runtime required for q-verification by 1.92 times compared to the conventional q-verification method using CGF.
Yusei Honda, Yutaka Masuda, Tohru Ishihara
IOLTS3
2022 An Accuracy Reconfigurable Vector Accelerator Based on Approximate Logarithmic Multipliers
abstract
The logarithmic approximate multiplier proposed by Mitchell provides an efficient alternative to accurate multipliers in terms of area and power consumption. However, its maximum error of 11.1% makes it difficult to deploy in applications requiring high accuracy. To widely reduce the error of the Mitchell multiplier, this paper proposes a novel operand decomposition method which decomposes one operand into multiple operands and calculates them using multiple Mitchell multipliers. Based on this operand decomposition, this paper also proposes an accuracy reconfigurable vector accelerator which can provide a required computational accuracy with a high parallelism. The proposed vector accelerator dramatically reduces the area by more than half from the accurate multiplier array while satisfying the required accuracy for various applications. The experimental results show that our proposed vector accelerator behaves well in image processing and robot localization.
Lingxiao Hou, Yutaka Masuda, Tohru Ishihara
ASP-DAC3
2022 Power-aware pruning for ultrafast, energy-efficient, and accurate optical neural network design
abstract
With the rapid progress of the integrated nanophotonics technology, the optical neural network (ONN) architecture has been widely investigated. Although the ONN inference is fast, conventional densely connected network structures consume large amounts of power in laser sources. We propose a novel ONN design method that finds an ultrafast, energy-efficient, and accurate ONN structure. The key idea is power-aware edge pruning that derives the near-optimal numbers of edges in the entire network. Optoelectronic circuit simulation demonstrates the correct functional behavior of the ONN. Furthermore, experimental evaluations using tensor-flow show the proposed methods achieved 98.28% power reduction without significant loss of accuracy.
Naoki Hattori, Yutaka Masuda, Tohru Ishihara, Akihiko Shinya, Masaya Notomi
DAC3
2022 DVFS Virtualization for Energy Minimization of Mixed-Criticality Dual-OS Platforms
abstract
A dual-OS platform can efficiently implement emerging mixed-criticality systems by consolidating a real-time OS (RTOS) and a general-purpose OS (GPOS). Although the dual-OS platform attracts increasing attention, it often suffers from energy inefficiency in the GPOS for guaranteeing real-time responses of the RTOS. This paper proposes an energy minimization method called DVFS virtualization, which allows running multiple DVFS policies dedicated to the RTOS and GPOS, respectively. The experimental evaluation using a commercial processor showed that the proposed hardware could change the supply voltage within 500 ns and reduce the energy consumption of typical applications by 60 % in the best case compared to conventional dual-OS platforms.
Takumi Komori, Yutaka Masuda, Tohru Ishihara
RTCSA3
2021 Critical Path Isolation and Bit-Width Scaling Are Highly Compatible for Voltage Over-Scalable Design
abstract
This work proposes a design methodology that saves the power under voltage over-scaling (VOS) operation. The key idea of the proposed design methodology is to combine critical path isolation (CPI) and bit-width scaling (BWS) under the constraint of computational quality, e.g., Peak Signal-to-Noise Ratio (PSNR). Conventional CPI inherently cannot reduce the delay of intrinsic critical paths (CPs), which may significantly restrict the power saving effect. On the other hand, the proposed methodology tries to reduce both intrinsic and non-intrinsic CPs. Therefore, our design dramatically reduces the supply voltage and power dissipation while satisfying the quality constraint. Moreover, for reducing co-design exploration space, the proposed methodology utilizes the exclusiveness of the paths targeted by CPI and BWS, where CPI aims at reducing the minimum supply voltage of non-intrinsic CP, and BWS focuses on intrinsic CPs in arithmetic units. From this key exclusiveness, the proposed design splits the simultaneous optimization problem into three sub-problems; (1) the determination of bit-width reduction, (2) the timing optimization for non-intrinsic CPs, and (3) investigating the minimum supply voltage of the BWS and CPI-applied circuit under quality constraint, for reducing power dissipation. Thanks to the problem splitting, the proposed methodology can efficiently find quality-constrained minimum-power design. Evaluation results show that CPI and BWS are highly compatible, and they significantly enhance the efficacy of VOS. In a case study of GPGPU processor, the proposed design saves the power dissipation by 42.7% for an image processing and by 51.2% for a neural network inference workload.
Yutaka Masuda, Jun Nagayama, TaiYu Cheng, Tohru Ishihara, Yoichi Momiyama, Masanori Hashimoto
DATE4
2021 Dynamic Verification of Approximate Computing Circuits using Coverage-based Grey-box Fuzzing
abstract
Approximate computing (AC) has recently emerged as a promising approach to the energy-efficient design of digital systems. For realizing the practical AC design, we need to verify whether the designed circuit can operate correctly under various operating conditions. Namely, the verification needs to efficiently find fatal logic errors or timing errors that violate the constraint of computational quality. This paper proposes a novel dynamic verification methodology of the AC circuit. The key idea of the proposed methodology is to incorporate a quality assessment capability into the Coverage-based Grey-box Fuzzing (CGF). CGF is one of the most promising techniques in the research field of software security testing. By repeating (1) mutation of test patterns, (2) execution of the program under test (PUT), and (3) aggregation of coverage information and feedback to the next test pattern generation, CGF can explore the verification space quickly and automatically. On the other hand, CGF originally cannot consider the computational quality by itself. For overcoming this quality unawareness in CGF, the proposed methodology additionally embeds the Design Under Test (DUT) mechanisms into the calculation part of computational quality. Thanks to the integration of CGF and DUT mechanism, the proposed framework realizes the quality-aware feedback loop in CGF and thus quickly enhances the verification coverage for test patterns that violate the quality constraint. In this work, we quantitatively compared the verification coverage of the approximate arithmetic circuits between the proposed methodology and the random test. In a case study of an approximate multiply-accumulate (MAC) unit, we experimentally confirmed that the proposed methodology achieves the target coverage three times faster than the random test.
Kazuki Yoshisue, Yutaka Masuda, Tohru Ishihara
IOLTS3
2019 BDD-based synthesis of optical logic circuits exploiting wavelength division multiplexing
abstract
Optical circuits using nanophotonic devices attract significant interest due to its ultra-high speed operation. As a consequence, the synthesis methods for the optical circuits also attract increasing attention. However, existing methods for synthesizing optical circuits mostly rely on straight-forward mappings from established data structures such as Binary Decision Diagram (BDD). The strategy of simply mapping a BDD to an optical circuit sometimes results in an explosion of size and involves significant power losses in branches and optical devices. To address these issues, this paper proposes a method for reducing the size of BDD-based optical logic circuits exploiting wavelength division multiplexing (WDM). The paper also proposes a method for reducing the number of branches in a BDD-based circuit, which reduces the power dissipation in laser sources. Experimental results obtained using a partial product accumulation circuit in parallel multipliers demonstrates significant advantages of our method over existing approaches in terms of area and power consumption.
Ryosuke Matsuo, Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera, Akihiko Shinya, Masaya Notomi
ASP-DAC3
2019 Area-efficient fully digital memory using minimum height standard cells for near-threshold voltage computing
Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera
Integr.2
2018 Independent N-Well And P-Well Biasing For Minimum Leakage Energy Operation
abstract
This paper proposes a method for minimizing leakage energy consumption under a specific supply voltage and a delay constraint by independently tuning threshold voltages of nMOSFETs and pMOSFETs with body-biasing. We first show a necessary and sufficient condition for the minimum leakage energy operation of a circuit under a delay constraint. We next show that the condition can be identified by a ratio of the leakage currents drawn through an nMOSFET and a pMOSFET in a leakage monitor circuit integrated with the targeting circuit. The leakage current ratio can be monitored at runtime using the leakage monitor. Assuming a constant supply voltage, it is thus possible to minimize the total energy consumption of the circuit by independent tuning of n-well and p-well bias voltages so that the leakage current ratio tracks the predetermined value while keeping the delay constraint. The proposed strategy is experimentally verified by measurements using a 32-bit RISC processor integrating the leakage monitor on the same die fabricated with a 65 nm CMOS process.
Yosuke Okamura, Tohru Ishihara, Hidetoshi Onodera
IOLTS2
2018 An Integrated Nanophotonic Parallel Adder
abstract
Integrated optical circuits with nanophotonic devices have attracted significant attention due to their low power dissipation and light-speed operation. With light interference and resonance phenomena, the nanophotonic device works as a voltage-controlled optical pass-gate like a pass-transistor. This article first introduces the concept of optical pass-gate logic and then proposes a parallel adder circuit based on optical pass-gate logic. Experimental results obtained with an optoelectronic circuit simulator show the advantages of our optical parallel adder circuit over a traditional CMOS-based parallel adder circuit.
Tohru Ishihara, Akihiko Shinya, Koji Inoue, Kengo Nozaki, Masaya Notomi
ACM J. Emerg. Technol. Comput. Syst.1
2016 A closed-form stability model for cross-coupled inverters operating in sub-threshold voltage region
abstract
A cross-coupled inverter which is an essential element of on-chip memory subsystems plays an important role in synchronous LSI circuits. In this paper, an analytical stability model for a cross-coupled inverter operating in a sub-threshold voltage region is proposed. The proposed model analytically shows that the minimum operating voltage of the cross-coupled inverter distributes normally in a high-s region if the distribution of the threshold voltage is Gaussian. The minimum supply voltage at which the yield of the cross-coupled inverter becomes a specific value can be accurately derived by a simple calculation using the model. Monte-Carlo simulation assuming a commercial 28 nm process technology demonstrates the accuracy and the validity of the proposed model. Based on the model, this paper shows strategies for variation tolerant memory design.
Tatsuya Kamakari, Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera
ASP-DAC3
2015 Microarchitectural-level statistical timing models for near-threshold circuit design
abstract
Near-threshold computing has emerged as a promising solution for drastically improving the energy efficiency of microprocessors. This paper proposes architectural-level statistical static timing analysis (SSTA) models for the near-threshold voltage computing where the path delay distribution is approximated as a lognormal distribution. First, we prove several important theorems that help consider architectural design strategies for high performance and energy efficient near-threshold computing. After that, we show the numerical experiments with Monte Carlo simulations using a commercial 28-nm process technology model and demonstrate that the properties presented in the theorems hold for the practical near-threshold logic circuits.
Jun Shiomi, Tohru Ishihara, Hidetoshi Onodera
ASP-DAC2
2013 DLIC: Decoded loop instructions caching for energy-aware embedded processors
Ji Gu, Hui Guo 0001, Tohru Ishihara
ACM Trans. Embed. Comput. Syst.3
2012 A flexible structure of standard cell and its optimization method for near-threshold voltage operation
abstract
With ever growing demands of mobile devices, low power consumption has become essential for VLSI circuits. Since standard cell libraries are typically used in many parts of VLSI circuits, their performance has a strong impact on realizing high speed and low power VLSI circuits. One of the most promising approaches for reducing the power consumption of the circuit is lowering the supply voltage. However this causes an increase of imbalance between rise and fall delays especially for cells having transistor stacks. For mitigating this imbalance, this paper proposes a structure of standard cells where the P/N ratio of each cell can be independently customized for near-threshold operation in VLSI circuits. The structure cancels the imbalance between rise and fall delays at the expense of cell area. The experiments with ISCAS'85 benchmark circuits demonstrate that the standard cell library consisting of the proposed cells reduces the power consumption of the benchmark circuits by 16% on average without increasing the circuit area, compared to that of the same circuit synthesized with a library which is not optimized for the near-threshold operation.
Shinichi Nishizawa, Tohru Ishihara, Hidetoshi Onodera
ICCD2
2011 Developing an integrated verification and debug methodology
abstract
As design complexity of LSI systems increase, so does the verification challenges. It is very important, yet difficult to find all design errors and correct them in a timely manner. This paper presents our experience with a new verification and debug methodology based on the combination of formal verification and automated debugging. This methodology, which is applied to the development of a DDR2 memory design targeted for an FGPA, is found to significantly reduce the verification and debug tasks typically performed.
Akitoshi Matsuda, Tohru Ishihara
DATE2
2011 An integrated optimization framework for reducing the energy consumption of embedded real-time applications
Hideki Takase, Lovic Gauthier, Hirotaka Kawashima, Noritoshi Atsumi, Tomohiro Tatematsu, Yoshitake Kobayashi, Shunitsu Kohara, Takenori Koshiro, Tohru Ishihara, Hiroyuki Tomiyama, Hiroaki Takada
ISLPED10
2010 Minimizing inter-task interferences in scratch-pad memory usage for reducing the energy consumption of multi-task systems
abstract
This paper presents a new technique for reducing the energy consumption of a multi-task system by sharing its scratchpad memory (SPM) space among the tasks. With this technique, tasks can interfere by using common areas of the SPM. However, this requires to update these areas during context switches, which involves considerable overheads. Hence, an integer linear programming formulation is used at compile time for finding the best assignment of memory objects to the SPM and their respective locations inside it. Experiments show that the technique achieves up to 85% energy reduction with 8Kb of SPM and surpasses other sharing approaches.
Lovic Gauthier, Tohru Ishihara, Hideki Takase, Hiroyuki Tomiyama, Hiroaki Takada
CASES2
2010 SRAM Leakage Reduction by Row/Column Redundancy Under Random Within-Die Delay Variation
abstract
Share of leakage in total power consumption of static RAM (SRAM) memories is increasing with technology scaling. Reverse body biasing increases threshold voltage (Vth), which exponentially reduces subthreshold leakage, but it increases SRAM access delay. Traditionally, when all cells of an SRAM block used to have almost the same delay, within-die variations are increasingly widening the delay distribution of cells even within a single SRAM block, and hence, most of these cells are substantially faster than the delay set for the entire block. Consequently, after the reverse body biasing and the resulting delay rise, only a small number of cells violate the original delay of the SRAM block; we propose to replace them with sufficient number of spare rows/columns of SRAM. Our experiments show that the leakage can be reduced by up to 40% in a 90-nm predictive technology by adding less than ten spare columns to an 8-kB SRAM array for a negligible penalty in delay, dynamic power, and area in the presence of 3% uncorrelated random delay variation.
Maziar Goudarzi, Tohru Ishihara
IEEE Trans. Very Large Scale Integr. Syst.2
2008 A Generalized Framework for System-Wide Energy Savings in Hard Real-Time Embedded Systems
abstract
A generalized dynamic energy performance scaling (DEPS) framework is proposed for exploring application-specific energy-saving potential in hard real-time embedded systems. This software-centric framework focuses on system-wide energy reduction and takes advantage of possible power control mechanisms to trade off performance for energy savings. Three existing technologies, i.e., dynamic hardware resource configuration (DHRC), dynamic voltage frequency scaling (DVFS), and dynamic power management (DPM) have been employed in this framework to achieve the maximal energy savings. Static and dynamic schemes of DEPS are proposed to deal with stable or variable workload in the embedded systems. Through a case study, its effectiveness has been validated.
Hiroyuki Tomiyama, Hiroaki Takada, Tohru Ishihara
EUC (1)4
2008 Instruction cache leakage reduction by changing register operands and using asymmetric sram cells
abstract
Share of leakage in cache memories is increasing with technology scaling. Studies show that most stored bits in instruction caches are zero, and hence, asymmetric SRAM cells which dissipate less leakage when storing 0, effectively reduce leakage with negligible performance penalty. We show that by carefully choosing register operands of instructions, it is possible to further increase the number of 0 bits, and hence, increase leakage savings in instruction cache. This compiler technique is performed off-line and introduces absolutely no delay penalty since processor registers are all the same. Experimental results of our benchmarks show up to 33% (averaging 30.35%) improvement in leakage.
Maziar Goudarzi, Tohru Ishihara
ACM Great Lakes Symposium on VLSI2
2008 Simultaneous optimization of memory configuration and code allocation for low power embedded systems
abstract
This paper proposes a hybrid memory architecture which consists of the following two regions; 1) a dynamic power conscious region which uses low Vdd and Vth and 2) a static power conscious region which uses high Vdd and Vth. This paper also proposes an optimization problem for finding the optimal memory division ratio, the code allocation, ²ratio and Vdd so as to minimize the total power consumption of the memory under constraints of static noise margin (SNM), memory access delay and area overhead. Experimental results demonstrate that the total power consumption can be reduced by 50.8% with 7.7% memory array area overhead without degradations of SNM and access delay.
Tadayuki Matsumura, Tohru Ishihara, Hiroto Yasuura
ACM Great Lakes Symposium on VLSI2
2008 Variation-Aware Software Techniques for Cache Leakage Reduction Using Value-Dependence of SRAM Leakage Due to Within-Die Process Variation
Maziar Goudarzi, Tohru Ishihara, Hamid Noori
HiPEAC2
2008 Row/column redundancy to reduce SRAM leakage in presence of random within-die delay variation
abstract
Traditionally, spare rows/columns have been used in two ways: either to replace too leaky cells to reduce leakage, or to substitute faulty cells to improve yield. In contrast, we first choose a higher threshold voltage (Vth) and/or gate-oxide thickness (Tox) for SRAM transistors at design time to reduce leakage, and then substitute the resulting too slow cells by spare rows/columns. We show that due to within-die delay variation of SRAM cells only a few cells violate target timing at higher Vth or Tox; we carefully choose the Vth and Tox values such that the original memory timing-yield remains intact for a negligible extra delay. On a commercial 90nm process assuming 3% variation in SRAM cell delay, we obtained 47% leakage reduction by adding only 5 redundant columns at negligible area, dynamic power and delay costs.
Maziar Goudarzi, Tohru Ishihara
ISLPED2
2007 A Software Technique to Improve Yield of Processor Chips in Presence of Ultra-Leaky SRAM Cells Caused by Process Variation
abstract
Exceptionally leaky transistors are increasingly more frequent in nano-scale technologies due to lower threshold voltage and its increased variation. Such leaky transistors may even change position with changes in the operating voltage and temperature, and hence, redundancy at circuit-level is not sufficient to tolerate such threats to yield. We show that in SRAM cells this leakage depends on the cell value and propose a first software-based runtime technique that suppresses such abnormal leakages by storing safe values in the corresponding cache lines before going to standby mode. Analysis shows the performance penalty is, in the worst case, linearly dependent to the number of so-cured cache lines while the energy saving linearly increases by the time spent in standby mode. Analysis and experimental results on commercial processors confirm that the technique is viable if the standby duration is more than a small fraction of a second.
Maziar Goudarzi, Tohru Ishihara, Hiroto Yasuura
ASP-DAC2
2007 Task scheduling for reliable cache architectures of multiprocessor systems
Makoto Sugihara, Tohru Ishihara, Kazuaki J. Murakami
DATE2
2005 A Way Memoization Technique for Reducing Power Consumption of Caches in Application Specific Integrated Processors
abstract
This paper presents a technique for eliminating redundant cache-tag and cache-way accesses to reduce power consumption. The basic idea is to keep a small number of most recently used (MRU) addresses in a memory address buffer (MAB) and to omit redundant tag and way accesses when there is a MAB-hit. Since the approach keeps only tag and set-index values in the MAB, the energy and area overheads are relatively small even for a MAB with a large number of entries. Furthermore, the approach does not sacrifice the performance. In other words, neither the cycle time nor the number of executed cycles increases. The proposed technique has been applied to the Fujitsu VLIW processor (FR-V) and its power saving has been estimated using NanoSim. Experiments for 32 kB 2-way set associative caches show the power consumption of I-cache and D-cache can be reduced by 40% and 50%, respectively.
Tohru Ishihara, Farzan Fallah
DATE1
2005 A cache-defect-aware code placement algorithm for improving the performance of processors
abstract
Yield improvement through exploiting fault-free sections of defective chips is a well-known technique (Koren and Singh (1990) and Stapper et al. (1980)). The idea is to partition the circuitry of a chip in a way that fault-free sections can function independently. Many fault tolerant techniques for improving the yield of processors with a cache memory have been proposed. In this paper, we propose a defect-aware code placement technique which offsets the performance degradation of a processor with a defective cache memory. To the best of our knowledge, this is the first compiler-based technique which offsets the performance degradation due to cache defects. Experiments demonstrate that the technique can compensate the performance degradation even when 5% of cache lines are faulty. In some cases the technique was able to offset the impact even in presence of 25% faulty cache-lines.
Tohru Ishihara, Farzan Fallah
ICCAD1
2005 A non-uniform cache architecture for low power system design
abstract
This paper proposes a non-uniform cache architecture for reducing the power consumption of memory systems. The non-uniform cache allows having different associativity values (i.e., the number of cache-ways) for different cache-sets. An algorithm determines the optimum number of cache-ways for each cache-set and generates object code suitable for the non-uniform cache memory. The paper also proposes a compiler technique for reducing redundant cache-way accesses and cache-tag accesses. Experiments demonstrate that our technique can reduce the power consumption of memory systems by up to 76% compared to the best result achieved by the conventional method
Tohru Ishihara, Farzan Fallah
ISLPED1
2001 A system level memory power optimization technique using multiple supply and threshold voltages
abstract
A system level approach for a memory power reduction is proposed in this paper. The basic idea is allocating frequently executed object codes into a small subprogram memory and optimizing supply voltage and threshold voltage of the subprogram memory. Since large scale memory contains a lot of direct paths from power supply to round, power dissipation caused by subthreshold leakage current is more serious than dynamic power dissipation. Our approach optimizes the size of subprogram memory, supply voltage, and threshold voltage so as to minimize memory power dissipation including static power dissipation caused by leakage current. A heuristic algorithm which determines code allocation, supply voltage, and threshold voltage simultaneously so as to minimize power dissipation of memories is proposed as well. Our experiments with some benchmark programs demonstrate significant energy reductions up to 80% over a program memory which does not employ our approach.
Tohru Ishihara, Kunihiro Asada
ASP-DAC1
2000 A Power Reduction Technique with Object Code Merging for Application Specific Embedded Processors
abstract
In this paper, a power reduction technique which merges frequently executed sequences of object codes into a set of single instructions is proposed. The merged sequence of object codes is restored by an instruction decompressor before decoding the object codes. The decompressor is implemented by a ROM. In many programs, only a few sequences of object codes are frequently executed. Therefore, merging these frequently executed sequences into a single instruction leads to a significant energy reduction. Our experiments with actual read only memory (ROM) modules and some benchmark program demonstrate significant energy reductions up to more than 65% at best case over an instruction memory without the object code merging.
Tohru Ishihara, Hiroto Yasuura
DATE1
1999 Way-predicting set-associative cache for high performance and low energy consumption
abstract
Set-Associative CacheThis paper proposes a new approach using way prediction for achieving high performance and low energy consumption of set-associative caches.By accessing only a single cache way predicted, instead of accessing all the ways in a set, the energy consumption can be reduced.This paper shows that the way-predicting set-associative cache improves the ED (energy-delay) product by SO-70% compared to a conventional set-associative cache. Conventional CacheTotal energy consumed for an access to a set-associative cache ( Ecache) can be approximated by the sum of following terms [9]:
Koji Inoue, Tohru Ishihara, Kazuaki J. Murakami
ISLPED2
1998 Power-Pro: Programmable Power Management Architecture
abstract
This paper presents Power-Pro architecture (Programmable Power Management Architecture), a novel processor architecture for power reduction. Power-Pro architecture has following two functionalities: (i) Supply voltage and clock frequency can be dynamically varied. (ii) Active data-path width can be dynamically adjusted to requirement of application programs. For the application programs which require less performance or less data-path width, Power-Pro architecture realize dramatic power reduction.
Tohru Ishihara, Hiroto Yasuura
ASP-DAC1
1998 Instruction Scheduling for Power Reduction in Processor-Based System Design
abstract
This paper proposes an instruction scheduling technique to reduce power consumed for off-chip driving. The technique minimizes the switching activity of a data bus between an on-chip cache and a main memory when instruction cache misses occur. The scheduling problem is formulated and a scheduling algorithm is also presented. Experimental results demonstrate the effectiveness and the efficiency of the proposed algorithm.
Hiroyuki Tomiyama, Tohru Ishihara, Akihiko Inoue, Hiroto Yasuura
DATE2
1998 Voltage scheduling problem for dynamically variable voltage processors
abstract
This paper presents a model of dynamically variable voltage processor and basic theorems for power-delay optimization. A static voltage scheduling problem is also proposed and formulated as an integer linear programming (ILP) problem. In the problem, we assume that a core processor can vary its supply voltage dynamically, but can use only a single voltage level at a time. For a given application program and a dynamically variable voltage processor, a voltage scheduling which minimizes energy consumption under an execution time constraint can be found.
Tohru Ishihara, Hiroto Yasuura
ISLPED1
1996 Basic experimentation on accuracy of power estimation for CMOS VLSI circuits
abstract
In this paper, we discuss on accuracy of several kinds of power dissipation model for CMOS VLSI circuits. Some researchers have proposed several efficient power estimation methods for CMOS circuits. However, we do not know how accurate they are because we have not established a method to compare the estimated results of power consumption with that of actual VLSI chip. To evaluate the accuracy of several kind of power dissipation model such as chip-level, block-level and gate-level etc., we examined as follows: (i) Measuring power consumption of actual micro-processors. (ii) Estimating power consumption with several kinds of power dissipation model. (iii) Comparing (i) with (ii). The experimental results show as follows: (1) Power estimation at gate level is accurate enough. (2) Estimating power of a clock tree independently makes estimation more accurate.
Tohru Ishihara, Hiroto Yasuura
ISLPED1