EDBT 2026 Demo / reviewers in the wild / expert
Hideki Ando
dblp:37/3370
· DBLP profile ↗
17ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0001-5916-8703ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Wakeup Logic of the Issue Queue for High Scalability
Hideki Ando, Hajime Shimada |
ISCA | 1 |
| 2024 | Localizing the Tag Comparisons in the Wakeup Logic to Reduce Energy Consumption of the Issue QueueabstractThere is a high demand to reduce the energy consumption of microprocessors. Among resources in a processor, the issue queue is one of the largest energy consumers, with much of the energy being consumed by the wakeup logic. The wakeup logic comprises the content-addressable memory, where tag comparisons are performed for all entries. Such global tag comparisons consume an appreciable amount of energy. This paper proposes a scheme called segmenting wakeup logic (SegWU), where the wakeup logic is segmented logically. An instruction is dispatched to an available entry in the segment selected based on the partial bits of the source tag values. At the wakeup time, a broadcast destination tag is compared only in the segment that corresponds to the partial bits of the destination tag value. Tag comparisons are not performed in the other segments. This localization of tag comparisons reduces the energy consumption. Our evaluation results for SPEC2017 benchmark programs show that SegWU reduces the number of tag comparisons by 90.3%. This lowers the IQ energy consumption, and the net reduction excluding the secondary effect (activity lowering of the front-end) and overhead is 3.0% in the base core energy consumption with no performance degradation. The overall energy reduction including the secondary effect is 6.9%. Kenichiro Mori, Sota Kosugi, Hiroto Yoshida, Hajime Shimada, Hideki Ando |
MICRO | 5 |
| 2022 | Segmenting Age Matrices to Improve Instruction Scheduling without Increasing Delay and AreaabstractCurrent superscalar processors have a special circuit called the age matrix (AM) in the issue queue (IQ), which selects the oldest ready instruction in the queue, to allow better instruction scheduling. However, the optimization level is insufficient, because the AM selects only the single oldest instruction, and the other instructions to be issued are selected randomly. In this paper, we propose a new AM organization and a scheme that uses it successfully, which we call the segmented AM (SegAM). In SegAM, the AM is physically segmented, and therefore, each AM segment is quadratically smaller than the original AM. Consequently, multiple AMs, which allow the multiple oldest instructions to be selected, can be inserted into the IQ with their total area and delay remaining unchanged or reduced from that of the single monolithic AM. To ensure that an AM segment selects the oldest ready instruction in the entire IQ, the scheme dispatches (i.e., writes) instructions to the IQ segment-by-segment, which orders the segments by age. Our evaluation results using SPEC2017 benchmark programs demonstrates that an IQ with three AMs, in which each AM is segmented into four, achieves higher performance than the conventional single monolithic AM by an average of 6.4% and 1.2% (up to 17.7% and 11.4%) for integer and floating-point programs, respectively, with reductions of 21% IQ delay, 8% IQ area, and 84% age matrix energy. Hideki Ando |
ICCD | 1 |
| 2019 | SWQUE: A Mode Switching Issue Queue with Priority-Correcting Circular QueueabstractThe improvement of single-thread performance is much needed. Among the many structures that comprise a processor, the issue queue (IQ) is one of the most important structures that influences high single-thread performance. Correctly assigning the issue priority and providing high capacity efficiency are key features, but no conventional IQ organizations do not sufficiently have these. Hideki Ando |
MICRO | 1 |
| 2018 | Rearranging Random Issue Queue with High IPC and Short DelayabstractSingle-thread performance has remained mostly static for more than a decade. Among structures in a processor, the issue queue (IQ) is a structure that significantly affects the performance. To achieve high performance, high IPC and a short delay are required for the IQ, which have failed to be achieved in conventional IQs. We propose a novel IQ organization that we call the rearranging random issue queue (RRQ). The RRQ realizes an age-aware instruction selection in the IQ where instructions are ordered randomly. The RRQ divides the IQ into small (OQ: old queue) and large portions (MQ: main queue), where instructions in the OQ are prioritized using a simple select logic. To achieve age-aware selection, a small number of the oldest instructions in the MQ are moved to the OQ every cycle. Our implementation of the RRQ does not complicate the IQ circuit, and hardly increases the delay of the IQ. Evaluation results obtained in architectural simulation show that the RRQ achieves IPC as high as the shifting queue with compaction that realizes the perfect age-aware selection. Our evaluation results also show that the performance of the RRQ significantly outweighs that of processors with an IQ that has the age matrix, which suffers a long delay of the IQ. Shinji Sakai, Taishi Suenaga, Ryota Shioya, Hideki Ando |
ICCD | 4 |
| 2018 | Performance Improvement by Prioritizing the Issue of the Instructions in Unconfident Branch SlicesabstractSingle-thread performance has hardly improved for more than a decade. One of the largest problems for performance improvements is branch misprediction. There are two approaches to reduce the penalty caused by this. One is to reduce the frequency of misprediction, and the other is to reduce the cycles consumed because of misprediction. Improving branch predictors is the former approach, and many studies on this topic have been done for several decades. However, the latter approach has been rarely studied. The present paper hence explores the latter approach. The cycles consumed because of misprediction are divided into the following two parts. The first part is the state recovery penalty, which consists of cycles consumed for rolling back the processor state. The second part is the misspeculation penalty, which are cycles consumed during useless speculative execution from the fetch of a mispredicted branch until the completion of the branch execution. We focus on reducing the misspeculation penalty. For this, we propose a scheme called PUBS, which allows the instructions in unconfident branch slices to be issued with highest priority from the issue queue (IQ). Here, a branch slice is a set consisting of a branch and the instructions this branch directly or indirectly depends on, and we call the branch slice unconfident if the associated branch prediction cannot be sufficiently trusted. By issuing instructions in unconfident branch slices as early as possible, the wait cycles of these instructions in the IQ are minimized and thus the misspeculation penalty is minimized. Our evaluation results using SPEC2006 benchmark programs show that the PUBS scheme improves the performance of the programs with difficult branch prediction by 7.8% on average (a maximum of 19.2%) using only 4.0KB hardware cost. Hideki Ando |
MICRO | 1 |
| 2014 | Energy efficiency improvement of renamed trace cache through the reduction of dependent path lengthabstractA renaming logic is a high-cost module in a superscalar processor, and it consumes significant energy. For mitigating this, renamed trace cache (RTC), which caches renamed operands, was proposed. However, conventional RTCs have several problems such as low capacity-efficiency, large hardware overhead and insufficient caching of renamed operands. We propose a semi-global renamed trace cache (SGRTC) that caches only renamed operands whose distances from producers outside traces are short, and it solves the problems of conventional RTCs. Evaluation results show that SGRTC achieves 64% lower energy consumption for renaming with a 0.2% performance overhead compared to a conventional processor. Ryota Shioya, Hideki Ando |
ICCD | 2 |
| 2014 | A Front-End Execution Architecture for High Energy EfficiencyabstractSmart phones and tablets have recently become widespread and dominant in the computer market. Users require that these mobile devices provide a high-quality experience and an even higher performance. Hence, major developers adopt out-of-order superscalar processors as application processors. However, these processors consume much more energy than in-order superscalar processors, because a large amount of energy is consumed by the hardware for dynamic instruction scheduling. We propose a Front-end Execution Architecture (FXA). FXA has two execution units: an out-of-order execution unit (OXU) and an in-order execution unit (IXU). The OXU is the execution core of a common out-of-order superscalar processor. In contrast, the IXU comprises functional units and a bypass network only. The IXU is placed at the processor front end and executes instructions without scheduling. Fetched instructions are first fed to the IXU, and the instructions that are already ready or become ready to execute by the resolution of their dependencies through operand bypassing in the IXU are executed in-order. Not ready instructions go through the IXU as a NOP, thereby, its pipeline is not stalled, and instructions keep flowing. The not-ready instructions are then dispatched to the OXU, and are executed out-of-order. The IXU does not include dynamic scheduling logic, and its energy consumption is consequently small. Evaluation results show that FXA can execute over 50% of instructions using IXU, thereby making it possible to shrink the energy-consuming OXU without incurring performance degradation. As a result, FXA achieves both a high performance and low energy consumption. We evaluated FXA compared with conventional out-of-order/in-order superscalar processors after ARM big. LITTLE architecture. The results show that FXA achieves performance improvements of 67% at the maximum and 7.4% on geometric mean in SPECCPU INT 2006 benchmark suite relative to a conventional superscalar processor (big), while reducing the energy consumption by 86% at the issue queue and 17% in the whole processor. The performance/energy ratio (the inverse of the energy-delay product) of FXA is 25% higher than that of a conventional superscalar processor (big) and 27% higher than that of a conventional in-order superscalar processor (LITTLE). Ryota Shioya, Masahiro Goshima, Hideki Ando |
MICRO | 3 |
| 2013 | MLP-aware dynamic instruction window resizing for adaptively exploiting both ILP and MLPabstractIt is difficult to improve the single-thread performance of a processor in memory-intensive programs because processors have hit the memory wall, i.e., the large speed discrepancy between the processors and the main memory. Exploiting memory-level parallelism (MLP) is an effective way to overcome this problem. One scheme for exploiting MLP is aggressive out-of-order execution. To achieve this, large instruction window resources (i.e., the reorder buffer, the issue queue, and the load/store queue) are required; however, simply enlarging these resources degrades the clock cycle time. While pipelining these resources can solve this problem, this leads to instruction issue delays, which prevents instruction-level parallelism (ILP) from being exploited effectively. As a result, the performance of compute-intensive programs is degraded dramatically. Yuya Kora, Kyohei Yamaguchi, Hideki Ando |
MICRO | 3 |
| 2011 | Evaluation of issue queue delay: Banking tag RAM and identifying correct critical pathabstractThe tradeoff between complexity and attained instructions per cycle is often an important issue in microarchitectural designs. In this design phase, quick quantification of the complexity (i.e., delay) of relevant structures is required. The issue queue is one of such complex structures for which it is difficult to estimate delay. In this paper, we evaluate the issue queue delay to aid microarchitectural design. Our study includes two features: a circuit design and evaluation. First, we introduce banking the tag RAM, which is one of the components comprising the issue queue, to reduce the delay. Unlike normal RAM, banking the tag RAM is not straightforward, because of its uniqueness in the organization of the issue queue. Second, we explore and identify a correct critical path in the issue queue. A previous study summed the critical path of each component in the issue queue to obtain the delay of the issue queue, but this does not provide the correct delay of the issue queue, because the critical paths of each component are not connected logically. In the evaluation assuming 32nm LSI technology, we obtained the delays of an issue queue with eight to 128 entries. The process of banking the tag RAM and identifying the correct critical path reduces the delay by up to 20%, compared with not banking the tag RAM and simply summing the critical path delay of each component. Kyohei Yamaguchi, Yuya Kora, Hideki Ando |
ICCD | 3 |
| 2009 | Reducing register file size through instruction pre-execution enhanced by value predictionabstractTwo-step physical register deallocation (TSD) is an architectural scheme, which enhances memory-level parallelism (MLP) by pre-executing instructions. Ideally, the TSD allows MLP under the unlimited number of physical registers to be exploited, and consequently only a small register file is necessary for MLP. In practice, however, the amount of MLP exploitable is limited, because there are cases where pre-execution is not performed or timing of pre-execution is delayed. This is caused by data dependencies among the pre-executed instructions. This paper proposes the use of value prediction to solve these problems. Our way of the value prediction usage has the advantage over the conventional way of the usage for enhancing ILP, that there is no need to recover from misspeculation. Our evaluation results using SPECfp2000 benchmark show that our scheme can achieve equivalent performance to that of the previous TSD scheme without value prediction, with 75% of the register file size. Yusuke Tanaka 0001, Hideki Ando |
ICCD | 2 |
| 2003 | Pipeline stage unification: a low-energy consumption technique for future mobile processorsabstractRecent mobile processors are required to exhibit both low-energy consumption and high performance. To satisfy these requirements, dynamic voltage scaling (DVS) is currently employed. However, its effectiveness will be limited in the future because of shrinking the variable supply voltage range. As an alternative, we previously proposed pipeline stage unification (PSU), which unifies multiple pipeline stages without reducing the supply voltage at a power-saving mode. This paper compares effectiveness of PSU to DVS in current and future process generations. Our evaluation results show PSU will reduce energy consumption by 27-34% more than DVS after about 10 years. Hajime Shimada, Hideki Ando, Toshio Shimada |
ISLPED | 2 |
| 2002 | A preactivating mechanism for a VT-CMOS cache using address predictionabstractIt has become an important requirement to achieve high performance and low-power consumption at the same time. The dynamic leakage cut-off (DLC) scheme, which controls transistors' threshold voltage by the line on demand, is a technique that potentially satisfies that requirement for a cache. Yet, conventional DLC causes access time to significantly lengthen, and consequently processor performance is unacceptably degraded. This paper proposes a mechanism that suppresses the performance degradation by preactivating cache lines using address prediction before access requests. Our evaluation results show significant performance improvements are achieved with little increase of power consumption. Ryo Fujioka, Kiyokazu Katayama, Ryotaro Kobayashi, Hideki Ando, Toshio Shimada |
ISLPED | 4 |
| 1999 | Distributed memoryless point convergence algorithm for mobile robots with limited visibilityabstractWe present a distributed algorithm for converging autonomous mobile robots with limited visibility toward a single point. Each robot is an omnidirectional mobile processor that repeatedly: 1) observes the relative positions of those robots that are visible; 2) computes its new position based on the observation using the given algorithm; 3) moves to that position. The robots' visibility is limited so that two robots can see each other if and only if they are within distance V of each other and there are no other robots between them. Our algorithm is memoryless in the sense that the next position of a robot is determined entirely from the positions of the robots that it can see at that moment. The correctness of the algorithm is proved formally under an abstract model of the robot system in which: 1) each robot is represented by a point that does not obstruct the view of other robots; 2) the robots' motion is instantaneous; 3) there are no sensor and control error; 4) the issue of collision is ignored. The results of computer simulation under a more realistic model give convincing indication that the algorithm, if implemented on physical robots, will be robust against sensor and control error. Hideki Ando, Yoshinobu Oasa, Ichiro Suzuki, Masafumi Yamashita |
IEEE Trans. Robotics Autom. | 1 |
| 1996 | Performance Comparison of ILP Machines with Cycle Time EvaluationabstractMany studies have investigated performance improvement through exploiting instruction-level parallelism (ILP) with a particular architecture. Unfortunately, these studies indicate performance improvement using the number of cycles that are required to execute a program, but do not quantitatively estimate the penalty imposed on the cycle time from the architecture. Since the performance of a microprocessor must be measured by its execution time, a cycle time evaluation is required as well as a cycle count speedup evaluation. Currently, superscalar machines are widely accepted as the machines which achieve the highest performance. On the other hand, because of hardware simplicity and instruction scheduling sophistication, there is a perception that the next generation of microprocessors will be implemented with a VLIW architecture. A simple VLIW machine, however, has a serious weakness regarding speculative execution. Thus, it is a question whether a simple VLIW machine really outperforms a superscalar machine. We recently proposed a mechanism called predicating that supports speculative execution for the VLIW machine, and showed a significant cycle count speedup over a scalar machine. Although the mechanism is simple, it is unknown how much it imposes a penalty on the cycle time, and how much the performance is improved as a result. This paper evaluates both the cycle count speedup and the cycle time for three ILP machines: a superscalar machine, a simple VLIW machine, and the VLIW machine with predicating. The evaluation results show that the simple VLIW machine slightly outperforms the superscalar machine, while the VLIW machine with predicating achieves a significant speedup of 1.41x over the superscalar machine. Tetsuya Hara, Hideki Ando, Chikako Nakanishi, Masao Nakaya |
ISCA | 2 |
| 1995 | Unconstrained Speculative Execution with Predicated State BufferingabstractSpeculative execution is execution of instructions before it is known whether these instructions should be executed. Compiler-based speculative execution has the potential to achieve both a high instruction per cycle rate and high clock rate. Pure compiler-based approaches, however, have greatly limited instruction scheduling due to a limited ability to handle side effects of speculative execution. Significant performance improvement is, thus, difficult in non-numerical applications. This paper proposes a new architectural mechanism, called predicating, which provides unconstrained speculative execution. Predicating removes restrictions which limit the compiler's ability to schedule instructions. Through our hardware support, the compiler is allowed to move instructions past multiple basic block boundaries from any succeeding control path. Predicating buffers the side effects of speculative execution with its predicate, and the buffered predicate efficiently commits or squashes the side effects. The mechanism also provides a speculative exception handling scheme. The scheme, called the future condition, properly postpones speculative exceptions and efficiently restarts the process. We show that our mechanism can be implemented through a modest amount of hardware with little complexity. The evaluation results show that our mechanism significantly improves performance, and achieves a 2.45x speedup over scalar machines. Hideki Ando, Chikako Nakanishi, Tetsuya Hara, Masao Nakaya |
ISCA | 1 |
| 1993 | Speculative Execution and Reducing Branch Penalty in a Parallel Issue MachineabstractParallel instruction issue is essential for performance improvement in current microprocessor designs. Just extra function units are, however, little beneficial in non-numerical applications since control dependence severely limits exploitation of instruction-level parallelism (ILP) and frequent branches consume ILP due to its long latency. Boosting is an interesting technique to reduce control dependence. It allows general speculative execution with little cycle time penalty. From the cost/performance point of view, we propose the efficient implementation of boosting, which requires the small support hardware and maximizes performance gain from boosting in the limited hardware. We also propose a new branch scheme to reduce the branch penalty which has a particularly big performance impact in a parallel issue machine. Our scheme fetches from both directions of the branch with small hardware cost through integration of a code movement and hardware support. We evaluate our schemes and find that they significantly contribute to performance improvement.> Hideki Ando, Chikako Nakanishi, Hirohisa Machida, Tetsuya Hara, Satoru Kishida, Masao Nakaya |
ICCD | 1 |