VLDB 2026 Research / reviewers in the wild / expert
Hiroshi Nakamura
dblp:94/5594
· DBLP profile ↗
79ranked-venue papers
8as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 60 · 5 first-authorSoftware engineering, systems software and programming languages · 11 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 since 2021Security and privacy · 5 · 3 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables secure computation over encrypted data, but its computational cost remains a major obstacle to practical deployment. To mitigate this overhead, many studies have explored GPU acceleration for the CKKS scheme, which is widely used for approximate arithmetic. In CKKS, CKKS parameters are configured for each workload by balancing multiplicative depth, security requirements, and performance. These parameters significantly affect ciphertext size, thereby determining how the memory footprint fits within the GPU memory hierarchy. Nevertheless, prior studies typically apply their proposed optimization methods uniformly, without considering differences in CKKS parameter configurations. In this work, we demonstrate that the optimal GPU optimization strategy for CKKS depends on the CKKS parameter configuration. We first classify prior optimizations by two aspects of dataflows which affect memory footprint and then conduct both qualitative and quantitative performance analyses. Our analysis shows that even on the same GPU architecture, the optimal strategy varies with CKKS parameters with performance differences of up to 1.98 $\times$ between strategies, and that the criteria for selecting an appropriate strategy differ across GPU architectures. Ai Nozaki, Takuya Kojima, Hiroshi Nakamura, Hideki Takase |
COMPSAC | 3 |
| 2025 | Exploring the Possibility of TypiClust for Low-Budget Federated Active LearningabstractFederated Active Learning (FAL) seeks to reduce the burden of annotation under the realistic constraints of federated learning by leveraging Active Learning (AL). As FAL settings make it more expensive to obtain ground truth labels, FAL strategies that work well in low-budget regimes, where the amount of annotation is very limited, are needed. In this work, we investigate the effectiveness of TypiClust, a successful low-budget AL strategy, in low-budget FAL settings. Our empirical results show that TypiClust works well even in low-budget FAL settings contrasted with relatively low performances of other methods, although these settings present additional challenges, such as data heterogeneity, compared to AL. In addition, we show that FAL settings cause distribution shifts in terms of typicality, but TypiClust is not very vulnerable to the shifts. We also analyze the sensitivity of TypiClust to feature extraction methods, and it suggests a way to perform FAL even in limited data situations. Yuta Ono, Hiroshi Nakamura, Hideki Takase |
COMPSAC | 2 |
| 2023 | An asynchronous federated learning focusing on updated models for decentralized systems with a practical frameworkabstractFederated learning (FL), a machine learning technique that preserves privacy by aggregating models from each device without exchanging personal data, has gained significant interest. This paper aims to establish an efficient and practical asynchronous FL method for decentralized systems. In asynchronous decentralized FL, the order of learning and aggregation is arbitrary. First, we explore the impact of this ordering on FL. Then, we propose to utilize the version information regarding model updates. Our strategy is to aggregate only the updated models after the previous round to improve the quality of the device’s model by avoiding the reaggregation of older models. Moreover, we also design a new practical framework for asynchronous decentralized FL by extending Flower framework. Our framework realizes effective communication ability by leveraging gRPC communication and thus can be applied to practical systems without the central server. Our evaluation shows the effectiveness of our methods that aggregate only the updated models from other devices. In addition, we show the impact of ordering on learning and aggregation according to situations. Yusuke Kanamori, Yusuke Yamasaki, Shintaro Hosoai, Hiroshi Nakamura, Hideki Takase |
COMPSAC | 4 |
| 2023 | ILP Based Mapping for Elastic CGRAsabstractIn recent years, the emergence of deep learning and the need for big data analysis have created a demand for computers with high computational performance and energy efficiency. Since conventional ASICs and general-purpose CPUs cannot meet this requirement, domain-specific architectures that constrain applications are currently the focus of attention. Coarse-Grained Reconfigurable Architecture (CGRA), one of the domain-specific architectures, is attracting attention because it is superior to CPUs and FPGAs in terms of computational performance and power efficiency [1]. As depicted in Fig. 1, CGRA comprises a two-dimensional array of Processing Elements (PEs) and provides flexibility to change the instructions executed on the PEs and the connections between them, depending on the software being executed. On the other hand, the mapping problem for CGRA is an NP-complete problem, and there exist tradeoffs between solution accuracy and execution time. For real time applications, it is indispensable to shorten the mapping time. Thus, in this study, we propose an Integer Linear Programming (ILP) based mapping method for Elastic CGRA that aims to shorten mapping time while preserving solution accuracy. We have also implemented and preliminary evaluated the proposed method. Makoto Saito, Takuya Kojima, Hideki Takase, Hiroshi Nakamura |
RTCSA | 4 |
| 2022 | GraphDEAR: An Accelerator Architecture for Exploiting Cache Locality in Graph Analytics ApplicationsabstractData structure is the key in Edge Computing where various types of data are continuously generated by ubiquitous devices. Within all common data structures, graphs are used to express relationships and dependencies among human identities, objects, and locations; and they are expected to become one of the most important data infrastructure in the near future. Furthermore, as graph processing often requires random accesses to vast memory spaces, conventional memory hierarchies with caches cannot perform efficiently. To alleviate such memory access bottlenecks in graph processing, we present a solution through vertex accesses scheduling and edge array re-ordering, in parallel with the execution of graph processing application to improve both temporal and spatial locality of memory accesses, especially for edge-centric graphs which are popular means in handling dynamic graphs. Our proposed architecture is evaluated and tested through both trace-based cache simulations and cycle-accurate FPGA-based prototyping. Evaluation results show that our proposal has a potential of significantly reducing the quantity of Miss-Per-Kilo-Instructions (MPKI) for Last Level Cache (LLC) by 56.27% on average. Masaaki Kondo, Yuan He 0002, Ryuichi Sakamoto, Hiroshi Nakamura |
PDP | 7 |
| 2021 | New LDoS Attack in Zigbee Network and its Possible CountermeasuresabstractLow-Rate DoS (LDoS) attacks degrade the quality of service with less traffic than ordinary DoS attacks. LDoS attacks can easily evade conventional counter-DoS detection mechanisms because their time-averaged flow is small. Thus, LDoS attack is getting a serious problem nowadays. With the recent spread of IoT devices, Zigbee, one of the IoT communication standards, attracts much attention. Zigbee is a low-power wireless communication protocol at the sacrifice of its transfer range and bandwidth. Since Zigbee consumes low power, it is widely adopted for inexpensive small sensors and IoT that run on batteries. The advantage of the low power consumption of Zigbee is due to its use of indirect transmission. However, we newly found that vulnerability to LDoS attacks existed in the indirect transmission. In this study, we first point out this vulnerability and describe new LDoS attack scenarios exploiting it. Next, we propose simple countermeasures against them which require low computational cost. Then, simulation experiments are conducted to evaluate the impact of the attacks and the effectiveness of the proposed countermeasures. The experimental results show that the new LDoS attacks can decrease the normal packet arrival rate to 0% by adjusting the timing of attacks. It is also revealed that our proposed countermeasures successfully prevent such cases and improve the minimum arrival rate to 80% at best. Satoshi Okada, Daisuke Miyamoto, Yuji Sekiya, Hiroshi Nakamura |
SMARTCOMP | 4 |
| 2019 | Power Management of Wireless Sensor Nodes with Coordinated Distributed Reinforcement LearningabstractEnergy Harvesting Wireless Sensor Nodes (EHWSNs) require adaptive energy management policies for uninterrupted perpetual operation in their physical environments. Contemporary online Reinforcement Learning (RL) solutions take an unrealistically long time exploring the environment to converge on working policies. Our work accelerates learning by partitioning the state-space for simultaneous exploration by multiple agents. We achieve this by using a novel coordinated e-greedy method and implement it via Distributed RL (DiRL) in an EHWSN network. Our simulation results show a four-fold increase in state-space penetration and reduction in time to achieve optimal operation by an order of magnitude (50x). Moreover, we also propose methods to reduce instances of disastrous outcomes associated with learning and exploration. This translates to reducing the downtimes of the nodes in simulations corresponding to a real-world scenario by one thirds. Shaswot Shresthamali, Masaaki Kondo, Hiroshi Nakamura |
ICCD | 3 |
| 2017 | Dual notch-type high-order frinction-free force observers for force sensorless fine force controlabstractIn order to realize a fine force sensorless force control, this paper proposes a new notch-type friction-free disturbance observer(DOB) and a new notch-type friction-free reaction force observer(FFRFO). Generally, a force sensorless force control system always has influence on static friction phenomenon. To overcome this problem, this paper proposes a new force sensorless control system using the notch-type friction free reaction force observer, whose inner system is the proposed acceleration control based on the notch-type friction free disturbance observer. The effectiveness of the proposed system is confirmed by experiments using an actual industrial robot. Toshimasa Miyazaki, Naoki Kamiya, Hiroshi Nakamura, Yuki Yokokura, Kiyoshi Ohishi |
IECON | 3 |
| 2017 | Energy-aware task scheduling for near real-time periodic tasks on heterogeneous multicore processorsabstractNear real-time periodic tasks, which are popular in multimedia streaming applications, have deadline periods that are longer than the input intervals thanks to buffering. For such applications, the conventional frame-based scheduling cannot realize optimal scheduling due to their shortsighted deadline assumption. To realize globally energy-efficient executions of these applications, we propose a novel task scheduling algorithm, which takes advantage of the long deadline period. We confirm our approach can take advantage of the longer deadline period and reduce the average power consumption by up to 65%. Takashi Nakada, Hiroyuki Yanagihashi, Hiroshi Nakamura, Kunimaro Imai, Hiroshi Ueki, Takashi Tsuchiya, Masanori Hayashikoshi |
VLSI-SoC | 3 |
| 2017 | Adaptive Power Management in Solar Energy Harvesting Sensor Node Using Reinforcement LearningabstractIn this paper, we present an adaptive power manager for solar energy harvesting sensor nodes. We use a simplified model consisting of a solar panel, an ideal battery and a general sensor node with variable duty cycle. Our power manager uses Reinforcement Learning (RL), specifically SARSA(λ) learning, to train itself from historical data. Once trained, we show that our power manager is capable of adapting to changes in weather, climate, device parameters and battery degradation while ensuring near-optimal performance without depleting or overcharging its battery. Our approach uses a simple but novel general reward function and leverages the use of weather forecast data to enhance performance. We show that our method achieves near perfect energy neutral operation (ENO) with less than 6% root mean square deviation from ENO as compared to more than 23% deviation that occur when using other approaches. Shaswot Shresthamali, Masaaki Kondo, Hiroshi Nakamura |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2016 | Fine force control without force sensor based on reaction force estimation system considering static friction and kinetic frictionabstractThis paper proposes a new sensorless force control system based on reaction force estimation system considering static friction and kinetic friction for industrial robot. As industrial robot uses the harmonic gear and the ball screw, the proposed force control system is structured by the resonance ratio control based speed controller. Moreover, as harmonic gear has a static friction and kinetic friction, the reaction force estimation system needs to reduce the influence on static friction and kinetic friction force. The proposed friction-free force estimation system has the two technics. The first one is the dither signal based static friction free force estimation, and the second one is the Stribeck model based kinetic friction free estimation. Using these technics, the friction-free reaction force estimation system is constructed. In this paper, the validity of proposed system is verified by the experimental results using the actual industrial robot. Hiroshi Nakamura, Kiyoshi Ohishi, Yuki Yokokura, Toshimasa Miyazaki, Akifumi Tsukamoto |
IECON | 1 |
| 2016 | An adaptive energy-efficient task scheduling under execution time variation based on statistical analysisabstractNear real-time data processing tasks, such as multimedia streaming applications, exhibit a common fact that their deadline periods are longer than their input intervals due to buffering. Therefore, it is possible to minimize their energy consumption without deadline violations. In this work, we propose an energy efficient slack-based task scheduling algorithm for such tasks by adapting to task size variations and applying DVFS with the help of statistical analysis. We confirmed that our proposal can further reduce the energy consumption when compared to oracle frame-based scheduling. Takashi Nakada, Tomoki Hatanaka, Hiroshi Nakamura, Hiroshi Ueki, Masanori Hayashikoshi, Toru Shimizu |
VLSI-SoC | 3 |
| 2015 | Suggestion-based interactive video digest design by user-system cooperative evolutionabstractThis paper proposes a suggestion-based interactive evolution method for video summarization, which attempts to enhance users' creative thought process without disturbing user edit operation. Although solutions of video summarization are time-varying, a few solutions can be evaluated in parallel. Therefore, the proposed method optimizes summarized video in background asynchronously with user operation. The background optimization is modeled as bi-objective optimization of user preference estimated by user operations and solution novelty. In this, Pareto solutions are stored to an archive and solutions are selected to suggest to the user based on α-domination. Experimental results using an eye-tracking device have revealed that the proposed method enhances convergent thinking rather than divergent thinking. Hiroshi Nakamura, Satoshi Ono |
CEC | 1 |
| 2015 | Immediate sleep: Reducing energy impact of peripheral circuits in STT-MRAM cachesabstractImplementing last level caches (LLCs) with STT-MRAM is a promising approach for designing energy efficient microprocessors due to high density and low leakage power of its memory cells. However, peripheral circuits of an STT-MRAM cache still suffer from leakage power because large and leaky transistors are required to drive large write current to STT-MRAM element. To overcome this problem, we propose a new power management scheme called Immediate Sleep (IS). IS immediately turns off a subarray of an STT-MRAM cache if the next access is predicted to be not critical in performance. Thus, IS can effectively reduce leakage energy with little impact on performance. Our experimental results show that our technique can save the leakage energy of an STT-MRAM LLC by 32% compared to an STT-MRAM LLC with the conventional scheme at the same performance. Eishi Arima, Hiroki Noguchi, Takashi Nakada, Shinobu Miwa, Susumu Takeda, Shinobu Fujita, Hiroshi Nakamura |
ICCD | 7 |
| 2015 | Runtime multi-optimizations for energy efficient on-chip interconnections1abstractOn-chip interconnection (or NoC) is a major performance and power contributor to modern and future multicore processors. So far, many optimization techniques have been developed to improve its bandwidth, latency and power consumption. But it is not clear how energy efficiency is affected since an optimization technique normally comes with overheads. This paper thus attempts to address when and how such optimization techniques should be applied and tuned to help achieve better energy efficiency. We firstly model the performance and energy impacts of representative NoC optimization techniques. These models help us more easily understand the consequences when applying these optimization techniques and their combinations under different circumstances. Moreover, based on such modeling, we propose and implement an adaptive control over these NoC optimization techniques to improve both performance and energy efficiency of the network. Our results show that, this proposal can achieve an average improvement of 26% and 57% on network performance and energy delay product, respectively. Yuan He 0002, Masaaki Kondo, Takashi Nakada, Hiroshi Sasaki 0001, Shinobu Miwa, Hiroshi Nakamura |
ICCD | 6 |
| 2015 | Position sensorless control with accelerometer for linear and curvilinear synchronous motorabstractIn the field of the manufacturing industry, it has been often discussed that high-mix low-volume manufacturing is required to meet a variety of customer needs. In such a situation, cellular manufacturing and U shaped production lines are often adopted. Linear synchronous motors (LSMs) are appropriate to realize U shaped automation lines because of its strong features such as flexibility of layout, high efficiency and position control performance. However, a problem seems to lie in the fact that LSMs need a position sensor which is expensive and has less durability. For this reason, the position sensorless control for LSMs is required. In addition, the method should be easy to apply to all kinds of LSMs. Specifically, the method should require no saliency and be apply to a LSM which has a complex track including curvilinear section. Based on such a background, this paper presents a position estimation method for a Linear and Curvilinear Synchronous Motor with an accelerometer which is inexpensive and easy to attach to LSMs. As the proposed method in this paper does not require any motor saliency, it can be quickly applied to all LSMs. In this paper, the proposed method is shown and verified by experimental results. Yoshiyasu Takase, Hiroshi Nakamura, Minoru Koga, Toru Shikayama, Akihito Toyota |
IECON | 2 |
| 2015 | Profile-based power shifting in interconnection networks with on/off linksabstractOverprovisioning hardware devices and coordinating their power budgets are proposed to improve the application performance of future power-constrained HPC systems. This coordination process is called power shifting. Meanwhile, recent studies have revealed that on/off links can save network power in HPC systems. Future HPC systems will thus adopt on/off links in addition to power shifting. This paper explores power shifting in interconnection networks with on/off links. Given that on/off links keep network power low at application runtime, we can transfer appreciable quantities of power budgets on networks to other devices before an application runs. We thus propose a profile-based power shifting technique that allows HPC users to transfer the power budget remaining on networks to other devices at the time of job dispatch. Experimental results show that the proposed technique appreciably improves application performance under various power constraints. Shinobu Miwa, Hiroshi Nakamura |
SC | 2 |
| 2014 | Normally-off computing project: Challenges and opportunitiesabstractNormally-Off is a way of computing which aggressively powers off components of computer systems when they need not to operate. Simple power gating cannot fully take the chances of power reduction because volatile memories lose data when power is turned off. Recently, new non-volatile memories (NVMs) have appeared. High attention has been paid to normally-off computing using these NVMs. In this paper, its expectation and challenges are addressed with a brief introduction of our project started in 2011. Hiroshi Nakamura, Takashi Nakada, Shinobu Miwa |
ASP-DAC | 1 |
| 2014 | Design and control methodology for fine grain power gating based on energy characterization and code profiling of microprocessorsabstractThis paper presents a design and control scheme of a microprocessor whose internal function units are power gated at instruction-by-instruction basis. Enabling/disabling the power gating is adaptively controlled under the support of on-chip leakage monitors and the operating system to minimize energy overhead due to sleep-in and wakeup. Measured results of the fabricated chip in the 65nm CMOS technology demonstrated that our approach reduces energy to 21-35% in the range of 25-85°C as compared to the non power-gated case. Energy dissipation was reduced by up to 15% as compared to the conventional fine-grain power gating technique in the same temperature range. Kimiyoshi Usami, Masaru Kudo, Kensaku Matsunaga, Tsubasa Kosaka, Yoshihiro Tsurui, Hideharu Amano, Hiroaki Kobayashi, Ryuichi Sakamoto, Mitaro Namiki, Masaaki Kondo, Hiroshi Nakamura |
ASP-DAC | 12 |
| 2014 | Design and evaluation of fine-grained power-gating for embedded microprocessorsabstractPower-performance efficiency is still remaining a primary concern for microprocessor designers. One of the sources of power inefficiency for recent LSI chips is increasing leakage power consumption. Power-gating is a well known technique to reduce leakage power consumption by switching off the power supply to idle logic blocks. Recently, fine-grained power-gating is emerged as a technique to minimize leakage current during the active processor cycles by switching on and off a logic blocks in much finer temporal/spatial granularity. Though fine-grained power-gating is useful, a comprehensive evaluation and analysis has not been conducted on a real LSI chips. In this paper, we evaluate fine-grained run-time power-gating for microprocessors' functional units using a real embedded microprocessor. We also introduce an architecture and compiler co-operative power-gating scheme which mitigates negative power reduction caused by the energy overhead associated with finegrained power-gating. The experimental results with a fabricated core shows that a hardware-based scheme saves power consumption of functional units by 44% and hardware compiler co-operative scheme further improves power efficiency by 5.9% when core temperature is 25 ˚C. Masaaki Kondo, Hiroaki Kobayashi, Ryuichi Sakamoto, Motoki Wada, Jun Tsukamoto, Mitaro Namiki, Hideharu Amano, Kensaku Matsunaga, Masaru Kudo, Kimiyoshi Usami, Toshiya Komoda, Hiroshi Nakamura |
DATE | 13 |
| 2014 | Proposal of position sensorless control and torque ripple compensation based on torque sensor feedbackabstractThis paper presents the following two methods based on torque sensor feedback: new position sensorless control method for surface-mounted permanent magnet synchronous machines (SPMSMs) and torque ripple compensation. The position sensorless techniques can be classified into two categories: the techniques based on the back electromotive force (back-EMF) and the techniques based on the high frequency injection (HFI). However, in general, the former category is weak in low speed, and the later one can not drive without the saliency. Furthermore, both of the categories have a problem that torque ripples are caused by transient position estimation errors. Meanwhile, in the field of robot research in recent years, the developments of torque sensors are active. This situation has suggested that small and reasonable torque sensors will be available in the foreseeable future. Therefore, we developed the torque sensor which can be attached to SPMSMs and the new position sensorless control method based on the torque sensor feedback. Using the proposed method, the position sensorlesss drive corresponding to the entire speed range including zero is achieved. Moreover, this method is applicable to SPMSMs which do not have saliency. In addition, the torque ripples are well suppressed. Effectiveness of the proposed method is verified by experimental results. Yoshiyasu Takase, Hiroshi Nakamura, Takashi Mamba |
IECON | 2 |
| 2013 | McRouter: Multicast within a router for high performance network-on-chipsabstractThe inevitable advent of the multi-core era has driven an increasing demand for low latency on-chip inter-connection networks (or NoCs). Being a critical part of the memory hierarchy for modern chip multi-processors (CMPs), these networks face stringent design constraints to provide fast communication with tight power budget. Modern NoC's first-order concern is clearly its latency, while we also find that internal bandwidth of its routers is relatively plentiful; thus, we present a low latency router design utilizing a technique we call “multicast within a router” or McRouter, which allows productive utilization of remaining bandwidth inside a NoC router. McRouter allows a single cycle transfer of flits which shortens the communication latency when there is enough remaining bandwidth within the router. The key idea is to transmit a header flit to all possible output ports (multicast) so that it is always transmitted to the correct output port without relying on route computation. In addition, we find it is affordable with marginal power overhead while still being a stand-alone design by maintaining portability and modularity (unlike look-ahead routing based designs). Our evaluation with application traffic shows that McRouter helps achieving system speed-ups of 1.28, 1.17 and 1.05 over the conventional router (CR), the VSA router (VSAR) and the prediction router (PR), respectively. Yuan He 0002, Hiroshi Sasaki 0001, Shinobu Miwa, Hiroshi Nakamura |
PACT | 4 |
| 2013 | D-MRAM cache: enhancing energy efficiency with 3T-1MTJ DRAM/MRAM hybrid memoryabstractThis paper describes a proposal of non-volatile cache architecture utilizing novel DRAM / MRAM cell-level hybrid structured memory (D-MRAM) that enables effective power reduction for high performance mobile SoCs without area overhead. Here, the key point to reduce active power is intermittent refresh process for the DRAM-mode. D-MRAM has advantage to reduce static power consumptions compared to the conventional SRAM, because there are no static leakage paths in the D-MRAM cell and it is not needed to supply voltage to its cells when used as the MRAM-mode. Besides, with advanced perpendicular magnetic tunnel junctions (p-MTJ), which decreases the write energy and latency without shortening its retention time, D-MRAM is capable of power reduction by replacing the traditional SRAM caches. Considering the 65-nm CMOS technology, the access latencies of 1MB memory macro are 2.2 ns / 1.5 ns for read / write in DRAM mode, and 2.2 ns / 4.5 ns in MRAM mode, while those of SRAM are 1.17 ns. The SPEC CPU2006 benchmarks have revealed that the energy per instruction (EPI) of the total cache memory can be dramatically reduced by 71 % on average, and the instruction per cycle (IPC) performance of the D-MRAM cache architecture degraded only by approximately 4 % on average in spite of its latency overhead. Hiroki Noguchi, Kumiko Nomura, Keiko Abe, Shinobu Fujita, Eishi Arima, Kyundong Kim, Takashi Nakada, Shinobu Miwa, Hiroshi Nakamura |
DATE | 9 |
| 2013 | Demonstration of a heterogeneous multi-core processor with 3-D inductive coupling linksabstractCube-1 is a heterogeneous multi-core processor which can achieve the required performance with the least energy consumption as possible. It can control the performance and energy with two levels: (1) the number of accelerators can be easily changed by increasing or decreasing the number of stacked chips after fabrication, as they are connected with inductive coupling links. (2) The supply voltage for PE array of the accelerator can be controlled by the host CPU so that the required performance can be obtained with a minimum supply voltage. Yusuke Koizumi, Noriyuki Miura, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPL | 11 |
| 2013 | A scalable 3D heterogeneous multi-core processor with inductive-coupling thruchip interface
Noriyuki Miura, Yusuke Koizumi, Eiichi Sasaki, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
Hot Chips Symposium | 12 |
| 2013 | Power capping of CPU-GPU heterogeneous systems through coordinating DVFS and task mappingabstractFuture computer systems are built under much stringent power budget due to the limitation of power delivery and cooling systems. To this end, sophisticated power management techniques are required. Power capping is a technique to limit the power consumption of a system to the predetermined level, and has been extensively studied in homogeneous systems. However, few studies about the power capping of CPU-GPU heterogeneous systems have been done yet. In this paper, we propose an efficient power capping technique through coordinating DVFS and task mapping in a single computing node equipped with GPUs. In CPU-GPU heterogeneous systems, settings of the device frequencies have to be considered with task mapping between the CPUs and the GPUs because the frequency scaling can incurs load imbalance between them. To guide the settings of DVFS and task mapping for avoiding power violation and the load imbalance, we develop new empirical models of the performance and the maximum power consumption of a CPU-GPU heterogeneous system. The models enable us to set near-optimal settings of the device frequencies and the task mapping in advance of the application execution. We evaluate the proposed technique with five data-parallel applications on a machine equipped with a single CPU and a single GPU. The experimental result shows that the performance achieved by the proposed power capping technique is comparable to the ideal one. Toshiya Komoda, Shingo Hayashi, Takashi Nakada, Shinobu Miwa, Hiroshi Nakamura |
ICCD | 5 |
| 2013 | Integrating Multi-GPU Execution in an OpenACC CompilerabstractGPUs have become promising computing devices in current and future computer systems due to its high performance, high energy efficiency, and low price. However, lack of high level GPU programming models hinders the wide spread of GPU applications. To resolve this issue, OpenACC is developed as the first industry standard of a directive-based GPU programming model and several implementations are now available. Although early evaluations of the OpenACC systems showed significant performance improvement with modest programming efforts, they also revealed the limitations of the systems. One of the biggest limitations is that the current OpenACC compilers do not automate the utilization of multiple GPUs. In this paper, we present an OpenACC compiler with the capability to execute single GPU OpenACC programs on multiple GPUs. By orchestrating the compiler and the runtime system, the proposed system can efficiently manage the necessary data movements among multiple GPUs memories. To enable advanced communication optimizations in the proposed system, we propose a small set of directives as extensions of OpenACC API. The directives allow programmers to express the patterns of memory accesses in the parallel loops to be offloaded. Inserting a few directives into an OpenACC program can reduce a large amount of unnecessary data movements and thus helps the proposed system drawing great performance from multi-GPU systems. We implemented and evaluated the prototype system on top of CUDA with three data parallel applications. The proposed system achieves up to 6.75x of the performance compared to OpenMP in the 1CPU with 2GPU machine, and up to 2.95x of the performance compared to OpenMP in the 2CPU with 3GPU machine. In addition, in two of the three applications, the multi-GPU OpenACC compiler outperforms the single GPU system where hand-written CUDA programs run. Toshiya Komoda, Shinobu Miwa, Hiroshi Nakamura, Naoya Maruyama |
ICPP | 3 |
| 2013 | Proposal of step climbing of wheeled robot using slip ratio controlabstractThis paper proposes a method which enables wheeled robot to climb a step that is higher than the radius of the wheel. In conventional method, this is impossible, because driving force of front wheel is not enough. This paper proposes a method utilizing the moment of impact between the wheel and the step to maximize the normal force, and hence the driving force of the front wheel. Effectiveness of the proposed method is verified by simulations and experiments. Masaki Higashino, Hiroshi Fujimoto, Yoshiyasu Takase, Hiroshi Nakamura |
IECON | 4 |
| 2013 | Performance modeling for designing NoC-based multiprocessorsabstractNetwork-on-Chip (NoC) based multiprocessors have become popular as a scalable alternative to classical bus architectures. The performance evaluation of NoC-based multiprocessors is largely based on simulation. However, precise simulation is extremely slow. Additionally, there are many design parameters that affect the total performance. Therefore, it is practically impossible to use the precise simulation for the design space exploration purposes. To alleviate this problem, prototyping NoC systems and estimating their performances are critically important. In this paper, we present a generalized novel performance model that combined with the simulations for designing NoC-based multiprocessors. We revealed that the performance impact of cache and network latencies are dominant. Moreover, network congestion rarely happens under near appropriate configuration. Thus, the performance model is mainly constructed using the hardware parameters and the statistics that obtained from a simple cache simulation that is separated from the network behavior. The proposed performance model is used not only to obtain fast and accurate performance, but also to guide the NoC-based multiprocessor design space exploration. The accuracy of our approach and its practical use are illustrated through simulation. The results showed that proposed model can estimate performance with only 3.4% error on average and 21% at worst. We also confirmed that our evaluation framework can estimate 360 times faster than the brute force full system simulation. Takashi Nakada, Shinobu Miwa, Keisuke Y. Yano, Hiroshi Nakamura |
RSP | 4 |
| 2012 | Scalability-based manycore partitioningabstractMulticore processors have been popular for years, and the industry is gradually shifting towards the era of manycore processors. Single-thread performance of microprocessors is not growing at a historical rate, but the existence of a number of active processes in the computer system and the continuing development of multi-threaded applications benefit from the growing core counts to sustain system throughput. This trend brings us a situation where a number of parallel applications simultaneously being executed on a single system. Since multi-threaded applications try to maximize its throughput by utilizing the whole system, each of them usually create equal or larger number of threads compared to underlying logical core counts. This introduces much greater number of threads to be co-scheduled in the entire system. However, each program has different characteristics (or scalability) and contends for shared resources, which are the CPU cores and memory hierarchies, with each other. Therefore, it is clear that OS thread scheduling will play a major role in achieving high system performance under such conditions. We develop a sophisticated scheduler that (1) dynamically predicts the scalability of programs via the use of hardware performance monitoring units, (2) decides the optimal number of cores to be allocated for each program, and (3) allocates the cores to programs while maximizing the system utilization to achieve fair and maximum performance. The evaluation results on a 48-core AMD Opteron system show improvements over the Linux scheduler for a variety of multiprogramming workloads. Hiroshi Sasaki 0001, Teruo Tanimoto, Koji Inoue, Hiroshi Nakamura |
PACT | 4 |
| 2012 | A multi-Vdd dynamic variable-pipeline on-chip router for CMPsabstractWe propose a multi-voltage (multi-Vdd) variable pipeline router to reduce the power consumption of Network-on-Chips (NoCs) designed for chip multi-processors (CMPs). Our multi-Vdd variable pipeline router adjusts its pipeline depth (i.e., communication latency) and supply voltage level in response to the applied workload. Unlike dynamic voltage and frequency scaling (DVFS) routers, the operating frequency is the same for all routers throughout the CMP; thus, there is no need to synchronize neighboring routers working at different frequencies. In this paper, we implemented the multi-Vdd variable pipeline router, which selects two supply voltage levels and pipeline modes, using a 65nm CMOS process and evaluated it using a full-system CMP simulator. Evaluation results show that although the application performance degraded by 1.0% to 2.1%, the standby power of NoCs reduced by 10.4% to 44.4%. Hiroki Matsutani, Yuto Hirata, Michihiro Koibuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
ASP-DAC | 5 |
| 2012 | CMA-Cube: A scalable reconfigurable accelerator with 3-D wireless inductive coupling interconnectabstractCMA-Cube is the second prototype of building block scalable reconfigurable accelerator using inductive coupling interconnect. It uses the wireless inductive coupling interconnect as a packet switching network which connects accelerators. As an accelerator core, CMA (Cool Mega Array), which consists of a large coarse-grained PE array with combinatorial circuits and tiny micro-controller, is applied. Evaluation results of Cube-1 Quad Core which consists of a host embedded CPU and three CMA-Cubes achieved 3.15 times performance acceleration as that without accelerators when JPEG decoder is executed. Yusuke Koizumi, Eiichi Sasaki, Hideharu Amano, Hiroki Matsutani, Yasuhiro Take, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPL | 11 |
| 2012 | Dynamic power control with a heterogeneous multi-core system using a 3-D wireless inductive coupling interconnectabstractCube-2 is a prototype of building block scalable reconfigurable accelerator using an inductive coupling interconnect. It is consisting of a ultra low leakage embedded processor Geyser and coarse-grained reconfigurable accelerators CMA (Cool Mega Array). A Geyser chip and multiple CMA chips are stacked, and a powerful network is formed by using the inductive coupling interconnect. The performance can be enhanced by increasing the number of CMA chips. JPEG decoder is implemented with a cooperation of Geyser and CMAs, and low power execution by controlling the power supply voltage of CMAs is demonstrated. Yusuke Koizumi, Hideharu Amano, Hiroki Matsutani, Noriyuki Miura, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPT | 10 |
| 2012 | A novel power-gating scheme utilizing data retentiveness on cachesabstractCaches are one of the most leakage consuming components in modern processor because of massive amount of transistors. To reduce leakage power of caches, several techniques using power-gating(PG) were proposed. Despite of its high leakage saving, a side effect of PG for caches is the loss of data during a sleep. If useful data is lost in sleep mode, it should be fetched again from a lower level memory. This consumes a considerable amount of energy, which very unfortunately mitigates the leakage saving. This paper proposes a new PG scheme considering data retentiveness of SRAM. After entering the sleep mode, data of an SRAM cell is not lost immediately and is usable by checking the validity of the data. Therefore, we utilize data retentiveness of SRAM to avoid energy overhead for data recovery, which results in further chance of leakage saving. To check availability, we introduce a simple hardware whose overhead is ignorable. We also examined leakage saving potential of our approach. For both L1 data and instruction caches, our scheme results in more than 2 times of smaller leakage energy compared to conventional PG scheme. Kyundong Kim, Seidai Takeda, Shinobu Miwa, Hiroshi Nakamura |
ACM Great Lakes Symposium on VLSI | 4 |
| 2012 | Stepwise sleep depth control for run-time leakage power savingabstractRecently, run-time sleep control scheme using multiple sleep modes have been studied. In those studies, each sleep mode has its own sleep depth. Deeper sleep mode provides higher leakage saving but incurs larger overhead energy.Use of multiple modes is helpful for further leakage saving if an appropriate mode is selected, but the best mode depends on the idle period whose length cannot be told in advance. Although the implementations how to realize different sleep depths have been well studied, few attention has been paid to the method of how to select the best sleep depth dynamically during execution. This paper proposes a simple but novel sleep control scheme, called stepwise sleep depth control, which aims to select the best depth among provided multiple sleep depths.Our scheme automatically applies deeper depth in a step-by-step manner after an idle state starts. It successfully reduces leakage energy while only a small modification is required for circuit implementation. This paper also proposes a methodology for optimizing control parameters of our sleep control scheme according to program behavior and temperature. Experimental result shows that stepwise sleep depth control applied to body biasing circuit improves net leakage saving of up to 43% for FPAlu at 1.0GHz, 75°C compared to conventional reverse body biasing. Seidai Takeda, Shinobu Miwa, Kimiyoshi Usami, Hiroshi Nakamura |
ACM Great Lakes Symposium on VLSI | 4 |
| 2011 | Geyser-2: The second prototype CPU with fine-grained run-time power gatingabstractGeyser-2 is the second prototype MIPS CPU which provides a fine-grained run-time power gating (PG) controlled by instructions. Geyser-l, the first prototype only provides the fine-grained run-time PG core. Although it demonstrated the leakage power reduction on a real chip, the operational frequency is limited at 60MHz because of the limitation of the I/O speed. Geyser-2 with cache and TLB mechanism is implemented to show (1) run-time PG works at least with 200MHz which is commonly used clock for embedded systems, and (2) it is also efficient on the environment with real application programs with an operating system. Daisuke Ikebuchi, Yoshiki Saito, M. Kamata, Naomi Seki, Yu Kojima, Hideharu Amano, Satoshi Koyama, Tatsunori Hashida, Y. Umahashi, D. Masuda, Kimiyoshi Usami, Mitaro Namiki, Seidai Takeda, Hiroshi Nakamura, Masaaki Kondo |
ASP-DAC | 16 |
| 2011 | Cool Mega-Array: A highly energy efficient reconfigurable acceleratorabstractA highly energy efficient reconfigurable accelerator called CMA (Cool Mega-Array) is proposed. It consists of a large Processing Element (PE) array without memory elements for maintain result of ALU and configuration data, a small simple programmable micro controller for data management, and the data memory. Unlike traditional coarse grained reconfigurable processors, the power consumption for hardware context switching, storing intermediate data in registers, and clock distribution for them are eliminated from PE array which occupies large area of a chip. Configuration registers are collected to small area of micro controller. The data flow graph mapped on the PE array is static during execution. Various application programs can be implemented by making the best use of flexible data management instructions with the micro controller. When the delay time in the PE array is longer than the data handling time with the micro controller, the supply voltage for the PE array is scaled to reduce the power consumption without degrading the performance. In the opposite case, wave pipelining is applied to enhance PE array performance. A prototype chip CMA-1 with 8 × 8 PE array with 24-bit data width was fabricated in 2.1 × 4.2mm265-nm CMOS technology, and achieves 2.4-GOPS/11.2-mW sustained performance. This energy efficiency is comparable to that of the most energy efficient accelerators that have been reported. Nobuaki Ozaki, Yoshihiro Yasuda, Yoshiki Saito, Daisuke Ikebuchi, Masayuki Kimura, Hideharu Amano, Hiroshi Nakamura, Kimiyoshi Usami, Mitaro Namiki, Masaaki Kondo |
FPT | 7 |
| 2011 | On-chip detection methodology for break-even time of power gated function units
Kimiyoshi Usami, Yuya Goto, Kensaku Matsunaga, Satoshi Koyama, Daisuke Ikebuchi, Hideharu Amano, Hiroshi Nakamura |
ISLPED | 7 |
| 2011 | Performance, Area, and Power Evaluations of Ultrafine-Grained Run-Time Power-Gating Routers for CMPsabstractThis paper proposes the ultrafine-grained run-time power gating of on-chip routers, in which the power supply to each router component (e.g., virtual-channel buffer, virtual-channel multiplexer, and crossbar multiplexer and output latch) can be individually controlled based on the applied workload. Since only the router components that are transferring a packet are activated, the leakage power of the on-chip network can be reduced to a near-optimal level. However, such techniques inherently increase the communication latency and degrade the application performance, since a certain amount of wakeup latency is required to activate the sleeping components. To mitigate this wakeup latency, an early wakeup method that can preliminarily detect the next packet arrival and activate the corresponding components is essential. We designed and implemented an ultrafine-grained power-gating router using a commercial 65 nm process. We propose four early wakeup methods and combine them with the power-gating router. The proposed router with the early wakeup methods is evaluated in terms of its application performance, area overhead, and leakage power reduction taking into account the on/off energy overhead. The simulation results showed that it reduces the leakage power by 54.4-59.9% on average even when the application programs are fully running, at the expense of 4.6% of the area and 0.7-3.7% of the performance overheads when we assume a 1 GHz operation. Hiroki Matsutani, Michihiro Koibuchi, Daisuke Ikebuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2010 | Geyser-1: a MIPS R3000 CPU core with fine-grained run-time power gatingabstractGeyser-1 is a MIPS CPU which provides a fine-grained run-time power gating (PG) controlled by instructions. Unlike traditional PGs, it uses special standard cells in which the virtual ground (VGND) is separated from the real ground, and a certain number of the sleep transistors are inserted for quick power shut-down and wake-up. In Geyser-1, the fine-grained run-time PG is applied to computational modules in the execution stage. The power shut-down and wakeup are controlled with architectural and software level. This implementation is the first available CPU with this type of run-time PG technique. Geyser-1 has both time and spatial fine-grained PG and works well with a real chip. Daisuke Ikebuchi, Naomi Seki, Yu Kojima, M. Kamata, Hideharu Amano, Toshiaki Shirai, Satoshi Koyama, Tatsunori Hashida, Y. Umahashi, Hiroki Masuda, Kimiyoshi Usami, Seidai Takeda, Hiroshi Nakamura, Mitaro Namiki, Masaaki Kondo |
ASP-DAC | 14 |
| 2010 | Ultra Fine-Grained Run-Time Power Gating of On-chip Routers for CMPsabstractThis paper proposes an ultra fine-grained run-time power gating of on-chip router, in which power supply to each router component (e.g., VC queue, crossbar MUX, and output latch) can be individually controlled in response to the applied workload. As only the router components which are just transferring a packet are activated, the leakage power of the on-chip network can be reduced to the near-optimal level. However, a certain amount of wakeup latency is required to activate the sleeping components, and the application performance will be degraded. In this paper, we estimate the wakeup latency for each component based on circuit simulations using a 65 nm process. Then we propose four early wakeup methods to overcome the wakeup latency. The proposed router with the early wakeup methods is evaluated in terms of the application performance, area, and leakage power. As a result, it reduces the leakage power by 78.9%, at the expense of the 4.3% area and 4.0% performance when we assume a 1 GHz operation. Hiroki Matsutani, Michihiro Koibuchi, Daisuke Ikebuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
NOCS | 5 |
| 2009 | Cooperative shared resource access control for low-power chip multiprocessorsabstractIn a single-chip multiprocessor (CMP), the last-level cache and its lower memory hierarchy components are typically shared by multiple processors. Conflicts in these resources lead to poor overall performance of the CMP and/or unpredictable performance of the individual cores. If applications on different cores have different performance constraints, even though these constraints can be satisfied by dynamic voltage and frequency scaling (DVFS) control of each core, conflicts in shared resources will lead to increased power consumption. Therefore, in the present paper, we derive a condition whereby, under resource conflicts, the total power consumption is minimized by a newly developed power consumption model and propose a method by which to minimize the power consumption of CMPs by cooperative access control of multiple shared resources and DVFS control. Experimental results reveal that the proposed technique can reduce power consumption by 15% on average in a dual-core CMP and by 13% in a quad-core CMP, as compared to the case in which only DVFS control is applied. Noriko Takagi, Hiroshi Sasaki 0001, Masaaki Kondo, Hiroshi Nakamura |
ISLPED | 4 |
| 2009 | Pangea: An Eager Database Replication Middleware guaranteeing Snapshot Isolation without Modification of Database ServersabstractRecently, several middleware-based approaches have been proposed. If we implement all functionalities of database replication only in a middleware layer, we can avoid the high cost of modifying existing database servers or scratch-building. However, it is a big challenge to propose middleware which can enhance performance and scalability without modification of database servers because the restriction may cause extra overhead. Unfortunately, many existing middleware-based approaches suffer from several shortcomings, i.e., some cause a hidden deadlock, some provide only table-level locking, some rely on total order communication tools, and others need to modify existing database servers. In this paper, we propose Pangea, a new eager database replication middleware guaranteeing snapshot isolation that solves the drawbacks of existing middleware by exploiting the property of the first updater wins rule. We have implemented the prototype of Pangea on top of PostgreSQL servers without modification. An advantage of Pangea is that it uses less than 2000 lines of C code. Our experimental results with the TPC-W benchmark reveal that, compared to an existing middleware guaranteeing snapshot isolation without modification of database servers, Pangea provides better performance in terms of throughput and scalability. Takeshi Mishima, Hiroshi Nakamura |
Proc. VLDB Endow. | 2 |
| 2009 | Energy-Efficient Dynamic Instruction Scheduling Logic Through Instruction GroupingabstractDynamic instruction scheduling logic is quite complex and dissipates significant energy in microprocessors that support superscalar and out-of-order execution. We propose a novel microarchitectural technique to reduce the complexity and energy consumption of the dynamic instruction scheduling logic. The proposed method groups several instructions as a single issue unit and reduces the required number of ports and the size of the structure. This paper describes the microarchitecture mechanisms and shows evaluation results for energy savings and performance. These results reveal that the proposed technique can greatly reduce energy with almost no performance degradation, compared to the conventional dynamic instruction scheduling logic. Hiroshi Sasaki 0001, Masaaki Kondo, Hiroshi Nakamura |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | A fine-grain dynamic sleep control scheme in MIPS R3000abstractA fine-grain dynamic power gating is proposed for saving the leakage power in MIPS R3000 by sleep control and applied to a processor pipeline. An execution unit is divided into four small units: multiplier, divider, shifter and other (CLU). The power of each unit is cut off dynamically, based on the operation. We tape-outed the prototype chip Geyser-0, which provides an R3000 Core with the power reduction technique, 16 KB caches and translation lookaside buffer (TLB) using 90 nm CMOS technology. The evaluation results of four benchmark programs for embedded applications show that 47% of the leakage power is reduced on average with 41% area overhead. Naomi Seki, Jo Kei, Daisuke Ikebuchi, Yu Kojima, Yohei Hasegawa, Hideharu Amano, Toshihiro Kashima, Seidai Takeda, Toshiaki Shirai, Mitsutaka Nakata, Kimiyoshi Usami, Tetsuya Sunata, Jun Kanai, Mitaro Namiki, Masaaki Kondo, Hiroshi Nakamura |
ICCD | 17 |
| 2008 | Detecting Inconsistent Values Caused by Interaction Faults Using Automatically Located Implicit RedundanciesabstractThis paper addresses the problem of detecting inconsistent values caused by interaction faults originated from an external system.This type of error occurs when a correctly formatted message that is not corrupted during transmission is generated with a field that contains incorrect data.When traditional schemes cannot be used, one alternative is resorting to receiver-based strategies that employ implicit redundancies - relations between events or data, often identified by a human expert.We propose an approach for detecting inconsistent values using implicit redundancies which are automatically located in examples of communications.We show that, even without adding any redundant information to the communication, the proposed approach can achieve a reasonable error detection coverage in fields where sequential relations exist.Other aspects, such as false alarms and latency, are also evaluated. Bogdan Tomoyuki Nassu, Takashi Nanya, Hiroshi Nakamura |
PRDC | 3 |
| 2007 | Interactive presentation: Task scheduling under performance constraints for reducing the energy consumption of the GALS multi-processor SoCabstractThe present paper focuses on applications that are periodic and have both latency and throughput constraints. For these applications, pipeline scheduling is effective for reducing energy consumption. Thus, the present paper proposes a pipelined task scheduling method for minimizing the energy consumption of GALS MP-SoC under latency and throughput constraints. First, we model target GALS MP-SoC architecture and application tasks. We then show that the energy optimization problem under this model belongs to the class of mixed-integer linear programming. Next, we propose a new scheduling method based on simulated annealing for the purpose of solving this problem quickly. Finally, experimental results demonstrate that the proposed method achieves a significant energy reduction on a real application under a practical architecture Ryo Watanabe, Masaaki Kondo, Masashi Imai, Hiroshi Nakamura, Takashi Nanya |
DATE | 4 |
| 2007 | Fast AbstractsabstractFast Abstracts are brief two page presentations, either on new ideas, opinion pieces, or a project update. They cover wide variety of issues within the field of dependable systems and networks. They are also designed to offer an opportunity for late-breaking results, partial results, or work in progress to be reported in a timely fashion. As such, they are lightly reviewed by the Fast Abstracts Program Committee and are not subjected to the rigorous referee process for regular DSN papers. The late submission deadline and expedited screening, along with the corresponding 5-minute talk during DSN, allow for very rapid dissemination and timely feedback from the community. Hiroshi Nakamura |
DSN | 1 |
| 2007 | Power reduction of chip multi-processors using shared resource control cooperating with DVFSabstractThis paper presents a novel power reduction method for chip multi-processors (CMPs) under real-time constraints. While the power consumption of processing units (PUs) on CMPs can be reduced without violating real-time constraints by dynamic voltage and frequency scaling (DVFS), the clock frequency of each PU cannot be determined independently because of the performance impact caused by the conflict for the shared resources. To minimize power consumption in this situation, we first derive an analytical model which provides the optimal priority and clock frequency setting, and then propose a method of controlling the priority of shared resource accesses in cooperation with DVFS. From the analytical model, in dual-core CMPs, we reveal that the total power consumption is minimized when the clock frequency of two PUs becomes the same. An experiment with a synthetic benchmark supports the validity of the analytical model and the evaluation results with real applications show that the proposed method reduces the power consumption by up to 15% and 6.7% on average compared with a conventional DVFS technique. Ryo Watanabe, Masaaki Kondo, Hiroshi Nakamura, Takashi Nanya |
ICCD | 3 |
| 2007 | A High Performance Cluster System Design by Adaptie Power ControlabstractThe first order design constraint in dense packaged clusters is power consumption. The currently developed cluster systems are conservatively designed so that the expected peak power does not exceed the power limit. However, practical power consumption seldom reaches the peak power. In this paper, we propose a new approach to design a high performance cluster system by an adaptive power control technique. Our approach is to integrate many computation nodes into a system whose total theoretical peak power exceeds the limit and to control runtime effective power by optimizing the number of working nodes and/or the clock frequency of the processors. We show the algorithm of the adaptive power control and performance evaluation by using a real cluster system. Evaluation results show that our proposed approach greatly improves performance as large as 46% compared to a conventional cluster system. Masaaki Kondo, Yoshimichi Ikeda, Hiroshi Nakamura |
IPDPS | 3 |
| 2007 | A Proposal of New Dependable Database Middleware with Consistency and Concurrency ControlabstractWe propose a new dependable database middleware that can synchronize off-the-shelf database servers for consistency and execute write queries concurrently for high throughput. Our proposal also helps to realize low cost system since both existing servers and client applications can be used without modification. We implemented a prototype using PostgreSQL without modification. Our experimental result reveals that our approach outperforms the conservative proposals. Takeshi Mishima, Hiroshi Nakamura |
PRDC | 2 |
| 2006 | MegaProto/E: power-aware high-performance cluster with commodity technologyabstractIn our research project named "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we have been developing a prototype cluster not based on ASIC or FPGA but instead only using commodity technology. Its packaging is extremely compact and dense, and its performance/power ratio is very high. Our previous prototype system named "MegaProto" demonstrated that one cluster unit, which consists of 16 commodity low-power processors, can be successfully implemented on just 1U height chassis and it is capable of up to 2.8 times higher performance/power ratio than ordinary high-performance dual-Xeon 1U server units. We have improved MegaProto by replacing the CPU and enhancing the I/O performance. The new cluster unit named "MegaProto/E" with 16 Transmeta Efficeon processors achieves 32 GFlops of peak performance, which is 2.2-fold greater than that of the original one. The cluster unit is equipped with an independent dual network of Gigabit Ethernet, including dual 24-port switches. The maximum power consumption of the cluster unit is 320 W, which is comparable with that of today's high-end PC servers for high performance clusters. Performance evaluation using NPB kernels and HPL shows that the performance of MegaProto/E exceeds that of a dual-Xeon server in all the benchmarks, and its performance ratio ranges from 1.3 to 3.7. These results reveal that our solution of implementing a number of ultra low-power processors in compact packaging is an excellent way to achieve extremely high performance in applications with a certain degree of parallelism. We are now building a multi-unit cluster with 128 CPUs (8 units) to prove that this advantage still holds with higher scalability Taisuke Boku, Mitsuhisa Sato, Daisuke Takahashi, Hiroshi Nakashima, Hiroshi Nakamura, Satoshi Matsuoka, Yoshihiko Hotta |
IPDPS | 5 |
| 2006 | Energy-efficient dynamic instruction scheduling logic through instruction groupingabstractDynamic instruction scheduling logic is quite complex and dissipates significant energy in microprocessors that support superscalar and out-of-order execution. We propose a novel microarchitectural technique to reduce the complexity and energy consumption of the dynamic instruction scheduling logic. The proposed method groups several instructions as a single issue unit and reduces the required number of ports and the size of the structure for dispatch, wakeup, select, and issue. The present paper describes the microarchitecture mechanisms and shows evaluation results for energy savings and performance. These results reveal that the proposed technique can greatly reduce energy with almost no performance degradation, compared to the conventional dynamic instruction scheduling logic. Hiroshi Sasaki 0001, Masaaki Kondo, Hiroshi Nakamura |
ISLPED | 3 |
| 2006 | A System Assisting Acquisition of Japanese Expressions Through Read-Write-Hear-Speaking and Comparing Between Use Cases of Relevant Expressions
Kohji Itoh, Hiroshi Nakamura, Shunsuke Unno, Jun'ichi Kakegawa |
KES (2) | 2 |
| 2005 | Secret sequence comparison on public grid computing resourcesabstractOnce a new gene has been sequenced, it must be verified whether or not it is similar to previously sequenced genes. In many cases, the organization that sequenced a potentially novel gene needs to keep the sequence itself in confidence. However, to compare the potentially novel sequence with known sequences, it must either be sent as a query to public databases, or these databases must be downloaded onto a local computer. In both cases, the potentially new sequence is exposed to the public. In this work, we propose a novel method to compare sequences without any exact sequence information leaks to the public. This method is based on our previous proposed method to find unique sequences on grid computing environments, which is well-parallelized in reasonable performance. In order to keep the exact sequence information in confidence, this method samples intervals (subsequences) from a sequence, and these intervals are hashed. Any key cryptosystem is not used. The hashed data are open to the public to verify the novelty of the sequence. The experimental results for 19797 h.sapiens genes show that the parallel implementation of this method performs reasonably well in terms of speed and memory usage. In this paper, the implementation on the world-wide testbeds of European Data Grid (EDG) and its results are described. Ken-ichi Kurata, Hiroshi Nakamura, Vincent Breton |
CCGRID | 2 |
| 2005 | A Small, Fast and Low-Power Register File by Bit-PartitioningabstractA large multi-ported register file is indispensable for exploiting instruction level parallelism (ILP) in today's dynamically scheduled superscalar processors. The number of ports and the size of the register file must be enlarged as the issue width and instruction window size increase. However, a larger register file causes longer access delays and more power consumption. To tackle these problems, we propose bit-partitioned register file which reduces the area, access time, and energy consumption of the register file. The proposed method relies on the fact that many operands do not need the full-bit width (typically a 32-bit or 64-bit width) of a register entry. Because the effective bit-width of most register operands is narrower than the full-bit width of a register entry, the upper bits of the register entries assigned to such narrow-width operands are useless. Thus, we propose to use of these useless upper bits for other operands by partitioning the register entries. In this paper, we show the mechanism of the proposed register file and evaluate its performance and power consumption. The evaluation results reveal that the proposed register file achieves higher instruction per cycle (IPC) in a smaller physical area, and consequently with shorter access time and less power consumption. Masaaki Kondo, Hiroshi Nakamura |
HPCA | 2 |
| 2005 | MegaProto: 1 TFlops/10kW Rack Is Feasible Even with Only Commodity TechnologyabstractIn our research project "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we claim that a million-scale parallel system could be built with densely mounted low-power commodity processors. "MegaProto" is a proof-of-concept low-power and highperformance cluster build only with commodity components to implement this claim. A one-rack system is composed of 32 motherboard "cluster units" of 1 U-height and commodity switches to interconnect them mutually as well as with other racks. Each cluster unit houses 16 low-power dollarbill- sized commodity PC-architecture daughterboards, together with a high bandwidth, 2 Gbps per processor embedded switched network based on Gigabit Ethernet. The peak performance of a one-rack system is 0.48 TFlops for the first version and will improve to 1.02 TFlops in the second version through a processor/daughterboard upgrade. The system consumes about 10 kW or less per rack, resulting in 100 MFlops/W power efficiency with a power-aware intrarack network of 32 Gbps bisection bandwidth, while additional 2.4 kW will boost this to sufficiently large 256 Gbps. Performance studies show that even the first version significantly outperforms a conventional high-end 1U server comprised of dual power-hungry processors in a majority of NPB programs. It is also investigated how the current automated DVS control could save power for the HPC parallel programs along with its limitation. Hiroshi Nakashima, Hiroshi Nakamura, Mitsuhisa Sato, Taisuke Boku, Satoshi Matsuoka, Daisuke Takahashi, Yoshihiko Hotta |
SC | 2 |
| 2005 | Multidimensional support vector machines for visualization of gene expression dataabstractMOTIVATION: Since DNA microarray experiments provide us with huge amount of gene expression data, they should be analyzed with statistical methods to extract the meanings of experimental results. Some dimensionality reduction methods such as Principal Component Analysis (PCA) are used to roughly visualize the distribution of high dimensional gene expression data. However, in the case of binary classification of gene expression data, PCA does not utilize class information when choosing axes. Thus clearly separable data in the original space may not be so in the reduced space used in PCA. RESULTS: For visualization and class prediction of gene expression data, we have developed a new SVM-based method called multidimensional SVMs, that generate multiple orthogonal axes. This method projects high dimensional data into lower dimensional space to exhibit properties of the data clearly and to visualize a distribution of the data roughly. Furthermore, the multiple axes can be used for class prediction. The basic properties of conventional SVMs are retained in our method: solutions of mathematical programming are sparse, and nonlinear classification is implemented implicitly through the use of kernel functions. The application of our method to the experimentally obtained gene expression datasets for patients' samples indicates that our algorithm is efficient and useful for visualization and class prediction. CONTACT: [email protected]. Daisuke Komura, Hiroshi Nakamura, Shuichi Tsutsumi, Hiroyuki Aburatani, Sigeo Ihara |
Bioinform. | 2 |
| 2004 | Secret sequence comparison in distributed computing environments by interval samplingabstractOnce a new gene has been sequenced, it must be verified whether or not it is similar to previously sequenced genes. In many cases, the organization that sequenced a potentially novel gene needs to keep the sequence itself in confidence. However, to compare the potentially novel sequence with known sequences, it must either be sent as a query to public databases, or these databases must be downloaded onto a local computer. In both cases, the potentially new sequence is exposed to the public. In this work, we propose a new method, called interval sampling, to compare sequences without leaking exact information about the new sequence. In order to keep the exact sequence information secret, this method samples intervals (subsequences) from a sequence, and these intervals are hashed. The hashed data are open to the public to verify the novelty of the sequence. We find that this method works well in parallel in a distributed computing environment, such as the Grid. The experimental results for 19797 h.sapiens genes and 25000 m.musculus genes show that the parallel implementation of this method performs reasonably well in terms of speed and memory usage. Ken-ichi Kurata, Vincent Breton, Hiroshi Nakamura |
CIBCB | 3 |
| 2004 | Skewed Checkpointing for Tolerating Multi-Node FailuresabstractLarge cluster systems have become widely utilized because they achieve a good performance/cost ratio especially in high performance computing. Although these cluster systems are distributed memory systems, coordinated checkpointing is a promising way to maintain high availability because the computing nodes are tightly connected to one another. However, as the number of computing nodes gets larger, the probability of multi-node failures increases. To tolerate multi-node failures, a large degree of redundancy is required in checkpointing, but this leads to performance degradation. Thus, we propose a new coordinated checkpointing called skewed checkpointing. In this method, checkpointing is skewed every time. Although each checkpointing itself contains only one degree of redundancy, this skewed checkpointing ensures /spl lfloor/log/sub 2/N/spl rfloor/ degrees of redundancy when the number of nodes is N. In this paper, we present the proposed method and an analysis of the performance overhead. Then, this method is applied to a cluster system and compared with other conventional checkpointing schemes. The results reveal the superiority of our method, especially for large cluster systems. Hiroshi Nakamura, Takuro Hayashida, Masaaki Kondo, Yuya Tajima, Masashi Imai, Takashi Nanya |
SRDS | 1 |
| 2004 | Grid as a bioinformatic tool
Nicolas Jacq, Christophe Blanchet, Christophe Combet, Emmanuel Cornillot, Laurent Duret, Ken-ichi Kurata, Hiroshi Nakamura, T. Silvestre, Vincent Breton |
Parallel Comput. | 7 |
| 2003 | Performance optimization of synchronous control units for datapaths with variable delay arithmetic unitsabstractNowadays, variable delay arithmetic units have been used for implementing a datapath of a target system in pursuit of performance improvement. However, adoption of variable delay arithmetic units requires modification of a typical synchronous control unit design methodology. A telescopic arithmetic unit based methodology is one of representative methodologies to design synchronous control units for variable delay datapaths. In this paper, we propose two optimization methods for it. Proposed optimization techniques will be analyzed in order to show their performance improvement effects explicitly. Euiseok Kim, Dong-Ik Lee, Hiroshi Saito, Hiroshi Nakamura, Jeong-Gun Lee, Takashi Nanya |
ASP-DAC | 4 |
| 2003 | Logic optimization for asynchronous speed independent controllers using transduction methodabstractAsynchronous speed independent (Sl) circuits based on an unbounded gate delay model often suffer from high area penalty. It happens due to the lack of efficient global optimization. This paper presents a boolean optimization method based on tranduction method to optimize asynchronous Sl circuits while preserving hazard-freeness. Hiroshi Saito, Hiroshi Nakamura, Takashi Nanya |
ASP-DAC | 2 |
| 2003 | A Method to Find Uniq e Sequences on Distrib ted Genomic DatabasesabstractThanks to the development of genetic engineering, various kinds of genomic information are being unveiled. Hence, it becomes feasible to analyze the entire genomic information all at once. On the other hand, the quantity of the genomic information stocked on databases is increasing day after day. In order to process the whole information, we have to develop an effective method to deal with lots of data. Therefore, it is indispensable not only to make an effective and rapid algorithm but also to use high-speed computer resource so as to analyze the biological information. For this purpose, as one of the most promised computing environments, the grid computing architecture has appeared recently. The European Data Grid (EDG) is one of the data-oriented grid computing environments [11]. In the field of bioinformatics, it is important to find unique sequences to succeed in molecular biological experiments [6]. Once unique sequences have been found they can be useful for target specific probes/primers design, gene sequence comparison and so on. In this paper, we propose a method to discover unique sequences from among genomic databases located in a distributed environment. Next, we implement this method upon the European Data Grid and show the calculation results for E. coli genomes. Ken-ichi Kurata, Vincent Breton, Hiroshi Nakamura |
CCGRID | 3 |
| 2003 | Distributed Synchronous Control Units for Dataflow Graphs under Allocation of Telescopic Arithmetic Units
Euiseok Kim, Hiroshi Saito, Jeong-Gun Lee, Dong-Ik Lee, Hiroshi Nakamura, Takashi Nanya |
DATE | 5 |
| 2002 | Formal Verification of a Pipelined Processor with New MemoryabstractRecently, model checkers have become commercially available. To investigate their ability, Solidify is selected as the representative of them and applied to a verification of a new processor. The processor adopts new memory hierarchy and new instructions. Its instruction issue is pipelined and in-order. Our experiment reveals that Solidify can verify the processor but drastic abstraction is indispensable for successful verification. The experimental results also suggest that it is quite hard to verify more complex out-of-order issue processors without very drastic and efficient abstraction. Through the experience, we also recognize the benefit of fully automatic verification. However, we suffer from the invariant problems. Experience is still important for this problem. Hiroshi Nakamura, Takanori Arai |
PRDC | 1 |
| 2000 | SCIMA: Software Controlled Integrated Memory Architecture for High Performance ComputingabstractProcessor performance has been improved due to clock acceleration and ILP extraction techniques. Performance of main memory, however, has not been improved so much. The performance gap between processor and memory will be growing further in the future. This is very serious problem in high performance computing because effective performance is limited by memory ability in most cases. In order to overcome this problem, we propose a new VLSI architecture called SCIMA which integrates software controllable memory into a processor chip. Most of data access is regular in high performance computing. The software controllable memory is more suitable for making good use of the regularity than conventional cache. This paper presents its architecture and performance evaluation. The evaluation results reveal the superiority of SCIMA compared with conventional cache-based architecture. Masaaki Kondo, Hideki Okawara, Hiroshi Nakamura, Taisuke Boku |
ICCD | 3 |
| 1999 | Performance of lattice QCD programs on CP-PACS
Sinya Aoki, R. Burkhalter, Kazuyuki Kanaya, T. Yoshié, Taisuke Boku, Hiroshi Nakamura, Yoshiyuki Yamashita |
Parallel Comput. | 6 |
| 1999 | CP-PACS: A massively parallel processor at the University of Tsukuba
Kisaburo Nakazawa, Hiroshi Nakamura, Taisuke Boku, Ikuo Nakata, Yoshiyuki Yamashita |
Parallel Comput. | 2 |
| 1999 | Augmenting Loop Tiling with Data Alignment for Improved Cache PerformanceabstractLoop blocking (tiling) is a well-known compiler optimization that helps improve cache performance by dividing the loop iteration space into smaller blocks (tiles); reuse of array elements within each tile is maximized by ensuring that the working set for the tile fits into the data cache. Padding is a data alignment technique that involves the insertion of dummy elements into a data structure for improving cache performance. In this work, we present DAT, a technique that augments loop tiling with data alignment, achieving improved efficiency (by ensuring that the cache is never under-utilized) as well as improved flexibility (by eliminating self-interference cache conflicts independent of the tile size). This results in a more stable and better cache performance than existing approaches, in addition to maximizing cache utilization, eliminating self-interference, and minimizing cross-interference conflicts. Further, while all previous efforts are targeted at programs characterized by the reuse of a single array, we also address the issue of minimizing conflict misses when several tiled arrays are involved. To validate our technique, we ran extensive experiments using both simulations as well as actual measurements on SUN Sparc5 and Sparc10 workstations. The results on benchmarks exhibiting varying memory access patterns demonstrate the effectiveness of our technique through consistently high hit ratios and improved performance across varying problem sizes. Preeti Ranjan Panda, Hiroshi Nakamura, Nikil Dutt, Alexandru Nicolau |
IEEE Trans. Computers | 2 |
| 1997 | Advanced processor design using hardware description language AIDLabstractIn order to design advanced processors in a short time, designers must simulate their designs and reflect the results to the designs at the very early stages. However, conventional hardware description languages (HDLs) do not have enough ability to describe designs easily and accurately at these stages. Thus, we have proposed a new HDL called AIDL (Architecture- and Implementation-level Description Language). In this paper, in order to evaluate the effectiveness of AIDL, we describe and compare three processors in both AIDL and VHDL descriptions. Takayuki Morimoto, Kazushi Saito, Hiroshi Nakamura, Taisuke Boku, Kisaburo Nakazawa |
ASP-DAC | 3 |
| 1997 | A Data Alignment Technique for Improving Cache PerformanceabstractWe address the problem of improving the data cache performance of numerical applications-specifically, those with blocked (or tiled) loops. We present DAT, a data alignment technique utilizing array-padding, to improve program performance through minimizing cache conflict misses. We describe algorithms for selecting tile sizes for maximizing data cache utilization, and computing pad sizes for eliminating self-interference conflicts in the chosen tile. We also present a generalization of the technique to handle applications with several tiled arrays. Our experimental results comparing our technique with previous published approaches on machines with different cache configurations show consistently good performance on several benchmark programs, for a variety of problem sizes. Preeti Ranjan Panda, Hiroshi Nakamura, Nikil Dutt, Alexandru Nicolau |
ICCD | 2 |
| 1997 | CP-PACS: A Massively Parallel Processor for Large Scale Scientific Calculations
Taisuke Boku, Ken'ichi Itakura, Hiroshi Nakamura, Kisaburo Nakazawa |
International Conference on Supercomputing | 3 |
| 1993 | A Scalar Architecture for Pseudo Vector Processing Based on Slide-Windowed RegistersabstractIn this paper, we present a new scalar architecture for high-speed vector processing. Without using cache memory, the proposed architecture tolerates main memory access latency by introducing slide-windowed floating-point registers with data preloading feature and pipelined memory. The architecture can hold upward compatibility with existing scalar architectures. In the new architecture, software can control the window structure. This is the advantage compared with our previous work of register-windows. Because of this advantage, registers are utilized more flexibly and computational efficiency is largely enhanced. Furthermore, this flexibility helps the compiler to generate efficient object codes easily. Hiroshi Nakamura, Taisuke Boku, Hideo Wada, Hiromitsu Imori, Ikuo Nakata, Yasuhiro Inagami, Kisaburo Nakazawa, Yoshiyuki Yamashita |
International Conference on Supercomputing | 1 |
| 1992 | Pseudo Vector Processor Based on Register-Windowed Superscalar PipelineabstractThe authors present a novel architecture for a high-speed pseudo vector processor based on a superscalar pipeline. Without using cache memory, the proposed architecture is able to overcome the penalty of memory access latency by introducing register windows with register preloading and pipelined memory. One outstanding feature of the proposed architecture is that it is upwardly compatible with existing scalar architectures. Performance evaluation of the proposed architecture using the Livermore Loop Kernels shows over 6 times higher performance than a usual superscalar processor and 1.2 times higher performance than a hypothetical extended model with a cache prefetching technique with a memory access latency of 20 CPU clock cycles. List vectors are also effectively handled in a similar architecture.> Kisaburo Nakazawa, Hiroshi Nakamura, Hiromitsu Imori, Shun Kawabe |
SC | 2 |
| 1990 | Practical design assistance at register transfer level using a data path verifierabstractA practical design assistance system at the register transfer level is proposed. The unique characteristic of this system is that users are allowed to modify a register-level design manually. At first, the designer gives an initial behavioral description and an initial structure of a data path to be designed. The initial structure is formed through the designers' intuition. The final design is obtained by modifying the initial design manually. Consistency between the data path and its behavioral specification is verified automatically. The verifier was implemented and applied to an ASIC chip design.> Hiroshi Nakamura, Yuji Kukimoto, Hidehiko Tanaka |
ICCD | 1 |
| 1987 | Multilevel QAM Modulation Techniques for Digital Microwave RadiosabstractA pilot carrier injection method is described together with feedback balance coding which reduces spectral power near the carrier. Robustness of carrier recovery using the pilot carrier injection method is theoretically estimated. The estimation suggests that recovered carrier SNR higher than 40 dB can be expected even under muitipath fading with notch depth of 45 dB located just at the carrier frequency. Signatures for multipath fading are estimated for a 64-QAM system with transversal equalizers as a countermeasure. Measured signatures agree reasonably well with the calculated ones. Dependences of signatures on modulation level, transversal equalizer tap number, and rolloff rate are also shown. Yoshimasa Daido, Sadao Takenaka, Eisuke Fukuda, Toshiaki Sakane, Hiroshi Nakamura |
IEEE J. Sel. Areas Commun. | 5 |
| 1986 | Theoretical Evalution of Signatures and CNR Penalties Caused by Modem Impairments in Multilevel QAM Digital Radio ModemsabstractA calculation method for symbol error probability is proposed for multilevel QAM systems. Carrier-to-noise ratio penalties are estimated by this method for seven typical impairment factors, including recovered carrier phase error, timing error, receive filter residual delay, etc. The validity of the method is verified by experiments on penalties caused by recovered carrier phase error and timing error. Based on the new method of analysis, signatures of a 64 QAM modem are estimated theoretically, using a pseudo-two-ray model of multipath fading. The dependence of signatures on the above-mentioned impairment factors is investigated in detail. A qualitative deduction of impairment factors from a measured signature is demonstrated. The measured signature agrees well with a calculated one in which reasonable values of the impairment factors are assumed. Yoshimasa Daido, Eisuke Fukuda, Yukio Takeda, Hiroshi Nakamura |
IEEE Trans. Commun. | 4 |
| 1984 | A New 4 GHz 90 Mbps Digital Radio System Using 64-QAM Modulation
Sadao Takenaka, Yukio Takeda, Toshiaki Sakane, Hiroshi Nakamura, N. Toyonaga |
ICC (2) | 4 |