EDBT 2026 Demo / reviewers in the wild / expert
Ryousei Takano
dblp:42/679
· DBLP profile ↗
32ranked-venue papers
6as first author
3since 2021 · last 2026
0000-0002-4341-3505ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 2Computer networks · 2Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The role of quantum computing in advancing scientific high-performance computing: A perspective from the ADAC institute
Gilles Buchs, Thomas L. Beck, Ryan S. Bennink, Daniel Claudino, Andrea Delgado 0002, Nur Aiman Fadel, Peter Groszkowski, Kathleen E. Hamilton, Travis S. Humble, Ang Li 0006, Phillip C. Lotshaw, Olli Mukkula, Ryousei Takano, In-Saeng Suh, Miwako Tsuji, Roel Van Beeumen, Ugo Varetto, Kazuya Yamazaki, Mikael P. Johansson |
Future Gener. Comput. Syst. | 14 |
| 2021 | An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural NetworksabstractDeep Neural Network (DNN) frameworks use distributed training to enable faster time to convergence and alleviate memory capacity limitations when training large models and/or using high dimension inputs. With the steady increase in datasets and model sizes, model/hybrid parallelism is deemed to have an important role in the future of distributed training of DNNs. We analyze the compute, communication, and memory requirements of Convolutional Neural Networks (CNNs) to understand the trade-offs between different parallelism approaches on performance and scalability. We leverage our model-driven analysis to be the basis for an oracle utility which can help in detecting the limitations and bottlenecks of different parallelism approaches at scale. We evaluate the oracle on six parallelization strategies, with four CNN models and multiple datasets (2D and 3D), on up to 1024 GPUs. The results demonstrate that the oracle has an average accuracy of about 86.74% when compared to empirical results, and as high as 97.57% for data parallelism. Albert Kahira, Truong Thao Nguyen, Leonardo Arturo Bautista-Gomez, Ryousei Takano, Rosa M. Badia, Mohamed Wahib |
HPDC | 4 |
| 2021 | Efficient MPI-AllReduce for large-scale deep learning on GPU-clustersabstractSummary Training models on large‐scale GPUs‐accelerated clusters are becoming a commonplace due to the increase in complexity and size in deep learning models. One of the main challenges for distributed training is the collective communication overhead for large message sizes: up to hundreds of MB. In this paper, we propose two hierarchical distributed memory multileader AllReduce algorithms optimized for GPU‐accelerated clusters (named lr_lr and lr_rab ), in which GPUs inside a computing node perform an intra‐node communication phase to gather and store results of local reduced values to designated GPUs (known as node leaders). Node leaders then keep a role as an inter‐node communicator. Each leader exchanges one part of reduced values to the leaders of the other nodes in parallel. Hence, we are capable of significantly reducing the time for injecting data into the inter‐node network. We also overlap the inter‐node and intra‐node communication by implementing our proposal in a pipelined manner. We evaluate those algorithms on the discrete‐event simulation Simgrid. We show that our algorithms, lr_lr and lr_rab , can cut down the execution time of an AllReduce microbenchmark that uses the logical ring algorithm ( lr ) by up to 45% and 51%, respectively. With the pipelined implementation, our lr_lr_pipe achieves 15% performance improvement when compared with lr_lr . In addition, the simulation result also projects power savings for the network devices of up to 23% and 32%. Truong Thao Nguyen, Mohamed Wahib, Ryousei Takano |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | Effect of an Incentive Implementation for Specifying Accurate Walltime in Job SchedulingabstractBackfill is a widely adopted scheduling technique in shared large scale systems. Accurate estimates of walltime of jobs benefit both users and operators of such systems because backfill uses the estimated walltime for scheduling decisions. However, reports on the accuracy analyses have shown that the accuracy is very low, which causes low utilization and long wait time. To overcome this situation, we propose to implement incentives for users to request accurate walltime in scheduling policy. We introduce a measure, named WRSA (Walltime Request Specification Accuracy), which represents the accuracy of requested walltime of each user and propose WRSA-aware backfill where jobs submitted by users with high WRSA are prioritized in scheduling. Through simulation using synthetic and real workloads, we confirm that utilization is improved up to 30% and the incentive for specifying accurate walltime is also improved against existing methods. Shin'ichiro Takizawa, Ryousei Takano |
HPC Asia | 2 |
| 2020 | Massively Parallel Causal Inference of Whole Brain Dynamics at Single Neuron ResolutionabstractEmpirical Dynamic Modeling (EDM) is a nonlinear time series causal inference framework. The latest implementation of EDM, cppEDM, has only been used for small datasets due to computational cost. With the growth of data collection capabilities, there is a great need to identify causal relationships in large datasets. We present mpEDM, a parallel distributed implementation of EDM optimized for modern GPU-centric supercomputers. We improve the original algorithm to reduce redundant computation and optimize the implementation to fully utilize hardware resources such as GPUs and SIMD units. As a use case, we run mpEDM on AI Bridging Cloud Infrastructure (ABCI) using datasets of an entire animal brain sampled at single neuron resolution to identify dynamical causation patterns across the brain. mpEDM is 1,530× faster than cppEDM and a dataset containing 101,729 neuron was analyzed in 199 seconds on 512 nodes. This is the largest EDM causal inference achieved to date. Wassapon Watanakeesuntorn, Keichi Takahashi, Kohei Ichikawa, Joseph Park, George Sugihara, Ryousei Takano, Jason H. Haga, Gerald M. Pao |
ICPADS | 6 |
| 2020 | Retraining Quantized Neural Network Models with Unlabeled DataabstractRunning neural network models on edge devices is attracting much attention by neural network researchers since edge computing technology is becoming more powerful than ever. However, deploying large neural network models on edge devices is challenging due to the limitation in available computing resources and storage space. Therefore, model compression techniques have been recently studied to reduce the model size and fit models on resource-limited edge devices. Compressing neural network models reduces the size of a model, but also degrades the accuracy of the model since it reduces the precision of weights in the model. Consequently, a retraining method is required to recover the accuracy of compressed models. Most existing retraining methods require the original labeled training datasets to retrain the models, but labeling is a time-consuming process. In particular, we cannot always access the original labeled datasets because of privacy policies and license limitations. In this paper, we propose a method to retrain a compressed neural network model with an unlabeled dataset that is different from the original labeled dataset. We compress the neural network model using quantization to decrease the size of the model. Subsequently, the compressed model is retrained by our proposed retraining method without using a labeled dataset to recover the accuracy of the model. We compared the proposed retraining method against the conventional retraining. The proposed method reduced the size of VGG-16 and ResNet-50 by 81.10% and 52.45%, respectively without significant accuracy loss. In addition, our proposed retraining method is clearly faster than the conventional retraining method. Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Chawanat Nakasan, Hidemoto Nakada, Ryousei Takano, Hajimu Iida |
IJCNN | 6 |
| 2020 | Scaling distributed deep learning workloads beyond the memory capacity with KARMAabstractThe dedicated memory of hardware accelerators can be insufficient to store all weights and/or intermediate states of large deep learning models. Although model parallelism is a viable approach to reduce the memory pressure issue, significant modification of the source code and considerations for algorithms are required. An alternative solution is to use out-of-core methods instead of, or in addition to, data parallelism. We propose a performance model based on the concurrency analysis of out-of-core training behavior, and derive a strategy that combines layer swapping and redundant recomputing. We achieve an average of 1. 52x speedup in six different models over the state-of-the-art out-of-core methods. We also introduce the first method to solve the challenging problem of out-of-core multi-node training by carefully pipelining gradient exchanges and performing the parameter updates on the host. Our data parallel out-of-core solution can outperform complex hybrid model parallelism in training large models, e.g. Megatron-LM and Turning-NLG. Mohamed Wahib, Truong Thao Nguyen, Aleksandr Drozd, Jens Domke, Lingqi Zhang 0001, Ryousei Takano, Satoshi Matsuoka |
SC | 7 |
| 2019 | Optimizing Weight Value Quantization for CNN InferenceabstractThe size and complexity of CNN models are increasing and as a result they are requiring more computational and memory resources to be used effectively. Use of a lower bit width numerical representation such as binary, ternary or several bit width has been studied extensively so as to reduce the required resources. However, the representation capability of such extremely low bit width is not always sufficient and the accuracy obtained for some CNN models and data is low. There are some prior studies that use moderate lower bit width with well-known numerical representations such as fixed point or logarithmic representation. It is not apparent, however, whether those representations are optimal for maintaining high accuracy. In this paper, we investigated the numerical quantization from the ground up, and introduced a novel "Variable Bin-size Quantization (VBQ)" representation in which quantization bin boundaries are optimized to obtain maximum accuracy for each CNN model. A genetic algorithm was employed to optimize the bin boundaries of VBQ. Additionally, since the appropriate bit width to obtain sufficient accuracy cannot be determined in advance, we attempted to use the parameters obtained by a training process using higher precision representation (FP32), and used quantization in inference only. This reduced the required large computational resource cost for training. During the process of tuning VBQ boundaries using a genetic algorithm, we discovered that the optimal distribution of bins can be approximated by an equation with two parameters. We then used simulated annealing for finding the optimal parameters of the equation for AlexNet and VGG16. As a result, AlexNet and VGG16 with our 4-bit quantization achieved top-5 accuracy at 74.8% and 86.3% respectively, which were comparable to 76.3% and 88.1% obtained by FP32. Thus, VBQ combined with the approximate equation and the simulated annealing scheme can achieve similar levels of accuracy with less resources and reduced computational cost compared to other current approaches. Wakana Nogami, Tsutomu Ikegami, Shin-ichi O'Uchi, Ryousei Takano, Tomohiro Kudoh |
IJCNN | 4 |
| 2019 | Topology-aware Sparse Allreduce for Large-scale Deep LearningabstractData parallelism is the dominant method used to scale-up deep learning (DL) training across multiple compute nodes. Collective communication of the local gradients between nodes is a critical bottleneck due to the significant increase in complexity and size of DL models. Researchers cope with this problem by one of the following solutions: a) optimizing the collective communication algorithm to account for the underlying network topology (topology-aware), and b) reducing the amount of transferred data. In the latter approach, sparse communication techniques communicate only the essential data, which helps to significantly cut down the communication cost. However, the diversity of message sizes, and unknown overlap of indicessets among compute nodes (i.e. irregular size), restricts their communication into the many-to-many scheme, which is ineffective in large-scale implementations. In this paper, we present allreduce algorithms that can exploit both sparse-communication and topology-aware techniques by mixing heterogeneous data formats (dense and sparse) with a trivial cost of computation. Truong Thao Nguyen, Mohamed Wahib, Ryousei Takano |
IPCCC | 3 |
| 2019 | A versatile software systolic execution model for GPU memory-bound kernelsabstractThis paper proposes a versatile high-performance execution model, inspired by systolic arrays, for memory-bound regular kernels running on CUDA-enabled GPUs. We formulate a systolic model that shifts partial sums by CUDA warp primitives for the computation. We also employ register files as a cache resource in order to operate the entire model efficiently. We demonstrate the effectiveness and versatility of the proposed model for a wide variety of stencil kernels that appear commonly in HPC, and also convolution kernels (increasingly important in deep learning workloads). Our algorithm outperforms the top reported state-of-the-art stencil implementations, including implementations with sophisticated temporal and spatial blocking techniques, on the two latest Nvidia architectures: Tesla V100 and P100. For 2D convolution of general filter sizes and shapes, our algorithm is on average 2.5× faster than Nvidia's NPP on V100 and P100 GPUs. Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
SC | 4 |
| 2019 | iFDK: a scalable framework for instant high-resolution image reconstructionabstractComputed Tomography (CT) is a widely used technology that requires compute-intense algorithms for image reconstruction. We propose a novel back-projection algorithm that reduces the projection computation cost to 1/6 of the standard algorithm. We also propose an efficient implementation that takes advantage of the heterogeneity of GPU-accelerated systems by overlapping the filtering and back-projection stages on CPUs and GPUs, respectively. Finally, we propose a distributed framework for high-resolution image reconstruction on state-of-the-art GPU-accelerated supercomputers. The framework relies on an elaborate interleave of MPI collective communication steps to achieve scalable communication. Evaluation on a single Tesla V100 GPU demonstrates that our back-projection kernel performs up to 1.6× faster than the standard FDK implementation. We also demonstrate the scalability and instantaneous CT capability of the distributed framework by using up to 2,048 V100 GPUs to solve 4K and 8K problems within 30 seconds and 2 minutes, respectively (including I/O). Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
SC | 4 |
| 2018 | Efficient Algorithms for the Summed Area Tables Primitive on GPUsabstractTwo-dimensional Summed Area Tables (SAT) is a fundamental primitive used in image processing and machine learning applications. We present a collection of optimization methods for computing SAT on CUDA-enabled GPUs. Conventional approaches rely on computing the prefix sum in one dimension in parallel, transposing the matrix, then computing the prefix sum for the other dimension in parallel. Additionally, conventional methods use the scratchpad memory as cache. We propose a collection of algorithms that are scalable with respect to problem size. We use the register cache technique instead of the scratchpad memory and also employ a naive serial scan on the thread level for computing the prefix sum for one of the dimensions. Using a novel transpose-in-registers method we increase the inter-thread parallelism and outperform conventional SAT implementations. In addition, we significantly reduce both the communication between threads and the number of arithmetic instructions. On an Nvidia Pascal P100 GPU and Volta V100, our evaluations demonstrate that our implementations outperform state of the art libraries and yield up to 2.3x and 3.2x speedup over OpenCV and Nvidia NPP libraries, respectively. Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
CLUSTER | 4 |
| 2018 | Image-Classifier Deep Convolutional Neural Network Training by 9-bit Dedicated Hardware to Realize Validation Accuracy and Energy Efficiency Superior to the Half Precision Floating Point FormatabstractWe propose a 9-bit floating point format for training image-classifier deep convolutional neural networks. The proposed floating point format has a 5-bit exponent, a 3-bit mantissa with the hidden most significant bit (MSB), and a sign bit. The 9-bit floating point format reduces not only the transistor count of the multiplier in the multipy-accumulate (MAC) unit, but also the data traffic for the forward and backward propagations and the weight update. Both of the reductions realize a power efficient training. To maintain the validation accuracy, the accumulator is implemented with an internal longer-bit-length floating point format while the multiplier accepts the 9-bit format. We examined this format in the training of the AlexNet and the ResNet-50 with the ILSVRC 2012 data set. The trained 9-bit AlexNet and ResNet-50 exhibited the validation accuracy superior to the 16-bit floating point format training by 1.2 % and 0.5 %, respectively. The transistor count in the 9-bit MAC unit is estimated to be reduced by 84% as compared to the 32-bit counterpart. Shin-ichi O'Uchi, Hiroshi Fuketa, Tsutomu Ikegami, Wakana Nogami, Takashi Matsukawa, Tomohiro Kudoh, Ryousei Takano |
ISCAS | 7 |
| 2017 | DEMU: A DPDK-based network latency emulatorabstractA network latency emulator allows IT architects to thoroughly investigate how network latencies impact workload performance. Software-based emulation tools have been widely used by researchers and engineers. It is possible to use commodity server computers for emulation and set up an emulation environment quickly without outstanding hardware cost. However, existing software-based tools built in the network stack of an operating system are not capable of supporting the bandwidth of today's standard interconnects (e.g., 10GbE) and emulating sub-milliseconds latencies likely caused by network virtualization in a datacenter. In this paper, we propose a network latency emulator (DEMU) supporting broad bandwidth traffic with sub-milliseconds accuracy, which is based on an emerging packet processing framework, DPDK. It avoids the overhead of the network stack by directly interacting with NIC hardware. Through experiments, we confirmed that DEMU can emulate latencies on the order of 10 μs for short-packet traffic at the line rate of 10GbE. The standard deviation of inserted delays was only 2-3 μs. This is a significant improvement from a network emulator built in the Linux Kernel (i.e., NetEm), which loses more than 50% of its packets for the same 10GbE traffic. For 1 Gbps traffic, the latency deviation of NetEm was approximately 20 μs, while that of our mechanism was 2 orders of magnitude smaller (i.e., only 0.3 μs). Shuhei Aketa, Takahiro Hirofuchi, Ryousei Takano |
LANMAN | 3 |
| 2017 | muMQ: A lightweight and scalable MQTT brokerabstractA message broker is an imperative component in IoT systems, and it works as a gateway between IoT devices and application platforms. With the growth of IoT devices today, these systems can easily overwhelm message brokers unless the software can fully utilize hardware resources such as multi-core facility. This paper presents muMQ, a high-performance MQTT broker running on Commercial-Off-The-Shelf hardware. It tackles the challenge to improve the performance of message brokering on a single machine by efficiently utilizing multi-core CPUs. First, muMQ exploits an event-driven I/O mechanism for multi-core scalability. Each CPU core equally handles dispatched TCP connections and locally processes MQTT logic. Second, muMQ adopts a user-level TCP/IP stack, mTCP with DPDK, to avoid the overhead of the in-kernel TCP/IP stack, including system call overhead and resource contention. We evaluate the effectiveness of our approach through experiments. The results show that muMQ can handle 512K or greater long-lived subscribers with no message loss; muMQ achieves a publish messaging rate at 930K messages per second, which is 5.38 times faster than an existing MQTT broker. We also confirm mTCP accelerates the performance by 1.8 times compared with muMQ using the in-kernel TCP/IP stack. Wiriyang Pipatsakulroj, Vasaka Visoottiviseth, Ryousei Takano |
LANMAN | 3 |
| 2016 | RAMinate: Hypervisor-based Virtualization for Hybrid Main Memory SystemsabstractIn the future, STT-MRAM will achieve larger capacity and comparable read/write performance, but incur orders of magnitude greater write energy than DRAM. To achieve large capacity as well as energy-efficiency, it is necessary to use both DRAM and STT-MRAM for the main memory of a computer. In this paper, we propose a hypervisor-based hybrid memory mechanism (RAMinate) that reduces write traffic to STT-MRAM by optimizing page locations between DRAM and STT-MRAM. In contrast to past studies, our mechanism works at the hypervisor level, not at the hardware or operating system level. It does not require any special program at the operating system level nor any design changes of the current memory controller at the hardware level. We developed a prototype of the proposed system by extending Qemu/KVM and conducted experiments with application benchmarks. We confirmed that our page replacement mechanism successfully worked for unmodified operating systems and dynamically diverted memory write traffic to DRAM. Our experiments also confirmed that our system successfully reduced write traffic to STT-MRAM by approximately 70% for tested workloads, which results in a 50% reduction in energy consumption in comparison to a DRAM-only system. Takahiro Hirofuchi, Ryousei Takano |
SoCC | 2 |
| 2016 | Flow-centric computing leveraged by photonic circuit switching for the post-moore eraabstractArtificial Intelligence (AI)-based big data analysis is emerging for broad application fields, including manufacturing, autonomous car, health care, agriculture, and so on. Data centers are under severe pressure to meet ever-increasing demands on clouds. On the other hand, the amount of data processing is bounded by the total power budget, where the reasonable number is up to 20 MW. Therefore, the energy efficiency becomes a critical issue both now and the future. A combination of special-purpose processing units communicating through huge bandwidth optical circuit switched network allows to dramatically improve the energy efficiency of big data processing. Based on this idea, we propose flow-centric computing, a software-defined data center architecture focusing on data flow processing. Compute, storage, and network resources are disaggregated and dynamically composes slices, i.e., data processing environments, based on workload-specific demands. We are conducting a feasibility study of the concept and developing system software technologies. Ryousei Takano, Tomohiro Kudoh |
NOCS | 1 |
| 2014 | Fast Live Migration with Small IO Performance Penalty by Exploiting SAN in ParallelabstractVirtualization techniques greatly benefit cloud computing. Live migration enables a datacenter to dynamically replace virtual machines (VMs) without disrupting services running on them. Efficient live migration is the key to improve the energy efficiency and resource utilization of a datacenter through dynamic placement of VMs. Recent studies have achieved efficient live migration by deleting the page cache of the guest OS to shrink the memory size of it before a migration. However, these studies do not solve the problem of IO performance penalty after a migration due to the loss of page cache. We propose an advanced memory transfer mechanism for live migration, which skips transferring the page cache to shorten total migration time while restoring it transparently from the guest OS via the SAN to prevent IO performance penalty. To start a migration, our mechanism collects the mapping information between page cache and disk blocks. During a migration, the source host skips transferring the page cache but transfers other memory content, while the destination host transfers the same data as the page cache from the disk blocks via the SAN. Experiments with web server and database workloads showed that our mechanism reduced total migration time with significantly small IO performance penalty. Soramichi Akiyama, Takahiro Hirofuchi, Ryousei Takano, Shinichi Honiden |
IEEE CLOUD | 3 |
| 2014 | Iris: An Inter-cloud Resource Integration System for Elastic Cloud Data CentersabstractThis paper proposes a new cloud computing service model, Hardware as a Service (HaaS), that is based on the idea of implementing ``elastic data centers'' that provide a data center administrator with resources located at different data centers as demand requires. To demonstrate the feasibility of the proposed model, we have developed what we call an Inter-cloud Resource Integration System (Iris) by using nested virtualization and OpenFlow technologies. Iris dynamically configures and provides a virtual infrastructure over inter-cloud resources, on which an IaaS cloud can run. Using Iris, we have confirmed an IaaS cloud can seamlessly extend and manage resources over multiple data centers. The experimental results on an emulated inter-cloud environment show that the overheads of the HaaS layer are acceptable when the network latency is less than 10 msec. We believe these results provide new insight to help establish inter-cloud computing. Ryousei Takano, Atsuko Takefusa, Hidemoto Nakada, Seiya Yanagita, Tomohiro Kudoh |
CLOSER | 1 |
| 2014 | Exploring the Performance Impact of Virtualization on an HPC CloudabstractThe feasibility of the cloud computing paradigm is examined from the High Performance Computing (HPC) viewpoint. The impact of virtualization is evaluated on our latest private cloud, the AIST Super Green Cloud, which provides elastic virtual clusters interconnected by Infini Band. Performance is measured by using typical HPC benchmark programs, both on physical and virtual cluster computing clusters. The results of the micro benchmarks indicate that the virtual clusters suffer from the scalability issue on almost all MPI collective functions. The relative performance gradually becomes worse as the number of nodes increases. On the other hand, the benchmarks based on actual applications, including LINPACK, OpenMX, and Graph 500, show that the virtualization overhead is about 5% even when the number of nodes increase to 128. This observation leads to our optimistic conclusions on the feasibility of the HPC Cloud. Nuttapong Chakthranont, Phonlawat Khunphet, Ryousei Takano, Tsutomu Ikegami |
CloudCom | 3 |
| 2013 | Fast Wide Area Live Migration with a Low Overhead through Page Cache TeleportationabstractLive migration of virtual machines over a wide area network has many use cases such as cross-data center load balancing, low carbon virtual private clouds, and disaster recovery of IT systems. An efficient wide area live migration method is required because cross-data center connections have a narrow bandwidth. Page cache occupies a large portion of the memory of a Virtual Machine (VM) when it executes data-intensive workloads. We propose a new live migration technique, page cache teleportation, which reduces the total migration time of wide area live migration and has a low overhead. It detects the restorable page cache in the guest memory that has the same contents as the corresponding disk blocks. The restorable page cache is not transferred via the WAN but is restored from the disk image before the VM resumes. In this way, the IO performance degradation reduces after the migration. Evaluations show that page cache teleportation reduces the total migration time of wide area live migration and has a lower performance overhead than existing approaches. Soramichi Akiyama, Takahiro Hirofuchi, Ryousei Takano, Shinichi Honiden |
CCGRID | 3 |
| 2012 | MiyakoDori: A Memory Reusing Mechanism for Dynamic VM ConsolidationabstractIn Infrastructure-as-a-Service datacenters, the placement of Virtual Machines (VMs) on physical hosts are dynamically optimized in response to resource utilization of the hosts. However, existing live migration techniques, used to move VMs between hosts, need to involve large data transfer and prevents dynamic consolidation systems from optimizing VM placements efficiently. In this paper, we propose a technique called “memory reusing” that reduces the amount of transferred memory of live migration. When a VM migrates to another host, the memory image of the VM is kept in the source host. When the VM migrates back to the original host later, the kept memory image will be “reused”, i.e. memory pages which are identical to the kept pages will not be transferred. We implemented a system named MiyakoDori that uses memory reusing in live migrations. Evaluations show that MiyakoDori significantly reduced the amount of transferred memory of live migrations and reduced 87% of unnecessary energy consumption when integrated with our dynamic VM consolidation system. Soramichi Akiyama, Takahiro Hirofuchi, Ryousei Takano, Shinichi Honiden |
IEEE CLOUD | 3 |
| 2012 | A distributed application execution system for an infrastructure with dynamically configured networksabstractWe have been developing a middleware suite called GridARS that enables co-allocation of computing and network resources from multiple administration sites. In such middleware, it is important to provide each user application with a slice which is a set of dynamically allocated resources distributed across sites. However, there are the following issues in constructing such a slice automatically: 1) multi-site administration heterogeneity, 2) dynamic determination of application configuration information, 3) distributed resource monitoring, and 4) asymmetric network reachability. We design and implement an application execution system that provides each application with a slice, that mimics a conventional computing cluster system over the dynamically allocated resources. From the demonstration of the proposed system on an emulated wide area network environment, we confirmed that: first, the proposed system can fully automate resource allocation, slice construction, application invocation, and resource monitoring, in coordination with GridARS. Second, the proposed system can setup a slice quickly, even if the allocated resources are widely distributed and their communication latencies are high. This is because the overhead for gathering and distributing contextualization information is small, and OS-level virtualization and stackable file system technologies accelerate the contextualization process at each node. Ryousei Takano, Hidemoto Nakada, Atsuko Takefusa, Tomohiro Kudoh |
CloudCom | 1 |
| 2012 | Cooperative VM migration for a virtualized HPC cluster with VMM-bypass I/O devicesabstractAn HPC cloud, a flexible and robust cloud computing service specially dedicated to high performance computing, is a promising future e-Science platform. In cloud computing, virtualization is widely used to achieve flexibility and security. Virtualization makes migration or checkpoint/restart of computing elements (virtual machines) easy, and such features are useful for realizing fault tolerance and server consolidations. However, in widely used virtualization schemes, I/O devices are also virtualized, and thus I/O performance is severely degraded. To cope with this problem, VMM-bypass I/O technologies, including PCI passthrough and SR-IOV, in which the I/O overhead can be significantly reduced, have been introduced. However, such VMM-bypass I/O technologies make it impossible to migrate or checkpoint/restart virtual machines, since virtual machines are directly attached to hardware devices. This paper proposes a novel and practical mechanism, called Symbiotic Virtualization (SymVirt), for enabling migration and checkpoint/restart on a virtualized cluster with VMM-bypass I/O devices, without the virtualization overhead during normal operations. SymVirt allows a VMM to cooperate with a message passing layer on the guest OS, then it realizes VM-level migration and checkpoint/restart by using a combination of a PCI hotplug and coordination of distributed VMMs. We have implemented the proposed mechanism on top of QEMU/KVM and the Open MPI system. All PCI devices, including Infiniband and Myrinet, are supported without implementing specific para-virtualized drivers; and it is not necessary to modify either of the MPI runtime and applications. Using the proposed mechanism, we demonstrate reactive and proactive FT mechanisms on a virtualized Infiniband cluster. We have confirmed the effectiveness using both a memory intensive micro benchmark and the NAS parallel benchmark. Moreover, we also show that postcopy live migration enables us to reduce the down time of an application as the memory footprint increases. Ryousei Takano, Hidemoto Nakada, Takahiro Hirofuchi, Yoshio Tanaka, Tomohiro Kudoh |
eScience | 1 |
| 2012 | On the use of virtualization technologies to support uninterrupted IT services: A case study with lessons learned from the Great East Japan EarthquakeabstractVirtualized IT infrastructures combined with virtual machine migration technologies have a potential to support IT services that are resilient to partial physical infrastructure failures caused by extreme events. This paper experimentally evaluates the migration of multiple VMs across long geographical distances - an activity that is required to move virtualized IT systems from a disaster site to a safe location. Taking into account the resource availability parameters observed after the Great East Japan Earthquake, experimental results show that if (1) service downtime in the order of minutes is acceptable, (2) VMs can be kept with small storage footprint, and (3) power and network are available for tens of minutes, it is possible to migrate tens of VMs from damaged sites to a very distant stable location. Maurício O. Tsugawa, Renato J. O. Figueiredo, José A. B. Fortes, Takahiro Hirofuchi, Hidemoto Nakada, Ryousei Takano |
ICC | 6 |
| 2011 | GridARS: A Grid Advanced Resource Management System Framework for IntercloudabstractIntercloud is a promising technology for data intensive applications. However, an important issue for Intercloud applications is orchestration of various virtualized and performance-assured resources, not only computers, but also network and storage, provided from multiple domains. We have been developing an advance reservation-based resource management framework, called Grid ARS, which can integrate heterogeneous resources and construct a performance-assured virtual infrastructure over Intercloud environment. Grid ARS provides four services that address resource management, resource allocation planning, provisioning and monitoring of the constructed virtual infrastructure. Grid ARS has been developed using common Web services technologies and standards. In this paper, we present overview of Grid ARS and its service components and describe Grid ARS demonstration challenges, demonstration at GLIF2010 and SC10 and OGF NSI interoperation in 2011. Atsuko Takefusa, Hidemoto Nakada, Ryousei Takano, Tomohiro Kudoh, Yoshio Tanaka |
CloudCom | 3 |
| 2010 | SSS: An Implementation of Key-Value Store Based MapReduce FrameworkabstractMapReduce has been very successful in implementing large-scale data-intensive applications. Because of its simple programming model, MapReduce has also begun being utilized as a programming tool for more general distributed and parallel applications, e.g., HPC applications. However, its applicability is limited due to relatively inefficient runtime performance and hence insufficient support for flexible workflow. In particular, the performance problem is not negligible in iterative MapReduce applications. On the other hand, today, HPC community is going to be able to utilize very fast and energy-efficient Solid State Drives (SSDs) with 10 Gbit/sec-class read/write performance. This fact leads us to the possibility to develop "High-Performance MapReduce'', so called. From this perspective, we have been developing a new MapReduce framework called "SSS'' based on distributed key-value store (KVS). In this paper, we first discuss the limitations of existing MapReduce implementations and present the design and implementation of SSS. Although our implementation of SSS is still in a prototype stage, we conduct two benchmarks for comparing the performance of SSS and Hadoop. The results indicate that SSS performs 1-10 times faster than Hadoop. Hirotaka Ogawa, Hidemoto Nakada, Ryousei Takano, Tomohiro Kudoh |
CloudCom | 3 |
| 2008 | High Performance Relay Mechanism for MPI Communication Libraries Run on Multiple Private IP Address ClustersabstractWe have been developing a Grid-enabled MPI communication library called GridMPI, which is designed to run on multiple clusters connected to a wide-area network. Some of these clusters may use private IP addresses. Therefore, some mechanism to enable communication between private IP address clusters is required. Such a mechanism should be widely adoptable, and should provide high communication performance. In this paper, we propose a message relay mechanism to support private IP address clusters in the manner of the Interoperable MPI (IMPI) standard. Therefore, any MPI implementations which follow the IMPI standard can communicate with the relay. Furthermore, we also propose a trunking method in which multiple pairs of relay nodes simultaneously communicate between clusters to improve the available communication bandwidth. While the relay mechanism introduces an one-way latency of about 25 musec, the extra overhead is negligible, since the communication latency through a wide area network is a few hundred times as large as this. By using trunking, the inter-cluster communication bandwidth can improve as the number of trunks increases. We confirmed the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. When relay nodes with 1 Gbps NICs are used, the performance of most of the NAS Parallel Benchmarks improved proportional to the number of trunks. Especially, using 8 trunks, FT and IS are 4.4 and 3.4 times faster, respectively, compared with the single trunk case. The results showed that the proposed method is effective for running MPI programs over high bandwidth-delay product networks. Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa, Yasufumi Yoshizawa |
CCGRID | 1 |
| 2007 | Effects of packet pacing for MPI programs in a Grid environmentabstractImproving the performance of TCP communication is the key to the successful deployment of MPI programs in a Grid environment in which multiple clusters are connected through high performance dedicated networks. To efficiently utilize the inter-cluster bandwidth, a traffic control mechanism is required so as not to allow the aggregate transmission bandwidth to exceed the inter-cluster bandwidth when multiple nodes communicate at one time. In this paper, we propose a traffic control method for MPI programs, in which an application or the MPI runtime controls the transmission rate based on the communication pattern by using certain MPI attributes. Packet pacing is used at each node preventing microscopic burst transmission to thus avoid congestion. We confirm the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. We show most of the NAS Parallel benchmarks improve the performance, since the proposed method reduces packet losses due to traffic congestion on the inter-cluster network. The results have indicated that it is feasible to connect multiple clusters and run large-scale scientific applications over distances up to 1000 kilometers, if an appropriate network is available. Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa |
CLUSTER | 1 |
| 2006 | Efficient MPI Collective Operations for Clusters in Long-and-Fast NetworksabstractSeveral MPI systems for grid environment, in which clusters are connected by wide-area networks, have been proposed. However, the algorithms of collective communication in such MPI systems assume relatively low bandwidth wide-area networks, and they are not designed for the fast wide-area networks that are becoming available. On the other hand, for cluster MPI systems, a beast algorithm by van de Geijn et al. and an allreduce algorithm by Rabenseifner have been proposed, which are efficient in a high bisection bandwidth environment. We modify those algorithms so as to effectively utilize fast wide-area inter-cluster networks and to control the number of nodes which can transfer data simultaneously through wide-area networks to avoid congestion. We confirmed the effectiveness of the modified algorithms by experiments using a 10 Gbps emulated WAN environment. The environment consists of two clusters, where each cluster consists of nodes with 1 Gbps Ethernet links and a switch with a 10 Gbps upper link. The two clusters are connected through a 10 Gbps WAN emulator which can insert latency. In a 10 millisecond latency environment, when the message size is 32 MB, the proposed beast and allreduce are 1.6 and 3.2 times faster, respectively, than the algorithms used in existing MPI systems for grid environment Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa |
CLUSTER | 4 |
| 2005 | TCP Adaptation for MPI on Long-and-Fat NetworksabstractTypical MPI applications work in phases of computation and communication, and messages are exchanged in relatively small chunks. This behavior is not optimal for TCP because TCP is designed only to handle a contiguous flow of messages efficiently. This behavior anomaly is well-known, but fixes are not integrated into today's TCP implementations, even though performance is seriously degraded, especially for MPI applications. This paper proposes three improvements in the Linux TCP stack: i.e., pacing at start-up, reducing Retransmit-Timeout time, and TCP parameter switching at the transition of computation phases in an MPI application. Evaluation of these improvements using the NAS parallel benchmarks shows that the BT, CG, IS, and SP benchmarks achieved 10 to 30 percent improvements. On the other hand, the FT and MG benchmarks showed no improvement because they have the steady communication that TCP assumes, and the LU benchmark became slightly worse because it has very little communication Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa |
CLUSTER | 4 |
| 2004 | GNET-1: gigabit Ethernet network testbedabstractGNET-1 is a fully programmable network testbed. It provides functions such as wide area network emulation, network instrumentation, traffic shaping, and traffic generation at gigabit Ethernet wire speeds by programming the core FPGA. GNET-1 is a powerful tool for developing network-aware grid software. It is also a network monitoring and traffic-shaping tool that provides high-performance communication over wide area networks. This work describes several sample uses of GNET-1 and presents its architecture. Yuetsu Kodama, Tomohiro Kudoh, Ryousei Takano, Hitoshi Sato, Osamu Tatebe, Satoshi Sekiguchi |
CLUSTER | 3 |