Tatsuhiro Chiba

dblp:23/6754 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-9458-899XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 8 since 2021Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Speeding up Model Loading with Fastsafetensors
abstract
The rapid increases in model parameter sizes introduces new challenges in pre-trained model loading. Currently, machine learning code often deserializes each parameter as a tensor object in host memory before copying it to device memory. We found that this approach underutilized storage throughput and significantly slowed down loading large models with a widely-used model file formats, safetensors. In this work, we present fastsafetensors, a Python library designed to optimize the deserialization of tensors in safetensors files. Our approach first copies groups of on-disk parameters to device memory, where they are directly instantiated as tensor objects. This design enables further optimization in low-level I/O and high-level tensor preprocessing, including parallelized copying, peer-to-peer DMA, and GPU offloading. Experimental results show performance improvements of 4.8x to 7.5x in loading models such as Llama (7, 13, and 70 billion parameters), Falcon (40 billion parameters), and the Bloom (176 billion parameters).
Takeshi Yoshimura, Tatsuhiro Chiba, Manish Sethi, Daniel G. Waddington, Swaminathan Sundararaman
CLOUD2
2024 Process-Based Efficient Power Level Exporter
abstract
In this paper, we present the Kepler framework, designed to address the critical need for precise power and energy measurement in on-prem cloud-native, containerized environments, with a specific focus on processes, containers, and Kubernetes pods. The framework aims to support other tools in making informed decisions regarding provisioning, scheduling, and energy-optimization in cloud environments. Our approach involves leveraging the Kepler framework to create power models using Hardware Counters (HC), and real- time system power metrics from hardware sensors like x86 Running Average Power Limit (RAPL). Unlike previous methods that create and validate power models using aggregated system metrics, we propose a versatile process-level power model trained with per-process metrics. Those metrics are collected via a series of experiments in a controlled environment, measuring the incremental power consumption of processes under different scenarios. The collected data is then utilized to create a power model to be used in a shared cloud environment, and to validate the created power models using different set of input metrics. Our results show a significant improvement in the model accuracy compared to prior works, when incorporating per- process metrics and real-time system power metrics into the power estimation process. For instance, using the simplest power model, which is based on CPU utilization ratio, resulted in a Sum of Squared Error (SSE) of 75. In contrast, a power model created using aggregated system metrics, as the related works, had an SSE of 175 without real-time power metrics, and 5.6 with our proposed model refinement by normalizing the model results with the real-time system power metrics. On the other hand, training the power model with per-process metrics from controlled experiments yielded an SSE as low as 1.68 using real- time system power metrics, representing a 70% improvement in model accuracy compared to using aggregated system metrics, and an SSE 8.7 without power metrics, representing a 95% improvement in model accuracy. Furthermore, the results show that Kepler has a notable lower overhead by utilizing extended Berkeley Packet Filter (eBPF) for HC collection than alternative methods.
Marcelo Amaral, Tatsuhiro Chiba, Rina Nakazawa, Sunyanan Choochotkaew, Tamar Eilam
CLOUD3
2024 Best-Effort Power Model Serving for Energy Quantification of Cloud Instances
abstract
Quantifying energy consumption is a fundamental element of green computing. Power models trained by resource utilization allow quantifying the energy number and enable energy-efficient resource management systems without raising the concerns of complexity, cost, and security. However, energy consuming behavior on different machines varies by several factors. In this paper, we address the challenges of power modeling for cloud instances where information about these factors is obscured or unseen in the training set, and propose a best-effort method to train and serve a power model as precise as possible by leveraging a large, industry-standard power database. The proposed method prioritizes the modeling precision, and offers similarity and uncertainty indicators to elucidate the confidence level when serving an unseen instance. The results have demonstrated feasibility and precision of the proposed method against comparable approaches.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Marcelo Amaral, Rina Nakazawa, Scott Trent, UmaMaheswari Devi, Tamar Eilam
MASCOTS2
2023 Kepler: A Framework to Calculate the Energy Consumption of Containerized Applications
abstract
Energy accounting is crucial in data centers for optimizing power provisioning, capping, and tuning. This paper introduces the Kepler framework, which estimates power consumption at the process, container, and Kubernetes pod levels. Kepler offers a set of power models applicable to various architectures and metrics. In this study, we propose a generic power model that utilizes hardware counters (HC) and realtime system power metrics (e.g., running average power limit (RAPL)) as independent variables in a regression model. Unlike previous approaches that rely on aggregate power consumption, our methodology measures individual process power consumption to train the power model. We provide step-by-step instructions to measure process power consumption in a controlled environment, considering the activation constant and load-dependent dynamic power consumption in different executions. By following the Greenhouse Gas (GHG) Protocol, our approach ensures the fair distribution of constant power among the user's processes. The results demonstrate significantly improved accuracy with a mean squared error (MSE) as low as 0.010 for the proposed method, compared with an MSE of 0.16 for a simple ratio approach and 0.92 when training the model using aggregated workload power.
Marcelo Amaral, Tatsuhiro Chiba, Rina Nakazawa, Sunyanan Choochotkaew, Tamar Eilam
CLOUD3
2023 Advancing Cloud Sustainability: A Versatile Framework for Container Power Model Training
abstract
Estimating power consumption in modern Cloud is important to account for the power consumed by each container. The challenge is that multiple customers are sharing the same hardware platform, where physical information is mostly obscured. In addition, there is the overhead in power consumption that the Cloud control plane induces. This paper addresses these challenges and introduces a pipeline framework for container power model training on the basis of available performance counters and other metrics. The proposed model utilizes machine learning techniques to predict the power consumed by the control plane and associated processes when running together with the user containers, and uses it for isolating the dynamic power consumed by the user-inducing workload. Applying the proposed power model does not require online power measurements, nor does it need machine information, or information on other tenants sharing the same machine. The results of cross-workload, cross-platform experiments demonstrated the higher accuracy of the model when predicting power consumption of unseen containers on unknown platforms, including on virtual machines.
Sunyanan Choochotkaew, Chen Wang 0039, Tatsuhiro Chiba, Marcelo Amaral, Tamar Eilam
MASCOTS4
2022 MicroLens: A Performance Analysis Framework for Microservices Using Hidden Metrics With BPF
abstract
Determining the root cause of performance regression for microservices is challenging. The topological cascading performance implications among microservices hide the source of the problem. Additionally, the lack of knowledge about application phases can potentially lead to false-positive critical service detection. Service resource utilization is an imperfect proxy for application performance, potentially leading to false positives. Therefore, in this work, we propose a new performance testing framework that leverages hidden Berkeley Packet Filter (BPF) kernel metrics to locate root causes of performance regression. The framework applies a systematic multi-level approach to analyze microservice performance without intrusive code instrumentation. First, the framework constructs an attributed graph with microservice requests, scores the services to identify the critical paths, and ranks the low-level metrics to highlight the root cause of performance regression. Through judiciously designed experiments, we evaluated the metric collection overhead, showing less than 18% more latency when the application is running across hosts and 9% within the same host. In addition, depending on the application, no overhead is experienced, while the state-of-the-art approach presented up to 1060% more latency. The microservice benchmark evaluation shows that MicroLens can successfully identify the set of root causes and that the causes vary when the application is running in different infrastructures.
Marcelo Amaral, Tatsuhiro Chiba, Scott Trent, Takeshi Yoshimura, Sunyanan Choochotkaew
CLOUD2
2022 Bypass Container Overlay Networks with Transparent BPF-driven Socket Replacement
abstract
Containerization on the cloud offers several crucial benefits. However, these benefits are negated by the effects of virtual network stack and address encapsulation, especially for workloads that require intense communication. Socket replacement is a promising approach to breach this wall without changing the underlay infrastructure by replacing a nested network stack with a simple host network stack. Current state-of-the-art approaches perform this replacement by preloading the overridden socket library in a containerized process. However, the preloading approach requires user effort to modify the deploying manifests and a compromised security policy configuration of privileged containers to access the host namespace. This paper introduces a new replacement framework where a secured control plane agent performs the replacement by utilizing low-overhead BPF kernel tracing technology. As a result, containers can obtain host-native network performance and neither modification nor escalated privileges are required for user containers. Experiments on multiple benchmarks including iPerf, MPI, memslap, and GROMACS have been conducted to confirm efficacy.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Scott Trent, Marcelo Amaral
CLOUD2
2022 AutoDECK: Automated Declarative Performance Evaluation and Tuning Framework on Kubernetes
abstract
Containerization and application variety bring many challenges in automating evaluations for performance tuning and comparison among infrastructure choices. Due to the tightly-coupled design of benchmarks and evaluation tools, the present automated tools on Kubernetes are limited to trivial microbenchmarks and cannot be extended to complex cloudnative architectures such as microservices and serverless, which are usually managed by customized operators for setting up workload dependencies. In this paper, we propose AutoDECK, a performance evaluation framework with a fully declarative manner. The proposed framework automates configuring, deploying, evaluating, summarizing, and visualizing the benchmarking workload. It seamlessly integrates mature Kubernetes-native systems and extends multiple functionalities such as tracking the image-build pipeline, and auto-tuning. We present five use cases of evaluations and analysis through various kinds of bench-marks including microbenchmarks and HPC/AI benchmarks. The evaluation results can also differentiate characteristics such as resource usage behavior and parallelism effectiveness between different clusters. Furthermore, the results demonstrate the benefit of integrating an auto-tuning feature in the proposed framework, as shown by the 10% transferred memory bytes in the Sysbench benchmark.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Scott Trent, Takeshi Yoshimura, Marcelo Amaral
CLOUD2
2022 Detecting Layered Bottlenecks in Microservices
abstract
We propose a method to detect both software and hardware bottlenecks in a web service consisting of microservices. A bottleneck is a resource that limits the maximum performance of the entire web service. Bottlenecks often include both software resources such as threads, locks, and channels, and hardware resources such as processors, memories, and disks. Bottlenecks form a layered structure since a single request can utilize multiple software resources and a hardware resource simultaneously. The microservice architecture makes the detection of layered bottlenecks challenging due to the lack of a uniform analysis perspective across languages, libraries, frameworks, and middle-ware.We detect layered bottlenecks in microservices by profiling numbers and status of working threads in each microservice and dependency among microservices via network connections. Our approach can be applied to various programming languages since it relies only on standard debugging tools. Nevertheless, our approach not only detects which microservice is a bottleneck but also enables us to understand why it becomes a bottleneck. This is enabled by a novel visualization method to show layered bottlenecks in microservices at a glance. We demonstrate that our approach successfully detects and visualizes layered bottlenecks in the state-of-the-art microservice benchmarks, DeathStarBench and Acme Air microservices. This enables us to optimize the microservices themselves to achieve a higher throughput per re-source utilization rate compared with simply scaling the number of replicas of microservices.
Tatsushi Inagaki, Yohei Ueda, Moriyoshi Ohara, Sunyanan Choochotkaew, Marcelo Amaral, Scott Trent, Tatsuhiro Chiba, Qi Zhang 0009
CLOUD7
2021 Run Wild: Resource Management System with Generalized Modeling for Microservices on Cloud
abstract
Microservice architecture competes with the traditional monolithic design by offering benefits of agility, flexibility, reusability resilience, and ease of use. Nevertheless, due to the increase in internal communication complexity, care must be taken for resource-usage scaling in harmony with placement scheduling, and request balancing to prevent cascading performance degradation across microservices. We prototype Run Wild, a resource management system that controls all mechanisms in the microservice-deployment process covering scaling, scheduling, and balancing to optimize for desirable performance on the dynamic cloud driven by an automatic, united, and consistent deployment plan. In this paper, we also highlight the significance of co-location aware metrics on predicting the resource usage and computing the deployment plan. We conducted experiments with an actual cluster on the IBM Cloud platform. RunWild reduced the 90th percentile response time by 11% and increased average throughput by 10% with more than 30% lower resource usage for widely used autoscaling benchmarks on Kubernetes clusters.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Scott Trent, Marcelo Amaral
CLOUD2
2020 ImageJockey: A Framework for Container Performance Engineering
abstract
Containerized applications have become widely used in modern software development due to their high flexibility and lightweight deployment. Currently, there exists a large set of publicly available container images with many different operating systems (OSs) and versions. As a result, developers typically select images carefully on the basis of memory footprint versus flexibility with available libraries. However, information regarding performance is insufficient. We have verified that different OS-based images with different versions vary the performance of certain applications running in them. Additionally, minor updates in container images, without changing its versions, also affect application performance. Therefore, to understand application performance, it is important to determine impact on performance in different container images along with the continuous monitoring of the minor changes. Since existing performance test frameworks do not encompass the required features to analyze performance regressions in container images, in this paper, we introduce ImageJockey, an original test framework to continuously evaluate the performance of a broad range of container images. Our framework enables experiments of container benchmarks with periodic image builds, simple container orchestration, metrics collection, and result visualization. We demonstrate the usefulness of our framework through case studies that analyze the performance characteristics of sixteen container images and nine popular benchmarks. The experimental results show that there is a noticeable performance variation due to the deployed environments and characteristics of Alpine and JDK images.
Takeshi Yoshimura, Rina Nakazawa, Tatsuhiro Chiba
CLOUD3
2020 Investigating Genome Analysis Pipeline Performance on GATK with Cloud Object Storage
abstract
Achieving fast, scalable, and cost-effective genome analytics is always important to open up a new frontier in biomedical and life science. Genome Analysis Toolkit (GATK), an industry-standard genome analysis tool, improves its scalability and performance by leveraging Spark and HDFS. Spark with HDFS has been a leading analytics platform in a past few years, however, the system cannot exploit full advantage of cloud elasticity in a recent modern cloud. In this paper we investigate performance characteristics of GATK using Spark with HDFS and identify scalability issues. Based on a quantitative analysis, we introduce a new approach to utilize Cloud Object Storage (COS) in GATK instead of HDFS, which can help decoupling compute and storage. We demonstrate how this approach can contribute to the improvement of the entire pipeline performance and cost saving. As a result, we demonstrate GATK with IBM COS can achieve up to 28% faster than GATK with HDFS. We also show that this approach can achieve up to 67 % cost saving in total, which includes the time for data loading and whole pipeline analysis.
Tatsuhiro Chiba, Takeshi Yoshimura
MASCOTS1
2019 EvFS: User-level, Event-Driven File System for Non-Volatile Memory
Takeshi Yoshimura, Tatsuhiro Chiba, Hiroshi Horii
HotStorage2
2019 ConfAdvisor: A Performance-centric Configuration Tuning Framework for Containers on Kubernetes
abstract
Configuration tuning of software is often a good option to improve application performance without any application code modifications. Although we can casually change configurations, it is not easy to apply optimal configurations, as optimal configurations require deep knowledge of the underlying system. This is problematic because applications with suboptimal configuration result in poor performance. As container and container management systems have emerged as an application platform on the cloud, configuration tuning becomes even more challenging because containers add more complexity to the application performance. We need to consider not only fundamental misconfiguration but also container image verification, deployment configuration, application characteristics awareness based on metrics and logs. Although previous knowledge regarding how we should tune configurations for a system software is sometimes available, knowledge about performance tuning practices is neither normalized nor reusable to expand on any advice for misconfiguration to the containers. Even in the cloud-native environment, there is no centralized service to deliver knowledge continuously to application containers nor a framework to develop a misconfiguration fix rule for a container throughout its lifetime. In this paper, we propose a performance-centric configuration tuning framework for containers on Kubernetes, named ConfAdvisor, that enables containers to achieve a higher performance by validating various misconfigurations adaptively. ConfAdivsor gives config tuning advice to application containers, images, and Kubernetes specs and also provides a development framework to build configuration validation rules. We present the design of ConfAdvisor and provide several case studies to tune application containers in the real world.
Tatsuhiro Chiba, Rina Nakazawa, Hiroshi Horii, Sahil Suneja, Seetharami Seelam
IC2E1
2018 Towards Selecting Best Combination of SQL-on-Hadoop Systems and JVMs
abstract
While Hadoop is the de facto standard big-data middleware, many frameworks have been developed on top of it. Since many SQL-on-Hadoop systems are available, we often consider which engine is best for our queries. We can choose not only query engines but also Java virtual machines (JVMs) as well. As their systems become more complex, however, it is not always true that a single system performs best at any time. Moreover, the performance of a mismatched system may degrade greatly. To exploit the best performance, it is important to know what type of queries are suitable for a system and then to schedule queries for the appropriate system. In this paper, we evaluated the TPC-DS benchmark on a combination of query engines (Spark and Tez) and JVMs (J9 and OpenJDK). We found that using different engines lead to a drawback of over 10 times and that using different JVMs leads to a drawback of 3 times. We also analyzed the characteristics of each combination and then proposed classification models for selecting the best combination of systems with a generated query plan. As a result, we achieved a performance improvement of up to two times in total with the classifier.
Tatsuhiro Chiba, Takeshi Yoshimura, Michihiro Horie, Hiroshi Horii
IEEE CLOUD1
2018 Column Cache: Buffer Cache for Columnar Storage on HDFS
abstract
Columnar storage is a data source for data analytics in distributed computing frameworks. For portability and scalability, columnar storage is built on top of existing distributed file systems with columnar data representations such as Parquet, RCFile, and ORC. However, these representations fail to utilize high-level information (e.g., columnar formats) for low-level disk buffer management in operating systems. As a result, data analytics workloads suffer from redundant memory buffers with expensive garbage collections, unnecessary disk readahead, and cache pollution in the operating system buffer cache.We propose column cache, which unifies and re-structures the buffers and caches of multiple software layers from columnar storage to operating systems. Column cache leverages high-level information such as file formats and query plans for enabling adaptive disk reads and cache eviction policies. We have developed a column cache prototype for Apache Parquet and observed that our prototype reduced redundant resource utilization in Apache Spark. Specifically, with our prototype, Spark showed a maximum speedup of 1.28x in TPC-DS workloads while increasing Linux page cache size by 18%, reducing total disk reads by 43%, and reducing garbage collection time in a Java virtual machine by 76%.
Takeshi Yoshimura, Tatsuhiro Chiba, Hiroshi Horii
IEEE BigData2
2016 Workload characterization and optimization of TPC-H queries on Apache Spark
abstract
Besides being an in-memory-oriented computing framework, Spark runs on top of Java Virtual Machines (JVMs), so JVM parameters must be tuned to improve Spark application performance. Misconfigured parameters and settings degrade performance. For example, using Java heaps that are too large often causes a long garbage collection pause time, which accounts for over 10-20% of application execution time. Moreover, recent computing nodes have many cores with simultaneous multi-threading technology and the processors on the node are connected via NUMA, so it is difficult to exploit best performance without taking into account of these hardware features. Thus, optimization in a full stack is also important. Not only JVM parameters but also OS parameters, Spark configuration, and application code based on CPU characteristics need to be optimized to take full advantage of underlying computing resources. In this paper, we used the TPC-H benchmark as our optimization case study and gathered many perspective logs such as application, JVM (e.g. GC and JIT), system utilization, and hardware events from a performance monitoring unit. We discuss current problems and introduce several JVM and OS parameter optimization approaches for accelerating Spark performance. As a result, our optimization exhibits 30-40% increase in speed on average and is up to 5x faster than the naive configuration.
Tatsuhiro Chiba, Tamiya Onodera
ISPASS1
2010 Dynamic Load-Balanced Multicast for Data-Intensive Applications on Clouds
abstract
Data-intensive parallel applications on clouds need to deploy large data sets from the cloud's storage facility to all compute nodes as fast as possible. Many multicast algorithms have been proposed for clusters and grid environments. The most common approach is to construct one or more spanning trees based on the network topology and network monitoring data in order to maximize available bandwidth and avoid bottleneck links. However, delivering optimal performance becomes difficult once the available bandwidth changes dynamically. In this paper, we focus on Amazon EC2/S3 (the most commonly used cloud platform today) and propose two high performance multicast algorithms. These algorithms make it possible to efficiently transfer large amounts of data stored in Amazon S3 to multiple Amazon EC2 nodes. The three salient features of our algorithms are (1) to construct an overlay network on clouds without network topology information, (2) to optimize the total throughput dynamically, and (3) to increase the download throughput by letting nodes cooperate with each other. The two algorithms differ in the way nodes cooperate: the first `non-steal' algorithm lets each node download an equal share of all data, while the second `steal' algorithm uses work stealing to counter the effect of heterogeneous download bandwidth. As a result, all nodes can download files from S3 quickly, even when the network performance changes while the algorithm is running. We evaluate our algorithms on EC2/S3, and show that they are scalable and consistently achieve high throughput. Both algorithms perform much better than having each node downloading all data directly from S3.
Tatsuhiro Chiba, Mathijs den Burger, Thilo Kielmann, Satoshi Matsuoka
CCGRID1
2007 High-Performance MPI Broadcast Algorithm for Grid Environments Utilizing Multi-lane NICs
abstract
The performance of MPI collective operations, such as broadcast and reduction, is heavily affected by network topologies, especially in grid environments. Many techniques to construct efficient broadcast trees have been proposed for grids. On the other hand, recent high performance computing nodes are often equipped with multi-lane network interface cards (NICs), most previous collective communication methods fail to harness effectively. Our new broadcast algorithm for grid environments harnesses almost all downward and upward bandwidths of multi-lane NICs; A message to be broadcast is split into two pieces, which are broadcast along two independent binary trees in a pipelined fashion, and swapped between both trees. The salient feature of our algorithm is generality; it works effectively on both large clusters and grid environments. It can be also applied to nodes with a single NIC, by making multiple sockets share the NIC. Experimentations on a emulated network environment show that we achieve higher performance than traditional methods, regardless of network topologies or the message sizes.
Tatsuhiro Chiba, Toshio Endo, Satoshi Matsuoka
CCGRID1