Takaaki Fukai

dblp:211/9065 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-4216-4807ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NecoFuzz: Effective Fuzzing of Nested Virtualization via Fuzz-Harness Virtual Machines
abstract
Nested virtualization is now widely supported by major cloud vendors, allowing users to leverage virtualization-based technologies in the cloud. However, supporting nested virtualization significantly increases host hypervisor complexity and introduces a new attack surface in cloud platforms. While many prior studies have explored hypervisor fuzzing, none has explicitly addressed nested virtualization due to the challenge of generating effective virtual machine (VM) instances with a vast state space as fuzzing inputs.
Reima Ishii, Takaaki Fukai, Takahiro Shinagawa
EuroSys2
2025 BadAML: Exploiting Legacy Firmware Interfaces to Compromise Confidential Virtual Machines
abstract
Confidential virtual machines (CVMs) are an emerging form of trusted execution environment that enable existing operating systems (OSs) to run securely without trusting cloud providers. To this end, CVMs employ hardware-based memory encryption for runtime confidentiality and cryptographic attestation to verify memory integrity at startup. However, we reveal a previously overlooked attack vector that allows malicious cloud providers to bypass CVM attestation and execute arbitrary code within users' CVMs regardless of specific CVM configurations. Our attack, BadAML, exploits the Advanced Configuration and Power Interface (ACPI), a legacy yet widely adopted firmware interface for machine configuration. Specifically, BadAML leverages ACPI Machine Language (AML) to inject arbitrary binary code into the guest OS kernel without affecting CVM attestation. Because ACPI remains an essential component even in virtualized environments, BadAML constitutes a powerful and portable attack vector independent of guest OS type and CVM technology. We demonstrate proof-of-concept exploits of BadAML in both Linux and Windows CVM environments. We then analyze possible mitigation measures, discussing their effectiveness and limitations. Finally, we introduce AML sandboxing, a practical defense that restricts memory access to safe regions under the CVM threat model; we present its design, implementation, and evaluation, demonstrating its effectiveness across 18 real-world cloud CVM instances.
Satoru Takekoshi, Manami Mori, Takaaki Fukai, Takahiro Shinagawa
CCS3
2025 Comprehensive Performance Evaluation of Microservices on Confidential Containers in Mutli-access Edge Computing Environments
abstract
Multi-access Edge Computing (MEC) in Edge-cloud computing continuum environments is an emerging technology that enables locality-aware low-latency processing for requests from edge devices. Confidential Containers (CoCo) are an emerging confidential computing technology designed for protecting in-use data in cloud datacenters. Although containers in a MEC layer can help reduce latency for requests from edge devices, the current increasing demands for security features such as confidential computing and memory integrity protection will seriously affect expected end-to-end latencies of microservice applications. In this paper, we build a CoCo environment using Kata containers on the MEC layer and attempt to evaluate microservice-level application behavior for the advanced security features. Especially, we focus on breaking down the latency into memory access level, microservice component level, and the end-to-end latency level. From the results, we observe that 9% overhead is incurred for each memory access when applying confidential computing with memory integrity protection. At the individual microservices level, we observed the latency for gRPC communication and memcached access are increased up to 40 %. We also reveal that the latency overheads in microservice-level and application-level are from twice to 4 times larger than those observed in memory access level.
Itsuki Nakai, Takaaki Fukai, Takahiro Hirofuchi, Yukinori Sato
IC2E2
2025 EFCC: Ethernet Frame Crafter & Capture for TSN Research
abstract
Time-Sensitive Network (TSN) has been considered one of the most viable solutions to meet the increasing demands for ultra-low latency in the 5G and post-5G eras. Due to the importance of timing constraints, network measurement tools are necessary for TSN research and operations to evaluate the performance limits of proofs-of-concept, troubleshoot failures, etc. However, current measurement solutions struggle to achieve important aspects of TSN networks, such as precise control of the transmission frame intervals and burst sizes. Additionally, they cannot generate a mix of multiple flows of various characteristics. In this paper, we present Ethernet Frame Crafter & Capture (EFCC), an FPGA-based network measurement tool to address the aforementioned issues. The proposed system allows the frame generation/capture functions to be configured independently for each port at a line rate of 1GbE. Furthermore, EFCC is open-source and can be ported to most commodity FPGA boards available in the market. The evaluation results demonstrate that the proposed frame generator allows the mixing of multiple flows of various characteristics and that the frame capture module can record the transmit and receive timestamps of frames with a precision of 8-ns/6.4-ns.
Akram Ben Ahmed, Takahiro Hirofuchi, Takaaki Fukai
LCN3
2024 FPGA-Based Network Switch Architecture Supporting Credit Based Shaper for Time Sensitive Networks
abstract
Time Sensitive Network (TSN) is one of the most auspicious solutions to respond to the increasing demands in ultra-low latency in the 5G and post-5G eras. It comes as an extension to the conventional IEEE 802.3 Ethernet networks by adding a set of novel open standards that aims to provide deterministic, reliable, high-bandwidth, and low-latency communication. In this paper, we present an open-source and light-weight FPGA-based network switch design and implementation supporting Credit Based Shaper (CBS) for Time Sensitive Networks. We present the key design components and implementation aspects of the proposed switch and discuss the preliminary evaluation results in a fair amount of detail to validate our proposal. The conducted experiments show that the proposed switch properly shapes the traffic by eliminating bursts in irregular traffic and efficiently forwards prioritized traffic as stipulated by the TSN requirements. In addition, we demonstrate that our hardware latency evaluation results conform with the CBS theoretical model and that the proposed switch consumes a very reasonable portion of the hardware resources on an affordable FPGA.
Akram Ben Ahmed, Takahiro Hirofuchi, Takaaki Fukai
ETFA3
2022 Analyzing I/O Performance of a Hierarchical HPC Storage System for Distributed Deep Learning
Takaaki Fukai, Kento Sato, Takahiro Hirofuchi
PDCAT1
2021 The 16, 384-node Parallelism of 3D-CNN Training on An Arm CPU based Supercomputer
abstract
As the computational cost and datasets available for deep neural network training continue to increase, there is a significant demand for fast distributed training on supercomputers. However, porting and tuning applications for new advanced supercomputers requires tremendous amount of development efforts. Therefore, we present software tuning best practice for a 3D-CNN model training on a new Arm CPU based supercomputer, Fugaku. We (i) tune computation in DL by a JIT translator for aarch64, (ii) optimize collective communication such as Allreduce for 6D mesh/torus network topology, (iii) tune I/O by data staging with compression and data loader with caching, and (iv) parallelize training in data and model parallelism. We apply the proposed methods to a CosmoFlow 3D-CNN model, and achieve the training in 30 minutes using 16,384 nodes consisting of 4096 data- and 4 model-parallelism. This is the fastest result of any CPU-based systems in MLPerf HPC v0.7 in the world.
Akihiro Tabuchi, Koichi Shirahata, Masafumi Yamazaki, Akihiko Kasagi, Takumi Honda, Kouji Kurihara, Kentaro Kawakami, Tsuguchika Tabaru, Naoto Fukumoto, Akiyoshi Kuroda, Takaaki Fukai, Kento Sato
HiPC11
2021 Live Migration in Bare-Metal Clouds
abstract
Live migration allows a running operating system (OS) to be moved to another physical machine with negligible downtime. Unfortunately, live migration is not supported in bare-metal clouds, which lease physical machines rather than virtual machines to offer maximum hardware performance. Since bare-metal clouds have no virtualization software, implementing live migration is difficult. Previous studies have proposed OS-level live migration; however, to prevent user intervention and broaden OS choices, live migration should be OS-independent. In addition, the overhead of live migration mechanisms should be as low as possible. This paper introduces BLMVisor, a live migration scheme for bare-metal clouds. To achieve OS-independent and lightweight live migration, BLMVisor utilizes a very thin hypervisor that exposes physical hardware devices to the guest OS directly rather than virtualizing the devices. The hypervisor captures, transfers, and reconstructs physical device states by monitoring access from the guest OS and controlling the physical devices with effective techniques. To minimize performance degradation, the hypervisor is mostly idle after completing the live migration. A performance evaluation confirmed that the OS performance with BLMVisor is comparable to that of a bare-metal machine.
Takaaki Fukai, Takahiro Shinagawa, Kazuhiko Kato
IEEE Trans. Cloud Comput.1
2018 Live migration on ARM-based micro-datacentres
abstract
Live migration, underpinned by virtualisation technologies, has enabled improved manageability and fault tolerance for servers. However, virtualised server infrastructures suffer from significant processing overheads, system inconsistencies, security issues and unpredictable performance which makes them unsuitable for low-power and resource-constraint computing devices that processing latency-sensitive, “Big-data”-type data. Consequently, we ask: “How do we eliminate the overhead of virtualisation whilst still retaining its benefits?” Motivated by this question, we investigate a practical approach for a bare-metal live migration scheme for ARM-based instances low-power servers and edge devices. In this paper, we position ARM-based bare-metal live migration as a technique that will underpin the efficiency on edge-computing and on Micro-datacentres. We also introduce our early work on identifying three key technical challenges and discuss their solutions.
Ilias Avramidis, Michael Mackay 0001, Fung Po Tso 0001, Takaaki Fukai, Takahiro Shinagawa
CCNC4
2018 FaultVisor2: Testing Hypervisor Device Drivers Against Real Hardware Failures
abstract
Hardware failures are inevitable, especially in cloud environments where there are many hardware devices. To improve the hypervisor's reliability, hypervisor device drivers must handle hardware failures appropriately. Our goal is to allow cloud vendors to test closed-source hypervisor device drivers against failures of their real hardware. Previous studies either require source code, can only test against virtual hardware, or cannot be applied to hypervisors. In this paper, we propose FaultVisor2, a hypervisor device driver testing framework that combines fault injection and nested virtualization. To test closed-source hypervisor device drivers, we inject pseudo faults to the I/O data returned from hardware to hypervisor device drivers. To test against real hardware, we allow the target hypervisors pass-through access to the physical hardware and manipulate I/O data of the target devices by intercepting I/O access. To apply to hypervisors, we exploit nested virtualization and run a small hypervisor underneath the target hypervisor to inject pseudo faults. We omit some nested virtualization functions, including nested paging virtualization, to achieve a close to real execution environment and reduce runtime overhead. In our experiment using the VMWare ESXi hypervisor, we found three types of errors which led to critical system failures.
Masanori Misono, Masahiro Ogino, Takaaki Fukai, Takahiro Shinagawa
CloudCom3
2017 BMCArmor: A Hardware Protection Scheme for Bare-Metal Clouds
abstract
Traditional infrastructure-as-a-service (IaaS) clouds provide virtual machines as servers. However, virtualization incurs a performance overhead and prevents maximum utilization of hardware functions, so several IaaS vendors have started new services called bare-metal clouds that provide physical rather than virtual machines, allowing users to have direct access to physical hardware in the cloud. Unfortunately, exposing physical hardware to users causes a hardware protection issue for cloud vendors. Since physical hardware uses non-volatile memory (NVM) to store firmware code and configuration data, this is also exposed to users. If the NVM is modified by malicious users, the hardware could be permanently corrupted or infected by malware without being noticed. This is difficult for cloud vendors to prevent because bare-metal clouds have no virtualization layer to protect their hardware. In this paper, we describe the types of attacks that are possible for bare-metal clouds and propose BMCArmor, a hardware protection scheme for baremetal clouds. BMCArmor uses a thin hypervisor that does not virtualize the hardware, just preventing access to NVM. Our experiments show that BMCArmor can successfully protect hardware while incurring little performance overhead.
Takaaki Fukai, Satoru Takekoshi, Kohei Azuma, Takahiro Shinagawa, Kazuhiko Kato
CloudCom1