Atsushi Koshiba

dblp:165/8394 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-5439-4357ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Proteus: Heterogeneous FPGA Virtualization
abstract
Cloud providers have widely adopted FPGAs to meet the high-performance, energy-efficient demands of cloud workloads. While they offer homogeneous FPGAs per service, recent FPGA products exhibit increasing heterogeneity in terms of vendors, capacity, off-chip memory, and performance. These diverse properties not only render applications incompatible across different FPGAs but also lead to performance disparities and resource inefficiency.
Felix Gust, Shu Anzai, Charalampos Mainas, Atsushi Koshiba, Pramod Bhatotia
EuroSys4
2025 TNIC: A Trusted NIC Architecture: A hardware-network substrate for building high-performance trustworthy distributed systems
abstract
We introduce TNIC, a trusted NIC architecture for building trustworthy distributed systems deployed in heterogeneous, untrusted (Byzantine) cloud environments. TNIC builds a minimal, formally verified, silicon root-of-trust at the network interface level. We strive for three primary design goals: (1) a host CPU-agnostic unified security architecture by providing trustworthy network-level isolation; (2) a minimalistic and verifiable TCB based on a silicon root-of-trust by providing two core properties of transferable authentication and non-equivocation; and (3) a hardware-accelerated trustworthy network stack leveraging SmartNICs. Based on the TNIC architecture and associated network stack, we present a generic set of programming APIs and a recipe for building high-performance, trustworthy, distributed systems for Byzantine settings. We formally verify the safety and security properties of our TNIC while demonstrating its use by building four trustworthy distributed systems. Our evaluation of TNIC shows up to 6× performance improvement compared to CPU-centric TEE systems.
Dimitra Giantsidi, Julian Pritzi, Felix Gust, Antonios Katsarakis, Atsushi Koshiba, Pramod Bhatotia
ASPLOS (2)5
2025 Funky: Cloud-Native FPGA Virtualization and Orchestration
abstract
The adoption of FPGAs in cloud-native environments is facing impediments due to FPGA limitations and CPU-oriented design of orchestrators, as they lack virtualization, isolation, and preemption support for FPGAs. Consequently, cloud providers offer no orchestration services for FPGAs, leading to low scalability, flexibility, and resiliency.
Atsushi Koshiba, Charalampos Mainas, Pramod Bhatotia
SoCC1
2025 F3: An FPGA-accelerated FaaS Framework
abstract
FPGAs provide a programmable, energy-efficient, and compute-intensive acceleration substrate; thus, on the one hand, they offer a compelling solution for optimizing serverless workloads in cloud environments. On the other hand, FPGAs also introduce significant challenges that directly contradict the serverless model in the cloud, including their low-level and complex programming APIs, lack of virtualization and isolation mechanisms, high reconfiguration and communication overheads, and absence of orchestration mechanisms.
Charalampos Mainas, Martin Lambeck, Bruno Scheufler, Laurent Bindschaedler, Atsushi Koshiba, Pramod Bhatotia
HPDC5
2024 vFPIO: A Virtual I/O Abstraction for FPGA-accelerated I/O Devices
Jiyang Chen, Harshavardhan Unnibhavi, Atsushi Koshiba, Pramod Bhatotia
USENIX ATC3
2023 ESSPER: Elastic and Scalable FPGA-Cluster System for High-Performance Reconfigurable Computing with Supercomputer Fugaku
abstract
FPGA clusters have yet to be a mainstream of HPC, even for accelerators, and several challenges exist in their architecture and system organization. This work presents ESSPER, a flexible and scalable FPGA cluster prototype system for reconfigurable HPC to meet the concept of customizability, scalability, and interoperability with existing HPC systems. Based on our classification of FPGA cluster architectures, we propose a new category of FPGA clusters with a host-FPGA bridging network using software-bridged APIs for the use of remote FPGAs. We have designed, implemented, verified, and demonstrated a proof-of-concept system of ESSPER, as a functional extension of the supercomputer Fugaku.
Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno
HPC Asia2
2022 ESSPER: Elastic and Scalable System for High-Performance Reconfigurable Computing with Software-bridged APIs
abstract
Many-core CPUs and GPUs, present mainstream architectures for HPC, are facing difficulty in maintaining the same performance improvement rate because of the recent slow-down in the semiconductor scaling, the dark silicon problem, and wasteful mechanisms required for accelerating general-purpose computing such as a branch predictor and an out-of-order mechanism. Also, the power efficiency of HPC systems is significantly important to achieve higher performance.
Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno
FPT2
2021 Virtual Circuit-Switching Network with Flexible Topology for High-Performance FPGA Cluster
abstract
As the performance of high-end FPGAs has increased in recent years, it’s getting more important to construct an FPGA cluster for both improved processing performance and power efficiency in data centers and supercomputers. For higher utilization of FPGA resources for various applications, we require a flexible inter-FPGA network which provides various topologies appropriately to different applications while a conventional direct-connection network (DCN) provides only a fixed topology, such like a 2D torus. In this paper, we propose a virtual circuit-switching network (VCSN) for a large-scale FPGA cluster to have a flexible inter-FPGA network, where communication links connecting FPGAs are virtualized on the top of Ethernet frames. We can easily configure the VCSN topology optimized for the application by modifying the destination MAC addresses registered in a table of a frame encoder. We present its efficient protocol, hardware implementation, demonstration with 100Gbps Ethernet, and performance comparison with a conventional direct-connection network for FPGAs. We show that VCSN has higher but acceptable latency and slightly higher throughput in comparison with DCN, so that numerical simulation running with a ring of FPGAs achieves comparable performance for DCN.
Tomohiro Ueno, Atsushi Koshiba, Kentaro Sano
ASAP2
2020 Performance Evaluation and Power Analysis of Teraflop-scale Fluid Simulation with Stratix 10 FPGA
abstract
Stream computing is a suitable approach to improve both performance and power efficiency of numerical computations with FPGAs. To achieve further performance gain, temporal and spatial parallelism were exploited: the first one deepens and the latter duplicates pipelines of streamed computation cores. These two types of parallelism were previously evaluated with Arria 10 FPGA. However, it has not been verified if they are also effective for the latest FPGA, Stratix 10, which has a larger amount of logic elements (i.e., 2.4X of Arria 10) and is equipped with a new feature to improve the maximum clock frequency (i.e., HyperFlex architecture). To show the scalability for such state-of-the-art FPGAs, in this paper, we firstly implemented a streamed fluid simulation accelerator with both parallelism types for Stratix 10. We then thoroughly evaluated it by obtaining computational performance (FLOPS), power efficiency (FLOPS/W), resource utilization, and maximum clock frequency (Fmax). From the results, we found that this implementation excessively used DSP blocks due to inefficient mapping of floating-point operations, which reduced Fmax and the number of pipelined cores. To improve the scalability, we optimized the implementation to reduce the DSP block usage by utilizing a Multiply-Add function in a single DSP block. As a result, the optimized fluid simulation achieves 1.06 TFLOPS and 12.6 GFLOPS/W, which is 1.36X and 1.24X higher than the non-optimized version, respectively. Moreover, we estimate that the fluid simulation with Stratix 10 could outperform GPU-based implementation with Tesla V100 by optimizing it for HyperFlex architecture.
Atsushi Koshiba, Kouki Watanabe, Takaaki Miyajima, Kentaro Sano
FPGA1
2018 TEE-KV: Secure Immutable Key-Value Store for Trusted Execution Environments
abstract
Trusted Execution Environments (TEEs) ensure strong data confidentiality for applications running in the TEEs even on untrusted servers. In particular, TEEs are expected to bring significant benefits to blockchain workloads for enterprise because it ensures confidentiality and correctness of transaction records without any heavy-weight data verification process such as proof-of-work. For example, Coco [3] improves both confidentiality and transaction throughput of existing blockchain protocols by utilizing TEE features.
Atsushi Koshiba, Zhongxin Guo, Mitaro Namiki, Lidong Zhou
SoCC1
2018 OpenCL Runtime for OS-Driven Task Pipelining on Heterogeneous Accelerators
abstract
Task pipelining on accelerators is suitable for streaming applications, while its performance can decrease due to frequent user/OS interactions caused by hardware control via device drivers. Our previous research has proposed PPM, an OS support that efficiently manages multiple accelerators by eliminating the user/OS interactions during the pipelined execution. To allow users to develop and execute pipeline applications using PPM, this paper introduces a customized OpenCL runtime library. When executing an OpenCL application, the runtime library dynamically analyzes OpenCL API calls and creates a data flow graph required for the PPM execution. With the runtime library, users easily execute applications written in common OpenCL pipeline model on PPM.
Atsushi Koshiba, Ryuichi Sakamoto, Mitaro Namiki
RTCSA1