Kiyofumi Tanaka

dblp:83/4136 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-0148-9250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Performance Evaluation of Multi-Head Logging in Flash File Systems
abstract
As NAND flash memory has become a widely used storage medium, traditional file systems face challenges in adapting to its unique characteristics. The Flash-Friendly File System (F2FS) is designed to optimize file and data management on NAND flash, addressing issues such as raw NAND interface compatibility, the "wandering tree" problem, high reading costs, and inefficient garbage collection. Through a log-structured foundation and various optimizations, including multi-head logging, node address table (NAT), and reserved in-place updated metadata sections, F2FS achieves enhanced performance, reliability, and prolonged lifespan. This study constructs a multi-head log model based on F2FS and introduces a new hardware-independent metric—Write Amplification Factor (WAF)—to analyze performance. The research demonstrates that separating node and data blocks is more effective in improving garbage collection efficiency than simple hot and cold data segregation.
Jingcheng Yuan, Kiyofumi Tanaka, Toshiaki Aoki
COMPSAC2
2025 Preliminary evaluation of SHAVER: sharing vector registers with an accelerator
Tomoaki Tanaka, Michiya Kato, Yasunori Osana, Takefumi Miyoshi, Jubee Tada, Kiyofumi Tanaka, Hironori Nakajo
J. Supercomput.6
2024 High Throughput and Low Bandwidth Demand: Accelerating CNN Inference Block-by-block on FPGAs
abstract
A multitude of accelerators have been designed to accelerate the inference of widely-used Convolutional Neural Networks (CNNs). They can primarily be classified into two architectures: the Overlay architecture, which accelerates layer by layer, and the Dataflow architecture, which accelerates the entire model. Each has its own strengths and weaknesses, and we will propose a novel architecture to capitalize on their strengths and mitigate their weaknesses. Our architecture allows optimization for different target network structures. Internally, it features multiple PE Arrays, and during runtime, it can flexibly switch between different interconnection modes among these PE Arrays: the serial execution mode, accelerating an entire block at once, and the parallel execution mode, accelerating a single layer at a time. Its minimum area requirement is close to that of typical Overlay architecture accelerators, while its throughput per unit area far exceeds that of existing state-of-the-art accelerators. When executing the widely-used 8-bit quantized MobileNetV2, we observed an exceptionally high throughput of 2331 FPS (frames per second) on the mid-range FPGA ZU7EV, while requiring only 2.89 GiB/s of off-chip memory bandwidth. Running other networks also demonstrated high efficiency and reasonable off-chip memory bandwidth requirements.
Yan Chen 0029, Kiyofumi Tanaka
DSD2
2024 A Flexible, Fast, Low Bandwidth Block-based Acceleration Architecture for CNN Inference on FPGAs
abstract
In recent years, Convolutional Neural Networks (CNNs) have been widely used in many fields. Many accelerators are designed to run CNN inference on hardware. They are classified into two types: Overlay and Dataflow. In this paper, we propose a new architecture of CNN Inference accelerators to leverage the strengths of each type. It is flexible, fast, and has low off-chip memory bandwidth requirements. The accelerator based on our new architecture can be configured to fit the CNN structure and FPGA capacity. It basically runs CNNs by blocks, whereas it has a fallback path to run single Convolution or Fully Connected layers efficiently. Our architecture demonstrates significant advantages when running networks composed of inverted residual blocks. We archived 1954FPS (frames per second) when running 8-bit quantized MobileNetV2 on a mid-range FPGA ZU7EV, and the off-chip memory bandwidth requirement is as low as 2.47GiB/s. It also runs on cost-optimized FPGAs, and archived 505FPS on ZU3EG. The throughputs of ours are higher than the other types, while the minimum FPGA capacity and the off-chip memory bandwidth requirement of ours are reasonable.
Yan Chen 0029, Kiyofumi Tanaka
FPGA2
2021 Efficient FPGA Design of Exception-Free Generic Elliptic Curve Cryptosystems
Kiyofumi Tanaka, Atsuko Miyaji, Yaoan Jin
ACNS (1)1
2021 A New Memory Consistency Model for Real-Time Multicore Processors
abstract
Several relaxed memory consistency models have been proposed to improve performance with reduced messages and data on software/hardware distributed shared memory. This paper proposes a new relaxed consistency model, Midway between Eager and Lazy release consistency (MEL), for real-time RISC-V multicore embedded processors. MEL ensures that memory accesses inside critical sections obtain the latest values with reduced communication latency by using ff_load and fu_store instructions [1]. We evaluate MEL with test programs in terms of necessary hardware resources.
Aye Myat Mon, Kiyofumi Tanaka
TENCON2
2020 Dependency-Driven Trace-Based Network-on-Chip Emulation on FPGAs
abstract
FPGA emulation is a promising approach to accelerating Network-on-Chip (NoC) modeling which has traditionally relied on software simulators. In most early studies of FPGA-based NoC emulators, only synthetic workloads like uniform and bit permutations were considered. Although a set of carefully designed synthetic workloads can reveal a relatively thorough coverage of the characteristics of the NoC under evaluation, they alone are insufficient, especially when the NoC needs to be optimized for specific applications. In such cases, trace-driven workloads are effective. However, there is a problem with conventional trace-driven workloads that has been pointed out by some recent studies: the network load and congestion may be distorted because dependencies between packets are not considered. These studies also provide infrastructures for extending existing software simulators to enforce dependencies between packets. Unfortunately, enforcing dependencies between packets is not trivial in the FPGA emulation approach. Therefore, although there are some recent FPGA-based NoC emulators supporting trace-driven workloads, most of them ignore packet dependencies. In this paper, we first clarify the challenges of supporting trace-driven workloads with dependencies between packets taken into account in the FPGA emulation approach. We then propose efficient methods and architectures to tackle these challenges and build an FPGA-based NoC emulator, which we call DNoC, based on the proposals. Our evaluation results show that (1) on a VC707 FPGA board, DNoC achieves an average speed of 10,753K cycles/s when emulating an 8x8 NoC with trace data collected from full-system simulation of the PARSEC benchmark suite, which is 274x higher than the speed reported in a recent related work on dependency-driven trace-based NoC emulation on FPGAs; (2) Compared to BookSim, one of the most popular NoC simulators, DNoC is 395x faster while providing the same results; (3) DNoC can scale to a 4,096-node NoC on a VC707 board, and the size of the largest NoC depends on only the on-chip memory capacity of the target FPGA.
Thiem Van Chu, Kenji Kise, Kiyofumi Tanaka
FPGA3
2019 Adaptive Local Assignment Algorithm for Scheduling Soft-Aperiodic Tasks on Multiprocessors
abstract
With the emergence of multiprocessors, embedded systems are nowadays capable of handling diverse and complicated applications. Besides static workloads such as periodic tasks, dynamic workloads such as aperiodic tasks happen more frequently and become challenging to researchers to deal with. In this study, an effective scheduling approach for aperiodic tasks in multiprocessors is presented. An adaptive Local Assignment Algorithm with the integration of servers is introduced to schedule mixture systems of periodic and aperiodic tasks. Servers are dedicatedly preserved for aperiodic tasks at runtime. This scheduling scheme guarantees the schedulability of 100% without significantly increasing the time complexity. Simulation results show that the proposed mechanism effectively improves the responsiveness of aperiodic tasks while maintaining relatively low runtime overhead. In addition, it causes much fewer scheduler invocations in comparison with the existing algorithms.
Doan Duy, Kiyofumi Tanaka
RTCSA2
2016 Configuration technique for adaptability of multicore processors on FPGA
abstract
To reach the efficient use of FPGA resources, we have developed a configuration environment for adapting a soft processor to a specific application program automatically. In our approach, an application program is analyzed and only components/forwarding paths which the instruction sequences in the program actually use are selected to build a customized soft processor in a minimum size. In this paper, we show, first, the processor core we designed as the base of configurable processors. Then, the functions of the configurator and how the configurator works are described. In the evaluation of the processors which are generated using this environment, a configured eight-core processor can be implemented in a relatively small FPGA device, while a two-core processor without configuration uses up the FPGA resources.
Tetsuo Miyauchi, Kiyofumi Tanaka
ASAP2
2005 Casablanca II: Implementation of a Real-Time RISC
abstract
We extended general-purpose RISC processor architecture and developed a new RISC core, Casablanca II, for supporting real-time processing in embedded systems. The processor core has multiple register-sets and achieves fast context-switching by automatically changing the active register-set and reducing overheads to save and restore the contents of the registers when exceptions or interruptions occur. In addition, the core has mechanisms for explicit data cache control, enabling data prefetching and fast DMA, which is invoked by executing extended instructions. In this paper, we describe the organization of Casablanca II developed by using an ASIC process and present preliminary evaluation of the processor.
Kiyofumi Tanaka
ASAP1
1999 Lightweight Hardware Distributed Shared Memory Supported by Generalized Combining
abstract
On a large scale parallel computer system, shared memory provides a general and convenient programming environment. The paper describes a lightweight method for constructing an efficient shared memory system supported by hierarchical coherence management and generalized combining. The hierarchical management technique and generalized combining cooperate with each other. We eliminate the following heavyweight and high cost factors: a large amount of directory memory which is proportional to the number of processors, a separate memory component for the directory, tag/state information, and a protocol processor. In our method, the amount of memory required for the directory is proportional to the logarithm of the number of processors. This implies that a single word for each memory block is sufficient for covering a massively parallel system and that the access costs of the directory are small. Moreover, our combining technique, generalized combining, does not expect the accidental events which existing combining networks do, that is, events that messages meet each other at a switching node. A switching node can combine succeeding messages with a preceding one even after the preceding message leaves the node. This can increase the rate of successful combining. We have developed a prototype parallel computer OCHANOMIZ-5, that implements this lightweight distributed shared memory and generalized combining with simple hardware. The results of evaluating the prototype's performance using several programs show that our methodology provides the advantages of parallelization.
Kiyofumi Tanaka, Takashi Matsumoto 0002, Kei Hiraki
HPCA1