Sohei Koyama

dblp:326/0078 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0002-1912-0075ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 44% High-performance computing · 44% Performance modeling and evaluation · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › file systems › distributed file system
parallel file system
1.012026
Lustre Query: Periodic Offline Metadata Monitoring from MDT Backups · HPDC 2026
Performance modeling and evaluation
capacity planning
0.312026
Lustre Query: Periodic Offline Metadata Monitoring from MDT Backups · HPDC 2026

Methods — techniques the papers use, named apart from their topics

offline metadata extraction · 1.0
YearPublicationVenuePosition
2026 Lustre Query: Periodic Offline Metadata Monitoring from MDT Backups
abstract
Parallel file systems in HPC manage metadata for billions of files across distributed storage servers. The structure of this metadata, how files distribute by size, how users concentrate across servers, which storage policies are actually in use, determines operational decisions about capacity planning, load balancing, and data migration. Despite decades of HPC storage research, these structural properties remain underreported in the published literature for HPC parallel file systems. Runtime I/O behavior has been profiled extensively at the application level. Aggregate monitoring captures quotas and throughput. But the metadata that accumulates on disk, the artifact of all user activity over the life of a system, has received little systematic study. Existing tools can extract inode-level detail in principle, but each imposes barriers that discourage routine analysis: online queries load the metadata server, database replicas require ETL pipelines, and low-level utilities demand scripting that few administrators undertake.
Sohei Koyama, Osamu Tatebe
HPDC1
2024 FINCHFS: Design of Ad-Hoc File System for I/O Heavy HPC Workloads
abstract
Although the performance improvements in parallel file systems have been significant, the rise of data science and deep learning using Python has introduced new I/O requirements. We redefine ad-hoc file systems as a complement to parallel file systems by specializing in requirements that cannot be met by improvements to parallel file systems. Ad-hoc file systems require extremely high metadata performance and the scalability to create and read large numbers of files in parallel in a single directory. Therefore, it is necessary to process a large number of small RPCs with low latency and high throughput, which cannot be achieved with existing ad-hoc file systems. In this research, we propose a new ad-hoc file system, FINCHFS, and realize the truly desired ad-hoc file system using intra-node shared-nothing architecture and number-aware hashing. The contribution of the intra-node shared-nothing architecture is the ability to run multiple iterative server processes on each node to achieve high IOPS using many-core. The contribution of number-aware hashing is to alleviate tail latency degradation due to the uneven distribution of requests to servers. In the evaluation using 64 nodes on the Pegasus supercomputer, we achieved 32.5 MIOPS in mdtest-hard-write.
Sohei Koyama, Kohei Hiraga, Osamu Tatebe
CLUSTER1