EDBT 2026 Demo / reviewers in the wild / expert
Yutaka Urino
dblp:71/7200
· DBLP profile ↗
6ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0001-8254-5562ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Effective switchless inter-FPGA memory networks
Truong Thao Nguyen, Kien Trung Pham, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
J. Parallel Distributed Comput. | 4 |
| 2022 | A Scalable Distributed Radix Sorter for FPGA Clusters using High-Bandwidth Memory NetworksabstractA modern FPGA card can be equipped with high bandwidth memory, such as HBM2. Since the amount of the memory is limited on an FPGA, highly parallel data processing becomes crucial on tightly coupled FPGAs by high-density optical integration, e.g., onboard Si-Photonics transceivers. This study presents a scalable distributed radix sorter, and implements it on an eight-FPGA cluster. Each custom Stratix10 MX2100 FPGA card has 819-Gbps memory bandwidth with two HBM2 memories and 800-Gbps network bandwidth with eight custom embedded optical modules. Existing FPGA sorter typically relies on a merge sort. However, it has a severe performance bottleneck at the final stage of data merge, which cannot make the best use of the high memory-to-memory bandwidth on the FPGA cluster. Instead, we implement a radix sort for a 32-bit key range consisting of eight 4-bit counting sorts optimized to the memory-network structure. Each counting sort needs memory read/write access only once through global and local pipelines. We demonstrated a sorting throughput of 37.2 GB/s. Yutaka Urino, Takanori Shimizu, Hiroshi Yamaguchi, Kenji Mizutani, Shigeru Nakamura, Tatsuya Usuki, Michihiro Koibuchi |
FCCM | 1 |
| 2022 | Scalable Low-Latency Inter-FPGA NetworksabstractA cutting-edge FPGA card can be equipped with many high-bandwidth I/Os by means of high-density optical integration, e.g., onboard Si-photonics transceivers, to provide high network bandwidth for memory-to-memory inter-FPGA communication. This study presents its scalable switchless net-work architecture by exploiting an indirect path, consisting of two one-hop paths, for enabling a diameter-2 network topology. It then takes a Kautz network topology with a diameter of two for connecting d(d + 1) FPGAs with a degree of$d$, which is close to the theoretical upper bound. The Kautz network topologies have bi-directional links and uni-directional links which form triangles. Uni-directional links introduce difficulty in avoiding channel buffer overflow because the existing link-level flow control assumes a bi-directional link. This study presents an indirect flow control along a uni-directional triangle embedded in the Kautz network topology. It then develops a combination of unicasts that forms multi-port collective communications to mitigate the influence of the startup latency on the execution time. Since a high-degree FPGA card introduces difficulty in storing many I/O ports at the panel of a 1- U compute server, we propose using WDM (Wavelength Division Multiplexing) as an alternative and present its efficient mapping onto arrayed waveguide grating (AWG). The required number of wavelengths becomes d on d+ 1 AWG equipments. Based on our experimental results with OPTWEB of custom Stratix10 FPGA cards, SimGrid simulation results show that our collective communication is 7 × faster than that of Dragonfly with 272 FPGAs. Kien Trung Pham, Truong Thao Nguyen, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IPDPS | 4 |
| 2021 | OPTWEB: A Lightweight Fully Connected Inter-FPGA Network for Efficient CollectivesabstractModern FPGA accelerators can be equipped with many high-bandwidth network I/Os, e.g., 64 x 50 Gbps, enabled by onboard optics or co-packaged optics. Some dozens of tightly coupled FPGA accelerators form an emerging computing platform for distributed data processing. However, a conventional indirect packet network using Ethernet's Intellectual Properties imposes an unacceptably large amount of the logic for handling such high-bandwidth interconnects on an FPGA. Besides the indirect network, another approach builds a direct packet network. Existing direct inter-FPGA networks have a low-radix network topology, e.g., 2-D torus. However, the low-radix network has the disadvantage of a large diameter and large average shortest path length that increases the latency of collectives. To mitigate both problems, we propose a lightweight, fully connected inter-FPGA network called OPTWEB for efficient collectives. Since all end-to-end separate communication paths are statically established using onboard optics, raw block data can be transferred with simple link-level synchronization. Once each source FPGA assigns a communication stream to a path by its internal switch logic between memory-mapped and stream interfaces for remote direct memory access (RDMA), a one-hop transfer is provided. Since each FPGA performs input/output of the remote memory access between all FPGAs simultaneously, multiple RDMAs efficiently form collectives. The OPTWEB network provides 0.71-μsec start-up latency of collectives among multiple Intel Stratix 10 MX FPGA cards with onboard optics. The OPTWEB network consumes 31.4 and 57.7 percent of adaptive logic modules for aggregate 400-Gbps and 800-Gbps interconnects on a custom Stratix 10 MX 2100 FPGA, respectively. The OPTWEB network reduces by 40 percent the cost compared to a conventional packet network. Kenji Mizutani, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IEEE Trans. Computers | 3 |
| 2020 | Wavelength-routing interconnect "Optical Hub" for parallel computing systemsabstractTo solve the inter-node bandwidth bottleneck in parallel computing systems, we propose a wavelength-routing inter-node interconnect "Optical Hub". The physical topology of Optical Hub is star network, which leads to advantages in term of its throughput, size, energy consumption and life-time cost. The logical topology is full-mesh network, which leads to advantages in term of its latency and reliability. We introduced multi-path routings, which expand the effective bandwidth with the full-mesh topology such as Optical Hub, by replacing conventional MPI functions with our wrapper functions. We simulated execution time of parallel benchmarks on the parallel computing system with Optical Hub using parallel computing simulator SimGrid. As a result, we have confirmed that the parallel computing system with Optical Hub can achieve higher performance and lower energy consumption than conventional ones. We also examined the scalability of Optical Hub and showed that recursive hierarchical configurations of Optical Hub can save cable count drastically in case of large number of nodes against Dragonfly networks. Yutaka Urino, Kenji Mizutani, Tatsuya Usuki, Shigeru Nakamura |
HPC Asia | 1 |
| 2013 | High performance PIN Ge photodetector and Si optical modulator with MOS junction for photonics-electronics convergence systemabstractWe report on a high speed silicon-waveguide-integrated PIN Ge photodetector of 45 GHz bandwidth, and a high efficiency of 0.3 V·cm silicon optical modulator with a metal-oxide-semiconductor (MOS) junction by applying the low optical loss and high conductivity poly-silicon gate. These OE/EO devices enable low drive voltage of around 1V, which would contribute to a high density optical interposer of the future photonics-electronics convergence system. Junichi Fujikata, Masataka Noguchi, Makoto Miura, Masashi Takahashi, Shigeki Takahashi, Tsuyoshi Horikawa, Yutaka Urino, Takahiro Nakamura, Yasuhiko Arakawa |
ASP-DAC | 7 |