EDBT 2026 Demo / reviewers in the wild / expert
Ryota Yasudo
dblp:157/0767
· DBLP profile ↗
22ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0003-2009-7105ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 11 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Scalable Accelerator for Local Score Computation of Structure Learning in Bayesian NetworksabstractA Bayesian network is a powerful tool for representing uncertainty in data, offering transparent and interpretable inference, unlike neural networks’ black-box mechanisms. To fully harness the potential of Bayesian networks, it is essential to learn the graph structure that appropriately represents variable interrelations within data. Score-based structure learning, which involves constructing collections of potentially optimal parent sets for each variable, is computationally intensive, especially when dealing with high-dimensional data in discrete random variables. Our proposed novel acceleration algorithm extracts high levels of parallelism, offering significant advantages even with reduced reusability of computational results. In addition, it employs an elastic data representation tailored for parallel computation, making it FPGA-friendly and optimizing module occupancy while ensuring uniform handling of diverse problem scenarios. Demonstrated on a Xilinx Alveo U50 FPGA, our implementation significantly outperforms optimal CPU algorithms and is several times faster than GPU implementations on an NVIDIA TITAN RTX. Furthermore, the results of performance modeling for the accelerator indicate that, for sufficiently large problem instances, it is weakly scalable, meaning that it effectively utilizes increased computational resources for parallelization. To our knowledge, this is the first study to propose a comprehensive methodology for accelerating score-based structure learning, blending algorithmic and architectural considerations. Ryota Miyagi, Ryota Yasudo, Kentaro Sano, Hideki Takase |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | A low-latency memory-cube network with dual diagonal mesh topology and bypassed pipelinesabstractSummary A memory cube network is an interconnection network composed of 3D stacked memories called memory cubes. By exploiting a packet switching, it can provide fast memory accesses to a large number of memory cubes. Although interconnection networks have been studied in many years for supercomputers and data centers, existing technologies are difficult to apply to memory cube networks. This is because the link length and the number of ports are limited, and hence the hop count increases. In this article, we propose a dual diagonal mesh (DDM), a layout‐oriented memory‐cube network. Furthermore, we propose the routing algorithm and the router architecture with bypassed pipelines for DDM. Our experimental results demonstrate that our routing and router architecture with bypassed pipelines reduces the memory access latency. We implement four router architectures and evaluate them with the traffic patterns derived from the NAS parallel benchmark. Masashi Oda, Kai Keida, Ryota Yasudo |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Bandit-based Variable Fixing for Binary Optimization on GPU Parallel ComputingabstractThis paper explores whether reinforcement learning is capable of enhancing metaheuristics for the quadratic unconstrained binary optimization (QUBO), which have recently attracted attention as a solver for a wide range of combinatorial optimization problems. In particular, we introduce a novel approach called the bandit-based variable fixing (BVF). The key idea behind BVF is to regard an execution of an arbitrary metaheuristic with a variable fixed as a play of a slot machine. Thus, BVF explores variables to fix with the maximum expected reward, and executes a metaheuristic at the same time. The bandit-based approach is then extended to fix multiple variables. To accelerate solving multi-armed bandit problem, we implement a parallel algorithm for BVF on a GPU. Our results suggest that our proposed BVF enhances original metaheuristics. Ryota Yasudo |
PDP | 1 |
| 2023 | Simple iterative trial search for the maximum independent set problem optimized for the GPUsabstractAbstract An independent set of a graph is a subset of the nodes such that no two nodes in it are adjacent. The maximum independent set (MIS) problem is an optimization problem to find a largest independent set. The main contribution of this article is to introduce a generic iterative trial search algorithm that we call iMIS for finding approximate solutions for the MIS problem. The generic algorithm iMIS is designed so that it can be implemented to run on GPUs very efficiently. Since the performance of the algorithm varies depending on its search strategy, we present a hybrid algorithm that combines three strategies of the iMIS such that best one of them for an input graph is automatically selected. We have implemented our hybrid algorithm to run on a multi‐GPU server with eight NVIDIA A100 GPUs. And evaluated the performance for 66 DIMACS benchmark graphs and 78 random graphs with up to 256M nodes. We have also evaluated the performance of three previously published algorithms for the MIS problem and an approach using Gurobi linear programming solver. The experimental results show that our hybrid algorithm can find larger independent sets for all 144 graphs compared to other methods. Tomohiro Imanaga, Koji Nakano, Ryota Yasudo, Yasuaki Ito, Yuya Kawamata, Ryota Katsuki, Yusuke Tabata, Takashi Yazane, Kenichiro Hamano |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | High-throughput FPGA implementation for quadratic unconstrained binary optimizationabstractAbstract Quadratic unconstrained binary optimization (QUBO) is a combinatorial optimization problem. Since various NP‐hard problems such as the traveling salesman problem can be formulated as a QUBO instance, QUBO is used with a wide range of applications. The main contribution of this article is to propose high‐throughput FPGA implementations for the QUBO solver. We perform the local search using different bit‐selection strategies based on the simulated annealing in the proposed implementation. The hardware is a pipeline structure with no pipeline hazards using multiple instances, where the bit‐flip operation is always performed every clock cycle. We implemented the proposed circuit on Xilinx UltraScale+ FPGA VU9P. The implementation result shows that the circuit can search solutions per second. Besides, by sharing the block RAM that stores a weight matrix, we implemented a dual annealer architecture that has two QUBO solvers into the FPGA. As a result, the dual annealer architecture can search solutions per second. Hiroshi Kagawa, Yasuaki Ito, Koji Nakano, Ryota Yasudo, Yuya Kawamata, Ryota Katsuki, Yusuke Tabata, Takashi Yazane, Kenichiro Hamano |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Designing low-diameter interconnection networks with multi-ported host-switch graphsabstractSummary A host‐switch graph was originally proposed as a graph that represents a network topology of a computer systems with 1‐port host computers and ‐port switches. It has been studied from both theoretical and practical aspects in terms of the diameter, the average shortest path length, and the performance of real applications. In recent high‐performance computing systems, however, a host computer is connected to multiple switches by using InfiniBand, NVSwitch, or Omni‐Path, and consequently they provide high bandwidths. Since a host‐switch graph cannot represent such systems, this article extends a host‐switch graph so that it can represent such systems. As a result, a host‐switch graph can include multi‐ported hosts. Furthermore, we propose to use multi‐port hosts for reducing the diameter. We show that the diameter minimization is equivalent to solving the degree diameter problem for bipartite graphs of diameter three. Our experimental results show that we can drastically reduce the diameter as well as increasing the bandwidth and improves performance of MPI applications by up to 162% as compared with networks with single‐ported hosts. Ryota Yasudo, Koji Nakano, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | Optimizing Application Mapping for Multi-FPGA Systems with Multi-ejection STDM SwitchesabstractMulti-FPGA systems have received an attention as a computing cluster for multi-access edge computing (MEC). Also, they can process time-critical jobs with their hardwired logic. For this purpose, the static time-division multiplexing (STDM) network is adopted because it enables to predict latency and bandwidth. However, the overall performance of the STDM network depends on the number of time slots. This paper proposes a new mapping tool that optimizes the application mapping so that the number of slots is minimized. Our tool handles multicasts and multi-ejection function which are effective techniques for STDM switches implemented on an FPGA cluster. For applications with all-to-all communication, our experimental results show that the tool reduces the number of time slots by 59–68% with both multicasts and multi-ejection switches. Kohei Ito, Ryota Yasudo, Hideharu Amano |
FPL | 2 |
| 2022 | Elastic Sample Filter: An FPGA-based Accelerator for Bayesian Network Structure Learningabstractproposed in 1985 by Judea Pearl [1], Ryota Miyagi, Ryota Yasudo, Kentaro Sano, Hideki Takase |
FPT | 2 |
| 2022 | GPU-accelerated scalable solver with bit permutated cyclic-min algorithm for quadratic unconstrained binary optimization
Ryota Yasudo, Koji Nakano, Yasuaki Ito, Ryota Katsuki, Yusuke Tabata, Takashi Yazane, Kenichiro Hamano |
J. Parallel Distributed Comput. | 1 |
| 2021 | Efficient implementations of Bloom filter using block RAMs and DSP slices on the FPGAabstractSummary This paper presents efficient FPGA implementations for the Bloom filter, in which a large set P of L‐byte patterns are registered beforehand. Our Bloom filter circuit performs the byte stream pattern test such that it receives an input byte stream t and outputs the bit stream in every clock cycle. Each bit of the output bit stream is 1 if an L‐byte sequence of t starting from the corresponding position is identical with one of the patterns in P. Our circuits use rolling hash functions to compute signatures of all patterns in P registered in Ultra RAMs of the Xilinx UltraScale+ FPGA VU9P. We present two types of implementations, DSP‐based implementation and RAM‐based implementation to compute rolling hash functions of L‐byte sequences using DSP slices and Block RAMs in the FPGA, respectively. The experimental results show that both DSP‐based and RAM‐based Bloom filter circuits for 4800K patterns of length 1024 can perform the byte stream pattern test for 1.1 Gps and 1.3 Gbps input byte streams, respectively, with false positive probability 10−12. Moreover, we can configure DSP‐based and RAM‐based Bloom filter circuits for 100K patterns to work for 54.9 Gbps and 62.2 Gbps input byte streams, respectively, with false positive probability 10−12. Takuma Wada, Naoki Matsumura, Ryota Yasudo, Koji Nakano, Yasuaki Ito |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Analytical Performance Estimation for Large-Scale Reconfigurable Dataflow PlatformsabstractNext-generation high-performance computing platforms will handle extreme data- and compute-intensive problems that are intractable with today’s technology. A promising path in achieving the next leap in high-performance computing is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators such as FPGAs, which have been shown to speed up compute-intensive tasks with reduced power consumption. However, assessing the feasibility of large-scale heterogeneous systems requires fast and accurate performance prediction. This article proposes Performance Estimation for Reconfigurable Kernels and Systems (PERKS), a novel performance estimation framework for reconfigurable dataflow platforms. PERKS makes use of an analytical model with machine and application parameters for predicting the performance of multi-accelerator systems and detecting their bottlenecks. Model calibration is automatic, making the model flexible and usable for different machine configurations and applications, including hypothetical ones. Our experimental results show that PERKS can predict the performance of current workloads on reconfigurable dataflow platforms with an accuracy above 91%. The results also illustrate how the modelling scales to large workloads, and how performance impact of architectural features can be estimated in seconds. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker, Ce Guo 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2020 | Dual-Plane Isomorphic Hypercube NetworkabstractWe propose a multi-plane isomorphic network that increases network throughput and reduces network latency by effectively configuring multi-plane networks. In the proposed network, each plane adopts the same graph topology but different switch-to-switch connections. We evaluate the dual-plane isomorphic hypercube network by graph analysis and cycle level simulation. Results of the graph analysis show that the dual-plane isomorphic 8-hypercube reduces the average shortest path length by 22% and improves throughput by 28% compared with the dual-plane hypercube. Similar improvements are confirmed from the results of the cycle level simulation. We also examine the dual-plane isomorphic folded-hypercube network. Finally, we discuss the effect of longer cable length caused by the isomorphic network on the network cost and latency. Takeo Hosomi, Ryota Yasudo, Michihiro Koibuchi, Shinji Shimojo |
HPC Asia | 2 |
| 2020 | Adaptive Bulk Search: Solving Quadratic Unconstrained Binary Optimization Problems on Multiple GPUsabstractThe quadratic unconstrained binary optimization (QUBO) is recently gathering attention in conjunction with quantum annealing (QA), since it is equivalent to finding the ground state of an Ising model. Due to the limitation of current QA systems, classical computers may outperform them. Researchers have thus been proposed to solve QUBO on FPGAs, GPUs, and special purpose processors. In this paper, we propose an adaptive bulk search (ABS), a framework for solving QUBO that can perform many searches in parallel on multiple GPUs. It supports fully-connected Ising models with up to 32k spins and 16-bit weights. In our ABS, a CPU host performs genetic algorithm (GA) while GPUs asynchronously perform local searches. A bottleneck for solving QUBO exists in the evaluation of the energy function, which requires computational cost for each solution. We show this can be reduced to in our ABS. The experimental results show that, with four NVIDIA GeForce RTX 2080 Ti GPUs, our framework can search up to 1.24 × 1012 solutions per second. We also show that our system quickly solves maximum cut and traveling salesman problems. Ryota Yasudo, Koji Nakano, Yasuaki Ito, Masaru Tatekawa, Ryota Katsuki, Takashi Yazane, Yoko Inaba |
ICPP | 1 |
| 2019 | Designing High-Performance Interconnection Networks with Host-Switch GraphsabstractThis paper aims at establishing a method for designing high-performance network topologies to bridge a gap between theoretical and practical studies. To this end, we present a novel graph called a host-switch graph, which consists of host vertices and switch vertices with maximum degree 1 and$r$, respectively. This graph represents a network topology of a practical parallel/distributed computer system with host computers connected by$r$-port switches. We discuss important metrics for designing high-performance interconnection networks: the host-to-host average shortest path length (h-ASPL) and the bisection width (BiW). In particular, we explore a method for constructing host-switch graphs with low h-ASPL and high BiW that connect the fixed number of hosts via any number of$r$-port switches. We demonstrate that the number of switches that provides the minimum h-ASPL can mathematically be approximated, and the minimum number of switches that provides a certain BiW can experimentally be approximated. On the basis of the approximations, we propose a randomized algorithm for searching host-switch graphs. We then apply the graphs to interconnection networks and compare them with typical network topologies. As compared with the torus, the dragonfly, and the fat-tree, our networks attain higher performance and smaller power and costs. Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Performance Prediction for Large-Scale Heterogeneous PlatformsabstractThis paper presents an approach for analysing, modelling and predicting application performance of large-scale heterogeneous platforms. Our approach combines analytical and statistical modelling techniques, and aims to: (1) identify and characterise code regions that are the most promising candidates to benefit from acceleration; (2) provide statistical models that predict application behaviour for unobserved inputs; and (3) predict performance gain with different system architectures. Ryota Yasudo, Ana Lucia Varbanescu, José Gabriel F. Coutinho, Wayne Luk, Hideharu Amano |
FCCM | 1 |
| 2018 | Performance Estimation for Exascale Reconfigurable Dataflow PlatformsabstractThe next generation high-performance computing platforms will need to support exascale computing. A promising path in achieving exascale is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators. However, assessing the feasibility of heterogeneous exascale systems requires fast and accurate performance prediction. This paper proposes PERKS, a novel performance estimation frame-work for reconfigurable dataflow platforms (RDPs). PERKS uses machine and application parameters to build an analytical model for predicting the performance of multi-accelerator systems. Moreover, model calibration is automatic, making the model flexible and usable for different machine configurations and applications. Our experimental results demonstrate that PERKS can predict the performance of current workloads and RDPs with an accuracy above 95%. We also demonstrate how the modelling scales to exascale workloads and exascale platforms. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker |
FPT | 1 |
| 2017 | HiRy: An Advanced Theory on Design of Deadlock-Free Adaptive Routing for Arbitrary TopologiesabstractRecently proposed irregular networks can reduce the latency for both on-chip and off-chip systems with a large number of computing nodes and thus can improve the performance of parallel application. However, these networks usually suffer from deadlocks in routing packets when using a naive minimal path routing algorithm. To solve this problem, we focus attention on a lately proposed theory that generalizes the turn model to maintain the network performance with deadlock-freedom. The theorems remain a challenge of applying themselves to arbitrary topologies including fully irregular networks. In this paper, we advance the theorems to completely general ones. To apply the idea of the turn model to arbitrary topologies, we introduce a concept of regions that define continuous directions of channels on an n-dimensional space. Moreover, we provide a feasible implementation of a deadlock-free routing method based on our advanced theorem. To reduce the latency and the number of required Virtual Channels (VCs) with this method, a heuristic approach is introduced to reduce the number of prohibited turns between channels. Experimental results show that the routing method based on our proposed theorem can improve the network throughput by up to 138 % compared to a conventional deterministic minimal routing method. Moreover, it can reduce the latency by up to 2.9 % compared to another fully adaptive routing method. Ryuta Kawano, Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICPADS | 2 |
| 2017 | Order/Radix Problem: Towards Low End-to-End Latency Interconnection NetworksabstractWe introduce a novel graph called a host-switch graph, which consists of host vertices and switch vertices. Using host-switch graphs, we formulate a graph problem called an order/radix problem (ORP) for designing low end-to-end latency interconnection networks. Our focus is on reducing the host-to-host average shortest path length (h-ASPL), since the shortest path length between hosts in a host-switch graph corresponds to the end-to-end latency of a network. We hence define ORP as follows: given order (the number of hosts) and radix (the number of ports per switch), find a host-switch graph with the minimum h-ASPL. We demonstrate that the optimal number of switches can mathematically be predicted. On the basis of the prediction, we carry out a randomized algorithm to find a host-switch graph with the minimum h-ASPL. Interestingly, our solutions include a host-switch graph such that switches have the different number of hosts. We then apply host-switch graphs to interconnection networks and evaluate them practically. As compared with the three conventional interconnection networks (the torus, the dragonfly, and the fat-tree), we demonstrate that our networks provide higher performance while the number of switches can decrease. Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
ICPP | 1 |
| 2017 | XYZ-Randomization using TSVs for Low-Latency Energy Efficient 3D-NoCsabstractIn this paper, we propose a method to design low latency and low energy networks for 3D Network-on-Chip (3D-NoC). Recent many-core processors require low-latency interconnection networks since the increasing number of cores limits the network performance. To achieve high performance in such many-core chips, small-world or random networks have been applied in the NoC field. However, the actual diameters and average shortest path lengths (ASPL) of these networks are far from the theoretical lower bound. In this work, we propose an approach based on the graph theory to design ultra low-latency topologies. We introduce a method to design a network that has low values of diameter and ASPL, with configurable upper bound of wire length, called opt ASPL. We also show that irregular topology, such as the topology used in opt ASPL, has a higher average energy consumption than general regular topology like 3D torus. In NoCs, energy budget and link length are limited, and thus such parameters must be carefully considered. Therefore, we introduce a multi-objective optimization for the ASPL and energy consumption called opt A/e which can obtain the Pareto optimal set useful for NoC designers. In a router with 64 nodes per chips and 4 chips stacked with a 3D-NoC, our proposed network optimized for energy consumption has a lower ASPL by 26.8% and a lower energy consumption by 10.9% compared to a 3D torus. Hiroshi Nakahara, Nguyen Anh Vu Doan, Ryota Yasudo, Hideharu Amano |
NOCS | 3 |
| 2017 | Scalable Networks-on-Chip with Elastic Links Demarcated by Decentralized RoutersabstractAs the number of cores on a chip increases, Networks-on-Chip (NoCs) that connect many cores would face long links to reduce hop counts. The long links become bottlenecks in terms of both energy and RC delays as technology advances. To alleviate the negative impact of long links, we propose decentralized routers for NoCs. A decentralized router consists of multiple submodules that are positioned on a link, and hence the long links are segmented. Furthermore, we illustrate the design of an entire network that uses decentralized routers to obtain a good tradeoff between hop counts and wire delays per hop. Decentralized routers are effective especially in high-radix topologies, such as the flattened butterfly, and energy-delay product is reduced by greater than 60 percent. As NoCs become larger and more complex, the benefit of the decentralized routers will become more significant. Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tadao Nakamura |
IEEE Trans. Computers | 1 |
| 2015 | On-Chip Decentralized Routers with Balanced Pipelines for Avoiding Interconnect BottleneckabstractTechnology scaling makes designers face difficulties dealing with wire delay of long global interconnects, especially for high-radix networks. In this context, we propose decentralization of on-chip packet routers. A decentralized router consists of submodules, each of which has particular functionality and they are scattered on a link, thereby long wires are segmented. Our starting point is from a conventional router architecture, and we illustrate four case studies to generalize our proposal. We also propose a new buffer design and how to balance pipelines of a router. A proof-of-concept is shown in 28-nm process technology. Our results demonstrate that the decentralization of an on-chip router enables Link Traversal (LT) stages to be eliminated, and the critical path delay is improved by up to 45% with the reduced area compared with a conventional router. As technology advances, the benefit of the decentralized routers become more substantial in the nano-scale era. Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tadao Nakamura |
NOCS | 1 |
| 2014 | Design of a low power NoC router using Marching Memory Through typeabstractPower consumption of Network-on-Chip (NoC) is becoming more important in many core processors. Input buffers utilized in routers consume a significant part of the total power of NoCs. In order to reduce this power consumption, a novel power efficient memory called Marching Memory Through type (MMTH) is introduced. By connecting transparent latches in tandem, MMTH achieves high speed operation with a low power consumption. MMTH, however, requires a certain overhead at read operation, and hence we propose a latency reduction scheme based on the look-ahead routing. The proposed router was designed in Renesas's 40nm process and compared with a standard router using conventional register-based FIFOs in terms of the network performance, application performance, and power consumption. The result of evaluation shows that the proposed router reduces the power consumption by 42.4% on average at 2GHz and the expense of only 0.5-2.0% performance overhead. Ryota Yasudo, Takahiro Kagami, Hideharu Amano, Yasunobu Nakase, Masashi Watanabe, Tsukasa Oishi, Toru Shimizu, Tadao Nakamura |
NOCS | 1 |