VLDB 2026 Research / reviewers in the wild / expert
Tomohiro Ueno
dblp:127/1212
· DBLP profile ↗
14ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-0228-0566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A case study in hardware specialization for Monte Carlo cross-section lookup
Kazutomo Yoshii, John R. Tramm, Bryce Allen, Tomohiro Ueno, Kentaro Sano, Andrew R. Siegel, Pete Beckman |
Parallel Comput. | 4 |
| 2025 | Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002 |
Future Gener. Comput. Syst. | 15 |
| 2024 | Flexible Systolic Array Platform on Virtual 2-D Multi-FPGA PlaneabstractSystolic arrays are a promising approach to achieving high-performance processing based on highly parallelized designs in various fields, such as AI and bioinformatics. Many previous studies have devoted considerable effort to exploring efficient circuit designs for specific processing. However, the increasing size of systolic arrays forces us to process increasingly large workloads by dividing them into smaller pieces. Therefore, we propose a systolic array platform based on a two-dimensional FPGA plane in which multiple FPGAs are connected by a virtual network. The systolic array realized by this system can be freely customized in shape and size according to the target. Distributed memory access through off-chip memory on each FPGA board and simple stream processing enable scalable performance. This paper presents a preliminary implementation based on the proposed systolic array platform and its performance evaluation. The evaluation results show that the proposed method improves the processing performance in proportion to the number of FPGAs. The results also show that the proposed platform is highly scalable due to the small circuit area required, and that the processing performance depends on the network bandwidth, which means that recent high-bandwidth FPGA boards can be expected to significantly improve the performance. Tomohiro Ueno, Emanuele Del Sozzo, Kentaro Sano |
HPC Asia | 1 |
| 2024 | Automated parallel execution of distributed task graphs with FPGA clustersabstractOver the years, Field Programmable Gate Arrays (FPGA) have been gaining popularity in the High Performance Computing (HPC) field, because their reconfigurability enables very fine-grained optimizations with low energy cost. However, the different characteristics, architectures, and network topologies of the clusters have hindered the use of FPGAs at a large scale. In this work, we present an evolution of OmpSs@FPGA, a high-level task-based programming model and extension to OmpSs-2, that aims at unifying all FPGA clusters by using a message-passing interface that is compatible with FPGA accelerators. These accelerators are programmed with C/C++ pragmas, and synthesized with High-Level Synthesis tools. The new framework includes a custom protocol to exchange messages between FPGAs, agnostic of the architecture and network type. On top of that, we present a new communication paradigm called Implicit Message Passing (IMP), where the user does not need to call any message-passing API. Instead, the runtime automatically infers data movement between nodes. We test classic message passing and IMP with three benchmarks on two different FPGA clusters. One is cloudFPGA, a disaggregated platform with AMD FPGAs that are only connected to the network through UDP/TCP/IP. The other is ESSPER, composed of CPU-attached Intel FPGAs that have a private network at the ethernet level. In both cases, we demonstrate that IMP with OmpSs@FPGA can increase the productivity of FPGA programmers at a large scale thanks to simplifying communication between nodes, without limiting the scalability of applications. We implement the N-body, Heat simulation and Cholesky decomposition benchmarks, and show that FPGA clusters get 2.6x and 2.4x better performance per watt than a CPU-only supercomputer for N-body and Heat. Juan Miguel De Haro Ruiz, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Tomohiro Ueno, Kentaro Sano, Burkhard Ringlein, François Abel, Beat Weiss |
Future Gener. Comput. Syst. | 5 |
| 2023 | ESSPER: Elastic and Scalable FPGA-Cluster System for High-Performance Reconfigurable Computing with Supercomputer FugakuabstractFPGA clusters have yet to be a mainstream of HPC, even for accelerators, and several challenges exist in their architecture and system organization. This work presents ESSPER, a flexible and scalable FPGA cluster prototype system for reconfigurable HPC to meet the concept of customizability, scalability, and interoperability with existing HPC systems. Based on our classification of FPGA cluster architectures, we propose a new category of FPGA clusters with a host-FPGA bridging network using software-bridged APIs for the use of remote FPGAs. We have designed, implemented, verified, and demonstrated a proof-of-concept system of ESSPER, as a functional extension of the supercomputer Fugaku. Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno |
HPC Asia | 4 |
| 2023 | VCSN: Virtual Circuit-Switching Network for Flexible and Simple-to-Operate Communication in HPC FPGA ClusterabstractFPGA clusters promise to play a critical role in high-performance computing (HPC) systems in the near future due to their flexibility and high power efficiency. The operation of large-scale general-purpose FPGA clusters on which multiple users run diverse applications requires flexible network topology to be divided and reconfigured. This paper proposes Virtual Circuit-Switching Network (VCSN) that provides an arbitrarily reconfigurable network topology and simple-to-operate network system among FPGA nodes. With virtualization, user logic on FPGAs can communicate with each other as if a circuit-switching network was available. This paper demonstrates that VCSN with 100 Gbps Ethernet achieves highly-efficient point-to-point communication among FPGAs due to its unique and efficient communication protocol. We compare VCSN with a direct connection network (DCN) that connects FPGAs directly. We also show a concrete procedure to realize collective communication on an FPGA cluster with VCSN. We demonstrate that the flexible virtual topology provided by VCSN can accelerate collective communication with simple operations. Furthermore, based on experimental results, we model and estimate communication performance by DCN and VCSN in a large FPGA cluster. The result shows that VCSN has the potential to accelerate gather communication up to about 1.97 times more than DCN. Tomohiro Ueno, Kentaro Sano |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Exploring Inter-tile Connectivity for HPC-oriented CGRA with Lower Resource UsageabstractThis research aims to explore the tradeoffs between routing flexibility and hardware resource usage, ultimately reducing the resource usage of our CGRA architecture while maintaining compute efficiency. we investigate statistics of connection usages among switch blocks for benchmark DFGs, propose several CGRA architecture with a reduced connection, and evaluate their hardware cost, routability of DFGs, and computational throughput for benchmarks. We found that the topology with horizontal plus diagonal connection saves about 30% of the resource usage while maintaining virtually the same routing flexibility as the full connectivity topology. Boma Anantasatya Adhi, Carlos Cortes, Tomohiro Ueno, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano |
FPT | 3 |
| 2022 | ESSPER: Elastic and Scalable System for High-Performance Reconfigurable Computing with Software-bridged APIsabstractMany-core CPUs and GPUs, present mainstream architectures for HPC, are facing difficulty in maintaining the same performance improvement rate because of the recent slow-down in the semiconductor scaling, the dark silicon problem, and wasteful mechanisms required for accelerating general-purpose computing such as a branch predictor and an out-of-order mechanism. Also, the power efficiency of HPC systems is significantly important to achieve higher performance. Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno |
FPT | 4 |
| 2021 | Virtual Circuit-Switching Network with Flexible Topology for High-Performance FPGA ClusterabstractAs the performance of high-end FPGAs has increased in recent years, it’s getting more important to construct an FPGA cluster for both improved processing performance and power efficiency in data centers and supercomputers. For higher utilization of FPGA resources for various applications, we require a flexible inter-FPGA network which provides various topologies appropriately to different applications while a conventional direct-connection network (DCN) provides only a fixed topology, such like a 2D torus. In this paper, we propose a virtual circuit-switching network (VCSN) for a large-scale FPGA cluster to have a flexible inter-FPGA network, where communication links connecting FPGAs are virtualized on the top of Ethernet frames. We can easily configure the VCSN topology optimized for the application by modifying the destination MAC addresses registered in a table of a frame encoder. We present its efficient protocol, hardware implementation, demonstration with 100Gbps Ethernet, and performance comparison with a conventional direct-connection network for FPGAs. We show that VCSN has higher but acceptable latency and slightly higher throughput in comparison with DCN, so that numerical simulation running with a ring of FPGAs achieves comparable performance for DCN. Tomohiro Ueno, Atsushi Koshiba, Kentaro Sano |
ASAP | 1 |
| 2018 | Performance Analysis of Hardware-Based Numerical Data Compression on Various Data FormatsabstractThe amount of data processed in high-performance computing has been growing rapidly. Accordingly, the cost of data movement in a large-scale computing system further increases, which has a huge effect on computing performance. To reduce the data movement cost, we have proposed hardware-based data compression for numerical data streams that can greatly reduce overhead has been proposed. Although our proposed hardware compressor works well for scientific computation in previous studies, the compression performance heavily depends on a type of target data. For practical use, we need to know the characteristics of the data compression and the relationship between data types and compression performance. In this paper, we clarify difference in compression performance among different data types by investigating hardware-based compression performance for data types of double, single, and half-precision floating-point and fixed-point with results of numerical simulation. We also propose an area-saving hardware design for double-precision floating-point data. Tomohiro Ueno, Kentaro Sano, Takashi Furusawa |
DCC | 1 |
| 2017 | Bandwidth Compression of Floating-Point Numerical Data Streams for FPGA-Based High-Performance ComputingabstractAlthough computational performance is often limited by insufficient bandwidth to/from an external memory, it is not easy to physically increase off-chip memory bandwidth. In this study, we propose a hardware-based bandwidth compression technique that can be applied to field-programmable gate array-- (FPGA) based high-performance computation with a logically wider effective memory bandwidth. Our proposed hardware approach can boost the performance of FPGA-based stream computations by applying a data compression technique to effectively transfer more data streams. To apply this data compression technique to bandwidth compression via hardware, several requirements must first be satisfied, including an acceptable level of compression performance and a sufficiently small hardware footprint. Our proposed hardware-based bandwidth compressor utilizes an efficient prediction-based data compression algorithm. Moreover, we propose a multichannel serializer and deserializer that enable applications to use multiple channels of computational data with the bandwidth compression. The serializer encodes compressed data blocks of multiple channels into a data stream, which is efficiently written to an external memory. Based on preliminary evaluation, we define an encoding format considering both high compression ratio and small hardware area. As a result, we demonstrate that our area saving bandwidth compressor increases performance of an FPGA-based fluid dynamics simulation by deploying more processing elements to exploit spatial parallelism with the enhanced memory bandwidth. Tomohiro Ueno, Kentaro Sano, Satoru Yamamoto |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2015 | A 58.3-to-65.4 GHz 34.2 mW sub-harmonically injection-locked PLL with a sub-sampling phase detectionabstractThis paper presents a low power and low noise sub-harmonically injection-locked PLL using a 20GHz sub-sampling PLL (SS-PLL) and a quadrature injection locked oscillator (QILO). Lower in-band phase noise and out-of-band phase noise have been achieved through the sub-sampling phase detection and sub-harmonic injection techniques, respectively. Implemented in a 65nm CMOS, this work can support all 60GHz channels and achieves a phase noise of -115dBc/Hz at 10MHz offset while consuming 20.2mW and 14mW from the 20GHz SS-PLL and the QILO, respectively. Teerachot Siriburanon, Tomohiro Ueno, Kento Kimura, Satoshi Kondo, Wei Deng 0001, Kenichi Okada 0001, Akira Matsuzawa |
ASP-DAC | 2 |
| 2015 | An HDL-synthesized gated-edge-injection PLL with a current output DACabstractThis paper presents a small area, low power, fully synthesizable PLL with a current output DAC and an interpolative-phase coupled oscillator using edge injection technique for on-chip clock generation. A prototype PLL is fabricated in a 65nm digital CMOS process, achieves a 1.7-ps integrated jitter at 0.9 GHz and consumes 0.78 mW leading to an FOM of -236.5 dB while only occupying an area of 0.0066 mm2. It achieves the best performance-area trade-off. Dongsheng Yang 0002, Wei Deng 0001, Tomohiro Ueno, Teerachot Siriburanon, Satoshi Kondo, Kenichi Okada 0001, Akira Matsuzawa |
ASP-DAC | 3 |
| 2014 | Bandwidth compression of multiple numerical data streams for high performance custom computingabstractBandwidth compression improves the performance of stream computing by enhancing an effective bandwidth. To apply the bandwidth compression to numerical applications such as numerical simulations, a compressor has to handle multiple data streams. In this paper, we describe a design of an FPGA-based bandwidth compressor for high performance stream computation. For synchronization of original data in multiple compressed streams with different bit-rate, we propose a data block transmission scheduler and explore a design space to reduce the size of their barrel shifters. Tomohiro Ueno, Ryo Ito, Kentaro Sano, Satoru Yamamoto |
ASAP | 1 |