EDBT 2026 Demo / reviewers in the wild / expert
Takefumi Miyoshi
dblp:12/3894
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0009-0536-9644ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Implementation and Evaluation of a Cryogenic DSP in 22-nm with Floating-Point Arithmetic for Estimating Qubit States
Yuki Koyama 0005, Takashi Imagawa, Ryo Kishida, Takefumi Miyoshi, Kazutoshi Kobayashi |
ISCAS | 4 |
| 2025 | Qu-Trefoil: Large-Scale Quantum Circuit Simulator Working on FPGA With SATA StoragesabstractQuantum circuits are fundamental components of quantum computing, and state-vector-based quantum circuit simulation is a widely used technique for tracking qubit behavior throughout circuit evolution. However, simulating a circuit with$n$qubits requires$2^{n+4}$bytes of memory, making simulations of more than 40 qubits feasible only on supercomputers. To address this limitation, we propose the Qu-Trefoil, a system designed for large-scale quantum circuit simulations on an FPGA-based platform called Trefoil. Trefoil is a multi-FPGA system connected to eight storage subsystems, each equipped with 32 SATA disks. Qu-Trefoil integrates a suite of HLS-based universal quantum gates, including Clifford gates (Hadamard (H), Pauli-Z (Z), Phase (S), Controlled-NOT (CNOT)), the T gate, and unitary matrix computation, along with HDL-designed modules for system-wide integration. Our extensive evaluation demonstrates the system's robustness and flexibility, covering quantum gate performance, chunk size, disk extensibility, and efficiency across different SATA generations. We successfully simulated quantum circuits with over 43 qubits, which required more than 128 TB of memory, in approximately 3.72 to 13.06 hours on a single storage subsystem equipped with one FPGA. This achievement represents a significant milestone in the advancement of quantum computing simulations. Furthermore, thanks to its unique architecture, Qu-Trefoil is more accessible, flexible, and cost-efficient than other existing simulators for large-scale quantum circuit simulations, making it a viable option for researchers with limited access to supercomputers. Kaijie Wei, Hideharu Amano, Ryohei Niwase, Yoshiki Yamaguchi, Takefumi Miyoshi |
IEEE Trans. Computers | 5 |
| 2025 | Simulation environment for reconfigurable virtual accelerators using a field programmable gate array development environment
Shunya Kawai, Eriko Maeda, Kazuki Yaguchi, Yasunori Osana, Takefumi Miyoshi, Hironori Nakajo |
J. Supercomput. | 5 |
| 2025 | Preliminary evaluation of SHAVER: sharing vector registers with an accelerator
Tomoaki Tanaka, Michiya Kato, Yasunori Osana, Takefumi Miyoshi, Jubee Tada, Kiyofumi Tanaka, Hironori Nakajo |
J. Supercomput. | 4 |
| 2022 | FPL Demo: A Flexible and Scalable Quantum-Classical Interface based on FPGAsabstractThis demonstration shows a Quantum-Classical interface (QC-IF) for quantum computing implemented on multiple FPGAs. Quantum computers need a controller to transmit/receive microwave to/from quantum devices. In order to explore various quantum devices, the controller requires flexibility to transmit/receive various wave shapes. In addition to that, scalability is also required for large-scale quantum computers. FPGAs are attractive platforms; however, some challenges exist in implementing QC-IF on FPGAs in terms of required specifications from physics, such as treating data and realizing scalability. This work demonstrates an implementation of QC-IF on FPGA with high bandwidth memory to treat large-volume data in high throughput. Furthermore, the IEEE1588-like clock synchronization mechanism is implemented to make multiple FPGAs synchronized. Takefumi Miyoshi, Keisuke Koike, Shinichi Morisaka, Hidehisa Shiomi, Kazuhisa Ogawa, Yutaka Tabuchi, Makoto Negoro |
FPL | 1 |
| 2013 | A fast handshake join implementation on FPGA with adaptive merging networkabstractOne of a critical design issues for implementing handshake-join hardware is result collection performed by a merging network. To address the issue, we introduce an adaptive merging network. Our implementation achieves over 3 million tuples per second when the selectivity is 0.1. The proposed implementation attains up to 5.2x higher throughput than original handshake-join hardware. In this demonstration, we apply the proposed technique to filter out malicious packets from packet streams. To the best of our knowledge, our system is the fastest handshake join implementation on FPGA. Yasin Oge, Takefumi Miyoshi, Hideyuki Kawashima, Tsutomu Yoshinaga |
SSDBM | 2 |
| 2011 | A Coarse Grain Reconfigurable Processor Architecture for Stream Processing EngineabstractThis paper proposes a processor architecture for DR-SPE, a dynamic reconfigurable stream processing engine. DR-SPE is special-purpose hardware for stream data processing, which achieves high processing performance by exploiting parallelism in the target query. It also handles query registration and execution order of operations at runtime. Available operations in DR-SPE are the same as those in Streams on Wires. In this paper, DR-SPE is implemented on a FPGA XC6VLX240T-1, and its performance is evaluated. The results of the evaluation show that DR-SPE achieves register modification within 506 μsec when the configuration path is driven at 1 Mbps, which is not achieved by Streams on Wires. DR-SPE also achieves flexibility and can support complicated queries by providing 10 × 10 operation units tiled onto an FPGA. DR-SPE achieves comparable operation throughput with Streams on Wires at the expense of requiring more LUTs. Takefumi Miyoshi, Hideyuki Kawashima, Yuta Terada, Tsutomu Yoshinaga |
FPL | 1 |
| 2009 | A Study of an Infrastructure for Research and Development of Many-Core ProcessorsabstractMany-core processors which have thousands of cores on a chip will be realized. We developed an infrastructure which accelerates the research and development of such many-core processors. This paper describes three main elements provided by our infrastructure. The first element is the definition of simple many-core processor architecture called M-Core. The second is SimMc, a software simulator of M-Core. The third is the software library MClib which helps the development of application programs for M-Core. The simulation speed of SimMc and the parallelization efficiency of M-Core are evaluated using some benchmark programs. We show that our infrastructure accelerates the research and development of many-core processors. Koh Uehara, Shimpei Sato, Takefumi Miyoshi, Kenji Kise |
PDCAT | 3 |