Hideki Takase

dblp:96/8189 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-2660-5927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables secure computation over encrypted data, but its computational cost remains a major obstacle to practical deployment. To mitigate this overhead, many studies have explored GPU acceleration for the CKKS scheme, which is widely used for approximate arithmetic. In CKKS, CKKS parameters are configured for each workload by balancing multiplicative depth, security requirements, and performance. These parameters significantly affect ciphertext size, thereby determining how the memory footprint fits within the GPU memory hierarchy. Nevertheless, prior studies typically apply their proposed optimization methods uniformly, without considering differences in CKKS parameter configurations. In this work, we demonstrate that the optimal GPU optimization strategy for CKKS depends on the CKKS parameter configuration. We first classify prior optimizations by two aspects of dataflows which affect memory footprint and then conduct both qualitative and quantitative performance analyses. Our analysis shows that even on the same GPU architecture, the optimal strategy varies with CKKS parameters with performance differences of up to 1.98 $\times$ between strategies, and that the criteria for selecting an appropriate strategy differ across GPU architectures.
Ai Nozaki, Takuya Kojima, Hiroshi Nakamura, Hideki Takase
COMPSAC4
2025 Exploring the Possibility of TypiClust for Low-Budget Federated Active Learning
abstract
Federated Active Learning (FAL) seeks to reduce the burden of annotation under the realistic constraints of federated learning by leveraging Active Learning (AL). As FAL settings make it more expensive to obtain ground truth labels, FAL strategies that work well in low-budget regimes, where the amount of annotation is very limited, are needed. In this work, we investigate the effectiveness of TypiClust, a successful low-budget AL strategy, in low-budget FAL settings. Our empirical results show that TypiClust works well even in low-budget FAL settings contrasted with relatively low performances of other methods, although these settings present additional challenges, such as data heterogeneity, compared to AL. In addition, we show that FAL settings cause distribution shifts in terms of typicality, but TypiClust is not very vulnerable to the shifts. We also analyze the sensitivity of TypiClust to feature extraction methods, and it suggests a way to perform FAL even in limited data situations.
Yuta Ono, Hiroshi Nakamura, Hideki Takase
COMPSAC3
2025 A Scalable Accelerator for Local Score Computation of Structure Learning in Bayesian Networks
abstract
A Bayesian network is a powerful tool for representing uncertainty in data, offering transparent and interpretable inference, unlike neural networks’ black-box mechanisms. To fully harness the potential of Bayesian networks, it is essential to learn the graph structure that appropriately represents variable interrelations within data. Score-based structure learning, which involves constructing collections of potentially optimal parent sets for each variable, is computationally intensive, especially when dealing with high-dimensional data in discrete random variables. Our proposed novel acceleration algorithm extracts high levels of parallelism, offering significant advantages even with reduced reusability of computational results. In addition, it employs an elastic data representation tailored for parallel computation, making it FPGA-friendly and optimizing module occupancy while ensuring uniform handling of diverse problem scenarios. Demonstrated on a Xilinx Alveo U50 FPGA, our implementation significantly outperforms optimal CPU algorithms and is several times faster than GPU implementations on an NVIDIA TITAN RTX. Furthermore, the results of performance modeling for the accelerator indicate that, for sufficiently large problem instances, it is weakly scalable, meaning that it effectively utilizes increased computational resources for parallelization. To our knowledge, this is the first study to propose a comprehensive methodology for accelerating score-based structure learning, blending algorithmic and architectural considerations.
Ryota Miyagi, Ryota Yasudo, Kentaro Sano, Hideki Takase
ACM Trans. Reconfigurable Technol. Syst.4
2024 Ph.D. Project: Field-Programmable Computing for Bayesian Network Structure Learning
abstract
A Bayesian network is a powerful and versatile framework for modeling uncertainty in data, providing trans-parent and interpretable inferences. To maximize the potential of Bayesian networks, it is crucial to accurately learn the graph structure that captures the interrelations among variables within data. However, the number of potential graphs grows exponentially with the number of variables, making the structure learning of large Bayesian networks challenging. Our project aims to establish scalable acceleration in structure learning while demonstrating the potential of field-programmable custom computing to achieve energy efficiency and high performance.
Ryota Miyagi, Hideki Takase
FCCM2
2023 F2MKD: Fog-enabled Federated Learning with Mutual Knowledge Distillation
abstract
Federated learning (FL) is a promising technology for achieving privacy-preserving distributed learning. Although most existing studies on FL have provided a single global model for all clients, the resulting single model cannot handle heterogeneous local environments. Therefore, we propose a novel FL framework called fog-enabled federated learning with mutual knowledge distillation (F2MKD), in which the server and each client are ascribed to distinctive models. The remarkable feature of F2MKD is the data collection at the fog servers while ensuring data confidentiality among the fog servers. This study paves the new way for client-friendly, collaborative distributed learning. The code is available at https://github.com/d3-ai/F2MKD.
Yusuke Yamasaki, Hideki Takase
CCNC2
2023 An asynchronous federated learning focusing on updated models for decentralized systems with a practical framework
abstract
Federated learning (FL), a machine learning technique that preserves privacy by aggregating models from each device without exchanging personal data, has gained significant interest. This paper aims to establish an efficient and practical asynchronous FL method for decentralized systems. In asynchronous decentralized FL, the order of learning and aggregation is arbitrary. First, we explore the impact of this ordering on FL. Then, we propose to utilize the version information regarding model updates. Our strategy is to aggregate only the updated models after the previous round to improve the quality of the device’s model by avoiding the reaggregation of older models. Moreover, we also design a new practical framework for asynchronous decentralized FL by extending Flower framework. Our framework realizes effective communication ability by leveraging gRPC communication and thus can be applied to practical systems without the central server. Our evaluation shows the effectiveness of our methods that aggregate only the updated models from other devices. In addition, we show the impact of ordering on learning and aggregation according to situations.
Yusuke Kanamori, Yusuke Yamasaki, Shintaro Hosoai, Hiroshi Nakamura, Hideki Takase
COMPSAC5
2023 Resource Allocation Methods among Server Clusters in a Resource Permeating Distributed Computing Platform for 5G Networks
abstract
With the spread and development of 5G technology, network configurations with MEC (Multi-Access Edge Computing) servers in the vicinity of 5G base stations are becoming more common. We have been researching and developing Giocci, a resource permeating distributed processing platform that offloads computation tasks on end devices to MEC servers and cloud servers. In this paper, we propose resource allocation methods to efficiently determine the server to which tasks are allocated in a network configuration that includes MEC servers. In order to construct the proposed methods, we first model the main functions of Giocci and define the resource allocation problem for this work. There are four proposed methods according to each objective and priority; a prioritized allocation by the average number of waiting tasks, by the communication delay, by the task response time, and by the cost of using computing resources. We initially implement these methods assuming that they are task allocation functions in Giocci. Experimental evaluations demonstrate that they could achieve appropriate resource allocation results for each objective. This research contributes to the smooth allocation of computational resources in 5G networks including MEC servers.
Daisuke Sasaki, Hiroki Kashiwazaki, Mitsuhiro Osaki, Kazuma Nishiuchi, Ikuo Nakagawa, Shunsuke Kikuchi, Yutaka Kikuchi, Shintaro Hosoai, Hideki Takase
COMPSAC9
2023 ILP Based Mapping for Elastic CGRAs
abstract
In recent years, the emergence of deep learning and the need for big data analysis have created a demand for computers with high computational performance and energy efficiency. Since conventional ASICs and general-purpose CPUs cannot meet this requirement, domain-specific architectures that constrain applications are currently the focus of attention. Coarse-Grained Reconfigurable Architecture (CGRA), one of the domain-specific architectures, is attracting attention because it is superior to CPUs and FPGAs in terms of computational performance and power efficiency [1]. As depicted in Fig. 1, CGRA comprises a two-dimensional array of Processing Elements (PEs) and provides flexibility to change the instructions executed on the PEs and the connections between them, depending on the software being executed. On the other hand, the mapping problem for CGRA is an NP-complete problem, and there exist tradeoffs between solution accuracy and execution time. For real time applications, it is indispensable to shorten the mapping time. Thus, in this study, we propose an Integer Linear Programming (ILP) based mapping method for Elastic CGRA that aims to shorten mapping time while preserving solution accuracy. We have also implemented and preliminary evaluated the proposed method.
Makoto Saito, Takuya Kojima, Hideki Takase, Hiroshi Nakamura
RTCSA3
2022 Elastic Sample Filter: An FPGA-based Accelerator for Bayesian Network Structure Learning
abstract
proposed in 1985 by Judea Pearl [1],
Ryota Miyagi, Ryota Yasudo, Kentaro Sano, Hideki Takase
FPT4
2021 Zytlebot : FPGA integrated ros-based autonomous mobile robot
abstract
The FPT 2021 Design Competition aims to improve the technology of utilizing FPGA and achieve level-5 autonomous driving. We developed FPGA Integrated ROS-Based autonomous mobile robot, ZytleBot, for the competition. ZytleBot collects environmental information with CMOS cameras, recognizes environments, decides its action on programmable SoC, and controls its actuator. As a result, ZytleBot can run road model courses, detect and adequately deal with traffic lights and obstacles. We used the robot development platform TurtleBot3 and the robot middleware ROS to develop the robot system quickly. In addition, we utilize FPGA to accelerate road-images processing and traffic lights recognition using the HOG feature and SVM classifier. As a result, traffic lights recognition with FPGA is 270 times faster than those only with CPU.
Ryota Miyagi, Naofumi Takagi, Sho Kinoshista, Masashi Oda, Hideki Takase
FPT5
2019 ZytleBot: FPGA Integrated Development Platform for ROS Based Autonomous Mobile Robot
abstract
ZytleBot is an autonomous driving robot with an FPGA-integrated development platform that uses the Xilinx programmable system-on-chip (SoC). ZytleBot can run a course, turn right/left at intersections, avoid obstacles, detect traffic signals, and stop. All judgments and calculations necessary for driving are performed on the embedded system mounted on the robot. In ZytleBot, the main autonomous driving system uses the Robot Operating System (ROS) running on a CPU, and high-load processing is offloaded to the FPGA to enable real-time operation. The FPGA preprocesses road surface images acquired from the camera and detects traffic signals. We demonstrate the running of ZytleBot on a miniature course to win the FPT'18 FPGA design competition 1. We also provide ZytleBot as a platform for the efficient development of FPGA-integrated ROS robots.
Yasuhiro Nitta, Sou Tamura, Hideki Takase
FPL3
2018 Design concept of a lightweight runtime environment for robot software components onto embedded devices: work-in-progress
abstract
Although ROS (Robotic Operating System) has attracted attention to enhance the productivity of robot software development, it is necessary to adopt the device with high function and large power consumption enough to install Linux. This paper designs a lightweight runtime environment of ROS nodes onto mid-range embedded devices. Our environment, that is named to mROS, consists of a real-time OS and TCP/IP protocol stack to provide a tiny ROS communication library. mROS provides the connectivity to host and other ROS nodes with the native ROS network protocol. One of advantages for mROS is that native ROS nodes can be ported from Linux-based systems to RTOS-based systems since APIs with the same name of native ROS can be used in the embedded program. Experimental results validate that the performance requirement of mROS can be achieved for the construction of distributed robot systems.
Hideki Takase, Tomoya Mori, Kazuyoshi Takagi, Naofumi Takagi
EMSOFT1
2018 A Study on Introducing FPGA to ROS Based Autonomous Driving System
abstract
We are developing an autonomous driving robot using programmable SoC. The robot under development does not communicate with the external PC and performs all judgment and control on the board mounted on the robot. We aim to realize a built-in autonomous driving system with low power consumption and high performance by offloading high-load processing with the FPGA. At present, it is used only for acquiring camera images on the FPGA, but we are planning to do hardware implementation of the system constructed by software. In addition, we used ROS (Robot Operating System) to construct the robot's autonomous driving system, and the components to be developed are reusable. This document describes the detailed configuration and future prospect of the robot currently under development.
Yasuhiro Nitta, Sou Tamura, Hideki Takase
FPT3
2013 A Buffering Method for Parallelized Loop with Non-Uniform Dependencies in High-Level Synthesis
Akihiro Suda, Hideki Takase, Kazuyoshi Takagi, Naofumi Takagi
ICA3PP (1)2
2011 An integrated optimization framework for reducing the energy consumption of embedded real-time applications
Hideki Takase, Lovic Gauthier, Hirotaka Kawashima, Noritoshi Atsumi, Tomohiro Tatematsu, Yoshitake Kobayashi, Shunitsu Kohara, Takenori Koshiro, Tohru Ishihara, Hiroyuki Tomiyama, Hiroaki Takada
ISLPED1
2010 Minimizing inter-task interferences in scratch-pad memory usage for reducing the energy consumption of multi-task systems
abstract
This paper presents a new technique for reducing the energy consumption of a multi-task system by sharing its scratchpad memory (SPM) space among the tasks. With this technique, tasks can interfere by using common areas of the SPM. However, this requires to update these areas during context switches, which involves considerable overheads. Hence, an integer linear programming formulation is used at compile time for finding the best assignment of memory objects to the SPM and their respective locations inside it. Experiments show that the technique achieves up to 85% energy reduction with 8Kb of SPM and surpasses other sharing approaches.
Lovic Gauthier, Tohru Ishihara, Hideki Takase, Hiroyuki Tomiyama, Hiroaki Takada
CASES3
2010 Partitioning and allocation of scratch-pad memory for priority-based preemptive multi-task systems
abstract
Scratch-pad memory has been employed as a partial or entire replacement for cache memory due to its better energy efficiency. In this paper, we propose scratch-pad memory management techniques for priority-based preemptive multi-task systems. Our techniques are applicable to a real-time environment. The three methods which we propose, i.e., spatial, temporal, and hybrid methods, bring about effective usage of the scratch-pad memory space, and achieve energy reduction in the instruction memory subsystems. We formulate each method as an integer programming problem that simultaneously determines (1) partitioning of scratch-pad memory space for the tasks, and (2) allocation of program code to scratch-pad memory space for each task. It is remarkable that periods and priorities of tasks are considered in the formulas. Additionally, we implement an RTOS-hardware cooperative support mechanism for a runtime code allocation to the scratch-pad memory space. We have made the experiments with the fully functional real-time operating system. The experimental results with four task sets have demonstrated the effectiveness of our techniques. Up to 73% energy reduction compared to a standard method was achieved.
Hideki Takase, Hiroyuki Tomiyama, Hiroaki Takada
DATE1