EDBT 2026 Demo / reviewers in the wild / expert
Takuya Kojima
dblp:33/5397
· DBLP profile ↗
14ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables secure computation over encrypted data, but its computational cost remains a major obstacle to practical deployment. To mitigate this overhead, many studies have explored GPU acceleration for the CKKS scheme, which is widely used for approximate arithmetic. In CKKS, CKKS parameters are configured for each workload by balancing multiplicative depth, security requirements, and performance. These parameters significantly affect ciphertext size, thereby determining how the memory footprint fits within the GPU memory hierarchy. Nevertheless, prior studies typically apply their proposed optimization methods uniformly, without considering differences in CKKS parameter configurations. In this work, we demonstrate that the optimal GPU optimization strategy for CKKS depends on the CKKS parameter configuration. We first classify prior optimizations by two aspects of dataflows which affect memory footprint and then conduct both qualitative and quantitative performance analyses. Our analysis shows that even on the same GPU architecture, the optimal strategy varies with CKKS parameters with performance differences of up to 1.98 $\times$ between strategies, and that the criteria for selecting an appropriate strategy differ across GPU architectures. Ai Nozaki, Takuya Kojima, Hiroshi Nakamura, Hideki Takase |
COMPSAC | 2 |
| 2023 | ILP Based Mapping for Elastic CGRAsabstractIn recent years, the emergence of deep learning and the need for big data analysis have created a demand for computers with high computational performance and energy efficiency. Since conventional ASICs and general-purpose CPUs cannot meet this requirement, domain-specific architectures that constrain applications are currently the focus of attention. Coarse-Grained Reconfigurable Architecture (CGRA), one of the domain-specific architectures, is attracting attention because it is superior to CPUs and FPGAs in terms of computational performance and power efficiency [1]. As depicted in Fig. 1, CGRA comprises a two-dimensional array of Processing Elements (PEs) and provides flexibility to change the instructions executed on the PEs and the connections between them, depending on the software being executed. On the other hand, the mapping problem for CGRA is an NP-complete problem, and there exist tradeoffs between solution accuracy and execution time. For real time applications, it is indispensable to shorten the mapping time. Thus, in this study, we propose an Integer Linear Programming (ILP) based mapping method for Elastic CGRA that aims to shorten mapping time while preserving solution accuracy. We have also implemented and preliminary evaluated the proposed method. Makoto Saito, Takuya Kojima, Hideki Takase, Hiroshi Nakamura |
RTCSA | 2 |
| 2023 | A Variation-Aware MTJ Store Energy Estimation Model for Edge Devices With Verify-and-Retryable Nonvolatile Flip-FlopsabstractWhile the spin-transfer torque (STT) magnetic tunnel junction (MTJ) is a promising technique for enabling nonvolatile flip-flops (NVFFs) to perform power gating to reduce leakage power without any data losses, the large store energy (the energy to make a store operation) of MTJs needs to be addressed. The nonvolatile cool mega array series is an edge-oriented coarse-grained reconfigurable accelerator that implements an improved MTJ-based NVFF with a verify-and-retryable store method that should ideally reduce the store energy under the presence of the switching time variation originating from the stochastic nature of the MTJs. However, the energy reduction effect of the method has not been formulated or evaluated thoroughly enough to make the best use of the method in actual applications. In this study, we propose an analytical model to estimate the store energy in typical operational conditions under the assumption of switching time variations following the normal distributions based on the measurements of a real chip fabricated with a 40-nm perpendicular MTJ/CMOS hybrid process. In contrast to the tedious measurement on each different condition, the proposed model allows for an instantaneous determination of the best storing method for minimizing the store energy, with an energy reduction of up to 69% compared with a conventional one-time attempt storing method. This model is expected to be used for system-level energy simulations and, ultimately, for design explorations in pursuit of energy-optimized memory. Aika Kamei, Hideharu Amano, Takuya Kojima, Daiki Yokoyama, Kimiyoshi Usami, Keizo Hiraga, Kenta Suzuki, Kazuhiro Bessho |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPCabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable architectures that inherit the performance and usability properties of Central Processing Units (CPUs) and the reconfigurability aspects of Field-Programmable Gate Arrays (FPGAs). Historically, CGRAs have been successfully used to accelerate embedded applications and are today also being considered to accelerate High-Performance Computing (HPC) applications in future supercomputers. However, embedded systems and supercomputers are two vastly different domains with different applications and constraints, and it is today not fully understood what CGRA design decisions adequately cater to the HPC market. One such unknown design decision is regarding the interconnect that facilitates intra-CGRA communication. Today, intra-CGRA communication comes in two flavors: using routers closely embedded into the compute units or using discrete routers outside the compute units. The former trades flexibility for a reduction in hardware cost, while the latter has greater flexibility but is more resource hungry. In this paper, we aspire to understand which of both designs best suits the CGRA HPC segment. We extend our previous methodology, which consists of both a parameterized CGRA design and an OpenMPcapable compiler, to accommodate both types of routing designs, including verification tests using RTL simulation. Our results show that the discrete router design can facilitate better use of processing elements (PEs) compared to embedded routers and can achieve up to 79.27% reduction in unnecessary PE occupancy for an aggressively unrolled stencil kernel on a 18 × 16 CGRA at a (estimated) hardware resource overhead cost of 6.3x. This reduction in PE occupancy can be used, for example, to exploit instruction-level parallelism (ILP) through even more aggressive unrolling. Boma Anantasatya Adhi, Carlos Cortes, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano |
CLUSTER | 4 |
| 2022 | Exploring Inter-tile Connectivity for HPC-oriented CGRA with Lower Resource UsageabstractThis research aims to explore the tradeoffs between routing flexibility and hardware resource usage, ultimately reducing the resource usage of our CGRA architecture while maintaining compute efficiency. we investigate statistics of connection usages among switch blocks for benchmark DFGs, propose several CGRA architecture with a reduced connection, and evaluate their hardware cost, routability of DFGs, and computational throughput for benchmarks. We found that the topology with horizontal plus diagonal connection saves about 30% of the resource usage while maintaining virtually the same routing flexibility as the full connectivity topology. Boma Anantasatya Adhi, Carlos Cortes, Tomohiro Ueno, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano |
FPT | 5 |
| 2022 | An efficient compilation of coarse-grained reconfigurable architectures utilizing pre-optimized sub-graph mappingsabstractIn recent years, IoT devices have become widespread, and energy-efficient coarse-grained reconfigurable architectures (CGRAs) have attracted attention. CGRAs comprise several processing units called processing elements (PEs) arranged in a two-dimensional array. The operations of PEs and the interconnections between them are adaptively changed depending on a target application, and this contributes to a higher energy efficiency compared to general-purpose processors. The application kernel executed on CGRAs is represented as a data flow graph (DFG), and CGRA compilers are responsible for mapping the DFG onto the PE array. Thus, mapping algorithms significantly influence the performance and power efficiency of CGRAs as well as the compile time. This paper proposes POCOCO, a compiler framework for CGRAs that can use pre-optimized subgraph mappings. This contributes to reducing the compiler optimization task. To leverage the subgraph mappings, we extend an existing mapping method based on a genetic algorithm. Experiments on three architectures demonstrated that the proposed method reduces the optimization time by 48%, on an average, for the best case of the three architectures. Ayaka Ohwada, Takuya Kojima, Hideharu Amano |
PDP | 2 |
| 2022 | Mapping-Aware Kernel Partitioning Method for CGRAs Assisted by Deep LearningabstractCoarse-grained reconfigurable architectures (CGRAs) provide high energy efficiency with word-level programmability rather than bit-level ones such as FPGAs. The coarser reconfigurability brings about higher energy efficiency and reduces the complexity of compiler tasks compared to the FPGAs. However, application mapping process for CGRAs is still time-consuming. When the compiler tries to map a large and complicated application data-flow-graph(DFG) onto the reconfigurable fabric, it tends to result in inefficient resource use or to fail in mapping. In case of failure, the compiler must divide it into several sub-DFGs and goes back to the same flow. In this work, we propose a novel partitioning method based on a genetic algorithm to eliminate the unmappable DFGs and improve the mapping quality. In order not to generate unmappable sub-DFGs, we also propose an estimation model which predicts the mappability and resource requirements using a DGCNN (Deep Graph Convolutional Neural Network). The genetic algorithm with this model can seek the most resource-efficient mapping without the back-end mapping process. Our model can predict the mappability with more than 98% accuracy and resource usage with a negligible error for two studied CGRAs. Besides, the proposed partitioning method demonstrates 53-75% of memory saving, 1.28-1.39x higher throughput, and better mapping quality over three comparative approaches. Takuya Kojima, Ayaka Ohwada, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | GenMap: A Genetic Algorithmic Approach for Optimizing Spatial Mapping of Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) are expected to be used for embedded systems, Internet of Things (IoT) devices, and edge computing thanks to their high-energy efficiency and programmability. In essence, a CGRA is an array of numerous processing elements. To exploit this abundant computation resource, a compiler for CGRAs has to fulfill more tasks compared that for general-purpose processors. Therefore, many studies have proposed optimization methods, especially for application mapping, because the performance and energy efficiency strongly depend on optimization at compile time. However, many works focus only on performance improvement or resource minimization, although such optimization objectives are not always appropriate when considering various use cases. In this work, we propose GenMap, an application mapping framework using multiobjective optimization based on a genetic algorithm so that users can set optimization criteria as needed. Besides, it provides aggressive power optimization using our dynamic power model and leakage minimization technique. The proposed method is applied to three fabricated CGRA chips for evaluation. Experimental results show that GenMap achieves 15.7% reduction of wire length while keeping processing element utilization when compared with conventional methods. In addition, according to real chip experiments, 12.1%-46.8% of energy consumption is reduced, and up to $2\times $ speedup is archived for several architectures when compared with other two approaches. Takuya Kojima, Nguyen Anh Vu Doan, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Demonstration of Low Power Stream Processing Using a Variable Pipelined CGRAabstractVPCMA (Variable Pipelined Cool Mega Array) is a low power CGRA (Coarse-Grained Reconfigurable Architecture) which we previously proposed in [1]. CC-SOTB2 is a real chip implementation of the VPCMA using Renesas 65-nm SOTB technology [2]. In this demonstration, we will show the power consumption of the CC-SOTB2 while performing a real image processing. Takuya Kojima, Naoki Ando, Yusuke Matsushita 0001, Hideharu Amano |
FPL | 1 |
| 2018 | A Configuration Data Multicasting Method for Coarse-Grained Reconfigurable ArchitecturesabstractThis paper proposes a novel configuration data compression technique for coarse-grained reconfigurable architectures (CGRAs). The proposed technique is based on a multicast configuration technique called RoMultiC, which reduces the configuration time by multicasting the same data to multiple PEs(Processing Elements) with two bit-maps. Scheduling algorithms for an optimizing the order of multicasting have been proposed. In general, configuration data for CGRAs can be divided into some fields like machine code formats. The proposed scheme confines a part of fields for multicasting so that the possibility of multicasting more PEs can be increased. This paper analyzes algorithms to find a configuration pattern which maximizes the number of multicasted PEs. We implemented the proposed scheme to CMA (Cool Mega Array), a straight forward CGRA as a case study. Experimental results show that the proposed method achieves 40.0% smaller configuration for an image processing application at maximum. Furthermore, it achieves 35.6% reduction of the power consumption for the configuration with a negligible area overhead. Takuya Kojima, Hideharu Amano |
FPL | 1 |
| 2017 | Body bias optimization for variable pipelined CGRAabstractVariable Pipeline Cool Mega Array (VPCMA) is an low power Coarse Grained Reconfigurable Architecture (CGRA) based on the concept of CMA (Cool Mega Array). It implements a pipeline structure that can be configured depending on performance requirements, and the silicon on thin buried oxide (SOTB) technology that allows to control its body bias voltage to balance performance and leakage power. In this paper, we propose a methodology to optimize exactly with an Integer Linear Program the VPCMA body bias while considering simultaneously its variable pipeline structure. For the studied applications, we evaluate that it is possible to achieve an average reduction of energy consumption of 19.3% and 11.8% when compared to respectively the zero bias (without body bias control) and the uniform (control of the whole PE array) cases, while respecting performance constraints. Besides, with appropriate body bias control, it is possible to extend the possible performance, hence enabling broader trade-off analyzes between consumption and performance. These promising results show that applying an adequate optimization technique for the body bias control while simultaneously considering pipeline structures can not only enable further power reduction than previous methods, but also allow more trade-off analysis possibilities. Takuya Kojima, Naoki Ando, Hayate Okuhara, Nguyen Anh Vu Doan, Hideharu Amano |
FPL | 1 |
| 2013 | Impression survey of the emotion expression humanoid robot with mental model based dynamic emotionsabstractThis paper describes the implementation in a walking humanoid robot of a mental model, allowing the dynamical change of the emotional state of the robot based on external stimuli; the emotional state affects the robot decisions and behavior, and it is expressed with both facial and whole-body patterns. The mental model is applied to KOBIAN-R, a 65-DoFs whole body humanoid robot designed for human-robot interaction and emotion expression. To evaluate the importance of the proposed system in the framework of human-robot interaction and communication, we conducted a survey by showing videos of the robot behaviors to a group of 30 subjects. The results show that the integration of dynamical emotion expression and locomotion makes the humanoid robot more appealing to humans, as it is perceived as more “favorable” and “useful”, and less “robot-like". Tatsuhiro Kishi, Takuya Kojima, Nobutsuna Endo, Matthieu Destephe, Takuya Otani, Lorenzo Jamone, Przemyslaw Kryczka, Gabriele Trovato, Kenji Hashimoto, Sarah Cosentino, Atsuo Takanishi |
ICRA | 2 |
| 2010 | Integration of emotion expression and visual tracking locomotion based on Vestibulo-Ocular ReflexabstractPersonal robots anticipated to become popular in the future are required to be active in joint work and community life with humans. These personal robots must recognize changing environment and must conduct adequate actions like human. Visual tracking can be said as a fundamental function from the view point of environmental sensing and reflex reaction against it. The authors developed a visual tracking motion algorithm by using upper body. Then, we integrated it with an online walking pattern generator and developed a visual tracking biped locomotion. Finally, we conducted an experimental evaluation with emotion expression. Nobutsuna Endo, Keita Endo, Kenji Hashimoto, Takuya Kojima, Fumiya Iida, Atsuo Takanishi |
RO-MAN | 4 |
| 1997 | An Evolutionary Algorithm Extended by Ecological Analogy and its Application to the Game of Go
Takuya Kojima, Kazuhiro Ueda, Saburo Nagano |
IJCAI (1) | 1 |