Hironori Nakajo

dblp:97/6826 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-7452-2125ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Kyress: a secure, scalable, and resource-efficient CRYSTALS-Kyber cryptosystem for low-cost embedded devices
abstract
The increasing need for post-quantum security has driven significant research into efficient implementations of lattice-based cryptography, particularly for resource-constrained embedded devices. While hardware accelerators can achieve high throughput, they often sacrifice flexibility and complicate software development, whereas software-only implementations struggle to meet the performance demands of real-time cryptographic workloads. However, few studies focus on hardware/software co-design approaches. In this paper, we present Kyress, a resource-balanced, secure, and scalable CRYSTALS-Kyber cryptosystem designed for low-cost embedded platforms. We implement three execution configurations, software-only experiments on a scalar processor, a combination of a scalar processor and a vector co-processor, and full hardware/software co-design on an FPGA platform as well as simulator for the evaluations. Experimental results show that Kyress achieves up to a $$6\,\times \,-\,9.76\,\times$$ speedup over state-of-the-art software implementations and delivers competitive performance compared with existing hardware/software accelerators, while requiring significantly fewer hardware resources.
Kien Tran-Hoang, Hironori Nakajo
J. Supercomput.2
2025 Simulation environment for reconfigurable virtual accelerators using a field programmable gate array development environment
Shunya Kawai, Eriko Maeda, Kazuki Yaguchi, Yasunori Osana, Takefumi Miyoshi, Hironori Nakajo
J. Supercomput.6
2025 Preliminary evaluation of SHAVER: sharing vector registers with an accelerator
Tomoaki Tanaka, Michiya Kato, Yasunori Osana, Takefumi Miyoshi, Jubee Tada, Kiyofumi Tanaka, Hironori Nakajo
J. Supercomput.7
2011 Detecting water waste activities for water-efficient living
abstract
Towards persuasive system for efficient use of water resource, we propose a method to detect "water waste" among water-related activities based on water sound analysis. We supposed two types of water-wastes: inter-activity water waste and intra-activity water waste. An evaluation with a variety of experimental conditions presents that the aggregate accuracies to identify the inter-activity water waste and the intra-activity waste are 96.3% and 92.6%, respectively.
Trang Thuy Vu, Akifumi Sokan, Hironori Nakajo, Kaori Fujinami, Jaakko Suutala, Pekka Siirtola, Tuomo Alasalmi, Ari Pitkänen, Juha Röning
UbiComp3
2011 Feature Selection and Activity Recognition to Detect Water Waste from Water Tap Usage
abstract
In this paper, water tap usage is examined based on water sound analysis. We focus on detecting "water waste" to make persuasion of water savings effective, where two types of water waste are defined: inter-activity water waste and intra-activity water waste. Based on a preliminary user survey, four types of basin-related activities are identified that occur with water waste. We apply a spectrum subtraction method for feature selection and propose cascaded classifiers for activity recognition. The result of an evaluation presents that the aggregate accuracies to identify inter-activity water waste and intra-activity one are 100.0 % and 81.1%, respectively.
Trang Thuy Vu, Akifumi Sokan, Hironori Nakajo, Kaori Fujinami, Jaakko Suutala, Pekka Siirtola, Tuomo Alasalmi, Ari Pitkänen, Juha Röning
RTCSA (2)3
2010 An enhancer of memory and network for applications with large-capacity data and non-continuous data accessing
Noboru Tanabe, Hirotaka Hakozaki, Hiroshi Ando, Yasunori Dohi, Zhengzhe Luo, Hironori Nakajo
J. Supercomput.6
2009 An Effective Replacement Strategy of Cache Memory for an SMT Processor
abstract
An SMT processor is designed to execute multiple threads simultaneously in order to gain higher performance with sharing resources such as ALUs and cache memory among several threads. However, sharing cache memory may cause thread conflict misses which degrades its performance. In this paper, an effective replacement strategy in which conflicts miss ratio among threads is controlled by limiting the range of replaceable cache blocks is proposed and designed in order to overcome the problem on cache memory of an SMT processor. The proposed replacement strategy shows 5.3% as high performance in average and up to 41.9% in maximum as a conventional pseudo LRU strategy. Moreover, hardware costs for implementing the proposed strategy are reduced by 0.74% compared with pseudo LRU strategy.
Yoshiyasu Ogasawara, Hironori Nakajo
DSD2
2009 Network Interface Architecture for Scalable Message Queue Processing
abstract
Most of scientists except computer scientists do not want to make efforts for performance tuning with rewriting their MPI applications. In addition, the number of processing elements which can be used by them is increasing year by year. On large-scale parallel systems, the number of accumulated messages on a message buffer tends to increase in some of their applications. Since searching message queue in MPI is time-consuming, system side scalable acceleration is needed for those systems. In this paper, a support function named LHS (Limited-length Head Separation) is proposed. Its performance in searching message buffer and hardware cost are evaluated. LHS accelerates searching message buffer by means of switching location to store limited-length heads of messages. It uses the effects such as increasing hit rate of cache on host with partial off-loading to hardware. Searching speed of message buffer when the order of message reception is different from the receiver's expectation is accelerated 14.3 times with LHS on FPGA-based network interface card (NIC) named DIMMnet-2. This absolute performance is 38.5 times higher than that of IBM BlueGene/P although the frequency is 8.5times slower than BlueGene/P. Hardware cost of LHS is significantly lower than that of ALPU, which is a hardware accelerator for searching message buffer. LHS has higher scalability than ALPU in the performance per frequency. Therefore, LHS is more suitable for larger parallel systems.
Noboru Tanabe, Atsushi Ohta, Pulung Waskito, Hironori Nakajo
ICPADS4
2008 An Enhancer of Memory and Network for Cluster and its Applications
abstract
Introduction of multi-core structures has not led to a decline in the rapid performance improvement of COTS CPU recently. On the other hand, the performance of memory and I/O systems is insufficient to catch up with that of COTS CPU. In this paper, with a view to realizing high-performance computer systems not only for HPC but also for Google-like servers, we propose concepts concerning memory systems and network systems with large extended memory. We introduce DIMMnet-3, which is a practical solution to enhance memory system and I/O system of PC, and Toshiba Cell Reference Set. Examples of the killer applications of this new type of hardware are presented. Communication mechanisms named LHS and LHC are also proposed. These are architectures for reducing latency for mixed messages with small controlling data and large acknowledge data. The latency evaluation of them is shown.
Noboru Tanabe, Hironori Nakajo
PDCAT2
2006 DIMMnet-2: A Reconfigurable Board Connected Into a Memory Slot
abstract
DIMMnet-2 is a reconfigurable board which can be connected into the DDR SDRAM memory slot of a common PC. It has two DDR SO-DIMM on-board memory and high speed InifiniBand interface. By using them, DIMMnet-2 enables various types of flexible co-processing with the host processor. The overview of DIMMnet-2 is shown as well as two case studies: an intelligent memory access control and an intelligent network interconnect.
Tomotaka Miyashiro, Akira Kitamura, Hironori Nakajo, Noboru Tanabe
FPL3
2005 Evaluation of Network Interface Controller on DIMMnet-2 Prototype Board
abstract
By recent performance improvement of interconnection networks for a PC cluster, standard I/O bus which connects network interface becomes the performance bottleneck. DIMMnet is a network interface which can solve the problem by using the memory bus instead of PCI bus or other I/O buses. The second generation network interface DIMMnet-2 can be connected with DDR-SDRAM slot by using the indirect accessing to memory and buffers. Although the current board is a prototype using an FPGA, the latency for 8 Bytes data transfer is only 0.441µs.
Akira Kitamura, Yasuo Miyabe, Tetsu Izawa, Tomotaka Miyashiro, Konosuke Watanabe, Tomohiro Otsuka, Hideharu Amano, Yoshihiro Hamada, Noboru Tanabe, Hironori Nakajo
PDCAT10
2000 MEMOnet : Network interface plugged into a memory slot
abstract
The communication architecture of the DIMMnet-1 network interface, based on MEMOnet, is described. MEMOnet is an architecture consisting of a network interface plugged into a memory slot. The DIMMnet-1 prototype will have two banks of PC133 based SO-DIMM slots and an 8 Gbps full duplex optical link or two 448 MB/s full duplex LVDS channel links. The software overhead incurred to generate a message is only I CPU cycle and the estimated hardware delay is less than 100 ns using the atomic on-the-fly sending with header TLB. The estimated achievable communication bandwidth with block on-the-fly sending with protection stampable window memory is 440 MB/s which was observed in our experiments writing to the DIMM area with a write combining attribute. This is 3.3 times higher than the maximum bandwidth of PCI. This high performance distributed computing environment is available using economical personal computers with DIMM slots.
Noboru Tanabe, Junji Yamamoto, Hiroaki Nishi, Tomohiro Kudoh, Yoshihiro Hamada, Hironori Nakajo, Hideharu Amano
CLUSTER6
1997 An I/O Network Architecture of the Distributed Shared-Memory Massively Parallel Computer JUMP-1
abstract
Article Free Access Share on An I/O network architecture of the distributed shared-memory massively parallel computer JUMP-1 Authors: Hironori Nakajo Department of Computer and Systems Engineering, Kobe University Department of Computer and Systems Engineering, Kobe UniversityView Profile , Satoshi Ohtani Department of Computer and Systems Engineering, Kobe University Department of Computer and Systems Engineering, Kobe UniversityView Profile , Takashi Matsumoto Department of Information Science, The University of Tokyo Department of Information Science, The University of TokyoView Profile , Masadi Kohata Department of Information and Computer Engineering, Okayama University of Science Department of Information and Computer Engineering, Okayama University of ScienceView Profile , Kei Hiraki Department of Information Science, The University of Tokyo Department of Information Science, The University of TokyoView Profile , Yukio Kaneda Department of Computer and Systems Engineering, Kobe University Department of Computer and Systems Engineering, Kobe UniversityView Profile Authors Info & Claims ICS '97: Proceedings of the 11th international conference on SupercomputingJuly 1997 Pages 253–260https://doi.org/10.1145/263580.263645Published:11 July 1997Publication History 4citation247DownloadsMetricsTotal Citations4Total Downloads247Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Hironori Nakajo, Satoshi Ohtani, Takashi Matsumoto 0002, Masadi Kohata, Kei Hiraki, Yukio Kaneda
International Conference on Supercomputing1
1996 A Simulation-based Evaluation of a Disk I/O Subsystem for a Massively Parallel Computer: JUMP-1
abstract
JUMP-1 is a distributed shared-memory massively parallel computer and is composed of multiple clusters of inter-connected network called RDT (Recursive Diagonal Torus). Each cluster in JUMP-1 consists of 4 element processors, secondary cache memories, and 2 MBP (Memory Based Processor) for high-speed synchronization and communication among clusters. The I/O subsystem is connected to a cluster via a high-speed serial link called STAFF-Link. The I/O buffer memory is mapped onto the JUMP-1 global shared-memory to permit each I/O access operation as memory access. In this paper we describe evaluation of the fundamental performance of the disk I/O subsystem using event-driven simulation, and estimated performance with a Video On Demand (VOD) application.
Hironori Nakajo, Satoshi Ohtani, Yukio Kaneda
ICDCS1