Tomoya Suzuki

dblp:60/833 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
abstract
When key-value (KV) stores use SSDs for storing a large number of items, oftentimes they also require large in-memory data structures including indices and caches to be traversed to reduce IOs. This paper considers offloading most of such data structures from the costly host DRAM to secondary memory whose latency is in the microsecond range, an order of magnitude longer than those of DIMM-mounted persistent memory and currently available CXL memory devices. While emerging microsecond-latency memory, such as one based on flash memory, is likely to cost much less than DRAM, it can significantly slow down pointer-chasing on those in-memory data structures of SSD-based KV stores if naively employed, although its impact has not been well studied. This paper analyzes and evaluates the impact of microsecond-level memory latency on the throughput of SSD-based KV operations. Our analysis finds that a well-known latency-hiding technique of software prefetching for long-latency memory from user-level threads is effective for SSD-based KV stores. The novelty of our analysis lies in modeling how the interplay between prefetching and IO affects performance, from which we derive an equation that well explains the throughput degradation due to long memory latency. The model tells us that the presence of IO in KV operations significantly enhances their tolerance to memory latency, and the throughput degradation is expected to be small even if the memory latency extends to a few microseconds, leading to a finding that SSD-based KV stores can be made latency-tolerant without devising new techniques for microsecond-latency memory. To confirm this through experiments, we design a microbenchmark as well as modify existing SSD-based KV stores so that they issue prefetches for long-latency memory from user-level threads, and run them while placing most of in-memory data structures on FPGA-based memory with adjustable microsecond latency. The results demonstrate that their KV operation throughputs for varying memory latency can be well explained by our model, and the modified KV stores achieve near-DRAM throughputs for up to a memory latency of around 5 microseconds. This suggests the possibility that SSD-based KV stores involving latency-sensitive in-memory data traversal can use microsecond-latency memory as a cost-effective alternative to the host DRAM.
Yosuke Bando, Akinobu Mita, Kazuhiro Hiwada, Shintaro Sano, Tomoya Suzuki, Yu Nakanishi, Kazutaka Tomida, Hirotsugu Kajihara, Akiyuki Kaneko, Daisuke Taki, Yukimasa Miyamoto, Tomokazu Yoshida, Tatsuo Shiozawa
Proc. ACM Manag. Data5
2023 Implementing and Evaluating E2LSH on Storage
abstract
Locality sensitive hashing (LSH) is one of the widely-used approaches to approximate nearest neighbor search (ANNS) in high-dimensional spaces. The first work on LSH for the Euclidean distance, E2LSH, showed how ANNS can be solved efficiently at a sublinear query time in the database size with theoretically-guaranteed accuracy, although it required a large hash index size. Since then, several LSH variants having much smaller index sizes have been proposed. Their query time is linear or superlinear, but they have been shown to run effectively faster because they require fewer I/Os when the index is stored on hard disk drives and because they also permit in-memory execution with modern DRAM capacity. In this paper, we show that E2LSH is regaining the advantage in query speed with the advent of modern flash storage devices such as solid-state drives (SSDs). We evaluate E2LSH on a modern single-node computing environment and analyze its computational cost and I/O cost, from which we derive storage performance requirements for its external memory execution. Our analysis indicates that E2LSH on a single consumer-grade SSD can run faster than the state-of-the-art small-index methods executed in-memory. It also indicates that E2LSH with emerging high-performance storage devices and interfaces can approach in-memory E2LSH speeds. We implement a simple adaptation of E2LSH to external memory, E2LSH-on-Storage (E2LSHoS), and evaluate it for practical large datasets of up to one billion objects using different combinations of modern storage devices and interfaces. We demonstrate that our E2LSHoS implementation runs much faster than small-index methods and can approach in-memory E2LSH speeds, and also that its query time scales sublinearly with the database size beyond the index size limit of in-memory E2LSH.
Yu Nakanishi, Kazuhiro Hiwada, Yosuke Bando, Tomoya Suzuki, Hirotsugu Kajihara, Shintaro Sano, Tatsuro Endo, Tatsuo Shiozawa
EDBT4
2021 Approaching DRAM performance by using microsecond-latency flash memory for small-sized random read accesses: a new access method and its graph applications
abstract
For applications in which small-sized random accesses frequently occur for datasets that exceed DRAM capacity, placing the datasets on SSD can result in poor application performance. For the read-intensive case we focus on in this paper, low latency flash memory with microsecond read latency is a promising solution. However, when they are used in large numbers to achieve high IOPS (Input/Output operations Per Second), the CPU processing involved in IO requests is an overhead. To tackle the problem, we propose a new access method combining two approaches: 1) optimizing issuance and completion of the IO requests to reduce the CPU overhead. 2) utilizing many contexts with lightweight context switches by stackless coroutines. These reduce the CPU overhead per request to less than 10 ns, enabling read access with DRAM-like overhead, while the access latency longer than DRAM can be hidden by the context switches. We apply the proposed method to graph algorithms such as BFS (Breadth First Search), which involves many small-sized random read accesses. In our evaluation, the large graph data is placed on microsecond-latency flash memories within prototype boards, and it is accessed by the proposed method. As a result, for the synthetic and real-world graphs, the execution times of the graph algorithms are 88--141% of those when all the data are placed in DRAM.
Tomoya Suzuki, Kazuhiro Hiwada, Hirotsugu Kajihara, Shintaro Sano, Shuou Nomura, Tatsuo Shiozawa
Proc. VLDB Endow.1
2019 Live Demonstration: FPGA-Based CNN Accelerator with Filter-Wise-Optimized Bit Precision
abstract
To enhance the efficiency for inference of deep convolutional neural network without noticeable degradation of the recognition accuracy, we have proposed a filter-wise optimized quantization with variable bit precision. In addition, we have proposed the hardware architecture that fully supports the quantized variable bit precision. By using this hardware architecture, the execution time is reduced proportionally to the reduced bit precision by the filter-wise quantization. In the demonstration, we present image classification experiments on an FPGA where the proposed architecture is implemented with less execution time than the fixed 16-bit precision model.
Kengo Nakata, Asuka Maki, Daisuke Miyashita, Fumihiko Tachibana, Tomoya Suzuki, Jun Deguchi
ISCAS5
2014 Visualizing Temporal Changes in Impressions from Tweets
abstract
Twitter is effective for connecting with celebrities, on-screen talent, and strangers, as well as with friends and acquaintances, and is a social networking service utilized by several generations. On Twitter, many users will offer constructive suggestions, while a small number of users will sometimes enact mental abuse. Some users may post lively tweets, while other users may express anger in their tweets. By carefully reading a larger number of an individual user's tweets, it is possible to judge the impressions from tweets received from an individual user's typical posts. Therefore, this paper proposes a web application system for visualizing Twitter users based on temporal changes in the impressions from the tweets posted by users on Twitter. When system users input the account name of a Twitter user and a period for analysis, the proposed system collects the user's tweets posted during the period, quantifies the impressions from each tweet, and generates pic and line charts illustrating the impression values of the tweets and their temporal changes.
Tadahiko Kumamoto, Hitomi Wada, Tomoya Suzuki
iiWAS3
2013 RoboCup 2013: Best Humanoid Award Winner JoiTech
Yuji Oshima, Dai Hirose, Syohei Toyoyama, Keisuke Kawano, Shibo Qin, Tomoya Suzuki, Kazumasa Shibata, Takashi Takuma, Minoru Asada
RoboCup6
2012 Predicting hit movie concepts using news articles
abstract
An incredible amount of movies are released every year in Japan, and movies that match the viewers' social needs are the ones that become the big hits. All movies share a common trait in that they all include various “concepts”. For example, the American film “The Matrix” includes the concepts “virtual space”, “artificial intelligence”, “domination”, and others. We have developed a system for predicting hit movie concepts that uses a case base comprised of an immense amount of “ Cause → Consequence” pairs derived from past movies. We consider past hit movies a “consequence” and a news item that has a strong relationship with such movies as a “cause”. When we input new social needs into the system, we can predict hit movies by referring to the case base. Experimental results showed that the proposed system could successfully predict certain parts of what made up a hit movie.
Kazuaki Shimamura, Shinichiro Ito, Tomohiro Takagi, Hiroshi Yoshida, Tomoya Suzuki, Kaoru Kato
FUZZ-IEEE5
2008 Motivation oriented action selection for understanding dynamics of objects
abstract
In this study, we propose a synthetic methodology that can enable a humanoid robot to understand the dynamics of objects in a psychological framework. The action selection of the robot is carried out based on a selection probability determined by several internal variables that draws upon the motivation mechanism of human beings. This makes the robot classify its action space by a sensory feedback caused by its own action. The system prioritizes actions that are expected to cause distinguishing sensory patterns. The action selection probability allows the robot to explore unknown spaces for understanding the dynamics of objects in a real environment. In our experiment, we demonstrate the performance of object clustering by multi-modal active sensing using the proposed action selection rule. Furthermore, we show that the system allows the robot to build knowledge about common object movements by repeating interactions with several different types of objects, and also to predict the movement of unknown objects.
Tomoya Suzuki, Sho Yano, Kenji Suzuki 0002
IROS1
2006 Robust Quantum Algorithms with epsilon-Biased Oracles
Tomoya Suzuki, Shigeru Yamashita, Masaki Nakanishi, Katsumasa Watanabe
COCOON1
2006 Bootstrap Prediction Intervals for Nonlinear Time-Series
Daisuke Haraki, Tomoya Suzuki, Tohru Ikeguchi
IDEAL2
2005 Virtual Human with Regard to Physical Contact and Eye Contact
Asami Takayama, Yusuke Sugimoto, Akio Okuie, Tomoya Suzuki, Kiyotaka Kato
ICEC4
2003 An Encoding Scheme for Generating lambda-Expressions in Genetic Programming
Kazuto Tominaga, Tomoya Suzuki, Kazuhiro Oka
GECCO2