Tatsuya Kubo

dblp:302/1883 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0005-1523-3063ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Clutch: High Performance Vector-Scalar Comparison using DRAM via Chunked Temporal Coding
abstract
Vector-scalar comparison is a fundamental computation primitive that compares each element in a vector against a single scalar value. It is widely used in a broad range of data-intensive workloads from databases to machine learning. Due to its low computational intensity, the execution of this operation tends to be memory-bound, especially for large vectors, thereby limiting the utilization of compute resources. Processing-using-DRAM (PuD) is an emerging computing paradigm that performs massively parallel bitwise operations directly within the DRAM array, alleviating off-chip data movement. Unfortunately, no prior work proposes an efficient PuD-based solution tailored to vector-scalar comparisons. Existing PuD-based approaches require many DRAM commands because the comparison's algorithmic complexity grows with operand bit-width in the bit-serial execution model, which is inherently induced by current PuD architectures. As a result, this command overhead becomes the dominant performance bottleneck, limiting application-level speed up. We propose Clutch, a novel data representation and comparison algorithm for accelerating vector-scalar comparisons in PuD systems with high efficiency and scalability. Our key idea is twofold. First, to reduce the number of DRAM commands required for comparison, Clutch adopts temporal coding for vectors, where each value is encoded as a sequence of leading ones. This enables lookup-based comparisons, where comparing against a scalar input simply involves accessing the corresponding DRAM row. Second, Clutch leverages our key insight that a divide-and-conquer approach enables scalable lookup-based comparisons without incurring a prohibitive memory footprint at high bit-precision. Specifically, Clutch partitions the operand into multiple multi-bit chunks which can be compared independently using compact lookup tables, and merges per-chunk results through a procedure designed to execute efficiently on PuD.Clutch provides a flexible tradeoff between throughput and memory usage by adjusting chunk count. Experimental results on two applications, predicate evaluation and decision tree inference, demonstrate that Clutch improves end-to-end application throughput (and energy efficiency) by an average of 12 × (69 ×) over highly-optimized CPU and GPU execution and 2.9 × (3.0 ×) over the state-of-the-art bit-serial PuD implementation. Notably, we present, to our knowledge, the first mapping of decision tree inference to PuD execution, extending PuD to a new application domain. Our results demonstrate that DRAM can serve as a high-performance and energy-efficient computing substrate for comparison-intensive workloads.
Daichi Tokuda, Tatsuya Kubo, Ismail Emir Yuksel, Ataberk Olgun, Haocong Luo, Tomoya Nagatani, Geraldo F. Oliveira, A. Giray Yaglikçi, Mohammad Sadrosadati, Onur Mutlu, Shinya Takamaeda-Yamazaki
ICS2
2026 PuDghost: Experimental Analysis of Computation Result Corruption in Processing-Using-Dram Operations on Real Dram Chips and Implications for Future Systems
Daichi Tokuda, Ismail Emir Yuksel, Tatsuya Kubo, Ataberk Olgun, Haocong Luo, Nisa Bostanci, Jikun Wang, A. Giray Yaglikçi, Shinya Takamaeda-Yamazaki, Onur Mutlu
ISCA3
2025 Hybrid Global/Local Search Algorithm for Optimization of Wireless Sensor Placement
abstract
The authors have investigated a method and algorithm for optimal placement of wireless sensors for carbon dioxide concentration observation. Specifically, the authors have investigated a solving method based on genetic algorithm (GA), in which a large number of wireless sensors are placed in advance, and the optimal placement is obtained by removing wireless sensors that provide similar observations. However, it is well known that the algorithm based on GA, which is a global search method, does not provide good accuracy of the solution. Therefore, in order to improve the accuracy of the solution, we propose a hybrid global/local search algorithm that combines GA with the n-Opt method, which is a local search method. We define a neighborhood solution on the wireless sensor placement problem to develop the hybrid global/local algorithm because the wireless sensor placement is a combinatorial optimization problem. From the evaluation results, the proposed method could obtain the optimal solution with a probability of more than 90% whereas the conventional method could obtain that with a probability of less than 20%.
Tatsuya Kubo, Tomoaki Matsuda, Shusuke Narieda
VTC2025-Fall1
2023 An Adaptive Virtual Node Management Method for Overlay Networks Based on Multiple Time Intervals
Tatsuya Kubo, Tomoya Kawakami
CISIS1
2021 An Enhanced Routing Method for Overlay Networks Based on Multiple Different Time Intervals
abstract
Due to the sensor devices price reduction and improvement of network connection quality, IoT(Internet of Things) is rapidly growing these days. IoT handles massive data that deeply related to time or location. Many peer-to-peer approaches that optimized this characteristic are existing. We have worked for a structured and virtualized overlay network that can efficiently handle time-consecutive specific interval queries. But this method had an unsolved problem that causes more load to the whole network than usual. In this work, we present an optimized routing method for an interval-queryable overlay network. To reduce redundant traffic, we expand the routing process to the physical node level. The simulation shows this method can reduce the load of the network or each node.
Tatsuya Kubo, Tomoya Kawakami
COMPSAC1