Tomomasa Yamasaki

dblp:237/0078 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-5610-8534ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Coflex: Enhancing HW-NAS with Sparse Gaussian Processes for Efficient and Scalable DNN Accelerator Design
abstract
Hardware-Aware Neural Architecture Search (HW-NAS) is an efficient approach to automatically co-optimizing neural network performance and hardware energy efficiency, making it particularly useful for the development of Deep Neural Network accelerators on the edge. However, the extensive search space and high computational cost pose significant challenges to its practical adoption. To address these limitations, we propose Coflex, a novel HW-NAS framework that integrates the Sparse Gaussian Process (SGP) with multi-objective Bayesian optimization. By leveraging sparse inducing points, Coflex reduces the GP kernel complexity from cubic to near-linear with respect to the number of training samples, without compromising optimization performance. This enables scalable approximation of large-scale search space, substantially decreasing computational overhead while preserving high predictive accuracy. We evaluate the efficacy of Coflex across various benchmarks, focusing on accelerator-specific architecture. Our experimental results show that Coflex outperforms state-of-the-art methods in terms of network accuracy and Energy-Delay-Product, while achieving a computational speed-up ranging from 1.9× to 9.5×.
Yinhui Ma, Tomomasa Yamasaki, Zhehui Wang, Tao Luo 0014, Bo Wang 0020
ICCAD2
2025 A 23.5 TOPS/W Depthwise Separable Convolution Accelerator for Event-based Depth Estimation
abstract
Recent efforts to improve energy efficiency in computer vision (CV) tasks, such as depth estimation, focus on integrating event-based cameras with lightweight networks using Depthwise Separable (DWS) Convolutions. Despite remarkable accuracies and hardware-friendly binary signals, existing accelerators have not fully leveraged this combination. This paper proposes a Separated Engine (SE) architecture for binary input feature maps (ifmaps) with dedicated arrays for each DWS stage, eliminating data storage between stages and enhancing array utilization. Additionally, an integrative dataflow incorporating Row, Weight, and Input Stationary (RS, WS, and IS) advantages is introduced to maximize data reuse, supported by an optimized mapping strategy that efficiently loads and updates ifmaps onto the first array. Our gate-level simulation results using 28nm CMOS technology demonstrated a 1.7x improvement in energy efficiency, achieving 23.5 TOPS/W, along with a 13x enhancement in area efficiency, reaching 1785.5 GOP/mm2.
Andres Brito, Tomomasa Yamasaki, Ulysse Rançon, Timothée Masquelier, Benoit Cottereau, Anh-Tuan Do, Bo Wang 0020
ISCAS2
2025 RBFleX-NAS: Training-Free Neural Architecture Search Using Radial Basis Function Kernel and Hyperparameter Detection
abstract
Neural architecture search (NAS) is an automated technique to design optimal neural network architectures for a specific workload. Conventionally, evaluating candidate networks in NAS involves extensive training, which requires significant time and computational resources. To address this, training-free NAS has been proposed to expedite network evaluation with minimal search time. However, state-of-the-art training-free NAS algorithms struggle to precisely distinguish well-performing networks from poorly performing networks, resulting in inaccurate performance predictions and consequently suboptimal top-one network accuracy. Moreover, they are less effective in activation function exploration. To tackle the challenges, this article proposes RBFleX-NAS, a novel training-free NAS framework that accounts for both activation outputs and input features of the last layer with a radial basis function (RBF) kernel. We also present a detection algorithm to identify optimal hyperparameters using the obtained activation outputs and input feature maps. We verify the efficacy of RBFleX-NAS over a variety of NAS benchmarks. RBFleX-NAS significantly outperforms state-of-the-art training-free NAS methods in terms of top-one accuracy, achieving this with short search time in NAS-Bench-201 and NAS-Bench-SSS. In addition, it demonstrates a higher Kendall correlation compared to layer-based training-free NAS algorithms. Furthermore, we propose the neural network activation function benchmark (NAFBee), a new activation design space that extends the activation type to encompass various commonly used functions. In this extended design space, RBFleX-NAS demonstrates its superiority by accurately identifying the best-performing network during activation function search, providing a significant advantage over other NAS algorithms.
Tomomasa Yamasaki, Zhehui Wang, Tao Luo 0014, Niangjun Chen, Bo Wang 0020
IEEE Trans. Neural Networks Learn. Syst.1
2023 LAXOR: A Bit-Accurate BNN Accelerator with Latch-XOR Logic for Local Computing
abstract
Binary Neural Network (BNN) accelerators are attractive solutions for Artificial Internet-of-Things (AIoT) applications thanks to the compact models and low computational cost while maintaining satisfactory classification performance. Various analog/mix-signal compute-in-memory macros have been proposed to boost the energy efficiency of binary convolution tasks. However, this approach incurs inaccurate computation due to its sensitivity to temperature, noise, and process variations. In this work, we present a full-digital BNN architecture that leverages a novel Latch-XOR logic array for local bitwise multiplication, suppressing massive data movement and achieving 4.2× lower energy per operation compared to the decoupled standard cell approach. An optimized population count circuitry is also proposed for data accumulation, which obtains 1.37× Energy-Delay-Area saving compared to Binary-Adder-Tree-based implementation. To enable seamless hardware-software co-optimization, we have developed an in-house simulator for design space exploration as well as flexible mapping with various network topologies and kernel sizes. Our experiment shows the Latch-XOR-based architecture in 28nm CMOS technology achieves an enhanced energy efficiency of 2315 TOPS/W, 3.4× higher compared to the state-of-the-art synthesized digital architecture. This manifests that the proposed accelerator is highly suited for AIoT applications.
Dongrui Li, Tomomasa Yamasaki, Aarthy Mani, Anh-Tuan Do, Niangjun Chen, Bo Wang 0020
ISLPED2