VLDB 2026 Research / reviewers in the wild / expert
Hiromitsu Awano
dblp:58/8743
· DBLP profile ↗
39ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 9 first-author · 20 since 2021Software engineering, systems software and programming languages · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 9 |
| 2026 | Biologically Constrained DNA Encoding With Triplet Networks for Similarity Image RetrievalabstractAs the volume of digital data continues to grow exponentially, DNA has emerged as a promising medium for long-term data storage due to its high density and durability. For enabling data retrieval via DNA's biochemical reactions, the encoding strategy plays a critical role. This paper proposes a training framework for a DNA encoder that improves both accuracy and training efficiency in content-based image retrieval by incorporating deep metric learning. In addition, we introduce loss functions that enforce biological constraints, specifically homopolymer length and GC content, thereby improving the biochemical stability of the generated DNA sequences. To evaluate the effectiveness of the proposed method, we conduct quantitative assessments based on image classification performance. Simulations on the CIFAR-10 and CIFAR-100 datasets demonstrate that our method achieves classification accuracy comparable to CNN-based baselines and a 20-fold speedup over the training time of the existing method. Moreover, the generated DNA sequences enable strict control of homopolymer length and maintain GC content within the optimal 40-60% range, significantly improving biological feasibility compared to baseline methods. Takefumi Koike, Hiromitsu Awano, Takashi Sato 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Random Telegraph Noise Observed on 65-nm Bulk pMOS Transistors at 3.8KabstractThis paper presents a detailed study on Random Telegraph Noise (RTN) behavior under cryogenic conditions. The study leverages a device array, BTIarray, to statistically measure RTN in a temperature range from room temperature down to 3.8 K. The measurement results indicate that while RTN's impact decreases in the low-temperature region at about 100 K, it becomes more pronounced at lower temperatures, especially in transistors with shorter channel lengths. This research advances the understanding of RTN in cryogenic environments, offering essential insights for future integrated circuit (IC) design. Takuma Kawakami, Takashi Sato 0001, Hiromitsu Awano |
ASP-DAC | 3 |
| 2025 | Weighted Range-Constrained Ising-Model Decoder for Quantum Error CorrectionabstractIsing model-based Quantum Error Correction decoders reduce topological complexity compared to classical decoders. However, the SOTA Ising decoder has a higher time complexity than union-find (UF) and a lower threshold than minimum-weight perfect-matching (MWPM). We propose the Weighted Range-Constrained Ising Model-Based (WRIM) decoder. WRIM uses a polygonal region to enclose flipped syndromes, ensuring the coverage of all potential error chains while optimizing coupling and external field coefficients. WRIM reduces the variable count by $97.8 x$, achieves microsecondlevel decoding, and has a worst-case time complexity of $O(n)$, outperforming UF. WRIM exhibits threshold behavior up to 10.7$\mathbf{1 1. 0 \%}$, surpassing the MWPM’s highest reported threshold. Hiromitsu Awano, Takashi Sato 0001 |
DAC | 2 |
| 2025 | Lookup Table-based Multiplication-free All-digital DNN Accelerator Featuring Self-Synchronous Pipeline AccumulationabstractDeep neural networks (DNNs) have been widely applied in our society, yet reducing power consumption due to large-scale matrix computations remains a critical challenge. MADDNESS is a known approach to improving energy efficiency by substituting matrix multiplication with table lookup operations. Previous research has employed large analog computing circuits to convert inputs into LUT addresses, which presents challenges to area efficiency and computational accuracy. This paper proposes a novel MADDNESS-based all-digital accelerator featuring a self-synchronous pipeline accumulator, resulting in a compact, energy-efficient, and PVT-invariant computation. Post-layout simulation using a commercial 22nm process showed that 2.5 × higher energy efficiency (174 TOPS/W) and 5× higher area efficiency (2.01 TOPS/mm2) can be achieved compared to the conventional accelerator. Hiroto Tagata, Takashi Sato 0001, Hiromitsu Awano |
DAC | 3 |
| 2025 | SOME: Symmetric One-Hot Matching Elector - A Lightweight Microsecond Decoder for Quantum Error CorrectionabstractConventional quantum error correction (QEC) de-coders such as Minimum-Weight Perfect Matching (MWPM) and Union-Find (UF) offer high thresholds and fast decoding, respectively, but both suffer from high topological complexity. In contrast, Ising model-based decoders reduce topological complexity but demand considerable decoding time. We propose the Symmetric One-Hot Matching Elector (SOME), a novel decoder that reformulates the QEC decoding task as a Quadratic Unconstrained Binary Optimization (QUBO) problem—termed the One-Hot QUBO (OHQ). Each variable in the QUBO represents whether a given pair of flipped syndromes is matched, while the error probabilities between the pair are encoded as interaction coefficients (weight). Constraints ensure that each flipped syndrome is matched exactly once. Valid solutions of OHQ correspond to self-inverse permutation matrices, characterized by symmetric one-hot encoding. To solve the OHQ efficiently, SOME reformulates the decoding task as the construction of permutation matrices that minimize the total weight. It initializes each candidate matrix from one of the minimum-weight syndrome pairs, then iteratively appends additional pairs in ascending order of weight, and finally selects the permutation matrix with the lowest total energy. SOME achieves up to a 99.9x reduction in variable count and reduces decoding times from milliseconds to microseconds on a single-threaded commodity CPU. OHQ also maintains performance up to a 10.5% physical error rate, surpassing the highest known threshold of MWPM. Geguang Miao, Shinichi Nishizawa, Hiromitsu Awano, Shinji Kimura, Takashi Sato 0001 |
ICCAD | 4 |
| 2025 | Zero-Aware Regularization for Energy-Efficient Inference on Akida Neuromorphic ProcessorabstractSpiking Neural Networks (SNNs) and their hardware accelerators have emerged as promising systems for advanced cognitive processing with low power consumption. Although the development of SNN hardware accelerators is particularly active, research on the intelligent use of these accelerators remains limited. This study focuses on the SNN accelerator Akida, a commercially available neuromorphic processor, and presents a novel training method designed to reduce inference energy by leveraging the unique architecture of the hardware. Specifically, we apply sparse constraints on neuron activations and synaptic connection weights, aiming to minimize the number of firing neurons by considering Akida's batch spike processing feature. Our proposed method was applied to a network consisting of three convolutional layers and two fully connected layers. In the MNIST image classification task, the activations became 76.1% sparser, and the weights became 22.1% sparser, resulting in a 13.8% reduction in energy consumption per image. Takehiro Habara, Takashi Sato 0001, Hiromitsu Awano |
ISCAS | 3 |
| 2025 | Gaitcloud: Leveraging Spatial-Temporal Information for Lidar-Base Gait Recognition With a True-3D Gait Representation
Hiromitsu Awano, Takashi Sato 0001 |
WACV | 2 |
| 2025 | Online Training and Inference System on Edge FPGA Using Delayed Feedback ReservoirabstractA delayed feedback reservoir (DFR) is a hardware-friendly reservoir computing system. Implementing DFRs in embedded hardware requires efficient online training. However, two main challenges prevent this: 1) hyperparameter selection, which is typically done by offline grid search, and 2) training of the output linear layer, which is memory-intensive. This article introduces a fast and accurate parameter optimization method for the reservoir layer utilizing backpropagation and gradient descent by adopting a modular DFR model. A truncated backpropagation strategy is proposed to reduce memory consumption associated with the expansion of the recursive structure while maintaining accuracy. The computation time is significantly reduced compared to grid search. In addition, an in-place Ridge regression for the output layer via 1-D Cholesky decomposition is presented, reducing memory usage to be 1/4. These methods enable the realization of an online edge training and inference system of DFR on an FPGA, reducing computation time by about 1/13 and power consumption by about 1/27 compared to software implementation on the same board. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Triplet Network-Based DNA Encoding for Enhanced Similarity Image RetrievalabstractWith the exponential growth of digital data, DNA is emerging as an attractive medium for storage and computing. Thus, design methods for encoding, storing, and searching digital data within DNA storage are of utmost importance. This paper introduces image classification as a measurable task for evaluating the performance of DNA encoders in similar image searches. Furthermore, we propose a novel triplet network-based DNA encoder to improve the accuracy and efficiency. The evaluation using the CIFAR-100 dataset demonstrates that the proposed encoder outperforms existing encoders in retrieving similar images, with an accuracy of 0.77, which is equivalent to 94% of the practical upper limit, and 16 times faster training time. Takefumi Koike, Hiromitsu Awano, Takashi Sato 0001 |
DAC | 2 |
| 2024 | Fast Parameter Optimization of Delayed Feedback Reservoir with Backpropagation and Gradient DescentabstractA delayed feedback reservoir (DFR) is a reservoir computing system well-suited for hardware implementations. However, achieving high accuracy in DFRs depends heavily on selecting appropriate hyperparameters. Conventionally, due to the presence of a non-linear circuit block in the DFR, the grid search has only been the preferred method, which is computationally intensive and time-consuming and thus performed offline. This paper presents a fast and accurate parameter optimization method for DFRs. To this end, we leverage the well-known backpropagation and gradient descent framework with the state-of-the-art DFR model for the first time to facilitate parameter optimization. We further propose a truncated backpropagation strategy applicable to the recursive dot-product reservoir representation to achieve the highest accuracy with reduced memory usage. With the proposed lightweight implementation, the computation time has been significantly reduced by up to 1/700 of the grid search. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
DATE | 2 |
| 2024 | DNA-Based Similar Image Retrieval via Triplet Network-Driven EncoderabstractWith the exponential growth of digital data, DNA has emerged as an attractive medium for storage and computing. Design methods for encoding, storing, and searching digital data within DNA storage are thus of utmost importance. This paper introduces image classification as a measurable task for evaluating the performance of DNA encoders in similar image searches. In addition, we propose a triplet network-based DNA encoder to improve accuracy and efficiency. The evaluation using the CIFAR-100 dataset demonstrates that the proposed encoder outperforms existing encoders in retrieving similar images, with an accuracy of 0.77, which is equivalent to 94 % of the practical upper limit, and achieves 16 times faster training time. Takefumi Koike, Hiromitsu Awano, Takashi Sato 0001 |
DATE | 2 |
| 2024 | S3M: Static Semi-Segmented Multipliers for Energy-Efficient DNN Inference AcceleratorsabstractApproximate multipliers offer an efficient approach to reduce power consumption in compute-intensive applications, such as Deep Neural Networks (DNNs). However, current 8-bit approximate multipliers struggle to maintain high accuracy across various DNN applications. In this paper, we highlight challenges in 8-bit multiplier designs with body approximation strategies and evaluate the effectiveness of input approximation methods. Recognizing that exact multipliers with quantization bit-widths below 8 bits have demonstrated superior performance, we aim to explore whether alternative input approximation methods can provide an even better tradeoff between accuracy and energy consumption. To this end, by exploiting the fact that weight operand values are smaller than activations and prepared offline in DNNs, we simplify a static segmented multiplier (SSM) into a static semi-segmented multiplier$(\mathbf{S}^{3}\mathbf{M})$, achieving a 31.58% reduction in power-delay product (PDP) compared to the original SSM, with similar classification accuracy. Additionally, we propose Coded$\mathbf{S}^{3}\mathbf{M}$with optimized memory usage and im-plement various multipliers on a systolic array-based accelerator. Experimental results show that the proposed$\mathbf{S}^{3}\mathbf{M}$and Coded$\mathbf{S}^{3}\mathbf{M}$outperform existing 8-bit approximate multipliers in DNN applications, effectively bridging the PDP and inference accuracy tradeoff observed across exact commercial IP multipliers of varied bit-widths without requiring time-consuming retraining. Consequently, the proposed multiplier designs provide enhanced computational solutions for energy-efficient DNN inference ac-celerators. Hiromitsu Awano, Longyang Lin, Masanori Hashimoto |
ICCD | 3 |
| 2024 | A Robust and Energy Efficient Hyperdimensional Computing System for Voltage-scaled CircuitsabstractVoltage scaling is one of the most promising approaches for energy efficiency improvement but also brings challenges to fully guaranteeing stable operation in modern VLSI. To tackle such issues, we further extend the DependableHD to the second version DependableHDv2 , a HyperDimensional Computing (HDC) system that can tolerate bit-level memory failure in the low voltage region with high robustness. DependableHDv2 introduces the concept of margin enhancement for model retraining and utilizes noise injection to improve the robustness, which is capable of application in most state-of-the-art HDC algorithms. We additionally propose the dimension-swapping technique, which aims at handling the stuck-at errors induced by aggressive voltage scaling in the memory cells. Our experiment shows that under 8% memory stuck-at error, DependableHDv2 exhibits a 2.42% accuracy loss on average, which achieves a 14.1× robustness improvement compared to the baseline HDC solution. The hardware evaluation shows that DependableHDv2 supports the systems to reduce the supply voltage from 430 mV to 340 mV for both item Memory and Associative Memory, which provides a 41.8% energy consumption reduction while maintaining competitive accuracy performance. Dehua Liang, Hiromitsu Awano, Noriyuki Miura, Jun Shiomi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | DependableHD: A Hyperdimensional Learning Framework for Edge-Oriented Voltage-Scaled CircuitsabstractVoltage scaling is one of the most promising approaches for energy efficiency improvement but also brings challenges to fully guaranteeing the stable operation in modern VLSI. To tackle such issues, we propose DependableHD, a learning framework based on HyperDimensional Computing (HDC), which supports the systems to tolerate bit-level memory failure in the low voltage region with high robustness. For the first time, DependableHD introduces the concept of margin enhancement for model retraining and utilizes noise injection to improve the robustness, which is capable of application in most state-of-the-art HDC algorithms. Our experiment shows that under 10% memory error, DependableHD exhibits a 1.22% accuracy loss on average, which achieves an 11.2× improvement compared to the baseline HDC solution. The hardware evaluation shows that DependableHD supports the systems to reduce the supply voltage from 400mV to 300mV, which provides a 50.41% energy consumption reduction while maintaining competitive accuracy performance. Dehua Liang, Hiromitsu Awano, Noriyuki Miura, Jun Shiomi |
ASP-DAC | 2 |
| 2023 | B2N2: Resource efficient Bayesian neural network accelerator using Bernoulli sampler on FPGAabstractA resource efficient hardware accelerator for Bayesian neural network (BNN) named B 2 N 2 , B ernoulli random number based B ayesian n eural network accelerator, is proposed. As neural networks expand their application into risk sensitive domains where mispredictions may cause serious social and economic losses, evaluating the NN’s confidence on its prediction has emerged as a critical concern. Among many uncertainty evaluation methods, BNN provides a theoretically grounded way to evaluate the uncertainty of NN’s output by treating network parameters as random variables . By exploiting the central limit theorem , we propose to replace costly Gaussian random number generators (RNG) with Bernoulli RNG which can be efficiently implemented on hardware since the possible outcome from Bernoulli distribution is binary. We demonstrate that B 2 N 2 implemented on Xilinx ZCU104 FPGA board consumes only 465 DSPs and 81661 LUTs which corresponds to 50.9% and 14.3% reductions compared to Gaussian-BNN (Hirayama et al., 2020) implemented on the same FPGA board for fair comparison. We further compare B 2 N 2 with VIBNN (Cai et al., 2018), which shows that B 2 N 2 successfully reduced DSPs and LUTs usages by 50.9% and 57.9%, respectively. Owing to the reduced hardware resources, B 2 N 2 improved energy efficiency by 7.50% and 57.5% compared to Gaussian-BNN (Hirayama et al., 2020) and VIBNN (Cai et al., 2018), respectively. Hiromitsu Awano, Masanori Hashimoto |
Integr. | 1 |
| 2023 | Modular DFR: Digital Delayed Feedback Reservoir Model for Enhancing Design FlexibilityabstractA delayed feedback reservoir (DFR) is a type of reservoir computing system well-suited for hardware implementations owing to its simple structure. Most existing DFR implementations use analog circuits that require both digital-to-analog and analog-to-digital converters for interfacing. However, digital DFRs emulate analog nonlinear components in the digital domain, resulting in a lack of design flexibility and higher power consumption. In this paper, we propose a novel modular DFR model that is suitable for fully digital implementations. The proposed model reduces the number of hyperparameters and allows flexibility in the selection of the nonlinear function, which improves the accuracy while reducing the power consumption. We further present two DFR realizations with different nonlinear functions, achieving 10× power reduction and 5.3× throughput improvement while maintaining equal or better accuracy. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | DistriHD: A Memory Efficient Distributed Binary Hyperdimensional Computing Architecture for Image ClassificationabstractHyper-Dimensional (HD) computing is a brain-inspired learning approach for efficient and fast learning on today's embedded devices. HD computing first encodes all data points to high-dimensional vectors called hypervectors and then efficiently performs the classification task using a well-defined set of operations. Although HD computing achieved reasonable performances in several practical tasks, it comes with huge memory requirements since the data point should be stored in a very long vector having thousands of bits. To alleviate this problem, we propose a novel HD computing architecture, called DistriHD which enables HD computing to be trained and tested using binary hypervectors and achieves high accuracy in single-pass training mode with significantly low hardware resources. DistriHD encodes data points to distributed binary hypervectors and eliminates the expensive item memory in the encoder, which significantly reduces the required hardware cost for inference. Our evaluation also shows that our model can achieve a$27.6\times$reduction in memory cost without hurting the classification accuracy. The hardware implementation also demonstrates that DistriHD achieves over$9.9\times$and$28.8\times$reduction in area and power, respectively. Dehua Liang, Jun Shiomi, Noriyuki Miura, Hiromitsu Awano |
ASP-DAC | 4 |
| 2022 | Respiratory Rate Estimation Based on WiFi Frame CaptureabstractThis paper presents a method that estimates the respiratory rate based on the frame capturing of wireless local area networks. The method uses beamforming feedback matrices (BFMs) contained in the captured frames, which is a rotation matrix of channel state information (CSI). BFMs are transmitted unencrypted and easily obtained using frame capturing, requiring no specific firmware or WiFi chipsets, unlike the methods that use CSI. Such properties of BFMs allow us to apply frame capturing to various sensing tasks, e.g., vital sensing. In the proposed method, principal component analysis is applied to BFMs to isolate the effect of the chest movement of the subject, and then, discrete Fourier transform is performed to extract respiratory rates in a frequency domain. Experimental evaluation results confirm that the frame-capture-based respiratory rate estimation can achieve estimation error lower than 3.5 breaths/minute. Takamochi Kanda, Takashi Sato 0001, Hiromitsu Awano, Sota Kondo, Koji Yamamoto 0001 |
CCNC | 3 |
| 2022 | Pay Attention via Binarization: Enhancing Explainability of Neural Networks via Binarization of ActivationabstractModern deep learning algorithms consist of highly complex artificial neural networks, making it extremely difficult for humans to track the inference process. While the social implementation of deep learning is progressing, the human and economic losses caused by inference errors are becoming more and more problematic, and there is a need for methods to explain the basis for the decisions of deep learning algorithms. Although, in an automated driving task, a method to visualize the regions that contribute to steering angle prediction using an attention mechanism has been proposed, its explanatory capability is still low. In this paper, we focus on the difference in the importance of each bit in the activation (i.e., the LSBs have the lowest weight while the MSBs have the highest weight), and propose a method to add attention only to the sign bits to further enhance the explanation. Our numerical experiment using the Udacity dataset revealed that the proposed method achieves 33% higher area under curve (AUC) in terms of the deletion metric. Yuma Tashiro, Hiromitsu Awano |
ISCAS | 2 |
| 2022 | Hardware-Friendly Delayed-Feedback Reservoir for Multivariate Time-Series ClassificationabstractReservoir computing (RC) is attracting attention as a machine-learning technique for edge computing. In time-series classification tasks, the number of features obtained using a reservoir depends on the length of the input series. Therefore, the features must be converted to a constant-length intermediate representation (IR), such that they can be processed by an output layer. Existing conversion methods involve computationally expensive matrix inversion that significantly increases the circuit size and requires processing power when implemented in hardware. In this article, we propose a simple but effective IR, namely, dot-product-based reservoir representation (DPRR), for RC based on the dot product of data features. Additionally, we propose a hardware-friendly delayed-feedback reservoir (DFR) consisting of a nonlinear element and delayed feedback loop with DPRR. The proposed DFR successfully classified multivariate time series data that has been considered particularly difficult to implement efficiently in hardware. In contrast to conventional DFR models that require analog circuits, the proposed model can be implemented in a fully digital manner suitable for high-level syntheses. A comparison with existing machine-learning methods via field-programmable gate array implementation using 12 multivariate time-series classification tasks confirmed the superior accuracy and small circuit size of the proposed method. Sosei Ikeda, Hiromitsu Awano, Takashi Sato 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | BloomCA: A Memory Efficient Reservoir Computing Hardware Implementation Using Cellular Automata and Ensemble Bloom FilterabstractIn this work, we propose a BloomCA which utilizes cellular automata (CA) and ensemble Bloom filter to organize an RC system by using only binary operations, which is suitable for hardware implementation. The rich pattern dynamics created by CA can map the input into high-dimensional space and provide more features for the classifier. Utilizing the ensemble Bloom filter as the classifier, the features can be memorized effectively. Our experiment reveals that applying the ensemble mechanism to Bloom filter endues a significant reduction in inference memory cost. Comparing with the state-of-the-art reference, the BloomCA achieves a 43× reduction for memory cost without hurting the accuracy. Our hardware implementation also demonstrates that BloomCA achieves over 21× and 43.64% reduction in area and power, respectively. Dehua Liang, Masanori Hashimoto, Hiromitsu Awano |
DATE | 3 |
| 2021 | Binary Neural Network in Robotic Manipulation: Flexible Object Manipulation for Humanoid Robot Using Partially Binarized Auto-Encoder on FPGAabstractA neural network based flexible object manipulation system for a humanoid robot on FPGA is proposed. Although the manipulations of flexible objects using robots attract ever increasing attention since these tasks are the basic and essential activities in our daily life, it has been put into practice only recently with the help of deep neural networks. However such systems have relied on GPU accelerators, which cannot be implemented into the space limited robotic body. Although field programmable gate arrays (FPGAs) are known to be energy efficient and suitable for embedded systems, the model size should be drastically reduced since FPGAs have limited on-chip memory. To this end, we propose "partially" binarized deep convolutional auto-encoder technique, where only an encoder part is binarized to compress model size without degrading the inference accuracy. The model implemented on Xilinx ZCU102 achieves 41.1 frames per second with a power consumption of 3.1 W, which corresponds to 10× and 3.7× improvements from the systems implemented on Core i7 6700K and RTX 2080 Ti, respectively. Satoshi Ohara, Tetsuya Ogata, Hiromitsu Awano |
IROS | 3 |
| 2020 | BYNQNet: Bayesian Neural Network with Quadratic Activations for Sampling-Free Uncertainty Estimation on FPGAabstractAn efficient inference algorithm for Bayesian neural network (BNN) named BYNQNet, Bayesian neural network with quadratic activations, and its FPGA implementation are proposed. As neural networks find applications in mission critical systems, uncertainty estimations in network inference become increasingly important. BNN is a theoretically grounded solution to deal with uncertainty in neural network by treating network parameters as random variables. However, an inference in BNN involves Monte Carlo (MC) sampling, i.e., a stochastic forwarding is repeated N times with randomly sampled network parameters, which results in N times slower inference compared to non-Bayesian approach. Although recent papers proposed sampling-free algorithms for BNN inference, they still require evaluation of complex functions such as a cumulative distribution function (CDF) of Gaussian distribution for propagating uncertainties through nonlinear activation functions such as ReLU and Heaviside, which requires considerable amount of resources for hardware implementation. Contrary to conventional BNN, BYNQNet employs quadratic nonlinear activation functions and hence the uncertainty propagation can be achieved using only polynomial operations. Our numerical experiment reveals that BYNQNet has comparative accuracy with MC-based BNN which requires N=10 forwardings. We also demonstrate that BYNQNet implemented on Xilinx PYNQ-Z1 FPGA board achieves the throughput of 131×103images per second and the energy efficiency of 44.7×103images per joule, which corresponds to 4.07× and 8.99× improvements from the state-of-the-art MC-based BNN accelerator. Hiromitsu Awano, Masanori Hashimoto |
DATE | 1 |
| 2020 | Implementation of pseudo-linear feedback shift register-based physical unclonable functions on silicon and sufficient Challenge-Response pair acquisition using Built-In Self-Test before shippingabstractWe implemented pseudo-linear feedback shift-register-based physical unclonable functions (PL-PUFs) on silicon and analyzed their performances in terms of reproducibility, uniqueness, and resistance to machine-learning attacks. A PL-PUF is compact and high-throughput PUF, slightly oversensitive to voltage fluctuations. To overcome this drawback, we developed a capturing signal generation circuit that was tolerant to the reproducibility degradation caused by supply voltage changes. We also implemented a Built-In Self-Test (BIST) circuit with an irreversible destruction mechanism to enable exceedingly fast challenge–response pairs (CPRs) for the PUFs before shipping. After the CPRs were evaluated, the BIST circuit became invulnerable to exploitation by attackers. Yasuhiro Ogasahara, Yohei Hori, Toshihiro Katashita, Tomoki Iizuka, Hiromitsu Awano, Hanpei Koike |
Integr. | 5 |
| 2019 | Fourℚ on ASIC: Breaking Speed Records for Elliptic Curve Scalar MultiplicationabstractAn ASIC cryptoprocessor for scalar multiplication (SM) on FourQ is proposed. By exploiting Karatsuba multiplication and lazy reduction techniques, the arithmetic units of the proposed processor are tailored for operations over quadratic extension field (Fp2). We also propose an automated instruction scheduling methodology based on a combinatorial optimization solver to fully exploit the available instruction-level parallelism. With the proposed processor fabricated by using a 65 nm silicon-on-thin-box (SOTB) CMOS process, we demonstrate that an SM can be computed in 10.1 μs when a typical operating voltage of 1.20V is applied, which corresponds to 3.66× acceleration compared to the conventional P-256 curve SM accelerator implemented on an ASIC platform and is the fastest ever reported. We also demonstrate that by lowering the supply voltage down to 0.32V, the lowest ever reported energy consumption of 0.327 μJ/SM is achieved. Hiromitsu Awano |
DATE | 1 |
| 2019 | PUFNet: A Deep Neural Network Based Modeling Attack for Physically Unclonable FunctionabstractA deep neural network named PUFNet is proposed, which reveals the unreported threats to device authentication systems based on physically unclonable functions (PUFs). We demonstrates that, by employing novel techniques developed in deep learning community, such as ReLU activation function and Xavier initialization technique, PUFNet successfully predicts the responses of double-arbiter PUF (DAPUF) to unseen challenges with probability of 88.4%, which is 21.1% higher than the conventional method. Those results highlight an important fact that PUF-based authentication schemes should be carefully designed considering the rapid evolving machine learning technology. Hiromitsu Awano, Tomoki Iizuka |
ISCAS | 1 |
| 2018 | Ising-PUF: A machine learning attack resistant PUF featuring lattice like arrangement of Arbiter-PUFsabstractA concept of Ising-PUF, a novel PUF structure that utilizes chaotic behavior of mutually interacting small PUFs, is proposed. Ising-PUF consists of a lattice like arrangement of small PUFs, each of which contains a spin register that stores the response of the small PUF, which also serves as a challenge of its neighbors. The spin patterns that develop along time determine the 1-bit response of the Ising-PUF. Utilizing state-memorizing nature of the spin registers, Ising-PUF attains a challenge hysteresis, i.e., allowing sequence of challenge inputs that continuously stimulate its chaotic behavior, which provides the drastically large challenge-to-response space. Experimental results demonstrate nearly ideal metrics; inter-chip Hamming distance (HD) of 50.1% and inter-environment HD of 2.26%. Further, Ising-PUF is remarkably tolerant to machine learning attacks, demonstrating that, even with a deep neural network using a 50k training cRPs, the prediction accuracy remains 50%, which is comparable to a random guess. Hiromitsu Awano, Takashi Sato 0001 |
DATE | 1 |
| 2017 | Efficient circuit failure probability calculation along product lifetime considering device agingabstractA device-aging simulation that efficiently estimates temporal degradation of failure probability of a circuit is proposed. As the size of transistors shrinks, consideration of device aging in addition to manufacturing variability has become an urgent issue for maintaining reliability of LSIs. Contrary to existing techniques that separately handle manufacturing variability and the device aging, we propose a simultaneous evaluation approach using an augmented reliability and subset simulation. By eliminating the repetitive failure-probability calculations at each device-age, the proposed method reduces the number of required circuit simulations to about 1/6 of that of the conventional method without compromising accuracy. Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
ASP-DAC | 1 |
| 2017 | Yield Enhancement by Repair Circuits for Ultra-Fine Pitch Stacked-Chip ConnectionsabstractIn order to realize high-performance LSIs, 3-D stacking technology, using vertical interconnections, is one of the promising solutions for shorter wiring. For high-performance processing systems, total bandwidth of interconnections should be more than 1Tbps. To accommodate this requirement within limited die area, intra and inter chip wires, TSVs and inter chip Vias should be of sub micron widths in diameter, with chip-to-chip misalignment up to 1um. To cope with these issues, we have studied yield estimation technique for such ultra-fine pitch inter-chip Vias (ICVs) and have proposed switch matrix arrays for repairing misconnections by misalignments. We have demonstrated 2.6 times enhancement of available number of ICVs in case of a 10nm process. Keitaro Koga, Hiromitsu Awano |
ATS | 2 |
| 2017 | Scalable Device Array for Statistical Characterization of BTI-Related ParametersabstractA device array circuit, scalable in terms of the number of transistors used, is proposed. The proposed array facilitates accurate and simultaneous bias voltage application to a large number of devices, making it suitable for the measurement-based statistical characterization of device degradation, known as bias temperature instability. Using the proposed array, the degradation measurement of thousands of transistors is made possible in a practical amount of time. The experimental results show that the defect-centric model can approximate the statistical variation in magnitudes of threshold voltage shifts (ΔVTH) and that the variance of ΔVTHbears an inverse relationship to the channel areas of transistors. The degradation variability under ac stress conditions is also presented for the first time. Hiromitsu Awano, Shumpei Morita, Takashi Sato 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | RTN in Scaled Transistors for On-Chip Random Seed GenerationabstractRandom numbers play a vital role in cryptography, where they are used to generate keys, nonce, one-time pads, and initialization vectors for symmetric encryption. The quality of random number generator (RNG) has significant implications on vulnerability and performance of these algorithms. A pseudo-RNG uses a deterministic algorithm to produce numbers with a distribution very similar to uniform. True RNGs (TRNGs), on the other hand, use some natural phenomenon/process to generate random bits. They are nondeterministic, because the next number to be generated cannot be determined in advance. In this paper, a novel on-chip noise source, random telegraph noise (RTN), is exploited for simple and reliable TRNG. RTN, a microscopic process of stochastic trapping/detrapping of charges, is usually considered as a noise and mitigated in design. Through physical modeling and silicon measurement, we demonstrate that RTN is appropriate for TRNG, especially in highly scaled MOSFETs. Due to the slow speed of RTN, we purpose the system for on-chip seed generation for random number. Our contributions are: 1) physical model calibration of RTN with comprehensive 65- and 180-nm transistor measurements; 2) the scaling trend of RTN, validated with silicon data down to 28 nm; 3) design principles to achieve 50% signal probability by using intrinsic RTN physical properties, without traditional postprocessing algorithms, the generated sequence passes the National Institute of Standards and Technology (NIST) tests; and 4) solutions to manage realistic issues in practice, including multilevel RTN signal, robustness to voltage and temperature fluctuations and the operation speed. Abinash Mohanty, Ketul Sutaria, Hiromitsu Awano, Takashi Sato 0001, Yu Cao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Efficient transistor-level timing yield estimation via line samplingabstractYield estimation has been and will continue to be the integral part in design flow, particularly under large process variability in advanced technology nodes. This paper proposes an efficient method that accelerates transistor-level statistical timing simulations which are intensively conducted in various design stages, such as in final timing verification, while considering thousands of random variables. The proposed method utilizes line sampling (LS), in which integration of randomly generated lines, not the random points, are evaluated. Numerical experiments show that the proposed method achieves 14× to 300× speed-up compared to the fastest one that has ever reported. Hiromitsu Awano, Takashi Sato 0001 |
DAC | 1 |
| 2016 | Workload-Aware Worst Path Analysis of Processor-Scale NBTI DegradationabstractAs technology further scales semiconductor devices, aging-induced device degradation has become one of the major threats to device reliability. In addition, aging mechanisms like the negative bias temperature instability (NBTI) is known to be sensitive to workload (i.e., signal probability) that is hard to be assumed at design phase. In this work, we analyze the workload dependence of NBTI degradation using a processor, and propose a novel technique to estimate the worst-case paths. In our approach, with careful examination, we exploit the fact that the deterministic nature of circuit structure limits the amount of NBTI degradation on different paths, and proposes a two-stage path extraction algorithm to identify the invariable critical paths in the processor. Through numerical experiment on a MIPS32 processor, we performed a detailed signal probability analysis, and successfully extracted 85 invariable critical paths out of the 24,978 path candidates, achieving nearly 300x reduction in the sheer number of paths. Song Bian 0001, Michihiro Shintani, Shumpei Morita, Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | Call Alternation Between Specific Pairs of Male Frogs Revealed by a Sound-Imaging Method in Their Natural Habitat
Ikkyu Aihara, Takeshi Mizumoto, Hiromitsu Awano, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2016 | Physically unclonable function using RTN-induced delay fluctuation in ring oscillatorsabstractThis paper proposes RTN-PUF, a novel PUF that utilizes random telegraph noise (RTN) of transistors as the physical uniqueness of individual devices. Our proposed RTN-PUF generates a response from a pair of ring oscillators (ROs) by comparing the numbers of frequency changes, which depend on the time constants of RTN. Due to the log-uniform distribution of the time constants, our RTN-PUF provides more stable responses than the existing manufacturing-variation-based PUFs. The numerical experiments show that the RTN-PUF reduces false negative errors by about 60 times compared to the conventional RO-based PUF. This facilitates to implement PUF into security purposes. Motoki Yoshinaga, Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
ISCAS | 2 |
| 2015 | ECRIPSE: an efficient method for calculating RTN-induced failure probability of an SRAM cell
Hiromitsu Awano, Masayuki Hiromoto, Takashi Sato 0001 |
DATE | 1 |
| 2011 | Use of a Sparse Structure to Improve Learning Performance of Recurrent Neural Networks
Hiromitsu Awano, Shun Nishide, Hiroaki Arie, Jun Tani, Toru Takahashi 0001, Hiroshi G. Okuno, Tetsuya Ogata |
ICONIP (3) | 1 |
| 2010 | Human-robot cooperation in arrangement of objects using confidence measure of neuro-dynamical systemabstractThe objective of our study was to develop dynamic collaboration between a human and a robot. Most conventional studies have created pre-designed rule-based collaboration systems to determine the timing and behavior of robots to participate in tasks. Our aim is to introduce the confidence of the task as a criterion for robots to determine their timing and behavior. In this paper, we report the effectiveness of applying reproduction accuracy as a measure for quantitatively evaluating confidence in an object arrangement task. Our method is comprised of three phases. First, we obtain human-robot interaction data through the Wizard of OZ method. Second, the obtained data are trained using a neuro-dynamical system, namely, the Multiple Time-scales Recurrent Neural Network (MTRNN). Finally, the prediction error in MTRNN is applied as a confidence measure to determine the robot's behavior. The robot participated in the task when its confidence was high, while it just observed when its confidence was low. Training data were acquired using an actual robot platform, Hiro. The method was evaluated using a robot simulator. The results revealed that motion trajectories could be precisely reproduced with a high degree of confidence, demonstrating the effectiveness of the method. Hiromitsu Awano, Tetsuya Ogata, Shun Nishide, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
SMC | 1 |