Yuko Hara-Azumi

dblp:21/1239 · also Yuko Hara · DBLP profile ↗
← Back
42ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-9486-5272ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 7 first-author · 6 since 2021Computer networks · 7 · 7 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiReS IDS: A Lightweight, Real-time, and Secure Intrusion Detection System for RISC-V Edge Devices
Songxuan Liu, Qingyu Zeng 0001, Jan Tobias Mühlberg, Yuko Hara-Azumi
ICC4
2025 Lightweight Real-Time Detection of DDoS Attacks on Iot Devices via Power Side-Channel Analysis
abstract
The proliferation of Internet of Things (IoT) devices has transformed numerous domains, such as smart homes and industrial automation, but has also introduced security vulnerabilities, notably to Distributed Denial of Service (DDoS) attacks. These attacks can severely disrupt IoT functionality by overwhelming devices with excessive traffic. In this study, we propose a novel, lightweight, real-time method to detect DDoS attacks through power side-channel analysis. This nonintrusive approach analyzes power consumption patterns to identify anomalous behaviors indicative of attacks. Our approach involves preprocessing the power consumption data collected in fixed window intervals, from which we extract pivotal features to be processed by a lightweight machine learning technique. Leveraging multi-threaded processing coupled with optimizations from the Open Neural Network Exchange runtime, the system ensures scalability across diverse IoT platforms. Experimental results on an actual microcontroller demonstrate that our system achieves 98.83 % accuracy in detecting DDoS attacks with a processing latency of 25.54 ms and memory usage of 14.92 KB. Additionally, the system maintains a low false positive rate ($\mathbf{1. 6 1 \%}$at the minimum) across various detection tasks, validating the effectiveness of utilizing power side-channel signals as a robust and practical defense within secure IoT architectures.
Qingyu Zeng 0001, Yunpei Xiao, Mingyu Yang 0001, Yuko Hara-Azumi
ICC4
2024 Tram-FL: Routing-based Model Training for Decentralized Federated Learning
abstract
In decentralized federated learning (DFL), the dual challenges of extensive inter-node communication and non-independent, identically distributed (non-IID) data impede the attainment of high-accuracy models while maintaining minimal communication traffic. We propose Tram-FL, which progressively refines a global model by transferring it sequentially amongst nodes. We also introduce a dynamic model routing algorithm for optimal route selection, aimed at enhancing model precision with minimal forwarding. Our experiments demonstrate that Tram-FL with the proposed routing delivers high model accuracy, outperforming baselines while reducing communication costs.
Kota Maejima, Takayuki Nishio, Asato Yamazaki, Yuko Hara-Azumi
CCNC4
2024 Multiplicative Masked M&M: An Attempt at Combined Countermeasures with Reduced Randomness
abstract
With the advancement of hardware security, combined attacks with techniques such as side-channel analysis (SCA) and fault analysis (FA) have prompted the development of combined countermeasures. However, these countermeasures often come with significant overhead. In this paper, we explore a solution to reduce the randomness requirement while maintaining security claims. We demonstrate the approach with Mask & Macs (M&M), a scheme that combines Boolean masking and MAC tag redundancy to provide SCA and DFA protection, addressing the challenge of high randomness requirement. We introduce a novel multiplicative masking scheme as a replacement for threshold implementation (TI) modules partially, leading to a reduction of over 50% in randomness requirement with minor increased FPGA resource overhead and latency. While the trade-off is beneficial, other limitations remain, and further research is needed to address these problems. This work provides a new perspective on improving combined countermeasures by exploring ways to reduce system overhead.
Haruka Hirata, Daiki Miyahara, Kazuo Sakiyama, Yuko Hara-Azumi, Yang Li 0001
TrustCom5
2024 Hardware/Software Cooperative Design Against Power Side-Channel Attacks on IoT Devices
abstract
With the growth of Internet of Things (IoT) era, the protection of secret information on IoT devices is becoming increasingly important. For IoT devices, attacks that target information leakage through physical side-channels (e.g., a power side-channel) are a major threat in many use cases because IoT devices can be accessed easily by a hostile third party. However, securing resource-constrained IoT devices against side-channel attacks is a challenging issue. Generally, it is difficult to satisfy the requirements on side-channel protection while maintaining the low-power and real-time constrains of IoT devices. In this paper, we propose a hardware/software cooperative design for cryptosystems that is suitable for resource-constrained IoT devices. Combining a security-oriented processor design (i.e., an instruction set architecture definition and its architectural structure) and careful implementations of masked software implementation for cipher algorithms can effectively improve the power-performance-area (PPA) while suppressing power side-channel leakage. In our evaluation, for three ciphers (Chaskey, Simon, and AES), we demonstrate that our work is superior to state-of-the-art works (two RISC-V processors and a small-scale low-power processor) in terms of both PPA and power side-channel protection.
Mingyu Yang 0001, Tanvir Ahmed 0004, Saya Inagaki, Kazuo Sakiyama, Yang Li 0022, Yuko Hara-Azumi
IEEE Internet Things J.6
2024 Hardware/Software Codesign of Real-Time Intrusion Detection System for Internet of Things Devices
abstract
The rapid expansion of the Internet of Things (IoT) has increased security concerns, thereby necessitating efficient intrusion detection systems (IDS). In this paper, we propose a real-time IoT IDS designed by combining a random forest (RF) classifier with an ensemble feature selection technique (EFST). The proposed IDS can be deployed on a small-scale field-programmable gate array (FPGA) board. The system utilizes a two-metric ensemble feature selection process to reduce computational complexity and enhance classification accuracy. In addition, the EFST aggressively extracts a limited number of features, thereby reducing the complexity of the RF model. Then, the tailored RF classifier is mapped onto an FPGA-based hardware accelerator to realize real-time detection. The proposed method was evaluated experimentally on the benchmark BoT-IoT dataset. The results demonstrate that the proposed IDS realizes significant improvements in terms of resource utilization and processing time compared to several state-of-the-art FPGA-based IDS implementations while maintaining sufficient detection accuracy. In particular, our implementation on the Xilinx PYNQ Z2 achieved 10.2×, 135.7×, and 8.43× speed-up compared to state-of-the-art IDSs running on an Intel Core i7 CPU, an ARM Cortex-A9 microprocessor, and a neural network-based accelerator on the PYNQ, respectively. In addition, our approach exhibits the lowest resource utilization among FPGA-based IDS solutions. These results demonstrate that this work contributes to developing secure and sustainable IoT ecosystems by integrating EFST, RF classification, and FPGA-based acceleration.
Qingyu Zeng 0001, Yuko Hara-Azumi
IEEE Internet Things J.2
2023 Impact of Quantization Noise on CNN-based Joint Source-Channel Coding and Modulation
abstract
This paper investigated the impact of a quantizer in analog-to-digital and digital-to-analog converters in communication devices on image quality when using deep learning-based joint source-channel coding modulation (JSCCM) for image transmission. In recent years, JSCCM, which efficiently encodes images and videos with low information entropy, has attracted great attention. JSCCM has a structure based on an autoencoder and determines the compression ratios for the image input by adjusting the number of IQ symbol output. The IQ symbol output from the encoder are allocated to symbol constellations with higher degrees of arbitrariness than those in typical square quadrature amplitude modulation and are therefore expected to be strongly affected by the quantization noise. In this paper, we employed quantization to the IQ symbol sequence and investigated its effect. Adjusting the quantizer's clipping ratio and the number of quantization bits, we examined the images' tolerance of the peak signal-to-noise ratio (PSNR). The simulation results showed that by adequately adjusting the clipping ratio, the image quality can be guaranteed to be equivalent to ideal conditions without quantization noise, and the number of required quantization bits that do not degrade the PSNR, was calculated.
Keigo Matsumoto, Yoshiaki Inoue, Yuko Hara-Azumi, Kazuki Maruta, Yu Nakayama, Daisuke Hisano
CCNC3
2023 Convergence Improvement by Parameters Exchange in Asynchronous Decentralized Federated Learning for Non-IID Data
abstract
Asynchronous decentralized federated learning is a promising distributed machine learning framework from the viewpoint of the server cost saving and communication bottleneck mitigation as well as privacy and security protection. Although several methods have been proposed for asynchronous decentralized federated learning, most existing works have a problem of performance degradation for non-independent and identically distributed (non-IID) data. Even the methods aimed for non-IID data do not achieve high accuracy or fast convergence for highly non-IID data in sparse network topologies. To address this issue, we propose a new parameters exchange approach "Skip" and employ it in combination with an existing approach "Swap" to let the distributed models efficiently learn non-local data. Then, we propose two novel asynchronous decentralized federated learning methods, Greedy-Skip & Swap SGD (GSS SGD) and Topology-aware-Skip & Swap SGD (TSS SGD), by combining Skip and Swap in a topology-agnostic and topology-aware fashion, respectively. Our evaluation demonstrated that our TSS SGD outperforms existing methods for highly non-IID data, in terms of the inference accuracy and convergence speed, regardless of the sparsity of topologies.
Asato Yamazaki, Takayuki Nishio, Yuko Hara-Azumi
SEAA3
2023 Power Side-channel Countermeasures for ARX Ciphers using High-level Synthesis
abstract
In the era of Internet of Things (IoT), edge devices are considerably diversified and are often designed using high-level synthesis (HLS) to improve design-productivity. A problem here is that HLS tools were originally developed in a security-unaware fashion, inducing vulnerabilities to power side-channel attacks (PSCA), which is a serious threat in IoT. Although PSCA vulnerabilities induced by HLS tools recently started to be discussed, the effects and applicability of existing methods for PSCA-resistant designs using HLS are limited so far. In this paper, we propose a novel HLS-based design method for PSCA-resistant ciphers in hardware. Particularly focusing on lightweight block ciphers composed of Addition-Rotation-XOR (ARX)-based permutations, we studied the effects of applying ''threshold implementation'', one of the provably secure countermeasures against PSCA, to behavioral descriptions of the ciphers. In addition, we tuned the scheduling optimization of HLS tools that might cause power side-channel leakage. In our experiment, using ARX-based ciphers (Chaskey, Simon, and Speck) as benchmarks, we implemented the unprotected and protected circuit on FPGA and evaluated the PSCA vulnerability using Welch's t-test. The results demonstrated that our proposed method can successfully mitigate vulnerabilities to PSCA for all benchmarks. From these results, we provide further discussion on the direction of PSCA countermeasures based on HLS.
Saya Inagaki, Mingyu Yang 0001, Yang Li 0001, Kazuo Sakiyama, Yuko Hara-Azumi
FPGA5
2023 Dynamic Split Computing-Aware Mixed-Precision Quantization for Efficient Deep Edge Intelligence
abstract
Deploying large deep neural networks (DNNs) on IoT and mobile devices poses a significant challenge due to hardware resource limitations. To address this challenge, an edge-cloud integration technique, called split computing (SC), is attractive in improving the inference time by splitting a single DNN model into two sub-models to be processed on an edge device and a server. Dynamic split computing (DSC) is a further emerging technique in SC to dynamically determine the split point depending on the communication conditions. In this work, we propose a DNN architecture optimization method for DSC. Our contributions are twofold. (1) First, we develop a DSC-aware mixed-precision quantization method exploiting neural architecture search (NAS). By NAS, we efficiently explore the optimal bitwidth of each layer from a huge design space to construct potential split points in the target DNN – with the more potential split points, the DNN architecture can more flexibly utilize one split point depending on the communication conditions. (2) Also, in order to improve the end-to-end inference time, we propose a new bitwidth-wise DSC (BW-DSC) algorithm to dynamically determine the optimal split point among the potential split points in the mixed-precision quantized DNN architecture. Our evaluation demonstrated that our work provides more effective split points than existing works while mitigating the inference accuracy degradation. Specifically in terms of the end-to-end inference time, our work achieved an average of 16.47% and up to 24.36% improvement compared with a state-of-the-art work.
Naoki Nagamatsu, Yuko Hara-Azumi
TrustCom2
2023 Implementation of Deep Joint Source-Channel Coding on 5G Systems for Image Transmission
abstract
Deep joint source-channel coding (JSCC) has been attracting attention for achieving task-oriented communication. It replaces traditional information source coding and channel coding with a deep learning-based autoencoder, directly mapping information sources such as images to IQ symbols. For images, it is claimed to avoid the cliff effect and achieve a higher peak signal noise ratio (PSNR) even in low SNR regions. While related work has assumed various propagation channel models and validated the effectiveness of Deep JSCC, there are few reports confirming its principles through experiments. Specifically, to the best of our knowledge, there are no reported examples of experiments of Deep JSCC in 5G systems. In this paper, we present a proof-of-concept of Deep JSCC in a 5G system. We modified commercially available 5G base stations (gNB) and 5G terminals to enable input and output of IQ data from external devices. We connect the 5G devices using coaxial cables and attenuators, transmit and receive JSCC signals, and evaluate the PSNR. The results demonstrate that even when communicating at power levels lower than the minimum receiver sensitivity specified in the receiver’s datasheet, the image can be successfully restored with less than 1 dB degradation in PSNR compared with the simulation result.
Keigo Matsumoto, Yoshiaki Inoue, Yuko Hara-Azumi, Kazuki Maruta, Yu Nakayama, Yoshinori Shinohara, Hiroki Ikeda, Daisuke Hisano
VTC Fall3
2023 Power Side-channel Attack Resistant Circuit Designs of ARX Ciphers Using High-level Synthesis
abstract
In the Internet of Things (IoT) era, edge devices have been considerably diversified and are often designed using high-level synthesis (HLS) for improved design productivity. However, HLS tools were originally developed in a security-unaware manner, resulting in vulnerabilities to power side-channel attacks (PSCAs), which are a serious threat to IoT systems. Currently, the impact and applicability of existing methods to PSCA-resistant designs using HLS are limited. In this article, we propose an effective HLS-based design method for PSCA-resistant ciphers implemented in hardware. In particular, we focus on lightweight block ciphers composed of addition/rotation/XOR (ARX)-based permutations to study the effects of the threshold implementation (which is one of the provably secure countermeasures against PSCAs) to the behavioral descriptions of ciphers along with the changes in HLS scheduling. The results obtained using Welch’s t-test demonstrate that our proposed method can successfully improve the resistance against PSCAs for all ARX-based ciphers used as benchmarks.
Saya Inagaki, Mingyu Yang 0001, Yang Li 0022, Kazuo Sakiyama, Yuko Hara-Azumi
ACM Trans. Embed. Comput. Syst.5
2022 Hardware SAT Solver-based Area-efficient Accelerator for Autonomous Driving
abstract
Today's embedded systems applications consisting of a variety of tasks are becoming larger and more complex. Hence, when multiple tasks need to be accelerated, designing a dedicated accelerator for each task would be difficult on small devices due to large area overhead. In this study, we propose an efficient accelerator for autonomous driving, which is a theme of a design competition held at International Conference on Field Programmable Technology. Focusing on two key tasks (path planning and object detection), we formulate each of them as a satisfiability problem (SAT) and use a hardware SAT solver as a common accelerator for these tasks. We present efficient problem formulation methods for solving these tasks on a small FPGA. Experimental results show the effectiveness of our work for these tasks.
Yusuke Inuma, Yuko Hara-Azumi
FPT2
2022 A Self-Attention Network for Deep JSCCM: The Design and FPGA Implementation
abstract
The deep joint source-channel coding and modulation (JSCCM) is a promising technology to realize efficient communication over extreme environments such as underwater area. In previous works, it is shown that deep convolutional neural networks (CNN) can successfully learn JSCCM encoder and decoder, outperforming conventional separation-based coding and modulation schemes in low signal-to-noise ratio settings. This paper proposes a new architecture for deep JSCCM based on the self-attention mechanism. We show that the proposed architecture achieves significant performance improvement compared with the CNN-based schemes while requiring a smaller network size in terms of the number of weight parameters. Furthermore, we present efficient hardware implementation of the proposed JSCCM encoder on a field programmable gate array (FPGA). In particular, we demonstrate that a systolic-array-like structure is effective for FPGA implementation of the proposed JSCCM scheme based on the self-attention mechanism.
Shohei Fujimaki, Yoshiaki Inoue, Daisuke Hisano, Kazuki Maruta, Yu Nakayama, Yuko Hara-Azumi
GLOBECOM6
2022 Gossip Swap SGD: Lightweight Decentralized Machine Learning for Non-Homogeneous Data Distribution
abstract
Developing lightweight training approaches for distributed machine learning is of key importance to deploy machine learning on Internet-of-Things (IoT) devices. Gossip SGD is attractive among distributed machine learning methods thanks to its applicability to various network topologies (irrespective of having a central entity) and simple parameters update strategy. However, it is known to be inefficient in non-homogeneous data distribution environments. Our work addresses this issue by introducing “parameters swapping” in the Gossip SGD during parameters update to let devices simply exchange their parameters. Our method, named Gossip Swap SGD, efficiently resolves a cause of insufficient accuracy or delayed convergence in parameters update for non-homogeneous data distribution. Our quantitative evaluation demonstrated that our method not only outperforms exiting methods in non-homogeneous data distribution but also has no degradation even in homogeneous data distribution.
Naoya Yokota, Yuko Hara-Azumi
SoMeT2
2022 Aquatic Fronthaul for Underwater-Ground Communication in 6G Mobile Communications
abstract
Underwater networks are expected to be service platforms for broad-sea and deep-sea activities. The significant challenge of underwater communication has been achieving high-speed and long-distance data transmission due to the high-attenuation and time-varying channel state in underwater environments. It is reasonable to get the underwater data above the water surface for establishing underwater-ground networks. However, it is still an unsolved issue to efficiently establish underwater-ground communication channel. To address this problem, we propose an aquatic fronthaul for underwater-ground communication, where floating aquatic relay nodes relay data from underwater drones/sensors to a ground radio unit. We propose a relocation algorithm for aquatic relay nodes to efficiently reconstruct the network according to the distribution of underwater nodes. The advantage of the proposed algorithm is robustness for the uncertainty of underwater node locations due to the difficulty in underwater localization. The performance of the proposed algorithm was evaluated with multi-agent simulations. The feasibility of the aquatic fronthaul network was confirmed via the experimental results with a Wi-Fi mesh network above the water.
Ayano Higuchi, Erina Takeshita, Daisuke Hisano, Yoshiaki Inoue, Kazuki Maruta, Takayuki Nishio, Yuko Hara-Azumi, Yu Nakayama
VTC Spring7
2022 Stochastic Image Transmission with CoAP for Extreme Environments
abstract
Communication in extreme environments is an important research topic for various use cases including environmental monitoring. A typical example is underwater acoustic communication for 6G mobile networks. The major challenges in such environments are extremely high-latency and high-error rate. They make real-time image transmission difficult using existing communication protocols. This is partly because frequent retransmission in noisy networks increases latency and leads to serious deterioration of real-timeness. To address this problem, this paper proposes a stochastic image transmission with Constrained Application Protocol (CoAP) for extreme environments. The goal of the proposed idea is to achieve approximate real-time image transmission without retransmission using CoAP over UDP. To this end, an image is divided into blocks, and value is assigned for each block based on the requirement. By the stochastic transmission of blocks, the reception probability is guaranteed without retransmission even when packets are lost in networks. We implemented the proposed scheme using Raspberry Pi 4 to demonstrate the feasibility. The performance of the proposed image transmission was confirmed from the experimental results.
Erina Takeshita, Asahi Sakaguchi, Daisuke Hisano, Yoshiaki Inoue, Kazuki Maruta, Yuko Hara-Azumi, Yu Nakayama
VTC Spring6
2022 Real-Time Resource Allocation in Passive Optical Network for Energy-Efficient Inference at GPU-Based Network Edge
abstract
In recent years, the advances in deep learning (DL) technology have greatly improved artificial intelligence (AI)-related research and services. Among them, real-time object recognition using network cameras has become an important technology for various applications. A large number of network cameras are being deployed for real-time object detection using DL models at GPU-based edge servers. A significant issue for widely deploying this type of systems is low-cost network deployment and low-latency data transmission. A promising option for efficiently accommodating numerous network cameras is time- and wavelength-division multiplexed passive optical network (TWDM-PON), which has prevailed in optical access network systems. The key challenge in a GPU-based inference system via TWDM-PON is to optimally allocate upstream wavelengths and bandwidths to enable real-time inference. To address this problem, this article proposes the concept of an inference system in which many cameras upload image data to a GPU-based edge server via TWDM-PON. A real-time resource allocation scheme for TWDM-PON is also proposed to guarantee low latency and time-synchronized data arrival at the edge. We formulated the wavelength and bandwidth allocation problem as a Boolean satisfiability problem (SAT) for fast computation. The performance of the proposed method is verified by computer simulation. The proposed scheme contributes to the increase in the batch size of arriving data at the edge server while ensuring low-latency data transmission. As a consequence, the computational efficiency of the GPU-based inference server is greatly improved by the increase in the batch size of data.
Yu Nakayama, Yukito Onodera, Anh Hoang Ngoc Nguyen, Yuko Hara-Azumi
IEEE Internet Things J.4
2022 Introduction to the Special Section on High-level Synthesis for FPGA: Next-generation Technologies and Applications
abstract
No abstract available.
Christian Pilato, Zhenman Fang, Yuko Hara-Azumi, Jim Hwang
ACM Trans. Design Autom. Electr. Syst.3
2021 Zero Correlation Error: A Metric for Finite-Length Bitstream Independence in Stochastic Computing
abstract
Stochastic computing (SC), with its probabilistic data representation format, has sparked renewed interest due to its ability to use very simple circuits to implement complex operations. Though unlike traditional binary computing, SC needs to carefully handle correlations that exist across data values to avoid the risk of unacceptably inaccurate results. With many SC circuits designed to operate under the assumption that input values are independent, it is important to provide the ability to accurately measure and characterize independence of SC bitstreams. We propose zero correlation error (ZCE), a metric that quantifies how independent two finite-length bitstreams are, and show that it addresses fundamental limitations in metrics currently used by the SC community. Through evaluation at both the functional unit level and application level, we demonstrate how ZCE can be an effective tool for analyzing SC bitstreams, simulating circuits and design space exploration.
Hsuan Hsiao, Joshua San Miguel, Yuko Hara-Azumi, Jason Helge Anderson
ASP-DAC3
2021 An FPGA-based Stochastic SAT Solver Leveraging Inter-Variable Dependencies
abstract
Control rules of Internet of Things (IoT) and embedded systems are often representable in Satisfiability (SAT) problem. This paper proposes a solution search acceleration technique applicable to hardware SAT solvers by leveraging inter-variable dependencies inherited in SAT-encoded realistic applications. We applied our technique on the FPGA implementation of AmoebaSAT, a state-of-the-art stochastic SAT algorithm. Our evaluation on SAT instances from two different IoT-oriented domains demonstrated a significant speedup in solution search against a state-of-the-art hardware SAT solver while incurring almost no effect on hardware implementation compared with the original (dependency-unaware) implementation.
Anh Hoang Ngoc Nguyen, Yuko Hara-Azumi
FPL2
2021 Deep Joint Source-Channel Coding and Modulation for Underwater Acoustic Communication
abstract
Underwater communication is a promising technology to provide ubiquitous network connectivity, where acoustic waves are used as the primary carrier for long-range communication. It has been a challenging research topic to efficiently transmit images with under-water acoustic communication (UAC), due to its inherently narrow bandwidth, strong signal attenuation, time-varying multipath propagation, and low propagation speed. In this paper, we present a new approach to addressing these limitations in UAC, namely the joint source-channel coding and modulation (JSCCM) based on a deep neural network (DNN). We develop a training method of DNN-based encoder and decoder, which directly encode/decode image-pixel values to modulated symbols, unlike conventional separation-based source and channel coding and modulation. Through numerical simulations, the deep JSCCM is confirmed to achieve significantly higher data-rate than conventional schemes.
Yoshiaki Inoue, Daisuke Hisano, Kazuki Maruta, Yuko Hara-Azumi, Yu Nakayama
GLOBECOM4
2021 Space- Time- Domain Adaptive Equalizer Employed Successive Interference Cancellation for Underwater Acoustic Communication
abstract
This paper proposes a space-time-domain successive interference cancellation-based adaptive equalizer (STD-SIC-AE) for underwater acoustic communication (UAC). The demand for high-capacity real-time video transmission underwater has increased for exploring ocean resources and marine research. UAC is capable of long-haul transmission and is the promising means for deep-sea exploration. However, the transmission capacity is limited because of the reflected wave from the sea surface and seafloor. It causes a multipath interference with a long propagation delay that is not easy to remove by the conventional space-time-domain equalizer. This is because that the finite impulse response (FIR) filter requires impractically huge taps. In this paper, by taking advantage of the fact that the interference (delay) wave is a direct wave that has already been received, a replica is generated by the received direct wave and the SIC is operated in the time domain. The numerical simulation verifies that our proposed STD-SIC-AE can significantly improve BER performance in terms of SNR and SIR even under higher-order modulation such as 16QAM.
Kosuke Suzuoki, Daisuke Hisano, Kazuki Maruta, Yoshiaki Inoue, Yuko Hara-Azumi, Yu Nakayama
VTC Fall5
2020 Real-Time Routing for Wireless Relay Fronthaul with Vehicle-Mounted Radio Units
abstract
The concept of vehicle-mounted crowdsourced radio units (CRUs) for a smart city has been proposed to utilize the power of citizens in the deployment of small cells of the centralized radio access network (C-RAN) architecture. Wireless relay fronthaul networking is a promising solution for efficient utilization of vehicle-mounted small cells. However, there have been no routing schemes that can satisfy the strict delay requirements of mobile fronthaul coping with the high dynamicity of vehicles. Thus, this paper proposes a real-time routing scheme for establishing wireless relay fronthaul with vehicle-mounted CRUs. The route optimization is formulated as a boolean satisfiability problem (SAT), and an FPGA-based SAT solver is employed for the fast computation. It can dynamically optimize the forwarding paths in real-time with the constraints of delay requirements. The performance of the proposed routing scheme is confirmed via computer simulations.
Yu Nakayama, Yuko Hara-Azumi, Anh Hoang Ngoc Nguyen, Daisuke Hisano, Yoshiaki Inoue, Takayuki Nishio, Kazuki Maruta
VTC Spring2
2019 SSA-AC: Static Significance Analysis for Approximate Computing
abstract
Recently, the quest to reduce energy consumption in digital systems has been the subject of a number of ongoing studies. One of the most researched focuses is approximate computing (AC) . AC is a new computing paradigm in both hardware and software designs that aim to achieve energy-efficient digital systems. Although a variety of AC techniques have been studied so far, the main question, “How (in which section) can a program or a circuit be approximated?,” has not been answered yet. This work addresses the above issue by developing a software framework Static Significance Analysis for Approximate Computing (SSA-AC) to analyze the target application program and guide the designers to identify parts of the program to which approximation can or cannot be applied. SSA-AC statically analyzes the significance of variables in the precise version of the program and thus needs no trial-and-error evaluation or specific test data. Experimental results show that SSA-AC can successfully extract the significance ranking of inputs/variables to be approximated in much shorter time than existing statistical works that are inevitably data dependent.
Sara Ayman Metwalli, Yuko Hara-Azumi
ACM Trans. Design Autom. Electr. Syst.2
2019 Approximate Data Reuse-based Accelerator Design for Embedded Processor
abstract
Due to increasing diversity and complexity of applications in embedded systems, accelerator designs trading-off area/energy-efficiency and design-productivity are becoming a further crucial issue. Targeting applications in the category of Recognition, Mining, and Synthesis (RMS), this study proposes a novel accelerator design to achieve a good trade-off in efficiency and design-productivity (or reusability) by introducing a new computing paradigm called “approximate computing” (AC). Leveraging from the facts that frequently executed parts of applications (i.e., hotspots) are conventionally the target of acceleration and that RMS applications are error-tolerant and often take similar input data repeatedly, our proposed accelerator reuses previous computational results of similar enough data to reduce computations. The proposed accelerator is composed of a simple controller and a dedicated memory to store limited sets of previous input data with corresponding computational results in a hotspot. Therefore, this accelerator can be applied to different and/or multiple hotspots/applications only through small extension of the controller, to achieve efficient accelerator design and resolve the design-productivity issue. We conducted quantitative evaluations using a representative RMS application (image compression) to demonstrate the effectiveness of our method over conventional ones with precise computing. Moreover, we provide important findings on parameter exploration for our accelerator design, offering a wider applicability of our accelerator to other applications.
Hisashi Osawa, Yuko Hara-Azumi
ACM Trans. Design Autom. Electr. Syst.2
2018 Simple Instruction-Set Computer for Area and Energy-Sensitive IoT Edge Devices
abstract
In this paper, we address a novel instruction set architecture (ISA) to realize a simple (small) yet performance-and-energy efficient processor targeting data-centric applications in the IoT. Focusing on the fact that many of our target applications are mainly composed of light operations which do not need complex arithmetic operations, our proposed processor, SubRISC, supports limited number/types of instructions and simple arithmetic units so that the resources are effectively used (i.e., energy waste can be mitigated) while achieving sufficient performance. We also define efficient instruction formats considering the features of the proposed ISA and target application domains. Our evaluations with two case studies demonstrated the effectiveness of SubRISC against two previous works and revealed important findings on how efficient ISAs desian in ioT.
Kaoru Saso, Yuko Hara-Azumi
ASAP2
2017 CGRA-ME: A unified framework for CGRA modelling and exploration
abstract
Coarse-grained reconfigurable arrays (CGRAs) are a style of programmable logic device situated between FPGAs and custom ASICs on the spectrum of programmability, performance, power and cost. CGRAs have been proposed by both academia and industry; however, prior works have been mainly self-contained without broad architectural exploration and comparisons with competing CGRAs. We present CGRA-ME - a unified CGRA framework that encompasses generic architecture description, architecture modelling, application mapping, and physical implementation. Within this framework, we discuss our architecture description language CGRA-ADL, a generic LLVM-based simulated annealing mapper, and a standard cell flow for physical implementation. An architecture exploration case study is presented, highlighting the capabilities of CGRA-ME by exploring a variety of architectures with varying functionality, interconnect, array size, and execution contexts through the mapping of application benchmarks and the production of standard cell designs.
S. Alexander Chin, Noriaki Sakamoto, Allan Rui, Jim Zhao, Jin Hee Kim, Yuko Hara-Azumi, Jason Helge Anderson
ASAP6
2017 One-instruction set computer-based multicore processors for energy-efficient streaming data processing
abstract
For architecture designs, flexibility of application-dependent optimization for better performance and energy-efficiency and productivity enhanced by application-independent versatility and reusability are both crucial but contradicting issues. In the IoT era, due to a more stringent energy constraint and more application diversity, such issues are becoming more difficult to satisfy. Even recent embedded processors prioritize the design-productivity over flexibility, leading to a lot of energy waste in unused resources for some applications. This paper proposes novel multicore processors to address the above two issues. Our processors are composed of application-independent tiny cores and application-dependent optimizable inter-core communications, which efficiently execute applications on a large amount of streaming data, in a pipeline manner. In this work, we utilize one of the simplest RISC processors, One-Instruction Set Computer (OISC), as a core. Our evaluation demonstrates that our processors outperform an existing RISC processor in terms of performance (throughput) and energy-efficiency, while having sufficient scalability, for two different types of applications.
Minato Yokota, Kaoru Saso, Yuko Hara-Azumi
RSP3
2016 Effect of LFSR seeding, scrambling and feedback polynomial on stochastic computing accuracy
Jason Helge Anderson, Yuko Hara-Azumi, Shigeru Yamashita
DATE2
2015 Timing speculation-aware instruction set extension for resource-constrained embedded systems
abstract
Performance, area, and power are important issues for many embedded systems. One area- and power-efficient way to improve performance is instruction set architecture (ISA) extension. Although existing works have introduced application-specific accelerators co-operating with a basic processor, most of them are still not suitable for embedded systems with stringent resource and/or power constraints because of excess, power-hungry resources in the basic processor. In this paper, we propose ISA extension for such stringently constrained embedded systems. Contrary to previous works, our work rather simplifies the basic processor by replacing original power-hungry resources with power-efficient alternatives. Then, considering the application features (not only input patterns but also instruction sequence), we extend software binary with new instructions executable on the simplified processor. These hardware and software extensions can jointly work well for timing speculation (TS). To the best of our knowledge, this is the first TS-aware ISA extension applicable to embedded systems with stringent area- and/or power-constraints. In our evaluation, we achieved 29.9% speedup in execution time and 1.5× aggressive clock scaling along with 8.7% and 48.3% reduction in circuit area and power-delay product, respectively, compared with the traditional worst-case design.
Tanvir Ahmed 0004, Yuko Hara-Azumi
ASAP2
2015 Profiling-driven multi-cycling in FPGA high-level synthesis
Stefan Hadjis, Andrew Canis, Ryoya Sobue, Yuko Hara-Azumi, Hiroyuki Tomiyama, Jason Helge Anderson
DATE4
2015 Synthesizable-from-C Embedded Processor Based on MIPS-ISA and OISC
abstract
We describe a lightweight open-source MIPS-ISA processor, wherein performance and area can be flexibly traded-off with one another. The processor contains an ultra-low-cost co-processor capable of executing programs comprised of SUBLEQ instructions (subtract and branch if the difference is ≤ 0), which recent work has shown to be sufficient for any computation. Area/performance trade-offs are realized by implementing a user-selectable subset of MIPS instructions with functionally equivalent SUBLEQ sub-routines that run on the coprocessor. Silicon area is reduced as more MIPS instructions are implemented with the co-processor, rather than "natively" using functional units within the host MIPS. The processor is described in the C language and synthesized to an FPGA hardware implementation with high-level synthesis (HLS). Since it is specified at a high level of abstraction, it is straightforward to tailor to any application. As such, the processor can be viewed as a family of processors with different area/performance/power characteristics. In an experimental study, we compare a variety of processor variants, wherein different subsets of MIPS instructions are handled by the co-processor. We also compare the proposed synthesizable processor with a hand-designed 5-pipeline-stage MIPS implementation, and achieve area reductions ranging from 2.5 - 4×.
Tanvir Ahmed 0004, Noriaki Sakamoto, Jason Helge Anderson, Yuko Hara-Azumi
EUC4
2014 Emulator-oriented tiny processors for unreliable post-silicon devices: A case study
abstract
Although various post-silicon devices have been invested in recent years, they still have a major issue of reliability. Because circuit area is an essential factor of reliability, especially for such unreliable post-silicon devices, it is desired to build small circuits which can reuse as many today's application programs as possible even if the performance is not very high. This paper presents the very first work to study novel, efficient techniques of emulating wider-bit guest processors (e.g., 32-bit) on a narrower-bit host processor (e.g., 8-bit) with very limited hardware resources while mitigating performance degradation. We propose three types of emulator-oriented tiny processors varying in available hardware resources and reliability enhancement approaches. Quantitative evaluation and discussions are done for comparing those three processors. We believe that this work will will be a good help of making breakthrough for further development of new device technologies and computers based on them.
Yuko Hara-Azumi, Masaya Kunimoto, Yasuhiko Nakashima
ASP-DAC1
2014 Better-Than-DMR Techniques for Yield Improvement
Shunichi Sanae, Yuko Hara-Azumi, Shigeru Yamashita, Yasuhiko Nakashima
FCCM2
2013 A clique-based approach to find binding and scheduling result in flow-based microfluidic biochips
abstract
Microfluidic biochips have been recently proposed to integrate all the necessary functions for biochemical analysis. There are several types of microfluidic biochips; among them there has been a great interest in flow-based microfluidic biochips, in which the flow of liquid is manipulated using integrated microvalves. By combining several microvalves, more complex resource units such as micropumps, switches and mixers can be built. For efficient execution, the flow of liquid routes in microfluidic biochips needs to be scheduled under some resource constraints or routing constraints. The execution time of the biochemical operations depends on the binding and scheduling results. The most previously developed binding and scheduling algorithms are based on heuristics, and there has been no method to obtain optimal results. Considering the above, this paper proposes an optimal method by casting the problem to a clique problem.
Trung Anh Dinh, Shigeru Yamashita, Tsung-Yi Ho, Yuko Hara-Azumi
ASP-DAC4
2013 VISA synthesis: Variation-aware Instruction Set Architecture synthesis
abstract
We present VISA: a novel Variation-aware Instruction Set Architecture synthesis approach that makes effective use of process variation from both software and hardware points of view. To achieve an efficient speedup, VISA selects custom instructions based on statistical static timing analysis (SSTA) for aggressive clocking. Furthermore, with minimum performance overhead, VISA dynamically detects and corrects timing faults resulting from aggressive clocking of the underlying processor. This hybrid software/hardware approach generates significant speedup without degrading the yield. Our experimental results on commonly used ISA synthesis benchmarks demonstrate that VISA achieves significant performance improvement compared with a traditional deterministic worst case-based approach (up to 78.0%) and an existing SSTA-based approach (up to 49.4%).
Yuko Hara-Azumi, Takuya Azumi, Nikil Dutt
ASP-DAC1
2013 Instruction-set extension under process variation and aging effects
abstract
We propose a novel custom instruction (CI) selection technique for process variation and transistor aging aware instruction-set architecture synthesis. For aggressive clocking, we select CIs based on statistical static timing analysis (SSTA), which achieves efficient speedup during target lifetime while mitigating degradation of timing yield (i.e., probability of satisfying the timing). Furthermore, we consider process variation and aging on not only CIs but also basic instructions (BIs). Even if basic functional units (BFUs), e.g., ALU, get slower due to aging, only a few BIs with critical propagation delay may violate the timing, whereas the other BIs running on the same BFU can still satisfy the timing. We then introduce “customized BFUs”, which execute only such aging-critical BIs. The customized BFUs, used as spare BFUs of the aging-critical BIs, can extend lifetime of the system. Combining the two approaches enables speedup as well as lifetime extension with no or negligibly small area/power overhead. Experiments demonstrate that our work outperforms conventional worst-case work (by an average speedup of about 49%) and existing SSTA-based work (16x or more lifetime extension with comparable speedup).
Yuko Hara-Azumi, Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori
DATE1
2012 Clock-constrained simultaneous allocation and binding for multiplexer optimization in high-level synthesis
abstract
This paper proposes a novel simultaneous allocation and binding method in high-level synthesis, which minimizes the circuit area including multiplexers (MUXs) under a clock constraint. Most existing works on binding minimize MUXs under given allocation by minimizing the number of interconnections, but do not care where the MUXs would be inserted in a circuit. As a result, they cannot guarantee the required clock frequency and often violate the clock constraint. On the contrary, our work globally optimizes binding and allocation for FUs and registers while meeting the clock constraint by considering where MUXs would be inserted. Our work is formulated as an ILP problem. Also, an effective ILP-based heuristic for non-small designs is presented. Experimental results demonstrate that our work satisfies the clock constraint with the minimum circuit area.
Yuko Hara-Azumi, Hiroyuki Tomiyama
ASP-DAC1
2008 CHStone: A benchmark program suite for practical C-based high-level synthesis
abstract
In general, standard benchmark suites are critically important for researchers to quantitatively evaluate their new ideas and algorithms. This paper presents CHStone, a suite of benchmark programs for C-based high-level synthesis. CHStone consists of a dozen of large, easy-to-use programs written in C, which are selected from various application domains. This paper also presents synthesis results which will be served as a baseline for researchers to compare their new techniques with. In addition, we present a case study on function-level transformation using a program in the CHStone suite.
Yuko Hara-Azumi, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada, Katsuya Ishii
ISCAS1
2007 Complexity-constrainted partitioning of sequential programs for efficient behavioral synthesis
abstract
This paper proposes a behavioral level partitioning method for efficient behavioral synthesis from a large sequential program consisting of a set of functions. Our method optimally determines functions to be inlined into the main module and ones to be synthesized into sub modules in such a way that the overall datapath is minimized while the complexity of individual modules is lower than a certain level. The partitioning problem is formulated as an integer programming problem. Experimental results show the effectiveness of the proposed method.
Yuko Hara-Azumi, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada, Katsuya Ishii
ACM Great Lakes Symposium on VLSI1
2006 Function Call Optimization in Behavioral Synthesis
abstract
Behavioral synthesis, which automatically synthesizes an RTL circuit from a sequential program, is one of promising technologies to improve the design productivity. However, behavioral synthesis has not become popular yet in industry since the quality of generated circuits is not satisfactory, especially in the synthesis from the large programs with a number of functions. This paper proposes a method to optimize function calls in behavioral synthesis. We formulate the optimization problem using integer linear programming. Our experimental results show that our method reduces the circuit area by 44.6%, compared with a traditional method
Yuko Hara-Azumi, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada
DSD1