Yang Yi 0002

dblp:19/1969-2 · also Cindy Yang Yi, Yang Cindy Yi · DBLP profile ↗
← Back
71ranked-venue papers
0as first author
28since 2021 · last 2026
0000-0002-1354-0204ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 17 since 2021Computer networks · 25 · 10 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 Finetune Chiplet Design Floorplan via KuanNet
abstract
Chiplet-based architectures require efficient through-silicon via (TSV) assignment to optimize interconnect performance and system integration. Unlike traditional 3-D integrated circuits, heterogeneous chiplet systems demand coordination across dies with varying sizes and functionalities, creating exponentially complex solution spaces that challenge existing optimization methods. This article introduces knowledge-unified attention neural network (KuanNet), a multiagent reinforcement learning (RL) framework integrating echo state networks (ESNs) with attention mechanisms for chiplet TSV assignment. The key innovation is a knowledge-unified architecture with temporal-static decomposition: temporal features shared across agents are processed through both ESN reservoirs and skip connections, while static features remain agent-private, enabling coordinated decisions with temporal memory and spatial awareness. Building on multiagent deep deterministic policy gradient (MADDPG) with K-head attention critics, KuanNet demonstrates superior optimization performance over the state-of-the-art baseline across standard benchmark circuits of varying scale and complexity. Ablation studies validate individual component contributions of the KuanNet architecture.
Yang Yi 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2026 Energy-Efficient Dynamic and Spatiotemporal Spectrum Access via Spiking Reservoir Computing
abstract
This work presents an energy-efficient reinforcement learning (RL) solution based on Neuromorphic Computing (NC) to enable opportunistic spectrum access in partially observable wireless environments. To improve the energy efficiency of the underlying spectrum access strategy, we explore Neuromorphic Computing and adopt spiking neural networks. Additionally, the time-dependent aspect of the problem and the necessity for sample efficiency drive us to liquid state machines, a variant of reservoir computing. Nevertheless, a priori hyperparameter optimization of the spiking reservoir is essential for handling state- and time-varying inputs in RL agents; yet, this can undermine model robustness and impede deployment. In response, we examine homeostatic regulation for self-modulating the small-world reservoir’s dynamics, thereby maintaining desired near-chaotic behavior throughout operation. The RL model for opportunistic spectrum access is evaluated under both dynamic spectrum access (DSA), where agents identify temporal spectrum holes for transmission, and spatiotemporal spectrum access (SSA), where agents also aim to minimize coverage overspill without coordination or sharing location data. Numerical analysis demonstrates that the proposed model outperforms existing learning models in the literature for both DSA and SSA, while significantly reducing power consumption.
Nima Mohammadi, Lingjia Liu 0001, Yifei Song 0001, Yang Yi 0002
IEEE Trans. Wirel. Commun.4
2025 SpikeSpec: An On-Chip Learning Neuromorphic Accelerator for Spectrum Sensing With Triplet-Boosting and Hardware Friendly Loss Function
abstract
Spectrum sensing (SS) is a pivotal function in next-generation (G) multiple-input-multiple-output (MIMO) communication systems, tasked with detecting and characterizing the occupancy or availability of frequency bands within the radio spectrum. Conventional SS techniques are hindered by challenges, such as hardware complexity and the signal-to-noise ratio (SNR) wall, leading to suboptimal performance in environments with high-noise levels. Recurrent neural networks (RNNs), particularly liquid state machines (LSMs), are highly effective for developing energy-efficient accelerators, as they efficiently capture temporal dependencies in primary user frequency spectrums with a reduced number of trainable parameters. This article introduces an innovative field programmable gate array (FPGA) accelerator based on LSMs, featuring fully integrated on-chip learning capabilities for SS. Our accelerator leverages reward-based spike timing-dependent plasticity (R-STDP) to discern temporal correlations within the related frequency band. Although traditional R-STDP methods face convergence difficulties during on-chip learning, the introduced approach overcomes this challenge with a novel, hardware-efficient loss function. This mechanism which is also hardware friendly, facilitates accelerated convergence with reduction of training period of about 46.67% in SS classification for hardware implementation. Moreover, we implemented a fully asynchronous, low-latency, unsupervised triplet-based spike-time-dependent-plasticity (STDP) learning mechanism in the LSM accelerator reservoir, which improves training accuracy by about 3.88% in high-noise channels while enhancing reconfigurability. Furthermore, we investigated various encoder mechanisms to identify the most efficient encoder architecture for our LSM, leading to an accuracy increase of 6.11%. Our optimized LSM architecture achieved 1.27 times and 2.01 times the LUT and register counts, respectively, compared to the basic fixed-reservoir LSM on the Virtex-707 FPGA.
Muhammad Farhan Azmine, Ruizhe Li 0006, Gauri Sharma, Yang Yi 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Dyna-ESN: Efficient Deep Reinforcement Learning for Partially Observable Dynamic Spectrum Access
abstract
This paper focuses on advancing reinforcement learning for challenging environments characterized by partial observability and non-stationarity, such as dynamic spectrum access (DSA). In the literature, the Deep Recurrent Q-Network was introduced to capitalize on the inherent temporal correlations present in DSA. Nevertheless, its practicality is still questionable due to sample inefficiency and slow convergence. We introduce Dyna-ESN, leveraging both model-based and model-free methods by employing Reservoir Computing for generative modeling. Specifically, we utilize Echo State Networks (ESNs) to synthesize samples for enhancing the sample efficiency of a model-free Deep Echo State Q-network, enabling effective operation of agent given limited genuine relevant samples obtained through interaction with environment. To mitigate potential adverse effects of synthetic samples, an evaluation algorithm guides the sample selection process, ensuring reliability. A sample augmentation technique is also introduced to allow agents to collect adequate samples despite controlling the sensing rate and duration of secondary transmissions. Our analysis explores trade-offs between data evaluation and sample efficiency, as well as the bias-variance trade-off of the model, identifying optimal design parameters. Evaluating the performance of Dyna-ESN in DSA scenarios demonstrates its performance benefits over existing methods, paving the way for more efficient and effective techniques in complex dynamic environments.
Hao-Hsuan Chang, Nima Mohammadi, Ramin Safavinejad, Yang Yi 0002, Lingjia Liu 0001
IEEE Trans. Wirel. Commun.4
2024 An In-Memory Power Efficient Computing Architecture with Emerging VGSOT MRAM Device
abstract
In this paper, we present a novel 2-Megabit (Mb) In-Memory Computing (IMC) architecture using advanced Voltage-Gated Spin-Orbit Torque (VGSOT) MRAM. This architecture is designed with 22nm FDSOI technology, and it offers nonvolatile storage, logic operations including AND, NAND, OR, NOR operations, and in-memory dot products for binary neural networks (BNNs). The compact IMC 3T1R bit-cell occupies 0.149 µm2, achieving high write and read speeds of 250-MHz and 1.72-GHz, respectively. Using a BNN model architecture comprising 784 neurons in the input layer, 512 neurons in the hidden layer, and 10 neurons in the output layer, impressive inference accuracies of 95.04% and 84.59% have been achieved when assessing the MNIST and FMNIST datasets, respectively. The proposed VGSOT MRAM architecture’s bit-cell area, read, and write power is 73.34%, 88.82%, and 38.78% less than 2T1R SOT-MRAM respectively.
Md Rubel Sarkar, Shirazush Salekin Chowdhury, Jeffrey S. Walling, Yang Yi 0002
ISCAS4
2024 Towards Energy-Efficient Spiking Neural Networks: A Robust Hybrid CMOS-Memristive Accelerator
abstract
Spiking Neural Networks (SNNs) are energy-efficient artificial neural network models that can carry out data-intensive applications. Energy consumption, latency, and memory bottleneck are some of the major issues that arise in machine learning applications due to their data-demanding nature. Memristor-enabled Computing-In-Memory (CIM) architectures have been able to tackle the memory wall issue, eliminating the energy and time-consuming movement of data. In this work we develop a scalable CIM-based SNN architecture with our fabricated two-layer memristor crossbar array. In addition to having an enhanced heat dissipation capability, our memristor exhibits substantial enhancement of 10% to 66% in design area, power and latency compared to state-of-the-art memristors. This design incorporates an inter-spike interval (ISI) encoding scheme due to its high information density to convert the incoming input signals into spikes. Furthermore, we include a time-to-first-spike (TTFS) based output processing stage for its energy-efficiency to carry out the final classification. With the combination of ISI, CIM and TTFS, this network has a competitive inference speed of 2μs/image and can successfully classify handwritten digits with 2.9mW of power and 2.51pJ energy per spike. The proposed architecture with the ISI encoding scheme can achieve ∼10% higher accuracy than those of other encoding schemes in the MNIST dataset.
Fabiha Nowshin, Hongyu An, Yang Yi 0002
ACM J. Emerg. Technol. Comput. Syst.3
2024 MERRC: A Memristor-Enabled Reconfigurable Low-Power Reservoir Computing Architecture at the Edge
abstract
The massive growth in the Internet of Things (IoT) has led to an increase in demand for devices receiving and transmitting data to and from the cloud during operations. Edge computing has been developed in an attempt to bring computations in proximity to the devices to overcome latency and cost overhead. IoT applications heavily reliant on machine learning (ML) tasks such as image or voice recognition can benefit from edge devices that facilitate real-time operations without costly data transmission back and forth from memory. In this work, we develop MERRC, a memristor-enabled reconfigurable architecture that incorporates processing-in-memory and reservoir computing to carry out ML tasks of image classification at the edge. This design uses a novel masking circuit to allow for image segmentation, which is combined with a delay-based reservoir to form a recurrent neural network. We further implement the final stage of classification using a fabricated memristor crossbar. Our hardware measurement results on the MNIST dataset of the delay-based reservoir with memristor crossbar arrays provide a recognition accuracy of 98%. On a more complex image classification dataset of CIFAR-10, MERRC shows a high accuracy of 88%, further highlighting the edge computing capabilities of the architecture.
Fabiha Nowshin, Yi Huang 0008, Md Rubel Sarkar, Qiangfei Xia, Yang Yi 0002
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 DNN-SNN Co-Learning for Sustainable Symbol Detection in 5G Systems on Loihi Chip
abstract
Performing symbol detection for multiple-input and multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems is challenging and resource-consuming. In this paper, we present a liquid state machine (LSM), a type of reservoir computing based on spiking neural networks (SNNs), to achieve energy-efficient and sustainable symbol detection on the Loihi chip for MIMO-OFDM systems. SNNs are more biological-plausible and energy-efficient than conventional deep neural networks (DNN) but have lower performance in terms of accuracy. To enhance the accuracy of SNNs, we propose a knowledge distillation training algorithm called DNN-SNN co-learning, which employs a bi-directional learning path between a DNN and an SNN. Specifically, the knowledge from the output and intermediate layer of the DNN is transferred to the SNN, and we exploit a decoder to convert the spikes in the intermediate layers of an SNN into real numbers to enable communication between the DNN and the SNN. Through the bi-directional learning path, the SNN can mimic the behavior of the DNN by learning the knowledge from the DNN. Conversely, the DNN can better adapt itself to the SNN by using the knowledge from the SNN. We introduce a new loss function to enable knowledge distillation on regression tasks. Our LSM is implemented on Intel's Loihi neuromorphic chip, a specialized hardware platform for SNN models. The experimental results on symbol detection in MIMO-OFDM systems demonstrate that our LSM on the Loihi chip is more precise than conventional symbol detection algorithms. Also, the model consumes approximately 6 times less energy per sample than other quantized DNN-based models with comparable accuracy.
Shiya Liu, Yibin Liang, Yang Yi 0002
IEEE Trans. Sustain. Comput.3
2023 Invited Paper: Accelerating Next-G Wireless Communications with FPGA-Based AI Accelerators
abstract
5G and beyond 5G wireless communication has revolutionized our daily lives. However, the increased bandwidth and data rate transfer in 5G present challenges, particularly in the domain of data receiving and recovery tasks. Orthogonal frequency-division multiplexing (OFDM) Symbol detection plays a pivotal role in ensuring efficient and error-free transmission in next-G communications. To ensure optimal performance and address potential bottlenecks and errors, propose a Field-Programmable Gate Array (FPGA)-based AI accelerator to accelerate symbol detection, which holds immense potential for enhancing signal recovery in MIMO systems and unlocking the full efficiency of 5G technology. To be more specific, we employ a hardware-verified Echo State Network (ESN) as the symbol detection method in the multiple-input and multiple-output (MIMO)-OFDM system. The ESN, a hardware-efficient Recurrent Neural Network, features a fixed reservoir network structure and fewer trainable parameters in the output layer. To validate our approach in real time, we construct a Software-defined Radio (SDR) platform. This platform allows us to collect datasets in real-world scenarios using antennas. Our experiment showcases a promising average bit-error rate (BER) of. 04 in the performance of our FPGA-based design under realistic conditions. We were able to achieve 3.3 times the throughput with an increase of only 21.4% LUT and 33 % FF in resource utilization for maximal processing speed design which is relatively lower in comparison with the ESN implementation for SISO systems. [1] In addition, we reduced the BRAM memory and DSP IP usage by 50% and 33.3% respectively.
Chunxiao Lin, Muhammad Farhan Azmine, Yang Yi 0002
ICCAD3
2023 Spiking Neural Encoding Schemes and STDP Training Algorithms for Edge Computing
abstract
To enhance real-time data processing, edge computing is utilized in a wider and wider range of applications. For the areas that require large bandwidth and low latency, edge computing even becomes a must. For instance, in the communication area, spectrum sharing within multiple users requires high accuracy of spectrum using prediction as well as low latency. For such tasks, neuromorphic computing, especially spiking neural networks (SNNs), can be a potential method because of its power and silicon area efficiency. In this paper, we have discussed various kinds of spiking neural encoding schemes and their integrated circuit (IC) implementations. We have also summarized the pair-based STDP and the triplet-based STDP learning rule, their mathematical models, and the triplet-based reconfigurable circuit implementation. The Pytorch simulation of different encoding schemes working with two STDP rules for the MNIST and a dynamic spectrum sensing dataset is also presented. It shows that multiplexing ISI-phase encoder can achieve at most 8.9% higher accuracy than other encoders, and TSTDP provides 2.7% higher accuracy than PSTDP for the MNIST dataset. What's more, for the task of spectrum sensing for edge computing, the multiplexing encoding is also 4.3% more accurate, and TSTDP is 0.3% more accurate for the spectrum utilization prediction.
Honghao Zheng, Yang Yi 0002
SEC2
2023 Enabling a New Methodology of Neural Coding: Multiplexing Temporal Encoding in Neuromorphic Computing
abstract
From rate to temporal encoding, spiking information processing has demonstrated advantages across diverse neuromorphic applications. In the aspects of data capacity and robustness, multiplexing encoding outperforms alternative encoding schemes. In this work, we aim to implement a new class of multiplexing temporal encoders, patterning stimuli in multiple timescales to improve the information processing capability, and robustness of systems deployed in noisy environments. Benefitted by the internal reference frame using subthreshold membrane oscillation (SMO), the encoded spike patterns are less sensitive to the input noise, increasing the encoder’s robustness. Our design results in a tremendous saving on power consumption and silicon area compared with the power-hungry analog-to-digital converters. Furthermore, a working prototype of the multiplexing temporal encoder built based on an interspike interval (ISI) encoding scheme is implemented on a silicon chip using the standard 180-nm CMOS process. To the best of our knowledge, our introduced encoder demonstrates the first integrated circuit (IC) implementation of neural encoding with multiplexing topology. Finally, the accuracy and efficiency of our design are evaluated through standard machine learning benchmarks, including Modified National Institute of Standards and Technology (MNIST), Canadian Institute For Advanced Research (CIFAR)-10, Street View House Number (SVHN), and spectrum sensing in high-speed communication networks. While our multiplexing temporal encoder demonstrates a higher classification accuracy across all the benchmarks, the power consumption and dissipated energy per spike reach merely$2.6~\mu \text {W}$and 95 fJ/spike, respectively, with an effective frame rate of 300 MHz. Compared with alternative encoding schemes, our multiplexing temporal encoder achieves at most 100% higher data capacity, 11.4% more accurate in classification, and 25% more robust against noise. Compared with the state-of-the-art designs, our work achieves up to$105 \times $power efficiency without significantly increasing the silicon area.
Honghao Zheng, Kangjun Bai, Yang Yi 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2022 Policy-based Fully Spiking Reservoir Computing for Multi-Agent Distributed Dynamic Spectrum Access
abstract
In the midst of the machine learning revolution, there is hope to thrive the ever-growing demand for limited spectrum resources imposed by the growth of wireless devices with a paradigm shift to more intelligent ways to manage and share the radio spectrum. This requirement mandates very energy-efficient solutions that can tackle the rapid changes of the wireless environment. This work considers spiking neural networks, which have been shown to drastically reduce the energy consumption compared to conventional neural networks in a reinforcement learning setup designed for the dynamic spectrum sharing scenario. Moreover, the temporal aspect of the problem and the necessity of sample efficiency motivates incorporating liquid state machines into this design. However, the agents’ state- and time-variant inputs impose a burden of a posteriori hyperparameter optimization for liquid state machines, rendering the deployment of reliable models whose reservoirs operate in favorable regimes very challenging in such a setting. Therefore, we employ a homeostatic learning rule for adaptively tuning small-world reservoir connections to maintain near-chaotic behavior during operation. Simulation results prove the performance of the introduced solution compared with several existing techniques.
Nima Mohammadi, Lingjia Liu 0001, Yang Yi 0002
ICC3
2022 Diagnosing Clinical Diseases using an Edge-Enabled Deep Learning Technology
abstract
Along with the development of high-speed communication networks, edge-enabled mobile devices have opened new possibilities for diagnosing health conditions or developing suitable treatment plans. While the latest deep learning technology has deployed to restructure and translate complex medical applications, the costly training operation using large-scale neural networks with tremendous amount of data remain the major challenge. In this work, we take advantages of reservoir computing to develop a reliable and low-cost medical diagnostic system for edge-enabled devices. Specifically, an echo state network (ESN) was trained to discover non-obvious correlation and likelihood from biomedical data with respect to various patients. Through the determination of cardiovascular and coronavirus diseases, numerical evaluations demonstrated advantage of ESN against the state-of-the-art. At particularly no computation overhead, ESN precisely described the prediction tasks of health conditions, offering improvements of up to 1000x in sample reduction, 175x in training speedup, and 15 percentage points in prediction accuracy.
Kangjun Bai, Yang Yi 0002
SEC2
2022 Spiking Reservoir Computing for Temporal Edge Intelligence on Loihi
abstract
Low latency and low energy consumption are the indispensable characteristics of Edge Computing applications. With the fusion of Edge Computing and Artificial Intelligence (AI) into Edge Intelligence, this need is more than ever. Of late, Spiking Neural Networks have shown a promise for low latency and low power AI when deployed on a neuromorphic hardware e.g., Intel's Loihi. In this paper, we present a Spiking Reservoir Computing model, based on the Legendre Memory Units which processes temporal data on Loihi hardware. Such a model is greatly suitable for the battery-powered AI enabled edge devices which call for a prompt processing of the temporal sensor-signals with high energy efficiency. We experiment our model with the ECG5000 dataset on the Loihi boards to show its efficacy.
Ramashish Gaurav, Terrence C. Stewart, Yang Yi 0002
SEC3
2022 Real-time Machine Learning for Symbol Detection in MIMO-OFDM Systems
abstract
Recently, there have been renewed interests in applying machine learning (ML) techniques to wireless systems. Nevertheless, ML-based approaches often require a large amount of data in training, and prior ML-based MIMO symbol detectors usually adopt offline learning approaches, which are not applicable to real-time signal processing. This paper adopts echo state network (ESN), a prominent type of reservoir computing (RC), to the real-time symbol detection task in MIMO-OFDM systems. Two novel ESN training methods, namely recursive-least-square and generalized adaptive weighted recursive-least-square, are introduced to enhance the performance of ESN training. Furthermore, a decision feedback mechanism is adopted to improve training efficiency and BER performance. Simulation studies show that the proposed methods perform better than previous conventional and ML-based MIMO symbol detectors. Finally, the effectiveness of our RC-based approach is validated with a software-defined radio (SDR) transceiver and extensive field tests in various real-world scenarios. To the best of our knowledge, this is the first real-time SDR implementation for ML-based MIMO-OFDM symbol detectors. Our work strongly indicates that ML-based signal processing could be a promising and critical approach for future wireless networks.
Yibin Liang, Lianjun Li 0001, Yang Yi 0002, Lingjia Liu 0001
INFOCOM3
2022 Delay-Aware Resource Allocation in Fog-Assisted IoT Networks Through Reinforcement Learning
abstract
Fog nodes in the vicinity of IoT devices are promising to provision low-latency services by offloading tasks from IoT devices to them. Mobile IoT is composed by mobile IoT devices, such as vehicles, wearable devices, and smartphones. Owing to the time-varying channel conditions, traffic loads, and computing loads, it is challenging to improve the Quality of Service (QoS) of mobile IoT devices. As task delay consists of both the transmission delay and computing delay, we investigate the resource allocation (i.e., including both radio resource and computation resource) in both the wireless channel and fog node to minimize the delay of all tasks while their QoS constraints are satisfied. We formulate the resource allocation problem into an integer nonlinear problem, where both the radio resource and computation resource are taken into account. As IoT tasks are dynamic, the resource allocation for different tasks are coupled with each other and the future information is impractical to be obtained. Therefore, we design an online reinforcement learning algorithm to make the suboptimal decision in real time based on the system’s experience replay data. The performance of the designed algorithm has been demonstrated by extensive simulation results.
Qiang Fan 0002, Jianan Bai 0001, Yang Yi 0002, Lingjia Liu 0001
IEEE Internet Things J.4
2022 Differential Privacy Meets Federated Learning Under Communication Constraints
abstract
The performance of federated learning systems is bottlenecked by communication costs and training variance. The communication overhead problem is usually addressed by three communication-reduction techniques, namely, model compression, partial device participation, and periodic aggregation, at the cost of increased training variance. Different from traditional distributed learning systems, federated learning suffers from data heterogeneity (since the devices sample their data from possibly different distributions), which induces additional variance among devices during training. Various variance-reduced training algorithms have been introduced to combat the effects of data heterogeneity, while they usually cost additional communication resources to deliver necessary control information. Additionally, data privacy remains a critical issue in FL and, thus, there have been attempts at bringing Differential Privacy to this framework as a mediator between utility and privacy requirements. This article investigates the tradeoffs between communication costs and training variance under a resource-constrained federated system theoretically and experimentally, and studies how communication reduction techniques interplay in a differentially private setting. The results provide important insights into designing practical privacy-aware federated learning systems.
Nima Mohammadi, Jianan Bai 0001, Qiang Fan 0002, Yifei Song 0001, Yang Yi 0002, Lingjia Liu 0001
IEEE Internet Things J.5
2022 Three-Dimensional Neuromorphic Computing System With Two-Layer and Low-Variation Memristive Synapses
abstract
Three-dimensional integrated circuits (3D-ICs) is a cutting-edge design methodology of placing the circuitry vertically aiming for a high-speed and energy-efficient system with the smallest design area. In this article, a novel 3-D neuromorphic system is proposed and analyzed, which utilizes the fabricated two-layer memristor as the electronic synapses in a spiking neural network (SNN). The two-layer structure of the memristors leads to a significant improvement in the design area ($2\times $), power consumption ($1.48 \times $), and latency ($2.58 \times $), compared to the traditional one-layer configuration. Meanwhile, the heat dissipation layers are added to our memristors reducing 30% cycle-to-cycle switching variation. Our memristive synapses are utilized for storing the exported weights of the SNNs that have threshold function as the activation function. The proposed neuromorphic system is evaluated using a hardware–software co-design approach importing the weights of SNNs into NeuroSIM. The simulation results demonstrate the significant improvement of memristive synapses on design area, power consumption, and latency, compared with the static random-access memory (SRAM) and other state-of-the-art memristive synapses (10%–66%).
Hongyu An, Mohammad Shah Al-Mamun, Marius Orlowski, Lingjia Liu 0001, Yang Yi 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Reservoir Computing Meets Extreme Learning Machine in Real-Time MIMO-OFDM Receive Processing
abstract
In this paper, we consider a real-time deep learning-based symbol detection approach for MIMO-OFDM systems. To exploit the temporal correlation of the wireless channel and the time-frequency structure of OFDM signals, a recurrent neural network (RNN) with deep feedforward output layers is introduced, where the recurrent layers and feedforward output layers are designed to process time-domain and frequency-domain information respectively. Reservoir computing (RC), a special type of RNN, and extreme learning machine (ELM), a special type of feedforward neural network, are chosen as the corresponding building blocks to facilitate over-the-air training. An online training loss objective is introduced to recursively update the neural weights in real-time. We believe this is the first work in the literature to realize real-time machine learning for MIMO-OFDM symbol detection, i.e., conducting NN-based symbol detection on an OFDM symbol basis. We demonstrate that (1) theIEEEstandardized WiFi training sequence can be directly applied as the real-time training sequence (2) the symbol detection performance can be further improved by using our theoretically derived pilot pattern. Evaluation results show that our RC-ELM-based symbol detection method outperforms traditional model-based techniques as well as state-of-the-art learning-based approaches in highly dynamic channel environments for real-time symbol detection.
Lianjun Li 0001, Lingjia Liu 0001, Zhou Zhou 0002, Yang Yi 0002
IEEE Trans. Commun.4
2022 Deep Echo State Q-Network (DEQN) and Its Application in Dynamic Spectrum Sharing for 5G and Beyond
abstract
Deep reinforcement learning (DRL) has been shown to be successful in many application domains. Combining recurrent neural networks (RNNs) and DRL further enables DRL to be applicable in non-Markovian environments by capturing temporal information. However, training of both DRL and RNNs is known to be challenging requiring a large amount of training data to achieve convergence. In many targeted applications, such as those used in the fifth-generation (5G) cellular communication, the environment is highly dynamic, while the available training data is very limited. Therefore, it is extremely important to develop DRL strategies that are capable of capturing the temporal correlation of the dynamic environment requiring limited training overhead. In this article, we introduce the deep echo state Q-network (DEQN) that can adapt to the highly dynamic environment in a short period of time with limited training data. We evaluate the performance of the introduced DEQN method under the dynamic spectrum sharing (DSS) scenario, which is a promising technology in 5G and future 6G networks to increase the spectrum utilization. Compared with conventional spectrum management policy that grants a fixed spectrum band to a single system for exclusive access, DSS allows the secondary system to share the spectrum with the primary system. Our work sheds light on the application of an efficient DRL framework in highly dynamic environments with limited available training data.
Hao-Hsuan Chang, Lingjia Liu 0001, Yang Yi 0002
IEEE Trans. Neural Networks Learn. Syst.3
2021 A Hybrid FPGA-ASIC Delayed Feedback Reservoir System to Enable Spectrum Sensing/Sharing for Low Power IoT Devices ICCAD Special Session Paper
abstract
The delayed feedback reservoir (DFR) network is a delay-dynamic architecture that incorporates time in its training and inference. This quality enables DFR networks to proficiently model time series in a scalable architecture with only one nonlinear neuron. Previous studies have highlighted the accuracy and energy efficiency of DFR networks in ASIC implementations; however, these approaches are limited by hardcoded weights and static reservoir architectures. In this work, we introduce a hybrid FPGA-ASIC DFR system that combines the flexibility of a FPGA platform with the energy efficiency of an ASIC. To be specific, the FPGA allows for dynamic reconfiguration and training of the readout weights during runtime, while the ASIC provides an analog activation function for the single neuron. The accuracy and energy consumption of the introduced system is demonstrated for the applications of NARMA10 as well as MIMO spectrum sensing which is a critical component of dynamic spectrum sharing/access for 5G/beyond-5G systems. Results showcase the potential to enable on-board intelligence for future wireless systems, especially for Internet of Things (IoT) devices in low-power environments.
Osaze Shears, Kangjun Bai, Lingjia Liu 0001, Yang Yi 0002
ICCAD4
2021 Edge Intelligence for Beyond-5G through Federated Learning
Shashank Jere, Yang Yi 0002
SEC2
2021 Multiagent Reinforcement Learning Meets Random Access in Massive Cellular Internet of Things
abstract
Internet of Things (IoT) has attracted considerable attention in recent years due to its potential of interconnecting a large number of heterogeneous wireless devices. However, it is usually challenging to provide reliable and efficient random access control when massive IoT devices are trying to access the network simultaneously. In this article, we investigate methods to introduce intelligent random access management for a massive cellular IoT network to reduce access latency and access failures. Toward this end, we introduce two novel frameworks, namely, local device selection (LDS) and intelligent preamble selection (IPS). LDS enables local communication between neighboring devices to provide cluster-wide cooperative congestion control, which leads to a better distribution of the access intensity under bursty traffics. Taking advantage of the capability of reinforcement learning in developing cooperative multiagent policies, IPS is introduced to enable the optimization of the preamble selection policy in each IoT clusters. To handle the exponentially growing action space in IPS, we design a novel reinforcement learning structure, named branching actor–critic, to ensure that the output size of the underlying neural networks only grows linearly with the number of action dimensions. Simulation results indicate that the introduced mechanism achieves much lower access delays with fewer access failures in various realistic scenarios of interests.
Jianan Bai 0001, Hao Song 0001, Yang Yi 0002, Lingjia Liu 0001
IEEE Internet Things J.3
2021 A Deep Reinforcement Learning Framework for Spectrum Management in Dynamic Spectrum Access
abstract
Dynamic spectrum access (DSA) has the great potential to alleviate spectrum shortage and promote network capacity. However, two fundamental technical issues have to be addressed, namely, interference coordination between DSA users and interference suppression for primary users (PUs). These two issues are very challenging since generally there is no powerful infrastructures in DSA networks to support centralized control. As a result, DSA users have to perform spectrum management individually, including spectrum access and power allocation, without accurate channel state information and centralized control. In this article, a novel spectrum management framework is proposed, in which Q-learning, a type of reinforcement learning, is utilized to enable DSA users to carry out effective spectrum management individually and intelligently. For more efficient process, neural networks (NNs) are employed to implement Q-learning processes, so-called deep Q-network (DQN). Furthermore, we also investigate the optimal way to construct DQN considering both the performance of wireless communications and the difficulty of NN training. Finally, extensive simulation studies are conducted to demonstrate the effectiveness of the proposed spectrum management framework.
Hao Song 0001, Lingjia Liu 0001, Jonathan D. Ashdown, Yang Yi 0002
IEEE Internet Things J.4
2021 A Cost-Efficient Digital ESN Architecture on FPGA for OFDM Symbol Detection
abstract
The echo state network (ESN) is a recently developed machine-learning paradigm whose processing capabilities rely on the dynamical behavior of recurrent neural networks. Its performance outperforms traditional recurrent neural networks in nonlinear system identification and temporal information processing applications. We design and implement a cost-efficient ESN architecture on field-programmable gate array (FPGA) that explores the full capacity of digital signal processor blocks on low-cost and low-power FPGA hardware. Specifically, our scalable ESN architecture on FPGA exploits Xilinx DSP48E1 units to cut down the need of configurable logic blocks. The proposed architecture includes a linear combination processor with negligible deployment of configurable logic blocks and a high-accuracy nonlinear function approximator. Our work is verified with the prediction task on the classical NARMA dataset and a symbol detection task for orthogonal frequency division multiplexing systems using a wireless communication testbed built on a software-defined radio platform. Experiments and performance measurement show that the new ESN architecture is capable of processing real-world data efficiently for low-cost and low-power applications.
Victor M. Gan, Yibin Liang, Lianjun Li 0001, Lingjia Liu 0001, Yang Yi 0002
ACM J. Emerg. Technol. Comput. Syst.5
2021 Robust Deep Reservoir Computing Through Reliable Memristor With Improved Heat Dissipation Capability
abstract
Deep neural networks (DNNs), a brain-inspired learning methodology, requires tremendous data for training before performing inference tasks. The recent studies demonstrate a strong positive correlation between the inference accuracy and the size of the DNNs and datasets, which leads to an inevitable demand for large DNNs. However, conventional memory techniques are not adequate to deal with the drastic growth of dataset and neural network size. Recently, a resistive memristor has been widely considered as the next generation memory device owing to its high density and low power consumption. Nevertheless, its high switching resistance variations (cycle-to-cycle) restrict its feasibility in deep learning. In this work, a novel memristor configuration with the enhanced heat dissipation feature is fabricated and evaluated to address this challenge. Our experimental results demonstrate our memristor reduces the resistance variation by ~ 30% and the inference accuracy increases correspondingly in a similar range. The accuracy increment is evaluated by our deep delay-feed-back reservoir computing (Deep-DFR) model. The design area, power consumption, and latency are reduced by ~48%, ~42%, and ~67%, respectively, compared to the conventional static random-access memory technique (6T). The performance of our memristor is improved at various degrees (~13%-73%) compared to the state-of-the-art memristors.
Hongyu An, Mohammad Shah Al-Mamun, Marius Orlowski, Lingjia Liu 0001, Yang Yi 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 Spatial-Temporal Hybrid Neural Network With Computing-in-Memory Architecture
abstract
Deep learning (DL) has gained unprecedented success in many real-world applications. However, DL poses difficulties for efficient hardware implementation due to the needs of a complex gradient-based learning algorithm and the required high memory bandwidth for synaptic weight storage, especially in today's data-intensive environment. Computing-in-memory (CIM) strategies have emerged as an alternative for realizing energy-efficient neuromorphic applications in silicon, reducing resources and energy required for neural computations. In this work, we exploit a CIM-based spatial-temporal hybrid neural network (STHNN) with a unique learning algorithm. To be specific, we integrate both multilayer perceptron and recurrent-based delay-dynamical system, making the network becomes linear separable while processing information in both spatial and temporal domains, better yet, reducing the memory bandwidth and hardware overhead through the CIM architecture. The prototype fabricated in 180 nm CMOS process is built of fully-analog components, yielding an average on-chip classification accuracy up to 86.9% on handprinted alphabet characters with a power consumption of 33 mW. Beyond that, through the handwritten digit database and the radio frequency fingerprinting dataset, software-based numerical evaluations offer 1.6 -to- 9.8 × and 1.9 -to- 4.4 × speedup, respectively, without significantly degrading its classification accuracy compared to the cutting-edge DL approaches.
Kangjun Bai, Lingjia Liu 0001, Yang Yi 0002
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 RCNet: Incorporating Structural Information Into Deep RNN for Online MIMO-OFDM Symbol Detection With Limited Training
abstract
In this paper, we investigate online learning-based MIMO-OFDM symbol detection strategies focusing on a special recurrent neural network (RNN) - reservoir computing (RC). We first introduce the Time-Frequency RC to take advantage of the structural information inherent in OFDM signals. Using the time domain RC and the time-frequency RC as building blocks, we provide two extensions of the shallow RC to RCNet: 1) Stacking multiple time domain RCs; 2) Stacking multiple time-frequency RCs into a deep structure. The combination of RNN dynamics, the time-frequency structure of MIMO-OFDM signals, and the deep network enables RCNet to handle the interference and nonlinear distortion of MIMO-OFDM signals to outperform existing methods. Unlike most existing NN-based detection strategies, RCNet is also shown to provide a good generalization performance even with a limited online training set (i.e, similar amount of reference signals/training as standard model-based approaches). Numerical experiments demonstrate that the introduced RCNet can offer a faster learning convergence and as much as 20% gain in bit error rate over a shallow RC structure by compensating for the nonlinear distortion of the MIMO-OFDM signal, such as due to power amplifier compression in the transmitter or due to finite quantization resolution in the receiver.
Zhou Zhou 0002, Lingjia Liu 0001, Shashank Jere, Jianzhong Zhang 0002, Yang Yi 0002
IEEE Trans. Wirel. Commun.5
2020 Deep Spiking Delayed Feedback Reservoirs and Its Application in Spectrum Sensing of MIMO-OFDM Dynamic Spectrum Sharing
abstract
In this paper, we introduce a deep spiking delayed feedback reservoir (DFR) model to combine DFR with spiking neuros: DFRs are a new type of recurrent neural networks (RNNs) that are able to capture the temporal correlations in time series while spiking neurons are energy-efficient and biologically plausible neurons models. The introduced deep spiking DFR model is energy-efficient and has the capability of analyzing time series signals. The corresponding field programmable gate arrays (FPGA)-based hardware implementation of such deep spiking DFR model is introduced and the underlying energy-efficiency and recourse utilization are evaluated. Various spike encoding schemes are explored and the optimal spike encoding scheme to analyze the time series has been identified. To be specific, we evaluate the performance of the introduced model using the spectrum occupancy time series data in MIMO-OFDM based cognitive radio (CR) in dynamic spectrum sharing (DSS) networks. In a MIMO-OFDM DSS system, available spectrum is very scarce and efficient utilization of spectrum is very essential. To improve the spectrum efficiency, the first step is to identify the frequency bands that are not utilized by the existing users so that a secondary user (SU) can use them for transmission. Due to the channel correlation as well as users' activities, there is a significant temporal correlation in the spectrum occupancy behavior of the frequency bands in different time slots. The introduced deep spiking DFR model is used to capture the temporal correlation of the spectrum occupancy time series and predict the idle/busy subcarriers in future time slots for potential spectrum access. Evaluation results suggest that our introduced model achieves higher area under curve (AUC) in the receiver operating characteristic (ROC) curve compared with the traditional energy detection-based strategies and the learning-based support vector machines (SVMs).
Kian Hamedani, Lingjia Liu 0001, Shiya Liu, Haibo He, Yang Yi 0002
AAAI5
2020 Deep Reservoir Computing Meets 5G MIMO-OFDM Systems in Symbol Detection
abstract
Conventional reservoir computing (RC) is a shallow recurrent neural network (RNN) with fixed high dimensional hidden dynamics and one trainable output layer. It has the nice feature of requiring limited training which is critical for certain applications where training data is extremely limited and costly to obtain. In this paper, we consider two ways to extend the shallow architecture to deep RC to improve the performance without sacrificing the underlying benefit: (1) Extend the output layer to a three layer structure which promotes a joint time-frequency processing to neuron states; (2) Sequentially stack RCs to form a deep neural network. Using the new structure of the deep RC we redesign the physical layer receiver for multiple-input multiple-output with orthogonal frequency division multiplexing (MIMO-OFDM) signals since MIMO-OFDM is a key enabling technology in the 5th generation (5G) cellular network. The combination of RNN dynamics and the time-frequency structure of MIMO-OFDM signals allows deep RC to handle miscellaneous interference in nonlinear MIMO-OFDM channels to achieve improved performance compared to existing techniques. Meanwhile, rather than deep feedforward neural networks which rely on a massive amount of training, our introduced deep RC framework can provide a decent generalization performance using the same amount of pilots as conventional model-based methods in 5G systems. Numerical experiments show that the deep RC based receiver can offer a faster learning convergence and effectively mitigate unknown non-linear radio frequency (RF) distortion yielding twenty percent gain in terms of bit error rate (BER) over the shallow RC structure.
Zhou Zhou 0002, Lingjia Liu 0001, Vikram Chandrasekhar, Jianzhong Zhang 0002, Yang Yi 0002
AAAI5
2020 Detection Through Deep Neural Networks: A Reservoir Computing Approach for MIMO-OFDM Symbol Detection
abstract
The Reservoir Computing, a neural computing framework suited for temporal information processing, utilizes a dynamic reservoir layer for high-dimensional encoding, enhancing the separability of the network. In this paper, we exploit a Deep Learning (DL)-based detection strategy for Multiple-input, Multiple-output Orthogonal Frequency-Division Multiplexing (MIMO-OFDM) symbol detection. To be specific, we introduce a Deep Echo State Network (DESN), a unique hierarchical processing structure with multiple time intervals, to enhance the memory capacity and accelerate the detection efficiency. The resulting hardware prototype with the hybrid memristor-CMOS co-design provides the in-memory computing and parallel processing capabilities, significantly reducing the hardware and power overhead. With the standard 180nm CMOS process and memristive synapses, the introduced DESN consumes merely 105mW of power consumption, exhibiting 16.7% power reduction compared to shallow ESN designs even with more dynamic layers and associated neurons. Furthermore, numerical evaluations demonstrate advantages of the DESN over state-of-the-art detection techniques in the literate for MIMO-OFDM systems even with a very limited training set, yielding a 47.8% improvement against conventional symbol detection techniques.
Kangjun Bai, Lingjia Liu 0001, Zhou Zhou 0002, Yang Yi 0002
ICCAD4
2020 Quantized Reservoir Computing on Edge Devices for Communication Applications
abstract
With the advance of edge computing, a fast and efficient machine learning model running on edge devices is needed. In this paper, we propose a novel quantization approach that reduces the memory and compute demands on edge devices without losing much accuracy. Also, we explore its application in communication such as symbol detection in 5G systems, attack detection of smart grid, and dynamic spectrum access. Conventional neural networks such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) could be exploited on these applications and achieve state-of-the-art performance. However, conventional neural networks consume a large amount of computation and storage resources, and thus do not fit well to edge devices. Reservoir computing (RC), which is a framework for computation derived from RNN, consists of a fixed reservoir layer and a trained readout layer. The advantages of RC compared to traditional RNNs are faster learning and lower training costs. Besides, RC has faster inference speed with fewer parameters and resistance to overfitting issues. These merits make the RC system more suitable for applications running on edge devices. We apply the proposed quantization approach to RC systems and demonstrate the proposed quantized RC system on Xilinx Zynq®-7000 FPGA board. On the sequential MNIST dataset, the quantized RC system utilizes 62%, 65%, and 64% less of DSP, FF, and LUT, respectively compared to the floating-point RNN. The inference speed is improved by 17 times with an 8% accuracy drop.
Shiya Liu, Lingjia Liu 0001, Yang Yi 0002
SEC3
2020 A Cross-Layer Optimization Framework for Distributed Computing in IoT Networks
abstract
In Internet-of-Thing (IoT) networks, enormous low-power IoT devices execute latency-sensitive yet computation intensive machine learning tasks. However, the energy is usually scarce for IoT devices, especially for some without battery and relying on solar power or other renewables forms. In this paper, we introduce a cross-layer optimization framework for distributed computing among low-power IoT devices. Specifically, a programming layer design for distributed IoT networks is presented by addressing the problems of application partition, task scheduling, and communication overhead mitigation. Furthermore, the associated federated learning and local differential privacy schemes are developed in the communication layer to enable distributed machine learning with privacy preservation. In addition, we illustrate a three-dimensional network architecture with various network components to facilitate efficient and reliable information exchange among IoT devices. Moreover, a model quantization design for IoT devices is illustrated to reduce the cost of information exchange. Finally, a parallel and scalable neuromorphic computing system for IoT devices is established to achieve energy-efficient distributed computing platforms in the hardware layer. Based on the introduced cross-layer optimization framework, IoT devices can execute their machine learning tasks in an energy-efficient way while guaranteeing data privacy and reducing communication costs.
Bodong Shang, Shiya Liu, Sidi Lu, Yang Yi 0002, Weisong Shi, Lingjia Liu 0001
SEC4
2020 Accelerating Model-Free Reinforcement Learning With Imperfect Model Knowledge in Dynamic Spectrum Access
abstract
Current studies that apply reinforcement learning (RL) to dynamic spectrum access (DSA) problems in wireless communications systems mainly focus on model-free RL (MFRL). However, in practice, MFRL requires a large number of samples to achieve good performance making it impractical in real-time applications such as DSA. Combining model-free and model-based RL can potentially reduce the sample complexity while achieving a similar level of performance as MFRL as long as the learned model is accurate enough. However, in a complex environment, the learned model is never perfect. In this article, we combine model-free and model-based RL, and introduce an algorithm that can work with an imperfectly learned model to accelerate the MFRL. Results show our algorithm achieves higher sample efficiency than the standard MFRL algorithm and the Dyna algorithm (a standard algorithm integrating model-based RL and MFRL) with much lower computation complexity than the Dyna algorithm. For the extreme case where the learned model is highly inaccurate, the Dyna algorithm performs even worse than the MFRL algorithm while our algorithm can still outperform the MFRL algorithm.
Lianjun Li 0001, Lingjia Liu 0001, Jianan Bai 0001, Hao-Hsuan Chang, Hao Chen 0010, Jonathan D. Ashdown, Jianzhong Zhang 0002, Yang Yi 0002
IEEE Internet Things J.8
2020 MIMO Spectrum Sensing for Cognitive Radio-Based Internet of Things
abstract
The emerging cognitive radio-based Internet-of-Things (CR-IoT) network provides a novel paradigm solution for IoT devices to efficiently utilize spectrum resources. Spectrum sensing is a critical problem in the CR-IoT network which has been investigated extensively under the Gaussian noise/interference. Since most of the interference in an IoT network is non-Gaussian, in this article, we introduce a novel spectrum sensing method for CR-IoT with additive Gaussian mixture noise/interference. The introduced method maps the observation signal matrix from the original input space to a high-dimensional feature space by a nonlinear Gaussian kernel function and then constructs a kernelized test statistic in the feature space. The approximate analytical expressions of the false alarm and detection probability of the proposed scheme are derived under Gaussian mixture noise, and the decision threshold can be determined according to false alarm probability. The simulation results show that the introduced multiple-input-multiple-output (MIMO) spectrum sensing method achieves good performance under Gaussian mixture noise/interference and significantly outperforms existing detectors.
Junlin Zhang, Lingjia Liu 0001, Mingqian Liu, Yang Yi 0002, Qinghai Yang, Fengkui Gong
IEEE Internet Things J.4
2020 Signal Estimation in Underlay Cognitive Networks for Industrial Internet of Things
abstract
Underlay cognitive radio (CR) holds the promise to address spectrum scarcity and let industrial wireless sensor networks obtain spectrum extension from shared frequency band resources. However, underlay CR devices should be capable of properly adjusting wireless transmission parameters according to the sensing of wireless environments. To realize the goal, in this article, two different signal-to-noise ratio (SNR) estimation methods are proposed for time-frequency overlapped signal estimations in the underlay CR-based industrial Internet of Things (IoT). In the first method, normalized higher order cumulant equations and the theoretical value of normalized higher order cumulants are adopted to estimate the SNR of component signals and the SNR of received signals. In the second one, the power of each component signals and the received signals is estimated based on the second-order time-varying moments. For the performance analysis, the Cramer-Rao lower bound of the SNR estimation for the time-frequency overlapped signals is derived. Simulation results show that the proposed method based on normalized higher order cumulants not only can effectively estimate the SNR of the time-frequency overlapped signals, but also has the strong robustness to the spectrum overlapped rate and the hybrid power ratio. The proposed method with second-order time-varying moments is able to accurately estimate the SNR of the time-frequency overlapped signals effectively, especially in the low-SNR region. These features are extremely useful in the industrial IoT, which usually operate in low-SNR regimes.
Mingqian Liu, Lingjia Liu 0001, Hao Song 0001, Yang Yi 0002, Fengkui Gong
IEEE Trans. Ind. Informatics5
2020 A Training-Efficient Hybrid-Structured Deep Neural Network With Reconfigurable Memristive Synapses
abstract
The continued success in the development of neuromorphic computing has immensely pushed today's artificial intelligence forward. Deep neural networks (DNNs), a brainlike machine learning architecture, rely on the intensive vector-matrix computation with extraordinary performance in data-extensive applications. Recently, the nonvolatile memory (NVM) crossbar array uniquely has unvailed its intrinsic vector-matrix computation with parallel computing capability in neural network designs. In this article, we design and fabricate a hybrid-structured DNN (hybrid-DNN), combining both depth-in-space (spatial) and depth-in-time (temporal) deep learning characteristics. Our hybrid-DNN employs memristive synapses working in a hierarchical information processing fashion and delay-based spiking neural network (SNN) modules as the readout layer. Our fabricated prototype in 130-nm CMOS technology along with experimental results demonstrates its high computing parallelism and energy efficiency with low hardware implementation cost, making the designed system a candidate for low-power embedded applications. From chaotic time-series forecasting benchmarks, our hybrid-DNN exhibits 1.16×-13.77× reduction on the prediction error compared to the state-of-the-art DNN designs. Moreover, our hybrid-DNN records 99.03% and 99.63% testing accuracy on the handwritten digit classification and the spoken digit recognition tasks, respectively.
Kangjun Bai, Qiyuan An, Lingjia Liu 0001, Yang Yi 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Deep-DFR: A Memristive Deep Delayed Feedback Reservoir Computing System with Hybrid Neural Network Topology
abstract
Deep neural networks (DNNs), the brain-like machine learning architecture, have gained immense success in data-extensive applications. In this work, a hybrid structured deep delayed feedback reservoir (Deep-DFR) computing model is proposed and fabricated. Our Deep-DFR employs memristive synapses working in a hierarchical information processing fashion with DFR modules as the readout layer, leading our proposed deep learning structure to be both depth-in-space and depth-in-time. Our fabricated prototype along with experimental results demonstrate its high energy efficiency with low hardware implementation cost. With applications on the image classification, MNIST and SVHN, our Deep-DFR yields a 1.26~7.69X reduction on the testing error compared to state-of-the-art DNN designs.
Kangjun Bai, Qiyuan An, Yang Yi 0002
DAC3
2019 Maximizing System Throughput in D2D Networks Using Alternative DC Programming
abstract
Power control plays an important role in improving the system throughput in communication system since co-channel interference is a major limitation to the system throughput. The power control problem of maximizing the system throughput in the multiuser and multichannel communication system is a highly complicated nonconvex problem since user are interfered with one another if operating in the same wireless channel. We reformulate the nonconvex objective function of this problem as a difference of two convex functions, which is called DC (difference of convex function) programming. To reduce the computation complexity in the high dimensional space, we introduce an alternative power allocation scheme to search in the low dimensional space, where each user updates its power sequentially. A global optimal power allocation is found by utilizing the branch-and- bound algorithm for each user while taking other users' power allocation as constant value. Furthermore, we incorporate each user's maximum power and minimum data rate constraint into the optimization framework. We found that the minimum data rate constraint of each user can be turned into multiple linear inequalities and then be added to the DC programming optimization framework. The simulation results show that our introduced method achieves the highest sum data rate compared to the state-of-the-art methods, including iterative water filling and geometric programming.
Hao-Hsuan Chang, Lingjia Liu 0001, Hao Song 0001, Alex Pidwerbetsky, Allan Berlinsky, Jonathan D. Ashdown, Kurt A. Turck, Yang Yi 0002
GLOBECOM8
2019 Monolithic 3D neuromorphic computing system with hybrid CMOS and memristor-based synapses and neurons
Hongyu An, M. Amimul Ehsan, Fangyang Shen, Yang Yi 0002
Integr.5
2019 Distributive Dynamic Spectrum Access Through Deep Reinforcement Learning: A Reservoir Computing-Based Approach
abstract
Dynamic spectrum access (DSA) is regarded as an effective and efficient technology to share radio spectrum among different networks. As a secondary user (SU), a DSA device will face two critical problems: 1) avoiding causing harmful interference to primary users (PUs) and 2) conducting effective interference coordination with other SUs. These two problems become even more challenging for a distributed DSA network where there is no centralized controllers for SUs. In this paper, we investigate communication strategies of a distributive DSA network under the presence of spectrum sensing errors. To be specific, we apply the powerful machine learning tool, deep reinforcement learning (DRL), for SUs to learn “appropriate” spectrum access strategies in a distributed fashion assuming NO knowledge of the underlying system statistics. Furthermore, a special type of recurrent neural network, called the reservoir computing (RC), is utilized to realize DRL by taking advantage of the underlying temporal correlation of the DSA network. Using the introduced machine learning-based strategy, SUs could make spectrum access decisions distributedly relying only on their own current and past spectrum sensing outcomes. Through extensive experiments, our results suggest that the RC-based spectrum access strategy can help the SU to significantly reduce the chances of collision with PUs and other SUs. We also show that our scheme outperforms the myopic method which assumes the knowledge of system statistics, and converges faster than the Q-learning method when the number of channels is large.
Hao-Hsuan Chang, Hao Song 0001, Yang Yi 0002, Jianzhong Zhang 0002, Haibo He, Lingjia Liu 0001
IEEE Internet Things J.3
2019 Green Massive Traffic Offloading for Cyber-Physical Systems over Heterogeneous Cellular Networks
Rachad Atat, Lingjia Liu 0001, Jinsong Wu 0001, Jonathan D. Ashdown, Yang Yi 0002
Mob. Networks Appl.5
2019 QoS-Aware D2D Cellular Networks With Spatial Spectrum Sensing: A Stochastic Geometry View
abstract
Spectrum access and interference management are amongst the most challenging issues in device-to-device (D2D) cellular networks. In order to address these issues, this paper introduces spatial spectrum sensing (SSS) for D2D cellular networks to facilitate cellular spectrum sharing by D2D users while providing a quality of service guarantee for cellular users. In order to assess the performance of the proposed scheme, we adopt a stochastic geometry approach in which the locations of base stations and D2D devices are modeled as independent Poisson point processes (PPPs). Assuming that the locations of the active cellular transmitters form another independent PPP, we characterize the area spectral efficiency of D2D networks under cellular users' outage probability constraint. The use of SSS prohibits D2D transmissions around the active cellular users because of which the locations of the active D2D transmitters are modeled as a Poisson hole process driven by the PPP of active cellular user locations. Our analysis carefully accounts for this spatial separation between active cellular users and active D2D devices. Extensive simulation and numerical results are presented to verify our analysis and demonstrate the advantages of SSS-based D2D cellular networks.
Hao Chen 0010, Lingjia Liu 0001, Harpreet S. Dhillon, Yang Yi 0002
IEEE Trans. Commun.4
2018 Enabling a new era of brain-inspired computing: energy-efficient spiking neural network with ring topology
abstract
The reservoir computing, an emerging computing paradigm, has proven its benefit to multifarious applications. In this work, we successfully designed and fabricated an analog delayed feedback reservoir (DFR) chip. Measurement results demonstrate its rich dynamic behaviors and high energy efficiency. System performance, as well as the robustness, are evaluated. The application of video frame recognition is investigated using a hybrid neural network, which employs the multilayer perceptron (MLP) training model as the readout layer of our designed DFR system, and yields 98% classification accuracy. Compared to results of using the MLP training only, our hybrid training model exhibits much higher recognition rate and accuracy.
Kangjun Bai, Kian Hamedani, Yang Yi 0002
DAC4
2018 Realizing Green Symbol Detection via Reservoir Computing: An Energy-Efficiency Perspective
abstract
Reservoir Computing (RC) is a class of machine learning approaches that is suitable for prediction tasks with low computational complexity. In this paper, an RC-based symbol detection for MIMO- OFDM systems is presented where RC is realized through the echo state network (ESN). Detailed energy-efficiency analysis is conducted to characterize the energy-efficiency of the introduced symbol detector. To be specific, the transmit power, the circuit power, and the computational power at both transmitter and receiver are jointly considered for the energy-efficiency analysis. The overall system energy-efficiency as well as the receiver energy-efficiency of the introduced RC-based symbol detector are compared with those of the popular linear minimum mean squared error (LMMSE)-based approach. Simulation and numerical results show that the RC-based symbol detector is a ``green'' solution compared to the traditional LMMSE-based method with lower energy consumption per information bit.
Rubayet Shafin Bradley Shafin, Lingjia Liu 0001, Jonathan D. Ashdown, John D. Matyjas, Michael J. Medley, Bryant T. Wysocki, Yang Yi 0002
ICC7
2018 Q-Learning for Non-Cooperative Channel Access Game of Cognitive Radio Networks
abstract
This paper investigates the channel access problem of cognitive radio networks. In the cognitive radio network, communication channels are assigned to primary users with priority while secondary users are able to detect the spectrum holes and switch among the channels for data transmission opportunities. The channel access problem of this kind of system can be formulated as a non-cooperative game. However, in prior works, the secondary users are usually assumed to be able to switch to any channel instantaneously, which is not possible in reality because the channel switching will incur transmission delays. In this paper, we formulate the channel access problem as a non-cooperative game where each channel can be used by only one user at a time. Moreover, considering the transmission delays, we limit the channel switching distance of the secondary users to a certain scope. In this case, the optimal channel access policy of each secondary user will depend on the long-term behaviors of primary users as well as the actions of other secondary users. For this non-cooperative game, we propose a multiagent Q-learning algorithm which requires neither the prior knowledge of channel dynamics nor the negotiations among players. Simulation examples are provided to demonstrate the effectiveness of the algorithm.
He Jiang 0004, Haibo He, Lingjia Liu 0001, Yang Yi 0002
IJCNN4
2018 A Physical Layer Security Scheme for Mobile Health Cyber-Physical Systems
abstract
Mobile health (m-Health) is one potential application of cyber-physical systems, where biomedical sensors and mobile devices interact tightly together to transmit medical data to an m-Health server. In this paper, we consider a three-tier hierarchical m-Health system: 1) the sensor network tier capturing vital signals; 2) the mobile computing network tier processing and routing the sensed data to a fixed remote location; and 3) the back-end network tier processing and analyzing the sensed medical data along with patient's medical history. Based on this architecture, a physical layer security scheme for the second tier is developed and network performance of the introduced scheme is analyzed under different metrics using stochastic geometry. To be specific, secure transmission range and average end-to-end delay are analyzed for two different strategies: the mobile device transmits: 1) to the nearest neighbor and 2) to the furthest neighbor. Furthermore, we consider the cases of full knowledge on eavesdroppers' locations and when such information is unavailable. Results show that transmitting to the nearest neighbor achieves the highest secure transmission distance with the lowest mean delay when full information on eavesdroppers is available.
Rachad Atat, Lingjia Liu 0001, Jonathan D. Ashdown, Michael J. Medley, John D. Matyjas, Yang Yi 0002
IEEE Internet Things J.6
2018 DFR: An Energy-efficient Analog Delay Feedback Reservoir Computing System for Brain-inspired Computing
abstract
Neuromorphic computing, which is built on a brain-inspired silicon chip, is uniquely applied to keep pace with the explosive escalation of algorithms and data density on machine learning. Reservoir computing, an emerging computing paradigm based on the recurrent neural network with proven benefits across multifaceted applications, offers an alternative training mechanism only at the readout stage. In this work, we successfully design and fabricate an energy-efficient analog delayed feedback reservoir (DFR) computing system, which is built upon a temporal encoding scheme, a nonlinear transfer function, and a dynamic delayed feedback loop. Measurement results demonstrate its high energy efficiency with rich dynamic behaviors, making the designed system a candidate for low power embedded applications. The system performance, as well as the robustness, are studied and analyzed through the Monte Carlo simulation. The chaotic time series prediction benchmark, NARMA10, is examined through the proposed DFR computing system, and exhibits a 36%−85% reduction on the error rate compared to state-of-the-art DFR computing system designs. To the best of our knowledge, our work represents the first analog integrated circuit (IC) implementation of the DFR computing system.
Kangjun Bai, Yang Yi 0002
ACM J. Emerg. Technol. Comput. Syst.2
2018 A Novel Approach for Using TSVs As Membrane Capacitance in Neuromorphic 3-D IC
abstract
An advanced neurophysiological computing system can incorporate a 3-D integration system composed of emerging nano-scale devices to provide massive parallelism having high speed, low cost, and energy efficient hardware implementation. Due to process technology constraints, a certain amount of redundant through silicon vias (TSVs) and dummy TSVs are always required in a 3-D integrated system. In this paper, we propose to use these redundant and dummy TSVs to supply the neuronal membrane capacitance that maps the membrane electrical activity in a hybrid 3-D neuromorphic system. This proposition could also serve the need of neuronal ion transportation dynamics. We also investigate two new methodologies that could significantly enhance the TSV capacitance in a 3-D neuromorphic system. The capacitance of these enhanced TSVs is studied with analytical models; the accuracy of the models are evaluated against 3-D field extracted values. The advantage of using the TSVs to mimic membrane capacitance in a 3-D neuromorphic chip is demonstrated through comparisons of both silicon area and energy consumption against their 2-D counterpart designs.
M. Amimul Ehsan, Hongyu An, Yang Yi 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Reservoir Computing Meets Smart Grids: Attack Detection Using Delayed Feedback Networks
abstract
A new method for attack detection of smart grids with wind power generators using reservoir computing (RC) is introduced in this paper. RC is an energy-efficient computing paradigm within the field of neuromorphic computing and the delayed feedback networks (DFNs) implementation of RC has shown superior performance in many classification tasks. The combination of temporal encoding, DFN, and a multilayer perceptron (MLP) as the output readout layer is shown to yield performance improvement over existing attack detection methods such as MLPs, support vector machines (SVM), and conventional state vector estimation (SVE) in terms of attack detection in smart grids. The proposed algorithms are shown to be more robust than MLP and SVE in dealing with different variables such as the amplitude of the attack, attack types, and the number of compromised measurements in smart grids. The attack detection rate for the proposed RC-based system is higher than 99%, based on the accuracy metric for the average of 10 000 simulations.
Kian Hamedani, Lingjia Liu 0001, Rachad Atat, Jinsong Wu 0001, Yang Yi 0002
IEEE Trans. Ind. Informatics5
2018 Brain-Inspired Wireless Communications: Where Reservoir Computing Meets MIMO-OFDM
abstract
Reservoir computing (RC) is a class of neuromorphic computing approaches that deals particularly well with time-series prediction tasks. It significantly reduces the training complexity of recurrent neural networks and is also suitable for hardware implementation whereby device physics are utilized in performing data processing. In this paper, the RC concept is applied to detecting a transmitted symbol in multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. Due to wireless propagation, the transmitted signal may undergo severe distortion before reaching the receiver. The nonlinear distortion introduced by the power amplifier at the transmitter may further complicate this process. Therefore, an efficient symbol detection strategy becomes critical. The conventional approach for symbol detection at the receiver requires accurate channel estimation of the underlying MIMO-OFDM system. However, in this paper, we introduce a novel symbol detection scheme where the estimation of the MIMO-OFDM channel becomes unnecessary. The introduced scheme utilizes an echo state network (ESN), which is a special class of RC. The ESN acts as a black box for system modeling purposes and can predict nonlinear dynamic systems in an efficient way. Simulation results for the uncoded bit error rate of nonlinear MIMO-OFDM systems show that the introduced scheme outperforms conventional symbol detection methods.
Somayeh Mosleh, Lingjia Liu 0001, Cenk Sahin, Yahong Rosa Zheng, Yang Yi 0002
IEEE Trans. Neural Networks Learn. Syst.5
2018 Enabling Sustainable Cyber Physical Security Systems through Neuromorphic Computing
abstract
As the novel paradigm in the field of machine learning, reservoir computing possesses exceptional performance, e.g., energy efficiency, in tasks in which the traditional von Neumann computing systems cannot incorporate. This makes reservoir computing an ideal candidate to enable the sustainable development of cyber-physical systems (CPS). In the realm of CPS, the tight interaction among physical objects places security threats under the spotlight of attention. For such systems, especially the power grid network, false data injection could potentially lead to catastrophic consequences such as blackouts in large geographical areas. In this paper, we will introduce a reservoir computing architecture, the delayed feedback system, and apply the reservoir computing architecture for anomaly detection. To be specific, detailed design of the three imperative components in the delayed feedback system will be discussed and the corresponding energy efficiency performance will be analyzed. The application of the reservoir computing architecture to anomaly detection in a smart grid network will be introduced.
Lingjia Liu 0001, Chenyuan Zhao, Kian Hamedani, Rachad Atat, Yang Yi 0002
IEEE Trans. Sustain. Comput.6
2017 Adaptation of Enhanced TSV Capacitance as Membrane Property in 3D Brain-inspired Computing System
abstract
Neurophysiological architecture using 3D integration technology offers a high device interconnection density as well as fast and energy efficient links among the neuron and synapses layers. In this paper, we propose to reconfigure the Through-Silicon-Vias (TSVs) to serve as the neuronal membrane capacitors that map the membrane electrical activities in a hybrid 3D neuromorphic system. We also investigate new methodology that could significantly enhance the TSV capacitance to achieve a high efficiency of signal processing through membrane. An optimal CAD framework is designed to optimally utilize such TSV devices, and resolve the signal-integrity issues arising at fast data rates during massively parallel data transmissions. The electrical performance of the 3D neuromorphic chip is compared against the ones of the 2D counterpart design to demonstrate the advantages of our design and methodology.
M. Amimul Ehsan, Hongyu An, Yang Yi 0002
DAC4
2017 Neuromorphic 3D Integrated Circuit: A Hybrid, Reliable and Energy Efficient Approach for Next Generation Computing
abstract
In this paper, we proposed to use 3D integration technology to create a neuromorphic hardware system that is compatible with current technology, provides high system speed, high density, massively parallel processing, low power consumption, and small footprint. The Through Silicon Vias (TSVs) used in the 3D neuromorphic structure provide high density integration and energy efficient links for transferring information through multiple neuron layers. This work details how a 3D neuromorphic system is benefited from the redundant TSV with substantial design-area reduction. We discussed the yield and reliability issues and explained the impact in neuromorphic 3D system design. A spiking neuron model is developed for the proposed 3D system. Furthermore, a new methodology have been proposed by introducing oxide around the bump that could significantly enhance the TSV capacitance in 3D Neuromorphic Computing (NC) system.
M. Amimul Ehsan, Yang Yi 0002
ACM Great Lakes Symposium on VLSI3
2017 Analog hardware implementation of spike-based delayed feedback reservoir computing system
abstract
The rate of enhancement is starting to saturate and slow down which indicates the end of Moore's prediction due to the fundamental performance limits of the chips. The need of breaking through the barrier has directed researchers into several directions, for instance, novel computing architecture. Reservoir computing, a novel concept in the field of machine learning, has emerged over the past few years. Combined the memory and spatio-temporal processing of recurrent neural networks, reservoir computing possesses the capability of processing temporal information. In this paper, we present an analog hardware implementation of delayed feedback reservoir computing system. We build a new class of computationally efficient spike timing-dependent encoders and delay-based reservoirs within reservoir networks. This approach allows us to avoid using power-consuming analog-to-digital converters (ADCs) and operational amplifiers (Op-AMPs), resulting in significant savings in power requirements and design area.
Chenyuan Zhao, Kian Hamedani, Yang Yi 0002
IJCNN4
2017 Energy Harvesting-Based D2D-Assisted Machine-Type Communications
abstract
Supporting massive numbers of machine-type communication (MTC) devices poses several challenges for future 5G networks, including network control, scheduling, and powering these devices. A potential solution is to offload MTC traffic onto device-to-device (D2D) communication links to better manage radio resources and reduce MTC devices' energy consumption. However, this approach requires D2D users to use their own limited energy to relay MTC traffic, which may be undesirable. This motivates us to exploit recent advancements in RF energy harvesting for powering D2D relay transmissions. In this paper, we consider a D2D communication as an underlay to the cellular network, where D2D users access a fraction of the spectrum occupied by cellular users. This underlay model presents a fundamental trade-off: to protect cellular users, the spectrum available to D2D users needs to be reduced, which limits the number of D2D transmissions, but increases the amount of time that D2D users can spend harvesting energy to support MTC traffic. We study this trade-off by characterizing the spectral efficiency of MTC, D2D, and cellular users using stochastic geometry. The optimal spectrum partition factor is characterized to achieve fairness and balance in the network, while increasing the average MTC spectral efficiency.
Rachad Atat, Lingjia Liu 0001, Nicholas Mastronarde, Yang Yi 0002
IEEE Trans. Commun.4
2017 Interspike-Interval-Based Analog Spike-Time-Dependent Encoder for Neuromorphic Processors
abstract
Von Neumann bottleneck, which refers to the limited throughput between the CPU and memory, has already become a major factor hindering the technical advances of computing systems. In recent years, neuromorphic systems have started to gain the increasing attentions as compact and energy-efficient computing platforms. As one of the most crucial components in the neuromorphic computing systems, neural encoder transforms the stimulus (input signals) into spike trains. In this paper, we adapt the temporal encoding scheme of interspike intervals (ISIs) and present an analog temporal neural encoder with its verification and recovery schemes. The proposed neural encoder allows efficient mapping of signal amplitude information into a spike-time sequence that represents the input data and offers perfect recovery for band-limited stimuli. With the novel iterative structure, the number of spikes increases exponentially with the number of neurons. From the measurements obtained from the fabricated neural encoder chip, our temporal encoder with ISI encoding is proved to be robust and error tolerant.
Chenyuan Zhao, Yang Yi 0002, Xin Fu 0001, Lingjia Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Cooperative Retransmission for Massive MTC under Spatiotemporally Correlated Interference
abstract
In a massive machine type communication (massive MTC) network, wireless connections between machine type devices (MTDs) and eNodeBs are unreliable due to the interference caused by uncoordinated access of other MTDs. Furthermore, it is with a high probability that the retransmission from the outage source MTD will fail again due to the spatiotemporally correlated interference. In this paper, we design and analyze a location-based cooperative strategy to improve the performance of massive MTC networks. In the cooperative strategy, an inactive MTD is selected as a relay if it has successfully decoded the packet and if it is located within a circular area around the eNodeB. Considering the spatial and temporal correlation of interference, the outage probability of the designed cooperative strategy is derived using stochastic geometry. Both the simulation and numerical results demonstrate that spatiotemporal correlation of interference significantly affects the performance analysis of cooperative massive MTC networks and our designed cooperative strategy can significantly reduce the outage probability compared to conventional retransmission.
Hao Chen 0010, Lingjia Liu 0001, Nicholas Mastronarde, Liangping Ma, Yang Yi 0002
GLOBECOM5
2016 Improving Spectral Efficiency of D2D Cellular Networks through RF Energy Harvesting
abstract
In this paper, we consider device-to-device (D2D) communication underlaying cellular networks, where D2D users harvest radio frequency (RF) power from uplink cellular transmissions. This paper addresses two important issues for energy harvesting-based D2D cellular networks. The first is how the energy harvested from cellular ambient RF signals affects the D2D spectral efficiency. Using tools from stochastic geometry, we investigate this issue by first characterizing the transmission probability of D2D transmitter as the probability of having enough battery power, and then obtaining analytic expressions of the spectral efficiency for D2D and cellular networks. The second issue is related to the cellular spectral efficiency: the more RF energy D2D users can harvest, the higher their transmission probability becomes, which leads to generating more interference in the cellular network. We carry out simulations to understand this impact of RF energy harvesting on the spectral efficiency of both D2D and cellular networks. Results show promising improvement to the whole network, in terms of the weighted spectral efficiency, when employing RF harvesting technology and when there are enough available channels in the network.
Rachad Atat, Lingjia Liu 0001, Jonathan D. Ashdown, Michael J. Medley, John D. Matyjas, Yang Yi 0002
GLOBECOM6
2016 Privacy Protection Scheme for eHealth Systems: A Stochastic Geometry Approach
abstract
The technological advancements in the health care system have made possible the massive integration of biomedical sensors for monitoring patients' health and disease progression. In this paper, we consider a three tier medical body area network (MBAN): intra-MBAN, inter-MBAN, and beyond-MBAN. The intra-MBAN transmits sensors' data to a controller, which in turn transmits them in the inter-MBAN tier to an access device like a PDA or a tablet device, usually connected to a patient's medical database. The access device then serves as a mean of communication between intra-MBAN and beyond-MBAN to access the hospital information systems. This widely deployed design in hospitals places security and privacy violation threats on the spotlight of attention, especially in the inter-MBAN tier. This has motivated us to optimize the average MBAN controller transmit power that minimizes the probability of eavesdroppers overhearing the communication, using tools from stochastic geometry. We analyze this privacy protection scheme for eHealth systems through simulations. Results show that the proposed scheme achieves higher privacy protection, but at the expense of reduced coverage.
Rachad Atat, Lingjia Liu 0001, Yang Yi 0002
GLOBECOM3
2016 Coordinated Data Assignment: A Novel Scheme for Big Data over Cached Cloud-RAN
abstract
A cloud radio access network (Cloud-RAN) is a network architecture that holds onto the promise of meeting the explosive growth of mobile data traffic. Cloud-RAN consists a central processor (CP) connecting to multiple multi-antenna base stations (BSs) via finite-capacity backhaul links. To reduce the backhaul traffic, BS-level caching technique is utilized in which the popular contents are pre-fetched in memories at each BS. This technique plays an important role in future wireless big data processing due to its simplicity, low cost, and natural integration with big data analytical tools. Considered the tradeoff between the backhaul and the transmission power cost, in this paper we define the network cost of the system as a normalized weighted sum. The problem of minimizing the network cost with respect to both the precoding matrix and the cache placement matrix is formulated subject to the quality of service (QoS), peak transmission power, and cache capacity constraints. The l0-norm in the objective function along with the QoS constraints renders the optimization problem non-convex. Additionally, since the entries of cache placement matrix take binary values, the optimization problem falls into a mixed integer nonlinear programming (MINLP) which is a NP-hard problem. An iterative coordinated data assignment algorithm is introduced which achieves a stationary point of the problem. Simulations are conducted to illustrate the performance of introduced algorithm. It suggests that the introduced scheme can significantly reduce the total network cost of the underlying Cloud-RAN network and demonstrate the importance of considering the designing of cache placement matrix.
Somayeh Mosleh, Lingjia Liu 0001, Hongyan Hou, Yang Yi 0002
GLOBECOM4
2016 Making neural encoding robust and energy efficient: an advanced analog temporal encoder for brain-inspired computing systems
abstract
Neural encoder is one of the key components in neuromorphic computing systems, whereby sensory information is transformed into spike coded trains. The design of temporal encoder has attracted a widespread attention in the field of neuromorphic computing in the past few years. The information in the temporal encoding scheme with inter-spike intervals can arise from correlations between spike times, which could not be incorporated in the traditional rate encoding scheme. In this paper, we propose a robust and energy efficient analog implementation of the spiking temporal encoder. We pattern the neural activities across multiple timescales and encode the sensory information using time dependent temporal scales. The concept of iteration structure is introduced to construct a neural encoder that greatly increases the information process ability of the proposed temporal encoder. Integrated with iteration technique and operational-amplifier-free design, the error rate of the output temporal codes is reduced to an extremely low level. A lower sampling rate accompanied by additional verification spikes is introduced in the schemes, which significantly reduces the power consumption of the encoding system. To the best of our knowledge, our proposed neuron circuit is the first analog hardware implementation of the neural encoder that could present the sensory data using inter-spike interval temporal encoding scheme. The simulation and measurement results show the proposed temporal encoder exhibits not only energy efficiency but also high accuracy.
Chenyuan Zhao, Yang Yi 0002
ICCAD3
2016 An energy efficient decoding scheme for nonlinear MIMO-OFDM network using reservoir computing
abstract
Reservoir computing (RC) is attracting widespread attention in several signal processing domains owing to its nonlinear stateful computation. It deals particularly well with time-series prediction tasks and reduces training complexity over recurrent neural networks. It is also suitable for hardware implementation whereby device physics are utilized in performing data processing. In this paper, the RC concept is applied to modeling a Multiple-Input Multiple-Output Orthogonal Frequency Division Multiplexing (MIMO-OFDM) system. Due to the harsh propagation environment, the transmitted signal undergoes severe distortion that must be compensated for at the receiver. The nonlinear distortion introduced by the power amplifier at the transmitter further complicates this process. An effective channel estimation scheme is therefore required. In this paper, we introduce a MIMO-OFDM channel estimation scheme utilizing Echo State Network (ESN). Echo State Networks are powerful recurrent neural networks that can predict time-series very well. They acts as a black-box for system modeling purposes and models nonlinear dynamic systems efficiently. Simulation results for the bit error rate of the nonlinear MIMO-OFDM system show that the introduced channel estimator outperforms commonly used channel estimation schemes.
Somayeh Mosleh, Cenk Sahin, Lingjia Liu 0001, R. Y. Zheng, Yang Yi 0002
IJCNN5
2016 Mitigating the Impact of Hardware Variability for GPGPUs Register File
abstract
As technology keeps scaling down, hardware variability, such as process variations (PV) and negative bias temperature instability (NBTI), emerges as a growing challenge in the modern GPGPUs (general-purpose computing on graphics processing units). PV induces significant delay variations statically, while NBTI dynamically slows down the GPGPUs. Each computing core (i.e., streaming multiprocessor) in GPGPUs supports thousands of simultaneously active threads, and requires a large register file. Such a sizable register file is very sensitive to the hardware variability, and becomes one of the major units in determining the core frequency. In this study, we propose a set of techniques that mitigate both the PV and NBTI impacts on GPGPUs register file. In order to mitigate the susceptibility to PV, we first develop a novel mechanism that classifies registers into fast and slow categories in the highly-banked register architecture to maximize the frequency improvement. We then leverage the unique features in GPGPU applications to effectively tolerate the extra access delay to the slow registers. Moreover, we propose to dynamically balance the utilization across registers to further tolerate the NBTI degradation. Our experimental results show that our proposed techniques optimize GPGPUs performance by 22 percent on average under both PV and NBTI effects.
Jingweijia Tan, Mingsong Chen 0001, Yang Yi 0002, Xin Fu 0001
IEEE Trans. Parallel Distributed Syst.3
2015 Channel estimation in wireless OFDM systems using reservoir computing
abstract
Reservoir Computing (RC) is a recent neurologically inspired concept for processing time dependent data that lends itself particularly well to hardware implementation by using the device physics to conduct information processing. In this paper, we apply RC to channel estimation in Orthogonal Frequency Division Multiplexing (OFDM) systems. Due to the multipath propagation environment between a transmitter and receiver, the received signal undergoes attenuation, time delay and phase shift. For mitigating these random effects and decoding the transmitted signal at the receiver, accurate channel estimation is vital. Statistical approaches for channel estimation assume that accurate channel information is available at the receiver. However, the time-variance of the channel complicates the channel estimation process by making the current estimation outdated. Recurrent Neural Networks (RNNs), which are analogous to the functioning of the human brain, are therefore utilized for channel prediction. Training algorithms for RNNs are categorized as gradient-descent methods, which often results in high computational complexity and leads to non-convergence due to the presence of bifurcations. In this paper, an Echo State Network (ESN), which is a class of RC approach, has been used for training a RNN to estimate the channel state information. Using this approach, the training and hence, the implementation complexity is significantly reduced. Simulation results show significant improvement in channel estimation accuracy for the proposed method.
Wafi Danesh, Chenyuan Zhao, Bryant T. Wysocki, Michael J. Medley, Ngwe Thawdar, Yang Yi 0002
CISDA6
2015 Neuromorphic encoding system design with chaos based CMOS analog neuron
abstract
Neuromorphic computing is a novel paradigm that inspired from the dynamic behavior of the biological brain. The encoding capability plays a vital role in information processing, especially for neural network based systems. In this paper, a compact, low power, and robust spiking-time-dependent encoder is designed with an accommodative Leaky Integrate and Fire (LIF) model based neuron cluster and a chaotic circuit with ring oscillators. Novel and fundamental methodologies, which represent data by using spike timing dependent encoding, has been developed. The information in signal amplitude has been mapped into a spike time sequence efficiently by time encoding, which represents the input data and offers perfect recovery for band limited stimuli. Time dependent temporal scales have been adopted to pattern the neural activities across multiple timescales and encode the sensory information. Furthermore, chaotic circuit based Pseudorandom Time Series Generator (PTSG) is designed to generate sampling clock. High resolution is provided with chaotic based sampling in the proposed encoding circuit. Detailed post layout simulation results and analysis of the designed circuit are presented.
Chenyuan Zhao, Wafi Danesh, Bryant T. Wysocki, Yang Yi 0002
CISDA4
2015 Spike-Time-Dependent Encoding for Neuromorphic Processors
abstract
This article presents our research towards developing novel and fundamental methodologies for data representation using spike-timing-dependent encoding. Time encoding efficiently maps a signal's amplitude information into a spike time sequence that represents the input data and offers perfect recovery for band-limited stimuli. In this article, we pattern the neural activities across multiple timescales and encode the sensory information using time-dependent temporal scales. The spike encoding methodologies for autonomous classification of time-series signatures are explored using near-chaotic reservoir computing. The proposed spiking neuron is compact, low power, and robust. A hardware implementation of these results is expected to produce an agile hardware implementation of time encoding as a signal conditioner for dynamical neural processor designs.
Chenyuan Zhao, Bryant T. Wysocki, Yifang Liu, Clare Thiem, Nathan R. McDonald, Yang Yi 0002
ACM J. Emerg. Technol. Comput. Syst.6
2015 Resource Allocation for Delay-Sensitive Traffic Over LTE-Advanced Relay Networks
abstract
Future wireless networks will face the dual challenge of supporting large traffic volumes while providing reliable service for delay-sensitive traffic. To meet the challenge, relay network has been introduced as a new network architecture for the fourth generation (4G) LTE-Advanced (LTE-A) networks. In this paper, we investigate resource allocation including subcarrier and power allocation for LTE-A relay networks under statistical quality of service (QoS) constraints. By dual decomposition, we derive the optimal subcarrier and power allocation strategies to maximize the effective capacity (EC) of the underlying LTE-A relay systems. Characteristics of optimal resource allocation strategies are identified, and a low-complexity suboptimal scheme is developed through optimizing the subcarrier and power allocation individually. Our result suggests that the optimal subcarrier and power allocation strategies depend heavily on the underlying QoS constraint. For example, in the low signal-to-interference-plus-noise (SINR) regime, when there are less stringent QoS constraints, base stations and relay stations tend to allocate all the power to the best available subcarrier. However, as QoS requirements become more stringent, both base stations and relay stations will spread their power over available subcarriers. On the other hand, in the high SINR regime, regardless of the QoS constraints, base stations and relay stations tend to equally allocate power among available subcarriers.
Lingjia Liu 0001, Hongxiang Li 0001, Jianzhong Zhang 0002, Yang Yi 0002
IEEE Trans. Wirel. Commun.5
2013 Adaptive resource allocation for heterogeneous traffic over heterogeneous relay networks
abstract
Future wireless communication networks will face the dual challenge of supporting large traffic volumes while providing reliable service for heterogeneous traffic types. In order to meet the ever increasing traffic demand, heterogeneous relay network is introduced as an enabling technology for the fourth generation (4G) mobile broadband networks. In this paper, we will investigate optimal power and subcarrier allocation strategies for heterogeneous relay networks under statistical quality of service (QoS) constraints. To be specific, we will characterize the effective capacity of a wireless relay system under QoS constraints. The properties of the optimal resource and subcarrier allocation strategies for delay-sensitive traffic over heterogeneous relay networks will also be identified. Two low complexity resource and subcarrier allocation algorithms will be introduced to optimize the effective capacity. Our results suggest that the optimal power and subcarrier allocation strategies depend heavily on the underlying QoS constraint. In the low signal-to-interference-plus-noise (SINR) regime, when there is no QoS constraint, both base stations and relay nodes will allocate all the transmit power to the best subcarrier. However, as the QoS requirement becomes more stringent, both base stations and relay nodes will spread their transmit power over multiple subcarriers.
Lingjia Liu 0001, Hongxiang Li 0001, Ying Li 0129, Yang Yi 0002
ICC5
2013 Modeling and characterizing GPGPU reliability in the presence of soft errors
Jingweijia Tan, Yang Yi 0002, Fangyang Shen, Xin Fu 0001
Parallel Comput.2
2011 Capacity of Multicarrier Multilayer Broadcast and Unicast Hybrid Cellular System with Independent Channel Coding over Subcarriers
abstract
In this paper, we discuss the hybrid capacity region of a generic multicarrier multilayer broadcast and unicast cellular system with independent channel coding over subcarriers. In particular, we analytically derive the capacity region and provide conditions to achieve its boundary. The simulation results show that the hybrid capacity regions are considerably higher than those of the traditional time division multiplexing scheme.
Siqian Liu, Hongxiang Li 0001, Guanying Ru, Weiyao Lin, Lingjia Liu 0001, Yang Yi 0002
VTC Fall6