Hun-Seok Kim

dblp:77/6653 · DBLP profile ↗
← Back
63ranked-venue papers
5as first author
39since 2021 · last 2026
0000-0002-6658-5502ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 27 · 5 first-author · 15 since 2021Systems, architecture and hardware · 20 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion
abstract
Multi-sensor fusion using LiDAR and RGB cameras significantly enhances 3D object detection task. However, conventional LiDAR sensors perform dense, stateless scans, ignoring the strong temporal continuity in real-world scenes. This leads to substantial sensing redundancy and excessive power consumption, limiting their practicality on resource-constrained platforms. To address this inefficiency, we propose a predictive, history-aware adaptive scanning framework that anticipates informative regions of interest (ROI) based on past observations. Our approach introduces a lightweight predictor network that distills historical spatial and temporal contexts into refined query embeddings. These embeddings guide a differentiable mask generator network, which leverages Gumbel-Softmax sampling to produce binary masks identifying critical ROIs for the upcoming frame. Our method significantly reduces unnecessary data acquisition by concentrating dense LiDAR scanning only within these ROIs and sparsely sampling elsewhere. Experiments on nuScenes and Lyft benchmarks demonstrate that our adaptive scanning strategy reduces LiDAR energy consumption by over 65% while maintaining competitive or even superior 3D object detection performance compared to traditional LiDAR-camera fusion methods with dense LiDAR scanning.
Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim
AAAI3
2026 Accelerating Block Low-Rank Foundation Model Inference on Memory-Constrained GPUs
Pierre Abillama, Changwoo Lee 0001, Juechu Dong, David T. Blaauw, Dennis Sylvester, Hun-Seok Kim
HPDC6
2026 QFEC: A 9.97 Gb/s Fully Configurable Quad-Mode Decoder for LDPC, Polar, Turbo, and Convolutional Codes
abstract
Rapidly evolving wireless channel conditions and communication standards demand adaptable forward error correction (FEC) decoders. Existing rigid architectures, designed for a single standard, exhibit limited adaptability in terms of throughput and/or coding gain, hindering the timely deployment of new applications. To overcome these limitations, we propose QFEC (quad-mode FEC decoder), a unified and highly configurable FEC decoder. QFEC enables a wide range of throughput vs. coding gain tradeoffs across QC-LDPC, Turbo, Polar, and Convolutional codes (CC) by providing full configurability for existing standards and proprietary systems. This ensures communication reliability under varying channel conditions and seamless support for both legacy and emerging communication protocols. Our hardware-unified approach leverages a shared memory and computation unit architecture that exploits the inherent commonalities in the iterative message-passing dataflow of all four code types. We attain outstanding flexibility at high data rates through an innovative combination of a fully customizable interconnect and a multi-mode computation datapath. The QFEC chip, fabricated in GF 12 nm FinFET technology, achieves 9.97 Gb/s throughput for the Optical Communication Terminal (OCT) standard and 6.52 Gb/s for 5G BG1, while consuming normalized energy efficiency (NEE) of 1.04 pJ/b and 1.53 pJ/b, respectively. This design can reach a maximum of 25.4 Gb/s using a proprietary QC-LDPC configuration. This design significantly surpasses existing solutions in flexibility by offering the broadest support for standards and parameters with a unified, efficient architecture. To the best of our knowledge, this is the first chip implementation of a fully flexible quad-mode FEC decoder for QC-LDPC, Polar, Turbo, CC codes.
Yufan Yue, Kuan-Yu Chen 0001, Xiangdong Wei, Tutu Ajayi, Hyunwon Chung, Ronald G. Dreslinski, David T. Blaauw, Hun-Seok Kim
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 End-to-end Neural-Classical Concatenated Feedback Coding with OFDM Guidance
abstract
In this paper, we propose an end-to-end (E2E) optimized concatenated feedback coding scheme that combines a neural inner encoder/decoder with a classical low-density parity-check (LDPC) outer code for reliable feedback-based communication over multipath fading channels. A key contribution of this work is the inclusion of the LDPC decoder within the training loop by exploiting the differentiability of the sum-product algorithm (SPA), allowing gradient-based optimization to flow across the entire feedback-based communication chain. To enhance robustness over multipath fading channels, we integrate orthogonal frequency-division multiplexing (OFDM) guidance as well as the bit-interleaved coded modulation (BICM) paradigm into the proposed concatenated feedback coding approach. The proposed framework supports practical blocklengths and fading scenarios, significantly outperforming both fully neural and non-E2E concatenated baselines.
Adnan Can Büyüksirin, Hun-Seok Kim
GLOBECOM2
2025 Learned Superposition Precoding For PAPR Reduction of OFDMA System
abstract
Orthogonal frequency-division multiple access (OFDMA) is the multiple access scheme used in Long term evolution (LTE), 5G new radio (NR), and the latest Wi-Fi protocols. OFDMA gains popularity because of its low computational complexity requirements for demodulation and one-tap equalization. But OFDMA is known to have a high peak-to-average power ratio (PAPR), which asks for higher dynamic range of amplifiers and reduces their power efficiency. In this paper, we propose a novel PAPR reduction scheme for OFDMA systems using a learned superposition precoder and a neural network (NN) based decoder. We define peak-to-noise ratio (PNR) as a metric to evaluate the tradeoff between PAPR and bit error rate (BER) performance. Our model has significantly lower BER at the same PNR compared with the original OFDM and various PAPR reduction schemes.
Boxuan Chang, Hun-Seok Kim
GLOBECOM2
2025 Secure Ranging for Proximity-Based Authentication Using Physical Unclonable Functions in OFDM
abstract
This paper presents a secure ranging system for proximity-based authentication using orthogonal frequency division multiplexing (OFDM) and Physically Unclonable Function (PUF). It specifically targets applications such as keyless entry systems that require accurate and tamper-resistant distance measurement. The proposed system utilizes PUF-based noiselike signatures embedded in the OFDM modulation process to mitigate distance manipulation attacks from adversaries. We implemented the design on a USRP X310 platform, leveraging both software and hardware-based processing to achieve real-time performance. Experimental evaluations conducted in wireless channels demonstrate that our PUF-based secure ranging system achieves decimeter-level accuracy and successfully resists distance reduction attacks. This approach offers high security and precision, making it a promising solution for secure and robust proximity-based authentication.
Demba Komma, Jiyoon Han, Mingyan Liu, Hun-Seok Kim
GLOBECOM4
2025 OSLA-FIM: Real-World Interference Tolerant Feedback-Based Communications
Winston Wang, Hun-Seok Kim
GLOBECOM2
2025 SAM-Guided Pseudo Label Enhancement for Multi-Modal 3D Semantic Segmentation
abstract
Multi-modal 3D semantic segmentation is vital for applications such as autonomous driving and virtual reality (VR). To effectively deploy these models in real-world scenarios, it is essential to employ cross-domain adaptation techniques that bridge the gap between training data and real-world data. Recently, self-training with pseudo-labels has emerged as a predominant method for cross-domain adaptation in multi-modal 3D semantic segmentation. However, generating reliable pseudo-labels necessitates stringent constraints, which often result in sparse pseudo-labels after pruning. This sparsity can potentially hinder performance improvement during the adaptation process. We propose an image-guided pseudo-label enhancement approach that leverages the complementary 2D prior knowledge from the Segment Anything Model (SAM) to introduce more reliable pseudo-labels, thereby boosting domain adaptation performance. Specifically, given a 3D point cloud and the SAM masks from its paired image data, we collect all 3D points covered by each SAM mask that potentially belong to the same object. Then our method refines the pseudo-labels within each SAM mask in two steps. First, we determine the class label for each mask using majority voting and employ various constraints to filter out unreliable mask labels. Next, we introduce Geometry-Aware Progressive Propagation (GAPP) which propagates the mask label to all 3D points within the SAM mask while avoiding outliers caused by 2D-3D misalignment. Experiments conducted across multiple datasets and domain adaptation scenarios demonstrate that our proposed method significantly increases the quantity of high-quality pseudo-labels and enhances the adaptation performance over baseline methods.
Mingyu Yang 0002, Jitong Lu, Hun-Seok Kim
ICRA3
2025 Demo: Networked iGYM for AR Exergames
abstract
iGYM is an augmented reality exercise game for inclusive play that allows people with and without wheelchairs to participate equally in a soccer game with a projected virtual field and ball. However, currently iGYM requires all players to be co-located and lacks capabilities for remote play. In this work, we describe a networked iGYM implementation that allows teams of players to play with each other remotely. The key networking challenge is meeting tight end-to-end latency requirements for interactive play over the Internet. We demonstrate a portable tabletop version of iGYM, implemented in Unity and ROS2, using a distributed authority model to keep track of ownership and propagate state updates about players, game objects, and scores. Each client receives local player tracking updates and spawns player peripersonal circles under its own authority; a shared ball object is synchronized under server ownership with clientside interpolation. The demo GUI will let attendees inject artificial network delay to explore its impact on interactivity.
Brandon McDonald, Michael Nebeling, Roland Graf, Hun-Seok Kim, Jiasi Chen
MobiCom4
2025 NBLoc: A Narrowband RF Localization System for Wide-Area Indoor Applications
abstract
We introduce NBLoc, a narrowband frequency hopping, long-range localization system designed for low-power Internet of Things (IoT) devices. Traditional high-accuracy localization systems typically require wide-bandwidth and power-demanding radio frequency (RF) circuits, leading to limitations such as short operational range and high power consumption to achieve decimeter-level accuracy. NBLoc overcomes these challenges by using narrowband symbols with a frequency-hopping mechanism across a wide bandwidth, enabling the localization of low-power tags over large areas. NBLoc features a novel custom-designed RF analog frontend (AFE) integrated circuit (IC), that eliminates the need for a conventional phase-locked loop, significantly reducing the cost and power consumption of the receiving tag. This advancement is enabled by NBLoc's thoughtful waveform design and specialized signal processing algorithms, which mitigate phase noise and uncertainty. Compared to previous solutions, NBLoc achieves lower power consumption and extended operational range due to its narrowband symbols while maintaining high localization accuracy by leveraging a 100 MHz localization bandwidth through frequency hopping. In NBLoc, system anchors transmit narrowband orthogonal symbols, hopping across the localization bandwidth in a predetermined pattern known to the tag. The tag, equipped with the custom low-power RF AFE IC, dynamically tunes its local oscillator (LO) frequency to match the hop pattern and capture these symbols, which are then used to estimate the channel impulse response (CIR). The tag calculates the time difference of arrival (TDOA) for each anchor pair from the CIRs, and determines its 2D location via multilateration. The system was implemented and tested using the low-power RF AFE IC in both line-of-sight (LOS) and non-line-of-sight (NLOS) environments, achieving decimeter-level accuracy across areas as large as 269 × 125 m$^{2}$.
Demba Komma, Chien-Wei Tseng, Andrea Bejarano-Carbo, Mingyu Yang 0002, David T. Blaauw, Hun-Seok Kim
IEEE Trans. Mob. Comput.6
2025 Hybrid Timestamping Using Crystal and RC Oscillators for Shock-Resistant Precision
abstract
Achieving precise timing in miniature systems attached to monarch butterflies is challenging due to the shock sensitivity of crystal oscillators (XOs) and the limited accuracy ofRCoscillators. This brief proposes a hybrid timestamping technique that combines both oscillators as timers to deliver shock-resistant, high-accuracy timing. Three algorithms are evaluated to fine-tune a multiplier (M), the ratio of the two timers’ speeds, for improved responsiveness and robustness against temperature and voltage variations. The direct ratioing algorithm proves the most effective, determining the correctMwithin a single wake-up cycle and reducing time shift error by 145 times in a 12-h test, compared to using anRCtimer alone. This work leverages the existing hardware and introduces new firmware, easily implementable using standard digital circuit design flows, to significantly enhance timing precision and shock resilience in millimeter-scale butterfly tracking systems, making a valuable contribution to the VLSI community.
Ehab A. Hamed, Gordy A. Carichner, Delbert A. Green II, Hun-Seok Kim, Inhee Lee 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2024 Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
abstract
This paper investigates efficient deep neural networks (DNNs) to replace dense unstructured weight matrices with structured ones that possess desired properties. The challenge arises because the optimal weight matrix structure in popular neural network models is obscure in most cases and may vary from layer to layer even in the same network. Prior structured matrices proposed for efficient DNNs were mostly hand-crafted without a generalized framework to systematically learn them. To address this issue, we propose a generalized and differentiable framework to learn efficient structures of weight matrices by gradient descent. We first define a new class of structured matrices that covers a wide range of structured matrices in the literature by adjusting the structural parameters. Then, the frequency-domain differentiable parameterization scheme based on the Gaussian-Dirichlet kernel is adopted to learn the structural parameters by proximal gradient descent. On the image and language tasks, our method learns efficient DNNs with structured matrices, achieving lower complexity and/or higher performance than prior approaches that employ low-rank, block-sparse, or block-low-rank matrices.
Changwoo Lee 0001, Hun-Seok Kim
ICLR2
2024 ParaBase: A Configurable Parallel Baseband Processor for Ultra-High-Speed Inter-Satellite Optical Communications
abstract
This paper presents ParaBase, a configurable baseband processing architecture that efficiently handles parallel sample streams targeting ultra-wide bandwidth inter-satellite optical communications for the target data rates exceeding 100 Gbps. We propose a parallelogram-style systolic accelerator, specifically designed for parallel processing for correlation kernels, preserving the hardware efficiency inherent in the systolic-array architecture. ParaBase supports end-to-end baseband processing through a set of heterogeneous configurable accelerators customized to their respective parallel processing requirements. It outperforms the SIMD-style architecture and surpasses previous baseband processors by ~7.0X in terms of energy efficiency for FIR filtering. The overall architecture reports energy efficiency of up to 2.9 TOPS/W and 121.8 Gbits/J while supporting a wide range of data rates from 1~128 Gbps via fast reconfiguration for various modulation schemes.
Seungkyu Choi, Huanshihong Deng, Kuan-Yu Chen 0001, Yufan Yue, David T. Blaauw, Hun-Seok Kim
ISLPED6
2024 BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference
abstract
Large-scale foundation models have demonstrated exceptional performance in language and vision tasks. However, the numerous dense matrix-vector operations involved in these large networks pose significant computational challenges during inference. To address these challenges, we introduce the Block-Level Adaptive STructured (BLAST) matrix, designed to learn and leverage efficient structures prevalent in the weight matrices of linear layers within deep learning models. Compared to existing structured matrices, the BLAST matrix offers substantial flexibility, as it can represent various types of structures that are either learned from data or computed from pre-existing weight matrices. We demonstrate the efficiency of using the BLAST matrix for compressing both language and vision tasks, showing that (i) for medium-sized models such as ViT and GPT-2, training with BLAST weights boosts performance while reducing complexity by 70\% and 40\%, respectively; and (ii) for large foundation models such as Llama-7B and DiT-XL, the BLAST matrix achieves a 2x compression while exhibiting the lowest performance degradation among all tested structured matrices. Our code is available at https://github.com/changwoolee/BLAST.
Changwoo Lee 0001, Soo Min Kwon, Qing Qu 0001, Hun-Seok Kim
NeurIPS4
2024 Canalis: A Throughput-Optimized Framework for Real-Time Stream Processing of Wireless Communication
abstract
Stream processing, which involves real-time computation of data as it is created or received, is vital for various applications, specifically wireless communication. The evolving protocols, the requirement for high-throughput, and the challenges of handling diverse processing patterns make it demanding. Traditional platforms grapple with meeting real-time throughput and latency requirements due to large data volume, sequential and indeterministic data arrival, and variable data rates, leading to inefficiencies in memory access and parallel processing. We present Canalis, a throughput-optimized framework designed to address these challenges, ensuring high-performance while achieving low energy consumption. Canalis is a hardware-software co-designed system. It includes a programmable spatial architecture, Flux Stream Processing Unit (FluxSPU), proposed by this work to enhance data throughput and energy efficiency. FluxSPU is accompanied by a software stack that eases the programming process. We evaluated Canalis with eight distinct benchmarks. When compared to CPU and GPU in mobile SoC to demonstrate the effectiveness of domain specialization, Canalis achieves an average speedup of 13.4 \(\times\) and 6.6 \(\times\) , and energy savings of 189.8 \(\times\) and 283.9 \(\times\) , respectively. In contrast to equivalent ASICs of the benchmarks, the average energy overhead of Canalis is within 2.4 \(\times\) , successfully maintaining generalizations without incurring significant overhead.
Kuan-Yu Chen 0001, Thomas Mason Nelson, Alireza Khadem, Morteza Fayazi, Sanjay Sri Vallabh Singapuram, Ronald G. Dreslinski, Nishil Talati, Hun-Seok Kim, David T. Blaauw
ACM Trans. Reconfigurable Technol. Syst.8
2023 Deep Joint Source-Channel Coding with Iterative Source Error Correction
abstract
In this paper, we propose an iterative source error correction (ISEC) decoding scheme for deep-learning-based joint source-channel coding (Deep JSCC). Given a noisy codeword received through the channel, we use a Deep JSCC encoder and decoder pair to update the codeword iteratively to find a (modified) maximum a-posteriori (MAP) solution. For efficient MAP decoding, we utilize a neural network-based denoiser to approximate the gradient of the log-prior density of the codeword space. Albeit the non-convexity of the optimization problem, our proposed scheme improves various distortion and perceptual quality metrics from the conventional one-shot (non-iterative) Deep JSCC decoding baseline. Furthermore, the proposed scheme produces more reliable source reconstruction results compared to the baseline when the channel noise characteristics do not match the ones used during training.
Changwoo Lee 0001, Hun-Seok Kim
AISTATS3
2023 SONA: An Accelerator for Transform-Domain Neural Networks with Sparse-Orthogonal Weights
abstract
Recent advances in model pruning have enabled sparsity-aware deep neural network accelerators that improve the energy-efficiency and performance of inference tasks. We introduce SONA, a novel transform-domain neural network accelerator in which convolution operations are replaced by element-wise multiplications with sparse-orthogonal weights. SONA employs an output stationary dataflow coupled with an energy-efficient memory organization to reduce the overhead of sparse-orthogonal transform-domain kernels that are concurrently processed without any conflicts. Weights in SONA are non-uniformly quantized with bit-sparse canonical-signed-digit representations to reduce multiplications to simple additions. Moreover, for sparse fully-connected layers (FCLs), SONA introduces column-based-block structured pruning, which is integrated into the same architecture that maintains full multiply-and-accumulate (MAC) array utilization. Compared to prior dense and sparse neural networks accelerators, SONA can reduce inference energy by$5.1\times$and$2.4 \times$and increase performance by$5.2\times$and$2.1\times$, respectively, for convolution layers. For sparse FCLs, SONA can reduce inference energy by$2.4\times$and increase performance by$2\times$compared to prior work.
Pierre Abillama, Zichen Fan, Yu Chen 0070, Hyochan An, Qirui Zhang 0001, Seungkyu Choi, David T. Blaauw, Dennis Sylvester, Hun-Seok Kim
ASAP9
2023 MMVC: Learned Multi-Mode Video Compression with Block-based Prediction Mode Selection and Density-Adaptive Entropy Coding
abstract
Learning-based video compression has been extensively studied over the past years, but it still has limitations in adapting to various motion patterns and entropy models. In this paper, we propose multi-mode video compression (MMVC), a block wise mode ensemble deep video compression framework that selects the optimal mode for feature domain prediction adapting to different motion patterns. Proposed multi-modes include ConvLSTM-based feature domain prediction, optical flow conditioned feature domain prediction, and feature propagation to address a wide range of cases from static scenes without apparent motions to dynamic scenes with a moving camera. We partition the feature space into blocks for temporal prediction in spatial block-based representations. For entropy coding, we consider both dense and sparse post-quantization residual blocks, and apply optional run-length coding to sparse residuals to improve the compression rate. In this sense, our method uses a dual-mode entropy coding scheme guided by a binary density map, which offers significant rate reduction surpassing the extra cost of transmitting the binary selection map. We validate our scheme with some of the most popular benchmarking datasets. Compared with state-of-the-art video compression schemes and standard codecs, our method yields better or competitive results measured with PSNR and MS-SSIM.
Bowen Liu 0001, Yu Chen 0070, Rakesh Chowdary Machineni, Hun-Seok Kim
CVPR5
2023 Search for Efficient Deep Visual-Inertial Odometry Through Neural Architecture Search
abstract
Recent deep learning based visual-inertial odometry (VIO) systems achieve impressive performance in various applications and challenging scenarios. However, it is difficult to deploy such VIO models directly on energy-constrained mobile platforms in real-time due to the extensive complexity of existing deep neural network (DNN) models. To address this issue, we propose to adopt the neural architecture search (NAS) technique to search for the most efficient VIO network architecture. Targeting the lowest number of operations and inference latency, our searched models achieve up to 97.4% complexity reduction with no performance degradation. The searched efficient visual encoder allows our VIO model to run at 83.3 frames per second on a single laptop CPU core. Moreover, the model complexity can be reduced by 99.1% when combined with a dynamic modality selection technique. Our searched efficient VIO models are available at https://github.com/unchenyu/NASVIO.
Yu Chen 0070, Mingyu Yang 0002, Hun-Seok Kim
ICASSP3
2023 Deep Learning-Based Joint Channel Coding and Frequency Modulation for Low Power Connectivity
abstract
Low-power, low-cost wireless communication is a fundamental requirement of Internet-of-Things (IoT) and massive machine-type communication (mMTC). Various low power connectivity standards such as Bluetooth and LoRa adopt non-coherent frequency modulation schemes as they exhibit significantly lower complexity and power consumption compared to coherent in-phase and quadrature (IQ) modulation schemes. In our paper, we propose a deep learning-based joint channel coding and modulation (JCM) scheme for digitally controlled oscillator (DCO)-based frequency modulation. The learned encoder takes an information bit sequence and produces DCO control samples that represent instantaneous frequency to modulate the radio frequency (RF) signal. The learned decoder recovers/decodes information bits from the received noisy samples without any preamble to assist time and frequency synchronization. We train and test the proposed scheme under significant phase noise and carrier frequency offset (CFO) of the DCO to successfully mitigate these practical impairments at the receiver.
Boxuan Chang, Hun-Seok Kim
ICC3
2023 Efficient Computation Sharing for Multi-Task Visual Scene Understanding
abstract
Solving multiple visual tasks using individual models can be resource-intensive, while multi-task learning can conserve resources by sharing knowledge across different tasks. Despite the benefits of multi-task learning, such techniques can struggle with balancing the loss for each task, leading to potential performance degradation. We present a novel computation- and parameter-sharing framework that balances efficiency and accuracy to perform multiple visual tasks utilizing individually-trained single-task transformers. Our method is motivated by transfer learning schemes to reduce computational and parameter storage costs while maintaining the desired performance. Our approach involves splitting the tasks into a base task and the other sub-tasks, and sharing a significant portion of activations and parameters/weights between the base and sub-tasks to decrease inter-task redundancies and enhance knowledge sharing. The evaluation conducted on NYUD-v2 and PASCAL-context datasets shows that our method is superior to the state-of-the-art transformer-based multi-task learning techniques with higher accuracy and reduced computational resources. Moreover, our method is extended to video stream inputs, further reducing computational costs by efficiently sharing information across the temporal domain as well as the task domain. Our codes are available at https://github.com/sarashoouri/EfficientMTL.
Sara Shoouri, Mingyu Yang 0002, Zichen Fan, Hun-Seok Kim
ICCV4
2023 TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language Processing
abstract
The combination of pre-trained models and task-specific fine-tuning schemes, such as BERT, has achieved great success in various natural language processing (NLP) tasks. However, the large memory and computation costs of such models make it challenging to deploy them in edge devices. Moreover, in real-world applications like chatbots, multiple NLP tasks need to be processed together to achieve higher response credibility. Running multiple NLP tasks with specialized models for each task increases the latency and memory cost latency linearly with the number of tasks. Though there have been recent works on parameter-shared tuning that aim to reduce the total parameter size by partially sharing weights among multiple tasks, computation remains intensive and redundant despite different tasks using the same input. In this work, we identify that a significant portion of activations and weights can be reused among different tasks, to reduce cost and latency for efficient multi-task NLP. Specifically, we propose TaskFusion, an efficient transfer learning software-hardware co-design that exploits delta sparsity in both weights and activations to boost data sharing among tasks. For training, TaskFusion uses ℓ1 regularization on delta activation to learn inter-task data redundancies. A novel hardware-aware sub-task inference algorithm is proposed to exploit the dual delta sparsity. We then designed a dedicated heterogeneous architecture to accelerate multi-task inference with an optimized scheduling to increase hardware utilization and reduce off-chip memory access. Extensive experiments demonstrate that TaskFusion can reduce the number of floating point operations (FLOPs) by over 73% in multi-task NLP with negligible accuracy loss, while adding a new task at the cost of only < 2% parameter size increase. With the proposed architecture and optimized scheduling, Task-Fusion can achieve 1.48--2.43× performance and 1.62--3.77× energy efficiency than those using state-of-the-art single-task accelerators for multi-task NLP applications.
Zichen Fan, Qirui Zhang 0001, Pierre Abillama, Sara Shoouri, Changwoo Lee 0001, David T. Blaauw, Hun-Seok Kim, Dennis Sylvester
ISCA7
2023 Global Localization of Energy-Constrained Miniature RF Emitters using Low Earth Orbit Satellites
abstract
Daily tracking of small objects or animals anywhere on earth for long time-periods is a long sought-after goal. Recently, the emergence of low earth orbit (LEO) satellites offers a unique pathway to achieve this goal. However, to date, LEO trackers have not achieved cm-size. While the integrated chip can be readily scaled to sub-cm size, the size of trackers remains limited by their battery and antenna size. To address these two fundamental size limiting factors, this paper presents a LEO satellite localization system that is specifically optimized to reduce antenna size and transmit power, thereby reducing battery size. To reduce power, a new cooperative waveform is designed which enhances the localization accuracy, combined with an increased packet length to enable low transmit power while maintaining packet energy. However, this long packet length introduces a intra-packet Doppler shift which we address by proposing a localization algorithm that includes a Doppler shift correction. The final result is a 50 kHz periodic BPSK signal with 23 dBm equivalent isotropic radiation power (EIRP), and 120 ms packet length (> 10 × longer than conventional), at a 60 s interval. The proposed solution enables 7 months operation on a 2.5 × 1.2 cm LiPo battery within a North American search area. To address the antenna size, the optimal transmit frequency was studied and a 1 cm loop antenna with 65% radiation efficiency was designed with internal matching to 50 Ohm. Using the proposed techniques, three satellite flyover experiments were performed to confirm the accuracy of the proposed tracking system and localization algorithms using a USRP-X310, a custom 1 cm-size antenna, and a commercial satellite cluster. The measured average localization error is 320 - 840 m depending on satellite trajectories, demonstrating an improved accuracy in real life measurements compared to prior art with experimental result while simultaneously achieving 15 -- 26 dB lower transmit power and > 3 × lower packet energy.
Demba Komma, Jaechan Lim, Zichen Fan, Chien-Wei Tseng, Hun-Seok Kim, David T. Blaauw
SenSys8
2023 Learning-Based Near-Orthogonal Superposition Code for MIMO Short Message Transmission
abstract
Massive machine type communication (mMTC) has attracted new coding schemes optimized for reliable short message transmission. In this paper, a novel deep learning-based near-orthogonal superposition (NOS) coding scheme is proposed to transmit short messages in multiple-input multiple-output (MIMO) channels for mMTC applications. In the proposed MIMO-NOS scheme, a neural network-based encoder is optimized via end-to-end learning with a corresponding neural network-based detector/decoder in a superposition-based auto-encoder framework including a MIMO channel. The proposed MIMO-NOS encoder spreads the information bits to multiple near-orthogonal high dimensional vectors to be combined (superimposed) into a single vector and reshaped for the space-time transmission. For the receiver, we propose a novel looped$K$-best tree-search algorithm with cyclic redundancy check (CRC) assistance to enhance the error correcting ability in the block-fading MIMO channel. For a comprehensive understanding of the proposed MIMO-NOS scheme, we further quantify the gain from individual components/modules in the framework, and analyze the decoding complexity measured by the floating point operations (FLOPs). Simulation results show the proposed MIMO-NOS scheme outperforms maximum likelihood (ML) MIMO detection combined with a polar code with CRC-assisted list decoding by 1 – 2 dB in various MIMO systems for short (32 – 64 bit) message transmission.
Chenghong Bian, Chin-Wei Hsu, Changwoo Lee 0001, Hun-Seok Kim
IEEE Trans. Commun.4
2023 Instantaneous Feedback-Based Opportunistic Symbol Length Adaptation for Reliable Communication
abstract
Although feedback cannot increase the channel capacity of memoryless channels, it can enhance the error rate performance and/or shorten the codeword length for the target performance. This work is based on an early work by Viterbi in 1965 that utilizes instantaneous feedback for reliable uncoded communications. We build on this work by incorporating convolutional codes as a new variable-symbol-length digital communication scheme using instantaneous feedback. In the proposed system, called Opportunistic Symbol Length Adaptation (OSLA), the symbol length opportunistically adapts to the noise realization observed within a sub-symbol interval to minimize the packet/codeword error rate. It is shown that the proposed OSLA scheme combined with tail-biting convolutional codes or turbo codes outperforms state-of-the-art non-feedback codes as well as a deep learning-based feedback scheme with up to 1.5 dB gain in noiseless and noisy feedback channels.
Chin-Wei Hsu, Achilleas Anastasopoulos, Hun-Seok Kim
IEEE Trans. Commun.3
2023 Hyper-Dimensional Modulation for Robust Short Packets in Massive Machine-Type Communications
abstract
In this paper, we introduce Hyper-Dimensional Modulation (HDM) for massive machine-type communications (mMTC). HDM enables robust communication of short packets by spreading information bits across many elements in a hyper-dimensional vector and superimposing a set of such non-orthogonal vectors. The proposed CRC-aided K-best decoding algorithm for HDM can achieve a very low packet error rate (PER) in additive white Gaussian noise (AWGN) channels for short packets. Furthermore, extended decoding algorithms are proposed to combat overwhelming interference in an mMTC network. Comprehensive simulation and real-world experiment results show that HDM outperforms sparse superposition codes in AWGN channels and state-of-the-art short codes such as polar and tail-biting convolutional codes in interference-heavy channels for short packet transmissions.
Chin-Wei Hsu, Hun-Seok Kim
IEEE Trans. Commun.2
2022 Squaring the circle: Executing Sparse Matrix Computations on FlexTPU - A TPU-Like Processor
abstract
Systolic arrays have been successful to accelerate dense linear algebra for deep neural networks (DNNs), but cannot handle sparse computations efficiently. Though early attempts have been made to perform sparse matrix operations on weight-pruned DNNs, handling highly sparse matrices with skewed nonzero distribution commonly seen in real-world graph analytics remains challenging. In this paper, we propose FlexTPU framework to repurpose tensor processing units (TPUs) to execute sparse matrix-vector operations (SpMV). First, we propose a lightweight Z-shape mapping of sparse matrices onto the systolic array to eliminate the processing of zeros as much as possible, regardless of the sparsity and nonzero distribution. On top of the mapping, we devise an SpMV dataflow executed by an array of PEs, which are a slightly modified version of the conventional TPU PE. Second, in contrast to the excess preprocessing mandatory for prior attempts, the Z-shape mapping facilitates on-the-fly matrix condensing from the widely-used compressed sparse matrix (e.g. CSR) representation. This is accomplished by a proposed sparse data loader that includes an on-chip row decoder and parallel nonzero loaders. We evaluate FlexTPU on a broad set of synthetic and real-world sparse matrices. The experimental result shows that FlexTPU achieves 3.55× speedup and 3.27× energy saving over a state-of-the-art design, Sparse-TPU. It performs even better on sparse matrices with power-law distributions. Compared to state-of-the-art library implementations on a CPU and a GPU, FlexTPU also achieves an average speedup of 2.4× and 4.3×, and energy saving of 130.4× and 495.3×, respectively. FlexTPU is also evaluated against a recent re configurable (chip multi-processor) CMP machine, Transmuter. FlexTPU outperforms Transmuter by achieving 5.12× speedup and 2.65× energy saving.
Xin He 0011, Kuan-Yu Chen 0001, Siying Feng, Hun-Seok Kim, David T. Blaauw, Ronald G. Dreslinski, Trevor N. Mudge
PACT4
2022 Efficient Deep Visual and Inertial Odometry with Adaptive Visual Modality Selection
Mingyu Yang 0002, Yu Chen 0070, Hun-Seok Kim
ECCV (38)3
2022 An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization
abstract
Sound source separation and localization for situational awareness enables a wide range of applications such as hearing enhancement and audio beam-forming. We present an end-to-end deep learning framework to separate and localize multiple audio sources from the mixture of multi-channels. The proposed framework jointly estimates the separated sources and their time difference of arrival (TDOA) at different microphones, then it obtains the direction-of-arrival (DOA) for each source. A new structure to reconstruct the mixed signal is introduced for joint optimization of source separation and TDOA estimation. In addition, a discriminator network is added during the training phase to further improve the separation quality. Experiment results demonstrate that the proposed method achieves state-of-the-art accuracy on source separation as well as DOA estimation.
Yu Chen 0070, Bowen Liu 0001, Zijian Zhang 0011, Hun-Seok Kim
ICASSP4
2022 Deep Joint Source-Channel Coding for Wireless Image Transmission with Adaptive Rate Control
abstract
We present a novel adaptive deep joint source-channel coding (JSCC) scheme for wireless image transmission. The proposed scheme supports multiple rates using a single deep neural network (DNN) model and learns to dynamically control the rate based on the channel condition and image contents. Specifically, a policy network is introduced to exploit the tradeoff space between the rate and signal quality. To train the policy network, the Gumbel-Softmax trick is adopted to make the policy network differentiable and hence the whole JSCC scheme can be trained end-to-end. To the best of our knowledge, this is the first deep JSCC scheme that can automatically adjust its rate using a single network model. Experiments show that our scheme successfully learns a reasonable policy that decreases channel bandwidth utilization for high SNR scenarios or simple image contents. For an arbitrary target rate, our rate-adaptive scheme using a single model achieves similar performance compared to an optimized model specifically trained for that fixed target rate. To reproduce our results, we make the source code publicly available at https://github.com/mingyuyng/Dynamic_JSCC.
Mingyu Yang 0002, Hun-Seok Kim
ICASSP2
2022 Deep Learning Based Near-Orthogonal Superposition Code for Short Message Transmission
abstract
Massive machine type communication (mMTC) has attracted new coding schemes optimized for reliable short message transmission. In this paper, a novel deep learning based near-orthogonal superposition (NOS) coding scheme is proposed for reliable transmission of short messages in the additive white Gaussian noise (AWGN) channel for mMTC applications. Similar to recent hyper-dimensional modulation (HDM), the NOS encoder spreads the information bits to multiple near-orthogonal high dimensional vectors to be combined (superimposed) into a single vector for transmission. The NOS decoder first estimates the information vectors and then performs a cyclic redundancy check (CRC)-assisted K-best tree-search algorithm to further reduce the packet error rate. The proposed NOS encoder and decoder are deep neural networks (DNNs) jointly trained as an auto-encoder and decoder pair to learn a new NOS coding scheme with near-orthogonal codewords. Simulation results show the proposed deep learning-based NOS scheme outperforms HDM and Polar code with CRC-aided list decoding for short (32-bit) message transmission.
Chenghong Bian, Mingyu Yang 0002, Chin-Wei Hsu, Hun-Seok Kim
ICC4
2022 Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous Processor
abstract
RF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design.
Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue
ISCAS25
2022 Improving Energy Efficiency of Convolutional Neural Networks on Multi-core Architectures through Run-time Reconfiguration
abstract
Convolutional neural networks (CNNs) are built with convolution layers which account for most of their computation time. The differences in the convolution kernel types (2D, point-wise, depth-wise), and input sizes lead to significant differences in their computation and memory demands. In this work, we exploit run-time reconfiguration to adapt to the differences in the characteristics of different convolution kernels on a low-power reconfigurable architecture, Transmuter. The architecture consists of light-weight cores interconnected by caches and crossbars that support run-time reconfiguration between different cache modes - shared or private, different dataflow modes - systolic or parallel, and different computation mapping schemes. To achieve run-time reconfiguration, we propose a decision-tree-based engine that selects the optimal Transmuter configuration at a low cost. The proposed method is evaluated on commonly-used CNN models such as ResNetl8, VGGII, AlexNet and MobileNetV3. Simulation results show that run-time reconfiguration helps improve the energy efficiency of Transmuter in the range of 3.1$\times-13.7\times$ across all networks.
Yan Xiong 0002, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski, Chaitali Chakrabarti
ISCAS4
2022 A Unified Forward Error Correction Accelerator for Multi-Mode Turbo, LDPC, and Polar Decoding
abstract
Forward error correction (FEC) is a critical component in communication systems as the errors induced by noisy channels can be corrected using the redundancy in the coded message. This paper introduces a novel multi-mode FEC decoder accelerator that can decode Turbo, LDPC, and Polar codes using a unified architecture. The proposed design explores the similarities in these codes to enable energy efficient decoding with minimal overhead in the total area of the unified architecture. Moreover, the proposed design is highly reconfigurable to support various existing and future FEC standards including 3GPP LTE/5G, and IEEE 802.11n WiFi. Implemented in GF 12nm FinFET technology, the design occupies 8.47mm2 of chip area attaining 25% logic and 49% memory area savings compared to a collection of single-mode designs. Running at 250MHz and 0.8V, the decoder achieves per-iteration throughput and energy efficiency of 690Mb/s and 44pJ/b for Turbo; 740Mb/s and 27.4pJ/b for LDPC; and 950Mb/s and 45.8pJ/b for Polar.
Yufan Yue, Tutu Ajayi, Xueyang Liu, Peiwen Xing, David T. Blaauw, Ronald G. Dreslinski, Hun-Seok Kim
ISLPED8
2021 Deep Learning in Latent Space for Video Prediction and Compression
abstract
Learning-based video compression has achieved substantial progress during recent years. The most influential approaches adopt deep neural networks (DNNs) to remove spatial and temporal redundancies by finding the appropriate lower-dimensional representations of frames in the video. We propose a novel DNN based framework that predicts and compresses video sequences in the latent vector space. The proposed method first learns the efficient lower-dimensional latent space representation of each video frame and then performs inter-frame prediction in that latent domain. The proposed latent domain compression of individual frames is obtained by a deep autoencoder trained with a generative adversarial network (GAN). To exploit the temporal correlation within the video frame sequence, we employ a convolutional long short-term memory (ConvLSTM) network to predict the latent vector representation of the future frame. We demonstrate our method with two applications; video compression and abnormal event detection that share the identical latent frame prediction network. The proposed method exhibits superior or competitive performance compared to the state-of-the-art algorithms specifically designed for either video compression or anomaly detection.1
Bowen Liu 0001, Yu Chen 0070, Hun-Seok Kim
CVPR4
2021 Instantaneous Feedback-based Opportunistic Symbol Length Adaptation for Reliable Communication
abstract
It is well known that although feedback cannot increase the channel capacity of memoryless channels, it can enhance reliability or shorten codeword length. This work is based on an early result by Viterbi in 1965 that utilizes instantaneous feedback for reliable communications. We build on this work by incorporating (tail-biting) convolutional codes and designing a system where the decoder interacts with the transmitter by sending feedback during the decoding process. The proposed system is called Opportunistic Symbol Length Adaptation (OSLA), in which the symbol length opportunistically adapts to noise realization of each symbol to ensure that the target reliability is achieved. It is shown that, combined with tail-biting convolutional codes, the proposed scheme outperforms state-of-the-art non-feedback codes, as well as a recently proposed deep learning-based feedback scheme with up to 1.5 dB gain in noise-less and noisy feedback channels.
Chin-Wei Hsu, Achilleas Anastasopoulos, Hun-Seok Kim
GLOBECOM3
2021 Deep Joint Source Channel Coding for Wireless Image Transmission with OFDM
abstract
We present a deep learning based joint source channel coding (JSCC) scheme for wireless image transmission over multipath fading channels with non-linear signal clipping. The proposed encoder and decoder use convolutional neural networks (CNN) and directly map the source images to complex-valued baseband samples for orthogonal frequency division multiplexing (OFDM) transmission. The proposed model-driven machine learning approach eliminates the need for separate source and channel coding while integrating an OFDM datapath to cope with multipath fading channels. The end-to-end JSCC communication system combines trainable CNN layers with non-trainable but differentiable layers representing the multipath channel model and OFDM signal processing blocks. Our results show that injecting domain expert knowledge by incorporating OFDM baseband processing blocks into the machine learning framework significantly enhances the overall performance compared to an unstructured CNN. Our method outperforms conventional schemes that employ state-of-the-art but separate source and channel coding such as BPG and LDPC with OFDM. Moreover, our method is shown to be robust against non-linear signal clipping in OFDM for various channel conditions that do not match the model parameter used during the training.
Mingyu Yang 0002, Chenghong Bian, Hun-Seok Kim
ICC3
2021 HTNN: Deep Learning in Heterogeneous Transform Domains with Sparse-Orthogonal Weights
abstract
Convolutional neural networks (CNNs) achieved great success on various tasks in recent years. Their applications to low power and low cost hardware platforms, however, have been often limited due to extensive complexity of convolution layers. We present a new class of transform domain deep neural networks (DNNs), where convolution operations are replaced by element-wise multiplications in heterogeneous transform domains. To further reduce the network complexity, we propose a framework to learn sparse-orthogonal weights in heterogeneous transform domains co-optimized with a hardware-efficient accelerator architecture to minimize the overhead of handling sparse weights. Furthermore, sparse-orthogonal weights are non-uniformly quantized with canonical-signed-digit (CSD) representations to substitute multiplications with simpler additions. The proposed approach reduces the complexity by a factor of 4.9– 6.8 $\times$ without compromising the DNN accuracy compared to equivalent CNNs that employ sparse (pruned) weights. The code is available at https://github.com/unchenyu/HTNN.
Yu Chen 0070, Bowen Liu 0001, Pierre Abillama, Hun-Seok Kim
ISLPED4
2021 mSAIL: milligram-scale multi-modal sensor platform for monarch butterfly migration tracking
abstract
Each fall, millions of monarch butterflies across the northern US and Canada migrate up to 4,000 km to overwinter in the exact same cluster of mountain peaks in central Mexico. To track monarchs precisely and study their navigation, a monarch tracker must obtain daily localization of the butterfly as it progresses on its 3-month journey. And, the tracker must perform this task while having a weight in the tens of milligram (mg) and measuring a few millimeters (mm) in size to avoid interfering with monarch's flight. This paper proposes mSAIL, 8 × 8 × 2.6 mm and 62 mg embedded system for monarch migration tracking, constructed using 8 prior custom-designed ICs providing solar energy harvesting, an ultra-low power processor, light/temperature sensors, power management, and a wireless transceiver, all integrated and 3D stacked on a micro PCB with an 8 × 8 mm printed antenna. The proposed system is designed to record and compress light and temperature data during the migration path while harvesting solar energy for energy autonomy, and wirelessly transmit the data at the overwintering site in Mexico, from which the daily location of the butterfly can be estimated using a deep learning-based localization algorithm. A 2-day trial experiment of mSAIL attached on a live butterfly in an outdoor botanical garden demonstrates the feasibility of individual butterfly localization and tracking.
Inhee Lee 0001, Roger Hsiao, Gordy A. Carichner, Chin-Wei Hsu, Mingyu Yang 0002, Sara Shoouri, Katherine Ernst, Tess Carichner, Yuyang Li 0001, Jaechan Lim, Cole R. Julick, Eunseong Moon, Jamie Phillips, Kristi L. Montooth, Delbert A. Green II, Hun-Seok Kim, David T. Blaauw
MobiCom17
2020 Transmuter: Bridging the Efficiency Gap using Memory and Dataflow Reconfiguration
abstract
With the end of Dennard scaling and Moore's law, it is becoming increasingly difficult to build hardware for emerging applications that meet power and performance targets, while remaining flexible and programmable for end users. This is particularly true for domains that have frequently changing algorithms and applications involving mixed sparse/dense data structures, such as those in machine learning and graph analytics. To overcome this, we present a flexible accelerator called Transmuter, in a novel effort to bridge the gap between General-Purpose Processors (GPPs) and Application-Specific Integrated Circuits (ASICs). Transmuter adapts to changing kernel characteristics, such as data reuse and control divergence, through the ability to reconfigure the on-chip memory type, resource sharing and dataflow at run-time within a short latency. This is facilitated by a fabric of light-weight cores connected to a network of reconfigurable caches and crossbars. Transmuter addresses a rapidly growing set of algorithms exhibiting dynamic data movement patterns, irregularity, and sparsity, while delivering GPU-like efficiencies for traditional dense applications. Finally, in order to support programmability and ease-of-adoption, we prototype a software stack composed of low-level runtime routines, and a high-level language library called TransPy, that cater to expert programmers and end-users, respectively.
Subhankar Pal, Siying Feng, Dong-Hyeon Park, Aporva Amarnath, Chi-Sheng Yang, Xin He 0011, Jonathan Beaumont, Kyle May, Yan Xiong 0002, Kuba Kaszyk, John Magnus Morton, Jiawen Sun, Michael F. P. O'Boyle, Murray Cole, Chaitali Chakrabarti, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski
PACT18
2020 Non-Orthogonal Modulation for Short Packets in Massive Machine Type Communications
abstract
Massive Machine Type Communication (mMTC) enables novel applications but its dense deployment and short packet properties lead to new challenges for physical layer design. This paper investigates hyper-dimensional modulation (HDM), a recently proposed novel non-orthogonal modulation, for short packet communications with superior interference tolerance in mMTC. We propose a new tree-based K-best decoding algorithm for HDM to improve the packet error rate performance in both additive white Gaussian noise (AWGN) and interference-limited scenarios. Simulation results show that the proposed algorithm can achieve 0.5 - 4 dB gain in AWGN and interference-limited channels compared to the Polar code with CRC (cyclic redundancy check)-aided list decoding.
Chin-Wei Hsu, Hun-Seok Kim
GLOBECOM2
2020 Unified Signal Compression Using Generative Adversarial Networks
abstract
We propose a unified compression framework that uses generative adversarial networks (GAN) to compress image and speech signals. The compressed signal is represented by a latent vector fed into a generator network which is trained to produce high quality signals that minimize a target objective function. To efficiently quantize the compressed signal, non-uniformly quantized optimal latent vectors are identified by iterative back-propagation with ADMM optimization performed for each iteration. Our experiments show that the proposed algorithm outperforms prior signal compression methods for both image and speech compression quantified in various metrics including bit rate, PSNR, and neural network based signal classification accuracy.
Bowen Liu 0001, Ang Cao, Hun-Seok Kim
ICASSP3
2020 Accelerating Deep Neural Network Computation on a Low Power Reconfigurable Architecture
abstract
Recent work on neural network architectures has focused on bridging the gap between performance/efficiency and programmability. We consider implementations of three popular neural networks, ResNet, AlexNet and ASGD weight-dropped Recurrent Neural Network (AWD RNN) on a low power programmable architecture, Transformer. The architecture consists of light-weight cores interconnected by caches and crossbars that support run-time reconfiguration between shared and private cache mode operations. We present efficient implementations of key neural network kernels and evaluate the performance of each kernel when operating in different cache modes. The best-performing cache modes are then used in the implementation of the end-to-end network. Simulation results show superior performance with ResNet, AlexNet and AWD RNN achieving 188.19 GOPS/W, 150.53 GOPS/W and 120.68 GOPS/W, respectively, in the 14 nm technology node.
Yan Xiong 0002, Jian Zhou 0012, Subhankar Pal, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski, Chaitali Chakrabarti
ISCAS5
2020 Interactive-Multiple-Model Algorithm Based on Minimax Particle Filtering
abstract
In this letter, we propose a new approach to tracking a target that maneuvers based on the multiple-constant-turns model. Usually, the interactive-multiple-model (IMM) algorithm based on the extended Kalman filter (IMM-EKF) is employed for this problem with successful tracking performance. Recently proposed IMM-particle filtering (IMM-PF) showed outperforming results over IMM-EKF for this nonlinear problem. The proposed approach in this letter is a new framework of PF that adopts the minimax strategy to IMM-PF. The minimax strategy results in the decreased variance of the weights of particles that provides the robustness against the degeneracy phenomenon (a common problem of generic PF). In this letter, we show outperforming results by IMM-minimax-PF over IMM-PF besides the IMM-EKF in terms of estimation accuracy and computational complexity.
Jaechan Lim, Hun-Seok Kim, Hyung-Min Park
IEEE Signal Process. Lett.2
2019 IoT2 - the Internet of Tiny Things: Realizing mm-Scale Sensors through 3D Die Stacking
abstract
The Internet of Things (IoT) is a rapidly evolving application space. One of the fascinating new fields in IoT research is mm-scale sensors, which make up the Internet of Tiny Things (IoT2). With their miniature size, these systems are poised to open up a myriad of new application domains. Enabled by the unique characteristics of cyber-physical systems and recent advances in low-power design and bare-die 3D chip stacking, mm-scale sensors are rapidly becoming a reality. In this paper, we will survey the challenges and solutions to 3D-stacked mm-scale design, highlighting low-power circuit issues ranging from low-power SRAM and miniature neural network accelerators to radio communication protocols and analog interfaces. We will discuss system-level challenges and illustrate several complete systems and their merging application spaces.
Sechang Oh 0001, Minchang Cho, Xiao Wu 0002, Yejoong Kim, Li-Xuan Chuo, Wootaek Lim, Pat Pannuto, Suyoung Bang, Kaiyuan Yang 0001, Hun-Seok Kim, Dennis Sylvester, David T. Blaauw
DATE10
2019 Collision-Tolerant Narrowband Communication Using Non-Orthogonal Modulation and Multiple Access
abstract
Ultra Narrowband (UNB) has recently received great attention for its potential to realize ultra- reliable, massive scale Low Power Wide Area Networks (LPWAN). Elaborate frequency planning and multiple access schemes have been regarded as an essential part of LPWAN because random frequency- and time-domain ALOHA accesses lead to significant network performance degradation due to inevitable packet collisions. In this paper, we propose a novel network scheme based on non-orthogonal modulation and multiple access (NOMMA) that is tolerant to packet collisions. The proposed scheme uses hyper-dimensional modulation (HDM) to outperform conventional orthogonal modulation and multiple access schemes with and without prior knowledge of interference in highly congested network scenarios. Simulation results show that HDM based NOMMA can achieve 70% higher network throughput than a conventional orthogonal modulation and multiple access scheme.
Chin-Wei Hsu, Hun-Seok Kim
GLOBECOM2
2019 iLPS: Local Positioning System with Simultaneous Localization and Wireless Communication
abstract
This paper presents a novel RF local positioning system, iLPS, specifically designed for challenging indoor non-lineof-sight (NLOS) scenarios and/or urban canyons where global positioning systems (GPS) fail to reliably operate. iLPS enables decimeter-level localization of numerous tags concurrently with wireless communication using frequency-shifting active reflector anchors and orthogonal frequency division multiple access (OFDMA) waveforms. OFDMA signals are devised with carefully assigned subcarriers so that each tag can estimate the time-difference-of-arrival (TDoA) by analyzing the channel impulse response (CIR) from the main and reflector anchors without interfering each other. The proposed active reflection scheme efficiently eliminates the stringent time synchronization requirement while providing the diversity gain for enhanced information decoding reliability at the tag. Significant challenges from NLOS multipaths and the usage of relatively narrow bandwidth of 80MHz in the ISM-band are successfully mitigated by machine learning assisted algorithms. Field trials with the prototype system on Universal Software Radio Peripheral (USRP) confirm that iLPS can achieve decimeter-level accuracy localization and concurrent wireless communication over up to 100m distances.
Mingyu Yang 0002, Li-Xuan Chuo, Karan Suri, Hun-Seok Kim
INFOCOM6
2019 Low Complexity, Hardware-Efficient Neighbor-Guided SGM Optical Flow for Low-Power Mobile Vision Applications
abstract
Accurate, low-latency, and energy-efficient optical flow estimation is a fundamental kernel function to enable several real-time vision applications on mobile platforms. This paper presents neighbor-guided semi-global matching (NG-fSGM), a new low-complexity optical flow algorithm tailored for low-power mobile applications. NG-fSGM obtains high accuracy optical flow by aggregating local matching costs over a semiglobal region, successfully resolving local ambiguity in textureless and occluded regions. The proposed NG-fSGM aggressively prunes the search space based on neighboring pixels' information to significantly lower the algorithm complexity from the original fSGM. As a result, NG-fSGM achieves 17.9× reduction in the number of computations and 8.37× reduction in memory space compared to the original fSGM without compromising its algorithm accuracy. A multicore architecture for NG-fSGM is implemented in hardware to quantify algorithm complexity and power consumption. The proposed architecture realizes NG-fSGM with overlapping blocks processed in parallel to enhance throughput and to lower power consumption. The eightcore architecture achieves 20 M pixel/s (66 frames/s for VGA) throughput with 9.6 mm2area at 679.2-mW power consumption in 28-nm node.
Ziyun Li 0001, Jiang Xiang, Luyao Gong, David T. Blaauw, Chaitali Chakrabarti, Hun-Seok Kim
IEEE Trans. Circuits Syst. Video Technol.6
2018 OuterSPACE: An Outer Product Based Sparse Matrix Multiplication Accelerator
abstract
Sparse matrices are widely used in graph and data analytics, machine learning, engineering and scientific applications. This paper describes and analyzes OuterSPACE, an accelerator targeted at applications that involve large sparse matrices. OuterSPACE is a highly-scalable, energy-efficient, reconfigurable design, consisting of massively parallel Single Program, Multiple Data (SPMD)-style processing units, distributed memories, high-speed crossbars and High Bandwidth Memory (HBM). We identify redundant memory accesses to non-zeros as a key bottleneck in traditional sparse matrix-matrix multiplication algorithms. To ameliorate this, we implement an outer product based matrix multiplication technique that eliminates redundant accesses by decoupling multiplication from accumulation. We demonstrate that traditional architectures, due to limitations in their memory hierarchies and ability to harness parallelism in the algorithm, are unable to take advantage of this reduction without incurring significant overheads. OuterSPACE is designed to specifically overcome these challenges. We simulate the key components of our architecture using gem5 on a diverse set of matrices from the University of Florida's SuiteSparse collection and the Stanford Network Analysis Project and show a mean speedup of 7.9× over Intel Math Kernel Library on a Xeon CPU, 13.0× against cuSPARSE and 14.0× against CUSP when run on an NVIDIA K40 GPU, while achieving an average throughput of 2.9 GFLOPS within a 24 W power budget in an area of 87 mm2.
Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siying Feng, Chaitali Chakrabarti, Hun-Seok Kim, David T. Blaauw, Trevor N. Mudge, Ronald G. Dreslinski
HPCA7
2018 HDM: Hyper-Dimensional Modulation for Robust Low-Power Communications
abstract
This paper introduces hyper-dimensional modulation (HDM), a new class of practical modulation scheme for robust communication among low-power and low- complexity devices. Unlike conventional orthogonal modulations, HDM conveys numerous information bits per symbol by combining hyper-dimensional vectors that are not strictly orthogonal to each other. Information bits are spread across many elements in the hyper-dimensional vector, thus HDM is tolerant of element-wise failures in high noise channels. Bit error rate (BER) evaluation results confirm that uncoded HDM with 256-dimension outperforms low density parity check (LDPC) and Polar codes with the same block length of 256. Analysis reveals HDM demodulation complexity is lower than that of LDPC and Polar decoders when the block length is relatively small. Moreover, HDM provides graceful tradeoffs between data rate and signal-to-noise ratio for robust short message communications among power- and complexity- constrained devices.
Hun-Seok Kim
ICC1
2018 Implementation and Evaluation of Bi-Directional WiFi Back-channel Communication
abstract
This paper presents implementation and validation of innovative back-channel wireless communication techniques for ultra-low power (ULP) devices. Back-channel schemes allow ULP devices that are not WiFi-compliant to communicate with already-deployed WiFi infrastructure without any hardware modification. This paper introduces an improved payload bit crafting procedure to create back-channel messages embedded in standard WiFi OFDM packets. In addition, the concept of bi-directional WiFi back-channel is newly introduced and validated in the prototype system. Using a commercial WiFi chip and Universal Software Radio Peripheral (USRP) X310 hardware platform, two downlink back-channel schemes as well as an uplink back-channel scheme are successfully demonstrated rendering realistic real-time communication performance. Field trials are performed for over-the-air wireless back-channel bi-directional communications. System characterization confirms that the proposed back-channel can operate at lower signal-to-noise ratio than what standard WiFi systems require.
Wenhao Peng, Yu Wang 0130, Li-Xuan Chuo, Karan Suri, David D. Wentzloff, Hun-Seok Kim
PIMRC8
2017 A Programmable Galois Field Processor for the Internet of Things
abstract
This paper investigates the feasibility of a unified processor architecture to enable error coding flexibility and secure communication in low power Internet of Things (IoT) wireless networks. Error coding flexibility for wireless communication allows IoT applications to exploit the large tradeoff space in data rate, link distance and energy-efficiency. As a solution, we present a light-weight Galois Field (GF) processor to enable energy-efficient block coding and symmetric/asymmetric cryptography kernel processing for a wide range of GF sizes (2m, m = 2, 3, ..., 233) and arbitrary irreducible polynomials. Program directed connections among primitive GF arithmetic units enable dynamically configured parallelism to efficiently perform either four-way SIMD 5- to 8-bit GF operations, including multiplicative inverse, or a wide bit-width (e.g., 32-bit) GF product in a single cycle. To illustrate our ideas, we synthesized our GF processor in a 28nm technology. Compared to a baseline software implementation optimized for a general purpose ARM M0+ processor, our processor exhibits a 5-20 x speedup for a range of error correction codes and symmetric/asymmetric cryptography applications. Additionally, our proposed GF processor consumes 431μW at 0.9V and 100MHz, and achieves 35.5pJ/b energy efficiency while executing AES operations at 12.2Mbps. We achieve this within an area of 0.01mm2.
Shengshuo Lu, David T. Blaauw, Ronald G. Dreslinski, Trevor N. Mudge, Hun-Seok Kim
ISCA7
2017 RF-Echo: A Non-Line-of-Sight Indoor Localization System Using a Low-Power Active RF Reflector ASIC Tag
abstract
Long-range low-power localization is a key technology that enables a host of new applications of wireless sensor nodes. We present RF-Echo, a new low-power RF localization solution that achieves decimeter accuracy in long range indoor non-line-of-sight (NLOS) scenarios. RF-Echo introduces a custom-designed active RF reflector ASIC (application specific integrated circuit) fabricated in a 180nm CMOS process which echoes a frequency-shifted orthogonal frequency division multiplexing (OFDM) signal originally generated from an anchor. The proposed technique is based on time-of-flight (ToF) estimation in the frequency domain that effectively eliminates inter-carrier and inter-symbol interference in multipath-rich indoor NLOS channels. RF-Echo uses a relatively narrow bandwidth of $\leq$80 MHz which does not require an expensive very high sampling rate analog-to-digital converter (ADC). Unlike ultra-wideband (UWB) systems, the active reflection scheme is designed to operate at a relatively low carrier frequency that can penetrate building walls and other blocking objects for challenging NLOS scenarios. Since the bandwidth at lower frequencies (2.4 GHz and sub-1 GHz) is severely limited, we propose novel signal processing algorithms as well as machine learning techniques to significantly enhance the localization resolution given the bandwidth constraint of the proposed system. The newly fabricated tag IC consumes 62.8 mW active power. The software defined radio (SDR) based anchor prototype is rapidly deployable without the need for accurate synchronization among anchors and tags. Field trials conducted in a university building confirm up to 85 m operation with decimeter accuracy for robust 2D localization.
Li-Xuan Chuo, Zhihong Luo, Dennis Sylvester, David T. Blaauw, Hun-Seok Kim
MobiCom5
2016 Software-Defined, WiFi and BLE Compliant Back-Channel for Ultra-Low Power Wireless Communication
abstract
In this paper, we present an innovative back-channel wireless communication concept. The proposed back- channel communication enables ultra-low power (ULP) devices that are neither WiFi (IEEE 802.11a/g/n) nor Bluetooth Low Energy (BLE) compliant to receive messages from both WiFi and BLE transmitters. This allows communication between heterogeneous devices beyond the boundary of WiFi and BLE standards without hardware modification on already-deployed infrastructure. Back- channel messages are created in frequency shift keying (FSK) modulation format by feeding carefully crafted bit sequences in the payload of the WiFi or BLE packets. Systematic algorithms are introduced to embed any desired back-channel messages in FSK format on WiFi and/or BLE compliant packets. Using a commercial off-the-shelf (COTS) narrowband FSK receiver, we demonstrate successful reception of back-channel messages from WiFi and BLE transmitters. Although this COTS narrowband FSK receiver is not specifically designed for the back-channel communication, less than 1% packet error rate (PER) is achieved, validating the concept of software defined back- channel communication among heterogeneous wireless devices.
Huajun Zhang 0001, David D. Wentzloff, Hun-Seok Kim
GLOBECOM3
2016 A low power software-defined-radio baseband processor for the Internet of Things
abstract
In this paper, we define a configurable Software Defined Radio (SDR) baseband processor design for the Internet of Things (IoT). We analyzed the fundamental algorithms in communications systems on IoT devices to enable a microarchitecture design that supports many IoT standards and custom nonstandard communications. Based on this analysis, we propose a custom SIMD execution model coupled with a scalar unit. We introduce several architectural optimizations to this design: streaming registers, variable bit width datapath, dedicated ALUs for critical kernels, and an optimized flexible reduction network. We employ voltage scaling and clock gating to further reduce the power, while more than a 100% time margin has been reserved for reliable operation in the near-threshold region. Together our architectural enhancements lead to a 71× power reduction compared to a classic general purpose SDR SIMD architecture. Our IoT SDR datapath has sub-mW power consumption based on SPICE simulation, and is placed and routed to fit within an area of 0.074mm2in a 28nm process. We implemented several essential elementary signal processing kernels and combined them to demonstrate two end-to-end upper bound systems, 802.15.4-OQPSK and Bluetooth Low Energy. Our full SDR baseband system consists of a configurable SIMD with a control plane MCU and memory. For comparison, the best commercial wireless transceiver consumes 23.8mW for the entire wireless system (digital/RF/ analog). We show that our digital system power is below 2mW, in other words only 8% of the total system power. The wireless system is dominated by RF/analog power comsumption, thus the price of flexibility that SDR affords is small. We believe this work is unique in demonstrating the value of baseband SDR in the low power IoT domain.
Shengshuo Lu, Hun-Seok Kim, David T. Blaauw, Ronald G. Dreslinski, Trevor N. Mudge
HPCA3
2016 Low complexity optical flow using neighbor-guided semi-global matching
abstract
This paper presents Neighbor-Guided SemiGlobal Matching (NG-fSGM), a new method for optical flow. It is based on SGM, a popular dynamic programming algorithm for stereo vision, where the disparity of each pixel is calculated by aggregating local matching costs over the entire image to resolve local ambiguity in texture-less and occluded regions. Unlike conventional SGM, NG-fSGM operates on a subset of the search space that has been aggressively pruned based on neighboring pixels' information. Our proposed method achieves a fast approximation of SGM with significantly simpler cost aggregation and flow computation. Compared to a prior SGM extension for optical flow, the proposed NG-fSGM provides about 9x reduction in the number of computations and 5x reduction in the memory requirement with only 0.17% accuracy degradation when evaluated with Middlebury benchmark test cases.
Jiang Xiang, Ziyun Li 0001, David T. Blaauw, Hun-Seok Kim, Chaitali Chakrabarti
ICIP4
2016 Energy-Autonomous Wireless Communication for Millimeter-Scale Internet-of-Things Sensor Nodes
abstract
This paper presents an energy-autonomous wireless communication system for ultra-small Internet-of-Things (IoT) platforms. In the proposed system, all necessary components, including the battery, energy-harvesting solar cells, and the RF antenna, are fully integrated within a millimeter-scale form factor. Designing an energy-optimized wireless communication system for such a miniaturized platform is challenging because of unique system constraints imposed by the ultra-small system dimension. The proposed system targets orders of magnitude improvement in wireless communication energy efficiency through a comprehensive system-level analysis that jointly optimizes various system parameters, such as node dimension, modulation scheme, synchronization protocol, RF/analog/digital circuit specifications, carrier frequency, and a miniaturized 3-D antenna. We propose a new protocol and modulation schemes that are specifically designed for energy-scarce ultra-small IoT nodes. These new schemes exploit abundant signal processing resources on gateway devices to simplify design for energy-scarce ultra-small sensor nodes. The proposed dynamic link adaptation guarantees that the ultra-small IoT node always operates in the most energy efficient mode for a given operating scenario. The outcome is a truly energy-optimized wireless communication system to enable various classes of new applications, such as implanted smart-dust devices.
Nikolaos Chiotellis, Li-Xuan Chuo, Carl Pfeiffer, Yao Shi 0001, Ronald G. Dreslinski, Anthony Grbic, Trevor N. Mudge, David D. Wentzloff, David T. Blaauw, Hun-Seok Kim
IEEE J. Sel. Areas Commun.11
2016 Back-Channel Wireless Communication Embedded in WiFi-Compliant OFDM Packets
abstract
This paper presents innovative back-channel wireless communication techniques for ultra-low power (ULP) devices. The concept of embedded back-channel communication is proposed to enable a variety of new applications by inter-connecting heterogeneous ULP devices through existing orthogonal frequency division multiplexing (OFDM)-based WiFi (IEEE 802.11a/g/n/ac) networks. The proposed back-channel communication allows ULP devices to decode messages embedded in WiFi OFDM packets even if these ULP devices are incapable of demodulating OFDM. The proposed back-channel signaling has unique properties that are easily detectable by non-WiFi ULP receivers consuming sub-mW of active power. The proposed scheme eliminates the need for specialized transmitter hardware or dedicated channel resources for embedded back-channel signal transmission. Instead, carefully sequenced data bit streams will generate back-channel messages from already-deployed WiFi infrastructure without any hardware modification. This paper demonstrates that WiFi OFDM back-channel communication is feasible in various modulation formats, such as pulse position modulation, pulse phase shift keying, or frequency shift keying. Systematic algorithms are unveiled to create back-channel messages in various modulation formats from a WiFi standard compliant datapath. Comprehensive bit error rate performance analysis of various WiFi back-channel communication schemes is derived and validated in realistic multi-path frequency selective fading channels.
Hun-Seok Kim, David D. Wentzloff
IEEE J. Sel. Areas Commun.1
2013 Forwarding metamorphosis: fast programmable match-action processing in hardware for SDN
abstract
In Software Defined Networking (SDN) the control plane is physically separate from the forwarding plane. Control software programs the forwarding plane (e.g., switches and routers) using an open interface, such as OpenFlow. This paper aims to overcomes two limitations in current switching chips and the OpenFlow protocol: i) current hardware switches are quite rigid, allowing ``Match-Action'' processing on only a fixed set of fields, and ii) the OpenFlow specification only defines a limited repertoire of packet processing actions. We propose the RMT (reconfigurable match tables) model, a new RISC-inspired pipelined architecture for switching chips, and we identify the essential minimal set of action primitives to specify how headers are processed in hardware. RMT allows the forwarding plane to be changed in the field without modifying hardware. As in OpenFlow, the programmer can specify multiple match tables of arbitrary width and depth, subject only to an overall resource limit, with each table configurable for matching on arbitrary fields. However, RMT allows the programmer to modify all header fields much more comprehensively than in OpenFlow. Our paper describes the design of a 64 port by 10 Gb/s switch chip implementing the RMT model. Our concrete design demonstrates, contrary to concerns within the community, that flexible OpenFlow hardware switch implementations are feasible at almost no additional cost or power.
Pat Bosshart, Glen Gibb, Hun-Seok Kim, George Varghese, Nick McKeown, Martin Izzard, Fernando A. Mujica, Mark Horowitz
SIGCOMM3
2012 Coding for jointly optimizing energy and peak current in deep sub-micron VLSI interconnects
abstract
In deep sub-micron processes, on-chip interconnect is becoming the delay bottleneck and predominant source of power consumption. Simultaneous switching of large buses pose a great challenge on peak current as well. In this paper, we present a novel bus coding technique, based on transition pattern codes (TPC), to perform joint optimization. A TPC scheme has been constructed employing a joint cost function on energy and peak current. The encoder and decoder of the code has been synthesized using a commercial 28nm process and the power, delay and area overhead has been evaluated. HSPICE simulations in 28nm show up to 70% reduction in peak current and 15% reduction in energy consumption compared to an uncoded bus.
Eric P. Kim, Hun-Seok Kim, Manish Goel
ISCAS2
2011 Power Optimized PA Clipping for MIMO-OFDM Systems
abstract
For a multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) system that is being pushed into power amplifier (PA) saturation, this letter investigates power-optimized PA clipping. Our goal is to identify the optimum clipping level for a MIMO-OFDM system that delivers the desired bit error rate (BER) with minimum power consumption in the PA. We present a complete theoretical framework resulting in an analytical expression for the BER of a MIMO-OFDM system subject to PA clipping. PA power saving is addressed by the total degradation metric, which shows that as much as 6dB power reduction can be achieved by proper choice of the clipping level.
Hun-Seok Kim, Babak Daneshrad
IEEE Trans. Wirel. Commun.1
2010 Energy-Constrained Link Adaptation for MIMO OFDM Wireless Communication Systems
abstract
We present a link adaptation strategy for multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) based wireless communications. Our objective is to choose the optimal mode that will maximize energy efficiency or data throughput subject to a given quality of service (QoS) constraint. We formulate the link adaptation problem as a convex optimization problem and expand the set of parameters under the control of the link adaptation protocol to include: number of spatial streams, number of transmit/receive antennas, use of spatial multiplexing or space time block coding (STBC), constellation size, bandwidth, transmit power and choice of maximum likelihood (ML) or zero-forcing (ZF) for MIMO decoding. Additionally, we increase the fidelity of the energy consumption modeling relative to the prior art. The resulting solution allows us to easily and quickly search the space of possible system parameters to deliver on the QoS with minimal energy consumption. Moreover, it provides us insight into where crossovers occur in the choice of the radio parameters. Application of the results to a generic MIMO-OFDM radio shows that the proposed strategy can provide an order of magnitude improvement in energy efficiency or data throughput relative to a static strategy.
Hun-Seok Kim, Babak Daneshrad
IEEE Trans. Wirel. Commun.1
2009 A Theoretical Treatment of PA Power Optimization in Clipped MIMO-OFDM Systems
abstract
This work presents a theoretical framework for determining the optimal clipping level for a multiple input multiple output (MIMO) orthogonal frequency division multiplexing (OFDM) system. We start by modeling of the clipping noise, and propose a maximum likelihood (ML) receiver for the resulting signal. The bit error rate (BER) for this ML receiver is then derived for MIMO spatial multiplexing, space time coding, and cyclic delay diversity schemes. A search through the tradeoff space between the BER performance penalty and power amplifier (PA) power consumption advantage reveals optimum clipping levels of the system. Using an IEEE 802.11n like system, we show that the optimum clipping as predicted by this approach can improve PA power efficiency as much as 70%.
Hun-Seok Kim, Babak Daneshrad
GLOBECOM1