VLDB 2026 Research / reviewers in the wild / expert
Xue Lin 0001
dblp:94/7236-1 · also Xue (Shelley) Lin
· DBLP profile ↗
131ranked-venue papers
15as first author
61since 2021 · last 2026
0000-0001-6210-8883ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 10 first-author · 24 since 2021Artificial intelligence and machine learning · 43 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 21 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diff-StyGS: 3D Gaussian Splatting Stylization via Tuning-Free Multi-view Sparse Diffusion
Zhenglun Kong, Yanzhi Wang 0001, Pu Zhao 0001, Xue Lin 0001 |
ICPR (11) | 6 |
| 2026 | AFL-PRF: Adaptive Federated Learning for Low-Quality Data: Enhancing Performance, Robustness, and FairnessabstractFederated learning (FL) enables collaborative model training across distributed clients while preserving privacy, yet its decentralized nature makes it vulnerable to poisoned updates and performance degradation under highly skewed data. Prior studies typically treat accuracy, robustness, and fairness separately, leaving open the challenge of a unified solution. We propose AFL-PRF, an adaptive federated learning framework that simultaneously enhances accuracy, robustness, and fairness in adversarial and heterogeneous environments. AFL-PRF integrates three key techniques. First, an exponential adaptive weighting mechanism dynamically scales client updates, suppressing poisoned or unreliable contributions while retaining meaningful signals from benign but low-quality clients. Second, a client prioritization strategy guided by the novel Weight Update Divergence (WUD) score promotes reliable updates and their benign neighbors, preventing malicious gradients from dominating aggregation. Third, sensitivity profiling identifies fully connected (FC) layers as highly vulnerable due to large weight variance, motivating a selective clipping strategy that filters extreme updates in these layers while preserving normal learning dynamics. Extensive experiments on benchmark datasets demonstrate that AFL-PRF consistently outperforms state-of-the-art baselines, achieving over 30% improvement in robustness and 20% enhancement in fairness, while maintaining superior predictive accuracy. By unifying adaptive weighting, client prioritization, and targeted clipping, AFL-PRF establishes a new benchmark for federated learning under poisoned and highly non-IID conditions. Pinrui Yu, Longtian Ye, Geng Yuan, Ningfang Mi, Xue Lin 0001 |
WACV | 6 |
| 2025 | Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance AssessmentabstractStructured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redundant weight groups at a coarse-grained granularity. Current structured pruning methods for LLMs typically depend on a singular granularity for assessing weight importance, resulting in notable performance degradation in downstream tasks. Intriguingly, our empirical investigations reveal that utilizing unstructured pruning, which achieves better performance retention by pruning weights at a finer granularity, \emph{i.e.}, individual weights, yields significantly varied sparse LLM structures when juxtaposed to structured pruning. This suggests that evaluating both holistic and individual assessments for weight importance are essential for LLM pruning. Building on this insight, we introduce the Hybrid-grained Weight Importance Assessment (HyWIA), a novel method that merges fine-grained and coarse-grained evaluations of weight importance for the pruning of LLMs. Leveraging an attention mechanism, HyWIA adaptively determines the optimal blend of granularity in weight importance assessments in an end-to-end pruning manner. Extensive experiments on LLaMA-V1/V2, Vicuna, Baichuan, and Bloom across various benchmarks demonstrate the effectiveness of HyWIA in pruning LLMs. For example, HyWIA surpasses the cutting-edge LLM-Pruner by an average margin of 2.82% in accuracy across seven downstream tasks when pruning LLaMA-7B by 50%. Jun Liu 0075, Zhenglun Kong, Pu Zhao 0001, Changdi Yang, Xuan Shen, Hao Tang 0005, Geng Yuan, Wei Niu 0002, Wenbin Zhang 0002, Xue Lin 0001, Yanzhi Wang 0001 |
AAAI | 10 |
| 2025 | LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network InferenceabstractFor FPGA-based neural network accelerators, digital signal processing (DSP) blocks have traditionally been the cornerstone for handling multiplications. This paper introduces LUTMUL, which harnesses the potential of look-up tables (LUTs) for performing multiplications. The availability of LUTs typically outnumbers that of DSPs by a factor of 100, offering a significant computational advantage. By exploiting this advantage of LUTs, our method demonstrates a potential boost in the performance of FPGA-based neural network accelerators with a reconfigurable dataflow architecture. Our approach challenges the conventional peak performance on DSP-based accelerators and sets a new benchmark for efficient neural network inference on FPGAs. Experimental results demonstrate that our design achieves the best inference speed among all FPGA-based accelerators, achieving a throughput of 1627 images per second and maintaining a top-1 accuracy of 70.95% on the ImageNet dataset. Yanyue Xie, Zhengang Li 0001, Dana Diaconu, Suranga Handagala, Miriam Leeser, Xue Lin 0001 |
ASP-DAC | 6 |
| 2025 | Pruning then Reweighting: Towards Data-Efficient Training of Diffusion ModelsabstractDespite the remarkable generation capabilities of Diffusion Models (DMs), conducting training and inference remains computationally expensive. Previous works have been devoted to accelerating diffusion sampling, but achieving data-efficient diffusion training has often been overlooked. In this work, we investigate efficient diffusion training from the perspective of dataset pruning. Inspired by the principles of data-efficient training for generative models such as generative adversarial networks (GANs), we first extend the data selection scheme used in GANs to DM training, where data features are encoded by a surrogate model, and a score criterion is then applied to select the coreset. To further improve the generation performance, we employ a class-wise reweighting approach, which derives class weights through distributionally robust optimization (DRO) over a pre-trained reference DM. For a pixel-wise DM (DDPM) on CIFAR-10, experiments demonstrate the superiority of our methodology over existing approaches and its effectiveness in image synthesis comparable to that of the original full-data model while achieving the speed-up between 2.34× and 8.32×. Additionally, our method could be generalized to latent DMs (LDMs), e.g., Masked Diffusion Transformer (MDT) and Stable Diffusion (SD), and achieves competitive generation capability on ImageNet. Sijia Liu 0001, Xue Lin 0001 |
ICASSP | 4 |
| 2025 | RoRA: Efficient Fine-Tuning of LLM with Reliability Optimization for Rank AdaptationabstractFine-tuning helps large language models (LLM) recover degraded information and enhance task performance. Although Low-Rank Adaptation (LoRA) is widely used and effective for fine-tuning, we have observed that its scaling factor can limit or even reduce performance as the rank size increases. To address this issue, we propose RoRA (Rank-adaptive Reliability Optimization), a simple yet effective method for optimizing LoRA’s scaling factor. By replacing α/r with $\alpha /\sqrt r $, RoRA ensures improved performance as rank size increases. Moreover, RoRA enhances low-rank adaptation in fine-tuning uncompressed models and excels in the more challenging task of accuracy recovery when fine-tuning pruned models. Extensive experiments demonstrate the effectiveness of RoRA in fine-tuning both uncompressed and pruned models. RoRA surpasses the state-of-the-art (SOTA) in average accuracy and robustness on LLaMA-7B/13B, LLaMA2-7B, and LLaMA3-8B, specifically outperforming LoRA and DoRA by 6.5% and 2.9% on LLaMA-7B, respectively. In pruned model fine-tuning, RoRA shows significant advantages; for SHEARED-LLAMA-1.3, a LLaMA-7B with 81.4% pruning, RoRA achieves 5.7% higher average accuracy than LoRA and 3.9% higher than DoRA. Jun Liu 0075, Zhenglun Kong, Peiyan Dong, Xuan Shen, Pu Zhao 0001, Hao Tang 0005, Geng Yuan, Wei Niu 0002, Wenbin Zhang 0002, Xue Lin 0001, Yanzhi Wang 0001 |
ICASSP | 10 |
| 2025 | Sparse Learning for State Space Models on MobileabstractTransformer models have been widely investigated in different domains by providing long-range dependency handling and global contextual awareness, driving the development of popular AI applications such as ChatGPT, Gemini, and Alexa.
State Space Models (SSMs) have emerged as strong contenders in the field of sequential modeling, challenging the dominance of Transformers. SSMs incorporate a selective mechanism that allows for dynamic parameter adjustment based on input data, enhancing their performance.
However, this mechanism also comes with increasing computational complexity and bandwidth demands, posing challenges for deployment on resource-constraint mobile devices.
To address these challenges without sacrificing the accuracy of the selective mechanism, we propose a sparse learning framework that integrates architecture-aware compiler optimizations. We introduce an end-to-end solution--$\mathbf{C}_4^n$ kernel sparsity, which prunes $n$ elements from every four contiguous weights, and develop a compiler-based acceleration solution to ensure execution efficiency for this sparsity on mobile devices.
Based on the kernel sparsity, our framework generates optimized sparse models targeting specific sparsity or latency requirements for various model sizes. We further leverage pruned weights to compensate for the remaining weights, enhancing downstream task performance.
For practical hardware acceleration, we propose $\mathbf{C}_4^n$-specific optimizations combined with a layout transformation elimination strategy.
This approach mitigates inefficiencies arising from fine-grained pruning in linear layers and improves performance across other operations.
Experimental results demonstrate that our method achieves superior task performance compared to other semi-structured pruning methods and achieves up-to 7$\times$ speedup compared to llama.cpp framework on mobile devices. Xuan Shen, Hangyu Zheng, Yifan Gong 0004, Zhenglun Kong, Changdi Yang, Zheng Zhan 0001, Yushu Wu, Xue Lin 0001, Yanzhi Wang 0001, Pu Zhao 0001, Wei Niu 0002 |
ICLR | 8 |
| 2025 | Taming Diffusion for Dataset Distillation with High RepresentativenessabstractRecent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of diffusion, it has been introduced to this field for generating distilled images. In this paper, we systematically investigate issues present in current diffusion-based dataset distillation methods, including inaccurate distribution matching, distribution deviation with random noise, and separate sampling. Building on this, we propose D$^3$HR, a novel diffusion-based framework to generate distilled datasets with high representativeness. Specifically, we adopt DDIM inversion to map the latents of the full dataset from a low-normality latent domain to a high-normality Gaussian domain, preserving information and ensuring structural consistency to generate representative latents for the distilled dataset. Furthermore, we propose an efficient sampling scheme to better align the representative latents with the high-normality Gaussian distribution. Our comprehensive experiments demonstrate that D$^3$HR can achieve higher accuracy across different model architectures compared with state-of-the-art baselines in dataset distillation. Source code: https://github.com/lin-zhao-resoLve/D3HR. Yushu Wu, Xinru Jiang, Jianyang Gu, Yanzhi Wang 0001, Xiaolin Xu 0001, Pu Zhao 0001, Xue Lin 0001 |
ICML | 8 |
| 2025 | FairSMOE: Mitigating Multi-Attribute Fairness Problem with Sparse Mixture-of-ExpertsabstractReal‐world datasets usually contain multiple attributes, making it essential to ensure fairness across all of them simultaneously. However, different attributes may vary in difficulty, and no existing approaches have effectively addressed this issue. Consequently, an attribute‐adaptive strategy is needed to achieve fairness for all attributes. Multi‐task Learning (MTL) leverages shared information to optimize multiple tasks concurrently, while Sparsely‐Gated Mixture‐of‐Experts (SMoE) can dynamically allocate computational resources to the most needed tasks. In this work, we formulate multi‐attribute fairness issue as an MTL problem and employ SMoE to achieve desirable performance across all attributes simultaneously. We first analyze the feasibility and find the potentiality by formalizing multi-attribute fairness problem into a MTL problem and mitigating it by using SMoE. However, vanilla SMoE could lead to over-utilization problem which causes sub-optimal performance. We then proposed an innovative SMoE framework for multi-attribute fair image classification, which further improves multi-attribute fairness by redesigning the MoE layer and routing policy with fairness consideration. Extensive experiments demonstrated the effectiveness. Taking a DeiT-Small as the backbone, we achieve 77.25% and 86.01% accuracy on the ISIC2019 and CelebA dataset respectively with Multi-attribute Predictive Quality Disparity (PQD) score of 0.801 and 0.787, beating current state-of-the-art methods Muffin, InfoFair and MultiFair. Changdi Yang, Zheng Zhan 0001, Ci Zhang, Yifan Gong 0004, Zichong Meng, Jun Liu 0075, Xuan Shen, Hao Tang 0005, Geng Yuan, Pu Zhao 0001, Xue Lin 0001, Yanzhi Wang 0001 |
IJCAI | 12 |
| 2025 | Can Adversarial Examples be Parsed to Reveal Victim Model Information?abstractNumerous adversarial attack methods have been developed to generate imperceptible image perturbations that cause erroneous predictions in state-of-the-art machine learning (ML) models, particularly deep neural networks (DNNs). Despite extensive research on adversarial examples, limited efforts have been made to explore the hidden characteristics carried by these perturbations. In this study, we investigate the feasibility of deducing information about the victim model (VM)—specifically, characteristics such as architecture type, kernel size, activation function, and weight sparsity—from adversarial examples. We approach this problem as a supervised learning task, where we aim to attribute categories of VM characteristics to individual adversarial examples. To facilitate this, we have assembled a dataset of adversarial attacks spanning seven types, generated from 135 victim models systematically varied across five architecture types, three kernel size configurations, three activation functions, and three levels of weight sparsity. We demonstrate that a supervised model parsing network (MPN) can effectively extract concealed details of the VM from adversarial examples. We also validate the practicality of this approach by evaluating the effects of various factors on parsing performance, such as different input formats and generalization to out-of-distribution cases. Furthermore, we highlight the connection between model parsing and attack transferability by showing how the MPN can uncover VM attributes in transfer attacks. Yuguang Yao, Jiancheng Liu, Yifan Gong 0004, Xiaoming Liu 0002, Yanzhi Wang 0001, Xue Lin 0001, Sijia Liu 0001 |
WACV | 6 |
| 2025 | Q-TempFusion: Quantization-Aware Temporal Multi-Sensor Fusion on Bird's-Eye View RepresentationabstractRecent advancements in bird's-eye view (BEV) perception models have highlighted the superior performance of LiDAR-camera fusion systems over single-modality approaches, garnering considerable interest in the field. Despite the progress, the integration of temporal information, a technique that has considerably benefitted camera-only BEV models, remains underexplored for LiDAR-camera fusion. This paper presents Q-TempFusion, a novel approach for temporal multi-sensor fusion designed to enhance the BEV model's inference speed while keeping high predictive performance compared with the current state-of-the-art. Moreover, we are the first to make the multi-modality BEV model profiling on hardware devices. To address the challenges of substantial memory demands and non-trivial latency that hinder deployment in on-vehicle systems, particularly when temporal dynamics are incorporated into complex multi-sensor models, we introduce an activation-aware quantization framework to generate the fully 8-bit quantized Q-TempFusion model based on the profiling result, which can be directly deployed to target devices with negligible detection performance degradation. Our experiments show that our Q-TempFusion (8-bit) achieves 70.3% mAP and 72.7% NDS with 3×~18× FPS improvement over leading multi-modality baselines and the Q-TempFusion (32-bit) achieves 72.1% mAP and 74.8% NDS, comparable to SOTA multi-modality approaches. The results suggest that Q-TempFusion is a promising step toward real-time multi-sensor BEV applications, setting a new benchmark for efficient and reliable perception. Pinrui Yu, Zhenglun Kong, Pu Zhao 0001, Peiyan Dong, Hao Tang 0005, Fei Sun 0002, Xue Lin 0001, Yanzhi Wang 0001 |
WACV | 7 |
| 2025 | Mobile-3DCNN: An Acceleration Framework for Ultra-Real-Time Execution of Large 3D CNNs on Mobile DevicesabstractIt is challenging to deploy 3D Convolutional Neural Networks (3D CNNs) on mobile devices, specifically if both real-time execution and high inference accuracy are in demand, because the increasingly large model size and complex model structure of 3D CNNs usually require tremendous computation and memory resources. Weight pruning is proposed to mitigate this challenge. However, existing pruning is either not compatible with modern parallel architectures, resulting in long inference latency or subject to significant accuracy degradation. This article proposes an end-to-end 3D CNN acceleration framework based on pruning/compilation co-design called Mobile-3DCNN that consists of two parts: a novel, fine-grained structured pruning enhanced by a prune/Winograd adaptive selection (that is mobile-hardware-friendly and can achieve high pruning accuracy), and a set of compiler optimization and code generation techniques enabled by our pruning (to fully transform the pruning benefit to real performance gains). The evaluation demonstrates that Mobile-3DCNN outperforms state-of-the-art end-to-end DNN acceleration frameworks that support 3D CNN execution on mobile devices, Alibaba Mobile Neural Networks and Pytorch-Mobile with speedup up to 34× with minor accuracy degradation, proving it is possible to execute high-accuracy large 3D CNNs on mobile devices in real-time (or even ultra-real-time). Wei Niu 0002, Mengshu Sun, Zhengang Li 0001, Jou-An Chen, Jiexiong Guan, Xipeng Shen, Jun Liu 0075, Yanzhi Wang 0001, Xue Lin 0001, Bin Ren 0002 |
ACM Trans. Archit. Code Optim. | 10 |
| 2025 | TSLA: A Task-Specific Learning Adaptation for Semantic Segmentation on Autonomous Vehicles PlatformabstractAutonomous driving platforms encounter diverse driving scenarios, each with varying hardware resources and precision requirements. Given the computational limitations of embedded devices, it is crucial to consider computing costs when deploying on target platforms like the DRIVE PX 2. Our objective is to customize the semantic segmentation network according to the computing power and specific scenarios of autonomous driving hardware. We implement dynamic adaptability through a three-tier control mechanism—width multiplier, classifier depth, and classifier kernel—allowing fine-grained control over model components based on hardware constraints and task requirements. This adaptability facilitates broad model scaling, targeted refinement of the final layers, and scenario-specific optimization of kernel sizes, leading to improved resource allocation and performance. Additionally, we leverage Bayesian Optimization with surrogate modeling to efficiently explore hyperparameter spaces under tight computational budgets. Our approach addresses scenario-specific and task-specific requirements through automatic parameter search, accommodating the unique computational complexity and accuracy needs of autonomous driving. It scales its multiply-accumulate operations (MACs) for task-specific learning adaptation (TSLA), resulting in alternative configurations tailored to diverse self-driving tasks. These TSLA customizations maximize computational capacity and model accuracy, optimizing hardware utilization. Jun Liu 0075, Zhenglun Kong, Pu Zhao 0001, Weihao Zeng 0002, Hao Tang 0005, Xuan Shen, Changdi Yang, Wenbin Zhang 0002, Geng Yuan, Wei Niu 0002, Xue Lin 0001, Yanzhi Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2024 | SuperFlow: A Fully-Customized RTL-to-GDS Design Automation Flow for Adiabatic Quantum- Flux - Parametron Superconducting CircuitsabstractSuperconducting circuits, like Adiabatic Quantum- Flux-Parametron (AQFP), offer exceptional energy efficiency but face challenges in physical design due to sophisticated spacing and timing constraints. Current design tools often neglect the importance of constraint adherence throughout the entire design flow. In this paper, we propose SuperFlow, a fully-customized RTL-to-GDS design flow tailored for AQFP devices. SuperFlow leverages a synthesis tool based on CMOS technology to transform any input RTL netlist to an AQFP-based netlist. Subsequently, we devise a novel place-and-route procedure that simultaneously con-siders wirelength, timing, and routability for AQFP circuits. The process culminates in the generation of the AQFP circuit layout, followed by a Design Rule Check (DR C) to identify and rectify any layout violations. Our experimental results demonstrate that SuperFlow achieves 12.8% wirelength improvement on average and 12.1 % better timing quality compared with previous state- of-the-art placers for AQFP circuits. Yanyue Xie, Peiyan Dong, Geng Yuan, Zhengang Li 0001, Masoud Zabihi, Chao Wu 0006, Sung-En Chang, Xue Lin 0001, Caiwen Ding, Nobuyuki Yoshikawa, Olivia Chen, Yanzhi Wang 0001 |
DATE | 9 |
| 2024 | Finding Needles in a Haystack: A Black-Box Approach to Invisible Watermark Detection
Minzhou Pan, Zhenting Wang, Xin Dong 0009, Vikash Sehwag, Lingjuan Lyu, Xue Lin 0001 |
ECCV (33) | 6 |
| 2024 | Rethinking Token Reduction for State Space ModelsabstractRecent advancements in State Space Models (SSMs) have attracted significant interest, particularly in models optimized for parallel training and handling long-range dependencies.Architectures like Mamba have scaled to billions of parameters with selective SSM.To facilitate broader applications using Mamba, exploring its efficiency is crucial.While token reduction techniques offer a straightforward post-training strategy, we find that applying existing methods directly to SSMs leads to substantial performance drops.Through insightful analysis, we identify the reasons for this failure and the limitations of current techniques.In response, we propose a tailored, unified post-training token reduction method for SSMs.Our approach integrates token importance and similarity, thus taking advantage of both pruning and merging, to devise a fine-grained intra-layer token reduction strategy.Extensive experiments show that our method improves the average accuracy by 5.7% to 13.1% on six benchmarks with Mamba-2 compared to existing methods, while significantly reducing computational demands and memory requirements.1 Zheng Zhan 0001, Yushu Wu, Zhenglun Kong, Changdi Yang, Yifan Gong 0004, Xuan Shen, Xue Lin 0001, Pu Zhao 0001, Yanzhi Wang 0001 |
EMNLP | 7 |
| 2024 | SDA: Low-Bit Stable Diffusion Acceleration on Edge FPGAsabstractThis paper introduces SDA, the first effort to adapt the expensive stable diffusion (SD) model for edge FPGA deployment. First, we apply quantization-aware training to quantize its weights to 4 -bit and activations to 8 -bit ($W 4 A 8$) with a negligible accuracy loss. Based on that, we propose a high-performance hybrid systolic array (hybridSA) architecture that natively executes convolution and attention operators across varying quantization bit-widths (e.g., $W 4 A 8$ and all 8 -bit $Q K^{T} V$ in attention). To improve computational efficiency, hybridSA integrates diverse DSP packing techniques into hybrid weightstationary and output-stationary dataflows that are optimized for convolution and attention. It also supports flexible dataflow transitions to address the distinct demands of its output sequence by subsequent nonlinear operators. Moreover, we observe that nonlinear operators become the new performance bottleneck after the acceleration of convolution and attention, and offload them onto the FPGA as well. To reduce the latency of each nonlinear operator, we pipeline its own execution at a fine granularity. To minimize the resource utilization of nonlinear operators, we carefully balance their execution with hybridSA in a coarse-grained pipeline. Experimental results demonstrate that our low-bit ($W 4 A 8$) SDA accelerator on the embedded AMDXilinx ZCU102 FPGA achieves a speedup of $97.3 \times$ (which takes about $\mathrm{2 . 1}$ minutes for one SD inference), compared to the original SD-v1.5 model on the ARM Cortex-A53 CPU (which takes about 3.5 hours for one SD inference). Our SDA project is open sourced here: https://github.com/Michaela1224/SDA_code. Geng Yang 0001, Yanyue Xie, Zhong Jia Xue, Sung-En Chang, Yanyu Li, Peiyan Dong, Jie Lei 0001, Weiying Xie, Yanzhi Wang 0001, Xue Lin 0001, Zhenman Fang |
FPL | 10 |
| 2024 | Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision TransformersabstractVision transformers (ViTs) have demonstrated their superior accuracy for computer vision tasks compared to convolutional neural networks (CNNs). However, ViT models are often computation-intensive for efficient deployment on resource-limited edge devices. This work proposes Quasar-ViT, a hardware-oriented quantization-aware architecture search framework for ViTs, to design efficient ViT models for hardware implementation while preserving the accuracy. First, Quasar-ViT trains a supernet using our row-wise flexible mixed-precision quantization scheme, mixed-precision weight entanglement, and supernet layer scaling techniques. Then, it applies an efficient hardware-oriented search algorithm, integrated with hardware latency and resource modeling, to determine a series of optimal subnets from supernet under different inference latency targets. Finally, we propose a series of model-adaptive designs on the FPGA platform to support the architecture search and mitigate the gap between the theoretical computation reduction and the practical inference speedup. Our searched models achieve 101.5, 159.6, and 251.6 frames-per-second (FPS) inference speed on the AMD/Xilinx ZCU102 FPGA with 80.4%, 78.6%, and 74.9% top-1 accuracy, respectively, for the ImageNet dataset, consistently outperforming prior works. Zhengang Li 0001, Alec Lu, Yanyue Xie, Zhenglun Kong, Mengshu Sun, Hao Tang 0005, Zhong Jia Xue, Peiyan Dong, Caiwen Ding, Yanzhi Wang 0001, Xue Lin 0001, Zhenman Fang |
ICS | 11 |
| 2024 | FasterVD: On Acceleration of Video Diffusion Models
Pinrui Yu, Timothy Rupprecht, Zhenglun Kong, Pu Zhao 0001, Yanyu Li, Octavia I. Camps, Xue Lin 0001, Yanzhi Wang 0001 |
IJCAI | 9 |
| 2024 | Adaptive Homogeneity-Based Client Selection Policy for Federated LearningabstractFederated learning (FL) is a distributed paradigm that enables multiple clients or edge devices to collaboratively train a model without sharing their local data. The FL system has to tackle a significant challenge due to the non-IID nature of client data. Traditional methods of client selection usually suffer from the problem of variable test accuracies as well as slow convergence because of their incapability in effectively handle data heterogeneity across clients. In this paper, we propose an Adaptive Homogeneity-Based Client Selection Policy (ASTraFL) to address this challenge. ASTraFL dynamically selects clients whose data distributions optimally complement the current state of the global model, focusing on increased homogeneity of the selected client data in each training round. Our experiments demonstrate that ASTraFL can accelerate convergence speed and ensure the learning process's robustness. Pinrui Yu, Geng Yuan, Xue Lin 0001, Ningfang Mi |
ISNCC | 4 |
| 2024 | HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image CompressionabstractThis paper investigates the challenging problem of learned image compression (LIC) with extreme low bitrates. Previous LIC methods based on transmitting quantized continuous features often yield blurry and noisy reconstruction due to the severe quantization loss. While previous LIC methods based on learned codebooks that discretize visual space usually give poor-fidelity reconstruction due to the insufficient representation power of limited codewords in capturing faithful details. We propose a novel dual-stream framework, HyrbidFlow, which combines the continuous-feature-based and codebook-based streams to achieve both high perceptual quality and high fidelity under extreme low bitrates. The codebook-based stream benefits from the high-quality learned codebook priors to provide high quality and clarity in reconstructed images. The continuous feature stream targets at maintaining fidelity details. To achieve the ultra low bitrate, a masked token-based transformer is further proposed, where we only transmit a masked portion of codeword indices and recover the missing indices through token generation guided by information from the continuous feature stream. We also develop a bridging correction network to merge the two streams in pixel decoding for final image reconstruction, where the continuous stream features rectify biases of the codebook-based pixel decoder to impose reconstructed fidelity details. Experimental results demonstrate superior performance across several datasets under extremely low bitrates, compared with existing single-stream codebook-based or continuous-feature-based LIC methods. Yanyue Xie, Wei Jiang 0001, Wei Wang 0311, Xue Lin 0001, Yanzhi Wang 0001 |
ACM Multimedia | 5 |
| 2024 | Search for Efficient Large Language ModelsabstractLarge Language Models (LLMs) have long held sway in the realms of artificial intelligence research.
Numerous efficient techniques, including weight pruning, quantization, and distillation, have been embraced to compress LLMs, targeting memory reduction and inference acceleration, which underscore the redundancy in LLMs.
However, most model compression techniques concentrate on weight optimization, overlooking the exploration of optimal architectures.
Besides, traditional architecture search methods, limited by the elevated complexity with extensive parameters, struggle to demonstrate their effectiveness on LLMs.
In this paper, we propose a training-free architecture search framework to identify optimal subnets that preserve the fundamental strengths of the original LLMs while achieving inference acceleration.
Furthermore, after generating subnets that inherit specific weights from the original LLMs, we introduce a reformation algorithm that utilizes the omitted weights to rectify the inherited weights with a small amount of calibration data.
Compared with SOTA training-free structured pruning works that can generate smaller networks, our method demonstrates superior performance across standard benchmarks.
Furthermore, our generated subnets can directly reduce the usage of GPU memory and achieve inference acceleration. Xuan Shen, Pu Zhao 0001, Yifan Gong 0004, Zhenglun Kong, Zheng Zhan 0001, Yushu Wu, Ming Lin 0002, Chao Wu 0006, Xue Lin 0001, Yanzhi Wang 0001 |
NeurIPS | 9 |
| 2024 | A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep LearningabstractDeep Neural Networks (DNNs) have been applied as an effective machine learning algorithm to tackle problems in different domains. However, the endeavor to train sophisticated DNN models can stretch from days into weeks, presenting substantial obstacles in the realm of research focused on large-scale DNN architectures. Distributed Deep Learning (DDL) contributes to accelerating DNN training by distributing training workloads across multiple computation accelerators, for example, graphics processing units (GPUs). Despite the considerable amount of research directed toward enhancing DDL training, the influence of data loading on GPU utilization and overall training efficacy remains relatively overlooked. It is non-trivial to optimize data-loading in DDL applications that need intensive central processing unit (CPU) and input/output (I/O) resources to process enormous training data. When multiple DDL applications are deployed on a system (e.g., Cloud and High-Performance Computing (HPC) system), the lack of a practical and efficient technique for data-loader allocation incurs GPU idleness and degrades the training throughput. Therefore, our work first focuses on investigating the impact of data-loading on the global training throughput. We then propose a throughput prediction model to predict the maximum throughput for an individual DDL training application. By leveraging the predicted results, A-Dloader is designed to dynamically allocate CPU and I/O resources to concurrently running DDL applications and use the data-loader allocation as a knob to reduce GPU idle intervals and thus improve the overall training throughput. We implement and evaluate A-Dloader in a DDL framework for a series of DDL applications arriving and completing across the runtime. Our experimental results show that A-Dloader can achieve a 28.9% throughput improvement and a 10% makespan improvement compared with allocating resources evenly across applications. Danlin Jia, Geng Yuan, Xue Lin 0001, Ningfang Mi |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | Hardware-Friendly 3-D CNN Acceleration With Balanced Kernel Group SparsityabstractBeing capable of extracting more information than 2D Convolutional Neural Networks (CNNs), 3D CNNs have been playing a vital role in video analysis tasks like human action recognition, but their massive operations hinder the real-time execution on edge devices with constrained computation and memory resources. Although various model compression techniques have been applied to accelerate 2D CNNs, there are rare efforts in investigating hardware-friendly pruning of 3D CNNs and acceleration on customizable edge platforms like FPGAs. This work starts from proposing a kernel group row-column (KGRC) weight sparsity pattern, which is fine-grained to achieve high pruning ratios with negligible accuracy loss, and balanced across kernel groups to achieve high computation parallelism on hardware. The reweighted pruning algorithm for this sparsity is then presented and performed on 3D CNNs, followed by quantization under different precisions. Along with model compression, FPGA-based accelerators with four modes are designed in support of the kernel group sparsity in multiple dimensions. The co-design framework of the pruning algorithm and the accelerator is tested on two representative 3D CNNs, namely C3D and R(2+1)D, with the Xilinx ZCU102 FPGA platform for action recognition. The experimental results indicate that the accelerator implementation with the KGRC sparsity and 8-bit quantization achieves a good balance between the speedup and model accuracy, leading to acceleration ratios of 4.12× for C3D and 3.85× for R(2+1)D compared with the 16-bit baseline designs supporting only dense models. Mengshu Sun, Kaidi Xu, Xue Lin 0001, Yongli Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Towards Real-Time Segmentation on the EdgeabstractThe research in real-time segmentation mainly focuses on desktop GPUs. However, autonomous driving and many other applications rely on real-time segmentation on the edge, and current arts are far from the goal. In addition, recent advances in vision transformers also inspire us to re-design the network architecture for dense prediction task. In this work, we propose to combine the self attention block with lightweight convolutions to form new building blocks, and employ latency constraints to search an efficient sub-network. We train an MLP latency model based on generated architecture configurations and their latency measured on mobile devices, so that we can predict the latency of subnets during search phase. To the best of our knowledge, we are the first to achieve over 74% mIoU on Cityscapes with semi-real-time inference (over 15 FPS) on mobile GPU from an off-the-shelf phone. Yanyu Li, Changdi Yang, Pu Zhao 0001, Geng Yuan, Wei Niu 0002, Jiexiong Guan, Hao Tang 0005, Minghai Qin, Qing Jin, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001 |
AAAI | 11 |
| 2023 | Pruning Parameterization with Bi-level Optimization for Efficient Semantic Segmentation on the EdgeabstractWith the ever-increasing popularity of edge devices, it is necessary to implement real-time segmentation on the edge for autonomous driving and many other applications. Vision Transformers (ViTs) have shown considerably stronger results for many vision tasks. However, ViTs with the fullattention mechanism usually consume a large number of computational resources, leading to difficulties for real- time inference on edge devices. In this paper, we aim to derive ViTs with fewer computations and fast inference speed to facilitate the dense prediction of semantic segmentation on edge devices. To achieve this, we propose a pruning parameterization method to formulate the pruning problem of semantic segmentation. Then we adopt a bi-level optimization method to solve this problem with the help of implicit gradients. Our experimental results demonstrate that we can achieve 38.9 mIoU on ADE20K val with a speed of 56.5 FPS on Samsung S21, which is the highest mIoU under the same computation constraint with real-time inference. Changdi Yang, Pu Zhao 0001, Yanyu Li, Wei Niu 0002, Jiexiong Guan, Hao Tang 0005, Minghai Qin, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001 |
CVPR | 9 |
| 2023 | Late Breaking Results: Fast Fair Medical Applications? Hybrid Vision Models Achieve the Fairness on the EdgeabstractAs edge devices become readily available and indispensable, there is an urgent need for effective and efficient intelligent applications to be deployed widespread. However, fairness has always been an issue, especially in edge medical applications. Compared to convolutional neuron networks (CNNs), Vision Transformer (ViT) has a better ability to extract global information, which will contribute to alleviating the unfairness problem. Typically, ViTs consume large amounts of computational and memory resources, which hinders their usage on edge. In this work, we propose a novel hardware-efficient Vision Model search framework for the fair dermatology classification, namely HeViFa. Experimental results show that HeViFa could search for a hybrid ViT model that reaches 173.1 FPS on a Samsung S21 mobile phone with 85.71% accuracy on the light skin dataset and 80.85% accuracy on the dark skin dataset. Note that HeViFa can reach both the highest accuracy and fairness under similar latency constrain on multiple edge devices (Samsung S21 mobile phone, iPhone 13 Pro and Raspberry PI). Changdi Yang, Yi Sheng 0001, Peiyan Dong, Zhenglun Kong, Yanyu Li, Pinrui Yu, Lei Yang 0018, Xue Lin 0001 |
DAC | 8 |
| 2023 | ESRU: Extremely Low-Bit and Hardware-Efficient Stochastic Rounding Unit Design for Low-Bit DNN TrainingabstractStochastic rounding is crucial in the low-bit (e.g., 8-bit) training of deep neural networks (DNNs) to achieve high accuracy. One of the drawbacks of prior studies is that they require a large number of high-precision stochastic rounding units (SRUs) to guarantee low-bit DNN accuracy, which involves considerable hardware overhead. In this paper, we use extremely low-bit SRUs (ESRUs) to save a large number of hardware resources during low-bit DNN training. However, a naively designed ESRU introduces a biased distribution of random numbers, causing accuracy degradation. To address this issue, we further propose an ESRU design with a plateau-shape distribution. The plateau-shape distribution in our ESRU design is implemented with the combination of an LFSR (linear-feedback shift register) and an inverted LFSR, which avoids LFSR packing and turns an inherent LFSR drawback into an advantage in our efficient ESRU design. Experimental results using state-of-the-art DNN models demonstrate that, compared to the prior 24-bit SRU with 24-bit pseudo-random number generators (PRNG), our 8-bit ESRU with 3-bit PRNG reduces the SRU hardware resource usage by 9.75x while achieving slightly higher accuracy. Sung-En Chang, Geng Yuan, Alec Lu, Mengshu Sun, Yanyu Li, Zhengang Li 0001, Yanyue Xie, Minghai Qin, Xue Lin 0001, Zhenman Fang, Yanzhi Wang 0001 |
DATE | 10 |
| 2023 | HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision TransformersabstractWhile vision transformers (ViTs) have continuously achieved new milestones in the field of computer vision, their sophisticated network architectures with high computation and memory costs have impeded their deployment on resource-limited edge devices. In this paper, we propose a hardware-efficient image-adaptive token pruning framework called HeatViT for efficient yet accurate ViT acceleration on embedded FPGAs. Based on the inherent computational patterns in ViTs, we first adopt an effective, hardware-efficient, and learnable head-evaluation token selector, which can be progressively inserted before transformer blocks to dynamically identify and consolidate the non-informative tokens from input images. Moreover, we implement the token selector on hardware by adding miniature control logic to heavily reuse existing hardware components built for the backbone ViT. To improve the hardware efficiency, we further employ 8-bit fixed-point quantization and propose polynomial approximations with regularization effect on quantization error for the frequently used nonlinear functions in ViTs. Compared to existing ViT pruning studies, under the similar computation cost, HeatViT can achieve 0.7% ~ 8.9% higher accuracy; while under the similar model accuracy, HeatViT can achieve more than 28.4% ~ 65.3% computation reduction, for various widely used ViTs, including DeiT-T, DeiT-S, DeiT-B, LV-ViT-S, and LV-ViT-M, on the ImageNet dataset. Compared to the baseline hardware accelerator, our implementations of HeatViT on the Xilinx ZCU102 FPGA achieve 3.46×~4.89× speedup with a trivial resource utilization overhead of 8%~11% more DSPs and 5%~8% more LUTs. Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Kenneth Liu, Zhenglun Kong, Zhengang Li 0001, Xue Lin 0001, Zhenman Fang, Yanzhi Wang 0001 |
HPCA | 9 |
| 2023 | Fast and Fair Medical AI on the Edge Through Neural Architecture Search for Hybrid Vision ModelsabstractAs edge devices become readily available and indispensable, there is an urgent need for effective and efficient intelligent applications to be deployed widespread. However, fairness has always been an issue, especially in edge medical applications. Although many approaches have been proposed to mitigate the unfairness problem, their edge performance is not desirable. By examining the fairness performance of different network architectures, we observed that compared to pure convolutional neuron network (CNN) architecture, hybrid models with CNN and Vision Transformer (ViT) have exhibited better performance in terms of fairness and accuracy. After further analyzing the feature maps of intermediate layers of CNNs, ViTs, and hybrid models, we found that ViT has a strong ability to extract global information, which contributes to alleviating the unfairness problem. However, ViTs consume large amounts of computational and memory resources, which hinders their application on edge devices. To address the challenges abovementioned, we propose the first hardware-oriented co-design NAS framework to explore hybrid ViT-CNN architecture for the fair dermatology classification, namely HeViFa, which can produce light-weight models for edge devices with low unfairness scores and high classification accuracy. Experimental results show that compared with FaHaNa-Small, HeViFa-Small could search for a hybrid ViT model that reaches 10.57% and 4.03% higher accuracy as well as 0.179 and 0.0403 higher PQD score on Mix and Fitzpatrick17k dataset, repectively, and speed up by 1.21 × on Samsung S21 mobile phone, 1.18 × on iPhone 13 Pro and 1.37 × on Raspberry Pi. Changdi Yang, Yi Sheng 0001, Peiyan Dong, Zhenglun Kong, Yanyu Li, Pinrui Yu, Lei Yang 0018, Xue Lin 0001, Yanzhi Wang 0001 |
ICCAD | 8 |
| 2023 | ASSET: Robust Backdoor Data Detection Across a Multiplicity of Deep Learning Paradigms
Minzhou Pan, Yi Zeng 0005, Lingjuan Lyu, Xue Lin 0001, Ruoxi Jia 0001 |
USENIX Security Symposium | 4 |
| 2023 | The Autonomous Vehicle Assistant (AVA): Emerging technology design supporting blind and visually impaired travelers in autonomous transportation
Paul D. S. Fink, Stacy A. Doore, Xue Lin 0001, Matthew Maring, Pu Zhao 0001, Aubree Nygaard, Grant Beals, Richard R. Corey, Raymond J. Perry, Katherine Freund, Velin D. Dimitrov, Nicholas A. Giudice |
Int. J. Hum. Comput. Stud. | 3 |
| 2022 | A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep LearningabstractDeep Neural Network (DNN) has been applied as an effective machine learning algorithm to tackle problems in different domains. However, training a sophisticated DNN model takes days to weeks and becomes a challenge in constructing research on large-scale DNN models. Distributed Deep Learning (DDL) contributes to accelerating DNN training by distributing training workloads across multiple computation accelerators (e.g., GPUs). Although a surge of research works has been devoted to optimizing DDL training, the impact of data-loading on GPU usage and training performance has been relatively under-explored. It is non-trivial to optimize data-loading in DDL applications that need intensive CPU and I/O resources to process enormous training data. When multiple DDL applications are deployed on a system (e.g., Cloud and HPC), the lack of a practical and efficient technique for data-loader allocation incurs GPU idleness and degrades the training throughput. Therefore, our work first focuses on investigating the impact of data-loading on the global training throughput. We then propose a throughput prediction model to predict the maximum throughput for an individual DDL training application. By leveraging the predicted results, A-Dloader is designed to dynamically allocate CPU and I/O resources to concurrently running DDL applications and use the data-loader allocation as a knob to reduce GPU idle intervals and thus improve the overall training throughput. We implement and evaluate A-Dloader in a DDL framework for a series of DDL applications arriving and completing across the runtime. Our experimental results show that A-Dloader can achieve a 23.5% throughput improvement and a 10% makespan improvement, compared to allocating resources evenly across applications. Danlin Jia, Geng Yuan, Xue Lin 0001, Ningfang Mi |
CLOUD | 3 |
| 2022 | Hardware-efficient stochastic rounding unit design for DNN training: late breaking resultsabstractStochastic rounding is crucial in the training of low-bit deep neural networks (DNNs) to achieve high accuracy. Unfortunately, prior studies require a large number of high-precision stochastic rounding units (SRUs) to guarantee the low-bit DNN accuracy, which involves considerable hardware overhead. In this paper, we propose an automated framework to explore hardware-efficient low-bit SRUs (ESRUs) that can still generate high-quality random numbers to guarantee the accuracy of low-bit DNN training. Experimental results using state-of-the-art DNN models demonstrate that, compared to the prior 24-bit SRU with 24-bit pseudo random number generator (PRNG), our 8-bit with 3-bit PRNG reduces the SRU resource usage by 9.75× while achieving a higher accuracy. Sung-En Chang, Geng Yuan, Alec Lu, Mengshu Sun, Yanyu Li, Zhengang Li 0001, Yanyue Xie, Minghai Qin, Xue Lin 0001, Zhenman Fang, Yanzhi Wang 0001 |
DAC | 10 |
| 2022 | FPGA-aware automatic acceleration framework for vision transformer with mixed-scheme quantization: late breaking resultsabstractVision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand impose urgent needs for new hardware accelerator design methodology. This work proposes an FPGA-aware automatic ViT acceleration framework based on the proposed mixed-scheme quantization. To the best of our knowledge, this is the first FPGA-based ViT acceleration framework exploring model quantization. Compared with state-of-the-art ViT quantization work (algorithmic approach only without hardware acceleration), our quantization achieves 0.31% to 1.25% higher Top-1 accuracy under the same bit-width. Compared with the 32-bit floating-point baseline FPGA accelerator, our accelerator achieves around 5.6× improvement on the frame rate (i.e., 56.4 FPS vs. 10.0 FPS) with 0.83% accuracy drop for DeiT-base. Mengshu Sun, Zhengang Li 0001, Alec Lu, Geng Yuan, Yanyue Xie, Hao Tang 0005, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin 0001, Zhenman Fang |
DAC | 11 |
| 2022 | Fault-Tolerant Deep Neural Networks for Processing-In-Memory based Autonomous Edge SystemsabstractIn-memory deep neural network (DNN) accelerators will be the key for energy-efficient autonomous edge systems. The resistive random access memory (ReRAM) is a potential solution for the non-CMOS-based in-memory computing platform for energy-efficient autonomous edge systems, thanks to its promising characteristics, such as near-zero leakage-power and non-volatility. However, due to the hardware instability of ReRAM, the weights of the DNN model may deviate from the originally trained weights, resulting in accuracy loss. To mitigate this undesirable accuracy loss, we propose two stochastic fault-tolerant training methods to generally improve the models' robustness without dealing with individual devices. Moreover, we propose Stability Score-a comprehensive metric that serves as an indicator to the instability problem. Extensive experiments demonstrate that the DNN models trained using our proposed stochastic fault-tolerant training method achieve superior performance, which provides better flexibility, scalability, and deployability of ReRAM on the autonomous edge systems. Siyue Wang, Geng Yuan, Yanyu Li, Xue Lin 0001, Bhavya Kailkhura |
DATE | 5 |
| 2022 | FILM-QNN: Efficient FPGA Acceleration of Deep Neural Networks with Intra-Layer, Mixed-Precision QuantizationabstractWith the trend to deploy Deep Neural Network (DNN) inference models on edge devices with limited resources, quantization techniques have been widely used to reduce on-chip storage and improve computation throughput. However, existing DNN quantization work deploying quantization below 8-bit may be either suffering from evident accuracy loss or facing a big gap between the theoretical improvement of computation throughput and the practical inference speedup. In this work, we propose a general framework, called FILM-QNN, to quantize and accelerate multiple DNN models across different embedded FPGA devices. First, we propose the novel intra-layer, mixed-precision quantization algorithm that assigns different precisions onto the filters of each layer. The candidate precision levels and assignment granularity are determined from our empirical study with the capability of preserving accuracy and improving hardware parallelism. Second, we apply multiple optimization techniques for the FPGA accelerator architecture in support of quantized computations, including DSP packing, weight reordering, and data packing, to enhance the overall throughput with the available resources. Moreover, a comprehensive resource model is developed to balance the allocation of FPGA computation resources (LUTs and DSPs) as well as data transfer and on-chip storage resources (BRAMs) to accelerate the computations in mixed precisions within each layer. Finally, to improve the portability of FILM-QNN, we implement it using Vivado High-Level Synthesis (HLS) on Xilinx PYNQ-Z2 and ZCU102 FPGA boards. Our experimental results of ResNet-18, ResNet-50, and MobileNet-V2 demonstrate that the implementations with intra-layer, mixed-precision (95% of 4-bit weights and 5% of 8-bit weights, and all 5-bit activations) can achieve comparable accuracy (70.47%, 77.25%, and 65.67% for the three models) as the 8-bit (and 32-bit) versions and comparable throughput (214.8 FPS, 109.1 FPS, and 537.9 FPS on ZCU102) as the 4-bit designs. Mengshu Sun, Zhengang Li 0001, Alec Lu, Yanyu Li, Sung-En Chang, Xue Lin 0001, Zhenman Fang |
FPGA | 7 |
| 2022 | Auto-ViT-Acc: An FPGA-Aware Automatic Acceleration Framework for Vision Transformer with Mixed-Scheme QuantizationabstractVision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand impose urgent needs for new hardware accelerator design methodology. This work proposes an FPGA-aware automatic ViT acceleration framework based on the proposed mixed-scheme quantization. To the best of our knowledge, this is the first FPGA-based ViT acceleration framework exploring model quantization. Compared with state-of-the-art ViT quantization work (algorithmic approach only without hardware acceleration), our quantization achieves 0.47% to 1.36% higher Top-l accuracy under the same bit-width. Compared with the 32-bit floating-point baseline FPGA accelerator, our accelerator achieves around 5.6x improvement on the frame rate (i.e., 56.8 FPS vs. 10.0 FPS) with 0.71% accuracy drop on ImageNet dataset for DeiT-base. Zhengang Li 0001, Mengshu Sun, Alec Lu, Geng Yuan, Yanyue Xie, Hao Tang 0005, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin 0001, Zhenman Fang |
FPL | 11 |
| 2022 | Reverse Engineering of Imperceptible Adversarial Image Perturbations
Yifan Gong 0004, Yuguang Yao, Xiaoming Liu 0002, Xue Lin 0001, Sijia Liu 0001 |
ICLR | 6 |
| 2022 | Pruning-as-Search: Efficient Neural Architecture Search via Channel Pruning and Structural ReparameterizationabstractNeural architecture search (NAS) and network pruning are widely studied efficient AI techniques, but not yet perfect. NAS performs exhaustive candidate architecture search, incurring tremendous search cost. Though (structured) pruning can simply shrink model dimension, it remains unclear how to decide the per-layer sparsity automatically and optimally. In this work, we revisit the problem of layer-width optimization and propose Pruning-as-Search (PaS), an end-to-end channel pruning method to search out desired sub-network automatically and efficiently. Specifically, we add a depth-wise binary convolution to learn pruning policies directly through gradient descent. By combining the structural reparameterization and PaS, we successfully searched out a new family of VGG-like and lightweight networks, which enable the flexibility of arbitrary width with respect to each layer instead of each stage. Experimental results show that our proposed architecture outperforms prior arts by around 1.0% top-1 accuracy under similar inference speed on ImageNet-1000 classification task. Furthermore, we demonstrate the effectiveness of our width search on complex tasks including instance segmentation and image translation. Code and models are released. Yanyu Li, Pu Zhao 0001, Geng Yuan, Xue Lin 0001, Yanzhi Wang 0001 |
IJCAI | 4 |
| 2022 | Learning to Generate Image Source-Agnostic Universal Adversarial PerturbationsabstractAdversarial perturbations are critical for certifying the robustness of deep learning models. A ``universal adversarial perturbation'' (UAP) can simultaneously attack multiple images, and thus offers a more unified threat model, obviating an image-wise attack algorithm. However, the existing UAP generator is underdeveloped when images are drawn from different image sources (e.g., with different image resolutions). Towards an authentic universality across image sources, we take a novel view of UAP generation as a customized instance of ``few-shot learning'', which leverages bilevel optimization and learning-to-optimize (L2O) techniques for UAP generation with improved attack success rate (ASR). We begin by considering the popular model agnostic meta-learning (MAML) framework to meta-learn a UAP generator. However, we see that the MAML framework does not directly offer the universal attack across image sources, requiring us to integrate it with another meta-learning framework of L2O. The resulting scheme for meta-learning a UAP generator (i) has better performance (50% higher ASR) than baselines such as Projected Gradient Descent, (ii) has better performance (37% faster) than the vanilla L2O and MAML frameworks (when applicable), and (iii) is able to simultaneously handle UAP generation for different victim models and data sources. Pu Zhao 0001, Parikshit Ram, Songtao Lu, Yuguang Yao, Djallel Bouneffouf 0001, Xue Lin 0001, Sijia Liu 0001 |
IJCAI | 6 |
| 2022 | GRIM: A General, Real-Time Deep Learning Inference Framework for Mobile Devices Based on Fine-Grained Structured Weight SparsityabstractIt is appealing but challenging to achieve real-time deep neural network (DNN) inference on mobile devices, because even the powerful modern mobile devices are considered as "resource-constrained" when executing large-scale DNNs. It necessitates the sparse model inference via weight pruning, i.e., DNN weight sparsity, and it is desirable to design a new DNN weight sparsity scheme that can facilitate real-time inference on mobile devices while preserving a high sparse model accuracy. This paper designs a novel mobile inference acceleration framework GRIM that is General to both convolutional neural networks (CNNs) and recurrent neural networks (RNNs) and that achieves Real-time execution and high accuracy, leveraging fine-grained structured sparse model Inference and compiler optimizations for Mobiles. We start by proposing a new fine-grained structured sparsity scheme through the Block-based Column-Row (BCR) pruning. Based on this new fine-grained structured sparsity, our GRIM framework consists of two parts: (a) the compiler optimization and code generation for real-time mobile inference; and (b) the BCR pruning optimizations for determining pruning hyperparameters and performing weight pruning. We compare GRIM with Alibaba MNN, TVM, TensorFlow-Lite, a sparse implementation based on CSR, PatDNN, and ESE (a representative FPGA inference acceleration framework for RNNs), and achieve up to 14.08× speedup. Wei Niu 0002, Zhengang Li 0001, Peiyan Dong, Gang Zhou 0002, Xuehai Qian, Xue Lin 0001, Yanzhi Wang 0001, Bin Ren 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Mobile or FPGA? A Comprehensive Evaluation on Energy Efficiency and a Unified Optimization FrameworkabstractEfficient deployment of Deep Neural Networks (DNNs) on edge devices (i.e., FPGAs and mobile platforms) is very challenging, especially under a recent witness of the increasing DNN model size and complexity. Model compression strategies, including weight quantization and pruning, are widely recognized as effective approaches to significantly reduce computation and memory intensities, and have been implemented in many DNNs on edge devices. However, most state-of-the-art works focus on ad hoc optimizations, and there lacks a thorough study to comprehensively reveal the potentials and constraints of different edge devices when considering different compression strategies. In this article, we qualitatively and quantitatively compare the energy efficiency of FPGA-based and mobile-based DNN executions using mobile GPU and provide a detailed analysis. Based on the observations obtained from the analysis, we propose a unified optimization framework using block-based pruning to reduce the weight storage and accelerate the inference speed on mobile devices and FPGAs, achieving high hardware performance and energy-efficiency gain while maintaining accuracy. Geng Yuan, Peiyan Dong, Mengshu Sun, Wei Niu 0002, Zhengang Li 0001, Yuxuan Cai 0001, Yanyu Li, Jun Liu 0075, Weiwen Jiang, Xue Lin 0001, Bin Ren 0002, Xulong Tang, Yanzhi Wang 0001 |
ACM Trans. Embed. Comput. Syst. | 10 |
| 2022 | Non-Structured DNN Weight Pruning - Is It Beneficial in Any Platform?abstractLarge deep neural network (DNN) models pose the key challenge to energy efficiency due to the significantly higher energy consumption of off-chip DRAM accesses than arithmetic or SRAM operations. It motivates the intensive research on model compression with two main approaches. Weight pruning leverages the redundancy in the number of weights and can be performed in a non-structured, which has higher flexibility and pruning rate but incurs index accesses due to irregular weights, or structured manner, which preserves the full matrix structure with a lower pruning rate. Weight quantization leverages the redundancy in the number of bits in weights. Compared to pruning, quantization is much more hardware-friendly and has become a "must-do" step for FPGA and ASIC implementations. Thus, any evaluation of the effectiveness of pruning should be on top of quantization. The key open question is, with quantization, what kind of pruning (non-structured versus structured) is most beneficial? This question is fundamental because the answer will determine the design aspects that we should really focus on to avoid the diminishing return of certain optimizations. This article provides a definitive answer to the question for the first time. First, we build ADMM-NN-S by extending and enhancing ADMM-NN, a recently proposed joint weight pruning and quantization framework, with the algorithmic supports for structured pruning, dynamic ADMM regulation, and masked mapping and retraining. Second, we develop a methodology for fair and fundamental comparison of non-structured and structured pruning in terms of both storage and computation efficiency. Our results show that ADMM-NN-S consistently outperforms the prior art: 1) it achieves 348× , 36× , and 8× overall weight pruning on LeNet-5, AlexNet, and ResNet-50, respectively, with (almost) zero accuracy loss and 2) we demonstrate the first fully binarized (for all layers) DNNs can be lossless in accuracy in many cases. These results provide a strong baseline and credibility of our study. Based on the proposed comparison framework, with the same accuracy and quantization, the results show that non-structured pruning is not competitive in terms of both storage and computation efficiency. Thus, we conclude that structured pruning has a greater potential compared to non-structured pruning. We encourage the community to focus on studying the DNN inference acceleration with structured sparsity. Sheng Lin 0001, Shaokai Ye, Zhezhi He, Linfeng Zhang 0001, Geng Yuan, Sia Huat Tan, Zhengang Li 0001, Deliang Fan, Xuehai Qian, Xue Lin 0001, Kaisheng Ma, Yanzhi Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2022 | StructADMM: Achieving Ultrahigh Efficiency in Structured Pruning for DNNsabstractWeight pruning methods of deep neural networks (DNNs) have been demonstrated to achieve a good model pruning rate without loss of accuracy, thereby alleviating the significant computation/storage requirements of large-scale DNNs. Structured weight pruning methods have been proposed to overcome the limitation of irregular network structure and demonstrated actual GPU acceleration. However, in prior work, the pruning rate (degree of sparsity) and GPU acceleration are limited (to less than 50%) when accuracy needs to be maintained. In this work, we overcome these limitations by proposing a unified, systematic framework of structured weight pruning for DNNs. It is a framework that can be used to induce different types of structured sparsity, such as filterwise, channelwise, and shapewise sparsity, as well as nonstructured sparsity. The proposed framework incorporates stochastic gradient descent (SGD; or ADAM) with alternating direction method of multipliers (ADMM) and can be understood as a dynamic regularization method in which the regularization target is analytically updated in each iteration. Leveraging special characteristics of ADMM, we further propose a progressive, multistep weight pruning framework and a network purification and unused path removal procedure, in order to achieve higher pruning rate without accuracy loss. Without loss of accuracy on the AlexNet model, we achieve 2.58× and 3.65× average measured speedup on two GPUs, clearly outperforming the prior work. The average speedups reach 3.15× and 8.52× when allowing a moderate accuracy loss of 2%. In this case, the model compression for convolutional layers is 15.0× , corresponding to 11.93× measured CPU speedup. As another example, for the ResNet-18 model on the CIFAR-10 data set, we achieve an unprecedented 54.2× structured pruning rate on CONV layers. This is 32× higher pruning rate compared with recent work and can further translate into 7.6× inference time speedup on the Adreno 640 mobile GPU compared with the original, unpruned DNN model. We share our codes and models at the link http://bit.ly/2M0V7DO. Tianyun Zhang, Shaokai Ye, Xiaoyu Feng, Kaiqi Zhang 0003, Zhengang Li 0001, Jian Tang 0008, Sijia Liu 0001, Xue Lin 0001, Yongpan Liu, Makan Fardad, Yanzhi Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2022 | Automatic Mapping of the Best-Suited DNN Pruning Schemes for Real-Time Mobile AccelerationabstractWeight pruning is an effective model compression technique to tackle the challenges of achieving real-time deep neural network (DNN) inference on mobile devices. However, prior pruning schemes have limited application scenarios due to accuracy degradation, difficulty in leveraging hardware acceleration, and/or restriction on certain types of DNN layers. In this article, we propose a general, fine-grained structured pruning scheme and corresponding compiler optimizations that are applicable to any type of DNN layer while achieving high accuracy and hardware inference performance. With the flexibility of applying different pruning schemes to different layers enabled by our compiler optimizations, we further probe into the new problem of determining the best-suited pruning scheme considering the different acceleration and accuracy performance of various pruning schemes. Two pruning scheme mapping methods—one -search based and the other is rule based—are proposed to automatically derive the best-suited pruning regularity and block size for each layer of any given DNN. Experimental results demonstrate that our pruning scheme mapping methods, together with the general fine-grained structured pruning scheme, outperform the state-of-the-art DNN optimization framework with up to 2.48 \( \times \) and 1.73 \( \times \) DNN inference acceleration on CIFAR-10 and ImageNet datasets without accuracy loss. Yifan Gong 0004, Geng Yuan, Zheng Zhan 0001, Wei Niu 0002, Zhengang Li 0001, Pu Zhao 0001, Yuxuan Cai 0001, Sijia Liu 0001, Bin Ren 0002, Xue Lin 0001, Xulong Tang, Yanzhi Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2021 | RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesabstractMobile devices are becoming an important carrier for deep learning tasks, as they are being equipped with powerful, high-end mobile CPUs and GPUs. However, it is still a challenging task to execute 3D Convolutional Neural Networks (CNNs) targeting for real-time performance, besides high inference accuracy. The reason is more complex model structure and higher model dimensionality overwhelm the available computation/storage resources on mobile devices. A natural way may be turning to deep learning weight pruning techniques. However, the direct generalization of existing 2D CNN weight pruning methods to 3D CNNs is not ideal for fully exploiting mobile parallelism while achieving high inference accuracy. This paper proposes RT3D, a model compression and mobile acceleration framework for 3D CNNs, seamlessly integrating neural network weight pruning and compiler code generation techniques. We propose and investigate two structured sparsity schemes i.e., the vanilla structured sparsity and kernel group structured (KGS) sparsity that are mobile acceleration friendly. The vanilla sparsity removes whole kernel groups, while KGS sparsity is a more fine-grained structured sparsity that enjoys higher flexibility while exploiting full on-device parallelism. We propose a reweighted regularization pruning algorithm to achieve the proposed sparsity schemes. The inference time speedup due to sparsity is approaching the pruning rate of the whole model FLOPs (floating point operations). RT3D demonstrates up to 29.1x speedup in end-to-end inference time comparing with current mobile frameworks supporting 3D CNNs, with moderate 1%~1.5% accuracy loss. The end-to-end inference time for 16 video frames could be within 150 ms, when executing representative C3D and R(2+1)D models on a cellphone. For the first time, real-time execution of 3D CNNs is achieved on off-the-shelf mobiles. Wei Niu 0002, Mengshu Sun, Zhengang Li 0001, Jou-An Chen, Jiexiong Guan, Xipeng Shen, Yanzhi Wang 0001, Sijia Liu 0001, Xue Lin 0001, Bin Ren 0002 |
AAAI | 9 |
| 2021 | Real-Time Mobile Acceleration of DNNs: From Computer Vision to Medical ApplicationsabstractWith the growth of mobile vision applications, there is a growing need to break through the current performance limitation of mobile platforms, especially for computationally intensive applications, such as object detection, action recognition, and medical diagnosis. To achieve this goal, we present our unified real-time mobile DNN inference acceleration framework, seamlessly integrating hardware-friendly, structured model compression with mobile-targeted compiler optimizations. We aim at an unprecedented, realtime performance of such large-scale neural network inference on mobile devices. A fine-grained block-based pruning scheme is proposed to be universally applicable to all types of DNN layers, such as convolutional layers with different kernel sizes and fully connected layers. Moreover, it is also successfully extended to 3D convolutions. With the assist of our compiler optimizations, the fine-grained block-based sparsity is fully utilized to achieve high model accuracy and high hardware acceleration simultaneously. To validate our framework, three representative fields of applications are implemented and demonstrated, object detection, activity detection, and medical diagnosis. All applications achieve real-time inference using an off-the-shelf smartphone, outperforming the representative mobile DNN inference acceleration frameworks by up to 6.7x in speed. The demonstrations of these applications can be found in the following link: https://bit.ly/39lWpYu. Hongjia Li 0003, Geng Yuan, Wei Niu 0002, Yuxuan Cai 0001, Mengshu Sun, Zhengang Li 0001, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001 |
ASP-DAC | 8 |
| 2021 | Intrinsic Examples: Robust Fingerprinting of Deep Neural Networks
Siyue Wang, Pu Zhao 0001, Xiao Wang 0028, Sang (Peter) Chin, Thomas Wahl, Yunsi Fei, Qi Alfred Chen, Xue Lin 0001 |
BMVC | 8 |
| 2021 | NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationabstractWith the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Prior methods towards this goal, including model compression and network architecture search (NAS), are largely performed independently, and do not fully consider compiler-level optimizations which is a must-do for mobile acceleration. In this work, we first propose (i) a general category of fine-grained structured pruning applicable to various DNN layers, and (ii) a comprehensive, compiler automatic code generation framework supporting different DNNs and different pruning schemes, which bridge the gap of model compression and NAS. We further propose NPAS, a compiler-aware unified network pruning and architecture search. To deal with large search space, we propose a meta-modeling procedure based on reinforcement learning with fast evaluation and Bayesian optimization, ensuring the total number of training epochs comparable with representative NAS frameworks. Our framework achieves 6.7ms, 5.9ms, and 3.9ms ImageNet inference times with 78.2%, 75% (MobileNet-V3 level), and 71% (MobileNet-V2 level) Top-1 accuracy respectively on an off-the-shelf mobile phone, consistently outperforming prior work. Zhengang Li 0001, Geng Yuan, Wei Niu 0002, Pu Zhao 0001, Yanyu Li, Yuxuan Cai 0001, Xuan Shen, Zheng Zhan 0001, Zhenglun Kong, Qing Jin, Zhiyu Chen 0003, Sijia Liu 0001, Kaiyuan Yang 0001, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001 |
CVPR | 16 |
| 2021 | Neural Pruning Search for Real-Time Object Detection of Autonomous VehiclesabstractObject detection plays an important role in self-driving cars for security development. However, mobile systems on self-driving cars with limited computation resources lead to difficulties for object detection. To facilitate this, we propose a compiler-aware neural pruning search framework to achieve high-speed inference on autonomous vehicles for 2D and 3D object detection. The framework automatically searches the pruning scheme and rate for each layer to find a best-suited pruning for optimizing detection accuracy and speed performance under compiler optimization. Our experiments demonstrate that for the first time, the proposed method achieves (close-to) real-time, 55ms and 97ms inference times for YOLOv4 based 2D object detection and PointPillars based 3D detection, respectively, on an off-the-shelf mobile phone with minor (or no) accuracy loss. Pu Zhao 0001, Geng Yuan, Yuxuan Cai 0001, Wei Niu 0002, Qi Liu 0017, Wujie Wen, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001 |
DAC | 9 |
| 2021 | Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkabstractDeep Neural Networks (DNNs) have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference on edge-computing platforms, e.g., ASICs, FPGAs, and embedded systems, are extensively investigated. Due to the huge model size and computation amount, model compression is a critical step to deploy DNN models on edge devices. This paper focuses on weight quantization, a hardware-friendly model compression approach that is complementary to weight pruning.Unlike existing methods that use the same quantization scheme for all weights, we propose the first solution that applies different quantization schemes for different rows of the weight matrix. It is motivated by (1) the distribution of the weights in the different rows are not the same; and (2) the potential of achieving better utilization of heterogeneous FPGA hardware resources. To achieve that, we first propose a hardware-friendly quantization scheme named sum-of-power-of-2 (SP2) suitable for Gaussian-like weight distribution, in which the multiplication arithmetic can be replaced with logic shifter and adder, thereby enabling highly efficient implementations with the FPGA LUT resources. In contrast, the existing fixed-point quantization is suitable for Uniform-like weight distribution and can be implemented efficiently by DSP. Then to fully explore the resources, we propose an FPGA-centric mixed scheme quantization (MSQ) with an ensemble of the proposed SP2 and the fixed-point schemes. Combining the two schemes can maintain, or even increase accuracy due to better matching with weight distributions.For the FPGA implementations, we develop a parameterized architecture with heterogeneous Generalized Matrix Multiplication (GEMM) cores-one using LUTs for computations with SP2 quantized weights and the other utilizing DSPs for fixed-point quantized weights. Given the partition ratio among the two schemes based on resource characterization, MSQ quantization training algorithm derives an optimally quantized model for the FPGA implementation. We evaluate our FPGA-centric quantization framework across multiple application domains. With optimal SP2/fixed-point ratios on two FPGA devices, i.e., Zynq XC7Z020 and XC7Z045, we achieve performance improvement of 2.1 × -4.1 × compared to solely exploiting DSPs for all multiplication operations. In addition, the CNN implementations with the proposed MSQ scheme can achieve higher accuracy and comparable hardware utilization efficiency compared to the state-of-the-art designs. Sung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi, Hayden Kwok-Hay So, Xuehai Qian, Yanzhi Wang 0001, Xue Lin 0001 |
HPCA | 8 |
| 2021 | Achieving on-Mobile Real-Time Super-Resolution with Neural Architecture and Pruning SearchabstractThough recent years have witnessed remarkable progress in single image super-resolution (SISR) tasks with the prosperous development of deep neural networks (DNNs), the deep learning methods are confronted with the computation and memory consumption issues in practice, especially for resource-limited platforms such as mobile devices. To overcome the challenge and facilitate the real-time deployment of SISR tasks on mobile, we combine neural architecture search with pruning search and propose an automatic search framework that derives sparse super-resolution (SR) models with high image quality while satisfying the real-time inference requirement. To decrease the search cost, we leverage the weight sharing strategy by introducing a supernet and decouple the search problem into three stages, including supernet construction, compiler-aware architecture and pruning search, and compiler-aware pruning ratio search. With the proposed framework, we are the first to achieve real-time SR inference (with only tens of milliseconds per frame) for implementing 720p resolution with competitive image quality (in terms of PSNR and SSIM) on mobile platforms (Samsung Galaxy S20). Zheng Zhan 0001, Yifan Gong 0004, Pu Zhao 0001, Geng Yuan, Wei Niu 0002, Yushu Wu, Tianyun Zhang, Malith Jayaweera, David R. Kaeli, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001 |
ICCV | 11 |
| 2021 | RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple PrecisionsabstractThis work proposes a novel Deep Neural Network (DNN) quantization framework, namely RMSMP, with a Row-wise Mixed-Scheme and Multi-Precision approach. Specifically, this is the first effort to assign mixed quantization schemes and multiple precisions within layers – among rows of the DNN weight matrix, for simplified operations in hardware inference, while preserving accuracy. Furthermore, this paper makes a different observation from the prior work that the quantization error does not necessarily exhibit the layer-wise sensitivity, and actually can be mitigated as long as a certain portion of the weights in every layer are in higher precisions. This observation enables layer-wise uniformality in the hardware implementation towards guaranteed inference acceleration, while still enjoying row-wise flexibility of mixed schemes and multiple precisions to boost accuracy. The candidates of schemes and precisions are derived practically and effectively with a highly hardware-informative strategy to reduce the problem search space.With the offline determined ratio of different quantization schemes and precisions for all the layers, the RMSMP quantization algorithm uses Hessian and variance based method to effectively assign schemes and precisions for each row. The proposed RMSMP is tested for the image classification and natural language processing (BERT) applications, and achieves the best accuracy performance among state-of-the-arts under the same equivalent precisions. The RMSMP is implemented on FPGA devices, achieving 3.65× speedup in the end-to-end inference time for ResNet-18 on ImageNet, comparing with the 4-bit Fixed-point baseline. Sung-En Chang, Yanyu Li, Mengshu Sun, Weiwen Jiang, Sijia Liu 0001, Yanzhi Wang 0001, Xue Lin 0001 |
ICCV | 7 |
| 2021 | Fast and Complete: Enabling Complete Neural Network Verification with Rapid and Massively Parallel Incomplete Verifiers
Kaidi Xu, Huan Zhang 0001, Shiqi Wang 0002, Suman Jana, Xue Lin 0001, Cho-Jui Hsieh |
ICLR | 6 |
| 2021 | Characteristic Examples: High-Robustness, Low-Transferability Fingerprinting of Neural NetworksabstractThis paper proposes Characteristic Examples for effectively fingerprinting deep neural networks, featuring high-robustness to the base model against model pruning as well as low-transferability to unassociated models. This is the first work taking both robustness and transferability into consideration for generating realistic fingerprints, whereas current methods lack practical assumptions and may incur large false positive rates. To achieve better trade-off between robustness and transferability, we propose three kinds of characteristic examples: vanilla C-examples, RC-examples, and LTRC-example, to derive fingerprints from the original base model. To fairly characterize the trade-off between robustness and transferability, we propose Uniqueness Score, a comprehensive metric that measures the difference between robustness and transferability, which also serves as an indicator to the false alarm problem. Extensive experiments demonstrate that the proposed characteristic examples can achieve superior performance when compared with existing fingerprinting methods. In particular, for VGG ImageNet models, using LTRC-examples gives 4X higher uniqueness score than the baseline method and does not incur any false positives. Siyue Wang, Xiao Wang 0028, Pu Zhao 0001, Xue Lin 0001 |
IJCAI | 5 |
| 2021 | Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness VerificationabstractBound propagation based incomplete neural network verifiers such as CROWN are very efficient and can significantly accelerate branch-and-bound (BaB) based complete verification of neural networks. However, bound propagation cannot fully handle the neuron split constraints introduced by BaB commonly handled by expensive linear programming (LP) solvers, leading to loose bounds and hurting verification efficiency. In this work, we develop $\beta$-CROWN, a new bound propagation based method that can fully encode neuron splits via optimizable parameters $\beta$ constructed from either primal or dual space. When jointly optimized in intermediate layers, $\beta$-CROWN generally produces better bounds than typical LP verifiers with neuron split constraints, while being as efficient and parallelizable as CROWN on GPUs. Applied to complete robustness verification benchmarks, $\beta$-CROWN with BaB is up to three orders of magnitude faster than LP-based BaB methods, and is notably faster than all existing approaches while producing lower timeout rates. By terminating BaB early, our method can also be used for efficient incomplete verification. We consistently achieve higher verified accuracy in many settings compared to powerful incomplete verifiers, including those based on convex barrier breaking techniques. Compared to the typically tightest but very costly semidefinite programming (SDP) based incomplete verifiers, we obtain higher verified accuracy with three orders of magnitudes less verification time. Our algorithm empowered the $\alpha,\!\beta$-CROWN (alpha-beta-CROWN) verifier, the winning tool in VNN-COMP 2021. Our code is available at http://PaperCode.cc/BetaCROWN. Shiqi Wang 0002, Huan Zhang 0001, Kaidi Xu, Xue Lin 0001, Suman Jana, Cho-Jui Hsieh, J. Zico Kolter |
NeurIPS | 4 |
| 2021 | MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the EdgeabstractRecently, a new trend of exploring sparsity for accelerating neural network training has emerged, embracing the paradigm of training on the edge. This paper proposes a novel Memory-Economic Sparse Training (MEST) framework targeting for accurate and fast execution on edge devices. The proposed MEST framework consists of enhancements by Elastic Mutation (EM) and Soft Memory Bound (&S) that ensure superior accuracy at high sparsity ratios. Different from the existing works for sparse training, this current work reveals the importance of sparsity schemes on the performance of sparse training in terms of accuracy as well as training speed on real edge devices. On top of that, the paper proposes to employ data efficiency for further acceleration of sparse training. Our results suggest that unforgettable examples can be identified in-situ even during the dynamic exploration of sparsity masks in the sparse training process, and therefore can be removed for further training speedup on edge devices. Comparing with state-of-the-art (SOTA) works on accuracy, our MEST increases Top-1 accuracy significantly on ImageNet when using the same unstructured sparsity scheme. Systematical evaluation on accuracy, training speed, and memory footprint are conducted, where the proposed MEST framework consistently outperforms representative SOTA works. A reviewer strongly against our work based on his false assumptions and misunderstandings. On top of the previous submission, we employ data efficiency for further acceleration of sparse training. And we explore the impact of model sparsity, sparsity schemes, and sparse training algorithms on the number of removable training examples. Our codes are publicly available at: https://github.com/boone891214/MEST. Geng Yuan, Wei Niu 0002, Zhengang Li 0001, Zhenglun Kong, Ning Liu 0007, Yifan Gong 0004, Zheng Zhan 0001, Chaoyang He 0001, Qing Jin, Siyue Wang, Minghai Qin, Bin Ren 0002, Yanzhi Wang 0001, Sijia Liu 0001, Xue Lin 0001 |
NeurIPS | 16 |
| 2021 | Work in Progress: Mobile or FPGA? A Comprehensive Evaluation on Energy Efficiency and a Unified Optimization FrameworkabstractEfficient deployment of Deep Neural Networks (DNNs) on edge devices (i.e., FPGAs and mobile platforms) is very challenging, especially under a recent witness of the increasing DNN model size and complexity. Although various optimization approaches have been proven to be effective in many DNNs on edge devices, most state-of-the-art work focuses on ad-hoc optimizations, and there lacks a thorough study to comprehensively reveal the potentials and constraints of different edge devices when considering different optimizations. In this paper, we qualitatively and quantitatively compare the energyefficiency of FPGA-based and mobile-based DNN executions, and provide detailed analysis. Geng Yuan, Peiyan Dong, Mengshu Sun, Wei Niu 0002, Zhengang Li 0001, Yuxuan Cai 0001, Jun Liu 0075, Weiwen Jiang, Xue Lin 0001, Bin Ren 0002, Xulong Tang, Yanzhi Wang 0001 |
RTAS | 9 |
| 2021 | Brief Industry Paper: Towards Real-Time 3D Object Detection for Autonomous Vehicles with Pruning SearchabstractIn autonomous driving, 3D object detection is es-sential as it provides basic knowledge about the environment. However, as deep learning based 3D detection methods are usually computation intensive, it is challenging to support realtime 3D object detection on edge-computing devices in selfdriving cars with limited computation and memory resources. To facilitate this, we propose a compiler-aware pruning search framework, to achieve real-time inference of 3D object detection on the resource-limited mobile devices. Specifically, a generator is applied to sample better pruning proposals in the search space based on current proposals with their performance, and an evaluator is adopted to evaluate the sampled pruning proposal performance. To accelerate the search, the evaluator employs Bayesian optimization with an ensemble of neural predictors. We demonstrate in experiments that for the first time, the pruning search framework can achieve real-time 3D object detection on mobile (Samsung Galaxy S20 phone) with state-of-the-art detection performance. Pu Zhao 0001, Wei Niu 0002, Geng Yuan, Yuxuan Cai 0001, Hsin-Hsuan Sung, Shaoshan Liu, Sijia Liu 0001, Xipeng Shen, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001 |
RTAS | 11 |
| 2021 | Dirty Road Can Attack: Security of Deep Learning based Automated Lane Centering under Physical-World Attack
Takami Sato, Junjie Shen 0001, Ningfei Wang, Yunhan Jia, Xue Lin 0001, Qi Alfred Chen |
USENIX Security Symposium | 5 |
| 2020 | PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesabstractModel compression techniques on Deep Neural Network (DNN) have been widely acknowledged as an effective way to achieve acceleration on a variety of platforms, and DNN weight pruning is a straightforward and effective method. There are currently two mainstreams of pruning methods representing two extremes of pruning regularity: non-structured, fine-grained pruning can achieve high sparsity and accuracy, but is not hardware friendly; structured, coarse-grained pruning exploits hardware-efficient structures in pruning, but suffers from accuracy drop when the pruning rate is high. In this paper, we introduce PCONV, comprising a new sparsity dimension, – fine-grained pruning patterns inside the coarse-grained structures. PCONV comprises two types of sparsities, Sparse Convolution Patterns (SCP) which is generated from intra-convolution kernel pruning and connectivity sparsity generated from inter-convolution kernel pruning. Essentially, SCP enhances accuracy due to its special vision properties, and connectivity sparsity increases pruning rate while maintaining balanced workload on filter computation. To deploy PCONV, we develop a novel compiler-assisted DNN inference framework and execute PCONV models in real-time without accuracy compromise, which cannot be achieved in prior work. Our experimental results show that, PCONV outperforms three state-of-art end-to-end DNN frameworks, TensorFlow-Lite, TVM, and Alibaba Mobile Neural Network with speedup up to 39.2 ×, 11.4 ×, and 6.3 ×, respectively, with no accuracy loss. Mobile devices can achieve real-time inference on large-scale DNNs. Fu-Ming Guo, Wei Niu 0002, Xue Lin 0001, Jian Tang 0008, Kaisheng Ma, Bin Ren 0002, Yanzhi Wang 0001 |
AAAI | 4 |
| 2020 | Towards Certificated Model Robustness Against Weight PerturbationsabstractThis work studies the sensitivity of neural networks to weight perturbations, firstly corresponding to a newly developed threat model that perturbs the neural network parameters. We propose an efficient approach to compute a certified robustness bound of weight perturbations, within which neural networks will not make erroneous outputs as desired by the adversary. In addition, we identify a useful connection between our developed certification method and the problem of weight quantization, a popular model compression technique in deep neural networks (DNNs) and a ‘must-try’ step in the design of DNN inference engines on resource constrained computing platforms, such as mobiles, FPGA, and ASIC. Specifically, we study the problem of weight quantization – weight perturbations in the non-adversarial setting – through the lens of certificated robustness, and we demonstrate significant improvements on the generalization ability of quantized networks through our robustness-aware quantization scheme. Tsui-Wei Weng, Pu Zhao 0001, Sijia Liu 0001, Xue Lin 0001, Luca Daniel |
AAAI | 5 |
| 2020 | Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient DescentabstractDespite the great achievements of the modern deep neural networks (DNNs), the vulnerability/robustness of state-of-the-art DNNs raises security concerns in many application domains requiring high reliability. Various adversarial attacks are proposed to sabotage the learning performance of DNN models. Among those, the black-box adversarial attack methods have received special attentions owing to their practicality and simplicity. Black-box attacks usually prefer less queries in order to maintain stealthy and low costs. However, most of the current black-box attack methods adopt the first-order gradient descent method, which may come with certain deficiencies such as relatively slow convergence and high sensitivity to hyper-parameter settings. In this paper, we propose a zeroth-order natural gradient descent (ZO-NGD) method to design the adversarial attacks, which incorporates the zeroth-order gradient estimation technique catering to the black-box attack scenario and the second-order natural gradient descent to achieve higher query efficiency. The empirical evaluations on image classification datasets demonstrate that ZO-NGD can obtain significantly lower model query complexities compared with state-of-the-art attack methods. Pu Zhao 0001, Siyue Wang, Xue Lin 0001 |
AAAI | 4 |
| 2020 | PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningabstractWith the emergence of a spectrum of high-end mobile devices, many applications that formerly required desktop-level computation capability are being transferred to these devices. However, executing Deep Neural Networks (DNNs) inference is still challenging considering the high computation and storage demands, specifically, if real-time performance with high accuracy is needed. Weight pruning of DNNs is proposed, but existing schemes represent two extremes in the design space: non-structured pruning is fine-grained, accurate, but not hardware friendly; structured pruning is coarse-grained, hardware-efficient, but with higher accuracy loss. Wei Niu 0002, Sheng Lin 0001, Xuehai Qian, Xue Lin 0001, Yanzhi Wang 0001, Bin Ren 0002 |
ASPLOS | 6 |
| 2020 | RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech RecognitionabstractRecurrent neural networks (RNNs) based automatic speech recognition has nowadays become promising and important on mobile devices such as smart phones. However, previous RNN compression techniques either suffer from hardware performance overhead due to irregularity or significant accuracy loss due to the preserved regularity for hardware friendliness. In this work, we propose RTMobile that leverages both a novel block-based pruning approach and compiler optimizations to accelerate RNN inference on mobile devices. Our proposed RTMobile is the first work that can achieve real-time RNN inference on mobile platforms. Experimental results demonstrate that RTMobile can significantly outperform existing RNN hardware acceleration methods in terms of both inference accuracy and time. Compared with prior work on FPGA, RTMobile using Adreno 640 embedded GPU on GRU can improve the energy-efficiency by 40× while maintaining the same inference time. Peiyan Dong, Siyue Wang, Wei Niu 0002, Chengming Zhang 0006, Sheng Lin 0001, Zhengang Li 0001, Yifan Gong 0004, Bin Ren 0002, Xue Lin 0001, Dingwen Tao |
DAC | 9 |
| 2020 | 3D CNN Acceleration on FPGA using Hardware-Aware PruningabstractThere have been many recent attempts to extend the successes of convolutional neural networks (CNNs) from 2-dimensional (2D) image classification to 3-dimensional (3D) video recognition by exploring 3D CNNs. Considering the emerging growth of mobile or Internet of Things (IoT) market, it is essential to investigate the deployment of 3D CNNs on edge devices. Previous works have implemented standard 3D CNNs (C3D) on hardware platforms, however, they have not exploited model compression for acceleration of inference. This work proposes a hardware-aware pruning approach that can fully adapt to the loop tiling technique of FPGA design and is applied onto a novel 3D network called R(2+1)D. Leveraging the powerful ADMM, the proposed pruning method achieves simultaneous high accuracy and significant acceleration of computation on FPGA. With layer-wise pruning rates up to 10× and negligible accuracy loss, the pruned model is implemented on a Xilinx ZCU102 FPGA board, where the pruned model achieves 2.6× speedup compared with the unpruned version, and 2.3× speedup and 2.3× power efficiency improvement compared with state-of-the-art FPGA implementation of C3D. Mengshu Sun, Pu Zhao 0001, Mehmet Güngör, Massoud Pedram, Miriam Leeser, Xue Lin 0001 |
DAC | 6 |
| 2020 | Adversarial T-Shirt! Evading Person Detectors in a Physical World
Kaidi Xu, Gaoyuan Zhang, Sijia Liu 0001, Quanfu Fan, Mengshu Sun, Hongge Chen, Yanzhi Wang 0001, Xue Lin 0001 |
ECCV (5) | 9 |
| 2020 | A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration FrameworkabstractWeight pruning of deep neural networks (DNNs) has been proposed to satisfy the limited storage and computing capability of mobile edge devices. However, previous pruning methods mainly focus on reducing the model size and/or improving performance without considering the privacy of user data. To mitigate this concern, we propose a privacy-preserving-oriented pruning and mobile acceleration framework that does not require the private training dataset. At the algorithm level of the proposed framework, a systematic weight pruning technique based on the alternating direction method of multipliers (ADMM) is designed to iteratively solve the pattern-based pruning problem for each layer with randomly generated synthetic data. In addition, corresponding optimizations at the compiler level are leveraged for inference accelerations on devices. With the proposed framework, users could avoid the time-consuming pruning process for non-experts and directly benefit from compressed models. Experimental results show that the proposed framework outperforms three state-of-art end-to-end DNN frameworks, i.e., TensorFlow-Lite, TVM, and MNN, with speedup up to 4.2×, 2.5×, and 2.0×, respectively, with almost no accuracy loss, while preserving data privacy. Yifan Gong 0004, Zheng Zhan 0001, Zhengang Li 0001, Wei Niu 0002, Wenhao Wang 0001, Bin Ren 0002, Caiwen Ding, Xue Lin 0001, Xiaolin Xu 0001, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 9 |
| 2020 | AdvMS: A Multi-Source Multi-Cost Defense Against Adversarial AttacksabstractDesigning effective defense against adversarial attacks is a crucial topic as deep neural networks have been proliferated rapidly in many security-critical domains such as malware detection and self-driving cars. Conventional defense methods, although shown to be promising, are largely limited by their single-source single-cost nature: The robustness promotion tends to plateau when the defenses are made increasingly stronger while the cost tends to amplify. In this paper, we study principles of designing multi-source and multi-cost schemes where defense performance is boosted from multiple defending components. Based on this motivation, we propose a multi-source and multi-cost defense scheme, Adversarially Trained Model Switching (AdvMS), that inherits advantages from two leading schemes: adversarial training and random model switching. We show that the multi-source nature of AdvMS mitigates the performance plateauing issue and the multi-cost nature enables improving robustness at a flexible and adjustable combination of costs over different factors which can better suit specific restrictions and needs in practice. Xiao Wang 0028, Siyue Wang, Xue Lin 0001, Sang (Peter) Chin |
ICASSP | 4 |
| 2020 | Towards an Efficient and General Framework of Robust Training for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have made significant advances on several fundamental inference tasks. As a result, there is a surge of interest in using these models for making potentially important decisions in high-regret applications. However, despite GNNs' impressive performance, it has been observed that carefully crafted perturbations on graph structures (or nodes attributes) lead them to make wrong predictions. Presence of these adversarial examples raises serious security concerns. Most of the existing robust GNN design/training methods are only applicable to white-box settings where model parameters are known and gradient based methods can be used by performing convex relaxation of the discrete graph domain. More importantly, these methods are not efficient and scalable which make them infeasible in time sensitive tasks and massive graph datasets. To overcome these limitations, we propose a general framework which leverages the greedy search algorithms and zeroth-order methods to obtain robust GNNs in a generic and an efficient manner. On several applications, we show that the proposed techniques are significantly less computationally expensive and, in some cases, more robust than the state-of-the-art methods making them suitable to large-scale problems which were out of the reach of traditional robust training methods. Kaidi Xu, Sijia Liu 0001, Mengshu Sun, Caiwen Ding, Bhavya Kailkhura, Xue Lin 0001 |
ICASSP | 7 |
| 2020 | New Passive and Active Attacks on Deep Neural Networks in Medical ApplicationsabstractSecurity of deep neural network (DNN) inference engines, i.e., trained DNN models on various platforms, has become one of the biggest challenges in deploying artificial intelligence in domains where privacy, safety, and reliability are of paramount importance, such as in medical applications. In addition to classic software attacks such as model inversion and evasion attacks, recently a new attack surface---implementation attacks which include both passive side-channel attacks and active fault injection and adversarial attacks---is arising, targeting implementation peculiarities of DNN to breach their confidentiality and integrity. This paper presents several novel passive and active attacks on DNN we have developed and tested over medical datasets. Our new attacks reveal a largely under-explored attack surface of DNN inference engines. Insights gained during attack exploration will provide valuable guidance for effectively protecting DNN execution against reverse-engineering and integrity violations. Cheng Gongye, Hongjia Li 0003, Majid Sabbagh, Geng Yuan, Xue Lin 0001, Thomas Wahl, Yunsi Fei |
ICCAD | 6 |
| 2020 | Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness
Pu Zhao 0001, Karthikeyan Natesan Ramamurthy, Xue Lin 0001 |
ICLR | 5 |
| 2020 | Towards Real-Time DNN Inference on Mobile Platforms with Model Pruning and Compiler OptimizationabstractHigh-end mobile platforms rapidly serve as primary computing devices for a wide range of Deep Neural Network (DNN) applications. However, the constrained computation and storage resources on these devices still pose significant challenges for real-time DNN inference executions. To address this problem, we propose a set of hardware-friendly structured model pruning and compiler optimization techniques to accelerate DNN executions on mobile devices. This demo shows that these optimizations can enable real-time mobile execution of multiple DNN applications, including style transfer, DNN coloring and super resolution. Wei Niu 0002, Pu Zhao 0001, Zheng Zhan 0001, Xue Lin 0001, Yanzhi Wang 0001, Bin Ren 0002 |
IJCAI | 4 |
| 2020 | Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondabstractLinear relaxation based perturbation analysis (LiRPA) for neural networks, which computes provable linear bounds of output neurons given a certain amount of input perturbation, has become a core component in robustness verification and certified defense. The majority of LiRPA-based methods focus on simple feed-forward networks and need particular manual derivations and implementations when extended to other architectures. In this paper, we develop an automatic framework to enable perturbation analysis on any neural network structures, by generalizing existing LiRPA algorithms such as CROWN to operate on general computational graphs. The flexibility, differentiability and ease of use of our framework allow us to obtain state-of-the-art results on LiRPA based certified defense on fairly complicated networks like DenseNet, ResNeXt and Transformer that are not supported by prior works. Our framework also enables loss fusion, a technique that significantly reduces the computational complexity of LiRPA for certified defense. For the first time, we demonstrate LiRPA based certified defense on Tiny ImageNet and Downscaled ImageNet where previous approaches cannot scale to due to the relatively large number of classes. Our work also yields an open-source library for the community to apply LiRPA to areas beyond certified defense without much LiRPA expertise, e.g., we create a neural network with a provably flat optimization landscape by applying LiRPA to network parameters. Our open source library is available at https://github.com/KaidiXu/auto_LiRPA. Kaidi Xu, Zhouxing Shi, Huan Zhang 0001, Kai-Wei Chang 0001, Minlie Huang, Bhavya Kailkhura, Xue Lin 0001, Cho-Jui Hsieh |
NeurIPS | 8 |
| 2020 | Exploring GPU acceleration of Deep Neural Networks using Block Circulant Matrices
Shi Dong 0002, Pu Zhao 0001, Xue Lin 0001, David R. Kaeli |
Parallel Comput. | 3 |
| 2019 | Universal Approximation Property and Equivalence of Stochastic Computing-Based Neural Networks and Binary Neural NetworksabstractLarge-scale deep neural networks are both memory and computation-intensive, thereby posing stringent requirements on the computing platforms. Hardware accelerations of deep neural networks have been extensively investigated. Specific forms of binary neural networks (BNNs) and stochastic computing-based neural networks (SCNNs) are particularly appealing to hardware implementations since they can be implemented almost entirely with binary operations. Despite the obvious advantages in hardware implementation, these approximate computing techniques are questioned by researchers in terms of accuracy and universal applicability. Also it is important to understand the relative pros and cons of SCNNs and BNNs in theory and in actual hardware implementations. In order to address these concerns, in this paper we prove that the “ideal” SCNNs and BNNs satisfy the universal approximation property with probability 1 (due to the stochastic behavior), which is a new angle from the original approximation property. The proof is conducted by first proving the property for SCNNs from the strong law of large numbers, and then using SCNNs as a “bridge” to prove for BNNs. Besides the universal approximation property, we also derive an appropriate bound for bit length M in order to provide insights for the actual neural network implementations. Based on the universal approximation property, we further prove that SCNNs and BNNs exhibit the same energy complexity. In other words, they have the same asymptotic energy consumption with the growth of network size. We also provide a detailed analysis of the pros and cons of SCNNs and BNNs for hardware implementations and conclude that SCNNs are more suitable. Yanzhi Wang 0001, Zheng Zhan 0001, Liang Zhao 0002, Jian Tang 0008, Siyue Wang, Bo Yuan 0001, Wujie Wen, Xue Lin 0001 |
AAAI | 9 |
| 2019 | ADMM attack: an enhanced adversarial attack for deep neural networks with undetectable distortionsabstractMany recent studies demonstrate that state-of-the-art Deep neural networks (DNNs) might be easily fooled by adversarial examples, generated by adding carefully crafted and visually imperceptible distortions onto original legal inputs through adversarial attacks. Adversarial examples can lead the DNN to misclassify them as any target labels. In the literature, various methods are proposed to minimize the different lp norms of the distortion. However, there lacks a versatile framework for all types of adversarial attacks. To achieve a better understanding for the security properties of DNNs, we propose a general framework for constructing adversarial examples by leveraging Alternating Direction Method of Multipliers (ADMM) to split the optimization approach for effective minimization of various lp norms of the distortion, including l0, l1, l2, and l∞ norms. Thus, the proposed general framework unifies the methods of crafting l0, l1, l2, and l∞ attacks. The experimental results demonstrate that the proposed ADMM attacks achieve both the high attack success rate and the minimal distortion for the misclassification compared with state-of-the-art attack methods. Pu Zhao 0001, Kaidi Xu, Sijia Liu 0001, Yanzhi Wang 0001, Xue Lin 0001 |
ASP-DAC | 5 |
| 2019 | ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Methods of MultipliersabstractModel compression is an important technique to facilitate efficient embedded and hardware implementations of deep neural networks (DNNs), a number of prior works are dedicated to model compression techniques. The target is to simultaneously reduce the model storage size and accelerate the computation, with minor effect on accuracy. Two important categories of DNN model compression techniques are weight pruning and weight quantization. The former leverages the redundancy in the number of weights, whereas the latter leverages the redundancy in bit representation of weights. These two sources of redundancy can be combined, thereby leading to a higher degree of DNN model compression. However, a systematic framework of joint weight pruning and quantization of DNNs is lacking, thereby limiting the available model compression ratio. Moreover, the computation reduction, energy efficiency improvement, and hardware performance overhead need to be accounted besides simply model size reduction, and the hardware performance overhead resulted from weight pruning method needs to be taken into consideration. To address these limitations, we present ADMM-NN, the first algorithm-hardware co-optimization framework of DNNs using Alternating Direction Method of Multipliers (ADMM), a powerful technique to solve non-convex optimization problems with possibly combinatorial constraints. The first part of ADMM-NN is a systematic, joint framework of DNN weight pruning and quantization using ADMM. It can be understood as a smart regularization technique with regularization target dynamically updated in each ADMM iteration, thereby resulting in higher performance in model compression than the state-of-the-art. The second part is hardware-aware DNN optimizations to facilitate hardware-level implementations. We perform ADMM-based weight pruning and quantization considering (i) the computation reduction and energy efficiency improvement, and (ii) the hardware performance overhead due to irregular sparsity. The first requirement prioritizes the convolutional layer compression over fully-connected layers, while the latter requires a concept of the break-even pruning ratio, defined as the minimum pruning ratio of a specific layer that results in no hardware performance degradation. Without accuracy loss, ADMM-NN achieves 85× and 24× pruning on LeNet-5 and AlexNet models, respectively, --- significantly higher than the state-of-the-art. The improvements become more significant when focusing on computation reduction. Combining weight pruning and quantization, we achieve 1,910× and 231× reductions in overall model size on these two benchmarks, when focusing on data storage. Highly promising results are also observed on other representative DNNs such as VGGNet and ResNet-50. We release codes and models at https://github.com/yeshaokai/admm-nn. Ao Ren, Tianyun Zhang, Shaokai Ye, Wenyao Xu, Xuehai Qian, Xue Lin 0001, Yanzhi Wang 0001 |
ASPLOS | 7 |
| 2019 | Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial ExamplesabstractImage compression-based approaches for defending against the adversarial-example attacks, which threaten the safety use of deep neural networks (DNN), have been investigated recently. However, prior works mainly rely on directly tuning parameters like compression rate, to blindly reduce image features, thereby lacking guarantee on both defense efficiency (i.e. accuracy of polluted images) and classification accuracy of benign images, after applying defense methods. To overcome these limitations, we propose a JPEG-based defensive compression framework, namely “feature distillation”, to effectively rectify adversarial examples without impacting classification accuracy on benign data. Our framework significantly escalates the defense efficiency with marginal accuracy reduction using a twostep method: First, we maximize malicious features filtering of adversarial input perturbations by developing defensive quantization in frequency domain of JPEG compression or decompression, guided by a semi-analytical method; Second, we suppress the distortions of benign features to restore classification accuracy through a DNN-oriented quantization refine process. Our experimental results show that proposed “feature distillation” can significantly surpass the latest input-transformation based mitigations such as Quilting and TV Minimization in three aspects, including defense efficiency (improve classification accuracy from ∼ 20% to ∼ 90% on adversarial examples), accuracy of benign images after defense (≤ 1% accuracy degradation), and processing time per image (∼ 259× Speedup). Moreover, our solution also can provide the best defense efficiency (∼ 60% accuracy) against the latest BPDA attack with least accuracy reduction (∼ 1%) on benign images among all other input-transformation based defense methods. Zihao Liu 0015, Qi Liu 0017, Tao Liu 0023, Nuo Xu 0013, Xue Lin 0001, Yanzhi Wang 0001, Wujie Wen |
CVPR | 5 |
| 2019 | Fault Sneaking Attack: a Stealthy Framework for Misleading Deep Neural NetworksabstractDespite the great achievements of deep neural networks (DNNs), the vulnerability of state-of-the-art DNNs raises security concerns of DNNs in many application domains requiring high reliability. We propose the fault sneaking attack on DNNs, where the adversary aims to misclassify certain input images into any target labels by modifying the DNN parameters. We apply ADMM (alternating direction method of multipliers) for solving the optimization problem of the fault sneaking attack with two constraints: 1) the classification of the other images should be unchanged and 2) the parameter modifications should be minimized. Specifically, the first constraint requires us not only to inject designated faults (misclassifications), but also to hide the faults for stealthy or sneaking considerations by maintaining model accuracy. The second constraint requires us to minimize the parameter modifications (using ℓ0 norm to measure the number of modifications and ℓ2 norm to measure the magnitude of modifications). Comprehensive experimental evaluation demonstrates that the proposed framework can inject multiple sneaking faults without losing the overall test accuracy performance. Pu Zhao 0001, Siyue Wang, Cheng Gongye, Yanzhi Wang 0001, Yunsi Fei, Xue Lin 0001 |
DAC | 6 |
| 2019 | ADMM-based Weight Pruning for Real-Time Deep Learning Acceleration on Mobile DevicesabstractDeep learning solutions are being increasingly deployed in mobile applications, at least for the inference phase. Due to the large model size and computational requirements, model compression for deep neural networks (DNNs) becomes necessary, especially considering the real-time requirement in embedded systems. In this paper, we extend the prior work on systematic DNN weight pruning using ADMM (Alternating Direction Method of Multipliers). We integrate ADMM regularization with masked mapping/retraining, thereby guaranteeing solution feasibility and providing high solution quality. Besides superior performance on representative DNN benchmarks (e.g., AlexNet, ResNet), we focus on two new applications facial emotion detection and eye tracking, and develop a top-down framework of DNN training, model compression, and acceleration in mobile devices. Experimental results show that with negligible accuracy degradation, the proposed method can achieve significant storage/memory reduction and speedup in mobile devices. Hongjia Li 0003, Ning Liu 0007, Sheng Lin 0001, Shaokai Ye, Tianyun Zhang, Xue Lin 0001, Wenyao Xu, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2019 | HSIM-DNN: Hardware Simulator for Computation-, Storage- and Power-Efficient Deep Neural NetworksabstractDeep learning that utilizes large-scale deep neural networks (DNNs) is effective in automatic high-level feature extraction but also computation and memory intensive. Constructing DNNs using block-circulant matrices can simultaneously achieve hardware acceleration and model compression while maintaining high accuracy. This paper proposes HSIM-DNN, an accurate hardware simulator on the C++ platform, to simulate the exact behavior of DNN hardware implementations and thereby facilitate the block-circulant matrix-based design of DNN training and inference procedures in hardware. Real FPGA implementations validate the simulator with various circulant block sizes and data bit lengths taking into account accuracy, compression ratio and power consumption, which provides excellent insights for hardware design. Mengshu Sun, Pu Zhao 0001, Yanzhi Wang 0001, Naehyuck Chang, Xue Lin 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2019 | E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAsabstractRecurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The two major types are Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) networks. It is a challenging task to have real-time, efficient, and accurate hardware RNN implementations because of the high sensitivity to imprecision accumulation and the requirement of special activation function implementations. Recently two works have focused on FPGA implementation of inference phase of LSTM RNNs with model compression. First, ESE uses a weight pruning based compressed RNN model but suffers from irregular network structure after pruning. The second work C-LSTM mitigates the irregular network limitation by incorporating block-circulant matrices for weight matrix representation in RNNs, thereby achieving simultaneous model compression and acceleration. A key limitation of the prior works is the lack of a systematic design optimization framework of RNN model and hardware implementations, especially when the block size (or compression ratio) should be jointly optimized with RNN type, layer size, etc. In this paper, we adopt the block-circulant matrixbased framework, and present the Efficient RNN (E-RNN) framework for FPGA implementations of the Automatic Speech Recognition (ASR) application. The overall goal is to improve performance/energy efficiency under accuracy requirement. We use the alternating direction method of multipliers (ADMM) technique for more accurate block-circulant training, and present two design explorations providing guidance on block size and reducing RNN training trials. Based on the two observations, we decompose E-RNN in two phases: Phase I on determining RNN model to reduce computation and storage subject to accuracy requirement, and Phase II on hardware implementations given RNN model, including processing element design/optimization, quantization, activation implementation, etc. 1 Experimental results on actual FPGA deployments show that E-RNN achieves a maximum energy efficiency improvement of 37.4× compared with ESE, and more than 2× compared with C-LSTM, under the same accuracy. Zhe Li 0001, Caiwen Ding, Siyue Wang, Wujie Wen, Youwei Zhuo, Qinru Qiu, Wenyao Xu, Xue Lin 0001, Xuehai Qian, Yanzhi Wang 0001 |
HPCA | 9 |
| 2019 | Adversarial Robustness vs. Model Compression, or Both?abstractIt is well known that deep neural networks (DNNs) are vulnerable to adversarial attacks, which are implemented by adding crafted perturbations onto benign examples. Min-max robust optimization based adversarial training can provide a notion of security against adversarial attacks. However, adversarial robustness requires a significantly larger capacity of the network than that for the natural training with only benign examples. This paper proposes a framework of concurrent adversarial training and weight pruning that enables model compression while still preserving the adversarial robustness and essentially tackles the dilemma of adversarial training. Furthermore, this work studies two hypotheses about weight pruning in the conventional setting and finds that weight pruning is essential for reducing the network model size in the adversarial setting; training a small model from scratch even with inherited initialization from the large model cannot achieve neither adversarial robustness nor high standard accuracy. Code is available at https://github.com/yeshaokai/Robustness-Aware-Pruning-ADMM. Shaokai Ye, Xue Lin 0001, Kaidi Xu, Sijia Liu 0001, Hao Cheng 0015, Jan-Henrik Lambrechts, Huan Zhang 0001, Aojun Zhou, Kaisheng Ma, Yanzhi Wang 0001 |
ICCV | 2 |
| 2019 | On the Design of Black-Box Adversarial Examples by Leveraging Gradient-Free Optimization and Operator Splitting MethodabstractRobust machine learning is currently one of the most prominent topics which could potentially help shaping a future of advanced AI platforms that not only perform well in average cases but also in worst cases or adverse situations. Despite the long-term vision, however, existing studies on black-box adversarial attacks are still restricted to very specific settings of threat models (e.g., single distortion metric and restrictive assumption on target model's feedback to queries) and/or suffer from prohibitively high query complexity. To push for further advances in this field, we introduce a general framework based on an operator splitting method, the alternating direction method of multipliers (ADMM) to devise efficient, robust black-box attacks that work with various distortion metrics and feedback settings without incurring high query complexity. Due to the black-box nature of the threat model, the proposed ADMM solution framework is integrated with zeroth-order (ZO) optimization and Bayesian optimization (BO), and thus is applicable to the gradient-free regime. This results in two new black-box adversarial attack generation methods, ZO-ADMM and BO-ADMM. Our empirical evaluations on image classification datasets show that our proposed approaches have much lower function query complexities compared to state-of-the-art attack methods, but achieve very competitive attack success rates. Pu Zhao 0001, Sijia Liu 0001, Nghia Hoang, Kaidi Xu, Bhavya Kailkhura, Xue Lin 0001 |
ICCV | 7 |
| 2019 | Structured Adversarial Attack: Towards General Implementation and Better Interpretability
Kaidi Xu, Sijia Liu 0001, Pu Zhao 0001, Huan Zhang 0001, Quanfu Fan, Deniz Erdogmus, Yanzhi Wang 0001, Xue Lin 0001 |
ICLR (Poster) | 9 |
| 2019 | Protecting Neural Networks with Hierarchical Random Switching: Towards Better Robustness-Accuracy Trade-off for Stochastic DefensesabstractDespite achieving remarkable success in various domains, recent studies have uncovered the vulnerability of deep neural networks to adversarial perturbations, creating concerns on model generalizability and new threats such as prediction-evasive misclassification or stealthy reprogramming. Among different defense proposals, stochastic network defenses such as random neuron activation pruning or random perturbation to layer inputs are shown to be promising for attack mitigation. However, one critical drawback of current defenses is that the robustness enhancement is at the cost of noticeable performance degradation on legitimate data, e.g., large drop in test accuracy.This paper is motivated by pursuing for a better trade-off between adversarial robustness and test accuracy for stochastic network defenses. We propose Defense Efficiency Score (DES), a comprehensive metric that measures the gain in unsuccessful attack attempts at the cost of drop in test accuracy of any defense. To achieve a better DES, we propose hierarchical random switching (HRS), which protects neural networks through a novel randomization scheme. A HRS-protected model contains several blocks of randomly switching channels to prevent adversaries from exploiting fixed model structures and parameters for their malicious purposes. Extensive experiments show that HRS is superior in defending against state-of-the-art white-box and adaptive adversarial misclassification attacks. We also demonstrate the effectiveness of HRS in defending adversarial reprogramming, which is the first defense against adversarial programs. Moreover, in most settings the average DES of HRS is at least 5X higher than current stochastic network defenses, validating its significantly improved robustness-accuracy trade-off. Xiao Wang 0028, Siyue Wang, Yanzhi Wang 0001, Brian Kulis, Xue Lin 0001, Sang (Peter) Chin |
IJCAI | 6 |
| 2019 | Topology Attack and Defense for Graph Neural Networks: An Optimization PerspectiveabstractGraph neural networks (GNNs) which apply the deep neural networks to graph data have achieved significant performance for the task of semi-supervised node classification. However, only few work has addressed the adversarial robustness of GNNs. In this paper, we first present a novel gradient-based attack method that facilitates the difficulty of tackling discrete graph data. When comparing to current adversarial attacks on GNNs, the results show that by only perturbing a small number of edge perturbations, including addition and deletion, our optimization-based attack can lead to a noticeable decrease in classification performance. Moreover, leveraging our gradient-based attack, we propose the first optimization-based adversarial training for GNNs. Our method yields higher robustness against both different gradient based and greedy attack methods without sacrifice classification accuracy on original graph. Kaidi Xu, Hongge Chen, Sijia Liu 0001, Tsui-Wei Weng, Mingyi Hong 0001, Xue Lin 0001 |
IJCAI | 7 |
| 2019 | ZO-AdaMM: Zeroth-Order Adaptive Momentum Method for Black-Box OptimizationabstractThe adaptive momentum method (AdaMM), which uses past gradients to update descent directions and learning rates simultaneously, has become one of the most popular first-order optimization methods for solving machine learning problems. However, AdaMM is not suited for solving black-box optimization problems, where explicit gradient forms are difficult or infeasible to obtain. In this paper, we propose a zeroth-order AdaMM (ZO-AdaMM) algorithm, that generalizes AdaMM to the gradient-free regime. We show that the convergence rate of ZO-AdaMM for both convex and nonconvex optimization is roughly a factor of $O(\sqrt{d})$ worse than that of the first-order AdaMM algorithm, where $d$ is problem size. In particular, we provide a deep understanding on why Mahalanobis distance matters in convergence of ZO-AdaMM and other AdaMM-type methods. As a byproduct, our analysis makes the first step toward understanding adaptive learning rate methods for nonconvex constrained optimization.Furthermore, we demonstrate two applications, designing per-image and universal adversarial attacks from black-box neural networks, respectively. We perform extensive experiments on ImageNet and empirically show that ZO-AdaMM converges much faster to a solution of high accuracy compared with $6$ state-of-the-art ZO optimization methods. Xiangyi Chen, Sijia Liu 0001, Kaidi Xu, Xingguo Li, Xue Lin 0001, Mingyi Hong 0001, David D. Cox |
NeurIPS | 5 |
| 2019 | Reduced-Complexity Deep Neural Networks Design Using Multi-Level CompressionabstractDeep Neural Network has achieved great success in many fields. However, many DNN models are both deep and large thereby causing high storage and energy consumption during the training and inference phases. This paper proposes multi-level compression framework. By utilizing cross-layer parameter-reducing techniques ranging from structure compression to weight compression to representation compression, the proposed compression strategy can enable order-of-magnitude reduction in network size for both training and inference with negligible accuracy loss, thereby leading to very high-efficiency and high-accuracy DNN models. Experiments show that the proposed strategy can achieve around 1.8K compression ratio in terms of dense matrices and around 30x for the overall model. Siyu Liao, Yi Xie 0001, Xue Lin 0001, Yanzhi Wang 0001, Min Zhang 0005, Bo Yuan 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2018 | Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization FrameworkabstractHardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural networks (DNNs). An algorithm-hardware co-optimization framework is developed, which is applicable to different DNN types, sizes, and application scenarios. The algorithm part adopts the general block-circulant matrices to achieve a fine-grained tradeoff of accuracy and compression ratio. It applies to both fully-connected and convolutional layers and contains a mathematically rigorous proof of the effectiveness of the method. The proposed algorithm reduces computational complexity per layer from O(n2) to O(n log n) and storage complexity from O(n2) to O(n), both for training and inference. The hardware part consists of highly efficient Field Programmable Gate Array (FPGA)-based implementations using effective reconfiguration, batch processing, deep pipelining, resource re-using, and hierarchical control. Experimental results demonstrate that the proposed framework achieves at least 152X speedup and 71X energy efficiency gain compared with IBM TrueNorth processor under the same test accuracy. It achieves at least 31X energy efficiency gain compared with the reference FPGA-based work. Yanzhi Wang 0001, Caiwen Ding, Zhe Li 0001, Geng Yuan, Siyu Liao, Bo Yuan 0001, Xuehai Qian, Jian Tang 0008, Qinru Qiu, Xue Lin 0001 |
AAAI | 11 |
| 2018 | A deep reinforcement learning framework for optimizing fuel economy of hybrid electric vehiclesabstractHybrid electric vehicles employ a hybrid propulsion system to combine the energy efficiency of electric motor and a long driving range of internal combustion engine, thereby achieving a higher fuel economy as well as convenience compared with conventional ICE vehicles. However, the relatively complicated powertrain structures of HEVs necessitate an effective power management policy to determine the power split between ICE and EM. In this work, we propose a deep reinforcement learning framework of the HEV power management with the aim of improving fuel economy. The DRL technique is comprised of an offline deep neural network construction phase and an online deep Q-learning phase. Unlike traditional reinforcement learning, DRL presents the capability of handling the high dimensional state and action space in the actual decision-making process, making it suitable for the HEV power management problem. Enabled by the DRL technique, the derived HEV power management policy is close to optimal, fully model-free, and independent of a prior knowledge of driving cycles. Simulation results based on actual vehicle setup over real-world and testing driving cycles demonstrate the effectiveness of the proposed framework on optimizing HEV fuel economy. Pu Zhao 0001, Yanzhi Wang 0001, Naehyuck Chang, Qi Zhu 0002, Xue Lin 0001 |
ASP-DAC | 5 |
| 2018 | Prediction-based fast thermoelectric generator reconfiguration for energy harvesting from vehicle radiatorsabstractThermoelectric generation (TEG) has increasingly drawn attention for being environmentally friendly. A few researches have focused on improving TEG efficiency at system level on vehicle radiators. The most recent reconfiguration algorithm shows improvement on performance but suffers from major drawback on computational time and energy overhead, and non-scalability in terms of array size and processing frequency. In this paper, we propose a novel TEG array reconfiguration algorithm that determines near-optimal configuration with an acceptable computational time. More precisely, with O(N) time complexity, our prediction-based fast TEG reconfiguration algorithm enables all modules to work at or near their maximum power points (MPP). Additionally, we incorporate prediction methods to further reduce the runtime and switching overhead during the reconfiguration process. Experimental results present 30% performance improvement, almost 100 χ reduction on switching overhead and 13 χ enhancement on computational speed compared to the baseline and prior work. The scalability of our algorithm makes it applicable to larger scale systems such as industrial boilers and heat exchangers. Feiyang Kang, Caiwen Ding, Ji Li 0006, Donkyu Baek, Shahin Nazarian, Xue Lin 0001, Paul Bogdan, Naehyuck Chang |
DATE | 8 |
| 2018 | Defensive dropout for hardening deep neural networks under adversarial attacks
Siyue Wang, Xiao Wang 0028, Pu Zhao 0001, Wujie Wen, David R. Kaeli, Sang (Peter) Chin, Xue Lin 0001 |
ICCAD | 7 |
| 2018 | An ADMM-Based Universal Framework for Adversarial Attacks on Deep Neural NetworksabstractDeep neural networks (DNNs) are known vulnerable to adversarial attacks. That is, adversarial examples, obtained by adding delicately crafted distortions onto original legal inputs, can mislead a DNN to classify them as any target labels. In a successful adversarial attack, the targeted mis-classification should be achieved with the minimal distortion added. In the literature, the added distortions are usually measured by $L_0$, $L_1$, $L_2$, and $L_\infty $ norms, namely, L_0, L_1, L_2, and L_∞ attacks, respectively. However, there lacks a versatile framework for all types of adversarial attacks. This work for the first time unifies the methods of generating adversarial examples by leveraging ADMM (Alternating Direction Method of Multipliers), an operator splitting optimization approach, such that $L_0$, $L_1$, $L_2$, and $L_\infty $ attacks can be effectively implemented by this general framework with little modifications. Comparing with the state-of-the-art attacks in each category, our ADMM-based attacks are so far the strongest, achieving both the 100% attack success rate and the minimal distortion. Pu Zhao 0001, Sijia Liu 0001, Yanzhi Wang 0001, Xue Lin 0001 |
ACM Multimedia | 4 |
| 2018 | Dynamic Reconfiguration of Thermoelectric Generators for Vehicle Radiators Energy Harvesting Under Location-Dependent Temperature Variations
Donkyu Baek, Caiwen Ding, Sheng Lin 0001, Donghwa Shin, Xue Lin 0001, Yanzhi Wang 0001, Youngjin Cho, Naehyuck Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2017 | Algorithm accelerations for luminescent solar concentrator-enhanced reconfigurable onboard photovoltaic systemabstractElectric vehicles (EVs) and hybrid electric vehicles (HEVs) are growing in popularity. Onboard photovoltaic (PV) systems have been proposed to overcome the limited all-electric driving range of EVs/HEVs. However, there exist obstacles to the wide adoption of onboard PV systems such as low efficiency, high cost, and low compatibility. To tackle these limitations, we propose to adopt the semiconductor nanomaterial-based luminescent solar concentrator (LSC)-enhanced PV cells into the onboard PV systems. In this paper, we investigate methods of accelerating the reconfiguration algorithm for the LSC-enhanced onboard PV system to reduce computational/energy overhead and capital cost. First, in the system design stage, we group LSC-enhanced PV cells into macrocells and reconfigure the onboard PV system based on macrocells. Second, we simplify the partial shading scenario by assuming an LSC-enhanced PV cell is either lighted or completely shaded (Algorithm 1). Third, we make use of the observation that the conversion efficiency of the charger is high and nearly constant as long as its input voltage exceeds a threshold value (Algorithm 2). We test and evaluate the effectiveness of the proposed two algorithms by comparing with the optimal PV array reconfiguration algorithm and simulating an LSC-enhanced reconfigurable onboard PV system using actually measured solar irradiance traces during vehicle driving. Experiments demonstrate the output power of algorithm 1 in the first scenario is 9.0% lower in average than that of the optimal PV array reconfiguration algorithm. In the second scenario, we observe an average of 1.16X performance improvement of the proposed algorithm 2. Caiwen Ding, Ji Li 0006, Naehyuck Chang, Xue Lin 0001, Yanzhi Wang 0001 |
ASP-DAC | 5 |
| 2017 | Energy-efficient, high-performance, highly-compressed deep neural network design using block-circulant matricesabstractDeep neural networks (DNNs) have emerged as the most powerful machine learning technique in numerous artificial intelligent applications. However, the large sizes of DNNs make themselves both computation and memory intensive, thereby limiting the hardware performance of dedicated DNN accelerators. In this paper, we propose a holistic framework for energy-efficient high-performance highly-compressed DNN hardware design. First, we propose block-circulant matrix-based DNN training and inference schemes, which theoretically guarantee Big-O complexity reduction in both computational cost (from O(n2) to O(n log n)) and storage requirement (from O(n2) to O(n)) of DNNs. Second, we dedicatedly optimize the hardware architecture, especially on the key fast Fourier transform (FFT) module, to improve the overall performance in terms of energy efficiency, computation performance and resource cost. Third, we propose a design flow to perform hardware-software co-optimization with the purpose of achieving good balance between test accuracy and hardware performance of DNNs. Based on the proposed design flow, two block-circulant matrix-based DNNs on two different datasets are implemented and evaluated on FPGA. The fixed-point quantization and the proposed block-circulant matrix-based inference scheme enables the network to achieve as high as 3.5 TOPS computation performance and 3.69 TOPS/W energy efficiency while the memory is saved by 108X ~ 116X with negligible accuracy degradation. Siyu Liao, Zhe Li 0001, Xue Lin 0001, Qinru Qiu, Yanzhi Wang 0001, Bo Yuan 0001 |
ICCAD | 3 |
| 2017 | Reconfigurable thermoelectric generators for vehicle radiators energy harvestingabstractConventional internal combustion engine vehicles (ICEV) generally have less than a 30% of fuel efficiency, and the most wasted energy is dissipated in the form of heat energy. The heat energy maintains the engine temperature for efficient combustion as a good aspect, but the amount of heat generation is excessive and eventually breaks the engine components unless advanced cooling system technologies are supported such as high-capacity radiators, elaborated water jackets, high-flow rate coolant pumps, etc. The excessive heat dissipation plays a key role on a poor fuel economy, but reclamation of the heat energy has not been a main focus of vehicle design. This work is first to propose a cross-layer, system-level solution to enhance thermoelectric generator (TEG) array efficiency introducing online reconfiguration of TEG modules. The proposed method is useful to any sort of TEG array to reclaim wasted heat energy because cooling and exhaust systems generally have different inlet and outlet temperatures. In this paper, we deploy the proposed method to vehicle radiator heat energy harvesting, which does not affect the vehicle performance while exhaust heat energy harvesting may disturb the combustion and emission control integrity. We introduce a novel TEG reconfiguration and maximize the TEG array output in spite of dynamic change of the coolant flow rate and temperature, which results in a huge variation in the coolant temperature distribution of inside the radiator. The proposed method enables all the TEG modules to run at or close to their maximum power points (MPP) under dynamically changing vehicle operating conditions. Experimental results show up to a 34% enhancement compared with a fixed array structure, which is a common practice. Donkyu Baek, Caiwen Ding, Sheng Lin 0001, Donghwa Shin, Xue Lin 0001, Yanzhi Wang 0001, Naehyuck Chang |
ISLPED | 6 |
| 2017 | CirCNN: accelerating and compressing deep neural networks using block-circulant weight matricesabstractLarge-scale deep neural networks (DNNs) are both compute and memory intensive. As the size of DNNs continues to grow, it is critical to improve the energy efficiency and performance while maintaining accuracy. For DNNs, the model size is an important factor affecting performance, scalability and energy efficiency. Weight pruning achieves good compression ratios but suffers from three drawbacks: 1) the irregular network structure after pruning, which affects performance and throughput; 2) the increased training complexity; and 3) the lack of rigirous guarantee of compression ratio and inference accuracy. Caiwen Ding, Siyu Liao, Yanzhi Wang 0001, Zhe Li 0001, Ning Liu 0007, Youwei Zhuo, Chao Wang 0051, Xuehai Qian, Yu Bai 0004, Geng Yuan, Jian Tang 0008, Qinru Qiu, Xue Lin 0001, Bo Yuan 0001 |
MICRO | 15 |
| 2016 | A Profit Optimization Framework of Energy Storage Devices in Data Centers: Hierarchical Structure and Hybrid TypesabstractThis paper investigates the hierarchical deployment and over-provisioning of energy storage devices (ESDs) in data ceners by (i) adopting a realistic power delivery architecture (from Intel) for centralized ESD structure as the starting point, (ii) presenting a novel and realistic power delivery architecture, borrowing the best features of the centralized ESD structure from Intel and distributed single-level ESD structures from Google and Microsoft, and supporting the case that different types of ESDs are employed for each of the data center, rack, and server levels, (iii) providing an optimal design (i.e., determining the ESD type, and ESD provisioning at each level) and control (i.e., scheduling the charging and discharging of various ESDs) framework to maximize the amortized profit of the hierarchical ESD structure. The amortized one-time capital cost (capex), operating cost (opex), and cost associated with battery aging and replacement are considered in the profit optimization. Constraints on ESD volume and realistic characteristics of ESDs and power conversion circuitries are accounted for in the framework. (iv) conducting experiments using real data center workload traces from Google based on realistic data center specifications, demonstrating the effectiveness of the proposed design and control framework. Xue Lin 0001, Massoud Pedram, Jian Tang 0008, Yanzhi Wang 0001 |
CLOUD | 1 |
| 2016 | A Reinforcement Learning-Based Power Management Framework for Green Computing Data CentersabstractVarious power management techniques have been exploited to reduce the energy consumption of data centers. In this work, we propose a reinforcement learning-based power management framework for data centers, which does not rely on any given stationary assumptions of the job arrival and job service processes. By carefully designing the state space, the action space, and the reward of a learning process, the objective of the reinforcement learning agent coincides with our goal of reducing the server pool energy consumption with reasonable average job response time. Real Google cluster data traces are used to verify the effectiveness of the proposed reinforcement learning-based data center power management framework. Xue Lin 0001, Yanzhi Wang 0001, Massoud Pedram |
IC2E | 1 |
| 2016 | Luminescent solar concentrator-based photovoltaic reconfiguration for hybrid and plug-in electric vehiclesabstractAlong with growing public concerns over the energy crisis, hybrid and plug-in electric vehicles (HPEVs) are becoming increasingly popular. However, the total carbon footprint cannot be significantly reduced yet due to the relatively high carbon footprint of batteries in HPEVs. On-board PV systems, which mount PV cells on hood, roof, trunk, and door panels of an HPEV, can assist propelling the vehicle and enable battery charging whenever there is sunlight, and therefore, better mileage can be achieved for HPEVs. A reconfigurable on-board PV system has been proposed to tackle the output power degradation under a non-uniform distribution of solar irradiance levels on different vehicle panels. However, there are still some limitations for mounting PV cells on HPEVs even with the reconfiguration technique such as low efficiency, high cost, and appearance. To address these limitations, we propose to use semiconductor nanomaterials-based luminescent solar concentrators (LSC)-enhanced PV cells for the reconfigurable on-board PV systems. We properly optimize the size of the LSC-enhanced PV cell, the size of macrocells, and the reconfiguration period to achieve a balance between system performance and computation complexity, energy overhead, and capital cost. Furthermore, due to the transparency and flexibility of LSC polymer, we consider employing LSC-enhanced PV cells on vehicle windows. Experiments demonstrate up to 2.49× performance improvement of the proposed LSC-based PV system comparing with the baseline PV system. Caiwen Ding, Hongjia Li 0003, Yanzhi Wang 0001, Naehyuck Chang, Xue Lin 0001 |
ICCD | 6 |
| 2016 | Power-aware virtual machine mapping in the data-center-on-a-chip paradigmabstractIt is projected that hundreds of cores can be integrated into a chip at the sub-20nm technology nodes. However, some challenges exist in the many-core architecture such as maintaining memory coherence, underutilized parallelism, and increased inter-core communication delay. This work proposes the data-center-on-a-chip (DCoC) paradigm employing virtualization technologies commonly used in today's data centers to reduce the overhead of maintaining memory coherence and inter-core communication and improve parallelism. In the DCoC paradigm, user applications with specific resource requirements need to be mapped onto different chips of a data center and different cores of a chip in the form of virtual machines (VMs). By a judicious VM mapping method, the data center performance can be maximized while satisfying the power budget and power density constraints of the chips and the resource requirements of VMs. To tackle the NP-hardness of the VM mapping problem, we propose a two-tier algorithm, which effectively solves the mapping problem with polynomial time complexity. Xue Lin 0001, Yuankun Xue, Paul Bogdan, Yanzhi Wang 0001, Siddharth Garg, Massoud Pedram |
ICCD | 1 |
| 2016 | Concurrent Task Scheduling and Dynamic Voltage and Frequency Scaling in a Real-Time Embedded System With Energy HarvestingabstractEnergy harvesting is a promising technique to overcome the limit on energy availability and increase the lifespan of battery-powered embedded systems. In this paper, the question of how one can achieve the prolonged lifespan1of a real-time embedded system with energy harvesting capability (RTES-EH) is investigated. The RTES-EH comprises a photovoltaic (PV) panel for energy harvesting, a supercapacitor for energy storage, and a real-time sensor node as the embedded load device. A global controller performs simultaneous optimal operating point tracking for the PV panel, state-of-charge (SoC) management for the supercapacitor, and energy-harvesting-aware real-time task scheduling with dynamic voltage and frequency scaling (DVFS) for the sensor node, while employing a precise solar irradiance prediction method. The controller employs a cascaded feedback control structure, where an outer supervisory control loop performs real-time task scheduling with DVFS in the sensor node while maintaining the optimal supercapacitor SoC for improved system availability, and an inner control loop tracks the optimal operating point of the PV panel on the fly. Experimental results show that the proposed global controller lowers the task instance drop rate by up to 63% compared with the baseline controller within the same service time (i.e., from sunrise to sunset). Xue Lin 0001, Yanzhi Wang 0001, Naehyuck Chang, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Hierarchical Deployment and Control of Energy Storage Devices in Data CentersabstractRecent work has presented hierarchical deployment of energy storage devices (ESDs) at the data center, rack, and server levels within a data center, along with a corresponding control framework for peak power shaving and energy cost reduction under (time-of-use) dynamic energy pricing policies. However, the prior work does not use a realistic power delivery architecture of the data center with hierarchical ESD structure, and fails to account for some key characteristics such as rate capacity effect of batteries and power losses in various AC/DC and DC/DC converters in the power delivery architecture. This paper aims to overcome these shortcomings by (i) adopting a realistic power delivery architecture (from Intel) for centralized ESD structure as the starting point, (ii) presenting a novel power delivery architecture for data centers with hierarchical ESD structure, borrowing the best features of the centralized ESD structure from Intel and the distributed single-level ESD structures from Google and Microsoft, (iii) providing a mathematical framework for the optimal design (i.e., ESD provisioning) and control (i.e., Scheduling the charging and discharging of various ESDs) of the hierarchical ESD structure to minimize overall energy cost under dynamic energy pricing functions. This framework accounts for constraints on ESD volume (for each level) and the overall (annually amortized) capital cost, and power losses due to the rate capacity effect and conversion circuitry. The ESD design problem is solved by using a search-based algorithm, whereas the ESD control problem is formulated and solved as a hierarchical convex optimization algorithm. Experiments have been conducted using real Google cluster workload based on realistic data center specifications, demonstrating the effectiveness of the proposed optimal design and control framework. Shuo Wang 0009, Yanzhi Wang 0001, Xue Lin 0001, Massoud Pedram |
CLOUD | 3 |
| 2015 | Negotiation-based task scheduling and storage control algorithm to minimize user's electric bills under dynamic pricesabstractDynamic energy pricing is a promising technique in the Smart Grid to alleviate the mismatch between electricity generation and consumption. Energy consumers are incentivized to shape their power demands, or more specifically, schedule their electricity-consuming applications (tasks) more prudently to minimize their electric bills. This has become a particularly interesting problem with the availability of residential photovoltaic (PV) power generation facilities and controllable energy storage systems. This paper addresses the problem of joint task scheduling and energy storage control for energy consumers with PV and energy storage facilities, in order to minimize the electricity bill. A general type of dynamic pricing scenario is assumed where the energy price is both time-of-use and power-dependent, and various energy loss components are considered including power dissipation in the power conversion circuitries as well as the rate capacity effect in the storage system. A negotiation-based iterative approach has been proposed for joint residential task scheduling and energy storage control that is inspired by the state-of-the-art Field-Programmable Gate Array (FPGA) routing algorithms. In each iteration, it rips-up and re-schedules all tasks under a fixed storage control scheme, and then derives a new charging/discharging scheme for the energy storage based on the latest task scheduling. The concept of congestion is introduced to dynamically adjust the schedule of each task based on the historical results as well as the current scheduling status, and a near-optimal storage control algorithm is effectively implemented by solving convex optimization problem(s) with polynomial time complexity. Experimental results demonstrate the proposed algorithm achieves up to 64.22% in the total energy cost reduction compared with the baseline methods. Ji Li 0006, Yanzhi Wang 0001, Xue Lin 0001, Shahin Nazarian, Massoud Pedram |
ASP-DAC | 3 |
| 2015 | Reinforcement learning-based control of residential energy storage systems for electric bill minimizationabstractIncorporating residential-level photovoltaic energy generation and energy storage systems have proved useful in utilizing renewable power and reducing electric bills for the residential energy consumer. This is particular true under dynamic energy prices, where consumers can use PV-based generation and controllable storage modules for peak shaving on their power demand profile from the grid. In general, accurate PV power generation and load power consumption predictions and accurate system modeling are required for the storage control algorithm in most previous works. In this work, the reinforcement learning technique is adopted for deriving the optimal control policy for the residential energy storage module, which does not depend on accurate predictions of future PV power generation and/or load power consumption results and only requires partial knowledge of system modeling. In order to achieve higher convergence rate and higher performance in non-Markovian environment, we employ the TD(Λ)-learning algorithm to derive the optimal energy storage system control policy, and carefully define the state and action spaces, and reward function in the TD(Λ)-learning algorithm such that the objective of the reinforcement learning algorithm coincides with our goal of electric bill minimization for the residential consumer. Simulation results over real-world PV power generation and load power consumption profiles demonstrate that the proposed reinforcement learning-based storage control algorithm can achieve up to 59.8% improvement in energy cost reduction. Chenxiao Guan, Yanzhi Wang 0001, Xue Lin 0001, Shahin Nazarian, Massoud Pedram |
CCNC | 3 |
| 2015 | Joint automatic control of the powertrain and auxiliary systems to enhance the electromobility in hybrid electric vehiclesabstractAutonomous driving has become a major goal of automobile manufacturers and an important driver for the vehicular technology. Hybrid electric vehicles (HEVs), which represent a trade-off between conventional internal combustion engine (ICE) vehicles and electric vehicles (EVs), have gained popularity due to their high fuel economy, low pollution, and excellent compatibility with the current fossil fuel dispensing and electric charging infrastructures. To facilitate autonomous driving, an autonomous HEV controller is needed for determining the power split between the powertrain components (including an ICE and an electric motor) while simultaneously managing the power consumption of auxiliary systems (e.g., air-conditioning and lighting systems) such that the overall electromobility is enhanced. Certain (partial) prior knowledge of the future driving profile is useful information for the automatic HEV control. In this paper, methods for predicting driving profile characteristics to enhance HEV power control are first presented. Based on the prediction results and the observed HEV system state (e.g. velocity, battery state-of-charge, propulsion power demand), we propose a reinforcement learning method to determine the power source split between the ICE and electric motor while also controlling the power consumptions of the air-conditioning and lighting systems in the automobile. Experimental results demonstrate significant improvement in the overall HEV system efficiency. Yanzhi Wang 0001, Xue Lin 0001, Massoud Pedram, Naehyuck Chang |
DAC | 2 |
| 2015 | Event-driven and sensorless photovoltaic system reconfiguration for electric vehicles
Xue Lin 0001, Yanzhi Wang 0001, Massoud Pedram, Naehyuck Chang |
DATE | 1 |
| 2015 | Machine Learning-Based Energy Management in a Hybrid Electric Vehicle to Minimize Total Operating CostabstractThis paper investigates the energy management problem in hybrid electric vehicles (HEVs) focusing on the minimization of the operating cost of an HEV, including both fuel and battery replacement cost. More precisely, the paper presents a nested learning framework in which both the optimal actions (which include the gear ratio selection and the use of internal combustion engine versus the electric motor to drive the vehicle) and limits on the range of the state-of-charge of the battery are learned on the fly. The inner-loop learning process is the key to minimization of the fuel usage whereas the outer-loop learning process is critical to minimization of the amortized battery replacement cost. Experimental results demonstrate a maximum of 48% operating cost reduction by the proposed HEV energy management policy. Xue Lin 0001, Paul Bogdan, Naehyuck Chang, Massoud Pedram |
ICCAD | 1 |
| 2015 | Optimizing fuel economy of hybrid electric vehicles using a Markov decision process modelabstractIn contrast to conventional internal combustion engine (ICE) propelled vehicles, hybrid electric vehicles (HEVs) can achieve both higher fuel economy and lower pollutant emissions. The HEV features a hybrid propulsion system consisting of one ICE and one or more electric motors (EMs). The use of both ICE and EM increases the complexity of HEV power management, and so advanced power management policy is required for achieving higher performance and lower fuel consumption. This work aims at minimizing the HEV fuel consumption over any driving cycles, about which no complete information is available to the HEV controller in advance. Therefore, this work proposes to model the HEV power management problem as a Markov decision process (MDP) and derives the optimal power management policy using the policy iteration technique. Simulation results over real-world and testing driving cycles demonstrate that the proposed optimal power management policy improves HEV fuel economy by 23.9% on average compared to the rule-based policy. Xue Lin 0001, Yanzhi Wang 0001, Paul Bogdan, Naehyuck Chang, Massoud Pedram |
Intelligent Vehicles Symposium | 1 |
| 2015 | Task Scheduling with Dynamic Voltage and Frequency Scaling for Energy Minimization in the Mobile Cloud Computing EnvironmentabstractMobile cloud computing (MCC) offers significant opportunities in performance enhancement and energy saving for mobile, battery-powered devices. Applications running on mobile devices may be represented by task graphs. This work investigates the problem of scheduling tasks (which belong to the same or possibly different applications) in the MCC environment. More precisely, the scheduling problem involves the following steps: (i) determining the tasks to be offloaded onto the cloud, (ii) mapping the remaining tasks onto (potentially heterogeneous) local cores in the mobile device, (iii) determining the frequencies for executing local tasks, and (iv) scheduling tasks on the cores (for in-house tasks) and the wireless communication channels (for offloaded tasks) such that the task-precedence requirements and the application completion time constraint are satisfied while the total energy dissipation in the mobile device is minimized. A novel algorithm is presented, which starts from a minimal-delay scheduling solution and subsequently performs energy reduction by migrating tasks among the local cores and the cloud and by applying the dynamic voltage and frequency scaling technique. A linear-time rescheduling algorithm is proposed for the task migration. Simulation results demonstrate significant energy reduction with the application completion time constraint satisfied. Xue Lin 0001, Yanzhi Wang 0001, Qing Xie 0001, Massoud Pedram |
IEEE Trans. Serv. Comput. | 1 |
| 2014 | Energy and Performance-Aware Task Scheduling in a Mobile Cloud Computing EnvironmentabstractMobile cloud computing (MCC) offers significant opportunities in performance enhancement and energy saving in mobile, battery-powered devices. An application running on a mobile device can be represented by a task graph. This work investigates the problem of scheduling tasks (which belong to the same or possibly different applications) in an MCC environment. More precisely, the scheduling problem involves the following steps: (i) determining the tasks to be offloaded on to the cloud, (ii) mapping the remaining tasks onto (potentially heterogeneous) cores in the mobile device, and (iii) scheduling all tasks on the cores (for in-house tasks) or the wireless communication channels (for offloaded tasks) such that the task-precedence requirements and the application completion time constraint are satisfied while the total energy dissipation in the mobile device is minimized. A novel algorithm is presented, which starts from a minimal-delay scheduling solution and subsequently performs energy reduction by migrating tasks among the local cores or between the local cores and the cloud. A linear-time rescheduling algorithm is proposed for the task migration. Simulation results show that the proposed algorithm can achieve a maximum energy reduction by a factor of 3.1 compared with the baseline algorithm. Xue Lin 0001, Yanzhi Wang 0001, Qing Xie 0001, Massoud Pedram |
IEEE CLOUD | 1 |
| 2014 | Semi-analytical current source modeling of FinFET devices operating in near/sub-threshold regime with independent gate control and considering process variationabstractOperating circuits in the near/sub-threshold regime can lower the circuit energy consumption at the expense of lowering the circuit speed. In addition near/sub-threshold can result in higher sensitivity to process-induced variations and transient noise. FinFETs have been proposed as an alternative to planar CMOS devices in sub-20nm CMOS technology nodes due to their more effective channel control, steep sub-threshold slope, high ON/OFF current ratio, low power consumption, and so on. Characteristics of FinFETs operating in the near/sub-threshold regime make it difficult to verify the timing of a circuit using conventional statistical static timing analysis (SSTA) techniques. Current source modeling (CSM) methods, which have been proposed to increase the accuracy of timing analysis in dealing with arbitrary shapes of the input signal waveforms, are the appropriate solution for performing SSTA on FinFET-based circuits. This paper thus extends the CSM to such circuits, operating in the near/sub-threshold voltage regime. In particular, FinFET devices with independent gate control and subject to process variations are modelled. The key idea of the proposed CSM approach is to combine non-linear analytical models and low-dimensional CSM lookup tables to simultaneously achieve high modeling accuracy and low time/space complexity. Tiansong Cui, Yanzhi Wang 0001, Xue Lin 0001, Shahin Nazarian, Massoud Pedram |
ASP-DAC | 3 |
| 2014 | Minimizing state-of-health degradation in hybrid electrical energy storage systems with arbitrary source and load profilesabstractHybrid electrical energy storage (HEES) systems consisting of heterogeneous electrical energy storage (EES) elements are proposed to exploit the strengths of different EES elements and hide their weaknesses. The cycle life of the EES elements is one of the most important metrics. The cycle life is directly related to the state-of-health (SoH), which is defined as the ratio of full charge capacity of an aged EES element to its designed (or nominal) capacity. The SoH degradation models of battery in the previous literature can only be applied to charging/discharging cycles with the same state-of-charge (SoC) swing. To address this shortcoming, this paper derives a novel SoH degradation model of battery for charging/discharging cycles with arbitrary patterns. Based on the proposed model, this paper presents a near-optimal charge management policy focusing on extending the cycle life of battery elements in the HEES systems while simultaneously improving the overall cycle efficiency. Yanzhi Wang 0001, Xue Lin 0001, Qing Xie 0001, Naehyuck Chang, Massoud Pedram |
DATE | 2 |
| 2014 | Energy optimal sizing of FinFET standard cells operating in multiple voltage regimes using adaptive independent gate controlabstractFinFET has been proposed as an alternative for bulk CMOS in the ultra-low power designs due to its more effective channel control, reduced random dopant fluctuation, higher ON/OFF current ratio, lower energy consumption, etc. The characteristics of FinFETs operating in the sub/near-threshold region are very different from those in the strong-inversion region. This paper introduces an analytical transregional FinFET model with high accuracy in both subthrehold and near-threshold regions. The unique feature of independent gate controls for FinFET devices is exploited for achieving a tradeoff between energy consumption and delay, and balancing the rise and fall times of FinFET gates. This paper proposes an effective design framework of FinFET standard cells based on the adaptive independent gate control method such that they can operate properly at all of subthreshold, near-threshold and super-threshold regions. The optimal voltage for independent gate control is derived so as to achieve equal rise and fall times or minimal energy-delay product at any supply voltage level. Yanzhi Wang 0001, Xue Lin 0001, Shahin Nazarian, Massoud Pedram |
ACM Great Lakes Symposium on VLSI | 3 |
| 2014 | Optimal power switch design methodology for ultra dynamic voltage scaling with a limited number of power railsabstractMany burst-mode applications require high performance for brief time periods between extended sections of low performance operation. Digital circuits supporting such burst-mode applications should work in both the near-threshold regime and the super-threshold regime for brief time periods. This work proposes the structure support of fine-grained ultra dynamic voltage scaling (UDVS) from the traditional strong-inversion region to the near-threshold region, with limitations on the number of power rails. The number, type, and size of the power switches are jointly optimized to minimize the overall energy consumption of the UDVS circuit block, meanwhile satisfying the target delay or frequency requirement at each DVS level. The proposed optimization framework properly accounts for the dynamic energy consumption as well as the leakage energy consumption through all the power switches during both the operation time and stand-by time of the circuit block. Experimental results on 22nm Predictive Technology Model demonstrate the effectiveness of the proposed optimization framework. Yanzhi Wang 0001, Xue Lin 0001, Massoud Pedram |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Reinforcement learning based power management for hybrid electric vehiclesabstractCompared to conventional internal combustion engine (ICE) propelled vehicles, hybrid electric vehicles (HEVs) can achieve both higher fuel economy and lower pollution emissions. The HEV consists of a hybrid propulsion system containing one ICE and one or more electric motors (EMs). The use of both ICE and EM increases the complexity of HEV power management, and therefore requires advanced power management policies to achieve higher performance and lower fuel consumption. Towards this end, our work aims at minimizing the HEV fuel consumption over any driving cycle (without prior knowledge of the cycle) by using a reinforcement learning technique. This is in clear contrast to prior work, which requires deterministic or stochastic knowledge of the driving cycles. In addition, the proposed reinforcement learning technique enables us to (partially) avoid reliance on complex HEV modeling while coping with driver specific behaviors. To our knowledge, this is the first work that applies the reinforcement learning technique to the HEV power management problem. Simulation results over real-world and testing driving cycles demonstrate the proposed HEV power management policy can improve fuel economy by 42%. Xue Lin 0001, Yanzhi Wang 0001, Paul Bogdan, Naehyuck Chang, Massoud Pedram |
ICCAD | 1 |
| 2014 | Power supply and consumption co-optimization of portable embedded systems with hybrid power supplyabstractEnergy efficiency has always been an important design criterion for portable embedded systems. To compensate for the shortcomings of electrochemical batteries such as low power density, limited cycle life, and the rate capacity effect, supercapacitors have been employed as complementary power supplies for electrochemical batteries, i.e., hybrid power supplies comprised of batteries and supercapacitors have been proposed. In this work, we consider a portable embedded system with a hybrid power supply and executing periodic real-time tasks. We perform system power management from both the power supply side and the power consumption side to maximize the system service time. Specifically, we use feedback control for maintaining the supercapacitor energy at a certain level by regulating the discharging current of the battery, such that the supercapacitor has the capability to buffer the load current fluctuation. At the power consumption side, we perform task scheduling to assist supercapacitor energy maintenance. Experimental results demonstrate that the proposed joint optimization framework of task scheduling and power supply control successfully prolongs the total service time by up to 57%. Xue Lin 0001, Yanzhi Wang 0001, Naehyuck Chang, Massoud Pedram |
ICCD | 1 |
| 2014 | Architecture and Control Algorithms for Combating Partial Shading in Photovoltaic SystemsabstractPartial shading is a serious obstacle to the effective utilization of photovoltaic (PV) systems since it can result in a significant degradation in the PV system output power. A PV system is organized as a series connection of PV modules, each module comprising a number of series-parallel connected PV cells. Backup PV cell employment and PV module reconfiguration techniques have been proposed to improve the performance of the PV system under the partial shading effects. However, these approaches are not very effective since they are costly in terms of their PV cell count and/or cell connectivity requirements. In contrast, this paper presents a cost-effective, reconfigurable PV module architecture with integrated switches in each PV cell. This paper also presents a dynamic programming algorithm to adaptively produce near-optimal reconfigurations of each PV module so as to maximize the PV system output power under any partial shading pattern. We implement a working prototype of reconfigurable PV module with 16 PV cells and confirm 45.2% output power level improvement. Using accurate PV cell models extracted from prototype measurement, we have demonstrated up to a factor of 2.36X output power improvement of a large-scale PV system comprised of three PV modules with 60 PV cells per module. Yanzhi Wang 0001, Xue Lin 0001, Younghyun Kim 0001, Naehyuck Chang, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Single-Source, Single-Destination Charge Migration in Hybrid Electrical Energy Storage SystemsabstractIn spite of extensive research it is still quite expensive to store electrical energy without converting it to a different form of energy. As of today, no single type of electrical energy storage (EES) element can fulfill all the desirable features of an ideal storage device, e.g., high-efficiency, high-power/energy capacity, low-cost, and long-cycle life. A hybrid EES system (HEES) consists of two or more heterogeneous EES elements, realizing the advantages of each EES element while hiding their weaknesses. HEES systems exhibit superior performance compared with homogeneous EES systems when appropriate charge allocation and replacement policies are developed and used. In addition, charge migration is mandatory because the optimal EES banks for charge allocation and replacement are in general different, and each EES bank has limited storage capacity. This paper formally describes the notion of charge migration efficiency and its optimization. We first define the charge migration architecture and the corresponding charge migration optimization problem. We provide a systematic solution for the single-source, single-destination charge migration problem considering the efficiency variation of the converters, the rate capacity and internal power loss of the storage element, the terminal voltage variation of the storage elements as a function of their state of charge, and so on. We also introduce the optimal solutions for both the time-constrained and -unconstrained versions of the charge migration problem formulations. Experimental results demonstrate significant charge migration efficiency improvement of up to 83.4%. Yanzhi Wang 0001, Xue Lin 0001, Younghyun Kim 0001, Qing Xie 0001, Massoud Pedram, Naehyuck Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Capital cost-aware design and partial shading-aware architecture optimization of a reconfigurable photovoltaic systemabstractPhotovoltaic (PV) systems are often subject to partial shading that significantly degrades the output power of the whole systems. Reconfiguration methods have been proposed to adaptively change the PV panel configuration according to the current partial shading pattern. The reconfigurable PV panel architecture integrates every PV cell with three programmable switches to facilitate the PV panel reconfiguration. The additional switches, however, increase the capital cost of the PV system. In this paper, we group a number of PV cells into a PV macro-cell, and the PV panel reconfiguration only changes the connections between adjacent PV macro-cells. The size and internal structure (i.e., the series-parallel connection of PV cells) of all PV macro-cells are the same and will not be changed after PV system installation in the field. Determining the optimal size of the PV macro-cell is the result of a trade-off between the decreased PV system capital cost and enhanced PV system performance. A larger PV macro-cell reduces the cost overhead whereas a smaller PV macro-cell achieves better performance. In this paper, we set out to calculate the optimal size of the PV macro-cells such that the maximum system performance can be achieved subject to an overall system cost limitation. This “design” problem is solved using an efficient search algorithm. In addition, we provide for in-field reconfigurability of the PV panel by enabling formation of series-connected groups of parallel-connected macro-cells. We ensure maximum output power for the PV system in response to any incurring partial shading pattern. This “architecture optimization” problem is solved using dynamic programming. Yanzhi Wang 0001, Xue Lin 0001, Massoud Pedram, Naehyuck Chang |
DATE | 2 |
| 2013 | Optimal control of a grid-connected hybrid electrical energy storage system for homesabstractIntegrating residential photovoltaic (PV) power generation and electrical energy storage (EES) systems into the Smart Grid is an effective way of utilizing renewable power and reducing the consumption of fossil fuels. This has become a particularly interesting problem with the introduction of dynamic electricity energy pricing models since electricity consumers can use their PV-based energy generation and EES systems for peak shaving on their power demand profile from the grid, and thereby, minimize their electricity bill. Due to the characteristics of a realistic electricity price function and the energy storage capacity limitation, the control algorithm for a residential EES system should accurately account for various energy loss components during operation. Hybrid electrical energy storage (HEES) systems are proposed to exploit the strengths of each type of EES element and hide its weaknesses so as to achieve a combination of performance metrics that is superior to those of any of its individual EES components. This paper introduces the problem of how best to utilize a HEES system for a residential Smart Grid user equipped with PV power generation facilities. The optimal control algorithm for the HEES system is developed, which aims at minimization of the total electricity cost over a billing period under a general electricity energy price function. The proposed algorithm is based on dynamic programming and has polynomial time complexity. Experimental results demonstrate that the proposed HEES system and optimal control algorithm achieves 73.9% average profit enhancement over baseline homogeneous EES systems. Yanzhi Wang 0001, Xue Lin 0001, Massoud Pedram, Sangyoung Park, Naehyuck Chang |
DATE | 2 |
| 2013 | Joint sizing and adaptive independent gate control for FinFET circuits operating in multiple voltage regimes using the logical effort methodabstractFinFET has been proposed as an alternative for bulk CMOS in current and future technology nodes due to more effective channel control, reduced random dopant fluctuation, high ON/OFF current ratio, lower energy consumption, etc. Key characteristics of FinFET operating in the sub/near-threshold region are very different from those in the strong-inversion region. This paper first introduces an analytical transregional FinFET model with high accuracy in both sub- and near-threshold regimes. Next, the paper extends the well-known and widely-adopted logical effort delay calculation and optimization method to FinFET circuits operating in multiple voltage (sub/near/super-threshold) regimes. More specifically, a joint optimization of gate sizing and adaptive independent gate control is presented and solved in order to minimize the delay of FinFET circuits operating in multiple voltage regimes. Experimental results on a 32nm Predictive Technology Model for FinFET demonstrate the effectiveness of the proposed logical effort-based delay optimization framework. Xue Lin 0001, Yanzhi Wang 0001, Massoud Pedram |
ICCAD | 1 |
| 2013 | A framework of concurrent task scheduling and dynamic voltage and frequency scaling in real-time embedded systems with energy harvestingabstractEnergy harvesting is a promising technique to overcome the limitation imposed by the finite energy capacity of batteries in conventional battery-powered embedded systems. In particular, the question of how one can achieve full energy autonomy (i.e., perpetual, battery-free operation) of a real-time embedded system with an energy harvesting capability (RTES-EH) by applying a global control strategy is investigated. The energy harvesting module is comprised of a Photovoltaic (PV) panel for harvesting energy and a supercapacitor for storing any excess energy. The global controller performs optimal operating point tracking for the PV panel, state-of-charge management for the supercapacitor, and energy-harvesting-aware real-time task scheduling with dynamic voltage and frequency scaling (DVFS) in the embedded load device. The controller, which accounts for dynamic V-I characteristics of the PV panel, terminal voltage variation and self-leakage of the supercapacitor, and power losses in voltage converters, employs a cascaded feedback control structure with an inner control loop determining the V-I operating point of the PV panel and an outer supervisory control loop performing real-time task scheduling and setting the voltage and frequency level in the embedded load device (to keep the state-of-charge of the supercapacitor in a desirable range). Experimental results show that the proposed global controller lowers the task drop rate in a RTES-EH by up to 60% compared with baseline controller within the same service time. Xue Lin 0001, Yanzhi Wang 0001, Siyu Yue, Naehyuck Chang, Massoud Pedram |
ISLPED | 1 |
| 2012 | Near-optimal, dynamic module reconfiguration in a photovoltaic system to combat partial shading effectsabstractPartial shading is a serious obstacle to effective utilization of photovoltaic (PV) systems since it can result in significant output power degradation for the system. A PV system is organized as a series connection of PV modules, each module comprising of a number of series-parallel connected cells. This paper presents modified PV cell structures with integrated switches, imbalanced cell connection topologies for PV modules, and a dynamic programming algorithm to produce near-optimal reconfigurations of each PV module with the goal of maximizing the system output power level under any partial shading patterns. Through simulations, we have demonstrated up to a factor of 2.3X improvement in the output power level of a PV system comprised of 3 PV modules with 60 PV cells per module. Xue Lin 0001, Yanzhi Wang 0001, Siyu Yue, Donghwa Shin, Naehyuck Chang, Massoud Pedram |
DAC | 1 |
| 2012 | State of health aware charge management in hybrid electrical energy storage systemsabstractThis paper is the first to present an efficient charge management algorithm focusing on extending the cycle life of battery elements in hybrid electrical energy storage (HEES) systems while simultaneously improving the overall cycle efficiency. In particular, it proposes to apply a crossover filter to the power source and load profiles. The goal of this filtering technique is to allow the battery banks to stably (i.e., with low variation) receive energy from the power source and/or provide energy to the load device, while leaving the spiky (i.e., with high variation) power supply or demand to be dealt with by the supercapacitor banks. To maximize the HEES system cycle efficiency, a mathematical problem is formulated and solved to determine the optimal charging/discharging current profiles and charge transfer interconnect voltage, taking into account the power loss of the EES elements and power converters. To minimize the state of health (SoH) degradation of the battery array in the HEES system, we make use of two facts: the SoH of battery is better maintained if (i) the SoC swing is smaller, and (ii) the same SoC swing occurs at lower average SoC. Now then using the supercapacitor bank to deal with the high-frequency component of the power supply or demand, we can reduce the SoC swing for the battery array and lower the SoC of the array. A secondary helpful effect is that, for fixed and given amount of energy delivered to the load device, an improvement in the overall charge cycle efficiency of the HEES system translates into a further reduction in both the average SoC and the SoC swing of the battery array. The proposed charge management algorithm for a Li-ion battery - supercapacitor bank HEES system is simulated and compared to a homogeneous EES system comprised of Li-ion batteries only. Experimental results show significant performance enhancements for the HEES system, an increase of up to 21.9% and 4.82x in terms of the cycle efficiency and cycle life, respectively. Qing Xie 0001, Xue Lin 0001, Yanzhi Wang 0001, Massoud Pedram, Donghwa Shin, Naehyuck Chang |
DATE | 2 |
| 2012 | Online fault detection and tolerance for photovoltaic energy harvesting systemsabstractPhotovoltaic energy harvesting systems (PV systems) are subject to PV cell faults, which decrease the efficiency of PV systems and even shorten the PV system lifespan. Manual PV cell fault detection and elimination are expensive and nearly impossible for remote PV systems, e.g., PV systems on satellites. Therefore, online fault detection techniques and fault tolerance solutions are needed that can detect and tolerate PV cell faults without manual intervention. In this work, we present an online fault detection and tolerance technique for remote PV systems, which is capable of dynamically locating faulty PV cells and tolerating PV cell faults. More precisely, we present a modified PV panel structure and an efficient algorithm for our online fault detection and tolerance. Our fault detection and tolerance technique reduces output power degradation due to PV cell faults in a PV system by up to 81.31%. Xue Lin 0001, Yanzhi Wang 0001, Di Zhu 0002, Naehyuck Chang, Massoud Pedram |
ICCAD | 1 |
| 2012 | Dynamic reconfiguration of photovoltaic energy harvesting system in hybrid electric vehiclesabstractPhotovoltaic (PV) energy harvesting system is a promising energy source for battery replenishment in hybrid electric vehicles (HEVs.) The PV cell array is installed on different parts of a vehicle body such as the engine hood, door panels, and the roof panel. Non-uniformity of the solar irradiance and temperature on the PV cell array is, however, a serious obstacle to efficient utilization of the PV system in HEVs because such variation, if not managed properly, can result in a significant degradation in the overall output power level of the PV system. This paper presents a dynamic PV array reconfiguration technique with structural support and a dynamic programming-based algorithm with polynomial time complexity to produce the near-optimal reconfiguration of the PV array on the HEV. The goal of this technique is to maximize the PV system output power under any solar irradiance and temperature distribution on the PV array. We demonstrate up to 6X improvement in the output power of a PV system against a conventional fixed configuration PV system. Yanzhi Wang 0001, Xue Lin 0001, Naehyuck Chang, Massoud Pedram |
ISLPED | 2 |