Yanyu Li

dblp:194/5818 · DBLP profile ↗
← Back
48ranked-venue papers
9as first author
47since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 7 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 18 since 2021Systems, architecture and hardware · 15 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2025 LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
abstract
Diffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications. The promising results come at the cost of slow inference, as each denoising step requires running the whole transformer model with a large amount of parameters. In this paper, we show that performing the full computation of the model at each diffusion step is unnecessary, as some computations can be skipped by lazily reusing the results of previous steps. Furthermore, we show that the lower bound of similarity between outputs at consecutive steps is notably high, and this similarity can be linearly approximated using the inputs. To verify our demonstrations, we propose the **LazyDiT**, a lazy learning framework that efficiently leverages cached results from earlier steps to skip redundant computations. Specifically, we incorporate lazy learning layers into the model, effectively trained to maximize laziness, enabling dynamic skipping of redundant computations. Experimental results show that LazyDiT outperforms the DDIM sampler across multiple diffusion transformer models at various resolutions. Furthermore, we implement our method on mobile devices, achieving better performance than DDIM with similar latency.
Xuan Shen, Zhao Song 0002, Yufa Zhou 0001, Bo Chen 0029, Yanyu Li, Yifan Gong 0004, Kai Zhang 0045, Hao Tan 0002, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu 0002, Pu Zhao 0001, Yanzhi Wang 0001, Jiuxiang Gu
AAAI5
2025 SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
abstract
Existing text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 10242px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 2562px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7× smaller than SDXL, 14× smaller than IF-XL).
Jierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun, Arpit Sahni, Aarush Gupta, Anujraaj Goyal, Dishani Lahiri, Yerlan Idelbayev, Junli Cao, Yanyu Li, Kwang-Ting Cheng, Shueng-Han Gary Chan, Mingming Gong, Sergey Tulyakov, Anil Kag, Yanwu Xu 0003, Jian Ren 0005
CVPR12
2025 SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
abstract
We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality. Project page at https://snap-research.github.io/snapgen-v/.
Yushu Wu, Yanyu Li, Yanwu Xu 0003, Anil Kag, Yang Sui 0001, Huseyin Coskun, Aleksei Lebedev, Ju Hu, Dimitris N. Metaxas, Yanzhi Wang 0001, Sergey Tulyakov, Jian Ren 0005
CVPR3
2025 Improving the Diffusability of Autoencoders
abstract
Latent diffusion models have emerged as the leading approach for generating high-quality images and videos, utilizing compressed latent representations to reduce the computational burden of the diffusion process. While recent advancements have primarily focused on scaling diffusion backbones and improving autoencoder reconstruction quality, the interaction between these components has received comparatively less attention. In this work, we perform a spectral analysis of modern autoencoders and identify inordinate high-frequency components in their latent spaces, which are especially pronounced in the autoencoders with a large bottleneck channel size. We hypothesize that this high-frequency component interferes with the coarse-to-fine nature of the diffusion synthesis process and hinders the generation quality. To mitigate the issue, we propose scale equivariance: a simple regularization strategy that aligns latent and RGB spaces across frequencies by enforcing scale equivariance in the decoder. It requires minimal code changes and only up to $20$K autoencoder fine-tuning steps, yet significantly improves generation quality, reducing FID by 19% for image generation on ImageNet-1K 256x256 and FVD by at least 44% for video generation on Kinetics-700 17x256x256. The source code is available at https://github.com/snap-research/diffusability.
Ivan Skorokhodov, Sharath Girish, Benran Hu 0001, Willi Menapace, Yanyu Li, Rameen Abdal, Sergey Tulyakov, Aliaksandr Siarohin
ICML5
2024 TextCraftor: Your Text Encoder can be Image Quality Controller
abstract
Diffusion-based text-to-image generative models, e.g., Stable Diffusion, have revolutionized the field of content generation, enabling significant advancements in areas like image editing and video synthesis. Despite their formidable capabilities, these models are not without their limitations. It is still challenging to synthesize an image that aligns well with the input text, and multiple runs with carefully crafted prompts are required to achieve satisfactory results. To mitigate these limitations, numerous studies have endeavored to fine-tune the pre-trained diffusion models, i.e., UNet, utilizing various technologies. Yet, amidst these efforts, a pivotal question of text-to-image diffusion model training has remained largely unexplored: Is it possible and feasible to fine-tune the text encoder to improve the performance of text-to-image diffusion models? Our findings reveal that, instead of replacing the CLIP text encoder used in Stable Diffusion with other large language models, we can enhance it through our proposed fine-tuning approach, TextCraftor, leading to substantial improvements in quantitative benchmarks and human assessments. Interestingly, our technique also empowers controllable image generation through the interpolation of different text encoders fine-tuned with various rewards. We also demonstrate that TextCraftor is orthogonal to UNet finetuning, and can be combined to further improve generative quality.
Yanyu Li, Anil Kag, Ju Hu, Yerlan Idelbayev, Dhritiman Sagar, Yanzhi Wang 0001, Sergey Tulyakov, Jian Ren 0005
CVPR1
2024 DACO: Pursuing Ultra-low Power Consumption via DNN-Adaptive CPU-GPU CO-optimization on Mobile Devices
abstract
As Deep Neural Networks (DNNs) become popular in mobile systems, their high computational and memory demands make them major power consumers, especially in limited-budget scenarios. In this paper, we propose DACO, a DNN-Adaptive CPU-GPU CO-optimization technique, to reduce the power consumption of DNNs. First, a resource-oriented classifier is proposed to quantify the computation/memory intensity of DNN models and classify them accordingly. Second, a set of rule-based policies is deduced for achieving the best-suited CPU-GPU system configuration in a coarse-grained manner. Combined with all the rules, a coarse-to-fine CPU-GPU auto-tuning approach is proposed to reach the Pareto-optimal speed and power consumption in DNN inference. Experimental results demonstrate that, compared with the existing approach, DACO could reduce power consumption by up to 71.9% while keeping an excellent DNN inference speed.
Yushu Wu, Chao Wu 0006, Geng Yuan, Yanyu Li, Weichao Guo, Jing Rao, Xipeng Shen, Bin Ren 0002, Yanzhi Wang 0001
DATE4
2024 Efficient Training with Denoised Neural Weights
Yifan Gong 0004, Zheng Zhan 0001, Yanyu Li, Yerlan Idelbayev, Andrey Zharkov, Kfir Aberman, Sergey Tulyakov, Yanzhi Wang 0001, Jian Ren 0003
ECCV (83)3
2024 SDA: Low-Bit Stable Diffusion Acceleration on Edge FPGAs
abstract
This paper introduces SDA, the first effort to adapt the expensive stable diffusion (SD) model for edge FPGA deployment. First, we apply quantization-aware training to quantize its weights to 4 -bit and activations to 8 -bit ($W 4 A 8$) with a negligible accuracy loss. Based on that, we propose a high-performance hybrid systolic array (hybridSA) architecture that natively executes convolution and attention operators across varying quantization bit-widths (e.g., $W 4 A 8$ and all 8 -bit $Q K^{T} V$ in attention). To improve computational efficiency, hybridSA integrates diverse DSP packing techniques into hybrid weightstationary and output-stationary dataflows that are optimized for convolution and attention. It also supports flexible dataflow transitions to address the distinct demands of its output sequence by subsequent nonlinear operators. Moreover, we observe that nonlinear operators become the new performance bottleneck after the acceleration of convolution and attention, and offload them onto the FPGA as well. To reduce the latency of each nonlinear operator, we pipeline its own execution at a fine granularity. To minimize the resource utilization of nonlinear operators, we carefully balance their execution with hybridSA in a coarse-grained pipeline. Experimental results demonstrate that our low-bit ($W 4 A 8$) SDA accelerator on the embedded AMDXilinx ZCU102 FPGA achieves a speedup of $97.3 \times$ (which takes about $\mathrm{2 . 1}$ minutes for one SD inference), compared to the original SD-v1.5 model on the ARM Cortex-A53 CPU (which takes about 3.5 hours for one SD inference). Our SDA project is open sourced here: https://github.com/Michaela1224/SDA_code.
Geng Yang 0001, Yanyue Xie, Zhong Jia Xue, Sung-En Chang, Yanyu Li, Peiyan Dong, Jie Lei 0001, Weiying Xie, Yanzhi Wang 0001, Xue Lin 0001, Zhenman Fang
FPL5
2024 HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion
abstract
Despite significant advances in large-scale text-to-image models, achieving hyper-realistic human image generation remains a desirable yet unsolved task. Existing models like Stable Diffusion and DALL·E 2 tend to generate human images with incoherent parts or unnatural poses. To tackle these challenges, our key insight is that human image is inherently structural over multiple granularities, from the coarse-level body skeleton to fine-grained spatial geometry. Therefore, capturing such correlations between the explicit appearance and latent structure in one model is essential to generate coherent and natural human images. To this end, we propose a unified framework, HyperHuman, that generates in-the-wild human images of high realism and diverse layouts. Specifically, 1) we first build a large-scale human-centric dataset, named HumanVerse, which consists of 340M images with comprehensive annotations like human pose, depth, and surface normal. 2) Next, we propose a Latent Structural Diffusion Model that simultaneously denoises the depth and surface normal along with the synthesized RGB image. Our model enforces the joint learning of image appearance, spatial relationship, and geometry in a unified network, where each branch in the model complements to each other with both structural awareness and textural richness. 3) Finally, to further boost the visual quality, we propose a Structure-Guided Refiner to compose the predicted conditions for more detailed generation of higher resolution. Extensive experiments demonstrate that our framework yields the state-of-the-art performance, generating hyper-realistic human images under diverse scenarios.
Jian Ren 0005, Aliaksandr Siarohin, Ivan Skorokhodov, Yanyu Li, Dahua Lin, Xihui Liu, Ziwei Liu 0002, Sergey Tulyakov
ICLR5
2024 Research on Image Aesthetic Assessment based on Graph Convolutional Network
abstract
Image aesthetic quality assessment (IAA) is aimed at predicting the general aesthetic evaluation of images by the public. However, the common neural network has a fixed size requirement for the input and ignores the close relation between image composition and aesthetic rating. Therefore, in this paper, we design and implement a concise and efficient IAA algorithm based on graph convolutional neural (GCN) networks, and apply two training schemes considering the original size of images to retain the original composition of the image to the extent. Our framework first transforms the images’ high-level semantic features into graph form, and a residual GCN-based module is then constructed for composition to learn aesthetic features, generating aesthetic score distributions. The experimental results demonstrate that graph construction and application of residual GCNs are effective for IAA model and our algorithm achieves the state-of-the-arts on the aesthetic dataset AVA.
Yanyu Li, Jiangbo Xu, Ruoyu Zou
ICME1
2024 E2GAN: Efficient Training of Efficient GANs for Image-to-Image Translation
abstract
One highly promising direction for enabling flexible real-time on-device image editing is utilizing data distillation by leveraging large-scale text-to-image diffusion models to generate paired datasets used for training generative adversarial networks (GANs). This approach notably alleviates the stringent requirements typically imposed by high-end commercial GPUs for performing image editing with diffusion models. However, unlike text-to-image diffusion models, each distilled GAN is specialized for a specific image editing task, necessitating costly training efforts to obtain models for various concepts. In this work, we introduce and address a novel research direction: can the process of distilling GANs from diffusion models be made significantly more efficient? To achieve this goal, we propose a series of innovative techniques. First, we construct a base GAN model with generalized features, adaptable to different concepts through fine-tuning, eliminating the need for training from scratch. Second, we identify crucial layers within the base GAN model and employ Low-Rank Adaptation (LoRA) with a simple yet effective rank search process, rather than fine-tuning the entire base model. Third, we investigate the minimal amount of data necessary for fine-tuning, further reducing the overall training time. Extensive experiments show that we can efficiently empower GANs with the ability to perform real-time high-quality image editing on mobile devices with remarkably reduced training and storage costs for each concept.
Yifan Gong 0004, Zheng Zhan 0001, Qing Jin, Yanyu Li, Yerlan Idelbayev, Andrey Zharkov, Kfir Aberman, Sergey Tulyakov, Yanzhi Wang 0001, Jian Ren 0005
ICML4
2024 FasterVD: On Acceleration of Video Diffusion Models
Pinrui Yu, Timothy Rupprecht, Zhenglun Kong, Pu Zhao 0001, Yanyu Li, Octavia I. Camps, Xue Lin 0001, Yanzhi Wang 0001
IJCAI7
2024 DEFCON: Deformable Convolutions Leveraging Interval Search and GPU Texture Hardware
abstract
Deformable convolutions can improve detection accuracy in Convolution Neural Networks (CNNs) by leveraging flexible spatial sampling in augmenting kernels with learnable offsets. However, the resulting irregular memory access patterns and additional pixel lookup overhead introduced by deformable layers pose inherent challenges when executed on high-throughput devices such as GPUs. To address these challenges, we introduce DEFCON, a systematic approach to optimizing deformable convolutions. DEFCON is designed to provide: (1) better placement of operators in the neural architecture using interval search, (2) reduced computational demands by leveraging lightweight operators, and (3) optimized inference by using GPU texture hardware. By performing an interval search, we reduce the number of deformable layers in our architecture. By leveraging the GPU’s texture hardware, we are able to use lightweight operators to improve the execution performance of layers, without sacrificing prediction accuracy. By combining these approaches, DEFCON increases the inference performance by 2.8× over YOLACT++ implementation, when run on an NVIDIA Jetson AGX Xavier GPU. Our work enables faster and more accurate predictions when performing deformable convolutions.
Malith Jayaweera, Yanyu Li, Yanzhi Wang 0001, Bin Ren 0002, David R. Kaeli
IPDPS2
2024 BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
abstract
Diffusion-based image generation models have achieved great success in recent years by showing the capability of synthesizing high-quality content. However, these models contain a huge number of parameters, resulting in a significantly large model size. Saving and transferring them is a major bottleneck for various applications, especially those running on resource-constrained devices. In this work, we develop a novel weight quantization method that quantizes the UNet from Stable Diffusion v1.5 to $1.99$ bits, achieving a model with $7.9\times$ smaller size while exhibiting even better generation quality than the original one. Our approach includes several novel techniques, such as assigning optimal bits to each layer, initializing the quantized model for better performance, and improving the training strategy to dramatically reduce quantization error. Furthermore, we extensively evaluate our quantized model across various benchmark datasets and through human evaluation to demonstrate its superior generation quality.
Yang Sui 0001, Yanyu Li, Anil Kag, Yerlan Idelbayev, Junli Cao, Ju Hu, Dhritiman Sagar, Bo Yuan 0001, Sergey Tulyakov, Jian Ren 0005
NeurIPS2
2024 SF-V: Single Forward Video Generation Model
abstract
Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models require multiple denoising steps during sampling, resulting in high computational costs. In this work, we propose a novel approach to obtain single-step video generation models by leveraging adversarial training to fine-tune pre-trained video diffusion models. We show that, through the adversarial training, the multi-steps video diffusion model, i.e., Stable Video Diffusion (SVD), can be trained to perform single forward pass to synthesize high-quality videos, capturing both temporal and spatial dependencies in the video data. Extensive experiments demonstrate that our method achieves competitive generation quality of synthesized videos with significantly reduced computational overhead for the denoising process (i.e., around $23\times$ speedup compared with SVD and $6\times$ speedup compared with existing works, with even better generation quality), paving the way for real-time video synthesis and editing.
Yanyu Li, Yushu Wu, Yanwu Xu 0003, Anil Kag, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Dimitris N. Metaxas, Sergey Tulyakov, Jian Ren 0005
NeurIPS2
2024 EQ-ViT: Algorithm-Hardware Co-Design for End-to-End Acceleration of Real-Time Vision Transformer Inference on Versal ACAP Architecture
abstract
While Vision Transformers (ViTs) have shown consistent progress in computer vision, deploying them for real-time decision-making scenarios (< 1 ms) is challenging. Current computing platforms like CPUs, GPUs, or FPGA-based solutions struggle to meet this deterministic low-latency real-time requirement, even with quantized ViT models. Some approaches use pruning or sparsity to reduce model size and latency, but this often results in accuracy loss. To address the aforementioned constraints, in this work, we propose EQ-ViT, an end-to-end acceleration framework with novel algorithm and architecture co-design features to enable real-time ViT acceleration on AMD Versal Adaptive Compute Acceleration Platform (ACAP). The contributions are four-fold. First, we perform in-depth kernel-level performance profiling & analysis and explain the bottlenecks for existing acceleration solutions on GPU, FPGA, and ACAP. Second, on the hardware level, we introduce a new spatial and heterogeneous accelerator architecture, EQ-ViT architecture. This architecture leverages the heterogeneous features of ACAP, where both FPGA and artificial intelligence engines (AIEs) coexist on the same system-on-chip (SoC). Third, On the algorithm level, we create a comprehensive quantization-aware training strategy, EQ-ViT algorithm. This strategy concurrently quantizes both weights and activations into 8-bit integers, aiming to improve accuracy rather than compromise it during quantization. Notably, the method also quantizes nonlinear functions for efficient hardware implementation. Fourth, we design EQ-ViT automation framework to implement the EQ-ViT architecture for four different ViT applications on the AMD Versal ACAP VCK190 board, achieving accuracy improvement with 2.4%, and average speedups of 315.0x, 3.39x, 3.38x, 14.92x, 59.5x, 13.1x over computing solutions of Intel Xeon 8375C vCPU, Nvidia A10G, A100, Jetson AGX Orin GPUs, and AMD ZCU102, U250 FPGAs. The energy efficiency gains are 62.2x, 15.33x, 12.82x, 13.31x, 13.5x, 21.9x.
Peiyan Dong, Jinming Zhuang, Zhuoping Yang, Shixin Ji, Yanyu Li, Dongkuan Xu, Heng Huang 0001, Jingtong Hu, Alex K. Jones, Yiyu Shi 0001, Yanzhi Wang 0001, Peipei Zhou 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Towards Real-Time Segmentation on the Edge
abstract
The research in real-time segmentation mainly focuses on desktop GPUs. However, autonomous driving and many other applications rely on real-time segmentation on the edge, and current arts are far from the goal. In addition, recent advances in vision transformers also inspire us to re-design the network architecture for dense prediction task. In this work, we propose to combine the self attention block with lightweight convolutions to form new building blocks, and employ latency constraints to search an efficient sub-network. We train an MLP latency model based on generated architecture configurations and their latency measured on mobile devices, so that we can predict the latency of subnets during search phase. To the best of our knowledge, we are the first to achieve over 74% mIoU on Cityscapes with semi-real-time inference (over 15 FPS) on mobile GPU from an off-the-shelf phone.
Yanyu Li, Changdi Yang, Pu Zhao 0001, Geng Yuan, Wei Niu 0002, Jiexiong Guan, Hao Tang 0005, Minghai Qin, Qing Jin, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001
AAAI1
2023 Pruning Parameterization with Bi-level Optimization for Efficient Semantic Segmentation on the Edge
abstract
With the ever-increasing popularity of edge devices, it is necessary to implement real-time segmentation on the edge for autonomous driving and many other applications. Vision Transformers (ViTs) have shown considerably stronger results for many vision tasks. However, ViTs with the fullattention mechanism usually consume a large number of computational resources, leading to difficulties for real- time inference on edge devices. In this paper, we aim to derive ViTs with fewer computations and fast inference speed to facilitate the dense prediction of semantic segmentation on edge devices. To achieve this, we propose a pruning parameterization method to formulate the pruning problem of semantic segmentation. Then we adopt a bi-level optimization method to solve this problem with the help of implicit gradients. Our experimental results demonstrate that we can achieve 38.9 mIoU on ADE20K val with a speed of 56.5 FPS on Samsung S21, which is the highest mIoU under the same computation constraint with real-time inference.
Changdi Yang, Pu Zhao 0001, Yanyu Li, Wei Niu 0002, Jiexiong Guan, Hao Tang 0005, Minghai Qin, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001
CVPR3
2023 Late Breaking Results: Fast Fair Medical Applications? Hybrid Vision Models Achieve the Fairness on the Edge
abstract
As edge devices become readily available and indispensable, there is an urgent need for effective and efficient intelligent applications to be deployed widespread. However, fairness has always been an issue, especially in edge medical applications. Compared to convolutional neuron networks (CNNs), Vision Transformer (ViT) has a better ability to extract global information, which will contribute to alleviating the unfairness problem. Typically, ViTs consume large amounts of computational and memory resources, which hinders their usage on edge. In this work, we propose a novel hardware-efficient Vision Model search framework for the fair dermatology classification, namely HeViFa. Experimental results show that HeViFa could search for a hybrid ViT model that reaches 173.1 FPS on a Samsung S21 mobile phone with 85.71% accuracy on the light skin dataset and 80.85% accuracy on the dark skin dataset. Note that HeViFa can reach both the highest accuracy and fairness under similar latency constrain on multiple edge devices (Samsung S21 mobile phone, iPhone 13 Pro and Raspberry PI).
Changdi Yang, Yi Sheng 0001, Peiyan Dong, Zhenglun Kong, Yanyu Li, Pinrui Yu, Lei Yang 0018, Xue Lin 0001
DAC5
2023 ESRU: Extremely Low-Bit and Hardware-Efficient Stochastic Rounding Unit Design for Low-Bit DNN Training
abstract
Stochastic rounding is crucial in the low-bit (e.g., 8-bit) training of deep neural networks (DNNs) to achieve high accuracy. One of the drawbacks of prior studies is that they require a large number of high-precision stochastic rounding units (SRUs) to guarantee low-bit DNN accuracy, which involves considerable hardware overhead. In this paper, we use extremely low-bit SRUs (ESRUs) to save a large number of hardware resources during low-bit DNN training. However, a naively designed ESRU introduces a biased distribution of random numbers, causing accuracy degradation. To address this issue, we further propose an ESRU design with a plateau-shape distribution. The plateau-shape distribution in our ESRU design is implemented with the combination of an LFSR (linear-feedback shift register) and an inverted LFSR, which avoids LFSR packing and turns an inherent LFSR drawback into an advantage in our efficient ESRU design. Experimental results using state-of-the-art DNN models demonstrate that, compared to the prior 24-bit SRU with 24-bit pseudo-random number generators (PRNG), our 8-bit ESRU with 3-bit PRNG reduces the SRU hardware resource usage by 9.75x while achieving slightly higher accuracy.
Sung-En Chang, Geng Yuan, Alec Lu, Mengshu Sun, Yanyu Li, Zhengang Li 0001, Yanyue Xie, Minghai Qin, Xue Lin 0001, Zhenman Fang, Yanzhi Wang 0001
DATE5
2023 MOC: Multi-Objective Mobile CPU-GPU Co-Optimization for Power-Efficient DNN Inference
abstract
With the emergence of DNN applications on mobile devices, plenty of attention has been attracted to their optimization. However, the impact of DNN inference tasks on device power consumption is still a lack of comprehensive study. In this work, we propose MOC, a Multi-Objective deep reinforcement learning-assisted DNN inference stage-adaptive CPU-GPU Co-optimization approach. We find through experiments that CPU-GPU parameters, including CPU core, CPU, and GPU frequency, could significantly impact the speed and power consumption of DNN inference. We empirically analyze various stages of DNN inference, including pre/post-processing and feed-forward calculating stages. Based on the analysis, a DNN demand-resource matching model is proposed to classify the DNNs into various categories. Next, a multi-objective deep reinforcement learning (MODRL)-assisted framework is proposed, which considers both the DNN type and hardware environment, to make decisions on DNN inference stage-adaptive CPU/GPU parameter tuning. Finally, a rule-based action refinement technique is introduced to tailor the search space of MOC. Extensive experiments show that, compared with existing works, MOC could substantially reduce the power consumption of DNN inference tasks by up to 74.4%, meanwhile delivering an excellent speed on mobile devices.
Yushu Wu, Yifan Gong 0004, Zheng Zhan 0001, Geng Yuan, Yanyu Li, Chao Wu 0006, Yanzhi Wang 0001
ICCAD5
2023 Fast and Fair Medical AI on the Edge Through Neural Architecture Search for Hybrid Vision Models
abstract
As edge devices become readily available and indispensable, there is an urgent need for effective and efficient intelligent applications to be deployed widespread. However, fairness has always been an issue, especially in edge medical applications. Although many approaches have been proposed to mitigate the unfairness problem, their edge performance is not desirable. By examining the fairness performance of different network architectures, we observed that compared to pure convolutional neuron network (CNN) architecture, hybrid models with CNN and Vision Transformer (ViT) have exhibited better performance in terms of fairness and accuracy. After further analyzing the feature maps of intermediate layers of CNNs, ViTs, and hybrid models, we found that ViT has a strong ability to extract global information, which contributes to alleviating the unfairness problem. However, ViTs consume large amounts of computational and memory resources, which hinders their application on edge devices. To address the challenges abovementioned, we propose the first hardware-oriented co-design NAS framework to explore hybrid ViT-CNN architecture for the fair dermatology classification, namely HeViFa, which can produce light-weight models for edge devices with low unfairness scores and high classification accuracy. Experimental results show that compared with FaHaNa-Small, HeViFa-Small could search for a hybrid ViT model that reaches 10.57% and 4.03% higher accuracy as well as 0.179 and 0.0403 higher PQD score on Mix and Fitzpatrick17k dataset, repectively, and speed up by 1.21 × on Samsung S21 mobile phone, 1.18 × on iPhone 13 Pro and 1.37 × on Raspberry Pi.
Changdi Yang, Yi Sheng 0001, Peiyan Dong, Zhenglun Kong, Yanyu Li, Pinrui Yu, Lei Yang 0018, Xue Lin 0001, Yanzhi Wang 0001
ICCAD5
2023 Rethinking Vision Transformers for MobileNet Size and Speed
abstract
With the success of Vision Transformers (ViTs) in computer vision tasks, recent arts try to optimize the performance and complexity of ViTs to enable efficient deployment on mobile devices. Multiple approaches are proposed to accelerate attention mechanism, improve inefficient designs, or incorporate mobile-friendly lightweight convolutions to form hybrid architectures. However, ViT and its variants still have higher latency or considerably more parameters than lightweight CNNs, even true for the years-old MobileNet. In practice, latency and size are both crucial for efficient deployment on resource-constraint hardware. In this work, we investigate a central question, can transformer models run as fast as MobileNet and maintain a similar size? We revisit the design choices of ViTs and propose a novel supernet with low latency and high parameter efficiency. We further introduce a novel fine-grained joint search strategy for transformer models that can find efficient architectures by optimizing latency and number of parameters simultaneously. The proposed models, EfficientFormerV2, achieve 3.5% higher top-1 accuracy than MobileNetV2 on ImageNet-1K with similar latency and parameters. This work demonstrate that properly designed and optimized vision transformers can achieve high performance even with MobileNet-level size and speed1.
Yanyu Li, Ju Hu, Georgios Evangelidis 0002, Kamyar Salahi, Yanzhi Wang 0001, Sergey Tulyakov, Jian Ren 0005
ICCV1
2023 SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds
abstract
Text-to-image diffusion models can create stunning images from natural language descriptions that rival the work of professional artists and photographers. However, these models are large, with complex network architectures and tens of denoising iterations, making them computationally expensive and slow to run. As a result, high-end GPUs and cloud-based inference are required to run diffusion models at scale. This is costly and has privacy implications, especially when user data is sent to a third party. To overcome these challenges, we present a generic approach that, for the first time, unlocks running text-to-image diffusion models on mobile devices in **less than 2 seconds**. We achieve so by introducing efficient network architecture and improving step distillation. Specifically, we propose an efficient UNet by identifying the redundancy of the original model and reducing the computation of the image decoder via data distillation. Further, we enhance the step distillation by exploring training strategies and introducing regularization from classifier-free guidance. Our extensive experiments on MS-COCO show that our model with $8$ denoising steps achieves better FID and CLIP scores than Stable Diffusion v$1.5$ with $50$ steps. Our work democratizes content creation by bringing powerful text-to-image diffusion models to the hands of users.
Yanyu Li, Huan Wang 0014, Qing Jin, Ju Hu, Pavlo Chemerys, Yun Fu 0001, Yanzhi Wang 0001, Sergey Tulyakov, Jian Ren 0005
NeurIPS1
2023 CanMethdb: a database for genome-wide DNA methylation annotation in cancers
abstract
MOTIVATION: DNA methylation within gene body and promoters in cancer cells is well documented. An increasing number of studies showed that cytosine-phosphate-guanine (CpG) sites falling within other regulatory elements could also regulate target gene activation, mainly by affecting transcription factors (TFs) binding in human cancers. This led to the urgent need for comprehensively and effectively collecting distinct cis-regulatory elements and TF-binding sites (TFBS) to annotate DNA methylation regulation. RESULTS: We developed a database (CanMethdb, http://meth.liclab.net/CanMethdb/) that focused on the upstream and downstream annotations for CpG-genes in cancers. This included upstream cis-regulatory elements, especially those involving distal regions to genes, and TFBS annotations for the CpGs and downstream functional annotations for the target genes, computed through integrating abundant DNA methylation and gene expression profiles in diverse cancers. Users could inquire CpG-target gene pairs for a cancer type through inputting a genomic region, a CpG, a gene name, or select hypo/hypermethylated CpG sets. The current version of CanMethdb documented a total of 38 986 060 CpG-target gene pairs (with 6 769 130 unique pairs), involving 385 217 CpGs and 18 044 target genes, abundant cis-regulatory elements and TFs for 33 TCGA cancer types. CanMethdb might help biologists perform in-depth studies of target gene regulations based on DNA methylations in cancer. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/chunquanlipathway/CanMethdb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianmei Zhao, Fengcui Qian, Xuecang Li, Zhengmin Yu, Yanyu Li, Yongsan Yang, Qi Pan, Qiuyu Wang, Jian Zhang 0084, Guohua Wang 0001, Chunquan Li 0002
Bioinform.9
2023 AirNN: Over-the-Air Computation for Neural Networks via Reconfigurable Intelligent Surfaces
abstract
Over-the-air analog computation allows offloading computation to the wireless environment through carefully constructed transmitted signals. In this paper, we design and implement the first-of-its-kind convolution that uses over-the-air computation and demonstrate it for inference tasks in a convolutional neural network (CNN). We engineer the ambient wireless propagation environment through reconfigurable intelligent surfaces (RIS) to design such an architecture, which we call ’AirNN’. AirNN leverages the physics of wave reflection to represent a digital convolution, an essential part of a CNN architecture, in the analog domain. In contrast to classical communication, where the receiver must react to the channel-induced transformation, generally represented as finite impulse response (FIR) filter, AirNN proactively creates the signal reflections to emulate specific FIR filters through RIS. AirNN involves two steps: first, the weights of the neurons in the CNN are drawn from a finite set of channel impulse responses (CIR) that correspond to realizable FIR filters. Second, each CIR is engineered through RIS, and reflected signals combine at the receiver to determine the output of the convolution. This paper presents a proof-of-concept of AirNN by experimentally demonstrating convolutions with over-the-air computation. We then validate the entire resulting CNN model accuracy via simulations for an example task of modulation classification.
Sara Garcia Sanchez, Guillem Reus Muns, Carlos Bocanegra, Yanyu Li, Ufuk Muncuk, M. Yousof Naderi, Yanzhi Wang 0001, Stratis Ioannidis, Kaushik R. Chowdhury
IEEE/ACM Trans. Netw.4
2022 NN-key: A Neural Network-Based Secret Key for Demapping OFDM Symbols
abstract
Generating custom modulation patterns as well as dynamically varying the mapping of the constellation points to their corresponding bit representations are some existing methods for mitigating eavesdropping attacks. In such cases, the custom symbol to bit mapping needs to be conveyed to the receiver through a secure and reliable channel. Instead of sending the representations of the modified symbols in regular information fields, we propose a machine learning-based approach, in which the modified symbols are encoded in the parameters of a light-weight neural network (NN). This NN is trained at the transmitter-side, sent as a secret key to the receiver, where it serves as a demapping block to recover the received symbols correctly. In addition, this paper explores the role of data augmentation during the training stage to increase the robustness of the NN with respect to the noise in the channel, as well as architecture compression to reduce transmission overhead. We validate the robustness of the proposed NN-based custom-modulation demapping approach by comparing it with demapping of a standard scheme (e.g., 16QAM), which reveals no appreciable loss in performance. We further quantitatively analyze the impact of channel and noise impairments on the demapping performance.
Nasim Soltani, Yanyu Li, Deniz Erdogmus, Yanzhi Wang 0001, Kaushik R. Chowdhury
CCNC2
2022 Hardware-efficient stochastic rounding unit design for DNN training: late breaking results
abstract
Stochastic rounding is crucial in the training of low-bit deep neural networks (DNNs) to achieve high accuracy. Unfortunately, prior studies require a large number of high-precision stochastic rounding units (SRUs) to guarantee the low-bit DNN accuracy, which involves considerable hardware overhead. In this paper, we propose an automated framework to explore hardware-efficient low-bit SRUs (ESRUs) that can still generate high-quality random numbers to guarantee the accuracy of low-bit DNN training. Experimental results using state-of-the-art DNN models demonstrate that, compared to the prior 24-bit SRU with 24-bit pseudo random number generator (PRNG), our 8-bit with 3-bit PRNG reduces the SRU resource usage by 9.75× while achieving a higher accuracy.
Sung-En Chang, Geng Yuan, Alec Lu, Mengshu Sun, Yanyu Li, Zhengang Li 0001, Yanyue Xie, Minghai Qin, Xue Lin 0001, Zhenman Fang, Yanzhi Wang 0001
DAC5
2022 FPGA-aware automatic acceleration framework for vision transformer with mixed-scheme quantization: late breaking results
abstract
Vision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand impose urgent needs for new hardware accelerator design methodology. This work proposes an FPGA-aware automatic ViT acceleration framework based on the proposed mixed-scheme quantization. To the best of our knowledge, this is the first FPGA-based ViT acceleration framework exploring model quantization. Compared with state-of-the-art ViT quantization work (algorithmic approach only without hardware acceleration), our quantization achieves 0.31% to 1.25% higher Top-1 accuracy under the same bit-width. Compared with the 32-bit floating-point baseline FPGA accelerator, our accelerator achieves around 5.6× improvement on the frame rate (i.e., 56.4 FPS vs. 10.0 FPS) with 0.83% accuracy drop for DeiT-base.
Mengshu Sun, Zhengang Li 0001, Alec Lu, Geng Yuan, Yanyue Xie, Hao Tang 0005, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin 0001, Zhenman Fang
DAC8
2022 Fault-Tolerant Deep Neural Networks for Processing-In-Memory based Autonomous Edge Systems
abstract
In-memory deep neural network (DNN) accelerators will be the key for energy-efficient autonomous edge systems. The resistive random access memory (ReRAM) is a potential solution for the non-CMOS-based in-memory computing platform for energy-efficient autonomous edge systems, thanks to its promising characteristics, such as near-zero leakage-power and non-volatility. However, due to the hardware instability of ReRAM, the weights of the DNN model may deviate from the originally trained weights, resulting in accuracy loss. To mitigate this undesirable accuracy loss, we propose two stochastic fault-tolerant training methods to generally improve the models' robustness without dealing with individual devices. Moreover, we propose Stability Score-a comprehensive metric that serves as an indicator to the instability problem. Extensive experiments demonstrate that the DNN models trained using our proposed stochastic fault-tolerant training method achieve superior performance, which provides better flexibility, scalability, and deployability of ReRAM on the autonomous edge systems.
Siyue Wang, Geng Yuan, Yanyu Li, Xue Lin 0001, Bhavya Kailkhura
DATE4
2022 Compiler-Aware Neural Architecture Search for On-Mobile Real-time Super-Resolution
Yushu Wu, Yifan Gong 0004, Pu Zhao 0001, Yanyu Li, Zheng Zhan 0001, Wei Niu 0002, Hao Tang 0005, Minghai Qin, Bin Ren 0002, Yanzhi Wang 0001
ECCV (19)4
2022 You Already Have It: A Generator-Free Low-Precision DNN Training Framework Using Stochastic Rounding
Geng Yuan, Sung-En Chang, Qing Jin, Alec Lu, Yanyu Li, Yushu Wu, Zhenglun Kong, Yanyue Xie, Peiyan Dong, Minghai Qin, Xulong Tang, Zhenman Fang, Yanzhi Wang 0001
ECCV (12)5
2022 FILM-QNN: Efficient FPGA Acceleration of Deep Neural Networks with Intra-Layer, Mixed-Precision Quantization
abstract
With the trend to deploy Deep Neural Network (DNN) inference models on edge devices with limited resources, quantization techniques have been widely used to reduce on-chip storage and improve computation throughput. However, existing DNN quantization work deploying quantization below 8-bit may be either suffering from evident accuracy loss or facing a big gap between the theoretical improvement of computation throughput and the practical inference speedup. In this work, we propose a general framework, called FILM-QNN, to quantize and accelerate multiple DNN models across different embedded FPGA devices. First, we propose the novel intra-layer, mixed-precision quantization algorithm that assigns different precisions onto the filters of each layer. The candidate precision levels and assignment granularity are determined from our empirical study with the capability of preserving accuracy and improving hardware parallelism. Second, we apply multiple optimization techniques for the FPGA accelerator architecture in support of quantized computations, including DSP packing, weight reordering, and data packing, to enhance the overall throughput with the available resources. Moreover, a comprehensive resource model is developed to balance the allocation of FPGA computation resources (LUTs and DSPs) as well as data transfer and on-chip storage resources (BRAMs) to accelerate the computations in mixed precisions within each layer. Finally, to improve the portability of FILM-QNN, we implement it using Vivado High-Level Synthesis (HLS) on Xilinx PYNQ-Z2 and ZCU102 FPGA boards. Our experimental results of ResNet-18, ResNet-50, and MobileNet-V2 demonstrate that the implementations with intra-layer, mixed-precision (95% of 4-bit weights and 5% of 8-bit weights, and all 5-bit activations) can achieve comparable accuracy (70.47%, 77.25%, and 65.67% for the three models) as the 8-bit (and 32-bit) versions and comparable throughput (214.8 FPS, 109.1 FPS, and 537.9 FPS on ZCU102) as the 4-bit designs.
Mengshu Sun, Zhengang Li 0001, Alec Lu, Yanyu Li, Sung-En Chang, Xue Lin 0001, Zhenman Fang
FPGA4
2022 Auto-ViT-Acc: An FPGA-Aware Automatic Acceleration Framework for Vision Transformer with Mixed-Scheme Quantization
abstract
Vision transformers (ViTs) are emerging with significantly improved accuracy in computer vision tasks. However, their complex architecture and enormous computation/storage demand impose urgent needs for new hardware accelerator design methodology. This work proposes an FPGA-aware automatic ViT acceleration framework based on the proposed mixed-scheme quantization. To the best of our knowledge, this is the first FPGA-based ViT acceleration framework exploring model quantization. Compared with state-of-the-art ViT quantization work (algorithmic approach only without hardware acceleration), our quantization achieves 0.47% to 1.36% higher Top-l accuracy under the same bit-width. Compared with the 32-bit floating-point baseline FPGA accelerator, our accelerator achieves around 5.6x improvement on the frame rate (i.e., 56.8 FPS vs. 10.0 FPS) with 0.71% accuracy drop on ImageNet dataset for DeiT-base.
Zhengang Li 0001, Mengshu Sun, Alec Lu, Geng Yuan, Yanyue Xie, Hao Tang 0005, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin 0001, Zhenman Fang
FPL8
2022 Real-Time Portrait Stylization on the Edge
abstract
In this work we demonstrate real-time portrait stylization, specifically, translating self-portrait into cartoon or anime style on mobile devices. We propose a latency-driven differentiable architecture search method, maintaining realistic generative quality. With our framework, we obtain 10× computation reduction on the generative model and achieve real-time video stylization on off-the-shelf smartphone using mobile GPUs.
Yanyu Li, Xuan Shen, Geng Yuan, Jiexiong Guan, Wei Niu 0002, Hao Tang 0005, Bin Ren 0002, Yanzhi Wang 0001
IJCAI1
2022 Pruning-as-Search: Efficient Neural Architecture Search via Channel Pruning and Structural Reparameterization
abstract
Neural architecture search (NAS) and network pruning are widely studied efficient AI techniques, but not yet perfect. NAS performs exhaustive candidate architecture search, incurring tremendous search cost. Though (structured) pruning can simply shrink model dimension, it remains unclear how to decide the per-layer sparsity automatically and optimally. In this work, we revisit the problem of layer-width optimization and propose Pruning-as-Search (PaS), an end-to-end channel pruning method to search out desired sub-network automatically and efficiently. Specifically, we add a depth-wise binary convolution to learn pruning policies directly through gradient descent. By combining the structural reparameterization and PaS, we successfully searched out a new family of VGG-like and lightweight networks, which enable the flexibility of arbitrary width with respect to each layer instead of each stage. Experimental results show that our proposed architecture outperforms prior arts by around 1.0% top-1 accuracy under similar inference speed on ImageNet-1000 classification task. Furthermore, we demonstrate the effectiveness of our width search on complex tasks including instance segmentation and image translation. Code and models are released.
Yanyu Li, Pu Zhao 0001, Geng Yuan, Xue Lin 0001, Yanzhi Wang 0001
IJCAI1
2022 EfficientFormer: Vision Transformers at MobileNet Speed
abstract
Vision Transformers (ViT) have shown rapid progress in computer vision tasks, achieving promising results on various benchmarks. However, due to the massive number of parameters and model design, e.g., attention mechanism, ViT-based models are generally times slower than lightweight convolutional networks. Therefore, the deployment of ViT for real-time applications is particularly challenging, especially on resource-constrained hardware such as mobile devices. Recent efforts try to reduce the computation complexity of ViT through network architecture search or hybrid design with MobileNet block, yet the inference speed is still unsatisfactory. This leads to an important question: can transformers run as fast as MobileNet while obtaining high performance? To answer this, we first revisit the network architecture and operators used in ViT-based models and identify inefficient designs. Then we introduce a dimension-consistent pure transformer (without MobileNet blocks) as a design paradigm. Finally, we perform latency-driven slimming to get a series of final models dubbed EfficientFormer. Extensive experiments show the superiority of EfficientFormer in performance and speed on mobile devices. Our fastest model, EfficientFormer-L1, achieves $79.2\%$ top-1 accuracy on ImageNet-1K with only $1.6$ ms inference latency on iPhone 12 (compiled with CoreML), which runs as fast as MobileNetV2$\times 1.4$ ($1.6$ ms, $74.7\%$ top-1), and our largest model, EfficientFormer-L7, obtains $83.3\%$ accuracy with only $7.0$ ms latency. Our work proves that properly designed transformers can reach extremely low latency on mobile devices while maintaining high performance.
Yanyu Li, Geng Yuan, Ju Hu, Georgios Evangelidis 0002, Sergey Tulyakov, Yanzhi Wang 0001, Jian Ren 0005
NeurIPS1
2022 Layer Freezing & Data Sieving: Missing Pieces of a Generic Framework for Sparse Training
abstract
Recently, sparse training has emerged as a promising paradigm for efficient deep learning on edge devices. The current research mainly devotes the efforts to reducing training costs by further increasing model sparsity. However, increasing sparsity is not always ideal since it will inevitably introduce severe accuracy degradation at an extremely high sparsity level. This paper intends to explore other possible directions to effectively and efficiently reduce sparse training costs while preserving accuracy. To this end, we investigate two techniques, namely, layer freezing and data sieving. First, the layer freezing approach has shown its success in dense model training and fine-tuning, yet it has never been adopted in the sparse training domain. Nevertheless, the unique characteristics of sparse training may hinder the incorporation of layer freezing techniques. Therefore, we analyze the feasibility and potentiality of using the layer freezing technique in sparse training and find it has the potential to save considerable training costs. Second, we propose a data sieving method for dataset-efficient training, which further reduces training costs by ensuring only a partial dataset is used throughout the entire training process. We show that both techniques can be well incorporated into the sparse training algorithm to form a generic framework, which we dub SpFDE. Our extensive experiments demonstrate that SpFDE can significantly reduce training costs while preserving accuracy from three dimensions: weight sparsity, layer freezing, and dataset sieving. Our code and models will be released.
Geng Yuan, Yanyu Li, Sheng Li 0019, Zhenglun Kong, Sergey Tulyakov, Xulong Tang, Yanzhi Wang 0001, Jian Ren 0005
NeurIPS2
2022 GREAP: a comprehensive enrichment analysis software for human genomic regions
abstract
The rapid development of genomic high-throughput sequencing has identified a large number of DNA regulatory elements with abundant epigenetics markers, which promotes the rapid accumulation of functional genomic region data. The comprehensively understanding and research of human functional genomic regions is still a relatively urgent work at present. However, the existing analysis tools lack extensive annotation and enrichment analytical abilities for these regions. Here, we designed a novel software, Genomic Region sets Enrichment Analysis Platform (GREAP), which provides comprehensive region annotation and enrichment analysis capabilities. Currently, GREAP supports 85 370 genomic region reference sets, which cover 634 681 107 regions across 11 different data types, including super enhancers, transcription factors, accessible chromatins, etc. GREAP provides widespread annotation and enrichment analysis of genomic regions. To reflect the significance of enrichment analysis, we used the hypergeometric test and also provided a Locus Overlap Analysis. In summary, GREAP is a powerful platform that provides many types of genomic region sets for users and supports genomic region annotations and enrichment analyses. In addition, we developed a customizable genome browser containing >400 000 000 customizable tracks for visualization. The platform is freely available at http://www.liclab.net/Greap/view/index.
Yongsan Yang, Fengcui Qian, Xuecang Li, Yanyu Li, Liwei Zhou, Qiuyu Wang, Xinyuan Zhou, Jian Zhang 0084, Zhengmin Yu, Ting Cui, Chenchen Feng, Desi Shang, Mengfei Sun, Yuexin Zhang, Huifang Tang, Chunquan Li 0002
Briefings Bioinform.4
2022 Automated deep learning-based wide-band receiver
Bahar Azari, Hai Cheng, Nasim Soltani, Haoqing Li 0001, Yanyu Li, Mauro Belgiovine, Tales Imbiriba, Salvatore D'Oro, Tommaso Melodia, Yanzhi Wang 0001, Pau Closas, Kaushik R. Chowdhury, Deniz Erdogmus
Comput. Networks5
2022 Mobile or FPGA? A Comprehensive Evaluation on Energy Efficiency and a Unified Optimization Framework
abstract
Efficient deployment of Deep Neural Networks (DNNs) on edge devices (i.e., FPGAs and mobile platforms) is very challenging, especially under a recent witness of the increasing DNN model size and complexity. Model compression strategies, including weight quantization and pruning, are widely recognized as effective approaches to significantly reduce computation and memory intensities, and have been implemented in many DNNs on edge devices. However, most state-of-the-art works focus on ad hoc optimizations, and there lacks a thorough study to comprehensively reveal the potentials and constraints of different edge devices when considering different compression strategies. In this article, we qualitatively and quantitatively compare the energy efficiency of FPGA-based and mobile-based DNN executions using mobile GPU and provide a detailed analysis. Based on the observations obtained from the analysis, we propose a unified optimization framework using block-based pruning to reduce the weight storage and accelerate the inference speed on mobile devices and FPGAs, achieving high hardware performance and energy-efficiency gain while maintaining accuracy.
Geng Yuan, Peiyan Dong, Mengshu Sun, Wei Niu 0002, Zhengang Li 0001, Yuxuan Cai 0001, Yanyu Li, Jun Liu 0075, Weiwen Jiang, Xue Lin 0001, Bin Ren 0002, Xulong Tang, Yanzhi Wang 0001
ACM Trans. Embed. Comput. Syst.7
2021 YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
abstract
The rapid development and wide utilization of object detection techniques have aroused attention on both accuracy and speed of object detectors. However, the current state-of-the-art object detection works are either accuracy-oriented using a large model but leading to high latency or speed-oriented using a lightweight model but sacrificing accuracy. In this work, we propose YOLObile framework, a real-time object detection on mobile devices via compression-compilation co-design. A novel block-punched pruning scheme is proposed for any kernel size. To improve computational efficiency on mobile devices, a GPU-CPU collaborative scheme is adopted along with advanced compiler-assisted optimizations. Experimental results indicate that our pruning scheme achieves 14x compression rate of YOLOv4 with 49.0 mAP. Under our YOLObile framework, we achieve 17 FPS inference speed using GPU on Samsung Galaxy S20. By incorporating our proposed GPU-CPU collaborative scheme, the inference speed is increased to 19.1 FPS, and outperforms the original YOLOv4 by 5x speedup. Source code is at: https://github.com/nightsnack/YOLObile.
Yuxuan Cai 0001, Hongjia Li 0003, Geng Yuan, Wei Niu 0002, Yanyu Li, Xulong Tang, Bin Ren 0002, Yanzhi Wang 0001
AAAI5
2021 A Compression-Compilation Co-Design Framework Towards Real-Time Object Detection on Mobile Devices
abstract
The rapid development and wide utilization of object detection techniques have aroused requirements for both accuracy and speed of object detectors. In this work, we propose a compression-compilation co-design framework to achieve real-time YOLOv4 inference on mobile devices. We propose a novel fine-grained structured pruning, which maintain high accuracy while achieving high hardware parallelism. Our pruned YOLOv4 achieves 48.9 mAP and 17 FPS inference speed on an off-the-shelf Samsung Galaxy S20 smartphone, which is 5.5x faster than the original state-of-the-art detector YOLOv4.
Yuxuan Cai 0001, Geng Yuan, Hongjia Li 0003, Wei Niu 0002, Yanyu Li, Xulong Tang, Bin Ren 0002, Yanzhi Wang 0001
AAAI5
2021 NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile Acceleration
abstract
With the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Prior methods towards this goal, including model compression and network architecture search (NAS), are largely performed independently, and do not fully consider compiler-level optimizations which is a must-do for mobile acceleration. In this work, we first propose (i) a general category of fine-grained structured pruning applicable to various DNN layers, and (ii) a comprehensive, compiler automatic code generation framework supporting different DNNs and different pruning schemes, which bridge the gap of model compression and NAS. We further propose NPAS, a compiler-aware unified network pruning and architecture search. To deal with large search space, we propose a meta-modeling procedure based on reinforcement learning with fast evaluation and Bayesian optimization, ensuring the total number of training epochs comparable with representative NAS frameworks. Our framework achieves 6.7ms, 5.9ms, and 3.9ms ImageNet inference times with 78.2%, 75% (MobileNet-V3 level), and 71% (MobileNet-V2 level) Top-1 accuracy respectively on an off-the-shelf mobile phone, consistently outperforming prior work.
Zhengang Li 0001, Geng Yuan, Wei Niu 0002, Pu Zhao 0001, Yanyu Li, Yuxuan Cai 0001, Xuan Shen, Zheng Zhan 0001, Zhenglun Kong, Qing Jin, Zhiyu Chen 0003, Sijia Liu 0001, Kaiyuan Yang 0001, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001
CVPR5
2021 Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
abstract
Deep Neural Networks (DNNs) have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference on edge-computing platforms, e.g., ASICs, FPGAs, and embedded systems, are extensively investigated. Due to the huge model size and computation amount, model compression is a critical step to deploy DNN models on edge devices. This paper focuses on weight quantization, a hardware-friendly model compression approach that is complementary to weight pruning.Unlike existing methods that use the same quantization scheme for all weights, we propose the first solution that applies different quantization schemes for different rows of the weight matrix. It is motivated by (1) the distribution of the weights in the different rows are not the same; and (2) the potential of achieving better utilization of heterogeneous FPGA hardware resources. To achieve that, we first propose a hardware-friendly quantization scheme named sum-of-power-of-2 (SP2) suitable for Gaussian-like weight distribution, in which the multiplication arithmetic can be replaced with logic shifter and adder, thereby enabling highly efficient implementations with the FPGA LUT resources. In contrast, the existing fixed-point quantization is suitable for Uniform-like weight distribution and can be implemented efficiently by DSP. Then to fully explore the resources, we propose an FPGA-centric mixed scheme quantization (MSQ) with an ensemble of the proposed SP2 and the fixed-point schemes. Combining the two schemes can maintain, or even increase accuracy due to better matching with weight distributions.For the FPGA implementations, we develop a parameterized architecture with heterogeneous Generalized Matrix Multiplication (GEMM) cores-one using LUTs for computations with SP2 quantized weights and the other utilizing DSPs for fixed-point quantized weights. Given the partition ratio among the two schemes based on resource characterization, MSQ quantization training algorithm derives an optimally quantized model for the FPGA implementation. We evaluate our FPGA-centric quantization framework across multiple application domains. With optimal SP2/fixed-point ratios on two FPGA devices, i.e., Zynq XC7Z020 and XC7Z045, we achieve performance improvement of 2.1 × -4.1 × compared to solely exploiting DSPs for all multiplication operations. In addition, the CNN implementations with the proposed MSQ scheme can achieve higher accuracy and comparable hardware utilization efficiency compared to the state-of-the-art designs.
Sung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi, Hayden Kwok-Hay So, Xuehai Qian, Yanzhi Wang 0001, Xue Lin 0001
HPCA2
2021 RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
abstract
This work proposes a novel Deep Neural Network (DNN) quantization framework, namely RMSMP, with a Row-wise Mixed-Scheme and Multi-Precision approach. Specifically, this is the first effort to assign mixed quantization schemes and multiple precisions within layers – among rows of the DNN weight matrix, for simplified operations in hardware inference, while preserving accuracy. Furthermore, this paper makes a different observation from the prior work that the quantization error does not necessarily exhibit the layer-wise sensitivity, and actually can be mitigated as long as a certain portion of the weights in every layer are in higher precisions. This observation enables layer-wise uniformality in the hardware implementation towards guaranteed inference acceleration, while still enjoying row-wise flexibility of mixed schemes and multiple precisions to boost accuracy. The candidates of schemes and precisions are derived practically and effectively with a highly hardware-informative strategy to reduce the problem search space.With the offline determined ratio of different quantization schemes and precisions for all the layers, the RMSMP quantization algorithm uses Hessian and variance based method to effectively assign schemes and precisions for each row. The proposed RMSMP is tested for the image classification and natural language processing (BERT) applications, and achieves the best accuracy performance among state-of-the-arts under the same equivalent precisions. The RMSMP is implemented on FPGA devices, achieving 3.65× speedup in the end-to-end inference time for ResNet-18 on ImageNet, comparing with the 4-bit Fixed-point baseline.
Sung-En Chang, Yanyu Li, Mengshu Sun, Weiwen Jiang, Sijia Liu 0001, Yanzhi Wang 0001, Xue Lin 0001
ICCV2
2021 TRlnc: a comprehensive database for human transcriptional regulatory information of lncRNAs
abstract
Long noncoding RNAs (lncRNAs) have been proven to play important roles in transcriptional processes and biological functions. With the increasing study of human diseases and biological processes, information in human H3K27ac ChIP-seq, ATAC-seq and DNase-seq datasets is accumulating rapidly, resulting in an urgent need to collect and process data to identify transcriptional regulatory regions of lncRNAs. We therefore developed a comprehensive database for human regulatory information of lncRNAs (TRlnc, http://bio.licpathway.net/TRlnc), which aimed to collect available resources of transcriptional regulatory regions of lncRNAs and to annotate and illustrate their potential roles in the regulation of lncRNAs in a cell type-specific manner. The current version of TRlnc contains 8 683 028 typical enhancers/super-enhancers and 32 348 244 chromatin accessibility regions associated with 91 906 human lncRNAs. These regions are identified from over 900 human H3K27ac ChIP-seq, ATAC-seq and DNase-seq samples. Furthermore, TRlnc provides the detailed genetic and epigenetic annotation information within transcriptional regulatory regions (promoter, enhancer/super-enhancer and chromatin accessibility regions) of lncRNAs, including common SNPs, risk SNPs, eQTLs, linkage disequilibrium SNPs, transcription factors, methylation sites, histone modifications and 3D chromatin interactions. It is anticipated that the use of TRlnc will help users to gain in-depth and useful insights into the transcriptional regulatory mechanisms of lncRNAs.
Yanyu Li, Xuecang Li, Yongsan Yang, Fengcui Qian, Zhidong Tang, Jianmei Zhao, Jian Zhang 0084, Xuefeng Bai 0002, Yong Jiang 0004, Jianyuan Zhou, Yuexin Zhang, Liwei Zhou, Jianjun Xie, Enmin Li, Qiuyu Wang, Chunquan Li 0002
Briefings Bioinform.1
2019 TRCirc: a resource for transcriptional regulation information of circRNAs
abstract
In recent years, high-throughput genomic technologies like chromatin immunoprecipitation sequencing (ChIp-seq) and transcriptome sequencing (RNA-seq) have been becoming both more refined and less expensive, making them more accessible. Many circular RNAs (circRNAs) that originate from back-spliced exons have been identified in various cell lines across different species. However, the regulatory mechanism for transcription of circRNAs remains unclear. Therefore, there is an urgent need to construct a database detailing the transcriptional regulation of circRNAs. TRCirc (http://www.licpathway.net/TRCirc) provides a resource for efficient retrieval, browsing and visualization of transcriptional regulation information of circRNAs. The current version of TRCirc documents 92 375 circRNAs and 161 transcription factors (TFs) from more than 100 cell types and together represent more than 765 000 TF-circRNA regulatory relationships. Furthermore, TRCirc provides other regulatory information about transcription of circRNAs, including their expression, methylation levels, H3K27ac signals in regulation regions and super-enhancers associated with circRNAs. TRCirc provides a convenient, user-friendly interface to search, browse and visualize detailed information about these circRNAs.
Zhidong Tang, Xuecang Li, Jianmei Zhao, Fengcui Qian, Chenchen Feng, Yanyu Li, Jian Zhang 0084, Yong Jiang 0004, Yongsan Yang, Qiuyu Wang, Chunquan Li 0002
Briefings Bioinform.6