Yang Sui 0001

dblp:77/10522-1 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
26since 2021 · last 2025
0000-0003-3020-0612ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 iTAP: An Incremental Task Graph Partitioner for Task-parallel Static Timing Analysis
abstract
Recent static timing analysis (STA) tools have utilized task dependency graph (TDG) parallelism to enhance the STA runtime performance. Although TDG parallelism shows promising speedup, the overhead of scheduling a TDG can become dominant as the TDG becomes larger. To minimize the scheduling overhead, several TDG partitioning algorithms have been proposed to reduce the TDG size without affecting its task parallelism. Despite improved performance, existing TDG partitioners all fall short of incremental partitioning, limiting their practical use in STA tools that support timing-driven operations. To overcome this limitation, we propose iTAP, an incremental TDG partitioner to fully leverage the power of TDG partitioning in task-parallel STA applications. Compared to a state-of-the-art full TDG partitioner, iTAP enhances the overall STA performance by up to 2.97×.
Boyang Zhang 0007, Che Chang, Cheng-Hsiang Chiu, Dian-Lun Lin, Yang Sui 0001, Chih-Chun Chang, Yi-Hua Chung, Wan-Luan Lee, Zizheng Guo 0001, Yibo Lin, Tsung-Wei Huang
ASP-DAC5
2025 DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
abstract
Video large language models (VLLMs) have significantly advanced recently in processing complex video content. Yet, their inference efficiency remains constrained because of the high computational cost stemming from the thousands of visual tokens generated from the video inputs. We empirically observe that, unlike single image inputs, VLLMs typically attend visual tokens from different frames at different decoding iterations. This makes a one-shot pruning strategy prone to removing important tokens by mistake. Motivated by this, we present DyCoke, a training-free token compression method to optimize token representation and accelerate VLLMs. DyCoke incorporates a plug-and-play temporal compression module to minimize temporal redundancy by merging redundant tokens across frames and applying dynamic KV cache reduction to prune spatially redundant tokens selectively. It ensures high-quality inference by dynamically retaining the critical tokens at each decoding step. Extensive experimental results demonstrate that DyCoke can outperform the prior SoTA counterparts, achieving 1.5× inference speedup, and 1.4× memory reduction against the baseline VLLM, while still improving the performance, with no training.
Keda Tao, Can Qin, Haoxuan You, Yang Sui 0001, Huan Wang 0014
CVPR4
2025 SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
abstract
We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality. Project page at https://snap-research.github.io/snapgen-v/.
Yushu Wu, Yanyu Li, Yanwu Xu 0003, Anil Kag, Yang Sui 0001, Huseyin Coskun, Aleksei Lebedev, Ju Hu, Dimitris N. Metaxas, Yanzhi Wang 0001, Sergey Tulyakov, Jian Ren 0005
CVPR6
2025 TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
abstract
Vision-Language Models (VLMs) demand substantial computational resources during inference, largely due to the extensive visual input tokens for representing visual information. Previous studies have noted that visual tokens tend to receive less attention than text tokens, suggesting their lower importance during inference and potential for pruning. However, their methods encounter several challenges: reliance on greedy heuristic criteria for token importance and incompatibility with FlashAttention and KV cache. To address these issues, we introduce TopV, a compatible TOken Pruning with inference Time Optimization for fast and low-memory VLM, achieving efficient pruning without additional training or fine-tuning. Instead of relying on attention scores, we formulate token pruning as an optimization problem, accurately identifying important visual tokens while remaining compatible with FlashAttention. Additionally, since we only perform this pruning once during the prefilling stage, it effectively reduces KV cache size. Our optimization framework incorporates a visual-aware cost function considering factors such as Feature Similarity, Relative Spatial Distance, and Absolute Central Distance, to measure the importance of each source visual token, enabling effective pruning of low-importance tokens. Extensive experiments demonstrate that our method outperforms previous token pruning methods, validating the effectiveness and efficiency of our approach.
Cheng Yang 0013, Yang Sui 0001, Jinqi Xiao, Lingyi Huang, Yu Gong 0003, Chendi Li, Jinghua Yan, Yu Bai 0009, P. Sadayappan, Xia Ben Hu, Bo Yuan 0001
CVPR2
2025 HoliTom: Holistic Token Merging for Fast Video Large Language Models
abstract
Video large language models (video LLMs) excel at video comprehension but face significant computational inefficiency due to redundant video tokens. Existing token pruning methods offer solutions. However, approaches operating within the LLM (inner-LLM pruning), such as FastV, incur intrinsic computational overhead in shallow layers. In contrast, methods performing token pruning before the LLM (outer-LLM pruning) primarily address spatial redundancy within individual frames or limited temporal windows, neglecting the crucial global temporal dynamics and correlations across longer video sequences. This leads to sub-optimal spatio-temporal reduction and does not leverage video compressibility fully. Crucially, the synergistic potential and mutual influence of combining these strategies remain unexplored. To further reduce redundancy, we introduce HoliTom, a novel training-free holistic token merging framework. HoliTom employs outer-LLM pruning through global redundancy-aware temporal segmentation, followed by spatial-temporal merging to reduce visual tokens by over 90%, significantly alleviating the LLM's computational burden. Complementing this, we introduce a robust inner-LLM token similarity-based merging approach, designed for superior performance and compatibility with outer-LLM pruning. Evaluations demonstrate our method's promising efficiency-performance trade-off on LLaVA-OneVision-7B, reducing computational costs to 6.9% of FLOPs while maintaining 99.1% of the original performance. Furthermore, we achieve a 2.28× reduction in Time-To-First-Token (TTFT) and a 1.32× acceleration in decoding throughput, highlighting the practical benefits of our integrated pruning approach for efficient video LLMs inference.
Kele Shao, Keda Tao, Can Qin, Haoxuan You, Yang Sui 0001, Huan Wang 0014
NeurIPS5
2025 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
abstract
Large-scale AI models, such as Large Language Models (LLMs) and Diffusion Models (DMs), have grown rapidly in size, creating significant challenges for efficient deployment on resource-constrained hardware. In this paper, we introduce Dynamic-Length Float (DFloat11), a lossless compression framework that reduces LLM and DM size by 30\% while preserving outputs that are bit-for-bit identical to the original model. DFloat11 is motivated by the low entropy in the BFloat16 weight representation of LLMs, which reveals significant inefficiency in the existing storage format. By applying entropy coding, DFloat11 assigns dynamic-length encodings to weights based on frequency, achieving near information-optimal compression without any loss of precision. To facilitate efficient inference with dynamic-length encodings, we develop a custom GPU kernel for fast online decompression. Our design incorporates the following: (i) compact, hierarchical lookup tables (LUTs) that fit within GPU SRAM for efficient decoding, (ii) a two-phase GPU kernel for coordinating thread read/write positions using lightweight auxiliary variables, and (iii) transformer-block-level decompression to minimize latency. Experiments on Llama 3.3, Qwen 3, Mistral 3, FLUX.1, and others validate our hypothesis that DFloat11 achieves around 30\% model size reduction while preserving bit-for-bit identical outputs. Compared to a potential alternative of offloading parts of an uncompressed model to the CPU to meet memory constraints, DFloat11 achieves 2.3--46.2$\times$ higher throughput in token generation. With a fixed GPU memory budget, DFloat11 enables 5.7--14.9$\times$ longer generation lengths than uncompressed models. Notably, our method enables lossless inference of Llama 3.1 405B, an 810GB model, on a single node equipped with 8$\times$80GB GPUs.
Tianyi Zhang 0011, Mohsen Hariri, Shaochen Zhong, Vipin Chaudhary, Yang Sui 0001, Xia Ben Hu, Anshumali Shrivastava
NeurIPS5
2025 Co-Exploring Structured Sparsification and Low-Rank Tensor Decomposition for Compact DNNs
abstract
Sparsification and low-rank decomposition are two important techniques to compress deep neural network (DNN) models. To date, these two popular yet distinct approaches are typically used in separate ways; while their efficient integration for better compression performance is little explored, especially for structured sparsification and decomposition. In this article, we perform systematic co-exploration on structured sparsification and decomposition toward compact DNN models. We first investigate and analyze several important design factors for joint structured sparsification and decomposition, including operational sequence, decomposition format, and optimization procedure. Based on the observations from our analysis, we then propose CEPD, a unified DNN compression framework that can co-explore the benefits of structured sparsification and tensor decomposition in an efficient way. Empirical experiments demonstrate the promising performance of our proposed solution. Notably, on the CIFAR-10 dataset, CEPD brings 0.72%-0.45% accuracy increase over the baseline ResNet-56 and MobileNetV2 models, respectively, and meanwhile, the computational costs are reduced by 43.0%-44.2%, respectively. On the ImageNet dataset, our approach can enable 0.10%-1.39% accuracy increase over the baseline ResNet-18 and ResNet-50 models with 59.4%-54.6% fewer parameters, respectively.
Yang Sui 0001, Miao Yin, Yu Gong 0003, Bo Yuan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Transferable Learned Image Compression-Resistant Adversarial Perturbations
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
BMVC1
2024 Invited: Algorithm and Hardware Co-Design for Energy-Efficient Neural SLAM
abstract
In this paper, we introduce a novel approach to enhancing neural network-based Simultaneous Localization and Mapping (SLAM) through the integration of model compression techniques and customized hardware architecture that focuses on micro-architectural and dataflow optimizations to improve computational efficiency and performance. Experiments across different scenarios demonstrate that the proposed approach achieves significant improvement.
Lingyi Huang, Cheng Yang 0013, Yu Gong 0003, Yang Sui 0001, Xiao Zang, Anthony Goeckner, Qi Zhu 0002, Bo Yuan 0001
DAC4
2024 Transferable Learned Image Compression-Resistant Adversarial Perturbations
abstract
With the rapid evolution of advanced image compression, DNN-based learned image compression has emerged as the promising approach for transmitting images in many security-critical applications, such as cloud-based face recognition and autonomous driving, due to its superior performance over traditional compression. There is a pressing need to fully investigate the robustness of a classification system post-processed by learned image compression. To bridge this research gap, we explore the adversarial attack on Learned Image Compression Classification System (LICCS) that targets image classification models that utilize learned image compressors as preprocessing modules. To perform an adversarial attack on an image within the LICCS, the goal is to introduce the adversarial perturbation δ to the source image X that causes the reconstructed adversarial examples gs(Q(ga(X+δ))) to be misclassified by the classification model, which can be formulated as follows:\begin{equation*}\begin{array}{ll} {\mathop {\arg \max }\limits_i f{{\left({{g_s}\left({Q\left({{g_a}\left({{\mathbf{X + \delta }}}\right)}\right)}\right)}\right)}_i} \ne y,}&{{\text{s}}{\text{.t}}{\text{.}}\parallel \delta {\parallel _p} \leq \varepsilon .} \end{array}\tag{1}\end{equation*}
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
DCC1
2024 Reconstruction Distortion of Learned Image Compression with Imperceptible Perturbations
abstract
In this paper, we introduce an imperceptible adversarial attack approach designed to effectively degrade the reconstruction quality of LIC, resulting in the reconstructed image being severely disrupted by noise where identifying any object in the reconstructed image is virtually impossible. More specifically, we generate adversarial examples by introducing a Frobenius norm-based loss function to maximize the discrepancy between original images and reconstructed images from adversarial examples in order to corrupt the reconstructed image severely.
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
DCC1
2024 Clean and Compact: Efficient Data-Free Backdoor Defense with Model Compactness
Huy Phan, Jinqi Xiao, Yang Sui 0001, Tianfang Zhang, Zijie Tang, Cong Shi 0004, Yan Wang 0003, Yingying Chen 0001, Bo Yuan 0001
ECCV (60)3
2024 MOPED: Efficient Motion Planning Engine with Flexible Dimension Support
abstract
Motion planning aims to compute the high-quality and collision-free robotic trajectory. To solve the planning problems defined in varying dimensional sizes, motion planners, especially sampling-based, are typically computation intensive because of the costly kernel operations, and computation inefficient due to the inherent sequential processing scheme, hindering their efficient deployment. To address these challenges and enable real-time highly efficient motion planning, this paper proposes MOPED, an algorithm and hardware co-design for sampling-based motion planning engine with flexible dimension support. At the algorithm level, MOPED proposes a two-stage processing scheme to reduce the frequency and unit cost of collision check. It also fully leverages the spatial information and unique property of planning process to enable low-cost approximated neighbor search. At the hardware level, MOPED proposes a correctness-ensured speculative processing scheme to overcome the serialization problem. It also develop a multi-level caching strategy to reduce data movement and resolve resource conflict. We demonstrate the effectiveness of MOPED via implementing a design example with CMOS 28nm technology via synthesizing. Compared with the baseline motion planning processors, MOPED brings significant improvement on throughput, energy efficiency and area efficiency.
Lingyi Huang, Yu Gong 0003, Yang Sui 0001, Xiao Zang, Bo Yuan 0001
HPCA3
2024 BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
abstract
Diffusion-based image generation models have achieved great success in recent years by showing the capability of synthesizing high-quality content. However, these models contain a huge number of parameters, resulting in a significantly large model size. Saving and transferring them is a major bottleneck for various applications, especially those running on resource-constrained devices. In this work, we develop a novel weight quantization method that quantizes the UNet from Stable Diffusion v1.5 to $1.99$ bits, achieving a model with $7.9\times$ smaller size while exhibiting even better generation quality than the original one. Our approach includes several novel techniques, such as assigning optimal bits to each layer, initializing the quantized model for better performance, and improving the training strategy to dramatically reduce quantization error. Furthermore, we extensively evaluate our quantized model across various benchmark datasets and through human evaluation to demonstrate its superior generation quality.
Yang Sui 0001, Yanyu Li, Anil Kag, Yerlan Idelbayev, Junli Cao, Ju Hu, Dhritiman Sagar, Bo Yuan 0001, Sergey Tulyakov, Jian Ren 0005
NeurIPS1
2024 Corner-to-Center long-range context model for efficient learned image compression
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Bo Yuan 0001, Zhenzhong Chen 0001
J. Vis. Commun. Image Represent.1
2023 CSTAR: Towards Compact and Structured Deep Neural Networks with Adversarial Robustness
abstract
Model compression and model defense for deep neural networks (DNNs) have been extensively and individually studied. Considering the co-importance of model compactness and robustness in practical applications, several prior works have explored to improve the adversarial robustness of the sparse neural networks. However, the structured sparse models obtained by the existing works suffer severe performance degradation for both benign and robust accuracy, thereby causing a challenging dilemma between robustness and structuredness of compact DNNs. To address this problem, in this paper, we propose CSTAR, an efficient solution that simultaneously impose Compactness, high STructuredness and high Adversarial Robustness on the target DNN models. By formulating the structuredness and robustness requirement within the same framework, the compressed DNNs can simultaneously achieve high compression performance and strong adversarial robustness. Evaluations for various DNN models on different datasets demonstrate the effectiveness of CSTAR. Compared with the state-of-the-art robust structured pruning, CSTAR shows consistently better performance. For instance, when compressing ResNet-18 on CIFAR-10, CSTAR achieves up to 20.07% and 11.91% improvement for benign accuracy and robust accuracy, respectively. For compressing ResNet-18 with 16x compression ratio on Imagenet, CSTAR obtains 8.58% benign accuracy gain and 4.27% robust accuracy gain compared to the existing robust structured pruning.
Huy Phan, Miao Yin, Yang Sui 0001, Bo Yuan 0001, Saman A. Zonouz
AAAI3
2023 HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
abstract
Low-rank compression is an important model compression strategy for obtaining compact neural network models. In general, because the rank values directly determine the model complexity and model accuracy, proper selection of layer-wise rank is very critical and desired. To date, though many low-rank compression approaches, either selecting the ranks in a manual or automatic way, have been proposed, they suffer from costly manual trials or unsatisfied compression performance. In addition, all of the existing works are not designed in a hardware-aware way, limiting the practical performance of the compressed models on real-world hardware platforms. To address these challenges, in this paper we propose HALOC, a hardware-aware automatic low-rank compression framework. By interpreting automatic rank selection from an architecture search perspective, we develop an end-to-end solution to determine the suitable layer-wise ranks in a differentiable and hardware-aware way. We further propose design principles and mitigation strategy to efficiently explore the rank space and reduce the potential interference problem. Experimental results on different datasets and hardware platforms demonstrate the effectiveness of our proposed approach. On CIFAR-10 dataset, HALOC enables 0.07% and 0.38% accuracy increase over the uncompressed ResNet-20 and VGG-16 models with 72.20% and 86.44% fewer FLOPs, respectively. On ImageNet dataset, HALOC achieves 0.9% higher top-1 accuracy than the original ResNet-18 model with 66.16% fewer FLOPs. HALOC also shows 0.66% higher top-1 accuracy increase than the state-of-the-art automatic low-rank compression solution with fewer computational and memory costs. In addition, HALOC demonstrates the practical speedups on different hardware platforms, verified by the measurement results on desktop GPU, embedded GPU and ASIC accelerator.
Jinqi Xiao, Chengming Zhang 0006, Yu Gong 0003, Miao Yin, Yang Sui 0001, Lizhi Xiang, Dingwen Tao, Bo Yuan 0001
AAAI5
2023 DSPIMM: A Fully Digital SParse In-Memory Matrix Vector Multiplier for Communication Applications
abstract
Channel decoders are key computing modules in wired/wireless communication systems. Recently neural network (NN)-based decoders have shown their promising error-correcting performance because of their end-to-end learning capability. However, compared with the traditional approaches, the emerging neural belief propagation (NBP) solution suffers higher storage and computational complexity, limiting its hardware performance. To address this challenge and develop a channel decoder that can achieve high decoding performance and hardware performance simultaneously, in this paper we take a first step towards exploring SRAM-based in-memory computing for efficient NBP channel decoding. We first analyze the unique sparsity pattern in the NBP processing, and then propose an efficient and fully Digital Sparse In-Memory Matrix vector Multiplier (DSPIMM) computing platform. Extensive experiments demonstrate that our proposed DSPIMM achieves significantly higher energy efficiency and throughput than the state-of-the-art counterparts.
Amitesh Sridharan, Fan Zhang 0069, Yang Sui 0001, Bo Yuan 0001, Deliang Fan
DAC3
2023 Invited Paper: In-Sensor Radio Frequency Computing for Energy-Efficient Intelligent Radar
abstract
Radio Frequency Neural Networks (RFNNs) have demonstrated advantages in realizing intelligent applications across various domains. However, as the model size of deep neural networks rapidly increases, implementing large-scale RFNN in practice requires an extensive number of RF interferometers and consumes a substantial amount of energy. To address this challenge, we propose to utilize low-rank decomposition to transform a large-scale RFNN into a compact RFNN while almost preserving its accuracy. Specifically, we develop a Tensor-Train RFNN (TT-RFNN) where each layer comprises a sequence of low-rank third-order tensors, leading to a notable reduction in parameter count, thereby optimizing RF interferometer utilization in comparison to the original large-scale RFNN. Additionally, considering the inherent physical errors when mapping TT-RFNN to RF device parameters in real-world deployment, from a general perspective, we construct the Robust TT-RFNN (RTT-RFNN) by incorporating a robustness solver on TT-RFNN to enhance its robustness. To adapt the RTT-RFNN to varying requirements of reshaping operations, we further provide a reconfigurable reshaping solution employing RF switch matrices. Empirical evaluations conducted on MNIST and CIFAR-10 datasets show the effectiveness of our proposed method.
Yang Sui 0001, Minning Zhu, Lingyi Huang, Chung-Tse Michael Wu, Bo Yuan 0001
ICCAD1
2023 DynGMP: Graph Neural Network-Based Motion Planning in Unpredictable Dynamic Environments
abstract
Neural networks have already demonstrated attractive performance for solving motion planning problems, especially in static and predictable environments. However, efficient neural planners that can adapt to unpredictable dynamic environments, a highly demanded scenario in many practical applications, are still under-explored. To fill this research gap and enrich the existing motion planning approaches, in this pa-per, we propose DynGMP, a graph neural network (GNN)-based planner that provides high-performance planning solutions in unpredictable dynamic environments. By fully leveraging the prior exploration experience and minimizing the replanning cost incurred by environmental change, DynGMP achieves high planning performance and efficiency simultaneously. Empirical evaluations across different environments show that DynGMP can achieve close to 100% success rate with fast planning speed and short path cost. Compared with existing non-learning and learning-based counterparts, DynGMP shows very significant planning performance improvement, e.g., at least 2.7×, 2.2×,$2.4\times$and$2\times$faster planning speed with low path distance in four environments, respectively.
Xiao Zang, Lingyi Huang, Yang Sui 0001, Jingjin Yu, Yingying Chen 0001, Bo Yuan 0001
IROS4
2023 ETTE: Efficient Tensor-Train-based Computing Engine for Deep Neural Networks
abstract
Tensor-train (TT) decomposition enables ultra-high compression ratio, making the deep neural network (DNN) accelerators based on this method very attractive. TIE, the state-of-the-art TT based DNN accelerator, achieved high performance by leveraging a compact inference scheme to remove unnecessary computations and memory access. However, TIE increases memory costs for stage-wise intermediate results and additional intra-layer data transfer, leading to limited speedups even the models are highly compressed.
Yu Gong 0003, Miao Yin, Lingyi Huang, Jinqi Xiao, Yang Sui 0001, Chunhua Deng, Bo Yuan 0001
ISCA5
2022 HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural Networks
abstract
High-order decomposition is a widely used model compression approach towards compact convolutional neural networks (CNNs). However, many of the existing solutions, though can efficiently reduce CNN model sizes, are very difficult to bring considerable saving for computational costs, especially when the compression ratio is not huge, thereby causing the severe computation inefficiency problem. To overcome this challenge, in this paper we propose efficient High-Order DEcomposed Convolution (HODEC). By performing systematic explorations on the underlying reason and mitigation strategy for the computation inefficiency, we develop a new decomposition and computation-efficient execution scheme, enabling simultaneous reductions in computational and storage costs. To demonstrate the effectiveness of HODEC, we perform empirical evaluations for various CNN models on different datasets. HODEC shows consistently outstanding compression and acceleration performance. For compressing ResNet-56 on CIFAR-10 dataset, HODEC brings 67% fewer parameters and 62% fewer FLOPs with 1.17% accuracy increase than the baseline model. For compressing ResNet-50 on ImageNet dataset, HODEC achieves 63% FLOPs reduction with 0.31% accuracy increase than the uncompressed model.
Miao Yin, Yang Sui 0001, Wanzhao Yang, Xiao Zang, Yu Gong 0003, Bo Yuan 0001
CVPR2
2021 Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization Framework
abstract
Advanced tensor decomposition, such as tensor train (TT) and tensor ring (TR), has been widely studied for deep neural network (DNN) model compression, especially for recurrent neural networks (RNNs). However, compressing convolutional neural networks (CNNs) using TT/TR always suffers significant accuracy loss. In this paper, we propose a systematic framework for tensor decomposition-based model compression using Alternating Direction Method of Multipliers (ADMM). By formulating TT decomposition-based model compression to an optimization problem with constraints on tensor ranks, we leverage ADMM technique to systemically solve this optimization problem in an iterative way. During this procedure, the entire DNN model is trained in the original structure instead of TT format, but gradually enjoys the desired low tensor rank characteristics. We then decompose this uncompressed model to TT format, and fine-tune it to finally obtain a high-accuracy TT-format DNN model. Our framework is very general, and it works for both CNNs and RNNs, and can be easily modified to fit other tensor decomposition approaches. We evaluate our proposed framework on different DNN models for image classification and video recognition tasks. Experimental results show that our ADMM-based TT-format models demonstrate very high compression performance with high accuracy. Notably, on CIFAR-100, with 2.3× and 2.4× compression ratios, our models have 1.96% and 2.21% higher top-1 accuracy than the original ResNet-20 and ResNet-32, respectively. For compressing ResNet-18 on ImageNet, our model achieves 2.47× FLOPs reduction without accuracy loss.
Miao Yin, Yang Sui 0001, Siyu Liao, Bo Yuan 0001
CVPR2
2021 Algorithm and Hardware Co-design for Deep Learning-powered Channel Decoder: A Case Study
abstract
Channel decoder is a key component module in many communication systems. Recently, neural networks-based channel decoders have been actively investigated because of the great potential of their data-driven decoding procedure. However, as the intersection among machine learning, information theory and hardware design, the efficient algorithm and hardware codesign of deep learning-powered channel decoder has not been well studied. This paper is a first step towards exploring the efficient DNN-enabled channel decoders, from a joint perspective of algorithm and hardware. We first revisit our recently proposed doubly residual neural decoder. By introducing the advanced architectural topology on the decoder design, the overall error-correcting performance can be significantly improved. Based on this algorithm, we further develop the corresponding systolic array-based hardware architecture for the DRN decoder. The corresponding FPGA implementation for our DRN decoder on short LDPC code is also developed.
Boyang Zhang 0007, Yang Sui 0001, Lingyi Huang, Siyu Liao, Chunhua Deng, Bo Yuan 0001
ICCAD2
2021 GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network Accelerator
abstract
The co-existence of activation sparsity and model sparsity in convolutional neural network (CNN) models makes sparsity-aware CNN hardware designs very attractive. The existing sparse CNN accelerators utilize intersection operation to search and identify the key positions of the matched entries between two sparse vectors, and hence avoid unnecessary computations. However, these state-of-the-art designs still suffer from three major architecture-level drawbacks, including 1) hardware cost for the intersection operation is high; 2) frequent stalls of computation phase due to strong data dependency between intersection and computation phases; and 3) unnecessary data transfer incurred by the explicit intersection operation.By leveraging the knowledge of the complete sparse 2-D convolution, this paper proposes two key ideas that overcome all of the three drawbacks. First, an implicit on-the-fly intersection is proposed to realize the optimal solution for intersection between one static stream and one dynamic stream, which is the case for sparse neural network inference. Second, by leveraging the global computation structure of 2-D convolution, we propose a specialized computation reordering to ensure that the activation is only transferred if necessary and only once.Based on these two key ideas, we develop GoSPA, an energy-efficient high-performance Globally Optimized SParse CNN Accelerator. GoSPA is implemented with CMOS 28nm technology. Compared with the state-of-the-art sparse CNN architecture, GoSPA achieves average 1.38×, 1.28×, 1.23×, 1.17×, 1.21× and 1.28× speedup on AlexNet, VGG, GoogLeNet, MobileNet, ResNet and ResNeXt workloads, respectively. Also, GoSPA achieves 5.38×, 4.96×, 4.79×, 5.02×, 4.86× and 2.06× energy efficiency improvement on AlexNet, VGG, GoogLeNet, MobileNet, ResNet and ResNeXt, respectively. In more comprehensive comparison including DRAM access, GoSPA also shows significant performance improvement over the existing designs.
Chunhua Deng, Yang Sui 0001, Siyu Liao, Xuehai Qian, Bo Yuan 0001
ISCA2
2021 CHIP: CHannel Independence-based Pruning for Compact Neural Networks
abstract
Filter pruning has been widely used for neural network compression because of its enabled practical acceleration. To date, most of the existing filter pruning works explore the importance of filters via using intra-channel information. In this paper, starting from an inter-channel perspective, we propose to perform efficient filter pruning using Channel Independence, a metric that measures the correlations among different feature maps. The less independent feature map is interpreted as containing less useful information$/$knowledge, and hence its corresponding filter can be pruned without affecting model capacity. We systematically investigate the quantification metric, measuring scheme and sensitiveness$/$reliability of channel independence in the context of filter pruning. Our evaluation results for different models on various datasets show the superior performance of our approach. Notably, on CIFAR-10 dataset our solution can bring $0.75\%$ and $0.94\%$ accuracy increase over baseline ResNet-56 and ResNet-110 models, respectively, and meanwhile the model size and FLOPs are reduced by $42.8\%$ and $47.4\%$ (for ResNet-56) and $48.3\%$ and $52.1\%$ (for ResNet-110), respectively. On ImageNet dataset, our approach can achieve $40.8\%$ and $44.8\%$ storage and computation reductions, respectively, with $0.15\%$ accuracy increase over the baseline ResNet-50 model. The code is available at https://github.com/Eclipsess/CHIP_NeurIPS2021.
Yang Sui 0001, Miao Yin, Yi Xie 0001, Huy Phan, Saman A. Zonouz, Bo Yuan 0001
NeurIPS1