VLDB 2026 Research / reviewers in the wild / expert
Miao Yin
dblp:199/1982
· DBLP profile ↗
28ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Systems, architecture and hardware · 9 · 8 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsabstractThe increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0× to 8.4× memory reduction and 1.7× to 75.0× speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs. Zhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu, Miao Yin, Gagan Agrawal, Wei Niu 0002 |
ASPLOS (2) | 5 |
| 2025 | AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory ReductionabstractThe advancements in large language models (LLMs) have propelled the improvement of video understanding tasks by incorporating LLMs with visual models. However, most existing LLM-based models (e.g., VideoLLaMA, VideoChat) are constrained to processing short-duration videos. Recent attempts to understand long-term videos by extracting and compressing visual features into a fixed memory size. Nevertheless, those methods leverage only visual modality to merge video tokens and overlook the correlation between visual and textual queries, leading to difficulties in effectively handling complex question-answering tasks. To address the challenges of long videos and complex prompts, we propose AdaCM2, which, for the first time, introduces an adaptive cross-modality memory reduction approach to video-text alignment in an auto-regressive manner on video streams. Our extensive experiments on various video understanding tasks, such as video captioning, video question answering, and video classification, demonstrate that AdaCM2achieves state-of-the-art performance across multiple datasets while significantly reducing memory usage. Notably, it achieves a 4.5% improvement across multiple tasks in the LVU dataset with a GPU memory consumption reduction of up to 65%. Yuanbin Man, Ying Huang 0008, Chengming Zhang 0006, Bingzhe Li, Wei Niu 0002, Miao Yin |
CVPR | 6 |
| 2025 | GaussianSpa: An "Optimizing-Sparsifying" Simplification Framework for Compact and High-Quality 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a mainstream for novel view synthesis, leveraging continuous aggregations of Gaussian functions to model scene geometry. However, 3DGS suffers from substantial memory requirements to store the large amount of Gaussians, hindering its efficiency and practicality. To address this challenge, we introduce GaussianSpa, an optimization-based simplification framework for compact and high-quality 3DGS. Specifically, we formulate the simplification objective as a constrained optimization problem associated with the 3DGS training. Correspondingly, we propose an efficient "optimizing-sparsifying" solution for the formulated problem, alternately solving two independent sub-problems and gradually imposing substantial sparsity onto the Gaussians in the 3DGS training process. We conduct quantitative and qualitative evaluations on various datasets, demonstrating the superiority of GaussianSpa over existing state-of-the-art approaches. Notably, GaussianSpa achieves an average PSNR improvement of 0.9 dB on the real-world Deep Blending dataset with 10× fewer Gaussians compared to the vanilla 3DGS. Our project page is available at https://noodle-lab.github.io/gaussianspa. Yangming Zhang 0001, Wenqi Jia 0003, Wei Niu 0002, Miao Yin |
CVPR | 4 |
| 2025 | Advancing Scientific Data Compression via Cross-Field PredictionabstractScientific applications generate massive amounts of data, leading to significant storage and I/O bottlenecks. Lossy compression is a crucial technique for reducing data volume in scientific computing, balancing storage efficiency and data fidelity. Such scientific data often consists of multiple fields representing various physical metrics, which exhibit inherent correlations. However, existing compression solutions predominantly rely on local information within a single field, overlooking cross-field correlations that could enhance compression efficiency. In this paper, we introduce a novel compression framework that integrates both local and cross-field information to enhance predictive accuracy. With the help of a carefully designed neural network, our approach effectively captures complex correlations across fields, significantly improving compression ratios. Additionally, our method can rely solely on cross-field prediction to reconstruct fields, achieving extremely high compression ratios under relaxed accuracy constraints. We evaluate our approach on two real-world scientific applications across 85 data fields, demonstrating compression ratio improvements of up to 103.4% for a single field and up to 19.3% overall. Youyuan Liu, Wenqi Jia 0003, Taolue Yang, Miao Yin, Sian Jin |
HPDC | 5 |
| 2025 | NeurLZ: An Online Neural Learning-based Method to Enhance Scientific Lossy CompressionabstractSZ3 (0.07MB, PSNR = 29.2) (c)NeurLZ (0.07MB, PSNR = 39.1) Wenqi Jia 0003, Zhewen Hu, Youyuan Liu, Boyuan Zhang 0002, Jinzhen Wang, Jinyang Liu 0003, Wei Niu 0002, Stavros Kalafatis, Junzhou Huang, Sian Jin, Daoce Wang, Jiannan Tian, Miao Yin |
ICS | 13 |
| 2025 | Co-Exploring Structured Sparsification and Low-Rank Tensor Decomposition for Compact DNNsabstractSparsification and low-rank decomposition are two important techniques to compress deep neural network (DNN) models. To date, these two popular yet distinct approaches are typically used in separate ways; while their efficient integration for better compression performance is little explored, especially for structured sparsification and decomposition. In this article, we perform systematic co-exploration on structured sparsification and decomposition toward compact DNN models. We first investigate and analyze several important design factors for joint structured sparsification and decomposition, including operational sequence, decomposition format, and optimization procedure. Based on the observations from our analysis, we then propose CEPD, a unified DNN compression framework that can co-explore the benefits of structured sparsification and tensor decomposition in an efficient way. Empirical experiments demonstrate the promising performance of our proposed solution. Notably, on the CIFAR-10 dataset, CEPD brings 0.72%-0.45% accuracy increase over the baseline ResNet-56 and MobileNetV2 models, respectively, and meanwhile, the computational costs are reduced by 43.0%-44.2%, respectively. On the ImageNet dataset, our approach can enable 0.10%-1.39% accuracy increase over the baseline ResNet-18 and ResNet-50 models with 59.4%-54.6% fewer parameters, respectively. Yang Sui 0001, Miao Yin, Yu Gong 0003, Bo Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on MobileabstractThis work is motivated by recent developments in Deep Neural Networks, particularly the Transformer architectures underlying applications such as ChatGPT, and the need for performing inference on mobile devices. Focusing on emerging transformers (specifically the ones with computationally efficient Swin-like architectures) and large models (e.g., Stable Diffusion and LLMs) based on transformers, we observe that layout transformations between the computational operators cause a significant slowdown in these applications. This paper presents SmartMem, a comprehensive framework for eliminating most layout transformations, with the idea that multiple operators can use the same tensor layout through careful choice of layout and implementation of operations. Our approach is based on classifying the operators into four groups, and considering combinations of producer-consumer edges between the operators. We develop a set of methods for searching such layouts. Another component of our work is developing efficient memory layouts for 2.5 dimensional memory commonly seen in mobile devices. Our experimental results show that SmartMem outperforms 5 state-of-the-art DNN execution frameworks on mobile devices across 18 varied neural networks, including CNNs, Transformers with both local and global attention, as well as LLMs. In particular, compared to DNNFusion, SmartMem achieves an average speedup of 2.8×, and outperforms TVM and MNN with speedups of 6.9× and 7.9×, respectively, on average. Wei Niu 0002, Md. Musfiqur Rahman Sanim, Zhihao Shu, Jiexiong Guan, Xipeng Shen, Miao Yin, Gagan Agrawal, Bin Ren 0002 |
ASPLOS (3) | 6 |
| 2024 | PointMLFF: Robust Point Cloud Analysis Based on Multi-Level Feature Fusionabstract3D point cloud is often affected by sensor noise, environmental interference, and incomplete collection, resulting in noise, missing, and anomalies in the point cloud. Existing work in handling such data primarily focuses on the coordinate information of the point cloud, overlooking its local structure and the interrelation between points, thereby reducing the accuracy of model predictions. To enhance the ability to capture point cloud features, we propose a high-precision and robust 3D point cloud analysis model called PointMLFF. Specifically, we introduce a 3D surface-based point local feature extractor, capturing the local geometric features of the point cloud by approximating the Taylor series to construct local spatial geometries. This method preserves both the absolute position information and the local shape information of the point cloud. Additionally, we propose a multi-level key-point attention module based on deep features, which constructs a point embedding space capable of perceiving abnormal changes in the point cloud by calculating the key-neighbor point attention and inter-key point attention in the feature space, significantly improving the model’s robustness. Extensive experiments show that PointMLFF outperforms most advanced methods in various downstream tasks. Notably, our method achieves a high classification accuracy of 88.9% on the challenging ScanObjectNN and surpasses others on abnormal point clouds. The visualization of partial segmentation results closely resembles the actual scenarios. Miao Yin, Jinshuo Zhang, Xiuyang Zhao |
IJCNN | 1 |
| 2024 | Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile DevicesabstractMobile devices have become essential enablers for AI applications, particularly in scenarios that require real-time performance. Vision Transformer (ViT) has become a fundamental cornerstone in this regard due to its high accuracy. Recent efforts have been dedicated to developing various transformer architectures that offer im- proved accuracy while reducing the computational requirements. However, existing research primarily focuses on reducing the theoretical computational complexity through methods such as local attention and model pruning, rather than considering realistic performance on mobile hardware. Although these optimizations reduce computational demands, they either introduce additional overheads related to data transformation (e.g., Reshape and Transpose) or irregular computation/data-access patterns. These result in significant overhead on mobile devices due to their limited bandwidth, which even makes the latency worse than vanilla ViT on mobile. In this paper, we present ECP-ViT, a real-time framework that employs the core-periphery principle inspired by the brain functional networks to guide self-attention in ViTs and enable the deployment of ViT models on smartphones. We identify the main bottleneck in transformer structures caused by data transformation and propose a hardware-friendly core-periphery guided self-attention to decrease computation demands. Additionally, we design the system optimizations for intensive data transformation in pruned models. ECP-ViT, with the proposed algorithm-system co-optimizations, achieves a speedup of 4.6× to 26.9× on mobile GPUs across four datasets: STL-10, CIFAR100, TinyImageNet, and ImageNet. Zhihao Shu, Xiaowei Yu 0001, Zihao Wu 0001, Wenqi Jia 0003, Yinchen Shi, Miao Yin, Tianming Liu 0001, Dajiang Zhu, Wei Niu 0002 |
NeurIPS | 6 |
| 2023 | CSTAR: Towards Compact and Structured Deep Neural Networks with Adversarial RobustnessabstractModel compression and model defense for deep neural networks (DNNs) have been extensively and individually studied. Considering the co-importance of model compactness and robustness in practical applications, several prior works have explored to improve the adversarial robustness of the sparse neural networks. However, the structured sparse models obtained by the existing works suffer severe performance degradation for both benign and robust accuracy, thereby causing a challenging dilemma between robustness and structuredness of compact DNNs. To address this problem, in this paper, we propose CSTAR, an efficient solution that simultaneously impose Compactness, high STructuredness and high Adversarial Robustness on the target DNN models. By formulating the structuredness and robustness requirement within the same framework, the compressed DNNs can simultaneously achieve high compression performance and strong adversarial robustness. Evaluations for various DNN models on different datasets demonstrate the effectiveness of CSTAR. Compared with the state-of-the-art robust structured pruning, CSTAR shows consistently better performance. For instance, when compressing ResNet-18 on CIFAR-10, CSTAR achieves up to 20.07% and 11.91% improvement for benign accuracy and robust accuracy, respectively. For compressing ResNet-18 with 16x compression ratio on Imagenet, CSTAR obtains 8.58% benign accuracy gain and 4.27% robust accuracy gain compared to the existing robust structured pruning. Huy Phan, Miao Yin, Yang Sui 0001, Bo Yuan 0001, Saman A. Zonouz |
AAAI | 2 |
| 2023 | HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural NetworksabstractLow-rank compression is an important model compression strategy for obtaining compact neural network models. In general, because the rank values directly determine the model complexity and model accuracy, proper selection of layer-wise rank is very critical and desired. To date, though many low-rank compression approaches, either selecting the ranks in a manual or automatic way, have been proposed, they suffer from costly manual trials or unsatisfied compression performance. In addition, all of the existing works are not designed in a hardware-aware way, limiting the practical performance of the compressed models on real-world hardware platforms. To address these challenges, in this paper we propose HALOC, a hardware-aware automatic low-rank compression framework. By interpreting automatic rank selection from an architecture search perspective, we develop an end-to-end solution to determine the suitable layer-wise ranks in a differentiable and hardware-aware way. We further propose design principles and mitigation strategy to efficiently explore the rank space and reduce the potential interference problem. Experimental results on different datasets and hardware platforms demonstrate the effectiveness of our proposed approach. On CIFAR-10 dataset, HALOC enables 0.07% and 0.38% accuracy increase over the uncompressed ResNet-20 and VGG-16 models with 72.20% and 86.44% fewer FLOPs, respectively. On ImageNet dataset, HALOC achieves 0.9% higher top-1 accuracy than the original ResNet-18 model with 66.16% fewer FLOPs. HALOC also shows 0.66% higher top-1 accuracy increase than the state-of-the-art automatic low-rank compression solution with fewer computational and memory costs. In addition, HALOC demonstrates the practical speedups on different hardware platforms, verified by the measurement results on desktop GPU, embedded GPU and ASIC accelerator. Jinqi Xiao, Chengming Zhang 0006, Yu Gong 0003, Miao Yin, Yang Sui 0001, Lizhi Xiang, Dingwen Tao, Bo Yuan 0001 |
AAAI | 4 |
| 2023 | GOHSP: A Unified Framework of Graph and Optimization-Based Heterogeneous Structured Pruning for Vision TransformerabstractThe recently proposed Vision transformers (ViTs) have shown very impressive empirical performance in various computer vision tasks, and they are viewed as an important type of foundation model. However, ViTs are typically constructed with large-scale sizes, which then severely hinder their potential deployment in many practical resources constrained applications. To mitigate this challenging problem, structured pruning is a promising solution to compress model size and enable practical efficiency. However, unlike its current popularity for CNNs and RNNs, structured pruning for ViT models is little explored. In this paper, we propose GOHSP, a unified framework of Graph and Optimization-based Structured Pruning for ViT models. We first develop a graph-based ranking for measuring the importance of attention heads, and the extracted importance information is further integrated to an optimization-based procedure to impose the heterogeneous structured sparsity patterns on the ViT models. Experimental results show that our proposed GOHSP demonstrates excellent compression performance. On CIFAR-10 dataset, our approach can bring 40% parameters reduction with no accuracy loss for ViT-Small model. On ImageNet dataset, with 30% and 35% sparsity ratio for DeiT-Tiny and DeiT-Small models, our approach achieves 1.65% and 0.76% accuracy increase over the existing structured pruning methods, respectively. Miao Yin, Burak Uzkent, Yilin Shen, Hongxia Jin, Bo Yuan 0001 |
AAAI | 1 |
| 2023 | COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsabstractAttention-based vision models, such as Vision Transformer (ViT) and its variants, have shown promising performance in various computer vision tasks. However, these emerging architectures suffer from large model sizes and high computational costs, calling for efficient model compression solutions. To date, pruning ViTs has been well studied, while other compression strategies that have been widely applied in CNN compression, e.g., model factorization, is little explored in the context of ViT compression. This paper explores an efficient method for compressing vision transformers to enrich the toolset for obtaining compact attention-based vision models. Based on the new insight on the multi-head attention layer, we develop a highly efficient ViT compression solution, which outperforms the state-of-the-art pruning methods. For compressing DeiT-small and DeiT-base models on ImageNet, our proposed approach can achieve $0.45%$ and $0.76%$ higher top-1 accuracy even with fewer parameters. Our finding can also be applied to improve the customization efficiency of text-to-image diffusion models, with much faster training (up to $2.6\times$ speedup) and lower extra storage cost (up to $1927.5\times$ reduction) than the existing works. Jinqi Xiao, Miao Yin, Yu Gong 0003, Xiao Zang, Jian Ren 0005, Bo Yuan 0001 |
ICML | 2 |
| 2023 | ETTE: Efficient Tensor-Train-based Computing Engine for Deep Neural NetworksabstractTensor-train (TT) decomposition enables ultra-high compression ratio, making the deep neural network (DNN) accelerators based on this method very attractive. TIE, the state-of-the-art TT based DNN accelerator, achieved high performance by leveraging a compact inference scheme to remove unnecessary computations and memory access. However, TIE increases memory costs for stage-wise intermediate results and additional intra-layer data transfer, leading to limited speedups even the models are highly compressed. Yu Gong 0003, Miao Yin, Lingyi Huang, Jinqi Xiao, Yang Sui 0001, Chunhua Deng, Bo Yuan 0001 |
ISCA | 2 |
| 2023 | GraphMP: Graph Neural Network-based Motion Planning with Efficient Graph SearchabstractMotion planning, which aims to find a high-quality collision-free path in the configuration space, is a fundamental task in robotic systems. Recently, learning-based motion planners, especially the graph neural network-powered, have shown promising planning performance. However, though the state-of-the-art GNN planner can efficiently extract and learn graph information, its inherent mechanism is not well suited for graph search process, hindering its further performance improvement. To address this challenge and fully unleash the potential of GNN in motion planning, this paper proposes GraphMP, a neural motion planner for both low and high-dimensional planning tasks. With the customized model architecture and training mechanism design, GraphMP can simultaneously perform efficient graph pattern extraction and graph search processing, leading to strong planning performance. Experiments on a variety of environments, ranging from 2D Maze to 14D dual KUKA robotic arm, show that our proposed GraphMP achieves significant improvement on path quality and planning speed over the state-of-the-art learning-based and classical planners; while preserving the competitive success rate. Xiao Zang, Miao Yin, Jinqi Xiao, Saman A. Zonouz, Bo Yuan 0001 |
NeurIPS | 2 |
| 2023 | TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker DecompositionabstractTucker decomposition is one of the SOTA CNN model compression techniques. However, unlike the FLOPs reduction, we observe very limited inference time reduction with Tucker-compressed models using existing GPU software such as cuDNN. To this end, we propose an efficient end-to-end framework that can generate highly accurate and compact CNN models via Tucker decomposition and optimized inference code on GPUs. Specifically, we propose an ADMM-based training algorithm that can achieve highly accurate Tucker-format models. We also develop a high-performance kernel for Tucker-format convolutions and analytical performance models to guide the selection of execution parameters. We further propose a co-design framework to determine the proper Tucker ranks driven by practical inference time (rather than FLOPs). Our evaluation on five modern CNNs with A100 demonstrates that our compressed models with our optimized code achieve up to 2.21× speedup over cuDNN, 1.12× speedup over TVM, and 3.27× over the original models using cuDNN with at most 0.05% accuracy loss. Lizhi Xiang, Miao Yin, Chengming Zhang 0006, Aravind Sukumaran-Rajam, P. Sadayappan, Bo Yuan 0001, Dingwen Tao |
PPoPP | 2 |
| 2023 | Online Reviews Sentiment Analysis and Product Feature Improvement with Deep LearningabstractThe text mining of online reviews is currently a popular research direction of e-commerce and is considered the next blue ocean. Online reviews can dig out consumer preferences and provide theoretical guidance for the improvement of product features. However, current research mostly focuses on sentiment analysis methods and rarely involves feature extraction and large-scale data recognition. This article uses word segmentation technology to create a new feature extraction method. With the long short-term memory neural network and latent Dirichlet allocation topic model, we propose a product feature improvement model—CESC (Consumer online reviews–Extract short text–Sentiment analysis–Cluster feature). The model can derive the product features and attitudes that consumers prefer based on consumer online reviews and use it to improve product features. According to the experimental results of three electronic products sold on the e-commerce platform, the model can effectively dig out consumer preferences for online reviews. Enterprises can improve the quality of products and services, better meet the needs of consumers, promote consumers’ consumption, and achieve the enterprises’ goals and values. Jihua Cao, Jie Li 0124, Miao Yin |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Projected Generative Adversarial Network for Point Cloud CompletionabstractAcquiring semantics directly from a point cloud is an important requirement for handling point cloud tasks. However, point clouds captured with laser scanner equipment are often incomplete due to the limitations posed by target occlusion and light reflection. Consequently, recovering the complete point clouds from partial and sparse ones is essential for further studies. In this paper, we model a novel projected generative adversarial network (PGAN) for point cloud completion. First, we present a multi-scale generator module (MSGM) to fully capture the local structures and global shape in the raw incompletion point cloud and generate the multi-scale complete point cloud. In contrast to existing point cloud feature extractors, our MSGM promotes a correlation between different regions of an incomplete point cloud and integrates the contextual information of the point cloud. Second, we observe that the existing point discriminator is inadequate to enhance the discrimination of the prediction point cloud. To address this problem, we project the completed point cloud to 2D maps and apply adversarial training to discriminate the geometrical shape from a specific viewpoint. Comprehensive experiments on the ShapeNet and ModelNet40 datasets show that the proposed method performs well against existing point cloud completion tasks. We also present an ablation study to demonstrate the advantages of the projected generative adversarial network. Xue Lin 0008, Dongmei Niu, Daole Wang, Miao Yin, Xiuyang Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | BATUDE: Budget-Aware Neural Network Compression Based on Tucker DecompositionabstractModel compression is very important for the efficient deployment of deep neural network (DNN) models on resource-constrained devices. Among various model compression approaches, high-order tensor decomposition is particularly attractive and useful because the decomposed model is very small and fully structured. For this category of approaches, tensor ranks are the most important hyper-parameters that directly determine the architecture and task performance of the compressed DNN models. However, as an NP-hard problem, selecting optimal tensor ranks under the desired budget is very challenging and the state-of-the-art studies suffer from unsatisfied compression performance and timing-consuming search procedures. To systematically address this fundamental problem, in this paper we propose BATUDE, a Budget-Aware TUcker DEcomposition-based compression approach that can efficiently calculate optimal tensor ranks via one-shot training. By integrating the rank selecting procedure to the DNN training process with a specified compression budget, the tensor ranks of the DNN models are learned from the data and thereby bringing very significant improvement on both compression ratio and classification accuracy for the compressed models. The experimental results on ImageNet dataset show that our method enjoys 0.33% top-5 higher accuracy with 2.52X less computational cost as compared to the uncompressed ResNet-18 model. For ResNet-50, the proposed approach enables 0.37% and 0.55% top-5 accuracy increase with 2.97X and 2.04X computational cost reduction, respectively, over the uncompressed model. Miao Yin, Huy Phan, Xiao Zang, Siyu Liao, Bo Yuan 0001 |
AAAI | 1 |
| 2022 | HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural NetworksabstractHigh-order decomposition is a widely used model compression approach towards compact convolutional neural networks (CNNs). However, many of the existing solutions, though can efficiently reduce CNN model sizes, are very difficult to bring considerable saving for computational costs, especially when the compression ratio is not huge, thereby causing the severe computation inefficiency problem. To overcome this challenge, in this paper we propose efficient High-Order DEcomposed Convolution (HODEC). By performing systematic explorations on the underlying reason and mitigation strategy for the computation inefficiency, we develop a new decomposition and computation-efficient execution scheme, enabling simultaneous reductions in computational and storage costs. To demonstrate the effectiveness of HODEC, we perform empirical evaluations for various CNN models on different datasets. HODEC shows consistently outstanding compression and acceleration performance. For compressing ResNet-56 on CIFAR-10 dataset, HODEC brings 67% fewer parameters and 62% fewer FLOPs with 1.17% accuracy increase than the baseline model. For compressing ResNet-50 on ImageNet dataset, HODEC achieves 63% FLOPs reduction with 0.31% accuracy increase than the uncompressed model. Miao Yin, Yang Sui 0001, Wanzhao Yang, Xiao Zang, Yu Gong 0003, Bo Yuan 0001 |
CVPR | 1 |
| 2022 | Robot Motion Planning as Video Prediction: A Spatio-Temporal Neural Network-based Motion PlannerabstractNeural network (NN)-based methods have emerged as an attractive approach for robot motion planning due to strong learning capabilities of NN models and their inherently high parallelism. Despite the current development in this direction, the efficient capture and processing of important sequential and spatial information, in a direct and simultaneous way, is still relatively under-explored. To overcome the challenge and unlock the potentials of neural networks for motion planning tasks, in this paper, we propose STP-Net, an end-to-end learning framework that can fully extract and leverage important spatio-temporal information to form an efficient neural motion planner. By interpreting the movement of the robot as a video clip, robot motion planning is transformed to a video prediction task that can be performed by STP-Net in both spatially and temporally efficient ways. Empirical evaluations across different seen and unseen environments show that, with nearly 100% accuracy (aka, success rate), STP-Net demonstrates very promising performance with respect to both planning speed and path cost. Compared with existing NN-based motion planners, STP-Net achieves at least 5×, 2.6× and 1.8× faster speed with lower path cost on 2D Random Forest, 2D Maze and 3D Random Forest environments, respectively. Furthermore, STP-Net can quickly and simultaneously compute multiple near-optimal paths in multi-robot motion planning tasks. Xiao Zang, Miao Yin, Lingyi Huang, Jingjin Yu, Saman A. Zonouz, Bo Yuan 0001 |
IROS | 2 |
| 2022 | Algorithm and Hardware Co-Design of Energy-Efficient LSTM Networks for Video Recognition With Hierarchical Tucker Tensor DecompositionabstractLong short-term memory (LSTM) is a type of powerful deep neural network that has been widely used in many sequence analysis and modeling applications. However, the large model size problem of LSTM networks make their practical deployment still very challenging, especially for the video recognition tasks that require high-dimensional input data. Aiming to overcome this limitation and fully unlock the potentials of LSTM models, in this paper we propose to perform algorithm and hardware co-design towards high-performance energy-efficient LSTM networks. At algorithm level, we propose to developfully decomposed hierarchical Tucker (FDHT)structure-based LSTM, namely FDHT-LSTM, which enjoys ultra-low model complexity while still achieving high accuracy. In order to fully reap such attractive algorithmic benefit, we further develop the corresponding customized hardware architecture to support the efficient execution of the proposed FDHT-LSTM model. With the delicate design of memory access scheme, the complicated matrix transformation can be efficiently supported by the underlying hardware without any access conflict in an on-the-fly way. Our evaluation results show that both the proposed ultra-compact FDHT-LSTM models and the corresponding hardware accelerator achieve very high performance. Compared with the state-of-the-art compressed LSTM models, FDHT-LSTM enjoys both order-of-magnitude reduction (more than$1000 \times$) in model size and significant accuracy improvement (0.6% to 12.7%) across different video recognition datasets. Meanwhile, compared with the state-of-the-art tensor decomposed model-oriented hardware TIE, our proposed FDHT-LSTM architecture achieve$2.5\times$,$1.46\times$and$2.41\times$increase in throughput, area efficiency and energy efficiency, respectively on LSTM-Youtube workload. For LSTM-UCF workload, our proposed design also outperforms TIE with$1.9\times$higher throughput,$1.83\times$higher energy efficiency and comparable area efficiency. Yu Gong 0003, Miao Yin, Lingyi Huang, Chunhua Deng, Bo Yuan 0001 |
IEEE Trans. Computers | 2 |
| 2021 | Doubly Residual Neural Decoder: Towards Low-Complexity High-Performance Channel DecodingabstractRecently deep neural networks have been successfully applied in channel coding to improve the decoding performance. However, the state-of-the-art neural channel decoders cannot achieve high decoding performance and low complexity simultaneously. To overcome this challenge, in this paper we propose doubly residual neural (DRN) decoder. By integrating both the residual input and residual learning to the design of neural channel decoder, DRN enables significant decoding performance improvement while maintaining low complexity. Extensive experiment results show that on different types of channel codes, our DRN decoder consistently outperform the state-of-the-art decoders in terms of decoding performance, model sizes and computational cost. Siyu Liao, Chunhua Deng, Miao Yin, Bo Yuan 0001 |
AAAI | 3 |
| 2021 | Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker StructureabstractRecurrent Neural Networks (RNNs) have been widely used in sequence analysis and modeling. However, when processing high-dimensional data, RNNs typically require very large model sizes, thereby bringing a series of deployment challenges. Although various prior works have been proposed to reduce the RNN model sizes, executing RNN models in the resource-restricted environments is still a very challenging problem. In this paper, we propose to develop extremely compact RNN models with fully decomposed hierarchical Tucker (FDHT) structure. The HT decomposition does not only provide much higher storage cost reduction than the other tensor decomposition approaches, but also brings better accuracy performance improvement for the compact RNN models. Meanwhile, unlike the existing tensor decomposition-based methods that can only decompose the input-to-hidden layer of RNNs, our proposed fully decomposition approach enables the comprehensive compression for the entire RNN models with maintaining very high accuracy. Our experimental results on several popular video recognition datasets show that, our proposed fully decomposed hierarchical tucker-based LSTM (FDHT-LSTM) is extremely compact and highly efficient. To the best of our knowledge, FDHT-LSTM, for the first time, consistently achieves very high accuracy with only few thousand parameters (3,132 to 8,808) on different datasets. Compared with the state-of-the-art compressed RNN models, such as TT-LSTM, TR-LSTM and BT-LSTM, our FDHT-LSTM simultaneously enjoys both order-of-magnitude (3,985× to 10,711×) fewer parameters and significant accuracy improvement (0.6% to 12.7%). Miao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang 0001, Bo Yuan 0001 |
CVPR | 1 |
| 2021 | Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkabstractAdvanced tensor decomposition, such as tensor train (TT) and tensor ring (TR), has been widely studied for deep neural network (DNN) model compression, especially for recurrent neural networks (RNNs). However, compressing convolutional neural networks (CNNs) using TT/TR always suffers significant accuracy loss. In this paper, we propose a systematic framework for tensor decomposition-based model compression using Alternating Direction Method of Multipliers (ADMM). By formulating TT decomposition-based model compression to an optimization problem with constraints on tensor ranks, we leverage ADMM technique to systemically solve this optimization problem in an iterative way. During this procedure, the entire DNN model is trained in the original structure instead of TT format, but gradually enjoys the desired low tensor rank characteristics. We then decompose this uncompressed model to TT format, and fine-tune it to finally obtain a high-accuracy TT-format DNN model. Our framework is very general, and it works for both CNNs and RNNs, and can be easily modified to fit other tensor decomposition approaches. We evaluate our proposed framework on different DNN models for image classification and video recognition tasks. Experimental results show that our ADMM-based TT-format models demonstrate very high compression performance with high accuracy. Notably, on CIFAR-100, with 2.3× and 2.4× compression ratios, our models have 1.96% and 2.21% higher top-1 accuracy than the original ResNet-20 and ResNet-32, respectively. For compressing ResNet-18 on ImageNet, our model achieves 2.47× FLOPs reduction without accuracy loss. Miao Yin, Yang Sui 0001, Siyu Liao, Bo Yuan 0001 |
CVPR | 1 |
| 2021 | CHIP: CHannel Independence-based Pruning for Compact Neural NetworksabstractFilter pruning has been widely used for neural network compression because of its enabled practical acceleration. To date, most of the existing filter pruning works explore the importance of filters via using intra-channel information. In this paper, starting from an inter-channel perspective, we propose to perform efficient filter pruning using Channel Independence, a metric that measures the correlations among different feature maps. The less independent feature map is interpreted as containing less useful information$/$knowledge, and hence its corresponding filter can be pruned without affecting model capacity. We systematically investigate the quantification metric, measuring scheme and sensitiveness$/$reliability of channel independence in the context of filter pruning. Our evaluation results for different models on various datasets show the superior performance of our approach. Notably, on CIFAR-10 dataset our solution can bring $0.75\%$ and $0.94\%$ accuracy increase over baseline ResNet-56 and ResNet-110 models, respectively, and meanwhile the model size and FLOPs are reduced by $42.8\%$ and $47.4\%$ (for ResNet-56) and $48.3\%$ and $52.1\%$ (for ResNet-110), respectively. On ImageNet dataset, our approach can achieve $40.8\%$ and $44.8\%$ storage and computation reductions, respectively, with $0.15\%$ accuracy increase over the baseline ResNet-50 model. The code is available at https://github.com/Eclipsess/CHIP_NeurIPS2021. Yang Sui 0001, Miao Yin, Yi Xie 0001, Huy Phan, Saman A. Zonouz, Bo Yuan 0001 |
NeurIPS | 2 |
| 2019 | Tensor Super-resolution for Seismic DataabstractIn this paper, we propose a novel method for generating high-granularity three-dimensional (3D) seismic data from low-granularity data based on tensor sparse coding, which jointly trains a high-granularity dictionary and a low-granularity dictionary. First, considering the high-dimensional properties of seismic data, we introduce tensor sparse coding to seismic data interpolation. Second, we propose that the dictionary pairs trained by low-granularity seismic data and high-granularity seismic data have the same sparse representation, which are used to recover high-granularity data with the high-granularity dictionary. Finally, experiments on the seismic data of an actual field show that the proposed method effectively perform seismic trace interpolation and can improve the resolution of seismic data imaging. Songjie Liao, Xiao-Yang Liu, Feng Qian 0005, Miao Yin, Guangmin Hu |
ICASSP | 4 |
| 2019 | High-performance Hardware Architecture for Tensor Singular Value Decomposition: Invited PaperabstractTensor provides a brief and natural representation for large-scale multidimensional data by way of appropriate low-rank approximations, thus we can discover significant latent structures of complex data and generalize data representation. To date, tensor has gained tremendous success in various science and technology fields, especially in machine learning and big data applications. However, tensor computation, especially tensor decomposition, is usually expensive due to the inherent large-size characteristic of tensors, and hence would potentially hinder their future wide deployment. In this paper, we develop a hardware architecture to accelerate tensor singular value decomposition (t-SVD), which is a new tensor decomposition technique that has been successfully applied to high-dimensional data classification and video recovery. Specifically, design consideration of each key computing unit is analyzed and discussed. Then, the proposed t-SVD hardware architecture is implemented and synthesized using CMOS 28nm technology. Comparison with real-world CPU-based implementations shows that the proposed hardware accelerator is expected to provide average 14× speedup on various t-SVD workloads. Chunhua Deng, Miao Yin, Xiao-Yang Liu, Xiaodong Wang 0001, Bo Yuan 0001 |
ICCAD | 2 |