EDBT 2026 Demo / reviewers in the wild / expert
Yu Gong 0003
dblp:76/3005-3
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-5465-9044ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language ModelabstractVision-Language Models (VLMs) demand substantial computational resources during inference, largely due to the extensive visual input tokens for representing visual information. Previous studies have noted that visual tokens tend to receive less attention than text tokens, suggesting their lower importance during inference and potential for pruning. However, their methods encounter several challenges: reliance on greedy heuristic criteria for token importance and incompatibility with FlashAttention and KV cache. To address these issues, we introduce TopV, a compatible TOken Pruning with inference Time Optimization for fast and low-memory VLM, achieving efficient pruning without additional training or fine-tuning. Instead of relying on attention scores, we formulate token pruning as an optimization problem, accurately identifying important visual tokens while remaining compatible with FlashAttention. Additionally, since we only perform this pruning once during the prefilling stage, it effectively reduces KV cache size. Our optimization framework incorporates a visual-aware cost function considering factors such as Feature Similarity, Relative Spatial Distance, and Absolute Central Distance, to measure the importance of each source visual token, enabling effective pruning of low-importance tokens. Extensive experiments demonstrate that our method outperforms previous token pruning methods, validating the effectiveness and efficiency of our approach. Cheng Yang 0013, Yang Sui 0001, Jinqi Xiao, Lingyi Huang, Yu Gong 0003, Chendi Li, Jinghua Yan, Yu Bai 0009, P. Sadayappan, Xia Ben Hu, Bo Yuan 0001 |
CVPR | 5 |
| 2025 | Crane: Inter-Layer Scheduling Framework for DNN Inference and Training Co-Support on Tiled ArchitectureabstractTiled architectures have emerged as a compelling platform for scaling deep neural network (DNN) execution, offering both compute density and communication efficiency.To harness their full potential, effective inter-layer scheduling is crucial for managing operation order, memory behavior, and compute resource coordination.However, current schedulers often fall short due to three persistent issues: incomplete treatment of core design factors, limited flexibility in handling diverse workload structures, and reliance on heuristic search algorithms with poor convergence.In this work, we trace these limitations to the absence of a unified and expressive scheduling representation.We introduce Crane, a framework that addresses these gaps through a hierarchical tableformat abstraction capable of encoding rich scheduling semantics.Crane supports both inference and training workloads, and reformulates scheduling as a mathematically structured optimization problem, enabling more complete and efficient exploration of the scheduling space.Evaluations show that Crane reduces energydelay product by up to 21.01× and improves scheduling speed by at least 2.82× over state-of-the-art baselines. Yu Gong 0003, Lingyi Huang, Haodong Chang, Rongjian Liang, Cheng Yang 0013, Zhexiang Tang, Jiang Hu 0001, Bo Yuan 0001 |
MICRO | 1 |
| 2025 | Co-Exploring Structured Sparsification and Low-Rank Tensor Decomposition for Compact DNNsabstractSparsification and low-rank decomposition are two important techniques to compress deep neural network (DNN) models. To date, these two popular yet distinct approaches are typically used in separate ways; while their efficient integration for better compression performance is little explored, especially for structured sparsification and decomposition. In this article, we perform systematic co-exploration on structured sparsification and decomposition toward compact DNN models. We first investigate and analyze several important design factors for joint structured sparsification and decomposition, including operational sequence, decomposition format, and optimization procedure. Based on the observations from our analysis, we then propose CEPD, a unified DNN compression framework that can co-explore the benefits of structured sparsification and tensor decomposition in an efficient way. Empirical experiments demonstrate the promising performance of our proposed solution. Notably, on the CIFAR-10 dataset, CEPD brings 0.72%-0.45% accuracy increase over the baseline ResNet-56 and MobileNetV2 models, respectively, and meanwhile, the computational costs are reduced by 43.0%-44.2%, respectively. On the ImageNet dataset, our approach can enable 0.10%-1.39% accuracy increase over the baseline ResNet-18 and ResNet-50 models with 59.4%-54.6% fewer parameters, respectively. Yang Sui 0001, Miao Yin, Yu Gong 0003, Bo Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Invited: Algorithm and Hardware Co-Design for Energy-Efficient Neural SLAMabstractIn this paper, we introduce a novel approach to enhancing neural network-based Simultaneous Localization and Mapping (SLAM) through the integration of model compression techniques and customized hardware architecture that focuses on micro-architectural and dataflow optimizations to improve computational efficiency and performance. Experiments across different scenarios demonstrate that the proposed approach achieves significant improvement. Lingyi Huang, Cheng Yang 0013, Yu Gong 0003, Yang Sui 0001, Xiao Zang, Anthony Goeckner, Qi Zhu 0002, Bo Yuan 0001 |
DAC | 3 |
| 2024 | DiMO-Sparse: Differentiable Modeling and Optimization of Sparse CNN Dataflow and Hardware ArchitectureabstractMany real-world CNNs exhibit sparsity, a characteristic that has primarily been utilized in manual design processes and has received little attention in existing automatic optimization techniques. To the best of our knowledge, this paper presents the first systematic investigation of automatic dataflow and hardware optimization for sparse CNN computation. A differentiable PPA (Power Performance Area) model incorporating stochastic modeling of sparse CNN workloads is developed to enable fast nonlinear optimization solving and massively parallel local search-based discretization. Experimental results on public domain testcases demonstrate the efficacy of the proposed approach, achieving an average of 5× and 10× better PPA than the previous work for two different sparsity patterns. Jianfeng Song, Rongjian Liang, Yu Gong 0003, Bo Yuan 0001, Jiang Hu 0001 |
DATE | 3 |
| 2024 | MOPED: Efficient Motion Planning Engine with Flexible Dimension SupportabstractMotion planning aims to compute the high-quality and collision-free robotic trajectory. To solve the planning problems defined in varying dimensional sizes, motion planners, especially sampling-based, are typically computation intensive because of the costly kernel operations, and computation inefficient due to the inherent sequential processing scheme, hindering their efficient deployment. To address these challenges and enable real-time highly efficient motion planning, this paper proposes MOPED, an algorithm and hardware co-design for sampling-based motion planning engine with flexible dimension support. At the algorithm level, MOPED proposes a two-stage processing scheme to reduce the frequency and unit cost of collision check. It also fully leverages the spatial information and unique property of planning process to enable low-cost approximated neighbor search. At the hardware level, MOPED proposes a correctness-ensured speculative processing scheme to overcome the serialization problem. It also develop a multi-level caching strategy to reduce data movement and resolve resource conflict. We demonstrate the effectiveness of MOPED via implementing a design example with CMOS 28nm technology via synthesizing. Compared with the baseline motion planning processors, MOPED brings significant improvement on throughput, energy efficiency and area efficiency. Lingyi Huang, Yu Gong 0003, Yang Sui 0001, Xiao Zang, Bo Yuan 0001 |
HPCA | 2 |
| 2023 | HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural NetworksabstractLow-rank compression is an important model compression strategy for obtaining compact neural network models. In general, because the rank values directly determine the model complexity and model accuracy, proper selection of layer-wise rank is very critical and desired. To date, though many low-rank compression approaches, either selecting the ranks in a manual or automatic way, have been proposed, they suffer from costly manual trials or unsatisfied compression performance. In addition, all of the existing works are not designed in a hardware-aware way, limiting the practical performance of the compressed models on real-world hardware platforms. To address these challenges, in this paper we propose HALOC, a hardware-aware automatic low-rank compression framework. By interpreting automatic rank selection from an architecture search perspective, we develop an end-to-end solution to determine the suitable layer-wise ranks in a differentiable and hardware-aware way. We further propose design principles and mitigation strategy to efficiently explore the rank space and reduce the potential interference problem. Experimental results on different datasets and hardware platforms demonstrate the effectiveness of our proposed approach. On CIFAR-10 dataset, HALOC enables 0.07% and 0.38% accuracy increase over the uncompressed ResNet-20 and VGG-16 models with 72.20% and 86.44% fewer FLOPs, respectively. On ImageNet dataset, HALOC achieves 0.9% higher top-1 accuracy than the original ResNet-18 model with 66.16% fewer FLOPs. HALOC also shows 0.66% higher top-1 accuracy increase than the state-of-the-art automatic low-rank compression solution with fewer computational and memory costs. In addition, HALOC demonstrates the practical speedups on different hardware platforms, verified by the measurement results on desktop GPU, embedded GPU and ASIC accelerator. Jinqi Xiao, Chengming Zhang 0006, Yu Gong 0003, Miao Yin, Yang Sui 0001, Lizhi Xiang, Dingwen Tao, Bo Yuan 0001 |
AAAI | 3 |
| 2023 | COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsabstractAttention-based vision models, such as Vision Transformer (ViT) and its variants, have shown promising performance in various computer vision tasks. However, these emerging architectures suffer from large model sizes and high computational costs, calling for efficient model compression solutions. To date, pruning ViTs has been well studied, while other compression strategies that have been widely applied in CNN compression, e.g., model factorization, is little explored in the context of ViT compression. This paper explores an efficient method for compressing vision transformers to enrich the toolset for obtaining compact attention-based vision models. Based on the new insight on the multi-head attention layer, we develop a highly efficient ViT compression solution, which outperforms the state-of-the-art pruning methods. For compressing DeiT-small and DeiT-base models on ImageNet, our proposed approach can achieve $0.45%$ and $0.76%$ higher top-1 accuracy even with fewer parameters. Our finding can also be applied to improve the customization efficiency of text-to-image diffusion models, with much faster training (up to $2.6\times$ speedup) and lower extra storage cost (up to $1927.5\times$ reduction) than the existing works. Jinqi Xiao, Miao Yin, Yu Gong 0003, Xiao Zang, Jian Ren 0005, Bo Yuan 0001 |
ICML | 3 |
| 2023 | ETTE: Efficient Tensor-Train-based Computing Engine for Deep Neural NetworksabstractTensor-train (TT) decomposition enables ultra-high compression ratio, making the deep neural network (DNN) accelerators based on this method very attractive. TIE, the state-of-the-art TT based DNN accelerator, achieved high performance by leveraging a compact inference scheme to remove unnecessary computations and memory access. However, TIE increases memory costs for stage-wise intermediate results and additional intra-layer data transfer, leading to limited speedups even the models are highly compressed. Yu Gong 0003, Miao Yin, Lingyi Huang, Jinqi Xiao, Yang Sui 0001, Chunhua Deng, Bo Yuan 0001 |
ISCA | 1 |
| 2022 | HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural NetworksabstractHigh-order decomposition is a widely used model compression approach towards compact convolutional neural networks (CNNs). However, many of the existing solutions, though can efficiently reduce CNN model sizes, are very difficult to bring considerable saving for computational costs, especially when the compression ratio is not huge, thereby causing the severe computation inefficiency problem. To overcome this challenge, in this paper we propose efficient High-Order DEcomposed Convolution (HODEC). By performing systematic explorations on the underlying reason and mitigation strategy for the computation inefficiency, we develop a new decomposition and computation-efficient execution scheme, enabling simultaneous reductions in computational and storage costs. To demonstrate the effectiveness of HODEC, we perform empirical evaluations for various CNN models on different datasets. HODEC shows consistently outstanding compression and acceleration performance. For compressing ResNet-56 on CIFAR-10 dataset, HODEC brings 67% fewer parameters and 62% fewer FLOPs with 1.17% accuracy increase than the baseline model. For compressing ResNet-50 on ImageNet dataset, HODEC achieves 63% FLOPs reduction with 0.31% accuracy increase than the uncompressed model. Miao Yin, Yang Sui 0001, Wanzhao Yang, Xiao Zang, Yu Gong 0003, Bo Yuan 0001 |
CVPR | 5 |
| 2022 | IMG-SMP: Algorithm and Hardware Co-Design for Real-time Energy-efficient Neural Motion PlanningabstractMotion planning is a fundamental and critical task in modern autonomous systems. Conventionally, motion planning is built on uniform sampling that causes long planning procedure. Recently, built upon the powerful learning and representation abilities of deep neural network (DNN), neural motion planners have attracted a lot of attention because of the better biased sampling strategy learned from data. However, the existing NN-based motion planners are facing several limitations, especially the insufficient exploit of critical spatial information and the high computational cost incurred by neural network models. To overcome these limitations, in this paper we propose IMG-SMP, an algorithm and hardware co-design framework for neural sampling-based motion planner. At the algorithm level, IMG-SMP is an end-to-end neural network that can efficiently capture and process the critical spatial correlation to ensure high planning performance. At the hardware level, by properly rescheduling the computing scheme, the dataflow of IMG-SMP architecture can eliminate the unnecessary computations without affecting planning quality. The IMG-SMP hardware accelerator is implemented and synthesized using CMOS 28nm technology. Evaluation results across different planning tasks show that our proposed hardware design achieves order-of-magnitude improvement over CPU and GPU solutions with respect to planning speed, area efficiency and energy efficiency. Lingyi Huang, Xiao Zang, Yu Gong 0003, Chunhua Deng, Jingang Yi, Bo Yuan 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Hardware Architecture of Graph Neural Network-Enabled Motion Planner (Invited Paper)abstractMotion planning aims to find a collision-free trajectory from the start to goal configurations of a robot. As a key cognition task for all the autonomous machines, motion planning is fundamentally required in various real-world robotic applications, such as 2-D/3-D autonomous navigation of unmanned mobile and aerial vehicles and high degree-of-freedom (DoF) autonomous manipulation of industry/medical robot arms and graspers. Lingyi Huang, Xiao Zang, Yu Gong 0003, Bo Yuan 0001 |
ICCAD | 3 |
| 2022 | Algorithm and Hardware Co-Design of Energy-Efficient LSTM Networks for Video Recognition With Hierarchical Tucker Tensor DecompositionabstractLong short-term memory (LSTM) is a type of powerful deep neural network that has been widely used in many sequence analysis and modeling applications. However, the large model size problem of LSTM networks make their practical deployment still very challenging, especially for the video recognition tasks that require high-dimensional input data. Aiming to overcome this limitation and fully unlock the potentials of LSTM models, in this paper we propose to perform algorithm and hardware co-design towards high-performance energy-efficient LSTM networks. At algorithm level, we propose to developfully decomposed hierarchical Tucker (FDHT)structure-based LSTM, namely FDHT-LSTM, which enjoys ultra-low model complexity while still achieving high accuracy. In order to fully reap such attractive algorithmic benefit, we further develop the corresponding customized hardware architecture to support the efficient execution of the proposed FDHT-LSTM model. With the delicate design of memory access scheme, the complicated matrix transformation can be efficiently supported by the underlying hardware without any access conflict in an on-the-fly way. Our evaluation results show that both the proposed ultra-compact FDHT-LSTM models and the corresponding hardware accelerator achieve very high performance. Compared with the state-of-the-art compressed LSTM models, FDHT-LSTM enjoys both order-of-magnitude reduction (more than$1000 \times$) in model size and significant accuracy improvement (0.6% to 12.7%) across different video recognition datasets. Meanwhile, compared with the state-of-the-art tensor decomposed model-oriented hardware TIE, our proposed FDHT-LSTM architecture achieve$2.5\times$,$1.46\times$and$2.41\times$increase in throughput, area efficiency and energy efficiency, respectively on LSTM-Youtube workload. For LSTM-UCF workload, our proposed design also outperforms TIE with$1.9\times$higher throughput,$1.83\times$higher energy efficiency and comparable area efficiency. Yu Gong 0003, Miao Yin, Lingyi Huang, Chunhua Deng, Bo Yuan 0001 |
IEEE Trans. Computers | 1 |