Cong Guo 0003

dblp:117/1754-3 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0002-4479-5525ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 8 first-author · 23 since 2021Software engineering, systems software and programming languages · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
abstract
The rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantization, there are abundant opportunities for results reuse, and thus it can be boosted with lookup tables (LUTs) based acceleration. However, existing LUT-based methods suffer from computation and hardware overheads for LUT construction, and rely solely on bit-serial computation, which is suboptimal for ternary-weight networks. We propose Platinum, a lightweight ASIC accelerator for integer weight mixed-precision matrix multiplication (mpGEMM) using LUTs. Platinum reduces LUT construction overhead via offline-generated construction paths and supports both general bit-serial and optimized ternaryweight execution through adaptive path switching. On BitNet b1.58-3B, Platinum achieves up to $73.6 \times, 4.09 \times$, and $2.15 \times$ speedups over SpikingEyeriss, Prosperity, and 16-thread T-MAC (CPU), respectively, along with energy reductions of $32.4 \times, 3.23 \times$, and $20.9 \times$, all within a $0.96 \mathrm{~mm}^{2}$ chip area. This demonstrates the potential of LUT-based ASICs as efficient, scalable solutions for ultra-low-bit neural networks on edge platforms.
Haoxuan Shan, Cong Guo 0003, Chiyue Wei, Junyao Zhang 0003, Hai Li 0001, Yiran Chen 0001
ASP-DAC2
2026 M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
abstract
Existing low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format. In this work, we explore strategies that introduce minimal metadata to recover accuracy lost during quantization while maintaining high bit efficiency across a wide range of large language models. We propose a complete algorithm-hardware co-design based on flexible metadata, featuring an online quantization with simple encoding. To support the proposed method efficiently, we implement a lightweight hardware unit and integrate it into the accelerator. Evaluation results demonstrate that our method substantially narrows the accuracy gap, achieving on average a 70.63% reduction in accuracy loss compared to MXFP4 and a 37.30% reduction relative to the latest NVFP4 on LLM benchmarks. Furthermore, our design delivers up to 1.91× speedup and 1.75× energy savings over state-of-the-art accelerators.
Weiming Hu 0005, Chen Zhang 0001, Cong Guo 0003, Yu Feng 0007, Tianchi Hu, Guanglin Li 0005, Guipeng Hu, Jingwen Leng
ASPLOS (2)5
2026 Frame Skipping Architecture for Video-Language Model Acceleration
Haoxuan Shan, Chiyue Wei, Cong Guo 0003, Yuzhe Fu, Hai Li 0001, Yiran Chen 0001
ACM Great Lakes Symposium on VLSI3
2026 FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
abstract
Three-dimensional (3D) point clouds are increasingly used in applications such as autonomous driving, robotics, and virtual reality (VR). Point-based neural networks (PNNs) have demonstrated strong performance in point cloud analysis, originally targeting small-scale inputs. However, as PNNs evolve to process large-scale point clouds with hundreds of thousands of points, all-to-all computation and global memory access in point cloud processing introduce substantial overhead, causing$O\left(n^{2}\right)$computational complexity and memory traffic where$n$is the number of points. Existing accelerators, primarily optimized for small-scale workloads, overlook this challenge and scale poorly due to inefficient partitioning and non-parallel architectures. To address these issues, we propose FractalCloud, a fractal-inspired hardware architecture for efficient large-scale 3D point cloud processing. FractalCloud introduces two key optimizations: (1) a co-designed Fractal method for shape-aware and hardware-friendly partitioning, and (2) block-parallel point operations that decompose and parallelize all point operations. A dedicated hardware design with on-chip fractal and flexible parallelism further enables fully parallel processing within limited memory resources. Implemented in 28 nm technology as a chip layout with a core area of$1.5 ~\text{mm}^{2}$, FractalCloud achieves$21.7 \times$speedup and$27 \times$energy reduction over state-of-the-art accelerators while maintaining network accuracy, demonstrating its scalability and efficiency for PNN inference. The code for FractalCloud is available at https://github.com/Yuzhe-Fu/FractalCloud.
Yuzhe Fu, Changchun Zhou 0001, Hancheng Ye, Bowen Duan 0003, Qiyu Huang, Chiyue Wei, Cong Guo 0003, Hai Li 0001, Yiran Chen 0001
HPCA7
2026 Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
abstract
Vision-Language Models (VLMs) have demonstrated strong performance on tasks such as video captioning and visual question answering. However, their growing scale and video-level inputs lead to significant computational and memory overhead, posing challenges for real-time deployment on hardware accelerators. While prior work attempts to reduce redundancy via token pruning or merging, these methods typically operate at coarse granularity and incur high runtime overhead due to global token-level operations. In this study, we propose Focus, a Streaming Concentration Architecture that efficiently accelerates VLM inference through progressive, fine-grained redundancy elimination. Focus introduces a multilevel concentration paradigm that hierarchically compresses vision-language inputs at three levels: (1) semantic-guided token pruning based on textual prompts, (2) spatial-temporal blocklevel concentration using localized comparisons, and (3) vectorlevel redundancy removal via motion-aware matching. All concentration steps are tightly co-designed with the architecture to support streaming-friendly, on-chip execution. Focus leverages GEMM tiling, convolution-style layout, and cross-modal attention to minimize off-chip access while enabling high throughput. Implemented as a modular unit within a systolic-array accelerator, Focus achieves$2.4 \times$speedup and$3.3 \times$reduction in energy, significantly outperforming state-of-the-art accelerator in both performance and energy efficiency. Full-stack implementation of Focus is open-sourced at https://github.com/dubcyfor3/Focus.
Chiyue Wei, Cong Guo 0003, Junyao Zhang 0003, Haoxuan Shan, Qinsi Wang, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001
HPCA2
2026 EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan 0003, Cong Guo 0003, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001
ISCA2
2026 A Full-Stack Framework for GNN Acceleration via Partition-Compiler-Architecture Co-Design
abstract
Graph Neural Networks (GNNs) have achieved remarkable success across domains such as recommendation and scientific computing, yet their practical deployment remains constrained by high execution cost. The diversity of GNN model structures and the sparsity of real-world graphs pose two fundamental challenges for hardware acceleration: supporting heterogeneous operator patterns and achieving high resource utilization under irregular data access. Existing accelerators often address only one aspect, either targeting specific models with hardwired pipelines or applying general architectures with limited efficiency. To address these challenges, we propose SWITCHBLADE, a full-stack framework for GNN acceleration through the coordinated design of partitioning, compilation, and architecture. SWITCHBLADE addresses these challenges through three key components. First, a phase-based intermediate representation unifies diverse GNN models by abstracting computation stages for model-independent code generation. Second, a fine-grained graph partitioner enhances data locality and reduces memory traffic by adapting to graph topology and model semantics. Third, the hardware architecture supports stream-level parallelism and decoupled execution to exploit cross-shard and inter-phase concurrency. Evaluation on representative models and datasets shows that SWITCHBLADE achieves up to 1.85× speedup and 19.03× energy savings over an NVIDIA V100 GPU, while outperforming state-of-the-art GNN accelerators across diverse full-graph workloads, demonstrating both high efficiency and broad model generality.
Yangjie Zhou 0001, Shuwen Lu, Cong Guo 0003, Jingwen Leng, Yufei Ma 0002, Yun Liang 0001, Minyi Guo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
abstract
Vector quantization (VQ), which treats a vector as a compression unit, gains increasing research interests for its potential to accelerate large language models (LLMs). Compared to conventional element-wise quantization methods, VQ algorithms can compress weight and KV cache tensors in LLMs with a greater ratio while maintaining the high model accuracy. However, translating a VQ algorithm’s memory reduction into the actual latency improvement is challenging. We profile and analyze the current approach of integrating VQ into computation kernels and show that its major inefficiency lies in the poor access efficiency of codebooks in VQ algorithms and uncoordinated computation dataflow. Meanwhile, the diversity of VQ algorithms (e.g., different vector sizes and entry counts) and LLMs, computation kernels (e.g matrix-matrix/vector multiplication and attention computation) makes it impractical to manually craft efficient kernel implementations for each specific case. In this work, we design and implement VQ-LLM, an efficient fused VQ kernel generation framework. We first introduce a software abstraction called codebook cache to optimize codebook access efficiency and support the integration of VQ with various computations. The codebook cache adaptively stores different entries across the GPU’s memory hierarchy, including off-chip global memory, on-chip shared memory, and registers. Centered around the codebook cache, we design an efficient computation engine that optimizes memory traffic during computations involving codebooks. This compute engine adopts the codebook-centric dataflow and fusion optimizations. Additionally, we provide adaptive heuristics to tailor parameter selection in our optimizations to diverse VQ configurations. Our optimizations achieve the latency reduction of $\mathbf{6 4. 3 6 \%}$ to $\mathbf{9 9. 1 \%}$ compared to existing open-source implementations. A final comparison with state-of-the-art element-wise quantization methods like AWQ and QoQ shows that our VQ-LLM is practically viable, achieving latencies close or even better latencies to those at equivalent bit-widths, potentially offering greater accuracy.
Zihan Liu 0002, Xinhao Luo, Junxian Guo, Wentao Ni, Yangjie Zhou 0001, Yue Guan 0003, Cong Guo 0003, Weihao Cui, Yu Feng 0007, Minyi Guo, Yuhao Zhu 0001, Minjia Zhang, Jingwen Leng
HPCA7
2025 M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
abstract
Large language models (LLMs) are one of the most important killer computer applications. The recent algorithmic advancement proposes a fine-grained group-wise quantization for LLMs, which treats a small set (e.g., 64) of values in a tensor as a compression unit. It effectively preserves the model accuracy without retraining, and has become the standard approach to efficiently deploy LLMs. On the other hand, there are works that propose various adaptive data types to better adapt to different distributions and further reduce the required bit length for LLMs. In this work, our detailed analysis unveils a key finding that while different tensors exhibit similar distributions, small groups can have markedly different distributions. As such, the group-level diversity requires a new level of adaptivity for which existing adaptive data types fail to provide.In this paper, we propose MANT, a mathematically adaptive numeric type, featuring a more flexible encoding paradigm with a wider range of data distribution and more efficient decoding-computation fusion mechanism to address these challenges. Based on MANT, we develop a supporting framework to assign the appropriate data type for each group adaptively. Meanwhile, the dynamically generated Key-Value (KV) caches in LLMs introduce further complexity for real-time quantization. To tackle this, we propose an efficient real-time quantization mechanism. Besides, we implement a specific processing element (PE) to efficiently support MANT and incorporate a real-time quantization unit. By integrating these components into a systolic array, MANT unifies the group-wise weight and KV cache quantization and addresses the associated challenges. Our evaluation shows achieving, on average, 2.99 × (up to 4.46 ×) speedup and 2.81 × (up to 4.10 ×) energy reduction to the state-of-the-art LLM accelerator.
Weiming Hu 0005, Cong Guo 0003, Yu Feng 0007, Renyang Guan, Zhendong Hua, Zihan Liu 0002, Yue Guan 0003, Minyi Guo, Jingwen Leng
HPCA3
2025 Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
abstract
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementations of SNNs leverage this sparsity pattern to avoid wasteful zero-value computations, yet this approach fails to fully capitalize on the potential efficiency of SNNs. This study introduces a novel sparsity paradigm called Product Sparsity, which leverages combinatorial similarities within matrix multiplication operations to reuse the inner product result and reduce redundant computations. Product Sparsity significantly enhances sparsity in SNNs without compromising the original computation results compared to traditional bit sparsity methods. For instance, in the SpikeBERT SNN model, Product Sparsity achieves a density of only 1.23% and reduces computation by $11 \times$, compared to bit sparsity, which has a density of 13.19%. To efficiently implement Product Sparsity, we propose Prosperity, an architecture that addresses the challenges of identifying and eliminating redundant computations in real-time. Compared to prior SNN accelerator PTB and the A100 GPU, Prosperity achieves an average speedup of $7.4 \times$ and $1.8 \times$, respectively, along with energy efficiency improvements of $8.0 \times$ and $193 \times$, respectively. The code for Prosperity is available at https://github.com/dubcyfor3/Prosperity.
Chiyue Wei, Cong Guo 0003, Shiyu Li 0001, Hao (Frank) Yang, Hai Li 0001, Yiran Chen 0001
HPCA2
2025 Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge
abstract
This paper presents Edge-based Mixture of Experts (MoE) Collaborative Computing (EMC2), an optimal computing system designed for autonomous vehicles (AVs) that simultaneously achieves low-latency and high-accuracy 3D object detection. Unlike conventional approaches, EMC2 incorporates a scenario-aware MoE architecture specifically optimized for edge platforms. By effectively fusing LiDAR and camera data, the system leverages the complementary strengths of sparse 3D point clouds and dense 2D images to generate robust multimodal representations. To enable this, EMC2 employs an adaptive multimodal data bridge that performs multi-scale preprocessing on sensor inputs, followed by a scenario-aware routing mechanism that dynamically dispatches features to dedicated expert models based on object visibility and distance. In addition, EMC2 integrates joint hardware-software optimizations, including hardware resource utilization optimization and computational graph simplification, to ensure efficient and real-time inference on resource-constrained edge devices. Experiments on open-source benchmarks clearly show the EMC2 advancements as an end-to-end system. On the KITTI dataset, it achieves an average accuracy improvement of 3.58% and a 159.06% inference speedup compared to 15 baseline methods on Jetson platforms, with similar performance gains on the nuScenes dataset, highlighting its capability to advance reliable, real-time 3D object detection tasks for AVs. The official implementation is available at https://github.com/LinshenLiu622/EMC2.
Linshen Liu, Boyan Su, Junyue Jiang, Guanlin Wu, Cong Guo 0003, Ceyu Xu, Hao (Frank) Yang
ICCV5
2025 Transitive Array: An Efficient GEMM Accelerator with Result Reuse
abstract
Deep Neural Networks (DNNs) and Large Language Models (LLMs) have revolutionized artificial intelligence, yet their deployment faces significant memory and computational challenges, especially in resource-constrained environments.Quantization techniques have mitigated some of these issues by reducing data precision, primarily focusing on General Matrix Multiplication (GEMM).This study introduces a novel sparsity paradigm, transitive sparsity, which leverages the reuse of previously computed results to substantially minimize computational overhead in GEMM operations.By representing transitive relations using a directed acyclic graph, we develop an efficient strategy for determining optimal execution orders, thereby overcoming inherent challenges related to execution dependencies and parallelism.Building on this foundation, we present the Transitive Array, a multiplication-free accelerator designed to exploit transitive sparsity in GEMM.Our architecture effectively balances computational workloads across multiple parallel lanes, ensuring high efficiency and optimal resource utilization.Comprehensive evaluations demonstrate that the Transitive Array achieves approximately 7.46× and 3.97× speedup and 2.31× and 1.65× energy reduction compared to state-of-the-art accelerators such as Olive and BitVert while maintaining comparable model accuracy on LLaMA models.
Cong Guo 0003, Chiyue Wei, Bowen Duan 0003, Song Han 0003, Hai Li 0001, Yiran Chen 0001
ISCA1
2025 Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-Aware Cache Compression
abstract
Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memory and computational demands, especially in resource-constrained environments.Quantization techniques have emerged as a critical solution, reducing data precision to enhance memory and computational efficiency.However, existing methods often suffer from high runtime overheads and potential accuracy degradation.To address these challenges, we propose Ecco, an entropy-based cache compression technique tailored for LLMs.Ecco combines group-wise and nonuniform quantization with pre-defined shared k-means patterns and Huffman coding to exploit the inherent entropy characteristics of LLM cache data.Recognizing the inefficiencies of traditional Huffman coding in terms of parallelism and latency, we introduce a novel parallel Huffman-based decoding process with a multi-stage pipeline design, reducing latency by two orders of magnitude and achieving throughput comparable to GPU L2 caches.Comprehensive evaluations demonstrate that Ecco achieves an up to 2.9× and 1.9× speedup over the state-of-the-art AWQ and SmoothQuant framework, 2.4× over the Olive accelerator, all while increasing memory capacity by nearly 4× and maintaining state-of-the-art LLM accuracy.These results underscore the effectiveness of our
Cong Guo 0003, Chiyue Wei, Junyao Zhang 0003, Changchun Zhou 0001, Edward Hanson, Jiaqi Zhang 0002, Xiaoxiao Liu 0001, Hai Li 0001, Yiran Chen 0001
ISCA2
2025 Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are gaining attention for their energy efficiency and biological plausibility, utilizing 0-1 activation sparsity through spike-driven computation.While existing SNN accelerators exploit this sparsity to skip zero computations, they often overlook the unique distribution patterns inherent in binary activations.In this work, we observe that particular patterns exist in spike activations, which we can utilize to reduce the substantial computation of SNN models.Based on these findings, we propose a novel pattern-based hierarchical sparsity framework, termed Phi, to optimize computation.Phi introduces a two-level sparsity hierarchy: Level 1 exhibits vector-wise sparsity by representing activations with pre-defined patterns, allowing for offline pre-computation with weights and significantly reducing most runtime computation.Level 2 features element-wise sparsity by complementing the Level 1 matrix, using a highly sparse matrix to further reduce computation while maintaining accuracy.We present an algorithm-hardware co-design approach.Algorithmically, we employ a k-means-based pattern selection method to identify representative patterns and introduce a pattern-aware fine-tuning technique to enhance Level 2 sparsity.Architecturally, we design Phi, a dedicated hardware architecture that efficiently processes the two levels of Phi sparsity on the fly.Extensive experiments demonstrate that Phi achieves a 3.45× speedup and a 4.93× improvement in energy efficiency compared to stateof-the-art SNN accelerators, showcasing the effectiveness of our framework in optimizing SNN computation.
Chiyue Wei, Bowen Duan 0003, Cong Guo 0003, Jingyang Zhang, Qingyue Song, Hai Li 0001, Yiran Chen 0001
ISCA3
2025 A Sample-Free Compilation Framework for Efficient Dynamic Tensor Computation
abstract
Dynamic-shape tensor computation poses challenges for shape-specific compilation due to variable input dimensions. Existing compilers rely on shape samples, incurring high tuning costs and performance degradation on unseen inputs. We present Helix, a dynamic tensor compilation framework with sample-free compilation and architecture-guided optimization to achieve both compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction to align tensor computations with architectural hierarchies, and bottom-up kernel construction to build efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime. This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations conducted on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174 × over the existing compilers and delivers 2.26 × and 3.29 × execution speedups over vendor libraries and dynamic-shape compilers, respectively.
Yangjie Zhou 0001, Weihao Cui, Zihan Liu 0002, Peng Chen 0035, Mohamed Wahib, Cong Guo 0003, Siyuan Feng 0007, Jintao Meng 0001, Haidong Lan, Jingwen Leng, Yun Lin 0001, Jin Song Dong 0001, Wenxi Zhu, Minwen Deng
SC8
2025 DSTC: Dual-Side Sparse Tensor Core for DNNs Acceleration on Modern GPU Architectures
abstract
Leveraging sparsity in deep neural network (DNN) models holds significant promise for accelerating model inference. However, current GPUs can only harness sparsity in model weights, leaving activations unutilized due to their dynamic and unpredictable nature, which poses a considerable challenge for exploitation. In our research, we introduce a novel architectural approach aimed at effectively leveraging dual-side sparsity, encompassing both weight and activation sparsity. Our methodology involves a systematic examination of previous sparsity-related architectures, and culminating in the proposal of an uncharted paradigm that combines outer-product computation primitive and bitmap-based encoding format. Our approach showcases feasibility through minimal modifications to existing production-scale inner-product-based Tensor Cores. We introduce a set of innovative ISA extensions and carefully co-design matrix-matrix multiplication and convolution algorithms, the two predominant computation patterns in contemporary DNN models, to exploit our novel dual-side sparse Tensor Core. Our evaluation demonstrates the efficacy of our design, unlocking the full potential of dual-side DNN sparsity and delivering performance enhancements of up to an order of magnitude while incurring only modest hardware overhead.
Chen Zhang 0001, Yang Wang 0053, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng, Zhigang Ji, Yuan Xie 0001, Ru Huang 0001
IEEE Trans. Computers4
2024 GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
abstract
Large-scale deep neural networks (DNNs), such as large language models (LLMs), have revolutionized the artificial intelligence (AI) field and become increasingly popular. However, training or fine-tuning such models requires substantial computational power and resources, where the memory capacity of a single acceleration device like a GPU is one of the most important bottlenecks. Owing to the prohibitively large overhead (e.g., 10×) of GPUs' native memory allocator, DNN frameworks like PyTorch and TensorFlow adopt a caching allocator that maintains a memory pool with a splitting mechanism for fast memory (de)allocation. Unfortunately, the caching allocator's efficiency degrades quickly for popular memory reduction techniques such as re-computation, offloading, distributed training, and low-rank adaptation. The primary reason is that those memory reduction techniques introduce frequent and irregular memory (de)allocation requests, leading to severe fragmentation problems for the splitting-based caching allocator. To mitigate this fragmentation problem, we propose a novel memory allocation framework based on low-level GPU virtual memory management called GPU memory lake (GMLake). GMLake employs a novel virtual memory stitching (VMS) mechanism, which can fuse or combine non-contiguous memory blocks with a virtual memory address mapping. GMLake can reduce average of 9.2 GB (up to 25 GB) GPU memory usage and 15% (up to 33%) fragmentation among eight LLM models on GPU A100 with 80 GB memory. GMLake is completely transparent to the DNN models and memory reduction techniques and ensures the seamless execution of resource-intensive deep-learning tasks. We have open-sourced GMLake at https://github.com/intelligent-machine-learning/glake/tree/main/GMLake.
Cong Guo 0003, Rui Zhang 0040, Jingwen Leng, Zihan Liu 0002, Minyi Guo, Shouren Zhao, Junping Zhao, Ke Zhang 0048
ASPLOS (2)1
2024 JUNO: Optimizing High-Dimensional Approximate Nearest Neighbour Search with Sparsity-Aware Algorithm and Ray-Tracing Core Mapping
abstract
Approximate nearest neighbor (ANN) search is a widely applied technique in modern intelligent applications, such as recommendation systems and vector databases. Therefore, efficient and high-throughput execution of ANN search has become increasingly important. In this paper, we first characterize the state-of-the-art product quantization-based method of ANN search and identify a significant source of inefficiency in the form of unnecessary pairwise distance calculations and accumulations. To improve efficiency, we propose Juno, an end-to-end ANN search system that adopts a carefully designed sparsity- and locality-aware search algorithm. We also present an efficient hardware mapping that utilizes ray tracing cores in modern GPUs with pipelined execution on tensor cores to execute our sparsity-aware ANN search algorithm. Our evaluations on four datasets from 1 to 100 million search points demonstrate 2.2×-8.5× improvements in search throughput. Moreover, our algorithmic enhancements alone achieve a maximal 2.6× improvement on the hardware without the acceleration of the RT core.
Zihan Liu 0002, Wentao Ni, Jingwen Leng, Yu Feng 0007, Cong Guo 0003, Quan Chen 0002, Chao Li 0009, Minyi Guo, Yuhao Zhu 0001
ASPLOS (2)5
2024 Accelerating Sparse DNNs Based on Tiled GEMM
abstract
Network pruning can reduce the computation cost of deep neural network (DNN) models. However, sparse models often produce randomly-distributed weights to maintain accuracy, leading to irregular computations. Consequently, unstructured sparse models cannot achieve meaningful speedup on commodity hardware built for dense matrix computations. Accelerators are usually modified or designed with structured sparsity-optimized architectures for exploiting sparsity. For example, the Ampere architecture introduces a sparse tensor core, which adopts the 2:4 sparsity pattern.We propose a pruning method that builds upon the insight that matrix multiplication generally breaks the large matrix into multiple smaller tiles for parallel execution. We present the “tile-wise” sparsity pattern, which maintains a structured sparsity pattern at the tile level for efficient execution but allows for irregular pruning at the global scale to maintain high accuracy. In addition, the tile-wise sparsity is implemented at the global memory level, and the 2:4 sparsity executes at the register level inside the sparse tensor core. We can combine these two patterns into a “tile-vector-wise” (TVW) sparsity pattern to explore more fine-grained sparsity and further accelerate the sparse DNN models. We evaluate the TVW on the GPU, achieving averages of 1:85×, 2:75×, and 22:18× speedups over the dense model, block sparsity, and unstructured sparsity.
Cong Guo 0003, Fengchen Xue, Jingwen Leng, Yuxian Qiu, Yue Guan 0003, Weihao Cui, Quan Chen 0002, Minyi Guo
IEEE Trans. Computers1
2023 AdaptGear: Accelerating GNN Training via Adaptive Subgraph-Level Kernels on GPUs
abstract
Graph neural networks (GNNs) are powerful tools for exploring and learning from graph structures and features. As such, achieving high-performance execution for GNNs becomes crucially important. Prior works have proposed to explore the sparsity (i.e., low density) in the input graph to accelerate GNNs, which uses the full-graph-level or block-level sparsity format. We show that they fail to balance the sparsity benefit and kernel execution efficiency. In this paper, we propose a novel system, referred to as AdaptGear, that addresses the challenge of optimizing GNNs performance by leveraging kernels tailored to the density characteristics at the subgraph level. Meanwhile, we also propose a method that dynamically chooses the optimal set of kernels for a given input graph. Our evaluation shows that AdaptGear can achieve a significant performance improvement, up to 6.49× (1.87× on average), over the state-of-the-art works on two mainstream NVIDIA GPUs across various datasets.
Yangjie Zhou 0001, Yaoxu Song, Jingwen Leng, Zihan Liu 0002, Weihao Cui, Zhendong Zhang 0004, Cong Guo 0003, Quan Chen 0002, Li Li 0012, Minyi Guo
CF7
2023 OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
abstract
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by 240× every two years, which outpaces the hardware progress and makes model inference increasingly costly. Model quantization is a promising approach to mitigate the widening gap between LLM size and hardware capacity. However, the existence of outliers, values with significant magnitudes, in LLMs makes existing quantization methods less effective. Prior outlier-aware quantization schemes adopt sparsity encoding techniques to separate outliers from normal values where the process requires global coordination (e.g., a global sparsity coordination list). This incurs complex encoding/decoding hardware logics and an extra orchestration controller for the computation between outlier and normal values. As such, it is not hardware-efficient and hence only achieves sub-optimal quantization benefits.
Cong Guo 0003, Weiming Hu 0005, Jingwen Leng, Chen Zhang 0001, Fan Yang 0024, Yunxin Liu 0001, Minyi Guo, Yuhao Zhu 0001
ISCA1
2022 Nesting Forward Automatic Differentiation for Memory-Efficient Deep Neural Network Training
abstract
An activation function is an element-wise mathematical function and plays a crucial role in deep neural networks (DNN). Many novel and sophisticated activation functions have been proposed to improve the DNN accuracy but also consume massive memory in the training process with back-propagation. In this study, we propose the nested forward automatic differentiation (Forward-AD), specifically for the element-wise activation function for memory-efficient DNN training. We deploy nested Forward-AD in two widely-used deep learning frameworks, TensorFlow and PyTorch, which support the static and dynamic computation graph, respectively. Our evaluation shows that nested Forward-AD reduces the memory footprint by up to 1.97× than the baseline model and outperforms the recomputation by 20% under the same memory reduction ratio.
Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Chen Zhang 0001, Quanlu Zhang, Yunxin Liu 0001, Fan Yang 0024, Minyi Guo
ICCD1
2022 SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation
Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Xiaotian Gao, Chen Zhang 0001, Yunxin Liu 0001, Fan Yang 0024, Yuhao Zhu 0001, Minyi Guo
ICLR1
2022 ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization
abstract
Quantization is a technique to reduce the computation and memory cost of DNN models, which are getting increasingly large. Existing quantization solutions use fixed-point integer or floating-point types, which have limited benefits, as both require more bits to maintain the accuracy of original models. On the other hand, variable-length quantization uses low-bit quantization for normal values and high-precision for a fraction of outlier values. Even though this line of work brings algorithmic benefits, it also introduces significant hardware overheads due to variable-length encoding and decoding.In this work, we propose a fixed-length a daptive n umerical data t ype called ANT to achieve low-bit quantization with tiny hardware overheads. Our data type ANT leverages two key innovations to exploit the intra-tensor and inter-tensor adaptive opportunities in DNN models. First, we propose a particular data type, flint, that combines the advantages of float and int for adapting to the importance of different values within a tensor. Second, we propose an adaptive framework that selects the best type for each tensor according to its distribution characteristics. We design a unified processing element architecture for ANT and show its ease of integration with existing DNN accelerators. Our design results in $2.8\times $ speedup and $2.5\times $ energy efficiency improvement over the state-of-the-art quantization accelerators.
Cong Guo 0003, Chen Zhang 0001, Jingwen Leng, Zihan Liu 0002, Fan Yang 0024, Yunxin Liu 0001, Minyi Guo, Yuhao Zhu 0001
MICRO1
2022 Towards Reliable AI Applications via Algorithm-Based Fault Tolerance on NVDLA
abstract
With the development of deep neural networks (DNNs), more complex accelerators have been designed for more sophisticated networks. Naturally, the complexity of accelerators makes them vulnerable to transient errors. Also, some DNN accelerators are widely used the safety-critical systems, such as autonomous vehicles. Therefore, the susceptibility to transient errors makes research on mitigation techniques more significant, and errors of accelerators should be limited to none. Some researchers proposed the modular redundancy method, which offers a highly reliable way but also considerably increases overhead. In this regard, algorithm-based solutions offer cheaper solutions. However, their implementation is primarily observed in software-based error injections. In this study, we propose a novel approach that focuses on implementing algorithm-based error detection (ABED) for RTL-level (hardware-based) error injections. Previous studies generally focused on the impact of soft errors in memory structures of embedded system-based accelerators. However, the main goal of this research is to study the impact of soft errors in processing elements and how to mitigate them. We implement an algorithm-based error detection that utilizes checksums for verifying convolution operations with low overhead. We first explain how to overcome the challenges of implementing ABED on FPGA-based accelerators, then how to implement it. We implement and evaluate our solution on an industry-level DNN accelerator called NVIDIA deep learning accelerator (NVDLA). In this study, our error injection method is constructed to test the most common soft error scenarios in processing units. The results of the research show that algorithm-based fault tolerance can detect all silent data corruptions (SDC) while maintaining a very low overhead (6-23%) on runtime.
Mustafa Sanic, Cong Guo 0003, Jingwen Leng, Minyi Guo, Weiyin Ma
MSN2
2021 Dual-side Sparse Tensor Core
abstract
Leveraging sparsity in deep neural network (DNN) models is promising for accelerating model inference. Yet existing GPUs can only leverage the sparsity from weights but not activations, which are dynamic, unpredictable, and hence challenging to exploit. In this work, we propose a novel architecture to efficiently harness the dual-side sparsity (i.e., weight and activation sparsity). We take a systematic approach to understand the (dis)advantages of previous sparsity-related architectures and propose a novel, unexplored paradigm that combines outer-product computation primitive and bitmap-based encoding format. We demonstrate the feasibility of our design with minimal changes to the existing production-scale inner-product-based Tensor Core. We propose a set of novel ISA extensions and co-design the matrix-matrix multiplication and convolution algorithms, which are the two dominant computation patterns in today’s DNN models, to exploit our new dual-side sparse Tensor Core. Our evaluation shows that our design can fully unleash the dual-side DNN sparsity and improve the performance by up to one order of magnitude with small hardware overhead.
Yang Wang 0053, Chen Zhang 0001, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng
ISCA4
2020 Balancing Efficiency and Flexibility for DNN Acceleration via Temporal GPU-Systolic Array Integration
abstract
The research interest in specialized hardware accelerators for deep neural networks (DNN) spikes recently owing to their superior performance and efficiency. However, today’s DNN accelerators primarily focus on accelerating specific "kernels" such as convolution and matrix multiplication, which are vital but only part of an end-to-end DNN-enabled application. Meaningful speedups over the entire application often require supporting computations that are, while massively parallel, ill-suited to DNN accelerators. Integrating a general-purpose processor such as a CPU or a GPU incurs significant data movement overhead and leads to resource under-utilization on the DNN accelerators.We propose Simultaneous Multi-mode Architecture (SMA), a novel architecture design and execution model that offers general-purpose programmability on DNN accelerators in order to accelerate end-to-end applications. The key to SMA is the temporal integration of the systolic execution model with the GPU-like SIMD execution model. The SMA exploits the common components shared between the systolic-array accelerator and the GPU, and provides lightweight reconfiguration capability to switch between the two modes in-situ. The SMA achieves up to 63% performance improvement while consuming 23% less energy than the baseline Volta architecture with TensorCore.
Cong Guo 0003, Yangjie Zhou 0001, Jingwen Leng, Yuhao Zhu 0001, Zidong Du, Quan Chen 0002, Chao Li 0009, Bin Yao 0002, Minyi Guo
DAC1
2020 Accelerating sparse DNN models without hardware-support via tile-wise sparsity
abstract
Network pruning can reduce the high computation cost of deep neural network (DNN) models. However, to maintain their accuracies, sparse models often carry randomly-distributed weights, leading to irregular computations. Consequently, sparse models cannot achieve meaningful speedup on commodity hardware (e.g., GPU) built for dense matrix computations. As such, prior works usually modify or design completely new sparsity-optimized architectures for exploiting sparsity. We propose an algorithm-software co-designed pruning method that achieves latency speedups on existing dense architectures. Our work builds upon the insight that the matrix multiplication generally breaks the large matrix into multiple smaller tiles for parallel execution. We propose a tiling-friendly “tile-wise” sparsity pattern, which maintains a regular pattern at the tile level for efficient execution but allows for irregular, arbitrary pruning at the global scale to maintain the high accuracy. We implement and evaluate the sparsity pattern on GPU tensor core, achieving a 1.95× speedup over the dense model.
Cong Guo 0003, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu, Yue Guan 0003, Zehuan Wang 0001, Xiaoying Jia 0001, Xipeng Li, Minyi Guo, Yuhao Zhu 0001
SC1
2019 Adversarial Defense Through Network Profiling Based Path Extraction
abstract
Recently, researchers have started decomposing deep neural network models according to their semantics or functions. Recent work has shown the effectiveness of decomposed functional blocks for defending adversarial attacks, which add small input perturbation to the input image to fool the DNN models. This work proposes a profiling-based method to decompose the DNN models to different functional blocks, which lead to the effective path as a new approach to exploring DNNs' internal organization. Specifically, the per-image effective path can be aggregated to the class-level effective path, through which we observe that adversarial images activate effective path different from normal images. We propose an effective path similarity-based method to detect adversarial images with an interpretable model, which achieve better accuracy and broader applicability than the state-of-the-art technique.
Yuxian Qiu, Jingwen Leng, Cong Guo 0003, Quan Chen 0002, Chao Li 0009, Minyi Guo, Yuhao Zhu 0001
CVPR3