EDBT 2026 Demo / reviewers in the wild / expert
Jun Liu 0117
dblp:95/3736-117
· DBLP profile ↗
18ranked-venue papers
5as first author
18since 2021 · last 2026
0009-0003-8280-9072ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 16 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FAST: A Scalable Framework for Accelerating Flexible Structured Sparse TrainingabstractSparse training is a critical approach to reducing the storage requirement while maintaining the model’s ability. However, it is non-trivial to apply the flexible structured sparsity (flex-SS) patterns during sparse training, which achieves Pareto optimality in terms of hardware efficiency and flexibility. we propose FAST, a fast and scalable framework that supports LLM training with flex-SS patterns. First, we propose a probability-based decoupling method that eliminates dependencies between tiles to generate the flex-SS mask efficiently. Second, we propose a weight-distribution-aware pivot search strategy that narrows down the available region of pivot candidates to reduce the communication overhead. Extensive experimental results show that FAST achieves up to 10.40× and 1.56× end-to-end training speedup compared with PyTorch and the SOTA framework. Shuaiheng Li, Jun Liu 0117, Yaoxiu Lian, Tianlang Zhao, Li Ding 0012, Guohao Dai 0001 |
DATE | 2 |
| 2026 | Endor: Exploit Nearly-Decode-Only Opportunities of LLM Reasoning on Near-Memory ArchitectureabstractReasoning with Large Language Models (LLMs) has become a pivotal research topic because their logical abilities significantly surpass those of standard LLMs. LLM reasoning typically forms multiple chains of thought, action-by-action, and selects the best one as the final answer. However, the inference overhead of LLM reasoning is more than an order of magnitude higher than that of LLM. Despite the emerging shift towards memory-optimized algorithms and near-memory hardware, we still face the following challenges: (1) Existing memory-centric algorithms (e.g., KV cache technique) have low computational utilization (< 4% on NVIDIA A100 GPU) due to intensive memory access for inter-action data. (2) Emerging hardware architectures (e.g., near-memory processing) fail to fully utilize the inherent parallelism due to dependencies among models, leading to low utilization of memory bandwidth.To tackle these challenges, we propose Endor, a hardware-algorithm co-design to accelerate the inference of LLM reasoning efficiently. We identify that the auto-regressive decoding of LLM reasoning changes from the token level to the action level in terms of the computing paradigm. At the algorithm level, we propose a "nearly-decode-only" method which encompasses an efficient inter-action cache reuse method and a prediction-based pipeline optimization to reduce computation overhead. At the hardware level, we propose Endor-NMP, a near-memory accelerator featuring a score-aware cache management architecture and a heterogeneous mapping dataflow. Endor fully exploits both interaction and intra-action parallelism to improve memory bandwidth utilization. Experimental results demonstrate that neither existing algorithms nor hardware can achieve the expected acceleration. Endor achieves an end-to-end average speedup of 2.97× and 2.52× compared to the NVIDIA A100 GPU and advanced LLM accelerators on multiple models and datasets. Jun Liu 0117, Tianlang Zhao, Jiancai Ye, Lin Li 0002, Li Ding 0012, Hao Zhou 0008, Zhenhua Zhu 0002, Xuefei Ning, Yuan Xie 0001, Yu Wang 0002, Guohao Dai 0001 |
DATE | 1 |
| 2026 | MARCA-v2: Mamba Accelerator With Complementary State-Space Model Sparsity and Reconfigurable ArchitectureabstractLarge Language models with state space model (SSM) especially Mamba have demonstrated remarkable capabilities in various domains. Compared to Transformers, Mamba reduces the quadratic computational complexity and achieves a higher algorithm performance. Current research on Mamba focuses primarily on integrating it with various application scenarios. However, there is limited research on optimizing for Mamba processing. Therefore, we profile the processing carefully and identify three main challenges in Mamba computations: (1) Large memory access overhead of element-wise operations in SSM. Based onMARCAarchitecture, the time proportion of SSM is still the bottleneck when the sequence length reaches 2048, accounting for 62.52% of the total runtime. Within the SSM, the memory access overhead of element-wise operations account for 97.17%, leading to consuming 96.56% of the time. (2) Inefficient sparse element-wise execution onMARCAarchitecture. SOTA architectures likeMARCApropose a reconfigurable reduction tree to accelerate dense element-wise operations but lack effective sparse support for sparse execution. When 30% of the elementwise operations are skipped, these skipped operations are still mapped to the PE array and trigger redundant execution cycles, incurring 1.43× performance gap with the ideal. (3) Large area overhead for nonlinear function unit. Exponential function and SiLU are two main nonlinear functions in SSM. Previous methods design specific unit for acceleration, leading to 38% and 18% area overheads of the PE. In response to these challenges, we propose a new Mamba accelerator with complementary state space model (SSM) sparsity and reconfigurable architecture,MARCA-v2, based onMARCA, to support fast and energy-efficient Mamba computations. Three novel techniques are as follows: (1) Complementary SSM sparsity with column-wise granularity. We first profile the numerical distributions of activations in SSM and propose a column-wise complementary static sparsity for SSM computation. To further enable lightweight and hardware-friendly sparse computation, we propose a δ-bitmap encoding scheme for compressed storage and introduce two abstractions for sparse element-wise operations. (2) Lightweight sparse element-wise architecture. Based on the hardware friendly sparsity algorithm andMARCAarchitecture, we design and integrate a lightweight Metadata Processing Unit (Meta-PU) into the existing pipeline, which decodes the sparsity metadata and dynamically generates control signals to guide PE arrays. The overall architecture can efficiently support both dense and sparse operations, maximizing speed and energy efficiency. (3) Reusable nonlinear function unit based on reconfigurable PE arrays. We decompose the exponential function and SiLU into several element-wise operations. Thus, the reconfigurable PEs are fully reused to execute nonlinear functions with negligible accuracy loss. We conduct extensive experiments on Mamba model families with different sizes. Experimental results show that in theprefillstage,MARCA-v2achieves 1.77-10.87×, 1.03- 1.08×, and 4.78-9.10× speedup and 8.29-33.47×/1.03-1.08×/4.78-9.10× energy efficiency improvement compared with Mamba-GPU,MARCAand Spada, respectively. In thedecodestage,MARCA-v2achieves 0.88-7.65×/1.00-1.01×/1.19-1.64× speedup and 3.11-27.01×/1.00-1.01×/1.19-1.64× energy efficiency improvement compared with Mamba-GPU,MARCAand Spada, respectively. Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Ningyi Xu, Guohao Dai 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | ViDA: Video Diffusion Transformer Acceleration with Differential Approximation and Adaptive DataflowabstractRecent advancements in Video Diffusion Transformer (VDiT) models have greatly promoted the development of video generation, as exemplified by Sora of OpenAI. However, there are still two challenges for VDiT: 1) There is still existing large inter-frame redundant computation. Previous works on reducing computation based on inter-frame similarity simply consider the Act-W operators. The remaining Act-Act operators still dominate the execution of VDiT (about 57%). 2) Operational intensity varies greatly, leading to under-utilization. There is a massive gap between the operational intensity of Act-W and Act-Act operators in VDiT with multiple frames. Previous works with the static hardware architecture and dataflow lead to under-utilization (<36.42%). Li Ding 0012, Jun Liu 0117, Shan Huang 0010, Guohao Dai 0001 |
ASP-DAC | 2 |
| 2025 | Accelerator for LLM-Enhanced GNN with Product Quantization and Unified IndexingabstractTo alleviate the vulnerability of graph neural networks (GNNs) on unseen graphs, many works propose to integrate large language models (LLMs) into GNNs, called graph foundation models (GFMs). The LLM-enhanced GNN, a typical integration method of GFMs, has achieved state-of-the-art performance in most graph-related tasks. However, intensive general matrix multiplications (GEMMs) overhead of LLMs poses a significant challenge to end-to-end inference latency. The introduction of LLMs brings 100× more workload than original GNNs, with GEMMs accounting for more than 99%, becoming the bottleneck of end-to-end inference. Jinhao Li 0006, Jun Liu 0117, Hao Zhou 0008, Guohao Dai 0001 |
ASP-DAC | 3 |
| 2025 | SG-Filter: Enhancing Similar Text Retrieval via Hierarchical Summarized-Semantic Index and Adaptive FilteringabstractSimilar Text Retrieval (STR) is an essential scenario in the field of information retrieval (IR). Unfortunately, existing mainstream vector-based retrieval methods cannot meet the recall rate requirements in STR scenarios (with a recall rate of less than 72%). This is because existing works have solely focused on the local information of text segments, that is, the text segments themselves ( i.e., semantic information ) and the relationships between them ( i.e., structured information ). Our key insight is that utilizing the global information of text segments ( i.e., summarized information ~. It includes the key expression of the documents to which the text segments belong and the relationship between documents. ) is crucial for improving the recall rate in STR, because the distinction of summarized information helps to filter out confusing vectors during retrieval. However, existing methods using summarized info still have a critical challenge. Their vectorization-based approaches fail to effectively model the global relationship in the summarized information, resulting in a further 79% deterioration in recall rate. Jiancai Ye, Jun Liu 0117, Maojia Sheng, Tao Yang 0042, Jinhao Li 0006, Yu Wang 0002, Guohao Dai 0001 |
CIKM | 2 |
| 2025 | Harnessing Conventional Video Processing Insights for Emerging 3D Video Generation Models: A Comprehensive Attention-aware WayabstractVideo Generation Models based on 3D full attention (3D-VGMs) have significantly enhanced video quality. However, their inference overhead remains substantial, primarily due to the high computational cost of the attention mechanism, which accounts for over 75% of computations. Inspired by the success of conventional video processing, where video compression exploits similarities among patches, we point out that the attention mechanism can also harness the benefits from similarities among tokens. Nonetheless, two critical problems arise: (1) How can similarities be efficiently acquired in real-time? (2) How can workload balance be maintained when similar tokens are randomly distributed? To address these problems and leverage similarities for 3DVGMs, we propose Simpicker, a comprehensive attentionaware algorithm-hardware co-design for 3D-VGMs. Our core methodology is to fully utilize similarities in attention through both coarse-grained and fine-grained approaches while adopting dynamic adaptive strategies to leverage them. From the algorithm perspective, we propose a speculation-based similarity exploitation algorithm, allowing real-time importance speculation on the frame level, which is coarse-grained, and the token level, which is fine-grained. From the micro-architecture perspective, we propose a buffered lookup table-based (LUT-based) multiplication architecture for FP-INT multiplication and further eliminate potential bank conflicts to accelerate unimportant attention computation. From the mapping perspective, SimPicker proposes an adaptive grouping strategy in speculation to tame workload imbalance caused by randomly distributed similar tokens and allow seamless integration of our algorithms. Extensive experiments show that Simpicker achieves an average of $5.21 \times 1.45 \times$ speedup and $17.92 \times 1.63 \times$ energy efficiency compared to the NVIDIA A100 GPU and the state-of-the-art accelerators. Tianlang Zhao, Jun Liu 0117, Xingyang Li, Li Ding 0012, Jinhao Li 0006, Shuaiheng Li, Jinbo Hu, Guohao Dai 0001 |
DAC | 2 |
| 2025 | FlightVGM: Efficient Video Generation Model Inference with Online Sparsification and Hybrid Precision on FPGAsabstractVideo Generation Model (VGM), as a representative of multi-modal large models, has revolutionized the productivity of video content creation. VGMs are compute-bound due to adopting the Diffusion Transformer (i.e., DiT) structure. Sparsification is a common method for accelerating compute-intensive models. Still, sparse VGMs cannot fully exploit the effective throughput (i.e., TOPS) of GPUs. FPGAs are good candidates for accelerating sparse deep learning models. However, existing FPGA accelerators still face low throughput ( < 2TOPS) on VGMs due to the significant gap in peak computing performance (PCP) with GPUs ( > 21× ). To achieve a higher throughput than GPUs, FPGA-based acceleration of sparse VGMs still faces the following challenges: large redundancy in activations, low performance of DSPs under hybrid precision, and under-utilization using static compilation for online compression. Jun Liu 0117, Shulin Zeng, Li Ding 0012, Widyadewi Soedarmadji, Hao Zhou 0008, Jinhao Li 0006, Jintao Li 0002, Yadong Dai, Kairui Wen, Yaqi Sun, Yu Wang 0002, Guohao Dai 0001 |
FPGA | 1 |
| 2025 | TB-STC: Transposable Block-wise N: M Structured Sparse Tensor CoreabstractThe computational and memory demands of Deep Learning (DL) models, from convolutional neural networks to Large Language Models (LLMs), are experiencing a notable surge. The sparsification (e.g., weight pruning and sparse attention) represents a significant approach to reducing latency and energy consumption. However, it is non-trivial to identify a good trade-off between model accuracy and hardware efficiency. Existing work has sought to mitigate the hardware complexity overhead through structured sparsity, yet the resulting accuracy loss remains considerable (e.g., more than 6% accuracy drop with 50% structured sparsity on OPT-6.7B and Llama2-7B).To address the above challenges, this paper proposes Transposable Block-wise Structured Sparsity (TBS). Our key insight is that the weight matrices of the forward and backward pass are transposed to each other during DL training. Exploiting this transposition property facilitates obtaining a structured sparsity pattern that is closer to the unstructured sparsity. In contrast, existing studies explore only one-dimensional structured sparsity. In light of these observations, we propose the transposable block-wise structured sparsity pattern with an efficient end-to-end sparse training method. This method improves accuracy by up to 2.58% over other structured sparsity studies under the same sparsity degree. At the micro-architecture level, we propose TB-STC, a Transposable Block-wise N:M Sparse Tensor Core to efficiently and flexibly facilitate the TBS pattern. TB-STC introduces an adaptive codec architecture for on-the-fly storage format conversion with a higher bandwidth utilization (1.47 ×), and implements an I/O-aware configurable architecture for sparsity-aware scheduling with a better computational utilization (1.57×). Compared with existing work, TB-STC improves the Energy-Delay Product (EDP) by an average of 3.82 × and offers an enhanced accuracy-EDP Pareto frontier across various sparse DL models. Jun Liu 0117, Shulin Zeng, Junbo Zhao 0007, Li Ding 0012, Jinhao Li 0006, Zhenhua Zhu 0002, Xuefei Ning, Chen Zhang 0001, Yu Wang 0002, Guohao Dai 0001 |
HPCA | 1 |
| 2024 | FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAsabstractTransformer-based Large Language Models (LLMs) have made a significant impact on various domains. However, LLMs' efficiency suffers from both heavy computation and memory overheads. Compression techniques like sparsification and quantization are commonly used to mitigate the gap between LLM's computation/memory overheads and hardware capacity. However, existing GPU and transformer-based accelerators cannot efficiently process compressed LLMs, due to the following unresolved challenges: low computational efficiency, underutilized memory bandwidth, and large compilation overheads. This paper proposes FlightLLM, enabling efficient LLMs inference with a complete mapping flow on FPGAs. In FlightLLM, we highlight an innovative solution that the computation and memory overhead of LLMs can be solved by utilizing FPGA-specific resources (e.g., DSP48 and heterogeneous memory hierarchy). We propose a configurable sparse DSP chain to support different sparsity patterns with high computation efficiency. Second, we propose an always-on-chip decode scheme to boost memory bandwidth with mixed-precision support. Finally, to make FlightLLM available for real-world LLMs, we propose a length adaptive compilation method to reduce the compilation overhead. Implemented on the Xilinx Alveo U280 FPGA, FlightLLM achieves 6.0× higher energy efficiency and 1.8× better cost efficiency against commercial GPUs (e.g., NVIDIA V100S) on modern LLMs (e.g., LLaMA2-7B) using vLLM and SmoothQuant under the batch size of one. FlightLLM beats NVIDIA A100 GPU with 1.2× higher throughput using the latest Versal VHK158 FPGA. Shulin Zeng, Jun Liu 0117, Guohao Dai 0001, Tianyu Fu 0004, Wenheng Ma, Hanbo Sun, Zixiao Huang 0001, Yadong Dai, Jintao Li 0002, Kairui Wen, Xuefei Ning, Yu Wang 0002 |
FPGA | 2 |
| 2024 | MARCA: Mamba Accelerator with Reconfigurable ArchitectureabstractState space model (SSM) especially Mamba has demonstrated remarkable capabilities in various domains. Compared to Transformers, Mamba reduces the quadratic computational complexity and achieves a higher algorithm accuracy (e.g., the accuracy of Mamba-2.8b is higher than OPT-6.7b). However, challenges still exist in accelerating Mamba computations. (1) Incompatibility between element-wise operations and Tensor Core. Linear operations (matrix multiplications) and element-wise operations are the two dominating operations in Mamba. The time proportion of element-wise operations escalates significantly (e.g., >60% with 2048 input length). These operations do not need reduction, which is not compatible with the existing Tensor Core-based architectures (e.g., 1/16 normalized speed). (2) Large area overhead for nonlinear function unit. The optimized nonlinear function unit like exponential unit still occupies >30% of the processing element (PE) area. (3) Large memory access but limited data sharing for element-wise operations. Linear and element-wise operations in Mamba exhibit large compute intensity variance (e.g., ~3 orders of magnitude) and large read/write ratio variance (e.g., >3 orders). Due to the limited data sharing in element-wise operations, it is useless to apply the existed methods like tiling to element-wise operations. Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Li Ding 0012, Ningyi Xu, Guohao Dai 0001 |
ICCAD | 4 |
| 2024 | Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous DequantizationabstractLarge language models (LLMs) have demonstrated impressive abilities in various domains while the inference cost is expensive. Many previous studies exploit quantization methods to reduce LLM inference cost by reducing latency and memory consumption. Applying 2-bit single-precision weight quantization brings >3% accuracy loss, so the state-of-the-art methods use mixed-precision methods for LLMs (e.g. Llama2-7b, etc.) to improve the accuracy. However, challenges still exist: (1) Uneven distribution in weight matrix. Weights are quantized by groups, while some groups contain weights with large range. Previous methods apply inter-weight mixed-precision quantization and neglect the range difference inside each weight matrix, resulting in >2.7% accuracy loss (e.g. LLM-MQ and APTQ). (2) Large speed degradation by adding sparse outliers. Reserving sparse outliers improves accuracy but slows down the speed affected by the outlier ratio (e.g. 1.5% outliers resulting in >30% speed degradation in SpQR). (3) Time-consuming dequantization operations on GPUs. Mainstream methods require a dequantization operation to perform computation on the quantized weights, and the 2-order dequantization operation is applied because scales of groups are also quantized. These dequantization operations lead to >50% execution time. Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Yaoxiu Lian, Guohao Dai 0001 |
ICCAD | 5 |
| 2023 | Processing-In-Hierarchical-Memory Architecture for Billion-Scale Approximate Nearest Neighbor SearchabstractGraph-based approximate nearest neighbor search (ANNS) algorithms achieve the best accuracy for fast high-recall searches on billion-scale datasets. Because of the irregular and large-volume data access, existing CPU-based systems suffer from heavy data movements when dealing with graph-based ANNS algorithms. Near-memory-computing (NMC) architectures have demonstrated great potential in boosting the performance of big-data processing. However, existing NMC architectures face two serious problems when processing graph-based ANNS algorithms: (1) the memory capacity of main memory level NMC (e.g., 64GB) cannot meet the storage requirement of ANNS on billion-scale datasets (e.g., 800GB), resulting in heavy data transfers between main memory and storage; (2) the contradiction between the irregular and fine-grained graph access and the page-level read granularity hinder the throughput of storage level NMC.This paper proposes Pyramid, the processing-in-hierarchical-memory architecture for graph-based ANNS on billion-scale datasets. Pyramid combines the internal bandwidth benefits of main memory level NMC with the capacity benefits of storage level NMC. A hierarchical graph-cluster-based ANNS is also proposed for Pyramid. It transforms the irregular data access on large-scale graphs into the irregular access on small-scale graphs at the main memory level and regular sequential in-cluster access at the storage level. Experimental results show that with the same recall of 0.9, Pyramid improves the throughput by 21.1~72.8× and 26.0~50.7× compared with existing CPU/GPU-based ANNS systems on million-scale and billion-scale datasets, respectively. Zhenhua Zhu 0002, Jun Liu 0117, Guohao Dai 0001, Shulin Zeng, Bing Li 0017, Huazhong Yang, Yu Wang 0002 |
DAC | 2 |
| 2023 | TSTC: Two-Level Sparsity Tensor Core Enabling both Algorithm Flexibility and Hardware EfficiencyabstractThe tensor cores in modern GPUs lead to significant performance improvement in matrix multiplication, which is the primary operation in deep learning. However, existing hardware architectures face unstructured sparsity in deep learning, resulting in algorithm inflexibility and hardware inefficiency. The previous tensor core architecture requires matrices to be pruned into 2:4 sparse patterns, leading to algorithm inflexibility. Customized accelerators introduce extra architectures (e.g., interconnection networks for dynamic data routing or buffers for avoiding data conflicts) for unstructured sparse matrices, leading to hardware inefficiency. To tackle the contradiction between algorithm inflexibility and hardware inefficiency, we propose Two-level Sparsity Tensor Core (TSTC) in this paper. TSTC points out that the unstructured sparsity which enables algorithm flexibility can be maintained at the coarse-grained level, while hardware efficiency which requires structured sparsity can be ensured at the fine-grained level. For algorithm flexibility, we propose Flexible Sparse Block (FSB) pattern. FSB enables unstructured sparse matrices can be divided into fine-grained blocks with different structured sparsity. As a result, using FSB leads to up to 7.29x speed up compared with other formats. For hardware efficiency, we propose Dynamic Extendible Reduction Network (DERN). DERN enables different structured sparse reductions by only extending the data width on the standard reduction network without introducing interconnections or buffers. DERN enables TSTC to achieve 7.19x more energy savings under a similar speed. We also propose the whole flow, which can automatically deploy different sparse deep learning algorithms to TSTC. According to extensive experiments, TSTC achieves 1.24 x ~7.69 x speedup and 3.68 x~4.17 x energy savings than the tensor core and the SOTA customized accelerator. Jun Liu 0117, Guohao Dai 0001, Lidong Guo, Xiangsheng Shi, Huazhong Yang, Yu Wang 0002 |
ICCAD | 1 |
| 2023 | DF-GAS: a Distributed FPGA-as-a-Service Architecture towards Billion-Scale Graph-based Approximate Nearest Neighbor SearchabstractEmbedding retrieval is a crucial task for recommendation systems. Graph-based approximate nearest neighbor search (GANNS) is the most commonly used method for retrieval, and achieves the best performance on billion-scale datasets. Unfortunately, the existing CPU- and GPU-based GANNS systems are difficult to optimize the throughput under the latency constraints on billion-scale datasets, due to the underutilized local memory bandwidth (5-45%) and the expensive remote data access overhead (∼ 85% of the total latency). In this paper, we first introduce a practically ideal GANNS architecture for billion-scale datasets, which facilitates a detailed analysis of the challenges and characteristics of distributed GANNS systems. Then, at the architecture level, we propose DF-GAS, a Distributed FPGA-as-a-Service (FPaaS) architecture for accelerating billion-scale Graph-based Approximate nearest neighbor Search. DF-GAS uses a feature-packing memory access engine and a data prefetching and delayed processing scheme to increase local memory bandwidth by 36-42% and reduce remote data access overhead by 76.2%, respectively. At the system level, we exploit the “full-graph + sub-graph” hybrid parallel search scheme on distributed FPaaS system. It achieves million-level query-per-second with sub-millisecond latency on billion-scale GANNS for the first time. Extensive evaluations on million-scale and billion-scale datasets show that DF-GAS achieves an average of 55.4 ×, 32.2 ×, 5.4 ×, and 4.4 × better latency-bounded throughput than CPUs, GPUs, and two state-of-the-art ANNS architectures, i.e., ANNA [23] and Vstore [27], respectively. Shulin Zeng, Zhenhua Zhu 0002, Jun Liu 0117, Guohao Dai 0001, Shuangchen Li, Xuefei Ning, Yuan Xie 0001, Huazhong Yang, Yu Wang 0002 |
MICRO | 3 |
| 2022 | Optimizing Graph-based Approximate Nearest Neighbor Search: Stronger and SmarterabstractApproximate Nearest Neighbor Search (ANNS) is widely used in many fields (e.g., recommender systems). In recent years, the graph-based ANNS methods have attracted the attention of many researchers due to their superiority compared to non-graph-based methods. Compared with traditional recommender systems, mobile recommender systems have higher latency requirements. The graph-based ANNS method faces the following challenges that make it difficult to meet the requirements. (1) Poor connectivity. Due to the limitation of the construction algorithm, the connectivity of the graph is poor, which in turn affects the search performance. (2) Redundant search. The existing search algorithm uses sufficiently long search steps for all queries to achieve high search accuracy. However, the query search steps follow the long-tailed distribution that brings the redundant search, e.g., for more than 40 % of the queries, 87.4 % of the search overhead is redundant. We propose two optimization strategies to tackle the above challenges. (1) Reverse connection enhancement strategy. In the graph construction process, we increase the in-degree of the point to be inserted to enhance the graph connectivity, while keeping the out-degree low to maintain the high search efficiency. (2) Query aware early termination strategy. We identify regional features to predict the number of remaining search steps to achieve dynamic search termination and reduce the redundant search overhead. Finally, we verify the proposed solutions on multiple representative datasets. Compared with the state-of-the-art graph-based algorithm, our solutions can improve the search speed up to 1.21x when the recall rate equals 0.95. Jun Liu 0117, Zhenhua Zhu 0002, Jingbo Hu, Hanbo Sun, Guohao Dai 0001, Huazhong Yang, Yu Wang 0002 |
MDM | 1 |
| 2022 | A Unified FPGA Virtualization Framework for General-Purpose Deep Neural Networks in the CloudabstractINFerence-as-a-Service (INFaaS) has become a primary workload in the cloud. However, existing FPGA-based Deep Neural Network (DNN) accelerators are mainly optimized for the fastest speed of a single task, while the multi-tenancy of INFaaS has not been explored yet. As the demand for INFaaS keeps growing, simply increasing the number of FPGA-based DNN accelerators is not cost-effective, while merely sharing these single-task optimized DNN accelerators in a time-division multiplexing way could lead to poor isolation and high-performance loss for INFaaS. On the other hand, current cloud-based DNN accelerators have excessive compilation overhead, especially when scaling out to multi-FPGA systems for multi-tenant sharing, leading to unacceptable compilation costs for both offline deployment and online reconfiguration. Therefore, it is far from providing efficient and flexible FPGA virtualization for public and private cloud scenarios. Aiming to solve these problems, we propose a unified virtualization framework for general-purpose deep neural networks in the cloud, enabling multi-tenant sharing for both the Convolution Neural Network (CNN), and the Recurrent Neural Network (RNN) accelerators on a single FPGA. The isolation is enabled by introducing a two-level instruction dispatch module and a multi-core based hardware resources pool. Such designs provide isolated and runtime-programmable hardware resources, which further leads to performance isolation for multi-tenant sharing. On the other hand, to overcome the heavy re-compilation overheads, a tiling-based instruction frame package design and a two-stage static-dynamic compilation, are proposed. Only the lightweight runtime information is re-compiled with ∼1 ms overhead, thus guaranteeing the private cloud’s performance. Finally, the extensive experimental results show that the proposed virtualized solutions achieve up to 3.12× and 6.18× higher throughput in the private cloud compared with the static CNN and RNN baseline designs, respectively. Shulin Zeng, Guohao Dai 0001, Hanbo Sun, Jun Liu 0117, Guangjun Ge, Kai Zhong 0007, Kaiyuan Guo, Yu Wang 0002, Huazhong Yang |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2021 | 3M-AI: A Multi-task and Multi-core Virtualization Framework for Multi-FPGA AI Systems in the CloudabstractWith the ever-growing demands for online Artificial Intelligence (AI), the hardware virtualization support for deep learning accelerators is vital for providing AI capability in the cloud. Three basic features, multi-task, dynamic workload, and remote access, are fundamental for hardware virtualization. However, most of the deep learning accelerators do not support concurrent execution of multiple tasks. Besides, the SOTA multi-DNN scheduling algorithm for NN accelerators neither consider the multi-task concurrent execution and resources allocation for the multi-core DNN accelerators. Moreover, existing GPU virtualized solutions could introduce a huge remote access latency overhead, resulting in a severe system performance drop. Shulin Zeng, Guohao Dai 0001, Hanbo Sun, Jun Liu 0117, Hongren Zheng, Yusong Wu, Yi Cai 0003, Yu Wang 0002, Huazhong Yang |
FPGA | 4 |