Ningyi Xu

dblp:88/9033 · also Ning-Yi Xu · DBLP profile ↗
← Back
63ranked-venue papers
3as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MoSs: Mixture of Scales for Efficient High-Resolution Autoregressive Image Generation
abstract
Since next-scale prediction was introduced as a new paradigm for autoregressive image generation, it has attracted extensive research interest. By progressively increasing resolution in a draft-to-refinement process, next-scale prediction demonstrates great potential in both generation quality and efficiency. However, at high resolutions, this paradigm faces a fundamental challenge: token sequences grow quadratically and accumulate across multiple scales, resulting in a key performance bottleneck. Our systematic study uncovers two critical observations: (1) most image regions have stabilized during early drafting stages, making later refinement across the full-scale image token-inefficient; (2) different scales inherently trade off efficiency and fidelity, suggesting that adaptive token dispatch on different scales can focus resources where they yield the greatest quality gains. Motivated by these insights, we propose a training-free Mixture of Scales (MoSs) method for efficient high-resolution autoregressive image generation. MoSs breaks the strict causal dependency across scales in the final refinement steps by parallelizing scales of different resolutions, each responsible for a subset of spatial regions. A lightweight frequency-based token dispatcher analyzes the drafted image and assigns regions to the appropriate scale. The outputs are then composited over the draft to produce the final high-resolution image. The scale-mixture method exhibits remarkable efficiency with little impact on generation quality on various models. For instance, our implementation achieves 2.05-4.96x speedup on the transformer backbone, up to 85.62% KV cache reduction, incurring only 0.1-2.4% loss on GenEval quality, based on the state-of-the-art Infinity model.
Yaoxiu Lian, Hao Liang 0003, Zhihong Gou, Guohao Dai 0001, Ningyi Xu
AAAI7
2026 ROMA: A Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
Guanting Huo, Hao Liang 0003, Shijie Cao, Ningyi Xu
ASP-DAC7
2026 RAMP: RTL-Level Emulation with Thousand-Core-Scale Parallelism
abstract
With the continuous increase in transistor counts on a single chip, the complexity of RTL verification has grown exponentially, and completing a full simulation flow often takes several months. In industrial practice, RTL simulation is typically divided into two stages: functional debugging and system verification. Functional debugging emphasizes fast compilation and is usually performed on multi-core CPUs, while system verification demands extremely high simulation speed and often relies on FPGA acceleration. However, the limited performance of CPU-based simulation has become a major bottleneck that restricts overall design productivity.To address this challenge, we propose RAMP, a scalable multi-core RTL simulation platform that balances fast compilation with high-throughput execution. RAMP leverages a specialized architecture and compilation strategy to accelerate both combinational logic evaluation and sequential logic synchronization. For combinational logic, it adopts a balanced DAG partitioning method together with highly efficient Boolean computation cores; for sequential logic, it employs a low-latency on-chip network (NoC) to achieve efficient state synchronization across cores. Experimental results demonstrate that RAMP achieves up to 12.9× speedup over state-of-the-art multi-core simulators.
Weigang Feng, Peijun Ma, Ningyi Xu
DATE7
2026 MARCA-v2: Mamba Accelerator With Complementary State-Space Model Sparsity and Reconfigurable Architecture
abstract
Large Language models with state space model (SSM) especially Mamba have demonstrated remarkable capabilities in various domains. Compared to Transformers, Mamba reduces the quadratic computational complexity and achieves a higher algorithm performance. Current research on Mamba focuses primarily on integrating it with various application scenarios. However, there is limited research on optimizing for Mamba processing. Therefore, we profile the processing carefully and identify three main challenges in Mamba computations: (1) Large memory access overhead of element-wise operations in SSM. Based onMARCAarchitecture, the time proportion of SSM is still the bottleneck when the sequence length reaches 2048, accounting for 62.52% of the total runtime. Within the SSM, the memory access overhead of element-wise operations account for 97.17%, leading to consuming 96.56% of the time. (2) Inefficient sparse element-wise execution onMARCAarchitecture. SOTA architectures likeMARCApropose a reconfigurable reduction tree to accelerate dense element-wise operations but lack effective sparse support for sparse execution. When 30% of the elementwise operations are skipped, these skipped operations are still mapped to the PE array and trigger redundant execution cycles, incurring 1.43× performance gap with the ideal. (3) Large area overhead for nonlinear function unit. Exponential function and SiLU are two main nonlinear functions in SSM. Previous methods design specific unit for acceleration, leading to 38% and 18% area overheads of the PE. In response to these challenges, we propose a new Mamba accelerator with complementary state space model (SSM) sparsity and reconfigurable architecture,MARCA-v2, based onMARCA, to support fast and energy-efficient Mamba computations. Three novel techniques are as follows: (1) Complementary SSM sparsity with column-wise granularity. We first profile the numerical distributions of activations in SSM and propose a column-wise complementary static sparsity for SSM computation. To further enable lightweight and hardware-friendly sparse computation, we propose a δ-bitmap encoding scheme for compressed storage and introduce two abstractions for sparse element-wise operations. (2) Lightweight sparse element-wise architecture. Based on the hardware friendly sparsity algorithm andMARCAarchitecture, we design and integrate a lightweight Metadata Processing Unit (Meta-PU) into the existing pipeline, which decodes the sparsity metadata and dynamically generates control signals to guide PE arrays. The overall architecture can efficiently support both dense and sparse operations, maximizing speed and energy efficiency. (3) Reusable nonlinear function unit based on reconfigurable PE arrays. We decompose the exponential function and SiLU into several element-wise operations. Thus, the reconfigurable PEs are fully reused to execute nonlinear functions with negligible accuracy loss. We conduct extensive experiments on Mamba model families with different sizes. Experimental results show that in theprefillstage,MARCA-v2achieves 1.77-10.87×, 1.03- 1.08×, and 4.78-9.10× speedup and 8.29-33.47×/1.03-1.08×/4.78-9.10× energy efficiency improvement compared with Mamba-GPU,MARCAand Spada, respectively. In thedecodestage,MARCA-v2achieves 0.88-7.65×/1.00-1.01×/1.19-1.64× speedup and 3.11-27.01×/1.00-1.01×/1.19-1.64× energy efficiency improvement compared with Mamba-GPU,MARCAand Spada, respectively.
Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Ningyi Xu, Guohao Dai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 LLSM: LLM-enhanced Logic Synthesis Model with EDA-guided CoT Prompting, Hybrid Embedding and AIG-tailored Acceleration
abstract
Machine learning-based methods have shown promising results in the field of Electronic Design Automation (EDA) like logic synthesis result prediction, enabling a shift-left in the overall EDA flow. Designers should fully optimize their Register Transfer Level (RTL) designs early because remedying low-quality RTL in downstream synthesis stages is extremely challenging. However, previous works mainly start modeling from the netlist level or layout level and apply Graph Neural Networks (GNNs) to make predictions.
Shan Huang 0010, Jinhao Li 0006, Jiancai Ye, Ningyi Xu, Guohao Dai 0001
ASP-DAC6
2025 A Cross-model Fusion-aware Framework for Optimizing (gather-matmul-scatter)s Workload
abstract
Modern deep learning models, such as Relation Graph Convolutional Network (RGCN), Sparse Convolutional Networks (SpConv), and Mixture of Experts Networks (MoE), are significantly dependent on the (gather-matmul-scatter) (abbreviated as (g-mm-s) ${ }_{\mathrm{s}}$) workload as their fundamental computational pattern. While existing works have made optimization attempts, several critical challenges remain unsolved, including domain-specific optimization migration, time-consuming exploration, and inefficient dataflow with dynamic inputs.To address these challenges, we introduce Efficient-GMS, a comprehensive framework that enhances ($\mathrm{g}-\mathrm{mm}-\mathrm{s})_{\text {s }}$ workload across diverse input scenarios. Our framework introduces (1) A Fusion-aware framework enabling cross-model optimization migration. We propose a comprehensive dataflow analysis that identifies shared computational patterns across models, enabling the development of four optimized dataflow patterns with vertical and horizontal fusion strategies. (2) Performance model-guided configuration space reduction. We develop a performance model to predict the relative execution efficiency across configurations, thereby reducing the search space and minimizing search time while ensuring optimal configuration selection. (3) Adaptive dataflow selection mechanism. We implement a lightweight heuristic model that dynamically selects optimal dataflow patterns based on the characteristics of the input and the hardware. Experimental results demonstrate that Efficient-GMS achieves significant performance gains, delivering an average end-to-end speedup of $1.46 \times$ in RGCN model, $1.32 \times$ in Sp-Conv-based model, and $1.15 \times$ in MoE model compared to state-of-the-art methods.
Yaoxiu Lian, Zhihong Gou, Yibo Han, Zhongming Yu, Sheng Yuan, Zhilin Pei, Xingcheng Zhang, Ningyi Xu, Guohao Dai 0001
DAC9
2025 Speculative Decoding for Verilog: Speed and Quality, All in One
abstract
The rapid advancement of large language models (LLMs) has revolutionized code generation tasks across various programming languages. However, the unique characteristics of programming languages, particularly those like Verilog with specific syntax and lower representation in training datasets, pose significant challenges for conventional tokenization and decoding approaches. In this paper, we introduce a novel application of speculative decoding for Verilog code generation, showing that it can improve both inference speed and output quality, effectively achieving speed and quality all in one. Unlike standard LLM tokenization schemes, which often fragment meaningful code structures, our approach aligns decoding stops with syntactically significant tokens, making it easier for models to learn the token distribution. This refinement addresses inherent tokenization issues and enhances the model’s ability to capture Verilog’s logical constructs more effectively. Our experimental results show that our method achieves up to a $5.05 \times$ speedup in Verilog code generation and increases pass@10 functional accuracy on RTLLM by up to $\mathbf{1 7. 1 9 \%}$ compared to conventional training strategies. These findings highlight speculative decoding as a promising approach to bridge the quality gap in code generation for specialized programming languages.
Changran Xu, Yi Liu 0081, Yunhao Zhou, Shan Huang 0010, Ningyi Xu, Qiang Xu 0001
DAC5
2025 DyLGNN: Efficient LM-GNN Fine-Tuning with Dynamic Node Partitioning, Low-Degree Sparsity, and Asynchronous Sub-Batch
abstract
Text-Attributed Graphs (TAGs) tasks involve both textual node information and graph topological structure. The top-k method, using Language Models (LMs) for text encoding and Graph Neural Networks (GNNs) for graph processing, offers the best accuracy while balancing memory and training time. However, challenges still exist: (1) Static sampling of k neighbors reduces performance. Using a fixed k can result in sampling too few or too many nodes, leading to a 3.2% accuracy loss across datasets. (2) Time-consuming processing for non-trainable nodes. After partitioning all nodes into with-gradient trainable and without-gradient non-trainable sets, the number of non-trainable nodes is ~9-10 x larger than trainable nodes, resulting in nearly 70% of the total time. (3) Time-consuming data movement. For processing non-trainable nodes, after the text strings are tokenized into tokens on the CPU side, the data movement from host memory to GPU takes 30%-40% of the time. In this paper, we propose DyLGNN, an efficient end-to-end LM-GNN fine-tuning framework through three innovations: (1) Heuristic Node Partitioning. We propose an algorithm that dynamically and adaptively selects “important” nodes to participate in the training process for downstream tasks. Compared to the static top-k method, we reduce the training memory usage by 24.0%. (2) Low-Degree Sparse Attention. We point out that the embedding of low-degree nodes has minimal impact on the final results (e.g. ~1.5% accuracy loss), therefore, We perform sparse attention computation on low-degree nodes to further reduce the computation caused by “unimportant” nodes, achieving an average 1.27 x speedup. (3) Asynchronous Sub-batch Pipeline. Within the top-k framework, we analyze the time breakdown of the LM inference component. Leveraging our heuristic node partitioning, which effectively minimizes memory demands, we can asynchronously execute data movement and computation, thereby overlapping the time required for data movement. This improves GPU utilization and results in an average 1.1x speedup. We conduct experiments on several common graph datasets, and by combining the three methods mentioned above, DyLGNN achieves a 22.0% reduction in memory usage and a 1.3x end-to-end speedup compared to the top-k strategy.
Jinhao Li 0006, Shan Huang 0010, Jiancai Ye, Ningyi Xu, Guohao Dai 0001
DATE6
2025 LoRA Decompose: Serving Fine-Tuned Models into LoRA-Like
abstract
Large language models (LLMs) achieve remarkable performance across diverse tasks but face increasing GPU-memory demands due to the growing variety and complexity of downstream tasks. Efficient inference has thus become essential, especially for resource-limited settings. In this paper, we propose LoRA Decompose, a novel compression approach based on a key insight: instruction-fine-tuned models share a common pretrained-like base component and differ primarily through low-rank, LoRA-like delta components. Leveraging this observation, we reformulate the inference problem as a constrained optimization task that jointly identifies a shared low-rank structure across multiple models, significantly reducing their memory footprints. We solve this optimization efficiently using a custom-designed block coordinate descent algorithm, converging quickly within a few iterations. Empirical experiments with Llama-2 7B and 13B models demonstrate that our method achieves a remarkable >32x GPU memory reduction while preserving task accuracy, allowing substantial efficiency gains for practical deployment.
Yibo Han, Tangzhi Xu, Zenan Li, Youshan Miao, Yuan Yao 0001, Ningyi Xu
ECAI7
2025 DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
abstract
Circuit representation learning has become pivotal in electronic design automation, enabling critical tasks such as testability analysis, logic reasoning, power estimation, and SAT solving. However, existing models face significant challenges in scaling to large circuits due to limitations like over-squashing in graph neural networks and the quadratic complexity of transformer-based models. To address these issues, we introduce \textbf{DeepGate4}, a scalable and efficient graph transformer specifically designed for large-scale circuits. DeepGate4 incorporates several key innovations: (1) an update strategy tailored for circuit graphs, which reduce memory complexity to sub-linear and is adaptable to any graph transformer; (2) a GAT-based sparse transformer with global and local structural encodings for AIGs; and (3) an inference acceleration CUDA kernel that fully exploit the unique sparsity patterns of AIGs. Our extensive experiments on the ITC99 and EPFL benchmarks show that DeepGate4 significantly surpasses state-of-the-art methods, achieving 15.5\% and 31.1\% performance improvements over the next-best models. Furthermore, the Fused-DeepGate4 variant reduces runtime by 35.1\% and memory usage by 46.8\%, making it highly efficient for large-scale circuit analysis. These results demonstrate the potential of DeepGate4 to handle complex EDA tasks while offering superior scalability and efficiency.
Shan Huang 0010, Jianyuan Zhong, Zhengyuan Shi, Guohao Dai 0001, Ningyi Xu, Qiang Xu 0001
ICLR6
2025 OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation Balance
abstract
Vision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different devices. The vision and language parts are inherently heterogeneous: their data distribution and model architecture differ significantly, which affects distributed training efficiency. To address this issue, we rebalance the computational load from data, model, and memory perspectives, achieving more balanced computation across devices. Specifically, for the data, instances are grouped into new balanced mini-batches within and across devices. A search-based method is employed for the model to achieve a more balanced partitioning. For memory optimization, we adaptively adjust the re-computation strategy for each partition to utilize the available memory fully. These three perspectives are not independent but are closely connected, forming an omniverse balanced training framework. Extensive experiments are conducted to validate the effectiveness of our method. Compared with the open-source training code of InternVL-Chat, training time is reduced greatly, achieving about 1.8$\times$ speed-up. Our method’s efficacy and generalizability are further validated across various models and datasets. Codes will be released at https://github.com/ModelTC/OmniBal.
Yongqiang Yao, Jingru Tan, Feizhao Zhang, Yazhe Niu, Xin Jin 0008, Bo Li 0126, Pengfei Liu 0003, Ruihao Gong, Dahua Lin, Ningyi Xu
ICML11
2025 SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
abstract
The application of vision-language models (VLMs) has achieved impressive success in various robotics tasks. However, there are few explorations for foundation models used in quadruped robot navigation through terrains in 3D environments. We introduce SARO (Space-Aware Robot System for Terrain Crossing), an innovative system composed of a high-level reasoning module, a closed-loop sub-task execution module, and a low-level control policy. It enables the robot to navigate across 3D terrains and reach the goal position. For high-level reasoning and execution, we propose a novel algorithmic system taking advantage of a VLM, with a design of task decomposition and a closed-loop sub-task execution mechanism. For low-level locomotion control, we utilize the Probability Annealing Selection (PAS) method to effectively train a control policy by reinforcement learning. Numerous experiments show that our whole system can accurately and robustly navigate across several 3D terrains, and its generalization ability ensures the applications in diverse indoor and outdoor scenarios and terrains. Appendix and Videos can be found in project page: https://saro-vlm.github.io/.
Shaoting Zhu, Derun Li, Linzhan Mou, Ningyi Xu, Hang Zhao 0021
ICRA5
2025 Hierachical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
abstract
Training Long-Context Large Language Models (LLMs) is challenging, as hybrid training with long-context and short-context data often leads to workload imbalances. Existing works mainly use data packing to alleviate this issue, but fail to consider imbalanced attention computation and wasted communication overhead. This paper proposes Hierarchical Balance Packing (HBP), which designs a novel batch-construction method and training recipe to address those inefficiencies. In particular, the HBP constructs multi-level data packing groups, each optimized with a distinct packing length. It assigns training samples to their optimal groups and configures each group with the most effective settings, including sequential parallelism degree and gradient checkpointing configuration. To effectively utilize multi-level groups of data, we design a dynamic training pipeline specifically tailored to HBP, including curriculum learning, adaptive sequential parallelism, and stable loss. Our extensive experiments demonstrate that our method significantly reduces training time over multiple datasets and open-source models while maintaining strong performance. For the largest DeepSeek-V2 (236B) MoE model, our method speeds up the training by 2.4$\times$ with competitive performance. Codes will be released at https://github.com/ModelTC/HBP.
Yongqiang Yao, Jingru Tan, Kaihuan Liang, Feizhao Zhang, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu
NeurIPS10
2025 A Point Transformer Accelerator With Distribution-Aware Heuristic Distance Calculation
abstract
Point clouds are an important form of 3-D data used in applications, such as computer vision and autonomous driving, but the irregular and disordered nature of point clouds makes processing them severely challenging. Recently, point-based neural networks for point clouds have been widely used in various 3-D applications. Notably, transformer-based models have demonstrated state-of-the-art accuracy. However, three significant challenges exist: 1) data interdependence hinders parallel execution in networks like Point Transformer; 2) the farthest point sampling (FPS) involves redundant memory access and computational overhead; and 3) intermediate results require repetitive memory access and calculations between FPS and K-nearest neighbor (kNN) operators. This limits Point Transformer’s processing speed to 17.80 frames/s on NVIDIA Jetson Orin, below the real-time requirement of around 30 frames/s. In this article, we introduce PTrAcc++, an innovative point transformer accelerator to address the aforementioned three challenges from the following three levels. On the computation graph level, our investigation reveals that the Point Transformer’s performance suffers minimal degradation when operating within a constrained receptive field. Leveraging this insight, PTrAcc++ strategically frees the MaxPool and attention-kNN layers, along with their associated data dependencies, achieving an inconsequential loss in accuracy. On the operator level, we identify that the variability for distance computation among accessed points during FPS iterations contributes to redundant memory accesses and computational overhead. PTrAcc++ proposes a distribution-aware heuristic for distance calculation to minimize unnecessary memory accesses and computational redundancies within the FPS operator. On the architecture level, we recognize that the transition down process (encompassing FPS and kNN operations) constitutes 71.77% of the total inference time, PTrAcc++ proposes an integrated FPS-kNN architecture to select error-driven k neighbors, reducing repeated memory accesses and distance recalculations of intermediate results. Through extensive experimentation, PTrAcc++ demonstrates remarkable performance improvements, achieving end-to-end speedups of up to$2.96\times $,$1.70\times $, and$1.19\times $when compared to the state-of-the-art acceleratorsPointAcc (Lin et al., 2021), MARS (Yang et al., 2023), and PTrAcc (Lian et al., 2023), respectively, across a variety of point cloud neural networks.
Yaoxiu Lian, Ke Hong, Yu Wang 0002, Ningyi Xu, Guohao Dai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 IPDR: An Inter-Chiplet Priority-Driven Deadlock Resolution for 2-D/2.5-D Multichiplet Systems
Yaoyao Ye, Jianfei Jiang 0001, Weiguang Sheng, Ningyi Xu, Yong Lian 0001, Guanghui He 0002
IEEE Trans. Very Large Scale Integr. Syst.8
2024 BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
abstract
The upscaling of Large Language Models (LLMs) has yielded impressive advances in natural language processing, yet it also poses significant deployment challenges.Weight quantization has emerged as a widely embraced solution to reduce memory and computational demands.This paper introduces BitDistiller, a framework that synergizes Quantization-Aware Training (QAT) with Knowledge Distillation (KD) to boost the performance of LLMs at ultra-low precisions (sub-4-bit).Specifically, BitDistiller first incorporates a tailored asymmetric quantization and clipping technique to maximally preserve the fidelity of quantized weights, and then proposes a novel Confidence-Aware Kullback-Leibler Divergence (CAKLD) objective, which is employed in a self-distillation manner to enable faster convergence and superior model performance.Empirical evaluations demonstrate that BitDistiller significantly surpasses existing methods in both 3-bit and 2-bit configurations on general language understanding and complex reasoning benchmarks.Notably, Bit-Distiller is shown to be more cost-effective, demanding fewer data and training resources.
Dayou Du, Shijie Cao, Ting Cao 0003, Xiaowen Chu 0001, Ningyi Xu
ACL (1)7
2024 Leveraging Enhanced Queries of Point Sets for Vectorized Map Construction
Zihao Liu 0018, Xiaoyu Zhang 0017, Guangwei Liu, Ji Zhao 0001, Ningyi Xu
ECCV (57)5
2024 Enhancing Vectorized Map Perception with Historical Rasterized Maps
Xiaoyu Zhang 0017, Guangwei Liu, Zihao Liu 0018, Ningyi Xu, Yun-Hui Liu 0001, Ji Zhao 0001
ECCV (17)4
2024 MARCA: Mamba Accelerator with Reconfigurable Architecture
abstract
State space model (SSM) especially Mamba has demonstrated remarkable capabilities in various domains. Compared to Transformers, Mamba reduces the quadratic computational complexity and achieves a higher algorithm accuracy (e.g., the accuracy of Mamba-2.8b is higher than OPT-6.7b). However, challenges still exist in accelerating Mamba computations. (1) Incompatibility between element-wise operations and Tensor Core. Linear operations (matrix multiplications) and element-wise operations are the two dominating operations in Mamba. The time proportion of element-wise operations escalates significantly (e.g., >60% with 2048 input length). These operations do not need reduction, which is not compatible with the existing Tensor Core-based architectures (e.g., 1/16 normalized speed). (2) Large area overhead for nonlinear function unit. The optimized nonlinear function unit like exponential unit still occupies >30% of the processing element (PE) area. (3) Large memory access but limited data sharing for element-wise operations. Linear and element-wise operations in Mamba exhibit large compute intensity variance (e.g., ~3 orders of magnitude) and large read/write ratio variance (e.g., >3 orders). Due to the limited data sharing in element-wise operations, it is useless to apply the existed methods like tiling to element-wise operations.
Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Li Ding 0012, Ningyi Xu, Guohao Dai 0001
ICCAD6
2024 VEGA: Implementing a Versatile and Efficient Deep Learning Processor with Graph-Based ALU
abstract
As neural networks advance, the diversity and latency proportion of non-matrix-multiplication operators (NMO) are on the rise. Providing a versatile and efficient acceleration for this intricate set of NMOs poses great challenges in hardware design. In this work, we analyze the algorithmic structure of NMOs and propose graph-based ALU (GALU) to improve efficiency. The key idea is to organize functional units into a dataflow graph with a configurable interconnection, which reduces computation time and memory access. Further, we provide architectural support to integrate GALU into a multi-thread processor called VEGA, which supports various NMO structures. At the hardware level, we devise swift interconnection reconfiguration (SIR) for GALU to reduce the latency caused by reconfiguration. We also design a fine-grained instruction scheduler to fully utilize SIR. At the software level, a three-stage compilation framework is developed to enhance the usability. Experiments demonstrate that GALU achieves a 2.27x speedup with only an 18.11 % increase in area overhead. Compared with NVIDIA Jetson Orin, the VEGA prototype achieves a 3.84x speedup on typical NMOs and achieves a 2.03 x end-to-end speedup at the network level.
Guanting Huo, Guanghui He 0002, Ningyi Xu
ICCD5
2024 Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
abstract
Efficient deployment of Large Language Models (LLMs) requires low-bit quantization to reduce model size and inference cost. Besides low-bit integer formats (e.g., INT8/INT4) used in previous quantization works, emerging low-bit floating-point formats (e.g., FP8/FP4) supported by advanced hardware like NVIDIA’s H100 GPU offer an alternative. Our study finds that introducing floating-point formats significantly improves LLMs quantization. We also discover that the optimal quantization format varies across layers. Therefore, we select the optimal format for each layer, which we call the Mixture of Formats Quantization (MoFQ) method. Our MoFQ method achieves better or comparable results over current methods in weight-only (W-only) and weight-activation (WA) post-training quantization scenarios across various tasks, with no additional hardware overhead.
Lingran Zhao, Shijie Cao, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004, Shanghang Zhang, Ningyi Xu
ICME10
2024 Hardware-oriented algorithms for softmax and layer normalization of large language models
Wenjie Li 0003, Dongxu Lyu, Gang Wang 0063, Aokun Hu, Ningyi Xu, Guanghui He 0002
Sci. China Inf. Sci.5
2024 CoDA: A Co-Design Framework for Versatile and Efficient Attention Accelerators
abstract
As a primary component of Transformers, attention mechanism suffers from quadratic computational complexity. To achieve efficient implementations, its hardware accelerator designs have aroused great research interest. However, most existing accelerators only support a single type of application and a single type of attention, making it difficult to meet the demands of diverse application scenarios. Additionally, they mainly focus on the dynamic pruning of attention matrices, which requires the deployment of pre-processing units, thereby reducing overall hardware efficiency. This paper presents CoDA which is an algorithm, dataflow and architecture co-design framework for versatile and efficient attention accelerators. The designed accelerator supports both NLP and CV applications, and can be configured into the mode supporting low-rank attention or low-rank plus sparse attention. We apply algorithmic transformations to low-rank attention to significantly reduce computational complexity. To prevent an increase in storage overhead resulting from the proposed algorithmic transformations, we carefully design the dataflows and adopt a block-wise fashion. Down-scaling softmax is further supported by architecture and dataflow co-design. Moreover, we propose a softmax sharing strategy to reduce the area cost. Our experiment results demonstrate that the proposed accelerator outperforms the state-of-the-art designs in terms of throughput, area efficiency and energy efficiency.
Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002
IEEE Trans. Computers3
2024 A Precision-Scalable Deep Neural Network Accelerator With Activation Sparsity Exploitation
abstract
To meet the demand in a wide range of practical applications, precision-scalable deep neural network (DNN) accelerators are becoming an unavoidable trend. On the other hand, it has been demonstrated that a DNN accelerator may achieve better computation efficiency through exploiting the sparsity. Therefore, DNN accelerators with both precision scalability and sparsity exploitation are expected to have better performance. In this article, we propose an efficient precision-scalable DNN accelerator that can exploit the sparsity of activations. The precision scalability is obtained from the decomposable multiplier which is inspired by the well-known design, Bit Fusion. Besides, a zero-skipping scheme is adopted to leverage the inherent sparsity of activations. We first modify the architecture of the conventional fusion unit (FU) to make it amenable to the zero-skipping scheme. Then, a segmentation approach is devised to tackle the memory access conflict. Furthermore, a sparsity-aware mapping method is proposed to balance the workload of processing elements (PEs). Moreover, we present a bit-splitting strategy which can take advantage of the sparsity in the bit level. Compared with the state-of-the-art precision-scalable designs, our proposed accelerator can provide speedups of$4.12\times $,$4.07\times $, and$6.62\times $in the precision modes$8b\times 8b$,$4b\times 4b$, and$2b\times 2b$, respectively. Meanwhile, it also achieves$3.92\times $peak area efficiency and competitive peak energy efficiency.
Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 INDM: Chiplet-Based Interconnect Network and Dataflow Mapping for DNN Accelerators
abstract
Chiplet-based deep neural network (DNN) accelerator is a promising solution to balance the performance and manufacturing cost. However, different from monolithic chips, interconnect network design and architectural partitioning for multiple chiplets would result in a huge design space and make it difficult to keep scalability and high hardware utilization. Moreover, how to efficiently map DNN workloads onto multiple DRAM dies and compute dies is another major challenge. To alleviate the above issues, in this work, we propose INDM, a chiplet-based interconnect network and dataflow mapping co-optimization for DNN accelerators. First, we propose an efficient hierarchical interconnect network composed of a multiring on-die network and a cluster-based interdie network, to facilitate the data reuse and traffic pattern in DNN workloads. Second, architectural partitioning and topology exploration for chiplet-based DNN accelerators are proposed to find the optimal architecture configurations. Third, an interdie communication-aware dataflow mapping is proposed to minimize traffic congestion during DNN layer switching. We implement the proposed chiplet-based interconnect network design and dataflow mapping algorithm for a set of popular DNN models, including VGG-16, ResNet-18, DarkNet-19, ResNet-50, and ResNet-101. Experimental results show that as compared with the state-of-the-art related work, such as NN-Baton and SIMBA, our work achieves 26.00%–73.81% energy-delay-product (EDP) reduction and 26.93%–79.78% latency reduction.
Xi Fan, Yaoyao Ye, Xuyan Wang, Guojie Xiong, Xianglun Leng, Ningyi Xu, Yong Lian 0001, Guanghui He 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Models
abstract
Large language models (LLMs) have sparked a new revolution in the field of natural language processing (NLP), and have garnered tremendous attention in both academic research and everyday life, thanks to their unprecedented performance in a wide range of applications. However, their deployment remains a significant challenge, primarily due to their intensive computational and memory requirements. Hardware acceleration and efficient quantization are promising solutions to address the two issues. In this paper, a quantization and hardware architecture co-design is presented for matrix-vector multiplications (MVMs) of LLMs. During quantization, we uniformly group weights and activations to ensure workload balance for hardware. To enhance the performance of quantization, we further propose two approaches called channel sorting and channel selection, which can be applied simultaneously. To support the proposed quantization scheme, we develop two precision-scalable MVM hardware architectures. They are specifically designed for high speed and high energy efficiency, respectively. Experimental results show that our proposed quantization scheme achieves state-of-the-art performance among all the reported post-training schemes that quantize both weights and activations into integers. Compared to MVM architecture of the state-of-the-art LLM accelerator OliVe, our design exhibits significant advantages in terms of area efficiency and energy efficiency.
Wenjie Li 0003, Aokun Hu, Ningyi Xu, Guanghui He 0002
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 M2M: A Fine-Grained Mapping Framework to Accelerate Multiple DNNs on a Multi-Chiplet Architecture
abstract
With the advancement of artificial intelligence, the collaboration of multiple deep neural networks (DNNs) has been crucial to existing embedded systems and cloud systems, especially for automatic driving applications as well as augmented and virtual reality (AR/VR) applications. To trade off between cost and performance, chiplet-based DNN accelerators have emerged as a promising solution for accelerating DNN workloads. However, most existing mapping methods for multiple DNNs target for the monolithic chip, which fail to solve the problems faced by the emerging multi-chiplet architecture, such as the problems of distributed memory access, complex heterogeneous interconnect network, and the scaling-up of computing resources. In this work, we propose M2M, a fine-grained mapping framework for accelerating multiple DNNs on a multi-chiplet architecture. It includes a temporal and spatial task scheduling for reconfigurable dataflow accelerators and a communication-aware task mapping in a heterogeneous interconnect network. To enhance communication efficiency and reduce the overall latency, we further propose a fine-tuned quality-of-service (QoS) policy for network-on-package (NoP) links. To the best of our knowledge, this is the first fine-grained mapping framework for multiple DNNs on a multi-chiplet architecture. We implemented the proposed fine-grained mapping framework using genetic algorithm and simulated annealing algorithm. Experimental results show that our work achieves 7.18%–61.09% latency reduction under vision, language, and mixed workloads when compared with the state-of-the-art related work.
Xuyan Wang, Yaoyao Ye, Dongxu Lyu, Guojie Xiong, Ningyi Xu, Yong Lian 0001, Guanghui He 0002
IEEE Trans. Very Large Scale Integr. Syst.6
2023 FLNA: An Energy-Efficient Point Cloud Feature Learning Accelerator with Dataflow Decoupling
abstract
Grid-based feature learning network plays a key role in recent point-cloud based 3D perception. However, high point sparsity and special operators lead to large memory footprint and long processing latency, posing great challenges to hardware acceleration. We propose FLNA, a novel feature learning accelerator with algorithm-architecture co-design. At algorithm level, the dataflow-decoupled graph is adopted to reduce 86% computation by exploiting inherent sparsity and concat redundancy. At hardware design level, we customize a pipelined architecture with block-wise processing, and introduce transposed SRAM strategy to save 82.1% access power. Implemented on a 40nm technology, FLNA achieves 13.4 − 43.3× speedup over RTX 2080Ti GPU. It rivals the state-of-the-art accelerator by 1.21× energy-efficiency improvement with 50.8% latency reduction.
Dongxu Lyu, Ningyi Xu, Guanghui He 0002
DAC4
2023 COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion Methodologies
abstract
Attention mechanism acceleration is becoming increasingly vital to achieve superior performance in deep learning tasks. Existing accelerators are commonly devised dedicatedly by exploring the potential sparsity in neural network (NN) models, which suffer from complicated training, tuning processes, and accuracy degradation. By systematically analyzing the inherent dataflow characteristics of attention mechanism, we propose the Co-Operative Systolic Array (COSA) to pursue higher computational efficiency for its acceleration. In COSA, two systolic arrays that can be dynamically configured into weight or output stationary modes are cascaded to enable efficient attention operation. Thus, hybrid dataflows are simultaneously supported in COSA. Furthermore, various fusion methodologies and an advanced softmax unit are designed. Experimental results show that the COSA-based accelerator can achieve 2.95-28.82× speedup compared with the existing designs, with up to 97.4% PE utilization rate and less memory access.
Zhican Wang, Gang Wang 0063, Honglan Jiang, Ningyi Xu, Guanghui He 0002
DAC4
2023 Adam Accumulation to Reduce Memory Footprints of Both Activations and Gradients for Large-Scale DNN Training
abstract
Running out of GPU memory has become a main bottleneck for large-scale DNN training. How to reduce the memory footprint during training has received intensive research attention. We find that previous gradient accumulation reduces activation memory but fails to be compatible with gradient memory reduction due to a contradiction between preserving gradients and releasing gradients. To address this issue, we propose a novel optimizer accumulation method for Adam, named Adam Accumulation (AdamA), which enables reducing both activation and gradient memory. Specifically, AdamA directly integrates gradients into optimizer states and accumulates optimizer states over micro-batches, so that gradients can be released immediately after use. We mathematically and experimentally demonstrate AdamA yields the same convergence properties as Adam. Evaluated on transformer-based models, AdamA achieves up to 23% memory reduction compared to gradient accumulation with less than 2% degradation in training throughput. Notably, AdamA can work together with memory reduction methods for optimizer states to fit 1.26×~3.14× larger models over PyTorch and DeepSpeed baseline on GPUs with different memory capacities.
Yibo Han, Shijie Cao, Guohao Dai 0001, Youshan Miao, Ting Cao 0003, Fan Yang 0024, Ningyi Xu
ECAI8
2023 A Point Transformer Accelerator with Fine-Grained Pipelines and Distribution-Aware Dynamic FPS
abstract
Recently, point-based point cloud neural networks have been applied to various 3D point cloud scenarios. Among them, transformer-based point cloud neural networks achieve state-of-the-art accuracy. However, there still exist three challenges that: (1) the data dependency between the transition down and feature extraction process hinders parallel execution in networks like Point Transformer; (2) farthest point sampling (FPS) operator has redundant memory access and computational overhead during the transition down process and (3) the intermediate results require repeated memory access and calculation between the FPS and kNN operators in the transition down process. As a result, typical networks like Point Transformer process on average 17.80 frames per second on NVIDIA Jetson Orin, which cannot meet the requirements of real-time perception (~30 frames per second). In this paper, we propose PTrAcc, a Point Transformer Accelerator with fine-grained pipelines and distribution-aware dynamic FPS. Computation graph level: Since we find that there is little accuracy loss with a narrowed receptive field in Point Transformer, PTrAcc removes the MaxPool and attention-kNN layers and their attached data dependencies with negligible accuracy loss to enable fine-grained pipelines. Consequently, the inference is accelerated by 1.05×. Operator level: Since the distribution of accessed points varies in different FPS iterations, PTrAcc introduces distribution-aware dynamic FPS to reduce redundant memory access and computation overhead based on the distribution. As a result, the speed of the FPS operations is increased by 1.35×. Architecture level: Since the transition down process (FPS, kNN) accounts for 71.77% of the total inference time, PTrAcc proposes a fused FPS-kNN architecture to reduce repeated memory access and distance calculation of intermediate results, and the process is accelerated by up to 2.15×. Extensive experimental results show that, PTrAcc achieves up to 1.63× and 2.38× end-to-end speedup over state-of-the-art accelerators, MARS [1] and PointAcc [2], on various point cloud neural networks, respectively.
Yaoxiu Lian, Ke Hong, Yu Wang 0002, Guohao Dai 0001, Ningyi Xu
ICCAD6
2023 SpOctA: A 3D Sparse Convolution Accelerator with Octree-Encoding-Based Map Search and Inherent Sparsity-Aware Processing
abstract
Point-cloud-based 3D perception has attracted great attention in various applications including robotics, autonomous driving and AR/VR. In particular, the 3D sparse convolution (SpConv) network has emerged as one of the most popular backbones due to its excellent performance. However, it poses severe challenges to real-time perception on general-purpose platforms, such as lengthy map search latency, high computation cost, and enormous memory footprint. In this paper, we propose SpOctA, a SpConv accelerator that enables high-speed and energy-efficient point cloud processing. SpOctA parallelizes the map search by utilizing algorithm-architecture co-optimization based on octree encoding, thereby achieving 8.8-21.2× search speedup. It also attenuates the heavy computational workload by exploiting inherent sparsity of each voxel, which eliminates computation redundancy and saves 44.4-79.1% processing latency. To optimize on-chip memory management, a SpConv-oriented non-uniform caching strategy is introduced to reduce external memory access energy by 57.6% on average. Implemented on a 40nm technology and extensively evaluated on representative benchmarks, SpOctA rivals the state-of-the-art SpConv accelerators by 1.1-6.9× speedup with 1.5-3.1× energy efficiency improvement,
Dongxu Lyu, Ningyi Xu, Guanghui He 0002
ICCAD5
2023 History-Detr: Optimize Query Initialization Strategy by Using Historical Information and Kinematics
abstract
Recent 3D object detectors leverage multi-frame data, including past and future data, to enhance performance. However, the method of temporal data fusion they employ has not fully tapped into its potential for improving performance. Existing works make use of multi-frame data which only fuse specific features according to ego-motion and cannot be directly applied to long sequences due to the huge computation and memory cost. We find that the present methods do not efficiently exploit history information including history predictions and object-motion. Building on our investigations, we present a novel hybrid query formulation comprised of the history queries and original queries. The history queries consist of inferred position and content queries obtained from the historical predictions and features, which take into account the motion of all objects in the current scene. What’s more, our method can be simply applied into other DETR-like models to boost performance without introducing huge computation and memory cost. As a result, our History-DETR results in a remarkable improvement(+1.1% NDS) under negligible inference time increase.
Weijie Luo, Zihao Liu 0018, Guohao Dai 0001, Ningyi Xu
MMAsia4
2021 A Low-Latency FPGA Implementation for Real-Time Object Detection
abstract
The advancement of object detection algorithms makes them widely used in autonomous systems. However, due to high computational complexity of Convolutional Neural Networks(CNN), stringent latency requirement is hard to meet for real-time object detection. To address this problem, a low-latency accelerator architecture is proposed in this paper. A fine-grained column-based pipeline architecture with padding skip technique is implemented to reduce the start-up time of pipeline. In order to cut down the computational time of CNN, double signed-multiplication correcting circuit is introduced. In addition, pooling unit with share buffer is proposed to reduce storage cost for pooling layer. To demonstrate our new architecture, we implement the YOLOv2-tiny deep neural network (you-only-look-once) with input size 1280×384 on ZC706 development board, improving the latency by 2.125× to 2.34× compared to previous FPGA accelerator for YOLOv2-tiny.
Lifu Cheng, Cen Li, Yongfu Li 0002, Guanghui He 0002, Ningyi Xu, Yong Lian 0001
ISCAS6
2020 Crane: Mitigating Accelerator Under-utilization Caused by Sparsity Irregularities in CNNs
abstract
Convolutional neural networks (CNNs) have achieved great success in numerous AI applications. To improve inference efficiency of CNNs, researchers have proposed various pruning techniques to reduce both computation intensity and storage overhead. These pruning techniques result in multi-level sparsity irregularities in CNNs. Together with that in activation matrices, which is induced by employment of ReLU activation function, all these sparsity irregularities cause a serious problem of computation resource under-utilization in sparse CNN accelerators. To mitigate this problem, we propose a method of load-balancing based on a workload stealing technique. We demonstrate that this method can be applied to two major inference data-flows, which cover all state-of-the-art sparse CNN accelerators. Based on this method, we present an accelerator, called Crane, which addresses all kinds of sparsity irregularities in CNNs. We perform a fair comparison between Crane and state-of-the-art prior approaches. Experimental results show that Crane improves performance by 27% ~ 88% and reduces energy consumption by 16% ~ 48%, respectively, compared to the counterparts.
Yijin Guan, Guangyu Sun 0003, Zhihang Yuan, Ningyi Xu, Jason Cong, Yuan Xie 0001
IEEE Trans. Computers5
2019 FlexSaaS: A Reconfigurable Accelerator for Web Search Selection
abstract
Web search engines deploy large-scale selection services on CPUs to identify a set of web pages that match user queries. An FPGA-based accelerator can exploit various levels of parallelism and provide a lower latency, higher throughput, more energy-efficient solution than commodity CPUs. However, maintaining such a customized accelerator in a commercial search engine is challenging because selection services are changed often. This article presents our design for FlexSaaS (Flexible Selection as a Service), an FPGA-based accelerator for web search selection. To address efficiency and flexibility challenges, FlexSaaS abstracts computing models and separates memory access from computation. Specifically, FlexSaaS (i) contains a reconfigurable number of matching processors that can handle various possible query plans, (ii) decouples index stream reading from matching computation to fetch and decode index files, and (iii) includes a universal memory accessor that hides the complex memory hierarchy and reduces host data access latency. Evaluated on FPGAs in the selection service of a commercial web search--the Bing web search engine—FlexSaaS can be evolved quickly to adapt to new updates. Compared to the software baseline, FlexSaaS on Arria 10 reduces average latency by 30% and increases throughput by 1.5×.
Shijie Cao, Lanshun Nie, Dechen Zhan, Ningyi Xu, Ramashis Das, Ming Wu 0007, Derek Chiou
ACM Trans. Reconfigurable Technol. Syst.5
2017 Using Data Compression for Optimizing FPGA-Based Convolutional Neural Network Accelerators
Yijin Guan, Ningyi Xu, Chen Zhang 0001, Zhihang Yuan, Jason Cong
APPT2
2017 FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid Templates
abstract
DNNs (Deep Neural Networks) have demonstrated great success in numerous applications such as image classification, speech recognition, video analysis, etc. However, DNNs are much more computation-intensive and memory-intensive than previous shallow models. Thus, it is challenging to deploy DNNs in both large-scale data centers and real-time embedded systems. Considering performance, flexibility, and energy efficiency, FPGA-based accelerator for DNNs is a promising solution. Unfortunately, conventional accelerator design flows make it difficult for FPGA developers to keep up with the fast pace of innovations in DNNs. To overcome this problem, we propose FP-DNN (Field Programmable DNN), an end-to-end framework that takes TensorFlow-described DNNs as input, and automatically generates the hardware implementations on FPGA boards with RTL-HLS hybrid templates. FP-DNN performs model inference of DNNs with our high-performance computation engine and carefully-designed communication optimization strategies. We implement CNNs, LSTM-RNNs, and Residual Nets with FPDNN, and experimental results show the great performance and flexibility provided by our proposed FP-DNN framework.
Yijin Guan, Hao Liang 0003, Ningyi Xu, Shaoshuai Shi, Xi Chen 0107, Guangyu Sun 0003, Wei Zhang 0012, Jason Cong
FCCM3
2017 ForeGraph: Exploring Large-scale Graph Processing on Multi-FPGA Architecture
Guohao Dai 0001, Yuze Chi, Ningyi Xu, Yu Wang 0002, Huazhong Yang
FPGA4
2017 FxpNet: Training a deep convolutional neural network in fixed-point representation
abstract
We introduce FxpNet, a framework to train deep convolutional neural networks with low bit-width arithmetics in both forward pass and backward pass. During training FxpNet further reduces the bit-width of stored parameters (also known as primal parameters) by adaptively updating their fixed-point formats. These primal parameters are usually represented in the full resolution of floating-point values in previous binarized and quantized neural networks. In FxpNet, during forward pass fixed-point primal weights and activations are first binarized before computation, while in backward pass all gradients are represented as low resolution fixed-point values and then accumulated to corresponding fixed-point primal parameters. To have highly efficient implementations in FPGAs, ASICs and other dedicated devices, FxpNet introduces Integer Batch Normalization (IBN) and Fixed-point ADAM (FxpADAM) methods to further reduce the required floating-point operations, which will save considerable power and chip area. The evaluation on CIFAR-10 dataset indicates the effectiveness that FxpNet with 12-bit primal parameters and 12-bit gradients achieves comparable prediction accuracy with state-of-the-art binarized and quantized neural networks.
Xi Chen 0107, Hucheng Zhou, Ningyi Xu
IJCNN4
2016 Going Deeper with Embedded FPGA Platform for Convolutional Neural Network
abstract
In recent years, convolutional neural network (CNN) based methods have achieved great success in a large number of applications and have been among the most powerful and widely used techniques in computer vision. However, CNN-based methods are com-putational-intensive and resource-consuming, and thus are hard to be integrated into embedded systems such as smart phones, smart glasses, and robots. FPGA is one of the most promising platforms for accelerating CNN, but the limited bandwidth and on-chip memory size limit the performance of FPGA accelerator for CNN.
Jiantao Qiu, Jie Wang 0022, Kaiyuan Guo, Boxun Li, Erjin Zhou, Tianqi Tang 0001, Ningyi Xu, Sen Song, Yu Wang 0002, Huazhong Yang
FPGA9
2016 ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable Hardware
abstract
Highly flexible software network functions (NFs) are crucial components to enable multi-tenancy in the clouds. However, software packet processing on a commodity server has limited capacity and induces high latency. While software NFs could scale out using more servers, doing so adds significant cost. This paper focuses on accelerating NFs with programmable hardware, i.e., FPGA, which is now a mature technology and inexpensive for datacenters. However, FPGA is predominately programmed using low-level hardware description languages (HDLs), which are hard to code and difficult to debug. More importantly, HDLs are almost inaccessible for most software programmers. This paper presents ClickNP, a FPGA-accelerated platform for highly flexible and high-performance NFs with commodity servers. ClickNP is highly flexible as it is completely programmable using high-level C-like languages, and exposes a modular programming abstraction that resembles Click Modular Router. ClickNP is also high performance. Our prototype NFs show that they can process traffic at up to 200 million packets per second with ultra-low latency ($< 2\mu$s). Compared to existing software counterparts, with FPGA, ClickNP improves throughput by 10x, while reducing latency by 10x. To the best of our knowledge, ClickNP is the first FPGA-accelerated platform for NFs, written completely in high-level language and achieving 40 Gbps line rate at any packet size.
Bojie Li, Kun Tan 0002, Layong Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, Peng Cheng 0005
SIGCOMM6
2015 Real-Time High-Quality Stereo Vision System in FPGA
abstract
Stereo vision is a well-known technique for acquiring depth information. In this paper, we propose a real-time high-quality stereo vision system in field-programmable gate array (FPGA). Using absolute difference-census cost initialization, cross-based cost aggregation, and semiglobal optimization, the system provides high-quality depth results for high-definition images. This is the first complete real-time hardware system that supports both cost aggregation on variable support regions and semiglobal optimization in FPGAs. Furthermore, the system is designed to be scaled with image resolution, disparity range, and parallelism degree for maximum parallel efficiency. We present the depth map quality on the Middlebury benchmark and some real-world scenarios with different image resolutions. The results show that our system performs the best among FPGA-based stereo vision systems and its accuracy is comparable with those of current top-performing software implementations. The first version of the system was demonstrated on an Altera Stratix-IV FPGA board, processing 1024 × 768 pixel images with 96 disparity levels at 67 frames/s. The system is then scaled up on a new Altera Stratix-V FPGA and the processing ability is enhanced to 1600 × 1200 pixel images with 128 disparity levels at 42 frames/s.
Ningyi Xu, Yu Wang 0002, Feng-Hsiung Hsu
IEEE Trans. Circuits Syst. Video Technol.3
2014 Energy efficient neural networks for big data analytics
abstract
The world is experiencing a data revolution to discover knowledge in big data. Large scale neural networks are one of the mainstream tools of big data analytics. Processing big data with large scale neural networks includes two phases: the training phase and the operation phase. Huge computing power is required to support the training phase. And the energy efficiency (power efficiency) is one of the major considerations of the operation phase. We first explore the computing power of GPUs for big data analytics and demonstrate an efficient GPU implementation of the training phase of large scale recurrent neural networks (RNNs). We then introduce a promising ultrahigh energy efficient implementation of neural networks' operation phase by taking advantage of the emerging memristor technique. Experiment results show that the proposed GPU implementation of RNNs is able to achieve 2 ~ 11× speed-up compared with the basic CPU implementation. And the scaled-up recurrent neural network trained with GPUs realizes an accuracy of 47% on the Microsoft Research Sentence Completion Challenge, the best result achieved by a single RNN on the same dataset. In addition, the proposed memristor-based implementation of neural networks demonstrates power efficiency of > 400 GFLOPS/W and achieves energy savings of 22× on the HMAX model compared with its pure digital implementation counterpart.
Yu Wang 0002, Boxun Li, Yiran Chen 0001, Ningyi Xu, Huazhong Yang
DATE5
2014 Large scale recurrent neural network on GPU
abstract
Large scale artificial neural networks (ANNs) have been widely used in data processing applications. The recurrent neural network (RNN) is a special type of neural network equipped with additional recurrent connections. Such a unique architecture enables the recurrent neural network to remember the past processed information and makes it an expressive model for nonlinear sequence processing tasks. However, the large computation complexity makes it difficult to effectively train a recurrent neural network and therefore significantly limits the research on the recurrent neural network in the last 20 years. In recent years, the use of graphics processing units (GPUs) becomes a significant advance to speed up the training process of large scale neural networks by taking advantage of the massive parallelism capabilities of GPUs. In this paper, we propose an efficient GPU implementation of the large scale recurrent neural network and demonstrate the power of scaling up the recurrent neural network with GPUs. We first explore the potential parallelism of the recurrent neural network and propose a fine-grained two-stage pipeline implementation. Experiment results show that the proposed GPU implementation can achieve 2 ~ 11 x speed-up compared with the basic CPU implementation with the Intel Math Kernel Library. We then use the proposed GPU implementation to scale up the recurrent neural network and improve its performance. The experiment results of the Microsoft Research Sentence Completion Challenge demonstrate that the large scale recurrent network without class layer is able to beat the traditional class-based modest-size recurrent network and achieve an accuracy of 47%, the best result achieved by a single recurrent neural network on the same dataset.
Boxun Li, Erjin Zhou, Jiayi Duan, Yu Wang 0002, Ningyi Xu, Huazhong Yang
IJCNN6
2013 Real-time high-quality stereo vision system in FPGA
abstract
Stereo vision is a well-known technique for acquiring depth information. In this paper, we present an FPGA-based real-time high-quality stereo vision system. By using AD-Census cost initialization, cross-based aggregation and semi-global optimization, the system provides high-quality depth results for highdefinition images. This is the first complete real-time hardware system that supports both cost aggregation on cross-based regions and semi-global optimization on FPGA. The system can adjust image resolution, parallelism degree, and support region size to achieve maximum efficiency flexibly during the implementation. We test the accuracy of the system on the Middlebury benchmark and some real-world scenarios with different image resolutions. The results show the accuracy is among the best of FPGA-based stereo vision systems and competitive with current top-performing software implementations. We demonstrate the system using an Altera Stratix-IV FPGA board, processing 1024 × 768 pixel images at 30 frames per second.
Ningyi Xu, Yu Wang 0002, Feng-Hsiung Hsu
FPT3
2012 Efficient Query Processing for Web Search Engine with FPGAs
abstract
Web search engines are now using tens of thousands of index servers that consume huge amount of power. In this paper, we investigate FPGAs as the implementation platform for power efficient index serving. We propose the architecture of an FPGA-based inverted index search engine, as well as implementations of essential components, including decoder, matcher and ranker. We successfully boot up the FPGA-based search engine and run experiments on real-world data from a commercial search engine. The targeted FPGA-based hardware index server could achieve up to 19.52X power efficiency and 7.17X price efficiency over an Intel Xeon server with highly optimized software. This is the first complete work using FPGAs to implement query processing for Web search engines.
Zhanxiang Zhao, Ningyi Xu, Lin-Tao Zhang, Feng-Hsiung Hsu
FCCM3
2011 Gemma in April: A matrix-like parallel programming architecture on OpenCL
abstract
Nowadays, Graphics Processing Unit (GPU), as a kind of massive parallel processor, has been widely used in general purposed computing tasks. Although there have been mature development tools, it is not a trivial task for programmers to write GPU programs. Based on this consideration, we propose a novel parallel computing architecture. The architecture includes a parallel programming model, named Gemma, and a programming framework, named April. Gemma is based on generalized matrix operations, and helps to alleviate the difficulty of describing parallel algorithms. April is a high-level framework that can compile and execute tasks described in Gemma with OpenCL. In particular, April can automatically 1) choose the best parallel algorithm and mapping scheme, and generate OpenCL kernels, 2) schedule Gemma tasks based on execution costs such as data storing and transferring. Our experimental results show that with competitive performance, April considerably reduces the programs' code length compared with OpenCL.
Tianji Wu, Di Wu 0013, Yu Wang 0002, Ningyi Xu, Huazhong Yang
DATE6
2011 A heterogeneous accelerator platform for multi-subject voxel-based brain network analysis
abstract
The research on understanding the human brain has attracted more and more attention. A promising method is to model the brain as a network based on modern imaging technologies and then to apply graph theory algorithms for analysis. In this work, we examine the computing bottleneck of this method, and propose a CPU-GPU heterogeneous platform to accelerate the process. We construct a statistical brain network from a sample of 198 people and get characteristics such as nodal degree and modularity. This is the first study of voxel-based brain networks on large samples. We also illustrate that domain-specific hardware platform can have a significant impact on neuroscience studies.
Yu Wang 0002, Mo Xu, Ling Ren 0001, Di Wu 0013, Yong He 0002, Ningyi Xu, Huazhong Yang
ICCAD7
2011 An FPGA-based accelerator for LambdaRank in Web search engines
abstract
In modern Web search engines, Neural Network (NN)-based learning to rank algorithms is intensively used to increase the quality of search results. LambdaRank is one such algorithm. However, it is hard to be efficiently accelerated by computer clusters or GPUs, because: (i) the cost function for the ranking problem is much more complex than that of traditional Back-Propagation(BP) NNs, and (ii) no coarse-grained parallelism exists in the algorithm. This article presents an FPGA-based accelerator solution to provide high computing performance with low power consumption. A compact deep pipeline is proposed to handle the complex computing in the batch updating. The area scales linearly with the number of hidden nodes in the algorithm. We also carefully design a data format to enable streaming consumption of the training data from the host computer. The accelerator shows up to 15.3X (with PCIe x4) and 23.9X (with PCIe x8) speedup compared with the pure software implementation on datasets from a commercial search engine.
Ningyi Xu, Xiongfei Cai, Yu Wang 0002, Feng-Hsiung Hsu
ACM Trans. Reconfigurable Technol. Syst.2
2010 FPMR: MapReduce framework on FPGA
abstract
Machine learning and data mining are gaining increasing attentions of the computing society. FPGA provides a highly parallel, low power, and flexible hardware platform for this domain, while the difficulty of programming FPGA greatly limits its prevalence. MapReduce is a parallel programming framework that could easily utilize inherent parallelism in algorithms. In this paper, we describe FPMR, a MapReduce framework on FPGA, which provides programming abstraction, hardware architecture, and basic building blocks to developers.
Bo Wang 0067, Yu Wang 0002, Ningyi Xu, Huazhong Yang
FPGA5
2010 LambdaRank acceleration for relevance ranking in web search engines (abstract only)
abstract
This paper describes a FPGA-based hardware acceleration system for LambdaRank algorithm. LambdaRank Algorithm is a Neural Network (NN)-based learning to rank algorithm. It is intensively used by web search engine companies to increase the search relevance. Since i) the cost function for the ranking problem is much more complex than that of traditional Back-Propagation(BP) NNs, and ii) no coarse-grained parallelism exists, LambdaRank is hard to be efficiently accelerated by GPU or computer clusters. We presents a FPGA-based accelerator solution to provide high computing performance. A compact deep pipeline is proposed to handle the complex computing in the batch updating. The area scales linearly with the number of hidden nodes in the NN model. We also carefully design a data format to enable streaming consumption of the training data from host computer. The accelerator shows up to 24.6 speedup compared with the pure software implementation on datasets from a commercial search engine.
Ningyi Xu, Xiongfei Cai, Yu Wang 0002, Feng-Hsiung Hsu
FPGA2
2010 A compression method for inverted index and its FPGA-based decompression solution
abstract
Reconfigurable computing based on FPGA is a promising solution to accelerate applications for web search engines. Due to the challenge of such data-intensive applications, data compression has become much more important. This paper proposes a data compression method for inverted indices, which combines the bit-level compression method - Huffman coding and a coarse-grained compression method, to achieve a balanced performance in compression ratio and decompression speed. Because an inverted index is only compressed once, the compression speed is not the major measurement for a compression method. The proposed method shows good to 21.61% compression ratio on inverted indices from a commercial search engine. This compression ratio is better than results by other existing compression methods. We also develop an efficient FPGA-based hardware decompression module, which could provide up to 996 MBps input bandwidth for the accelerator system.
Ningyi Xu, Zenglin Xia, Feng-Hsiung Hsu
FPT2
2010 Making Human Connectome Faster: GPU Acceleration of Brain Network Analysis
abstract
The research on complex Brain Networks plays a vital role in understanding the connectivity patterns of the human brain and disease-related alterations. Recent studies have suggested a noninvasive way to model and analyze human brain networks by using multi-modal imaging and graph theoretical approaches. Both the construction and analysis of the Brain Networks require tremendous computation. As a result, most current studies of the Brain Networks are focused on a coarse scale based on Brain Regions. Networks on this scale usually consist around 100 nodes. The more accurate and meticulous voxel-base Brain Networks, on the other hand, may consist 20K to 100K nodes. In response to the difficulties of analyzing large-scale networks, we propose an acceleration framework for voxel-base Brain Network Analysis based on Graphics Processing Unit (GPU). Our GPU implementations of Brain Network construction and modularity achieve 24x and 80x speedup respectively, compared with single-core CPU. Our work makes the processing time affordable to analyze multiple large-scale Brain Networks.
Di Wu 0013, Tianji Wu, Yu Wang 0002, Yong He 0002, Ningyi Xu, Huazhong Yang
ICPADS6
2010 Efficient PageRank and SpMV Computation on AMD GPUs
abstract
Google's famous PageRank algorithm is widely used to determine the importance of web pages in search engines. Given the large number of web pages on the World Wide Web, efficient computation of PageRank becomes a challenging problem. We accelerated the power method for computing PageRank on AMD GPUs. The core component of the power method is the Sparse Matrix-Vector Multiplication (SpMV). Its performance is largely determined by the characteristics of the sparse matrix, such as sparseness and distribution of non-zero values. Based on careful analysis on the web linkage matrices, we design a fast and scalable SpMV routine with three passes, using a modified Compressed Sparse Row format. Our PageRank computation achieves 15x speedup on a Radeon 5870 Graphic Card compared with a PhenomII 965 CPU at 3.4GHz. Our method can easily adapt to large scale data sets. We also compare the performance of the same method on the OpenCL platform with our low-level implementation.
Tianji Wu, Bo Wang 0067, Feng Yan 0003, Yu Wang 0002, Ningyi Xu
ICPP6
2009 FPGA-based acceleration of neural network for ranking in web search engine with a streaming architecture
abstract
Web search engine companies are intensively running learning to rank algorithms to improve the search relevance. Neural network (NN)-based approaches, such as LambdaRank, can significantly increase the ranking quality. While, their training is very slow on a single computer and inherent coarse-grained parallelism could be hardly utilized by computer clusters. Thus an efficient implementation is necessary to timely generate acceptable NN models on frequently updated training datasets. This paper presents our work in accelerator. A SIMD streaming architecture is proposed to i) efficiently map the query-level NN computation and data structure to FPGA, ii) fully exploit the inherent fine-grained parallelism, and iii) provide scalability to large scale datasets. The accelerator shows up to 17.9X speedup over the software implementation on datasets from a commercial search engine.
Ningyi Xu, Xiongfei Cai, Yu Wang 0002, Feng-Hsiung Hsu
FPL2
2009 RankBoost Acceleration on both NVIDIA CUDA and ATI Stream Platforms
abstract
NVIDIA CUDA and ATI Stream are the two major general-purpose GPU (GPGPU) computing technologies. We implemented RankBoost, a web relevance ranking algorithm, on both NVIDIA CUDA and ATI Stream platforms to accelerate the algorithm and illustrate the differences between these two technologies. It shows that the performances of GPU programs are highly dependent on the utilization of GPU's hardware memory architectural features. In this work, we accelerated RankBoost algorithm on both platforms, and we achieved 22.9X speedup on CUDA and 9.2X speedup on ATI Stream respectively. Then we made a comparison on the differences of memory architecture between NVIDIA CUDA and ATI Stream.
Bo Wang 0067, Tianji Wu, Feng Yan 0003, Ningyi Xu, Yu Wang 0002
ICPADS5
2009 FTL design exploration in reconfigurable high-performance SSD for server applications
abstract
Solid-state disks (SSDs) are becoming widely used in personal computers and are expected to replace a great portion of magnetic disks in servers and supercomputers. Although many high-speed SSDs are present in the market, both the design of hardware architecture and the details of the flash translation layer (FTL) are not well known. Meanwhile, in the systems requiring high-end storages, specially tuned SSDs can perform better than the generic ones, because the applications in such environment are usually fixed.
Ji-Yong Shin, Zenglin Xia, Ningyi Xu, Xiongfei Cai, Seung Ryoul Maeng, Feng-Hsiung Hsu
ICS3
2009 Parallel Inference for Latent Dirichlet Allocation on Graphics Processing Units
abstract
The recent emergence of Graphics Processing Units (GPUs) as general-purpose parallel computing devices provides us with new opportunities to develop scalable learning methods for massive data. In this work, we consider the problem of parallelizing two inference methods on GPUs for latent Dirichlet Allocation (LDA) models, collapsed Gibbs sampling (CGS) and collapsed variational Bayesian (CVB). To address limited memory constraints on GPUs, we propose a novel data partitioning scheme that effectively reduces the memory cost. Furthermore, the partitioning scheme balances the computational cost on each multiprocessor and enables us to easily avoid memory access conflicts. We also use data streaming to handle extremely large datasets. Extensive experiments showed that our parallel inference methods consistently produced LDA models with the same predictive power as sequential training methods did but with 26x speedup for CGS and 196x speedup for CVB on a GPU with 30 multiprocessors; actually the speedup is almost linearly scalable with the number of multiprocessors available. The proposed partitioning scheme and data streaming can be easily ported to many other models in machine learning.
Feng Yan 0003, Ningyi Xu, Yuan Qi 0001
NIPS2
2009 FPGA Acceleration of RankBoost in Web Search Engines
abstract
Search relevance is a key measurement for the usefulness of search engines. Shift of search relevance among search engines can easily change a search company's market cap by tens of billions of dollars. With the ever-increasing scale of the Web, machine learning technologies have become important tools to improve search relevance ranking. RankBoost is a promising algorithm in this area, but it is not widely used due to its long training time. To reduce the computation time for RankBoost, we designed a FPGA-based accelerator system and its upgraded version. The accelerator, plugged into a commodity PC, increased the training speed on MSN search engine data up to 1800x compared to the original software implementation on a server. The proposed accelerator has been successfully used by researchers in the search relevance ranking.
Ningyi Xu, Xiongfei Cai, Lei Zhang 0001, Feng-Hsiung Hsu
ACM Trans. Reconfigurable Technol. Syst.1
2008 Distributed RankBoost Acceleration Using FPGA and MPI for Web Relevance Ranking
abstract
Web search engine ranks web pages according to their relevance to user queries, which is critical for the success of commercial search engines. Rank Boost algorithm is promising in Web relevance ranking area, while its computation complexity makes our existing implementations (including single node software-based implementation and a FPGA-based accelerator) too slow to reflect the dynamics of the Web. Moreover, previous implementations can not handle the huge web-scale data. As such, in this paper, we present the RankBoost implementation on a MPI-based distributed FPGA-based accelerators. Our results show that the combination of the coarse parallel efficiency of distributed system and the fine parallel efficiency of reconfigurable hardware accelerators can significantly increase the computing performance.
Ningyi Xu, Feng-Hsiung Hsu, Xiongfei Cai, Zenglin Xia
ICPADS2
2007 FPGA-based Accelerator Design for RankBoost in Web Search Engines
abstract
Search relevance is a key measurement for the usefulness of search engines. Shift of search relevance among search engines can easily change a search company's market cap by tens of billions of dollars. With the ever-increasing scale of the Web, machine learning technologies have become important tools to improve search relevance ranking. RankBoost is a promising algorithm in this area, but it is not widely used due to its long training time. To reduce the computation time for RankBoost, we designed a FPGA-based accelerator system. The accelerator, plugged into a commodity PC, increased the training speed on MSN search engine data by 2 orders of magnitude compared to the original software implementation on a server. The proposed accelerator has been successfully used by researchers in the search relevance ranking.
Ningyi Xu, Xiongfei Cai, Lei Zhang 0001, Feng-Hsiung Hsu
FPT1
2005 The design and implementation of a DVB receiving chip with PCI interface
abstract
A DVB receiving chip with PCI interface for PC is presented. The chip supports DVB protocols and integrates useful interfaces, including I2C, SmartCard and PCI. A card with this chip could change PC into digital TV terminal. The architecture of FPGA prototype system together with some main design issues is introduced. The experimental result shows that the chip could accomplish required functionalities.
Ningyi Xu, Guanghui He 0002, Zucheng Zhou
ASP-DAC1