Sitao Huang

dblp:126/3480 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0001-7669-1467ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 4 first-author · 18 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HARK: Hierarchical Agentic Retrieval with Keyframing for Video Understanding (Student Abstract)
abstract
Current video understanding models struggle with temporal reasoning and efficient processing while balancing detail preservation with computational efficiency. We propose a hierarchical memory system that segments videos into action and scene units, combined with question-aware agentic keyframe selection. Our method achieves 70.3% overall accuracy on VideoMME short video benchmarks.
Jingcheng Li, Ye Qiao, Sitao Huang
AAAI3
2026 Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs (Student Abstract)
abstract
Extending LLM context windows is key for long-range tasks. RoPE-based position interpolation (PI) scales input length without retraining, and post-training quantization (PTQ) enables efficient deployment; however, combining PI with PTQ degrades accuracy due to long-context aliasing, dynamic-range dilation, axis-grid anisotropy, and outlier shifts that induce position-dependent logit noise. We give the first systematic analysis of PI+PTQ and propose two diagnostics: Interpolation Pressure (per-band phase-scaling sensitivity) and Tail Inflation Ratio (outlier shift from short to long contexts). We then introduce Q-ROAR, a RoPE-aware, weight-only stabilization that bands RoPE dimensions and lightly searches per-band scales for W_Q,W_K, with an optional symmetric variant. Q-ROAR needs only a tiny long-context dev set and no fine-tuning or kernel changes, recovering up to 0.7% accuracy and more than 14% GovReport perplexity reduction while preserving short-context performance.
Ye Qiao, Sitao Huang
AAAI2
2026 APEX-Q: Arbitrary-dimension Product-EXtension Quantization for Accelerated LLM Deployment (Student Abstract)
abstract
We present APEX-Q, a flexible product quantization framework for compressing large language models. Unlike prior multi-codebook quantization methods with fixed partitions, APEX-Q supports arbitrary-dimensional tensor quantization, better capturing weight redundancy. It achieves performance on par with 4-bit and 8-bit baselines, enables post-training quantization without retraining, and reveals key trade-offs across subvector dimensions, codebook sizes, and hardware efficiency. APEX-Q thus provides a unified, hardware-friendly approach to scalable LLM deployment.
Ye Qiao, Sitao Huang, Hyoukjun Kwon
AAAI3
2026 TeLLMe: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
abstract
With the emergence of wearable devices and other embedded systems, deploying large language models (LLMs) on edge platforms becomes an urgent need. However, it is challenging because of their high computational and memory demands. Although recent low-bitwidth quantization methods (e.g., BitNet, DeepSeek) compress weights to as low as 1.58 bits with minimal accuracy loss, edge deployment is still constrained by limited on-chip resources, power budgets, and the often-neglected long latency of the prefill stage. We present TeLLMe, the first table-lookup-based ternary LLM accelerator for low-power edge FPGAs that fully supports both prefill and autoregressive decoding using 1.58-bit weights and 8-bit activations. TeLLMe incorporates our proposed novel techniques including (1) a table-lookup-based ternary matrix multiplication (TLMM) engine utilizing grouped activations and online precomputation for low resource utilization and high throughput; (2) a fine-grained URAM-based weight buffer management scheme supporting weight loading from global memory and compute engine weight access; (3) a streaming dataflow architecture that fuses floating-point element-wise operations with linear computations to hide latency; (4) a reversed-reordered prefill stage attention with fused attention operation for high memory efficiency; and (5) a resource-efficient specialized decoding stage attention. Under a 5W power budget, TeLLMe delivers up to 25 tokens/s decoding throughput and 0.45s to 0.96s Time-to-First-Token (TTFT) for 64–128 token prompts, marking a significant energy-efficiency advancement in LLM inference on edge FPGAs.
Ye Qiao, Sitao Huang
FPGA5
2026 Characterizing State Space Model and Hybrid Language Model Performance with Long Context
abstract
Emerging applications such as AR are driving demands for machine intelligence capable of processing continuous and/or long-context inputs on local devices. However, currently dominant models based on Transformer architecture suffers from the quadratic computational and memory overhead, which hinders applications required to process long contexts. This has spurred a paradigm shift towards new architectures like State Space Models (SSMs) and SSM-Transformer hybrid models, which provide near-linear scaling. The near-linear scaling enabled efficient handling of millions of tokens while delivering high performance in recent studies. Although such works present promising results, their workload characteristics in terms of computational performance and hardware resource requirements are not yet thoroughly explored, which limits our understanding of their implications to the system level optimizations. To address this gap, we present a comprehensive, comparative benchmarking of carefully selected Transformers, SSMs, and hybrid models specifically for long-context inference on consumer and embedded GPUs. Our analysis shows that SSMs are well-suited for on-device AI on consumer and embedded GPUs for long context inferences. While Transformers are up to $1.9 \times$ faster at short sequences ($ \lt 8 \mathrm{~K}$ tokens), SSMs demonstrate a dramatic performance inversion, becoming up to $4 \times$ faster at very long contexts ($\sim 57 \mathrm{~K}$ tokens), thanks to their linear computational complexity and $\boldsymbol{\sim} \mathbf{6 4 \%}$ reduced memory footprint. Our operator-level analysis reveals that custom SSM kernels like selective scan despite being hardware-aware to minimize memory IO, dominate the inference runtime on edge platforms, accounting for over $55 \%$ of latency due to their sequential, element-wise nature. To foster further research, we are sharing CPU/GPU profiling traces and have made our characterization framework SSM-Scope open-sourced at https://github.com/sapmitra/ssm-scope
Saptarshi Mitra, Rachid Karami, Haocheng Xu, Sitao Huang, Hyoukjun Kwon
ISPASS4
2025 SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training
abstract
The growth rate of the GPU memory capacity has not been able to keep up with that of the size of large language models (LLMs), hindering the model training process. In particular, activations-the intermediate tensors produced during forward propagation and reused in backward propagation-dominate the GPU memory use. This leads to high training overheads such as expensive weight update costs due to the small micro-batch size. To address this challenge, we propose SSDTrain, an adaptive activation offloading framework to high-capacity NVMe SSDs. SSDTrain reduces GPU memory usage without impacting performance by fully overlapping data transfers with computation. SSDTrain is compatible with popular deep learning frameworks like PyTorch, Megatron, and DeepSpeed, and it employs techniques such as tensor deduplication and forwarding to further enhance efficiency. We extensively experimented with popular LLMs like GPT, BERT, and T5. Results demonstrate that SSDTrain reduces 47% of the activation peak memory usage. At the same time, SSDTrain perfectly overlaps the I/O with the computation and incurs negligible overhead. Compared with keeping activations in GPU memory and layerwise full recomputation, SSDTrain achieves the best memory savings with negligible throughput loss. We further analyze how the reduced activation memory use may be leveraged to increase throughput by increasing micro-batch size and reducing pipeline parallelism bubbles.
Kun Wu 0002, Jeongmin Brian Park, Xiaofan Zhang 0001, Mert Hidayetoglu, Vikram S. Mailthody, Sitao Huang, Steven S. Lumetta, Wen-Mei W. Hwu
DAC6
2025 COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
abstract
Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and high demands in computation, memory, and communication limit their deployment to edge platforms for local, secure inference. Binary transformers offer a compact, low-complexity solution for edge deployment with reduced bandwidth needs and acceptable accuracy. However, existing binary transformers perform inefficiently on current hardware due to the lack of binary specific optimizations. To address this, we introduce COBRA, an algorithm-architecture co-optimized binary Transformer accelerator for edge computing. COBRA features a real 1-bit binary multiplication unit, enabling matrix operations with -1, 0, and +1 values, surpassing ternary methods. With further hardware-friendly optimizations in the attention block, COBRA achieves up to 3,894.7 GOPS throughput and 448.7 GOPS/Watt energy efficiency on edge FPGAs, delivering a 311× energy efficiency improvement over GPUs and a 3.5× throughput improvement over the state-of-the-art binary accelerator, with only negligible inference accuracy degradation.
Ye Qiao, Yunzhe Deng, Sitao Huang
ICCAD6
2025 RSEND: Retinex-based Squeeze and Excitation Network with Dark Region Detection for Efficient Low Light Image Enhancement
abstract
Images captured under low-light scenarios often suffer from low quality. Efficient low-light image enhancement with mobile computing has become an urgent need. Previous CNN-based low-light image enhancement methods often involve using Retinex theory. Nevertheless, most of them do not perform well in complicated datasets like LOL-v2 while using too much computational resources. Besides, some of these methods require sophisticated training at different stages, making the procedure even more time-consuming and tedious. In this paper, we propose an accurate, concise, and one-stage Retinex theory-based framework with a novel dark region detection module and Squeeze and Excitation blocks for enhanced detail retention, RSEND, for efficient low-light image enhancement. RSEND first divides the low-light image into the illumination map and reflectance map, then detects the different dark regions in the illumination map and performs light enhancement. After this step, it refines the enhanced gray-scale image and does element-wise matrix multiplication with the reflectance map. By denoising the output it has from the previous step, it obtains the final result. In all the steps, RSEND utilizes Squeeze and Excitation network to better capture the details. Comprehensive quantitative and qualitative experiments show that our efficient Retinex model significantly outperforms other CNN-based state-of-the-art models, achieving a PSNR improvement ranging from 1.69 dB to 3.63 dB in different datasets. Compared to Transformer-based models, RSEND achieves higher PSNR values ranging from 1.22 dB to 2.44 dB in the LOL-v2-real dataset. Importantly, RSEND achieves these performance improvements with remarkable efficiency, utilizing only 0.41 million parameters, which represents a substantial reduction (3.93–9.78×) in computational resources compared to existing state-of-the-art methods. The code can be found at https://github.com/jeffconqueror/RSEND/tree/main.
Jingcheng Li, Ye Qiao, Haocheng Xu, Sitao Huang
IJCNN4
2025 MONAS: Efficient Zero-Shot Neural Architecture Search for MCUs
abstract
Neural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined ac-curacy optimization goals. However, previous approaches often involve time-consuming training on super networks or intensive architecture sampling and evaluations. Although various zero-cost proxies correlated with CNN model accuracy have been proposed for efficient architecture search without training, their lack of hardware consideration makes it challenging to target highly resource-constrained edge devices such as microcontroller units (MCUs). To address these challenges, we introduce MONAS, a novel hardware-aware zero-shot NAS framework specifically designed for MCUs in edge computing. MONAS incorporates hardware optimality considerations into the search process through our proposed MCU hardware latency estimation model. By combining this with specialized performance indicators (proxies), MONAS identifies optimal neural architectures without incurring heavy training and evaluation costs, optimizing for both hardware latency and accuracy under resource constraints. MONAS achieves up to a 1104× improvement in search efficiency over previous work targeting MCUs and can discover CNN models with over 3.23× faster inference on MCUs while maintaining similar accuracy compared to more general NAS approaches.
Ye Qiao, Haocheng Xu, Sitao Huang
IJCNN4
2025 HeteroBench: Multi-kernel Benchmarks for Heterogeneous Systems
Hongzheng Tian, Alok Mishra 0002, Rolando P. Hong Enriquez, Dejan S. Milojicic, Eitan Frachtenberg, Sitao Huang
ICPE7
2025 A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
abstract
Path planning is a critical task for autonomous driving, aiming to generate smooth, collision-free, and feasible paths based on input perception and localization information. The planning task is both highly time-sensitive and computationally intensive, posing significant challenges to resource-constrained autonomous driving hardware. In this article, we propose an end-to-end framework for accelerating path planning on FPGA platforms. This framework focuses on accelerating quadratic programming (QP) solving, which is the core of optimization-based path planning and has the most computationally-intensive workloads. Our method leverages a hardware-friendly alternating direction method of multipliers (ADMM) to solve QP problems while employing a highly parallelizable preconditioned conjugate gradient (PCG) method for solving the associated linear systems. We analyze the sparse patterns of matrix operations in QP and design customized storage schemes along with efficient sparse matrix multiplication and sparse matrix-vector multiplication units. Our customized design significantly reduces resource consumption for data storage and computation while dramatically speeding up matrix operations. Additionally, we propose a multi-level dataflow optimization strategy. Within individual operators, we achieve acceleration through parallelization and pipelining. For different operators in an algorithm, we analyze inter-operator data dependencies to enable fine-grained pipelining. At the system level, we map different steps of the planning process to the CPU and FPGA and pipeline these steps to enhance end-to-end throughput. We implement and validate our design on the AMD ZCU102 platform. Our implementation achieves state-of-the-art performance in both latency and energy efficiency compared with existing works, including an average 1.48× speedup over the best FPGA-based design, a 2.89× speedup compared with the state-of-the-art QP solver on an Intel i7-11800H CPU, a 5.62× speedup over an ARM Cortex-A57 embedded CPU, and a 1.56× speedup over state-of-the-art GPU-based work. Furthermore, our design delivers a 2.05× improvement in throughput compared with the state-of-the-art FPGA-based design.
Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang
ACM Trans. Archit. Code Optim.7
2024 Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
abstract
Relational graph neural networks (RGNNs) are graph neural networks with dedicated structures for modeling the different types of nodes and edges in heterogeneous graphs. While RGNNs have been increasingly adopted in many real-world applications due to their versatility and accuracy, they pose performance and system design challenges: inherent memory-intensive computation patterns, the gap between the programming interface and kernel APIs, and heavy programming effort required to optimize kernels caused by their coupling with data layout and heterogeneity. To systematically address these challenges, we propose Hector, a novel two-level intermediate representation and its code generator framework that (a) captures the key properties of RGNN models, and opportunities to reduce memory accesses in inter-operator scheduling and materialization, (b) generates code with flexible data access schemes to eliminate redundant data copies, and (c) decouples model semantics, data layout, and operators-specific optimizations from each other to reduce programming effort. By building on one general matrix multiply (GEMM) template and a node/edge traversal template, Hector achieves up to 9.9× speed-up in inference and 43.7× speed-up in training compared with the state-of-the-art public systems on select models, RGCN, RGAT and HGT, when running heterogeneous graphs provided by Deep Graph Library (DGL) and Open Graph Benchmark (OGB). In addition, Hector does not trigger any out-of-memory (OOM) exception in these tests. We also propose linear operator reordering and compact materialization to further accelerate the system by up to 3.8×. As an indicator of the reduction of programming effort, Hector takes in 51 lines of code expressing the three models and generates a total of 8K lines of CUDA and C++ code. Through profiling, we found that higher memory efficiency allows Hector to accommodate larger input and therefore attain higher throughput in forward propagation, while backward propagation is bound by latency introduced by atomic updates and outer products.
Kun Wu 0002, Mert Hidayetoglu, Xiang Song 0003, Sitao Huang, Da Zheng 0004, Israt Nisa, Wen-Mei W. Hwu
ASPLOS (3)4
2024 MicroNAS: Zero-Shot Neural Architecture Search for MCUs
abstract
Neural architecture search (NAS) effectively discovers new convolutional neural network (CNN) architectures, particularly for accuracy optimization. However, prior approaches often require resource-intensive training on super networks or extensive architecture evaluations, limiting practical applications. To address these challenges, we propose MicroNAS, a hardware-aware zero-shot NAS framework designed for microcontroller units (MCVs) in edge computing. MicroNAS considers target hardware optimality during the search, utilizing specialized performance indicators to identify optimal neural architectures without heavy computational costs. Compared to previous works, MicroNAS achieves up to$1104\times$improvement in search efficiency and discovers models with over$3.23\times$faster MCU inference while maintaining similar accuracy.
Ye Qiao, Haocheng Xu, Sitao Huang
DATE4
2024 Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-Optimizations
abstract
Path planning is a critical task in autonomous driving systems, with quadratic programming being the most time-consuming component. Solving quadratic programming problems using a CPU not only takes a long time but can also lead to high power consumption and costs. In this work, we propose an FPGA-based acceleration method for quadratic programming based path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear equations, which proves to be more scalable and hardware-friendly than the original direct method. We propose optimizations for better memory management, and boost processing throughput and reduce execution time by task level and operator level parallelism with hardware pipelining. Our FPGA-based implementation achieves up to 1.8× speedup and 3.2× power reduction compared with the Intel i5 CPU, 3.1× speedup compared with ARM Cortex-A57.
Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang
FPGA7
2024 A Sparsity-Aware Autonomous Path Planning Accelerator with Algorithm-Architecture Co-Design
abstract
Path planning is a critical task in autonomous driving systems that is most susceptible to real-time constraints but often demands computationally intensive mathematical solvers, two contradictory goals. This conflict makes the computing of path planning a paramount challenge. At the heart of most path planners is the quadratic programming (QP) solver, which places excessive demands on the CPU in real-world autonomous driving applications. In this paper, we present an FPGA-based acceleration framework for path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear systems, which are customized to be more hardware-friendly than prior works. Specific memory management and parallel processing were tailored to the matrix pattern, and the incorporation of pipelining was executed to enhance throughput and execution speed. Our FPGA-based implementation achieves state-of-the-art performance against existing works, including an average 1.98× speedup compared with the state-of-the-art QP solver on Intel i7-11800H CPU, 3.90× speedup over an ARM Cortex-A57 embedded CPU, and 12.3× speedup over an NVIDIA RTX 3090 GPU.
Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang
ICCAD7
2024 HyperDetect: A Real-Time Hyperdimensional Solution for Intrusion Detection in IoT Networks
abstract
Network-based security has emerged as an increasingly critical challenge in the domain of the Internet of Things (IoT). A number of network intrusion detection systems (NIDS), typically relying on sophisticated machine learning (ML) algorithms, have been proposed to monitor network traffic and detect malicious activity. However, these NIDS designs require extensive memory and computational power, exceeding the capability of today’s IoT devices, and often fail to provide timely detection of network attacks. To tackle this issue, we propose HyperDetect, the first attempt at NIDS modeling that leverages the highly efficient and parallel operations of brain-inspired hyperdimensional computing (HDC). Our innovative model updating method effectively mitigates model saturation and significantly reduces the number of retraining iterations needed to reach convergence. Additionally, we employ a novel dynamic encoding technique to regenerate insignificant dimensions, considerably lowering the dimensionalities required to achieve high-quality performance and further accelerating the learning process. HyperDetect delivers on average 5.02× faster training and 31.83× faster inference compared to state-of-the-art (SOTA) learning approaches on a wide range of network intrusion classification tasks. We also extensively evaluate HyperDetect on embedded hardware to demonstrate its low-latency and resource-efficient characteristics.
Junyao Wang 0001, Haocheng Xu, Yonatan Gizachew Achamyeleh, Sitao Huang, Mohammad Abdullah Al Faruque
IEEE Internet Things J.4
2023 Late Breaking Results: Scalable and Efficient Hyperdimensional Computing for Network Intrusion Detection
abstract
Cybersecurity has emerged as a critical challenge for the industry. With the large complexity of the security landscape, sophisticated and costly deep learning models often fail to provide timely detection of cyber threats on edge devices. Brain-inspired hyperdimensional computing (HDC) has been introduced as a promising solution to address this issue. However, existing HDC approaches use static encoders and require very high dimensionality and hundreds of training iterations to achieve reasonable accuracy. This results in a serious loss of learning efficiency and causes huge latency for detecting attacks. In this paper, we propose CyberHD, an innovative HDC learning framework that identifies and regenerates insignificant dimensions to capture complicated patterns of cyber threats with remarkably lower dimensionality. Additionally, the holographic distribution of patterns in high dimensional space provides CyberHD with notably high robustness against hardware errors.
Junyao Wang 0001, Hanning Chen, Mariam Issa, Sitao Huang, Mohsen Imani
DAC4
2023 DistHD: A Learner-Aware Dynamic Encoding Method for Hyperdimensional Classification
abstract
The Internet of Things (IoT) has become an emerging trend that connects heterogeneous devices and enables them with new capabilities. Many applications exploit machine learning methodology to dissect collected data, and edge computing was introduced to enhance the efficiency and scalability in resource-constrained computing environments. Unfortunately, popular deep learning algorithms involve intensive computations that are overcomplicated for edge devices. Brain-inspired Hyperdimensional Computing (HDC) has been considered a promising approach to address this issue. However, existing HDC methods use static encoders, and thus require extremely high dimensionality and hundreds of training iterations to achieve reasonable accuracy. This results in a huge loss of efficiency and severely impedes the application of HDC algorithms in power-limited machines. In this paper, we propose DistHD, a novel HDC framework with a unique dynamic encoding technique consisting of two parts: top-2 classification and dimension regeneration. Our top-2 classification provides top-2 labels for each data sample based on cosine similarity, and dimension regeneration identifies and regenerates dimensions that mislead the classification and reduce the learning quality. The highly parallel algorithm of DistHD effectively accelerates the learning process and achieves the desired accuracy with considerably lower dimensionality. Our evaluation on a wide range of practical classification tasks shows that DistHD is capable of achieving on average 2.12% higher accuracy than state-of-the-art (SOTA) HDC approaches while reducing dimensionality by 8.0×. It delivers 5.97× faster training and 8.09× faster inference than SOTA learning algorithms. Additionally, the holographic distribution of patterns in high dimensional space provides DistHD with 12.90× higher robustness against hardware errors than SOTA DNNs. DistHD has been open-sourced to enable future research in this field.1
Junyao Wang 0001, Sitao Huang, Mohsen Imani
DAC2
2023 CARMA: Context-Aware Runtime Reconfiguration for Energy-Efficient Sensor Fusion
abstract
Autonomous systems (AS) are systems that can adapt and change their behaviors in response to unanticipated events and include systems such as aerial drones, autonomous vehicles, and ground/aquatic robots. AS require a wide array of sensors, deep learning models, and powerful hardware platforms to perceive the environment and safely operate in real-time. However, in many contexts, some sensing modalities negatively impact perception while increasing the system's overall energy consumption. Since AS are often energy-constrained edge devices, energy-efficient sensor fusion methods have been proposed. However, existing methods either fail to adapt to changing scenario conditions or to optimize system-wide energy efficiency. We propose CARMA, a context-aware sensor fusion approach that uses context to dynamically reconfigure the computation flow on a field-programmable gate array (FPGA) at runtime. By clock gating unused sensors and model sub-components, CARMA significantly reduces the energy used by a multi-sensory object detector without compromising performance. We use a deep learning processor unit (DPU) based reconfiguration approach to minimize the latency of model reconfiguration. We evaluate multiple context identification strategies, propose a novel system-wide energy-performance joint optimization, and evaluate scenario-specific perception performance. Across challenging real-world sensing contexts, CARMA outperforms state-of-the-art methods with up to 1.3× speedup and 73% lower energy consumption.
Arnav Vaibhav Malawade, DongHwan Seong, Mohammad Abdullah Al Faruque, Sitao Huang
ISLPED7
2023 Sgap: towards efficient sparse tensor algebra compilation for GPU
Genghan Zhang, Yuetong Zhao, Yanting Tao, Zhongming Yu, Guohao Dai 0001, Sitao Huang, Yuan Wen, Pavlos Petoumenos, Yu Wang 0002
CCF Trans. High Perform. Comput.6
2021 Mixed Precision Quantization for ReRAM-based DNN Inference Accelerators
abstract
ReRAM-based accelerators have shown great potential for accelerating DNN inference because ReRAM crossbars can perform analog matrix-vector multiplication operations with low latency and energy consumption. However, these crossbars require the use of ADCs which constitute a significant fraction of the cost of MVM operations. The overhead of ADCs can be mitigated via partial sum quantization. However, prior quantization flows for DNN inference accelerators do not consider partial sum quantization which is not highly relevant to traditional digital architectures. To address this issue, we propose a mixed precision quantization scheme for ReRAM-based DNN inference accelerators where weight quantization, input quantization, and partial sum quantization are jointly applied for each DNN layer. We also propose an automated quantization flow powered by deep reinforcement learning to search for the best quantization configuration in the large design space. Our evaluation shows that the proposed mixed precision quantization scheme and quantization flow reduce inference latency and energy consumption by up to 3.89x and 4.84x, respectively, while only losing 1.18% in DNN inference accuracy.
Sitao Huang, Aayush Ankit, Plínio Silveira, Rodrigo Antunes, Sai Rahul Chalamalasetti, Izzat El Hajj, Dong Eun Kim, Glaucimar Aguiar, Pedro Bruel, Sergey Serebryakov, Can Li 0024, Paolo Faraboschi, John Paul Strachan, Deming Chen, Kaushik Roy 0001, Wen-Mei W. Hwu, Dejan S. Milojicic
ASP-DAC1
2021 Mind mappings: enabling efficient algorithm-accelerator mapping space search
abstract
Modern day computing increasingly relies on specialization to satiate growing performance and efficiency requirements. A core challenge in designing such specialized hardware architectures is how to perform mapping space search, i.e., search for an optimal mapping from algorithm to hardware. Prior work shows that choosing an inefficient mapping can lead to multiplicative-factor efficiency overheads. Additionally, the search space is not only large but also non-convex and non-smooth, precluding advanced search techniques. As a result, previous works are forced to implement mapping space search using expert choices or sub-optimal search heuristics.
Kartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra, Angshuman Parashar, Christopher W. Fletcher
ASPLOS3
2021 Extending HLS with High-Level Descriptive Language for Configurable Algorithm-Level Spatial Structure Design
abstract
High-level synthesis (HLS) tools have greatly improved the development efficiency of FPGA accelerators in many application areas. With the HLS tools, FPGA designers can focus more on algorithm specifications using software languages such as C/C++, OpenCL, and Python. However, due to the fact that CPU-oriented software languages are designed to describe sequential execution, the repurposing of these languages yields insufficient support for describing parallel data execution and flexible spatial structures on FPGA architecture. To strengthen HLS's ability to describe configurable algorithmlevel spatial structures, we propose fusing hardware-friendly design patterns, namely high-level descriptive language, into imperative programming model on Python.
Chengyue Wang 0002, Sitao Huang, Wen-Mei W. Hwu, Deming Chen
FCCM2
2021 PyLog: An Algorithm-Centric Python-Based FPGA Programming and Synthesis Flow
abstract
The exploding complexity and computation efficiency requirements of applications are stimulating a strong demand for hardware acceleration with heterogeneous platforms such as FPGAs. However, a high-quality FPGA design is very hard to create as it requires FPGA expertise and a long design iteration time. In contrast, software applications are typically developed in a short development cycle, in high-level languages like Python, which is at a much higher level of abstraction than all existing hardware design flows. To close this gap between hardware design flows and software applications, and simplify FPGA programming, we create PyLog, a high-level, algorithm-centric Python-based programming and synthesis flow for FPGA. PyLog is powered by a set of compiler optimization passes and a type inference system to generate high-quality design. It abstracts away the implementation details and allows designers to focus on algorithm specification. PyLog captures more high-level computation patterns for better optimization than traditional HLS systems. PyLog also has a runtime for running PyLog code directly on FPGA platform without any extra code development. Evaluation shows that PyLog significantly improves FPGA design productivity and generates highly efficient FPGA designs that outperform highly optimized CPU and FPGA version by 3.17× and 1.24× on average.
Sitao Huang, Kun Wu 0002, Hyunmin Jeong, Chengyue Wang 0002, Deming Chen, Wen-Mei W. Hwu
FPGA1
2021 Improved GPU Implementations of the Pair-HMM Forward Algorithm for DNA Sequence Alignment
abstract
With the rise of Next-Generation Sequencing (NGS) technology, clinical sequencing services become more accessible but are also facing new challenges. The surging demand motivates developments of more efficient algorithms for computational genomics and their hardware acceleration. In this work, we use GPU to accelerate the DNA variant calling and its related alignment problem. The Pair-Hidden Markov Model (Pair-HMM) is one of the most popular and compute-intensive models used in variant calling. As a critical part of the Pair-HMM, the forward algorithm is not only a computational but data-intensive algorithm. Multiple previous works have been done in efforts to accelerate the computation of the forward algorithm by the massive parallelization of the workload. In this paper, we bring advanced GPU implementations with various optimizations, such as efficient host-device communication, task parallelization, pipelining, and memory management, to tackle this challenging task. Our design has shown a speedup of 783X comparing to the Java baseline on Intel single-core CPU, 31.88X to the C++ baseline on IBM Power8 multicore CPU, and 1.53X - 2.21X to the previous state-of-the-art GPU implementations over various genomics datasets.
Enliang Li, Subho S. Banerjee, Sitao Huang, Ravishankar K. Iyer, Deming Chen
ICCD3
2021 Chimera: A Hybrid Machine Learning-Driven Multi-Objective Design Space Exploration Tool for FPGA High-Level Synthesis
Mang Yu, Sitao Huang, Deming Chen
IDEAL2
2021 Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture
abstract
Graph Convolutional Networks (GCNs) are increasingly adopted in large-scale graph-based recommender systems. Training GCN requires the minibatch generator traversing graphs and sampling the sparsely located neighboring nodes to obtain their features. Since real-world graphs often exceed the capacity of GPU memory, current GCN training systems keep the feature table in host memory and rely on the CPU to collect sparse features before sending them to the GPUs. This approach, however, puts tremendous pressure on host memory bandwidth and the CPU. This is because the CPU needs to (1) read sparse features from memory, (2) write features into memory as a dense format, and (3) transfer the features from memory to the GPUs. In this work, we propose a novel GPU-oriented data communication approach for GCN training, where GPU threads directly access sparse features in host memory through zero-copy accesses without much CPU help. By removing the CPU gathering stage, our method significantly reduces the consumption of the host resources and data access latency. We further present two important techniques to achieve high host memory access efficiency by the GPU: (1) automatic data access address alignment to maximize PCIe packet efficiency, and (2) asynchronous zero-copy access and kernel execution to fully overlap data transfer with training. We incorporate our method into PyTorch and evaluate its effectiveness using several graphs with sizes up to 111 million nodes and 1.6 billion edges. In a multi-GPU training setup, our method is 65--92% faster than the conventional data transfer method, and can even match the performance of all-in-GPU-memory training for some graphs that fit in GPU memory.
Seungwon Min, Kun Wu 0002, Sitao Huang, Mert Hidayetoglu, Jinjun Xiong, Eiman Ebrahimi, Deming Chen, Wen-Mei W. Hwu
Proc. VLDB Endow.3
2021 PyLog: An Algorithm-Centric Python-Based FPGA Programming and Synthesis Flow
abstract
The exploding complexity and computation efficiency requirements of applications are stimulating a strong demand for hardware acceleration with heterogeneous platforms such as FPGAs. However, a high-quality FPGA design is very hard to create as it requires FPGA expertise and a long design iteration time. In contrast, software applications are typically developed in a short development cycle, with high-level languages like Python, which is at a much higher level of abstraction than all existing hardware design flows. To close this gap and simplify FPGA programming, we create PyLog, a high-level, algorithm-centric programming and synthesis flow for FPGA. PyLog features a set of compiler optimization passes and a type inference system to generate high-quality design. It abstracts away the implementation details, and allows designers to focus on algorithm specification. PyLog takes in Python functions and generates complete optimized FPGA system design. PyLog also has a runtime that allows users to run the PyLog code directly on the target FPGA platform without any extra code development. The whole design flow is automated. The evaluation shows that PyLog significantly improves FPGA design productivity and generates highly efficient FPGA designs that outperform highly optimized CPU and FPGA versions by 3.17x and 1.24x on average.
Sitao Huang, Kun Wu 0002, Hyunmin Jeong, Chengyue Wang 0002, Deming Chen, Wen-Mei W. Hwu
IEEE Trans. Computers1
2019 Automatic Generation of Warp-Level Primitives and Atomic Instructions for Fast and Portable Parallel Reduction on GPUs
abstract
Since the advent of GPU computing, GPU hardware has evolved at a fast pace. Since application performance heavily depends on the latest hardware improvements, performance portability is extremely challenging for GPU application library developers. Portability becomes even more difficult when new low-level instructions are added to the ISA (e.g., warp shuffle instructions) or the microarchitectural support for existing instructions is improved (e.g., atomic instructions). Library developers, besides re-tuning the code for new hardware features, deal with the performance portability issue by hand-writing multiple algorithm versions that leverage different instruction sets and microarchitectures. High-level programming frameworks and Domain Specific Languages (DSLs) do not typically support lowlevel instructions (e.g., warp shuffle and atomic instructions), so it is painful or even impossible for these programming systems to take advantage of the latest architectural improvements. In this work, we design a new set of high-level APIs and qualifiers, as well as specialized Abstract Syntax Tree (AST) transformations for high-level programming languages and DSLs. Our transformations enable warp shuffle instructions and atomic instructions (on global and shared memories) to be easily generated. We show a practical implementation of these transformations by building on Tangram, a high-level kernel synthesis framework. Using our new language and compiler extensions, we implement parallel reduction, a fundamental building block used in a wide range of algorithms. Parallel reduction is representative of the performance portability challenge, as its performance heavily depends on the latest hardware improvements. We compare our synthesized parallel reduction to another high-level programming framework and a hand-written high-performance library across three generations of GPU architectures, and show up to 7.8× speedup (2× on average) over hand-written code.
Simon Garcia de Gonzalo, Sitao Huang, Juan Gómez-Luna, Simon D. Hammond, Onur Mutlu, Wen-Mei W. Hwu
CGO2
2019 FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge
abstract
While embedded FPGAs are attractive platforms for DNN acceleration on edge-devices due to their low latency and high energy efficiency, the scarcity of resources of edge-scale FPGA devices also makes it challenging for DNN deployment. In this paper, we propose a simultaneous FPGA/DNN co-design methodology with both bottom-up and top-down approaches: a bottom-up hardware-oriented DNN model search for high accuracy, and a top-down FPGA accelerator design considering DNN-specific characteristics. We also build an automatic co-design flow, including an Auto-DNN engine to perform hardware-oriented DNN model search, as well as an Auto-HLS engine to generate synthesizable C code of the FPGA accelerator for explored DNNs. We demonstrate our co-design approach on an object detection task using PYNQ-Z1 FPGA. Results show that our proposed DNN model and accelerator outperform the state-of-the-art FPGA designs in all aspects including Intersection-over-Union (IoU) (6.2% higher), frames per second (FPS) (2.48× higher), power consumption (40% lower), and energy efficiency (2.5× higher). Compared to GPU-based solutions, our designs deliver similar accuracy but consume far less energy.
Cong Hao, Xiaofan Zhang 0001, Sitao Huang, Jinjun Xiong, Kyle Rupnow, Wen-Mei W. Hwu, Deming Chen
DAC4
2019 Analysis and Optimization of I/O Cache Coherency Strategies for SoC-FPGA Device
abstract
Unlike traditional PCIe-based FPGA accelerators, heterogeneous SoC-FPGA devices provide tighter integrations between software running on CPUs and hardware accelerators. Modern heterogeneous SoC-FPGA platforms support multiple I/O cache coherence options between CPUs and FPGAs, but these options can have inadvertent effects on the achieved bandwidths depending on applications and data access patterns. To provide the most efficient communications between CPUs and accelerators, understanding the data transaction behaviors and selecting the right I/O cache coherence method is essential. In this paper, we use Xilinx Zynq UltraScale+ as the SoC platform to show how certain I/O cache coherence method can perform better or worse in different situations, ultimately affecting the overall accelerator performances as well. Based on our analysis, we further explore possible software and hardware modifications to improve the I/O performances with different I/O cache coherence options. With our proposed modifications, the overall performance of SoC design can be averagely improved by 20%.
Seungwon Min, Sitao Huang, Mohamed El-Hadedy 0001, Jinjun Xiong, Deming Chen, Wen-Mei W. Hwu
FPL2
2019 Analysis and Modeling of Collaborative Execution Strategies for Heterogeneous CPU-FPGA Architectures
abstract
Heterogeneous CPU-FPGA systems are evolving towards tighter integration between CPUs and FPGAs for improved performance and energy efficiency. At the same time, programmability is also improving with High Level Synthesis tools (e.g., OpenCL Software Development Kits), which allow programmers to express their designs with high-level programming languages, and avoid time-consuming and error-prone register-transfer level (RTL) programming. In the traditional loosely-coupled accelerator mode, FPGAs work as offload accelerators, where an entire kernel runs on the FPGA while the CPU thread waits for the result. However, tighter integration of the CPUs and the FPGAs enables the possibility of fine-grained collaborative execution, i.e., having both devices working concurrently on the same workload. Such collaborative execution makes better use of the overall system resources by employing both CPU threads and FPGA concurrency, thereby achieving higher performance. In this paper, we explore the potential of collaborative execution between CPUs and FPGAs using OpenCL High Level Synthesis. First, we compare various collaborative techniques (namely, data partitioning and task partitioning), and evaluate the tradeoffs between them. We observe that choosing the most suitable partitioning strategy can improve performance by up to 2x. Second, we study the impact of a common optimization technique, kernel duplication, in a collaborative CPU-FPGA context. We show that the general trend is that kernel duplication improves performance until the memory bandwidth saturates. Third, we provide new insights that application developers can use when designing CPU-FPGA collaborative applications to choose between different partitioning strategies. We find that different partitioning strategies pose different tradeoffs (e.g., task partitioning enables more kernel duplication, while data partitioning has lower communication overhead and better load balance), but they generally outperform execution on conventional CPU-FPGA systems where no collaborative execution strategies are used. Therefore, we advocate even more integration in future heterogeneous CPU-FPGA systems (e.g., OpenCL 2.0 features, such as fine-grained shared virtual memory).
Sitao Huang, Li-Wen Chang, Izzat El Hajj, Simon Garcia de Gonzalo, Juan Gómez-Luna, Sai Rahul Chalamalasetti, Mohamed El-Hadedy 0001, Dejan S. Milojicic, Onur Mutlu, Deming Chen, Wen-Mei W. Hwu
ICPE1
2018 Towards Neural Phrase-based Machine Translation
Po-Sen Huang, Chong Wang 0002, Sitao Huang, Dengyong Zhou, Li Deng 0001
ICLR (Poster)3
2017 Hardware Acceleration of the Pair-HMM Algorithm for DNA Variant Calling
Sitao Huang, Gowthami Jayashri Manikandan, Anand Ramachandran 0001, Kyle Rupnow, Wen-Mei W. Hwu, Deming Chen
FPGA1
2017 Collaborative Computing for Heterogeneous Integrated Systems
abstract
Computing systems today typically employ, in addition to powerful CPUs, various types of specialized devices such as Graphics Processing Units (GPUs) and Field-Programmable Gate Arrays (FPGAs). Such heterogeneous systems are evolving towards tighter integration of devices for improved performance and reduced energy consumption. Compared to traditional use of GPUs and FPGAs as offload accelerators, this tight integration enables close collaboration between processors which is important for better utilization of system resources and higher performance. Programming interfaces are also adapting rapidly to tightly integrated heterogeneous platforms by introducing features such as shared virtual memory, memory coherence, and system-wide atomics, making collaborative computing among different devices even more practical.
Li-Wen Chang, Juan Gómez-Luna, Izzat El Hajj, Sitao Huang, Deming Chen, Wen-Mei W. Hwu
ICPE4
2016 Acceleration of the Pair-HMM Algorithm for DNA Variant Calling
abstract
In this project, we propose an SoC solution to accelerate the Pair-HMM's forward algorithm which is the key performance bottleneck in the GATK's HaplotypeCaller tool for DNA variant calling. We develop two versions of the Pair-HMM accelerator: one using High Level Synthesis (HLS), and another ring-based manual RTL implementation. We investigate the performance of the manual RTL design and HLS design in terms of design flexibility and overall run-time. We achieve a significant speed-up of up to 19x through the HLS implementation and speed-up of up to 95x through the RTL implementation of the algorithm.
Gowthami Jayashri Manikandan, Sitao Huang, Kyle Rupnow, Wen-Mei W. Hwu, Deming Chen
FCCM2
2014 Accelerating frequent item counting with FPGA
abstract
Frequent item counting is one of the most important operations in time series data mining algorithms, and the space saving algorithm is a widely used approach to solving this problem. With the rapid rising of data input speeds, the most challenging problem in frequent item counting is to meet the requirement of wire-speed processing. In this paper, we propose a streaming oriented PE-ring framework on FPGA for counting frequent items. Compared with the best existing FPGA implementation, our basic PE-ring framework saves 50% lookup table resources cost and achieves the same throughput in a more scalable way. Furthermore, we adopt SIMD-like cascaded filter for further performance improvements, which outperforms the previous work by up to 3.24 times in some data distributions.
Sitao Huang, Lanjun Wang, Yu Wang 0002, Huazhong Yang
FPGA3
2013 Accelerating subsequence similarity search based on dynamic time warping distance with FPGA
abstract
Subsequence search, especially subsequence similarity search, is one of the most important subroutines in time series data mining algorithms, and there is increasing evidence that Dynamic Time Warping (DTW) is the best distance metric. However, in spite of the great effort in software speedup techniques, including early abandoning strategies, lower bound, indexing, computation-reuse, DTW still cost too much time for many applications, e.g. 80% of the total time. Since DTW is a 2-Dimension sequential dynamic search with quite high data dependency, it is hard to use parallel hardware to accelerate it. In this work, we propose a novel framework for FPGA based subsequence similarity search and a novel PE-ring structure for DTW calculation. This framework utilizes the data reusability of continuous DTW calculations to reduce the bandwidth and exploit the coarse-grain parallelism; meanwhile guarantees the accuracy with a two-phase precision reduction. The PE-ring supports on-line updating patterns of arbitrary lengths, and utilizes the hard-wired synchronization of FPGA to realize the fine-grained parallelism. It also achieves flexible parallelism degree to do performance-cost trade-off. The experimental results show that we can achieve several orders of magnitude speedup in accelerating subsequence similarity search compared with the best software and current GPU/FPGA implementations in different datasets.
Sitao Huang, Lanjun Wang, Yu Wang 0002, Huazhong Yang
FPGA2