Sai Qian Zhang

dblp:164/7945 · DBLP profile ↗
← Back
46ranked-venue papers
13as first author
29since 2021 · last 2026
0000-0002-4815-9235ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 7 since 2021Computer networks · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation
abstract
Zining Liu, Yunhai Hu, Tianhua Xia, BO Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zining Liu, Yunhai Hu, Tianhua Xia, B. O. Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang
ACL (1)7
2026 Segment Only Where You Look: Leveraging Human Gaze Behavior for Efficient Computer Vision Applications in Augmented Reality
abstract
Augmented reality (AR) comprises groundbreaking technologies that are reshaping the landscape of human interaction. Image segmentation, which divides a user front-scene frame into more manageable parts for analysis, is of paramount importance since this technique enables AR systems to extract digital information precisely from the real world by identifying and isolating specific objects in the user's surroundings. Despite its importance, the segmentation task imposes substantial computational demands and processing delays on AR devices, significantly degrading the user experience.
Tianhua Xia, Sai Qian Zhang
ASPLOS (2)3
2026 FLAME: A Framework Exploring Execution Strategies for Multi-Cycle Operations in CGRA
abstract
Effective mapping of dataflow graphs onto Coarse-Grained Reconfigurable Arrays necessitates compiler-architecture co-design, yet existing approaches frequently assume single-cycle operations despite real-world applications often involving multi-cycle operations that constrain achievable clock frequencies. To address this, we propose FLAME, a novel framework supporting three execution strategies (exclusive, distributed, inclusive) specifically designed for multi-cycle operations, with co-designed compiler and hardware support. Our evaluations demonstrate that FLAME not only surpasses prior methods in performance and but also enables flexible exploration of these operations. The framework achieves average speedups of 2.21× over baseline CGRA and 1.49× over prior state-of-the-art framework while highlighting the distinct characteristics of each strategy.
Jiajun Qin, Cheng Tan 0002, Ruihong Yin, Tianhua Xia, Sai Qian Zhang, Bei Yu 0001
DATE5
2026 ECHO: Efficient Head-Orientation-Guided Real-Time Sound Spatialization for Virtual Reality
Tianhua Xia, Sai Qian Zhang
ISCA3
2025 H4H: Hybrid Convolution-Transformer Architecture Search for NPU-CIM Heterogeneous Systems for AR/VR Applications
abstract
Low-latency and low-power edge AI is crucial for Augmented/Virtual Reality applications. Recent advances demonstrate that hybrid models, combining convolution layers (CNN) and transformers (ViT), often achieve a superior accuracy/performance tradeoff on various computer vision and machine learning (ML) tasks. However, hybrid ML models can present system challenges for latency and energy efficiency due to their diverse nature in dataflow and memory access patterns. In this work, we leverage architecture heterogeneity from Neural Processing Units (NPU) and Compute-In-Memory (CIM) and explore diverse execution schemas for efficient hybrid model executions. We introduce H4H-NAS, a two-stage Neural Architecture Search (NAS) framework to automate the design of hybrid CNN/ViT models for heterogeneous edge systems featuring both NPU and CIM. We propose a two-phase incremental supernet training in our NAS to resolve gradient conflicts between sampled subnets caused by different block types in a hybrid model search space. Our H4H-NAS approach is also powered by a performance estimator built with NPU performance results measured on real silicon, and CIM performance based on industry IPs. H4H-NAS searches hybrid CNN-ViT models with fine granularity and achieves significant (up to 1.34%) top-1 accuracy improvement on ImageNet-1k. Moreover, results from our algorithm/hardware co-design reveal up to 56.08% overall latency and 41.72% energy improvements by introducing heterogeneous computing over baseline solutions. Overall, our framework guides the design of hybrid network architectures and system architectures for NPU+CIM heterogeneous systems.
Yiwei Zhao 0001, Sai Qian Zhang, Syed Shakib Sarwar, Kleber Stangherlin, Jorge Gomez 0001, Jae-sun Seo, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li 0001
ASP-DAC3
2025 PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMs
abstract
Large language models (LLMs) have revolutionized natural language processing (NLP) domain by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in LLMs significantly contribute to inference latency and present unique challenges that have not been encountered previously. Addressing these challenges requires accelerators that combine efficiency, flexibility, and support for user-defined precision. Our analysis reveals that Coarse-Grained Reconfigurable Arrays (CGRAs) provide an effective solution, offering a balance of performance and flexibility tailored to domain-specific workloads.
Jiajun Qin, Tianhua Xia, Cheng Tan 0002, Jeff Zhang 0001, Sai Qian Zhang
ASPLOS (2)5
2025 Foveated Instance Segmentation
abstract
Instance segmentation is essential for augmented reality and virtual reality (AR/VR) as it enables precise object recognition and interaction, enhancing the integration of virtual and real-world elements for an immersive experience. However, the high computational overhead of segmentation limits its application on resource-constrained AR/VR devices, causing large processing latency and degrading user experience. In contrast to conventional scenarios, AR/VR users typically focus on only a few regions within their field of view before shifting perspective, allowing segmentation to be concentrated on gaze-specific areas. This insight drives the need for efficient segmentation methods that prioritize processing instance of interest, reducing computational load and enhancing real-time performanceIn this paper, we present a foveated instance segmentation (FovealSeg) framework that leverages real-time user gaze data to perform instance segmentation exclusively on instance of interest, resulting in substantial computational savings. Evaluation results show that FSNet achieves an IoU of 0.56 on ADE20K and 0.54 on LVIS, notably outperforming the baseline. The code is available at https://github.com/SAI-Lab-NYU/Foveated-Instance-Segmentation
Hongyi Zeng, Tianhua Xia, Sai Qian Zhang
CVPR6
2025 HAAN: A Holistic Approach for Accelerating Normalization Operations in Large Language Models
abstract
Large language models (LLMs) have revolutionized natural language processing (NLP) tasks by achieving state-of-the-art performance across a range of benchmarks. Central to the success of these models is the integration of sophisticated architectural components aimed at improving training stability, convergence speed, and generalization capabilities. Among these components, normalization operation, such as layer normalization (LayerNorm), emerges as a pivotal technique, offering substantial benefits to the overall model performance. However, previous studies have indicated that normalization operations can substantially elevate processing latency and energy usage. In this work, we adopt the principles of algorithm and hardware co-design, introducing a holistic normalization accelerating method named HAAN. The evaluation results demonstrate that HAAN can achieve significantly better hardware performance compared to state-of-the-art solutions.
Tianfan Peng, Tianhua Xia, Jiajun Qin, Sai Qian Zhang
DATE4
2025 DFT Gaze: Distilled and Fine-Tuned Gaze Estimation for Personalization on Tiny Devices
abstract
Real-time personalized gaze estimation on AR/VR devices requires both accuracy and efficiency, especially when adapting to individual users with limited personal data. This task is challenging due to low-latency requirements, the presence of dataset biases from dominant gaze directions, and risk of catastrophic forgetting during adaptation. We present Distilled and Fine-Tuned (DFT) Gaze, a lightweight model for personalized gaze estimation. Distilled from a larger teacher model, DFT Gaze reduces model size while retaining essential visual features through knowledge distillation, without relying on gaze-specific supervision. During fine-tuning, it integrates gaze-specific supervision with Adapters, reaching 281K parameters for efficient adaptation and online updates on edge devices. To mitigate dataset biases and reduce catastrophic forgetting, we introduce a clustering-based sampling that balances gaze distribution for better generalization and improves adaptation to individual gaze patterns, even with only 5 personal images. DFT Gaze outperforms state-of-the-art methods on the MPIIFaceGaze dataset for personalized gaze estimation. Despite having the smallest model size at 281K parameters, it maintains low gaze errors across other datasets, including MPIIGaze, OpenEDS2020, and AEA. At 10× smaller than its teacher model, DFT Gaze achieves fast inference, a low parameter count, and effective adaptation, making it well-suited for real-time applications in resource-constrained environments.
He-Yen Hsieh, Ziyun Li 0001, Sai Qian Zhang, Wei-Te Mark Ting, Kao-Den Chang, Barbara De Salvo, Chiao Liu, H. T. Kung 0001
ICIP3
2025 A3FR: Agile 3D Gaussian Splatting with Incremental Gaze Tracked Foveated Rendering in Virtual Reality
abstract
Virtual reality (VR) significantly transforms immersive digital interfaces, greatly enhancing education, professional practices, and entertainment by increasing user engagement and opening up new possibilities in various industries.Among its numerous applications, image rendering is crucial.Nevertheless, rendering methodologies like 3D Gaussian Splatting impose high computational demands, driven predominantly by user expectations for superior visual quality.This results in notable processing delays for realtime image rendering, which greatly affects the user experience.Additionally, VR devices such as head-mounted displays (HMDs) are intricately linked to human visual behavior, leveraging knowledge from perception and cognition to improve user experience.These insights have spurred the development of foveated rendering, a technique that dynamically adjusts rendering resolution based on the user's gaze direction.The resultant solution, known as gazetracked foveated rendering, significantly reduces the computational burden of the rendering process.Although gaze-tracked foveated rendering can reduce rendering costs, the computational overhead of the gaze tracking process itself can sometimes outweigh the rendering savings, leading to increased processing latency.To address this issue, we propose an efficient rendering framework called A3FR, designed to minimize the latency of gaze-tracked foveated rendering via the parallelization of gaze tracking and foveated rendering processes.For the rendering algorithm, we utilize 3D Gaussian Splatting, a state-ofthe-art neural rendering technique.Evaluation results demonstrate that A3FR can reduce end-to-end rendering latency by up to 2× while maintaining visual quality.
Shuo Xin, Sai Qian Zhang
ICS3
2025 Process Only Where You Look: Hardware and Algorithm Co-optimization for Efficient Gaze-Tracked Foveated Rendering in Virtual Reality
abstract
Virtual reality (VR) plays a crucial role in advancing immersive, interactive experiences that transform learning, work, and entertainment by enhancing user engagement and expanding possibilities across various fields.Image rendering is one of the most crucial application in VR, as it produces high-quality, realistic visuals that are vital for maintaining immersive user experiences and preventing visual discomfort or motion sickness.However, the cost of image rendering in VR environment is considerable, primarily due to the demands of high-quality visual experiences from users.This challenge is even greater in real-time applications, where maintaining low latency further increases the complexity of the rendering process.On the other hand, VR devices, such as head-mounted displays (HMDs), are intrinsically linked to human behavior, using insights from perception and cognition to enhance user experience.In this work, we aim to reduce the high computational costs of the rendering process in VR by leveraging natural human eye dynamics and focusing on processing only where you look (POLO).This involves co-optimizing AI algorithms with underlying hardware for greater efficiency.We introduce POLONet, an efficient multitask deep learning framework designed to track human eye movements with minimal latency.Integrated with the POLO accelerator as a plug-in for VR HMD SoCs, this approach significantly lowers image rendering costs, achieving up to a 3.9× reduction in end-to-end latency compared to the latest gaze tracking methods.
Wenxuan Liu 0006, Kenneth Chen, Qi Sun 0003, Sai Qian Zhang
ISCA5
2025 FrameVoting: A Robust and Fast Method of Using Gaze Estimations to Identify Objects of Interest
abstract
We introduce FrameVoting, a voting-based method for real-time, gaze-driven object identification. It is a training-free method that incurs small computation and low processing latency, making the method ideal for wearable devices. In FrameVoting, the Point of Gaze (PoG) in each frame is used to define a potential region of interest. Regions across multiple frames are compared using the Sum of Absolute Differences (SAD) as a similarity measure. Each frame votes for the region from each of the other frames that is most similar to the region in the current frame, and only the region receiving the most votes is considered as the user’s region of interest and sent to a classifier for inference. FrameVoting thus eliminates the need for frame-by-frame bounding box retrieval and object detection required by traditional methods, thereby reducing computation overhead and latency. The method is robust, as it eliminates the need for threshold-tuning to determine whether gaze estimations are focused on a specific object. Further, the method is efficient and fast, as inference is only performed on the most-voted region, and the SAD computation is highly parallelizable. Our experiments on the AEA Dataset demonstrate that FrameVoting reduces the frequency of inferences by 95.6% compared to frame-by-frame object detection, while still accurately identifying the user’s objects of interest in real-time at 30 fps on a Raspberry Pi 5.
Kao-Den Chang, He-Yen Hsieh, H. T. Kung 0001, Ziyun Li 0001, Sai Qian Zhang
ISCAS5
2025 Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
abstract
Running Large Language Models (LLMs) on edge devices is crucial for reducing latency, improving real-time processing, and enhancing privacy.By performing inference directly on the device, data does not need to be sent to the cloud, ensuring faster responses and reducing reliance on network connectivity.However, implementing LLMs on edge devices presents challenges, particularly with managing key-value (KV) caches, which plays a pivotal role in LLM serving.As the input text lengthens, the size of the KV cache increases linearly with the sequence length, leading to a significant memory footprint and data access costs.On the other hand, edge devices have limited memory and computational power, making it hard to store and efficiently access the large caches needed for LLM inference.To mitigate the substantial overhead caused by KV cache, we propose using embedded DRAM (eDRAM) as the primary storage for LLM serving in edge device, which offers higher storage density compared to SRAM.However, to ensure data integrity, eDRAM needs periodic refresh operations, which are power-intensive.To reduce eDRAM costs and improve overall system performance, we propose Kelle, a software-hardware co-design solution optimized for deploying LLMs on eDRAM-based edge systems.Combined with our fine-grained memory eviction, recomputation, and refresh control algorithms, the Kelle accelerator delivers a 3.9× speedup and 4.5× energy savings compared to existing baseline solutions.
Tianhua Xia, Sai Qian Zhang
MICRO2
2025 CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models
abstract
As Vision-Language Models (VLMs) are increasingly deployed in split-DNN configurations--with visual encoders (e.g., ResNet, ViT) operating on user devices and sending intermediate features to the cloud--there is a growing privacy risk from semantic information leakage. Existing approaches to reconstructing images from these intermediate features often result in blurry, semantically ambiguous images. To directly address semantic leakage, we propose CapRecover, a cross-modality inversion framework that recovers high-level semantic content, such as labels or captions, directly from intermediate features without image reconstruction. We evaluate CapRecover on multiple datasets and victim models, demonstrating strong performance in semantic recovery. Specifically, CapRecover achieves up to 92.71% Top-1 label accuracy on CIFAR-10 and generates fluent captions from ResNet50 features on COCO2017 with ROUGE-L scores up to 0.52. Our analysis further reveals that deeper convolutional layers encode significantly more semantic information compared to shallow layers. To mitigate semantic leakage, we introduce a simple yet effective protection method: adding random noise to intermediate features at each layer and removing the noise in the next layer. Experimental results show that this approach prevents semantic leakage without additional training costs. Our code is available at https://jus1mple.github.io/Image2CaptionAttack.
Kedong Xiu, Sai Qian Zhang
ACM Multimedia2
2025 DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
abstract
Speculative decoding (SD) has emerged as a powerful method for accelerating autoregressive generation in large language models (LLMs), yet its integration into vision-language models (VLMs) remains underexplored. We introduce DREAM, a novel speculative decoding framework tailored for VLMs that combines three key innovations: (1) a cross-attention-based mechanism to inject intermediate features from the target model into the draft model for improved alignment, (2) adaptive intermediate feature selection based on attention entropy to guide efficient draft model training, and (3) visual token compression to reduce draft model latency. DREAM enables efficient, accurate, and parallel multimodal decoding with significant throughput improvement. Experiments across a diverse set of recent popular VLMs, including LLaVA, Pixtral, SmolVLM and Gemma3, demonstrate up to 3.6x speedup over conventional decoding and significantly outperform prior SD baselines in both inference throughput and speculative draft acceptance length across a broad range of multimodal benchmarks.
Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang
NeurIPS9
2025 QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
abstract
Vision-Language Models (VLMs) are integral to tasks such as image captioning and visual question answering, but their high computational cost, driven by large memory footprints and processing time, limits their scalability and real-time applicability. In this work, we propose leveraging Singular-Value Decomposition (SVD) over the joint query (Q), key (K), and value (V) weight matrices to reduce KV cache size and computational overhead. We in addition introduce an efficient rank allocation strategy that dynamically adjusts the SVD rank based on its impact on VLM accuracy, achieving a significant reduction in both memory usage and computational cost. Finally, we extend this approach by applying quantization to both VLM weights and activations, resulting in a highly efficient VLM. Our method outperforms previous approaches that rely solely on quantization or SVD by achieving more than $10$% accuracy improvement while consuming less hardware cost, making it better for real-time deployment on resource-constrained devices. We open source our code at https://github.com/SAI-Lab-NYU/QSVD.
Sai Qian Zhang
NeurIPS3
2025 ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
abstract
Photorealistic Codec Avatars (PCA), which generate high-fidelity human face renderings, are increasingly being used in Virtual Reality (VR) environments to enable immersive communication and interaction through deep learning–based generative models. However, these models impose significant computational demands, making real-time inference challenging on resource-constrained VR devices such as head-mounted displays (HMDs), where latency and power efficiency are critical. To address this challenge, we propose an efficient post-training quantization (PTQ) method tailored for Codec Avatar models, enabling low-precision execution without compromising output quality. In addition, we design a custom hardware accelerator that can be integrated into the system-on-chip (SoC) of VR devices to further enhance processing efficiency. Building on these components, we introduce ESCA, a full-stack optimization framework that accelerates PCA inference on edge VR platforms. Experimental results demonstrate that ESCA boosts FovVideoVDP quality scores by up to +0.39 over the best 4-bit baseline, delivers up to 3.36× latency reduction, and sustains a rendering rate of 100 frames per second in end-to-end tests, satisfying real-time VR requirements. These results demonstrate the feasibility of deploying high-fidelity codec avatars on resource-constrained devices, opening the door to more immersive and portable VR experiences.
Mingzhi Zhu, Ding Shang, Sai Qian Zhang
NeurIPS3
2025 DiTAS: Quantizing Diffusion Transformers via Enhanced Activation Smoothing
abstract
Diffusion Transformers (DiTs) have recently attracted significant interest from both industry and academia due to their enhanced capabilities in visual generation, surpassing the performance of traditional diffusion models that employ U-Net. However, the improved performance of DiTs comes at the expense of higher parameter counts and implementation costs, which significantly limits their deployment on resource-constrained devices like mobile phones. We propose DiTAS, a data-free post-training quantization (PTQ) method for efficient DiT inference. DiTAS relies on the proposed temporal-aggregated smoothing techniques to mitigate the impact of the channel-wise outliers within the input activations, leading to much lower quantization error under extremely low bitwidth. To further enhance the performance of the quantized DiT, we adopt the layer-wise grid search strategy to optimize the smoothing factor. Moreover, we integrate a training-free LoRA module for weight quantization, leveraging alternating optimization to minimize quantization errors without additional fine-tuning. Experimental results demonstrate that our approach enables 4-bit weight, 8-bit activation (W4A8) quantization for DiTs while maintaining comparable performance as the full-precision model. Code is available at https://github.com/DZY122/DiTAS
Zhenyuan Dong, Sai Qian Zhang
WACV2
2025 FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Efficient Foveated Rendering in Virtual Reality
abstract
Leveraging real-time eye tracking, foveated rendering optimizes hardware efficiency and enhances visual quality virtual reality (VR). This approach leverages eye-tracking techniques to determine where the user is looking, allowing the system to render high-resolution graphics only in the foveal region-the small area of the retina where visual acuity is highest, while the peripheral view is rendered at lower resolution. However, modern deep learning-based gaze-tracking solutions often exhibit a long-tail distribution of tracking errors, which can degrade user experience and reduce the benefits of foveated rendering by causing misalignment and decreased visual quality. This paper introduces FovealNet, an advanced AI-driven gaze tracking framework designed to optimize system performance by strategically enhancing gaze tracking accuracy. To further reduce the implementation cost of the gaze tracking algorithm, FovealNet employs an event-based cropping method that eliminates over 64.8% of irrelevant pixels from the input image. Additionally, it incorporates a simple yet effective token-pruning strategy that dynamically removes tokens on the fly without compromising tracking accuracy. Finally, to support different runtime rendering configurations, we propose a system performance-aware multi-resolution training strategy, allowing the gaze tracking DNN to adapt and optimize overall system performance more effectively. Evaluation results demonstrate that FovealNet achieves at least 1.42× speed up compared to previous methods and 13% increase in perceptual quality for foveated output. The code is available at https://github.com/wl3181/FovealNet.
Wenxuan Liu 0006, Budmonde Duinkharjav, Qi Sun 0003, Sai Qian Zhang
IEEE Trans. Vis. Comput. Graph.4
2024 CAMEL: Co-Designing AI Models and eDRAMs for Efficient On-Device Learning
abstract
On-device learning allows AI models to adapt to user data, thereby enhancing service quality on edge platforms. However, training AI on resource-limited devices poses significant challenges due to the demanding computing workload and the substantial memory consumption and data access required by deep neural networks (DNNs). To address these issues, we propose utilizing embedded dynamic random-access memory (eDRAM) as the primary storage medium for transient training data. In comparison to static random-access memory (SRAM), eDRAM provides higher storage density and lower leakage power, resulting in reduced access cost and power leakage. Nevertheless, to maintain the integrity of the stored data, periodic power-hungry refresh operations could potentially degrade system performance. To minimize the occurrence of expensive eDRAM refresh operations, it is beneficial to shorten the lifetime of stored data during the training process. To achieve this, we adopt the principles of algorithm and hardware co-design, introducing a family of reversible DNN architectures that effectively decrease data lifetime and storage costs throughout training. Additionally, we present a highly efficient on-device training engine named CAMEL, which leverages eDRAM as the primary on-chip memory. This engine enables efficient on-device training with significantly reduced memory usage and off-chip DRAM traffic while maintaining superior training accuracy. We evaluate our CAMEL system on multiple DNNs with different datasets, demonstrating a 2.5× speedup of the training process and 2.8× training energy savings than the other baseline hardware platforms.
Sai Qian Zhang, Thierry Tambe, Nestor Cuevas, Gu-Yeon Wei, David Brooks 0001
HPCA1
2024 Murmuration: On-the-fly DNN Adaptation for SLO-Aware Distributed Inference in Dynamic Edge Environments
abstract
The proliferation of Virtual and Augmented Reality (VR/AR) and the Internet of Things (IoT) applications is driving the demand for efficient Deep Neural Network (DNN) inference at the edge. These applications often impose stringent Service Level Objectives (SLOs), such as latency or accuracy, that must be met under the constraints of limited resources and dynamic network conditions. In this study, we explore a novel approach to DNN inference across multiple edge devices, incorporating both model customization and partitioning dynamically, to better align with these constraints and SLOs. Unlike conventional methods that employ a single fixed DNN network, our system, termed Murmuration, combines one-shot Neural Architecture Search (NAS) and Reinforcement Learning (RL) to dynamically customize and partition DNN models. This approach adapts in real-time to the capabilities of the edge devices, network conditions, and varying SLO requirements. The design of Murmuration allows it to effectively navigate the large search space defined by DNN models, network delays, and bandwidth, offering a significant improvement in managing trade-offs between accuracy and latency.
Jieyu Lin, Minghao Li 0009, Sai Qian Zhang, Alberto Leon-Garcia
ICPP3
2024 Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and Inference
abstract
The attention mechanism is a pivotal element within the transformer architecture, making a substantial contribution to its exceptional performance. Within this attention mechanism, Softmax is an imperative component that enables the model to assess the degree of correlation between various segments of the input. Yet, prior research has shown that Softmax operations can significantly increase processing latency and energy consumption in the transformer network due to their internal nonlinear operations and data dependencies. In this work, we proposed Hyft, a hardware efficient floating point Softmax accelerator for both training and inference. Hyft aims to reduce the implementation cost of different nonlinear arithmetic operations within softmax by adaptively converting intermediate results into the most suitable numeric format for each specific operation, leading to reconfigurable accelerator with hybrid numeric format. The evaluation results highlight that Hyft achieves a remarkable 10X reduction in hardware resource utilization and a 6x reduction in processing latency, all while maintaining a negligible impact on transformer accuracy.
Tianhua Xia, Sai Qian Zhang
ISLPED2
2024 JointNF: Enhancing DNN Performance through Adaptive N: M Pruning across both Weight and Activation
abstract
Balancing accuracy and hardware efficiency remains a challenge with traditional pruning methods. N:M sparsity is a recent approach offering a compromise, allowing up to N non-zero weights in a group of M consecutive weights. However, N:M pruning enforces a uniform sparsity level of N/M across all layers, which does not align well sparse nature of deep neural networks (DNNs). To achieve a more flexible sparsity pattern and a higher overall sparsity level, we present JointNF, a novel joint N:M and structured pruning algorithm to enable fine-grained structured pruning with adaptive sparsity levels across the DNN layers. Moreover, we show for the first time that N:M pruning can also be applied over the input activation for further performance enhancement.
Sai Qian Zhang, Thierry Tambe, Gu-Yeon Wei, David Brooks 0001
ISLPED1
2024 Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical Framework
abstract
Augmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration.
Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong
ACM Trans. Design Autom. Electr. Syst.3
2022 A Multi-Agent Reinforcement Learning Approach for Efficient Client Selection in Federated Learning
abstract
Federated learning (FL) is a training technique that enables client devices to jointly learn a shared model by aggregating locally computed models without exposing their raw data. While most of the existing work focuses on improving the FL model accuracy, in this paper, we focus on the improving the training efficiency, which is often a hurdle for adopting FL in real world applications. Specifically, we design an efficient FL framework which jointly optimizes model accuracy, processing latency and communication efficiency, all of which are primary design considerations for real implementation of FL. Inspired by the recent success of Multi Agent Reinforcement Learning (MARL) in solving complex control problems, we present FedMarl, a federated learning framework that relies on trained MARL agents to perform efficient run-time client selection. Experiments show that FedMarl can significantly improve model accuracy with much lower processing latency and communication cost.
Sai Qian Zhang, Jieyu Lin, Qi Zhang 0008
AAAI1
2022 SphereFed: Hyperspherical Federated Learning
Xin Dong 0009, Sai Qian Zhang, H. T. Kung 0001
ECCV (26)2
2022 FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic Rounding
abstract
Block Floating Point (BFP) can efficiently support quantization for Deep Neural Network (DNN) training by providing a wide dynamic range via a shared exponent across a group of values. In this paper, we propose a Fast First, Accurate Second Training (FAST) system for DNNs, where the weights, activations, and gradients are represented in BFP. FAST supports matrix multiplication with variable precision BFP input operands, enabling incremental increases in DNN precision throughout training. By increasing the BFP precision across both training iterations and DNN layers, FAST can greatly shorten the training time while reducing overall hardware resource usage. Our FAST Multipler-Accumulator (fMAC) supports dot product computations under multiple BFP precisions. We validate our FAST system on multiple DNNs with different datasets, demonstrating a 2-6× speedup in training on a single-chip platform over prior work based on mixed-precision or block floating point number systems while achieving similar performance in validation accuracy.
Sai Qian Zhang, Bradley McDanel, H. T. Kung 0001
HPCA1
2021 Training for multi-resolution inference using reusable quantization terms
abstract
Low-resolution uniform quantization (e.g., 4-bit bitwidth) for both Deep Neural Network (DNN) weights and data has emerged as an important technique for efficient inference. Departing from conventional quantization, we describe a novel training approach to support inference at multiple resolutions by reusing a single set of quantization terms (the same set of nonzero bits in values). The proposed approach streamlines the training and supports dynamic selection of resolution levels during inference. We evaluate the method on a diverse range of applications including multiple CNNs on ImageNet, an LSTM on Wikitext-2, and YOLO-v5 on COCO. We show that models resulting from our multi-resolution training can support up to 10 resolutions with only a moderate performance reduction (e.g., ≤ 1%) compared to training them individually. Lastly, using an FPGA, we compare our multi-resolution multiplier-accumulator (mMAC) against other conventional MAC designs and evaluate the inference performance. We show that the mMAC design broadens the choices in trading off cost, efficiency, and latency across a range of computational budgets.
Sai Qian Zhang, Bradley McDanel, H. T. Kung 0001, Xin Dong 0009
ASPLOS1
2021 Saturation RRAM Leveraging Bit-Level Sparsity Resulting from Term Quantization
abstract
The proposed saturation RRAM for in-memory computing of a pre-trained Convolutional Neural Network (CNN) inference imposes a limit on the maximum analog value output from each bitline in order to reduce analog-to-digital (A/D) conversion costs. The proposed scheme uses term quantization (TQ) to enable flexible bit annihilation at any position for a value in the context of a group of weights values in RRAM. This enables a drastic reduction in the required ADC resolution while still maintaining CNN model accuracy. Specifically, we show that the A/D conversion errors after TQ have a minimum impact on the classification accuracy of the inference task. For instance, for a 64×64 RRAM, reducing the ADC resolution from 6 bits to 4 bits enables a 1.58× reduction in the total system power, without a significant impact to classification accuracy.
Bradley McDanel, Sai Qian Zhang, H. T. Kung 0001
ISCAS2
2020 RTN: Reparameterized Ternary Network
abstract
To deploy deep neural networks on resource-limited devices, quantization has been widely explored. In this work, we study the extremely low-bit networks which have tremendous speed-up, memory saving with quantized activation and weights. We first bring up three omitted issues in extremely low-bit networks: the squashing range of quantized values; the gradient vanishing during backpropagation and the unexploited hardware acceleration of ternary networks. By reparameterizing quantized activation and weights vector with full precision scale and offset for fixed ternary vector, we decouple the range and magnitude from direction to extenuate above problems. Learnable scale and offset can automatically adjust the range of quantized values and sparsity without gradient vanishing. A novel encoding and computation pattern are designed to support efficient computing for our reparameterized ternary network (RTN). Experiments on ResNet-18 for ImageNet demonstrate that the proposed RTN finds a much better efficiency between bitwidth and accuracy and achieves up to 26.76% relative accuracy improvement compared with state-of-the-art methods. Moreover, we validate the proposed computation pattern on Field Programmable Gate Arrays (FPGA), and it brings 46.46 × and 89.17 × savings on power and area compared with the full precision convolution.
Yuhang Li 0001, Xin Dong 0009, Sai Qian Zhang, Haoli Bai, Yuanpeng Chen, Wei Wang 0059
AAAI3
2020 Adaptive Distributed Convolutional Neural Network Inference at the Network Edge with ADCNN
abstract
The emergence of the Internet of Things (IoT) has led to a remarkable increase in the volume of data generated at the network edge. In order to support real-time smart IoT applications, massive amounts of data generated from edge devices need to be processed using methods such as deep neural networks (DNNs) with low latency. To improve application performance and minimize resource cost, enterprises have begun to adopt Edge computing, a computation paradigm that advocates processing input data locally at the network edge. However, as edge nodes are often resource-constrained, running data-intensive DNN inference tasks on each individual edge node often incurs high latency, which seriously limits the practicality and effectiveness of this model.
Sai Qian Zhang, Jieyu Lin, Qi Zhang 0008
ICPP1
2020 Succinct and Robust Multi-Agent Communication With Temporal Message Control
abstract
Recent studies have shown that introducing communication between agents can significantly improve overall performance in cooperative Multi-agent reinforcement learning (MARL). However, existing communication schemes often require agents to exchange an excessive number of messages at run-time under a reliable communication channel, which hinders its practicality in many real-world situations. In this paper, we present \textit{Temporal Message Control} (TMC), a simple yet effective approach for achieving succinct and robust communication in MARL. TMC applies a temporal smoothing technique to drastically reduce the amount of information exchanged between agents. Experiments show that TMC can significantly reduce inter-agent communication overhead without impacting accuracy. Furthermore, TMC demonstrates much better robustness against transmission loss than existing approaches in lossy networking environments.
Sai Qian Zhang, Qi Zhang 0008, Jieyu Lin
NeurIPS1
2020 Term quantization: furthering quantization at run time
abstract
We present a novel technique, called Term Quantization (TQ), for furthering quantization at run time for improved computational efficiency of deep neural networks (DNNs) already quantized with conventional quantization methods. TQ operates on power-of-two terms in expressions of values. In computing a dot-product computation, TQ dynamically selects a fixed number of largest terms to use from values of the two vectors. By exploiting weight and data distributions typically present in DNNs, TQ has a minimal impact on DNN model performance (e.g., accuracy or perplexity). We use TQ to facilitate tightly synchronized processor arrays, such as systolic arrays, for efficient parallel processing. We evaluate TQ on an MLP for MNIST, multiple CNNs for ImageNet and an LSTM for Wikitext-2. We demonstrate significant reductions in inference computation costs (between 3-10×) compared to conventional uniform quantization for the same level of model performance.
H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang
SC3
2019 Maestro: A Memory-on-Logic Architecture for Coordinated Parallel Use of Many Systolic Arrays
abstract
We present the Maestro memory-on-logic 3D-IC architecture for coordinated parallel use of a plurality of systolic arrays (SAs) in performing deep neural network (DNN) inference. Maestro reduces under-utilization common for a single large SA by allowing parallel use of many smaller SAs on DNN weight matrices of varying shapes and sizes. In order to buffer immediate results in memory blocks (MBs) and provide coordinated high-bandwidth communication between SAs and MBs in transferring weights and results Maestro employs three innovations. (1) An SA on the logic die can access its corresponding MB on the memory die in short distance using 3D-IC interconnects, (2) through an efficient switch based on H-trees, an SA can access any MB with low latency, and (3) the switch can combine partial results from SAs in an elementwise fashion before writing back to a destination MB. We describe the Maestro architecture, including a circuit and layout design, detail scheduling of the switch, analyze system performance for real-time inference applications using input with batch size equal to one, and showcase applications for deep learning inference, with ShiftNet for computer vision and recent Transformer models for natural language processing. For the same total number of systolic cells, Maestro, with multiple smaller SAs, leads to 16x and 12x latency improvements over a single large SA on ShiftNet and Transformer, respectively. Compared to a floating-point GPU implementation of ShiftNet and Transform, a baseline Maestro system with 4,096 SAs (each with 8x8 systolic cells) provides significant latency improvements of 30x and 47x, respectively.
H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang, Xin Dong 0009, Chih-Chiang Chen
ASAP3
2019 Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
abstract
This paper describes a novel approach of packing sparse convolutional neural networks into a denser format for efficient implementations using systolic arrays. By combining multiple sparse columns of a convolutional filter matrix into a single dense column stored in the systolic array, the utilization efficiency of the systolic array can be substantially increased (e.g., 8x) due to the increased density of nonzero weights in the resulting packed filter matrix. In combining columns, for each row, all filter weights but the one with the largest magnitude are pruned. The remaining weights are retrained to preserve high accuracy. We study the effectiveness of this joint optimization for both high utilization efficiency and classification accuracy with ASIC and FPGA designs based on efficient bit-serial implementations of multiplier-accumulators. We demonstrate that in mitigating data privacy concerns the retraining can be accomplished with only fractions of the original dataset (e.g., 10% for CIFAR-10). We present analysis and empirical evidence on the superior performance of our column combining approach against prior arts under metrics such as energy efficiency (3x) and inference latency (12x).
H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang
ASPLOS3
2019 Full-stack optimization for accelerating CNNs using powers-of-two weights with FPGA validation
abstract
We present a full-stack optimization framework for accelerating inference of CNNs (Convolutional Neural Networks) and validate the approach with a field-programmable gate array (FPGA) implementation. By jointly optimizing CNN models, computing architectures, and hardware implementations, our full-stack approach achieves unprecedented performance in the trade-off space characterized by inference latency, energy efficiency, hardware utilization, and inference accuracy. An FPGA implementation is used as the validation vehicle for our design, achieving a 2.28ms inference latency for the ImageNet benchmark. Our implementation shines in that it has 9x higher energy efficiency compared to other implementations while achieving comparable latency. A highlight of our approach which contributes to the achieved high energy efficiency is an efficient Selector-Accumulator (SAC) architecture for implementing CNNs with powers-of-two weights. Compared to an FPGA implementation for a traditional 8-bit MAC, SAC substantially reduces required hardware resources (4.85x fewer lookup tables) and power consumption (2.48x).
Bradley McDanel, Sai Qian Zhang, H. T. Kung 0001, Xin Dong 0009
ICS2
2019 Systolic Building Block for Logic-on-Logic 3D-IC Implementations of Convolutional Neural Networks
abstract
We present a building block architecture for systolic array 3D-IC implementations of convolutional neural network (CNN) inference. The building block can be part of a library offered by a chip design service provider to support efficient CNN implementations. We describe how the building block can form systolic arrays for implementing low-latency, energy-efficient CNN inference for models of any size, while incorporating advanced packaging features such as “logic-on-logic” 3D-IC (micro-bump/TSV, monolithic 3D or other 3D technology). We present delay and power analysis for 2D and 3D implementations, and argue that as systolic arrays scale in size, 3D implementations based on, e.g., micro-bump/TSV, lead to significant performance improvements over 2D implementations.
H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang, C. T. Wang, Jin Cai, Victor C. Y. Chang, M. F. Chen, Jack Yuan-Chen Sun, Douglas Yu
ISCAS3
2019 Efficient Communication in Multi-Agent Reinforcement Learning via Variance Based Control
abstract
Multi-agent reinforcement learning (MARL) has recently received considerable attention due to its applicability to a wide range of real-world applications. However, achieving efficient communication among agents has always been an overarching problem in MARL. In this work, we propose Variance Based Control (VBC), a simple yet efficient technique to improve communication efficiency in MARL. By limiting the variance of the exchanged messages between agents during the training phase, the noisy component in the messages can be eliminated effectively, while the useful part can be preserved and utilized by the agents for better performance. Our evaluation using multiple MARL benchmarks indicates that our method achieves $2-10\times$ lower in communication overhead than state-of-the-art MARL algorithms, while allowing agents to achieve better overall performance.
Sai Qian Zhang, Qi Zhang 0008, Jieyu Lin
NeurIPS1
2018 Adaptive Tiling: Applying Fixed-size Systolic Arrays To Sparse Convolutional Neural Networks
abstract
We introduce adaptive tiling, a method of partitioning layers in a sparse convolutional neural network (CNN) into blocks of filters and channels, called tiles, each implementable with a fixed-size systolic array. By allowing a tile to adapt its size so that it can cover a large sparse area, we minimize the total number of tiles, or equivalently, the number of systolic array calls required to perform CNN inference. The proposed scheme resolves a challenge of applying systolic array architectures, traditionally designed for dense matrices, to sparse CNNs. To validate the approach, we construct a highly sparse Lasso-Mobile network by pruning MobileNet trained with an l1 regularization penalty, and demonstrate that adaptive tiling can lead to a 2- 3x reduction in systolic array calls, on Lasso-Mobile, for several benchmark datasets.
H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang
ICPR3
2017 TCAM space-efficient routing in a software defined network
Sai Qian Zhang, Qi Zhang 0008, Ali Tizghadam, Byungchul Park, Hadi Bannazadeh, Raouf Boutaba, Alberto Leon-Garcia
Comput. Networks1
2016 Joint NFV placement and routing for multicast service on SDN
abstract
Network function visualization (NFV) has emerged as a promising paradigm in networking, where the hardware-based middleboxes are replaced with software-based virtualized entities typically running on the cloud to provide specific functionalities. By deploying NFV, network services become more adaptive and cost-effective. Many multicast services such as real-time multimedia streaming and intrusion detection require appropriate services chaining; however, NFVs placement in the network as well as traffic routing strategy to guarantee that the multicast flows traverse through the services chain before reaching the end user is still an open problem. In this paper, we present an algorithm to solve this problem.
Sai Qian Zhang, Ali Tizghadam, Byungchul Park, Hadi Bannazadeh, Alberto Leon-Garcia
NOMS1
2015 Network Function Virtualization enabled multicast routing on SDN
abstract
Many multicast services such as live multimedia distribution and real-time event monitoring require constructing a multicast mechanism that involves network functions (e.g. firewall, video transcoding). Network Function Virtualization (NFV) is a concept that proposes using virtualization to implement network functions on infrastructure building block (such as high volume servers, virtual machines), where software provides the functionality of existing purpose-built network equipment. We present an approach for building the multicast mechanism whereby multicast flows are processed by NFV before reaching their end users. We propose a routing algorithm and a method for building an appropriate multicast topology.
Sai Qian Zhang, Qi Zhang 0008, Hadi Bannazadeh, Alberto Leon-Garcia
ICC1
2015 Aurora: Adaptive Block Replication in Distributed File Systems
abstract
Distributed file systems such as Google File System and Hadoop Distributed File System have been used to store large volumes of data in Cloud data centers. These systems divide data sets in blocks of fixed size and replicate them over multiple machines to achieve both reliability and efficiency. Recent studies have shown that data blocks tend to have a wide disparity in data popularity. In this context, the naive block replication schemes used by these systems often cause an uneven load distribution across machines, which reduces the overall I/O throughput of the system. While many replication algorithms have been proposed, existing solutions have not carefully studied the placement of data blocks that balances the load across machines, while ensuring node and rack-level reliability requirements are satisfied. In this paper, we study the dynamic data replication problem with the goal of balancing machine load while ensuring machine and rack-level reliability requirements are met. We propose several local search algorithms that provide constant approximation guarantees, yet simple and practical for implementation. We further present Aurora, a dynamic block placement mechanism that implements these algorithms in the Hadoop Distributed File System with minimal overhead. Through experiments using workload traces from Yahoo! and Facebook, we show Aurora reduces machine load imbalance by up to 26.9% compared to existing solutions, while satisfying node and rack-level reliability requirements.
Qi Zhang 0008, Sai Qian Zhang, Alberto Leon-Garcia, Raouf Boutaba
ICDCS2
2015 Fast Network Flow Resumption for Live Virtual Machine Migration on SDN
abstract
Virtual machine (VM) migration occurs very frequently in cloud computing. VM Migration enables a running OS, including memory and storage to move from one physical host to another physical host. A particular case of interest is live migration where the process of migrating the full state from one OS to the other should happen continuously and without any connection disruption. In order to have a seamless VM migration process the system has to be able to resume network connectivity very quickly. Fast resumption has proved to be a challenging problem. In this paper, we present a scheme to efficiently to migrate VMs networking resources using SDN techniques to achieve fast network flow resumption on SDN. We formulate the problem by an integer programming problem, we prove its NP-completeness and we propose a heuristic algorithm to solve the problem. Software simulation and real testbed implementation are done to demonstrate the performance of the flow migration scheme.
Sai Qian Zhang, Pouya Yasrebi, Ali Tizghadam, Hadi Bannazadeh, Alberto Leon-Garcia
ICNP1
2015 Kaleidoscope: Real-time content delivery in software defined infrastructures
abstract
Real-time content delivery services such as live media streaming, news casting and real-time event subscription/publication systems have become popular in recent years. Unlike traditional content delivery applications, real-time content delivery requires live content to be processed and delivered to end users in a timely and efficient manner. Furthermore, as both content producers and consumers may change over-time, it is a challenge to provision resources for these applications to achieve high service quality while minimizing total operational costs. Fortunately, the recent development of Cloud computing and Software Defined Networking (SDN) enables efficient implementation of real-time content delivery systems. Recently, the concept of Software Defined Infrastructure aims at combining Cloud computing and SDN to provide an unified framework for application deployment and management. In this paper, we present Kaleidoscope, an architecture for real-time content delivery in Software Defined Infrastructures. Kaleidoscope leverages network virtualization, SDN-based broadcasting and dynamic cloud resource provisioning to achieve high resource efficiency and service performance. Specifically, we present a resource management scheme that controls Cloud resource allocation and network configuration at run-time in accordance with service demand. Experiments show that Kaleidoscope is able to achieve lower resource cost while providing high service quality.
Qi Zhang 0008, Sai Qian Zhang, Jieyu Lin, Hadi Bannazadeh, Alberto Leon-Garcia
IM2
2015 Routing Algorithms for Network Function Virtualization Enabled Multicast Topology on SDN
abstract
Many multicast services such as live multimedia distribution and real-time event monitoring require multicast mechanisms that involve network functions (e.g., firewall and video transcoding). Network function virtualization (NFV) is a concept that proposes using virtualization to implement network functions on infrastructure building block (such as high volume servers and virtual machines), where software provides the functionality of existing purpose-built network equipment. We present an approach for building the multicast mechanism whereby multicast flows are processed by NFV before reaching their end users. We propose a routing algorithm and a method for building an appropriate multicast topology.
Sai Qian Zhang, Qi Zhang 0008, Hadi Bannazadeh, Alberto Leon-Garcia
IEEE Trans. Netw. Serv. Manag.1