Ziyun Li 0001

dblp:164/4195-1 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-6070-6310ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 H4H: Hybrid Convolution-Transformer Architecture Search for NPU-CIM Heterogeneous Systems for AR/VR Applications
abstract
Low-latency and low-power edge AI is crucial for Augmented/Virtual Reality applications. Recent advances demonstrate that hybrid models, combining convolution layers (CNN) and transformers (ViT), often achieve a superior accuracy/performance tradeoff on various computer vision and machine learning (ML) tasks. However, hybrid ML models can present system challenges for latency and energy efficiency due to their diverse nature in dataflow and memory access patterns. In this work, we leverage architecture heterogeneity from Neural Processing Units (NPU) and Compute-In-Memory (CIM) and explore diverse execution schemas for efficient hybrid model executions. We introduce H4H-NAS, a two-stage Neural Architecture Search (NAS) framework to automate the design of hybrid CNN/ViT models for heterogeneous edge systems featuring both NPU and CIM. We propose a two-phase incremental supernet training in our NAS to resolve gradient conflicts between sampled subnets caused by different block types in a hybrid model search space. Our H4H-NAS approach is also powered by a performance estimator built with NPU performance results measured on real silicon, and CIM performance based on industry IPs. H4H-NAS searches hybrid CNN-ViT models with fine granularity and achieves significant (up to 1.34%) top-1 accuracy improvement on ImageNet-1k. Moreover, results from our algorithm/hardware co-design reveal up to 56.08% overall latency and 41.72% energy improvements by introducing heterogeneous computing over baseline solutions. Overall, our framework guides the design of hybrid network architectures and system architectures for NPU+CIM heterogeneous systems.
Yiwei Zhao 0001, Sai Qian Zhang, Syed Shakib Sarwar, Kleber Stangherlin, Jorge Gomez 0001, Jae-sun Seo, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li 0001
ASP-DAC11
2025 DFT Gaze: Distilled and Fine-Tuned Gaze Estimation for Personalization on Tiny Devices
abstract
Real-time personalized gaze estimation on AR/VR devices requires both accuracy and efficiency, especially when adapting to individual users with limited personal data. This task is challenging due to low-latency requirements, the presence of dataset biases from dominant gaze directions, and risk of catastrophic forgetting during adaptation. We present Distilled and Fine-Tuned (DFT) Gaze, a lightweight model for personalized gaze estimation. Distilled from a larger teacher model, DFT Gaze reduces model size while retaining essential visual features through knowledge distillation, without relying on gaze-specific supervision. During fine-tuning, it integrates gaze-specific supervision with Adapters, reaching 281K parameters for efficient adaptation and online updates on edge devices. To mitigate dataset biases and reduce catastrophic forgetting, we introduce a clustering-based sampling that balances gaze distribution for better generalization and improves adaptation to individual gaze patterns, even with only 5 personal images. DFT Gaze outperforms state-of-the-art methods on the MPIIFaceGaze dataset for personalized gaze estimation. Despite having the smallest model size at 281K parameters, it maintains low gaze errors across other datasets, including MPIIGaze, OpenEDS2020, and AEA. At 10× smaller than its teacher model, DFT Gaze achieves fast inference, a low parameter count, and effective adaptation, making it well-suited for real-time applications in resource-constrained environments.
He-Yen Hsieh, Ziyun Li 0001, Sai Qian Zhang, Wei-Te Mark Ting, Kao-Den Chang, Barbara De Salvo, Chiao Liu, H. T. Kung 0001
ICIP2
2025 FrameVoting: A Robust and Fast Method of Using Gaze Estimations to Identify Objects of Interest
abstract
We introduce FrameVoting, a voting-based method for real-time, gaze-driven object identification. It is a training-free method that incurs small computation and low processing latency, making the method ideal for wearable devices. In FrameVoting, the Point of Gaze (PoG) in each frame is used to define a potential region of interest. Regions across multiple frames are compared using the Sum of Absolute Differences (SAD) as a similarity measure. Each frame votes for the region from each of the other frames that is most similar to the region in the current frame, and only the region receiving the most votes is considered as the user’s region of interest and sent to a classifier for inference. FrameVoting thus eliminates the need for frame-by-frame bounding box retrieval and object detection required by traditional methods, thereby reducing computation overhead and latency. The method is robust, as it eliminates the need for threshold-tuning to determine whether gaze estimations are focused on a specific object. Further, the method is efficient and fast, as inference is only performed on the most-voted region, and the SAD computation is highly parallelizable. Our experiments on the AEA Dataset demonstrate that FrameVoting reduces the frequency of inferences by 95.6% compared to frame-by-frame object detection, while still accurately identifying the user’s objects of interest in real-time at 30 fps on a Raspberry Pi 5.
Kao-Den Chang, He-Yen Hsieh, H. T. Kung 0001, Ziyun Li 0001, Sai Qian Zhang
ISCAS4
2024 Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical Framework
abstract
Augmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration.
Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong
ACM Trans. Design Autom. Electr. Syst.7
2024 Thermally Constrained Codesign of Heterogeneous 3-D Integration of Compute-in-Memory, Digital ML Accelerator, and RISC-V Cores for Mixed ML and Non-ML Workloads
abstract
Heterogeneous 3-D (H3D) integration not only reduces the chip form factor and fabrication cost but also allows the merging of diverse compute paradigms that suit different applications. This is especially attractive when modern algorithms, such as the augmented reality/virtual reality (AR/VR) workloads, consist of mixed machine learning (ML) and non-ML workloads. To date, codesign that considers the thermal, latency, and power constraints of H3D hardware is largely unexplored. In this work, a thermally aware framework for H3D hardware design is developed to evaluate the thermal, latency, and power trade-offs for a heterogeneous system with compute-in-memory (CIM), digital ML cores, and RISC-V cores. The framework solves for runtime tunable operating points described as the optimal speedup factor, the number of activated RISC-V cores, the cooling coefficient, and the activity rate based on user-defined criteria, achieving up to 135 TOPS and 215 TOPS/W under$74~^{\circ }$C for the AR/VR workloads.
Yuan-Chun Luo, Anni Lu, Janak Sharda, Moritz Scherer 0001, Jorge Gomez 0001, Syed Shakib Sarwar, Ziyun Li 0001, Reid Frederick Pinkham, Barbara De Salvo, Shimeng Yu
IEEE Trans. Very Large Scale Integr. Syst.7
2022 SplitNets: Designing Neural Architectures for Efficient Distributed Computing on Head-Mounted Systems
abstract
We design deep neural networks (DNNs) and corresponding networks' splittings to distribute DNNs' workload to camera sensors and a centralized aggregator on head mounted devices to meet system performance targets in inference accuracy and latency under the given hardware resource constraints. To achieve an optimal balance among computation, communication, and performance, a split-aware neural architecture search framework, SplitNets, is introduced to conduct model designing, splitting, and communication reduction simultaneously. We further extend the framework to multi-view systems for learning to fuse inputs from multiple camera sensors with optimal performance and systemic efficiency. We validate SplitNets for single-view system on ImageNet as well as multi-view system on 3D classification, and show that the SplitNets framework achieves state-of-the-art (SOTA) performance and system latency compared with existing approaches.
Xin Dong 0009, Barbara De Salvo, Meng Li 0004, Chiao Liu, Zhongnan Qu, H. T. Kung 0001, Ziyun Li 0001
CVPR7
2022 EyeCoD: eye tracking system acceleration via flatcam-based algorithm & accelerator co-design
abstract
Eye tracking has become an essential human-machine interaction modality for providing immersive experience in numerous virtual and augmented reality (VR/AR) applications desiring high throughput (e.g., 240 FPS), small-form, and enhanced visual privacy. However, existing eye tracking systems are still limited by their: (1) large form-factor largely due to the adopted bulky lens-based cameras; (2) high communication cost required between the camera and backend processor; and (3) potentially concerned low visual privacy, thus prohibiting their more extensive applications. To this end, we propose, develop, and validate a lensless FlatCambased eye tracking algorithm and accelerator co-design framework dubbed EyeCoD to enable eye tracking systems with a much reduced form-factor and boosted system efficiency without sacrificing the tracking accuracy, paving the way for next-generation eye tracking solutions. On the system level, we advocate the use of lensless FlatCams instead of lens-based cameras to facilitate the small form-factor need in mobile eye tracking systems, which also leaves rooms for a dedicated sensing-processor co-design to reduce the required camera-processor communication latency. On the algorithm level, EyeCoD integrates a predict-then-focus pipeline that first predicts the region-of-interest (ROI) via segmentation and then only focuses on the ROI parts to estimate gaze directions, greatly reducing redundant computations and data movements. On the hardware level, we further develop a dedicated accelerator that (1) integrates a novel workload orchestration between the aforementioned segmentation and gaze estimation models, (2) leverages intra-channel reuse opportunities for depth-wise layers, (3) utilizes input feature-wise partition to save activation memory size, and (4) develops a sequential-write-parallel-read input buffer to alleviate the bandwidth requirement for the activation global buffer. On-silicon measurement and extensive experiments validate that our EyeCoD consistently reduces both the communication and computation costs, leading to an overall system speedup of 10.95×, 3.21×, and 12.85× over general computing platforms including CPUs and GPUs, and a prior-art eye tracking processor called CIS-GEP, respectively, while maintaining the tracking accuracy. Codes are available at https://github.com/RICE-EIC/EyeCoD.
Haoran You, Cheng Wan 0005, Yang Zhao 0013, Zhongzhi Yu, Yonggan Fu, Jiayi Yuan 0001, Shang Wu 0003, Yongan Zhang, Chaojian Li, Vivek Boominathan, Ashok Veeraraghavan, Ziyun Li 0001, Yingyan (Celine) Lin
ISCA13
2019 Low Complexity, Hardware-Efficient Neighbor-Guided SGM Optical Flow for Low-Power Mobile Vision Applications
abstract
Accurate, low-latency, and energy-efficient optical flow estimation is a fundamental kernel function to enable several real-time vision applications on mobile platforms. This paper presents neighbor-guided semi-global matching (NG-fSGM), a new low-complexity optical flow algorithm tailored for low-power mobile applications. NG-fSGM obtains high accuracy optical flow by aggregating local matching costs over a semiglobal region, successfully resolving local ambiguity in textureless and occluded regions. The proposed NG-fSGM aggressively prunes the search space based on neighboring pixels' information to significantly lower the algorithm complexity from the original fSGM. As a result, NG-fSGM achieves 17.9× reduction in the number of computations and 8.37× reduction in memory space compared to the original fSGM without compromising its algorithm accuracy. A multicore architecture for NG-fSGM is implemented in hardware to quantify algorithm complexity and power consumption. The proposed architecture realizes NG-fSGM with overlapping blocks processed in parallel to enhance throughput and to lower power consumption. The eightcore architecture achieves 20 M pixel/s (66 frames/s for VGA) throughput with 9.6 mm2area at 679.2-mW power consumption in 28-nm node.
Ziyun Li 0001, Jiang Xiang, Luyao Gong, David T. Blaauw, Chaitali Chakrabarti, Hun-Seok Kim
IEEE Trans. Circuits Syst. Video Technol.1
2016 Low complexity optical flow using neighbor-guided semi-global matching
abstract
This paper presents Neighbor-Guided SemiGlobal Matching (NG-fSGM), a new method for optical flow. It is based on SGM, a popular dynamic programming algorithm for stereo vision, where the disparity of each pixel is calculated by aggregating local matching costs over the entire image to resolve local ambiguity in texture-less and occluded regions. Unlike conventional SGM, NG-fSGM operates on a subset of the search space that has been aggressively pruned based on neighboring pixels' information. Our proposed method achieves a fast approximation of SGM with significantly simpler cost aggregation and flow computation. Compared to a prior SGM extension for optical flow, the proposed NG-fSGM provides about 9x reduction in the number of computations and 5x reduction in the memory requirement with only 0.17% accuracy degradation when evaluated with Middlebury benchmark test cases.
Jiang Xiang, Ziyun Li 0001, David T. Blaauw, Hun-Seok Kim, Chaitali Chakrabarti
ICIP2