Yihan Pang

dblp:241/5847 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-0524-6934ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 32% Energy-efficient computing · 29% Cloud and datacenter computing · 28%
Computer graphics and multimedia
2 papers
Virtual and augmented reality · 100%

Topics — the 5 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
computation offloading
0.912025
Ada: A Distributed, Power-Aware, Real-Time Scene Provider for XR · IEEE Trans. Vis. Comput. Graph. 2025
Parallel and multicore computing › parallel scheduling
heterogeneous multiprocessor scheduling
0.912025
MLCD: Machine Learning-Based Code Version and Device Selection for Heterogeneous Systems · IEEE Trans. Computers 2025
Energy-efficient computing › energy-quality tradeoff
energy-latency-accuracy trade-off
0.812024
Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy · ISMAR 2024
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.312025
MLCD: Machine Learning-Based Code Version and Device Selection for Heterogeneous Systems · IEEE Trans. Computers 2025
Cloud and datacenter computing
resource management
0.112019
Quantifying Memory Underutilization in HPC Systems and Using it to Improve Performance via Architecture Support · MICRO 2019

Methods — techniques the papers use, named apart from their topics

computation offloading · 1.7GPU acceleration · 1.7pareto optimization · 1.5design space exploration · 1.5machine learning · 0.9active learning · 0.9workload characterization · 0.4
YearPublicationVenuePosition
2025 RemoteVIO: Offloading Head Tracking in an End-to-End XR System
abstract
Power consumption, and the resulting limitation to computational load, is a first-order constraint in designing comfortable all-day-wear extended reality (XR) devices that can provide rich immersive experiences. This paper concerns reducing XR device power consumption by offloading head tracking, one of the top CPU and power consumers, to a remote server. We present RemoteVIO, the first open-source end-to-end XR system that offloads head tracking (visual inertial odometry or VIO) to a remote server. Our work distinguishes itself from past studies on computation offloading in XR by properly addressing two under-explored but critical aspects: 1) a comprehensive evaluation of user experience in a complete end-to-end XR system and 2) a quantification of the net power savings on real hardware.
Qinjun Jiang, Yihan Pang, William Sentosa, Muhammad Huzaifa, Jeffrey Zhang 0004, Javier Perez-Ramirez, David Gonzalez-Aguirre, Brighten Godfrey, Sarita V. Adve
MMSys2
2025 MLCD: Machine Learning-Based Code Version and Device Selection for Heterogeneous Systems
abstract
Heterogeneous systems with hardware accelerators are increasingly common, and various optimized implementations/algorithms exist for computation kernels. However, no single best combination ofcode version and device(C&D) can outper-form others across all input cases, demanding a method to select the best C&D pair based on input. We presentmachinelearning-basedcode version anddevice selection method, namedMLCD, that uses input data characteristics to select the best C&D pair dynamically. We also apply active learning to reduce the number of samples needed to construct the model. Demonstrated on two different CPU-GPU systems, MLCD achieves near-optimal speed-up regardless of which systems tested. Concretely, reporting results from system one with mid-end hardwares, it achieves 99.9%, 95.6%, 99.9%, and 98.6% of the optimal acceleration attainable through the ideal choice of C&D pairs in General Matrix Multiply, PageRank, N-body Simulation, and K-Motif Counting, respectively. MLCD achieves a speed-up of 2.57×, 1.58×, 2.68×, and 1.09× compared to baselines without MLCD. Additionally, MLCD handles end-to-end applications, achieving up to 10% and 46% speed-up over GPU-only and CPU-only solutions with Graph Neural Networks. Furthermore, it achieves 7.28× average speed-up in execution latency over the state-of-the-art approach and determines suitable code versions for unseen input 108− 1010× faster.
Kaiwen Cao, Hanchen Ye, Yihan Pang, Deming Chen
IEEE Trans. Computers3
2025 Ada: A Distributed, Power-Aware, Real-Time Scene Provider for XR
abstract
Real-time scene provisioning-reconstructing and delivering scene data to requesting XR applications during runtime-is central to enabling spatial computing in modern XR systems. However, existing solutions struggle to balance latency, power and scene fidelity under XR device constraints, and often rely on designs that are either closed, application-specific designs, or both. We present Ada, the first open distributed, power-aware, application-agnostic real-time scene provisioning system. Through computation offloading along with algorithmic and system innovations, Ada provides high-fidelity scenes with stable performance across all evaluated scene sizes and with low power consumption. To isolate the benefits of Ada's algorithmic and design innovations over the closest prior work [82], which is on-device and CPU-based, we configure a comparable on-device, CPU-based variant of Ada (AdaLocal-CPU). We show this variant achieves up to 6.8× lower scene request latency and higher scene fidelity compared to the prior work. Furthermore, Ada's final distributed GPU-accelerated implementation reduces latency by an additional 2×, highlighting the benefits of GPU acceleration and distributed computing. Additionally, Ada also lowers the incremental power cost of scene provisioning by 24% compared to the best on-device variant (AdaLocal-GPU). Finally, Ada flexibly adapts to diverse latency, power, scene fidelity, and network bandwidth requirements.
Yihan Pang, Sushant Kondguli, Shenlong Wang, Sarita V. Adve
IEEE Trans. Vis. Comput. Graph.1
2024 Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy
abstract
Extended Reality (XR) enables immersive experiences through untethered headsets but suffers from stringent battery and resource constraints. Energy-efficient design is crucial to ensure both longevity and high performance in XR devices. However, latency and accuracy are often prioritized over energy, leading to a gap in achieving energy efficiency. This paper examines scene reconstruction, a key building block for immersive XR experiences, and demonstrates how energy efficiency can be achieved by navigating the trilemma of energy, latency, and accuracy. We explore three classes of energy-oriented optimizations, covering the algorithm, execution, and data, that reveal a broad de-sign space through configurable parameters. Our resulting 72 designs expose a wide range of latency and energy trade-offs, with a smaller range of accuracy loss. We identify a Pareto-optimal curve and show that the designs on the curve are achievable only through synergistic co-optimization of all three optimization classes and by considering the latency and accuracy needs of downstream scene reconstruction consumers. Our analysis covering various use cases and measurements on an embedded class system shows that, relative to the baseline, our designs offer energy benefits of up to $60 \times$ with potential latency range of $4 \times$ slowdown to $2 \times$ speedup. Detailed exploration of a use case across representative data sequences from ScanNet showed about $25 \times$ energy savings with $1.5 \times$ latency reduction and negligible reconstruction quality loss.
Boyuan Tian, Yihan Pang, Muhammad Huzaifa, Shenlong Wang, Sarita V. Adve
ISMAR2
2019 Quantifying Memory Underutilization in HPC Systems and Using it to Improve Performance via Architecture Support
abstract
A system's memory size is often dictated by worst-case workloads with highest memory requirements; this causes memory to be underutilized in the common case when the system is not running its worst-case workloads. Cognizant of this memory underutilization problem, many prior works have studied memory utilization and explored how to improve it in the context of cloud.
Gagandeep Panwar, Da Zhang 0004, Yihan Pang, Mai Dahshan, Nathan DeBardeleben, Binoy Ravindran, Xun Jian 0002
MICRO3
2019 Cross-ISA execution of SIMD regions for improved performance
abstract
We investigate the effectiveness of executing SIMD workloads on multiprocessors with heterogeneous Instruction Set Architecture (ISA) cores. Heterogeneous ISAs offer an intriguing clock speed/parallelism tradeoff for workloads with frequent usage of SIMD instructions. We consider dynamic migration of SIMD and non-SIMD workloads across ISA-different cores to exploit this trade-off. We present the necessary modifications for a general compiler/run-time infrastructure to transform the dynamic program state of SIMD regions at run-time from one ISA format to another for cross-ISA migration and execution. Additionally, we present a SIMD-aware scheduling policy that makes cross-ISA migration decisions that improve system throughput. We prototype a heterogeneous-ISA system using an Intel Xeon x86-64 server and a Cavium ThunderX ARMv8 server and evaluate the effectiveness of our infrastructure and scheduling policy. Our results reveal that cross-ISA execution migration within SIMD regions can yield throughput gains up to 36% compared to traditional homogeneous ISA systems.
Yihan Pang, Robert Lyerly, Binoy Ravindran
SYSTOR1