Pingyi Huo

dblp:322/4220 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-9769-2570ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 TPNM: A CXL Based General Purpose Tiered Process Near Memory Framework
abstract
Near-Memory Processing (NMP) has gained significant attention for its potential to accelerate various workloads. However, NMP performance suffers from challenges related to data locality and scalability, particularly in disaggregated datacenter environments. To address these issues, this paper presents TPNM (Tiered Processing Near Memory), a novel framework for near-data processing in disaggregated memory settings. TPNM leverages Compute Express Link (CXL) technology to enable processing both within memory devices and at the fabric switch level, creating a tiered approach to data processing. Evaluations across diverse workloads demonstrate significant performance improvements over baseline and existing near-data processing approaches. In particular, TPNM reduces latency by up to 4.76x compared to CPU-only baselines and by as much as 62 % compared to existing NMP-based solutions.
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Meena Arunachalam, Mahmut T. Kandemir, Narayanan Vijaykrishnan
ISPASS1
2025 Multi-Dimensional ML-Pipeline Optimization in Cost-Effective Disaggregated Datacenter
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Nandhini Chandramoorthy, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Narayanan Vijaykrishnan
MICRO1
2025 SVDE: Serverless Framework for Low-Latency Video Analytic Queries With Hardware Disaggregation
abstract
Video analytics applications are gaining popularity among serverless environments. One issue that appears in existing serverless platforms is that they do not fully exploit opportunities to efficiently handle video chunking as they assume that all video chunks display similar computation and communication overheads. Those overheads vary and benefit from fine-grained handling. Scheduling non-uniform chunks is even more challenging when the hardware resources are heterogeneous and the network between the resources is non-uniform. To address these challenges, we propose SVDE, a heterogeneous serverless cloud framework for massive video processing workloads. SVDE employs a trained decision tree regression to efficiently decide where to process each video chunk, by holistically considering the non-uniform chunk sizes, heterogeneity of the node hardware, queuing status of each node, and an unbalanced network. Furthermore, we develop an efficient operator backend that will be open-sourced as part of SVDE.Compared to prior works, SVDE achieves up to 3.2× speedup on ten real-world video workloads due to its holistic scheduling decision-making, while our operator backend outperforms the popular Pytorch JIT backend by 5×.
Pingyi Huo, Theodore Michailidis, Prapti Panigrahi, Kiwan Maeng, Jishen Zhao, Narayanan Vijaykrishnan
IEEE Trans. Computers1
2024 PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
abstract
Deep Learning Recommendation Models (DLRMs) have become increasingly popular and prevalent in today's datacenters, consuming most of the AI inference cycles. The performance of DLRMs is heavily influenced by available band-width due to their large vector sizes in embedding tables and concurrent accesses. To achieve substantial improvements over existing solutions, novel approaches towards DLRM optimization are needed, especially, in the context of emerging interconnect technologies like CXL. This study delves into exploring CXL-enabled systems, implementing a process-in-fabric-switch (PIFS) solution to accelerate DLRMs while optimizing their memory and bandwidth scalability. We present an in-depth characterization of industry-scale DLRM workloads running on CXL-ready systems, identifying the predominant bottlenecks in existing CXL systems. We, therefore, propose PIFS-Rec, a PIFS-based scheme that implements near-data processing through downstream ports of the fabric switch. PIFS-Rec achieves a latency that is 3.89 x lower than Pond, an industry-standard CXL-based system, and also outperforms BEACON, a state-of-the-art scheme, by 2.03x.
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Krishnakumar Nair, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Narayanan Vijaykrishnan
MICRO1
2024 QoS-Diff: Adaptive Auto-tuning Framework for Low-latency Diffusion Model Inference
Pingyi Huo, Ajay Narayanan Sridhar, Md Fahim Faysal Khan, Kiwan Maeng, Narayanan Vijaykrishnan
MMAsia1
2023 ISVABI: In-Storage Video Analytics Engine with Block Interface
abstract
The wide use of cameras in the past decade has increased the need to process video data significantly. Due to the large volume of video data, analyzing videos to extract useful information has become a critical challenge. Several prior works have tried to accelerate video analytics workloads by offloading some operations to embedded processors within storage devices.
Joshua Fixelle, Pingyi Huo, Mircea R. Stan, Michael P. Mesnier, Narayanan Vijaykrishnan
LCTES3
2022 ISKEVA: in-SSD key-value database engine for video analytics applications
abstract
Key-value databases are widely used to store the features or metadata generated from the neural network based video processing platforms. Due to the large volumes of video data, these databases use solid state drives (SSDs) as the primary data storage platform, and user query-based filtering, and retrieval operations on data incur large volume of data movement between the SSD and the host processor. In this paper, we present an in-SSD key-value database which uses the embedded CPU core, and DRAM memory on the SSD to support various queries with predicates and reduce the data movement between SSD and host processor significantly. We augment the SSD flash translation layer with key-value database functions and auxiliary data structures to support the user queries using the embedded core and DRAM memory on SSD. The proposed key-value store prototype on the Cosmos plus OpenSSD board reduces data movement between host processor and SSD by 14.57x, achieves an application-level speedup by 1.16x, and reduced energy consumption by 56% across different types of user queries.
Joshua Fixelle, Nagadastagiri Challapalle, Pingyi Huo, Zhaoyan Shen, Zili Shao, Mircea R. Stan, Narayanan Vijaykrishnan
LCTES4