Joongun Park

dblp:227/0785 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0000-8406-0817ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Compositional AI Beyond LLMs: System Implications of Neuro-Symbolic-Probabilistic Architectures
abstract
Large Language Models (LLMs) have driven remarkable progress in artificial intelligence (AI), but their rapid growth faces challenges of unsustainable computation, limited robustness, and poor explainability. Compositional AI, which integrates LLMs with symbolic reasoning and probabilistic inference, has emerged as a promising paradigm to enable interpretability, robustness, trustworthiness, and data-efficient learning. Recent neuro-symbolic-probabilistic systems demonstrate strong potential in agentic applications, advancing reasoning and cognitive capabilities toward human-like intelligence.
Zishen Wan, Hanchen Yang 0001, Jiayi Qian, Ritik Raj, Joongun Park, Arijit Raychowdhury, Tushar Krishna
ASPLOS (1)5
2026 SCALE: Tackling Communication Bottlenecks in Confidential Distributed Machine Learning
abstract
Machine Learning (ML) has become a cornerstone in numerous applications, creating the need for secure and efficient distributed ML frameworks. However, maintaining data privacy in these systems poses significant challenges, particularly in distributed environments where user data and model parameters must frequently be transmitted between GPUs. Confidential GPU computing technologies, such as NVIDIA's Confidential Computing (CC) mode, offer hardware-based enterprise solutions designed to protect ML workloads in untrusted environments (e.g., public clouds). These technologies leverage heterogeneous systems that combine Confidential Virtual Machines (CVMs) with GPU-based Trusted Execution Environments (TEEs). Nevertheless, confidential computing introduces considerable performance overhead due to its complex heterogeneous architecture and the high-throughput data flows required across TEE security boundaries. For example, encrypted communication occurs both between CVMs and GPU TEEs, and among multiple GPU TEEs, resulting in significant latency compared to native PCIe or high-speed interconnects such as NVLink. Our extensive evaluation shows that these overheads become particularly severe during collective communication operations, which suffer from encryption-induced delays that negatively impact end-to-end training performance. To address this, we propose a co-encryption design that leverages underutilized GPU resources, optimizes encryption and authentication, and introduces a communication algorithm tailored for confidential settings. We evaluate our design using real ML workloads and execution traces collected from four HGX H100/H200 clusters. While CC mode was not available on current NVIDIA software stacks, we incorporate encryptionaware modeling based on hardware specifications to estimate secure communication overheads. Our results demonstrate a$\mathbf{4 0 - 7 0 \%}$reduction in communication-related security costs.
Joongun Park, Yongqin Wang, Hanjiang Wu, Tushar Krishna
HPCA1
2026 Closing the Efficiency Gap: AI Datacenter Co-design Roadmap for Scalable Training of LLMs
abstract
The massive compute, memory, and networking needs for LLM training necessitate a fundamental rethinking of datacenter architectures to ensure scalability, efficiency, and cost-effectiveness. In particular, the design of the network fabric for AI datacenters for emerging LLMs (such as MoEs) remains a crucial and challenging open question, spanning technology choices (that determine the size of the high-bandwidth domain), topology, and software optimizations (collective algorithms and overlap strategies). This necessitates an agile framework to traverse the co-design space. This work introduces Calculon-MoE, a tool that jointly explores FLOPS, HBM bandwidth and capacity, multiple network topologies (Two-tiered vs. FullFlat optical), the size of scale-up domain, and popular parallelism/optimization strategies used in LLMs. Our validation studies demonstrate that our LLM/MoE runtime predictions are within 10% of real-world measurements. Using Calculon-MoE, we conduct a suite of case studies to develop an actionable roadmap for data centers. For example, the results point to the promise of Fullflat network architectures, which provide uniform high-bandwidth, low-latency connectivity between all nodes and demonstrate their positive impacts on performance and scalability. We also quantify the benefits of overlapping compute and communication, hardware-accelerated collectives, widening the scale-up domain, and higher memory bandwidth and capacity. Our study spans both sparse (mixture of experts) and dense transformer-based LLMs, revealing how system design and optimization choices affect system efficiency and overall throughput in both cases.
Jesmin Jahan Tithi, Hanjiang Wu, Joongun Park, Avishaii Abuhatzera, Fabrizio Petrini, Tushar Krishna
ICS3
2026 Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor Graphs
Changhai Man, Joongun Park, Hanjiang Wu, Srinivas Sridharan 0002, Tushar Krishna
ISCA2
2025 NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
abstract
Neuro-Symbolic AI (NSAI) is an emerging paradigm that integrates neural networks with symbolic reasoning to enhance the transparency, reasoning capabilities, and data efficiency of AI systems. Recent NSAI systems have gained traction due to their exceptional performance in reasoning tasks and human-AI collaborative scenarios. Despite these algorithmic advancements, executing NSAI tasks on existing hardware (e.g., CPUs, GPUs, TPUs) remains challenging, due to their heterogeneous computing kernels, high memory intensity, and unique memory access patterns. Moreover, current NSAI algorithms exhibit significant variation in operation types and scales, making them incompatible with existing ML accelerators. These challenges highlight the need for a versatile and flexible acceleration framework tailored to NSAI workloads. In this paper, we propose NSFlow, an FPGA-based acceleration framework designed to achieve high efficiency, scalability, and versatility across NSAI systems. NSFlow features a design architecture generator that identifies workload data dependencies and creates optimized dataflow architectures, as well as a reconfigurable array with flexible compute units, re-organizable memory, and mixed-precision capabilities. Evaluating across NSAI workloads, NSFlow achieves $31 \times$ speedup over Jetson TX2, more than $2 \times$ over GPU, $8 \times$ speedup over TPU-like systolic array, and more than $3 \times$ over Xilinx DPU. NSFlow also demonstrates enhanced scalability, with only $4 \times$ runtime increase when symbolic workloads scale by $150 \times$. To the best of our knowledge, NSFlow is the first framework to enable real-time generalizable NSAI algorithms acceleration, demonstrating a promising solution for next-generation cognitive systems.
Hanchen Yang 0001, Zishen Wan, Ritik Raj, Joongun Park, Ananda Samajdar, Arijit Raychowdhury, Tushar Krishna
DAC4
2025 Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
abstract
The rapid scaling of Large Language Models (LLMs) has pushed training workloads far beyond the limits of single-node analysis, demanding a deeper understanding of how these models behave across large-scale, multi-GPU systems.In this paper, we present a comprehensive characterization of LLM training across diverse real-world workloads and hardware platforms, including NVIDIA H100/H200 and AMD MI250 GPUs.We analyze dense and sparse models under various parallelism strategies -tensor, pipeline, data, and expert -and evaluate their effects on hardware utilization, power consumption, and thermal behavior.We further evaluate the effectiveness of optimizations such as activation recomputation and compute-communication overlap.Our findings show that performance is not determined solely by scaling hardware capacity.Scale-up systems with fewer, higher-memory GPUs can outperform scale-out systems in communication-bound regimes, but only under carefully tuned configurations; in other cases, scale-out deployments achieve superior throughput.We also show that certain parallelism combinations, such as tensor with pipeline, lead to bandwidth underutilization due to inefficient data chunking, while increasing microbatch sizes beyond a certain point induces bursty execution and peak power excursions that worsen thermal throttling.These insights reveal how training performance is shaped by complex interactions between hardware, system topology, and model execution.We conclude by offering recommendations for system and hardware design to improve the scalability and reliability of future LLM systems and workloads.The source code of this project is available at https:/
Seokjin Go, Joongun Park, Spandan More, Hanjiang Wu, Irene Wang, Aaron Jezghani, Tushar Krishna, Divya Mahajan 0001
MICRO2
2024 Supporting Trusted Virtual Machines with Hardware-Based Secure Remote Memory
abstract
Although recent studies have been improving the performance of RDMA-based memory disaggregation systems, their security aspect has not been thoroughly investigated. For secure disaggregated memory, the memory-providing node must protect its memory from memory-requesting nodes, and the memory-requesting node requires the confidentiality and integrity protection of its memory contents in the remote node, even when the privileged software is compromised. To provide protection of remote memory, this study proposes a hardware-assisted memory disaggregation system. The proposed trusted disaggregated memory combines the current trusted hardware-based virtual machine (VM) and a new dedicated hardware engine for trusted memory disaggregation. The processor with supports for trusted VM protects the context of a user VM within the local system, while the proposed hardware engine provides an efficient isolation and protection of remote memory pages, guaranteeing the confidentiality and integrity of remote memory pages. In the secure memory disaggregation system, fast address translation and access validation are supported with the cooperation of the hardware engine and guest OS in a trusted virtual machine. In addition, the proposed system hides the memory access patterns observable from remote nodes, supporting obliviousness. Our evaluation with an FPGA-based prototype implementation shows that such fine-grained secure disaggregated memory is feasible with comparable performance to the latest software-based technique without security support.
Taekyung Heo, Seunghyo Kang, Soojin Hwang, Joongun Park, Jaehyuk Huh 0001
ISMM5
2024 Hardware-hardened Sandbox Enclaves for Trusted Serverless Computing
abstract
In cloud-based serverless computing, an application consists of multiple functions provided by mutually distrusting parties. For secure serverless computing, the hardware-based trusted execution environment (TEE) can provide strong isolation among functions. However, not only protecting each function from the host OS and other functions, but also protecting the host system from the functions, is critical for the security of the cloud servers. Such an emerging trusted serverless computing poses new challenges: Each TEE must be isolated from the host system bi-directionally, and the system calls from it must be validated. In addition, the resource utilization of each TEE must be accountable in a mutually trusted way. However, the current TEE model cannot efficiently represent such trusted serverless applications. To overcome the lack of such hardware support, this article proposes an extended TEE model called Cloister , designed for trusted serverless computing. Cloister proposes four new key techniques. First, it extends the hardware-based memory isolation in SGX to confine a deployed function only within its TEE (enclave). Second, it proposes a trusted monitor enclave that filters and validates system calls from enclaves. Third, it provides a trusted resource accounting mechanism for enclaves that is agreeable to both service developers and cloud providers. Finally, Cloister accelerates enclave loading by redesigning its memory verification for fast function deployment. Using an emulated Intel SGX platform with the proposed extensions, this article shows that trusted serverless applications can be effectively supported with small changes in the SGX hardware.
Joongun Park, Seunghyo Kang, Taehoon Kim 0001, Jongse Park, Youngjin Kwon, Jaehyuk Huh 0001
ACM Trans. Archit. Code Optim.1
2020 Nested Enclave: Supporting Fine-grained Hierarchical Isolation with SGX
abstract
Although hardware-based trusted execution environments (TEEs) have evolved to provide strong isolation with efficient hardware supports, their current monolithic model poses challenges in representing common software structures with modules produced from potentially untrusted 3rd parties. For better mapping of such modular software designs to trusted execution environments, it is necessary to extend the current monolithic model to a hierarchical one, which provides multiple inner TEEs within a TEE. For such hierarchical compartmentalization within a TEE, this paper proposes a novel hierarchical TEE called nested enclave, which extends the enclave support from Intel SGX. Inspired by the multi-level security model, nested enclave provides multiple inner enclaves sharing the same outer enclave. Inner enclaves can access the context of the outer enclave, but they are protected from the outer enclave and non-enclave execution. Peer inner enclaves are isolated from each other while accessing the execution environment of the shared outer enclave. Both of the inner and outer enclaves are protected from vulnerable privileged software and physical attacks. Such fine-grained nested enclaves allow secure multitiered environments using software modules from untrusted 3rd parties. The security-sensitive modules run on the inner enclave with the higher security level, while the 3rd party modules on the outer enclave. It can be further extended to provide a separate inner module for each user to process privacy-sensitive data while sharing the same library with efficient hardwareprotected communication channels. This study investigates three case scenarios implemented with an emulated nested enclave support, proving the feasibility and security improvement of the nested enclave model.
Joongun Park, Naegyeong Kang, Taehoon Kim 0001, Youngjin Kwon, Jaehyuk Huh 0001
ISCA1
2019 ShieldStore: Shielded In-memory Key-value Storage with SGX
abstract
The shielded computation of hardware-based trusted execution environments such as Intel Software Guard Extensions (SGX) can provide secure cloud computing on remote systems under untrusted privileged system software. However, hardware overheads for securing protected memory restrict its capacity to a modest size of several tens of megabytes, and more demands for protected memory beyond the limit cause costly demand paging. Although one of the widely used applications benefiting from the enhanced security of SGX, is the in-memory key-value store, its memory requirements are far larger than the protected memory limit. Furthermore, the main data structures commonly use fine-grained data items such as pointers and keys, which do not match well with the coarse-grained paging of the SGX memory extension technique. To overcome the memory restriction, this paper proposes a new in-memory key-value store designed for SGX with application-specific data security management. The proposed key-value store, called ShieldStore, maintains the main data structures in unprotected memory with each key-value pair individually encrypted and integrity-protected by its secure component running inside an enclave. Based on the enclave protection by SGX, ShieldStore provides secure data operations much more efficiently than the baseline SGX key-value store, achieving 8--11 times higher throughput with 1 thread, and 24--30 times higher throughput with 4 threads.
Taehoon Kim 0001, Joongun Park, Jae-Woo Chang, Seungheun Jeon, Jaehyuk Huh 0001
EuroSys2
2018 Secure In-memory Key-Value Storage with SGX
abstract
No abstract available.
Taehoon Kim 0001, Joongun Park, Jae-Woo Chang, Seungheun Jeon, Jaehyuk Huh 0001
SoCC2