Juseong Park

dblp:311/2384 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generating High Dimensional User-Specific Wireless Channels Using Diffusion Models
abstract
Deep neural network (DNN)-based algorithms are emerging as an important tool for many physical and MAC layer functions in future wireless communication systems, including for large multi-antenna channels. However, training such models typically requires a large dataset of high-dimensional channel measurements, which are very difficult and expensive to obtain. This paper introduces a novel method for generating synthetic wireless channel data using diffusion-based models to produce user-specific channels that accurately reflect real-world wireless environments. Our approach employs a conditional denoising diffusion implicit model (cDDIM) framework, effectively capturing the relationship between user location and multi-antenna channel characteristics. We generate synthetic high fidelity channel samples using user positions as conditional inputs, creating larger augmented datasets to overcome measurement scarcity. The utility of this method is demonstrated through its efficacy in training various downstream tasks such as channel compression and beam alignment. Our diffusion-based augmentation approach achieves over a 1-2 dB gain in NMSE for channel compression, and an 11 dB SNR boost in beamforming compared to prior methods, such as noise addition or the use of generative adversarial networks (GANs).
Taekyun Lee, Juseong Park, Hyeji Kim, Jeffrey G. Andrews
IEEE Trans. Wirel. Commun.2
2026 Self-Nomination: Deep Learning for Decentralized CSI Feedback Reduction in MU-MIMO Systems
abstract
This paper introduces a novel deep learning-based user-side feedback reduction framework, termedself-nomination. The goal of self-nomination is to reduce the number of users (UEs) feeding back channel state information (CSI) to the base station (BS), by letting each UE decide whether to feed back based on its estimated likelihood of being scheduled and its potential contribution to precoding in a multiuser MIMO (MU-MIMO) downlink. Unlike SNR- or SINR-based thresholding methods, the proposed approach uses rich spatial channel statistics and learns nontrivial correlation effects that affect eventual MU-MIMO scheduling decisions. To train the self-nomination network under an average feedback constraint, we propose two different strategies: one based on direct optimization with gradient approximations, and another using policy gradient-based optimization with a stochastic Bernoulli policy to handle non-differentiable scheduling. The framework also supports proportional-fair scheduling by incorporating dynamic user weights. Numerical results confirm that the proposed self-nomination method significantly reduces CSI feedback overhead. Compared to baseline feedback methods, self-nomination can reduce feedback by as much as 65%, saving not only bandwidth but also allowing many UEs to avoid feedback altogether (and thus, potentially enter a sleep mode). Self-nomination achieves this significant savings with negligible reduction in sum-rate or fairness.
Juseong Park, Foad Sohrabi, Jinfeng Du, Jeffrey G. Andrews
IEEE Trans. Wirel. Commun.1
2025 HYPERF: End-to-End Autotuning Framework for High-Performance Computing
Juseong Park, Yongwon Shin, Oh-Kyoung Kwon, Hyojin Sung
HPDC1
2025 End-to-End Deep Learning for TDD MIMO Systems in the 6G Upper Midbands
abstract
This paper proposes and analyzes novel deep learning methods for downlink (DL) single-user multiple-input multiple-output (MIMO) and multi-user MIMO (MU-MIMO) systems operating in time division duplex mode. A motivating application is the 6G upper midbands (7-24 GHz), where the base station (BS) antenna arrays are large, user equipment array sizes are moderate, and theoretically optimal approaches are practically infeasible for several reasons. To deal with uplink (UL) pilot overhead and low signal power issues, we introduce the channel-adaptive pilot, as part of the novel analog channel state information feedback mechanism. Deep neural network (DNN)-generated pilots are used to linearly transform the UL channel matrix into lower-dimensional latent vectors. Meanwhile, the BS employs a second DNN that processes the received UL pilots to directly generate near-optimal DL precoders. The training is end-to-end which exploits synergies between the two DNNs. For MU-MIMO precoding, we propose a DNN structure inspired by theoretically optimum linear precoding. The proposed methods are evaluated against genie-aided upper bounds and conventional approaches, using realistic upper midband datasets. Numerical results demonstrate the potential of our approach to achieve significantly increased sum-rate, particularly at moderate to high signal-to-noise ratio and when UL pilot overhead is constrained.
Juseong Park, Foad Sohrabi, Amitava Ghosh, Jeffrey G. Andrews
IEEE Trans. Wirel. Commun.1
2024 NavCim: Comprehensive Design Space Exploration for Analog Computing-in-Memory Architectures
abstract
Analog computing-in-memory (ACiM) technology has shown strong potential for neural network accelerators, addressing von-Neumann performance bottlenecks with in-memory data processing and computation. Understanding the ACiM design space, including its trade-offs and constraints, and systematically and effectively exploring it for optimal performance is essential to turn the promise into a viable product. Recent research demonstrated that multi-objective searches for ACiM architectures with heterogeneous tiles can simultaneously optimize power, performance, and area (PPA), outperforming existing tiled ACiM proposals. In this paper, we propose NavCim, a comprehensive ACiM design space exploration mechanism that advances the prior work in terms of search efficiency, search space coverage, and optimization metrics. NavCim introduces predictive modeling of ACiM hardware performance and uses the PPA prediction models instead of running simulators, significantly reducing search overheads. Faster searches enable NavCim to extend the architecture and model search spaces with an evolutionary search process to optimize architectures with more than two different tile sizes for multiple input models. With accuracy-aware searches, NavCim considers PPA and model accuracy together as optimization goals to achieve more balanced trade-offs. The experimental searches show that NavCim leverages predictive models to reduce search time by up to 7.3x without compromising the quality of search results. It also successfully identifies heterogeneous ACiM architectures that can efficiently execute multiple models on a single chip, improving accuracy by up to 19% over the prior work.
Juseong Park, Boseok Kim, Hyojin Sung
PACT1
2023 PIMFlow: Compiler and Runtime Support for CNN Models on Processing-in-Memory DRAM
abstract
Processing-in-Memory (PIM) has evolved over decades into a feasible solution to addressing the exacerbating performance bottleneck with main memory by placing computational logic in or near memory. Recent proposals from DRAM manufacturers highlighted the HW constraint-aware design of PIM-enabled DRAM with specialized MAC logic, providing an order of magnitude speedup for memory-intensive operations in DL models. Although the main target for PIM acceleration did not initially include convolutional neural networks due to their high compute intensity, recent CNN models are increasingly adopting computationally lightweight implementation. Motivated by the potential for the software stack to enable CNN models on DRAM-PIM hardware without invasive changes, we propose PIMFlow, an end-to-end compiler and runtime support, to accelerate CNN models on a PIM-enabled GPU memory. PIMFlow transforms model graphs to create inter-node parallelism across GPU and PIM, explores possible task- and data-parallel execution scenarios for optimal execution time, and provides a code-generating back-end and execution engine for DRAM-PIM. PIMFlow achieves up to 82% end-to-end speedup and reduces energy consumption by 26% on average for CNN model inferences.
Yongwon Shin, Juseong Park, Sungjun Cho, Hyojin Sung
CGO2
2023 Multi-Objective Architecture Search and Optimization for Heterogeneous Neuromorphic Architecture
abstract
Neuro-inspired in-memory computing offers a promising solution to overcome the limitations of traditional von Neumann architectures by emulating brain activities. This approach takes advantage of parallel processing while minimizing power consumption and area overhead. However, maximizing the performance of neuromorphic hardware, particularly for increasingly deep and complex NN models, is a challenging task. Existing design space exploration methods focus primarily on layer placement and resource allocation, ignoring hardware-level configurations that directly influence performance, power, and area (PPA) trade-offs. Additionally, current tiled neuromorphic architectures lack support for size heterogeneity, making optimal resource utilization difficult. To address these challenges, we propose a multi-objective architecture search and optimization mechanism for neuromorphic architectures. Our approach introduces heterogeneous architectures with multiple tile/processing element/synaptic array sizes and provides a comprehensive end-to-end design automation tool to support them. We define a heterogeneous neuromorphic architecture, as exemplified by a “big-tile, little-tile” architecture with mesh interconnects. Our search mechanism expands the search space by considering candidates for both homogeneous and heterogeneous architectures and performs Pareto-front searches guided by user-defined weights or constraints on performance metrics. We also implement a hierarchical beam search technique to explore the vast search space of heterogeneous architecture candidates more effectively. Our mechanism identifies numerous heterogeneous architectures that outperform the baseline for different convolutional neural network (CNN) models. For EfficientNetB0, we achieve PPA improvements of 40.1%, 19.3%, and 4.4% over the baseline. Our tool is available at https://github.com/wntjd9805/hetero-neurosim-search.
Juseong Park, Yongwon Shin, Hyojin Sung
ICCAD1
2022 Deep Reinforcement Learning Approach for UAV-Assisted Mobile Edge Computing Networks
abstract
This paper studies a deep reinforcement learning (DRL) approach for the unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) networks where a UAV-mounted server offloads computation tasks of mobile users (MUs). We aim at minimizing the energy consumption of the MUs by adjusting UAV mobility, UAV-MU association, computation resource allocation, and task offloading rules. This requires an online and joint optimization of different types of variables constructing heterogeneous solution spaces. To realize real-time optimization strategies, we propose an online DRL method based on the twin-delayed deep deterministic policy gradient (TD3) framework. The joint optimization of heterogeneous action variables is tackled by a novel actor neural network that partitions the high-dimensional action set into several solution spaces. In addition, the proposed TD3 framework achieves adaptability to new task offloading requests through our proposed training and execution strategy. Numerical results verify the effectiveness of the proposed DRL architecture over benchmark schemes.
Juseong Park, Hoon Lee, Inkyu Lee
GLOBECOM2