Jungwoo Park

dblp:195/5411 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 19% Knowledge representation and reasoning · 19% Trustworthy machine learning · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.532026
The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models · AAAI 2026
Monet: Mixture of Monosemantic Experts for Transformers · ICLR 2025
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
analogical reasoning
1.012026
The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational reasoning
1.012026
The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models · AAAI 2026
Machine learning › Efficient and distributed learning › model compression › quantization › low-bit quantization
4-bit quantization
0.912025
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models · ACL (1) 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Monet: Mixture of Monosemantic Experts for Transformers · ICLR 2025
Machine learning › Trustworthy machine learning
language model interpretability
0.912025
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information · ACL (1) 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
Monet: Mixture of Monosemantic Experts for Transformers · ICLR 2025
Natural language and speech › Question answering and dialogue systems
medical reasoning
0.912025
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards · EMNLP 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Monet: Mixture of Monosemantic Experts for Transformers · ICLR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models · ACL (1) 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
temporal knowledge
0.912025
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information · ACL (1) 2025
Memory systems
cache management
0.412019
MH Cache: A Mult Stephen Jarvisi-retention STT-RAM-based Low-power Last-level Cache for Mobile Hardware Rendering Systems · ACM Trans. Archit. Code Optim. 2019
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.412019
MH Cache: A Mult Stephen Jarvisi-retention STT-RAM-based Low-power Last-level Cache for Mobile Hardware Rendering Systems · ACM Trans. Archit. Code Optim. 2019
Memory systems › cache
STT-RAM cache
0.412019
MH Cache: A Mult Stephen Jarvisi-retention STT-RAM-based Low-power Last-level Cache for Mobile Hardware Rendering Systems · ACM Trans. Archit. Code Optim. 2019
Machine learning › Efficient and distributed learning › model deployment
on-device deployment
0.312025
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

layer-wise probing · 1.0hidden representation patching · 1.0stepwise process reward modeling · 0.9reinforcement learning · 0.9muon optimizer · 0.9mixture of experts · 0.9embedding projection · 0.9circuit analysis · 0.9attention head analysis · 0.9RMSNorm · 0.9write-intensity measurement · 0.4multi-retention cache management · 0.4
YearPublicationVenuePosition
2026 The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
abstract
Analogical reasoning is at the core of human cognition, serving as an important foundation for a variety of intellectual activities. While prior work has shown that LLMs can represent task patterns and surface-level concepts, it remains unclear whether these models can encode high-level relational concepts and apply them to novel situations through structured comparisons. In this work, we explore this fundamental aspect using proportional and story analogies, and identify three key findings. First, LLMs effectively encode the underlying relationships between analogous entities; both attributive and relational information propagate through mid-upper layers in correct cases, whereas reasoning failures reflect missing relational information within these layers. Second, unlike humans, LLMs often struggle not only when relational information is missing, but also when attempting to apply it to new entities. In such cases, strategically patching hidden representations at critical token positions can facilitate information transfer to a certain extent. Lastly, successful analogical reasoning in LLMs is marked by strong structural alignment between analogous situations, whereas failures often reflect degraded or misplaced alignment. Overall, our findings reveal that LLMs exhibit emerging but limited capabilities in encoding and applying high-level relational concepts, highlighting both parallels and gaps with human cognition.
Taewhoo Lee, Minju Song, Chanwoong Yoon, Jungwoo Park, Jaewoo Kang
AAAI4
2025 Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
abstract
Extreme activation outliers in Large Language Models (LLMs) critically degrade quantization performance, hindering efficient on-device deployment. While channel-wise operations and adaptive gradient scaling are recognized causes, practical mitigation remains challenging. We introduce Outlier-Safe Pre-Training (OSP), a practical guideline that proactively prevents outlier formation, rather than relying on post-hoc mitigation. OSP combines three key innovations: (1) the Muon optimizer, eliminating privileged bases while maintaining training efficiency, (2) Single-Scale RMSNorm, preventing channel-wise amplification, and (3) a learnable embedding projection, redistributing activation magnitudes. We validate OSP by training a 1.4B-parameter model on 1 trillion tokens, which is the first production-scale LLM trained without such outliers. Under aggressive 4-bit quantization, our OSP model achieves a 35.7 average score across 10 benchmarks (versus 26.5 for an Adam-trained model), with only a 2% training overhead. Remarkably, OSP models exhibit near-zero excess kurtosis (0.04) compared to extreme values (1818.56) in standard models, fundamentally altering LLM quantization behavior. Our work demonstrates that outliers are not inherent to LLMs but are consequences of training strategies, paving the way for more efficient LLM deployment.
Jungwoo Park, Taewhoo Lee, Chanwoong Yoon, Hyeon Hwang, Jaewoo Kang
ACL (1)1
2025 Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
abstract
While the ability of language models to elicit facts has been widely investigated, how they handle temporally changing facts remains underexplored. We discover Temporal Heads, specific attention heads that primarily handle temporal knowledge, through circuit analysis. We confirm that these heads are present across multiple models, though their specific locations may vary, and their responses differ depending on the type of knowledge and its corresponding years. Disabling these heads degrades the model’s ability to recall time-specific knowledge while maintaining its general capabilities without compromising time-invariant and question-answering performances. Moreover, the heads are activated not only numeric conditions (“In 2004”) but also textual aliases (“In the year ...”), indicating that they encode a temporal dimension beyond simple numerical representation. Furthermore, we expand the potential of our findings by demonstrating how temporal knowledge can be edited by adjusting the values of these heads.
Yein Park, Chanwoong Yoon, Jungwoo Park, Minbyul Jeong, Jaewoo Kang
ACL (1)3
2025 Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
abstract
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yong Hoe Koo, Ko Minhyeok, Qingyu Chen, Mark Gerstein, Michael Moor, Jaewoo Kang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yonghoe Koo, Minhyeok Ko, Qingyu Chen 0001, Mark Gerstein, Michael Moor, Jaewoo Kang
EMNLP3
2025 Monet: Mixture of Monosemantic Experts for Transformers
abstract
Understanding the internal computations of large language models (LLMs) is crucial for aligning them with human values and preventing undesirable behaviors like toxic content generation. However, mechanistic interpretability is hindered by *polysemanticity*—where individual neurons respond to multiple, unrelated concepts. While Sparse Autoencoders (SAEs) have attempted to disentangle these features through sparse dictionary learning, they have compromised LLM performance due to reliance on post-hoc reconstruction loss. To address this issue, we introduce **Mixture of Monosemantic Experts for Transformers (Monet)** architecture, which incorporates sparse dictionary learning directly into end-to-end Mixture-of-Experts pretraining. Our novel expert decomposition method enables scaling the expert count to 262,144 per layer while total parameters scale proportionally to the square root of the number of experts. Our analyses demonstrate mutual exclusivity of knowledge across experts and showcase the parametric knowledge encapsulated within individual experts. Moreover, **Monet** allows knowledge manipulation over domains, languages, and toxicity mitigation without degrading general performance. Our pursuit of transparent LLMs highlights the potential of scaling expert counts to enhance mechanistic interpretability and directly resect the internal knowledge to fundamentally adjust model behavior.
Jungwoo Park, Ahn Young Jin, Kee-Eung Kim, Jaewoo Kang
ICLR1
2025 ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
abstract
Large language models (LLMs) have brought significant changes to many aspects of our lives. However, assessing and ensuring their chronological knowledge remains challenging. Existing approaches fall short in addressing the temporal adaptability of knowledge, often relying on a fixed time-point view. To overcome this, we introduce ChroKnowBench, a benchmark dataset designed to evaluate chronologically accumulated knowledge across three key aspects: multiple domains, time dependency, temporal state. Our benchmark distinguishes between knowledge that evolves (e.g., personal history, scientific discoveries, amended laws) and knowledge that remain constant (e.g., mathematical truths, commonsense facts). Building on this benchmark, we present ChroKnowledge (Chronological Categorization of Knowledge), a novel sampling-based framework for evaluating LLMs' non-parametric chronological knowledge. Our evaluation led to the following observations: (1) The ability of eliciting temporal knowledge varies depending on the data format that model was trained on. (2) LLMs partially recall knowledge or show a cut-off at temporal boundaries rather than recalling all aspects of knowledge correctly. Thus, we apply our ChroKnowPrompt, an in-depth prompting to elicit chronological knowledge by traversing step-by-step through the surrounding time spans. We observe that it successfully recalls objects across both open-source and proprietary LLMs, demonstrating versatility, though it faces challenges with dynamic datasets and unstructured formats.
Yein Park, Chanwoong Yoon, Jungwoo Park, Minbyul Jeong, Jaewoo Kang
ICLR3
2024 SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
abstract
Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip gestures that represent different phonemes. Prior approaches have sought to distinguish fine-grained visemes by aligning visual and auditory semantics, but often fell short of full synchronization. To address this, we present SyncVSR, an end-to-end learning framework that leverages quantized audio for frame-level crossmodal supervision. By integrating a projection layer that synchronizes visual representation with acoustic data, our encoder learns to generate discrete audio tokens from a video sequence in a non-autoregressive manner. SyncVSR shows versatility across tasks, languages, and modalities at the cost of a forward pass. Our empirical evaluations show that it not only achieves state-of-the-art results but also reduces data usage by up to ninefold.
Youngjin Ahn, Jungwoo Park, Sangha Park, Kee-Eung Kim
INTERSPEECH2
2020 SALE: Smartly Allocating Low-Cost Many-Bit ECC for Mitigating Read and Write Errors in STT-RAM Caches
abstract
Spin-transfer torque RAM (STT-RAM) is a future technology for ON-chip caches. However, it suffers from high read and write error rates. Concurrently dealing with these errors is quite challenging and incurs large performance overhead. This article proposes a smartly allocating low-cost many-bit ECC (SALE) scheme, which makes use of the low-cost many-bit error correction coding (ECC) to overcome this performance overhead. The low-cost many-bit ECC can fix many errors with low logic complexity and latency overheads. However, it requires a large number of parity bits. Therefore, SALE smartly uses low-cost many-bit ECC for only a certain type of cache lines and manages the corresponding large number of parity bits in the data array. SALE also introduces an ECC-free partition to reduce the ECC storage requirement for the STT-RAM caches. The cache lines belonging to an ECC-free partition do not have dedicated storage space for the ECC parity bits, thereby reducing the ECC storage requirement for the STT-RAM caches. Our experimental results demonstrate that SALE achieves performance close to that of an error-free cache by improving performance by 13% (16%) over the baseline scheme in single-core (quad-core) systems while requiring 50% less storage space for the ECC parity bits.
Muhammad Avais Qureshi, Jungwoo Park, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2019 MH Cache: A Mult Stephen Jarvisi-retention STT-RAM-based Low-power Last-level Cache for Mobile Hardware Rendering Systems
abstract
Mobile devices have become the most important devices in our life. However, they are limited in battery capacity. Therefore, low-power computing is crucial for their long lifetime. A spin-transfer torque RAM (STT-RAM) has become emerging memory technology because of its low leakage power consumption. We herein propose MH cache, a multi-retention STT-RAM-based cache management scheme for last-level caches (LLC) to reduce their power consumption for mobile hardware rendering systems. We analyzed the memory access patterns of processes and observed how rendering methods affect process behaviors. We propose a cache management scheme that measures write-intensity of each process dynamically and exploits it to manage a power-efficient multi-retention STT-RAM-based cache. Our proposed scheme uses variable threshold for a process’ write-intensity to determine cache line placement. We explain how to deal with the following issue to implement our proposed scheme. Our experimental results show that our techniques significantly reduce the LLC power consumption by 32% and 32.2% in single- and quad-core systems, respectively, compared to a full STT-RAM LLC.
Jungwoo Park, Myoungjun Lee, Soontae Kim, Minho Ju, Jeongkyu Hong
ACM Trans. Archit. Code Optim.1
2017 A Way-Filtering-Based Dynamic Logical-Associative Cache Architecture for Low-Energy Consumption
abstract
Last-level caches (LLCs) help improve performance but suffer from energy overhead because of their large sizes. An effective solution to this problem is to selectively power down several cache ways, which, however, reduces cache associativity and performance and thus limits its effectiveness in reducing energy consumption. To overcome this limitation, we propose a new cache architecture that can logically increase cache associativity of way-powered-down LLCs. Our proposed scheme is designed to be dynamic in activating an appropriate number of cache ways in order to eliminate the need for static profiling to determine an energy-optimized cache configuration. The experimental results show that our proposed dynamic scheme reduces the energy consumption of LLCs by 34% and 40% on single- and dual-core systems, respectively, compared with the best performing conventional static cache configuration. The overall system energy consumption including CPU, L2 cache, and DRAM is reduced by 9.2% on quad-core systems.
Jungwoo Park, Jongmin Lee 0002, Soontae Kim
IEEE Trans. Very Large Scale Integr. Syst.1