Jaemin Jung

dblp:09/5571 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Exploring Dynamic Memory Allocation of CXL Memory Pools in Enterprise In-Memory Database Management Systems
Donghun Lee 0001, Minseon Ahn, Jaemin Jung, Norman May, Daniel Ritter 0001, Heekwon Park, Changho Choi, Yang-Seok Ki
EDBT4
2026 SCORE: Scaling Audio Generation Using Standardized COmposite REwards
abstract
The goal of this letter is to enhance Text-to-Audio (T2A) generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing approaches often fail to achieve a reliable balance between semantic alignment and perceptual quality. We address this by adopting Inference-Time Scaling for the first time in T2A generation, and propose a multi-reward guidance that places each criterion on equal footing. By standardizing each reward to zero mean and unit variance before a weighted summation, the method reduces the scale mismatch that otherwise demands weight tuning, providing stable guidance at equal weights while leaving the weight as an interpretable control over the desired aspect. Moreover, we introduce a new text-audio alignment metric using an audio language model for more robust evaluation. Empirically, our method improves both semantic alignment and perceptual quality, achieving improved performance over naive generation and existing reward guidance techniques.
Jaemin Jung, Jaehun Kim, Inkyu Shin, Joon Son Chung
IEEE Signal Process. Lett.1
2025 Test-Time Augmentation for Pose-invariant Face Recognition
abstract
The goal of this paper is to enhance face recognition performance by augmenting head poses during the testing phase. Existing methods often rely on training on frontalised images or learning pose-invariant representations, yet both approaches typically require re-training and testing for each dataset, involving a substantial amount of effort. In contrast, this study proposes Pose-TTA, a novel approach that aligns faces at inference time without additional training. To achieve this, we employ a portrait animator that transfers the source image identity into the pose of a driving image. Instead of frontalising a side-profile face - which can introduce distortion - Pose-TTA generates matching side-profile images for comparison, thereby reducing identity information loss. Furthermore, we propose a weighted feature aggregation strategy to address any distortions or biases arising from the synthetic data, thus enhancing the reliability of the augmented images. Extensive experiments on diverse datasets and with various pre-trained face recognition models demonstrate that PoseTTA consistently improves inference performance. Moreover, our method is straightforward to integrate into existing face recognition pipelines, as it requires no retraining or fine-tuning of the underlying recognition models.
Jaemin Jung, Youngjoon Jang 0001, Joon Son Chung
FG1
2025 VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
abstract
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the field. To address this, we present a novel audio generation pipeline named VoiceDiT. This pipeline includes three key components: (1) the creation of a large-scale synthetic speech dataset for pre-training and a refined real-world speech dataset for fine-tuning, (2) the Dual-DiT, a model designed to efficiently preserve aligned speech information while accurately reflecting environmental conditions, and (3) a diffusion-based Image-to-Audio Translator that allows the model to bridge the gap between audio and image, facilitating the generation of environmental sound that aligns with the multi-modal prompts. Extensive experimental results demonstrate that VoiceDiT outperforms previous models on real-world datasets, showcasing significant improvements in both audio quality and modality integration. Synthesized samples are available on our demo page: https://mm.kaist.ac.kr/projects/voicedit/
Jaemin Jung, Junseok Ahn, Chaeyoung Jung, Tan Dat Nguyen, Youngjoon Jang 0001, Joon Son Chung
ICASSP1
2024 VoxMM: Rich Transcription of Conversations in the Wild
abstract
This paper presents a multi-modal dataset that contains rich transcriptions of spoken conversations. As diverse multi-modal and multi-task models emerge, there is a growing need for multi-modal training and evaluation datasets accompanied by rich metadata. However, there is no universal dataset that addresses these requirements for the diverse tasks partially due to the cost of annotation. To overcome this limitation, we develop a semi-automatic pipeline that makes the annotation more feasible. The resulting dataset is VoxMM, a multi-modal, multi-domain dataset. VoxMM incorporates video, audio, and text modalities. In terms of labels, it offers a wide array of metadata such as speaker labels, transcriptions, gender, and more. VoxMM supports both the training and the evaluation of any-to-any modality mapping models. It also offers a more accurate representation of real-world scenarios, bridging the gap between controlled laboratory experiments and the varying performances in the real-world. We present initial benchmarks on automatic speech recognition and speaker diarisation. The VoxMM dataset can be downloaded from https://mm.kaist.ac.kr/projects/voxmm
Doyeop Kwak, Jaemin Jung, Kihyun Nam, Youngjoon Jang 0001, Jee-Weon Jung, Shinji Watanabe 0001, Joon Son Chung
ICASSP2
2024 Artificial Intelligence-Driven Video Indexing for Rapid Surveillance Footage Summarization and Review
Jaemin Jung, Soonyong Park, Harim Kim, Chang Ha Lee, Charmgil Hong
IJCAI1
2024 Bridging the Gap Between Audio and Text Using Parallel-Attention for User-Defined Keyword Spotting
abstract
This letter proposes a novel user-defined keyword spotting framework that accurately detects audio keywords based on text enrollment. Since audio data possesses additional acoustic information compared to text, there are discrepancies between these two modalities. To address this challenge, we present ParallelKWS, which utilises self- and cross-attention in a parallel architecture to effectively capture information both within and across the two modalities. We further propose a phoneme duration-based alignment loss that enforces the sequential correspondence between audio and text features. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art performance on several benchmark datasets in both seen and unseen domains, without incorporating extra data beyond the dataset used in previous studies.
Youkyum Kim, Jaemin Jung, Byeong-Yeol Kim, Joon Son Chung
IEEE Signal Process. Lett.2
2023 Metric Learning for User-Defined Keyword Spotting
abstract
The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits their transferability to unseen terms. The ability to define custom keywords has advantages in terms of user experience.In this paper, we propose a metric learning-based training strategy for user-defined keyword spotting. In particular, we make the following contributions: (1) we construct a large-scale keyword dataset with an existing speech corpus and propose a filtering method to remove data that degrade model training; (2) we propose a metric learning-based two-stage training strategy, and demonstrate that the proposed method improves the performance on the user-defined keyword spotting task by enriching their representations; (3) to facilitate the fair comparison in the user-defined KWS field, we propose unified evaluation protocol and metrics.Our proposed system does not require an incremental training on the user-defined keywords, and outperforms previous works by a significant margin on the Google Speech Commands dataset using the proposed as well as the existing metrics.
Jaemin Jung, Youkyum Kim, Youshin Lim, Byeong-Yeol Kim, Youngjoon Jang 0001, Joon Son Chung
ICASSP1
2022 Enabling CXL Memory Expansion for In-Memory Database Management Systems
abstract
Limited memory volume is always a performance bottleneck in an in-memory database management system (IMDBMS) as the data size keeps increasing. To overcome the physical memory limitation, heterogeneous and disaggregated computing platforms are proposed, such as Gen-Z, CCIX, OpenCAPI, and CXL. In this work, we introduce flexible CXL memory expansion using a CXL type 3 prototype and evaluate its performance in an IMDBMS. Our evaluation shows that CXL memory devices interfaced with PCIe Gen5 are appropriate for memory expansion with nearly no throughput degradation in OLTP workloads and less than 8% throughput degradation in OLAP workloads. Thus, CXL memory is a good candidate for memory expansion with lower TCO in IMDBMSs.
Minseon Ahn, Donghun Lee 0001, Jaemin Jung, Oliver Rebholz, Vincent Pham, Krishna T. Malladi, Yang-Seok Ki
DaMoN6
2020 Optimizing Data Movement with Near-Memory Acceleration of In-memory DBMS
Donghun Lee 0001, Minseon Ahn, Jaemin Jung, Kang-Woo Choi, Vincent Pham, Oliver Rebholz, Krishna T. Malladi, Yang-Seok Ki
EDBT6
2020 Virtualize and share non-volatile memories in user space
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt
CCF Trans. High Perform. Comput.2
2019 vNVML: An Efficient User Space Library for Virtualizing and Sharing Non-Volatile Memories
abstract
The emerging non-volatile memory (NVM) has attractive characteristics such as DRAM-like, low-latency together with the non-volatility of storage devices. Recently, byte-addressable, memory bus-attached NVM has become available. This paper addresses the problem of combining a smaller, faster byte-addressable NVM with a larger, slower storage device, like SSD, to create the impression of a larger and faster byte-addressable NVM which can be shared across many applications. In this paper, we propose vNVML, a user space library for virtualizing and sharing NVM. vNVML provides for applications transaction like memory semantics that ensures write ordering and persistency guarantees across system failures. vNVML exploits DRAM for read caching, to enable improvements in performance and potentially to reduce the number of writes to NVM, extending the NVM lifetime. vNVML is implemented and evaluated with realistic workloads to show that our library allows applications to share NVM, both in a single O/S and when docker like containers are employed. The results from the evaluation show that vNVML incurs less than 10% overhead while providing the benefits of an expanded virtualized NVM space to the applications, allowing applications to safely share the virtual NVM.
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt
MSST2
2018 Barrier-Enabled IO Stack for Flash Storage
Youjip Won, Jaemin Jung, Gyeongyeol Choi, Joontaek Oh, Seongbae Son, Joo Young Hwang, Sangyeun Cho
FAST2
2018 Bringing Order to Chaos: Barrier-Enabled I/O Stack for Flash Storage
abstract
This work is dedicated to eliminating the overhead required for guaranteeing the storage order in the modern IO stack. The existing block device adopts a prohibitively expensive approach in ensuring the storage order among write requests: interleaving the write requests with Transfer-and-Flush . For exploiting the cache barrier command for flash storage, we overhaul the IO scheduler, the dispatch module, and the filesystem so that these layers are orchestrated to preserve the ordering condition imposed by the application with which the associated data blocks are made durable. The key ingredients of Barrier-Enabled IO stack are Epoch-based IO scheduling , Order-Preserving Dispatch , and Dual-Mode Journaling . Barrier-enabled IO stack can control the storage order without Transfer-and-Flush overhead. We implement the barrier-enabled IO stack in server as well as in mobile platforms. SQLite performance increases by 270% and 75%, in server and in smartphone, respectively. In a server storage, BarrierFS brings as much as by 43 × and by 73× performance gain in MySQL and SQLite, respectively, against EXT4 via relaxing the durability of a transaction.
Youjip Won, Joontaek Oh, Jaemin Jung, Gyeongyeol Choi, Seongbae Son, Joo Young Hwang, Sangyeun Cho
ACM Trans. Storage3
2016 nvramdisk: A Transactional Block Device Driver for Non-Volatile RAM
abstract
In this work, we developed nvramdisk, a transactional block device driver for byte-addressable NVRAM. nvramdisk effectively addresses the key technical challenges in using a section of NVRAM as a transactional persistent block device. nvramdisk adopts (i) shadow block, (ii) mapping table journaling, and (iii) type-dependent ordering guarantee to provide atomicity, consistency, integrity and durability in write operations on nvramdisk imposed block device. We fully implemented nvramdisk device driver on Linux OS and port it on the desktop computer as well as Android smartphones. In memcachedb, locating the database table in nvramdisk brings ×1.9 insertions/sec and updates/sec performance gain against locating the database table in a high-end SSD (FusionIO ioDrive2). SQLite performance increases by ×2.9, from 743 ins/sec to 2,184 ins/sec, in smartphone(Samsung Galaxy S4) and ×15, from 730 ins/sec to 12390 ins/sec in PC. nvramdisk yields 26 percent higher random write performance against Persistent Memory Block Driver. The overhead of supporting transaction accompanies 6 percent performance penalty in memcachedb operations.
Jaemin Jung, Youjip Won
IEEE Trans. Computers1
2015 HEAPO: Heap-Based Persistent Object Store
abstract
In this work, we developed a Heap-Based Persistent Object Store (HEAPO) to manage persistent objects in byte-addressable Nonvolatile RAM (NVRAM). HEAPO defines its own persistent heap layout, the persistent object format, name space organization, object sharing and protection mechanism, and undo-only log-based crash recovery, all of which are effectively tailored for NVRAM. We put our effort into developing a lightweight and flexible layer to exploit the DRAM-like access latency of NVRAM. To address this objective, we developed (i) a native management layer for NVRAM to eliminate redundancy between in-core and on-disk copies of the metadata, (ii) an expandable object format, (iii) a burst trie-based global name space with local name space caching, (iv) static address binding, and (v) minimal logging for undo-only crash recovery. We implemented HEAPO at commodity OS (Linux 2.6.32) and measured the performance. By eliminating metadata redundancy, HEAPO improved the speed of creating, attaching, and expanding an object by 1.3×, 4.5×, and 3.8×, respectively, compared to memory-mapped file-based persistent object store. Burst trie-based name space organization of HEAPO yielded 7.6× better lookup performance compared to hashed B-tree-based name space of EXT4. We modified memcachedb to use HEAPO in maintaining its search structure. For hash table update, HEAPO-based memcachedb yielded 3.4× performance improvement against original memcachedb implementation which uses mmap() over ramdisk approach to maintain the key-value store in memory.
Taeho Hwang, Jaemin Jung, Youjip Won
ACM Trans. Storage2
2010 FRASH: Exploiting storage class memory in hybrid file system for hierarchical storage
abstract
In this work, we develop a novel hybrid file system, FRASH, for storage-class memory and NAND Flash. Despite the promising physical characteristics of storage-class memory, its scale is an order of magnitude smaller than the current storage device scale. This fact makes it less than desirable for use as an independent storage device. We carefully analyze in-memory and on-disk file system objects in a log-structured file system, and exploit memory and storage aspects of the storage-class memory to overcome the drawbacks of the current log-structured file system. FRASH provides a hybrid view storage-class memory. It harbors an in-memory data structure as well as a on-disk structure. It provides nonvolatility to key data structures which have been maintained in-memory in a legacy log-structured file system. This approach greatly improves the mount latency and effectively resolves the robustness issue. By maintaining on-disk structure in storage-class memory, FRASH provides byte-addressability to the file system object and metadata for page, and subsequently greatly improves the I/O performance compared to the legacy log-structured approach. While storage-class memory offers byte granularity, it is still far slower than its DRAM counter part. We develop a copy-on-mount technique to overcome the access latency difference between main memory and storage-class memory. Our file system was able to reduce the mount time by 92% and file system I/O performance was increased by 16%.
Jaemin Jung, Youjip Won, Eun-ki Kim, Hyungjong Shin, Byeonggil Jeon
ACM Trans. Storage1
2007 FRASH: Hierarchical File System for FRAM and Flash
Eun-ki Kim, Hyungjong Shin, Byung-Gil Jeon, Seokhee Han, Jaemin Jung, Youjip Won
ICCSA (1)5