Hyojun Kim

dblp:89/2022 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
5since 2021 · last 2026
0000-0002-9210-5357ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 first-authorArtificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 38% Learning paradigms · 22% Vision and language · 19%
Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 48% Memory systems · 34% Embedded and real-time systems · 18%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 67% Operating systems · 33%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025
Natural language and speech › Language models and text generation › LLM agents
web agents
0.912025
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025
Security and privacy of machine learning
adversarial attack
0.612022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Security and privacy of machine learning › adversarial attack
textual adversarial attack
0.612022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Program synthesis and code generation › code model
code model robustness
0.612022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.512021
SS-IL: Separated Softmax for Incremental Learning · ICCV 2021
Machine learning › Learning paradigms › continual learning
class-incremental learning
0.512021
SS-IL: Separated Softmax for Incremental Learning · ICCV 2021
Memory systems › non-volatile memory
phase change memory
0.422014
Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches · ACM Trans. Storage 2014
Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches · FAST 2014
Embedded and real-time systems
mobile computing
0.232012
Revisiting storage for smartphones · ACM Trans. Storage 2012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Revisiting storage for smartphones · FAST 2012
Storage systems › flash and SSD › flash memory
flash storage
0.222012
Revisiting storage for smartphones · ACM Trans. Storage 2012
BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008
Storage systems › storage architecture
enterprise storage
0.212014
Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches · FAST 2014
Memory systems
non-volatile memory
0.212014
Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches · FAST 2014
Memory systems › cache management
storage caching
0.212014
Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches · ACM Trans. Storage 2014
Storage systems › storage hierarchy
tiered storage
0.212014
Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches · ACM Trans. Storage 2014
Embedded and real-time systems › mobile computing
smartphone storage
0.222012
Revisiting storage for smartphones · FAST 2012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Information retrieval › document retrieval › domain-specific retrieval
code search
0.212022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Information retrieval › document retrieval › domain-specific retrieval › code search
semantic code search
0.212022
TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.112021
SS-IL: Separated Softmax for Incremental Learning · ICCV 2021
Operating systems › resource management › memory management
buffer cache
0.112012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Operating systems › resource management › memory management
cache replacement
0.112012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Storage systems
flash and SSD
0.112012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Storage systems › storage devices › storage media
mobile storage
0.112012
Revisiting storage for smartphones · FAST 2012
Storage systems
buffer management
0.112008
BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008
Storage systems › flash and SSD
flash memory
0.012008
BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008
Storage systems › i/o optimization
write optimization
0.012008
BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008

Methods — techniques the papers use, named apart from their topics

contextual semantic filtering · 1.7beam search · 1.7process reward model · 0.9preference pairs · 0.9task-wise knowledge distillation · 0.5separated softmax · 0.5exemplar memory · 0.5trace-driven simulation · 0.3OS-level implementation · 0.3workload trace analysis · 0.2performance study · 0.2storage characterization · 0.1pilot solution design · 0.1measurement study · 0.1buffer management · 0.1
YearPublicationVenuePosition
2026 Efficient Bootstrapping in Fully Homomorphic Encryption for Matrix Arithmetic
Eric Crockett 0001, Craig Gentry, Hyojun Kim, Yeongmin Lee
CRYPTO (2)3
2025 SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
abstract
Suyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Suyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001
NAACL (Long Papers)3
2025 Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
abstract
Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks. Yet, specialized reward models for web navigation that can be utilized during both training and test-time have been absent until now. Despite the importance of speed and cost-effectiveness, prior works have utilized MLLMs as reward models, which poses significant constraints for real-world deployment. To address this, in this work, we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels. Next, we also introduce the WebRewardBench, the first meta-evaluation benchmark for evaluating PRMs. In our experiments, we observe that our Web-Shepherd achieves about 30 points better accuracy compared to using GPT-4o on WebRewardBench. Furthermore, when testing on WebArena-lite by using GPT-4o-mini as the policy and Web-Shepherd as the verifier, we achieve 10.9 points better performance, in 10x less cost compared to using GPT-4o-mini as the verifier. Our model, dataset, and code are publicly available at https://github.com/kyle8581/Web-Shepherd.
Hyungjoo Chae, Seonghwan Kim, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Beong-woo Kwak, Dongjin Kang, Jinyoung Yeo
NeurIPS15
2022 TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search
abstract
As pre-trained models have shown successful performance in program language processing as well as natural language processing, adversarial attacks on these models also attract attention.However, previous works on blackbox adversarial attacks generated adversarial examples in a very inefficient way with simple greedy search.They also failed to find out better adversarial examples because it was hard to reduce the search space without performance loss.In this paper, we propose TABS, an efficient beam search black-box adversarial attack method.We adopt beam search to find out better adversarial examples, and contextual semantic filtering to effectively reduce the search space.Contextual semantic filtering reduces the number of candidate adversarial words considering the surrounding context and the semantic similarity.Our proposed method shows good performance in terms of attack success rate, the number of queries, and semantic similarity in attacking models for two tasks: NL code search classification and retrieval tasks.
YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001
EMNLP2
2021 SS-IL: Separated Softmax for Incremental Learning
abstract
We consider class incremental learning (CIL) problem, in which a learning agent continuously learns new classes from incrementally arriving training data batches and aims to predict well on all the classes learned so far. The main challenge of the problem is the catastrophic forgetting, and for the exemplar-memory based CIL methods, it is generally known that the forgetting is commonly caused by the classification score bias that is injected due to the data imbalance between the new classes and the old classes (in the exemplar-memory). While several methods have been proposed to correct such score bias by some additional post-processing, e.g., score re-scaling or balanced fine-tuning, no systematic analysis on the root cause of such bias has been done. To that end, we analyze that computing the softmax probabilities by combining the output scores for all old and new classes could be the main cause of the bias. Then, we propose a new method, dubbed as Separated Softmax for Incremental Learning (SS-IL), that consists of separated softmax (SS) output layer combined with task-wise knowledge distillation (TKD) to resolve such bias. Throughout our extensive experimental results on several large-scale CIL benchmark datasets, we show our SS-IL achieves strong state-of-the-art accuracy through attaining much more balanced prediction scores across old and new classes, without any additional post-processing.
Hongjoon Ahn, Jihwan Kwak, Subin Lim, Hyeonsu Bang, Hyojun Kim, Taesup Moon
ICCV5
2018 A 3.2-GHz Supply Noise-Insensitive PLL Using a Gate-Voltage-Boosted Source-Follower Regulator and Residual Noise Cancellation
Youngwoo Jo, Hyojun Kim, SeongHwan Cho
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Finding Consistency in an Inconsistent World: Towards Deep Semantic Understanding of Scale-out Distributed Databases
Neville Carvalho, Hyojun Kim, Maohua Lu, Prasenjit Sarkar, Rohit Shekhar, Tarun Thakur, Pin Zhou, Remzi H. Arpaci-Dusseau
HotStorage2
2014 Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches
Hyojun Kim, Sangeetha Seshadri, Clem Dickey, Lawrence Chiu
FAST1
2014 Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches
abstract
Storage systems based on Phase Change Memory (PCM) devices are beginning to generate considerable attention in both industry and academic communities. But whether the technology in its current state will be a commercially and technically viable alternative to entrenched technologies such as flash-based SSDs remains undecided. To address this, it is important to consider PCM SSD devices not just from a device standpoint, but also from a holistic perspective. This article presents the results of our performance study of a recent all-PCM SSD prototype. The average latency for a 4KiB random read is 6.7μs, which is about 16× faster than a comparable eMLC flash SSD. The distribution of I/O response times is also much narrower than flash SSD for both reads and writes. Based on the performance measurements and real-world workload traces, we explore two typical storage use cases: tiering and caching. We report that the IOPS/$ of a tiered storage system can be improved by 12--66% and the aggregate elapsed time of a server-side caching solution can be improved by up to 35% by adding PCM. Our results show that (even at current price points) PCM storage devices show promising performance as a new component in enterprise storage systems.
Hyojun Kim, Sangeetha Seshadri, Clem Dickey, Lawrence Chiu
ACM Trans. Storage1
2013 Fjord: Informed storage management for smartphones
abstract
Smartphone applications are becoming more sophisticated and require high storage performance. Unfortunately, the OS storage software stack is not well engineered to support flash-based storage used in smartphones. On top of that, storage software stack is configured to be too conservative due to the fear of sudden power failures. We believe that this conservatism with respect to data reliability is misplaced considering that many of the popular apps (e.g., Web browsing, Facebook, Gmail) that run on today's smartphones are cloud-backed, and the local storage on the smartphone is often used as a cache for cloud data. In this paper, we propose Informed Storage Management framework, named Fjord, for mobile platforms. The key insight is to use system-wide dynamic context information to improve the storage performance on mobile platforms. We implement a set of mechanisms (write buffering, logging, and fine-grained reliability control), and through judicious use of these mechanisms based on system context, we show how we can achieve significant improvement in storage performance. As proof of concept, we implement Fjord in two Android smartphones and experimentally validate the performance advantage of informed storage management with multiple smartphone applications.
Hyojun Kim, Umakishore Ramachandran
MSST1
2012 Revisiting storage for smartphones
Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu
FAST1
2012 Why are state-of-the-art flash-based multi-tiered storage systems performing poorly for HTTP video streaming?
abstract
MLC flash memory is a promising technology for building a high-performance and cost-effective video streaming system when it is used as an intermediate level cache in a multi-tiered storage hierarchy. Therefore, we were quite surprised when through extensive measurements we found that two state-of-the-art flash-based multi-tiered storage systems (namely, flashcache and ZFS) have quite disappointing performance for HTTP video streaming using the DASH protocol. We have conducted a thorough analysis to understand the reasons for the poor performance of these two systems. In a nutshell, unless attention is paid to the unique performance characteristics of flash memory-based SSDs, we could end up with suboptimal or even poor performance as we discovered through experimentation with these two systems. Based on the analysis, we present design guidelines for building a cost-effective high-performance HTTP video streaming server.
Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran
NOSSDAV2
2012 What is a good buffer cache replacement scheme for mobile flash storage?
abstract
Smartphones are becoming ubiquitous and powerful. The Achilles' heel in such devices that limits performance is the storage. Low-end flash memory is the storage technology of choice in such devices due to energy, size, and cost considerations. In this paper, we take a critical look at the performance of flash on smartphones for mobile applications. Specifically, we ask the question whether the state-of-the-art buffer cache replacement schemes proposed thus far (both flash-agnostic and flash-aware ones) are the right ones for mobile flash storage. To answer this question, we first expose the limitations of current buffer cache performance evaluation methods, and propose a novel evaluation framework that is a hybrid between trace-driven simulation and real implementation of such schemes inside an operating system. Such an evaluation reveals some unexpected and surprising insights on the performance of buffer management schemes that contradicts conventional wisdom. Armed with this knowledge, we propose a new buffer cache replacement scheme called SpatialClock.
Hyojun Kim, Moonkyung Ryu, Umakishore Ramachandran
SIGMETRICS1
2012 Revisiting storage for smartphones
abstract
Conventional wisdom holds that storage is not a big contributor to application performance on mobile devices. Flash storage (the type most commonly used today) draws little power, and its performance is thought to exceed that of the network subsystem. In this article, we present evidence that storage performance does indeed affect the performance of several common applications such as Web browsing, maps, application install, email, and Facebook. For several Android smartphones, we find that just by varying the underlying flash storage, performance over WiFi can typically vary between 100% and 300% across applications; in one extreme scenario, the variation jumped to over 2000%. With a faster network (set up over USB), the performance variation rose even further. We identify the reasons for the strong correlation between storage and application performance to be a combination of poor flash device performance, random I/O from application databases, and heavy-handed use of synchronous writes. Based on our findings, we implement and evaluate a set of pilot solutions to address the storage performance deficiencies in smartphones.
Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu
ACM Trans. Storage1
2011 Impact of flash memory on video-on-demand storage: analysis of tradeoffs
abstract
There is no doubt that video-on-demand (VoD) services are very popular these days. However, disk storage is a serious bottleneck limiting the scalability of a VoD server. Disk throughput degrades dramatically due to seek time overhead when the server is called upon to serve a large number of simultaneous video streams. To address the performance problem of disk, buffer cache algorithms that utilize RAM have been proposed. Interval caching is a state-of-the-art caching algorithm for a VoD server. Flash Memory Solid-State Drive (SSD) is a relatively new storage technology. Its excellent random read performance, low power consumption, and sharply dropping cost per gigabyte are opening new opportunities to efficiently use the device for enterprise systems. On the other hand, it has deficiencies such as poor small random write performance and limited number of erase operations. In this paper, we analyze tradeoffs and potential impact that flash memory SSD can have for a VoD server. Performance of various commercially available flash memory SSD models is studied. We find that low-end flash memory SSD provides better performance than the high-end one while costing less than the high-end one when the I/O request size is large, which is typical for a VoD server. Because of the wear problem and asymmetric read/write performance of flash memory SSD, we claim that interval caching cannot be used with it. Instead, we propose using file-level Least Frequently Used (LFU) due to the highly skewed video access pattern of the VoD workload. We compare the performance of interval caching with RAM and file-level LFU with flash memory by simulation experiments. In addition, from the cost-effectiveness analysis of three different storage configurations, we find that flash memory with hard disk drive is the most cost-effective solution compared to DRAM with hard disk drive or hard disk drive only.
Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran
MMSys2
2009 FlashLite: A User-Level Library to Enhance Durability of SSD for P2P File Sharing
abstract
Peer-to-peer file sharing is popular, but it generates random write traffic to storage due to the nature of swarming. NAND flash memory based Solid-State Drive (SSD) technology is available as an alternative to hard drives for notebook and tablet PCs. As it turns out, random write is extremely detrimental to the lifetime of SSD drives. This paper focuses on the following problem, namely, P2P file downloading when the target of the download is an SSD drive. We make three contributions: first, analysis of write patterns of downloading program to establish the premise of the problem; second, development of a simple yet powerful technique called FlashLite to combat this problem, by automatically converting the random writes to sequential writes; third, showing through performance evaluation using modified eMule file downloading program that FlashLite does change random writes to sequential, and most importantly eliminates about 94% of erase operations of the original eMule program.
Hyojun Kim, Umakishore Ramachandran
ICDCS1
2009 Design and implementation of MLC NAND flash-based DBMS for mobile devices
Ki Yong Lee, Hyojun Kim, Kyoung-Gu Woo, Yon Dohn Chung, Myoung-Ho Kim
J. Syst. Softw.2
2008 BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage
Hyojun Kim, Seongjun Ahn
FAST1
2007 A Page Padding Method for Fragmented Flash Storage
Hyojun Kim, Jin-Hyuk Kim, ShinHo Choi, HyunRyong Jung, JaeGyu Jung
ICCSA (1)1
2007 SWL: a search-while-load demand paging scheme with NAND flash memory
abstract
As mobile phones become increasingly multifunctional, the number and size of applications installed in phones are rapidly increasing. Consequently, mobile phones require more hardware resources such as NOR/NAND flash memory and DRAM, and their production cost is accordingly becoming higher. One candidate solution to reduce production cost is demand paging using MMU. However, demand paging causes unpredictably long page fault latency, and as such mobile phone manufacturers are reluctant to deploy this scheme. In this paper, we present a method that reduces the long latency of page faults by performing page fault handling in a parallelized manner, considering the characteristics of NAND-Type flash memory. We also discuss how to modify the existing page cache replacement policies so that they can exploit the benefits of the parallelized page fault handler. Experimental results show that the parallelized page fault handler improves the worst case latency of page faults significantly, by up to roughly 20%, and that the modified page cache replacement policies improve both the average and worst instruction fetch time.
Jihyun In, Ilhoon Shin, Hyojun Kim
LCTES3
2006 MNFS: mobile multimedia file system for NAND flash based storage device
abstract
In this work, we present a novel mobile multimedia file system, MNFS, which is specifically designed for NAND flash memory. It is designed for mobile multimedia devices such as MP3 player, Personal Media Player (PMP), digital camcorder, etc. Our file system has three novel features important in mobile multimedia applications: (1) predictable and uniform write latency, (2) quick file system mount, and (3) small memory footprint. We implement the proto-type file system on ARM9 embedded platform. In experiments, MNFS exhibits uniform I/O latency for sequential write operation. It is mountable within 0.2 seconds, and available with only 34Kbytes heap memory for 128Mbytes volume. Compared to YAFFS which is the de facto standard file system for NAND flash memory, the mounting time is 30 times faster and the heap memory usage is only 5% of YAFFS usage.
Hyojun Kim, Youjip Won
CCNC1