EDBT 2026 Demo / reviewers in the wild / expert
Hyojun Kim
dblp:89/2022
· DBLP profile ↗
21ranked-venue papers
10as first author
5since 2021 · last 2026
0000-0002-9210-5357ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-authorArtificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 38% Learning paradigms · 22% Vision and language · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Storage systems · 48% Memory systems · 34% Embedded and real-time systems · 18% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Software engineering, system software, and programming languages
2 papers |
Program synthesis and code generation · 67% Operating systems · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model |
0.9 | 1 | 2025 | Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025 |
Natural language and speech › Language models and text generation › LLM agents
web agents |
0.9 | 1 | 2025 | Web-Shepherd: Advancing PRMs for Reinforcing Web Agents · NeurIPS 2025 |
Security and privacy of machine learning
adversarial attack |
0.6 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Security and privacy of machine learning › adversarial attack
textual adversarial attack |
0.6 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Program synthesis and code generation › code model
code model robustness |
0.6 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.5 | 1 | 2021 | SS-IL: Separated Softmax for Incremental Learning · ICCV 2021 |
Machine learning › Learning paradigms › continual learning
class-incremental learning |
0.5 | 1 | 2021 | SS-IL: Separated Softmax for Incremental Learning · ICCV 2021 |
Memory systems › non-volatile memory
phase change memory |
0.4 | 2 | 2014 | Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches · ACM Trans. Storage 2014 Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches · FAST 2014 |
Embedded and real-time systems
mobile computing |
0.2 | 3 | 2012 | Revisiting storage for smartphones · ACM Trans. Storage 2012 What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012 Revisiting storage for smartphones · FAST 2012 |
Storage systems › flash and SSD › flash memory
flash storage |
0.2 | 2 | 2012 | Revisiting storage for smartphones · ACM Trans. Storage 2012 BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008 |
Storage systems › storage architecture
enterprise storage |
0.2 | 1 | 2014 | Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches · FAST 2014 |
Memory systems
non-volatile memory |
0.2 | 1 | 2014 | Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches · FAST 2014 |
Memory systems › cache management
storage caching |
0.2 | 1 | 2014 | Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches · ACM Trans. Storage 2014 |
Storage systems › storage hierarchy
tiered storage |
0.2 | 1 | 2014 | Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering Approaches · ACM Trans. Storage 2014 |
Embedded and real-time systems › mobile computing
smartphone storage |
0.2 | 2 | 2012 | Revisiting storage for smartphones · FAST 2012 What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012 |
Information retrieval › document retrieval › domain-specific retrieval
code search |
0.2 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Information retrieval › document retrieval › domain-specific retrieval › code search
semantic code search |
0.2 | 1 | 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam Search · EMNLP 2022 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.1 | 1 | 2021 | SS-IL: Separated Softmax for Incremental Learning · ICCV 2021 |
Operating systems › resource management › memory management
buffer cache |
0.1 | 1 | 2012 | What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012 |
Operating systems › resource management › memory management
cache replacement |
0.1 | 1 | 2012 | What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012 |
Storage systems
flash and SSD |
0.1 | 1 | 2012 | What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012 |
Storage systems › storage devices › storage media
mobile storage |
0.1 | 1 | 2012 | Revisiting storage for smartphones · FAST 2012 |
Storage systems
buffer management |
0.1 | 1 | 2008 | BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008 |
Storage systems › flash and SSD
flash memory |
0.0 | 1 | 2008 | BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008 |
Storage systems › i/o optimization
write optimization |
0.0 | 1 | 2008 | BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage · FAST 2008 |
Methods — techniques the papers use, named apart from their topics
contextual semantic filtering · 1.7beam search · 1.7process reward model · 0.9preference pairs · 0.9task-wise knowledge distillation · 0.5separated softmax · 0.5exemplar memory · 0.5trace-driven simulation · 0.3OS-level implementation · 0.3workload trace analysis · 0.2performance study · 0.2storage characterization · 0.1pilot solution design · 0.1measurement study · 0.1buffer management · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Bootstrapping in Fully Homomorphic Encryption for Matrix Arithmetic
Eric Crockett 0001, Craig Gentry, Hyojun Kim, Yeongmin Lee |
CRYPTO (2) | 3 |
| 2025 | SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented DataabstractSuyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Suyoung Bae, YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001 |
NAACL (Long Papers) | 3 |
| 2025 | Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsabstractWeb navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks.
Yet, specialized reward models for web navigation that can be utilized during both training and test-time have been absent until now. Despite the importance of speed and cost-effectiveness, prior works have utilized MLLMs as reward models, which poses significant constraints for real-world deployment. To address this, in this work, we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels. Next, we also introduce the WebRewardBench, the first meta-evaluation benchmark for evaluating PRMs. In our experiments, we observe that our Web-Shepherd achieves about 30 points better accuracy compared to using GPT-4o on WebRewardBench.
Furthermore, when testing on WebArena-lite by using GPT-4o-mini as the policy and Web-Shepherd as the verifier, we achieve 10.9 points better performance, in 10x less cost compared to using GPT-4o-mini as the verifier.
Our model, dataset, and code are publicly available at https://github.com/kyle8581/Web-Shepherd. Hyungjoo Chae, Seonghwan Kim, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Beong-woo Kwak, Dongjin Kang, Jinyoung Yeo |
NeurIPS | 15 |
| 2022 | TABS: Efficient Textual Adversarial Attack for Pre-trained NL Code Model Using Semantic Beam SearchabstractAs pre-trained models have shown successful performance in program language processing as well as natural language processing, adversarial attacks on these models also attract attention.However, previous works on blackbox adversarial attacks generated adversarial examples in a very inefficient way with simple greedy search.They also failed to find out better adversarial examples because it was hard to reduce the search space without performance loss.In this paper, we propose TABS, an efficient beam search black-box adversarial attack method.We adopt beam search to find out better adversarial examples, and contextual semantic filtering to effectively reduce the search space.Contextual semantic filtering reduces the number of candidate adversarial words considering the surrounding context and the semantic similarity.Our proposed method shows good performance in terms of attack success rate, the number of queries, and semantic similarity in attacking models for two tasks: NL code search classification and retrieval tasks. YunSeok Choi, Hyojun Kim, Jee-Hyong Lee 0001 |
EMNLP | 2 |
| 2021 | SS-IL: Separated Softmax for Incremental LearningabstractWe consider class incremental learning (CIL) problem, in which a learning agent continuously learns new classes from incrementally arriving training data batches and aims to predict well on all the classes learned so far. The main challenge of the problem is the catastrophic forgetting, and for the exemplar-memory based CIL methods, it is generally known that the forgetting is commonly caused by the classification score bias that is injected due to the data imbalance between the new classes and the old classes (in the exemplar-memory). While several methods have been proposed to correct such score bias by some additional post-processing, e.g., score re-scaling or balanced fine-tuning, no systematic analysis on the root cause of such bias has been done. To that end, we analyze that computing the softmax probabilities by combining the output scores for all old and new classes could be the main cause of the bias. Then, we propose a new method, dubbed as Separated Softmax for Incremental Learning (SS-IL), that consists of separated softmax (SS) output layer combined with task-wise knowledge distillation (TKD) to resolve such bias. Throughout our extensive experimental results on several large-scale CIL benchmark datasets, we show our SS-IL achieves strong state-of-the-art accuracy through attaining much more balanced prediction scores across old and new classes, without any additional post-processing. Hongjoon Ahn, Jihwan Kwak, Subin Lim, Hyeonsu Bang, Hyojun Kim, Taesup Moon |
ICCV | 5 |
| 2018 | A 3.2-GHz Supply Noise-Insensitive PLL Using a Gate-Voltage-Boosted Source-Follower Regulator and Residual Noise Cancellation
Youngwoo Jo, Hyojun Kim, SeongHwan Cho |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Finding Consistency in an Inconsistent World: Towards Deep Semantic Understanding of Scale-out Distributed Databases
Neville Carvalho, Hyojun Kim, Maohua Lu, Prasenjit Sarkar, Rohit Shekhar, Tarun Thakur, Pin Zhou, Remzi H. Arpaci-Dusseau |
HotStorage | 2 |
| 2014 | Evaluating phase change memory for enterprise storage systems: a study of caching and tiering approaches
Hyojun Kim, Sangeetha Seshadri, Clem Dickey, Lawrence Chiu |
FAST | 1 |
| 2014 | Evaluating Phase Change Memory for Enterprise Storage Systems: A Study of Caching and Tiering ApproachesabstractStorage systems based on Phase Change Memory (PCM) devices are beginning to generate considerable attention in both industry and academic communities. But whether the technology in its current state will be a commercially and technically viable alternative to entrenched technologies such as flash-based SSDs remains undecided. To address this, it is important to consider PCM SSD devices not just from a device standpoint, but also from a holistic perspective. This article presents the results of our performance study of a recent all-PCM SSD prototype. The average latency for a 4KiB random read is 6.7μs, which is about 16× faster than a comparable eMLC flash SSD. The distribution of I/O response times is also much narrower than flash SSD for both reads and writes. Based on the performance measurements and real-world workload traces, we explore two typical storage use cases: tiering and caching. We report that the IOPS/$ of a tiered storage system can be improved by 12--66% and the aggregate elapsed time of a server-side caching solution can be improved by up to 35% by adding PCM. Our results show that (even at current price points) PCM storage devices show promising performance as a new component in enterprise storage systems. Hyojun Kim, Sangeetha Seshadri, Clem Dickey, Lawrence Chiu |
ACM Trans. Storage | 1 |
| 2013 | Fjord: Informed storage management for smartphonesabstractSmartphone applications are becoming more sophisticated and require high storage performance. Unfortunately, the OS storage software stack is not well engineered to support flash-based storage used in smartphones. On top of that, storage software stack is configured to be too conservative due to the fear of sudden power failures. We believe that this conservatism with respect to data reliability is misplaced considering that many of the popular apps (e.g., Web browsing, Facebook, Gmail) that run on today's smartphones are cloud-backed, and the local storage on the smartphone is often used as a cache for cloud data. In this paper, we propose Informed Storage Management framework, named Fjord, for mobile platforms. The key insight is to use system-wide dynamic context information to improve the storage performance on mobile platforms. We implement a set of mechanisms (write buffering, logging, and fine-grained reliability control), and through judicious use of these mechanisms based on system context, we show how we can achieve significant improvement in storage performance. As proof of concept, we implement Fjord in two Android smartphones and experimentally validate the performance advantage of informed storage management with multiple smartphone applications. Hyojun Kim, Umakishore Ramachandran |
MSST | 1 |
| 2012 | Revisiting storage for smartphones
Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu |
FAST | 1 |
| 2012 | Why are state-of-the-art flash-based multi-tiered storage systems performing poorly for HTTP video streaming?abstractMLC flash memory is a promising technology for building a high-performance and cost-effective video streaming system when it is used as an intermediate level cache in a multi-tiered storage hierarchy. Therefore, we were quite surprised when through extensive measurements we found that two state-of-the-art flash-based multi-tiered storage systems (namely, flashcache and ZFS) have quite disappointing performance for HTTP video streaming using the DASH protocol. We have conducted a thorough analysis to understand the reasons for the poor performance of these two systems. In a nutshell, unless attention is paid to the unique performance characteristics of flash memory-based SSDs, we could end up with suboptimal or even poor performance as we discovered through experimentation with these two systems. Based on the analysis, we present design guidelines for building a cost-effective high-performance HTTP video streaming server. Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran |
NOSSDAV | 2 |
| 2012 | What is a good buffer cache replacement scheme for mobile flash storage?abstractSmartphones are becoming ubiquitous and powerful. The Achilles' heel in such devices that limits performance is the storage. Low-end flash memory is the storage technology of choice in such devices due to energy, size, and cost considerations. In this paper, we take a critical look at the performance of flash on smartphones for mobile applications. Specifically, we ask the question whether the state-of-the-art buffer cache replacement schemes proposed thus far (both flash-agnostic and flash-aware ones) are the right ones for mobile flash storage. To answer this question, we first expose the limitations of current buffer cache performance evaluation methods, and propose a novel evaluation framework that is a hybrid between trace-driven simulation and real implementation of such schemes inside an operating system. Such an evaluation reveals some unexpected and surprising insights on the performance of buffer management schemes that contradicts conventional wisdom. Armed with this knowledge, we propose a new buffer cache replacement scheme called SpatialClock. Hyojun Kim, Moonkyung Ryu, Umakishore Ramachandran |
SIGMETRICS | 1 |
| 2012 | Revisiting storage for smartphonesabstractConventional wisdom holds that storage is not a big contributor to application performance on mobile devices. Flash storage (the type most commonly used today) draws little power, and its performance is thought to exceed that of the network subsystem. In this article, we present evidence that storage performance does indeed affect the performance of several common applications such as Web browsing, maps, application install, email, and Facebook. For several Android smartphones, we find that just by varying the underlying flash storage, performance over WiFi can typically vary between 100% and 300% across applications; in one extreme scenario, the variation jumped to over 2000%. With a faster network (set up over USB), the performance variation rose even further. We identify the reasons for the strong correlation between storage and application performance to be a combination of poor flash device performance, random I/O from application databases, and heavy-handed use of synchronous writes. Based on our findings, we implement and evaluate a set of pilot solutions to address the storage performance deficiencies in smartphones. Hyojun Kim, Nitin Agrawal 0001, Cristian Ungureanu |
ACM Trans. Storage | 1 |
| 2011 | Impact of flash memory on video-on-demand storage: analysis of tradeoffsabstractThere is no doubt that video-on-demand (VoD) services are very popular these days. However, disk storage is a serious bottleneck limiting the scalability of a VoD server. Disk throughput degrades dramatically due to seek time overhead when the server is called upon to serve a large number of simultaneous video streams. To address the performance problem of disk, buffer cache algorithms that utilize RAM have been proposed. Interval caching is a state-of-the-art caching algorithm for a VoD server. Flash Memory Solid-State Drive (SSD) is a relatively new storage technology. Its excellent random read performance, low power consumption, and sharply dropping cost per gigabyte are opening new opportunities to efficiently use the device for enterprise systems. On the other hand, it has deficiencies such as poor small random write performance and limited number of erase operations. In this paper, we analyze tradeoffs and potential impact that flash memory SSD can have for a VoD server. Performance of various commercially available flash memory SSD models is studied. We find that low-end flash memory SSD provides better performance than the high-end one while costing less than the high-end one when the I/O request size is large, which is typical for a VoD server. Because of the wear problem and asymmetric read/write performance of flash memory SSD, we claim that interval caching cannot be used with it. Instead, we propose using file-level Least Frequently Used (LFU) due to the highly skewed video access pattern of the VoD workload. We compare the performance of interval caching with RAM and file-level LFU with flash memory by simulation experiments. In addition, from the cost-effectiveness analysis of three different storage configurations, we find that flash memory with hard disk drive is the most cost-effective solution compared to DRAM with hard disk drive or hard disk drive only. Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran |
MMSys | 2 |
| 2009 | FlashLite: A User-Level Library to Enhance Durability of SSD for P2P File SharingabstractPeer-to-peer file sharing is popular, but it generates random write traffic to storage due to the nature of swarming. NAND flash memory based Solid-State Drive (SSD) technology is available as an alternative to hard drives for notebook and tablet PCs. As it turns out, random write is extremely detrimental to the lifetime of SSD drives. This paper focuses on the following problem, namely, P2P file downloading when the target of the download is an SSD drive. We make three contributions: first, analysis of write patterns of downloading program to establish the premise of the problem; second, development of a simple yet powerful technique called FlashLite to combat this problem, by automatically converting the random writes to sequential writes; third, showing through performance evaluation using modified eMule file downloading program that FlashLite does change random writes to sequential, and most importantly eliminates about 94% of erase operations of the original eMule program. Hyojun Kim, Umakishore Ramachandran |
ICDCS | 1 |
| 2009 | Design and implementation of MLC NAND flash-based DBMS for mobile devices
Ki Yong Lee, Hyojun Kim, Kyoung-Gu Woo, Yon Dohn Chung, Myoung-Ho Kim |
J. Syst. Softw. | 2 |
| 2008 | BPLRU: A Buffer Management Scheme for Improving Random Writes in Flash Storage
Hyojun Kim, Seongjun Ahn |
FAST | 1 |
| 2007 | A Page Padding Method for Fragmented Flash Storage
Hyojun Kim, Jin-Hyuk Kim, ShinHo Choi, HyunRyong Jung, JaeGyu Jung |
ICCSA (1) | 1 |
| 2007 | SWL: a search-while-load demand paging scheme with NAND flash memoryabstractAs mobile phones become increasingly multifunctional, the number and size of applications installed in phones are rapidly increasing. Consequently, mobile phones require more hardware resources such as NOR/NAND flash memory and DRAM, and their production cost is accordingly becoming higher. One candidate solution to reduce production cost is demand paging using MMU. However, demand paging causes unpredictably long page fault latency, and as such mobile phone manufacturers are reluctant to deploy this scheme. In this paper, we present a method that reduces the long latency of page faults by performing page fault handling in a parallelized manner, considering the characteristics of NAND-Type flash memory. We also discuss how to modify the existing page cache replacement policies so that they can exploit the benefits of the parallelized page fault handler. Experimental results show that the parallelized page fault handler improves the worst case latency of page faults significantly, by up to roughly 20%, and that the modified page cache replacement policies improve both the average and worst instruction fetch time. Jihyun In, Ilhoon Shin, Hyojun Kim |
LCTES | 3 |
| 2006 | MNFS: mobile multimedia file system for NAND flash based storage deviceabstractIn this work, we present a novel mobile multimedia file system, MNFS, which is specifically designed for NAND flash memory. It is designed for mobile multimedia devices such as MP3 player, Personal Media Player (PMP), digital camcorder, etc. Our file system has three novel features important in mobile multimedia applications: (1) predictable and uniform write latency, (2) quick file system mount, and (3) small memory footprint. We implement the proto-type file system on ARM9 embedded platform. In experiments, MNFS exhibits uniform I/O latency for sequential write operation. It is mountable within 0.2 seconds, and available with only 34Kbytes heap memory for 128Mbytes volume. Compared to YAFFS which is the de facto standard file system for NAND flash memory, the mounting time is 30 times faster and the heap memory usage is only 5% of YAFFS usage. Hyojun Kim, Youjip Won |
CCNC | 1 |