Moonkyung Ryu

dblp:67/8693 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 46% Reinforcement learning · 32% Trustworthy machine learning · 9%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 87% Embedded and real-time systems · 13%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
text-to-image generation
1.722025
Preference Adaptive and Sequential Text-to-Image Generation · ICML 2025
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation · CVPR 2025
Machine learning › Generative modeling
diffusion model
1.522025
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation · CVPR 2025
Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models · NeurIPS 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.912025
Preference Adaptive and Sequential Text-to-Image Generation · ICML 2025
Collaborative and social computing › collaborative design
co-creation
0.912025
Preference Adaptive and Sequential Text-to-Image Generation · ICML 2025
Machine learning › Reinforcement learning
policy optimization
0.722023
Variational Model-based Policy Optimization · IJCAI 2021
Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models · NeurIPS 2023
Natural language and speech › Question answering and dialogue systems
dialogue management
0.712023
A Mixture-of-Expert Approach to RL-based Dialogue Management · ICLR 2023
Machine learning › Generative modeling › diffusion model › diffusion model training
diffusion model fine-tuning
0.712023
Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models · NeurIPS 2023
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.712023
Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models · NeurIPS 2023
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.712023
Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models · NeurIPS 2023
Machine learning › Reinforcement learning
model-based reinforcement learning
0.512021
Variational Model-based Policy Optimization · IJCAI 2021
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.412020
CAQL: Continuous Action Q-Learning · ICLR 2020
Natural language and speech › Language models and text generation › alignment
preference alignment
0.312025
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation · CVPR 2025
Natural language and speech › Language models and text generation › alignment
reward hacking
0.312025
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation · CVPR 2025
Storage systems
flash and SSD
0.322012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Flashboost: design of flash memory buffer cache mechanism for video-on-demand · ACM Multimedia 2010
Content delivery and video streaming › adaptive video streaming
HTTP adaptive streaming
0.212013
FlashStream: a multi-tiered storage architecture for adaptive HTTP streaming · ACM Multimedia 2013
Storage systems › flash and SSD
SSD cache
0.212013
FlashStream: a multi-tiered storage architecture for adaptive HTTP streaming · ACM Multimedia 2013
Storage systems › storage hierarchy
tiered storage
0.212013
FlashStream: a multi-tiered storage architecture for adaptive HTTP streaming · ACM Multimedia 2013
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.112021
Variational Model-based Policy Optimization · IJCAI 2021
Operating systems › resource management › memory management
buffer cache
0.112012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Operating systems › resource management › memory management
cache replacement
0.112012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Embedded and real-time systems
mobile computing
0.012012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Embedded and real-time systems › mobile computing
smartphone storage
0.012012
What is a good buffer cache replacement scheme for mobile flash storage? · SIGMETRICS 2012
Multimedia systems and quality of experience › video streaming
video-on-demand
0.012010
Flashboost: design of flash memory buffer cache mechanism for video-on-demand · ACM Multimedia 2010

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 3.1expectation-maximization · 2.2large multimodal language model · 1.7reward model · 0.9region-aware fine-tuning · 0.9mixture of experts · 0.7variational lower bound · 0.5actor-critic · 0.5write granularity optimization · 0.3qos-sensitive caching · 0.3trace-driven simulation · 0.3OS-level implementation · 0.3interval caching · 0.2
YearPublicationVenuePosition
2025 Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation
abstract
Text-to-image (T2I) generation has made significant advances in recent years, but challenges still remain in the generation of perceptual artifacts, misalignment with complex prompts, and safety. The prevailing approach to address these issues involves collecting human feedback on generated images, training reward models to estimate human feedback, and then fine-tuning T2I models based on the reward models to align them with human preferences. However, while existing reward fine-tuning methods can produce images with higher rewards, they may change model behavior in unexpected ways. For example, fine-tuning for one quality aspect (e.g., safety) may degrade other aspects (e.g., prompt alignment), or may lead to reward hacking (e.g., finding a way to increase rewards without having the intended effect). In this paper, we propose Focus-N-Fix, the first region-aware fine-tuning method that trains models to correct only previously problematic image regions. The resulting fine-tuned model generates images with the same high-level structure as the original model but shows significant improvements in regions where the original model was deficient in safety (over-sexualization and violence), plausibility, or other criteria. Our experiments demonstrate that Focus-N-Fix improves these localized quality aspects with little or no degradation to others and typically imperceptible changes in the rest of the image. Disclaimer: This paper contains images that may be overly sexual, violent, offensive or harmful.
Xiaoying Xing, Avinab Saha, Junfeng He, Susan Hao, Paul Vicol, Moonkyung Ryu, Gang Li 0021, Sahil Singla 0005, Sarah Young, Yinxiao Li, Feng Yang 0008, Deepak Ramachandran
CVPR6
2025 Preference Adaptive and Sequential Text-to-Image Generation
abstract
We address the problem of interactive text-to-image (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions. Using human raters, we create a novel dataset of sequential preferences, which we leverage, together with large-scale open-source (non-sequential) datasets. We construct user-preference and user-choice models using an EM strategy and identify varying user preference types. We then leverage a large multimodal language model (LMM) and a value-based RL approach to suggest an adaptive and diverse slate of prompt expansions to the user. Our Preference Adaptive and Sequential Text-to-image Agent (PASTA) extends T2I models with adaptive multi-turn capabilities, fostering collaborative co-creation and addressing uncertainty or underspecification in a user’s intent. We evaluate PASTA using human raters, showing significant improvement compared to baseline methods. We also open-source our sequential rater dataset and simulated user-rater interactions to support future research in user-centric multi-turn T2I systems.
Ofir Nabati, Guy Tennenholtz, Moonkyung Ryu, Deepak Ramachandran, Yinlam Chow, Craig Boutilier
ICML4
2023 A Mixture-of-Expert Approach to RL-based Dialogue Management
Yinlam Chow, Aza Tulepbergenov, Ofir Nachum, Dhawal Gupta, Moonkyung Ryu, Mohammad Ghavamzadeh, Craig Boutilier
ICLR5
2023 Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Olivia Watkins, Hao Liu 0055, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee 0001, Kimin Lee
NeurIPS5
2021 Variational Model-based Policy Optimization
abstract
Model-based reinforcement learning (RL) algorithms allow us to combine model-generated data with those collected from interaction with the real system in order to alleviate the data efficiency problem in RL. However, designing such algorithms is often challenging because the bias in simulated data may overshadow the ease of data generation. A potential solution to this challenge is to jointly learn and improve model and policy using a universal objective function. In this paper, we leverage the connection between RL and probabilistic inference, and formulate such an objective function as a variational lower-bound of a log-likelihood. This allows us to use expectation maximization (EM) and iteratively fix a baseline policy and learn a variational distribution, consisting of a model and a policy (E-step), followed by improving the baseline policy given the learned variational distribution (M-step). We propose model-based and model-free policy iteration (actor-critic) style algorithms for the E-step and show how the variational distribution learned by them can be used to optimize the M-step in a fully model-based fashion. Our experiments on a number of continuous control tasks show that our model-based (E-step) algorithm, called variational model-based policy optimization (VMBPO), is more sample-efficient and robust to hyper-parameter tuning than its model-free (E-step) counterpart. Using the same control tasks, we also compare VMBPO with several state-of-the-art model-based and model-free RL algorithms and show its sample efficiency and performance.
Yinlam Chow, Brandon Cui, Moonkyung Ryu, Mohammad Ghavamzadeh
IJCAI3
2020 CAQL: Continuous Action Q-Learning
Moonkyung Ryu, Yinlam Chow, Christian Tjandraatmadja, Craig Boutilier
ICLR1
2013 FlashStream: a multi-tiered storage architecture for adaptive HTTP streaming
abstract
Video streaming on the Internet is popular and the need to store and stream video content using CDNs is continually on the rise thanks to services such as Hulu and Netflix. Adaptive HTTP streaming using the deployed CDN infrastructure has become the de facto standard for meeting the increasing demand for video streaming on the Internet. The storage architecture that is used for storing and streaming the video content is the focus of this study. Hard-disk as the storage medium has been the norm for enterprise-class storage servers for the longest time. More recently, multi-tiered storage servers (incorporating SSDs) such as Sun's ZFS and Facebook's flashcache offer an alternative to disk-based storage servers for enterprise applications. Both these systems use the SSD as a cache between the DRAM and the hard disk. The thesis of our work is that the current-state-of-the art in multi-tiered storage systems, architected for general-purpose enterprise workloads, do not cater to the unique needs of adaptive HTTP streaming. We present FlashStream, a multi-tiered storage architecture that addresses the unique needs of adaptive HTTP streaming. Like ZFS and flashcache, it also incorporates SSDs as a cache between the DRAM and the hard disk. The key architectural elements of FlashStream include optimal write granularity to overcome the write amplification effect of flash memory SSDs and a QoS-sensitive caching strategy that monitors the activity of the flash memory SSDs to ensure that video streaming performance is not hampered by the caching activity. We have implemented FlashStream and experimentally compared it with ZFS and flashcache for adaptive HTTP streaming workloads. We show that FlashStream outperforms both these systems for the same hardware configuration. Specifically, it is better by a factor of two compared to its nearest competitor, namely ZFS. In addition, we have compared FlashStream with a traditional two-level storage architecture (DRAM + HDDs), and have shown that, for the same investment cost, FlashStream provides 33% better performance and 94% better energy efficiency.
Moonkyung Ryu, Umakishore Ramachandran
ACM Multimedia1
2012 Why are state-of-the-art flash-based multi-tiered storage systems performing poorly for HTTP video streaming?
abstract
MLC flash memory is a promising technology for building a high-performance and cost-effective video streaming system when it is used as an intermediate level cache in a multi-tiered storage hierarchy. Therefore, we were quite surprised when through extensive measurements we found that two state-of-the-art flash-based multi-tiered storage systems (namely, flashcache and ZFS) have quite disappointing performance for HTTP video streaming using the DASH protocol. We have conducted a thorough analysis to understand the reasons for the poor performance of these two systems. In a nutshell, unless attention is paid to the unique performance characteristics of flash memory-based SSDs, we could end up with suboptimal or even poor performance as we discovered through experimentation with these two systems. Based on the analysis, we present design guidelines for building a cost-effective high-performance HTTP video streaming server.
Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran
NOSSDAV1
2012 What is a good buffer cache replacement scheme for mobile flash storage?
abstract
Smartphones are becoming ubiquitous and powerful. The Achilles' heel in such devices that limits performance is the storage. Low-end flash memory is the storage technology of choice in such devices due to energy, size, and cost considerations. In this paper, we take a critical look at the performance of flash on smartphones for mobile applications. Specifically, we ask the question whether the state-of-the-art buffer cache replacement schemes proposed thus far (both flash-agnostic and flash-aware ones) are the right ones for mobile flash storage. To answer this question, we first expose the limitations of current buffer cache performance evaluation methods, and propose a novel evaluation framework that is a hybrid between trace-driven simulation and real implementation of such schemes inside an operating system. Such an evaluation reveals some unexpected and surprising insights on the performance of buffer management schemes that contradicts conventional wisdom. Armed with this knowledge, we propose a new buffer cache replacement scheme called SpatialClock.
Hyojun Kim, Moonkyung Ryu, Umakishore Ramachandran
SIGMETRICS2
2011 Impact of flash memory on video-on-demand storage: analysis of tradeoffs
abstract
There is no doubt that video-on-demand (VoD) services are very popular these days. However, disk storage is a serious bottleneck limiting the scalability of a VoD server. Disk throughput degrades dramatically due to seek time overhead when the server is called upon to serve a large number of simultaneous video streams. To address the performance problem of disk, buffer cache algorithms that utilize RAM have been proposed. Interval caching is a state-of-the-art caching algorithm for a VoD server. Flash Memory Solid-State Drive (SSD) is a relatively new storage technology. Its excellent random read performance, low power consumption, and sharply dropping cost per gigabyte are opening new opportunities to efficiently use the device for enterprise systems. On the other hand, it has deficiencies such as poor small random write performance and limited number of erase operations. In this paper, we analyze tradeoffs and potential impact that flash memory SSD can have for a VoD server. Performance of various commercially available flash memory SSD models is studied. We find that low-end flash memory SSD provides better performance than the high-end one while costing less than the high-end one when the I/O request size is large, which is typical for a VoD server. Because of the wear problem and asymmetric read/write performance of flash memory SSD, we claim that interval caching cannot be used with it. Instead, we propose using file-level Least Frequently Used (LFU) due to the highly skewed video access pattern of the VoD workload. We compare the performance of interval caching with RAM and file-level LFU with flash memory by simulation experiments. In addition, from the cost-effectiveness analysis of three different storage configurations, we find that flash memory with hard disk drive is the most cost-effective solution compared to DRAM with hard disk drive or hard disk drive only.
Moonkyung Ryu, Hyojun Kim, Umakishore Ramachandran
MMSys1
2010 DynaStream: Adaptive Overlay Management for Peer-to-Peer Video Streaming
abstract
A number of mesh-based peer-to-peer video streaming systems have been proposed, however, they have not paid careful attention to the impact of the number of data streams on streaming quality. In this paper, we first explore the effect of the choice of the number of data streams to the streaming quality. We propose a simple, practical, and fully distributed mechanism, Loss Rate Window, that maintains the overlay topology adaptive to peers' uplink bandwidth and network fluctuation. The mechanism allows each node to make an entirely local decision to adaptively control the right number of data streams to support on its uplink without explicitly knowing or measuring its uplink bandwidth. This mechanism is general, and can be applied to any kind of mesh-based peer-to-peer video streaming system to improve the quality of the video streaming service.
Moonkyung Ryu, Umakishore Ramachandran
ICCCN1
2010 Flashboost: design of flash memory buffer cache mechanism for video-on-demand
abstract
A magnetic disk is a serious bottleneck which limits the scalability of a video server due to its head seek overhead. For a video server, Interval Caching is a state-of-the-art caching mechanism that addresses the problem utilizing RAM as a buffer cache to serve more video streams. Flash memory SSD (Solid State Drive) is a brand new storage device which has very different traits from old storage devices like RAM or disks. The objective of this research is to investigate the applicability and potential impact that flash memory SSD has for a video server. Moreover, I will propose a novel buffer cache mechanism which exploits characteristics of flash memory SSD at maximum to improve the scalability of a video server.
Moonkyung Ryu
ACM Multimedia1