Hyokun Yun

dblp:45/9671 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-6169-4761ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 41% Reinforcement learning · 19% Efficient and distributed learning · 14%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 67% Data mining · 26% Machine learning and data management · 7%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 62% High-performance computing · 24% Parallel and multicore computing · 14%

Topics — the 29 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
web agents
1.922026
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users · AAAI 2026
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025
Natural language and speech › Language models and text generation
alignment
1.722025
Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025
Aligning Large Language Models with Implicit Preferences from User-Generated Content · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model training
data mixing
0.912025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025
Machine learning › Reinforcement learning
multi-turn reinforcement learning
0.912025
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025
Machine learning › Reinforcement learning
preference learning
0.912025
Aligning Large Language Models with Implicit Preferences from User-Generated Content · ACL (1) 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025
Machine learning › Efficient and distributed learning › data curation
training data curation
0.912025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025
Machine learning › Learning paradigms
multi-task learning
0.812024
Robust Multi-Task Learning with Excess Risks · ICML 2024
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
robustness to label noise
0.812024
Robust Multi-Task Learning with Excess Risks · ICML 2024
Machine learning › Optimization for machine learning › multi-task optimization
task balancing
0.812024
Robust Multi-Task Learning with Excess Risks · ICML 2024
Machine learning › Optimization for machine learning
distributed optimization
0.412019
Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training
0.412019
Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.412019
Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model › logistic regression
multinomial logistic regression
0.412019
Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019
Machine learning › Efficient and distributed learning
active learning
0.312018
Deep Active Learning for Named Entity Recognition · ICLR (Poster) 2018
Natural language and speech › Information extraction and text analysis
named entity recognition
0.312018
Deep Active Learning for Named Entity Recognition · ICLR (Poster) 2018
Human-AI interaction › GUI agent
GUI grounding
0.312026
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users · AAAI 2026
Natural language and speech › Language models and text generation
LLM agents
0.312025
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.212016
WordRank: Learning Word Embeddings via Robust Ranking · EMNLP 2016
Data mining › text mining
topic modeling
0.212015
A Scalable Asynchronous Distributed Algorithm for Topic Modeling · WWW 2015
Distributed systems › distributed algorithms
asynchronous distributed algorithm
0.212015
A Scalable Asynchronous Distributed Algorithm for Topic Modeling · WWW 2015
Distributed systems
distributed algorithms
0.212015
A Scalable Asynchronous Distributed Algorithm for Topic Modeling · WWW 2015
Information retrieval › ranking
learning to rank
0.212014
Ranking via Robust Binary Classification · NIPS 2014
Information retrieval
retrieval models
0.212014
Ranking via Robust Binary Classification · NIPS 2014
Information retrieval › ranking › learning to rank
robust ranking
0.212014
Ranking via Robust Binary Classification · NIPS 2014
Parallel and multicore computing › parallel computing › parallel machine learning
asynchronous parallel training
0.112019
Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019
Machine learning and data management
matrix completion
0.112014
NOMAD: Nonlocking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion · Proc. VLDB Endow. 2014
Distributed systems
distributed coordination
0.112014
NOMAD: Nonlocking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion · Proc. VLDB Endow. 2014

Methods — techniques the papers use, named apart from their topics

large language model · 2.0benchmark evaluation · 2.0uncertainty quantification · 0.9policy gradient · 0.9pairwise preference classification · 0.9end-to-end reinforcement learning · 0.9direct preference optimization · 0.9adaptive data mixing · 0.9taylor approximation · 0.8excess risk minimization · 0.8nomad algorithm · 0.4gibbs sampling · 0.4fenwick tree · 0.4token-ring communication · 0.4hybrid parallelism · 0.4stochastic gradient descent · 0.4serializable updates · 0.4non-blocking communication · 0.4
YearPublicationVenuePosition
2026 RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
abstract
To achieve successful assistance with long-horizon web-based tasks, AI agents must be able to sequentially follow real-world user instructions over a long period. Unlike existing web-based agent benchmarks, sequential instruction following in the real world poses significant challenges beyond performing a single, clearly defined task. For instance, real-world human instructions can be ambiguous, require different levels of AI assistance, and may evolve over time, reflecting changes in the user's mental state. To address this gap, we introduce RealWebAssist, a novel benchmark designed to evaluate sequential instruction-following in realistic scenarios involving long-horizon interactions with the web, visual GUI grounding, and understanding ambiguous real-world user instructions. RealWebAssist includes a dataset of sequential instructions collected from real-world human users. Each user instructs a web-based assistant to perform a series of tasks on multiple websites. A successful agent must reason about the true intent behind each instruction, keep track of the mental state of the user, understand user-specific routines, and ground the intended tasks to actions on the correct GUI elements. Our experimental results show that state-of-the-art models struggle to understand and ground user instructions, posing critical challenges in following real-world user instructions for long-horizon web assistance.
Suyu Ye, Haojun Shi, Darren Shih, Hyokun Yun, Tanya G. Roosta, Tianmin Shu
AAAI4
2025 AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
abstract
Nicholas E. Corrado, Julian Katz-Samuels, Adithya M Devraj, Hyokun Yun, Chao Zhang, Yi Xu, Yi Pan, Bing Yin, Trishul Chilimbi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nicholas Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun, Yi Xu 0011, Trishul Chilimbi
ACL (1)4
2025 Aligning Large Language Models with Implicit Preferences from User-Generated Content
abstract
Learning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from humans or advanced LLMs, which is costly and difficult to scale. In this work, we present PUGC, a novel framework that leverages implicit human Preferences in unlabeled User-Generated Content (UGC) to generate preference data. Although UGC is not explicitly created to guide LLMs in generating human-preferred responses, it often reflects valuable insights and implicit preferences from its creators that has the potential to address readers’ questions. PUGC transforms UGC into user queries and generates responses from the policy model. The UGC is then leveraged as a reference text for response scoring, aligning the model with these implicit preferences. This approach improves the quality of preference data while enabling scalable, domain-specific alignment. Experimental results on Alpaca Eval 2 show that models trained with DPO and PUGC achieve a 9.37% performance improvement over traditional methods, setting a 35.93% state-of-the-art length-controlled win rate using Mistral-7B-Instruct. Further studies highlight gains in reward quality, domain-specific alignment effectiveness, robustness against UGC quality, and theory of mind capabilities. Our code and dataset are available at https://zhaoxuan.info/PUGC.github.io/.
Zhaoxuan Tan, Zheng Li 0018, Hyokun Yun, Ming Zeng 0001, Zhihan Zhang 0001, Yifan Gao 0001, Ruijie Wang 0004, Priyanka Nigam, Meng Jiang 0001
ACL (1)5
2025 Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment
abstract
Large Language Models (LLMs) have seen widespread adoption due to their remarkable natural language capabilities. However, when deploying them in real-world settings, it is important to align LLMs to generate texts according to acceptable human standards. Methods such as Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO) have enabled significant progress in refining LLMs using human preference data. However, the privacy concerns inherent in utilizing such preference data have yet to be adequately studied. In this paper, we investigate the vulnerability of LLMs aligned using two widely used methods - DPO and PPO - to membership inference attacks (MIAs). Our study has two main contributions: first, we theoretically motivate that DPO models are more vulnerable to MIA compared to PPO models; second, we introduce a novel reference-based attack framework specifically for analyzing preference data called PREMIA (Preference data MIA). Using PREMIA and existing baselines we empirically show that DPO models have a relatively heightened vulnerability towards MIA.
Qizhang Feng, Siva Rajesh Kasa, Santhosh Kumar Kasa, Hyokun Yun, Choon Hui Teo, Sravan Babu Bodapati
AISTATS4
2025 WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
abstract
Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, Lihong Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhepei Wei, Wenlin Yao, Changlong Yu, Puyang Xu, Chao Zhang 0014, Hyokun Yun, Lihong Li 0001
EMNLP11
2025 Ask a Strong LLM Judge when Your Reward Model is Uncertain
abstract
Reward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs. By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF. In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF.
Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao
NeurIPS11
2024 Robust Multi-Task Learning with Excess Risks
abstract
Multi-task learning (MTL) considers learning a joint model for multiple tasks by optimizing a convex combination of all task losses. To solve the optimization problem, existing methods use an adaptive weight updating scheme, where task weights are dynamically adjusted based on their respective losses to prioritize difficult tasks. However, these algorithms face a great challenge whenever label noise is present, in which case excessive weights tend to be assigned to noisy tasks that have relatively large Bayes optimal errors, thereby overshadowing other tasks and causing performance to drop across the board. To overcome this limitation, we propose Multi-Task Learning with Excess Risks (ExcessMTL), an excess risk-based task balancing method that updates the task weights by their distances to convergence instead. Intuitively, ExcessMTL assigns higher weights to worse-trained tasks that are further from convergence. To estimate the excess risks, we develop an efficient and accurate method with Taylor approximation. Theoretically, we show that our proposed algorithm achieves convergence guarantees and Pareto stationarity. Empirically, we evaluate our algorithm on various MTL benchmarks and demonstrate its superior performance over existing methods in the presence of label noise. Our code is available at https://github.com/yifei-he/ExcessMTL.
Shiji Zhou, Hyokun Yun, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi, Han Zhao 0002
ICML4
2022 MICO: Selective Search with Mutual Information Co-training
abstract
In contrast to traditional exhaustive search, selective search first clusters documents into several groups before all the documents are searched exhaustively by a query, to limit the search executed within one group or only a few groups. Selective search is designed to reduce the latency and computation in modern large-scale search systems. In this study, we propose MICO, a Mutual Information CO-training framework for selective search with minimal supervision using the search logs. After training, MICO does not only cluster the documents, but also routes unseen queries to the relevant clusters for efficient retrieval. In our empirical experiments, MICO significantly improves the performance on multiple metrics of selective search and outperforms a number of existing competitive baselines.
Zhanyu Wang, Hyokun Yun, Choon Hui Teo, Trishul Chilimbi
COLING3
2019 Scaling Multinomial Logistic Regression via Hybrid Parallelism
abstract
We study the problem of scaling Multinomial Logistic Regression (MLR) to datasets with very large number of data points in the presence of large number of classes. At a scale where neither data nor the parameters are able to fit on a single machine, we argue that simultaneous data and model parallelism (Hybrid Parallelism) is inevitable. The key challenge in achieving such a form of parallelism in MLR is the log-partition function which needs to be computed across all K classes per data point, thus making model parallelism non-trivial. To overcome this problem, we propose a reformulation of the original objective that exploits double-separability, an attractive property that naturally leads to hybrid parallelism. Our algorithm (DS-MLR) is asynchronous and completely de-centralized, requiring minimal communication across workers while keeping both data and parameter workloads partitioned. Unlike standard data parallel approaches, DS-MLR avoids bulk-synchronization by maintaining local normalization terms on each worker and accumulating them incrementally using a token-ring topology. We demonstrate the versatility of DS-MLR under various scenarios in data and model parallelism, through an empirical study consisting of real-world datasets. In particular, to demonstrate scaling via hybrid parallelism, we created a new benchmark dataset (Reddit-Full) by pre-processing 1.7 billion reddit user comments spanning the period 2007-2015. We used DS-MLR to solve an extreme multi-class classification problem of classifying 211 million data points into their corresponding subreddits. Reddit-Full is a massive data set with data occupying 228 GB and 44 billion parameters occupying 358 GB. To the best of our knowledge, no other existing methods can handle MLR in this setting.
Parameswaran Raman, Sriram Srinivasan 0004, Shin Matsushima, Hyokun Yun, S. V. N. Vishwanathan
KDD5
2018 Deep Active Learning for Named Entity Recognition
Yanyao Shen, Hyokun Yun, Zachary C. Lipton, Yakov Kronrod, Anima Anandkumar
ICLR (Poster)2
2017 Distributed Stochastic Optimization of Regularized Risk via Saddle-Point Problem
Shin Matsushima, Hyokun Yun, S. V. N. Vishwanathan
ECML/PKDD (1)2
2016 WordRank: Learning Word Embeddings via Robust Ranking
abstract
Embedding words in a vector space has gained a lot of attention in recent years.While stateof-the-art methods provide efficient computation of word similarities via a low-dimensional matrix embedding, their motivation is often left unclear.In this paper, we argue that word embedding can be naturally viewed as a ranking problem due to the ranking nature of the evaluation metrics.Then, based on this insight, we propose a novel framework Wor-dRank that efficiently estimates word representations via robust ranking, in which the attention mechanism and robustness to noise are readily achieved via the DCG-like ranking losses.The performance of WordRank is measured in word similarity and word analogy benchmarks, and the results are compared to the state-of-the-art word embedding techniques.Our algorithm is very competitive to the state-of-the-arts on large corpora, while outperforms them by a significant margin when the training set is limited (i.e., sparse and noisy).With 17 million tokens, WordRank performs almost as well as existing methods using 7.2 billion tokens on a popular word similarity benchmark.Our multi-node distributed implementation of WordRank is publicly available for general usage.
Shihao Ji 0001, Hyokun Yun, Pinar Yanardag Delul, Shin Matsushima, S. V. N. Vishwanathan
EMNLP2
2015 A Scalable Asynchronous Distributed Algorithm for Topic Modeling
abstract
Learning meaningful topic models with massive document collections which contain millions of documents and billions of tokens is challenging because of two reasons. First, one needs to deal with a large number of topics (typically on the order of thousands). Second, one needs a scalable and efficient way of distributing the computation across multiple machines. In this paper, we present a novel algorithm F+Nomad LDA which simultaneously tackles both these problems. In order to handle large number of topics we use an appropriately modified Fenwick tree. This data structure allows us to sample from a multinomial distribution over T items in O(log T) time. Moreover, when topic counts change the data structure can be updated in O(log T) time. In order to distribute the computation across multiple processors, we present a novel asynchronous framework inspired by the Nomad algorithm of Yun et al, 2014. We show that F+Nomad LDA significantly outperforms recent state-of-the-art topic modeling approaches on massive problems which involve millions of documents, billions of words, and thousands of topics.
Hsiang-Fu Yu, Cho-Jui Hsieh, Hyokun Yun, S. V. N. Vishwanathan, Inderjit S. Dhillon
WWW3
2014 Ranking via Robust Binary Classification
Hyokun Yun, Parameswaran Raman, S. V. N. Vishwanathan
NIPS1
2014 NOMAD: Nonlocking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion
abstract
We develop an efficient parallel distributed algorithm for matrix completion, named NOMAD (Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion). NOMAD is a decentralized algorithm with non-blocking communication between processors. One of the key features of NOMAD is that the ownership of a variable is asynchronously transferred between processors in a decentralized fashion. As a consequence it is a lock-free parallel algorithm. In spite of being asynchronous, the variable updates of NOMAD are serializable, that is, there is an equivalent update ordering in a serial implementation. NOMAD outperforms synchronous algorithms which require explicit bulk synchronization after every iteration: our extensive empirical evaluation shows that not only does our algorithm perform well in distributed setting on commodity hardware, but also outperforms state-of-the-art algorithms on a HPC cluster both in multi-core and distributed memory settings.
Hyokun Yun, Hsiang-Fu Yu, Cho-Jui Hsieh, S. V. N. Vishwanathan, Inderjit S. Dhillon
Proc. VLDB Endow.1