EDBT 2026 Demo / reviewers in the wild / expert
Hyokun Yun
dblp:45/9671
· DBLP profile ↗
15ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-6169-4761ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 41% Reinforcement learning · 19% Efficient and distributed learning · 14% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 67% Data mining · 26% Machine learning and data management · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 62% High-performance computing · 24% Parallel and multicore computing · 14% |
Topics — the 29 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › LLM agents
web agents |
1.9 | 2 | 2026 | RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users · AAAI 2026 WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025 |
Natural language and speech › Language models and text generation
alignment |
1.7 | 2 | 2025 | Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025 Aligning Large Language Models with Implicit Preferences from User-Generated Content · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
data mixing |
0.9 | 1 | 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025 |
Machine learning › Reinforcement learning
multi-turn reinforcement learning |
0.9 | 1 | 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025 |
Machine learning › Reinforcement learning
preference learning |
0.9 | 1 | 2025 | Aligning Large Language Models with Implicit Preferences from User-Generated Content · ACL (1) 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › data curation
training data curation |
0.9 | 1 | 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025 |
Machine learning › Learning paradigms
multi-task learning |
0.8 | 1 | 2024 | Robust Multi-Task Learning with Excess Risks · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
robustness to label noise |
0.8 | 1 | 2024 | Robust Multi-Task Learning with Excess Risks · ICML 2024 |
Machine learning › Optimization for machine learning › multi-task optimization
task balancing |
0.8 | 1 | 2024 | Robust Multi-Task Learning with Excess Risks · ICML 2024 |
Machine learning › Optimization for machine learning
distributed optimization |
0.4 | 1 | 2019 | Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019 |
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training |
0.4 | 1 | 2019 | Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019 |
Machine learning › Efficient and distributed learning › distributed training
model parallelism |
0.4 | 1 | 2019 | Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model › logistic regression
multinomial logistic regression |
0.4 | 1 | 2019 | Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019 |
Machine learning › Efficient and distributed learning
active learning |
0.3 | 1 | 2018 | Deep Active Learning for Named Entity Recognition · ICLR (Poster) 2018 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.3 | 1 | 2018 | Deep Active Learning for Named Entity Recognition · ICLR (Poster) 2018 |
Human-AI interaction › GUI agent
GUI grounding |
0.3 | 1 | 2026 | RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users · AAAI 2026 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning · EMNLP 2025 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 1 | 2016 | WordRank: Learning Word Embeddings via Robust Ranking · EMNLP 2016 |
Data mining › text mining
topic modeling |
0.2 | 1 | 2015 | A Scalable Asynchronous Distributed Algorithm for Topic Modeling · WWW 2015 |
Distributed systems › distributed algorithms
asynchronous distributed algorithm |
0.2 | 1 | 2015 | A Scalable Asynchronous Distributed Algorithm for Topic Modeling · WWW 2015 |
Distributed systems
distributed algorithms |
0.2 | 1 | 2015 | A Scalable Asynchronous Distributed Algorithm for Topic Modeling · WWW 2015 |
Information retrieval › ranking
learning to rank |
0.2 | 1 | 2014 | Ranking via Robust Binary Classification · NIPS 2014 |
Information retrieval
retrieval models |
0.2 | 1 | 2014 | Ranking via Robust Binary Classification · NIPS 2014 |
Information retrieval › ranking › learning to rank
robust ranking |
0.2 | 1 | 2014 | Ranking via Robust Binary Classification · NIPS 2014 |
Parallel and multicore computing › parallel computing › parallel machine learning
asynchronous parallel training |
0.1 | 1 | 2019 | Scaling Multinomial Logistic Regression via Hybrid Parallelism · KDD 2019 |
Machine learning and data management
matrix completion |
0.1 | 1 | 2014 | NOMAD: Nonlocking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion · Proc. VLDB Endow. 2014 |
Distributed systems
distributed coordination |
0.1 | 1 | 2014 | NOMAD: Nonlocking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion · Proc. VLDB Endow. 2014 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.0benchmark evaluation · 2.0uncertainty quantification · 0.9policy gradient · 0.9pairwise preference classification · 0.9end-to-end reinforcement learning · 0.9direct preference optimization · 0.9adaptive data mixing · 0.9taylor approximation · 0.8excess risk minimization · 0.8nomad algorithm · 0.4gibbs sampling · 0.4fenwick tree · 0.4token-ring communication · 0.4hybrid parallelism · 0.4stochastic gradient descent · 0.4serializable updates · 0.4non-blocking communication · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World UsersabstractTo achieve successful assistance with long-horizon web-based tasks, AI agents must be able to sequentially follow real-world user instructions over a long period. Unlike existing web-based agent benchmarks, sequential instruction following in the real world poses significant challenges beyond performing a single, clearly defined task. For instance, real-world human instructions can be ambiguous, require different levels of AI assistance, and may evolve over time, reflecting changes in the user's mental state. To address this gap, we introduce RealWebAssist, a novel benchmark designed to evaluate sequential instruction-following in realistic scenarios involving long-horizon interactions with the web, visual GUI grounding, and understanding ambiguous real-world user instructions. RealWebAssist includes a dataset of sequential instructions collected from real-world human users. Each user instructs a web-based assistant to perform a series of tasks on multiple websites. A successful agent must reason about the true intent behind each instruction, keep track of the mental state of the user, understand user-specific routines, and ground the intended tasks to actions on the correct GUI elements. Our experimental results show that state-of-the-art models struggle to understand and ground user instructions, posing critical challenges in following real-world user instructions for long-horizon web assistance. Suyu Ye, Haojun Shi, Darren Shih, Hyokun Yun, Tanya G. Roosta, Tianmin Shu |
AAAI | 4 |
| 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsabstractNicholas E. Corrado, Julian Katz-Samuels, Adithya M Devraj, Hyokun Yun, Chao Zhang, Yi Xu, Yi Pan, Bing Yin, Trishul Chilimbi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nicholas Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun, Yi Xu 0011, Trishul Chilimbi |
ACL (1) | 4 |
| 2025 | Aligning Large Language Models with Implicit Preferences from User-Generated ContentabstractLearning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from humans or advanced LLMs, which is costly and difficult to scale. In this work, we present PUGC, a novel framework that leverages implicit human Preferences in unlabeled User-Generated Content (UGC) to generate preference data. Although UGC is not explicitly created to guide LLMs in generating human-preferred responses, it often reflects valuable insights and implicit preferences from its creators that has the potential to address readers’ questions. PUGC transforms UGC into user queries and generates responses from the policy model. The UGC is then leveraged as a reference text for response scoring, aligning the model with these implicit preferences. This approach improves the quality of preference data while enabling scalable, domain-specific alignment. Experimental results on Alpaca Eval 2 show that models trained with DPO and PUGC achieve a 9.37% performance improvement over traditional methods, setting a 35.93% state-of-the-art length-controlled win rate using Mistral-7B-Instruct. Further studies highlight gains in reward quality, domain-specific alignment effectiveness, robustness against UGC quality, and theory of mind capabilities. Our code and dataset are available at https://zhaoxuan.info/PUGC.github.io/. Zhaoxuan Tan, Zheng Li 0018, Hyokun Yun, Ming Zeng 0001, Zhihan Zhang 0001, Yifan Gao 0001, Ruijie Wang 0004, Priyanka Nigam, Meng Jiang 0001 |
ACL (1) | 5 |
| 2025 | Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM AlignmentabstractLarge Language Models (LLMs) have seen widespread adoption due to their remarkable natural language capabilities. However, when deploying them in real-world settings, it is important to align LLMs to generate texts according to acceptable human standards. Methods such as Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO) have enabled significant progress in refining LLMs using human preference data. However, the privacy concerns inherent in utilizing such preference data have yet to be adequately studied. In this paper, we investigate the vulnerability of LLMs aligned using two widely used methods - DPO and PPO - to membership inference attacks (MIAs). Our study has two main contributions: first, we theoretically motivate that DPO models are more vulnerable to MIA compared to PPO models; second, we introduce a novel reference-based attack framework specifically for analyzing preference data called PREMIA (Preference data MIA). Using PREMIA and existing baselines we empirically show that DPO models have a relatively heightened vulnerability towards MIA. Qizhang Feng, Siva Rajesh Kasa, Santhosh Kumar Kasa, Hyokun Yun, Choon Hui Teo, Sravan Babu Bodapati |
AISTATS | 4 |
| 2025 | WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement LearningabstractZhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, Lihong Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhepei Wei, Wenlin Yao, Changlong Yu, Puyang Xu, Chao Zhang 0014, Hyokun Yun, Lihong Li 0001 |
EMNLP | 11 |
| 2025 | Ask a Strong LLM Judge when Your Reward Model is UncertainabstractReward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs.
By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF.
In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF. Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao |
NeurIPS | 11 |
| 2024 | Robust Multi-Task Learning with Excess RisksabstractMulti-task learning (MTL) considers learning a joint model for multiple tasks by optimizing a convex combination of all task losses. To solve the optimization problem, existing methods use an adaptive weight updating scheme, where task weights are dynamically adjusted based on their respective losses to prioritize difficult tasks. However, these algorithms face a great challenge whenever label noise is present, in which case excessive weights tend to be assigned to noisy tasks that have relatively large Bayes optimal errors, thereby overshadowing other tasks and causing performance to drop across the board. To overcome this limitation, we propose Multi-Task Learning with Excess Risks (ExcessMTL), an excess risk-based task balancing method that updates the task weights by their distances to convergence instead. Intuitively, ExcessMTL assigns higher weights to worse-trained tasks that are further from convergence. To estimate the excess risks, we develop an efficient and accurate method with Taylor approximation. Theoretically, we show that our proposed algorithm achieves convergence guarantees and Pareto stationarity. Empirically, we evaluate our algorithm on various MTL benchmarks and demonstrate its superior performance over existing methods in the presence of label noise. Our code is available at https://github.com/yifei-he/ExcessMTL. Shiji Zhou, Hyokun Yun, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi, Han Zhao 0002 |
ICML | 4 |
| 2022 | MICO: Selective Search with Mutual Information Co-trainingabstractIn contrast to traditional exhaustive search, selective search first clusters documents into several groups before all the documents are searched exhaustively by a query, to limit the search executed within one group or only a few groups. Selective search is designed to reduce the latency and computation in modern large-scale search systems. In this study, we propose MICO, a Mutual Information CO-training framework for selective search with minimal supervision using the search logs. After training, MICO does not only cluster the documents, but also routes unseen queries to the relevant clusters for efficient retrieval. In our empirical experiments, MICO significantly improves the performance on multiple metrics of selective search and outperforms a number of existing competitive baselines. Zhanyu Wang, Hyokun Yun, Choon Hui Teo, Trishul Chilimbi |
COLING | 3 |
| 2019 | Scaling Multinomial Logistic Regression via Hybrid ParallelismabstractWe study the problem of scaling Multinomial Logistic Regression (MLR) to datasets with very large number of data points in the presence of large number of classes. At a scale where neither data nor the parameters are able to fit on a single machine, we argue that simultaneous data and model parallelism (Hybrid Parallelism) is inevitable. The key challenge in achieving such a form of parallelism in MLR is the log-partition function which needs to be computed across all K classes per data point, thus making model parallelism non-trivial. To overcome this problem, we propose a reformulation of the original objective that exploits double-separability, an attractive property that naturally leads to hybrid parallelism. Our algorithm (DS-MLR) is asynchronous and completely de-centralized, requiring minimal communication across workers while keeping both data and parameter workloads partitioned. Unlike standard data parallel approaches, DS-MLR avoids bulk-synchronization by maintaining local normalization terms on each worker and accumulating them incrementally using a token-ring topology. We demonstrate the versatility of DS-MLR under various scenarios in data and model parallelism, through an empirical study consisting of real-world datasets. In particular, to demonstrate scaling via hybrid parallelism, we created a new benchmark dataset (Reddit-Full) by pre-processing 1.7 billion reddit user comments spanning the period 2007-2015. We used DS-MLR to solve an extreme multi-class classification problem of classifying 211 million data points into their corresponding subreddits. Reddit-Full is a massive data set with data occupying 228 GB and 44 billion parameters occupying 358 GB. To the best of our knowledge, no other existing methods can handle MLR in this setting. Parameswaran Raman, Sriram Srinivasan 0004, Shin Matsushima, Hyokun Yun, S. V. N. Vishwanathan |
KDD | 5 |
| 2018 | Deep Active Learning for Named Entity Recognition
Yanyao Shen, Hyokun Yun, Zachary C. Lipton, Yakov Kronrod, Anima Anandkumar |
ICLR (Poster) | 2 |
| 2017 | Distributed Stochastic Optimization of Regularized Risk via Saddle-Point Problem
Shin Matsushima, Hyokun Yun, S. V. N. Vishwanathan |
ECML/PKDD (1) | 2 |
| 2016 | WordRank: Learning Word Embeddings via Robust RankingabstractEmbedding words in a vector space has gained a lot of attention in recent years.While stateof-the-art methods provide efficient computation of word similarities via a low-dimensional matrix embedding, their motivation is often left unclear.In this paper, we argue that word embedding can be naturally viewed as a ranking problem due to the ranking nature of the evaluation metrics.Then, based on this insight, we propose a novel framework Wor-dRank that efficiently estimates word representations via robust ranking, in which the attention mechanism and robustness to noise are readily achieved via the DCG-like ranking losses.The performance of WordRank is measured in word similarity and word analogy benchmarks, and the results are compared to the state-of-the-art word embedding techniques.Our algorithm is very competitive to the state-of-the-arts on large corpora, while outperforms them by a significant margin when the training set is limited (i.e., sparse and noisy).With 17 million tokens, WordRank performs almost as well as existing methods using 7.2 billion tokens on a popular word similarity benchmark.Our multi-node distributed implementation of WordRank is publicly available for general usage. Shihao Ji 0001, Hyokun Yun, Pinar Yanardag Delul, Shin Matsushima, S. V. N. Vishwanathan |
EMNLP | 2 |
| 2015 | A Scalable Asynchronous Distributed Algorithm for Topic ModelingabstractLearning meaningful topic models with massive document collections which contain millions of documents and billions of tokens is challenging because of two reasons. First, one needs to deal with a large number of topics (typically on the order of thousands). Second, one needs a scalable and efficient way of distributing the computation across multiple machines. In this paper, we present a novel algorithm F+Nomad LDA which simultaneously tackles both these problems. In order to handle large number of topics we use an appropriately modified Fenwick tree. This data structure allows us to sample from a multinomial distribution over T items in O(log T) time. Moreover, when topic counts change the data structure can be updated in O(log T) time. In order to distribute the computation across multiple processors, we present a novel asynchronous framework inspired by the Nomad algorithm of Yun et al, 2014. We show that F+Nomad LDA significantly outperforms recent state-of-the-art topic modeling approaches on massive problems which involve millions of documents, billions of words, and thousands of topics. Hsiang-Fu Yu, Cho-Jui Hsieh, Hyokun Yun, S. V. N. Vishwanathan, Inderjit S. Dhillon |
WWW | 3 |
| 2014 | Ranking via Robust Binary Classification
Hyokun Yun, Parameswaran Raman, S. V. N. Vishwanathan |
NIPS | 1 |
| 2014 | NOMAD: Nonlocking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completionabstractWe develop an efficient parallel distributed algorithm for matrix completion, named NOMAD (Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion). NOMAD is a decentralized algorithm with non-blocking communication between processors. One of the key features of NOMAD is that the ownership of a variable is asynchronously transferred between processors in a decentralized fashion. As a consequence it is a lock-free parallel algorithm. In spite of being asynchronous, the variable updates of NOMAD are serializable, that is, there is an equivalent update ordering in a serial implementation. NOMAD outperforms synchronous algorithms which require explicit bulk synchronization after every iteration: our extensive empirical evaluation shows that not only does our algorithm perform well in distributed setting on commodity hardware, but also outperforms state-of-the-art algorithms on a HPC cluster both in multi-core and distributed memory settings. Hyokun Yun, Hsiang-Fu Yu, Cho-Jui Hsieh, S. V. N. Vishwanathan, Inderjit S. Dhillon |
Proc. VLDB Endow. | 1 |