VLDB 2026 Research / reviewers in the wild / expert
Hyunsoo Cho
dblp:86/125
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 40% Representation and self-supervised learning · 17% Knowledge representation and reasoning · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Transaction processing and concurrency control · 70% Indexing and storage engines · 30% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
in-context learning |
1.2 | 2 | 2023 | Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context Learners · AAAI 2023 Ground-Truth Labels Matter: A Deeper Look into Input-Label Demonstrations · EMNLP 2022 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization |
1.0 | 1 | 2026 | CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
fairness |
1.0 | 1 | 2026 | Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
moral judgment |
1.0 | 1 | 2026 | Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › prompting
persona prompting |
1.0 | 1 | 2026 | Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.0 | 1 | 2026 | CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026 |
Machine learning › Learning theory › generalization
combinatorial generalization |
0.7 | 1 | 2023 | MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.7 | 1 | 2023 | MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action · ICML 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.7 | 1 | 2023 | Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context Learners · AAAI 2023 |
Machine learning › Generative modeling
generative adversarial network |
0.7 | 1 | 2023 | Finding the Global Semantic Representation in GAN through Fréchet Mean · ICLR 2023 |
Machine learning › Representation and self-supervised learning › equivariance
group equivariance |
0.7 | 1 | 2023 | MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action · ICML 2023 |
Machine learning › Representation and self-supervised learning
latent space |
0.7 | 1 | 2023 | Finding the Global Semantic Representation in GAN through Fréchet Mean · ICLR 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation |
0.7 | 1 | 2023 | Finding the Global Semantic Representation in GAN through Fréchet Mean · ICLR 2023 |
Natural language and speech › Information extraction and text analysis
text classification |
0.7 | 1 | 2023 | CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification |
0.7 | 1 | 2023 | CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels · ACL (1) 2023 |
Transaction processing and concurrency control › concurrency control
multiversion concurrency control |
0.6 | 2 | 2021 | Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020 Rethink the Scan in MVCC Databases · SIGMOD Conference 2021 |
Machine learning › Time series and sequential data
anomaly detection |
0.5 | 1 | 2021 | Masked Contrastive Learning for Anomaly Detection · IJCAI 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.5 | 1 | 2021 | Masked Contrastive Learning for Anomaly Detection · IJCAI 2021 |
Indexing and storage engines
access methods |
0.5 | 1 | 2021 | Rethink the Scan in MVCC Databases · SIGMOD Conference 2021 |
Indexing and storage engines
multiversion index |
0.5 | 1 | 2021 | Rethink the Scan in MVCC Databases · SIGMOD Conference 2021 |
Transaction processing and concurrency control › concurrency control › multiversion concurrency control
garbage collection |
0.4 | 1 | 2020 | Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020 |
Transaction processing and concurrency control › transaction models
long-lived transactions |
0.4 | 1 | 2020 | Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020 |
Transaction processing and concurrency control › isolation levels
snapshot isolation |
0.4 | 1 | 2020 | Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020 |
Transaction processing and concurrency control
versioning |
0.4 | 1 | 2020 | Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020 |
Methods — techniques the papers use, named apart from their topics
role-play at scale · 1.0persona prompting · 1.0prompting · 0.7prompt engineering · 0.7linear probing · 0.7linear discriminative analysis · 0.7in-context learning · 0.7group action · 0.7fréchet mean · 0.7clustering · 0.7version search structure · 0.5pointer augmentation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CUB: Benchmarking Context Utilisation Techniques for Language ModelsabstractIncorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples. Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein |
ACL (1) | 6 |
| 2026 | Inertia in Moral and Value Judgments of Large Language ModelsabstractLarge Language Models (LLMs) behave nondeterministically, and prompting has become a common method for steering their outputs.A popular strategy is to assign a persona to the model to produce more varied, contextsensitive responses, similar to how responses vary across human individuals.Against the expectation that persona prompting yields a wide range of opinions, our experiments show that LLMs keep consistent value orientations.We observe a persistent inertia in their responses, where certain moral and value dimensions (especially harm avoidance and fairness) stay skewed in one direction across persona settings.To study this, we use role-play at scale, which pairs randomized persona prompts with a macro-level analysis of model outputs.Our results point to strong internal biases and value preferences in LLMs, which we call value orientation and inertia.These models warrant scrutiny and adjustment before use in applications where balanced outputs matter. Bruce W. Lee, Yeongheon Lee, Hyunsoo Cho |
ACL (1) | 3 |
| 2025 | Analyzing the latent space of GAN through local dimension estimation for disentanglement evaluation
Jaewoong Choi, Geonho Hwang, Hyunsoo Cho, Myungjoo Kang |
Pattern Recognit. | 3 |
| 2023 | Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context LearnersabstractThrough in-context learning (ICL), large-scale language models are effective few-shot learners without additional model fine-tuning. However, the ICL performance does not scale well with the number of available training sample as it is limited by the inherent input length constraint of the underlying language model. Meanwhile, many studies have revealed that language models are also powerful feature extractors, allowing them to be utilized in a black-box manner and enabling the linear probing paradigm, where lightweight discriminators are trained on top of the pre-extracted input representations. This paper proposes prompt-augmented linear probing (PALP), a hybrid of linear probing and ICL, which leverages the best of both worlds. PALP inherits the scalability of linear probing and the capability of enforcing language models to derive more meaningful representations via tailoring input into a more conceivable form. Throughout in-depth investigations on various datasets, we verified that PALP significantly closes the gap between ICL in the data-hungry scenario and fine-tuning in the data-abundant scenario with little training overhead, potentially making PALP a strong alternative in a black-box scenario. Hyunsoo Cho, Hyuhng Joon Kim, Junyeob Kim, Sang-Woo Lee 0001, Sang-goo Lee, Kang Min Yoo, Taeuk Kim |
AAAI | 1 |
| 2023 | CELDA: Leveraging Black-box Language Model as Enhanced Classifier without LabelsabstractUtilizing language models (LMs) without internal access is becoming an attractive paradigm in the field of NLP as many cutting-edge LMs are released through APIs and boast a massive scale.The de-facto method in this type of black-box scenario is known as prompting, which has shown progressive performance enhancements in situations where data labels are scarce or unavailable.Despite their efficacy, they still fall short in comparison to fully supervised counterparts and are generally brittle to slight modifications.In this paper, we propose Clustering-Enhanced Linear Discriminative Analysis (CELDA), a novel approach that improves the text classification accuracy with a very weak-supervision signal (i.e., name of the labels).Our framework draws a precise decision boundary without accessing weights or gradients of the LM model or data labels.The core ideas of CELDA are twofold: (1) extracting a refined pseudo-labeled dataset from an unlabeled dataset, and (2) training a lightweight and robust model on the top of LM, which learns an accurate decision boundary from an extracted noisy dataset.Throughout in-depth investigations on various datasets, we demonstrated that CELDA reaches new state-of-theart in weakly-supervised text classification and narrows the gap with a fully-supervised model.Additionally, our proposed methodology can be applied universally to any LM and has the potential to scale to larger models, making it a more viable option for utilizing large LMs. Hyunsoo Cho, Youna Kim, Sang-goo Lee |
ACL (1) | 1 |
| 2023 | Finding the Global Semantic Representation in GAN through Fréchet Mean
Jaewoong Choi, Geonho Hwang, Hyunsoo Cho, Myungjoo Kang |
ICLR | 3 |
| 2023 | MAGANet: Achieving Combinatorial Generalization by Modeling a Group ActionabstractCombinatorial generalization refers to the ability to collect and assemble various attributes from diverse data to generate novel unexperienced data. This ability is considered a necessary passing point for achieving human-level intelligence. To achieve this ability, previous unsupervised approaches mainly focused on learning the disentangled representation, such as the variational autoencoder. However, recent studies discovered that the disentangled representation is insufficient for combinatorial generalization and is not even correlated. In this regard, we propose a novel framework for data generation that can robustly generalize under these distribution shift situations. Instead of representing each data, our model discovers the fundamental transformation between a pair of data by simulating a group action. To test the combinatorial generalizability, we evaluated our model in two settings: Recombination-to-Element and Recombination-to-Range. The experiments demonstrated that our method has quantitatively and qualitatively superior generalizability and generates better images than traditional models. Geonho Hwang, Jaewoong Choi, Hyunsoo Cho, Myungjoo Kang |
ICML | 3 |
| 2022 | Ground-Truth Labels Matter: A Deeper Look into Input-Label DemonstrationsabstractKang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang-goo Lee, Taeuk Kim. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee 0001, Sang-goo Lee, Taeuk Kim |
EMNLP | 4 |
| 2021 | Masked Contrastive Learning for Anomaly DetectionabstractDetecting anomalies is one fundamental aspect of a safety-critical software system, however, it remains a long-standing problem. Numerous branches of works have been proposed to alleviate the complication and have shown promising results. In particular, self-supervised learning based methods are spurring interest due to their capability of learning diverse representations without additional labels. Among self-supervised learning tactics, contrastive learning is one specific framework showing pronounced results in various fields including anomaly detection. However, the primary objective of contrastive learning is to learn task-agnostic features without any labels, which is not entirely suited to discern anomalies. In this paper, we propose a task-specific variant of contrastive learning named masked contrastive learning, which is more befitted for anomaly detection. Moreover, we propose a new inference method dubbed self-ensemble inference that further boosts performance by leveraging the ability learned through auxiliary self-supervision tasks. By combining our models, we can outperform previous state-of-the-art methods by a significant margin on various benchmark datasets. Hyunsoo Cho, Jinseok Seol, Sang-goo Lee |
IJCAI | 1 |
| 2021 | Rethink the Scan in MVCC DatabasesabstractA scan is one of the fundamental operations in databases for retrieving tuples from tables, and research on access methods has been of importance to query optimization. However, our community is aware of the inconvenient truth that its performance may plummet amid steep increases in search costs when acting on MVCC databases since multi-versioning may forfeit all the benefits of using database indexes. An execution plan for a query on multi-versioned data often comprises a series of point lookup operations, of which each internally executes a linear traversal of record versions. Therefore, the generated plan is surprisingly worse than a full table (or version store) scan, mainly due to redundant access to database pages. To address such an all-or-nothing approach, we propose version weaver (vWeaver), a light-weight access method for record versions, that expedites a scan on record versions with each being augmented by just a few pointer fields. vWeaver incrementally constructs a version search structure over even an append-only version store (e.g., undo space) and allows a scan to traverse new version search structures for fast lookup. We applied vWeaver to in-memory and disk-based MVCC databases and demonstrated that the systems with vWeaver generally improved scan performance under various workloads with negligible space overhead. Jong-Bin Kim, Kihwang Kim, Hyunsoo Cho, Jaeseon Yu, Sooyong Kang, Hyungsoo Jung 0001 |
SIGMOD Conference | 3 |
| 2020 | Long-lived Transactions Made Less HarmfulabstractMany systems use snapshot isolation, or something similar, as defaults, and multi-version concurrency control (MVCC) remains essential to offering such point-in-time consistency. One major issue in MVCC is the timely removal of unnecessary versions of data items, especially in the presence of long-lived transactions (LLTs). We have observed that the latest versions of MySQL and PostgreSQL are still vulnerable to LLTs. Our analysis of existing proposals suggests that new solutions to this matter must provide rigorous rules for completely identifying unnecessary versions, and elaborate designs for version cleaning lest old versions required for LLTs should suspend garbage collection. In this paper, we formalize such rules into our version pruning theorem and version classification, of which all form theoretical foundations for our new version management system, vDriver, that bases its record versioning on a new principle: Single In-row Remaining Off-row (SIRO) versioning. We implemented a prototype of vDriver and integrated it with MySQL-8.0 and PostgreSQL-12.0. The experimental evaluation demonstrated that the engines with Driver continue to perform the reclamation of dead versions in the face of LLTs while retaining transaction throughput with reduced space consumption. Jong-Bin Kim, Hyunsoo Cho, Kihwang Kim, Jaeseon Yu, Sooyong Kang, Hyungsoo Jung 0001 |
SIGMOD Conference | 2 |
| 2018 | Automatic Generation of Multiple-Choice Fill-in-the-Blank Question Using Document Embedding
Junghyuk Park, Hyunsoo Cho, Sang-goo Lee |
AIED (2) | 2 |
| 2017 | Improving Flash Storage Performance by Caching Address Mapping Table in Host Memory
Wookhan Jeong, Hyunsoo Cho, Yongmyung Lee, Jaegyu Lee, SongHo Yoon, Joo Young Hwang, Dong-Gi Lee |
HotStorage | 2 |