Hyunsoo Cho

dblp:86/125 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 40% Representation and self-supervised learning · 17% Knowledge representation and reasoning · 11%
Databases, data mining, and information retrieval
2 papers
Transaction processing and concurrency control · 70% Indexing and storage engines · 30%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
1.222023
Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context Learners · AAAI 2023
Ground-Truth Labels Matter: A Deeper Look into Input-Label Demonstrations · EMNLP 2022
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization
1.012026
CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
fairness
1.012026
Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
moral judgment
1.012026
Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › prompting
persona prompting
1.012026
Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
CUB: Benchmarking Context Utilisation Techniques for Language Models · ACL (1) 2026
Machine learning › Learning theory › generalization
combinatorial generalization
0.712023
MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.712023
MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action · ICML 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context Learners · AAAI 2023
Machine learning › Generative modeling
generative adversarial network
0.712023
Finding the Global Semantic Representation in GAN through Fréchet Mean · ICLR 2023
Machine learning › Representation and self-supervised learning › equivariance
group equivariance
0.712023
MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action · ICML 2023
Machine learning › Representation and self-supervised learning
latent space
0.712023
Finding the Global Semantic Representation in GAN through Fréchet Mean · ICLR 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation
0.712023
Finding the Global Semantic Representation in GAN through Fréchet Mean · ICLR 2023
Natural language and speech › Information extraction and text analysis
text classification
0.712023
CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels · ACL (1) 2023
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.712023
CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels · ACL (1) 2023
Transaction processing and concurrency control › concurrency control
multiversion concurrency control
0.622021
Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020
Rethink the Scan in MVCC Databases · SIGMOD Conference 2021
Machine learning › Time series and sequential data
anomaly detection
0.512021
Masked Contrastive Learning for Anomaly Detection · IJCAI 2021
Machine learning › Representation and self-supervised learning
contrastive learning
0.512021
Masked Contrastive Learning for Anomaly Detection · IJCAI 2021
Indexing and storage engines
access methods
0.512021
Rethink the Scan in MVCC Databases · SIGMOD Conference 2021
Indexing and storage engines
multiversion index
0.512021
Rethink the Scan in MVCC Databases · SIGMOD Conference 2021
Transaction processing and concurrency control › concurrency control › multiversion concurrency control
garbage collection
0.412020
Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020
Transaction processing and concurrency control › transaction models
long-lived transactions
0.412020
Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020
Transaction processing and concurrency control › isolation levels
snapshot isolation
0.412020
Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020
Transaction processing and concurrency control
versioning
0.412020
Long-lived Transactions Made Less Harmful · SIGMOD Conference 2020

Methods — techniques the papers use, named apart from their topics

role-play at scale · 1.0persona prompting · 1.0prompting · 0.7prompt engineering · 0.7linear probing · 0.7linear discriminative analysis · 0.7in-context learning · 0.7group action · 0.7fréchet mean · 0.7clustering · 0.7version search structure · 0.5pointer augmentation · 0.5
YearPublicationVenuePosition
2026 CUB: Benchmarking Context Utilisation Techniques for Language Models
abstract
Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples.
Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein
ACL (1)6
2026 Inertia in Moral and Value Judgments of Large Language Models
abstract
Large Language Models (LLMs) behave nondeterministically, and prompting has become a common method for steering their outputs.A popular strategy is to assign a persona to the model to produce more varied, contextsensitive responses, similar to how responses vary across human individuals.Against the expectation that persona prompting yields a wide range of opinions, our experiments show that LLMs keep consistent value orientations.We observe a persistent inertia in their responses, where certain moral and value dimensions (especially harm avoidance and fairness) stay skewed in one direction across persona settings.To study this, we use role-play at scale, which pairs randomized persona prompts with a macro-level analysis of model outputs.Our results point to strong internal biases and value preferences in LLMs, which we call value orientation and inertia.These models warrant scrutiny and adjustment before use in applications where balanced outputs matter.
Bruce W. Lee, Yeongheon Lee, Hyunsoo Cho
ACL (1)3
2025 Analyzing the latent space of GAN through local dimension estimation for disentanglement evaluation
Jaewoong Choi, Geonho Hwang, Hyunsoo Cho, Myungjoo Kang
Pattern Recognit.3
2023 Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context Learners
abstract
Through in-context learning (ICL), large-scale language models are effective few-shot learners without additional model fine-tuning. However, the ICL performance does not scale well with the number of available training sample as it is limited by the inherent input length constraint of the underlying language model. Meanwhile, many studies have revealed that language models are also powerful feature extractors, allowing them to be utilized in a black-box manner and enabling the linear probing paradigm, where lightweight discriminators are trained on top of the pre-extracted input representations. This paper proposes prompt-augmented linear probing (PALP), a hybrid of linear probing and ICL, which leverages the best of both worlds. PALP inherits the scalability of linear probing and the capability of enforcing language models to derive more meaningful representations via tailoring input into a more conceivable form. Throughout in-depth investigations on various datasets, we verified that PALP significantly closes the gap between ICL in the data-hungry scenario and fine-tuning in the data-abundant scenario with little training overhead, potentially making PALP a strong alternative in a black-box scenario.
Hyunsoo Cho, Hyuhng Joon Kim, Junyeob Kim, Sang-Woo Lee 0001, Sang-goo Lee, Kang Min Yoo, Taeuk Kim
AAAI1
2023 CELDA: Leveraging Black-box Language Model as Enhanced Classifier without Labels
abstract
Utilizing language models (LMs) without internal access is becoming an attractive paradigm in the field of NLP as many cutting-edge LMs are released through APIs and boast a massive scale.The de-facto method in this type of black-box scenario is known as prompting, which has shown progressive performance enhancements in situations where data labels are scarce or unavailable.Despite their efficacy, they still fall short in comparison to fully supervised counterparts and are generally brittle to slight modifications.In this paper, we propose Clustering-Enhanced Linear Discriminative Analysis (CELDA), a novel approach that improves the text classification accuracy with a very weak-supervision signal (i.e., name of the labels).Our framework draws a precise decision boundary without accessing weights or gradients of the LM model or data labels.The core ideas of CELDA are twofold: (1) extracting a refined pseudo-labeled dataset from an unlabeled dataset, and (2) training a lightweight and robust model on the top of LM, which learns an accurate decision boundary from an extracted noisy dataset.Throughout in-depth investigations on various datasets, we demonstrated that CELDA reaches new state-of-theart in weakly-supervised text classification and narrows the gap with a fully-supervised model.Additionally, our proposed methodology can be applied universally to any LM and has the potential to scale to larger models, making it a more viable option for utilizing large LMs.
Hyunsoo Cho, Youna Kim, Sang-goo Lee
ACL (1)1
2023 Finding the Global Semantic Representation in GAN through Fréchet Mean
Jaewoong Choi, Geonho Hwang, Hyunsoo Cho, Myungjoo Kang
ICLR3
2023 MAGANet: Achieving Combinatorial Generalization by Modeling a Group Action
abstract
Combinatorial generalization refers to the ability to collect and assemble various attributes from diverse data to generate novel unexperienced data. This ability is considered a necessary passing point for achieving human-level intelligence. To achieve this ability, previous unsupervised approaches mainly focused on learning the disentangled representation, such as the variational autoencoder. However, recent studies discovered that the disentangled representation is insufficient for combinatorial generalization and is not even correlated. In this regard, we propose a novel framework for data generation that can robustly generalize under these distribution shift situations. Instead of representing each data, our model discovers the fundamental transformation between a pair of data by simulating a group action. To test the combinatorial generalizability, we evaluated our model in two settings: Recombination-to-Element and Recombination-to-Range. The experiments demonstrated that our method has quantitatively and qualitatively superior generalizability and generates better images than traditional models.
Geonho Hwang, Jaewoong Choi, Hyunsoo Cho, Myungjoo Kang
ICML3
2022 Ground-Truth Labels Matter: A Deeper Look into Input-Label Demonstrations
abstract
Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang-goo Lee, Taeuk Kim. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee 0001, Sang-goo Lee, Taeuk Kim
EMNLP4
2021 Masked Contrastive Learning for Anomaly Detection
abstract
Detecting anomalies is one fundamental aspect of a safety-critical software system, however, it remains a long-standing problem. Numerous branches of works have been proposed to alleviate the complication and have shown promising results. In particular, self-supervised learning based methods are spurring interest due to their capability of learning diverse representations without additional labels. Among self-supervised learning tactics, contrastive learning is one specific framework showing pronounced results in various fields including anomaly detection. However, the primary objective of contrastive learning is to learn task-agnostic features without any labels, which is not entirely suited to discern anomalies. In this paper, we propose a task-specific variant of contrastive learning named masked contrastive learning, which is more befitted for anomaly detection. Moreover, we propose a new inference method dubbed self-ensemble inference that further boosts performance by leveraging the ability learned through auxiliary self-supervision tasks. By combining our models, we can outperform previous state-of-the-art methods by a significant margin on various benchmark datasets.
Hyunsoo Cho, Jinseok Seol, Sang-goo Lee
IJCAI1
2021 Rethink the Scan in MVCC Databases
abstract
A scan is one of the fundamental operations in databases for retrieving tuples from tables, and research on access methods has been of importance to query optimization. However, our community is aware of the inconvenient truth that its performance may plummet amid steep increases in search costs when acting on MVCC databases since multi-versioning may forfeit all the benefits of using database indexes. An execution plan for a query on multi-versioned data often comprises a series of point lookup operations, of which each internally executes a linear traversal of record versions. Therefore, the generated plan is surprisingly worse than a full table (or version store) scan, mainly due to redundant access to database pages. To address such an all-or-nothing approach, we propose version weaver (vWeaver), a light-weight access method for record versions, that expedites a scan on record versions with each being augmented by just a few pointer fields. vWeaver incrementally constructs a version search structure over even an append-only version store (e.g., undo space) and allows a scan to traverse new version search structures for fast lookup. We applied vWeaver to in-memory and disk-based MVCC databases and demonstrated that the systems with vWeaver generally improved scan performance under various workloads with negligible space overhead.
Jong-Bin Kim, Kihwang Kim, Hyunsoo Cho, Jaeseon Yu, Sooyong Kang, Hyungsoo Jung 0001
SIGMOD Conference3
2020 Long-lived Transactions Made Less Harmful
abstract
Many systems use snapshot isolation, or something similar, as defaults, and multi-version concurrency control (MVCC) remains essential to offering such point-in-time consistency. One major issue in MVCC is the timely removal of unnecessary versions of data items, especially in the presence of long-lived transactions (LLTs). We have observed that the latest versions of MySQL and PostgreSQL are still vulnerable to LLTs. Our analysis of existing proposals suggests that new solutions to this matter must provide rigorous rules for completely identifying unnecessary versions, and elaborate designs for version cleaning lest old versions required for LLTs should suspend garbage collection. In this paper, we formalize such rules into our version pruning theorem and version classification, of which all form theoretical foundations for our new version management system, vDriver, that bases its record versioning on a new principle: Single In-row Remaining Off-row (SIRO) versioning. We implemented a prototype of vDriver and integrated it with MySQL-8.0 and PostgreSQL-12.0. The experimental evaluation demonstrated that the engines with Driver continue to perform the reclamation of dead versions in the face of LLTs while retaining transaction throughput with reduced space consumption.
Jong-Bin Kim, Hyunsoo Cho, Kihwang Kim, Jaeseon Yu, Sooyong Kang, Hyungsoo Jung 0001
SIGMOD Conference2
2018 Automatic Generation of Multiple-Choice Fill-in-the-Blank Question Using Document Embedding
Junghyuk Park, Hyunsoo Cho, Sang-goo Lee
AIED (2)2
2017 Improving Flash Storage Performance by Caching Address Mapping Table in Host Memory
Wookhan Jeong, Hyunsoo Cho, Yongmyung Lee, Jaegyu Lee, SongHo Yoon, Joo Young Hwang, Dong-Gi Lee
HotStorage2