VLDB 2026 Research / reviewers in the wild / expert
Sang Keun Choe
dblp:211/6811
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 32% Trustworthy machine learning · 27% Learning paradigms · 14% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 77% High-performance computing · 23% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.2 | 2 | 2023 | Making Scalable Meta Learning Practical · NeurIPS 2023 Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning · OSDI 2021 |
Machine learning › Trustworthy machine learning › Data-centric AI
data valuation |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Learning paradigms › continual learning
gradient projection |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability › training data attribution
influence function |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › large-scale learning
scalable training |
0.9 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
automatic differentiation |
0.7 | 1 | 2023 | Betty: An Automatic Differentiation Library for Multilevel Optimization · ICLR 2023 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.7 | 1 | 2023 | Making Scalable Meta Learning Practical · NeurIPS 2023 |
Mathematical optimization
bilevel optimization |
0.7 | 1 | 2023 | Betty: An Automatic Differentiation Library for Multilevel Optimization · ICLR 2023 |
Mathematical optimization
multilevel optimization |
0.7 | 1 | 2023 | Betty: An Automatic Differentiation Library for Multilevel Optimization · ICLR 2023 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.5 | 1 | 2021 | Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning · OSDI 2021 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions · NeurIPS 2025 |
Machine learning › Optimization for machine learning
implicit differentiation |
0.2 | 1 | 2023 | Making Scalable Meta Learning Practical · NeurIPS 2023 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2021 | Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning · OSDI 2021 |
Methods — techniques the papers use, named apart from their topics
automatic differentiation · 1.3influence functions · 0.9gradient projection · 0.9implicit differentiation · 0.7first-order gradient · 0.7distributed training · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsabstractLarge language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve its scalability with an efficient gradient projection strategy called LoGra that leverages the gradient structure in backpropagation. We then provide a theoretical motivation of gradient projection approaches to influence functions to promote trust in the data valuation process. Lastly, we lower the barrier to implementing data valuation systems by introducing LogIX, a software package that can transform existing training code into data valuation code with minimal effort. In our data valuation experiments, LoGra achieves competitive accuracy against more expensive baselines while showing up to 6,500x improvement in throughput and 5x reduction in GPU memory usage when applied to Llama3-8B-Instruct and the 1B-token dataset. Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff G. Schneider, Eduard H. Hovy, Roger B. Grosse, Eric P. Xing |
NeurIPS | 1 |
| 2023 | Betty: An Automatic Differentiation Library for Multilevel Optimization
Sang Keun Choe, Willie Neiswanger, Pengtao Xie, Eric P. Xing |
ICLR | 1 |
| 2023 | Making Scalable Meta Learning PracticalabstractDespite its flexibility to learn diverse inductive biases in machine learning programs, meta learning (i.e.,\ learning to learn) has long been recognized to suffer from poor scalability due to its tremendous compute/memory costs, training instability, and a lack of efficient distributed training support. In this work, we focus on making scalable meta learning practical by introducing SAMA, which combines advances in both implicit differentiation algorithms and systems. Specifically, SAMA is designed to flexibly support a broad range of adaptive optimizers in the base level of meta learning programs, while reducing computational burden by avoiding explicit computation of second-order gradient information, and exploiting efficient distributed training techniques implemented for first-order gradients. Evaluated on multiple large-scale meta learning benchmarks, SAMA showcases up to 1.7/4.8x increase in throughput and 2.0/3.8x decrease in memory consumption respectively on single-/multi-GPU setups compared to other baseline meta learning algorithms. Furthermore, we show that SAMA-based data optimization leads to consistent improvements in text classification accuracy with BERT and RoBERTa large language models, and achieves state-of-the-art results in both small- and large-scale data pruning on image classification tasks, demonstrating the practical applicability of scalable meta learning across language and vision domains. Sang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger, Pengtao Xie, Emma Strubell, Eric P. Xing |
NeurIPS | 1 |
| 2021 | Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang 0025, Gregory R. Ganger, Eric P. Xing |
OSDI | 2 |
| 2019 | On Leveraging the Visual Modality for Neural Machine TranslationabstractLeveraging the visual modality effectively for Neural Machine Translation (NMT) remains an open problem in computational linguistics.Recently, Caglayan et al. posit that the observed gains are limited mainly due to the very simple, short, repetitive sentences of the Multi30k dataset (the only multimodal MT dataset available at the time), which renders the source text sufficient for context.In this work, we further investigate this hypothesis on a new large scale multimodal Machine Translation (MMT) dataset, How2, which has 1.57 times longer mean sentence length than Multi30k and no repetition.We propose and evaluate three novel fusion techniques, each of which is designed to ensure the utilization of visual context at different stages of the Sequence-to-Sequence transduction pipeline, even under full linguistic context.However, we still obtain only marginal gains under full linguistic context and posit that visual embeddings extracted from deep vision models (ResNet for Multi30k, ResNext for How2) do not lend themselves to increasing the discriminativeness between the vocabulary elements at token level prediction in NMT.We demonstrate this qualitatively by analyzing attention distribution and quantitatively through Principal Component Analysis, arriving at the conclusion that it is the quality of the visual embeddings rather than the length of sentences, which need to be improved in existing MMT datasets. Vikas Raunak, Sang Keun Choe, Quanyang Lu, Florian Metze |
INLG | 2 |
| 2018 | Cover Song Identification Using Song-to-Song Cross-Similarity Matrix with Convolutional Neural NetworkabstractIn this paper, we propose a cover song identification algorithm using a convolutional neural network (CNN). We first train the CNN model to classify any non-/cover relationship, by feeding a cross-similarity matrix that is generated from a pair of songs as an input. Our main idea is to use the CNN output-the cover-probabilities of one song to all other candidate songs-as a new representation vector for measuring the distance between songs. Based on this, the present algorithm searches cover songs by applying several ranking methods: 1. sorting without using the representation vectors; 2. the cosine distance between the representation vectors; and 3. the correlation between the vectors. In our experiment, the proposed algorithm significantly outperformed the algorithms used in recent studies, by achieving a mean average precision (MAP) of 93.18% in a dataset consisting of 3,300 cover-pairs and 496,200 non-cover-pairs. Juheon Lee, Sungkyun Chang, Sang Keun Choe, Kyogu Lee |
ICASSP | 3 |