VLDB 2026 Research / reviewers in the wild / expert
Jun Woo Chung
dblp:353/7742
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0004-3414-6283ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 53% Indexing and storage engines · 27% Machine learning and data management · 20% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning › model intellectual property protection
model watermarking |
1.0 | 1 | 2026 | Robust Watermarking on Gradient Boosting Decision Trees · AAAI 2026 |
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search |
0.9 | 1 | 2025 | Locality-Sensitive Indexing for Graph-Based Approximate Nearest Neighbor Search · SIGIR 2025 |
Information retrieval › similarity search › nearest neighbor search › approximate nearest neighbor search
graph-based ANNS |
0.9 | 1 | 2025 | Locality-Sensitive Indexing for Graph-Based Approximate Nearest Neighbor Search · SIGIR 2025 |
Indexing and storage engines
index maintenance |
0.9 | 1 | 2025 | Locality-Sensitive Indexing for Graph-Based Approximate Nearest Neighbor Search · SIGIR 2025 |
Machine learning › Trustworthy machine learning
machine unlearning |
0.7 | 1 | 2023 | Machine Unlearning in Gradient Boosting Decision Trees · KDD 2023 |
Algorithms and data structures › data structure design › search structures › hashing
locality-sensitive hashing |
0.3 | 1 | 2025 | Locality-Sensitive Indexing for Graph-Based Approximate Nearest Neighbor Search · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
watermark embedding · 2.0in-place fine-tuning · 2.0proximity graph · 1.7locality-sensitive hashing · 1.7lazy update · 1.3gradient boosting decision tree · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Watermarking on Gradient Boosting Decision TreesabstractGradient Boosting Decision Trees (GBDTs) are widely used in industry and academia for their high accuracy and efficiency, particularly on structured data. However, the subject of watermarking GBDT models remains underexplored, especially compared to neural networks. In this work, we present the first robust watermarking framework tailored to GBDT models, utilizing in-place fine-tuning to embed imperceptible and resilient watermarks. We propose four embedding strategies, each designed to minimize impact on model accuracy while ensuring watermark robustness. Through experiments across diverse datasets, we demonstrate that our methods achieve high watermark embedding rates, low accuracy degradation, and strong resistance to post-deployment fine-tuning. Jun Woo Chung, Yingjie Lao, Weijie Zhao 0001 |
AAAI | 1 |
| 2025 | Locality-Sensitive Indexing for Graph-Based Approximate Nearest Neighbor SearchabstractThe burgeoning size of modern text datasets has heightened the need for efficient text retrieval systems. For such applications, Approximate Nearest Neighbor (ANN) search algorithms, and in particular graph-based methods have long been established as the leading approach in terms of recall and search speed. However, the data and execution dependencies of vertices increase the construction workload and complicate maintenance processes for the constructed index. In this paper, we present Locality-Sensitive Indexing for Graph-Based Search (or LIGS), which utilizes independent locality-sensitive hashing algorithms to simulate a proximity graph, on which a standard graph search can be performed. We show that LIGS offers substantially faster maintenance (insertion/deletion) speeds and better conservation of graph quality compared to state-of-the-art graph-based ANN methods, demonstrating LIGS as a promising alternative for maintenance-heavy scenarios. Jun Woo Chung, Huawei Lin 0001, Weijie Zhao 0001 |
SIGIR | 1 |
| 2023 | Machine Unlearning in Gradient Boosting Decision TreesabstractVarious machine learning applications take users' data to train the models. Recently enforced legislation requires companies to remove users' data upon requests, i.e.,the right to be forgotten. In the context of machine learning, the trained model potentially memorizes the training data. Machine learning algorithms have to be able to unlearn the user data that are requested to delete to meet the requirement. Gradient Boosting Decision Trees (GBDT) is a widely deployed model in many machine learning applications. However, few studies investigate the unlearning on GBDT. This paper proposes a novel unlearning framework for GBDT. To the best of our knowledge, this is the first work that considers machine unlearning on GBDT. It is not straightforward to transfer the unlearning methods of DNN to GBDT settings. We formalized the machine unlearning problem and its relaxed version. We propose an unlearning framework that efficiently and effectively unlearns a given collection of data without retraining the model from scratch. We introduce a collection of techniques, including random split point selection and random partitioning layers training, to the training process of the original tree models to ensure that the trained model requires few subtree retrainings during the unlearning. We investigate the intermediate data and statistics to store as an auxiliary data structure during the training so that we can immediately determine if a subtree is required to be retrained without touching the original training dataset. Furthermore, a lazy update technique is proposed as a trade-off between unlearning time and model functionality. We experimentally evaluate our proposed methods on public datasets. The empirical results confirm the effectiveness of our framework. Huawei Lin 0001, Jun Woo Chung, Yingjie Lao, Weijie Zhao 0001 |
KDD | 2 |