VLDB 2026 Research / reviewers in the wild / expert
Woohwan Jung
dblp:193/7295
· DBLP profile ↗
18ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-4561-2214ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Single Model Ensemble Framework for Neural Machine Translation Using Pivot TranslationabstractDespite the recent remarkable advances in neural machine translation, translation quality for low-resource language pairs remains subpar. Ensembling multiple systems is a widely adopted technique to enhance performance, often accomplished by combining probability distributions. However, previous approaches face the challenge of high computational costs for training multiple models. Furthermore, for black-box models, averaging token-level probabilities at each decoding step is not feasible. To address the problems of multi-model ensemble methods, we present a pivot-based single model ensemble. The proposed strategy consists of two steps: pivot-based candidate generation and post-hoc aggregation. In the first step, we generate candidates through pivot translation. This can be achieved with only a single model and facilitates knowledge transfer from high-resource pivot languages, resulting in candidates that are not only diverse but also more accurate. Next, in the aggregation step, we select k high-quality candidates from the generated candidates and merge them to generate a final translation that outperforms the existing candidates. Our experimental results show that our method produces translations of superior quality by leveraging candidates from pivot translation to capture the subtle nuances of the source sentence. Seokjin Oh, Keonwoong Noh, Woohwan Jung |
LREC | 3 |
| 2025 | DP-FROST: Differentially Private Fine-tuning of Pre-trained Models with Freezing Model ParametersabstractTraining models with differential privacy has received a lot of attentions since differential privacy provides theoretical guarantee of privacy preservation. For a task in a specific domain, since a large-scale pre-trained model in the same domain contains general knowledge of the task, using such a model requires less effort in designing and training the model. However, differentially privately fine-tuning such models having a large number of trainable parameters results in large degradation of utility. Thus, we propose methods that effectively fine-tune the large-scale pre-trained models with freezing unimportant parameters for downstream tasks while satisfying differential privacy. To select the parameters to be fine-tuned, we propose several efficient methods based on the gradients of model parameters. We show the effectiveness of the proposed method by performing experiments with real datasets. Daeyoung Hong, Woohwan Jung, Kyuseok Shim |
COLING | 2 |
| 2025 | Improving Detail in Pluralistic Image Inpainting with Feature DequantizationabstractPluralistic Image Inpainting (PII) offers multiple plausible solutions for restoring missing parts of images and has been successfully applied to various applications including image editing and object removal. Recently, VQGANbased methods have been proposed and have shown that they significantly improve the structural integrity in the generated images. Nevertheless, the state-of-the-art VQGANbased model PUT faces a critical challenge: degradation of detail quality in output images due to feature quantization. Feature quantization restricts the latent space and causes information loss, which negatively affects the detail quality essential for image inpainting. To tackle the problem, we propose the FDM (Feature Dequantization Module) specifically designed to restore the detail quality of images by compensating for the information loss. Furthermore, we develop an efficient training method for FDM which drastically reduces training costs. We empirically demonstrate that our method significantly enhances the detail quality of the generated images with negligible training and inference overheads. The code is available at https://github.com/hyudsl/FDM Kyungri Park, Woohwan Jung |
WACV | 2 |
| 2025 | Cardinality Estimation of LIKE Predicate Queries using Deep LearningabstractCardinality estimation of LIKE predicate queries has an important role in the query optimization of database systems. Traditional approaches generally use a summary of text data with some statistical assumptions. Recently, the deep learning model for cardinality estimation of LIKE predicate queries has been investigated. To provide more accurate cardinality estimates and reduce the maximum estimation errors, we propose a deep learning model that utilizes the extended N -gram table and the conditional regression header. We next investigate how to efficiently generate training data. Our LEADER (LikE predicate trAining Data gEneRation) algorithms utilize the shareable results across the relational queries corresponding to the LIKE predicates. By analyzing the queries corresponding to LIKE predicates, we develop an efficient join method and utilize the join order for fast query execution and maximal sharing of shareable results . Extensive experiments with real-life datasets confirm the efficiency of the proposed training data generation algorithms and the effectiveness of the proposed model. Suyong Kwon, Kyuseok Shim, Woohwan Jung |
Proc. ACM Manag. Data | 3 |
| 2024 | Beyond Reference: Evaluating High Quality Translations Better than Human ReferencesabstractIn Machine Translation (MT) evaluations, the conventional approach is to compare a translated sentence against its human-created reference sentence.MT metrics provide an absolute score (e.g., from 0 to 1) to a candidate sentence based on the similarity with the reference sentence.Thus, existing MT metrics give the maximum score to the reference sentence.However, this approach overlooks the potential for a candidate sentence to exceed the reference sentence in terms of quality.In particular, recent advancements in Large Language Models (LLMs) have highlighted this issue, as LLMgenerated sentences often exceed the quality of human-written sentences.To address the problem, we introduce the Residual score Metric (RESUME), which evaluates the relative quality between reference and candidate sentences.RESUME assigns a positive score to candidate sentences that outperform their reference sentences, and a negative score when they fall short.By adding the residual scores from RE-SUME to the absolute scores from MT metrics, it can be possible to allocate higher scores to candidate sentences than what reference sentences are received from MT metrics.Experimental results demonstrate that RESUME enhances the alignments between MT metrics and human judgments both at the segment-level and the system-level. Keonwoong Noh, Seokjin Oh, Woohwan Jung |
EMNLP | 3 |
| 2024 | Improving Domain-Specific ASR with LLM-Generated Contextual Descriptions
Jiwon Suh, Injae Na, Woohwan Jung |
INTERSPEECH | 3 |
| 2024 | Encoder-Based Multimodal Ensemble Learning for High Compatibility and Accuracy in Phishing Website Detection
Jemin Ahn, Dorian Akhavan, Woohwan Jung, Kyungtae Kang, Junggab Son |
SecureComm (3) | 3 |
| 2023 | THUNDER: Named Entity Recognition Using a Teacher-Student Model with Dual Classifiers for Strong and Weak SupervisionsabstractStrong and weak supervisions have complementary characteristics. However, utilizing both supervisions for named entity recognition (NER) has not been extensively studied. Moreover, the existing works address only incomplete annotations and neglects inaccurate annotations during NER model training. To effectively utilize weak labels, we introduce an auxiliary classifier that learns from weak labels. Furthermore, we adopt the teacher-student framework to handle both incomplete and inaccurate weak labels. A teacher model is first trained using both strongly and weakly supervised data, and next generates pseudo labels to replace weak labels. Then, the student model is trained so that the main classifier learns from both strong labels and confident pseudo labels while the auxiliary classifier learns from less confident pseudo labels. We also incorporate data augmentation through ChatGPT to generate additional annotated sentences to improve model performance and generalization capabilities. The experimental results with different weak supervisions demonstrate that our proposed method surpasses existing techniques. Seongwoong Oh, Woohwan Jung, Kyuseok Shim |
ECAI | 2 |
| 2023 | Enhancing Low-resource Fine-grained Named Entity Recognition by Leveraging Coarse-grained DatasetsabstractNamed Entity Recognition (NER) frequently suffers from the problem of insufficient labeled data, particularly in fine-grained NER scenarios.Although K-shot learning techniques can be applied, their performance tends to saturate when the number of annotations exceeds several tens of labels.To overcome this problem, we utilize existing coarse-grained datasets that offer a large number of annotations.A straightforward approach to address this problem is prefinetuning, which employs coarse-grained data for representation learning.However, it cannot directly utilize the relationships between finegrained and coarse-grained entities, although a fine-grained entity type is likely to be a subcategory of a coarse-grained entity type.We propose a fine-grained NER model with a Fineto-Coarse(F2C) mapping matrix to leverage the hierarchical structure explicitly.In addition, we present an inconsistency filtering method to eliminate coarse-grained entities that are inconsistent with fine-grained entity types to avoid performance degradation.Our experimental results show that our method outperforms both K-shot learning and supervised learning methods when dealing with a small number of fine-grained annotations.Code is available at https://github.com/sue991/CoFiNER. Su Ah Lee, Seokjin Oh, Woohwan Jung |
EMNLP | 3 |
| 2023 | Collecting Geospatial Data Under Local Differential Privacy With Improving Frequency EstimationabstractGeospatial data provides a lot of benefits for personalized services. However, since the geospatial data contains sensitive information about personal activities, collecting the raw data has a potential risk of leaking private information from the collectors. Recently, local differential privacy (LDP), which protects the privacy of users without trusting the collector, has been adopted to preserve privacy in many real applications. In this paper, we investigate the problem of collecting the locations of individual users under LDP, and propose a perturbation mechanism designed carefully to minimize the expected error of perturbed locations according to the privacy budget and the data domain. The frequency distribution of perturbed locations inevitably has a large error. To tackle the problem, we also propose a postprocessing algorithm to estimate the original frequency distribution of collected data by using convex optimization. By experiments with various real datasets, we show the effectiveness of the proposed algorithms. Daeyoung Hong, Woohwan Jung, Kyuseok Shim |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Cardinality Estimation of Approximate Substring Queries using Deep LearningabstractCardinality estimation of an approximate substring query is an important problem in database systems. Traditional approaches build a summary from the text data and estimate the cardinality using the summary with some statistical assumptions. Since deep learning models can learn underlying complex data patterns effectively, they have been successfully applied and shown to outperform traditional methods for cardinality estimations of queries in database systems. However, since they are not yet applied to approximate substring queries, we investigate a deep learning approach for cardinality estimation of such queries. Although the accuracy of deep learning models tends to improve as the train data size increases, producing a large train data is computationally expensive for cardinality estimation of approximate substring queries. Thus, we develop efficient train data generation algorithms by avoiding unnecessary computations and sharing common computations. We also propose a deep learning model as well as a novel learning method to quickly obtain an accurate deep learning-based estimator. Extensive experiments confirm the superiority of our data generation algorithms and deep learning model with the novel learning method. Suyong Kwon, Woohwan Jung, Kyuseok Shim |
Proc. VLDB Endow. | 2 |
| 2021 | Collecting Geospatial Data with Local Differential Privacy for Personalized ServicesabstractGeospatial data provides a lot of benefits for personalized services. However, since the geospatial data contains sensitive information about personal activities, collecting the raw data has a potential risk of leaking private information from the collectors. Recently, local differential privacy (LDP), which protects the privacy of users without trusting the collector, has been adopted to preserve privacy in many real applications. However, most of existing LDP algorithms focus on obtaining aggregated values such as mean and histogram from the collected data. In this paper, we investigate the problem of collecting the locations of individual users under LDP, and propose a perturbation mechanism designed carefully to reduce the error of each perturbed location according to the privacy budget and the domain size. In addition, we show the effectiveness of the proposed algorithm through experiments on various real datasets. Daeyoung Hong, Woohwan Jung, Kyuseok Shim |
ICDE | 2 |
| 2021 | TIDY: Publishing a Time Interval Dataset With Differential PrivacyabstractLog data from mobile devices generally contain a series of events with temporal information including time intervals which consist of the start and finish times. However, the problem of releasing differentially private time interval datasets has not been tackled yet. A time interval dataset can be represented by a two dimensional (2D) histogram. Most of the methods to publish 2D histograms partition the data into rectangular spaces to reduce the aggregated noise error for range queries. However, the existing algorithms to publish 2D histograms suffer from the structural error when applied to time interval datasets. To reduce the aggregated noise errors and suppress the increase in the structural error, we propose the TIDY (publishing Time Intervals via Differential privacY) algorithm. We use the frequency vectors as a compact representation of the time interval dataset. After applying the Laplace mechanism to the frequency vectors, we improve the utility of the frequency vectors based on a maximum likelihood estimation. We also develop a new partitioning method adapted for the frequency vectors to balance the trade-off between the noise and structural errors. Our empirical study on real-life and synthetic datasets confirms that TIDY outperforms the existing algorithms for 2D histograms. Woohwan Jung, Suyong Kwon, Kyuseok Shim |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | T-REX: A Topic-Aware Relation Extraction ModelabstractDocument-level relation extraction (RE) has recently received a lot of attention. However, existing models for document-level RE have similar structures to the models for sentence-level RE. Thus, they still do not consider some unique characteristics of the new problem setting. For example, in Wikipedia, there is a title for each page and it usually represents the topic entity that is mainly described on the page. In many cases, the topic entity is omitted in the text. Thus, existing RE models often fail to find the relations with the omitted topic entity. To tackle the problem, we propose a Topic-aware Relation EXtraction (T-REX) model. To extract the relations with the (possibly omitted) topic entity, the proposed model first encodes the topic entity by aggregating the information of all its mentions in the document. Then it finds the relations between the topic entity and each mention of other entities. Finally, the output layer combines the mention-wise results and outputs all relations expressed in the document. Our performance study with a large-scale dataset confirms the effectiveness of the T-REX model. Woohwan Jung, Kyuseok Shim |
CIKM | 1 |
| 2020 | Dual Supervision Framework for Relation Extraction with Distant Supervision and Human AnnotationabstractRelation extraction (RE) has been extensively studied due to its importance in real-world applications such as knowledge base construction and question answering.Most of the existing works train the models on either distantly supervised data or human-annotated data.To take advantage of the high accuracy of human annotation and the cheap cost of distant supervision, we propose the dual supervision framework which effectively utilizes both types of data.However, simply combining the two types of data to train a RE model may decrease the prediction accuracy since distant supervision has labeling bias.We employ two separate prediction networks HA-Net and DS-Net to predict the labels by human annotation and distant supervision, respectively, to prevent the degradation of accuracy by the incorrect labeling of distant supervision.Furthermore, we propose an additional loss term called disagreement penalty to enable HA-Net to learn from distantly supervised labels.In addition, we exploit additional networks to adaptively assess the labeling bias by considering contextual information.Our performance study on sentence-level and document-level REs confirms the effectiveness of the dual supervision framework. Woohwan Jung, Kyuseok Shim |
COLING | 1 |
| 2020 | TIDY: Publishing a Time Interval Dataset with Differential Privacy (Extended abstract)abstractLog data from mobile devices usually contain a series of events with time intervals. However, the problem of releasing differentially private time interval data has not been tackled yet. We propose the TIDY (publishing Time Intervals via Differential privacY) algorithm to release time interval data under differential privacy. We use the frequency vectors as a compact representation of the time interval data to reduce the aggregated noise. We also develop a new partitioning method adapted for the frequency vectors to balance the trade-off between the noise and structural errors. Our experiments confirm that TIDY outperforms the existing algorithms for releasing 2D histograms. Woohwan Jung, Suyong Kwon, Kyuseok Shim |
ICDE | 1 |
| 2019 | Crowdsourced Truth Discovery in the Presence of Hierarchies for Knowledge Fusion
Woohwan Jung, Kyuseok Shim |
EDBT | 1 |
| 2017 | Integration of graphs from different data sources using crowdsourcing
Woohwan Jung, Kyuseok Shim |
Inf. Sci. | 2 |