VLDB 2026 Research / reviewers in the wild / expert
Jun-Hyung Park
dblp:16/716
· DBLP profile ↗
16ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0002-7900-3743ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Incorporating Domain Knowledge into Materials TokenizationabstractWhile language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language processing.However, these methods frequently produce excessive fragmentation and semantic loss, failing to maintain the structural and semantic integrity of material concepts.To address this issue, we propose MATTER, a novel tokenization approach that integrates material knowledge into tokenization.Based on MatDetector trained on our materials knowledge base and a re-ranking method prioritizing material concepts in token merging, MATTER maintains the structural integrity of identified material concepts and prevents fragmentation during tokenization, ensuring their semantic meaning remains intact.The experimental results demonstrate that MATTER outperforms existing tokenization methods, achieving an average performance gain of 4% and 2% in the generation and classification tasks, respectively.These results underscore the importance of domain knowledge for tokenization strategies in scientific text processing. 1 Yerim Oh, Jun-Hyung Park, Sung-Ho Kim 0010, SangKeun Lee 0001 |
ACL (1) | 2 |
| 2025 | Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
Nayeon Kim 0002, Eojin Jeon, Jun-Hyung Park, SangKeun Lee 0001 |
PAKDD (3) | 3 |
| 2025 | Continual debiasing: A bias mitigation framework for natural language understanding systems
Jun-Hyung Park, SangKeun Lee 0001 |
Expert Syst. Appl. | 3 |
| 2024 | MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionabstractChemical representation learning has gained increasing interest due to the limited availability of supervised data in fields such as drug and materials design.This interest particularly extends to chemical language representation learning, which involves pre-training Transformers on SMILES sequences -textual descriptors of molecules.Despite its success in molecular property prediction, current practices often lead to overfitting and limited scalability due to early convergence.In this paper, we introduce a novel chemical language representation learning framework, called MolTRES, to address these issues.MolTRES incorporates generator-discriminator training, allowing the model to learn from more challenging examples that require structural understanding.In addition, we enrich molecular representations by transferring knowledge from scientific literature by integrating external materials embedding.Experimental results show that our models outperform existing state-of-the-art models on popular molecular property prediction tasks. github.com/irishev/MolTRES Jun-Hyung Park, Yeachan Kim, Hyuntae Park, SangKeun Lee 0001 |
EMNLP | 1 |
| 2023 | Dynamic Structure Pruning for Compressing CNNsabstractStructure pruning is an effective method to compress and accelerate neural networks. While filter and channel pruning are preferable to other structure pruning methods in terms of realistic acceleration and hardware compatibility, pruning methods with a finer granularity, such as intra-channel pruning, are expected to be capable of yielding more compact and computationally efficient networks. Typical intra-channel pruning methods utilize a static and hand-crafted pruning granularity due to a large search space, which leaves room for improvement in their pruning performance. In this work, we introduce a novel structure pruning method, termed as dynamic structure pruning, to identify optimal pruning granularities for intra-channel pruning. In contrast to existing intra-channel pruning methods, the proposed method automatically optimizes dynamic pruning granularities in each layer while training deep neural networks. To achieve this, we propose a differentiable group learning method designed to efficiently learn a pruning granularity based on gradient-based learning of filter groups. The experimental results show that dynamic structure pruning achieves state-of-the-art pruning performance and better realistic acceleration on a GPU compared with channel pruning. In particular, it reduces the FLOPs of ResNet50 by 71.85% without accuracy degradation on the ImageNet dataset. Our code is available at https://github.com/irishev/DSP. Jun-Hyung Park, Yeachan Kim, Joon-Young Choi, SangKeun Lee 0001 |
AAAI | 1 |
| 2023 | SMoP: Towards Efficient and Effective Prompt Tuning with Sparse Mixture-of-PromptsabstractPrompt tuning has emerged as a successful parameter-efficient alternative to the full finetuning of language models.However, prior works on prompt tuning often utilize long soft prompts of up to 100 tokens to improve performance, overlooking the inefficiency associated with extended inputs.In this paper, we propose a novel prompt tuning method SMoP (Sparse Mixture-of-Prompts) that utilizes short soft prompts for efficient training and inference while maintaining performance gains typically induced from longer soft prompts.To achieve this, SMoP employs a gating mechanism to train multiple short soft prompts specialized in handling different subsets of the data, providing an alternative to relying on a single long soft prompt to cover the entire data.Experimental results demonstrate that SMoP outperforms baseline methods while reducing training and inference costs.We release our code at https://github.com/jyjohnchoi/SMoP. Joon-Young Choi, Jun-Hyung Park, Wing-Lam Mok, SangKeun Lee 0001 |
EMNLP | 3 |
| 2023 | Leap-of-Thought: Accelerating Transformers via Dynamic Token RoutingabstractComputational inefficiency in transformers has been a long-standing challenge, hindering the deployment in resource-constrained or realtime applications.One promising approach to mitigate this limitation is to progressively remove less significant tokens, given that the sequence length strongly contributes to the inefficiency.However, this approach entails a potential risk of losing crucial information due to the irrevocable nature of token removal.In this paper, we introduce Leap-of-Thought (LoT), a novel token reduction approach that dynamically routes tokens within layers.Unlike previous work that irrevocably discards tokens, LoT enables tokens to 'leap' across layers.This ensures that all tokens remain accessible in subsequent layers while reducing the number of tokens processed within layers.We achieve this by pairing the transformer with dynamic token routers, which learn to selectively process tokens essential for the task.Evaluation results clearly show that LoT achieves substantial improvement on computational efficiency.Specifically, LoT attains up to 25× faster inference time without a significant loss in accuracy 1 . Yeachan Kim, Jun-Hyung Park, SangKeun Lee 0001 |
EMNLP | 3 |
| 2023 | DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationabstractTowards human-level visual understanding, visual commonsense generation has been introduced to generate commonsense inferences beyond images.However, current research on visual commonsense generation has overlooked an important human cognitive ability: generating descriptive and diverse inferences.In this work, we propose a novel visual commonsense generation framework, called DIVE, which aims to improve the descriptiveness and diversity of generated inferences.DIVE involves two methods, generic inference filtering and contrastive retrieval learning, which address the limitations of existing visual commonsense resources and training objectives.Experimental results verify that DIVE outperforms state-ofthe-art models for visual commonsense generation in terms of both descriptiveness and diversity, while showing a superior quality in generating unique and novel inferences.Notably, DIVE achieves human-level descriptiveness and diversity on Visual Commonsense Graphs.Furthermore, human evaluations confirm that DIVE aligns closely with human judgments on descriptiveness and diversity 1 . Jun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon, SangKeun Lee 0001 |
EMNLP | 1 |
| 2022 | Break it Down into BTS: Basic, Tiniest Subword Units for KoreanabstractWe introduce Basic, Tiniest Subword (BTS) units for the Korean language, which are inspired by the invention principle of Hangeul, the Korean writing system.Instead of relying on 51 Korean consonant and vowel letters, we form the letters from BTS units by adding strokes or combining them.To examine the impact of BTS units on Korean language processing, we develop a novel BTSbased word embedding framework that is readily applicable to various models.Our experiments reveal that BTS units significantly improve the performance of Korean word embedding on all intrinsic and extrinsic tasks in our evaluation.In particular, BTS-based word embedding outperforms the state-of-theart Korean word embedding by 11.8% in word analogy.We further investigate the unique advantages provided by BTS units through indepth analysis.Our code is available at https: //github.com/irishev/BTS. Nayeon Kim 0002, Jun-Hyung Park, Joon-Young Choi, Eojin Jeon, Youjin Kang, SangKeun Lee 0001 |
EMNLP | 2 |
| 2022 | Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkabstractPre-trained language models have achieved remarkable successes in natural language processing tasks, coming at the cost of increasing model size.To address this issue, knowledge distillation (KD) has been widely applied to compress language models.However, typical KD approaches for language models have overlooked the difficulty of training examples, suffering from incorrect teacher prediction transfer and sub-efficient training.In this paper, we propose a novel KD framework, Tutor-KD, which improves the distillation effectiveness by controlling the difficulty of training examples during pre-training.We introduce a tutor network that generates samples that are easy for the teacher but difficult for the student, with training on a carefully designed policy gradient method.Experimental results show that Tutor-KD significantly and consistently outperforms the state-of-the-art KD methods with variously sized student models on the GLUE benchmark, demonstrating that the tutor can effectively generate training examples for the student 1 . Jun-Hyung Park, Wing-Lam Mok, Joon-Young Choi, SangKeun Lee 0001 |
EMNLP | 2 |
| 2022 | Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingabstractMasked language modeling (MLM) has been widely used for pre-training effective bidirectional representations, but incurs substantial training costs.In this paper, we propose a novel concept-based curriculum masking (CCM) method to efficiently pre-train a language model.CCM has two key differences from existing curriculum learning approaches to effectively reflect the nature of MLM.First, we introduce a carefully-designed linguistic difficulty criterion that evaluates the MLM difficulty of each token.Second, we construct a curriculum that gradually masks words related to the previously masked words by retrieving a knowledge graph.Experimental results show that CCM significantly improves pre-training efficiency.Specifically, the model trained with CCM shows comparative performance with the original BERT on the General Language Understanding Evaluation benchmark at half of the training cost.Code is available at https://github.com/KoreaMGLEE/Concept- based-curriculum-masking. Jun-Hyung Park, Kang-Min Kim, SangKeun Lee 0001 |
EMNLP | 2 |
| 2022 | Examining the impact of adaptive convolution on natural language understanding
Jun-Hyung Park, Byung-Ju Choi, SangKeun Lee 0001 |
Expert Syst. Appl. | 1 |
| 2022 | Quantized Sparse Training: A Unified Trainable Framework for Joint Pruning and Quantization in DNNsabstractDeep neural networks typically have extensive parameters and computational operations. Pruning and quantization techniques have been widely used to reduce the complexity of deep models. Both techniques can be jointly used for realizing significantly higher compression ratios. However, separate optimization processes and difficulties in choosing the hyperparameters limit the application of both the techniques simultaneously. In this study, we propose a novel compression framework, termed as quantized sparse training, that prunes and quantizes networks jointly in a unified training process. We integrate pruning and quantization into a gradient-based optimization process based on the straight-through estimator. Quantized sparse training enables us to simultaneously train, prune, and quantize a network from scratch. The empirical results validate the superiority of the proposed methodology over the recent state-of-the-art baselines with respect to both the model size and accuracy. Specifically, quantized sparse training achieves a 135 KB model size in the case of VGG16, without any accuracy degradation, which is 40% of the model size feasible based on the state-of-the-art pruning and quantization approach. Jun-Hyung Park, Kang-Min Kim, SangKeun Lee 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2019 | meChat: In-device Conversational Photo Sharing ServiceabstractIn this demo, we demonstrate an in-device conversational photo sharing service, termed meChat, which helps users share in-device photos easily in messaging applications by searching conversation-related photos automatically. In particular, meChat understands the semantics of on-going conversation and in-device photos by projecting both of them into a single semantic space. Subsequently, it retrieves in-device photos related to the conversation context. Through this process, meChat makes it easy for the users to share the photos while communicating in messaging applications. In addition, it is worth noting that meChat works in a stand-alone, privacy-protecting manner without sending out any in-device photos and conversations to the external servers. This is different from existing photo services (e.g., Google Photos) which resort on the cloud server. Woo-Jong Ryu, Yoonjoo Ahn, Song-Eun Lee, Kang-Min Kim, Jun-Hyung Park, SangKeun Lee 0001 |
MobiSys | 6 |
| 2009 | Realization of a vibro-tactile glove type mouseabstractIn this paper, we suggested a glove type mouse using a gyroscope sensor, an acceleration sensor, and six pin-type vibro-tactile modules. It was designed as a USB HID(human interface device) so that it can be automatically recognized and installed as a general mouse when it is plugged in PC USB socket. This mouse recognizes a user's wrist movement by the gyroscope sensor and the acceleration sensor in the glove and transmits coordinate value to the PC through wireless Bluetooth. This vibro-tactile glove type mouse accommodates all circuits and devices in the glove, implementing a wearable system. A user can use general spatial mouse without any driver or application program. However, since tactile devices are not included in the USB HID, we made an application program for vibro-tactile display so that a PC program can transmit specific vibro-tactile information to the user to represent gray scale of pictures, braille codes, directions to go, and so on. Jun-Hyung Park, Ju Seok Jeong, Tae-Jeong Jang |
VRST | 1 |
| 2004 | Detection Techniques for ELF Executable File Using Assembly Instruction Searching
Jun-Hyung Park, Minsoo Kim 0002, BongNam Noh |
ICCSA (1) | 1 |