EDBT 2026 Demo / reviewers in the wild / expert
Jinyoung Moon
dblp:34/391
· DBLP profile ↗
24ranked-venue papers
3as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HiCM²: Hierarchical Compact Memory Modeling for Dense Video CaptioningabstractWith the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical dense memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical dense memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets. Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 0001, Seong Tae Kim 0001 |
AAAI | 3 |
| 2025 | Depth-aware Pedestrian Situation Classification for Enhanced Safety Monitoring in Smart CitiesabstractAs urban environments grow increasingly complex, accurate classification of pedestrian situations becomes vital for public safety in smart cities. However, few approaches address pedestrian situation assessment beyond general-purpose detection and tracking. This paper presents a novel pedestrian situation classification method that employs depth information to improve accuracy. Specifically, we devise a depth-aware ground region refinement technique that selectively focuses on ground regions near pedestrians during classification, effectively filtering out distant areas that could mislead the classification. Experimental results on a public dataset recorded in school zones demonstrate improvements in classification performance, particularly in precision. This depth-aware approach could improve pedestrian safety monitoring in smart cities by enabling more accurate situation assessment in complex visual environments. Dae Hoe Kim, Jinyoung Moon |
AVSS | 2 |
| 2025 | PredDETR: An End-to-End Transformer-Based Model for One-Stage Pedestrian Trajectory PredictionabstractThis work remedies the drawbacks of existing methods in pedestrian trajectory prediction for video surveillance. A traditional two-stage approach, which first detects and tracks pedestrians before forecasting their subsequent trajectories, relies solely on historical trajectories to predict future movements, overlooking the rich contextual information within videos. Although some one-stage methods have been introduced to streamline this process, they still struggle with complex model architectures due to their reliance on pixel-wise flow estimation. To resolve these issues, we propose a simple yet effective Transformer-based one-stage model, called PredDETR, that directly anticipates the future trajectories of multiple pedestrians from videos. Following the end-to-end paradigm of DETR, our PredDETR model not only associates the bounding boxes of each pedestrian across subsequent frames but also predicts their future trajectories using an additional predictive decoder. Taking past raw video frames as input, the proposed PredDETR directly forecasts the future locations of each pedestrian in a non-autoregressive manner. The empirical results validate that, despite its simpler design, PredDETR is a compelling method compared to a previous one-stage approach. Yongjin Kwon, Sungchan Oh, Jinyoung Moon, Yeonseung Chung |
AVSS | 3 |
| 2025 | Proactive Pedestrian Safety System by Fusing Trajectory Prediction With Semantic Ground Region ClassificationabstractThis paper presents a proactive pedestrian risk assessment system for traffic environments that integrates trajectory prediction with situation classification. Unlike end-to-end approaches that function as black boxes when predicting pedestrians' crossing intentions, our system employs trajectory forecasting combined with ground region classification of predicted paths. The proposed methodology first predicts future pedestrian trajectories using an attention-based recurrent neural network, then classifies the predicted situation using accumulated segmentation maps to assess potential pedestrian risk. Experimental evaluations demonstrate that our system outperforms black box approaches across multiple evaluation metrics. We also present validation results using real-world surveillance footage captured in urban environments, demonstrating the system’s real-time capability and practical applicability for integration into smart city infrastructure. Sungchan Oh, Dae Hoe Kim, Je-Seok Ham, Jinyoung Moon |
AVSS | 4 |
| 2025 | Weakly Supervised Video Scene Graph Generation via Natural Language SupervisionabstractExisting Video Scene Graph Generation (VidSGG) studies are trained in a fully supervised manner, which requires all frames in a video to be annotated, thereby incurring high annotation cost compared to Image Scene Graph Generation (ImgSGG). Although the annotation cost of VidSGG can be alleviated by adopting a weakly supervised approach commonly used for ImgSGG (WS-ImgSGG) that uses image captions, there are two key reasons that hinder such a naive adoption: 1) Temporality within video captions, i.e., unlike image captions, video captions include temporal markers (e.g., before, while, then, after) that indicate time-related details, and 2) Variability in action duration, i.e., unlike human actions in image captions, human actions in video captions unfold over varying duration. To address these issues, we propose a Natural Language-based Video Scene Graph Generation (NL-VSGG) framework that only utilizes the readily available video captions for training a VidSGG model. NL-VSGG consists of two key modules: Temporality-aware Caption Segmentation (TCS) module and Action Duration Variability-aware caption-frame alignment (ADV) module. Specifically, TCS segments the video captions into multiple sentences in a temporal order based on a Large Language Model (LLM), and ADV aligns each segmented sentence with appropriate frames considering the variability in action duration. Our approach leads to a significant enhancement in performance compared to simply applying the WS-ImgSGG pipeline to VidSGG on the Action Genome dataset. As a further benefit of utilizing the video captions as weak supervision, we show that the VidSGG model trained by NL-VSGG is able to predict a broader range of action classes that are not included in the training data, which makes our framework practical in reality. Kibum Kim 0001, Kanghoon Yoon, Yeonjun In, Jaehyeong Jeon, Jinyoung Moon, Chanyoung Park 0001 |
ICLR | 5 |
| 2025 | WatchoutPed: A dataset and model for Vulnerable Pedestrian Anticipation in surveillance videosabstractThis paper addresses the crucial challenge of Vulnerable Pedestrian Anticipation (VPA) in urban environments, utilizing surveillance video data to enhance pedestrian safety. VPA is crucial for identifying pedestrians in potentially dangerous situations, such as walking alongside roads or crossing crosswalks, where the risk of vehicular collisions is elevated. To advance research in this field, we introduce two primary components: the WatchoutPed dataset and the Vulnerable Pedestrian Anticipation Network (VPANet), a baseline network especially designed for VPA. The WatchoutPed dataset has been meticulously enriched with extensive annotations through an innovative auto-labeling technique that integrates ground region analysis with pedestrian state estimation, thus providing a solid foundation for VPA research. Complementing this, the VPANet is engineered to process visual and non-visual inputs extracted from past frames in surveillance footage, enabling it to predict the future state of pedestrians as either safe or unsafe. Tested on the WatchoutPed dataset, VPANet achieves an impressive 89% accuracy, outperforming current methods. Furthermore, we demonstrate the effectiveness of our auto-labeling approach. Notably, the accuracy of VPANet, when trained with the auto-generated annotations from the WatchoutPed, closely parallels that achieved with human-verified annotations, with a negligible variance of less than 1%. The broader implications of our work are significant for the development of smart urban safety infrastructures. Integrating these insights into intelligent crosswalk systems could greatly enhance the monitoring of pedestrian activity near crosswalks, enabling the timely alerting of drivers to the presence of vulnerable pedestrians, and thereby proactively preventing potential vehicular accidents. Je-Seok Ham, Dae Hoe Kim, Jinyoung Moon |
Knowl. Based Syst. | 3 |
| 2024 | PedRiskNet: Classifying Pedestrian Situations in Surveillance Videos for Pedestrian Safety Monitoring in Smart CitiesabstractIn the context of smart cities, pedestrian safety enhancement through visual AI-based technology is crucial. While pedestrian detection and tracking have been extensively studied, classifying pedestrian situations for immediate risk assessment remains challenging. Existing methods often fail to consider the broader context of unsafe situations of pedestrians or rely solely on pedestrian detection. To address this gap, we propose a novel pedestrian situation classification method incorporating ground region estimation and multi-modal fusion. By utilizing semantic segmentation and temporal consistency, we estimate ground region maps to mitigate occlusion effects. The proposed method fuses multi-modal features from the local surroundings and ground regions, allowing for accurate classification of pedestrian situations. Experimental results on a public dataset recorded in school zones demonstrate superior performance compared to baseline models. This approach holds significant potential for improving pedestrian safety in smart cities, enabling proactive measures and interventions to mitigate risks and enhance overall road safety. Dae Hoe Kim, Jinyoung Moon |
AVSS | 2 |
| 2024 | Trajectory Prediction using Attentive Visual FeaturesabstractThis research paper presents a new method for predicting future trajectories of objects in video. Previous studies have mainly focused on the spatial attributes of objects, such as their bounding boxes or coordinates, while often overlooking visual features of the objects and surroundings. Our approach overcomes this limitation by incorporating visual feature extraction networks with trajectory prediction networks, resulting in a significant improvement in predictive accuracy. We conducted extensive testing on trajectory datasets captured from both first-person and bird’s eye views, which validated our method and demonstrated a notable improvement in prediction accuracy. These results confirm the effectiveness of our integrated visual feature extraction in improving trajectory prediction models and emphasize the importance of considering dynamic relationships between objects and their surroundings for more precise predictions. Sungchan Oh, Jinyoung Moon |
AVSS | 2 |
| 2024 | Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalabstractThere has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a multitasking problem of event localization and event captioning to consider inter-task relations. However, addressing both tasks using only visual input is challenging due to the lack of semantic content. In this study, we address this by proposing a novel framework inspired by the cognitive information processing of humans. Our model utilizes external memory to incorporate prior knowledge. The memory retrieval method is proposed with cross-modal video-to-text matching. To effectively incorporate retrieved text features, the versatile encoder and the decoder with visual and textual cross-attention modules are designed. Comparative experiments have been conducted to show the effectiveness of the proposed method on ActivityNet Captions and YouCook2 datasets. Experimental results show promising performance of our model without extensive pretraining from a large video dataset. Our code is available at https://github.com/ailab-kyunghee/CM2_DVC. Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 0001, Seong Tae Kim 0001 |
CVPR | 3 |
| 2024 | LLM4SGG: Large Language Models for Weakly Supervised Scene Graph GenerationabstractWeakly-Supervised Scene Graph Generation (WSSGG) research has recently emerged as an alternative to the fully-supervised approach that heavily relies on costly annotations. In this regard, studies on WSSGG have utilized image captions to obtain unlocalized triplets while primarily focusing on grounding the unlocalized triplets over image regions. However, they have overlooked the two issues involved in the triplet formation process from the captions: 1) Semantic over-simplification issue arises when extracting triplets from captions, where fine- grained predicates in captions are undesirably converted into coarse-grained predi-cates, resulting in a long-tailed predicate distribution, and 2) Low-density scene graph issue arises when aligning the triplets in the caption with entity/predicate classes of interest, where many triplets are discarded and not used in training, leading to insufficient supervision. To tackle the two issues, we propose a new approach, i.e., Large Language Modelfor weakly-supervised SGG (LLM4SGG), where we mitigate the two issues by leveraging the LLM' sin-depth understanding of language and reasoning ability during the extraction of triplets from captions and alignment of entity/predicate classes with target data. To further engage the LLM in these processes, we adopt the idea of Chain-of-Thought and the in-context few-shot learning strategy. To validate the effectiveness of LLM4SGG, we conduct extensive experiments on Visual Genome and GQA datasets, showing significant improvements in both Recall@K and mean Recall@K compared to the state-of-the-art WSSGG methods. A further appeal is that LLM4SGG is data-efficient, enabling effective model training with a small amount of training images. Our code is available on https://github.com/rlqjall07/torch-LLM4SGG Kibum Kim 0001, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In, Jinyoung Moon, Donghyun Kim 0006, Chanyoung Park 0001 |
CVPR | 5 |
| 2024 | Adaptive Self-training Framework for Fine-grained Scene Graph GenerationabstractScene graph generation (SGG) models have suffered from inherent problems regarding the benchmark datasets such as the long-tailed predicate distribution and missing annotation problems. In this work, we aim to alleviate the long-tailed problem of SGG by utilizing unannotated triplets. To this end, we introduce a **S**elf-**T**raining framework for **SGG** **(ST-SGG)** that assigns pseudo-labels for unannotated triplets based on which the SGG models are trained. While there has been significant progress in self-training for image recognition, designing a self-training framework for the SGG task is more challenging due to its inherent nature such as the semantic ambiguity and the long-tailed distribution of predicate classes. Hence, we propose a novel pseudo-labeling technique for SGG, called **C**lass-specific **A**daptive **T**hresholding with **M**omentum **(CATM)**, which is a model-agnostic framework that can be applied to any existing SGG models. Furthermore, we devise a graph structure learner (GSL) that is beneficial when adopting our proposed self-training framework to the state-of-the-art message-passing neural network (MPNN)-based SGG models. Our extensive experiments verify the effectiveness of ST-SGG on various SGG models, particularly in enhancing the performance on fine-grained predicate classes. Kibum Kim 0001, Kanghoon Yoon, Yeonjun In, Jinyoung Moon, Donghyun Kim 0006, Chanyoung Park 0001 |
ICLR | 4 |
| 2023 | Unbiased Heterogeneous Scene Graph Generation with Relation-Aware Message Passing Neural NetworkabstractRecent scene graph generation (SGG) frameworks have focused on learning complex relationships among multiple objects in an image. Thanks to the nature of the message passing neural network (MPNN) that models high-order interactions between objects and their neighboring objects, they are dominant representation learning modules for SGG. However, existing MPNN-based frameworks assume the scene graph as a homogeneous graph, which restricts the context-awareness of visual relations between objects. That is, they overlook the fact that the relations tend to be highly dependent on the objects with which the relations are associated. In this paper, we propose an unbiased heterogeneous scene graph generation (HetSGG) framework that captures relation-aware context using message passing neural networks. We devise a novel message passing layer, called relation-aware message passing neural network (RMP), that aggregates the contextual information of an image considering the predicate type between objects. Our extensive evaluations demonstrate that HetSGG outperforms state-of-the-art methods, especially outperforming on tail predicate classes. The source code for HetSGG is available at https://github.com/KanghoonYoon/hetsgg-torch Kanghoon Yoon, Kibum Kim 0001, Jinyoung Moon, Chanyoung Park 0001 |
AAAI | 3 |
| 2023 | Exploiting recollection effects for memory-based video object segmentation
Enki Cho, Minkuk Kim, Hyungil Kim, Jinyoung Moon, Seong Tae Kim 0001 |
Image Vis. Comput. | 4 |
| 2023 | Learning to Discriminate Information for Online Action Detection: Analysis and ApplicationabstractOnline action detection, which aims to identify an ongoing action from a streaming video, is an important subject in real-world applications. For this task, previous methods use recurrent neural networks for modeling temporal relations in an input sequence. However, these methods overlook the fact that the input image sequence includes not only the action of interest but background and irrelevant actions. This would induce recurrent units to accumulate unnecessary information for encoding features on the action of interest. To overcome this problem, we propose a novel recurrent unit, named Information Discrimination Unit (IDU), which explicitly discriminates the information relevancy between an ongoing action and others to decide whether to accumulate the input information. This enables learning more discriminative representations for identifying an ongoing action. In this paper, we further present a new recurrent unit, called Information Integration Unit (IIU), for action anticipation. Our IIU exploits the outputs from IDN as pseudo action labels as well as RGB frames to learn enriched features of observed actions effectively. In experiments on TVSeries and THUMOS-14, the proposed methods outperform state-of-the-art methods by a significant margin in online action detection and action anticipation. Moreover, we demonstrate the effectiveness of the proposed units by conducting comprehensive ablation studies. Hyunjun Eun, Jinyoung Moon, Seokeon Choi, Yoonhyung Kim, Chanho Jung, Changick Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning to combine the modalities of language and video for temporal moment localizationabstractTemporal moment localization aims to retrieve the best video segment matching a moment specified by a query. The existing methods generate the visual and semantic embeddings independently and fuse them without full consideration of the long-term temporal relationship between them. To address these shortcomings, we introduce a novel recurrent unit, cross-modal long short-term memory (CM-LSTM), by mimicking the human cognitive process of localizing temporal moments that focuses on the part of a video segment related to the part of a query, and accumulates the contextual information across the entire video recurrently. In addition, we devise a two-stream attention mechanism for both attended and unattended video features by the input query to prevent necessary visual information from being neglected. To obtain more precise boundaries, we propose a two-stream attentive cross-modal interaction network (TACI) that generates two 2D proposal maps obtained globally from the integrated contextual features, which are generated by using CM-LSTM, and locally from boundary score sequences and then combines them into a final 2D map in an end-to-end manner. On the TML benchmark dataset, ActivityNet-Captions, the TACI outperforms state-of-the-art TML methods with [email protected] of 45.50% and 27.23% for [email protected] and [email protected], respectively. In addition, we show that the revised state-of-the-arts methods by replacing original LSTM with our CM-LSTM achieves performance gains. Jungkyoo Shin, Jinyoung Moon |
Comput. Vis. Image Underst. | 2 |
| 2021 | Reproducibility Companion Paper: Knowledge Enhanced Neural Fashion Trend ForecastingabstractThis companion paper supports the replication of the fashion trend forecasting experiments with the KERN (Knowledge Enhanced Recurrent Network) method that we presented in the ICMR 2020. We provide an artifact that allows the replication of the experiments using a Python implementation. The artifact is easy to deploy with simple installation, training and evaluation. We reproduce the experiments conducted in the original paper and obtain similar performance as previously reported. The replication results of the experiments support the main claims in the original paper. Yunshan Ma 0002, Yujuan Ding, Xun Yang 0001, Lizi Liao, Wai Keung Wong, Tat-Seng Chua, Jinyoung Moon, Hong-Han Shuai |
ICMR | 7 |
| 2021 | Temporal filtering networks for online action detection
Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, Changick Kim |
Pattern Recognit. | 2 |
| 2020 | Learning to Discriminate Information for Online Action DetectionabstractFrom a streaming video, online action detection aims to identify actions in the present. For this task, previous methods use recurrent networks to model the temporal sequence of current action frames. However, these methods overlook the fact that an input image sequence includes background and irrelevant actions as well as the action of interest. For online action detection, in this paper, we propose a novel recurrent unit to explicitly discriminate the information relevant to an ongoing action from others. Our unit, named Information Discrimination Unit (IDU), decides whether to accumulate input information based on its relevance to the current action. This enables our recurrent network with IDU to learn a more discriminative representation for identifying ongoing actions. In experiments on two benchmark datasets, TVSeries and THUMOS-14, the proposed method outperforms state-of-the-art methods by a significant margin. Moreover, we demonstrate the effectiveness of our recurrent unit by conducting comprehensive ablation studies. Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, Changick Kim |
CVPR | 2 |
| 2020 | SRG: Snippet Relatedness-Based Temporal Action Proposal GeneratorabstractRecent temporal action proposal generation approaches have suggested integrating segment- and snippet score-based methodologies to produce proposals with high recall and accurate boundaries. In this paper, different from such a hybrid strategy, we focus on the potential of the snippet score-based approach. Specifically, we propose a new snippet score-based method, named Snippet Relatedness-based Generator (SRG), with a novel concept of “snippet relatedness”. Snippet relatedness represents which snippets are related to a specific action instance. To effectively learn this snippet relatedness, we present “pyramid non-local operations” for locally and globally capturing long-range dependencies among snippets. By employing these components, SRG first produces a 2D relatedness score map that enables the generation of various temporal intervals reliably covering most action instances with high overlap. Then, SRG evaluates the action confidence scores of these temporal intervals and refines their boundaries to obtain temporal action proposals. On THUMOS-14 and ActivityNet-1.3 datasets, SRG outperforms state-of-the-art methods for temporal action proposal generation. Furthermore, compared to competing proposal generators, SRG leads to significant improvements in temporal action detection. Hyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung, Changick Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Detecting user attention to video segments using interval EEG features
Jinyoung Moon, Yongjin Kwon, Jongyoul Park, Wan Chul Yoon |
Expert Syst. Appl. | 1 |
| 2017 | Hierarchically linked infinite hidden Markov model based trajectory analysis and semantic region retrieval in a trajectory dataset
Yongjin Kwon, Kyuchang Kang, Junho Jin, Jinyoung Moon, Jongyoul Park |
Expert Syst. Appl. | 4 |
| 2015 | Challenging Issues in Visual Information Understanding Researches
Kyuchang Kang, Yongjin Kwon, Jinyoung Moon, Changseok Bae |
MMM (2) | 3 |
| 2015 | Recognition of Meaningful Human Actions for Video Annotation Using EEG Based User Responses
Jinyoung Moon, Yongjin Kwon, Kyuchang Kang, Changseok Bae, Wan Chul Yoon |
MMM (2) | 1 |
| 2003 | ebXML BP Modeling Toolkit abstractCollaboration in a business system requires a business process specification defining the procedure of the business scenario. The business process specification is generated from a business process model. ebXML, which is the XML-based B2B standard framework for organizations of any size using the Internet, recommends process analysts and modelers to use the UN/CEFACT (United Nations Centre for the Facilitation of Procedures and Practices for Administration, Commerce and Transport) Modeling Methodology (UMM). The artifacts of the modeling are UML (Unified Modeling Language) diagrams and worksheets. They can be transformed into an ebXML business process (BP) specification and other business models. The artifacts and transformed results are registered in the business library for being shared with ebXML systems and being re-used in other modeling tools. This paper reports on our implementation of an ebXML BP modeling toolkit that accepts the architecture suggested in the ebXML and considers the required functions of modeling. These functions include modeling business processes based on UMM, generating a BP specification, transforming the processes using a metaframework, and registering them in the ebXML registry. The business process modeling toolkit is made up of the business process modeler, business process editor, and built-in registry client. The business process modeler not only models the business process with UML diagrams but also generates the business process specification, reverses it, and exports XMI of the business process. The business process editor is used only for editing the business process specification. The built-in registry client stores the business process model, the business process specification, or the XMI document, and searches and loads them. Jinyoung Moon, Daeha Lee, Chankyu Park, Hyunkyu Cho |
EDOC | 1 |