VLDB 2026 Research / reviewers in the wild / expert
Seonghyeon Moon
dblp:218/5632
· DBLP profile ↗
15ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FCC: Fully Connected Correlation for One-Shot SegmentationabstractOne-shot segmentation (OSS) aims to segment the target object in a query image using only one set of support image and mask. Therefore, having strong prior information for the target object using the support set is essential to guide the initial training of OSS, which leads to the success of one-shot segmentation in challenging cases, such as when the target object shows considerable variation in appearance, texture, or scale across the support and query images. To enrich this prior knowledge, we introduce FCC (Fully Connected Correlation) which integrates pixel-level correlations between support and query features, capturing associations that reveal target-specific patterns and correspondences in both same-layers and cross-layers. FCC captures previously inaccessible target information, effectively addressing the limitations of support mask. Our approach consistently demonstrates state-of-the-art performance in the PASCAL, COCO, and domain shift tests, while also notably accelerating model convergence. We conducted an ablation study and cross-layer correlation analysis to validate FCC’s core methodology. These findings reveal the effectiveness of FCC in enhancing prior information and overall model performance for OSS1. Seonghyeon Moon, Haein Kong, Muhammad Haris Khan, Mubbasir Kapadia, Yuewei Lin |
WACV | 1 |
| 2025 | Judging From Support-Set: A New Way To Utilize Few-Shot Segmentation For Segmentation Refinement ProcessabstractSegmentation refinement enhances coarse masks generated by segmentation algorithms, aiming for detailed and accurate contours of target objects. Despite advancements in segmentation refinement research, no method exists to evaluate its success, which is critical for reliable applications. To address this gap, we propose Judging From Support-set (JFS), leveraging a few-shot segmentation (FSS) model in a novel evaluation pipeline. Traditional FSS aims to locate target objects in query images using support set information. In JFS, coarse and refined masks from segmentation refinement methods become support masks for the FSS model, with the existing support mask serving as the test set. This setup evaluates the quality of refined segmentation. We validate JFS using the SAM Enhanced Pseudo-Labels (SEPL) and SegGPT on the PASCAL dataset, demonstrating its potential to reliably judge segmentation refinement success and foster innovation in image processing. Seonghyeon Moon, Qingze Tony Liu, Haein Kong, Muhammad Haris Khan |
ICIP | 1 |
| 2025 | Construction regulatory document digitalization with layout knowledge-informed object detection and semantic text recognition
Seonghyeon Moon, Yuguang Fu |
Adv. Eng. Informatics | 2 |
| 2024 | Learning from Synthetic Human Group ActivitiesabstractThe study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation, we introduce M3 Act, a synthetic data generator for multi-view multi-group multi-person human atomic actions and group activities. Powered by Unity Engine, M3 Act features mul-tiple semantic groups, highly diverse and photorealistic images, and a comprehensive set of annotations, which facilitates the learning of human-centered tasks across single-person, multi-person, and multi-group conditions. We demonstrate the advantages of M3 Act across three core experiments. The results suggest our synthetic dataset can significantly improve the performance of several downstream methods and replace real-world datasets to reduce cost. Notably, M3 Act improves the state-of-the-art MOTRv2 on DanceTrack dataset, leading to a hop on the leaderboard from 10thto 2ndplace. Moreover, M3 Act opens new research for controllable 3D group activity generation. We define multiple metrics and propose a competitive baseline for the novel task. Our code and data are available at our project page: http://cjerry1243.github.io/M3Act. Che-Jui Chang, Danrui Li, Deep Patel, Parth Goel, Honglu Zhou, Seonghyeon Moon, Samuel S. Sohn, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
CVPR | 6 |
| 2023 | MSI: Maximize Support-Set Information for Few-Shot SegmentationabstractFSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract information relevant to the target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature excision through a limiting support mask introduces an information bottleneck in several challenging FSS cases, e.g., for small targets and/or inaccurate target boundaries. To this end, we present a novel method (MSI), which maximizes the support-set information by exploiting two complementary sources of features to generate super correlation maps. We validate the effectiveness of our approach by instantiating it into three recent and strong FSS methods. Experimental results on several publicly available FSS benchmarks show that our proposed method consistently improves performance by visible margins and leads to faster convergence. Our code and trained models are available at: https://github.com/moonsh/MSI-Maximize-Support-Set-Information Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia |
ICCV | 1 |
| 2023 | Development of a real-time noise estimation model for construction sitesabstractAs construction noise negatively affects the health and quality of life of stakeholders, field managers need to properly monitor and manage noise. Thus, the authors developed a model that estimates real-time noise levels at a construction site and the surroundings to enable preemptive responses to noise-related issues. To accurately estimate noise, necessary field data were collected using an unmanned aerial vehicle (UAV) and noise sensors. The noise estimation model was composed of two sub-models: the noise-customized spatial interpolation model and the noise propagation model. The noise-customized spatial interpolation model was developed to estimate the internal noise of the construction site using a few sensor noise levels. Meanwhile, the noise propagation model was developed to estimate the noise level outside the construction site using internal noise estimation results, obstacles, weather information, and noise sources information. The model was evaluated through field tests at a construction technology demonstration center, environments identical to real construction sites in South Korea. The model showed satisfactory performance, with an accuracy of 96.71% and a root mean square error (RMSE) of 2.62 for the internal construction site noise and an accuracy of 96.03% and an RMSE of 2.70 for outside the construction site. To facilitate the usage of the noise estimation results for field managers, the research team visualized the results using the Unity 3D Engine. The results will enable field managers to assess workers’ long-term noise exposure and respond to potential civil complaints, gearing up to realize environmental, social, and governance (ESG) goals in the construction industry. Gitaek Lee, Seonghyeon Moon, Jae-Hyun Hwang, Seokho Chi |
Adv. Eng. Informatics | 2 |
| 2022 | MUSE-VAE: Multi-Scale VAE for Environment-Aware Long Term Trajectory PredictionabstractAccurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging stochastic forecasting problem. In this work, we propose MUSEVAE, a new probabilistic modeling framework based on a cascade of Conditional VAEs, which tackles the long-term, uncertain trajectory prediction task using a coarse-to-fine multi-factor forecasting architecture. In its Macro stage, the model learns a joint pixel-space representation of two key factors, the underlying environment and the agent movements, to predict the long and short term motion goals. Conditioned on them, the Micro stage learns a fine-grained spatio-temporal representation for the prediction of individual agent trajectories. The VAE backbones across the two stages make it possible to naturally account for the joint uncertainty at both levels of granularity. As a result, MUSEVAE offers diverse and simultaneously more accurate predictions compared to the current state-of-the-art. We demonstrate these assertions through a comprehensive set of experiments on nuScenes and SDD benchmarks as well as PFSD, a new synthetic dataset, which challenges the forecasting ability of models on complex agent-environment interaction scenarios. Mihee Lee, Samuel S. Sohn, Seonghyeon Moon, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001 |
CVPR | 3 |
| 2022 | HM: Hybrid Masking for Few-Shot Segmentation
Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia |
ECCV (20) | 1 |
| 2022 | Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow PredictionabstractWith the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end, we propose Fourier Isovists, a novel input representation based on egocentric visibility, which consistently improves all existing models. We also propose GeoInteractNet (GINet), which couples the layers between a multi-scale attention network (M-SCAN) and a convolutional encoder-decoder network (CED). M-SCAN approximates a super-resolution map of where humans are likely to interact on the way to their goals and produces multi-scale attention maps. The CED then uses these maps in either its encoder's inputs or its decoder's attention gates, which allows GINet to produce super-resolution predictions with substantially higher accuracy than existing models even with Fourier Isovists. In order to evaluate the scalability of models to large and complex environments, which the only existing LTCFP dataset is unsuitable for, a new synthetic crowd dataset with both real and synthetic environments has been generated. In its nascent state, LTCFP has much to gain from our key contributions. The Supplementary Materials, dataset, and code are available at sssohn.github.io/GeoInteractNet. Samuel S. Sohn, Seonghyeon Moon, Honglu Zhou, Mihee Lee, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
IJCAI | 2 |
| 2022 | Automated system for construction specification review using natural language processingabstractExisting attempts to automate construction document analysis are limited in understanding the varied semantic properties of different documents. Due to the semantic conflicts, the construction specification review process is still conducted manually in practice despite the promising performance of the existing approaches. This research aimed to develop an automated system for reviewing construction specifications by analyzing the different semantic properties using natural language processing techniques. The proposed method analyzed varied semantic properties of 56 different specifications from five different countries in terms of vocabulary, sentence structure, and the organizing styles of provisions. First, the authors developed a semantic thesaurus for construction terms including 208 word-replacement rules based on Word2Vec embedding to understand the different vocabularies. Second, the authors developed a named entity recognition model based on bi-directional long short-term memory with a conditional random field layer, which identified the required keywords from given provisions with an averaged F1 score of 0.928. Third, the authors developed a provision-pairing model based on Doc2Vec embedding, which identified the most relevant provisions with an average accuracy of 84.4%. The web-based prototype demonstrated that the proposed system can facilitate the construction specification review process by reducing the time spent, supplementing the reviewer’s experience, enhancing accuracy, and achieving consistency. The results contribute to risk management in the construction industry, with practitioners being able to review construction specifications thoroughly in spite of tight schedules and few available experts. Seonghyeon Moon, Gitaek Lee, Seokho Chi |
Adv. Eng. Informatics | 1 |
| 2022 | Corrigendum to "Automated system for construction specification review using natural language processing" [Adv. Eng. Inf. 51 (2022) 101495]
Seonghyeon Moon, Gitaek Lee, Seokho Chi |
Adv. Eng. Informatics | 1 |
| 2022 | A2X: An end-to-end framework for assessing agent and environment interactions in multimodal human trajectory prediction
Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
Comput. Graph. | 3 |
| 2021 | A2X: An Agent and Environment Interaction Benchmark for Multimodal Human Trajectory PredictionabstractIn recent years, human trajectory prediction (HTP) has garnered attention in computer vision literature. Although this task has much in common with the longstanding task of crowd simulation, there is little from crowd simulation that has been borrowed, especially in terms of evaluation protocols. The key difference between the two tasks is that HTP is concerned with forecasting multiple steps at a time and capturing the multimodality of real human trajectories. A majority of HTP models are trained on the same few datasets, which feature small, transient interactions between real people and little to no interaction between people and the environment. Unsurprisingly, when tested on crowd egress scenarios, these models produce erroneous trajectories that accelerate too quickly and collide too frequently, but the metrics used in HTP literature cannot convey these particular issues. To address these challenges, we propose (1) the A2X dataset, which has simulated crowd egress and complex navigation scenarios that compensate for the lack of agent-to-environment interaction in existing real datasets, and (2) evaluation metrics that convey model performance with more reliability and nuance. A subset of these metrics are novel multiverse metrics, which are better-suited for multimodal models than existing metrics. The dataset is available at: https://mubbasir.github.io/HTP-benchmark/. Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
MIG | 3 |
| 2020 | Laying the Foundations of Deep Long-Term Crowd Flow Prediction
Samuel S. Sohn, Honglu Zhou, Seonghyeon Moon, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
ECCV (29) | 3 |
| 2020 | Deep Integration of Physical Humanoid Control and Crowd NavigationabstractMany multi-agent navigation approaches make use of simplified representations such as a disk. These simplifications allow for fast simulation of thousands of agents but limit the simulation accuracy and fidelity. In this paper, we propose a fully integrated physical character control and multi-agent navigation method. In place of sample complex online planning methods, we extend the use of recent deep reinforcement learning techniques. This extension improves on multi-agent navigation models and simulated humanoids by combining Multi-Agent and Hierarchical Reinforcement Learning. We train a single short term goal-conditioned low-level policy to provide directed walking behaviour. This task-agnostic controller can be shared by higher-level policies that perform longer-term planning. The proposed approach produces reciprocal collision avoidance, robust navigation, and emergent crowd behaviours. Furthermore, it offers several key affordances not previously possible in multi-agent navigation including tunable character morphology and physically accurate interactions with agents and the environment. Our results show that the proposed method outperforms prior methods across environments and tasks, as well as, performing well in terms of zero-shot generalization over different numbers of agents and computation time. M. Brandon Haworth, Glen Berseth, Seonghyeon Moon, Petros Faloutsos, Mubbasir Kapadia |
MIG | 3 |