Jinyoung Park 0001

dblp:03/1524-1 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-7129-2141ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Difficulty-aware Balancing Margin Loss for Long-tailed Recognition
abstract
When trained with severely imbalanced data, deep neural networks often struggle to accurately recognize classes with few samples. Previous studies in long-tailed recognition have attempted to rebalance biased learning using known sample distributions, primarily addressing different classification difficulties at the class level. However, these approaches often overlook the instance difficulty variation within each class. In this paper, we propose a difficulty-aware balancing margin (DBM) loss, which considers both class imbalance and instance difficulty. DBM loss comprises two components: a class-wise margin to mitigate learning bias caused by imbalanced class frequencies, and an instance-wise margin assigned to hard positive samples based on their individual difficulty. DBM loss improves class discriminativity by assigning larger margins to more difficult samples. Our method effortlessly combine with existing approaches and consistently improves performance across various long-tailed recognition benchmarks.
Minseok Son, Inyong Koo, Jinyoung Park 0001, Changick Kim
AAAI3
2024 Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition
Muhammad Adi Nugroho, Sangmin Woo, Jinyoung Park 0001, Yooseung Wang, Changick Kim
ECCV (48)4
2024 VideoMamba: Spatio-Temporal Selective State Space Model
Jinyoung Park 0001, Hee-Seon Kim, Kangwook Ko, Minbeom Kim, Changick Kim
ECCV (25)1
2024 Anchoring Vision and Language Knowledge for Weakly Supervised Group Activity Recognition
abstract
The emergence of Foundation Vision-Language Models (VLMs) has ignited a surge of research in the computer vision field due to their robust baseline performance. Inspired by this, we propose the Anchoring Vision-Language Network (AnViL-Net), which integrates a vision language model for the challenging task of Weakly-Supervised Group Activity Recognition (WSGAR). Our network effectively incorporates VLMs into WSGAR, addressing the challenges posed by dynamic actor motions and domain-specific activity classes. AnViL-Net leverages highly generalized VLM vision features as anchors for extracting visual features. Additionally, semantically meaningful VLM language features serve as anchors for inferring the semantic relationships between actors and their activities. We demonstrate the effectiveness of AnViL-Net on multiple group activity datasets, achieving competitive state-of-the-art results.
Muhammad Adi Nugroho, Jinyoung Park 0001, Changick Kim
VCIP2
2024 Sketch-based Video Object Localization
abstract
We introduce Sketch-based Video Object Localization (SVOL), a new task aimed at localizing spatio-temporal object boxes in video queried by the input sketch. We first outline the challenges in the SVOL task and build the Sketch-Video Attention Network (SVANet) with the following design principles: (i) to consider temporal information of video and bridge the domain gap between sketch and video; (ii) to accurately identify and localize multiple objects simultaneously; (iii) to handle various styles of sketches; (iv) to be classification-free. In particular, SVANet is equipped with a Cross-modal Transformer that models the interaction between learnable object tokens, query sketch, and video through attention operations, and learns upon a per-frame set matching strategy that enables frame-wise prediction while utilizing global video context. We evaluate SVANet on a newly curated SVOL dataset. By design, SVANet successfully learns the mapping between the query sketches and video objects, achieving state-of-the-art results on the SVOL benchmark. We further confirm the effectiveness of SVANet via extensive ablation studies and visualizations. Lastly, we demonstrate its transfer capability on unseen datasets and novel categories, suggesting its high scalability in real-world applications. Codes are available at https://github.com/sangminwoo/SVOL.
Sangmin Woo, So-Yeong Jeon, Jinyoung Park 0001, Minji Son, Changick Kim
WACV3
2023 Multi-modal Social Group Activity Recognition in Panoramic Scene
abstract
Group Activity Recognition (GAR) is a challenging problem in computer vision due to the intricate dynamics and interactions among individuals. The existing methods utilize RGB videos face challenges in panoramic environments with numerous individuals and social groups. In this paper, we propose Multimodal Group Activity Recognition network (MGAR-net), that leverages the combined power of RGB and LiDAR modalities. Our approach effectively utilizes information from both modalities thus robustly and accurately captures individual relationships and detects social groups in face of optical challenges. By harnessing the capability of LiDAR with our new fusion module, called Distance Aware Fusion Module (DAFM), MGAR-net acquires valuable 3D structure information. We conduct experiments on the JRDB-Act dataset, which contains challenging scenarios with numerous people. The results demonstrate that LiDAR data provide valuable information for social grouping and recognizing individual action and group activities, particularly in crowded group settings. For social grouping, our MGAR-net improve performance by about 12% compared to the existing state-of-the-art models in terms of the AP metric.
Sangmin Woo, Jinyoung Park 0001, Muhammad Adi Nugroho, Changick Kim
VCIP4
2022 DAT: Domain Adaptive Transformer for Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) for semantic segmentation aims to predict class annotations on an unlabeled target dataset by training on a rich labeled source dataset. It is crucial in UDA semantic segmentation to decrease the domain gap by learning domain invariant feature representations across both domains. In this paper, we propose a novel transformer-based network, called a domain adaptive transformer (DAT), using a self-training scheme. We introduce domain invariant attention (DIA), which enables the DAT to exploit high-level domain invariant features at the patch level. Moreover, an entropy-based selective pseudo-labeling algorithm provides the DAT with reliable pseudo-labels of target samples for domain adaptive self-training, which corrects the noisy pseudo-labels online. We show that our DAT greatly improves the domain adaptability and achieves state-of-the-art results on the SYNTHIA-to-Cityscapes benchmark.
Jinyoung Park 0001, Minseok Son, Changick Kim
ICIP1