EDBT 2026 Demo / reviewers in the wild / expert
Jinwoo Choi 0001
dblp:47/2621-1
· DBLP profile ↗
23ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0001-7043-0610ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ${\text{CA}^{2}\text{ST}}$: Cross-Attention in Audio, Space, and Time for Holistic Video RecognitionabstractWe propose Cross-Attention in Audio, Space, and Time (C$\text{A}^{2}$A2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a balanced spatio-temporal understanding of videos. To address this, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), using only RGB input. In each layer of CAST, Bottleneck Cross-Attention (B-CA) enables spatial and temporal experts to exchange information and make synergistic predictions. For holistic video understanding, we extend CAST by integrating an audio expert, forming Cross-Attention in Visual and Audio (CAVA). We validate the CAST on benchmarks with different characteristics, EPIC-KITCHENS-100, Something-Something-V2, Kinetics-400, ActivityNet, and HD-EPIC to show balanced performance. We also validate the CAVA on audio-visual action recognition benchmarks, including UCF-101, VGG-Sound, KineticsSound, EPIC- SOUNDS, and HD-EPIC-SOUNDS. CAVA shows favorable performance on these datasets, demonstrating the effective information exchange among multiple experts within the B-CA module. In addition, C$\text{A}^{2}$A2ST combines CAST and CAVA by employing spatial, temporal, and audio experts through cross-attention, achieving balanced and holistic video understanding. Jongseo Lee, Joohyun Chang, Dongho Lee, Jinwoo Choi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | HiCM²: Hierarchical Compact Memory Modeling for Dense Video CaptioningabstractWith the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical dense memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical dense memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets. Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 0001, Seong Tae Kim 0001 |
AAAI | 4 |
| 2025 | MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal RepresentationsabstractIn this work, we tackle action-scene hallucination in Video Large Language Models (Video-LLMs), where models incorrectly predict actions based on the scene context or scenes based on observed actions. We observe that existing Video-LLMs often suffer from action-scene hallucination due to two main factors. First, existing Video-LLMs intermingle spatial and temporal features by applying an attention operation across all tokens. Second, they use standard Rotary Position Embedding (RoPE), which causes the text tokens to overemphasize certain types of tokens depending on their sequential orders. To address these issues, we introduce MASH-VLM, Mitigating Action-Scene Hallucination in Video-LLMs through disentangled spatial-temporal representations. Our approach includes two key innovations: (1) DST-attention, a novel attention mechanism that disentangles spatial and temporal tokens within the LLM by using masked attention to restrict direct interactions be tween spatial and temporal tokens; (2) Harmonic-RoPE, which extends the dimensionality of the positional IDs, allowing spatial and temporal tokens to maintain balanced positions relative to the text tokens. To evaluate the action-scene hallucination in Video-LLMs, we introduce the UN-SCENE benchmark with 1,320 videos and 4,078 QA pairs. MASH-VLM achieves state-of-the-art performance on the UNSCENE benchmark, as well as on existing video understanding benchmarks. Kyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee, Jinwoo Choi 0001 |
CVPR | 6 |
| 2025 | Universal Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaptation for semantic segmentation (UDA-SS) aims to transfer knowledge from labeled source data to unlabeled target data. However, traditional UDA-SS methods assume that category settings between source and target domains are known, which is unrealistic in real-world scenarios. This leads to performance degradation if private private classes exist. To address this limitation, we propose Universal Domain Adaptation for Semantic Segmentation (UniDA-SS), achieving robust adaptation even without prior knowledge of category settings. We define the problem in the UniDA-SS scenario as low confidence scores of common classes in the target domain, which leads to confusion with private classes. To solve this problem, we propose UniMAP: UniDA-SS with Image Matching and Prototype-based Distinction, a novel framework composed of two key components. First, Domain-Specific Prototype-based Distinction (DSPD) divides each class into two domain-specific prototypes, enabling finer separation of domain-specific features and enhancing the identification of common classes across domains. Second, Target-based Image Matching (TIM) selects a source image containing the most common-class pixels based on the target pseudo-label and pairs it in a batch to promote effective learning of common classes. We also introduce a new UniDA-SS benchmark and demonstrate through various experiments that UniMAP significantly outperforms baselines. The code is available at https://github.com/KU-VGI/UniMAP. Seun-An Choe, Keon-Hee Park, Jinwoo Choi 0001, Gyeong-Moon Park |
CVPR | 3 |
| 2025 | ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
Jongseo Lee, Kyungho Bae, Kyle Min 0001, Gyeong-Moon Park, Jinwoo Choi 0001 |
ICCV | 5 |
| 2025 | Disentangled Concepts Speak Louder Than Words: Explainable Video Action RecognitionabstractEffective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature—intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets—KTH, Penn Action, HAA500, and UCF101—demonstrate that DANCE significantly improves explanation clarity with competitive performance. Through a user study, we validate the superior interpretability of DANCE. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis. Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim 0001, Jinwoo Choi 0001 |
NeurIPS | 5 |
| 2024 | Open-Set Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaptation (UDA) for semantic segmentation aims to transfer the pixel-wise knowledge from the labeled source domain to the unlabeled target do-main. However, current UDA methods typically assume a shared label space between source and target, limiting their applicability in real-world scenarios where novel cat-egories may emerge in the target domain. In this paper, we introduce Open-Set Domain Adaptation for Semantic Segmentation (OSDA -SS) for the first time, where the target domain includes unknown classes. We identify two major problems in the OSDA -SS scenario as follows: 1) the existing UDA methods struggle to predict the exact boundary of the unknown classes, and 2) they fail to accurately predict the shape of the unknown classes. To address these issues, we propose Boundary and Unknown Shape-Aware open-set domain adaptation, coined BUS. Our BUS can accu-rately discern the boundaries between known and unknown classes in a contrastive manner using a novel dilation-erosion-based contrastive loss. In addition, we propose OpenReMix, a new domain mixing augmentation method that guides our model to effectively learn domain and size-invariant features for improving the shape detection of the known and unknown classes. Through extensive experiments, we demonstrate that our proposed BUS effectively detects unknown classes in the challenging OSDA-SS sce-nario compared to the previous methods by a large margin. The code is available at https://github.com/KHUAGI/BUS. Seun-An Choe, Ah-Hyung Shin, Keon-Hee Park, Jinwoo Choi 0001, Gyeong-Moon Park |
CVPR | 4 |
| 2024 | Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalabstractThere has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a multitasking problem of event localization and event captioning to consider inter-task relations. However, addressing both tasks using only visual input is challenging due to the lack of semantic content. In this study, we address this by proposing a novel framework inspired by the cognitive information processing of humans. Our model utilizes external memory to incorporate prior knowledge. The memory retrieval method is proposed with cross-modal video-to-text matching. To effectively incorporate retrieved text features, the versatile encoder and the decoder with visual and textual cross-attention modules are designed. Comparative experiments have been conducted to show the effectiveness of the proposed method on ActivityNet Captions and YouCook2 datasets. Experimental results show promising performance of our model without extensive pretraining from a large video dataset. Our code is available at https://github.com/ailab-kyunghee/CM2_DVC. Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 0001, Seong Tae Kim 0001 |
CVPR | 4 |
| 2024 | DEVIAS: Learning Disentangled Video Representations of Action and Scene
Kyungho Bae, Geo Ahn, Youngrae Kim 0001, Jinwoo Choi 0001 |
ECCV (68) | 4 |
| 2024 | Multi-teacher Invariance Distillation for Domain-Generalized Action Recognition
Abhishek Maiti, Yuliang Zou, Jinwoo Choi 0001 |
ICPR (29) | 4 |
| 2024 | GLAD: Global-Local View Alignment and Background Debiasing for Unsupervised Video Domain Adaptation with Large Domain GapabstractIn this work, we tackle the challenging problem of unsupervised video domain adaptation (UVDA) for action recognition. We specifically focus on scenarios with a substantial domain gap, in contrast to existing works primarily deal with small domain gaps between labeled source domains and unlabeled target domains. To establish a more realistic setting, we introduce a novel UVDA scenario, denoted as Kinetics→BABEL, with a more considerable domain gap in terms of both temporal dynamics and background shifts. To tackle the temporal shift, i.e., action duration difference between the source and target domains, we propose a global-local view alignment approach. To mitigate the background shift, we propose to learn temporal order sensitive representations by temporal order learning and background invariant representations by background augmentation. We empirically validate that the proposed method shows significant improvement over the existing methods on the Kinetics→BABEL dataset with a large domain gap. The code is available at https://github.com/KHU-VLL/GLAD. Hyogun Lee, Kyungho Bae, Seong Jong Ha, Yumin Ko, Gyeong-Moon Park, Jinwoo Choi 0001 |
WACV | 6 |
| 2024 | Background debiased class incremental learning for video action recognition
Le Quan Nguyen, Jinwoo Choi 0001, Lien Minh Dang, Hyeonjoon Moon |
Image Vis. Comput. | 2 |
| 2024 | Video domain adaptation for semantic segmentation using perceptual consistency matching
Ihsan Ullah 0005, Sion An, Myeongkyun Kang, Philip Chikontwe, Hyunki Lee, Jinwoo Choi 0001, Sanghyun Park 0004 |
Neural Networks | 6 |
| 2023 | CAST: Cross-Attention in Space and Time for Video Action RecognitionabstractRecognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), that achieves a balanced spatio-temporal understanding of videos using only RGB input. Our proposed bottleneck cross-attention mechanism enables the spatial and temporal expert models to exchange information and make synergistic predictions, leading to improved performance. We validate the proposed method with extensive experiments on public benchmarks with different characteristics: EPIC-Kitchens-100, Something-Something-V2, and Kinetics-400. Our method consistently shows favorable performance across these datasets, while the performance of existing methods fluctuates depending on the dataset characteristics. The code is available at https://github.com/KHU-VLL/CAST. Dongho Lee, Jongseo Lee, Jinwoo Choi 0001 |
NeurIPS | 3 |
| 2023 | Learning representational invariances for data-efficient action recognition
Yuliang Zou, Jinwoo Choi 0001, Qitong Wang 0001, Jia-Bin Huang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Contrastive encoder pre-training-based clustered federated learning for heterogeneous data
Ye Lin Tun 0001, Minh N. H. Nguyen, Chu Myaet Thwal, Jinwoo Choi 0001, Choong Seon Hong |
Neural Networks | 4 |
| 2022 | Self-Supervised Cross-Video Temporal Learning for Unsupervised Video Domain AdaptationabstractWe address the task of unsupervised domain adaptation (UDA) for videos with self-supervised learning. While UDA for images is a widely studied problem, UDA for videos is relatively unexplored. In this paper, we propose a novel self-supervised loss for the task of video UDA. The method is motivated by inverted reasoning. Many works on video classification have shown success with representations based on events in videos, e.g., ‘reaching’, ‘picking’, and ‘drinking’ events for ‘drinking coffee’. We argue that if we have event-based representations, we should be able to predict the relative distances between clips in videos. Inverting that, we propose a self-supervised task to predict the difference of the distance between two clips from the source video and the distance between two clips from the target video. We hope that such a task would encourage learning event-based representations of the videos, which is known to be beneficial for classification. Since we predict the difference of clip distances between clips from source videos and target videos, we ‘tie’ the two domains and expect to achieve well-adapted representations. We combine this purely self-supervised loss and the source classification loss to learn the model parameters. We give extensive empirical results on challenging video UDA benchmarks, i.e., UCF-HMDB and EPIC-Kitchens. The presented qualitative and quantitative results support our motivations and method. Jinwoo Choi 0001, Jia-Bin Huang 0001 |
ICPR | 1 |
| 2020 | Shuffle and Attend: Video Domain Adaptation
Jinwoo Choi 0001, Samuel Schulter, Jia-Bin Huang 0001 |
ECCV (12) | 1 |
| 2020 | Unsupervised and Semi-Supervised Domain Adaptation for Action Recognition from DronesabstractWe address the problem of human action classification in drone videos. Due to the high cost of capturing and labeling large-scale drone videos with diverse actions, we present unsupervised and semi-supervised domain adaptation approaches that leverage both the existing fully annotated action recognition datasets and unannotated (or only a few annotated) videos from drones. To study the emerging problem of drone-based action recognition, we create a new dataset, NEC-DRONE, containing 5,250 videos to evaluate the task. We tackle both problem settings with 1) same and 2) different action label sets for the source (e.g., Kinectics dataset) and target domains (drone videos). We present a combination of video and instance-based adaptation methods, paired with either a classifier or an embedding-based framework to transfer the knowledge from source to target. Our results show that the proposed adaptation approach substantially improves the performance on these challenging and practical tasks. We further demonstrate the applicability of our method for learning cross-view action recognition on the Charades-Ego dataset. We provide qualitative analysis to understand the behaviors of our approaches. Jinwoo Choi 0001, Manmohan Krishna Chandraker, Jia-Bin Huang 0001 |
WACV | 1 |
| 2019 | Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action RecognitionabstractHuman activities often occur in specific scene contexts, e.g., playing basketball on a basketball court. Training a model using existing video datasets thus inevitably captures and leverages such bias (instead of using the actual discriminative cues). The learned representation may not generalize well to new action classes or different tasks. In this paper, we propose to mitigate scene bias for video representation learning. Specifically, we augment the standard cross-entropy loss for action classification with 1) an adversarial loss for scene types and 2) a human mask confusion loss for videos where the human actors are masked out. These two losses encourage learning representations that are unable to predict the scene types and the correct actions when there is no evidence. We validate the effectiveness of our method by transferring our pre-trained model to three different tasks, including action classification, temporal localization, and spatio-temporal action detection. Our results show consistent improvement over the baseline model without debiasing. Jinwoo Choi 0001, Chen Gao 0003, Joseph Messou, Jia-Bin Huang 0001 |
NeurIPS | 1 |
| 2017 | HOMER: An Interactive System for Home Based Stroke RehabilitationabstractDelivering long term, unsupervised stroke rehabilitation in the home is a complex challenge that requires robust, low cost, scalable, and engaging solutions. We present HOMER, an interactive system that uses novel therapy artifacts, a computer vision approach, and a tablet interface to provide users with a flexible solution suitable for home based rehabilitation. HOMER builds on our prior work developing systems for lightly supervised rehabilitation use in the clinic, by identifying key features for functional movement analysis, adopting a simplified classification assessment approach, and supporting transferability of therapy outcomes to daily living experiences through the design of novel rehabilitation artifacts. A small pilot study with unimpaired subjects indicates the potential of the system in effectively assessing movement and establishing a creative environment for training. Aisling Kelliher, Jinwoo Choi 0001, Jia-Bin Huang 0001, Thanassis Rikakis, Kris Makoto Kitani |
ASSETS | 2 |
| 2009 | Multiple channel division for efficient distributed video codingabstractA novel concept of channel division to improve the performance of distributed video coding is proposed in this work. At the decoder, the proposed algorithm partitions each side information frame into multiple regions with different expected distortions. It is shown that the partitioning is equivalent to the division of the virtual noisy channel from the encoder to the decoder into multiple channels. Then, the proposed algorithm analyzes the noise characteristics of those multiple channels, and allocates a limited bit budget to those channels adaptively to improve the rate-distortion performance. Simulation results demonstrate that the proposed algorithm provides up to 5 dB better performance than the conventional DVC algorithms, as well as reduces the decoding complexity significantly. Young-Yoon Lee, Jinwoo Choi 0001, Chang-Su Kim 0001, Sang Uk Lee |
ICIP | 3 |
| 2009 | Flexible complexity control between encoder and decoder for video codingabstractWe propose a novel video codec that can distribute computational complexity between the encoder and the decoder in a flexible manner. Since motion estimation is the most computationally intensive part in video coding, it is important to share the motion estimation task. In the proposed algorithm, the decoder estimates motion vectors by performing a partial three step search. The estimated motion vectors are sent back to the encoder via a feedback channel. The encoder then refines the received motion vectors. In this work, four operation modes corresponding to the encoder to decoder complexity ratios are presented. Experimental results demonstrate that the proposed algorithm allocates the complexity between the encoder and the decoder effectively, while providing promising compression performance. Jinwoo Choi 0001, Chang-Su Kim 0001, Sang Uk Lee |
MMSP | 1 |