VLDB 2026 Research / reviewers in the wild / expert
Muhammad Adi Nugroho
dblp:214/8126
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-9360-5441ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modality mixer exploiting complementary information for multi-modal action recognition
Sangmin Woo, Muhammad Adi Nugroho, Changick Kim |
Comput. Vis. Image Underst. | 3 |
| 2024 | Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition
Muhammad Adi Nugroho, Sangmin Woo, Jinyoung Park 0001, Yooseung Wang, Changick Kim |
ECCV (48) | 1 |
| 2024 | Anchoring Vision and Language Knowledge for Weakly Supervised Group Activity RecognitionabstractThe emergence of Foundation Vision-Language Models (VLMs) has ignited a surge of research in the computer vision field due to their robust baseline performance. Inspired by this, we propose the Anchoring Vision-Language Network (AnViL-Net), which integrates a vision language model for the challenging task of Weakly-Supervised Group Activity Recognition (WSGAR). Our network effectively incorporates VLMs into WSGAR, addressing the challenges posed by dynamic actor motions and domain-specific activity classes. AnViL-Net leverages highly generalized VLM vision features as anchors for extracting visual features. Additionally, semantically meaningful VLM language features serve as anchors for inferring the semantic relationships between actors and their activities. We demonstrate the effectiveness of AnViL-Net on multiple group activity datasets, achieving competitive state-of-the-art results. Muhammad Adi Nugroho, Jinyoung Park 0001, Changick Kim |
VCIP | 1 |
| 2023 | Towards Good Practices for Missing Modality Robust Action RecognitionabstractStandard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances degrade drastically if any modality is missing in the inference stage. We ask: how can we train a model that is robust to missing modalities? This paper seeks a set of good practices for multi-modal action recognition, with a particular interest in circumstances where some modalities are not available at an inference time. First, we show how to effectively regularize the model during training (e.g., data augmentation). Second, we investigate on fusion methods for robustness to missing modalities: we find that transformer-based fusion shows better robustness for missing modality than summation or concatenation. Third, we propose a simple modular network, ActionMAE, which learns missing modality predictive coding by randomly dropping modality features and tries to reconstruct them with the remaining modality features. Coupling these good practices, we build a model that is not only effective in multi-modal action recognition but also robust to modality missing. Our model achieves the state-of-the-arts on multiple benchmarks and maintains competitive performances even in missing modality scenarios. Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim |
AAAI | 4 |
| 2023 | Audio-Visual Glance Network for Efficient Video RecognitionabstractDeep learning has made significant strides in video understanding tasks, but the computation required to classify lengthy and massive videos using clip-level video classifiers remains impractical and prohibitively expensive. To address this issue, we propose Audio-Visual Glance Network (AVGN), which leverages the commonly available audio and visual modalities to efficiently process the spatio-temporally important parts of a video. AVGN firstly divides the video into snippets of image-audio clip pair and employs lightweight unimodal encoders to extract global visual features and audio features. To identify the important temporal segments, we use an Audio-Visual Temporal Saliency Transformer (AV-TeST) that estimates the saliency scores of each frame. To further increase efficiency in the spatial dimension, AVGN processes only the important patches instead of the whole images. We use an Audio-Enhanced Spatial Patch Attention (AESPA) module to produce a set of enhanced coarse visual features, which are fed to a policy network that produces the coordinates of the important patches. This approach enables us to focus only on the most important spatio-temporally parts of the video, leading to more efficient video recognition. Moreover, we incorporate various training techniques and multi-modal feature fusion to enhance the robustness and effectiveness of our AVGN. By combining these strategies, our AVGN sets new state-of-the-art performance in multiple video recognition benchmarks while achieving faster processing speed. Muhammad Adi Nugroho, Sangmin Woo, Changick Kim |
ICCV | 1 |
| 2023 | Multi-modal Social Group Activity Recognition in Panoramic SceneabstractGroup Activity Recognition (GAR) is a challenging problem in computer vision due to the intricate dynamics and interactions among individuals. The existing methods utilize RGB videos face challenges in panoramic environments with numerous individuals and social groups. In this paper, we propose Multimodal Group Activity Recognition network (MGAR-net), that leverages the combined power of RGB and LiDAR modalities. Our approach effectively utilizes information from both modalities thus robustly and accurately captures individual relationships and detects social groups in face of optical challenges. By harnessing the capability of LiDAR with our new fusion module, called Distance Aware Fusion Module (DAFM), MGAR-net acquires valuable 3D structure information. We conduct experiments on the JRDB-Act dataset, which contains challenging scenarios with numerous people. The results demonstrate that LiDAR data provide valuable information for social grouping and recognizing individual action and group activities, particularly in crowded group settings. For social grouping, our MGAR-net improve performance by about 12% compared to the existing state-of-the-art models in terms of the AP metric. Sangmin Woo, Jinyoung Park 0001, Muhammad Adi Nugroho, Changick Kim |
VCIP | 5 |
| 2023 | AHFu-Net: Align, Hallucinate, and Fuse Network for Missing Multimodal Action RecognitionabstractIn this work, we explore the multimodal action recognition problem, specifically in the context of RGB-Depth modalities scenario, where a subset of the learning modalities is missing at inference time. To address this issue, we construct a hallucination network to generate missing modality information from the available modality at inference time. We propose key components of an effective spatio-temporal encoder for strong unimodal performance with Local Patch Temporal Transformer (LPTT) and Spatial Encoder Transformer (SET), alignment of multi-modal features, and fusion strategy with our Multimodal Bottleneck Transformer Fusion Module (MMBTF). We incorporate these ideas into a novel framework named AHFu-Net (Align, Hallucinate, and Fuse network) for RGB-Depth action recognition. Our experiments demonstrate that AHFu Net achieves state-of-the-art performance while maintaining high accuracy in the case of missing modality on multimodal datasets of NTU-RGB+D and NWUCLA. Muhammad Adi Nugroho, Sangmin Woo, Changick Kim |
VCIP | 1 |
| 2023 | Modality Mixer for Multi-modal Action RecognitionabstractIn multi-modal action recognition, it is important to consider not only the complementary nature of different modalities but also global action content. In this paper, we propose a novel network, named Modality Mixer (M-Mixer) network, to leverage complementary information across modalities and temporal context of an action for multi-modal action recognition. We also introduce a simple yet effective recurrent unit, called Multi-modal Contextualization Unit (MCU), which is a core component of M-Mixer. Our MCU temporally encodes a sequence of one modality (e.g., RGB) with action content features of other modalities (e.g., depth, IR). This process encourages M-Mixer to exploit global action content and also to supplement complementary information of other modalities. As a result, our proposed method outperforms state-of-the-art methods on NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA datasets. Moreover, we demonstrate the effectiveness of M-Mixer by conducting comprehensive ablation studies. Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim |
WACV | 4 |
| 2023 | Cross-modal alignment and translation for missing modality action recognition
Yeonju Park, Sangmin Woo, Muhammad Adi Nugroho, Changick Kim |
Comput. Vis. Image Underst. | 4 |