VLDB 2026 Research / reviewers in the wild / expert
Ryosuke Kawamura
dblp:175/9824
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Camera Self-Calibration in Sports Motion Capture: Leveraging Human and Stick Poses
Changsoo Jung, Ryosuke Kawamura, Hon Yung Wong |
FG | 3 |
| 2025 | Custom Condition Generation for Zero-Shot Human-Scene Interactions SynthesisabstractExisting methods for creating human interactions within scenes show promise for common interactions, but often fail with less frequent ones. To overcome this, we introduce a new approach that creates tailored conditions for generating these interactions without previously seen examples. This method leverages the strengths of both large language models (LLMs) and vision-language models (VLMs). Unlike the GenZI, the current state-of-the-art approach, which struggles with rare interactions due to its reliance on VLM inpainting, our method follows a three-step process: first, we generate a preliminary human posture using VLMs and then estimate this posture in three dimensions. Next, we refine the conditions to fit the specific scene and interaction by analyzing the inputs with both LLMs and VLMs. Finally, we fine-tune the placement, orientation, and posture of the human figure using specific optimization techniques. Our experimental results show that this method performs well across a wide range of interactions, including those that are less common. Ryosuke Kawamura, Zoltán Ádám Milacski, Fernando De la Torre, László A. Jeni, Koichiro Niinuma |
FG | 1 |
| 2025 | RN-Sam: Road Network-Aided Sam Optimization for Road Segmentation In Satellite ImageryabstractRoad segmentation in satellite imagery is critical for various applications, and the Segment Anything Model (SAM) has recently been applied to this task, as with other remote sensing applications. However, despite its advancements, applying SAM to road segmentation poses notable challenges. First, inaccurate prompts can degrade SAM’s performance. Second, existing methods often lack practical evaluation in cross-region scenario. To address these issues, this paper introduces RN-SAM, a novel framework that leverages OpenStreetMap (OSM) as a reliable source of auxiliary road network information to enhance SAM for road segmentation and incorporates a new dataset designed for robust cross-region evaluations. The proposed framework comprises two phases: fine-tuning SAM with road network-based prompts and applying test-time adaptation using OSM road network data. Experimental results on our dataset demonstrate that the proposed framework significantly enhances SAM’s performance for road segmentation and outperforms existing methods in both same-region and cross-region scenarios. Ryosuke Kawamura, Pablo Guarda, Pradeep Narwade, Koichiro Niinuma |
ICIP | 1 |
| 2025 | GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text ContextsabstractThe connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities makes capturing this connection challenging with a fixed set of descriptors. Specifically, closed vocabulary scene encoders, which require learning text-scene associations from scratch, have been favored in the literature, often resulting in inaccurate motion grounding. In this paper, we propose a method that integrates an open vocabulary scene encoder into the architecture, establishing a robust connection between text and scene. Our two-step approach starts with pretraining the scene encoder through knowledge distillation from an existing open vocabulary semantic image segmentation model, ensuring a shared text-scene feature space. Subsequently, the scene encoder is fine-tuned for conditional motion generation, incorporating two novel regularization losses that regress the category and size of the goal object. Our methodology achieves up to a 30% reduction in the goal object distance metric compared to the prior state-of-the-art baseline model on the HUMANISE dataset. This improvement is demonstrated through evaluations conducted using three implementations of our framework, a perceptual study, and an open vocabulary experiment. Additionally, our method is designed to accommodate future 2D open vocabulary segmentation methods for distillation in a plug-and-play manner. Zoltán Ádám Milacski, Koichiro Niinuma, Ryosuke Kawamura, Fernando De la Torre, László A. Jeni |
WACV | 3 |
| 2024 | Is Internal State Feedback in an E-Learning Environment Acceptable to People?abstractIn on-demand e-learning environments, the lack of direct intervention can lead to a decline in learners' engagement. To address this issue, systems that estimate the learners' attitudes and provide feedback have been proposed. However, the acceptability of such systems has not been sufficiently researched. In this study, we investigated the acceptability by people to an e-learning system with internal state feedback, for future personalized learning support. To this end, we developed a system that estimates and visualizes the learner's internal state in real-time. The system was exhibited in a public space for free use, and users' impressions were analyzed. To estimate the learners' internal state, we developed a machine-learning model that recognizes learners' alertness from facial videos. The system was deployed in an exhibition space, and 131 responses were collected. These responses were coded and analyzed using a co-occurrence network. The result indicated that learners tend to dislike the system due to feelings of being observed by supervisors. In contrast, instructors expressed favorable options toward the introduction of the system. Atsushi Ashida, Ryosuke Kawamura, Shizuka Shirai, Noriko Takemura, Mehrasa Alizadeh, Hideaki Hayashi, Hajime Nagahara |
ICCE | 2 |
| 2024 | Synthetic Video Generation for Weakly Supervised Cross-Domain Video Anomaly Detection
Pradeep Narwade, Ryosuke Kawamura, Gaurav Gajbhiye, Koichiro Niinuma |
ICPR (15) | 2 |
| 2024 | MIDAS: Mixing Ambiguous Data with Soft Labels for Dynamic Facial Expression RecognitionabstractDynamic facial expression recognition (DFER) is an important task in the field of computer vision. To apply automatic DFER in practice, it is necessary to accurately recognize ambiguous facial expressions, which often appear in data in the wild. In this paper, we propose MIDAS, a data augmentation method for DFER, which augments ambiguous facial expression data with soft labels consisting of probabilities for multiple emotion classes. In MIDAS, the training data are augmented by convexly combining pairs of video frames and their corresponding emotion class labels, which can also be regarded as an extension of mixup to soft-labeled video data. This simple extension is remarkably effective in DFER with ambiguous facial expression data. To evaluate MIDAS, we conducted experiments on the DFEW dataset. The results demonstrate that the model trained on the data augmented by MIDAS outperforms the existing state-of-the-art method trained on the original dataset. Ryosuke Kawamura, Hideaki Hayashi, Noriko Takemura, Hajime Nagahara |
WACV | 1 |
| 2022 | Uncertainty Prediction for Facial Action Units Recognition under Degraded ConditionsabstractFacial action units (AUs) represent muscular activities, and their recognition from facial images can capture various psychological states, such as people’s interests as consumers and mental health states. However, degradation of conditions, such as occlusions by hand, often occurs and affects the accuracy of AUs recognition in the real world. Most existing studies on degraded conditions have adopted the approach using additional training images and advanced structures of neural networks to improve the robustness of AUs recognition from a degraded facial image. However, such an approach cannot deal with cases in which evidence of the AUs is completely or almost invisible. Therefore, we propose a novel method to address the degraded conditions by predicting the uncertainties of the AUs recognition caused by them. Our method interpolates the high-uncertainty data using surrounding data to reduce the influence of the degraded conditions, and visualizes the conditions causing the uncertainties to handle cases where the conditions are very poor and need to be improved. In the evaluation experiments, the public datasets BP4D+ and DISFA were modified to degrade them for testing. By evaluating the modified test data, we demonstrated that the maximum improvement with our method was 12% for BP4D+ and 17% for DISFA, and that our method can prevent the decrease in accuracy owing to degraded conditions. We also presented some visualization examples which demonstrate that our method can reasonably predict the conditions and uncertainties. Junya Saito, Sachihiro Youoku, Ryosuke Kawamura, Akiyoshi Uchida, Kentaro Murase, Xiaoyu Mi |
ICMLA | 3 |
| 2021 | Facial Action Unit Detection Based on Teacher-Student Learning Framework for Partially Occluded Facial ImagesabstractFacial action unit (AU) detection is an important task in facial expression analysis. However, occlusion is a major hindrance in practical applications of AU detection as it interferes with extracting features from facial images and makes it difficult to capture the occurrence of AUs. To address this problem, we first construct a database for AU detection with occlusion by synthesizing occlusion objects such as bangs (i.e., fringe), glasses, and hands on facial images. Then we apply a teacher-student learning framework with two types of loss functions to AU detection for occluded facial images. To improve our model's robustness to occlusion, we propose a loss function for order regularization which considers the relationship between facial images as well as conventional distillation loss. The results of our experiment with our database for occluded images demonstrate that our method is effective for detecting AUs with occluded facial images. Ryosuke Kawamura, Kentaro Murase |
FG | 1 |