VLDB 2026 Research / reviewers in the wild / expert
Go Irie
dblp:98/7454
· DBLP profile ↗
63ranked-venue papers
13as first author
30since 2021 · last 2025
0000-0002-4309-4700ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 13 first-author · 22 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal ModelsabstractThis paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed Unsolvable Problem Detection (UPD). Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM’s ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. Atsuyuki Miyai, Jingyang Zhang, Yifei Ming, Qing Yu 0013, Go Irie, Yixuan Li 0001, Hai Li 0001, Ziwei Liu 0002, Kiyoharu Aizawa |
ACL (1) | 6 |
| 2025 | Multi-Task Learning for Ultrasonic Echo-based Depth Estimation with Audible Frequency RecoveryabstractWhile depth maps of indoor scenes are often essential for a variety of applications, measuring depth maps usually requires dedicated depth sensors, which are not always available. Echo-based depth estimation has been explored as a promising alternative solution. However, most existing methods assume the use of audible echoes, with the major problem that prevents their use in quiet spaces or in situations where the generation of audible sound is prohibited. In this paper, we explore depth estimation based on ultrasonic echoes, which has scarcely been explored so far. The key idea of our method is to learn a depth estimation model that can exploit useful, but missing information in the audible frequency band. To this end, we perform multi-task learning that requires estimation of depth maps from ultrasound echoes while simultaneously restoring the audible frequency range. Furthermore, to evaluate the performance with real echo data, we develop a data collection device and collect a real sound dataset. Experimental results on this real echo dataset and public simulation benchmark dataset demonstrate that our method outperforms existing methods. Our real echo dataset and the code will be publicly available if the paper is accepted. Junpei Honma, Akisato Kimura, Go Irie |
ICASSP | 3 |
| 2025 | A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language ModelsabstractOut-of-distribution (OOD) detection is a task that detects OOD samples during inference to ensure the safety of deployed models. However, conventional benchmarks have reached performance saturation, making it difficult to compare recent OOD detection methods. To address this challenge, we introduce three novel OOD detection benchmarks that enable a deeper understanding of method characteristics and reflect real-world conditions. First, we present ImageNet-X, designed to evaluate performance under challenging semantic shifts. Second, we propose ImageNet-FS-X for full-spectrum OOD detection, assessing robustness to covariate shifts (feature distribution shifts). Finally, we propose Wilds-FS-X, which extends these evaluations to real-world datasets, offering a more comprehensive testbed. Our experiments reveal that recent CLIP-based OOD detection methods struggle to varying degrees across the three proposed benchmarks, and none of them consistently outperforms the others. We hope the community goes beyond specific benchmarks and includes more challenging conditions reflecting real-world scenarios. The code is https://github.com/hoshi23/OOD-X-Benchmarks. Shiho Noda, Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
ICIP | 4 |
| 2025 | Approximate Domain Unlearning for Vision-Language ModelsabstractPre-trained Vision-Language Models (VLMs) exhibit strong generalization capabilities, enabling them to recognize a wide range of objects across diverse domains without additional training. However, they often retain irrelevant information beyond the requirements of specific target downstream tasks, raising concerns about computational efficiency and potential information leakage. This has motivated growing interest in approximate unlearning, which aims to selectively remove unnecessary knowledge while preserving overall model performance. Existing approaches to approximate unlearning have primarily focused on {\em class unlearning}, where a VLM is retrained to fail to recognize specified object classes while maintaining accuracy for others. However, merely forgetting object classes is often insufficient in practical applications. For instance, an autonomous driving system should accurately recognize {\em real} cars, while avoiding misrecognition of {\em illustrated} cars depicted in roadside advertisements as {\em real} cars, which could be hazardous. In this paper, we introduce {\em Approximate Domain Unlearning (ADU)}, a novel problem setting that requires reducing recognition accuracy for images from specified domains (e.g., {\em illustration}) while preserving accuracy for other domains (e.g., {\em real}). ADU presents new technical challenges: due to the strong domain generalization capability of pre-trained VLMs, domain distributions are highly entangled in the feature space, making naive approaches based on penalizing target domains ineffective. To tackle this limitation, we propose a novel approach that explicitly disentangles domain distributions and adaptively captures instance-specific domain information. Extensive experiments on four multi-domain benchmark datasets demonstrate that our approach significantly outperforms strong baselines built upon state-of-the-art VLM tuning techniques, paving the way for practical and fine-grained unlearning in VLMs. Code : https://kodaikawamura.github.io/Domain_Unlearning/. Kodai Kawamura, Yuta Goto, Rintaro Yanagi, Hirokatsu Kataoka, Go Irie |
NeurIPS | 5 |
| 2025 | Open-set domain adaptation with visual-language foundation models
Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
Comput. Vis. Image Underst. | 2 |
| 2025 | GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution DetectionabstractAbstract Zero-shot OOD detection is a task that detects OOD images during inference with only in-distribution (ID) class names. Existing methods assume ID images contain a single, centered object, and do not consider the more realistic multi-object scenarios, where both ID and OOD objects are present. To meet the needs of many users, the detection method must have the flexibility to adapt the type of ID images. To this end, we present Global-Local Maximum Concept Matching (GL-MCM), which incorporates local image scores as an auxiliary score to enhance the separability of global and local visual features. Due to the simple ensemble score function design, GL-MCM can control the type of ID images with a single weight parameter. Experiments on ImageNet and multi-object benchmarks demonstrate that GL-MCM outperforms baseline zero-shot methods and is comparable to fully supervised methods. Furthermore, GL-MCM offers strong flexibility in adjusting the target type of ID images. The code is available via https://github.com/AtsuMiyai/GL-MCM . Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
Int. J. Comput. Vis. | 3 |
| 2024 | Linear Calibration Approach to Knowledge-free Group Robust Classification
Ryota Ishizaki, Shunya Yamagami, Yuta Goto, Go Irie |
BMVC | 4 |
| 2024 | Region-based Entropy Separation for One-shot Test-Time Adaptation
Kodai Kawamura, Shunya Yamagami, Go Irie |
BMVC | 3 |
| 2024 | Acoustic-based 3D Human Pose Estimation Robust to Human Position
Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa |
BMVC | 3 |
| 2024 | Estimating Indoor Scene Depth Maps From Ultrasonic EchoesabstractMeasuring 3D geometric structures of indoor scenes requires dedicated depth sensors, which are not always available. Echo-based depth estimation has recently been studied as a promising alternative solution. All previous studies have assumed the use of echoes in the audible range. However, one major problem is that audible echoes cannot be used in quiet spaces or other situations where producing audible sounds is prohibited. In this paper, we consider echo-based depth estimation using inaudible ultrasonic echoes. While ultrasonic waves provide high measurement accuracy in theory, the actual depth estimation accuracy when ultrasonic echoes are used has remained unclear, due to its disadvantage of being sensitive to noise and susceptible to attenuation. We first investigate the depth estimation accuracy when the frequency of the sound source is restricted to the high-frequency band, and found that the accuracy decreased when the frequency was limited to ultrasonic ranges. Based on this observation, we propose a novel deep learning method to improve the accuracy of ultrasonic echo-based depth estimation by using audible echoes as auxiliary data only during training. Experimental results with a public dataset demonstrate that our method improves the estimation accuracy. Junpei Honma, Akisato Kimura, Go Irie |
ICIP | 3 |
| 2024 | Pose-Invariant Learning for Efficient Person Identification from Hyperspectral Hand ImagesabstractWhile person identification from multi/hyperspectral images of hands has advantages such as contactless and high flexibility in capturing images, it remains a difficult task because individual characteristics are not as clear as those of a face or fingerprints. The state-of-the-art method uses a 3D CNN classifier to capture detailed spectral information. However, this is computationally expensive and is negatively affected by undesired spectral variations caused by changes in hand pose. We propose a new method to address these problems. The key technical components of the proposed method are in the introduction of adversarial learning to learn pose-invariant features and in usage of the separable convolutions to decouple the operations in the channel and spatial directions to improve efficiency. Furthermore, these technical components are integrated into a unified supervised contrastive learning framework, which is suitable for person identification. Experimental results demonstrate that our method achieves higher accuracy than the existing method while significantly reducing computational complexity. Keigo Kunikata, Amane Kashino, Yota Yamamoto, Yukinobu Taniguchi, Yoko Sogabe, Ayumi Matsumoto, Masaki Kitahara, Go Irie |
ICIP | 8 |
| 2024 | Cross-Action Cross-Subject Skeleton Action Recognition Via Simultaneous Action-Subject Learning With Two-Step Feature RemovalabstractIn this paper, we tackle a novel skeleton-based action recognition problem named Cross-Action Cross-Subject (CACS) Skeleton Action Recognition, where we can access the data of only a part of the target action classes for each training subject. Existing skeleton-based action recognition methods suffer from solving this problem because there are scarce clues to resolve the cross-entanglement of action and subject information, and the trained model will confuse those two features. To solve this challenging problem, we propose a method that consists of simultaneous action-subject learning with feature removal. In our method, 1) we use two data augmentation techniques, Bone Randomization and Phase Randomization, to roughly remove unnecessary features for respective recognitions, and then, 2) we introduce a debiased learning approach to remove the confusing features by minimizing mutual information with an action-subject-shared discriminator network. Extensive experiments on three datasets demonstrate that our method is consistently effective for several CACS problems. Yu Mitsuzumi, Akisato Kimura, Go Irie, Atsushi Nakazawa |
ICIP | 3 |
| 2024 | Bivariate Mixup for 2D Contact Point Localization with Piezoelectric Microphone ArrayabstractAiming at capturing interactions between a human hand and a desk when seated at a desk, we address the task of localizing the position of a hand touching a desk from the contact sound.Our framework uses an array of contact microphones called piezoelectric devices mounted on the desk to regress the 2D coordinates of the hand position from the resulting contact sound.Training a qualified regression model requires a number of high-quality training samples, but collecting accurate ground truth coordinates is often costly.To address this problem, we propose a novel data augmentation technique customized to 2D regression problems named Bivariate Mixup (BiMixup).BiMixup is a generalization of Mixup, which is formulated as a univariate linear interpolation, to a bivariate quadratic interpolation.Experimental results using a real sound dataset show that our BiMixup reduces the error by 12.2 % compared to the original Mixup. Shogo Yonezawa, Yukinobu Taniguchi, Go Irie |
MMAsia | 3 |
| 2024 | Black-Box ForgettingabstractLarge-scale pre-trained models (PTMs) provide remarkable zero-shot classification capability covering a wide variety of object classes. However, practical applications do not always require the classification of all kinds of objects, and leaving the model capable of recognizing unnecessary classes not only degrades overall accuracy but also leads to operational disadvantages. To mitigate this issue, we explore the selective forgetting problem for PTMs, where the task is to make the model unable to recognize only the specified classes, while maintaining accuracy for the rest. All the existing methods assume ''white-box'' settings, where model information such as architectures, parameters, and gradients is available for training. However, PTMs are often ''black-box,'' where information on such models is unavailable for commercial reasons or social responsibilities. In this paper, we address a novel problem of selective forgetting for black-box models, named Black-Box Forgetting, and propose an approach to the problem. Given that information on the model is unavailable, we optimize the input prompt to decrease the accuracy of specified classes through derivative-free optimization. To avoid difficult high-dimensional optimization while ensuring high forgetting performance, we propose Latent Context Sharing, which introduces common low-dimensional latent components among multiple tokens for the prompt. Experiments on four standard benchmark datasets demonstrate the superiority of our method with reasonable baselines. The code is available at https://github.com/yusukekwn/Black-Box-Forgetting. Yusuke Kuwana, Yuta Goto, Takashi Shibata 0001, Go Irie |
NeurIPS | 4 |
| 2024 | Phase Randomization: A data augmentation for domain adaptation in human action recognition
Yu Mitsuzumi, Go Irie, Akisato Kimura, Atsushi Nakazawa |
Pattern Recognit. | 2 |
| 2024 | Self-Labeling Framework for Open-Set Domain Adaptation With Few Labeled SamplesabstractUnsupervised domain adaptation (UDA) is extremely effective for transferring knowledge from a label-rich source domain to a label-scarce target domain. Because the target domain is unlabeled and may contain additional novel classes, open-set domain adaptation (ODA) has been suggested as a possible solution to detect these novel classes in the training phase. However, existing ODA methods rely heavily on abundant fully labeled source data, which are expensive to collect in specific applications and may also contain novel classes. In this study, we propose a novel self-labeling framework with prototypical contrastive learning and mutual information maximization to achieve ODA even when the amount of labeled data is very small, which is a new problem setting named few-shot ODA (FODA). We use self-supervised prototypical contrastive learning to train the network to learn the representations of source and target samples and maximize the mutual information between labels and input data to simultaneously recognize known and novel classes in the source and target domains. We evaluated our strategy in several domain adaptation environments and found that our method performed far better than existing approaches. Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
IEEE Trans. Multim. | 2 |
| 2023 | Listening Human Behavior: 3D Human Pose Estimation with Acoustic SignalsabstractGiven only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actions, we explore how low-level acoustic signals can provide enough clues to estimate 3D human poses by active acoustic sensing with a single pair of microphones and loudspeakers (see Fig. 1). This is a challenging task since sound is much more diffractive than other signals and therefore covers up the shape of objects in a scene. Accordingly, we introduce a framework that encodes multichannel audio features into 3D human poses. Aiming to capture subtle sound changes to reveal detailed pose information, we explicitly extract phase features from the acoustic signals together with typical spectrum features and feed them into our human pose estimation network. Also, we show that reflected or diffracted sounds are easily influenced by subjects' physique differences e.g., height and muscularity, which deteriorates prediction accuracy. We reduce these gaps by using a subject discriminator to improve accuracy. Our experiments suggest that with the use of only low-dimensional acoustic information, our method outperforms baseline methods. The datasets and codes used in this project will be publicly available. Yuto Shibata, Yutaka Kawashima, Mariko Isogawa, Go Irie, Akisato Kimura, Yoshimitsu Aoki |
CVPR | 4 |
| 2023 | Noise-Avoidance Sampling for Annotation Missing Object DetectionabstractExcellent results can be achieved using object detection with fully supervised training on large well-annotated datasets. However, the problem of missing annotations in real-world datasets can considerably reduce the performance of object detectors. In this study, we thoroughly analyze the effect of missing annotations on both positive and negative samples in object detector training. To mitigate the negative impact caused by annotation missing problem, we propose a simple yet effective method, noise-avoidance sampling, to distinguish noisy training samples and subsequently reduce their negative impact. Experiments are conducted on the PASCAL VOC 07+12 dataset with varying levels of missing annotations. The results reveal that the proposed method achieves comparable or superior performance with state-of-the-art methods. Jiafeng Mao, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
ICIP | 3 |
| 2023 | Text-to-Image Fashion Retrieval with Fabric TexturesabstractIn this study, we proposed text-to-image fashion image retrieval that captures the texture of clothing fabrics. A fabric’s texture is a major factor governing the comfort and appearance of clothes and significantly influences user preferences. However, unlike patterns and shapes that can readily be captured from a global image of the entire piece of clothing, extracting the fine and ambiguous characteristics of textures is considerably more challenging. The key concept is that by focusing on the "local" regions of clothing, detailed fabric textures can be more accurately captured. To this end, we propose a framework for learning cross-modal features from both global (the entire garment) and local (a close-up detail) image-text pairs. To verify the idea, we constructed a new dataset named Global and Local FACAD (G&L FACAD) by modifying the existing large-scale public FACAD dataset used for fashion retrieval. The experimental results confirm that the retrieval accuracy is significantly improved compared to the baselines. The code is available at https://github.com/SuzukiDaichi-git/texture_aware_fashion_retrieval.git. Daichi Suzuki, Go Irie, Kiyoharu Aizawa |
ICMR | 2 |
| 2023 | LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt LearningabstractWe present a novel vision-language prompt learning approach for few-shot out-of-distribution (OOD) detection. Few-shot OOD detection aims to detect OOD images from classes that are unseen during training using only a few labeled in-distribution (ID) images. While prompt learning methods such as CoOp have shown effectiveness and efficiency in few-shot ID classification, they still face limitations in OOD detection due to the potential presence of ID-irrelevant information in text embeddings. To address this issue, we introduce a new approach called $\textbf{Lo}$cal regularized $\textbf{Co}$ntext $\textbf{Op}$timization (LoCoOp), which performs OOD regularization that utilizes the portions of CLIP local features as OOD features during training. CLIP's local features have a lot of ID-irrelevant nuisances ($\textit{e.g.}$, backgrounds), and by learning to push them away from the ID class text embeddings, we can remove the nuisances in the ID class text embeddings and enhance the separation between ID and OOD. Experiments on the large-scale ImageNet OOD detection benchmarks demonstrate the superiority of our LoCoOp over zero-shot, fully supervised detection methods and prompt learning methods. Notably, even in a one-shot setting -- just one label per class, LoCoOp outperforms existing zero-shot and fully supervised detection methods. The code is available via https://github.com/AtsuMiyai/LoCoOp. Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa |
NeurIPS | 3 |
| 2023 | Rethinking Rotation in Self-Supervised Contrastive Learning: Adaptive Positive or Negative Data AugmentationabstractRotation is frequently listed as a candidate for data augmentation in contrastive learning but seldom provides satisfactory improvements. We argue that this is because the rotated image is always treated as either positive or negative. The semantics of an image can be rotation-invariant or rotation-variant, so whether the rotated image is treated as positive or negative should be determined based on the content of the image. Therefore, we propose a novel augmentation strategy, adaptive Positive or Negative Data Augmentation (PNDA), in which an original and its rotated image are a positive pair if they are semantically close and a negative pair if they are semantically different. To achieve PNDA, we first determine whether rotation is positive or negative on an image-by-image basis in an unsupervised way. Then, we apply PNDA to contrastive learning frameworks. Our experiments showed that PNDA improves the performance of contrastive learning. The code is available at https://github.com/AtsuMiyai/rethinking_rotation. Atsuyuki Miyai, Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa |
WACV | 4 |
| 2022 | Self-Labeling Framework for Novel Category Discovery over DomainsabstractUnsupervised domain adaptation (UDA) has been highly successful in transferring knowledge acquired from a label-rich source domain to a label-scarce target domain. Open-set domain adaptation (open-set DA) and universal domain adaptation (UniDA) have been proposed as solutions to the problem concerning the presence of additional novel categories in the target domain. Existing open-set DA and UniDA approaches treat all novel categories as one unified unknown class and attempt to detect this unknown class during the training process. However, the features of the novel categories learned by these methods are not discriminative. This limits the applicability of UDA in the further classification of these novel categories into their original categories, rather than assigning them to a single unified class. In this paper, we propose a self-labeling framework to cluster all target samples, including those in the ''unknown'' categories. We train the network to learn the representations of target samples via self-supervised learning (SSL) and to identify the seen and unseen (novel) target-sample categories simultaneously by maximizing the mutual information between labels and input data. We evaluated our approach under different DA settings and concluded that our method generally outperformed existing ones by a wide margin. Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa |
AAAI | 3 |
| 2022 | Co-Attention-Guided Bilinear Model for Echo-Based Depth EstimationabstractEchoes reflect a geometric structure of a scene surrounding a sound source. In this paper, we address the problem of estimating depth maps of indoor scenes based on echoes. First, we experimentally show that fusing multiple acoustic features, especially spectrogram and angular spectrum, can improve estimation accuracy. We then propose a novel bilinear model that incorporates dense co-attention for effective feature fusion. Our model is able to obtain a compact fused feature while capturing the second-order correlations of intra-and inter-features. Thorough evaluations on two datasets demonstrate the superiority of the proposed method over the state-of-the-art echo-based depth estimation and feature fusion methods. Go Irie, Takashi Shibata 0001, Akisato Kimura |
ICASSP | 1 |
| 2021 | Generalized Domain AdaptationabstractMany variants of unsupervised domain adaptation (UDA) problems have been proposed and solved individually. Its side effect is that a method that works for one variant is often ineffective for or not even applicable to another, which has prevented practical applications. In this paper, we give a general representation of UDA problems, named Generalized Domain Adaptation (GDA). GDA covers the major variants as special cases, which allows us to organize them in a comprehensive framework. Moreover, this generalization leads to a new challenging setting where existing methods fail, such as when domain labels are unknown, and class labels are only partially given to each domain. We propose a novel approach to the new setting. The key to our approach is self-supervised class-destructive learning, which enables the learning of class-invariant representations and domain-adversarial classifiers without using any domain labels. Extensive experiments using three benchmark datasets demonstrate that our method outperforms the state-of-the-art UDA methods in the new setting and that it is competitive in existing UDA variations as well. Yu Mitsuzumi, Go Irie, Daiki Ikami, Takashi Shibata 0001 |
CVPR | 2 |
| 2021 | Disentangling Subject-Dependent/-Independent Representations for 2D Motion RetargetingabstractWe consider the problem of 2D motion retargeting, which is to transfer the motion of one 2D skeleton to another skeleton of a different body shape. Existing methods decompose the input motion skeleton into dynamic (motion) and static (body shape, viewpoint angle, and emotion) features and synthesize a new skeleton by mixing up the features extracted from the different data. However, the resulting motion skeletons do not reflect subject-dependent factors that can stylize motion, such as skill and expressions, leading to unattractive results. In this work, we propose a novel network to separate subject-dependent and -independent motion features and to reconstruct a new skeleton with or without subject-dependent motion features. The core of our approach is adversarial feature disentanglement. The motion features and a subject classifier are trained simultaneously such that subject-dependent motion features do allow for between-subject discrimination, whereas subject-independent features cannot. The presence or absence of individuality is readily controlled by a simple summation of the motion features. Our method shows superior performance to the state-of-the-art method in terms of reconstruction error and can generate new skeletons while maintaining individuality. Fanglu Xie, Go Irie, Tatsushi Matsubayashi |
ICASSP | 2 |
| 2021 | Deep Reinforcement Image Matching with Self-TerminationabstractDeep reinforcement learning-based image matching sequentially searches only the promising regions in the reference image that match the query, leading to a significantly small number of steps compared to traditional methods. Since existing methods do not have any function to judge whether the target region has been successfully identified or not, they continue to search until the preset maximum number of search steps is reached. In this paper, we propose a deep image matching network that can terminate the matching process by itself. Our network is designed to have a halting module that identifies whether the current reference region matches the query based on the image features and the search history. The entire network is effectively trained end-to-end in a framework of deep reinforcement learning that incorporates a new loss function to evaluate the accuracy of the termination decision. Experimental results demonstrate that our method can achieve highly competitive or better matching accuracy with fewer search steps than the existing methods. Onkar Krishna, Go Irie, Xiaomeng Wu, Akisato Kimura, Kunio Kashino |
ICIP | 2 |
| 2021 | Learning with Selective ForgettingabstractLifelong learning aims to train a highly expressive model for a new task while retaining all knowledge for previous tasks. However, many practical scenarios do not always require the system to remember all of the past knowledge. Instead, ethical considerations call for selective and proactive forgetting of undesirable knowledge in order to prevent privacy issues and data leakage. In this paper, we propose a new framework for lifelong learning, called Learning with Selective Forgetting, which is to update a model for the new task with forgetting only the selected classes of the previous tasks while maintaining the rest. The key is to introduce a class-specific synthetic signal called mnemonic code. The codes are "watermarked" on all the training samples of the corresponding classes when the model is updated for a new task. This enables us to forget arbitrary classes later by only using the mnemonic codes without using the original data. Experiments on common benchmark datasets demonstrate the remarkable superiority of the proposed method over several existing methods. Takashi Shibata 0001, Go Irie, Daiki Ikami, Yu Mitsuzumi |
IJCAI | 2 |
| 2021 | Constrained Weight Optimization for Learning without Activation Normalization
Daiki Ikami, Go Irie, Takashi Shibata 0001 |
WACV | 2 |
| 2021 | Computational attention model for children, adults and the elderly
Onkar Krishna, Kiyoharu Aizawa, Go Irie |
Multim. Tools Appl. | 3 |
| 2021 | Joint object recognition and pose estimation using multiple-anchor triplet learning of canonical plane
Shunsuke Yoneda, Kouki Ueno, Go Irie, Masashi Nishiyama, Yoshio Iwai |
Pattern Recognit. Lett. | 3 |
| 2020 | Cascaded Transposed Long-Range Convolutions for Monocular Depth Estimation
Go Irie, Daiki Ikami, Takahito Kawanishi, Kunio Kashino |
ACCV (3) | 1 |
| 2020 | Adaptive Spotting: Deep Reinforcement Object Search in 3D Point Clouds
Onkar Krishna, Go Irie, Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino |
ACCV (3) | 2 |
| 2020 | Multi-task Curriculum Framework for Open-Set Semi-supervised Learning
Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa |
ECCV (12) | 3 |
| 2020 | A Generative Self-Ensemble Approach To Simulated+Unsupervised LearningabstractIn this paper, we consider Simulated and Unsupervised (S+U) learning which is a problem of learning from labeled synthetic and unlabeled real images. After translating the synthetic images to real ones, existing S+U learning methods use only the labeled synthetic images for training a predictor (e.g., a regression function) and ignore the target real images, which may result in unsatisfactory prediction performance. Our approach utilizes both synthetic and real images to train the predictor. The main idea of ours is to involve a self-ensemble learning framework into S+U learning. More specifically, we require the prediction results for an unlabeled real image to be consistent between “teacher” and “student” predictors, even after some perturbations are added to the image. Furthermore, aiming at generating diverse perturbations along the underlying data manifold, we introduce one-to-many image translation between synthetic and real images. Evaluation experiments on an appearance-based gaze estimation task demonstrate that the proposed ideas can improve the prediction accuracy and our full method can outperform existing S+U learning methods. Yu Mitsuzumi, Go Irie, Akisato Kimura, Atsushi Nakazawa |
ICIP | 2 |
| 2020 | Translating Adult's Focus of Attention to Elderly'sabstractPredicting which part of a scene elderly people would pay attention to could be useful in assisting their daily activities, such as driving, walking, and searching. Many computational models for predicting focus of attention (FoA) have been developed. However, most of them focus on mimicking adult FoA and do not work well for predicting elderly's, due to age-related changes in human vision. Is it possible to leverage the prediction results made by an FoA model of general adults to accurately predict elderly's FoA, rather than training a new network from scratch? In this paper, we consider a novel problem of translating adult's FoA to elderly's and propose an approach based on deep image-to-image translation. Our model is trained by minimizing both Kullback-Leibler divergence and adversarial loss to approximate the joint probability distribution of adult and elderly FoA. Experiments on two datasets demonstrate that our model gives remarkable prediction accuracy. Onkar Krishna, Go Irie, Takahito Kawanishi, Kunio Kashino, Kiyoharu Aizawa |
ICPR | 2 |
| 2019 | Delving Deep into Least Square Regression Model for Subspace Clustering
Masataka Yamaguchi, Go Irie, Takahito Kawanishi, Kunio Kashino |
BMVC | 2 |
| 2019 | Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio SignalsabstractSounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of semantic segmentation with given multichannel audio signals. Our approach uses a convolutional neural network that is designed to directly output semantic segmentation results by taking audio features as its inputs. A bilinear feature fusion scheme is incorporated that efficiently models underlying higher-order interactions between audio and visual sources. Experimental evaluations with both synthetic and real sound datasets show that our approach can recover the desired segmented images reasonably well. Go Irie, Mirela Ostrek, Hirokazu Kameoka, Akisato Kimura, Takahito Kawanishi, Kunio Kashino |
ICASSP | 1 |
| 2019 | Learning Search Path for Region-level Image MatchingabstractFinding a region of an image which matches to a query from a large number of candidates is a fundamental problem in image processing. The exhaustive nature of the sliding window approach has encouraged works that can reduce the run time by skipping unnecessary windows or pixels that do not play a substantial role in search results. However, such a pruning-based approach still needs to evaluate the non-ignorable number of candidates, which leads to a limited efficiency improvement. We propose an approach to learn efficient search paths from data. Our model is based on a CNN-LSTM architecture which is designed to sequentially determine a prospective location to be searched next based on the history of the locations attended. We propose a reinforcement learning algorithm to train the model in an end-to-end manner, which allows to jointly learn the search paths and deep image features for matching. These properties together significantly reduce the number of windows to be evaluated and makes it robust to background clutters. Our model gives remarkable matching accuracy with the reduced number of windows and run time on MNIST and FlickrLogos-32 datasets. Onkar Krishna, Go Irie, Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino |
ICASSP | 2 |
| 2019 | Subspace Structure-Aware Spectral Clustering for Robust Subspace ClusteringabstractSubspace clustering is the problem of partitioning data drawn from a union of multiple subspaces. The most popular subspace clustering framework in recent years is the graph clustering-based approach, which performs subspace clustering in two steps: graph construction and graph clustering. Although both steps are equally important for accurate clustering, the vast majority of work has focused on improving the graph construction step rather than the graph clustering step. In this paper, we propose a novel graph clustering framework for robust subspace clustering. By incorporating a geometry-aware term with the spectral clustering objective, we encourage our framework to be robust to noise and outliers in given affinity matrices. We also develop an efficient expectation-maximization-based algorithm for optimization. Through extensive experiments on four real-world datasets, we demonstrate that the proposed method outperforms existing methods. Masataka Yamaguchi, Go Irie, Takahito Kawanishi, Kunio Kashino |
ICCV | 2 |
| 2019 | Robust Learning for Deep Monocular Depth EstimationabstractExisting methods for deep monocular depth estimation are often trained with basic loss functions such as mean absolute error (MAE) or reverse Huber (BerHu). We revisit several basic loss functions to explore possibilities for improvement and show that the final depth estimation accuracy is dominated by pixels with small errors, which account for the vast majority. Based on this observation, we propose a new robust loss function that suppresses the contributions of pixels with higher errors by taking their square root. The loss function called the square root Huber (Ruber) is designed to be first-order differentiable on ℝ>0, so it can be directly applied to the end-to-end learning of a general type of neural network. Unlike the widely-used robust loss function called Huber, the Ruber loss function facilitates further refinement of the pixels with smaller errors by giving larger gradient values to the errors that are close to zero. Moreover, we show that the estimation accuracy can be further improved by introducing a second training step based on an edge-preserving loss. Experimental results with public indoor scene datasets demonstrate that our method outperforms major loss functions and can yield better accuracies than existing approaches in terms of root mean square error (RMSE). Go Irie, Takahito Kawanishi, Kunio Kashino |
ICIP | 1 |
| 2019 | Weakly Supervised Triplet Learning of Canonical Plane Transformation for Joint Object Recognition and Pose EstimationabstractWe propose a method for jointly performing object recognition and pose estimation using training samples of canonical planes to reduce the effort of data label supervision. Collecting a sufficient number of training samples is important to realizing high performance. However, labeling pose parameters is time consuming. We thus train our network model using only object class labels without explicitly labeling pose parameters. To recognize objects and estimate their poses, we design a network with a spatial transformer in a contrastive learning manner such that the canonical plane of an object is always transformed to a certain pose and the features are consistent with those of the object class. Experiments show that our method has improved accuracy in object recognition and lower error in pose estimation compared with simply using triplet learning or a spatial transformer network on a publicly available dataset. Kouki Ueno, Go Irie, Masashi Nishiyama, Yoshio Iwai |
ICIP | 2 |
| 2018 | Query Expansion with Diffusion On Mutual Rank GraphsabstractIn query expansion for object retrieval, there is substantial danger of query drift, where irrelevant information is inferred from pseudo-relevant images to enrich the query. To address this issue, we propose a query expansion method from the viewpoint of diffusion. It explores the structure of highly ranked images in a topological space, assuming that false positives reside on different manifolds from the query. For this purpose, a mutual rank graph is defined on pseudo-relevant images, and their distribution is learned by diffusing their query similarities through the graph. The relevance of a database image can thus be obtained by marginalizing over the learned distribution. The mutual rank graph accounts for varying local density in the image space, leading to great robustness as regards query drift and high generalization ability. The proposed method experimentally shows a consistent boost in the performance of object retrieval with handcrafted features on standard benchmarks. Xiaomeng Wu, Go Irie, Kaoru Hiramatsu, Kunio Kashino |
ICASSP | 2 |
| 2018 | Weighted Generalized Mean Pooling for Deep Image RetrievalabstractSpatial pooling over convolutional activations (e.g., max pooling or sum pooling) has been shown to be successful in learning deep representations for image retrieval. However, most pooling techniques assume that every activation is equally important, and as a result they suffer from the presence of uninformative image regions that play a negative role as regards matching or lead to the confusion of particular visual instances. To address this issue, we propose a trainable building block that steers pooling to local information important to the task at hand. The method formulates pooling as a weighted generalized mean (wGeM), in which weights are learned on activations, reflecting the discriminative power of each activation in image matching. Embedding wGeM in a deep network improves image representation and boosts retrieval performance on standard benchmarks. wGeM does not require any bounding box annotations, but instead learns the latent probabilities of activations from scratch. It even goes beyond objectness, and learns to look at important visual details rather than the whole region of the object of interest. Xiaomeng Wu, Go Irie, Kaoru Hiramatsu, Kunio Kashino |
ICIP | 2 |
| 2017 | Cross-modal transfer with neural word vectors for image feature learningabstractNeural word vector (NWV) such as word2vec is a powerful text representation tool that can encode extensive semantic information into compact vectors. This ability poses an interesting question in relation to image processing research - Can we learn better semantic image features from NWVs? We empirically explore this question in the context of semantic content-based image retrieval (CBIR). In this paper, we consider cross-modal transfer learning (CMT) to improve initial convolutional neural network (CNN) image features by using NWVs. We first show that NWVs can improve semantic CBIR performance compared to classical word vectors, even if it is with simple CMT models, i.e., canonical correlation analysis (CCA). Next, inspired by a characteristic property of NWVs, we propose a new CMT model and demonstrate that it can improve CBIR performance even further. Go Irie, Taichi Asami, Shuhei Tarashima, Takayuki Kurozumi, Tetsuya Kinebuchi |
ICASSP | 1 |
| 2016 | Joint object discovery and segmentation with image-wise reconstruction errorabstractWe tackle the problem of joint discovery and segmentation of the object of interest from noisy image sets collected via web crawling (e.g., Figure 1). Existing methods [1] [2] [3] employ region-wise comparison in order to separate noise images (images not containing target objects) from the rest, which may be a bottleneck for scaling up to larger datasets. Our idea to avoid such computationally intensive operations is to use image-wise reconstruction errors. Specifically, based on the assumption that images containing target objects are easier to be reconstructed by a pool of target objects than noise images, we first reconstruct each image using a small number of similar target objects. The resulting error is then combined with some other criteria (e.g., saliency) so as to delineate only target object regions. Experimental evaluations on a noisy image dataset [1] demonstrate that our approach achieves state-of-the-art results on every subset of the dataset 5-7 times faster than existing methods. Shuhei Tarashima, Jingjing Pan, Go Irie, Takayuki Kurozumi, Tetsuya Kinebuchi |
ICIP | 3 |
| 2016 | Attribute Discovery for Person Re-Identification
Takayuki Umeda, Yongqing Sun, Go Irie, Kyoko Sudo, Tetsuya Kinebuchi |
MMM (2) | 3 |
| 2015 | Alternating Co-Quantization for Cross-Modal HashingabstractThis paper addresses the problem of unsupervised learning of binary hash codes for efficient cross-modal retrieval. Many unimodal hashing studies have proven that both similarity preservation of data and maintenance of quantization quality are essential for improving retrieval performance with binary hash codes. However, most existing cross-modal hashing methods mainly have focused on the former, and the latter still remains almost untouched. We propose a method to minimize the binary quantization errors, which is tailored to cross-modal hashing. Our approach, named Alternating Co-Quantization (ACQ), alternately seeks binary quantizers for each modality space with the help of connections to other modality data so that they give minimal quantization errors while preserving data similarities. ACQ can be coupled with various existing cross-modal dimension reduction methods such as Canonical Correlation Analysis (CCA) and substantially boosts their retrieval performance in the Hamming space. Extensive experiments demonstrate that ACQ can outperform several state-of-the-art methods, even when it is combined with simple CCA. Go Irie, Hiroyuki Arai, Yukinobu Taniguchi |
ICCV | 1 |
| 2015 | Guest Editorial: Challenges and Perspectives for Affective Analysis in MultimediaabstractThe articles in this special section focus on new areas of development in the multimedia industry. Mohammad Soleymani 0001, Yi-Hsuan Yang, Go Irie, Alan Hanjalic |
IEEE Trans. Affect. Comput. | 3 |
| 2014 | Locally Linear Hashing for Extracting Non-linear ManifoldsabstractPrevious efforts in hashing intend to preserve data variance or pairwise affinity, but neither is adequate in capturing the manifold structures hidden in most visual data. In this paper, we tackle this problem by reconstructing the locally linear structures of manifolds in the binary Hamming space, which can be learned by locality-sensitive sparse coding. We cast the problem as a joint minimization of reconstruction error and quantization loss, and show that, despite its NP-hardness, a local optimum can be obtained efficiently via alternative optimization. Our method distinguishes itself from existing methods in its remarkable ability to extract the nearest neighbors of the query from the same manifold, instead of from the ambient space. On extensive experiments on various image benchmarks, our results improve previous state-of-the-art by 28-74% typically, and 627% on the Yale face data. Go Irie, Zhenguo Li, Xiao-Ming Wu 0003, Shih-Fu Chang |
CVPR | 1 |
| 2014 | Efficient Label PropagationabstractLabel propagation is a popular graph-based semi-supervised learning framework. So as to obtain the optimal labeling scores, the label propagation algorithm requires an inverse matrix which incurs the high computational cost of O(n^3+cn^2), where n and c are the numbers of data points and labels, respectively. This paper proposes an efficient label propagation algorithm that guarantees exactly the same labeling results as those yielded by optimal labeling scores. The key to our approach is to iteratively compute lower and upper bounds of labeling scores to prune unnecessary score computations. This idea significantly reduces the computational cost to O(cnt) where t is the average number of iterations for each label and t << n in practice. Experiments demonstrate the significant superiority of our algorithm over existing label propagation methods. Yasuhiro Fujiwara, Go Irie |
ICML | 2 |
| 2014 | Scaling Manifold Ranking Based Image RetrievalabstractManifold Ranking is a graph-based ranking algorithm being successfully applied to retrieve images from multimedia databases. Given a query image, Manifold Ranking computes the ranking scores of images in the database by exploiting the relationships among them expressed in the form of a graph. Since Manifold Ranking effectively utilizes the global structure of the graph, it is significantly better at finding intuitive results compared with current approaches. Fundamentally, Manifold Ranking requires an inverse matrix to compute ranking scores and so needs O ( n 3 ) time, where n is the number of images. Manifold Ranking, unfortunately, does not scale to support databases with large numbers of images. Our solution, Mogul , is based on two ideas: (1) It efficiently computes ranking scores by sparse matrices, and (2) It skips unnecessary score computations by estimating upper bounding scores. These two ideas reduce the time complexity of Mogul to O ( n ) from O ( n 3 ) of the inverse matrix approach. Experiments show that Mogul is much faster and gives significantly better retrieval quality than a state-of-the-art approximation approach. Yasuhiro Fujiwara, Go Irie, Shari Kuroyama, Makoto Onizuka |
Proc. VLDB Endow. | 2 |
| 2013 | A Bayesian Approach to Multimodal Visual Dictionary LearningabstractDespite significant progress, most existing visual dictionary learning methods rely on image descriptors alone or together with class labels. However, Web images are often associated with text data which may carry substantial information regarding image semantics, and may be exploited for visual dictionary learning. This paper explores this idea by leveraging relational information between image descriptors and textual words via co-clustering, in addition to information of image descriptors. Existing co-clustering methods are not optimal for this problem because they ignore the structure of image descriptors in the continuous space, which is crucial for capturing visual characteristics of images. We propose a novel Bayesian co-clustering model to jointly estimate the underlying distributions of the continuous image descriptors as well as the relationship between such distributions and the textual words through a unified Bayesian inference. Extensive experiments on image categorization and retrieval have validated the substantial value of the proposed joint modeling in improving visual dictionary learning, where our model shows superior performance over several recent methods. Go Irie, Dong Liu 0001, Zhenguo Li, Shih-Fu Chang |
CVPR | 1 |
| 2013 | Fast image/video collection summarization with local clusteringabstractImage/video collection summarization is an emerging paradigm to provide an overview of contents stored in massive databases. Existing algorithms require at least O(N) time to generate a summary, which cannot be applied to online scenarios. Assuming that contents are represented as a sparse graph, we propose a fast image/video collection summarization algorithm using local graph clustering. After a query node is specified, our algorithm first finds a small sub-graph near the query without looking at the whole graph, and then selects fewer number of nodes diverse to each other. Our algorithm thus provides a summary in nearly constant time in the number of contents. Experimental results demonstrate that our algorithm is more than 1500 times faster than a state-of-the-art method, with comparable summarization quality. Shuhei Tarashima, Go Irie, Ken Tsutsuguchi, Hiroyuki Arai, Yukinobu Taniguchi |
ACM Multimedia | 2 |
| 2013 | Travel route recommendation using geotagged photos
Takeshi Kurashima, Tomoharu Iwata, Go Irie, Ko Fujimura |
Knowl. Inf. Syst. | 3 |
| 2012 | Improving Item Recommendation Based on Social Tag Ranking
Taiga Yoshida, Go Irie, Takashi Satou, Akira Kojima, Suguru Higashino |
MMM | 2 |
| 2011 | Fast Algorithm for Affinity Propagation
Yasuhiro Fujiwara, Go Irie, Tomoe Kitahara |
IJCAI | 2 |
| 2010 | Travel route recommendation using geotags in photo sharing sitesabstractThe ability to create geotagged photos enables people to share their personal experiences as tourists at specific locations and times. Assuming that the collection of each photographer's geotagged photos is a sequence of visited locations, photo-sharing sites are important sources for gathering the location histories of tourists. By following their location sequences, we can find representative and diverse travel routes that link key landmarks. In this paper, we propose a travel route recommendation method that makes use of the photographers' histories as held by Flickr. Recommendations are performed by our photographer behavior model, which estimates the probability of a photographer visiting a landmark. We incorporate user preference and present location information into the probabilistic behavior model by combining topic models and Markov models. We demonstrate the effectiveness of the proposed method using a real-life dataset holding information from 71,718 photographers taken in the United States in terms of the prediction accuracy of travel behavior. Takeshi Kurashima, Tomoharu Iwata, Go Irie, Ko Fujimura |
CIKM | 3 |
| 2010 | Automatic trailer generationabstractThis paper presents a content-based movie trailer generation method, named Vid2Trailer (V2T). Since trailers are intended to advertise movies, they must show specific symbols such as the title logo and the main theme music. Moreover, it is expected to attract viewers by its visual and audio content. V2T satisfies these two requirements when creating a trailer from the original movie content. First, the title logo and the main theme music are extracted. Second, impressive speech and video segments are extracted by using an affective content analysis technique. Third, all of the extracted components are concatenated into the form of a trailer; to realize this, we propose a method that estimates the affective impact of shot sequences, and introduce an algorithm that arranges a set of shots so as to maximize the affective impact of the sequence. Experiments show that our V2T is more appropriate to trailer generation than conventional techniques. Go Irie, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 1 |
| 2010 | Facial Parameters and Their Influence on Subjective Impression in the Context of Keyframe Extraction from Home Video Contents
Uwe Kowalik, Go Irie, Yasuhiko Miyazaki, Akira Kojima |
MMM | 2 |
| 2010 | Affective Audio-Visual Words and Latent Topic Driving Model for Realizing Movie Affective Scene ClassificationabstractThis paper presents a novel method for movie affective scene classification that outputs the emotion (in the form of labels) that the scene is likely to arouse in viewers. Since the affective preferences of users play an important role in movie selection, affective scene classification has the potential to develop more attractive user-centric movie search and browsing applications. Two main issues in designing movie affective scene classification are considered. One is “how to extract features that are strongly related to the viewer's emotions”, and the other is “how to map the extracted features to the emotion categories”. For the former, we propose a method to extract emotion-category-specific audio-visual features named affective audio-visual words (AAVWs). For the latter issue, we propose a classification model named latent topic driving model (LTDM). Assuming that viewers' emotions are dynamically changed by the movie scene sequences, LTDM models emotions as Markovian dynamic systems driven by the sequential stimuli of the movie content. Experiments on 206 movie scenes extracted from 24 movie titles and the corresponding labels of eight emotion categories given by 16 subjects show that our method outperforms conventional approaches in terms of the subject agreement rate. Go Irie, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa |
IEEE Trans. Multim. | 1 |
| 2009 | Affective video segment retrieval for consumer generated videos based on correlation between emotions and emotional audio eventsabstractA novel affective video segment retrieval method based on the correlation between emotion and emotional audio events (EAEs) is presented. The proposed method focuses on retrieving three types of affective video segments, joy, sadness and excitement, by utilizing correlations between emotions and EAEs. The correlation between these emotions and EAEs is investigated by a subjective evaluation. The proposed method detects EAEs and rates each EAE in terms of emotion levels. The EAEs are detected by using the generalized state-space model (GSSM) and low-level audio features. Experiments conducted on consumer generated videos (CGVs) show that the proposed EAE detection outperforms conventional HMM and GMM based methods in terms of accuracy, the agreement rate of the retrieved affective video segments reaches 73.3%. Go Irie, Kota Hidaka, Takashi Satou, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 1 |
| 2009 | A degree-of-edit ranking for consumer generated video retrievalabstractWe introduce degree-of-edit (DoE) ranking to focus on ldquohow much a CGV is editedrdquo as a ranking measure for consumer generated video (CGV) retrieval; a method to estimate DoE ranking is proposed. In the proposed method, the DoE score of a CGV is estimated by using low-level features such as the number of shot boundaries and time ratio of music. We evaluate the rank correlation between DoE ranking determined by subjects and by our method. To demonstrate its performance in a practical scenario, a user test is performed on over 22,000 CGVs in the context of CGV search. The obtained results show that our method significantly improves conventional CGV ranking results in terms of availabilities of interesting and high-quality CGVs. Go Irie, Kota Hidaka, Takashi Satou, Toshihiko Yamasaki, Kiyoharu Aizawa |
ICME | 1 |
| 2009 | Latent topic driving model for movie affective scene classificationabstractThis paper proposes a latent topic driving model (LTDM) as a novel approach to movie affective scene classification. LTDM is a discriminative model of emotions driven by movie affective contents. Unlike existing methods, our approach is based on movie topic extraction via the latent Dirichlet allocation (LDA) and emotion dynamics modeling with reference to Plutchik's emotion theory. The classification procedure starts by segmenting movie scenes into movie shots, each of which is represented by a histogram of quantized affect-related audio-visual features. LDA is applied to detect topics of each movie shot. Emotions for the current movie shot are estimated based on both the topics of the shot and emotion transition weights determined by Plutchik's emotion theory. We conduct experiments using 206 movie scenes extracted from 24 movie titles (total 6 hours 20 min. 12 sec.) and the labels of eight emotion categories given by 16 subjects are collected. The results show that LTDM outperforms conventional modeling approaches in terms of the subject agreement rate. Go Irie, Kota Hidaka, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 1 |